EDBT 2026 Demo / reviewers in the wild / expert
Anna Hilsmann
dblp:57/1174
· DBLP profile ↗
37ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0002-2086-0951ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Security and privacy · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RIPE: Reinforcement Learning on Unlabeled Image Pairs for Robust Keypoint ExtractionabstractWe introduce RIPE, an innovative reinforcement learning-based framework for weakly-supervised training of a keypoint extractor that excels in both detection and description tasks. In contrast to conventional training regimes that depend heavily on artificial transformations, pre-generated models, or 3D data, RIPE requires only a binary label indicating whether paired images represent the same scene. This minimal supervision significantly expands the pool of training data, enabling the creation of a highly generalized and robust keypoint extractor. RIPE utilizes the encoder's intermediate layers for the description of the keypoints with a hyper-column approach to integrate information from different scales. Additionally, we propose an auxiliary loss to enhance the discriminative capability of the learned descriptors. Comprehensive evaluations on standard benchmarks demonstrate that RIPE simplifies data preparation while achieving competitive performance compared to state-of-the-art techniques, marking a significant advancement in robust keypoint extraction and description. To support further research, we have made our code publicly available at https://github.com/fraunhoferhhi/RIPE. Johannes Künzel, Anna Hilsmann, Peter Eisert |
ICCV | 2 |
| 2025 | Batch-Aware Active Learning for Object DetectionabstractWe propose a Batch-Aware Active Learning (BAAL) framework to optimize the training of object detection models, reducing annotation costs while maintaining strong model performance. The framework adapts different uncertainty sampling strategies to the specific challenges of object detection, including multi-class labelling and spatial localization. By combining uncertainty with diversity, leveraging feature representations and clustering, our method ensures diverse and informative batch selection. The non-invasive, plug-and-play design supports seamless integration with any object detection model without architectural modifications. Evaluations on COCO and Pascal VOC datasets with SSD, Faster R-CNN, YOLOv8, and RetinaNet demonstrate that our approach is not only efficient and robust but also comparable to, and in some cases exceeds, current state-of-the-art solutions. Mykyta Kovalenko, Peter Eisert, Anna Hilsmann, Sebastian Bosse |
ICIP | 3 |
| 2025 | CGS-GAN: 3D Consistent Gaussian Splatting GANs for High Resolution Human Head SynthesisabstractRecently, 3D GANs based on 3D Gaussian splatting have been proposed for high quality synthesis of human heads. However, existing methods stabilize training and enhance rendering quality from steep viewpoints by conditioning the random latent vector on the current camera position. This compromises 3D consistency, as we observe significant identity changes when re-synthesizing the 3D head with each camera shift. Conversely, fixing the camera to a single viewpoint yields high-quality renderings for that perspective but results in poor performance for novel views. Removing view-conditioning typically destabilizes GAN training, often causing the training to collapse. In response to these challenges, we introduce CGS-GAN, a novel 3D Gaussian Splatting GAN framework that enables stable training and high-quality 3D-consistent synthesis of human heads without relying on view-conditioning. To ensure training stability, we introduce a multi-view regularization technique that enhances generator convergence with minimal computational overhead. Additionally, we adapt the conditional loss used in existing 3D Gaussian splatting GANs and propose a generator architecture designed to not only stabilize training but also facilitate efficient rendering and straightforward scaling, enabling output resolutions up to $2048^2$. To evaluate the capabilities of CGS-GAN, we curate a new dataset derived from FFHQ. This dataset enables very high resolutions, focuses on larger portions of the human head, reduces view-dependent artifacts for improved 3D consistency, and excludes images where subjects are obscured by hands or other objects. As a result, our approach achieves very high rendering quality, supported by competitive FID scores, while ensuring consistent 3D scene generation. Florian Barthel, Wieland Morgenstern, Paul Hinzer, Anna Hilsmann, Peter Eisert |
NeurIPS | 4 |
| 2025 | AI-based Denoising and Interpolation of Magnetic UXO DataabstractMagnetic surveys are a key tool in detecting buried objects such as unexploded ordnance (UXO), where dense magnetic maps must be reconstructed from sparsely sampled gradiometer data. We present a deep learning-based approach that outperforms classical interpolation methods in both accuracy and speed. Trained on synthetic magnetic fields simulating realistic UXO signatures and measurement noise, our modified U-Net with ResNet-34 encoding reconstructs high-resolution magnetic maps from sparse inputs. Compared to state-of-the-art gridding methods, our model achieves 3–5% higher reconstruction accuracy on average while operating up to 80× faster than SOTA algorithms, enabling more efficient and interpretable UXO detection in real-world survey conditions. Mykyta Kovalenko, David Przewozny, Paul Chojecki, Anna Hilsmann, Peter Eisert, Sebastian Bosse |
SMC | 4 |
| 2025 | Adaptive and Temporally Consistent Gaussian Surfels for Multi-View Dynamic Reconstructionabstract3D Gaussian Splatting has recently achieved notable success in novel view synthesis for dynamic scenes and ge-ometry reconstruction in static scenes. Building on these advancements, early methods have been developed for dy-namic surface reconstruction by globally optimizing entire sequences. However, reconstructing dynamic scenes with significant topology changes, emerging or disappearing ob-jects, and rapid movements remains a substantial chal-lenge, particularly for long sequences. To address these issues, we propose AT-GS, a novel method for reconstructing high-quality dynamic surfaces from multi-view videos through per-frame incremental optimization. To avoid local minima across frames, we introduce a unified and adaptive gradient-aware densification strategy that integrates the strengths of conventional cloning and splitting techniques. Additionally, we reduce temporal jittering in dy-namic surfaces by ensuring consistency in curvature maps across consecutive frames. Our method achieves superior accuracy and temporal coherence in dynamic surface re-construction, delivering high-fidelity space-time novel view synthesis, even in complex and challenging scenes. Extensive experiments on diverse multi-view video datasets demonstrate the effectiveness of our approach, showing clear advantages over baseline methods. Project page: https://fraunhoferhhi.github.io/AT-GS Decai Chen, Brianne Oberson, Ingo Feldmann, Oliver Schreer, Anna Hilsmann, Peter Eisert |
WACV | 5 |
| 2025 | 3DGS.zip: A survey on 3D Gaussian Splatting Compression MethodsabstractAbstract 3D Gaussian Splatting (3DGS) has emerged as a cutting‐edge technique for real‐time radiance field rendering, offering state‐of‐the‐art performance in terms of both quality and speed. 3DGS models a scene as a collection of three‐dimensional Gaussians, with additional attributes optimized to conform to the scene's geometric and visual properties. Despite its advantages in rendering speed and image fidelity, 3DGS is limited by its significant storage and memory demands. These high demands make 3DGS impractical for mobile devices or headsets, reducing its applicability in important areas of computer graphics. To address these challenges and advance the practicality of 3DGS, this state‐of‐the‐art report (STAR) provides a comprehensive and detailed examination of two complementary yet fundamentally distinct strategies: compression and compaction. Compression techniques focus on reducing the file size by encoding Gaussian attributes more efficiently. In contrast, compaction methods directly optimize the scene's structure by optimizing the number of Gaussian primitives. Notably, while methods in both categories aim to maintain or improve quality, each while minimizing its respective attributes—file size for compression and the number of Gaussians for compaction—compaction does not necessarily lead to smaller file sizes; it specifically targets improved efficiency during rendering, making it distinct from compression. We introduce the basic mathematical concepts underlying the analyzed methods, as well as key implementation details and design choices. Our report thoroughly discusses similarities and differences among the methods, as well as their respective advantages and disadvantages. We establish a consistent framework for comparing the surveyed methods based on key performance metrics and datasets. Specifically, since these methods have been developed in parallel and over a short period of time, currently, no comprehensive comparison exists. This survey, for the first time, presents a unified framework to evaluate 3DGS compression techniques. To facilitate the continuous monitoring of emerging methodologies, we maintain a dedicated website that will be regularly updated with new techniques and revisions of existing findings. Overall, this STAR provides an intuitive starting point for researchers interested in exploring the rapidly growing field of 3DGS compression. By comprehensively categorizing and evaluating existing compression and compaction strategies, our work advances the understanding and practical application of 3DGS in computationally constrained environments. Milena T. Bagdasarian, Paul Knoll, Yi-Hsin Li, Florian Barthel, Anna Hilsmann, Peter Eisert, Wieland Morgenstern |
Comput. Graph. Forum | 5 |
| 2025 | Real-time fusion of stereo vision and hyperspectral imaging for objective decision support during surgeryabstractWe present a real-time stereo hyperspectral imaging (stereo-HSI) system for intraoperative tissue and organ analysis that integrates multispectral snapshot imaging with stereo vision to support clinical decision-making. The system visualize both RGB and high-dimensional spectral data while simultaneously reconstructing 3D surfaces, offering a compact, non-contact solution for seamless integration into surgical workflows. A modular processing pipeline enables robust demosaicing, spectral and spatial fusion, and pixel-wise medical assessment, including perfusion and tissue classification. Our spectral warping algorithm leverages a custom learned mapping, our white-balance network method is the first for snapshot MSI cameras, and our fusion CNN employs spectral-attention modules to exploit the rich hyperspectral domain. Clinical feasibility was demonstrated in 57 surgical procedures, including kidney transplantation, parotidectomy, and neck dissection, achieving high spatial and spectral resolution under standard surgical lighting conditions. The system enables visualization of oxygenation and tissue composition in real-time, offering surgeons a novel tool for image-guided interventions. This study establishes the stereo-HSI platform as a clinically viable and effective method for enhancing intraoperative insight and surgical precision. Eric L. Wisotzky, Jost Triller, Michael Knoke, Brigitta Globke, Anna Hilsmann, Peter Eisert |
Comput. Vis. Image Underst. | 5 |
| 2024 | SPVLoc: Semantic Panoramic Viewport Matching for 6D Camera Localization in Unseen Environments
Niklas Gard, Anna Hilsmann, Peter Eisert |
ECCV (73) | 2 |
| 2024 | Compact 3D Scene Representation via Self-Organizing Gaussian Grids
Wieland Morgenstern, Florian Barthel, Anna Hilsmann, Peter Eisert |
ECCV (85) | 3 |
| 2024 | Multi-View Gesture Recognition in Conflict SituationsabstractNon-verbal cues play a crucial role in social interactions and can influence conflict dynamics. For law enforcement officers, recognizing these cues is essential for effective deescalation, yet traditional training may not fully address their complexity. This paper focuses on body gestures and presents an automated system for recognizing specific body gestures relevant to social conflict situations, aiming to foster awareness for unconsciously performed body gestures and thereby enabling the training of de-escalation strategies. Karam Tomotaki-Dawoud, Birgit Nierula, Farelle Toumaleu Siewe, Daniel Johannes Meyer, Andreas Bock, Marianne Heinze, Daniela Knuth, Denis Martin, Julia Schander, Anna Hilsmann, Peter Eisert, Sebastian Bosse |
ISM | 11 |
| 2024 | Animatable Virtual Humans: Learning Pose-Dependent Human Representations in UV Space for Interactive Performance SynthesisabstractWe propose a novel representation of virtual humans for highly realistic real-time animation and rendering in 3D applications. We learn pose dependent appearance and geometry from highly accurate dynamic mesh sequences obtained from state-of-the-art multiview-video reconstruction. Learning pose-dependent appearance and geometry from mesh sequences poses significant challenges, as it requires the network to learn the intricate shape and articulated motion of a human body. However, statistical body models like SMPL provide valuable a-priori knowledge which we leverage in order to constrain the dimension of the search space, enabling more efficient and targeted learning and to define pose-dependency. Instead of directly learning absolute pose-dependent geometry, we learn the difference between the observed geometry and the fitted SMPL model. This allows us to encode both pose-dependent appearance and geometry in the consistent UV space of the SMPL model. This approach not only ensures a high level of realism but also facilitates streamlined processing and rendering of virtual humans in real-time scenarios. Wieland Morgenstern, Milena T. Bagdasarian, Anna Hilsmann, Peter Eisert |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Fooling State-of-the-art Deepfake Detection with High-quality DeepfakesabstractDue to the rising threat of deepfakes to security and privacy, it is most important to develop robust and reliable detectors. In this paper, we examine the need for high-quality samples in the training datasets of such detectors. Accordingly, we show that deepfake detectors proven to generalize well on multiple research datasets still struggle in real-world scenarios with well-crafted fakes. First, we propose a novel autoencoder for face swapping alongside an advanced face blending technique, which we utilize to generate 90 high-quality deepfakes. Second, we feed those fakes to a state-of-the-art detector, causing its performance to decrease drastically. Moreover, we fine-tune the detector on our fakes and demonstrate that they contain useful clues for the detection of manipulations. Overall, our results provide insights into the generalization of deepfake detectors and suggest that their training datasets should be complemented by high-quality fakes since training on mere research data is insufficient. Arian Beckmann, Anna Hilsmann, Peter Eisert |
IH&MMSec | 2 |
| 2023 | Unsupervised learning of style-aware facial animation from real acting performancesabstractThis paper presents a novel approach for text/speech-driven animation of a photo-realistic head model based on blend-shape geometry, dynamic textures, and neural rendering. Training a VAE for geometry and texture yields a parametric model for accurate capturing and realistic synthesis of facial expressions from a latent feature vector. Our animation method is based on a conditional CNN that transforms text or speech into a sequence of animation parameters. In contrast to previous approaches, our animation model learns disentangling/synthesizing different acting-styles in an unsupervised manner, requiring only phonetic labels that describe the content of training sequences. For realistic real-time rendering, we train a U-Net that refines rasterization-based renderings by computing improved pixel colors and a foreground matte. We compare our framework qualitatively/quantitatively against recent methods for head modeling as well as facial animation and evaluate the perceived rendering/animation quality in a user-study, which indicates large improvements compared to state-of-the-art approaches. Wolfgang Paier, Anna Hilsmann, Peter Eisert |
Graph. Model. | 2 |
| 2023 | Imposing temporal consistency on deep monocular body shape and pose estimationabstractAccurate and temporally consistent modeling of human bodies is essential for a wide range of applications, including character animation, understanding human social behavior, and AR/VR interfaces. Capturing human motion accurately from a monocular image sequence remains challenging; modeling quality is strongly influenced by temporal consistency of the captured body motion. Our work presents an elegant solution to integrating temporal constraints during fitting. This increases both temporal consistency and robustness during optimization. In detail, we derive parameters of a sequence of body models, representing shape and motion of a person. We optimize these parameters over the complete image sequence, fitting a single consistent body shape while imposing temporal consistency on the body motion, assuming body joint trajectories to be linear over short time. Our approach enables the derivation of realistic 3D body models from image sequences, including jaw pose, facial expression, and articulated hands. Our experiments show that our approach accurately estimates body shape and motion, even for challenging movements and poses. Further, we apply it to the particular application of sign language analysis, where accurate and temporally consistent motion modelling is essential, and show that the approach is well-suited to this kind of application. Alexandra Zimmer, Anna Hilsmann, Wieland Morgenstern, Peter Eisert |
Comput. Vis. Media | 2 |
| 2022 | CASAPose: Class-Adaptive and Semantic-Aware Multi-Object Pose Estimation
Niklas Gard, Anna Hilsmann, Peter Eisert |
BMVC | 2 |
| 2021 | Zero in on Shape: A Generic 2D-3D Instance Similarity Metric Learned from Synthetic DataabstractWe present a network architecture which compares RGB images and untextured 3D models by the similarity of the represented shape. Our system is optimised for zero-shot retrieval, meaning it can recognise shapes never shown in training. We use a view-based shape descriptor and a siamese network to learn object geometry from pairs of 3D models and 2D images.Due to scarcity of datasets with exact photograph-mesh correspondences, we train our network with only synthetic data.Our experiments investigate the effect of different qualities and quantities of training data on retrieval accuracy and present insights from bridging the domain gap. We show that increasing the variety of synthetic data improves retrieval accuracy and that our system’s performance in zero-shot mode can match that of the instance-aware mode, as far as narrowing down the search to the top 10% of objects. Maciej Janik, Niklas Gard, Anna Hilsmann, Peter Eisert |
ICIP | 3 |
| 2020 | EEG-Based Assessment of Perceived Realness in Stylized Face ImagesabstractIn this paper, we investigate the perception of realness in rendered face images experimentally using electroencephalography. To this end, we presented ten subjects with 36 character images based on six different faces (varying in gender and emotional expression) rendered at six different levels of realness ranging from abstract, cartoon-like renderings to real photographs. In the first psychophysical part of our study, we asked participants to rate perceived realness, appeal, familiarity, reassurance, and attractiveness for the presented characters. In the second part, we recorded the electroencephalogram when presenting the character images at a stimulation frequency of fstim= 5 Hz. We show that the amplitudes of the odd harmonics of the elicited steady-state visual evoked potential correlate with the psychophysical responses (|ρ| = 0.83, p < 0.05). Milena T. Bagdasarian, Anna Hilsmann, Peter Eisert, Gabriel Curio, Klaus-Robert Müller, Thomas Wiegand 0001, Sebastian Bosse |
QoMEX | 2 |
| 2020 | Going beyond free viewpoint: creating animatable volumetric video of human performancesabstractAn end‐to‐end pipeline for the creation of high‐quality animatable volumetric video of human performances is presented. Going beyond the application of free‐viewpoint video, the authors allow re‐animation and alteration of an actor's performance through the enrichment of the captured data with semantics and animation properties. Hybrid geometry‐ and video‐based animation methods are applied that allow a direct animation of the high‐quality data itself instead of creating a CG model that resembles the captured data. Semantic enrichment and animation are achieved by establishing temporal consistency followed by automatic rigging of each 3D frame using a parametric human body model. The hybrid approach combines the flexibility of classical CG animation with the realism of real captured data. For the face, coarse movements are modelled in the geometry only, while very fine and subtle details, often lacking in purely geometric methods, are captured in video textures, which can interactively be combined to form new facial expressions. On top of that, regions that are challenging to synthesise, such as the teeth or the eyes, are learned and filled in realistically in an autoencoder‐based approach. This study covers the full pipeline from capturing, volumetric video production, and enrichment with semantics for the final hybrid animation. Anna Hilsmann, Philipp Fechteler, Wieland Morgenstern, Wolfgang Paier, Ingo Feldmann, Oliver Schreer, Peter Eisert |
IET Comput. Vis. | 1 |
| 2020 | Interactive facial animation with deep neural networksabstractCreating realistic animations of human faces is still a challenging task in computer graphics. While computer graphics (CG) models capture much variability in a small parameter vector, they usually do not meet the necessary visual quality. This is due to the fact, that geometry‐based animation often does not allow fine‐grained deformations and fails in difficult areas (mouth, eyes) to produce realistic renderings. Image‐based animation techniques avoid these problems by using dynamic textures that capture details and small movements that are not explained by geometry. This comes at the cost of high‐memory requirements and limited flexibility in terms of animation because dynamic texture sequences need to be concatenated seamlessly, which is not always possible and prone to visual artefacts. In this study, the authors present a new hybrid animation framework that exploits recent advances in deep learning to provide an interactive animation engine that can be used via a simple and intuitive visualisation for facial expression editing. The authors describe an automatic pipeline to generate training sequences that consist of dynamic textures plus sequences of consistent three‐dimensional face models. Based on this data, they train a variational autoencoder to learn a low‐dimensional latent space of facial expressions that is used for interactive facial animation. Wolfgang Paier, Anna Hilsmann, Peter Eisert |
IET Comput. Vis. | 2 |
| 2020 | Accurate and robust neural networks for face morphing attack detectionabstractArtificial neural networks tend to use only what they need for a task. For example, to recognize a rooster, a network might only considers the rooster’s red comb and wattle and ignores the rest of the animal. This makes them vulnerable to attacks on their decision making process and can worsen their generality. Thus, this phenomenon has to be considered during the training of networks, especially in safety and security related applications. In this paper, we propose neural network training schemes, which are based on different alternations of the training data, to increase robustness and generality. Precisely, we limit the amount and position of information available to the neural network for the decision making process and study their effects on the accuracy, generality, and robustness against semantic and black box attacks for the particular example of face morphing attacks. In addition, we exploit layer-wise relevance propagation (LRP) to analyze the differences in the decision making process of the differently trained neural networks. A face morphing attack is an attack on a biometric facial recognition system, where the system is fooled to match two different individuals with the same synthetic face image. Such a synthetic image can be created by aligning and blending images of the two individuals that should be matched with this image. We train neural networks for face morphing attack detection using our proposed training schemes and show that they lead to an improvement of robustness against attacks on neural networks. Using LRP, we show that the improved training forces the networks to develop and use reliable models for all regions of the analyzed image. This redundancy in representation is of crucial importance to security related applications. Clemens Seibold, Wojciech Samek, Anna Hilsmann, Peter Eisert |
J. Inf. Secur. Appl. | 3 |
| 2019 | Interactive and Multimodal-based Augmented Reality for Remote Assistance using a Digital Surgical MicroscopeabstractWe present an interactive and multimodal-based augmented reality system for computer-assisted surgery in the context of ear, nose and throat (ENT) treatment. The proposed processing pipeline uses fully digital stereoscopic imaging devices, which support multispectral and white light imaging to generate high resolution image data, and consists of five modules. Input/output data handling, a hybrid multimodal image analysis and a bi-directional interactive augmented reality (AR) and mixed reality (MR) interface for local and remote surgical assistance are of high relevance for the complete framework. The hybrid multimodal 3D scene analysis module uses different wavelengths to classify tissue structures and combines this spectral data with metric 3D information. Additionally, we propose a zoom-independent intraoperative tool for virtual ossicular prosthesis insertion (e.g. stapedectomy) guaranteeing very high metric accuracy in sub-millimeter range (1/10 mm). A bi-directional interactive AR/MR communication module guarantees low latency, while consisting surgical information and avoiding informational overload. Display agnostic AR/MR visualization can show our analyzed data synchronized inside the digital binocular, the 3D display or any connected head-mounted-display (HMD). In addition, the analyzed data can be enriched with annotations by involving external clinical experts using AR/MR and furthermore an accurate registration of preoperative data. The benefits of such a collaborative surgical system are manifold and will lead to a highly improved patient outcome through an easier tissue classification and reduced surgery risk. Eric L. Wisotzky, Jean-Claude Rosenthal, Peter Eisert, Anna Hilsmann, Falko Schmid, Armin Schneider, Florian C. Uecker |
VR | 4 |
| 2019 | Markerless Multiview Motion Capture with 3D Shape Model AdaptationabstractAbstract In this paper, we address simultaneous markerless motion and shape capture from 3D input meshes of partial views onto a moving subject. We exploit a computer graphics model based on kinematic skinning as template tracking model. This template model consists of vertices, joints and skinning weights learned a priori from registered full‐body scans, representing true human shape and kinematics‐based shape deformations. Two data‐driven priors are used together with a set of constraints and cues for setting up sufficient correspondences. A Gaussian mixture model‐based pose prior of successive joint configurations is learned to soft‐constrain the attainable pose space to plausible human poses. To make the shape adaptation robust to outliers and non‐visible surface regions and to guide the shape adaptation towards realistically appearing human shapes, we use a mesh‐Laplacian‐based shape prior. Both priors are learned/extracted from the training set of the template model learning phase. The output is a model adapted to the captured subject with respect to shape and kinematic skeleton as well as the animation parameters to resemble the observed movements. With example applications, we demonstrate the benefit of such footage. Experimental evaluations on publicly available datasets show the achieved natural appearance and accuracy. Philipp Fechteler, Anna Hilsmann, Peter Eisert |
Comput. Graph. Forum | 2 |
| 2019 | Projection Distortion-based Object Tracking in Shader Lamp ScenariosabstractShader lamp systems augment the real environment by projecting new textures on known target geometries. In dynamic scenes, object tracking maintains the illusion if the physical and virtual objects are well aligned. However, traditional trackers based on texture or contour information are often distracted by the projected content and tend to fail. In this paper, we present a model-based tracking strategy, which directly takes advantage from the projected content for pose estimation in a projector-camera system. An iterative pose estimation algorithm captures and exploits visible distortions caused by object movements. In a closed-loop, the corrected pose allows the update of the projection for the subsequent frame. Synthetic frames simulating the projection on the model are rendered and an optical flow-based method minimizes the difference between edges of the rendered and the camera image. Since the thresholds automatically adapt to the synthetic image, a complicated radiometric calibration can be avoided. The pixel-wise linear optimization is designed to be easily implemented on the GPU. Our approach can be combined with a regular contour-based tracker and is transferable to other problems, like the estimation of the extrinsic pose between projector and camera. We evaluate our procedure with real and synthetic images and obtain very precise registration results. Niklas Gard, Anna Hilsmann, Peter Eisert |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | Animatable 3D Model Generation from 2D Monocular Visual DataabstractIn this paper, we present an approach for creating animatable 3D models from temporal monocular image acquisitions of non-rigid objects. During deformation, the object of interest is captured with only a single camera under full perspective projection. The aim of the presented framework is to obtain a shape deformation model in terms of joints and skinning weights that can finally be used for animating the model vertices. First, the monocular rigid shape estimation problem is solved by computing a template model of the object in rest pose from an image sequence. Next, the unknown external camera parameters and the deformation for each vertex are estimated alternately in a sequential approach. The resulting consistent non-rigid shape geometries are used to compute a kinematic skeleton control structure including skinning weights and optimized shape. For that, a completely data-driven optimization scheme is used, which iterates over three steps: (a) optimization of pose for each frame as well as joint parameters consistent over the entire sequence, (b) optimization of rest pose vertices to enhance the shape and (c) optimization of skinning weights for improved deformation characteristics. With experimental results on publicly available synthetic as well as real-world datasets, we demonstrate the quality of the proposed approach. The resulting models with fixed topology and rigged with skeleton and skinning weights can be animated in existing render engines. Philipp Fechteler, Lisa Kausch, Anna Hilsmann, Peter Eisert |
ICIP | 3 |
| 2018 | Surface tracking assessment and interaction in texture spaceabstractIn this paper, we present a novel approach for assessing and interacting with surface tracking algorithms targeting video manipulation in post-production. As tracking inaccuracies are unavoidable, we enable the user to provide small hints to the algorithms instead of correcting erroneous results afterwards. Based on 2D mesh warp-based optical flow estimation, we visualize results and provide tools for user feedback in a consistent reference system, texture space. In this space, accurate tracking results are reflected by static appearance, and errors can easily be spotted as apparent change. A variety of established tools can be utilized to visualize and assess the change between frames. User interaction to improve tracking results becomes more intuitive in texture space, as it can focus on a small region rather than a moving object. We show how established tools can be implemented for interaction in texture space to provide a more intuitive interface allowing more effective and accurate user feedback. Johannes Furch, Anna Hilsmann, Peter Eisert |
Comput. Vis. Media | 2 |
| 2017 | Detection of Face Morphing Attacks by Deep Learning
Clemens Seibold, Wojciech Samek, Anna Hilsmann, Peter Eisert |
IWDW | 3 |
| 2017 | Model-based motion blur estimation for the improvement of motion trackingabstractVideo tracking is an important task in many automated or semi-automated applications, like cinematic post production, surveillance or traffic monitoring. Most established video tracking methods fail or lead to an inaccurate estimate when motion blur occurs in the video, as they assume, that the object appears constantly sharp in the video. In this paper, we present a novel motion tracking method with explicit modeling of motion blur, estimating the continuous motion of a rigid 3-D object with known geometry in a monocular video as well as the sharp object texture. Instead of treating motion blur as a potential source of errors, we take advantage of it and consider motion blur as an additional information source, providing information about the motion of the tracked object during the exposure. In an analysis-by-synthesis approach we explicitly model the effects of motion blur reconstructing the captured frames, in order to accomplish a more accurate estimation. We design our algorithm to be capable to run in parallel on the GPU using the common rendering pipeline and considering each frame individually to handle also long videos. We tested our approach on both synthetic and real videos. In both cases, we achieve significant improvements of accuracy and reductions of frame reconstruction error compared to the estimated motion of a rigid body tracker, without motion blur handling. Clemens Seibold, Anna Hilsmann, Peter Eisert |
Comput. Vis. Image Underst. | 2 |
| 2017 | A Hybrid Approach for Facial Performance Analysis and EditingabstractAs of today, fine-grained editing of facial performances in movie and video production requires either retouching every single frame or creating a highly detailed CGI model of the actor, both of which is restricted to high-budget productions. In this paper, we present an example-based approach for facial performance editing that achieves realistic results with standard equipment and very little manual intervention. Based on a model-free surface tracking approach, temporally consistent dynamic texture sequences are extracted from multiple video streams. Using such geometry-plus-texture sequences allows transferring facial expressions/performances between videos, which enables the editor, for example, to change the facial expression while leaving the rest of the video untouched. Moreover, concatenating and/or looping these sequences in a motion-graph-like manner offers a convenient way of composing novel facial performances from multiple source videos. Finally, we present a blending method to seamlessly concatenate texture/mesh sequences and to insert a composed facial performance in a target video. Wolfgang Paier, Markus Kettern, Anna Hilsmann, Peter Eisert |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Real-time avatar animation with dynamic face texturingabstractIn this paper, we present a system to capture and animate a highly realistic avatar model of a user in real-time. The animated human model consists of a rigged 3D mesh and a texture map. The system is based on KinectV2 input which captures the skeleton of the current pose of the subject in order to animate the human shape model. An additional high-resolution RGB camera is used to capture the face for updating the texture map on each frame. With this combination of image based rendering with computer graphics we achieve photo-realistic animations in real-time. Additionally, this approach is well suited for networked scenarios, because of the low per frame amount of data to animate the model, which consists of motion capture parameters and a video frame. With experimental results, we demonstrate the high degree of realism of the presented approach. Philipp Fechteler, Wolfgang Paier, Anna Hilsmann, Peter Eisert |
ICIP | 3 |
| 2015 | A framework for image-based asset generation and animationabstractCreating digital animatable models of real-world objects and characters is important for many applications, ranging from highly expensive movie productions to low-cost real-time applications like computer games and augmented reality. However, achieving real photorealism with convincing appearance and deformation behavior requires sophisticated capturing, elaborate manual modeling and time-consuming simulation. This can only be achieved in well funded film productions, while in low-cost applications, animated objects usually lack visual quality. In this paper, we present a new framework for image-based animatable asset generation which avoids these time-consuming processes both in the modeling and the simulation stage. Real-time photo-realistic animation is enabled by the use of captured images and shifting computational complexity to an a-priori training phase. Our paper covers the complete pipeline of content creation, asset generation and representation, and a real-time animation and rendering implementation. Johannes Furch, Anna Hilsmann, Peter Eisert |
ICIP | 2 |
| 2015 | Interactive Scene Flow Editing for Improved Image-based Rendering and Virtual Spacetime NavigationabstractHigh-quality stereo and optical flow maps are essential for a multitude of tasks in visual media production, e.g. virtual camera navigation, disparity adaptation or scene editing. Rather than estimating stereo and optical flow separately, scene flow is a valid alternative since it combines both spatial and temporal information and recently surpassed the former two in terms of accuracy. However, since automated scene flow estimation is non-accurate in a number of situations, resulting rendering artifacts have to be corrected manually in each output frame, an elaborate and time-consuming task. We propose a novel workflow to edit the scene flow itself, catching the problem at its source and yielding a more flexible instrument for further processing. By integrating user edits in early stages of the optimization, we allow the use of approximate scribbles instead of accurate editing, thereby reducing interaction times. Our results show that editing the scene flow improves the quality of visual results considerably while requiring vastly less editing effort. Kai Ruhl, Martin Eisemann, Anna Hilsmann, Peter Eisert, Marcus A. Magnor |
ACM Multimedia | 3 |
| 2013 | Pose Space Image Based RenderingabstractAbstract This paper introduces a new image‐based rendering approach for articulated objects with complex pose‐dependent appearance, such as clothes. Our approach combines body‐pose‐dependent appearance and geometry to synthesize images of new poses from a database of examples. A geometric model allows animation and view interpolation, while small details as well as complex shading and reflection properties are modeled by pose‐dependent appearance examples in a database. Correspondences between the images are represented as mesh‐based warps, both in the spatial and intensity domain. For rendering, these warps are interpolated inpose space, i.e. the space of body poses, using scattered data interpolation methods. Warp estimation as well as geometry reconstruction is performed in an offline procedure, thus shifting computational complexity to an a‐priori training phase. Anna Hilsmann, Philipp Fechteler, Peter Eisert |
Comput. Graph. Forum | 1 |
| 2011 | Template-free Shape From Texture with Perspective Camerasabstract10 S. Anna Hilsmann, David C. Schneider, Peter Eisert |
BMVC | 1 |
| 2010 | Virtual jewel rendering for augmented reality environmentsabstractVirtual Mirrors are augmented reality applications that capture a viewer's image and add 2D or 3D overlaid virtual elements, such as cloth patterns, or replace existing pieces of garments, such as shoes, by virtual ones. The purpose of our study is to focus on the addition of jewels that present complex rendering issues due to reflection and translucency, especially for real-time applications. We propose two complementary approaches which are both suited for particular objects. The image-based rendering technique relies on an image data-set of photos taken on a semi-circle of camera positions and a 3D reconstruction of the jewel. This approach is exploited to demonstrate real jewels with sophisticated refractions and highlights. The analytical 3D rendering technique can also visualize configurable virtual jewels and is based on the decomposition of the jewel into parts without non-local reflection or transparency. A graph-based rendering chain takes in consideration some of the self-reflection and lighting effects. Peter Eisert, Christian Jacquemin, Anna Hilsmann |
ICIP | 3 |
| 2010 | Patch-based reconstruction and rendering of human headsabstractReconstructing the 3D shape of human faces is an intensively researched topic. Most approaches aim at generating a closed surface representation of geometry, i.e. a mesh, which is texture-mapped for rendering. However, if free viewpoint rendering is the primary purpose of the reconstruction, representations other than meshes are possible. In this paper a coarse patch-based approach to both reconstruction and rendering is explored and applied not only to the face but the whole human head. The approach has advantages on parts of the scene that are traditionally difficult to reconstruct and render, which is the case for hair when it comes to human heads. In the paper, reconstruction of a patch is posed as a parameter estimation problem which is solved in a generic image-based optimization framework using the Levenberg-Marquard algorithm. In order to improve robustness, the Huber error metric is used and a geometric regularization strategy is introduced. Initial values for the optimization, which are crucial for the method's success, are obtained by triangulation of SIFT feature points and a recursive expansion scheme. David C. Schneider, Anna Hilsmann, Peter Eisert |
ICIP | 2 |
| 2010 | Realistic cloth augmentation in single view video under occlusions
Anna Hilsmann, David C. Schneider, Peter Eisert |
Comput. Graph. | 1 |
| 2008 | Optical flow based tracking and retexturing of garmentsabstractIn this paper, we present a method for tracking and retexturing of garments that exploits the entire image information using the optical flow constraint instead of working with distinct features. In a hierarchical framework we refine the motion model with every level. The motion model is used to regularize the optical flow field such that finding the best transformation amounts in minimizing an error function that can be solved in a least squares sense. Knowledge about the position and deformation of the garment in 2D allows us to erase the old texture and replace it by a new one with correct deformation and shading properties without 3D reconstruction. Additionally, it provides an estimation of the irradiance such that the new texture can be illuminated realistically. Anna Hilsmann, Peter Eisert |
ICIP | 1 |