Bodo Rosenhahn

dblp:09/2973 · DBLP profile ↗
← Back
151ranked-venue papers
13as first author
44since 2021 · last 2026
0000-0003-3861-1424ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 103 · 8 first-author · 22 since 2021Artificial intelligence and machine learning · 91 · 7 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 since 2021Systems, architecture and hardware · 5 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Improved Convex Decomposition with Ensembling and Negative Primitives
abstract
Describing a scene in terms of primitives - geometrically simple shapes that offer a parsimonious but accurate abstraction of structure - is an established and difficult fitting problem. Different scenes require different numbers of primitives, and these primitives interact strongly. Existing methods are evaluated by comparing predicted depth, normals, and segmentation against ground truth. The state of the art method involves a learned regression procedure to predict a start point consisting of a fixed number of primitives, followed by a descent method to refine the geometry and remove redundant primitives. CSG (Constructive Solid Geometry) representations are significantly enhanced by a set-differencing operation. Our representation incorporates negative primitives, which are differenced from the positive primitives. These notably enrich the geometry that the model can encode, while complicating the fitting problem. This paper presents a method that can (a) incorporate these negative primitives and (b) choose the overall number of positive and negative primitives by ensembling. Extensive experiments on the standard NYUv2 dataset confirm that (a) this approach results in substantial improvements in depth representation and segmentation over SOTA and (b) negative primitives improve fitting accuracy. Our method is robustly applicable across datasets: in a first, we evaluate primitive prediction for LAION images.
Vaibhav Vavilala, Florian Kluger, Seemandhar Jain, Bodo Rosenhahn, Anand Bhattad, David A. Forsyth
3DV4
2026 Improving 3D Foot Motion Reconstruction in Markerless Monocular Human Motion Capture
abstract
State-of-the-art methods can recover accurate overall 3D human body motion from in-the-wild videos. However, they often fail to capture fine-grained articulations, especially in the feet, which are critical for applications such as gait analysis and animation. This limitation results from training datasets with inaccurate foot annotations and limited foot motion diversity. We address this gap with FootMR, a Foot Motion Refinement method that refines foot motion estimated by an existing human recovery model through lifting 2D foot keypoint sequences to 3D. By avoiding direct image input, FootMR circumvents inaccurate image-3D annotation pairs and can instead leverage large-scale motion capture data. To resolve ambiguities of 2D-to-3D lifting, FootMR incorporates knee and foot motion as context and predicts only residual foot motion. Generalization to extreme foot poses is further improved by representing joints in global rather than parent-relative rotations and applying extensive data augmentation. To support evaluation of foot motion reconstruction, we introduce MOOF, a$2D$dataset of complex foot movements. Experiments on MOOF, MOYO, and RICH show that FootMR outperforms state-of-the-art methods, reducing ankle joint angle error on MOYO by up to 30 % over the best video-based approach. Our code and dataset are available for research purposes at twehrbein.github.io/footmr-website/.
Tom Wehrbein, Bodo Rosenhahn
3DV2
2026 From Gameplay Traces to Game Mechanics: Causal Induction with Large Language Models
Mohit Jiwatode, Alexander Dockhorn, Bodo Rosenhahn
ICPR (13)3
2026 SPAN: Learning Similarity Between Scene Graphs and Images With Transformers
abstract
Learning similarity between scene graphs and images aims to estimate a similarity score given a scene graph and an image. There is currently no research dedicated to this task, although it is critical for scene graph generation and downstream applications. Scene graph generation is conventionally evaluated by Recall$@K$@K and mean Recall$@K$@K, which measure the ratio of predicted triplets that appear in the human-labeled triplet set. However, such triplet-oriented metrics fail to demonstrate the overall semantic difference between a scene graph and an image and are sensitive to annotation bias and noise. Using generated scene graphs in the downstream applications is therefore limited. To address this issue, for the first time, we propose a Scene graPh-imAge coNtrastive learning framework, SPAN, that can measure the similarity between scene graphs and images. Our novel framework consists of a graph Transformer and an image Transformer to align scene graphs and their corresponding images in the shared latent space. We introduce a novel graph serialization technique that transforms a scene graph into a sequence with structural encodings. Based on our framework, we propose R-Precision measuring image retrieval accuracy as a new evaluation metric for scene graph generation. We establish new benchmarks on the Visual Genome and Open Images datasets. Extensive experiments are conducted to verify the effectiveness of SPAN, which shows great potential as a scene graph encoder.
Yuren Cong, Wentong Liao, Bodo Rosenhahn, Michael Ying Yang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 QPM: Discrete Optimization for Globally Interpretable Image Classification
abstract
Understanding the classifications of deep neural networks, e.g. used in safety-critical situations, is becoming increasingly important. While recent models can locally explain a single decision, to provide a faithful global explanation about an accurate model’s general behavior is a more challenging open task. Towards that goal, we introduce the Quadratic Programming Enhanced Model (QPM), which learns globally interpretable class representations. QPM represents every class with a binary assignment of very few, typically 5, features, that are also assigned to other classes, ensuring easily comparable contrastive class representations. This compact binary assignment is found using discrete optimization based on predefined similarity measures and interpretability constraints. The resulting optimal assignment is used to fine-tune the diverse features, so that each of them becomes the shared general concept between the assigned classes. Extensive evaluations show that QPM delivers unprecedented global interpretability across small and large-scale datasets while setting the state of the art for the accuracy of interpretable models.
Thomas Norrenbrock, Timo Kaiser, Sovan Biswas, Ramesh Manuvinakurike, Bodo Rosenhahn
ICLR5
2025 UncertainSAM: Fast and Efficient Uncertainty Quantification of the Segment Anything Model
abstract
The introduction of the Segment Anything Model (SAM) has paved the way for numerous semantic segmentation applications. For several tasks, quantifying the uncertainty of SAM is of particular interest. However, the ambiguous nature of the class-agnostic foundation model SAM challenges current uncertainty quantification (UQ) approaches. This paper presents a theoretically motivated uncertainty quantification model based on a Bayesian entropy formulation jointly respecting aleatoric, epistemic, and the newly introduced task uncertainty. We use this formulation to train USAM, a lightweight post-hoc UQ method. Our model traces the root of uncertainty back to under-parameterised models, insufficient prompts or image ambiguities. Our proposed deterministic USAM demonstrates superior predictive capabilities on the SA-V, MOSE, ADE20k, DAVIS, and COCO datasets, offering a computationally cheap and easy-to-use UQ alternative that can support user-prompting, enhance semi-supervised pipelines, or balance the tradeoff between accuracy and cost efficiency.
Timo Kaiser, Thomas Norrenbrock, Bodo Rosenhahn
ICML3
2025 Explainable Reinforcement Learning via Dynamic Mixture Policies
abstract
Learning control policies using deep reinforcement learning has shown great success for a variety of applications, including robotics and automated driving. A key area limiting the adaptation of RL in the real world is the lack of trust in the decision-making process of such policies. Therefore, explainability is a requirement of any RL agent operating in the real world. In this work, we propose a family of control policies that are explainable-by-design regarding individual observation components on object-based scene representations. By estimating diagonal squashed Gaussian and categorical mixture distributions on sub-spaces of the decomposed observations, we develop stochastic policies with easy-to-read explanations of the decision-making process. Our design is generally applicable to any RL algorithm using stochastic policies. We showcase the explainability on an extensive suite of single-and multi-agent simulations, set-and sequence-based high-level scenes, and discrete and continuous action spaces, with performance at least on-par or better compared to standard policy architectures. In additional experiments, we analyze the robustness of our approach to its single additional hyper-parameter and examine its potential for very low computational requirements with tiny policies.
Maximilian Schier, Frederik Schubert, Bodo Rosenhahn
ICRA3
2025 CHiQPM: Calibrated Hierarchical Interpretable Image Classification
abstract
Globally interpretable models are a promising approach for trustworthy AI in safety-critical domains. Alongside global explanations, detailed local explanations are a crucial complement to effectively support human experts during inference. This work proposes the Calibrated Hierarchical QPM (CHiQPM) which offers uniquely comprehensive global and local interpretability, paving the way for human-AI complementarity. CHiQPM achieves superior global interpretability by contrastively explaining the majority of classes and offers novel hierarchical explanations that are more similar to how humans reason and can be traversed to offer a built-in interpretable Conformal prediction (CP) method. Our comprehensive evaluation shows that CHiQPM achieves state-of-the-art accuracy as a point predictor, maintaining 99% accuracy of non-interpretable models. This demonstrates a substantial improvement, where interpretability is incorporated without sacrificing overall accuracy. Furthermore, its calibrated set prediction is competitively efficient to other CP methods, while providing interpretable predictions of coherent sets along its hierarchical explanation.
Thomas Norrenbrock, Timo Kaiser, Sovan Biswas, Neslihan Kose, Ramesh Manuvinakurike, Bodo Rosenhahn
NeurIPS6
2025 Utilizing Uncertainty in 2D Pose Detectors for Probabilistic 3D Human Mesh Recovery
Tom Wehrbein, Marco Rudolph, Bodo Rosenhahn, Bastian Wandt
WACV3
2025 Attribute-Centric Compositional Text-to-Image Generation
abstract
Abstract Despite the recent impressive breakthroughs in text-to-image generation, generative models have difficulty in capturing the data distribution of underrepresented attribute compositions while over-memorizing overrepresented attribute compositions, which raises public concerns about their robustness and fairness. To tackle this challenge, we propose ACTIG, an attribute-centric compositional text-to-image generation framework. We present an attribute-centric feature augmentation and a novel image-free training scheme, which greatly improves model’s ability to generate images with underrepresented attributes. We further propose an attribute-centric contrastive loss to avoid overfitting to overrepresented attribute compositions. We validate our framework on the CelebA-HQ and CUB datasets. Extensive experiments show that the compositional generalization of ACTIG is outstanding, and our framework outperforms previous works in terms of image quality and text-image consistency. The source code and trained models are publicly available at https://github.com/yrcong/ACTIG .
Yuren Cong, Martin Renqiang Min, Li Erran Li, Bodo Rosenhahn, Michael Ying Yang
Int. J. Comput. Vis.4
2025 Guest Editorial: Special Issue on Multimodal Learning
Michael Ying Yang, Paolo Rota, Massimiliano Mancini, Pietro Morerio, Bodo Rosenhahn, Vittorio Murino
Int. J. Comput. Vis.5
2025 Cell Tracking According to Biological Needs - Strong Mitosis-Aware Multi- Hypothesis Tracker With Aleatoric Uncertainty
abstract
Cell tracking and segmentation enable biologists to extract insights from large-scale microscopy time-lapse data. Driven by local accuracy metrics, current tracking approaches often suffer from a lack of long-term consistency and an inability to correctly reconstruct lineage trees. To address this issue, we introduce a novel assignment strategy consisting of two key components. First, we propose an uncertainty estimation technique for motion estimation frameworks. This method relaxes single-point motion representations into probabilistic spatial densities using problem-specific test-time augmentations. Second, we leverage these spatial densities to define a novel mitosis-aware assignment problem formulation. This formulation allows multi-hypothesis trackers to model cell divisions and resolve false associations and mitosis detections based on long-term conflicts. Our framework integrates explicit biological knowledge into assignment costs and combines it with learned representations derived from spatial densities. We evaluate our approach on nine competitive datasets and demonstrate that it substantially outperforms the current state-of-the-art on biologically inspired metrics, achieving improvements by a factor of approximately six and providing new insights into the behavior of motion estimation uncertainty.
Timo Kaiser, Maximilian Schier, Bodo Rosenhahn
IEEE Trans. Medical Imaging3
2024 PARSAC: Accelerating Robust Multi-Model Fitting with Parallel Sample Consensus
abstract
We present a real-time method for robust estimation of multiple instances of geometric models from noisy data. Geometric models such as vanishing points, planar homographies or fundamental matrices are essential for 3D scene analysis. Previous approaches discover distinct model instances in an iterative manner, thus limiting their potential for speedup via parallel computation. In contrast, our method detects all model instances independently and in parallel. A neural network segments the input data into clusters representing potential model instances by predicting multiple sets of sample and inlier weights. Using the predicted weights, we determine the model parameters for each potential instance separately in a RANSAC-like fashion. We train the neural network via task-specific loss functions, i.e. we do not require a ground-truth segmentation of the input data. As suitable training data for homography and fundamental matrix fitting is scarce, we additionally present two new synthetic datasets. We demonstrate state-of-the-art performance on these as well as multiple established datasets, with inference times as small as five milliseconds per image.
Florian Kluger, Bodo Rosenhahn
AAAI2
2024 Q-SENN: Quantized Self-Explaining Neural Networks
abstract
Explanations in Computer Vision are often desired, but most Deep Neural Networks can only provide saliency maps with questionable faithfulness. Self-Explaining Neural Networks (SENN) extract interpretable concepts with fidelity, diversity, and grounding to combine them linearly for decision-making. While they can explain what was recognized, initial realizations lack accuracy and general applicability. We propose the Quantized-Self-Explaining Neural Network “Q-SENN”. Q-SENN satisfies or exceeds the desiderata of SENN while being applicable to more complex datasets and maintaining most or all of the accuracy of an uninterpretable baseline model, outperforming previous work in all considered metrics. Q-SENN describes the relationship between every class and feature as either positive, negative or neutral instead of an arbitrary number of possible relations, enforcing more binary human-friendly features. Since every class is assigned just 5 interpretable features on average, Q-SENN shows convincing local and global interpretability. Additionally, we propose a feature alignment method, capable of aligning learned features with human language-based concepts without additional supervision. Thus, what is learned can be more easily verbalized. The code is published: https://github.com/ThomasNorr/Q-SENN
Thomas Norrenbrock, Marco Rudolph, Bodo Rosenhahn
AAAI3
2024 Segment Any Object Model (SAOM): Real-To-Simulation Fine-Tuning Strategy For Multi-Class Multi-Instance Segmentation
abstract
Multi-class multi-instance segmentation is the task of identifying masks for multiple object classes and multiple instances of the same class within an image. The foundational Segment Anything Model (SAM) is designed for promptable multi-class multi-instance segmentation but tends to output part or sub-part masks in the “everything” mode for various real-world applications. Whole object segmentation masks play a crucial role for indoor scene understanding, especially in robotics applications. We propose a new domain invariant Real-to-Simulation (Real-Sim) fine-tuning strategy for SAM. We use object images and ground truth data collected from Ai2Thor simulator during fine-tuning (real-to-sim). To allow our Segment Any Object Model (SAOM) to work in the “everything” mode, we propose the novel nearest neighbour assignment method, updating point embeddings for each ground-truth mask. SAOM is evaluated on our own dataset collected from Ai2Thor simulator. SAOM significantly improves on SAM, with a $28 \%$ increase in mIoU and a $25 \%$ increase in mAcc for 54 frequently-seen indoor object classes. Moreover, our Real-to-Simulation fine-tuning strategy demonstrates promising generalization performance in real environments without being trained on the real-world data (sim-to-real). The dataset and the code are available here.
Mariia Khan, Yue Qiu 0001, Yuren Cong, Bodo Rosenhahn, Jumana M. Abu-Khalaf, David Suter
ICIP4
2024 FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
abstract
Text-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts. A major challenge in this task is to ensure that all frames in the edited video are visually consistent. Most recent works apply advanced text-to-image diffusion models to this task by inflating 2D spatial attention in the U-Net into spatio-temporal attention. Although temporal context can be added through spatio-temporal attention, it may introduce some irrelevant information for each patch and therefore cause inconsistency in the edited video. In this paper, for the first time, we introduce optical flow into the attention module in diffusion model's U-Net to address the inconsistency issue for text-to-video editing. Our method, FLATTEN, enforces the patches on the same flow path across different frames to attend to each other in the attention module, thus improving the visual consistency in the edited videos. Additionally, our method is training-free and can be seamlessly integrated into any diffusion based text-to-video editing methods and improve their visual consistency. Experiment results on existing text-to-video editing benchmarks show that our proposed method achieves the new state-of-the-art performance. In particular, our method excels in maintaining the visual consistency in the edited videos.
Yuren Cong, Mengmeng Xu 0006, Christian Simon, Shoufa Chen, Jiawei Ren 0001, Yanping Xie, Juan-Manuel Pérez-Rúa, Bodo Rosenhahn, Tao Xiang 0002, Sen He 0001
ICLR8
2024 Mastering Zero-Shot Interactions in Cooperative and Competitive Simultaneous Games
abstract
The combination of self-play and planning has achieved great successes in sequential games, for instance in Chess and Go. However, adapting algorithms such as AlphaZero to simultaneous games poses a new challenge. In these games, missing information about concurrent actions of other agents is a limiting factor as they may select different Nash equilibria or do not play optimally at all. Thus, it is vital to model the behavior of the other agents when interacting with them in simultaneous games. To this end, we propose Albatross: AlphaZero for Learning Bounded-rational Agents and Temperature-based Response Optimization using Simulated Self-play. Albatross learns to play the novel equilibrium concept of a Smooth Best Response Logit Equilibrium (SBRLE), which enables cooperation and competition with agents of any playing strength. We perform an extensive evaluation of Albatross on a set of cooperative and competitive simultaneous perfect-information games. In contrast to AlphaZero, Albatross is able to exploit weak agents in the competitive game of Battlesnake. Additionally, it yields an improvement of 37.6% compared to previous state of the art in the cooperative Overcooked benchmark.
Yannik Mahlau, Frederik Schubert, Bodo Rosenhahn
ICML3
2024 Indoor Scene Change Understanding (SCU): Segment, Describe, and Revert Any Change
abstract
Understanding of scene changes is crucial for embodied AI applications, such as visual room rearrangement, where the agent must revert changes by restoring the objects to their original locations or states. Visual changes between two scenes, pre- and post-rearrangement, encompass two tasks: scene change detection (locating changes) and image difference captioning (describing changes). While previous methods, focused on sequential 2D images, have addressed these tasks separately, it is essential to emphasize the significance of their combination. Therefore, we propose a new Scene Change Understanding (SCU) task for simultaneous change detection and description. Moreover, we go beyond change language description generation and aim to generate rearrangement instructions for the robotic agent to revert changes. To solve this task, we propose a novel method - EmbSCU, which allows to compare instance-level change object masks (for 53 frequently-seen indoor object classes) before and after changes and generate rearrangement language instructions for the agent. EmbSCU is built on our Segment Any Object Model (SAOMv2) - a fine-tuned version of Segment Anything Model (SAM), adapted to obtain instance-level object masks for both foreground and background objects in indoor embodied environments. EmbSCU is evaluated on our own dataset of sequential 2D image pairs before and after changes, collected from the Ai2Thor simulator. The proposed framework achieves promising results in both change detection and change description. Moreover, EmbSCU demonstrates positive generalization results on real-world scenes without using any real-life data during training. The dataset and the code are available here.
Mariia Khan, Yue Qiu 0001, Yuren Cong, Bodo Rosenhahn, David Suter, Jumana M. Abu-Khalaf
IROS4
2024 Robust Shape Fitting for 3D Scene Abstraction
abstract
Humans perceive and construct the world as an arrangement of simple parametric models. In particular, we can often describe man-made environments using volumetric primitives such as cuboids or cylinders. Inferring these primitives is important for attaining high-level, abstract scene descriptions. Previous approaches for primitive-based abstraction estimate shape parameters directly and are only able to reproduce simple objects. In contrast, we propose a robust estimator for primitive fitting, which meaningfully abstracts complex real-world environments using cuboids. A RANSAC estimator guided by a neural network fits these primitives to a depth map. We condition the network on previously detected parts of the scene, parsing it one-by-one. To obtain cuboids from single RGB images, we additionally optimise a depth estimation CNN end-to-end. Naively minimising point-to-primitive distances leads to large or spurious cuboids occluding parts of the scene. We thus propose an improved occlusion-aware distance metric correctly handling opaque scenes. Furthermore, we present a neural network based cuboid solver which provides more parsimonious scene abstractions while also reducing inference time. The proposed algorithm does not require labour-intensive labels, such as cuboid annotations, for training. Results on the NYU Depth v2 dataset demonstrate that the proposed algorithm successfully abstracts cluttered real-world 3D scene layouts.
Florian Kluger, Eric Brachmann, Michael Ying Yang, Bodo Rosenhahn
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data
abstract
Interpreting transcriptome data is an important yet challenging aspect of bioinformatic analysis. While gene set enrichment analysis is a standard tool for interpreting regulatory changes, we utilize deep learning techniques, specifically autoencoder architectures, to learn latent variables that drive transcriptome signals. We investigate whether simple, variational autoencoder (VAE), and beta-weighted VAE are capable of learning reduced representations of transcriptomes that retain critical biological information. We propose a novel VAE that utilizes priors from biological data to direct the network to learn a representation of the transcriptome that is based on understandable biological concepts. After benchmarking five different autoencoder architectures, we found that each succeeded in reducing the transcriptomes to 50 latent dimensions, which captured enough variation for accurate reconstruction. The simple, fully connected autoencoder, performs best across the benchmarks, but lacks the characteristic of having directly interpretable latent dimensions. The beta-weighted, prior-informed VAE implementation is able to solve the benchmarking tasks, and provide semantically accurate latent features equating to biological pathways. This study opens a new direction for differential pathway analysis in transcriptomics with increased transparency and interpretability.
Bin Liu 0073, Bodo Rosenhahn, Thomas Illig, David S. DeLuca
PLoS Comput. Biol.2
2024 The voraus-AD Dataset for Anomaly Detection in Robot Applications
abstract
During the operation of industrial robots, unusual events may endanger the safety of humans and the quality of production. When collecting data to detect such cases, it is not ensured that data from all potentially occurring errors is included as unforeseeable events may happen over time. Therefore, anomaly detection (AD) delivers a practical solution, using only normal data to learn to detect unusual events. We introduce a dataset that allows training and benchmarking of anomaly detection methods for robotic applications based on machine data which will be made publicly available to the research community. As a typical robot task the dataset includes a pick-and-place application which involves movement, actions of the end effector, and interactions with the objects of the environment. Since several of the contained anomalies are not task-specific but general, evaluations on our dataset are transferable to other robotics applications as well. In addition, we present multivariate time-series flow (MVT-Flow) as a new baseline method for anomaly detection: It relies on deep-learning-based density estimation with normalizing flows, tailored to the data domain by taking its structure into account for the architecture. Our evaluation shows that MVT-Flow outperforms baselines from previous work by a large margin of 6.2% in area under receiving operator characteristic.
Jan Thieß Brockmann, Marco Rudolph, Bodo Rosenhahn, Bastian Wandt
IEEE Trans. Robotics3
2023 Deep Reinforcement Learning for Autonomous Driving using High-Level Heterogeneous Graph Representations
abstract
Graph networks have recently been used for decision making in automated driving tasks for their ability to capture a variable number of traffic participants. Current high-level graph-based approaches, however, do not model the entire road network and thus must rely on handcrafted features for vehicle-to-vehicle edges encompassing the road topology indirectly. We propose an entity-relation framework that intuitively models the road network and the traffic participants in a heterogeneous graph, representing all relevant information. Our novel architecture transforms the heterogeneous road-vehicle graph into a simpler graph of homogeneous node and edge types to allow effective training for deep reinforcement learning while introducing minimal prior knowledge. Unlike previous approaches, the vehicle-to-vehicle edges of this reduced graph are fully learnable and can therefore encode traffic rules without explicit feature design, an important step towards a holistic reinforcement learning model for automated driving. We show that our proposed method outperforms precomputed handcrafted features on intersection scenarios while also learning the semantics of right-of-way rules.
Maximilian Schier, Christoph Reinders, Bodo Rosenhahn
ICRA3
2023 Color-aware Deep Temporal Backdrop Duplex Matting System
abstract
Deep learning-based alpha matting showed tremendous improvements in recent years, yet, feature film production studios still rely on classical chroma keying including costly post-production steps. This perceived discrepancy can be explained by some missing links necessary for production which are currently not adequately addressed in the alpha matting community, in particular foreground color estimation or color spill compensation. We propose a neural network-based temporal multi-backdrop production system that combines beneficial features from chroma keying and alpha matting. Given two consecutive frames with different background colors, our one-encoder-dual-decoder network predicts foreground colors and alpha values using a patch-based overlap-blend approach. The system is able to handle imprecise backdrops, dynamic cameras, and dynamic foregrounds and has no restrictions on foreground colors. We compare our method to state-of-the-art algorithms using benchmark datasets and a video sequence captured by a demonstrator setup. We verify that a dual backdrop input is superior to the usually applied trimap-based approach. In addition, the proposed studio set is actor friendly, and produces high-quality, temporal consistent alpha and color estimations that include a superior color spill compensation.
Hendrik Hachmann, Bodo Rosenhahn
MMSys2
2023 Asymmetric Student-Teacher Networks for Industrial Anomaly Detection
abstract
Industrial defect detection is commonly addressed with anomaly detection (AD) methods where no or only incomplete data of potentially occurring defects is available. This work discovers previously unknown problems of student-teacher approaches for AD and proposes a solution, where two neural networks are trained to produce the same output for the defect-free training examples. The core assumption of student-teacher networks is that the distance between the outputs of both networks is larger for anomalies since they are absent in training. However, previous methods suffer from the similarity of student and teacher architecture, such that the distance is undesirably small for anomalies. For this reason, we propose asymmetric student-teacher networks (AST). We train a normalizing flow for density estimation as a teacher and a conventional feed-forward network as a student to trigger large distances for anomalies: The bijectivity of the normalizing flow enforces a divergence of teacher outputs for anomalies compared to normal data. Outside the training distribution the student cannot imitate this divergence due to its fundamentally different architecture. Our AST network compensates for wrongly estimated likelihoods by a normalizing flow, which was alternatively used for anomaly detection in previous work. We show that our method produces state-of-the-art results on the two currently most relevant defect detection datasets MVTec AD and MVTec 3D-AD regarding image-level anomaly detection on RGB and 3D data.
Marco Rudolph, Tom Wehrbein, Bodo Rosenhahn, Bastian Wandt
WACV3
2023 AdaCC: cumulative cost-sensitive boosting for imbalanced classification
abstract
Abstract Class imbalance poses a major challenge for machine learning as most supervised learning models might exhibit bias towards the majority class and under-perform in the minority class. Cost-sensitive learning tackles this problem by treating the classes differently, formulated typically via a user-defined fixed misclassification cost matrix provided as input to the learner. Such parameter tuning is a challenging task that requires domain knowledge and moreover, wrong adjustments might lead to overall predictive performance deterioration. In this work, we propose a novel cost-sensitive boosting approach for imbalanced data that dynamically adjusts the misclassification costs over the boosting rounds in response to model’s performance instead of using a fixed misclassification cost matrix. Our method, called AdaCC, is parameter-free as it relies on the cumulative behavior of the boosting model in order to adjust the misclassification costs for the next boosting round and comes with theoretical guarantees regarding the training error. Experiments on 27 real-world datasets from different domains with high class imbalance demonstrate the superiority of our method over 12 state-of-the-art cost-sensitive boosting approaches exhibiting consistent improvements in different measures, for instance, in the range of [0.3–28.56%] for AUC, [3.4–21.4%] for balanced accuracy, [4.8–45%] for gmean and [7.4–85.5%] for recall.
Vasileios Iosifidis, Symeon Papadopoulos, Bodo Rosenhahn, Eirini Ntoutsi
Knowl. Inf. Syst.3
2023 RelTR: Relation Transformer for Scene Graph Generation
abstract
Different objects in the same scene are more or less related to each other, but only a limited number of these relationships are noteworthy. Inspired by Detection Transformer, which excels in object detection, we view scene graph generation as a set prediction problem. In this article, we propose an end-to-end scene graph generation model Relation Transformer (RelTR), which has an encoder-decoder architecture. The encoder reasons about the visual feature context while the decoder infers a fixed-size set of triplets subject-predicate-object using different types of attention mechanisms with coupled subject and object queries. We design a set prediction loss performing the matching between the ground truth and predicted triplets for the end-to-end training. In contrast to most existing scene graph generation methods, RelTR is a one-stage method that predicts sparse scene graphs directly only using visual appearance without combining entities and labeling all possible predicates. Extensive experiments on the Visual Genome, Open Images V6, and VRD datasets demonstrate the superior performance and fast inference of our model.
Yuren Cong, Michael Ying Yang, Bodo Rosenhahn
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Wor(l)d-GAN: Toward Natural-Language-Based PCG in Minecraft
abstract
This article presents Wor(l)d-GAN, a method to perform data-driven procedural content generation via machine learning inMinecraftfrom a single example. Based on a 3-D generative adversarial network (GAN) architecture, we are able to create arbitrarily sized world snippets from a given sample. Our method applies dense representations used in natural language processing in two ways. First, we proposeblock2vecrepresentations based onword2vec. Second, we use the pretrained large language model bidirectional encoder representations from transformers (BERT) to generate representations directly from the token names. These representations make Wor(l)d-GAN independent of the number of different blocks, which can vary a lot inMinecraft, and enable the generation of larger levels. We evaluate our approach on creations from the community as well as structures generated with theMinecraftWorld Generator under several metrics. Wor(l)d-GAN enables its users to generateMinecraftworlds based on parts of their creations.
Maren Awiszus, Frederik Schubert, Bodo Rosenhahn
IEEE Trans. Games3
2022 Text to Image Generation with Semantic-Spatial Aware GAN
abstract
Text-to-image synthesis (T2I) aims to generate photorealistic images which are semantically consistent with the text descriptions. Existing methods are usually built upon conditional generative adversarial networks (GANs) and initialize an image from noise with sentence embedding, and then refine the features with fine-grained word embedding iteratively. A close inspection of their generated images reveals a major limitation: even though the generated image holistically matches the description, individual image regions or parts of somethings are often not recognizable or consistent with words in the sentence, e.g. “a white crown”. To address this problem, we propose a novel framework Semantic-Spatial Aware GAN for synthesizing images from input text. Concretely, we introduce a simple and effective Semantic-Spatial Aware block, which (1) learns semantic-adaptive transformation conditioned on text to effectively fuse text features and image features, and (2) learns a semantic mask in a weakly-supervised way that depends on the current text-image fusion process in order to guide the transformation spatially. Experiments on the challenging COCO and CUB bird datasets demonstrate the advantage of our method over the recent state-of-the-art approaches, regarding both visual fidelity and alignment with input text description. Code available at https://github.com/wtliao/text2image.
Wentong Liao, Michael Ying Yang, Bodo Rosenhahn
CVPR4
2022 LMGP: Lifted Multicut Meets Geometry Projections for Multi-Camera Multi-Object Tracking
abstract
Multi-Camera Multi-Object Tracking is currently drawing attention in the computer vision field due to its superior performance in real-world applications such as video surveillance with crowded scenes or in wide spaces. In this work, we propose a mathematically elegant multi-camera multiple object tracking approach based on a spatial-temporal lifted multicut formulation. Our model utilizes state-of-the-art tracklets produced by single-camera trackers as proposals. As these tracklets may contain ID-Switch errors, we refine them through a novel pre-clustering obtained from 3D geometry projections. As a result, we derive a better tracking graph without ID switches and more precise affinity costs for the data association phase. Tracklets are then matched to multi-camera trajectories by solving a global lifted multicut formulation that incorporates short and long-range temporal interactions on tracklets located in the same camera as well as inter-camera ones. Experimental results on the WildTrack dataset yield near-perfect performance, outperforming state-of-the-art trackers on Campus while being on par on the PETS-09 dataset. We will release our implementations at this link https://github.com/nhmduy/LMGP.
Duy M. H. Nguyen, Roberto Henschel, Bodo Rosenhahn, Daniel Sonntag, Paul Swoboda
CVPR3
2022 ChimeraMix: Image Classification on Small Datasets via Masked Feature Mixing
abstract
Deep convolutional neural networks require large amounts of labeled data samples. For many real-world applications, this is a major limitation which is commonly treated by augmentation methods. In this work, we address the problem of learning deep neural networks on small datasets. Our proposed architecture called ChimeraMix learns a data augmentation by generating compositions of instances. The generative model encodes images in pairs, combines the features guided by a mask, and creates new samples. For evaluation, all methods are trained from scratch without any additional data. Several experiments on benchmark datasets, e.g. ciFAIR-10, STL-10, and ciFAIR-100, demonstrate the superior performance of ChimeraMix compared to current state-of-the-art methods for classification on small datasets. Code is available at https://github.com/creinders/ChimeraMix.
Christoph Reinders, Frederik Schubert, Bodo Rosenhahn
IJCAI3
2022 Mixed Integer Linear Programming for Optimizing a Hopfield Network
Bodo Rosenhahn
ECML/PKDD (5)1
2022 Constrained Mean Shift Clustering
abstract
In this paper, we present Constrained Mean Shift (CMS), a novel approach for mean shift clustering under sparse supervision using cannot-link constraints. The constraints provide a guidance in constrained clustering indicating that the respective pair should not be assigned to the same cluster. Our method introduces a density-based integration of the constraints to generate individual distributions of the sampling points per cluster. We also alleviate the (in general very sensitive) mean shift bandwidth parameter by proposing an adaptive bandwidth adjustment which is especially useful for clustering imbalanced data sets. Several experiments show that our approach achieves better performance compared to state-of-the-art methods both clustering synthetic data sets as well as clustering encoded features of real-world image data sets.
Maximilian Schier, Christoph Reinders, Bodo Rosenhahn
SDM3
2022 Fully Convolutional Cross-Scale-Flows for Image-based Defect Detection
abstract
In industrial manufacturing processes, errors frequently occur at unpredictable times and in unknown manifestations. We tackle the problem of automatic defect detection without requiring any image samples of defective parts. Recent works model the distribution of defect-free image data, using either strong statistical priors or overly simplified data representations. In contrast, our approach handles fine-grained representations incorporating the global and local image context while flexibly estimating the density. To this end, we propose a novel fully convolutional cross-scale normalizing flow (CS-Flow) that jointly processes multiple feature maps of different scales. Using normalizing flows to assign meaningful likelihoods to input samples allows for efficient defect detection on image-level. Moreover, due to the preserved spatial arrangement the latent space of the normalizing flow is interpretable which enables to localize defective regions in the image. Our work sets a new state-of-the-art in image-level defect detection on the benchmark datasets Magnetic Tile Defects and MVTec AD showing a 100% AUROC on 4 out of 15 classes.
Marco Rudolph, Tom Wehrbein, Bodo Rosenhahn, Bastian Wandt
WACV3
2022 TOAD-GAN: A Flexible Framework for Few-Shot Level Generation in Token-Based Games
abstract
This work presentsToken-basedOne-shotArbitraryDimensionGenerativeAdversarialNetwork (TOAD-GAN), a novel procedural content generation algorithm that generates token-based video game levels from only one example. We show that the created levels can be of arbitrary size, and the patterns of the training levels are well captured. The method can be extended with user interaction during the generation process to achieve certain token layouts and interpretations of the same base level by different generators. Our method is further evaluated with an extensive ablation study and level similarity metrics on theSuper Mario Bros.benchmark. Finally, we extend our method to mix the style of multiple input levels, turning it into a framework for few-shot level generation.
Frederik Schubert, Maren Awiszus, Bodo Rosenhahn
IEEE Trans. Games3
2021 World-GAN: a Generative Model for Minecraft Worlds
abstract
This work introduces World-GAN, the first method to perform data-driven Procedural Content Generation via Machine Learning in Minecraft from a single example. Based on a 3D Generative Adversarial Network (GAN) architecture, we are able to create arbitrarily sized world snippets from a given sample. We evaluate our approach on creations from the community as well as structures generated with the Minecraft World Generator. Our method is motivated by the dense representations used in Natural Language Processing (NLP) introduced with word2vec [1]. The proposed block2vec representations make World-GAN independent from the number of different blocks, which can vary a lot in Minecraft, and enable the generation of larger levels. Finally, we demonstrate that changing this new representation space allows us to change the generated style of an already trained generator. World-GAN enables its users to generate Minecraft worlds based on parts of their creations.
Maren Awiszus, Frederik Schubert, Bodo Rosenhahn
CoG3
2021 Context-Aware Layout to Image Generation With Enhanced Object Appearance
abstract
A layout to image (L2I) generation model aims to generate a complicated image containing multiple objects (things) against natural background (stuff), conditioned on a given layout. Built upon the recent advances in generative adversarial networks (GANs), existing L2I models have made great progress. However, a close inspection of their generated images reveals two major limitations: (1) the object-to-object as well as object-to-stuff relations are often broken and (2) each object’s appearance is typically distorted lacking the key defining characteristics associated with the object class. We argue that these are caused by the lack of context-aware object and stuff feature encoding in their generators, and location-sensitive appearance representation in their discriminators. To address these limitations, two new modules are proposed in this work. First, a context-aware feature transformation module is introduced in the generator to ensure that the generated feature encoding of either object or stuff is aware of other coexisting objects/stuff in the scene. Second, instead of feeding location-insensitive image features to the discriminator, we use the Gram matrix computed from the feature maps of the generated object images to preserve location-sensitive information, resulting in much enhanced object appearance. Extensive experiments show that the proposed method achieves state-of-the-art performance on the COCO-Thing-Stuff and Visual Genome benchmarks. Code available at: https://github.com/wtliao/layout2img.
Sen He 0001, Wentong Liao, Michael Ying Yang, Yongxin Yang, Yi-Zhe Song, Bodo Rosenhahn, Tao Xiang 0002
CVPR6
2021 Cuboids Revisited: Learning Robust 3D Shape Fitting to Single RGB Images
abstract
Humans perceive and construct the surrounding world as an arrangement of simple parametric models. In particular, man-made environments commonly consist of volumetric primitives such as cuboids or cylinders. Inferring these primitives is an important step to attain high-level, abstract scene descriptions. Previous approaches directly estimate shape parameters from a 2D or 3D input, and are only able to reproduce simple objects, yet unable to accurately parse more complex 3D scenes. In contrast, we propose a robust estimator for primitive fitting, which can meaningfully abstract real-world environments using cuboids. A RANSAC estimator guided by a neural network fits these primitives to 3D features, such as a depth map. We condition the network on previously detected parts of the scene, thus parsing it one-by-one. To obtain 3D features from a single RGB image, we additionally optimise a feature extraction CNN in an end-to-end manner. However, naively minimising point-to-primitive distances leads to large or spurious cuboids occluding parts of the scene behind. We thus propose an occlusion-aware distance metric correctly handling opaque scenes. The proposed algorithm does not require labour-intensive labels, such as cuboid annotations, for training. Results on the challenging NYU Depth v2 dataset demonstrate that the proposed algorithm successfully abstracts cluttered real-world 3D scene layouts.
Florian Kluger, Hanno Ackermann, Eric Brachmann, Michael Ying Yang, Bodo Rosenhahn
CVPR5
2021 CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the Wild
abstract
Human pose estimation from single images is a challenging problem in computer vision that requires large amounts of labeled training data to be solved accurately. Unfortunately, for many human activities (e.g. outdoor sports) such training data does not exist and is hard or even impossible to acquire with traditional motion capture systems. We propose a self-supervised approach that learns a single image 3D pose estimator from unlabeled multi-view data. To this end, we exploit multi-view consistency constraints to disentangle the observed 2D pose into the underlying 3D pose and camera rotation. In contrast to most existing methods, we do not require calibrated cameras and can therefore learn from moving cameras. Nevertheless, in the case of a static camera setup, we present an optional extension to include constant relative camera rotations over multiple views into our framework. Key to the success are new, unbiased reconstruction objectives that mix information across views and training samples. The proposed approach is evaluated on two benchmark datasets (Human3.6M and MPII-INF-3DHP) and on the in-the-wild SkiPose dataset.
Bastian Wandt, Marco Rudolph, Petrissa Zell, Helge Rhodin, Bodo Rosenhahn
CVPR5
2021 Spatial-Temporal Transformer for Dynamic Scene Graph Generation
abstract
Dynamic scene graph generation aims at generating a scene graph of the given video. Compared to the task of scene graph generation from images, it is more challenging because of the dynamic relationships between objects and the temporal dependencies between frames allowing for a richer semantic interpretation. In this paper, we propose Spatial-temporal Transformer (STTran), a neural network that consists of two core modules: (1) a spatial encoder that takes an input frame to extract spatial context and reason about the visual relationships within a frame, and (2) a temporal decoder which takes the output of the spatial encoder as input in order to capture the temporal dependencies between frames and infer the dynamic relationships. Furthermore, STTran is flexible to take varying lengths of videos as input without clipping, which is especially important for long videos. Our method is validated on the benchmark dataset Action Genome (AG). The experimental results demonstrate the superior performance of our method in terms of dynamic scene graphs. Moreover, a set of ablative studies is conducted and the effect of each proposed module is justified. Code available at: https://github.com/yrcong/STTran.
Yuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn, Michael Ying Yang
ICCV4
2021 Disentangled Lifespan Face Synthesis
abstract
A lifespan face synthesis (LFS) model aims to generate a set of photo-realistic face images of a person’s whole life, given only one snapshot as reference. The generated face image given a target age code is expected to be age-sensitive reflected by bio-plausible transformations of shape and texture, while being identity preserving. This is extremely challenging because the shape and texture characteristics of a face undergo separate and highly nonlinear transformations w.r.t. age. Most recent LFS models are based on generative adversarial networks (GANs) whereby age code conditional transformations are applied to a latent face representation. They benefit greatly from the recent advancements of GANs. However, without explicitly disentangling their latent representations into the texture, shape and identity factors, they are fundamentally limited in modeling the nonlinear age-related transformation on texture and shape whilst preserving identity. In this work, a novel LFS model is proposed to disentangle the key face characteristics including shape, texture and identity so that the unique shape and texture age transformations can be modeled effectively. This is achieved by extracting shape, texture and identity features separately from an encoder. Critically, two transformation modules, one conditional convolution based and the other channel attention based, are designed for modeling the nonlinear shape and texture feature transformations respectively. This is to accommodate their rather distinct aging processes and ensure that our synthesized images are both age-sensitive and identity preserving. Extensive experiments show that our LFS model is clearly superior to the state-of-the-art alternatives. Codes and demo are available on our project website: https://senhe.github.io/projects/iccv_2021_lifespan_face.
Sen He 0001, Wentong Liao, Michael Ying Yang, Yi-Zhe Song, Bodo Rosenhahn, Tao Xiang 0002
ICCV5
2021 Making Higher Order MOT Scalable: An Efficient Approximate Solver for Lifted Disjoint Paths
abstract
We present an efficient approximate message passing solver for the lifted disjoint paths problem (LDP), a natural but NP-hard model for multiple object tracking (MOT). Our tracker scales to very large instances that come from long and crowded MOT sequences. Our approximate solver enables us to process the MOT15/16/17 benchmarks without sacrificing solution quality and allows for solving MOT20, which has been out of reach up to now for LDP solvers due to its size and complexity. On all these four standard MOT benchmarks we achieve performance comparable or better than current state-of-the-art methods including a tracker based on an optimal LDP solver.
Andrea Hornáková, Timo Kaiser, Paul Swoboda, Michal Rolínek, Bodo Rosenhahn, Roberto Henschel
ICCV5
2021 Probabilistic Monocular 3D Human Pose Estimation with Normalizing Flows
abstract
3D human pose estimation from monocular images is a highly ill-posed problem due to depth ambiguities and occlusions. Nonetheless, most existing works ignore these ambiguities and only estimate a single solution. In contrast, we generate a diverse set of hypotheses that represents the full posterior distribution of feasible 3D poses. To this end, we propose a normalizing flow based method that exploits the deterministic 3D-to-2D mapping to solve the ambiguous inverse 2D-to-3D problem. Additionally, uncertain detections and occlusions are effectively modeled by incorporating uncertainty information of the 2D detector as condition. Further keys to success are a learned 3D pose prior and a generalization of the best-of-M loss. We evaluate our approach on the two benchmark datasets Human3.6M and MPI-INF-3DHP, outperforming all comparable methods in most metrics. The implementation is available on GitHub1.
Tom Wehrbein, Marco Rudolph, Bodo Rosenhahn, Bastian Wandt
ICCV3
2021 Exploring Dynamic Context for Multi-path Trajectory Prediction
abstract
To accurately predict future positions of different agents in traffic scenarios is crucial for safely deploying intelligent autonomous systems in the real-world environment. However, it remains a challenge due to the behavior of a target agent being affected by other agents dynamically and there being more than one socially possible paths the agent could take. In this paper, we propose a novel framework, named Dynamic Context Encoder Network (DCENet). In our framework, first, the spatial context between agents is explored by using self-attention architectures. Then, the two-stream encoders are trained to learn temporal context between steps by taking the respective observed trajectories and the extracted dynamic spatial context as input. The spatial-temporal context is encoded into a latent space using a Conditional Variational Auto-Encoder (CVAE) module. Finally, a set of future trajectories for each agent is predicted conditioned on the learned spatial-temporal context by sampling from the latent space, repeatedly. DCENet is evaluated on one of the most popular challenging benchmarks for trajectory forecasting Trajnet and reports a new state-of-the-art performance. It also demonstrates superior performance evaluated on the benchmark inD for mixed traffic at intersections. A series of ablation studies is conducted to validate the effectiveness of each proposed module. Our code is available at https://github.com/wtliao/DCENet.
Hao Cheng 0008, Wentong Liao, Xuejiao Tang, Michael Ying Yang, Monika Sester, Bodo Rosenhahn
ICRA6
2021 Same Same But DifferNet: Semi-Supervised Defect Detection with Normalizing Flows
abstract
The detection of manufacturing errors is crucial in fabrication processes to ensure product quality and safety standards. Since many defects occur very rarely and their characteristics are mostly unknown a priori, their detection is still an open research question. To this end, we propose DifferNet: It leverages the descriptiveness of features extracted by convolutional neural networks to estimate their density using normalizing flows. Normalizing flows are well-suited to deal with low dimensional data distributions. However, they struggle with the high dimensionality of images. Therefore, we employ a multi-scale feature extractor which enables the normalizing flow to assign meaningful likelihoods to the images. Based on these likelihoods we develop a scoring function that indicates defects. Moreover, propagating the score back to the image enables pixel-wise localization. To achieve a high robustness and performance we exploit multiple transformations in training and evaluation. In contrast to most other methods, ours does not require a large number of training samples and performs well with as low as 16 images. We demonstrate the superior performance over existing approaches on the challenging and newly proposed MVTec AD [4] and Magnetic Tile Defects [14] datasets.
Marco Rudolph, Bastian Wandt, Bodo Rosenhahn
WACV3
2020 Image Captioning Through Image Transformer
Sen He 0001, Wentong Liao, Hamed Rezazadegan Tavakoli, Michael Ying Yang, Bodo Rosenhahn, Nicolas Pugeault
ACCV (4)5
2020 CONSAC: Robust Multi-Model Fitting by Conditional Sample Consensus
abstract
We present a robust estimator for fitting multiple parametric models of the same form to noisy measurements. Applications include finding multiple vanishing points in man-made scenes, fitting planes to architectural imagery, or estimating multiple rigid motions within the same sequence. In contrast to previous works, which resorted to hand-crafted search strategies for multiple model detection, we learn the search strategy from data. A neural network conditioned on previously detected models guides a RANSAC estimator to different subsets of all measurements, thereby finding model instances one after another. We train our method supervised, as well as, self-supervised. For supervised training of the search strategy, we contribute a new dataset for vanishing point estimation. Leveraging this dataset, the proposed algorithm is superior with respect to other robust estimators, as well as, to designated vanishing point estimation algorithms. For self-supervised learning of the search, we evaluate the proposed algorithm on multi-homography estimation and demonstrate an accuracy that is superior to state-of-the-art methods.
Florian Kluger, Eric Brachmann, Hanno Ackermann, Carsten Rother, Michael Ying Yang, Bodo Rosenhahn
CVPR6
2020 FairNN - Conjoint Learning of Fair Representations for Fair Decisions
Tongxin Hu, Vasileios Iosifidis, Wentong Liao, Michael Ying Yang, Eirini Ntoutsi, Bodo Rosenhahn
DS7
2020 NODIS: Neural Ordinary Differential Scene Understanding
Yuren Cong, Hanno Ackermann, Wentong Liao, Michael Ying Yang, Bodo Rosenhahn
ECCV (20)5
2020 Weakly-Supervised Learning of Human Dynamics
Petrissa Zell, Bodo Rosenhahn, Bastian Wandt
ECCV (26)2
2020 Lifted Disjoint Paths with Application in Multiple Object Tracking
abstract
We present an extension to the disjoint paths problem in which additional lifted edges are introduced to provide path connectivity priors. We call the resulting optimization problem the lifted disjoint paths problem. We show that this problem is NP-hard by reduction from integer multicommodity flow and 3-SAT. To enable practical global optimization, we propose several classes of linear inequalities that produce a high-quality LP-relaxation. Additionally, we propose efficient cutting plane algorithms for separating the proposed linear inequalities. The lifted disjoint path problem is a natural model for multiple object tracking and allows an elegant mathematical formulation for long range temporal interactions. Lifted edges help to prevent id switches and to re-identify persons. Our lifted disjoint paths tracker achieves nearly optimal assignments with respect to input detections. As a consequence, it leads on all three main benchmarks of the MOT challenge, improving significantly over state-of-the-art.
Andrea Hornáková, Roberto Henschel, Bodo Rosenhahn, Paul Swoboda
ICML3
2020 Temporally Consistent Horizon Lines
abstract
The horizon line is an important geometric feature for many image processing and scene understanding tasks in computer vision. For instance, in navigation of autonomous vehicles or driver assistance, it can be used to improve 3D reconstruction as well as for semantic interpretation of dynamic environments. While both algorithms and datasets exist for single images, the problem of horizon line estimation from video sequences has not gained attention. In this paper, we show how convolutional neural networks are able to utilise the temporal consistency imposed by video sequences in order to increase the accuracy and reduce the variance of horizon line estimates. A novel CNN architecture with an improved residual convolutional LSTM is presented for temporally consistent horizon line estimation. We propose an adaptive loss function that ensures stable training as well as accurate results. Furthermore, we introduce an extension of the KITTI dataset which contains precise horizon line labels for 43699 images across 72 video sequences. A comprehensive evaluation shows that the proposed approach consistently achieves superior performance compared with existing methods.
Florian Kluger, Hanno Ackermann, Michael Ying Yang, Bodo Rosenhahn
ICRA4
2020 Learning inverse dynamics for human locomotion analysis
Petrissa Zell, Bodo Rosenhahn
Neural Comput. Appl.2
2020 Accurate Long-Term Multiple People Tracking Using Video and Body-Worn IMUs
abstract
Most modern approaches for video-based multiple people tracking rely on human appearance to exploit similarities between person detections. Consequently, tracking accuracy degrades if this kind of information is not discriminative or if people change apparel. In contrast, we present a method to fuse video information with additional motion signals from body-worn inertial measurement units (IMUs). In particular, we propose a neural network to relate person detections with IMU orientations, and formulate a graph labeling problem to obtain a tracking solution that is globally consistent with the video and inertial recordings. The fusion of visual and inertial cues provides several advantages. The association of detection boxes in the video and IMU devices is based on motion, which is independent of a person's outward appearance. Furthermore, inertial sensors provide motion information irrespective of visual occlusions. Hence, once detections in the video are associated with an IMU device, intermediate positions can be reconstructed from corresponding inertial sensor data, which would be unstable using video only. Since no dataset exists for this new setting, we release a dataset of challenging tracking sequences, containing video and IMU recordings together with ground-truth annotations. We evaluate our approach on our new dataset, achieving an average IDF1 score of 91.2%. The proposed method is applicable to any situation that allows one to equip people with inertial sensors.
Roberto Henschel, Timo von Marcard, Bodo Rosenhahn
IEEE Trans. Image Process.3
2019 RepNet: Weakly Supervised Training of an Adversarial Reprojection Network for 3D Human Pose Estimation
abstract
This paper addresses the problem of 3D human pose estimation from single images. While for a long time human skeletons were parameterized and fitted to the observation by satisfying a reprojection error, nowadays researchers directly use neural networks to infer the 3D pose from the observations. However, most of these approaches ignore the fact that a reprojection constraint has to be satisfied and are sensitive to overfitting. We tackle the overfitting problem by ignoring 2D to 3D correspondences. This efficiently avoids a simple memorization of the training data and allows for a weakly supervised training. One part of the proposed reprojection network (RepNet) learns a mapping from a distribution of 2D poses to a distribution of 3D poses using an adversarial training approach. Another part of the network estimates the camera. This allows for the definition of a network layer that performs the reprojection of the estimated 3D pose back to 2D which results in a reprojection loss function. Our experiments show that RepNet generalizes well to unknown data and outperforms state-of-the-art methods when applied to unseen data. Moreover, our implementation runs in real-time on a standard desktop PC.
Bastian Wandt, Bodo Rosenhahn
CVPR2
2019 Machine learning for measurement-based bandwidth estimation
Sukhpreet Kaur Khangura, Markus Fidler, Bodo Rosenhahn
Comput. Commun.3
2019 Occlusion-Aware Method for Temporally Consistent Superpixels
abstract
A wide variety of computer vision applications rely on superpixel or supervoxel algorithms as a preprocessing step. This underlines the overall importance that these approaches have gained in recent years. However, most methods show a lack of temporal consistency or fail in producing temporally stable superpixels. In this paper, we present an approach to generate temporally consistent superpixels for video content. Our method is formulated as a contour-evolving expectation-maximization framework, which utilizes an efficient label propagation scheme to encourage the preservation of superpixel shapes and their relative positioning over time. By explicitly detecting the occlusion of superpixels and the disocclusion of new image regions, our framework is able to terminate and create superpixels whose corresponding image region becomes hidden or newly appears. Additionally, the occluded parts of superpixels are incorporated in the further optimization. This increases the compliance of the superpixel flow with the optical flow present in the scene. Using established benchmark suites, we show the performance of our approach in comparison to state-of-the-art supervoxel and superpixel algorithms for video content.
Matthias Reso, Jörn Jachalsky, Bodo Rosenhahn, Jörn Ostermann
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Region-based Cycle-Consistent Data Augmentation for Object Detection
abstract
Roads constitute a major part of the lives of everybody. Heavy use, for instance by cars and especially trucks, and even soil movement lead to visible damages. While major roads are regularly inspected, smaller roads often lack attention. It is therefore of great interest to have camera-based systems which can automatically detect and even classify damages.This report presents a system developed by the authors as part of the Road Damage Detection and Classification Challenge at the 2018 IEEE Big Data Cup [1]. Further contributions made here are techniques to augment the small set of training data. As a major contribution we also propose refinements to the dataset and evaluation metric to improve the challenge.
Florian Kluger, Christoph Reinders, Kevin Raetz, Philipp Schelske, Bastian Wandt, Hanno Ackermann, Bodo Rosenhahn
IEEE BigData7
2018 Detail-Aware Image Decomposition for an HEVC-Based Texture Synthesis Framework
abstract
Modern video coding standards like High Efficiency Video Coding (HEVC) provide superior coding efficiency. However, this does not state true for complex and hard to predict textures which require high bit rates to achieve a high quality. To overcome this limitation of HEVC, texture synthesis frameworks were proposed in previous works. However, these frameworks only result in good reconstruction quality if the decomposition into synthesizable and non-synthesizable regions is either known or trivial. The frameworks fail for more challenging content, e.g. for content with fine non-synthesizable details within synthesizable regions. To enable texture synthesis-based video coding with high quality for this content, we propose sophisticated detail-aware decomposition techniques in this paper. These techniques are based on an initial coarse segmentation step followed by a refinement step that detects even small differences in the previously segmented region. With this new approach, we are able to achieve average luma BD-rate gains of 13.77% over HEVC and 3.03% over the closest related work from the literature. Furthermore, the considerably improved visual quality in addition to the bit rate savings is confirmed by comprehensive subjective tests.
Bastian Wandt, Thorsten Laude, Bodo Rosenhahn, Jörn Ostermann
DCC3
2018 Recovering Accurate 3D Human Pose in the Wild Using IMUs and a Moving Camera
Timo von Marcard, Roberto Henschel, Michael J. Black, Bodo Rosenhahn, Gerard Pons-Moll
ECCV (10)4
2018 Deep Learning for Vehicle Detection in Aerial Images
abstract
The detection of vehicles in aerial images is widely applied in many domains. In this paper, we propose a novel double focal loss convolutional neural network framework (DFL-CNN). In the proposed framework, the skip connection is used in the CNN structure to enhance the feature learning. Also, the focal loss function is used to substitute for conventional cross entropy loss function in both of the region proposed network and the final classifier. We further introduce the first large-scale vehicle detection dataset ITCVD with ground truth annotations for all the vehicles in the scene. The experimental results show that our DFL-CNN outperforms the baselines on vehicle detection.
Michael Ying Yang, Wentong Liao, Xinbo Li, Bodo Rosenhahn
ICIP4
2018 Object Recognition from very few Training Examples for Enhancing Bicycle Maps
abstract
In recent years, data-driven methods have shown great success for extracting information about the infrastructure in urban areas. These algorithms are usually trained on large datasets consisting of thousands or millions of labeled training examples. While large datasets have been published regarding cars, for cyclists very few labeled data is available although appearance, point of view, and positioning of even relevant objects differ. Unfortunately, labeling data is costly and requires a huge amount of work. In this paper, we thus address the problem of learning with very few labels. The aim is to recognize particular traffic signs in crowdsourced data to collect information which is of interest to cyclists. We propose a system for object recognition that is trained with only 15 examples per class on average. To achieve this, we combine the advantages of convolutional neural networks and random forests to learn a patch-wise classifier. In the next step, we map the random forest to a neural network and transform the classifier to a fully convolutional network. Thereby, the processing of full images is significantly accelerated and bounding boxes can be predicted. Finally, we integrate data of the Global Positioning System (GPS) to localize the predictions on the map. In comparison to Faster R-CNN and other networks for object recognition or algorithms for transfer learning, we considerably reduce the required amount of labeled data. We demonstrate good performance on the recognition of traffic signs for cyclists as well as their localization in maps.
Christoph Reinders, Hanno Ackermann, Michael Ying Yang, Bodo Rosenhahn
Intelligent Vehicles Symposium4
2018 Physical High Dynamic Range Imaging with Conventional Sensors
abstract
This paper aims at simplified high dynamic range (HDR) image generation with non-modified, conventional camera sensors. One typical HDR approach is exposure bracketing, e.g. with varying shutter speeds. It requires to capture the same scene multiple times at different exposure times. These pictures are then merged into a single HDR picture which typically is converted back to an 8-bit image by using tone-mapping. Existing works on HDR imaging focus on image merging and tone mapping whereas we aim at simplified image acquisition. The proposed algorithm can be used in consumer-level cameras without hardware modifications at sensor level. Based on intermediate samplings of each sensor element during the total (pre-defined) exposure time, we extrapolate the luminance of sensor elements which are saturated after the total exposure time. Compared to existing HDR approaches which typically require three different images with carefully determined exposure times, we only take one image at the longest exposure time. The shortened total time between start and end of image acquisition can reduce ghosting artifacts. The experimental evaluation demonstrates the effectiveness of the algorithm.
Holger Meuel, Hanno Ackermann, Bodo Rosenhahn, Jörn Ostermann
PCS3
2018 Extending HEVC with a Texture Synthesis Framework using Detail-aware Image Decomposition
abstract
In recent years, there has been a tremendous improvement in video coding algorithms. This improvement resulted in 2013 in the standardization of the first version of High Efficiency Video Coding (HEVC) which now forms the state-of-theart with superior coding efficiency. Nevertheless, the development of video coding algorithms did not stop as HEVC still has its limitations. Especially for complex textures HEVC reveals one of its limitations. As these textures are hard to predict, very high bit rates are required to achieve a high quality. Texture synthesis was proposed as solution for this limitation in previous works. However, previous texture synthesis frameworks only prevailed if the decomposition into synthesizable and non-synthesizable regions was either known or very easy. In this paper, we address this scenario with a texture synthesis framework based on detail-aware image decomposition techniques. Our techniques are based on a multiple-steps coarse-to-fine approach in which an initial decomposition is refined with awareness for small details. The efficiency of our approach is evaluated objectively and subjectively: BD-rate gains of up to 28.81% over HEVC and up to 12.75% over the closest related work were achieved. Our subjective tests indicate an improved visual quality in addition to the bit rate savings.
Bastian Wandt, Thorsten Laude, Bodo Rosenhahn, Jörn Ostermann
PCS3
2018 3D braid guide hair reconstruction using electroluminescent wires
Hendrik Hachmann, Maren Awiszus, Bodo Rosenhahn
Vis. Comput.3
2017 Extending HEVC using texture synthesis
abstract
The High Efficiency Video Coding (HEVC) standard provides superior coding efficiency compared to its predecessors. Nevertheless, the encoding of complex and thus hardly to predict textures either requires high bit rates or results in low quality of the reconstructed signal. To compensate for this limitation of HEVC, we propose a sophisticated texture synthesis framework which solves multiple lacks of previous texture synthesis approaches. By easing the bit rate cost for synthesizable regions and reallocating the freed bit rate resources to non-synthesizable regions, for high-value soccer content we are able to achieve average BD-rate gains of 21.9% for all-intra, 17.6% for low delay, and 16.3% for random access, respectively, while maintaining the same objective quality for the latter. Subjective tests for the synthesizable regions confirm the objectively measured convincing results. The general applicability of our method is confirmed for other types of content.
Bastian Wandt, Thorsten Laude, Yiqun Liu 0003, Bodo Rosenhahn, Jörn Ostermann
VCIP4
2017 Global Consistency Priors for Joint Part-Based Object Tracking and Image Segmentation
abstract
Tracking of previously unseen, articulated objects is an active research area. Recently, Deformable Parts Models (DPMs) have been used to improve the online tracking performance for bounding-box trackers. In this paper, we extend the DPM with global priors which enforce consistency with foreground/background segmentation cues. We propose a Dual Decomposition approach and show how to efficiently solve the high-order coupling constraints as a feasible sub-problem. The proposed approach is evaluated on the VOT online tracking benchmark, outperforming the baseline in both tracking accuracy and robustness. We further show that in presence of stable image segmentation cues, the flexibility of a generic DPM generated from a single reference frame can be improved by introducing the concept of part visibility, the visibility-aware DPM (VDPM). This allows for fine-grained articulated object tracking using an automatically generated DPM from a single template image.
Oliver Müller 0003, Bodo Rosenhahn
WACV2
2017 Sparse Inertial Poser: Automatic 3D Human Pose Estimation from Sparse IMUs
abstract
We address the problem of making human motion capture in the wild more practical by using a small set of inertial sensors attached to the body. Since the problem is heavily under-constrained, previous methods either use a large number of sensors, which is intrusive, or they require additional video input. We take a different approach and constrain the problem by: (i) making use of a realistic statistical body model that includes anthropometric constraints and (ii) using a joint optimization framework to fit the model to orientation and acceleration measurements over multiple frames. The resulting tracker Sparse Inertial Poser (SIP) enables motion capture using only 6 sensors (attached to the wrists, lower legs, back and head) and works for arbitrary human motions. Experiments on the recently released TNT15 dataset show that, using the same number of sensors, SIP achieves higher accuracy than the dataset baseline without using any video data. We further demonstrate the effectiveness of SIP on newly recorded challenging motions in outdoor scenarios such as climbing or jumping over a wall.
Timo von Marcard, Bodo Rosenhahn, Michael J. Black, Gerard Pons-Moll
Comput. Graph. Forum2
2016 Exploiting View-Specific Appearance Similarities Across Classes for Zero-Shot Pose Prediction: A Metric Learning Approach
abstract
Viewpoint estimation, especially in case of multiple object classes, remains an important and challenging problem. First, objects under different views undergo extreme appearance variations, often making within-class variance larger than between-class variance. Second, obtaining precise ground truth for real-world images, necessary for training supervised viewpoint estimation models, is extremely difficult and time consuming. As a result, annotated data is often available only for a limited number of classes. Hence it is desirable to share viewpoint information across classes. Additional complexity arises from unaligned pose labels between classes, i.e. a side view of a car might look more like a frontal view of a toaster, than its side view. To address these problems, we propose a metric learning approach for joint class prediction and pose estimation. Our approach allows to circumvent the problem of viewpoint alignment across multiple classes, and does not require dense viewpoint labels. Moreover, we show, that the learned metric generalizes to new classes, for which the pose labels are not available, and therefore makes it possible to use only partially annotated training sets, relying on the intrinsic similarities in the viewpoint manifolds. We evaluate our approach on two challenging multi-class datasets, 3DObjects and PASCAL3D+.
Alina Kuznetsova, Sung Ju Hwang, Bodo Rosenhahn, Leonid Sigal
AAAI3
2016 Multicamera Calibration from Visible and Mirrored Epipoles
abstract
Multicamera rigs are used in a large number of 3D Vision applications, such as 3D modeling, motion capture or telepresence and a robust calibration is of utmost importance in order to achieve a high accuracy results. In many practical configurations the cameras in a rig are arranged in such a way, that they can observe each other, in other words a number of epipoles correspond to the real image points. In this paper we propose a solution for the automatic recovery of the external calibration of a multicamera system by enforcing only simple geometrical constraints, arising from the epipole visibility, without using any calibration object, such as checkerboards, laser pointers or similar. Additionally, we introduce an extension of the method that handles the case of epipoles being visible in the reflection of a planar mirror, which makes the algorithm suitable for the calibration of any multicamera system, irrespective of the number of cameras and their actual mutual visibility, and furthermore we remark that it requires only one or a few images per camera and therefore features a high speed and usability. We produce an evidence of the algorithm effectiveness by presenting a wide set of tests performed on synthetic as well as real datasets and we compare the results with those obtained using a traditional LED-based algorithm. The real datasets have been captured using a multicamera Virtual Reality (VR) rig and a spherical dome configuration for 3D reconstruction.
Andrey Bushnevskiy, Lorenzo Sorgi, Bodo Rosenhahn
CVPR3
2016 PSyCo: Manifold Span Reduction for Super Resolution
abstract
The main challenge in Super Resolution (SR) is to discover the mapping between the low-and high-resolution manifolds of image patches, a complex ill-posed problem which has recently been addressed through piecewise linear regression with promising results. In this paper we present a novel regression-based SR algorithm that benefits from an extended knowledge of the structure of both manifolds. We propose a transform that collapses the 16 variations induced from the dihedral group of transforms (i.e. rotations, vertical and horizontal reflections) and antipodality (i.e. diametrically opposed points in the unitary sphere) into a single primitive. The key idea of our transform is to study the different dihedral elements as a group of symmetries within the high-dimensional manifold. We obtain the respective set of mirror-symmetry axes by means of a frequency analysis of the dihedral elements, and we use them to collapse the redundant variability through a modified symmetry distance. The experimental validation of our algorithm shows the effectiveness of our approach, which obtains competitive quality with a dictionary of as little as 32 atoms (reducing other methods' dictionaries by at least a factor of 32) and further pushing the state-of-the-art with a 1024 atoms dictionary.
Eduardo Pérez-Pellitero, Jordi Salvador, Javier Ruiz Hidalgo, Bodo Rosenhahn
CVPR4
2016 Multimode camera calibration
abstract
Camera calibration denotes the task of estimating the projective mapping between 3D world and the camera image plane. Most of the modern cameras have multiple image and video acquisition modes, which differ in image resolution, field of view or aspect ratio. Each mode should be treated as an independent device, and therefore independently calibrated. This straightforward solution implies the acquisition of the calibration dataset and execution of the calibration routine for each of the camera modes, a time-consuming and error-prone task, especially in a multi-camera scenario. In this paper we propose a multimode camera calibration approach that noticeably relieves this task. We show how the calibration model of one camera mode can be easily transferred to the other modes without the need of multiple calibrations. The evaluation test performed on off-the-shelf consumer cameras shows that the proposed method not only reduces the required amount of user interaction, but also allows for the calibration accuracy improvement.
Andrey Bushnevskiy, Lorenzo Sorgi, Bodo Rosenhahn
ICIP3
2016 Bi-layer dictionary learning for remote sensing image classification
abstract
With the widely application of high-resolution remote sensing images, its classification has attracted a lot of attention. Usually, some different categories share common patterns, which make these categories look similar. This makes the classification of such categories a challenging task. In this paper, we propose a novel dictionary learning based bilayer classification algorithm to solve this problem. Using SIFT descriptor, instead of directly classifying an image, we separate the classification in two steps. In the first step, the similar categories are clustered to be as a new category for the first classification layer. In this step, the inter-class variation are maximized. The second layer is designed to classify the similar categories clustered in the same group. Experimental results show the superiority of our method compared to the state-of-the-art methods using UCMerced LandUse dataset.
Michael Ying Yang, Saif Dawood Salman Al-Shaikhli, Yanpeng Cao, Bodo Rosenhahn
IGARSS5
2016 Moving object tracking for aerial video coding using linear motion prediction and block matching
abstract
Region of Interest (ROI) coding is a common method for data reduction in scenarios where bandwidth is crucial like in aerial video surveillance from Unmanned Aerial Vehicles (UAVs). In order to save bits, non-ROI areas are typically reduced in quality or not transmitted at all and thus, an accurate ROI classification is mandatory. Moving objects (MOs) are often considered as ROIs and consequently have to be accurately detected onboard. However, common detection approaches either rely on computationally demanding processing which is not available at small UAVs with only limited energy, are model based or cannot provide a sufficient detection precision. While not detected MOs lead to a degraded representation at the decoder, erroneously detected MOs lead to an unnecessary high bit rate. We tackle all these issues utilizing an efficient object proposal computation. Based on a dual-threshold strategy applied to image differences, we propose a linear prediction-supported block matcher. Compared to a simple thresholding approach, it shows superior performance and is robust to threshold tuning. By integrating superpixels into the framework, we further recover the complete shape of the MOs. Finally, an efficient tracking-by-detection system is employed to produce accurate detections from the proposals, thereby recovering missed MOs and denying wrong proposals, making the coding more efficient. We achieve an improved detection precision of up to 76 % compared to a simple difference image-based approach. By using a general ROI coding framework we reduce the bit rate of our test set by 70 % compared to common HEVC.
Holger Meuel, Luis Angerstein, Roberto Henschel, Bodo Rosenhahn, Jörn Ostermann
PCS4
2016 Half hypersphere confinement for piecewise linear regression
abstract
Recent research in piecewise linear regression for Super-Resolution has shown the positive impact of training regressors with densely populated clusters whose datapoints are tight in the Euclidean space. In this paper we further research how to improve the locality condition during the training of regressors and how to better select them during testing time. We study the characteristics of the metrics best suited for the piecewise regression algorithms, in which comparisons are usually made between normalized vectors that lie on the unitary hypersphere. Even though Euclidean distance has been widely used for this purpose, it is suboptimal since it does not handle antipodal points (i.e. diametrically opposite points) properly, as vectors with same module and angle but opposite directions are, for linear regression purposes, identical. Therefore, we propose the usage of antipodally invariant metrics and introduce the Half Hypersphere Confinement (HHC), a fast alternative to Multidimensional Scaling (MDS) that allows to map antipodally invariant distances in the Euclidean space with very little approximation error By doing so, we enable the usage of fast search structures based on Euclidean distances without undermining their speed gains with complex distance transformations. The performance of our method, which we named HHC Regression (HHCR), applied to SuperResolution (SR) improves both in quality (PSNR) and it is faster than any other state-of-the-art method. Additionally, under an application-agnostic interpretation of our regression framework, we also test our algorithm for denoising and depth upscaling with promising results.
Eduardo Pérez-Pellitero, Jordi Salvador, Javier Ruiz Hidalgo, Bodo Rosenhahn
WACV4
2016 Recognizing human actions using novel space-time volume binary patterns
Florian Baumann, Arne Ehlers, Bodo Rosenhahn
Neurocomputing3
2016 Human Pose Estimation from Video and IMUs
abstract
In this work, we present an approach to fuse video with sparse orientation data obtained from inertial sensors to improve and stabilize full-body human motion capture. Even though video data is a strong cue for motion analysis, tracking artifacts occur frequently due to ambiguities in the images, rapid motions, occlusions or noise. As a complementary data source, inertial sensors allow for accurate estimation of limb orientations even under fast motions. However, accurate position information cannot be obtained in continuous operation. Therefore, we propose a hybrid tracker that combines video with a small number of inertial units to compensate for the drawbacks of each sensor type: on the one hand, we obtain drift-free and accurate position information from video data and, on the other hand, we obtain accurate limb orientations and good performance under fast motions from inertial sensors. In several experiments we demonstrate the increased performance and stability of our human motion tracker.
Timo von Marcard, Gerard Pons-Moll, Bodo Rosenhahn
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 3D Reconstruction of Human Motion from Monocular Image Sequences
abstract
This article tackles the problem of estimating non-rigid human 3D shape and motion from image sequences taken by uncalibrated cameras. Similar to other state-of-the-art solutions we factorize 2D observations in camera parameters, base poses and mixing coefficients. Existing methods require sufficient camera motion during the sequence to achieve a correct 3D reconstruction. To obtain convincing 3D reconstructions from arbitrary camera motion, our method is based on a-priorly trained base poses. We show that strong periodic assumptions on the coefficients can be used to define an efficient and accurate algorithm for estimating periodic motion such as walking patterns. For the extension to non-periodic motion we propose a novel regularization term based on temporal bone length constancy. In contrast to other works, the proposed method does not use a predefined skeleton or anthropometric constraints and can handle arbitrary camera motion. We achieve convincing 3D reconstructions, even under the influence of noise and occlusions. Multiple experiments based on a 3D error metric demonstrate the stability of the proposed method. Compared to other state-of-the-art methods our algorithm shows a significant improvement.
Bastian Wandt, Hanno Ackermann, Bodo Rosenhahn
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Antipodally Invariant Metrics for Fast Regression-Based Super-Resolution
abstract
Dictionary-based super-resolution (SR) algorithms usually select dictionary atoms based on the distance or similarity metrics. Although the optimal selection of the nearest neighbors is of central importance for such methods, the impact of using proper metrics for SR has been overlooked in literature, mainly due to the vast usage of Euclidean distance. In this paper, we present a very fast regression-based algorithm, which builds on the densely populated anchored neighborhoods and sublinear search structures. We perform a study of the nature of the features commonly used for SR, observing that those features usually lie in the unitary hypersphere, where every point has a diametrically opposite one, i.e., its antipode, with same module and angle, but the opposite direction. Even though, we validate the benefits of using antipodally invariant metrics, most of the binary splits use Euclidean distance, which does not handle antipodes optimally. In order to benefit from both the worlds, we propose a simple yet effective antipodally invariant transform that can be easily included in the Euclidean distance calculation. We modify the original spherical hashing algorithm with this metric in our antipodally invariant spherical hashing scheme, obtaining the same performance as a pure antipodally invariant metric. We round up our contributions with a novel feature transform that obtains a better coarse approximation of the input image thanks to iterative backprojection. The performance of our method, which we named antipodally invariant SR, improves quality (Peak Signal to Noise Ratio) and it is faster than any other state-of-the-art method.
Eduardo Pérez-Pellitero, Jordi Salvador, Javier Ruiz Hidalgo, Bodo Rosenhahn
IEEE Trans. Image Process.4
2016 Editorial
Renhong Wang 0001, Bodo Rosenhahn
Vis. Comput.2
2015 On-The-Fly Handwriting Recognition Using a High-Level Representation
Christoph Reinders, Florian Baumann, Björn Scheuermann 0002, Arne Ehlers, Nicole Mühlpforte, Alfred Effenberg, Bodo Rosenhahn
CAIP (1)7
2015 Expanding object detector's Horizon: Incremental learning framework for object detection in videos
abstract
Over the last several years it has been shown that image-based object detectors are sensitive to the training data and often fail to generalize to examples that fall outside the original training sample domain (e.g., videos). A number of domain adaptation (DA) techniques have been proposed to address this problem. DA approaches are designed to adapt a fixed complexity model to the new (e.g., video) domain. We posit that unlabeled data should not only allow adaptation, but also improve (or at least maintain) performance on the original and other domains by dynamically adjusting model complexity and parameters. We call this notion domain expansion. To this end, we develop a new scalable and accurate incremental object detection algorithm, based on several extensions of large-margin embedding (LME). Our detection model consists of an embedding space and multiple class prototypes in that embedding space, that represent object classes; distance to those prototypes allows us to reason about multi-class detection. By incrementally detecting object instances in video and adding confident detections into the model, we are able to dynamically adjust the complexity of the detector over time by instantiating new prototypes to span all domains the model has seen. We test performance of our approach by expanding an object detector trained on ImageNet to detect objects in egocentric videos of Activity Daily Living (ADL) dataset and challenging videos from YouTube Objects (YTO) dataset.
Alina Kuznetsova, Sung Ju Hwang, Bodo Rosenhahn, Leonid Sigal
CVPR3
2015 Fast label propagation for real-time superpixels for video content
abstract
Many recent superpixel algorithms for video content rely on dense optical flow vectors to propagate segmentation results from one frame to the next. In this paper, we assess the impact of the optical flow quality on the over-segmentation quality. Our evaluation shows that it is indispensable for videos with large object displacement and camera motion. But due to the high computational costs high-quality, dense optical flow is not suitable for real-time applications. Therefore, we propose a fast propagation scheme that is based on sparse feature tracking and mesh-based image warping. In a thorough evaluation, we compare our proposed scheme to the results of other state-of-the-art propagation methods using established benchmarks. The results show that our method speeds up the propagation process by a factor of 100 while producing a comparable segmentation quality.
Matthias Reso, Jörn Jachalsky, Bodo Rosenhahn, Jörn Ostermann
ICIP3
2015 A novel dictionary learning method for remote sensing image classification
abstract
With the widely application of high-resolution remote sensing images, its classification has attracted a lot of attention. Most classification methods focus on various combination of features and ignore the similarities between different categories. In this paper we present a modification by combining ScSPM [1] with a dictionary learning method DL-COPAR [2], which separates the particularity and commonality atoms of class-specific sub-dictionaries. With this over-complete dictionary, the sparse representation of a query image can be specified to capture salient and unique properties. Experimental results on two remote sensing datasets show that, this modification achieves state-of-the-art classification accuracy, when merely SIFT feature is applied.
Michael Ying Yang, Saif Dawood Salman Al-Shaikhli, Bodo Rosenhahn
IGARSS4
2015 Hyperspectral image classification using Gaussian process models
abstract
Hyperspectral image processing has been a very dynamic area in remote sensing and other applications since last decades. Hyperspectral images provide abundant spectral information to identify and distinguish spectrally similar materials. Recent advances in kernel machines promote the novel use of Gaussian processes (GP) for classifying hyper-spectral images. Many sophisticated kernel functions have been provided for kernel-based methods. However, different kernel functions has different performance in different applications. This paper introduces GP models with different kernel functions for classifying hyperspectral images. We first provided the mathematical formulation of GP models for classification. Then, several popular kernel functions and their hyperparaeters selection for GP models are introduced. The experiment are performed on three benchmark datasets to evaluate the performances of different kernel functions in terms of classification accuracy. Their performances are compared with each other and discussed in detailed.
Michael Ying Yang, Wentong Liao, Bodo Rosenhahn
IGARSS3
2015 Correspondence between variational methods and Hidden Markov Models
abstract
This paper establishes a duality between the calculus of variations, an increasingly common method for trajectory planning, and Hidden Markov Models (HMMs), a common probabilistic graphical model with applications in artificial intelligence and machine learning. This duality allows findings from each field to be applied to the other, namely providing an efficient and robust global optimization tool and machine learning algorithms for variational problems, and fast local solution methods for large state-space HMMs.
Jens R. Ziehn, Miriam Ruf, Bodo Rosenhahn, Dieter Willersinn, Jürgen Beyerer, Heinrich Gotzig
Intelligent Vehicles Symposium3
2015 Sequential Boosting for Learning a Random Forest Classifier
abstract
This paper introduces a novel tree induction algorithm called sequential Random Forest (sRF) to improve the detection accuracy of a standard Random Forest classifier. Observations have shown that the overall performance of a forest is strongly influenced by the number of training samples. The main idea is to sequentially adapt the number of training samples per class so that each tree better complements the existing trees in the whole forest. Further, we propose a weighted majority voting with respect to a class and tree specific error rate for decreasing the influence of poorly performing trees. The sRF algorithm shows competing results in comparison to state-of-the-art approaches using two datasets for object recognition, two standard machine learning datasets and three datasets for human action recognition.
Florian Baumann, Arne Ehlers, Bodo Rosenhahn
WACV3
2015 A Global-to-Local Framework for Infrared and Visible Image Sequence Registration
abstract
Based on the development of image registration, sequence registration can be done by computing the transformations between consecutive frames. To take into account the accumulated error, global registration method is usually employed as a global error minimizing approach. However, in real surveillance applications, the visible sequence and infrared sequence may be taken at different times, or from different viewpoints, and may have different dynamic contents. Therefore, global registration is only an approximate estimation for two sequences, resulting in inferior local contents. In this paper we present a novel integrated global-to-local framework that addresses the problems of dynamic infrared and visible image sequence registration. We propose to maximize the sum of the mutual information of two sequences for the global homography estimation. Then, frame-to-frame registration is performed to estimate the per-frame local homography. Finally, a smoothing strategy is adopted to smooth the local homographies in the temporal domain to enforce temporal consistency. We evaluate our proposed framework by comparing it to the state-of-the art sequence registration algorithm. Our method achieves improved performance on the public benchmark dataset.
Michael Ying Yang, Yu Qiang, Bodo Rosenhahn
WACV3
2015 Descriptor evaluation and feature regression for multimodal image analysis
Xuanzi Yong, Michael Ying Yang, Yanpeng Cao, Bodo Rosenhahn
Mach. Vis. Appl.4
2014 Fast Super-Resolution via Dense Local Training and Inverse Regressor Search
Eduardo Pérez-Pellitero, Jordi Salvador, Iban Torres-Xirau, Javier Ruiz Hidalgo, Bodo Rosenhahn
ACCV (3)5
2014 Superpixels for Video Content Using a Contour-Based EM Optimization
Matthias Reso, Jörn Jachalsky, Bodo Rosenhahn, Jörn Ostermann
ACCV (4)3
2014 Computation strategies for volume local binary patterns applied to action recognition
abstract
Volume Local Binary Patterns are a well-known feature type to describe object characteristics in the spatiotemporal domain. Apart from the computation of a binary pattern further steps are required to create a discriminative feature. In this paper we propose different computation methods for Volume Local Binary Patterns. These methods are evaluated in detail and the best strategy is shown. A Random Forest is used to find discriminative patterns. The proposed methods are applied to the well-known and publicly available KTH dataset and Weizman dataset for single-view action recognition and to the IXMAS dataset for multiview action recognition. Furthermore, a comparison of the proposed framework to state-of-the-art methods is given.
Florian Baumann, Arne Ehlers, Bodo Rosenhahn
AVSS3
2014 Localization accuracy of interest point detectors with different scale space representations
abstract
The detection of scale invariant image features is a fundamental task for computer vision applications like object recognition or re-identification. Features are localized by computing extrema of the gradients in the Laplacian of Gaussian (LoG) scale space. The most popular detector for scale invariant features is the SIFT detector which uses the Difference of Gaussians (DoG) pyramid as an approximation of the LoG. Recently, the alternative interest point (ALP) detector demonstrated its strength in fast computation on highly parallel architectures like the GPU. It uses the LoG scale space representation for the localization of interest points. This paper evaluates the localization accuracy of ALP in comparison to SIFT. By using synthetic images, it is demonstrated that both localization approaches show a systematic error which is dependent on the subpixel position of the feature. The error increases with the scale of the detected feature. However, using the LoG instead of the DoG representation reduces the maximum systematic error by 77 %. For the evaluation with natural images, benchmark data sets are used. The repeatability criterion evaluates the accuracy of the detectors. The LoG based detector results in up to 16 % higher repeatability. The comparisons are completed with a reference feature localization which uses a signal based approach for the gradient approximation. Based on this approach, a new feature selection criterion is proposed.
Kai Cordes, Bodo Rosenhahn, Jörn Ostermann
AVSS2
2014 Learning an Image-Based Motion Context for Multiple People Tracking
abstract
We present a novel method for multiple people tracking that leverages a generalized model for capturing interactions among individuals. At the core of our model lies a learned dictionary of interaction feature strings which capture relationships between the motions of targets. These feature strings, created from low-level image features, lead to a much richer representation of the physical interactions between targets compared to hand-specified social force models that previous works have introduced for tracking. One disadvantage of using social forces is that all pedestrians must be detected in order for the forces to be applied, while our method is able to encode the effect of undetected targets, making the tracker more robust to partial occlusions. The interaction feature strings are used in a Random Forest framework to track targets according to the features surrounding them. Results on six publicly available sequences show that our method outperforms state-of-the-art approaches in multiple people tracking.
Laura Leal-Taixé, Michele Fenzi, Alina Kuznetsova, Bodo Rosenhahn, Silvio Savarese
CVPR4
2014 Posebits for Monocular Human Pose Estimation
abstract
We advocate the inference of qualitative information about 3D human pose, called posebits, from images. Posebits represent Boolean geometric relationships between body parts (e.g., left-leg in front of right-leg or hands close to each other). The advantages of posebits as a mid-level representation are 1) for many tasks of interest, such qualitative pose information may be sufficient (e.g., semantic image retrieval), 2) it is relatively easy to annotate large image corpora with posebits, as it simply requires answers to yes/no questions, and 3) they help resolve challenging pose ambiguities and therefore facilitate the difficult talk of image-based 3D pose estimation. We introduce posebits, a posebit database, a method for selecting useful posebits for pose estimation and a structural SVM model for posebit inference. Experiments show the use of posebits for semantic image retrieval and for improving 3D pose estimation.
Gerard Pons-Moll, David J. Fleet, Bodo Rosenhahn
CVPR3
2014 Brain tumor classification using sparse coding and dictionary learning
abstract
Brain tumor classification is considered as one of the most challenging tasks in medical imaging. In this paper, a novel approach for multi-class brain tumor classification based on sparse coding and dictionary learning is proposed. We propose an individual (per-class) dictionary learning and sparse coding classification using K-SVD algorithm. This approach combines topological and texture features to build and learn a dictionary. Experimental results demonstrate that the sparse coding based classification outperforms other state-of-the-art methods.
Saif Dawood Salman Al-Shaikhli, Michael Ying Yang, Bodo Rosenhahn
ICIP3
2014 Motion Binary Patterns for Action Recognition
abstract
In this paper, we propose a novel feature type to recognize human actions from video data. By combining the benefit of Volume Local Binary Patterns and Optical Flow, a simple and efficient descriptor is constructed. Motion Binary Patterns (MBP) are computed in spatio-temporal domain while static object appearances as well as motion information are gathered. Histograms are used to learn a Random Forest classifier which is applied to the task of human action recognition. The proposed framework is evaluated on the well-known, publicly available KTH dataset, Weizman dataset and on the IXMAS dataset for multi-view action recognition. The results demonstrate state-of-the-art accuracies in comparison to other methods. 1
Florian Baumann, Jie Lao, Arne Ehlers, Bodo Rosenhahn
ICPRAM4
2014 Video segmentation with joint object and trajectory labeling
abstract
Unsupervised video object segmentation is a challenging problem because it involves a large amount of data and object appearance may significantly change over time. In this paper, we propose a bottom-up approach for the combination of object segmentation and motion segmentation using a novel graphical model, which is formulated as inference in a conditional random field (CRF) model. This model combines object labeling and trajectory clustering in a unified probabilistic framework. The CRF contains binary variables representing the class labels of image pixels as well as binary variables indicating the correctness of trajectory clustering, which integrates dense local interaction and sparse global constraint. An optimization scheme based on a coordinate ascent style procedure is proposed to solve the inference problem. We evaluate our proposed framework by comparing it to other video and motion segmentation algorithms. Our method achieves improved performance on state-of-the-art benchmark datasets.
Michael Ying Yang, Bodo Rosenhahn
WACV2
2014 Multi-Sensor Fusion for Video Segmentation
abstract
Video Segmentation is a fundamental task in computer vision. In many sequences, appearance does not provide enough information to solve the problem. Time-of-Flight cameras provide additional information, namely depth, that can be integrated as an additional feature in a segmentation approach. Typically, the depth information is less sensitive to environment changes. Combined with appearance, this has the potential to be a more robust segmentation method. Motivated by the fact that a simple combination of two information sources might not be the best solution, a novel scheme based on Dempster's theory of evidence is proposed. In contrast to existing methods, the use of Dempster's theory of evidence allows to model inaccuracy and uncertainty. The inaccuracy of the information is influenced by an adaptive weight, that provides a measurement of how reliable a certain information might be. The proposed method is compared with others on a publicly available set of image sequences. The experiments show that the use of the proposed feature fusion improves the segmentation.
Björn Scheuermann 0002, Bodo Rosenhahn
Int. J. Pattern Recognit. Artif. Intell.2
2014 Estimating layout of cluttered indoor scenes using trajectory-based priors
Muhammad Shoaib 0007, Michael Ying Yang, Bodo Rosenhahn, Jörn Ostermann
Image Vis. Comput.3
2013 Bayesian region selection for adaptive dictionary-based Super-Resolution
abstract
The performance of dictionary-based super-resolution (SR) strongly depends on the\ncontents of the training dataset. Nevertheless, many dictionary-based SR methods randomly select patches from of a larger set of training images to build their dictionaries\n[\n8\n,\n14\n,\n19\n,\n20\n], thus relying on patches being diverse enough. This paper describes\na dictionary building method for SR based on adaptively selecting an optimal subset of\npatches out of the training images. Each training image is divided into sub-image entities,\nnamed regions, of such a size that texture consistency is preserved and high-frequency\n(HF) energy is present. For each input patch to super-resolve, the best-fitting region is\nfound through a Bayesian selection. In order to handle the high number of regions in\nthe training dataset, a local Naive Bayes Nearest Neighbor (NBNN) approach is used.\nTrained with this adapted subset of patches, sparse coding SR is applied to recover the\nhigh-resolution image. Experimental results demonstrate that using our adaptive algo-\nrithm produces an improvement in SR performance with respect to non-adaptive training.
Eduardo Pérez-Pellitero, Jordi Salvador, Javier Ruiz Hidalgo, Bodo Rosenhahn
BMVC4
2013 High-Resolution Feature Evaluation Benchmark
Kai Cordes, Bodo Rosenhahn, Jörn Ostermann
CAIP (1)2
2013 Cleaning Up Multiple Detections Caused by Sliding Window Based Object Detectors
Arne Ehlers, Björn Scheuermann 0002, Florian Baumann, Bodo Rosenhahn
CIARP (1)4
2013 Multi-sensor Fusion Using Dempster's Theory of Evidence for Video Segmentation
Björn Scheuermann 0002, Sotirios Gkoutelitsas, Bodo Rosenhahn
CIARP (2)3
2013 Class Generative Models Based on Feature Regression for Pose Estimation of Object Categories
abstract
In this paper, we propose a method for learning a class representation that can return a continuous value for the pose of an unknown class instance using only 2D data and weak 3D labeling information. Our method is based on generative feature models, i.e., regression functions learned from local descriptors of the same patch collected under different viewpoints. The individual generative models are then clustered in order to create class generative models which form the class representation. At run-time, the pose of the query image is estimated in a maximum a posteriori fashion by combining the regression functions belonging to the matching clusters. We evaluate our approach on the EPFL car dataset and the Pointing'04 face dataset. Experimental results show that our method outperforms by 10% the state-of-the-art in the first dataset and by 9% in the second.
Michele Fenzi, Laura Leal-Taixé, Bodo Rosenhahn, Jörn Ostermann
CVPR3
2013 Slice Sampling Particle Belief Propagation
abstract
Inference in continuous label Markov random fields is a challenging task. We use particle belief propagation (PBP) for solving the inference problem in continuous label space. Sampling particles from the belief distribution is typically done by using Metropolis-Hastings (MH) Markov chain Monte Carlo (MCMC) methods which involves sampling from a proposal distribution. This proposal distribution has to be carefully designed depending on the particular model and input data to achieve fast convergence. We propose to avoid dependence on a proposal distribution by introducing a slice sampling based PBP algorithm. The proposed approach shows superior convergence performance on an image denoising toy example. Our findings are validated on a challenging relational 2D feature tracking application.
Oliver Müller 0003, Michael Ying Yang, Bodo Rosenhahn
ICCV3
2013 Temporally Consistent Superpixels
abstract
Super pixel algorithms represent a very useful and increasingly popular preprocessing step for a wide range of computer vision applications, as they offer the potential to boost efficiency and effectiveness. In this regards, this paper presents a highly competitive approach for temporally consistent super pixels for video content. The approach is based on energy-minimizing clustering utilizing a novel hybrid clustering strategy for a multi-dimensional feature space working in a global color subspace and local spatial subspaces. Moreover, a new contour evolution based strategy is introduced to ensure spatial coherency of the generated super pixels. For a thorough evaluation the proposed approach is compared to state of the art super voxel algorithms using established benchmarks and shows a superior performance.
Matthias Reso, Jörn Jachalsky, Bodo Rosenhahn, Jörn Ostermann
ICCV3
2013 "RegionCut" - Interactive multi-label segmentation utilizing cellular automaton
abstract
This paper addresses the problem of interactive image segmentation. We propose an extension of the GrowCut framework which follows Cellular Automaton theory and is comparable to a label propagation algorithm. Therefore, user labels are propagated according to Cellular Automaton until convergency. A common problem of GrowCut is the time consuming user initialization which requires distributed seeds. Our main contribution focuses on determining such an initialization utilizing GMMs and spherical coordinates. Furthermore we propose a new weight function based on the mean image gradient. According to our evaluation, our extensions result in a simplified user interaction and in better results in terms of accuracy and running time. Our experiments show that our method can compete with state-of-the-art energy minimization frameworks.
Oliver Jakob Arndt, Björn Scheuermann 0002, Bodo Rosenhahn
WACV3
2012 Non-rigid Self-calibration of a Projective Camera
Hanno Ackermann, Bodo Rosenhahn
ACCV (4)2
2012 Learning Object Appearance from Occlusions Using Structure and Motion Recovery
Kai Cordes, Björn Scheuermann 0002, Bodo Rosenhahn, Jörn Ostermann
ACCV (3)3
2012 Efficient Pixel-Grouping Based on Dempster's Theory of Evidence for Image Segmentation
Björn Scheuermann 0002, Markus Schlosser, Bodo Rosenhahn
ACCV (1)3
2012 Branch-and-price global optimization for multi-view multi-target tracking
abstract
We present a new algorithm to jointly track multiple objects in multi-view images. While this has been typically addressed separately in the past, we tackle the problem as a single global optimization. We formulate this assignment problem as a min-cost problem by defining a graph structure that captures both temporal correlations between objects as well as spatial correlations enforced by the configuration of the cameras. This leads to a complex combinatorial optimization problem that we solve using Dantzig-Wolfe decomposition and branching. Our formulation allows us to solve the problem of reconstruction and tracking in a single step by taking all available evidence into account. In several experiments on multiple people tracking and 3D human pose tracking, we show our method outperforms state-of-the-art approaches.
Laura Leal-Taixé, Gerard Pons-Moll, Bodo Rosenhahn
CVPR3
2012 Multi-scale Clustering of Frame-to-Frame Correspondences for Motion Segmentation
Ralf Dragon, Bodo Rosenhahn, Jörn Ostermann
ECCV (2)2
2012 Compensating Motion Artifacts of 3D in vivo SD-OCT Scans
Oliver Müller 0003, Sabine Donner, Tobias Klinder, Ivonne Bartsch, Alexander Krüger, Alexander Heisterkamp, Bodo Rosenhahn
MICCAI (1)7
2012 Region-based pose tracking with occlusions using 3D models
Christian Schmaltz, Bodo Rosenhahn, Thomas Brox, Joachim Weickert
Mach. Vis. Appl.2
2011 PCA Enhanced Training Data for Adaboost
Arne Ehlers, Florian Baumann, Ralf Spindler, Birgit Glasmacher, Bodo Rosenhahn
CAIP (1)5
2011 Outdoor human motion capture using inverse kinematics and von mises-fisher sampling
abstract
Human motion capturing (HMC) from multiview image sequences is an extremely difficult problem due to depth and orientation ambiguities and the high dimensionality of the state space. In this paper, we introduce a novel hybrid HMC system that combines video input with sparse inertial sensor input. Employing an annealing particle-based optimization scheme, our idea is to use orientation cues derived from the inertial input to sample particles from the manifold of valid poses. Then, visual cues derived from the video input are used to weight these particles and to iteratively derive the final pose. As our main contribution, we propose an efficient sampling procedure where the particles are derived analytically using inverse kinematics on the orientation cues. Additionally, we introduce a novel sensor noise model to account for uncertainties based on the von Mises-Fisher distribution. Doing so, orientation constraints are naturally fulfilled and the number of needed particles can be kept very small. More generally, our method can be used to sample poses that fulfill arbitrary orientation or positional kinematic constraints. In the experiments, we show that our system can track even highly dynamic motions in an outdoor environment with changing illumination, background clutter, and shadows.
Gerard Pons-Moll, Andreas Baak, Juergen Gall, Laura Leal-Taixé, Meinard Müller, Hans-Peter Seidel, Bodo Rosenhahn
ICCV7
2011 Model Based 3D Segmentation and OCT Image Undistortion of Percutaneous Implants
Oliver Müller 0003, Sabine Donner, Tobias Klinder, Ralf Dragon, Ivonne Bartsch, Frank Witte, Alexander Krüger, Alexander Heisterkamp, Bodo Rosenhahn
MICCAI (3)9
2010 A Linear Solution to 1-Dimensional Subspace Fitting under Incomplete Data
Hanno Ackermann, Bodo Rosenhahn
ACCV (2)2
2010 Feature Quarrels: The Dempster-Shafer Evidence Theory for Image Segmentation Using a Variational Framework
Björn Scheuermann 0002, Bodo Rosenhahn
ACCV (2)2
2010 Multilinear pose and body shape estimation of dressed subjects from image sets
abstract
In this paper we propose a multilinear model of human pose and body shape which is estimated from a database of registered 3D body scans in different poses. The model is generated by factorizing the measurements into pose and shape dependent components. By combining it with an ICP based registration method, we are able to estimate pose and body shape of dressed subjects from single images. If several images of the subject are available, shape and poses can be optimized simultaneously for all input images. Additionally, while estimating pose and shape, we use the model as a virtual calibration pattern and also recover the parameters of the perspective camera model the images were created with.
Nils Hasler, Hanno Ackermann, Bodo Rosenhahn, Thorsten Thormählen, Hans-Peter Seidel
CVPR3
2010 Multisensor-fusion for 3D full-body human motion capture
abstract
In this work, we present an approach to fuse video with orientation data obtained from extended inertial sensors to improve and stabilize full-body human motion capture. Even though video data is a strong cue for motion analysis, tracking artifacts occur frequently due to ambiguities in the images, rapid motions, occlusions or noise. As a complementary data source, inertial sensors allow for drift-free estimation of limb orientations even under fast motions. However, accurate position information cannot be obtained in continuous operation. Therefore, we propose a hybrid tracker that combines video with a small number of inertial units to compensate for the drawbacks of each sensor type: on the one hand, we obtain drift-free and accurate position information from video data and, on the other hand, we obtain accurate limb orientations and good performance under fast motions from inertial sensors. In several experiments we demonstrate the increased performance and stability of our human motion tracker.
Gerard Pons-Moll, Andreas Baak, Thomas Helten, Meinard Müller, Hans-Peter Seidel, Bodo Rosenhahn
CVPR6
2010 NF-Features - No-Feature-Features for Representing Non-textured Regions
Ralf Dragon, Muhammad Shoaib 0007, Bodo Rosenhahn, Jörn Ostermann
ECCV (2)3
2010 Learning skeletons for shape and pose
abstract
In this paper a method for estimating a rigid skeleton, including skinning weights, skeleton connectivity, and joint positions, given a sparse set of example poses is presented. In contrast to other methods, we are able to simultaneously take examples of different subjects into account, which improves the robustness of the estimation. It is additionally possible to generate a skeleton that primarily describes variations in body shape instead of pose. The shape skeleton can then be combined with a regular pose varying skeleton. That way pose and body shape can be controlled simultaneously but separately. As this skeleton is technically still just a skinned rigid skeleton, compatibility with major modelling packages and game engines is retained. We further present an approach for synthesizing a suitable bind shape that additionally improves the accuracy of the generated model.
Nils Hasler, Thorsten Thormählen, Bodo Rosenhahn, Hans-Peter Seidel
SI3D3
2010 Foreword
Marcus A. Magnor, Bodo Rosenhahn, Holger Theisel
Comput. Graph.2
2010 Optimization and Filtering for Human Motion Capture
abstract
Local optimization and filtering have been widely applied to model-based 3D human motion capture. Global stochastic optimization has recently been proposed as promising alternative solution for tracking and initialization. In order to benefit from optimization and filtering, we introduce a multi-layer framework that combines stochastic optimization, filtering, and local optimization. While the first layer relies on interacting simulated annealing and some weak prior information on physical constraints, the second layer refines the estimates by filtering and local optimization such that the accuracy is increased and ambiguities are resolved over time without imposing restrictions on the dynamics. In our experimental evaluation, we demonstrate the significant improvements of the multi-layer framework and provide quantitative 3D pose tracking results for the complete HumanEva-II dataset. The paper further comprises a comparison of global stochastic optimization with particle filtering, annealed particle filtering, and local optimization.
Juergen Gall, Bodo Rosenhahn, Thomas Brox, Hans-Peter Seidel
Int. J. Comput. Vis.2
2010 Combined Region and Motion-Based 3D Tracking of Rigid and Articulated Objects
abstract
In this paper, we propose the combined use of complementary concepts for 3D tracking: region fitting on one side and dense optical flow as well as tracked SIFT features on the other. Both concepts are chosen such that they can compensate for the shortcomings of each other. While tracking by the object region can prevent the accumulation of errors, optical flow and SIFT can handle larger transformations. Whereas segmentation works best in case of homogeneous objects, optical flow computation and SIFT tracking rely on sufficiently structured objects. We show that a sensible combination yields a general tracking system that can be applied in a large variety of scenarios without the need to manually adjust weighting parameters.
Thomas Brox, Bodo Rosenhahn, Juergen Gall, Daniel Cremers
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Trajectory reconstruction for affine structure-from-motion by global and local constraints
abstract
The problem of reconstructing a 3D scene from a moving camera can be solved by means of the so-called Factorization method. It directly computes a global solution without the need to merge several partial reconstructions. However, if the trajectories are not complete, i.e. not every feature point could be observed in all the images, this method cannot be used. We use a Factorization-style algorithm for recovering the unobserved feature positions in a non-incremental way. This method uniformly utilizes all data and finds a global solution without any need of sequential or hierarchical merging. Two contributions are made in this work: Firstly, partially known trajectories are completed by minimizing the distance between the subspace and the trajectory within an affine subspace associated with the trajectory. This amounts to imposing a global constraint on the data. Secondly, we propose to further include local constraints derived from epipolar geometry into the estimation. It is shown how to simultaneously optimize both constraints. By using simulated and real image sequences we show the improvements achieved with our algorithm.
Hanno Ackermann, Bodo Rosenhahn
CVPR2
2009 Motion capture using joint skeleton tracking and surface estimation
abstract
This paper proposes a method for capturing the performance of a human or an animal from a multi-view video sequence. Given an articulated template model and silhouettes from a multi-view image sequence, our approach recovers not only the movement of the skeleton, but also the possibly non-rigid temporal deformation of the 3D surface. While large scale deformations or fast movements are captured by the skeleton pose and approximate surface skinning, true small scale deformations or non-rigid garment motion are captured by fitting the surface to the silhouette. We further propose a novel optimization scheme for skeleton-based pose estimation that exploits the skeleton's tree structure to split the optimization problem into a local one and a lower dimensional global one. We show on various sequences that our approach can capture the 3D motion of animals and humans accurately even in the case of rapid movements and wide apparel like skirts.
Juergen Gall, Carsten Stoll, Edilson de Aguiar, Christian Theobalt, Bodo Rosenhahn, Hans-Peter Seidel
CVPR5
2009 Markerless Motion Capture with unsynchronized moving cameras
abstract
In this work we present an approach for markerless motion capture (MoCap) of articulated objects, which are recorded with multiple unsynchronized moving cameras. Instead of using fixed (and expensive) hardware synchronized cameras, this approach allows us to track people with off-the-shelf handheld video cameras. To prepare a sequence for motion capture, we first reconstruct the static background and the position of each camera using Structure-from-Motion (SfM). Then the cameras are registered to each other using the reconstructed static background geometry. Camera synchronization is achieved via the audio streams recorded by the cameras in parallel. Finally, a markerless MoCap approach is applied to recover positions and joint configurations of subjects. Feature tracks and dense background geometry are further used to stabilize the MoCap. The experiments show examples with highly challenging indoor and outdoor scenes.
Nils Hasler, Bodo Rosenhahn, Thorsten Thormählen, Michael Wand 0001, Juergen Gall, Hans-Peter Seidel
CVPR2
2009 Stabilizing motion tracking using retrieved motion priors
abstract
In this paper, we introduce a novel iterative motion tracking framework that combines 3D tracking techniques with motion retrieval for stabilizing markerless human motion capturing. The basic idea is to start human tracking without prior knowledge about the performed actions. The resulting 3D motion sequences, which may be corrupted due to tracking errors, are locally classified according to available motion categories. Depending on the classification result, a retrieval system supplies suitable motion priors, which are then used to regularize and stabilize the tracking in the next iteration step. Experiments with the HumanEVA-II benchmark show that tracking and classification are remarkably improved after few iterations.
Andreas Baak, Bodo Rosenhahn, Meinard Müller, Hans-Peter Seidel
ICCV2
2009 Ball joints for Marker-less human Motion Capture
abstract
This work presents an approach for the modeling and numerical optimization of ball joints within a Marker-less Motion Capture (MoCap) framework. In skeleton based approaches, kinematic chains are commonly used to model 1 DoF revolute joints. A 3 DoF joint (e.g. a shoulder or hip) is consequently modeled by concatenating three consecutive 1 DoF revolute joints. Obviously such a representation is not optimal and singularities can occur. Therefore, we propose to model 3 DoF joints with spherical joints or ball joints using the representation of a twist and its exponential mapping (known from 1 DoF revolute joints). The exact modeling and numerical optimization of ball joints requires additionally the adjoint transform and the logarithm of the exponential mapping. Experiments with simulated and real data demonstrate that ball joints can better represent arbitrary rotations than the concatenation of 3 revolute joints. Moreover, we demonstrate that the 3 revolute joints representation is very similar to the Euler angles representation and has the same limitations in terms of singularities.
Gerard Pons-Moll, Bodo Rosenhahn
WACV2
2009 Estimating body shape of dressed humans
Nils Hasler, Carsten Stoll, Bodo Rosenhahn, Thorsten Thormählen, Hans-Peter Seidel
Comput. Graph.3
2009 A Statistical Model of Human Pose and Body Shape
abstract
Abstract Generation and animation of realistic humans is an essential part of many projects in today's media industry. Especially, the games and special effects industry heavily depend on realistic human animation. In this work a unified model that describes both, human pose and body shape is introduced which allows us to accurately model muscle deformations not only as a function of pose but also dependent on the physique of the subject. Coupled with the model's ability to generate arbitrary human body shapes, it severely simplifies the generation of highly realistic character animations. A learning based approach is trained on approximately 550 full body 3D laser scans taken of 114 subjects. Scan registration is performed using a non‐rigid deformation technique. Then, a rotation invariant encoding of the acquired exemplars permits the computation of a statistical model that simultaneously encodes pose and body shape. Finally, morphing or generating meshes according to several constraints simultaneously can be achieved by training semantically meaningful regressors.
Nils Hasler, Carsten Stoll, Martin Sunkel, Bodo Rosenhahn, Hans-Peter Seidel
Comput. Graph. Forum4
2008 Drift-free tracking of rigid and articulated objects
abstract
Model-based 3D tracker estimate the position, rotation, and joint angles of a given model from video data of one or multiple cameras. They often rely on image features that are tracked over time but the accumulation of small errors results in a drift away from the target object. In this work, we address the drift problem for the challenging task of human motion capture and tracking in the presence of multiple moving objects where the error accumulation becomes even more problematic due to occlusions. To this end, we propose an analysis-by-synthesis framework for articulated models. It combines the complementary concepts of patch-based and region-based matching to track both structured and homogeneous body parts. The performance of our method is demonstrated for rigid bodies, body parts, and full human bodies where the sequences contain fast movements, self-occlusions, multiple moving objects, and clutter. We also provide a quantitative error analysis and comparison with other model-based approaches.
Juergen Gall, Bodo Rosenhahn, Hans-Peter Seidel
CVPR2
2008 Markerless motion capture of man-machine interaction
abstract
This work deals with modeling and markerless tracking of athletes interacting with sports gear. In contrast to classical markerless tracking, the interaction with sports gear comes along with joint movement restrictions due to additional constraints: while humans can generally use all their joints, interaction with the equipment imposes a coupling between certain joints. A cyclist who performs a cycling pattern is one example: The feet are supposed to stay on the pedals, which are again restricted to move along a circular trajectory in 3D-space. In this paper, we present a markerless motion capture system that takes the lower-dimensional pose manifold into account by modeling the motion restrictions via soft constraints during pose optimization. Experiments with two different models, a cyclist and a snowboarder, demonstrate the applicability of the method. Moreover, we present motion capture results for challenging outdoor scenes including shadows and strong illumination changes.
Bodo Rosenhahn, Christian Schmaltz, Thomas Brox, Joachim Weickert, Daniel Cremers, Hans-Peter Seidel
CVPR1
2007 Scaled Motion Dynamics for Markerless Motion Capture
abstract
This work proposes a way to use a-priori knowledge on motion dynamics for markerless human motion capture (MoCap). Specifically, we match tracked motion patterns to training patterns in order to predict states in successive frames. Thereby, modeling the motion by means of twists allows for a proper scaling of the prior. Consequently, there is no need for training data of different frame rates or velocities. Moreover, the method allows to combine very different motion patterns. Experiments in indoor and outdoor scenarios demonstrate the continuous tracking of familiar motion patterns in case of artificial frame drops or in situations insufficiently constrained by the image data.
Bodo Rosenhahn, Thomas Brox, Hans-Peter Seidel
CVPR1
2007 Three-Dimensional Shape Knowledge for Joint Image Segmentation and Pose Tracking
Bodo Rosenhahn, Thomas Brox, Joachim Weickert
Int. J. Comput. Vis.1
2007 A system for articulated tracking incorporating a clothing model
Bodo Rosenhahn, Uwe G. Kersting, Katie Powell 0001, Reinhard Klette, Gisela Klette, Hans-Peter Seidel
Mach. Vis. Appl.1
2006 High Accuracy Optical Flow Serves 3-D Pose Tracking: Exploiting Contour and Flow Based Constraints
Thomas Brox, Bodo Rosenhahn, Daniel Cremers, Hans-Peter Seidel
ECCV (2)2
2006 A Comparison of Shape Matching Methods for Contour Based Pose Estimation
Bodo Rosenhahn, Thomas Brox, Daniel Cremers, Hans-Peter Seidel
IWCIA1
2006 Robust Pose Estimation with 3D Textured Models
Juergen Gall, Bodo Rosenhahn, Hans-Peter Seidel
PSIVT2
2006 Collinearity and Coplanarity Constraints for Structure from Motion
Reinhard Klette, Bodo Rosenhahn
PSIVT3
2006 Target Calibration and Tracking Using Conformal Geometric Algebra
Yilan Zhao, Robert J. Valkenburg, Reinhard Klette, Bodo Rosenhahn
PSIVT4
2005 Automatic Human Model Generation
Bodo Rosenhahn, Reinhard Klette
CAIP1
2004 Pose Estimation of Free-Form Objects
Bodo Rosenhahn, Gerald Sommer
ECCV (1)1
2004 Geometric Algebra for Pose Estimation and Surface Morphing in Human Motion Estimation
Bodo Rosenhahn, Reinhard Klette
IWCIA1
2004 Free-Form Pose Estimation by Using Twist Representations
Bodo Rosenhahn, Christian Perwass, Gerald Sommer
Algorithmica1
2004 Pose Estimation of 3D Free-Form Contours
Bodo Rosenhahn, Christian Perwass, Gerald Sommer
Int. J. Comput. Vis.1
2003 Modeling Adaptive Deformations during Free-Form Pose Estimation
Bodo Rosenhahn, Christian Perwass, Gerald Sommer
CAIP1
2002 A geometric approach for the analysis and computation of the intrinsic camera parameters
Eduardo Bayro-Corrochano, Bodo Rosenhahn
Pattern Recognit.2
1999 Computing the Intrinsic Camera Parameters Using Pascal's Theorem
Bodo Rosenhahn, Eduardo Bayro-Corrochano
CAIP1