Didier Stricker

dblp:02/5478 · DBLP profile ↗
← Back
200ranked-venue papers
0as first author
95since 2021 · last 2026
0000-0002-5708-6023ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 156 · 72 since 2021Artificial intelligence and machine learning · 80 · 48 since 2021Human-computer interaction and ubiquitous computing · 19 · 4 since 2021Databases, data management, data science and information retrieval · 12 · 9 since 2021Systems, architecture and hardware · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 NURBGen: High-Fidelity Text-to-CAD Generation Through LLM-Driven NURBS Modeling
abstract
Generating editable 3D CAD models from natural language remains challenging, as existing text-to-CAD systems either produce meshes or rely on scarce design-history data. We present NURBGen, the first framework to generate high fidelity 3D CAD models directly from text using Non-Uniform Rational B-Splines (NURBS). To achieve this, we fine-tune a large language model (LLM) to translate free-form texts into JSON representations containing NURBS surface parameters (i.e, control points, knot vectors, degrees, and rational weights) which can be directly converted into BRep format using Python. We further propose a hybrid representation that combines untrimmed NURBS with analytic primitives to handle trimmed surfaces and degenerate regions more robustly, while reducing token complexity. Additionally, we introduce partABC, a curated subset of the ABC dataset consisting of individual CAD components, annotated with detailed captions using an automated annotation pipeline. NURBGen demonstrates strong performance on diverse prompts, surpassing prior methods in geometric fidelity and dimensional accuracy, as confirmed by expert evaluations.
Mohammad Sadil Khan, Didier Stricker, Muhammad Zeshan Afzal
AAAI3
2026 PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion
Mahdi Chamseddine, Didier Stricker, Jason R. Rambach
ICPR (1)2
2026 Efficient Post-hoc Calibration in Object Detection Without Held-Out Data
Nikolas Ebert, Didier Stricker, Oliver Wasenmüller
ICPR (1)2
2026 MILE: Mixture of Incremental LoRA Experts for Continual Semantic Segmentation Across Domains and Modalities
Shishir Muralidhara, Didier Stricker, René Schuster
ICPR (10)2
2026 Sensor Generalization for Adaptive Sensing in Event-Based Object Detection via Joint Distribution Training
abstract
Bio-inspired event cameras have recently attracted significant research due to their asynchronous and low-latency capabilities. These features provide a high dynamic range and significantly reduce motion blur. However, because of the novelty in the nature of their output signals, there is a gap in the variability of available data and a lack of extensive analysis of the parameters characterizing their signals. This paper addresses these issues by providing readers with an in-depth understanding of how intrinsic parameters affect the performance of a model trained on event data, specifically for object detection. We also use our findings to expand the capabilities of the downstream model towards sensor-agnostic robustness.
Aheli Saha, René Schuster, Didier Stricker
ICPRAM3
2026 TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
abstract
Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of driving poses. Nevertheless, generating temporally coherent long-form content remains challenging. Existing approaches are constrained by computational and memory limitations, as they are typically trained on short video segments, thus performing effectively only over limited frame lengths and hindering their potential for extended coherent generation. To address these constraints, we propose TalkingPose, a novel diffusion-based framework specifically designed for producing long-form, temporally consistent human upper-body animations. TalkingPose leverages driving frames to precisely capture expressive facial and hand movements, transferring these seamlessly to a target actor through a stable diffusion backbone. To ensure continuous motion and enhance temporal coherence, we introduce a feedback-driven mechanism built upon image-based diffusion models. Notably, this mechanism does not incur additional computational costs or require secondary training stages, enabling the generation of animations with unlimited duration. Additionally, we introduce a comprehensive, large-scale dataset to serve as a new benchmark for human upper-body animation. Project page: https://dfki-av.github.io/TalkingPose
Alireza Javanmardi, Pragati Jaiswal, Tewodros Habtegebrial, Christen Millerdurai, Shaoxiang Wang, Alain Pagani, Didier Stricker
WACV7
2026 PoseAdapt: Sustainable Human Pose Estimation via Continual Learning Benchmarks and Toolkit
abstract
Human pose estimators are typically retrained from scratch or naively fine-tuned whenever keypoint sets, sensing modalities, or deployment domains change—an inefficient, compute-intensive practice that rarely matches field constraints. We present PoseAdapt, an open-source framework and benchmark suite for continual pose model adaptation. PoseAdapt defines domain-incremental and class-incremental tracks that simulate realistic changes in density, lighting, and sensing modality, as well as skeleton growth. The toolkit supports two workflows: (i) Strategy Benchmarking, which lets researchers implement continual learning (CL) methods as plugins and evaluate them under standardized protocols; and (ii) Model Adaptation, which allows practitioners to adapt strong pretrained models to new tasks with minimal supervision. We evaluate representative regularization-based methods in single-step and sequential settings. Benchmarks enforce a fixed lightweight backbone, no access to past data, and tight per-step budgets. This isolates adaptation strategy effects, highlighting the difficulty of maintaining accuracy under strict resource limits. PoseAdapt connects modern CL techniques with practical pose estimation needs, enabling adaptable models that improve over time without repeated full retraining.
Muhammad Saif Ullah Khan, Didier Stricker
WACV2
2026 IMKD: Intensity-Aware Multi-Level Knowledge Distillation for Camera-Radar Fusion
abstract
High-performance Radar-Camera 3D object detection can be achieved by leveraging knowledge distillation without using LiDAR at inference time. However, existing distillation methods typically transfer modality-specific features directly to each sensor, which can distort their unique characteristics and degrade their individual strengths. To address this, we introduce IMKD, a radar-camera fusion framework based on multi-level knowledge distillation that preserves each sensor’s intrinsic characteristics while amplifying their complementary strengths. IMKD applies a three-stage, intensity-aware distillation strategy to enrich the fused representation across the architecture: (1) LiDAR-to-Radar intensity-aware feature distillation to enhance radar representations with fine-grained structural cues, (2) LiDAR-to-Fused feature intensity-guided distillation to selectively highlight useful geometry and depth information at the fusion level, fostering complementarity between the modalities rather than forcing them to align, and (3) Camera-Radar intensity-guided fusion mechanism that facilitates effective feature alignment and calibration. Extensive experiments on the nuScenes benchmark show that IMKD reaches 67.0% NDS and 61.0% mAP, outperforming all prior distillation-based radar-camera fusion methods. Our code and models are available at: https://github.com/dfki-av/IMKD/.
Shashank Mishra 0001, Karan Patil, Didier Stricker, Jason R. Rambach
WACV3
2026 Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° Scenes
abstract
Despite recent advances in single-object front-facing inpainting using NeRF and 3D Gaussian Splatting (3DGS), inpainting in complex 360° scenes remains largely underexplored. This is primarily due to three key challenges: (i) identifying target objects in the 3D field of 360° environments, (ii) dealing with severe occlusions in multi-object scenes, which makes it hard to define regions to inpaint, and (iii) maintaining consistent and high-quality appearance across views effectively.To tackle these challenges, we propose Inpaint360GS, a flexible 360° editing framework based on 3DGS that supports multi-object removal and high-fidelity inpainting in 3D space. By distilling 2D segmentation into 3D and leveraging virtual camera views for contextual guidance, our method enables accurate object-level editing and consistent scene completion. We further introduce a new dataset tailored for 360° inpainting, addressing the lack of ground truth object-free scenes. Experiments demonstrate that Inpaint360GS outperforms existing baselines and achieves state-of-the-art performance. Project page: https://dfki-av.github.io/inpaint360gs/
Shaoxiang Wang, Shihong Zhang, Christen Millerdurai, Rüdiger Westermann, Didier Stricker, Alain Pagani
WACV5
2026 Multi-View Face and Gesture Animation with Dynamic Gaussians
abstract
Creating photorealistic 3D human avatars with realistic upper-body motion remains challenging. Existing approaches either focus on the head and overlook hand gestures, or reconstruct the full body but fail to preserve fine-grained facial fidelity and hand pose accuracy. As a result, current methods struggle to capture the subtle dynamics of facial expressions and hand gestures that are crucial for natural human communication. While methods based on full-body parametric models enable avatar reconstruction from monocular or multi-view inputs, they often lack accurate facial animation and detailed hand articulation. To address these limitations, we propose MVFGA, a novel multi-view-consistent pipeline for generating realistic upper-body avatars. Our approach models the face and hands separately and fuses them with a parametric upper-body mesh model, enabling the capture of fine-grained facial expressions and hand poses for accurate upper-body avatar reconstruction. We then splat 3D Gaussians onto the obtained mesh, enabling high-quality rendering of dynamic avatars from novel viewpoints. Furthermore, we introduce MVFGA-MoCap, a multi-view upper-body motion capture dataset featuring controlled facial expression sequences, diverse hand gestures, and free-form communication. Experiments show that MVFGA generates visually realistic avatars with high-fidelity facial expressions and hand motions, outperforming baselines for upper-body avatar animation. Project page: https://dfki-av.github.io/MVFGA/
Alireza Javanmardi, Vippin Kumar Jeetmal, Christen Millerdurai, Alain Pagani, Didier Stricker
Comput. Graph. Forum5
2025 MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
abstract
Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MARVEL-40M+, an extensive dataset with 40 million text annotations for over 8.9 million 3D assets aggregated from seven major 3D datasets. Our contribution is a novel multi-stage annotation pipeline that integrates open-source pretrained multi-view VLMs and LLMs to automatically produce multi-level descriptions, ranging from detailed (150-200 words) to concise semantic tags (10-20 words). This structure supports both fine-grained 3D reconstruction and rapid prototyping. Furthermore, we incorporate human metadata from source datasets into our annotation pipeline to add domain-specific information in our annotation and reduce VLM hallucinations. Additionally, we develop MARVEL-FX3D, a two-stage text-to-3D pipeline. We fine-tune Stable Diffusion with our annotations and use a pretrained image-to-3D network to generate 3D textured meshes within 15s. Extensive evaluations show that MARVEL-40M+ significantly outperforms existing datasets in annotation quality and linguistic diversity, achieving win rates of 72.41% by GPT-4 and 73.40% by human evaluators. Project page is available at https://sankalpsinha-cmos.github.io/MARVEL/.
Sankalp Sinha, Mohammad Sadil Khan, Shino Sam, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
CVPR5
2025 TorchAdapt: Towards Light-Agnostic Real-Time Visual Perception
Khurram Azeem Hashmi, Karthik Palyakere Suresh, Didier Stricker, Muhammad Zeshan Afzal
ICCV3
2025 STEP-DETR: Advancing DETR-based Semi-Supervised Object Detection with Super Teacher and Pseudo-Label Guided Text Queries
Tahira Shehzadi, Khurram Azeem Hashmi, Shalini Sarode, Didier Stricker, Muhammad Zeshan Afzal
ICCV4
2025 SemiTabDETR: End-to-End Semi-supervised Table Detection with Transformer-Based Enhanced Query Approach
Tahira Shehzadi, Didier Stricker, Muhammad Zeshan Afzal
ICDAR (1)2
2025 1D-DiffNPHR: 1D Diffusion Neural Parametric Head Reconstruction Using a Single Image
Pragati Jaiswal, Tewodros Habtegebrial, Didier Stricker
ICPRAM3
2025 HI²: Sparse-View 3D Object Reconstruction with a Hybrid Implicit Initialization
Pragati Jaiswal, Didier Stricker
ICPRAM2
2025 Domain-Incremental Semantic Segmentation for Autonomous Driving Under Adverse Driving Conditions
Shishir Muralidhara, René Schuster, Didier Stricker
ICPRAM3
2025 Object-Centric 2D Gaussian Splatting: Background Removal and Occlusion-Aware Pruning for Compact Object Models
Marcel Rogge, Didier Stricker
ICPRAM2
2025 3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
abstract
Multimodal Large Language Models (MLLMs) have made significant progress in tasks such as image captioning and question answering. However, while these models can generate realistic captions, they often struggle with providing precise instructions, particularly when it comes to localizing and disambiguating objects in complex 3D environments. This capability is critical as MLLMs become more integrated with collaborative robotic systems. In scenarios where a target object is surrounded by similar objects (distractors), robots must deliver clear, spatially-aware instructions to guide humans effectively. We refer to this challenge as contextual object localization and disambiguation, which imposes stricter constraints than conventional 3D dense captioning, especially regarding ensuring target exclusivity. In response, we propose simple yet effective techniques to enhance the model's ability to localize and disambiguate target objects. Our approach not only achieves state-of-the-art performance on conventional metrics that evaluate sentence similarity, but also demonstrates improved 3D spatial understanding through 3D visual grounding model. https://birdy666.github.io/projects/3d_spatial_understanding_in_mllms/
Chun-Peng Chang, Alain Pagani, Didier Stricker
ICRA3
2025 JENGA: Object selection and pose estimation for robotic grasping from a stack
abstract
Vision-based robotic object grasping is typically investigated in the context of isolated objects or unstructured object sets in bin picking scenarios. However, there are several settings, such as construction or warehouse automation, where a robot needs to interact with a structured object formation such as a stack. In this context, we define the problem of selecting suitable objects for grasping along with estimating an accurate 6DoF pose of these objects. To address this problem, we propose a camera-IMU based approach that prioritizes unobstructed objects on the higher layers of stacks and introduce a dataset for benchmarking and evaluation, along with a suitable evaluation metric that combines object selection with pose accuracy. Experimental results show that although our method can perform quite well, this is a challenging problem if a completely error-free solution is needed. Finally, we show results from the deployment of our method for a brick-picking application in a construction scenario.
Sai Srinivas Jeevanandam, Sandeep Inuganti, Shreedhar Govil, Didier Stricker, Jason R. Rambach
IROS4
2025 Rendering Togetherness: Embodied Social Synchronization in Multi-User VR
abstract
Implementing multi-person interactions in Virtual Reality (VR) poses two interrelated challenges: (1) accurately capturing and rendering the kinematics of social interactions, and (2) fostering social connectedness, a key component of effective group communication. In physical settings, social connectedness is closely linked to interpersonal motor coordination. Whether similar mechanisms operate in VR, and how best to render them remains an open question. This study investigated group synchronization in VR by manipulating two key factors: visual coupling and joint commitment to synchronize, both known to influence group synchrony and social connectedness. Data from this VR experiment were compared with a previous study conducted in a real-world setting. In both contexts, visual coupling and joint commitment enhanced group synchrony and the feeling of social connectedness. However, their effects on individual kinematic features (e.g., movement frequency and amplitude) diverged between real and virtual environments, suggesting different coordination strategies were employed. These findings demonstrate that while multi-user VR can support the emergence of collective movement and foster social bonds, technical constraints, such as limited motion fidelity and restricted field of view, can shape how users adapt their movements to achieve joint action. This has important theoretical and practical implications for the design and modeling of collective motion in VR.
Julia Ayache, Julie LaRoche, Maxime Hérissé, Pierre Jean, Anna Katharina Hebborn, Didier Stricker, Benoît G. Bardy
ISMAR6
2025 Supplementary Material AnonyNoise: Anonymizing Event Data with Smart Noise to Outsmart Re-Identification and Preserve Privacy
abstract
In this supplementary material, we provide a more detailed overview of AnonyNoise, a method developed for predicting data-dependent noise aimed at preventing reidentification. The document is structured as follows: First, we detail the training parameters used in our implementation for the three datasets: DVS-Gesture [4], SEE [29], and Event-ReId [1], in order to ensure reproducibility. Next, we present numerical results from an inversion attack on our method, comparing its effectiveness to Gaussian noise when evaluated using a denoising network. This comparison provides insight into the robustness of AnonyNoise in contrast to traditional noise techniques in preventing data recovery and re-identification attempts. We moreover provide our statement regarding our responsibility to human subjects in the datasets used during our experiments. Lastly, we include an expanded set of visual examples across all datasets, including the results from image reconstruction attacks.
Katharina Bendig, René Schuster, Nicole Thiemer, Karen Joisten, Didier Stricker
WACV5
2025 Beyond Boxes: Mask-Guided Spatio-Temporal Feature Aggregation for Video Object Detection
abstract
The primary challenge in Video Object Detection (VOD) is effectively exploiting temporal information to enhance object representations. Traditional strategies, such as aggregating region proposals, often suffer from feature variance due to the inclusion of background information. We introduce a novel instance mask-based feature aggregation approach, significantly refining this process and deepening the understanding of object dynamics across video frames. We present FAIM, a new VOD method that enhances temporal Feature Aggregation by leveraging Instance Mask features. In particular, we propose the lightweight Instance Feature Extraction Module (IFEM) to learn instance mask features and the Temporal Instance Classification Aggregation Module (TICAM) to aggregate instance mask and classification features across video frames. Using YOLOX as a base detector, FAIM achieves 87.9% mAP on the ImageNet VID dataset at 33 FPS on a single 2080Ti GPU, setting a new benchmark for the speed-accuracy trade-off. Additional experiments on multiple datasets validate that our approach is robust, method-agnostic, and effective in multi-object tracking, demonstrating its broader applicability to video understanding tasks.
Khurram Azeem Hashmi, Talha Uddin Sheikh, Didier Stricker, Muhammad Zeshan Afzal
WACV3
2025 Modality-Incremental Learning with Disjoint Relevance Mapping Networks for Image-Based Semantic Segmentation
abstract
In autonomous driving, environment perception has significantly advanced with the utilization of deep learning techniques for diverse sensors such as cameras, depth sensors, or infrared sensors. The diversity in the sensor stack increases the safety and contributes to robustness against adverse weather and lighting conditions. However, the variance in data acquired from different sensors poses challenges. In the context of continual learning (CL), incremental learning is especially challenging for considerably large domain shifts, e.g. different sensor modalities. This amplifies the problem of catastrophic forgetting. To address this issue, we formulate the concept of modality-incremental learning and examine its necessity, by contrasting it with existing incremental learning paradigms. We propose the use of a modified Relevance Mapping Network (RMN) to incrementally learn new modalities while preserving performance on previously learned modalities, in which relevance maps are disjoint. Experimental results demonstrate that the prevention of shared connections in this approach helps alleviate the problem of forgetting within the constraints of a strict continual learning framework.
Niharika Hegde, Shishir Muralidhara, René Schuster, Didier Stricker
WACV4
2025 Uni-SLAM: Uncertainty-Aware Neural Implicit SLAM for Real-Time Dense Indoor Scene Reconstruction
abstract
Neural implicit fields have recently emerged as a powerful representation method for multi-view surface reconstruction due to their simplicity and state-of-the-art performance. However, reconstructing thin structures of indoor scenes while ensuring real-time performance remains a challenge for dense visual SLAM systems. Previous methods do not consider varying quality of input RGB-D data and employ fixed-frequency mapping process to reconstruct the scene, which could result in the loss of valuable information in some frames. In this paper, we propose Uni-SLAM, a decoupled 3D spatial representation based on hash grids for indoor reconstruction. We introduce a novel defined predictive uncertainty to reweight the loss function, along with strategic local-to-global bundle adjustment. Experiments on synthetic and real-world datasets demonstrate that our system achieves state-of-the-art tracking and mapping accuracy while maintaining real-time performance. It significantly improves over current methods with a 25% reduction in depth L1 error and a 66.86% completion rate within 1 cm on the Replica dataset, reflecting a more accurate reconstruction of thin structures. Project page: https://shaoxiang777.github.io/project/uni-slam/
Shaoxiang Wang, Yaxu Xie, Chun-Peng Chang, Christen Millerdurai, Alain Pagani, Didier Stricker
WACV6
2025 EventEgo3D++: 3D Human Motion Capture from A Head-Mounted Event Camera
abstract
Monocular egocentric 3D human motion capture remains a significant challenge, particularly under conditions of low lighting and fast movements, which are common in head-mounted device applications. Existing methods that rely on RGB cameras often fail under these conditions. To address these limitations, we introduce EventEgo3D++, the first approach that leverages a monocular event camera with a fisheye lens for 3D human motion capture. Event cameras excel in high-speed scenarios and varying illumination due to their high temporal resolution, providing reliable cues for accurate 3D human motion capture. EventEgo3D++ leverages the LNES representation of event streams to enable precise 3D reconstructions. We have also developed a mobile head-mounted device (HMD) prototype equipped with an event camera, capturing a comprehensive dataset that includes real event observations from both controlled studio environments and in-the-wild settings, in addition to a synthetic dataset. Additionally, to provide a more holistic dataset, we include allocentric RGB streams that offer different perspectives of the HMD wearer, along with their corresponding SMPL body model. Our experiments demonstrate that EventEgo3D++ achieves superior 3D accuracy and robustness compared to existing solutions, even in challenging conditions. Moreover, our method supports real-time 3D pose updates at a rate of 140Hz. This work is an extension of the EventEgo3D approach (CVPR 2024) and further advances the state of the art in egocentric 3D human motion capture. For more details, visit the project page at https://eventego3d.mpi-inf.mpg.de.
Christen Millerdurai, Hiroyasu Akada, Jian Wang 0042, Diogo C. Luvizon, Alain Pagani, Didier Stricker, Christian Theobalt, Vladislav Golyanik
Int. J. Comput. Vis.6
2024 G3FA: Geometry-guided GAN for Face Animation
Alireza Javanmardi, Alain Pagani, Didier Stricker
BMVC3
2024 MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
abstract
3D visual grounding involves matching natural language descriptions with their corresponding objects in 3D spaces. Existing methods often face challenges with accuracy in object recognition and struggle in interpreting complex linguistic queries, particularly with descriptions that involve multiple anchors or are view-dependent. In response, we present the MiKASA (Multi-Key-Anchor Scene-Aware) Transformer. Our novel end-to-end trained model integrates a self-attention-based scene-aware object encoder and an original multi-key-anchor technique, enhancing object recognition accuracy and the understanding of spatial relationships. Furthermore, MiKASA improves the explainability of decision-making, facilitating error diagnosis. Our model achieves the highest overall accuracy in the Referit3D challenge for both the Sr3D and Nr3D datasets, particularly excelling by a large margin in categories that require viewpoint-dependent descriptions. The source code and additional resources for this project are available on GitHub: https://github.com/dfki-av/MiKASA-3DVG
Chun-Peng Chang, Shaoxiang Wang, Alain Pagani, Didier Stricker
CVPR4
2024 HiPose: Hierarchical Binary Surface Encoding and Correspondence Pruning for RGB-D 6DoF Object Pose Estimation
abstract
In this work, we present a novel dense-correspondence method for 6DoF object pose estimation from a single RGB-D image. While many existing data-driven methods achieve impressive performance, they tend to be time-consuming due to their reliance on rendering-based refinement approaches. To circumvent this limitation, we present HiPose, which establishes 3D-3D correspondences in a coarse-to-fine manner with a hierarchical binary surface encoding. Unlike previous dense-correspondence methods, we estimate the correspondence surface by employing point-to-surface matching and iteratively constricting the surface until it becomes a correspondence point while gradually removing outliers. Extensive experiments on public benchmarks LM-O, YCB-V, and T-Less demonstrate that our method surpasses all refinement-free methods and is even on par with expensive refinement-based approaches. Crucially, our approach is computationally efficient and enables real-time critical applications with high accuracy requirements.
Yongliang Lin, Yongzhi Su, Praveen Nathan, Sandeep Inuganti, Yan Di, Martin Sundermeyer, Fabian Manhardt, Didier Stricker, Jason R. Rambach, Yu Zhang 0018
CVPR8
2024 Sparse Semi-DETR: Sparse Learnable Queries for Semi-Supervised Object Detection
abstract
In this paper, we address the limitations of the DETR-based semi-supervised object detection (SSOD) framework, particularly focusing on the challenges posed by the quality of object queries. In DETR-based SSOD, the one-to-one as-signment strategy provides inaccurate pseudo-labels, while the one-to-many assignments strategy leads to overlapping predictions. These issues compromise training efficiency and degrade model performance, especially in detecting small or occluded objects. We introduce Sparse Semi-DETR, a novel transformer-based, end-to-end semi-supervised object detection solution to overcome these challenges. Sparse Semi-DETR incorporates a Query Refinement Module to enhance the quality of object queries, significantly improving detection capabilities for small and partially obscured objects. Additionally, we integrate a Reliable Pseudo-Label Filtering Module that selectively filters high-quality pseudo-labels, thereby enhancing detection accuracy and consistency. On the MS-COCO and Pascal VOC object detection benchmarks, Sparse Semi-DETR achieves a significant improvement over current state-of-the-art methods that highlight Sparse Semi-DETR's effectiveness in semi-supervised object detection, particularly in challenging scenarios involving small or partially obscured objects.
Tahira Shehzadi, Khurram Azeem Hashmi, Didier Stricker, Muhammad Zeshan Afzal
CVPR3
2024 SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and its Downstream Tasks
abstract
Scene graphs have been recently introduced into 3D spatial understanding as a comprehensive representation of the scene. The alignment between 3D scene graphs is the first step of many downstream tasks such as scene graph aided point cloud registration, mosaicking, overlap checking, and robot navigation. In this work, we treat 3D scene graph alignment as a partial graph-matching problem and propose to solve it with a graph neural network. We reuse the geometric features learned by a point cloud registration method and associate the clustered point-level geometric features with the node-level semantic feature via our designed feature fusion module. Partial matching is enabled by using a learnable method to select the top-k similar node pairs. Subsequent downstream tasks such as point cloud registration are achieved by running a pre-trained registration network within the matched regions. We further propose a point-matching rescoring method, that uses the node-wise alignment of the 3D scene graph to reweight the matching candidates from a pre-trained point cloud registration method. It reduces the false point correspondences estimated especially in low-overlapping cases. Experiments show that our method improves the alignment accuracy by lOrv20% in low-overlap and random transformation scenarios and outperforms the existing work in multiple downstream tasks. Our code and models are available here.
Yaxu Xie, Alain Pagani, Didier Stricker
CVPR3
2024 Enhanced Bank Check Security: Introducing a Novel Dataset and Transformer-Based Approach for Detection and Verification
Muhammad Saif Ullah Khan, Tahira Shehzadi, Rabeya Noor, Didier Stricker, Muhammad Zeshan Afzal
DAS4
2024 UnSupDLA: Towards Unsupervised Document Layout Analysis
Talha Uddin Sheikh, Tahira Shehzadi, Khurram Azeem Hashmi, Didier Stricker, Muhammad Zeshan Afzal
DAS4
2024 CLEO: Continual Learning of Evolving Ontologies
Shishir Muralidhara, Saqib Bukhari, Georg Schneider 0008, Didier Stricker, René Schuster
ECCV (54)4
2024 In-Domain Inversion for Improved 3D Face Alignment on Asymmetrical Expressions
abstract
Facial landmark detection, often termed as face alignment, is a well-studied research problem in computer vision. Nonetheless, face alignment on asymmetrical expressions has been overlooked in the literature, particularly for unusual gestures observed in individuals with unilateral facial paralysis. In this paper, we explore in-domain inversion in a semi-supervised approach for face alignment and target the detection of 3D landmarks on symmetrical and extremely asymmetrical facial expressions due to paralysis. Our approach first leverages unlabeled face data to synthesize face images, while learning a compressed representation in the latent space. Then, it integrates in-domain inversion in the self-supervised stage, to make the latent space semantically meaningful. This is exploited in the supervised stage by a 2D face landmark detector, trained on labeled data. Finally, we extend the pipeline to 3D face alignment and regress the depth coordinate from the intermediate latent space and the predicted 2D landmarks. We evaluate and compare our method to related work on publicly available datasets, and demonstrate that our approach outperforms the state of the art in the detection of 3D facial landmarks in our newly introduced dataset of facial paralysis, ParFace. Our implementation and dataset are available at https://github.com/jilliam/ParFace.
Jilliam María Díaz Barros, Jason R. Rambach, Pramod Murthy, Didier Stricker
FG4
2024 SynthSL: Expressive Humans for Sign Language Image Synthesis
abstract
Around 5% of the world's population live with disabling hearing loss. Despite recent advancements to improve accessibility to the Deaf community, research on sign language is still limited. In this work, we introduce a large-scale synthetic dataset on sign language, SynthSL, targeted to sign language production, recognition and translation. Using state-of-the-start methods for human body modelling, SynthSL aims to augment current datasets by providing additional ground truth data such as depth and normal maps, rendered models, segmentation masks and 2D/3D body joints. We additionally explore a generative architecture for the synthesis of sign images and propose a new generator based on Swin Transformers, conditioned on given body poses and appearance. We believe that an increase on the publicly available data on sign language would boost research and close the performance gap with related topics on human body synthesis. Our code, models and dataset are available at https://github.com/jilliam/SynthSL.
Jilliam María Díaz Barros, Jameel Malik, Abdalla Arafa, Didier Stricker
FG5
2024 Embedding Layout in Text for Document Understanding Using Large Language Models
Mohammad Minouei, Mohammad Reza Soheili, Didier Stricker
ICDAR (1)3
2024 A Hybrid Approach for Document Layout Analysis in Document Images
Tahira Shehzadi, Didier Stricker, Muhammad Zeshan Afzal
ICDAR (4)2
2024 Towards End-to-End Semi-supervised Table Detection with Semantic Aligned Matching Transformer
Tahira Shehzadi, Shalini Sarode, Didier Stricker, Muhammad Zeshan Afzal
ICDAR (5)3
2024 CICA: Content-Injected Contrastive Alignment for Zero-Shot Document Image Classification
Sankalp Sinha, Muhammad Saif Ullah Khan, Talha Uddin Sheikh, Didier Stricker, Muhammad Zeshan Afzal
ICDAR (4)4
2024 GenFormer - Generated Images Are All You Need to Improve Robustness of Transformers on Small Datasets
Sven Oehri, Nikolas Ebert, Didier Stricker, Oliver Wasenmüller
ICPR (2)4
2024 ShapeAug: Occlusion Augmentation for Event Camera Data
Katharina Bendig, René Schuster, Didier Stricker
ICPRAM3
2024 ShapeAug++: More Realistic Shape Augmentation for Event Data
Katharina Bendig, René Schuster, Didier Stricker
ICPRAM3
2024 CaRaCTO: Robust Camera-Radar Extrinsic Calibration with Triple Constraint Optimization
Mahdi Chamseddine, Jason R. Rambach, Didier Stricker
ICPRAM3
2024 Learned Fusion: 3D Object Detection Using Calibration-Free Transformer Feature Fusion
Michael Fürst, Rahul Jakkamsetty, René Schuster, Didier Stricker
ICPRAM4
2024 Achieving RGB-D Level Segmentation Performance from a Single ToF Camera
Pranav Sharma, Jigyasa Katrolia, Jason R. Rambach, Bruno Mirbach, Didier Stricker
ICPRAM5
2024 Text2CAD: Generating Sequential CAD Designs from Beginner-to-Expert Level Text Prompts
abstract
Prototyping complex computer-aided design (CAD) models in modern softwares can be very time-consuming. This is due to the lack of intelligent systems that can quickly generate simpler intermediate parts. We propose Text2CAD, the first AI framework for generating text-to-parametric CAD models using designer-friendly instructions for all skill levels. Furthermore, we introduce a data annotation pipeline for generating text prompts based on natural language instructions for the DeepCAD dataset using Mistral and LLaVA-NeXT. The dataset contains $\sim170$K models and $\sim660$K text annotations, from abstract CAD descriptions (e.g., _generate two concentric cylinders_) to detailed specifications (e.g., _draw two circles with center_ $(x,y)$ and _radius_ $r_{1}$, $r_{2}$, \textit{and extrude along the normal by} $d$...). Within the Text2CAD framework, we propose an end-to-end transformer-based auto-regressive network to generate parametric CAD models from input texts. We evaluate the performance of our model through a mixture of metrics, including visual quality, parametric precision, and geometrical accuracy. Our proposed framework shows great potential in AI-aided design applications. Project page is available at https://sadilkhan.github.io/text2cad-project/.
Mohammad Sadil Khan, Sankalp Sinha, Talha Uddin Sheikh, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
NeurIPS4
2024 RMS-FlowNet++: Efficient and Robust Multi-scale Scene Flow Estimation for Large-Scale Point Clouds
abstract
Abstract The proposed RMS-FlowNet++ is a novel end-to-end learning-based architecture for accurate and efficient scene flow estimation that can operate on high-density point clouds. For hierarchical scene flow estimation, existing methods rely on expensive Farthest-Point-Sampling (FPS) to sample the scenes, must find large correspondence sets across the consecutive frames and/or must search for correspondences at a full input resolution. While this can improve the accuracy, it reduces the overall efficiency of these methods and limits their ability to handle large numbers of points due to memory requirements. In contrast to these methods, our architecture is based on an efficient design for hierarchical prediction of multi-scale scene flow. To this end, we develop a special flow embedding block that has two advantages over the current methods: First, a smaller correspondence set is used, and second, the use of Random-Sampling (RS) is possible. In addition, our architecture does not need to search for correspondences at a full input resolution. Exhibiting high accuracy, our RMS-FlowNet++ provides a faster prediction than state-of-the-art methods, avoids high memory requirements and enables efficient scene flow on dense point clouds of more than 250K points at once. Our comprehensive experiments verify the accuracy of RMS-FlowNet++ on the established FlyingThings3D data set with different point cloud densities and validate our design choices. Furthermore, we demonstrate that our model has a competitive ability to generalize to the real-world scenes of the KITTI data set without fine-tuning.
Ramy Battrawy, René Schuster, Didier Stricker
Int. J. Comput. Vis.3
2024 End-to-end semi-supervised approach with modulated object queries for table detection in documents
Iqraa Ehsan, Tahira Shehzadi, Didier Stricker, Muhammad Zeshan Afzal
Int. J. Document Anal. Recognit.3
2023 EgoFlowNet: Non-Rigid Scene Flow from Point Clouds with Ego-Motion Support
Ramy Battrawy, René Schuster, Didier Stricker
BMVC3
2023 On Motion Artifacts Arising when Integrating Inertial Sensors into Loose Clothing Such as a Working Jacket
abstract
Inertial human motion capture (IHMC) has become a robust tool to estimate human kinematics in the wild such as industrial facilities. In contrast to optical motion capture, where occlusions might take place, the kinematics of a worker can be continuously provided. This is for instance a prerequisite for an ergonomic assessments of the workers. State-of-the-art IHMC solutions require inertial sensors to be tightly attached to body segments. This requires an additional setup time and lowers the practicability and ease of use when it comes to an industrial application. In contrast, sensors integrated into loose clothing such as a working jacket, may yield corrupted kinematics estimates due to the additional motion of loose clothing. In this work we present a study of orientations deviations obtained from kinematics estimates using tightly attached inertial sensors and into a working jacket integrated ones. We performed a quantitative analysis using data from the two hardware setups worn by 19 subjects performing different industry related tasks and measures of their body shapes. Using this data we approximated probability distributions of the deviation angles for each person and body segment. Applying different statistical measures we could gain insights to questions like, how severe orientation deviations are, if there is an influence of body shapes on the distribution and how probability distributions of the deviation angles can indicate physical motion limitations of a sensor attached to a segment.
Michael Lorenz, Rebecca Keilhauer, Takayuki Akiyama, Takehiro Niikura, Didier Stricker, Bertram Taetz, Gabriele Bleser-Taetz
CoDIT5
2023 I2MVFormer: Large Language Model Generated Multi-View Document Supervision for Zero-Shot Image Classification
abstract
Recent works have shown that unstructured text (doc-uments) from online sources can serve as useful auxiliary information for zero-shot image classification. However, these methods require access to a high-quality source like Wikipedia and are limited to a single source of information. Large Language Models (LLM) trained on web-scale text show impressive abilities to repurpose their learned knowledge for a multitude of tasks. In this work, we provide a novel perspective on using an LLM to provide text supervision for a zero-shot image classification model. The LLM is provided with a few text descriptions from different annota-tors as examples. The LLM is conditioned on these exam-ples to generate multiple text descriptions for each class (re-ferred to as views). Our proposed model, I2MVFormer, learns multi-view semantic embeddings for zero-shot image classification with these class views. We show that each text view of a class provides complementary information allowing a model to learn a highly discriminative class embed-ding. Moreover, we show that I2MVFormer is better at consuming the multi-view text supervision from LLM compared to baseline models. I2MVFormer establishes a new state-of-the-art on three public benchmark datasets for zero-shot image classification with unsupervised semantic embeddings. Code available at https://github.com/ferjad/I2DFormer
Muhammad Ferjad Naeem, Muhammad Gul Zain Ali Khan, Yongqin Xian, Muhammad Zeshan Afzal, Didier Stricker, Luc Van Gool, Federico Tombari
CVPR5
2023 U-RED: Unsupervised 3D Shape Retrieval and Deformation for Partial Point Clouds
abstract
In this paper, we propose U-RED, an Unsupervised shape REtrieval and Deformation pipeline that takes an arbitrary object observation as input, typically captured by RGB images or scans, and jointly retrieves and deforms the geometrically similar CAD models from a pre-established database to tightly match the target. Considering existing methods typically fail to handle noisy partial observations, U-RED is designed to address this issue from two aspects. First, since one partial shape may correspond to multiple potential full shapes, the retrieval method must allow such an ambiguous one-to-many relationship. Thereby U-RED learns to project all possible full shapes of a partial target onto the surface of a unit sphere. Then during inference, each sampling on the sphere will yield a feasible retrieval. Second, since real-world partial observations usually contain noticeable noise, a reliable learned metric that measures the similarity between shapes is necessary for stable retrieval. In U-RED, we design a novel point-wise residual-guided metric that allows noise-robust comparison. Extensive experiments on the synthetic datasets PartNet, ComplementMe and the real-world dataset Scan2CAD demonstrate that U-RED surpasses existing state-of-the-art approaches by 47.3%, 16.7% and 31.6% respectively under Chamfer Distance.
Yan Di, Chenyangguang Zhang, Ruida Zhang, Fabian Manhardt, Yongzhi Su, Jason R. Rambach, Didier Stricker, Xiangyang Ji, Federico Tombari
ICCV7
2023 FeatEnHancer: Enhancing Hierarchical Features for Object Detection and Beyond Under Low-Light Vision
abstract
Extracting useful visual cues for the downstream tasks is especially challenging under low-light vision. Prior works create enhanced representations by either correlating visual quality with machine perception or designing illumination-degrading transformation methods that require pre-training on synthetic datasets. We argue that optimizing enhanced image representation pertaining to the loss of the downstream task can result in more expressive representations. Therefore, in this work, we propose a novel module, FeatEnHancer, that hierarchically combines multiscale features using multi-headed attention guided by task-related loss function to create suitable representations. Furthermore, our intra-scale enhancement improves the quality of features extracted at each scale or level, as well as combines features from different scales in a way that reflects their relative importance for the task at hand. FeatEnHancer is a general-purpose plug-and-play module and can be incorporated into any low-light vision pipeline. We show with extensive experimentation that the enhanced representation produced with FeatEnHancer significantly and consistently improves results in several low-light vision tasks, including dark object detection (+5.7 mAP on ExDark), face detection (+1.5 mAP on DARK FACE), nighttime semantic segmentation (+5.1 mIoU on ACDC), and video object detection (+1.8 mAP on DarkVision), highlighting the effectiveness of enhancing hierarchical features under low-light vision.
Khurram Azeem Hashmi, Goutham Kallempudi, Didier Stricker, Muhammad Zeshan Afzal
ICCV3
2023 Introducing Language Guidance in Prompt-based Continual Learning
abstract
Continual Learning aims to learn a single model on a sequence of tasks without having access to data from previous tasks. The biggest challenge in the domain still remains catastrophic forgetting: a loss in performance on seen classes of earlier tasks. Some existing methods rely on an expensive replay buffer to store a chunk of data from previous tasks. This, while promising, becomes expensive when the number of tasks becomes large or data can not be stored for privacy reasons. As an alternative, prompt-based methods have been proposed that store the task information in a learnable prompt pool. This prompt pool instructs a frozen image encoder on how to solve each task. While the model faces a disjoint set of classes in each task in this setting, we argue that these classes can be encoded to the same embedding space of a pre-trained language encoder. In this work, we propose Language Guidance for Prompt-based Continual Learning (LGCL) as a plug-in for prompt-based methods. LGCL is model agnostic and introduces language guidance at the task level in the prompt pool and at the class level on the output feature of the vision encoder. We show with extensive experimentation that LGCL consistently improves the performance of prompt-based continual learning methods to set a new state-of-the art. LGCL achieves these performance improvements without needing any additional learnable parameters.
Muhammad Gul Zain Ali Khan, Muhammad Ferjad Naeem, Luc Van Gool, Didier Stricker, Federico Tombari, Muhammad Zeshan Afzal
ICCV4
2023 Towards End-to-End Semi-Supervised Table Detection with Deformable Transformer
Tahira Shehzadi, Khurram Azeem Hashmi, Didier Stricker, Marcus Liwicki, Muhammad Zeshan Afzal
ICDAR (2)3
2023 Confidence-Aware Clustered Landmark Filtering For Hybrid 3D Face Tracking
abstract
The detection of facial landmarks in 2D images has received a great attention in the last decade, as it is a key step for several computer-vision-related applications. Most of the approaches are focused on still images, and are extended to videos by using a tracking-by-detection scheme. In this work, we propose a frame-to-frame tracking module based on grouped-landmark Kalman filters that can be integrated into existing deep-learning-based 3D face alignment pipelines. This method improves the landmark accuracy in cases with large occlusion, extreme head poses and blurriness that affect existing approaches. Our experiments on the Menpo 3DA-2D benchmark show improvements on model-free and 3D-model-based face alignment approaches.
Jilliam María Díaz Barros, Didier Stricker, Jason R. Rambach
ICIP3
2023 On the Future of Training Spiking Neural Networks
Katharina Bendig, René Schuster, Didier Stricker
ICPRAM3
2023 Multi-task Fusion for Efficient Panoptic-Part Segmentation
Sravan Kumar Jagadeesh, René Schuster, Didier Stricker
ICPRAM3
2023 EvLiDAR-Flow: Attention-Guided Fusion Between Point Clouds and Events for Scene Flow Estimation
Ankit Sonthalia, Ramy Battrawy, René Schuster, Didier Stricker
ICPRAM4
2023 Severity of Catastrophic Forgetting in Object Detection for Autonomous Driving
Christian Witte, René Schuster, Syed Saqib Bukhari, Patrick Trampert, Didier Stricker, Georg Schneider 0008
ICPRAM5
2023 Structure PLP-SLAM: Efficient Sparse Mapping and Localization using Point, Line and Plane for Monocular, RGB-D and Stereo Cameras
abstract
This paper presents a visual SLAM system that uses both points and lines for robust camera localization, and simultaneously performs a piece-wise planar reconstruction (PPR) of the environment to provide a structural map in real-time. One of the biggest challenges in parallel tracking and mapping with a monocular camera is to keep the scale consistent when reconstructing the geometric primitives. This further introduces difficulties in graph optimization of the bundle adjustment (BA) step. We solve these problems by proposing several run-time optimizations on the reconstructed lines and planes. Our system is able to run with depth and stereo sensors in addition to the monocular setting. Our proposed SLAM tightly incorporates the semantic and geometric features to boost both frontend pose tracking and backend map optimization. We evaluate our system exhaustively on various datasets, and show that we outperform state-of-the-art methods in terms of trajectory precision. The code of PLP-SLAM has been made available in open-source for the research community (https://github.com/PeterFWS/Structure-PLP-SLAM).
Fangwen Shu, Alain Pagani, Didier Stricker
ICRA4
2023 THOR-Net: End-to-end Graformer-based Realistic Two Hands and Object Reconstruction with Self-supervision
abstract
Realistic reconstruction of two hands interacting with objects is a new and challenging problem that is essential for building personalized Virtual and Augmented Reality environments. Graph Convolutional networks (GCNs) allow for the preservation of the topologies of hands poses and shapes by modeling them as a graph. In this work, we propose the THOR-Net which combines the power of GCNs, Transformer, and self-supervision to realistically reconstruct two hands and an object from a single RGB image. Our network comprises two stages; namely the features extraction stage and the reconstruction stage. In the features extraction stage, a Keypoint RCNN is used to ex-tract 2D poses, features maps, heatmaps, and bounding boxes from a monocular RGB image. Thereafter, this 2D information is modeled as two graphs and passed to the two branches of the reconstruction stage. The shape re-construction branch estimates meshes of two hands and an object using our novel coarse-to-fine GraFormer shape network. The 3D poses of the hands and objects are re-constructed by the other branch using a GraFormer network. Finally, a self-supervised photometric loss is used to directly regress the realistic textured of each vertex in the hands’ meshes. Our approach achieves State-of-the-art results in Hand shape estimation on the HO-3D dataset (10.0mm) exceeding ArtiBoost (10.8mm). It also surpasses other methods in hand pose estimation on the challenging two hands and object (H2O) dataset by 5mm on the left-hand pose and 1 mm on the right-hand pose. THOR-Net code will be available at https://github.com/ATAboukhadra/THOR-Net.
Ahmed Tawfik Aboukhadra, Jameel Malik, Ahmed Elhayek, Nadia Robertini, Didier Stricker
WACV5
2023 Attribution-aware Weight Transfer: A Warm-Start Initialization for Class-Incremental Semantic Segmentation
abstract
In class-incremental semantic segmentation (CISS), deep learning architectures suffer from the critical problems of catastrophic forgetting and semantic background shift. Although recent works focused on these issues, existing classifier initialization methods do not address the background shift problem and assign the same initialization weights to both background and new foreground class classifiers. We propose to address the background shift with a novel classifier initialization method which employs gradient-based attribution to identify the most relevant weights for new classes from the classifier’s weights for the previous background and transfers these weights to the new classifier. This warm-start weight initialization provides a general solution applicable to several CISS methods. Furthermore, it accelerates learning of new classes while mitigating forgetting. Our experiments demonstrate significant improvement in mIoU compared to the state-of-the-art CISS methods on the Pascal-VOC 2012, ADE20K and Cityscapes datasets.
Dipam Goswami, René Schuster, Joost van de Weijer 0001, Didier Stricker
WACV4
2023 BoxMask: Revisiting Bounding Box Supervision for Video Object Detection
abstract
We present a new, simple yet effective approach to up- lift video object detection. We observe that prior works operate on instance-level feature aggregation that imminently neglects the refined pixel-level representation, resulting in confusion among objects sharing similar appearance or motion characteristics. To address this limitation, we propose BoxMask, which effectively learns discriminative representations by incorporating class-aware pixel-level information. We simply consider bounding box-level annotations as a coarse mask for each object to supervise our method. The proposed module can be effortlessly integrated into any region-based detector to boost detection. Extensive experiments on ImageNet VID and EPIC KITCHENS datasets demonstrate consistent and significant improvement when we plug our BoxMask module into numerous recent state-of-the-art methods. The code will be available at https://github.com/khurramHashmi/BoxMask.
Khurram Azeem Hashmi, Alain Pagani, Didier Stricker, Muhammad Zeshan Afzal
WACV3
2023 Learning Attention Propagation for Compositional Zero-Shot Learning
abstract
Compositional zero-shot learning aims to recognize unseen compositions of seen visual primitives of object classes and their states. While all primitives (states and objects) are observable during training in some combination, their complex interaction makes this task especially hard. For example, wet changes the visual appearance of a dog very differently from a bicycle. Furthermore, we argue that relationships between compositions go beyond shared states or objects. A cluttered office can contain a busy table; even though these compositions don’t share a state or object, the presence of a busy table can guide the presence of a cluttered office. We propose a novel method called Compositional Attention Propagated Embedding (CAPE) as a solution. The key intuition to our method is that a rich dependency structure exists between compositions arising from complex interactions of primitives in addition to other dependencies between compositions. CAPE learns to identify this structure and propagates knowledge between them to learn class embedding for all seen and unseen compositions. In the challenging generalized compositional zero-shot setting, we show that our method outperforms previous baselines to set a new state-of-the-art on three publicly available benchmarks.
Muhammad Gul Zain Ali Khan, Muhammad Ferjad Naeem, Luc Van Gool, Alain Pagani, Didier Stricker, Muhammad Zeshan Afzal
WACV5
2022 Spatio-Temporal Learnable Proposals for End-to-End Video Object Detection
Khurram Azeem Hashmi, Didier Stricker, Muhammad Zeshan Afzal
BMVC2
2022 SOMSI: Spherical Novel View Synthesis with Soft Occlusion Multi-Sphere Images
abstract
Spherical novel view synthesis (SNVS) is the task of estimating 360○views at dynamic novel views given a set of 360○input views. Prior arts learn multi-sphere image (MSI) representations that enable fast rendering times but are only limited to modelling low-dimensional color values. Modelling high-dimensional appearance features in MSI can result in better view synthesis, but it is not feasible to represent high-dimensional features in a large number (> 64) of MSI spheres. We propose a novel MSI representation called Soft Occlusion MSI (SOMSI) that enables modelling high-dimensional appearance features in MSI while retaining the fast rendering times of a standard MSI. Our key insight is to model appearance features in a smaller set (e.g. 3) of occlusion levels instead of larger number of MSI levels. Experiments on both synthetic and real-world scenes demonstrate that using SOMSI can provide a good balance between accuracy and run-time. SOMSI can produce considerably better results compared to MSI based MODS [1], while having similar fast rendering time. SOMSI view synthesis quality is on-par with state-of-the-art NeRF [24] like model while being 2 orders of magnitude faster. For code, additional results and data, please visit https://tedyhabtegebrial.github.io/somsi.
Tewodros Habtegebrial, Christiano Couto Gava, Marcel Rogge, Didier Stricker, Varun Jampani
CVPR4
2022 ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose Estimation
abstract
Establishing correspondences from image to 3D has been a key task of 6DoF object pose estimation for a long time. To predict pose more accurately, deeply learned dense maps replaced sparse templates. Dense methods also improved pose estimation in the presence of occlusion. More recently researchers have shown improvements by learning object fragments as segmentation. In this work, we present a discrete descriptor, which can represent the object surface densely. By incorporating a hierarchical binary grouping, we can encode the object surface very efficiently. Moreover, we propose a coarse to fine training strategy, which enables fine-grained correspondence prediction. Finally, by matching predicted codes with object surface and using a PnP solver, we estimate the 6DoF pose. Results on the public LM-O and YCB-V datasets show major improvement over the state of the art w.r.t. ADD(-S) metric, even surpassing RGB-D based methods in some cases.
Yongzhi Su, Mahdi Saleh, Torben Fetzer, Jason R. Rambach, Nassir Navab, Benjamin Busam, Didier Stricker, Federico Tombari
CVPR7
2022 Controlling Continuous Locomotion in Virtual Reality with Bare Hands Using Hand Gestures
abstract
Abstract Moving around in a virtual world is one of the essential interactions for Virtual Reality (VR) applications. The current standard for moving in VR is using a controller. Recently, VR Head Mounted Displays integrate new input modalities such as hand tracking which allows the investigation of different techniques to move in VR. This work explores different techniques for bare-handed locomotion since it could offer a promising alternative to existing freehand techniques. The presented techniques enable continuous movement through an immersive virtual environment. The proposed techniques are compared to each other in terms of efficiency, usability, perceived workload, and user preference.
Alexander Schäfer 0001, Gerd Reis, Didier Stricker
EuroXR3
2022 Towards Inertial Human Motion Tracking with Drift-Free Absolute Orientations using only Sparse Sources of Heading Information
Michael Lorenz, Gabriele Bleser-Taetz, Didier Stricker, Bertram Taetz
FUSION3
2022 Self-Superflow: Self-Supervised Scene Flow Prediction in Stereo Sequences
abstract
In recent years, deep neural networks showed their exceeding capabilities in addressing many computer vision tasks including scene flow prediction. However, most of the advances are dependent on the availability of a vast amount of dense per pixel ground truth annotations, which are very difficult to obtain for real life scenarios. Therefore, synthetic data is often relied upon for supervision, resulting in a representation gap between the training and test data. Even though a great quantity of unlabeled real world data is available, there is a huge lack in self-supervised methods for scene flow prediction. Hence, we explore the extension of a self-supervised loss based on the Census transform and occlusion-aware bidirectional displacements for the problem of scene flow prediction. Regarding the KITTI scene flow benchmark, our method outperforms the corresponding supervised pre-training of the same network and shows improved generalization capabilities while achieving much faster convergence.
Katharina Bendig, René Schuster, Didier Stricker
ICIP3
2022 Autoencoder Attractors for Uncertainty Estimation
abstract
The reliability assessment of a machine learning model’s prediction is an important quantity for the deployment in safety critical applications. Not only can it be used to detect novel sceneries, either as out-of-distribution or anomaly sample, but it also helps to determine deficiencies in the training data distribution. A lot of promising research directions have either proposed traditional methods like Gaussian processes or extended deep learning based approaches, for example, by interpreting them from a Bayesian point of view. In this work we propose a novel approach for uncertainty estimation based on autoencoder models: The recursive application of a previously trained autoencoder model can be interpreted as a dynamical system storing training examples as attractors. While input images close to known samples will converge to the same or similar attractor, input samples containing unknown features are unstable and converge to different training samples by potentially removing or changing characteristic features. The use of dropout during training and inference leads to a family of similar dynamical systems, each one being robust on samples close to the training distribution but unstable on new features. Either the model reliably removes these features or the resulting instability can be exploited to detect problematic input samples. We evaluate our approach on several dataset combinations as well as on an industrial application for occupant classification in the vehicle interior for which we additionally release a new synthetic dataset.
Steve Dias Da Cruz, Bertram Taetz, Thomas Stifter, Didier Stricker
ICPR4
2022 Autoencoder for Synthetic to Real Generalization: From Simple to More Complex Scenes
abstract
Learning on synthetic data and transferring the resulting properties to their real counterparts is an important challenge for reducing costs and increasing safety in machine learning. In this work, we focus on autoencoder architectures and aim at learning latent space representations that are invariant to inductive biases caused by the domain shift between simulated and real images showing the same scenario. We train on synthetic images only, present approaches to increase generalizability and improve the preservation of the semantics to real datasets of increasing visual complexity. We show that pre-trained feature extractors (e.g. VGG) can be sufficient for generalization on images of lower complexity, but additional improvements are required for visually more complex scenes. To this end, we demonstrate a new sampling technique, which matches semantically important parts of the image, while randomizing the other parts, leads to salient feature extraction and a neglection of unimportant parts. This helps the generalization to real data and we further show that our approach outperforms fine-tuned classification models.
Steve Dias Da Cruz, Bertram Taetz, Thomas Stifter, Didier Stricker
ICPR4
2022 Object Permanence in Object Detection Leveraging Temporal Priors at Inference Time
abstract
Object permanence is the concept that objects do not suddenly disappear in the physical world. Humans understand this concept at young ages and know that another person is still there, even though it is temporarily occluded. Neural networks currently often struggle with this challenge. Thus, we introduce explicit object permanence into two stage detection approaches drawing inspiration from particle filters. At the core, our detector uses the predictions of previous frames as additional proposals for the current one at inference time. Experiments confirm the feedback loop improving detection performance by a up to 10.3 mAP with little computational overhead.Our approach is suited to extend two-stage detectors for stabilized and reliable detections even under heavy occlusion. Additionally, the ability to apply our method without retraining an existing model promises wide application in real-world tasks.
Michael Fürst, Priyash Bhugra, René Schuster, Didier Stricker
ICPR4
2022 RMS-FlowNet: Efficient and Robust Multi-Scale Scene Flow Estimation for Large-Scale Point Clouds
abstract
The proposed RMS-FlowNet is a novel end-to-end learning-based architecture for accurate and efficient scene flow estimation which can operate on point clouds of high density. For hierarchical scene flow estimation, the existing methods depend on either expensive Farthest-Point-Sampling (FPS) or structure-based scaling which decrease their ability to handle a large number of points. Unlike these methods, we base our fully supervised architecture on Random-Sampling (RS) for multiscale scene flow prediction. To this end, we propose a novel flow embedding design which can predict more robust scene flow in conjunction with RS. Exhibiting high accuracy, our RMS-FlowNet provides a faster prediction than state-of-the-art methods and works efficiently on consecutive dense point clouds of more than 250K points at once. Our comprehensive experiments verify the accuracy of RMS-FlowNet on the established FlyingThings3D data set with different point cloud densities and validate our design choices. Additionally, we show that our model presents a competitive ability to generalize towards the real-world scenes of KITTI data set without fine-tuning.
Ramy Battrawy, René Schuster, Mohammad-Ali Nikouei Mahani, Didier Stricker
ICRA4
2022 Towards Artefact Aware Human Motion Capture using Inertial Sensors Integrated into Loose Clothing
abstract
Inertial motion capture has become an attractive alternative to optical motion capture for human joint angle estimation outside the laboratory. Usually inertial sensors are assumed to be tightly fixed to the body segments, which can be cumbersome regarding setup-time and ease-of-use. However, integrating the sensors directly into loose clothing, usually, results in additional clothing motion relative to the motion of the underlying bones that should be captured. In this work we propose the Difference Mapping distributions approach that corrects the segment orientations of a given inertial motion capture system that assumes tightly coupled sensors. The approach allows to reduce the joint angle errors due to clothing artefacts by at least 77.2% for people with similar morphology performing a similar task as seen in the training data, including an ergonomic assessments scenario at work places with ten participants. Moreover, we show that the uncertainty of the distribution can be used to measure the reliability of the predicted map if e.g. the motion is further away from the training data to allow for an artefact aware inertial motion tracking approach. The experimental data for this study is available online under [1].
Michael Lorenz, Gabriele Bleser-Taetz, Takayuki Akiyama, Takehiro Niikura, Didier Stricker, Bertram Taetz
ICRA5
2022 Classification of Manual Versus Autonomous Driving based on Machine Learning of Eye Movement Patterns
abstract
Recent advances in autonomous driving systems raise new questions about how to enhance the communication and takeover control between the system and the driver. Eye tracking technologies have shown their feasibility to recognize whether the driver’s gaze is directed ‘on-road’ or ‘off-road’. However, this binary information alone is not sufficient to infer the driver’s engagement in the driving task. In the present work, we take the next step and investigate how driving modes (autopilot, navigation system, and printed map) associated with different levels of engagement can be categorized from drivers’ gaze patterns. Using gaze data recorded in these three driving tasks along with several state-of-the-art machine learning methods, we demonstrate that the driving modes are associated with different gaze patterns. We achieved an average accuracy of 90.1% for binary and 80.3% for multi-class driving mode classification. Our findings pave the way for enhancing driver monitoring systems in (semi-) autonomous cars.
Iuliia Brishtel, Stephan Krauß, Jason R. Rambach, Igor Vozniak, Didier Stricker
SMC6
2022 Fusion Point Pruning for Optimized 2D Object Detection with Radar-Camera Fusion
abstract
Object detection is one of the most important perception tasks for advanced driver assistant systems and autonomous driving. Due to its complementary features and moderate cost, radar-camera fusion is of particular interest in the automotive industry but comes with the challenge of how to optimally fuse the heterogeneous data sources. To solve this for 2D object detection, we propose two new techniques to project the radar detections onto the image plane, exploiting additional uncertainty information. We also introduce a new technique called fusion point pruning, which automatically finds the best fusion points of radar and image features in the neural network architecture. These new approaches combined surpass the state of the art in 2D object detection performance for radar-camera fusion models, evaluated with the nuScenes dataset. We further find that the utilization of radar-camera fusion is especially beneficial for night scenes.
Lukas Stäcker, Philipp Heidenreich, Jason R. Rambach, Didier Stricker
WACV4
2022 HandVoxNet++: 3D Hand Shape and Pose Estimation Using Voxel-Based Neural Networks
abstract
3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. Existing methods addressing it directly regress hand meshes via 2D convolutional neural networks, which leads to artifacts due to perspective distortions in the images. To address the limitations of the existing methods, we develop HandVoxNet++, i.e., a voxel-based deep network with 3D and graph convolutions trained in a fully supervised manner. The input to our network is a 3D voxelized-depth-map-based on the truncated signed distance function (TSDF). HandVoxNet++ relies on two hand shape representations. The first one is the 3D voxelized grid of hand shape, which does not preserve the mesh topology and which is the most accurate representation. The second representation is the hand surface that preserves the mesh topology. We combine the advantages of both representations by aligning the hand surface to the voxelized hand shape either with a new neural Graph-Convolutions-based Mesh Registration (GCN-MeshReg) or classical segment-wise Non-Rigid Gravitational Approach (NRGA++) which does not rely on training data. In extensive evaluations on three public benchmarks, i.e., SynHand5M, depth-based HANDS19 challenge and HO-3D, the proposed HandVoxNet++ achieves the state-of-the-art performance. In this journal extension of our previous approach presented at CVPR 2020, we gain 41.09% and 13.7% higher shape alignment accuracy on SynHand5M and HANDS19 datasets, respectively. Our method is ranked first on the HANDS19 challenge dataset (Task 1: Depth-Based 3D Hand Pose Estimation) at the moment of the submission of our results to the portal in August 2020.
Jameel Malik, Soshi Shimada, Ahmed Elhayek, Sk Aziz Ali, Christian Theobalt, Vladislav Golyanik, Didier Stricker
IEEE Trans. Pattern Anal. Mach. Intell.7
2021 TICaM: A Time-of-flight In-car Cabin Monitoring Dataset
Jigyasa Katrolia, Ahmed El-Sherif, Hartmut Feld, Bruno Mirbach, Jason R. Rambach, Didier Stricker
BMVC6
2021 PlaneRecNet: Multi-Task Learning with Cross-Task consistency for Piece-Wise Plane Detection and Reconstruction from a Single RGB Image
Yaxu Xie, Fangwen Shu, Jason R. Rambach, Alain Pagani, Didier Stricker
BMVC5
2021 Simultaneous Bi-directional Structured Light Encoding for Practical Uncalibrated Profilometry
Torben Fetzer, Gerd Reis, Didier Stricker
CAIP (1)3
2021 Joint Global ICP for Improved Automatic Alignment of Full Turn Object Scans
Torben Fetzer, Gerd Reis, Didier Stricker
CAIP (1)3
2021 Fast Projector-Driven Structured Light Matching in Sub-pixel Accuracy Using Bilinear Interpolation Assumption
Torben Fetzer, Gerd Reis, Didier Stricker
CAIP (1)3
2021 RPSRNet: End-to-End Trainable Rigid Point Set Registration Network Using Barnes-Hut 2D-Tree Representation
abstract
We propose RPSRNet - a novel end-to-end trainable deep neural network for rigid point set registration. For this task, we use a novel 2D-tree representation for the input point sets and a hierarchical deep feature embedding in the neural network. An iterative transformation refinement module of our network boosts the feature matching accuracy in the intermediate stages. We achieve an inference speed of ∼12-15 ms to register a pair of input point clouds as large as ∼250K. Extensive evaluations on (i) KITTI LiDAR-odometry and (ii) ModelNet-40 datasets show that our method outperforms prior state-of-the-art methods – e.g., on the KITTI dataset, DCP-v2 by 1.3 and 1.5 times, and PointNetLK by 1.8 and 1.9 times better rotational and translational accuracy respectively. Evaluation on ModelNet40 shows that RPSRNet is more robust than other benchmark methods when the samples contain a significant amount of noise and disturbance. RPSRNet accurately registers point clouds with non-uniform sampling densities, e.g., LiDAR data, which cannot be processed by many existing deep-learning-based registration methods.
Sk Aziz Ali, Kerem Kahraman, Gerd Reis, Didier Stricker
CVPR4
2021 Pose Tracking vs. Pose Estimation of AR Glasses with Convolutional, Recurrent, and Non-local Neural Networks: A Comparison
Ahmet Firintepe, Sarfaraz Habib, Alain Pagani, Didier Stricker
EuroXR4
2021 Semantic Segmentation in Depth Data: A Comparative Evaluation of Image and Point Cloud Based Methods
abstract
The problem of semantic segmentation from depth images can be addressed by segmenting directly in the image domain or at 3D point cloud level. In this paper, we attempt for the first time to provide a study and experimental comparison of the two approaches. Through experiments on three datasets, namely SUN RGB-D, NYUdV2 and TICaM, we extensively compare various semantic segmentation algorithms, the input to which includes images and point clouds derived from them. Based on this, we offer analysis of the performance and computational cost of these algorithms that can provide guidelines on when each method should be preferred.
Jigyasa Katrolia, Lars Krämer, Jason R. Rambach, Bruno Mirbach, Didier Stricker
ICIP5
2021 A Blended Attention-CTC Network Architecture for Amharic Text-image Recognition
Birhanu Belay, Tewodros Habtegebrial, Marcus Liwicki, Gebeyehu Belay, Didier Stricker
ICPRAM5
2021 PlaneSegNet: Fast and Robust Plane Estimation Using a Single-stage Instance Segmentation CNN
abstract
Instance segmentation of planar regions in indoor scenes benefits visual SLAM and other applications such as augmented reality (AR) where scene understanding is required. Existing methods built upon two-stage frameworks show satisfactory accuracy but are limited by low frame rates. In this work, we propose a real-time deep neural architecture that estimates piece-wise planar regions from a single RGB image. Our model employs a variant of a fast single-stage CNN architecture to segment plane instances. Considering the particularity of the target detected, we propose Fast Feature Non-maximum Suppression (FF-NMS) to reduce the suppression errors resulted from overlapping bounding boxes of planes. We also utilize a Residual Feature Augmentation module in the Feature Pyramid Network (FPN) . Our method achieves significantly higher frame-rates and comparable segmentation accuracy against two-stage methods. We automatically label over 70,000 images as ground truth from the Stanford 2D-3D-Semantics dataset. Moreover, we incorporate our method with a state-of-the-art planar SLAM and validate its benefits.
Yaxu Xie, Jason R. Rambach, Fangwen Shu, Didier Stricker
ICRA4
2021 Autoencoder Based Inter-Vehicle Generalization for In-Cabin Occupant Classification
abstract
Common domain shift problem formulations consider the integration of multiple source domains, or the target domain during training. Regarding the generalization of machine learning models between different car interiors, we formulate the criterion of training in a single vehicle: without access to the target distribution of the vehicle the model would be deployed to, neither with access to multiple vehicles during training. We performed an investigation on the SVIRO dataset for occupant classification on the rear bench and propose an autoencoder based approach to improve the transferability. The autoencoder is on par with commonly used classification models when trained from scratch and sometimes out-performs models pre-trained on a large amount of data. Moreover, the autoencoder can transform images from unknown vehicles into the vehicle it was trained on. These results are corroborated by an evaluation on real infrared images from two vehicle interiors.
Steve Dias Da Cruz, Bertram Taetz, Oliver Wasenmüller, Thomas Stifter, Didier Stricker
IV5
2021 Illumination Normalization by Partially Impossible Encoder-Decoder Cost Function
abstract
Images recorded during the lifetime of computer vision based systems undergo a wide range of illumination and environmental conditions affecting the reliability of previously trained machine learning models. Image normalization is hence a valuable preprocessing component to enhance the models' robustness. To this end, we introduce a new strategy for the cost function formulation of encoder-decoder networks to average out all the unimportant information in the input images (e.g. environmental features and illumination changes) to focus on the reconstruction of the salient features (e.g. class instances). Our method exploits the availability of identical sceneries under different illumination and environmental conditions for which we formulate a partially impossible reconstruction target: the input image will not convey enough information to reconstruct the target in its entirety. Its applicability is assessed on three publicly available datasets. We combine the triplet loss as a regular- izer in the latent space representation and a nearest neighbour search to improve the generalization to unseen illuminations and class instances. The importance of the aforementioned post-processing is highlighted on an automotive application. To this end, we release a synthetic dataset of sceneries from three different passenger compartments where each scenery is rendered under ten different illumination and environmental conditions: https://sviro.kl.dfki.de.
Steve Dias Da Cruz, Bertram Taetz, Thomas Stifter, Didier Stricker
WACV4
2021 A Deep Temporal Fusion Framework for Scene Flow Using a Learnable Motion Model and Occlusions
abstract
Motion estimation is one of the core challenges in computer vision. With traditional dual-frame approaches, occlusions and out-of-view motions are a limiting factor, especially in the context of environmental perception for vehicles due to the large (ego-) motion of objects. Our work pro-poses a novel data-driven approach for temporal fusion of scene flow estimates in a multi-frame setup to overcome the issue of occlusion. Contrary to most previous methods, we do not rely on a constant motion model, but instead learn a generic temporal relation of motion from data. In a second step, a neural network combines bi-directional scene flow estimates from a common reference frame, yielding a refined estimate and a natural byproduct of occlusion masks. This way, our approach provides a fast multi-frame extension for a variety of scene flow estimators, which outperforms the underlying dual-frame approaches.
René Schuster, Christian Unger, Didier Stricker
WACV3
2021 SSGP: Sparse Spatial Guided Propagation for Robust and Generic Interpolation
abstract
Interpolation of sparse pixel information towards a dense target resolution finds its application across multiple disciplines in computer vision. State-of-the-art interpolation of motion fields applies model-based interpolation that makes use of edge information extracted from the target image. For depth completion, data-driven learning approaches are widespread. Our work is inspired by latest trends in depth completion that tackle the problem of dense guidance for sparse information. We extend these ideas and create a generic cross-domain architecture that can be applied for a multitude of interpolation problems like optical flow, scene flow, or depth completion. In our experiments, we show that our proposed concept of Sparse Spatial Guided Propagation (SSGP) achieves improvements to robustness, accuracy, or speed compared to specialized algorithms.
René Schuster, Oliver Wasenmüller, Christian Unger, Didier Stricker
WACV4
2021 SLAM in the Field: An Evaluation of Monocular Mapping and Localization on Challenging Dynamic Agricultural Environment
abstract
This paper demonstrates a system capable of combining a sparse, indirect, monocular visual SLAM, with both offline and real-time Multi-View Stereo (MVS) reconstruction algorithms. This combination overcomes many obstacles encountered by autonomous vehicles or robots employed in agricultural environments, such as overly repetitive patterns, need for very detailed reconstructions, and abrupt movements caused by uneven roads. Furthermore, the use of a monocular SLAM makes our system much easier to integrate with an existing device, as we do not rely on a LiDAR (which is expensive and power consuming), or stereo camera (whose calibration is sensitive to external perturbation e.g. camera being displaced). To the best of our knowledge, this paper presents the first evaluation results for monocular SLAM, and our work further explores unsupervised depth estimation on this specific application scenario by simulating RGB-D SLAM to tackle the scale ambiguity, and shows our approach produces reconstructions that are helpful to various agricultural tasks. Moreover, we highlight that our experiments provide meaningful insight to improve monocular SLAM systems under agricultural settings.
Fangwen Shu, Paul Lesur, Yaxu Xie, Alain Pagani, Didier Stricker
WACV5
2020 Intrinsic Dynamic Shape Prior for Dense Non-Rigid Structure from Motion
abstract
While dense non-rigid structure from motion (NRSfM) has been extensively studied from the perspective of the reconstructability problem over the recent years, almost no attempts have been undertaken to bring it into the practical realm. The reasons for the slow dissemination are the severe ill-posedness, high sensitivity to motion and deformation cues, and the difficulty to obtain reliable point tracks in the vast majority of practical scenarios.To fill this gap, we propose a new framework that first extracts prior knowledge from an input image sequence with NRSfM. Our Dynamic Shape Prior Reconstruction (DSPR) approach then uses the obtained 3D reconstructions as a dynamic shape prior for sequential surface recovery in scenarios with recurrence. DSPR can be combined with existing dense NRSfM techniques while its energy functional is optimised with multi-start gradient descent at real-time frame rates for new incoming point tracks. The proposed versatile framework with a new core NRSfM approach outperforms several other methods in the ability to handle inaccurate and noisy point tracks, provided we have access to a representative (in terms of the deformation variety) image sequence. Comprehensive experiments highlight convergence properties and the accuracy of DSPR under different disturbing effects. We also perform a joint study of tracking and reconstruction and show applications to shape compression and heart reconstruction under occlusions. We achieve state-of-the-art metrics (accuracy and compression ratios) in different scenarios.
Vladislav Golyanik, André Jonas, Didier Stricker, Christian Theobalt
3DV3
2020 HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation From a Single Depth Map
abstract
3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. The state-of-the-art methods directly regress 3D hand meshes from 2D depth images via 2D convolutional neural networks, which leads to artefacts in the estimations due to perspective distortions in the images. In contrast, we propose a novel architecture with 3D convolutions trained in a weakly-supervised manner. The input to our method is a 3D voxelized depth map, and we rely on two hand shape representations. The first one is the 3D voxelized grid of the shape which is accurate but does not preserve the mesh topology and the number of mesh vertices. The second representation is the 3D hand surface which is less accurate but does not suffer from the limitations of the first representation. We combine the advantages of these two representations by registering the hand surface to the voxelized hand shape. In the extensive experiments, the proposed approach improves over the state of the art by47.8% on the SynHand5M dataset. Moreover, our augmentation policy for voxelized depth maps further enhances the accuracy of 3D hand pose estimation on real data. Our method produces visually more reasonable and realistic hand shapes on NYU and BigHand2.2M datasets compared to the existing approaches.
Jameel Malik, Ibrahim Abdelaziz, Ahmed Elhayek, Soshi Shimada, Sk Aziz Ali, Vladislav Golyanik, Christian Theobalt, Didier Stricker
CVPR8
2020 Foldmatch: Accurate and High Fidelity Garment Fitting Onto 3D Scans
abstract
In this paper, we propose a new template fitting method that can capture fine details of garments in target 3D scans of dressed human bodies. Matching the high fidelity details of such loose/tight-fit garments is a challenging task as they express intricate folds, creases, wrinkle patterns, and other high fidelity surface details. Our proposed method of non-rigid shape fitting - FoldMatch - uses physics-based particle dynamics to explicitly model the deformation of loose-fit garments and wrinkle vector fields for capturing clothing details. The 3D scan point cloud behaves as a collection of astrophysical particles, which attracts the points in template mesh and defines the template motion model. We use this point-based motion model to derive regularized deformation gradients for the template mesh. We show the parameterization of the wrinkle vector fields helps in the accurate shape fitting. Our method shows better performance than the state-of-the-art methods. We define several deformation and shape matching quality measurement metrics to evaluate FoldMatch on synthetic and real data sets.
Sk Aziz Ali, Sikang Yan, Wolfgang Dornisch, Didier Stricker
ICIP4
2020 Ghost Target Detection in 3D Radar Data using Point Cloud based Deep Neural Network
abstract
Ghost targets are targets that appear at wrong locations in radar data and are caused by the presence of multiple indirect reflections between the target and the sensor. In this work, we introduce the first point based deep learning approach for ghost target detection in 3D radar point clouds. This is done by extending the PointNet network architecture by modifying its input to include radar point features beyond location and introducing skip connetions. We compare different input modalities and analyze the effects of the changes we introduced. We also propose an approach for automatic labeling of ghost targets 3D radar data using lidar as reference. The algorithm is trained and tested on real data in various driving scenarios and the tests show promising results in classifying real and ghost radar targets.
Mahdi Chamseddine, Jason R. Rambach, Didier Stricker, Oliver Wasenmüller
ICPR3
2020 HPERL: 3D Human Pose Estimation from RGB and LiDAR
abstract
In-the-wild human pose estimation has a huge potential for various fields, ranging from animation and action recognition to intention recognition and prediction for autonomous driving. The current state-of-the-art is focused only on RGB and RGB-D approaches for predicting the 3D human pose. However, not using precise LiDAR depth information limits the performance and leads to very inaccurate absolute pose estimation. With LiDAR sensors becoming more affordable and common on robots and autonomous vehicle setups, we propose an end-to-end architecture using RGB and LiDAR to predict the absolute 3D human pose with unprecedented precision. Additionally, we introduce a weakly-supervised approach to generate 3D predictions using 2D pose annotations from PedX [1]. This allows for many new opportunities in the field of 3D human pose estimation.
Michael Fürst, Shriya T. P. Gupta, René Schuster, Oliver Wasenmüller, Didier Stricker
ICPR5
2020 ResFPN: Residual Skip Connections in Multi-Resolution Feature Pyramid Networks for Accurate Dense Pixel Matching
abstract
Dense pixel matching is required for many computer vision algorithms such as disparity, optical flow or scene flow estimation. Feature Pyramid Networks (FPN) have proven to be a suitable feature extractor for CNN-based dense matching tasks. FPN generates well localized and semantically strong features at multiple scales. However, the generic FPN is not utilizing its full potential, due to its reasonable but limited localization accuracy. Thus, we present ResFPN - a multi-resolution feature pyramid network with multiple residual skip connections, where at any scale, we leverage the information from higher resolution maps for stronger and better localized features. In our ablation study, we demonstrate the effectiveness of our novel architecture with clearly higher accuracy than FPN. In addition, we verify the superior accuracy of ResFPN in many different pixel matching applications on established datasets like KITTI, Sintel, and FlyingThings3D.
Rishav, René Schuster, Ramy Battrawy, Oliver Wasenmüller, Didier Stricker
ICPR5
2020 Using Automatic Features for Text-image Classification in Amharic Documents
Birhanu Belay, Tewodros Habtegebrial, Gebeyehu Belay, Didier Stricker
ICPRAM4
2020 DeepLiDARFlow: A Deep Learning Architecture For Scene Flow Estimation Using Monocular Camera and Sparse LiDAR
abstract
Scene flow is the dense 3D reconstruction of motion and geometry of a scene. Most state-of-the-art methods use a pair of stereo images as input for full scene reconstruction. These methods depend a lot on the quality of the RGB images and perform poorly in regions with reflective objects, shadows, ill-conditioned light environment and so on. LiDAR measurements are much less sensitive to the aforementioned conditions but LiDAR features are in general unsuitable for matching tasks due to their sparse nature. Hence, using both LiDAR and RGB can potentially overcome the individual disadvantages of each sensor by mutual improvement and yield robust features which can improve the matching process. In this paper, we present DeepLiDARFlow, a novel deep learning architecture which fuses high level RGB and LiDAR features at multiple scales in a monocular setup to predict dense scene flow. Its performance is much better in the critical regions where image-only and LiDAR-only methods are inaccurate. We verify our DeepLiDARFlow using the established data sets KITTI and FlyingThings3D and we show strong robustness compared to several state-of-the-art methods which used other input modalities. The code of our paper is available at https://github.com/dfki-av/DeepLiDARFlow.
Rishav, Ramy Battrawy, René Schuster, Oliver Wasenmüller, Didier Stricker
IROS5
2020 The More, the Merrier? A Study on In-Car IR-based Head Pose Estimation
abstract
Deep learning methods have proven useful for head pose estimation, but the effect of their depth, type and input resolution based on infrared (IR) images still need to be explored. In this paper, we present a study on in-car head pose estimation on the IR images of the AutoPOSE dataset, where we extract 64 x 64 and 128 x 128 pixel cropped head images. We propose the novel networks Head Orientation Network (HON) and ResNetHG and compare them with state-of-the-art methods like the HPN model from DriveAHead on different input resolutions. In addition, we evaluate multiple depths within our HON and ResNetHG networks and their effect on the accuracy. Our experiments show that higher resolution images lead to lower estimation errors. Furthermore, we show that deep learning methods with fewer layers perform better on head orientation regression based on IR images. Our HON and ResNetHG18 architectures outperform the state-of-the-art on IR images on four different metrics, where we achieve a reduction of the residual error of up to 74%.
Ahmet Firintepe, Mohamed Selim, Alain Pagani, Didier Stricker
IV4
2020 Generative View Synthesis: From Single-view Semantics to Novel-view Images
abstract
Content creation, central to applications such as virtual reality, can be tedious and time-consuming. Recent image synthesis methods simplify this task by offering tools to generate new views from as little as a single input image, or by converting a semantic map into a photorealistic image. We propose to push the envelope further, and introduce Generative View Synthesis (GVS) that can synthesize multiple photorealistic views of a scene given a single semantic map. We show that the sequential application of existing techniques, e.g., semantics-to-image translation followed by monocular view synthesis, fail at capturing the scene's structure. In contrast, we solve the semantics-to-image translation in concert with the estimation of the 3D layout of the scene, thus producing geometrically consistent novel views that preserve semantic structures. We first lift the input 2D semantic map onto a 3D layered representation of the scene in feature space, thereby preserving the semantic labels of 3D geometric structures. We then project the layered features onto the target views to generate the final novel-view images. We verify the strengths of our method and compare it with several advanced baselines on three different datasets. Our approach also allows for style manipulation and image editing operations, such as the addition or removal of objects, with simple manipulations of the input style images and semantic maps respectively. For code and additional results, visit the project page at https://gvsnet.github.io
Tewodros Habtegebrial, Varun Jampani, Orazio Gallo, Didier Stricker
NeurIPS4
2020 HMDPose: A large-scale trinocular IR Augmented Reality Glasses Pose Dataset
abstract
Augmented Reality Glasses usually implement an Inside-Out tracking. In case of a driving scenario or glasses with less computation capabilities, an Outside-In tracking approach is required. However, to the best of our knowledge, no public datasets exist that collects images of users wearing AR glasses. To address this problem, we present HMDPose, an infrared trinocular dataset of four different AR Head-mounted displays captured in a car. It contains sequences of 14 subjects captured by three different cameras running at 60 FPS each, adding up to more than 3,000,000 labeled images in total. We provide a ground truth 6DoF-pose, captured by a submillimeter accurate marker-based tracker. We make HMDPose publicly available for non-profit, academic use and non-commercial benchmarking on ags.cs.uni-kl.de/datasets/hmdpose/.
Ahmet Firintepe, Alain Pagani, Didier Stricker
VRST3
2020 SVIRO: Synthetic Vehicle Interior Rear Seat Occupancy Dataset and Benchmark
abstract
We release SVIRO, a synthetic dataset for sceneries in the passenger compartment of ten different vehicles, in order to analyze machine learning-based approaches for their generalization capacities and reliability when trained on a limited number of variations (e.g. identical backgrounds and textures, few instances per class). This is in contrast to the intrinsically high variability of common benchmark datasets, which focus on improving the state-of-the-art of general tasks. Our dataset contains bounding boxes for object detection, instance segmentation masks, keypoints for pose estimation and depth images for each synthetic scenery as well as images for each individual seat for classification. The advantage of our use-case is twofold: The proximity to a realistic application to benchmark new approaches under novel circumstances while reducing the complexity to a more tractable environment, such that applications and theoretical questions can be tested on a more challenging dataset as toy problems. The data and evaluation server are available under https://sviro.kl.dfki.de.
Steve Dias Da Cruz, Oliver Wasenmüller, Hans-Peter Beise, Thomas Stifter, Didier Stricker
WACV5
2020 Stable Intrinsic Auto-Calibration from Fundamental Matrices of Devices with Uncorrelated Camera Parameters
abstract
Auto-Calibration is an important task in computer vision and is necessary for many visual applications. Methods like photogrammetry, depth map estimation, metrology, augmented/mixed reality or odometry are strongly dependent on well calibrated devices. While classical calibration relies on tools like checkerboards or additional scene information, auto-calibration only takes epipolar relations into account. Classical calibration is often impractical, tends to de-adjust over time and distributes the error over the entire, limited working volume. Auto-calibration, on the other hand, does not require any information other than the image content itself, has a virtually unlimited working range and usually achieves highest accuracy at the objects' surfaces. Unfortunately, auto-calibration methods are sensitive to errors in the fundamental matrix and need good initialization to converge to the global solution. In practice this leads to difficulties if optical parameters like principal point or focal length are unconstrained. In such situations, even state-of- the-art auto-calibration methods tend to diverge and do not yield a valid calibration. This work assesses reasons for this behavior, in particular for the initialization method of Bougnoux [3] and Lourakis' state-of-the-art auto-calibration method [21]. Based on the analysis, a more stable method is proposed. A continuous and smooth energy functional is introduced, providing superior convergence properties. I.e. it can not diverge, converges faster, and has a significantly enlarged convergence region with respect to the global minimum. Finally, a thorough evaluation has been conducted and a detailed comparison with the state of the art is presented.
Torben Fetzer, Gerd Reis, Didier Stricker
WACV3
2020 SceneFlowFields++: Multi-frame Matching, Visibility Prediction, and Robust Interpolation for Scene Flow Estimation
René Schuster, Oliver Wasenmüller, Christian Unger, Georg Kuschk, Didier Stricker
Int. J. Comput. Vis.5
2019 DispVoxNets: Non-Rigid Point Set Alignment with Supervised Learning Proxies
abstract
We introduce a supervised-learning framework for nonrigid point set alignment of a new kind - Displacements on Voxels Networks (DispVoxNets) - which abstracts away from the point set representation and regresses 3D displacement fields on regularly sampled proxy 3D voxel grids. Thanks to recently released collections of deformable objects with known intra-state correspondences, DispVoxNets learn a deformation model and further priors (e.g., weak point topology preservation) for different object categories such as cloths, human bodies and faces. DispVoxNets cope with large deformations, noise and clustered outliers more robustly than the state-of-the-art. At test time, our approach runs orders of magnitude faster than previous techniques. All properties of DispVoxNets are ascertained numerically and qualitatively in extensive experiments and comparisons to several previous methods.
Soshi Shimada, Vladislav Golyanik, Edith Tretschk, Didier Stricker, Christian Theobalt
3DV4
2019 Convolutional Recurrent Neural Network for Bubble Detection in a Portable Continuous Bladder Irrigation Monitor
Xiaoying Tan, Gerd Reis, Didier Stricker
AIME3
2019 A Compact Light Field Camera for Real-Time Depth Estimation
Yuriy Anisimov, Oliver Wasenmüller, Didier Stricker
CAIP (1)3
2019 SDC - Stacked Dilated Convolution: A Unified Descriptor Network for Dense Matching Tasks
abstract
Dense pixel matching is important for many computer vision tasks such as disparity and flow estimation. We present a robust, unified descriptor network that considers a large context region with high spatial variance. Our network has a very large receptive field and avoids striding layers to maintain spatial resolution. These properties are achieved by creating a novel neural network layer that consists of multiple, parallel, stacked dilated convolutions (SDC). Several of these layers are combined to form our SDC descriptor network. In our experiments, we show that our SDC features outperform state-of-the-art feature descriptors in terms of accuracy and robustness. In addition, we demonstrate the superior performance of SDC in state-of-the-art stereo matching, optical flow and scene flow algorithms on several famous public benchmarks.
René Schuster, Oliver Wasenmüller, Christian Unger, Didier Stricker
CVPR4
2019 Inertial Motion Capture Using Adaptive Sensor Fusion and Joint Angle Drift Correction
Hammad T. Butt, Manthan Pancholi, Mathias Musahl, Pramod Murthy, Maria Alejandra Sanchez, Didier Stricker
FUSION6
2019 Accelerated Gravitational Point Set Alignment With Altered Physical Laws
abstract
This work describes Barnes-Hut Rigid Gravitational Approach (BH-RGA) - a new rigid point set registration method relying on principles of particle dynamics. Interpreting the inputs as two interacting particle swarms, we directly minimise the gravitational potential energy of the system using non-linear least squares. Compared to solutions obtained by solving systems of second-order ordinary differential equations, our approach is more robust and less dependent on the parameter choice. We accelerate otherwise exhaustive particle interactions with a Barnes-Hut tree and efficiently handle massive point sets in quasilinear time while preserving the globally multiply-linked character of interactions. Among the advantages of BH-RGA is the possibility to define boundary conditions or additional alignment cues through varying point masses. Systematic experiments demonstrate that BH-RGA surpasses performances of baseline methods in terms of the convergence basin and accuracy when handling incomplete, noisy and perturbed data. The proposed approach also positively compares to the competing method for the alignment with prior matches.
Vladislav Golyanik, Christian Theobalt, Didier Stricker
ICCV3
2019 Amharic Text Image Recognition: Database, Algorithm, and Analysis
abstract
This paper introduces a dataset for an exotic, but very interesting script, Amharic. Amharic follows a unique syllabic writing system which uses 33 consonant characters with their 7 vowels variants of each. Some labialized characters derived by adding diacritical marks on consonants and or removing part of it. These associated diacritics on consonant characters are relatively smaller in size and challenging to distinguish the derived (vowel and labialized) characters. In this paper we tackle the problem of Amharic text-line image recognition. In this work, we propose a recurrent neural network based method to recognize Amharic text-line images. The proposed method uses Long Short Term Memory (LSTM) networks together with CTC (Connectionist Temporal Classification). Furthermore, in order to overcome the lack of annotated data, we introduce a new dataset that contains 337,332 Amharic text-line images which is made freely available at http://www.dfki.uni-kl.de/~belay/. The performance of the proposed Amharic OCR model is tested by both printed and synthetically generated datasets, and promising results are obtained.
Birhanu Belay, Tewodros Habtegebrial, Marcus Liwicki, Gebeyehu Belay, Didier Stricker
ICDAR5
2019 Face It!: A Pipeline for Real-Time Performance-Driven Facial Animation
abstract
This paper presents a new lightweight approach for real-time performance-driven facial animation from monocular videos. We transfer facial expressions from 2D images to a 3D virtual character, by estimating the rigid head pose and non-rigid face deformation from detected and tracked 2D facial landmarks. We map the input face into the facial expression space of the 3D head model using blendshape models and formulate a lightweight energy-based optimization problem, which is solved by non-linear least squares at 18 FPS on a single CPU. Our method robustly handles varying head poses and different facial expressions, including moderately asymmetric ones. Compared to related methods, our approach does not require training data, specialised camera setups or graphics cards, and is suitable for embedded systems. We support our claims with several experiments.
Jilliam María Díaz Barros, Vladislav Golyanik, Kiran Varanasi, Didier Stricker
ICIP4
2019 Factored Convolutional Neural Network for Amharic Character Image Recognition
abstract
In this paper we propose a novel CNN based approach for Amharic character image recognition. The proposed method is designed by leveraging the structure of Amharic graphemes. Amharic characters could be decomposed in to a consonant and a vowel. As a result of this consonant-vowel combination structure, Amharic characters lie within a matrix structure called 'Fidel Gebeta'. The rows and columns of 'Fidel Gebeta' correspond to a character's consonant and the vowel components, respectively. The proposed method has a CNN architecture with two classifiers that detect the row/consonant and column/vowel components of a character. The two classifiers share a common feature space before they fork-out at their last layers. The method achieves state-of-the-art result on a synthetically generated dataset. The proposed method achieves 94.97% overall character recognition accuracy.
Birhanu Belay, Tewodros Habtegebrial, Marcus Liwicki, Gebeyehu Belay, Didier Stricker
ICIP5
2019 Simple Domain Adaptation for CAD based Object Recognition
abstract
We present a simple method of domain adaptation between synthetic images and real images - by high quality rendering of the 3D models and correlation alignment. Using this method, we solve the problem of 3D object recognition in 2D images by fine-tuning existing pretrained CNN models for the object categories using the rendered images. Experimentally, we show that our rendering pipeline along with the correlation alignment improve the recognition accuracy of existing CNN based recognition trained on rendered images - by a canonical renderer - by a large margin. Using the same idea we present a general image classifier of common objects which is trained only on the 3D models from the publicly available databases, and show that a small number of training models are sufficient to capture different variations within and across the classes.
Kripasindhu Sarkar, Didier Stricker
ICPRAM2
2019 LiDAR-Flow: Dense Scene Flow Estimation from Sparse LiDAR and Stereo Images
abstract
We propose a new approach called LiDAR-Flow to robustly estimate a dense scene flow by fusing a sparse LiDAR with stereo images. We take the advantage of the high accuracy of LiDAR to resolve the lack of information in some regions of stereo images due to textureless objects, shadows, ill-conditioned light environment and many more. Additionally, this fusion can overcome the difficulty of matching unstructured 3D points between LiDAR-only scans. Our LiDAR-Flow approach consists of three main steps; each of them exploits LiDAR measurements. First, we build strong seeds from LiDAR to enhance the robustness of matches between stereo images. The imagery part seeks the motion matches and increases the density of scene flow estimation. Then, a consistency check employs LiDAR seeds to remove the possible mismatches. Finally, LiDAR measurements constraint the edge-preserving interpolation method to fill the remaining gaps. In our evaluation we investigate the individual processing steps of our LiDAR-Flow approach and demonstrate the superior performance compared to image-only approach.
Ramy Battrawy, René Schuster, Oliver Wasenmüller, Qing Rao, Didier Stricker
IROS5
2019 PWOC-3D: Deep Occlusion-Aware End-to-End Scene Flow Estimation
abstract
In the last few years, convolutional neural networks (CNNs) have demonstrated increasing success at learning many computer vision tasks including dense estimation problems such as optical flow and stereo matching. However, the joint prediction of these tasks, called scene flow, has traditionally been tackled using slow classical methods based on primitive assumptions which fail to generalize. The work presented in this paper overcomes these drawbacks efficiently (in terms of speed and accuracy) by proposing PWOC-3D, a compact CNN architecture to predict scene flow from stereo image sequences in an end-to-end supervised setting. Further, large motion and occlusions are well-known problems in scene flow estimation. PWOC-3D employs specialized design decisions to explicitly model these challenges. In this regard, we propose a novel self-supervised strategy to predict occlusions from images (learned without any labeled occlusion data). Leveraging several such constructs, our network achieves competitive results on the KITTI benchmark and the challenging FlyingThings3D dataset. Especially on KITTI, PWOC-3D achieves the second place among end-to-end deep learning methods with 48 times fewer parameters than the top-performing method.
Rohan Saxena, René Schuster, Oliver Wasenmüller, Didier Stricker
IV4
2019 DeLiO: Decoupled LiDAR Odometry
abstract
Most LiDAR odometry algorithms estimate the transformation between two consecutive frames by estimating the rotation and translation in an intervening fashion. In this paper, we propose our Decoupled LiDAR Odometry (DeLiO), which - for the first time - decouples the rotation estimation completely from the translation estimation. In particular, the rotation is estimated by extracting the surface normals from the input point clouds and tracking their characteristic pattern on a unit sphere. Using this rotation the point clouds are unrotated so that the underlying transformation is pure translation, which can be easily estimated using a line cloud approach. An evaluation is performed on the KITTI dataset and the results are compared against state-of-the-art algorithms.
Queens Maria Thomas, Oliver Wasenmüller, Didier Stricker
IV3
2019 Simple and effective deep hand shape and pose regression from a single depth image
Jameel Malik, Ahmed Elhayek, Fabrizio Nunnari, Didier Stricker
Comput. Graph.4
2019 Flow Fields: Dense Correspondence Fields for Highly Accurate Large Displacement Optical Flow Estimation
Christian Bailer, Bertram Taetz, Didier Stricker
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 NRGA: Gravitational Approach for Non-rigid Point Set Registration
abstract
Recovery of correspondences between point sets which differ by some non-rigid transformation is an ill-posed problem. Many existing methods underperform on noisy or corrupted input data. In this study, a novel physics-based approach -- Non-Rigid Gravitational Approach (NRGA) -- for non-rigid point set registration is introduced which is robust to the mentioned artifacts. Thereafter, a distributed N-body simulation and iterative Procrustes alignment non-rigidly transform and register the template point set. Furthermore, in the force field evolution, per-point Gaussian curvature serves as a shape matching descriptor whereas the displacement fields are regularized by coherent collective motion. The optimal alignment is referred to as the state of minimum gravitational potential energy between the point sets. A thorough experimental evaluation and comparison are provided with widely used state-of-the-art methods on 2D and 3D data sets. Experiments show NRGA's robustness against uniform outliers and missing data.
Sk Aziz Ali, Vladislav Golyanik, Didier Stricker
3DV3
2018 DeepHPS: End-to-end Estimation of 3D Hand Pose and Shape by Learning from Synthetic Depth
abstract
Articulated hand pose and shape estimation is an important problem for vision-based applications such as augmented reality and animation.In contrast to the existing methods which optimize only for joint positions, we propose a fully supervised deep network which learns to jointly estimate a full 3D hand mesh representation and pose from a single depth image.To this end, a CNN architecture is employed to estimate parametric representations i.e. hand pose, bone scales and complex shape parameters. Then, a novel hand pose and shape layer, embedded inside our deep framework, produces 3D joint positions and hand mesh. Lack of sufficient training data with varying hand shapes limits the generalized performance of learning based methods. Also, manually annotating real data is suboptimal. Therefore, we present SynHand5M: a million-scale synthetic benchmark with accurate joint annotations, segmentation masks and mesh files of depth maps. Among model based learning (hybrid) methods, we show improved results on two of the public benchmarks i.e. NYU and ICVL. Also, by employing a joint training strategy with real and synthetic data, we recover 3D hand mesh and pose from real images in 30ms.
Jameel Malik, Ahmed Elhayek, Fabrizio Nunnari, Kiran Varanasi, Kiarash Tamaddon, Alexis Héloir, Didier Stricker
3DV7
2018 Structured Low-Rank Matrix Factorization for Point-Cloud Denoising
abstract
In this work we address the problem of point-cloud denoising where we assume that a given point-cloud comprises (noisy) points that were sampled from an underlying surface that is to be denoised. We phrase the point-cloud denoising problem in terms of a dictionary learning framework. To this end, for a given point-cloud we (robustly) extract planar patches covering the entire point-cloud, where each patch contains a (noisy) description of the local structure of the underlying surface. Based on the general assumption that many of the local patches (in the noise-free point-cloud) contain redundant information (e.g. due to smoothness of the surface, or due to repetitive structures), we find a low-dimensional affine subspace that (approximately) explains the extracted (noisy) patches. Computationally, this is achieved by solving a structured low-rank matrix factorization problem, where we impose smoothness on the patch dictionary and sparsity on the coefficients. We experimentally demonstrate that our method outperforms existing denoising approaches in various noise scenarios.
Kripasindhu Sarkar, Florian Bernard 0001, Kiran Varanasi, Christian Theobalt, Didier Stricker
3DV5
2018 Learning 3D Shapes as Multi-layered Height-Maps Using 2D Convolutional Networks
Kripasindhu Sarkar, Basavaraj Hampiholi, Kiran Varanasi, Didier Stricker
ECCV (16)4
2018 Dense Scene Reconstruction from Spherical Light Fields
abstract
Spherical images are suitable for reconstruction of large, outdoor scenes. We present a method for dense 3D reconstruction based on spherical light fields (SLFs). In contrast to light fields generated from perspective cameras, SLFs are intrinsically distorted. Our approach is the first to exploit the internal structure of SLFs while properly handling distortion. Additionally, we adopt a variational formulation and our algorithm produces high quality, locally smooth reconstructions.
Christiano Couto Gava, Didier Stricker, Soichiro Yokota
ICIP2
2018 FlowFields++: Accurate Optical Flow Correspondences Meet Robust Interpolation
abstract
Optical Flow algorithms are of high importance for many applications. Recently, the Flow Field algorithm and its modifications have shown remarkable results, as they have been evaluated with top accuracy on different data sets. In our analysis of the algorithm we have found that it produces accurate sparse matches, but there is room for improvement in the interpolation. Thus, we propose in this paper FlowFields++, where we combine the accurate matches of Flow Fields with a robust interpolation. In addition, we propose improved variational optimization as post-processing. Our new algorithm is evaluated on the challenging KITTI and MPI Sintel data sets with public top results on both benchmarks.
René Schuster, Christian Bailer, Oliver Wasenmüller, Didier Stricker
ICIP4
2018 Improving Time-of-Flight Sensor for Specular Surfaces with Shape from Polarization
abstract
Time-of-Flight (ToF) sensors can obtain depth values for diffuse objects. However, the essential problem is that the sensor can not receive active light from specular surfaces due to specular reflections. In this paper, we propose a new depth reconstruction framework for specular objects that combines ToF cues and Shape from Polarization (SfP). To overcome the ill-posedness of SfP with a single view, we integrate superpixel segmentation with planarity constraints for every superpixel. Experimental results demonstrate the effectiveness of the depth reconstruction algorithm for both controlled environment data and real vehicle data in a parking area.
Tomonari Yoshida, Vladislav Golyanik, Oliver Wasenmüller, Didier Stricker
ICIP4
2018 Fusion of Keypoint Tracking and Facial Landmark Detection for Real-Time Head Pose Estimation
abstract
In this paper, we address the problem of extreme head pose estimation from intensity images, in a monocular setup. We introduce a novel fusion pipeline to integrate into a dedicated Kalman Filter the pose estimated from a tracking scheme in the prediction stage and the pose estimated from a detection scheme in the correction stage. To that end, the measurement covariance of the Kalman Filter is updated in every frame. The tracking scheme is performed using a set of keypoints extracted in the area of the head along with a simple 3D geometric model. The detection scheme, on the other hand, relies on the alignment of facial landmarks in each frame combined with 3D features extracted on a head mesh. The head pose in each scheme is estimated by minimizing the reprojection error from the 3D-2D correspondences. By combining both frameworks, we extend the applicability of head pose estimation from facial landmarks to cases where these features are no longer visible. We compared the proposed method to other related approaches, showing that it can achieve state-of-the-art performance. We also demonstrate that our approach is suitable for cases with extreme head rotations and (self-) occlusions, besides being suitable for real time applications.
Jilliam María Díaz Barros, Bruno Mirbach, Frederic Garcia, Kiran Varanasi, Didier Stricker
WACV5
2018 3D Shape Processing by Convolutional Denoising Autoencoders on Local Patches
abstract
We propose a system for surface completion and inpainting of 3D shapes using denoising autoencoders with convolutional layers, learnt on local patches. Our method uses height map based local patches parameterized using 3D mesh quadrangulation of the low resolution input shape. This provides us sufficient amount of local 3D patch dataset to learn deep generative Convolutional Neural Networks (CNNs) for the task of repairing moderate sized holes. We design generative networks specifically suited for the 3D encoding following ideas from the recent progress in 2D inpainting, and show our results to be better than the previous methods of surface inpainting that use linear dictionary. We validate our method on both synthetic shapes and real world scans.
Kripasindhu Sarkar, Kiran Varanasi, Didier Stricker
WACV3
2018 SceneFlowFields: Dense Interpolation of Sparse Scene Flow Correspondences
abstract
While most scene flow methods use either variational optimization or a strong rigid motion assumption, we show for the first time that scene flow can also be estimated by dense interpolation of sparse matches. To this end, we find sparse matches across two stereo image pairs that are detected without any prior regularization and perform dense interpolation preserving geometric and motion boundaries by using edge information. A few iterations of variational energy minimization are performed to refine our results, which are thoroughly evaluated on the KITTI benchmark and additionally compared to state-of-the-art on MPI Sintel. For application in an automotive context, we further show that an optional ego-motion model helps to boost performance and blends smoothly into our approach to produce a segmentation of the scene into static and dynamic parts.
René Schuster, Oliver Wasenmüller, Georg Kuschk, Christian Bailer, Didier Stricker
WACV5
2017 Fast and Efficient Depth Map Estimation from Light Fields
abstract
The paper presents an algorithm for depth map estimation from the light field images in relatively small amount of time, using only single thread on CPU. The proposed method improves existing principle of line fitting in 4- dimensional light field space. Line fitting is based on color values comparison using kernel density estimation. Our method utilizes result of Semi-Global Matching (SGM) with Census transform-based matching cost as a border initialization for line fitting. It provides a significant reduction of computations needed to find the best depth match. With the suggested evaluation metric we show that proposed method is applicable for efficient depth map estimation while preserving low computational time compared to others.
Yuriy Anisimov, Didier Stricker
3DV2
2017 Scalable Dense Monocular Surface Reconstruction
abstract
This paper reports on a novel template-free monocular non-rigid surface reconstruction approach. Existing techniques using motion and deformation cues rely on multiple prior assumptions, are often computationally expensive and do not perform equally well across the variety of data sets. In contrast, the proposed Scalable Monocular Surface Reconstruction (SMSR) combines strengths of several algorithms, i.e., it is scalable with the number of points, can handle sparse and dense settings as well as different types of motions and deformations. We estimate camera pose by singular value thresholding and proximal gradient. Our formulation adopts alternating direction method of multipliers which converges in linear time for large point track matrices. In the proposed SMSR, trajectory space constraints are integrated by smoothing of the measurement matrix. In the extensive experiments, SMSR is demonstrated to consistently achieve state-of-the-art accuracy on a wide variety of data sets.
Mohammad Dawud Ansari, Vladislav Golyanik, Didier Stricker
3DV3
2017 Multiframe Scene Flow with Piecewise Rigid Motion
abstract
We introduce a novel multiframe scene flow approach that jointly optimizes the consistency of the patch appearances and their local rigid motions from RGB-D image sequences. In contrast to the competing methods, we take advantage of an oversegmentation of the reference frame and robust optimization techniques. We formulate scene flow recovery as a global non-linear least squares problem which is iteratively solved by a damped Gauss-Newton approach. As a result, we obtain a qualitatively new level of accuracy in RGB-D based scene flow estimation which can potentially run in real-time. Our method can handle challenging cases with rigid, piecewise rigid, articulated and moderate non-rigid motion, and does not rely on prior knowledge about the types of motions and deformations. Extensive experiments on synthetic and real data show that our method outperforms state-of-the-art.
Vladislav Golyanik, Robert Maier 0001, Matthias Nießner, Didier Stricker, Jan Kautz
3DV5
2017 High Dimensional Space Model for Dense Monocular Surface Recovery
abstract
Dense surface reconstruction from monocular image sequences - known as Non-Rigid Structure from Motion (NRSfM) - is a highly ill-posed inverse problem. The objective of NRSfM is to learn 3D shapes from 2D point tracks in an unsupervised manner. While existing methods rely on low-rank models, we propose the concept of High Dimensional Space Model (HDSM). In HDSM, time-varying geometry is encoded by a high-dimensional static structure projected into different metric subspaces. To express non-rigid deformations, instead of directly modelling in the 3D space, we gradually increase space dimensionality as the complexity of the scene increases. HDSM allows for a compact representation with deformation localisation and can be interpreted as a generalisation of the previously proposed models for NRSfM. Relying on HDSM, we develop an algorithm for dense monocular surface recovery. Experiments show that the proposed method achieves high accuracy while allowing for the fine-grained control.
Vladislav Golyanik, Didier Stricker
3DV2
2017 Simultaneous Hand Pose and Skeleton Bone-Lengths Estimation from a Single Depth Image
abstract
Articulated hand pose estimation is a challenging task for human-computer interaction. The state-of-the-art hand pose estimation algorithms work only with one or a few subjects for which they have been calibrated or trained. Particularly, the hybrid methods based on learning followed by model fitting or model based deep learning do not explicitly consider varying hand shapes and sizes. In this work, we introduce a novel hybrid algorithm for estimating the 3D hand pose as well as bone-lengths of the hand skeleton at the same time, from a single depth image. The proposed CNN architecture learns hand pose parameters and scale parameters associated with the bone-lengths simultaneously. Subsequently, a new hybrid forward kinematics layer employs both parameters to estimate 3D joint positions of the hand. For end-to-end training, we combine three public datasets NYU, ICVL and MSRA-2015 in one unified format to achieve large variation in hand shapes and sizes. Among hybrid methods, our method shows improved accuracy over the state-of-the-art on the combined dataset and the ICVL dataset that contain multiple subjects. Also, our algorithm is demonstrated to work well with unseen images.
Jameel Malik, Ahmed Elhayek, Didier Stricker
3DV3
2017 Learning Quadrangulated Patches for 3D Shape Parameterization and Completion
abstract
We propose a novel 3D shape parameterization by surface patches, that are oriented by 3D mesh quadrangulation of the shape. By encoding 3D surface detail on local patches, we learn a patch dictionary that identifies principal surface features of the shape. Unlike previous methods, we are able to encode surface patches of variable size as determined by the user. We propose novel methods for dictionary learning and patch reconstruction based on the query of a noisy input patch with holes. We evaluate the patch dictionary towards various applications in 3D shape inpainting, denoising and compression. Our method is able to predict missing vertices and inpaint moderately sized holes. We demonstrate a complete pipeline for reconstructing the 3D mesh from the patch encoding. We validate our shape parameterization and reconstruction methods on both synthetic shapes and real world scans. We show that our patch dictionary performs successful shape completion of complicated surface textures.
Kripasindhu Sarkar, Kiran Varanasi, Didier Stricker
3DV3
2017 Fast dense feature extraction with convolutional neural networks that have pooling or striding layers
Christian Bailer, Tewodros Habtegebrial, Kiran Varanasi, Didier Stricker
BMVC4
2017 Introduction to Coherent Depth Fields for Dense Monocular Surface Recovery
Vladislav Golyanik, Torben Fetzer, Didier Stricker
BMVC3
2017 Sparse-MVRVMs Tree for Fast and Accurate Head Pose Estimation in the Wild
Mohamed Selim, Alain Pagani, Didier Stricker
CAIP (1)3
2017 CNN-Based Patch Matching for Optical Flow with Thresholded Hinge Embedding Loss
abstract
Learning based approaches have not yet achieved their full potential in optical flow estimation, where their performance still trails heuristic approaches. In this paper, we present a CNN based patch matching approach for optical flow estimation. An important contribution of our approach is a novel thresholded loss for Siamese networks. We demonstrate that our loss performs clearly better than existing losses. It also allows to speed up training by a factor of 2 in our tests. Furthermore, we present a novel way for calculating CNN based features for different image scales, which performs better than existing methods. We also discuss new ways of evaluating the robustness of trained features for the application of patch matching for optical flow. An interesting discovery in our paper is that low-pass filtering of feature maps can increase the robustness of features created by CNNs. We proved the competitive performance of our approach by submitting it to the KITTI 2012, KITTI 2015 and MPI-Sintel evaluation portals where we obtained state-of-the-art results on all three datasets.
Christian Bailer, Kiran Varanasi, Didier Stricker
CVPR3
2017 Hardware architecture of Bidirectional Long Short-Term Memory Neural Network for Optical Character Recognition
abstract
Optical Character Recognition is conversion of printed or handwritten text images into machine-encoded text. It is a building block of many processes such as machine translation, text-to-speech conversion and text mining. Bidirectional Long Short-Term Memory Neural Networks have shown a superior performance in character recognition with respect to other types of neural networks. In this paper, to the best of our knowledge, we propose the first hardware architecture of Bidirectional Long Short-Term Memory Neural Network with Connectionist Temporal Classification for Optical Character Recognition. Based on the new architecture, we present an FPGA hardware accelerator that achieves 459 times higher throughput than state-of-the-art. Visual recognition is a typical task on mobile platforms that usually use two scenarios either the task runs locally on embedded processor or offloaded to a cloud to be run on high performance machine. We show that computationally intensive visual recognition task benefits from being migrated to our dedicated hardware accelerator and outperforms high-performance CPU in terms of runtime, while consuming less energy than low power systems with negligible loss of recognition accuracy.
Vladimir Rybalkin, Norbert Wehn, Mohammad Reza Yousefi, Didier Stricker
DATE4
2017 Eyes of Things
abstract
Responsible Research and Innovation (RRI) is an approach that anticipates and assesses potential implications and societal expectations with regard to research and innovation, with the aim to foster the design of inclusive and sustainable research and innovation. While RRI includes many aspects, in certain types of projects ethics and particularly privacy, is arguably the most sensitive topic. The objective in Horizon 2020 innovation project Eyes of Things (EoT) is to build a small high-performance, low-power, computer vision platform (similar to a smart camera) that can work independently and also embedded into all types of artefacts. In this paper, we describe the actions taken within the project related to ethics and privacy. A privacy-by-design approach has been followed, and work continues now in four platform demonstrators.
Noelia Vállez, José Luis Espinosa-Aranda, Jose M. Rico-Saavedra, Javier Parra-Patino, Oscar Déniz-Suárez, Alain Pagani, Stephan Krauß, Ruben Reiser, Didier Stricker, David Moloney, Aubrey K. Dunne, Dexmont Peña, Martin Wäny, Matteo Sorci, Tim Llewellynn, Christian Fedorczak, Thierry Larmoire, Elodie Roche, Marco Herbst, Andre Seirafi, Kasra Seirafi
IC2E9
2017 Real-time monocular 6-DOF head pose estimation from salient 2D points
abstract
We propose a real-time and robust approach to estimate the full 3D head pose from extreme head poses using a monocular system. To this end, we first model the head using a simple geometric shape initialized using facial landmarks, i.e., eye corners, extracted from the face. Next, 2D salient points are detected within the region defined by the projection of the visible surface of the geometric head model onto the image, and projected back to the head model to generate the corresponding 3D features. Optical flow is used to find the respective 2D correspondences in the next video frame. Assuming that the monocular system is calibrated, it is then possible to solve the Perspective-n-Point (PnP) problem of estimating the head pose given a set of 3D features on the geometric model surface and their corresponding 2D correspondences from optical flow in the next frame. The experimental evaluation shows that the performance of the proposed approach achieves, and in some cases improves the state-of-the-art performance with a major advantage of not requiring facial landmarks (except for initialization). As a result, our method also applies to real scenarios in which facial landmarks-based methods fail due to self-occlusions.
Jilliam María Díaz Barros, Frederic Garcia, Bruno Mirbach, Didier Stricker
ICIP4
2017 Towards scheduling hard real-time image processing tasks on a single GPU
abstract
Graphics Processing Units (GPU) are becoming the key hardware accelerators in the emerging image processing applications such as self-driving cars and mobile augmented reality systems. As GPUs execute launched workloads non-preemptively, their usage in safety-critical systems with hard real-time constraints is impeded. The existing solutions for scheduling real-time tasks on a single GPU focus on soft real-time systems. In this paper, we consider real-time systems with a single dedicated GPU handling sporadic tasks with hard deadlines and propose a scheduling approach based on time division multiplexing called the GPU-TDMh - a lightweight middleware framework located between the application and the GPU driver layers. We evaluate the proposed approach on a matrix multiplication benchmark on a heterogeneous platform. The experiments demonstrate the effectiveness of our method as well as superiority over the non-preemptive online scheduling policies.
Vladislav Golyanik, Mitra Nasri, Didier Stricker
ICIP3
2017 Time-of-flight sensor depth enhancement for automotive exhaust gas
abstract
The Time-of-Flight (ToF) sensor has been envisioned as a candidate of next generation sensors for intelligent vehicles. One of the problems in automotive environment is that the sensor outputs wrong values if exhaust gas exists in the scene. In this paper, we provide two new contributions to the signal processing aspects of the ToF sensor for automotive use. First, we present the sensor characteristics and models for exhaust gas to cope with them. Second, we develop a depth enhancement algorithm to reject the influence of exhaust gas from multiple images. Experimental results demonstrate the effectiveness of the depth enhancement algorithm for both static data (including ground truth) and on-vehicle data acquired by the sensor mounted on a car.
Tomonari Yoshida, Oliver Wasenmüller, Didier Stricker
ICIP3
2017 Training a whole-book LSTM-based recognizer with an optimal training set
abstract
Despite the recent progress in OCR technologies, whole-book recognition, is still a challenging task, in particular in case of old and historical books, that the unknown font faces or low quality of paper and print contributes to the challenge. Therefore, pre-trained recognizers and generic methods do not usually perform up to required standards, and usually the performance degrades for larger scale recognition tasks, such as of a book. Such reportedly low error-rate methods turn out to require a great deal of manual correction. Generally, such methodologies do not make effective use of concepts such redundancy in whole-book recognition. In this work, we propose to train Long Short Term Memory (LSTM) networks on a minimal training set obtained from the book to be recognized. We show that clustering all the sub-words in the book, and using the sub-word cluster centers as the training set for the LSTM network, we can train models that outperform any identical network that is trained with randomly selected pages of the book. In our experiments, we also show that although the sub-word cluster centers are equivalent to about 8 pages of text for a 101- page book, a LSTM network trained on such a set performs competitively compared to an identical network that is trained on a set of 60 randomly selected pages of the book.
Mohammad Reza Soheili, Mohammad Reza Yousefi, Ehsanollah Kabir, Didier Stricker
ICMV4
2017 Addressing security challenges in industrial augmented reality systems
abstract
In context of Industry 4.0 Augmented Reality (AR) is frequently mentioned as the upcoming interface technology for human-machine communication and collaboration. Many prototypes have already arisen in both the consumer market and in the industrial sector. According to numerous experts it will take only few years until AR will reach the maturity level to be deployed in productive applications. Especially for industrial usage it is required to assess security risks and challenges this new technology implicates. Thereby we focus on plant operators, Original Equipment Manufacturers (OEMs) and component vendors as stakeholders. Starting from several industrial AR use cases and the structure of contemporary AR applications, in this paper we identify security assets worthy of protection and derive the corresponding security goals. Afterwards we elaborate the threats industrial AR applications are exposed to and develop an edge computing architecture for future AR applications which encompasses various measures to reduce security risks for our stakeholders.
Michael Langfinger, Didier Stricker, Hans D. Schotten
INDIN3
2017 Accurate 3D Reconstruction of Dynamic Scenes from Monocular Image Sequences with Severe Occlusions
abstract
The paper introduces an accurate solution to dense orthographic Non-Rigid Structure from Motion (NRSfM) in scenarios with severe occlusions or, likewise, inaccurate correspondences. We integrate a shape prior term into variational optimisation framework. It allows to penalize irregularities of the time-varying structure on the per-pixel level if correspondence quality indicator such as an occlusion tensor is available. We make a realistic assumption that several non-occluded views of the scene are sufficient to estimate an initial shape prior, though the entire observed scene may exhibit non-rigid deformations. Experiments on synthetic and real image data show that the proposed framework significantly outperforms state of the art methods for correspondence establishment in combination with the state of the art NRSfM methods. Together with the profound insights into optimisation methods, implementation details for heterogeneous platforms are provided.
Vladislav Golyanik, Torben Fetzer, Didier Stricker
WACV3
2017 Dense Batch Non-Rigid Structure from Motion in a Second
abstract
In this paper, we show how to minimise a quadratic function on a set of orthonormal matrices using an efficient semidefinite programming solver with application to dense non-rigid structure from motion. Thanks to the proposed technique, a new form of the convex relaxation for the Metric Projections (MP) algorithm is obtained. The modification results in an efficient single-core CPU implementation enabling dense factorisations of long image sequences with tens of thousands of points into camera pose and non-rigid shape in seconds, i.e., at least two orders of magnitude faster than the runtimes reported in the literature so far. The proposed implementation can be useful for interactive or real-time robotic and other applications, where monocular non-rigid reconstruction is required. In a narrow sense, our paper complements research on MP, though the proposed convex relaxation methodology can also be useful in other computer vision tasks. The experimental part providing runtime evaluation and qualitative analysis concludes the paper.
Vladislav Golyanik, Didier Stricker
WACV2
2017 Introduction to the Special Section on Augmented Video
abstract
Merging computer-generated content with real-world visual data is one of the main challenges in fields like augmented reality or visual effects and is increasingly important in broadcasting, gaming, medical, automotive, maintenance, and learning applications. Although augmented reality (AR) has been investigated for a long time, it has recently emerged as a hot topic, with significant commercial interest. One reason for that is the availability of new camera-equipped devices like smart phones or tablets that enable see-through capabilities. Combined with powerful graphics capabilities, sensors, and tracking methods, AR now becomes available to everyone. In addition, glasses-based systems, like Microsoft’s HoloLens, allow for hands-free visualization, enabling many new applications.
Peter Eisert, Yebin Liu, Kyuong Mu Lee, Didier Stricker, Graham A. Thomas
IEEE Trans. Circuits Syst. Video Technol.4
2016 NRSfM-Flow: Recovering Non-Rigid Scene Flow from Monocular Image Sequences
Vladislav Golyanik, Aman S. Mathur, Didier Stricker
BMVC3
2016 Gravitational Approach for Point Set Registration
abstract
In this paper a new astrodynamics inspired rigid point set registration algorithm is introduced-the Gravitational Approach (GA). We formulate point set registration as a modified N-body problem with additional constraints and obtain an algorithm with unique properties which is fully scalable with the number of processing cores. In GA, a template point set moves in a viscous medium under gravitational forces induced by a reference point set. Pose updates are completed by numerically solving the differential equations of Newtonian mechanics. We discuss techniques for efficient implementation of the new algorithm and evaluate it on several synthetic and real-world scenarios. GA is compared with the widely used Iterative Closest Point and the state of the art rigid Coherent Point Drift algorithms. Experiments evidence that the new approach is robust against noise and can handle challenging scenarios with structured outliers.
Vladislav Golyanik, Sk Aziz Ali, Didier Stricker
CVPR3
2016 Joint pre-alignment and robust rigid point set registration
abstract
We present an elegant solution to joint pre-alignment and rigid point set registration, given prior matches. Instead of performing pre-alignment and the actual registration in the separate steps, prior matches explicitly influence the registration procedure in our approach. This results in several advantages. Firstly, our approach solves the pre-alignment task - an approximate resolving of rotation and translation - with an insufficient number of prior correspondences, when other methods fail. Secondly, it produces more accurate rigid registrations of noisy point sets than the state of the art Coherent Point Drift method. Combined with application specific methods for correspondence establishment, we demonstrate superiority of our approach in several synthetic and real-world scenarios.
Vladislav Golyanik, Bertram Taetz, Didier Stricker
ICIP3
2016 Adding Model Constraints to CNN for Top View Hand Pose Recognition in Range Images
abstract
A new dataset for hand-pose is introduced. The dataset includes the top view images of the palm by Time of Flight (ToF) camera. It is recorded in an experimental setting with twelve participants for six hand-poses. An evaluation on the dataset is carried out with a dedicated Convolutional Neural Network (CNN) architecture for Hand Pose Recognition (HPR). This architecture uses a model-layer. The small size model layer creates a funnel shape network which adds a priori knowledge and constrains the network by modelling the degree of freedom of the palm, such that it learns palm features. It is demonstrated that this network performs better than a similar network without the prior added. A two-phase learning scheme which allows training the model on full dataset even when the classification problem is confined to a subset of the classes is described. The best model performs at an accuracy of 92%. Finally, we show the feature transfer capability of the network and compare the extracted features from various networks and discuss usefulness for various applications.
Aditya Tewari, Frédéric Grandidier, Bertram Taetz, Didier Stricker
ICPRAM4
2016 Learning to Fuse: A Deep Learning Approach to Visual-Inertial Camera Pose Estimation
abstract
Camera pose estimation is the cornerstone of Augmented Reality applications. Pose tracking based on camera images exclusively has been shown to be sensitive to motion blur, occlusions, and illumination changes. Thus, a lot of work has been conducted over the last years on visual-inertial pose tracking using acceleration and angular velocity measurements from inertial sensors in order to improve the visual tracking. Most proposed systems use statistical filtering techniques to approach the sensor fusion problem, that require complex system modelling and calibrations in order to perform adequately. In this work we present a novel approach to sensor fusion using a deep learning method to learn the relation between camera poses and inertial sensor measurements. A long short-term memory model (LSTM) is trained to provide an estimate of the current pose based on previous poses and inertial measurements. This estimates then appropriately combined with the output of a visual tracking system using a linear Kalman Filter to provide a robust final pose estimate. Our experimental results confirm the applicability and tracking performance improvement gained from the proposed sensor fusion system.
Jason R. Rambach, Aditya Tewari, Alain Pagani, Didier Stricker
ISMAR4
2016 Augmented Reality 3D Discrepancy Check in Industrial Applications
abstract
Discrepancy check is a well-known task in industrial Augmented Reality (AR). In this paper we present a new approach consisting of three main contributions: First, we propose a new two-step depth mapping algorithm for RGB-D cameras, which fuses depth images with given camera pose in real-time into a consistent 3D model. In a rigorous evaluation with two public benchmarks we show that our mapping outperforms the state-of-the-art in accuracy. Second, we propose a semi-automatic alignment algorithm, which rapidly aligns a reference model to the reconstruction. Third, we propose an algorithm for 3D discrepancy check based on pre-computed distances. In a systematic evaluation we show the superior performance of our approach compared to state-of-the-art 3D discrepancy checks.
Oliver Wasenmüller, Marcel Meyer, Didier Stricker
ISMAR3
2016 Extended coherent point drift algorithm with correspondence priors and optimal subsampling
abstract
The problem of dense point set registration, given a sparse set of prior correspondences, often arises in computer vision tasks. Unlike in the rigid case, integrating prior knowledge into a registration algorithm is especially demanding in the non-rigid case due to the high variability of motion and deformation. In this paper we present the Extended Coherent Point Drift registration algorithm. It enables, on the one hand, to couple correspondence priors into the dense registration procedure in a closed form and, on the other hand, to process large point sets in reasonable time through adopting an optimal coarse-to-fine strategy. Combined with a suitable keypoint extractor during the preprocessing step, our method allows for non-rigid registrations with increased accuracy for point sets with structured outliers. We demonstrate advantages of our approach against other non-rigid point set registration methods in synthetic and real-world scenarios.
Vladislav Golyanik, Bertram Taetz, Gerd Reis, Didier Stricker
WACV4
2016 Occlusion-aware video registration for highly non-rigid objects
abstract
This paper addresses the problem of video registration for dense non-rigid structure from motion under suboptimal conditions, such as noise, self-occlusions, considerable external occlusions or specularities, i.e. the computation of optical flow between the reference image and each of the subsequent images in a video sequence when the camera observes a highly deformable object. We tackle this challenging task by improving previously proposed variational optimization techniques for multi-frame optical flow (MFOF) through detection, tracking and handling of uncertain flow field estimates. This is based on a novel Bayesian inference approach incorporated into the MFOF. At the same time, computational costs are significantly reduced through iterative pre-computation of the flow fields. As shown through experiments, the resulting method performs superior to other state-of-the-art (MF)OF methods on video sequences showing a highly non-rigidly deforming object with considerable occlusions.
Bertram Taetz, Gabriele Bleser-Taetz, Vladislav Golyanik, Didier Stricker
WACV4
2016 CoRBS: Comprehensive RGB-D benchmark for SLAM using Kinect v2
abstract
In scientific evaluation public datasets and benchmarks are indispensable to perform objective assessment. In this paper we present a new Comprehensive RGB-D Benchmark for SLAM (CoRBS). In contrast to state-of-the-art RGB-D SLAM benchmarks, we provide the combination of real depth and color data together with a ground truth trajectory of the camera and a ground truth 3D model of the scene. Our novel benchmark allows for the first time to independently evaluate the localization as well as the mapping part of RGB-D SLAM systems with real data. We obtained the ground truth for the trajectory using an external motion capture system and for the scene geometry via an external 3D scanner, each with sub-millimeter precision. With precise calibration and systematic validation we ensured the high quality of CoRBS. Our dataset contains twenty image sequences of four different scenes captured with a Kinect v2. We provide all data in a global coordinate system to enable direct evaluation without any further alignment or calibration.
Oliver Wasenmüller, Marcel Meyer, Didier Stricker
WACV3
2015 Real-Time Head Pose Estimation Using Multi-variate RVM on Faces in the Wild
Mohamed Selim, Alain Pagani, Didier Stricker
CAIP (2)3
2015 Flow Fields: Dense Correspondence Fields for Highly Accurate Large Displacement Optical Flow Estimation
abstract
Modern large displacement optical flow algorithms usually use an initialization by either sparse descriptor matching techniques or dense approximate nearest neighbor fields. While the latter have the advantage of being dense, they have the major disadvantage of being very outlier-prone as they are not designed to find the optical flow, but the visually most similar correspondence. In this article we present a dense correspondence field approach that is much less outlier-prone and thus much better suited for optical flow estimation than approximate nearest neighbor fields. Our approach does not require explicit regularization, smoothing (like median filtering) or a new data term. Instead we solely rely on patch matching techniques and a novel multi-scale matching strategy. We also present enhancements for outlier filtering. We show that our approach is better suited for large displacement optical flow estimation than modern descriptor matching techniques. We do so by initializing EpicFlow with our approach instead of their originally used state-of-the-art descriptor matching technique. We significantly outperform the original EpicFlow on MPI-Sintel, KITTI 2012, KITTI 2015 and Middlebury. In this extended article of our former conference publication we further improve our approach in matching accuracy as well as runtime and present more experiments and insights.
Christian Bailer, Bertram Taetz, Didier Stricker
ICCV3
2015 Binarization-free OCR for historical documents using LSTM networks
abstract
A primary preprocessing block of almost any typical OCR system is binarization, through which it is intended to remove unwanted part of the input image, and only keep a binarized and cleaned-up version for further processing. The binarization step does not, however, always perform perfectly, and it can happen that binarization artifacts result in important information loss, by for instance breaking or deforming character shapes. In historical documents, due to a more dominant presence of noise and other sources of degradations, the performance of binarization methods usually deteriorates; as a result the performance of the recognition pipeline is hindered by such preprocessing phases. In this paper, we propose to skip the binarization step by directly training a 1D Long Short Term Memory (LSTM) network on gray-level text lines. We collect a large set of historical Fraktur documents, from publicly available online sources, and form train and test sets for performing experiments on both gray-level and binarized text lines. In order to observe the impact of resolution, the experiments are carried out on two identical sets of low and high resolutions. Overall, using gray-level text lines, the 1D LSTM network can reach 24% and 19% lower error rates on the low- and high-resolution sets, respectively, compared to the case of using binarization in the recognition pipeline.
Mohammad Reza Yousefi, Mohammad Reza Soheili, Thomas M. Breuel, Ehsanollah Kabir, Didier Stricker
ICDAR5
2015 Cognitive Augmented Reality
Nils Petersen, Didier Stricker
Comput. Graph.2
2015 A novel confidence-based multiclass boosting algorithm for mobile physical activity monitoring
Attila Reiss, Gustaf Hendeby, Didier Stricker
Pers. Ubiquitous Comput.3
2014 Spherical Light Fields
Bernd Krolla, Maximilian Diebold, Bastian Goldlücke, Didier Stricker
BMVC4
2014 A Superior Tracking Approach: Building a Strong Tracker through Fusion
Christian Bailer, Alain Pagani, Didier Stricker
ECCV (7)3
2014 Robust and Accurate Non-parametric Estimation of Reflectance Using Basis Decomposition and Correction Functions
Tobias Nöll, Johannes Köhler 0002, Didier Stricker
ECCV (2)3
2014 Accurate and robust spherical camera pose estimation using consistent points
abstract
This paper addresses the problem of multi-view camera pose estimation of high resolution, full spherical images. A novel approach to simultaneously retrieve camera poses along with a sparse point cloud is designed for large scale scenes. We introduce the concept of consistent points that allows to dynamically select the most reliable 3D points for nonlinear pose refinement. In contrast to classical bundle adjustment approaches, we propose to reduce the parameter search space while jointly optimizing camera poses and scene geometry. Our method notably improves accuracy and robustness of camera pose estimation, as shown by experiments carried out on real image data.
Christiano Couto Gava, Bernd Krolla, Didier Stricker
ICMV3
2014 Sub-word image clustering in Farsi printed books
abstract
Most OCR systems are designed for the recognition of a single page. In case of unfamiliar font faces, low quality papers and degraded prints, the performance of these products drops sharply. However, an OCR system can use redundancy of word occurrences in large documents to improve recognition results. In this paper, we propose a sub-word image clustering method for the applications dealing with large printed documents. We assume that the whole document is printed by a unique unknown font with low quality print. Our proposed method finds clusters of equivalent sub-word images with an incremental algorithm. Due to the low print quality, we propose an image matching algorithm for measuring the distance between two sub-word images, based on Hamming distance and the ratio of the area to the perimeter of the connected components. We built a ground-truth dataset of more than 111000 sub-word images to evaluate our method. All of these images were extracted from an old Farsi book. We cluster all of these sub-words, including isolated letters and even punctuation marks. Then all centers of created clusters are labeled manually. We show that all sub-words of the book can be recognized with more than 99.7% accuracy by assigning the label of each cluster center to all of its members.
Mohammad Reza Soheili, Ehsanollah Kabir, Didier Stricker
ICMV3
2014 LSTM-Based Early Recognition of Motion Patterns
abstract
In this paper a method for Early Recognition (ER) of Motion Templates (MTs) is presented. We define ER as an algorithm to provide recognition results before a motion sequence is completed. In our experiments we apply Long Short-Term Memory (LSTM) and optimize the training for the task of recognizing the motion template as early as possible. The evaluation has shown that the recognition accuracy for a frame-by-frame classification the LSTM achieves a recognition accuracy of 88% if no training data of the person him/herself is included, and 92% if the training data also contains motion sequences of the person. Furthermore, the average earliness - the number of time frames it takes before the LSTM correctly classifies a motion pattern - is around 24.77 frames, which is less than a second with the used tracking technology, i.e., the Microsoft Kinect.
Marcus Liwicki, Didier Stricker, Christopher Schölzel, Seiichi Uchida
ICPR3
2013 A competitive approach for human activity recognition on smartphones
Attila Reiss, Gustaf Hendeby, Didier Stricker
ESANN3
2013 Real-time modeling and tracking manual workflows from first-person vision
abstract
Recognizing previously observed actions in video sequences can lead to Augmented Reality manuals that (1) automatically follow the progress of the user and (2) can be created from video examples of the workflow. Modeling is challenging, as the environment is susceptible to change drastically due to user interaction and camera motion may not provide sufficient translation to robustly estimate geometry. We propose a piecewise homographic transform that projects the given video material onto a series of distinct planar subsets of the scene. These subsets are selected by segmenting the largest image region that is consistent with a homographic model and contains a given region of interest. We are then able to model the state of the environment and user actions using simple 2D region descriptors. The model elegantly handles estimation errors due to incomplete observation and is robust towards occlusions, e.g., due to the user's hands. We demonstrate the effectiveness of our approach quantitatively and compare it to the current state of the art. Further, we show how we apply the approach to visualize automatically assessed correctness criteria during run-time.
Nils Petersen, Alain Pagani, Didier Stricker
ISMAR3
2013 A full-spherical device for simultaneous geometry and reflectance acquisition
abstract
We present OrcaM, a device for exploring new methods in the field of simultaneous acquisition of geometry, color and reflectance properties. OrcaM employs a full-spherical construction, a movable projector-camera unit, 633 individually controllable LEDs and a height-adjustable turntable with a glass carrier. In contrast to state of the art hardware layouts, this design allows data acquisition from all possible directions in a single pass without any user interaction. In this paper we report the challenges we encountered during development. We describe the used calibration algorithms that constitute the basis for all future reconstruction methods and present results computed with the methods we currently use.
Johannes Köhler 0002, Tobias Nöll, Gerd Reis, Didier Stricker
WACV4
2013 Algorithms for 3D Shape Scanning with a Depth Camera
abstract
We describe a method for 3D object scanning by aligning depth scans that were taken from around an object with a Time-of-Flight (ToF) camera. These ToF cameras can measure depth scans at video rate. Due to comparably simple technology, they bear potential for economical production in big volumes. Our easy-to-use, cost-effective scanning solution, which is based on such a sensor, could make 3D scanning technology more accessible to everyday users. The algorithmic challenge we face is that the sensor's level of random noise is substantial and there is a nontrivial systematic bias. In this paper, we show the surprising result that 3D scans of reasonable quality can also be obtained with a sensor of such low data quality. Established filtering and scan alignment techniques from the literature fail to achieve this goal. In contrast, our algorithm is based on a new combination of a 3D superresolution method with a probabilistic scan alignment approach that explicitly takes into account the sensor's noise characteristics.
Yan Cui 0011, Sebastian Schuon, Sebastian Thrun, Didier Stricker, Christian Theobalt
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Unsupervised motion pattern learning for motion segmentation
Gabriele Bleser-Taetz, Marcus Liwicki, Didier Stricker
ICPR4
2012 Learning task structure from video examples for workflow tracking and authoring
abstract
We present a robust real-time capable and simple framework for segmenting video sequences and live-streams of manual workflows into the comprising single tasks. Using classifiers trained on these segments we can follow a user that is performing the workflow in real-time as well as learn task variants from additional video examples. Our proposed method neither requires object detection nor high-level features. Instead we propose a novel measure derived from image distance that evaluates image properties jointly without prior segmentation. Our method can cope with repetitive and free-hand activities and the results are in many cases comparable or equal to manual task segmentation. One important application of our method is the automatic creation of a step-by-step task documentation from a video demonstration. The entire process to automatically create a fully functional augmented reality manual will be explained in detail and results are shown.
Nils Petersen, Didier Stricker
ISMAR2
2012 An Integrated Mobile System for Long-Term Aerobic Activity Monitoring and Support in Daily Life
abstract
This paper presents a mobile and unobtrusive platform that enables the accurate monitoring of physical activities in daily life, and is integrated into a healthcare system supporting out-of-hospital services. The main focus of the paper is to describe and evaluate a complete data processing chain for recognizing different activities, and estimating their intensity level. To develop a mobile platform that is more applicable in real-life scenarios than systems in previous work, an additional other activity class is introduced to deal with various everyday, household and fitness activities. Evaluation of the proposed methods is done on a dataset recorded by 9 subjects performing 16 different activities, while wearing the mobile platform consisting of three IMUs, a HR-monitor and a mobile companion unit. Another important aspect of this work is the integration of the mobile platform with an EHR, providing access for both the clinician (e.g. to enter a patient's medical record or to set up a care plan) and the patient (e.g. to watch assigned educational material). Feedback about the patient's daily progress is also given both to the patient and the clinician, preserving or even increasing the patient's motivation to follow the defined care plan, and providing valuable information on program adherence for the clinician.
Attila Reiss, Didier Stricker, Ilias Lamprinos
TrustCom2
2011 Exploring and extending the boundaries of physical activity recognition
abstract
This paper discusses several aspects and practical issues of physical activity recognition. Many existent activity recognition applications only include the few and well known basic activities, thus limiting the applicability of these systems. One of the main goals of this paper is to point out the importance of extending activity recognition with background activities, and to demonstrate its effects. Another practical issue of activity recognition is, that the comparison of different approaches is often not possible not only because of the lack of common datasets, but also because of the usage of different evaluation methods. This work argues, that usually subject independent validation techniques should be applied for the evaluation of activity recognition systems. The statements of this work are demonstrated on a dataset with 8 subjects and 13 different activities. Finally, an empirical study carried out within this work shows the feasibility of using more complex classifiers (required due to the extended activity recognition problem) for mobile applications.
Attila Reiss, Didier Stricker
SMC3
2011 Unsupervised model generation for motion monitoring
abstract
This paper addresses two fundamental requirements of full body motion monitoring: (a) the ability to sense the input of the user and (b) the means to interpret the captured input. Appropriate technology in both areas is required for an interactive virtual reality system to provide feedback in a useful and natural way. This paper combines technologies for both areas: It develops a sensor fusion approach for capturing user input based on miniature on-body inertial and magnetic motion sensors. Furthermore, it presents work in progress to automatically generate models for motion patterns from the captured input. The technology is then used and evaluated in the context of a personalized virtual rehabilitation trainer application.
Gabriele Bleser-Taetz, Gustaf Hendeby, Attila Reiss, Didier Stricker
SMC5
2011 Efficient Packing of Arbitrary Shaped Charts for Automatic Texture Atlas Generation
abstract
Abstract Texture atlases are commonly used as representations for mesh parameterizations in numerous applications including texture and normal mapping. Therefore, packing is an important post‐processing step that tries to place and orient the single parameterizations in a way that the available space is used as efficiently as possible. However, since packing is NP hard, only heuristics can be used in practice to find near‐optimal solutions. In this publication we introduce the new search space of modulo valid packings. The key idea thereby is to allow the texture charts to wrap around in the atlas. By utilizing this search space we propose a new algorithm that can be used in order to automatically pack texture atlases. In the evaluation section we show that our algorithm achieves solutions with a significantly higher packing efficiency when compared to the state of the art, especially for complex packing problems.
Tobias Nöll, Didier Stricker
Comput. Graph. Forum2
2010 SIFT in perception-based color space
abstract
Scale Invariant Feature Transform (SIFT) has been proven to be the most robust local invariant feature descriptor. However, SIFT is designed mainly for grayscale images. Many local features can be misclassified if their color information is ignored. Motivated by perceptual principles, this paper addresses a new color space, called perception-based color space, in which the associated metric approximates perceived distances and color displacements and captures illumination invariant relationship. Instead of using grayscale values to represent the input image, the proposed approach builds the SIFT descriptors in the new color space, resulting in a descriptor that is more robust than the standard SIFT with respect to color and illumination variations. The evaluation results support the potential of the proposed approach.
Yan Cui 0011, Alain Pagani, Didier Stricker
ICIP3
2010 3D discrepancy check via Augmented Reality
abstract
For many tasks like markerless model-based camera tracking it is essential that the 3D model of a scene accurately represents the real geometry of the scene. It is therefore very important to detect deviations between a 3D model and a scene. We present an innovative approach which is based on the insight that camera tracking can not only be used for Augmented Reality visualization but also to solve the correspondence problem between 3D measurements of a real scene and their corresponding positions in the 3D model. We combine a time-of-flight camera (which acquires depth images in real time) with a custom 2D camera (used for the camera tracking) and developed an analysis-by-synthesis approach to detect deviations between a scene and a 3D model of the scene.
Svenja Kahn, Harald Wuest, Didier Stricker, Dieter W. Fellner
ISMAR3
2009 Learning Local Patch Orientation with a Cascade of Sparse Regressors
abstract
We present a new method for infering the local 3D orientation of keypoints from their appearance. The method is based on the idea that the relation between keypoint appearance and pose can be learnt efficiently with an adequate regressor. Using one reference view of a keypoint, it is possible to train a keypoint-specific regressor that takes the point appearance as input and delivers the local perspective transformation as output. We show that an elegant choice of regressor is a set of sparse regressors applied sequentially in a cascade. In our case, we use a set of parametrized multivariate relevance vector machines (MVRVM) to learn the local 8-dimensional homography from the patch normalized pixel values. We show that using a cascade of regressors, ranging from coarse pose approximation to fine rectifications, considerably speeds up the identification and pose estimation process. Moreover, we show that our method improves the precision of classical points detectors, as the location of the point is rectified together with the homography. The resulting system is able to recover the orientation of patches in real time.
Alain Pagani, Didier Stricker
BMVC2
2009 Integral P-channels for fast and robust region matching
abstract
We present a new method for matching a region between an input and a query image, based on the P-channel representation of pixel-based image features such as grayscale and color information, local gradient orientation and local spatial coordinates. We introduce the concept of integral P-channels, which conciliates the concepts of P-channel and integral images. Using integral images, the P-channel representation of a given region is extracted with a few arithmetic operations. This enables a fast nearest-neighbor search in all possible target regions. We present extensive experimental results and show that our approach compares favorably to existing methods for region matching such as histograms or region covariance.
Alain Pagani, Didier Stricker, Michael Felsberg
ICIP2
2009 Continuous natural user interface: Reducing the gap between real and digital world
abstract
Augmented reality (AR) presentation enables the creation of natural user interfaces that employ the whole user's environment as interaction device. Additionally, by using hand based 3D interaction with gestures that have a physical meaning like grabbing, dragging, and dropping this leads to a user experience that is intuitive, since close to the real world's behavior. We propose a novel approach to an AR-based natural user interface, that goes one step further by enabling the contents of the interface to switch domains from a virtual instance in AR to a physical instance in the real-world. All instances stay associated and changes made to the physical instance will be reflected on the virtual one. Because the behavior of our interface in AR is in key aspects consistent with the real-world, the gap between those domains is made less salient. To demonstrate our concept, we have implemented an exemplary industrial use case. Our main contribution is the methodology for an intuitive interface we call continuous natural user interface (CNUI). Additionaly, we conducted a user study to investigate the acceptance of this kind of interface. Results indicate an ergonomic ease and after a training period also an increased performance when using our system.
Nils Petersen, Didier Stricker
ISMAR2
2009 Advanced tracking through efficient image processing and visual-inertial sensor fusion
Gabriele Bleser-Taetz, Didier Stricker
Comput. Graph.2
2008 Using the marginalised particle filter for real-time visual-inertial sensor fusion
abstract
The use of a particle filter (PF) for camera pose estimation is an ongoing topic in the robotics and computer vision community, especially since the FastSLAM algorithm has been utilised for simultaneous localisation and mapping (SLAM) applications with a single camera. The major problem in this context consists in the poor proposal distribution of the camera pose particles obtained from the weak motion model of a camera moved freely in 3D space. While the FastSLAM 2.0 extension is one possibility to improve the proposal distribution, this paper addresses the question of how to use measurements from low-cost inertial sensors (gyroscopes and accelerometers) to compensate for the missing control information. However, the integration of inertial data requires the additional estimation of sensor biases, velocities and potentially accelerations, resulting in a state dimension, which is not manageable by a standard PF. Therefore, the contribution of this paper consists in developing a real-time capable sensor fusion strategy based upon the marginalised particle filter (MPF) framework. The performance of the proposed strategy is evaluated in combination with a marker-based tracking system and results from a comparison with previous visual-inertial fusion strategies based upon the extended Kalman filter (EKF), the standard PF and the MPF are presented.
Gabriele Bleser-Taetz, Didier Stricker
ISMAR2
2008 Advanced tracking through efficient image processing and visual-inertial sensor fusion
abstract
We present a new visual-inertial tracking device for augmented and virtual reality applications. The paper addresses two fundamental issues of such systems. The first one concerns the definition and modelling of the sensor fusion. Much work has been done in this area and several models for exploiting the data of the gyroscopes and linear accelerometers have been proposed. However, the respective advantages of each model and in particular the benefits of the integration of the accelerometer data in the filter are still unclear. The paper therefore provides an evaluation of different models with special investigation of the effects of using accelerometers on the tracking performance. The second contribution is about the development of an image processing approach that does not require special landmarks but uses natural features. Our solution relies on a 3D model of the scene that enables to predict the appearances of the features by rendering the model using the prediction data of the sensor fusion filter. The feature localisation is robust and accurate mainly because local lighting is also estimated. The final system is evaluated with help of ground-truth and real data. High stability and accuracy is demonstrated also for large environments.
Gabriele Bleser-Taetz, Didier Stricker
VR2
2007 Feature Management for Efficient Camera Tracking
Harald Wuest, Alain Pagani, Didier Stricker
ACCV (1)3
2007 Adaptable Model-Based Tracking Using Analysis-by-Synthesis Techniques
Harald Wuest, Folker Wientapper, Didier Stricker
CAIP3
2007 Identifying differences between CAD and physical mock-ups using AR
abstract
Since the last ten years product development in automotive industry is changing radically. Most physical mock-ups have vanished and are now replaced by digital ones. But they are still needed for final evaluations or issues, which cannot be adequately simulated. During their production, deviations from the CAD model may be made. Since digital and real mock-up must match for the further product development, the transfer of differences between physical and digital mock-up to the CAD format is a crucial issue. In this paper an augmented reality (AR) based tool-chain is presented, which allows matching the CAD data with real mock-ups and documents the differences between them. Essential functions like measurement and online construction are provided, allowing the end-users to create information in AR space and feeding them back into the CAD model.
Sabine Webel, Mario Becker, Didier Stricker, Harald Wuest
ISMAR3
2006 Online camera pose estimation in partially known and dynamic scenes
abstract
One of the key requirements of augmented reality systems is a robust real-time camera pose estimation. In this paper we present a robust approach, which does neither depend on offline pre-processing steps nor on pre-knowledge of the entire target scene. The connection between the real and the virtual world is made by a given CAD model of one object in the scene. However, the model is only needed for initialization. A line model is created out of the object rendered from a given camera pose and registrated onto the image gradient for finding the initial pose. In the tracking phase, the camera is not restricted to the modeled part of the scene anymore. The scene structure is recovered automatically during tracking. Point features are detected in the images and tracked from frame to frame using a brightness invariant template matching algorithm. Several template patches are extracted from different levels of an image pyramid and are used to make the 2D feature tracking capable for large changes in scale. Occlusion is detected already on the 2D feature tracking level. The features' 3D locations are roughly initialized by linear triangulation and then refined recursively over time using techniques of the Extended Kalman Filter framework. A quality manager handles the influence of a feature on the estimation of the camera pose. As structure and pose recovery are always performed under uncertainty, statistical methods for estimating and propagating uncertainty have been incorporated consequently into both processes. Finally, validation results on synthetic as well as on real video sequences are presented.
Gabriele Bleser-Taetz, Harald Wuest, Didier Stricker
ISMAR3
2005 Adaptive Line Tracking with Multiple Hypotheses for Augmented Reality
abstract
We present a real-time model-based line tracking approach with adaptive learning of image edge features that can handle partial occlusion and illumination changes. A CAD (VRML) model of the object to track is needed. First, the visible edges of the model with respect to the camera pose estimate are sorted out by a visibility test performed on standard graphics hardware. For every sample point of every projected visible 3D model line a search for gradient maxima in the image is then carried out in a direction perpendicular to that line. Multiple hypotheses of these maxima are considered as putative matches. The camera pose is updated by minimizing the distances between the projection of all sample points of the visible 3D model lines and the most likely matches found in the image. The state of every edge's visual properties is updated after each successful camera pose estimation. We evaluated the algorithm and showed the improvements compared to other tracking approaches.
Harald Wuest, Florent Vial, Didier Stricker
ISMAR3
2003 Lessons learned on the way to industrial augmented reality applications, a retrospective on ARVIKA
Jens Weidenhausen, Christian Knöpfle, Didier Stricker
Comput. Graph.3
1999 An Optically Based Direct Manipulation Interface for Human-Computer Interaction in an Augmented World
Gudrun Klinker, Didier Stricker, Dirk Reiners
EGVE2
1999 Optically based direct manipulation for augmented reality
Gudrun Klinker, Didier Stricker, Dirk Reiners
Comput. Graph.2