Rainer Stiefelhagen

dblp:31/4699 · DBLP profile ↗
← Back
255ranked-venue papers
11as first author
94since 2021 · last 2026
0000-0001-8046-4945ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 141 · 6 first-author · 45 since 2021Artificial intelligence and machine learning · 137 · 6 first-author · 60 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 1 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 29 · 1 first-author · 4 since 2021Systems, architecture and hardware · 15 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 4 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 HybriDLA: Hybrid Generation for Document Layout Analysis
abstract
Conventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for early-generation documents with a small, predetermined number of regions, this paradigm struggles with contemporary documents, which exhibit diverse element counts and increasingly complex layouts. To address challenges posed by modern documents, we present HybriDLA, a novel generative framework that unifies diffusion and autoregressive decoding within a single layer. The diffusion component iteratively refines bounding-box hypotheses, whereas the autoregressive component injects semantic and contextual awareness, enabling precise region prediction even in highly varied layouts. To further enhance detection quality, we design a multi-scale feature-fusion encoder that captures both fine-grained and high-level visual cues. This architecture elevates performance to 83.5% mean Average Precision (mAP). Extensive experiments on the DocLayNet and M6Doc benchmarks demonstrate that HybriDLA sets a state-of-the-art performance, outperforming previous approaches.
Yufan Chen 0001, Omar Moured, Ruiping Liu 0001, Junwei Zheng, Kunyu Peng, Jiaming Zhang 0001, Rainer Stiefelhagen
AAAI7
2026 Mitigating Label Noise using Prompt-Based Hyperbolic Meta-Learning in Open-Set Domain Generalization
Kunyu Peng, Di Wen 0006, M. Saquib Sarfraz, Yufan Chen 0001, Junwei Zheng, David Schneider 0006, Kailun Yang 0001, Alina Roitberg, Rainer Stiefelhagen
Int. J. Comput. Vis.10
2025 Every Component Counts: Rethinking the Measure of Success for Medical Semantic Segmentation in Multi-Instance Segmentation Tasks
abstract
We present Connected-Component (CC)-Metrics, a novel semantic segmentation evaluation protocol, targeted to align existing semantic segmentation metrics to a multi-instance detection scenario in which each connected component matters. We motivate this setup in the common medical scenario of semantic metastases segmentation in a full-body PET/CT. We show how existing semantic segmentation metrics suffer from a bias towards larger connected components contradicting the clinical assessment of scans in which tumor size and clinical relevance are uncorrelated. To rebalance existing segmentation metrics, we propose to evaluate them on a per-component basis thus giving each tumor the same weight irrespective of its size. To match predictions to ground-truth segments, we employ a proximity-based matching criterion, evaluating common metrics locally at the component of interest. Using this approach, we break free of biases introduced by large metastasis for overlap-based metrics such as Dice or Surface Dice. CC-Metrics also improves distance-based metrics such as Hausdorff Distances which are uninformative for small changes that do not influence the maximum or 95th percentile, and avoids pitfalls introduced by directly combining counting-based metrics with overlap-based metrics as it is done in Panoptic Quality.
Alexander Jaus, Constantin Seibold, Simon Reiß, Zdravko Marinov, Zeling Ye, Stefan Krieg 0002, Jens Kleesiek, Rainer Stiefelhagen
AAAI9
2025 A Low-Fidelity Prototyping Method for Blind Users: A Case Study on Designing Two-Dimensional Tactile Displays
abstract
Design workshops are often inaccessible to blind people due to a strong reliance on visual design methods. This paper presents an adapted low-fidelity prototyping method to enable blind persons to actively participate in the early design stages of developing assistive devices. We chose the development of a 2D tactile display as a case study. Therefore, we conducted four workshops with blind participants to develop a haptic toolkit and adapt the structure of prototyping workshops based on the preferences of the target group. The resulting toolkit and workshop structure were applied in six prototyping workshops with 12 blind participants to design a new 2D tactile display to test the new method. Our findings show that the participation of blind individuals is highly relevant for structuring accessible prototyping workshops, deciding group size, and developing accessible prototyping materials. Furthermore, our work lays the foundation for further investigation of the role of sighted assistants and the design of workshops that truly center blind persons as primary contributors.
Sara Alzalabny, Karin Müller 0001, Kathrin Maria Gerling, Thorsten Schwarz, Bastian E. Rapp, Rainer Stiefelhagen
ASSETS6
2025 Scene-agnostic Pose Regression for Visual Localization
abstract
Absolute Pose Regression (APR) predicts 6D camera poses but lacks the adaptability to unknown environments without retraining, while Relative Pose Regression (RPR) generalizes better yet requires a large image retrieval database. Visual Odometry (VO) generalizes well in unseen environments but suffers from accumulated error in open trajectories. To address this dilemma, we introduce a new task, Scene-agnostic Pose Regression (SPR), which can achieve accurate pose regression in a flexible way while eliminating the need for retraining or databases. To benchmark SPR, we created a large-scale dataset, 360SPR, with over 200K photorealistic panoramas, 3.6M pinhole images and camera poses in 270 scenes at three different sensor heights. Furthermore, a SPR-Mamba model is initially proposed to address SPR in a dual-branch manner. Extensive experiments and studies demonstrate the effectiveness of our SPR paradigm, dataset, and model. In the unknown scenes of both 360SPR and 360Loc datasets, our method consistently outperforms APR, RPR and VO. The dataset and code are available at SPR.
Junwei Zheng, Ruiping Liu 0001, Yufan Chen 0001, Zhenfang Chen, Kailun Yang 0001, Jiaming Zhang 0001, Rainer Stiefelhagen
CVPR7
2025 Is Visual in-Context Learning for Compositional Medical Tasks Within Reach?
abstract
In this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks during test time without re-training. Unlike previous approaches, our focus is on training in-context learners to adapt to sequences of tasks, rather than individual tasks. Our goal is to solve complex tasks that involve multiple intermediate steps using a single model, allowing users to define entire vision pipelines flexibly at test time. To achieve this, we first examine the properties and limitations of visual in-context learning architectures, with a particular focus on the role of codebooks. We then introduce a novel method for training in-context learners using a synthetic compositional task generation engine. This engine bootstraps task sequences from arbitrary segmentation datasets, enabling the training of visual in-context learners for compositional tasks. Additionally, we investigate different masking-based training objectives to gather insights into how to train models better for solving complex, compositional tasks. Our exploration not only provides important insights especially for multi-modal medical task sequences but also highlights challenges that need to be addressed.
Simon Reiß, Zdravko Marinov, Alexander Jaus, Constantin Seibold, M. Saquib Sarfraz, Erik Rodner, Rainer Stiefelhagen
ICCV7
2025 SFDLA: Source-Free Document Layout Analysis
Sebastian Tewes, Yufan Chen 0001, Omar Moured, Jiaming Zhang 0001, Rainer Stiefelhagen
ICDAR (1)5
2025 RefChartQA: Grounding Visual Answer on Chart Images Through Instruction Tuning
Alexander Vogel, Omar Moured, Yufan Chen 0001, Jiaming Zhang 0001, Rainer Stiefelhagen
ICDAR (4)5
2025 Graph-based Document Structure Analysis
abstract
When reading a document, glancing at the spatial layout of a document is an initial step to understand it roughly. Traditional document layout analysis (DLA) methods, however, offer only a superficial parsing of documents, focusing on basic instance detection and often failing to capture the nuanced spatial and logical relationships between instances. These limitations hinder DLA-based models from achieving a gradually deeper comprehension akin to human reading. In this work, we propose a novel graph-based Document Structure Analysis (gDSA) task. This task requires that model not only detects document elements but also generates spatial and logical relations in form of a graph structure, allowing to understand documents in a holistic and intuitive manner. For this new task, we construct a relation graph-based document structure analysis dataset(GraphDoc) with 80K document images and 4.13M relation annotations, enabling training models to complete multiple tasks like reading order, hierarchical structures analysis, and complex inter-element relationship inference. Furthermore, a document relation graph generator (DRGG) is proposed to address the gDSA task, which achieves performance with 57.6% at $mAP_g$@$0.5$ for a strong benchmark baseline on this novel task and dataset. We hope this graphical representation of document structure can mark an innovative advancement in document structure analysis and understanding. The new dataset and code will be made publicly available.
Yufan Chen 0001, Ruiping Liu 0001, Junwei Zheng, Di Wen 0006, Kunyu Peng, Jiaming Zhang 0001, Rainer Stiefelhagen
ICLR7
2025 Exploring Self-supervised Skeleton-based Action Recognition in Occluded Environments
abstract
To integrate action recognition into autonomous robotic systems, it is essential to address challenges such as person occlusions—a common yet often overlooked scenario in existing self-supervised skeleton-based action recognition methods. In this work, we propose IosPSTL, a simple and effective self-supervised learning framework designed to handle occlusions. IosPSTL combines a cluster-agnostic KNN imputer with an Occluded Partial Spatio-Temporal Learning (OPSTL) strategy. First, we pre-train the model on occluded skeleton sequences. Then, we introduce a cluster-agnostic KNN imputer that performs semantic grouping using k-means clustering on sequence embeddings. It imputes missing skeleton data by applying K-Nearest Neighbors in the latent space, leveraging nearby sample representations to restore occluded joints. This imputation generates more complete skeleton sequences, which significantly benefits downstream self-supervised models. To further enhance learning, the OPSTL module incorporates Adaptive Spatial Masking (ASM) to make better use of intact, high-quality skeleton sequences during training. Our method achieves state-of-the-art performance on the occluded versions of the NTU-60 and NTU-120 datasets, demonstrating its robustness and effectiveness under challenging conditions. Code is available at https://github.com/cyfml/OPSTL.
Kunyu Peng, Alina Roitberg, David Schneider 0006, Jiaming Zhang 0001, Junwei Zheng, Yufan Chen 0001, Ruiping Liu 0001, Kailun Yang 0001, Rainer Stiefelhagen
IJCNN10
2025 VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and Invisibility
abstract
We propose VISO-Grasp, a novel vision-language-informed system designed to systematically address visibility constraints for grasping in severely occluded environments. By leveraging Foundation Models (FMs) for spatial reasoning and active view planning, our framework constructs and updates an instance-centric representation of spatial relationships, enhancing grasp success under challenging occlusions. Furthermore, this representation facilitates active Next-Best-View (NBV) planning and optimizes sequential grasping strategies when direct grasping is infeasible. Additionally, we introduce a multi-view uncertainty-driven grasp fusion mechanism that refines grasp confidence and directional uncertainty in real-time, ensuring robust and stable grasp execution. Extensive real-world experiments demonstrate that VISO-Grasp achieves a success rate of 87.5% in target-oriented grasping with the fewest grasp attempts outperforming baselines. To the best of our knowledge, VISO-Grasp is the first unified framework integrating FMs into target-aware active view planning and 6-DoF grasping in environments with severe occlusions and entire invisibility constraints. Code is available at: https://github.com/YitianShi/vMF-Contact
Yitian Shi, Di Wen 0006, Guanqi Chen, Edgar Welte, Kunyu Peng, Rainer Stiefelhagen, Rania Rayyes
IROS7
2025 Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
abstract
Physical environments and circumstances are fundamentally dynamic, yet current 3D datasets and evaluation benchmarks tend to concentrate on either dynamic scenarios or dynamic situations in isolation, resulting in incomplete comprehension. To overcome these constraints, we introduce Situat3DChange, an extensive dataset supporting three situation-aware change understanding tasks following the perception-action model: 121K question-answer pairs, 36K change descriptions for perception tasks, and 17K rearrangement instructions for the action task. To construct this large-scale dataset, Situat3DChange leverages 11K human observations of environmental changes to establish shared mental models and shared situational awareness for human-AI collaboration. These observations, enriched with egocentric and allocentric perspectives as well as categorical and coordinate spatial relations, are integrated using an LLM to support understanding of situated changes. To address the challenge of comparing pairs of point clouds from the same scene with minor changes, we propose SCReasoner, an efficient 3D MLLM approach that enables effective point cloud comparison with minimal parameter overhead and no additional tokens required for the language decoder. Comprehensive evaluation on Situat3DChange tasks highlights both the progress and limitations of MLLMs in dynamic scene and situation understanding. Additional experiments on data scaling and cross-domain transfer demonstrate the task-agnostic effectiveness of using Situat3DChange as a training dataset for MLLMs. The established dataset and source code are publicly available at: https://github.com/RuipingL/Situat3DChange.
Ruiping Liu 0001, Junwei Zheng, Yufan Chen 0001, Kunyu Peng, Kailun Yang 0001, Jiaming Zhang 0001, Marc Pollefeys, Rainer Stiefelhagen
NeurIPS9
2025 HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios
abstract
Action segmentation is a core challenge in high-level video understanding, aiming to partition untrimmed videos into segments and assign each a label from a predefined action set. Existing methods primarily address single-person activities with fixed action sequences, overlooking multi-person scenarios. In this work, we pioneer textual reference-guided human action segmentation in multi-person settings, where a textual description specifies the target person for segmentation. We introduce the first dataset for Referring Human Action Segmentation, i.e., RHAS133, built from 133 movies and annotated with 137 fine-grained actions with 33h video data, together with textual descriptions for this new task. Benchmarking existing action segmentation methods on RHAS133 using VLM-based feature extractors reveals limited performance and poor aggregation of visual cues for the target person. To address this, we propose a holistic-partial aware Fourier-conditioned diffusion framework, i.e., HopaDIFF, leveraging a novel cross-input gate attentional xLSTM to enhance holistic-partial long-range reasoning and a novel Fourier condition to introduce more fine-grained control to improve the action segmentation generation. HopaDIFF achieves state-of-the-art results on RHAS133 in diverse evaluation settings. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF.
Kunyu Peng, Junchao Huang, Xiangsheng Huang, Di Wen 0006, Junwei Zheng, Yufan Chen 0001, Kailun Yang 0001, Chongqing Hao, Rainer Stiefelhagen
NeurIPS10
2025 mmWalk: Towards Multi-modal Multi-view Walking Assistance
abstract
Walking assistance in extreme or complex environments remains a significant challenge for people with blindness or low vision (BLV), largely due to the lack of a holistic scene understanding. Motivated by the real-world needs of the BLV community, we build mmWalk, a simulated multi-modal dataset that integrates multi-view sensor and accessibility-oriented features for outdoor safe navigation. Our dataset comprises $120$ manually controlled, scenario-categorized walking trajectories with $62k$ synchronized frames. It contains over $559k$ panoramic images across RGB, depth, and semantic modalities. Furthermore, to emphasize real-world relevance, each trajectory involves outdoor corner cases and accessibility-specific landmarks for BLV users. Additionally, we generate mmWalkVQA, a VQA benchmark with over $69k$ visual question-answer triplets across $9$ categories tailored for safe and informed walking assistance. We evaluate state-of-the-art Vision-Language Models (VLMs) using zero- and few-shot settings and found they struggle with our risk assessment and navigational tasks. We validate our mmWalk-finetuned model on real-world datasets and show the effectiveness of our dataset for advancing multi-modal walking assistance.
Kedi Ying, Ruiping Liu 0001, Chongyan Chen, Mingzhe Tao, Hao Shi 0004, Kailun Yang 0001, Jiaming Zhang 0001, Rainer Stiefelhagen
NeurIPS8
2025 Exploring Video-Based Driver Activity Recognition under Noisy Labels
abstract
As an open research topic in the field of deep learning, learning with noisy labels has attracted much attention and grown rapidly over the past ten years. Learning with label noise is crucial for driver distraction behavior recognition, as real-world video data often contains mislabeled samples, impacting model reliability and performance. However, label noise learning is barely explored in the driver activity recognition field. In this paper, we propose the first label noise learning approach for the driver activity recognition task. Based on the cluster assumption, we initially enable the model to learn clustering-friendly low-dimensional representations from given videos and assign the resultant embeddings into clusters. We subsequently perform co-refinement within each cluster to smooth the classifier outputs. Furthermore, we propose a flexible sample selection strategy that combines two selection criteria without relying on any hyperparameters to filter clean samples from the training dataset. We also incorporate a self-adaptive parameter into the sample selection process to enforce balancing across classes. A comprehensive variety of experiments on the public Drive&Act dataset for all granularity levels demonstrates the superior performance of our method in comparison with other label-denoising methods derived from the image classification field. The source code is available at https://github.com/ilonafan/DAR-noisy-labels.
Linjuan Fan, Di Wen 0006, Kunyu Peng, Kailun Yang 0001, Jiaming Zhang 0001, Ruiping Liu 0001, Yufan Chen 0001, Junwei Zheng, Rainer Stiefelhagen
SMC11
2025 Snap, Segment, Deploy: A Visual Data and Detection Pipeline for Wearable Industrial Assistants
abstract
Industrial assembly requires rapid adaptation to complex procedures under constrained computing, connectivity, and privacy conditions, rendering cloud-based solutions impractical. We present an on-device assistant for real-time, semi-hands-free guidance, integrating lightweight detection, speech recognition, and retrieval-augmented response generation. To enable scalable training without manual labeling, we construct the Gear8 dataset via an automated pipeline and introduce a two-stage refinement strategy to enhance robustness against domain shift. Experiments show improved generalization under diverse corruptions. User studies confirm notable gains in efficiency and error reduction, underscoring the system’s suitability for real-world industrial deployment. The Gear8 dataset, models, and code are publicly available at: https://github.com/Kratos-Wen/Gear8.
Di Wen 0006, Junwei Zheng, Ruiping Liu 0001, Kunyu Peng, Rainer Stiefelhagen
SMC6
2025 @BENCH: Benchmarking Vision-Language Models for Human-centered Assistive Technology
abstract
As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However, benchmarking VLMs for ATs remains under-explored. To bridge this gap, we first create a novel AT benchmark (@ Bench). Guided by a pre-design user study with PVIs, our benchmark includes the five most crucial vision-language tasks: Panoptic Segmentation, Depth Estimation, Optical Character Recognition (OCR), Image Captioning, and Visual Question Answering (VQA). Besides, we propose a novel AT model (@MODEL) that addresses all tasks simultaneously and can be expanded to more assistive functions for helping PVIs. Our framework exhibits outstanding performance across tasks by integrating multi-modal information, and it offers PVIs a more comprehensive assistance. Extensive experiments prove the effectiveness and generalizability of our framework.
Junwei Zheng, Ruiping Liu 0001, Jiaming Zhang 0001, Sven Matthiesen, Rainer Stiefelhagen
WACV7
2024 Navigating Open Set Scenarios for Skeleton-Based Action Recognition
abstract
In real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions poses challenges due to the lack of visual background cues and the distinct sparse structure of body pose sequences. In this paper, we tackle the unexplored Open-Set Skeleton-based Action Recognition (OS-SAR) task and formalize the benchmark on three skeleton-based datasets. We assess the performance of seven established open-set approaches on our task and identify their limits and critical generalization issues when dealing with skeleton information.To address these challenges, we propose a distance-based cross-modality ensemble method that leverages the cross-modal alignment of skeleton joints, bones, and velocities to achieve superior open-set recognition performance. We refer to the key idea as CrossMax - an approach that utilizes a novel cross-modality mean max discrepancy suppression mechanism to align latent spaces during training and a cross-modality distance-based logits refinement method during testing. CrossMax outperforms existing approaches and consistently yields state-of-the-art results across all datasets and backbones. We will release the benchmark, code, and models to the community.
Kunyu Peng, Junwei Zheng, Ruiping Liu 0001, David Schneider 0006, Jiaming Zhang 0001, Kailun Yang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg
AAAI9
2024 Strike the Balance: On-the-Fly Uncertainty Based User Interactions for Long-Term Video Object Segmentation
Stéphane Vujasinovic, Stefan Becker, Sebastian Bullinger, Norbert Scherer-Negenborn, Michael Arens, Rainer Stiefelhagen
ACCV (2)6
2024 OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
Jiale Wei, Junwei Zheng, Ruiping Liu 0001, Jie Hu 0039, Jiaming Zhang 0001, Rainer Stiefelhagen
ACCV (10)6
2024 RoDLA: Benchmarking the Robustness of Document Layout Analysis Models
abstract
Before developing a Document Layout Analysis (DLA) model in real-world applications, conducting comprehensive robustness testing is essential. However, the robustness of DLA models remains underexplored in the literature. To address this, we are the first to introduce a robustness benchmark for DLA models, which includes 450K document images of three datasets. To cover realistic corruptions, we propose a perturbation taxonomy with 12 common document perturbations with 3 severity levels inspired by realworld document processing. Additionally, to better understand document perturbation impacts, we propose two metrics, Mean Perturbation Effect (mPE) for perturbation assessment and Mean Robustness Degradation (mRD) for robustness evaluation. Furthermore, we introduce a self-titled model, i.e., Robust Document Layout Analyzer (RoDLA), which improves attention mechanisms to boost extraction of robust features. Experiments on the proposed benchmarks (PubLayNet-P, DocLayNet-P, andM6Doc-P) demonstrate that RoDLA obtains state-of-the-art mRD scores of 115.7, 135.4, and 150.4, respectively. Compared to previous methods, RoDLA achieves notable improvements in mAP of +3.8%, +7.1% and +12.1%, respectively.
Yufan Chen 0001, Jiaming Zhang 0001, Kunyu Peng, Junwei Zheng, Ruiping Liu 0001, Philip Torr 0001, Rainer Stiefelhagen
CVPR7
2024 Occlusion-Aware Seamless Segmentation
Yihong Cao, Jiaming Zhang 0001, Hao Shi 0004, Kunyu Peng, Yuhongxuan Zhang, Hui Zhang 0023, Rainer Stiefelhagen, Kailun Yang 0001
ECCV (19)7
2024 Statewide Visual Geolocalization in the Wild
Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen
ECCV (36)5
2024 Referring Atomic Video Action Recognition
Kunyu Peng, Jia Fu 0001, Kailun Yang 0001, Di Wen 0006, Yufan Chen 0001, Ruiping Liu 0001, Junwei Zheng, Jiaming Zhang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg
ECCV (19)10
2024 Open Panoramic Segmentation
Junwei Zheng, Ruiping Liu 0001, Yufan Chen 0001, Kunyu Peng, Chengzhi Wu, Kailun Yang 0001, Jiaming Zhang 0001, Rainer Stiefelhagen
ECCV (39)8
2024 SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
abstract
Tu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Danni Liu, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao, Fabian Peller-Konrad, Tobias Röddiger, Alexander Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Tu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao 0002, Fabian Tërnava, Tobias Röddiger, Alex Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues
EMNLP15
2024 Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
abstract
Self-supervised representation learning for human action recognition has developed rapidly in recent years. Most of the existing works are based on skeleton data while using a multi-modality setup. These works overlooked the differences in performance among modalities, which led to the propagation of erroneous knowledge between modalities while only three fundamental modalities, i.e., joints, bones, and motions are used, hence no additional modalities are explored.In this work, we first propose an Implicit Knowledge Exchange Module (IKEM) which alleviates the propagation of erroneous knowledge between low-performance modalities. Then, we further propose three new modalities to enrich the complementary information between modalities. Finally, to maintain efficiency when introducing new modalities, we propose a novel teacher-student framework to distill the knowledge from the secondary modalities into the mandatory modalities considering the relationship constrained by anchors, positives, and negatives, named relational cross-modality knowledge distillation. The experimental results demonstrate the effectiveness of our approach, unlocking the efficient use of skeleton-based multi-modality data. Source code will be made publicly available at https://github.com/desehuileng0o0/IKEM.
Yiping Wei, Kunyu Peng, Alina Roitberg, Jiaming Zhang 0001, Junwei Zheng, Ruiping Liu 0001, Yufan Chen 0001, Kailun Yang 0001, Rainer Stiefelhagen
ICASSP9
2024 STS New Methods for Creating Accessible Material in Higher Education - Introduction to the Special Thematic Session
Michaela Hanousková, Boris Janca, Lukás Másilko, Karin Müller 0001, Svatoslav Ondra, Radek Pavlicek, Petr Penáz, Andrea Petz, Thorsten Schwarz, Rainer Stiefelhagen
ICCHP (1)10
2024 ChartFormer: A Large Vision Language Model for Converting Chart Images into Tactile Accessible SVGs
Omar Moured, Sara Alzalabny, Anas Osman, Thorsten Schwarz, Karin Müller 0001, Rainer Stiefelhagen
ICCHP (1)6
2024 Alt4Blind: A User Interface to Simplify Charts Alt-Text Creation
Omar Moured, Shahid Ali Farooqui, Karin Müller 0001, Sharifeh Fadaeijouybari, Thorsten Schwarz, Mohammed Javed, Rainer Stiefelhagen
ICCHP (1)7
2024 ACCSAMS: Automatic Conversion of Exam Documents to Accessible Learning Material for Blind and Visually Impaired
David Wilkening, Omar Moured, Thorsten Schwarz, Karin Müller 0001, Rainer Stiefelhagen
ICCHP (1)5
2024 AltChart: Enhancing VLM-Based Chart Summarization Through Multi-pretext Tasks
Omar Moured, Jiaming Zhang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen
ICDAR (1)4
2024 Towards Unifying Anatomy Segmentation: Automated Generation of a Full-Body CT Dataset
abstract
In this paper, we present a method for generating automated anatomy segmentation datasets using a sequential process that involves nnU-Net-based pseudo-labeling and anatomy-guided pseudo-label refinement. By combining various fragmented knowledge bases, we generate a dataset of whole-body CT scans with 142 voxel-level labels for 533 volumes providing comprehensive anatomical coverage. We validate its usefulness via Human expert evaluation and medical validity. This dataset enables the analysis of whole-body anatomy segmentation for cancer patients. Besides the DAP Atlas dataset, we release our trained anatomy segmentation models capable of predicting 142 anatomical structures on CT data.
Alexander Jaus, Constantin Seibold, Kelsey Hermann, Negar Shahamiri, Alexandra Walter, Kristina Giske, Johannes Haubold, Jens Kleesiek, Rainer Stiefelhagen
ICIP9
2024 SynthAct: Towards Generalizable Human Action Recognition based on Synthetic Data
abstract
Synthetic data generation is a proven method for augmenting training sets without the need for extensive setups, yet its application in human activity recognition is underexplored. This is particularly crucial for human-robot collaboration in household settings, where data collection is often privacy-sensitive. In this paper, we introduce SynthAct, a synthetic data generation pipeline designed to significantly minimize the reliance on real-world data. Leveraging modern 3D pose estimation techniques, SynthAct can be applied to arbitrary 2D or 3D video action recordings, making it applicable for uncontrolled in-the-field recordings by robotic agents or smarthome monitoring systems. We present two SynthAct datasets: AMARV, a large synthetic collection with over 800k multi-view action clips, and Synthetic Smarthome, mirroring the Toyota Smarthome dataset. SynthAct generates a rich set of data, including RGB videos and depth maps from four synchronized views, 3D body poses, normal maps, segmentation masks and bounding boxes. We validate the efficacy of our datasets through extensive synthetic-to-real experiments on NTU RGB+D and Toyota Smarthome. SynthAct is available on our project page4.
David Schneider 0006, Marco Keller, Zeyun Zhong, Kunyu Peng, Alina Roitberg, Jürgen Beyerer, Rainer Stiefelhagen
ICRA7
2024 MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments
abstract
People with Visual Impairments (PVI) typically recognize objects through haptic perception. Knowing objects and materials before touching is desired by the target users but under-explored in the field of human-centered robotics. To fill this gap, in this work, a wearable vision-based robotic system, MATERobot, is established for PVI to recognize materials and object categories beforehand. To address the computational constraints of mobile platforms, we propose a lightweight yet accurate model MATEViT to perform pixel-wise semantic segmentation, simultaneously recognizing both objects and materials. Our methods achieve respective 40.2% and 51.1% of mIoU on COCOStuff-10K and DMS datasets, surpassing the previous method with +5.7% and +7.0% gains. Moreover, on the field test with participants, our wearable system reaches a score of 28 in the NASA-Task Load Index, indicating low cognitive demands and ease of use. Our MATERobot demonstrates the feasibility of recognizing material property through visual cues and offers a promising step towards improving the functionality of wearable robots for PVI. The source code has been made publicly available at MATERobot.
Junwei Zheng, Jiaming Zhang 0001, Kailun Yang 0001, Kunyu Peng, Rainer Stiefelhagen
ICRA5
2024 Skeleton-Based Human Action Recognition with Noisy Labels
abstract
Understanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and annotation of activity sequences is time-consuming and the resulting labels are often noisy. If not effectively addressed, label noise negatively affects the model’s training, resulting in lower recognition quality. Despite its importance, addressing label noise for skeleton-based action recognition has been overlooked so far. In this study, we bridge this gap by implementing a framework that augments well-established skeleton-based human action recognition methods with label-denoising strategies from various research areas to serve as the initial benchmark. Observations reveal that these baselines yield only marginal performance when dealing with sparse skeleton data. Consequently, we introduce a novel methodology, NoiseEraSAR, which integrates global sample selection, co-teaching, and Cross-Modal Mixture-of-Experts (CM-MOE) strategies, aimed at mitigating the adverse impacts of label noise. Our proposed approach demonstrates better performance on the established benchmark, setting new state-of-the-art standards. The source code for this study will be made accessible at https://github.com/xuyizdby/NoiseEraSAR.
Kunyu Peng, Di Wen 0006, Ruiping Liu 0001, Junwei Zheng, Yufan Chen 0001, Jiaming Zhang 0001, Alina Roitberg, Kailun Yang 0001, Rainer Stiefelhagen
IROS10
2024 Chart4Blind: An Intelligent Interface for Chart Accessibility Conversion
abstract
In a world driven by data visualization, ensuring the inclusive accessibility of charts for Blind and Visually Impaired (BVI) individuals remains a significant challenge. Charts are usually presented as raster graphics without textual and visual metadata needed for an equivalent exploration experience for BVI people. Additionally, converting these charts into accessible formats requires considerable effort from sighted individuals. Digitizing charts with metadata extraction is just one aspect of the issue; transforming it into accessible modalities, such as tactile graphics, presents another difficulty. To address these disparities, we propose Chart4Blind, an intelligent user interface that converts bitmap image representations of line charts into universally accessible formats. Chart4Blind achieves this transformation by generating Scalable Vector Graphics (SVG), Comma-Separated Values (CSV), and alternative text exports, all comply with established accessibility standards. Through interviews and a formal user study, we demonstrate that even inexperienced sighted users can make charts accessible in an average of 4 minutes using Chart4Blind, achieving a System Usability Scale rating of 90%. In comparison to existing approaches, Chart4Blind provides a comprehensive solution, generating end-to-end accessible SVGs suitable for assistive technologies such as embossed prints (papers and laser cut), 2D tactile displays, and screen readers. For additional information, including open-source codes and demos, please visit our project page https://moured.github.io/chart4blind/.
Omar Moured, Morris Baumgarten-Egemole, Karin Müller 0001, Alina Roitberg, Thorsten Schwarz, Rainer Stiefelhagen
IUI6
2024 Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation
abstract
Integrating information from multiple modalities enhances the robustness of scene perception systems in autonomous vehicles, providing a more comprehensive and reliable sensory framework. However, the modality incompleteness in multi-modal segmentation remains under-explored. In this work, we establish a task called Modality-Incomplete Scene Segmentation (MISS), which encompasses both system-level modality absence and sensor-level modality errors. To avoid the predominant modality reliance in multi-modal fusion, we introduce a Missing-aware Modal Switch (MMS) strategy to proactively manage missing modalities during training. Utilizing bit-level batch-wise sampling enhances the model’s performance in both complete and incomplete testing scenarios. Furthermore, we introduce the Fourier Prompt Tuning (FPT) method to incorporate representative spectral information into a limited number of learnable prompts that maintain robustness against all MISS scenarios. Akin to fine-tuning effects but with fewer tunable parameters (1.1%). Extensive experiments prove the efficacy of our proposed approach, showcasing an improvement of 5.84% mIoU over the prior state-of-the-art parameter-efficient methods in modality missing. The source code is publicly available at https://github.com/RuipingL/MISS.
Ruiping Liu 0001, Jiaming Zhang 0001, Kunyu Peng, Yufan Chen 0001, Junwei Zheng, M. Saquib Sarfraz, Kailun Yang 0001, Rainer Stiefelhagen
IV9
2024 Anatomy-Guided Pathology Segmentation
Alexander Jaus, Constantin Seibold, Simon Reiß, Lukas Heine, Anton Schily, Moon S. Kim 0002, Fin Hendrik Bahnsen, Ken Herrmann, Rainer Stiefelhagen, Jens Kleesiek
MICCAI (8)9
2024 Towards Video-based Activated Muscle Group Estimation in the Wild
abstract
In this paper, we tackle the new task of video-based Activated Muscle Group Estimation (AMGE) aiming at identifying active muscle regions during physical activity in the wild.To this intent, we provide the MuscleMap dataset featuring >15𝐾 video clips with 135 different activities and 20 labeled muscle groups.This dataset opens the vistas to multiple video-based applications in sports and rehabilitation medicine under flexible environment constraints.The proposed MuscleMap dataset is constructed with YouTube videos, specifically targeting High-Intensity Interval Training (HIIT) physical exercise in the wild.To make the AMGE model applicable in real-life situations, it is crucial to ensure that the model can generalize well to numerous types of physical activities not present during training and involving new combinations of activated muscles.To achieve this, our benchmark also covers an evaluation setting where the model is exposed to activity types excluded from the training set.Our experiments reveal that the generalizability of existing architectures adapted for the AMGE task remains a challenge.Therefore, we also propose a new approach, TransM 3 E, which employs a multi-modality feature fusion mechanism between both the video transformer model and the skeleton-based graph convolution model with novel cross-modal knowledge distillation executed on multiclassification tokens.The proposed method surpasses all popular video classification models when dealing with both, previously seen and new types of physical activities.The database and code can be found at https://github.com/KPeng9510/MuscleMap.
Kunyu Peng, David Schneider 0006, Alina Roitberg, Kailun Yang 0001, Jiaming Zhang 0001, Chen Deng, M. Saquib Sarfraz, Rainer Stiefelhagen
ACM Multimedia9
2024 Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain Scheduler
abstract
In Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories are present at test time. The challenges of this task arise from the dual need to generalize across diverse domains and accurately quantify category novelty, which is critical for applications in dynamic environments. Recently, meta-learning techniques have demonstrated superior results in OSDG, effectively orchestrating the meta-train and -test tasks by employing varied random categories and predefined domain partition strategies. These approaches prioritize a well-designed training schedule over traditional methods that focus primarily on data augmentation and the enhancement of discriminative feature learning. The prevailing meta-learning models in OSDG typically utilize a predefined sequential domain scheduler to structure data partitions. However, a crucial aspect that remains inadequately explored is the influence brought by strategies of domain schedulers during training. In this paper, we observe that an adaptive domain scheduler benefits more in OSDG compared with prefixed sequential and random domain schedulers. We propose the Evidential Bi-Level Hardest Domain Scheduler (EBiL-HaDS) to achieve an adaptive domain scheduler. This method strategically sequences domains by assessing their reliabilities in utilizing a follower network, trained with confidence scores learned in an evidential manner, regularized by max rebiasing discrepancy, and optimized in a bilevel manner. We verify our approach on three OSDG benchmarks, i.e., PACS, DigitsDG, and OfficeHome. The results show that our method substantially improves OSDG performance and achieves more discriminative embeddings for both the seen and unseen categories, underscoring the advantage of a judicious domain scheduler for the generalizability to unseen domains and unseen categories. The source code is publicly available at https://github.com/KPeng9510/EBiL-HaDS.
Kunyu Peng, Di Wen 0006, Kailun Yang 0001, Ao Luo, Yufan Chen 0001, Jia Fu 0001, M. Saquib Sarfraz, Alina Roitberg, Rainer Stiefelhagen
NeurIPS9
2024 Muscles in Time: Learning to Understand Human Motion In-Depth by Simulating Muscle Activations
abstract
Exploring the intricate dynamics between muscular and skeletal structures is pivotal for understanding human motion. This domain presents substantial challenges, primarily attributed to the intensive resources required for acquiring ground truth muscle activation data, resulting in a scarcity of datasets.In this work, we address this issue by establishing Muscles in Time (MinT), a large-scale synthetic muscle activation dataset.For the creation of MinT, we enriched existing motion capture datasets by incorporating muscle activation simulations derived from biomechanical human body models using the OpenSim platform, a common framework used in biomechanics and human motion research.Starting from simple pose sequences, our pipeline enables us to extract detailed information about the timing of muscle activations within the human musculoskeletal system.Muscles in Time contains over nine hours of simulation data covering 227 subjects and 402 simulated muscle strands. We demonstrate the utility of this dataset by presenting results on neural network-based muscle activation estimation from human pose sequences with two different sequence-to-sequence architectures.
David Schneider 0006, Simon Reiß, Marco Kugler, Alexander Jaus, Kunyu Peng, Susanne Sutschet, M. Saquib Sarfraz, Sven Matthiesen, Rainer Stiefelhagen
NeurIPS9
2024 360BEV: Panoramic Semantic Mapping for Indoor Bird's-Eye View
abstract
Seeing only a tiny part of the whole is not knowing the full circumstance. Bird’s-eye-view (BEV) perception, a process of obtaining allocentric maps from egocentric views, is restricted when using a narrow Field of View (FoV) alone. In this work, mapping from 360° panoramas to BEV semantics, the 360BEV task, is established for the first time to achieve holistic representations of indoor scenes in a top-down view. Instead of relying on narrow-FoV image sequences, a panoramic image with depth information is sufficient to generate a holistic BEV semantic map. To benchmark 360BEV, we present two indoor datasets, 360BEV-Matterport and 360BEV-Stanford, both of which include egocentric panoramic images and semantic segmentation labels, as well as allocentric semantic maps. Besides delving deep into different mapping paradigms, we propose a dedicated solution for panoramic semantic mapping, namely 360Mapper. Through extensive experiments, our methods achieve 44.32% and 45.78% mIoU on both datasets respectively, surpassing previous counterparts with gains of +7.60% and +9.70% in mIoU.1
Zhifeng Teng, Jiaming Zhang 0001, Kailun Yang 0001, Kunyu Peng, Hao Shi 0004, Simon Reiß, Rainer Stiefelhagen
WACV8
2024 Recognizing affective states from the expressive behavior of tennis players using convolutional neural networks
abstract
This study describes an AI model by leveraging advanced Convolutional Neural Networks (CNNs) to recognize affective states in real-world sports settings, particularly tennis matches. In contrast to prior studies that primarily utilized data acquired from actors and rudimentary statistical methods, the present research emphasizes the analysis of bodily expressions in real-life contexts, aiming for a more naturalistic representation of human emotions. Our CNN-based models demonstrate an accuracy rate of up to 68.9%, outperforming or matching human observers in many instances. Intriguingly, both the machine learning models and human observers exhibited a shared propensity to more effectively identify negative affective states, which may be attributed to the more intense and straightforward expression of these states. These results not only advance the state of the art in affective state recognition but also pave the way for broader applications, including in healthcare and automotive safety sectors, thereby constituting a significant advancement in the development of sophisticated and universally applicable emotional recognition systems.
Darko Jekauc, Diana Burkart, Julian Fritsch, Marc Hesenius, Ole Meyer, M. Saquib Sarfraz, Rainer Stiefelhagen
Knowl. Based Syst.7
2024 Deep Interactive Segmentation of Medical Images: A Systematic Review and Taxonomy
abstract
Interactive segmentation is a crucial research area in medical image analysis aiming to boost the efficiency of costly annotations by incorporating human feedback. This feedback takes the form of clicks, scribbles, or masks and allows for iterative refinement of the model output so as to efficiently guide the system towards the desired behavior. In recent years, deep learning-based approaches have propelled results to a new level causing a rapid growth in the field with 121 methods proposed in the medical imaging domain alone. In this review, we provide a structured overview of this emerging field featuring a comprehensive taxonomy, a systematic review of existing methods, and an in-depth analysis of current practices. Based on these contributions, we discuss the challenges and opportunities in the field. For instance, we find that there is a severe lack of comparison across methods which needs to be tackled by standardized baselines and benchmarks.
Zdravko Marinov, Paul F. Jaeger, Jan Egger, Jens Kleesiek, Rainer Stiefelhagen
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Behind Every Domain There is a Shift: Adapting Distortion-Aware Vision Transformers for Panoramic Semantic Segmentation
abstract
In this paper, we address panoramic semantic segmentation which is under-explored due to two critical challenges: (1) image distortions and object deformations on panoramas; (2) lack of semantic annotations in the$360^\circ$imagery. To tackle these problems, first, we propose the upgraded Transformer for Panoramic Semantic Segmentation, ie, Trans4PASS+, equipped withDeformable Patch Embedding (DPE)andDeformable MLP (DMLPv2)modules for handling object deformations and image distortions whenever (before or after adaptation) and wherever (shallow or deep levels). Second, we enhance theMutual Prototypical Adaptation (MPA)strategy via pseudo-label rectification for unsupervised domain adaptive panoramic segmentation. Third, aside from Pinhole-to-Panoramic (Pin2Pan) adaptation, we create a new dataset (SynPASS) with 9,080 panoramic images, facilitating Synthetic-to-Real (Syn2Real) adaptation scheme in$360^\circ$imagery. Extensive experiments are conducted, which cover indoor and outdoor scenarios, and each of them is investigated withPin2PanandSyn2Realregimens. Trans4PASS+ achieves state-of-the-art performances on four domain adaptive panoramic semantic segmentation benchmarks. Code is available athttps://github.com/jamycheung/Trans4PASS.
Jiaming Zhang 0001, Kailun Yang 0001, Hao Shi 0004, Simon Reiß, Kunyu Peng, Chaoxiang Ma, Haodong Fu, Philip Torr 0001, Kaiwei Wang, Rainer Stiefelhagen
IEEE Trans. Pattern Anal. Mach. Intell.10
2024 CoBEV: Elevating Roadside 3D Object Detection With Depth and Height Complementarity
abstract
Roadside camera-driven 3D object detection is a crucial task in intelligent transportation systems, which extends the perception range beyond the limitations of vision-centric vehicles and enhances road safety. While previous studies have limitations in using only depth or height information, we find both depth and height matter and they are in fact complementary. The depth feature encompasses precise geometric cues, whereas the height feature is primarily focused on distinguishing between various categories of height intervals, essentially providing semantic context. This insight motivates the development of Complementary-BEV (CoBEV), a novel end-to-end monocular 3D object detection framework that integrates depth and height to construct robust BEV representations. In essence, CoBEV estimates each pixel's depth and height distribution and lifts the camera features into 3D space for lateral fusion using the newly proposed two-stage complementary feature selection (CFS) module. A BEV feature distillation framework is also seamlessly integrated to further enhance the detection accuracy from the prior knowledge of the fusion-modal CoBEV teacher. We conduct extensive experiments on the public 3D detection benchmarks of roadside camera-based DAIR-V2X-I and Rope3D, as well as the private Supremind-Road dataset, demonstrating that CoBEV not only achieves the accuracy of the new state-of-the-art, but also significantly advances the robustness of previous methods in challenging long-distance scenarios and noisy camera disturbance, and enhances generalization by a large margin in heterologous settings with drastic changes in scene and camera parameters. For the first time, the vehicle AP score of a camera model reaches 80% on DAIR-V2X-I in terms of easy mode. The source code will be made publicly available at CoBEV.
Hao Shi 0004, Chengshan Pang, Jiaming Zhang 0001, Kailun Yang 0001, Huajian Ni, Yining Lin, Rainer Stiefelhagen, Kaiwei Wang
IEEE Trans. Image Process.8
2024 EchoTrack: Auditory Referring Multi-Object Tracking for Autonomous Driving
abstract
This paper introduces the task of Auditory Referring Multi-Object Tracking (AR-MOT), which dynamically tracks specific objects in a video sequence based on audio expressions and appears as a challenging problem in autonomous driving. Due to the lack of semantic modeling capacity in audio and video, existing works have mainly focused on text-based multi-object tracking, which often comes at the cost of tracking quality, interaction efficiency, and even the safety of assistance systems, limiting the application of such methods in autonomous driving. In this paper, we delve into the problem of AR-MOT from the perspective of audio-video fusion and audio-video tracking. We put forward EchoTrack, an end-to-end AR-MOT framework with dual-stream vision transformers. The dual streams are intertwined with our Bidirectional Frequency-domain Cross-attention Fusion Module (Bi-FCFM), which bidirectionally fuses audio and video features from both frequency- and spatiotemporal domains. Moreover, we propose the Audio-visual Contrastive Tracking Learning (ACTL) regime to extract homogeneous semantic features between expressions and visual objects by learning homogeneous features between different audio and video objects effectively. Aside from the architectural design, we establish the first set of large-scale AR-MOT benchmarks, including Echo-KITTI, Echo-KITTI+, and Echo-BDD. Extensive experiments on the established benchmarks demonstrate the effectiveness of the proposed EchoTrack and its components. The source code and datasets are available athttps://github.com/lab206/EchoTrack.
Jiacheng Lin, Kunyu Peng, Zhiyong Li 0001, Rainer Stiefelhagen, Kailun Yang 0001
IEEE Trans. Intell. Transp. Syst.6
2024 TransKD: Transformer Knowledge Distillation for Efficient Semantic Segmentation
abstract
Semantic segmentation benchmarks in the realm of autonomous driving are dominated by large pre-trained transformers, yet their widespread adoption is impeded by substantial computational costs and prolonged training durations. To lift this constraint, we look at efficient semantic segmentation from a perspective of comprehensive knowledge distillation and aim to bridge the gap between multi-source knowledge extractions and transformer-specific patch embeddings. We put forward the Transformer-based Knowledge Distillation (TransKD) framework which learns compact student transformers by distilling both feature maps and patch embeddings of large teacher transformers, bypassing the long pre-training process and reducing the FLOPs by >85.0%. Specifically, we propose two fundamental modules to realize feature map distillation and patch embedding distillation, respectively: 1) Cross Selective Fusion (CSF) enables knowledge transfer between cross-stage features via channel attention and feature map distillation within hierarchical transformers; 2) Patch Embedding Alignment (PEA) performs dimensional transformation within the patchifying process to facilitate the patch embedding distillation. Furthermore, we introduce two optimization modules to enhance the patch embedding distillation from different perspectives: 1) Global-Local Context Mixer (GL-Mixer) extracts both global and local information of a representative embedding; 2) Embedding Assistant (EA) acts as an embedding method to seamlessly bridge teacher and student models with the teacher’s number of channels. Experiments on Cityscapes, ACDC, NYUv2, and Pascal VOC2012 datasets show that TransKD outperforms state-of-the-art distillation frameworks and rivals the time-consuming pre-training method. The source code is publicly available athttps://github.com/RuipingL/TransKD.
Ruiping Liu 0001, Kailun Yang 0001, Alina Roitberg, Jiaming Zhang 0001, Kunyu Peng, Huayao Liu, Yaonan Wang 0001, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.8
2023 READMem: Robust Embedding Association for a Diverse Memory in Unconstrained Video Object Segmentation
Stéphane Vujasinovic, Sebastian Bullinger, Stefan Becker, Norbert Scherer-Negenborn, Michael Arens, Rainer Stiefelhagen
BMVC6
2023 Delivering Arbitrary-Modal Semantic Segmentation
abstract
Multimodal fusion can make semantic segmentation more robust. However, fusing an arbitrary number of modalities remains underexplored. To delve into this problem, we create the Deliver arbitrary-modal segmentation benchmark, covering Depth, LiDAR, multiple Views, Events, and RGB. Aside from this, we provide this dataset in four severe weather conditions as well as five sensor failure cases to exploit modal complementarity and resolve partial outages. To make this possible, we present the arbitrary cross-modal segmentation model CMNEXT. It encompasses a Self-Query Hub (SQ-Hub) designed to extract effective information from any modality for subsequent fusion with the RGB representation and adds only negligible amounts of parameters ~0.1M) per additional modality. On top, to efficiently and flexibly harvest discriminative cues from the auxiliary modalities, we introduce the simple Parallel Pooling Mixer (PPX). With extensive experiments on a total of six benchmarks, our CMNEXT achieves state-of-the-art performance on the Deliver, Kitti-360, MFNet, NYU Depth V2, UrbanLF, and MCubeS datasets, allowing to scale from 1 to 81 modalities. On the freshly collected Deliver, the quad-modal CMNEXT reaches up to 66.30% in mIoU with a +9.10% gain as compared to the mono-modal baseline.11The Deliver dataset and our code will be made publicly available at https://jamycheung.github.io/DELIVER.html.
Jiaming Zhang 0001, Ruiping Liu 0001, Hao Shi 0004, Kailun Yang 0001, Simon Reiß, Kunyu Peng, Haodong Fu, Kaiwei Wang, Rainer Stiefelhagen
CVPR9
2023 Uncertainty-Aware Vision-Based Metric Cross-View Geolocalization
abstract
This paper proposes a novel method for vision-based metric cross-view geolocalization (CVGL) that matches the camera images captured from a ground-based vehicle with an aerial image to determine the vehicle's geo-pose. Since aerial images are globally available at low cost, they represent a potential compromise between two established paradigms of autonomous driving, i.e. using expensive high-definition prior maps or relying entirely on the sensor data captured at runtime. We present an end-to-end differentiable model that uses the ground and aerial images to predict a probability distribution over possible vehicle poses. We combine multiple vehicle datasets with aerial images from orthophoto providers on which we demonstrate the feasibility of our method. Since the ground truth poses are often inaccurate w.r.t. the aerial images, we implement a pseudo-label approach to produce more accurate ground truth poses and make them publicly available. While previous works require training data from the target region to achieve reasonable localization accuracy (i.e. same-area evaluation), our approach overcomes this limitation and outperforms previous results even in the strictly more challenging cross-area case. We improve the previous state-of-the-art by a large margin even without ground or aerial data from the test region, which highlights the model's potential for global-scale application. We further integrate the uncertainty-aware predictions in a tracking framework to determine the vehicle's trajectory over time resulting in a mean position error on KITTI-360 of 0.78m.
Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen
CVPR5
2023 Decoupled Semantic Prototypes enable learning from diverse annotation types for semi-weakly segmentation in expert-driven domains
abstract
A vast amount of images and pixel-wise annotations allowed our community to build scalable segmentation solutions for natural domains. However, the transfer to expert-driven domains like microscopy applications or medical healthcare remains difficult as domain experts are a critical factor due to their limited availability for providing pixel-wise annotations. To enable affordable segmentation solutions for such domains, we need training strategies which can simultaneously handle diverse annotation types and are not bound to costly pixel-wise annotations. In this work, we analyze existing training algorithms towards their flexibility for different annotation types and scalability to small annotation regimes. We conduct an extensive evaluation in the challenging domain of organelle segmentation and find that existing semi- and semi-weakly supervised training algorithms are not able to fully exploit diverse annotation types. Driven by our findings, we introduce Decoupled Semantic Prototypes (DSP) as a training method for semantic segmentation which enables learning from annotation types as diverse as image-level-, point-, bounding box-, and pixel-wise annotations and which leads to remarkable accuracy gains over existing solutions for semi-weakly segmentation.
Simon Reiß, Constantin Seibold, Alexander Freytag, Erik Rodner, Rainer Stiefelhagen
CVPR5
2023 Line Graphics Digitization: A Step Towards Full Automation
Omar Moured, Jiaming Zhang 0001, Alina Roitberg, Thorsten Schwarz, Rainer Stiefelhagen
ICDAR (5)5
2023 Quantized Distillation: Optimizing Driver Activity Recognition Models for Resource-Constrained Environments
abstract
Deep learning-based models are at the top of most driver observation benchmarks due to their remarkable accuracies but come with a high computational cost, while the resources are often limited in real-world driving scenarios. This paper presents a lightweight framework for resource- efficient driver activity recognition. We enhance 3D MobileNet, a speed-optimized neural architecture for video classification, with two paradigms for improving the trade-off between model accuracy and computational efficiency: knowledge distillation and model quantization. Knowledge distillation prevents large drops in accuracy when reducing the model size by harvesting knowledge from a large teacher model (I3D) via soft labels instead of using the original ground truth. Quantization further drastically reduces the memory and computation requirements by representing the model weights and activations using lower precision integers. Extensive experiments on a public dataset for in-vehicle monitoring during autonomous driving show that our proposed framework leads to an 3- fold reduction in model size and 1.4-fold improvement in inference time compared to an already speed-optimized architecture. Our code is available at https://github.com/calvintanama/qd-driver-activity-reco.
Calvin Tanama, Kunyu Peng, Zdravko Marinov, Rainer Stiefelhagen, Alina Roitberg
IROS4
2023 Guiding the Guidance: A Comparative Analysis of User Guidance Signals for Interactive Segmentation of Volumetric Images
Zdravko Marinov, Rainer Stiefelhagen, Jens Kleesiek
MICCAI (3)2
2023 Trans4Map: Revisiting Holistic Bird's-Eye-View Mapping from Egocentric Images to Allocentric Semantics with Vision Transformers
abstract
Humans have an innate ability to sense their surroundings, as they can extract the spatial representation from the egocentric perception and form an allocentric semantic map via spatial transformation and memory updating. However, endowing mobile agents with such a spatial sensing ability is still a challenge, due to two difficulties: (1) the previous convolutional models are limited by the local receptive field, thus, struggling to capture holistic long-range dependencies during observation; (2) the excessive computational budgets required for success, often lead to a separation of the mapping pipeline into stages, resulting the en-tire mapping process inefficient. To address these issues, we propose an end-to-end one-stage Transformer-based frame-work for Mapping, termed Trans4Map. Our egocentric-to-allocentric mapping process includes three steps: (1) the efficient transformer extracts the contextual features from a batch of egocentric images; (2) the proposed Bidirectional Allocentric Memory (BAM) module projects egocentric features into the allocentric memory; (3) the map de-coder parses the accumulated memory and predicts the top-down semantic segmentation map. In contrast, Trans4Map achieves state-of-the-art results, reducing 67.2% parameters, yet gaining a +3.25% mIoU and a +4.09% mBF1 improvements on the Matterport3D dataset.1
Jiaming Zhang 0001, Kailun Yang 0001, Kunyu Peng, Rainer Stiefelhagen
WACV5
2023 Anticipative Feature Fusion Transformer for Multi-Modal Action Anticipation
abstract
Although human action anticipation is a task which is inherently multi-modal, state-of-the-art methods on well known action anticipation datasets leverage this data by applying ensemble methods and averaging scores of uni-modal anticipation networks. In this work we introduce transformer based modality fusion techniques, which unify multi-modal data at an early stage. Our Anticipative Feature Fusion Transformer (AFFT) proves to be superior to popular score fusion approaches and presents state-of-the-art results outperforming previous methods on EpicKitchens-100 and EGTEA Gaze+. Our model is easily extensible and allows for adding new modalities without architectural changes. Consequently, we extracted audio features on EpicKitchens-100 which we add to the set of commonly used features in the community.1
Zeyun Zhong, David Schneider 0006, Michael Voit, Rainer Stiefelhagen, Jürgen Beyerer
WACV4
2023 Panoramic Panoptic Segmentation: Insights Into Surrounding Parsing for Mobile Agents via Unsupervised Contrastive Learning
abstract
In this work, we introduce panoramic panoptic segmentation, as the most holistic scene understanding, both in terms of Field of View (FoV) and image-level understanding for standard camera-based input. A complete surrounding understanding provides a maximum of information to a mobile agent. This is essential information for any intelligent vehicle to make informed decisions in a safety-critical dynamic environment such as real-world traffic. In order to overcome the lack of annotated panoramic images, we propose a framework which allows model training on standard pinhole images and transfers the learned features to the panoramic domain in a cost-minimizing way. The domain shift from pinhole to panoramic images is non-trivial as large objects and surfaces are heavily distorted close to the image border regions and look different across the two domains. Using our proposed method with dense contrastive learning, we manage to achieve significant improvements over a non-adapted approach. Depending on the efficient panoptic segmentation architecture, we can improve 3.5–6.5% measured in Panoptic Quality (PQ) over non-adapted models on our established Wild Panoramic Panoptic Segmentation (WildPPS) dataset. Furthermore, our efficient framework does not need access to the images of the target domain, making it a feasible domain generalization approach suitable for a limited hardware setting. As additional contributions, we publish WildPPS: The first panoramic panoptic image dataset to foster progress in surrounding perception and explore a novel training procedure combining supervised and contrastive training.
Alexander Jaus, Kailun Yang 0001, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.3
2023 CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation With Transformers
abstract
Scene understanding based on image segmentation is a crucial component of autonomous vehicles. Pixel-wise semantic segmentation of RGB images can be advanced by exploiting complementary features from the supplementary modality (${X}$-modality). However, covering a wide variety of sensors with a modality-agnostic model remains an unresolved problem due to variations in sensor characteristics among different modalities. Unlike previous modality-specific methods, in this work, we propose a unified fusion framework, CMX, for RGB-X semantic segmentation. To generalize well across different modalities, that often include supplements as well as uncertainties, a unified cross-modal interaction is crucial for modality fusion. Specifically, we design a Cross-Modal Feature Rectification Module (CM-FRM) to calibrate bi-modal features by leveraging the features from one modality to rectify the features of the other modality. With rectified feature pairs, we deploy a Feature Fusion Module (FFM) to perform sufficient exchange of long-range contexts before mixing. To verify CMX, for the first time, we unify five modalities complementary to RGB, i.e., depth, thermal, polarization, event, and LiDAR. Extensive experiments show that CMX generalizes well to diverse multi-modal fusion, achieving state-of-the-art performances on five RGB-Depth benchmarks, as well as RGB-Thermal, RGB-Polarization, and RGB-LiDAR datasets. Besides, to investigate the generalizability to dense-sparse data fusion, we establish an RGB-Event semantic segmentation benchmark based on the EventScape dataset, on which CMX sets the new state-of-the-art. The source code of CMX is publicly available athttps://github.com/huaaaliu/RGBX_Semantic_Segmentation.
Jiaming Zhang 0001, Huayao Liu, Kailun Yang 0001, Xinxin Hu, Ruiping Liu 0001, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.6
2023 Delving Deep Into One-Shot Skeleton-Based Action Recognition With Diverse Occlusions
abstract
Occlusions areuniversal disruptions constantly present in the real world. Especially for sparse representations, such as human skeletons, a few occluded points might destroy the geometrical and temporal continuity critically affecting the results. Yet, the research of data-scarce recognition from skeleton sequences, such as one-shot action recognition, does not explicitly consider occlusions despite their everyday pervasiveness. In this work, we explicitly tackle body occlusions forSkeleton-basedOne-shotActionRecognition (SOAR). We mainly consider two occlusion variants: 1) random occlusions and 2) more realistic occlusions caused by diverse everyday objects, which we generate by projecting the existing IKEA 3D furniture models into the camera coordinate system of the 3D skeletons with different geometric parameters, (e.g., rotation and displacement). We leverage the proposed pipeline to blend out portions of skeleton sequences of the three popular action recognition datasets (NTU-120, NTU-60 and Toyota Smart Home) and formalize the first benchmark for SOAR from partially occluded body poses. This is the first benchmark which considers occlusions for data-scarce action recognition. Another key property of our benchmark are the more realistic occlusions generated by everyday objects, as even in standard recognition from 3D skeletons, only randomly missing joints were considered. We re-evaluate existing state-of-the-art frameworks for SOAR in the light of this new task and further introduceTrans4SOAR– a new transformer-based model which leverages three data streams and mixed attention fusion mechanism to alleviate the adverse effects caused by occlusions. While our experiments demonstrate a clear decline in accuracy with missing skeleton portions, this effect is smaller withTrans4SOAR, which outperforms other architectures on all datasets. Although we specifically focus onocclusions,Trans4SOARadditionally yields state-of-the-art in thestandardSOAR without occlusion, surpassing the best published approach by 2.85% on NTU-120.
Kunyu Peng, Alina Roitberg, Kailun Yang 0001, Jiaming Zhang 0001, Rainer Stiefelhagen
IEEE Trans. Multim.5
2022 Reference-Guided Pseudo-Label Generation for Medical Semantic Segmentation
abstract
Producing densely annotated data is a difficult and tedious task for medical imaging applications. To address this problem, we propose a novel approach to generate supervision for semi-supervised semantic segmentation. We argue that visually similar regions between labeled and unlabeled images likely contain the same semantics and therefore should share their label. Following this thought, we use a small number of labeled images as reference material and match pixels in an unlabeled image to the semantic of the best fitting pixel in a reference set. This way, we avoid pitfalls such as confirmation bias, common in purely prediction-based pseudo-labeling. Since our method does not require any architectural changes or accompanying networks, one can easily insert it into existing frameworks. We achieve the same performance as a standard fully supervised model on X-ray anatomy segmentation, albeit using 95% fewer labeled images. Aside from an in-depth analysis of different aspects of our proposed method, we further demonstrate the effectiveness of our reference-guided learning paradigm by comparing our approach against existing methods for retinal fluid segmentation with competitive performance as we improve upon recent work by up to 15% mean IoU.
Constantin Seibold, Simon Reiß, Jens Kleesiek, Rainer Stiefelhagen
AAAI4
2022 MatchFormer: Interleaving Attention in Transformers for Feature Matching
Jiaming Zhang 0001, Kailun Yang 0001, Kunyu Peng, Rainer Stiefelhagen
ACCV (3)5
2022 Detailed Annotations of Chest X-Rays via CT Projection for Report Understanding
Constantin Seibold, Simon Reiß, M. Saquib Sarfraz, Matthias A. Fink, Victoria Mayer, Jan Sellner, Moon S. Kim 0002, Klaus H. Maier-Hein, Jens Kleesiek, Rainer Stiefelhagen
BMVC10
2022 Bending Reality: Distortion-aware Transformers for Adapting to Panoramic Semantic Segmentation
abstract
Panoramic images with their 360° directional view encompass exhaustive information about the surrounding space, providing a rich foundation for scene understanding. To unfold this potential in the form of robust panoramic segmentation models, large quantities of expensive, pixel-wise annotations are crucial for success. Such annotations are available, but predominantly for narrow-angle, pinhole-camera images which, off the shelf, serve as sub-optimal resources for training panoramic models. Distortions and the distinct image-feature distribution in 360° panoramas impede the transfer from the annotation-rich pinhole domain and therefore come with a big dent in performance. To get around this domain difference and bring together semantic annotations from pinhole- and 360° surround-visuals, we propose to learn object deformations and panoramic image distortions in the Deformable Patch Embedding (DPE) and Deformable MLP (DMLP) components which blend into our Transformer for PAnoramic Semantic Segmentation (Trans4PASS) model. Finally, we tie together shared semantics in pinhole- and panoramic feature embeddings by generating multi-scale prototype features and aligning them in our Mutual Prototypical Adaptation (MPA) for unsupervised domain adaptation. On the indoor Stanford2D3D dataset, our Trans4PASS with MPA maintains comparable performance to fully-supervised state-of-the-arts, cutting the need for over 1,400 labeled panoramas. On the outdoor DensePASS dataset, we break state-of-the-art by 14.39% mIoU and set the new bar at 56.38%.
Jiaming Zhang 0001, Kailun Yang 0001, Chaoxiang Ma, Simon Reiß, Kunyu Peng, Rainer Stiefelhagen
CVPR6
2022 Hierarchical Nearest Neighbor Graph Embedding for Efficient Dimensionality Reduction
abstract
Dimensionality reduction is crucial both for visualization and preprocessing high dimensional data for machine learning. We introduce a novel method based on a hierarchy built on 1-nearest neighbor graphs in the original space which is used to preserve the grouping properties of the data distribution on multiple levels. The core of the proposal is an optimization-free projection that is competitive with the latest versions of t-SNE and UMAP in performance and visualization quality while being an order of magnitude faster at run-time. Furthermore, its interpretable mechanics, the ability to project new data, and the natural separation of data clusters in visualizations make it a general purpose unsupervised dimension reduction technique. In the paper, we argue about the soundness of the proposed method and evaluate it on a diverse collection of datasets with sizes varying from 1 K to 11M samples and dimensions from 28 to 16K. We perform comparisons with other state-of-the-art methods on multiple metrics and target dimensions high-lighting its efficiency and performance. Code is available at https://github.com/koulakis/h-nne
M. Saquib Sarfraz, Marios Koulakis, Constantin Seibold, Rainer Stiefelhagen
CVPR4
2022 Graph-Constrained Contrastive Regularization for Semi-weakly Volumetric Segmentation
Simon Reiß, Constantin Seibold, Alexander Freytag, Erik Rodner, Rainer Stiefelhagen
ECCV (21)5
2022 Revisiting Click-Based Interactive Video Object Segmentation
abstract
While current methods for interactive Video Object Segmentation (iVOS) rely on scribble-based interactions to generate precise object masks, we propose a Click-based interactive Video Object Segmentation (CiVOS) framework to simplify the required user workload as much as possible. CiVOS builds on de-coupled modules reflecting user interaction and mask propagation. The interaction module converts click-based interactions into an object mask, which is then inferred to the remaining frames by the propagation module. Additional user interactions allow for a refinement of the object mask. The approach is extensively evaluated on the popular interactive DAVIS dataset, but with an inevitable adaptation of scribble-based interactions with click-based counterparts. We consider several strategies for generating clicks during our evaluation to reflect various user inputs and adjust the DAVIS performance metric to perform a hardware-independent comparison. The presented CiVOS pipeline achieves competitive results, although requiring a lower user workload.
Stéphane Vujasinovic, Sebastian Bullinger, Stefan Becker, Norbert Scherer-Negenborn, Michael Arens, Rainer Stiefelhagen
ICIP6
2022 Towards Automatic Parsing of Structured Visual Content through the Use of Synthetic Data
abstract
Structured Visual Content (SVC) such as graphs, flow charts, or the like are used by authors to illustrate various concepts. While such depictions allow the average reader to better understand the contents, images containing SVCs are typically not machine-readable. This, in turn, not only hinders automated knowledge aggregation, but also the perception of displayed information for visually impaired people. In this work, we propose a synthetic dataset, containing SVCs in the form of images as well as ground truths. We show the usage of this dataset by an application that automatically extracts a graph representation from an SVC image. This is done by training a model via common supervised learning methods. As there currently exist no large-scale public datasets for the detailed analysis of SVC, we propose the Synthetic SVC (SSVC) dataset comprising 12,000 images with respective bounding box annotations and detailed graph representations. Our dataset enables the development of strong models for the interpretation of SVCs while skipping the time-consuming dense data annotation.We evaluate our model on both synthetic and manually annotated data and show the transferability of synthetic to real via various metrics, given the presented application. Here, we evaluate that this proof of concept is possible to some extend and lay down a solid baseline for this task. We discuss the limitations of our approach for further improvements. Our utilized metrics can be used as a tool for future comparisons in this domain. To enable further research on this task, the dataset is publicly available at https://bit.ly/3jN1pJJ.
Lukas Schölch, Jonas Steinhäuser, Maximilian Beichter, Constantin Seibold, Kailun Yang 0001, Merlin Knaeble, Thorsten Schwarz, Alexander Maedche, Rainer Stiefelhagen
ICPR9
2022 Continuous Self-Localization on Aerial Images Using Visual and Lidar Sensors
abstract
This paper proposes a novel method for geo-tracking, i.e. continuous metric self-localization in outdoor environments by registering a vehicle's sensor information with aerial imagery of an unseen target region. Geo- tracking methods offer the potential to supplant noisy signals from global navigation satellite systems (GNSS) and expensive and hard to maintain prior maps that are typically used for this purpose. The proposed geo-tracking method aligns data from on-board cameras and lidar sensors with geo-registered orthophotos to continuously localize a vehicle. We train a model in a metric learning setting to extract visual features from ground and aerial images. The ground features are projected into a top-down perspective via the lidar points and are matched with the aerial features to determine the relative pose between vehicle and orthophoto. Our method is the first to utilize on-board cameras in an end-to-end differentiable model for metric self-localization on unseen orthophotos. It exhibits strong generalization, is robust to changes in the environment and requires only geo-poses as ground truth. We evaluate our approach on the KITTI-360 dataset and achieve a mean absolute position error (APE) of 0.94m. We further compare with previous approaches on the KITTI odometry dataset and achieve state-of-the-art results on the geo-tracking task.3
Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen
IROS5
2022 Multimodal Generation of Novel Action Appearances for Synthetic-to-Real Recognition of Activities of Daily Living
abstract
Domain shifts, such as appearance changes, are a key challenge in real-world applications of activity recognition models, which range from assistive robotics and smart homes to driver observation in intelligent vehicles. For example, while simulations are an excellent way of economical data collection, a Synthetic→Real domain shift leads to > 60% drop in accuracy when recognizing Activities of Daily Living (ADLs). We tackle this challenge and introduce an activity domain generation framework which creates novel ADL appearances (novel domains) from different existing activity modalities (source domains) inferred from video training data. Our frame-work computes human poses, heatmaps of body joints, and optical flow maps and uses them alongside the original RGB videos to learn the essence of source domains in order to generate completely new ADL domains. The model is optimized by maximizing the distance between the existing source appearances and the generated novel appearances while ensuring that the semantics of an activity is preserved through an additional classification loss. While source data multimodality is an important concept in this design, our setup does not rely on multi-sensor setups, (i.e., all source modalities are inferred from a single video only.) The newly created activity domains are then integrated in the training of the ADL classification networks, resulting in models far less susceptible to changes in data distributions. Extensive experiments on the Synthetic→Real benchmark Sims4Action demonstrate the potential of the domain generation paradigm for cross-domain ADL recognition, setting new state-of-the-art results. Our code is publicly available at https://github.com/Zrrr1997/syn2real_DG.
Zdravko Marinov, David Schneider 0006, Alina Roitberg, Rainer Stiefelhagen
IROS4
2022 TransDARC: Transformer-based Driver Activity Recognition with Latent Space Feature Calibration
abstract
Traditional video-based human activity recognition has experienced remarkable progress linked to the rise of deep learning, but this effect was slower as it comes to the downstream task of driver behavior understanding. Understanding the situation inside the vehicle cabin is essential for Advanced Driving Assistant System (ADAS) as it enables identifying distraction, predicting driver's intent and leads to more convenient human-vehicle interaction. At the same time, driver observation systems face substantial obstacles as they need to capture different granularities of driver states, while the complexity of such secondary activities grows with the rising automation and increased driver freedom. Furthermore, a model is rarely deployed under conditions identical to the ones in the training set, as sensor placements and types vary from vehicle to vehicle, constituting a substantial obstacle for real-life deployment of data-driven models. In this work, we present a novel vision-based framework for recognizing secondary driver behaviours based on visual transformers and an additional augmented feature distribution calibration module. This module operates in the latent feature-space enriching and diversifying the training set at feature-level in order to improve generalization to novel data appearances, (e.g., sensor changes) and general feature quality. Our framework consistently leads to better recognition rates, surpassing previous state-of-the-art results of the public Drive&Act benchmark on all granularity levels. Our code will be made publicly available at https://github.com/KPeng9510/TransDARC.
Kunyu Peng, Alina Roitberg, Kailun Yang 0001, Jiaming Zhang 0001, Rainer Stiefelhagen
IROS5
2022 A Comparative Analysis of Decision-Level Fusion for Multimodal Driver Behaviour Understanding
abstract
Visual recognition inside the vehicle cabin leads to safer driving and more intuitive human-vehicle interaction but such systems face substantial obstacles as they need to capture different granularities of driver behaviour while dealing with highly limited body visibility and changing illumination. Multimodal recognition mitigates a number of such issues: prediction outcomes of different sensors complement each other due to different modality-specific strengths and weaknesses. While several late fusion methods have been considered in previously published frameworks, they constantly feature different architecture backbones and building blocks making it very hard to isolate the role of the chosen late fusion strategy itself.This paper presents an empirical evaluation of different paradigms for decision-level late fusion in video-based driver observation. We compare seven different mechanisms for joining the results of single-modal classifiers which have been both popular, (e.g. score averaging) and not yet considered (e.g. rank-level fusion) in the context of driver observation evaluating them based on different criteria and benchmark settings. This is the first systematic study of strategies for fusing outcomes of multimodal predictors inside the vehicles, conducted with the goal to provide guidance for fusion scheme selection.
Alina Roitberg, Kunyu Peng, Zdravko Marinov, Constantin Seibold, David Schneider 0006, Rainer Stiefelhagen
IV6
2022 Breaking with Fixed Set Pathology Recognition Through Report-Guided Contrastive Training
Constantin Seibold, Simon Reiß, M. Saquib Sarfraz, Rainer Stiefelhagen, Jens Kleesiek
MICCAI (5)4
2022 Transfer Beyond the Field of View: Dense Panoramic Semantic Segmentation via Unsupervised Domain Adaptation
abstract
Autonomous vehicles clearly benefit from the expanded Field of View (FoV) of 360° sensors, but modern semantic segmentation approaches rely heavily on annotated training data which is rarely available forpanoramicimages. We look at this problem from the perspective of domain adaptation and bringpanoramicsemantic segmentation to a setting, where labelled training data originates from a different distribution of conventionalpinholecamera images. To achieve this, we formalize the task of unsupervised domain adaptation for panoramic semantic segmentation and collect DensePass - a novel densely annotated dataset for panoramic segmentation under cross-domain conditions, specifically built to study the Pinhole$\rightarrow$PANORAMIC domain shift and accompanied with pinhole camera training examples obtained from Cityscapes. DensePass covers both, labelled- and unlabelled 360° images, with the labelled data comprising 19 classes which explicitly fit the categories available in the source (i.e.pinhole) domain. Since data-driven models are especially susceptible to changes in data distribution, we introduce P2PDA - a generic framework for Pinhole$\rightarrow$Panoramic semantic segmentation which addresses the challenge of domain divergence with different variants of attention-augmented domain adaptation modules, enabling the transfer in output-, feature-, and feature confidence spaces. P2PDA intertwines uncertainty-aware adaptation using confidence values regulated on-the-fly through attention heads with discrepant predictions. Our framework facilitates context exchange when learning domain correspondences and dramatically improves the adaptation performance of accuracy- and efficiency-focused models. Comprehensive experiments verify that our framework clearly surpasses unsupervised domain adaptation- and specialized panoramic segmentation approaches as well as state-of-the-art semantic segmentation methods.
Jiaming Zhang 0001, Chaoxiang Ma, Kailun Yang 0001, Alina Roitberg, Kunyu Peng, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.6
2022 MASS: Multi-Attentional Semantic Segmentation of LiDAR Data for Dense Top-View Understanding
abstract
At the heart of all automated driving systems is the ability to sense the surroundings,e.g.,through semantic segmentation of LiDAR sequences, which experienced a remarkable progress due to the release of large datasets such as SemanticKITTI and nuScenes-LidarSeg. While most previous works focus onsparsesegmentation of the LiDAR input,denseoutput masks provide self-driving cars with almost complete environment information. In this paper, we introduce MASS - a Multi-Attentional Semantic Segmentation model specifically built for dense top-view understanding of the driving scenes. Our framework operates on pillar- and occupancy features and comprises three attention-based building blocks: (1) a keypoint-driven graph attention, (2) an LSTM-based attention computed from a vector embedding of the spatial input, and (3) a pillar-based attention, resulting in a dense 360° segmentation mask. With extensive experiments on both, SemanticKITTI and nuScenes-LidarSeg, we quantitatively demonstrate the effectiveness of our model, outperforming the state of the art by 19.0% on SemanticKITTI and reaching 30.4% in mIoU on nuScenes-LidarSeg, where MASS is the first work addressing the dense segmentation task. Furthermore, our multi-attention model is shown to be very effective for 3D object detection validated on the KITTI-3D dataset, showcasing its high generalizability to other tasks related to 3D vision.
Kunyu Peng, Juncong Fei, Kailun Yang 0001, Alina Roitberg, Jiaming Zhang 0001, Frank Bieder, Philipp Heidenreich, Christoph Stiller, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.9
2022 Is My Driver Observation Model Overconfident? Input-Guided Calibration Networks for Reliable and Interpretable Confidence Estimates
abstract
Driver observation models are rarely deployed under perfect conditions. In practice, illumination, camera placement and type differ from the ones present during training and unforeseen behaviours may occur at any time. While observing the human behind the steering wheel leads to more intuitive human-vehicle-interaction and safer driving, it requires recognition algorithms which do not only predict the correct driver state, but also determine their prediction quality through realistic and interpretable confidence measures. Reliable uncertainty estimates are crucial for building trust and are a serious obstacle for deploying activity recognition networks in real driving systems. In this work, we for the first time examine how well the confidence values of modern driver observation models indeed match the probability of the correct outcome and show that raw neural network-based approaches tend to significantly overestimate their prediction quality. To correct this misalignment between the confidence values and the actual uncertainty, we consider two strategies. First, we enhance two activity recognition models often used for driver observation with temperature scaling – an off-the-shelf method for confidence calibration in image classification. Then, we introduce Calibrated Action Recognition with Input Guidance (CARING) – a novel approach leveraging an additional neural network to learn scaling the confidences depending on the video representation. Extensive experiments on the Drive&Act dataset demonstrate that both strategies drastically improve the quality of model confidences, while our CARING model outperforms both, the original architectures and their temperature scaling enhancement, leading to best uncertainty estimates.
Alina Roitberg, Kunyu Peng, David Schneider 0006, Kailun Yang 0001, Marios Koulakis, Manuel Martínez 0001, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.7
2022 Omnisupervised Omnidirectional Semantic Segmentation
abstract
Modern efficient Convolutional Neural Networks (CNNs) are able to perform semantic segmentation both swiftly and accurately, which covers typically separate detection tasks desired by Intelligent Vehicles (IV) in a unified way. Most of the current semantic perception frameworks are designed to work with pinhole cameras and benchmarked against public datasets with narrow Field-of-View (FoV) images. However, there is a large accuracy downgrade when a pinhole-yielded CNN is taken to omnidirectional imagery, causing it unreliable for surrounding perception. In this paper, we propose an omnisupervised learning framework for efficient CNNs, which bridges multiple heterogeneous data sources that are already available in the community, bypassing the labor-intensive process to have manually annotated panoramas, while improving their reliability in unseen omnidirectional domains. Being omnisupervised, the efficient CNN exploits both labeled pinhole images and unlabeled panoramas. The framework is based on our specialized ensemble method that considers the wide-angle and wrap-around features of omnidirectional images, to automatically generate panoramic labels for data distillation. A comprehensive variety of experiments demonstrates that the proposed solution helps to attain significant generalizability gains in panoramic imagery domains. Our approach outperforms state-of-the-art efficient segmenters on highly unconstrained IDD20K and PASS datasets.
Kailun Yang 0001, Xinxin Hu, Yicheng Fang, Kaiwei Wang, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.5
2022 Trans4Trans: Efficient Transformer for Transparent Object and Semantic Scene Segmentation in Real-World Navigation Assistance
abstract
Transparent objects, such as glass walls and doors, constitute architectural obstacles hindering the mobility of people with low vision or blindness. For instance, the open space behind glass doors is inaccessible, unless it is correctly perceived and interacted with. However, traditional assistive technologies rarely cover the segmentation of these safety-critical transparent objects. In this paper, we build a wearable system with a novel dual-head Transformer for Transparency (Trans4Trans) perception model, which can segment general- and transparent objects. The two dense segmentation results are further combined with depth information in the system to help users navigate safely and assist them to negotiate transparent obstacles. We propose a lightweight Transformer Parsing Module (TPM) to perform multi-scale feature interpretation in the transformer-based decoder. Benefiting from TPM, the double decoders can perform joint learning from corresponding datasets to pursue robustness, meanwhile maintain efficiency on a portable GPU, with negligible calculation increase. The entire Trans4Trans model is constructed in a symmetrical encoder-decoder architecture, which outperforms state-of-the-art methods on the test sets of Stanford2D3D and Trans10K-v2 datasets, obtaining mIoU of 45.13% and 75.14%, respectively. Through a user study and various pre-tests conducted in indoor and outdoor scenes, the usability and reliability of our assistive system have been extensively verified. Meanwhile, the Tran4Trans model has outstanding performances on driving scene datasets. On Cityscapes, ACDC, and DADA-seg datasets corresponding to common environments, adverse weather, and traffic accident scenarios, mIoU scores of 81.5%, 76.3%, and 39.2% are obtained, demonstrating its high efficiency and robustness for real-world transportation applications.
Jiaming Zhang 0001, Kailun Yang 0001, Angela Constantinescu, Kunyu Peng, Karin Müller 0001, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.6
2022 Exploring Event-Driven Dynamic Context for Accident Scene Segmentation
abstract
The robustness of semantic segmentation on edge cases of traffic scene is a vital factor for the safety of intelligent transportation. However, most of the critical scenes of traffic accidents are extremely dynamic and previously unseen, which seriously harm the performance of semantic segmentation methods. In addition, the delay of the traditional camera during high-speed driving will further reduce the contextual information in the time dimension. Therefore, we propose to extract dynamic context from event-based data with a higher temporal resolution to enhance static RGB images, even for those from traffic accidents with motion blur, collisions, deformations, overturns,etc.Moreover, in order to evaluate the segmentation performance in traffic accidents, we provide a pixel-wise annotated accident dataset, namely DADA-seg, which contains a variety of critical scenarios from traffic accidents. Our experiments indicate that event-based data can provide complementary information to stabilize semantic segmentation under adverse conditions by preserving fine-grained motion of fast-moving foreground (crash objects) in accidents. Our approach achieves +8.2% performance gain on the proposed accident dataset, exceeding more than 20 state-of-the-art semantic segmentation methods. The proposal has been demonstrated to be consistently effective for models learned on multiple source databases including Cityscapes, KITTI-360, BDD, and ApolloScape.
Jiaming Zhang 0001, Kailun Yang 0001, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.3
2021 UBR²S: Uncertainty-Based Resampling and Reweighting Strategy for Unsupervised Domain Adaptation
Tobias Ringwald, Rainer Stiefelhagen
BMVC2
2021 Capturing Omni-Range Context for Omnidirectional Segmentation
abstract
Convolutional Networks (ConvNets) excel at semantic segmentation and have become a vital component for perception in autonomous driving. Enabling an all-encompassing view of street-scenes, omnidirectional cameras present themselves as a perfect fit in such systems. Most segmentation models for parsing urban environments operate on common, narrow Field of View (FoV) images. Transferring these models from the domain they were designed for to 360° perception, their performance drops dramatically, e.g., by an absolute 30.0% (mIoU) on established test-beds. To bridge the gap in terms of FoV and structural distribution between the imaging domains, we introduce Efficient Concurrent Attention Networks (ECANets), directly capturing the inherent long-range dependencies in omnidirectional imagery. In addition to the learned attention-based contextual priors that can stretch across 360° images, we upgrade model training by leveraging multi-source and omni-supervised learning, taking advantage of both: Densely labeled and unlabeled data originating from multiple datasets. To foster progress in panoramic image segmentation, we put forward and extensively evaluate models on Wild PAnoramic Semantic Segmentation (WildPASS), a dataset designed to capture diverse scenes from all around the globe. Our novel model, training regimen and multi-source prediction fusion elevate the performance (mIoU) to new state-of-the-art results on the public PASS (60.2%) and the fresh WildPASS (69.0%) benchmarks.1
Kailun Yang 0001, Jiaming Zhang 0001, Simon Reiß, Xinxin Hu, Rainer Stiefelhagen
CVPR5
2021 Every Annotation Counts: Multi-Label Deep Supervision for Medical Image Segmentation
abstract
Pixel-wise segmentation is one of the most data and an-notation hungry tasks in our field. Providing representative and accurate annotations is often mission-critical especially for challenging medical applications. In this paper, we propose a semi-weakly supervised segmentation algorithm to overcome this barrier. Our approach is based on a new formulation of deep supervision and student-teacher model and allows for easy integration of different supervision signals. In contrast to previous work, we show that care has to be taken how deep supervision is integrated in lower layers and we present multi-label deep supervision as the most important secret ingredient for success. With our novel training regime for segmentation that flexibly makes use of images that are either fully labeled, marked with bounding boxes, just global labels, or not at all, we are able to cut the requirement for expensive labels by 94.22% – narrowing the gap to the best fully supervised baseline to only 5% mean IoU. Our approach is validated by extensive experiments on retinal fluid segmentation and we provide an in-depth analysis of the anticipated effect each annotation type can have in boosting segmentation performance.
Simon Reiß, Constantin Seibold, Alexander Freytag, Erik Rodner, Rainer Stiefelhagen
CVPR5
2021 Temporally-Weighted Hierarchical Clustering for Unsupervised Action Segmentation
abstract
Action segmentation refers to inferring boundaries of semantically consistent visual concepts in videos and is an important requirement for many video understanding tasks. For this and other video understanding tasks, supervised approaches have achieved encouraging performance but require a high volume of detailed frame-level annotations. We present a fully automatic and unsupervised approach for segmenting actions in a video that does not require any training. Our proposal is an effective temporally-weighted hierarchical clustering algorithm that can group semantically consistent frames of the video. Our main finding is that representing a video with a 1-nearest neighbor graph by taking into account the time progression is sufficient to form semantically and temporally consistent clusters of frames where each cluster may represent some action in the video. Additionally, we establish strong unsupervised baselines for action segmentation and show significant performance improvements over published unsupervised methods on five challenging action segmentation datasets. Our code is available.1
M. Saquib Sarfraz, Naila Murray, Vivek Sharma 0001, Ali Diba, Luc Van Gool, Rainer Stiefelhagen
CVPR6
2021 Affect-DML: Context-Aware One-Shot Recognition of Human Affect using Deep Metric Learning
abstract
Human affect recognition is a well-established research area with numerous applications, e.g. in psychological care, but existing methods assume that all emotions-of-interest are given a priori as annotated training examples. However, the rising granularity and refinements of the human emotional spectrum through novel psychological theories and the increased consideration of emotions in context brings considerable pressure to data collection and labeling work. In this paper, we conceptualize one-shot recognition of emotions in context - a new problem aimed at recognizing human affect states in finer particle level from a single support sample. To address this challenging task, we follow the deep metric learning paradigm and introduce a multi-modal emotion embedding approach which minimizes the distance of the same-emotion embeddings by leveraging complementary information of human appearance and the semantic scene context obtained through a semantic segmentation network. All streams of our context-aware model are optimized jointly using weighted triplet loss and weighted cross entropy loss. We conduct thorough experiments on both, categorical and numerical emotion recognition tasks of the Emotic dataset adapted to our one-shot recognition problem, revealing that categorizing human affect from a single example is a hard task. Still, all variants of our model clearly outperform the random baseline, while leveraging the semantic scene context consistently improves the learnt representations, setting state-of-the-art results in one-shot emotion recognition. To foster research of more universal representations of human affect states, we will make our benchmark and models publicly available to the community under https://github.com/KPeng9510/Affect-DML.
Kunyu Peng, Alina Roitberg, David Schneider 0006, Marios Koulakis, Kailun Yang 0001, Rainer Stiefelhagen
FG6
2021 Vi2CLR: Video and Image for Visual Contrastive Learning of Representation
abstract
In this paper, we introduce a novel self-supervised visual representation learning method which understands both images and videos in a joint learning fashion. The proposed neural network architecture and objectives are designed to obtain two different Convolutional Neural Networks for solving visual recognition tasks in the domain of videos and images. Our method called Video/Image for Visual Contrastive Learning of Representation(Vi2CLR) uses unlabeled videos to exploit dynamic and static visual cues for self-supervised and instances similarity/dissimilarity learning. Vi2CLR optimization pipeline consists of visual clustering part and representation learning based on groups of similar positive instances within a cluster and negative ones from other clusters and learning visual clusters and their distances. We show how a joint self-supervised visual clustering and instance similarity learning with 2D (image) and 3D (video) CovNet encoders yields such robust and near to supervised learning performance.We extensively evaluate the method on downstream tasks like large scale action recognition, image and object classification on datasets like Kinetics, ImageNet, Pascal VOC’07 and UCF101 and achieve outstanding results compared to state-of-the-art self-supervised methods.
Ali Diba, Vivek Sharma 0001, Reza Safdari, Dariush Lotfi, M. Saquib Sarfraz, Rainer Stiefelhagen, Luc Van Gool
ICCV6
2021 Let's Play for Action: Recognizing Activities of Daily Living by Learning from Life Simulation Video Games
abstract
Recognizing Activities of Daily Living (ADL) is a vital process for intelligent assistive robots, but collecting large annotated datasets requires time-consuming temporal labeling and raises privacy concerns, e.g., if the data is collected in a real household. In this work, we explore the concept of constructing training examples for ADL recognition by playing life simulation video games and introduce the SIMS4ACTION dataset created with the popular commercial game THE SIMS 4. We build SIMS4ACTION by specifically executing actions-of-interest in a "top-down" manner, while the gaming circumstances allow us to freely switch between environments, camera angles and subject appearances. While ADL recognition on gaming data is interesting from the theoretical perspective, the key challenge arises from transferring it to the real-world applications, such as smart-homes or assistive robotics. To meet this requirement, SIMS4ACTION is accompanied with a GAMING→REAL benchmark, where the models are evaluated on real videos derived from an existing ADL dataset. We integrate two modern algorithms for video-based activity recognition in our framework, revealing the value of life simulation video games as an inexpensive and far less intrusive source of training data. However, our results also indicate that tasks involving a mixture of gaming and real data are challenging, opening a new research direction. We will make our dataset publicly available at https://github.com/aroitberg/sims4action.
Alina Roitberg, David Schneider 0006, Aulia Djamal, Constantin Seibold, Simon Reiß, Rainer Stiefelhagen
IROS6
2021 ISSAFE: Improving Semantic Segmentation in Accidents by Fusing Event-based Data
abstract
Ensuring the safety of all traffic participants is a prerequisite for bringing intelligent vehicles closer to practical applications. The assistance system should not only achieve high accuracy under normal conditions, but obtain robust perception against extreme situations. However, traffic accidents that involve object collisions, deformations, overturns, etc., yet unseen in most training sets, will largely harm the performance of existing semantic segmentation models. To tackle this issue, we present a rarely addressed task regarding semantic segmentation in accidental scenarios, along with an accident dataset DADA-seg. It contains 313 various accident sequences with 40 frames each, of which the time windows are located before and during a traffic accident. Every 11th frame is manually annotated for benchmarking the segmentation performance. Furthermore, we propose a novel event-based multi-modal segmentation architecture ISSAFE. Our experiments indicate that event-based data can provide complementary information to stabilize semantic segmentation under adverse conditions by preserving fine-grain motion of fast-moving foreground (crash objects) in accidents. Our approach achieves +8.2% mIoU performance gain on the proposed evaluation set, exceeding more than 10 state-of-the-art segmentation methods. The proposed ISSAFE architecture is demonstrated to be consistently effective for models learned on multiple source databases including Cityscapes, KITTI-360, BDD and ApolloScape.
Jiaming Zhang 0001, Kailun Yang 0001, Rainer Stiefelhagen
IROS3
2021 DR-TANet: Dynamic Receptive Temporal Attention Network for Street Scene Change Detection
abstract
Street scene change detection continues to capture researchers' interests in the computer vision community. It aims to identify the changed regions of the paired street-view images captured at different times. The state-of-the-art network based on the encoder-decoder architecture leverages the feature maps at the corresponding level between two channels to gain sufficient information of changes. Still, the efficiency of feature extraction, feature correlation calculation, even the whole network requires further improvement. This paper proposes the temporal attention and explores the impact of the dependency-scope size of temporal attention on the performance of change detection. In addition, based on the Temporal Attention Module (TAM), we introduce a more efficient and light-weight version - Dynamic Receptive Temporal Attention Module (DRTAM) and propose the Concurrent Horizontal and Vertical Attention (CHVA) to improve the accuracy of the network on specific challenging entities. On street scene datasets ‘GSV’, ‘TSUNAMI’ and ‘VL-CMU-CD’, our approach gains excellent performance, establishing new state-of-the-art scores without bells and whistles, while maintaining high efficiency applicable in autonomous vehicles.
Kailun Yang 0001, Rainer Stiefelhagen
IV3
2021 Panoramic Panoptic Segmentation: Towards Complete Surrounding Understanding via Unsupervised Contrastive Learning
abstract
In this work, we introduce panoramic panoptic segmentation as the most holistic scene understanding both in terms of field of view and image level understanding for standard camera based input. A complete surrounding understanding provides a maximum of information to the agent, which is essential for any intelligent vehicle in order to make informed decisions in a safety-critical dynamic environment such as real-world traffic. In order to overcome the lack of annotated panoramic images, we propose a framework which allows model training on standard pinhole images and transfers the learned features to a different domain. Using our proposed method, we manage to achieve significant improvements of over 5% measured in PQ over non-adapted models on our Wild Panoramic Panoptic Segmentation (WildPPS) dataset. We show that our proposed Panoramic Robust Feature (PRF) framework is not only suitable to improve performance on panoramic images but can be beneficial whenever model training and deployment are executed on data taken from different distributions. As an additional contribution, we publish WildPPS: The first panoramic panoptic image dataset to foster progress in surrounding perception.
Alexander Jaus, Kailun Yang 0001, Rainer Stiefelhagen
IV3
2021 From Driver Talk To Future Action: Vehicle Maneuver Prediction by Learning from Driving Exam Dialogs
abstract
A rapidly growing amount of content posted online inherently holds knowledge about concepts of interest, i.e. driver actions. We leverage methods at the intersection of vision and language to surpass costly annotation and present the first automated framework for anticipating driver intention by learning from recorded driving exam conversations. We query YouTube and collect a dataset of posted mock road tests comprising student-teacher dialogs and video data, which we use for learning to foresee the next maneuver without any additional supervision. However, instructional conversations give us very loose labels, while casual chat results in a high amount of noise. To mitigate this effect, we propose a technique for automatic detection of smalltalk based on the likelihood of spoken words being present in everyday dialogs. While visually recognizing driver's intention by learning from natural dialogs only is a challenging task, learning from less but better data via our smalltalk refinement consistently improves performance.
Alina Roitberg, Simon Reiß, Rainer Stiefelhagen
IV3
2021 Unsupervised Meta-Domain Adaptation for Fashion Retrieval
abstract
Cross-domain fashion item retrieval naturally arises when unconstrained consumer images are used to query for fashion items in a collection of high-quality photographs provided by retailers. To perform this task, approaches typically leverage both consumer and shop domains from a given dataset to learn a domain invariant representation, allowing these images of different nature to be directly compared. When consumer images are not available beforehand, such training is impossible. In this paper, we focus on this challenging and yet practical scenario, and we propose instead to leverage representations learned for cross-domain retrieval from another source dataset and to adapt them to the target dataset for this particular setting. More precisely, we bypass the lack of consumer images and directly target the more challenging meta-domain gap which occurs between consumer images and shop images, independently of their dataset. Assuming that datasets share some similar fashion items, we cluster their shop images and leverage the clusters to automatically generate pseudo-labels. Those are used to associate consumer and shop images across datasets, which in turn allows to learn meta-domain-invariant representations suitable for cross-domain retrieval in the target dataset. The features and code are available at https://github.com/vivoutlaw/UDMA.
Vivek Sharma 0001, Naila Murray, Diane Larlus, M. Saquib Sarfraz, Rainer Stiefelhagen, Gabriela Csurka
WACV5
2021 Adaptiope: A Modern Benchmark for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) deals with the adaptation process of a given source domain with labeled training data to a target domain for which only unannotated data is available. This is a challenging task as the domain shift leads to degraded performance on the target domain data if not addressed. In this paper, we analyze commonly used UDA classification datasets and discover systematic problems with regard to dataset setup, ground truth ambiguity and annotation quality. We manually clean the most popular UDA dataset in the research area (Office-31) and quantify the negative effects of inaccurate annotations through thorough experiments. Based on these insights, we collect the Adaptiope dataset - a large scale, diverse UDA dataset with synthetic, product and real world data - and show that its transfer tasks provide a challenge even when considering recent UDA algorithms. Our datasets are available at https://gitlab.com/tringwald/adaptiope.
Tobias Ringwald, Rainer Stiefelhagen
WACV2
2021 Is Context-Aware CNN Ready for the Surroundings? Panoramic Semantic Segmentation in the Wild
abstract
Semantic segmentation, unifying most navigational perception tasks at the pixel level has catalyzed striking progress in the field of autonomous transportation. Modern Convolution Neural Networks (CNNs) are able to perform semantic segmentation both efficiently and accurately, particularly owing to their exploitation of wide context information. However, most segmentation CNNs are benchmarked against pinhole images with limited Field of View (FoV). Despite the growing popularity of panoramic cameras to sense the surroundings, semantic segmenters have not been comprehensively evaluated on omnidirectional wide-FoV data, which features rich and distinct contextual information. In this paper, we propose a concurrent horizontal and vertical attention module to leverage width-wise and height-wise contextual priors markedly available in the panoramas. To yield semantic segmenters suitable for wide-FoV images, we present a multi-source omni-supervised learning scheme with panoramic domain covered in the training via data distillation. To facilitate the evaluation of contemporary CNNs in panoramic imagery, we put forward the Wild PAnoramic Semantic Segmentation (WildPASS) dataset, comprising images from all around the globe, as well as adverse and unconstrained scenes, which further reflects perception challenges of navigation applications in the real world. A comprehensive variety of experiments demonstrates that the proposed methods enable our high-efficiency architecture to attain significant accuracy gains, outperforming the state of the art in panoramic imagery domains.
Kailun Yang 0001, Xinxin Hu, Rainer Stiefelhagen
IEEE Trans. Image Process.3
2020 Self-guided Multiple Instance Learning for Weakly Supervised Disease Classification and Localization in Chest Radiographs
Constantin Seibold, Jens Kleesiek, Heinz-Peter Schlemmer, Rainer Stiefelhagen
ACCV (5)4
2020 Travelling more independently: A Requirements Analysis for Accessible Journeys to Unknown Buildings for People with Visual Impairments
abstract
It is much more difficult for people with visual impairments to plan and implement a journey to unknown places than for sighted people, because in addition to the usual travel arrangements, they also need to know whether the different parts of the travel chain are accessible at all. The need for information is presumably therefore very high and ranges from knowledge about the accessibility of public transport as well as outdoor and indoor environments. However, to the best of our knowledge, there is no study that examines in-depth requirements of both the planning of a trip and its implementation, looking separately at the various special needs of people with low vision and blindness. In this paper, we present a survey with 106 people with visual impairments, in which we examine the strategies they use to prepare for a journey to unknown buildings, how they orient themselves in unfamiliar buildings and what materials they use. Our analysis shows that requirements for people with blindness and low vision differ. The feedback from the participants reveals that there is a large information gap, especially for orientation in buildings, regarding maps, accessibility of buildings and supporting systems. In particular, there is a lack of availability of indoor maps.
Christin Engel, Karin Müller 0001, Angela Constantinescu, Claudia Loitsch, Vanessa Petrausch, Gerhard Weber 0002, Rainer Stiefelhagen
ASSETS7
2020 CLEVR: A Customizable Interactive Learning Environment for Users with Low Vision in Virtual Reality
abstract
Technological advances enable the development of new low vision aids. One such new technology is virtual reality (VR). Existing VR applications offer a variety of adjustments to the content and the user's environment to support people with low vision. Yet, an interaction concept to conveniently activate and control these aids is missing. Therefore, we designed and implemented an interaction concept based on user and expert feedback and evaluated it in a user study. Our application offers various aids to support residual vision and a radial menu for intuitive use of these aids. The results of the user study show that the VR application allows simple task solving comparable to a desktop computer. Furthermore, the user study gives an insight into the effects of different visual impairments on the usage of VR, since the field of view has a larger impact than visual acuity. Overall, our results indicate that VR is a suitable aid to support the individual needs of users with low vision.
Adrian Heinrich Hoppe, Julia K. Anken, Thorsten Schwarz, Rainer Stiefelhagen, Florian van de Camp
ASSETS4
2020 Unsupervised Domain Adaptation by Uncertain Feature Alignment
Tobias Ringwald, Rainer Stiefelhagen
BMVC2
2020 Anchor-free Small-scale Multispectral Pedestrian Detection
Alexander Wolpert, Michael Teutsch, M. Saquib Sarfraz, Rainer Stiefelhagen
BMVC4
2020 Understanding what you feel: A Mobile Audio-Tactile System for Graphics Used at Schools with Students with Visual Impairment
abstract
A lot of information is nowadays presented graphically. However, students with blindness do not have access to visual information. Providing an alternative text is not always the appropriate solution as exploring graphics to discover information independently is a fundamental part of the learning process. In this work, we introduce a mobile audio-tactile learning environment, which facilitates the incorporation of real educational material. We evaluate our system by comparing three methods of interaction with tactile graphics: A tactile graphic augmented by (1) a document with key index information in Braille, (2) a digital document with key index information and (3) the TPad system, an audio-tactile solution meeting the specific needs within the school context. Our study shows that the TPad system is suitable for educational environments. Moreover, compared to the other methods TPad is faster to explore tactile graphics and it suggests a promising effect on the memorization of information.
Giuseppe Melfi, Karin Müller 0001, Thorsten Schwarz, Gerhard Jaworek, Rainer Stiefelhagen
CHI5
2020 Large Scale Holistic Video Understanding
Ali Diba, Mohsen Fayyaz, Vivek Sharma 0001, Manohar Paluri, Juergen Gall, Rainer Stiefelhagen, Luc Van Gool
ECCV (5)6
2020 Clustering based Contrastive Learning for Improving Face Representations
abstract
A good clustering algorithm can discover natural groupings in data. These groupings, if used wisely, provide a form of weak supervision for learning representations. In this work, we present Clustering-based Contrastive Learning (CCL), a new clustering-based representation learning approach that uses labels obtained from clustering along with video constraints to learn discriminative face features. We demonstrate our method on the challenging task of learning representations for video face clustering. Through several ablation studies, we analyze the impact of creating pair-wise positive and negative labels from different sources. Experiments on three challenging video face clustering datasets: BBT-0101, BF-0502, and ACCIO show that CCL achieves a new state-of-the-art on all datasets.
Vivek Sharma 0001, Makarand Tapaswi, M. Saquib Sarfraz, Rainer Stiefelhagen
FG4
2020 Image-Based Recognition of Braille Using Neural Networks on Mobile Devices
Christopher Baumgärtner, Thorsten Schwarz, Rainer Stiefelhagen
ICCHP (1)3
2020 Can We Unify Perception and Localization in Assisted Navigation? An Indoor Semantic Visual Positioning System for Visually Impaired People
Haoye Chen, Kailun Yang 0001, Manuel Martínez 0001, Karin Müller 0001, Rainer Stiefelhagen
ICCHP (1)6
2020 AccessibleMaps: Addressing Gaps in Maps for People with Visual and Mobility Impairments
Claudia Loitsch, Karin Müller 0001, Christin Engel, Gerhard Weber 0002, Rainer Stiefelhagen
ICCHP (2)5
2020 Developing a Magnification Prototype Based on Head and Eye-Tracking for Persons with Low Vision
Thorsten Schwarz, Arsalan Akbarioroumieh, Giuseppe Melfi, Rainer Stiefelhagen
ICCHP (1)4
2020 Bring the Environment to Life: A Sonification Module for People with Visual Impairments to Improve Situation Awareness
abstract
Digital navigation tools for helping people with visual impairments have become increasingly popular in recent years. While conventional navigation solutions give routing instructions to the user, systems such as GoogleMaps, BlindSquare, or Soundscape offer additional information about the surroundings and, thereby, improve the orientation of people with visual impairments. However, these systems only provide information about static environments, while dynamic scenes comprising objects such as bikes, dogs, and persons are not considered. In addition, both the routing and the information about the environment are usually conveyed by speech. We address this gap and implement a mobile system that combines object identification with a sonification interface. Our system can be used in three different scenarios of macro and micro navigation: orientation, obstacle avoidance, and exploration of known and unknown routes. Our proposed system leverages popular computer vision methods to localize 18 static and dynamic object classes in real-time. At the heart of our system is a mixed reality sonification interface which is adaptable to the user's needs and is able to transmit the recognized semantic information to the user. The system is designed in a user-centered approach. An exploratory user study conducted by us showed that our object-to-sound mapping with auditory icons is intuitive. On average, users perceived our system as useful and indicated that they want to know more about their environment, apart from wayfinding and points of interest.
Angela Constantinescu, Karin Müller 0001, Monica-Laura Haurilet, Vanessa Petrausch, Rainer Stiefelhagen
ICMI5
2020 Detective: An Attentive Recurrent Model for Sparse Object Detection
abstract
In this work, we present Detective - an attentive object detector that identifies objects in images in a sequential manner. Our network is based on an encoder-decoder architecture, where the encoder is a convolutional neural network, and the decoder is a convolutional recurrent neural network coupled with an attention mechanism. At each iteration, our decoder focuses on the relevant parts of the image using an attention mechanism, and then estimates the object's class and the bounding box coordinates. Current object detection models generate dense predictions and rely on post-processing to remove duplicate predictions. Detective is a sparse object detector that generates a single bounding box per object instance. However, training a sparse object detector is challenging, as it requires the model to reason at the instance level and not just at the class and spatial levels. We propose a training mechanism based on the Hungarian algorithm and a loss that balances the localization and classification tasks. This allows Detective to achieve promising results on the PASCAL VOC dataset for object detection. Our experiments demonstrate that sparse object detection is possible and has great potential for future developments in applications where the order of the desired objects is of interest.
Amine Kechaou, Manuel Martínez 0001, Monica-Laura Haurilet, Rainer Stiefelhagen
ICPR4
2020 Uncertainty-sensitive Activity Recognition: A Reliability Benchmark and the CARING Models
abstract
Beyond assigning the correct class, an activity recognition model should also be able to determine, how certain it is in its predictions. We present the first study of how well the confidence values of modern action recognition architectures indeed reflect the probability of the correct outcome and propose a learning-based approach for improving it. First, we extend two popular action recognition datasets with a reliability benchmark in form of the expected calibration error and reliability diagrams. Since our evaluation highlights that confidence values of standard action recognition architectures do not represent the uncertainty well, we introduce a new approach which learns to transform the model output into realistic confidence estimates through an additional calibration network. The main idea of our Calibrated Action Recognition with Input Guidance (CARING) model is to learn an optimal scaling parameter depending on the video representation. We compare our model with the native action recognition networks and the temperature scaling approach - a wide spread calibration method utilized in image classification. While temperature scaling alone drastically improves the reliability of the confidence values, our CARING method consistently leads to the best uncertainty estimates in all benchmark settings.
Alina Roitberg, Monica-Laura Haurilet, Manuel Martínez 0001, Rainer Stiefelhagen
ICPR4
2020 Multi-Task Learning for Calorie Prediction on a Novel Large-Scale Recipe Dataset Enriched with Nutritional Information
abstract
A rapidly growing amount of content posted online, such as food recipes, opens doors to new exciting applications at the intersection of vision and language. In this work, we aim to estimate the calorie amount of a meal directly from an image by learning from recipes people have published on the Internet, thus skipping time-consuming manual data annotation. Since there are few large-scale publicly available datasets captured in unconstrained environments, we propose the pic2kcal benchmark comprising 308 000 images from over 70 000 recipes including photographs, ingredients, and instructions. To obtain nutritional information of the ingredients and automatically determine the ground-truth calorie value, we match the items in the recipes with structured information from a food item database. We evaluate various neural networks for regression of the calorie quantity and extend them with the multi-task paradigm. Our learning procedure combines the calorie estimation with prediction of proteins, carbohydrates, and fat amounts as well as a multi-label ingredient classification. Our experiments demonstrate clear benefits of multi-task learning for calorie estimation, surpassing the single-task calorie regression by 9.9%. To encourage further research on this task, we make the code for generating the dataset and the models publicly available.
Robin Ruede, Verena Heusser, Lukas Frank, Alina Roitberg, Monica-Laura Haurilet, Rainer Stiefelhagen
ICPR6
2020 DS-PASS: Detail-Sensitive Panoramic Annular Semantic Segmentation through SwaftNet for Surrounding Sensing
abstract
Semantically interpreting the traffic scene is crucial for autonomous transportation and robotics systems. However, state-of-the-art semantic segmentation pipelines are dominantly designed to work with pinhole cameras and train with narrow Field-of-View (FoV) images. In this sense, the perception capacity is severely limited to offer higher-level confidence for upstream navigation tasks. In this paper, we propose a network adaptation framework to achieve Panoramic Annular Semantic Segmentation (PASS), which allows to re-use conventional pinhole-view image datasets, enabling modern segmentation networks to comfortably adapt to panoramic images. Specifically, we adapt our proposed SwaftNet to enhance the sensitivity to details by implementing attention-based lateral connections between the detail-critical encoder layers and the context-critical decoder layers. We benchmark the performance of efficient segmenters on panoramic segmentation with our extended PASS dataset, demonstrating that the proposed real-time SwaftNet outperforms state-of-the-art efficient networks. Furthermore, we assess real-world performance when deploying the Detail-Sensitive PASS (DS-PASS) system on a mobile robot and an instrumented vehicle, as well as the benefit of panoramic semantics for visual odometry, showing the robustness and potential to support diverse navigational applications.
Kailun Yang 0001, Xinxin Hu, Kaite Xiang, Kaiwei Wang, Rainer Stiefelhagen
IV6
2020 In Defense of Multi-Source Omni-Supervised Efficient ConvNet for Robust Semantic Segmentation in Heterogeneous Unseen Domains
abstract
Semantic segmentation renders a unified way of surrounding perception, where most of driving scene detection tasks can be covered by running a single efficient ConvNet through a forward pass. However, current frameworks posit the closed-world paradigm expressed as a single source of distribution over a predetermined set of visual classes, forgetting that a deep model must be deployed in the wild facing unseen domains and unforeseen hazards. In spite of being accurate in its comfort zone, the segmentation model may not generalize well to a new domain. In addition, a model trained with single dataset is heavily limited in terms of recognizable classes. In this paper, we propose an omni-supervised learning framework for semantic segmentation which is able to leverage heterogeneous data sources. Our omni-supervised training framework incorporates all available labeled and unlabeled data, meanwhile bridges multiple training sets to be capable of recognizing more classes that are needed for autonomous navigation application at hand in the new domain. A comprehensive variety of experiments shows that with the proposed multi-source omni-supervised learning solution, an efficient ConvNet like our ERF-PSPNet attains significant robustness gains in open domains that are of critical relevance to real deployment of vision algorithms. Our approach surpasses the state of the art on the highly unconstrained PASS and IDD20K datasets.
Kailun Yang 0001, Xinxin Hu, Kaiwei Wang, Rainer Stiefelhagen
IV4
2020 Deep Classification-driven Domain Adaptation for Cross-Modal Driver Behavior Recognition
abstract
We encounter a wide range of obstacles when integrating computer vision algorithms into applications inside the vehicle cabin, e.g. variations in illumination, sensor-type and -placement. Thus, designing domain-invariant representations is crucial for employing such models in practice. Still, the vast majority of driver activity recognition algorithms are developed under the assumption of a static domain, i.e. an identical distribution of training- and test data. In this work, we aim to bring driver monitoring to a setting, where domain shifts can occur at any time and explore generative models which learn a shared representation space of the source and target domain. First, we formulate the problem of unsupervised domain adaptation for driver activity recognition, where a model trained on labeled examples from the source domain (i.e. color images) is intended to adjust to a different target domain (i.e. infrared images) where only unlabeled data is available during training. To address this problem, we leverage current progress in image-to-image translation and adopt multiple strategies for learning a joint latent space of the source and target distribution and a mapping function to the domain of interest. As our long-term goal is a robust cross-domain classification, we enhance a Variational Auto-Encoder (VAE) for image translation with a classification-driven optimization strategy. Our model for classification-driven domain transfer leads to the best cross-domain recognition results and outperforms a conventional classification approach in color-to-infrared recognition by 13.75%.
Simon Reiß, Alina Roitberg, Monica-Laura Haurilet, Rainer Stiefelhagen
IV4
2020 Open Set Driver Activity Recognition
abstract
A common obstacle for applying computer vision models inside the vehicle cabin is the dynamic nature of the surrounding environment, as unforeseen situations may occur at any time. Driver monitoring has been widely researched in the context of closed set recognition i.e. under the premise that all categories are known a priori. Such restrictions represent a significant bottleneck in real-life, as the driver observation models are intended to handle the uncertainty of an open world. In this work, we aim to introduce the concept of open sets to the area of driver observation, where methods have been evaluated only on a static set of classes in the past. First, we formulate the problem of open set recognition for driver monitoring, where a model is intended to identify behaviors previously unseen by the classifier and present a novel Open-Drive&Act benchmark. We combine current closed set models with multiple strategies for novelty detection adopted from general action classification [1] in a generic open set driver behavior recognition framework. In addition to conventional approaches, we employ the prominent I3D architecture extended with modules for assessing its uncertainty via Monte-Carlo dropout. Our experiments demonstrate clear benefits of uncertainty-sensitive models, while leveraging the uncertainty of all the output neurons in a voting-like fashion leads to the best recognition results. To create an avenue for future work, we make Open-Drive&Act public at www.github.com/aroitberg/open-set-driver-activity-recognition.
Alina Roitberg, Chaoxiang Ma, Monica-Laura Haurilet, Rainer Stiefelhagen
IV4
2020 ShiSha: Enabling Shared Perspective With Face-to-Face Collaboration Using Redirected Avatars in Virtual Reality
abstract
The importance of remote collaboration grows in an interconnected world as the reasons to avoid travel increase. The spatial rendering and collaboration capabilities of virtual and augmented reality systems are well suited for tasks such as support or training. Users can take a shared perspective to build a common understanding. Also, users may engage in face-to-face cooperation to support interpersonal communication. However, a shared perspective and face-to-face collaboration are both desirable but naturally exclude each other. We place all users at the same location to provide a shared perspective. To avoid overlapping body parts, the avatars of the other connected users are shifted to the side. A redirected body pose modification corrects the resulting inconsistencies. The implemented system is compared to a baseline of two users standing in the same location and working with overlapping avatars. The results of a user study show that the proposed modifications provide an easy to use, efficient collaboration and yield higher co-presence and the feeling of teamwork. Applying redirection techniques to other users opens up novel ways to increase social presence for local or remote collaboration.
Adrian Heinrich Hoppe, Florian van de Camp, Rainer Stiefelhagen
Proc. ACM Hum. Comput. Interact.3
2019 Towards a Standardized Grammar for Navigation Systems for Persons with Visual Impairments
abstract
Pedestrian navigation systems are rarely accessible or suit the needs of persons with visual impairments. They usually lack a standardized grammar for their speech instructions, forcing users to learn new types of instructions for each new system. Thus, we propose (1) a German grammar with syntax rules and vocabulary for mobile pedestrian navigation systems that take into account the special requirements of people with visual impairments. (2) a set of rules for specifying what should be spoken and when, given GPS accuracy in a city [18]. We describe (3) the methodology used to obtain the grammar as well as (4) a qualitative evaluation with orientation and mobility experts and with people with visual impairments who deployed our grammar during a user study. Our approach is the first of its kind, as there is no such grammar neither for German nor for English, as far as we know. It serves as a contribution to standardize pedestrian navigation speech instructions for people with visual impairments.
Angela Constantinescu, Vanessa Petrausch, Karin Müller 0001, Rainer Stiefelhagen
ASSETS4
2019 Content and Colour Distillation for Learning Image Translations with the Spatial Profile Loss
M. Saquib Sarfraz, Constantin Seibold, Haroon Khalid, Rainer Stiefelhagen
BMVC4
2019 It's Not About the Journey; It's About the Destination: Following Soft Paths Under Question-Guidance for Visual Reasoning
abstract
Visual Reasoning remains a challenging task, as it has to deal with long-range and multi-step object relationships in the scene. We present a new model for Visual Reasoning, aimed at capturing the interplay among individual objects in the image represented as a scene graph. As not all graph components are relevant for the query, we introduce the concept of a question-based visual guide, which constrains the potential solution space by learning an optimal traversal scheme, where the final destination nodes alone are used to produce the answer. We show, that finding relevant semantic structures facilitates generalization to new tasks by introducing a novel problem of knowledge transfer: training on one question type and answering questions from a different domain without any training data. Furthermore, we report state-of-the-art results for Visual Reasoning on multiple query types and diverse image and video datasets.
Monica-Laura Haurilet, Alina Roitberg, Rainer Stiefelhagen
CVPR3
2019 Efficient Parameter-Free Clustering Using First Neighbor Relations
abstract
We present a new clustering method in the form of a single clustering equation that is able to directly discover groupings in the data. The main proposition is that the first neighbor of each sample is all one needs to discover large chains and finding the groups in the data. In contrast to most existing clustering algorithms our method does not require any hyper-parameters, distance thresholds and/or the need to specify the number of clusters. The proposed algorithm belongs to the family of hierarchical agglomerative methods. The technique has a very low computational overhead, is easily scalable and applicable to large practical problems. Evaluation on well known datasets from different domains ranging between 1077 and 8.1 million samples shows substantial performance gains when compared to the existing clustering techniques.
M. Saquib Sarfraz, Vivek Sharma 0001, Rainer Stiefelhagen
CVPR3
2019 Self-Supervised Learning of Face Representations for Video Face Clustering
abstract
Analyzing the story behind TV series and movies often requires understanding who the characters are and what they are doing. With improving deep face models, this may seem like a solved problem. However, as face detectors get better, clustering/identification needs to be revisited to address increasing diversity in facial appearance. In this paper, we address video face clustering using unsupervised methods. Our emphasis is on distilling the essential information, identity, from the representations obtained using deep pre-trained face networks. We propose a self-supervised Siamese network that can be trained without the need for video/track based supervision, and thus can also be applied to image collections. We evaluate our proposed method on three video face clustering datasets. The experiments show that our methods outperform current state-of-the-art methods on all datasets. Video face clustering is lacking a common benchmark as current works are often evaluated with different metrics and/or different sets of face tracks. The datasets and code are available at https://github.com/vivoutlaw/SSIAM.
Vivek Sharma 0001, Makarand Tapaswi, M. Saquib Sarfraz, Rainer Stiefelhagen
FG4
2019 DynamoNet: Dynamic Action and Motion Network
abstract
In this paper, we are interested in self-supervised learning the motion cues in videos using dynamic motion filters for a better motion representation to finally boost human action recognition in particular. Thus far, the vision community has focused on spatio-temporal approaches using standard filters, rather we here propose dynamic filters that adaptively learn the video-specific internal motion representation by predicting the short-term future frames. We name this new motion representation, as dynamic motion representation (DMR) and is embedded inside of 3D convolutional network as a new layer, which captures the visual appearance and motion dynamics throughout entire video clip via end-to-end network learning. Simultaneously, we utilize these motion representation to enrich video classification. We have designed the frame prediction task as an auxiliary task to empower the classification problem. With these overall objectives, to this end, we introduce a novel unified spatio-temporal 3D-CNN architecture (DynamoNet) that jointly optimizes the video classification and learning motion representation by predicting future frames as a multi-task learning problem. We conduct experiments on challenging human action datasets: Kinetics 400, UCF101, HMDB51. The experiments using the proposed DynamoNet show promising results on all the datasets.
Ali Diba, Vivek Sharma 0001, Luc Van Gool, Rainer Stiefelhagen
ICCV4
2019 Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous Vehicles
abstract
We introduce the novel domain-specific Drive&Act benchmark for fine-grained categorization of driver behavior. Our dataset features twelve hours and over 9.6 million frames of people engaged in distractive activities during both, manual and automated driving. We capture color, infrared, depth and 3D body pose information from six views and densely label the videos with a hierarchical annotation scheme, resulting in 83 categories. The key challenges of our dataset are: (1) recognition of fine-grained behavior inside the vehicle cabin; (2) multi-modal activity recognition, focusing on diverse data streams; and (3) a cross view recognition benchmark, where a model handles data from an unfamiliar domain, as sensor type and placement in the cabin can change between vehicles. Finally, we provide challenging benchmarks by adopting prominent methods for video- and body pose-based action recognition.
Manuel Martin, Alina Roitberg, Monica-Laura Haurilet, Matthias Horne, Simon Reiß, Michael Voit, Rainer Stiefelhagen
ICCV7
2019 WiSe - Slide Segmentation in the Wild
abstract
We address the task of segmenting presentation slides, where the examined page was captured as a live photo during lectures. Slides are important document types used as visual components accompanying presentations in a variety of fields ranging from education to business. However, automatic analysis of presentation slides has not been researched sufficiently, and, so far, only preprocessed images of already digitalized slide documents were considered. We aim to introduce the task of analyzing unconstrained photos of slides taken during lectures and present a novel dataset for Page Segmentation with slides captured in the Wild (WiSe). Our dataset covers pixel-wise annotations of 25 classes on 1300 pages, allowing overlapping regions (i.e., multi-class assignments). To evaluate the performance, we define multiple benchmark metrics and baseline methods for our dataset. We further implement two different deep neural network approaches previously used for segmenting natural images and adopt them for the task. Our evaluation results demonstrate the effectiveness of the deep learning-based methods, surpassing the baseline methods by over 30%. To foster further research of slide analysis in unconstrained photos, we make the WiSe dataset publicly available to the community.
Monica-Laura Haurilet, Alina Roitberg, Manuel Martínez 0001, Rainer Stiefelhagen
ICDAR4
2019 3D Object Trajectory Reconstruction using Instance-Aware Multibody Structure from Motion and Stereo Sequence Constraints
abstract
Three-dimensional environment perception is a key element of autonomous driving and driver assistance systems. A common image based approach to determine three-dimensional scene information is stereo matching, which is limited by the stereo camera baseline. In contrast to stereo matching based methods, we present an approach to reconstruct three-dimensional object trajectories combining temporal adjacent views for object point triangulation. We track two-dimensional object shapes on pixel level exploiting instance-aware semantic segmentation techniques and optical flow cues. We apply Structure from Motion (SfM) to object and background images to determine initial camera poses relative to object instances as well as background structures and refine the initial SfM results by integrating stereo camera constraints using factor graphs. We compute object trajectories using stereo sequence constraints of object and background reconstructions. We show qualitative results using publicly available video data of driving sequences. Due to the lack of suitable ground truth, we create a synthetic benchmark dataset of stereo sequences with vehicles in urban environments. Our algorithm achieves an average trajectory error of 0.09 meter using the dataset. The dataset is on our website1publicly available.
Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen
IV4
2019 End-to-end Prediction of Driver Intention using 3D Convolutional Neural Networks
abstract
Despite extraordinary progress of Advanced Driver Assistance Systems (ADAS), an alarming number of over 1,2 million people are still fatally injured in traffic accidents every year1. Human error is mostly responsible for such casualties, as by the time the ADAS system has alarmed the driver, it is often too late. We present a vision-based system based on deep neural networks with 3D convolutions and residual learning for anticipating the future maneuver based on driver observation. While previous work focuses on hand-crafted features (e.g. head pose), our model predicts the intention directly from video in an end-to-end fashion. Our architecture consists of three components: a neural network for extraction of optical flow, a 3D residual network for maneuver classification and a Long Short-Term Memory network (LSTM) for handling temporal data of varying length. To evaluate our idea, we conduct thorough experiments on the publicly available Brain4Cars benchmark, which covers both inside and outside views for future maneuver anticipation. Our model is able to predict driver intention with an accuracy of 83,12% and 4,07s before the beginning of the maneuver, outperforming state-of-the-art approaches, while considering the inside view only.
Patrick Gebert, Alina Roitberg, Monica-Laura Haurilet, Rainer Stiefelhagen
IV4
2019 Self-supervised Face-Grouping on Graphs
abstract
We propose a novel self-supervised method for fine-tuning deep face representations called Face-Grouping on Graphs. We apply our method to automatic face grouping, where characters are to be separated based on their identity. To solve this problem, a graph structure with positive and negative edges over a set of face-tracks based on their temporal overlap and similarity constraints is in- duced, which requires no manual labor. We compute feature repre- sentations over sub-sequences of each track (sub-tracks) in order to obtain robust features whilst being able to utilize information contained in face variance. Each sub-track is given the ability to exchange information with adjacent sub-tracks via a typed graph neural network running over the induced graph. This allows us to push each representation in a direction in feature space that groups all representations of the same character together and separates representations of different characters. We show that our method is capable of improving clustering accuracy on popular video face clustering datasets The Big Bang Theory and Buffy the Vampire Slayer by 4.9% and 17.0% respectively compared to baseline performance, and 0.52% respective 5.55% com- pared to state-of-the-art methods. Additionally, we achieve 19.0% absolute increase in B3 F-Score on Harry Potter 1 (ACCIO) over other state-of-the-art unsupervised methods. We provide perfor- mance metrics on all episodes of The Big Bang Theory and Buffy the Vampire Slayer to enable further comparison in the future.
Veith Röthlingshöfer, Vivek Sharma 0001, Rainer Stiefelhagen
ACM Multimedia3
2019 SPaSe - Multi-Label Page Segmentation for Presentation Slides
abstract
We introduce the first benchmark dataset for slide-page segmentation. Presentation slides are one of the most prominent document types used to exchange ideas across the web, educational institutes and businesses. This document format is marked with a complex layout which contains a rich variety of graphical (e.g. diagram, logo), textual (e.g. heading, affiliation) and structural components (e.g. enumeration, legend). This vast and popular knowledge source is still unattainable by modern machine learning technique due to lack of annotated data. To tackle this issue, we introduce SPaSe (Slide Page Segmentation), a novel dataset containing in total 2000 slides with dense, pixel-wise annotations of 25 classes. We show that slide segmentation reveals some interesting properties that characterize this task. Unlike the common image segmentation problem, disjoint classes tend to have a high overlap of regions, thus posing this segmentation task as a multi-label problem. Furthermore, many of the frequently encountered classes in slides are location sensitive (e.g. title, footnote). Hence, we believe our dataset represents a challenging and interesting benchmark for novel segmentation models. Finally, we evaluate state-of-the-art deep segmentation models on our dataset and show that it is suitable for developing deep learning models without any need of pre-training. Our dataset will be released to the public to foster further research on this interesting task.
Monica-Laura Haurilet, Ziad Al-Halah, Rainer Stiefelhagen
WACV3
2018 Informed Democracy: Voting-based Novelty Detection for Action Recognition
Alina Roitberg, Ziad Al-Halah, Rainer Stiefelhagen
BMVC3
2018 A Pose-Sensitive Embedding for Person Re-Identification With Expanded Cross Neighborhood Re-Ranking
abstract
Person re-identification is a challenging retrieval task that requires matching a person's acquired image across non-overlapping camera views. In this paper we propose an effective approach that incorporates both the fine and coarse pose information of the person to learn a discriminative embedding. In contrast to the recent direction of explicitly modeling body parts or correcting for misalignment based on these, we show that a rather straightforward inclusion of acquired camera view and/or the detected joint locations into a convolutional neural network helps to learn a very effective representation. To increase retrieval performance, re-ranking techniques based on computed distances have recently gained much attention. We propose a new unsupervised and automatic re-ranking framework that achieves state-of-the-art re-ranking performance. We show that in contrast to the current state-of-the-art re-ranking methods our approach does not require to compute new rank lists for each image pair (e.g., based on reciprocal neighbors) and performs well by using simple direct rank list based comparison or even by just using the already computed euclidean distances between the images. We show that both our learned representation and our re-ranking method achieve state-of-the-art performance on a number of challenging surveillance image and video datasets. Code is available at https://github.com/pse-ecn.
M. Saquib Sarfraz, Arne Schumann, Andreas Eberle, Rainer Stiefelhagen
CVPR4
2018 Classification-Driven Dynamic Image Enhancement
abstract
Convolutional neural networks rely on image texture and structure to serve as discriminative features to classify the image content. Image enhancement techniques can be used as preprocessing steps to help improve the overall image quality and in turn improve the overall effectiveness of a CNN. Existing image enhancement methods, however, are designed to improve the perceptual quality of an image for a human observer. In this paper, we are interested in learning CNNs that can emulate image enhancement and restoration, but with the overall goal to improve image classification and not necessarily human perception. To this end, we present a unified CNN architecture that uses a range of enhancement filters that can enhance image-specific details via end-to-end dynamic filter learning. We demonstrate the effectiveness of this strategy on four challenging benchmark datasets for fine-grained, object, scene, and texture classification: CUB-200-2011, PASCAL-VOC2007, MIT-Indoor, and DTD. Experiments using our proposed enhancement show promising results on all the datasets. In addition, our approach is capable of improving the performance of all generic CNN architectures.
Vivek Sharma 0001, Ali Diba, Davy Neven, Michael S. Brown, Luc Van Gool, Rainer Stiefelhagen
CVPR6
2018 3D Vehicle Trajectory Reconstruction in Monocular Video Data Using Environment Structure Constraints
Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen
ECCV (10)4
2018 Visual Shoreline Detection for Blind and Partially Sighted People
Daniel Koester, Tobias Allgeyer, Rainer Stiefelhagen
ICCHP (2)3
2018 UML4ALL Syntax - A Textual Notation for UML Diagrams
Claudia Loitsch, Karin Müller 0001, Stephan Seifermann, Jörg Henß, Sebastian Dieter Krach, Gerhard Jaworek, Rainer Stiefelhagen
ICCHP (1)7
2018 An Inclusive and Accessible LaTeX Editor
Giuseppe Melfi, Thorsten Schwarz, Rainer Stiefelhagen
ICCHP (1)3
2018 Prototype Development of a Low-Cost Vibro-Tactile Navigation Aid for the Visually Impaired
Vanessa Petrausch, Thorsten Schwarz, Rainer Stiefelhagen
ICCHP (2)3
2018 Optical Braille Recognition
Thorsten Schwarz, Reiner Dolp, Rainer Stiefelhagen
ICCHP (1)3
2018 Accessible EPUB: Making EPUB 3 Documents Universal Accessible
Thorsten Schwarz, Sachin Rajgopal, Rainer Stiefelhagen
ICCHP (1)3
2018 Body Pose and Context Information for Driver Secondary Task Detection
abstract
Distraction of the driver by secondary tasks is already dangerous while driving manually but especially in handover situations in an automated mode this can lead to critical situations. Currently, these tasks are not taken into account in most modern cars. We present a system that detects typical distracting secondary tasks in an efficient modular way. We first determine the body pose of the driver and afterwards use recurrent neuronal networks to estimate actions based on sequences of the captured body poses. Our system uses knowledge about the surroundings of the driver that is unique to the car environment. Our evaluation shows that this approach achieves better results than other state of the art systems for action recognition on our dataset.
Manuel Martin, Johannes Popp, Mathias Anneken, Michael Voit, Rainer Stiefelhagen
Intelligent Vehicles Symposium5
2018 Personal Perspective: Using Modified World Views to Overcome Real-Life Limitations in Virtual Reality
abstract
Virtual Reality opens up new possibilities as it allows to overcome real-life limitations and create novel experiences. While interacting with other people, it is beneficial to share a common view point. We modify the virtual world to allow face-to-face interaction with another person, while still retaining an optimal point of view on presented data. This is done by adapting the virtual environment independently for each user, using translation, rotation and scaling. The presented modification of the world gives a natural solution to the problems of collaborative analysis of content. It is therefore beneficial for usage in human-human interaction scenarios that support cooperative work.
Adrian Heinrich Hoppe, Florian van de Camp, Rainer Stiefelhagen
VR3
2017 Deep View-Sensitive Pedestrian Attribute Inference in an end-to-end Model
M. Saquib Sarfraz, Arne Schumann, Rainer Stiefelhagen
BMVC4
2017 Automatic Discovery, Association Estimation and Learning of Semantic Attributes for a Thousand Categories
abstract
Attribute-based recognition models, due to their impressive performance and their ability to generalize well on novel categories, have been widely adopted for many computer vision applications. However, usually both the attribute vocabulary and the class-attribute associations have to be provided manually by domain experts or large number of annotators. This is very costly and not necessarily optimal regarding recognition performance, and most importantly, it limits the applicability of attribute-based models to large scale data sets. To tackle this problem, we propose an end-to-end unsupervised attribute learning approach. We utilize online text corpora to automatically discover a salient and discriminative vocabulary that correlates well with the human concept of semantic attributes. Moreover, we propose a deep convolutional model to optimize class-attribute associations with a linguistic prior that accounts for noise and missing data in text. In a thorough evaluation on ImageNet, we demonstrate that our model is able to efficiently discover and learn semantic attributes at a large scale. Furthermore, we demonstrate that our model outperforms the state-of-the-art in zero-shot learning on three data sets: ImageNet, Animals with Attributes and aPascal/aYahoo. Finally, we enable attribute-based learning on ImageNet and will share the attributes and associations for future research.
Ziad Al-Halah, Rainer Stiefelhagen
CVPR2
2017 Marlin: A High Throughput Variable-to-Fixed Codec Using Plurally Parsable Dictionaries
abstract
We present Marlin, a variable-to-fixed (VF) codec optimized for decoding speed. Marlin builds upon a novel way of constructing VF dictionaries that maximizes efficiency for a given dictionary size. On a lossless image coding experiment, Marlin achieves a compression ratio of 1.94 at 2494MiB/s. Marlin is as fast as state-of-the-art high-throughput codecs (e.g., Snappy, 1.24 at 2643MiB/s), and its compression ratio is close to the best entropy codecs (e.g., FiniteStateEntropy, 2.06 at 523MiB/s). Therefore, Marlin enables efficient and high throughput encoding for memoryless sources, which was not possible until now.
Manuel Martínez 0001, Monica-Laura Haurilet, Rainer Stiefelhagen, Joan Serra-Sagristà
DCC3
2017 Fashion Forward: Forecasting Visual Style in Fashion
abstract
What is the future of fashion? Tackling this question from a data-driven vision perspective, we propose to forecast visual style trends before they occur. We introduce the first approach to predict the future popularity of styles discovered from fashion images in an unsupervised manner. Using these styles as a basis, we train a forecasting model to represent their trends over time. The resulting model can hypothesize new mixtures of styles that will become popular in the future, discover style dynamics (trendy vs. classic), and name the key visual attributes that will dominate tomorrow's fashion. We demonstrate our idea applied to three datasets encapsulating 80,000 fashion products sold across six years on Amazon. Results indicate that fashion forecasting benefits greatly from visual analysis, much more than textual or meta-data cues surrounding products.
Ziad Al-Halah, Rainer Stiefelhagen, Kristen Grauman
ICCV2
2017 Breathing Rate Monitoring during Sleep from a Depth Camera under Real-Life Conditions
abstract
Computer vision is a non-invasive way to supervise patients in the bed. We introduce a novel algorithm that monitors breathing rate from a depth camera placed above the bed. While visually registering breathing rate has raised significant interest in the health care field, most published approaches are evaluated only on constrained or simulated settings. Conversely, we evaluate our method in a real dataset consisting of 3,239 segments collected from 67 sleep laboratory patients. Our method introduces three novel contributions: a dynamic Region of Interest (RoI) which is aligned to the bed, a confidence metric based on patient agitation, and the Early Fourier Fusion strategy. Overall, our camera based method is accurate on 85.9% of the segments. This performance is similar to the obtained from a chest sensor (88.7%). Most importantly, we report the performance impact related to different sleep conditions, like apnea, position and staging.
Manuel Martínez 0001, Rainer Stiefelhagen
WACV2
2017 Deep Perceptual Mapping for Cross-Modal Face Recognition
M. Saquib Sarfraz, Rainer Stiefelhagen
Int. J. Comput. Vis.2
2016 Automatic generation of scene-specific person trackers
abstract
The large variety of influencing factors requires the manual creation and optimization of person trackers for different environments as well as different views, such as in a distributed network of cameras. The manual creation and adjustments are time-consuming as well as prone to error and therefore expensive. We propose a system that uses basic computer-vision building blocks to automatically create and optimize a person tracker, using only a few annotated camera frames. An evaluation shows that our system creates target-oriented trackers from largely very basic methods, that can compare to manually, published person trackers. The system has the potential to drastically reduce the cost of creating person trackers for new environments and optimizing it for every single camera in a camera network.
Gerrit Holzbach, Florian van de Camp, Rainer Stiefelhagen
AVSS3
2016 Recovering the Missing Link: Predicting Class-Attribute Associations for Unsupervised Zero-Shot Learning
abstract
Collecting training images for all visual categories is not only expensive but also impractical. Zero-shot learning (ZSL), especially using attributes, offers a pragmatic solution to this problem. However, at test time most attribute-based methods require a full description of attribute associations for each unseen class. Providing these associations is time consuming and often requires domain specific knowledge. In this work, we aim to carry out attribute-based zero-shot classification in an unsupervised manner. We propose an approach to learn relations that couples class embeddings with their corresponding attributes. Given only the name of an unseen class, the learned relationship model is used to automatically predict the class-attribute associations. Furthermore, our model facilitates transferring attributes across data sets without additional effort. Integrating knowledge from multiple sources results in a significant additional improvement in performance. We evaluate on two public data sets: Animals with Attributes and aPascal/aYahoo. Our approach outperforms state-of the-art methods in both predicting class-attribute associations and unsupervised ZSL by a large margin.
Ziad Al-Halah, Makarand Tapaswi, Rainer Stiefelhagen
CVPR3
2016 MovieQA: Understanding Stories in Movies through Question-Answering
abstract
We introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text. The dataset consists of 14,944 questions about 408 movies with high semantic diversity. The questions range from simpler "Who" did "What" to "Whom", to "Why" and "How" certain events occurred. Each question comes with a set of five possible answers, a correct one and four deceiving answers provided by human annotators. Our dataset is unique in that it contains multiple sources of information – video clips, plots, subtitles, scripts, and DVS [32]. We analyze our data through various statistics and methods. We further extend existing QA techniques to show that question-answering with such open-ended semantics is hard. We make this data set public along with an evaluation benchmark to encourage inspiring work in this challenging domain.
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba 0001, Raquel Urtasun, Sanja Fidler
CVPR3
2016 Zebra Crossing Detection from Aerial Imagery Across Countries
Daniel Koester, Björn Lunt, Rainer Stiefelhagen
ICCHP (2)3
2016 User Requirements Regarding Information Included in Audio-Tactile Maps for Individuals with Blindness
Konstantinos Papadopoulos 0001, Konstantinos Charitakis, Lefkothea Kartasidou, Georgios Kouroupetroglou, Suad Sakalli Gumus, Efstratios Stylianidis, Rainer Stiefelhagen, Karin Müller 0001, Engin Yilmaz, Gerhard Jaworek, Christos Polimeras, Utku Sayin, Nikolaos Oikonomidis, Nikolaos Lithoxopoulos
ICCHP (2)7
2016 Mobile Interactive Image Sonification for the Blind
Torsten Wörtwein, Boris Schauerte, Karin Müller 0001, Rainer Stiefelhagen
ICCHP (1)4
2016 Sleep position classification from a depth camera using Bed Aligned Maps
abstract
Sleep position is an important feature used to assess the quality and quantity of an individual's sleep. Furthermore, it is related to sleep disorders like sleep apnoea and snoring, and needs to be tracked in nursery homes to avoid pressure ulcers. Therefore, a gravity sensor attached to the chest is generally used to register body position during sleep studies. We suggest a non-intrusive and cost-efficient approach to detect the sleep position based on a single depth camera. Compared to alternative state-of-the-art approaches, ours require no calibration, and has been evaluated on a real setting comprising 78 patients from a sleep laboratory. We use the Bed Aligned Maps to extract a low resolution descriptor from a depth map which is aligned to the bed position, We perform classification using Convolutional Neural Networks, achieving an accuracy of 94.0%, thus outperforming current state-of-the-art algorithms and even the contact sensor from the sleep laboratory which achieves an accuracy of 91.9%.
Timo Grimm, Manuel Martínez 0001, Andreas Benz, Rainer Stiefelhagen
ICPR4
2016 Predicting lane keeping behavior of visually distracted drivers using inverse suboptimal control
abstract
Driver distraction strongly contributes to crash-risk. Therefore, assistance systems that warn drivers if their distraction poses a hazard to road safety, promise a great safety benefit. Current approaches either seek to detect critical situations using environmental sensors or estimate a driver's attention state solely from his/her behavior. However, this neglects that driving situation, driver deficiencies and compensation strategies altogether determine the risk of an accident. This work proposes to use inverse suboptimal control to predict these aspects in visually distracted lane keeping. In contrast to other approaches, this allows a situation-dependent assessment of the risk posed by distraction. Real traffic data of seven drivers are used for evaluation of the predictive power of our approach. For comparison, a baseline was built using established behavior models. In the evaluation our method achieves a consistently lower prediction error over speed and track-topology variations. Additionally, our approach generalizes better to driving speeds unseen in training phase.
Felix Schmitt 0001, Hans-Joachim Bieg, Dietrich Manstetten, Michael Herman, Rainer Stiefelhagen
Intelligent Vehicles Symposium5
2016 Exact Maximum Entropy Inverse Optimal Control for modeling human attention switching and control
abstract
Maximum Causal Entropy (MCE) Inverse Optimal Control (IOC) has become an effective tool for modeling human behavior in many control tasks. Its advantage over classic techniques for estimating human policies is the transferability of the inferred objectives: Behavior can be predicted in variations of the control task by policy computation using a relaxed optimality criterion. However, exact policy inference is often computationally intractable in control problems with imperfect state observation. In this work, we present a model class that allows modeling human control of two tasks of which only one be perfectly observed at a time requiring attention switching. We show how efficient and exact objective and policy inference via MCE can be conducted for these control problems. Both MCE-IOC and Maximum Causal Likelihood (MCL)-IOC, a variant of the original MCE approach, as well as Direct Policy Estimation (DPE) are evaluated using simulated and real behavioral data. Prediction error and generalization over changes in the control process are both considered in the evaluation. The results show a clear advantage of both IOC methods over DPE, especially in the transfer over variation of the control process. MCE and MCL performed similar when training on a large set of simulated data, but differed significantly on small sets and real data.
Felix Schmitt 0001, Hans-Joachim Bieg, Dietrich Manstetten, Michael Herman, Rainer Stiefelhagen
SMC5
2016 Naming TV characters by watching and analyzing dialogs
abstract
Person identification in TV series has been a popular research topic over the last decade. In this area, most approaches either use manually annotated data or extract character supervision from a combination of subtitles and transcripts. However, both approaches have key drawbacks that hinder application of these methods at a large scale - manual annotation is expensive and transcripts are often hard to obtain. We investigate the topic of automatically labeling all character appearances in TV series using information obtained solely from subtitles. This task is extremely difficult as the dialogs between characters provide very sparse and weakly supervised data. We address these challenges by exploiting recent advances in face descriptors and Multiple Instance Learning methods. We propose methods to create MIL bags and evaluate and discuss several MIL techniques. The best combination achieves an average precision over 80% on three diverse TV series. We demonstrate that only using subtitles provides good results on identifying characters in TV series and wish to encourage the community towards this problem.
Monica-Laura Haurilet, Makarand Tapaswi, Ziad Al-Halah, Rainer Stiefelhagen
WACV4
2016 HeHOP: Highly efficient head orientation and position estimation
abstract
Continuous head pose estimation is an important visual component for human-computer interaction. However, an accurate and computationally efficient method to estimate the head orientation and position remains a challenging task in computer vision. We propose a Highly efficient Head Orientation and Position estimation (HeHOP) approach based on depth data which uses a stage-by-stage regression framework. At each stage, binary features are obtained from local areas of depth information. A global linear mapping is used to predict the head orientation and position update using the binary features. We evaluate our method on the BIWI dataset containing depth images labeled with head orientation and position. The results show that our approach is robust against occlusions and achieves state-of-the-art performance in terms of accuracy, has a low miss rate, and is several times faster than previous methods.
Anke Schwarz, Zhuang Lin, Rainer Stiefelhagen
WACV3
2016 Transfer metric learning for action similarity using high-level semantics
Ziad Al-Halah, Lukas Rybok, Rainer Stiefelhagen
Pattern Recognit. Lett.3
2015 Message from general chairs
abstract
AVSS is the premier annual international conference in the field of video and signal-based surveillance that brings together experts from academia, industry, and government to advance theories, methods, systems, and applications related to surveillance.
Jürgen Beyerer, Rainer Stiefelhagen
AVSS2
2015 Transferring attributes for person re-identification
abstract
Person re-identification is an important computer vision task with many applications in areas such as surveillance or multimedia. Approaches relying on handcrafted image features struggle with many factors (e.g. lighting, camera angle) which lead to a large variety in visual appearance for the same individual. Features based on semantic attributes of a person's appearance can help with some of these challenges. In this work we describe an approach that integrates such attributes with existing re-identification methods based on low-level features. We start by training a set of attribute classifiers and present a metric learning approach that uses these attributes for person re-identification. The method is then applied to a second dataset without attributes labels by transferring the attributes classifiers. Performance on the target dataset can be increased by applying a whitening transformation prior to transfer. We present experiments on publicly available datasets and demonstrate the performance improvement gained by this added re-identification cue.
Arne Schumann, Rainer Stiefelhagen
AVSS2
2015 Deep Perceptual Mapping for Thermal to Visible Face Recogntion
abstract
Cross modal face matching between the thermal and visible spectrum is a much desired capability for night-time surveillance and security applications. Due to a very large modality gap, thermal-to-visible face recognition is one of the most challenging face matching problem. In this paper, we present an approach to bridge this modality gap by a significant margin. Our approach captures the highly non-linear relationship between the two modalities by using a deep neural network. Our model attempts to learn a non-linear mapping from visible to thermal spectrum while preserving the identity information. We show substantive performance improvement on a difficult thermal-visible face dataset. The presented approach improves the state-of-the-art by more than 10% in terms of Rank-1 identification and bridge the drop in performance due to the modality gap by more than 40%.
M. Saquib Sarfraz, Rainer Stiefelhagen
BMVC2
2015 Book2Movie: Aligning video scenes with book chapters
abstract
Film adaptations of novels often visually display in a few shots what is described in many pages of the source novel. In this paper we present a new problem: to align book chapters with video scenes. Such an alignment facilitates finding differences between the adaptation and the original source, and also acts as a basis for deriving rich descriptions from the novel for the video clips. We propose an efficient method to compute an alignment between book chapters and video scenes using matching dialogs and character identities as cues. A major consideration is to allow the alignment to be non-sequential. Our suggested shortest path based approach deals with the non-sequential alignments and can be used to determine whether a video scene was part of the original book. We create a new data set involving two popular novel-to-film adaptations with widely varying properties and compare our method against other text-to-video alignment baselines. Using the alignment, we present a qualitative analysis of describing the video through rich narratives obtained from the novel.
Makarand Tapaswi, Martin Bäuml, Rainer Stiefelhagen
CVPR3
2015 Color decorrelation helps visual saliency detection
abstract
We present how color decorrelation allows visual saliency models to achieve higher performance when predicting where people look in images. For this purpose, we decorrelate the color information of each image, which leads to an image-specific color space with decorrelated color components. This way, we are able to improve the performance of several well-known visual saliency algorithms such as, for example, Itti and Koch's model and Hou and Zhang's spectral residual saliency. We show the advantage of color decorrelation on three eye-tracking datasets (Kootstra, Toronto, and MIT) with respect to three evaluation measures (AUC, CC, and NSS).
Boris Schauerte, Torsten Wörtwein, Rainer Stiefelhagen
ICIP3
2015 Multimodal Public Speaking Performance Assessment
abstract
The ability to speak proficiently in public is essential for many professions and in everyday life. Public speaking skills are difficult to master and require extensive training. Recent developments in technology enable new approaches for public speaking training that allow users to practice in engaging and interactive environments. Here, we focus on the automatic assessment of nonverbal behavior and multimodal modeling of public speaking behavior. We automatically identify audiovisual nonverbal behaviors that are correlated to expert judges' opinions of key performance aspects. These automatic assessments enable a virtual audience to provide feedback that is essential for training during a public speaking performance. We utilize multimodal ensemble tree learners to automatically approximate expert judges' evaluations to provide post-hoc performance assessments to the speakers. Our automatic performance evaluation is highly correlated with the experts' opinions with r = 0.745 for the overall performance assessments. We compare multimodal approaches with single modalities and find that the multimodal ensembles consistently outperform single modalities.
Torsten Wörtwein, Mathieu Chollet, Boris Schauerte, Louis-Philippe Morency, Rainer Stiefelhagen, Stefan Scherer
ICMI5
2015 Interactive Web-based Image Sonification for the Blind
abstract
In this demonstration, we show a web-based sonification platform that allows blind users to interactively experience various information using two nowadays widespread technologies: modern web browsers that implement high-level JavaScript APIs and touch-sensitive displays. This way, blind users can easily access information such as, for example, maps or graphs. Our current prototype provides various sonifications that can be switched depending on the image type and user preference. The prototype runs in Chrome and Firefox on PCs, smart phones, and tablets.
Torsten Wörtwein, Boris Schauerte, Karin Müller 0001, Rainer Stiefelhagen
ICMI4
2015 Pedestrian intention recognition using Latent-dynamic Conditional Random Fields
abstract
We present a novel approach for pedestrian intention recognition for advanced video-based driver assistance systems using a Latent-dynamic Conditional Random Field model. The model integrates pedestrian dynamics and situational awareness using observations from a stereo-video system for pedestrian detection and human head pose estimation. The model is able to capture both intrinsic and extrinsic class dynamics. Evaluation of our method is performed on a public available dataset addressing scenarios of lateral approaching pedestrians that might cross the road, turn into the road or stop at the curbside. During experiments, we demonstrate that the proposed approach leads to better stability and class separation compared to state-of-the-art pedestrian intention recognition approaches.
Andreas T. Schulz, Rainer Stiefelhagen
Intelligent Vehicles Symposium2
2015 Accio: A Data Set for Face Track Retrieval in Movies Across Age
abstract
Video face recognition is a very popular task and has come a long way. The primary challenges such as illumination, resolution and pose are well studied through multiple data sets. However there are no video-based data sets dedicated to study the effects of aging on facial appearance. We present a challenging face track data set, Harry Potter Movies Aging Data set (Accio1), to study and develop age invariant face recognition methods for videos. Our data set not only has strong challenges of pose, illumination and distractors, but also spans a period of ten years providing substantial variation in facial appearance. We propose two primary tasks: within and across movie face track retrieval; and two protocols which differ in their freedom to use external data. We present baseline results for the retrieval performance using a state-of-the-art face track descriptor. Our experiments show clear trends of reduction in performance as the age gap between the query and database increases. We will make the data set publicly available for further exploration in age-invariant video face recognition.
Esam Ghaleb, Makarand Tapaswi, Ziad Al-Halah, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICMR5
2015 How to Transfer? Zero-Shot Object Recognition via Hierarchical Transfer of Semantic Attributes
abstract
Attribute based knowledge transfer has proven very successful in visual object analysis and learning previously unseen classes. However, the common approach learns and transfers attributes without taking into consideration the embedded structure between the categories in the source set. Such information provides important cues on the intraattribute variations. We propose to capture these variations in a hierarchical model that expands the knowledge source with additional abstraction levels of attributes. We also provide a novel transfer approach that can choose the appropriate attributes to be shared with an unseen class. We evaluate our approach on three public datasets: a Pascal, Animals with Attributes and CUB-200-2011 Birds. The experiments demonstrate the effectiveness of our model with significant improvement over state-of-the-art.
Ziad Al-Halah, Rainer Stiefelhagen
WACV2
2015 3D Pictorial Structures for Human Pose Estimation with Supervoxels
abstract
Pictorial structures provide a powerful framework for human pose estimation, in particular in the domain of 2D data. However, solving pictorial structures directly in 3D drastically increases its complexity and it quickly exceeds tractable dimensions. In this paper, we propose a discretization-by-segmentation approach by applying super voxels to 3D pictorial structures which significantly reduces the search space. The proposed 3D pictorial structures approach achieves 3D errors of 115 mm and 135 mm on the Human Eva-I and UMPM datasets and PCP scores of 78% and 75%, respectively. Due to the search space reduction, the overall pose estimation runtime is below 100 ms which is up to four orders of magnitude faster than comparable 3D pictorial structure approaches. The presented approach is not limited to human pose estimation, but provides a general and efficient solution for 3D pictorial structures.
Alexander Schick, Rainer Stiefelhagen
WACV2
2014 Real Time Head Model Creation and Head Pose Estimation on Consumer Depth Cameras
abstract
Head pose estimation is an important part of the human perception and is therefore also relevant to make interaction with computer systems more natural. However, accurate estimation of the pose in a wide range is a challenging computer vision problem. We present an accurate approach for head pose estimation on consumer depth cameras that works in a wide pose range without prior knowledge about the tracked person and without prior training of a detector. Our algorithm builds and registers a 3D head model with the iterative closest point algorithm. To track the head pose using this head model an initialization with a known pose is necessary. Instead of providing such an initialization manually we determine the initial pose using features of the head and improve this pose over time. An evaluation shows that our algorithm works in real time with limited resources and achieves superior accuracy compared to other state of the art systems. Our main contribution is the combination of features of the head and the head model generation to build a detector that gives accurate results in a wide pose range.
Manuel Martínez 0001, Florian van de Camp, Rainer Stiefelhagen
3DV3
2014 A time pooled track kernel for person identification
abstract
We present a novel method for comparing tracks by means of a time pooled track kernel. In contrast to spatial or feature-space pooling, the track kernel pools base kernel results within tracks over time. It includes as special cases frame-wise classification on the one hand and the normalized sum kernel on the other hand. We also investigate non-Mercer instantiations of the track kernel and obtain good results despite its Gram matrices not being positive semidefinite. Second, the track kernel matrices in general require less memory than single frame kernels, allowing to process larger datasets without resorting to subsampling. Finally, the track kernel formulation allows for very fast testing compared to frame-wise classification which is important in settings where user feedback is obtained and quick iterations of re-training and re-testing are required. We apply our approach to the task of video-based person identification in large scale settings and obtain state-of-the art results.
Martin Bäuml, Makarand Tapaswi, Rainer Stiefelhagen
AVSS3
2014 StoryGraphs: Visualizing Character Interactions as a Timeline
abstract
We present a novel way to automatically summarize and represent the storyline of a TV episode by visualizing character interactions as a chart. We also propose a scene detection method that lends itself well to generate over-segmented scenes which is used to partition the video. The positioning of character lines in the chart is formulated as an optimization problem which trades between the aesthetics and functionality of the chart. Using automatic person identification, we present StoryGraphs for 3 diverse TV series encompassing a total of 22 episodes. We define quantitative criteria to evaluate StoryGraphs and also compare them against episode summaries to evaluate their ability to provide an overview of the episode.
Makarand Tapaswi, Martin Bäuml, Rainer Stiefelhagen
CVPR3
2014 Cognitive Evaluation of Haptic and Audio Feedback in Short Range Navigation Tasks
Manuel Martínez 0001, Angela Constantinescu, Boris Schauerte, Daniel Koester, Rainer Stiefelhagen
ICCHP (2)5
2014 Kinect unbiased
abstract
Since its release, Kinect has been the de facto standard for low-cost RGB-D sensors. An infrared laser ray shot through an holographic diffraction grating projects a fixed dot pattern which is captured using an infrared camera. The pseudo-random pattern ensures that a simple block matching algorithm suffices to provide reliable depth estimates, allowing a cost-effective implementation. In this paper, we analyze the software limitations of Kinect's method, which allows us to propose algorithms that provide better precision. First, we analyze the dot pattern: we measure its pincushion distortion and its effect on the dot density, which is smaller towards the edges of the image. Then, we analyze the behavior of Block Matching algorithms, we show how Kinect's Block Matching implementation is; in general; limited by the dot density of the pattern, and a significant spatial bias is introduced as a result. We propose an efficient approach to estimate the disparity of each dot, allowing us to produce a point cloud with better spatial resolution than Block Matching algorithms.
Manuel Martínez 0001, Rainer Stiefelhagen
ICIP2
2014 Cleaning up after a face tracker: False positive removal
abstract
Automatic person identification in TV series has gained popularity over the years. While most of the works rely on using face-based recognition, errors during tracking such as false positive face tracks are typically ignored. We propose a variety of methods to remove false positive face tracks and categorize the methods into confidence- and context-based. We evaluate our methods on a large TV series data set and show that up to 75% of the false positive face tracks are removed at the cost of 3.6% true positive tracks. We further show that the proposed method is general and applicable to other detectors or trackers.
Makarand Tapaswi, Cemal Cagn Corez, Martin Bäuml, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICIP5
2014 What to Transfer? High-Level Semantics in Transfer Metric Learning for Action Similarity
abstract
Learning from few examples is considered a very challenging task where transfer learning proved to be beneficial. Such a learning framework exploits previous experiences and knowledge to compensate for the lack of training data in a novel domain. Knowledge representation plays a vital role in the type and performance of transfer learning approaches, as well as its robustness against negative transfer effect. This aspect is usually not considered in most of the proposed transfer learning methodologies, where the focus is either on the transfer type or on the representation. In this work, we study the use of various high-level semantics in transfer metric learning. We propose a generic transfer metric learning framework, and analyze the effect of different semantic similarity spaces on transfer type and efficiency against negative transfer. Furthermore, we introduce a hierarchical knowledge representation model based on the embedded structure in the attribute semantic space. The evaluation of the framework on challenging transfer settings in the context of action similarity demonstrates the effectiveness of our approach.
Ziad Al-Halah, Lukas Rybok, Rainer Stiefelhagen
ICPR3
2014 Manifold Alignment for Person Independent Appearance-Based Gaze Estimation
abstract
We show that dually supervised manifold embedding can improve the performance of machine learning based person-independent and thus calibration-free gaze estimation. For this purpose, we perform a manifold embedding for each person in the training dataset and then learn a linear transformation that aligns the individual, person-dependent manifolds. We evaluate the effect of manifold alignment on the recently presented Columbia dataset, where we analyze the influence on 6 regression methods and 8 feature variants. Using manifold alignment, we are able to improve the person-independent gaze estimation performance by up to 31.2 % compared to the best approach without manifold alignment.
Timo Schneider, Boris Schauerte, Rainer Stiefelhagen
ICPR3
2014 "Look at this!" learning to guide visual saliency in human-robot interaction
abstract
We learn to direct visual saliency in multimodal (i.e., pointing gestures and spoken references) human-robot interaction to highlight and segment arbitrary referent objects. For this purpose, we train a conditional random field to integrate features that reflect low-level visual saliency, the likelihood of salient objects, the probability that a given pixel is pointed at, and - if available - spoken information about the target object's visual appearance. As such, this work integrates several of our ideas and approaches, ranging from multi-scale spectral saliency detection, spatially debiased salient object detection, computational attention in human-robot interaction to learning robust color term models. We demonstrate that this machine learning driven integration outperforms the previously reported results on two datasets, one dataset without and one with spoken object references. In summary, for automatically detected pointing gestures and automatically extracted object references, our approach improves the rate at which the correct object is included in the initial focus of attention by 10.37% in the absence and 25.21% in the presence of spoken target object information.
Boris Schauerte, Rainer Stiefelhagen
IROS2
2014 Story-based Video Retrieval in TV series using Plot Synopses
abstract
We present a novel approach to search for plots in the storyline of structured videos such as TV series. To this end, we propose to align natural language descriptions of the videos, such as plot synopses, with the corresponding shots in the video. Guided by subtitles and person identities the alignment problem is formulated as an optimization task over all possible assignments and solved efficiently using dynamic programming. We evaluate our approach on a novel dataset comprising of the complete season 5 of Buffy the Vampire Slayer, and show good alignment performance and the ability to retrieve plots in the storyline.
Makarand Tapaswi, Martin Bäuml, Rainer Stiefelhagen
ICMR3
2014 "Important stuff, everywhere!" Activity recognition with salient proto-objects as context
abstract
Object information is an important cue to discriminate between activities that draw part of their meaning from context. Most of current work either ignores this information or relies on specific object detectors. However, such object detectors require a significant amount of training data and complicate the transfer of the action recognition framework to novel domains with different objects and object-action relationships. Motivated by recent advances in saliency detection, we propose to employ salient proto-objects for unsupervised discovery of object- and object-part candidates and use them as a contextual cue for activity recognition. Our experimental evaluation on three publicly available data sets shows that the integration of proto-objects and simple motion features substantially improves recognition performance, outperforming the state-of-the-art.
Lukas Rybok, Boris Schauerte, Ziad Al-Halah, Rainer Stiefelhagen
WACV4
2014 An evaluation of the compactness of superpixels
Alexander Schick, Mika Fischer, Rainer Stiefelhagen
Pattern Recognit. Lett.3
2013 Person tracking-by-detection with efficient selection of part-detectors
abstract
In this paper we introduce a new person tracking-by-detection approach based on a particle filter. We leverage detection and appearance cues and apply explicit occlusion reasoning. The approach samples efficiently from a large set of available person part-detectors in order to increase runtime performance while retaining accuracy. The tracking approach is evaluated and compared to the state of the art on the CAVIAR surveillance dataset as well as on a multimedia dataset consisting of six episodes of the TV series The Big Bang Theory. The results demonstrate the versatility of the approach on very different types of data and its robustness to camera movement and non-pedestrian body poses.
Arne Schumann, Martin Bäuml, Rainer Stiefelhagen
AVSS3
2013 RPM: Random Points Matching for Pair wise Face-Similarity
M. Saquib Sarfraz, Muhammad Adnan Siddique, Rainer Stiefelhagen
BMVC3
2013 "BAM!" Depth-Based Body Analysis in Critical Care
Manuel Martínez 0001, Boris Schauerte, Rainer Stiefelhagen
CAIP (1)3
2013 Semi-supervised Learning with Constraints for Person Identification in Multimedia Data
abstract
We address the problem of person identification in TV series. We propose a unified learning framework for multi-class classification which incorporates labeled and unlabeled data, and constraints between pairs of features in the training. We apply the framework to train multinomial logistic regression classifiers for multi-class face recognition. The method is completely automatic, as the labeled data is obtained by tagging speaking faces using subtitles and fan transcripts of the videos. We demonstrate our approach on six episodes each of two diverse TV series and achieve state-of-the-art performance.
Martin Bäuml, Makarand Tapaswi, Rainer Stiefelhagen
CVPR3
2013 "Wow!" Bayesian surprise for salient acoustic event detection
abstract
We extend our previous work and present how Bayesian surprise can be applied to detect salient acoustic events. Therefore, we use the Gamma distribution to model each frequencies spectrogram distribution. Then, we use the Kullback-Leibler divergence of the posterior and prior distribution to calculate how “unexpected” and thus surprising newly observed audio samples are. This way, we are able to efficiently detect arbitrary, unexpected and thus surprising acoustic events. Complementing our qualitative system evaluations for (humanoid) robots, we demonstrate the effectiveness and practical applicability of the approach on the CLEAR 2007 acoustic event detection data.
Boris Schauerte, Rainer Stiefelhagen
ICASSP2
2013 How the distribution of salient objects in images influences salient object detection
abstract
We investigate the spatial distribution of salient objects in images. First, we empirically show that the centroid locations of salient objects correlate strongly with a centered, half-Gaussian model. This is an important insight, because it provides a justification for the integration of such a center bias in salient object detection algorithms. Second, we assess the influence of the center bias on salient object detection. Therefore, we integrate an explicit center bias into Cheng's state-of-the-art salient object detection algorithm. This way, first, we quantify the influence of the Gaussian center bias on salient object detection, second, improve the performance with respect to several established evaluation measures, and, third, derive a state-of-the-art unbiased salient object detection algorithm.
Boris Schauerte, Rainer Stiefelhagen
ICIP2
2013 GlueTK: a framework for multi-modal, multi-display human-machine-interaction
abstract
As new input modalities allow interaction not only in front of a single display, but enable interaction in the whole room, application developers face new challenges. They have to handle many new input modalities, each with its own interface and requirements for pre-processing, deal with multiple displays, and applications that are distributed across multiple machines. We present glueTK, a framework that abstracts from the complexities of these input modalities, allows the design of interfaces for a wide range of display sizes, and makes the distribution across multiple machines transparent to the developer as well as the user. With an example application we demonstrate the wide range of input modalities glueTK can support and the functionality this enables. GlueTK moves away from the focus on point and touch like input modalities, enabling the design of applications tailored towards interactive rooms instead of the traditional desktop environment.
Florian van de Camp, Rainer Stiefelhagen
IUI2
2012 Contextual Constraints for Person Retrieval in Camera Networks
abstract
We use contextual constraints for person retrieval in camera networks. We start by formulating a set of general positive and negative constraints on the identities of person tracks in camera networks, such as a person cannot appear twice in the same frame. We then show how these constraints can be used to improve person retrieval. First, we use the constraints to obtain training data in an unsupervised way to learn a general metric that is better suited to discriminate between different people than the Euclidean distance. Second, starting from an initial query track, we enhance the query-set using the constraints to obtain additional positive and negative samples for the query. Third, we formulate the person retrieval task as an energy minimization problem, integrate track scores and constraints in a common framework and jointly optimize the retrieval over all interconnected tracks. We evaluate our approach on the CAVIAR dataset and achieve 22% relative performance improvement in terms of mean average precision over standard retrieval where each track is treated independently.
Martin Bäuml, Makarand Tapaswi, Arne Schumann, Rainer Stiefelhagen
AVSS4
2012 Face Alignment Using a Ranking Model based on Regression Trees
abstract
In this work, we exploit the regression trees-based ranking model, which has been successfully applied in the domain of web-search ranking, to build appearance models for face alignment. The model is an ensemble of regression trees which is learned with gradient boosting. The MCT (Modified Census Transform) as well as its unbinarized version PCT (Pseudo Census Transform) are used as features due to their robustness to illumination changes. To avoid the overfitting problem in gradient boosting, we use random trees to initialize the boosting. The Nelder Mead’s simplex method is applied for fitting the learned model. We compare the proposed regression trees-based pointwise ranking model to pairwise ranking model. Experiments show that the proposed model improves both robustness and accuracy for face alignment.
Hua Gao, Hazim Kemal Ekenel, Rainer Stiefelhagen
BMVC3
2012 "Knock! Knock! Who is it?" probabilistic person identification in TV-series
abstract
We describe a probabilistic method for identifying characters in TV series or movies. We aim at labeling every character appearance, and not only those where a face can be detected. Consequently, our basic unit of appearance is a person track (as opposed to a face track). We model each TV series episode as a Markov Random Field, integrating face recognition, clothing appearance, speaker recognition and contextual constraints in a probabilistic manner. The identification task is then formulated as an energy minimization problem. In order to identify tracks without faces, we learn clothing models by adapting available face recognition results. Within a scene, as indicated by prior analysis of the temporal structure of the TV series, clothing features are combined by agglomerative clustering. We evaluate our approach on the first 6 episodes of The Big Bang Theory and achieve an absolute improvement of 20% for person identification and 12% for face recognition.
Makarand Tapaswi, Martin Bäuml, Rainer Stiefelhagen
CVPR3
2012 Quaternion-Based Spectral Saliency Detection for Eye Fixation Prediction
Boris Schauerte, Rainer Stiefelhagen
ECCV (2)2
2012 An Assistive Vision System for the Blind That Helps Find Lost Things
Boris Schauerte, Manuel Martínez 0001, Angela Constantinescu, Rainer Stiefelhagen
ICCHP (2)4
2012 Vision-based handwriting recognition for unrestricted text input in mid-air
abstract
We propose a vision-based system that recognizes handwriting in mid-air. The system does not depend on sensors or markers attached to the users and allows unrestricted character and word input from any position. It is the result of combining handwriting recognition based on Hidden Markov Models with multi-camera 3D hand tracking. We evaluated the system for both quantitative and qualitative aspects. The system achieves recognition rates of 86.15% for character and 97.54% for small-vocabulary isolated word recognition. Limitations are due to slow and low-resolution cameras or physical strain. Overall, the proposed handwriting recognition system provides an easy-to-use and accurate text input modality without placing restrictions on the users.
Alexander Schick, Daniel Morlock, Christoph Amma, Tanja Schultz, Rainer Stiefelhagen
ICMI5
2012 A ranking model for face alignment with Pseudo Census Transform
Hua Gao, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICPR3
2012 Breath rate monitoring during sleep using near-ir imagery and PCA
Manuel Martínez 0001, Rainer Stiefelhagen
ICPR2
2012 Robust multi-pose face tracking by multi-stage tracklet association
Markus Roth, Martin Bäuml, Ramakant Nevatia, Rainer Stiefelhagen
ICPR4
2012 Learning robust color name models from web images
Boris Schauerte, Rainer Stiefelhagen
ICPR2
2012 Measuring and evaluating the compactness of superpixels
Alexander Schick, Mika Fischer, Rainer Stiefelhagen
ICPR3
2012 Multimodal saliency-based attention: A lazy robot's approach
abstract
We extend our work on an integrated object-based system for saliency-driven overt attention and knowledge-driven object analysis. We present how we can reduce the amount of necessary head movement during scene analysis while still focusing all salient proto-objects in an order that strongly favors proto-objects with a higher saliency. Furthermore, we integrated motion saliency and as a consequence adaptive predictive gaze control to allow for efficient gazing behavior on the ARMAR-III robot head. To evaluate our approach, we first collected a new data set that incorporates two robotic platforms, three scenarios, and different scene complexities. Second, we introduce measures for the effectiveness of active overt attention mechanisms in terms of saliency cumulation and required head motion. This way, we are able to objectively demonstrate the effectiveness of the proposed multicriterial focus of attention selection.
Benjamin Kühn, Boris Schauerte, Kristian Kroschel, Rainer Stiefelhagen
IROS4
2012 Predicting human gaze using quaternion DCT image signature saliency and face detection
abstract
We combine and extend the previous work on DCT-based image signatures and face detection to determine the visual saliency. To this end, we transfer the scalar definition of image signatures to quaternion images and thus introduce a novel saliency method using quaternion type-II DCT image signatures. Furthermore, we use MCT-based face detection to model the important influence of faces on the visual saliency using rotated elliptical Gaussian weight functions and evaluate several integration schemes. In order to demonstrate the performance of the proposed methods, we evaluate our approach on the Bruce-Tsotsos (Toronto) [2] and Cerf (FIFA) [3] benchmark eye-tracking data sets. Additionally, we present evaluation results on the Bruce-Tsotsos data set of the most important spectral saliency approaches. We achieve state-of-the-art results in terms of the well-established area under curve (AUC) measure on the Bruce-Tsotsos data set and come close to the ideal AUC on the Cerf data set - with less than one millisecond to calculate the bottom-up QDCT saliency map.
Boris Schauerte, Rainer Stiefelhagen
WACV2
2012 Best of Automatic Face and Gesture Recognition 2011
Rainer Stiefelhagen, Marian Stewart Bartlett, Kevin W. Bowyer
Image Vis. Comput.1
2011 Evaluation of local features for person re-identification in image sequences
abstract
In this paper we present a comparative study of local features for the task of person (re) identification. A combination of state of the art interest point detectors and descriptors is evaluated. The experiments are performed on a novel dataset which we make publicly available for future research in this area. The results indicate that there are significant differences between the evaluated descriptors, with GLOH and SIFT outperforming both Shape Context and SURF descriptors. The evaluated interest point descriptors perform equally well, with a slight advantage for the Hessian-Laplace detector. The Harris-Affine and Hessian-Affine affine invariant region detectors do not provide any performance advantage and therefore do not justify their additional computational expense.
Martin Bäuml, Rainer Stiefelhagen
AVSS2
2011 AVSS 2011 demo session: Interactive person-retrieval in a distributed camera network
abstract
Summary form only given. Two fundamental pillars of Software Engineering practice are formalism and structure. Formalism allows engineers to reason rigorously about the system in hand; structure allows them to understand its purposes and behaviours. In the constructive activity of system development structure must therefore take precedence. The central role of formalism is to check and verify — or, where necessary, correct — the products of more informal modes of thought. In this talk these ideas are explored in the context of an illustrative system. The large structure of the system functionality is discussed, together with the nature of the components of that structure. Informal criteria of functional simplicity are presented. The inescapable mismatch between an intelligible functional structure and implementable software architecture is exposed. The role of formalism in these concerns is suggested.
Martin Bäuml, Rainer Stiefelhagen
AVSS2
2011 Part-based clothing segmentation for person retrieval
abstract
Recent advances have shown that clothing appearance provides important features for person re-identification and retrieval in surveillance and multimedia data. However, the regions from which such features are extracted are usually only very crudely segmented, due to the difficulty of segmenting highly articulated entities such as persons. In order to overcome the problem of unconstrained poses, we propose a segmentation approach based on a large number of part detectors. Our approach is able to separately segment a person's upper and lower clothing regions, taking into account the person's body pose. We evaluate our approach on the task of character retrieval on a new challenging data set and present promising results.
Martin Bäuml, Rainer Stiefelhagen
AVSS3
2011 Boosting Pseudo Census Transform Features for Face Alignment
abstract
Face alignment using deformable face model has attracted broad interest in recent years for its wide range of applications in facial analysis. Previous work has shown that discriminative deformable models have better generalization capacity compared to generative models [8, 9]. In this paper, we present a new discriminative face model based on boosting pseudo census transform features. This feature is considered to be less sensitive to illumination changes, which yields a more robust alignment algorithm. The alignment is based on maximizing the scores of boosted strong classifier, which indicate whether the current alignment is a correct or incorrect one. The proposed approach has been evaluated extensively on several databases. The experimental results show that our approach generalizes better on unseen data compared to the Haar feature-based approach. Moreover, its training procedure is much faster due to the low dimensionality of the configuration space of the proposed feature.
Hua Gao, Hazim Kemal Ekenel, Mika Fischer, Rainer Stiefelhagen
BMVC4
2011 Tue-SeA Real-Time Speech Command Detector for a Smart Control Room
abstract
In this work we present an online ASR system that is able to discriminate voice commands directed to an operationable screen from irrelevant speech segments. For classification of the sound segments we explored several features that are based on prosody as well as properties generated during the decoding process. For a vocabulary of 259 words and more than 10k possible commands, our realtime Verbal Command Detector managed to detect 88.3% of the commands in our evaluation data while maintaining a low False Positive Rate (FPR) of 1.5%. On an evaluation task using an episode of Star Trek, our system was able to detect 91.2% of all commands with a FPR of 1.8% with only minor adjustments. The system is part of and used in the Smart Control Room at the Fraunhofer IOSB in Karlsruhe [1], an experimental smart environment that uses multiple input modalities for crisis response.
Daniel Reich, Felix Putze, Dominic Heger, Joris IJsselmuiden, Rainer Stiefelhagen, Tanja Schultz
INTERSPEECH5
2011 Combined intention, activity, and motion recognition for a humanoid household robot
abstract
In this paper, a multi-level approach to intention, activity, and motion recognition for a humanoid robot is proposed. Our system processes images from a monocular camera and combines this information with domain knowledge. The recognition works on-line and in real-time, it is independent of the test person, but limited to predefined view-points. Main contributions of this paper are the extensible, multi-level modeling of the robot's vision system, the efficient activity and motion recognition, and the asynchronous information fusion based on generic processing of mid-level recognition results. The complementarity of the activity and motion recognition renders the approach robust against misclassifications. Experimental results on a real-world data set of complex kitchen tasks, e.g., Prepare Cereals or Lay Table, prove the performance and robustness of the multi-level recognition approach.
Dirk Gehrig, Peter Krauthausen, Lukas Rybok, Hilde Kuehne, Uwe D. Hanebeck, Tanja Schultz, Rainer Stiefelhagen
IROS7
2011 Multimodal saliency-based attention for object-based scene analysis
abstract
Multimodal attention is a key requirement for humanoid robots in order to navigate in complex environments and act as social, cognitive human partners. To this end, robots have to incorporate attention mechanisms that focus the processing on the potentially most relevant stimuli while controlling the sensor orientation to improve the perception of these stimuli. In this paper, we present our implementation of audio-visual saliency-based attention that we integrated in a system for knowledge-driven audio-visual scene analysis and object-based world modeling. For this purpose, we introduce a novel isophote-based method for proto-object segmentation of saliency maps, a surprise-based auditory saliency definition, and a parametric 3-D model for multimodal saliency fusion. The applicability of the proposed system is demonstrated in a series of experiments.
Boris Schauerte, Benjamin Kühn, Kristian Kroschel, Rainer Stiefelhagen
IROS4
2011 Person re-identification in TV series using robust face recognition and user feedback
Mika Fischer, Hazim Kemal Ekenel, Rainer Stiefelhagen
Multim. Tools Appl.3
2010 Multi-pose Face Recognition for Person Retrieval in Camera Networks
abstract
In this paper, we study the use of facial appearance features for the re-identification of persons using distributed camera networks in a realistic surveillance scenario. In contrast to features commonly used for person reidentification, such as whole body appearance, facial features offer the advantage of remaining stable over much larger intervals of time. The challenge in using faces for such applications, apart from low captured face resolutions, is that their appearance across camera sightings is largely influenced by lighting and viewing pose. Here, a number of techniques to address these problems are presented and evaluated on a database of surveillance-type recordings. A system for online capture and interactive retrieval is presented that allows to search for sightings of particular persons in the video database. Evaluation results are presented on surveillance data recorded with four cameras over several days. A mean average precision of 0.60 was achieved for inter-camera retrieval using just a single track as query set, and up to 0.86 after relevance feedback by an operator.
Martin Bäuml, Keni Bernardin, Mika Fischer, Hazim Kemal Ekenel, Rainer Stiefelhagen
AVSS5
2010 Automatic Frequency Band Selection for Illumination Robust Face Recognition
abstract
Varying illumination conditions cause a dramatic change in facial appearance that leads to a significant drop in face recognition algorithms' performance. In this paper, to overcome this problem, we utilize an automatic frequency band selection scheme. The proposed approach is incorporated to a local appearance-based face recognition algorithm, which employs discrete cosine transform (DCT) for processing local facial regions. From the extracted DCT coefficients, the approach determines to the ones that should be used for classification. Extensive experiments conducted on the extended Yale face database B have shown that benefiting from frequency information provides robust face recognition under changing illumination conditions.
Hazim Kemal Ekenel, Rainer Stiefelhagen
ICPR2
2010 Multi-resolution Local Appearance-Based Face Verification
abstract
Facial analysis based on local regions/blocks usually outperforms holistic approaches because it is less sensitive to local deformations and occlusions. Moreover, modeling local features enables us to avoid the problem of high dimensionality of feature space. In this paper, we model the local face blocks with Gabor features and project them into a discriminant identity space. The similarity score of a face pair is determined by fusion of the local classifiers. To acquire complementary information in different scales of face images, we integrate the local decisions from various image resolutions. The proposed multi-resolution block based face verification system is evaluated on the experiment 4 of Face Recognition Grand Challenge (FRGC) version 2.0. We obtained 92.5% verification [email protected]% FAR, which is the highest performance reported on this experiment so far in the literature.
Hua Gao, Hazim Kemal Ekenel, Mika Fischer, Rainer Stiefelhagen
ICPR4
2010 Multi-view Based Estimation of Human Upper-Body Orientation
abstract
The knowledge about the body orientation of humans can improve speed and performance of many service components of a smart-room. Since many of such components run in parallel, an estimator to acquire this knowledge needs a very low computational complexity. In this paper we address these two points with a fast and efficient algorithm using the smart-room's multiple camera output. The estimation is based on silhouette information only and is performed for each camera view separately. The single view results are fused within a Bayesian filter framework. We evaluate our system on a subset of videos from the CLEAR 2007 dataset and achieve an average correct classification rate of 87.8%, while the estimation itself just takes 12 ms when four cameras are used.
Lukas Rybok, Michael Voit, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICPR4
2010 Interactive person-retrieval in TV series and distributed surveillance video
abstract
Tracking and identifying persons in videos are important building blocks in many applications. For browsing of multimedia data or interactive investigation of surveillance footage it is not even necessary to uniquely identify a person. Rather it often suffices to find occurrences of a person indicated by the user with an exemplary image sequence. We present two systems in which the search for a specific person can be initiated by a sample image sequence and then be further refined by interactive feedback by the operator. In the first system, episodes of TV series have been processed offline and can be searched for occurrences of the different characters. The second system tracks people online in multiple cameras and makes the sequences immediately searchable from a central station
Martin Bäuml, Mika Fischer, Keni Bernardin, Hazim Kemal Ekenel, Rainer Stiefelhagen
ACM Multimedia5
2010 A video-based door monitoring system using local appearance-based face models
Hazim Kemal Ekenel, Johannes Stallkamp, Rainer Stiefelhagen
Comput. Vis. Image Underst.3
2009 Open-Set Face Recognition-Based Visitor Interface System
Hazim Kemal Ekenel, Lorant Szasz-Toth, Rainer Stiefelhagen
ICVS3
2009 A System for Probabilistic Joint 3D Head Tracking and Pose Estimation in Low-Resolution, Multi-view Environments
Michael Voit, Rainer Stiefelhagen
ICVS2
2009 Multimodal identity tracking in a smart room
Keni Bernardin, Hazim Kemal Ekenel, Rainer Stiefelhagen
Pers. Ubiquitous Comput.3
2009 Multimodal identity tracking in a smart room
Keni Bernardin, Hazim Kemal Ekenel, Rainer Stiefelhagen
Pers. Ubiquitous Comput.3
2008 Dynamic Integration of Generalized Cues for Person Tracking
Kai Nickel, Rainer Stiefelhagen
ECCV (4)2
2008 Face recognition for smart interactions
abstract
In this paper, face recognition systems that have been developed for smart interactions at the interACT Research Center is presented. The face recognition efforts at the interACT Research Center consist of development of a fast and robust face recognition algorithm and fully automatic face recognition systems that can be deployed for real-life smart interaction applications. The face recognition algorithm is based on appearances of local facial regions that are represented with discrete cosine transform coefficients. Many fully automatic face recognition systems have been developed based on this algorithm. Among these systems two of the portable ones will be shown as interactive demos. Moreover, demo videos will be shown for the other systems.
Hazim Kemal Ekenel, Mika Fischer, Hua Gao, Lorant Szasz-Toth, Rainer Stiefelhagen
FG5
2008 Tracking identities and attention in smart environments - contributions and progress in the CHIL project
abstract
To provide intelligent services in a smart environments it is necessary to acquire information about the room, the people in it and their interactions. This includes, for example, the number of people, their identities, locations, postures, body and head orientations, among others. This paper gives an overview of the perceptual technology evaluations that were conducted in the CHIL project, specifically those held in the CLEAR 2006 and 2007 evaluation workshops. We then summarize the main achievements and lessons learnt in the project in the areas of person tracking, person identification and head pose estimation, all of which are critical perception components in order to build perceptive smart environments.
Rainer Stiefelhagen, Keni Bernardin, Hazim Kemal Ekenel, Michael Voit
FG1
2008 Deducing the visual focus of attention from head pose estimation in dynamic multi-view meeting scenarios
abstract
This paper presents our work on recognizing the visual focus of attention during dynamic meeting scenarios. We collected a new dataset of meetings, in which acting participants were to follow a predefined script of events, to enforce focus shifts of the remaining, unaware meeting members. Including the whole room, all in all, a total of 35 potential focus targets were annotated, of which some were moved or introduced spontaneously during the meeting. On this dynamic dataset, we present a new approach to deduce the visual focus by means of head orientation as a first clue and show, that our system recognizes the correct visual target in over 57% of all frames, compared to 47% when mapping head pose to the first-best intersecting focus target directly.
Michael Voit, Rainer Stiefelhagen
ICMI2
2008 Data Collection for the CHIL CLEAR 2007 Evaluation Campaign
Nicolas Moreau, Djamel Mostefa, Rainer Stiefelhagen, Susanne Burger, Khalid Choukri
LREC3
2008 Probabilistic integration of sparse audio-visual cues for identity tracking
abstract
In the context of smart environments, the ability to track and identify persons is a key factor, determining the scope and flexibility of analytical components or intelligent services that can be provided. While some amount of work has been done concerning the camera-based tracking of multiple users in a variety of scenarios, technologies for acoustic and visual identification, such as face or voice ID, are unfortunately still subjected to severe limitations when distantly placed sensors have to be used. Because of this, reliable cues for identification can be hard to obtain without user cooperation, especially when multiple users are involved.
Keni Bernardin, Rainer Stiefelhagen, Alex Waibel
ACM Multimedia2
2008 A context-aware virtual secretary in a smart office environment
abstract
A lot of the communication at the workplace - via the phone as well as face-to-face - occurs in inappropriate contexts, disturbing meetings and conversations, invading personal and corporate privacy, and more broadly breaking social norms. This is because both, callers and visitors in front of closed office doors, face the same problem: they can only guess the other person's current availability for a conversation.
Maria Danninger, Rainer Stiefelhagen
ACM Multimedia2
2007 Automatic Person Detection and Tracking using Fuzzy Controlled Active Cameras
abstract
This paper presents an automatic system for the monitoring of indoor environments using pan-tilt-zoomable cameras. A combination of Haar-feature classifier-based detection and color histogram filtering is used to achieve reliable initialization of person tracks even in the presence of camera movement. A combination of adaptive color and KLT feature trackers for face and upper body allows for robust tracking and track recovery in the presence of occlusion or interference. The continuous recomputation of camera parameters, coupled with a fuzzy controlling scheme allow for smooth tracking of moving targets as well as acquisition of stable facial close ups, similar to the natural behavior of a human cameraman. The system is tested on a series of natural indoor monitoring scenarios and shows a high degree of naturalness, flexibility and robustness.
Keni Bernardin, Florian van de Camp, Rainer Stiefelhagen
CVPR3
2007 Multi-modal Person Identification in a Smart Environment
abstract
In this paper, we present a detailed analysis of multimodal fusion for person identification in a smart environment. The multi-modal system consists of a video-based face recognition system and a speaker identification system. We investigated different score normalization, modality weighting and modality combination schemes during the fusion of the individual modalities. We introduced two new modality weighting schemes, namely, the cumulative ratio of correct matches (CRCM) and distance-to-second-closest (DT2ND) measures. In addition, we also assessed the effects of the well-known score normalization and classifier combination methods on the identification performance. Experimental results obtained on the CLEAR 2007 evaluation corpus, which contains audio-visual recordings from different smart rooms, show that CRCM-based modality weighting improves the correct identification rates significantly.
Hazim Kemal Ekenel, Mika Fischer, Qin Jin, Rainer Stiefelhagen
CVPR4
2007 State Synchronous Modeling on Phone Boundary for Audio Visual Speech Recognition and Application to Muti-View Face Images
abstract
Visual speech cues are known to improve the performance of automatic speech recognition (ASR). However, many researchers have used speaker's frontal pose mainly. We therefore introduce a new database for large vocabulary audio visual automatic speech recognition (AV-ASR), which contains not only frontal face images but also face images taken from different angles (multi-view face images). Another contribution of this paper is to present a new algorithm which can model audio and visual characteristics between phones. Finally we conducted large vocabulary continuous speech recognition experiments on the new database using the new algorithm. Experimental results show that the proposed AV-ASR system achieved high accuracy even if there are mismatches of the views between training and test data.
Ken'ichi Kumatani, Rainer Stiefelhagen
ICASSP (4)2
2007 Video-based Face Recognition on Real-World Data
abstract
In this paper, we present the classification sub-system of a real-time video-based face identification system which recognizes people entering through the door of a laboratory. Since the subjects are not asked to cooperate with the system but are allowed to behave naturally, this application scenario poses many challenges. Continuous, uncontrolled variations of facial appearance due to illumination, pose, expression, and occlusion need to be handled to allow for successful recognition. Faces are classified by a local appearance-based face recognition algorithm. The obtained confidence scores from each classification are progressively combined to provide the identity estimate of the entire sequence. We introduce three different measures to weight the contribution of each individual frame to the overall classification decision. They are distance- to-model (DTM), distance-to-second-closest (DT2ND), and their combination. Both a k-nearest neighbor approach and a set of Gaussian mixtures are evaluated to produce individual frame scores. We have conducted closed set and open set identification experiments on a database of 41 subjects. The experimental results show that the proposed system is able to reach high correct recognition rates in a difficult scenario.
Johannes Stallkamp, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICCV3
2007 Face Recognition for Smart Interactions
abstract
In this paper an overview of face recognition research activities at the interACT Research Center is given. The face recognition efforts at the interACT Research Center consist of development of a fast and robust face recognition algorithm and fully automatic face recognition systems that can be deployed for real-life smart interaction applications. The face recognition algorithm is based on appearances of local facial regions that are represented with discrete cosine transform coefficients. Three fully automatic face recognition systems have been developed that are based on this algorithm. The first one is the "door monitoring system" that observes the entrance of a room and identifies the subjects while they are entering the room. The second one is the "portable face recognition system" that aims at environment-free face recognition and recognizes the user of a machine. The third system, "3D face recognition system", performs fully automatic face recognition on 3D range data.
Hazim Kemal Ekenel, Johannes Stallkamp, Hua Gao, Mika Fischer, Rainer Stiefelhagen
ICME5
2007 Audio-visual multi-person tracking and identification for smart environments
abstract
This paper presents a novel system for the automatic and unobtrusive tracking and identification of multiple persons in an indoor environment. Information from several fixed cameras is fused in a particle filter framework to simultaneously track multiple occupants. A set of steerable fuzzy-controlled pan-tilt-zoom cameras serves to smoothly track persons of interest and opportunistically capture facial close-ups for face identification. In parallel, speech segmentation, sound source localization and speaker identification are performed using several far-field microphones and arrays. The information coming asynchronously and sporadically from several sources, such as track updates and spatio-temporally localized visual and acoustic identification cues, is fused at higher level to gradually refine the global scene model and increase the system's confidence in the set of recognized identities. The system has been trained on a small set of users' faces and/or voices and showed good performance in natural meeting scenarios at quickly acquiring their identities and complementing the ID information missing in single modalities.
Keni Bernardin, Rainer Stiefelhagen
ACM Multimedia2
2007 Visual recognition of pointing gestures for human-robot interaction
Kai Nickel, Rainer Stiefelhagen
Image Vis. Comput.2
2007 3-D Face Recognition Using Local Appearance-Based Models
abstract
In this paper, we present a local appearance-based approach for 3-D face recognition. In the proposed algorithm, we first register the 3-D point clouds to provide a dense correspondence between faces. Afterwards, we analyze two mapping techniques—the closest-point mapping and the ray-casting mapping, to construct depth images from the corresponding well-registered point clouds. The depth images that are obtained are then divided into local regions where the discrete cosine transformation is performed to extract local information. The local features are combined at the feature level for classification. Experimental results on the FRGC version 2.0 face database show that the proposed algorithm performs superior to the well-known face recognition algorithms.
Hazim Kemal Ekenel, Hua Gao, Rainer Stiefelhagen
IEEE Trans. Inf. Forensics Secur.3
2007 Enabling Multimodal Human-Robot Interaction for the Karlsruhe Humanoid Robot
abstract
In this paper, we present our work in building technologies for natural multimodal human-robot interaction. We present our systems for spontaneous speech recognition, multimodal dialogue processing, and visual perception of a user, which includes localization, tracking, and identification of the user, recognition of pointing gestures, as well as the recognition of a person's head orientation. Each of the components is described in the paper and experimental results are presented. We also present several experiments on multimodal human-robot interaction, such as interaction using speech and gestures, the automatic determination of the addressee during human-human-robot interaction, as well on interactive learning of dialogue strategies. The work and the components presented here constitute the core building blocks for audiovisual perception of humans and multimodal human-robot interaction used for the humanoid robot developed within the German research project (Sonderforschungsbereich) on humanoid cooperative robots.
Rainer Stiefelhagen, Hazim Kemal Ekenel, Christian Fügen, Petra Gieselmann, Hartwig Holzapfel, Florian Kraft, Kai Nickel, Michael Voit, Alex Waibel
IEEE Trans. Robotics1
2006 Tracking of the Articulated Upper Body on Multi-View Stereo Image Sequences
abstract
We propose a novel method for tracking an articulated model in a 3D-point cloud. The tracking problem is formulated as the registration of two point sets, one of them parameterised by the model’s state vector and the other acquired from a 3D-sensor system. Finding the correct parameter vector is posed as a linear estimation problem, which is solved by means of a scaled unscented Kalman filter. Our method draws on concepts from the widely used iterative closest point registration algorithm (ICP), basing the measurement model on point correspondences established between the synthesised model point cloud and the measured 3D-data. We apply the algorithm to kinematically track a model of the human upper body on a point cloud obtained through stereo image processing from one or more stereo cameras. We determine torso position and orientation as well as joint angles of shoulders and elbows. The algorithm has been successfully tested on thousands of frames of real image data. Challenging sequences of several minutes length where tracked correctly. Complete processing time remains below one second per frame.
Julius Ziegler, Kai Nickel, Rainer Stiefelhagen
CVPR (1)3
2006 MyConnector: analysis of context cues to predict human availability for communication
abstract
In this thriving world of mobile communications, the difficulty of communication is no longer contacting someone, but rather contacting people in a socially appropriate manner. Ideally, senders should have some understanding of a receiver's availability in order to make contact at the right time, in the right contexts, and with the optimal communication medium.We describe the design and implementation of MyConnector, an adaptive and context-aware service designed to facilitate efficient and appropriate communication, based on each party's availability. One of the chief design questions of such a service is to produce technologies with sufficient contextual awareness to decide upon a person's availability for communication. We present results from a pilot study comparing a number of context cues and their predictive power for gauging one's availability.
Maria Danninger, Tobias Kluge, Rainer Stiefelhagen
ICMI3
2006 Tracking head pose and focus of attention with multiple far-field cameras
abstract
In this work we present our recent approach on estimating head orientations and foci of attention of multiple people in a smart room, which is equipped with several cameras to monitor the room. In our approach, we estimate each person's head orientation with respect to the room coordinate system by using all camera views. We implemented a Neural Network to estimate head pose on every single camera view, a Bayes filter is then applied to integrate every estimate into one final, joint hypothesis. Using this scheme, we can track peoples' horizontal head orientations in a full 360° range at almost all positions within the room. The tracked head orientations are then used to determine who is looking at whom, i.e. people's focus of attention. We report experimental results on one meeting video, that was recorded in the smart room.
Michael Voit, Rainer Stiefelhagen
ICMI2
2006 Audio-visual perception of a lecturer in a smart seminar room
Rainer Stiefelhagen, Keni Bernardin, Hazim Kemal Ekenel, John W. McDonough, Kai Nickel, Michael Voit, Matthias Wölfel
Signal Process.1
2005 The connector: facilitating context-aware communication
abstract
We present the Connector, a context-aware service that intelligently connects people. It maintains an awareness of its users' activities, preoccupations and social relationships to mediate a proper connection at the right time between them. In addition to providing users with important contextual cues about the availability of potential callees, the Connector adapts the behavior of the contactee's device automatically in order to avoid inappropriate interruptions.To acquire relevant context information, perceptual components analyze sensor input obtained from a smart mobile phone and --- if available --- from a variety of audio-visual sensors built into a smart meeting room environment. The Connector also uses any available multimodal interface (e.g. a speech interface to the smart phone, steerable camera-projector, targeted loudspeakers) in the smart meeting room, to deliver information to users in the most unobtrusive way possible.
Maria Danninger, G. Flaherty, Keni Bernardin, Hazim Kemal Ekenel, Thilo Köhler, Robert G. Malkin, Rainer Stiefelhagen, Alex Waibel
ICMI7
2005 A joint particle filter for audio-visual speaker tracking
abstract
In this paper, we present a novel approach for tracking a lecturer during the course of his speech. We use features from multiple cameras and microphones, and process them in a joint particle filter framework. The filter performs sampled projections of 3D location hypotheses and scores them using features from both audio and video. On the video side, the features are based on foreground segmentation, multi-view face detection and upper body detection. On the audio side, the time delays of arrival between pairs of microphones are estimated with a generalized cross correlation function. Computationally expensive features are evaluated only at the particles' projected positions in the respective camera images, thus the complexity of the proposed algorithm is low. We evaluated the system on data that was recorded during actual lectures. The results of our experiments were 36 cm average error for video only tracking, 46 cm for audio only, and 31 cm for the combined audio-video system.
Kai Nickel, Tobias Gehrig, Rainer Stiefelhagen, John W. McDonough
ICMI3
2004 Implementation and evaluation of a constraint-based multimodal fusion system for speech and 3D pointing gestures
abstract
This paper presents an architecture for fusion of multimodal input streams for natural interaction with a humanoid robot as well as results from a user study with our system. The presented fusion architecture consists of an application independent parser of input events, and application specific rules. In the presented user study, people could interact with a robot in a kitchen scenario, using speech and gesture input. In the study, we could observe that our fusion approach is very tolerant against falsely detected pointing gestures. This is because we use speech as the main modality and pointing gestures mainly for disambiguation of objects. In the paper we also report about the temporal correlation of speech and gesture events as observed in the user study.
Hartwig Holzapfel, Kai Nickel, Rainer Stiefelhagen
ICMI3
2004 Identifying the addressee in human-human-robot interactions based on head pose and speech
abstract
In this work we investigate the power of acoustic and visual cues, and their combination, to identify the addressee in a human-human-robot interaction. Based on eighteen audio-visual recordings of two human beings and a (simulated) robot we discriminate the interaction of the two humans from the interaction of one human with the robot. The paper compares the result of three approaches. The first approach uses purely acoustic cues to find the addressees. Low level, feature based cues as well as higher-level cues are examined. In the second approach we test whether the human's head pose is a suitable cue. Our results show that visually estimated head pose is a more reliable cue for the identification of the addressee in the human-human-robot interaction. In the third approach we combine the acoustic and visual cues which results in significant improvements.
Michael Katzenmaier, Rainer Stiefelhagen, Tanja Schultz
ICMI2
2004 Natural human-robot interaction using speech, head pose and gestures
abstract
In this paper we present our ongoing work in building technologies for natural multimodal human-robot interaction. We present our systems for spontaneous speech recognition, multimodal dialogue processing and visual perception of a user, which includes the recognition of pointing gestures as well as the recognition of a person's head orientation. Each of the components is described in the paper and experimental results are presented. In order to demonstrate and measure the usefulness of such technologies for human-robot interaction, all components have been integrated on a mobile robot platform and have been used for real-time human-robot interaction in a kitchen scenario.
Rainer Stiefelhagen, Christian Fügen, Petra Gieselmann, Hartwig Holzapfel, Kai Nickel, Alex Waibel
IROS1
2003 SMaRT: the Smart Meeting Room Task at ISL
abstract
As computational and communications systems become increasingly smaller, faster, more powerful, and more integrated, the goal of interactive, integrated meeting support rooms is slowly becoming reality. It is already possible, for instance, to rapidly locate task-related information during a meeting, filter it, and share it with remote users. Unfortunately, the technologies that provide such capabilities are as obstructive as they are useful - they force humans to focus on the tool rather than the task. Thus the veneer of utility often hides the true costs of use, which are longer, less focused human interactions. To address this issue, we present our current research efforts towards SMaRT: the Smart Meeting Room Task. The goal of SMaRT is to provide meeting support services that do not require explicit human-computer interaction. Instead, by monitoring the activities in the meeting room using both video and audio analysis, the room is able to react appropriately to users' needs and allow the users to focus on their own goals.
Alex Waibel, Tanja Schultz, Michael Bett, Matthias Denecke, Robert G. Malkin, Ivica Rogina, Rainer Stiefelhagen, Jie Yang 0001
ICASSP (4)7
2003 Pointing gesture recognition based on 3D-tracking of face, hands and head orientation
abstract
In this paper, we present a system capable of visually detecting pointing gestures and estimating the 3D pointing direction in real-time. In order to acquire input features for gesture recognition, we track the positions of a person's face and hands on image sequences provided by a stereo-camera. Hidden Markov Models (HMMs), trained on different phases of sample pointing gestures, are used to classify the 3D-trajectories in order to detect the occurrence of a gesture. When analyzing sample pointing gestures, we noticed that humans tend to look at the pointing target while performing the gesture. In order to utilize this behavior, we additionally measured head orientation by means of a magnetic sensor in a similar scenario. By using head orientation as an additional feature, we observed significant gains in both recall and precision of pointing gestures. Moreover, the percentage of correctly identified pointing targets improved significantly from 65% to 83%. For estimating the pointing direction, we comparatively used three approaches: 1) The line of sight between head and hand, 2) the forearm orientation, and 3) the head orientation.
Kai Nickel, Rainer Stiefelhagen
ICMI2
2002 Towards Vision-Based 3-D People Tracking in a Smart Room
abstract
This paper presents our work on building a real time distributed system to track 3D locations of people in an indoor environment, such as a smart room, using multiple calibrated cameras. In our system, each camera is connected to a dedicated computer on which foreground regions in the camera image are detected. This is done using an adaptive background model. These detected foreground regions are broadcasted to a tracking agent, which computes believed 3D locations of persons based on the detected image regions. We have implemented both a best-hypothesis heuristic tracking approach as well as a probabilistic multi-hypothesis tracker to find the object tracks from these 3D locations. The two tracking approaches are evaluated on a sequence of two people walking in a conference room recorded with three cameras. The results suggest that the probabilistic tracker shows comparable performance to the heuristic tracker.
Dirk Focken, Rainer Stiefelhagen
ICMI2
2002 Tracking Focus of Attention in Meetings
abstract
The author presents an overview of his work on tracking focus of attention in meeting situations. He has developed a system capable of estimating participants' focus of attention from multiple cues. In the system he employs an omni-directional camera to simultaneously track the faces of participants sitting around a meeting table and uses neural networks to estimate their head poses. In addition, he uses microphones to detect who is speaking. The system predicts participants' focus of attention from acoustic and visual information separately, and then combines the output of the audio- and video-based focus of attention predictors. In addition he reports recent experimental results: In order to determine how well we can predict a subject's focus of attention solely on the basis of his or her head orientation, he has conducted an experiment in which he recorded head and eye orientations of participants in a meeting using special tracking equipment. The results demonstrate that head orientation was a sufficient indicator of the subjects' focus target in 89% of the time. Furthermore he discusses how the neural networks used to estimate head orientation can be adapted to work in new locations and under new illumination conditions.
Rainer Stiefelhagen
ICMI1
2002 Modeling focus of attention for meeting indexing based on multiple cues
abstract
A user's focus of attention plays an important role in human-computer interaction applications, such as a ubiquitous computing environment and intelligent space, where the user's goal and intent have to be continuously monitored. We are interested in modeling people's focus of attention in a meeting situation. We propose to model participants' focus of attention from multiple cues. We have developed a system to estimate participants' focus of attention from gaze directions and sound sources. We employ an omnidirectional camera to simultaneously track participants' faces around a meeting table and use neural networks to estimate their head poses. In addition, we use microphones to detect who is speaking. The system predicts participants' focus of attention from acoustic and visual information separately. The system then combines the output of the audio- and video-based focus of attention predictors. We have evaluated the system using the data from three recorded meetings. The acoustic information has provided 8% relative error reduction on average compared to only using one modality. The focus of attention model can be used as an index for a multimedia meeting record. It can also be used for analyzing a meeting.
Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel
IEEE Trans. Neural Networks1
2000 Simultaneous Tracking of Head Poses in a Panoramic View
abstract
In this paper we present an approach to simultaneously estimate gaze directions of multiple people in the view of a panoramic camera. Human faces are located and tracked using a probabilistic skin-color model and motion detection. Neural networks are used to estimate head poses of the detected faces. With this approach, it is possible to simultaneously track the locations of multiple people around a meeting table and estimate their gaze directions using only a panoramic camera. We have achieved an accuracy of 9 degrees for head pan estimation and 6 degrees for tilt estimation for a multi-user system.
Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel
ICPR1
2000 Towards Unrestricted Lip Reading
abstract
Lip reading provides useful information in speech perception and language understanding, especially when the auditory speech is degraded. However, many current automatic lip reading systems impose some restrictions on users. In this paper, we present our research efforts in the Interactive System Laboratory, towards unrestricted lip reading. We first introduce a top–down approach to automatically track and extract lip regions. This technique makes it possible to acquire visual information in real-time without limiting the user's freedom of movement. We then discuss normalization algorithms to preprocess images for different lightning conditions (global illumination and side illumination). We also compare different visual preprocessing methods such as raw image, Linear Discriminant Analysis (LDA), and Principle Component Analysis (PCA). We demonstrate the feasibility of the proposed methods by the development of a modular system for flexible human–computer interaction via both visual and acoustic speech. The system is based on an extension of the existing state-of-the-art speech recognition system, a modular Multiple State–Time Delayed Neural Network (MS–TDNN) system. We have developed adaptive combination methods at several different levels of the recognition network. The system can automatically track a speaker and extract his/her lip region in real-time. The system has been evaluated under different noisy conditions such as white noise, music, and mechanical noise. The experimental results indicate that the system can achieve up to 55% error reduction using visual information in addition to the acoustic signal.
Uwe Meier, Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel
Int. J. Pattern Recognit. Artif. Intell.2
1999 Modeling focus of attention for meeting indexing
abstract
Visual cues, such as gesturing, looking at each other or monitoring each others facial expressions, play an important role in meetings.Such information can be used for indexing of multimedia meeting recordings.In this paper, we present an approach to detect who is looking at whom during a meeting.Our proposal is to employ Hidden Markov Models to characterize participants' focus of attention by using gaze information as well as knowledge about the number and positions of people present in a meeting.The number and positions of the participants faces are detected in the field of view of a panoramic camera.We use neural networks to estimate the directions of participants' gaze from camera images.We discuss the implementation of the approach in detail including system architecture, data collection, and evaluation.The system has achieved an accuracy rate of up to 93 % in detecting focus of attention on test sequences taken from meetings.We have used focus of attention as an index in a multimedia meeting browser.
Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel
ACM Multimedia (1)1
1998 Visual Tracking for Multimodal Human Computer Interaction
abstract
In this paper, we present visual tracking techniques for multimodal human computer interaction.First, we discuss techniques for tracking human faces in which human skin-color is used as a major feature.An adaptive stochastic model has been developed to characterize the skin-color distributions.Based on the maximum likelihood method, the model parameters can be adapted for different people and different lighting conditions.The feasibility of the model has been demonstrated by the development of a real-time face tracker.The system has achieved a rate of 30-t-frames/second using a low-end workstation with a framegrabber and a camera.We also present a top-down approach for tracking facial features such as eyes, nostrils, and lip comers.These real-time visual tracking techniques have been successfully applied to many applications such as gaze tracking, and lipreading.The face tracker has been combined with a microphone array for extracting speech signal from a specific person.The gaze tracker has been combined with a speech recognizer in a multimodal interface for controlling a panoramic image viewer.
Jie Yang 0001, Rainer Stiefelhagen, Uwe Meier, Alex Waibel
CHI2
1997 Gaze tracking for multimodal human-computer interaction
abstract
This paper discusses the problem of gaze tracking and its applications to multimodal human-computer interaction. The function of a gaze tracking system can be either passive or active. For example, a system can identify user's message target by monitoring the user's gaze, or the user could use his gaze to directly control an application or launch actions. We have developed a real-time gaze tracking system that estimates the 3D position and rotation (pose) of a user's head. We demonstrate the applications of the gaze tracker to human-computer interaction by two examples. The first example shows that gaze tracker can help speech recognition systems by switching language model and grammar based on user's gaze information. The second example illustrates the combination of the gaze tracker and a speech recognizer to view a panorama image.
Rainer Stiefelhagen, Jie Yang 0001
ICASSP1
1997 Real-time lip-tracking for lipreading
abstract
This paper presents a new approach to lip tracking for lipreading.Instead of only tracking features on lips, we propose to track lips along with other facial features such as pupils and nostril.In the new approach, the face is rst located in an image using a stochastic skin-color model, the eyes, lip-corners and nostrils are then located and tracked inside the facial region.The new approach can eectively improve the robustness of lip-tracking and simplify automatic detection and recovery of tracking failure.The feasibility of the proposed approach has been demonstrated by implementation of a lip tracking system.The system has been tested by a database that contains 900 image sequences of dierent speakers spelling words.The system has successfully extract lip regions from the image sequences to obtain training data for the audio-visual speech recognition system.The system has been also applied to extract the lip region in real-time from live video images to obtain the visual input for an audio-visual speech recognition system.On test sequences we have achieved a reduction of the number of frames with tracking failures by a factor of two using detection and prediction of outliers in the set of found features.
Rainer Stiefelhagen, Uwe Meier, Jie Yang 0001
EUROSPEECH1