VLDB 2026 Research / reviewers in the wild / expert
Alina Roitberg
dblp:159/9887
· DBLP profile ↗
35ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0003-4724-9164ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Label Noise using Prompt-Based Hyperbolic Meta-Learning in Open-Set Domain Generalization
Kunyu Peng, Di Wen 0006, M. Saquib Sarfraz, Yufan Chen 0001, Junwei Zheng, David Schneider 0006, Kailun Yang 0001, Alina Roitberg, Rainer Stiefelhagen |
Int. J. Comput. Vis. | 9 |
| 2025 | AttentionLeak: What Does Human Attention Reveal About Information Visualisation?
Malte Sönnichsen, Mayar Elfares, Yao Wang 0018, Ralf Küsters, Alina Roitberg, Andreas Bulling |
ICDAR (4) | 5 |
| 2025 | Exploring Self-supervised Skeleton-based Action Recognition in Occluded EnvironmentsabstractTo integrate action recognition into autonomous robotic systems, it is essential to address challenges such as person occlusions—a common yet often overlooked scenario in existing self-supervised skeleton-based action recognition methods. In this work, we propose IosPSTL, a simple and effective self-supervised learning framework designed to handle occlusions. IosPSTL combines a cluster-agnostic KNN imputer with an Occluded Partial Spatio-Temporal Learning (OPSTL) strategy. First, we pre-train the model on occluded skeleton sequences. Then, we introduce a cluster-agnostic KNN imputer that performs semantic grouping using k-means clustering on sequence embeddings. It imputes missing skeleton data by applying K-Nearest Neighbors in the latent space, leveraging nearby sample representations to restore occluded joints. This imputation generates more complete skeleton sequences, which significantly benefits downstream self-supervised models. To further enhance learning, the OPSTL module incorporates Adaptive Spatial Masking (ASM) to make better use of intact, high-quality skeleton sequences during training. Our method achieves state-of-the-art performance on the occluded versions of the NTU-60 and NTU-120 datasets, demonstrating its robustness and effectiveness under challenging conditions. Code is available at https://github.com/cyfml/OPSTL. Kunyu Peng, Alina Roitberg, David Schneider 0006, Jiaming Zhang 0001, Junwei Zheng, Yufan Chen 0001, Ruiping Liu 0001, Kailun Yang 0001, Rainer Stiefelhagen |
IJCNN | 3 |
| 2025 | AdaptiveClick: Click-Aware Transformer With Adaptive Focal Loss for Interactive Image SegmentationabstractInteractive image segmentation (IIS) has emerged as a promising technique for decreasing annotation time. Substantial progress has been made in pre- and post-processing for IIS, but the critical issue of interaction ambiguity, notably hindering segmentation quality, has been under-researched. To address this, we introduce ADAPTIVE CLICK - a click-aware transformer incorporating an adaptive focal loss (AFL) that tackles annotation inconsistencies with tools for mask- and pixel-level ambiguity resolution. To the best of our knowledge, AdaptiveClick is the first transformer-based, mask-adaptive segmentation framework for IIS. The key ingredient of our method is the click-aware mask-adaptive transformer decoder (CAMD), which enhances the interaction between click and image features. Additionally, AdaptiveClick enables pixel-adaptive differentiation of hard and easy samples in the decision space, independent of their varying distributions. This is primarily achieved by optimizing a generalized AFL with a theoretical guarantee, where two adaptive coefficients control the ratio of gradient values for hard and easy pixels. Our analysis reveals that the commonly used Focal and BCE losses can be considered special cases of the proposed AFL. With a plain ViT backbone, extensive experimental results on nine datasets demonstrate the superiority of AdaptiveClick compared to state-of-the-art methods. The source code is publicly available at https://github.com/lab206/AdaptiveClick. Jiacheng Lin, Kailun Yang 0001, Alina Roitberg, Siyu Li 0002, Zhiyong Li 0001, Shutao Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Navigating Open Set Scenarios for Skeleton-Based Action RecognitionabstractIn real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions poses challenges due to the lack of visual background cues and the distinct sparse structure of body pose sequences. In this paper, we tackle the unexplored Open-Set Skeleton-based Action Recognition (OS-SAR) task and formalize the benchmark on three skeleton-based datasets. We assess the performance of seven established open-set approaches on our task and identify their limits and critical generalization issues when dealing with skeleton information.To address these challenges, we propose a distance-based cross-modality ensemble method that leverages the cross-modal alignment of skeleton joints, bones, and velocities to achieve superior open-set recognition performance. We refer to the key idea as CrossMax - an approach that utilizes a novel cross-modality mean max discrepancy suppression mechanism to align latent spaces during training and a cross-modality distance-based logits refinement method during testing. CrossMax outperforms existing approaches and consistently yields state-of-the-art results across all datasets and backbones. We will release the benchmark, code, and models to the community. Kunyu Peng, Junwei Zheng, Ruiping Liu 0001, David Schneider 0006, Jiaming Zhang 0001, Kailun Yang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg |
AAAI | 10 |
| 2024 | Referring Atomic Video Action Recognition
Kunyu Peng, Jia Fu 0001, Kailun Yang 0001, Di Wen 0006, Yufan Chen 0001, Ruiping Liu 0001, Junwei Zheng, Jiaming Zhang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg |
ECCV (19) | 11 |
| 2024 | Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-SupervisionabstractSelf-supervised representation learning for human action recognition has developed rapidly in recent years. Most of the existing works are based on skeleton data while using a multi-modality setup. These works overlooked the differences in performance among modalities, which led to the propagation of erroneous knowledge between modalities while only three fundamental modalities, i.e., joints, bones, and motions are used, hence no additional modalities are explored.In this work, we first propose an Implicit Knowledge Exchange Module (IKEM) which alleviates the propagation of erroneous knowledge between low-performance modalities. Then, we further propose three new modalities to enrich the complementary information between modalities. Finally, to maintain efficiency when introducing new modalities, we propose a novel teacher-student framework to distill the knowledge from the secondary modalities into the mandatory modalities considering the relationship constrained by anchors, positives, and negatives, named relational cross-modality knowledge distillation. The experimental results demonstrate the effectiveness of our approach, unlocking the efficient use of skeleton-based multi-modality data. Source code will be made publicly available at https://github.com/desehuileng0o0/IKEM. Yiping Wei, Kunyu Peng, Alina Roitberg, Jiaming Zhang 0001, Junwei Zheng, Ruiping Liu 0001, Yufan Chen 0001, Kailun Yang 0001, Rainer Stiefelhagen |
ICASSP | 3 |
| 2024 | SynthAct: Towards Generalizable Human Action Recognition based on Synthetic DataabstractSynthetic data generation is a proven method for augmenting training sets without the need for extensive setups, yet its application in human activity recognition is underexplored. This is particularly crucial for human-robot collaboration in household settings, where data collection is often privacy-sensitive. In this paper, we introduce SynthAct, a synthetic data generation pipeline designed to significantly minimize the reliance on real-world data. Leveraging modern 3D pose estimation techniques, SynthAct can be applied to arbitrary 2D or 3D video action recordings, making it applicable for uncontrolled in-the-field recordings by robotic agents or smarthome monitoring systems. We present two SynthAct datasets: AMARV, a large synthetic collection with over 800k multi-view action clips, and Synthetic Smarthome, mirroring the Toyota Smarthome dataset. SynthAct generates a rich set of data, including RGB videos and depth maps from four synchronized views, 3D body poses, normal maps, segmentation masks and bounding boxes. We validate the efficacy of our datasets through extensive synthetic-to-real experiments on NTU RGB+D and Toyota Smarthome. SynthAct is available on our project page4. David Schneider 0006, Marco Keller, Zeyun Zhong, Kunyu Peng, Alina Roitberg, Jürgen Beyerer, Rainer Stiefelhagen |
ICRA | 5 |
| 2024 | Skeleton-Based Human Action Recognition with Noisy LabelsabstractUnderstanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and annotation of activity sequences is time-consuming and the resulting labels are often noisy. If not effectively addressed, label noise negatively affects the model’s training, resulting in lower recognition quality. Despite its importance, addressing label noise for skeleton-based action recognition has been overlooked so far. In this study, we bridge this gap by implementing a framework that augments well-established skeleton-based human action recognition methods with label-denoising strategies from various research areas to serve as the initial benchmark. Observations reveal that these baselines yield only marginal performance when dealing with sparse skeleton data. Consequently, we introduce a novel methodology, NoiseEraSAR, which integrates global sample selection, co-teaching, and Cross-Modal Mixture-of-Experts (CM-MOE) strategies, aimed at mitigating the adverse impacts of label noise. Our proposed approach demonstrates better performance on the established benchmark, setting new state-of-the-art standards. The source code for this study will be made accessible at https://github.com/xuyizdby/NoiseEraSAR. Kunyu Peng, Di Wen 0006, Ruiping Liu 0001, Junwei Zheng, Yufan Chen 0001, Jiaming Zhang 0001, Alina Roitberg, Kailun Yang 0001, Rainer Stiefelhagen |
IROS | 8 |
| 2024 | Chart4Blind: An Intelligent Interface for Chart Accessibility ConversionabstractIn a world driven by data visualization, ensuring the inclusive accessibility of charts for Blind and Visually Impaired (BVI) individuals remains a significant challenge. Charts are usually presented as raster graphics without textual and visual metadata needed for an equivalent exploration experience for BVI people. Additionally, converting these charts into accessible formats requires considerable effort from sighted individuals. Digitizing charts with metadata extraction is just one aspect of the issue; transforming it into accessible modalities, such as tactile graphics, presents another difficulty. To address these disparities, we propose Chart4Blind, an intelligent user interface that converts bitmap image representations of line charts into universally accessible formats. Chart4Blind achieves this transformation by generating Scalable Vector Graphics (SVG), Comma-Separated Values (CSV), and alternative text exports, all comply with established accessibility standards. Through interviews and a formal user study, we demonstrate that even inexperienced sighted users can make charts accessible in an average of 4 minutes using Chart4Blind, achieving a System Usability Scale rating of 90%. In comparison to existing approaches, Chart4Blind provides a comprehensive solution, generating end-to-end accessible SVGs suitable for assistive technologies such as embossed prints (papers and laser cut), 2D tactile displays, and screen readers. For additional information, including open-source codes and demos, please visit our project page https://moured.github.io/chart4blind/. Omar Moured, Morris Baumgarten-Egemole, Karin Müller 0001, Alina Roitberg, Thorsten Schwarz, Rainer Stiefelhagen |
IUI | 4 |
| 2024 | Towards Video-based Activated Muscle Group Estimation in the WildabstractIn this paper, we tackle the new task of video-based Activated Muscle Group Estimation (AMGE) aiming at identifying active muscle regions during physical activity in the wild.To this intent, we provide the MuscleMap dataset featuring >15𝐾 video clips with 135 different activities and 20 labeled muscle groups.This dataset opens the vistas to multiple video-based applications in sports and rehabilitation medicine under flexible environment constraints.The proposed MuscleMap dataset is constructed with YouTube videos, specifically targeting High-Intensity Interval Training (HIIT) physical exercise in the wild.To make the AMGE model applicable in real-life situations, it is crucial to ensure that the model can generalize well to numerous types of physical activities not present during training and involving new combinations of activated muscles.To achieve this, our benchmark also covers an evaluation setting where the model is exposed to activity types excluded from the training set.Our experiments reveal that the generalizability of existing architectures adapted for the AMGE task remains a challenge.Therefore, we also propose a new approach, TransM 3 E, which employs a multi-modality feature fusion mechanism between both the video transformer model and the skeleton-based graph convolution model with novel cross-modal knowledge distillation executed on multiclassification tokens.The proposed method surpasses all popular video classification models when dealing with both, previously seen and new types of physical activities.The database and code can be found at https://github.com/KPeng9510/MuscleMap. Kunyu Peng, David Schneider 0006, Alina Roitberg, Kailun Yang 0001, Jiaming Zhang 0001, Chen Deng, M. Saquib Sarfraz, Rainer Stiefelhagen |
ACM Multimedia | 3 |
| 2024 | Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain SchedulerabstractIn Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories are present at test time. The challenges of this task arise from the dual need to generalize across diverse domains and accurately quantify category novelty, which is critical for applications in dynamic environments. Recently, meta-learning techniques have demonstrated superior results in OSDG, effectively orchestrating the meta-train and -test tasks by employing varied random categories and predefined domain partition strategies. These approaches prioritize a well-designed training schedule over traditional methods that focus primarily on data augmentation and the enhancement of discriminative feature learning.
The prevailing meta-learning models in OSDG typically utilize a predefined sequential domain scheduler to structure data partitions. However, a crucial aspect that remains inadequately explored is the influence brought by strategies of domain schedulers during training.
In this paper, we observe that an adaptive domain scheduler benefits more in OSDG compared with prefixed sequential and random domain schedulers. We propose the Evidential Bi-Level Hardest Domain Scheduler (EBiL-HaDS) to achieve an adaptive domain scheduler. This method strategically sequences domains by assessing their reliabilities in utilizing a follower network, trained with confidence scores learned in an evidential manner, regularized by max rebiasing discrepancy, and optimized in a bilevel manner. We verify our approach on three OSDG benchmarks, i.e., PACS, DigitsDG, and OfficeHome. The results show that our method substantially improves OSDG performance and achieves more discriminative embeddings for both the seen and unseen categories, underscoring the advantage of a judicious domain scheduler for the generalizability to unseen domains and unseen categories. The source code is publicly available at https://github.com/KPeng9510/EBiL-HaDS. Kunyu Peng, Di Wen 0006, Kailun Yang 0001, Ao Luo, Yufan Chen 0001, Jia Fu 0001, M. Saquib Sarfraz, Alina Roitberg, Rainer Stiefelhagen |
NeurIPS | 8 |
| 2024 | TransKD: Transformer Knowledge Distillation for Efficient Semantic SegmentationabstractSemantic segmentation benchmarks in the realm of autonomous driving are dominated by large pre-trained transformers, yet their widespread adoption is impeded by substantial computational costs and prolonged training durations. To lift this constraint, we look at efficient semantic segmentation from a perspective of comprehensive knowledge distillation and aim to bridge the gap between multi-source knowledge extractions and transformer-specific patch embeddings. We put forward the Transformer-based Knowledge Distillation (TransKD) framework which learns compact student transformers by distilling both feature maps and patch embeddings of large teacher transformers, bypassing the long pre-training process and reducing the FLOPs by >85.0%. Specifically, we propose two fundamental modules to realize feature map distillation and patch embedding distillation, respectively: 1) Cross Selective Fusion (CSF) enables knowledge transfer between cross-stage features via channel attention and feature map distillation within hierarchical transformers; 2) Patch Embedding Alignment (PEA) performs dimensional transformation within the patchifying process to facilitate the patch embedding distillation. Furthermore, we introduce two optimization modules to enhance the patch embedding distillation from different perspectives: 1) Global-Local Context Mixer (GL-Mixer) extracts both global and local information of a representative embedding; 2) Embedding Assistant (EA) acts as an embedding method to seamlessly bridge teacher and student models with the teacher’s number of channels. Experiments on Cityscapes, ACDC, NYUv2, and Pascal VOC2012 datasets show that TransKD outperforms state-of-the-art distillation frameworks and rivals the time-consuming pre-training method. The source code is publicly available athttps://github.com/RuipingL/TransKD. Ruiping Liu 0001, Kailun Yang 0001, Alina Roitberg, Jiaming Zhang 0001, Kunyu Peng, Huayao Liu, Yaonan Wang 0001, Rainer Stiefelhagen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Line Graphics Digitization: A Step Towards Full Automation
Omar Moured, Jiaming Zhang 0001, Alina Roitberg, Thorsten Schwarz, Rainer Stiefelhagen |
ICDAR (5) | 3 |
| 2023 | Quantized Distillation: Optimizing Driver Activity Recognition Models for Resource-Constrained EnvironmentsabstractDeep learning-based models are at the top of most driver observation benchmarks due to their remarkable accuracies but come with a high computational cost, while the resources are often limited in real-world driving scenarios. This paper presents a lightweight framework for resource- efficient driver activity recognition. We enhance 3D MobileNet, a speed-optimized neural architecture for video classification, with two paradigms for improving the trade-off between model accuracy and computational efficiency: knowledge distillation and model quantization. Knowledge distillation prevents large drops in accuracy when reducing the model size by harvesting knowledge from a large teacher model (I3D) via soft labels instead of using the original ground truth. Quantization further drastically reduces the memory and computation requirements by representing the model weights and activations using lower precision integers. Extensive experiments on a public dataset for in-vehicle monitoring during autonomous driving show that our proposed framework leads to an 3- fold reduction in model size and 1.4-fold improvement in inference time compared to an already speed-optimized architecture. Our code is available at https://github.com/calvintanama/qd-driver-activity-reco. Calvin Tanama, Kunyu Peng, Zdravko Marinov, Rainer Stiefelhagen, Alina Roitberg |
IROS | 5 |
| 2023 | Delving Deep Into One-Shot Skeleton-Based Action Recognition With Diverse OcclusionsabstractOcclusions areuniversal disruptions constantly present in the real world. Especially for sparse representations, such as human skeletons, a few occluded points might destroy the geometrical and temporal continuity critically affecting the results. Yet, the research of data-scarce recognition from skeleton sequences, such as one-shot action recognition, does not explicitly consider occlusions despite their everyday pervasiveness. In this work, we explicitly tackle body occlusions forSkeleton-basedOne-shotActionRecognition (SOAR). We mainly consider two occlusion variants: 1) random occlusions and 2) more realistic occlusions caused by diverse everyday objects, which we generate by projecting the existing IKEA 3D furniture models into the camera coordinate system of the 3D skeletons with different geometric parameters, (e.g., rotation and displacement). We leverage the proposed pipeline to blend out portions of skeleton sequences of the three popular action recognition datasets (NTU-120, NTU-60 and Toyota Smart Home) and formalize the first benchmark for SOAR from partially occluded body poses. This is the first benchmark which considers occlusions for data-scarce action recognition. Another key property of our benchmark are the more realistic occlusions generated by everyday objects, as even in standard recognition from 3D skeletons, only randomly missing joints were considered. We re-evaluate existing state-of-the-art frameworks for SOAR in the light of this new task and further introduceTrans4SOAR– a new transformer-based model which leverages three data streams and mixed attention fusion mechanism to alleviate the adverse effects caused by occlusions. While our experiments demonstrate a clear decline in accuracy with missing skeleton portions, this effect is smaller withTrans4SOAR, which outperforms other architectures on all datasets. Although we specifically focus onocclusions,Trans4SOARadditionally yields state-of-the-art in thestandardSOAR without occlusion, surpassing the best published approach by 2.85% on NTU-120. Kunyu Peng, Alina Roitberg, Kailun Yang 0001, Jiaming Zhang 0001, Rainer Stiefelhagen |
IEEE Trans. Multim. | 2 |
| 2022 | Multimodal Generation of Novel Action Appearances for Synthetic-to-Real Recognition of Activities of Daily LivingabstractDomain shifts, such as appearance changes, are a key challenge in real-world applications of activity recognition models, which range from assistive robotics and smart homes to driver observation in intelligent vehicles. For example, while simulations are an excellent way of economical data collection, a Synthetic→Real domain shift leads to > 60% drop in accuracy when recognizing Activities of Daily Living (ADLs). We tackle this challenge and introduce an activity domain generation framework which creates novel ADL appearances (novel domains) from different existing activity modalities (source domains) inferred from video training data. Our frame-work computes human poses, heatmaps of body joints, and optical flow maps and uses them alongside the original RGB videos to learn the essence of source domains in order to generate completely new ADL domains. The model is optimized by maximizing the distance between the existing source appearances and the generated novel appearances while ensuring that the semantics of an activity is preserved through an additional classification loss. While source data multimodality is an important concept in this design, our setup does not rely on multi-sensor setups, (i.e., all source modalities are inferred from a single video only.) The newly created activity domains are then integrated in the training of the ADL classification networks, resulting in models far less susceptible to changes in data distributions. Extensive experiments on the Synthetic→Real benchmark Sims4Action demonstrate the potential of the domain generation paradigm for cross-domain ADL recognition, setting new state-of-the-art results. Our code is publicly available at https://github.com/Zrrr1997/syn2real_DG. Zdravko Marinov, David Schneider 0006, Alina Roitberg, Rainer Stiefelhagen |
IROS | 3 |
| 2022 | TransDARC: Transformer-based Driver Activity Recognition with Latent Space Feature CalibrationabstractTraditional video-based human activity recognition has experienced remarkable progress linked to the rise of deep learning, but this effect was slower as it comes to the downstream task of driver behavior understanding. Understanding the situation inside the vehicle cabin is essential for Advanced Driving Assistant System (ADAS) as it enables identifying distraction, predicting driver's intent and leads to more convenient human-vehicle interaction. At the same time, driver observation systems face substantial obstacles as they need to capture different granularities of driver states, while the complexity of such secondary activities grows with the rising automation and increased driver freedom. Furthermore, a model is rarely deployed under conditions identical to the ones in the training set, as sensor placements and types vary from vehicle to vehicle, constituting a substantial obstacle for real-life deployment of data-driven models. In this work, we present a novel vision-based framework for recognizing secondary driver behaviours based on visual transformers and an additional augmented feature distribution calibration module. This module operates in the latent feature-space enriching and diversifying the training set at feature-level in order to improve generalization to novel data appearances, (e.g., sensor changes) and general feature quality. Our framework consistently leads to better recognition rates, surpassing previous state-of-the-art results of the public Drive&Act benchmark on all granularity levels. Our code will be made publicly available at https://github.com/KPeng9510/TransDARC. Kunyu Peng, Alina Roitberg, Kailun Yang 0001, Jiaming Zhang 0001, Rainer Stiefelhagen |
IROS | 2 |
| 2022 | A Comparative Analysis of Decision-Level Fusion for Multimodal Driver Behaviour UnderstandingabstractVisual recognition inside the vehicle cabin leads to safer driving and more intuitive human-vehicle interaction but such systems face substantial obstacles as they need to capture different granularities of driver behaviour while dealing with highly limited body visibility and changing illumination. Multimodal recognition mitigates a number of such issues: prediction outcomes of different sensors complement each other due to different modality-specific strengths and weaknesses. While several late fusion methods have been considered in previously published frameworks, they constantly feature different architecture backbones and building blocks making it very hard to isolate the role of the chosen late fusion strategy itself.This paper presents an empirical evaluation of different paradigms for decision-level late fusion in video-based driver observation. We compare seven different mechanisms for joining the results of single-modal classifiers which have been both popular, (e.g. score averaging) and not yet considered (e.g. rank-level fusion) in the context of driver observation evaluating them based on different criteria and benchmark settings. This is the first systematic study of strategies for fusing outcomes of multimodal predictors inside the vehicles, conducted with the goal to provide guidance for fusion scheme selection. Alina Roitberg, Kunyu Peng, Zdravko Marinov, Constantin Seibold, David Schneider 0006, Rainer Stiefelhagen |
IV | 1 |
| 2022 | Transfer Beyond the Field of View: Dense Panoramic Semantic Segmentation via Unsupervised Domain AdaptationabstractAutonomous vehicles clearly benefit from the expanded Field of View (FoV) of 360° sensors, but modern semantic segmentation approaches rely heavily on annotated training data which is rarely available forpanoramicimages. We look at this problem from the perspective of domain adaptation and bringpanoramicsemantic segmentation to a setting, where labelled training data originates from a different distribution of conventionalpinholecamera images. To achieve this, we formalize the task of unsupervised domain adaptation for panoramic semantic segmentation and collect DensePass - a novel densely annotated dataset for panoramic segmentation under cross-domain conditions, specifically built to study the Pinhole$\rightarrow$PANORAMIC domain shift and accompanied with pinhole camera training examples obtained from Cityscapes. DensePass covers both, labelled- and unlabelled 360° images, with the labelled data comprising 19 classes which explicitly fit the categories available in the source (i.e.pinhole) domain. Since data-driven models are especially susceptible to changes in data distribution, we introduce P2PDA - a generic framework for Pinhole$\rightarrow$Panoramic semantic segmentation which addresses the challenge of domain divergence with different variants of attention-augmented domain adaptation modules, enabling the transfer in output-, feature-, and feature confidence spaces. P2PDA intertwines uncertainty-aware adaptation using confidence values regulated on-the-fly through attention heads with discrepant predictions. Our framework facilitates context exchange when learning domain correspondences and dramatically improves the adaptation performance of accuracy- and efficiency-focused models. Comprehensive experiments verify that our framework clearly surpasses unsupervised domain adaptation- and specialized panoramic segmentation approaches as well as state-of-the-art semantic segmentation methods. Jiaming Zhang 0001, Chaoxiang Ma, Kailun Yang 0001, Alina Roitberg, Kunyu Peng, Rainer Stiefelhagen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | MASS: Multi-Attentional Semantic Segmentation of LiDAR Data for Dense Top-View UnderstandingabstractAt the heart of all automated driving systems is the ability to sense the surroundings,e.g.,through semantic segmentation of LiDAR sequences, which experienced a remarkable progress due to the release of large datasets such as SemanticKITTI and nuScenes-LidarSeg. While most previous works focus onsparsesegmentation of the LiDAR input,denseoutput masks provide self-driving cars with almost complete environment information. In this paper, we introduce MASS - a Multi-Attentional Semantic Segmentation model specifically built for dense top-view understanding of the driving scenes. Our framework operates on pillar- and occupancy features and comprises three attention-based building blocks: (1) a keypoint-driven graph attention, (2) an LSTM-based attention computed from a vector embedding of the spatial input, and (3) a pillar-based attention, resulting in a dense 360° segmentation mask. With extensive experiments on both, SemanticKITTI and nuScenes-LidarSeg, we quantitatively demonstrate the effectiveness of our model, outperforming the state of the art by 19.0% on SemanticKITTI and reaching 30.4% in mIoU on nuScenes-LidarSeg, where MASS is the first work addressing the dense segmentation task. Furthermore, our multi-attention model is shown to be very effective for 3D object detection validated on the KITTI-3D dataset, showcasing its high generalizability to other tasks related to 3D vision. Kunyu Peng, Juncong Fei, Kailun Yang 0001, Alina Roitberg, Jiaming Zhang 0001, Frank Bieder, Philipp Heidenreich, Christoph Stiller, Rainer Stiefelhagen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Is My Driver Observation Model Overconfident? Input-Guided Calibration Networks for Reliable and Interpretable Confidence EstimatesabstractDriver observation models are rarely deployed under perfect conditions. In practice, illumination, camera placement and type differ from the ones present during training and unforeseen behaviours may occur at any time. While observing the human behind the steering wheel leads to more intuitive human-vehicle-interaction and safer driving, it requires recognition algorithms which do not only predict the correct driver state, but also determine their prediction quality through realistic and interpretable confidence measures. Reliable uncertainty estimates are crucial for building trust and are a serious obstacle for deploying activity recognition networks in real driving systems. In this work, we for the first time examine how well the confidence values of modern driver observation models indeed match the probability of the correct outcome and show that raw neural network-based approaches tend to significantly overestimate their prediction quality. To correct this misalignment between the confidence values and the actual uncertainty, we consider two strategies. First, we enhance two activity recognition models often used for driver observation with temperature scaling – an off-the-shelf method for confidence calibration in image classification. Then, we introduce Calibrated Action Recognition with Input Guidance (CARING) – a novel approach leveraging an additional neural network to learn scaling the confidences depending on the video representation. Extensive experiments on the Drive&Act dataset demonstrate that both strategies drastically improve the quality of model confidences, while our CARING model outperforms both, the original architectures and their temperature scaling enhancement, leading to best uncertainty estimates. Alina Roitberg, Kunyu Peng, David Schneider 0006, Kailun Yang 0001, Marios Koulakis, Manuel Martínez 0001, Rainer Stiefelhagen |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Affect-DML: Context-Aware One-Shot Recognition of Human Affect using Deep Metric LearningabstractHuman affect recognition is a well-established research area with numerous applications, e.g. in psychological care, but existing methods assume that all emotions-of-interest are given a priori as annotated training examples. However, the rising granularity and refinements of the human emotional spectrum through novel psychological theories and the increased consideration of emotions in context brings considerable pressure to data collection and labeling work. In this paper, we conceptualize one-shot recognition of emotions in context - a new problem aimed at recognizing human affect states in finer particle level from a single support sample. To address this challenging task, we follow the deep metric learning paradigm and introduce a multi-modal emotion embedding approach which minimizes the distance of the same-emotion embeddings by leveraging complementary information of human appearance and the semantic scene context obtained through a semantic segmentation network. All streams of our context-aware model are optimized jointly using weighted triplet loss and weighted cross entropy loss. We conduct thorough experiments on both, categorical and numerical emotion recognition tasks of the Emotic dataset adapted to our one-shot recognition problem, revealing that categorizing human affect from a single example is a hard task. Still, all variants of our model clearly outperform the random baseline, while leveraging the semantic scene context consistently improves the learnt representations, setting state-of-the-art results in one-shot emotion recognition. To foster research of more universal representations of human affect states, we will make our benchmark and models publicly available to the community under https://github.com/KPeng9510/Affect-DML. Kunyu Peng, Alina Roitberg, David Schneider 0006, Marios Koulakis, Kailun Yang 0001, Rainer Stiefelhagen |
FG | 2 |
| 2021 | Let's Play for Action: Recognizing Activities of Daily Living by Learning from Life Simulation Video GamesabstractRecognizing Activities of Daily Living (ADL) is a vital process for intelligent assistive robots, but collecting large annotated datasets requires time-consuming temporal labeling and raises privacy concerns, e.g., if the data is collected in a real household. In this work, we explore the concept of constructing training examples for ADL recognition by playing life simulation video games and introduce the SIMS4ACTION dataset created with the popular commercial game THE SIMS 4. We build SIMS4ACTION by specifically executing actions-of-interest in a "top-down" manner, while the gaming circumstances allow us to freely switch between environments, camera angles and subject appearances. While ADL recognition on gaming data is interesting from the theoretical perspective, the key challenge arises from transferring it to the real-world applications, such as smart-homes or assistive robotics. To meet this requirement, SIMS4ACTION is accompanied with a GAMING→REAL benchmark, where the models are evaluated on real videos derived from an existing ADL dataset. We integrate two modern algorithms for video-based activity recognition in our framework, revealing the value of life simulation video games as an inexpensive and far less intrusive source of training data. However, our results also indicate that tasks involving a mixture of gaming and real data are challenging, opening a new research direction. We will make our dataset publicly available at https://github.com/aroitberg/sims4action. Alina Roitberg, David Schneider 0006, Aulia Djamal, Constantin Seibold, Simon Reiß, Rainer Stiefelhagen |
IROS | 1 |
| 2021 | From Driver Talk To Future Action: Vehicle Maneuver Prediction by Learning from Driving Exam DialogsabstractA rapidly growing amount of content posted online inherently holds knowledge about concepts of interest, i.e. driver actions. We leverage methods at the intersection of vision and language to surpass costly annotation and present the first automated framework for anticipating driver intention by learning from recorded driving exam conversations. We query YouTube and collect a dataset of posted mock road tests comprising student-teacher dialogs and video data, which we use for learning to foresee the next maneuver without any additional supervision. However, instructional conversations give us very loose labels, while casual chat results in a high amount of noise. To mitigate this effect, we propose a technique for automatic detection of smalltalk based on the likelihood of spoken words being present in everyday dialogs. While visually recognizing driver's intention by learning from natural dialogs only is a challenging task, learning from less but better data via our smalltalk refinement consistently improves performance. Alina Roitberg, Simon Reiß, Rainer Stiefelhagen |
IV | 1 |
| 2020 | Uncertainty-sensitive Activity Recognition: A Reliability Benchmark and the CARING ModelsabstractBeyond assigning the correct class, an activity recognition model should also be able to determine, how certain it is in its predictions. We present the first study of how well the confidence values of modern action recognition architectures indeed reflect the probability of the correct outcome and propose a learning-based approach for improving it. First, we extend two popular action recognition datasets with a reliability benchmark in form of the expected calibration error and reliability diagrams. Since our evaluation highlights that confidence values of standard action recognition architectures do not represent the uncertainty well, we introduce a new approach which learns to transform the model output into realistic confidence estimates through an additional calibration network. The main idea of our Calibrated Action Recognition with Input Guidance (CARING) model is to learn an optimal scaling parameter depending on the video representation. We compare our model with the native action recognition networks and the temperature scaling approach - a wide spread calibration method utilized in image classification. While temperature scaling alone drastically improves the reliability of the confidence values, our CARING method consistently leads to the best uncertainty estimates in all benchmark settings. Alina Roitberg, Monica-Laura Haurilet, Manuel Martínez 0001, Rainer Stiefelhagen |
ICPR | 1 |
| 2020 | Multi-Task Learning for Calorie Prediction on a Novel Large-Scale Recipe Dataset Enriched with Nutritional InformationabstractA rapidly growing amount of content posted online, such as food recipes, opens doors to new exciting applications at the intersection of vision and language. In this work, we aim to estimate the calorie amount of a meal directly from an image by learning from recipes people have published on the Internet, thus skipping time-consuming manual data annotation. Since there are few large-scale publicly available datasets captured in unconstrained environments, we propose the pic2kcal benchmark comprising 308 000 images from over 70 000 recipes including photographs, ingredients, and instructions. To obtain nutritional information of the ingredients and automatically determine the ground-truth calorie value, we match the items in the recipes with structured information from a food item database. We evaluate various neural networks for regression of the calorie quantity and extend them with the multi-task paradigm. Our learning procedure combines the calorie estimation with prediction of proteins, carbohydrates, and fat amounts as well as a multi-label ingredient classification. Our experiments demonstrate clear benefits of multi-task learning for calorie estimation, surpassing the single-task calorie regression by 9.9%. To encourage further research on this task, we make the code for generating the dataset and the models publicly available. Robin Ruede, Verena Heusser, Lukas Frank, Alina Roitberg, Monica-Laura Haurilet, Rainer Stiefelhagen |
ICPR | 4 |
| 2020 | Deep Classification-driven Domain Adaptation for Cross-Modal Driver Behavior RecognitionabstractWe encounter a wide range of obstacles when integrating computer vision algorithms into applications inside the vehicle cabin, e.g. variations in illumination, sensor-type and -placement. Thus, designing domain-invariant representations is crucial for employing such models in practice. Still, the vast majority of driver activity recognition algorithms are developed under the assumption of a static domain, i.e. an identical distribution of training- and test data. In this work, we aim to bring driver monitoring to a setting, where domain shifts can occur at any time and explore generative models which learn a shared representation space of the source and target domain. First, we formulate the problem of unsupervised domain adaptation for driver activity recognition, where a model trained on labeled examples from the source domain (i.e. color images) is intended to adjust to a different target domain (i.e. infrared images) where only unlabeled data is available during training. To address this problem, we leverage current progress in image-to-image translation and adopt multiple strategies for learning a joint latent space of the source and target distribution and a mapping function to the domain of interest. As our long-term goal is a robust cross-domain classification, we enhance a Variational Auto-Encoder (VAE) for image translation with a classification-driven optimization strategy. Our model for classification-driven domain transfer leads to the best cross-domain recognition results and outperforms a conventional classification approach in color-to-infrared recognition by 13.75%. Simon Reiß, Alina Roitberg, Monica-Laura Haurilet, Rainer Stiefelhagen |
IV | 2 |
| 2020 | Open Set Driver Activity RecognitionabstractA common obstacle for applying computer vision models inside the vehicle cabin is the dynamic nature of the surrounding environment, as unforeseen situations may occur at any time. Driver monitoring has been widely researched in the context of closed set recognition i.e. under the premise that all categories are known a priori. Such restrictions represent a significant bottleneck in real-life, as the driver observation models are intended to handle the uncertainty of an open world. In this work, we aim to introduce the concept of open sets to the area of driver observation, where methods have been evaluated only on a static set of classes in the past. First, we formulate the problem of open set recognition for driver monitoring, where a model is intended to identify behaviors previously unseen by the classifier and present a novel Open-Drive&Act benchmark. We combine current closed set models with multiple strategies for novelty detection adopted from general action classification [1] in a generic open set driver behavior recognition framework. In addition to conventional approaches, we employ the prominent I3D architecture extended with modules for assessing its uncertainty via Monte-Carlo dropout. Our experiments demonstrate clear benefits of uncertainty-sensitive models, while leveraging the uncertainty of all the output neurons in a voting-like fashion leads to the best recognition results. To create an avenue for future work, we make Open-Drive&Act public at www.github.com/aroitberg/open-set-driver-activity-recognition. Alina Roitberg, Chaoxiang Ma, Monica-Laura Haurilet, Rainer Stiefelhagen |
IV | 1 |
| 2019 | It's Not About the Journey; It's About the Destination: Following Soft Paths Under Question-Guidance for Visual ReasoningabstractVisual Reasoning remains a challenging task, as it has to deal with long-range and multi-step object relationships in the scene. We present a new model for Visual Reasoning, aimed at capturing the interplay among individual objects in the image represented as a scene graph. As not all graph components are relevant for the query, we introduce the concept of a question-based visual guide, which constrains the potential solution space by learning an optimal traversal scheme, where the final destination nodes alone are used to produce the answer. We show, that finding relevant semantic structures facilitates generalization to new tasks by introducing a novel problem of knowledge transfer: training on one question type and answering questions from a different domain without any training data. Furthermore, we report state-of-the-art results for Visual Reasoning on multiple query types and diverse image and video datasets. Monica-Laura Haurilet, Alina Roitberg, Rainer Stiefelhagen |
CVPR | 2 |
| 2019 | Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous VehiclesabstractWe introduce the novel domain-specific Drive&Act benchmark for fine-grained categorization of driver behavior. Our dataset features twelve hours and over 9.6 million frames of people engaged in distractive activities during both, manual and automated driving. We capture color, infrared, depth and 3D body pose information from six views and densely label the videos with a hierarchical annotation scheme, resulting in 83 categories. The key challenges of our dataset are: (1) recognition of fine-grained behavior inside the vehicle cabin; (2) multi-modal activity recognition, focusing on diverse data streams; and (3) a cross view recognition benchmark, where a model handles data from an unfamiliar domain, as sensor type and placement in the cabin can change between vehicles. Finally, we provide challenging benchmarks by adopting prominent methods for video- and body pose-based action recognition. Manuel Martin, Alina Roitberg, Monica-Laura Haurilet, Matthias Horne, Simon Reiß, Michael Voit, Rainer Stiefelhagen |
ICCV | 2 |
| 2019 | WiSe - Slide Segmentation in the WildabstractWe address the task of segmenting presentation slides, where the examined page was captured as a live photo during lectures. Slides are important document types used as visual components accompanying presentations in a variety of fields ranging from education to business. However, automatic analysis of presentation slides has not been researched sufficiently, and, so far, only preprocessed images of already digitalized slide documents were considered. We aim to introduce the task of analyzing unconstrained photos of slides taken during lectures and present a novel dataset for Page Segmentation with slides captured in the Wild (WiSe). Our dataset covers pixel-wise annotations of 25 classes on 1300 pages, allowing overlapping regions (i.e., multi-class assignments). To evaluate the performance, we define multiple benchmark metrics and baseline methods for our dataset. We further implement two different deep neural network approaches previously used for segmenting natural images and adopt them for the task. Our evaluation results demonstrate the effectiveness of the deep learning-based methods, surpassing the baseline methods by over 30%. To foster further research of slide analysis in unconstrained photos, we make the WiSe dataset publicly available to the community. Monica-Laura Haurilet, Alina Roitberg, Manuel Martínez 0001, Rainer Stiefelhagen |
ICDAR | 2 |
| 2019 | End-to-end Prediction of Driver Intention using 3D Convolutional Neural NetworksabstractDespite extraordinary progress of Advanced Driver Assistance Systems (ADAS), an alarming number of over 1,2 million people are still fatally injured in traffic accidents every year1. Human error is mostly responsible for such casualties, as by the time the ADAS system has alarmed the driver, it is often too late. We present a vision-based system based on deep neural networks with 3D convolutions and residual learning for anticipating the future maneuver based on driver observation. While previous work focuses on hand-crafted features (e.g. head pose), our model predicts the intention directly from video in an end-to-end fashion. Our architecture consists of three components: a neural network for extraction of optical flow, a 3D residual network for maneuver classification and a Long Short-Term Memory network (LSTM) for handling temporal data of varying length. To evaluate our idea, we conduct thorough experiments on the publicly available Brain4Cars benchmark, which covers both inside and outside views for future maneuver anticipation. Our model is able to predict driver intention with an accuracy of 83,12% and 4,07s before the beginning of the maneuver, outperforming state-of-the-art approaches, while considering the inside view only. Patrick Gebert, Alina Roitberg, Monica-Laura Haurilet, Rainer Stiefelhagen |
IV | 2 |
| 2018 | Informed Democracy: Voting-based Novelty Detection for Action Recognition
Alina Roitberg, Ziad Al-Halah, Rainer Stiefelhagen |
BMVC | 1 |
| 2015 | Multimodal Human Activity Recognition for Industrial Manufacturing Processes in Robotic WorkcellsabstractWe present an approach for monitoring and interpreting human activities based on a novel multimodal vision-based interface, aiming at improving the efficiency of human-robot interaction (HRI) in industrial environments. Multi-modality is an important concept in this design, where we combine inputs from several state-of-the-art sensors to provide a variety of information, e.g. skeleton and fingertip poses. Based on typical industrial workflows, we derived multiple levels of human activity labels, including large-scale activities (e.g. assembly) and simpler sub-activities (e.g. hand gestures), creating a duration- and complexity-based hierarchy. We train supervised generative classifiers for each activity level and combine the output of this stage with a trained Hierarchical Hidden Markov Model (HHMM), which models not only the temporal aspects between the activities on the same level, but also the hierarchical relationships between the levels. Alina Roitberg, Nikhil Somani, Alexander Clifford Perzylo, Markus Rickert 0001, Alois C. Knoll |
ICMI | 1 |