EDBT 2026 Demo / reviewers in the wild / expert
Simon Reiß
dblp:259/5361
· DBLP profile ↗
20ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0003-1953-6211ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Every Component Counts: Rethinking the Measure of Success for Medical Semantic Segmentation in Multi-Instance Segmentation TasksabstractWe present Connected-Component (CC)-Metrics, a novel semantic segmentation evaluation protocol, targeted to align existing semantic segmentation metrics to a multi-instance detection scenario in which each connected component matters. We motivate this setup in the common medical scenario of semantic metastases segmentation in a full-body PET/CT. We show how existing semantic segmentation metrics suffer from a bias towards larger connected components contradicting the clinical assessment of scans in which tumor size and clinical relevance are uncorrelated. To rebalance existing segmentation metrics, we propose to evaluate them on a per-component basis thus giving each tumor the same weight irrespective of its size. To match predictions to ground-truth segments, we employ a proximity-based matching criterion, evaluating common metrics locally at the component of interest. Using this approach, we break free of biases introduced by large metastasis for overlap-based metrics such as Dice or Surface Dice. CC-Metrics also improves distance-based metrics such as Hausdorff Distances which are uninformative for small changes that do not influence the maximum or 95th percentile, and avoids pitfalls introduced by directly combining counting-based metrics with overlap-based metrics as it is done in Panoptic Quality. Alexander Jaus, Constantin Seibold, Simon Reiß, Zdravko Marinov, Zeling Ye, Stefan Krieg 0002, Jens Kleesiek, Rainer Stiefelhagen |
AAAI | 3 |
| 2025 | Is Visual in-Context Learning for Compositional Medical Tasks Within Reach?abstractIn this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks during test time without re-training. Unlike previous approaches, our focus is on training in-context learners to adapt to sequences of tasks, rather than individual tasks. Our goal is to solve complex tasks that involve multiple intermediate steps using a single model, allowing users to define entire vision pipelines flexibly at test time. To achieve this, we first examine the properties and limitations of visual in-context learning architectures, with a particular focus on the role of codebooks. We then introduce a novel method for training in-context learners using a synthetic compositional task generation engine. This engine bootstraps task sequences from arbitrary segmentation datasets, enabling the training of visual in-context learners for compositional tasks. Additionally, we investigate different masking-based training objectives to gather insights into how to train models better for solving complex, compositional tasks. Our exploration not only provides important insights especially for multi-modal medical task sequences but also highlights challenges that need to be addressed. Simon Reiß, Zdravko Marinov, Alexander Jaus, Constantin Seibold, M. Saquib Sarfraz, Erik Rodner, Rainer Stiefelhagen |
ICCV | 1 |
| 2024 | SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic GradingabstractTu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Danni Liu, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao, Fabian Peller-Konrad, Tobias Röddiger, Alexander Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Tu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao 0002, Fabian Tërnava, Tobias Röddiger, Alex Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues |
EMNLP | 6 |
| 2024 | Anatomy-Guided Pathology Segmentation
Alexander Jaus, Constantin Seibold, Simon Reiß, Lukas Heine, Anton Schily, Moon S. Kim 0002, Fin Hendrik Bahnsen, Ken Herrmann, Rainer Stiefelhagen, Jens Kleesiek |
MICCAI (8) | 3 |
| 2024 | Muscles in Time: Learning to Understand Human Motion In-Depth by Simulating Muscle ActivationsabstractExploring the intricate dynamics between muscular and skeletal structures is pivotal for understanding human motion. This domain presents substantial challenges, primarily attributed to the intensive resources required for acquiring ground truth muscle activation data, resulting in a scarcity of datasets.In this work, we address this issue by establishing Muscles in Time (MinT), a large-scale synthetic muscle activation dataset.For the creation of MinT, we enriched existing motion capture datasets by incorporating muscle activation simulations derived from biomechanical human body models using the OpenSim platform, a common framework used in biomechanics and human motion research.Starting from simple pose sequences, our pipeline enables us to extract detailed information about the timing of muscle activations within the human musculoskeletal system.Muscles in Time contains over nine hours of simulation data covering 227 subjects and 402 simulated muscle strands. We demonstrate the utility of this dataset by presenting results on neural network-based muscle activation estimation from human pose sequences with two different sequence-to-sequence architectures. David Schneider 0006, Simon Reiß, Marco Kugler, Alexander Jaus, Kunyu Peng, Susanne Sutschet, M. Saquib Sarfraz, Sven Matthiesen, Rainer Stiefelhagen |
NeurIPS | 2 |
| 2024 | 360BEV: Panoramic Semantic Mapping for Indoor Bird's-Eye ViewabstractSeeing only a tiny part of the whole is not knowing the full circumstance. Bird’s-eye-view (BEV) perception, a process of obtaining allocentric maps from egocentric views, is restricted when using a narrow Field of View (FoV) alone. In this work, mapping from 360° panoramas to BEV semantics, the 360BEV task, is established for the first time to achieve holistic representations of indoor scenes in a top-down view. Instead of relying on narrow-FoV image sequences, a panoramic image with depth information is sufficient to generate a holistic BEV semantic map. To benchmark 360BEV, we present two indoor datasets, 360BEV-Matterport and 360BEV-Stanford, both of which include egocentric panoramic images and semantic segmentation labels, as well as allocentric semantic maps. Besides delving deep into different mapping paradigms, we propose a dedicated solution for panoramic semantic mapping, namely 360Mapper. Through extensive experiments, our methods achieve 44.32% and 45.78% mIoU on both datasets respectively, surpassing previous counterparts with gains of +7.60% and +9.70% in mIoU.1 Zhifeng Teng, Jiaming Zhang 0001, Kailun Yang 0001, Kunyu Peng, Hao Shi 0004, Simon Reiß, Rainer Stiefelhagen |
WACV | 6 |
| 2024 | Behind Every Domain There is a Shift: Adapting Distortion-Aware Vision Transformers for Panoramic Semantic SegmentationabstractIn this paper, we address panoramic semantic segmentation which is under-explored due to two critical challenges: (1) image distortions and object deformations on panoramas; (2) lack of semantic annotations in the$360^\circ$imagery. To tackle these problems, first, we propose the upgraded Transformer for Panoramic Semantic Segmentation, ie, Trans4PASS+, equipped withDeformable Patch Embedding (DPE)andDeformable MLP (DMLPv2)modules for handling object deformations and image distortions whenever (before or after adaptation) and wherever (shallow or deep levels). Second, we enhance theMutual Prototypical Adaptation (MPA)strategy via pseudo-label rectification for unsupervised domain adaptive panoramic segmentation. Third, aside from Pinhole-to-Panoramic (Pin2Pan) adaptation, we create a new dataset (SynPASS) with 9,080 panoramic images, facilitating Synthetic-to-Real (Syn2Real) adaptation scheme in$360^\circ$imagery. Extensive experiments are conducted, which cover indoor and outdoor scenarios, and each of them is investigated withPin2PanandSyn2Realregimens. Trans4PASS+ achieves state-of-the-art performances on four domain adaptive panoramic semantic segmentation benchmarks. Code is available athttps://github.com/jamycheung/Trans4PASS. Jiaming Zhang 0001, Kailun Yang 0001, Hao Shi 0004, Simon Reiß, Kunyu Peng, Chaoxiang Ma, Haodong Fu, Philip Torr 0001, Kaiwei Wang, Rainer Stiefelhagen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Delivering Arbitrary-Modal Semantic SegmentationabstractMultimodal fusion can make semantic segmentation more robust. However, fusing an arbitrary number of modalities remains underexplored. To delve into this problem, we create the Deliver arbitrary-modal segmentation benchmark, covering Depth, LiDAR, multiple Views, Events, and RGB. Aside from this, we provide this dataset in four severe weather conditions as well as five sensor failure cases to exploit modal complementarity and resolve partial outages. To make this possible, we present the arbitrary cross-modal segmentation model CMNEXT. It encompasses a Self-Query Hub (SQ-Hub) designed to extract effective information from any modality for subsequent fusion with the RGB representation and adds only negligible amounts of parameters ~0.1M) per additional modality. On top, to efficiently and flexibly harvest discriminative cues from the auxiliary modalities, we introduce the simple Parallel Pooling Mixer (PPX). With extensive experiments on a total of six benchmarks, our CMNEXT achieves state-of-the-art performance on the Deliver, Kitti-360, MFNet, NYU Depth V2, UrbanLF, and MCubeS datasets, allowing to scale from 1 to 81 modalities. On the freshly collected Deliver, the quad-modal CMNEXT reaches up to 66.30% in mIoU with a +9.10% gain as compared to the mono-modal baseline.11The Deliver dataset and our code will be made publicly available at https://jamycheung.github.io/DELIVER.html. Jiaming Zhang 0001, Ruiping Liu 0001, Hao Shi 0004, Kailun Yang 0001, Simon Reiß, Kunyu Peng, Haodong Fu, Kaiwei Wang, Rainer Stiefelhagen |
CVPR | 5 |
| 2023 | Decoupled Semantic Prototypes enable learning from diverse annotation types for semi-weakly segmentation in expert-driven domainsabstractA vast amount of images and pixel-wise annotations allowed our community to build scalable segmentation solutions for natural domains. However, the transfer to expert-driven domains like microscopy applications or medical healthcare remains difficult as domain experts are a critical factor due to their limited availability for providing pixel-wise annotations. To enable affordable segmentation solutions for such domains, we need training strategies which can simultaneously handle diverse annotation types and are not bound to costly pixel-wise annotations. In this work, we analyze existing training algorithms towards their flexibility for different annotation types and scalability to small annotation regimes. We conduct an extensive evaluation in the challenging domain of organelle segmentation and find that existing semi- and semi-weakly supervised training algorithms are not able to fully exploit diverse annotation types. Driven by our findings, we introduce Decoupled Semantic Prototypes (DSP) as a training method for semantic segmentation which enables learning from annotation types as diverse as image-level-, point-, bounding box-, and pixel-wise annotations and which leads to remarkable accuracy gains over existing solutions for semi-weakly segmentation. Simon Reiß, Constantin Seibold, Alexander Freytag, Erik Rodner, Rainer Stiefelhagen |
CVPR | 1 |
| 2022 | Reference-Guided Pseudo-Label Generation for Medical Semantic SegmentationabstractProducing densely annotated data is a difficult and tedious task for medical imaging applications. To address this problem, we propose a novel approach to generate supervision for semi-supervised semantic segmentation. We argue that visually similar regions between labeled and unlabeled images likely contain the same semantics and therefore should share their label. Following this thought, we use a small number of labeled images as reference material and match pixels in an unlabeled image to the semantic of the best fitting pixel in a reference set. This way, we avoid pitfalls such as confirmation bias, common in purely prediction-based pseudo-labeling. Since our method does not require any architectural changes or accompanying networks, one can easily insert it into existing frameworks. We achieve the same performance as a standard fully supervised model on X-ray anatomy segmentation, albeit using 95% fewer labeled images. Aside from an in-depth analysis of different aspects of our proposed method, we further demonstrate the effectiveness of our reference-guided learning paradigm by comparing our approach against existing methods for retinal fluid segmentation with competitive performance as we improve upon recent work by up to 15% mean IoU. Constantin Seibold, Simon Reiß, Jens Kleesiek, Rainer Stiefelhagen |
AAAI | 2 |
| 2022 | Detailed Annotations of Chest X-Rays via CT Projection for Report Understanding
Constantin Seibold, Simon Reiß, M. Saquib Sarfraz, Matthias A. Fink, Victoria Mayer, Jan Sellner, Moon S. Kim 0002, Klaus H. Maier-Hein, Jens Kleesiek, Rainer Stiefelhagen |
BMVC | 2 |
| 2022 | Bending Reality: Distortion-aware Transformers for Adapting to Panoramic Semantic SegmentationabstractPanoramic images with their 360° directional view encompass exhaustive information about the surrounding space, providing a rich foundation for scene understanding. To unfold this potential in the form of robust panoramic segmentation models, large quantities of expensive, pixel-wise annotations are crucial for success. Such annotations are available, but predominantly for narrow-angle, pinhole-camera images which, off the shelf, serve as sub-optimal resources for training panoramic models. Distortions and the distinct image-feature distribution in 360° panoramas impede the transfer from the annotation-rich pinhole domain and therefore come with a big dent in performance. To get around this domain difference and bring together semantic annotations from pinhole- and 360° surround-visuals, we propose to learn object deformations and panoramic image distortions in the Deformable Patch Embedding (DPE) and Deformable MLP (DMLP) components which blend into our Transformer for PAnoramic Semantic Segmentation (Trans4PASS) model. Finally, we tie together shared semantics in pinhole- and panoramic feature embeddings by generating multi-scale prototype features and aligning them in our Mutual Prototypical Adaptation (MPA) for unsupervised domain adaptation. On the indoor Stanford2D3D dataset, our Trans4PASS with MPA maintains comparable performance to fully-supervised state-of-the-arts, cutting the need for over 1,400 labeled panoramas. On the outdoor DensePASS dataset, we break state-of-the-art by 14.39% mIoU and set the new bar at 56.38%. Jiaming Zhang 0001, Kailun Yang 0001, Chaoxiang Ma, Simon Reiß, Kunyu Peng, Rainer Stiefelhagen |
CVPR | 4 |
| 2022 | Graph-Constrained Contrastive Regularization for Semi-weakly Volumetric Segmentation
Simon Reiß, Constantin Seibold, Alexander Freytag, Erik Rodner, Rainer Stiefelhagen |
ECCV (21) | 1 |
| 2022 | Breaking with Fixed Set Pathology Recognition Through Report-Guided Contrastive Training
Constantin Seibold, Simon Reiß, M. Saquib Sarfraz, Rainer Stiefelhagen, Jens Kleesiek |
MICCAI (5) | 2 |
| 2021 | Capturing Omni-Range Context for Omnidirectional SegmentationabstractConvolutional Networks (ConvNets) excel at semantic segmentation and have become a vital component for perception in autonomous driving. Enabling an all-encompassing view of street-scenes, omnidirectional cameras present themselves as a perfect fit in such systems. Most segmentation models for parsing urban environments operate on common, narrow Field of View (FoV) images. Transferring these models from the domain they were designed for to 360° perception, their performance drops dramatically, e.g., by an absolute 30.0% (mIoU) on established test-beds. To bridge the gap in terms of FoV and structural distribution between the imaging domains, we introduce Efficient Concurrent Attention Networks (ECANets), directly capturing the inherent long-range dependencies in omnidirectional imagery. In addition to the learned attention-based contextual priors that can stretch across 360° images, we upgrade model training by leveraging multi-source and omni-supervised learning, taking advantage of both: Densely labeled and unlabeled data originating from multiple datasets. To foster progress in panoramic image segmentation, we put forward and extensively evaluate models on Wild PAnoramic Semantic Segmentation (WildPASS), a dataset designed to capture diverse scenes from all around the globe. Our novel model, training regimen and multi-source prediction fusion elevate the performance (mIoU) to new state-of-the-art results on the public PASS (60.2%) and the fresh WildPASS (69.0%) benchmarks.1 Kailun Yang 0001, Jiaming Zhang 0001, Simon Reiß, Xinxin Hu, Rainer Stiefelhagen |
CVPR | 3 |
| 2021 | Every Annotation Counts: Multi-Label Deep Supervision for Medical Image SegmentationabstractPixel-wise segmentation is one of the most data and an-notation hungry tasks in our field. Providing representative and accurate annotations is often mission-critical especially for challenging medical applications. In this paper, we propose a semi-weakly supervised segmentation algorithm to overcome this barrier. Our approach is based on a new formulation of deep supervision and student-teacher model and allows for easy integration of different supervision signals. In contrast to previous work, we show that care has to be taken how deep supervision is integrated in lower layers and we present multi-label deep supervision as the most important secret ingredient for success. With our novel training regime for segmentation that flexibly makes use of images that are either fully labeled, marked with bounding boxes, just global labels, or not at all, we are able to cut the requirement for expensive labels by 94.22% – narrowing the gap to the best fully supervised baseline to only 5% mean IoU. Our approach is validated by extensive experiments on retinal fluid segmentation and we provide an in-depth analysis of the anticipated effect each annotation type can have in boosting segmentation performance. Simon Reiß, Constantin Seibold, Alexander Freytag, Erik Rodner, Rainer Stiefelhagen |
CVPR | 1 |
| 2021 | Let's Play for Action: Recognizing Activities of Daily Living by Learning from Life Simulation Video GamesabstractRecognizing Activities of Daily Living (ADL) is a vital process for intelligent assistive robots, but collecting large annotated datasets requires time-consuming temporal labeling and raises privacy concerns, e.g., if the data is collected in a real household. In this work, we explore the concept of constructing training examples for ADL recognition by playing life simulation video games and introduce the SIMS4ACTION dataset created with the popular commercial game THE SIMS 4. We build SIMS4ACTION by specifically executing actions-of-interest in a "top-down" manner, while the gaming circumstances allow us to freely switch between environments, camera angles and subject appearances. While ADL recognition on gaming data is interesting from the theoretical perspective, the key challenge arises from transferring it to the real-world applications, such as smart-homes or assistive robotics. To meet this requirement, SIMS4ACTION is accompanied with a GAMING→REAL benchmark, where the models are evaluated on real videos derived from an existing ADL dataset. We integrate two modern algorithms for video-based activity recognition in our framework, revealing the value of life simulation video games as an inexpensive and far less intrusive source of training data. However, our results also indicate that tasks involving a mixture of gaming and real data are challenging, opening a new research direction. We will make our dataset publicly available at https://github.com/aroitberg/sims4action. Alina Roitberg, David Schneider 0006, Aulia Djamal, Constantin Seibold, Simon Reiß, Rainer Stiefelhagen |
IROS | 5 |
| 2021 | From Driver Talk To Future Action: Vehicle Maneuver Prediction by Learning from Driving Exam DialogsabstractA rapidly growing amount of content posted online inherently holds knowledge about concepts of interest, i.e. driver actions. We leverage methods at the intersection of vision and language to surpass costly annotation and present the first automated framework for anticipating driver intention by learning from recorded driving exam conversations. We query YouTube and collect a dataset of posted mock road tests comprising student-teacher dialogs and video data, which we use for learning to foresee the next maneuver without any additional supervision. However, instructional conversations give us very loose labels, while casual chat results in a high amount of noise. To mitigate this effect, we propose a technique for automatic detection of smalltalk based on the likelihood of spoken words being present in everyday dialogs. While visually recognizing driver's intention by learning from natural dialogs only is a challenging task, learning from less but better data via our smalltalk refinement consistently improves performance. Alina Roitberg, Simon Reiß, Rainer Stiefelhagen |
IV | 2 |
| 2020 | Deep Classification-driven Domain Adaptation for Cross-Modal Driver Behavior RecognitionabstractWe encounter a wide range of obstacles when integrating computer vision algorithms into applications inside the vehicle cabin, e.g. variations in illumination, sensor-type and -placement. Thus, designing domain-invariant representations is crucial for employing such models in practice. Still, the vast majority of driver activity recognition algorithms are developed under the assumption of a static domain, i.e. an identical distribution of training- and test data. In this work, we aim to bring driver monitoring to a setting, where domain shifts can occur at any time and explore generative models which learn a shared representation space of the source and target domain. First, we formulate the problem of unsupervised domain adaptation for driver activity recognition, where a model trained on labeled examples from the source domain (i.e. color images) is intended to adjust to a different target domain (i.e. infrared images) where only unlabeled data is available during training. To address this problem, we leverage current progress in image-to-image translation and adopt multiple strategies for learning a joint latent space of the source and target distribution and a mapping function to the domain of interest. As our long-term goal is a robust cross-domain classification, we enhance a Variational Auto-Encoder (VAE) for image translation with a classification-driven optimization strategy. Our model for classification-driven domain transfer leads to the best cross-domain recognition results and outperforms a conventional classification approach in color-to-infrared recognition by 13.75%. Simon Reiß, Alina Roitberg, Monica-Laura Haurilet, Rainer Stiefelhagen |
IV | 1 |
| 2019 | Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous VehiclesabstractWe introduce the novel domain-specific Drive&Act benchmark for fine-grained categorization of driver behavior. Our dataset features twelve hours and over 9.6 million frames of people engaged in distractive activities during both, manual and automated driving. We capture color, infrared, depth and 3D body pose information from six views and densely label the videos with a hierarchical annotation scheme, resulting in 83 categories. The key challenges of our dataset are: (1) recognition of fine-grained behavior inside the vehicle cabin; (2) multi-modal activity recognition, focusing on diverse data streams; and (3) a cross view recognition benchmark, where a model handles data from an unfamiliar domain, as sensor type and placement in the cabin can change between vehicles. Finally, we provide challenging benchmarks by adopting prominent methods for video- and body pose-based action recognition. Manuel Martin, Alina Roitberg, Monica-Laura Haurilet, Matthias Horne, Simon Reiß, Michael Voit, Rainer Stiefelhagen |
ICCV | 5 |