VLDB 2026 Research / reviewers in the wild / expert
Benjamin Risse
dblp:124/2684
· DBLP profile ↗
21ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0001-5691-4029ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 13 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FixationFormer: Direct Utilization of Expert Gaze Trajectories for Chest X-Ray Classification
Daniel Beckmann, Benjamin Risse |
ICPR (3) | 2 |
| 2026 | Dynamic Personality Adaptation in Large Language Models via State Machines
Leon Pielage, Ole Hätscher, Mitja Back, Bernhard Marschall, Benjamin Risse |
ICPR (6) | 5 |
| 2025 | LangOcc: Open Vocabulary Occupancy Estimation via Volume RenderingabstractThe 3D occupancy estimation task has become an important challenge in the area of vision-based autonomous driving recently. However, most existing camera-based methods rely on costly 3D voxel labels or LiDAR scans for training, limiting their practicality and scalability. Moreover, most methods are tied to a predefined set of classes which they can detect. In this work we present a novel approach for open vocabulary occupancy estimation called LangOcc, that is trained only via camera images, and can detect arbitrary semantics via vision-language alignment. In particular, we distill the knowledge of the strong vision-language aligned encoder CLIP into a 3D occupancy model via differentiable volume rendering. Our model estimates vision-language aligned features in a 3D voxel grid using only images. It is trained in a weakly-supervised manner by rendering our estimations back to$2 D$space, where features can easily be aligned with CLIP. This training mechanism automatically supervises the scene geometry, allowing for a straight-forward and powerful training method without any explicit geometry supervision. LangOcc outperforms LiDAR-supervised competitors in open vocabulary occupancy with a mAP of 22.7 by a large margin ($+4.3 \%$), solely relying on vision-based training. We also achieve a mIoU score of 11.84 on the Occ3D-nuScenes dataset, surpassing previous vision-only semantic occupancy estimation methods ($+1.71 \%$), despite not being limited to a specific set of categories. Simon Boeder, Fabian Gigengack, Benjamin Risse |
3DV | 3 |
| 2025 | GaussianFlowOcc: Sparse and Weakly Supervised Occupancy Estimation using Gaussian Splatting and Temporal Flow
Simon Boeder, Fabian Gigengack, Benjamin Risse |
ICCV | 3 |
| 2025 | Interactive High-Quality Skin Lesion Generation using Diffusion Models for VR-based Dermatological Education
Leon Pielage, Paul Schmidle, Bernhard Marschall, Benjamin Risse |
IUI | 4 |
| 2025 | Momentum-SAM: Sharpness Aware Minimization without Computational OverheadabstractThe recently proposed optimization algorithm for deep neural networks Sharpness Aware Minimization (SAM) suggests perturbing parameters before gradient calculation by a gradient ascent step to guide the optimization into parameter space regions of flat loss.
While significant generalization improvements and thus reduction of overfitting could be demonstrated, the computational costs are doubled due to the additionally needed gradient calculation, making SAM unfeasible in case of limited computationally capacities.
Motivated by Nesterov Accelerated Gradient (NAG) we propose Momentum-SAM (MSAM), which perturbs parameters in the direction of the accumulated momentum vector to achieve low sharpness without significant computational overhead or memory demands over SGD or Adam.
We evaluate MSAM in detail and reveal insights on separable mechanisms of NAG, SAM and MSAM regarding training optimization and generalization. Marlon Becker, Frederick Altrock, Benjamin Risse |
NeurIPS | 3 |
| 2025 | OccFlowNet: Occupancy Estimation via Differentiable Rendering and Occupancy FlowabstractSemantic occupancy has recently gained significant traction as a prominent 3D scene representation. However, most existing camera-based methods rely on large and costly datasets with fine-grained 3D voxel labels for training, which limits their practicality and scalability. Furthermore, approaches in this domain lack the modelling of scene dynamics. In this work we present a novel approach to occupancy estimation inspired by neural radiance field (NeRF) using supervision in 2D based on 3D labels provided by LiDAR, that offers a more natural way of supervision than voxel labels. In particular, we employ differentiable volumetric rendering to predict depth and semantic maps and train a 3D network based on supervision in 2D space only. To enhance geometric accuracy and increase the supervisory signal, we introduce temporal rendering of adjacent time steps. Additionally, we introduce occupancy flow as a mechanism to handle dynamic objects in the scene and ensure their temporal consistency. Through extensive experimentation we demonstrate that supervision in 2D with LiDAR can achieve state-of-the-art performance compared to methods using voxel labels, and when combining it with voxel supervision in 3D, temporal rendering and occupancy flow, we outperform all previous occupancy estimation models significantly. We conclude that the proposed rendering supervision and occupancy flow advances occupancy estimation. Simon Boeder, Benjamin Risse |
WACV | 2 |
| 2025 | Investigating Imaging, Annotation and Self-Supervision for the Classification of Continuously Developing Cells in Histological Whole Slide ImagesabstractThe analysis of individual cells is increasingly automated through deep learning techniques. This is particularly relevant for high-resolution whole slide images (WSIs), which can contain thousands of cells, making manual evaluation impractical. This increase in automation, however, requires higher levels of standardisation (with respect to the scanning hardware, settings and staining) and is further aggravated by the dynamics of the underlying cellular processes, rendering unique cell classifications difficult. To address these difficulties we investigated the entire processing pipeline (from imaging over annotation to model training) and study its underlying trade-offs. In particular, we created a new dataset comprising of more than 6, 300 labelled and 500, 000 unlabelled cells scanned using two different scan settings, resulting in fully registered image pairs with varying level of detail and quality. Using these alternative dataset versions we analysed the impact of inter- and intra-variability between three different annotators and addressed the challenge of limited labelled data by comparing the impact of different self-supervised pretraining strategies. Overall, our analyses provide new insights into the dependencies between imaging, annotation, self-supervision and deep learning-based classification, especially in the context of continuously developing cells and demonstrate the beneficial impact of these considerations on the overall classification accuracy. Code is available at https://zivgitlab.uni-muenster.de/cvmls/icdc and the data will be shared upon qualified request due to data privacy laws. Jacqueline Kockwelp, Joachim Wistuba, Sabine Kliesch, Jörg Gromoll, Benjamin Risse |
WACV | 6 |
| 2024 | Learning Proposal Distributions in Simulated Annealing via Template Networks: A Case Study in Nanophotonic Inverse Design
Marlon Becker, Marco Butz, David Lemli, Carsten Schuck, Benjamin Risse |
ICPR (8) | 5 |
| 2024 | Towards a Dynamic Vision Sensor-based Insect Camera TrapabstractThis paper introduces a visual real-time insect monitoring approach capable of detecting and tracking tiny and fast-moving objects in cluttered wildlife conditions using an RGB-DVS stereo-camera system. By building on the intrinsic benefits of event vision data acquisition, we demonstrate that insect presence can be detected at an extremely high temporal rate (on average more than 40 times real-time) while surpassing the spatial and spectral sensitivity of conventional colour-based sensing. Our DVS-based detection and tracking algorithm extracts insect locations over time, and we evaluated our system based on 81104 manually annotated stereo-frames with 34453 insect appearances featuring highly varying scenes and imaging conditions (including clutter, wind-induced motion, etc.). Comparing our algorithm to two state-of-the-art deep learning algorithms reveals superior results in both detection performance and computational speed. Using the DVS as a trigger for the temporally synchronised RGB camera, we are able to correctly identify 73% of images with and without insects which can be increased to 76% with parameters optimised for different scenes. Overall, our study suggests that DVS-based sensing can be used for visual insect monitoring by enabling reliable real-time insect detection in wildlife conditions while significantly reducing the necessity for data storage, manual labour and energy. Eike Gebauer, Pierre Ouvrard, Adrien Sicard, Benjamin Risse |
WACV | 5 |
| 2024 | Solving the Plane-Sphere Ambiguity in Top-Down Structure-from-MotionabstractDrone-based land surveys and tracking applications with a moving camera require three-dimensional reconstructions from videos recorded using a downward facing camera and are usually generated by Structure-from-Motion (SfM) algorithms. Unfortunately, monocular SfM pipelines can fail in the presence of lens distortion due to a critical configuration resulting in a plane-sphere ambiguity which is characterized by severe curvatures of the reconstructions and erroneous relative camera pose estimations. We propose a 4-point minimal solver for the relative pose estimation for two views sharing the same radial distortion parameters (i.e. from the same camera) with a viewing direction perpendicular to the ground plane. To extract 3D reconstructions from continuous videos, the relative pose of pairwise frames is estimated by using the solver with RANSAC and the Sampson error where globally consistent distortion parameters are determined by taking the medial of all values. Moreover, we propose an additional regularizer for the final bundle adjustment to remove any remaining curvature of the reconstruction if necessary. We tested our methods on synthetic and real-world data and our results demonstrate a significant reduction of curvature and more accurate relative pose estimations. Our algorithm can be easily integrated into existing pipelines and is therefore a practical solution to resolve the plane-sphere ambiguity in a variety of top-down SfM applications. Lars Haalck, Benjamin Risse |
WACV | 2 |
| 2024 | Tracking Tiny Insects in Cluttered Natural Environments using Refinable Recurrent Neural NetworksabstractVisual tracking of tiny and low-contrast objects such as insects in cluttered natural environments is a very challenging computer vision task. This is particularly true for machine learning algorithms, which usually require distinct visual foreground features to reliably identify the object of interest. Here, we propose a novel deep learning-based tracking framework capable of detecting tiny and visually camouflaged ants (covering only a few pixels) in complex and dynamic high-resolution videos. In particular, we introduce refinable recurrent Hourglass Networks, which combine color and temporal information to continuously detect insects recorded using a freely moving camera. Moreover, this architecture provides comprehensible heatmaps of positional estimations and a seamless integration of optional user-input to further refine the tracking results if necessary. We evaluated our algorithm on an extremely challenging wildlife ant dataset with a resolution of 1024 × 1024 and report a mean deviation of 19 pixels from the ground truth (object ≈ 30 px) without any user input. By providing only 0.6% manual locations this accuracy can be improved to a mean deviation of 9 pixels. A comparison to a well known deep learning-based single frame detection algorithm (YOLOv7), two state-of-the-art tracking methods (ToMP and KeepTrack), a probabilistic tracking framework and a comprehensive ablation study reveal superior performances in all our experiments. Our tracking framework therefore provides a foundation for challenging tiny singleobject tracking scenarios and a practical and interactive solution for biologists and ecologists. Lars Haalck, Benjamin Risse |
WACV | 3 |
| 2024 | VR-based Competence Training at Scale: Teaching Clinical Skills in the Context of Virtual Brain Death ExaminationabstractTeaching medical practical and soft skills in clinical routines is increasingly difficult, and manikin or actor-based simulations have gained popularity in the last decades. These simulations, however, hardly scale with the demand, are commonly insufficient to train crucial clinical competencies, and cannot portray complex visual and dynamic symptomatologies as required in, for example, brain death examinations. In this paper, we explore the requirements and challenges of integrating a large-scale high-throughput VR setup into a real medical curriculum and describe our approaches and implementation. Therefore we extend and evaluate an interactive virtual reality-based simulation for training brain death diagnostics in a virtual intensive care environment, featuring a fully reactive simulated patient. To enable the required scalability we integrated the simulation into a dedicated hardware and software framework, enabling 12 simultaneous VR trainings which are controlled by a centralized server system. Using this setup we continuously collected feedback on the application's usability and realism from hundreds of students to gain first insights into the applicability of large-scale VR-based learning systems in real course designs. After integrating this feedback, we conducted a controlled curricular study in which we compared the virtual brain death simulation with the classical manikin-based training approach. Our results indicate that the immersive learning experience is perceived to be more realistic and engaging and is overall preferred by the students while also providing the same learning effect as the alternatives. Pascal Kockwelp, Marcel Meyerheim, Dimitar Valkov, Marvin Mergen, Anna Junga, Antonio Krüger, Bernhard Marschall, Markus Holling, Benjamin Risse |
Proc. ACM Hum. Comput. Interact. | 9 |
| 2023 | EyeGuide - From Gaze Data to Instance Segmentation
Jacqueline Kockwelp, Jörg Gromoll, Joachim Wistuba, Benjamin Risse |
BMVC | 4 |
| 2022 | Narrowing Attention in Capsule Networks
Benjamin Risse |
ICPR | 2 |
| 2021 | Embedded Dense Camera Trajectories in Multi-Video Image Mosaics by Geodesic Interpolation-based ReintegrationabstractDense registrations of huge image sets are still challenging due to exhaustive matchings and computationally expensive optimisations. Moreover, the resultant image mosaics often suffer from structural errors such as drift. Here, we propose a novel algorithm to generate global large-scale registrations from thousands of images extracted from multiple videos to derive high-resolution image mosaics which include full frame rate camera trajectories. Our algorithm does not require any initialisations and ensures the effective integration of all available image data by combining efficient and highly parallelised key-frame and loop-closure mechanisms with a novel geodesic interpolation-based reintegration strategy. As a consequence, global refinement can be done in a fraction of iterations compared to traditional optimisation strategies, while effectively avoiding drift and convergence towards inappropriate solutions. We compared our registration strategy with state-of-the-art algorithms and quantitative evaluations revealed millimetre spatial and high angular accuracy. Applicability is demonstrated by registering more than 110,000 frames from multiple scan recordings and provide dense camera trajectories in a globally referenced coordinate system as used for drone-based mappings, ecological studies, object tracking and land surveys. Lars Haalck, Benjamin Risse |
WACV | 2 |
| 2021 | Resolving Colliding Larvae by Fitting ASM to Random Walker-Based Pre-SegmentationsabstractDrosophila melanogaster is an important model organism for research in neuro- and behavioral biology. Automated studies of their locomotion are crucial to link sensory input and neural processing to motor output which has led to numerous vision-based tracking systems. However, most of these approaches share the inability to segment the contours of colliding animals causing identity losses, appearing and disappearing animals, and the absence of posture and motion related measurements during the time of the collision. We present a novel collision resolution algorithm enabling an accurate contour segmentation of multiple touching Drosophila larvae. Our algorithm utilizes an adapted active shape model (ASM) to learn a low dimensional posture space which is fitted to random-walker generated pre-segmentations. We evaluate our collision resolution algorithm using three publicly available datasets and compare it with the current state-of-the-art methods. In addition, we introduce a refined dataset enabling a segmentation evaluation on the level of pixel accuracy. The results demonstrate that our approach outperforms the state-of-the-art approaches in both accuracy and computational time. We will incorporate this algorithm into our widely used tracking program to improve the statistical strength of the behavioral quantification and allow marker-free studies of interacting Drosophila larvae. Ang Bian, Xiaoyi Jiang 0001, Dimitri Berh, Benjamin Risse |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | From skylight input to behavioural output: A computational model of the insect polarised light compassabstractMany insects navigate by integrating the distances and directions travelled on an outward path, allowing direct return to the starting point. Fundamental to the reliability of this process is the use of a neural compass based on external celestial cues. Here we examine how such compass information could be reliably computed by the insect brain, given realistic constraints on the sky polarisation pattern and the insect eye sensor array. By processing the degree of polarisation in different directions for different parts of the sky, our model can directly estimate the solar azimuth and also infer the confidence of the estimate. We introduce a method to correct for tilting of the sensor array, as might be caused by travel over uneven terrain. We also show that the confidence can be used to approximate the change in sun position over time, allowing the compass to remain fixed with respect to 'true north' during long excursions. We demonstrate that the compass is robust to disturbances and can be effectively used as input to an existing neural model of insect path integration. We discuss the plausibility of our model to be mapped to known neural circuits, and to be implemented for robot navigation. Evripidis Gkanias, Benjamin Risse, Michael Mangan, Barbara Webb |
PLoS Comput. Biol. | 2 |
| 2017 | FIMTrack: An open source tracking and locomotion analysis software for small animalsabstractImaging and analyzing the locomotion behavior of small animals such as Drosophila larvae or C. elegans worms has become an integral subject of biological research. In the past we have introduced FIM, a novel imaging system feasible to extract high contrast images. This system in combination with the associated tracking software FIMTrack is already used by many groups all over the world. However, so far there has not been an in-depth discussion of the technical aspects. Here we elaborate on the implementation details of FIMTrack and give an in-depth explanation of the used algorithms. Among others, the software offers several tracking strategies to cover a wide range of different model organisms, locomotion types, and camera properties. Furthermore, the software facilitates stimuli-based analysis in combination with built-in manual tracking and correction functionalities. All features are integrated in an easy-to-use graphical user interface. To demonstrate the potential of FIMTrack we provide an evaluation of its accuracy using manually labeled data. The source code is available under the GNU GPLv3 at https://github.com/i-git/FIMTrack and pre-compiled binaries for Windows and Mac are available at http://fim.uni-muenster.de. Benjamin Risse, Dimitri Berh, Nils Otto, Christian Klämbt, Xiaoyi Jiang 0001 |
PLoS Comput. Biol. | 1 |
| 2013 | Biomedical Imaging: A Computer Vision Perspective
Xiaoyi Jiang 0001, Mohammad Dawood, Fabian Gigengack, Benjamin Risse, Sönke Schmid, Daniel Tenbrinck, Klaus P. Schäfers |
CAIP (1) | 4 |
| 2013 | Stereo and Motion Based 3D High Density Object Tracking
Junli Tao, Benjamin Risse, Xiaoyi Jiang 0001 |
PSIVT | 2 |