Benjamin Kiefer

dblp:284/3037 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FAM-HRI: Foundation-Model Assisted Multimodal Human-Robot Interaction Combining Gaze and Speech
abstract
Effective Human-Robot Interaction (HRI) is crucial for enhancing accessibility and usability in real-world robotics applications. However, existing solutions often rely on gesture-only or language-only commands, making interaction inefficient and ambiguous, particularly for users with physical impairments. In this paper, we introduce FAM-HRI, an efficient multimodal framework for HRI that integrates language and gaze inputs via foundation models. By leveraging lightweight Meta ARIA glasses, our system captures real-time multimodal signals and utilizes large language models (LLMs) to fuse user intention with scene context, enabling intuitive and precise robot manipulation. Our method accurately determines the gaze fixation time interval, reducing noise caused by the gaze dynamic nature. Experimental evaluations demonstrate that FAM-HRI achieves a high success rate in task execution while maintaining a low interaction time, providing a practical solution for individuals with limited physical mobility or motor impairments. To support the community, we have released our system design, algorithms, and solutions at https://github.com/laiyuzhi/FAM-HRI.
Yuzhi Lai, Shenghai Yuan 0001, Peizheng Li, Benjamin Kiefer, Tianchen Deng, Andreas Zell
IEEE Trans Autom. Sci. Eng.5
2024 Robust Single-Cam Surround View Object Detection and Localization Using Memory Maps
Yitong Quan, Benjamin Kiefer, Martin Messmer, Charan Ram Akupati, Rainer Graser, Andreas Zell
ICPR (30)2
2024 Real-Time Horizon Locking on Unmanned Surface Vehicles
abstract
The expanding use of automated vision, assistance systems, and augmented reality applications in marine settings calls for reliable and accurate horizon detection and locking. Traditional methods utilizing Inertial Measurement Units (IMU) or feature-based computer vision techniques often yield inconsistent results, particularly when unmanned surface vehicles or boats are subject to high-speed movement or choppy waters. Addressing this, our work introduces a computer vision (CV)-based solution for real-time horizon locking. Employing real-time semantic segmentation, we accurately differentiate between sky, land or water in the frame, enabling computational locking of the horizon’s position. This stable visual reference significantly improves the performance and reliability of on-board systems for autonomous navigation, augmented reality overlays, and multi-object tracking. Supported by a dataset collected under various marine conditions, our method has proven to achieve high accuracy with low computational latency, making it a promising avenue for wide-scale implementation on automated and semi-automated systems.
Benjamin Kiefer, Andreas Zell
IROS1
2023 Fast Region of Interest Proposals on Maritime UAVs
abstract
Unmanned aerial vehicles assist in maritime search and rescue missions by flying over large search areas to autonomously search for objects or people. Reliably detecting objects of interest requires fast models to employ on embedded hardware. Moreover, with increasing distance to the ground station only part of the video data can be transmitted. In this work, we consider the problem of finding meaningful region of interest proposals in a video stream on an embedded GPU. Current object or anomaly detectors are not suitable due to their slow speed, especially on limited hardware and for large image resolutions. Lastly, objects of interest, such as pieces of wreckage, are often not known a priori. Therefore, we propose an end-to-end future frame prediction model running in real-time on embedded GPUs to generate region proposals. We analyze its performance on large-scale maritime data sets and demonstrate its benefits over traditional and modern methods.
Benjamin Kiefer, Andreas Zell
ICRA1
2023 Memory Maps for Video Object Detection and Tracking on UAVs
abstract
This paper introduces a novel approach to video object detection detection and tracking on Unmanned Aerial Vehicles (UAVs). By incorporating metadata, the proposed approach creates a memory map of object locations in actual world coordinates, providing a more robust and interpretable representation of object locations in both, image space and the real world. We use this representation to boost confidences, resulting in improved performance for several temporal computer vision tasks, such as video object detection, short and long-term single and multi-object tracking, and video anomaly detection. These findings confirm the benefits of metadata in enhancing the capabilities of UAVs in the field of temporal computer vision and pave the way for further advancements in this area.
Benjamin Kiefer, Yitong Quan, Andreas Zell
IROS1
2023 HyperPosePDF Hypernetworks Predicting the Probability Distribution on SO(3)
abstract
Pose estimation of objects in images is an essential problem in virtual and augmented reality and robotics. Traditional solutions use depth cameras, which can be expensive, and working solutions require long processing times. This work focuses on the more difficult task when only RGB information is available. To this end, we predict not only the pose of an object but the complete probability density function (pdf) on the rotation manifold. This is the most general way to approach the pose estimation problem and is particularly useful in analysing object symmetries. In this work, we leverage implicit neural representations for the task of pose estimation and show that hypernetworks can be used to predict the rotational pdf. Furthermore, we analyse the Fourier embedding on SO(3) and evaluate the effectiveness of an initial Fourier embedding that proved successful. Our HyperPosePDF outperforms the current SOTA approaches on the SYMSOL dataset.
Timon Höfer, Benjamin Kiefer, Martin Messmer, Andreas Zell
WACV2
2022 Leveraging Synthetic Data in Object Detection on Unmanned Aerial Vehicles
abstract
Acquiring data to train deep learning-based object detectors on Unmanned Aerial Vehicles (UAVs) is expensive, time-consuming and may even be prohibited by law in specific environments. On the other hand, synthetic data is fast and cheap to access. In this work, we explore the potential use of synthetic data in object detection from UAVs across various application environments. For that, we extend the open-source framework DeepGTAV to work for UAV scenarios. We capture various large-scale high-resolution synthetic data sets in several domains to demonstrate their use in real-world object detection from UAVs by analyzing multiple training strategies across several models. Furthermore, we analyze several different data generation and sampling parameters to provide actionable engineering advice for further scientific research. The DeepGTAV framework is available at https://git.io/Jyf5j.
Benjamin Kiefer, David Ott, Andreas Zell
ICPR1
2022 Gaining Scale Invariance in UAV Bird's Eye View Object Detection by Adaptive Resizing
abstract
This work introduces a new preprocessing step for object detection applicable to UAV bird’s eye view imagery, which we call Adaptive Resizing. By design, it helps alleviate the challenges coming with the vast variances in objects’ scales, naturally inherent to UAV data sets. Furthermore, it improves inference speed by two to three times on average. We test this extensively on UAVDT, VisDrone, and on a new data set we captured ourselves and achieve consistent improvements while being considerably faster. Moreover, we show how to apply this method to generic UAV object detection tasks. Additionally, we successfully test our approach on a height transfer task where we train on some interval of altitudes and test on a different one. Furthermore, we introduce a small, fast detector meant for deployment to an embedded GPU. Code is available at https://github.com/cogsys-tuebingen/adaptive_resizer.
Martin Messmer, Benjamin Kiefer, Andreas Zell
ICPR2
2022 SeaDronesSee: A Maritime Benchmark for Detecting Humans in Open Water
abstract
Unmanned Aerial Vehicles (UAVs) are of crucial importance in search and rescue missions in maritime environments due to their flexible and fast operation capabilities. Modern computer vision algorithms are of great interest in aiding such missions. However, they are dependent on large amounts of real-case training data from UAVs, which is only available for traffic scenarios on land. Moreover, current object detection and tracking data sets only provide limited environmental information or none at all, neglecting a valuable source of information. Therefore, this paper introduces a large-scaled visual object detection and tracking benchmark (SeaDronesSee) aiming to bridge the gap from land-based vision systems to sea-based ones. We collect and annotate over 54,000 frames with 400,000 instances captured from various altitudes and viewing angles ranging from 5 to 260 meters and 0 to 90° degrees while providing the respective meta information for altitude, viewing angle and other meta data. We evaluate multiple state-of-the-art computer vision algorithms on this newly established benchmark serving as baselines. We provide an evaluation server where researchers can upload their prediction and compare their results on a central leaderboard1.
Leon Amadeus Varga, Benjamin Kiefer, Martin Messmer, Andreas Zell
WACV2