Timm Linder

dblp:72/8555 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0001-8532-0262ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021Systems, architecture and hardware · 7 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
abstract
Vision foundation models (VFMs) trained on large-scale image datasets provide high-quality features that have significantly advanced$2 D$visual recognition. However, their potential in 3D scene segmentation remains largely untapped, despite the common availability of$2 D$images alongside 3D point cloud datasets. While significant research has been dedicated to 2D-3D fusion, recent state-of-the-art 3D methods predominantly focus on 3D data, leaving the integration of VFMs into 3D models underexplored. In this work, we challenge this trend by introducing DITR, a generally applicable approach that extracts$2 D$foundation model features, projects them to 3D, and finally injects them into a 3D point cloud segmentation model. DITR achieves state-of-the-art results on both indoor and outdoor 3D semantic segmentation benchmarks. To enable the use of VFMs even when images are unavailable during inference, we additionally propose to pretrain 3D models by distilling 2D foundation models. By initializing the 3D backbone with knowledge distilled from 2D VFMs, we create a strong basis for downstream 3D segmentation tasks, ultimately boosting performance across various datasets.
Karim Knaebel, Kadir Yilmaz, Daan de Geus, Alexander Hermans, David B. Adrian, Timm Linder, Bastian Leibe
3DV6
2025 Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
abstract
Fisheye cameras offer robots the ability to capture human movements across a wider field of view (FOV) than standard pinhole cameras, making them particularly useful for applications in human-robot interaction and automotive contexts. However, accurately detecting human poses in fisheye images is challenging due to the curved distortions inherent to fisheye optics. While various methods for undistorting fisheye images have been proposed, their effectiveness and limitations for poses that cover a wide FOV has not been systematically evaluated in the context of absolute human pose estimation from monocular fisheye images. To address this gap, we evaluate the impact of pinhole, equidistant and double sphere camera models, as well as cylindrical projection methods, on 3D human pose estimation accuracy. We find that in close-up scenarios, pinhole projection is inadequate, and the optimal projection method varies with the FOV covered by the human pose. The usage of advanced fisheye models like the double sphere model significantly enhances 3D human pose estimation accuracy. We propose a heuristic for selecting the appropriate projection model based on the detection bounding box to enhance prediction quality. Additionally, we introduce and evaluate on our novel FISHnCHIPS dataset, which features 3D human skeleton annotations in fisheye images, including images from unconventional angles, such as extreme close-ups, ground-mounted cameras, and wide-FOV poses, available at: https://www.vision.rwth-aachen.de/fishnchips.
Stephanie Käs, Sven Peter, Henrik Thillmann, Anton Burenko, David B. Adrian, Dennis Mack, Timm Linder, Bastian Leibe
ICRA7
2025 UPTor: Unified 3D Human Pose Dynamics and Trajectory Prediction for Human-Robot Interaction
abstract
We introduce a unified approach to forecast the dynamics of human keypoints along with the motion trajectory based on a short sequence of input poses. While many studies address either full-body pose prediction or motion trajectory prediction, only a few attempt to merge them. We propose a motion transformation technique to simultaneously predict full-body pose and trajectory key-points in a global coordinate frame. We utilize an off-the-shelf 3D human pose estimation module, a graph attention network to encode the skeleton structure, and a compact, non-autoregressive transformer suitable for real-time motion prediction for human-robot interaction and human-aware navigation. We introduce a human navigation dataset “DARKO” with specific focus on navigational activities that are relevant for human-aware mobile robot navigation. We perform extensive evaluation on Human3.6M, CMU-Mocap, and our DARKO dataset. In comparison to prior work, we show that our approach is compact, real-time, and accurate in predicting human navigation motion across all datasets. Result animations, our dataset, and code will be available at https://nisarganc.github.io/UPTor-page/
Nisarga Nilavadi, Andrey Rudenko, Timm Linder
ICRA3
2025 How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction?
abstract
Gestures enable non-verbal human-robot communication, especially in noisy environments like agile production. Traditional deep learning-based gesture recognition relies on task-specific architectures using images, videos, or skeletal pose estimates as input. Meanwhile, Vision Foundation Models (VFMs) and Vision Language Models (VLMs) with their strong generalization abilities offer potential to reduce system complexity by replacing dedicated task-specific modules. This study investigates adapting such models for dynamic, full-body gesture recognition, comparing V-JEPA (a state-of-the-art VFM), Gemini Flash 2.0 (a multimodal VLM), and HD-GCN (a top-performing skeleton-based approach). We introduce NUGGET, a dataset tailored for human-robot communication in intralogistics environments, to evaluate the different gesture recognition approaches. In our experiments, HD-GCN achieves best performance, but V-JEPA comes close with a simple, task-specific classification head—thus paving a possible way towards reducing system complexity, by using it as a shared multi-task model. In contrast, Gemini struggles to differentiate gestures based solely on textual descriptions in the zero-shot setting, highlighting the need of further research on suitable input representations for gestures.
Stephanie Käs, Anton Burenko, Louis Markert, Onur Alp Culha, Dennis Mack, Timm Linder, Bastian Leibe
RO-MAN6
2021 Cross-Modal Analysis of Human Detection for Robotics: An Industrial Case Study
abstract
Advances in sensing and learning algorithms have led to increasingly mature solutions for human detection by robots, particularly in selected use-cases such as pedestrian detection for self-driving cars or close-range person detection in consumer settings. Despite this progress, the simple question which sensor-algorithm combination is best suited for a person detection task at handƒ remains hard to answer. In this paper, we tackle this issue by conducting a systematic cross-modal analysis of sensor-algorithm combinations typically used in robotics. We compare the performance of state-of-the-art person detectors for 2D range data, 3D lidar, and RGB-D data as well as selected combinations thereof in a challenging industrial use-case.We further address the related problems of data scarcity in the industrial target domain, and that recent research on human detection in 3D point clouds has mostly focused on autonomous driving scenarios. To leverage these methodological advances for robotics applications, we utilize a simple, yet effective multi-sensor transfer learning strategy by extending a strong image-based RGB-D detector to provide cross-modal supervision for lidar detectors in the form of weak 3D bounding box labels.Our results show a large variance among the different approaches in terms of detection performance, generalization, frame rates and computational requirements. As our use-case contains difficulties representative for a wide range of service robot applications, we believe that these results point to relevant open challenges for further research and provide valuable support to practitioners for the design of their robot system.
Timm Linder, Narunas Vaskevicius, Robert Schirmer, Kai Oliver Arras
IROS1
2020 Metric-Scale Truncation-Robust Heatmaps for 3D Human Pose Estimation
abstract
Heatmap representations have formed the basis of 2D human pose estimation systems for many years, but their generalizations for 3D pose have only recently been considered. This includes 2.5D volumetric heatmaps, whose X and Y axes correspond to image space and the Z axis to metric depth around the subject. To obtain metric-scale predictions, these methods must include a separate, explicit post-processing step to resolve scale ambiguity. Further, they cannot encode body joint positions outside of the image boundaries, leading to incomplete pose estimates in case of image truncation. We address these limitations by proposing metric-scale truncation-robust (MeTRo) volumetric heatmaps, whose dimensions are defined in metric 3D space near the subject, instead of being aligned with image space. We train a fully-convolutional network to estimate such heatmaps from monocular RGB in an end-to-end manner. This reinterpretation of the heatmap dimensions allows us to estimate complete metric-scale poses without test-time knowledge of the focal length or person distance and without relying on anthropometric heuristics in post-processing. Furthermore, as the image space is decoupled from the heatmap space, the network can learn to reason about joints beyond the image boundary. Using ResNet-50 without any additional learned layers, we obtain state-of-the-art results on the Human3.6M and MPI-INF-3DHP benchmarks. As our method is simple and fast, it can become a useful component for real-time top-down multi-person pose estimation systems. We make our code publicly available to facilitate further research.
István Sárándi, Timm Linder, Kai Oliver Arras, Bastian Leibe
FG2
2020 Accurate detection and 3D localization of humans using a novel YOLO-based RGB-D fusion approach and synthetic training data
abstract
While 2D object detection has made significant progress, robustly localizing objects in 3D space under presence of occlusion is still an unresolved issue. Our focus in this work is on real-time detection of human 3D centroids in RGB-D data. We propose an image-based detection approach which extends the YOLO v3 architecture with a 3D centroid loss and mid-level feature fusion to exploit complementary information from both modalities. We employ a transfer learning scheme which can benefit from existing large-scale 2D object detection datasets, while at the same time learning end-to-end 3D localization from our highly randomized, diverse synthetic RGB-D dataset with precise 3D groundtruth. We further propose a geometrically more accurate depth-aware crop augmentation for training on RGB-D data, which helps to improve 3D localization accuracy. In experiments on our challenging intralogistics dataset, we achieve state-of-the-art performance even when learning 3D localization just from synthetic data.
Timm Linder, Kilian Y. Pfeiffer, Narunas Vaskevicius, Robert Schirmer, Kai Oliver Arras
ICRA1
2016 On multi-modal people tracking from mobile platforms in very crowded and dynamic environments
abstract
Tracking people is a key technology for robots and intelligent systems in human environments. Many person detectors, filtering methods and data association algorithms for people tracking have been proposed in the past 15+ years in both the robotics and computer vision communities, achieving decent tracking performances from static and mobile platforms in real-world scenarios. However, little effort has been made to compare these methods, analyze their performance using different sensory modalities and study their impact on different performance metrics. In this paper, we propose a fully integrated real-time multi-modal laser/RGB-D people tracking framework for moving platforms in environments like a busy airport terminal. We conduct experiments on two challenging new datasets collected from a first-person perspective, one of them containing very dense crowds of people with up to 30 individuals within close range at the same time. We consider four different, recently proposed tracking methods and study their impact on seven different performance metrics, in both single and multi-modal settings. We extensively discuss our findings, which indicate that more complex data association methods may not always be the better choice, and derive possible future research directions.
Timm Linder, Stefan Breuers, Bastian Leibe, Kai Oliver Arras
ICRA1
2015 Real-time full-body human gender recognition in (RGB)-D data
abstract
Understanding social context is an important skill for robots that share a space with humans. In this paper, we address the problem of recognizing gender, a key piece of information when interacting with people and understanding human social relations and rules. Unlike previous work which typically considered faces or frontal body views in image data, we address the problem of recognizing gender in RGB-D data from side and back views as well. We present a large, gender-balanced, annotated, multi-perspective RGB-D dataset with full-body views of over a hundred different persons captured with both the Kinect v1 and Kinect v2 sensor. We then learn and compare several classifiers on the Kinect v2 data using a HOG baseline, two state-of-the-art deep-learning methods, and a recent tessellation-based learning approach. Originally developed for person detection in 3D data, the latter is able to learn the best selection, location and scale of a set of simple point cloud features. We show that for gender recognition, it outperforms the other approaches for both standing and walking people while being very efficient to compute with classification rates up to 150 Hz.
Timm Linder, Sven Wehner, Kai Oliver Arras
ICRA1
2015 Real-time full-body human attribute classification in RGB-D using a tessellation boosting approach
abstract
Robots that cooperate and interact with humans require the capacity to detect and track people, analyze their behavior and understand human social relations and rules. A key piece of information for such tasks are human attributes like gender, age, hair or clothing. In this paper, we address the problem of recognizing such attributes in RGB-D data from varying full-body views. To this end, we extend a recent tessellation boosting approach which learns the best selection, location and scale of a set of simple RGB-D features. The approach outperforms the original approach and a HOG baseline for five human attributes including gender, has long hair, has long trousers, has long sleeves and has jacket. Experiments on a multi-perspective RGB-D dataset with full-body views of over a hundred different persons show that the method is able to robustly recognize multiple attributes across different view directions and distances to the sensor with accuracies up to 90%. Our methods runs in real-time, achieving a classification rate of around 300 Hz for a single attribute.
Timm Linder, Kai Oliver Arras
IROS1
2014 Multi-model hypothesis tracking of groups of people in RGB-D data
Timm Linder, Kai Oliver Arras
FUSION1
2014 Removing Motion Blur using Natural Image Statistics
abstract
We tackle deconvolution of motion blur in hand-held consumer photography with a Bayesian framework combing sparse gradient and color priors for regularization. We develop a closed-form optimization utilizing iterated re-weighted least squares (IRLS) with a Gaussian approximation of the regularization priors. The model parameters of the priors can be learned from a set of natural images because they resemble common image statistics. We throughly evaluate and discuss the effect of different regularization factors and make suggestions for reasonable values. Both gradient and color priors are current state-of-the-art. In natural images the magnitude of gradients resembles a kurtotic hyper-Laplacian distribution, and the two-color model exploits the observation that locally any color is a linear approximation between some primary and secondary colors. Our contribution is integrating both priors into a single optimization framework and providing a more detailed derivation of their optimization functions. Our re-implementation reveals different model parameters than previously published, and the effectiveness of the color priors alone are explicitly examined. Finally, we propose a context-adaptive parameterization of the regularization factors in order to avoid over-smoothing the deconvolution result in highly textured areas.
Johannes Herwig, Timm Linder, Josef Pauli
ICPRAM2
2014 Hybreed: A software framework for developing context-aware hybrid recommender systems
Tim Hussein, Timm Linder, Werner Gaulke, Jürgen Ziegler 0001
User Model. User Adapt. Interact.2
2011 Generating route instructions with varying levels of detail
abstract
In this paper, we present a technique for adaptive generation of personalized route instructions based on the driver's knowledge of particular route sections. We evaluated the mechanism with two empirical studies, both attesting significant preference for the adaptively generated presentations over an established online service (Google Maps).
Jürgen Ziegler 0001, Tim Hussein, Daniel Münter, Jens Hofmann 0002, Timm Linder
AutomotiveUI5