EDBT 2026 Demo / reviewers in the wild / expert
Timm Linder
dblp:72/8555
· DBLP profile ↗
14ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0001-8532-0262ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021Systems, architecture and hardware · 7 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DINO in the Room: Leveraging 2D Foundation Models for 3D SegmentationabstractVision foundation models (VFMs) trained on large-scale image datasets provide high-quality features that have significantly advanced$2 D$visual recognition. However, their potential in 3D scene segmentation remains largely untapped, despite the common availability of$2 D$images alongside 3D point cloud datasets. While significant research has been dedicated to 2D-3D fusion, recent state-of-the-art 3D methods predominantly focus on 3D data, leaving the integration of VFMs into 3D models underexplored. In this work, we challenge this trend by introducing DITR, a generally applicable approach that extracts$2 D$foundation model features, projects them to 3D, and finally injects them into a 3D point cloud segmentation model. DITR achieves state-of-the-art results on both indoor and outdoor 3D semantic segmentation benchmarks. To enable the use of VFMs even when images are unavailable during inference, we additionally propose to pretrain 3D models by distilling 2D foundation models. By initializing the 3D backbone with knowledge distilled from 2D VFMs, we create a strong basis for downstream 3D segmentation tasks, ultimately boosting performance across various datasets. Karim Knaebel, Kadir Yilmaz, Daan de Geus, Alexander Hermans, David B. Adrian, Timm Linder, Bastian Leibe |
3DV | 6 |
| 2025 | Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye ImagesabstractFisheye cameras offer robots the ability to capture human movements across a wider field of view (FOV) than standard pinhole cameras, making them particularly useful for applications in human-robot interaction and automotive contexts. However, accurately detecting human poses in fisheye images is challenging due to the curved distortions inherent to fisheye optics. While various methods for undistorting fisheye images have been proposed, their effectiveness and limitations for poses that cover a wide FOV has not been systematically evaluated in the context of absolute human pose estimation from monocular fisheye images. To address this gap, we evaluate the impact of pinhole, equidistant and double sphere camera models, as well as cylindrical projection methods, on 3D human pose estimation accuracy. We find that in close-up scenarios, pinhole projection is inadequate, and the optimal projection method varies with the FOV covered by the human pose. The usage of advanced fisheye models like the double sphere model significantly enhances 3D human pose estimation accuracy. We propose a heuristic for selecting the appropriate projection model based on the detection bounding box to enhance prediction quality. Additionally, we introduce and evaluate on our novel FISHnCHIPS dataset, which features 3D human skeleton annotations in fisheye images, including images from unconventional angles, such as extreme close-ups, ground-mounted cameras, and wide-FOV poses, available at: https://www.vision.rwth-aachen.de/fishnchips. Stephanie Käs, Sven Peter, Henrik Thillmann, Anton Burenko, David B. Adrian, Dennis Mack, Timm Linder, Bastian Leibe |
ICRA | 7 |
| 2025 | UPTor: Unified 3D Human Pose Dynamics and Trajectory Prediction for Human-Robot InteractionabstractWe introduce a unified approach to forecast the dynamics of human keypoints along with the motion trajectory based on a short sequence of input poses. While many studies address either full-body pose prediction or motion trajectory prediction, only a few attempt to merge them. We propose a motion transformation technique to simultaneously predict full-body pose and trajectory key-points in a global coordinate frame. We utilize an off-the-shelf 3D human pose estimation module, a graph attention network to encode the skeleton structure, and a compact, non-autoregressive transformer suitable for real-time motion prediction for human-robot interaction and human-aware navigation. We introduce a human navigation dataset “DARKO” with specific focus on navigational activities that are relevant for human-aware mobile robot navigation. We perform extensive evaluation on Human3.6M, CMU-Mocap, and our DARKO dataset. In comparison to prior work, we show that our approach is compact, real-time, and accurate in predicting human navigation motion across all datasets. Result animations, our dataset, and code will be available at https://nisarganc.github.io/UPTor-page/ Nisarga Nilavadi, Andrey Rudenko, Timm Linder |
ICRA | 3 |
| 2025 | How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction?abstractGestures enable non-verbal human-robot communication, especially in noisy environments like agile production. Traditional deep learning-based gesture recognition relies on task-specific architectures using images, videos, or skeletal pose estimates as input. Meanwhile, Vision Foundation Models (VFMs) and Vision Language Models (VLMs) with their strong generalization abilities offer potential to reduce system complexity by replacing dedicated task-specific modules. This study investigates adapting such models for dynamic, full-body gesture recognition, comparing V-JEPA (a state-of-the-art VFM), Gemini Flash 2.0 (a multimodal VLM), and HD-GCN (a top-performing skeleton-based approach). We introduce NUGGET, a dataset tailored for human-robot communication in intralogistics environments, to evaluate the different gesture recognition approaches. In our experiments, HD-GCN achieves best performance, but V-JEPA comes close with a simple, task-specific classification head—thus paving a possible way towards reducing system complexity, by using it as a shared multi-task model. In contrast, Gemini struggles to differentiate gestures based solely on textual descriptions in the zero-shot setting, highlighting the need of further research on suitable input representations for gestures. Stephanie Käs, Anton Burenko, Louis Markert, Onur Alp Culha, Dennis Mack, Timm Linder, Bastian Leibe |
RO-MAN | 6 |
| 2021 | Cross-Modal Analysis of Human Detection for Robotics: An Industrial Case StudyabstractAdvances in sensing and learning algorithms have led to increasingly mature solutions for human detection by robots, particularly in selected use-cases such as pedestrian detection for self-driving cars or close-range person detection in consumer settings. Despite this progress, the simple question which sensor-algorithm combination is best suited for a person detection task at handƒ remains hard to answer. In this paper, we tackle this issue by conducting a systematic cross-modal analysis of sensor-algorithm combinations typically used in robotics. We compare the performance of state-of-the-art person detectors for 2D range data, 3D lidar, and RGB-D data as well as selected combinations thereof in a challenging industrial use-case.We further address the related problems of data scarcity in the industrial target domain, and that recent research on human detection in 3D point clouds has mostly focused on autonomous driving scenarios. To leverage these methodological advances for robotics applications, we utilize a simple, yet effective multi-sensor transfer learning strategy by extending a strong image-based RGB-D detector to provide cross-modal supervision for lidar detectors in the form of weak 3D bounding box labels.Our results show a large variance among the different approaches in terms of detection performance, generalization, frame rates and computational requirements. As our use-case contains difficulties representative for a wide range of service robot applications, we believe that these results point to relevant open challenges for further research and provide valuable support to practitioners for the design of their robot system. Timm Linder, Narunas Vaskevicius, Robert Schirmer, Kai Oliver Arras |
IROS | 1 |
| 2020 | Metric-Scale Truncation-Robust Heatmaps for 3D Human Pose EstimationabstractHeatmap representations have formed the basis of 2D human pose estimation systems for many years, but their generalizations for 3D pose have only recently been considered. This includes 2.5D volumetric heatmaps, whose X and Y axes correspond to image space and the Z axis to metric depth around the subject. To obtain metric-scale predictions, these methods must include a separate, explicit post-processing step to resolve scale ambiguity. Further, they cannot encode body joint positions outside of the image boundaries, leading to incomplete pose estimates in case of image truncation. We address these limitations by proposing metric-scale truncation-robust (MeTRo) volumetric heatmaps, whose dimensions are defined in metric 3D space near the subject, instead of being aligned with image space. We train a fully-convolutional network to estimate such heatmaps from monocular RGB in an end-to-end manner. This reinterpretation of the heatmap dimensions allows us to estimate complete metric-scale poses without test-time knowledge of the focal length or person distance and without relying on anthropometric heuristics in post-processing. Furthermore, as the image space is decoupled from the heatmap space, the network can learn to reason about joints beyond the image boundary. Using ResNet-50 without any additional learned layers, we obtain state-of-the-art results on the Human3.6M and MPI-INF-3DHP benchmarks. As our method is simple and fast, it can become a useful component for real-time top-down multi-person pose estimation systems. We make our code publicly available to facilitate further research. István Sárándi, Timm Linder, Kai Oliver Arras, Bastian Leibe |
FG | 2 |
| 2020 | Accurate detection and 3D localization of humans using a novel YOLO-based RGB-D fusion approach and synthetic training dataabstractWhile 2D object detection has made significant progress, robustly localizing objects in 3D space under presence of occlusion is still an unresolved issue. Our focus in this work is on real-time detection of human 3D centroids in RGB-D data. We propose an image-based detection approach which extends the YOLO v3 architecture with a 3D centroid loss and mid-level feature fusion to exploit complementary information from both modalities. We employ a transfer learning scheme which can benefit from existing large-scale 2D object detection datasets, while at the same time learning end-to-end 3D localization from our highly randomized, diverse synthetic RGB-D dataset with precise 3D groundtruth. We further propose a geometrically more accurate depth-aware crop augmentation for training on RGB-D data, which helps to improve 3D localization accuracy. In experiments on our challenging intralogistics dataset, we achieve state-of-the-art performance even when learning 3D localization just from synthetic data. Timm Linder, Kilian Y. Pfeiffer, Narunas Vaskevicius, Robert Schirmer, Kai Oliver Arras |
ICRA | 1 |
| 2016 | On multi-modal people tracking from mobile platforms in very crowded and dynamic environmentsabstractTracking people is a key technology for robots and intelligent systems in human environments. Many person detectors, filtering methods and data association algorithms for people tracking have been proposed in the past 15+ years in both the robotics and computer vision communities, achieving decent tracking performances from static and mobile platforms in real-world scenarios. However, little effort has been made to compare these methods, analyze their performance using different sensory modalities and study their impact on different performance metrics. In this paper, we propose a fully integrated real-time multi-modal laser/RGB-D people tracking framework for moving platforms in environments like a busy airport terminal. We conduct experiments on two challenging new datasets collected from a first-person perspective, one of them containing very dense crowds of people with up to 30 individuals within close range at the same time. We consider four different, recently proposed tracking methods and study their impact on seven different performance metrics, in both single and multi-modal settings. We extensively discuss our findings, which indicate that more complex data association methods may not always be the better choice, and derive possible future research directions. Timm Linder, Stefan Breuers, Bastian Leibe, Kai Oliver Arras |
ICRA | 1 |
| 2015 | Real-time full-body human gender recognition in (RGB)-D dataabstractUnderstanding social context is an important skill for robots that share a space with humans. In this paper, we address the problem of recognizing gender, a key piece of information when interacting with people and understanding human social relations and rules. Unlike previous work which typically considered faces or frontal body views in image data, we address the problem of recognizing gender in RGB-D data from side and back views as well. We present a large, gender-balanced, annotated, multi-perspective RGB-D dataset with full-body views of over a hundred different persons captured with both the Kinect v1 and Kinect v2 sensor. We then learn and compare several classifiers on the Kinect v2 data using a HOG baseline, two state-of-the-art deep-learning methods, and a recent tessellation-based learning approach. Originally developed for person detection in 3D data, the latter is able to learn the best selection, location and scale of a set of simple point cloud features. We show that for gender recognition, it outperforms the other approaches for both standing and walking people while being very efficient to compute with classification rates up to 150 Hz. Timm Linder, Sven Wehner, Kai Oliver Arras |
ICRA | 1 |
| 2015 | Real-time full-body human attribute classification in RGB-D using a tessellation boosting approachabstractRobots that cooperate and interact with humans require the capacity to detect and track people, analyze their behavior and understand human social relations and rules. A key piece of information for such tasks are human attributes like gender, age, hair or clothing. In this paper, we address the problem of recognizing such attributes in RGB-D data from varying full-body views. To this end, we extend a recent tessellation boosting approach which learns the best selection, location and scale of a set of simple RGB-D features. The approach outperforms the original approach and a HOG baseline for five human attributes including gender, has long hair, has long trousers, has long sleeves and has jacket. Experiments on a multi-perspective RGB-D dataset with full-body views of over a hundred different persons show that the method is able to robustly recognize multiple attributes across different view directions and distances to the sensor with accuracies up to 90%. Our methods runs in real-time, achieving a classification rate of around 300 Hz for a single attribute. Timm Linder, Kai Oliver Arras |
IROS | 1 |
| 2014 | Multi-model hypothesis tracking of groups of people in RGB-D data
Timm Linder, Kai Oliver Arras |
FUSION | 1 |
| 2014 | Removing Motion Blur using Natural Image StatisticsabstractWe tackle deconvolution of motion blur in hand-held consumer photography with a Bayesian framework combing sparse gradient and color priors for regularization. We develop a closed-form optimization utilizing iterated re-weighted least squares (IRLS) with a Gaussian approximation of the regularization priors. The model parameters of the priors can be learned from a set of natural images because they resemble common image statistics. We throughly evaluate and discuss the effect of different regularization factors and make suggestions for reasonable values. Both gradient and color priors are current state-of-the-art. In natural images the magnitude of gradients resembles a kurtotic hyper-Laplacian distribution, and the two-color model exploits the observation that locally any color is a linear approximation between some primary and secondary colors. Our contribution is integrating both priors into a single optimization framework and providing a more detailed derivation of their optimization functions. Our re-implementation reveals different model parameters than previously published, and the effectiveness of the color priors alone are explicitly examined. Finally, we propose a context-adaptive parameterization of the regularization factors in order to avoid over-smoothing the deconvolution result in highly textured areas. Johannes Herwig, Timm Linder, Josef Pauli |
ICPRAM | 2 |
| 2014 | Hybreed: A software framework for developing context-aware hybrid recommender systems
Tim Hussein, Timm Linder, Werner Gaulke, Jürgen Ziegler 0001 |
User Model. User Adapt. Interact. | 2 |
| 2011 | Generating route instructions with varying levels of detailabstractIn this paper, we present a technique for adaptive generation of personalized route instructions based on the driver's knowledge of particular route sections. We evaluated the mechanism with two empirical studies, both attesting significant preference for the adaptively generated presentations over an established online service (Google Maps). Jürgen Ziegler 0001, Tim Hussein, Daniel Münter, Jens Hofmann 0002, Timm Linder |
AutomotiveUI | 5 |