Jason R. Rambach

dblp:191/7799 · also Jason Raphael Rambach · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0001-8122-6789ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion
Mahdi Chamseddine, Didier Stricker, Jason R. Rambach
ICPR (1)3
2026 IMKD: Intensity-Aware Multi-Level Knowledge Distillation for Camera-Radar Fusion
abstract
High-performance Radar-Camera 3D object detection can be achieved by leveraging knowledge distillation without using LiDAR at inference time. However, existing distillation methods typically transfer modality-specific features directly to each sensor, which can distort their unique characteristics and degrade their individual strengths. To address this, we introduce IMKD, a radar-camera fusion framework based on multi-level knowledge distillation that preserves each sensor’s intrinsic characteristics while amplifying their complementary strengths. IMKD applies a three-stage, intensity-aware distillation strategy to enrich the fused representation across the architecture: (1) LiDAR-to-Radar intensity-aware feature distillation to enhance radar representations with fine-grained structural cues, (2) LiDAR-to-Fused feature intensity-guided distillation to selectively highlight useful geometry and depth information at the fusion level, fostering complementarity between the modalities rather than forcing them to align, and (3) Camera-Radar intensity-guided fusion mechanism that facilitates effective feature alignment and calibration. Extensive experiments on the nuScenes benchmark show that IMKD reaches 67.0% NDS and 61.0% mAP, outperforming all prior distillation-based radar-camera fusion methods. Our code and models are available at: https://github.com/dfki-av/IMKD/.
Shashank Mishra 0001, Karan Patil, Didier Stricker, Jason R. Rambach
WACV4
2025 JENGA: Object selection and pose estimation for robotic grasping from a stack
abstract
Vision-based robotic object grasping is typically investigated in the context of isolated objects or unstructured object sets in bin picking scenarios. However, there are several settings, such as construction or warehouse automation, where a robot needs to interact with a structured object formation such as a stack. In this context, we define the problem of selecting suitable objects for grasping along with estimating an accurate 6DoF pose of these objects. To address this problem, we propose a camera-IMU based approach that prioritizes unobstructed objects on the higher layers of stacks and introduce a dataset for benchmarking and evaluation, along with a suitable evaluation metric that combines object selection with pose accuracy. Experimental results show that although our method can perform quite well, this is a challenging problem if a completely error-free solution is needed. Finally, we show results from the deployment of our method for a brick-picking application in a construction scenario.
Sai Srinivas Jeevanandam, Sandeep Inuganti, Shreedhar Govil, Didier Stricker, Jason R. Rambach
IROS5
2025 Resolving Symmetry Ambiguity in Correspondence-Based Methods for Instance-Level Object Pose Estimation
abstract
Estimating the 6D pose of an object from a single RGB image is a critical task that becomes additionally challenging when dealing with symmetric objects. Recent approaches typically establish one-to-one correspondences between image pixels and 3D object surface vertices. However, the utilization of one-to-one correspondences introduces ambiguity for symmetric objects. To address this, we propose SymCode, a symmetry-aware surface encoding that encodes the object surface vertices based on one-to-many correspondences, eliminating the problem of one-to-one correspondence ambiguity. We also introduce SymNet, a fast end-to-end network that directly regresses the 6D pose parameters without solving a PnP problem. We demonstrate faster runtime and comparable accuracy achieved by our method on the T-LESS and IC-BIN benchmarks of mostly symmetric objects. The code is available at https://github.com/lyltc1/SymNet.
Yongliang Lin, Yongzhi Su, Sandeep Inuganti, Yan Di, Naeem Ajilforoushan, Hanqing Yang 0002, Yu Zhang 0018, Jason R. Rambach
IEEE Trans. Image Process.8
2024 HiPose: Hierarchical Binary Surface Encoding and Correspondence Pruning for RGB-D 6DoF Object Pose Estimation
abstract
In this work, we present a novel dense-correspondence method for 6DoF object pose estimation from a single RGB-D image. While many existing data-driven methods achieve impressive performance, they tend to be time-consuming due to their reliance on rendering-based refinement approaches. To circumvent this limitation, we present HiPose, which establishes 3D-3D correspondences in a coarse-to-fine manner with a hierarchical binary surface encoding. Unlike previous dense-correspondence methods, we estimate the correspondence surface by employing point-to-surface matching and iteratively constricting the surface until it becomes a correspondence point while gradually removing outliers. Extensive experiments on public benchmarks LM-O, YCB-V, and T-Less demonstrate that our method surpasses all refinement-free methods and is even on par with expensive refinement-based approaches. Crucially, our approach is computationally efficient and enables real-time critical applications with high accuracy requirements.
Yongliang Lin, Yongzhi Su, Praveen Nathan, Sandeep Inuganti, Yan Di, Martin Sundermeyer, Fabian Manhardt, Didier Stricker, Jason R. Rambach, Yu Zhang 0018
CVPR9
2024 In-Domain Inversion for Improved 3D Face Alignment on Asymmetrical Expressions
abstract
Facial landmark detection, often termed as face alignment, is a well-studied research problem in computer vision. Nonetheless, face alignment on asymmetrical expressions has been overlooked in the literature, particularly for unusual gestures observed in individuals with unilateral facial paralysis. In this paper, we explore in-domain inversion in a semi-supervised approach for face alignment and target the detection of 3D landmarks on symmetrical and extremely asymmetrical facial expressions due to paralysis. Our approach first leverages unlabeled face data to synthesize face images, while learning a compressed representation in the latent space. Then, it integrates in-domain inversion in the self-supervised stage, to make the latent space semantically meaningful. This is exploited in the supervised stage by a 2D face landmark detector, trained on labeled data. Finally, we extend the pipeline to 3D face alignment and regress the depth coordinate from the intermediate latent space and the predicted 2D landmarks. We evaluate and compare our method to related work on publicly available datasets, and demonstrate that our approach outperforms the state of the art in the detection of 3D facial landmarks in our newly introduced dataset of facial paralysis, ParFace. Our implementation and dataset are available at https://github.com/jilliam/ParFace.
Jilliam María Díaz Barros, Jason R. Rambach, Pramod Murthy, Didier Stricker
FG2
2024 CaRaCTO: Robust Camera-Radar Extrinsic Calibration with Triple Constraint Optimization
Mahdi Chamseddine, Jason R. Rambach, Didier Stricker
ICPRAM2
2024 Achieving RGB-D Level Segmentation Performance from a Single ToF Camera
Pranav Sharma, Jigyasa Katrolia, Jason R. Rambach, Bruno Mirbach, Didier Stricker
ICPRAM3
2024 Single Frame Semantic Segmentation Using Multi-Modal Spherical Images
abstract
In recent years, the research community has shown a lot of interest to panoramic images that offer a 360° directional perspective. Multiple data modalities can be fed, and complimentary characteristics can be utilized for more robust and rich scene interpretation based on semantic segmentation, to fully realize the potential. Existing research, however, mostly concentrated on pinhole RGB-X semantic segmentation. In this study, we propose a transformer-based cross-modal fusion architecture to bridge the gap between multi-modal fusion and omnidirectional scene perception. We employ distortion-aware modules to address extreme object deformations and panorama distortions that result from equirectangular representation. Additionally, we conduct cross-modal interactions for feature rectification and information exchange before merging the features in order to communicate long-range contexts for bi-modal and tri-modal feature streams. In thorough tests using combinations of four different modality types in three indoor panoramic-view datasets, our technique achieved state-of-the-art mIoU performance: 60.60% on Stanford2D3DS [2] (RGB-HHA), 71.97% on Structured3D [44] (RGB-D-N), and 35.92% on Matterport3D [5] (RGB-D).1
Suresh Guttikonda, Jason R. Rambach
WACV2
2023 U-RED: Unsupervised 3D Shape Retrieval and Deformation for Partial Point Clouds
abstract
In this paper, we propose U-RED, an Unsupervised shape REtrieval and Deformation pipeline that takes an arbitrary object observation as input, typically captured by RGB images or scans, and jointly retrieves and deforms the geometrically similar CAD models from a pre-established database to tightly match the target. Considering existing methods typically fail to handle noisy partial observations, U-RED is designed to address this issue from two aspects. First, since one partial shape may correspond to multiple potential full shapes, the retrieval method must allow such an ambiguous one-to-many relationship. Thereby U-RED learns to project all possible full shapes of a partial target onto the surface of a unit sphere. Then during inference, each sampling on the sphere will yield a feasible retrieval. Second, since real-world partial observations usually contain noticeable noise, a reliable learned metric that measures the similarity between shapes is necessary for stable retrieval. In U-RED, we design a novel point-wise residual-guided metric that allows noise-robust comparison. Extensive experiments on the synthetic datasets PartNet, ComplementMe and the real-world dataset Scan2CAD demonstrate that U-RED surpasses existing state-of-the-art approaches by 47.3%, 16.7% and 31.6% respectively under Chamfer Distance.
Yan Di, Chenyangguang Zhang, Ruida Zhang, Fabian Manhardt, Yongzhi Su, Jason R. Rambach, Didier Stricker, Xiangyang Ji, Federico Tombari
ICCV6
2023 Confidence-Aware Clustered Landmark Filtering For Hybrid 3D Face Tracking
abstract
The detection of facial landmarks in 2D images has received a great attention in the last decade, as it is a key step for several computer-vision-related applications. Most of the approaches are focused on still images, and are extended to videos by using a tracking-by-detection scheme. In this work, we propose a frame-to-frame tracking module based on grouped-landmark Kalman filters that can be integrated into existing deep-learning-based 3D face alignment pipelines. This method improves the landmark accuracy in cases with large occlusion, extreme head poses and blurriness that affect existing approaches. Our experiments on the Menpo 3DA-2D benchmark show improvements on model-free and 3D-model-based face alignment approaches.
Jilliam María Díaz Barros, Didier Stricker, Jason R. Rambach
ICIP4
2022 ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose Estimation
abstract
Establishing correspondences from image to 3D has been a key task of 6DoF object pose estimation for a long time. To predict pose more accurately, deeply learned dense maps replaced sparse templates. Dense methods also improved pose estimation in the presence of occlusion. More recently researchers have shown improvements by learning object fragments as segmentation. In this work, we present a discrete descriptor, which can represent the object surface densely. By incorporating a hierarchical binary grouping, we can encode the object surface very efficiently. Moreover, we propose a coarse to fine training strategy, which enables fine-grained correspondence prediction. Finally, by matching predicted codes with object surface and using a PnP solver, we estimate the 6DoF pose. Results on the public LM-O and YCB-V datasets show major improvement over the state of the art w.r.t. ADD(-S) metric, even surpassing RGB-D based methods in some cases.
Yongzhi Su, Mahdi Saleh, Torben Fetzer, Jason R. Rambach, Nassir Navab, Benjamin Busam, Didier Stricker, Federico Tombari
CVPR4
2022 Classification of Manual Versus Autonomous Driving based on Machine Learning of Eye Movement Patterns
abstract
Recent advances in autonomous driving systems raise new questions about how to enhance the communication and takeover control between the system and the driver. Eye tracking technologies have shown their feasibility to recognize whether the driver’s gaze is directed ‘on-road’ or ‘off-road’. However, this binary information alone is not sufficient to infer the driver’s engagement in the driving task. In the present work, we take the next step and investigate how driving modes (autopilot, navigation system, and printed map) associated with different levels of engagement can be categorized from drivers’ gaze patterns. Using gaze data recorded in these three driving tasks along with several state-of-the-art machine learning methods, we demonstrate that the driving modes are associated with different gaze patterns. We achieved an average accuracy of 90.1% for binary and 80.3% for multi-class driving mode classification. Our findings pave the way for enhancing driver monitoring systems in (semi-) autonomous cars.
Iuliia Brishtel, Stephan Krauß, Jason R. Rambach, Igor Vozniak, Didier Stricker
SMC4
2022 Fusion Point Pruning for Optimized 2D Object Detection with Radar-Camera Fusion
abstract
Object detection is one of the most important perception tasks for advanced driver assistant systems and autonomous driving. Due to its complementary features and moderate cost, radar-camera fusion is of particular interest in the automotive industry but comes with the challenge of how to optimally fuse the heterogeneous data sources. To solve this for 2D object detection, we propose two new techniques to project the radar detections onto the image plane, exploiting additional uncertainty information. We also introduce a new technique called fusion point pruning, which automatically finds the best fusion points of radar and image features in the neural network architecture. These new approaches combined surpass the state of the art in 2D object detection performance for radar-camera fusion models, evaluated with the nuScenes dataset. We further find that the utilization of radar-camera fusion is especially beneficial for night scenes.
Lukas Stäcker, Philipp Heidenreich, Jason R. Rambach, Didier Stricker
WACV3
2021 TICaM: A Time-of-flight In-car Cabin Monitoring Dataset
Jigyasa Katrolia, Ahmed El-Sherif, Hartmut Feld, Bruno Mirbach, Jason R. Rambach, Didier Stricker
BMVC5
2021 PlaneRecNet: Multi-Task Learning with Cross-Task consistency for Piece-Wise Plane Detection and Reconstruction from a Single RGB Image
Yaxu Xie, Fangwen Shu, Jason R. Rambach, Alain Pagani, Didier Stricker
BMVC3
2021 Semantic Segmentation in Depth Data: A Comparative Evaluation of Image and Point Cloud Based Methods
abstract
The problem of semantic segmentation from depth images can be addressed by segmenting directly in the image domain or at 3D point cloud level. In this paper, we attempt for the first time to provide a study and experimental comparison of the two approaches. Through experiments on three datasets, namely SUN RGB-D, NYUdV2 and TICaM, we extensively compare various semantic segmentation algorithms, the input to which includes images and point clouds derived from them. Based on this, we offer analysis of the performance and computational cost of these algorithms that can provide guidelines on when each method should be preferred.
Jigyasa Katrolia, Lars Krämer, Jason R. Rambach, Bruno Mirbach, Didier Stricker
ICIP3
2021 PlaneSegNet: Fast and Robust Plane Estimation Using a Single-stage Instance Segmentation CNN
abstract
Instance segmentation of planar regions in indoor scenes benefits visual SLAM and other applications such as augmented reality (AR) where scene understanding is required. Existing methods built upon two-stage frameworks show satisfactory accuracy but are limited by low frame rates. In this work, we propose a real-time deep neural architecture that estimates piece-wise planar regions from a single RGB image. Our model employs a variant of a fast single-stage CNN architecture to segment plane instances. Considering the particularity of the target detected, we propose Fast Feature Non-maximum Suppression (FF-NMS) to reduce the suppression errors resulted from overlapping bounding boxes of planes. We also utilize a Residual Feature Augmentation module in the Feature Pyramid Network (FPN) . Our method achieves significantly higher frame-rates and comparable segmentation accuracy against two-stage methods. We automatically label over 70,000 images as ground truth from the Stanford 2D-3D-Semantics dataset. Moreover, we incorporate our method with a state-of-the-art planar SLAM and validate its benefits.
Yaxu Xie, Jason R. Rambach, Fangwen Shu, Didier Stricker
ICRA2
2020 Ghost Target Detection in 3D Radar Data using Point Cloud based Deep Neural Network
abstract
Ghost targets are targets that appear at wrong locations in radar data and are caused by the presence of multiple indirect reflections between the target and the sensor. In this work, we introduce the first point based deep learning approach for ghost target detection in 3D radar point clouds. This is done by extending the PointNet network architecture by modifying its input to include radar point features beyond location and introducing skip connetions. We compare different input modalities and analyze the effects of the changes we introduced. We also propose an approach for automatic labeling of ghost targets 3D radar data using lidar as reference. The algorithm is trained and tested on real data in various driving scenarios and the tests show promising results in classifying real and ghost radar targets.
Mahdi Chamseddine, Jason R. Rambach, Didier Stricker, Oliver Wasenmüller
ICPR2
2016 Learning to Fuse: A Deep Learning Approach to Visual-Inertial Camera Pose Estimation
abstract
Camera pose estimation is the cornerstone of Augmented Reality applications. Pose tracking based on camera images exclusively has been shown to be sensitive to motion blur, occlusions, and illumination changes. Thus, a lot of work has been conducted over the last years on visual-inertial pose tracking using acceleration and angular velocity measurements from inertial sensors in order to improve the visual tracking. Most proposed systems use statistical filtering techniques to approach the sensor fusion problem, that require complex system modelling and calibrations in order to perform adequately. In this work we present a novel approach to sensor fusion using a deep learning method to learn the relation between camera poses and inertial sensor measurements. A long short-term memory model (LSTM) is trained to provide an estimate of the current pose based on previous poses and inertial measurements. This estimates then appropriately combined with the output of a visual tracking system using a linear Kalman Filter to provide a robust final pose estimate. Our experimental results confirm the applicability and tracking performance improvement gained from the proposed sensor fusion system.
Jason R. Rambach, Aditya Tewari, Alain Pagani, Didier Stricker
ISMAR1
2015 Collaborative multi-camera face recognition and tracking
abstract
In this paper, a framework for collaborative face recognition from video sequences in a multi-camera environment is proposed. Collaboration between cameras allows for higher recognition performance in both the common and non-common field-of-view (FOV) cases. For the latter, the appearance of an object in a nearby camera is predicted using the last tracked position of the object paired with a time-of-arrival model between camera pairs. An experiment using four cameras in an office environment confirms the applicability and performance gains of the proposed framework.
Jason R. Rambach, Marco F. Huber, Mark Ryan Balthasar, Abdelhak M. Zoubir
AVSS1