EDBT 2026 Demo / reviewers in the wild / expert
Omid Hosseini Jafari
dblp:151/9744
· DBLP profile ↗
7ranked-venue papers
4as first author
1since 2021 · last 2024
0000-0003-0002-2608ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
3D vision · 62% Image recognition and object detection · 14% Segmentation and scene understanding · 12% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › stereo vision
stereo rectification |
0.8 | 1 | 2024 | Flow-Guided Online Stereo Rectification for Wide Baseline Stereo · CVPR 2024 |
Computer vision › 3D vision
stereo vision |
0.8 | 1 | 2024 | Flow-Guided Online Stereo Rectification for Wide Baseline Stereo · CVPR 2024 |
Computer vision › 3D vision › stereo vision › stereo matching
wide-baseline stereo |
0.8 | 1 | 2024 | Flow-Guided Online Stereo Rectification for Wide Baseline Stereo · CVPR 2024 |
Robotics › Autonomous driving
perception |
0.5 | 2 | 2024 | Bounding Boxes, Segmentations and Object Coordinates: How Important is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios? · ICCV 2017 Flow-Guided Online Stereo Rectification for Wide Baseline Stereo · CVPR 2024 |
Computer vision › Image recognition and object detection
pedestrian detection |
0.4 | 2 | 2016 | Real-time RGB-D based template matching pedestrian detection · ICRA 2016 Real-time RGB-D based people detection and tracking for mobile robots and head-worn cameras · ICRA 2014 |
Computer vision › 3D vision › motion estimation
3d motion estimation |
0.3 | 1 | 2017 | Bounding Boxes, Segmentations and Object Coordinates: How Important is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios? · ICCV 2017 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.3 | 1 | 2017 | Bounding Boxes, Segmentations and Object Coordinates: How Important is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios? · ICCV 2017 |
Computer vision › 3D vision
scene flow estimation |
0.3 | 1 | 2017 | Bounding Boxes, Segmentations and Object Coordinates: How Important is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios? · ICCV 2017 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.3 | 1 | 2017 | Analyzing modular CNN architectures for joint depth prediction and semantic segmentation · ICRA 2017 |
Computer vision › Image recognition and object detection
template matching |
0.2 | 1 | 2016 | Real-time RGB-D based template matching pedestrian detection · ICRA 2016 |
Computer vision › 3D vision
depth estimation |
0.1 | 1 | 2017 | Analyzing modular CNN architectures for joint depth prediction and semantic segmentation · ICRA 2017 |
Computer vision › 3D vision › 3d object detection
depth-based detection |
0.1 | 1 | 2016 | Real-time RGB-D based template matching pedestrian detection · ICRA 2016 |
Computer vision › Video understanding and tracking › multi-object tracking
multi-person tracking |
0.1 | 1 | 2014 | Real-time RGB-D based people detection and tracking for mobile robots and head-worn cameras · ICRA 2014 |
Methods — techniques the papers use, named apart from their topics
vertical optical flow · 0.8stereo correlation volume · 0.8cross-image attention · 0.8HOG detector · 0.4multi-task learning · 0.3convolutional neural network · 0.3conditional random field · 0.3CNN · 0.3depth-based template matching · 0.2multi-hypothesis tracking · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Flow-Guided Online Stereo Rectification for Wide Baseline StereoabstractStereo rectification is widely considered “solved” due to the abundance of traditional approaches to perform recti-fication. However, autonomous vehicles and robots in-the-wild require constant re-calibration due to exposure to var-ious environmental factors, including vibration, and structural stress, when cameras are arranged in a wide-baseline configuration. Conventional rectification methods fail in these challenging scenarios: especially for larger vehicles, such as autonomous freight trucks and semi-trucks, the resulting incorrect rectification severely affects the quality of downstream tasks that use stereo/multi-view data. To tackle these challenges, we propose an online rectification approach that operates at real-time rates while achieving high accuracy. We propose a novel learning-based online cal-ibration approach that utilizes stereo correlation volumes built from a feature representation obtained from cross-image attention. Our model is trained to minimize vertical optical flow as proxy rectification constraint, and predicts the relative rotation between the stereo pair. The method is real-time and even outperforms conventional methods used for offline calibration, and substantially improves downstream stereo depth, post-rectification. We release two public datasets (https://light.princeton.edu/online-stereo-recification/), a synthetic and experimental wide baseline dataset, to foster further research. Anush Kumar, Fahim Mannan, Omid Hosseini Jafari, Shile Li, Felix Heide |
CVPR | 3 |
| 2018 | iPose: Instance-Aware 6D Pose Estimation of Partly Occluded Objects
Omid Hosseini Jafari, Siva Karthik Mustikovela, Karl Pertsch, Eric Brachmann, Carsten Rother |
ACCV (3) | 1 |
| 2018 | Deep Object Co-segmentation
Weihao Li 0005, Omid Hosseini Jafari, Carsten Rother |
ACCV (3) | 2 |
| 2017 | Bounding Boxes, Segmentations and Object Coordinates: How Important is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios?abstractExisting methods for 3D scene flow estimation often fail in the presence of large displacement or local ambiguities, e.g., at texture-less or reflective surfaces. However, these challenges are omnipresent in dynamic road scenes, which is the focus of this work. Our main contribution is to overcome these 3D motion estimation problems by exploiting recognition. In particular, we investigate the importance of recognition granularity, from coarse 2D bounding box estimates over 2D instance segmentations to fine-grained 3D object part predictions. We compute these cues using CNNs trained on a newly annotated dataset of stereo images and integrate them into a CRF-based model for robust 3D scene flow estimation - an approach we term Instance Scene Flow. We analyze the importance of each recognition cue in an ablation study and observe that the instance segmentation cue is by far strongest, in our setting. We demonstrate the effectiveness of our method on the challenging KITTI 2015 scene flow benchmark where we achieve state-of-the-art performance at the time of submission. Aseem Behl, Omid Hosseini Jafari, Siva Karthik Mustikovela, Hassan Abu Alhaija, Carsten Rother, Andreas Geiger 0001 |
ICCV | 2 |
| 2017 | Analyzing modular CNN architectures for joint depth prediction and semantic segmentationabstractThis paper addresses the task of designing a modular neural network architecture that jointly solves different tasks. As an example we use the tasks of depth estimation and semantic segmentation given a single RGB image. The main focus of this work is to analyze the cross-modality influence between depth and semantic prediction maps on their joint refinement. While most of the previous works solely focus on measuring improvements in accuracy, we propose a way to quantify the cross-modality influence. We show that there is a relationship between final accuracy and cross-modality influence, although not a simple linear one. Hence a larger cross-modality influence does not necessarily translate into an improved accuracy. We find that a beneficial balance between the cross-modality influences can be achieved by network architecture and conjecture that this relationship can be utilized to understand different network design choices. Towards this end we propose a Convolutional Neural Network (CNN) architecture that fuses the state-of-the-art results for depth estimation and semantic labeling. By balancing the cross-modality influences between depth and semantic prediction, we achieve improved results for both tasks using the NYU-Depth v2 benchmark. Omid Hosseini Jafari, Oliver Groth, Alexander Kirillov, Michael Ying Yang, Carsten Rother |
ICRA | 1 |
| 2016 | Real-time RGB-D based template matching pedestrian detectionabstractPedestrian detection is one of the most popular topics in computer vision and robotics. Considering challenging issues in multiple pedestrian detection, we present a real-time depth-based template matching people detector. In this paper, we propose different approaches for training the depth-based template. We train multiple templates for handling issues due to various upper-body orientations of the pedestrians and different levels of detail in depth-map of the pedestrians with various distances from the camera. And, we take into account the degree of reliability for different regions of sliding window by proposing the weighted template approach. Furthermore, we combine the depth-detector with an appearance based detector as a verifier to take advantage of the appearance cues for dealing with the limitations of depth data. We evaluate our method on the challenging ETH dataset sequence. We show that our method outperforms the state-of-the-art approaches. Omid Hosseini Jafari, Michael Ying Yang |
ICRA | 1 |
| 2014 | Real-time RGB-D based people detection and tracking for mobile robots and head-worn camerasabstractWe present a real-time RGB-D based multiperson detection and tracking system suitable for mobile robots and head-worn cameras. Our approach combines RGB-D visual odometry estimation, region-of-interest processing, ground plane estimation, pedestrian detection, and multi-hypothesis tracking components into a robust vision system that runs at more than 20fps on a laptop. As object detection is the most expensive component in any such integration, we invest significant effort into taking maximum advantage of the available depth information. In particular, we propose to use two different detectors for different distance ranges. For the close range (up to 5–7m), we present an extremely fast depth-based upper-body detector that allows video-rate system performance on a single CPU core when applied to Kinect sensors. In order to cover also farther distance ranges, we optionally add an appearance-based full-body HOG detector (running on the GPU) that exploits scene geometry to restrict the search space. Our approach can work with both Kinect RGB-D input for indoor settings and with stereo depth input for outdoor scenarios. We quantitatively evaluate our approach on challenging indoor and outdoor sequences and show state-of-the-art performance in a large variety of settings. Our code is publicly available. Omid Hosseini Jafari, Dennis Mitzel, Bastian Leibe |
ICRA | 1 |