Ivan Shugurov

dblp:236/5918 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-5413-6622ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 85% Image recognition and object detection · 8% Face, body and person analysis · 5%
Computer graphics and multimedia
1 paper
Rendering · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
object pose estimation
2.142022
DPODv2: Dense Correspondence-Based 6 DoF Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022
WeLSA: Learning to Predict 6D Pose from Weakly Labeled Data Using Shape Alignment · ECCV (8) 2022
OSOP: A Multi-Stage One Shot Object Pose Estimation Framework · CVPR 2022
Computer vision › 3D vision
human mesh recovery
0.912025
PHD: Personalized 3D Human Body Fitting with Point Diffusion · ICCV 2025
Computer vision › 3D vision › point cloud analysis › point cloud learning
point cloud feature learning
0.812024
RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud Registration · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › 3D vision
point cloud registration
0.812024
RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud Registration · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › 3D vision › local feature descriptor
rotation-invariant descriptor
0.812024
RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud Registration · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › Image recognition and object detection
object detection
0.722022
OSOP: A Multi-Stage One Shot Object Pose Estimation Framework · CVPR 2022
DPOD: 6D Pose Object Detector and Refiner · ICCV 2019
Computer vision › 3D vision › correspondence estimation
dense correspondence
0.612022
DPODv2: Dense Correspondence-Based 6 DoF Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Computer vision › 3D vision
3d object detection
0.412019
DPOD: 6D Pose Object Detector and Refiner · ICCV 2019
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.412019
DPOD: 6D Pose Object Detector and Refiner · ICCV 2019
Computer vision › 3D vision › pose estimation
correspondence-based pose estimation
0.412019
DPOD: 6D Pose Object Detector and Refiner · ICCV 2019
Computer vision › 3D vision
3d human pose estimation
0.312025
PHD: Personalized 3D Human Body Fitting with Point Diffusion · ICCV 2025
Machine learning › Learning paradigms
weakly supervised learning
0.212022
WeLSA: Learning to Predict 6D Pose from Weakly Labeled Data Using Shape Alignment · ECCV (8) 2022
Rendering
differentiable rendering
0.212022
DPODv2: Dense Correspondence-Based 6 DoF Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022

Methods — techniques the papers use, named apart from their topics

deep learning · 1.5ablation study · 1.1pnp · 1.0point distillation sampling · 0.9point diffusion transformer · 0.9point pair features · 0.8global aggregation · 0.8template matching · 0.6shape alignment · 0.6multi-view pose refinement · 0.6CNN · 0.6
YearPublicationVenuePosition
2025 PHD: Personalized 3D Human Body Fitting with Point Diffusion
abstract
We introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these methods often refine poses using constraints derived from the 2D image to improve alignment, this process compromises 3D accuracy by failing to jointly account for person-specific body shapes and the plausibility of 3D poses. In contrast, our pipeline decouples this process by first calibrating the user's body shape and then employing a personalized pose fitting process conditioned on that shape. To achieve this, we develop a body shape-conditioned 3D pose prior, implemented as a Point Diffusion Transformer, which iteratively guides the pose fitting via a Point Distillation Sampling loss. This learned 3D pose prior effectively mitigates errors arising from an over-reliance on 2D constraints. Consequently, our approach improves not only pelvis-aligned pose accuracy but also absolute pose accuracy -- an important metric often overlooked by prior work. Furthermore, our method is highly data-efficient, requiring only synthetic data for training, and serves as a versatile plug-and-play module that can be seamlessly integrated with existing 3D pose estimators to enhance their performance. Project page: https://phd-pose.github.io/
Hsuan-I Ho, Po-Chen Wu, Ivan Shugurov, Chengcheng Tang, Abhay Mittal, Sizhe An, Manuel Kaufmann, Linguang Zhang
ICCV4
2024 RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud Registration
abstract
Successful point cloud registration relies on accurate correspondences established upon powerful descriptors. However, existing neural descriptors either leverage a rotation-variant backbone whose performance declines under large rotations, or encode local geometry that is less distinctive. To address this issue, we introduce RIGA to learn descriptors that are Rotation-Invariant by design and Globally-Aware. From the Point Pair Features (PPFs) of sparse local regions, rotation-invariant local geometry is encoded into geometric descriptors. Global awareness of 3D structures and geometric context is subsequently incorporated, both in a rotation-invariant fashion. More specifically, 3D structures of the whole frame are first represented by our global PPF signatures, from which structural descriptors are learned to help geometric descriptors sense the 3D world beyond local regions. Geometric context from the whole scene is then globally aggregated into descriptors. Finally, the description of sparse regions is interpolated to dense point descriptors, from which correspondences are extracted for registration. To validate our approach, we conduct extensive experiments on both object- and scene-level data. With large rotations, RIGA surpasses the state-of-the-art methods by a margin of 8${}^\circ$in terms of the Relative Rotation Error on ModelNet40 and improves the Feature Matching Recall by at least 5 percentage points on 3DLoMatch.
Hao Yu 0010, Ji Hou, Zheng Qin 0002, Mahdi Saleh, Ivan Shugurov, Kai Wang 0037, Benjamin Busam, Slobodan Ilic
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 OSOP: A Multi-Stage One Shot Object Pose Estimation Framework
abstract
We present a novel one-shot method for object detection and 6 DoF pose estimation, that does not require training on target objects. At test time, it takes as input a target image and a textured 3D query model. The core idea is to represent a 3D model with a number of 2D templates rendered from different viewpoints. This enables CNN-based direct dense feature extraction and matching. The object is first localized in 2D, then its approximate viewpoint is estimated, followed by dense 2D-3D correspondence prediction. The final pose is computed with PnP. We evaluate the method on LineMOD, Occlusion, Homebrewed, YCB-V and TLESS datasets and report very competitive performance in comparison to the state-of-the-art methods trained on synthetic data, even though our method is not trained on the object models used for testing.
Ivan Shugurov, Benjamin Busam, Slobodan Ilic
CVPR1
2022 WeLSA: Learning to Predict 6D Pose from Weakly Labeled Data Using Shape Alignment
Shishir Reddy Vutukur, Ivan Shugurov, Benjamin Busam, Andreas Hutter, Slobodan Ilic
ECCV (8)2
2022 DPODv2: Dense Correspondence-Based 6 DoF Pose Estimation
abstract
We propose a three-stage 6 DoF object detection method called DPODv2 (Dense Pose Object Detector) that relies on dense correspondences. We combine a 2D object detector with a dense correspondence estimation network and a multi-view pose refinement method to estimate a full 6 DoF pose. Unlike other deep learning methods that are typically restricted to monocular RGB images, we propose a unified deep learning network allowing different imaging modalities to be used (RGB or Depth). Moreover, we propose a novel pose refinement method, that is based on differentiable rendering. The main concept is to compare predicted and rendered correspondences in multiple views to obtain a pose which is consistent with predicted correspondences in all views. Our proposed method is evaluated rigorously on different data modalities and types of training data in a controlled setup. The main conclusions is that RGB excels in correspondence estimation, while depth contributes to the pose accuracy if good 3D-3D correspondences are available. Naturally, their combination achieves the overall best performance. We perform an extensive evaluation and an ablation study to analyze and validate the results on several challenging datasets. DPODv2 achieves excellent results on all of them while still remaining fast and scalable independent of the used data modality and the type of training data.
Ivan Shugurov, Sergey Zakharov, Slobodan Ilic
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 DPOD: 6D Pose Object Detector and Refiner
abstract
In this paper we present a novel deep learning method for 3D object detection and 6D pose estimation from RGB images. Our method, named DPOD (Dense Pose Object Detector), estimates dense multi-class 2D-3D correspondence maps between an input image and available 3D models. Given the correspondences, a 6DoF pose is computed via PnP and RANSAC. An additional RGB pose refinement of the initial pose estimates is performed using a custom deep learning-based refinement scheme. Our results and comparison to a vast number of related works demonstrate that a large number of correspondences is beneficial for obtaining high-quality 6D poses both before and after refinement. Unlike other methods that mainly use real data for training and do not train on synthetic renderings, we perform evaluation on both synthetic and real training data demonstrating superior results before and after refinement when compared to all recent detectors. While being precise, the presented approach is still real-time capable.
Sergey Zakharov, Ivan Shugurov, Slobodan Ilic
ICCV2