Eldar Insafutdinov

dblp:172/1246 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
3D vision · 50% Face, body and person analysis · 23% Video understanding and tracking · 20%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 61% Virtual and augmented reality · 39%

Topics — the 24 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
human pose estimation
1.142018
PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018
Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications · CVPR 2017
EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016
Computer vision › Video understanding and tracking › object tracking
3d object tracking
0.912025
Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025
Computer vision › 3D vision
3d reconstruction
0.912025
Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025
Computer vision › 3D vision › 3d reconstruction
dynamic 3d reconstruction
0.912025
Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025
Computer vision › 3D vision
scene flow estimation
0.912025
Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025
Computer vision › 3D vision › 3d object detection › 3d object localization
3d bounding box estimation
0.612022
Lifting 2D Object Locations to 3D by Discounting LiDAR Outliers across Objects and Views · ICRA 2022
Computer vision › 3D vision
3d object detection
0.612022
Lifting 2D Object Locations to 3D by Discounting LiDAR Outliers across Objects and Views · ICRA 2022
Computer vision › Face, body and person analysis › human pose estimation
articulated pose estimation
0.622017
Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications · CVPR 2017
ArtTrack: Articulated Multi-Person Tracking in the Wild · CVPR 2017
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
neural surface reconstruction
0.612022
SNeS: Learning Probably Symmetric Neural Surfaces from Incomplete Data · ECCV (32) 2022
Computer vision › Face, body and person analysis › human pose estimation
multi-person pose estimation
0.522016
DeeperCut: A Deeper, Stronger, and Faster Multi-person Pose Estimation Model · ECCV (6) 2016
DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation · CVPR 2016
Computer vision › Video understanding and tracking › object tracking
articulated object tracking
0.312018
PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018
Computer vision › Video understanding and tracking › multi-object tracking
multi-person pose tracking
0.312018
PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018
Computer vision › Image recognition and object detection › point set representation
point cloud representation
0.312018
Unsupervised Learning of Shape and Pose with Differentiable Point Clouds · NeurIPS 2018
Computer vision › 3D vision › pose estimation
shape and pose estimation
0.312018
Unsupervised Learning of Shape and Pose with Differentiable Point Clouds · NeurIPS 2018
Computer vision › Video understanding and tracking
multi-object tracking
0.312017
Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications · CVPR 2017
Computer vision › Video understanding and tracking › multi-object tracking
multi-person tracking
0.312017
ArtTrack: Articulated Multi-Person Tracking in the Wild · CVPR 2017
Computer vision › Segmentation and scene understanding › image segmentation
semantic and instance segmentation
0.312017
Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications · CVPR 2017
Computer vision › 3D vision
depth estimation
0.312025
Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
whole-body pose estimation
0.212016
EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016
Virtual and augmented reality › tracking
egocentric motion capture
0.212016
EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016
Computer animation and physical simulation › motion capture
markerless motion capture
0.212016
EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016
Computer animation and physical simulation
motion capture
0.212016
EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016
Computer vision › 3D vision
point cloud processing
0.212022
Lifting 2D Object Locations to 3D by Discounting LiDAR Outliers across Objects and Views · ICRA 2022
Computer vision › 3D vision › 3d reconstruction › geometric reconstruction
symmetry-based reconstruction
0.212022
SNeS: Learning Probably Symmetric Neural Surfaces from Incomplete Data · ECCV (32) 2022

Methods — techniques the papers use, named apart from their topics

pointmap regression · 0.9evaluation server · 0.7benchmark dataset · 0.7temporal consistency · 0.6self-supervision · 0.6rotation optimization · 0.6neural implicit surface · 0.6ensemble distillation · 0.3convolutional network · 0.3convolutional neural network · 0.3generative pose estimation · 0.2fisheye stereo · 0.2convnet-based body-part detection · 0.2
YearPublicationVenuePosition
2025 Flash3D: Feed-Forward Generalisable 3D Scene Reconstruction from a Single Image
abstract
We propose Flash3D, a method for scene reconstruction and novel view synthesis from a single image which is both very generalisable and efficient. For generalisability, we start from a 'foundation' model for monocular depth estimation and extend it to a full 3D shape and appearance reconstructor. For efficiency, we base this extension on feed-forward Gaussian Splatting. Specifically, we predict a first layer of 3D Gaussians at the predicted depth, and then add additional layers of Gaussians that are offset in space, allowing the model to complete the reconstruction behind occlusions and truncations. Flash3D is very efficient, trainable on a single GPU in a day, and thus accessible to most researchers. It achieves state-of-the-art results when trained and tested on RealEstate10k. When transferred to unseen datasets like NYU it outperforms competitors by a large margin. More impressively, when transferred to KITTI, Flash3D achieves better PSNR than methods trained specifically on that dataset. In some instances, it even outperforms recent methods that use multiple views as input. Code, models, demo, and more results are available at https://www.robots.ox.ac.uk/~vgg/research/flash3d/.
Stanislaw Szymanowicz, Eldar Insafutdinov, Chuanxia Zheng, Dylan Campbell, João F. Henriques, Christian Rupprecht 0001, Andrea Vedaldi
3DV2
2025 Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction
abstract
DUSt3R has recently shown that one can reduce many tasks in multi-view geometry, including estimating camera intrinsics and extrinsics, reconstructing the scene in 3D, and establishing image correspondences, to the prediction of a pair of viewpoint-invariant point maps, i.e., pixel-aligned point clouds defined in a common reference frame. This formulation is elegant and powerful, but unable to tackle dynamic scenes. To address this challenge, we introduce the concept of Dynamic Point Maps (DPM), extending standard point maps to support 4D tasks such as motion segmentation, scene flow estimation, 3D object tracking, and 2D correspondence. Our key intuition is that, when time is introduced, there are several possible spatial and time references that can be used to define the point maps. We identify a minimal subset of such combinations that can be regressed by a network to solve the sub tasks mentioned above. We train a DPM predictor on a mixture of synthetic and real data and evaluate it across diverse benchmarks for video depth prediction, dynamic point cloud reconstruction, 3D scene flow and object pose tracking, achieving state-of-the-art performance. Code, models and additional results are available at https://www.robots.ox.ac.uk/~vgg/research/dynamic-point-maps/.
Edgar Sucar, Zihang Lai, Eldar Insafutdinov, Andrea Vedaldi
ICCV3
2025 SEED4D: A Synthetic Ego-Exo Dynamic 4D Data Generator, Driving Dataset and Benchmark
abstract
Models for egocentric 3D and 4D reconstruction, including few-shot interpolation and extrapolation settings, can benefit from having images from exocentric viewpoints as supervision signals. No existing dataset provides the necessary mixture of complex, dynamic, and multi-view data. To facilitate the development of 3D and 4D reconstruction methods in the autonomous driving context, we propose a Synthetic Ego-Exo Dynamic 4D (SEED4D) data generator and dataset. We present a customizable, easy-to-use data generator for spatio-temporal multi-view data creation. Our open-source data generator allows the creation of synthetic data for camera setups commonly used in the NuScenes, KITTI360, and Waymo datasets. Additionally, SEED4D encompasses two large-scale multi-view synthetic urban scene datasets. Our static (3D) dataset encompasses 212k inward- and outward-facing vehicle images from 2k scenes, while our dynamic (4D) dataset contains 16.8M images from 10k trajectories, each sampled at 100 points in time with egocentric images, exocentric images, and LiDAR data. The datasets and the data generator can be found here.
Marius Kästingschäfer, Théo Gieruc, Sebastian Bernhard, Dylan Campbell, Eldar Insafutdinov, Eyvaz Najafli, Thomas Brox
WACV5
2022 SNeS: Learning Probably Symmetric Neural Surfaces from Incomplete Data
Eldar Insafutdinov, Dylan Campbell, João F. Henriques, Andrea Vedaldi
ECCV (32)1
2022 Lifting 2D Object Locations to 3D by Discounting LiDAR Outliers across Objects and Views
abstract
We present a system for automatic converting of 2D mask object predictions and raw LiDAR point clouds into full 3D bounding boxes of objects. Because the LiDAR point clouds are partial, directly fitting bounding boxes to the point clouds is meaningless. Instead, we suggest that obtaining good results requires sharing information between all objects in the dataset jointly, over multiple frames. We then make three improvements to the baseline. First, we address ambiguities in predicting the object rotations via direct optimization in this space while still backpropagating rotation prediction through the model. Second, we explicitly model outliers and task the network with learning their typical patterns, thus better discounting them. Third, we enforce temporal consistency when video data is available. With these contributions, our method significantly outperforms previous work despite the fact that those methods use significantly more complex pipelines, 3D models and additional human-annotated external sources of prior information.
Robert McCraith, Eldar Insafutdinov, Lukás Neumann, Andrea Vedaldi
ICRA2
2019 360-Degree Textures of People in Clothing from a Single Image
abstract
In this paper we predict a full 3D avatar of a person from a single image. We infer texture and geometry in the UV-space of the SMPL model using an image-to-image translation method. Given partial texture and segmentation layout maps derived from the input view, our model predicts the complete segmentation map, the complete texture map, and a displacement map. The predicted maps can be applied to the SMPL model in order to naturally generalize to novel poses, shapes, and even new clothing. In order to learn our model in a common UV-space, we non-rigidly register the SMPL model to thousands of 3D scans, effectively encoding textures and geometries as images in correspondence. This turns a difficult 3D inference task into a simpler image-to-image translation one. Results on rendered scans of people and images from the DeepFashion dataset demonstrate that our method can reconstruct plausible 3D avatars from a single image. We further use our model to digitally change pose, shape, swap garments between people and edit clothing. To encourage research in this direction we will make the source code available for research purpose [5].
Verica Lazova, Eldar Insafutdinov, Gerard Pons-Moll
3DV2
2018 PoseTrack: A Benchmark for Human Pose Estimation and Tracking
abstract
Existing systems for video-based pose estimation and tracking struggle to perform well on realistic videos with multiple people and often fail to output body-pose trajectories consistent over time. To address this shortcoming this paper introduces PoseTrack which is a new large-scale benchmark for video-based human pose estimation and articulated tracking. Our new benchmark encompasses three tasks focusing on i) single-frame multi-person pose estimation, ii) multi-person pose estimation in videos, and iii) multi-person articulated tracking. To establish the benchmark, we collect, annotate and release a new dataset that features videos with multiple people labeled with person tracks and articulated pose. A public centralized evaluation server is provided to allow the research community to evaluate on a held-out test set. Furthermore, we conduct an extensive experimental study on recent approaches to articulated pose tracking and provide analysis of the strengths and weaknesses of the state of the art. We envision that the proposed benchmark will stimulate productive research both by providing a large and representative training dataset as well as providing a platform to objectively evaluate and compare the proposed methods. The benchmark is freely accessible at https://posetrack.net/.
Mykhaylo Andriluka, Umar Iqbal 0001, Eldar Insafutdinov, Leonid Pishchulin, Anton Milan, Juergen Gall, Bernt Schiele
CVPR3
2018 Unsupervised Learning of Shape and Pose with Differentiable Point Clouds
abstract
We address the problem of learning accurate 3D shape and camera pose from a collection of unlabeled category-specific images. We train a convolutional network to predict both the shape and the pose from a single image by minimizing the reprojection error: given several views of an object, the projections of the predicted shapes to the predicted camera poses should match the provided views. To deal with pose ambiguity, we introduce an ensemble of pose predictors which we then distill to a single "student" model. To allow for efficient learning of high-fidelity shapes, we represent the shapes by point clouds and devise a formulation allowing for differentiable projection of these. Our experiments show that the distilled ensemble of pose predictors learns to estimate the pose accurately, while the point cloud representation allows to predict detailed shape models.
Eldar Insafutdinov, Alexey Dosovitskiy
NeurIPS1
2017 ArtTrack: Articulated Multi-Person Tracking in the Wild
abstract
In this paper we propose an approach for articulated tracking of multiple people in unconstrained videos. Our starting point is a model that resembles existing architectures for single-frame pose estimation but is substantially faster. We achieve this in two ways: (1) by simplifying and sparsifying the body-part relationship graph and leveraging recent methods for faster inference, and (2) by offloading a substantial share of computation onto a feed-forward convolutional architecture that is able to detect and associate body joints of the same person even in clutter. We use this model to generate proposals for body joint locations and formulate articulated tracking as spatio-temporal grouping of such proposals. This allows to jointly solve the association problem for all people in the scene by propagating evidence from strong detections through time and enforcing constraints that each proposal can be assigned to one person only. We report results on a public MPII Human Pose benchmark and on a new MPII Video Pose dataset of image sequences with multiple people. We demonstrate that our model achieves state-of-the-art results while using only a fraction of time and is able to leverage temporal information to improve state-of-the-art for crowded scenes.
Eldar Insafutdinov, Mykhaylo Andriluka, Leonid Pishchulin, Siyu Tang 0001, Evgeny Levinkov, Bjoern Andres, Bernt Schiele
CVPR1
2017 Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications
abstract
We state a combinatorial optimization problem whose feasible solutions define both a decomposition and a node labeling of a given graph. This problem offers a common mathematical abstraction of seemingly unrelated computer vision tasks, including instance-separating semantic segmentation, articulated human body pose estimation and multiple object tracking. Conceptually, it generalizes the unconstrained integer quadratic program and the minimum cost lifted multicut problem, both of which are NP-hard. In order to find feasible solutions efficiently, we define two local search algorithms that converge monotonously to a local optimum, offering a feasible solution at any time. To demonstrate the effectiveness of these algorithms in tackling computer vision tasks, we apply them to instances of the problem that we construct from published data, using published algorithms. We report state-of-the-art application-specific accuracy in the three above-mentioned applications.
Evgeny Levinkov, Jonas Uhrig, Siyu Tang 0001, Mohamed Omran, Eldar Insafutdinov, Alexander Kirillov, Carsten Rother, Thomas Brox, Bernt Schiele, Bjoern Andres
CVPR5
2016 DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation
abstract
This paper considers the task of articulated human pose estimation of multiple people in real world images. We propose an approach that jointly solves the tasks of detection and pose estimation: it infers the number of persons in a scene, identifies occluded body parts, and disambiguates body parts between people in close proximity of each other. This joint formulation is in contrast to previous strategies, that address the problem by first detecting people and subsequently estimating their body pose. We propose a partitioning and labeling formulation of a set of body-part hypotheses generated with CNN-based part detectors. Our formulation, an instance of an integer linear program, implicitly performs non-maximum suppression on the set of part candidates and groups them to form configurations of body parts respecting geometric and appearance constraints. Experiments on four different datasets demonstrate state-of-the-art results for both single person and multi person pose estimation.
Leonid Pishchulin, Eldar Insafutdinov, Siyu Tang 0001, Bjoern Andres, Mykhaylo Andriluka, Peter V. Gehler, Bernt Schiele
CVPR2
2016 DeeperCut: A Deeper, Stronger, and Faster Multi-person Pose Estimation Model
Eldar Insafutdinov, Leonid Pishchulin, Bjoern Andres, Mykhaylo Andriluka, Bernt Schiele
ECCV (6)1
2016 EgoCap: egocentric marker-less motion capture with two fisheye cameras
abstract
Marker-based and marker-less optical skeletal motion-capture methods use an outside-in arrangement of cameras placed around a scene, with viewpoints converging on the center. They often create discomfort with marker suits, and their recording volume is severely restricted and often constrained to indoor scenes with controlled backgrounds. Alternative suit-based systems use several inertial measurement units or an exoskeleton to capture motion with an inside-in setup, i.e. without external sensors. This makes capture independent of a confined volume, but requires substantial, often constraining, and hard to set up body instrumentation. Therefore, we propose a new method for real-time, marker-less, and egocentric motion capture: estimating the full-body skeleton pose from a lightweight stereo pair of fisheye cameras attached to a helmet or virtual reality headset - an optical inside-in method, so to speak. This allows full-body motion capture in general indoor and outdoor scenes, including crowded scenes with many people nearby, which enables reconstruction in larger-scale activities. Our approach combines the strength of a new generative pose estimation framework for fisheye views with a ConvNet-based body-part detector trained on a large new dataset. It is particularly useful in virtual reality to freely roam and interact, while seeing the fully motion-captured virtual body.
Helge Rhodin, Christian Richardt, Dan Casas, Eldar Insafutdinov, Mohammad Shafiei, Hans-Peter Seidel, Bernt Schiele, Christian Theobalt
ACM Trans. Graph.4