VLDB 2026 Research / reviewers in the wild / expert
Eldar Insafutdinov
dblp:172/1246
· DBLP profile ↗
13ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
3D vision · 50% Face, body and person analysis · 23% Video understanding and tracking · 20% | |
| Computer graphics and multimedia
1 paper |
Computer animation and physical simulation · 61% Virtual and augmented reality · 39% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
human pose estimation |
1.1 | 4 | 2018 | PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018 Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications · CVPR 2017 EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016 |
Computer vision › Video understanding and tracking › object tracking
3d object tracking |
0.9 | 1 | 2025 | Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025 |
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025 |
Computer vision › 3D vision › 3d reconstruction
dynamic 3d reconstruction |
0.9 | 1 | 2025 | Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025 |
Computer vision › 3D vision
scene flow estimation |
0.9 | 1 | 2025 | Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025 |
Computer vision › 3D vision › 3d object detection › 3d object localization
3d bounding box estimation |
0.6 | 1 | 2022 | Lifting 2D Object Locations to 3D by Discounting LiDAR Outliers across Objects and Views · ICRA 2022 |
Computer vision › 3D vision
3d object detection |
0.6 | 1 | 2022 | Lifting 2D Object Locations to 3D by Discounting LiDAR Outliers across Objects and Views · ICRA 2022 |
Computer vision › Face, body and person analysis › human pose estimation
articulated pose estimation |
0.6 | 2 | 2017 | Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications · CVPR 2017 ArtTrack: Articulated Multi-Person Tracking in the Wild · CVPR 2017 |
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
neural surface reconstruction |
0.6 | 1 | 2022 | SNeS: Learning Probably Symmetric Neural Surfaces from Incomplete Data · ECCV (32) 2022 |
Computer vision › Face, body and person analysis › human pose estimation
multi-person pose estimation |
0.5 | 2 | 2016 | DeeperCut: A Deeper, Stronger, and Faster Multi-person Pose Estimation Model · ECCV (6) 2016 DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation · CVPR 2016 |
Computer vision › Video understanding and tracking › object tracking
articulated object tracking |
0.3 | 1 | 2018 | PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018 |
Computer vision › Video understanding and tracking › multi-object tracking
multi-person pose tracking |
0.3 | 1 | 2018 | PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018 |
Computer vision › Image recognition and object detection › point set representation
point cloud representation |
0.3 | 1 | 2018 | Unsupervised Learning of Shape and Pose with Differentiable Point Clouds · NeurIPS 2018 |
Computer vision › 3D vision › pose estimation
shape and pose estimation |
0.3 | 1 | 2018 | Unsupervised Learning of Shape and Pose with Differentiable Point Clouds · NeurIPS 2018 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.3 | 1 | 2017 | Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications · CVPR 2017 |
Computer vision › Video understanding and tracking › multi-object tracking
multi-person tracking |
0.3 | 1 | 2017 | ArtTrack: Articulated Multi-Person Tracking in the Wild · CVPR 2017 |
Computer vision › Segmentation and scene understanding › image segmentation
semantic and instance segmentation |
0.3 | 1 | 2017 | Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications · CVPR 2017 |
Computer vision › 3D vision
depth estimation |
0.3 | 1 | 2025 | Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction · ICCV 2025 |
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
whole-body pose estimation |
0.2 | 1 | 2016 | EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016 |
Virtual and augmented reality › tracking
egocentric motion capture |
0.2 | 1 | 2016 | EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016 |
Computer animation and physical simulation › motion capture
markerless motion capture |
0.2 | 1 | 2016 | EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016 |
Computer animation and physical simulation
motion capture |
0.2 | 1 | 2016 | EgoCap: egocentric marker-less motion capture with two fisheye cameras · ACM Trans. Graph. 2016 |
Computer vision › 3D vision
point cloud processing |
0.2 | 1 | 2022 | Lifting 2D Object Locations to 3D by Discounting LiDAR Outliers across Objects and Views · ICRA 2022 |
Computer vision › 3D vision › 3d reconstruction › geometric reconstruction
symmetry-based reconstruction |
0.2 | 1 | 2022 | SNeS: Learning Probably Symmetric Neural Surfaces from Incomplete Data · ECCV (32) 2022 |
Methods — techniques the papers use, named apart from their topics
pointmap regression · 0.9evaluation server · 0.7benchmark dataset · 0.7temporal consistency · 0.6self-supervision · 0.6rotation optimization · 0.6neural implicit surface · 0.6ensemble distillation · 0.3convolutional network · 0.3convolutional neural network · 0.3generative pose estimation · 0.2fisheye stereo · 0.2convnet-based body-part detection · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Flash3D: Feed-Forward Generalisable 3D Scene Reconstruction from a Single ImageabstractWe propose Flash3D, a method for scene reconstruction and novel view synthesis from a single image which is both very generalisable and efficient. For generalisability, we start from a 'foundation' model for monocular depth estimation and extend it to a full 3D shape and appearance reconstructor. For efficiency, we base this extension on feed-forward Gaussian Splatting. Specifically, we predict a first layer of 3D Gaussians at the predicted depth, and then add additional layers of Gaussians that are offset in space, allowing the model to complete the reconstruction behind occlusions and truncations. Flash3D is very efficient, trainable on a single GPU in a day, and thus accessible to most researchers. It achieves state-of-the-art results when trained and tested on RealEstate10k. When transferred to unseen datasets like NYU it outperforms competitors by a large margin. More impressively, when transferred to KITTI, Flash3D achieves better PSNR than methods trained specifically on that dataset. In some instances, it even outperforms recent methods that use multiple views as input. Code, models, demo, and more results are available at https://www.robots.ox.ac.uk/~vgg/research/flash3d/. Stanislaw Szymanowicz, Eldar Insafutdinov, Chuanxia Zheng, Dylan Campbell, João F. Henriques, Christian Rupprecht 0001, Andrea Vedaldi |
3DV | 2 |
| 2025 | Dynamic Point Maps: A Versatile Representation for Dynamic 3D ReconstructionabstractDUSt3R has recently shown that one can reduce many tasks in multi-view geometry, including estimating camera intrinsics and extrinsics, reconstructing the scene in 3D, and establishing image correspondences, to the prediction of a pair of viewpoint-invariant point maps, i.e., pixel-aligned point clouds defined in a common reference frame. This formulation is elegant and powerful, but unable to tackle dynamic scenes. To address this challenge, we introduce the concept of Dynamic Point Maps (DPM), extending standard point maps to support 4D tasks such as motion segmentation, scene flow estimation, 3D object tracking, and 2D correspondence. Our key intuition is that, when time is introduced, there are several possible spatial and time references that can be used to define the point maps. We identify a minimal subset of such combinations that can be regressed by a network to solve the sub tasks mentioned above. We train a DPM predictor on a mixture of synthetic and real data and evaluate it across diverse benchmarks for video depth prediction, dynamic point cloud reconstruction, 3D scene flow and object pose tracking, achieving state-of-the-art performance. Code, models and additional results are available at https://www.robots.ox.ac.uk/~vgg/research/dynamic-point-maps/. Edgar Sucar, Zihang Lai, Eldar Insafutdinov, Andrea Vedaldi |
ICCV | 3 |
| 2025 | SEED4D: A Synthetic Ego-Exo Dynamic 4D Data Generator, Driving Dataset and BenchmarkabstractModels for egocentric 3D and 4D reconstruction, including few-shot interpolation and extrapolation settings, can benefit from having images from exocentric viewpoints as supervision signals. No existing dataset provides the necessary mixture of complex, dynamic, and multi-view data. To facilitate the development of 3D and 4D reconstruction methods in the autonomous driving context, we propose a Synthetic Ego-Exo Dynamic 4D (SEED4D) data generator and dataset. We present a customizable, easy-to-use data generator for spatio-temporal multi-view data creation. Our open-source data generator allows the creation of synthetic data for camera setups commonly used in the NuScenes, KITTI360, and Waymo datasets. Additionally, SEED4D encompasses two large-scale multi-view synthetic urban scene datasets. Our static (3D) dataset encompasses 212k inward- and outward-facing vehicle images from 2k scenes, while our dynamic (4D) dataset contains 16.8M images from 10k trajectories, each sampled at 100 points in time with egocentric images, exocentric images, and LiDAR data. The datasets and the data generator can be found here. Marius Kästingschäfer, Théo Gieruc, Sebastian Bernhard, Dylan Campbell, Eldar Insafutdinov, Eyvaz Najafli, Thomas Brox |
WACV | 5 |
| 2022 | SNeS: Learning Probably Symmetric Neural Surfaces from Incomplete Data
Eldar Insafutdinov, Dylan Campbell, João F. Henriques, Andrea Vedaldi |
ECCV (32) | 1 |
| 2022 | Lifting 2D Object Locations to 3D by Discounting LiDAR Outliers across Objects and ViewsabstractWe present a system for automatic converting of 2D mask object predictions and raw LiDAR point clouds into full 3D bounding boxes of objects. Because the LiDAR point clouds are partial, directly fitting bounding boxes to the point clouds is meaningless. Instead, we suggest that obtaining good results requires sharing information between all objects in the dataset jointly, over multiple frames. We then make three improvements to the baseline. First, we address ambiguities in predicting the object rotations via direct optimization in this space while still backpropagating rotation prediction through the model. Second, we explicitly model outliers and task the network with learning their typical patterns, thus better discounting them. Third, we enforce temporal consistency when video data is available. With these contributions, our method significantly outperforms previous work despite the fact that those methods use significantly more complex pipelines, 3D models and additional human-annotated external sources of prior information. Robert McCraith, Eldar Insafutdinov, Lukás Neumann, Andrea Vedaldi |
ICRA | 2 |
| 2019 | 360-Degree Textures of People in Clothing from a Single ImageabstractIn this paper we predict a full 3D avatar of a person from a single image. We infer texture and geometry in the UV-space of the SMPL model using an image-to-image translation method. Given partial texture and segmentation layout maps derived from the input view, our model predicts the complete segmentation map, the complete texture map, and a displacement map. The predicted maps can be applied to the SMPL model in order to naturally generalize to novel poses, shapes, and even new clothing. In order to learn our model in a common UV-space, we non-rigidly register the SMPL model to thousands of 3D scans, effectively encoding textures and geometries as images in correspondence. This turns a difficult 3D inference task into a simpler image-to-image translation one. Results on rendered scans of people and images from the DeepFashion dataset demonstrate that our method can reconstruct plausible 3D avatars from a single image. We further use our model to digitally change pose, shape, swap garments between people and edit clothing. To encourage research in this direction we will make the source code available for research purpose [5]. Verica Lazova, Eldar Insafutdinov, Gerard Pons-Moll |
3DV | 2 |
| 2018 | PoseTrack: A Benchmark for Human Pose Estimation and TrackingabstractExisting systems for video-based pose estimation and tracking struggle to perform well on realistic videos with multiple people and often fail to output body-pose trajectories consistent over time. To address this shortcoming this paper introduces PoseTrack which is a new large-scale benchmark for video-based human pose estimation and articulated tracking. Our new benchmark encompasses three tasks focusing on i) single-frame multi-person pose estimation, ii) multi-person pose estimation in videos, and iii) multi-person articulated tracking. To establish the benchmark, we collect, annotate and release a new dataset that features videos with multiple people labeled with person tracks and articulated pose. A public centralized evaluation server is provided to allow the research community to evaluate on a held-out test set. Furthermore, we conduct an extensive experimental study on recent approaches to articulated pose tracking and provide analysis of the strengths and weaknesses of the state of the art. We envision that the proposed benchmark will stimulate productive research both by providing a large and representative training dataset as well as providing a platform to objectively evaluate and compare the proposed methods. The benchmark is freely accessible at https://posetrack.net/. Mykhaylo Andriluka, Umar Iqbal 0001, Eldar Insafutdinov, Leonid Pishchulin, Anton Milan, Juergen Gall, Bernt Schiele |
CVPR | 3 |
| 2018 | Unsupervised Learning of Shape and Pose with Differentiable Point CloudsabstractWe address the problem of learning accurate 3D shape and camera pose from a collection of unlabeled category-specific images. We train a convolutional network to predict both the shape and the pose from a single image by minimizing the reprojection error: given several views of an object, the projections of the predicted shapes to the predicted camera poses should match the provided views. To deal with pose ambiguity, we introduce an ensemble of pose predictors which we then distill to a single "student" model. To allow for efficient learning of high-fidelity shapes, we represent the shapes by point clouds and devise a formulation allowing for differentiable projection of these. Our experiments show that the distilled ensemble of pose predictors learns to estimate the pose accurately, while the point cloud representation allows to predict detailed shape models. Eldar Insafutdinov, Alexey Dosovitskiy |
NeurIPS | 1 |
| 2017 | ArtTrack: Articulated Multi-Person Tracking in the WildabstractIn this paper we propose an approach for articulated tracking of multiple people in unconstrained videos. Our starting point is a model that resembles existing architectures for single-frame pose estimation but is substantially faster. We achieve this in two ways: (1) by simplifying and sparsifying the body-part relationship graph and leveraging recent methods for faster inference, and (2) by offloading a substantial share of computation onto a feed-forward convolutional architecture that is able to detect and associate body joints of the same person even in clutter. We use this model to generate proposals for body joint locations and formulate articulated tracking as spatio-temporal grouping of such proposals. This allows to jointly solve the association problem for all people in the scene by propagating evidence from strong detections through time and enforcing constraints that each proposal can be assigned to one person only. We report results on a public MPII Human Pose benchmark and on a new MPII Video Pose dataset of image sequences with multiple people. We demonstrate that our model achieves state-of-the-art results while using only a fraction of time and is able to leverage temporal information to improve state-of-the-art for crowded scenes. Eldar Insafutdinov, Mykhaylo Andriluka, Leonid Pishchulin, Siyu Tang 0001, Evgeny Levinkov, Bjoern Andres, Bernt Schiele |
CVPR | 1 |
| 2017 | Joint Graph Decomposition & Node Labeling: Problem, Algorithms, ApplicationsabstractWe state a combinatorial optimization problem whose feasible solutions define both a decomposition and a node labeling of a given graph. This problem offers a common mathematical abstraction of seemingly unrelated computer vision tasks, including instance-separating semantic segmentation, articulated human body pose estimation and multiple object tracking. Conceptually, it generalizes the unconstrained integer quadratic program and the minimum cost lifted multicut problem, both of which are NP-hard. In order to find feasible solutions efficiently, we define two local search algorithms that converge monotonously to a local optimum, offering a feasible solution at any time. To demonstrate the effectiveness of these algorithms in tackling computer vision tasks, we apply them to instances of the problem that we construct from published data, using published algorithms. We report state-of-the-art application-specific accuracy in the three above-mentioned applications. Evgeny Levinkov, Jonas Uhrig, Siyu Tang 0001, Mohamed Omran, Eldar Insafutdinov, Alexander Kirillov, Carsten Rother, Thomas Brox, Bernt Schiele, Bjoern Andres |
CVPR | 5 |
| 2016 | DeepCut: Joint Subset Partition and Labeling for Multi Person Pose EstimationabstractThis paper considers the task of articulated human pose estimation of multiple people in real world images. We propose an approach that jointly solves the tasks of detection and pose estimation: it infers the number of persons in a scene, identifies occluded body parts, and disambiguates body parts between people in close proximity of each other. This joint formulation is in contrast to previous strategies, that address the problem by first detecting people and subsequently estimating their body pose. We propose a partitioning and labeling formulation of a set of body-part hypotheses generated with CNN-based part detectors. Our formulation, an instance of an integer linear program, implicitly performs non-maximum suppression on the set of part candidates and groups them to form configurations of body parts respecting geometric and appearance constraints. Experiments on four different datasets demonstrate state-of-the-art results for both single person and multi person pose estimation. Leonid Pishchulin, Eldar Insafutdinov, Siyu Tang 0001, Bjoern Andres, Mykhaylo Andriluka, Peter V. Gehler, Bernt Schiele |
CVPR | 2 |
| 2016 | DeeperCut: A Deeper, Stronger, and Faster Multi-person Pose Estimation Model
Eldar Insafutdinov, Leonid Pishchulin, Bjoern Andres, Mykhaylo Andriluka, Bernt Schiele |
ECCV (6) | 1 |
| 2016 | EgoCap: egocentric marker-less motion capture with two fisheye camerasabstractMarker-based and marker-less optical skeletal motion-capture methods use an outside-in arrangement of cameras placed around a scene, with viewpoints converging on the center. They often create discomfort with marker suits, and their recording volume is severely restricted and often constrained to indoor scenes with controlled backgrounds. Alternative suit-based systems use several inertial measurement units or an exoskeleton to capture motion with an inside-in setup, i.e. without external sensors. This makes capture independent of a confined volume, but requires substantial, often constraining, and hard to set up body instrumentation. Therefore, we propose a new method for real-time, marker-less, and egocentric motion capture: estimating the full-body skeleton pose from a lightweight stereo pair of fisheye cameras attached to a helmet or virtual reality headset - an optical inside-in method, so to speak. This allows full-body motion capture in general indoor and outdoor scenes, including crowded scenes with many people nearby, which enables reconstruction in larger-scale activities. Our approach combines the strength of a new generative pose estimation framework for fisheye views with a ConvNet-based body-part detector trained on a large new dataset. It is particularly useful in virtual reality to freely roam and interact, while seeing the fully motion-captured virtual body. Helge Rhodin, Christian Richardt, Dan Casas, Eldar Insafutdinov, Mohammad Shafiei, Hans-Peter Seidel, Bernt Schiele, Christian Theobalt |
ACM Trans. Graph. | 4 |