VLDB 2026 Research / reviewers in the wild / expert
Ignas Budvytis
dblp:08/8939
· DBLP profile ↗
36ranked-venue papers
5as first author
20since 2021 · last 2025
0000-0002-4278-198XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NPL-MVPS: Neural Point-Light Multi-View Photometric StereoabstractIn this work we present a novel multi-view photometric stereo (MVPS) method. Like many works in 3D reconstruction we are leveraging neural shape representations and learnt renderers. However, our work differs from the state-of-the-art multi-view PS methods such as PS-NeRF [47] or Supernormal [4] in that we explicitly leverage per-pixel intensity renderings rather than relying mainly on estimated normals. We model point light attenuation and explicitly raytrace cast shadows in order to best approximate the incoming radiance for each point. The estimated incoming radiance is used as input to a fully neural material renderer that uses minimal prior assumptions and it is jointly optimised with the surface. Estimated normals and segmentation maps are also incorporated in order to maximise the surface accuracy. Our method is among the first (along with Supernormal [4]) to outperform the classical MVPS approach proposed by the DiLiGenT-MV benchmark and achieves average 0.2mm Chamfer distance for objects imaged at approx 1.5m distance away with approximate 400 × 400 resolution. Moreover, our method shows high robustness to the sparse MVPS setup (6 views, 6 lights) greatly outperforming the SOTA competitor (0.38mm vs 0.61mm), illustrating the importance of neural rendering in multi-view photometric stereo. Fotios Logothetis, Ignas Budvytis, Roberto Cipolla |
WACV | 2 |
| 2024 | DiaLoc: An Iterative Approach to Embodied Dialog LocalizationabstractMultimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localization task understudied. The few existing dialogbased localization approaches assume the availability of entire dialog prior to Iocalizaiton, which is impractical for deployed dialog-based localization. In this paper, we propose DiaLoc, a new dialog-based localization framework which aligns with a real human operator behavior. Specifically, we produce an iterative refinement of location predictions which can visualize current pose believes after each dialog turn. DiaLoc effectively utilizes the multimodal data for multi-shot localization, where a fusion encoder fuses vision and dialog information iteratively. We achieve state-of-the-art results on embodied dialog-based localization task, in single-shot (+7.08% in Acc5@valUnseen) and multishot settings (+10.85% in Acc5@valUnseen). DiaLoc narrows the gap between simulation and real-world applications, opening doors for future research on collaborative localization and navigation. Chao Zhang 0023, Mohan Li, Ignas Budvytis, Stephan Liwicki |
CVPR | 3 |
| 2024 | A Neural Height-Map Approach for the Binocular Photometric Stereo ProblemabstractIn this work we propose a novel, highly practical, binocular photometric stereo (PS) framework, which has same acquisition speed as single view PS, however significantly improves the quality of the estimated geometry.As in recent neural multi-view shape estimation frameworks such as NeRF [29], SIREN [35] and inverse graphics approaches to multi-view photometric stereo (e.g. PS-NeRF [38]) we formulate shape estimation task as learning of a differentiable surface and texture representation by minimising surface normal discrepancy for normals estimated from multiple varying light images for two views as well as discrepancy between rendered surface intensity and observed images. Our method differs from typical multi-view shape estimation approaches in two key ways. First, our surface is represented not as a volume but as a neural heightmap where heights of points on a surface are computed by a deep neural network. Second, instead of predicting an average intensity as PS-NeRF or introducing lambertian material assumptions as Guo et al. [7], we use a learnt BRDF and perform near-field per point intensity rendering.Our method achieves the state-of-the-art performance on the DiLiGenT-MV dataset adapted to binocular stereo setup as well as a new binocular photometric stereo dataset - LUCES-ST. Fotios Logothetis, Ignas Budvytis, Roberto Cipolla |
WACV | 2 |
| 2023 | Sparse Multi-Object Render-and-Compare
Florian Langer, Ignas Budvytis, Roberto Cipolla |
BMVC | 2 |
| 2023 | HuManiFlow: Ancestor-Conditioned Normalising Flows on SO(3) Manifolds for Human Pose and Shape Distribution EstimationabstractMonocular 3D human pose and shape estimation is an illposed problem since multiple 3D solutions can explain a 2D image of a subject. Recent approaches predict a probability distribution over plausible 3D pose and shape parameters conditioned on the image. We show that these approaches exhibit a trade-off between three key properties: (i) accuracy - the likelihood of the ground-truth 3D solution under the predicted distribution, (ii) sample-input consistency - the extent to which 3D samples from the predicted distribution match the visible 2D image evidence, and (iii) sample diversity - the range of plausible 3D solutions modelled by the predicted distribution. Our method, HuManiFlow, predicts simultaneously accurate, consistent and diverse distributions. We use the human kinematic tree to factorise full body pose into ancestor-conditioned per-body-part pose distributions in an autoregressive manner. Per-body-part distributions are implemented using normalising flows that respect the manifold structure of SO(3), the Lie group of per-body-part poses. We show that ill-posed, but ubiquitous, 3D point estimate losses reduce sample diversity, and employ only probabilistic training losses. HuManiFlow outperforms state-of-the-art probabilistic approaches on the 3DPW and SSP-3D datasets. Akash Sengupta, Ignas Budvytis, Roberto Cipolla |
CVPR | 2 |
| 2023 | SFD2: Semantic-Guided Feature Detection and DescriptionabstractVisual localization is a fundamental task for various applications including autonomous driving and robotics. Prior methods focus on extracting large amounts of often redundant locally reliable features, resulting in limited efficiency and accuracy, especially in large-scale environments under challenging conditions. Instead, we propose to extract globally reliable features by implicitly embedding high-level semantics into both the detection and description processes. Specifically, our semantic-aware detector is able to detect keypoints from reliable regions (e.g. building, traffic lane) and suppress unreliable areas (e.g. sky, car) implicitly instead of relying on explicit semantic labels. This boosts the accuracy of keypoint matching by reducing the number of features sensitive to appearance changes and avoiding the need of additional segmentation networks at test time. Moreover, our descriptors are augmented with semantics and have stronger discriminative ability, providing more inliers at test time. Particularly, experiments on long-term large-scale visual localization Aachen Day Night and RobotCar-Seasons datasets demonstrate that our model outperforms previous local features and gives competitive accuracy to advanced matchers but is about 2 and 3 times faster when using 2k and 4k keypoints, respectively. Code is available at https://github.com/feixue94/sfd2. Ignas Budvytis, Roberto Cipolla |
CVPR | 2 |
| 2023 | IMP: Iterative Matching and Pose Estimation with Adaptive PoolingabstractPrevious methods solve feature matching and pose estimation using a two-stage process by first finding matches and then estimating the pose. As they ignore the geometric relationships between the two tasks, they focus on either improving the quality of matches or filtering potential outliers, leading to limited efficiency or accuracy. In contrast, we propose an iterative matching and pose estimation framework (IMP) leveraging the geometric connections between the two tasks: a few good matches are enough for a roughly accurate pose estimation; a roughly accurate pose can be used to guide the matching by providing geometric constraints. To this end, we implement a geometry-aware recurrent attention-based module which jointly outputs sparse matches and camera poses. Specifically, for each iteration, we first implicitly embed geometric information into the module via a pose-consistency loss, allowing it to predict geometry-aware matches progressively. Second, we introduce an efficient IMP, called EIMP, to dynamically discard keypoints without potential matches, avoiding redundant updating and significantly reducing the quadratic time complexity of attention computation in transformers. Experiments on YFCC100m, Scannet, and Aachen Day-Night datasets demonstrate that the proposed method outperforms previous approaches in terms of accuracy and efficiency. Code is available at https://github.com/feixue94/imp-release Ignas Budvytis, Roberto Cipolla |
CVPR | 2 |
| 2023 | A CNN Based Approach for the Point-Light Photometric Stereo Problem
Fotios Logothetis, Roberto Mecca, Ignas Budvytis, Roberto Cipolla |
Int. J. Comput. Vis. | 3 |
| 2022 | IronDepth: Iterative Refinement of Single-View Depth using Surface Normal and its Uncertainty
Gwangbin Bae, Ignas Budvytis, Roberto Cipolla |
BMVC | 2 |
| 2022 | SPARC: Sparse Render-and-Compare for CAD model alignment in a single RGB Image
Florian Langer, Gwangbin Bae, Ignas Budvytis, Roberto Cipolla |
BMVC | 3 |
| 2022 | Multi-View Depth Estimation by Fusing Single-View Depth Probability with Multi-View GeometryabstractMulti-view depth estimation methods typically require the computation of a multi-view cost-volume, which leads to huge memory consumption and slow inference. Furthermore, multi-view matching can fail for texture-less surfaces, reflective surfaces and moving objects. For such failure modes, single-view depth estimation methods are often more reliable. To this end, we propose MaGNet, a novel framework for fusing single-view depth probability with multi-view geometry, to improve the accuracy, robustness and efficiency of multi-view depth estimation. For each frame, MaGNet estimates a single-view depth probability distribution, parameterized as a pixel-wise Gaussian. The distribution estimated for the reference frame is then used to sample per-pixel depth candidates. Such probabilistic sampling enables the network to achieve higher accuracy while evaluating fewer depth candidates. We also propose depth consistency weighting for the multi-view matching score, to ensure that the multi-view depth is consistent with the single-view predictions. The proposed method achieves state-of-the-art performance on ScanNet [8], 7- Scenes [38] and KITTI [15]. Qualitative evaluation demonstrates that our method is more robust against challenging artifacts such as texture-less/reflective surfaces and moving objects. Our code and model weights are available at https://github.com/baegwangbin/MaGNet. Gwangbin Bae, Ignas Budvytis, Roberto Cipolla |
CVPR | 2 |
| 2022 | Efficient Large-scale Localization by Global Instance RecognitionabstractHierarchical frameworks consisting of both coarse and fine localization are often used as the standard pipeline for large-scale visual localization. Despite their promising performance in simple environments, they still suffer from low efficiency and accuracy in large-scale scenes, especially under challenging conditions. In this paper, we propose an efficient and accurate large-scale localization framework based on the recognition of buildings, which are not only discriminative for coarse localization but also robust for fine localization. Specifically, we assign each building instance a global ID and perform pixel-wise recognition of these global instances in the localization process. For coarse localization, we employ an efficient reference search strategy to find candidates progressively from the local map observing recognized instances instead of the whole database. For fine localization, predicted labels are further used for instance-wise feature detection and matching, allowing our model to focus on fewer but more robust keypoints for establishing correspondences. The experiments in long-term large-scale localization datasets including Aachen and RobotCar-Seasons demonstrate that our method outperforms previous approaches consistently in terms of both efficiency and accuracy. Ignas Budvytis, Daniel Olmeda Reino, Roberto Cipolla |
CVPR | 2 |
| 2021 | Lifted Semantic Graph Embedding for Omnidirectional Place RecognitionabstractTypical place recognition is dependent on the visual appearance and camera position of query images, without explicit use of domain knowledge and geometric relationships between key features in the scene. We exploit semantic grouping of pixels, and camera-pose robust scene graphs to perform structure-based visual localization for place recognition. In particular, we first formulate place recognition as an image retrieval task. Then, we lift the omnidirectional input images into 3D space, and compute a rotation and translation invariant semantic graph embedding to encode query and reference images. Finally, place information is obtained through graph similarity matching. Our graph representation is a simple addition to standard image embeddings with minimal overhead, but contains awareness of objects and their geometric relationships. In our experiments, we show improvement over typical place recognition, especially in environments with repetitions and dynamic appearance changes. Chao Zhang 0023, Ignas Budvytis, Stephan Liwicki, Roberto Cipolla |
3DV | 2 |
| 2021 | Leveraging Geometry for Shape Estimation from a Single RGB Image
Florian Langer, Ignas Budvytis, Roberto Cipolla |
BMVC | 2 |
| 2021 | LUCES: A Dataset for Near-Field Point Light Source Photometric Stereo
Roberto Mecca, Fotios Logothetis, Ignas Budvytis, Roberto Cipolla |
BMVC | 3 |
| 2021 | Probabilistic Estimation of 3D Human Shape and Pose with a Semantic Local Parametric Model
Akash Sengupta, Ignas Budvytis, Roberto Cipolla |
BMVC | 2 |
| 2021 | Probabilistic 3D Human Shape and Pose Estimation From Multiple Unconstrained Images in the WildabstractThis paper addresses the problem of 3D human body shape and pose estimation from RGB images. Recent progress in this field has focused on single images, video or multi-view images as inputs. In contrast, we propose a new task: shape and pose estimation from a group of multiple images of a human subject, without constraints on subject pose, camera viewpoint or background conditions between images in the group. Our solution to this task predicts distributions over SMPL body shape and pose parameters conditioned on the input images in the group. We probabilistically combine predicted body shape distributions from each image to obtain a final multi-image shape prediction. We show that the additional body shape information present in multi-image input groups improves 3D human shape estimation metrics compared to single-image inputs on the SSP-3D dataset and a private dataset of tape-measured humans. In addition, predicting distributions over 3D bodies allows us to quantify pose prediction uncertainty, which is useful when faced with challenging input images with significant occlusion. Our method demonstrates meaningful pose uncertainty on the 3DPW dataset and is competitive with the state-of-the-art in terms of pose estimation metrics. Akash Sengupta, Ignas Budvytis, Roberto Cipolla |
CVPR | 2 |
| 2021 | Estimating and Exploiting the Aleatoric Uncertainty in Surface Normal EstimationabstractSurface normal estimation from a single image is an important task in 3D scene understanding. In this paper, we address two limitations shared by the existing methods: the inability to estimate the aleatoric uncertainty and lack of detail in the prediction. The proposed network estimates the per-pixel surface normal probability distribution. We introduce a new parameterization for the distribution, such that its negative log-likelihood is the angular loss with learned attenuation. The expected value of the angular error is then used as a measure of the aleatoric uncertainty. We also present a novel decoder framework where pixel-wise multi-layer perceptrons are trained on a subset of pixels sampled based on the estimated uncertainty. The proposed uncertainty-guided sampling prevents the bias in training towards large planar surfaces and improves the quality of prediction, especially near object boundaries and on small structures. Experimental results show that the proposed method outperforms the state-of-the-art in ScanNet [4] and NYUv2 [33], and that the estimated uncertainty correlates well with the prediction error. Code is available at https://github.com/baegwangbin/surface_normal_uncertainty. Gwangbin Bae, Ignas Budvytis, Roberto Cipolla |
ICCV | 2 |
| 2021 | PX-NET: Simple and Efficient Pixel-Wise Training of Photometric Stereo NetworksabstractRetrieving accurate 3D reconstructions of objects from the way they reflect light is a very challenging task in computer vision. Despite more than four decades since the definition of the Photometric Stereo problem, most of the literature has had limited success when global illumination effects such as cast shadows, self-reflections and ambient light come into play, especially for specular surfaces. Recent approaches have leveraged the capabilities of deep learning in conjunction with computer graphics in order to cope with the need of a vast number of training data to invert the image irradiance equation and retrieve the geometry of the object. However, rendering global illumination effects is a slow process which can limit the amount of training data that can be generated.In this work we propose a novel pixel-wise training procedure for normal prediction by replacing the training data (observation maps) of globally rendered images with independent per-pixel generated data. We show that global physical effects can be approximated on the observation map domain and this simplifies and speeds up the data creation procedure. Our network, PX-NET, achieves state-of-the-art performance compared to other pixelwise methods on synthetic datasets, as well as the DiLiGenT real dataset on both dense and sparse light settings. Fotios Logothetis, Ignas Budvytis, Roberto Mecca, Roberto Cipolla |
ICCV | 2 |
| 2021 | Hierarchical Kinematic Probability Distributions for 3D Human Shape and Pose Estimation from Images in the WildabstractThis paper addresses the problem of 3D human body shape and pose estimation from an RGB image. This is often an ill-posed problem, since multiple plausible 3D bodies may match the visual evidence present in the input - particularly when the subject is occluded. Thus, it is desirable to estimate a distribution over 3D body shape and pose conditioned on the input image instead of a single 3D re-construction. We train a deep neural network to estimate a hierarchical matrix-Fisher distribution over relative 3D joint rotation matrices (i.e. body pose), which exploits the human body’s kinematic tree structure, as well as a Gaussian distribution over SMPL body shape parameters. To further ensure that the predicted shape and pose distributions match the visual evidence in the input image, we implement a differentiable rejection sampler to impose a reprojection loss between ground-truth 2D joint coordinates and samples from the predicted distributions, projected onto the image plane. We show that our method is competitive with the state-of-the-art in terms of 3D shape and pose metrics on the SSP-3D and 3DPW datasets, while also yielding a structured probability distribution over 3D body shape and pose, with which we can meaningfully quantify prediction uncertainty and sample multiple plausible 3D reconstructions to explain a given input image. Akash Sengupta, Ignas Budvytis, Roberto Cipolla |
ICCV | 2 |
| 2020 | Efficient Large-Scale Semantic Visual Localization in 2D Maps
Tomás Vojír, Ignas Budvytis, Roberto Cipolla |
ACCV (3) | 2 |
| 2020 | Rotation Equivariant Orientation Estimation for Omnidirectional Localization
Chao Zhang 0023, Ignas Budvytis, Stephan Liwicki, Roberto Cipolla |
ACCV (4) | 2 |
| 2020 | A CNN Based Approach for the Near-Field Photometric Stereo Problem
Fotios Logothetis, Ignas Budvytis, Roberto Mecca, Roberto Cipolla |
BMVC | 2 |
| 2020 | Synthetic Training for Accurate 3D Human Pose and Shape Estimation in the Wild
Akash Sengupta, Roberto Cipolla, Ignas Budvytis |
BMVC | 3 |
| 2020 | Deep Multi-view Stereo for Dense 3D Reconstruction from Monocular Endoscopic Video
Gwangbin Bae, Ignas Budvytis, Chung-Kwong Yeung, Roberto Cipolla |
MICCAI (3) | 2 |
| 2019 | Large scale joint semantic re-localisation and scene understanding via globally unique instance coordinate regression
Ignas Budvytis, Marvin Teichmann, Tomás Vojír, Roberto Cipolla |
BMVC | 1 |
| 2018 | Semantic Localisation via Globally Unique Instance Segmentation
Ignas Budvytis, Patrick Sauer, Roberto Cipolla |
BMVC | 1 |
| 2017 | Real-time Factored ConvNets: Extracting the X Factor in Human Parsing
James Charles, Ignas Budvytis, Roberto Cipolla |
BMVC | 2 |
| 2017 | Indirect deep structured learning for 3D human body shape and pose prediction
Vince Tan, Ignas Budvytis, Roberto Cipolla |
BMVC | 2 |
| 2014 | Mixture of Trees Probabilistic Graphical Model for Video Segmentation
Vijay Badrinarayanan, Ignas Budvytis, Roberto Cipolla |
Int. J. Comput. Vis. | 2 |
| 2013 | Semi-Supervised Video Segmentation Using Tree Structured Graphical ModelsabstractWe present a novel patch-based probabilistic graphical model for semi-supervised video segmentation. At the heart of our model is a temporal tree structure that links patches in adjacent frames through the video sequence. This permits exact inference of pixel labels without resorting to traditional short time window-based video processing or instantaneous decision making. The input to our algorithm is labeled key frame(s) of a video sequence and the output is pixel-wise labels along with their confidences. We propose an efficient inference scheme that performs exact inference over the temporal tree, and optionally a per frame label smoothing step using loopy BP, to estimate pixel-wise labels and their posteriors. These posteriors are used to learn pixel unaries by training a Random Decision Forest in a semi-supervised manner. These unaries are used in a second iteration of label inference to improve the segmentation quality. We demonstrate the efficacy of our proposed algorithm using several qualitative and quantitative tests on both foreground/background and multiclass video segmentation problems using publicly available and our own datasets. Vijay Badrinarayanan, Ignas Budvytis, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | MoT - Mixture of Trees Probabilistic Graphical Model for Video SegmentationabstractWe present a novel mixture of trees (MoT) graphical model for video segmentation.Each component in this mixture represents a tree structured temporal linkage between super-pixels from the first to the last frame of a video sequence.Our time-series model explicitly captures the uncertainty in temporal linkage between adjacent frames which improves segmentation accuracy.We provide a variational inference scheme for this model to estimate super-pixel labels and their confidences in nearly realtime.The efficacy of our approach is demonstrated via quantitative comparisons on the challenging SegTrack joint segmentation and tracking dataset [23]. Ignas Budvytis, Vijay Badrinarayanan, Roberto Cipolla |
BMVC | 1 |
| 2012 | Making a Shallow Network Deep: Conversion of a Boosting Classifier into a Decision Tree by Boolean Optimisation
Tae-Kyun Kim 0001, Ignas Budvytis, Roberto Cipolla |
Int. J. Comput. Vis. | 2 |
| 2011 | Semi-supervised video segmentation using tree structured graphical modelsabstractWe present a novel, implementation friendly and occlusion aware semi-supervised video segmentation algorithm using tree structured graphical models, which delivers pixel labels along with their uncertainty estimates. Our motivation to employ supervision is to tackle a task-specific segmentation problem where the semantic objects are pre-defined by the user. The video model we propose for this problem is based on a tree structured approximation of a patch based undirected mixture model, which includes a novel time-series and a soft label Random Forest classifier participating in a feedback mechanism. We demonstrate the efficacy of our model in cutting out foreground objects and multi-class segmentation problems in lengthy and complex road scene sequences. Our results have wide applicability, including harvesting labelled video data for training discriminative models, shape/pose/articulation learning and large scale statistical analysis to develop priors for video segmentation. Ignas Budvytis, Vijay Badrinarayanan, Roberto Cipolla |
CVPR | 1 |
| 2010 | Label propagation in complex video sequences using semi-supervised learningabstractWe propose a novel directed graphical model for label propagation in lengthy and complex video sequences. Given hand-labelled start and end frames of a video sequence, a variational EM based inference strategy propagates either one of several class labels or assigns an unknown class (void) label to each pixel in the video. These labels are used to train a multi-class classifier. The pixel labels estimated by this classifier are injected back into the Bayesian network for another iteration of label inference. The novel aspect of this iterative scheme, as compared to a recent approach [1], is its ability to handle occlusions. This is attributed to a hybrid of generative propagation and discriminative classification in a pseudo time-symmetric video model. The end result is a conservative labelling of the video; large parts of the static scene are labelled into known classes, and a void label is assigned to moving objects and remaining parts of the static scene. These labels can be used as ground truth data to learn the static parts of a scene from videos of it or more generally for semantic video segmentation. We demonstrate the efficacy of the proposed approach using extensive qualitative and quantitative tests over six challenging sequences. We bring out the advantages and drawbacks of our approach, both to encourage its repeatability and motivate future research directions. Ignas Budvytis, Vijay Badrinarayanan, Roberto Cipolla |
BMVC | 1 |
| 2010 | Making a Shallow Network Deep: Growing a Tree from Decision Regions of a Boosting ClassifierabstractThis paper presents a novel way to speed up the classification time of a boosting classifier. We make the shallow (flat) network deep (hierarchical) by growing a tree from the decision regions of a given boosting classifier. This provides many short paths for speeding up and preserves the reasonably smooth decision regions of the boosting classifier for good generalisation. We express the conversion as a Boolean optimisation problem, which has been previously studied for circuit design but limited to a small number of binary variables. In this work, a novel optimisation method is proposed for several tens of variables, i.e. weak-learners of a boosting classifier. The method is then used in a two stage cascade allowing the speed-up of a boosting classifier with any larger number of weak-learners. Experiments on the synthetic and face image data sets show that the obtained tree significantly speeds up both a standard boosting classifier and Fast-exit, a prior-art for fast boosting classification, at the same accuracy. The proposed method as a general meta-algorithm is also shown useful for a boosting cascade, since it speeds up individual stage classifiers by different gains. The proposed method is further demonstrated for rapid object tracking and segmentation problems. Tae-Kyun Kim 0001, Ignas Budvytis, Roberto Cipolla |
BMVC | 2 |