EDBT 2026 Demo / reviewers in the wild / expert
Slobodan Ilic
dblp:76/20
· DBLP profile ↗
109ranked-venue papers
8as first author
21since 2021 · last 2025
0000-0002-3413-1936ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 90 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 74 · 6 first-author · 10 since 2021Systems, architecture and hardware · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SeaLion: Semantic Part-Aware Latent Point Diffusion Models for 3D GenerationabstractDenoising diffusion probabilistic models have achieved significant success in point cloud generation, enabling numerous downstream applications, such as generative data augmentation and 3D model editing. However, little attention has been given to generating point clouds with pointwise segmentation labels, as well as to developing evaluation metrics for this task. Therefore, in this paper, we present SeaLion, a novel diffusion model designed to generate high-quality and diverse point clouds with fine-grained segmentation labels. Specifically, we introduce the semantic part-aware latent point diffusion technique, which leverages the intermediate features of the generative models to jointly predict the noise for perturbed latent points and associated part segmentation labels during the denoising process, and subsequently decodes the latent points to point clouds conditioned on part segmentation labels. To effectively evaluate the quality of generated point clouds, we introduce a novel point cloud pairwise distance calculation method named part-aware Chamfer distance (p-CD). This method enables existing metrics, such as 1-NNA, to measure both the local structural quality and inter-part coherence of generated point clouds. Experiments on the large-scale synthetic dataset ShapeNet and real-world medical dataset IntrA, demonstrate that SeaLion achieves remarkable performance in generation quality and diversity, outperforming the existing state-of-the-art model, DiffFacto, by 13.33% and 6.52% on 1-NNA (p-CD) across the two datasets. Experimental analysis shows that SeaLion can be trained semi-supervised, thereby reducing the demand for labeling efforts. Lastly, we validate the applicability of SeaLion in generative data augmentation for training segmentation models and the capability of SeaLion to serve as a tool for part-aware 3D shape editing. Dekai Zhu, Yan Di, Stefan Gavranovic, Slobodan Ilic |
CVPR | 4 |
| 2025 | RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose EstimationabstractTypical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the correct template often leads to inaccurate pose predictions. To address this, we reformulate template-based object pose estimation as a ray alignment problem, where the viewing directions from multiple posed template images are learned to align with a non-posed query image. Inspired by recent progress in diffusion-based camera pose estimation, we embed this formulation into a diffusion transformer architecture that aligns a query image with a set of posed templates. We reparameterize object rotation using object-centered camera rays and model object translation by extending scale-invariant translation estimation to dense translation offsets. Our model leverages geometric priors from the templates to guide accurate query pose inference. A coarse-to-fine training strategy based on narrowed template sampling improves performance without modifying the network architecture. Extensive experiments across multiple benchmark datasets show competitive results of our method compared to state-of-the-art approaches in unseen object pose estimation. Junwen Huang 0001, Shishir Reddy Vutukur, Peter KT Yu, Nassir Navab, Slobodan Ilic, Benjamin Busam |
ICCV | 5 |
| 2025 | Multi-Layer Feature Exchange Transformer for Multi-View 6D Object Pose Estimation in Robot Bin PickingabstractAccurate 6D object pose estimation is crucial in industrial automation, particularly in robotic bin picking, where objects are often textureless, reflective, and arranged in cluttered environments. Multi-view pose estimation methods offer significant advantages over single-view methods by providing more comprehensive information, effectively handling occlusions and lack of features, and resolving depth ambiguities. However, current multi-view methods often rely on late-stage information fusion, limiting their ability to fully exploit complementary multi-view data. This paper presents a novel approach to enhance multiview 6D pose estimation by introducing a Feature Exchange Transformer (FET) for early-stage feature fusion. This approach leverages self-attention and epipolar cross-attention mechanisms to enable multi-layer feature aggregation across views. Additionally, we introduce a coarse-to-fine strategy for an efficient feature exchange at multiple network layers. Our method, implemented on top of EpiSurfEmb[1], enhances the utilization of multi-view information, leading to significant improvements in pose estimation accuracy and robustness, especially in challenging bin-picking scenarios. We evaluate our approach on the ROBI dataset, demonstrating that it outperforms both the baseline EpiSurfEmb and other state-of-the-art multi-view pose estimation methods. Momen Khalil, Vincent Dietrich, Slobodan Ilic |
ICRA | 3 |
| 2025 | Spiral: Semantic-Aware Progressive LiDAR Scene Generation and UnderstandingabstractLeveraging diffusion models, 3D LiDAR scene generation has achieved great success in both range-view and voxel-based representations. While recent voxel-based approaches can generate both geometric structures and semantic labels, existing range-view methods are limited to producing unlabeled LiDAR scenes. Relying on pretrained segmentation models to predict the semantic maps often results in suboptimal cross-modal consistency. To address this limitation while preserving the advantages of range-view representations, such as computational efficiency and simplified network design, we propose Spiral, a novel range-view LiDAR diffusion model that simultaneously generates depth, reflectance images, and semantic maps. Furthermore, we introduce novel semantic-aware metrics to evaluate the quality of the generated labeled range-view data. Experiments on SemanticKITTI and nuScenes datasets demonstrate that Spiral achieves state-of-the-art performance with the smallest parameter size, outperforming two-step methods that combine the best available generative and segmentation models. Additionally, we validate that Spiral’s generated range images can be effectively used for synthetic data augmentation in the downstream segmentation training, significantly reducing the labeling effort on LiDAR data. Dekai Zhu, Yixuan Hu, Youquan Liu, Dongyue Lu, Lingdong Kong, Slobodan Ilic |
NeurIPS | 6 |
| 2024 | NeRF-Feat: 6D Object Pose Estimation using Feature RenderingabstractObject Pose Estimation is a crucial component in robotic grasping and augmented reality. Learning based approaches typically require training data from a highly accurate CAD model or labeled training data acquired using a complex setup. We address this by learning to estimate pose from weakly labeled data without a known CAD model. We propose to use a NeRF to learn object shape implicitly which is later used to learn view-invariant features in conjunction with CNN using a contrastive loss. While NeRF helps in learning features that are view-consistent, CNN ensures that the learned features respect symmetry. During inference, CNN is used to predict view-invariant features which can be used to establish correspondences with the implicit 3d model in NeRF. The correspondences are then used to estimate the pose in the reference frame of NeRF. Our approach can also handle symmetric objects unlike other approaches using a similar training setup. Specifically, we learn viewpoint invariant, discriminative features using NeRF which are later used for pose estimation. We evaluated our approach on LM, LM-Occlusion, and T-Less dataset and achieved benchmark accuracy despite using weakly labeled data. Shishir Reddy Vutukur, Heike Brock, Benjamin Busam, Tolga Birdal, Andreas Hutter, Slobodan Ilic |
3DV | 6 |
| 2024 | MatchU: Matching Unseen Objects for 6D Pose Estimation from RGB-D ImagesabstractRecent learning methods for object pose estimation require resource-intensive training for each individual object instance or category, hampering their scalability in real applications when confronted with previously unseen objects. In this paper, we propose MatchU, a Fuse-Describe-Match strategy for 6D pose estimation from RGB-D images. MatchU is a generic approach that fuses 2D texture and 3D geometric cues for 6D pose prediction of unseen objects. We rely on learning geometric 3D descriptors that are rotation-invariant by design. By encoding pose-agnostic geometry, the learned descriptors naturally generalize to unseen objects and capture symmetries. To tackle ambiguous associations using 3D geometry only, we fuse additional RGB information into our descriptor. This is achieved through a novel attention-based mechanism that fuses cross-modal information, together with a matching loss that leverages the latent space learned from RGB data to guide the descriptor learning process. Extensive experiments reveal the generalizability of both the RGB-D fusion strategy as well as the descriptor efficacy. Benefiting from the novel designs, MatchU surpasses all existing methods by a significant margin in terms of both accuracy and speed, even without the requirement of expensive re-training or rendering. Junwen Huang 0001, Hao Yu 0010, Kuan-Ting Yu, Nassir Navab, Slobodan Ilic, Benjamin Busam |
CVPR | 5 |
| 2024 | Deep Intra-operative Illumination Calibration of Hyperspectral CamerasabstractAbstract Hyperspectral imaging (HSI) is emerging as a promising novel imaging modality with various potential surgical applications. Currently available cameras, however, suffer from poor integration into the clinical workflow because they require the lights to be switched off, or the camera to be manually recalibrated as soon as lighting conditions change. Given this critical bottleneck, the contribution of this paper is threefold: (1) We demonstrate that dynamically changing lighting conditions in the operating room dramatically affect the performance of HSI applications, namely physiological parameter estimation, and surgical scene segmentation. (2) We propose a novel learning-based approach to automatically recalibrating hyperspectral images during surgery and show that it is sufficiently accurate to replace the tedious process of white reference-based recalibration. (3) Based on a total of 742 HSI cubes from a phantom, porcine models, and rats we show that our recalibration method not only outperforms previously proposed methods, but also generalizes across species, lighting conditions, and image processing tasks. Due to its simple workflow integration as well as high accuracy, speed, and generalization capabilities, our method could evolve as a central component in clinical surgical HSI. Alex Baumann, Leonardo Ayala, Alexander Studier-Fischer, Jan Sellner, Berkin Özdemir, Karl-Friedrich Kowalewski, Slobodan Ilic, Silvia Seidlitz, Lena Maier-Hein |
MICCAI (6) | 7 |
| 2024 | RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud RegistrationabstractSuccessful point cloud registration relies on accurate correspondences established upon powerful descriptors. However, existing neural descriptors either leverage a rotation-variant backbone whose performance declines under large rotations, or encode local geometry that is less distinctive. To address this issue, we introduce RIGA to learn descriptors that are Rotation-Invariant by design and Globally-Aware. From the Point Pair Features (PPFs) of sparse local regions, rotation-invariant local geometry is encoded into geometric descriptors. Global awareness of 3D structures and geometric context is subsequently incorporated, both in a rotation-invariant fashion. More specifically, 3D structures of the whole frame are first represented by our global PPF signatures, from which structural descriptors are learned to help geometric descriptors sense the 3D world beyond local regions. Geometric context from the whole scene is then globally aggregated into descriptors. Finally, the description of sparse regions is interpolated to dense point descriptors, from which correspondences are extracted for registration. To validate our approach, we conduct extensive experiments on both object- and scene-level data. With large rotations, RIGA surpasses the state-of-the-art methods by a margin of 8${}^\circ$in terms of the Relative Rotation Error on ModelNet40 and improves the Feature Matching Recall by at least 5 percentage points on 3DLoMatch. Hao Yu 0010, Ji Hou, Zheng Qin 0002, Mahdi Saleh, Ivan Shugurov, Kai Wang 0037, Benjamin Busam, Slobodan Ilic |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Rotation-Invariant Transformer for Point Cloud MatchingabstractThe intrinsic rotation invariance lies at the core of matching point clouds with handcrafted descriptors. However, it is widely despised by recent deep matchers that obtain the rotation invariance extrinsically via data augmentation. As the finite number of augmented rotations can never span the continuous$SO(3)$space, these methods usually show instability when facing rotations that are rarely seen. To this end, we introduce RoITr, a Rotation-Invariant Transformer to cope with the pose variations in the point cloud matching task. We contribute both on the local and global levels. Starting from the local level, we introduce an attention mechanism embedded with Point Pair Feature (PPF)-based coordinates to describe the pose-invariant geometry, upon which a novel attention-based encoder-decoder architecture is constructed. We further propose a global transformer with rotation-invariant cross-frame spatial awareness learned by the self-attention mechanism, which significantly improves the feature distinctiveness and makes the model robust with respect to the low overlap. Experiments are conducted on both the rigid and non-rigid public benchmarks, where RoITr outperforms all the state-of-the-art models by a considerable margin in the low-overlapping scenarios. Especially when the rotations are enlarged on the challenging 3DLoMatch benchmark, RoITr surpasses the existing methods by at least 13 and 5 percentage points in terms of Inlier Ratio and Registration Recall, respectively. Code is publicly available11https://github.com/haoyu94/RoITr. Hao Yu 0010, Zheng Qin 0002, Ji Hou, Mahdi Saleh, Dongsheng Li 0001, Benjamin Busam, Slobodan Ilic |
CVPR | 7 |
| 2023 | On the Importance of Accurate Geometry Data for Dense 3D Vision TasksabstractLearning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks. These are typically not compared nor discussed in the literature due to a lack of multi-modal datasets. Texture-less regions are problematic for structure from motion and stereo, reflective material poses issues for active sensing, and distances for translucent objects are intricate to measure with existing hardware. Training on inaccurate or corrupt data induces model bias and hampers generalisation capabilities. These effects remain unnoticed if the sensor measurement is considered as ground truth during the evaluation. This paper investigates the effect of sensor errors for the dense 3D vision tasks of depth estimation and reconstruction. We rigorously show the significant impact of sensor characteristics on the learned predictions and notice generalisation issues arising from various technologies in everyday household environments. For evaluation, we introduce a carefully designed dataset11dataset available at https://github.com/Junggy/HAMMER-dataset comprising measurements from commodity sensors, namely D-ToF, I-ToF, passive/active stereo, and monocular RGB+P. Our study quantifies the considerable sensor noise impact and paves the way to improved dense vision estimates and targeted data fusion. Patrick Ruhkamp, Guangyao Zhai, Nikolas Brasch, Yannick Verdie, Jifei Song, Yiren Zhou, Anil Armagan, Slobodan Ilic, Ales Leonardis, Nassir Navab, Benjamin Busam |
CVPR | 10 |
| 2023 | GeoTransformer: Fast and Robust Point Cloud Registration With Geometric TransformerabstractWe study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods have shown great potential through bypassing the detection of repeatable keypoints which is difficult to do especially in low-overlap scenarios. They seek correspondences over downsampled superpoints, which are then propagated to dense points. Superpoints are matched based on whether their neighboring patches overlap. Such sparse and loose matching requires contextual features capturing the geometric structure of the point clouds. We propose Geometric Transformer, or GeoTransformer for short, to learn geometric feature for robust superpoint matching. It encodes pair-wise distances and triplet-wise angles, making it invariant to rigid transformation and robust in low-overlap cases. The simplistic design attains surprisingly high matching accuracy such that no RANSAC is required in the estimation of alignment transformation, leading to 100 times acceleration. Extensive experiments on rich benchmarks encompassing indoor, outdoor, synthetic, multiway and non-rigid demonstrate the efficacy of GeoTransformer. Notably, our method improves the inlier ratio by 18 ∼ 31 percentage points and the registration recall by over 7 points on the challenging 3DLoMatch benchmark. Zheng Qin 0002, Hao Yu 0010, Yulan Guo, Yuxing Peng 0001, Slobodan Ilic, Dewen Hu, Kai Xu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Digital Staining of White Blood Cells With Confidence EstimationabstractChemical staining of the blood smears is one of the crucial components of blood analysis. It is an expensive, lengthy and sensitive process, often prone to produce slight variations in colour and seen structures due to a lack of unified protocols across laboratories. Even though the current developments in deep generative modeling offer an opportunity to replace the chemical process with a digital one, there are specific safety-ensuring requirements due to the severe consequences of mistakes in a medical setting. Therefore digital staining system would profit from an additional confidence estimation quantifying the quality of the digitally stained white blood cell. To this aim, during the staining generation, we disentangle the latent space of the Generative Adversarial Network, obtaining separate representation s of the white blood cell and the staining. We estimate the generated image's confidence of white blood cell structure and staining quality by corrupting these representations with noise and quantifying the information retained between multiple outputs. We show that confidence estimated in this way correlates with image quality measured in terms of LPIPS values calculated for the generated and ground truth stained images. We validate our method by performing digital staining of images captured with a Differential Inference Contrast microscope on a dataset composed of white blood cells of 24 patients. The high absolute value of the correlation between our confidence score and LPIPS demonstrates the effectiveness of our method, opening the possibility of predicting the quality of generated output and ensuring trustworthiness in medical safety-critical setup. Agnieszka Tomczak, Slobodan Ilic, Gaby Marquardt, Thomas Engel 0006, Nassir Navab, Shadi Albarqouni |
IEEE Trans. Medical Imaging | 2 |
| 2022 | OSOP: A Multi-Stage One Shot Object Pose Estimation FrameworkabstractWe present a novel one-shot method for object detection and 6 DoF pose estimation, that does not require training on target objects. At test time, it takes as input a target image and a textured 3D query model. The core idea is to represent a 3D model with a number of 2D templates rendered from different viewpoints. This enables CNN-based direct dense feature extraction and matching. The object is first localized in 2D, then its approximate viewpoint is estimated, followed by dense 2D-3D correspondence prediction. The final pose is computed with PnP. We evaluate the method on LineMOD, Occlusion, Homebrewed, YCB-V and TLESS datasets and report very competitive performance in comparison to the state-of-the-art methods trained on synthetic data, even though our method is not trained on the object models used for testing. Ivan Shugurov, Benjamin Busam, Slobodan Ilic |
CVPR | 4 |
| 2022 | WeLSA: Learning to Predict 6D Pose from Weakly Labeled Data Using Shape Alignment
Shishir Reddy Vutukur, Ivan Shugurov, Benjamin Busam, Andreas Hutter, Slobodan Ilic |
ECCV (8) | 5 |
| 2022 | What Can We Learn About a Generated Image Corrupting Its Latent Representation?
Agnieszka Tomczak, Aarushi Gupta, Slobodan Ilic, Nassir Navab, Shadi Albarqouni |
MICCAI (6) | 3 |
| 2022 | Deep Bingham Networks: Dealing with Uncertainty and Ambiguity in Pose Estimation
Haowen Deng, Mai Bui 0001, Nassir Navab, Leonidas J. Guibas, Slobodan Ilic, Tolga Birdal |
Int. J. Comput. Vis. | 5 |
| 2022 | DPODv2: Dense Correspondence-Based 6 DoF Pose EstimationabstractWe propose a three-stage 6 DoF object detection method called DPODv2 (Dense Pose Object Detector) that relies on dense correspondences. We combine a 2D object detector with a dense correspondence estimation network and a multi-view pose refinement method to estimate a full 6 DoF pose. Unlike other deep learning methods that are typically restricted to monocular RGB images, we propose a unified deep learning network allowing different imaging modalities to be used (RGB or Depth). Moreover, we propose a novel pose refinement method, that is based on differentiable rendering. The main concept is to compare predicted and rendered correspondences in multiple views to obtain a pose which is consistent with predicted correspondences in all views. Our proposed method is evaluated rigorously on different data modalities and types of training data in a controlled setup. The main conclusions is that RGB excels in correspondence estimation, while depth contributes to the pose accuracy if good 3D-3D correspondences are available. Naturally, their combination achieves the overall best performance. We perform an extensive evaluation and an ablation study to analyze and validate the results on several challenging datasets. DPODv2 achieves excellent results on all of them while still remaining fast and scalable independent of the used data modality and the type of training data. Ivan Shugurov, Sergey Zakharov, Slobodan Ilic |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | DistillPose: Lightweight Camera Localization Using Auxiliary LearningabstractWe propose a lightweight retrieval-based pipeline to predict 6DOF camera poses from RGB images. Our pipeline uses a convolutional neural network (CNN) to encode a query image as a feature vector. A nearest neighbor lookup finds the pose-wise nearest database image. A siamese convolutional neural network regresses the relative pose from the nearest neighboring database image to the query image. The relative pose is then applied to the nearest neighboring absolute pose to obtain the query image’s final absolute pose prediction. Our model is a distilled version of NN-Net [1] that reduces its parameters by 98.87%, information retrieval feature vector size by 87.5%, and inference time by 89.18% without a significant decrease in localization accuracy. Yehya Abouelnaga, Mai Bui 0001, Slobodan Ilic |
IROS | 3 |
| 2021 | CoFiNet: Reliable Coarse-to-fine Correspondences for Robust PointCloud RegistrationabstractWe study the problem of extracting correspondences between a pair of point clouds for registration. For correspondence retrieval, existing works benefit from matching sparse keypoints detected from dense points but usually struggle to guarantee their repeatability. To address this issue, we present CoFiNet - Coarse-to-Fine Network which extracts hierarchical correspondences from coarse to fine without keypoint detection. On a coarse scale and guided by a weighting scheme, our model firstly learns to match down-sampled nodes whose vicinity points share more overlap, which significantly shrinks the search space of a consecutive stage. On a finer scale, node proposals are consecutively expanded to patches that consist of groups of points together with associated descriptors. Point correspondences are then refined from the overlap areas of corresponding patches, by a density-adaptive matching module capable to deal with varying point density. Extensive evaluation of CoFiNet on both indoor and outdoor standard benchmarks shows our superiority over existing methods. Especially on 3DLoMatch where point clouds share less overlap, CoFiNet significantly outperforms state-of-the-art approaches by at least 5% on Registration Recall, with at most two-third of their parameters. Hao Yu 0010, Mahdi Saleh, Benjamin Busam, Slobodan Ilic |
NeurIPS | 5 |
| 2021 | Variational Level Set Evolution for Non-Rigid 3D Reconstruction From a Single Depth CameraabstractWe present a framework for real-time 3D reconstruction of non-rigidly moving surfaces captured with a single RGB-D camera. Based on the variational level set method, it warps a given truncated signed distance field (TSDF) to a target TSDF via gradient flow without explicit correspondence search. We optimize an energy that contains a data term which steers towards voxel-wise alignment. To ensure geometrically consistent reconstructions, we develop and compare different strategies, namely an approximately Killing vector field regularizer, gradient flow in Sobolev space and newly devised accelerated optimization. The underlying TSDF evolution makes our approach capable of capturing rapid motions, topological changes and interacting agents, but entails loss of data association. To recover correspondences, we propose to utilize the lowest-frequency Laplacian eigenfunctions of the TSDFs, which encode inherent deformation patterns. For moderate motions we are able to obtain implicit associations via a term that imposes voxel-wise eigenfunction alignment. This is not sufficient for larger motions, so we explicitly estimate voxel correspondences via signature matching of lower-dimensional eigenfunction embeddings. We carry out qualitative and quantitative evaluation of our geometric reconstruction fidelity and voxel correspondence accuracy, demonstrating advantages over related techniques in handling topological changes and fast motions. Miroslava Slavcheva, Maximilian Baust, Slobodan Ilic |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Multi-Task Multi-Domain Learning for Digital Staining and Classification of LeukocytesabstractThis paper addresses digital staining and classification of the unstained white blood cell images obtained with a differential contrast microscope. We have data coming from multiple domains that are partially labeled and partially matching across the domains. Using unstained images removes time-consuming staining procedures and could facilitate and automatize comprehensive diagnostics. To this aim, we propose a method that translates unstained images to realistically looking stained images preserving the inter-cellular structures, crucial for the medical experts to perform classification. We achieve better structure preservation by adding auxiliary tasks of segmentation and direct reconstruction. Segmentation enforces that the network learns to generate correct nucleus and cytoplasm shape, while direct reconstruction enforces reliable translation between the matching images across domains. Besides, we build a robust domain agnostic latent space by injecting the target domain label directly to the generator, i.e., bypassing the encoder. It allows the encoder to extract features independently of the target domain and enables an automated domain invariant classification of the white blood cells. We validated our method on a large dataset composed of leukocytes of 24 patients, achieving state-of-the-art performance on both digital staining and classification tasks. Agnieszka Tomczak, Slobodan Ilic, Gaby Marquardt, Thomas Engel 0006, Frank Forster, Nassir Navab, Shadi Albarqouni |
IEEE Trans. Medical Imaging | 2 |
| 2020 | 3D Object Detection and Pose Estimation of Unseen Objects in Color Images with Local Surface Embeddings
Giorgia Pitteri, Aurélie Bugeau, Slobodan Ilic, Vincent Lepetit |
ACCV (1) | 3 |
| 2020 | 6D Camera Relocalization in Ambiguous Scenes via Continuous Multimodal Inference
Mai Bui 0001, Tolga Birdal, Haowen Deng, Shadi Albarqouni, Leonidas J. Guibas, Slobodan Ilic, Nassir Navab |
ECCV (18) | 6 |
| 2020 | Generic Primitive Detection in Point Clouds Using Novel Minimal Quadric FitsabstractWe present a novel and effective method for detecting 3D primitives in cluttered, unorganized point clouds, without axillary segmentation or type specification. We consider the quadric surfaces for encapsulating the basic building blocks of our environments - planes, spheres, ellipsoids, cones or cylinders, in a unified fashion. Moreover, quadrics allow us to model higher degree of freedom shapes, such as hyperboloids or paraboloids that could be used in non-rigid settings. We begin by contributing two novel quadric fits targeting 3D point sets that are endowed with tangent space information. Based upon the idea of aligning the quadric gradients with the surface normals, our first formulation is exact and requires as low as four oriented points. The second fit approximates the first, and reduces the computational effort. We theoretically analyze these fits with rigor, and give algebraic and geometric arguments. Next, by re-parameterizing the solution, we devise a new local Hough voting scheme on the null-space coefficients that is combined with RANSAC, reducing the complexity from O(N4) to O(N3) (three points). To the best of our knowledge, this is the first method capable of performing a generic cross-type multi-object primitive detection in difficult scenes without segmentation. Our extensive qualitative and quantitative results show that our method is efficient and flexible, as well as being accurate. Tolga Birdal, Benjamin Busam, Nassir Navab, Slobodan Ilic, Peter F. Sturm |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | On Object Symmetries and 6D Pose Estimation from ImagesabstractObjects with symmetries are common in our daily life and in industrial contexts, but are often ignored in the recent literature on 6D pose estimation from images. In this paper, we study in an analytical way the link between the symmetries of a 3D object and its appearance in images. We explain why symmetrical objects can be a challenge when training machine learning algorithms that aim at estimating their 6D pose from images. We propose an efficient and simple solution that relies on the normalization of the pose rotation. Our approach is general and can be used with any 6D pose estimation algorithm. Moreover, our method is also beneficial for objects that are 'almost symmetrical', i.e. objects for which only a detail breaks the symmetry. We validate our approach within a Faster-RCNN framework on a synthetic dataset made with objects from the T-Less dataset, which exhibit various types of symmetries, as well as real sequences from T-Less. Giorgia Pitteri, Michaël Ramamonjisoa, Slobodan Ilic, Vincent Lepetit |
3DV | 3 |
| 2019 | 3D Local Features for Direct Pairwise RegistrationabstractWe present a novel, data driven approach for solving the problem of registration of two point cloud scans. Our approach is direct in the sense that a single pair of corresponding local patches already provides the necessary transformation cue for the global registration. To achieve that, we first endow the state of the art PPF-FoldNet auto-encoder (AE) with a pose-variant sibling, where the discrepancy between the two leads to pose-specific descriptors. Based upon this, we introduce RelativeNet, a relative pose estimation network to assign correspondence-specific orientations to the keypoints, eliminating any local reference frame computations. Finally, we devise a simple yet effective hypothesize-and-verify algorithm to quickly use the predictions and align two point sets. Our extensive quantitative and qualitative experiments suggests that our approach outperforms the state of the art in challenging real datasets of pairwise registration and that augmenting the keypoints with local pose information leads to better generalization and a dramatic speed-up. Haowen Deng, Tolga Birdal, Slobodan Ilic |
CVPR | 3 |
| 2019 | DeceptionNet: Network-Driven Domain RandomizationabstractWe present a novel approach to tackle domain adaptation between synthetic and real data. Instead, of employing "blind" domain randomization, i.e., augmenting synthetic renderings with random backgrounds or changing illumination and colorization, we leverage the task network as its own adversarial guide toward useful augmentations that maximize the uncertainty of the output. To this end, we design a min-max optimization scheme where a given task competes against a special deception network to minimize the task error subject to the specific constraints enforced by the deceiver. The deception network samples from a family of differentiable pixel-level perturbations and exploits the task architecture to find the most destructive augmentations. Unlike GAN-based approaches that require unlabeled data from the target domain, our method achieves robust mappings that scale well to multiple target distributions from source data alone. We apply our framework to the tasks of digit recognition on enhanced MNIST variants, classification and object pose estimation on the Cropped LineMOD dataset as well as semantic segmentation on the Cityscapes dataset and compare it to a number of domain adaptation approaches, thereby demonstrating similar results with superior generalization capabilities. Sergey Zakharov, Wadim Kehl, Slobodan Ilic |
ICCV | 3 |
| 2019 | DPOD: 6D Pose Object Detector and RefinerabstractIn this paper we present a novel deep learning method for 3D object detection and 6D pose estimation from RGB images. Our method, named DPOD (Dense Pose Object Detector), estimates dense multi-class 2D-3D correspondence maps between an input image and available 3D models. Given the correspondences, a 6DoF pose is computed via PnP and RANSAC. An additional RGB pose refinement of the initial pose estimates is performed using a custom deep learning-based refinement scheme. Our results and comparison to a vast number of related works demonstrate that a large number of correspondences is beneficial for obtaining high-quality 6D poses both before and after refinement. Unlike other methods that mainly use real data for training and do not train on synthetic renderings, we perform evaluation on both synthetic and real training data demonstrating superior results before and after refinement when compared to all recent detectors. While being precise, the presented approach is still real-time capable. Sergey Zakharov, Ivan Shugurov, Slobodan Ilic |
ICCV | 3 |
| 2019 | Seeing Beyond Appearance - Mapping Real Images into Geometrical Domains for Unsupervised CAD-based RecognitionabstractWhile convolutional neural networks are dominating the field of computer vision, one usually does not have access to the large amount of domain-relevant data needed for their training. Therefore, it has become common practice to use available synthetic samples along domain adaptation schemes to prepare algorithms for the target domain. Tackling this problem from a different angle, we introduce a pipeline to map unseen target samples into the synthetic domain used to train task-specific methods. Denoising the data and retaining only the features these recognition algorithms are familiar with, our solution greatly improves their performance. As this mapping is easier to learn than the opposite one (i.e., to generate realistic features to augment the source samples), we demonstrate how our whole solution can be trained purely on augmented synthetic data and still performs better than methods trained with domain-relevant information (e.g., real images or realistic textures for the 3D models). Applying our approach to object recognition from texture-less CAD data, we present a custom generative network which fully utilizes the purely geometrical information to learn robust features and to achieve a more refined mapping for unseen color images. Benjamin Planche, Sergey Zakharov, Ziyan Wu 0001, Andreas Hutter, Harald Kosch, Slobodan Ilic |
IROS | 6 |
| 2018 | Survey of Higher Order Rigid Body Motion Interpolation Methods for Keyframe Animation and Continuous-Time Trajectory EstimationabstractIn this survey we carefully analyze the characteristics of higher order rigid body motion interpolation methods to obtain a continuous trajectory from a discrete set of poses. We first discuss the tradeoff between continuity, local control and approximation of classical Euclidean interpolation schemes such as Bezier and B-splines. The benefits of the manifold of unit quaternions SU(2), a double-cover of rotation matrices SO(3), as rotation parameterization are presented, which allow for an elegant formulation of higher order orientation interpolation with easy analytic derivatives, made possible through the Lie Algebra su(2) of pure quaternions and the cumulative form of cubic B-splines. The same construction scheme is then applied for joint interpolation in the full rigid body pose space, which had previously been done for the matrix representation SE(3) and its twists, but not for the more efficient unit dual quaternion DH1 and its screw motions. Both suffer from the effects of coupling translation and rotation that have mostly been ignored by previous work. We thus conclude that split interpolation in ℝ3× SU(2) is preferable for most applications. Our final runtime experiments show that joint interpolation in SE(3) is 2 times and in DH11.3 times slower - which furthermore justifies our suggestion from a practical point of view. Adrian Haarbach, Tolga Birdal, Slobodan Ilic |
3DV | 3 |
| 2018 | Patch-Based Non-rigid 3D Reconstruction from a Single Depth StreamabstractWe propose an approach for 3D reconstruction and tracking of dynamic surfaces using a single depth sensor, without any prior knowledge of the scene. It is robust to rapid inter-frame motions due to the probabilistic expectation-maximization non-rigid registration framework. Our pipeline subdivides each input depth image into non-rigidly connected surface patches, and deforms it towards the canonical pose by estimating a rigid transformation for each patch. The combination of a data term imposing similarity between model and data, and a regularizer enforcing as-rigid-as-possible motion of neighboring patches ensures that we can handle large deformations, while coping with sensor noise. We employ a surfel-based fusion technique, which lets us circumvent the repeated conversion between mesh and signed distance field representations which are used by related techniques. Furthermore, a robust keyframe-based scheme allows us to keep track of correspondences throughout the entire sequence. Through a variety of qualitative and quantitative experiments, we demonstrate resistance to larger motion and achieving lower reconstruction errors than related approaches. Carmel Kozlov, Miroslava Slavcheva, Slobodan Ilic |
3DV | 3 |
| 2018 | Keep it Unreal: Bridging the Realism Gap for 2.5D Recognition with Geometry Priors OnlyabstractWith the increasing availability of large databases of 3D CAD models, methods for depth-based recognition of localized objects can be trained on an uncountable number of synthetically rendered images. However, discrepancies with the real data acquired from various depth sensors still noticeably impede progress. Previous works adopted unsupervised approaches to generate more realistic depth data, but they all require real scans for training, even if unlabeled. This still represents a strong requirement, especially when considering real-life/industrial settings where real training images are hard or impossible to acquire, but texture-less 3D models are available. We thus propose a novel approach leveraging only CAD models to bridge the realism gap. Purely trained on synthetic data, playing against an extensive augmentation pipeline in an unsupervised manner, our generative adversarial network learns to effectively segment depth images and recover the clean synthetic-looking depth information even from partial occlusions. As our solution is not only fully decoupled from the real domains but also from the task-specific analytics, the pre-processed scans can be handed to any kind and number of recognition methods also trained on synthetic data. Through various experiments, we demonstrate how this simplifies their training and consistently enhances their performance, with results on par with the same methods trained on real data, and better than usual approaches doing the reverse mapping. Sergey Zakharov, Benjamin Planche, Ziyan Wu 0001, Andreas Hutter, Harald Kosch, Slobodan Ilic |
3DV | 6 |
| 2018 | Scene Coordinate and Correspondence Learning for Image-Based Localization
Mai Bui 0001, Shadi Albarqouni, Slobodan Ilic, Nassir Navab |
BMVC | 3 |
| 2018 | A Minimalist Approach to Type-Agnostic Detection of Quadrics in Point CloudsabstractThis paper proposes a segmentation-free, automatic and efficient procedure to detect general geometric quadric forms in point clouds, where clutter and occlusions are inevitable. Our everyday world is dominated by man-made objects which are designed using 3D primitives (such as planes, cones, spheres, cylinders, etc.). These objects are also omnipresent in industrial environments. This gives rise to the possibility of abstracting 3D scenes through primitives, thereby positions these geometric forms as an integral part of perception and high level 3D scene understanding. As opposed to state-of-the-art, where a tailored algorithm treats each primitive type separately, we propose to encapsulate all types in a single robust detection procedure. At the center of our approach lies a closed form 3D quadric fit, operating in both primal & dual spaces and requiring as low as 4 oriented-points. Around this fit, we design a novel, local null-space voting strategy to reduce the 4-point case to 3. Voting is coupled with the famous RANSAC and makes our algorithm orders of magnitude faster than its conventional counterparts. This is the first method capable of performing a generic cross-type multi-object primitive detection in difficult scenes. Results on synthetic and real datasets support the validity of our method. Tolga Birdal, Benjamin Busam, Nassir Navab, Slobodan Ilic, Peter F. Sturm |
CVPR | 4 |
| 2018 | PPFNet: Global Context Aware Local Features for Robust 3D Point MatchingabstractWe present PPFNet - Point Pair Feature NETwork for deeply learning a globally informed 3D local feature descriptor to find correspondences in unorganized point clouds. PPFNet learns local descriptors on pure geometry and is highly aware of the global context, an important cue in deep learning. Our 3D representation is computed as a collection of point-pair-features combined with the points and normals within a local vicinity. Our permutation invariant network design is inspired by PointNet and sets PPFNet to be ordering-free. As opposed to voxelization, our method is able to consume raw point clouds to exploit the full sparsity. PPFNet uses a novel N-tuple loss and architecture injecting the global information naturally into the local descriptor. It shows that context awareness also boosts the local feature representation. Qualitative and quantitative evaluations of our network suggest increased recall, improved robustness and invariance as well as a vital step in the 3D descriptor extraction performance. Haowen Deng, Tolga Birdal, Slobodan Ilic |
CVPR | 3 |
| 2018 | SobolevFusion: 3D Reconstruction of Scenes Undergoing Free Non-Rigid MotionabstractWe present a system that builds 3D models of non-rigidly moving surfaces from scratch in real time using a single RGB-D stream. Our solution is based on the variational level set method, thus it copes with arbitrary geometry, including topological changes. It warps a given truncated signed distance field (TSDF) to a target TSDF via gradient flow. Unlike previous approaches that define the gradient using an L2inner product, our method relies on gradient flow in Sobolev space. Its favourable regularity properties allow for a more straightforward energy formulation that is faster to compute and that achieves higher geometric detail, mitigating the over-smoothing effects introduced by other regularization schemes. In addition, the coarse-to-fine evolution behaviour of the flow is able to handle larger motions, making few frames sufficient for a high-fidelity reconstruction. Last but not least, our pipeline determines voxel correspondences between partial shapes by matching signatures in a low-dimensional embedding of their Laplacian eigenfunctions, and is thus able to reliably colour the output model. A variety of quantitative and qualitative evaluations demonstrate the advantages of our technique. Miroslava Slavcheva, Maximilian Baust, Slobodan Ilic |
CVPR | 3 |
| 2018 | PPF-FoldNet: Unsupervised Learning of Rotation Invariant 3D Local Descriptors
Haowen Deng, Tolga Birdal, Slobodan Ilic |
ECCV (5) | 3 |
| 2018 | When Regression Meets Manifold Learning for Object Recognition and Pose EstimationabstractIn this work, we propose a method for object recognition and pose estimation from depth images using convolutional neural networks. Previous methods addressing this problem rely on manifold learning to learn low dimensional viewpoint descriptors and employ them in a nearest neighbor search on an estimated descriptor space. In comparison we create an efficient multi-task learning framework combining manifold descriptor learning and pose regression. By combining the strengths of manifold learning using triplet loss and pose regression, we could either estimate the pose directly reducing the complexity compared to NN search, or use the learned descriptor for the NN descriptor matching. By in depth experimental evaluation of the novel loss function we observed that the view descriptors learned by the network are much more discriminative resulting in almost 30% increase regarding relative pose accuracy compared to related works. On the other hand, regarding directly regressed poses we obtained important improvement compared to simple pose regression. By leveraging the advantages of both manifold learning and regression tasks, we are able to improve the current state-of-the-art for object recognition and pose retrieval. Mai Bui 0001, Sergey Zakharov, Shadi Albarqouni, Slobodan Ilic, Nassir Navab |
ICRA | 4 |
| 2018 | Bayesian Pose Graph Optimization via Bingham Distributions and Tempered Geodesic MCMCabstractWe introduce Tempered Geodesic Markov Chain Monte Carlo (TG-MCMC) algorithm for initializing pose graph optimization problems, arising in various scenarios such as SFM (structure from motion) or SLAM (simultaneous localization and mapping). TG-MCMC is first of its kind as it unites global non-convex optimization on the spherical manifold of quaternions with posterior sampling, in order to provide both reliable initial poses and uncertainty estimates that are informative about the quality of solutions. We devise theoretical convergence guarantees and extensively evaluate our method on synthetic and real benchmarks. Besides its elegance in formulation and theory, we show that our method is robust to missing data, noise and the estimated uncertainties capture intuitive properties of the data. Tolga Birdal, Umut Simsekli, M. Onur Eken, Slobodan Ilic |
NeurIPS | 4 |
| 2018 | SDF-2-SDF Registration for Real-Time 3D Reconstruction from RGB-D Data
Miroslava Slavcheva, Wadim Kehl, Nassir Navab, Slobodan Ilic |
Int. J. Comput. Vis. | 4 |
| 2018 | Almost constant-time 3D nearest-neighbor lookup using implicit octreesabstractA recurring problem in 3D applications is nearest-neighbor lookups in 3D point clouds. In this work, a novel method for exact and approximate 3D nearest-neighbor lookups is proposed that allows lookup times that are, contrary to previous approaches, nearly independent of the distribution of data and query points, allowing to use the method in real-time scenarios. The lookup times of the proposed method outperform prior art sometimes by several orders of magnitude. This speedup is bought at the price of increased costs for creating the indexing structure, which, however, can typically be done in an offline phase. Additionally, an approximate variant of the method is proposed that significantly reduces the time required for data structure creation and further improves lookup times, outperforming all other methods and yielding almost constant lookup times. The method is based on a recursive spatial subdivision using an octree that uses the underlying Voronoi tessellation as splitting criteria, thus avoiding potentially expensive backtracking. The resulting octree is represented implicitly using a hash table, which allows finding the leaf node a query point belongs to with a runtime that is logarithmic in the tree depth. The method is also trivially extendable to 2D nearest neighbor lookups. Bertram Drost, Slobodan Ilic |
Mach. Vis. Appl. | 2 |
| 2018 | Tracking-by-Detection of 3D Human Shapes: From Surfaces to Volumesabstract3D Human shape tracking consists in fitting a template model to temporal sequences of visual observations. It usually comprises an association step, that finds correspondences between the model and the input data, and a deformation step, that fits the model to the observations given correspondences. Most current approaches follow the Iterative-Closest-Point (ICP) paradigm, where the association step is carried out by searching for the nearest neighbors. It fails when large deformations occur and errors in the association tend to propagate over time. In this paper, we propose a discriminative alternative for the association, that leverages random forests to infer correspondences in one shot. Regardless the choice of shape parameterizations, being surface or volumetric meshes, we convert 3D shapes to volumetric distance fields and thereby design features to train the forest. We investigate two ways to draw volumetric samples: voxels of regular grids and cells from Centroidal Voronoi Tessellation (CVT). While the former consumes considerable memory and in turn limits us to learn only subject-specific correspondences, the latter yields much less memory footprint by compactly tessellating the interior space of a shape with optimal discretization. This facilitates the use of larger cross-subject training databases, generalizes to different human subjects and hence results in less overfitting and better detection. The discriminative correspondences are successfully integrated to both surface and volumetric deformation frameworks that recover human shape poses, which we refer to as 'tracking-by-detection of 3D human shapes.' It allows for large deformations and prevents tracking errors from being accumulated. When combined with ICP for refinement, it proves to yield better accuracy in registration and more stability when tracking over time. Evaluations on existing datasets demonstrate the benefits with respect to the state-of-the-art. Chun-Hao P. Huang, Benjamin Allain, Edmond Boyer, Jean-Sébastien Franco, Federico Tombari, Nassir Navab, Slobodan Ilic |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2017 | Real-Time 3D Model Tracking in Color and Depth on a Single CPU CoreabstractWe present a novel method to track 3D models in color and depth data. To this end, we introduce approximations that accelerate the state-of-the-art in region-based tracking by an order of magnitude while retaining similar accuracy. Furthermore, we show how the method can be made more robust in the presence of depth data and consequently formulate a new joint contour and ICP tracking energy. We present better results than the state-of-the-art while being much faster then most other methods and achieving all of the above on a single CPU core. Wadim Kehl, Federico Tombari, Slobodan Ilic, Nassir Navab |
CVPR | 3 |
| 2017 | KillingFusion: Non-rigid 3D Reconstruction without CorrespondencesabstractWe introduce a geometry-driven approach for real-time 3D reconstruction of deforming surfaces from a single RGB-D stream without any templates or shape priors. To this end, we tackle the problem of non-rigid registration by level set evolution without explicit correspondence search. Given a pair of signed distance fields (SDFs) representing the shapes of interest, we estimate a dense deformation field that aligns them. It is defined as a displacement vector field of the same resolution as the SDFs and is determined iteratively via variational minimization. To ensure it generates plausible shapes, we propose a novel regularizer that imposes local rigidity by requiring the deformation to be a smooth and approximately Killing vector field, i.e. generating nearly isometric motions. Moreover, we enforce that the level set property of unity gradient magnitude is preserved over iterations. As a result, KillingFusion reliably reconstructs objects that are undergoing topological changes and fast inter-frame motion. In addition to incrementally building a model from scratch, our system can also deform complete surfaces. We demonstrate these capabilities on several public datasets and introduce our own sequences that permit both qualitative and quantitative comparison to related approaches. Miroslava Slavcheva, Maximilian Baust, Daniel Cremers, Slobodan Ilic |
CVPR | 4 |
| 2017 | CAD Priors for Accurate and Flexible Instance ReconstructionabstractWe present an efficient and automatic approach for accurate instance reconstruction of big 3D objects from multiple, unorganized and unstructured point clouds, in presence of dynamic clutter and occlusions. In contrast to conventional scanning, where the background is assumed to be rather static, we aim at handling dynamic clutter where the background drastically changes during object scanning. Currently, it is tedious to solve this problem with available methods unless the object of interest is first segmented out from the rest of the scene. We address the problem by assuming the availability of a prior CAD model, roughly resembling the object to be reconstructed. This assumption almost always holds in applications such as industrial inspection or reverse engineering. With aid of this prior acting as a proxy, we propose a fully enhanced pipeline, capable of automatically detecting and segmenting the object of interest from scenes and creating a pose graph, online, with linear complexity. This allows initial scan alignment to the CAD model space, which is then refined without the CAD constraint to fully recover a high fidelity 3D reconstruction, accurate up to the sensor noise level. We also contribute a novel object detection method, local implicit shape models (LISM) and give a fast verification scheme. We evaluate our method on multiple datasets, demonstrating the ability to accurately reconstruct objects from small sizes up to 125m3. Tolga Birdal, Slobodan Ilic |
ICCV | 2 |
| 2017 | SSD-6D: Making RGB-Based 3D Detection and 6D Pose Estimation Great AgainabstractWe present a novel method for detecting 3D model instances and estimating their 6D poses from RGB data in a single shot. To this end, we extend the popular SSD paradigm to cover the full 6D pose space and train on synthetic model data only. Our approach competes or surpasses current state-of-the-art methods that leverage RGBD data on multiple challenging datasets. Furthermore, our method produces these results at around 10Hz, which is many times faster than the related methods. For the sake of reproducibility, we make our trained networks and detection code publicly available. Wadim Kehl, Fabian Manhardt, Federico Tombari, Slobodan Ilic, Nassir Navab |
ICCV | 4 |
| 2017 | A point sampling algorithm for 3D matching of irregular geometriesabstractWe present a 3D mesh re-sampling algorithm, carefully tailored for 3D object detection using point pair features (PPF). Computing a sparse representation of objects is critical for the success of state-of-the-art object detection, recognition and pose estimation methods. Yet, sparsity needs to preserve fidelity. To this end, we develop a simple, yet very effective point sampling strategy for detection of any CAD model through geometric hashing. Our approach relies on rendering the object coordinates from a set of views evenly distributed on a sphere. Actual sampling takes place on 2D domain over these renderings; the resulting samples are efficiently merged in 3D with the aid of a special voxel structure and relaxed with Lloyd iterations. The generated vertices are not concentrated only on critical points, as in many keypoint extraction algorithms, and there is even spacing between selected vertices. This is valuable for quantization based detection methods, such as geometric hashing of point pair features. The algorithm is fast and can easily handle the elongated/acute triangles and sharp edges typically existent in industrial CAD models, while automatically pruning the invisible structures. We do not introduce structural changes such as smoothing or interpolation and sample the normals on the original CAD model, achieving the maximum fidelity. We demonstrate the strength of this approach on 3D object detection in comparison to similar sampling algorithms. Tolga Birdal, Slobodan Ilic |
IROS | 2 |
| 2017 | 3D object instance recognition and pose estimation using triplet loss with dynamic marginabstractIn this paper, we address the problem of 3D object instance recognition and pose estimation of localized objects in cluttered environments using convolutional neural networks. Inspired by the descriptor learning approach of Wohlhart et al. [1], we propose a method that introduces the dynamic margin in the manifold learning triplet loss function. Such a loss function is designed to map images of different objects under different poses to a lower-dimensional, similarity-preserving descriptor space on which efficient nearest neighbor search algorithms can be applied. Introducing the dynamic margin allows for faster training times and better accuracy of the resulting low-dimensional manifolds. Furthermore, we contribute the following: adding in-plane rotations (ignored by the baseline method) to the training, proposing new background noise types that help to better mimic realistic scenarios and improve accuracy with respect to clutter, adding surface normals as another powerful image modality representing an object surface leading to better performance than merely depth, and finally implementing an efficient online batch generation that allows for better variability during the training phase. We perform an exhaustive evaluation to demonstrate the effects of our contributions. Additionally, we assess the performance of the algorithm on the large BigBIRD dataset [2] to demonstrate good scalability properties of the pipeline with respect to the number of models. Sergey Zakharov, Wadim Kehl, Benjamin Planche, Andreas Hutter, Slobodan Ilic |
IROS | 5 |
| 2017 | X-Ray PoseNet: 6 DoF Pose Estimation for Mobile X-Ray DevicesabstractPrecise reconstruction of 3D volumes from X-ray projections requires precisely pre-calibrated systems where accurate knowledge of the systems geometric parameters is known ahead. However, when dealing with mobile X-ray devices such calibration parameters are unknown. Joint estimation of the systems calibration parameters and 3d reconstruction is a heavily unconstrained problem, especially when the projections are arbitrary. In industrial applications, that we target here, nominal CAD models of the object to be reconstructed are usually available. We rely on this prior information and employ Deep Learning to learn the mapping between simulated X-ray projections and its pose. Moreover, we introduce the reconstruction loss in addition to the pose loss to further improve the reconstruction quality. Finally, we demonstrate the generalization capabilities of our method in case where poses can be learned on instances of the objects belonging to the same class, allowing pose estimation of unseen objects from the same category, thus eliminating the need for the actual CAD model. We performed exhaustive evaluation demonstrating the quality of our results on both synthetic and real data. Mai Bui 0001, Shadi Albarqouni, Michael Schrapp, Nassir Navab, Slobodan Ilic |
WACV | 5 |
| 2016 | X-Tag: A Fiducial Tag for Flexible and Accurate Bundle AdjustmentabstractIn this paper we design a novel planar 2D fiducial marker and develop fast detection algorithm aiming easy camera calibration and precise 3D reconstruction at the marker locations via the bundle adjustment. Even though an abundance of planar fiducial markers have been made and used in various tasks, none of them has properties necessary to solve the aforementioned tasks. Our marker, X-tag, enjoys a novel design, coupled with very efficient and robust detection scheme, resulting in a reduced number of false positives. This is achieved by constructing markers with random circular features in the image domain and encoding them using two true perspective invariants: cross-ratios and intersection preservation constraints. To detect the markers, we developed an effective search scheme, similar to Geometric Hashing and Hough Voting, in which the marker decoding is cast as a retrieval problem. We apply our system to the task of camera calibration and bundle adjustment. With qualitative and quantitative experiments, we demonstrate the robustness and accuracy of X-tag in spite of blur, noise, perspective and radial distortions, and showcase camera calibration, bundle adjustment and 3d fusion of depth data from precise extrinsic camera poses. Tolga Birdal, Ievgeniia Dobryden, Slobodan Ilic |
3DV | 3 |
| 2016 | An Octree-Based Approach towards Efficient Variational Range Data Fusion
Wadim Kehl, Tobias Holl, Federico Tombari, Slobodan Ilic, Nassir Navab |
BMVC | 4 |
| 2016 | SDF-TAR: Parallel Tracking and Refinement in RGB-D Data using Volumetric Registration
Miroslava Slavcheva, Slobodan Ilic |
BMVC | 2 |
| 2016 | Volumetric 3D Tracking by DetectionabstractIn this paper, we propose a new framework for 3D tracking by detection based on fully volumetric representations. On one hand, 3D tracking by detection has shown robust use in the context of interaction (Kinect) and surface tracking. On the other hand, volumetric representations have recently been proven efficient both for building 3D features and for addressing the 3D tracking problem. We leverage these benefits by unifying both families of approaches into a single, fully volumetric tracking-by-detection framework. We use a centroidal Voronoi tessellation (CVT) representation to compactly tessellate shapes with optimal discretization, construct a feature space, and perform the tracking according to the correspondences provided by trained random forests. Our results show improved tracking and training computational efficiency and improved memory performance. This in turn enables the use of larger training databases than state of the art approaches, which we leverage by proposing a cross-tracking subject training scheme to benefit from all subject sequences for all tracking situations, thus yielding better detection and less overfitting. Chun-Hao Huang, Benjamin Allain, Jean-Sébastien Franco, Nassir Navab, Slobodan Ilic, Edmond Boyer |
CVPR | 5 |
| 2016 | Deep Learning of Local RGB-D Patches for 3D Object Detection and 6D Pose Estimation
Wadim Kehl, Fausto Milletari, Federico Tombari, Slobodan Ilic, Nassir Navab |
ECCV (3) | 4 |
| 2016 | SDF-2-SDF: Highly Accurate 3D Object Reconstruction
Miroslava Slavcheva, Wadim Kehl, Nassir Navab, Slobodan Ilic |
ECCV (1) | 4 |
| 2016 | Online inspection of 3D parts via a locally overlapping camera networkabstractThe raising standards in manufacturing demands reliable and fast industrial quality control mechanisms. This paper proposes an accurate, yet easy to install multi-view, close range optical metrology system, which is suited to online operation. The system is composed of multiple static, locally overlapping cameras forming a network. Initially, these cameras are calibrated to obtain a global coordinate frame. During run-time, the measurements are performed via a novel geometry extraction techniques coupled with an elegant projective registration framework, where 3D to 2D fitting energies are minimized. Finally, a non-linear regression is carried out to compensate for the uncontrollable errors. We apply our pipeline to inspect various geometrical structures found on automobile parts. While presenting the implementation of an involved 3D metrology system, we also demonstrate that the resulting inspection is as accurate as 0.2 mm, repeatable and much faster, compared to the existing methods such as coordinate measurement machines (CMM) or ATOS. Tolga Birdal, Emrah Bala, Tolga Eren, Slobodan Ilic |
WACV | 4 |
| 2016 | Geodesic pixel neighborhoods for 2D and 3D scene understanding
Vladimir Haltakov, Christian Unger, Slobodan Ilic |
Comput. Vis. Image Underst. | 3 |
| 2016 | A Bayesian Approach to Multi-view 4D Modeling
Chun-Hao Huang, Cedric Cagniart, Edmond Boyer, Slobodan Ilic |
Int. J. Comput. Vis. | 4 |
| 2016 | Parsing human skeletons in an operating room
Vasileios Belagiannis, Xinchao Wang, Horesh Ben Shitrit, Kiyoshi Hashimoto, Ralf Stauder, Yoshimitsu Aoki, Michael Kranzfelder, Armin Schneider, Pascal Fua, Slobodan Ilic, Hubertus Feußner, Nassir Navab |
Mach. Vis. Appl. | 10 |
| 2016 | 3D Pictorial Structures Revisited: Multiple Human Pose EstimationabstractWe address the problem of 3D pose estimation of multiple humans from multiple views. The transition from single to multiple human pose estimation and from the 2D to 3D space is challenging due to a much larger state space, occlusions and across-view ambiguities when not knowing the identity of the humans in advance. To address these problems, we first create a reduced state space by triangulation of corresponding pairs of body parts obtained by part detectors for each camera view. In order to resolve ambiguities of wrong and mixed parts of multiple humans after triangulation and also those coming from false positive detections, we introduce a 3D pictorial structures (3DPS) model. Our model builds on multi-view unary potentials, while a prior model is integrated into pairwise and ternary potential functions. To balance the potentials' influence, the model parameters are learnt using a Structured SVM (SSVM). The model is generic and applicable to both single and multiple human pose estimation. To evaluate our model on single and multiple human pose estimation, we rely on four different datasets. We first analyse the contribution of the potentials and then compare our results with related work where we demonstrate superior performance. Vasileios Belagiannis, Sikandar Amin, Mykhaylo Andriluka, Bernt Schiele, Nassir Navab, Slobodan Ilic |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2015 | Point Pair Features Based Object Detection and Pose Estimation RevisitedabstractWe present a revised pipe-line of the existing 3D object detection and pose estimation framework based on point pair feature matching. This framework proposed to represent 3D target object using self-similar point pairs, and then matching such model to 3D scene using efficient Hough-like voting scheme operating on the reduced pose parameter space. Even though this work produces great results and motivated a large number of extensions, it had some general shortcoming like relatively high dimensionality of the search space, sensitivity in establishing 3D correspondences, having performance drops in presence of many outliers and low density surfaces. In this paper, we explain and address these drawbacks and propose new solutions within the existing framework. In particular, we propose to couple the object detection with a coarse-to-fine segmentation, where each segment is subject to disjoint pose estimation. During matching, we apply a weighted Hough voting and an interpolated recovery of pose parameters. Finally, all the generated hypothesis are tested via an occlusion-aware ranking and sorted. We argue that such a combined pipeline simultaneously boosts the detection rate and reduces the complexity, while improving the accuracy of the resulting pose. Thanks to such enhanced pose retrieval, our verification doesn't necessitate ICP and thus achieves better compromise of speed vs. accuracy. We demonstrate our method on existing datasets as well as on our scenes. We conclude that via the new pipe-line, point pair features can now be used in more challenging scenarios. Tolga Birdal, Slobodan Ilic |
3DV | 2 |
| 2015 | Local Hough Transform for 3D Primitive DetectionabstractDetecting primitive geometric shapes, such as cylinders, planes and spheres, in 3D point clouds is an important building block for many high-level vision tasks. One approach for this detection is the Hough Transform, where features vote for parameters that explain them. However, as the voting space grows exponentially with the number of parameters, a full voting scheme quickly becomes impractical. Solutions in the literature, such as decomposing the global voting space, often degrade the robustness w.r.t. Noise, clutter or multiple primitive instances. We instead propose a local Hough Transform, which votes on sub-manifolds of the original parameter space. For this, only those parameters are considered that align a given scene point with the primitive. The voting then recovers the locally best fitting primitive. This local detection scheme is embedded in a coarse-to-fine detection pipeline, which refines the found candidates and removes duplicates. The evaluation shows high robustness against clutter and noise, as well as competitive results w.r.t. Prior art. Bertram Drost, Slobodan Ilic |
3DV | 2 |
| 2015 | Hashmod: A Hashing Method for Scalable 3D Object DetectionabstractWe present a scalable method for detecting objects and estimating their 3D poses in RGB-D data. To this end, we rely on an efficient representation of object views and employ hashing techniques to match these views against the input frame in a scalable way. While a similar approach already exists for 2D detection, we show how to extend it to estimate the 3D pose of the detected objects. In particular, we explore different hashing strategies and identify the one which is more suitable to our problem. We show empirically that the complexity of our method is sublinear with the number of objects and we enable detection and pose estimation of many 3D objects with high accuracy while outperforming the state-of-the-art in terms of runtime. Wadim Kehl, Federico Tombari, Nassir Navab, Slobodan Ilic, Vincent Lepetit |
BMVC | 4 |
| 2015 | Universal Hough dictionaries for object trackingabstractWe propose a novel approach to online visual tracking that combines the robustness of sparse coding with the flexibility of voting-based methods. Our algorithm relies on a dictionary that is learned once and for all from a large set of training patches extracted from images unrelated to the test sequences. In this way we obtain basis functions, also known as atoms, that can be sparsely combined to reconstruct local image content. In order to adapt the generic knowledge encoded in the dictionary to the specific object being tracked, we associate a set of votes and local object appearances to each atom: this is the only information being updated during online tracking. In each frame of the sequence the object's bounding box position is retrieved through a voting strategy. Our method exhibits robustness towards occlusions, sudden local and global illumination changes as well as shape changes. We test our method on 50 standard sequences obtaining results comparable or superior to the state of the art. Fausto Milletari, Wadim Kehl, Federico Tombari, Slobodan Ilic, Seyed-Ahmad Ahmadi, Nassir Navab |
BMVC | 4 |
| 2015 | Toward user-specific tracking by detection of human shapes in multi-camerasabstractHuman shape tracking consists in fitting a template model to temporal sequences of visual observations. It usually comprises an association step, that finds correspondences between the model and the input data, and a deformation step, that fits the model to the observations given correspondences. Most current approaches find their common ground with the Iterative-Closest-Point (ICP) algorithm, which facilitates the association step with local distance considerations. It fails when large deformations occur, and errors in the association tend to propagate over time. In this paper, we propose a discriminative alternative for the association, that leverages random forests to infer correspondences in one shot. It allows for large deformations and prevents tracking errors from accumulating. The approach is successfully integrated to a surface tracking framework that recovers human shapes and poses jointly. When combined with ICP, this discriminative association proves to yield better accuracy in registration, more stability when tracking over time, and faster convergence. Evaluations on existing datasets demonstrate the benefits with respect to the state-of-the-art. Chun-Hao Huang, Edmond Boyer, Bibiana do Canto Angonese, Nassir Navab, Slobodan Ilic |
CVPR | 5 |
| 2015 | A Versatile Learning-Based 3D Temporal Tracker: Scalable, Robust, OnlineabstractThis paper proposes a temporal tracking algorithm based on Random Forest that uses depth images to estimate and track the 3D pose of a rigid object in real-time. Compared to the state of the art aimed at the same goal, our algorithm holds important attributes such as high robustness against holes and occlusion, low computational cost of both learning and tracking stages, and low memory consumption. These are obtained (a) by a novel formulation of the learning strategy, based on a dense sampling of the camera viewpoints and learning independent trees from a single image for each camera view, as well as, (b) by an insightful occlusion handling strategy that enforces the forest to recognize the object's local and global structures. Due to these attributes, we report state-of-the-art tracking accuracy on benchmark datasets, and accomplish remarkable scalability with the number of targets, being able to simultaneously track the pose of over a hundred objects at 30~fps with an off-the-shelf CPU. In addition, the fast learning time enables us to extend our algorithm as a robust online tracker for model-free 3D objects under different viewpoints and appearance changes as demonstrated by the experiments. David Joseph Tan, Federico Tombari, Slobodan Ilic, Nassir Navab |
ICCV | 3 |
| 2015 | Efficient Learning of Linear Predictors for Template Tracking
Stefan Holzer, Slobodan Ilic, David Joseph Tan, Marc Pollefeys, Nassir Navab |
Int. J. Comput. Vis. | 2 |
| 2014 | Extended Co-occurrence HOG with Dense Trajectories for Fine-Grained Activity Recognition
Hirokatsu Kataoka, Kiyoshi Hashimoto, Kenji Iwata, Yutaka Satoh, Nassir Navab, Slobodan Ilic, Yoshimitsu Aoki |
ACCV (5) | 6 |
| 2014 | Geodesic pixel neighborhoods for multi-class image segmentation
Vladimir Haltakov, Christian Unger, Slobodan Ilic |
BMVC | 3 |
| 2014 | Coloured signed distance fields for full 3D object reconstruction
Wadim Kehl, Nassir Navab, Slobodan Ilic |
BMVC | 3 |
| 2014 | Deformable Template Tracking in 1ms
David Joseph Tan, Stefan Holzer, Nassir Navab, Slobodan Ilic |
BMVC | 4 |
| 2014 | A Stochastic Cost Function for Stereo Vision
Christian Unger, Slobodan Ilic |
BMVC | 2 |
| 2014 | 3D Pictorial Structures for Multiple Human Pose EstimationabstractIn this work, we address the problem of 3D pose estimation of multiple humans from multiple views. This is a more challenging problem than single human 3D pose estimation due to the much larger state space, partial occlusions as well as across view ambiguities when not knowing the identity of the humans in advance. To address these problems, we first create a reduced state space by triangulation of corresponding body joints obtained from part detectors in pairs of camera views. In order to resolve the ambiguities of wrong and mixed body parts of multiple humans after triangulation and also those coming from false positive body part detections, we introduce a novel 3D pictorial structures (3DPS) model. Our model infers 3D human body configurations from our reduced state space. The 3DPS model is generic and applicable to both single and multiple human pose estimation. In order to compare to the state-of-the art, we first evaluate our method on single human 3D pose estimation on HumanEva-I [22] and KTH Multiview Football Dataset II [8] datasets. Then, we introduce and evaluate our method on two datasets for multiple human 3D pose estimation. In order to compare to the state-of-the art, we first evaluate our method on single human 3D pose estimation on HumanEva-I [22] and KTH Multiview Football Dataset II [8] datasets. Then, we introduce and evaluate our method on two datasets for multiple human 3D pose estimation. Vasileios Belagiannis, Sikandar Amin, Mykhaylo Andriluka, Bernt Schiele, Nassir Navab, Slobodan Ilic |
CVPR | 6 |
| 2014 | Human Shape and Pose Tracking Using KeyframesabstractThis paper considers human tracking in multi-view setups and investigates a robust strategy that learns online key poses to drive a shape tracking method. The interest arises in realistic dynamic scenes where occlusions or segmentation errors occur. The corrupted observations present missing data and outliers that deteriorate tracking results. We propose to use key poses of the tracked person as multiple reference models. In contrast to many existing approaches that rely on a single reference model, multiple templates represent a larger variability of human poses. They provide therefore better initial hypotheses when tracking with noisy data. Our approach identifies these reference models online as distinctive keyframes during tracking. The most suitable one is then chosen as the reference at each frame. In addition, taking advantage of the proximity between successive frames, an efficient outlier handling technique is proposed to prevent from associating the model to irrelevant outliers. The two strategies are successfully experimented with a surface deformation framework that recovers both the pose and the shape. Evaluations on existing datasets also demonstrate their benefits with respect to the state of the art. Chun-Hao Huang, Edmond Boyer, Nassir Navab, Slobodan Ilic |
CVPR | 4 |
| 2014 | Multi-forest Tracker: A Chameleon in TrackingabstractIn this paper, we address the problem of object tracking in intensity images and depth data. We propose a generic framework that can be used either for tracking 2D templates in intensity images or for tracking 3D objects in depth images. To overcome problems like partial occlusions, strong illumination changes and motion blur, that notoriously make energy minimization-based tracking methods get trapped in a local minimum, we propose a learning-based method that is robust to all these problems. We use random forests to learn the relation between the parameters that defines the object's motion, and the changes they induce on the image intensities or the point cloud of the template. It follows that, to track the template when it moves, we use the changes on the image intensities or point cloud to predict the parameters of this motion. Our algorithm has an extremely fast tracking performance running at less than 2 ms per frame, and is robust to partial occlusions. Moreover, it demonstrates robustness to strong illumination changes when tracking templates using intensity images, and robustness in tracking 3D objects from arbitrary viewpoints even in the presence of motion blur that causes missing or erroneous data in depth images. Extensive experimental evaluation and comparison to the related approaches strongly demonstrates the benefits of our method. David Joseph Tan, Slobodan Ilic |
CVPR | 2 |
| 2014 | Parking assistance using dense motion-stereo - Real-time parking slot detection, collision warning and augmented parking
Christian Unger, Eric Wahl, Slobodan Ilic |
Mach. Vis. Appl. | 3 |
| 2014 | Optimal Local Searching for Fast and Robust Textureless 3D Object Tracking in Highly Cluttered BackgroundsabstractEdge-based tracking is a fast and plausible approach for textureless 3D object tracking, but its robustness is still very challenging in highly cluttered backgrounds due to numerous local minima. To overcome this problem, we propose a novel method for fast and robust textureless 3D object tracking in highly cluttered backgrounds. The proposed method is based on optimal local searching of 3D-2D correspondences between a known 3D object model and 2D scene edges in an image with heavy background clutter. In our searching scheme, searching regions are partitioned into three levels (interior, contour, and exterior) with respect to the previous object region, and confident searching directions are determined by evaluating candidates of correspondences on their region levels; thus, the correspondences are searched among likely candidates in only the confident directions instead of searching through all candidates. To ensure the confident searching direction, we also adopt the region appearance, which is efficiently modeled on a newly defined local space (called a searching bundle). Experimental results and performance evaluations demonstrate that our method fully supports fast and robust textureless 3D object tracking even in highly cluttered backgrounds. Byung-Kuk Seo, Hanhoon Park, Jong-Il Park, Stefan Hinterstoißer, Slobodan Ilic |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2013 | Robust Human Body Shape and Pose TrackingabstractIn this paper we address the problem of marker-less human performance capture from multiple camera videos. We consider in particular the recovery of both shape and parametric motion information as often required in applications that produce and manipulate animated 3D contents using multiple videos. To this aim, we propose an approach that jointly estimates skeleton joint positions and surface deformations by fitting a reference surface model to 3D point reconstructions. The approach is Based on a probabilistic deformable surface registration framework coupled with a bone binding energy. The former makes soft assignments between the model and the observations while the latter guides the skeleton fitting. The main benefit of this strategy lies in its ability to handle outliers and erroneous observations frequently present in multiview data. For the same purpose, we also introduce a learning Based method that partition the point cloud observations into different rigid body parts that further discriminate input data into classes in addition to reducing the complexity of the association between the model and the observations. We argue that such combination of a learning Based matching and of a probabilistic fitting framework efficiently handle unreliable observations with fake geometries or missing data and hence, it reduces the need for tedious manual interventions. A thorough evaluation of the method is presented that includes comparisons with related works on most publicly available multiview datasets. Chun-Hao Huang, Edmond Boyer, Slobodan Ilic |
3DV | 3 |
| 2013 | Multi-task Forest for Human Pose Estimation in Depth ImagesabstractIn this paper, we address the problem of human body pose estimation from depth data. Previous works Based on random forests relied either on a classification strategy to infer the different body parts or on a regression approach to predict directly the joint positions. To permit the inference of very generic poses, those approaches did not consider additional information during the learning phase, e.g. the performed activity. In the present work, we introduce a novel approach to integrate additional information at training time that actually improves the pose prediction during the testing. Our main contribution is a multi-task forest that aims at solving a joint regression-classification task: each foreground pixel from a depth image is associated to its relative displacements to the 3D joint positions as well as the activity class. Integrating activity information in the objective function during forest training permits a better partitioning of the 3D pose space that leads to a better modelling of the posterior. Thereby, our approach provides an improved pose prediction, and as a by-product, can give an estimate of the performed activity. We performed experiments on a dataset performed by 10 people associated with the ground truth body poses from a motion capture system. To demonstrate the benefits of our approach, poses are divided into 10 different activities for the training phase. Results on this dataset show that our multi-task forest provides improved human pose estimation compared to a pure regression forest approach. Joé Lallemand, Olivier Pauly, Loren Arthur Schwarz, David Joseph Tan, Slobodan Ilic |
3DV | 5 |
| 2013 | 3D Semantic Parameterization for Human Shape Modeling: Application to 3D AnimationabstractStatistical human body models, like SCAPE, capture static 3D human body shapes and poses and are applied to many Computer Vision problems. Defined in a statistical context, their parameters do not explicitly capture semantics of the human body shapes such as height, weight, limb length, etc. Having a set of semantic parameters would allow users and automated algorithms to sample the space of possible body shape variations in a more intuitive way. Therefore, in this paper we propose a method for re-parameterization of statistical human body models such that shapes are controlled by a small set of intuitive semantic parameters. These parameters are learned directly from the available statistical human body model. In order to apply any arbitrary animation to our human body shape model we perform retargeting. From any set of 3D scans, a semantic parametrized model can be generated and animated with the presented methods using any animation data. We quantitatively show that our semantic parameterization is more reliable than standard semantic parameterizations, and show a number of animations retargeted to our semantic body shape model. Christian Rupprecht 0001, Olivier Pauly, Christian Theobalt, Slobodan Ilic |
3DV | 4 |
| 2013 | Classification of images in fog and fog-free scenes for use in vehiclesabstractToday modern vehicles are often equipped with a camera, which captures the scene in front of the vehicle. The recognition of weather conditions with this camera can help to improve many applications as well as establish new ones. In this article we will show how it is possible to distinguish between scenes with clear and foggy weather situations. The proposed method uses only gray-scale images as input signal and is running in real time. Using spectral features and a simple linear classifier, we can achieve high detection rates in both daytime and night-time scenes. Furthermore, we will show that in our application area these features outperform others. Mario Pavlic, Gerhard Rigoll, Slobodan Ilic |
Intelligent Vehicles Symposium | 3 |
| 2013 | Multilayer Adaptive Linear Predictors for Real-Time TrackingabstractEnlarging or reducing the template size by adding new parts or removing parts of the template according to their suitability for tracking requires the ability to deal with the variation of the template size. For instance, real-time template tracking using linear predictors, although fast and reliable, requires using templates of a fixed size and does not allow online modification of the predictor. To solve this problem, we propose the Adaptive Linear Predictors (ALPs), which enable fast online modifications of prelearned linear predictors. Instead of applying a full matrix inversion for every modification of the template shape, as standard approaches to learning linear predictors do, we just perform a fast update of this inverse. This allows us to learn the ALPs in a much shorter time than standard learning approaches while performing equally well. Additionally, we propose a multilayer approach to detect occlusions and use ALPs to effectively handle them. This allows us to track large templates and modify them according to the present occlusions. We performed an exhaustive evaluation of our approach and compared it to standard linear predictors and other state-of-the-art approaches. Stefan Holzer, Slobodan Ilic, Nassir Navab |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Model Based Training, Detection and Pose Estimation of Texture-Less 3D Objects in Heavily Cluttered Scenes
Stefan Hinterstoißer, Vincent Lepetit, Slobodan Ilic, Stefan Holzer, Gary R. Bradski, Kurt Konolige, Nassir Navab |
ACCV (1) | 3 |
| 2012 | Efficient Learning of Linear Predictors Using Dimensionality Reduction
Stefan Holzer, Slobodan Ilic, David Joseph Tan, Nassir Navab |
ACCV (3) | 2 |
| 2012 | Segmentation Based Particle Filtering for Real-Time 2D Object Tracking
Vasileios Belagiannis, Falk Schubert, Nassir Navab, Slobodan Ilic |
ECCV (4) | 4 |
| 2012 | Online Learning of Linear Predictors for Real-Time Tracking
Stefan Holzer, Marc Pollefeys, Slobodan Ilic, David Joseph Tan, Nassir Navab |
ECCV (1) | 3 |
| 2012 | Scene understanding from a moving camera for object detection and free space estimationabstractModern vehicles are equipped with multiple cameras which are already used in various practical applications. Advanced driver assistance systems (ADAS) are of particular interest because of the safety and comfort features they offer to the driver. Camera based scene understanding is an important scientific problem that has to be addressed in order to provide the information needed for camera based driver assistance systems. While frontal cameras are widely used, there are applications where cameras observing lateral space can deliver better results. Fish eye cameras mounted in the side mirrors are particularly interesting, because they can observe a big area on the side of the vehicle and can be used for several applications for which the traditional front facing cameras are not suitable. We present a general method for scene understanding using 3D reconstruction of the environment around the vehicle. It is based on pixel-wise image labeling using a conditional random field (CRF). Our method is able to create a simple 3D model of the scene and also to provide semantic labels of the different objects and areas in the image, like for example cars, sidewalks, and buildings. We demonstrate how our method can be used for two applications that are of high importance for various driver assistance systems - car detection and free space estimation. We show that our system is able to perform in real time for speeds of up to 63 km/h. Vladimir Haltakov, Heidrun Belzner, Slobodan Ilic |
Intelligent Vehicles Symposium | 3 |
| 2012 | Image based fog detection in vehiclesabstractModern vehicles are equipped with many cameras and their use in many practical applications is extensive. Detecting the presence of fog from images of a camera mounted in vehicles is a very challenging task with the potential to be used in many practical applications. Approaches introduced until now analyze properties of local objects in the image like lane markings, traffic signs, back lights of vehicles in front or head lights of approaching vehicles. By contrast to all these related works we propose to use image descriptors and a classification procedure in order to distinguish images with fog present from those free of fog. These image descriptors are global and describe the entire image using Gabor filters at different frequencies, scales and orientations. Our experiments demonstrated hight potential of the proposed method for fog detection on daytime images. Mario Pavlic, Heidrun Belzner, Gerhard Rigoll, Slobodan Ilic |
Intelligent Vehicles Symposium | 4 |
| 2012 | Gradient Response Maps for Real-Time Detection of Textureless ObjectsabstractWe present a method for real-time 3D object instance detection that does not require a time-consuming training stage, and can handle untextured objects. At its core, our approach is a novel image representation for template matching designed to be robust to small image transformations. This robustness is based on spread image gradient orientations and allows us to test only a small subset of all possible pixel locations when parsing the image, and to represent a 3D object with a limited set of templates. In addition, we demonstrate that if a dense depth sensor is available we can extend our approach for an even better performance also taking 3D surface normal orientations into account. We show how to take advantage of the architecture of modern computers to build an efficient but very discriminant representation of the input images that can be used to consider thousands of templates in real time. We demonstrate in many experiments on real data that our method is much faster and more robust with respect to background clutter than current state-of-the-art methods. Stefan Hinterstoißer, Cedric Cagniart, Slobodan Ilic, Peter F. Sturm, Nassir Navab, Pascal Fua, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenesabstractWe present a method for detecting 3D objects using multi-modalities. While it is generic, we demonstrate it on the combination of an image and a dense depth map which give complementary object information. It works in real-time, under heavy clutter, does not require a time consuming training stage, and can handle untextured objects. It is based on an efficient representation of templates that capture the different modalities, and we show in many experiments on commodity hardware that our approach significantly outperforms state-of-the-art methods on single modalities. Stefan Hinterstoißer, Stefan Holzer, Cedric Cagniart, Slobodan Ilic, Kurt Konolige, Nassir Navab, Vincent Lepetit |
ICCV | 4 |
| 2011 | RGB-D camera-based parallel tracking and meshingabstractCompared to standard color cameras, RGB-D cameras are designed to additionally provide the depth of imaged pixels which in turn results in a dense colored 3D point cloud representing the environment from a certain viewpoint. We present a real-time tracking method that performs motion estimation of a consumer RGB-D camera with respect to an unknown environment while at the same time reconstructing this environment as a dense textured mesh. Unlike parallel tracking and mapping performed with a standard color or grey scale camera, tracking with an RGB-D camera allows a correctly scaled camera motion estimation. Therefore, there is no need for measuring the environment by any additional tool or equipping the environment by placing objects in it with known sizes. The tracking can be directly started and does not require any preliminary known and/or constrained camera motion. The colored point clouds obtained from every RGB-D image are used to create textured meshes representing the environment from a certain camera view and the real-time estimated camera motion is used to correctly align these meshes over time in order to combine them into a dense reconstruction of the environment. We quantitatively evaluated the proposed method using real image sequences of a challenging scenario and their corresponding ground truth motion obtained with a mechanical measurement arm. We also compared it to a commonly used state-of-the-art method where only the color information is used. We show the superiority of the proposed tracking in terms of accuracy, robustness and usability. We also demonstrate its usage in several Augmented Reality scenarios where the tracking allows a reliable camera motion estimation and the meshing increases the realism of the augmentations by correctly handling their occlusions. Sebastian Lieberknecht, Andrea Huber, Slobodan Ilic, Selim Benhimane |
ISMAR | 3 |
| 2011 | Efficient stereo matching for moving cameras and decalibrated rigsabstractIn vehicular applications based on motion-stereo using monocular side-looking cameras, pairs of images must usually be rectified very well, to allow the application of dense stereo methods. But also long-term installations of stereo rigs in vehicles require approaches that cope with the decalibration of the cameras. The need for such methods is further underlined by the fact that offline camera calibration is a costly and time-consuming procedure at vehicle production sites. In this paper we propose an approach for dense stereo matching that overcomes issues arising from inaccurately rectified images. For this, we significantly increase the search range for correspondences, but still preserve a high efficiency of the method to allow operation on platforms with highly limited processing resources. We demonstrate the performance of our ideas quantitatively using well known stereo datasets and qualitatively using real video sequences of a motion-stereo application. Christian Unger, Eric Wahl, Slobodan Ilic |
Intelligent Vehicles Symposium | 3 |
| 2010 | Free-form mesh tracking: A patch-based approachabstractIn this paper, we consider the problem of tracking nonrigid surfaces and propose a generic data-driven mesh deformation framework. In contrast to methods using strong prior models, this framework assumes little on the observed surface and hence easily generalizes to most free-form surfaces while effectively handling large deformations. To this aim, the reference surface is divided into elementary surface cells or patches. This strategy ensures robustness by providing natural integration domains over the surface for noisy data, while enabling to express simple patch-level rigidity constraints. In addition, we associate to this scheme a robust numerical optimization that solves for physically plausible surface deformations given arbitrary constraints. In order to demonstrate the versatility of the proposed framework, we conducted experiments on open and closed surfaces, with possibly non-connected components, that undergo large deformations and fast motions. We also performed quantitative and qualitative evaluations in multi-cameras and monocular environments, and with different types of data including 2D correspondences and 3D point clouds. Cedric Cagniart, Edmond Boyer, Slobodan Ilic |
CVPR | 3 |
| 2010 | Model globally, match locally: Efficient and robust 3D object recognitionabstractThis paper addresses the problem of recognizing free-form 3D objects in point clouds. Compared to traditional approaches based on point descriptors, which depend on local information around points, we propose a novel method that creates a global model description based on oriented point pair features and matches that model locally using a fast voting scheme. The global model description consists of all model point pair features and represents a mapping from the point pair feature space to the model, where similar features on the model are grouped together. Such representation allows using much sparser object and scene point clouds, resulting in very fast performance. Recognition is done locally using an efficient voting scheme on a reduced two-dimensional search space. We demonstrate the efficiency of our approach and show its high recognition performance in the case of noise, clutter and partial occlusions. Compared to state of the art approaches we achieve better recognition rates, and demonstrate that with a slight or even no sacrifice of the recognition performance our method is much faster then the current state of the art approaches. Bertram Drost, Markus Ulrich, Nassir Navab, Slobodan Ilic |
CVPR | 4 |
| 2010 | Dominant orientation templates for real-time detection of texture-less objectsabstractWe present a method for real-time 3D object detection that does not require a time consuming training stage, and can handle untextured objects. At its core, is a novel template representation that is designed to be robust to small image transformations. This robustness based on dominant gradient orientations lets us test only a small subset of all possible pixel locations when parsing the image, and to represent a 3D object with a limited set of templates. We show that together with a binary representation that makes evaluation very fast and a branch-and-bound approach to efficiently scan the image, it can detect untextured objects in complex situations and provide their 3D pose in real-time. Stefan Hinterstoißer, Vincent Lepetit, Slobodan Ilic, Pascal Fua, Nassir Navab |
CVPR | 3 |
| 2010 | Adaptive linear predictors for real-time trackingabstractEnlarging or reducing the template size by adding new parts, or removing parts of the template, according to their suitability for tracking, requires the ability to deal with the variation of the template size. For instance, real-time template tracking using linear predictors, although fast and reliable, requires using templates of fixed size and does not allow on-line modification of the predictor. To solve this problem we propose the Adaptive Linear Predictors (ALPs) which enable fast online modifications of pre-learned linear predictors. Instead of applying a full matrix inversion for every modification of the template shape as standard approaches to learning linear predictors do, we just perform a fast update of this inverse. This allows us to learn the ALPs in a much shorter time than standard learning approaches while performing equally well. We performed exhaustive evaluation of our approach and compared it to standard linear predictors and other state of the art approaches. Stefan Holzer, Slobodan Ilic, Nassir Navab |
CVPR | 2 |
| 2010 | Probabilistic Deformable Surface Tracking from Multiple Videos
Cedric Cagniart, Edmond Boyer, Slobodan Ilic |
ECCV (4) | 3 |
| 2009 | Distance transform templates for object detection and pose estimationabstractWe propose a new approach for detecting low textured planar objects and estimating their 3D pose. Standard matching and pose estimation techniques often depend on texture and feature points. They fail when there is no or only little texture available. Edge-based approaches mostly can deal with these limitations but are slow in practice when they have to search for six degrees of freedom. We overcome these problems by introducing the distance transform templates, generated by applying the distance transform to standard edge based templates. We obtain robustness against perspective transformations by training a classifier for various template poses. In addition, spatial relations between multiple contours on the template are learnt and later used for outlier removal. At runtime, the classifier provides the identity and a rough 3D pose of the distance transform template, which is further refined by a modified template matching algorithm that is also based on the distance transform. We qualitatively and quantitatively evaluate our approach on synthetic and real-life examples and demonstrate robust real-time performance. Stefan Holzer, Stefan Hinterstoißer, Slobodan Ilic, Nassir Navab |
CVPR | 3 |
| 2008 | Object labeling for recognition using vocabulary treesabstractWe propose an approach to object recognition using vocabulary tree which, instead of finding the closest image in the database to the given query image, finds object labels representing the most similar objects to the query image. We can also recognize object pose if pose labels are associated to the database images. Slobodan Ilic |
ICPR | 1 |
| 2007 | Comparing Timoshenko Beam to Energy Beam for Fitting Noisy Data
Slobodan Ilic |
ACCV (1) | 1 |
| 2007 | Non-Linear Beam Model for Tracking Large DeformationsabstractIn this paper we investigate physics-based plane beam model, frequently used in mechanical and civil engineering, to track large non-linear deformations in images. Such models do not only contribute to robust and precise tracking, in the presence of clutter and partial occlusions, but also allow to compute the forces that produce observed deformations. We verify the correctness of the recovered forces by using them in a simulation and compare the results to the original image displacements. We apply this method to track deformations of the pole vault, the rat whiskers and the car antenna. Slobodan Ilic, Pascal Fua |
ICCV | 1 |
| 2007 | Implicit Meshes for Effective Silhouette Handling
Slobodan Ilic, Mathieu Salzmann, Pascal Fua |
Int. J. Comput. Vis. | 1 |
| 2007 | Surface Deformation Models for Nonrigid 3D Shape RecoveryabstractThree-dimensional detection and shape recovery of a nonrigid surface from video sequences require deformation models to effectively take advantage of potentially noisy image data. Here, we introduce an approach to creating such models for deformable 3D surfaces. We exploit the fact that the shape of an inextensible triangulated mesh can be parameterized in terms of a small subset of the angles between its facets. We use this set of angles to create a representative set of potential shapes, which we feed to a simple dimensionality reduction technique to produce low-dimensional 3D deformation models. We show that these models can be used to accurately model a wide range of deforming 3D surfaces from video sequences acquired under realistic conditions. Mathieu Salzmann, Julien Pilet, Slobodan Ilic, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | Implicit Meshes for Surface ReconstructionabstractDeformable 3D models can be represented either as traditional explicit surfaces, such as triangulated meshes, or as implicit surfaces. Explicit surfaces are widely accepted because they are simple to deform and render, but fitting them involves minimizing a nondifferentiable distance function. By contrast, implicit surfaces allow fitting by minimizing a differentiable algebraic distance, but are harder to meaningfully deform and render. Here, we propose a method that combines the strength of both approaches. It relies on a technique that can turn a completely arbitrary triangulated mesh, such as one taken from the Web, into an implicit surface that closely approximates it and can deform in tandem with it. This allows both automated algorithms to take advantage of the attractive properties of implicit surfaces for fitting purposes and people to use standard deformation tools they feel comfortable for interaction and animation purposes. We demonstrate the applicability of our technique to modeling the human upper-body, including face, neck, shoulders, and ears, from noisy stereo and silhouette data. Slobodan Ilic, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Physically Valid Shape Parameterization for Monocular 3-D Deformable Surface TrackingabstractCVLAB Mathieu Salzmann, Slobodan Ilic, Pascal Fua |
BMVC | 2 |
| 2005 | Implicit Surfaces Make for Better SilhouettesabstractThis paper advocates an implicit-surface representation of generic 3-D surfaces to take advantage of occluding edges in a very robust way. This lets us exploit silhouette constraints in uncontrolled environments that may involve occlusions and changing or cluttered backgrounds, which limit the applicability of most silhouette based methods. This desirable behavior is completely independent from the way the surface deformations are parametrized. To show this, we demonstrate our technique in three very different cases: modeling the deformations of a piece of paper represented by an ordinary triangulated mesh; tracking a person's shoulders whose deformations are expressed in terms of Dirichlet free form deformations; reconstructing the shape of a human face parametrized in terms of a principal component analysis model. Slobodan Ilic, Mathieu Salzmann, Pascal Fua |
CVPR (1) | 1 |
| 2004 | Accurate Face Models from Uncalibrated and Ill-Lit Video Sequences
Miodrag Dimitrijevic, Slobodan Ilic, Pascal Fua |
CVPR (2) | 2 |
| 2003 | Implicit Meshes for Modeling and ReconstructionabstractExplicit surfaces, such as triangulations or wireframe models, have been extensively used to represent the deformable 3D models that are used to fit 3D point and 2D silhouette data. The resulting approaches, however, suffer from the fact that fitting typically involves finding the facets that are closest to the 3D data points or most likely to be silhouette facets. This requires searching, which is slow, and dealing with the non-differentiability of the distance function. By contrast, implicit surface representations allow fitting without search, since one can simply evaluate a differentiable field function at every data point. However, implicit representations are not necessarily the most intuitive ones and users, such as graphics designers, tend to prefer explicit models. Slobodan Ilic, Pascal Fua |
CVPR (2) | 1 |
| 2002 | Using Dirichlet Free Form Deformation to Fit Deformable Models to Noisy 3-D Data
Slobodan Ilic, Pascal Fua |
ECCV (2) | 1 |