Ales Leonardis

dblp:l/AlesLeonardis · DBLP profile ↗
← Back
173ranked-venue papers
14as first author
46since 2021 · last 2026
0000-0003-0773-3277ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 143 · 12 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 98 · 7 first-author · 32 since 2021Systems, architecture and hardware · 14 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Force-Aware 3D Contact Modeling for Stable Grasp Generation
abstract
Contact-based grasp generation plays a crucial role in various applications. Recent methods typically focus on the geometric structure of objects, producing grasps with diverse hand poses and plausible contact points. However, these approaches often overlook the physical attributes of the grasp, specifically the contact force, leading to reduced stability of the grasp. In this paper, we focus on stable grasp generation using explicit contact force predictions. First, we define a force-aware contact representation by transforming the normal force value into discrete levels and encoding it using a one-hot vector. Next, we introduce force-aware stability constraints. We define the stability problem as an acceleration minimization task and explicitly relate stability with contact geometry by formulating the underlying physical constraints. Finally, we present a pose optimizer that systematically integrates our contact representation and stability constraints to enable stable grasp generation. We show that these constraints can help identify key contact points for stability which provide effective initialization and guidance for optimization towards a stable grasp. Experiments are carried out on two public benchmarks, showing that our method brings about 20% improvement in stability metrics and adapts well to novel objects.
Zhuo Chen 0028, Zhongqun Zhang, Yihua Cheng, Ales Leonardis, Hyung Jin Chang
AAAI4
2026 3D Superquadric Splatting
abstract
Gaussian Splatting has proven to be an effective algorithm for novel view synthesis and 3D reconstruction from multi-view images. However, the underlying volumetric primitive—the ellipsoidal Gaussian—has limited expressive capabilities, leading to difficulties in 3D modelling (especially geometry such as edges, corners, and high curvature). To address this limitation, in this paper, we introduce Superquadric Splats (SQS), an extended class of volumetric primitives, as a superset of Gaussian Splats, to model more detailed geometry. We treat superquadrics as volumetric distance functions rather than level-set surfaces. A non-trivial differentiable rendering pipeline is developed to support this. Extensive experimental analysis on multiple datasets validates the effectiveness of the proposed SQS approach, showing both enhanced visual and geometric performance compared to Gaussian-based splatting (with more than 1dB in PSNR and prominent geometric improvement). Project page can be found at: https://daniel-macswayne.github.io/3DSQS/
Daniel MacSwayne, Ales Leonardis, Jianbo Jiao
WACV2
2025 Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
abstract
With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in representing object shape and movement using 3D bounding boxes. Additionally, the reliance on object templates at test time limits their generalisability to unseen objects. To address these challenges, we propose to leverage superquadrics as an alternative 3D object representation to bounding boxes and demonstrate their effectiveness on both template-free object reconstruction and action recognition tasks. Moreover, as we find that pure appearance-based methods can outperform the unified methods, the potential benefits from 3D geometric information remain unclear. Therefore, we study the compositionality of actions by considering a more challenging task where the training combinations of verbs and nouns do not overlap with the testing split. We extend H2O and FPHA datasets with compositional splits and design a novel collaborative learning framework that can explicitly reason about the geometric relations between hands and the manipulated object. Through extensive quantitative and qualitative evaluations, we demonstrate significant improvements over the state-of-the-arts in (compositional) action recognition.
Tze Ho Elden Tse, Runyang Feng, Linfang Zheng, Yixing Gao 0001, Jihie Kim, Ales Leonardis, Hyung Jin Chang
AAAI7
2024 NCRF: Neural Contact Radiance Fields for Free-Viewpoint Rendering of Hand-Object Interaction
abstract
Modeling hand-object interactions is a fundamentally challenging task in 3D computer vision. Despite remarkable progress that has been achieved in this field, existing methods still fail to synthesize the hand-object interaction photo-realistically, suffering from degraded rendering quality caused by the heavy mutual occlusions between the hand and the object, and inaccurate hand-object pose estimation. To tackle these challenges, we present a novel free-viewpoint rendering framework, Neural Contact Radiance Field (NCRF), to reconstruct hand-object interactions from a sparse set of videos. In particular, the proposed NCRF framework consists of two key components: (a) A contact optimization field that predicts an accurate contact field from 3D query points for achieving desirable contact between the hand and the object. (b) A hand-object neural radiance field to learn an implicit hand-object representation in a static canonical space, in concert with the specifically designed hand-object motion field to produce observation-to-canonical correspondences. We jointly learn these key components where they mutually help and regularize each other with visual and geometric constraints, producing a high-quality hand-object reconstruction that achieves photorealistic novel view synthesis. Extensive experiments on HO3D and DexYCB datasets show that our approach outperforms the current state-of-the-art in terms of both rendering quality and pose estimation accuracy.
Zhongqun Zhang, Jifei Song, Eduardo Pérez-Pellitero, Yiren Zhou, Hyung Jin Chang, Ales Leonardis
3DV6
2024 Improving Object Detection via Local-global Contrastive Learning
Danai Triantafyllidou, Sarah Parisot, Ales Leonardis, Steven McDonagh 0001
BMVC3
2024 GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose Refinement
abstract
Object pose refinement is essential for robust object pose estimation. Previous work has made significant progress to-wards instance-level object pose refinement. Yet, category-level pose refinement is a more challenging problem due to large shape variations within a category and the discrep-ancies between the target object and the shape prior. To address these challenges, we introduce a novel architecture for category-level object pose refinement. Our approach in-tegrates an HS-Iayer and learnable affine transformations, which aims to enhance the extraction and alignment of Geometric information. Additionally, we introduce a cross-cloud transformation mechanism that efficiently merges di-verse data sources. Finally, we push the limits of our model by incorporating the shape prior information for translation and size error prediction. We conducted extensive ex-periments to demonstrate the effectiveness of the proposed framework. Through extensive quantitative experiments, we demonstrate significant improvement over the baseline method by a large margin across all metrics.11Project page: https://lynne-zheng-linfang.github.io/georef.github.io
Linfang Zheng, Tze Ho Elden Tse, Chen Wang 0123, Yinghan Sun, Hua Chen 0007, Ales Leonardis, Wei Zhang 0013, Hyung Jin Chang
CVPR6
2024 Multi-task Learning with 3D-Aware Regularization
abstract
Deep neural networks have become the standard solution for designing models that can perform multiple dense computer vision tasks such as depth estimation and semantic segmentation thanks to their ability to capture complex correlations in high dimensional feature space across tasks. However, the cross-task correlations that are learned in the unstructured feature space can be extremely noisy and susceptible to overfitting, consequently hurting performance. We propose to address this problem by introducing a structured 3D-aware regularizer which interfaces multiple tasks through the projection of features extracted from an image encoder to a shared 3D feature space and decodes them into their task output space through differentiable rendering. We show that the proposed method is architecture agnostic and can be plugged into various prior multi-task backbones to improve their performance; as we evidence using standard benchmarks NYUv2 and PASCAL-Context.
Wei-Hong Li 0001, Steven McDonagh 0001, Ales Leonardis, Hakan Bilen
ICLR3
2024 bit2bit: 1-bit quanta video reconstruction via self-supervised photon prediction
abstract
Quanta image sensors, such as SPAD arrays, are an emerging sensor technology, producing 1-bit arrays representing photon detection events over exposures as short as a few nanoseconds. In practice, raw data are post-processed using heavy spatiotemporal binning to create more useful and interpretable images at the cost of degrading spatiotemporal resolution. In this work, we propose bit2bit, a new method for reconstructing high-quality image stacks at the original spatiotemporal resolution from sparse binary quanta image data. Inspired by recent work on Poisson denoising, we developed an algorithm that creates a dense image sequence from sparse binary photon data by predicting the photon arrival location probability distribution. However, due to the binary nature of the data, we show that the assumption of a Poisson distribution is inadequate. Instead, we model the process with a Bernoulli lattice process from the truncated Poisson. This leads to the proposal of a novel self-supervised solution based on a masked loss function. We evaluate our method using both simulated and real data. On simulated data from a conventional video, we achieve 34.35 mean PSNR with extremely photon-sparse binary input (<0.06 photons per pixel per frame). We also present a novel dataset containing a wide range of real SPAD high-speed videos under various challenging imaging conditions. The scenes cover strong/weak ambient light, strong motion, ultra-fast events, etc., which will be made available to the community, on which we demonstrate the promise of our approach. Both reconstruction quality and throughput substantially surpass the state-of-the-art methods (e.g., Quanta Burst Photography (QBP)). Our approach significantly enhances the visualization and usability of the data, enabling the application of existing analysis techniques.
Yehe Liu, Alexander Krull, Hector Basevi, Ales Leonardis, Michael W. Jenkins
NeurIPS4
2024 Image Denoising and the Generative Accumulation of Photons
abstract
We present a fresh perspective on shot noise corrupted images and noise removal. By viewing image formation as the sequential accumulation of photons on a detector grid, we show that a network trained to predict where the next photon could arrive is in fact solving the minimum mean square error (MMSE) denoising task. This new perspective allows us to make three contributions: i. We present a new strategy for self-supervised denoising, ii. We present a new method for sampling from the posterior of possible solutions by iteratively sampling and adding small numbers of photons to the image. iii. We derive a full generative model by starting this process from an empty canvas. We call this approach generative accumulation of photons (GAP). We evaluate our method quantitatively and qualitatively on 4 new fluorescence microscopy datasets, which will be made available to the community. We find that it outperforms its baselines or performs on-par.
Alexander Krull, Hector Basevi, Benjamin Salmon, Andre Zeug, Franziska Müller 0004, Samuel Tonks, Leela Muppala, Ales Leonardis
WACV8
2024 Wavelet-based network for high dynamic range imaging
abstract
High dynamic range (HDR) imaging from multiple low dynamic range (LDR) images has been suffering from ghosting artifacts caused by scene and objects motion. Existing methods, such as optical flow based and end-to-end deep learning based solutions, are error-prone either in detail restoration or ghosting artifacts removal. Comprehensive empirical evidence shows that ghosting artifacts caused by large foreground motion are mainly low-frequency signals and the details are mainly high-frequency signals. In this work, we propose a novel frequency-guided end-to-end deep neural network (FHDRNet) to conduct HDR fusion in the frequency domain, and Discrete Wavelet Transform (DWT) is used to decompose inputs into different frequency bands. The low-frequency signals are used to avoid specific ghosting artifacts, while the high-frequency signals are used for preserving details. Using a U-Net as the backbone, we propose two novel modules: merging module and frequency-guided upsampling module. The merging module applies the attention mechanism to the low-frequency components to deal with the ghost caused by large foreground motion. The frequency-guided upsampling module reconstructs details from multiple frequency-specific components with rich details. In addition, a new RAW dataset is created for training and evaluating multi-frame HDR imaging algorithms in the RAW domain. Extensive experiments are conducted on public datasets and our RAW dataset, showing that the proposed FHDRNet achieves state-of-the-art performance.
Tianhong Dai, Wei Li 0002, Xilei Cao, Jianzhuang Liu, Xu Jia 0012, Ales Leonardis, Youliang Yan, Shanxin Yuan
Comput. Vis. Image Underst.6
2024 Label-efficient object detection via region proposal network pre-training
abstract
Self-supervised pre-training, based on the pretext task of instance discrimination, has fueled the recent advance in label-efficient object detection. However, existing studies focus on pre-training only a feature extractor network to learn transferable representations for downstream detection tasks. This leads to the necessity of training multiple detection-specific modules from scratch in the fine-tuning phase. We argue that the region proposal network (RPN), a common detection-specific module, can additionally be pre-trained towards reducing the localization error of multi-stage detectors. In this work, we propose a simple pretext task that provides an effective pre-training for the RPN, towards efficiently improving downstream object detection performance. We evaluate the efficacy of our approach on benchmark object detection tasks and additional downstream tasks, including instance segmentation and few-shot detection. In comparison with multi-stage detectors without RPN pre-training, our approach is able to consistently improve downstream task performance, with largest gains found in label-scarce settings.
Nanqing Dong, Linus Ericsson, Yongxin Yang, Ales Leonardis, Steven McDonagh 0001
Neurocomputing4
2024 Structure and Intensity Unbiased Translation for 2D Medical Image Segmentation
abstract
Data distribution gaps often pose significant challenges to the use of deep segmentation models. However, retraining models for each distribution is expensive and time-consuming. In clinical contexts, device-embedded algorithms and networks, typically unretrainable and unaccessable post-manufacture, exacerbate this issue. Generative translation methods offer a solution to mitigate the gap by transferring data across domains. However, existing methods mainly focus on intensity distributions while ignoring the gaps due to structure disparities. In this paper, we formulate a new image-to-image translation task to reduce structural gaps. We propose a simple, yet powerful Structure-Unbiased Adversarial (SUA) network which accounts for both intensity and structural differences between the training and test sets for segmentation. It consists of a spatial transformation block followed by an intensity distribution rendering module. The spatial transformation block is proposed to reduce the structural gaps between the two images. The intensity distribution rendering module then renders the deformed structure to an image with the target intensity distribution. Experimental results show that the proposed SUA method has the capability to transfer both intensity distribution and structural content between multiple pairs of datasets and is superior to prior arts in closing the gaps for improving segmentation.
Tianyang Miller, Shaoming Zheng, Jun Cheng 0003, Xi Jia, Joseph Bartlett, Xinxing Cheng, Zhaowen Qiu, Huazhu Fu, Jiang Liu 0001, Ales Leonardis, Jinming Duan 0001
IEEE Trans. Pattern Anal. Mach. Intell.10
2024 Unveiling the Power of Visible-Thermal Video Object Segmentation
abstract
Despite recent progress, Video Object Segmentation (VOS) remains challenging in complex situations such as low light and dark scenes. In this paper, we tackle the visibility limitations by introducing thermal information as auxillary for VOS. Specifically, we generate a hybrid benchmark dataset for Visible-Thermal VOS, named VisT300, which contains 300 challenging videos with visible light and thermal frames and corresponding object mask annotations. Besides, a Visible-Thermal integration Network, named as VTiNet, is proposed to use both cross-modal and cross-frame propagation for accurate video object segmentation. It is advantageous in two aspects: 1) effective cross-modal feature fusion and propagation for strong expressions on visible, thermal, and fused modalities; 2) effective modality-sensitive memory bank enables preserving the most valuable historical contexts in each modality. Extensive experiments demonstrate our VTiNet outperforms the state-of-the-art VOS works by a large margin (over 5% than RGB SotAs in Mean J&F). Our preliminary research clearly recovers that importing complementary modalities can effectively increase the strength of models to achieve robust segmentation in challenging scenarios. Data and code are released at https://github.com/yjybuaa/vtinet, and we hope this work will promote the progress of visible-thermal VOS.
Mingqi Gao 0003, Runmin Cong, Chengjie Wang 0001, Feng Zheng 0001, Ales Leonardis
IEEE Trans. Circuits Syst. Video Technol.6
2024 Weakly-Supervised RGBD Video Object Segmentation
abstract
Depth information opens up new opportunities for video object segmentation (VOS) to be more accurate and robust in complex scenes. However, the RGBD VOS task is largely unexplored due to the expensive collection of RGBD data and time-consuming annotation of segmentation. In this work, we first introduce a new benchmark for RGBD VOS, named DepthVOS, which contains 350 videos (over 55k frames in total) annotated with masks and bounding boxes. We futher propose a novel, strong baseline model - Fused Color-Depth Network (FusedCDNet), which can be trained solely under the supervision of bounding boxes, while being used to generate masks with a bounding box guideline only in the first frame. Thereby, the model possesses three major advantages: a weakly-supervised training strategy to overcome the high-cost annotation, a cross-modal fusion module to handle complex scenes, and weakly-supervised inference to promote ease of use. Extensive experiments demonstrate that our proposed method performs on par with top fully-supervised algorithms. We will open-source our project on https://github.com/yjybuaa/depthvos/ to facilitate the development of RGBD VOS.
Mingqi Gao 0003, Feng Zheng 0001, Xiantong Zhen, Rongrong Ji, Ling Shao 0001, Ales Leonardis
IEEE Trans. Image Process.7
2023 On the Importance of Accurate Geometry Data for Dense 3D Vision Tasks
abstract
Learning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks. These are typically not compared nor discussed in the literature due to a lack of multi-modal datasets. Texture-less regions are problematic for structure from motion and stereo, reflective material poses issues for active sensing, and distances for translucent objects are intricate to measure with existing hardware. Training on inaccurate or corrupt data induces model bias and hampers generalisation capabilities. These effects remain unnoticed if the sensor measurement is considered as ground truth during the evaluation. This paper investigates the effect of sensor errors for the dense 3D vision tasks of depth estimation and reconstruction. We rigorously show the significant impact of sensor characteristics on the learned predictions and notice generalisation issues arising from various technologies in everyday household environments. For evaluation, we introduce a carefully designed dataset11dataset available at https://github.com/Junggy/HAMMER-dataset comprising measurements from commodity sensors, namely D-ToF, I-ToF, passive/active stereo, and monocular RGB+P. Our study quantifies the considerable sensor noise impact and paves the way to improved dense vision estimates and targeted data fusion.
Patrick Ruhkamp, Guangyao Zhai, Nikolas Brasch, Yannick Verdie, Jifei Song, Yiren Zhou, Anil Armagan, Slobodan Ilic, Ales Leonardis, Nassir Navab, Benjamin Busam
CVPR11
2023 Tunable Convolutions with Parametric Multi-Loss Optimization
abstract
Behavior of neural networks is irremediably determined by the specific loss and data used during training. However it is often desirable to tune the model at inference time based on external factors such as preferences of the user or dynamic characteristics of the data. This is especially important to balance the perception-distortion trade-off of ill-posed image-to-image translation tasks. In this work, we propose to optimize a parametric tunable convolutional layer, which includes a number of different kernels, using a parametric multi-loss, which includes an equal number of objectives. Our key insight is to use a shared set of parameters to dynamically interpolate both the objectives and the kernels. During training, these parameters are sampled at random to explicitly optimize all possible combinations of objectives and consequently disentangle their effect into the corresponding kernels. During inference, these parameters become interactive inputs of the model hence enabling reliable and consistent control over the model behavior. Extensive experimental results demonstrate that our tunable convolutions effectively work as a drop-in replacement for traditional convolutions in existing neural networks at virtually no extra computational cost, outperforming state-of-the-art control strategies in a wide range of applications; including image denoising, deblurring, super-resolution, and style transfer.
Matteo Maggioni, Thomas Tanay, Francesca Babiloni, Steven McDonagh 0001, Ales Leonardis
CVPR5
2023 Efficient View Synthesis and 3D-based Multi-Frame Denoising with Multiplane Feature Representations
abstract
While current multi-frame restoration methods combine information from multiple input images using 2D alignment techniques, recent advances in novel view synthesis are paving the way for a new paradigm relying on volu-metric scene representations. In this work, we introduce the first 3D-based multi-frame denoising method that significantly outperforms its 2D-based counterparts with lower computational requirements. Our method extends the mul-tiplane image (MPI) framework for novel view synthesis by introducing a learnable encoder-renderer pair manipulating multiplane representations in feature space. The encoder fuses information across views and operates in a depth-wise manner while the renderer fuses information across depths and operates in a view-wise manner. The two modules are trained end-to-end and learn to separate depths in an unsupervised way, giving rise to Multiplane Feature (MPF) representations. Experiments on the Spaces and Real Forward-Facing datasets as well as on raw burst data validate our approach for view synthesis, multi-frame denoising, and view synthesis under noisy conditions.
Thomas Tanay, Ales Leonardis, Matteo Maggioni
CVPR2
2023 Resource-Efficient RGBD Aerial Tracking
abstract
Aerial robots are now able to fly in complex environments, and drone-captured data gains lots of attention in object tracking. However, current research on aerial perception has mainly focused on limited categories, such as pedestrian or vehicle, and most scenes are captured in urban environments from a birds-eye view. Recently, UAVs equipped with depth cameras have been also deployed for more complex applications, while RGBD aerial tracking is still unexplored. Compared with traditional RGB object tracking, adding depth information can more effectively deal with more challenging scenes such as target and background interference. To this end, in this paper, we explore RGBD aerial tracking in an overhead space, which can greatly enlarge the development of drone-based visual perception. To boost the research, we first propose a large-scale benchmark for RGBD aerial tracking, containing 1,000 drone-captured RGBD videos with dense annotations. Then, as drone-based applications require for real-time processing with limited computational resources, we also propose an efficient RGBD tracker named EMT. Our tracker runs at over 100 fps on GPU, and 25 fps on the edge platform of NVidia Jetson NX Xavier, benefiting from its efficient multimodal fusion and feature matching. Extensive experiments show that our EMT achieves promising tracking performance. All resources are available at https://github.com/yjybuaa/RGBDAerialTracking.
Shang Gao 0012, Zhe Li 0008, Feng Zheng 0001, Ales Leonardis
CVPR5
2023 HS-Pose: Hybrid Scope Feature Extraction for Category-level Object Pose Estimation
abstract
In this paper, we focus on the problem of category-level object pose estimation, which is challenging due to the large intra-category shape variation. 3D graph convolution (3D-GC) based methods have been widely used to extract local geometric features, but they have limitations for complex shaped objects and are sensitive to noise. Moreover, the scale and translation invariant properties of 3D-GC restrict the perception of an object's size and translation information. In this paper, we propose a simple network structure, the HS-layer, which extends 3D-GC to extract hybrid scope latent features from point cloud data for category-level object pose estimation tasks. The proposed HS-layer: 1) is able to perceive local-global geometric structure and global information, 2) is robust to noise, and 3) can encode size and translation information. Our experiments show that the simple replacement of the 3D-GC layer with the proposed HS-layer on the baseline method (GPV-Pose) achieves a significant improvement, with the performance increased by 14.5% on 5°2cm metric and 10.3% on IoU75. Our method outperforms the state-of-the-art methods by a large margin (8.3% on 5°2cm, 6.9% on IoU75) on REAL275 dataset and runs in real-time (50 FPS)11Codeisavailable: https://github.com/Lynne-Zheng-Linfang/HS-Pose.
Linfang Zheng, Chen Wang 0123, Yinghan Sun, Esha Dasgupta, Hua Chen 0007, Ales Leonardis, Wei Zhang 0013, Hyung Jin Chang
CVPR6
2023 Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes
abstract
The success of deep learning models on structured data has generated significant interest in extending their application to non-Euclidean domains. In this work, we introduce a novel intrinsic operator suitable for representation learning on 3D meshes. Our operator is specifically tailored to adapt its behavior to the irregular structure of the underlying graph and effectively utilize its long-range dependencies, while at the same time ensuring computational efficiency and ease of optimization. In particular, inspired by the framework of Spiral Convolution, which extracts and transforms the vertices in the 3D mesh following a local spiral ordering, we propose a general operator that dynamically adjusts the length of the spiral trajectory and the parameters of the transformation for each processed vertex and mesh. Then, we use polyadic decomposition to factorize its dense weight tensor into a sequence of lighter linear layers that separately process features and vertices information, hence significantly reducing the computational complexity without introducing any stringent inductive biases. Notably, we leverage dynamic gating to achieve spatial adaptivity and induce global reasoning with constant time complexity benefitting from an efficient dynamic pooling mechanism based on Summed-Area-tables. Used as a drop-in replacement on existing architectures for shape correspondence our operator significantly improves the performance-efficiency trade-off, and in 3D shape generation with morphable models achieves state-of-the-art performance with a three-fold reduction in the number of parameters required. Project page: https://github.com/Fb2221/DFC
Francesca Babiloni, Matteo Maggioni, Thomas Tanay, Jiankang Deng, Ales Leonardis, Stefanos Zafeiriou
ICCV5
2023 G-DAIC: A Gaze Initialized Framework for Description and Aesthetic-Based Image Cropping
abstract
We propose a new gaze-initialised optimisation framework to generate aesthetically pleasing image crops based on user description. We extended the existing description-based image cropping dataset by collecting user eye movements corresponding to the image captions. To best leverage the contextual information to initialise the optimisation framework using the collected gaze data, this work proposes two gaze-based initialisation strategies, Fixed Grid and Region Proposal. In addition, we propose the adaptive Mixed scaling method to find the optimal output despite the size of the generated initialisation region and the described part of the image. We address the runtime limitation of the state-of-the-art method by implementing the Early termination strategy to reduce the number of iterations required to produce the output. Our experiments show that G-DAIC reduced the runtime by 92.11%, and the quantitative and qualitative experiments demonstrated that the proposed framework produces higher quality and more accurate image crops w.r.t. user intention.
Nora Horanyi, Ales Leonardis, Hyung Jin Chang
Proc. ACM Hum. Comput. Interact.3
2023 Diagnosing and Preventing Instabilities in Recurrent Video Processing
abstract
Recurrent models are a popular choice for video enhancement tasks such as video denoising or super-resolution. In this work, we focus on their stability as dynamical systems and show that they tend to fail catastrophically at inference time on long video sequences. To address this issue, we (1) introduce a diagnostic tool which produces input sequences optimized to trigger instabilities and that can be interpreted as visualizations of temporal receptive fields, and (2) propose two approaches to enforce the stability of a model during training: constraining the spectral norm or constraining the stable rank of its convolutional layers. We then introduce Stable Rank Normalization for Convolutional layers (SRN-C), a new algorithm that enforces these constraints. Our experimental results suggest that SRN-C successfully enforces stablility in recurrent video processing models without a significant performance loss.
Thomas Tanay, Aivar Sootla, Matteo Maggioni, Puneet K. Dokania, Philip Torr 0001, Ales Leonardis, Gregory Slabaugh
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Model-Based Image Signal Processors via Learnable Dictionaries
abstract
Digital cameras transform sensor RAW readings into RGB images by means of their Image Signal Processor (ISP). Computational photography tasks such as image denoising and colour constancy are commonly performed in the RAW domain, in part due to the inherent hardware design, but also due to the appealing simplicity of noise statistics that result from the direct sensor readings. Despite this, the availability of RAW images is limited in comparison with the abundance and diversity of available RGB data. Recent approaches have attempted to bridge this gap by estimating the RGB to RAW mapping: handcrafted model-based methods that are interpretable and controllable usually require manual parameter fine-tuning, while end-to-end learnable neural networks require large amounts of training data, at times with complex training procedures, and generally lack interpretability and parametric control. Towards addressing these existing limitations, we present a novel hybrid model-based and data-driven ISP that builds on canonical ISP operations and is both learnable and interpretable. Our proposed invertible model, capable of bidirectional mapping between RAW and RGB domains, employs end-to-end learning of rich parameter representations, i.e. dictionaries, that are free from direct parametric supervision and additionally enable simple and plausible data augmentation. We evidence the value of our data generation process by extensive experiments under both RAW image reconstruction and RAW image denoising tasks, obtaining state-of-the-art performance in both. Additionally, we show that our ISP can learn meaningful mappings from few data samples, and that denoising models trained with our dictionary-based data augmentation are competitive despite having only few or zero ground-truth labels.
Marcos V. Conde, Steven McDonagh 0001, Matteo Maggioni, Ales Leonardis, Eduardo Pérez-Pellitero
AAAI4
2022 Imagining Hidden Supporting Objects using Volumetric Conditional GANs and Differentiable Stability Scores
Hector Basevi, Ales Leonardis
BMVC2
2022 Disentangling 3D Attributes from a Single 2D Image: Human Pose, Shape and Garment
Xinghui Li, Benjamin Busam, Yiren Zhou, Ales Leonardis, Shanxin Yuan
BMVC5
2022 HDR Reconstruction from Bracketed Exposures and Events
Richard Shaw, Sibi Catley-Chandar, Ales Leonardis, Eduardo Pérez-Pellitero
BMVC3
2022 Content-Diverse Comparisons improve IQA
William Thong, José Costa Pereira, Sarah Parisot, Ales Leonardis, Steven McDonagh 0001
BMVC4
2022 Collaborative Learning for Hand and Object Reconstruction with Attention-guided Graph Convolution
abstract
Estimating the pose and shape of hands and objects under interaction finds numerous applications including aug-mented and virtual reality. Existing approaches for hand and object reconstruction require explicitly defined physical constraints and known objects, which limits its application domains. Our algorithm is agnostic to object models, and it learns the physical rules governing hand-object interaction. This requires automatically inferring the shapes and physi-cal interaction of hands and (potentially unknown) objects. We seek to approach this challenging problem by proposing a collaborative learning strategy where two-branches of deep networks are learning from each other. Specifically, we transfer hand mesh information to the object branch and vice versa for the hand branch. The resulting optimi-sation (training) problem can be unstable, and we address this via two strategies: (i) attention-guided graph convo-lution which helps identify and focus on mutual occlusion and (ii) unsupervised associative loss which facilitates the transfer of information between the branches. Experiments using four widely-used benchmarks show that our frame-work achieves beyond state-of-the-art accuracy in 3D pose estimation, as well as recovers dense 3D hand and object shapes. Each technical component above contributes meaningfully in the ablation study.
Tze Ho Elden Tse, Kwang In Kim, Ales Leonardis, Hyung Jin Chang
CVPR3
2022 CroMo: Cross-Modal Learning for Monocular Depth Estimation
abstract
Learning-based depth estimation has witnessed recent progress in multiple directions; from self-supervision using monocular video to supervised methods offering highest accuracy. Complementary to supervision, further boosts to performance and robustness are gained by combining information from multiple signals. In this paper we systematically investigate key trade-offs associated with sensor and modality design choices as well as related model training strategies. Our study leads us to a new method, capable of connecting modality-specific advantages from polarisation, Time-of-Flight and structured-light inputs. We propose a novel pipeline capable of estimating depth from monocular polarisation for which we evaluate various training signals. The inversion of differentiable analytic models thereby connects scene geometry with polarisation and ToF signals and enables self-supervised and cross-modal learning. In the absence of existing multimodal datasets, we examine our approach with a custom-made multi-modal camera rig and collect CroMo; the first dataset to consist of synchronized stereo polarisation, indirect ToF and structured-light depth, captured at video rates. Extensive experiments on challenging video scenes confirm both qualitative and quantitative pipeline advantages where we are able to outperform competitive monocular depth estimation methods.
Yannick Verdie, Jifei Song, Barnabé Mas, Benjamin Busam, Ales Leonardis, Steven McDonagh 0001
CVPR5
2022 SQN: Weakly-Supervised Semantic Segmentation of Large-Scale 3D Point Clouds
Qingyong Hu, Bo Yang 0027, Guangchi Fang, Yulan Guo, Ales Leonardis, Agathoniki Trigoni, Andrew Markham
ECCV (27)5
2022 S2Contact: Graph-Based Network for 3D Hand-Object Contact Estimation with Semi-supervised Learning
Tze Ho Elden Tse, Zhongqun Zhang, Kwang In Kim, Ales Leonardis, Feng Zheng 0001, Hyung Jin Chang
ECCV (1)4
2022 Towards Generic 3D Tracking in RGBD Videos: Benchmark and Baseline
Zhongqun Zhang, Zhe Li 0008, Hyung Jin Chang, Ales Leonardis, Feng Zheng 0001
ECCV (22)5
2022 TP-AE: Temporally Primed 6D Object Pose Tracking with Auto-Encoders
abstract
Fast and accurate tracking of an object's motion is one of the key functionalities of a robotic system for achieving reliable interaction with the environment. This paper focuses on the instance-level six-dimensional (6D) pose tracking problem with a symmetric and textureless object under occlusion. We propose a Temporally Primed 6D pose tracking framework with Auto-Encoders (TP-AE) to tackle the pose tracking problem. The framework consists of a prediction step and a temporally primed pose estimation step. The prediction step aims to quickly and efficiently generate a guess on the object's real-time pose based on historical information about the target object's motion. Once the prior prediction is obtained, the temporally primed pose estimation step embeds the prior pose into the RGB-D input, and leverages auto-encoders to reconstruct the target object with higher quality under occlusion, thus improving the framework's performance. Extensive experiments show that the proposed 6D pose tracking method can accurately estimate the 6D pose of a symmetric and textureless object under occlusion, and significantly outperforms the state-of-the-art on T-LESS dataset while running in real-time at 26 FPS.
Linfang Zheng, Ales Leonardis, Tze Ho Elden Tse, Nora Horanyi, Hua Chen 0007, Wei Zhang 0013, Hyung Jin Chang
ICRA2
2022 Residual Contrastive Learning for Image Reconstruction: Learning Transferable Representations from Noisy Images
abstract
This paper is concerned with contrastive learning (CL) for low-level image restoration and enhancement tasks. We propose a new label-efficient learning paradigm based on residuals, residual contrastive learning (RCL), and derive an unsupervised visual representation learning framework, suitable for low-level vision tasks with noisy inputs. While supervised image reconstruction aims to minimize residual terms directly, RCL alternatively builds a connection between residuals and CL by defining a novel instance discrimination pretext task, using residuals as the discriminative feature. Our formulation mitigates the severe task misalignment between instance discrimination pretext tasks and downstream image reconstruction tasks, present in existing CL frameworks. Experimentally, we find that RCL can learn robust and transferable representations that improve the performance of various downstream tasks, such as denoising and super resolution, in comparison with recent self-supervised methods designed specifically for noisy inputs. Additionally, our unsupervised pre-training can significantly reduce annotation costs whilst maintaining performance competitive with fully-supervised image reconstruction.
Nanqing Dong, Matteo Maggioni, Yongxin Yang, Eduardo Pérez-Pellitero, Ales Leonardis, Steven McDonagh 0001
IJCAI5
2022 Conditional Patch-Based Domain Randomization: Improving Texture Domain Randomization Using Natural Image Patches
abstract
Using Domain Randomized synthetic data for training deep learning systems is a promising approach for addressing the data and the labeling requirements for supervised techniques to bridge the gap between simulation and the real world. We propose a novel approach for generating and applying class-specific Domain Randomization textures by using randomly cropped image patches from real-world data. In evaluation against the current Domain Randomization texture application techniques, our approach outperforms the highest performing technique by 4.94 AP and 6.71 AP when solving object detection and semantic segmentation tasks on the YCB-M [1] real-world robotics dataset. Our approach is a fast and inexpensive way of generating Domain Randomized textures while avoiding the need to handcraft texture distributions currently being used.
Mohammad Ani, Hector Basevi, Ales Leonardis
IROS3
2022 Prompting for Multi-Modal Tracking
abstract
Multi-modal tracking gains attention due to its ability to be more accurate and robust in complex scenarios compared to traditional RGB-based tracking. Its key lies in how to fuse multi-modal data and reduce the gap between modalities. However, multi-modal tracking still severely suffers from data deficiency, thus resulting in the insufficient learning of fusion modules. Instead of building such a fusion module, in this paper, we provide a new perspective on multi-modal tracking by attaching importance to the multi-modal visual prompts. We design a novel multi-modal prompt tracker (ProTrack), which can transfer the multi-modal inputs to a single modality by the prompt paradigm. By best employing the tracking ability of pre-trained RGB trackers learning at scale, our ProTrack can achieve high-performance multi-modal tracking by only altering the inputs, even without any extra training on multi-modal data. Extensive experiments on 5 benchmark datasets demonstrate the effectiveness of the proposed ProTrack.
Zhe Li 0008, Feng Zheng 0001, Ales Leonardis, Jingkuan Song
ACM Multimedia4
2022 A Continual Learning Survey: Defying Forgetting in Classification Tasks
abstract
Artificial neural networks thrive in solving the classification problem for a particular rigid task, acquiring knowledge through generalized learning behaviour from a distinct training phase. The resulting network resembles a static entity of knowledge, with endeavours to extend this knowledge without targeting the original task resulting in a catastrophic forgetting. Continual learning shifts this paradigm towards networks that can continually accumulate knowledge over different tasks without the need to retrain from scratch. We focus on task incremental classification, where tasks arrive sequentially and are delineated by clear boundaries. Our main contributions concern: (1) a taxonomy and extensive overview of the state-of-the-art; (2) a novel framework to continually determine the stability-plasticity trade-off of the continual learner; (3) a comprehensive experimental comparison of 11 state-of-the-art continual learning methods; and (4) baselines. We empirically scrutinize method strengths and weaknesses on three benchmarks, considering Tiny Imagenet and large-scale unbalanced iNaturalist and a sequence of recognition datasets. We study the influence of model capacity, weight decay and dropout regularization, and the order in which the tasks are presented, and qualitatively compare methods in terms of required memory, computation time, and storage.
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia 0012, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Learning Frequency Domain Priors for Image Demoireing
abstract
Image demoireing is a multi-faceted image restoration task involving both moire pattern removal and color restoration. In this paper, we raise a general degradation model to describe an image contaminated by moire patterns, and propose a novel multi-scale bandpass convolutional neural network (MBCNN) for single image demoireing. For moire pattern removal, we propose a multi-block-size learnable bandpass filters (M-LBFs), based on a block-wise frequency domain transform, to learn the frequency domain priors of moire patterns. We also introduce a new loss function named Dilated Advanced Sobel loss (D-ASL) to better sense the frequency information. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, and then performs local fine tuning of the color per pixel. To determine the most appropriate frequency domain transform, we investigate several transforms including DCT, DFT, DWT, learnable non-linear transform and learnable orthogonal transform. We finally adopt the DCT. Our basic model won the AIM2019 demoireing challenge. Experimental results on three public datasets show that our method outperforms state-of-the-art methods by a large margin.
Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Xiang Tian 0002, Jiyong Zhang 0001, Yaoqi Sun, Lin Liu 0016, Ales Leonardis, Gregory Slabaugh
IEEE Trans. Pattern Anal. Mach. Intell.8
2022 Repurposing existing deep networks for caption and aesthetic-guided image cropping
Nora Horanyi, Kedi Xia, Kwang Moo Yi, Abhishake Kumar Bojja, Ales Leonardis, Hyung Jin Chang
Pattern Recognit.5
2022 FlexHDR: Modeling Alignment and Exposure Uncertainties for Flexible HDR Imaging
abstract
High dynamic range (HDR) imaging is of fundamental importance in modern digital photography pipelines and used to produce a high-quality photograph with well exposed regions despite varying illumination across the image. This is typically achieved by merging multiple low dynamic range (LDR) images taken at different exposures. However, over-exposed regions and misalignment errors due to poorly compensated motion result in artefacts such as ghosting. In this paper, we present a new HDR imaging technique that specifically models alignment and exposure uncertainties to produce high quality HDR results. We introduce a strategy that learns to jointly align and assess the alignment and exposure reliability using an HDR-aware, uncertainty-driven attention map that robustly merges the frames into a single high quality HDR image. Further, we introduce a progressive, multi-stage image fusion approach that can flexibly merge any number of LDR images in a permutation-invariant manner. Experimental results show our method can produce better quality HDR images with up to 1.1dB PSNR improvement to the state-of-the-art, and subjective improvements in terms of better detail, colours, and fewer artefacts.
Sibi Catley-Chandar, Thomas Tanay, Lucas Vandroux, Ales Leonardis, Gregory Slabaugh, Eduardo Pérez-Pellitero
IEEE Trans. Image Process.4
2022 Learning a Model-Driven Variational Network for Deformable Image Registration
abstract
Data-driven deep learning approaches to image registration can be less accurate than conventional iterative approaches, especially when training data is limited. To address this issue and meanwhile retain the fast inference speed of deep learning, we propose VR-Net, a novel cascaded variational network for unsupervised deformable image registration. Using a variable splitting optimization scheme, we first convert the image registration problem, established in a generic variational framework, into two sub-problems, one with a point-wise, closed-form solution and the other one being a denoising problem. We then propose two neural layers (i.e. warping layer and intensity consistency layer) to model the analytical solution and a residual U-Net (termed generalized denoising layer) to formulate the denoising problem. Finally, we cascade the three neural layers multiple times to form our VR-Net. Extensive experiments on three (two 2D and one 3D) cardiac magnetic resonance imaging datasets show that VR-Net outperforms state-of-the-art deep learning methods on registration accuracy, whilst maintaining the fast inference speed of deep learning and the data-efficiency of variational models.
Xi Jia, Alexander Thorley, Wei Chen 0092, Huaqi Qiu, LinLin Shen, Iain B. Styles, Hyung Jin Chang, Ales Leonardis, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Daniel Rueckert, Jinming Duan 0001
IEEE Trans. Medical Imaging8
2021 Wild ToFu: Improving Range and Quality of Indirect Time-of-Flight Depth with RGB Fusion in Challenging Environments
abstract
Indirect Time-of-Flight (I-ToF) imaging is a widespread way of depth estimation for mobile devices due to its small size and affordable price. Previous works have mainly focused on quality improvement for I-ToF imaging especially curing the effect of Multi Path Interference (MPI). These investigations are typically done in specifically constrained scenarios at close distance, indoors and under little ambient light. Surprisingly little work has investigated I-ToF quality improvement in real-life scenarios where strong ambient light and far distances pose difficulties due to an extreme amount of induced shot noise and signal sparsity, caused by the attenuation with limited sensor power and light scattering. In this work, we propose a new learning based end-to-end depth prediction network which takes noisy raw I-ToF signals as well as an RGB image and fuses their latent representation based on a multi step approach involving both implicit and explicit alignment to predict a high quality long range depth map aligned to the RGB viewpoint. We test our approach on challenging real-world scenes and show more than 40% RMSE improvement on the final depth map compared to the baseline approach [33].
Nikolas Brasch, Ales Leonardis, Nassir Navab, Benjamin Busam
3DV3
2021 Depth-only Object Tracking
Ales Leonardis, Joni-Kristian Kämäräinen
BMVC3
2021 FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation Mechanism
abstract
In this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net) with efficient category-level feature extraction for 6D pose estimation. First, we design an orientation aware autoencoder with 3D graph convolution for latent feature extraction. Thanks to the shift and scale-invariance properties of 3D graph convolution, the learned latent feature is insensitive to point shift and object size. Then, to efficiently decode category-level rotation information from the latent feature, we propose a novel decoupled rotation mechanism that employs two decoders to complementarily access the rotation information. For translation and size, we estimate them by two residuals: the difference between the mean of object points and ground truth translation, and the difference between the mean size of the category and ground truth size, respectively. Finally, to increase the generalization ability of the FS-Net, we propose an on-line box-cage based 3D deformation mechanism to augment the training data. Extensive experiments on two benchmark datasets show that the proposed method achieves state-of-the-art performance in both category- and instance-level 6D object pose estimation. Especially in category-level pose estimation, without extra synthetic data, our method outperforms existing methods by 6.3% on the NOCS-REAL dataset1.
Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, LinLin Shen, Ales Leonardis
CVPR6
2021 DepthTrack: Unveiling the Power of RGBD Tracking
abstract
RGBD (RGB plus depth) object tracking is gaining momentum as RGBD sensors have become popular in many application fields such as robotics. However, the best RGBD trackers are extensions of the state-of-the-art deep RGB trackers. They are trained with RGB data and the depth channel is used as a sidekick for subtleties such as occlusion detection. This can be explained by the fact that there are no sufficiently large RGBD datasets to 1) train "deep depth trackers" and to 2) challenge RGB trackers with sequences for which the depth cue is essential. This work introduces a new RGBD tracking dataset - Depth-Track - that has twice as many sequences (200) and scene types (40) than in the largest existing dataset, and three times more objects (90). In addition, the average length of the sequences (1473), the number of deformable objects (16) and the number of annotated tracking attributes (15) have been increased. Furthermore, by running the SotA RGB and RGBD trackers on DepthTrack, we propose a new RGBD tracking baseline, namely DeT, which reveals that deep RGBD tracking indeed benefits from genuine training data. The code and dataset is available at https://github.com/xiaozai/DeT.
Jani Käpylä, Feng Zheng 0001, Ales Leonardis, Joni-Kristian Kämäräinen
ICCV5
2021 Motion-aware ensemble of three-mode trackers for unmanned aerial vehicles
Kyuewang Lee, Hyung Jin Chang, Jongwon Choi 0002, Byeongho Heo, Ales Leonardis, Jin Young Choi 0002
Mach. Vis. Appl.5
2020 G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector Features
abstract
In this paper, we propose a novel real-time 6D object pose estimation framework, named G2L-Net. Our network operates on point clouds from RGB-D detection in a divide-and-conquer fashion. Specifically, our network consists of three steps. First, we extract the coarse object point cloud from the RGB-D image by 2D detection. Second, we feed the coarse object point cloud to a translation localization network to perform 3D segmentation and object translation prediction. Third, via the predicted segmentation and translation, we transfer the fine object point cloud into a local canonical coordinate, in which we train a rotation localization network to estimate initial object rotation. In the third step, we define point-wise embedding vector features to capture viewpoint-aware information. To calculate more accurate rotation, we adopt a rotation residual estimator to estimate the residual between initial rotation and ground truth, which can boost initial pose estimation performance. Our proposed G2L-Net is real-time despite the fact multiple steps are stacked via the proposed coarse-to-fine framework. Extensive experiments on two benchmark datasets show that G2L-Net achieves state-of-the-art performance in terms of both accuracy and speed.
Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, Ales Leonardis
CVPR5
2020 A Multi-Hypothesis Approach to Color Constancy
abstract
Contemporary approaches frame the color constancy problem as learning camera specific illuminant mappings. While high accuracy can be achieved on camera specific data, these models depend on camera spectral sensitivity and typically exhibit poor generalisation to new devices. Additionally, regression methods produce point estimates that do not explicitly account for potential ambiguities among plausible illuminant solutions, due to the ill-posed nature of the problem. We propose a Bayesian framework that naturally handles color constancy ambiguity via a multi-hypothesis strategy. Firstly, we select a set of candidate scene illuminants in a data-driven fashion and apply them to a target image to generate a set of corrected images. Secondly, we estimate, for each corrected image, the likelihood of the light source being achromatic using a camera-agnostic CNN. Finally, our method explicitly learns a final illumination estimate from the generated posterior probability distribution. Our likelihood estimator learns to answer a camera-agnostic question and thus enables effective multi-camera training by disentangling illuminant estimation from the supervised learning task. We extensively evaluate our proposed approach and additionally set a benchmark for novel sensor generalisation without re-training. Our method provides state-of-the-art accuracy on multiple public datasets (up to 11% median angular error improvement) while maintaining real-time execution.
Daniel Hernández Juárez, Sarah Parisot, Benjamin Busam, Ales Leonardis, Gregory Slabaugh, Steven McDonagh 0001
CVPR4
2020 Unsupervised Model Personalization While Preserving Privacy and Scalability: An Open Problem
abstract
This work investigates the task of unsupervised model personalization, adapted to continually evolving, unlabeled local user images. We consider the practical scenario where a high capacity server interacts with a myriad of resource-limited edge devices, imposing strong requirements on scalability and local data privacy. We aim to address this challenge within the continual learning paradigm and provide a novel Dual User-Adaptation framework (DUA) to explore the problem. This framework flexibly disentangles user-adaptation into model personalization on the server and local data regularization on the user device, with desirable properties regarding scalability and privacy constraints. First, on the server, we introduce incremental learning of task-specific expert models, subsequently aggregated using a concealed unsupervised user prior. Aggregation avoids retraining, whereas the user prior conceals sensitive raw user data, and grants unsupervised adaptation. Second, local user-adaptation incorporates a domain adaptation point of view, adapting regularizing batch normalization parameters to the user data. We explore various empirical user configurations with different priors in categories and a tenfold of transforms for MIT Indoor Scene recognition, and classify numbers in a combined MNIST and SVHN setup. Extensive experiments yield promising results for data-driven local adaptation and elicit user priors for server adaptation to depend on the model rather than user data. Hence, although user-adaptation remains a challenging open problem, the DUA framework formalizes a principled foundation for personalizing both on server and user device, while maintaining privacy and scalability.
Matthias De Lange, Xu Jia 0012, Sarah Parisot, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars
CVPR4
2020 Image Demoireing with Learnable Bandpass Filters
abstract
Image demoireing is a multi-faceted image restoration task involving both texture and color restoration. In this paper, we propose a novel multiscale bandpass convolutional neural network (MBCNN) to address this problem. As an end-to-end solution, MBCNN respectively solves the two sub-problems. For texture restoration, we propose a learnable bandpass filter (LBF) to learn the frequency prior for moire texture removal. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, then performs local fine tuning of the color per pixel. Through an ablation study, we demonstrate the effectiveness of the different components of MBCNN. Experimental results on two public datasets show that our method outperforms state-of-the-art methods by a large margin (more than 2dB in terms of PSNR).
Bolun Zheng, Shanxin Yuan, Gregory Slabaugh, Ales Leonardis
CVPR4
2020 More Classifiers, Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning
Yu Liu 0012, Sarah Parisot, Gregory Slabaugh, Xu Jia 0012, Ales Leonardis, Tinne Tuytelaars
ECCV (26)5
2020 Many-Shot from Low-Shot: Learning to Annotate Using Mixed Supervision for Object Detection
Carlo Biffi, Steven McDonagh 0001, Philip Torr 0001, Ales Leonardis, Sarah Parisot
ECCV (8)4
2020 Wavelet-Based Dual-Branch Network for Image Demoiréing
Lin Liu 0016, Jianzhuang Liu, Shanxin Yuan, Gregory Slabaugh, Ales Leonardis, Wengang Zhou 0001, Qi Tian 0001
ECCV (13)5
2020 Quantifying the Use of Domain Randomization
abstract
Synthetic image generation provides the ability to efficiently produce large quantities of labeled data, which addresses both the data volume requirements of state-of-the-art vision systems and the expense of manually labeling data. However, systems trained on synthetic data typically under-perform systems trained on realistic data due to mismatch between the synthetic and realistic data distributions. Domain Randomization (DR) is a method of broadening a synthetic data distribution to encompass a realistic data distribution and provides better performance when the exact characteristics of the realistic data distribution are not known or cannot be simulated. However, there is no consensus in the literature on the best method of performing DR. We propose a novel method of ranking DR methods by directly measuring the difference between realistic and DR data distributions. This avoids the need to measure task-specific performance and the associated expense of training and evaluation. We compare different methods for measuring distribution differences, including the Wasserstein and Fréchet Inception distances. We also examine the effect of performing this evaluation directly on images and features generated by an image classification backbone. Finally, we show that the ranking generated by our method is reflected in actual task performance.
Mohammad Ani, Hector Basevi, Ales Leonardis
ICPR3
2020 PointPoseNet: Point Pose Network for Robust 6D Object Pose Estimation
abstract
In this paper, we propose a novel pipeline to estimate 6D object pose from RGB-D images of known objects present in complex scenes. The pipeline directly operates on raw point clouds extracted from RGB-D scans. Specifically, our method takes the point cloud as input and regresses the point-wise unit vectors pointing to the 3D keypoints. We then use these vectors to generate keypoint hypotheses from which the 6D object pose hypotheses are computed. Finally, we select the best 6D object pose from the hypotheses based on a proposed scoring mechanism with geometry constraints. Extensive experiments show that the proposed method is robust against the variety in object shape and appearance as well as occlusions between objects, and that our method outperforms the state-of-the-art methods on the LINEMOD and Occlusion LINEMOD datasets.
Wei Chen 0092, Jinming Duan 0001, Hector Basevi, Hyung Jin Chang, Ales Leonardis
WACV5
2020 Spatially-Adaptive Filter Units for Compact and Efficient Deep Neural Networks
Domen Tabernik, Matej Kristan, Ales Leonardis
Int. J. Comput. Vis.3
2018 Spatially-Adaptive Filter Units for Deep Neural Networks
abstract
Classical deep convolutional networks increase receptive field size by either gradual resolution reduction or application of hand-crafted dilated convolutions to prevent increase in the number of parameters. In this paper we propose a novel displaced aggregation unit (DAU) that does not require hand-crafting. In contrast to classical filters with units (pixels) placed on a fixed regular grid, the displacement of the DAUs are learned, which enables filters to spatially-adapt their receptive field to a given problem. We extensively demonstrate the strength of DAUs on a classification and semantic segmentation tasks. Compared to ConvNets with regular filter, ConvNets with DAUs achieve comparable performance at faster convergence and up to 3-times reduction in parameters. Furthermore, DAUs allow us to study deep networks from novel perspectives. We study spatial distributions of DAU filters and analyze the number of parameters allocated for spatial coverage in a filter.
Domen Tabernik, Matej Kristan, Ales Leonardis
CVPR3
2018 Learning to Exploit Stability for 3D Scene Parsing
abstract
Human scene understanding uses a variety of visual and non-visual cues to perform inference on object types, poses, and relations. Physics is a rich and universal cue which we exploit to enhance scene understanding. We integrate the physical cue of stability into the learning process using a REINFORCE approach coupled to a physics engine, and apply this to the problem of producing the 3D bounding boxes and poses of objects in a scene. We first show that applying physics supervision to an existing scene understanding model increases performance, produces more stable predictions, and allows training to an equivalent performance level with fewer annotated training examples. We then present a novel architecture for 3D scene parsing named Prim R-CNN, learning to predict bounding boxes as well as their 3D size, translation, and rotation. With physics supervision, Prim R-CNN outperforms existing scene understanding approaches on this problem. Finally, we show that applying physics supervision on unlabeled real images improves real domain transfer of models training on synthetic data.
Yilun Du, Hector Basevi, Ales Leonardis, William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001
NeurIPS4
2018 Region-sequence based six-stream CNN features for general and fine-grained human action recognition in videos
Miao Ma, Naresh Marturi, Yibin Li 0001, Ales Leonardis, Rustam Stolkin
Pattern Recognit.4
2018 Robust Fusion of Color and Depth Data for RGB-D Target Tracking Using Adaptive Range-Invariant Depth Models and Spatio-Temporal Consistency Constraints
abstract
This paper presents a novel robust method for single target tracking in RGB-D images, and also contributes a substantial new benchmark dataset for evaluating RGB-D trackers. While a target object's color distribution is reasonably motion-invariant, this is not true for the target's depth distribution, which continually varies as the target moves relative to the camera. It is therefore nontrivial to design target models which can fully exploit (potentially very rich) depth information for target tracking. For this reason, much of the previous RGB-D literature relies on color information for tracking, while exploiting depth information only for occlusion reasoning. In contrast, we propose an adaptive range-invariant target depth model, and show how both depth and color information can be fully and adaptively fused during the search for the target in each new RGB-D image. We introduce a new, hierarchical, two-layered target model (comprising local and global models) which uses spatio-temporal consistency constraints to achieve stable and robust on-the-fly target relearning. In the global layer, multiple features, derived from both color and depth data, are adaptively fused to find a candidate target region. In ambiguous frames, where one or more features disagree, this global candidate region is further decomposed into smaller local candidate regions for matching to local-layer models of small target parts. We also note that conventional use of depth data, for occlusion reasoning, can easily trigger false occlusion detections when the target moves rapidly toward the camera. To overcome this problem, we show how combining target information with contextual information enables the target's depth constraint to be relaxed. Our adaptively relaxed depth constraints can robustly accommodate large and rapid target motion in the depth direction, while still enabling the use of depth data for highly accurate reasoning about occlusions. For evaluation, we introduce a new RGB-D benchmark dataset with per-frame annotated attributes and extensive bias analysis. Our tracker is evaluated using two different state-of-the-art methodologies, VOT and object tracking benchmark, and in both cases it significantly outperforms four other state-of-the-art RGB-D trackers from the literature.
Rustam Stolkin, Ales Leonardis
IEEE Trans. Cybern.4
2017 Rolling Shutter Correction in Manhattan World
abstract
A vast majority of consumer cameras operate the rolling shutter mechanism, which often produces distorted images due to inter-row delay while capturing an image. Recent methods for monocular rolling shutter compensation utilize blur kernel, straightness of line segments, as well as angle and length preservation. However, they do not incorporate scene geometry explicitly for rolling shutter correction, therefore, information about the 3D scene geometry is often distorted by the correction process. In this paper we propose a novel method which leverages geometric properties of the scene-in particular vanishing directions-to estimate the camera motion during rolling shutter exposure from a single distorted image. The proposed method jointly estimates the orthogonal vanishing directions and the rolling shutter camera motion. We performed extensive experiments on synthetic and real datasets which demonstrate the benefits of our approach both in terms of qualitative and quantitative results (in terms of a geometric structure fitting) as well as with respect to computation time.
Pulak Purkait, Christopher Zach, Ales Leonardis
ICCV3
2017 Beyond Standard Benchmarks: Parameterizing Performance Evaluation in Visual Object Tracking
abstract
Object-to-camera motion produces a variety of apparent motion patterns that significantly affect performance of short-term visual trackers. Despite being crucial for designing robust trackers, their influence is poorly explored in standard benchmarks due to weakly defined, biased and overlapping attribute annotations. In this paper we propose to go beyond pre-recorded benchmarks with post-hoc annotations by presenting an approach that utilizes omnidirectional videos to generate realistic, consistently annotated, short-term tracking scenarios with exactly parameterized motion patterns. We have created an evaluation system, constructed a fully annotated dataset of omnidirectional videos and generators for typical motion patterns. We provide an in-depth analysis of major tracking paradigms which is complementary to the standard benchmarks and confirms the expressiveness of our evaluation approach.
Luka Cehovin, Alan Lukezic, Ales Leonardis, Matej Kristan
ICCV3
2017 Visual stability prediction for robotic manipulation
abstract
Understanding physical phenomena is a key competence that enables humans and animals to act and interact under uncertain perception in previously unseen environments containing novel objects and their configurations. Developmental psychology has shown that such skills are acquired by infants from observations at a very early stage. In this paper, we contrast a more traditional approach of taking a model-based route with explicit 3D representations and physical simulation by an end-to-end approach that directly predicts stability from appearance. We ask the question if and to what extent and quality such a skill can directly be acquired in a data-driven way — bypassing the need for an explicit simulation at run-time. We present a learning-based approach based on simulated data that predicts stability of towers comprised of wooden blocks under different conditions and quantities related to the potential fall of the towers. We first evaluate the approach on synthetic data and compared the results to human judgments on the same stimuli. Further, we extend this approach to reason about future states of such towers that in return enables successful stacking.
Wenbin Li 0003, Ales Leonardis, Mario Fritz
ICRA2
2017 Dynamic multi-level appearance models and adaptive clustered decision trees for single target tracking
Rustam Stolkin, Ales Leonardis
Pattern Recognit.3
2016 Distractor-Supported Single Target Tracking in Extremely Cluttered Scenes
Linbo Qiao, Rustam Stolkin, Ales Leonardis
ECCV (4)4
2016 Towards deep compositional networks
abstract
Hierarchical feature learning based on convolutional neural networks (CNN) has recently shown significant potential in various computer vision tasks. While allowing high-quality discriminative feature learning, the downside of CNNs is the lack of explicit structure in features, which often leads to overfitting, absence of reconstruction from partial observations and limited generative abilities. Explicit structure is inherent in hierarchical compositional models, however, these lack the ability to optimize a well-defined cost function. We propose a novel analytic model of a basic unit in a layered hierarchical model with both explicit compositional structure and a well-defined discriminative cost function. Our experiments on two datasets show that the proposed compositional model performs on a par with standard CNNs on discriminative tasks, while, due to explicit modeling of the structure in the feature units, affording a straight-forward visualization of parts and faster inference due to separability of the units.
Domen Tabernik, Matej Kristan, Jeremy L. Wyatt, Ales Leonardis
ICPR4
2016 Hierarchical spatial model for 2D range data based room categorization
abstract
The next generation service robots are expected to co-exist with humans in their homes. Such a mobile robot requires an efficient representation of space, which should be compact and expressive, for effective operation in real-world environments. In this paper we present a novel approach for 2D ground-plan-like laser-range-data-based room categorization that builds on a compositional hierarchical representation of space, and show how an additional abstraction layer, whose parts are formed by merging partial views of the environment followed by graph extraction, can achieve improved categorization performance. A new algorithm is presented that finds a dictionary of exemplar elements from a multi-category set, based on the affinity measure defined among pairs of elements. This algorithm is used for part selection in new layer construction. Room categorization experiments have been performed on a challenging publicly available dataset, which has been extended in this work. State-of-the-art results were obtained by achieving the most balanced performance over all categories.
Peter Ursic, Ales Leonardis, Danijel Skocaj, Matej Kristan
ICRA2
2016 Part-based room categorization for household service robots
abstract
A service robot that operates in a previously-unseen home environment should be able to recognize the functionality of the rooms it visits, such as a living room, a bathroom, etc. We present a novel part-based model and an approach for room categorization using data obtained from a visual sensor. Images are represented with sets of unordered parts that are obtained by object-agnostic region proposals, and encoded using state-of-the-art image descriptor extractor - a convolutional neural network (CNN). An approach is proposed that learns category-specific discriminative parts for the part-based model. The proposed approach was compared to the state-of-the-art CNN trained specifically for place recognition. Experimental results show that the proposed approach outperforms the holistic CNN by being robust to image degradation, such as occlusions, modifications of image scaling, and aspect changes. In addition, we report non-negligible annotation errors and image duplicates in a popular dataset for place categorization and discuss annotation ambiguities.
Peter Ursic, Rok Mandeljc, Ales Leonardis, Matej Kristan
ICRA3
2016 Task-relevant grasp selection: A joint solution to planning grasps and manipulative motion trajectories
abstract
This paper addresses the problem of jointly planning both grasps and subsequent manipulative actions. Previously, these two problems have typically been studied in isolation, however joint reasoning is essential to enable robots to complete real manipulative tasks. In this paper, the two problems are addressed jointly and a solution that takes both into consideration is proposed. To do so, a manipulation capability index is defined, which is a function of both the task execution waypoints and the object grasping contact points. We build on recent state-of-the-art grasp-learning methods, to show how this index can be combined with a likelihood function computed by a probabilistic model of grasp selection, enabling the planning of grasps which have a high likelihood of being stable, but which also maximise the robot's capability to deliver a desired post-grasp task trajectory. We also show how this paradigm can be extended, from a single arm and hand, to enable efficient grasping and manipulation with a bi-manual robot. We demonstrate the effectiveness of the approach using experiments on a simulated as well as a real robot.
Amir M. Ghalamzan E., Nikos Mavrakis, Marek Sewer Kopicki, Rustam Stolkin, Ales Leonardis
IROS5
2016 Robust visual tracking using template anchors
abstract
Deformable part models exhibit excellent performance in tracking non-rigidly deforming targets, but are usually outperformed by holistic models when the target does not deform or in the presence of uncertain visual data. The reason is that part-based models require estimation of a larger number of parameters compared to holistic models and since the updating process is self-supervised, the errors in parameter estimation are amplified with time, leading to a faster accuracy reduction than in holistic models. On the other hand, the robustness of part-based trackers is generally greater than in holistic trackers. We address the problem of self-supervised estimation of a large number of parameters by introducing controlled graduation in estimation of the free parameters. We propose decomposing the visual model into several sub-models, each describing the target at a different level of detail. The sub-models interact during target localization and, depending on the visual uncertainty, serve for cross-sub-model supervised updating. A new tracker is proposed based on this model which exhibits the qualities of part-based as well as holistic models. The tracker is tested on the highly-challenging VOT2013 and VOT2014 benchmarks, outperforming the state-of-the-art.
Luka Cehovin, Ales Leonardis, Matej Kristan
WACV2
2016 A local-global coupled-layer puppet model for robust online human pose tracking
Miao Ma, Naresh Marturi, Yibin Li 0001, Rustam Stolkin, Ales Leonardis
Comput. Vis. Image Underst.5
2016 A Novel Performance Evaluation Methodology for Single-Target Trackers
abstract
This paper addresses the problem of single-target tracker performance evaluation. We consider the performance measures, the dataset and the evaluation system to be the most important components of tracker evaluation and propose requirements for each of them. The requirements are the basis of a new evaluation methodology that aims at a simple and easily interpretable tracker comparison. The ranking-based methodology addresses tracker equivalence in terms of statistical significance and practical differences. A fully-annotated dataset with per-frame annotations with several visual attributes is introduced. The diversity of its visual properties is maximized in a novel way by clustering a large number of videos according to their visual attributes. This makes it the most sophistically constructed and annotated dataset to date. A multi-platform evaluation system allowing easy integration of third-party trackers is presented as well. The proposed evaluation methodology was tested on the VOT2014 challenge on the new dataset and 38 trackers, making it the largest benchmark to date. Most of the tested trackers are indeed state-of-the-art since they outperform the standard baselines, resulting in a highly-challenging benchmark. An exhaustive analysis of the dataset from the perspective of tracking difficulty is carried out. To facilitate tracker comparison a new performance visualization technique is proposed.
Matej Kristan, Jiri Matas, Ales Leonardis, Tomás Vojír, Roman P. Pflugfelder, Gustavo Fernández, Georg Nebehay, Fatih Porikli, Luka Cehovin
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Visual Object Tracking Performance Measures Revisited
abstract
The problem of visual tracking evaluation is sporting a large variety of performance measures, and largely suffers from lack of consensus about which measures should be used in experiments. This makes the cross-paper tracker comparison difficult. Furthermore, as some measures may be less effective than others, the tracking results may be skewed or biased toward particular tracking aspects. In this paper, we revisit the popular performance measures and tracker performance visualizations and analyze them theoretically and experimentally. We show that several measures are equivalent from the point of information they provide for tracker comparison and, crucially, that some are more brittle than the others. Based on our analysis, we narrow down the set of potential measures to only two complementary ones, describing accuracy and robustness, thus pushing toward homogenization of the tracker evaluation methodology. These two measures can be intuitively interpreted and visualized and have been employed by the recent visual object tracking challenges as the foundation for the evaluation methodology.
Luka Cehovin, Ales Leonardis, Matej Kristan
IEEE Trans. Image Process.2
2015 Single target tracking using adaptive clustered decision trees and dynamic multi-level appearance models
abstract
This paper presents a method for single target tracking of arbitrary objects in challenging video sequences. Targets are modeled at three different levels of granularity (pixel level, parts-based level and bounding box level), which are cross-constrained to enable robust model relearning. The main contribution is an adaptive clustered decision tree method which dynamically selects the minimum combination of features necessary to sufficiently represent each target part at each frame, thereby providing robustness with computational efficiency. The adaptive clustered decision tree is implemented in two separate parts of the tracking algorithm: firstly to enable robust matching at the parts-based level between successive frames; and secondly to select the best superpixels for learning new parts of the target. We have tested the tracker using two different tracking benchmarks (VOT2013-2014 and CVPR2013 tracking challenges), based on two different test methodologies, and show it to be significantly more robust than the best state-of-the-art methods from both of those tracking challenges, while also offering competitive tracking precision.
Rustam Stolkin, Ales Leonardis
CVPR3
2015 Compositional Hierarchical Representation of Shape Manifolds for Classification of Non-manifold Shapes
abstract
We address the problem of statistical learning of shape models which are invariant to translation, rotation and scale in compositional hierarchies when data spaces of measurements and shape spaces are not topological manifolds. In practice, this problem is observed while modeling shapes having multiple disconnected components, e.g. partially occluded shapes in cluttered scenes. We resolve the aforementioned problem by first reformulating the relationship between data and shape spaces considering the interaction between Receptive Fields (RFs) and Shape Manifolds (SMs) in a compositional hierarchical shape vocabulary. Then, we suggest a method to model the topological structure of the SMs for statistical learning of the geometric transformations of the shapes that are defined by group actions on the SMs. For this purpose, we design a disjoint union topology using an indexing mechanism for the formation of shape models on SMs in the vocabulary, recursively. We represent the topological relationship between shape components using graphs, which are aggregated to construct a hierarchical graph structure for the shape vocabulary. To this end, we introduce a framework to implement the indexing mechanisms for the employment of the vocabulary for structural shape classification. The proposed approach is used to construct invariant shape representations. Results on benchmark shape classification outperform state-of-the-art methods.
Mete Ozay, Ümit Rusen Aktas, Jeremy L. Wyatt, Ales Leonardis
ICCV4
2015 Adding discriminative power to a generative hierarchical compositional model using histograms of compositions
Domen Tabernik, Ales Leonardis, Marko Boben, Danijel Skocaj, Matej Kristan
Comput. Vis. Image Underst.2
2014 Multi-target tracking in team-sports videos via multi-level context-conditioned latent behaviour models
Rustam Stolkin, Ales Leonardis
BMVC3
2014 A Graph Theoretic Approach for Object Shape Representation in Compositional Hierarchies Using a Hybrid Generative-Descriptive Model
Ümit Rusen Aktas, Mete Ozay, Ales Leonardis, Jeremy L. Wyatt
ECCV (3)3
2014 Object Categorization from Range Images Using a Hierarchical Compositional Representation
abstract
This paper proposes a novel hierarchical compositional representation of 3D shape that can accommodate a large number of object categories and enables efficient learning and inference. The hierarchy starts with simple pre-defined parts on the first layer, after which subsequent layers are learned recursively by taking the most statistically significant compositions of parts from the previous layer. Our representation is able to scale because of its very economical use of memory and because subparts of the representation are shared. We apply our representation to 3D multi-class object categorization. Object categories are represented by histograms of compositional parts, which are then used as inputs to an SVM classifier. We present results for two datasets, Aim Shape [1] and the Washington RGB-D Object Dataset [2], and demonstrate the competitive performance of our method.
Vladislav Kramarev, Sebastian Zurek, Jeremy L. Wyatt, Ales Leonardis
ICPR4
2014 A hierarchical approach for joint multi-view object pose estimation and categorization
abstract
We propose a joint object pose estimation and categorization approach which extracts information about object poses and categories from the object parts and compositions constructed at different layers of a hierarchical object representation algorithm, namely Learned Hierarchy of Parts (LHOP) [7]. In the proposed approach, we first employ the LHOP to learn hierarchical part libraries which represent entity parts and compositions across different object categories and views. Then, we extract statistical and geometric features from the part realizations of the objects in the images in order to represent the information about object pose and category at each different layer of the hierarchy. Unlike the traditional approaches which consider specific layers of the hierarchies in order to extract information to perform specific tasks, we combine the information extracted at different layers to solve a joint object pose estimation and categorization problem using distributed optimization algorithms. We examine the proposed generative-discriminative learning approach and the algorithms on two benchmark 2-D multi-view image datasets. The proposed approach and the algorithms outperform state-of-the-art classification, regression and feature extraction algorithms. In addition, the experimental results shed light on the relationship between object categorization, pose estimation and the part realizations observed at different layers of the hierarchy.
Mete Ozay, Krzysztof Walas, Ales Leonardis
ICRA3
2014 Is my new tracker really better than yours?
abstract
The problem of visual tracking evaluation is sporting an abundance of performance measures, which are used by various authors, and largely suffers from lack of consensus about which measures should be preferred. This is hampering the cross-paper tracker comparison and faster advancement of the field. In this paper we provide an overview of the popular measures and performance visualizations and their critical theoretical and experimental analysis. We show that several measures are equivalent from the point of information they provide for tracker comparison and, crucially, that some are more brittle than the others. Based on our analysis we narrow down the set of potential measures to only two complementary ones that can be intuitively interpreted and visualized, thus pushing towards homogenization of the tracker evaluation methodology.
Luka Cehovin, Matej Kristan, Ales Leonardis
WACV3
2014 Online Discriminative Kernel Density Estimator With Gaussian Kernels
abstract
We propose a new method for a supervised online estimation of probabilistic discriminative models for classification tasks. The method estimates the class distributions from a stream of data in the form of Gaussian mixture models (GMMs). The reconstructive updates of the distributions are based on the recently proposed online kernel density estimator (oKDE). We maintain the number of components in the model low by compressing the GMMs from time to time. We propose a new cost function that measures loss of interclass discrimination during compression, thus guiding the compression toward simpler models that still retain discriminative properties. The resulting classifier thus independently updates the GMM of each class, but these GMMs interact during their compression through the proposed cost function. We call the proposed method the online discriminative kernel density estimator (odKDE). We compare the odKDE to oKDE, batch state-of-the-art kernel density estimators (KDEs), and batch/incremental support vector machines (SVM) on the publicly available datasets. The odKDE achieves comparable classification performance to that of best batch KDEs and SVM, while allowing online adaptation from large datasets, and produces models of lower complexity than the oKDE.
Matej Kristan, Ales Leonardis
IEEE Trans. Cybern.2
2013 A Web-Service for Object Detection Using Hierarchical Models
Domen Tabernik, Luka Cehovin, Matej Kristan, Marko Boben, Ales Leonardis
ICVS5
2013 Learning Compositional Hierarchies of a Sensorimotor System
Jure Zabkar, Ales Leonardis
IDA2
2013 Robust Visual Tracking Using an Adaptive Coupled-Layer Visual Model
abstract
This paper addresses the problem of tracking objects which undergo rapid and significant appearance changes. We propose a novel coupled-layer visual model that combines the target's global and local appearance by interlacing two layers. The local layer in this model is a set of local patches that geometrically constrain the changes in the target's appearance. This layer probabilistically adapts to the target's geometric deformation, while its structure is updated by removing and adding the local patches. The addition of these patches is constrained by the global layer that probabilistically models the target's global visual properties, such as color, shape, and apparent local motion. The global visual properties are updated during tracking using the stable patches from the local layer. By this coupled constraint paradigm between the adaptation of the global and the local layer, we achieve a more robust tracking through significant appearance changes. We experimentally compare our tracker to 11 state-of-the-art trackers. The experimental results on challenging sequences confirm that our tracker outperforms the related trackers in many cases by having a smaller failure rate as well as better accuracy. Furthermore, the parameter analysis shows that our tracker is stable over a range of parameter values.
Luka Cehovin, Matej Kristan, Ales Leonardis
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Deep Hierarchies in the Primate Visual Cortex: What Can We Learn for Computer Vision?
abstract
Computational modeling of the primate visual system yields insights of potential relevance to some of the challenges that computer vision is facing, such as object recognition and categorization, motion detection and activity recognition, or vision-based navigation and manipulation. This paper reviews some functional principles and structures that are generally thought to underlie the primate visual cortex, and attempts to extract biological principles that could further advance computer vision research. Organized for a computer vision audience, we present functional principles of the processing hierarchies present in the primate visual system considering recent discoveries in neurophysiology. The hierarchical processing in the primate visual system is characterized by a sequence of different levels of processing (on the order of 10) that constitute a deep hierarchy in contrast to the flat vision architectures predominantly used in today's mainstream computer vision. We hope that the functional description of the deep hierarchies realized in the primate visual system provides valuable insights for the design of computer vision algorithms, fostering increasingly productive interaction between biological and computer vision research.
Norbert Krüger, Peter Janssen, Sinan Kalkan, Markus Lappe, Ales Leonardis, Justus H. Piater, Antonio Jose Rodríguez-Sánchez, Laurenz Wiskott
IEEE Trans. Pattern Anal. Mach. Intell.5
2012 Learning statistically relevant edge structure improves low-level visual descriptors
Domen Tabernik, Matej Kristan, Marko Boben, Ales Leonardis
ICPR4
2012 Room classification using a hierarchical representation of space
abstract
Mobile robots need an effective spatial model for the successful operation in real-world environment. The model should be compact and simultaneously possess large expressive power. Moreover, it should scale well. In this paper we propose a new hierarchical representation of space, whose compositional structure is learned based on statistically significant observations. We have focused on a two dimensional space, since many robots perceive their surroundings in two dimensions with the use of a laser range finder or a sonar. We also propose the use of a low-level image descriptor for addressing the room classification problem, by which we demonstrate the performance of our representation. Using only the lower layers of the hierarchy, we obtain state-of-the-art classification results on demanding datasets.
Peter Ursic, Matej Kristan, Danijel Skocaj, Ales Leonardis
IROS4
2012 Modeling binding and cross-modal learning in Markov logic networks
Alen Vrecko, Ales Leonardis, Danijel Skocaj
Neurocomputing2
2011 An adaptive coupled-layer visual model for robust visual tracking
abstract
This paper addresses the problem of tracking objects which undergo rapid and significant appearance changes. We propose a novel coupled-layer visual model that combines the target's global and local appearance. The local layer in this model is a set of local patches that geometrically constrain the changes in the target's appearance. This layer probabilistically adapts to the target's geometric deformation, while its structure is updated by removing and adding the local patches. The addition of the patches is constrained by the global layer that probabilistically models target's global visual properties such as color, shape and apparent local motion. The global visual properties are updated during tracking using the stable patches from the local layer. By this coupled constraint paradigm between the adaptation of the global and the local layer, we achieve a more robust tracking through significant appearance changes. Indeed, the experimental results on challenging sequences confirm that our tracker outperforms the related state-of-the-art trackers by having smaller failure rate as well as better accuracy.
Luka Cehovin, Matej Kristan, Ales Leonardis
ICCV3
2011 Hyperlinking reality via camera phones
Dusan Omercevic, Ales Leonardis
Mach. Vis. Appl.2
2011 Multivariate online kernel density estimation with Gaussian kernels
Matej Kristan, Ales Leonardis, Danijel Skocaj
Pattern Recognit.2
2010 A Coarse-to-Fine Taxonomy of Constellations for Fast Multi-class Object Detection
Sanja Fidler, Marko Boben, Ales Leonardis
ECCV (5)3
2010 Online Discriminative Kernel Density Estimation
abstract
We propose a new method for online estimation of probabilistic discriminative models. The method is based on the recently proposed online Kernel Density Estimation (oKDE) framework which produces Gaussian mixture models and allows adaptation using only a single data point at a time. The oKDE builds reconstructive models from the data, and we extend it to take into account the interclass discrimination through a new distance function between the classifiers. We arrive at an online discriminative Kernel Density Estimator (odKDE). We compare the odKDE to oKDE, batch state-of-the-art KDEs and support vector machine (SVM) on a standard database. The odKDE achieves comparable classification performance to that of best batch KDEs and SVM, while allowing online adaptation, and produces models of lower complexity than the oKDE.
Matej Kristan, Ales Leonardis
ICPR2
2010 Self-supervised cross-modal online learning of basic object affordances for developmental robotic systems
abstract
For a developmental robotic system to function successfully in the real world, it is important that it be able to form its own internal representations of affordance classes based on observable regularities in sensory data. Usually successful classifiers are built using labeled training data, but it is not always realistic to assume that labels are available in a developmental robotics setting. There does, however, exist an advantage in this setting that can help circumvent the absence of labels: co-occurrence of correlated data across separate sensory modalities over time. The main contribution of this paper is an online classifier training algorithm based on Kohonen's learning vector quantization (LVQ) that, by taking advantage of this co-occurrence information, does not require labels during training, either dynamically generated or otherwise. We evaluate the algorithm in experiments involving a robotic arm that interacts with various household objects on a table surface where camera systems extract features for two separate visual modalities. It is shown to improve its ability to classify the affordances of novel objects over time, coming close to the performance of equivalent fully-supervised algorithms.
Barry Ridge, Danijel Skocaj, Ales Leonardis
ICRA3
2010 Towards correct and informative evaluation methodology for texture classification under varying viewpoint and illumination
Ondrej Drbohlav, Ales Leonardis
Comput. Vis. Image Underst.2
2010 A framework for visual-context-aware object detection in still images
Roland Perko, Ales Leonardis
Comput. Vis. Image Underst.2
2010 Online kernel density estimation for interactive learning
Matej Kristan, Danijel Skocaj, Ales Leonardis
Image Vis. Comput.3
2010 Online pattern recognition and machine learning techniques for computer-vision: Theory and applications
Bogdan Raducanu, Jordi Vitrià, Ales Leonardis
Image Vis. Comput.3
2010 A Two-Stage Dynamic Model for Visual Tracking
abstract
We propose a new dynamic model which can be used within blob trackers to track the target's center of gravity. A strong point of the model is that it is designed to track a variety of motions which are usually encountered in applications such as pedestrian tracking, hand tracking, and sports. We call the dynamic model a two-stage dynamic model due to its particular structure, which is a composition of two models: a liberal model and a conservative model. The liberal model allows larger perturbations in the target's dynamics and is able to account for motions in between the random-walk dynamics and the nearly constant-velocity dynamics. On the other hand, the conservative model assumes smaller perturbations and is used to further constrain the liberal model to the target's current dynamics. We implement the two-stage dynamic model in a two-stage probabilistic tracker based on the particle filter and apply it to two separate examples of blob tracking: 1) tracking entire persons and 2) tracking of a person's hands. Experiments show that, in comparison to the widely used models, the proposed two-stage dynamic model allows tracking with smaller number of particles in the particle filter (e.g., 25 particles), while achieving smaller errors in the state estimation and a smaller failure rate. The results suggest that the improved performance comes from the model's ability to actively adapt to the target's motion during tracking.
Matej Kristan, Stanislav Kovacic, Ales Leonardis, Janez Pers
IEEE Trans. Syst. Man Cybern. Part B3
2009 Optimization Framework for Learning a Hierarchical Shape Vocabulary for Object Class Detection
abstract
This paper proposes a stochastic optimization framework for unsupervised learning of a hierarchical vocabulary of object shape intended for object class detection. We build on the approach by [6], which has two drawbacks: 1.) learning is performed strictly bottom-up; and 2.) the selection of vocabulary shapes is done solely on their frequency of appearance. This makes the method prone to overfitting of certain parts of object shape while losing the more discriminative shape information. The idea of this paper is to cast the vocabulary learning into an optimization framework that iteratively improves the hierarchy as a whole. Optimization is two-fold: one that learns and selects the vocabulary of shapes at each layer in a bottom-up phase and the other that extends/improves it by top-down feedback from the higher layers. The algorithm then loops between the two learning stages several times. We have evaluated the proposed learning approach for object class detection on 11 diverse object classes taken from the standard recognition data sets. Compared to the original approach [6], we obtain a 3 times more compact vocabulary, a 2:5 times faster inference, and a 10% higher detection performance at the expense of 5 times longer training time (25min vs 5min). The approach attains a competitive detection performance with respect to the current state-of-the-art at both, faster inference as well as shorter training times.
Sanja Fidler, Marko Boben, Ales Leonardis
BMVC3
2009 Learning contextual rules for priming object categories in images
abstract
In this paper we introduce and exploit the concept of contextual rules in the field of object detection. These rules are defined as associations between different object likelihood maps and are learned from given examples. The contextual rules can be used to prime regions where a target object category occurs in an image given areas of other object categories. The principal idea is to locate several basic object categories in an image and then use this information to infer object likelihood maps for other object categories. The proposed framework itself is general and not limited to specific object categories. For demonstrating our approach, we use likely occurrences of pedestrians and windows in urban scenes, extracted by a technique employing visual context, and use them to prime for shop logos.
Roland Perko, Lucas Paletta, Ales Leonardis
ICIP3
2009 A computer vision integration model for a multi-modal cognitive system
abstract
We present a general method for integrating visual components into a multi-modal cognitive system. The integration is very generic and can work with an arbitrary set of modalities. We illustrate our integration approach with a specific instantiation of the architecture schema that focuses on integration of vision and language: a cognitive system able to collaborate with a human, learn and display some understanding of its surroundings. As examples of cross-modal interaction we describe mechanisms for clarification and visual learning.
Alen Vrecko, Danijel Skocaj, Nick Hawes, Ales Leonardis
IROS4
2009 Evaluating multi-class learning strategies in a generative hierarchical framework for object detection
abstract
Multiple object class learning and detection is a challenging problem due to the large number of object classes and their high visual variability. Specialized detectors usually excel in performance, while joint representations optimize sharing and reduce inference time --- but are complex to train. Conveniently, sequential learning of categories cuts down training time by transferring existing knowledge to novel classes, but cannot fully exploit the richness of shareability and might depend on ordering in learning. In hierarchical frameworks these issues have been little explored. In this paper, we show how different types of multi-class learning can be done within one generative hierarchical framework and provide a rigorous experimental analysis of various object class learning strategies as the number of classes grows. Specifically, we propose, evaluate and compare three important types of multi-class learning: 1.) independent training of individual categories, 2.) joint training of classes, 3.) sequential learning of classes. We explore and compare their computational behavior (space and time) and detection performance as a function of the number of learned classes on several recognition data sets.
Sanja Fidler, Marko Boben, Ales Leonardis
NIPS3
2009 Editorial Special Issue ECCV 2006
Horst Bischof, Ales Leonardis
Int. J. Comput. Vis.2
2009 A local-motion-based probabilistic model for visual tracking
Matej Kristan, Janez Pers, Stanislav Kovacic, Ales Leonardis
Pattern Recognit.4
2008 Similarity-based cross-layered hierarchical representation for object categorization
abstract
This paper proposes a new concept in hierarchical representations that exploits features of different granularity and specificity coming from all layers of the hierarchy. The concept is realized within a cross-layered compositional representation learned from the visual data. We show how similarity connections among discrete labels within and across hierarchical layers can be established in order to produce a set of layer-independent shape-terminals, i.e. shapinals. We thus break the traditional notion of hierarchies and show how the category-specific layers can make use of all the necessary features stemming from all hierarchical layers. This, on the one hand, brings higher generalization into the representation, yet on the other hand, it also encodes the notion of scales directly into the hierarchy, thus enabling a multi-scale representation of object categories. By focusing on shape information only, the approach is tested on the Caltech 101 dataset demonstrating good performance in comparison with other state-of-the-art methods.
Sanja Fidler, Marko Boben, Ales Leonardis
CVPR3
2008 Robust Object Detection with Interleaved Categorization and Segmentation
Bastian Leibe, Ales Leonardis, Bernt Schiele
Int. J. Comput. Vis.2
2008 Unsupervised Learning of a Hierarchy of Topological Maps Using Omnidirectional Images
abstract
This paper presents a novel appearance-based method for path-based map learning by a mobile robot equipped with an omnidirectional camera. In particular, we focus on an unsupervised construction of topological maps, which provide an abstraction of the environment in terms of visual aspects. An unsupervised clustering algorithm is used to represent the images in multiple subspaces, forming thus a sensory grounded representation of the environment's appearance. By introducing transitional fields between clusters we are able to obtain a partitioning of the image set into distinctive visual aspects. By abstracting the low-level sensory data we are able to efficiently reconstruct the overall topological layout of the covered path. After the high level topology is estimated, we repeat the procedure on the level of visual aspects to obtain local topological maps. We demonstrate how the resulting representation can be used for modeling indoor and outdoor environments, how it successfully detects previously visited locations and how it can be used for the estimation of the current visual aspect and the retrieval of the relative position within the current visual aspect.
Ales Stimec, Matjaz Jogan, Ales Leonardis
Int. J. Pattern Recognit. Artif. Intell.3
2008 Incremental and robust learning of subspace representations
Danijel Skocaj, Ales Leonardis
Image Vis. Comput.2
2008 Selecting features for object detection using an AdaBoost-compatible evaluation function
Luka Fürst, Sanja Fidler, Ales Leonardis
Pattern Recognit. Lett.3
2007 Incremental LDA Learning by Combining Reconstructive and Discriminative Approaches
abstract
Incremental subspace methods have proven to enable efficient training if large amounts of training data have to be processed or if not all data is available in advance. In this paper we focus on incremental LDA learning which provides good classification results while it assures a compact data representation. In contrast to existing incremental LDA methods we additionally consider reconstructive information when incrementally building the LDA subspace. Hence, we get a more flexible representation that is capable to adapt to new data. Moreover, this allows to add new instances to existing classes as well as to add new classes. The experimental results show that the proposed approach outperforms other incremental LDA methods even approaching classification results obtained by batch learning. 1
Martina Uray, Danijel Skocaj, Peter M. Roth, Horst Bischof, Ales Leonardis
BMVC5
2007 Towards Scalable Representations of Object Categories: Learning a Hierarchy of Parts
abstract
This paper proposes a novel approach to constructing a hierarchical representation of visual input that aims to enable recognition and detection of a large number of object categories. Inspired by the principles of efficient indexing (bottom-up), robust matching (top-down), and ideas of compositionality, our approach learns a hierarchy of spatially flexible compositions, i.e. parts, in an unsupervised, statistics-driven manner. Starting with simple, frequent features, we learn the statistically most significant compositions (parts composed of parts), which consequently define the next layer. Parts are learned sequentially, layer after layer, optimally adjusting to the visual data. Lower layers are learned in a category-independent way to obtain complex, yet sharable visual building blocks, which is a crucial step towards a scalable representation. Higher layers of the hierarchy, on the other hand, are constructed by using specific categories, achieving a category representation with a small number of highly generalizable parts that gained their structural flexibility through composition within the hierarchy. Built in this way, new categories can be efficiently and continuously added to the system by adding a small number of parts only in the higher layers. The approach is demonstrated on a large collection of images and a variety of object categories. Detection results confirm the effectiveness and robustness of the learned parts.
Sanja Fidler, Ales Leonardis
CVPR2
2007 Catadioptric Image-based Rendering for Mobile Robot Localization
abstract
We present an approach to view-based mobile robot localization using a X-slits image based rendering (IBR) method for creating novel views from a set of input images. The input images are acquired by a non-central catadioptric sensor mounted on a robot moving on a straight line. We propose to use the IBR for column ordering only, where occlusions in the horizontal direction are modeled and the sensor can be non-central. For the column matching between a query view at an unknown position and virtual views created by IBR, we use correlation of columns.
Hynek Bakstein, Ales Leonardis
ICCV2
2007 High-Dimensional Feature Matching: Employing the Concept of Meaningful Nearest Neighbors
abstract
Matching of high-dimensional features using nearest neighbors search is an important part of image matching methods which are based on local invariant features. In this work we highlight effects pertinent to high-dimensional spaces that are significant for matching, yet have not been explicitly accounted for in previous work. In our approach, we require every nearest neighbor to be meaningful, that is, sufficiently close to a query feature such that it is an out-lier to a background feature distribution. We estimate the background feature distribution from the extended neighborhood of a query feature given by its k nearest neighbors. Based on the concept of meaningful nearest neighbors, we develop a novel high-dimensional feature matching method and evaluate its performance by conducting image matching on two challenging image data sets. A superior performance in terms of accuracy is shown in comparison to several state-of-the-art approaches. Additionally, to make search for k nearest neighbors more efficient, we develop a novel approximate nearest neighbors search method based on sparse coding with an overcomplete basis set that provides a ten-fold speed-up over an exhaustive search even for high dimensional spaces and retains excellent approximation to an exact nearest neighbors search.
Dusan Omercevic, Ondrej Drbohlav, Ales Leonardis
ICCV3
2007 Learning Hierarchical Representations of Object Categories for Robot Vision
Ales Leonardis, Sanja Fidler
ISRR1
2007 Weighted and robust learning of subspace representations
Danijel Skocaj, Ales Leonardis, Horst Bischof
Pattern Recognit.2
2006 Hierarchical Statistical Learning of Generic Parts of Object Structure
abstract
With the growing interest in object categorization various methods have emerged that perform well in this challenging task, yet are inherently limited to only a moderate number of object classes. In pursuit of a more general categorization system this paper proposes a way to overcome the computational complexity encompassing the enormous number of different object categories by exploiting the statistical properties of the highly structured visual world. Our approach proposes a hierarchical acquisition of generic parts of object structure, varying from simple to more complex ones, which stem from the favorable statistics of natural images. The parts recovered in the individual layers of the hierarchy can be used in a top-down manner resulting in a robust statistical engine that could be efficiently used within many of the current categorization systems. The proposed approach has been applied to large image datasets yielding important statistical insights into the generic parts of object structure.
Sanja Fidler, Gregor Berginc, Ales Leonardis
CVPR (1)3
2006 Structural descriptions in human-assisted robot visual learning
abstract
The paper presents an approach to using structural descriptions, obtained through a human-robot tutoring dialogue, as labels for the visual object models a robot learns. The paper shows how structural descriptions enable relating models for different aspects of one and the same object, and how being able to relate descriptions for visual models and discourse referents enables incremental updating of model descriptions through dialogue (either robot- or human initiated). The approach has been implemented in an integrated architecture for human-assisted robot visual learning.
Geert-Jan M. Kruijff, John D. Kelleher, Gregor Berginc, Ales Leonardis
HRI4
2006 Combining Reconstructive and Discriminative Subspace Methods for Robust Classification and Regression by Subsampling
abstract
Linear subspace methods that provide sufficient reconstruction of the data, such as PCA, offer an efficient way of dealing with missing pixels, outliers, and occlusions that often appear in the visual data. Discriminative methods, such as LDA, which, on the other hand, are better suited for classification tasks, are highly sensitive to corrupted data. We present a theoretical framework for achieving the best of both types of methods: An approach that combines the discrimination power of discriminative methods with the reconstruction property of reconstructive methods which enables one to work on subsets of pixels in images to efficiently detect and reject the outliers. The proposed approach is therefore capable of robust classification with a high-breakdown point. We also show that subspace methods, such as CCA, which are used for solving regression tasks, can be treated in a similar manner. The theoretical results are demonstrated on several computer vision tasks showing that the proposed approach significantly outperforms the standard discriminative methods in the case of missing pixels and images containing occlusions and outliers.
Sanja Fidler, Danijel Skocaj, Ales Leonardis
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Panoramic volumes for robot localization
abstract
We propose a method for visual robot localization using a panoramic image volume as the representation from which we can generate views from virtual viewpoints and match them to the current view. We use a geometric image-based rendering formalism in combination with a subspace representation of images, which allows us to synthesize views at arbitrary virtual viewpoints from a compact low-dimensional representation.
Matej Artac, Matjaz Jogan, Ales Leonardis, Hynek Bakstein
IROS3
2004 Illumination insensitive recognition using eigenspaces
Horst Bischof, Horst Wildenauer, Ales Leonardis
Comput. Vis. Image Underst.3
2004 A generalisation of model selection criteria
Bojan Kverh, Ales Leonardis
Pattern Anal. Appl.2
2003 Weighted and Robust Incremental Method for Subspace Learning
abstract
Visual learning is expected to be a continuous and robust process, which treats input images and pixels selectively. In this paper, we present a method for subspace learning, which takes these considerations into account. We present an incremental method, which sequentially updates the principal subspace considering weighted influence of individual images as well as individual pixels within an image. This approach is further extended to enable determination of consistencies in the input data and imputation of the values in inconsistent pixels using the previously acquired knowledge, resulting in a novel incremental, weighted and robust method for subspace learning.
Danijel Skocaj, Ales Leonardis
ICCV2
2003 A Framework for Robust and Incremental Self-Localization of a Mobile Robot
Matjaz Jogan, Matej Artac, Danijel Skocaj, Ales Leonardis
ICVS4
2003 Kernel and subspace methods for computer vision
Ales Leonardis, Horst Bischof
Pattern Recognit.1
2003 Karhunen-Loéve expansion of a set of rotated templates
abstract
In this paper, we propose a novel method for efficiently calculating the eigenvectors of uniformly rotated images of a set of templates. As we show, the images can be optimally approximated by a linear series of eigenvectors which can be calculated without actually decomposing the sample covariance matrix.
Matjaz Jogan, Emil Zagar, Ales Leonardis
IEEE Trans. Image Process.3
2002 A Robust PCA Algorithm for Building Representations from Panoramic Images
Danijel Skocaj, Horst Bischof, Ales Leonardis
ECCV (4)3
2002 Mobile Robot Localization using an Incremental Eigenspace Model
abstract
When using appearance-based recognition for self-localization of mobile robots, the images obtained during the exploration of the environment need to be efficiently stored in the memory. PCA offers means for representing the images in a low-dimensional subspace, which allows for efficient matching and recognition. For active exploration it is necessary to use an incremental method for the computation of the subspace. We propose to use an incremental PCA algorithm with the updating of partial image representations in a way that allows the robot to discard the acquired images immediately after the update. Such a model is open-ended, meaning that we can easily update it with new images. We show that the performance of the proposed method is comparable to the performance of the batch method in terms of compression, computational cost and the precision of localization. We also show that by applying the repetitive learning, the subspace converges to that constructed with the batch method.
Matej Artac, Matjaz Jogan, Ales Leonardis
ICRA3
2002 Multiple eigenspaces
Ales Leonardis, Horst Bischof, Jasna Maver
Pattern Recognit.1
2002 Recognizing 2-tone images in grey-level parametric eigenspaces
Jasna Maver, Ales Leonardis
Pattern Recognit. Lett.2
2001 Illumination Insensitive Eigenspaces
Horst Bischof, Horst Wildenauer, Ales Leonardis
ICCV3
2001 View-based object representations using RBF networks
Horst Bischof, Ales Leonardis
Image Vis. Comput.2
2000 Robust Localization Using Panoramic View-Based Recognition
abstract
The results of earlier studies on the possibility of spatial localization from panoramic images have shown good prospects for view-based methods. The major advantages of these methods are a wide field-of-view, capability of modeling cluttered environments, and flexibility in the learning phase. The redundant information captured in similar views is efficiently handled by the eigenspace approach. However, the standard approaches are sensitive to noise and occlusion. We present a method of view-based localization in a robust framework that solves these problems to a large degree. Experimental results on a large set of real panoramic images demonstrate the effectiveness of the approach and the level of achieved robustness.
Matjaz Jogan, Ales Leonardis
ICPR2
2000 Multiple Eigenspaces by MDL
abstract
We propose an approach to constructing multiple eigenspaces from a set of training images based on the minimum description length (MDL) principle. The main idea is to systematically build a redundant set of eigenspaces, which are treated as hypotheses that are then subject to a selection procedure. The selection procedure, based on the MDL principle, selects the final resulting set of eigenspaces as an optimal representation of the training set. We have tested the proposed method on a number of standard image sets, and the significance of the approach with respect to the recognition rate has been clearly demonstrated.
Ales Leonardis, Horst Bischof
ICPR1
2000 Fuzzy C-Means in an MDL-Framework
abstract
In this paper we present a minimum description length (MDL) framework for fuzzy clustering algorithms. This framework enables us to find an optimal number of cluster centers. We applied our approach to the fuzzy c-means algorithm for which we designed a computationally efficient procedure. We report the results of our approach on a 2D clustering problem and on RGB color image segmentation.
Alexander Selb, Horst Bischof, Ales Leonardis
ICPR3
2000 Range Image Acquisition of Objects with Non-Uniform Albedo Using Structured Light Range Sensor
abstract
We present an approach to acquisition of range images of objects with non-uniform albedo using a structured light sensor. The main idea is to systematically vary the intensity of the light projector and to form high dynamic scale radiance maps. The range images are then formed from these radiance maps. We tested the method on the objects which have surfaces with very different reflectance properties. We demonstrate that the range images obtained from the high dynamic scale radiance maps are of much better quality, than those obtained directly from the original images of a limited dynamic scale.
Danijel Skocaj, Ales Leonardis
ICPR2
2000 Recognizing Objects by Their Appearance Using Eigenimages
Horst Bischof, Ales Leonardis
SOFSEM2
2000 Robust Recognition Using Eigenimages
Ales Leonardis, Horst Bischof
Comput. Vis. Image Underst.1
1999 Panoramic Eigenimages for Spatial Localisation
Matjaz Jogan, Ales Leonardis
CAIP2
1999 Registration of Range Images Based on Segmented Data
Bojan Kverh, Ales Leonardis
CAIP2
1999 MDL Principle for Robust Vector Quantisation
Horst Bischof, Ales Leonardis, Alexander Selb
Pattern Anal. Appl.2
1998 Robust Recognition of Scaled Eigenimages through a Hierarchical Approach
abstract
Recently, we have proposed a new approach to estimation of the coefficients of eigenimages, which is robust against occlusion, varying background, and other types of non-Gaussian noise. In this paper we show that our method for estimating the coefficients can be applied to convolved and subsampled images yielding the same value of the coefficients. This enables an efficient multiresolution approach, where the values of the coefficients can directly be propagated through the scales. This property is used to extend our robust method to the problem of scaled images. We performed extensive experimental evaluations to confirm our theoretical results.
Horst Bischof, Ales Leonardis
CVPR2
1998 MDL-based design of vector quantizers
abstract
We develop a framework for vector quantization networks based on the minimum description length (MDL) principle. This MDL framework is used to derive conditions for the removal of superfluous units from the network. We design a computationally efficient algorithm for finding the optimal number of reference vectors as well as their positions. We illustrate our approach on 2D clustering problems and present applications on image coding.
Horst Bischof, Ales Leonardis
ICPR2
1998 A robust subspace classifier
abstract
In this paper we study the problem of missing features and the issues of robustness of subspace classification methods. We propose a new robust method for subspace classification which can cope with missing features and/or outliers. The main idea of our method is to use a robust projection of the patterns onto a subspace. We demonstrate our approach on cervicomotography data and compare our results to the results obtained by using various decision tree algorithms.
Horst Bischof, Ales Leonardis, Florian Pezzei
ICPR2
1998 Proper scale for modeling visual data
Franc Solina, Ales Leonardis
Image Vis. Comput.2
1998 An efficient MDL-based construction of RBF networks
Ales Leonardis, Horst Bischof
Neural Networks1
1998 Planning Sequences Of Views For 3-D Object Recognition And Pose Determination
Stanislav Kovacic, Ales Leonardis, Franjo Pernus
Pattern Recognit.2
1998 Finding optimal neural networks for land use classification
abstract
The authors present a fully automatic and computationally efficient algorithm based on the minimum description length principle (MDL) for optimizing multilayer perceptron (MLP) classifiers. They demonstrate their method on the problem of multispectral Landsat image classification. They compare their results with a hand-designed MLP and a Gaussian maximum likelihood classifier, in which their method produces better classification accuracy with a smaller number of hidden units.
Horst Bischof, Ales Leonardis
IEEE Trans. Geosci. Remote. Sens.2
1997 Computational Complexity Reduction in Eigenspace Approaches
Ales Leonardis, Horst Bischof
CAIP1
1997 Stereo Matching Using M-Estimators
Christian Menard, Ales Leonardis
CAIP2
1997 Planning Multiple Views for 3-D Object Recognition and Pose Determination
Franjo Pernus, Ales Leonardis, Stanislav Kovacic
CAIP2
1997 Superquadrics for Segmenting and Modeling Range Data
abstract
We present an approach to reliable and efficient recovery of part-descriptions in terms of superquadric models from range data. We show that superquadrics can directly be recovered from unsegmented data, thus avoiding any presegmentation steps (e.g. in terms of surfaces). The approach is based on the recover-and-select paradigm. We present several experiments on real and synthetic range images, where we demonstrate the stability of the results with respect to viewpoint and noise.
Ales Leonardis, Ales Jaklic, Franc Solina
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Dealing with occlusions in the eigenspace approach
abstract
The basic limitations of the current appearance-based matching methods using eigenimages are non-robust estimation of coefficients and inability to cope with problems related to occlusions and segmentation. In this paper we present a new approach which successfully solves these problems. The major novelty of our approach lies in the way how the coefficients of the eigenimages are determined. Instead of computing the coefficients by a projection of the data onto the eigenimages, we extract them by a hypothesize-and-test paradigm using subsets of image points. Competing hypotheses are then subject to a selection procedure based on the Minimum Description Length principle. The approach enables us not only to reject outliers and to deal with occlusions but also to simultaneously use multiple classes of eigenimages.
Ales Leonardis, Horst Bischof
CVPR1
1996 Automatic Selection of Reference Views for Image-Based Scene Representations
Václav Hlavác, Ales Leonardis, Tomás Werner
ECCV (1)2
1996 Robust stereo on multiple resolutions
abstract
Stereo computation is one of the vision problems where the presence of outliers cannot be neglected. Most standard algorithms make unrealistic assumptions about noise distributions, which leads to erroneous results that cannot be corrected in subsequent postprocessing stages. In this paper we present a modification of the standard area-based correlation approach so that it can tolerate a significant number of outliers. The approach exhibits a robust behavior not only in the presence of mismatches but also in the case of depth discontinuities. The confidence measure of the correlation and the number of outliers provide two complementary sources of information which, when implemented in a multiresolution framework, result in a robust and efficient method. We present the results of this approach on a number of synthetic and real images.
Christian Menard, Ales Leonardis
ICPR2
1996 Selection of reference views for image-based representation
abstract
Recently, much attention has been devoted to image-based scene representations. They allow one to construct an arbitrary view of a 3D scene by the interpolation (transfer) from a sparse set of real 2D (reference) images, rather than by rendering an explicit 3D model. While many authors address mainly the purely geometrical aspect of the task, we focus on the problem of how to select the optimal set of reference views. Selection of reference views from a dense set of real primary views is posed as a selection and fitting of parametric models. The selected set must minimize a weighted sum of the number of reference views and the total fit error. We propose two different algorithms for solving this optimization problem. The experimental results on synthetic and real data indicate the feasibility of the approach for 1-DOF camera movement. We discuss the possibility to extend one of the algorithms for more general case.
Tomás Werner, Václav Hlavác, Ales Leonardis, Tomás Pajdla
ICPR3
1996 Choosing Reference Views for Image-Based Representation
Tomás Werner, Václav Hlavác, Ales Leonardis, Tomás Pajdla
SOFSEM3
1996 Detection of diffuse and specular interface reflections and inter-reflections by color image segmentation
Ruzena Bajcsy, Sang Wook Lee, Ales Leonardis
Int. J. Comput. Vis.3
1995 Recognition and Pose Determinatn of 3-D Objects Using Multiple Views
Ales Leonardis, Stanislav Kovacic, Franjo Pernus
CAIP1
1995 ExSel++: A General Framework to Extract Parametric Models
Markus Stricker, Ales Leonardis
CAIP2
1995 Grasping Arbitrarily Shaped 3-D Objects from a Pile
abstract
Presents a reliable and robust approach to the problem of grasping arbitrarily shaped 3-D objects from a pile. The approach adheres to the paradigm of purposive vision, which says that one should only extract as much information as it is needed to perform a certain task, e.g. grasping, while a complete and precise recovery of the shape of the objects is not necessary. The authors show that planar patches obtained by the recover-and-select paradigm contain enough information to enable generating object hypotheses and to estimate grasping points for the objects. The authors present some results for objects with polyhedral as well as with curved surfaces obtained on real range images.
Marjan Trobina, Ales Leonardis
ICRA2
1995 Segmentation of range images as the search for geometric parametric models
Ales Leonardis, Alok Gupta, Ruzena Bajcsy
Int. J. Comput. Vis.1
1994 A Direct Recovery of Superquadric Models in Range Images Using Recover-and-Select Paradigm
Ales Leonardis, Franc Solina, Alenka Macerl
ECCV (1)1
1994 Planning the Optimal Set of Views Using the Max-Min Principle
Jasna Maver, Ales Leonardis, Franc Solina
ECCV (2)2
1994 A Direct Part-Level Segmentation of Range Images Using Volumetric Models
abstract
Volumetric part models play an important part in robotic applications such as grasping, path planning, object avoidance, and modeling kinematic chains. The authors present a novel method for reliable and efficient recovery of part-descriptions in terms of superquadric models from range images. In contrast to usual approaches which perform the recovery of volumetric models in several steps (from curves, surfaces to volumes), the authors show that a direct recovery is possible. This is achieved by combining two existing methods: recover-and-select paradigm and recovery of superquadric models. A redundant set of superquadratics is initiated in the image and only the recovered models resulting in the simplest overall description are selected. The author show the results on several real range images.>
Franc Solina, Ales Leonardis, Alenka Macerl
ICRA2
1994 Two-dimensional object recognition using multiresolution non-information-preserving shape features
Franjo Pernus, Ales Leonardis, Stanislav Kovacic
Pattern Recognit. Lett.2
1993 Planning the Next View Using the Max-Min Principle
Jasna Maver, Ales Leonardis, Franc Solina
CAIP2
1992 Finding Parametric Curves in an Image
Ales Leonardis, Ruzena Bajcsy
ECCV1
1992 Noninformation-preserving shape features at multiple resolution
abstract
The performance of any classification method is limited by the quality of the feature measurements provided. One way to improve the classification is by extending the feature set and selecting the best features out of this set. In this paper a set of noninformation-preserving features based on a multiresolution curve analysis which makes shape features explicit at multiple scales is proposed. An automatic procedure is designed to construct an optimal binary tree which is used to classify the unknown objects.>
Franjo Pernus, Ales Leonardis, Stanko Kovacic
ICPR (2)2
1992 Selective scene modeling
abstract
Proposes an efficient architecture for selective image modeling. The authors give an example in which models of different scale are reconstructed in parallel. It is shown that this redundant representation can effectively be pruned using the criterion of minimum description length. Models that are selected in the final description indicate the appropriate scale of observation.>
Franc Solina, Ales Leonardis
ICPR (1)2
1990 Segmentation as the search for the best description of the image in terms of primitives
abstract
A paradigm is presented for the segmentation of images into piecewise continuous patches. Data aggregation is performed via model recovery in terms of variable-order bivariate polynomials using iterative regression. All the recovered models are candidates for the final description of the data. Selection of the models is achieved through a maximization of the quadratic Boolean problem. The procedure can be adapted to prefer certain kinds of descriptions (one which describes more data points, or has smaller error, or has a lower order model). A fast optimization procedure for model selection is discussed. The approach combines model extraction and model selection in a dynamic way. Partial recovery of the models is followed by the optimization (selection) procedure where only the best models are allowed to develop further. The results are comparable with the results obtained when using the selection module only after all the models are fully recovered, while the computational complexity is significantly reduced. The procedure was tested on real range and intensity images. >
Ales Leonardis, Alok Gupta, Ruzena Bajcsy
ICCV1
1990 Color image segmentation with detection of highlights and local illumination induced by inter-reflections
abstract
A computational model for color image segmentation is proposed based on the physical properties of sensors, illumination lights, and surface reflectances. For image segmentation to depend only on the change of material surface, variations of surface reflection due to illumination, shading, shadows, and highlights should be discounted from a measured image. The authors use the dichromatic model for dielectric materials and develop a metric space of intensity, hue, and saturation in order to interpret various light-surface interactions better. Using the established model for surface reflection, the authors perform color image segmentation based on the material change with detection of highlights and small interreflections between adjacent objects. A reference plate is used to whiten global illumination. The detected interreflections represent local variation of illumination.>
Ruzena Bajcsy, Sang Wook Lee, Ales Leonardis
ICPR (1)3