VLDB 2026 Research / reviewers in the wild / expert
Lars Petersson
dblp:18/5485
· DBLP profile ↗
83ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0002-0103-1904ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 25 since 2021Systems, architecture and hardware · 12 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Subspace-Guided Knowledge Distillation for Efficient Model TransferabstractCompact models can be effectively trained via Knowledge Distillation (KD), where a lightweight student model learns to replicate the behavior of a larger, high-performing teacher. A persistent challenge in KD lies in the misalignment between the representational spaces of teacher and student networks, especially when they differ in architecture or capacity. To address this, we propose Subspace-Driven Knowledge Distillation (SDMD), a novel framework that mitigates representational disparity by projecting features into an indefinite inner product space. This relaxation from traditional Hilbert spaces enables more flexible geometric alignment, capturing transformations such as rotations and reflections that are often necessary for accurate knowledge transfer. By learning a subspace that bridges the semantic gap between teacher and student, SDMD facilitates more effective distillation without increasing model complexity. We validate SDMD through extensive experiments on large-scale image classification (ImageNet-1K) and object detection (COCO), where it consistently outperforms existing distillation methods. Notably, SDMD-trained models not only achieve state-of-the-art results in distilled settings but also surpass the performance of equivalent models trained from scratch, highlighting the strength of our subspace-based alignment strategy. Zeeshan Hayder, Ali Cheraghian, Lars Petersson, Mehrtash Harandi |
WACV | 3 |
| 2026 | PointCaM: Cut-and-Mix for open-set point cloud learning
Shi Qiu 0001, Weihao Li 0005, Saeed Anwar, Mehrtash Harandi, Nick Barnes, Lars Petersson |
Comput. Vis. Image Underst. | 7 |
| 2026 | MFGS: Mask-free Gaussian separation for 3D object reconstructionabstractAccurate 3D reconstruction from multi-view images is a fundamental problem in computer vision. A common acquisition strategy involves placing an object on a rotating turntable while moving the camera to capture it from various viewpoints. In such scenarios, object moves relative to the background, many existing reconstruction methods rely on object masks to separate the foreground from the background. The quality of these masks significantly affects the final reconstruction, yet obtaining high-quality and consistent masks is a challenging and laborious process, especially when controlled environments like green screens are unavailable. To address this limitation, we introduce Mask-free Gaussian Separation (MFGS), a novel method that performs simultaneous object reconstruction and segmentation without requiring any input masks. Our approach builds on Gaussian Splatting and automatically disentangles the scene by extending each Gaussian primitive with a learnable parameter that represents its probability of belonging to the dynamic foreground object. This separation is optimized in a self-supervised manner, optimized by the object and camera transformation constraints. We evaluated MFGS on new synthetic and real-world datasets designed to reflect this challenging capture scenario. Experimental results demonstrate that our mask-free approach significantly outperforms existing methods. Notably, MFGS surpasses the performance of the state-of-the-art method(2DGS) that relies on high-quality segmentation masks, achieving a 27% improvement in novel view synthesis and a 7% improvement in geometry reconstruction. Jinguang Tong, Xuesong Li 0001, Sundaram Muthu, Fahira A. Maken, Lars Petersson, Hongdong Li |
Pattern Recognit. | 5 |
| 2025 | GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstructionabstract3D modeling of highly reflective objects remains challenging due to strong view-dependent appearances. While previous SDF-based methods can recover high-quality meshes, they are often time-consuming and tend to produce over-smoothed surfaces. In contrast, 3D Gaussian Splatting (3DGS) offers the advantage of high speed and detailed real-time rendering, but extracting surfaces from the Gaussians can be noisy due to the lack of geometric constraints. To bridge the gap between these approaches, we propose a novel reconstruction method called GS-2DGS for reflective objects based on 2D Gaussian Splatting (2DGS). Our approach combines the rapid rendering capabilities of Gaussian Splatting with additional geometric information from foundation models. Experimental results on synthetic and real datasets demonstrate that our method significantly outperforms Gaussian-based techniques in terms of reconstruction and relighting and achieves performance comparable to SDF-based methods while being an order of magnitude faster. Code is available at https://github.com/hirotong/GS2DGS Jinguang Tong, Xuesong Li 0001, Fahira A. Maken, Sundaram Muthu, Lars Petersson, Hongdong Li |
CVPR | 5 |
| 2025 | Open Set Label Shift with Test Time Out-of-Distribution ReferenceabstractOpen set label shift (OSLS) occurs when label distributions change from a source to a target distribution, and the target distribution has an additional out-of-distribution (OOD) class. In this work, we build estimators for both source and target open set label distributions using a source domain in-distribution (ID) classifier and an ID/OOD classifier. With reasonable assumptions on the ID/OOD classifier, the estimators are assembled into a sequence of three stages: 1) an estimate of the source label distribution of the OOD class, 2) an EM algorithm for Maximum Likelihood estimates (MLE) of the target label distribution, and 3) an estimate of the target label distribution of OOD class under relaxed assumptions on the OOD classifier. The sampling errors of estimates in 1) and 3) are quantified with a concentration inequality. The estimation result allows us to correct the ID classifier trained on the source distribution to the target distribution without retraining. Experiments on a variety of open set label shift settings demonstrate the effectiveness of our model. Our code is available at https://github.com/ChangkunYe/OpenSetLabelShift. Changkun Ye, Russell Tsuchida, Lars Petersson, Nick Barnes |
CVPR | 3 |
| 2025 | DGNS: Deformable Gaussian Splatting and Dynamic Neural Surface for Monocular Dynamic 3D ReconstructionabstractDynamic scene reconstruction from monocular video is essential for real-world applications. We introduce DGNS, a hybrid framework integrating Deformable Gaussian Splatting and Dynamic Neural Surfaces, effectively addressing dynamic novel-view synthesis and 3D geometry reconstruction simultaneously. During training, depth maps generated by the deformable Gaussian splatting module guide the ray sampling for faster processing and provide depth supervision within the dynamic neural surface module to improve geometry reconstruction. Conversely, the dynamic neural surface directs the distribution of Gaussian primitives around the surface, enhancing rendering quality. In addition, we propose a depth-filtering approach to further refine depth supervision. Extensive experiments conducted on public datasets demonstrate that DGNS achieves state-of-the-art performance in 3D reconstruction, along with competitive results in novel-view synthesis. Xuesong Li 0001, Jinguang Tong, Vivien Rolland, Lars Petersson |
ACM Multimedia | 5 |
| 2025 | BioNet and NeFF: Crop Biomass Prediction from Point Clouds to Drone ImageryabstractCrop biomass offers crucial insights into plant health and yield, making it essential for crop science, farming systems, and agricultural research. However, current measurement methods, which are labor-intensive, destructive, and imprecise, hinder large-scale quantification of this trait. To address this limitation, we present a biomass prediction network (BioNet), designed for adaptation across different data modalities, including point clouds and drone imagery. Our BioNet, utilizing a sparse 3D convolutional neural network (CNN) and a transformer-based prediction module, processes point clouds and other 3D data representations to predict biomass. To further extend BioNet for drone imagery, we integrate a neural feature field (NeFF) module, enabling 3D structure reconstruction and the transformation of 2D semantic features from vision foundation models into the corresponding 3D surfaces. For the point cloud modality, BioNet demonstrates superior performance on two public datasets, with an approximate 6.1% relative improvement (RI) over the state-of-the-art. In the RGB image modality, the combination of BioNet and NeFF achieves a 7.9% RI. Additionally, the NeFF-based approach utilizes inexpensive, portable drone-mounted cameras, providing a scalable solution for large field applications. Xuesong Li 0001, Zeeshan Hayder, Ali Zia, Connor Cassidy, Shiming Liu, Warwick Stiller, Eric A. Stone, Warren Conaty, Lars Petersson, Vivien Rolland |
WACV | 9 |
| 2025 | Facial Expression Recognition with Controlled Privacy Preservation and Feature CompensationabstractFacial expression recognition (FER) systems raise significant privacy concerns due to the potential exposure of sensitive identity information. This paper presents a study on removing identity information while preserving FER capabilities. Drawing on the observation that lowfrequency components predominantly contain identity information and high-frequency components capture expression, we propose a novel two-stream framework that applies privacy enhancement to each component separately. We introduce a controlled privacy enhancement mechanism to optimize performance and a feature compensator to enhance task-relevant features without compromising privacy. Furthermore, we propose a novel privacy-utility trade-off, providing a quantifiable measure of privacy preservation efficacy in closed-set FER tasks. Extensive experiments on the benchmark CREMA-D dataset demonstrate that our framework achieves 78.84 % recognition accuracy with a privacy (facial identity) leakage ratio of only 2.01 %, highlighting its potential for secure and reliable video-based FER applications. We encourage the readers to visit the project page: https://fengxxu.github.io/ppfer/. David Ahmedt-Aristizabal, Lars Petersson, Dadong Wang, Xun Li 0004 |
WACV | 3 |
| 2025 | Attention-Based Real Image RestorationabstractDeep convolutional neural networks perform better on images containing spatially invariant degradations, also known as synthetic degradations; however, their performance is limited on real-degraded photographs and requires multiple-stage network modeling. To advance the practicability of restoration algorithms, this article proposes a novel single-stage blind real image restoration network ( Net) by employing a modular architecture. We use a residual on the residual structure to ease low-frequency information flow and apply feature attention to exploit the channel dependencies. Furthermore, the evaluation in terms of quantitative metrics and visual quality for four restoration tasks, i.e., denoising, super-resolution, raindrop removal, and JPEG compression on 11 real degraded datasets against more than 30 state-of-the-art algorithms, demonstrates the superiority of our Net. We also present the comparison on three synthetically generated degraded datasets for denoising to showcase our method's capability on synthetics denoising. The codes, trained models, and results are available on https://github.com/saeed-anwar/R2Net. Saeed Anwar, Nick Barnes, Lars Petersson |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Backpropagation-free Network for 3D Test-time AdaptationabstractReal-world systems often encounter new data over time, which leads to experiencing target domain shifts. Existing Test- Time Adaptation (TTA) methods tend to apply computationally heavy and memory-intensive backpropagation-based approaches to handle this. Here, we propose a novel method that uses a backpropagation-free approach for TTA for the specific case of 3D data. Our model uses a two-stream architecture to maintain knowledge about the source domain as well as complementary target-domain-specific information. The backpropagation-free property of our model helps address the well-known forgetting prob-lem and mitigates the error accumulation issue. The pro-posed method also eliminates the need for the usually noisy process of pseudo-labeling and reliance on costly self-supervised training. Moreover, our method leverages sub-space learning, effectively reducing the distribution vari-ance between the two domains. Furthermore, the source-domain-specific and the target-domain-specific streams are aligned using a novel entropy-based adaptive fusion strat-egy. Extensive experiments on popular benchmarks demon-strate the effectiveness of our method. The code will be available at https://github.com/abie-e/BFTT3D. Yanshuo Wang, Ali Cheraghian, Zeeshan Hayder, Sameera Ramasinghe, Shafin Rahman, David Ahmedt-Aristizabal, Xuesong Li 0001, Lars Petersson, Mehrtash Harandi |
CVPR | 9 |
| 2024 | Canonical Shape Projection Is All You Need for 3D Few-Shot Class Incremental Learning
Ali Cheraghian, Zeeshan Hayder, Sameera Ramasinghe, Shafin Rahman, Javad Jafaryahya, Lars Petersson, Mehrtash Harandi |
ECCV (41) | 6 |
| 2024 | Continual Test-time Domain Adaptation via Dynamic Sample SelectionabstractThe objective of Continual Test-time Domain Adaptation (CTDA) is to gradually adapt a pre-trained model to a sequence of target domains without accessing the source data. This paper proposes a Dynamic Sample Selection (DSS) method for CTDA. DSS consists of dynamic thresholding, positive learning, and negative learning processes. Traditionally, models learn from unlabeled unknown environment data and equally rely on all samples’ pseudo-labels to update their parameters through self-training. However, noisy predictions exist in these pseudo-labels, so all samples are not equally trustworthy. Therefore, in our method, a dynamic thresholding module is first designed to select suspected low-quality from high-quality samples. The selected low-quality samples are more likely to be wrongly predicted. Therefore, we apply joint positive and negative learning on both high- and low-quality samples to reduce the risk of using wrong information. We conduct extensive experiments that demonstrate the effectiveness of our proposed method for CTDA in the image domain, outperforming the state-of-the-art results. Furthermore, our approach is also evaluated in the 3D point cloud domain, showcasing its versatility and potential for broader applicability. Yanshuo Wang, Ali Cheraghian, Shafin Rahman, David Ahmedt-Aristizabal, Lars Petersson, Mehrtash Harandi |
WACV | 6 |
| 2024 | Label Shift Estimation for Class-Imbalance Problem: A Bayesian ApproachabstractAs a type of distribution shift, label shift occurs when the source and target domains have different label distributions $\mathbb{P}(Y)$ but identical conditional distributions of data given labels $\mathbb{P}(X|Y)$. Under a Bayesian framework, we propose a novel Maximum A Posteriori (MAP) model and a novel posterior sampling model for the label shift problem. We prove the MAP objective admits a unique optimum and derive an EM algorithm that converges to the global optimum. We propose a novel Adaptive Prior Learning (APL) model to adaptively select prior parameters given data. We use the Markov Chain Monte Carlo (MCMC) method in our posterior sampling model to estimate and correct for label shift. Our methods can effectively resolve class imbalance problems on large-scale datasets without fine-tuning the classifier. Experiments show that our model outperforms existing methods on a variety of label shift settings. Our code is available at https://github.com/ChangkunYe/MAPLS/. Changkun Ye, Russell Tsuchida, Lars Petersson, Nick Barnes |
WACV | 3 |
| 2024 | Curved Geometric Networks for Visual Anomaly RecognitionabstractLearning a latent embedding to understand the underlying nature of data distribution is often formulated in Euclidean spaces with zero curvature. However, the success of the geometry constraints, posed in the embedding space, indicates that curved spaces might encode more structural information, leading to better discriminative power and hence richer representations. In this work, we investigate the benefits of the curved space for analyzing anomalous, open-set, or out-of-distribution (OOD) objects in data. This is achieved by considering embeddings via three geometry constraints, namely, spherical geometry (with positive curvature), hyperbolic geometry (with negative curvature), or mixed geometry (with both positive and negative curvatures). Three geometric constraints can be chosen interchangeably in a unified design, given the task at hand. Tailored for the embeddings in the curved space, we also formulate functions to compute the anomaly score. Two types of geometric modules (i.e., geometric-in-one (GiO) and geometric-in-two (GiT) models) are proposed to plug in the original Euclidean classifier, and anomaly scores are computed from the curved embeddings. We evaluate the resulting designs under a diverse set of visual recognition scenarios, including image detection (multiclass OOD detection and one-class anomaly detection) and segmentation (multiclass anomaly segmentation and one-class anomaly segmentation). The empirical results show the effectiveness of our proposal through consistent improvement over various scenarios. The code is made available at https://github.com/JHome1/GiO-GiT. Pengfei Fang, Weihao Li 0005, Junlin Han, Lars Petersson, Mehrtash Harandi |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | GOSS: towards generalized open-set semantic segmentationabstractAbstract In this paper, we extend Open-set Semantic Segmentation (OSS) into a new image segmentation task called Generalized Open-set Semantic Segmentation (GOSS). Previously, with well-known OSS, the intelligent agents only detect unknown regions without further processing, limiting their perception capacity of the environment. It stands to reason that further analysis of the detected unknown pixels would be beneficial for agents’ decision-making. Therefore, we propose GOSS, which holistically unifies the abilities of two well-defined segmentation tasks, i.e. OSS and generic segmentation. Specifically, GOSS classifies pixels as belonging to known classes, and clusters (or groups) of pixels of unknown class are labelled as such. We propose a metric that balances the pixel classification and clustering aspects to evaluate this newly expanded task. Moreover, we build benchmark tests on existing datasets and propose neural architectures as baselines. Our experiments on multiple benchmarks demonstrate the effectiveness of our baselines. Code is made available at https://github.com/JHome1/GOSS_Segmentor . Weihao Li 0005, Junlin Han, Jiyang Zheng, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
Vis. Comput. | 7 |
| 2024 | Publisher Correction: GOSS: towards generalized open-set semantic segmentation
Weihao Li 0005, Junlin Han, Jiyang Zheng, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
Vis. Comput. | 7 |
| 2023 | Hyperbolic Audio-visual Zero-shot LearningabstractAudio-visual zero-shot learning aims to classify samples consisting of a pair of corresponding audio and video sequences from classes that are not present during training. An analysis of the audio-visual data reveals a large degree of hyperbolicity, indicating the potential benefit of using a hyperbolic transformation to achieve curvature-aware geometric learning, with the aim of exploring more complex hierarchical data structures for this task. The proposed approach employs a novel loss function that incorporates cross-modality alignment between video and audio features in the hyperbolic space. Additionally, we explore the use of multiple adaptive curvatures for hyperbolic projections. The experimental results on this very challenging task demonstrate that our proposed hyperbolic approach for zero-shot learning outperforms the SOTA method on three datasets: VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL achieving a harmonic mean (HM) improvement of around 3.0%, 7.0%, and 5.3%, respectively. Zeeshan Hayder, Junlin Han, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
ICCV | 6 |
| 2023 | Weakly-supervised Point Cloud Instance Segmentation with Geometric PriorsabstractThis paper investigates how to leverage more readily acquired annotations, i.e., 3D bounding boxes instead of dense point-wise labels, for instance segmentation. We propose a Weakly-supervised point cloud Instance Segmentation framework with Geometric Priors (WISGP) that allows segmentation models to be trained with 3D bounding boxes of instances. Considering intersections among bounding boxes in a scene would result in ambiguous la- bels, we first group points into two sets, i.e., univocal and equivocal sets, indicating the certainty of a 3D point belonging to an instance, respectively. Specifically, 3D points with clear labels belong to the univocal set while the rest are grouped into the equivocal set. To assign reliable labels to points in the equivocal set, we design a Geometry-guided Label Propagation (GLP) scheme that progressively propagates labels to linked points based on geometric structure, e.g., polygon meshes and superpoints. Afterwards, we train an instance segmentation model with the univocal points and equivocal points labeled by GLP, and then employ it to assign pseudo labels for the remainder of the unlabeled points. Lastly, we retrain the model with all the labeled points to achieve better instance segmentation performance. Experiments on large-scale datasets ScanNet-v2 and S3DIS demonstrate that WISGP is superior to competing weakly-supervised algorithms and even on par with a few fully-supervised ones. Heming Du, Xin Yu 0002, Farookh Khadeer Hussain, Mohammad Ali Armin, Lars Petersson, Weihao Li 0005 |
WACV | 5 |
| 2023 | Poincaré Kernels for Hyperbolic Representations
Pengfei Fang, Mehrtash Harandi, Zhen-Zhong Lan, Lars Petersson |
Int. J. Comput. Vis. | 4 |
| 2022 | Transcribing Natural Languages for the Deaf via Neural Editing ProgramsabstractThis work studies the task of glossification, of which the aim is to em transcribe natural spoken language sentences for the Deaf (hard-of-hearing) community to ordered sign language glosses. Previous sequence-to-sequence language models trained with paired sentence-gloss data often fail to capture the rich connections between the two distinct languages, leading to unsatisfactory transcriptions. We observe that despite different grammars, glosses effectively simplify sentences for the ease of deaf communication, while sharing a large portion of vocabulary with sentences. This has motivated us to implement glossification by executing a collection of editing actions, e.g. word addition, deletion, and copying, called editing programs, on their natural spoken language counterparts. Specifically, we design a new neural agent that learns to synthesize and execute editing programs, conditioned on sentence contexts and partial editing results. The agent is trained to imitate minimal editing programs, while exploring more widely the program space via policy gradients to optimize sequence-wise transcription quality. Results show that our approach outperforms previous glossification models by a large margin, improving the BLEU-4 score from 16.45 to 18.89 on RWTH-PHOENIX-WEATHER-2014T and from 18.38 to 21.30 on CSL-Daily. Dongxu Li 0003, Liu Liu 0009, Yiran Zhong, Lars Petersson, Hongdong Li |
AAAI | 6 |
| 2022 | Blind Image Decomposition
Junlin Han, Weihao Li 0005, Pengfei Fang, Chunyi Sun, Mohammad Ali Armin, Lars Petersson, Hongdong Li |
ECCV (18) | 7 |
| 2022 | Declarative nets that are equilibrium models
Russell Tsuchida, Suk Yee Yong, Mohammad Ali Armin, Lars Petersson, Cheng Soon Ong |
ICLR | 4 |
| 2022 | You Only Cut Once: Boosting Data Augmentation with a Single CutabstractWe present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying YOCO improves the diversity of the augmentation per sample and encourages neural networks to recognize objects from partial information. YOCO enjoys the properties of parameter-free, easy usage, and boosting almost all augmentations for free. Thorough experiments are conducted to evaluate its effectiveness. We first demonstrate that YOCO can be seamlessly applied to varying data augmentations, neural network architectures, and brings performance gains on CIFAR and ImageNet classification tasks, sometimes surpassing conventional image-level augmentation by large margins. Moreover, we show YOCO benefits contrastive pre-training toward a more powerful representation that can be better transferred to multiple downstream tasks. Finally, we study a number of variants of YOCO and empirically analyze the performance for respective settings. Junlin Han, Pengfei Fang, Weihao Li 0005, Mohammad Ali Armin, Ian D. Reid 0001, Lars Petersson, Hongdong Li |
ICML | 7 |
| 2022 | Efficient Gaussian Process Model on Class-Imbalanced Datasets for Generalized Zero-Shot LearningabstractZero-Shot Learning (ZSL) models aim to classify object classes that are not seen during the training process. However, the problem of class imbalance is rarely discussed, despite its presence in several ZSL datasets. In this paper, we propose a Neural Network model that learns a latent feature embedding and a Gaussian Process (GP) regression model that predicts latent feature prototypes of unseen classes. A calibrated classifier is then constructed for ZSL and Generalized ZSL tasks. Our Neural Network model is trained efficiently with a simple training strategy that mitigates the impact of class-imbalanced training data. The model has an average training time of 5 minutes and can achieve state-of-the-art (SOTA) performance on imbalanced ZSL benchmark datasets like AWA2, AWA1 and APY, while having relatively good performance on the SUN and CUB datasets. Changkun Ye, Nick Barnes, Lars Petersson, Russell Tsuchida |
ICPR | 3 |
| 2022 | Biomass Prediction with 3D Point Clouds from LiDARabstractWith population growth and a shrinking rural workforce, agricultural technologies have become increasingly important. Above-ground biomass (AGB) is a key trait relevant to breeding, agronomy and crop physiology field experiments. However, measuring the biomass of a cereal plot requires cutting, drying and weighing processes, which are laborious, expensive and destructive tasks. This paper proposes a non-destructive and high-throughput method to predict biomass from field samples based on Light Detection and Ranging (LiDAR). Unlike previous methods that are based on the density of a point cloud or plant height, our biomass prediction network (BioNet) additionally considers plant structure. Our BioNet contains three modules: 1) a completion module to predict missing points due to canopy occlusion; 2) a regularization module to regularize the neural representation of the whole plot; and 3) a projection module to learn the salient structures from a bird’s eye view of the point cloud. An attention-based fusion block is used to achieve final biomass predictions. In addition, the complete dataset, including hand-measured biomass and LiDAR data, is made available to the community. Experiments show that our BioNet achieves ≈ 33% improvement over current state-of-the-art methods. Liyuan Pan, Liu Liu 0009, Anthony G. Condon, Gonzalo M. Estavillo, Robert Coe, Geoff Bull, Eric A. Stone, Lars Petersson, Vivien Rolland |
WACV | 8 |
| 2022 | Towards a Robust Differentiable Architecture Search under Label NoiseabstractNeural Architecture Search (NAS) is the game changer in designing robust neural architectures. Architectures designed by NAS outperform or compete with the best manual network designs in terms of accuracy, size, memory footprint and FLOPs. That said, previous studies focus on developing NAS algorithms for clean high quality data, a restrictive and somewhat unrealistic assumption. In this paper, focusing on the differentiable NAS algorithms, we show that vanilla NAS algorithms suffer from a performance loss if class labels are noisy. To combat this issue, we make use of the principle of information bottleneck as a regularizer. This leads us to develop a noise injecting operation that is included during the learning process, preventing the network from learning from noisy samples. Our empirical evaluations show that the noise injecting operation does not degrade the performance of the NAS algorithm if the data is indeed clean. In contrast, if the data is noisy, the architecture learned by our algorithm comfortably outperforms algorithms specifically equipped with sophisticated mechanisms to learn in the presence of label noise. In contrast to many algorithms designed to work in the presence of noisy labels, prior knowledge about the properties of the noise and its characteristics are not required for our algorithm. Christian Simon, Piotr Koniusz, Lars Petersson, Mehrtash Harandi |
WACV | 3 |
| 2022 | Zero-Shot Learning on 3D Point Cloud Objects and Beyond
Ali Cheraghian, Shafin Rahman, Townim F. Chowdhury, Dylan Campbell, Lars Petersson |
Int. J. Comput. Vis. | 5 |
| 2022 | Attention in Attention Networks for Person RetrievalabstractThis paper generalizes the Attention in Attention (AiA) mechanism, in P. Fang et al., 2019 by employing explicit mapping in reproducing kernel Hilbert spaces to generate attention values of the input feature map. The AiA mechanism models the capacity of building inter-dependencies among the local and global features by the interaction of inner and outer attention modules. Besides a vanilla AiA module, termed linear attention with AiA, two non-linear counterparts, namely, second-order polynomial attention and Gaussian attention, are also proposed to utilize the non-linear properties of the input features explicitly, via the second-order polynomial kernel and Gaussian kernel approximation. The deep convolutional neural network, equipped with the proposed AiA blocks, is referred to as Attention in Attention Network (AiA-Net). The AiA-Net learns to extract a discriminative pedestrian representation, which combines complementary person appearance and corresponding part features. Extensive ablation studies verify the effectiveness of the AiA mechanism and the use of non-linear features hidden in the feature map for attention design. Furthermore, our approach outperforms current state-of-the-art by a considerable margin across a number of benchmarks. In addition, state-of-the-art performance is also achieved in the video person retrieval task with the assistance of the proposed AiA blocks. Pengfei Fang, Jieming Zhou, Soumava Kumar Roy, Pan Ji, Lars Petersson, Mehrtash Harandi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Semantic-Aware Knowledge Distillation for Few-Shot Class-Incremental LearningabstractFew-shot class incremental learning (FSCIL) portrays the problem of learning new concepts gradually, where only a few examples per concept are available to the learner. Due to the limited number of examples for training, the techniques developed for standard incremental learning cannot be applied verbatim to FSCIL. In this work, we introduce a distillation algorithm to address the problem of FSCIL and propose to make use of semantic information during training. To this end, we make use of word embeddings as semantic information which is cheap to obtain and which facilitate the distillation process. Furthermore, we propose a method based on an attention mechanism on multiple parallel embeddings of visual data to align visual and semantic vectors, which reduces issues related to catastrophic forgetting. Via experiments on MiniImageNet, CUB200, and CIFAR100 dataset, we establish new state-of-the-art results by outperforming existing approaches. Ali Cheraghian, Shafin Rahman, Pengfei Fang, Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
CVPR | 5 |
| 2021 | Reinforced Attention for Few-Shot Learning and BeyondabstractFew-shot learning aims to correctly recognize query samples from unseen classes given a limited number of support samples, often by relying on global embeddings of images. In this paper, we propose to equip the backbone network with an attention agent, which is trained by reinforcement learning. The policy gradient algorithm is employed to train the agent towards adaptively localizing the representative regions on feature maps over time. We further design a reward function based on the prediction of the held-out data, thus helping the attention mechanism to generalize better across the unseen classes. The extensive experiments show, with the help of the reinforced attention, that our embedding network has the capability to progressively generate a more discriminative representation in few-shot learning. Moreover, experiments on the task of image classification also show the effectiveness of the proposed design. Pengfei Fang, Weihao Li 0005, Tong Zhang 0023, Christian Simon, Mehrtash Harandi, Lars Petersson |
CVPR | 7 |
| 2021 | Contextually Plausible and Diverse 3D Human Motion PredictionabstractWe tackle the task of diverse 3D human motion prediction, that is, forecasting multiple plausible future 3D poses given a sequence of observed 3D poses. In this context, a popular approach consists of using a Conditional Variational Autoencoder (CVAE). However, existing approaches that do so either fail to capture the diversity in human motion, or generate diverse but semantically implausible continuations of the observed motion. In this paper, we address both of these problems by developing a new variational framework that accounts for both diversity and context of the generated future motion. To this end, and in contrast to existing approaches, we condition the sampling of the latent variable that acts as source of diversity on the representation of the past observation, thus encouraging it to carry relevant information. Our experiments demonstrate that our approach yields motions not only of higher quality while retaining diversity, but also that preserve the contextual information contained in the observed motion. Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Lars Petersson, Stephen Gould, Mathieu Salzmann |
ICCV | 3 |
| 2021 | Synthesized Feature based Few-Shot Class-Incremental Learning on a Mixture of SubspacesabstractFew-shot class incremental learning (FSCIL) aims to incrementally add sets of novel classes to a well-trained base model in multiple training sessions with the restriction that only a few novel instances are available per class. While learning novel classes, FSCIL methods gradually forget base (old) class training and overfit to a few novel class samples. Existing approaches have addressed this problem by computing the class prototypes from the visual or semantic word vector domain. In this paper, we propose addressing this problem using a mixture of subspaces. Subspaces define the cluster structure of the visual domain and help to describe the visual and semantic domain considering the overall distribution of the data. Additionally, we propose to employ a variational autoencoder (VAE) to generate synthesized visual samples for augmenting pseudo-feature while learning novel classes incrementally. The combined effect of the mixture of subspaces and synthesized features reduces the forgetting and overfitting problem of FSCIL. Extensive experiments on three image classification datasets show that our proposed method achieves competitive results compared to state-of-the-art methods. Ali Cheraghian, Shafin Rahman, Sameera Ramasinghe, Pengfei Fang, Christian Simon, Lars Petersson, Mehrtash Harandi |
ICCV | 6 |
| 2021 | Kernel Methods in Hyperbolic SpacesabstractEmbedding data in hyperbolic spaces has proven beneficial for many advanced machine learning applications such as image classification and word embeddings. However, working in hyperbolic spaces is not without difficulties as a result of its curved geometry (e.g., computing the Frechet mean of a set of points requires an iterative algorithm). Furthermore, in Euclidean spaces, one can resort to kernel machines that not only enjoy rich theoretical properties but that can also lead to superior representational power (e.g., infinite-width neural networks). In this paper, we introduce positive definite kernel functions for hyperbolic spaces. This brings in two major advantages, 1. kernelization will pave the way to seamlessly benefit from kernel machines in conjunction with hyperbolic embeddings, and 2. the rich structure of the Hilbert spaces associated with kernel machines enables us to simplify various operations involving hyperbolic data. That said, identifying valid kernel functions on curved spaces is not straightforward and is indeed considered an open problem in the learning community. Our work addresses this gap and develops several valid positive definite kernels in hyperbolic spaces, including the universal ones (e.g., RBF). We comprehensively study the proposed kernels on a variety of challenging tasks including few-shot learning, zero-shot learning, person reidentification and knowledge distillation, showing the superiority of the kernelization for hyperbolic representations. Pengfei Fang, Mehrtash Harandi, Lars Petersson |
ICCV | 3 |
| 2021 | Single Underwater Image Restoration by Contrastive LearningabstractUnderwater image restoration attracts significant attention due to its importance in unveiling the underwater world. This paper elaborates on a novel method that achieves state-of-the-art results for underwater image restoration based on the unsupervised image-to-image translation framework. We design our method by leveraging from contrastive learning and generative adversarial networks to maximize mutual information between raw and restored images. Additionally, we release a large-scale real underwater image dataset to support both paired and unpaired training modules. Extensive experiments with comparisons to recent approaches further demonstrate the superiority of our proposed method. Junlin Han, Mehrdad Shoeiby, Timothy J. Malthus, Elizabeth J. Botha, Janet M. Anstee, Saeed Anwar, Lars Petersson, Mohammad Ali Armin |
IGARSS | 8 |
| 2021 | Set Augmented Triplet Loss for Video Person Re-IdentificationabstractModern video person re-identification (re-ID) machines are often trained using a metric learning approach, supervised by a triplet loss. The triplet loss used in video re-ID is usually based on so-called clip features, each aggregated from a few frame features. In this paper, we propose to model the video clip as a set and instead study the distance between sets in the corresponding triplet loss. In contrast to the distance between clip representations, the distance between clip sets considers the pair-wise similarity of each element (i.e., frame representation) between two sets. This allows the network to directly optimize the feature representation at a frame level. Apart from the commonly-used set distance metrics (e.g., ordinary distance and Hausdorff distance), we further propose a hybrid distance metric, tailored for the set-aware triplet loss. Also, we propose a hard positive set construction strategy using the learned class prototypes in a batch. Our proposed method achieves state-of-the-art results across several standard benchmarks, demonstrating the advantages of the proposed method. Pengfei Fang, Pan Ji, Lars Petersson, Mehrtash Harandi |
WACV | 3 |
| 2021 | Discrepant collaborative training by Sinkhorn divergences
Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
Image Vis. Comput. | 3 |
| 2021 | Multi-FAN: multi-spectral mosaic super-resolution via multi-scale feature aggregation network
Mehrdad Shoeiby, Mohammad Sadegh Ali Akbarian, Saeed Anwar, Lars Petersson |
Mach. Vis. Appl. | 4 |
| 2020 | Channel Recurrent Attention Networks for Video Pedestrian Retrieval
Pengfei Fang, Pan Ji, Jieming Zhou, Lars Petersson, Mehrtash Harandi |
ACCV (6) | 4 |
| 2020 | A Stochastic Conditioning Scheme for Diverse Human Motion PredictionabstractHuman motion prediction, the task of predicting future 3D human poses given a sequence of observed ones, has been mostly treated as a deterministic problem. However, human motion is a stochastic process: Given an observed sequence of poses, multiple future motions are plausible. Existing approaches to modeling this stochasticity typically combine a random noise vector with information about the previous poses. This combination, however, is done in a deterministic manner, which gives the network the flexibility to learn to ignore the random noise. Alternatively, in this paper, we propose to stochastically combine the root of variations with previous pose information, so as to force the model to take the noise into account. We exploit this idea for motion prediction by incorporating it into a recurrent encoder-decoder network with a conditional variational autoencoder block that learns to exploit the perturbations. Our experiments on two large-scale motion prediction datasets demonstrate that our model yields high-quality pose sequences that are much more diverse than those from state-of-the-art stochastic motion prediction techniques. Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Lars Petersson, Stephen Gould |
CVPR | 4 |
| 2020 | Transferring Cross-Domain Knowledge for Video Sign Language RecognitionabstractWord-level sign language recognition (WSLR) is a fundamental task in sign language interpretation. It requires models to recognize isolated sign words from videos. However, annotating WSLR data needs expert knowledge, thus limiting WSLR dataset acquisition. On the contrary, there are abundant subtitled sign news videos on the internet. Since these videos have no word-level annotation and exhibit a large domain gap from isolated signs, they cannot be directly used for training WSLR models. We observe that despite the existence of a large domain gap, isolated and news signs share the same visual concepts, such as hand gestures and body movements. Motivated by this observation, we propose a novel method that learns domain-invariant visual concepts and fertilizes WSLR models by transferring knowledge of subtitled news sign to them. To this end, we extract news signs using a base WSLR model, and then design a classifier jointly trained on news and isolated signs to coarsely align these two domain features. In order to learn domain-invariant features within each class and suppress domain-specific features, our method further resorts to an external memory to store the class centroids of the aligned news signs. We then design a temporal attention based on the learnt descriptor to improve recognition performance. Experimental results on standard WSLR datasets show that our method outperforms previous state-of-the-art methods significantly. We also demonstrate the effectiveness of our method on automatically localizing signs from sign news, achieving 28.1 for [email protected]. Dongxu Li 0003, Xin Yu 0002, Lars Petersson, Hongdong Li |
CVPR | 4 |
| 2020 | Transductive Zero-Shot Learning for 3D Point Cloud ClassificationabstractZero-shot learning, the task of learning to recognize new classes not seen during training, has received considerable attention in the case of 2D image classification. However despite the increasing ubiquity of 3D sensors, the corresponding 3D point cloud classification problem has not been meaningfully explored and introduces new challenges. This paper extends, for the first time, transductive ZeroShot Learning (ZSL) and Generalized Zero-Shot Learning (GZSL) approaches to the domain of 3D point cloud classification. To this end, a novel triplet loss is developed that takes advantage of unlabeled test data. While designed for the task of 3D point cloud classification, the method is also shown to be applicable to the more common use-case of 2D image classification. An extensive set of experiments is carried out, establishing state-of-the-art for ZSL and GZSL in the 3D point cloud domain, as well as demonstrating the applicability of the approach to the image domain.1 Ali Cheraghian, Shafin Rahman, Dylan Campbell, Lars Petersson |
WACV | 4 |
| 2020 | Learning from Noisy Labels via Discrepant Collaborative TrainingabstractNoise is ubiquitous in the world around us. Difficulty in estimating the noise within a dataset makes learning from such a dataset a difficult and challenging task. In this paper, we propose a novel and effective learning framework in order to alleviate the adverse effects of noise within a dataset. Towards this aim, we modify a collaborative training framework to utilize discrepancy constraints between respective feature extractors enabling the learning of distinct, yet discriminative features, pacifying the adverse effects of noise. Empirical results of our proposed algorithm, Discrepant Collaborative Training (DCT), achieve competitive results against several current state-of-the-art algorithms across MNIST, CIFAR10 and CIFAR100, as well as large fine-grained image classification datasets such as CUBS-200-2011 and CARS196 for different levels of noise. Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
WACV | 3 |
| 2020 | Super-resolved Chromatic Mapping of Snapshot Mosaic Image Sensors via a Texture Sensitive Residual NetworkabstractThis paper introduces a novel method to simultaneously super-resolve and colour-predict images acquired by snapshot mosaic sensors. These sensors allow for spectral images to be acquired using low-power, small form factor, solid-state CMOS sensors that can operate at video frame rates without the need for complex optical setups. Despite their desirable traits, their main drawback stems from the fact that the spatial resolution of the imagery acquired by these sensors is low. Moreover, chromatic mapping in snapshot mosaic sensors is not straightforward since the bands delivered by the sensor tend to be narrow and unevenly distributed across the range in which they operate. We tackle this drawback as applied to chromatic mapping by using a residual channel attention network equipped with a texture sensitive block. Our method significantly outperforms the traditional approach of interpolating the image and, afterwards, applying a colour matching function. This work establishes state-of-the-art in this domain while also making available to the research community a dataset containing 296 registered stereo multi-spectral/RGB images pairs. Mehrdad Shoeiby, Lars Petersson, Mohammad Ali Armin, Mohammad Sadegh Ali Akbarian, Antonio Robles-Kelly |
WACV | 2 |
| 2020 | Cross-Correlated Attention Networks for Person Re-Identification
Jieming Zhou, Soumava Kumar Roy, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
Image Vis. Comput. | 5 |
| 2020 | Globally-Optimal Inlier Set Maximisation for Camera Pose and Correspondence EstimationabstractEstimating the 6-DoF pose of a camera from a single image relative to a 3D point-set is an important task for many computer vision applications. Perspective-n-point solvers are routinely used for camera pose estimation, but are contingent on the provision of good quality 2D-3D correspondences. However, finding cross-modality correspondences between 2D image points and a 3D point-set is non-trivial, particularly when only geometric information is known. Existing approaches to the simultaneous pose and correspondence problem use local optimisation, and are therefore unlikely to find the optimal solution without a good pose initialisation, or introduce restrictive assumptions. Since a large proportion of outliers and many local optima are common for this problem, we instead propose a robust and globally-optimal inlier set maximisation approach that jointly estimates the optimal camera pose and correspondences. Our approach employs branch-and-bound to search the 6D space of camera poses, guaranteeing global optimality without requiring a pose prior. The geometry of SE(3) is used to find novel upper and lower bounds on the number of inliers and local optimisation is integrated to accelerate convergence. The algorithm outperforms existing approaches on challenging synthetic and real datasets, reliably finding the global optimum, with a GPU implementation greatly reducing runtime. Dylan Campbell, Lars Petersson, Laurent Kneip, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Mitigating the Hubness Problem for Zero-Shot Learning of 3D Objects
Ali Cheraghian, Shafin Rahman, Dylan Campbell, Lars Petersson |
BMVC | 4 |
| 2019 | The Alignment of the Spheres: Globally-Optimal Spherical Mixture Alignment for Camera Pose EstimationabstractDetermining the position and orientation of a calibrated camera from a single image with respect to a 3D model is an essential task for many applications. When 2D-3D correspondences can be obtained reliably, perspective-n-point solvers can be used to recover the camera pose. However, without the pose it is non-trivial to find cross-modality correspondences between 2D images and 3D models, particularly when the latter only contains geometric information. Consequently, the problem becomes one of estimating pose and correspondences jointly. Since outliers and local optima are so prevalent, robust objective functions and global search strategies are desirable. Hence, we cast the problem as a 2D-3D mixture model alignment task and propose the first globally-optimal solution to this formulation under the robust L2 distance between mixture distributions. We derive novel bounds on this objective function and employ branch-and-bound to search the 6D space of camera poses, guaranteeing global optimality without requiring a pose estimate. To accelerate convergence, we integrate local optimization, implement GPU bound computations, and provide an intuitive way to incorporate side information such as semantic labels. The algorithm is evaluated on challenging synthetic and real datasets, outperforming existing approaches and reliably converging to the global optimum. Dylan Campbell, Lars Petersson, Laurent Kneip, Hongdong Li, Stephen Gould |
CVPR | 2 |
| 2019 | Bilinear Attention Networks for Person RetrievalabstractThis paper investigates a novel Bilinear attention (Bi-attention) block, which discovers and uses second order statistical information in an input feature map, for the purpose of person retrieval. The Bi-attention block uses bilinear pooling to model the local pairwise feature interactions along each channel, while preserving the spatial structural information. We propose an Attention in Attention (AiA) mechanism to build inter-dependency among the second order local and global features with the intent to make better use of, or pay more attention to, such higher order statistical relationships. The proposed network, equipped with the proposed Bi-attention is referred to as Bilinear ATtention network (BAT-net). Our approach outperforms current state-of-the-art by a considerable margin across the standard benchmark datasets (e.g., CUHK03, Market-1501, DukeMTMC-reID and MSMT17). Pengfei Fang, Jieming Zhou, Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
ICCV | 4 |
| 2019 | 3DCapsule: Extending the Capsule Architecture to Classify 3D Point CloudsabstractThis paper introduces the 3DCapsule, which is a 3D extension of the recently introduced Capsule concept that makes it applicable to unordered point sets. The original Capsule relies on the existence of a spatial relationship between the elements in the feature map it is presented with, whereas in point permutation invariant formulations of 3D point set classification methods, such relationships are typically lost. Here, a new layer called ComposeCaps is introduced that, in lieu of a spatially relevant feature mapping, learns a new mapping that can be exploited by the 3DCapsule. Previous works in the 3D point set classification domain have focused on other parts of the architecture, whereas instead, the 3DCapsule is a drop-in replacement of the commonly used fully connected classifier. It is demonstrated via an ablation study, that when the 3DCapsule is applied to recent 3D point set classification architectures, it consistently shows an improvement, in particular when subjected to noisy data. Similarly, the ComposeCaps layer is evaluated and demonstrates an improvement over the baseline. In an apples-to-apples comparison against state-of-the-art methods, again, better performance is demonstrated by the 3DCapsule. Ali Cheraghian, Lars Petersson |
WACV | 2 |
| 2018 | VIENA ^2 : A Driving Anticipation Dataset
Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Basura Fernando, Lars Petersson, Lars Andersson |
ACCV (1) | 5 |
| 2018 | Improving Object Localization With Fitness NMS and Bounded IoU LossabstractWe demonstrate that many detection methods are designed to identify only a sufficiently accurate bounding box, rather than the best available one. To address this issue we propose a simple and fast modification to the existing methods called Fitness NMS. This method is tested with the DeNet model and obtains a significantly improved MAP at greater localization accuracies without a loss in evaluation rate, and can be used in conjunction with Soft NMS for additional improvements. Next we derive a novel bounding box regression loss based on a set of IoU upper bounds that better matches the goal of IoU maximization while still providing good convergence properties. Following these novelties we investigate RoI clustering schemes for improving evaluation rates for the DeNet wide model variants and provide an analysis of localization performance at various input image dimensions. We obtain a MAP of 33.6%@79Hz and 41.8%@5Hz for MSCOCO and a Titan X (Maxwell). Source code available from: https://github.com/lachlants/denet. Lachlan Tychsen-Smith, Lars Petersson |
CVPR | 2 |
| 2018 | Effective Use of Synthetic Data for Urban Scene Semantic Segmentation
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004 |
ECCV (2) | 4 |
| 2018 | Incorporating Network Built-in Priors in Weakly-Supervised Semantic SegmentationabstractPixel-level annotations are expensive and time consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recently, CNN-based methods have proposed to fine-tune pre-trained networks using image tags. Without additional information, this leads to poor localization accuracy. This problem, however, was alleviated by making use of objectness priors to generate foreground/background masks. Unfortunately these priors either require pixel-level annotations/bounding boxes, or still yield inaccurate object boundaries. Here, we propose a novel method to extract accurate masks from networks pre-trained for the task of object recognition, thus forgoing external objectness modules. We first show how foreground/background masks can be obtained from the activations of higher-level convolutional layers of a network. We then show how to obtain multi-class masks by the fusion of foreground/background ones with information extracted from a weakly-supervised localization network. Our experiments evidence that exploiting these masks in conjunction with a weakly-supervised training loss yields state-of-the-art tag-based weakly-supervised semantic segmentation results. Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004, Stephen Gould |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Encouraging LSTMs to Anticipate Actions Very EarlyabstractIn contrast to the widely studied problem of recognizing an action given a complete sequence, action anticipation aims to identify the action from only partially available videos. As such, it is therefore key to the success of computer vision applications requiring to react as early as possible, such as autonomous navigation. In this paper, we propose a new action anticipation method that achieves high prediction accuracy even in the presence of a very small percentage of a video sequence. To this end, we develop a multi-stage LSTM architecture that leverages context-aware and action-aware features, and introduce a novel loss function that encourages the model to predict the correct class as early as possible. Our experiments on standard benchmark datasets evidence the benefits of our approach; We outperform the state-of-the-art action anticipation methods for early prediction by a relative increase in accuracy of 22.0% on JHMDB-21, 14.0% on UT-Interaction and 49.9% on UCF-101. Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Basura Fernando, Lars Petersson, Lars Andersson |
ICCV | 5 |
| 2017 | Globally-Optimal Inlier Set Maximisation for Simultaneous Camera Pose and Feature CorrespondenceabstractEstimating the 6-DoF pose of a camera from a single image relative to a pre-computed 3D point-set is an important task for many computer vision applications. Perspective-n-Point (PnP) solvers are routinely used for camera pose estimation, provided that a good quality set of 2D-3D feature correspondences are known beforehand. However, finding optimal correspondences between 2D key-points and a 3D point-set is non-trivial, especially when only geometric (position) information is known. Existing approaches to the simultaneous pose and correspondence problem use local optimisation, and are therefore unlikely to find the optimal solution without a good pose initialisation, or introduce restrictive assumptions. Since a large proportion of outliers are common for this problem, we instead propose a globally-optimal inlier set cardinality maximisation approach which jointly estimates optimal camera pose and optimal correspondences. Our approach employs branch-and-bound to search the 6D space of camera poses, guaranteeing global optimality without requiring a pose prior. The geometry of SE(3) is used to find novel upper and lower bounds for the number of inliers and local optimisation is integrated to accelerate convergence. The evaluation empirically supports the optimality proof and shows that the method performs much more robustly than existing approaches, including on a large-scale outdoor data-set. Dylan Campbell, Lars Petersson, Laurent Kneip, Hongdong Li |
ICCV | 2 |
| 2017 | Bringing Background into the Foreground: Making All Classes Equal in Weakly-Supervised Video Semantic SegmentationabstractPixel-level annotations are expensive and time-consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recent years have seen great progress in weakly-supervised semantic segmentation, whether from a single image or from videos. However, most existing methods are designed to handle a single background class. In practical applications, such as autonomous navigation, it is often crucial to reason about multiple background classes. In this paper, we introduce an approach to doing so by making use of classifier heatmaps. We then develop a two-stream deep architecture that jointly leverages appearance and motion, and design a loss based on our heatmaps to train it. Our experiments demonstrate the benefits of our classifier heatmaps and of our two-stream architecture on challenging urban scene datasets and on the YouTube-Objects benchmark, where we obtain state-of-the-art results. Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004 |
ICCV | 4 |
| 2017 | DeNet: Scalable Real-Time Object Detection with Directed Sparse SamplingabstractWe define the object detection from imagery problem as estimating a very large but extremely sparse bounding box dependent probability distribution. Subsequently we identify a sparse distribution estimation scheme, Directed Sparse Sampling, and employ it in a single end-to-end CNN based detection model. This methodology extends and formalizes previous state-of-the-art detection models with an additional emphasis on high evaluation rates and reduced manual engineering. We introduce two novelties, a corner based region-of-interest estimator and a deconvolution based CNN model. The resulting model is scene adaptive, does not require manually defined reference bounding boxes and produces highly competitive results on MSCOCO, Pascal VOC 2007 and Pascal VOC 2012 with real-time evaluation rates. Further analysis suggests our model performs particularly well when finegrained object localization is desirable. We argue that this advantage stems from the significantly larger set of available regions-of-interest relative to other methods. Source-code is available from: https://github.com/lachlants/denet. Lachlan Tychsen-Smith, Lars Petersson |
ICCV | 2 |
| 2016 | GOGMA: Globally-Optimal Gaussian Mixture AlignmentabstractGaussian mixture alignment is a family of approaches that are frequently used for robustly solving the point-set registration problem. However, since they use local optimisation, they are susceptible to local minima and can only guarantee local optimality. Consequently, their accuracy is strongly dependent on the quality of the initialisation. This paper presents the first globally-optimal solution to the 3D rigid Gaussian mixture alignment problem under the L2 distance between mixtures. The algorithm, named GOGMA, employs a branch-and-bound approach to search the space of 3D rigid motions SE(3), guaranteeing global optimality regardless of the initialisation. The geometry of SE(3) was used to find novel upper and lower bounds for the objective function and local optimisation was integrated into the scheme to accelerate convergence without voiding the optimality guarantee. The evaluation empirically supported the optimality proof and showed that the method performed much more robustly on two challenging datasets than an existing globally-optimal registration solution. Dylan Campbell, Lars Petersson |
CVPR | 2 |
| 2016 | Sample and Filter: Nonparametric Scene Parsing via Efficient FilteringabstractScene parsing has attracted a lot of attention in computer vision. While parametric models have proven effective for this task, they cannot easily incorporate new training data. By contrast, nonparametric approaches, which bypass any learning phase and directly transfer the labels from the training data to the query images, can readily exploit new labeled samples as they become available. Unfortunately, because of the computational cost of their label transfer procedures, state-of-the-art nonparametric methods typically filter out most training images to only keep a few relevant ones to label the query. As such, these methods throw away many images that still contain valuable information and generally obtain an unbalanced set of labeled samples. In this paper, we introduce a nonparametric approach to scene parsing that follows a sample-andfilter strategy. More specifically, we propose to sample labeled superpixels according to an image similarity score, which allows us to obtain a balanced set of samples. We then formulate label transfer as an efficient filtering procedure, which lets us exploit more labeled samples than existing techniques. Our experiments evidence the benefits of our approach over state-of-the-art nonparametric methods on two benchmark datasets. Mohammad Najafi, Sarah Taghavi Namin, Mathieu Salzmann, Lars Petersson |
CVPR | 4 |
| 2016 | Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, Stephen Gould, José M. Álvarez 0004 |
ECCV (8) | 4 |
| 2016 | Latent structural SVM with marginal probabilities for weakly labeled structured learningabstractIn the last years, the increasing availability of annotated data has facilitated the great success of supervised learning in real-world applications such as semantic labeling. However, the vast majority of data is nowadays unlabeled or partially annotated. In this paper, we develop an Expected Marginal Latent Structural SVM (EM-LSSVM) framework for performing structured learning in the presence of weakly (partially) annotated data by incorporating the uncertainty of the unobserved data as marginals. Experimental results on semantic labeling show the potential of the proposed method. In particular, we learn the parameters of a CRF where large amounts of noisy and unobserved data are available. Comparison against state of the art demonstrates the applicability of our algorithm to practical applications. Shahin Namin, José M. Álvarez 0004, Laurent Kneip, Lars Petersson |
ICIP | 4 |
| 2016 | 2D-3D semantic segmentation using cardinality as higher-order lossabstractMulti-modal scene analysis is a growing field of importance as additional sensors, such as 3D LIDAR, is becoming a common complement to image capturing systems. However, while additional sensory data potentially can make the analysis more accurate, it also comes with a host of associated issues. For example, inconsistencies in the data between sensors resulting from, e.g., misalignment, moving objects, or parallax effects, can severely affect the performance. Additionally, real-world scenes tend to have an inherent imbalance in the number of items of each class which typically suppresses the performance of infrequent classes. In this paper, we address those two issues specifically by a) using a cardinality loss function designed to target inconsistencies at training time, and b) devising an average per class loss function addressing the imbalance issue. Shahin Namin, José M. Álvarez 0004, Lars Petersson |
ICPR | 3 |
| 2015 | An Adaptive Data Representation for Robust Point-Set Registration and MergingabstractThis paper presents a framework for rigid point-set registration and merging using a robust continuous data representation. Our point-set representation is constructed by training a one-class support vector machine with a Gaussian radial basis function kernel and subsequently approximating the output function with a Gaussian mixture model. We leverage the representation's sparse parametrisation and robustness to noise, outliers and occlusions in an efficient registration algorithm that minimises the L2 distance between our support vector -- parametrised Gaussian mixtures. In contrast, existing techniques, such as Iterative Closest Point and Gaussian mixture approaches, manifest a narrower region of convergence and are less robust to occlusions and missing data, as demonstrated in the evaluation on a range of 2D and 3D datasets. Finally, we present a novel algorithm, GMMerge, that parsimoniously and equitably merges aligned mixture models, allowing the framework to be used for reconstruction and mapping. Dylan Campbell, Lars Petersson |
ICCV | 2 |
| 2015 | Cutting Edge: Soft Correspondences in Multimodal Scene ParsingabstractExploiting multiple modalities for semantic scene parsing has been shown to improve accuracy over the single modality scenario. Existing methods, however, assume that corresponding regions in two modalities have the same label. In this paper, we address the problem of data misalignment and label inconsistencies, e.g., due to moving objects, in semantic labeling, which violate the assumption of existing techniques. To this end, we formulate multimodal semantic labeling as inference in a CRF, and introduce latent nodes to explicitly model inconsistencies between two domains. These latent nodes allow us not only to leverage information from both domains to improve their labeling, but also to cut the edges between inconsistent regions. To eliminate the need for hand tuning the parameters of our model, we propose to learn intra-domain and inter-domain potential functions from training data. We demonstrate the benefits of our approach on two publicly available datasets containing 2D imagery and 3D point clouds. Thanks to our latent nodes and our learning strategy, our method outperforms the state-of-the-art in both cases. Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann, Lars Petersson |
ICCV | 4 |
| 2015 | Modeling the cost and coverage of an ad-hoc asset management system based on existing fleet vehiclesabstractMonitoring road assets such as road signs, utility poles, features of the road itself or other structures close to where vehicles are driving is important. Such assets need to be monitored in order to maintain them and minimize accident fatalities caused by non-compliance [1]. However, traditional surveying methods that utilize dedicated vehicles equipped with high-end expensive sensors turn out to be very costly and hence, surveys can only be carried out every few years. This paper explores the feasibility of equipping existing fleet vehicles, such as taxis, with low-end, low-quality sensors that traverse the road network through their normal daily activities. The cost and coverage of such a new approach is modeled with the help of a dataset T-Drive from Microsoft that provides taxi trajectories for more than 10,000 taxis in Beijing. The paper further estimates the optimal, from a cost perspective, number of taxis needed to survey the region by considering the cost of explicitly surveying areas that have not been covered by the random trajectories of the taxis. Dana Pordel, Lars Petersson, Shahin Namin, Adrian Rebola-Pardo |
Intelligent Vehicles Symposium | 2 |
| 2015 | A Multi-modal Graphical Model for Scene AnalysisabstractIn this paper, we introduce a multi-modal graphical model to address the problems of semantic segmentation using 2D-3D data exhibiting extensive many-to-one correspondences. Existing methods often impose a hard correspondence between the 2D and 3D data, where the 2D and 3D corresponding regions are forced to receive identical labels. This results in performance degradation due to misalignments, 3D-2D projection errors and occlusions. We address this issue by defining a graph over the entire set of data that models soft correspondences between the two modalities. This graph encourages each region in a modality to leverage the information from its corresponding regions in the other modality to better estimate its class label. We evaluate our method on a publicly available dataset and beat the state-of-the-art. Additionally, to demonstrate the ability of our model to support multiple correspondences for objects in 3D and 2D domains, we introduce a new multi-modal dataset, which is composed of panoramic images and LIDAR data, and features a rich set of many-to-one correspondences. Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann, Lars Petersson |
WACV | 4 |
| 2014 | Non-associative Higher-Order Markov Networks for Point Cloud Classification
Mohammad Najafi, Sarah Taghavi Namin, Mathieu Salzmann, Lars Petersson |
ECCV (5) | 4 |
| 2014 | Multi-view terrain classification using panoramic imagery and LIDARabstractThe focus of this work is addressing the challenges of performing object recognition in real world scenes as captured by a commercial, state-of-the-art, surveying vehicle equipped with a 360° panoramic camera in conjunction with a 3D laser scanner (LIDAR). Even with state-of-the-art surveying equipment, there is colour saturation and very dark regions in images, as well as some degree of time-varying misalignment between the point cloud data and imagery due to, for instance, imperfect tracking of sensor pose. Moreover, there are frequent occlusions due to both static and moving objects. These issues are inherently difficult to avoid and therefore need to be dealt with in a more robust fashion. This is where the contribution of the paper is; that is, the development of a consensus method that can intelligently incorporate feature responses from multiple views and reject those that are not very descriptive. It is shown that the overall performance in a ten class problem is increased from 70.5% for a simple 2D-3D classification system, to 77.5%. Subsequently, an enhanced CRF which has become robust using the misclassifications of training data and equipped with the probabilities of the adjacent points, was applied to the system and further improved its performance to 82.9%. The experiments were performed on a challenging dataset captured both in summer and winter. Sarah Taghavi Namin, Mohammad Najafi, Lars Petersson |
IROS | 3 |
| 2014 | Creating robust high-throughput traffic sign detectors using centre-surround HOG statistics
Gary Overett, Lachlan Tychsen-Smith, Lars Petersson, Niklas Pettersson, Lars Andersson |
Mach. Vis. Appl. | 3 |
| 2013 | Classification of natural scene multi spectral images using a new enhanced CRFabstractIn this paper, a new enhanced CRF for discriminating between different materials in natural scenes using terrestrial multi spectral imaging is established. Most of the existing formulations of the CRF often suffer from over smoothing and loss of small detail, thereby deteriorating the information from the underlying unary classifier in areas with a high spatial frequency. This work specifically addresses this issue by incorporating a new pairwise potential that is better at taking local context into account. Certain materials are very unlikely to appear next to each other in the scene and such configurations are penalised by employing the confusion matrix of the unary classifier. Similarly, horizontal as well as vertical configurations, which may be more or less likely for certain combinations of materials, are regarded in this formulation. Furthermore, the proposed pairwise potential also considers the length of boundaries between regions to account for the segmentation granularity issues and also uses class probabilities of the neighbouring regions to make up for the uncertainty of the unary classifier results. Seven band terrestrial multi spectral imaging were used due to its potential in distinguishing between different materials and objects. The proposed approach was evaluated using cross-validation, resulting in an average accuracy of 88.9% which is about 17% more than the accuracy of a standard CRF, which demonstrates the superiority of our approach in preserving local details. Mohammad Najafi, Sarah Taghavi Namin, Lars Petersson |
IROS | 3 |
| 2012 | A computationally efficient low-bandwidth method for very-large-scale mapping of road signs with multiple vehicles
Ashkan Amirsadri, Adrian N. Bishop, Jonghyuk Kim, Jochen Trumpf, Lars Petersson |
FUSION | 5 |
| 2012 | Classification of materials in natural scenes using multi-spectral imagesabstractIn this paper, a method suitable for distinguishing between different materials occurring in natural scenes using a multi-spectral camera is devised. Such a capability is useful in autonomous robot applications to help negotiating the environment as well as, e.g. applications intended to create large scale inventories of assets in the proximity of roads. The utilised sensor records a seven band multi-spectral image, of which six bands are in the visible range and one in the NIR (near-infrared) range. Many materials appearing similar if viewed by a common RGB camera, will show discriminating properties if viewed by a camera capturing a greater number of separated wavelengths. The approach in this paper is to combine the discriminating strength of the multi-spectral signature in each pixel and the corresponding nature of the surrounding texture. Local features, considering seven bands in each pixel and texture features such as GLCM and Fourier spectrum features are exploited to make the system more robust to different lighting conditions. Then classifiers built using SVM and AdaBoost are evaluated with very promising results, an average classification accuracy of 91.9% and 89.1%, respectively for a ten class problem. Sarah Taghavi Namin, Lars Petersson |
IROS | 2 |
| 2011 | Large scale sign detection using HOG feature variantsabstractIn this paper we present two variant formulations of the well-known Histogram of Oriented Gradients (HOG) features and provide a comparison of these features on a large scale sign detection problem. The aim of this research is to find features capable of driving further improvements atop a preexisting detection framework used commercially to detect traffic signs on the scale of entire national road networks (1000's of kilometres of video). We assume the computationally efficient framework of a cascade of boosted weak classifiers. Rather than comparing features on the general problem of detection we compare their merits in the final stages of a cascaded detection problem where a feature's ability to reduce error is valued more highly than computational efficiency. Results show the benefit of the two new features on a New Zealand speed sign detection problem. We also note the importance of using non-sign training and validation instances taken from the same video data that contains the training and validation positives. This is attributed to the potential for the more powerful HOG features to overfit on specific local patterns which may be present in alternative video data. Gary Overett, Lars Petersson |
Intelligent Vehicles Symposium | 2 |
| 2008 | Statistical Threat Assessment for General Road Scenes Using Monte Carlo SamplingabstractThis paper presents a threat-assessment algorithm for general road scenes. A road scene consists of a number of objects that are known, and the threat level of the scene is based on their current positions and velocities. The future driver inputs of the surrounding objects are unknown and are modeled as random variables. In order to capture realistic driver behavior, a dynamic driver model is implemented as a probabilistic prior, which computes the likelihood of a potential maneuver. A distribution of possible future scenarios can then be approximated using a Monte Carlo sampling. Based on this distribution, different threat measures can be computed, e.g., probability of collision or time to collision. Since the algorithm is based on the Monte Carlo sampling, it is computationally demanding, and several techniques are presented to increase performance without increasing computational load. The algorithm is intended both for online safety applications in a vehicle and for offline data analysis. Andreas Eidehall, Lars Petersson |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2007 | Improved Response Modelling on Weak Classifiers for BoostingabstractThis paper demonstrates a method of increasing the quality of weak classifiers in the boosting context by using improved response modelling. The new method improves upon the results of a recent response binning approach proposed by Rasolzadeh et al. (2006). For experimental purposes the improved method is applied to the familiar Haar features as used by Viola and Jones in their face/pedestrian detection systems. However, the methods benefits are general and therefore not restricted to this particular feature type. Unlike many previous methods, this method is suitable for modelling multi-modal responses and is highly resistant to overfitting. It does this by adaptively choosing suitable support regions around the values taken by the standard response binning method. More accurate models are produced, with particular improvement around the final decision boundary. It is shown that the new method can be trained with one tenth of the training data required to achieve similar results on previous methods. This substantially lowers the overall training time of the system. The method's ability to consistently produce better hypotheses over a variety of pedestrian detection tasks is shown. Gary Overett, Lars Petersson |
ICRA | 2 |
| 2005 | A Sign Reading Driver Assistance System Using Eye GazeabstractCars are becoming, in effect, a robotic system with an embedded human. It is not possible to know what the driver is thinking. We can, however, monitor their gaze and compare it with information in their view-field to make an inference. In this paper we present a complete system that reads speed signs in real-time, compares the driver gaze, and provides immediate feedback if the sign has been missed by the driver. This paper focuses on correlating measures of driver gaze direction with the position of signs in the road scene and improving recognition of signs through image enhancement. Luke Fletcher, Lars Petersson, Nick Barnes, David J. Austin, Alexander Zelinsky |
ICRA | 2 |
| 2005 | Generic fusion of visual cues applied to real-world object segmentationabstractFusion of information from different complementary sources may be necessary to achieve a robust sensing system that degrades gracefully under various conditions. Many approaches use a specific tailor-made combination of algorithms that do not easily allow the inclusion of more, or other, types of algorithms. In this paper, we explore a variant of a generic algorithm for fusing visual cues to the task of object segmentation in a video stream. The fusion algorithm combines the output of several segmentation algorithms in a straight forward way by using a Bayesian approach and a particle filter to track several hypotheses. Segmentation algorithms can be added or removed without changing the over all structure of the system. It was of particular interest to investigate if the method was suitable when realistic real-world scenes with much noise was analysed. The system has been tested on image sequences taken from a moving vehicle where stationary and moving objects are successfully segmented from the background. In conclusion, the fusion algorithm explored is well suited to this problem domain and is easily adopted. The context of this work is on-line pedestrian detection to be deployed in cars. Fredrik Arnell, Lars Petersson |
IROS | 2 |
| 2004 | An Interactive Driver Assistance System Monitoring the Scene in and out of the VehicleabstractThis paper presents a framework for interactive driver assistance systems including techniques for fast speed sign detection and classification, car detection and tracking, and lane departure warning. In addition, the driver's actions are monitored. The integrated system uses information extracted from the road scene (speed signs, position within the lane, relative position to other cars, etc.) together with information about the driver's state such as eye gaze and head pose, to issue adequate warnings. A touch screen monitor presents relevant information and allows the driver to interact with the system. The research is focused around robust on-line algorithms. Initial results of online speed sign detection and car tracking are presented in the context of a driver assistance system. Lars Petersson, Luke Fletcher, Nick Barnes, Alexander Zelinsky |
ICRA | 1 |
| 2003 | Driver assistance: an integration of vehicle monitoring and controlabstractAbout 1.17 million people die in road crashes around the world each year. It is estimated that up to 30% of these fatalities are caused by fatigue and inattention. There are systems able to detect what is happening outside of the car, e.g., lane tracking, obstacle detection, pedestrian detection etc. Further on, there are also means for monitoring the actions of the driver. A natural step is to fuse the available data from within and outside of the car, and suggest a suitable response. This paper discusses driver assistance systems, lists a set of necessary core competencies of such a system and in particular presents a system for force-feedback in the steering wheel when crossing lanes. The presented system utilises a robust lane tracker which is experimentally evaluated for the purpose of driver assistance. In addition, preliminary results from simultaneous driver monitoring and lane tracking are presented that indicates a good correlation between the two, i.e. the driver's gaze direction and the structure of the road. These data can in turn be used for more advanced driver assistance systems in the future. Lars Petersson, Nicholas Apostoloff, Alexander Zelinsky |
ICRA | 1 |
| 2002 | Systems Integration for Real-World Manipulation TasksabstractA system developed to demonstrate integration of a number of key research areas such as localization, recognition, visual tracking, visual servoing and grasping is presented together with the underlying methodology adopted to facilitate the integration. Through sequencing of basic skills, provided by the above mentioned competencies, the system has the potential to carry out flexible grasping for fetch and carry in realistic environments. Through careful fusion of reactive and deliberative control and use of multiple sensory modalities a significant flexibility is achieved. Experimental verification of the integrated system is presented. Lars Petersson, Patric Jensfelt, Dennis Tell, M. Strandberg, Danica Kragic, Henrik I. Christensen |
ICRA | 1 |
| 2001 | DCA: a distributed control architecture for roboticsabstractMany control applications are by nature distributed, not only over different processes but also over several processors. Managing such a system with respect to the startup of processes, internal communications and state changes quickly becomes a very complex task. The paper presents a distributed control architecture which supports a formal model of computation as described by Lyons and Arib (1989). The architecture is primarily intended for robot control but has a wide range of potential applications. We motivate the design and implementation of the architecture by discussing the desired properties of a robot system capable of doing real-time tasks like manipulation. This leads to functionality such as a process algebra controlling the life-cycle of the processes, grouping and distribution of processes and internal communication transparent to location. Our implementation does not in itself introduce any bottlenecks due to a tree structure with local control over processes which gives an efficient and scalable architecture. At the end, an example scenario in which a fairly advanced problem like opening a door using a mobile robot with a manipulator arm is demonstrated in the presented framework. Lars Petersson, David J. Austin, Henrik Christenseni |
IROS | 1 |
| 2000 | High-level control of a mobile manipulator for door openingabstractIn this paper, off-the-shelf algorithms for force/torque control are used in the context of mobile manipulation, in particular, the task of opening a door is studied. To make the solution robust, as few assumptions as possible are made. By using relaxation of forces as the basic level of control more complex information can be derived from the resulting motion. In our system, the radius and centre of rotation of the door are estimated online. This enables the complete system to have a higher degree of autonomy in an unknown environment. In addition, the redundancy of the robot is exploited in such a way to drive the system towards a desired configuration. The framework of hybrid dynamic systems is used to implement the algorithm which gives a theoretically sound framework for analysing the system with respect to safety and functionality. The integration of the above approaches results in a system which can robustly locate and grasp the handle and then open the door. Lars Petersson, David J. Austin, Danica Kragic |
IROS | 1 |
| 1999 | A hybrid control architecture for mobile manipulationabstractWe present a scheme for mobile manipulation by introducing a mobile manipulation control architecture (MMCA). This architecture is motivated by a need for a systematic control structure for robotic manipulation within a behavior based framework. The control structure enables integration of the manipulator into a behavior based control structure for the platform. Furthermore, our suggested MMCA is designed in such a way that it supports design and performance analysis from both a manipulator dynamics and a hybrid automata perspective. Lars Petersson, Magnus Egerstedt, Henrik I. Christensen |
IROS | 1 |