Lars Petersson

dblp:18/5485 · DBLP profile ↗
← Back
83ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0002-0103-1904ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 25 since 2021Systems, architecture and hardware · 12 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Subspace-Guided Knowledge Distillation for Efficient Model Transfer
abstract
Compact models can be effectively trained via Knowledge Distillation (KD), where a lightweight student model learns to replicate the behavior of a larger, high-performing teacher. A persistent challenge in KD lies in the misalignment between the representational spaces of teacher and student networks, especially when they differ in architecture or capacity. To address this, we propose Subspace-Driven Knowledge Distillation (SDMD), a novel framework that mitigates representational disparity by projecting features into an indefinite inner product space. This relaxation from traditional Hilbert spaces enables more flexible geometric alignment, capturing transformations such as rotations and reflections that are often necessary for accurate knowledge transfer. By learning a subspace that bridges the semantic gap between teacher and student, SDMD facilitates more effective distillation without increasing model complexity. We validate SDMD through extensive experiments on large-scale image classification (ImageNet-1K) and object detection (COCO), where it consistently outperforms existing distillation methods. Notably, SDMD-trained models not only achieve state-of-the-art results in distilled settings but also surpass the performance of equivalent models trained from scratch, highlighting the strength of our subspace-based alignment strategy.
Zeeshan Hayder, Ali Cheraghian, Lars Petersson, Mehrtash Harandi
WACV3
2026 PointCaM: Cut-and-Mix for open-set point cloud learning
Shi Qiu 0001, Weihao Li 0005, Saeed Anwar, Mehrtash Harandi, Nick Barnes, Lars Petersson
Comput. Vis. Image Underst.7
2026 MFGS: Mask-free Gaussian separation for 3D object reconstruction
abstract
Accurate 3D reconstruction from multi-view images is a fundamental problem in computer vision. A common acquisition strategy involves placing an object on a rotating turntable while moving the camera to capture it from various viewpoints. In such scenarios, object moves relative to the background, many existing reconstruction methods rely on object masks to separate the foreground from the background. The quality of these masks significantly affects the final reconstruction, yet obtaining high-quality and consistent masks is a challenging and laborious process, especially when controlled environments like green screens are unavailable. To address this limitation, we introduce Mask-free Gaussian Separation (MFGS), a novel method that performs simultaneous object reconstruction and segmentation without requiring any input masks. Our approach builds on Gaussian Splatting and automatically disentangles the scene by extending each Gaussian primitive with a learnable parameter that represents its probability of belonging to the dynamic foreground object. This separation is optimized in a self-supervised manner, optimized by the object and camera transformation constraints. We evaluated MFGS on new synthetic and real-world datasets designed to reflect this challenging capture scenario. Experimental results demonstrate that our mask-free approach significantly outperforms existing methods. Notably, MFGS surpasses the performance of the state-of-the-art method(2DGS) that relies on high-quality segmentation masks, achieving a 27% improvement in novel view synthesis and a 7% improvement in geometry reconstruction.
Jinguang Tong, Xuesong Li 0001, Sundaram Muthu, Fahira A. Maken, Lars Petersson, Hongdong Li
Pattern Recognit.5
2025 GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction
abstract
3D modeling of highly reflective objects remains challenging due to strong view-dependent appearances. While previous SDF-based methods can recover high-quality meshes, they are often time-consuming and tend to produce over-smoothed surfaces. In contrast, 3D Gaussian Splatting (3DGS) offers the advantage of high speed and detailed real-time rendering, but extracting surfaces from the Gaussians can be noisy due to the lack of geometric constraints. To bridge the gap between these approaches, we propose a novel reconstruction method called GS-2DGS for reflective objects based on 2D Gaussian Splatting (2DGS). Our approach combines the rapid rendering capabilities of Gaussian Splatting with additional geometric information from foundation models. Experimental results on synthetic and real datasets demonstrate that our method significantly outperforms Gaussian-based techniques in terms of reconstruction and relighting and achieves performance comparable to SDF-based methods while being an order of magnitude faster. Code is available at https://github.com/hirotong/GS2DGS
Jinguang Tong, Xuesong Li 0001, Fahira A. Maken, Sundaram Muthu, Lars Petersson, Hongdong Li
CVPR5
2025 Open Set Label Shift with Test Time Out-of-Distribution Reference
abstract
Open set label shift (OSLS) occurs when label distributions change from a source to a target distribution, and the target distribution has an additional out-of-distribution (OOD) class. In this work, we build estimators for both source and target open set label distributions using a source domain in-distribution (ID) classifier and an ID/OOD classifier. With reasonable assumptions on the ID/OOD classifier, the estimators are assembled into a sequence of three stages: 1) an estimate of the source label distribution of the OOD class, 2) an EM algorithm for Maximum Likelihood estimates (MLE) of the target label distribution, and 3) an estimate of the target label distribution of OOD class under relaxed assumptions on the OOD classifier. The sampling errors of estimates in 1) and 3) are quantified with a concentration inequality. The estimation result allows us to correct the ID classifier trained on the source distribution to the target distribution without retraining. Experiments on a variety of open set label shift settings demonstrate the effectiveness of our model. Our code is available at https://github.com/ChangkunYe/OpenSetLabelShift.
Changkun Ye, Russell Tsuchida, Lars Petersson, Nick Barnes
CVPR3
2025 DGNS: Deformable Gaussian Splatting and Dynamic Neural Surface for Monocular Dynamic 3D Reconstruction
abstract
Dynamic scene reconstruction from monocular video is essential for real-world applications. We introduce DGNS, a hybrid framework integrating Deformable Gaussian Splatting and Dynamic Neural Surfaces, effectively addressing dynamic novel-view synthesis and 3D geometry reconstruction simultaneously. During training, depth maps generated by the deformable Gaussian splatting module guide the ray sampling for faster processing and provide depth supervision within the dynamic neural surface module to improve geometry reconstruction. Conversely, the dynamic neural surface directs the distribution of Gaussian primitives around the surface, enhancing rendering quality. In addition, we propose a depth-filtering approach to further refine depth supervision. Extensive experiments conducted on public datasets demonstrate that DGNS achieves state-of-the-art performance in 3D reconstruction, along with competitive results in novel-view synthesis.
Xuesong Li 0001, Jinguang Tong, Vivien Rolland, Lars Petersson
ACM Multimedia5
2025 BioNet and NeFF: Crop Biomass Prediction from Point Clouds to Drone Imagery
abstract
Crop biomass offers crucial insights into plant health and yield, making it essential for crop science, farming systems, and agricultural research. However, current measurement methods, which are labor-intensive, destructive, and imprecise, hinder large-scale quantification of this trait. To address this limitation, we present a biomass prediction network (BioNet), designed for adaptation across different data modalities, including point clouds and drone imagery. Our BioNet, utilizing a sparse 3D convolutional neural network (CNN) and a transformer-based prediction module, processes point clouds and other 3D data representations to predict biomass. To further extend BioNet for drone imagery, we integrate a neural feature field (NeFF) module, enabling 3D structure reconstruction and the transformation of 2D semantic features from vision foundation models into the corresponding 3D surfaces. For the point cloud modality, BioNet demonstrates superior performance on two public datasets, with an approximate 6.1% relative improvement (RI) over the state-of-the-art. In the RGB image modality, the combination of BioNet and NeFF achieves a 7.9% RI. Additionally, the NeFF-based approach utilizes inexpensive, portable drone-mounted cameras, providing a scalable solution for large field applications.
Xuesong Li 0001, Zeeshan Hayder, Ali Zia, Connor Cassidy, Shiming Liu, Warwick Stiller, Eric A. Stone, Warren Conaty, Lars Petersson, Vivien Rolland
WACV9
2025 Facial Expression Recognition with Controlled Privacy Preservation and Feature Compensation
abstract
Facial expression recognition (FER) systems raise significant privacy concerns due to the potential exposure of sensitive identity information. This paper presents a study on removing identity information while preserving FER capabilities. Drawing on the observation that lowfrequency components predominantly contain identity information and high-frequency components capture expression, we propose a novel two-stream framework that applies privacy enhancement to each component separately. We introduce a controlled privacy enhancement mechanism to optimize performance and a feature compensator to enhance task-relevant features without compromising privacy. Furthermore, we propose a novel privacy-utility trade-off, providing a quantifiable measure of privacy preservation efficacy in closed-set FER tasks. Extensive experiments on the benchmark CREMA-D dataset demonstrate that our framework achieves 78.84 % recognition accuracy with a privacy (facial identity) leakage ratio of only 2.01 %, highlighting its potential for secure and reliable video-based FER applications. We encourage the readers to visit the project page: https://fengxxu.github.io/ppfer/.
David Ahmedt-Aristizabal, Lars Petersson, Dadong Wang, Xun Li 0004
WACV3
2025 Attention-Based Real Image Restoration
abstract
Deep convolutional neural networks perform better on images containing spatially invariant degradations, also known as synthetic degradations; however, their performance is limited on real-degraded photographs and requires multiple-stage network modeling. To advance the practicability of restoration algorithms, this article proposes a novel single-stage blind real image restoration network ( Net) by employing a modular architecture. We use a residual on the residual structure to ease low-frequency information flow and apply feature attention to exploit the channel dependencies. Furthermore, the evaluation in terms of quantitative metrics and visual quality for four restoration tasks, i.e., denoising, super-resolution, raindrop removal, and JPEG compression on 11 real degraded datasets against more than 30 state-of-the-art algorithms, demonstrates the superiority of our Net. We also present the comparison on three synthetically generated degraded datasets for denoising to showcase our method's capability on synthetics denoising. The codes, trained models, and results are available on https://github.com/saeed-anwar/R2Net.
Saeed Anwar, Nick Barnes, Lars Petersson
IEEE Trans. Neural Networks Learn. Syst.3
2024 Backpropagation-free Network for 3D Test-time Adaptation
abstract
Real-world systems often encounter new data over time, which leads to experiencing target domain shifts. Existing Test- Time Adaptation (TTA) methods tend to apply computationally heavy and memory-intensive backpropagation-based approaches to handle this. Here, we propose a novel method that uses a backpropagation-free approach for TTA for the specific case of 3D data. Our model uses a two-stream architecture to maintain knowledge about the source domain as well as complementary target-domain-specific information. The backpropagation-free property of our model helps address the well-known forgetting prob-lem and mitigates the error accumulation issue. The pro-posed method also eliminates the need for the usually noisy process of pseudo-labeling and reliance on costly self-supervised training. Moreover, our method leverages sub-space learning, effectively reducing the distribution vari-ance between the two domains. Furthermore, the source-domain-specific and the target-domain-specific streams are aligned using a novel entropy-based adaptive fusion strat-egy. Extensive experiments on popular benchmarks demon-strate the effectiveness of our method. The code will be available at https://github.com/abie-e/BFTT3D.
Yanshuo Wang, Ali Cheraghian, Zeeshan Hayder, Sameera Ramasinghe, Shafin Rahman, David Ahmedt-Aristizabal, Xuesong Li 0001, Lars Petersson, Mehrtash Harandi
CVPR9
2024 Canonical Shape Projection Is All You Need for 3D Few-Shot Class Incremental Learning
Ali Cheraghian, Zeeshan Hayder, Sameera Ramasinghe, Shafin Rahman, Javad Jafaryahya, Lars Petersson, Mehrtash Harandi
ECCV (41)6
2024 Continual Test-time Domain Adaptation via Dynamic Sample Selection
abstract
The objective of Continual Test-time Domain Adaptation (CTDA) is to gradually adapt a pre-trained model to a sequence of target domains without accessing the source data. This paper proposes a Dynamic Sample Selection (DSS) method for CTDA. DSS consists of dynamic thresholding, positive learning, and negative learning processes. Traditionally, models learn from unlabeled unknown environment data and equally rely on all samples’ pseudo-labels to update their parameters through self-training. However, noisy predictions exist in these pseudo-labels, so all samples are not equally trustworthy. Therefore, in our method, a dynamic thresholding module is first designed to select suspected low-quality from high-quality samples. The selected low-quality samples are more likely to be wrongly predicted. Therefore, we apply joint positive and negative learning on both high- and low-quality samples to reduce the risk of using wrong information. We conduct extensive experiments that demonstrate the effectiveness of our proposed method for CTDA in the image domain, outperforming the state-of-the-art results. Furthermore, our approach is also evaluated in the 3D point cloud domain, showcasing its versatility and potential for broader applicability.
Yanshuo Wang, Ali Cheraghian, Shafin Rahman, David Ahmedt-Aristizabal, Lars Petersson, Mehrtash Harandi
WACV6
2024 Label Shift Estimation for Class-Imbalance Problem: A Bayesian Approach
abstract
As a type of distribution shift, label shift occurs when the source and target domains have different label distributions $\mathbb{P}(Y)$ but identical conditional distributions of data given labels $\mathbb{P}(X|Y)$. Under a Bayesian framework, we propose a novel Maximum A Posteriori (MAP) model and a novel posterior sampling model for the label shift problem. We prove the MAP objective admits a unique optimum and derive an EM algorithm that converges to the global optimum. We propose a novel Adaptive Prior Learning (APL) model to adaptively select prior parameters given data. We use the Markov Chain Monte Carlo (MCMC) method in our posterior sampling model to estimate and correct for label shift. Our methods can effectively resolve class imbalance problems on large-scale datasets without fine-tuning the classifier. Experiments show that our model outperforms existing methods on a variety of label shift settings. Our code is available at https://github.com/ChangkunYe/MAPLS/.
Changkun Ye, Russell Tsuchida, Lars Petersson, Nick Barnes
WACV3
2024 Curved Geometric Networks for Visual Anomaly Recognition
abstract
Learning a latent embedding to understand the underlying nature of data distribution is often formulated in Euclidean spaces with zero curvature. However, the success of the geometry constraints, posed in the embedding space, indicates that curved spaces might encode more structural information, leading to better discriminative power and hence richer representations. In this work, we investigate the benefits of the curved space for analyzing anomalous, open-set, or out-of-distribution (OOD) objects in data. This is achieved by considering embeddings via three geometry constraints, namely, spherical geometry (with positive curvature), hyperbolic geometry (with negative curvature), or mixed geometry (with both positive and negative curvatures). Three geometric constraints can be chosen interchangeably in a unified design, given the task at hand. Tailored for the embeddings in the curved space, we also formulate functions to compute the anomaly score. Two types of geometric modules (i.e., geometric-in-one (GiO) and geometric-in-two (GiT) models) are proposed to plug in the original Euclidean classifier, and anomaly scores are computed from the curved embeddings. We evaluate the resulting designs under a diverse set of visual recognition scenarios, including image detection (multiclass OOD detection and one-class anomaly detection) and segmentation (multiclass anomaly segmentation and one-class anomaly segmentation). The empirical results show the effectiveness of our proposal through consistent improvement over various scenarios. The code is made available at https://github.com/JHome1/GiO-GiT.
Pengfei Fang, Weihao Li 0005, Junlin Han, Lars Petersson, Mehrtash Harandi
IEEE Trans. Neural Networks Learn. Syst.5
2024 GOSS: towards generalized open-set semantic segmentation
abstract
Abstract In this paper, we extend Open-set Semantic Segmentation (OSS) into a new image segmentation task called Generalized Open-set Semantic Segmentation (GOSS). Previously, with well-known OSS, the intelligent agents only detect unknown regions without further processing, limiting their perception capacity of the environment. It stands to reason that further analysis of the detected unknown pixels would be beneficial for agents’ decision-making. Therefore, we propose GOSS, which holistically unifies the abilities of two well-defined segmentation tasks, i.e. OSS and generic segmentation. Specifically, GOSS classifies pixels as belonging to known classes, and clusters (or groups) of pixels of unknown class are labelled as such. We propose a metric that balances the pixel classification and clustering aspects to evaluate this newly expanded task. Moreover, we build benchmark tests on existing datasets and propose neural architectures as baselines. Our experiments on multiple benchmarks demonstrate the effectiveness of our baselines. Code is made available at https://github.com/JHome1/GOSS_Segmentor .
Weihao Li 0005, Junlin Han, Jiyang Zheng, Pengfei Fang, Mehrtash Harandi, Lars Petersson
Vis. Comput.7
2024 Publisher Correction: GOSS: towards generalized open-set semantic segmentation
Weihao Li 0005, Junlin Han, Jiyang Zheng, Pengfei Fang, Mehrtash Harandi, Lars Petersson
Vis. Comput.7
2023 Hyperbolic Audio-visual Zero-shot Learning
abstract
Audio-visual zero-shot learning aims to classify samples consisting of a pair of corresponding audio and video sequences from classes that are not present during training. An analysis of the audio-visual data reveals a large degree of hyperbolicity, indicating the potential benefit of using a hyperbolic transformation to achieve curvature-aware geometric learning, with the aim of exploring more complex hierarchical data structures for this task. The proposed approach employs a novel loss function that incorporates cross-modality alignment between video and audio features in the hyperbolic space. Additionally, we explore the use of multiple adaptive curvatures for hyperbolic projections. The experimental results on this very challenging task demonstrate that our proposed hyperbolic approach for zero-shot learning outperforms the SOTA method on three datasets: VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL achieving a harmonic mean (HM) improvement of around 3.0%, 7.0%, and 5.3%, respectively.
Zeeshan Hayder, Junlin Han, Pengfei Fang, Mehrtash Harandi, Lars Petersson
ICCV6
2023 Weakly-supervised Point Cloud Instance Segmentation with Geometric Priors
abstract
This paper investigates how to leverage more readily acquired annotations, i.e., 3D bounding boxes instead of dense point-wise labels, for instance segmentation. We propose a Weakly-supervised point cloud Instance Segmentation framework with Geometric Priors (WISGP) that allows segmentation models to be trained with 3D bounding boxes of instances. Considering intersections among bounding boxes in a scene would result in ambiguous la- bels, we first group points into two sets, i.e., univocal and equivocal sets, indicating the certainty of a 3D point belonging to an instance, respectively. Specifically, 3D points with clear labels belong to the univocal set while the rest are grouped into the equivocal set. To assign reliable labels to points in the equivocal set, we design a Geometry-guided Label Propagation (GLP) scheme that progressively propagates labels to linked points based on geometric structure, e.g., polygon meshes and superpoints. Afterwards, we train an instance segmentation model with the univocal points and equivocal points labeled by GLP, and then employ it to assign pseudo labels for the remainder of the unlabeled points. Lastly, we retrain the model with all the labeled points to achieve better instance segmentation performance. Experiments on large-scale datasets ScanNet-v2 and S3DIS demonstrate that WISGP is superior to competing weakly-supervised algorithms and even on par with a few fully-supervised ones.
Heming Du, Xin Yu 0002, Farookh Khadeer Hussain, Mohammad Ali Armin, Lars Petersson, Weihao Li 0005
WACV5
2023 Poincaré Kernels for Hyperbolic Representations
Pengfei Fang, Mehrtash Harandi, Zhen-Zhong Lan, Lars Petersson
Int. J. Comput. Vis.4
2022 Transcribing Natural Languages for the Deaf via Neural Editing Programs
abstract
This work studies the task of glossification, of which the aim is to em transcribe natural spoken language sentences for the Deaf (hard-of-hearing) community to ordered sign language glosses. Previous sequence-to-sequence language models trained with paired sentence-gloss data often fail to capture the rich connections between the two distinct languages, leading to unsatisfactory transcriptions. We observe that despite different grammars, glosses effectively simplify sentences for the ease of deaf communication, while sharing a large portion of vocabulary with sentences. This has motivated us to implement glossification by executing a collection of editing actions, e.g. word addition, deletion, and copying, called editing programs, on their natural spoken language counterparts. Specifically, we design a new neural agent that learns to synthesize and execute editing programs, conditioned on sentence contexts and partial editing results. The agent is trained to imitate minimal editing programs, while exploring more widely the program space via policy gradients to optimize sequence-wise transcription quality. Results show that our approach outperforms previous glossification models by a large margin, improving the BLEU-4 score from 16.45 to 18.89 on RWTH-PHOENIX-WEATHER-2014T and from 18.38 to 21.30 on CSL-Daily.
Dongxu Li 0003, Liu Liu 0009, Yiran Zhong, Lars Petersson, Hongdong Li
AAAI6
2022 Blind Image Decomposition
Junlin Han, Weihao Li 0005, Pengfei Fang, Chunyi Sun, Mohammad Ali Armin, Lars Petersson, Hongdong Li
ECCV (18)7
2022 Declarative nets that are equilibrium models
Russell Tsuchida, Suk Yee Yong, Mohammad Ali Armin, Lars Petersson, Cheng Soon Ong
ICLR4
2022 You Only Cut Once: Boosting Data Augmentation with a Single Cut
abstract
We present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying YOCO improves the diversity of the augmentation per sample and encourages neural networks to recognize objects from partial information. YOCO enjoys the properties of parameter-free, easy usage, and boosting almost all augmentations for free. Thorough experiments are conducted to evaluate its effectiveness. We first demonstrate that YOCO can be seamlessly applied to varying data augmentations, neural network architectures, and brings performance gains on CIFAR and ImageNet classification tasks, sometimes surpassing conventional image-level augmentation by large margins. Moreover, we show YOCO benefits contrastive pre-training toward a more powerful representation that can be better transferred to multiple downstream tasks. Finally, we study a number of variants of YOCO and empirically analyze the performance for respective settings.
Junlin Han, Pengfei Fang, Weihao Li 0005, Mohammad Ali Armin, Ian D. Reid 0001, Lars Petersson, Hongdong Li
ICML7
2022 Efficient Gaussian Process Model on Class-Imbalanced Datasets for Generalized Zero-Shot Learning
abstract
Zero-Shot Learning (ZSL) models aim to classify object classes that are not seen during the training process. However, the problem of class imbalance is rarely discussed, despite its presence in several ZSL datasets. In this paper, we propose a Neural Network model that learns a latent feature embedding and a Gaussian Process (GP) regression model that predicts latent feature prototypes of unseen classes. A calibrated classifier is then constructed for ZSL and Generalized ZSL tasks. Our Neural Network model is trained efficiently with a simple training strategy that mitigates the impact of class-imbalanced training data. The model has an average training time of 5 minutes and can achieve state-of-the-art (SOTA) performance on imbalanced ZSL benchmark datasets like AWA2, AWA1 and APY, while having relatively good performance on the SUN and CUB datasets.
Changkun Ye, Nick Barnes, Lars Petersson, Russell Tsuchida
ICPR3
2022 Biomass Prediction with 3D Point Clouds from LiDAR
abstract
With population growth and a shrinking rural workforce, agricultural technologies have become increasingly important. Above-ground biomass (AGB) is a key trait relevant to breeding, agronomy and crop physiology field experiments. However, measuring the biomass of a cereal plot requires cutting, drying and weighing processes, which are laborious, expensive and destructive tasks. This paper proposes a non-destructive and high-throughput method to predict biomass from field samples based on Light Detection and Ranging (LiDAR). Unlike previous methods that are based on the density of a point cloud or plant height, our biomass prediction network (BioNet) additionally considers plant structure. Our BioNet contains three modules: 1) a completion module to predict missing points due to canopy occlusion; 2) a regularization module to regularize the neural representation of the whole plot; and 3) a projection module to learn the salient structures from a bird’s eye view of the point cloud. An attention-based fusion block is used to achieve final biomass predictions. In addition, the complete dataset, including hand-measured biomass and LiDAR data, is made available to the community. Experiments show that our BioNet achieves ≈ 33% improvement over current state-of-the-art methods.
Liyuan Pan, Liu Liu 0009, Anthony G. Condon, Gonzalo M. Estavillo, Robert Coe, Geoff Bull, Eric A. Stone, Lars Petersson, Vivien Rolland
WACV8
2022 Towards a Robust Differentiable Architecture Search under Label Noise
abstract
Neural Architecture Search (NAS) is the game changer in designing robust neural architectures. Architectures designed by NAS outperform or compete with the best manual network designs in terms of accuracy, size, memory footprint and FLOPs. That said, previous studies focus on developing NAS algorithms for clean high quality data, a restrictive and somewhat unrealistic assumption. In this paper, focusing on the differentiable NAS algorithms, we show that vanilla NAS algorithms suffer from a performance loss if class labels are noisy. To combat this issue, we make use of the principle of information bottleneck as a regularizer. This leads us to develop a noise injecting operation that is included during the learning process, preventing the network from learning from noisy samples. Our empirical evaluations show that the noise injecting operation does not degrade the performance of the NAS algorithm if the data is indeed clean. In contrast, if the data is noisy, the architecture learned by our algorithm comfortably outperforms algorithms specifically equipped with sophisticated mechanisms to learn in the presence of label noise. In contrast to many algorithms designed to work in the presence of noisy labels, prior knowledge about the properties of the noise and its characteristics are not required for our algorithm.
Christian Simon, Piotr Koniusz, Lars Petersson, Mehrtash Harandi
WACV3
2022 Zero-Shot Learning on 3D Point Cloud Objects and Beyond
Ali Cheraghian, Shafin Rahman, Townim F. Chowdhury, Dylan Campbell, Lars Petersson
Int. J. Comput. Vis.5
2022 Attention in Attention Networks for Person Retrieval
abstract
This paper generalizes the Attention in Attention (AiA) mechanism, in P. Fang et al., 2019 by employing explicit mapping in reproducing kernel Hilbert spaces to generate attention values of the input feature map. The AiA mechanism models the capacity of building inter-dependencies among the local and global features by the interaction of inner and outer attention modules. Besides a vanilla AiA module, termed linear attention with AiA, two non-linear counterparts, namely, second-order polynomial attention and Gaussian attention, are also proposed to utilize the non-linear properties of the input features explicitly, via the second-order polynomial kernel and Gaussian kernel approximation. The deep convolutional neural network, equipped with the proposed AiA blocks, is referred to as Attention in Attention Network (AiA-Net). The AiA-Net learns to extract a discriminative pedestrian representation, which combines complementary person appearance and corresponding part features. Extensive ablation studies verify the effectiveness of the AiA mechanism and the use of non-linear features hidden in the feature map for attention design. Furthermore, our approach outperforms current state-of-the-art by a considerable margin across a number of benchmarks. In addition, state-of-the-art performance is also achieved in the video person retrieval task with the assistance of the proposed AiA blocks.
Pengfei Fang, Jieming Zhou, Soumava Kumar Roy, Pan Ji, Lars Petersson, Mehrtash Harandi
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Semantic-Aware Knowledge Distillation for Few-Shot Class-Incremental Learning
abstract
Few-shot class incremental learning (FSCIL) portrays the problem of learning new concepts gradually, where only a few examples per concept are available to the learner. Due to the limited number of examples for training, the techniques developed for standard incremental learning cannot be applied verbatim to FSCIL. In this work, we introduce a distillation algorithm to address the problem of FSCIL and propose to make use of semantic information during training. To this end, we make use of word embeddings as semantic information which is cheap to obtain and which facilitate the distillation process. Furthermore, we propose a method based on an attention mechanism on multiple parallel embeddings of visual data to align visual and semantic vectors, which reduces issues related to catastrophic forgetting. Via experiments on MiniImageNet, CUB200, and CIFAR100 dataset, we establish new state-of-the-art results by outperforming existing approaches.
Ali Cheraghian, Shafin Rahman, Pengfei Fang, Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi
CVPR5
2021 Reinforced Attention for Few-Shot Learning and Beyond
abstract
Few-shot learning aims to correctly recognize query samples from unseen classes given a limited number of support samples, often by relying on global embeddings of images. In this paper, we propose to equip the backbone network with an attention agent, which is trained by reinforcement learning. The policy gradient algorithm is employed to train the agent towards adaptively localizing the representative regions on feature maps over time. We further design a reward function based on the prediction of the held-out data, thus helping the attention mechanism to generalize better across the unseen classes. The extensive experiments show, with the help of the reinforced attention, that our embedding network has the capability to progressively generate a more discriminative representation in few-shot learning. Moreover, experiments on the task of image classification also show the effectiveness of the proposed design.
Pengfei Fang, Weihao Li 0005, Tong Zhang 0023, Christian Simon, Mehrtash Harandi, Lars Petersson
CVPR7
2021 Contextually Plausible and Diverse 3D Human Motion Prediction
abstract
We tackle the task of diverse 3D human motion prediction, that is, forecasting multiple plausible future 3D poses given a sequence of observed 3D poses. In this context, a popular approach consists of using a Conditional Variational Autoencoder (CVAE). However, existing approaches that do so either fail to capture the diversity in human motion, or generate diverse but semantically implausible continuations of the observed motion. In this paper, we address both of these problems by developing a new variational framework that accounts for both diversity and context of the generated future motion. To this end, and in contrast to existing approaches, we condition the sampling of the latent variable that acts as source of diversity on the representation of the past observation, thus encouraging it to carry relevant information. Our experiments demonstrate that our approach yields motions not only of higher quality while retaining diversity, but also that preserve the contextual information contained in the observed motion.
Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Lars Petersson, Stephen Gould, Mathieu Salzmann
ICCV3
2021 Synthesized Feature based Few-Shot Class-Incremental Learning on a Mixture of Subspaces
abstract
Few-shot class incremental learning (FSCIL) aims to incrementally add sets of novel classes to a well-trained base model in multiple training sessions with the restriction that only a few novel instances are available per class. While learning novel classes, FSCIL methods gradually forget base (old) class training and overfit to a few novel class samples. Existing approaches have addressed this problem by computing the class prototypes from the visual or semantic word vector domain. In this paper, we propose addressing this problem using a mixture of subspaces. Subspaces define the cluster structure of the visual domain and help to describe the visual and semantic domain considering the overall distribution of the data. Additionally, we propose to employ a variational autoencoder (VAE) to generate synthesized visual samples for augmenting pseudo-feature while learning novel classes incrementally. The combined effect of the mixture of subspaces and synthesized features reduces the forgetting and overfitting problem of FSCIL. Extensive experiments on three image classification datasets show that our proposed method achieves competitive results compared to state-of-the-art methods.
Ali Cheraghian, Shafin Rahman, Sameera Ramasinghe, Pengfei Fang, Christian Simon, Lars Petersson, Mehrtash Harandi
ICCV6
2021 Kernel Methods in Hyperbolic Spaces
abstract
Embedding data in hyperbolic spaces has proven beneficial for many advanced machine learning applications such as image classification and word embeddings. However, working in hyperbolic spaces is not without difficulties as a result of its curved geometry (e.g., computing the Frechet mean of a set of points requires an iterative algorithm). Furthermore, in Euclidean spaces, one can resort to kernel machines that not only enjoy rich theoretical properties but that can also lead to superior representational power (e.g., infinite-width neural networks). In this paper, we introduce positive definite kernel functions for hyperbolic spaces. This brings in two major advantages, 1. kernelization will pave the way to seamlessly benefit from kernel machines in conjunction with hyperbolic embeddings, and 2. the rich structure of the Hilbert spaces associated with kernel machines enables us to simplify various operations involving hyperbolic data. That said, identifying valid kernel functions on curved spaces is not straightforward and is indeed considered an open problem in the learning community. Our work addresses this gap and develops several valid positive definite kernels in hyperbolic spaces, including the universal ones (e.g., RBF). We comprehensively study the proposed kernels on a variety of challenging tasks including few-shot learning, zero-shot learning, person reidentification and knowledge distillation, showing the superiority of the kernelization for hyperbolic representations.
Pengfei Fang, Mehrtash Harandi, Lars Petersson
ICCV3
2021 Single Underwater Image Restoration by Contrastive Learning
abstract
Underwater image restoration attracts significant attention due to its importance in unveiling the underwater world. This paper elaborates on a novel method that achieves state-of-the-art results for underwater image restoration based on the unsupervised image-to-image translation framework. We design our method by leveraging from contrastive learning and generative adversarial networks to maximize mutual information between raw and restored images. Additionally, we release a large-scale real underwater image dataset to support both paired and unpaired training modules. Extensive experiments with comparisons to recent approaches further demonstrate the superiority of our proposed method.
Junlin Han, Mehrdad Shoeiby, Timothy J. Malthus, Elizabeth J. Botha, Janet M. Anstee, Saeed Anwar, Lars Petersson, Mohammad Ali Armin
IGARSS8
2021 Set Augmented Triplet Loss for Video Person Re-Identification
abstract
Modern video person re-identification (re-ID) machines are often trained using a metric learning approach, supervised by a triplet loss. The triplet loss used in video re-ID is usually based on so-called clip features, each aggregated from a few frame features. In this paper, we propose to model the video clip as a set and instead study the distance between sets in the corresponding triplet loss. In contrast to the distance between clip representations, the distance between clip sets considers the pair-wise similarity of each element (i.e., frame representation) between two sets. This allows the network to directly optimize the feature representation at a frame level. Apart from the commonly-used set distance metrics (e.g., ordinary distance and Hausdorff distance), we further propose a hybrid distance metric, tailored for the set-aware triplet loss. Also, we propose a hard positive set construction strategy using the learned class prototypes in a batch. Our proposed method achieves state-of-the-art results across several standard benchmarks, demonstrating the advantages of the proposed method.
Pengfei Fang, Pan Ji, Lars Petersson, Mehrtash Harandi
WACV3
2021 Discrepant collaborative training by Sinkhorn divergences
Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi
Image Vis. Comput.3
2021 Multi-FAN: multi-spectral mosaic super-resolution via multi-scale feature aggregation network
Mehrdad Shoeiby, Mohammad Sadegh Ali Akbarian, Saeed Anwar, Lars Petersson
Mach. Vis. Appl.4
2020 Channel Recurrent Attention Networks for Video Pedestrian Retrieval
Pengfei Fang, Pan Ji, Jieming Zhou, Lars Petersson, Mehrtash Harandi
ACCV (6)4
2020 A Stochastic Conditioning Scheme for Diverse Human Motion Prediction
abstract
Human motion prediction, the task of predicting future 3D human poses given a sequence of observed ones, has been mostly treated as a deterministic problem. However, human motion is a stochastic process: Given an observed sequence of poses, multiple future motions are plausible. Existing approaches to modeling this stochasticity typically combine a random noise vector with information about the previous poses. This combination, however, is done in a deterministic manner, which gives the network the flexibility to learn to ignore the random noise. Alternatively, in this paper, we propose to stochastically combine the root of variations with previous pose information, so as to force the model to take the noise into account. We exploit this idea for motion prediction by incorporating it into a recurrent encoder-decoder network with a conditional variational autoencoder block that learns to exploit the perturbations. Our experiments on two large-scale motion prediction datasets demonstrate that our model yields high-quality pose sequences that are much more diverse than those from state-of-the-art stochastic motion prediction techniques.
Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Lars Petersson, Stephen Gould
CVPR4
2020 Transferring Cross-Domain Knowledge for Video Sign Language Recognition
abstract
Word-level sign language recognition (WSLR) is a fundamental task in sign language interpretation. It requires models to recognize isolated sign words from videos. However, annotating WSLR data needs expert knowledge, thus limiting WSLR dataset acquisition. On the contrary, there are abundant subtitled sign news videos on the internet. Since these videos have no word-level annotation and exhibit a large domain gap from isolated signs, they cannot be directly used for training WSLR models. We observe that despite the existence of a large domain gap, isolated and news signs share the same visual concepts, such as hand gestures and body movements. Motivated by this observation, we propose a novel method that learns domain-invariant visual concepts and fertilizes WSLR models by transferring knowledge of subtitled news sign to them. To this end, we extract news signs using a base WSLR model, and then design a classifier jointly trained on news and isolated signs to coarsely align these two domain features. In order to learn domain-invariant features within each class and suppress domain-specific features, our method further resorts to an external memory to store the class centroids of the aligned news signs. We then design a temporal attention based on the learnt descriptor to improve recognition performance. Experimental results on standard WSLR datasets show that our method outperforms previous state-of-the-art methods significantly. We also demonstrate the effectiveness of our method on automatically localizing signs from sign news, achieving 28.1 for [email protected].
Dongxu Li 0003, Xin Yu 0002, Lars Petersson, Hongdong Li
CVPR4
2020 Transductive Zero-Shot Learning for 3D Point Cloud Classification
abstract
Zero-shot learning, the task of learning to recognize new classes not seen during training, has received considerable attention in the case of 2D image classification. However despite the increasing ubiquity of 3D sensors, the corresponding 3D point cloud classification problem has not been meaningfully explored and introduces new challenges. This paper extends, for the first time, transductive ZeroShot Learning (ZSL) and Generalized Zero-Shot Learning (GZSL) approaches to the domain of 3D point cloud classification. To this end, a novel triplet loss is developed that takes advantage of unlabeled test data. While designed for the task of 3D point cloud classification, the method is also shown to be applicable to the more common use-case of 2D image classification. An extensive set of experiments is carried out, establishing state-of-the-art for ZSL and GZSL in the 3D point cloud domain, as well as demonstrating the applicability of the approach to the image domain.1
Ali Cheraghian, Shafin Rahman, Dylan Campbell, Lars Petersson
WACV4
2020 Learning from Noisy Labels via Discrepant Collaborative Training
abstract
Noise is ubiquitous in the world around us. Difficulty in estimating the noise within a dataset makes learning from such a dataset a difficult and challenging task. In this paper, we propose a novel and effective learning framework in order to alleviate the adverse effects of noise within a dataset. Towards this aim, we modify a collaborative training framework to utilize discrepancy constraints between respective feature extractors enabling the learning of distinct, yet discriminative features, pacifying the adverse effects of noise. Empirical results of our proposed algorithm, Discrepant Collaborative Training (DCT), achieve competitive results against several current state-of-the-art algorithms across MNIST, CIFAR10 and CIFAR100, as well as large fine-grained image classification datasets such as CUBS-200-2011 and CARS196 for different levels of noise.
Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi
WACV3
2020 Super-resolved Chromatic Mapping of Snapshot Mosaic Image Sensors via a Texture Sensitive Residual Network
abstract
This paper introduces a novel method to simultaneously super-resolve and colour-predict images acquired by snapshot mosaic sensors. These sensors allow for spectral images to be acquired using low-power, small form factor, solid-state CMOS sensors that can operate at video frame rates without the need for complex optical setups. Despite their desirable traits, their main drawback stems from the fact that the spatial resolution of the imagery acquired by these sensors is low. Moreover, chromatic mapping in snapshot mosaic sensors is not straightforward since the bands delivered by the sensor tend to be narrow and unevenly distributed across the range in which they operate. We tackle this drawback as applied to chromatic mapping by using a residual channel attention network equipped with a texture sensitive block. Our method significantly outperforms the traditional approach of interpolating the image and, afterwards, applying a colour matching function. This work establishes state-of-the-art in this domain while also making available to the research community a dataset containing 296 registered stereo multi-spectral/RGB images pairs.
Mehrdad Shoeiby, Lars Petersson, Mohammad Ali Armin, Mohammad Sadegh Ali Akbarian, Antonio Robles-Kelly
WACV2
2020 Cross-Correlated Attention Networks for Person Re-Identification
Jieming Zhou, Soumava Kumar Roy, Pengfei Fang, Mehrtash Harandi, Lars Petersson
Image Vis. Comput.5
2020 Globally-Optimal Inlier Set Maximisation for Camera Pose and Correspondence Estimation
abstract
Estimating the 6-DoF pose of a camera from a single image relative to a 3D point-set is an important task for many computer vision applications. Perspective-n-point solvers are routinely used for camera pose estimation, but are contingent on the provision of good quality 2D-3D correspondences. However, finding cross-modality correspondences between 2D image points and a 3D point-set is non-trivial, particularly when only geometric information is known. Existing approaches to the simultaneous pose and correspondence problem use local optimisation, and are therefore unlikely to find the optimal solution without a good pose initialisation, or introduce restrictive assumptions. Since a large proportion of outliers and many local optima are common for this problem, we instead propose a robust and globally-optimal inlier set maximisation approach that jointly estimates the optimal camera pose and correspondences. Our approach employs branch-and-bound to search the 6D space of camera poses, guaranteeing global optimality without requiring a pose prior. The geometry of SE(3) is used to find novel upper and lower bounds on the number of inliers and local optimisation is integrated to accelerate convergence. The algorithm outperforms existing approaches on challenging synthetic and real datasets, reliably finding the global optimum, with a GPU implementation greatly reducing runtime.
Dylan Campbell, Lars Petersson, Laurent Kneip, Hongdong Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Mitigating the Hubness Problem for Zero-Shot Learning of 3D Objects
Ali Cheraghian, Shafin Rahman, Dylan Campbell, Lars Petersson
BMVC4
2019 The Alignment of the Spheres: Globally-Optimal Spherical Mixture Alignment for Camera Pose Estimation
abstract
Determining the position and orientation of a calibrated camera from a single image with respect to a 3D model is an essential task for many applications. When 2D-3D correspondences can be obtained reliably, perspective-n-point solvers can be used to recover the camera pose. However, without the pose it is non-trivial to find cross-modality correspondences between 2D images and 3D models, particularly when the latter only contains geometric information. Consequently, the problem becomes one of estimating pose and correspondences jointly. Since outliers and local optima are so prevalent, robust objective functions and global search strategies are desirable. Hence, we cast the problem as a 2D-3D mixture model alignment task and propose the first globally-optimal solution to this formulation under the robust L2 distance between mixture distributions. We derive novel bounds on this objective function and employ branch-and-bound to search the 6D space of camera poses, guaranteeing global optimality without requiring a pose estimate. To accelerate convergence, we integrate local optimization, implement GPU bound computations, and provide an intuitive way to incorporate side information such as semantic labels. The algorithm is evaluated on challenging synthetic and real datasets, outperforming existing approaches and reliably converging to the global optimum.
Dylan Campbell, Lars Petersson, Laurent Kneip, Hongdong Li, Stephen Gould
CVPR2
2019 Bilinear Attention Networks for Person Retrieval
abstract
This paper investigates a novel Bilinear attention (Bi-attention) block, which discovers and uses second order statistical information in an input feature map, for the purpose of person retrieval. The Bi-attention block uses bilinear pooling to model the local pairwise feature interactions along each channel, while preserving the spatial structural information. We propose an Attention in Attention (AiA) mechanism to build inter-dependency among the second order local and global features with the intent to make better use of, or pay more attention to, such higher order statistical relationships. The proposed network, equipped with the proposed Bi-attention is referred to as Bilinear ATtention network (BAT-net). Our approach outperforms current state-of-the-art by a considerable margin across the standard benchmark datasets (e.g., CUHK03, Market-1501, DukeMTMC-reID and MSMT17).
Pengfei Fang, Jieming Zhou, Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi
ICCV4
2019 3DCapsule: Extending the Capsule Architecture to Classify 3D Point Clouds
abstract
This paper introduces the 3DCapsule, which is a 3D extension of the recently introduced Capsule concept that makes it applicable to unordered point sets. The original Capsule relies on the existence of a spatial relationship between the elements in the feature map it is presented with, whereas in point permutation invariant formulations of 3D point set classification methods, such relationships are typically lost. Here, a new layer called ComposeCaps is introduced that, in lieu of a spatially relevant feature mapping, learns a new mapping that can be exploited by the 3DCapsule. Previous works in the 3D point set classification domain have focused on other parts of the architecture, whereas instead, the 3DCapsule is a drop-in replacement of the commonly used fully connected classifier. It is demonstrated via an ablation study, that when the 3DCapsule is applied to recent 3D point set classification architectures, it consistently shows an improvement, in particular when subjected to noisy data. Similarly, the ComposeCaps layer is evaluated and demonstrates an improvement over the baseline. In an apples-to-apples comparison against state-of-the-art methods, again, better performance is demonstrated by the 3DCapsule.
Ali Cheraghian, Lars Petersson
WACV2
2018 VIENA ^2 : A Driving Anticipation Dataset
Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Basura Fernando, Lars Petersson, Lars Andersson
ACCV (1)5
2018 Improving Object Localization With Fitness NMS and Bounded IoU Loss
abstract
We demonstrate that many detection methods are designed to identify only a sufficiently accurate bounding box, rather than the best available one. To address this issue we propose a simple and fast modification to the existing methods called Fitness NMS. This method is tested with the DeNet model and obtains a significantly improved MAP at greater localization accuracies without a loss in evaluation rate, and can be used in conjunction with Soft NMS for additional improvements. Next we derive a novel bounding box regression loss based on a set of IoU upper bounds that better matches the goal of IoU maximization while still providing good convergence properties. Following these novelties we investigate RoI clustering schemes for improving evaluation rates for the DeNet wide model variants and provide an analysis of localization performance at various input image dimensions. We obtain a MAP of 33.6%@79Hz and 41.8%@5Hz for MSCOCO and a Titan X (Maxwell). Source code available from: https://github.com/lachlants/denet.
Lachlan Tychsen-Smith, Lars Petersson
CVPR2
2018 Effective Use of Synthetic Data for Urban Scene Semantic Segmentation
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004
ECCV (2)4
2018 Incorporating Network Built-in Priors in Weakly-Supervised Semantic Segmentation
abstract
Pixel-level annotations are expensive and time consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recently, CNN-based methods have proposed to fine-tune pre-trained networks using image tags. Without additional information, this leads to poor localization accuracy. This problem, however, was alleviated by making use of objectness priors to generate foreground/background masks. Unfortunately these priors either require pixel-level annotations/bounding boxes, or still yield inaccurate object boundaries. Here, we propose a novel method to extract accurate masks from networks pre-trained for the task of object recognition, thus forgoing external objectness modules. We first show how foreground/background masks can be obtained from the activations of higher-level convolutional layers of a network. We then show how to obtain multi-class masks by the fusion of foreground/background ones with information extracted from a weakly-supervised localization network. Our experiments evidence that exploiting these masks in conjunction with a weakly-supervised training loss yields state-of-the-art tag-based weakly-supervised semantic segmentation results.
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004, Stephen Gould
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Encouraging LSTMs to Anticipate Actions Very Early
abstract
In contrast to the widely studied problem of recognizing an action given a complete sequence, action anticipation aims to identify the action from only partially available videos. As such, it is therefore key to the success of computer vision applications requiring to react as early as possible, such as autonomous navigation. In this paper, we propose a new action anticipation method that achieves high prediction accuracy even in the presence of a very small percentage of a video sequence. To this end, we develop a multi-stage LSTM architecture that leverages context-aware and action-aware features, and introduce a novel loss function that encourages the model to predict the correct class as early as possible. Our experiments on standard benchmark datasets evidence the benefits of our approach; We outperform the state-of-the-art action anticipation methods for early prediction by a relative increase in accuracy of 22.0% on JHMDB-21, 14.0% on UT-Interaction and 49.9% on UCF-101.
Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Basura Fernando, Lars Petersson, Lars Andersson
ICCV5
2017 Globally-Optimal Inlier Set Maximisation for Simultaneous Camera Pose and Feature Correspondence
abstract
Estimating the 6-DoF pose of a camera from a single image relative to a pre-computed 3D point-set is an important task for many computer vision applications. Perspective-n-Point (PnP) solvers are routinely used for camera pose estimation, provided that a good quality set of 2D-3D feature correspondences are known beforehand. However, finding optimal correspondences between 2D key-points and a 3D point-set is non-trivial, especially when only geometric (position) information is known. Existing approaches to the simultaneous pose and correspondence problem use local optimisation, and are therefore unlikely to find the optimal solution without a good pose initialisation, or introduce restrictive assumptions. Since a large proportion of outliers are common for this problem, we instead propose a globally-optimal inlier set cardinality maximisation approach which jointly estimates optimal camera pose and optimal correspondences. Our approach employs branch-and-bound to search the 6D space of camera poses, guaranteeing global optimality without requiring a pose prior. The geometry of SE(3) is used to find novel upper and lower bounds for the number of inliers and local optimisation is integrated to accelerate convergence. The evaluation empirically supports the optimality proof and shows that the method performs much more robustly than existing approaches, including on a large-scale outdoor data-set.
Dylan Campbell, Lars Petersson, Laurent Kneip, Hongdong Li
ICCV2
2017 Bringing Background into the Foreground: Making All Classes Equal in Weakly-Supervised Video Semantic Segmentation
abstract
Pixel-level annotations are expensive and time-consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recent years have seen great progress in weakly-supervised semantic segmentation, whether from a single image or from videos. However, most existing methods are designed to handle a single background class. In practical applications, such as autonomous navigation, it is often crucial to reason about multiple background classes. In this paper, we introduce an approach to doing so by making use of classifier heatmaps. We then develop a two-stream deep architecture that jointly leverages appearance and motion, and design a loss based on our heatmaps to train it. Our experiments demonstrate the benefits of our classifier heatmaps and of our two-stream architecture on challenging urban scene datasets and on the YouTube-Objects benchmark, where we obtain state-of-the-art results.
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004
ICCV4
2017 DeNet: Scalable Real-Time Object Detection with Directed Sparse Sampling
abstract
We define the object detection from imagery problem as estimating a very large but extremely sparse bounding box dependent probability distribution. Subsequently we identify a sparse distribution estimation scheme, Directed Sparse Sampling, and employ it in a single end-to-end CNN based detection model. This methodology extends and formalizes previous state-of-the-art detection models with an additional emphasis on high evaluation rates and reduced manual engineering. We introduce two novelties, a corner based region-of-interest estimator and a deconvolution based CNN model. The resulting model is scene adaptive, does not require manually defined reference bounding boxes and produces highly competitive results on MSCOCO, Pascal VOC 2007 and Pascal VOC 2012 with real-time evaluation rates. Further analysis suggests our model performs particularly well when finegrained object localization is desirable. We argue that this advantage stems from the significantly larger set of available regions-of-interest relative to other methods. Source-code is available from: https://github.com/lachlants/denet.
Lachlan Tychsen-Smith, Lars Petersson
ICCV2
2016 GOGMA: Globally-Optimal Gaussian Mixture Alignment
abstract
Gaussian mixture alignment is a family of approaches that are frequently used for robustly solving the point-set registration problem. However, since they use local optimisation, they are susceptible to local minima and can only guarantee local optimality. Consequently, their accuracy is strongly dependent on the quality of the initialisation. This paper presents the first globally-optimal solution to the 3D rigid Gaussian mixture alignment problem under the L2 distance between mixtures. The algorithm, named GOGMA, employs a branch-and-bound approach to search the space of 3D rigid motions SE(3), guaranteeing global optimality regardless of the initialisation. The geometry of SE(3) was used to find novel upper and lower bounds for the objective function and local optimisation was integrated into the scheme to accelerate convergence without voiding the optimality guarantee. The evaluation empirically supported the optimality proof and showed that the method performed much more robustly on two challenging datasets than an existing globally-optimal registration solution.
Dylan Campbell, Lars Petersson
CVPR2
2016 Sample and Filter: Nonparametric Scene Parsing via Efficient Filtering
abstract
Scene parsing has attracted a lot of attention in computer vision. While parametric models have proven effective for this task, they cannot easily incorporate new training data. By contrast, nonparametric approaches, which bypass any learning phase and directly transfer the labels from the training data to the query images, can readily exploit new labeled samples as they become available. Unfortunately, because of the computational cost of their label transfer procedures, state-of-the-art nonparametric methods typically filter out most training images to only keep a few relevant ones to label the query. As such, these methods throw away many images that still contain valuable information and generally obtain an unbalanced set of labeled samples. In this paper, we introduce a nonparametric approach to scene parsing that follows a sample-andfilter strategy. More specifically, we propose to sample labeled superpixels according to an image similarity score, which allows us to obtain a balanced set of samples. We then formulate label transfer as an efficient filtering procedure, which lets us exploit more labeled samples than existing techniques. Our experiments evidence the benefits of our approach over state-of-the-art nonparametric methods on two benchmark datasets.
Mohammad Najafi, Sarah Taghavi Namin, Mathieu Salzmann, Lars Petersson
CVPR4
2016 Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, Stephen Gould, José M. Álvarez 0004
ECCV (8)4
2016 Latent structural SVM with marginal probabilities for weakly labeled structured learning
abstract
In the last years, the increasing availability of annotated data has facilitated the great success of supervised learning in real-world applications such as semantic labeling. However, the vast majority of data is nowadays unlabeled or partially annotated. In this paper, we develop an Expected Marginal Latent Structural SVM (EM-LSSVM) framework for performing structured learning in the presence of weakly (partially) annotated data by incorporating the uncertainty of the unobserved data as marginals. Experimental results on semantic labeling show the potential of the proposed method. In particular, we learn the parameters of a CRF where large amounts of noisy and unobserved data are available. Comparison against state of the art demonstrates the applicability of our algorithm to practical applications.
Shahin Namin, José M. Álvarez 0004, Laurent Kneip, Lars Petersson
ICIP4
2016 2D-3D semantic segmentation using cardinality as higher-order loss
abstract
Multi-modal scene analysis is a growing field of importance as additional sensors, such as 3D LIDAR, is becoming a common complement to image capturing systems. However, while additional sensory data potentially can make the analysis more accurate, it also comes with a host of associated issues. For example, inconsistencies in the data between sensors resulting from, e.g., misalignment, moving objects, or parallax effects, can severely affect the performance. Additionally, real-world scenes tend to have an inherent imbalance in the number of items of each class which typically suppresses the performance of infrequent classes. In this paper, we address those two issues specifically by a) using a cardinality loss function designed to target inconsistencies at training time, and b) devising an average per class loss function addressing the imbalance issue.
Shahin Namin, José M. Álvarez 0004, Lars Petersson
ICPR3
2015 An Adaptive Data Representation for Robust Point-Set Registration and Merging
abstract
This paper presents a framework for rigid point-set registration and merging using a robust continuous data representation. Our point-set representation is constructed by training a one-class support vector machine with a Gaussian radial basis function kernel and subsequently approximating the output function with a Gaussian mixture model. We leverage the representation's sparse parametrisation and robustness to noise, outliers and occlusions in an efficient registration algorithm that minimises the L2 distance between our support vector -- parametrised Gaussian mixtures. In contrast, existing techniques, such as Iterative Closest Point and Gaussian mixture approaches, manifest a narrower region of convergence and are less robust to occlusions and missing data, as demonstrated in the evaluation on a range of 2D and 3D datasets. Finally, we present a novel algorithm, GMMerge, that parsimoniously and equitably merges aligned mixture models, allowing the framework to be used for reconstruction and mapping.
Dylan Campbell, Lars Petersson
ICCV2
2015 Cutting Edge: Soft Correspondences in Multimodal Scene Parsing
abstract
Exploiting multiple modalities for semantic scene parsing has been shown to improve accuracy over the single modality scenario. Existing methods, however, assume that corresponding regions in two modalities have the same label. In this paper, we address the problem of data misalignment and label inconsistencies, e.g., due to moving objects, in semantic labeling, which violate the assumption of existing techniques. To this end, we formulate multimodal semantic labeling as inference in a CRF, and introduce latent nodes to explicitly model inconsistencies between two domains. These latent nodes allow us not only to leverage information from both domains to improve their labeling, but also to cut the edges between inconsistent regions. To eliminate the need for hand tuning the parameters of our model, we propose to learn intra-domain and inter-domain potential functions from training data. We demonstrate the benefits of our approach on two publicly available datasets containing 2D imagery and 3D point clouds. Thanks to our latent nodes and our learning strategy, our method outperforms the state-of-the-art in both cases.
Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann, Lars Petersson
ICCV4
2015 Modeling the cost and coverage of an ad-hoc asset management system based on existing fleet vehicles
abstract
Monitoring road assets such as road signs, utility poles, features of the road itself or other structures close to where vehicles are driving is important. Such assets need to be monitored in order to maintain them and minimize accident fatalities caused by non-compliance [1]. However, traditional surveying methods that utilize dedicated vehicles equipped with high-end expensive sensors turn out to be very costly and hence, surveys can only be carried out every few years. This paper explores the feasibility of equipping existing fleet vehicles, such as taxis, with low-end, low-quality sensors that traverse the road network through their normal daily activities. The cost and coverage of such a new approach is modeled with the help of a dataset T-Drive from Microsoft that provides taxi trajectories for more than 10,000 taxis in Beijing. The paper further estimates the optimal, from a cost perspective, number of taxis needed to survey the region by considering the cost of explicitly surveying areas that have not been covered by the random trajectories of the taxis.
Dana Pordel, Lars Petersson, Shahin Namin, Adrian Rebola-Pardo
Intelligent Vehicles Symposium2
2015 A Multi-modal Graphical Model for Scene Analysis
abstract
In this paper, we introduce a multi-modal graphical model to address the problems of semantic segmentation using 2D-3D data exhibiting extensive many-to-one correspondences. Existing methods often impose a hard correspondence between the 2D and 3D data, where the 2D and 3D corresponding regions are forced to receive identical labels. This results in performance degradation due to misalignments, 3D-2D projection errors and occlusions. We address this issue by defining a graph over the entire set of data that models soft correspondences between the two modalities. This graph encourages each region in a modality to leverage the information from its corresponding regions in the other modality to better estimate its class label. We evaluate our method on a publicly available dataset and beat the state-of-the-art. Additionally, to demonstrate the ability of our model to support multiple correspondences for objects in 3D and 2D domains, we introduce a new multi-modal dataset, which is composed of panoramic images and LIDAR data, and features a rich set of many-to-one correspondences.
Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann, Lars Petersson
WACV4
2014 Non-associative Higher-Order Markov Networks for Point Cloud Classification
Mohammad Najafi, Sarah Taghavi Namin, Mathieu Salzmann, Lars Petersson
ECCV (5)4
2014 Multi-view terrain classification using panoramic imagery and LIDAR
abstract
The focus of this work is addressing the challenges of performing object recognition in real world scenes as captured by a commercial, state-of-the-art, surveying vehicle equipped with a 360° panoramic camera in conjunction with a 3D laser scanner (LIDAR). Even with state-of-the-art surveying equipment, there is colour saturation and very dark regions in images, as well as some degree of time-varying misalignment between the point cloud data and imagery due to, for instance, imperfect tracking of sensor pose. Moreover, there are frequent occlusions due to both static and moving objects. These issues are inherently difficult to avoid and therefore need to be dealt with in a more robust fashion. This is where the contribution of the paper is; that is, the development of a consensus method that can intelligently incorporate feature responses from multiple views and reject those that are not very descriptive. It is shown that the overall performance in a ten class problem is increased from 70.5% for a simple 2D-3D classification system, to 77.5%. Subsequently, an enhanced CRF which has become robust using the misclassifications of training data and equipped with the probabilities of the adjacent points, was applied to the system and further improved its performance to 82.9%. The experiments were performed on a challenging dataset captured both in summer and winter.
Sarah Taghavi Namin, Mohammad Najafi, Lars Petersson
IROS3
2014 Creating robust high-throughput traffic sign detectors using centre-surround HOG statistics
Gary Overett, Lachlan Tychsen-Smith, Lars Petersson, Niklas Pettersson, Lars Andersson
Mach. Vis. Appl.3
2013 Classification of natural scene multi spectral images using a new enhanced CRF
abstract
In this paper, a new enhanced CRF for discriminating between different materials in natural scenes using terrestrial multi spectral imaging is established. Most of the existing formulations of the CRF often suffer from over smoothing and loss of small detail, thereby deteriorating the information from the underlying unary classifier in areas with a high spatial frequency. This work specifically addresses this issue by incorporating a new pairwise potential that is better at taking local context into account. Certain materials are very unlikely to appear next to each other in the scene and such configurations are penalised by employing the confusion matrix of the unary classifier. Similarly, horizontal as well as vertical configurations, which may be more or less likely for certain combinations of materials, are regarded in this formulation. Furthermore, the proposed pairwise potential also considers the length of boundaries between regions to account for the segmentation granularity issues and also uses class probabilities of the neighbouring regions to make up for the uncertainty of the unary classifier results. Seven band terrestrial multi spectral imaging were used due to its potential in distinguishing between different materials and objects. The proposed approach was evaluated using cross-validation, resulting in an average accuracy of 88.9% which is about 17% more than the accuracy of a standard CRF, which demonstrates the superiority of our approach in preserving local details.
Mohammad Najafi, Sarah Taghavi Namin, Lars Petersson
IROS3
2012 A computationally efficient low-bandwidth method for very-large-scale mapping of road signs with multiple vehicles
Ashkan Amirsadri, Adrian N. Bishop, Jonghyuk Kim, Jochen Trumpf, Lars Petersson
FUSION5
2012 Classification of materials in natural scenes using multi-spectral images
abstract
In this paper, a method suitable for distinguishing between different materials occurring in natural scenes using a multi-spectral camera is devised. Such a capability is useful in autonomous robot applications to help negotiating the environment as well as, e.g. applications intended to create large scale inventories of assets in the proximity of roads. The utilised sensor records a seven band multi-spectral image, of which six bands are in the visible range and one in the NIR (near-infrared) range. Many materials appearing similar if viewed by a common RGB camera, will show discriminating properties if viewed by a camera capturing a greater number of separated wavelengths. The approach in this paper is to combine the discriminating strength of the multi-spectral signature in each pixel and the corresponding nature of the surrounding texture. Local features, considering seven bands in each pixel and texture features such as GLCM and Fourier spectrum features are exploited to make the system more robust to different lighting conditions. Then classifiers built using SVM and AdaBoost are evaluated with very promising results, an average classification accuracy of 91.9% and 89.1%, respectively for a ten class problem.
Sarah Taghavi Namin, Lars Petersson
IROS2
2011 Large scale sign detection using HOG feature variants
abstract
In this paper we present two variant formulations of the well-known Histogram of Oriented Gradients (HOG) features and provide a comparison of these features on a large scale sign detection problem. The aim of this research is to find features capable of driving further improvements atop a preexisting detection framework used commercially to detect traffic signs on the scale of entire national road networks (1000's of kilometres of video). We assume the computationally efficient framework of a cascade of boosted weak classifiers. Rather than comparing features on the general problem of detection we compare their merits in the final stages of a cascaded detection problem where a feature's ability to reduce error is valued more highly than computational efficiency. Results show the benefit of the two new features on a New Zealand speed sign detection problem. We also note the importance of using non-sign training and validation instances taken from the same video data that contains the training and validation positives. This is attributed to the potential for the more powerful HOG features to overfit on specific local patterns which may be present in alternative video data.
Gary Overett, Lars Petersson
Intelligent Vehicles Symposium2
2008 Statistical Threat Assessment for General Road Scenes Using Monte Carlo Sampling
abstract
This paper presents a threat-assessment algorithm for general road scenes. A road scene consists of a number of objects that are known, and the threat level of the scene is based on their current positions and velocities. The future driver inputs of the surrounding objects are unknown and are modeled as random variables. In order to capture realistic driver behavior, a dynamic driver model is implemented as a probabilistic prior, which computes the likelihood of a potential maneuver. A distribution of possible future scenarios can then be approximated using a Monte Carlo sampling. Based on this distribution, different threat measures can be computed, e.g., probability of collision or time to collision. Since the algorithm is based on the Monte Carlo sampling, it is computationally demanding, and several techniques are presented to increase performance without increasing computational load. The algorithm is intended both for online safety applications in a vehicle and for offline data analysis.
Andreas Eidehall, Lars Petersson
IEEE Trans. Intell. Transp. Syst.2
2007 Improved Response Modelling on Weak Classifiers for Boosting
abstract
This paper demonstrates a method of increasing the quality of weak classifiers in the boosting context by using improved response modelling. The new method improves upon the results of a recent response binning approach proposed by Rasolzadeh et al. (2006). For experimental purposes the improved method is applied to the familiar Haar features as used by Viola and Jones in their face/pedestrian detection systems. However, the methods benefits are general and therefore not restricted to this particular feature type. Unlike many previous methods, this method is suitable for modelling multi-modal responses and is highly resistant to overfitting. It does this by adaptively choosing suitable support regions around the values taken by the standard response binning method. More accurate models are produced, with particular improvement around the final decision boundary. It is shown that the new method can be trained with one tenth of the training data required to achieve similar results on previous methods. This substantially lowers the overall training time of the system. The method's ability to consistently produce better hypotheses over a variety of pedestrian detection tasks is shown.
Gary Overett, Lars Petersson
ICRA2
2005 A Sign Reading Driver Assistance System Using Eye Gaze
abstract
Cars are becoming, in effect, a robotic system with an embedded human. It is not possible to know what the driver is thinking. We can, however, monitor their gaze and compare it with information in their view-field to make an inference. In this paper we present a complete system that reads speed signs in real-time, compares the driver gaze, and provides immediate feedback if the sign has been missed by the driver. This paper focuses on correlating measures of driver gaze direction with the position of signs in the road scene and improving recognition of signs through image enhancement.
Luke Fletcher, Lars Petersson, Nick Barnes, David J. Austin, Alexander Zelinsky
ICRA2
2005 Generic fusion of visual cues applied to real-world object segmentation
abstract
Fusion of information from different complementary sources may be necessary to achieve a robust sensing system that degrades gracefully under various conditions. Many approaches use a specific tailor-made combination of algorithms that do not easily allow the inclusion of more, or other, types of algorithms. In this paper, we explore a variant of a generic algorithm for fusing visual cues to the task of object segmentation in a video stream. The fusion algorithm combines the output of several segmentation algorithms in a straight forward way by using a Bayesian approach and a particle filter to track several hypotheses. Segmentation algorithms can be added or removed without changing the over all structure of the system. It was of particular interest to investigate if the method was suitable when realistic real-world scenes with much noise was analysed. The system has been tested on image sequences taken from a moving vehicle where stationary and moving objects are successfully segmented from the background. In conclusion, the fusion algorithm explored is well suited to this problem domain and is easily adopted. The context of this work is on-line pedestrian detection to be deployed in cars.
Fredrik Arnell, Lars Petersson
IROS2
2004 An Interactive Driver Assistance System Monitoring the Scene in and out of the Vehicle
abstract
This paper presents a framework for interactive driver assistance systems including techniques for fast speed sign detection and classification, car detection and tracking, and lane departure warning. In addition, the driver's actions are monitored. The integrated system uses information extracted from the road scene (speed signs, position within the lane, relative position to other cars, etc.) together with information about the driver's state such as eye gaze and head pose, to issue adequate warnings. A touch screen monitor presents relevant information and allows the driver to interact with the system. The research is focused around robust on-line algorithms. Initial results of online speed sign detection and car tracking are presented in the context of a driver assistance system.
Lars Petersson, Luke Fletcher, Nick Barnes, Alexander Zelinsky
ICRA1
2003 Driver assistance: an integration of vehicle monitoring and control
abstract
About 1.17 million people die in road crashes around the world each year. It is estimated that up to 30% of these fatalities are caused by fatigue and inattention. There are systems able to detect what is happening outside of the car, e.g., lane tracking, obstacle detection, pedestrian detection etc. Further on, there are also means for monitoring the actions of the driver. A natural step is to fuse the available data from within and outside of the car, and suggest a suitable response. This paper discusses driver assistance systems, lists a set of necessary core competencies of such a system and in particular presents a system for force-feedback in the steering wheel when crossing lanes. The presented system utilises a robust lane tracker which is experimentally evaluated for the purpose of driver assistance. In addition, preliminary results from simultaneous driver monitoring and lane tracking are presented that indicates a good correlation between the two, i.e. the driver's gaze direction and the structure of the road. These data can in turn be used for more advanced driver assistance systems in the future.
Lars Petersson, Nicholas Apostoloff, Alexander Zelinsky
ICRA1
2002 Systems Integration for Real-World Manipulation Tasks
abstract
A system developed to demonstrate integration of a number of key research areas such as localization, recognition, visual tracking, visual servoing and grasping is presented together with the underlying methodology adopted to facilitate the integration. Through sequencing of basic skills, provided by the above mentioned competencies, the system has the potential to carry out flexible grasping for fetch and carry in realistic environments. Through careful fusion of reactive and deliberative control and use of multiple sensory modalities a significant flexibility is achieved. Experimental verification of the integrated system is presented.
Lars Petersson, Patric Jensfelt, Dennis Tell, M. Strandberg, Danica Kragic, Henrik I. Christensen
ICRA1
2001 DCA: a distributed control architecture for robotics
abstract
Many control applications are by nature distributed, not only over different processes but also over several processors. Managing such a system with respect to the startup of processes, internal communications and state changes quickly becomes a very complex task. The paper presents a distributed control architecture which supports a formal model of computation as described by Lyons and Arib (1989). The architecture is primarily intended for robot control but has a wide range of potential applications. We motivate the design and implementation of the architecture by discussing the desired properties of a robot system capable of doing real-time tasks like manipulation. This leads to functionality such as a process algebra controlling the life-cycle of the processes, grouping and distribution of processes and internal communication transparent to location. Our implementation does not in itself introduce any bottlenecks due to a tree structure with local control over processes which gives an efficient and scalable architecture. At the end, an example scenario in which a fairly advanced problem like opening a door using a mobile robot with a manipulator arm is demonstrated in the presented framework.
Lars Petersson, David J. Austin, Henrik Christenseni
IROS1
2000 High-level control of a mobile manipulator for door opening
abstract
In this paper, off-the-shelf algorithms for force/torque control are used in the context of mobile manipulation, in particular, the task of opening a door is studied. To make the solution robust, as few assumptions as possible are made. By using relaxation of forces as the basic level of control more complex information can be derived from the resulting motion. In our system, the radius and centre of rotation of the door are estimated online. This enables the complete system to have a higher degree of autonomy in an unknown environment. In addition, the redundancy of the robot is exploited in such a way to drive the system towards a desired configuration. The framework of hybrid dynamic systems is used to implement the algorithm which gives a theoretically sound framework for analysing the system with respect to safety and functionality. The integration of the above approaches results in a system which can robustly locate and grasp the handle and then open the door.
Lars Petersson, David J. Austin, Danica Kragic
IROS1
1999 A hybrid control architecture for mobile manipulation
abstract
We present a scheme for mobile manipulation by introducing a mobile manipulation control architecture (MMCA). This architecture is motivated by a need for a systematic control structure for robotic manipulation within a behavior based framework. The control structure enables integration of the manipulator into a behavior based control structure for the platform. Furthermore, our suggested MMCA is designed in such a way that it supports design and performance analysis from both a manipulator dynamics and a hybrid automata perspective.
Lars Petersson, Magnus Egerstedt, Henrik I. Christensen
IROS1