EDBT 2026 Demo / reviewers in the wild / expert
Peyman Moghadam
dblp:06/2605
· DBLP profile ↗
45ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0002-8169-3560ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 4 first-author · 21 since 2021Systems, architecture and hardware · 16 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting
Binh Long Nguyen, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
ICPR (5) | 5 |
| 2026 | DFedDG2: Distribution-Guided Gossip-Based Generalizable and Communication-Efficient Decentralized Federated LearningabstractTraditional Federated Learning (FL) focuses on collaborative global model training while ensuring privacy and personalization. Decentralized Federated Learning (DFL), a variant of FL, allows clients to independently manage and optimize local models without a central server. DFL reduces centralized communication bottlenecks and vulnerability to server failures or attacks. However, because the optimization dynamics change and there is no global model, generalization can suffer, making effective learning under data and model heterogeneity a critical challenge in DFL. Despite growing interest in DFL, the lack of distributional and uncertainty modeling in the literature limits reliability and effective generalization in non-IID settings. In this work, we propose DFedDG2, a personalized federated learning framework that operates within a peer-to-peer protocol. DFedDG2 offers the technical advantage of modeling each client’s local data distribution and exchanging this information with one-hop neighbors. In the decentralized network, clients perform distribution-aware gossip, where statistically similar clients exert greater influence to drive global alignment. This likelihood-weighted mixing fuses only a handful of vectors and scalars, significantly reducing communication costs while aligning semantic spaces across the network and enabling personalized training at the edge. In addition, theoretically, we prove that DFedDG2 achieves a sublinear convergence rate while the consensus error decays at a geometric rate under well-principled properties of gossip. Unlike the label-only non-IID experiments in DFL literature, we conduct extensive experiments on multiple data non-IID scenarios, topology variations, model heterogeneity, and uncertainty quantification, demonstrating the practical advantages of DFedDG2. Our results show that DFedDG2 not only achieves communication efficiency but also provides better generalization and improved reliability compared to state-of-the-art approaches. Biprodip Pal, Stanislav Funiak, Jiajun Liu 0004, Peyman Moghadam, Md. Saiful Islam 0003, Alan Wee-Chung Liew |
IEEE Internet Things J. | 4 |
| 2025 | Always Skip AttentionabstractWe highlight a curious empirical result within modern Vision Transformers (ViTs). Specifically, self-attention catastrophically fails to train unless it is used in conjunction with a skip connection. This is in contrast to other elements of a ViT that continue to exhibit good performance (albeit suboptimal) when skip connections are removed. Further, we show that this critical dependence on skip connections is a relatively new phenomenon, with previous deep architectures (\eg, CNNs) exhibiting good performance in their absence. In this paper, we theoretically characterize that the self-attention mechanism is fundamentally ill-conditioned and is, therefore, uniquely dependent on skip connections for regularization. Additionally, we propose Token Graying -- a simple yet effective complement (to skip connections) that further improves the condition of input tokens. We validate our approach in both supervised and self-supervised training methods. Yiping Ji, Hemanth Saratchandran, Peyman Moghadam, Simon Lucey |
ICCV | 3 |
| 2025 | Shape-Space Deformer: Unified Visuo-Tactile Representations for Robotic Manipulation of Deformable ObjectsabstractAccurate modelling of object deformations is crucial for a wide range of robotic manipulation tasks, where interacting with soft or deformable objects is essential. Current methods struggle to generalise to unseen forces or adapt to new objects, limiting their utility in real-world applications. We propose Shape-Space Deformer, a unified representation for encoding a diverse range of object deformations using template augmentation to achieve robust, fine-grained reconstructions that are resilient to outliers and unwanted artefacts. Our method improves generalization to unseen forces and can rapidly adapt to novel objects, significantly outperforming existing approaches. We perform extensive experiments to test a range of force generalisation settings and evaluate our method's ability to reconstruct unseen deformations. Our results demonstrate significant improvements in reconstruction accuracy and robustness. Our approach is suitable for real-time performance, making it ready for downstream manipulation applications. Sean M. V. Collins, Brendan Tidd, Mahsa Baktash, Peyman Moghadam |
ICRA | 4 |
| 2025 | SOLVR: Submap Oriented LiDAR-Visual Re-LocalisationabstractThis paper proposes SOLVR, a unified pipeline for learning based LiDAR-Visual re-localisation which performs place recognition and 6-DoF registration across sensor modalities. We propose a strategy to align the input sensor modalities by leveraging stereo image streams to produce metric depth predictions with pose information, followed by fusing multiple scene views from a local window using a probabilistic occupancy framework to expand the limited field-of-view of the camera. Additionally, SOLVR adopts a flexible definition of what constitutes positive examples for different training losses, allowing us to simultaneously optimise place recognition and registration performance. Furthermore, we replace RANSAC with a registration function that weights a simple least-squares fitting with the estimated inlier likelihood of sparse keypoint correspondences, improving performance in scenarios with a low inlier ratio between the query and retrieved place. Our experiments on the KITTI and KITTI360 datasets show that SOLVR achieves state-of-the-art performance for LiDAR-Visual place recognition and registration, particularly improving registration accuracy over larger distances between the query and retrieved place. Joshua Knights, Sebastián Barbas Laina, Peyman Moghadam, Stefan Leutenegger |
ICRA | 3 |
| 2025 | M2Distill: Multi-Modal Distillation for Lifelong Imitation LearningabstractLifelong imitation learning for manipulation tasks poses significant challenges due to distribution shifts that occur in incremental learning steps. Existing methods often rely on unsupervised skill discovery to construct an ever-growing skill library or distillation from multiple policies, which can lead to scalability issues as diverse manipulation tasks are continually introduced and may fail to ensure a consistent latent space throughout the learning process, leading to catastrophic forgetting of previously learned skills. In this paper, we introduce M2Distill, a multimodal distillation-based method for lifelong imitation learning focusing on preserving consistent latent space across vision, language, and action distributions throughout the learning process. By regulating the shifts in latent representations across different modalities from previous to current steps, and reducing discrepancies in Gaussian Mixture Model (GMM) policies between consecutive learning steps, we ensure that the learned policy retains its ability to perform previously learned tasks while seamlessly integrating new skills. Evaluations on the LIBERO lifelong imitation learning benchmark suites, including LIBERO-OBJECT, LIBERO-GOAL, and LIBERO-SPATIAL, demonstrate that our method consistently outperforms prior state-of-the-art methods across all evaluated metrics. Kaushik Roy 0008, Akila Dissanayakc, Brendan Tidd, Peyman Moghadam |
ICRA | 4 |
| 2025 | Inductive Graph Few-shot Class Incremental LearningabstractNode classification with Graph Neural Networks (GNN) under a fixed set of labels is well studied, while Graph Few-Shot Class Incremental Learning (GFSCIL), which involves learning a GNN classifier as graph nodes and classes growing over time sporadically, has received much less attention despite its importance. We introduce inductive GFSCIL that continually learns novel classes with newly emerging nodes while maintaining performance on old classes without accessing previous data. This addresses the practical concern of transductive GFSCIL, which requires storing the entire graph with historical data. Compared to the transductive GFSCIL, the inductive setting exacerbates catastrophic forgetting due to inaccessible previous data during incremental training, in addition to the overfitting issue caused by label sparsity. Thus, we propose a novel method, called Topology-based class Augmentation and Prototype calibration (TAP). To be specific, it first performs a topology-based class augmentation method, helping replicate the setting of disjoint subgraphs with nodes of novel classes received in incremental sessions, to enhance backbone versatility. In incremental learning, given the limited number of novel class samples, we propose an iterative prototype calibration to improve the separation of class prototypes. Furthermore, as backbone fine-tuning poses the feature distribution drift, prototypes of old classes start failing over time, we propose the prototype shift method for old classes to compensate for the drift. We showcase the proposed method on four datasets. Yayong Li, Peyman Moghadam, Can Peng, Piotr Koniusz |
WSDM | 2 |
| 2025 | Flashbacks to harmonize stability and plasticity in continual learningabstractWe introduce Flashback Learning (FL), a novel method designed to harmonize the stability and plasticity of models in Continual Learning (CL). Unlike prior approaches that primarily focus on regularizing model updates to preserve old information while learning new concepts, FL explicitly balances this trade-off through a bidirectional form of regularization. This approach effectively guides the model to swiftly incorporate new knowledge while actively retaining its old knowledge. FL operates through a two-phase training process and can be seamlessly integrated into various CL methods, including replay, parameter regularization, distillation, and dynamic architecture techniques. In designing FL, we use two distinct knowledge bases: one to enhance plasticity and another to improve stability. FL ensures a more balanced model by utilizing both knowledge bases to regularize model updates. Theoretically, we analyze how the FL mechanism enhances the stability-plasticity balance. Empirically, FL demonstrates tangible improvements over baseline methods within the same training budget. By integrating FL into at least one representative baseline from each CL category, we observed an average accuracy improvement of up to 4.91% in Class-Incremental and 3.51% in Task-Incremental settings on standard image classification benchmarks. Additionally, measurements of the stability-to-plasticity ratio confirm that FL effectively enhances this balance. FL also outperforms state-of-the-art CL methods on more challenging datasets like ImageNet. The codes of this article will be available at https://github.com/csiro-robotics/Flashback-Learning. Leila Mahmoodi, Peyman Moghadam, Munawar Hayat, Christian Simon, Mehrtash Harandi |
Neural Networks | 2 |
| 2025 | Spatioformer: A Geo-Encoded Transformer for Large-Scale Plant Species Richness PredictionabstractEarth observation (EO) data have shown promise in predicting species richness of vascular plants ($\alpha $-diversity), but extending this approach to large spatial scales is challenging because geographically distant regions may exhibit different compositions of plant species ($\beta $-diversity), resulting in a location-dependent relationship between richness and spectral measurements. In order to handle such geolocation dependence, we propose Spatioformer, where a novel geolocation encoder is coupled with the transformer model to encode geolocation context into remote sensing imagery. The Spatioformer model compares favorably to state-of-the-art models in richness predictions on a large-scale ground-truth richness dataset harmonized Australian vegetation plot (HAVPlot) that consists of 68 170 in situ richness samples covering diverse landscapes across Australia. The results demonstrate that geolocational information is advantageous in predicting species richness from satellite observations over large spatial scales. With Spatioformer, plant species richness maps over Australia are compiled from the Landsat archive for the years from 2015 to 2023. The richness maps produced in this study reveal the spatiotemporal dynamics of plant species richness in Australia, providing supporting evidence to inform effective planning and policy development for plant diversity conservation. Regions of high richness prediction uncertainties are identified, highlighting the need for future in situ surveys to be conducted in these areas to enhance the prediction accuracy. Yiqing Guo, Karel Mokany, Shaun R. Levick, Jinyan Yang, Peyman Moghadam |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | TULIP: Transformer for Upsampling of LiDAR Point CloudsabstractLiDAR Upsampling is a challenging task for the perception systems of robots and autonomous vehicles, due to the sparse and irregular structure of large-scale scene contexts. Recent works propose to solve this problem by converting LiDAR data from 3D Euclidean space into an image super-resolution problem in 2D image space. Although their methods can generate high-resolution range images with fine-grained details, the resulting 3D point clouds often blur out details and predict invalid points. In this paper, we propose TULIP, a new method to reconstruct high-resolution LiDAR point clouds from low-resolution LiDAR input. We also follow a range image-based approach but specifically modify the patch and window geometries of a Swin- Transformer-based network to better fit the characteristics of range images. We conducted several experiments on three public real-world and simulated datasets. TULIP outperforms state-of-the-art methods in all relevant metrics and generates robust and more realistic point clouds than prior works. The code is available at https://github.com/ethz-asl/TULIP.git. Patrick Pfreundschuh, Roland Siegwart, Marco Hutter 0001, Peyman Moghadam, Vaishakh Patil |
CVPR | 5 |
| 2024 | Pre-training with Random Orthogonal Projection Image ModelingabstractMasked Image Modeling (MIM) is a powerful self-supervised strategy for visual pre-training without the use of labels. MIM applies random crops to input images, processes them with an encoder, and then recovers the masked inputs with a decoder, which encourages the network to capture and learn structural information about objects and scenes. The intermediate feature representations obtained from MIM are suitable for fine-tuning on downstream tasks. In this paper, we propose an Image Modeling framework based on random orthogonal projection instead of binary masking as in MIM. Our proposed Random Orthogonal Projection Image Modeling (ROPIM) reduces spatially-wise token information under guaranteed bound on the noise variance and can be considered as masking entire spatial image area under locally varying masking degrees. Since ROPIM uses a random subspace for the projection that realizes the masking step, the readily available complement of the subspace can be used during unmasking to promote recovery of removed information. In this paper, we show that using random orthogonal projection leads to superior performance compared to crop-based masking. We demonstrate state-of-the-art results on several popular benchmarks. Maryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr Koniusz |
ICLR | 2 |
| 2024 | Reg-NF: Efficient Registration of Implicit Surfaces within Neural FieldsabstractNeural fields, coordinate-based neural networks, have recently gained popularity for implicitly representing a scene. In contrast to classical methods that are based on explicit representations such as point clouds, neural fields provide a continuous scene representation able to represent 3D geometry and appearance in a way which is compact and ideal for robotics applications. However, limited prior methods have investigated registering multiple neural fields by directly utilising these continuous implicit representations. In this paper, we present Reg-NF, a neural fields-based registration that optimises for the relative 6-DoF transformation between two arbitrary neural fields, even if those two fields have different scale factors. Key components of Reg-NF include a bidirectional registration loss, multi-view surface sampling, and utilisation of volumetric signed distance functions (SDFs). We showcase our approach on a new neural field dataset for evaluating registration problems. We provide an exhaustive set of experiments and ablation studies to identify the performance of our approach, while also discussing limitations to provide future direction to the research community on open challenges in utilizing neural fields in unconstrained environments. Stephen Hausler, David Hall 0003, Sutharsan Mahendren, Peyman Moghadam |
ICRA | 4 |
| 2024 | Multivariate prototype representation for domain-generalized incremental learningabstractDeep learning models often suffer from catastrophic forgetting when fine-tuned with samples of new classes. This issue becomes even more challenging when there is a domain shift between training and testing data. In this paper, we address the critical yet less explored Domain-Generalized Class-Incremental Learning (DGCIL) task. We propose a DGCIL approach designed to memorize old classes, adapt to new classes, and reliably classify objects from unseen domains. Specifically, our loss formulation maintains classification boundaries while suppressing domain-specific information for each class. Without storing old exemplars, we employ knowledge distillation and estimate the drift of old class prototypes as incremental training progresses. Our prototype representations are based on multivariate Normal distributions , with means and covariances continually adapted to reflect evolving model features, providing effective representations for old classes. We then sample pseudo-features for these old classes from the adapted Normal distributions using Cholesky decomposition . Unlike previous pseudo-feature sampling strategies that rely solely on average mean prototypes, our method captures richer semantic variations. Experiments on several benchmarks demonstrate the superior performance of our method compared to the state of the art. Can Peng, Piotr Koniusz, Kaiyu Guo, Brian C. Lovell, Peyman Moghadam |
Comput. Vis. Image Underst. | 5 |
| 2024 | FactoFormer: Factorized Hyperspectral Transformers With Self-Supervised PretrainingabstractHyperspectral images (HSIs) contain rich spectral and spatial information. Motivated by the success of transformers in the field of natural language processing and computer vision where they have shown the ability to learn long-range dependencies within input data, recent research has focused on using transformers for HSIs. However, current state-of-the-art hyperspectral transformers only tokenize the input HSI sample along the spectral dimension, resulting in the underutilization of spatial information. Moreover, transformers are known to be data-hungry and their performance relies heavily on large-scale pretraining, which is challenging due to limited annotated hyperspectral data. Therefore, the full potential of HSI transformers has not been fully realized. To overcome these limitations, we propose a novel factorized spectral–spatial transformer that incorporates factorized self-supervised pretraining procedures, leading to significant improvements in performance. The factorization of the inputs allows the spectral and spatial transformers to better capture the interactions within the hyperspectral data cubes. Inspired by masked image modeling (MIM) pretraining, we also devise efficient masking strategies for pretraining each of the spectral and spatial transformers. We conduct experiments on six publicly available datasets for the HSI classification task and demonstrate that our model achieves state-of-the-art performance in all the datasets. The code for our model will be made available athttps://github.com/csiro-robotics/FactoFormer. Shaheer Mohamed, Maryam Haghighat, Tharindu Fernando, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Learning Partial Correlation based Deep Visual Representation for Image ClassificationabstractVisual representation based on covariance matrix has demonstrates its efficacy for image classification by characterising the pairwise correlation of different channels in convolutional feature maps. However, pairwise correlation will become misleading once there is another channel correlating with both channels of interest, resulting in the “confounding” effect. For this case, “partial correlation” which removes the confounding effect shall be estimated instead. Nevertheless, reliably estimating partial correlation requires to solve a symmetric positive definite matrix optimisation, known as sparse inverse covariance estimation (SICE). How to incorporate this process into CNN remains an open issue. In this work, we formulate SICE as a novel structured layer of CNN. To ensure end-to-end trainability, we develop an iterative method to solve the above matrix optimisation during forward and backward propagation steps. Our work obtains a partial correlation based deep visual representation and mitigates the small sample problem often encountered by covariance matrix estimation in CNN. Computationally, our model can be effectively trained with GPU and works well with a large number of channels of advanced CNNs. Experiments show the efficacy and superior classification performance of our deep visual representation compared to covariance matrix based counterparts. Saimunur Rahman, Piotr Koniusz, Lei Wang 0001, Luping Zhou, Peyman Moghadam, Changming Sun |
CVPR | 5 |
| 2023 | Wild-Places: A Large-Scale Dataset for Lidar Place Recognition in Unstructured Natural EnvironmentsabstractMany existing datasets for lidar place recognition are solely representative of structured urban environments, and have recently been saturated in performance by deep learning based approaches. Natural and unstructured environments present many additional challenges for the tasks of long-term localisation but these environments are not represented in currently available datasets. To address this we introduce Wild-Places, a challenging large-scale dataset for lidar place recognition in unstructured, natural environments. Wild-Places contains eight lidar sequences collected with a handheld sensor payload over the course of fourteen months, containing a total of 63K undistorted lidar submaps along with accurate 6DoF ground truth. This dataset contains multi-ple revisits both within and between sequences, allowing for both intra-sequence (i.e., loop closure detection) and inter-sequence (i.e., re-localisation) tasks. We also benchmark several state-of-the-art approaches to demonstrate the challenges that this dataset introduces, particularly the case of long-term place recognition due to natural environments changing over time. Our dataset and code is available at https://csiro-robotics.github.io/Wild-Places Joshua Knights, Kavisha Vidanapathirana, Milad Ramezani, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
ICRA | 6 |
| 2023 | Uncertainty-Aware Lidar Place Recognition in Novel EnvironmentsabstractState-of-the-art lidar place recognition models exhibit unreliable performance when tested on environments different from their training dataset, which limits their use in complex and evolving environments. To address this issue, we investigate the task of uncertainty-aware lidar place recognition, where each predicted place must have an associated uncertainty that can be used to identify and reject incorrect predictions. We introduce a novel evaluation protocol and present the first comprehensive benchmark for this task, testing across five uncertainty estimation techniques and three large-scale datasets. Our results show that an Ensembles approach is the highest performing technique, consistently improving the performance of lidar place recognition and uncertainty estimation in novel environments, though it incurs a computational cost. Code is publicly available at https://github.com/csiro-robotics/Uncertainty-LPR. Keita Mason, Joshua Knights, Milad Ramezani, Peyman Moghadam, Dimity Miller |
IROS | 4 |
| 2023 | Deep Robust Multi-Robot Re-Localisation in Natural EnvironmentsabstractThe success of re-localisation has crucial implications for the practical deployment of robots operating within a prior map or relative to one another in real-world scenarios. Using single-modality, place recognition and localisation can be compromised in challenging environments such as forests. To address this, we propose a strategy to prevent lidar-based re-localisation failure using lidar-image cross-modality. Our solution relies on self-supervised 2D-3D feature matching to predict alignment and misalignment. Leveraging a deep network for lidar feature extraction and relative pose estimation between point clouds, we train a model to evaluate the estimated transformation. A model predicting the presence of misalignment is learned by analysing image-lidar similarity in the embedding space and the geometric constraints available within the region seen in both modalities in Euclidean space. Experimental results using real datasets (offline and online modes) demonstrate the effectiveness of the proposed pipeline for robust re-localisation in unstructured, natural environments. Milad Ramezani, Ethan Griffiths, Maryam Haghighat, Alex Pitt, Peyman Moghadam |
IROS | 5 |
| 2023 | L3DMC: Lifelong Learning Using Distillation via Mixed-Curvature Space
Kaushik Roy 0008, Peyman Moghadam, Mehrtash Harandi |
MICCAI (2) | 2 |
| 2023 | Measuring Situational Awareness Latency in Human-Robot Teaming ExperimentsabstractA human supervisor’s Situational Awareness (SA) is a critical aspect for successful Human-Robot Teaming (HRT). SA has been estimated using different techniques; however, many of those are associated with various biases, including recall and overgeneralisation biases. A key SA metric is latency, the delay between the time the robotic system requires supervisor assistance and the time the supervisor identifies that need in HRT experiments. Eye movements are increasingly used to assess SA across a range of domains, enabling objective and continuous SA assessment. However, to date, only a small number of features have been evaluated for estimating different types of SA latencies. In this paper, we investigated how two types of SA latencies (perceptual and comprehending) correlate with eye movement data collected during a remote field experiment, where a human supervisor directed a team of robots in a smart farming context. We identified 39 instances of SA latencies (13 perceptual and 26 comprehending). These instances were used to identify how a human supervisor’s SA is affected by task context, and to evaluate correlations between five eye movement features and SA latencies. Two eye movement features related to fixation duration and saccade duration demonstrated very strong correlations ($r \approx - 0.8$ and $r \approx 0.85$). Our findings can be extended to estimate the real-time likelihood of the human experiencing SA latency. Hashini Senaratne, Alex Pitt, Fletcher Talbot, Peyman Moghadam, Pavan Sikka, Gerard David Howard, Jason Williams 0002, Dana Kulic, Cécile Paris |
RO-MAN | 4 |
| 2023 | Subspace distillation for continual learningabstractAn ultimate objective in continual learning is to preserve knowledge learned in preceding tasks while learning new tasks. To mitigate forgetting prior knowledge, we propose a novel knowledge distillation technique that takes into the account the manifold structure of the latent/output space of a neural network in learning novel tasks. To achieve this, we propose to approximate the data manifold up-to its first order, hence benefiting from linear subspaces to model the structure and maintain the knowledge of a neural network while learning novel concepts. We demonstrate that the modeling with subspaces provides several intriguing properties, including robustness to noise and therefore effective for mitigating Catastrophic Forgetting in continual learning. We also discuss and show how our proposed method can be adopted to address both classification and segmentation problems. Empirically, we observe that our proposed method outperforms various continual learning methods on several challenging datasets including Pascal VOC, and Tiny-Imagenet. Furthermore, we show how the proposed method can be seamlessly combined with existing learning approaches to improve their performances. The codes of this article will be available at https://github.com/csiro-robotics/SDCL. Kaushik Roy 0008, Christian Simon, Peyman Moghadam, Mehrtash Harandi |
Neural Networks | 3 |
| 2023 | Exploiting Field Dependencies for Learning on Categorical DataabstractTraditional approaches for learning on categorical data underexploit the dependencies between columns (a.k.a. fields) in a dataset because they rely on the embedding of data points driven alone by the classification/regression loss. In contrast, we propose a novel method for learning on categorical data with the goal of exploiting dependencies between fields. Instead of modelling statistics of features globally (i.e., by the covariance matrix of features), we learn a global field dependency matrix that captures dependencies between fields and then we refine the global field dependency matrix at the instance-wise level with different weights (so-called local dependency modelling) w.r.t. each field to improve the modelling of the field dependencies. Our algorithm exploits the meta-learning paradigm, i.e., the dependency matrices are refined in the inner loop of the meta-learning algorithm without the use of labels, whereas the outer loop intertwines the updates of the embedding matrix (the matrix performing projection) and global dependency matrix in a supervised fashion (with the use of labels). Our method is simple yet it outperforms several state-of-the-art methods on six popular dataset benchmarks. Detailed ablation studies provide additional insights into our method. Zhibin Li 0002, Piotr Koniusz, Lu Zhang 0062, Daniel Edward Pagendam, Peyman Moghadam |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | LoGG3D-Net: Locally Guided Global Descriptor Learning for 3D Place RecognitionabstractRetrieval-based place recognition is an efficient and effective solution for re-localization within a pre-built map, or global data association for Simultaneous Localization and Mapping (SLAM). The accuracy of such an approach is heavily dependant on the quality of the extracted scene-level representation. While end-to-end solutions - which learn a global descriptor from input point clouds - have demonstrated promising results, such approaches are limited in their ability to enforce desirable properties at the local feature level. In this paper, we introduce a local consistency loss to guide the network towards learning local features which are consistent across revisits, hence leading to more repeatable global descriptors resulting in an overall improvement in 3D place recognition performance. We formulate our approach in an end-to-end trainable architecture called LoGG3D-Net. Experiments on two large-scale public benchmarks (KITTI and MulRan) show that our method achieves mean F1maxscores of 0.939 and 0.968 on KITTI and MulRan respectively, achieving state-of-the-art performance while operating in near real-time. The open-source implementation is available at: https://github.com/csiro-robotics/LoGG3D-Net. Kavisha Vidanapathirana, Milad Ramezani, Peyman Moghadam, Sridha Sridharan, Clinton Fookes |
ICRA | 3 |
| 2022 | Quantitative Assessment of DESIS Hyperspectral Data for Plant Biodiversity Estimation in AustraliaabstractDiversity of terrestrial plants plays a key role in maintaining a stable, healthy, and productive ecosystem. Though remote sensing has been seen as a promising and cost-effective proxy for estimating plant diversity, there is a lack of quantitative studies on how confidently plant diversity can be inferred from spaceborne hyperspectral data. In this study, we assessed the ability of hyperspectral data captured by the DLR Earth Sensing Imaging Spectrometer (DESIS) for estimating plant species richness in the Southern Tablelands and Snowy Mountains regions in southeast Australia. Spectral features were firstly extracted from DESIS spectra with principal component analysis, canonical correlation analysis, and partial least squares analysis. Then regression was conducted between the extracted features and plant species richness with ordinary least squares regression, kernel ridge regression, and Gaussian process regression. Results were assessed with the coefficient of correlation$(r)$and Root-Mean-Square Error (RMSE), based on a two-fold cross validation scheme. With the best performing model,$r$is 0.71 and RMSE is 5.99 for the Southern Tablelands region, while$r$is 0.62 and RMSE is 6.20 for the Snowy Mountains region. The assessment results reported in this study provide supports for future studies on understanding the relationship between spaceborne hyperspectral measurements and terrestrial plant biodiversity. Yiqing Guo, Karel Mokany, Cindy Ong, Peyman Moghadam, Simon Ferrier, Shaun R. Levick |
IGARSS | 4 |
| 2022 | InCloud: Incremental Learning for Point Cloud Place RecognitionabstractPlace recognition is a fundamental component of robotics, and has seen tremendous improvements through the use of deep learning models in recent years. Networks can experience significant drops in performance when deployed in unseen or highly dynamic environments, and require additional training on the collected data. However naively fine-tuning on new training distributions can cause severe degradation of performance on previously visited domains, a phenomenon known as catastrophic forgetting. In this paper we address the problem of incremental learning for point cloud place recognition and introduce InCloud, a structure-aware distillation-based approach which preserves the higher-order structure of the network's embedding space. We introduce several challenging new benchmarks on four popular and large-scale LiDAR datasets (Oxford, MulRan, In-house and KITTI) showing broad improvements in point cloud place recognition performance over a variety of network architectures. To the best of our knowledge, this work is the first to effectively apply incremental learning for point cloud place recognition. Data pre-processing, training and evaluation code for this paper can be found at https://github.com/csiro-robotics/InCloud. Joshua Knights, Peyman Moghadam, Milad Ramezani, Sridha Sridharan, Clinton Fookes |
IROS | 2 |
| 2022 | A real-time edge-AI system for reef surveysabstractCrown-of-Thorn Starfish (COTS) outbreaks are a major cause of coral loss on the Great Barrier Reef (GBR) and substantial surveillance and control programs are ongoing to manage COTS populations to ecologically sustainable levels. In this paper, we present a comprehensive real-time machine learning-based underwater data collection and curation system on edge devices for COTS monitoring. In particular, we leverage the power of deep learning-based object detection techniques, and propose a resource-efficient COTS detector that performs detection inferences on the edge device to assist marine experts with COTS identification during the data collection phase. The preliminary results show that several strategies for improving computational efficiency (e.g., batch-wise processing, frame skipping, model input size) can be combined to run the proposed detection model on edge hardware with low resource consumption and low information loss. Yang Li 0184, Jiajun Liu 0004, Branislav Kusy, Ross Marchant, Brendan Do, Torsten Merz, Joey Crosswell, Andrew D. L. Steven, Lachlan Tychsen-Smith, David Ahmedt-Aristizabal, Jeremy Oorloff, Peyman Moghadam, Russ Babcock, Megha Malpani, Ard Oerlemans |
MobiCom | 12 |
| 2022 | Elasticity Meets Continuous-Time: Map-Centric Dense 3D LiDAR SLAMabstractMap-centric SLAM utilizes elasticity as a means of loop closure. This approach reduces the cost of loop closure while still providing large-scale fusion-based dense maps, when compared to trajectory-centric SLAM approaches. In this article, we present a novel framework, namedElasticLiDAR++, for multimodal map-centric SLAM. Having the advantages of a map-centric approach, our method exhibits new features to overcome the shortcomings of existing systems associated with multimodal (LiDAR-inertial-visual) sensor fusion and LiDAR motion distortion. This is accomplished through the use of a local continuous-time trajectory representation. Also, our surface resolution preserving matching algorithm and normal-inverse-Wishart-based surfel fusion model enables nonredundant yet dense mapping. Furthermore, we present a robust metric loop closure model to make the approach stable regardless of where the loop closure occurs. Finally, we demonstrate our approach through both simulation and real data experiments using multiple sensor payload configurations and environments to illustrate its utility and robustness. Chanoh Park, Peyman Moghadam, Jason Williams 0002, Soohwan Kim, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Robotics | 2 |
| 2021 | Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order PoolingabstractPlace Recognition enables the estimation of a globally consistent map and trajectory by providing non-local constraints in Simultaneous Localisation and Mapping (SLAM). This paper presents Locus, a novel place recognition method using 3D LiDAR point clouds in large-scale environments. We propose a method for extracting and encoding topological and temporal information related to components in a scene and demonstrate how the inclusion of this auxiliary information in place description leads to more robust and discriminative scene representations. Second-order pooling along with a non- linear transform is used to aggregate these multi-level features to generate a fixed-length global descriptor, which is invariant to the permutation of input features. The proposed method outperforms state-of-the-art methods on the KITTI dataset. Furthermore, Locus is demonstrated to be robust across several challenging situations such as occlusions and viewpoint changes in 3D LiDAR point clouds. The open-source implementation is available at: https://github.com/csiro-robotics/locus. Kavisha Vidanapathirana, Peyman Moghadam, Ben Harwood, Muming Zhao, Sridha Sridharan, Clinton Fookes |
ICRA | 2 |
| 2020 | Temporally Coherent Embeddings for Self-Supervised Video Representation LearningabstractThis paper presents TCE: Temporally Coherent Embeddings for self-supervised video representation learning. The proposed method exploits inherent structure of unlabeled video data to explicitly enforce temporal coherency in the embedding space, rather than indirectly learning it through ranking or predictive proxy tasks. In the same way that high-level visual information in the world changes smoothly, we believe that nearby frames in learned representations will benefit from demonstrating similar properties. Using this assumption, we train our TCE model to encode videos such that adjacent frames exist close to each other and videos are separated from one another. Using TCE we learn robust representations from large quantities of unlabeled video data. We thoroughly analyse and evaluate our self-supervised learned TCE models on a downstream task of video action recognition using multiple challenging benchmarks (Kinetics400, UCF101, HMDB51). With a simple but effective 2D-CNN backbone and only RGB stream inputs, TCE pre-trained representations outperform all previous self-supervised 2D-CNN and 3D-CNN pre-trained on UCF101. The code and pre-trained models for this paper can be downloaded at: https://github.com/csiro-robotics/TCE. Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, Peyman Moghadam |
ICPR | 6 |
| 2020 | Scalable learning for bridging the species gap in image-based plant phenotyping
Daniel Ward, Peyman Moghadam |
Comput. Vis. Image Underst. | 2 |
| 2018 | Deep Leaf Segmentation Using Synthetic Data
Daniel Ward, Peyman Moghadam, Nicolas Hudson |
BMVC | 2 |
| 2018 | Non-rigid Reconstruction with a Single Moving RGB-D CameraabstractWe present a novel non-rigid reconstruction method using a moving RGB-D camera. Current approaches use only non-rigid part of the scene and completely ignore the rigid background. Non-rigid parts often lack sufficient geometric and photometric information for tracking large frame-to-frame motion. Our approach uses camera pose estimated from the rigid background for foreground tracking. This enables robust foreground tracking in situations where large frame-to-frame motion occurs. Moreover, we are proposing a multi-scale deformation graph which improves non-rigid tracking without compromising the quality of the reconstruction. We are also contributing a synthetic dataset which is made publically available for evaluating non-rigid reconstruction methods. The dataset provides frame-by-frame ground truth geometry of the scene, the camera trajectory, and masks for background foreground. Experimental results show that our approach is more robust in handling larger frame-to-frame motions and provides better reconstruction compared to state-of-the-art approaches. Shafeeq Elanattil, Peyman Moghadam, Sridha Sridharan, Clinton Fookes, Mark Cox |
ICPR | 2 |
| 2018 | Elastic LiDAR Fusion: Dense Map-Centric Continuous-Time SLAMabstractThe concept of continuous-time trajectory representation has brought increased accuracy and efficiency to multi-modal sensor fusion in modern SLAM. However, regardless of these advantages, its offline property caused by the requirement of global batch optimization is critically hindering its relevance for real-time and life-long applications. In this paper, we present a dense map-centric SLAM method based on a continuous-time trajectory to cope with this problem. The proposed system locally functions in a similar fashion to conventional Continuous-Time SLAM (CT-SLAM). However, it removes the need for global trajectory optimization by introducing map deformation. The computational complexity of the proposed approach for loop closure does not depend on the operation time, but only on the size of the space it explored before the loop closure. It is therefore more suitable for long term operation compared to the conventional CT-SLAM. Furthermore, the proposed method reduces uncertainty in the reconstructed dense map by using probabilistic surface element (surfel) fusion. We demonstrate that the proposed method produces globally consistent maps without global batch trajectory optimization, and effectively reduces LiDAR noise by surfel fusion. Chanoh Park, Peyman Moghadam, Soohwan Kim, Alberto Elfes, Clinton Fookes, Sridha Sridharan |
ICRA | 2 |
| 2015 | Energetics-informed hexapod gait transitions across terrainsabstractLegged robots offer the potential of locomotion across various types of terrains. Different terrains require different gait patterns to enable greater traversal efficiency. Consequently, as a legged robot transitions from one type of terrain to another, the gait pattern should be adapted so as to maximise traction and energy efficiency. This paper explores the use of power consumption as estimated by the robot in real-time for guiding this gait transition in the case of statically-stable locomotion. While moving, the robot autonomously assesses its power consumption, relates it to the traction, and switches between gaits so as to maximise efficiency. In this way, the robot only needs proprioceptive sensors and consequently does not require velocity estimation, ground imaging or profiling to maintain efficient locomotion across different terrains. The approach has been tested on a hexapod robot traversing a variety of terrain types and stiffness, including concrete, grass, mulch and leaf litter. The experimental results show that gait switching on energetics alone enables traction maintenance and efficient locomotion across different terrains. We also present comparisons between the power consumption metric used in this work and cost of transport which is used in the literature for characterising energetics for legged locomotion. Navinda Kottege, Callum Parkinson, Peyman Moghadam, Alberto Elfes, Surya P. N. Singh |
ICRA | 3 |
| 2014 | Combining motion and appearance for scene segmentationabstractImage segmentation is a key topic in computer vision, serving as a pre-step in a number of robotics tasks, including object recognition, obstacle avoidance and topological localization. In the literature, image segmentation has been employed as auxiliary information in order to improve optical flow performance. In this work, an alternative approach is proposed, in which optical flow information is used to aid image segmentation, aiming at scene understading for mobile robots. The proposed system performs dense optical flow analysis, followed by clustering of the optical flow vectors in a four dimensional space (formed by the x and y positions, angle and magnitude of each vector). Results from the clustering are used as ‘seeds’ in the segmentation process, performed by watershed segmentation in our implementation. In addition, the flow ‘image’ is combined with the original image, generating an image better suited for watershed segmentation, reducing the local minima effect often seen in this type of segmentation algorithms. The main pipeline considers the use of multi-modality cameras (visible and thermal-infrared). Since they see substantially different information, multi-modality further improves the amount of features of the resulting flows. Experimental results in urban and semi-urban scenarios with efficient segmentation illustrate the applicability of the method. Paulo Vinicius Koerich Borges, Peyman Moghadam |
ICRA | 2 |
| 2013 | Line-based extrinsic calibration of range and image sensorsabstractCreating rich representations of environments requires integration of multiple sensing modalities with complementary characteristics such as range and imaging sensors. To precisely combine multisensory information, the rigid transformation between different sensor coordinate systems (i.e., extrinsic parameters) must be estimated. The majority of existing extrinsic calibration techniques require one or multiple planar calibration patterns (such as checkerboards) to be observed simultaneously from the range and imaging sensors. The main limitation of these approaches is that they require modifying the scene with artificial targets. In this paper, we present a novel algorithm for extrinsically calibrating a range sensor with respect to an image sensor with no requirement of external artificial targets. The proposed method exploits natural linear features in the scene to precisely determine the rigid transformation between the coordinate frames. First, a set of 3D lines (plane intersection and boundary line segments) are extracted from the point cloud, and a set of 2D line segments are extracted from the image. Correspondences between the 3D and 2D line segments are used as inputs to an optimization problem which requires jointly estimating the relative translation and rotation between the coordinate frames. The proposed method is not limited to any particular types or configurations of sensors. To demonstrate robustness, efficiency and generality of the presented algorithm, we include results using various sensor configurations. Peyman Moghadam, Michael Bosse, Robert Zlot |
ICRA | 1 |
| 2013 | 3D thermal mapping of building interiors using an RGB-D and thermal cameraabstractThe building sector is the dominant consumer of energy and therefore a major contributor to anthropomorphic climate change. The rapid generation of photorealistic, 3D environment models with incorporated surface temperature data has the potential to improve thermographic monitoring of building energy efficiency. In pursuit of this goal, we propose a system which combines a range sensor with a thermal-infrared camera. Our proposed system can generate dense 3D models of environments with both appearance and temperature information, and is the first such system to be developed using a low-cost RGB-D camera. The proposed pipeline processes depth maps successively, forming an ongoing pose estimate of the depth camera and optimizing a voxel occupancy map. Voxels are assigned 4 channels representing estimates of their true RGB and thermal-infrared intensity values. Poses corresponding to each RGB and thermal-infrared image are estimated through a combination of timestamp-based interpolation and a predetermined knowledge of the extrinsic calibration of the system. Raycasting is then used to color the voxels to represent both visual appearance using RGB, and an estimate of the surface temperature. The output of the system is a dense 3D model which can simultaneously represent both RGB and thermal-infrared data using one of two alternative representation schemes. Experimental results demonstrate that the system is capable of accurately mapping difficult environments, even in complete darkness. Stephen Vidas, Peyman Moghadam, Michael Bosse |
ICRA | 2 |
| 2012 | Assessing the vulnerability of magnetic gestural authentication to video-based shoulder surfing attacksabstractSecure user authentication on mobile phones is crucial, as they store highly sensitive information. Common approaches to authenticate a user on a mobile phone are based either on entering a PIN, a password, or drawing a pattern. However, these authentication methods are vulnerable to the shoulder surfing attack. The risk of this attack has increased since means for recording high-resolution videos are cheaply and widely accessible. If the attacker can videotape the authentication process, PINs, passwords, and patterns do not even provide the most basic level of security. In this project, we assessed the vulnerability of a magnetic gestural authentication method to the video-based shoulder surfing attack. We chose a scenario that is favourable to the attack-er. In a real world environment, we videotaped the interactions of four users performing magnetic signatures on a phone, in the presence of HD cameras from four different angles. We then recruited 22 participants and asked them to watch the videos and try to forge the signatures. The results revealed that with a certain threshold, i.e, th=1.67, none of the forging attacks was successful, whereas at this level all eligible login attempts were successfully recognized. The qualitative feedback also indicated that users found the magnetic gestural signature authentication method to be more secure than PIN-based and 2D signature methods. Alireza Sahami Shirazi, Peyman Moghadam, Hamed Ketabdar, Albrecht Schmidt 0001 |
CHI | 2 |
| 2012 | Pingu: A New Miniature Wearable Device for Ubiquitous Computing EnvironmentsabstractAround Device Interaction (ADI) is recently introduced in the field of Human Computer Interaction (HCI)to provide touch less, more intuitive way of interaction using space beyond the physical boundary of the computing devices. In this paper, we introduce a new ADI input device called Pingu in the form factor of a fingering that allows users to interact with any nearby computing device with wireless connectivity in a ubiquitous environment. Fingering form factor is chosen for our prototype design, as it is socially acceptable and is commonly worn in everyday social contexts, and based on the previous research, the information entropy of interaction by fingers is greater than the entropy for any other parts of the human body. The current Pingu prototype is consisted of an extensive set of sensors, visual and vibrot actile feedback mechanisms with wireless connectivity that make it a unique input device for human-computer or human-human interaction in the form of gestures, tactile and touch. Its usage can range from advanced, tiny and novel gestural interaction with a variety of devices to mobile and networked sensing, and social computing. We present a few potential applications of Pingu such as social interaction, context recognition, in-car interaction, and physical activity analysis. Hamed Ketabdar, Peyman Moghadam, Mehran Roshandel |
CISIS | 2 |
| 2012 | Road direction detection based on vanishing-point trackingabstractWe present a novel approach for vision-based road direction detection for autonomous Unmanned Ground Vehicles (UGVs). The proposed method utilizes only monocular vision information similar to human perception to detect road directions with respect to the vehicle. The algorithm searches for a global feature of the roads due to perspective projection (so-called vanishing point) to distinguish road directions. The proposed approach consists of two stages. The first stage estimates the vanishing-point locations from single frames. The second stage uses a Rao-Blackwellised particle filter to track initial vanishing-point estimations over a sequence of images in order to provide more robust estimation. Simultaneously, the direction of the road ahead of the vehicle is predicted, which is prerequisite information for vehicle steering and path planning. The proposed approach assumes minimum prior knowledge about the environment and can cope with complex situations such as ground cover variations, different illuminations, and cast shadows. Its performance is evaluated on video sequences taken during test run of the DARPA Grand Challenge. Peyman Moghadam, Dong Jun Feng |
IROS | 1 |
| 2012 | Fast Vanishing-Point Detection in Unstructured EnvironmentsabstractVision-based road detection in unstructured environments is a challenging problem as there are hardly any discernible and invariant features that can characterize the road or its boundaries in such environments. However, a salient and consistent feature of most roads or tracks regardless of type of the environments is that their edges, boundaries, and even ruts and tire tracks left by previous vehicles on the path appear to converge into a single point known as the vanishing point. Hence, estimating this vanishing point plays a pivotal role in the determination of the direction of the road. In this paper, we propose a novel methodology based on image texture analysis for the fast estimation of the vanishing point in challenging and unstructured roads. The key attributes of the methodology consist of the optimal local dominant orientation method that uses joint activities of only four Gabor filters to precisely estimate the local dominant orientation at each pixel location in the image plane, the weighting of each pixel based on its dominant orientation, and an adaptive distance-based voting scheme for the estimation of the vanishing point. A series of quantitative and qualitative analyses are presented using natural data sets from the Defense Advanced Research Projects Agency Grand Challenge projects to demonstrate the effectiveness and the accuracy of the proposed methodology. Peyman Moghadam, Janusz A. Starzyk, W. Sardha Wijesoma |
IEEE Trans. Image Process. | 1 |
| 2010 | Towards a fully-autonomous vision-based vehicle navigation system in outdoor environmentsabstractColour Stereo visions are the primary perception system of the most Unmanned Ground Vehicles (UGVs), which can provide not only 3D perception of the terrain but also its colour and texture information. The downside with present stereo vision technologies and processing algorithms is that they are limited by the cameras' field of view and maximum range, which causes the vehicles to get caught in cul-de-sacs. The philosophy underlying the proposed framework in this paper is to use the near-field stereo vision information associated with the terrain appearance to train a classifier to classify the far-field terrain well beyond the stereo range for each incoming image. We propose an online, self-supervised learning method to learn far-field terrain traversability with the ability to adapt to unknown environments without using hand-labelled training data. The method described in this paper enhances current near-to-far learning techniques by automating the task of selecting which learning strategy to be used from among several strategies based on the nature of the incoming real-time input training data. Promising results obtained using real datasets from the DARPA-LAGR program is presented and the performance is evaluated using hand-labelled ground truth. Peyman Moghadam, W. Sardha Wijesoma, M. D. P. Moratuwage |
ICARCV | 1 |
| 2010 | Collaborative multi-vehicle localization and mapping in high clutter environmentsabstractAmong today's robotics applications, exploration missions in dynamic, high clutter and uncertain environmental conditions is quite common. Autonomous multi-vehicle systems come in handy for such exploration missions since a team of autonomous vehicles can explore an environment more efficiently and reliably than a single autonomous vehicle (AV). In order to improve the navigation accuracy, especially in the absence of a priori feature maps, various simultaneous localization and mapping (SLAM) algorithms are widely used in such applications. As for multi-vehicle scenarios, collaborative multi-vehicle simultaneous localization and mapping algorithm (CSLAM) is an effective strategy. However use of multiple AVs poses additional scaling problems such as inter-vehicle map fusion, and data association which needs to be addressed. Although existing CSLAM algorithms are shown to perform quite adequately in simulations, their performance is much less to be desired in high clutter scenarios that is inevitable in actual environments. In this paper, we present an approach to improve the performance of a CSLAM algorithm in the presence of high clutter, by combining an effective clutter filter framework based on Random Finite Sets (RFS). The performance of the improved CSLAM algorithm is evaluated using simulations under varying clutter conditions. M. D. P. Moratuwage, W. Sardha Wijesoma, Bharath Kalyan, Nicholas M. Patrikalakis, Peyman Moghadam |
ICARCV | 5 |
| 2009 | Online, Self-Supervised Vision-Based Terrain Classification in Unstructured EnvironmentsabstractOutdoor, unstructured and cross-country environments introduce several challenging problems such as highly complex scene geometry, ground cover variation, uncontrolled lighting, weather conditions and shadows for vision-based terrain classification of Unmanned Ground Vehicles (UGVs). Color stereo vision is mostly used for UGVs, but the present stereo vision technologies and processing algorithms are limited by cameras' field of view and maximum range, which causes the vehicles to get caught in cul-de-sacs that could possibly be avoided if the vehicle had access to information or could make inferences about the terrain well beyond the range of the vision system. The philosophy underlying the proposed strategy in this paper is to use the near-field stereo information associated with the terrain appearance to train a classifier to classify the far-field terrain well beyond the stereo range for each incoming image. To date, strategies based on this concept are limited to using single model construction and classification per frame. Although this single-model-per-frame approach can adapt to the changing environments concurrently, it lacks memory or history of past information. The approach described in this study is to use an online, self-supervised learning algorithm that exploits multiple frames to develop adaptive models that can classify different terrains the robot traverses. Preliminary but promising results of the paradigm proposed is presented using real data sets from the DARPA-LAGR project, which is the current gold standard for vision-based terrain classification using machine-learning techniques. This is followed by a proposal for future work on the development of robust terrain classifiers based on the proposed methodology. Peyman Moghadam, W. Sardha Wijesoma |
SMC | 1 |
| 2008 | Improving path planning and mapping based on stereo vision and lidarabstract2D laser range finders have been widely used in mobile robot navigation. However, their use is limited to simple environments containing objects of regular geometry and shapes. Stereo vision, instead, provides 3D structural data of complex objects. In this paper, measurements from a stereo vision camera system and a 2D laser range finder are fused to dynamically plan and navigate a mobile robot in cluttered and complex environments. A robust estimator is used to detect obstacles and ground plane in 3D world model in front of the robot based on disparity information from stereo vision system. Based on this 3D world model, 2D cost map is generated. A separate 2D cost map is also generated by 2D laser range finder. Then we use a grid-based occupancy map approach to fuse the complementary information provided by the 2D laser range finder and stereo vision system. Since the two sensors may detect different parts of an object, two different fusion strategies are addressed here. The final occupancy grid map is simultaneously used for obstacle avoidance and path planning. Experimental results obtained form a Point Grey's Bumblebee stereo camera and a SICK LDOEM laser range finder mounted on a Packbot robot are provided to demonstrate the effectiveness of the proposed lidar and stereo vision fusion strategy for mobile robot navigation. Peyman Moghadam, W. Sardha Wijesoma, Dong Jun Feng |
ICARCV | 1 |