Yasir Latif

dblp:123/5170 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0002-2529-5322ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 4 first-author · 7 since 2021Systems, architecture and hardware · 16 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Event-based Structure-from-Orbit
abstract
Event sensors offer high temporal resolution visual sensing, which makes them ideal for perceiving fast visual phe-nomena without suffering from motion blur. Certain applications in robotics and vision-based navigation require 3D perception of an object undergoing circular or spinning motion in front of a static camera, such as recovering the angular velocity and shape of the object. The setting is equiv-alent to observing a static object with an orbiting camera. In this paper, we propose event-based structure-from-orbit (eSfO), where the aim is to simultaneously reconstruct the 3D structure of a fast spinning object observed from a static event camera, and recover the equivalent orbital motion of the camera. Our contributions are threefold: since state-of-the-art event feature trackers cannot handle periodic self-occlusion due to the spinning motion, we develop a novel event feature tracker based on spatio-temporal clustering and data association that can better track the helical trajectories of valid features in the event data. The feature tracks are then fed to our novel factor graph-based structure-from-orbit back-end that calculates the orbital motion parameters (e.g., spin rate, relative rotational axis) that minimize the reprojection error. For evaluation, we produce a new event dataset of objects under spinning motion. Comparisons against ground truth indicate the efficacy of eSfO.
Ethan Elms, Yasir Latif, Tae Ha Park, Tat-Jun Chin
CVPR2
2024 Test-Time Certifiable Self-Supervision to Bridge the Sim2Real Gap in Event-Based Satellite Pose Estimation
abstract
Deep learning plays a critical role in vision-based satellite pose estimation. However, the scarcity of real data from the space environment means that deep models need to be trained using synthetic data, which raises the Sim2Real domain gap problem. A major cause of the Sim2Real gap are novel lighting conditions encountered during test time. Event sensors have been shown to provide some robustness against lighting variations in vision-based pose estimation. However, challenging lighting conditions due to strong directional light can still cause undesirable effects in the output of commercial off-the-shelf event sensors, such as noisy/spurious events and inhomogeneous event densities on the object. Such effects are non-trivial to simulate in software, thus leading to Sim2Real gap in the event domain. To close the Sim2Real gap in event-based satellite pose estimation, the paper proposes a test-time self-supervision scheme with a certifier module. Self-supervision is enabled by an optimisation routine that aligns a dense point cloud of the predicted satellite pose with the event data to attempt to rectify the inaccurately estimated pose. The certifier attempts to verify the corrected pose, and only certified test-time inputs are backpropagated via implicit differentiation to refine the predicted landmarks, thus improving the pose estimates and closing the Sim2Real gap. Results show that the our method outperforms established test-time adaptation schemes.
Abdul Mohsi Jawaid, Rajat Talak, Yasir Latif, Luca Carlone, Tat-Jun Chin
IROS3
2023 Towards Bridging the Space Domain Gap for Satellite Pose Estimation using Event Sensing
abstract
Deep models trained using synthetic data require domain adaptation to bridge the gap between the simulation and target environments. State-of-the-art domain adaptation methods often demand sufficient amounts of (unlabelled) data from the target domain. However, this need is difficult to fulfil when the target domain is an extreme environment, such as space. In this paper, our target problem is close proximity satellite pose estimation, where it is costly to obtain images of satellites from actual rendezvous missions. We demonstrate that event sensing offers a promising solution to generalise from the simulation to the target domain under stark illumination differences. Our main contribution is an event-based satellite pose estimation technique, trained purely on synthetic event data with basic data augmentation to improve robustness against practical (noisy) event sensors. Underpinning our method is a novel dataset with carefully calibrated ground truth, comprising of real event data obtained by emulating satellite rendezvous scenarios in the lab under drastic lighting conditions. Results on the dataset showed that our event-based satellite pose estimation method, trained only on synthetic data without adaptation, could generalise to the target domain effectively.
Abdul Mohsi Jawaid, Ethan Elms, Yasir Latif, Tat-Jun Chin
ICRA3
2023 MapFlow: latent transition via normalizing flow for unsupervised domain adaptation
abstract
Abstract Unsupervised domain adaptation (UDA) aims at enhancing the generalizability of the classification model learned from the labeled source domain to an unlabeled target domain. An established approach to UDA is to constrain the classifier on an intermediate representation that is distributionally invariant across domains. However, recent theoretical and empirical research has revealed that relying only on invariance fails to guarantee a small target error, thus making equality in the distribution of representations unnecessary. In this paper, we propose to relax invariant representation learning by finding a general relationship between the source and target representations, which allows an interchange of the more discriminative domain information. To this end, we formalize the MapFlow framework, which explicitly constructs an invertible mapping between the target encoded distribution and variationally induced source representation. Empirical results on public benchmark datasets show the desirable performance of our proposed algorithm compared to state-of-the-art methods.
Hossein Askari, Yasir Latif, Hongfu Sun
Mach. Learn.2
2022 Asynchronous Optimisation for Event-based Visual Odometry
abstract
Event cameras open up new possibilities for robotic perception due to their low latency and high dynamic range. On the other hand, developing effective event-based vision algorithms that fully exploit the beneficial properties of event cameras remains work in progress. In this paper, we focus on event-based visual odometry (VO). While existing event-driven VO pipelines have adopted continuous-time representations to asynchronously process event data, they either assume a known map, restrict the camera to planar trajectories, or integrate other sensors into the system. Towards map-free event-only monocular VO in SE(3), we propose an asynchronous structure-from-motion optimisation back-end. Our formulation is underpinned by a principled joint optimisation problem involving non-parametric Gaussian Process motion modelling and incremental maximum a posteriori inference. A high-performance incremental computation engine is employed to reason about the camera trajectory with every incoming event. We demonstrate the robustness of our asynchronous back-end in comparison to frame-based methods which depend on accurate temporal accumulation of measurements.
Daqi Liu, Álvaro Parra Bustos, Yasir Latif, Bo Chen 0009, Tat-Jun Chin, Ian D. Reid 0001
ICRA3
2021 Learning to Predict Repeatability of Interest Points
abstract
Many robotics applications require interest points that are highly repeatable under varying viewpoints and lighting conditions. However, this requirement is very challenging as the environment changes continuously and indefinitely, leading to appearance changes of interest points with respect to time. This paper proposes to predict the repeatability of an interest point as a function of time, which can tell us the lifespan of the interest point considering daily or seasonal variation. The repeatability predictor (RP) is formulated as a regressor trained on repeated interest points from multiple viewpoints over a long period of time. Through comprehensive experiments, we demonstrate that our RP can estimate when a new interest point is repeated, and also highlight an insightful analysis about this problem. For further comparison, we apply our RP to the map summarization under visual localization framework, which builds a compact representation of the full context map given the query time. The experimental result shows a careful selection of potentially repeatable interest points predicted by our RP can significantly mitigate the degeneration of localization accuracy from map summarization.
Anh-Dzung Doan, Daniyar Turmukhambetov, Yasir Latif, Tat-Jun Chin, Soohyun Bae
ICRA3
2021 Visual localization under appearance change: filtering approaches
Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin, Yu Liu 0029, Shin-Fang Ch'ng, Thanh-Toan Do, Ian D. Reid 0001
Neural Comput. Appl.2
2020 SPRINT: Subgraph Place Recognition for INtelligent Transportation
abstract
Visual place recognition is an important problem in mobile robotics which aims to localize a robot using image information alone. Recent methods have shown promising results for place recognition under varying environmental conditions by exploiting the sequential nature of the image acquisition process. We show that by using k nearest neighbours based image retrieval as the backend, and exploiting the structure of the image acquisition process which introduces temporal relations between images in the database, the location of possible matches can be restricted to a subset of all the images seen so far. In effect, the original problem space can thus be restricted to a significantly smaller subspace, reducing the inference time significantly. This is particularly important for scalable place recognition over databases containing millions of images. We present large scale experiments using publicly sourced data that show the computational performance of the proposed method under varying environmental conditions.
Yasir Latif, Anh-Dzung Doan, Tat-Jun Chin, Ian D. Reid 0001
ICRA1
2019 Scalable Place Recognition Under Appearance Change for Autonomous Driving
abstract
A major challenge in place recognition for autonomous driving is to be robust against appearance changes due to short-term (e.g., weather, lighting) and long-term (seasons, vegetation growth, etc.) environmental variations. A promising solution is to continuously accumulate images to maintain an adequate sample of the conditions and incorporate new changes into the place recognition decision. However, this demands a place recognition technique that is scalable on an ever growing dataset. To this end, we propose a novel place recognition technique that can be efficiently retrained and compressed, such that the recognition of new queries can exploit all available data (including recent changes) without suffering from visible growth in computational cost. Underpinning our method is a novel temporal image matching technique based on Hidden Markov Models. Our experiments show that, compared to state-of-the-art techniques, our method has much greater potential for large-scale place recognition for autonomous driving.
Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin, Yu Liu 0029, Thanh-Toan Do, Ian D. Reid 0001
ICCV2
2019 Real-Time Monocular Object-Model Aware Sparse SLAM
abstract
Simultaneous Localization And Mapping (SLAM) is a fundamental problem in mobile robotics. While sparse point-based SLAM methods provide accurate camera localization, the generated maps lack semantic information. On the other hand, state of the art object detection methods provide rich information about entities present in the scene from a single image. This work incorporates a real-time deep-learned object detector to the monocular SLAM framework for representing generic objects as quadrics that permit detections to be seamlessly integrated while allowing the real-time performance. Finer reconstruction of an object, learned by a CNN network, is also incorporated and provides a shape prior for the quadric leading further refinement. To capture the structure of the scene, additional planar landmarks are detected by a CNN-based plane detector and modelled as independent landmarks in the map. Extensive experiments support our proposed inclusion of semantic objects and planar structures directly in the bundle-adjustment of SLAM - Semantic SLAM- that enriches the reconstructed map semantically, while significantly improving the camera localization.
Mehdi Hosseinzadeh 0003, Kejie Li, Yasir Latif, Ian D. Reid 0001
ICRA3
2018 Structure Aware SLAM Using Quadrics and Planes
Mehdi Hosseinzadeh 0003, Yasir Latif, Trung Pham, Niko Sünderhauf, Ian D. Reid 0001
ACCV (3)2
2018 Learning Deeply Supervised Good Features to Match for Dense Monocular Reconstruction
Chamara Saroj Weerasekera, Ravi Garg, Yasir Latif, Ian D. Reid 0001
ACCV (5)3
2018 Addressing Challenging Place Recognition Tasks Using Generative Adversarial Networks
abstract
Place recognition is an essential component of Simultaneous Localization And Mapping (SLAM). Under severe appearance change, reliable place recognition is a difficult perception task since the same place is perceptually very different in the morning, at night, or over different seasons. This work addresses place recognition as a domain translation task. Using a pair of coupled Generative Adversarial Networks (GANs), we show that it is possible to generate the appearance of one domain (such as summer) from another (such as winter) without requiring image-to-image correspondences across the domains. Mapping between domains is learned from sets of images in each domain without knowing the instance-to-instance correspondence by enforcing a cyclic consistency constraint. In the process, meaningful feature spaces are learned for each domain, the distances in which can be used for the task of place recognition. Experiments show that learned features correspond to visual similarity and can be effectively used for place recognition across seasons.
Yasir Latif, Ravi Garg, Michael Milford, Ian D. Reid 0001
ICRA1
2017 RRD-SLAM: Radial-distorted rolling-shutter direct SLAM
abstract
In this paper, we present a monocular direct semi-dense SLAM (Simultaneous Localization And Mapping) method that can handle both radial distortion and rolling-shutter distortion. Such distortions are common in, but not restricted to, situations when an inexpensive wide-angle lens and a CMOS sensor are used, and leads to significant inaccuracy in the map and trajectory estimates if not modeled correctly. The apparent naive solution of simply undistorting the images using pre-calibrated parameters does not apply to this case since rows in the undistorted image are no longer captured at the same time. To address this we develop an algorithm that incorporates radial distortion into an existing state-of-the-art direct semi-dense SLAM system that takes rolling-shutters into account. We propose a method for finding the generalized epipolar curve for each rolling-shutter radially distorted image. Our experiments demonstrate the efficacy of our approach and compare it favorably with the state-of-the-art in direct semi-dense rolling-shutter SLAM.
Jae-Hak Kim, Yasir Latif, Ian D. Reid 0001
ICRA2
2017 Dense monocular reconstruction using surface normals
abstract
This paper presents an efficient framework for dense 3D scene reconstruction using input from a moving monocular camera. Visual SLAM (Simultaneous Localisation and Mapping) approaches based solely on geometric methods have proven to be quite capable of accurately tracking the pose of a moving camera and simultaneously building a map of the environment in real-time. However, most of them suffer from the 3D map being too sparse for practical use. The missing points in the generated map correspond mainly to areas lacking texture in the input images, and dense mapping systems often rely on hand-crafted priors like piecewise-planarity or piecewise-smooth depth. These priors do not always provide the required level of scene understanding to accurately fill the map. On the other hand, Convolutional Neural Networks (CNNs) have had great success in extracting high-level information from images and regressing pixel-wise surface normals, semantics, and even depth. In this work we leverage this high-level scene context learned by a deep CNN in the form of a surface normal prior. We show, in particular, that using the surface normal prior leads to better reconstructions than the weaker smoothness prior.
Chamara Saroj Weerasekera, Yasir Latif, Ravi Garg, Ian D. Reid 0001
ICRA2
2017 Meaningful maps with object-oriented semantic mapping
abstract
For intelligent robots to interact in meaningful ways with their environment, they must understand both the geometric and semantic properties of the scene surrounding them. The majority of research to date has addressed these mapping challenges separately, focusing on either geometric or semantic mapping. In this paper we address the problem of building environmental maps that include both semantically meaningful, object-level entities and point- or mesh-based geometrical representations. We simultaneously build geometric point cloud models of previously unseen instances of known object classes and create a map that contains these object models as central entities. Our system leverages sparse, feature-based RGB-D SLAM, image-based deep-learning object detection and 3D unsupervised segmentation.
Niko Sünderhauf, Trung T. Pham, Yasir Latif, Michael Milford, Ian D. Reid 0001
IROS3
2016 Measuring the performance of single image depth estimation methods
abstract
We consider the question of benchmarking the performance of methods used for estimating the depth of a scene from a single image. We describe various measures that have been used in the past, discuss their limitations and demonstrate that each is deficient in one or more ways. We propose a new measure of performance for depth estimation that overcomes these deficiencies, and has a number of desirable properties. We show that in various cases of interest the new measure enables visualisation of the performance of a method that is otherwise obfuscated by existing metrics. Our proposed method is capable of illuminating the relative performance of different algorithms on different kinds of data, such as the difference in efficacy of a method when estimating the depth of the ground plane versus estimating the depth of other generic scene structure. We showcase the method by comparing a number of existing single-view methods against each other and against more traditional depth estimation methods such as binocular stereo.
Cesar Dario Cadena Lerma, Yasir Latif, Ian D. Reid 0001
IROS2
2016 Past, Present, and Future of Simultaneous Localization and Mapping: Toward the Robust-Perception Age
abstract
Simultaneous localization and mapping (SLAM) consists in the concurrent construction of a model of the environment (the map), and the estimation of the state of the robot moving within it. The SLAM community has made astonishing progress over the last 30 years, enabling large-scale real-world applications and witnessing a steady transition of this technology to industry. We survey the current state of SLAM and consider future directions. We start by presenting what is now the de-facto standard formulation for SLAM. We then review related work, covering a broad set of topics including robustness and scalability in long-term mapping, metric and semantic representations for mapping, theoretical performance guarantees, active SLAM and exploration, and other new frontiers. This paper simultaneously serves as a position paper and tutorial to those who are users of SLAM. By looking at the published research with a critical eye, we delineate open challenges and new research issues, that still deserve careful scientific investigation. The paper also contains the authors' take on two questions that often animate discussions during robotics conferences: Do robots need SLAM? and Is SLAM solved?
Cesar Dario Cadena Lerma, Luca Carlone, Henry Carrillo, Yasir Latif, Davide Scaramuzza 0001, José Neira, Ian D. Reid 0001, John J. Leonard
IEEE Trans. Robotics4
2015 Hierarchical Higher-Order Regression Forest Fields: An Application to 3D Indoor Scene Labelling
abstract
This paper addresses the problem of semantic segmentation of 3D indoor scenes reconstructed from RGB-D images. Traditionally label prediction for 3D points is tackled by employing graphical models that capture scene features and complex relations between different class labels. However, the existing work is restricted to pairwise conditional random fields, which are insufficient when encoding rich scene context. In this work we propose models with higher-order potentials to describe complex relational information from the 3D scenes. Specifically, we relax the labelling problem to a regression, and generalize the higher-order associative P n Potts model to a new family of arbitrary higher-order models based on regression forests. We show that these models, like the robust P n models, can still be decomposed into the sum of pairwise terms by introducing auxiliary variables. Moreover, our proposed higher-order models also permit extension to hierarchical random fields, which allows for the integration of scene context and features computed at different scales. Our potential functions are constructed based on regression forests encoding Gaussian densities that admit efficient inference. The parameters of our model are learned from training data using a structured learning approach. Results on two datasets show clear improvements over current state-of-the-art methods.
Trung-Thanh Pham, Ian D. Reid 0001, Yasir Latif, Stephen Gould
ICCV3
2015 On the monotonicity of optimality criteria during exploration in active SLAM
abstract
In this paper we investigate the monotonicity of various optimality criteria during the exploration phase of an active SLAM algorithm. Optimality criteria such as A-opt, D-opt or E-opt are used in active SLAM to account for uncertainty in the map or the robot's pose, and these criteria are usually part of utility functions which help active SLAM algorithms decide where the robot should move next. The monotonicity of the optimality criteria is of utmost importance. During the exploration phase, i.e. when the robot is traversing new territory or cannot perform a loop closure, the most common way of estimating the pose of the robot is through dead-reckoning. Correctly accounting for the uncertainty is important for an active SLAM algorithm and in particular for a dead-reckoning scenario, where by definition the uncertainty in the robot's pose grows. If monotonicity does not hold in this scenario, active SLAM algorithms can execute actions under the false belief that the uncertainty has reduced. We show analytically and experimentally some conditions in which the A-opt and E-opt criteria lose monotonicity in a dead-reckoning scenario, where the propagation of the robot's pose is done using a linearized framework. We also show analytically and experimentally that under the same conditions the D-opt does not lose monotonicity and, in general for the linearized framework under consideration, D-opt does not break monotonicity.
Henry Carrillo, Yasir Latif, María Luisa Rodríguez-Arévalo, José Neira, José A. Castellanos 0001
ICRA2
2014 Place categorization using sparse and redundant representations
abstract
Place categorization addresses the problem of determining the semantic label of the current position of a robot, given a snapshot of the environment as well as previously labeled information about different places that the robot has already seen. State-of-the-art approaches use machine learning techniques that require extensive and often time consuming training. This work proposes a novel formulation by posing place categorization as an efficient ℓ1-minimization problem, leading to both a faster training phase and to performance comparable to state-of-the-art methods. The formulation allows online robot operation particularly in the case when the training phase has to be learned on-the-fly and in an active manner. To validate the performance of the proposed method, extensive experimental results carried out on real data under different lighting conditions as well as structural changes in the environment are provided.
Henry Carrillo, Yasir Latif, José Neira, José A. Castellanos 0001
IROS2
2014 Robust graph SLAM back-ends: A comparative analysis
abstract
In this work, we provide an in-depth analysis of several recent robust Simultaneous Localization And Mapping (SLAM) back-end techniques that aim to recover the correct graph estimate in the presence of outliers in loop closure constraints. We present a benchmark dataset for evaluation of such methods by augmenting the KITTI Vision Benchmark with ground truth as well as generated loop closure hypotheses and present a detailed analysis of recently proposed robust SLAM methods using this benchmark. We also look into how these methods achieve the desired robustness and what are the implications for the SLAM problem. We discuss the issues involved in using the output of these robust back-ends for tasks such as path planning and how they can be addressed. The problem of robustness needs to be addressed adequately in order to have a complete and reliable solution to the SLAM problem.
Yasir Latif, Cesar Dario Cadena Lerma, José Neira
IROS1
2013 Speeded-up SURF: Design of an efficient multiscale feature detector
abstract
We present a fast and highly performant multiscale feature detector which is based on the established SURF algorithm. The additional speed-up is achieved by linearizing the SURF detector in a way that its detection characteristics are preserved. Our evaluations show that the proposed suSURF detector is roughly 30% faster without significant sacrifices in feature quality. The key points detected by suSURF are compatible with standard SURF features and those found by other determinant of Hessian (DoH) based detectors.
Florian Schweiger, Georg Schroth, Robert Huitl, Yasir Latif, Eckehard G. Steinbach
ICIP4
2012 Fast minimum uncertainty search on a graph map representation
abstract
This paper addresses the problem of path planning considering uncertainty criteria over the belief space. Specifically, we propose a path planning algorithm that uses a novel determinant-based measure of uncertainty and a reduced representation of the environment, in order to obtain the minimum uncertainty path from a roadmap. Our proposal does not require a priori knowledge of the environment due to the construction of the roadmap via a graph-based SLAM algorithm. We report experimental results of our proposal in four datasets that show its feasibility to obtain the minimum uncertainty path towards an autonomous navigation framework and we also show an improvement in the computation time with respect to the state of the art.
Henry Carrillo, Yasir Latif, José Neira, José A. Castellanos 0001
IROS2
2012 Realizing, reversing, recovering: Incremental robust loop closing over time using the iRRR algorithm
abstract
The ability to reconsider information over time allows to detect failures and is crucial for long term robust autonomous robot applications. This applies to loop closure decisions in localization and mapping systems. This paper describes a method to analyze all available information up to date in order to robustly remove past incorrect loop closures from the optimization process. The main novelties of our algorithm are: 1. incrementally reconsidering loop closures and 2. handling multi-session, spatially related or unrelated experiments. We validate our proposal in real multi-session experiments showing better results than those obtained by state of the art methods.
Yasir Latif, Cesar Dario Cadena Lerma, José Neira
IROS1