Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ben Harwood

dblp:191/4628 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot navigation and mapping · 60% Representation and self-supervised learning · 17% 3D vision · 15%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping › place recognition
LiDAR-based place recognition
0.512021
Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order Pooling · ICRA 2021
Robotics › Robot navigation and mapping › SLAM
loop closure detection
0.512021
Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order Pooling · ICRA 2021
Robotics › Robot navigation and mapping
place recognition
0.512021
Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order Pooling · ICRA 2021
Robotics › Robot navigation and mapping
SLAM
0.512021
Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order Pooling · ICRA 2021
Machine learning › Representation and self-supervised learning › representation learning › metric learning
deep metric learning
0.312017
Smart Mining for Deep Metric Learning · ICCV 2017
Machine learning › Deep learning architectures and training
hard example mining
0.312017
Smart Mining for Deep Metric Learning · ICCV 2017
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.312017
Smart Mining for Deep Metric Learning · ICCV 2017
Computer vision › 3D vision
approximate nearest neighbor
0.212016
FANNG: Fast Approximate Nearest Neighbour Graphs · CVPR 2016
Computer vision › 3D vision
feature matching
0.212016
FANNG: Fast Approximate Nearest Neighbour Graphs · CVPR 2016
Algorithms and data structures › similarity search › nearest neighbor search
approximate nearest neighbor search
0.212016
FANNG: Fast Approximate Nearest Neighbour Graphs · CVPR 2016
Algorithms and data structures › similarity search
nearest neighbor search
0.212016
FANNG: Fast Approximate Nearest Neighbour Graphs · CVPR 2016

Methods — techniques the papers use, named apart from their topics

spatiotemporal encoding · 0.5higher-order pooling · 0.5graph-based search · 0.5global descriptor learning · 0.5GPU implementation · 0.5triplet network · 0.3smart mining · 0.3adaptive controller · 0.3
YearPublicationVenuePosition
2021 Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order Pooling
abstract
Place Recognition enables the estimation of a globally consistent map and trajectory by providing non-local constraints in Simultaneous Localisation and Mapping (SLAM). This paper presents Locus, a novel place recognition method using 3D LiDAR point clouds in large-scale environments. We propose a method for extracting and encoding topological and temporal information related to components in a scene and demonstrate how the inclusion of this auxiliary information in place description leads to more robust and discriminative scene representations. Second-order pooling along with a non- linear transform is used to aggregate these multi-level features to generate a fixed-length global descriptor, which is invariant to the permutation of input features. The proposed method outperforms state-of-the-art methods on the KITTI dataset. Furthermore, Locus is demonstrated to be robust across several challenging situations such as occlusions and viewpoint changes in 3D LiDAR point clouds. The open-source implementation is available at: https://github.com/csiro-robotics/locus.
Kavisha Vidanapathirana, Peyman Moghadam, Ben Harwood, Muming Zhao, Sridha Sridharan, Clinton Fookes
ICRA3
2020 Temporally Coherent Embeddings for Self-Supervised Video Representation Learning
abstract
This paper presents TCE: Temporally Coherent Embeddings for self-supervised video representation learning. The proposed method exploits inherent structure of unlabeled video data to explicitly enforce temporal coherency in the embedding space, rather than indirectly learning it through ranking or predictive proxy tasks. In the same way that high-level visual information in the world changes smoothly, we believe that nearby frames in learned representations will benefit from demonstrating similar properties. Using this assumption, we train our TCE model to encode videos such that adjacent frames exist close to each other and videos are separated from one another. Using TCE we learn robust representations from large quantities of unlabeled video data. We thoroughly analyse and evaluate our self-supervised learned TCE models on a downstream task of video action recognition using multiple challenging benchmarks (Kinetics400, UCF101, HMDB51). With a simple but effective 2D-CNN backbone and only RGB stream inputs, TCE pre-trained representations outperform all previous self-supervised 2D-CNN and 3D-CNN pre-trained on UCF101. The code and pre-trained models for this paper can be downloaded at: https://github.com/csiro-robotics/TCE.
Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, Peyman Moghadam
ICPR2
2018 Deep Metric Learning and Image Classification with Nearest Neighbour Gaussian Kernels
abstract
We present a Gaussian kernel loss function and training algorithm for convolutional neural networks that can be directly applied to both distance metric learning and image classification problems. Our method treats all training features from a deep neural network as Gaussian kernel centres and computes loss by summing the influence of a feature's nearby centres in the feature embedding space. Our approach is made scalable by treating it as an approximate nearest neighbour search problem. We show how to make end-to-end learning feasible, resulting in a well formed embedding space, in which semantically related instances are likely to be located near one another, regardless of whether or not the network was trained on those classes. Our approach outperforms state-of-the-art deep metric learning approaches on embedding learning challenges, as well as conventional softmax classification on several datasets.
Benjamin J. Meyer 0001, Ben Harwood, Tom Drummond
ICIP2
2017 Smart Mining for Deep Metric Learning
abstract
To solve deep metric learning problems and producing feature embeddings, current methodologies will commonly use a triplet model to minimise the relative distance between samples from the same class and maximise the relative distance between samples from different classes. Though successful, the training convergence of this triplet model can be compromised by the fact that the vast majority of the training samples will produce gradients with magnitudes that are close to zero. This issue has motivated the development of methods that explore the global structure of the embedding and other methods that explore hard negative/positive mining. The effectiveness of such mining methods is often associated with intractable computational requirements. In this paper, we propose a novel deep metric learning method that combines the triplet model and the global structure of the embedding space. We rely on a smart mining procedure that produces effective training samples for a low computational cost. In addition, we propose an adaptive controller that automatically adjusts the smart mining hyper-parameters and speeds up the convergence of the training process. We show empirically that our proposed method allows for fast and more accurate training of triplet ConvNets than other competing mining methods. Additionally, we show that our method achieves new state-of-the-art embedding results for CUB-200-2011 and Cars196 datasets.
Ben Harwood, Gustavo Carneiro 0001, Ian D. Reid 0001, Tom Drummond
ICCV1
2016 FANNG: Fast Approximate Nearest Neighbour Graphs
abstract
We present a new method for approximate nearest neighbour search on large datasets of high dimensional feature vectors, such as SIFT or GIST descriptors. Our approach constructs a directed graph that can be efficiently explored for nearest neighbour queries. Each vertex in this graph represents a feature vector from the dataset being searched. The directed edges are computed by exploiting the fact that, for these datasets, the intrinsic dimensionality of the local manifold-like structure formed by the elements of the dataset is significantly lower than the embedding space. We also provide an efficient search algorithm that uses this graph to rapidly find the nearest neighbour to a query with high probability. We show how the method can be adapted to give a strong guarantee of 100% recall where the query is within a threshold distance of its nearest neighbour. We demonstrate that our method is significantly more efficient than existing state of the art methods. In particular, our GPU implementation can deliver 90% recall for queries on a data set of 1 million SIFT descriptors at a rate of over 1.2 million queries per second on a Titan X. Finally we also demonstrate how our method scales to datasets of 5M and 20M entries.
Ben Harwood, Tom Drummond
CVPR1