VLDB 2026 Research / reviewers in the wild / expert
Edward Johns
dblp:68/9968 · also Edward David Johns
· DBLP profile ↗
38ranked-venue papers
12as first author
16since 2021 · last 2025
0000-0002-8914-8786ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 12 first-author · 16 since 2021Systems, architecture and hardware · 22 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Instant Policy: In-Context Imitation Learning via Graph DiffusionabstractFollowing the impressive capabilities of in-context learning with large transformers, In-Context Imitation Learning (ICIL) is a promising opportunity for robotics. We introduce Instant Policy, which learns new tasks instantly from just one or two demonstrations, achieving ICIL through two key components. First, we introduce inductive biases through a graph representation and model ICIL as a graph generation problem using a learned diffusion process, enabling structured reasoning over demonstrations, observations, and actions. Second, we show that such a model can be trained using pseudo-demonstrations – arbitrary trajectories generated in simulation – as a virtually infinite pool of training data. Our experiments, in both simulation and reality, show that Instant Policy enables rapid learning of various everyday robot tasks. We also show how it can serve as a foundation for cross-embodiment and zero-shot transfer to language-defined tasks. Vitalis Vosylius, Edward Johns |
ICLR | 2 |
| 2025 | R+X: Retrieval and Execution from Everyday Human VideosabstractWe present$\mathbf{R}+\mathbf{X}$, a framework which enables robots to learn skills from long, unlabelled, first-person videos of humans performing everyday tasks. Given a language command from a human,$\mathbf{R}+\mathbf{X}$first retrieves short video clips containing relevant behaviour, and then executes the skill by conditioning an in-context imitation learning method (KAT) on this behaviour. By leveraging a Vision Language Model (VLM) for retrieval,$\mathbf{R}+\mathbf{X}$does not require any manual annotation of the videos, and by leveraging in-context learning for execution, robots can perform commanded skills immediately, without requiring a period of training on the retrieved videos. Experiments studying a range of everyday household tasks show that$\mathbf{R}+\mathbf{X}$succeeds at translating unlabelled human videos into robust robot skills, and that$\mathbf{R}+\mathbf{X}$outperforms several recent alternative methods. Appendix and videos are available at https://www.robot-learning.uk/r-plus-x. Georgios Papagiannis, Norman Di Palo, Pietro Vitiello, Edward Johns |
ICRA | 4 |
| 2025 | One-Shot Dual-Arm Imitation LearningabstractWe introduce One-Shot Dual-Arm Imitation Learning (ODIL), which enables dual-arm robots to learn precise and coordinated everyday tasks from just a single demonstration of the task. ODIL uses a new three-stage visual servoing (3-VS) method for precise alignment between the end-effector and target object, after which replay of the demonstration trajectory is sufficient to perform the task. This is achieved without requiring prior task or object knowledge, or additional data collection and training following the single demonstration. Furthermore, we propose a new dual-arm coordination paradigm for learning dual-arm tasks from a single demonstration. ODIL was tested on a real-world dual-arm robot, demonstrating state-of-the-art performance across six precise and coordinated tasks in both 4-DoF and 6-DoF settings, and showing robustness in the presence of distractor objects and partial occlusions. Videos are available at: https://www.robot-learning.uk/one-shot-dual-arm. Edward Johns |
ICRA | 2 |
| 2025 | Neural Stochastic Flows: Solver-Free Modelling and Inference for SDE SolutionsabstractStochastic differential equations (SDEs) are well suited to modelling noisy and/or irregularly-sampled time series, which are omnipresent in finance, physics, and machine learning applications. Traditional approaches require costly simulation of numerical solvers when sampling between arbitrary time points. We introduce Neural Stochastic Flows (NSFs) and their latent dynamic versions, which learns (latent) SDE transition laws directly using conditional normalising flows, with architectural constraints that preserve properties inherited from stochastic flow. This enables sampling between arbitrary states in a single step, providing up to two orders of magnitude speedup for distant time points. Experiments on synthetic SDE simulations and real-world tracking and video data demonstrate that NSF maintains distributional accuracy comparable to numerical approaches while dramatically reducing computation for arbitrary time-point sampling, enabling applications where numerical solvers remain prohibitively expensive. Naoki Kiyohara, Edward Johns, Yingzhen Li |
NeurIPS | 2 |
| 2024 | Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language ModelsabstractWe introduce Dream2Real, a robotics framework which integrates vision-language models (VLMs) trained on 2D data into a 3D object rearrangement pipeline. This is achieved by the robot autonomously constructing a 3D representation of the scene, where objects can be rearranged virtually and an image of the resulting arrangement rendered. These renders are evaluated by a VLM, so that the arrangement which best satisfies the user instruction is selected and recreated in the real world with pick-and-place. This enables language-conditioned rearrangement to be performed zero-shot, without needing to collect a training dataset of example arrangements. Results on a series of real-world tasks show that this framework is robust to distractors, controllable by language, capable of understanding complex multi-object relations, and readily applicable to both tabletop and 6-DoF rearrangement tasks. Videos are available on our webpage at: https://www.robot-learning.uk/dream2real. Ivan Kapelyukh, Ignacio Alzugaray, Edward Johns |
ICRA | 4 |
| 2024 | Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment CollaborationabstractLarge, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io. Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin |
ICRA | 58 |
| 2024 | DINOBot: Robot Manipulation via Retrieval and Alignment with Vision Foundation ModelsabstractWe propose DINOBot, a novel imitation learning framework for robot manipulation, which leverages the image-level and pixel-level capabilities of features extracted from Vision Transformers trained with DINO. When interacting with a novel object, DINOBot first uses these features to retrieve the most visually similar object experienced during human demonstrations, and then uses this object to align its endeffector with the novel object to enable effective interaction. Through a series of real-world experiments on everyday tasks, we show that exploiting both the image-level and pixel-level properties of vision foundation models enables unprecedented learning efficiency and generalisation. Videos and code are available at https://www.robot-learning.uk/dinobot. Norman Di Palo, Edward Johns |
ICRA | 2 |
| 2024 | Adapting Skills to Novel Grasps: A Self-Supervised ApproachabstractIn this paper, we study the problem of adapting manipulation trajectories involving grasped objects (e.g. tools) defined for a single grasp pose to novel grasp poses. A common approach to address this is to define a new trajectory for each possible grasp explicitly, but this is highly inefficient. Instead, we propose a method to adapt such trajectories directly while only requiring a period of self-supervised data collection, during which a camera observes the robot’s end-effector moving with the object rigidly grasped. Importantly, our method requires no prior knowledge of the grasped object (such as a 3D CAD model), it can work with RGB images, depth images, or both, and it requires no camera calibration. Through a series of real-world experiments involving 1360 evaluations, we find that self-supervised RGB data consistently outperforms alternatives that rely on depth images including several state-of-the-art pose estimation methods. Compared to the best-performing baseline, our method results in an average of 28.5% higher success rate when adapting manipulation trajectories to novel grasps on several everyday tasks. The appendix accompanying the paper and videos of the experiments are available on our webpage at www.robot-learning.uk/adapting-skills. Georgios Papagiannis, Kamil Dreczkowski, Vitalis Vosylius, Edward Johns |
IROS | 4 |
| 2023 | Learning Tethered Perching for Aerial RobotsabstractAerial robots have a wide range of applications, such as collecting data in hard-to-reach areas. This requires the longest possible operation time. However, because currently available commercial batteries have limited specific energy of roughly 300 W h kg-1, a drone's flight time is a bottleneck for sustainable long-term data collection. Inspired by birds in nature, a possible approach to tackle this challenge is to perch drones on trees, and environmental or man-made structures, to save energy whilst in operation. In this paper, we propose an algorithm to automatically generate trajectories for a drone to perch on a tree branch, using the proposed tethered perching mechanism with a pendulum-like structure. This enables a drone to perform an energy-optimised, controlled 180° flip to safely disarm upside down. To fine-tune a set of reachable trajectories, a soft actor critic-based reinforcement algorithm is used. Our experimental results show the feasibility of the set of trajectories with successful perching. Our findings demonstrate that the proposed approach enables energy-efficient landing for long-term data collection tasks. Fabian Hauf, Basaran Bahadir Kocer, Alan Slatter, Hai-Nguyen Nguyen, Oscar Pang, Ronald Clark, Edward Johns, Mirko Kovac |
ICRA | 7 |
| 2022 | Bootstrapping Semantic Segmentation with Regional Contrast
Shikun Liu, Shuaifeng Zhi, Edward Johns, Andrew J. Davison |
ICLR | 3 |
| 2022 | Demonstrate Once, Imitate Immediately (DOME): Learning Visual Servoing for One-Shot Imitation LearningabstractWe present DOME, a novel method for one-shot imitation learning, where a task can be learned from just a single demonstration and then be deployed immediately, without any further data collection or training. DOME does not require prior task or object knowledge, and can perform the task in novel object configurations and with distractors. At its core, DOME uses an image-conditioned object segmentation network followed by a learned visual servoing network, to move the robot's end-effector to the same relative pose to the object as during the demonstration, after which the task can be completed by replaying the demonstration's end-effector velocities. We show that DOME achieves near 100% success rate on 7 real-world everyday tasks, and we perform several studies to thoroughly understand each individual component of DOME. Videos and supplementary material are available at: https://www.robot-learning.uk/dome. Eugene Valassakis, Georgios Papagiannis, Norman Di Palo, Edward Johns |
IROS | 4 |
| 2022 | AGO-Net: Association-Guided 3D Point Cloud Object Detection NetworkabstractThe human brain can effortlessly recognize and localize objects, whereas current 3D object detection methods based on LiDAR point clouds still report inferior performance for detecting occluded and distant objects: The point cloud appearance varies greatly due to occlusion, and has inherent variance in point densities along the distance to sensors. Therefore, designing feature representations robust to such point clouds is critical. Inspired by human associative recognition, we propose a novel 3D detection framework that associates intact features for objects via domain adaptation. We bridge the gap between the perceptual domain, where features are derived from real scenes with sub-optimal representations, and the conceptual domain, where features are extracted from augmented scenes that consist of non-occlusion objects with rich detailed information. A feasible method is investigated to construct conceptual scenes without external datasets. We further introduce an attention-based re-weighting module that adaptively strengthens the feature adaptation of more informative regions. The network's feature enhancement ability is exploited without introducing extra cost during inference, which is plug-and-play in various 3D detection frameworks. We achieve new state-of-the-art performance on the KITTI 3D detection benchmark in both accuracy and speed. Experiments on nuScenes and Waymo datasets also validate the versatility of our method. Liang Du 0004, Xiaoqing Ye, Xiao Tan 0001, Edward Johns, Errui Ding, Xiangyang Xue 0001, Jianfeng Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Benchmarking Domain Randomisation for Visual Sim-to-Real TransferabstractDomain randomisation is a very popular method for visual sim-to-real transfer in robotics, due to its simplicity and ability to achieve transfer without any real-world images at all. Nonetheless, a number of design choices must be made to achieve optimal transfer. In this paper, we perform a comprehensive benchmarking study on these different choices, with two key experiments evaluated on a real-world object pose estimation task. First, we study the rendering quality, and nd that a small number of high-quality images is superior to a large number of low-quality images. Second, we study the type of randomisation, and nd that both distractors and textures are important for generalisation to novel environments. Raghad Alghonaim, Edward Johns |
ICRA | 2 |
| 2021 | Coarse-to-Fine Imitation Learning: Robot Manipulation from a Single DemonstrationabstractWe introduce a simple new method for visual imitation learning, which allows a novel robot manipulation task to be learned from a single human demonstration, without requiring any prior knowledge of the object being interacted with. Our method models imitation learning as a state estimation problem, with the state defined as the end-effector’s pose at the point where object interaction begins, as observed from the demonstration. By modelling a manipulation task as a coarse, approach trajectory followed by a fine, interaction trajectory, this state estimator can be trained in a self-supervised manner, by automatically moving the end-effector’s camera around the object. At test time, the end-effector is moved to the estimated state through a linear path, at which point the demonstration’s end-effector velocities are simply repeated, enabling convenient acquisition of a complex interaction trajectory without actually needing to explicitly learn a policy. Real-world experiments on 8 everyday tasks show that our method can learn a diverse range of skills from just a single human demonstration, whilst also yielding a stable and interpretable controller. Videos at: www.robot-learning.uk/coarse-to-fine-imitation-learning. Edward Johns |
ICRA | 1 |
| 2021 | Hybrid ICPabstractICP algorithms typically involve a fixed choice of data association method and a fixed choice of error metric. In this paper, we propose Hybrid ICP, a novel and flexible ICP variant which dynamically optimises both the data association method and error metric based on the live image of an object and the current ICP estimate. We show that when used for object pose estimation, Hybrid ICP is more accurate and more robust to noise than other commonly used ICP variants. We also consider the setting where ICP is applied sequentially with a moving camera, and we study the trade-off between the accuracy of each ICP estimate and the number of ICP estimates available within a fixed amount of time. Kamil Dreczkowski, Edward Johns |
IROS | 2 |
| 2021 | Coarse-to-Fine for Sim-to-Real: Sub-Millimetre Precision Across Wide Task SpacesabstractIn this paper, we study the problem of zero-shot sim-to-real when the task requires both highly precise control with sub-millimetre error tolerance, and wide task space generalisation. Our framework involves a coarse-to-fine controller, where trajectories begin with classical motion planning using ICP-based pose estimation, and transition to a learned end-to-end controller which maps images to actions and is trained in simulation with domain randomisation. In this way, we achieve precise control whilst also generalising the controller across wide task spaces, and keeping the robustness of vision-based, end-to-end control. Real-world experiments on a range of different tasks show that, by exploiting the best of both worlds, our framework significantly outperforms purely motion planning methods, and purely learning-based methods. Furthermore, we answer a range of questions on best practices for precise sim-to-real transfer, such as how different image sensor modalities and image feature representations perform. Eugene Valassakis, Norman Di Palo, Edward Johns |
IROS | 3 |
| 2020 | Shape Adaptor: A Learnable Resizing Module
Shikun Liu, Zhe Lin 0001, Yilin Wang 0002, Jianming Zhang 0001, Federico Perazzi, Edward Johns |
ECCV (12) | 6 |
| 2020 | Sim-to-Real Transfer for Optical Tactile SensingabstractDeep learning and reinforcement learning methods have been shown to enable learning of flexible and complex robot controllers. However, the reliance on large amounts of training data often requires data collection to be carried out in simulation, with a number of sim-to-real transfer methods being developed in recent years. In this paper, we study these techniques for tactile sensing using the TacTip optical tactile sensor, which consists of a deformable tip with a camera observing the positions of pins inside this tip. We designed a model for soft body simulation which was implemented using the Unity physics engine, and trained a neural network to predict the locations and angles of edges when in contact with the sensor. Using domain randomisation techniques for sim-to-real transfer, we show how this framework can be used to accurately predict edges with less than 1 mm prediction error in real-world testing, without any real-world data at all. Nathan F. Lepora, Edward Johns |
ICRA | 3 |
| 2020 | Physics-Based Dexterous Manipulations with Estimated Hand Poses and Residual Reinforcement LearningabstractDexterous manipulation of objects in virtual environments with our bare hands, by using only a depth sensor and a state-of-the-art 3D hand pose estimator (HPE), is challenging. While virtual environments are ruled by physics, e.g. object weights and surface frictions, the absence of force feedback makes the task challenging, as even slight inaccuracies on finger tips or contact points from HPE may make the interactions fail. Prior arts simply generate contact forces in the direction of the fingers' closures, when finger joints penetrate virtual objects. Although useful for simple grasping scenarios, they cannot be applied to dexterous manipulations such as inhand manipulation. Existing reinforcement learning (RL) and imitation learning (IL) approaches train agents that learn skills by using task-specific rewards, without considering any online user input. In this work, we propose to learn a model that maps noisy input hand poses to target virtual poses, which introduces the needed contacts to accomplish the tasks on a physics simulator. The agent is trained in a residual setting by using a model-free hybrid RL+IL approach. A 3D hand pose estimation reward is introduced leading to an improvement on HPE accuracy when the physics-guided corrected target poses are remapped to the input space. As the model corrects HPE errors by applying minor but crucial joint displacements for contacts, this helps to keep the generated motion visually close to the user input. Since HPE sequences performing successful virtual interactions do not exist, a data generation scheme to train and evaluate the system is proposed. We test our framework in two applications that use hand pose estimates for dexterous manipulations: hand-object interactions in VR and hand-object motion reconstruction in-the-wild. Experiments show that the proposed method outperforms various RL/IL baselines and the simple prior art of enforcing hand closure, both in task success and hand pose accuracy. Guillermo Garcia-Hernando, Edward Johns, Tae-Kyun Kim 0001 |
IROS | 2 |
| 2020 | Crossing the Gap: A Deep Dive into Zero-Shot Sim-to-Real Transfer for DynamicsabstractZero-shot sim-to-real transfer of tasks with complex dynamics is a highly challenging and unsolved problem. A number of solutions have been proposed in recent years, but we have found that many works do not present a thorough evaluation in the real world, or underplay the significant engineering effort and task-specific fine tuning that is required to achieve the published results. In this paper, we dive deeper into the sim-to-real transfer challenge, investigate why this is such a difficult problem, and present objective evaluations of a number of transfer methods across a range of real-world tasks. Surprisingly, we found that a method which simply injects random forces into the simulation performs just as well as more complex methods, such as those which randomise the simulator's dynamics parameters, or adapt a policy online using recurrent network architectures. Eugene Valassakis, Edward Johns |
IROS | 3 |
| 2019 | End-To-End Multi-Task Learning With AttentionabstractWe propose a novel multi-task learning architecture, which allows learning of task-specific feature-level attention. Our design, the Multi-Task Attention Network (MTAN), consists of a single shared network containing a global feature pool, together with a soft-attention module for each task. These modules allow for learning of task-specific features from the global features, whilst simultaneously allowing for features to be shared across different tasks. The architecture can be trained end-to-end and can be built upon any feed-forward neural network, is simple to implement, and is parameter efficient. We evaluate our approach on a variety of datasets, across both image-to-image predictions and image classification tasks. We show that our architecture is state-of-the-art in multi-task learning compared to existing methods, and is also less sensitive to various weighting schemes in the multi-task loss function. Code is available at https://github.com/lorenmt/mtan. Shikun Liu, Edward Johns, Andrew J. Davison |
CVPR | 2 |
| 2019 | Self-Supervised Generalisation with Meta Auxiliary LearningabstractLearning with auxiliary tasks can improve the ability of a primary task to generalise. However, this comes at the cost of manually labelling auxiliary data. We propose a new method which automatically learns appropriate labels for an auxiliary task, such that any supervised learning task can be improved without requiring access to any further data. The approach is to train two neural networks: a label-generation network to predict the auxiliary labels, and a multi-task network to train the primary task alongside the auxiliary task. The loss for the label-generation network incorporates the loss of the multi-task network, and so this interaction between the two networks can be seen as a form of meta learning with a double gradient. We show that our proposed method, Meta AuXiliary Learning (MAXL), outperforms single-task learning on 7 image datasets, without requiring any additional data. We also show that MAXL outperforms several other baselines for generating auxiliary labels, and is even competitive when compared with human-defined auxiliary labels. The self-supervised nature of our method leads to a promising new direction towards automated generalisation. Source code can be found at \url{https://github.com/lorenmt/maxl}. Shikun Liu, Andrew J. Davison, Edward Johns |
NeurIPS | 3 |
| 2017 | Application-oriented design space exploration for SLAM algorithmsabstractIn visual SLAM, there are many software and hardware parameters, such as algorithmic thresholds and GPU frequency, that need to be tuned; however, this tuning should also take into account the structure and motion of the camera. In this paper, we determine the complexity of the structure and motion with a few parameters calculated using information theory. Depending on this complexity and the desired performance metrics, suitable parameters are explored and determined. Additionally, based on the proposed structure and motion parameters, several applications are presented, including a novel active SLAM approach which guides the camera in such a way that the SLAM algorithm achieves the desired performance metrics. Real-world and simulated experimental results demonstrate the effectiveness of the proposed design space and its applications. Sajad Saeedi G., Luigi Nardi, Edward Johns, Bruno Bodin, Paul H. J. Kelly, Andrew J. Davison |
ICRA | 3 |
| 2016 | Pairwise Decomposition of Image Sequences for Active Multi-view RecognitionabstractA multi-view image sequence provides a much richer capacity for object recognition than from a single image. However, most existing solutions to multi-view recognition typically adopt hand-crafted, model-based geometric methods, which do not readily embrace recent trends in deep learning. We propose to bring Convolutional Neural Networks to generic multi-view recognition, by decomposing an image sequence into a set of image pairs, classifying each pair independently, and then learning an object classifier by weighting the contribution of each pair. This allows for recognition over arbitrary camera trajectories, without requiring explicit training over the potentially infinite number of camera paths and lengths. Building these pairwise relationships then naturally extends to the next-best-view problem in an active recognition framework. To achieve this, we train a second Convolutional Neural Network to map directly from an observed image to next viewpoint. Finally, we incorporate this into a trajectory optimisation task, whereby the best recognition confidence is sought for a given trajectory length. We present state-of-the-art results in both guided and unguided multi-view recognition on the ModelNet dataset, and show how our method can be used with depth images, greyscale images, or both. Edward Johns, Stefan Leutenegger, Andrew J. Davison |
CVPR | 1 |
| 2016 | Deep learning a grasp function for grasping under gripper pose uncertaintyabstractThis paper presents a new method for parallel-jaw grasping of isolated objects from depth images, under large gripper pose uncertainty. Whilst most approaches aim to predict the single best grasp pose from an image, our method first predicts a score for every possible grasp pose, which we denote the grasp function. With this, it is possible to achieve grasping robust to the gripper's pose uncertainty, by smoothing the grasp function with the pose uncertainty function. Therefore, if the single best pose is adjacent to a region of poor grasp quality, that pose will no longer be chosen, and instead a pose will be chosen which is surrounded by a region of high grasp quality. To learn this function, we train a Convolutional Neural Network which takes as input a single depth image of an object, and outputs a score for each grasp pose across the image. Training data for this is generated by use of physics simulation and depth image simulation with 3D object meshes, to enable acquisition of sufficient data without requiring exhaustive real-world experiments. We evaluate with both synthetic and real experiments, and show that the learned grasp score is more robust to gripper pose uncertainty than when this uncertainty is not accounted for. Edward Johns, Stefan Leutenegger, Andrew J. Davison |
IROS | 1 |
| 2016 | Robust Image Descriptors for Real-Time Inter-Examination Retargeting in Gastrointestinal Endoscopy
Menglong Ye, Edward Johns, Benjamin M. Walter, Alexander Meining, Guang-Zhong Yang |
MICCAI (1) | 2 |
| 2015 | Becoming the expert - interactive multi-class machine teachingabstractCompared to machines, humans are extremely good at classifying images into categories, especially when they possess prior knowledge of the categories at hand. If this prior information is not available, supervision in the form of teaching images is required. To learn categories more quickly, people should see important and representative images first, followed by less important images later - or not at all. However, image-importance is individual-specific, i.e. a teaching image is important to a student if it changes their overall ability to discriminate between classes. Further, students keep learning, so while image-importance depends on their current knowledge, it also varies with time. In this work we propose an Interactive Machine Teaching algorithm that enables a computer to teach challenging visual concepts to a human. Our adaptive algorithm chooses, online, which labeled images from a teaching set should be shown to the student as they learn. We show that a teaching strategy that probabilistically models the student's ability and progress, based on their correct and incorrect answers, produces better `experts'. We present results using real human participants across several varied and challenging real-world datasets. Edward Johns, Oisin Mac Aodha, Gabriel J. Brostow |
CVPR | 1 |
| 2014 | Pairwise Probabilistic Voting: Fast Place Recognition without RANSAC
Edward Johns, Guang-Zhong Yang |
ECCV (2) | 1 |
| 2014 | Online Scene Association for Endoscopic Navigation
Menglong Ye, Edward Johns, Stamatia Giannarou, Guang-Zhong Yang |
MICCAI (2) | 2 |
| 2014 | Generative Methods for Long-Term Place Recognition in Dynamic Scenes
Edward Johns, Guang-Zhong Yang |
Int. J. Comput. Vis. | 1 |
| 2013 | Dynamic scene models for incremental, long-term, appearance-based localisationabstractIn this paper we present a new appearance-based localisation system that is able to deal with dynamic elements in the scene. By independently modelling the properties of local features observed in a scene over long periods of time, we show that feature appearances and geometric relationships can be learned more accurately than when representing a location by a single image. We also present a new dataset consisting of a 6 km outdoor path traversed once per month for a period of 5 months, which contains several challenges including short-term and long-term dynamic behaviour, lateral deviations in the path, repetitive scene appearances and strong illumination changes. We show superior performance of the dynamic mapping system compared to state-of-the-art techniques on our dataset. Edward Johns, Guang-Zhong Yang |
ICRA | 1 |
| 2013 | Feature Co-occurrence Maps: Appearance-based localisation throughout the dayabstractIn this paper we present a new method, Feature Co-occurrence Maps, for appearance-based localisation over the course of a day. We show that by quantising local features in both feature and image space, discriminative statistics can be learned on the co-occurrences of features at different times of the day. This allows for matching at any time, without requiring individual images to be stored representing each time of day, and matching is performed efficiently by simultaneously matching to the entire database. We further show how matching along image sequences can be incorporated into the system and adapt existing methods by allowing for non-zero acceleration. Results on a 20km outdoor dataset show improved performance in precision-recall over state of the art. Edward Johns, Guang-Zhong Yang |
ICRA | 1 |
| 2012 | An Intelligent Food-Intake Monitoring System Using Wearable SensorsabstractThe prevalence of obesity worldwide presents a great challenge to existing healthcare systems. There is a general need for pervasive monitoring of the dietary behaviour of those who are at risk of co-morbidities. Currently, however, there is no accurate method of assessing the nutritional intake of people in their home environment. Traditional methods require subjects to manually respond to questionnaires for analysis, which is subjective, prone to errors, and difficult to ensure consistency and compliance. In this paper, we present a wearable sensor platform that autonomously provides detailed information regarding a subject's dietary habits. The sensor consists of a microphone and a camera and is worn discretely on the ear. Sound features are extracted in real-time and if a chewing activity is classified, the camera captures a video sequence for further analysis. From this sequence, a number of key frames are extracted to represent important episodes during the course of a meal. Results show a high classification rate of chewing activities, and the visual log demonstrates a detailed overview of the subject's food intake that is difficult to quantify from manually-acquired food records. Edward Johns, Louis Atallah, Claire Pettitt, Benny P. L. Lo, Gary S. Frost, Guang-Zhong Yang |
BSN | 2 |
| 2011 | Place Recognition and Online Learning in Dynamic Scenes with Spatio-Temporal LandmarksabstractThis paper presents a new framework for visual place recognition that incrementally learns models of each place and offers adaptability to dynamic elements in a scene. Traditional bag-of-features image-retrieval approaches to place recognition treat images in a holistic manner and are typically not capable of dealing with sub-scene dynamics, such as structural changes to a building facade or the rearrangement of furniture in a room. However, by treating local features as observations of real-world landmarks in a scene that are consistently observed, such dynamics can be accurately modelled at a local level, and the spatio-temporal properties of each landmark can be independently updated online. We propose a framework for place recognition that models each scene by sequentially learning landmarks from a set of images, and in the long term adapts the model to dynamic behaviour. Results on both indoor and outdoor datasets show an improvement in recognition performance and efficiency when compared to the traditional bag-offeatures image retrieval approach. Edward Johns, Guang-Zhong Yang |
BMVC | 1 |
| 2011 | From images to scenes: Compressing an image cluster into a single scene model for place recognitionabstractThe recognition of a place depicted in an image typically adopts methods from image retrieval in large-scale databases. First, a query image is described as a “bag-of-features” and compared to every image in the database. Second, the most similar images are passed to a geometric verification stage. However, this is an inefficient approach when considering that some database images may be almost identical, and many image features may not repeatedly occur. We address this issue by clustering similar database images to represent distinct scenes, and tracking local features that are consistently detected to form a set of real-world landmarks. Query images are then matched to landmarks rather than features, and a probabilistic model of landmark properties is learned from the cluster to appropriately verify or reject putative feature matches. We present novelties in both a bag-of-features retrieval and geometric verification stage based on this concept. Results on a database of 200K images of popular tourist destinations show improvements in both recognition performance and efficiency compared to traditional image retrieval methods. Edward Johns, Guang-Zhong Yang |
ICCV | 1 |
| 2011 | Global localization in a dense continuous topological mapabstractVision-based topological maps for mobile robot localization traditionally consist of a set of images captured along a path, with a query image then compared to every individual map image. This paper introduces a new approach to topological mapping, whereby the map consists of a set of landmarks that are detected across multiple images, spanning the continuous space between nodal images. Matches are then made to landmarks, rather than to individual images, enabling a topological map of far greater density than traditionally possible, without sacrificing computational speed. Furthermore, by treating each landmark independently, a probabilistic approach to localization can be employed by taking into account the learned discriminative properties of each landmark. An optimization stage is then used to adjust the map according to speed and localization accuracy requirements. Results for global localization show a greater positive location identification rate compared to the traditional topological map, together with enabling a greater localization resolution in the denser topological map, without requiring a decrease in frame rate. Edward Johns, Guang-Zhong Yang |
ICRA | 1 |
| 2011 | A scene-associated training method for mobile robot speech recognition in multisource reverberated environmentsabstractIn this paper, we present a new technique for social mobile robot speech recognition based on scene-associated training models. The key contribution of the paper is a real-time framework that reduces the effect of room reverberation and ambient noise, a challenging problem in speech recognition. In classical approaches, anechoic sound is used to train the model, with the main focus on removing reverberation or noise from the sound. Our technique differs in that we train a number of speech recognizers directly from the reverberated sound, by associating each recognizer with a unique visual scene, to deal with the varying reverberation properties of different rooms. By extracting local features from a captured image and recognizing a scene, the robot can use the appropriate speech recognizer that is trained for the particular structural properties of that scene. We tested our method by using a baseline speech recognition model (HTK) across a variety of rooms and different levels of background noise. The results show that the association between a visual scene and a corresponding speech recognizer greatly improves the robot's speech recognition accuracy, together with increasing the computational speed of recognition, compared to competing techniques. Edward Johns, Guang-Zhong Yang |
IROS | 2 |
| 2010 | Scene association for mobile robot navigationabstractAccurate, efficient and robust location recognition is a fundamental task for any mobile robot. This paper presents a new approach using visual features to efficiently represent a series of locations along a path in an indoor environment. In the training stage, local features which are detected across multiple images from a single tour are combined to represent a real-world landmark, modelled by the expected variance of its descriptor. Those landmarks which represent the scene in the most efficient and discriminative manner are then retained, and this selection is optimized with respect to the scale of the environment. In the recognition stage, features detected in an image are matched to the landmarks in memory, based upon a novel similarity measure drawing from feature co-occurrence statistics. Edward Johns, Guang-Zhong Yang |
IROS | 1 |