VLDB 2026 Research / reviewers in the wild / expert
Walterio W. Mayol-Cuevas
dblp:64/3609 · also Walterio W. Mayol
· DBLP profile ↗
75ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0001-8973-1931ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 1 first-author · 6 since 2021Systems, architecture and hardware · 25 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 16 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EvoStruggle: A Dataset Capturing the Evolution of Struggle Across Activities and Skill Levels
Shijia Feng, Michael Wray, Walterio W. Mayol-Cuevas |
ICPR (5) | 3 |
| 2026 | From Detection to Anticipation: Online Understanding of Struggles across Various Tasks and ActivitiesabstractUnderstanding human skill performance is essential for intelligent assistive systems, with struggle recognition offering a natural cue for identifying user difficulties. While prior work focuses on offline struggle classification and localization, real-time applications require models capable of detecting and anticipating struggle online. We reformulate struggle localization as an online detection task and further extend it to anticipation, predicting struggle moments before they occur. We adapt two off-the-shelf models as baselines for online struggle detection and anticipation. Online struggle detection achieves 70-80% per-frame mAP, while struggle anticipation up to 2 seconds ahead yields comparable performance with slight drops. We further examine generalization across tasks and activities and analyse the impact of skill evolution. Despite larger domain gaps in activity-level generalization, models still outperform random baselines by 4-20%. Our feature-based models run at up to 143 FPS, and the whole pipeline, including feature extraction, operates at around 20 FPS, sufficient for real-time assistive applications. Shijia Feng, Michael Wray, Walterio W. Mayol-Cuevas |
WACV | 3 |
| 2025 | Focal Plane Visual Feature Generation and Matching on a Pixel Processor Array
Laurie Bose, Jianing Chen 0005, Piotr Dudek, Walterio W. Mayol-Cuevas |
ICCV | 5 |
| 2025 | CULTURE3D: A Large-Scale and Diverse Dataset of Cultural Landmarks and Terrains for Gaussian-Based Scene RenderingabstractCurrent state-of-the-art 3D reconstruction models face limitations in building extra-large scale outdoor scenes, primarily due to the lack of sufficiently large-scale and detailed datasets. In this paper, we present a extra-large fine-grained dataset with 10 billion points composed of 41,006 drone-captured high-resolution aerial images, covering 20 diverse and culturally significant scenes from worldwide locations such as Cambridge Uni main buildings, the Pyramids, and the Forbidden City Palace. Compared to existing datasets, ours offers significantly larger scale and higher detail, uniquely suited for fine-grained 3D applications. Each scene contains an accurate spatial layout and comprehensive structural information, supporting detailed 3D reconstruction tasks. By reconstructing environments using these detailed images, our dataset supports multiple applications, including outputs in the widely adopted COLMAP format, establishing a novel benchmark for evaluating state-of-the-art large-scale Gaussian Splatting methods.The dataset's flexibility encourages innovations and supports model plug-ins, paving the way for future 3D breakthroughs. All datasets and code will be open-sourced for community use. Steve Zhang, Weizhe Lin, Aaron Zhang, Walterio W. Mayol-Cuevas, Junxiao Shen |
ICCV | 5 |
| 2025 | Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly VideosabstractAbstract Determining when people are struggling allows for a finer-grained understanding of actions that complements conventional action classification and error detection. Struggle detection, as defined in this paper, is a distinct and important task that can be identified without explicit step or activity knowledge. We introduce the first struggle dataset with three real-world problem-solving activities that are labelled by both expert and crowd-source annotators. Video segments were scored w.r.t. their level of struggle using a forced choice 4-point scale. This dataset contains 5.1 hours of video from 73 participants. We conducted a series of experiments to identify the most suitable modelling approaches for struggle determination. Additionally, we compared various deep learning models, establishing baseline results for struggle classification, struggle regression, and struggle label distribution learning. Our results indicate that struggle detection in video can achieve up to $$88.24\%$$ 88.24 % accuracy in binary classification, while detecting the level of struggle in a four-way classification setting performs lower, with an overall accuracy of $$52.45\%$$ 52.45 % . Our work is motivated toward a more comprehensive understanding of action in video and potentially the improvement of assistive systems that analyse struggle and can better support users during manual activities. Shijia Feng, Michael Wray, Brian Sullivan, Youngkyoon Jang, Casimir J. H. Ludwig, Iain D. Gilchrist, Walterio W. Mayol-Cuevas |
Int. J. Comput. Vis. | 7 |
| 2023 | Environment modeling and localization from datasets of omnidirectional scenes using machine learning techniquesabstractAbstract This work presents a framework to create a visual model of the environment which can be used to estimate the position of a mobile robot by means of artificial intelligence techniques. The proposed framework retrieves the structure of the environment from a dataset composed of omnidirectional images captured along it. These images are described by means of global-appearance approaches. The information is arranged in two layers, with different levels of granularity. The first layer is obtained by means of classifiers and the second layer is composed of a set of data fitting neural networks. Subsequently, the model is used to estimate the position of the robot, in a hierarchical fashion, by comparing the image captured from the unknown position with the information in the model. Throughout this work, five classifiers are evaluated (Naïve Bayes, SVM, random forest, linear discriminant classifier and a classifier based on a shallow neural network) along with three different global-appearance descriptors (HOG, gist, and a descriptor calculated from an intermediate layer of a pre-trained CNN). The experiments have been tackled with some publicly available datasets of omnidirectional images captured indoors with the presence of dynamic changes. Several parameters are used to assess the efficiency of the proposal: the ability of the algorithm to estimate coarsely the position (hit ratio), the average error (cm) and the necessary computing time. The results prove the efficiency of the framework to model the environment and localize the robot from the knowledge extracted from a set of omnidirectional images with the proposed artificial intelligence techniques. Sergio Cebollada, Luis Payá, Adrián Peidró, Walterio W. Mayol-Cuevas, Óscar Reinoso |
Neural Comput. Appl. | 4 |
| 2021 | Weighted Node Mapping and Localisation on a Pixel Processor ArrayabstractThis paper implements and demonstrates visual route mapping and localisation upon a Pixel Processor Array (PPA). The PPA sensor comprises of an array of Processing Elements (PEs), each of which can capture and process visual information directly. This provides significant parallel processing power allowing novel ways in which information can be processed on-sensor. Our method predicts the correct node within a topological map generated from an image sequence by measuring image similarities, spatial coherence, and exploiting the parallel nature of the PPA. Our implementation runs at +300Hz on large public datasets with +2K locations requiring 2.5W at 500 GOPS/W. We compare vs traditionally implemented methods demonstrating better F-1 performance even on simulation. As far as we are aware, we present the first on-sensor mapping and localisation system running entirely on-sensor. Hector Castillo-Elizalde, Laurie Bose, Walterio W. Mayol-Cuevas |
ICRA | 4 |
| 2021 | The Object at Hand: Automated Editing for Mixed Reality Video Guidance from Hand-Object InteractionsabstractIn this paper, we concern with the problem of how to automatically extract the steps that compose real-life hand activities. This is a key competence towards processing, monitoring and providing video guidance in Mixed Reality systems. We use egocentric vision to observe hand-object interactions in real-world tasks and automatically decompose a video into its constituent steps. Our approach combines hand-object interaction (HOI) detection, object similarity measurement and a finite state machine (FSM) representation to automatically edit videos into steps. We use a combination of Convolutional Neural Networks (CNNs) and the FSM to discover, edit cuts and merge segments while observing real hand activities. We evaluate quantitatively and qualitatively our algorithm on two datasets: the GTEA [19], and a new dataset we introduce for Chinese Tea making. Results show our method is able to segment hand-object interaction videos into key step segments with high levels of precision. Walterio W. Mayol-Cuevas |
ISMAR | 2 |
| 2021 | Agile reactive navigation for a non-holonomic mobile robot using a pixel processor arrayabstractAbstract This paper presents an agile reactive navigation strategy for driving a non‐holonomic ground vehicle around a pre‐set course of gates in a cluttered environment using a low‐cost processor array sensor. This enables machine vision tasks to be performed directly upon the sensor's image plane, rather than using a separate general‐purpose computer. The authors demonstrate a small ground vehicle running through or avoiding multiple gates at high speed using minimal computational resources. To achieve this, target tracking algorithms are developed for the Pixel Processing Array and captured images are then processed directly on the vision sensor acquiring target information for controlling the ground vehicle. The algorithm can run at up to 2000 fps outdoors and 200 fps at indoor illumination levels. Conducting image processing at the sensor level avoids the bottleneck of image transfer encountered in conventional sensors. The real‐time performance of on‐board image processing and robustness is validated through experiments. Experimental results demonstrate the algorithm's ability to enable a ground vehicle to navigate at an average speed of 2.20 m/s for passing through multiple gates and 3.88 m/s for a ‘slalom’ task in an environment featuring significant visual clutter. Laurie Bose, Colin Greatwood, Jianing Chen 0005, Rui Fan 0001, Tom Richardson 0002, Stephen J. Carey, Piotr Dudek, Walterio W. Mayol-Cuevas |
IET Image Process. | 9 |
| 2020 | High-speed Light-weight CNN Inference via Strided Convolutions on a Pixel Processor Array
Laurie Bose, Jianing Chen 0005, Stephen J. Carey, Piotr Dudek, Walterio W. Mayol-Cuevas |
BMVC | 6 |
| 2020 | Action Modifiers: Learning From Adverbs in Instructional Videos
Hazel Doughty, Ivan Laptev, Walterio W. Mayol-Cuevas, Dima Damen |
CVPR | 3 |
| 2020 | Fully Embedding Fast Convolutional Networks on Pixel Processor Arrays
Laurie Bose, Piotr Dudek, Jianing Chen 0005, Stephen J. Carey, Walterio W. Mayol-Cuevas |
ECCV (29) | 5 |
| 2020 | Centroids Triplet Network and Temporally-Consistent Embeddings for In-Situ Object RecognitionabstractThis work proposes learning to recognize objects from a small number of training examples collected and deployed in-situ. That is, from data collected where the objects are commonly placed or being used, perhaps after first encountering them, the learning algorithm immediately is able to recognize them again. We refer to this method-ology as in-situ learning, and it opposes to the conventional methodology of using complex data acquisition mechanisms, such as rotating tables or synthetic data, to build a large-scale dataset for training convolutional neural networks (ConvNets). To learn in-situ, we propose a novel loss function that generates discriminative features for known and unseen objects, by utilizing a regularization term that reduces the distance between features and their manifold centroid. Additionally, we propose a temporal filter that is particularly useful to quickly react to appearing objects on the scene, which depending on the distance between neighboring video-frame features, it applies a weighted average between the current and the previous frame. Our framework achieves state-of-the-art accuracy for in-situ and on-the-fly learning, for the case of known objects achieves an average increase in accuracy of 3.01%, an increase of 3.3% for novel objects, and an average increase of 7.07% for the combined case, compared with the closest baseline. Utilizing the temporal filtering, led to a further increase in accuracy against nuisances of 7.32% for the known and novels objects case. Miguel Lagunes-Fortiz, Dima Damen, Walterio W. Mayol-Cuevas |
IROS | 3 |
| 2020 | Live Demonstration: CNN Inference on the Focal Plane with a Pixel Processor ArrayabstractWe present a novel method of CNN inference on a pixel processor array device, demonstrating it using a handwritten digit (digits 0-9) classification task, with all steps of the neural network computation performed on the focal plane. The vision chip that we deploy (SCAMP-7) has a 256×256 array of processor elements (PE) integrated within the image sensor. The algorithm runs at over 3000 frames per second (FPS) and over 90% classification accuracy, with the sensor chip only outputting ten scalar values corresponding to the classification scores. Stephen J. Carey, Laurie Bose, Tom Richardson 0002, Walterio W. Mayol-Cuevas, Jianing Chen 0005, Piotr Dudek |
ISCAS | 4 |
| 2019 | The Pros and Cons: Rank-Aware Temporal Attention for Skill Determination in Long VideosabstractWe present a new model to determine relative skill from long videos, through learnable temporal attention modules. Skill determination is formulated as a ranking problem, making it suitable for common and generic tasks. However, for long videos, parts of the video are irrelevant for assessing skill, and there may be variability in the skill exhibited throughout a video. We therefore propose a method which assesses the relative overall level of skill in a long video by attending to its skill-relevant parts. Our approach trains temporal attention modules, learned with only video-level supervision, using a novel rank-aware loss function. In addition to attending to task-relevant video parts, our proposed loss jointly trains two attention modules to separately attend to video parts which are indicative of higher (pros) and lower (cons) skill. We evaluate our approach on the EPIC-Skills dataset and additionally annotate a larger dataset from YouTube videos for skill determination with five previously unexplored tasks. Our method outperforms previous approaches and classic softmax attention on both datasets by over 4% pairwise accuracy, and as much as 12% on individual tasks. We also demonstrate our model’s ability to attend to rank-aware parts of the video. Hazel Doughty, Walterio W. Mayol-Cuevas, Dima Damen |
CVPR | 2 |
| 2019 | A Camera That CNNs: Towards Embedded Neural Networks on Pixel Processor ArraysabstractWe present a convolutional neural network implementation for pixel processor array (PPA) sensors. PPA hardware consists of a fine-grained array of general-purpose processing elements, each capable of light capture, data storage, program execution, and communication with neighboring elements. This allows images to be stored and manipulated directly at the point of light capture, rather than having to transfer images to external processing hardware. Our CNN approach divides this array up into 4x4 blocks of processing elements, essentially trading-off image resolution for increased local memory capacity per 4x4 ”pixel”. We implement parallel operations for image addition, subtraction and bit-shifting images in this 4x4 block format. Using these components we formulate how to perform ternary weight convolutions upon these images, compactly store results of such convolutions, perform max-pooling, and transfer the resulting sub-sampled data to an attached micro-controller. We train ternary weight filter CNNs for digit recognition and a simple tracking task, and demonstrate inference of these networks upon the SCAMP5 PPA system. This work represents a first step towards embedding neural network processing capability directly onto the focal plane of a sensor. Laurie Bose, Piotr Dudek, Jianing Chen 0005, Stephen J. Carey, Walterio W. Mayol-Cuevas |
ICCV | 5 |
| 2019 | Learning Discriminative Embeddings for Object Recognition on-the-flyabstractWe address the problem of learning to recognize new objects on-the-fly efficiently. When using CNNs, a typical approach for learning new objects is by fine-tuning the model. However, this approach relies on the assumption that the original training set is available and requires high-end computational resources for training the ever-growing dataset efficiently, which can be unfeasible for robots with limited hardware. To overcome these limitations, we propose a new architecture that: 1) Instead of predicting labels, it learns to generate discriminative and separable embeddings of an object's viewpoints by using a Supervised Triplet Loss, which is easier to implement than current smart mining techniques and the trained model can be applied to unseen objects. 2) Infers an object's identity efficiently by utilizing a lightweight classifier in the features embedding space, this keeps the inference time in the order of milliseconds and can be retrained efficiently when new objects are learned. We evaluate our approach on four real-world images datasets used for Robotics and Computer Vision applications: Amazon Robotics Challenge 2017 by MIT-Princeton, T-LESS, ToyBoX, and CORe50 datasets. Code available at [1]. Miguel Lagunes-Fortiz, Dima Damen, Walterio W. Mayol-Cuevas |
ICRA | 3 |
| 2019 | Rebellion and Obedience: The Effects of Intention Prediction in Cooperative Handheld RobotsabstractWithin this work, we explore intention inference for user actions in the context of a handheld robot setup. Handheld robots share the shape and properties of handheld tools while being able to process task information and aid manipulation. Here, we propose an intention prediction model to enhance cooperative task solving. The model derives intention from the combined information about the user's gaze pattern and task knowledge. Within experimental studies, the model is validated through a comparison of user frustration for the case where the robot follows the predicted location of the user's intended action versus doing the opposite (rebellion). The proposed model yields real-time capabilities and reliable accuracy up to 1.5 s prior to predicted actions being executed. Janis Stolzenwald, Walterio W. Mayol-Cuevas |
IROS | 2 |
| 2018 | Who's Better? Who's Best? Pairwise Deep Ranking for Skill DeterminationabstractThis paper presents a method for assessing skill from video, applicable to a variety of tasks, ranging from surgery to drawing and rolling pizza dough. We formulate the problem as pairwise (who's better?) and overall (who's best?) ranking of video collections, using supervised deep ranking. We propose a novel loss function that learns discriminative features when a pair of videos exhibit variance in skill, and learns shared features when a pair of videos exhibit comparable skill levels. Results demonstrate our method is applicable across tasks, with the percentage of correctly ordered pairs of videos ranging from 70% to 83% for four datasets. We demonstrate the robustness of our approach via sensitivity analysis of its parameters. We see this work as effort toward the automated organization of how-to video collections and overall, generic skill determination in video. Hazel Doughty, Dima Damen, Walterio W. Mayol-Cuevas |
CVPR | 3 |
| 2018 | Where can i do this? Geometric Affordances from a Single Example with the Interaction TensorabstractThis paper introduces and evaluates a new tensor field representation to express the geometric affordance of one object relative to another, a key competence for Cognitive and Autonomous robots. We expand the bisector surface representation to one that is weight-driven and that retains the provenance of surface points with directional vectors. We also incorporate the notion of affordance keypoints which allow for faster decisions at a point of query and with a compact and straightforward descriptor. Using a single interaction example, we are able to generalize to previously-unseen scenarios; both synthetic and also real scenes captured with RGB-D sensors. Evaluations also include crowdsourcing comparisons that confirm the validity of our affordance proposals, which agree on average 84 % of the time with human judgments, that is 20-40 % better than the baseline methods. Eduardo Ruiz, Walterio W. Mayol-Cuevas |
ICRA | 2 |
| 2018 | Perspective Correcting Visual Odometry for Agile MAVs using a Pixel Processor ArrayabstractThis paper presents a visual odometry approach using a Pixel Processor Array (PPA) camera, specifically, the SCAMP-5 vision chip. In this device, each pixel is capable of storing data and performing computation, enabling a variety of computer vision tasks to be carried out directly upon the sensor itself. In this work the PPA performs HDR edge detection, perspective correction and image alignment based odometry, allowing the position and heading of a MAV to be tracked at several hundred frames per second. We evaluate our PPA based approach by direct comparison with a motion capture system for a variety of trajectories. These include rapid accelerations that would incur significant motion blur at low frame rates, and lighting conditions that would typically lead to under or over exposure of image detail. Such challenging conditions would often lead to unusable images when relying on traditional image sensors. Colin Greatwood, Laurie Bose, Tom Richardson 0002, Walterio W. Mayol-Cuevas, Jianing Chen 0005, Stephen J. Carey, Piotr Dudek |
IROS | 4 |
| 2018 | I Can See Your Aim: Estimating User Attention from Gaze for Handheld Robot CollaborationabstractThis paper explores the estimation of user attention in the setting of a cooperative handheld robot - a robot designed to behave as a handheld tool but that has levels of task knowledge. We use a tool-mounted gaze tracking system, which, after modelling via a pilot study, we use as a proxy for estimating the attention of the user. This information is then used for cooperation with users in a task of selecting and engaging with objects on a dynamic screen. Via a video game setup, we test various degrees of robot autonomy from fully autonomous, where the robot knows what it has to do and acts, to no autonomy where the user is in full control of the task. Our results measure performance and subjective metrics and show how the attention model benefits the interaction and preference of users. Janis Stolzenwald, Walterio W. Mayol-Cuevas |
IROS | 2 |
| 2017 | Visual Odometry for Pixel Processor ArraysabstractWe present an approach of estimating constrained egomotion on a Pixel Processor Array (PPA). These devices embed processing and data storage capability into the pixels of the image sensor, allowing for fast and low power parallel computation directly on the image-plane. Rather than the standard visual pipeline whereby whole images are transferred to an external general processing unit, our approach performs all computation upon the PPA itself, with the camera's estimated motion as the only information output. Our approach estimates 3D rotation and a 1D scale-less estimate of translation. We introduce methods of image scaling, rotation and alignment which are performed solely upon the PPA itself and form the basis for conducting motion estimation. We demonstrate the algorithms on a SCAMP-5 vision chip, achieving frame rates >1000Hz at ~2W power consumption. Laurie Bose, Jianing Chen 0005, Stephen J. Carey, Piotr Dudek, Walterio W. Mayol-Cuevas |
ICCV | 5 |
| 2017 | Trespassing the Boundaries: Labeling Temporal Bounds for Object Interactions in Egocentric VideoabstractManual annotations of temporal bounds for object interactions (i.e. start and end times) are typical training input to recognition, localization and detection algorithms. For three publicly available egocentric datasets, we uncover inconsistencies in ground truth temporal bounds within and across annotators and datasets. We systematically assess the robustness of state-of-the-art approaches to changes in labeled temporal bounds, for object interaction recognition. As boundaries are trespassed, a drop of up to 10% is observed for both Improved Dense Trajectories and Two- Stream Convolutional Neural Network. We demonstrate that such disagreement stems from a limited understanding of the distinct phases of an action, and propose annotating based on the Rubicon Boundaries, inspired by a similarly named cognitive model, for consistent temporal bounds of object interactions. Evaluated on a public dataset, we report a 4% increase in overall accuracy, and an increase in accuracy for 55% of classes when Rubicon Boundaries are used for temporal annotations. Davide Moltisanti, Michael Wray, Walterio W. Mayol-Cuevas, Dima Damen |
ICCV | 3 |
| 2017 | O-POCO: Online point cloud compression mapping for visual odometry and SLAMabstractThis paper presents O-POCO, a visual odometry and SLAM system that makes online decisions regarding what to map and what to ignore. It takes a point cloud from classical SfM and aims to sample it on-line by selecting map features useful for future 6D relocalisation. We use the camera's traveled trajectory to compartamentalize the point cloud, along with visual and spatial information to sample and compress the map. We propose and evaluate a number of different information layers such as the descriptor information's relative entropy, map-feature occupancy grid, and the point cloud's geometry error. We compare our proposed system against both SfM, and online and offline ORB-SLAM using publicly available datasets in addition to our own. Results show that our online compression strategy is capable of outperforming the baseline even for conditions when the number of features per key-frame used for mapping is four times less. Luis Contreras 0001, Walterio W. Mayol-Cuevas |
ICRA | 2 |
| 2017 | Compression of topological models and localization using the global appearance of visual informationabstractIn this work, a clustering approach to obtain compact topological models of an environment is developed and evaluated. The usefulness of these models is tested by studying their utility to solve the robot localization problem subsequently. Omnidirectional visual information and global appearance descriptors are used both to create and compress the models and to estimate the position of the robot. Comparing to the methods based on the extraction and description of landmarks, global appearance approaches permit building models that can be handled and interpreted more intuitively and using relatively straightforward algorithms to estimate the position of the robot. The proposed algorithms are tested with a set of panoramic images captured with a catadioptric vision sensor in a large environment under real working conditions. The results show that it is possible to compress substantially the visual information contained in topological models to arrive to a balance between the computational cost and the accuracy of the localization process. Luis Payá, Walterio W. Mayol-Cuevas, Sergio Cebollada, Óscar Reinoso |
ICRA | 2 |
| 2017 | Tracking control of a UAV with a parallel visual processorabstractThis paper presents a vision-based control strategy for tracking a ground target using a novel vision sensor featuring a processor for each pixel element. This enables computer vision tasks to be carried out directly on the focal plane in a highly efficient manner rather than using a separate general purpose computer. The strategy enables a small, agile quadrotor Unmanned Air Vehicle (UAV) to track the target from close range using minimal computational effort and with low power consumption. To evaluate the system we target a vehicle driven by chaotic dual-pendulum trajectories. Target proximity and the large, unpredictable accelerations of the vehicle cause challenges for the UAV in keeping it within the downward facing camera's field of view (FoV). A state observer is used to smooth out predictions of the target's location and, importantly, estimate velocity. Experimental results also demonstrate that it is possible to continue to re-acquire and follow the target during short periods of loss in target visibility. The tracking algorithm exploits the parallel nature of the visual sensor, enabling high rate image processing ahead of any communication bottleneck with the UAV controller. With the vision chip carrying out the most intense visual information processing, it is computationally trivial to compute all of the controls for tracking onboard. This work is directed toward visual agile robots that are power efficient and that ferry only useful data around the information and control pathways. Colin Greatwood, Laurie Bose, Tom Richardson 0002, Walterio W. Mayol-Cuevas, Jianing Chen 0005, Stephen J. Carey, Piotr Dudek |
IROS | 4 |
| 2016 | Dominant plane recognition in interior scenes from a single imageabstractRecognition of dominant planes is an important task used in areas such as robot navigation, augmented reality, 3D reconstruction, among others. There are several approaches for recognizing planar structures, however, most of these approaches are based on processing two or more images captured from different camera views or on processing 3D data in the form of point clouds associated with the camera images. An alternative is to process a single image seeking to interpret areas of the images where the planar structure may be observed, thus removing parallax dependency, but adding the challenge of having to correctly interpret image ambiguities. Motivated by the latter, this work presents initial results of a novel methodology for dominant planes recognition in a single image by combining three key strategies: a learning algorithm, a segmentation scheme and a contour detection method. We constraint our approach to work with interior scenes as an attempt to identify key elements that may help in the recognition process in this sort of scenes. In this sense, our results show a recognition accuracy of 60.17% with an error of 3.14%, which indicate the feasibility of our approach. Juan Antonio de Jesús Osuna-Coutiño, José Martínez-Carranza, Miguel O. Arias-Estrada, Walterio W. Mayol-Cuevas |
ICPR | 4 |
| 2016 | Inverse kinematics and design of a novel 6-DoF handheld robot armabstractWe present a novel 6-DoF cable driven manipulator for handheld robotic tasks. Based on a coupled tendon approach, the arm is optimized to maximize movement speed and configuration space while reducing the total mass of the arm. We propose a space carving approach to design optimal link geometry maximizing structural strength and joint limits while minimizing link mass. The design improves on similar non-handheld tendon-driven manipulators and reduces the required number of actuators to one per DoF. As the manipulator has one redundant joint, we present a 5-DoF inverse kinematics solution for the end effector pose. The inverse kinematics is solved by splitting the 6-DoF problem into two coupled 3-DoF problems and merging their results. A method for gracefully degrading the output of the inverse kinematics is described for cases where the desired end effector pose is outside the configuration space. This is useful for settings where the user is in the control loop and can help the robot to get closer to the desired location. The design of the handheld robot is offered as open source. While our results and tools are aimed at handheld robotics, the design and approach is useful to non-handheld applications. Austin Gregg-Smith, Walterio W. Mayol-Cuevas |
ICRA | 2 |
| 2016 | Investigating spatial guidance for a cooperative handheld robotabstractIn this paper we address the question of how to provide feedback information to guide users of a handheld robotic device when performing a spatial exploration task. We consider various feedback methods for communicating a five degree of freedom target pose to a user including a stereoscopic VR display, a monocular see-through AR display and a 2D screen as well as robot arm gesturing. The spatial exploration task with the handheld robot arm was compared against a baseline of a passive handheld wand. We compared the performance of each of the methods with 21 volunteers using a repeated measures ANOVA experimental design, and recorded users' opinions via a NASA Task Load Index Survey. The robot assisted reaching feedback methods significantly outperform manual reaching with the wand for all three visual feedback methods. However, there is little difference between each of the three visual feedback methods when using the robot. The completion time of the task varies with changing difficulty when using the wand but remains stable when assisted by the robot. These results convey useful information for the design of cooperative handheld robots. Austin Gregg-Smith, Walterio W. Mayol-Cuevas |
ICRA | 2 |
| 2016 | You-Do, I-Learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance
Dima Damen, Teesid Leelasawassuk, Walterio W. Mayol-Cuevas |
Comput. Vis. Image Underst. | 3 |
| 2015 | The design and evaluation of a cooperative handheld robotabstractThis paper concerns itself with a relatively unexplored type of personal robot that operates in the tool space. Handheld robots aim to cooperate with the user to solve tasks and improve what tools can offer enhanced by actuation, sensing, and importantly, task knowledge. To this end, we devised a new lightweight robotic platform that has 4 DoF and uses a cable driven continuum structure. Feedback from the robot to the user is provided in an intuitive, implicit manner by the robot end effector pointing towards the goal, avoiding pointing, and/or refusing to perform an action when it conflicts with the task specification. We evaluate two generic tasks involving aiming in space and picking/placing objects with a number of volunteers. Repeated measures ANOVA is used to analyse results to show in which conditions an increased level of automation in the handheld robot improves task performance or user perception of task load. The robot is offered as an open robotics platform[1] and the results indicate directions to improve on feedback and interaction mechanisms. Austin Gregg-Smith, Walterio W. Mayol-Cuevas |
ICRA | 2 |
| 2015 | Inverse depth for accurate photometric and geometric error minimisation in RGB-D dense visual odometryabstractIn this paper we present a dense visual odometry system for RGB-D cameras performing both photometric and geometric error minimisation to estimate the camera motion between frames. Contrary to most works in the literature, we parametrise the geometric error by the inverse depth instead of the depth, which translates into a better fit of the distribution of the geometric error to the used robust cost functions. We also provide a unified evaluation under the same framework of different estimators and ways of computing the scale of the residuals which can be found spread along the related literature. For the comparison of our approach with state-of-the-art approaches we use the popular dataset from the TUM for RGB-D benchmarking. Our approach shows to be competitive with state-of-the-art methods in terms of drift in meters per second, even compared to methods performing loop closure too. When comparing to approaches performing pure odometry like ours, our method outperforms them in the majority of the tested datasets. Additionally we show that our approach is able to work in real time and we provide a qualitative evaluation on our own sequences showing a low drift in the 3D reconstructions. Daniel Gutiérrez-Gómez, Walterio W. Mayol-Cuevas, Josechu J. Guerrero |
ICRA | 2 |
| 2015 | What should I landmark? Entropy of normals in depth juts for place recognition in changing environments using RGB-D dataabstractOne open problem in the fields of place recognition and mapping is to be able to recognise a revisited place when its appearance and layout have changed between visits. In this paper, we investigate this problem in the context of RGB-D mapping in indoor environments. We propose to segment the scene in juts (neighbourhood of 3D points with normals that stick out from the surroundings) and look at low-level features, like textureness or entropy of the normals. These could differentiate those zones of the scene that change or move along time from those that are likely to remain static. We also present a method which improves the matching between images of the same place taken at different times by pruning details basing on these features. We evaluate on a number of communal areas and also on some scenes captured 6 months apart. Experiments with our approach, show an increase up to 70% in inlier matching ratio at the cost of pruning only less than 20% of correct matches, without the need of performing geometric verification. Daniel Gutiérrez-Gómez, Walterio W. Mayol-Cuevas, Josechu J. Guerrero |
ICRA | 2 |
| 2015 | Improving MAV control by predicting aerodynamic effects of obstaclesabstractBuilding on our previous work [1], in this paper we demonstrate how it is possible to improve flight control of a MAV that experiences aerodynamic disturbances caused by objects on its path. Predictions based on low resolution depth images taken at a distance are incorporated into the flight control loop on the throttle channel as this is adjusted to target undisrupted level flight. We demonstrate that a statistically significant improvement (p ≪ 0.001) is possible for some common obstacles such as boxes and steps, compared to using conventional feedback-only control. Our approach and results are encouraging toward more autonomous MAV exploration strategies. John Bartholomew, Andrew Calway, Walterio W. Mayol-Cuevas |
IROS | 3 |
| 2015 | Trajectory-driven point cloud compression techniques for visual SLAMabstractWe develop and evaluate methods based on a novel data compression strategy for visual SLAM that uses traveled trajectory analysis. Beyond compressing scene structure based purely on geometry, we aim at developing compact map representations that are useful for re-exploration while preserving scene structure. Our work is evaluated on data collected from a visual sensor and exploits the information intrinsic to the trajectory of exploration together with the visual information of map points. We perform rigorous statistical evaluation and Pareto analysis to show how this approach compares with three widely used baseline compression methods: k-means on point geometry, keyframes and random sampling. Results indicate that compressing maps to levels of 25% or even less of the original data is possible, while preserving good 6D visual relocalisation performance. Luis Contreras 0001, Walterio W. Mayol-Cuevas |
IROS | 2 |
| 2015 | Correspondence, Matching and Recognition
Tilo Burghardt, Dima Damen, Walterio W. Mayol-Cuevas, Majid Mirmehdi |
Int. J. Comput. Vis. | 3 |
| 2014 | You-Do, I-Learn: Discovering Task Relevant Objects and their Modes of Interaction from Multi-User Egocentric Video
Dima Damen, Teesid Leelasawassuk, Osian Haines, Andrew Calway, Walterio W. Mayol-Cuevas |
BMVC | 5 |
| 2014 | Learning to predict obstacle aerodynamics from depth images for Micro Air VehiclesabstractMany applications of Micro Air Vehicles (MAVs) require them to operate in cluttered environments, flying in constrained spaces and close to obstacles. Such obstacles affect the airflow around the MAV and can thereby affect its flight characteristics. We describe a system for predicting these effects at a distance, using depth images obtained from an RGB-D sensor. Predictions are based on learning from prior experience gathered during training flights. We show that aerodynamic effects caused by obstacles are consistent, and demonstrate that it is practical to make predictions from experience without running a computationally expensive aerodynamic simulation. Our approach uses a Gaussian process regression, it requires minimal parameter tuning and is able to predict the acceleration that will be expected at a distance in the future. The method produces estimates within 12ms without any code optimisation and the results indicate good prediction ability with mean errors within 4–10cm/s2on a database of various obstacles. John Bartholomew, Andrew Calway, Walterio W. Mayol-Cuevas |
ICRA | 3 |
| 2014 | Recognition and reconstruction of transparent objects for augmented realityabstractDealing with real transparent objects for AR is challenging due to their lack of texture and visual features as well as the drastic changes in appearance as the background, illumination and camera pose change. The few existing methods for glass object detection usually require a carefully controlled environment, specialized illumination hardware or ignore information from different viewpoints. In this work, we explore the use of a learning approach for classifying transparent objects from multiple images with the aim of both discovering such objects and building a 3D reconstruction to support convincing augmentations. We extract, classify and group small image patches using a fast graph-based segmentation and employ a probabilistic formulation for aggregating spatially consistent glass regions. We demonstrate our approach via analysis of the performance of glass region detection and example 3D reconstructions that allow virtual objects to interact with them. Alan Torres-Gomez, Walterio W. Mayol-Cuevas |
ISMAR | 2 |
| 2013 | Topological Map Building and Path Estimation Using Global-appearance Image DescriptorsabstractVisual-based navigation has been a source of numerous researches in the field of mobile robotics. In this paper we present a topological map building and localization algorithm using wide-angle scenes. Global-appearance descriptors are used in order to optimally represent the visual information. First, we build a topological graph that represents the navigation environment. Each node of the graph is a different position within the area, and it is composed of a collection of images that covers the complete field of view. We use the information provided by a camera that is mounted on the mobile robot when it travels along some routes between the nodes in the graph. With this aim, we estimate the relative position of each node using the visual information stored. Once the map is built, we propose a localization system that is able to estimate the location of the mobile not only in the nodes but also on intermediate positions using the visual information. The approach has been evaluated and shows good performance in real indoor scenarios under realistic illumination conditions. Francisco Amorós, Luis Payá, Óscar Reinoso, Walterio W. Mayol-Cuevas, Andrew Calway |
ICINCO (2) | 4 |
| 2013 | Fast place recognition with plane-based mapsabstractThis paper presents a new method for recognizing places in indoor environments based on the extraction of planar regions from range data provided by a hand-held RGB-D sensor. We propose to build a plane-based map (PbMap) consisting of a set of 3D planar patches described by simple geometric features (normal vector, centroid, area, etc.). This world representation is organized as a graph where the nodes represent the planar patches and the edges connect planes that are close by. This map structure permits to efficiently select subgraphs representing the local neighborhood of observed planes, that will be compared with other subgraphs corresponding to local neighborhoods of planes acquired previously. To find a candidate match between two subgraphs we employ an interpretation tree that permits working with partially observed and missing planes. The candidates from the interpretation tree are further checked out by a rigid registration test, which also gives us the relative pose between the matched places. The experimental results indicate that the proposed approach is an efficient way to solve this problem, working satisfactorily even when there are substantial changes in the scene (lifelong maps). Eduardo Fernández-Moral, Walterio W. Mayol-Cuevas, Vicente Arévalo, Javier González 0001 |
ICRA | 2 |
| 2013 | Enhancing 6D visual relocalisation with depth camerasabstractRelocalisation in 6D is relevant to a variety of Robotics applications and in particular to agile cameras exploring a 3D environment. While the use of geometry has commonly helped to validate appearance as a back-end process in several relocalisation systems before, we are interested in using 3D information to assist fast pose relocalisation computation as part of a front-end task. Our approach rapidly searches for a reduced number of visual descriptors, previously observed and stored in a database, that can be used to effectively compute the camera pose corresponding to the current view. We guide the search by means of constructing validated candidate sets using a 3D test involving the depth information obtained with an RGB-D camera (e.g. stereo of with structured light). Our experiments demonstrate that this process returns a compact quality set that works better for the pose estimation stage than when using a typical Nearest-Neighbor search over appearance only. The improvements are observed in terms of percentage of relocalised frames and speed, where the latter goes up to two orders of magnitude w.r.t. the conventional search. José Martínez-Carranza, Andrew Calway, Walterio W. Mayol-Cuevas |
IROS | 3 |
| 2012 | Real-time Learning and Detection of 3D Texture-less Objects: A Scalable ApproachabstractWe present a method for the learning and detection of multiple rigid texture-less 3D objects intended to operate at frame rate speeds for video input. The method is geared for fast and scalable learning and detection by combining tractable extraction of edgelet constellations with library lookup based on rotation- and scale-invariant descriptors. The approach learns object views in real-time, and is generative- enabling more objects to be learnt without the need for re-training. During testing, a random sample of edgelet constellations is tested for the presence of known objects. We perform testing of single and multi-object detection on a 30 objects dataset showing detections of any of them within milliseconds from the object’s visibility. The results show the scalability of the approach and its framerate performance. 1 Dima Damen, Pished Bunnun, Andrew Calway, Walterio W. Mayol-Cuevas |
BMVC | 4 |
| 2012 | 6D Relocalisation for RGBD Cameras Using Synthetic View RegressionabstractWith the advent of real-time dense scene reconstruction from handheld cameras, one key aspect to enable robust operation is the ability to relocalise in a previously mapped environment or after loss of measurement. Tasks such as operating on a workspace, where moving objects and occlusions are likely, require a recovery competence in order to be useful. For RGBD cameras, this must also include the ability to relocalise in areas with reduced visual texture. This paper describes a method for relocalisation of a freely moving RGBD camera in small workspaces. The approach combines both 2D image and 3D depth information to estimate the full 6D camera pose. The method uses a general regression over a set of synthetic views distributed throughout an informed estimate of possible camera viewpoints. The resulting relocalisation is accurate and works faster than framerate and the system’s performance is demonstrated through a comparison against visual and geometric feature matching relocalisation techniques on sequences with moving objects and minimal texture. 1 Andrew P. Gee, Walterio W. Mayol-Cuevas |
BMVC | 2 |
| 2012 | MUSTARD: a multi user see through AR displayabstractWe present MUSTARD, a multi-user dynamic random hole see-through display, capable of delivering viewer dependent information for objects behind a glass cabinet. Multiple viewers are allowed to observe both the physical object(s) being augmented and their location dependent annotations at the same time. The system consists of two liquid-crystal (LC) panels within which physical objects can be placed. The back LC panel serves as a dynamic mask while the front panel serves as the data. We first describe the principle of MUSTARD and then examine various functions that can be used to minimize crosstalk between multiple viewer positions. We compare different conflict management strategies using PSNR and the quality mean opinion score of HDR-VDP2. Finally, through a user-study we show that users can clearly identify images and objects even when the images are shown with strong conflicting regions; demonstrating that our system works even in the most extreme of circumstances. Abhijit Karnik, Walterio W. Mayol-Cuevas, Sriram Subramanian |
CHI | 2 |
| 2012 | What are we doing here? Egocentric activity recognition on the move for contextual mappingabstractAimed at contextual mapping of environments by exploration, this paper proposes a method that recognises human activity observed from a moving camera and references this information to a previously mapped environment. We first introduce a novel method that uses sparse features and dense optical flow, to perform dense background subtraction for an agile camera. With the ego-motion disambiguated, we present a method capable of recognising both external and egocentric human activity. When combined with visual simultaneous localisation and mapping (SLAM), this enables augmentation of visual maps with activity tags, highlighting areas of interest within large environments. Association of activity with location introduces the contextual element of purpose to each area of interest. Sudeep Sundaram, Walterio W. Mayol-Cuevas |
ICRA | 2 |
| 2012 | Predicting Micro Air Vehicle landing behaviour from visual textureabstractWe introduce a framework to predict the landing behaviour of a Micro Air Vehicle (MAV) from the appearance of the landing surface. We approach this problem by learning a mapping from visual texture observed from an onboard camera to the landing behaviour on a set of sample materials. In this case we exemplify our framework by predicting the yaw angle of the MAV after landing. Our framework demonstrates the applicability of established texture classification methods usually tested on stationary camera setups for the more challenging case of textures observed from a MAV. Results for supervised training demonstrate good estimation of the landing behaviour and motivate future work to implement autonomous decision making strategies and other behaviour predictions based on imagery. John Bartholomew, Andrew Calway, Walterio W. Mayol-Cuevas |
IROS | 3 |
| 2012 | Egocentric Real-time Workspace Monitoring using an RGB-D cameraabstractWe describe an integrated system for personal workspace monitoring based around an RGB-D sensor. The approach is egocentric, facilitating full flexibility, and operates in real-time, providing object detection and recognition, and 3D trajectory estimation whilst the user undertakes tasks in the workspace. A prototype on-body system developed in the context of work-flow analysis for industrial manipulation and assembly tasks is described. The system is evaluated on two tasks with multiple users, and results indicate that the method is effective, giving good accuracy performance. Dima Damen, Andrew P. Gee, Walterio W. Mayol-Cuevas, Andrew Calway |
IROS | 3 |
| 2012 | Integrating 3D object detection, modelling and tracking on a mobile phoneabstractThis paper presents a complete system on a camera phone that integrates a texture less 3D object detector together with in-situ modelling and tracking. The result is a suite intended for AR applications on the move where objects can be captured, tracked and with the automated detection providing a bridge to either initialize tracking or resume it after measurement loss. The object detector training is online and in-situ and benefits from tight integration with the modelling and tracking processes. Pished Bunnun, Dima Damen, Andrew Calway, Walterio W. Mayol-Cuevas |
ISMAR | 4 |
| 2012 | PiVOT: personalized view-overlays for tabletopsabstractWe present PiVOT, a tabletop system aimed at supporting mixed-focus collaborative tasks. Through two view-zones, PiVOT provides personalized views to individual users while presenting an unaffected and unobstructed shared view to all users. The system supports multiple personalized views which can be present at the same spatial location and yet be only visible to the users it belongs to. The system also allows the creation of personal views that can be either 2D or (auto-stereoscopic) 3D images. We first discuss the motivation and the different implementation principles required for realizing such a system, before exploring different designs able to address the seemingly opposing challenges of shared and personalized views. We then implement and evaluate a sample prototype to validate our design ideas and present a set of sample applications to demonstrate the utility of the system. Abhijit Karnik, Diego Martínez 0001, Walterio W. Mayol-Cuevas, Sriram Subramanian |
UIST | 3 |
| 2011 | Sensor suites for assistive arm prostheticsabstractThis paper introduces a sensor suite framework for the partial automation of prosthetic arm control allowing high level control with a reduction of cognitive burden placed upon the user. Automation aims to replicate the hand eye co-ordination through the synergy of a virtual 7DOF arm prosthesis together with the development of a gaze tracking system. The interactions between elements of the suite are detailed and a selection of sensors implemented to control a simple simulation. Use of the novel tongue control system is used to provide discrete input to the system. Initial tests are made of of the system together with a users ability to learn to use the system with promising user feedback on ease of interaction and potential for reduced cognitive burden. Martin Buckley, Ravi Vaidyanathan, Walterio W. Mayol-Cuevas |
CBMS | 3 |
| 2011 | A topometric system for wide area augmented reality
Andrew P. Gee, Matthew Webb, Ponciano Jorge Escamilla-Ambrosio, Walterio W. Mayol-Cuevas, Andrew Calway |
Comput. Graph. | 4 |
| 2011 | Adaptive Sampling for Feature Detection, Tracking, and Recognition on Mobile PlatformsabstractLocal image features have become ubiquitous for a wide range of computer vision tasks. For embedded and low power devices, speed and memory efficiency is of main concern, and therefore, there have been several recent attempts to improve these issues. In this paper, we are concerned with the early components of the object recognition pipeline, namely, feature detection, feature description, and feature tracking. In particular, we propose a novel approach to speed up feature detectors and to inform feature tracking that speeds up the recognition process by using the concept of adaptive sampling. We select two examples of visual algorithms to be modified by adaptive sampling and present comparative results with and without modifications. We show how processing time and memory footprint can benefit by this approach with little impact on overall output quality. We implement our proposed methods on a chipset commonly found on smartphones and we discuss the obtained improvements. Mosalam Ebrahimi, Walterio W. Mayol-Cuevas |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Application of multiple-wireless to a visual localisation system for emergency servicesabstractIn this paper we discuss the application of multiple-wireless technology to a practical context-enhanced service system called ViewNet. ViewNet develops technologies to support enhanced coordination and cooperation between operation teams in the emergency services and the police. Distributed localisation of users and mapping of environments implemented over a secure wireless network enables teams of operatives to search and map an incident area rapidly and in full coordination with each other and with a control centre. Sensing is based on fusing absolute positioning systems (UWB and GPS) with relative localisation and mapping from on-body or hand-held vision and inertial sensors. This paper focuses on the case for multiple-wireless capabilities in such a system and the benefits it can provide. We describe our work of developing a software API to support both WLAN and TETRA in ViewNet. It also provides a basis for incorporating future wireless technologies into ViewNet. Costas Efthymiou, Sedat Görmüs, Zhong Fan, Andrew Calway, Walterio W. Mayol-Cuevas, Angela Doufexi |
PIMRC | 5 |
| 2009 | Improving Image Sets Through Sense Disambiguation and Context RankingabstractCurrent approaches to automatic, class specific, image retrieval from the World Wide Web (WWW) by linguistic query often make use of an image's internal characteristics and file meta-data to augment and improve result accuracy. We propose that, in extension, improvement can be achieved in relevance, noise-reduction and completeness through sense disambiguation and contextual meta-data prepossessing. Our schemes exploits a linguistic ontology identifying query relevant homographs used to construct sense specific keyword sets allowing for enhanced image search and result ranking via the calculation of relatedness between query homographs and image context prior to any additional filtering. Within the paper we investigate different schemes for keyword set construction; ontology exclusive and authority extended, along with three differing ranking mechanisms. Anthony Roger Buck, Walterio W. Mayol-Cuevas |
SMC | 2 |
| 2009 | On the Choice and Placement of Wearable Vision SensorsabstractThis paper discusses two of the most important design considerations for a wearable device with visual sensing: what kind of sensor to use and where to place it. While nature and computer vision have explored a wide range of imaging techniques, wearables have mostly viewed the world through conventional narrow-view passive cameras designed for nonwearable applications, which are attached to the wearer's head. The rationale presented here for sensor selection and the novel methodology developed for objectively studying sensor placement have informed the development of a number of visual wearables. Walterio W. Mayol-Cuevas, Ben Tordoff, David William Murray 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2008 | Appearance Based Indexing for Relocalisation in Real-Time Visual SLAMabstractPrevious work on visual SLAM has shown that indexing on space and scale facilitates the use of feature descriptors for matching in real-time systems and that this can significantly increase robustness. However, the performance gains necessarily diminish as uncertainty about camera position increases. In this paper we address this issue by introducing a further level of indexing based on appearance, using low order Haar wavelet coefficients. This enables fast look up of descriptors even when the camera is lost, hence allowing efficient relocalisation. Results of experiments on a range of real world test cases demonstrate that the method is effective, including single frame relocalisation rates up to 90\% using relatively low numbers of descriptor comparisons. Denis Chekhlov, Walterio W. Mayol-Cuevas, Andrew Calway |
BMVC | 2 |
| 2008 | OutlinAR: an assisted interactive model building system with reduced computational effortabstractThis paper presents a system that allows online building of 3D wireframe models through a combination of user interaction and automated methods from a handheld camera-mouse. Crucially, the model being built is used to concurrently compute camera pose, permitting extendable tracking while enabling the user to edit the model interactively. In contrast to other model building methods that are either off-line and/or automated but computationally intensive, the aim here is to have a system that has low computational requirements and that enables the user to define what is relevant (and what is not) at the time the model is being built. OutlinAR hardware is also developed which simply consists of the combination of a camera with a wide field of view lens and a wheeled computer mouse. Pished Bunnun, Walterio W. Mayol-Cuevas |
ISMAR | 2 |
| 2008 | Tracking with general regression
Walterio W. Mayol-Cuevas, David William Murray 0001 |
Mach. Vis. Appl. | 1 |
| 2008 | Discovering Higher Level Structure in Visual SLAMabstractIn this paper, we describe a novel method for discovering and incorporating higher level map structure in a real-time visual simultaneous localization and mapping (SLAM) system. Previous approaches use sparse maps populated by isolated features such as 3-D points or edgelets. Although this facilitates efficient localization, it yields very limited scene representation and ignores the inherent redundancy among features resulting from physical structure in the scene. In this paper, higher level structure, in the form of lines and surfaces, is discovered concurrently with SLAM operation, and then, incorporated into the map in a rigorous manner, attempting to maintain important cross-covariance information and allow consistent update of the feature parameters. This is achieved by using a bottom-up process, in which subsets of low-level features are ldquofolded inrdquo to a parameterization of an associated higher level feature, thus collapsing the state space as well as building structure into the map. We demonstrate and analyze the effects of the approach for the cases of line and plane discovery, both in simulation and within a real-time system operating with a handheld camera in an office environment. Andrew P. Gee, Denis Chekhlov, Andrew Calway, Walterio W. Mayol-Cuevas |
IEEE Trans. Robotics | 4 |
| 2007 | Discovering Planes and Collapsing the State Space in Visual SLAMabstractRecent advances in real-time visual SLAM have been based primarily on mapping isolated 3-D points. This presents difficulties when seeking to extend operation to wide areas, as the system state becomes large, requiring increasing computational effort. In this paper we present a novel approach to this problem in which planar structural components are embedded within the state to represent mapped points lying on a common plane. This collapses the state size, reducing computation and improving scalability, as well as giving a higher level scene description. Critically, the plane parameters are augmented into the SLAM state in a proper fashion, maintaining inherent uncertainties via a full covariance representation. Results for simulated data and for real-time operation demonstrate that the approach is effective. 1 Andrew P. Gee, Denis Chekhlov, Walterio W. Mayol-Cuevas, Andrew Calway |
BMVC | 3 |
| 2007 | Robust Feature Descriptors for Efficient Vision-Based Tracking
Gerardo Carrera, Jesús Savage, Walterio W. Mayol-Cuevas |
CIARP | 3 |
| 2007 | Robust Real-Time Visual SLAM Using Scale Prediction and Exemplar Based Feature DescriptionabstractTwo major limitations of real-time visual SLAM algorithms are the restricted range of views over which they can operate and their lack of robustness when faced with erratic camera motion or severe visual occlusion. In this paper we describe a visual SLAM algorithm which addresses both of these problems. The key component is a novel feature description method which is both fast and capable of repeat-able correspondence matching over a wide range of viewing angles and scales. This is achieved in real-time by using a SIFT-like spatial gradient descriptor in conjunction with efficient scale prediction and exemplar based feature representation. Results are presented illustrating robust realtime SLAM operation within an office environment. Denis Chekhlov, Mark Pupilli, Walterio W. Mayol-Cuevas, Andrew Calway |
CVPR | 3 |
| 2007 | Ninja on a Plane: Automatic Discovery of Physical Planes for Augmented Reality Using Visual SLAMabstractMost work in visual augmented reality (AR) employs predefined markers or models that simplify the algorithms needed for sensor positioning and augmentation but at the cost of imposing restrictions on the areas of operation and on interactivity. This paper presents a simple game in which an AR agent has to navigate using real planar surfaces on objects that are dynamically added to an unprepared environment. An extended Kalman filter (EKF) simultaneous localisation and mapping (SLAM) framework with automatic plane discovery is used to enable the player to interactively build a structured map of the game environment using a single, agile camera. By using SLAM, we are able to achieve real-time interactivity and maintain rigorous estimates of the system's uncertainty, which enables the effects of high quality estimates to be propagated to other features (points and planes) even if they are outside the camera's current field of view. Denis Chekhlov, Andrew P. Gee, Andrew Calway, Walterio W. Mayol-Cuevas |
ISMAR | 4 |
| 2004 | Interaction between hand and wearable camera in 2d and 3d environmentsabstractThis paper is concerned with allowing the user of a wearable, portable, vision system to interact with the visual information using hand movements and gestures. Two example scenarios are explored. The first, in 2D, uses the wearer’s hand to both guide an active wearable camera and to highlight objects of interest using a grasping vector. The second is based in 3D, and builds on earlier work which recovers 3D scene structure at video-rate, allowing real-time purposive redirection of the camera to any scene point. Here, a range of hand gestures are used to highlight and select 3D points within the structure and in this instance used to insert 3D graphical objects into the scene. Structure recovery, gesture recognition, scene annotation and augmentation are achieved in parallel and at video-rate. Walterio W. Mayol-Cuevas, Andrew J. Davison, Ben Tordoff, Nicholas Molton, David William Murray 0001 |
BMVC | 1 |
| 2003 | Real-Time Localisation and Mapping with Wearable Active VisionabstractWe present a general method for real-time, vision-only single-camera simultaneous localisation and mapping (SLAM) - an algorithm which is applicable to the localisation of any camera moving through a scene - and study its application to the localisation of a wearable robot with active vision. Starting from very sparse initial scene knowledge, a map of natural point features spanning a section of a room is generated on-the-fly as the motion of the camera is simultaneously estimated in full 3D. Naturally this permits the annotation of the scene with rigidly-registered graphics, but further it permits automatic control of the robot's active camera: for instance, fixation on a particular object can be maintained during extended periods of arbitrary user motion, then shifted at will to another object which has potentially been out of the field of view. This kind of functionality is the key to the understanding or "management" of a workspace which the robot needs to have in order to assist its wearer usefully in tasks. We believe that the techniques and technology developed are of particular immediate value in scenarios of remote collaboration, where a remote expert is able to annotate, through the robot, the environment the wearer is working in. Andrew J. Davison, Walterio W. Mayol-Cuevas, David William Murray 0001 |
ISMAR | 2 |
| 2003 | Real-Time Visual Workspace Localisation and Mapping for a Wearable RobotabstractThis demo showcases breakthrough results in the general field real-time simultaneous localization and mapping (SLAM) using vision and in particular its vital role in enabling a wearable robot to assists its user. In our approach, a wearable active vision system ("wearable robot") is mounted at the shoulder. As the wearer moves around his environment, typically browsing a workspace in which a task must be completed, the robot acquires images continuously and generates a map of natural visual features on-the-fly while estimating its ego-motion. Andrew J. Davison, Walterio W. Mayol-Cuevas, David William Murray 0001 |
ISMAR | 2 |
| 2003 | Applying Active Vision and SLAM to Wearables
Walterio W. Mayol-Cuevas, Andrew J. Davison, Ben Tordoff, David William Murray 0001 |
ISRR | 1 |
| 2002 | Head pose estimation for wearable robot controlabstractRecent advances in wearable sensing allow active control of the orientation of a body-mounted camera worn by a remote user. In this paper we consider the control of the active camera from head movements. In the context of teleoperation, these may be the head movements of a remote operator, perhaps acting as the wearer’s assistant. The movements are likely to be larger than those in video-conference applications, and so frontal facial features are insufficient. The paper presents a model which incrementally combines a fixed 3D shape model with specific features found on the observed head. Robust methods, including the incorporation of a colour model, are used to mitigate the effect of mismatching, the main contribution being the use of both interest point and colour features within a single random-sampling framework. Ben Tordoff, Walterio W. Mayol-Cuevas, Teófilo Emídio de Campos, David William Murray 0001 |
BMVC | 2 |
| 2002 | Designing a Miniature Wearable Visual RobotabstractWe report on two methods we have developed to aid in the design of a wearable visual robot-a body mounted robot for which the main sensor is a camera. Specifically, we have first refined the analysis of sensor placement through the computation of the field of view and body motion using a 3D model of the human form. Second we have improved the design of the robot's morphology with the help of an optimization algorithm based on the Pareto front, within constraints set by the overall choice of robot kinematic chain and the need to specify obtainable actuators and sensors. The methods could be of use for the design and performance evaluation of rather different kinds of wearable robots and devices. Walterio W. Mayol-Cuevas, Ben Tordoff, David William Murray 0001 |
ICRA | 1 |
| 2002 | Wearable Visual Robots
Walterio W. Mayol-Cuevas, Ben Tordoff, David William Murray 0001 |
Pers. Ubiquitous Comput. | 1 |
| 2000 | Towards wearable active vision platformsabstractThe paper describes the design and construction of a wearable active vision platform which is able to achieve substantial decoupling of the camera motion from the wearer's motion. Design issues in sensor placement, robot kinematics and their relation to wearability are discussed and the prototype platform's performance is evaluated in a number of important visual tasks. The paper also considers potential application scenarios for this kind of wearable visual robot. Walterio W. Mayol-Cuevas, Ben Tordoff, David William Murray 0001 |
SMC | 1 |
| 1998 | Design of a walking machine structure using evolutionary strategiesabstractCommonly, the design of a walking machine (WM) use to imitate the nature like multilegged creatures, which try to do the best approaching to a spider, a biped, quadruped or any other living animal. Trying to imitate the way that nature has generated animals seems to be more interesting and promising than imitating special results of the natural process of evolution; if we could imitate the motor that generates walking beings, then we may found better solutions, even in the scenarios where life doesn't exist. In this paper we present an algorithm using evolutionary strategies, to aid in key steps of the design of a walking machine structure. Mathematical models are developed together with objective functions. J. Juarez-Guerrero, Stalin Muñoz Gutiérrez, Walterio W. Mayol-Cuevas |
SMC | 3 |
| 1998 | A first approach to tactile texture recognitionabstractTactile texture recognition seems to be a very important research area with various interesting applications which include medical, geological and autonomous mobile robotics; despite the great amount of tasks that biological beings solve using the sense of touch, not much work is done in this area (except for pressure-related sensing), and there are only a few papers in literature that uses dynamic tactile sensing strategies, i.e. the kind of exploration that we and the biological beings use to recognize a texture. In this paper we present a system for tactile texture recognition, using sound-understanding techniques. We develop a sensing "pen" with an electret piezoelectric microphone, covered by a rugged material. The pen is rubbed over the material that we want to identify, the sound produced is segmented, the FFT is obtained and the result is introduced to a learning vector quantization technique (LVQ). We explore 18 common materials which includes surfaces from glass to a real human beard. We achieve more than 93% of recognition over the 18 textures, and when the system makes an error it gives a similar texture as result. Walterio W. Mayol-Cuevas, J. Juarez-Guerrero, Stalin Muñoz Gutiérrez |
SMC | 1 |