EDBT 2026 Demo / reviewers in the wild / expert
Patricio A. Vela
dblp:65/1102
· DBLP profile ↗
71ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0002-6888-7002ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 1 first-author · 21 since 2021Systems, architecture and hardware · 37 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 since 2021Databases, data management, data science and information retrieval · 6Applied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Schema-Guided Scene-Graph Reasoning Based on Multi-Agent Large Language Model SystemabstractScene graphs have emerged as a structured and serializable environment representation for grounded spatial reasoning with Large Language Models (LLMs). In this work, we propose SG2, an iterative Schema-Guided Scene-Graph reasoning framework based on multi-agent LLMs. The agents are grouped into two modules: a (1) Reasoner module for abstract task planning and graph information queries generation, and a (2) Retriever module for extracting corresponding graph information based on code-writing following the queries. Two modules collaborate iteratively, enabling sequential reasoning and adaptive attention to graph information. The scene graph schema, prompted to both modules, serves to not only streamline both reasoning and retrieval process, but also guide the cooperation between two modules. This eliminates the need to prompt LLMs with full graph data, reducing the chance of hallucination due to irrelevant information. Through experiments in multiple simulation environments, we show that our framework surpasses existing LLM-based approaches and baseline single-agent, tool-based Reason-while-Retrieve strategy in numerical Q&A and planning tasks. Yiye Chen, Harpreet Sawhney, Nicholas Gyde, Yanan Jian, Jack Saunders, Patricio A. Vela, Ben Lundell |
AAAI | 6 |
| 2025 | Dynamic Gap: Safe Gap-based Navigation in Dynamic EnvironmentsabstractThis paper extends the family of gap-based local planners to unknown dynamic environments through generating provably collision-free properties for hierarchical navigation systems. Existing perception-informed local planners that operate in dynamic environments rely on emergent or empirical robustness for collision avoidance as opposed to performing formal analysis of dynamic obstacles. In addition to this, the obstacle tracking that is performed in these existent planners is often achieved with respect to a global inertial frame, subjecting such tracking estimates to transformation errors from odometry drift. The proposed local planner, dynamic gap, shifts the tracking paradigm to modeling how the free space, represented as gaps, evolves over time. Gap crossing and closing conditions are developed to aid in determining the feasibility of passage through gaps, and a breadth of simulation benchmarking is performed against other navigation planners in the literature where the proposed dynamic gap planner achieves the highest success rate out of all planners tested in all environments. Max Asselmeier, Dhruv Ahuja, Abdel Zaro, Ahmad Abuaish, Ye Zhao 0002, Patricio A. Vela |
ICRA | 6 |
| 2025 | Task-driven SLAM Benchmarking for Robot NavigationabstractA critical use case of SLAM for mobile robots is to support localization during task-directed navigation. Current SLAM benchmarks overlook the importance of repeatability (precision) despite its impact on real-world deployments. TaskSLAM-Bench, a task-driven approach to SLAM benchmarking, addresses this gap. It employs precision as a key metric, accounts for SLAM’s mapping capabilities, and has easy-to-meet requirements. Simulated and real-world evaluation of SLAM methods provide insights into the navigation performance of modern visual and LiDAR SLAM solutions. The outcomes show that passive stereo SLAM precision may match that of 2D LiDAR SLAM in indoor environments. TaskSLAM-Bench complements existing benchmarks and offers richer assessment of SLAM performance in navigation-focused scenarios. Publicly available code permits in-situ SLAM testing in custom environments with properly equipped robots. Yanwei Du, Shiyu Feng, Carlton G. Cort, Patricio A. Vela |
IROS | 4 |
| 2025 | OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGBabstractTo address the challenge of short-term object pose tracking in dynamic environments with monocular RGB input, we introduce a large-scale synthetic dataset Omni-Pose6D, crafted to mirror the diversity of real-world conditions. We additionally present a benchmarking framework for a comprehensive comparison of pose tracking algorithms. We propose a pipeline featuring an uncertainty-aware keypoint refinement network, employing probabilistic modeling to refine pose estimation. Comparative evaluations demonstrate that our approach achieves performance superior to existing baselines on real datasets, underscoring the effectiveness of our synthetic dataset and refinement technique in enhancing tracking precision in dynamic contexts. Our contributions set a new precedent for the development and assessment of object pose tracking methodologies in complex scenes. Yunzhi Lin, Yipu Zhao, Fu-Jen Chu, Weiyao Wang 0001, Patricio A. Vela, Matt Feiszli, Kevin J. Liang |
IROS | 7 |
| 2024 | Hierarchical Experience-informed Navigation for Multi-modal Quadrupedal Rebar Grid TraversalabstractThis study focuses on a layered, experience-based, multi-modal contact planning framework for agile quadrupedal locomotion over a constrained rebar environment. To this end, our hierarchical planner incorporates locomotion-specific modules into the high-level contact sequence planner and performs kinodynamically-aware trajectory optimization as the low-level motion planner. Through quantitative analysis of the experience accumulation process and experimental validation of the kinodynamic feasibility of the generated locomotion trajectories, we demonstrate that the planning heuristic of experience offers an effective way of providing candidate footholds for a legged contact planner. Additionally, we introduce a guiding torso path heuristic at the global planning level to enhance the navigation success rate in the presence of environmental obstacles. Our results indicate that the torso-path guided experience accumulation requires significantly fewer offline trials to successfully reach the goal compared to regular experience accumulation. Finally, our planning framework is validated in both dynamics simulations and real hardware implementations on a quadrupedal robot provided by Skymul Inc. Max Asselmeier, Jane Ivanova, Ziyi Zhou 0004, Patricio A. Vela, Ye Zhao 0002 |
ICRA | 4 |
| 2023 | WDiscOOD: Out-of-Distribution Detection via Whitened Linear Discriminant AnalysisabstractDeep neural networks are susceptible to generating overconfident yet erroneous predictions when presented with data beyond known concepts. This challenge underscores the importance of detecting out-of-distribution (OOD) samples in the open world. In this work, we propose a novel feature-space OOD detection score based on class-specific and class-agnostic information. Specifically, the approach utilizes Whitened Linear Discriminant Analysis to project features into two subspaces the discriminative and residual subspaces - for which the in-distribution (ID) classes are maximally separated and closely clustered, respectively. The OOD score is then determined by combining the deviation from the input data to the ID pattern in both subspaces. The efficacy of our method, named WDiscOOD, is verified on the large-scale ImageNet-1k benchmark, with six OOD datasets that cover a variety of distribution shifts. WDiscOOD demonstrates superior performance on deep classifiers with diverse backbone architectures, including CNN and vision transformer. Furthermore, we also show that WDiscOOD more effectively detects novel concepts in representation spaces trained with contrastive objectives, including supervised contrastive loss and multi-modality contrastive loss. Yiye Chen, Yunzhi Lin, Ruinian Xu, Patricio A. Vela |
ICCV | 4 |
| 2023 | Planning with Sequence Models through Iterative Energy Minimization
Yilun Du, Yiye Chen, Josh Tenenbaum, Patricio A. Vela |
ICLR | 5 |
| 2023 | Keypoint-GraspNet: Keypoint-based 6-DoF Grasp Generation from the Monocular RGB-D inputabstractThe success of 6-DoF grasp learning with point cloud input is tempered by the computational costs resulting from their unordered nature and pre-processing needs for reducing the point cloud to a manageable size. These properties lead to failure on small objects with low point cloud cardinality. Instead of point clouds, this manuscript explores grasp generation directly from the RGB-D image input. The approach, called Keypoint-GraspNet (KGN), operates in perception space by detecting projected gripper keypoints in the image, then recovering their SE(3) poses with a$\mathrm{P}n\mathrm{P}$algorithm. Training of the network involves a synthetic dataset derived from primitive shape objects with known continuous grasp families. Trained with only single-object synthetic data, Keypoint-GraspNet achieves superior result on our single-object dataset, comparable performance with state-of-art baselines on a multi-object test set, and outperforms the most competitive baseline on small objects. Keypoint-GraspNet is more than 3x faster than tested point cloud methods. Robot experiments show high success rate, demonstrating KGN's practical potential. Yiye Chen, Yunzhi Lin, Ruinian Xu, Patricio A. Vela |
ICRA | 4 |
| 2023 | GPF-BG: A Hierarchical Vision-Based Planning Framework for Safe Quadrupedal NavigationabstractSafe quadrupedal navigation through unknown environments is a challenging problem. This paper proposes a hierarchical vision-based planning framework (GPF-BG) integrating our previous Global Path Follower (GPF) navigation system and a gap-based local planner using Bézier curves, so called$B$ézier Gap (BG). This BG-based trajectory synthesis can generate smooth trajectories and guarantee safety for point-mass robots. With a gap analysis extension based on non-point, rectangular geometry, safety is guaranteed for an idealized quadrupedal motion model and significantly improved for an actual quadrupedal robot model. Stabilized perception space improves performance under oscillatory internal body motions that impact sensing. Simulation-based and real experiments under different benchmarking configurations test safe navigation performance. GPF-BG has the best safety outcomes across all experiments. Shiyu Feng, Ziyi Zhou 0004, Justin S. Smith, Max Asselmeier, Ye Zhao 0002, Patricio A. Vela |
ICRA | 6 |
| 2023 | Parallel Inversion of Neural Radiance Fields for Robust Pose EstimationabstractWe present a parallelized optimization method based on fast Neural Radiance Fields (NeRF) for estimating 6-DoF pose of a camera with respect to an object or scene. Given a single observed RGB image of the target, we can predict the translation and rotation of the camera by minimizing the residual between pixels rendered from a fast NeRF model and pixels in the observed image. We integrate a momentum-based camera extrinsic optimization procedure into Instant Neural Graphics Primitives, a recent exceptionally fast NeRF implementation. By introducing parallel Monte Carlo sampling into the pose estimation task, our method overcomes local minima and improves efficiency in a more extensive search space. We also show the importance of adopting a more robust pixel-based loss function to reduce error. Experiments demonstrate that our method can achieve improved generalization and robustness on both synthetic and real-world benchmarks. Yunzhi Lin, Thomas Müller 0013, Jonathan Tremblay, Bowen Wen, Stephen Tyree, Alex Evans, Patricio A. Vela, Stanley T. Birchfield |
ICRA | 7 |
| 2023 | AeriaLPiPS: A Local Planner for Aerial Vehicles with Geometric Collision CheckingabstractReal-time navigation in non-trivial environments by micro aerial vehicles (MAVs) predominantly relies on modelling the MAV with idealized geometry, such as a sphere. Simplified, conservative representations increase the likelihood of a planner failing to identify valid paths. That likelihood increases the more a robot's geometry differs from the idealized version. Few current approaches consider these situations; we are unaware of any that do so using perception space representations. This work introduces the egocan, a perception space obstacle representation using line-of-sight free space estimates, and 3D Gap, a perception space approach to gap finding for identifying goal-directed, collision-free directions of travel through 3D space. Both are integrated, with real-time considerations in mind, to define a local planner module of a hierarchical navigation system. The result is Aerial Local Planning in Perception Space (AeriaLPiPS). AeriaLPiPS is shown to be capable of safely navigating a MAV with non-idealized geometry through various environments, including those impassable by traditional real-time approaches. The open source implementation of this work is available at github.com/ivaROS/AeriaLPiPS. Justin S. Smith, Patricio A. Vela |
ICRA | 2 |
| 2023 | KGNv2: Separating Scale and Pose Prediction for Keypoint-Based 6-DoF Grasp Synthesis on RGB-D InputabstractWe propose an improved keypoint approach for 6-DoF grasp pose synthesis from RGB-D input. Keypoint-based grasp detection from image input demonstrated promising results in a previous study, where the visual information provided by color imagery compensates for noisy or imprecise depth measurements. However, it relies heavily on accurate keypoint prediction in image space. We devise a new grasp generation network that reduces the dependency on precise keypoint estimation. Given an RGB-D input, the network estimates both the grasp pose and the camera-grasp length scale. Re-design of the keypoint output space mitigates the impact of keypoint prediction noise on Perspective-n-Point (PnP) algorithm solutions. Experiments show that the proposed method outperforms the baseline by a large margin, validating its design. Though trained only on simple synthetic objects, our method demonstrates sim-to-real capacity through competitive results in real-world robot experiments. Yiye Chen, Ruinian Xu, Yunzhi Lin, Patricio A. Vela |
IROS | 5 |
| 2023 | Multi-Gait Locomotion Planning and Tracking for Tendon-Actuated Terrestrial Soft Robot (TerreSoRo)abstractThe adaptability of soft robots makes them ideal candidates to maneuver through unstructured environments. However, locomotion challenges arise due to complexities in modeling the body mechanics, actuation, and robot-environment dynamics. These factors contribute to the gap between their potential and actual autonomous field deployment. A closed-loop path planning framework for soft robot locomotion is critical to close the real-world realization gap. This paper presents a generic path planning framework applied to TerreSoRo (Tetra-Limb Terrestrial Soft Robot) with pose feedback. It employs a gait-based, lattice trajectory planner to facilitate navigation in the presence of obstacles. The locomotion gaits are synthesized using a data-driven optimization approach that allows for learning from the environment. The trajectory planner employs a greedy breadth-first search strategy to obtain a collision-free trajectory. The synthesized trajectory is a sequence of rotate-then-translate gait pairs. The control architecture integrates high-level and low-level controllers with real-time localization (using an overhead webcam). Terre-SoRo successfully navigates environments with obstacles where path re-planning is performed. To best of our knowledge, this is the first instance of real-time, closed-loop path planning of a non-pneumatic soft robot. Arun Niddish Mahendran, Caitlin Freeman, Alexander H. Chang, Michael McDougall, Patricio A. Vela, Vishesh Vikas |
IROS | 5 |
| 2022 | Keypoint-Based Category-Level Object Pose Tracking from an RGB Sequence with Uncertainty EstimationabstractWe propose a single-stage, category-level 6-DoF pose estimation algorithm that simultaneously detects and tracks instances of objects within a known category. Our method takes as input the previous and current frame from a monocular RGB video, as well as predictions from the previous frame, to predict the bounding cuboid and 6- DoF pose (up to scale). Internally, a deep network predicts distributions over object keypoints (vertices of the bounding cuboid) in image coordinates, after which a novel probabilistic filtering process integrates across estimates before computing the final pose using PnP. Our framework allows the system to take previous uncertainties into consideration when predicting the current frame, resulting in predictions that are more accurate and stable than single frame methods. Extensive experiments show that our method outperforms existing approaches on the challenging Objectron benchmark of annotated object videos. We also demonstrate the usability of our work in an augmented reality setting. Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, Stanley T. Birchfield |
ICRA | 4 |
| 2022 | Single-Stage Keypoint- Based Category-Level Object Pose Estimation from an RGB ImageabstractPrior work on 6-DoF object pose estimation has largely focused on instance-level processing, in which a textured CAD model is available for each object being detected. Category-level 6- DoF pose estimation represents an important step toward developing robotic vision systems that operate in unstructured, real-world scenarios. In this work, we propose a single-stage, keypoint-based approach for category-level object pose estimation that operates on unknown object instances within a known category using a single RGB image as input. The proposed network performs 2D object detection, detects 2D keypoints, estimates 6- DoF pose, and regresses relative bounding cuboid dimensions. These quantities are estimated in a sequential fashion, leveraging the recent idea of convGRU for propagating information from easier tasks to those that are more difficult. We favor simplicity in our design choices: generic cuboid vertex coordinates, single-stage network, and monocular RGB input. We conduct extensive experiments on the challenging Objectron benchmark, outperforming state-of-the-art methods on the 3D IoU metric (27.6% higher than the MobilePose single-stage approach and 7.1 % higher than the related two-stage approach). Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, Stanley T. Birchfield |
ICRA | 4 |
| 2021 | A Joint Network for Grasp Detection Conditioned on Natural Language CommandsabstractWe consider the task of grasping a target object based on a natural language command query. Previous work primarily focused on localizing the object given the query, which requires a separate grasp detection module to grasp it. The cascaded application of two pipelines incurs errors in overlapping multi-object cases due to ambiguity in the individal outputs. This work proposes a model named Command Grasping Network (CGNet) to directly output command satisficing grasps from RGB image and textual command inputs. A dataset with ground truth (image, command, grasps) tuple is generated based on the VMRD dataset to train the proposed network. Experimental results on the generated test set show that CGNet outperforms a cascaded object-retrieval and grasp detection baseline by a large margin. Three physical experiments demonstrate the functionality and performance of CGNet. Yiye Chen, Ruinian Xu, Yunzhi Lin, Patricio A. Vela |
ICRA | 4 |
| 2021 | Ego-centric Stereo Navigation Using Stixel WorldabstractThis paper explores the use of passive, stereo sensing for vision-based navigation. The traditional approach uses dense depth algorithms, which can be computationally costly or potentially inaccurate. These drawbacks compound when including the additional computational demands associated to the sensor fusion, collision checking, and path planning modules that interpret the dense depth measurements. These problems can be avoided through the use of the stixel representation, a compact and sparse visual representation for local free-space. When integrated into a Planning in Perception Space based hierarchical navigation framework, stixels permit fast and scalable navigation for different robot geometries. Computational studies quantify the processing performance and demonstrate the favorable scaling properties over comparable dense depth methods. Navigation benchmarking demonstrates more consistent performance across high and low performance compute hardware for PiPS-based stixel navigation versus traditional hierarchical navigation. Shiyu Feng, Fanzhe Lyu, Jin Ha Hwang, Patricio A. Vela |
ICRA | 4 |
| 2021 | Simultaneous Multi-Level Descriptor Learning and Semantic Segmentation for Domain-Specific RelocalizationabstractThis paper presents a semi-supervised framework for multi-level description learning aiming for robust and accurate camera relocalization across large perception variations. Our proposed network, namely DLSSNet, simultaneously learns weakly-supervised semantic segmentation and local feature description in the hierarchy. Therefore, the augmented descriptors, trained in an end-to-end manner, provide a more stable high-level representation for local feature dis-ambiguity. To facilitate end-to-end semantic description learning, the descriptor segmentation module is proposed to jointly learn semantic descriptors and cluster centers using standard semantic segmentation loss. We show that our model can be easily fine-tuned for domain-specific usage without any further semantic annotations, instead, requiring only 2D-2D pixel correspondences. The learned descriptors, trained with our proposed pipeline, can boost the cross-season localization performance against other state-of-the-arts. Yiye Chen, Cédric Pradalier, Patricio A. Vela |
ICRA | 4 |
| 2021 | Shape-centric Modeling for Soft Robot Inchworm LocomotionabstractSoft robot modeling tends to prioritize soft robot dynamics in order to recover how they might behave. Soft robot design tends to focus on how to use compliant elements with actuation to effect certain canonical movement profiles. For soft robot locomotors, these profiles should lead to locomotion. Naturally, there is a gap between the emphasis of computational modeling and the needs of locomotion design. This paper proposes to consider modeling and computation efforts directed more toward understanding soft robot-world interactions with locomotion in mind. With a SMA-actuated inchworm as the soft robot to model and control, the framework is a combination of shape identification and geometric modeling that culminates in control equations of motion. When applied to the task of gait-based locomotion, the equations operate in a low dimensional shape-based gait space. Simulated and experimentally applied gaits for an inchworm model showed qualitatively similar outcomes, while the measured net displacement per gait cycle coincided within 9%. This result advances the idea that a shape-centric approach to soft robot modeling for control and locomotion may provide predictive locomotive models. Alexander H. Chang, Caitlin Freeman, Arun Niddish Mahendran, Vishesh Vikas, Patricio A. Vela |
IROS | 5 |
| 2021 | Multi-view Fusion for Multi-level Robotic Scene UnderstandingabstractWe present a system for multi-level scene awareness for robotic manipulation. Given a sequence of camera-inhand RGB images, the system calculates three types of information: 1) a point cloud representation of all the surfaces in the scene, for the purpose of obstacle avoidance. 2) the rough pose of unknown objects from categories corresponding to primitive shapes (e.g., cuboids and cylinders), and 3) full 6-DoF pose of known objects. By developing and fusing recent techniques in these domains, we provide a rich scene representation for robot awareness. We demonstrate the importance of each of these modules, their complementary nature, and the potential benefits of the system in the context of robotic manipulation. Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, Stanley T. Birchfield |
IROS | 4 |
| 2021 | NavTuner: Learning a Scene-Sensitive Family of Navigation PoliciesabstractThe advent of deep learning has inspired research into end-to-end learning for a variety of problem domains in robotics. For navigation, the resulting methods may not have the generalization properties desired let alone match the performance of traditional methods. Instead of learning a navigation policy, we explore learning an adaptive policy in the parameter space of an existing navigation module. Having adaptive parameters provides the navigation module with a family of policies that can be dynamically reconfigured based on the local scene structure and addresses the common assertion in machine learning that engineered solutions are inflexible. Of the methods tested, reinforcement learning (RL) is shown to provide a significant performance boost to a modern navigation method through reduced sensitivity of its success rate to environmental clutter. The outcomes indicate that RL as a meta-policy learner, or dynamic parameter tuner, effectively robustifies algorithms sensitive to external, measurable nuisance factors. Haoxin Ma, Justin S. Smith, Patricio A. Vela |
IROS | 3 |
| 2020 | Using Synthetic Data and Deep Networks to Recognize Primitive Shapes for Object GraspingabstractA segmentation-based architecture is proposed to decompose objects into multiple primitive shapes from monocular depth input for robotic manipulation. The backbone deep network is trained on synthetic data with 6 classes of primitive shapes generated by a simulation engine. Each primitive shape is designed with parametrized grasp families, permitting the pipeline to identify multiple grasp candidates per shape primitive region. The grasps are priority ordered via proposed ranking algorithm, with the first feasible one chosen for execution. On task-free grasping of individual objects, the method achieves a 94% success rate. On task-oriented grasping, it achieves a 76% success rate. Overall, the method supports the hypothesis that shape primitives can support task-free and task-relevant grasp prediction. Yunzhi Lin, Chao Tang 0001, Fu-Jen Chu, Patricio A. Vela |
ICRA | 4 |
| 2020 | egoTEB: Egocentric, Perception Space Navigation Using Timed-Elastic-BandsabstractThe TEB hierarchical planner for real-time navigation through unknown environments is highly effective at balancing collision avoidance with goal directed motion. Designed over several years and publications, it implements a multi-trajectory optimization based synthesis method for identifying topologically distinct trajectory candidates through navigable space. Unfortunately, the underlying factor graph approach to the optimization problem induces a mismatch between grid-based representations and the optimization graph, which leads to several time and optimization inefficiencies. This paper explores the impact of using egocentric, perception space representations for the local planning map. Doing so alleviates many of the identified issues related to TEB and leads to a new method called egoTEB. Timing experiments and Monte Carlo evaluations in benchmark worlds quantify the benefits of egoTEB for navigation through uncertain environments. Justin S. Smith, Ruoyang Xu, Patricio A. Vela |
ICRA | 3 |
| 2020 | Closed-Loop Benchmarking of Stereo Visual-Inertial SLAM Systems: Understanding the Impact of Drift and Latency on Tracking AccuracyabstractVisual-inertial SLAM is essential for robot navigation in GPS-denied environments, e.g. indoor, underground. Conventionally, the performance of visual-inertial SLAM is evaluated with open-loop analysis, with a focus on the drift level of SLAM systems. In this paper, we raise the question on the importance of visual estimation latency in closed-loop navigation tasks, such as accurate trajectory tracking. To understand the impact of both drift and latency on visualinertial SLAM systems, a closed-loop benchmarking simulation is conducted, where a robot is commanded to follow a desired trajectory using the feedback from visual-inertial estimation. By extensively evaluating the trajectory tracking performance of representative state-of-the-art visual-inertial SLAM systems, we reveal the importance of latency reduction in visual estimation module of these systems. The findings suggest directions of future improvements for visual-inertial SLAM. Yipu Zhao, Justin S. Smith, Sambhu H. Karumanchi, Patricio A. Vela |
ICRA | 4 |
| 2020 | Synthesis of Control Barrier Functions Using a Supervised Machine Learning ApproachabstractControl barrier functions are mathematical constructs used to guarantee safety for robotic systems. When integrated as constraints in a quadratic programming optimization problem, instantaneous control synthesis with real-time performance demands can be achieved for robotics applications. Prevailing use has assumed full knowledge of the safety barrier functions, however there are cases where the safe regions must be estimated online from sensor measurements. In these cases, the corresponding barrier function must be synthesized online. This paper describes a learning framework for estimating control barrier functions from sensor data. Doing so affords system operation in unknown state space regions without compromising safety. Here, a support vector machine classifier provides the barrier function specification as determined by sets of safe and unsafe states obtained from sensor measurements. Theoretical safety guarantees are provided. Experimental ROS-based simulation results for an omnidirectional robot equipped with LiDAR demonstrate safe operation. Mohit Srinivasan, Amogh Dabholkar, Samuel Coogan 0001, Patricio A. Vela |
IROS | 4 |
| 2020 | Robust Monocular Edge Visual Odometry through Coarse-to-Fine Data AssociationabstractThis work describes a monocular visual odometry framework, which exploits the best attributes of edge features for illumination-robust camera tracking, while at the same time ameliorating the performance degradation of edge mapping. In the front-end, an ICP-based edge registration provides robust motion estimation and coarse data association under lighting changes. In the back-end, a novel edge-guided data association pipeline searches for the best photometrically matched points along geometrically possible edges through template matching, so that the matches can be further refined in later bundle adjustment. The core of our proposed data association strategy lies in a point-to-edge geometric uncertainty analysis, which analytically derives (1) a probabilistic search length formula that significantly reduces the search space and (2) a geometric confidence metric for mapping degradation detection based on the predicted depth uncertainty. Moreover, a match confidence based patch size adaption strategy is integrated into our pipeline to reduce matching ambiguity. We present extensive analysis and evaluation of our proposed system on synthetic and real- world benchmark datasets under the influence of illumination changes and large camera motions, where our proposed system outperforms current state-of-art algorithms. Patricio A. Vela, Cédric Pradalier |
IROS | 2 |
| 2020 | Good Feature Matching: Toward Accurate, Robust VO/VSLAM With Low LatencyabstractAnalysis of state-of-the-art visual odometry/visual simultaneous localization and mapping (VSLAM) system exposes a gap in balancing performance (accuracy and robustness) and efficiency (latency). Feature-based systems exhibit good performance, yet have higher latency due to explicit data association; direct and semidirect systems have lower latency, but are inapplicable in some target scenarios or exhibit lower accuracy than feature-based ones. This article aims to fill the performance-efficiency gap with an enhancement applied to feature-based VSLAM. We present good feature matching, an active map-to-frame feature matching method. Feature matching effort is tied to submatrix selection, which has combinatorial time complexity and requires choosing a scoring metric. Via simulation, the Max-logDet matrix revealing metric is shown to perform best. For real-time applicability, the combination of deterministic selection and randomized acceleration is studied. The proposed algorithm is integrated into monocular and stereo feature-based VSLAM systems. Extensive evaluations on multiple benchmarks and compute hardware quantify the latency reduction and the accuracy and robustness preservation. Yipu Zhao, Patricio A. Vela |
IEEE Trans. Robotics | 2 |
| 2019 | Every Hop is an Opportunity: Quickly Classifying and Adapting to Terrain During Targeted HoppingabstractPractical use of robots in diverse domains requires programming for, or adapting to, each domain and its unique characteristics. Failure to do so compromises the ability of the robot to achieve task-relevant objectives. Here we describe how the learned terrain reaction force profiles of a hopping robot serve the additional objectives of classifying terrain and quickly learning control strategies to accomplish a jumping task on novel terrain. We show that the reaction forces experienced during closed-loop jumping are sufficient to discriminate between three different terrain types (granular, trampoline, and rigid) when using the learned models as discriminators. Building on this, we show that applying the classification to unknown terrain types leads to faster task completion, where the task objective is to meet a specific jump height. The classification experiments, utilizing real-world jumping data, achieve 95% prediction accuracy. The online learning experiments leverage simulation as there is more control over the terrain properties. Terrain-informed learning achieves the target hop heights more than 2x faster than without terrain knowledge when the prediction is correct, and 1.5x faster when the prediction is incorrect. Thus, applying the closest approximately known terrain knowledge facilitates low shot learning when hopping on unknown terrain. Alexander H. Chang, Christian Hubicki, Aaron D. Ames, Patricio A. Vela |
ICRA | 4 |
| 2019 | Low-latency Visual SLAM with Appearance-Enhanced Local Map BuildingabstractA local map module is often implemented in modern VO/VSLAM systems to improve data association and pose estimation. Conventionally, the local map contents are determined by co-visibility. While co-visibility is cheap to establish, it utilizes the relatively-weak temporal prior (i.e. seen before, likely to be seen now), therefore admitting more features into the local map than necessary. This paper describes an enhancement to co-visibility local map building by incorporating a strong appearance prior, which leads to a more compact local map and latency reduction in downstream data association. The appearance prior collected from the current image influences the local map contents: only the map features visually similar to the current measurements are potentially useful for data association. To that end, mapped features are indexed and queried with Multi-index Hashing (MIH). An online hash table selection algorithm is developed to further reduce the query overhead of MIH and the local map size. The proposed appearance-based local map building method is integrated into a state-of-the-art VO/VSLAM system. When evaluated on two public benchmarks, the size of the local map, as well as the latency of real-time pose tracking in VO/VSLAM are significantly reduced. Meanwhile, the VO/VSLAM mean performance is preserved or improves. Yipu Zhao, Wenkai Ye, Patricio A. Vela |
ICRA | 3 |
| 2018 | Good Line Cutting: Towards Accurate Pose Tracking of Line-Assisted VO/VSLAM
Yipu Zhao, Patricio A. Vela |
ECCV (2) | 2 |
| 2018 | Hands-Free Assistive Manipulator Using Augmented Reality and Tongue Drive SystemabstractA human-in-the-loop system is proposed to enable hands-free collaborative manipulation for people with physical disabilities. Studies show that the cognitive burden of interfacing with a robotic assistant decreases with increased robot autonomy. Incorporating modern advances in perception with augmented reality, this paper describes a framework for obtaining high-level intents from the user to specify manipulation tasks for execution. Augmented reality glasses provide an egocentric perspective to the robot. The glasses also provide visual feedback to users on a virtual menu showing a summary of robot affordances. The system processes the vision input to interpret the users environment. A Tongue Drive System serves as the input modality for triggering task execution by the robotic arm. Several manipulation experiments are performed with comparison to Cartesian control. The outcomes are also compared to reported state-of-the-art approaches. The results demonstrate competitive performance with minimal user input requirements. Fu-Jen Chu, Ruinian Xu, Zhenxuan Zhang, Patricio A. Vela, Maysam Ghovanloo |
IROS | 4 |
| 2018 | Good Feature Selection for Least Squares Pose Optimization in VO/VSLAMabstractThis paper aims to select features that contribute most to the pose estimation in VO/VSLAM. Unlike existing feature selection works that are focused on efficiency only, our method significantly improves the accuracy of pose tracking, while introducing little overhead. By studying the impact of feature selection towards least squares pose optimization, we demonstrate the applicability of improving accuracy via good feature selection. To that end, we introduce the Max-logDet metric to guide the feature selection, which is connected to the conditioning of least squares pose optimization problem. We then describe an efficient algorithm for approximately solving the NP-hard Max-logDet problem. Integrating Max-logDet feature selection into a state-of-the-art visual SLAM system leads to accuracy improvements with low overhead, as demonstrated via evaluation on a public benchmark. Yipu Zhao, Patricio A. Vela |
IROS | 2 |
| 2017 | Learning to jump in granular media: Unifying optimal control synthesis with Gaussian process-based regressionabstractThe varied and complex dynamics of deformable terrain are significant impediments toward real-world viability of locomotive robotics, particularly for legged machines. We explore vertical jumping on granular media (GM) as a model task for legged locomotion on uncharacterized deformable terrain. By integrating (Gaussian process) GP-based regression and evaluation to estimate ground forcing as a function of state, a one-dimensional jumper acquires the ability to learn forcing profiles exerted by its environment in tandem to achieving its control objective. The GP-based dynamical model initially assumes a baseline rigid, non-compliant surface. As part of an iterative procedure, the optimizer employing this model generates an optimal control to achieve a target jump height while respecting known hardware limitations of the robot model. Trajectory and forcing data recovered from evaluation on the true GM surface model simulation is applied to train the GP, and in turn, provide the optimizer a more richly informed dynamical model of the environment. After three iterations, predicted optimal control trajectories coincide with execution results, within 1.2% jumping height error, as the GP-based approximation converges to the true GM model. Alexander H. Chang, Christian Hubicki, Jeff J. Aguilar, Daniel I. Goldman, Aaron D. Ames, Patricio A. Vela |
ICRA | 6 |
| 2017 | Closed-loop path following of traveling wave rectilinear motion through obstacle-strewn terrainabstractHigh-level, closed-loop traversal through obstacle-strewn environments remains an open and challenging endeavor for snake-like robotic platforms. Rectilinear forms of locomotion, despite their unique mobility advantages compared to other gait shapes, have seen relatively little progress toward this objective. A dynamical exposition of traveling wave rectilinear motion is reviewed and applied to generate a functional mapping from gait parameter space to corresponding averaged steady-behavior body velocities. We demonstrate the system dynamics, in this case, resemble that of a fixed forward-velocity unicycle where average body curvature presents itself as a versatile control input, modulating angular body velocity in a linear manner. Target body velocities computed to track non-trivial planned paths, then, are mapped to average body curvature commands, enabling autonomous path following, by a physical, multi-link robotic snake, through a series of distinct obstacle arrangements. Alexander H. Chang, Patricio A. Vela |
ICRA | 2 |
| 2017 | PiPS: Planning in perception spaceabstractPath planning for mobile robots requires rapidly finding collision-free trajectories in an uncertain and changing environment. Full collision checking with detailed, online-revised representations of the robot and world imposes a delay that undermines reactive obstacle avoidance. As a result, reactive vision-based approaches make various assumptions to arrive at simplified representations, such as circular or spherical robot shapes reducible to point masses, or obstacles that always rise from the ground. We seek to avoid these problems by modeling the robot directly in perception space so that collisionfree trajectories can be sought in a consistent representation with minimal processing needs. Here perception space refers to the depth space image measurements available by modern consumer range sensors. We hallucinate a robot navigating through the world and synthesize depth images of its path for comparison against the directly sensed depth images of the local world. The approach performs collision checking in a 3D volume but only requires 2D image comparisons. Experiments show that an implementation is able to negotiate an obstacle course consisting of miscellaneous objects in real-time. Justin S. Smith, Patricio A. Vela |
ICRA | 2 |
| 2016 | Learning binary features online from motion dynamics for incremental loop-closure detection and place recognitionabstractThis paper proposes a simple yet effective approach to learn visual features online for improving loop-closure detection and place recognition, based on bag-of-words frameworks. The approach learns a codeword in the bag-of-words model from a pair of matched features from two consecutive frames, such that the codeword has temporally-derived perspective invariance to camera motion. The learning algorithm is efficient: the binary descriptor is generated from the mean image patch, and the mask is learned based on discriminative projection by minimizing the intra-class distances among the learned feature and the two original features. A codeword is generated by packaging the learned descriptor and mask, with a masked Hamming distance defined to measure the distance between two codewords. The geometric properties of the learned codewords are then mathematically justified. In addition, hypothesis constraints are imposed through temporal consistency in matched codewords, which improves precision. The approach, integrated in an incremental bag-of-words system, is validated on multiple benchmark data sets and compared to state-of-the-art methods. Experiments demonstrate improved precision/recall outperforming state of the art with little loss in runtime. Guangcong Zhang, Mason J. Lilly, Patricio A. Vela |
ICRA | 3 |
| 2016 | A Stochastic Approach to Diffeomorphic Point Set Registration with Landmark ConstraintsabstractThis work presents a deformable point set registration algorithm that seeks an optimal set of radial basis functions to describe the registration. A novel, global optimization approach is introduced composed of simulated annealing with a particle filter based generator function to perform the registration. It is shown how constraints can be incorporated into this framework. A constraint on the deformation is enforced whose role is to ensure physically meaningful fields (i.e., invertible). Further, examples in which landmark constraints serve to guide the registration are shown. Results on 2D and 3D data demonstrate the algorithm's robustness to noise and missing information. Ivan Kolesov, Jehoon Lee, Gregory C. Sharp, Patricio A. Vela, Allen R. Tannenbaum |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Good features to track for visual SLAMabstractNot all measured features in SLAM/SfM contribute to accurate localization during the estimation process, thus it is sensible to utilize only those that do. This paper describes a method for selecting a subset of features that are of high utility for localization in the SLAM/SfM estimation process. It is derived by examining the observability of SLAM and, being complimentary to the estimation process, it easily integrates into existing SLAM systems. The measure of estimation utility is formulated with temporal and instantaneous observability indices. Efficient computation strategies for the observability indices are described based on incremental singular value decomposition (SVD) and greedy selection for the temporal and instantaneous observability indices, respectively. The greedy selection is near-optimal since the observability index is (approximately) submodular. The proposed method improves localization and data association. Controlled synthetic experiments with ground truth demonstrate the improved localization accuracy, and real-time SLAM experiments demonstrate the improved data association. Guangcong Zhang, Patricio A. Vela |
CVPR | 2 |
| 2015 | Incorporating frictional anisotropy in the design of a robotic snake through the exploitation of scalesabstractThe scales on the skin of a snake are an integral part of the snake's locomotive capabilities. It stands to reason that the integration of scales into the design of robotic snakes would open new properties to exploit. In this work, we present a robotic snake design that incorporates rigid scales in the casing of each module. To validate the impact of the scales, locomotion is tested under three conditions: with scales, covered in cloth with scales, and covered in cloth without scales. The performance of the snake robot, in each of the aforementioned scenarios, is evaluated based on its forward displacement while executing each of two pre-programmed gaits: inchworm and lateral undulation. The lateral undulation gait is tested under two additional conditions: pitched and un-pitched. Tracks of the experimental runs are presented followed by a statistical analysis demonstrating an increase in locomotive performance when incorporating scales in the chassis design. Miguel Moises Serrano, Alexander H. Chang, Guangcong Zhang, Patricio A. Vela |
ICRA | 4 |
| 2015 | Optimally observable and minimal cardinality monocular SLAMabstractThis paper utilizes system observability to guide monocular SLAM. Instead of providing all measured features then performing data-driven outlier rejection (such as with RANSAC), we propose to identify only the minimal subset of features which form an optimally observable SLAM subsystem for localization. Modeling the SLAM system as a discrete time system with piece-wise linear SE〈3〉 motion, complete observability conditions are derived and a means to test the observability conditioning of candidate feature point groupings is proposed. Based on the conditioning, an efficient algorithm for picking the optimally observable feature subset is derived by incorporating the image geometric measures. The proposed monocular SLAM algorithm, called Optimally Observable and Minimal Cardinality (OOMC) SLAM is formulated as an EKF process. OOMC SLAM is first validated using a 6-DOF localization experiment; the results demonstrate accuracy comparable to the state-of-art SLAM algorithm with significantly improved computational efficiency. A longer sequence on a 620-meter trajectory is also tested. The algorithm achieves 0.9178% relative error against the GPS ground truth. Guangcong Zhang, Patricio A. Vela |
ICRA | 2 |
| 2015 | Construction performance monitoring via still images, time-lapse photos, and video streams: Now, tomorrow, and the future
Man-Woo Park, Patricio A. Vela, Mani Golparvar Fard |
Adv. Eng. Informatics | 3 |
| 2015 | Bayesian Nonparametric Adaptive Control Using Gaussian ProcessesabstractMost current model reference adaptive control (MRAC) methods rely on parametric adaptive elements, in which the number of parameters of the adaptive element are fixed a priori, often through expert judgment. An example of such an adaptive element is radial basis function networks (RBFNs), with RBF centers preallocated based on the expected operating domain. If the system operates outside of the expected operating domain, this adaptive element can become noneffective in capturing and canceling the uncertainty, thus rendering the adaptive controller only semiglobal in nature. This paper investigates a Gaussian process-based Bayesian MRAC architecture (GP-MRAC), which leverages the power and flexibility of GP Bayesian nonparametric models of uncertainty. The GP-MRAC does not require the centers to be preallocated, can inherently handle measurement noise, and enables MRAC to handle a broader set of uncertainties, including those that are defined as distributions over functions. We use stochastic stability arguments to show that GP-MRAC guarantees good closed-loop performance with no prior domain knowledge of the uncertainty. Online implementable GP inference methods are compared in numerical simulations against RBFN-MRAC with preallocated centers and are shown to provide better tracking and improved long-term learning. Girish Chowdhary 0001, Hassan A. Kingravi, Jonathan P. How, Patricio A. Vela |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2014 | Forage RRT - An efficient approach to task-space goal planning for high dimensional systemsabstractAchieving efficient end-effector planning for manipulators in real world workspaces is challenging due to the fact that planning is performed in manipulator joint space, while the planning goal is given in end-effector or tool space. For manipulator planning, the problem becomes a joint path planning and inverse kinematics problem to be resolved efficiently, in spite of the potentially infinite number of inverse solutions to the end-effector goal state and the nonlinear relationship between joint configurations and world obstacles. The Forage-RRT algorithm described in this paper attempts to efficiently and quickly resolve the end-effector or tool planning problem. Using ideas from foraging theory, Forage-RRT implements a diffusion-based search strategy with two rates of diffusion, one high and one low, which simulate both long jumps (coarse search) and focused small area exploration (fine search) in the joint space, respectively. During coarse search, it is important to keep track of past fine searches, therefore the traditional RRT algorithm is augmented with a goal heap that keeps track of potential focused search regions and discards them when they result in failure. By mixing between two search distributions with different spreads, the search space is rapidly covered and potentially fruitful avenues are finely explored. The search coverage advantages of foraging identified in the biological research literature are demonstrated here for end-effector based, manipulator path planning. Leo Keselman, Erik I. Verriest, Patricio A. Vela |
ICRA | 3 |
| 2013 | Reduced Set KPCA for Improving the Training and Execution Speed of Kernel MachinesabstractThis paper presents a practical, and theoretically well-founded, approach to improve the speed of kernel manifold learning algorithms relying on spectral decomposition. Utilizing recent insights in kernel smoothing and learning with integral operators, we propose Reduced Set KPCA (RSKPCA), which also suggests an easy-to-implement method to remove or replace samples with minimal effect on the empirical operator. A simple data point selection procedure is given to generate a substitute density for the data, with accuracy that is governed by a user-tunable parameter ℓ. The effect of the approximation on the quality of the KPCA solution, in terms of spectral and operator errors, can be shown directly in terms of the density estimate error and as a function of the parameter ℓ. We show in experiments that RSKPCA can improve both training and evaluation time of KPCA by up to an order of magnitude, and compares favorably to the widely-used Nystrom and density-weighted Nystrom methods. Alexander G. Gray, Hassan A. Kingravi, Patricio A. Vela |
SDM | 3 |
| 2013 | Optimized selection of key frames for monocular videogrammetric surveying of civil infrastructure
Abbas Rashidi, Fei Dai 0003, Ioannis K. Brilakis, Patricio A. Vela |
Adv. Eng. Informatics | 4 |
| 2013 | A comparative study of efficient initialization methods for the k-means clustering algorithm
M. Emre Celebi 0001, Hassan A. Kingravi, Patricio A. Vela |
Expert Syst. Appl. | 3 |
| 2013 | Joint CT/CBCT deformable registration and CBCT enhancement for cancer radiotherapy
Yifei Lou, Tianye Niu, Xun Jia, Patricio A. Vela, Lei Zhu 0001, Allen R. Tannenbaum |
Medical Image Anal. | 4 |
| 2013 | Interactive Medical Image Segmentation Using PDE Control of Active ContoursabstractSegmentation of injured or unusual anatomic structures in medical imagery is a problem that has continued to elude fully automated solutions. In this paper, the goal of easy-to-use and consistent interactive segmentation is transformed into a control synthesis problem. A nominal level set partial differential equation (PDE) is assumed to be given; this open-loop system achieves correct segmentation under ideal conditions, but does not agree with a human expert's ideal boundary for real image data. Perturbing the state and dynamics of a level set PDE via the accumulated user input and an observer-like system leads to desirable closed-loop behavior. The input structure is designed such that a user can stabilize the boundary in some desired state without needing to understand any mathematical parameters. Effectiveness of the technique is illustrated with applications to the challenging segmentations of a patellar tendon in magnetic resonance and a shattered femur in computed tomography. Peter Karasev, Ivan Kolesov, Karl D. Fritscher, Patricio A. Vela, Phillip Mitchell, Allen R. Tannenbaum |
IEEE Trans. Medical Imaging | 4 |
| 2012 | A modified KLT multiple objects tracking framework based on global segmentation and adaptive template
Kang Xue, Patricio A. Vela, Yue Liu 0005, Yongtian Wang |
ICPR | 2 |
| 2012 | Information propagation applied to robot-assisted evacuationabstractInspired by large fatality rates due to fires in crowded areas and the increasing presence of robots in dangerous emergency situations, we have implemented a model of information propagation among evacuees. Information about the locations of exits and the relative confidence of the individual in the location of the exit disseminated through a simulated crowd of people during an evacuation modeled after The Station Nightclub fire of 2003. True believers were added to this system as individuals who refused to accept exit information from others, instead preferring to head to their own exit. This system was then tested to find what percentage of true believers most likely existed in the actual fire. Using this true believer percentage, robots were added to the environment to guide evacuees to the nearest exit. The number of people who believed a robot's instructions was varied to find what percentage of people need to trust these robots in order to exploit information propagation and thus increase survivability. As a lower bound, we have found that 30% of the evacuees should believe a robot's instructions to significantly increase survival rates. Paul Robinette, Patricio A. Vela, Ayanna M. Howard |
ICRA | 2 |
| 2012 | Reproducing Kernel Hilbert Space Approach for the Online Update of Radial Bases in Neuro-Adaptive ControlabstractClassical work in model reference adaptive control for uncertain nonlinear dynamical systems with a radial basis function (RBF) neural network adaptive element does not guarantee that the network weights stay bounded in a compact neighborhood of the ideal weights when the system signals are not persistently exciting (PE). Recent work has shown, however, that an adaptive controller using specifically recorded data concurrently with instantaneous data guarantees boundedness without PE signals. However, the work assumes fixed RBF network centers, which requires domain knowledge of the uncertainty. Motivated by reproducing kernel Hilbert space theory, we propose an online algorithm for updating the RBF centers to remove the assumption. In addition to proving boundedness of the resulting neuro-adaptive controller, a connection is made between PE signals and kernel methods. Simulation results show improved performance. Hassan A. Kingravi, Girish Chowdhary 0001, Patricio A. Vela, Eric N. Johnson |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2011 | PTZ camera-based adaptive panoramic and multi-layered background modelabstractIn this paper, we present a novel approach for constructing an adaptive panoramic and multi-layered background model for Pan-tilt-zoom (PTZ) camera that provides fast registration of the observed frame and localizes the foreground targets with arbitrary camera position and scale (optical zoom). Our method consists of two stages. (1) An adaptive panoramic background mixture model is generated off-line for foreground detection. (2) A layered correspondence is generated off-line from frames captured at different optical zoom values of the camera, and a correspondence propagation method is used to register the observed frame with the panoramic background online. We demonstrate the advantages of the proposed adaptive panoramic and multi-layered background model within wide field of view (FOV) and over large scale range. Kang Xue, Gbolabo Ogunmakin, Yue Liu 0005, Patricio A. Vela, Yongtian Wang |
ICIP | 4 |
| 2011 | A performance evaluation of vision and radio frequency tracking methods for interacting workforce
Jochen Teizer, Patricio A. Vela, Zhongke Shi |
Adv. Eng. Informatics | 4 |
| 2011 | Kernel Map Compression for Speeding the Execution of Kernel-Based MethodsabstractThe use of Mercer kernel methods in statistical learning theory provides for strong learning capabilities, as seen in kernel principal component analysis and support vector machines. Unfortunately, after learning, the computational complexity of execution through a kernel is of the order of the size of the training set, which is quite large for many applications. This paper proposes a two-step procedure for arriving at a compact and computationally efficient execution procedure. After learning in the kernel space, the proposed extension exploits the universal approximation capabilities of generalized radial basis function neural networks to efficiently approximate and replace the projections onto the empirical kernel map used during execution. Sample applications demonstrate significant compression of the kernel representation with graceful performance loss. Omar Arif, Patricio A. Vela |
IEEE Trans. Neural Networks | 2 |
| 2010 | Visual tracking and segmentation using Time-of-Flight sensorabstractTime-of-Flight (TOF) sensors provide range information at each pixel in addition to intensity information. They are becoming more widely available and more affordable. This paper examines the utility of dense TOF range data for image segmentation and tracking. Energy based formulations for image segmentation are used, which consist of a data term and a smoothness term. The paper proposes novel methods to incorporate range information, obtained from the TOF sensor, into the data and the smoothness term of the energy. Graph cut is used to minimize the energy. Omar Arif, Wayne Daley, Patricio A. Vela, Jochen Teizer, John M. Stewart |
ICIP | 3 |
| 2010 | Pre-image Problem in Manifold Learning and Dimensional Reduction MethodsabstractManifold learning and dimensional reduction methods provide a low dimensional embedding for a collection of training samples. These methods are based on the eigenvalue decomposition of the kernel matrix formed using the training samples. In the embedding is extended to new test samples using the Nystrom approximation method. This paper addresses the pre-image problem for these methods, which is to find the mapping back from the embedding space to the input space for new test points. The relationship of these learning methods to kernel principal component analysis and the connection of the out-of-sample problem to the pre-image problem is used to provide the pre-image. Omar Arif, Patricio A. Vela, Wayne Daley |
ICMLA | 2 |
| 2010 | Tracking multiple workers on construction sites using video cameras
Omar Arif, Patricio A. Vela, Jochen Teizer, Zhongke Shi |
Adv. Eng. Informatics | 3 |
| 2010 | A Probabilistic Contour Observer for Online Visual TrackingabstractThis paper presents an online, recursive filtering strategy for contour-based tracking. Approaching the tracking problem from an estimation perspective leads to an observer design for the visual track signal associated with an individual target in an image sequence. The track state of the observer is decomposed into group and shape components that describe the gross location and the nonrigid shape, respectively, of the object. A probabilistic representation describes the shape nonparametrically. The constitutive components of the observer are detailed, which include a dynamical prediction model and a correction mechanism. Incorporating the probabilistic observer into the tracking process leads to improved performance and segmentations. The improvements are validated through application of the observer to recorded imagery with evaluation via objective measures of quality. Ibrahima J. Ndiour, Jochen Teizer, Patricio A. Vela |
SIAM J. Imaging Sci. | 3 |
| 2009 | Robust Density Comparison for Visual TrackingabstractThis paper presents a technique to robustly compare two distributions represented by samples, without explicitly estimating the density. The method is based on mapping the distributions into a reproducing kernel Hilbert space, where eigenvalue decomposition is performed. Retention of only the top M eigenvectors minimizes the effect of noise on density comparison. A sample application of the technique is visual tracking, where an object is tracked by minimizing the distance between a model distribution and candidate distributions. Omar Arif, Patricio A. Vela |
BMVC | 2 |
| 2009 | Non-rigid object localization and segmentation using eigenspace representationabstractThis paper presents a novel non-rigid object localization and segmentation algorithm using an eigenspace representation. Previous approaches to eigenspace methods for object tracking use vectorized image regions as observations, whereas the proposed method uses each individual pixel as an observation. Localization using the pixel-wise eigenspace representation is robust to noise and occlusions. A unique feature of the approach is that it permits segmentation in addition to localization. Localization and segmentation are carried out by deriving a similarity function in the eigenspace. The algorithm is tested on synthetic and real world tracking examples to demonstrate the performance. Omar Arif, Patricio A. Vela |
ICCV | 2 |
| 2009 | Kernel map compression using generalized radial basis functionsabstractThe use of Mercer kernel methods in statistical learning theory provides for strong learning capabilities, as seen in kernel principal component analysis and support vector machines. Unfortunately the computational complexity of the resulting method is of the order of the training set, which is quite large for many applications. This paper proposes a two step procedure for arriving at a compact and computationally efficient learning procedure. After learning, the second step takes advantage of the universal approximation capabilities of generalized radial basis function neural networks to efficiently approximate the empirical kernel maps. Sample applications demonstrate significant compression of the kernel representation with graceful performance loss. Omar Arif, Patricio A. Vela |
ICCV | 2 |
| 2009 | Kernel covariance image region description for object trackingabstractWe propose a nonlinear covariance region descriptor for target tracking. The target object appearance and spatial information is represented using a covariance matrix in a target derived Hilbert space using kernel principal component analysis. A similarity measure is derived, which computes the similarity of a candidate image region to the learned covariance matrix. A variational technique is provided to maximize the similarity measure, which iteratively finds the best matched object region. Tracking performance is demonstrated on a variety of sequences containing noise, occlusions, illumination changes, background clutter, etc. Omar Arif, Patricio A. Vela |
ICIP | 2 |
| 2009 | A probabilistic shape filter for online contour trackingabstractOnline contour-based tracking is considered through the estimation perspective. We propose a recursive dynamic filtering solution to the tracking problem. The state of the target is described by a pose state which represents the ensemble movement and a shape state which represents the local deformations. The shape state of the filter is described implicitly by a probability field with prediction and correction mechanisms expressed accordingly. The filtering procedure decouples the pose and shape estimation. Experiments conducted with objective measures of quality demonstrate improved tracking. Ibrahima J. Ndiour, Omar Arif, Jochen Teizer, Patricio A. Vela |
ICIP | 4 |
| 2009 | Noise estimation and adaptive filtering during visual trackingabstractThis paper proposes a procedure to characterize segmentation-based visual tracking performance with respect to imaging noise. It identifies how imaging noise affects the target segmentation as measured through local shape metrics (Sobolev and Laplace metrics). Such a procedure would be an important calibration step prior to implementing a visual tracking filter for a given need. We utilize the Bhattacharyya coefficient between the target and background intensity distributions to estimate the segmentation error. An empirical study is conducted to establish a correspondence between the Bhattacharyya coefficient and the segmentation error. The correspondence is used to adaptively filter temporally correlated segmentations. Preliminary results show improved performance when compared to fixed gains. Ibrahima J. Ndiour, Patricio A. Vela |
ICIP | 2 |
| 2009 | Robust Target Localization and Segmentation Using Graph Cut, KPCA and Mean-ShiftabstractThis paper presents an algorithm for object localization and segmentation. The algorithm uses machine learning, and statistical and combinatorial optimization tools to build a tracker that is robust to noise and occlusions. The method is based on a novel energy formulation and its dual use for object localization and segmentation. The energy uses kernel principal component analysis to incorporate shape and appearance constraints of the target object and the background. The energy arising from the procedure is equivalent to an un-normalized density function, thus providing a probabilistic interpretation to the procedure. Mean-shift optimization finds the most probable location of the target object. Graph-cut maximization on the localized object window in the image generates the globally optimal segmentation. Omar Arif, Patricio A. Vela |
ICMLA | 2 |
| 2009 | Personnel tracking on construction sites using video cameras
Jochen Teizer, Patricio A. Vela |
Adv. Eng. Informatics | 2 |
| 2008 | Geometric Observers for Dynamically Evolving CurvesabstractThis paper proposes a deterministic observer framework for visual tracking based on non-parametric implicit (level-set) curve descriptions. The observer is continuous-discrete, with continuous-time system dynamics and discrete-time measurements. Its state-space consists of an estimated curve position augmented by additional states (e.g., velocities) associated with every point on the estimated curve. Multiple simulation models are proposed for state prediction. Measurements are performed through standard static segmentation algorithms and optical-flow computations. Special emphasis is given to the geometric formulation of the overall dynamical system. The discrete-time measurements lead to the problem of geometric curve interpolation and the discrete-time filtering of quantities propagated along with the estimated curve. Interpolation and filtering are intimately linked to the correspondence problem between curves. Correspondences are established by a Laplace-equation approach. The proposed scheme is implemented completely implicitly (by Eulerian numerical solutions of transport equations) and thus naturally allows for topological changes and subpixel accuracy on the computational grid. Marc Niethammer, Patricio A. Vela, Allen R. Tannenbaum |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Layered Active Contours for TrackingabstractPresented at British Machine Vision Conference 2007, University of Warwick, UK, September 10-13, 2007. Gallagher Pryor, Patricio A. Vela, Tauseef ur Rehman, Allen R. Tannenbaum |
BMVC | 2 |
| 2005 | On the Evolution of Vector Distance Functions of Closed Curves
Marc Niethammer, Patricio A. Vela, Allen R. Tannenbaum |
Int. J. Comput. Vis. | 2 |
| 2003 | Control of biomimetic locomotion via averaging theoryabstractBased on a recently developed "generalized averaging theory", we present a generic approach for the design of stabilizing feedback controller for biomimetic locomotive systems. The control laws exponentially stabilize in the average, and they apply to a very wide class of systems. Two examples are given: a "kinematic biped" that demonstrates how our theory handles discontinuities, and the snakeboard, which is an underactuated mechanical system with drift. Patricio A. Vela, Joel W. Burdick |
ICRA | 1 |
| 2002 | Trajectory Stabilization for a Planar Carangiform Robot FishabstractConsiders the task of trajectory stabilization for a fish-like robot by means of feedback. We use oscillatory control inputs and apply correction signals at the endpoints of each periodic input signal. Such a strategy can be proven to cause the system to converge to a desired trajectory. We present a specific model of a planar carangiform fish, and verify the stabilization results with simulations and with experiment on a planar robotic fish system that is propelled using carangiform-like movements. Kristi A. Morgansen, Patricio A. Vela, Joel W. Burdick |
ICRA | 2 |