VLDB 2026 Research / reviewers in the wild / expert
Garrett Warnell
dblp:173/5902 · also Garrett A. Warnell
· DBLP profile ↗
33ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0003-3846-8787ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 2 first-author · 15 since 2021Systems, architecture and hardware · 15 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VertiCoder: Self-Supervised Kinodynamic Representation Learning on Vertically Challenging TerrainabstractWe present Verticoder, a self-supervised representation learning approach for robot mobility on vertically challenging terrain. Using the same pre-training process, Ver-ticodercan handle four different downstream tasks, in-cluding forward kinodynamics learning, inverse kinodynamics learning, behavior cloning, and patch reconstruction with a single representation. Verticoder uses a TransformerEn-coder to learn the local context of its surroundings by random masking and next patch reconstruction. We show that Verti-coderachieves better performance across all four different tasks compared to specialized End - to- End models with 77 % fewer parameters. We also show Verticoder's comparable performance against state-of-the-art kinodynamic modeling and planning approaches in real-world robot deployment. These results underscore the efficacy of Verticoder in mitigating overfitting and fostering more robust generalization across diverse environmental contexts and downstream vehicle kin-odynamic tasks11https://github.com/mhnazeri/VertiCoder. Mohammad Nazeri, Aniket Datar, Anuj Pokhrel, Chenhui Pan, Garrett Warnell, Xuesu Xiao |
ICRA | 5 |
| 2024 | Wait, That Feels Familiar: Learning to Extrapolate Human Preferences for Preference-Aligned Path PlanningabstractAutonomous mobility tasks such as last-mile delivery require reasoning about operator-indicated preferences over terrains on which the robot should navigate to ensure both robot safety and mission success. However, coping with out of distribution data from novel terrains or appearance changes due to lighting variations remains a fundamental problem in visual terrain-adaptive navigation. Existing solutions either require labor-intensive manual data re-collection and labeling or use hand-coded reward functions that may not align with operator preferences. In this work, we posit that operator preferences for visually novel terrains, which the robot should adhere to, can often be extrapolated from established terrain preferences within the inertial-proprioceptive-tactile domain. Leveraging this insight, we introduce Preference extrApolation for Terrain-awarE Robot Navigation (PATERN), a novel framework for extrapolating operator terrain preferences for visual navigation. PATERN learns to map inertial-proprioceptive-tactile measurements from the robot’s observations to a representation space and performs nearest-neighbor search in this space to estimate operator preferences over novel terrains. Through physical robot experiments in outdoor environments, we assess PATERN’s capability to extrapolate preferences and generalize to novel terrains and challenging lighting conditions. Compared to baseline approaches, our findings indicate that PATERN1robustly generalizes to diverse terrains and varied lighting conditions, while navigating in a preference-aligned manner. Haresh Karnan, Elvin Yang, Garrett Warnell, Joydeep Biswas, Peter Stone 0001 |
ICRA | 3 |
| 2023 | Learning Perceptual Hallucination for Multi-Robot Navigation in Narrow HallwaysabstractWhile current systems for autonomous robot navigation can produce safe and efficient motion plans in static environments, they usually generate suboptimal behaviors when multiple robots must navigate together in confined spaces. For example, when two robots meet each other in a narrow hallway, they may either turn around to find an alternative route or collide with each other. This paper presents a new approach to navigation that allows two robots to pass each other in a narrow hallway without colliding, stopping, or waiting. Our approach, Perceptual Hallucination for Hallway Passing (PHHP), learns to synthetically generate virtual obstacles (i.e., perceptual hallucination) to facilitate passing in narrow hallways by multiple robots that utilize otherwise standard autonomous navigation systems. Our experiments on physical robots in a variety of hallways show improved performance compared to multiple baselines. Jin Soo Park, Xuesu Xiao, Garrett Warnell, Harel Yedidsion, Peter Stone 0001 |
ICRA | 3 |
| 2022 | Skeletal Feature Compensation for Imitation Learning with Embodiment MismatchabstractLearning from demonstrations in the wild (e.g. YouTube videos) is a tantalizing goal in imitation learning. However, for this goal to be achieved, imitation learning algorithms must deal with the fact that the demonstrators and learners may have bodies that differ from one another. This condition — “embodiment mismatch” — is ignored by many recent imitation learning algorithms. Our proposed imitation learning technique, SILEM (Skeletal feature compensation for Imitation Learning with Embodiment Mismatch), addresses a particular type of embodiment mismatch by introducing a learned affine transform to compensate for differences in the skeletal features obtained from the learner and expert. We create toy domains based on PyBullet's HalfCheetah and Ant to assess SILEM's benefits for this type of embodiment mismatch. We also provide qualitative and quantitative results on more realistic problems — teaching simulated humanoid agents, including Atlas from Boston Dynamics, to walk by observing human demonstrations. Eddy Hudson, Garrett Warnell, Faraz Torabi, Peter Stone 0001 |
ICRA | 2 |
| 2022 | Adversarial Imitation Learning from Video Using a State ObserverabstractThe imitation learning research community has recently made significant progress towards the goal of enabling artificial agents to imitate behaviors from video demonstrations alone. However, current state-of-the-art approaches developed for this problem exhibit high sample complexity due, in part, to the high-dimensional nature of video observations. Towards addressing this issue, we introduce here a new algorithm called Visual Generative Adversarial Imitation from Observation using a State Observer (VGAIfO-SO). At its core, VGAIfO-SO seeks to address sample inefficiency using a novel, self-supervised state observer, which provides estimates of lower-dimensional proprioceptive state representations from high-dimensional images. We show experimentally in several continuous control environments that VGAIfO-SO is more sample efficient than other IfO algorithms at learning from video-only demonstrations and can sometimes even achieve performance close to the Generative Adversarial Imitation from Observation (GAIfO) algorithm that has privileged access to the demonstrator's proprioceptive state information. Haresh Karnan, Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
ICRA | 3 |
| 2022 | VOILA: Visual-Observation-Only Imitation Learning for Autonomous NavigationabstractWhile imitation learning for vision-based au-tonomous mobile robot navigation has recently received a great deal of attention in the research community, existing approaches typically require state-action demonstrations that were gathered using the deployment platform. However, what if one cannot easily outfit their platform to record these demonstration signals or-worse yet-the demonstrator does not have access to the platform at all? Is imitation learning for vision-based autonomous navigation even possible in such scenarios? In this work, we hypothesize that the answer is yes and that recent ideas from the Imitation from Observation (IfO) literature can be brought to bear such that a robot can learn to navigate using only ego-centric video collected by a demonstrator, even in the presence of viewpoint mismatch. To this end, we introduce a new algorithm, Visual-Observation-only Imitation Learning for Autonomous navigation (VOILA), that can successfully learn navigation policies from a single video demonstration collected from a physically different agent. We evaluate VOILA in the AirSim simulator and show that VOILA not only successfully imitates the expert, but that it also learns navigation policies that can generalize to novel environments. Further, we demonstrate the effectiveness of VOILA in a real-world setting by showing that it allows a wheeled Jackal robot to successfully imitate a human walking in an environment while recording video with a handheld mobile phone camera. Haresh Karnan, Garrett Warnell, Xuesu Xiao, Peter Stone 0001 |
ICRA | 2 |
| 2022 | Visual Representation Learning for Preference-Aware Path PlanningabstractAutonomous mobile robots deployed in outdoor environments must reason about different types of terrain for both safety (e.g., prefer dirt over mud) and deployer preferences (e.g., prefer dirt path over flower beds). Most existing solutions to this preference-aware path planning problem use semantic segmentation to classify terrain types from camera images, and then ascribe costs to each type. Unfortunately, there are three key limitations of such approaches - they 1) require preenumeration of the discrete terrain types, 2) are unable to handle hybrid terrain types (e.g., grassy dirt), and 3) require expensive labelled data to train visual semantic segmentation. We introduce Visual Representation Learning for Preference-Aware Path Planning (VRL-PAP), an alternative approach that overcomes all three limitations: VRL-PAP leverages un-labelled human demonstrations of navigation to autonomously generate triplets for learning visual representations of terrain that are viewpoint invariant and encode terrain types in a continuous representation space. The learned representations are then used along with the same unlabelled human navigation demonstrations to learn a mapping from the representation space to terrain costs. At run time, VRL-PAP maps from images to representations and then representations to costs to perform preference-aware path planning. We present empirical results from challenging outdoor settings that demonstrate VRL-PAP 1) is successfully able to pick paths that reflect demonstrated preferences, 2) is comparable in execution to geometric navigation with a highly detailed manually annotated map (without requiring such annotations), 3) is able to generalize to novel terrain types with minimal additional unlabeled demonstrations. Kavan Singh Sikand, Sadegh Rabiee, Adam Uccello, Xuesu Xiao, Garrett Warnell, Joydeep Biswas |
ICRA | 5 |
| 2022 | VI-IKD: High-Speed Accurate Off-Road Navigation using Learned Visual-Inertial Inverse KinodynamicsabstractOne of the key challenges in high-speed off-road navigation on ground vehicles is that the kinodynamics of the vehicle-terrain interaction can differ dramatically depending on the terrain. Previous approaches to addressing this challenge have considered learning an inverse kinodynamics (IKD) model, conditioned on inertial information of the vehicle to sense the kinodynamic interactions. In this paper, we hypothesize that to enable accurate high-speed off-road navigation using a learned IKD model, in addition to inertial information from the past, one must also anticipate the kinodynamic interactions of the vehicle with the terrain in the future. To this end, we introduce Visual-Inertial Inverse Kinodynamics (VI-IKD), a novel learning based IKD model that is conditioned on visual information from a terrain patch ahead of the robot in addition to past inertial information, enabling it to anticipate kinodynamic interactions in the future. We validate the effectiveness of VI-IKD in accurate high-speed off-road navigation experimentally on a scale 1/5 UT-AlphaTruck off-road autonomous vehicle in both indoor and outdoor environments and show that compared to other state-of-the-art approaches, VI-IKD enables more accurate and robust off-road navigation on a variety of different terrains at speeds of up to 3.5m/s. Haresh Karnan, Kavan Singh Sikand, Pranav Atreya, Sadegh Rabiee, Xuesu Xiao, Garrett Warnell, Peter Stone 0001, Joydeep Biswas |
IROS | 6 |
| 2022 | Lucid dreaming for experience replay: refreshing past states with the current policy
Yunshu Du, Garrett Warnell, Assefaw Hadish Gebremedhin, Peter Stone 0001, Matthew E. Taylor |
Neural Comput. Appl. | 2 |
| 2021 | Goal Blending for Responsive Shared Autonomy in a Navigating VehicleabstractHuman-robot shared autonomy techniques for vehicle navigation hold promise for reducing a human driver’s workload, ensuring safety, and improving navigation efficiency. However, because typical techniques achieve these improvements by effectively removing human control at critical moments, these approaches often exhibit poor responsiveness to human commands—especially in cluttered environments. In this paper, we propose a novel goal-blending shared autonomy (GBSA) system, which aims to improve responsiveness in shared autonomy systems by blending human and robot input during the selection of local navigation goals as opposed to low-level motor (servo-level) commands. We validate the proposed approach by performing a human study involving an intelligent wheelchair and compare GBSA to a representative servo-level shared control system that uses a policy-blending approach. The results of both quantitative performance analysis and a subjective survey show that GBSA exhibits significantly better system responsiveness and induces higher user satisfaction than the existing approach. Yu-Sian Jiang, Garrett Warnell, Peter Stone 0001 |
AAAI | 2 |
| 2021 | APPLI: Adaptive Planner Parameter Learning From InterventionsabstractWhile classical autonomous navigation systems can typically move robots from one point to another safely and in a collision-free manner, these systems may fail or produce suboptimal behavior in certain scenarios. The current practice in such scenarios is to manually re-tune the system’s parameters, e.g. max speed, sampling rate, inflation radius, to optimize performance. This practice requires expert knowledge and may jeopardize performance in the originally good scenarios. Meanwhile, it is relatively easy for a human to identify those failure or suboptimal cases and provide a teleoperated intervention to correct the failure or suboptimal behavior. In this work, we seek to learn from those human interventions to improve navigation performance. In particular, we propose Adaptive Planner Parameter Learning from Interventions (APPLI), in which multiple sets of navigation parameters are learned during training and applied based on a confidence measure to the underlying navigation system during deployment. In our physical experiments, the robot achieves better performance compared to the planner with static default parameters, and even dynamic parameters learned from a full human demonstration. We also show APPLI’s generalizability in another unseen physical test course, and a suite of 300 simulated navigation environments. Xuesu Xiao, Bo Liu 0042, Garrett Warnell, Peter Stone 0001 |
ICRA | 4 |
| 2021 | APPLR: Adaptive Planner Parameter Learning from ReinforcementabstractClassical navigation systems typically operate using a fixed set of hand-picked parameters (e.g. maximum speed, sampling rate, inflation radius, etc.) and require heavy expert re-tuning in order to work in new environments. To mitigate this requirement, it has been proposed to learn parameters for different contexts in a new environment using human demonstrations collected via teleoperation. However, learning from human demonstration limits deployment to the training environment, and limits overall performance to that of a potentially-suboptimal demonstrator. In this paper, we introduce APPLR, Adaptive Planner Parameter Learning from Reinforcement, which allows existing navigation systems to adapt to new scenarios by using a parameter selection scheme discovered via reinforcement learning (RL) in a wide variety of simulation environments. We evaluate APPLR on a robot in both simulated and physical experiments, and show that it can outperform both a fixed set of hand-tuned parameters and also a dynamic parameter tuning scheme learned from human demonstration. Zifan Xu, Gauraang Dhamankar, Anirudh Nair, Xuesu Xiao, Garrett Warnell, Bo Liu 0042, Peter Stone 0001 |
ICRA | 5 |
| 2021 | DEALIO: Data-Efficient Adversarial Learning for Imitation from ObservationabstractIn imitation learning from observation (IfO), a learning agent seeks to imitate a demonstrating agent using only observations of the demonstrated behavior without access to the control signals generated by the demonstrator. Recent methods based on adversarial imitation learning have led to state-of-the-art performance on IfO problems, but they typically suffer from high sample complexity due to a reliance on data-inefficient, model-free reinforcement learning algorithms. This issue makes them impractical to deploy in real-world settings, where gathering samples can incur high costs in terms of time, energy, and risk. In this work, we hypothesize that we can incorporate ideas from model-based reinforcement learning with adversarial methods for IfO in order to increase the data efficiency of these methods without sacrificing performance. Specifically, we consider time-varying linear Gaussian policies, and propose a method that integrates the linear-quadratic regulator with path integral policy improvement into an existing adversarial IfO framework. The result is a more data-efficient IfO algorithm with better performance, which we show empirically in four simulation domains: using far fewer interactions with the environment, the proposed method exhibits similar or better performance than the existing technique. Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
IROS | 2 |
| 2021 | Recent advances in leveraging human guidance for sequential decision-making tasks
Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 3 |
| 2021 | Grounded action transformation for sim-to-real reinforcement learningabstractAbstract Reinforcement learning in simulation is a promising alternative to the prohibitive sample cost of reinforcement learning in the physical world. Unfortunately, policies learned in simulation often perform worse than hand-coded policies when applied on the target, physical system. Grounded simulation learning (gsl) is a general framework that promises to address this issue by altering the simulator to better match the real world (Farchy et al. 2013 in Proceedings of the 12th international conference on autonomous agents and multiagent systems (AAMAS)). This article introduces a new algorithm for gsl—Grounded Action Transformation (GAT)—and applies it to learning control policies for a humanoid robot. We evaluate our algorithm in controlled experiments where we show it to allow policies learned in simulation to transfer to the real world. We then apply our algorithm to learning a fast bipedal walk on a humanoid robot and demonstrate a 43.27% improvement in forward walk velocity compared to a state-of-the art hand-coded walk. This striking empirical success notwithstanding, further empirical analysis shows that gat may struggle when the real world has stochastic state transitions. To address this limitation we generalize gat to the stochasticgat (sgat) algorithm and empirically show that sgat leads to successful real world transfer in situations where gat may fail to find a good policy. Our results contribute to a deeper understanding of grounded simulation learning and demonstrate its effectiveness for applying reinforcement learning to learn robot control policies entirely in simulation. Josiah Hanna, Siddharth Desai, Haresh Karnan, Garrett Warnell, Peter Stone 0001 |
Mach. Learn. | 4 |
| 2020 | Stochastic Grounded Action Transformation for Robot Learning in SimulationabstractRobot control policies learned in simulation do not often transfer well to the real world. Many existing solutions to this sim-to-real problem, such as the Grounded Action Transformation (GAT) algorithm, seek to correct for- or ground-these differences by matching the simulator to the real world. However, the efficacy of these approaches is limited if they do not explicitly account for stochasticity in the target environment. In this work, we analyze the problems associated with grounding a deterministic simulator in a stochastic real world environment, and we present examples where GAT fails to transfer a good policy due to stochastic transitions in the target domain. In response, we introduce the Stochastic Grounded Action Transformation (SGAT) algorithm, which models this stochasticity when grounding the simulator. We find experimentally-for both simulated and physical target domains-that SGAT can find policies that are robust to stochasticity in the target domain. Siddharth Desai, Haresh Karnan, Josiah Hanna, Garrett Warnell, Peter Stone 0001 |
IROS | 4 |
| 2020 | Reinforced Grounded Action Transformation for Sim-to-Real TransferabstractRobots can learn to do complex tasks in simulation, but often, learned behaviors fail to transfer well to the real world due to simulator imperfections (the "reality gap"). Some existing solutions to this sim-to-real problem, such as Grounded Action Transformation (gat), use a small amount of real-world experience to minimize the reality gap by "grounding" the simulator. While very effective in certain scenarios, gat is not robust on problems that use complex function approximation techniques to model a policy. In this paper, we introduce Reinforced Grounded Action Transformation (rgat), a new sim-to-real technique that uses Reinforcement Learning (RL) not only to update the target policy in simulation, but also to perform the grounding step itself. This novel formulation allows for end-to-end training during the grounding step, which, compared to gat, produces a better grounded simulator. Moreover, we show experimentally in several MuJoCo domains that our approach leads to successful transfer for policies modeled using neural networks. Haresh Karnan, Siddharth Desai, Josiah Hanna, Garrett Warnell, Peter Stone 0001 |
IROS | 4 |
| 2020 | An Imitation from Observation Approach to Transfer Learning with Dynamics MismatchabstractWe examine the problem of transferring a policy learned in a source environment to a target environment with different dynamics, particularly in the case where it is critical to reduce the amount of interaction with the target environment during learning. This problem is particularly important in sim-to-real transfer because simulators inevitably model real-world dynamics imperfectly. In this paper, we show that one existing solution to this transfer problem-- grounded action transformation --is closely related to the problem of imitation from observation (IfO): learning behaviors that mimic the observations of behavior demonstrations. After establishing this relationship, we hypothesize that recent state-of-the-art approaches from the IfO literature can be effectively repurposed for grounded transfer learning. To validate our hypothesis we derive a new algorithm -- generative adversarial reinforced action transformation (GARAT) -- based on adversarial imitation from observation techniques. We run experiments in several domains with mismatched dynamics, and find that agents trained with GARAT achieve higher returns in the target environment compared to existing black-box transfer methods. Siddharth Desai, Ishan Durugkar, Haresh Karnan, Garrett Warnell, Josiah Hanna, Peter Stone 0001 |
NeurIPS | 4 |
| 2020 | Agents teaching agents: a survey on inter-agent transfer learning
Felipe Leno da Silva, Garrett Warnell, Anna Helena Reali Costa, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 2 |
| 2019 | Imitation Learning from Video by Leveraging ProprioceptionabstractClassically, imitation learning algorithms have been developed for idealized situations, e.g., the demonstrations are often required to be collected in the exact same environment and usually include the demonstrator's actions. Recently, however, the research community has begun to address some of these shortcomings by offering algorithmic solutions that enable imitation learning from observation (IfO), e.g., learning to perform a task from visual demonstrations that may be in a different environment and do not include actions. Motivated by the fact that agents often also have access to their own internal states (i.e., proprioception), we propose and study an IfO algorithm that leverages this information in the policy learning process. The proposed architecture learns policies over proprioceptive state representations and compares the resulting trajectories visually to the demonstration data. We experimentally test the proposed technique on several MuJoCo domains and show that it outperforms other imitation from observation algorithms by a large margin. Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
IJCAI | 2 |
| 2019 | Recent Advances in Imitation Learning from ObservationabstractImitation learning is the process by which one agent tries to learn how to perform a certain task using information generated by another, often more-expert agent performing that same task. Conventionally, the imitator has access to both state and action information generated by an expert performing the task (e.g., the expert may provide a kinesthetic demonstration of object placement using a robotic arm). However, requiring the action information prevents imitation learning from a large number of existing valuable learning resources such as online videos of humans performing tasks. To overcome this issue, the specific problem of imitation from observation (IfO) has recently garnered a great deal of attention, in which the imitator only has access to the state information (e.g., video frames) generated by the expert. In this paper, we provide a literature review of methods developed for IfO, and then point out some open research problems and potential future work. Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
IJCAI | 2 |
| 2019 | Parsimonious Online Learning with Kernels via Sparse Projections in Function SpaceabstractDespite their attractiveness, popular perception is that techniques for nonparametric function approximation do not scale to streaming data due to an intractable growth in the amount of storage they require. To solve this problem in a memory-affordable way, we propose an online technique based on functional stochastic gradient descent in tandem with supervised sparsification based on greedy function subspace projections. The method, called parsimonious online learning with kernels (POLK), provides a controllable tradeoff between its solution accuracy and the amount of memory it requires. We derive conditions under which the generated function sequence converges almost surely to the optimal function, and we establish that the memory requirement remains finite. We evaluate POLK for kernel multi-class logistic regression and kernel hinge-loss classification on three canonical data sets: a synthetic Gaussian mixture model, the MNIST hand-written digits, and the Brodatz texture database. On all three tasks, we observe a favorable trade-off of objective function evaluation, classification performance, and complexity of the nonparametric regressor extracted by the proposed method. Alec Koppel, Garrett Warnell, Ethan Stump, Alejandro Ribeiro |
J. Mach. Learn. Res. | 2 |
| 2018 | Deep TAMER: Interactive Agent Shaping in High-Dimensional State SpacesabstractWhile recent advances in deep reinforcement learning have allowed autonomous learning agents to succeed at a variety of complex tasks, existing algorithms generally require a lot oftraining data. One way to increase the speed at which agent sare able to learn to perform tasks is by leveraging the input of human trainers. Although such input can take many forms, real-time, scalar-valued feedback is especially useful in situations where it proves difficult or impossible for humans to provide expert demonstrations. Previous approaches have shown the usefulness of human input provided in this fashion (e.g., the TAMER framework), but they have thus far not considered high-dimensional state spaces or employed the use of deep learning. In this paper, we do both: we propose DeepTAMER, an extension of the TAMER framework that leverages the representational power of deep neural networks inorder to learn complex tasks in just a short amount of time with a human trainer. We demonstrate Deep TAMER’s success by using it and just 15 minutes of human-provided feedback to train an agent that performs better than humans on the Atari game of Bowling - a task that has proven difficult for even state-of-the-art reinforcement learning methods. Garrett Warnell, Nicholas R. Waytowich, Vernon Lawhern, Peter Stone 0001 |
AAAI | 1 |
| 2018 | Inferring User Intention using Gaze in VehiclesabstractMotivated by the desire to give vehicles better information about their drivers, we explore human intent inference in the setting of a human driver riding in a moving vehicle. Specifically, we consider scenarios in which the driver intends to go to or learn about a specific point of interest along the vehicle's route, and an autonomous system is tasked with inferring this point of interest using gaze cues. Because the scene under observation is highly dynamic --- both the background and objects in the scene move independently relative to the driver --- such scenarios are significantly different from the static scenes considered by most literature in the eye tracking community. In this paper, we provide a formulation for this new problem of determining a point of interest in a dynamic scenario. We design an experimental framework to systematically evaluate initial solutions to this novel problem, and we propose our own solution called dynamic interest point detection (DIPD). We experimentally demonstrate the success of DIPD when compared to baseline nearest-neighbor or filtering approaches. Yu-Sian Jiang, Garrett Warnell, Peter Stone 0001 |
ICMI | 2 |
| 2018 | Behavioral Cloning from ObservationabstractHumans often learn how to perform tasks via imitation: they observe others perform a task, and then very quickly infer the appropriate actions to take based on their observations. While extending this paradigm to autonomous agents is a well-studied problem in general, there are two particular aspects that have largely been overlooked: (1) that the learning is done from observation only (i.e., without explicit action information), and (2) that the learning is typically done very quickly. In this work, we propose a two-phase, autonomous imitation learning technique called behavioral cloning from observation (BCO), that aims to provide improved performance with respect to both of these aspects. First, we allow the agent to acquire experience in a self-supervised fashion. This experience is used to develop a model which is then utilized to learn a particular task by observing an expert perform that task without the knowledge of the specific actions taken. We experimentally compare BCO to imitation learning methods, including the state-of-the-art, generative adversarial imitation learning (GAIL) technique, and we show comparable task performance in several different simulation domains while exhibiting increased learning speed after expert trajectories become available. Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
IJCAI | 2 |
| 2018 | A Study of Human-Robot Copilot Systems for En-route Destination ChangingabstractIn this paper, we introduce the problem of enroute destination changing for a self-driving car, and we study the effectiveness of human-robot copilot systems as a solution. The copilot system is one in which the autonomous vehicle not only handles low-level vehicle control, but also continually monitors the intent of the human passenger in order to respond to dynamic changes in desired destination. We specifically consider a vehicle parking task, where the vehicle must respond to the user's intent to drive to and park next to a particular roadside sign board, and we study a copilot system that detects the passenger's intended destination based on gaze. We conduct a human study to investigate, in the context of our parking task, (a) if there is benefit in using a copilot system over manual driving, and (b) if copilot systems that use eye tracking to detect the intended destination have any benefit compared to those that use a more traditional, keyboard-based system. We find that the answers to both of these questions are affirmative: our copilot systems can complete the autonomous parking task more efficiently than human drivers can, and our copilot system that utilizes gaze information enjoys an increased success rate over one that utilizes typed input. Yu-Sian Jiang, Garrett Warnell, Eduardo Munera Sánchez, Peter Stone 0001 |
RO-MAN | 2 |
| 2017 | Parsimonious Online Learning with Kernels via sparse projections in function spaceabstractWe consider stochastic nonparametric regression problems in a reproducing kernel Hilbert space (RKHS), an extension of expected risk minimization to nonlinear function estimation. Popular perception is that kernel methods are inapplicable to online settings, since the generalization of stochastic methods to kernelized function spaces require memory storage that is cubic in the iteration index (“the curse of kernelization”). We alleviate this intractability in two ways: (1) we consider the use of functional stochastic gradient method (FSGD) which operates on a subset of training examples at each step; and (2), we extract parsimonious approximations of the resulting stochastic sequence via a greedy sparse subspace projection scheme based on kernel orthogonal matching pursuit (KOMP). We establish that this method converges almost surely in both diminishing and constant algorithm step-size regimes for a specific selection of sparse approximation budget. The method is evaluated on a kernel multi-class support vector machine problem, where data samples are generated from class-dependent Gaussian mixture models. Alec Koppel, Garrett Warnell, Ethan Stump, Alejandro Ribeiro |
ICASSP | 2 |
| 2016 | Online learning for characterizing unknown environments in ground robotic vehicle modelsabstractIn pursuit of increasing the operational tempo of a ground robotics platform in unknown domains, we consider the problem of predicting the distribution of structural state-estimation error due to poorly-modeled platform dynamics as well as environmental effects. Such predictions are a critical component of any modern control approach that utilizes uncertainty information to provide robustness in control design. We use an online learning algorithm based on matrix factorization techniques to fit a statistical model of error that provides enough expressive power to enable prediction directly from motion control signals and low-level visual features. Moreover, we empirically demonstrate that this technique compares favorably to predictors that do not incorporate this information. Alec Koppel, Jonathan Fink, Garrett Warnell, Ethan Stump, Alejandro Ribeiro |
IROS | 3 |
| 2016 | Ray Saliency: Bottom-Up Visual Saliency for a Rotating and Zooming Camera
Garrett Warnell, Philip David, Rama Chellappa |
Int. J. Comput. Vis. | 1 |
| 2015 | Integrability-regularized phase unwrapping via sparse error correctionabstractWe propose a new formulation of the classical two-dimensional phase unwrapping problem. Using a sparse-error, gradient-domain measurement model, we simultaneously seek the absolute phase and sparse gradient errors that minimize a novel energy functional that strongly encourages the integrability of the corrected gradient field. Our approach can be cast as a generalized lasso problem, and we compute the solution using the alternating direction method of multipliers (ADMM) algorithm. Adopting a commonly-used inter-ferometric synthetic aperture radar noise model, we evaluate our technique for several synthetic surfaces. Garrett Warnell, Vishal M. Patel, Rama Chellappa |
ICIP | 1 |
| 2015 | D4L: Decentralized dynamic discriminative dictionary learningabstractWe consider discriminative dictionary learning in a distributed online setting, where a team of networked robots aims to jointly learn both a common basis of the feature space and a classifier over this basis from sequentially observed signals. We formulate this problem as a distributed stochastic program with a non-convex objective and present a block variant of the Arrow-Hurwicz saddle point algorithm to solve it. Only neighboring nodes in the communications network need to exchange information, and we penalize the discrepency between the individual feature basis and classifiers using Lagrange multipliers. The application we consider is for a team of robots to collaboratively recognize objects of interest in dynamic environments. As a preliminary performance benchmark, we consider the problem of learning a texture classifier across a network of robots moving around an urban setting where separate training examples are sequentially observed at each robot. Results are shown for both a standard texture dataset and a new dataset from an urban training facility, and we compare the performance of the standard centralized construction to the new distributed algorithm for the case when distinct samples from all classes are seen by the robots. These experiments yield comparable performance between the decentralized and the centralized cases, demonstrating the proposed method's practical utility. Alec Koppel, Garrett Warnell, Ethan Stump, Alejandro Ribeiro |
IROS | 2 |
| 2015 | Adaptive-Rate Compressive Sensing Using Side InformationabstractWe provide two novel adaptive-rate compressive sensing (CS) strategies for sparse, time-varying signals using side information. The first method uses extra cross-validation measurements, and the second one exploits extra low-resolution measurements. Unlike the majority of current CS techniques, we do not assume that we know an upper bound on the number of significant coefficients that comprises the images in the video sequence. Instead, we use the side information to predict the number of significant coefficients in the signal at the next time instant. We develop our techniques in the specific context of background subtraction using a spatially multiplexing CS camera such as the single-pixel camera. For each image in the video sequence, the proposed techniques specify a fixed number of CS measurements to acquire and adjust this quantity from image to image. We experimentally validate the proposed methods on real surveillance video sequences. Garrett Warnell, Sourabh Bhattacharya, Rama Chellappa, Tamer Basar |
IEEE Trans. Image Process. | 1 |
| 2012 | Adaptive rate compressive sensing for background subtractionabstractWe study the problem of adaptive compressive sensing (CS) of a time-varying signal with slowly changing sparsity and rapidly varying support. We are specifically interested in visual surveillance applications such as background subtraction and tracking. Classical CS theory assumes prior knowledge of signal sparsity in order to determine the number of sensor measurements needed to ensure adequate signal reconstruction. However, when dealing with time-varying signals such as video, prior information regarding the exact sparsity may be difficult to obtain. Assuming a sensor that is able to take an adaptive number of compressive measurements, we present an algorithm based on cross validation that quantitatively evaluates the current measurement rate and adjusts it as needed. Garrett Warnell, Dikpal Reddy, Rama Chellappa |
ICASSP | 1 |