EDBT 2026 Demo / reviewers in the wild / expert
Ashish Kapoor
dblp:93/161
· DBLP profile ↗
107ranked-venue papers
25as first author
20since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 17 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 11 first-author · 2 since 2021Systems, architecture and hardware · 16 · 8 since 2021Human-computer interaction and ubiquitous computing · 14 · 4 first-authorComputer networks · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ConBaT: Control Barrier Transformer for Safe Robot Learning from DemonstrationsabstractLarge-scale self-supervised models have recently revolutionized our ability to perform a variety of tasks within the vision and language domains. However, using such models for autonomous systems is challenging because of safety requirements: besides executing correct actions, an autonomous agent must also avoid the high cost and potentially fatal critical mistakes. Traditionally, self-supervised training mainly focuses on imitating previously observed behaviors, and the training demonstrations carry no notion of which behaviors should be explicitly avoided. In this work, we propose Control Barrier Transformer (ConBaT), an approach that learns safe behaviors from demonstrations in a self-supervised fashion. ConBaT is inspired by the concept of control barrier functions in control theory and uses a causal transformer that learns to predict safe robot actions autoregressively using a critic that requires minimal safety data labeling. During deployment, we employ a lightweight online optimization to find actions that ensure future states lie within the learned safe set. We apply our approach to different simulated control tasks and show that our method results in safer control policies compared to other classical and learning-based methods such as imitation learning, reinforcement learning, and model predictive control. Sai Vemprala, Rogerio Bonatti, Chuchu Fan, Ashish Kapoor |
ICRA | 5 |
| 2024 | EvDNeRF: Reconstructing Event Data with Dynamic Neural Radiance FieldsabstractWe present EvDNeRF, a pipeline for generating event data and training an event-based dynamic NeRF, for the purpose of faithfully reconstructing eventstreams on scenes with rigid and non-rigid deformations that may be too fast to capture with a standard camera. Event cameras register asynchronous per-pixel brightness changes at MHz rates with high dynamic range, making them ideal for observing fast motion with almost no motion blur. Neural radiance fields (NeRFs) offer visual-quality geometric-based learnable rendering, but prior work with events has only considered reconstruction of static scenes. Our EvDNeRF can predict eventstreams of dynamic scenes from a static or moving viewpoint between any desired timestamps, thereby allowing it to be used as an event-based simulator for a given scene. We show that by training on varied batch sizes of events, we can improve test-time predictions of events at fine time resolutions, outperforming baselines that pair standard dynamic NeRFs with event generators. We release our simulated and real datasets, as well as code for multi-view event-based data generation and the training and evaluation of EvDNeRF models1. Anish Bhattacharya, Ratnesh Madaan, Fernando Cladera Ojeda, Sai Vemprala, Rogerio Bonatti, Kostas Daniilidis, Ashish Kapoor, Vijay Kumar 0001, Nikolai Matni, Jayesh K. Gupta |
WACV | 7 |
| 2024 | Foundation Models for Aerial RoboticsabstractDeveloping machine intelligence abilities in robots and autonomous systems is an expensive and time-consuming process. Existing solutions are tailored to specific applications and are harder to generalize. Furthermore, scarcity of training data adds a layer of complexity in deploying deep machine learning models. We present a new platform for General Robot Intelligence Development (GRID) to address both of these issues. The platform enables robots to learn, compose and adapt skills to their physical capabilities, environmental constraints and goals. One of the components of GRID is a state-of-the-art simulation system that models robot physics and the machine intelligence processes. This system in turn is tightly coupled with a multitude of Foundation Models that enables rapid prototyping, design, debugging and refinement of robot AI models. GRID is designed from the ground up to be extensible to accommodate new types of robots, vehicles, hardware platforms and software protocols. In addition, the modular design enables various deep ML components and existing foundation models to be easily usable in a wider variety of robot-centric problems. We demonstrate the platform in various aerial robotics scenarios and demonstrate how the platform dramatically accelerates development of machine intelligent robots. Ashish Kapoor |
WSDM | 1 |
| 2023 | Formal Verification of Floating-Point DivisionabstractVerification of complex datapath circuits such as floating-point dividers are known to be a challenging problem. In this paper, we present a formal verification methodology to verify floating-point (FP) dividers. In general, floating-point division unit builds around a fixed-point division implementation. Our solution performs a two-step verification.The first step verifies the fixed-point division implementation. We target fixed-point division algorithms that compute a fixed number of quotient bits in each iteration. This step uses a combination of equivalence checking and assertion-based property checking techniques. We used property checking to show correctness of radix-2 restoring division and equivalence checking to show equivalence between the radix-2 restoring division and a prescaled radix-4 non-restoring division.The second step uses equivalence checking to compare the floating-point divider with a software golden reference of FP division. In this step we assume the fixed-point division is working correctly to make the proof tractable. Using the proposed steps, verification of single precision FP divider took 1 hour 30 minutes and double precision FP divider took 7 hours and 30 minutes. Ashish Kapoor, Warren E. Ferguson, Himanshu Jain, Sudipta Kundu |
ARITH | 1 |
| 2023 | Is Imitation All You Need? Generalized Decision-Making with Dual-Phase TrainingabstractWe introduce DualMind, a generalist agent designed to tackle various decision-making tasks that addresses challenges posed by current methods, such as overfitting behaviors and dependence on task-specific fine-tuning. DualMind uses a novel "Dual-phase" training strategy that emulates how humans learn to act in the world. The model first learns fundamental common knowledge through a self-supervised objective tailored for control tasks and then learns how to make decisions based on different contexts through imitating behaviors conditioned on given prompts. DualMind can handle tasks across domains, scenes, and embodiments using just a single set of model weights and can execute zero-shot prompting without requiring task-specific finetuning. We evaluate DualMind on MetaWorld [40] and Habitat [31] through extensive experiments and demonstrate its superior generalizability compared to previous techniques, outperforming other generalist agents by over 50% and 70% on Habitat and MetaWorld, respectively. On the 45 tasks in MetaWorld, DualMind achieves over 30 tasks at a 90% success rate. Our source code is available at https://github.com/yunyikristy/DualMind. Yao Wei 0002, Yanchao Sun, Ruijie Zheng, Sai Vemprala, Rogerio Bonatti, Ratnesh Madaan, Zhongjie Ba, Ashish Kapoor |
ICCV | 9 |
| 2023 | SMART: Self-supervised Multi-task pretrAining with contRol Transformers
Yanchao Sun, Ratnesh Madaan, Rogerio Bonatti, Furong Huang, Ashish Kapoor |
ICLR | 6 |
| 2023 | ClimaX: A foundation model for weather and climateabstractRecent data-driven approaches based on machine learning aim to directly solve a downstream forecasting or projection task by learning a data-driven functional mapping using deep neural networks. However, these networks are trained using curated and homogeneous climate datasets for specific spatiotemporal tasks, and thus lack the generality of currently used computationally intensive physics-informed numerical models for weather and climate modeling. We develop and demonstrate ClimaX, a flexible and generalizable deep learning model for weather and climate science that can be trained using heterogeneous datasets spanning different variables, spatio-temporal coverage, and physical groundings. ClimaX extends the Transformer architecture with novel encoding and aggregation blocks that allow effective use of available compute and data while maintaining general utility. ClimaX is pretrained with a self-supervised learning objective on climate datasets derived from CMIP6. The pretrained ClimaX can then be fine-tuned to address a breadth of climate and weather tasks, including those that involve atmospheric variables and spatio-temporal scales unseen during pretraining. Compared to existing data-driven baselines, we show that this generality in ClimaX results in superior performance on benchmarks for weather forecasting and climate projections, even when pretrained at lower resolutions and compute budgets. Our source code is available at https://github.com/microsoft/ClimaX. Johannes Brandstetter, Ashish Kapoor, Jayesh K. Gupta, Aditya Grover |
ICML | 3 |
| 2023 | LATTE: LAnguage Trajectory TransformErabstractNatural language is one of the most intuitive ways to express human intent. However, translating instructions and commands towards robotic motion generation and deployment in the real world is far from being an easy task. The challenge of combining a robot's inherent low-level geometric and kinodynamic constraints with a human's high-level semantic instructions traditionally is solved using task-specific solutions with little generalizability between hardware platforms, often with the use of static sets of target actions and commands. This work instead proposes a flexible language-based framework that allows a user to modify generic robotic trajectories. Our method leverages pre-trained language models (BERT and CLIP) to encode the user's intent and target objects directly from a free-form text input and scene images, fuses geometrical features generated by a transformer encoder network, and finally outputs trajectories using a transformer decoder, without the need of priors related to the task or robot information. We significantly extend our own previous work presented in [1] by expanding the trajectory parametrization space to 3D and velocity as opposed to just XY movements. In addition, we now train the model to use actual images of the objects in the scene for context (as opposed to textual descriptions), and we evaluate the system in a diverse set of scenarios beyond manipulation, such as aerial and legged robots. Our simulated and real-life experiments demonstrate that our transformer model can successfully follow human intent, modifying the shape and speed of trajectories within multiple environments. Codebase avail-able at: https://github.com/arthurfenderbucker/LaTTe-Language-Trajectory-TransformEr.git. Arthur Bucker, Luis Figueredo 0001, Sami Haddadin, Ashish Kapoor, Sai Vemprala, Rogerio Bonatti |
ICRA | 4 |
| 2023 | PACT: Perception-Action Causal Transformer for Autoregressive Robotics Pre-TrainingabstractRobotics has long been a field riddled with complex systems architectures whose modules and connections, whether traditional or learning-based, require significant human expertise and prior knowledge. Inspired by large pre-trained language models, this work introduces a paradigm for pretraining a general purpose representation that can serve as a starting point for multiple tasks on a given robot. We present the Perception-Action Causal Transformer (PACT), a generative transformer-based architecture that aims to build representations directly from robot data in a self-supervised fashion. Through autoregressive prediction of states and actions over time, our model implicitly encodes dynamics and behaviors for a particular robot. Our experimental evaluation focuses on the domain of mobile agents, where we show that this robot-specific representation can function as a single starting point to achieve distinct tasks such as safe navigation, localization and mapping. We evaluate two form factors: a wheeled robot that uses a LiDAR sensor as perception input (MuSHR), and a simulated agent that uses first-person RGB images (Habitat). We show that finetuning small task-specific networks on top of the larger pretrained model results in significantly better performance compared to training a single model from scratch for all tasks simultaneously, and comparable performance to training a separate large model for each task independently. By sharing a common good-quality representation across tasks we can lower overall model capacity and speed up the real-time deployment of such systems. Rogerio Bonatti, Sai Vemprala, Felipe Vieira Frujeri, Ashish Kapoor |
IROS | 6 |
| 2022 | Reshaping Robot Trajectories Using Natural Language Commands: A Study of Multi-Modal Data Alignment Using TransformersabstractNatural language is the most intuitive medium for us to interact with other people when expressing commands and instructions. However, using language is seldom an easy task when humans need to express their intent towards robots, since most of the current language interfaces require rigid templates with a static set of action targets and commands. In this work, we provide a flexible language-based interface for human-robot collaboration, which allows a user to reshape existing trajectories for an autonomous agent. We take advantage of recent advancements in the field of large language models (BERT and CLIP) to encode the user command, and then combine these features with trajectory information using multi-modal attention transformers. We train the model using imitation learning over a dataset containing robot trajectories modified by language commands, and treat the trajectory generation process as a sequence prediction problem, analogously to how language generation architectures operate. We evaluate the system in multiple simulated trajectory scenarios, and show a significant performance increase of our model over baseline approaches. In addition, our real-world experiments with a robot arm show that users significantly prefer our natural language interface over traditional methods such as kinesthetic teaching or cost-function programming. Our study shows how the field of robotics can take advantage of large pre-trained language models towards creating more intuitive interfaces between robots and machines. Project webpage: https://arthurfenderbucker.github.io/NL_trajectory_reshaper/ Arthur Bucker, Luis Figueredo 0001, Sami Haddadin, Ashish Kapoor, Rogerio Bonatti |
IROS | 4 |
| 2022 | Learning to Simulate Realistic LiDARsabstractSimulating realistic sensors is a challenging part in data generation for autonomous systems, often involving carefully handcrafted sensor design, scene properties, and physics modeling. To alleviate this, we introduce a pipeline for data-driven simulation of a realistic LiDAR sensor. We propose a model that learns a mapping between RGB images and corresponding LiDAR features such as raydrop or perpoint intensities directly from real datasets. We show that our model can learn to encode realistic effects such as dropped points on transparent surfaces or high intensity returns on reflective materials. When applied to naively raycasted point clouds provided by off-the-shelf simulator software, our model enhances the data by predicting intensities and removing points based on the scene's appearance to match a real LiDAR sensor. We use our technique to learn models of two distinct LiDAR sensors and use them to improve simulated LiDAR data accordingly. Through a sample task of vehicle segmentation, we show that enhancing simulated point clouds with our technique improves downstream task performance. Benoît Guillard, Sai Vemprala, Jayesh K. Gupta, Ondrej Miksik, Vibhav Vineet, Pascal Fua, Ashish Kapoor |
IROS | 7 |
| 2022 | COMPASS: Contrastive Multimodal Pretraining for Autonomous SystemsabstractLearning representations that generalize across tasks and domains is challenging yet necessary for autonomous systems. Although task-driven approaches are appealing, de-signing models specific to each application can be difficult in the face of limited data, especially when dealing with highly variable multimodal input spaces arising from different tasks in different environments. We introduce the first general-purpose pretraining pipeline, COntrastive Multimodal Pretraining for AutonomouS Systems (COMPASS), to overcome the limitations of task-specific models and existing pretraining approaches. COMPASS constructs a multimodal graph by considering the essential information for autonomous systems and the proper-ties of different modalities. Through this graph, multimodal signals are connected and mapped into two factorized spatio-temporal latent spaces: a “motion pattern space” and a “current state space.” By learning from multimodal correspondences in each latent space, COMPASS creates state representations that models necessary information such as temporal dynamics, geometry, and semantics. We pretrain COMPASS on a large-scale multimodal simulation dataset TartanAir [1] and evaluate it on drone navigation, vehicle racing, and visual odometry tasks. The experiments indicate that COMPASS can tackle all three scenarios and can also generalize to unseen environments and real-world data.11Our code implementation can be found at https://github.com/microsoft/COMPASS Sai Vemprala, Jayesh K. Gupta, Yale Song, Daniel McDuff, Ashish Kapoor |
IROS | 7 |
| 2022 | Learning Modular Simulations for Homogeneous SystemsabstractComplex systems are often decomposed into modular subsystems for engineering tractability. Although various equation based white-box modeling techniques make use of such structure, learning based methods have yet to incorporate these ideas broadly. We present a modular simulation framework for modeling homogeneous multibody dynamical systems, which combines ideas from graph neural networks and neural differential equations. We learn to model the individual dynamical subsystem as a neural ODE module. Full simulation of the composite system is orchestrated via spatio-temporal message passing between these modules. An arbitrary number of modules can be combined to simulate systems of a wide variety of coupling topologies. We evaluate our framework on a variety of systems and show that message passing allows coordination between multiple modules over time for accurate predictions and in certain cases, enables zero-shot generalization to new system configurations. Furthermore, we show that our models can be transferred to new system configurations with lower data requirement and training effort, compared to those trained from scratch. Jayesh K. Gupta, Sai Vemprala, Ashish Kapoor |
NeurIPS | 3 |
| 2022 | 3DB: A Framework for Debugging Computer Vision ModelsabstractWe introduce 3DB: an extendable, unified framework for testing and debugging vision models using photorealistic simulation. We demonstrate, through a wide range of use cases, that 3DB allows users to discover vulnerabilities in computer vision systems and gain insights into how models make decisions. 3DB captures and generalizes many robustness analyses from prior work, and enables one to study their interplay. Finally, we find that the insights generated by the system transfer to the physical world. 3DB will be released as a library alongside a set of examples and documentation. We attach 3DB to the submission. Guillaume Leclerc, Hadi Salman, Andrew Ilyas, Sai Vemprala, Logan Engstrom, Vibhav Vineet, Kai Yuanqing Xiao, Pengchuan Zhang, Shibani Santurkar, Greg Yang, Ashish Kapoor, Aleksander Madry |
NeurIPS | 11 |
| 2022 | Sample-Efficient Safe Learning for Online Nonlinear Control with Control Barrier Functions
Ashish Kapoor |
WAFR | 3 |
| 2021 | Quantum algorithms for reinforcement learning with a generative modelabstractReinforcement learning studies how an agent should interact with an environment to maximize its cumulative reward. A standard way to study this question abstractly is to ask how many samples an agent needs from the environment to learn an optimal policy for a $\gamma$-discounted Markov decision process (MDP). For such an MDP, we design quantum algorithms that approximate an optimal policy ($\pi^*$), the optimal value function ($v^*$), and the optimal $Q$-function ($q^*$), assuming the algorithms can access samples from the environment in quantum superposition. This assumption is justified whenever there exists a simulator for the environment; for example, if the environment is a video game or some other program. Our quantum algorithms, inspired by value iteration, achieve quadratic speedups over the best-possible classical sample complexities in the approximation accuracy ($\epsilon$) and two main parameters of the MDP: the effective time horizon ($\frac{1}{1-\gamma}$) and the size of the action space ($A$). Moreover, we show that our quantum algorithm for computing $q^*$ is optimal by proving a matching quantum lower bound. Daochen Wang, Aarthi Sundaram, Robin Kothari, Ashish Kapoor, Martin Rötteler |
ICML | 4 |
| 2021 | Adversarial Attacks on Optimization based PlannersabstractTrajectory planning is a key piece in the algorithmic architecture of a robot. Trajectory planners typically use iterative optimization schemes for generating smooth trajectories that avoid collisions and are optimal for tracking given the robot’s physical specifications. Starting from an initial estimate, the planners iteratively refine the solution so as to satisfy the desired constraints. In this paper, we show that such iterative optimization based planners can be vulnerable to adversarial attacks that force the planner either to fail completely, or significantly increase the time required to find a solution. The key insight here is that an adversary in the environment can directly affect the optimization cost function of a planner. We demonstrate how the adversary can adjust its own state configurations to result in poorly conditioned eigenstructure of the objective leading to failures. We apply our method against two state of the art trajectory planners and demonstrate that an adversary can consistently exploit certain weaknesses of an iterative optimization scheme. Sai Vemprala, Ashish Kapoor |
ICRA | 2 |
| 2021 | Modeling Affect-based Intrinsic Rewards for Exploration and LearningabstractPositive affect has been linked to increased interest, curiosity and satisfaction in human learning. In reinforcement learning, extrinsic rewards are often sparse and difficult to define, intrinsically motivated learning can help address these challenges. We argue that positive affect is an important intrinsic reward that effectively helps drive exploration that is useful in gathering experiences. We present a novel approach leveraging a task-independent reward function trained on spontaneous smile behavior that reflects the intrinsic reward of positive affect. To evaluate our approach we trained several downstream computer vision tasks on data collected with our policy and several baseline methods. We show that the policy based on our affective rewards successfully increases the duration of episodes, the area explored and reduces collisions. The impact is the increased speed of learning for several downstream computer vision tasks. Dean Zadok, Daniel McDuff, Ashish Kapoor |
ICRA | 3 |
| 2021 | Unadversarial Examples: Designing Objects for Robust VisionabstractWe study a class of computer vision settings wherein one can modify the design of the objects being recognized. We develop a framework that leverages this capability---and deep networks' unusual sensitivity to input perturbations---to design ``robust objects,'' i.e., objects that are explicitly optimized to be confidently classified. Our framework yields improved performance on standard benchmarks, a simulated robotics environment, and physical-world experiments. Hadi Salman, Andrew Ilyas, Logan Engstrom, Sai Vemprala, Aleksander Madry, Ashish Kapoor |
NeurIPS | 6 |
| 2021 | Representation Learning for Event-based Visuomotor PoliciesabstractEvent-based cameras are dynamic vision sensors that provide asynchronous measurements of changes in per-pixel brightness at a microsecond level. This makes them significantly faster than conventional frame-based cameras, and an appealing choice for high-speed robot navigation. While an interesting sensor modality, this asynchronously streamed event data poses a challenge for machine learning based computer vision techniques that are more suited for synchronous, frame-based data. In this paper, we present an event variational autoencoder through which compact representations can be learnt directly from asynchronous spatiotemporal event data. Furthermore, we show that such pretrained representations can be used for event-based reinforcement learning instead of end-to-end reward driven perception. We validate this framework of learning event-based visuomotor policies by applying it to an obstacle avoidance scenario in simulation. Compared to techniques that treat event data as images, we show that representations learnt from event streams result in faster policy training, adapt to different control capacities, and demonstrate a higher degree of robustness to environmental changes and sensor noise. Sai Vemprala, Sami Mian, Ashish Kapoor |
NeurIPS | 3 |
| 2020 | Learning Visuomotor Policies for Aerial Navigation Using Cross-Modal RepresentationsabstractMachines are a long way from robustly solving open-world perception-control tasks, such as first-person view (FPV) aerial navigation. While recent advances in end-to- end Machine Learning, especially Imitation Learning and Reinforcement appear promising, they are constrained by the need of large amounts of difficult-to-collect labeled real- world data. Simulated data, on the other hand, is easy to generate, but generally does not render safe behaviors in diverse real-life scenarios. In this work we propose a novel method for learning robust visuomotor policies for real-world deployment which can be trained purely with simulated data. We develop rich state representations that combine supervised and unsupervised environment data. Our approach takes a cross-modal perspective, where separate modalities correspond to the raw camera data and the system states relevant to the task, such as the relative pose of gates to the drone in the case of drone racing. We feed both data modalities into a novel factored architecture, which learns a joint lowdimensional embedding via Variational Auto Encoders. This compact representation is then fed into a control policy, which we trained using imitation learning with expert trajectories in a simulator. We analyze the rich latent spaces learned with our proposed representations, and show that the use of our cross-modal architecture significantly improves control policy performance as compared to end-to-end learning or purely unsupervised feature extractors. We also present real-world results for drone navigation through gates in different track configurations and environmental conditions. Our proposed method, which runs fully onboard, can successfully generalize the learned representations and policies across simulation and reality, significantly outperforming baseline approaches. Rogerio Bonatti, Ratnesh Madaan, Vibhav Vineet, Sebastian A. Scherer, Ashish Kapoor |
IROS | 5 |
| 2020 | Safety Considerations in Deep Control Policies with Safety Barrier Certificates Under UncertaintyabstractRecent advances in Deep Machine Learning have shown promise in solving complex perception and control loops via methods such as reinforcement and imitation learning. However, guaranteeing safety for such learned deep policies has been a challenge due to issues such as partial observability and difficulties in characterizing the behavior of the neural networks. While a lot of emphasis in safe learning has been placed during training, it is non-trivial to guarantee safety at deployment or test time. This paper extends how under mild assumptions, Safety Barrier Certificates can be used to guarantee safety with deep control policies despite uncertainty arising due to perception and other latent variables. Specifically for scenarios where the dynamics are smooth and uncertainty has a finite support, the proposed framework wraps around an existing deep control policy and generates safe actions by dynamically evaluating and modifying the policy from the embedded network. Our framework utilizes control barrier functions to create spaces of control actions that are safe under uncertainty, and when the original actions are found to be in violation of the safety constraint, uses quadratic programming to minimally modify the original actions to ensure they lie in the safe set. Representations of the environment are built through Euclidean signed distance fields that are then used to infer the safety of actions and to guarantee forward invariance. We implement this method in simulation in a drone-racing environment and show that our method results in safer actions compared to a baseline that only relies on imitation learning to generate control actions. Tom Hirshberg, Sai Vemprala, Ashish Kapoor |
IROS | 3 |
| 2020 | TartanAir: A Dataset to Push the Limits of Visual SLAMabstractWe present a challenging dataset, the TartanAir, for robot navigation tasks and more. The data is collected in photo-realistic simulation environments with the presence of moving objects, changing light and various weather conditions. By collecting data in simulations, we are able to obtain multi-modal sensor data and precise ground truth labels such as the stereo RGB image, depth image, segmentation, optical flow, camera poses, and LiDAR point cloud. We set up large numbers of environments with various styles and scenes, covering challenging viewpoints and diverse motion patterns that are difficult to achieve by using physical data collection platforms. In order to enable data collection at such a large scale, we develop an automatic pipeline, including mapping, trajectory sampling, data processing, and data verification. We evaluate the impact of various factors on visual SLAM algorithms using our data. The results of state-of-the-art algorithms reveal that the visual SLAM problem is far from solved. Methods that show good performance on established datasets such as KITTI do not perform well in more difficult scenarios. Although we use the simulation, our goal is to push the limits of Visual SLAM algorithms in the real world by providing a challenging benchmark for testing new methods, while also using a large diverse training data for learning-based methods. Our dataset is available at http://theairlab.org/tartanair-dataset. Delong Zhu 0001, Yaoyu Hu, Yuheng Qiu, Chen Wang 0033, Yafei Hu, Ashish Kapoor, Sebastian A. Scherer |
IROS | 8 |
| 2020 | Multi-Robot Collision Avoidance under Uncertainty with Probabilistic Safety Barrier CertificatesabstractSafety in terms of collision avoidance for multi-robot systems is a difficult challenge under uncertainty, non-determinism, and lack of complete information. This paper aims to propose a collision avoidance method that accounts for both measurement uncertainty and motion uncertainty. In particular, we propose Probabilistic Safety Barrier Certificates (PrSBC) using Control Barrier Functions to define the space of admissible control actions that are probabilistically safe with formally provable theoretical guarantee. By formulating the chance constrained safety set into deterministic control constraints with PrSBC, the method entails minimally modifying an existing controller to determine an alternative safe controller via quadratic programming constrained to PrSBC constraints. The key advantage of the approach is that no assumptions about the form of uncertainty are required other than finite support, also enabling worst-case guarantees. We demonstrate effectiveness of the approach through experiments on realistic simulation environments. Ashish Kapoor |
NeurIPS | 3 |
| 2020 | Do Adversarially Robust ImageNet Models Transfer Better?abstractTransfer learning is a widely-used paradigm in deep learning, where models pre-trained on standard datasets can be efficiently adapted to downstream tasks. Typically, better pre-trained models yield better transfer results, suggesting that initial accuracy is a key aspect of transfer learning performance. In this work, we identify another such aspect: we find that adversarially robust models, while less accurate, often perform better than their standard-trained counterparts when used for transfer learning. Specifically, we focus on adversarially robust ImageNet classifiers, and show that they yield improved accuracy on a standard suite of downstream classification tasks. Further analysis uncovers more differences between robust and standard models in the context of transfer learning. Our results are consistent with (and in fact, add to) recent hypotheses stating that robustness leads to improved feature representations. Code and models is available in the supplementary material. Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor, Aleksander Madry |
NeurIPS | 4 |
| 2020 | Denoised Smoothing: A Provable Defense for Pretrained ClassifiersabstractWe present a method for provably defending any pretrained image classifier against $\ell_p$ adversarial attacks. This method, for instance, allows public vision API providers and users to seamlessly convert pretrained non-robust classification services into provably robust ones. By prepending a custom-trained denoiser to any off-the-shelf image classifier and using randomized smoothing, we effectively create a new classifier that is guaranteed to be $\ell_p$-robust to adversarial examples, without modifying the pretrained classifier. Our approach applies to both the white-box and the black-box settings of the pretrained classifier. We refer to this defense as denoised smoothing, and we demonstrate its effectiveness through extensive experimentation on ImageNet and CIFAR-10. Finally, we use our approach to provably defend the Azure, Google, AWS, and ClarifAI image classification APIs. Our code replicating all the experiments in the paper can be found at: https://github.com/microsoft/denoised-smoothing. Hadi Salman, Mingjie Sun, Greg Yang, Ashish Kapoor, J. Zico Kolter |
NeurIPS | 4 |
| 2020 | Synthetic Examples Improve Generalization for Rare ClassesabstractThe ability to detect and classify rare occurrences in images has important applications - for example, counting rare and endangered species when studying biodiversity, or detecting infrequent traffic scenarios that pose a danger to self-driving cars. Few-shot learning is an open problem: current computer vision systems struggle to categorize objects they have seen only rarely during training, and collecting a sufficient number of training examples of rare events is often challenging and expensive, and sometimes outright impossible. We explore in depth an approach to this problem: complementing the few available training images with ad-hoc simulated data.Our testbed is animal species classification, which has a real-world long-tailed distribution. We present two natural world simulators, and analyze the effect of different axes of variation in simulation, such as pose, lighting, model, and simulation method, and we prescribe best practices for efficiently incorporating simulated data for real-world performance gain. Our experiments reveal that synthetic data can considerably reduce error rates for classes that are rare, that as the amount of simulated data is increased, accuracy on the target class improves, and that high variation of simulated data provides maximum performance gain. Sara Beery, Dan Morris 0001, James Piavis, Ashish Kapoor, Markus Meister, Neel Joshi, Pietro Perona |
WACV | 5 |
| 2020 | BIRDSAI: A Dataset for Detection and Tracking in Aerial Thermal Infrared VideosabstractMonitoring of protected areas to curb illegal activities like poaching and animal trafficking is a monumental task. To augment existing manual patrolling efforts, unmanned aerial surveillance using visible and thermal infrared (TIR) cameras is increasingly being adopted. Automated data acquisition has become easier with advances in unmanned aerial vehicles (UAVs) and sensors like TIR cameras, which allow surveillance at night when poaching typically occurs. However, it is still a challenge to accurately and quickly process large amounts of the resulting TIR data. In this paper, we present the first large dataset collected using a TIR camera mounted on a fixed-wing UAV in multiple African protected areas. This dataset includes TIR videos of humans and animals with several challenging scenarios like scale variations, background clutter due to thermal reflections, large camera rotations, and motion blur. Additionally, we provide another dataset with videos synthetically generated with the publicly available Microsoft AirSim simulation platform using a 3D model of an African savanna and a TIR camera model. Through our benchmarking experiments on state-of-the-art detectors, we demonstrate that leveraging the synthetic data in a domain adaptive setting can significantly improve detection performance. We also evaluate various recent approaches for single and multi-object tracking. With the increasing popularity of aerial imagery for monitoring and surveillance purposes, we anticipate this unique dataset to be used to develop and evaluate techniques for object detection, tracking, and domain adaptation for aerial, TIR videos. Elizabeth Bondi-Kelly, Raghav Jain, Palash Aggrawal, Saket Anand, Robert Hannaford, Ashish Kapoor, James Piavis, Shital Shah, Lucas Joppa, Bistra Dilkina, Milind Tambe |
WACV | 6 |
| 2019 | Visceral Machines: Risk-Aversion in Reinforcement Learning with Intrinsic Physiological Rewards
Daniel McDuff, Ashish Kapoor |
ICLR (Poster) | 2 |
| 2019 | Inverse Optimal Planning for Air Traffic ControlabstractWe envision a system that concisely describes the rules of air traffic control, assists human operators and supports dense autonomous air traffic around commercial airports. We develop a method to learn the rules of air traffic control from real data as a cost function via maximum entropy inverse reinforcement learning. This cost function is used as a penalty for a search-based motion planning method that discretizes both the control and the state space. We illustrate the methodology by showing that our approach can learn to imitate the airport arrival routes and separation rules of dense commercial air traffic. The resulting trajectories are shown to be safe, feasible, and efficient. Kate Tolstaya, Alejandro Ribeiro, Vijay Kumar 0001, Ashish Kapoor |
IROS | 4 |
| 2019 | Bias Correction of Learned Generative Models using Likelihood-Free Importance WeightingabstractA learned generative model often produces biased statistics relative to the underlying data distribution. A standard technique to correct this bias is importance sampling, where samples from the model are weighted by the likelihood ratio under model and true distributions. When the likelihood ratio is unknown, it can be estimated by training a probabilistic classifier to distinguish samples from the two distributions. We employ this likelihood-free importance weighting method to correct for the bias in generative models. We find that this technique consistently improves standard goodness-of-fit metrics for evaluating the sample quality of state-of-the-art deep generative models, suggesting reduced bias. Finally, we demonstrate its utility on representative applications in a) data augmentation for classification using generative adversarial networks, and b) model-based policy evaluation using off-policy data. Aditya Grover, Jiaming Song, Ashish Kapoor, Kenneth Tran, Alekh Agarwal, Eric Horvitz, Stefano Ermon |
NeurIPS | 3 |
| 2019 | Characterizing Bias in Classifiers using Generative ModelsabstractModels that are learned from real-world data are often biased because the data used to train them is biased. This can propagate systemic human biases that exist and ultimately lead to inequitable treatment of people, especially minorities. To characterize bias in learned classifiers, existing approaches rely on human oracles labeling real-world examples to identify the "blind spots" of the classifiers; these are ultimately limited due to the human labor required and the finite nature of existing image examples. We propose a simulation-based approach for interrogating classifiers using generative adversarial models in a systematic manner. We incorporate a progressive conditional generative model for synthesizing photo-realistic facial images and Bayesian Optimization for an efficient interrogation of independent facial image classification systems. We show how this approach can be used to efficiently characterize racial and gender biases in commercial systems. Daniel McDuff, Yale Song, Ashish Kapoor |
NeurIPS | 4 |
| 2019 | Explorations and Lessons Learned in Building an Autonomous Formula SAE Car from SimulationsabstractThis paper describes the exploration and learnings during the process of developing a self-driving algorithm in simulation, followed by deployment on a real car. We specifically concentrate on the Formula Student Driverless competition. In such competitions, a formula race car, designed and built by students, is challenged to drive through previously unseen tracks that are marked by traffic cones. We explore and highlight the challenges associated with training a deep neural network that uses a single camera as input for inferring car steering angles in real-time. The paper explores in-depth creation of simulation, usage of simulations to train and validate the software stack and then finally the engineering challenges associated with the deployment of the system in real-world. Dean Zadok, Tom Hirshberg, Amir Biran, Kira Radinsky, Ashish Kapoor |
SIMULTECH | 5 |
| 2018 | AirSim-W: A Simulation Environment for Wildlife Conservation with UAVsabstractIncreases in poaching levels have led to the use of unmanned aerial vehicles (UAVs or drones) to count animals, locate animals in parks, and even find poachers. Finding poachers is often done at night through the use of long wave thermal infrared cameras mounted on these UAVs. Unfortunately, monitoring the live video stream from the conservation UAVs all night is an arduous task. In order to assist in this monitoring task, new techniques in computer vision have been developed. This work is based on a dataset which took approximately six months to label. However, further improvement in detection and future testing of autonomous flight require not only more labeled training data, but also an environment where algorithms can be safely tested. In order to meet both goals efficiently, we present AirSim-W, a simulation environment that has been designed specifically for the domain of wildlife conservation. This includes (i) creation of an African savanna environment in Unreal Engine, (ii) integration of a new thermal infrared model based on radiometry, (iii) API code expansions to follow objects of interest or fly in zig-zag patterns to generate simulated training data, and (iv) demonstrated detection improvement using simulated data generated by AirSim-W. With these additional simulation features, AirSim-W will be directly useful for wildlife conservation research. Elizabeth Bondi-Kelly, Debadeepta Dey, Ashish Kapoor, James Piavis, Shital Shah, Fei Fang 0001, Bistra Dilkina, Robert Hannaford, Arvind Iyer, Lucas Joppa, Milind Tambe |
COMPASS | 3 |
| 2018 | Learn-to-Score: Efficient 3D Scene Exploration by Predicting View Utility
Benjamin Hepp, Debadeepta Dey, Sudipta N. Sinha, Ashish Kapoor, Neel Joshi, Otmar Hilliges |
ECCV (15) | 4 |
| 2018 | Verifying Controllers Against Adversarial Examples with Bayesian OptimizationabstractRecent successes in reinforcement learning have lead to the development of complex controllers for realworld robots. As these robots are deployed in safety-critical applications and interact with humans, it becomes critical to ensure safety in order to avoid causing harm. A first step in this direction is to test the controllers in simulation. To be able to do this, we need to capture what we mean by safety and then efficiently search the space of all behaviors to see if they are safe. In this paper, we present an active-testing framework based on Bayesian Optimization. We specify safety constraints using logic and exploit structure in the problem in order to test the system for adversarial counter examples that violate the safety specifications. These specifications are defined as complex boolean combinations of smooth functions on the trajectories and, unlike reward functions in reinforcement learning, are expressive and impose hard constraints on the system. In our framework, we exploit regularity assumptions on individual functions in form of a Gaussian Process (GP) prior. We combine these into a coherent optimization framework using problem structure. The resulting algorithm is able to provably verify complex safety specifications or alternatively find counter examples. Experimental results show that the proposed method is able to find adversarial examples quickly. Shromona Ghosh, Felix Berkenkamp, Gireeja Ranade, Shaz Qadeer, Ashish Kapoor |
ICRA | 5 |
| 2018 | Near Real-Time Detection of Poachers from Drones in AirSimabstractThe unrelenting threat of poaching has led to increased development of new technologies to combat it. One such example is the use of thermal infrared cameras mounted on unmanned aerial vehicles (UAVs or drones) to spot poachers at night and report them to park rangers before they are able to harm any animals. However, monitoring the live video stream from these conservation UAVs all night is an arduous task. Therefore, we discuss SPOT (Systematic Poacher deTector), a novel application that augments conservation drones with the ability to automatically detect poachers and animals in near real time. SPOT illustrates the feasibility of building upon state-of-the-art AI techniques, such as Faster RCNN, to address the challenges of automatically detecting animals and poachers in infrared images. This paper reports (i) the design of SPOT, (ii) efficient processing techniques to ensure usability in the field, (iii) evaluation of SPOT based on historical videos and a real-world test run by the end-users, Air Shepherd, in the field, and (iv) the use of AirSim for live demonstration of SPOT. The promising results from a field test have led to a plan for larger-scale deployment in a national park in southern Africa. While SPOT is developed for conservation drones, its design and novel techniques have wider application for automated detection from UAV videos. Elizabeth Bondi-Kelly, Ashish Kapoor, Debadeepta Dey, James Piavis, Shital Shah, Robert Hannaford, Arvind Iyer, Lucas Joppa, Milind Tambe |
IJCAI | 2 |
| 2018 | Enabling a Nationwide Radio Frequency Inventory Using the Spectrum ObservatoryabstractKnowledge about active radio transmitters is critical for multiple applications: spectrum regulators can use this information to assign spectrum, licensees can identify spectrum usage patterns and provision their future needs, and dynamic spectrum access applications can efficiently pick operating frequency. To achieve these goals, we need a system that continuously senses and characterizes the radio spectrum. Current measurement systems, however, do not scale over time, frequency and space and cannot perform transmitter detection. We address these challenges with theSpectrum Observatory, an end-to-end system for spectrum measurement and characterization. This paper details the design and integration of the Spectrum Observatory, and describes and evaluates the first unsupervised method for detailed characterization of arbitrary transmitters calledTxMiner. We evaluate TxMiner on real-world spectrum measurements collected by the Spectrum Observatory between 30 MHz and 6 GHz and show that it identifies transmitters robustly. Furthermore, we demonstrate the Spectrum Observatory’s capabilities to map the number of active transmitters and their frequency and temporal characteristics, to detect rogue transmitters, and identify opportunities for dynamic spectrum access. Mariya Zheleva, Ranveer Chandra, Aakanksha Chowdhery, Paul Garnett, Anoop Gupta, Ashish Kapoor, Matt Valerio |
IEEE Trans. Mob. Comput. | 6 |
| 2017 | Wi-Fly: Widespread Opportunistic Connectivity via Commercial Air TransportabstractMore than half of the world's population face barriers in accessing the Internet. A recent ITU study estimates that 2.6 billion people cannot afford connectivity and that 3.8 billion do not have access. Recent proposals for providing low-cost connectivity include fielding of drones and long-lasting balloons in the stratosphere. We propose a more economical alternative, which we refer to as Wi-Fly, that leverages existing commercial planes to provide Internet connectivity to remote regions. In Wi-Fly we enable communication between a lightweight Wi-Fi device on commercial planes and ground stations, resulting in connectivity in regions that do not otherwise have low-cost Internet connectivity. Wi-Fly leverages existing ADS-B signals from planes as a control channel to ensure that there is a strong link from the plane to the ground, and that the stations intelligently wake up and associate to the appropriate AP. For our experimentation, we have customized two airplanes to conduct measurements. Through empirical experiments with test flights and simulations, we show that Wi-Fly and its extensions have the potential to provide connectivity to the most remote regions of the world at a significantly lower cost than existing alternatives. Talal Ahmad, Ranveer Chandra, Ashish Kapoor, Eric Horvitz |
HotNets | 3 |
| 2017 | Submodular Trajectory Optimization for Aerial 3D ScanningabstractDrones equipped with cameras are emerging as a powerful tool for large-scale aerial 3D scanning, but existing automatic flight planners do not exploit all available information about the scene, and can therefore produce inaccurate and incomplete 3D models. We present an automatic method to generate drone trajectories, such that the imagery acquired during the flight will later produce a high-fidelity 3D model. Our method uses a coarse estimate of the scene geometry to plan camera trajectories that: (1) cover the scene as thoroughly as possible; (2) encourage observations of scene geometry from a diverse set of viewing angles; (3) avoid obstacles; and (4) respect a user-specified flight time budget. Our method relies on a mathematical model of scene coverage that exhibits an intuitive diminishing returns property known as submodularity. We leverage this property extensively to design a trajectory planning algorithm that reasons globally about the non-additive coverage reward obtained across a trajectory, jointly with the cost of traveling between views. We evaluate our method by using it to scan three large outdoor scenes, and we perform a quantitative evaluation using a photorealistic video game simulator. Mike Roberts 0001, Shital Shah, Debadeepta Dey, Anh Truong, Sudipta N. Sinha, Ashish Kapoor, Pat Hanrahan, Neel Joshi |
ICCV | 6 |
| 2017 | Safety-Aware Algorithms for Adversarial Contextual BanditabstractIn this work we study the safe sequential decision making problem under the setting of adversarial contextual bandits with sequential risk constraints. At each round, nature prepares a context, a cost for each arm, and additionally a risk for each arm. The learner leverages the context to pull an arm and receives the corresponding cost and risk associated with the pulled arm. In addition to minimizing the cumulative cost, for safety purposes, the learner needs to make safe decisions such that the average of the cumulative risk from all pulled arms should not be larger than a pre-defined threshold. To address this problem, we first study online convex programming in the full information setting where in each round the learner receives an adversarial convex loss and a convex constraint. We develop a meta algorithm leveraging online mirror descent for the full information setting and then extend it to contextual bandit with sequential risk constraints setting using expert advice. Our algorithms can achieve near-optimal regret in terms of minimizing the total cost, while successfully maintaining a sub-linear growth of accumulative risk constraint violation. We support our theoretical results by demonstrating our algorithm on a simple simulated robotics reactive control task. Debadeepta Dey, Ashish Kapoor |
ICML | 3 |
| 2017 | Learning to gather information via imitationabstractThe budgeted information gathering problem - where a robot with a fixed fuel budget is required to maximize the amount of information gathered from the world - appears in practice across a wide range of applications in autonomous exploration and inspection with mobile robots. Although there is an extensive amount of prior work investigating effective approximations of the problem, these methods do not address the fact that their performance is heavily dependent on distribution of objects in the world. In this paper, we attempt to address this issue by proposing a novel data-driven imitation learning framework. We present an efficient algorithm, EXPLORE, that trains a policy on the target distribution to imitate a clairvoyant oracle - an oracle that has full information about the world and computes non-myopic solutions to maximize information gathered. We validate the approach on a spectrum of results on a number of 2D and 3D exploration problems that demonstrates the ability of EXPLORE to adapt to different object distributions. Additionally, our analysis provides theoretical insight into the behavior of EXPLORE. Our approach paves the way forward for efficiently applying data-driven methods to the domain of information gathering. Sanjiban Choudhury, Ashish Kapoor, Gireeja Ranade, Debadeepta Dey |
ICRA | 2 |
| 2017 | No-regret replanning under uncertaintyabstractThis paper explores the problem of path planning under uncertainty. Specifically, we consider online receding horizon based planners that need to operate in a latent environment where the latent information can be modelled via Gaussian Processes. Online path planning in latent environments is challenging since the robot needs to explore the environment to get a more accurate model of latent information for better planning later and also achieves the task as quick as possible. We propose UCB style algorithms that are popular in the bandit settings and show how those analyses can be adapted to the online robotic path planning problems. The proposed algorithm trades-off exploration and exploitation in near-optimal manner and has appealing no-regret properties. We demonstrate the efficacy of the framework on the application of aircraft flight path planning when the winds are partially observed. Niteesh Sood, Debadeepta Dey, Gireeja Ranade, Siddharth Prakash, Ashish Kapoor |
ICRA | 6 |
| 2017 | Fast second-order cone programming for safe mission planningabstractThis paper considers the problem of safe mission planning of dynamic systems operating under uncertain environments. Much of the prior work on achieving robust and safe control requires solving second-order cone programs (SOCP). Unfortunately, existing general purpose SOCP methods are often infeasible for real-time robotic tasks due to high memory and computational requirements imposed by existing general optimization methods. The key contribution of this paper is a fast and memory-efficient algorithm for SOCP that would enable robust and safe mission planning on-board robots in realtime. Our algorithm does not have any external dependency, can efficiently utilize warm start provided in safe planning settings, and in fact leads to significant speed up over standard optimization packages (like SDPT3) for even standard SOCP problems. For example, for a standard quadrotor problem, our method leads to speedup of 1000× over SDPT3 without any deterioration in the solution quality. Our method is based on two insights: a) SOCPs can be interpreted as optimizing a function over a polytope with infinite sides, b) a linear function can be efficiently optimized over this polytope. We combine the above observations with a novel utilization of Wolfe's algorithm [1] to obtain an efficient optimization method that can be easily implemented on small embedded devices. In addition to the above mentioned algorithm, we also design a two-level sensing method based on Gaussian Process for complex obstacles with non-linear boundaries such as a cylinder. Kai Zhong 0007, Prateek Jain 0002, Ashish Kapoor |
ICRA | 3 |
| 2017 | FarmBeats: An IoT Platform for Data-Driven Agriculture
Deepak Vasisht, Zerina Kapetanovic, Jongho Won, Xinxin Jin, Ranveer Chandra, Sudipta N. Sinha, Ashish Kapoor, Madhusudhan Sudarshan, Sean Stratman |
NSDI | 7 |
| 2016 | Asking for a second opinion: Re-querying of noisy multi-class labelsabstractIn this paper, we propose a new maximum margin-based, active learning algorithm for identifying incorrectly labeled training data. The algorithm combines a round-robin approach for investigating each class with a simple, yet effective ranking metric called maximum negative margin (MNM). Samples are given to an expert for re-evaluation to determine if they are indeed mislabeled. We also propose using five active learning metrics, including uncertainty sampling with margin sampling (USMS) and minimum margin, for the noisy label task which have previously been used in the standard active learning setting for identifying new samples to label. USMS is very competitive with maximum negative margin. In addition, we consider other information theoretic objective criteria for this new task including uncertainty sampling with entropy, query-by-committee with voting entropy, and K-nearest neighbor with voting entropy, but these consistently perform worse than MNM and USMS. The MNM noisy label active learning algorithm can be useful in several different scenarios including data cleansing as a preprocessing step before training and identifying mislabeled examples in the test set. Jack W. Stokes, Ashish Kapoor, Debajyoti Ray |
ICASSP | 2 |
| 2016 | Quantum Perceptron ModelsabstractWe demonstrate how quantum computation can provide non-trivial improvements in the computational and statistical complexity of the perceptron model. We develop two quantum algorithms for perceptron learning. The first algorithm exploits quantum information processing to determine a separating hyperplane using a number of steps sublinear in the number of data points $N$, namely $O(\sqrt{N})$. The second algorithm illustrates how the classical mistake bound of $O(\frac{1}{\gamma^2})$ can be further improved to $O(\frac{1}{\sqrt{\gamma}})$ through quantum means, where $\gamma$ denotes the margin. Such improvements are achieved through the application of quantum amplitude amplification to the version space interpretation of the perceptron model. Ashish Kapoor, Nathan Wiebe, Krysta M. Svore |
NIPS | 1 |
| 2016 | A Joint Gaussian Process Model for Active Visual Recognition with Expertise Estimation in Crowdsourcing
Chengjiang Long, Gang Hua 0001, Ashish Kapoor |
Int. J. Comput. Vis. | 3 |
| 2016 | Scalable Semisupervised Functional Neurocartography Reveals Canonical Neurons in Behavioral NetworksabstractLarge-scale data collection efforts to map the brain are underway at multiple spatial and temporal scales, but all face fundamental problems posed by high-dimensional data and intersubject variability. Even seemingly simple problems, such as identifying a neuron/brain region across animals/subjects, become exponentially more difficult in high dimensions, such as recognizing dozens of neurons/brain regions simultaneously. We present a framework and tools for functional neurocartography-the large-scale mapping of neural activity during behavioral states. Using a voltage-sensitive dye (VSD), we imaged the multifunctional responses of hundreds of leech neurons during several behaviors to identify and functionally map homologous neurons. We extracted simple features from each of these behaviors and combined them with anatomical features to create a rich medium-dimensional feature space. This enabled us to use machine learning techniques and visualizations to characterize and account for intersubject variability, piece together a canonical atlas of neural activity, and identify two behavioral networks. We identified 39 neurons (18 pairs, 3 unpaired) as part of a canonical swim network and 17 neurons (8 pairs, 1 unpaired) involved in a partially overlapping preparatory network. All neurons in the preparatory network rapidly depolarized at the onsets of each behavior, suggesting that it is part of a dedicated rapid-response network. This network is likely mediated by the S cell, and we referenced VSD recordings to an activity atlas to identify multiple cells of interest simultaneously in real time for further experiments. We targeted and electrophysiologically verified several neurons in the swim network and further showed that the S cell is presynaptic to multiple neurons in the preparatory network. This study illustrates the basic framework to map neural activity in high dimensions with large-scale recordings and how to extract the rich information necessary to perform analyses in light of intersubject variability. Edward Paxon Frady, Ashish Kapoor, Eric Horvitz, William B. Kristan Jr. |
Neural Comput. | 2 |
| 2015 | Identifying and Accounting for Task-Dependent Bias in CrowdsourcingabstractModels for aggregating contributions by crowd workers have been shown to be challenged by the rise of task-specific biases or errors. Task-dependent errors in assessment may shift the majority opinion of even large numbers of workers to an incorrect answer. We introduce and evaluate probabilistic models that can detect and correct task-dependent bias automatically. First, we show how to build and use probabilistic graphical models for jointly modeling task features, workers' biases, worker contributions and ground truth answers of tasks so that task-dependent bias can be corrected. Second, we show how the approach can perform a type of transfer learning among workers to address the issue of annotation sparsity. We evaluate the models with varying complexity on a large data set collected from a citizen science project and show that the models are effective at correcting the task-dependent worker bias. Finally, we investigate the use of active learning to guide the acquisition of expert assessments to enable automatic detection and correction of worker bias. Ece Kamar, Ashish Kapoor, Eric Horvitz |
HCOMP | 2 |
| 2015 | The Activity Platform
Helen J. Wang, Alexander Moshchuk, Michael Gamon, Shamsi T. Iqbal, Eli T. Brown, Ashish Kapoor, Christopher Meek, Eric Yawei Chen, Yuan Tian 0001, Jaime Teevan, Mary Czerwinski, Susan T. Dumais |
HotOS | 6 |
| 2015 | On Greedy Maximization of EntropyabstractSubmodular function maximization is one of the key problems that arise in many machine learning tasks. Greedy selection algorithms are the proven choice to solve such problems, where prior theoretical work guarantees (1 - 1/e) approximation ratio. However, it has been empirically observed that greedy selection provides almost optimal solutions in practice. The main goal of this paper is to explore and answer why the greedy selection does significantly better than the theoretical guarantee of (1 - 1/e). Applications include, but are not limited to, sensor selection tasks which use both entropy and mutual information as a maximization criteria. We give a theoretical justification for the nearly optimal approximation ratio via detailed analysis of the curvature of these objective functions for Gaussian RBF kernels. Dravyansh Sharma, Ashish Kapoor, Amit Deshpande 0001 |
ICML | 2 |
| 2015 | Games of drones: an affective game with cyber-physical systemsabstractGames have transformed from PC-based games to more immersive controllers, but gameplay has been limited to a digital screen or a virtual reality headset. In this demo, we present a novel indoor game that allows a player to maneuver the drone altitude by sensing their emotional state using affective sensors. This game demo introduces two innovations in sensing and actuation for next-generation indoor gaming: maneuvering physical objects such as drones for gameplay and sensing the player's emotional state using affective sensors as a control input. Therefore, the immersive experience of players is heightened same as that of real world. We begin to address the systems challenges in building the gaming platform for "Game-of-drones". Aakanksha Chowdhery, Ashish Kapoor, Shivani Bahl |
IPSN | 3 |
| 2015 | A Deep Hybrid Model for Weather ForecastingabstractWeather forecasting is a canonical predictive challenge that has depended primarily on model-based methods. We explore new directions with forecasting weather as a data-intensive challenge that involves inferences across space and time. We study specifically the power of making predictions via a hybrid approach that combines discriminatively trained predictive models with a deep neural network that models the joint statistics of a set of weather-related variables. We show how the base model can be enhanced with spatial interpolation that uses learned long-range spatial dependencies. We also derive an efficient learning and inference procedure that allows for large scale optimization of the model parameters. We evaluate the methods with experiments on real-world meteorological data that highlight the promise of the approach. Aditya Grover, Ashish Kapoor, Eric Horvitz |
KDD | 2 |
| 2014 | Active Learning with Model SelectionabstractMost active learning methods avoid model selection by training models of one type (SVMs, boosted trees, etc.) using one pre-defined set of model hyperparameters. We propose an algorithm that actively samples data to simultaneously train a set of candidate models (different model types and/or different hyperparameters) and also select the best model from this set. The algorithm actively samples points for training that are most likely to improve the accuracy of the more promising candidate models, and also samples points for model selection---all samples count against the same labeling budget. This exposes a natural trade-off between the focused active sampling that is most effective for training models, and the unbiased sampling that is better for model selection. We empirically demonstrate on six test problems that this algorithm is nearly as effective as an active learning oracle that knows the optimal model in advance. Alnur Ali, Rich Caruana, Ashish Kapoor |
AAAI | 3 |
| 2014 | Blind Image Quality Assessment Using Semi-supervised Rectifier NetworksabstractIt is often desirable to evaluate images quality with a perceptually relevant measure that does not require a reference image. Recent approaches to this problem use human provided quality scores with machine learning to learn a measure. The biggest hurdles to these efforts are: 1) the difficulty of generalizing across diverse types of distortions and 2) collecting the enormity of human scored training data that is needed to learn the measure. We present a new blind image quality measure that addresses these difficulties by learning a robust, nonlinear kernel regression function using a rectifier neural network. The method is pre-trained with unlabeled data and fine-tuned with labeled data. It generalizes across a large set of images and distortion types without the need for a large amount of labeled data. We evaluate our approach on two benchmark datasets and show that it not only outperforms the current state of the art in blind image quality estimation, but also outperforms the state of the art in non-blind measures. Furthermore, we show that our semi-supervised approach is robust to using varying amounts of labeled data. Huixuan Tang, Neel Joshi, Ashish Kapoor |
CVPR | 3 |
| 2014 | Airplanes aloft as a sensor network for wind forecasting
Ashish Kapoor, Zachary Horvitz, Spencer Laube, Eric Horvitz |
IPSN | 1 |
| 2014 | Mining text snippets for images on the webabstractImages are often used to convey many different concepts or illustrate many different stories. We propose an algorithm to mine multiple diverse, relevant, and interesting text snippets for images on the web. Our algorithm scales to all images on the web. For each image, all webpages that contain it are considered. The top-K text snippet selection problem is posed as combinatorial subset selection with the goal of choosing an optimal set of snippets that maximizes a combination of relevancy, interestingness, and diversity. The relevancy and interestingness are scored by machine learned models. Our algorithm is run at scale on the entire image index of a major search engine resulting in the construction of a database of images with their corresponding text snippets. We validate the quality of the database through a large-scale comparative study. We showcase the utility of the database through two web-scale applications: (a) augmentation of images on the web as webpages are browsed and (b)~an image browsing experience (similar in spirit to web browsing) that is enabled by interconnecting semantically related images (which may not be visually related) through shared concepts in their corresponding text snippets. Anitha Kannan, Simon Baker, Krishnan Ramnath, Juliet Fiss, Dahua Lin, Lucy Vanderwende, Rizwan Ansary, Ashish Kapoor, Qifa Ke, Matthew Uyttendaele, Xin-Jing Wang, Lei Zhang 0001 |
KDD | 8 |
| 2014 | Active learning for sparse bayesian multilabel classificationabstractWe study the problem of active learning for multilabel classification. We focus on the real-world scenario where the average number of positive (relevant) labels per data point is small leading to positive label sparsity. Carrying out mutual information based near-optimal active learning in this setting is a challenging task since the computational complexity involved is exponential in the total number of labels. We propose a novel inference algorithm for the sparse Bayesian multilabel model of [17]. The benefit of this alternate inference scheme is that it enables a natural approximation of the mutual information objective. We prove that the approximation leads to an identical solution to the exact optimization problem but at a fraction of the optimization cost. This allows us to carry out efficient, non-myopic, and near-optimal active learning for sparse multilabel classification. Extensive experiments reveal the effectiveness of the method. Deepak Vasisht, Andreas Damianou, Manik Varma, Ashish Kapoor |
KDD | 4 |
| 2014 | An Interactive Approach to Solving Correspondence Problems
Stefanie Jegelka, Ashish Kapoor, Eric Horvitz |
Int. J. Comput. Vis. | 2 |
| 2014 | Collaborative Personalization of Image Enhancement
Ashish Kapoor, Juan C. Caicedo, Dani Lischinski, Sing Bing Kang |
Int. J. Comput. Vis. | 1 |
| 2013 | Food and Mood: Just-in-Time Support for Emotional EatingabstractBehavior modification in health is difficult, as habitual behaviors are extremely well-learned, by definition. This research is focused on building a persuasive system for behavior modification around emotional eating. In this paper, we make strides towards building a just-in-time support system for emotional eating in three user studies. The first two studies involved participants using a custom mobile phone application for tracking emotions, food, and receiving interventions. We found lots of individual differences in emotional eating behaviors and that most participants wanted personalized interventions, rather than a pre-determined intervention. Finally, we also designed a novel, wearable sensor system for detecting emotions using a machine learning approach. This system consisted of physiological sensors which were placed into women's brassieres. We tested the sensing system and found positive results for emotion detection in this mobile, wearable system. Erin A. Carroll, Mary Czerwinski, Asta Roseway, Ashish Kapoor, Paul Johns, Kael Rowan, m. c. schraefel |
ACII | 4 |
| 2013 | On Recovering Structure of AffectabstractThis paper presents novel human computation experiments geared towards uncovering the structure of affect. Using Mechanical Turk workers across 2 separate studies, we empirically verified some of the popular beliefs about the structure of affect, but also provide some new evidence. We replicate and reveal not only the statistical structure of the dimensions of affect, but also the effect of cultural influences. We close with a proposition for a framework for doing this kind of large scale research and provide recommendations and opportunities for innovations in research around emotional theory. Ashish Kapoor, Mary Czerwinski, Diana L. MacLean, Alex Zolotovitski |
ACII | 1 |
| 2013 | Active Visual Recognition with Expertise Estimation in CrowdsourcingabstractWe present a noise resilient probabilistic model for active learning of a Gaussian process classifier from crowds, i.e., a set of noisy labelers. It explicitly models both the overall label noises and the expertise level of each individual labeler in two levels of flip models. Expectation propagation is adopted for efficient approximate Bayesian inference of our probabilistic model for classification, based on which, a generalized EM algorithm is derived to estimate both the global label noise and the expertise of each individual labeler. The probabilistic nature of our model immediately allows the adoption of the prediction entropy and estimated expertise for active selection of data sample to be labeled, and active selection of high quality labelers to label the data, respectively. We apply the proposed model for three visual recognition tasks, i.e., object category recognition, gender recognition, and multi-modal activity recognition, on three datasets with real crowd-sourced labels from Amazon Mechanical Turk. The experiments clearly demonstrated the efficacy of the proposed model. Chengjiang Long, Gang Hua 0001, Ashish Kapoor |
ICCV | 3 |
| 2013 | Lifelong Learning for Acquiring the Wisdom of the Crowd
Ece Kamar, Ashish Kapoor, Eric Horvitz |
IJCAI | 2 |
| 2012 | Learning to Learn: Algorithmic Inspirations from Human Problem SolvingabstractWe harness the ability of people to perceive and interact with visual patterns in order to enhance the performance of a machine learning method. We show how we can collect evidence about how people optimize the parameters of an ensemble classification system using a tool that provides a visualization of misclassification costs. Then, we use these observations about human attempts to minimize cost in order to extend the performance of a state-of-the-art ensemble classification system. The study highlights opportunities for learning from evidence collected about human problem solving to refine and extend automated learning and inference. Ashish Kapoor, Bongshin Lee, Desney S. Tan, Eric Horvitz |
AAAI | 1 |
| 2012 | Performance and Preferences: Interactive Refinement of Machine Learning ProceduresabstractProblem-solving procedures have been typically aimed at achieving well-defined goals or satisfying straightforward preferences. However, learners and solvers may often generate rich multiattribute results with procedures guided by sets of controls that define different dimensions of quality. We explore methods that enable people to explore and express preferences about the operation of classification models in supervised multiclass learning. We leverage a leave-one-out confusion matrix that provides users with views and real-time controls of a model space. The approach allows people to consider in an interactive manner the global implications of local changes in decision boundaries. We focus on kernel classifiers and show the effectiveness of the methodology on a variety of tasks. Ashish Kapoor, Bongshin Lee, Desney S. Tan, Eric Horvitz |
AAAI | 1 |
| 2012 | AffectAura: an intelligent system for emotional memoryabstractWe present AffectAura, an emotional prosthetic that allows users to reflect on their emotional states over long periods of time. We designed a multimodal sensor set-up for continuous logging of audio, visual, physiological and contextual data, a classification scheme for predicting user affective state and an interface for user reflection. The system continuously predicts a user's valence, arousal and engage-ment, and correlates this with information on events, communications and data interactions. We evaluate the interface through a user study consisting of six users and over 240 hours of data, and demonstrate the utility of such a reflection tool. We show that users could reason forward and backward in time about their emotional experiences using the interface, and found this useful. Daniel McDuff, Amy K. Karlson, Ashish Kapoor, Asta Roseway, Mary Czerwinski |
CHI | 3 |
| 2012 | Memory constrained face recognitionabstractReal-time recognition may be limited by scarce memory and computing resources for performing classification. Although, prior research has addressed the problem of training classifiers with limited data and computation, few efforts have tackled the problem of memory constraints on recognition. We explore methods that can guide the allocation of limited storage resources for classifying streaming data so as to maximize discriminatory power. We focus on computation of the expected value of information with nearest neighbor classifiers for online face recognition. Experiments on real-world datasets show the effectiveness and power of the approach. The methods provide a principled approach to vision under bounded resources, and have immediate application to enhancing recognition capabilities in consumer devices with limited memory. Ashish Kapoor, Simon Baker, Sumit Basu, Eric Horvitz |
CVPR | 1 |
| 2012 | Context-Based Automatic Local Image Enhancement
Sung Ju Hwang, Ashish Kapoor, Sing Bing Kang |
ECCV (1) | 2 |
| 2012 | Multilabel Classification using Bayesian Compressed SensingabstractIn this paper, we present a Bayesian framework for multilabel classification using compressed sensing. The key idea in compressed sensing for multilabel classification is to first project the label vector to a lower dimensional space using a random transformation and then learn regression functions over these projections. Our approach considers both of these components in a single probabilistic model, thereby jointly optimizing over compression as well as learning tasks. We then derive an efficient variational inference scheme that provides joint posterior distribution over all the unobserved labels. The two key benefits of the model are that a) it can naturally handle datasets that have missing labels and b) it can also measure uncertainty in prediction. The uncertainty estimate provided by the model naturally allows for active learning paradigms where an oracle provides information about labels that promise to be maximally informative for the prediction task. Our experiments show significant boost over prior methods in terms of prediction performance over benchmark datasets, both in the fully labeled and the missing labels case. Finally, we also highlight various useful active learning scenarios that are enabled by the probabilistic model. Ashish Kapoor, Raajay Viswanathan, Prateek Jain 0002 |
NIPS | 1 |
| 2012 | Riffled Independence for Efficient Inference with Partial RankingsabstractDistributions over rankings are used to model data in a multitude of real world settings such as preference analysis and political elections. Modeling such distributions presents several computational challenges, however, due to the factorial size of the set of rankings over an item set. Some of these challenges are quite familiar to the artificial intelligence community, such as how to compactly represent a distribution over a combinatorially large space, and how to efficiently perform probabilistic inference with these representations. With respect to ranking, however, there is the additional challenge of what we refer to as human task complexity users are rarely willing to provide a full ranking over a long list of candidates, instead often preferring to provide partial ranking information. Simultaneously addressing all of these challenges i.e., designing a compactly representable model which is amenable to efficient inference and can be learned using partial ranking data is a difficult task, but is necessary if we would like to scale to problems with nontrivial size. In this paper, we show that the recently proposed riffled independence assumptions cleanly and efficiently address each of the above challenges. In particular, we establish a tight mathematical connection between the concepts of riffled independence and of partial rankings. This correspondence not only allows us to then develop efficient and exact algorithms for performing inference tasks using riffled independence based represen- tations with partial rankings, but somewhat surprisingly, also shows that efficient inference is not possible for riffle independent models (in a certain sense) with observations which do not take the form of partial rankings. Finally, using our inference algorithm, we introduce the first method for learning riffled independence based models from partially ranked data. Jonathan Huang, Ashish Kapoor, Carlos Guestrin |
J. Artif. Intell. Res. | 2 |
| 2011 | Effective End-User Interaction with Machine LearningabstractEnd-user interactive machine learning is a promising tool for enhancing human productivity and capabilities with large unstructured data sets. Recent work has shown that we can create end-user interactive machine learning systems for specific applications. However, we still lack a generalized understanding of how to design effective end-user interaction with interactive machine learning systems. This work presents three explorations in designing for effective end-user interaction with machine learning in CueFlik, a system developed to support Web image search. These explorations demonstrate that interactions designed to balance the needs of end-users and machine learning algorithms can significantly improve the effectiveness of end-user interactive machine learning. Saleema Amershi, James Fogarty, Ashish Kapoor, Desney S. Tan |
AAAI | 3 |
| 2011 | CueT: human-guided fast and accurate network alarm triageabstractNetwork alarm triage refers to grouping and prioritizing a stream of low-level device health information to help operators find and fix problems. Today, this process tends to be largely manual because existing tools cannot easily evolve with the network. We present CueT, a system that uses interactive machine learning to learn from the triaging decisions of operators. It then uses that learning in novel visualizations to help them quickly and accurately triage alarms. Unlike prior interactive machine learning systems, CueT handles a highly dynamic environment where the groups of interest are not known a-priori and evolve constantly. A user study with real operators and data from a large network shows that CueT significantly improves the speed and accuracy of alarm triage compared to the network's current practice. Saleema Amershi, Bongshin Lee, Ashish Kapoor, Ratul Mahajan, Blaine Christian |
CHI | 3 |
| 2011 | Collaborative personalization of image enhancementabstractWhile most existing enhancement tools for photographs have universal auto-enhancement functionality, recent research shows that users can have personalized preferences. In this paper, we explore whether such personalized preferences in image enhancement tend to cluster and whether users can be grouped according to such preferences. To this end, we analyze a comprehensive data set of image enhancements collected from 336 users via Amazon Mechanical Turk. We find that such clusters do exist and can be used to derive methods to learn statistical preference models from a group of users. We also present a probabilistic framework that exploits the ideas behind collaborative filtering to automatically enhance novel images for new users. Experiments show that inferring clusters in image enhancement preferences results in better prediction of image enhancement preferences and outperforms generic auto-correction tools. Juan C. Caicedo, Ashish Kapoor, Sing Bing Kang |
CVPR | 2 |
| 2011 | Learning a blind measure of perceptual image qualityabstractIt is often desirable to evaluate an image based on its quality. For many computer vision applications, a perceptually meaningful measure is the most relevant for evaluation; however, most commonly used measure do not map well to human judgements of image quality. A further complication of many existing image measure is that they require a reference image, which is often not available in practice. In this paper, we present a “blind” image quality measure, where potentially neither the groundtruth image nor the degradation process are known. Our method uses a set of novel low-level image features in a machine learning framework to learn a mapping from these features to subjective image quality scores. The image quality features stem from natural image measure and texture statistics. Experiments on a standard image quality benchmark dataset shows that our method outperforms the current state of art. Huixuan Tang, Neel Joshi, Ashish Kapoor |
CVPR | 3 |
| 2011 | Honest signals in video conferencingabstractWe propose a novel system to analyze gestural and nonverbal cues of participants in video conferencing. These cues have previously been referred to as “honest signals” and are usually associated with the underlying cognitive state of the participants. The presented system analyzes a set of audio-visual, non-linguistic features in real time from the audio and video streams of two participants in a video conference. We show how these features can be used to compute indicators of the overall quality and type of conversation being held. The system also provides visual feedback to the participants, who then have the choice of modifying their conversational style in order to achieve the desired outcome of the video conference. Experiments on real-life data show that the system can predict the type of conversation with high accuracy using the non-linguistic signals only. Qualitative user studies highlight the positive effects of increased awareness amongst the participants about their own gestural and non-verbal cues. Byungki Byun, Anurag Awasthi, Philip A. Chou, Ashish Kapoor, Bongshin Lee, Mary Czerwinski |
ICME | 4 |
| 2011 | Human-Guided Machine Learning for Fast and Accurate Network Alarm Triage
Saleema Amershi, Bongshin Lee, Ashish Kapoor, Ratul Mahajan, Blaine Christian |
IJCAI | 3 |
| 2011 | Using Multiple Models to Understand Data
Kayur Patel, Steven Mark Drucker, James Fogarty, Ashish Kapoor, Desney S. Tan |
IJCAI | 4 |
| 2011 | Active Graph Reachability Reduction for Network Security and Software Engineering
Alice X. Zheng, John Dunagan, Ashish Kapoor |
IJCAI | 3 |
| 2011 | Efficient Probabilistic Inference with Partial Ranking Queries
Jonathan Huang, Ashish Kapoor, Carlos Guestrin |
UAI | 2 |
| 2010 | Examining multiple potential models in end-user interactive concept learningabstractEnd-user interactive concept learning is a technique for interacting with large unstructured datasets, requiring insights from both human-computer interaction and machine learning. This note re-examines an assumption implicit in prior interactive machine learning research, that interaction should focus on the question "what class is this object?". We broaden interaction to include examination of multiple potential models while training a machine learning system. We evaluate this approach and find that people naturally adopt revision in the interactive machine learning process and that this improves the quality of their resulting models for difficult concepts. Saleema Amershi, James Fogarty, Ashish Kapoor, Desney S. Tan |
CHI | 3 |
| 2010 | Interactive optimization for steering machine classificationabstractInterest has been growing within HCI on the use of machine learning and reasoning in applications to classify such hidden states as user intentions, based on observations. HCI researchers with these interests typically have little expertise in machine learning and often employ toolkits as relatively fixed "black boxes" for generating statistical classifiers. However, attempts to tailor the performance of classifiers to specific application requirements may require a more sophisticated understanding and custom-tailoring of methods. We present ManiMatrix, a system that provides controls and visualizations that enable system builders to refine the behavior of classification systems in an intuitive manner. With ManiMatrix, users directly refine parameters of a confusion matrix via an interactive cycle of re-classification and visualization. We present the core methods and evaluate the effectiveness of the approach in a user study. Results show that users are able to quickly and effectively modify decision boundaries of classifiers to tai-lor the behavior of classifiers to problems at hand. Ashish Kapoor, Bongshin Lee, Desney S. Tan, Eric Horvitz |
CHI | 1 |
| 2010 | Personalization of image enhancementabstractWe address the problem of incorporating user preference in automatic image enhancement. Unlike generic tools for automatically enhancing images, we seek to develop methods that can first observe user preferences on a training set, and then learn a model of these preferences to personalize enhancement of unseen images. The challenge of designing such system lies at intersection of computer vision, learning, and usability; we use techniques such as active sensor selection and distance metric learning in order to solve the problem. The experimental evaluation based on user studies indicates that different users do have different preferences in image enhancement, which suggests that personalization can further help improve the subjective quality of generic image enhancements. Sing Bing Kang, Ashish Kapoor, Dani Lischinski |
CVPR | 2 |
| 2010 | Visual recognition and detection under bounded computational resourcesabstractVisual recognition and detection are computationally intensive tasks and current research efforts primarily focus on solving them without considering the computational capability of the devices they run on. In this paper we explore the challenge of deriving methods that consider constraints on computation, appropriately schedule the next best computation to perform and finally have the capability of producing reasonable results at any time when a solution is required. We specifically derive an approach for the task of object category localization and classification in cluttered, natural scenes that can not only produce anytime results but also utilize the principle of value-of-information in order to provide the most recognition bang for the computational buck. Experiments on two standard object detection challenges show that the proposed framework can triage computation effectively and attain state-of-the-art results when allowed to run till completion. Additionally, the real benefit of the proposed framework is highlighted in the experiments where we demonstrate that the method can provide reasonable recognition results even if the procedure needs to terminate before completion. Sudheendra Vijayanarasimhan, Ashish Kapoor |
CVPR | 2 |
| 2010 | Joint People, Event, and Location Recognition in Personal Photo Collections Using Cross-Domain Context
Dahua Lin, Ashish Kapoor, Gang Hua 0001, Simon Baker |
ECCV (1) | 2 |
| 2010 | Modeling Long-Term Search Engine Usage
Ryen W. White, Ashish Kapoor, Susan T. Dumais |
UMAP | 2 |
| 2010 | Gaussian Processes for Object CategorizationabstractDiscriminative methods for visual object category recognition are typically non-probabilistic, predicting class labels but not directly providing an estimate of uncertainty. Gaussian Processes (GPs) provide a framework for deriving regression techniques with explicit uncertainty models; we show here how Gaussian Processes with covariance functions defined based on a Pyramid Match Kernel (PMK) can be used for probabilistic object category recognition. Our probabilistic formulation provides a principled way to learn hyperparameters, which we utilize to learn an optimal combination of multiple covariance functions. It also offers confidence estimates at test points, and naturally allows for an active learning paradigm in which points are optimally selected for interactive labeling. We show that with an appropriate combination of kernels a significant boost in classification performance is possible. Further, our experiments indicate the utility of active learning with probabilistic predictive models, especially when the amount of training data labels that may be sought for a category is ultimately very small. Ashish Kapoor, Kristen Grauman, Raquel Urtasun, Trevor Darrell |
Int. J. Comput. Vis. | 1 |
| 2009 | EnsembleMatrix: interactive visualization to support machine learning with multiple classifiersabstractMachine learning is an increasingly used computational tool within human-computer interaction research. While most researchers currently utilize an iterative approach to refining classifier models and performance, we propose that ensemble classification techniques may be a viable and even preferable alternative. In ensemble learning, algorithms combine multiple classifiers to build one that is superior to its components. In this paper, we present EnsembleMatrix, an interactive visualization system that presents a graphical view of confusion matrices to help users understand relative merits of various classifiers. EnsembleMatrix allows users to directly interact with the visualizations in order to explore and build combination models. We evaluate the efficacy of the system and the approach in a user study. Results show that users are able to quickly combine multiple classifiers operating on multiple feature sets to produce an ensemble classifier with accuracy that approaches best-reported performance classifying images in the CalTech-101 dataset. Justin Talbot, Bongshin Lee, Ashish Kapoor, Desney S. Tan |
CHI | 3 |
| 2009 | Co-training with noisy perceptual observationsabstractMany perception problems involve datasets that are naturally comprised of multiple streams or modalities for which supervised training data is only sparsely available. In cases where there is a degree of conditional independence between such views, a class of semi-supervised learning techniques that are based on maximizing view agreement over unlabeled data has been proven successful in a wide range of machine learning domains. However, these `co-training' or `multi-view' learning methods have had relatively limited application in vision, due in part to the assumption of constant per-channel noise models. In this paper we propose a probabilistic heteroscedastic approach to co-training that simultaneously discovers the amount of noise on a per-sample basis, while solving the classification task. This results in high performance in the presence of occlusion or other complex observation noise processes. We demonstrate our approach in two domains, multi-view object recognition from low-fidelity sensor networks and audio-visual classification. C. Mario Christoudias, Raquel Urtasun, Ashish Kapoor, Trevor Darrell |
CVPR | 3 |
| 2009 | Active learning for large multi-class problemsabstractScarcity and infeasibility of human supervision for large scale multi-class classification problems necessitates active learning. Unfortunately, existing active learning methods for multi-class problems are inherently binary methods and do not scale up to a large number of classes. In this paper, we introduce a probabilistic variant of the K-nearest neighbor method for classification that can be seamlessly used for active learning in multi-class scenarios. Given some labeled training data, our method learns an accurate metric/kernel function over the input space that can be used for classification and similarity search. Unlike existing metric/kernel learning methods, our scheme is highly scalable for classification problems and provides a natural notion of uncertainty over class labels. Further, we use this measure of uncertainty to actively sample training examples that maximize discriminating capabilities of the model. Experiments on benchmark datasets show that the proposed method learns appropriate distance metrics that lead to state-of-the-art performance for object categorization problems. Furthermore, our active learning method effectively samples training examples, resulting in significant accuracy gains over random sampling for multi-class problems involving a large number of classes. Prateek Jain 0002, Ashish Kapoor |
CVPR | 2 |
| 2009 | Which faces to tag: Adding prior constraints into active learningabstractWe introduce an algorithm that guides the user to tag faces in the best possible order during a face recognition assisted tagging scenario. In particular, we extend the active learning paradigm to take advantage of constraints known a priori. For example, in the context of personal photo collections, if two faces come from the same source photograph, we know that they must be of different people. Similarly, in the context of video, we know that the faces from a single track must be of the same person. Given a set of unlabeled images and constraints, we use a probabilistic discriminative model that models the posterior distributions by propagating label information using a message passing scheme. The uncertainty estimate provided by the model naturally allows for active learning paradigms where the user is consulted after each iteration to tag additional faces. Our experiments show that performing active learning while incorporating a priori constraints provides a significant boost in many real-world face recognition tasks. Ashish Kapoor, Gang Hua 0001, Amir Akbarzadeh, Simon Baker |
ICCV | 1 |
| 2009 | Breaking Boundaries Between Induction Time and Diagnosis Time Active Information AcquisitionabstractThere has been a clear distinction between induction or training time and diagnosis time active information acquisition. While active learning during induction focuses on acquiring data that promises to provide the best classification model, the goal at diagnosis time focuses completely on next features to observe about the test case at hand in order to make better predictions about the case. We introduce a model and inferential methods that breaks this distinction. The methods can be used to extend case libraries under a budget but, more fundamentally, provide a framework for guiding agents to collect data under scarce resources, focused by diagnostic challenges. This extension to active learning leads to a new class of policies for real-time diagnosis, where recommended information-gathering sequences include actions that simultaneously seek new data for the case at hand and for cases in the training set. Ashish Kapoor, Eric Horvitz |
NIPS | 1 |
| 2009 | Overview based example selection in end user interactive concept learningabstractInteraction with large unstructured datasets is difficult because existing approaches, such as keyword search, are not always suited to describing concepts corresponding to the distinctions people want to make within datasets. One possible solution is to allow end users to train machine learning systems to identify desired concepts, a strategy known as interactive concept learning. A fundamental challenge is to design systems that preserve end user flexibility and control while also guiding them to provide examples that allow the machine learning system to effectively learn the desired concept. This paper presents our design and evaluation of four new overview based approaches to guiding example selection. We situate our explorations within CueFlik, a system examining end user interactive concept learning in Web image search. Our evaluation shows our approaches not only guide end users to select better training examples than the best performing previous design for this application, but also reduce the impact of not knowing when to stop training the system. We discuss challenges for end user interactive concept learning systems and identify opportunities for future research on the effective design of such systems. Saleema Amershi, James Fogarty, Ashish Kapoor, Desney S. Tan |
UIST | 3 |
| 2008 | CueFlik: interactive concept learning in image searchabstractWeb image search is difficult in part because a handful of keywords are generally insufficient for characterizing the visual properties of an image. Popular engines have begun to provide tags based on simple characteristics of images (such as tags for black and white images or images that contain a face), but such approaches are limited by the fact that it is unclear what tags end users want to be able to use in examining Web image search results. This paper presents CueFlik, a Web image search application that allows end users to quickly create their own rules for re ranking images based on their visual characteristics. End users can then re rank any future Web image search results according to their rule. In an experiment we present in this paper, end users quickly create effective rules for such concepts as "product photos", "portraits of people", and "clipart". When asked to conceive of and create their own rules, participants create such rules as "sports action shot" with images from queries for "basketball" and "football". CueFlik represents both a promising new approach to Web image search and an important study in end user interactive machine learning. James Fogarty, Desney S. Tan, Ashish Kapoor, Simon A. J. Winder |
CHI | 3 |
| 2008 | Experience sampling for building predictive user models: a comparative studyabstractExperience sampling has been employed for decades to collect assessments of subjects' intentions, needs, and affective states. In recent years, investigators have employed automated experience sampling to collect data to build predictive user models. To date, most procedures have relied on random sampling or simple heuristics. We perform a comparative analysis of several automated strategies for guiding experience sampling, spanning a spectrum of sophistication, from a random sampling procedure to increasingly sophisticated active learning. The more sophisticated methods take a decision-theoretic approach, centering on the computation of the expected value of information of a probe, weighing the cost of the short-term disruptiveness of probes with their benefits in enhancing the long-term performance of predictive models. We test the different approaches in a field study, focused on the task of learning predictive models of the cost of interruption. Ashish Kapoor, Eric Horvitz |
CHI | 1 |
| 2008 | Combining brain computer interfaces with vision for object categorizationabstractHuman-aided computing proposes using information measured directly from the human brain in order to perform useful tasks. In this paper, we extend this idea by fusing computer vision-based processing and processing done by the human brain in order to build more effective object categorization systems. Specifically, we use an electroencephalograph (EEG) device to measure the subconscious cognitive processing that occurs in the brain as users see images, even when they are not trying to explicitly classify them. We present a novel framework that combines a discriminative visual category recognition system based on the Pyramid Match Kernel (PMK) with information derived from EEG measurements as users view images. We propose a fast convex kernel alignment algorithm to effectively combine the two sources of information. Our approach is validated with experiments using real-world data, where we show significant gains in classification accuracy. We analyze the properties of this information fusion method by examining the relative contributions of the two modalities, the errors arising from each source, and the stability of the combination in repeated experiments. Ashish Kapoor, Pradeep Shenoy, Desney S. Tan |
CVPR | 1 |
| 2008 | Complementary computing for visual tasks: Meshing computer vision with human visual processingabstractWe explore the opportunity to harness electroencephalograph (EEG) signals generated during human visual processing to enhance computer vision systems. We review the challenging task of categorizing objects, such as faces, in images and then describe methods that can be used to combine the complementary competencies of human and machine computation to achieve improved recognition performance. We present the results of several experiments where brain signals, recorded from people examining images, are used to enhance the performance of vision systems on categorization tasks. We find that significant gains in classification accuracy can be achieved with the human-aided vision systems. Ashish Kapoor, Desney S. Tan, Pradeep Shenoy, Eric Horvitz |
FG | 1 |
| 2007 | Active Learning with Gaussian Processes for Object CategorizationabstractDiscriminative methods for visual object category recognition are typically non-probabilistic, predicting class labels but not directly providing an estimate of uncertainty. Gaussian Processes (GPs) are powerful regression techniques with explicit uncertainty models; we show here how Gaussian Processes with covariance functions defined based on a Pyramid Match Kernel (PMK) can be used for probabilistic object category recognition. The uncertainty model provided by GPs offers confidence estimates at test points, and naturally allows for an active learning paradigm in which points are optimally selected for interactive labeling. We derive a novel active category learning method based on our probabilistic regression model, and show that a significant boost in classification performance is possible, especially when the amount of training data for a category is ultimately very small. Ashish Kapoor, Kristen Grauman, Raquel Urtasun, Trevor Darrell |
ICCV | 1 |
| 2007 | Selective Supervision: Guiding Supervised Learning with Decision-Theoretic Active Learning
Ashish Kapoor, Eric Horvitz, Sumit Basu |
IJCAI | 1 |
| 2007 | On Discarding, Caching, and Recalling Samples in Active Learning
Ashish Kapoor, Eric Horvitz |
UAI | 1 |
| 2007 | Automatic prediction of frustration
Ashish Kapoor, Winslow Burleson, Rosalind W. Picard |
Int. J. Hum. Comput. Stud. | 1 |
| 2006 | Located Hidden Random Fields: Learning Discriminative Parts for Object Detection
Ashish Kapoor, John M. Winn |
ECCV (3) | 1 |
| 2005 | The audio epitome: a new representation for modeling and classifying auditory phenomenaabstractThe paper presents a novel representation for auditory environments that can be used for classifying events of interest, such as speech, cars, etc., and potentially used to classify the environments themselves. We propose a novel discriminative framework that is based on the audio epitome, an audio extension of the image representation developed by N. Jojic et al. (see Proc. Int. Conf. Comp. Vision, 2003). We also develop an informative patch sampling procedure to train the epitomes. This procedure reduces the computational complexity and increases the quality of the epitome. For classification, the training data is used to learn distributions over the epitomes to model the different classes; the distributions for new inputs are then compared to these models. On a task of distinguishing between 4 auditory classes in the context of environmental sounds (car, speech, birds, utensils), our method outperforms the conventional approaches of nearest neighbor and mixture of Gaussians on three out of the four classes. Ashish Kapoor, Sumit Basu |
ICASSP (5) | 1 |
| 2005 | Multimodal affect recognition in learning environmentsabstractWe propose a multi-sensor affect recognition system and evaluate it on the challenging task of classifying interest (or disinterest) in children trying to solve an educational puzzle on the computer. The multimodal sensory information from facial expressions and postural shifts of the learner is combined with information about the learner's activity on the computer. We propose a unified approach, based on a mixture of Gaussian Processes, for achieving sensor fusion under the problematic conditions of missing channels and noisy labels. This approach generates separate class labels corresponding to each individual modality. The final classification is based upon a hidden random variable, which probabilistically combines the sensors. The multimodal Gaussian Process approach achieves accuracy of over 86%, significantly outperforming classification using the individual modalities, and several other combination schemes. Ashish Kapoor, Rosalind W. Picard |
ACM Multimedia | 1 |
| 2005 | Hyperparameter and Kernel Learning for Graph Based Semi-Supervised ClassificationabstractThere have been many graph-based approaches for semi-supervised clas- sification. One problem is that of hyperparameter learning: performance depends greatly on the hyperparameters of the similarity graph, trans- formation of the graph Laplacian and the noise model. We present a Bayesian framework for learning hyperparameters for graph-based semi- supervised classification. Given some labeled data, which can contain inaccurate labels, we pose the semi-supervised classification as an in- ference problem over the unknown labels. Expectation Propagation is used for approximate inference and the mean of the posterior is used for classification. The hyperparameters are learned using EM for evidence maximization. We also show that the posterior mean can be written in terms of the kernel matrix, providing a Bayesian classifier to classify new points. Tests on synthetic and real datasets show cases where there are significant improvements in performance over the existing approaches. Ashish Kapoor, Yuan Qi 0001, Hyungil Ahn, Rosalind W. Picard |
NIPS | 1 |
| 2001 | Audio Driven Facial Animation For Audio-Visual RealityabstractIn this paper, we demonstrate a morphing based automated audio driven facial animation system. Based on an incoming audio stream, a face image is animated with full lip synchronization and expression. An animation sequence using optical flow between visemes is constructed, given an incoming audio stream and still pictures of a face speaking different visemes. Rules are formulated based on coarticulation and the duration of a viseme to control the continuity in terms of shape and extent of lip opening. In addition to this new viseme-expression combinations are synthesized to be able to generate animations with new facial expressions. Finally various applications of this system are discussed in the context of creating audio-visual reality. 1. Tanveer A. Faruquie, Ashish Kapoor, Rohit J. Kate, Nitendra Rajput, L. Venkata Subramaniam |
ICME | 2 |