Ashutosh Saxena

dblp:82/6189 · DBLP profile ↗
← Back
76ranked-venue papers
16as first author
2since 2021 · last 2024
0000-0002-6657-2285ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 14 first-author · 1 since 2021Systems, architecture and hardware · 22 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-authorSecurity and privacy · 8 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
43 papers
3D vision · 20% Video understanding and tracking · 18% Robot manipulation · 16%
Computer graphics and multimedia
4 papers
Multimedia analysis and retrieval · 53% Audio and music processing · 23% Computational photography and imaging · 17%

Topics — the 30 heaviest of 85, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d scene understanding
1.162016
Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016
3D Reasoning from Blocks to Stability · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Hallucinated Humans as the Hidden Context for Labeling 3D Scenes · CVPR 2013
Robotics › Robot manipulation
grasping
0.662012
Learning hardware agnostic grasps for a universal jamming gripper · ICRA 2012
Efficient grasping from RGBD images: Learning using a new rectangle representation · ICRA 2011
Towards Holistic Scene Understanding: Feedback Enabled Cascaded Classification Models · NIPS 2010
Computer vision › Video understanding and tracking
action recognition
0.522018
Watch-n-Patch: Unsupervised Learning of Actions and Relations · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Watch-n-patch: Unsupervised understanding of actions and relations · CVPR 2015
Computer vision › Video understanding and tracking › action recognition
action relation modeling
0.522018
Watch-n-Patch: Unsupervised Learning of Actions and Relations · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Watch-n-patch: Unsupervised understanding of actions and relations · CVPR 2015
Machine learning › Deep learning architectures and training
recurrent neural network
0.522016
Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016
Structural-RNN: Deep Learning on Spatio-Temporal Graphs · CVPR 2016
Computer vision › Image recognition and object detection
object detection
0.542012
Toward Holistic Scene Understanding: Feedback Enabled Cascaded Classification Models · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011
FeCCM for scene understanding: Helping the robot to learn multiple tasks · ICRA 2011
Robotics › Autonomous driving
maneuver anticipation
0.522016
Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016
Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models · ICCV 2015
Computer vision › Video understanding and tracking › action segmentation
unsupervised action segmentation
0.522016
Watch-Bot: Unsupervised learning for reminding humans of forgotten actions · ICRA 2016
Watch-n-patch: Unsupervised understanding of actions and relations · CVPR 2015
Computer vision › 3D vision
depth estimation
0.482012
3-D Depth Reconstruction from a Single Still Image · Int. J. Comput. Vis. 2008
Make3D: Depth Perception from a Single Still Image · AAAI 2008
3-D Reconstruction from Sparse Views using Monocular Vision · ICCV 2007
Robotics › Robot manipulation
object rearrangement
0.432012
Learning to place new objects · ICRA 2012
Learning to place objects: Organizing a room · ICRA 2012
Learning Object Arrangements in 3D Scenes using Human Context · ICML 2012
Robotics › Motion planning and robot control
robot learning
0.422016
Watch-Bot: Unsupervised learning for reminding humans of forgotten actions · ICRA 2016
Learning to place new objects · ICRA 2012
Computer vision › Segmentation and scene understanding
scene understanding
0.432012
Toward Holistic Scene Understanding: Feedback Enabled Cascaded Classification Models · IEEE Trans. Pattern Anal. Mach. Intell. 2012
$\theta$-MRF: Capturing Spatial and Semantic Structure in the Parameters for Scene Understanding · NIPS 2011
FeCCM for scene understanding: Helping the robot to learn multiple tasks · ICRA 2011
Computer vision › Segmentation and scene understanding › multimodal segmentation
RGB-D segmentation
0.422015
3D Reasoning from Blocks to Stability · IEEE Trans. Pattern Anal. Mach. Intell. 2015
3D-Based Reasoning with Blocks, Support, and Stability · CVPR 2013
Computer vision › Segmentation and scene understanding
semantic segmentation
0.422016
Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.462012
3-D Depth Reconstruction from a Single Still Image · Int. J. Comput. Vis. 2008
Make3D: Depth Perception from a Single Still Image · AAAI 2008
Learning 3-D Scene Structure from a Single Still Image · ICCV 2007
Robotics › Robot manipulation › service robot
assistive robotics
0.322018
Anticipating Human Activities Using Object Affordances for Reactive Robotic Response · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Watch-n-Patch: Unsupervised Learning of Actions and Relations · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Video understanding and tracking › activity recognition
activity detection
0.212016
Anticipating Human Activities Using Object Affordances for Reactive Robotic Response · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Machine learning › Transfer learning and domain adaptation › domain alignment
cross-domain feature alignment
0.212016
Learning Transferrable Representations for Unsupervised Domain Adaptation · NIPS 2016
Computer vision › Video understanding and tracking › action anticipation
driver action prediction
0.212016
Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM
0.212016
Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016
Machine learning › Graph learning › graph neural network
spatio-temporal graph
0.212016
Structural-RNN: Deep Learning on Spatio-Temporal Graphs · CVPR 2016
Machine learning › Graph learning › graph neural network › dynamic graph neural network
spatio-temporal graph neural network
0.212016
Structural-RNN: Deep Learning on Spatio-Temporal Graphs · CVPR 2016
Machine learning › Representation and self-supervised learning › transferable representation
transferable representation learning
0.212016
Learning Transferrable Representations for Unsupervised Domain Adaptation · NIPS 2016
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.212016
Learning Transferrable Representations for Unsupervised Domain Adaptation · NIPS 2016
Human-robot interaction
assistive robotics
0.212016
Watch-Bot: Unsupervised learning for reminding humans of forgotten actions · ICRA 2016
Computer vision › Video understanding and tracking
action segmentation
0.212015
Watch-n-patch: Unsupervised understanding of actions and relations · CVPR 2015
Robotics › Autonomous driving
driver behavior modeling
0.212015
Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models · ICCV 2015
Natural language and speech › Information extraction and text analysis › lexical resources › lexical resource construction
lexicon induction
0.212015
Environment-Driven Lexicon Induction for High-Level Instructions · ACL (1) 2015
Computer vision › Video understanding and tracking › multimodal video understanding › audio-visual video understanding › audio-visual video parsing
video parsing
0.212015
Unsupervised Semantic Parsing of Video Collections · ICCV 2015
Computer vision › 3D vision
3d reconstruction
0.232015
3-D Reconstruction from Sparse Views using Monocular Vision · ICCV 2007
Learning 3-D Scene Structure from a Single Still Image · ICCV 2007
3D Reasoning from Blocks to Stability · IEEE Trans. Pattern Anal. Mach. Intell. 2015

Methods — techniques the papers use, named apart from their topics

supervised learning · 0.8markov random field · 0.5unsupervised learning · 0.5action-object co-occurrence modeling · 0.5probabilistic topic model · 0.3co-occurrence and temporal relation modeling · 0.3sequence-to-sequence prediction · 0.2sensory fusion · 0.2recurrent neural network · 0.2graph neural network · 0.2visual and language cues · 0.2joint generative model · 0.2probabilistic modeling · 0.2plane parameter estimation · 0.1machine learning · 0.1distribution modeling · 0.1stereo vision · 0.1monocular cues · 0.1
YearPublicationVenuePosition
2024 Noise level estimation using locality preserving natural image statistics
Gitam Shikkenawis, Suman K. Mitra, Ashutosh Saxena
Pattern Recognit.3
2023 Generation of 8 × 8 S-boxes using 4 × 4 optimal S-boxes
Vikas Tiwari, Appala Naidu Tentu, Ashutosh Saxena
Int. J. Inf. Comput. Secur.4
2018 Watch-n-Patch: Unsupervised Learning of Actions and Relations
abstract
There is a large variation in the activities that humans perform in their everyday lives. We consider modeling these composite human activities which comprises multiple basic level actions in a completely unsupervised setting. Our model learns high-level co-occurrence and temporal relations between the actions. We consider the video as a sequence of short-term action clips, which contains human-words and object-words. An activity is about a set of action-topics and object-topics indicating which actions are present and which objects are interacting with. We then propose a new probabilistic model relating the words and the topics. It allows us to model long-range action relations that commonly exist in the composite activities, which is challenging in previous works. We apply our model to the unsupervised action segmentation and clustering, and to a novel application that detects forgotten actions, which we call action patching. For evaluation, we contribute a new challenging RGB-D activity video dataset recorded by the new Kinect v2, which contains several human daily activities as compositions of multiple actions interacting with different objects. Moreover, we develop a robotic system that watches and reminds people using our action patching algorithm. Our robotic setup can be easily deployed on any assistive robots.
Chenxia Wu, Jiemi Zhang, Ozan Sener, Bart Selman, Silvio Savarese, Ashutosh Saxena
IEEE Trans. Pattern Anal. Mach. Intell.6
2017 Deep multimodal embedding: Manipulating novel objects with point-clouds, language and trajectories
abstract
A robot operating in a real-world environment needs to perform reasoning over a variety of sensor modalities such as vision, language and motion trajectories. However, it is extremely challenging to manually design features relating such disparate modalities. In this work, we introduce an algorithm that learns to embed point-cloud, natural language, and manipulation trajectory data into a shared embedding space with a deep neural network. To learn semantically meaningful spaces throughout our network, we use a loss-based margin to bring embeddings of relevant pairs closer together while driving less-relevant cases from different modalities further apart. We use this both to pre-train its lower layers and fine-tune our final embedding space, leading to a more robust representation. We test our algorithm on the task of manipulating novel objects and appliances based on prior experience with other objects. On a large dataset, we achieve significant improvements in both accuracy and inference time over the previous state of the art. We also perform end-to-end experiments on a PR2 robot utilizing our learned embedding space.
Jaeyong Sung, Ian Lenz, Ashutosh Saxena
ICRA3
2017 Learning to represent haptic feedback for partially-observable tasks
abstract
The sense of touch, being the earliest sensory system to develop in a human body [1], plays a critical part of our daily interaction with the environment. In order to successfully complete a task, many manipulation interactions require incorporating haptic feedback. However, manually designing a feedback mechanism can be extremely challenging. In this work, we consider manipulation tasks that need to incorporate tactile sensor feedback in order to modify a provided nominal plan. To incorporate partial observation, we present a new framework that models the task as a partially observable Markov decision process (POMDP) and learns an appropriate representation of haptic feedback which can serve as the state for a POMDP model. The model, that is parametrized by deep recurrent neural networks, utilizes variational Bayes methods to optimize the approximate posterior. Finally, we build on deep Q-learning to be able to select the optimal action in each state without access to a simulator. We test our model on a PR2 robot for multiple tasks of turning a knob until it clicks.
Jaeyong Sung, John Kenneth Salisbury Jr., Ashutosh Saxena
ICRA3
2016 Structural-RNN: Deep Learning on Spatio-Temporal Graphs
abstract
Deep Recurrent Neural Network architectures, though remarkably capable at modeling sequences, lack an intuitive high-level spatio-temporal structure. That is while many problems in computer vision inherently have an underlying high-level structure and can benefit from it. Spatiotemporal graphs are a popular tool for imposing such high-level intuitions in the formulation of real world problems. In this paper, we propose an approach for combining the power of high-level spatio-temporal graphs and sequence learning success of Recurrent Neural Networks (RNNs). We develop a scalable method for casting an arbitrary spatio-temporal graph as a rich RNN mixture that is feedforward, fully differentiable, and jointly trainable. The proposed method is generic and principled as it can be used for transforming any spatio-temporal graph through employing a certain set of well defined steps. The evaluations of the proposed approach on a diverse set of problems, ranging from modeling human motion to object interactions, shows improvement over the state-of-the-art with a large margin. We expect this method to empower new approaches to problem formulation through high-level spatio-temporal graphs and Recurrent Neural Networks.
Ashesh Jain, Amir Zamir, Silvio Savarese, Ashutosh Saxena
CVPR4
2016 Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture
abstract
Anticipating the future actions of a human is a widely studied problem in robotics that requires spatio-temporal reasoning. In this work we propose a deep learning approach for anticipation in sensory-rich robotics applications. We introduce a sensory-fusion architecture which jointly learns to anticipate and fuse information from multiple sensory streams. Our architecture consists of Recurrent Neural Networks (RNNs) that use Long Short-Term Memory (LSTM) units to capture long temporal dependencies. We train our architecture in a sequence-to-sequence prediction manner, and it explicitly learns to predict the future given only a partial temporal context. We further introduce a novel loss layer for anticipation which prevents over-fitting and encourages early anticipation. We use our architecture to anticipate driving maneuvers several seconds before they happen on a natural driving data set of 1180 miles. The context for maneuver anticipation comes from multiple sensors installed on the vehicle. Our approach shows significant improvement over the state-of-the-art in maneuver anticipation by increasing the precision from 77.4% to 90.5% and recall from 71.2% to 87.4%.
Ashesh Jain, Avi Singh, Hema Swetha Koppula, Shane Soh, Ashutosh Saxena
ICRA5
2016 Watch-Bot: Unsupervised learning for reminding humans of forgotten actions
abstract
We present a robotic system that watches a human using a Kinect v2 RGB-D sensor, detects what he forgot to do while performing an activity, and if necessary reminds the person using a laser pointer to point out the related object. Our simple setup can be easily deployed on any assistive robot. Our approach is based on a learning algorithm trained in a purely unsupervised setting, which does not require any human annotations. This makes our approach scalable and applicable to variant scenarios. Our model learns the action/object co-occurrence and action temporal relations in the activity, and uses the learned rich relationships to infer the forgotten action and the related object. We show that our approach not only improves the unsupervised action segmentation and action cluster assignment performance, but also effectively detects the forgotten actions on a challenging human activity RGB-D video dataset. In robotic experiments, we show that our robot is able to remind people of forgotten actions successfully.
Chenxia Wu, Jiemi Zhang, Bart Selman, Silvio Savarese, Ashutosh Saxena
ICRA5
2016 Learning Transferrable Representations for Unsupervised Domain Adaptation
abstract
Supervised learning with large scale labelled datasets and deep layered models has caused a paradigm shift in diverse areas in learning and recognition. However, this approach still suffers from generalization issues under the presence of a domain shift between the training and the test data distribution. Since unsupervised domain adaptation algorithms directly address this domain shift problem between a labelled source dataset and an unlabelled target dataset, recent papers have shown promising results by fine-tuning the networks with domain adaptation loss functions which try to align the mismatch between the training and testing data distributions. Nevertheless, these recent deep learning based domain adaptation approaches still suffer from issues such as high sensitivity to the gradient reversal hyperparameters and overfitting during the fine-tuning stage. In this paper, we propose a unified deep learning framework where the representation, cross domain transformation, and target label inference are all jointly optimized in an end-to-end fashion for unsupervised domain adaptation. Our experiments show that the proposed method significantly outperforms state-of-the-art algorithms in both object recognition and digit classification experiments by a large margin. We will make our learned models as well as the source code available immediately upon acceptance.
Ozan Sener, Hyun Oh Song, Ashutosh Saxena, Silvio Savarese
NIPS3
2016 MDPs with Unawareness in Robotics
Nan Rong, Joseph Y. Halpern, Ashutosh Saxena
UAI3
2016 Modeling 3D Environments through Hidden Human Context
abstract
The idea of modeling object-object relations has been widely leveraged in many scene understanding applications. However, as the objects are designed by humans and for human usage, when we reason about a human environment, we reason about it through an interplay between the environment, objects and humans. In this paper, we model environments not only through objects, but also through latent human poses and human-object interactions. In order to handle the large number of latent human poses and a large variety of their interactions with objects, we present Infinite Latent Conditional Random Field (ILCRF) that models a scene as a mixture of CRFs generated from Dirichlet processes. In each CRF, we model objects and object-object relations as existing nodes and edges, and hidden human poses and human-object relations as latent nodes and edges. ILCRF generatively models the distribution of different CRF structures over these latent nodes and edges. We apply the model to the challenging applications of 3D scene labeling and robotic scene arrangement. In extensive experiments, we show that our model significantly outperforms the state-of-the-art results in both applications. We further use our algorithm on a robot for arranging objects in a new scene using the two applications aforementioned.
Hema Swetha Koppula, Ashutosh Saxena
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Anticipating Human Activities Using Object Affordances for Reactive Robotic Response
abstract
An important aspect of human perception is anticipation, which we use extensively in our day-to-day activities when interacting with other humans as well as with our surroundings. Anticipating which activities will a human do next (and how) can enable an assistive robot to plan ahead for reactive responses. Furthermore, anticipation can even improve the detection accuracy of past activities. The challenge, however, is two-fold: We need to capture the rich context for modeling the activities and object affordances, and we need to anticipate the distribution over a large space of future human activities. In this work, we represent each possible future using an anticipatory temporal conditional random field (ATCRF) that models the rich spatial-temporal relations through object affordances. We then consider each ATCRF as a particle and represent the distribution over the potential futures using a set of particles. In extensive evaluation on CAD-120 human activity RGB-D dataset, we first show that anticipation improves the state-of-the-art detection results. We then show that for new subjects (not seen in the training set), we obtain an activity anticipation accuracy (defined as whether one of top three predictions actually happened) of 84.1, 74.4 and 62.2 percent for an anticipation time of 1, 3 and 10 seconds respectively. Finally, we also show a robot using our algorithm for performing a few reactive responses.
Hema Swetha Koppula, Ashutosh Saxena
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Environment-Driven Lexicon Induction for High-Level Instructions
abstract
Dipendra Kumar Misra, Kejia Tao, Percy Liang, Ashutosh Saxena. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Dipendra Misra, Kejia Tao, Percy Liang, Ashutosh Saxena
ACL (1)4
2015 Watch-n-patch: Unsupervised understanding of actions and relations
abstract
We focus on modeling human activities comprising multiple actions in a completely unsupervised setting. Our model learns the high-level action co-occurrence and temporal relations between the actions in the activity video. We consider the video as a sequence of short-term action clips, called action-words, and an activity is about a set of action-topics indicating which actions are present in the video. Then we propose a new probabilistic model relating the action-words and the action-topics. It allows us to model long-range action relations that commonly exist in the complex activity, which is challenging to capture in the previous works. We apply our model to unsupervised action segmentation and recognition, and also to a novel application that detects forgotten actions, which we call action patching. For evaluation, we also contribute a new challenging RGB-D activity video dataset recorded by the new Kinect v2, which contains several human daily activities as compositions of multiple actions interacted with different objects. The extensive experiments show the effectiveness of our model.
Chenxia Wu, Jiemi Zhang, Silvio Savarese, Ashutosh Saxena
CVPR4
2015 Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models
abstract
Advanced Driver Assistance Systems (ADAS) have made driving safer over the last decade. They prepare vehicles for unsafe road conditions and alert drivers if they perform a dangerous maneuver. However, many accidents are unavoidable because by the time drivers are alerted, it is already too late. Anticipating maneuvers beforehand can alert drivers before they perform the maneuver and also give ADAS more time to avoid or prepare for the danger. In this work we anticipate driving maneuvers a few seconds before they occur. For this purpose we equip a car with cameras and a computing device to capture the driving context from both inside and outside of the car. We propose an Autoregressive Input-Output HMM to model the contextual information alongwith the maneuvers. We evaluate our approach on a diverse data set with 1180 miles of natural freeway and city driving and show that we can anticipate maneuvers 3.5 seconds before they occur with over 80% F1-score in real-time.
Ashesh Jain, Hema Swetha Koppula, Bharad Raghavan, Shane Soh, Ashutosh Saxena
ICCV5
2015 Unsupervised Semantic Parsing of Video Collections
abstract
Human communication typically has an underlying structure. This is reflected in the fact that in many user generated videos, a starting point, ending, and certain objective steps between these two can be identified. In this paper, we propose a method for parsing a video into such semantic steps in an unsupervised way. The proposed method is capable of providing a semantic "storyline" of the video composed of its objective steps. We accomplish this utilizing both visual and language cues in a joint generative model. The proposed method can also provide a textual description for each of identified semantic steps and video segments. We evaluate this method on a large number of complex YouTube videos and show results of unprecedented quality for this new and impactful problem.
Ozan Sener, Amir Zamir, Silvio Savarese, Ashutosh Saxena
ICCV4
2015 PlanIt: A crowdsourcing approach for learning to plan paths from large scale preference feedback
abstract
We consider the problem of learning user preferences over robot trajectories for environments rich in objects and humans. This is challenging because the criterion defining a good trajectory varies with users, tasks and interactions in the environment. We represent trajectory preferences using a cost function that the robot learns and uses it to generate good trajectories in new environments. We design a crowdsourcing system - PlanIt, where non-expert users label segments of the robot's trajectory. PlanIt allows us to collect a large amount of user feedback, and using the weak and noisy labels from PlanIt we learn the parameters of our model. We test our approach on 122 different environments for robotic navigation and manipulation tasks. Our extensive experiments show that the learned cost function generates preferred trajectories in human environments. Our crowdsourcing system is publicly available for the visualization of the learned costs and for providing preference feedback: http://planit.cs.cornell.edu
Ashesh Jain, Debarghya Das, Jayesh K. Gupta, Ashutosh Saxena
ICRA4
2015 Robobarista: Object Part Based Transfer of Manipulation Trajectories from Crowd-Sourcing in 3D Pointclouds
Jaeyong Sung, Seok Hyun Jin, Ashutosh Saxena
ISRR (2)3
2015 3D Reasoning from Blocks to Stability
abstract
Objects occupy physical space and obey physical laws. To truly understand a scene, we must reason about the space that objects in it occupy, and how each objects is supported stably by each other. In other words, we seek to understand which objects would, if moved, cause other objects to fall. This 3D volumetric reasoning is important for many scene understanding tasks, ranging from segmentation of objects to perception of a rich 3D, physically well-founded, interpretations of the scene. In this paper, we propose a new algorithm to parse a single RGB-D image with 3D block units while jointly reasoning about the segments, volumes, supporting relationships, and object stability. Our algorithm is based on the intuition that a good 3D representation of the scene is one that fits the depth data well, and is a stable, self-supporting arrangement of objects (i.e., one that does not topple). We design an energy function for representing the quality of the block representation based on these properties. Our algorithm fits 3D blocks to the depth values corresponding to image segments, and iteratively optimizes the energy function. Our proposed algorithm is the first to consider stability of objects in complex arrangements for reasoning about the underlying structure of the scene. Experimental results show that our stability-reasoning framework improves RGB-D segmentation and scene volumetric representation.
Zhaoyin Jia, Andrew C. Gallagher, Ashutosh Saxena, Tsuhan Chen
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Physically Grounded Spatio-temporal Object Affordances
Hema Swetha Koppula, Ashutosh Saxena
ECCV (3)2
2014 Learning haptic representation for manipulating deformable food objects
abstract
Manipulation of complex deformable semi-solids such as food objects is an important skill for personal robots to have. In this work, our goal is to model and learn the physical properties of such objects. We design actions involving use of tools such as forks and knives that obtain haptic data containing information about the physical properties of the object. We then design appropriate features and use supervised learning to map these features to certain physical properties (hardness, plasticity, elasticity, tensile strength, brittleness, adhesiveness). Additionally, we present a method to compactly represent the robot's beliefs about the object's properties using a generative model, which we use to plan appropriate manipulation actions. We extensively evaluate our approach on a dataset including haptic data from 12 categories of food (including categories not seen before by the robot) obtained in 941 experiments. Our robot prepared a salad during 60 sequential robotic experiments where it made a mistake in only 4 instances.
Mevlana Gemici, Ashutosh Saxena
IROS2
2014 Synthesizing manipulation sequences for under-specified tasks using unrolled Markov Random Fields
abstract
Many tasks in human environments require performing a sequence of navigation and manipulation steps involving objects. In unstructured human environments, the location and configuration of the objects involved often change in unpredictable ways. This requires a high-level planning strategy that is robust and flexible in an uncertain environment. We propose a novel dynamic planning strategy, which can be trained from a set of example sequences. High level tasks are expressed as a sequence of primitive actions or controllers (with appropriate parameters). Our score function, based on Markov Random Field (MRF), captures the relations between environment, controllers, and their arguments. By expressing the environment using sets of attributes, the approach generalizes well to unseen scenarios. We train the parameters of our MRF using a maximum margin learning method. We provide a detailed empirical validation of our overall framework demonstrating successful plan strategies for a variety of tasks.
Jaeyong Sung, Bart Selman, Ashutosh Saxena
IROS3
2013 3D-Based Reasoning with Blocks, Support, and Stability
abstract
3D volumetric reasoning is important for truly understanding a scene. Humans are able to both segment each object in an image, and perceive a rich 3D interpretation of the scene, e.g., the space an object occupies, which objects support other objects, and which objects would, if moved, cause other objects to fall. We propose a new approach for parsing RGB-D images using 3D block units for volumetric reasoning. The algorithm fits image segments with 3D blocks, and iteratively evaluates the scene based on block interaction properties. We produce a 3D representation of the scene based on jointly optimizing over segmentations, block fitting, supporting relations, and object stability. Our algorithm incorporates the intuition that a good 3D representation of the scene is the one that fits the data well, and is a stable, self-supporting (i.e., one that does not topple) arrangement of objects. We experiment on several datasets including controlled and real indoor scenarios. Results show that our stability-reasoning framework improves RGB-D segmentation and scene volumetric representation.
Zhaoyin Jia, Andrew C. Gallagher, Ashutosh Saxena, Tsuhan Chen
CVPR3
2013 Hallucinated Humans as the Hidden Context for Labeling 3D Scenes
abstract
For scene understanding, one popular approach has been to model the object-object relationships. In this paper, we hypothesize that such relationships are only an artifact of certain hidden factors, such as humans. For example, the objects, monitor and keyboard, are strongly spatially correlated only because a human types on the keyboard while watching the monitor. Our goal is to learn this hidden human context (i.e., the human-object relationships), and also use it as a cue for labeling the scenes. We present Infinite Factored Topic Model (IFTM), where we consider a scene as being generated from two types of topics: human configurations and human-object relationships. This enables our algorithm to hallucinate the possible configurations of the humans in the scene parsimoniously. Given only a dataset of scenes containing objects but not humans, we show that our algorithm can recover the human object relationships. We then test our algorithm on the task of attribute and object labeling in 3D scenes and show consistent improvements over the state-of-the-art.
Hema Swetha Koppula, Ashutosh Saxena
CVPR3
2013 Learning Spatio-Temporal Structure from RGB-D Videos for Human Activity Detection and Anticipation
abstract
We consider the problem of detecting past activities as well as anticipating which activity will happen in the future and how. We start by modeling the rich spatio-temporal relations between human poses and objects (called affordances) using a conditional random field (CRF). However, because of the ambiguity in the temporal segmentation of the sub-activities that constitute an activity, in the past as well as in the future, multiple graph structures are possible. In this paper, we reason about these alternate possibilities by reasoning over multiple possible graph structures. We obtain them by approximating the graph with only additive features, which lends to efficient dynamic programming. Starting with this proposal graph structure, we then design moves to obtain several other likely graph structures. We then show that our approach improves the state-of-the-art significantly for detecting past activities as well as for anticipating future activities, on a dataset of 120 activity videos collected from four subjects.
Hema Swetha Koppula, Ashutosh Saxena
ICML (3)2
2013 Discovering Different Types of Topics: Factored Topic Models
Ashutosh Saxena
IJCAI2
2013 Anticipating human activities for reactive robotic response
abstract
An important aspect of human perception is anticipation, which we use extensively in our day-to-day activities when interacting with other humans as well as with our surroundings. Anticipating which activities will a human do next (and how to do them) can enable an assistive robot to plan ahead for reactive responses in the human environments. In this work, our goal is to enable robots to predict the future activities as well as the details of how a human is going to perform them in short-term (e.g., 1-10 seconds). For example, if a robot has seen a person move his hand to a coffee mug, it is possible he would move the coffee mug to a few potential places such as his mouth, to a kitchen sink or just move it to a different location on the table. If a robot can anticipate this, then it would rather not start pouring milk into the coffee when the person is moving his hand towards the mug, thus avoiding a spill. We represent each possible future using an anticipatory temporal conditional random field (ATCRF) that models the rich spatial-temporal relations through object affordances. We then consider each ATCRF as a particle and represent the distribution over the potential futures using a set of particles. We evaluate our anticipation approach extensively on CAD-120 human activity dataset, which contains 120 RGB-D videos of daily human activities, such as microwaving food, taking medicine, etc. For robotic evaluation, we measure how many times the robot anticipates and performs the correct reactive response. The accompanying video shows a PR2 robot performing assistive tasks based on the anticipations generated by our proposed method.
Hema Swetha Koppula, Ashutosh Saxena
IROS2
2013 Tangled: Learning to untangle ropes with RGB-D perception
abstract
In this paper, we address the problem of manipulating deformable objects such as ropes. Starting with an RGB-D view of a tangled rope, our goal is to infer its knot structure and then choose appropriate manipulation actions that result in the rope getting untangled. We design appropriate features and present an inference algorithm based on particle filters to infer the rope's structure. Our learning algorithm is based on max-margin learning. We then choose an appropriate manipulation action based on the current knot structure and other properties such as slack in the rope. We then repeatedly perform perception and manipulation until the rope is untangled. We evaluate our algorithm extensively on a dataset having five different types of ropes and 10 different types of knots. We then perform robotic experiments, in which our bimanual manipulator (PR2) untangles ropes successfully 76.9% of the time.
Wen Hao Lui, Ashutosh Saxena
IROS2
2013 Beyond Geometric Path Planning: Learning Context-Driven Trajectory Preferences via Sub-optimal Feedback
Ashesh Jain, Shikhar Sharma 0001, Ashutosh Saxena
ISRR3
2013 Learning Trajectory Preferences for Manipulators via Iterative Improvement
abstract
We consider the problem of learning good trajectories for manipulation tasks. This is challenging because the criterion defining a good trajectory varies with users, tasks and environments. In this paper, we propose a co-active online learning framework for teaching robots the preferences of its users for object manipulation tasks. The key novelty of our approach lies in the type of feedback expected from the user: the human user does not need to demonstrate optimal trajectories as training data, but merely needs to iteratively provide trajectories that slightly improve over the trajectory currently proposed by the system. We argue that this co-active preference feedback can be more easily elicited from the user than demonstrations of optimal trajectories, which are often challenging and non-intuitive to provide on high degrees of freedom manipulators. Nevertheless, theoretical regret bounds of our algorithm match the asymptotic rates of optimal trajectory algorithms. We also formulate a score function to capture the contextual information and demonstrate the generalizability of our algorithm on a variety of household tasks, for whom, the preferences were not only influenced by the object being manipulated but also by the surrounding environment.
Ashesh Jain, Brian Wojcik, Thorsten Joachims, Ashutosh Saxena
NIPS4
2012 Learning the right model: Efficient max-margin learning in Laplacian CRFs
abstract
An important modeling decision made while designing Conditional Random Fields (CRFs) is the choice of the potential functions over the cliques of variables. Laplacian potentials are useful because they are robust potentials and match image statistics better than Gaussians. Moreover, energies with Laplacian terms remain convex, which simplifies inference. This makes Laplacian potentials an ideal modeling choice for some applications. In this paper, we study max-margin parameter learning in CRFs with Laplacian potentials (LCRFs). We first show that structured hinge-loss [35] is non-convex for LCRFs and thus techniques used by previous works are not applicable. We then present the first approximate max-margin algorithm for LCRFs. Finally, we make our learning algorithm scalable in the number of training images by using dual-decomposition techniques. Our experiments on single-image depth estimation show that even with simple features, our approach achieves comparable to state-of-art results.
Dhruv Batra, Ashutosh Saxena
CVPR2
2012 Co-evolutionary predictors for kinematic pose inference from RGBD images
abstract
Markerless pose inference of arbitrary subjects is a primary problem for a variety of applications, including robot vision and teaching by demonstration. Unsupervised kinematic pose inference is an ideal method for these applications as it provides a robust, training-free approach with minimal reliance on prior information. However, these methods have been considered intractable for complex models. This paper presents a general framework for inferring poses from a single depth image given an arbitrary kinematic structure without prior training. A co-evolutionary algorithm, consisting of pose and predictor populations, is applied to overcome the traditional limitations in kinematic pose inference. Evaluated on test sets of 256 synthetic and 52 real images, our algorithm shows consistent pose inference for 34 and 78 degree of freedom models with point clouds containing over 40,000 points, even in cases of significant self-occlusion. Compared to various baselines, the co-evolutionary algorithm provides at least a 3.5-fold increase in pose accuracy and a two-fold reduction in computational effort for articulated models.
Daniel Le Ly, Ashutosh Saxena, Hod Lipson
GECCO2
2012 Learning Object Arrangements in 3D Scenes using Human Context
Marcus Lim, Ashutosh Saxena
ICML3
2012 Learning to place objects: Organizing a room
abstract
In this video, we consider the task of a personal robot organizing a room by placing objects stably as well as in semantically preferred locations. While this includes many sub-tasks such as grasping an object, moving to a placing position, localizing itself and placing the object in a proper location and orientation, it is the last problem - how and where to place the objects - that is our focus in this work and has not been widely studied yet. We formulate the placing task as a learning problem. By computing appearance and shape features from the input (point-clouds) that can capture the stability and semantics, our algorithm can identify good placements for multiple objects. In this video, we put together the placing algorithm with other sub-tasks to enable a robot organize a room in several scenarios, such as loading a bookshelf, a fridge, a waste bin and blackboard with various objects.
Gaurab Basu, Ashutosh Saxena
ICRA3
2012 Learning hardware agnostic grasps for a universal jamming gripper
abstract
Grasping has been studied from various perspectives including planning, control, and learning. In this paper, we take a learning approach to predict successful grasps for a universal jamming gripper. A jamming gripper is comprised of a flexible membrane filled with granular material, and it can quickly harden or soften to grip objects of varying shape by modulating the air pressure within the membrane. Although this gripper is easy to control, developing a physical model of its gripping mechanism is difficult because it undergoes significant deformation during use. Thus, many grasping approaches based on physical models (such as based on form- and force-closure) would be challenging to apply to a jamming gripper. Here we instead use a supervised learning algorithm and design both visual and shape features for capturing the properties of good grasps. We show that given target object data from an RGBD sensor, our algorithm can predict successful grasps for the jamming gripper without requiring a physical model. It can therefore be applied to both a parallel plate gripper and a jamming gripper without modification. We demonstrate that our learning algorithm enables both grippers to pick up a wide variety of objects, including objects from outside the training set. Through robotic experiments we are then able to define the type of objects each gripper is best suited for handling.
John R. Amend, Hod Lipson, Ashutosh Saxena
ICRA4
2012 Learning to place new objects
abstract
The ability to place objects in an environment is an important skill for a personal robot. An object should not only be placed stably, but should also be placed in its preferred location/orientation. For instance, it is preferred that a plate be inserted vertically into the slot of a dish-rack as compared to being placed horizontally in it. Unstructured environments such as homes have a large variety of object types as well as of placing areas. Therefore our algorithms should be able to handle placing new object types and new placing areas. These reasons make placing a challenging manipulation task. In this work, we propose using a supervised learning approach for finding good placements given point-clouds of the object and the placing area. Our method combines the features that capture support, stability and preferred configurations, and uses a shared sparsity structure in its the parameters. Even when neither the object nor the placing area is seen previously in the training set, our learning algorithm predicts good placements. In robotic experiments, our method enables the robot to stably place known objects with a 98% success rate and 98% when also considering semantically preferred orientations. In the case of placing a new object into a new placing area, the success rate is 82% and 72%.
Changxi Zheng, Marcus Lim, Ashutosh Saxena
ICRA4
2012 Unstructured human activity detection from RGBD images
abstract
Being able to detect and recognize human activities is essential for several applications, including personal assistive robotics. In this paper, we perform detection and recognition of unstructured human activity in unstructured environments. We use a RGBD sensor (Microsoft Kinect) as the input sensor, and compute a set of features based on human pose and motion, as well as based on image and point-cloud information. Our algorithm is based on a hierarchical maximum entropy Markov model (MEMM), which considers a person's activity as composed of a set of sub-activities. We infer the two-layered graph structure using a dynamic programming approach. We test our algorithm on detecting and recognizing twelve different activities performed by four people in different environments, such as a kitchen, a living room, an office, etc., and achieve good performance even when the person was not seen before in the training set.1
Jaeyong Sung, Colin Ponce, Bart Selman, Ashutosh Saxena
ICRA4
2012 Low-power parallel algorithms for single image based obstacle avoidance in aerial robots
abstract
For an aerial robot, perceiving and avoiding obstacles are necessary skills to function autonomously in a cluttered unknown environment. In this work, we use a single image captured from the onboard camera as input, produce obstacle classifications, and use them to select an evasive maneuver. We present a Markov Random Field based approach that models the obstacles as a function of visual features and non-local dependencies in neighboring regions of the image. We perform efficient inference using new low-power parallel neuromorphic hardware, where belief propagation updates are done using leaky integrate and fire neurons in parallel, while consuming less than 1 W of power. In outdoor robotic experiments, our algorithm was able to consistently produce clean, accurate obstacle maps which allowed our robot to avoid a wide variety of obstacles, including trees, poles and fences.
Ian Lenz, Mevlana Gemici, Ashutosh Saxena
IROS3
2012 Toward Holistic Scene Understanding: Feedback Enabled Cascaded Classification Models
abstract
Scene understanding includes many related subtasks, such as scene categorization, depth estimation, object detection, etc. Each of these subtasks is often notoriously hard, and state-of-the-art classifiers already exist for many of them. These classifiers operate on the same raw image and provide correlated outputs. It is desirable to have an algorithm that can capture such correlation without requiring any changes to the inner workings of any classifier. We propose Feedback Enabled Cascaded Classification Models (FE-CCM), that jointly optimizes all the subtasks while requiring only a "black box" interface to the original classifier for each subtask. We use a two-layer cascade of classifiers, which are repeated instantiations of the original ones, with the output of the first layer fed into the second layer as input. Our training method involves a feedback step that allows later classifiers to provide earlier classifiers information about which error modes to focus on. We show that our method significantly improves performance in all the subtasks in the domain of scene understanding, where we consider depth estimation, scene categorization, event categorization, object detection, geometric labeling, and saliency detection. Our method also improves performance in two robotic applications: an object-grasping robot and an object-finding robot.
Adarsh Kowdle, Ashutosh Saxena, Tsuhan Chen
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 A security analysis of smartphone data flow and feasible solutions for lawful interception
abstract
Smartphones providing proprietary encryption schemes, albeit offering a novel paradigm to privacy, are becoming a bone of contention for certain sovereignties. These sovereignties have raised concerns about their security agencies not having any control on the encrypted data leaving their jurisdiction and the ensuing possibility of it being misused by people with malicious intents. Such smartphones have typically two types of customers, independent users who use it to access public mail servers and corporates/enterprises whose employees use it to access corporate emails in an encrypted form. The threat issues raised by security agencies concern mainly the enterprise servers where the encrypted data leaves the jurisdiction of the respective sovereignty while on its way to the global smartphone router. In this paper, we have analyzed such email message transfer mechanisms in smartphones and proposed some feasible solutions, which, if accepted and implemented by entities involved, can lead to a possible win-win situation for both the parties, viz., the smartphone provider who does not want to lose the customers and these sovereignties who can avoid the worry of encrypted data leaving their jurisdiction.
Mithun Paul, Nitin Singh Chauhan, Ashutosh Saxena
IAS3
2011 A DRM framework towards preventing digital piracy
abstract
Digital piracy is a major challenge faced by content publishers and software vendors today. This paper presents a Digital Rights Management (DRM) framework to secure digital content and software applications. The DRM framework uses cryptographic techniques and supports protection of digital content viz., PDF, image and audio files by enforcing user rights such as view, copy, play or print as applicable. The framework is extendable to safeguard libraries of software applications on multiple operating systems. The design offers protection to various file formats with a DRM license that can be upgraded for additional rights or be renewed to get an extended validity. The DRM framework also accommodates offline use of protected content by a one-time (initial) setup and a user license stored locally. Finally, the paper analyzes the design for DRM's crucial requirements like security, flexibility, efficiency and interoperability.
Ravi Sankar Veerubhotla, Ashutosh Saxena
IAS2
2011 Autonomous MAV flight in indoor environments using single image perspective cues
abstract
We consider the problem of autonomously flying Miniature Aerial Vehicles (MAVs) in indoor environments such as home and office buildings. The primary long range sensor in these MAVs is a miniature camera. While previous approaches first try to build a 3D model in order to do planning and control, our method neither attempts to build nor requires a 3D model. Instead, our method first classifies the type of indoor environment the MAV is in, and then uses vision algorithms based on perspective cues to estimate the desired direction to fly. We test our method on two MAV platforms: a co-axial miniature helicopter and a toy quadrotor. Our experiments show that our vision algorithms are quite reliable, and they enable our MAVs to fly in a variety of corridors and staircases.
Cooper Bills, Joyce Chen, Ashutosh Saxena
ICRA3
2011 Efficient grasping from RGBD images: Learning using a new rectangle representation
abstract
Given an image and an aligned depth map of an object, our goal is to estimate the full 7-dimensional gripper configuration—its 3D location, 3D orientation and the gripper opening width. Recently, learning algorithms have been successfully applied to grasp novel objects—ones not seen by the robot before. While these approaches use low-dimensional representations such as a ‘grasping point’ or a ‘pair of points’ that are perhaps easier to learn, they only partly represent the gripper configuration and hence are sub-optimal. We propose to learn a new ‘grasping rectangle’ representation: an oriented rectangle in the image plane. It takes into account the location, the orientation as well as the gripper opening width. However, inference with such a representation is computationally expensive. In this work, we present a two step process in which the first step prunes the search space efficiently using certain features that are fast to compute. For the remaining few cases, the second step uses advanced features to accurately select a good grasp. In our extensive experiments, we show that our robot successfully uses our algorithm to pick up a variety of novel objects.
Stephen Moseson, Ashutosh Saxena
ICRA3
2011 FeCCM for scene understanding: Helping the robot to learn multiple tasks
abstract
Helping a robot to understand a scene can include many sub-tasks, such as scene categorization, object detection, geometric labeling, etc. Each sub-task is notoriously hard, and state-of-art classifiers exist for many sub-tasks. It is desirable to have an algorithm that can capture such correlation without requiring to make any changes to the inner workings of any classifier, and therefore make the perception for a robot better. We have recently proposed a generic model (Feedback Enabled Cascaded Classification Model) that enables us to easily take state-of-art classifiers as black-boxes and improve performance. In this video, we show that we can use our FeCCM model to quickly combine existing classifiers for various sub-tasks, and build a shoe finder robot in a day. The video shows our robot using FeCCM to find a shoe on request.
T. P. Wong, Norris Xu, Ashutosh Saxena
ICRA4
2011 Robotic Object Detection: Learning to Improve the Classifiers Using Sparse Graphs for Path Planning
abstract
Object detection is a basic skill for a robot to perfor-m tasks in human environments. In order to build a good object classifier, a large training set of la-beled images is required; this is typically collect-ed and labeled (often painstakingly) by a human. This method is not scalable and therefore limits the robot’s detection performance. We propose an algorithm for a robot to collect more data in the environment during its training phase so that in the future it could detect objects more reli-ably. The first step is to plan a path for collecting additional training images, which is hard because a previously visited location affects the decision for the future locations. One key component of our work is path planning by building a sparse graph that captures these dependencies. The other key component is our learning algorithm that weighs the errors made in robot’s data collection process while updating the classifier. In our experiments, we show that our algorithms enable the robot to im-prove its object classifiers significantly.
Zhaoyin Jia, Ashutosh Saxena, Tsuhan Chen
IJCAI2
2011 Semantic Labeling of 3D Point Clouds for Indoor Scenes
abstract
Inexpensive RGB-D cameras that give an RGB image together with depth data have become widely available. In this paper, we use this data to build 3D point clouds of full indoor scenes such as an office and address the task of semantic labeling of these 3D point clouds. We propose a graphical model that captures various features and contextual relations, including the local visual appearance and shape cues, object co-occurence relationships and geometric relationships. With a large number of object classes and relations, the model’s parsimony becomes important and we address that by using multiple types of edge potentials. The model admits efficient approximate inference, and we train it using a maximum-margin learning approach. In our experiments over a total of 52 3D scenes of homes and offices (composed from about 550 views, having 2495 segments labeled with 27 object classes), we get a performance of 84.06% in labeling 17 object classes for offices, and 73.38% in labeling 17 object classes for home scenes. Finally, we applied these algorithms successfully on a mobile robot for the task of finding objects in large cluttered rooms.
Hema Swetha Koppula, Abhishek Anand, Thorsten Joachims, Ashutosh Saxena
NIPS4
2011 $\theta$-MRF: Capturing Spatial and Semantic Structure in the Parameters for Scene Understanding
abstract
For most scene understanding tasks (such as object detection or depth estimation), the classifiers need to consider contextual information in addition to the local features. We can capture such contextual information by taking as input the features/attributes from all the regions in the image. However, this contextual dependence also varies with the spatial location of the region of interest, and we therefore need a different set of parameters for each spatial location. This results in a very large number of parameters. In this work, we model the independence properties between the parameters for each location and for each task, by defining a Markov Random Field (MRF) over the parameters. In particular, two sets of parameters are encouraged to have similar values if they are spatially close or semantically close. Our method is, in principle, complementary to other ways of capturing context such as the ones that use a graphical model over the labels instead. In extensive evaluation over two different settings, of multi-class object detection and of multiple scene understanding tasks (scene categorization, depth estimation, geometric labeling), our method beats the state-of-the-art methods in all the four tasks.
Ashutosh Saxena, Tsuhan Chen
NIPS2
2011 A neural network approach for data masking
Vishal Anjaiah Gujjary, Ashutosh Saxena
Neurocomputing2
2010 Learning to open new doors
abstract
We consider the problem of enabling a robot to autonomously open doors, including novel ones that the robot has not previously seen. Given the large variation in the appearances and locations of doors and door handles, this is a challenging perception and control problem; but this capability will significantly enlarge the range of environments that our robots can autonomously navigate through. In this paper, we focus on the case of doors with door handles. We propose an approach that, rather than trying to build a full 3d model of the door/door handle-which is challenging because of occlusion, specularity of many door handles, and the limited accuracy of our 3d sensors-instead uses computer vision to choose a manipulation strategy. Specifically, it uses an image of the door handle to identify a small number of “3d key locations,” such as the axis of rotation of the door handle, and the location of the end-point of the door-handle. These key locations then completely define a trajectory for the robot end-effector (hand) that successfully turns the door handle and opens the door. Evaluated on a large set of doors that the robot had not previously seen, it successfully opened 31 out of 34 doors. We also show that this approach of using vision to identify a small number of key locations also generalizes to a range of other tasks, including turning a thermostat knob, pulling open a drawer, and pushing elevator buttons.
Ellen Klingbeil, Ashutosh Saxena, Andrew Y. Ng
IROS2
2010 Towards Holistic Scene Understanding: Feedback Enabled Cascaded Classification Models
abstract
In many machine learning domains (such as scene understanding), several related sub-tasks (such as scene categorization, depth estimation, object detection) operate on the same raw data and provide correlated outputs. Each of these tasks is often notoriously hard, and state-of-the-art classifiers already exist for many sub-tasks. It is desirable to have an algorithm that can capture such correlation without requiring to make any changes to the inner workings of any classifier. We propose Feedback Enabled Cascaded Classification Models (FE-CCM), that maximizes the joint likelihood of the sub-tasks, while requiring only a ‘black-box’ interface to the original classifier for each sub-task. We use a two-layer cascade of classifiers, which are repeated instantiations of the original ones, with the output of the first layer fed into the second layer as input. Our training method involves a feedback step that allows later classifiers to provide earlier classifiers information about what error modes to focus on. We show that our method significantly improves performance in all the sub-tasks in two different domains: (i) scene understanding, where we consider depth estimation, scene categorization, event categorization, object detection, geometric labeling and saliency detection, and (ii) robotic grasping, where we consider grasp point detection and object classification.
Adarsh Kowdle, Ashutosh Saxena, Tsuhan Chen
NIPS3
2010 MDPs with Unawareness
Joseph Y. Halpern, Nan Rong, Ashutosh Saxena
UAI3
2009 Reactive grasping using optical proximity sensors
abstract
We propose a system for improving grasping using fingertip optical proximity sensors that allows us to perform online grasp adjustments to an initial grasp point without requiring premature object contact or regrasping strategies. We present novel optical proximity sensors that fit inside the fingertips of a Barrett Hand, and demonstrate their use alongside a probabilistic model for robustly combining sensor readings and a hierarchical reactive controller for improving grasps online. This system can be used to complement existing grasp planning algorithms, or be used in more interactive settings where a human indicates the location of objects. Finally, we perform a series of experiments using a Barrett hand equipped with our sensors to grasp a variety of common objects with mixed geometries and surface textures.
Kaijen Hsiao, Paul Nangeroni, Manfred Huber, Ashutosh Saxena, Andrew Y. Ng
ICRA4
2009 Learning 3-D object orientation from images
abstract
We propose a learning algorithm for estimating the 3-D orientation of objects. Orientation learning is a difficult problem because the space of orientations is non-Euclidean, and in some cases (such as quaternions) the representation is ambiguous, in that multiple representations exist for the same physical orientation. Learning is further complicated by the fact that most man-made objects exhibit symmetry, so that there are multiple ldquocorrectrdquo orientations. In this paper, we propose a new representation for orientations-and a class of learning and inference algorithms using this representation-that allows us to learn orientations for symmetric or asymmetric objects as a function of a single image. We extensively evaluate our algorithm for learning orientations of objects from six categories.
Ashutosh Saxena, Justin Driemeyer, Andrew Y. Ng
ICRA1
2009 Learning sound location from a single microphone
abstract
We consider the problem of estimating the incident angle of a sound, using only a single microphone. The ability to perform monaural (single-ear) localization is important to many animals; indeed, monaural cues are also the primary method by which humans decide if a sound comes from the front or back, as well as estimate its elevation. Such monaural localization is made possible by the structure of the pinna (outer ear), which modifies sound in a way that is dependent on its incident angle. In this paper, we propose a machine learning approach to monaural localization, using only a single microphone and an ldquoartificial pinnardquo (that distorts sound in a direction-dependent way). Our approach models the typical distribution of natural and artificial sounds, as well as the direction-dependent changes to sounds induced by the pinna. Our experimental results also show that the algorithm is able to fairly accurately localize a wide range of sounds, such as human speech, dog barking, waterfall, thunder, and so on. In contrast to microphone arrays, this approach also offers the potential of significantly more compact, as well as lower cost and power, devices for sounds localization.
Ashutosh Saxena, Andrew Y. Ng
ICRA1
2009 Autonomous indoor helicopter flight using a single onboard camera
abstract
We consider the problem of autonomously flying a helicopter in indoor environments. Navigation in indoor settings poses two major challenges. First, real-time perception and response is crucial because of the high presence of obstacles. Second, the limited free space in such a setting places severe restrictions on the size of the aerial vehicle, resulting in a frugal payload budget. We autonomously fly a miniature RC helicopter in small known environments using an on-board light-weight camera as the only sensor. We use an algorithm that combines data-driven image classification with optical flow techniques on the images captured by the camera to achieve real-time 3D localization and navigation. We perform successful autonomous test flights along trajectories in two different indoor settings. Our results demonstrate that our method is capable of autonomous flight even in narrow indoor spaces with sharp corners.
Sai Prashanth Soundararaj, Arvind K. Sujeeth, Ashutosh Saxena
IROS3
2009 Danger theory based SYN flood attack detection in autonomic network
abstract
In the context of autonomic environment, we present a simple yet, effective Danger Theory based method to detect TCP SYN Flooding attack. An autonomous communication network consists of self-managed (i.e. self-configuring, self-awareness, self-optimization, self-healing and self-protection, collectively denoted as self-*) entities. These self-* properties ensure functioning of the network without or very minimum human intervention. In such an environment, security of the system is very challenging as there is no dedicated authority to monitor malicious activities and each entity, the computing device, has to monitor itself. Denial of service (DoS) attack, in particular flooding attack, is one of the most frequent and devastating attacks on networks. Traditionally, the detection of flooding attacks is achieved by a network-based intrusion detection system (IDS), mainly relying on the statistical characteristics of network data with fine tuning from a human administrator by monitoring the traffic continuously. Obviously, such facility is not assumed in autonomic networks. We, therefore, propose a danger theory based approach that can detect DoS attack in an automatic manner. The proposed scheme is able to detect SYN flood attack in its early stage, thereby enabling to control the damage. To empirically validate our proposal, we conduct experiments in a simulated environment and the results are encouraging. We assert that the work will be useful in designing the security of autonomic networks.
Sanjay Rawat 0001, Ashutosh Saxena
SIN2
2009 An efficient and secure protocol for DTV broadcasts
abstract
Bilinear pairing based mutual authentication and key agreement protocol for DTV broadcast encryption is presented in this paper. The protocol facilitates key agreement with lesser communication between set-top box and smart card with forward secrecy and is resilient to replay, forgery, man-in-the-middle and insider attacks and we provide the security analysis for it. The protocol is especially attractive for conditional access system including gaming, betting, shopping and banking services and where the user' smart card have low computational power. The protocol also provides flexible password change option to the users.
Ashutosh Saxena
SIN1
2009 Application security code analysis: a step towards software assurance
abstract
The last few years have witnessed a rapid growth in cyber attacks, with daily new vulnerabilities being discovered in computer applications. Various security-related technologies, e.g., anti-virus programs, Intrusion Detection Systems (IDSs)/Intrusion Prevention Systems (IPSs), firewalls, etc., are deployed to minimise the number of attacks and incurred losses. However, such technologies are not enough to completely eliminate the attacks to some extent; they can only minimise them. Therefore, software assurance is becoming a priority and an important characteristic of the software development life cycle. Application code analysis is gaining importance, as it can help in writing safe code during the development phase by detecting bugs that may lead to vulnerabilities. As a result, tremendous research on code analysis has been carried out by industry and academia and there exist many commercial and open source tools and approaches for this purpose. These have their own pros and cons. Therefore, the main objective of this article is to explore the state-of-the-art in code analysis and a few major tools which benefit not only security professionals, but also novice Information Technology (IT) professionals. We study the tools and techniques under the basic four types of analysis (Static Source Code (SSC), Static Binary Code (SBC), Dynamic Source Code (DSC) and Dynamic Binary Code (DBC) analysis) and briefly discuss them.
Sanjay Rawat 0001, Ashutosh Saxena
Int. J. Inf. Comput. Secur.2
2009 Make3D: Learning 3D Scene Structure from a Single Still Image
abstract
We consider the problem of estimating detailed 3D structure from a single still image of an unstructured environment. Our goal is to create 3D models that are both quantitatively accurate as well as visually pleasing. For each small homogeneous patch in the image, we use a Markov Random Field (MRF) to infer a set of "plane parameters" that capture both the 3D location and 3D orientation of the patch. The MRF, trained via supervised learning, models both image depth cues as well as the relationships between different parts of the image. Other than assuming that the environment is made up of a number of small planes, our model makes no explicit assumptions about the structure of the scene; this enables the algorithm to capture much more detailed 3D structure than does prior art and also give a much richer experience in the 3D flythroughs created using image-based rendering, even for scenes with significant nonvertical structure. Using this approach, we have created qualitatively correct 3D models for 64.9 percent of 588 images downloaded from the Internet. We have also extended our model to produce large-scale 3D models from a few images.
Ashutosh Saxena, Andrew Y. Ng
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 A Fast Data Collection and Augmentation Procedure for Object Recognition
Benjamin Sapp, Ashutosh Saxena, Andrew Y. Ng
AAAI2
2008 Make3D: Depth Perception from a Single Still Image
Ashutosh Saxena, Andrew Y. Ng
AAAI1
2008 Learning Grasp Strategies with Partial Shape Information
Ashutosh Saxena, Lawson L. S. Wong, Andrew Y. Ng
AAAI1
2008 Cascaded Classification Models: Combining Models for Holistic Scene Understanding
abstract
One of the original goals of computer vision was to fully understand a natural scene. This requires solving several problems simultaneously, including object detection, labeling of meaningful regions, and 3d reconstruction. While great progress has been made in tackling each of these problems in isolation, only recently have researchers again been considering the difficult task of assembling various methods to the mutual benefit of all. We consider learning a set of such classification models in such a way that they both solve their own problem and help each other. We develop a framework known as Cascaded Classification Models (CCM), where repeated instantiations of these classifiers are coupled by their input/output variables in a cascade that improves performance at each level. Our method requires only a limited “black box” interface with the models, allowing us to use very sophisticated, state-of-the-art classifiers without having to look under the hood. We demonstrate the effectiveness of our method on a large set of natural images by combining the subtasks of scene categorization, object detection, multiclass image segmentation, and 3d scene reconstruction.
Geremy Heitz, Stephen Gould, Ashutosh Saxena, Daphne Koller
NIPS3
2008 3-D Depth Reconstruction from a Single Still Image
abstract
We consider the task of 3-d depth estimation from a single still image. We take a supervised learning approach to this problem, in which we begin by collecting a training set of monocular images (of unstructured indoor and outdoor environments which include forests, sidewalks, trees, buildings, etc.) and their corresponding ground-truth depthmaps. Then, we apply supervised learning to predict the value of the depthmap as a function of the image. Depth estimation is a challenging problem, since local features alone are insufficient to estimate depth at a point, and one needs to consider the global context of the image. Our model uses a hierarchical, multiscale Markov Random Field (MRF) that incorporates multiscale local- and global-image features, and models the depths and the relation between depths at different points in the image. We show that, even on unstructured scenes, our algorithm is frequently able to recover fairly accurate depthmaps. We further propose a model that incorporates both monocular cues and stereo (triangulation) cues, to obtain significantly more accurate depth estimates than is possible using either monocular or stereo cues alone.
Ashutosh Saxena, Sung H. Chung, Andrew Y. Ng
Int. J. Comput. Vis.1
2007 Threshold SKI Protocol for ID-based Cryptosystems
abstract
Traditional public key cryptography uses certificates to bind the users with their public keys and are considered the best alternative for key distribution, but requires to have a very involved key management process. Identity based cryptography makes the key management easier but suffers from the key escrow problem and requires secure channel to issue the private keys to the users. Key issuing protocols deal with secret key issuing (SKI) process to overcome the two problems. We present an efficient and secure key issuing protocol which enables the identity based cryptosystems to be more acceptable and applicable in the real world. In the protocol, neither key generating center nor key privacy authority can impersonate the users to obtain the private keys. Performance and security analysis are being carried out for the protocol and is shown that it is efficient and secure against replay, man-in-the-middle and insider attacks.
Ashutosh Saxena
IAS1
2007 Learning 3-D Scene Structure from a Single Still Image
abstract
We consider the problem of estimating detailed 3D structure from a single still image of an unstructured environment. Our goal is to create 3D models which are both quantitatively accurate as well as visually pleasing. For each small homogeneous patch in the image, we use a Markov random field (MRF) to infer a set of "plane parameters" that capture both the 3D location and 3D orientation of the patch. The MRF, trained via supervised learning, models both image depth cues as well as the relationships between different parts of the image. Inference in our model is tractable, and requires only solving a convex optimization problem. Other than assuming that the environment is made up of a number of small planes, our model makes no explicit assumptions about the structure of the scene; this enables the algorithm to capture much more detailed 3D structure than does prior art (such as Saxena et ah, 2005, Delage et ah, 2005, and Hoiem et el, 2005), and also give a much richer experience in the 3D flythroughs created using image-based rendering, even for scenes with significant non-vertical structure. Using this approach, we have created qualitatively correct 3D models for 64.9% of 588 images downloaded from the Internet, as compared to Hoiem et al.'s performance of 33.1%. Further, our models are quantitatively more accurate than either Saxena et al. or Hoiem et al.
Ashutosh Saxena, Andrew Y. Ng
ICCV1
2007 3-D Reconstruction from Sparse Views using Monocular Vision
abstract
We consider the task of creating a 3-d model of a large novel environment, given only a small number of images of the scene. This is a difficult problem, because if the images are taken from very different viewpoints or if they contain similar-looking structures, then most geometric reconstruction methods will have great difficulty finding good correspondences. Further, the reconstructions given by most algorithms include only points in 3-d that were observed in two or more images; a point observed only in a single image would not be reconstructed. In this paper, we show how monocular image cues can be combined with triangulation cues to build a photo-realistic model of a scene given only a few images—even ones taken from very different viewpoints or with little overlap. Our approach begins by over-segmenting each image into small patches (superpixels). It then simultaneously tries to infer the 3-d position and orientation of every superpixel in every image. This is done using a Markov Random Field (MRF) which simultaneously reasons about monocular cues and about the relations between multiple image patches, both within the same image and across different images (triangulation cues). MAP inference in our model is efficiently approximated using a series of linear programs, and our algorithm scales well to a large number of images.
Ashutosh Saxena, Andrew Y. Ng
ICCV1
2007 Depth Estimation Using Monocular and Stereo Cues
Ashutosh Saxena, Jamie Schulte, Andrew Y. Ng
IJCAI1
2007 A Vision-Based System for Grasping Novel Objects in Cluttered Environments
Ashutosh Saxena, Lawson L. S. Wong, Morgan Quigley, Andrew Y. Ng
ISRR1
2006 Robotic Grasping of Novel Objects
abstract
We consider the problem of grasping novel objects, specifically ones that are being seen for the first time through vision. We present a learning algorithm that neither requires, nor tries to build, a 3-d model of the object. Instead it predicts, directly as a function of the images, a point at which to grasp the object. Our algorithm is trained via supervised learning, using synthetic images for the training set. We demonstrate on a robotic manipulation platform that this approach successfully grasps a wide variety of objects, such as wine glasses, duct tape, markers, a translucent box, jugs, knife-cutters, cellphones, keys, screwdrivers, staplers, toothbrushes, a thick coil of wire, a strangely shaped power horn, and others, none of which were seen in the training set.
Ashutosh Saxena, Justin Driemeyer, Justin Kearns, Andrew Y. Ng
NIPS1
2006 A novel remote user authentication scheme using bilinear pairings
Manik Lal Das, Ashutosh Saxena, Ved Prakash Gulati, Deepak B. Phatak
Comput. Secur.2
2005 High speed obstacle avoidance using monocular vision and reinforcement learning
abstract
We consider the task of driving a remote control car at high speeds through unstructured outdoor environments. We present an approach in which supervised learning is first used to estimate depths from single monocular images. The learning algorithm can be trained either on real camera images labeled with ground-truth distances to the closest obstacles, or on a training set consisting of synthetic graphics images. The resulting algorithm is able to learn monocular vision cues that accurately estimate the relative depths of obstacles in a scene. Reinforcement learning/policy search is then applied within a simulator that renders synthetic scenes. This learns a control policy that selects a steering direction as a function of the vision system's output. We present results evaluating the predictive ability of the algorithm both on held out test data, and in actual autonomous driving experiments.
Jeff Michels, Ashutosh Saxena, Andrew Y. Ng
ICML2
2005 In Use Parameter Estimation of Inertial Sensors by Detecting Multilevel Quasi-static States
Ashutosh Saxena, Vadim Gerasimov, Sébastien Ourselin
KES (4)1
2005 Learning Depth from Single Monocular Images
abstract
We consider the task of depth estimation from a single monocular image. We take a supervised learning approach to this problem, in which we begin by collecting a training set of monocular images (of unstructured outdoor environments which include forests, trees, buildings, etc.) and their corresponding ground-truth depthmaps. Then, we apply supervised learning to predict the depthmap as a function of the image. Depth estimation is a challenging problem, since local features alone are insufficient to estimate depth at a point, and one needs to consider the global context of the image. Our model uses a discriminatively-trained Markov Random Field (MRF) that incorporates multiscale local- and global-image features, and models both depths at individual points as well as the relation between depths at different points. We show that, even on unstructured scenes, our algorithm is frequently able to recover fairly accurate depthmaps.
Ashutosh Saxena, Sung H. Chung, Andrew Y. Ng
NIPS1
2004 On Reduction of Bootstrapping Information Using Digital Multisignature
Sadybakasov Ulanbek, Ashutosh Saxena, Atul Negi
CIT2
2004 Non-linear Dimensionality Reduction by Locally Linear Isomaps
Ashutosh Saxena, Abhinav Gupta 0001, Amitabha Mukerjee
ICONIP1