Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Rahul Sukthankar

dblp:57/3775 · DBLP profile ↗
← Back
114ranked-venue papers
6as first author
5since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 87 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 79 · 3 first-author · 3 since 2021Systems, architecture and hardware · 9 · 1 first-author · 1 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
60 papers
Video understanding and tracking · 33% 3D vision · 19% Representation and self-supervised learning · 7%
Computer graphics and multimedia
15 papers
Multimedia analysis and retrieval · 30% Geometric modeling and processing · 23% Computational photography and imaging · 17%
Databases, data mining, and information retrieval
12 papers
Information retrieval · 51% Recommender systems · 36% Machine learning and data management · 9%

Topics — the 30 heaviest of 155, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
2.9162020
Speech2Action: Cross-Modal Supervision for Action Recognition · CVPR 2020
Relational Action Forecasting · CVPR 2019
AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions · CVPR 2018
Computer vision › 3D vision
human mesh recovery
1.432021
THUNDR: Transformer-based 3D HUmaN Reconstruction with Markers · ICCV 2021
Neural Descent for Visual 3D Human Pose and Shape · CVPR 2021
GHUM & GHUML: Generative 3D Human Shape and Articulated Pose Models · CVPR 2020
Computer vision › 3D vision
depth estimation
1.122022
UFO Depth: Unsupervised learning with flow-based odometry optimization for metric depth estimation · ICRA 2022
Semi-Supervised Learning for Multi-Task Scene Understanding by Neural Graph Consensus · AAAI 2021
Robotics › Robot navigation and mapping
visual navigation
0.722020
Cognitive Mapping and Planning for Visual Navigation · Int. J. Comput. Vis. 2020
Cognitive Mapping and Planning for Visual Navigation · CVPR 2017
Machine learning › Learning paradigms
semi-supervised learning
0.622021
Semi-Supervised Learning for Multi-Task Scene Understanding by Neural Graph Consensus · AAAI 2021
Semi-supervised Learning with Weakly-Related Unlabeled Data: Towards Better Text Categorization · NIPS 2008
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.612022
Discrete Representations Strengthen Vision Transformer Robustness · ICLR 2022
Machine learning › Representation and self-supervised learning
discrete representation
0.612022
Discrete Representations Strengthen Vision Transformer Robustness · ICLR 2022
Machine learning › Representation and self-supervised learning
vector quantization
0.612022
Discrete Representations Strengthen Vision Transformer Robustness · ICLR 2022
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
vision transformer robustness
0.612022
Discrete Representations Strengthen Vision Transformer Robustness · ICLR 2022
Computer vision › Video understanding and tracking › action detection
temporal action localization
0.522018
Rethinking the Faster R-CNN Architecture for Temporal Action Localization · CVPR 2018
Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images · ACM Multimedia 2015
Computer vision › 3D vision
3d human reconstruction
0.512021
THUNDR: Transformer-based 3D HUmaN Reconstruction with Markers · ICCV 2021
Machine learning › Learning paradigms
multi-task learning
0.512021
Semi-Supervised Learning for Multi-Task Scene Understanding by Neural Graph Consensus · AAAI 2021
Computer vision › Segmentation and scene understanding
semantic segmentation
0.512021
Semi-Supervised Learning for Multi-Task Scene Understanding by Neural Graph Consensus · AAAI 2021
Computer vision › 3D vision › 3d human reconstruction
single-view human reconstruction
0.512021
Neural Descent for Visual 3D Human Pose and Shape · CVPR 2021
Computer vision › 3D vision
surface normal estimation
0.512021
Semi-Supervised Learning for Multi-Task Scene Understanding by Neural Graph Consensus · AAAI 2021
Computer vision › Video understanding and tracking › action detection
spatio-temporal action localization
0.522018
AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions · CVPR 2018
Spatiotemporal Deformable Part Models for Action Detection · CVPR 2013
Computer vision › Image recognition and object detection
object detection
0.522019
Customizing Object Detectors for Indoor Robots · ICRA 2019
Rethinking the Faster R-CNN Architecture for Temporal Action Localization · CVPR 2018
Computer vision › Segmentation and scene understanding › image segmentation
co-segmentation
0.522017
Video Object Discovery and Co-Segmentation with Extremely Weak Supervision · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Video Object Discovery and Co-segmentation with Extremely Weak Supervision · ECCV (4) 2014
Computer vision › Video understanding and tracking › video analytics › video object analysis › object-centric video understanding
video object discovery
0.522017
Video Object Discovery and Co-Segmentation with Extremely Weak Supervision · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Video Object Discovery and Co-segmentation with Extremely Weak Supervision · ECCV (4) 2014
Computer vision › Video understanding and tracking
video classification
0.422016
Labeling the Features Not the Samples: Efficient Video Classification with Minimal Supervision · AAAI 2016
Large-Scale Video Classification with Convolutional Neural Networks · CVPR 2014
Computer vision › Vision and language
cross-modal supervision
0.412020
Speech2Action: Cross-Modal Supervision for Action Recognition · CVPR 2020
Computer vision › Video understanding and tracking › action recognition › action recognition under limited supervision
weakly supervised action recognition
0.412020
Speech2Action: Cross-Modal Supervision for Action Recognition · CVPR 2020
Computer vision › 3D vision › 3d reconstruction › learning-based 3d reconstruction
weakly supervised reconstruction
0.412020
Weakly Supervised 3D Human Pose and Shape Reconstruction with Normalizing Flows · ECCV (6) 2020
Geometric modeling and processing › shape modeling
human body modeling
0.412020
GHUM & GHUML: Generative 3D Human Shape and Articulated Pose Models · CVPR 2020
Computer vision › Video understanding and tracking
action detection
0.432015
Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images · ACM Multimedia 2015
Spatiotemporal Deformable Part Models for Action Detection · CVPR 2013
Spatio-temporal Shape and Flow Correlation for Action Recognition · CVPR 2007
Computer vision › Video understanding and tracking
action anticipation
0.412019
Relational Action Forecasting · CVPR 2019
Computer vision › Video understanding and tracking › action recognition › temporal action recognition
early action recognition
0.412019
Relational Action Forecasting · CVPR 2019
Robotics › Robot manipulation
grasping
0.412019
Customizing Object Detectors for Indoor Robots · ICRA 2019
Computer vision › Video understanding and tracking › action anticipation
multi-person action forecasting
0.412019
Relational Action Forecasting · CVPR 2019
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.422016
Labeling the Features Not the Samples: Efficient Video Classification with Minimal Supervision · AAAI 2016
Feature seeding for action recognition · ICCV 2011

Methods — techniques the papers use, named apart from their topics

weak supervision · 1.1optical flow · 1.0unsupervised learning · 0.7odometry optimization · 0.6discrete tokenization · 0.6convolutional neural network · 0.6neural graph consensus · 0.5meta-learning · 0.5ensemble teacher-student training · 0.5differentiable rendering · 0.5collaborative filtering · 0.5variational autoencoder · 0.4pose-space deformation · 0.4blend skinning · 0.4homography estimation · 0.3spectral clustering · 0.2inpainting · 0.2linear dynamical systems · 0.2
YearPublicationVenuePosition
2022 Discrete Representations Strengthen Vision Transformer Robustness
Chengzhi Mao, Lu Jiang 0004, Mostafa Dehghani 0001, Carl Vondrick, Rahul Sukthankar, Irfan A. Essa
ICLR5
2022 UFO Depth: Unsupervised learning with flow-based odometry optimization for metric depth estimation
abstract
We propose an efficient method for unsupervised learning of metric depth estimation from a single image in the context of unconstrained videos captured from UAVs. We combine the accuracy of an analytical solution based on odometry with the power of deep learning. First, we show how to correct the noisy odometric measurements by optimizing the alignment between the derotated optical flow and the projected linear speed in the image. Then, we detail an analytical depth estimation method based on optical flow and corrected camera velocities. Subsequently, the improved depth and camera veloc-ities obtained analytically are used, as additional cost terms, for training our novel unsupervised learning architecture for metric depth estimation. We extensively test on a recent UAV dataset, which we significantly extend by adding completely novel scenes. We outperform by significant margins different kinds of state-of-the-art approaches, ranging from analytical and unsupervised solutions to transformer-based architectures that require heavy computation and pre-training. The resulting algorithm could be deployed on embedded devices, being a good candidate for practical robotics use cases, such as obstacle avoidance and safe landing for UAV s.
Vlad Licaret, Victor Robu, Alina Marcu, Dragos Costea, Emil Slusanschi, Rahul Sukthankar, Marius Leordeanu
ICRA6
2021 Semi-Supervised Learning for Multi-Task Scene Understanding by Neural Graph Consensus
abstract
We address the challenging problem of semi-supervised learning in the context of multiple visual interpretations of the world by finding consensus in a graph of neural networks. Each graph node is a scene interpretation layer, while each edge is a deep net that transforms one layer at one node into another from a different node. During the supervised phase edge networks are trained independently. During the next unsupervised stage edge nets are trained on the pseudo-ground truth provided by consensus among multiple paths that reach the nets' start and end nodes. These paths act as ensemble teachers for any given edge and strong consensus is used for high-confidence supervisory signal. The unsupervised learning process is repeated over several generations, in which each edge becomes a "student" and also part of different ensemble "teachers" for training other students. By optimizing such consensus between different paths, the graph reaches consistency and robustness over multiple interpretations and generations, in the face of unknown labels. We give theoretical justifications of the proposed idea and validate it on a large dataset. We show how prediction of different representations such as depth, semantic segmentation, surface normals and pose from RGB input could be effectively learned through self-supervised consensus in our graph. We also compare to state-of-the-art methods for multi-task and semi-supervised learning and show superior performance.
Marius Leordeanu, Mihai Cristian Pîrvu, Dragos Costea, Alina Marcu, Emil Slusanschi, Rahul Sukthankar
AAAI6
2021 Neural Descent for Visual 3D Human Pose and Shape
abstract
We present deep neural network methodology to reconstruct the 3d pose and shape of people, including hand gestures and facial expression, given an input RGB image. We rely on a recently introduced, expressive full body statistical 3d human model, GHUM, trained end-to-end, and learn to reconstruct its pose and shape state in a self-supervised regime. Central to our methodology, is a learning to learn and optimize approach, referred to as HUman Neural Descent (HUND), which avoids both second-order differentiation when training the model parameters, and expensive state gradient descent in order to accurately minimize a semantic differentiable rendering loss at test time. Instead, we rely on novel recurrent stages to update the pose and shape parameters such that not only losses are minimized effectively, but the process is meta-regularized in order to ensure endprogress. HUND’s symmetry between training and testing makes it the first 3d human sensing architecture to natively support different operating regimes including self-supervised ones. In diverse tests, we show that HUND achieves very competitive results in datasets like H3.6M and 3DPW, as well as good quality 3d reconstructions for complex imagery collected in-the-wild.
Andrei Zanfir, Eduard Gabriel Bazavan, Mihai Zanfir, William T. Freeman, Rahul Sukthankar, Cristian Sminchisescu
CVPR5
2021 THUNDR: Transformer-based 3D HUmaN Reconstruction with Markers
abstract
We present THUNDR, a transformer-based deep neural network methodology to reconstruct the 3d pose and shape of people, given monocular RGB images. Key to our methodology is an intermediate 3d marker representation, where we aim to combine the predictive power of model-free-output architectures and the regularizing, anthropometrically-preserving properties of a statistical human surface model like GHUM—a recently introduced, expressive full body statistical 3d human model, trained end-to-end. Our novel transformer-based prediction pipeline can focus on image regions relevant to the task, supports self-supervised regimes, and ensures that solutions are consistent with human anthropometry. We show state-of-the-art results on Human3.6M and 3DPW, for both the fully-supervised and the self-supervised models, for the task of inferring 3d human shape, joint positions, and global translation. Moreover, we observe very solid 3d reconstruction performance for difficult human poses collected in the wild.
Mihai Zanfir, Andrei Zanfir, Eduard Gabriel Bazavan, William T. Freeman, Rahul Sukthankar, Cristian Sminchisescu
ICCV5
2020 Speech2Action: Cross-Modal Supervision for Action Recognition
abstract
Is it possible to guess human action from dialogue alone? In this work we investigate the link between spoken words and actions in movies. We note that movie screenplays describe actions, as well as contain the speech of characters and hence can be used to learn this correlation with no additional supervision. We train a BERT-based Speech2Action classifier on over a thousand movie screenplays, to predict action labels from transcribed speech segments. We then apply this model to the speech segments of a large unlabelled movie corpus (188M speech segments from 288K movies). Using the predictions of this model, we obtain weak action labels for over 800K video clips. By training on these video clips, we demonstrate superior action recognition performance on standard action recognition benchmarks, without using a single manually labelled action example.
Arsha Nagrani, Chen Sun 0002, David A. Ross, Rahul Sukthankar, Cordelia Schmid, Andrew Zisserman
CVPR4
2020 GHUM & GHUML: Generative 3D Human Shape and Articulated Pose Models
abstract
We present a statistical, articulated 3D human shape modeling pipeline, within a fully trainable, modular, deep learning framework. Given high-resolution complete 3D body scans of humans, captured in various poses, together with additional closeups of their head and facial expressions, as well as hand articulation, and given initial, artist designed, gender neutral rigged quad-meshes, we train all model parameters including non-linear shape spaces based on variational auto-encoders, pose-space deformation correctives, skeleton joint center predictors, and blend skinning functions, in a single consistent learning loop. The models are simultaneously trained with all the 3d dynamic scan data (over 60,000 diverse human configurations in our new dataset) in order to capture correlations and ensure consistency of various components. Models support facial expression analysis, as well as body (with detailed hand) shape and pose estimation. We provide fully train-able generic human models of different resolutions- the moderate-resolution GHUM consisting of 10,168 vertices and the low-resolution GHUML(ite) of 3,194 vertices-, run comparisons between them, analyze the impact of different components and illustrate their reconstruction from image data. The models will be available for research.
Eduard Gabriel Bazavan, Andrei Zanfir, William T. Freeman, Rahul Sukthankar, Cristian Sminchisescu
CVPR5
2020 Weakly Supervised 3D Human Pose and Shape Reconstruction with Normalizing Flows
Andrei Zanfir, Eduard Gabriel Bazavan, William T. Freeman, Rahul Sukthankar, Cristian Sminchisescu
ECCV (6)5
2020 SelfieDroneStick: A Natural Interface for Quadcopter Photography
abstract
A physical selfie stick extends the user's reach, enabling the acquisition of personal photos that include more of the background scene. Similarly, a quadcopter can capture photos from vantage points unattainable by the user; but teleoperating a quadcopter to good viewpoints is a difficult task. This paper presents a natural interface for quadcopter photography, the SelfieDroneStick that allows the user to guide the quadcopter to the optimal vantage point based on the phone's sensors. Users specify the composition of their desired long-range selfies using their smartphone, and the quadcopter autonomously flies to a sequence of vantage points from where the desired shots can be taken. The robot controller is trained from a combination of real-world images and simulated flight data. This paper describes two key innovations required to deploy deep reinforcement learning models on a real robot: 1) an abstract state representation for transferring learning from simulation to the hardware platform, and 2) reward shaping and staging paradigms for training the controller. Both of these improvements were found to be essential in learning a robot controller from simulation that transfers successfully to the real robot.
Saif Alabachi, Gita Reese Sukthankar, Rahul Sukthankar
IROS3
2020 D3D: Distilled 3D Networks for Video Action Recognition
abstract
State-of-the-art methods for action recognition commonly use two networks: the spatial stream, which takes RGB frames as input, and the temporal stream, which takes optical flow as input. In recent work, both streams are 3D Convolutional Neural Networks, which use spatiotemporal filters. These filters can respond to motion, and therefore should allow the network to learn motion representations, removing the need for optical flow. However, we still see significant benefits in performance by feeding optical flow into the temporal stream, indicating that the spatial stream is "missing" some of the signal that the temporal stream captures. In this work, we first investigate whether motion representations are indeed missing in the spatial stream, and show that there is significant room for improvement. Second, we demonstrate that these motion representations can be improved using distillation, that is, by tuning the spatial stream to mimic the temporal stream, effectively combining both models into a single stream. Finally, we show that our Distilled 3D Network (D3D) achieves performance on par with the two-stream approach, with no need to compute optical flow during inference.
Jonathan C. Stroud, David A. Ross, Chen Sun 0002, Jia Deng 0001, Rahul Sukthankar
WACV5
2020 Cognitive Mapping and Planning for Visual Navigation
Saurabh Gupta 0001, Varun Tolani, James Davidson, Sergey Levine, Rahul Sukthankar, Jitendra Malik
Int. J. Comput. Vis.5
2019 An Efficient 3D CNN for Action/Object Segmentation in Video
Rui Hou 0008, Chen Chen 0001, Rahul Sukthankar, Mubarak Shah
BMVC3
2019 Relational Action Forecasting
abstract
This paper focuses on multi-person action forecasting in videos. More precisely, given a history of H previous frames, the goal is to detect actors and to predict their future actions for the next T frames. Our approach jointly models temporal and spatial interactions among different actors by constructing a recurrent graph, using actor proposals obtained with Faster R-CNN as nodes. Our method learns to select a subset of discriminative relations without requiring explicit supervision, thus enabling us to tackle challenging visual data. We refer to our model as Discriminative Relational Recurrent Network (DRRN). Evaluation of action prediction on AVA demonstrates the effectiveness of our proposed method compared to simpler baselines. Furthermore, we significantly improve performance on the task of early action classification on J-HMDB, from the previous SOTA of 48% to 60%.
Chen Sun 0002, Abhinav Shrivastava, Carl Vondrick, Rahul Sukthankar, Kevin Murphy 0002, Cordelia Schmid
CVPR4
2019 Customizing Object Detectors for Indoor Robots
abstract
Object detection models based on convolutional neural networks (CNNs) demonstrate impressive performance when trained on large-scale labeled datasets. While a generic object detector trained on such a dataset performs adequately in applications where the input data is similar to user photographs, the detector performs poorly on small objects, particularly ones with limited training data or imaged from uncommon viewpoints. Also, a specific room will have many objects that are missed by standard object detectors, frustrating a robot that continually operates in the same indoor environment.This paper describes a system for rapidly creating customized object detectors. Data is collected from a quadcopter that is teleoperated with an interactive interface. Once an object is selected, the quadcopter autonomously photographs the object from multiple viewpoints to collect data to train DUNet (Dense Upscaled Network), our proposed model for learning customized object detectors from scratch given limited data. Our experiments compare the performance of learning models from scratch with DUNet vs. fine tuning existing state of the art object detectors, both on our indoor robotics domain and on standard datasets.
Saif Alabachi, Gita Reese Sukthankar, Rahul Sukthankar
ICRA3
2018 Rethinking the Faster R-CNN Architecture for Temporal Action Localization
abstract
We propose TAL-Net, an improved approach to temporal action localization in video that is inspired by the Faster RCNN object detection framework. TAL-Net addresses three key shortcomings of existing approaches: (1) we improve receptive field alignment using a multi-scale architecture that can accommodate extreme variation in action durations; (2) we better exploit the temporal context of actions for both proposal generation and action classification by appropriately extending receptive fields; and (3) we explicitly consider multi-stream feature fusion and demonstrate that fusing motion late is important. We achieve state-of-the-art performance for both action proposal and localization on THUMOS'14 detection benchmark and competitive performance on ActivityNet challenge.
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A. Ross, Jia Deng 0001, Rahul Sukthankar
CVPR6
2018 AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions
abstract
This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 437 15-minute video clips, where actions are localized in space and time, resulting in 1.59M action labels with multiple labels per person occurring frequently. The key characteristics of our dataset are: (1) the definition of atomic visual actions, rather than composite actions; (2) precise spatio-temporal annotations with possibly multiple annotations for each person; (3) exhaustive annotation of these atomic actions over 15-minute video clips; (4) people temporally linked across consecutive segments; and (5) using movies to gather a varied set of action representations. This departs from existing datasets for spatio-temporal action recognition, which typically provide sparse annotations for composite actions in short video clips. AVA, with its realistic scene and action complexity, exposes the intrinsic difficulty of action recognition. To benchmark this, we present a novel approach for action localization that builds upon the current state-of-the-art methods, and demonstrates better performance on JHMDB and UCF101-24 categories. While setting a new state of the art on existing datasets, the overall results on AVA are low at 15.8% mAP, underscoring the need for developing new approaches for video understanding.
Chunhui Gu, Chen Sun 0002, David A. Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, Jitendra Malik
CVPR10
2018 Actor-Centric Relation Network
Chen Sun 0002, Abhinav Shrivastava, Carl Vondrick, Kevin Murphy 0002, Rahul Sukthankar, Cordelia Schmid
ECCV (11)5
2018 Guest Editorial
Lamberto Ballan, Shih-Fu Chang, Gang Hua 0001, Thomas Mensink, Greg Mori, Rahul Sukthankar
Comput. Vis. Image Underst.6
2017 Real-Time Temporal Action Localization in Untrimmed Videos by Sub-Action Discovery
Rui Hou 0008, Rahul Sukthankar, Mubarak Shah
BMVC2
2017 Cognitive Mapping and Planning for Visual Navigation
abstract
We introduce a neural architecture for navigation in novel environments. Our proposed architecture learns to map from first-person views and plans a sequence of actions towards goals in the environment. The Cognitive Mapper and Planner (CMP) is based on two key ideas: a) a unified joint architecture for mapping and planning, such that the mapping is driven by the needs of the planner, and b) a spatial memory with the ability to plan given an incomplete set of observations about the world. CMP constructs a top-down belief map of the world and applies a differentiable neural net planner to produce the next action at each time step. The accumulated belief of the world enables the agent to track visited regions of the environment. Our experiments demonstrate that CMP outperforms both reactive strategies and standard memory-based architectures and performs well in novel environments. Furthermore, we show that CMP can also achieve semantically specified goals, such as "go to a chair".
Saurabh Gupta 0001, James Davidson, Sergey Levine, Rahul Sukthankar, Jitendra Malik
CVPR4
2017 Robust Adversarial Reinforcement Learning
abstract
Deep neural networks coupled with fast simulation and improved computational speeds have led to recent successes in the field of reinforcement learning (RL). However, most current RL-based approaches fail to generalize since: (a) the gap between simulation and real world is so large that policy-learning approaches fail to transfer; (b) even if policy learning is done in real world, the data scarcity leads to failed generalization from training to test scenarios (e.g., due to different friction or object masses). Inspired from H-infinity control methods, we note that both modeling errors and differences in training and test scenarios can just be viewed as extra forces/disturbances in the system. This paper proposes the idea of robust adversarial reinforcement learning (RARL), where we train an agent to operate in the presence of a destabilizing adversary that applies disturbance forces to the system. The jointly trained adversary is reinforced – that is, it learns an optimal destabilization policy. We formulate the policy learning as a zero-sum, minimax objective function. Extensive experiments in multiple environments (InvertedPendulum, HalfCheetah, Swimmer, Hopper, Walker2d and Ant) conclusively demonstrate that our method (a) improves training stability; (b) is robust to differences in training/test conditions; and c) outperform the baseline even in the absence of the adversary.
Lerrel Pinto, James Davidson, Rahul Sukthankar, Abhinav Gupta 0001
ICML3
2017 The THUMOS challenge on action recognition for videos "in the wild"
Haroon Idrees, Amir Zamir, Yu-Gang Jiang 0001, Alex Gorban, Ivan Laptev, Rahul Sukthankar, Mubarak Shah
Comput. Vis. Image Underst.6
2017 Behavior Discovery and Alignment of Articulated Object Classes from Unstructured Video
abstract
We propose an automatic system for organizing the content of a collection of unstructured videos of an articulated object class (e.g., tiger, horse). By exploiting the recurring motion patterns of the class across videos, our system: (1) identifies its characteristic behaviors, and (2) recovers pixel-to-pixel alignments across different instances. Our system can be useful for organizing video collections for indexing and retrieval. Moreover, it can be a platform for learning the appearance or behaviors of object classes from Internet video. Traditional supervised techniques cannot exploit this wealth of data directly, as they require a large amount of time-consuming manual annotations. The behavior discovery stage generates temporal video intervals, each automatically trimmed to one instance of the discovered behavior, clustered by type. It relies on our novel motion representation for articulated motion based on the displacement of ordered pairs of trajectories. The alignment stage aligns hundreds of instances of the class to a great accuracy despite considerable appearance variations (e.g., an adult tiger and a cub). It uses a flexible thin plate spline deformation model that can vary through time. We carefully evaluate each step of our system on a new, fully annotated dataset. On behavior discovery, we outperform the state-of-the-art improved dense trajectory feature descriptor. On spatial alignment, we outperform the popular SIFT Flow algorithm.
Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari
Int. J. Comput. Vis.3
2017 Video Object Discovery and Co-Segmentation with Extremely Weak Supervision
abstract
We present a spatio-temporal energy minimization formulation for simultaneous video object discovery and co-segmentation across multiple videos containing irrelevant frames. Our approach overcomes a limitation that most existing video co-segmentation methods possess, i.e., they perform poorly when dealing with practical videos in which the target objects are not present in many frames. Our formulation incorporates a spatio-temporal auto-context model, which is combined with appearance modeling for superpixel labeling. The superpixel-level labels are propagated to the frame level through a multiple instance boosting algorithm with spatial reasoning, based on which frames containing the target object are identified. Our method only needs to be bootstrapped with the frame-level labels for a few video frames (e.g., usually 1 to 3) to indicate if they contain the target objects or not. Extensive experiments on four datasets validate the efficacy of our proposed method: 1) object segmentation from a single video on the SegTrack dataset, 2) object co-segmentation from multiple videos on a video co-segmentation dataset, and 3) joint object discovery and co-segmentation from multiple videos containing irrelevant frames on the MOViCS dataset and XJTU-Stevens, a new dataset that we introduce in this paper. The proposed method compares favorably with the state-of-the-art in all of these experiments.
Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Zhenxing Niu, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Labeling the Features Not the Samples: Efficient Video Classification with Minimal Supervision
abstract
Feature selection is essential for effective visual recognition. We propose an efficient joint classifier learning and feature selection method that discovers sparse, compact representations of input features from a vast sea of candidates, with an almost unsupervised formulation. Our method requires only the following knowledge, which we call the feature sign - whether or not a particular feature has on average stronger values over positive samples than over negatives. We show how this can be estimated using as few as a single labeled training sample per class. Then, using these feature signs, we extend an initial supervised learning problem into an (almost) unsupervised clustering formulation that can incorporate new data without requiring ground truth labels. Our method works both as a feature selection mechanism and as a fully competitive classifier. It has important properties, low computational cost annd excellent accuracy, especially in difficult cases of very limited training data. We experiment on large-scale recognition in video and show superior speed and performance to established feature selection approaches such as AdaBoost, Lasso, greedy forward-backward selection, and powerful classifiers such as SVM.
Marius Leordeanu, Alexandra Radu, Shumeet Baluja, Rahul Sukthankar
AAAI4
2016 Discovering the Physical Parts of an Articulated Object Class from Multiple Videos
abstract
We propose a motion-based method to discover the physical parts of an articulated object class (e.g. head/torso/leg of a horse) from multiple videos. The key is to find object regions that exhibit consistent motion relative to the rest of the object, across multiple videos. We can then learn a location model for the parts and segment them accurately in the individual videos using an energy function that also enforces temporal and spatial consistency in part motion. Unlike our approach, traditional methods for motion segmentation or non-rigid structure from motion operate on one video at a time. Hence they cannot discover a part unless it displays independent motion in that particular video. We evaluate our method on a new dataset of 32 videos of tigers and horses, where we significantly outperform a recent motion segmentation method on the task of part discovery (obtaining roughly twice the accuracy).
Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari
CVPR3
2015 The Virtues of Peer Pressure: A Simple Method for Discovering High-Value Mistakes
Shumeet Baluja, Michele Covell, Rahul Sukthankar
CAIP (2)3
2015 MatchNet: Unifying feature and metric learning for patch-based matching
abstract
Motivated by recent successes on learning feature representations and on learning feature comparison functions, we propose a unified approach to combining both for training a patch matching system. Our system, dubbed Match-Net, consists of a deep convolutional network that extracts features from patches and a network of three fully connected layers that computes a similarity between the extracted features. To ensure experimental repeatability, we train MatchNet on standard datasets and employ an input sampler to augment the training set with synthetic exemplar pairs that reduce overfitting. Once trained, we achieve better computational efficiency during matching by disassembling MatchNet and separately applying the feature computation and similarity networks in two sequential stages. We perform a comprehensive set of experiments on standard datasets to carefully study the contributions of each aspect of MatchNet, with direct comparisons to established methods. Our results confirm that our unified approach improves accuracy over previous state-of-the-art results on patch matching datasets, while reducing the storage requirement for descriptors. We make pre-trained MatchNet publicly available.
Xufeng Han, Thomas K. Leung, Yangqing Jia, Rahul Sukthankar, Alexander C. Berg
CVPR4
2015 Articulated motion discovery using pairs of trajectories
abstract
We propose an unsupervised approach for discovering characteristic motion patterns in videos of highly articulated objects performing natural, unscripted behaviors, such as tigers in the wild. We discover consistent patterns in a bottom-up manner by analyzing the relative displacements of large numbers of ordered trajectory pairs through time, such that each trajectory is attached to a different moving part on the object. The pairs of trajectories descriptor relies entirely on motion and is more discriminative than state-of-the-art features that employ single trajectories. Our method generates temporal video intervals, each automatically trimmed to one instance of the discovered behavior, and clusters them by type (e.g., running, turning head, drinking water). We present experiments on two datasets: dogs from YouTube-Objects and a new dataset of National Geographic tiger videos. Results confirm that our proposed descriptor outperforms existing appearance- and trajectory-based descriptors (e.g., HOG and DTFs) on both datasets and enables us to segment unconstrained animal video into intervals containing single behaviors.
Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari
CVPR3
2015 Robust video segment proposals with painless occlusion handling
abstract
We propose a robust algorithm to generate video segment proposals. The proposals generated by our method can start from any frame in the video and are robust to complete occlusions. Our method does not assume specific motion models and even has a limited capability to generalize across videos. We build on our previous least squares tracking framework, where image segment proposals are generated and tracked using learned appearance models. The innovation in our new method lies in the use of two efficient moves, the merge move and free addition, to efficiently start segments from any frame and track them through complete occlusions, without much additional computation. Segment size interpolation is used for effectively detecting occlusions. We propose a new metric for evaluating video segment proposals on the challenging VSB-100 benchmark and present state-of-the-art results. Preliminary results are also shown for the potential use of our framework to track segments across different videos.
Zhengyang Wu 0003, Fuxin Li, Rahul Sukthankar, James M. Rehg
CVPR3
2015 Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images
abstract
We address the problem of fine-grained action localization from temporally untrimmed web videos. We assume that only weak video-level annotations are available for training. The goal is to use these weak labels to identify temporal segments corresponding to the actions, and learn models that generalize to unconstrained web videos. We find that web images queried by action names serve as well-localized highlights for many actions, but are noisily labeled. To solve this problem, we propose a simple yet effective method that takes weak video labels and noisy image labels as input, and generates localized action frames as output. This is achieved by cross-domain transfer between video frames and web images, using pre-trained deep convolutional neural networks. We then use the localized action frames to train action recognition models with long short-term memory networks. We collect a fine-grained sports action data set FGA-240 of more than 130,000 YouTube videos. It has 240 fine-grained actions under 85 sports activities. Convincing results are shown on the FGA-240 data set, as well as the THUMOS 2014 localization data set with untrimmed training videos.
Chen Sun 0002, Sanketh Shetty, Rahul Sukthankar, Ramakant Nevatia
ACM Multimedia3
2015 Exploring the Benefits of Context in 3D Gesture Recognition for Game-Based Virtual Environments
abstract
We present a systematic exploration of how to utilize video game context (e.g., player and environmental state) to modify and augment existing 3D gesture recognizers to improve accuracy for large gesture sets. Specifically, our work develops and evaluates three strategies for incorporating context into 3D gesture recognizers. These strategies include modifying the well-known Rubine linear classifier to handle unsegmented input streams and per-frame retraining using contextual information (CA-Linear); a GPU implementation of dynamic time warping (DTW) that reduces the overhead of traditional DTW by utilizing context to evaluate only relevant time sequences inside of a multithreaded kernel (CA-DTW); and a multiclass SVM with per-class probability estimation that is combined with a contextually based prior probability distribution (CA-SVM). We evaluate each strategy using a Kinect-based third-person perspective VE game prototype that combines parkour-style navigation with hand-to-hand combat. Using a simple gesture collection application to collect a set of 57 gestures and the game prototype that implements 37 of these gestures, we conduct three experiments. In the first experiment, we evaluate the effectiveness of several established classifiers on our gesture set and demonstrate state-of-the-art results using our proposed method. In our second experiment, we generate 500 random scenarios having between 5 and 19 of the 57 gestures in context. We show that the contextually aware classifiers CA-Linear, CA-DTW, and CA-SVM significantly outperform their non--contextually aware counterparts by 37.74%, 36.04%, and 20.81%, respectively. On the basis of the results of the second experiment, we derive upper-bound expectations for in-game performance for the three CA classifiers: 96.61%, 86.79%, and 96.86%, respectively. Finally, our third experiment is an in-game evaluation of the three CA classifiers with and without context. Our results show that through the use of context, we are able to achieve an average in-game recognition accuracy of 89.67% with CA-Linear compared to 65.10% without context, 79.04% for CA-DTW compared to 58.1% without context, and 90.85% with CA-SVM compared to 75.2% without context.
Eugene M. Taranta II, Thaddeus K. Simons, Rahul Sukthankar, Joseph J. LaViola Jr.
ACM Trans. Interact. Intell. Syst.3
2014 Recognition of Complex Events: Exploiting Temporal Dynamics between Underlying Concepts
abstract
While approaches based on bags of features excel at low-level action classification, they are ill-suited for recognizing complex events in video, where concept-based temporal representations currently dominate. This paper proposes a novel representation that captures the temporal dynamics of windowed mid-level concept detectors in order to improve complex event recognition. We first express each video as an ordered vector time series, where each time step consists of the vector formed from the concatenated confidences of the pre-trained concept detectors. We hypothesize that the dynamics of time series for different instances from the same event class, as captured by simple linear dynamical system (LDS) models, are likely to be similar even if the instances differ in terms of low-level visual features. We propose a two-part representation composed of fusing: (1) a singular value decomposition of block Hankel matrices (SSID-S) and (2) a harmonic signature (HS) computed from the corresponding eigen-dynamics matrix. The proposed method offers several benefits over alternate approaches: our approach is straightforward to implement, directly employs existing concept detectors and can be plugged into linear classification frameworks. Results on standard datasets such as NIST's TRECVID Multimedia Event Detection task demonstrate the improved accuracy of the proposed method.
Subhabrata Bhattacharya, Mahdi M. Kalayeh, Rahul Sukthankar, Mubarak Shah
CVPR3
2014 Large-Scale Video Classification with Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have been established as a powerful class of models for image recognition problems. Encouraged by these results, we provide an extensive empirical evaluation of CNNs on large-scale video classification using a new dataset of 1 million YouTube videos belonging to 487 classes. We study multiple approaches for extending the connectivity of a CNN in time domain to take advantage of local spatio-temporal information and suggest a multiresolution, foveated architecture as a promising way of speeding up the training. Our best spatio-temporal networks display significant performance improvements compared to strong feature-based baselines (55.3% to 63.9%), but only a surprisingly modest improvement compared to single-frame models (59.3% to 60.9%). We further study the generalization performance of our best model by retraining the top layers on the UCF-101 Action Recognition dataset and observe significant performance improvements compared to the UCF-101 baseline model (63.3% up from 43.9%).
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas K. Leung, Rahul Sukthankar, Li Fei-Fei 0001
CVPR5
2014 DaMN - Discriminative and Mutually Nearest: Exploiting Pairwise Category Proximity for Video Action Recognition
Rui Hou 0008, Amir Zamir, Rahul Sukthankar, Mubarak Shah
ECCV (3)3
2014 Video Object Discovery and Co-segmentation with Extremely Weak Supervision
Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Nanning Zheng 0001
ECCV (4)3
2014 Generalized Boundaries from Multiple Image Interpretations
abstract
Boundary detection is a fundamental computer vision problem that is essential for a variety of tasks, such as contour and region segmentation, symmetry detection and object recognition and categorization. We propose a generalized formulation for boundary detection, with closed-form solution, applicable to the localization of different types of boundaries, such as object edges in natural images and occlusion boundaries from video. Our generalized boundary detection method (Gb) simultaneously combines low-level and mid-level image representations in a single eigenvalue problem and solves for the optimal continuous boundary orientation and strength. The closed-form solution to boundary detection enables our algorithm to achieve state-of-the-art results at a significantly lower computational cost than current methods. We also propose two complementary novel components that can seamlessly be combined with Gb: first, we introduce a soft-segmentation procedure that provides region input layers to our boundary detection algorithm for a significant improvement in accuracy, at negligible computational cost; second, we present an efficient method for contour grouping and reasoning, which when applied as a final post-processing stage, further increases the boundary detection performance.
Marius Leordeanu, Rahul Sukthankar, Cristian Sminchisescu
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Classification of Cinematographic Shots Using Lie Algebra and its Application to Complex Event Recognition
abstract
In this paper, we propose a discriminative representation of a video shot based on its camera motion and demonstrate how the representation can be used for high level multimedia tasks like complex event recognition. In our technique, we assume that a homography exists between a pair of subsequent frames in a given shot. Using purely image-based methods, we compute homography parameters that serve as coarse indicators of the ambient camera motion. Next, using Lie algebra, we map the homography matrices to an intermediate vector space that preserves the intrinsic geometric structure of the transformation. The mappings are stacked temporally to generate vector time-series per shot. To extract meaningful features from time-series, we propose an efficient linear dynamical system based technique. The extracted temporal features are further used to train linear SVMs as classifiers for a particular shot class. In addition to demonstrating the efficacy of our method on a novel dataset, we extend its applicability to recognize complex events in large scale videos under unconstrained scenarios. Our empirical evaluations on eight cinematographic shot classes show that our technique performs close to approaches that involve extraction of 3-D trajectories using computationally prohibitive structure from motion techniques.
Subhabrata Bhattacharya, Ramin Mehran, Rahul Sukthankar, Mubarak Shah
IEEE Trans. Multim.3
2013 CrowdCam: Instantaneous Navigation of Crowd Images Using Angled Graph
abstract
We present a near real-time algorithm for interactively exploring a collectively captured moment without explicit 3D reconstruction. Our system favors immediacy and local coherency to global consistency. It is common to represent photos as vertices of a weighted graph, where edge weights measure similarity or distance between pairs of photos. We introduce Angled Graphs as a new data structure to organize collections of photos in a way that enables the construction of visually smooth paths. Weighted angled graphs extend weighted graphs with angles and angle weights which penalize turning along paths. As a result, locally straight paths can be computed by specifying a photo and a direction. The weighted angled graphs of photos used in this paper can be regarded as the result of discretizing the Riemannian geometry of the high dimensional manifold of all possible photos. Ultimately, our system enables everyday people to take advantage of each others' perspectives in order to create on-the-spot spatiotemporal visual experiences similar to the popular bullet-time sequence. We believe that this type of application will greatly enhance shared human experiences spanning from events as personal as parents watching their children's football game to highly publicized red carpet galas.
Aydin Arpa, Luca Ballan, Rahul Sukthankar, Gabriel Taubin, Marc Pollefeys, Ramesh Raskar
3DV3
2013 Discriminative Segment Annotation in Weakly Labeled Video
abstract
The ubiquitous availability of Internet video offers the vision community the exciting opportunity to directly learn localized visual concepts from real-world imagery. Unfortunately, most such attempts are doomed because traditional approaches are ill-suited, both in terms of their computational characteristics and their inability to robustly contend with the label noise that plagues uncurated Internet content. We present CRANE, a weakly supervised algorithm that is specifically designed to learn under such conditions. First, we exploit the asymmetric availability of real-world training data, where small numbers of positive videos tagged with the concept are supplemented with large quantities of unreliable negative data. Second, we ensure that CRANE is robust to label noise, both in terms of tagged videos that fail to contain the concept as well as occasional negative videos that do. Finally, CRANE is highly parallelizable, making it practical to deploy at large scale without sacrificing the quality of the learned solution. Although CRANE is general, this paper focuses on segment annotation, where we show state-of-the-art pixel-level segmentation results on two datasets, one of which includes a training set of spatiotemporal segments from more than 20,000 videos.
Kevin D. Tang, Rahul Sukthankar, Jay Yagnik, Li Fei-Fei 0001
CVPR2
2013 Spatiotemporal Deformable Part Models for Action Detection
abstract
Deformable part models have achieved impressive performance for object detection, even on difficult image datasets. This paper explores the generalization of deformable part models from 2D images to 3D spatiotemporal volumes to better study their effectiveness for action detection in video. Actions are treated as spatiotemporal patterns and a deformable part model is generated for each action from a collection of examples. For each action model, the most discriminative 3D sub volumes are automatically selected as parts and the spatiotemporal relations between their locations are learned. By focusing on the most distinctive parts of each action, our models adapt to intra-class variation and show robustness to clutter. Extensive experiments on several video datasets demonstrate the strength of spatiotemporal DPMs for classifying and localizing actions.
Yicong Tian, Rahul Sukthankar, Mubarak Shah
CVPR2
2013 Multi-armed recommendation bandits for selecting state machine policies for robotic systems
abstract
We investigate the problem of selecting a state-machine from a library to control a robot. We are particularly interested in this problem when evaluating such state machines on a particular robotics task is expensive. As a motivating example, we consider a problem where a simulated vacuuming robot must select a driving state machine well-suited for a particular (unknown) room layout. By borrowing concepts from collaborative filtering (recommender systems such as Netflix and Amazon.com), we present a multi-armed bandit formulation that incorporates recommendation techniques to efficiently select state machines for individual room layouts. We show that this formulation outperforms the individual approaches (recommendation, multi-armed bandits) as well as the baseline of selecting the `average best' state machine across all rooms.
Pyry Matikainen, P. Michael Furlong, Rahul Sukthankar, Martial Hebert
ICRA3
2013 Exploring the Trade-off Between Accuracy and Observational Latency in Action Recognition
Christopher Ellis, Syed Zain Masood, Marshall F. Tappen, Joseph J. LaViola Jr., Rahul Sukthankar
Int. J. Comput. Vis.5
2012 D-Nets: Beyond patch-based image descriptors
abstract
Despite much research on patch-based descriptors, SIFT remains the gold standard for finding correspondences across images and recent descriptors focus primarily on improving speed rather than accuracy. In this paper we propose Descriptor-Nets (D-Nets), a computationally efficient method that significantly improves the accuracy of image matching by going beyond patch-based approaches. D-Nets constructs a network in which nodes correspond to traditional sparsely or densely sampled keypoints, and where image content is sampled from selected edges in this net. Not only is our proposed representation invariant to cropping, translation, scale, reflection and rotation, but it is also significantly more robust to severe perspective and non-linear distortions. We present several variants of our algorithm, including one that tunes itself to the image complexity and an efficient parallelized variant that employs a fixed grid. Comprehensive direct comparisons against SIFT and ORB on standard datasets demonstrate that D-Nets dominates existing approaches in terms of precision and recall while retaining computational efficiency.
Felix von Hundelshausen, Rahul Sukthankar
CVPR2
2012 Model recommendation for action recognition
abstract
Simply choosing one model out of a large set of possibilities for a given vision task is a surprisingly difficult problem, especially if there is limited evaluation data with which to distinguish among models, such as when choosing the best “walk” action classifier from a large pool of classifiers tuned for different viewing angles, lighting conditions, and background clutter. In this paper we suggest that this problem of selecting a good model can be recast as a recommendation problem, where the goal is to recommend a good model for a particular task based on how well a limited probe set of models appears to perform. Through this conceptual remapping, we can bring to bear all the collaborative filtering techniques developed for consumer recommender systems (e.g., Netflix, Amazon.com). We test this hypothesis on action recognition, and find that even when every model has been directly rated on a training set, recommendation finds better selections for the corresponding test set than the best performers on the training set.
Pyry Matikainen, Rahul Sukthankar, Martial Hebert
CVPR2
2012 Efficient Closed-Form Solution to Generalized Boundary Detection
Marius Leordeanu, Rahul Sukthankar, Cristian Sminchisescu
ECCV (4)2
2012 Importance-weighted label prediction for active learning with noisy annotations
Liyue Zhao, Gita Reese Sukthankar, Rahul Sukthankar
ICPR3
2012 Classification of plant structures from uncalibrated image sequences
abstract
This paper demonstrates the feasibility of recovering fine-scale plant structure in 3D point clouds by leveraging recent advances in structure from motion and 3D point cloud segmentation techniques. The proposed pipeline is designed to be applicable to a broad variety of agricultural crops. A particular agricultural application is described, motivated by the need to estimate crop yield during the growing season. The structure of grapevines is classified into leaves, branches, and fruit using a combination of shape and color features, smoothed using a conditional random field (CRF). Our experiments show a classification accuracy (AUC) of 0.98 for grapes prior to ripening (while still green) and 0.96 for grapes during ripening (changing color), significantly improving over the baseline performance achieved using established methods.
Debadeepta Dey, Lily B. Mummert, Rahul Sukthankar
WACV3
2012 Unsupervised Learning for Graph Matching
Marius Leordeanu, Rahul Sukthankar, Martial Hebert
Int. J. Comput. Vis.2
2011 Violence Detection in Video Using Computer Vision Techniques
Enrique Bermejo Nievas, Oscar Déniz-Suárez, Gloria Bueno García, Rahul Sukthankar
CAIP (2)4
2011 A probabilistic representation for efficient large scale visual recognition tasks
abstract
In this paper, we present an efficient alternative to the traditional vocabulary based on bag-of-visual words (BoW) used for visual classification tasks. Our representation is both conceptually and computationally superior to the bag-of-visual words: (1) We iteratively generate a Maximum Likelihood estimate of an image given a set of characteristic features in contrast to the BoW methods where an image is represented as a histogram of visual words, (2) We randomly sample a set of characteristic features instead of employing computation-intensive clustering algorithms used during the vocabulary generation step of BoW methods. Our performance compares favorably to the state-of-the-art on experiments over three challenging human action and a scene categorization dataset, demonstrating the universal applicability of our method.
Subhabrata Bhattacharya, Rahul Sukthankar, Rong Jin 0001, Mubarak Shah
CVPR2
2011 Prop-free pointing detection in dynamic cluttered environments
abstract
Vision-based prop-free pointing detection is challenging both from an algorithmic and a systems standpoint. From a computer vision perspective, accurately determining where multiple users are pointing is difficult in cluttered environments with dynamic scene content. Standard approaches relying on appearance models or background subtraction to segment users operate poorly in this domain. We propose a method that focuses on motion analysis to detect pointing gestures and robustly estimate the pointing direction. Our algorithm is self-initializing; as the user points, we analyze the observed motion from two cameras and infer rotation centers that best explain the observed motion. From these, we group pixel-level flow into dominant pointing vectors that each originate from a rotation center and merge across views to obtain 3D pointing vectors. However, our proposed algorithm is computationally expensive, posing systems challenges even with current computing infrastructure. We achieve interactive speeds by exploiting coarse-grained parallelization over a cluster of computers. In unconstrained environments, we obtain an average angular precision of 2.7°.
Pyry Matikainen, Padmanabhan Pillai, Lily B. Mummert, Rahul Sukthankar, Martial Hebert
FG4
2011 Feature seeding for action recognition
abstract
Progress in action recognition has been in large part due to advances in the features that drive learning-based methods. However, the relative sparsity of training data and the risk of overfitting have made it difficult to directly search for good features. In this paper we suggest using synthetic data to search for robust features that can more easily take advantage of limited data, rather than using the synthetic data directly as a substitute for real data. We demonstrate that the features discovered by our selection method, which we call seeding, improve performance on an action classification task on real data, even though the synthetic data from which the features are seeded differs significantly from the real data, both in terms of appearance and the set of action classes.
Pyry Matikainen, Rahul Sukthankar, Martial Hebert
ICCV2
2011 Fast and accurate global motion compensation
Oscar Déniz-Suárez, Gloria Bueno García, Enrique Bermejo Nievas, Rahul Sukthankar
Pattern Recognit.4
2011 A holistic approach to aesthetic enhancement of photographs
abstract
This article presents an interactive application that enables users to improve the visual aesthetics of their digital photographs using several novel spatial recompositing techniques. This work differs from earlier efforts in two important aspects: (1) it focuses on both photo quality assessment and improvement in an integrated fashion, (2) it enables the user to make informed decisions about improving the composition of a photograph. The tool facilitates interactive selection of one or more than one foreground objects present in a given composition, and the system presents recommendations for where it can be relocated in a manner that optimizes a learned aesthetic metric while obeying semantic constraints. For photographic compositions that lack a distinct foreground object, the tool provides the user with crop or expansion recommendations that improve the aesthetic appeal by equalizing the distribution of visual weights between semantically different regions. The recomposition techniques presented in the article emphasize learning support vector regression models that capture visual aesthetics from user data and seek to optimize this metric iteratively to increase the image appeal. The tool demonstrates promising aesthetic assessment and enhancement results on variety of images and provides insightful directions towards future research.
Subhabrata Bhattacharya, Rahul Sukthankar, Mubarak Shah
ACM Trans. Multim. Comput. Commun. Appl.2
2010 Optimizing one-shot recognition with micro-set learning
abstract
For object category recognition to scale beyond a small number of classes, it is important that algorithms be able to learn from a small amount of labeled data per additional class. One-shot recognition aims to apply the knowledge gained from a set of categories with plentiful data to categories for which only a single exemplar is available for each. As with earlier efforts motivated by transfer learning, we seek an internal representation for the domain that generalizes across classes. However, in contrast to existing work, we formulate the problem in a fundamentally new manner by optimizing the internal representation for the one-shot task using the notion of micro-sets. A micro-set is a sample of data that contains only a single instance of each category, sampled from the pool of available data, which serves as a mechanism to force the learned representation to explicitly address the variability and noise inherent in the one-shot recognition task. We optimize our learned domain features so that they minimize an expected loss over micro-sets drawn from the training set and show that these features generalize effectively to previously unseen categories. We detail a discriminative approach for optimizing one-shot recognition using micro-sets and present experiments on the Animals with Attributes and Caltech-101 datasets that demonstrate the benefits of our formulation.
Kevin D. Tang, Marshall F. Tappen, Rahul Sukthankar, Christoph H. Lampert
CVPR3
2010 Food recognition using statistics of pairwise local features
abstract
Food recognition is difficult because food items are de-formable objects that exhibit significant variations in appearance. We believe the key to recognizing food is to exploit the spatial relationships between different ingredients (such as meat and bread in a sandwich). We propose a new representation for food items that calculates pairwise statistics between local features computed over a soft pixel-level segmentation of the image into eight ingredient types. We accumulate these statistics in a multi-dimensional histogram, which is then used as a feature vector for a discriminative classifier. Our experiments show that the proposed representation is significantly more accurate at identifying food than existing methods.
Shulin Yang, Dean Pomerleau, Rahul Sukthankar
CVPR4
2010 Representing Pairwise Spatial and Temporal Relations for Action Recognition
Pyry Matikainen, Martial Hebert, Rahul Sukthankar
ECCV (1)3
2010 Motif Discovery and Feature Selection for CRF-based Activity Recognition
abstract
Due to their ability to model sequential data without making unnecessary independence assumptions, conditional random fields (CRFs) have become an increasingly popular discriminative model for human activity recognition. However, how to represent signal sensor data to achieve the best classification performance within a CRF model is not obvious. This paper presents a framework for extracting motif features for CRF-based classification of IMU (inertial measurement unit) data. To do this, we convert the signal data into a set of motifs, approximately repeated symbolic sub sequences, for each dimension of IMU data. These motifs leverage structure in the data and serve as the basis to generate a large candidate set of features from the multi-dimensional raw data. By measuring reductions in the conditional log-likelihood error of the training samples, we can select features and train a CRF classifier to recognize human activities. An evaluation of our classifier on the CMU Multi-Modal Activity Database reveals that it outperforms the CRF-classifier trained on the raw features as well as other standard classifiers used in prior work.
Liyue Zhao, Gita Reese Sukthankar, Rahul Sukthankar
ICPR4
2010 A framework for photo-quality assessment and enhancement based on visual aesthetics
abstract
We present an interactive application that enables users to improve the visual aesthetics of their digital photographs using spatial recomposition. Unlike earlier work that focuses either on photo quality assessment or interactive tools for photo editing, we enable the user to make informed decisions about improving the composition of a photograph and to implement them in a single framework. Specifically, the user interactively selects a foreground object and the system presents recommendations for where it can be moved in a manner that optimizes a learned aesthetic metric while obeying semantic constraints. For photographic compositions that lack a distinct foreground object, our tool provides the user with cropping or expanding recommendations that improve its aesthetic quality. We learn a support vector regression model for capturing image aesthetics from user data and seek to optimize this metric during recomposition. Rather than prescribing a fully-automated solution, we allow user-guided object segmentation and inpainting to ensure that the final photograph matches the user's criteria. Our approach achieves 86% accuracy in predicting the attractiveness of unrated images, when compared to their respective human rankings. Additionally, 73% of the images recomposited using our tool are ranked more attractive than their original counterparts by human raters.
Subhabrata Bhattacharya, Rahul Sukthankar, Mubarak Shah
ACM Multimedia2
2010 Exploiting multi-level parallelism for low-latency activity recognition in streaming video
abstract
Video understanding is a computationally challenging task that is critical not only for traditionally throughput-oriented applications such as search but also latency-sensitive interactive applications such as surveillance, gaming, videoconferencing, and vision-based user interfaces. Enabling these types of video processing applications will require not only new algorithms and techniques, but new runtime systems that optimize latency as well as throughput. In this paper, we present a runtime system called Sprout that achieves low latency by exploiting the parallelism inherent in video understanding applications. We demonstrate the utility of our system on an activity recognition application that employs a robust new descriptor called MoSIFT, which explicitly augments appearance features with motion information. MoSIFT outperforms previous recognition techniques, but like other state-of-the-art techniques, it is computationally expensive -- a sequential implementation runs 100 times slower than real time. We describe the implementation of the activity recognition application on Sprout, and show that it can accurately recognize activities at full frame rate (25 fps) and low latency on a challenging airport surveillance video corpus.
Ming-yu Chen 0001, Lily B. Mummert, Padmanabhan Pillai, Alex Hauptmann 0001, Rahul Sukthankar
MMSys5
2010 Volumetric Features for Video Event Detection
Yan Ke, Rahul Sukthankar, Martial Hebert
Int. J. Comput. Vis.2
2010 A Boosting Framework for Visuality-Preserving Distance Metric Learning and Its Application to Medical Image Retrieval
abstract
Similarity measurement is a critical component in content-based image retrieval systems, and learning a good distance metric can significantly improve retrieval performance. However, despite extensive study, there are several major shortcomings with the existing approaches for distance metric learning that can significantly affect their application to medical image retrieval. In particular, "similarity" can mean very different things in image retrieval: resemblance in visual appearance (e.g., two images that look like one another) or similarity in semantic annotation (e.g., two images of tumors that look quite different yet are both malignant). Current approaches for distance metric learning typically address only one goal without consideration of the other. This is problematic for medical image retrieval where the goal is to assist doctors in decision making. In these applications, given a query image, the goal is to retrieve similar images from a reference library whose semantic annotations could provide the medical professional with greater insight into the possible interpretations of the query image. If the system were to retrieve images that did not look like the query, then users would be less likely to trust the system; on the other hand, retrieving images that appear superficially similar to the query but are semantically unrelated is undesirable because that could lead users toward an incorrect diagnosis. Hence, learning a distance metric that preserves both visual resemblance and semantic similarity is important. We emphasize that, although our study is focused on medical image retrieval, the problem addressed in this work is critical to many image retrieval systems. We present a boosting framework for distance metric learning that aims to preserve both visual and semantic similarities. The boosting framework first learns a binary representation using side information, in the form of labeled pairs, and then computes the distance as a weighted Hamming distance using the learned binary representation. A boosting algorithm is presented to efficiently learn the distance function. We evaluate the proposed algorithm on a mammographic image reference library with an Interactive Search-Assisted Decision Support (ISADS) system and on the medical image data set from ImageCLEF. Our results show that the boosting framework compares favorably to state-of-the-art approaches for distance metric learning in retrieval accuracy, with much lower computational cost. Additional evaluation with the COREL collection shows that our algorithm works well for regular image data sets.
Liu Yang 0001, Rong Jin 0001, Lily B. Mummert, Rahul Sukthankar, Adam Goode, Steven C. H. Hoi, Mahadev Satyanarayanan
IEEE Trans. Pattern Anal. Mach. Intell.4
2009 PFID: Pittsburgh fast-food image dataset
abstract
We introduce the first visual dataset of fast foods with a total of 4,545 still images, 606 stereo pairs, 303 3600videos for structure from motion, and 27 privacy-preserving videos of eating events of volunteers. This work was motivated by research on fast food recognition for dietary assessment. The data was collected by obtaining three instances of 101 foods from 11 popular fast food chains, and capturing images and videos in both restaurant conditions and a controlled lab setting. We benchmark the dataset using two standard approaches, color histogram and bag of SIFT features in conjunction with a discriminative classifier. Our dataset and the benchmarks are designed to stimulate research in this area and will be released freely to the research community.
Kapil Dhingra, Lei Yang 0063, Rahul Sukthankar, Jie Yang 0001
ICIP5
2009 The 1st workshop on large-scale multimedia retrieval and mining (LS-MMRM'09)
abstract
This workshop, as the first of its kind, aims to bring together researchers and industrial practitioners interested in large-scale multimedia data retrieval and mining. The workshop will provide a venue for the participants to explore a variety of aspects and applications on how advanced multimedia analysis techniques can be leveraged to address the challenges in large-scale data collections.
John R. Smith, Qi Tian 0001, Rahul Sukthankar
ACM Multimedia4
2009 An Integer Projected Fixed Point Method for Graph Matching and MAP Inference
abstract
Graph matching and MAP inference are essential problems in computer vision and machine learning. We introduce a novel algorithm that can accommodate both problems and solve them efficiently. Recent graph matching algorithms are based on a general quadratic programming formulation, that takes in consideration both unary and second-order terms reflecting the similarities in local appearance as well as in the pairwise geometric relationships between the matched features. In this case the problem is NP-hard and a lot of effort has been spent in finding efficiently approximate solutions by relaxing the constraints of the original problem. Most algorithms find optimal continuous solutions of the modified problem, ignoring during the optimization the original discrete constraints. The continuous solution is quickly binarized at the end, but very little attention is put into this final discretization step. In this paper we argue that the stage in which a discrete solution is found is crucial for good performance. We propose an efficient algorithm, with climbing and convergence properties, that optimizes in the discrete domain the quadratic score, and it gives excellent results either by itself or by starting from the solution returned by any graph matching algorithm. In practice it outperforms state-or-the art algorithms and it also significantly improves their performance if used in combination. When applied to MAP inference, the algorithm is a parallel extension of Iterated Conditional Modes (ICM) with climbing and convergence properties that make it a compelling alternative to the sequential ICM. In our experiments on MAP inference our algorithm proved its effectiveness by outperforming ICM and Max-Product Belief Propagation.
Marius Leordeanu, Martial Hebert, Rahul Sukthankar
NIPS3
2009 SLIPstream: scalable low-latency interactive perception on streaming data
abstract
A critical problem in implementing interactive perception applications is the considerable computational cost of current computer vision and machine learning algorithms, which typically run one to two orders of magnitude too slowly to be used interactively. Fortunately, many of these algorithms exhibit coarse-grained task and data parallelism that can be exploited across machines. The SLIPstream project focuses on building a highly-parallel runtime system called Sprout that can harness the computing power of a cluster to execute perception applications with low latency. This paper makes the case for using clusters for perception applications, describes the architecture of the Sprout runtime, and presents two compute-intensive yet interactive applications.
Padmanabhan Pillai, Lily B. Mummert, Steven W. Schlosser, Rahul Sukthankar, Casey Helfrich
NOSSDAV4
2008 Semi-Supervised Clustering via Learnt Codeword Distances
abstract
10.5244/C.22.90
Dhruv Batra, Rahul Sukthankar, Tsuhan Chen
BMVC2
2008 Fast Motion Consistency through Matrix Quantization
abstract
Determining the motion consistency between two video clips is a key component for many applications such as video event detection and human pose estimation. Shechtman and Irani recently proposed a method for measuring the motion consistency between two videos by representing the motion about each point with a space-time Harris matrix of spatial and temporal derivatives. A motion-consistency measure can be accurately estimated without explicitly calculating the optical flow from the videos, which could be noisy. However, the motion consistency calculation is computationally expensive and it must be evaluated between all possible pairs of points between the two videos. We propose a novel quantization method for the space-time Harris matrices that reduces the consistency calculation to a fast table lookup for any arbitrary consistency measure. We demonstrate that for the continuous rank drop consistency measure used by Shechtman and Irani, our quantization method is much faster and achieves the same accuracy as the existing approximation.
Pyry Matikainen, Rahul Sukthankar, Martial Hebert, Yan Ke
BMVC2
2008 Learning class-specific affinities for image labelling
abstract
Spectral clustering and eigenvector-based methods have become increasingly popular in segmentation and recognition. Although the choice of the pairwise similarity metric (or affinities) greatly influences the quality of the results, this choice is typically specified outside the learning framework. In this paper, we present an algorithm to learn class-specific similarity functions. Mapping our problem in a Conditional Random Fields (CRF) framework enables us to pose the task of learning affinities as parameter learning in undirected graphical models. There are two significant advances over previous work. First, we learn the affinity between a pair of data-points as a function of a pairwise feature and (in contrast with previous approaches) the classes to which these two data-points were mapped, allowing us to work with a richer class of affinities. Second, our formulation provides a principled probabilistic interpretation for learning all of the parameters that define these affinities. Using ground truth segmentations and labellings for training, we learn the parameters with the greatest discriminative power (in an MLE sense) on the training data. We demonstrate the power of this learning algorithm in the setting of joint segmentation and recognition of object classes. Specifically, even with very simple appearance features, the proposed method achieves state-of-the-art performance on standard datasets.
Dhruv Batra, Rahul Sukthankar, Tsuhan Chen
CVPR2
2008 Unifying discriminative visual codebook generation with classifier training for object category recognition
abstract
The idea of representing images using a bag of visual words is currently popular in object category recognition. Since this representation is typically constructed using unsupervised clustering, the resulting visual words may not capture the desired information. Recent work has explored the construction of discriminative visual codebooks that explicitly consider object category information. However, since the codebook generation process is still disconnected from that of classifier training, the set of resulting visual words, while individually discriminative, may not be those best suited for the classifier. This paper proposes a novel optimization framework that unifies codebook generation with classifier training. In our approach, each image feature is encoded by a sequence of ldquovisual bitsrdquo optimized for each category. An image, which can contain objects from multiple categories, is represented using aggregates of visual bits for each category. Classifiers associated with different categories determine how well a given image corresponds to each category. Based on the performance of these classifiers on the training data, we augment the visual words by generating additional bits. The classifiers are then updated to incorporate the new representation. These two phases are repeated until the desired performance is achieved. Experiments compare our approach to standard clustering-based methods and with state-of-the-art discriminative visual codebook generation. The significant improvements over previous techniques clearly demonstrate the value of unifying representation and classification into a single optimization framework.
Liu Yang 0001, Rong Jin 0001, Rahul Sukthankar, Frédéric Jurie
CVPR3
2008 Semi-supervised Learning with Weakly-Related Unlabeled Data: Towards Better Text Categorization
abstract
The cluster assumption is exploited by most semi-supervised learning (SSL) methods. However, if the unlabeled data is merely weakly related to the target classes, it becomes questionable whether driving the decision boundary to the low density regions of the unlabeled data will help the classification. In such case, the cluster assumption may not be valid; and consequently how to leverage this type of unlabeled data to enhance the classification accuracy becomes a challenge. We introduce Semi-supervised Learning with Weakly-Related Unlabeled Data" (SSLW), an inductive method that builds upon the maximum-margin approach, towards a better usage of weakly-related unlabeled information. Although the SSLW could improve a wide range of classification tasks, in this paper, we focus on text categorization with a small training pool. The key assumption behind this work is that, even with different topics, the word usage patterns across different corpora tends to be consistent. To this end, SSLW estimates the optimal word-correlation matrix that is consistent with both the co-occurrence information derived from the weakly-related unlabeled documents and the labeled documents. For empirical evaluation, we present a direct comparison with a number of state-of-the-art methods for inductive semi-supervised learning and text categorization; and we show that SSLW results in a significant improvement in categorization accuracy, equipped with a small training set and an unlabeled resource that is weakly related to the test beds."
Liu Yang 0001, Rong Jin 0001, Rahul Sukthankar
NIPS3
2007 Spatio-temporal Shape and Flow Correlation for Action Recognition
abstract
This paper explores the use of volumetric features for action recognition. First, we propose a novel method to correlate spatio-temporal shapes to video clips that have been automatically segmented. Our method works on over-segmented videos, which means that we do not require background subtraction for reliable object segmentation. Next, we discuss and demonstrate the complementary nature of shape- and flow-based features for action recognition. Our method, when combined with a recent flow-based correlation technique, can detect a wide range of actions in video, as demonstrated by results on a long tennis video. Although not specifically designed for whole-video classification, we also show that our method's performance is competitive with current action classification techniques on a standard video classification dataset.
Yan Ke, Rahul Sukthankar, Martial Hebert
CVPR2
2007 Beyond Local Appearance: Category Recognition from Pairwise Interactions of Simple Features
abstract
We present a discriminative shape-based algorithm for object category localization and recognition. Our method learns object models in a weakly-supervised fashion, without requiring the specification of object locations nor pixel masks in the training data. We represent object models as cliques of fully-interconnected parts, exploiting only the pairwise geometric relationships between them. The use of pairwise relationships enables our algorithm to successfully overcome several problems that are common to previously-published methods. Even though our algorithm can easily incorporate local appearance information from richer features, we purposefully do not use them in order to demonstrate that simple geometric relationships can match (or exceed) the performance of state-of-the-art object recognition algorithms.
Marius Leordeanu, Martial Hebert, Rahul Sukthankar
CVPR3
2007 Discriminative Cluster Refinement: Improving Object Category Recognition Given Limited Training Data
abstract
A popular approach to problems in image classification is to represent the image as a bag of visual words and then employ a classifier to categorize the image. Unfortunately, a significant shortcoming of this approach is that the clustering and classification are disconnected. Since the clustering into visual words is unsupervised, the representation does not necessarily capture the aspects of the data that are most useful for classification. More seriously, the semantic relationship between clusters is lost, causing the overall classification performance to suffer. We introduce "discriminative cluster refinement" (DCR), a method that explicitly models the pairwise relationships between different visual words by exploiting their co-occurrence information. The assigned class labels are used to identify the co-occurrence patterns that are most informative for object classification. DCR employs a maximum-margin approach to generate an optimal kernel matrix for classification. One important benefit of DCR is that it integrates smoothly into existing bag-of-words information retrieval systems by employing the set of visual words generated by any clustering method. While DCR could improve a broad class of information retrieval systems, this paper focuses on object category recognition. We present a direct comparison with a state-of-the art method on the PASCAL 2006 database and show that cluster refinement results in a significant improvement in classification accuracy given a small number of training examples.
Liu Yang 0001, Rong Jin 0001, Caroline Pantofaru, Rahul Sukthankar
CVPR4
2007 Semi-supervised Collaborative Text Classification
Rong Jin 0001, Rahul Sukthankar
ECML3
2007 Event Detection in Crowded Videos
abstract
Real-world actions occur often in crowded, dynamic environments. This poses a difficult challenge for current approaches to video event detection because it is difficult to segment the actor from the background due to distracting motion from other objects in the scene. We propose a technique for event recognition in crowded videos that reliably identifies actions in the presence of partial occlusion and background clutter. Our approach is based on three key ideas: (1) we efficiently match the volumetric representation of an event against oversegmented spatio-temporal video volumes; (2) we augment our shape-based features using flow; (3) rather than treating an event template as an atomic entity, we separately match by parts (both in space and time), enabling robustness against occlusions and actor variability. Our experiments on human actions, such as picking up a dropped object or waving in a crowd show reliable detection with few false positives.
Yan Ke, Rahul Sukthankar, Martial Hebert
ICCV2
2007 Interactive Search of Adipocytes in Large Collections of Digital Cellular Images
abstract
In the field of lipid research, the measurement of adipocyte size is an important but difficult problem. We describe an imaging-based solution that combines precise investigator control with semi-automated quantitation. By using unfixed live cells, we avoid many complications that arise in trying to isolate individual adipocytes. Instead, we image a small drop of live adipocyte suspension under a microscope, and then quantitate the image using an open-source software tool called FatFind. Since we have developed FatFind on the open-source Diamond distributed search platform, it inherits the scaling, parallelism and remote access attributes of Diamond. This paper reports on the design, implementation, and evaluation of FatFind.
Adam Goode, Anil Tarachandani, Lily B. Mummert, Rahul Sukthankar, Casey Helfrich, Alice Stefanni, Limor Fix, Jeffrey Saltzman 0002, Mahadev Satyanarayanan
ICME5
2007 Bayesian Active Distance Metric Learning
Liu Yang 0001, Rong Jin 0001, Rahul Sukthankar
UAI3
2007 Feature-based Part Retrieval for Interactive 3D Reassembly
abstract
We propose a novel framework for 3D reassembly, the task of assembling a solid object from its broken pieces. The primary challenge in this under-explored problem is to robustly establish compatibility between parts from one object. Feature-based techniques have shown success in domains such as 3D similarity search; unfortunately, the global features typically employed to quantify whole-object similarity are unsuitable for identifying part-level compatibility. Therefore, we propose the use of local features which, in conjunction with robust matching, have become popular for object recognition in 2D images. This paper demonstrates that an analogous framework can be successful for 3D reassembly. Automating part-level compatibility enables the construction of an interactive system for 3D reassembly, where the user can easily assemble a desired object from a large collection of pieces (many of which are irrelevant) by iteratively selecting compatible parts. We evaluate our approach on a simulated database of broken objects and show that it scales well in the presence of noise and extraneous pieces
Devi Parikh, Rahul Sukthankar, Tsuhan Chen
WACV2
2007 Shadow Elimination and Blinding Light Suppression for Interactive Projected Displays
abstract
A major problem with interactive displays based on front projection is that users cast undesirable shadows on the display surface. This paper demonstrates that shadows can be muted by redundantly illuminating the display surface using multiple projectors, all mounted at different locations. However, this technique alone does not eliminate shadows: Multiple projectors create multiple dark regions on the surface (penumbral occlusions) and cast undesirable light onto the users. These problems can be solved by eliminating shadows and suppressing the light that falls on occluding users by actively modifying the projected output. This paper categorizes various methods that can be used to achieve redundant illumination, shadow elimination, and blinding light suppression and evaluates their performance.
Jay Summet, Matthew Flagg, Tat-Jen Cham, James M. Rehg, Rahul Sukthankar
IEEE Trans. Vis. Comput. Graph.5
2006 An Efficient Algorithm for Local Distance Metric Learning
Liu Yang 0001, Rong Jin 0001, Rahul Sukthankar, Yi Liu 0054
AAAI3
2006 Correlated Label Propagation with Application to Multi-label Learning
abstract
Many computer vision applications, such as scene analysis and medical image interpretation, are ill-suited for traditional classification where each image can only be associated with a single class. This has stimulated recent work in multi-label learning where a given image can be tagged with multiple class labels. A serious problem with existing approaches is that they are unable to exploit correlations between class labels. This paper presents a novel framework for multi-label learning termed Correlated Label Propagation (CLP) that explicitly models interactions between labels in an efficient manner. As in standard label propagation, labels attached to training data points are propagated to test data points; however, unlike standard algorithms that treat each label independently, CLP simultaneously co-propagates multiple labels. Existing work eschews such an approach since naive algorithms for label co-propagation are intractable. We present an algorithm based on properties of submodular functions that efficiently finds an optimal solution. Our experiments demonstrate that CLP leads to significant gains in precision/recall against standard techniques on two real-world computer vision tasks involving several hundred labels.
Feng Kang, Rong Jin 0001, Rahul Sukthankar
CVPR (2)3
2006 Distributed localization of networked cameras
abstract
Camera networks are perhaps the most common type of sensor network and are deployed in a variety of real-world applications including surveillance, intelligent environments and scientific remote monitoring. A key problem in deploying a network of cameras is calibration, i.e., determining the location and orientation of each sensor so that observations in an image can be mapped to locations in the real world. This paper proposes a fully distributed approach for camera network calibration. The cameras collaborate to track an object that moves through the environment and reason probabilistically about which camera poses are consistent with the observed images. This reasoning employs sophisticated techniques for handling the difficult nonlinearities imposed by projective transformations, as well as the dense correlations that arise between distant cameras. Our method requires minimal overlap of the cameras' fields of view and makes very few assumptions about the motion of the object. In contrast to existing approaches, which are centralized, our distributed algorithm scales easily to very large camera networks. We evaluate the system on a real camera network with 25 nodes as well as simulated camera networks of up to 50 cameras and demonstrate that our approach performs well even when communication is lossy.
Stanislav Funiak, Carlos Guestrin, Mark A. Paskin, Rahul Sukthankar
IPSN4
2006 Distributed Inference in Dynamical Systems
abstract
We present a robust distributed algorithm for approximate probabilistic inference in dynamical systems, such as sensor networks and teams of mobile robots. Using assumed density filtering, the network nodes maintain a tractable representation of the belief state in a distributed fashion. At each time step, the nodes coordinate to condition this distribution on the observations made throughout the network, and to advance this estimate to the next time step. In addition, we identify a significant challenge for probabilistic inference in dynamical systems: message losses or network partitions can cause nodes to have inconsistent beliefs about the current state of the system. We address this problem by developing distributed algorithms that guarantee that nodes will reach an informative consistent distribution when communication is re-established. We present a suite of experimental results on real-world sensor data for two real sensor network deployments: one with 25 cameras and another with 54 temperature sensors.
Stanislav Funiak, Carlos Guestrin, Mark A. Paskin, Rahul Sukthankar
NIPS4
2005 Computer Vision for Music Identification
abstract
We describe how certain tasks in the audio domain can be effectively addressed using computer vision approaches. This paper focuses on the problem of music identification, where the goal is to reliably identify a song given a few seconds of noisy audio. Our approach treats the spectrogram of each music clip as a 2D image and transforms music identification into a corrupted sub-image retrieval problem. By employing pairwise boosting on a large set of Viola-Jones features, our system learns compact, discriminative, local descriptors that are amenable to efficient indexing. During the query phase, we retrieve the set of song snippets that locally match the noisy sample and employ geometric verification in conjunction with an EM-based "occlusion" model to identify the song that is most consistent with the observed signal. We have implemented our algorithm in a practical system that can quickly and accurately recognize music from short audio samples in the presence of distortions such as poor recording quality and significant ambient noise. Our experiments demonstrate that this approach significantly outperforms the current state-of-the-art in content-based music identification.
Yan Ke, Derek Hoiem, Rahul Sukthankar
CVPR (1)3
2005 Computer Vision for Music Identification: Video Demonstration
abstract
This paper describes a demonstration video for our music identification system. The goal of music identification is to reliably recognize a song from a small sample of noisy audio. This problem is challenging because the recording is often corrupted by noise and because the audio sample will only match a small portion of the target song. Additionally, a practical music identification system should scale (in both accuracy and speed) to databases containing hundreds of thousands of songs. Recently, the music identification problem has attracted considerable attention. However, the task remains unsolved, particularly for noisy real-world queries. We cast music identification into an equivalent sub-image retrieval framework: identify the portion of a spectrogram image from the database that best matches a given query snippet. Our approach treats the spectrogram of each music clip as a 2D image and transforms music identification into a corrupted sub-image retrieval problem.
Yan Ke, Derek Hoiem, Rahul Sukthankar
CVPR (2)3
2005 Dynamic load balancing for distributed search
abstract
This paper examines how computation can be mapped across the nodes of a distributed search system to effectively utilize available resources. We specifically address computationally intensive search of complex data, such as content-based retrieval of digital images or sounds, where sophisticated algorithms must be evaluated on the objects of interest. Since these problems require significant computation, we distribute the search over a collection of compute nodes, such as active storage devices, intermediate processors and host computers. A key challenge with mapping the desired computation to the available resources is that the most efficient distribution depends on several factors: relative power and number of compute nodes; network bandwidth between the compute nodes; the cost of evaluating query predicates; and the selectivity of the given query. This wide range of variables renders manual partitioning of the computation infeasible, particularly since some of the parameters (e.g., available network bandwidth) can change during the course of a search. This paper proposes several techniques for dynamic partitioning of computation, and demonstrates that they can significantly improve efficiency for distributed search applications.
Larry Huston, Alex Nizhner, Padmanabhan Pillai, Rahul Sukthankar, Peter Steenkiste
HPDC4
2005 SOLAR: sound object localization and retrieval in complex audio environments
abstract
The ability to identify sounds in complex audio environments is highly useful for multimedia retrieval, security, and many mobile robotic applications, but very little work has been done in this area. We present the SOLAR system, a system capable of finding sound objects, such as dog barks or car horns, in complex audio data extracted from movies. SOLAR avoids the need for segmentation by scanning over the audio data in fixed increments and classifying each short audio window separately. SOLAR employs boosted decision tree classifiers to select suitable features for modeling each sound object and to discriminate between the object of interest and all other sounds. We demonstrate the effectiveness of our approach with experiments on thirteen sound object classes trained using only tens of positive examples and tested on hours of audio data extracted from popular movies.
Derek Hoiem, Yan Ke, Rahul Sukthankar
ICASSP (5)3
2005 Efficient Visual Event Detection Using Volumetric Features
abstract
This paper studies the use of volumetric features as an alternative to popular local descriptor approaches for event detection in video sequences. Motivated by the recent success of similar ideas in object detection on static images, we generalize the notion of 2D box features to 3D spatio-temporal volumetric features. This general framework enables us to do real-time video analysis. We construct a realtime event detector for each action of interest by learning a cascade of filters based on volumetric features that efficiently scans video sequences in space and time. This event detector recognizes actions that are traditionally problematic for interest point methods - such as smooth motions where insufficient space-time interest points are available. Our experiments demonstrate that the technique accurately detects actions on real-world sequences and is robust to changes in viewpoint, scale and action speed. We also adapt our technique to the related task of human action classification and confirm that it achieves performance comparable to a current interest point based human activity recognizer on a standard database of human activities.
Yan Ke, Rahul Sukthankar, Martial Hebert
ICCV2
2005 Evaluating keypoint methods for content-based copyright protection of digital images
abstract
This paper evaluates the effectiveness of keypoint methods for content-based protection of digital images. These methods identify a set of "distinctive" regions (termed keypoints) in an image and encode them using descriptors that are robust to expected image transformations. To determine whether particular images were derived from a protected image, the keypoints for both images are generated and their descriptors matched. We describe a comprehensive set of experiments to examine how keypoint methods cope with three real-world challenges: (1) loss of keypoints due to cropping; (2) matching failures caused by approximate nearest-neighbor indexing schemes; (3) degraded descriptors due to significant image distortions. While keypoint methods perform very well in general, this paper identifies cases where the accuracy of such methods degrades.
Larry Huston, Rahul Sukthankar, Yan Ke
ICME2
2005 A Robust Visual Odometry and Precipice Detection System Using Consumer-grade Monocular Vision
abstract
We describe a monocular robot vision system which accomplishes accurate 3-DOF dead-reckoning, closed loop motion control, and precipice and obstacle detection, all in dynamic environments, using a single, consumer-grade web cam and typical laptop computer hardware. Simultaneous translation and rotation are accurately measured, and the camera need not be placed at the robot’s center of rotation. The algorithm is straightforward to implement and robust to noisy measurements. The software is based on open source computer vision libraries and is itself open source. It has been tested in a wide variety of real-world environments and on several different mobile robot platforms.
Jason Campbell, Rahul Sukthankar, Illah R. Nourbakhsh, Aroon Pahwa
ICRA2
2005 IrisNet: an internet-scale architecture for multimedia sensors
abstract
Most current sensor network research explores the use of extremely simple sensors on small devices called motes and focuses on over-coming the resource constraints of these devices. In contrast, our research explores the challenges of multimedia sensors and is motivated by the fact that multimedia devices, such as cameras, are rapidly becoming inexpensive, yet their use in a sensor network presents a number of unique challenges. For example, the data rates involved with multimedia sensors are orders of magnitude greater than those for sensor motes and this data cannot easily be processed by traditional sensor network techniques that focus on scalar data. In addition, the richness of the data generated by multimedia sensors makes them useful for a wide variety of applications. This paper presents an overview of IRISNET, a sensor network architecture that enables the creation of a planetary-scale infrastructure of multimedia sensors that can be shared by a large number of applications. To ensure the efficient collection of sensor readings, IRISNET enables the application-specific processing of sensor feeds on the significant computation resources that are typically attached to multimedia sensors. IRISNET enables the storage of sensor readings close to their source by providing a convenient and extensible distributed XML database infrastructure. Finally, IRISNET provides a number of multimedia processing primitives that enable the effective processing of sensor feeds in-network and at-sensor.
Jason Campbell, Phillip B. Gibbons, Suman Nath, Padmanabhan Pillai, Srinivasan Seshan, Rahul Sukthankar
ACM Multimedia6
2004 Visual Odometry Using Commodity Optical Flow
Jason Campbell, Rahul Sukthankar, Illah R. Nourbakhsh
AAAI2
2004 A Flexible Projector-Camera System for Multi-Planar Displays
Mark Ashdown, Matthew Flagg, Rahul Sukthankar, James M. Rehg
CVPR (2)3
2004 Object-Based Image Retrieval Using the Statistical Structure of Images
Derek Hoiem, Rahul Sukthankar, Henry Schneiderman, Larry Huston
CVPR (2)2
2004 PCA-SIFT: A More Distinctive Representation for Local Image Descriptors
Yan Ke, Rahul Sukthankar
CVPR (2)2
2004 Diamond: A Storage Architecture for Early Discard in Interactive Search
Larry Huston, Rahul Sukthankar, Rajiv Wickremesinghe, Mahadev Satyanarayanan, Gregory R. Ganger, Erik Riedel, Anastasia Ailamaki
FAST2
2004 SnapFind: brute force interactive image retrieval
abstract
SnapFind is an image retrieval system that enables efficient interactive search of large data sets by exploiting active disk technology. In contrast to earlier approaches, where data is typically pre-indexed for efficient retrieval according to a fixed scheme, SnapFind provides users with the flexibility to search non-indexed data in a brute force manner. The query is translated into a customized searchlet that is executed in parallel by processors near the storage devices. This enables the majority of irrelevant images to be discarded where they are stored. Partial results are displayed during search execution allowing users to interactively refine the query without waiting for search termination. This paper argues that algorithms with user-adjustable parameters are preferable to black-box image retrieval techniques.
Larry Huston, Rahul Sukthankar, Derek Hoiem
ICIG2
2004 Techniques for evaluating optical flow for visual odometry in extreme terrain
abstract
Motion vision (visual odometry, the estimation of camera egomotion) is a well researched field, yet has seen relatively limited use despite strong evidence from biological systems that vision can be extremely valuable for navigation. The limited use of such vision techniques has been attributed to a lack of good algorithms and insufficient computer power, but both of those problems were resolved as long as a decade ago. A gap presently yawns between theory and practice, perhaps due to perceptions of robot vision as less reliable and more complex than other types of sensing. We present an experimental methodology for assessing the real world precision and reliability of visual odometry techniques in both normal and extreme terrain. This paper evaluates the performance of a mobile robot equipped with a simple vision system in common outdoor and indoor environments, including grass, pavement, ice, and carpet. Our results show that motion vision algorithms can be robust and effective, and suggest a number of directions for further development.
Jason Campbell, Rahul Sukthankar, Illah R. Nourbakhsh
IROS2
2004 An efficient parts-based near-duplicate and sub-image retrieval system
abstract
We introduce a system for near-duplicate detection and sub-image retrieval. Such a system is useful for finding copyright violations and detecting forged images. We define near-duplicate as images altered with common transformations such as changing contrast, saturation, scaling, cropping, framing, etc. Our system builds a parts-based representation of images using distinctive local descriptors which give high quality matches even under severe transformations. To cope with the large number of features extracted from the images, we employ locality-sensitive hashing to index the local descriptors. This allows us to make approximate similarity queries that only examine a small fraction of the database. Although locality-sensitive hashing has excellent theoretical performance properties, a standard implementation would still be unacceptably slow for this application. We show that, by optimizing layout and access to the index data on disk, we can efficiently query indices containing millions of keypoints. Our system achieves near-perfect accuracy (100% precision at 99.85% recall) on the tests presented in Meng et al. [16], and consistently strong results on our own, significantly more challenging experiments. Query times are interactive even for collections of thousands of images.
Yan Ke, Rahul Sukthankar, Larry Huston
ACM Multimedia2
2003 Shadow Elimination and Occluder Light Suppression for Multi-Projector Displays
abstract
Two related problems of front projection displays, which occur when users obscure a projector, are: (i) undesirable shadows cast on the display by the users, and (ii) projected light falling on and distracting the users. This paper provides a computational framework for solving these two problems based on multiple overlapping projectors and cameras. The overlapping projectors are automatically aligned to display the same dekeystoned image. The system detects when and where shadows are cast by occluders and is able to determine the pixels, which are occluded in different projectors. Through a feedback control loop, the contributions of unoccluded pixels from other projectors are boosted in the shadowed regions, thereby eliminating the shadows. In addition, pixels, which are being occluded, are blanked, thereby preventing the projected light from falling on a user when they occlude the display. This can be accomplished even when the occluders are not visible to the camera. The paper presents results from a number of experiments demonstrating that the system converges rapidly with low steady-state errors.
Tat-Jen Cham, James M. Rehg, Rahul Sukthankar, Gita Reese Sukthankar
CVPR (2)3
2002 Projected light displays using visual feedback
abstract
A system of coordinated projectors and cameras enables the creation of projected light displays that are robust to environmental disturbances. This paper describes approaches for tackling both geometric and photometric aspects of the problem: (1) the projected image remains stable even when the system components (projector, camera or screen) are moved; (2) the display automatically removes shadows caused by users moving between a projector and the screen, while simultaneously suppressing projected light on the user. The former can be accomplished without knowing the positions of the system components. The latter can be achieved without direct observation of the occluder. We demonstrate that the system responds quickly to environmental disturbances and achieves low steady-state errors.
James M. Rehg, Matthew Flagg, Tat-Jen Cham, Rahul Sukthankar, Gita Reese Sukthankar
ICARCV4
2002 Scalable Alignment of Large-Format Multi-Projector Displays Using Camera Homography Trees
abstract
This paper presents a vision-based geometric alignment system for aligning the projectors in an arbitrarily large display wall. Existing algorithms typically rely on a single camera view and degrade in accuracy as the display resolution exceeds the camera resolution by several orders of magnitude. Naive approaches to integrating multiple zoomed camera views fail since small errors in aligning adjacent views propagate quickly over the display surface to create glaring discontinuities. Our algorithm builds and refines a camera homography tree to automatically register any number of uncalibrated camera images; the resulting system is both faster and significantly more accurate than competing approaches, reliably achieving alignment errors of 0.55 pixels on a 24-projector display in under 9 minutes. Detailed experiments compare our system to two recent display wall alignment algorithms, both on our 18 Megapixel display wall and in simulation. These results indicate that our approach achieves sub-pixel accuracy even on displays with hundreds of projectors.
Rahul Sukthankar, Grant Wallace, Kai Li 0001
IEEE Visualization2
2001 Dynamic Shadow Elimination for Multi-Projector Displays
abstract
A major problem with interactive displays based on front-projection is that users cast undesirable shadows on the display surface. This situation is only partially addressed by mounting a single projector at an extreme angle and pre-warping the projected image to undo keystoning distortions. This paper demonstrates that shadows can be muted by redundantly illuminating the display surface using multiple projectors, all mounted at different locations. However, this technique alone does not eliminate shadows: multiple projectors create multiple dark regions on the surface (penumbral occlusions). We solve the problem by using cameras to automatically identify occlusions as they occur and dynamically adjust each projector's output so that additional light is projected onto each partially-occluded patch. The system is self-calibrating: relevant homographies relating projectors, cameras and the display surface are recovered by observing the distortions induced in projected calibration patterns. The resulting redundantly-projected display retains the high image quality of a single-projector system while dynamically correcting for all penumbral occlusions. Our initial two-projector implementation operates at 3 Hz.
Rahul Sukthankar, Tat-Jen Cham, Gita Reese Sukthankar
CVPR (2)1
2001 Self-Calibrating Camera Projector Systems for Interactive Displays and Presentations
abstract
The authors demonstrate a self-calibrating system that employs uncalibrated cameras and microportable projectors to create novel interactive displays and presentations. Three benefits of ther system are detailed.
Rahul Sukthankar, Tat-Jen Cham, Gita Reese Sukthankar, James M. Rehg, David Hsu, Thomas K. Leung
ICCV1
2001 Smarter Presentations: Exploiting Homography in Camera-Projector Systems
Rahul Sukthankar, Robert G. Stockton, Matthew D. Mullin
ICCV1
2000 Memory-Based Face Recognition for Visitor Identification
abstract
We show that a simple, memory-based technique for appearance-based face recognition, motivated by the real-world task of visitor identification, can outperform more sophisticated algorithms that use principal components analysis (PCA) and neural networks. This technique is closely related to correlation templates; however, we show that the use of novel similarity measures greatly improves performance. We also show that augmenting the memory base with additional, synthetic face images results in further improvements in performance. Results of extensive empirical testing on two standard face recognition datasets are presented, and direct comparisons with published work show that our algorithm achieves comparable (or superior) results. Our system is incorporated into an automated visitor identification system that has been operating successfully in an outdoor environment since January 1999.
Terence Sim, Rahul Sukthankar, Matthew D. Mullin, Shumeet Baluja
FG2
2000 AutomaticKeystone Correction for Camera-Assisted Presentation Interfaces
Rahul Sukthankar, Robert G. Stockton, Matthew D. Mullin
ICMI1
2000 Complete Cross-Validation for Nearest Neighbor Classifiers
Matthew D. Mullin, Rahul Sukthankar
ICML2
2000 JKanji: Wavelet-Based Interactive Kanji Completion
abstract
JKanji is an interactive character completion system that provides stroke-order-independent recognition of complex handwritten glyphs such as Japanese kanji or Chinese hanzi. As the user enters each stroke, JKanji offers a menu of likely completions, generated from a robust multiscale matching algorithm augmented with a statistical language model. Unlike many existing systems, JKanji can incrementally incorporate new training examples, either to adapt to the idiosyncrasies of a particular user, or to increase its vocabulary. On a kanji input task with a vocabulary of 6369 kanji and English characters, JKanji has demonstrated 93%-96% recognition accuracy and up to 80% reduction in the number of input strokes. JKanji is computationally efficient, processing images at 5-10 Hz on an inexpensive portable computer and is well-suited for integration into personal digital assistants as an input method. JKanji's recognition system also processes low-quality digital camera images.
Robert G. Stockton, Rahul Sukthankar
ICPR2
2000 Applying Machine Learning for High-Performance Named-Entity Extraction
abstract
This paper describes a machine learning approach to building an efficient and accurate name spotting system. Finding names in free text is an important task in many text‐based applications. Most previous approaches were based on hand‐crafted modules encoding language and genre‐specific knowledge. These approaches had at least two shortcomings: They required large amounts of time and expertise to develop and were not easily portable to new languages and genres. This paper describes an extensible system that automatically combines weak evidence from different, easily available sources: parts‐of‐speech tags, dictionaries, and surface‐level syntactic information such as capitalization and punctuation. Individually, each piece of evidence is insufficient for robust name detection. However, the combination of evidence, through standard machine learning techniques, yields a system that achieves performance equivalent to the best existing hand‐crafted approaches.
Shumeet Baluja, Vibhu O. Mittal, Rahul Sukthankar
Comput. Intell.3
1998 Multiple Adaptive Agents for Tactical Driving
Rahul Sukthankar, Shumeet Baluja, John A. Hancock
Appl. Intell.1
1997 Evolving an intelligent vehicle for tactical reasoning in traffic
abstract
Recent research in automated highway systems has ranged from low-level vision-based controllers to high-level route-guidance software. However there is currently no system for tactical-level reasoning. Such a system should address tasks such as passing cars, making exits on time, and merging into a traffic stream. Our approach to this intermediate-level planning combines a distributed reasoning system (PolySAPIENT) with a novel evolutionary optimization strategy (PBIL). PBIL automatically tunes PolySAPIENT module parameters in simulation by evaluating candidate modules on various traffic scenarios. Since the control interface to the simulated vehicles is identical to that on the Carnegie Mellon Navlab vehicles, modules developed using this process can be directly ported to existing hardware. This method is currently being applied to the automated highway system domain; it also generalizes to many complex robotics tasks where multiple interacting modules must simultaneously be configured without individual module feedback.
Rahul Sukthankar, Shumeet Baluja, John A. Hancock
ICRA1