EDBT 2026 Demo / reviewers in the wild / expert
Vignesh Ramanathan
dblp:30/10073
· DBLP profile ↗
25ranked-venue papers
12as first author
7since 2021 · last 2024
0000-0002-0119-4420ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 12 first-author · 7 since 2021Artificial intelligence and machine learning · 21 · 9 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Context Diffusion: In-Context Aware Image Generation
Ivona Najdenkoska, Animesh Sinha, Abhimanyu Dubey, Dhruv Mahajan 0001, Vignesh Ramanathan, Filip Radenovic |
ECCV (77) | 5 |
| 2023 | PartDistillation: Learning Parts from Instance SegmentationabstractWe present a scalable framework to learn part segmentation from object instance labels. State-of-the-art instance segmentation models contain a surprising amount of part information. However, much of this information is hidden from plain view. For each object instance, the part information is noisy, inconsistent, and incomplete. PartDistillation transfers the part information of an instance segmentation model into a part segmentation model through self-supervised self-training on a large dataset. The resulting segmentation model is robust, accurate, and generalizes well. We evaluate the model on various part segmentation datasets. Our model outperforms supervised part segmentation in zero-shot generalization performance by a large margin. Our model outperforms when finetuned on target datasets compared to supervised counterpart and other baselines especially in few-shot regime. Finally, our model provides a wider coverage of rare parts when evaluated over 10K object classes. Code is at https://github.com/facebookresearch/PartDistillation. Jang Hyun Cho, Philipp Krähenbühl, Vignesh Ramanathan |
CVPR | 3 |
| 2023 | Filtering, Distillation, and Hard Negatives for Vision-Language Pre-TrainingabstractVision-language models trained with contrastive learning on large-scale noisy data are becoming increasingly popular for zero-shot recognition problems. In this paper we improve the following three aspects of the contrastive pre-training pipeline: dataset noise, model initialization and the training objective. First, we propose a straightforward filtering strategy titled Complexity, Action, and Text-spotting (CAT) that significantly reduces dataset size, while achieving improved performance across zero-shot vision-language tasks. Next, we propose an approach titled Concept Distillation to leverage strong unimodal representations for contrastive training that does not increase training complexity while outperforming prior work. Finally, we modify the traditional contrastive alignment objective, and propose an importance-sampling approach to up-sample the importance of hard-negatives without adding additional complexity. On an extensive zero-shot benchmark of 29 tasks, our Distilled and Hard-negative Training (DiHT) approach improves on 20 tasks compared to the baseline. Furthermore, for few-shot linear probing, we propose a novel approach that bridges the gap between zero-shot and few-shot performance, substantially improving over prior work. Models are available at github.com/facebookresearch/diht. Filip Radenovic, Abhimanyu Dubey, Abhishek Kadian, Todor Mihaylov, Simon Vandenhende, Yi Wen 0006, Vignesh Ramanathan, Dhruv Mahajan 0001 |
CVPR | 8 |
| 2023 | PACO: Parts and Attributes of Common ObjectsabstractObject models are gradually progressing from predicting just category labels to providing detailed descriptions of object instances. This motivates the need for large datasets which go beyond traditional object masks and provide richer annotations such as part masks and attributes. Hence, we introduce PACO: Parts and Attributes of Common Objects. It spans 75 object categories, 456 object-part categories and 55 attributes across image (LVIS) and video (Eg04D) datasets. We provide 641K part masks an-notated across 260K object boxes, with roughly half of them exhaustively annotated with attributes as well. We design evaluation metrics and provide benchmark results for three tasks on the dataset: part mask segmentation, object and part attribute prediction and zero-shot instance detection. Dataset, models, and code are open-sourced at https://github.com/jacebookresearch/paco. Vignesh Ramanathan, Anmol Kalia, Vladan Petrovic, Yi Wen 0006, Baixue Zheng, Baishan Guo, Rui Wang 0067, Aaron Marquez, Rama Kovvuri, Abhishek Kadian, Amir Mousavi, Yiwen Song, Abhimanyu Dubey, Dhruv Mahajan 0001 |
CVPR | 1 |
| 2021 | Weakly Supervised Instance Segmentation for Videos With Temporal Mask ConsistencyabstractWeakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) partial segmentation of objects and (b) missing object predictions. We show that these issues can be better addressed by training with weakly labeled videos instead of images. In videos, motion and temporal consistency of predictions across frames provide complementary signals which can help segmentation. We are the first to explore the use of these video signals to tackle weakly supervised instance segmentation. We propose two ways to leverage this information in our model. First, we adapt inter-pixel relation network (IRN) [1] to effectively incorporate motion information during training. Second, we introduce a new MaskConsist module, which addresses the problem of missing object instances by transferring stable predictions between neighboring frames during training. We demonstrate that both approaches together improve the instance segmentation metric AP50on video frames of two datasets: Youtube-VIS and Cityscapes by 5% and 3% respectively. Qing Liu 0017, Vignesh Ramanathan, Dhruv Mahajan 0001, Alan L. Yuille, Zhenheng Yang |
CVPR | 2 |
| 2021 | Adaptive Methods for Real-World Domain GeneralizationabstractInvariant approaches have been remarkably successful in tackling the problem of domain generalization, where the objective is to perform inference on data distributions different from those used in training. In our work, we investigate whether it is possible to leverage domain information from the unseen test samples themselves. We propose a domain-adaptive approach consisting of two steps: a) we first learn a discriminative domain embedding from unsupervised training examples, and b) use this domain embedding as supplementary information to build a domain-adaptive model, that takes both the input as well as its domain into account while making predictions. For unseen domains, our method simply uses few unlabelled test examples to construct the domain embedding. This enables adaptive classification on any unseen domain. Our approach achieves state-of-the-art performance on various domain generalization benchmarks. In addition, we introduce the first real-world, large-scale domain generalization benchmark, Geo-YFCC, containing 1.1M samples over 40 training, 7 validation and 15 test domains, orders of magnitude larger than prior work. We show that the existing approaches either do not scale to this dataset or underperform compared to the simple baseline of training a model on the union of data from all training domains. In contrast, our approach achieves a significant 1% improvement. Abhimanyu Dubey, Vignesh Ramanathan, Alex Pentland, Dhruv Mahajan 0001 |
CVPR | 2 |
| 2021 | PreDet: Large-scale weakly supervised pre-training for detectionabstractState-of-the-art object detection approaches typically rely on pre-trained classification models to achieve better performance and faster convergence. We hypothesize that classification pre-training strives to achieve translation invariance, and consequently ignores the localization aspect of the problem. We propose a new large-scale pre-training strategy for detection, where noisy class labels are available for all images, but not bounding-boxes. In this setting, we augment standard classification pre-training with a new detection-specific pretext task. Motivated by the noise-contrastive learning based self-supervised approaches, we design a task that forces bounding boxes with high-overlap to have similar representations in different views of an image, compared to non-overlapping boxes. We redesign Faster R-CNN modules to perform this task efficiently. Our experimental results show significant improvements over existing weakly-supervised and self-supervised pre-training approaches in both detection accuracy as well as fine-tuning speed. Vignesh Ramanathan, Rui Wang 0067, Dhruv Mahajan 0001 |
ICCV | 1 |
| 2020 | DLWL: Improving Detection for Lowshot Classes With Weakly Labelled DataabstractLarge detection datasets have a long tail of lowshot classes with very few bounding box annotations. We wish to improve detection for lowshot classes with weakly labelled web-scale datasets only having image-level labels. This requires a detection framework that can be jointly trained with limited number of bounding box annotated images and large number of weakly labelled images. Towards this end, we propose a modification to the FRCNN model to automatically infer label assignment for objects proposals from weakly labelled images during training. We pose this label assignment as a Linear Program with constraints on the number and overlap of object instances in an image. We show that this can be solved efficiently during training for weakly labelled images. Compared to just training with few annotated examples, augmenting with weakly labelled examples in our framework provides significant gains. We demonstrate this on the LVIS dataset 3.5 gain in AP as well as different lowshot variants of the COCO dataset. We provide a thorough analysis of the effect of amount of weakly labelled and fully labelled data required to train the detection model. Our DLWL framework can also outperform self-supervised baselines like omni-supervision for lowshot classes. Vignesh Ramanathan, Rui Wang 0067, Dhruv Mahajan 0001 |
CVPR | 1 |
| 2020 | QUICKSAL: A small and sparse visual saliency model for efficient inference in resource constrained hardwareabstractVisual saliency is an important problem in the field of cognitive science and computer vision with applications such as surveillance, adaptive compressing, detecting unknown objects and scene understanding. In this paper, we propose a small and sparse neural network model for performing salient object segmentation that is suitable for use in mobile and embedded applications. Our model is built using depthwise separable convolutions and bottleneck inverted residuals which have been proven to perform very memory efficient inference and can be easily implemented using standard functions available in all deep learning frameworks. The multiscale features extracted along the layers with deep residuals allow our network to learn high quality saliency maps. We present the quantitative results of our QUICKSAL model with multiple levels of model sparsity ranging from 0% to ~96%, with the non-zero parameter count varying from ~3.3M to ~0.14M respectively - on publicly available benchmark datasets - showing that our highly constrained approach is comparable to other state-of-the-art approaches (parameter count ~35M). We also present qualitative results on camouflage images and show that our model can successfully distinguish between the salient and non-salient parts even when both seem blended together. Vignesh Ramanathan, Pritesh Dwivedi, Bharath Katabathuni, Anirban Chakraborty 0001, Chetan Singh Thakur |
WACV | 1 |
| 2019 | Activity Driven Weakly Supervised Object DetectionabstractWeakly supervised object detection aims at reducing the amount of supervision required to train detection models. Such models are traditionally learned from images/videos labelled only with the object class and not the object bounding box. In our work, we try to leverage not only the object class labels but also the action labels associated with the data. We show that the action depicted in the image/video can provide strong cues about the location of the associated object. We learn a spatial prior for the object dependent on the action (e.g. "ball" is closer to "leg of the person" in "kicking ball"), and incorporate this prior to simultaneously train a joint object detection and action classification model. We conducted experiments on both video datasets and image datasets to evaluate the performance of our weakly supervised object detection model. Our approach outperformed the current state-of-the-art (SOTA) method by more than 6% in mAP on the Charades video dataset. Zhenheng Yang, Dhruv Mahajan 0001, Deepti Ghadiyaram, Ramakant Nevatia, Vignesh Ramanathan |
CVPR | 5 |
| 2018 | What Makes a Video a Video: Analyzing Temporal Information in Video Understanding Models and DatasetsabstractThe ability to capture temporal information has been critical to the development of video understanding models. While there have been numerous attempts at modeling motion in videos, an explicit analysis of the effect of temporal information for video understanding is still missing. In this work, we aim to bridge this gap and ask the following question: How important is the motion in the video for recognizing the action? To this end, we propose two novel frameworks: (i) class-agnostic temporal generator and (ii) motion-invariant frame selector to reduce/remove motion for an ablation analysis without introducing other artifacts. This isolates the analysis of motion from other aspects of the video. The proposed frameworks provide a much tighter estimate of the effect of motion (from 25% to 6% on UCF101 and 15% to 5% on Kinetics) compared to baselines in our analysis. Our analysis provides critical insights about existing models like C3D, and how it could be made to achieve comparable results with a sparser set of frames. De-An Huang, Vignesh Ramanathan, Dhruv Mahajan 0001, Lorenzo Torresani, Manohar Paluri, Li Fei-Fei 0001, Juan Carlos Niebles |
CVPR | 2 |
| 2018 | Exploring the Limits of Weakly Supervised Pretraining
Dhruv Mahajan 0001, Ross B. Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li 0001, Ashwin Bharambe, Laurens van der Maaten |
ECCV (2) | 3 |
| 2017 | Learning to Learn from Noisy Web VideosabstractUnderstanding the simultaneously very diverse and intricately fine-grained set of possible human actions is a critical open problem in computer vision. Manually labeling training videos is feasible for some action classes but doesnt scale to the full long-tailed distribution of actions. A promising way to address this is to leverage noisy data from web queries to learn new actions, using semi-supervised or webly-supervised approaches. However, these methods typically do not learn domain-specific knowledge, or rely on iterative hand-tuned data labeling policies. In this work, we instead propose a reinforcement learning-based formulation for selecting the right examples for training a classifier from noisy web search results. Our method uses Q-learning to learn a data labeling policy on a small labeled training dataset, and then uses this to automatically label noisy web data for new visual concepts. Experiments on the challenging Sports-1M action recognition benchmark as well as on additional fine-grained and newly emerging action classes demonstrate that our method is able to learn good labeling policies for noisy data and use this to learn accurate visual concept classifiers. Serena Yeung-Levy, Vignesh Ramanathan, Olga Russakovsky, Liyue Shen, Greg Mori, Li Fei-Fei 0001 |
CVPR | 2 |
| 2016 | Social LSTM: Human Trajectory Prediction in Crowded SpacesabstractPedestrians follow different trajectories to avoid obstacles and accommodate fellow pedestrians. Any autonomous vehicle navigating such a scene should be able to foresee the future positions of pedestrians and accordingly adjust its path to avoid collisions. This problem of trajectory prediction can be viewed as a sequence generation task, where we are interested in predicting the future trajectory of people based on their past positions. Following the recent success of Recurrent Neural Network (RNN) models for sequence prediction tasks, we propose an LSTM model which can learn general human movement and predict their future trajectories. This is in contrast to traditional approaches which use hand-crafted functions such as Social forces. We demonstrate the performance of our method on several public datasets. Our model outperforms state-of-the-art methods on some of these datasets. We also analyze the trajectories predicted by our model to demonstrate the motion behaviour learned by our model. Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei 0001, Silvio Savarese |
CVPR | 3 |
| 2016 | Detecting Events and Key Actors in Multi-person VideosabstractMulti-person event recognition is a challenging task, often with many people active in the scene but only a small subset contributing to an actual event. In this paper, we propose a model which learns to detect events in such videos while automatically "attending" to the people responsible for the event. Our model does not use explicit annotations regarding who or where those people are during training and testing. In particular, we track people in videos and use a recurrent neural network (RNN) to represent the track features. We learn time-varying attention weights to combine these features at each time-instant. The attended features are then processed using another RNN for event detection/ classification. Since most video datasets with multiple people are restricted to a small number of videos, we also collected a new basketball dataset comprising 257 basketball games with 14K event annotations corresponding to 11 event classes. Our model outperforms state-of-the-art methods for both event classification and detection on this new dataset. Additionally, we show that the attention mechanism is able to consistently localize the relevant players. Vignesh Ramanathan, Jonathan Huang, Sami Abu-El-Haija, Alexander N. Gorban, Kevin Murphy 0002, Li Fei-Fei 0001 |
CVPR | 1 |
| 2015 | Learning semantic relationships for better action retrieval in imagesabstractHuman actions capture a wide variety of interactions between people and objects. As a result, the set of possible actions is extremely large and it is difficult to obtain sufficient training examples for all actions. However, we could compensate for this sparsity in supervision by leveraging the rich semantic relationship between different actions. A single action is often composed of other smaller actions and is exclusive of certain others. We need a method which can reason about such relationships and extrapolate unobserved actions from known actions. Hence, we propose a novel neural network framework which jointly extracts the relationship between actions and uses them for training better action retrieval models. Our model incorporates linguistic, visual and logical consistency based cues to effectively identify these relationships. We train and test our model on a largescale image dataset of human actions. We show a significant improvement in mean AP compared to different baseline methods including the HEX-graph approach from Deng et al. [8]. Vignesh Ramanathan, Jia Deng 0001, Zhen Li 0028, Kunlong Gu, Yang Song 0009, Samy Bengio, Charles Rosenberg 0001, Li Fei-Fei 0001 |
CVPR | 1 |
| 2015 | Learning Temporal Embeddings for Complex Video AnalysisabstractIn this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are sequences of temporally and semantically coherent images. We leverage this information to learn temporal embeddings for video frames by associating frames with the temporal context that they appear in. To do this, we propose a scheme for incorporating temporal context based on past and future frames in videos, and compare this to other contextual representations. In addition, we show how data augmentation using multi-resolution samples and hard negatives helps to significantly improve the quality of the learned embeddings. We evaluate various design decisions for learning temporal embeddings, and show that our embeddings can improve performance for multiple video tasks such as retrieval, classification, and temporal order recovery in unconstrained Internet video. Vignesh Ramanathan, Kevin D. Tang, Greg Mori, Li Fei-Fei 0001 |
ICCV | 1 |
| 2014 | Socially-Aware Large-Scale Crowd ForecastingabstractIn crowded spaces such as city centers or train stations, human mobility looks complex, but is often influenced only by a few causes. We propose to quantitatively study crowded environments by introducing a dataset of 42 million trajectories collected in train stations. Given this dataset, we address the problem of forecasting pedestrians' destinations, a central problem in understanding large-scale crowd mobility. We need to overcome the challenges posed by a limited number of observations (e.g. sparse cameras), and change in pedestrian appearance cues across different cameras. In addition, we often have restrictions in the way pedestrians can move in a scene, encoded as priors over origin and destination (OD) preferences. We propose a new descriptor coined as Social Affinity Maps (SAM) to link broken or unobserved trajectories of individuals in the crowd, while using the OD-prior in our framework. Our experiments show improvement in performance through the use of SAM features and OD prior. To the best of our knowledge, our work is one of the first studies that provides encouraging results towards a better understanding of crowd behavior at the scale of million pedestrians. Alexandre Alahi, Vignesh Ramanathan, Li Fei-Fei 0001 |
CVPR | 2 |
| 2014 | Linking People in Videos with "Their" Names Using Coreference Resolution
Vignesh Ramanathan, Armand Joulin, Percy Liang, Li Fei-Fei 0001 |
ECCV (1) | 1 |
| 2014 | Pixel rearrangement based statistical restoration scheme reducing embedding noise
Arijit Sur, Vignesh Ramanathan, Jayanta Mukhopadhyay |
Multim. Tools Appl. | 2 |
| 2013 | Social Role Discovery in Human EventsabstractWe deal with the problem of recognizing social roles played by people in an event. Social roles are governed by human interactions, and form a fundamental component of human event description. We focus on a weakly supervised setting, where we are provided different videos belonging to an event class, without training role labels. Since social roles are described by the interaction between people in an event, we propose a Conditional Random Field to model the inter-role interactions, along with person specific social descriptors. We develop tractable variational inference to simultaneously infer model weights, as well as role assignment to all people in the videos. We also present a novel YouTube social roles dataset with ground truth role annotations, and introduce annotations on a subset of videos from the TRECVID-MED11 [1] event kits for evaluation purposes. The performance of the model is compared against different baseline methods on these datasets. Vignesh Ramanathan, Bangpeng Yao, Li Fei-Fei 0001 |
CVPR | 1 |
| 2013 | Video Event Understanding Using Natural Language DescriptionsabstractHuman action and role recognition play an important part in complex event understanding. State-of-the-art methods learn action and role models from detailed spatio temporal annotations, which requires extensive human effort. In this work, we propose a method to learn such models based on natural language descriptions of the training videos, which are easier to collect and scale with the number of actions and roles. There are two challenges with using this form of weak supervision: First, these descriptions only provide a high-level summary and often do not directly mention the actions and roles occurring in a video. Second, natural language descriptions do not provide spatio temporal annotations of actions and roles. To tackle these challenges, we introduce a topic-based semantic relatedness (SR) measure between a video description and an action and role label, and incorporate it into a posterior regularization objective. Our event recognition system based on these action and role models matches the state-of-the-art method on the TRECVID-MED11 event kit, despite weaker supervision. Vignesh Ramanathan, Percy Liang, Li Fei-Fei 0001 |
ICCV | 1 |
| 2012 | Shifting Weights: Adapting Object Detectors from Image to VideoabstractTypical object detectors trained on images perform poorly on video, as there is a clear distinction in domain between the two types of data. In this paper, we tackle the problem of adapting object detectors learned from images to work well on videos. We treat the problem as one of unsupervised domain adaptation, in which we are given labeled data from the source domain (image), but only unlabeled data from the target domain (video). Our approach, self-paced domain adaptation, seeks to iteratively adapt the detector by re-training the detector with automatically discovered target domain examples, starting with the easiest first. At each iteration, the algorithm adapts by considering an increased number of target domain examples, and a decreased number of source domain examples. To discover target domain examples from the vast amount of video data, we introduce a simple, robust approach that scores trajectory tracks instead of bounding boxes. We also show how rich and expressive features specific to the target domain can be incorporated under the same framework. We show promising results on the 2011 TRECVID Multimedia Event Detection and LabelMe Video datasets that illustrate the benefit of our approach to adapt object detectors to video. Kevin D. Tang, Vignesh Ramanathan, Li Fei-Fei 0001, Daphne Koller |
NIPS | 2 |
| 2011 | Quadtree decomposition based extended vector space model for image retrievalabstractBag of visual words approach for image retrieval does not exploit the spatial distribution of visual words in an image. Previous attempts to incorporate the spatial distribution include modification of visual vocabulary using visual phrases along with visual words and use of spatial pyramid matching (SPM) techniques for comparing two images. This paper proposes a novel extended vector space based image retrieval technique which takes into account the spatial occurrence (context) of a visual word in an image along with the co-occurrence of other visual words in a pre-defined region (block) of the image obtained by quadtree decomposition of the image up to a fixed level of resolution. Experiments show a 19.22% increase in Mean Average Precision (MAP) over the BoW approach for the Caltech 101 database. Vignesh Ramanathan, Shaunak Mishra, Pabitra Mitra |
WACV | 1 |
| 2010 | Fast Non-Local Means (NLM) Computation With Probabilistic Early TerminationabstractA speed up technique for the non-local means (NLM) image denoising algorithm based on probabilistic early termination (PET) is proposed. A significant amount of computation in the NLM scheme is dedicated to the distortion calculation between pixel neighborhoods. The proposed PET scheme adopts a probability model to achieve early termination. Specifically, the distortion computation can be terminated and the corresponding contributing pixel can be rejected earlier, if the expected distortion value is too high to be of significance in weighted averaging. Performance comparative with several fast NLM schemes is provided to demonstrate the effectiveness of the proposed algorithm. Vignesh Ramanathan, Byung Tae Oh, C.-C. Jay Kuo |
IEEE Signal Process. Lett. | 1 |