Yashaswi Verma

dblp:117/4872 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
5since 2021 · last 2025
0000-0003-2317-2641ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 12 · 8 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Multi-view multi-label canonical correlation analysis for cross-modal multimedia retrieval
Asha Rani 0002, Rushil Kaushal Sanghavi, Yashaswi Verma
Multim. Tools Appl.3
2024 Activity-based Early Autism Diagnosis Using A Multi-Dataset Supervised Contrastive Learning Approach
abstract
Autism Spectrum Disorder (ASD) is a neurological disorder. Its primary symptoms include difficulty in verbal/non-verbal communication and rigid/repetitive behavior. Traditional methods of autism diagnosis require multiple visits to a human specialist. However, this process is generally time-consuming and may result in a delayed (early) intervention. In this paper, we present a data-driven approach to automate autism diagnosis using video clips of subjects performing simple activities recorded in a weakly constrained environment. This task is particularly challenging since the available training data is small, videos from the two categories ("ASD" and “Control”) are generally perceptually indistinguishable, and there is no clear understanding of what features would be beneficial in this task. To address these, we present a novel multi-dataset supervised contrastive learning technique to learn discriminative features simultaneously from multiple video datasets with significantly diverse distributions. Extensive empirical analyses demonstrate the promise of our approach compared to competing techniques on this challenging task.
Asha Rani 0002, Yashaswi Verma
WACV2
2023 Early-stage autism diagnosis using action videos and contrastive feature learning
Asha Rani 0002, Pankaj Yadav, Yashaswi Verma
Multim. Syst.3
2022 Cross-modal Retrieval Using Contrastive Learning of Visual-Semantic Embeddings
abstract
Contrastive learning is a powerful technique to learn representations that are semantically distinctive and geometrically invariant. While most of the earlier approaches have demonstrated its effectiveness on single-modality learning tasks such as image classification, recently there have been a few attempts towards extending this idea to multi-modal data. In this paper, we propose two loss functions based on normalized cross-entropy to perform the task of learning joint visual-semantic embedding using batch contrastive training. In a batch, for a given anchor point from one modality, we consider its negatives only from another modality, and define our first contrastive loss based on the expected violations incurred by all the negatives. Next, we update this loss and define the second contrastive loss based on the violation incurred only by the hardest negative. We compare our results with existing visual-semantic embedding methods on cross-modal image-to-text and text-to-image retrieval tasks using the MS-COCO and Flickr30K datasets, where we achieve competitive results and are outperformed only by adaptations of the n-pairs symmetric angular loss for multi-modal data. We have also shared our code and pre-trained models for reproducibility.
Yashaswi Verma
ICPR2
2022 Surprising Effectiveness of Random Feature Embeddings in eXtreme Classification
abstract
The goal of eXtreme Multi-label Learning (XML) is to automatically annotate a given data point with the most relevant subset of labels from an extremely large vocabulary of labels (e.g., a million labels). Lately, many attempts have been made to address this problem that achieve reasonable performance on benchmark datasets. In this paper, rather than coming-up with an altogether new method, our objective is to present and validate a simple baseline for this task. Precisely, we investigate a global and structure preserving feature embedding technique using random projections whose learning phase is independent of training samples and label vocabulary. Further, we show how an ensemble of multiple such learners can be used to achieve further boost in prediction accuracy with only linear increase in training and prediction time. Experiments on three public XML benchmarks show that the proposed approach obtains competitive accuracy compared with many existing methods. Additionally, it also provides around 6572× speed-up ratio in terms of training time and around 14.7× reduction in model-size compared to the closest competitors on the largest public dataset. We have also shared our code for reproducibility.
Yashaswi Verma
ICPR1
2020 Recurrent Image Annotation with Explicit Inter-label Dependencies
Ayushi Dutta, Yashaswi Verma, C. V. Jawahar
ECCV (29)2
2019 Diverse image annotation with missing labels
Yashaswi Verma
Pattern Recognit.1
2018 Automatic image annotation: the quirks and what works
Ayushi Dutta, Yashaswi Verma, C. V. Jawahar
Multim. Tools Appl.2
2017 A support vector approach for cross-modal search of images and texts
Yashaswi Verma, C. V. Jawahar
Comput. Vis. Image Underst.1
2017 Image Annotation by Propagating Labels from Semantic Neighbourhoods
Yashaswi Verma, C. V. Jawahar
Int. J. Comput. Vis.1
2016 A Robust Distance with Correlated Metric Learning for Multi-Instance Multi-Label Data
abstract
In multi-instance data, every object is a bag that contains multiple elements or instances. Each bag may be assigned to one or more classes, such that it has at least one instance corresponding to every assigned class. However, since the annotations are at bag-level, there is no direct association between the instances within a bag and the assigned class labels, hence making the problem significantly challenging. While existing methods have mostly focused on Bag-to-Bag or Class-to-Bag distances, in this paper, we address the multiple instance learning problem using a novel Bag-to-Class distance measure. This is based on two observations: (a) existence of outliers is natural in multi-instance data, and (b) there may exist multiple instances within a bag that belong to a particular class. In order to address these, in the proposed distance measure (a) we employ L1-distance that brings robustness against outliers, and (b) rather than considering only the most similar instance-pair during distance computation as done by existing methods, we consider a subset of instances within a bag while determining its relevance to a given class. We parameterize the proposed distance measure using class-specific distance metrics, and propose a novel metric learning framework that explicitly captures inter-class correlations within the learned metrics. Experiments on two popular datasets demonstrate the effectiveness of the proposed distance measure and metric learning.
Yashaswi Verma, C. V. Jawahar
ACM Multimedia1
2015 Exploring Locally Rigid Discriminative Patches for Learning Relative Attributes
abstract
Relative attributes help in comparing two images based on their visual properties. These are of great interest as they have been shown to be useful in several vision related problems such as recognition, retrieval, and understanding image collections in general. In the recent past, quite a few techniques have been proposed for the relative attribute learning task that give reasonable performance. However, these have focused either on the algorithmic aspect or the representational aspect. In this work, we revisit these ap-proaches and integrate their broader ideas to develop simple baselines. These not only take care of the algorithmic aspects, but also take a step towards analyzing a simple yet domain independent patch-based representation for this task. This representation can capture local shape in an image, as well as spatially rigid correspondences across regions in an image pair. The baselines are extensively evaluated on three challenging relative attribute datasets (OSR, LFW-10 and UT-Zap50K). Experiments demonstrate that they achieve promising results on the OSR and LFW-10 datasets, and perform better than the current state-of-the-art on the UT-Zap50K dataset. Moreover, they also provide some interesting insights about the problem, that could be helpful in developing the future techniques in this domain. 1
Yashaswi Verma, C. V. Jawahar
BMVC1
2015 A Probabilistic Approach for Image Retrieval Using Descriptive Textual Queries
abstract
We address the problem of image retrieval using textual queries. In particular, we focus on descriptive queries that can be either in the form of simple captions (e.g., ``a brown cat sleeping on a sofa''), or even long descriptions with multiple sentences. We present a probabilistic approach that seamlessly integrates visual and textual information for the task. It relies on linguistically and syntactically motivated mid-level textual patterns (or phrases) that are automatically extracted from available descriptions. At the time of retrieval, the given query is decomposed into such phrases, and images are ranked based on their joint relevance with these phrases. Experiments on two popular datasets (UIUC Pascal Sentence and IAPR-TC12 benchmark) demonstrate that our approach effectively retrieves semantically meaningful images, and outperforms baseline methods.
Yashaswi Verma, C. V. Jawahar
ACM Multimedia1
2014 Im2Text and Text2Im: Associating Images and Texts for Cross-Modal Retrieval
Yashaswi Verma, C. V. Jawahar
BMVC1
2014 Relative Parts: Distinctive Parts for Learning Relative Attributes
abstract
The notion of relative attributes as introduced by Parikh and Grauman (ICCV, 2011) provides an appealing way of comparing two images based on their visual properties (or attributes) such as "smiling" for face images, "naturalness" for outdoor images, etc. For learning such attributes, a Ranking SVM based formulation was proposed that uses globally represented pairs of annotated images. In this paper, we extend this idea towards learning relative attributes using local parts that are shared across categories. First, instead of using a global representation, we introduce a part-based representation combining a pair of images that specifically compares corresponding parts. Then, with each part we associate a locally adaptive "significance-coefficient" that represents its discriminative ability with respect to a particular attribute. For each attribute, the significance-coefficients are learned simultaneously with a max-margin ranking model in an iterative manner. Compared to the baseline method, the new method is shown to achieve significant improvement in relative attribute prediction accuracy. Additionally, it is also shown to improve relative feedback based interactive image search.
Ramachandruni N. Sandeep, Yashaswi Verma, C. V. Jawahar
CVPR2
2013 Exploring SVM for Image Annotation in Presence of Confusing Labels
abstract
We address the problem of automatic image annotation in large vocabulary datasets. In such datasets, for a given label, there could be several other labels that act as its confusing labels. Three possible factors for this are (i) incomplete-labeling (“cars ” vs. “vehicle”), (ii) label-ambiguity (“flowers ” vs. “blooms”), and (iii) structural-overlap (“lion ” vs. “tiger”). While previous studies in this domain have mostly focused on nearest-neighbour based models, we show that even the conventional one-vs-rest SVM significantly outperforms several benchmark models. We also demonstrate that with a simple modification in the hinge-loss of SVM, it is possible to significantly improve its performance. In particular, we introduce a tolerance-parameter in the hinge-loss. This makes the new model more tolerant against the errors in the classification of samples tagged with confusing labels as compared to other samples. This tolerance parameter is automatically determined using visual similarity and dataset statistics. Experimental evaluations demonstrate that our method (referred to as SVM with Variable Tolerance or SVM-VT) shows promising results on the task of image annotation on three challenging datasets, and establishes a baseline for such models in this domain. 1
Yashaswi Verma, C. V. Jawahar
BMVC1
2012 Choosing Linguistics over Vision to Describe Images
abstract
In this paper, we address the problem of automatically generating human-like descriptions for unseen images, given a collection of images and their corresponding human-generated descriptions. Previous attempts for this task mostly rely on visual clues and corpus statistics, but do not take much advantage of the semantic information inherent in the available image descriptions. Here, we present a generic method which benefits from all these three sources (i.e. visual clues, corpus statistics and available descriptions) simultaneously, and is capable of constructing novel descriptions. Our approach works on syntactically and linguistically motivated phrases extracted from the human descriptions. Experimental evaluations demonstrate that our formulation mostly generates lucid and semantically correct descriptions, and significantly outperforms the previous methods on automatic evaluation metrics. One of the significant advantages of our approach is that we can generate multiple interesting descriptions for an image. Unlike any previous work, we also test the applicability of our method on a large dataset containing complex images with rich descriptions.
Yashaswi Verma, C. V. Jawahar
AAAI2
2012 Image Annotation Using Metric Learning in Semantic Neighbourhoods
Yashaswi Verma, C. V. Jawahar
ECCV (3)1