Aditay Tripathi

dblp:202/6464 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0002-1667-9623ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
abstract
While Large Vision Language Models (LVLMs) are increasingly deployed in real-world applications, their ability to interpret abstract visual inputs remains limited. Specifically, they struggle to comprehend hand-drawn sketches, a modality that offers an intuitive means of expressing concepts that are difficult to describe textually. We identify the primary bottleneck as the absence of a large-scale dataset that jointly models sketches, photorealistic images, and corresponding natural language instructions. To address this, we present two key contributions: (1) a new, large-scale dataset of image-sketch-instruction triplets designed to facilitate both pretraining and instruction tuning, and (2) O3SLM, an LVLM trained on this dataset. Comprehensive evaluations on multiple sketch-based tasks: (a) object localization, (b) counting, (c) image retrieval i.e., (SBIR and fine-grained SBIR), and (d) visual question answering (VQA); while incorporating the three existing sketch datasets, namely QuickDraw!, Sketchy, and Tu-Berlin, along with our generated SketchVCL dataset, show that O3SLM achieves state-of-the-art performance, substantially outperforming existing LVLMs in sketch comprehension and reasoning.
Rishi Gupta, Mukilan Karuppasamy, Shyam Marjit, Aditay Tripathi, Anirban Chakraborty 0001
AAAI4
2024 Dynamic Data Selection for Efficient SSL via Coarse-to-Fine Refinement
Aditay Tripathi, Pradeep Shenoy, Anirban Chakraborty 0001
ECCV (58)1
2024 Enhancing 3D Referential Grounding by Learning Coarse Spatial Relationships
Soham Joshi, Aditay Tripathi, Viswanath Gopalakrishnan, Anirban Chakraborty 0001
ICPR (30)2
2024 Query-guided Attention in Vision Transformers for Localizing Objects Using a Single Sketch
abstract
In this study, we explore sketch-based object localization on natural images. Given a crude hand-drawn object sketch, the task is to locate all instances of that object in the target image. This problem proves difficult due to the abstract nature of hand-drawn sketches, variations in the style and quality of sketches, and the large domain gap between the sketches and the natural images. Existing solutions address this using attention-based frameworks to merge query information into image features. Yet, these methods often integrate query features after independently learning image features, causing inadequate alignment and as a result incorrect localization. In contrast, we propose a novel sketch-guided vision transformer encoder that uses cross-attention after each block of the transformer-based image encoder to learn query-conditioned image features, leading to stronger alignment with the query sketch. Further, at the decoder’s output, object and sketch features are refined better to align the representation of objects with the sketch query, thereby improving localization. The proposed model also generalizes to the object categories not seen during training, as the target image features learned by the proposed model are query-aware. Our framework can utilize multiple sketch queries via a trainable novel sketch fusion strategy. The model is evaluated on the images from the public benchmark, MS-COCO, using the sketch queries from QuickDraw! and Sketchy datasets. Compared with existing localization methods, the proposed approach gives a 6.6% and 8.0% improvement in mAP for seen objects using sketch queries from QuickDraw! and Sketchy datasets, respectively, and a 12.2% improvement in AP@50 for large objects that are ‘unseen’ during training. The code is available at https://vcl-iisc.github.io/locformer/.
Aditay Tripathi, Anand Mishra 0001, Anirban Chakraborty 0001
WACV1
2024 Multimodal query-guided object localization
Aditay Tripathi, Rajath R. Dani, Anand Mishra 0001, Anirban Chakraborty 0001
Multim. Tools Appl.1
2023 Edges to Shapes to Concepts: Adversarial Augmentation for Robust Vision
abstract
Recent work has shown that deep vision models tend to be overly dependent on low-level or “texture” features, leading to poor generalization. Various data augmentation strategies have been proposed to overcome this so-called texture bias in DNNs. We propose a simple, lightweight adversarial augmentation technique that explicitly incentivizes the network to learn holistic shapes for accurate prediction in an object classification setting. Our augmentations superpose edgemaps from one image onto another image with shuffled patches, using a randomly determined mixing proportion, with the image label of the edgemap image. To classify these augmented images, the model needs to not only detect and focus on edges but distinguish between relevant and spurious edges. We show that our augmentations significantly improve classification accuracy and robustness measures on a range of datasets and neural architectures. As an example, for ViT-S, We obtain absolute gains on classification accuracy gains up to 6%. We also obtain gains of up to 28% and 8.5% on natural adversarial and out-of-distribution datasets like ImageNet-A (for ViT-B) and ImageNet-R (for ViT-S), respectively. Analysis using a range of probe datasets shows substantially increased shape sensitivity in our trained models, explaining the observed improvement in robustness and classification accuracy.
Aditay Tripathi, Rishubh Singh, Anirban Chakraborty 0001, Pradeep Shenoy
CVPR1
2023 Grounding Scene Graphs on Natural Images via Visio-Lingual Message Passing
abstract
This paper presents a framework for jointly grounding objects that follow certain semantic relationship constraints given in a scene graph. A typical natural scene contains several objects, often exhibiting visual relationships of varied complexities between them. These inter-object relationships provide strong contextual cues towards improving grounding performance compared to a traditional object query-only-based localization task. A scene graph is an efficient and structured way to represent all the objects and their semantic relationships in the image. In an attempt towards bridging these two modalities representing scenes and utilizing contextual information for improving object localization, we rigorously study the problem of grounding scene graphs on natural images. To this end, we propose a novel graph neural network-based approach referred to as Visio-Lingual Message Passing Graph Neural Network (VL-MPAG Net). In VL-MPAG Net, we first construct a directed graph with object proposals as nodes and an edge between a pair of nodes representing a plausible relation between them. Then a three-step inter-graph and intra-graph message passing is performed to learn the context- dependent representation of the proposals and query objects. These object representations are used to score the proposals to generate object localization. The proposed method significantly outperforms the baselines on four public datasets.
Aditay Tripathi, Anand Mishra 0001, Anirban Chakraborty 0001
WACV1
2020 Improving Multi-hop Question Answering over Knowledge Graphs using Knowledge Base Embeddings
abstract
Knowledge Graphs (KG) are multi-relational graphs consisting of entities as nodes and relations among them as typed edges.Goal of the Question Answering over KG (KGQA) task is to answer natural language queries posed over the KG.Multi-hop KGQA requires reasoning over multiple edges of the KG to arrive at the right answer.KGs are often incomplete with many missing links, posing additional challenges for KGQA, especially for multi-hop KGQA.Recent research on multihop KGQA has attempted to handle KG sparsity using relevant external text, which isn't always readily available.In a separate line of research, KG embedding methods have been proposed to reduce KG sparsity by performing missing link prediction.Such KG embedding methods, even though highly relevant, have not been explored for multi-hop KGQA so far.We fill this gap in this paper and propose EmbedKGQA.EmbedKGQA is particularly effective in performing multi-hop KGQA over sparse KGs.EmbedKGQA also relaxes the requirement of answer selection from a prespecified neighborhood, a sub-optimal constraint enforced by previous multi-hop KGQA methods.Through extensive experiments on multiple benchmark datasets, we demonstrate EmbedKGQA's effectiveness over other stateof-the-art baselines.
Apoorv Saxena, Aditay Tripathi, Partha P. Talukdar
ACL2
2020 Sketch-Guided Object Localization in Natural Images
Aditay Tripathi, Rajath R. Dani, Anand Mishra 0001, Anirban Chakraborty 0001
ECCV (6)1
2018 Adversarial Learning of Raw Speech Features for Domain Invariant Speech Recognition
abstract
Recent advances in neural network based acoustic modelling have shown significant improvements in automatic speech recognition (ASR) performance. In order for acoustic models to be able to handle large acoustic variability, large amounts of labeled data is necessary, which are often expensive to obtain. This paper explores the application of adversarial training to learn features from raw speech that are invariant to acoustic variability. This acoustic variability is referred to as a domain shift in this paper. The experimental study presented in this paper leverages the architecture of Domain Adversarial Neural Networks (DANNs) [1] which uses data from two different domains. The DANN is a Y-shaped network that consists of a multi-layer CNN feature extractor module that is common to a label (senone) classifier and a so-called domain classifier. The utility of DANNs is evaluated on multiple datasets with domain shifts caused due to differences in gender and speaker accents. Promising empirical results indicate the strength of adversarial training for unsupervised domain adaptation in ASR, thereby emphasizing the ability of DANNs to learn domain invariant features from raw speech.
Aditay Tripathi, Aanchan Mohan, Saket Anand, Maneesh Kumar Singh 0001
ICASSP1
2017 Asymmetric stacked autoencoder
abstract
Traditional stacked autoencoders have an equal number of encoders and decoders. However, while fine-tuned as a deep neural network the decoder portion is detached and never used. This begs the question: ‘do we need equal number of decoders and encoders’? In this study we explore asymmetric autoencoders — unequal number of encoders and decoders. We specifically address two tasks — 1. Classification capacity as deep neural network and 2. Compressibility of stacked autoencoder. For both the problems, our asymmetric autoencoders have several encoders but a single decoders. We find that such autoencoders are more accurate compared to traditional symmetrically stacked autoencoders for classification accuracy and also yield slightly better results on compression problems.
Angshul Majumdar, Aditay Tripathi
IJCNN2