Aakansha Mishra

dblp:158/9046 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
8since 2021 · last 2024
0000-0002-0792-3006ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Learning Representations from Explainable and Connectionist Approaches for Visual Question Answering
abstract
Reasoning conditioned on visual and linguistic information has gained immense importance in recent times. The prior art in Visual Question Answering (VQA) has been predominantly connectionist in nature. To resolve the issues of connectionist AI models, Symbolic models were proposed that allowed for explainable visual reasoning. In addition to semantic parsing, such models worked towards visual parsing resulting in scene graphs that provided scope for accurate reasoning conditioned on the explainable scene graphs. However, the real scenarios of VQA cannot always be segregated exclusively into connectionist (neural networks) and conceptual modalities. Rather, they are always dependent on the relationships and interactions between the two modalities. In this work, the authors proposed a question-guided attention mechanism that combines the approach of explainable visual reasoning through scene graphs with a cross-modality-based multi-head attention mechanism. The contributions of con-nectionist and conceptual modalities are learned through the semantic parsing of questions in each VQA task. The novel method is tested with the VQA2.0 and GQA and it resulted in 65.31% and 63.06% accuracy, respectively, which is better than the state-of-the-art in explainable AI.
Aakansha Mishra, Srinivas Soumitri Miriyala, Vikram Nelvoy Rajendiran
ICASSP1
2024 Efficient Visual Question Answering on Embedded Devices: Cross-Modality Attention With Evolutionary Quantization
abstract
Visual Question Answering (VQA) lies at the intersection of vision and language domains necessitating learning representations from multiple modalities. While the model development for VQA has witnessed tremendous growth, the efforts for its deployment on embedded devices have been lagging limiting its true potential. In this work, the authors address this challenge by designing a novel hardware-friendly architecture for VQA based on the transformer model with cross-modality attention. The memory footprint of the VQA model is optimized for on-device deployment using a distributed framework for Post Training Quantization (PTQ) formulated as a Non-Linear Programming (NLP) problem. The NLP problem is solved using an Evolutionary algorithm to determine the low-bit representation of the VQA model with minimal accuracy drop compared to the full precision model. The quantized model for VQA with a marginal accuracy drop of less than 2%, resulted in 4 times memory improvement, and over 2 times latency improvement, enabling its successful deployment on the Samsung Galaxy S23 device. The comprehensive study explores the potential of the proposed generic end-to-end pipeline from VQA model development to its deployment.
Aakansha Mishra, Aditya Agarwala, Utsav Tiwari, Vikram Nelvoy Rajendiran, Srinivas Soumitri Miriyala
ICIP1
2024 Visual Question Answering with Cascade of Self- and Co-Attention Blocks
Aakansha Mishra, Ashish Anand, Prithwijit Guha
ICPR (19)1
2024 Efficient Adapter on Pre-trained Visual Feature Reliance in Medical Visual Question Answering
Aakansha Mishra, Prateek Keserwani, Vikram Nelvoy Rajendiran, Ashok K. Senapati
ICPR (28)1
2024 Meta-Learned Attribute Self-Interaction Network for Continual and Generalized Zero-Shot Learning
abstract
Zero-shot learning (ZSL) is a promising approach to generalizing a model to categories unseen during training by leveraging class attributes, but challenges remain. Recently, methods using generative models to combat bias towards classes seen during training have pushed state of the art, but these generative models can be slow or computationally expensive to train. Also, these generative models assume that the attribute vector of each unseen class is available a priori at training, which is not always practical. Additionally, while many previous ZSL methods assume a one-time adaptation to unseen classes, in reality, the world is always changing, necessitating a constant adjustment of deployed models. Models unprepared to handle a sequential stream of data are likely to experience catastrophic forgetting. We propose a Meta-learned Attribute self-Interaction Network (MAIN) for continual ZSL. By pairing attribute self-interaction trained using meta-learning with inverse regularization of the attribute encoder, we are able to outperform state-of-the-art results without leveraging the unseen class attributes while also being able to train our models substantially faster (> 100×) than expensive generative-based approaches. We demonstrate this with experiments on five standard ZSL datasets (CUB, aPY, AWA1, AWA2, and SUN) in the generalized zero-shot learning and continual (fixed/dynamic) zero-shot learning settings. Extensive ablations and analyses demonstrate the efficacy of various components proposed.
Vinay Kumar Verma, Nikhil Mehta 0002, Kevin J. Liang, Aakansha Mishra, Lawrence Carin
WACV4
2023 Sub-Band Contrastive Learning-Based Knowledge Distillation For Sound Classification
abstract
Knowledge distillation(KD) technique is widely known for its outstanding ability to train a compact student model under supervision of a cumbersome pre-trained teacher network. The traditional KD technique focuses only on distilling dark knowledge using logits of a teacher network while neglecting the information regarding contrastive representation. To this end, we propose a new KD loss function that enables a student network to learn informative contrastive distribution and fine grained information from spectrogram representation of a signal thus enhancing performance of a student network for sound classification task. The experiments are conducted on two benchmark sound classification datasets, viz. ESC-10 and Audio MNIST, which illustrates that the student network trained using the proposed KD loss function outperformed the competitive KD techniques.
Achyut Mani Tripathi, Aakansha Mishra
ICASSP2
2022 Revamped Knowledge Distillation for Sound Classification
abstract
This paper presents a novel knowledge distillation technique that inherits knowledge from multiple deep Environment Sound Classification (ESC) models trained on spectrogram features created by dividing the spectrogram into multiple subband spectrogram. The deep models trained on sub-band spectrograms prevent information loss while performing knowledge distillation from a teacher model to a student model receiving the full spectrogram as an input. The student models' performance is evaluated on two benchmark sound datasets, viz. the ESC-10 and Audio MNIST datasets. The impact of teacher models trained with different number of sub-band features and four ensemble techniques has been investigated thoroughly to enhance the final accuracy of the student model supervised by the proposed knowledge distillation framework. Experiments and results shows that the accuracy of the student model is comparable and competitive to state-of-the-art methods for sound classification. Moreover, the student model trained on the Audio MNIST dataset attains an hitherto unpublished accuracy of 98.25%, a new benchmark for the Audio MNIST dataset. Additionally, Grad- CAM visualization of the spectrogram features is generated to identify the spectrogram's relevant regions and understand why the model classifies a signal into a specific class.
Achyut Mani Tripathi, Aakansha Mishra
IJCNN2
2021 Environment sound classification using an attention-based residual neural network
Achyut Mani Tripathi, Aakansha Mishra
Neurocomputing2
2020 Multi-stage Attention based Visual Question Answering
abstract
Recent developments in the field of Visual Question Answering (VQA) have witnessed promising improvements in performance through contributions in attention based networks. Most such approaches have focused on unidirectional attention that leverage over attention from textual domain (question) on visual space. These approaches mostly focused on learning high-quality attention in the visual space. In contrast, this work proposes an alternating bi-directional attention framework. First, a question to image attention helps to learn the robust visual space embedding, and second, an image to question attention helps to improve the question embedding. This attention mechanism is realized in an alternating fashion i.e. question-to-image followed by image-to-question and is repeated for maximizing performance. We believe that this process of alternating attention generation helps both the modalities and leads to better representations for the VQA task. This proposal is benchmark on TDIUC dataset and against state-of-art approaches. Our ablation analysis shows that alternate attention is the key to achieve high performance in VOA.
Aakansha Mishra, Ashish Anand, Prithwijit Guha
ICPR1
2020 CQ-VQA: Visual Question Answering on Categorized Questions
abstract
This paper proposes CQ-VQA, a novel two-level hierarchical but end-to-end model to solve the task of visual question answering (VQA). The first level of CQ-VQA, referred to as Question Categorizer (QC), classifies questions to reduce the potential answer search space. The QC uses attended and fused features of the input question and image. The second level, referred to as Answer Predictor (AP), comprises of a set of distinct classifiers corresponding to each question category. Depending on the question category predicted by QC, only one of the classifiers of AP remains active. The loss functions of QC and AP are aggregated together to make it an end-to-end model. The proposed model (CQ-VQA) is evaluated on the TDIUC dataset and is benchmarked against state-of-the-art approaches. Results indicate a competitive or better performance of CQ-VQA.
Aakansha Mishra, Ashish Anand, Prithwijit Guha
IJCNN1