EDBT 2026 Demo / reviewers in the wild / expert
Abhirama Subramanyam Penamakuri
dblp:331/3275 · also Abhirama Subramanyam V. B. Penamakuri
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-3646-8492ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 70% Question answering and dialogue systems · 16% Efficient and distributed learning · 14% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 74% Knowledge graphs · 26% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
visual question answering |
1.5 | 2 | 2025 | When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs · EMNLP 2025 Answer Mining from a Pool of Images: Towards Retrieval-Based Visual Question Answering · IJCAI 2023 |
Natural language and speech › Question answering and dialogue systems
medical dialogue systems |
1.0 | 1 | 2026 | PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis · AAAI 2026 |
Computer vision › Vision and language › vision-language model › domain-specific vision-language model
medical vision-language model |
1.0 | 1 | 2026 | PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.9 | 1 | 2025 | When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs · EMNLP 2025 |
Computer vision › Vision and language › visual question answering
knowledge-based visual question answering |
0.8 | 1 | 2024 | Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant · EMNLP 2024 |
Computer vision › Vision and language › visual question answering
text-based visual question answering |
0.8 | 1 | 2024 | Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant · EMNLP 2024 |
Information retrieval
multimodal retrieval |
0.7 | 1 | 2023 | Answer Mining from a Pool of Images: Towards Retrieval-Based Visual Question Answering · IJCAI 2023 |
Medical and health informatics
computer-aided diagnosis |
0.3 | 1 | 2026 | PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis · AAAI 2026 |
Computer vision › Vision and language
vision-language model |
0.3 | 1 | 2025 | When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 2.0fine-tuning · 2.0dialogue simulation · 2.0visual text recognition · 1.5large multimodal model · 1.5relevance encoder · 1.3BART · 1.3unlabeled data training · 0.9knowledge distillation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient DiagnosisabstractTraditionally, AI research in medical diagnosis has largely centered on image analysis. While this has led to notable advancements, the absence of patient-reported symptoms continues to hinder diagnostic accuracy. To address this, we propose a Pre-Consultation Dialogue Framework (PCDF) that mimics real-world diagnostic procedures, where doctors iteratively query patients before reaching a conclusion. Specifically, we simulate diagnostic dialogues between two vision–language models (VLMs): a DocVLM, which generates follow-up questions based on the image and dialogue history, and a PatientVLM, which responds using a symptom profile derived from the ground-truth diagnosis. We additionally conducted a small-scale clinical validation of the synthetic symptoms generated by our framework, with licensed clinicians confirming their clinical relevance, symptom coverage, and overall realism. These findings indicate that the resulting DocVLM–PatientVLM interactions form coherent, multi-turn consultations paired with images and diagnoses, which we then use to fine-tune the DocVLM. This dialogue-based supervision leads to substantial gains over image-only training, highlighting the value of realistic symptom elicitation for diagnosis. K. Lokesh, Abhirama Subramanyam Penamakuri, Uday Agarwal, Apoorva Challa, Shreya K. Gowda, Somesh Gupta, Anand Mishra 0001 |
AAAI | 2 |
| 2025 | When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMsabstractLarge Vision-Language Models (L-VLMs) have demonstrated remarkable performance in various vision and language tasks, including visual question answering (VQA).However, their high computational cost makes them impractical for resource-constrained settings and inference-heavy applications.In contrast, Small Vision-Language Models (S-VLMs) offer efficiency but suffer from a significant performance gap compared to their larger counterparts.In this work, we introduce the Model Parity Aligner (MPA), a novel framework designed to systematically improve S-VLMs by leveraging unlabeled images and effective knowledge transfer from L-VLMs.Instead of traditional knowledge distillation methods that rely on labeled training data, MPA employs a strategic parity-based approach that precisely identifies the knowledge disparities between S-VLMs and L-VLMs, and optimizes training by targeting only these disparities.We conduct extensive experiments on four diverse VQA benchmarks, namely TextVQA, ST-VQA, ChartQA, and OKVQA, each of which required specialized reasoning capabilities such as text recognition, chart interpretation, and commonsense and factual understanding.Our results demonstrate that MPA consistently enhances the performance of S-VLM on all benchmarks, reducing the performance gap while maintaining computational efficiency.We make our code publicly available.Q: What time is it on the clock?A: 5:00 Q: Which airline is represented by the blue and white plane Abhirama Subramanyam Penamakuri, Navlika Singh, Piyush Arora, Anand Mishra 0001 |
EMNLP | 1 |
| 2025 | Audiopedia: Audio QA with KnowledgeabstractIn this paper, we introduce Audiopedia, a novel task called Audio Question Answering with Knowledge, which requires both audio comprehension and external knowledge reasoning. Unlike traditional Audio Question Answering (AQA) benchmarks that focus on simple queries answerable from audio alone, Audiopedia targets knowledge-intensive questions. We define three sub-tasks: (i) Single Audio Question Answering (s-AQA), where questions are answered based on a single audio sample, (ii) Multi-Audio Question Answering (m-AQA), which requires reasoning over multiple audio samples, and (iii) Retrieval-Augmented Audio Question Answering (r-AQA), which involves retrieving relevant audio to answer the question. We benchmark large audio language models (LALMs) on these sub-tasks and observe suboptimal performance. To address this, we propose a generic framework that can be adapted to any LALM, equipping them with knowledge reasoning capabilities. Our framework has two components: (i) Audio Entity Linking (AEL) and (ii) Knowledge-Augmented Audio Large Multimodal Model (KA2LM), which together improve performance on knowledge-intensive AQA tasks. To our knowledge, this is the first work to address advanced speech audio understanding through knowledge-intensive tasks like Audiopedia. We make our data publicly available: https://abhiram4572.github.io/projects/audiopedia/. Abhirama Subramanyam Penamakuri, Kiran Chhatre, Akshat Jain |
ICASSP | 1 |
| 2024 | Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal AssistantabstractWe revisit knowledge-aware text-based visual question answering, also known as Text-KVQA in the light of modern advancements in large multimodal models (LMMs), and make the following contributions: (i) We propose VisTEL -a principled approach to perform visual text entity linking.The proposed VisTEL module harnesses a state-of-the-art visual text recognition engine and the power of a large multimodal model to jointly reason using textual and visual context obtained using surrounding cues in the image to link the visual text entity to the correct knowledge base entity.(ii) We present KaLMA -knowledge-aware large multimodal assistant that augments an LMM with knowledge associated with visual text entity in the image to arrive at an accurate answer.Further, we provide a comprehensive experimental analysis and comparison of our approach with traditional visual question answering, pre-large multimodal models, and large multimodal models, as well as prior top-performing approaches.Averaging over three splits of Text-KVQA, our proposed approach surpasses the previous best approach by a substantial 23.3% on an absolute scale and establishes a new state of the art.We make our implementation publicly available. Abhirama Subramanyam Penamakuri, Anand Mishra 0001 |
EMNLP | 1 |
| 2023 | Answer Mining from a Pool of Images: Towards Retrieval-Based Visual Question AnsweringabstractWe study visual question answering in a setting where the answer has to be mined from a pool of relevant and irrelevant images given as a context. For such a setting, a model must first retrieve relevant images from the pool and answer the question from these retrieved images. We refer to this problem as retrieval-based visual question answering (or RETVQA in short). The RETVQA is distinctively different and more challenging than the traditionally-studied Visual Question Answering (VQA), where a given question has to be answered with a single relevant image in context. Towards solving the RETVQA task, we propose a unified Multi Image BART (MI-BART) that takes a question and retrieved images using our relevance encoder for free-form fluent answer generation. Further, we introduce the largest dataset in this space, namely RETVQA, which has the following salient features: multi-image and retrieval requirement for VQA, metadata-independent questions over a pool of heterogeneous images, expecting a mix of classification-oriented and open-ended generative answers. Our proposed framework achieves an accuracy of 76.5% and a fluency of 79.3% on the proposed dataset, namely RETVQA and also outperforms state-of-the-art methods by 4.9% and 11.8% on the image segment of the publicly available WebQA dataset on the accuracy and fluency metrics, respectively. Abhirama Subramanyam Penamakuri, Manish Gupta 0001, Mithun Das Gupta, Anand Mishra 0001 |
IJCAI | 1 |