VLDB 2026 Research / reviewers in the wild / expert
Akash Ghosh
dblp:157/1699
· DBLP profile ↗
14ranked-venue papers
8as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Background Matters: Breaking Medical Vision Language Models by Transferable AttackabstractVision-Language Models (VLMs) are increasingly used in clinical diagnostics, yet their robustness to adversarial attacks remains largely unexplored, posing serious risks.Existing medical attacks focus on secondary objectives such as model stealing or adversarial finetuning, while transferable attacks from natural images introduce visible distortions that clinicians can easily detect.To address this, we propose MedFocusLeak, a highly transferable black-box multimodal attack that induces incorrect yet clinically plausible diagnoses while keeping perturbations imperceptible.The method injects coordinated perturbations into non-diagnostic background regions and employs an attention-distraction mechanism to shift the model's focus away from pathological areas.Extensive evaluations across six medical imaging modalities show that MedFocusLeak achieves state-of-the-art performance, generating misleading yet realistic diagnostic outputs across diverse VLMs.We further introduce a unified evaluation framework with novel metrics that jointly capture attack success and image fidelity, revealing a critical weakness in the reasoning capabilities of modern clinical VLMs.The code associated with this project is available at MedFocusLeak. Akash Ghosh, Subhadip Baidya, Sriparna Saha 0001, Xiuying Chen |
ACL (1) | 1 |
| 2026 | CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical ReasoningabstractEric Onyame, Akash Ghosh, Subhadip Baidya, Sriparna Saha, Xiuying Chen, Chirag Agarwal. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Eric Onyame, Akash Ghosh, Subhadip Baidya, Sriparna Saha 0001, Xiuying Chen, Chirag Agarwal |
ACL (1) | 2 |
| 2026 | RADO: Trustworthy Radiology Impression Generation Using Safety and Faithfulness-Based Preference OptimizationabstractRadiology impression generation involves producing concise, clinically meaningful summaries from detailed imaging findings such as CT and MRI scans, serving as a critical aid in diagnosis and treatment planning. However, recent studies highlight a severe shortage of radiologists, particularly in low and middle-income countries, where there is fewer than one radiologist per 100,000 people, making timely expert interpretation a significant challenge. While advancements in AI, especially large language models (LLMs), offer promising potential to automate this task, current systems often suffer from hallucinations, omissions of key clinical details, and a lack of linguistic clarity, thereby raising serious concerns about their safety and reliability in real-world clinical settings. In this work, we attempted to address this issue by introducing RADO , a novel framework for radiology impression generation that integrates safety, faithfulness, and linguistic refinement rewards for preference optimization. To support robust evaluation, we introduce RIB , a real-world benchmark dataset curated and annotated by radiologists, spanning 1,429 annotated CT and MRI findings and impressions across 27 study types. RADO enforces critical safety and factuality constraints via carefully designed reward models and achieves state-of-the-art performance across multiple automatic and human evaluation metrics. Our framework significantly outperforms existing baselines, demonstrating improved factual consistency, reduced omissions, and higher clinical relevance, thus advancing the safety and reliability of generative AI in high-stakes medical applications. The code and dataset associated with the work are made available at RADO . Akash Ghosh, Nitesh Patnaik, Adity Prakash, Rishi Raj, Sriparna Saha 0001 |
ACM Trans. Comput. Heal. | 1 |
| 2025 | Infogen: Generating Complex Statistical Infographics from DocumentsabstractAkash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Akash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha 0001 |
ACL (1) | 1 |
| 2025 | M3Retrieve: Benchmarking Multimodal Retrieval for MedicineabstractWith the increasing use of Retrieval-Augmented Generation (RAG), strong retrieval models have become more important than ever.In healthcare, multimodal retrieval models that combine information from both text and images offer major advantages for many downstream tasks such as question answering, cross-modal retrieval, and multimodal summarization, since medical data often includes both formats.However, there is currently no standard benchmark to evaluate how well these models perform in medical settings.To address this gap, we introduce M3Retrieve, a Multimodal Medical Retrieval Benchmark.M3Retrieve, spans 5 domains,16 medical fields, and 4 distinct tasks, with over 1.2 Million text documents and 164K multimodal queries, all collected under approved licenses.We evaluate leading multimodal retrieval models on this benchmark to explore the challenges specific to different medical specialities and to understand their impact on retrieval performance.By releasing M3Retrieve, we aim to enable systematic evaluation, foster model innovation, and accelerate research toward building more capable and reliable multimodal retrieval systems for medical applications. Arkadeep Acharya, Akash Ghosh, Pradeepika Verma, Kitsuchart Pasupa, Sriparna Saha 0001, Priti Singh |
EMNLP | 2 |
| 2025 | DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian CultureabstractArijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Nishant Mishra, Sriparna Saha. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Arijit Maji, Raghvendra Kumar 0003, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Sriparna Saha 0001 |
EMNLP | 3 |
| 2025 | Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of SportsabstractPunit Kumar Singh, Nishant Kumar, Akash Ghosh, Kunal Pasad, Khushi Soni, Manisha Jaishwal, Sriparna Saha, Syukron Abu Ishaq Alfarozi, Asres Temam Abagissa, Kitsuchart Pasupa, Haiqin Yang, Jose G Moreno. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Punit Kumar Singh, Akash Ghosh, Kunal Pasad, Khushi Soni, Manisha Jaishwal, Sriparna Saha 0001, Syukron Abu Ishaq Alfarozi, Asres Temam Abagissa, Kitsuchart Pasupa, Haiqin Yang, José G. Moreno 0001 |
EMNLP | 3 |
| 2025 | Reasoning and Planning for Multimodal Large Language Models: A Multilingual and Cross-Domain ExplorationabstractRecent advancements in Multimodal Large Language Models (MLLMs), coupled with the progress of reinforcement learning, have substantially enhanced reasoning and decision-making across modalities, including text, vision, audio, and video. This tutorial introduces the fundamental principles, methodologies, and practical applications of MLLM reasoning, with a particular emphasis on strengthening reasoning capabilities in multilingual and cross-domain settings. We further discuss the key challenges and limitations of current multimodal reasoning approaches, as well as future directions for advancing the field. By highlighting how MLLMs support enhanced reasoning and planning in cross-lingual and cross-domain contexts, this session aims to equip researchers and practitioners with the conceptual foundations and practical tools needed to effectively integrate MLLM reasoning into their work. Sarmistha Das 0001, Akash Ghosh, Sriparna Saha 0001, Koustava Goswami, K. J. Joseph |
ACM Multimedia | 2 |
| 2024 | CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in HealthcareabstractIn the era of modern healthcare, swiftly generating medical question summaries is crucial for informed and timely patient care. Despite the increasing complexity and volume of medical data, existing studies have focused solely on text-based summarization, neglecting the integration of visual information. Recognizing the untapped potential of combining textual queries with visual representations of medical conditions, we introduce the Multimodal Medical Question Summarization (MMQS) Dataset. This dataset, a major contribution of our work, pairs medical queries with visual aids, facilitating a richer and more nuanced understanding of patient needs. We also propose a framework, utilizing the power of Contrastive Language Image Pretraining(CLIP) and Large Language Models(LLMs), consisting of four modules that identify medical disorders, generate relevant context, filter medical concepts, and craft visually aware summaries. Our comprehensive framework harnesses the power of CLIP, a multimodal foundation model, and various general-purpose LLMs, comprising four main modules: the medical disorder identification module, the relevant context generation module, the context filtration module for distilling relevant medical concepts and knowledge, and finally, a general-purpose LLM to generate visually aware medical question summaries. Leveraging our MMQS dataset, we showcase how visual cues from images enhance the generation of medically nuanced summaries. This multimodal approach not only enhances the decision-making process in healthcare but also fosters a more nuanced understanding of patient queries, laying the groundwork for future research in personalized and responsive medical care. Disclaimer: The article features graphic medical imagery, a result of the subject's inherent requirements. Akash Ghosh, Arkadeep Acharya, Raghav Jain, Sriparna Saha 0001, Aman Chadha, Setu Sinha |
AAAI | 1 |
| 2024 | From Sights to Insights: Towards Summarization of Multimodal Clinical DocumentsabstractThe advancement of Artificial Intelligence is pivotal in reshaping healthcare, enhancing diagnostic precision, and facilitating personalized treatment strategies.One major challenge for healthcare professionals is quickly navigating through long clinical documents to provide timely and effective solutions.Doctors often struggle to draw quick conclusions from these extensive documents.To address this issue and save time for healthcare professionals, an effective summarization model is essential.Most current models assume the data is only textbased.However, patients often include images of their medical conditions in clinical documents.To effectively summarize these multimodal documents, we introduce EDI-Summ, an innovative Image-Guided Encoder-Decoder Model.This model uses modality-aware contextual attention on the encoder and an image cross-attention mechanism on the decoder, enhancing the BART base model to create detailed visual-guided summaries.We have tested our model extensively on three multimodal clinical benchmarks involving multimodal question and dialogue summarization tasks.Our analysis demonstrates that EDI-Summ outperforms state-of-the-art large language and vision-aware models in these summarization tasks. Akash Ghosh, Mohit Tomar, Abhisek Tiwari, Sriparna Saha 0001, Jatin Salve, Setu Sinha |
ACL (1) | 1 |
| 2024 | How Robust Are the QA Models for Hybrid Scientific Tabular Data? A Study Using Customized DatasetabstractQuestion-answering (QA) on hybrid scientific tabular and textual data deals with scientific information, and relies on complex numerical reasoning. In recent years, while tabular QA has seen rapid progress, understanding their robustness on scientific information is lacking due to absence of any benchmark dataset. To investigate the robustness of the existing state-of-the-art QA models on scientific hybrid tabular data, we propose a new dataset, “SciTabQA”, consisting of 822 question-answer pairs from scientific tables and their descriptions. With the help of this dataset, we assess the state-of-the-art Tabular QA models based on their ability (i) to use heterogeneous information requiring both structured data (table) and unstructured data (text) and (ii) to perform complex scientific reasoning tasks. In essence, we check the capability of the models to interpret scientific tables and text. Our experiments show that “SciTabQA” is an innovative dataset to study question-answering over scientific heterogeneous data. We benchmark three state-of-the-art Tabular QA models, and find that the best F1 score is only 0.462. Akash Ghosh, Venkata Sahith Bathini, Niloy Ganguly, Pawan Goyal 0002, Mayank Singh 0001 |
LREC/COLING | 1 |
| 2024 | MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical Queries
Akash Ghosh, Arkadeep Acharya, Prince Jha, Sriparna Saha 0001, Aniket Gaudgaul, Rajdeep Majumdar, Aman Chadha, Raghav Jain, Setu Sinha, Shivani Agarwal 0005 |
ECIR (5) | 1 |
| 2024 | Parameter-Efficient Instruction Tuning of Large Language Models For Extreme Financial Numeral LabellingabstractSubhendu Khatuya, Rajdeep Mukherjee, Akash Ghosh, Manjunath Hegde, Koustuv Dasgupta, Niloy Ganguly, Saptarshi Ghosh, Pawan Goyal. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Subhendu Khatuya, Rajdeep Mukherjee, Akash Ghosh, Manjunath Hegde, Koustuv Dasgupta, Niloy Ganguly, Saptarshi Ghosh 0001, Pawan Goyal 0002 |
NAACL-HLT | 3 |
| 2018 | Semantic Clone Detection: Can Source Code Comments Help?abstractProgrammers reuse code to increase their productivity, which leads to large fragments of duplicate or near-duplicate code in the code base. The current code clone detection techniques for finding semantic clones utilize Program Dependency Graphs (PDG), which are expensive and resource-intensive. PDG and other clone detection techniques utilize code and have completely ignored the comments - due to ambiguity of English language, but in terms of program comprehension, comments carry the important domain knowledge. We empirically evaluated the accuracy of detecting clones with both code and comments on a JHotDraw package. Results show that detecting code clones in the presence of comments, Latent Dirichlet Allocation (LDA), gave 84% precision and 94% recall, while in the presence of a PDG, using GRAPLE, we got 55% precision and 29% recall. These results indicate that comments can be used to find semantic clones. We recommend utilizing comments with LDA to find clones at the file level and code with PDG for finding clones at the function level. These findings necessitate a need to reexamine the assumptions regarding semantic clone detection techniques. Akash Ghosh, Sandeep Kaur Kuttal |
VL/HCC | 1 |