EDBT 2026 Demo / reviewers in the wild / expert
Madhusudan Ghosh
dblp:311/4546
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-8330-2703ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mask-to-Correct⁺: Leveraging Retriever Diversity for Masking-guided Faithful Fact CorrectionabstractThe rapid spread of misinformation on social media highlights the need for robust, automated fact correction frameworks.However, existing works rely on supervised learning from manually annotated claim-evidence pairs, which are scarce and prone to biases, limiting their generalization across domains.Moreover, these methods overlook semantic faithfulness in their correction process.To address these challenges, we propose Mask-to-Correct (M 2 C), a training-free, inference-only Retrieval Augmented Generation (RAG) based framework that leverages diversity-aware masking to identify erroneous spans of claims and evaluate the faithfulness of corrections using retrieved evidence.However, the effectiveness of RAG heavily depends on the choice of retriever, which may vary across queries.To mitigate this, we further introduce M 2 C + , an ensemblebased framework that combines corrections across multiple rankers to reduce retrieval bias and improve robustness.Extensive experiments on the benchmark datasets demonstrate that our proposed frameworks consistently outperform all baselines, achieving up to 14% improvement in SARI scores, without using gold evidence. Payel Santra, Lavisha Sharma, Madhusudan Ghosh, Partha Basuchowdhuri |
ACL (1) | 3 |
| 2025 | HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and RankersabstractLeveraging both labeled (input-output associations) and unlabeled data (wider contextual grounding) may provide complementary benefits in retrieval augmented generation (RAG). However, effectively combining evidence from these heterogeneous sources is challenging as the respective similarity scores are not inter-comparable. Additionally, aggregating beliefs from the outputs of multiple rankers can improve the effectiveness of RAG. Our proposed method first aggregates the top-documents from a number of IR models using a standard rank fusion technique for each source (labeled and unlabeled). Next, we standardize the retrieval score distributions within each source by applying z-score transformation before merging the top-retrieved documents from the two sources. We evaluate our approach on the fact verification task, demonstrating that it consistently improves over the best-performing individual ranker or source and also shows better out-of-domain generalization. Payel Santra, Madhusudan Ghosh, Debasis Ganguly, Partha Basuchowdhuri, Sudip Kumar Naskar |
CIKM | 2 |
| 2025 | In-Context Learning as an Effective Estimator of Functional Correctness of LLM-Generated CodeabstractWhen applying LLM-based code generation to software development projects that follow a feature-driven or rapid application development approach, it becomes necessary to estimate the functional correctness of the generated code in the absence of test cases. Just as a user selects a relevant document from a ranked list of retrieved ones, a software generation workflow requires a developer to choose (and potentially refine) a generated solution from a ranked list of alternative solutions, ordered by their posterior likelihoods. This implies that estimating the quality of a ranked list - akin to estimating ''relevance'' for query performance prediction (QPP) in IR - is also crucial for generative software development, where quality is defined in terms of ''functional correctness''. In this paper, we propose an in-context learning (ICL) based approach for code quality estimation. Our findings demonstrate that providing few-shot examples of functionally correct code from a training set enhances the performance of existing QPP approaches as well as a zero-shot-based approach for code quality estimation. Susmita Das 0002, Madhusudan Ghosh, Priyanka Swami, Debasis Ganguly, Gül Çalikli |
SIGIR | 2 |
| 2024 | "The Absence of Evidence is Not the Evidence of Absence": Fact Verification via Information Retrieval-Based In-Context Learning
Payel Santra, Madhusudan Ghosh, Debasis Ganguly, Partha Basuchowdhuri, Sudip Kumar Naskar |
DaWaK | 2 |
| 2023 | Extracting Methodology Components from AI Research Papers: A Data-driven Factored Sequence Labeling ApproachabstractExtraction of methodology component names from scientific articles is a challenging task due to the diversified contexts around the occurrences of these entities, and the different levels of granularity and containment relationships exhibited by these entities. We hypothesize that standard sequence labeling approaches may not adequately model the dependence of methodology name mentions with their contexts, due to the problems of their large, fast evolving, and domain-specific vocabulary. As a solution, we propose a factored approach, where the mention-context dependencies are represented in a more fine-grained manner, thus allowing the model parameters to better adjust to the different characteristic patterns inherent within the data. In particular, we experiment with two variants of this factored approach - one that uses the per-entity category information derived from an ontology, and the other that makes use of the topology of the sentence embedding space to infer a category for each entity constituting that sentence. We demonstrate that both these factored variants of SciBERT outperform their non-factored counterpart, a state-of-the-art model for scientific concept extraction. Madhusudan Ghosh, Debasis Ganguly, Partha Basuchowdhuri, Sudip Kumar Naskar |
CIKM | 1 |
| 2022 | DeepGLSTM: Deep Graph Convolutional Network and LSTM based approach for predicting drug-target binding affinityabstractDevelopment of new drugs is an expensive and time-consuming process. Due to the world-wide SARS-CoV-2 outbreak, it is essential that new drugs for SARS-CoV-2 are developed as soon as possible. Drug repurposing techniques can reduce the time span needed to develop new drugs by probing the list of existing FDA-approved drugs and their properties to reuse them for combating the new disease. We propose a novel architecture DeepGLSTM, which is a Graph Convolutional network and LSTM based method that predicts binding affinity values between the FDA-approved drugs and the viral proteins of SARS-CoV-2. Our proposed model has been trained on Davis, KIBA (Kinase Inhibitor Bioactivity), DTC (Drug Target Commons), Metz, ToxCast and STITCH datasets. We use our novel architecture to predict a Combined Score (calculated using Davis and KIBA score) of 2,304 FDA-approved drugs against 5 viral proteins. On the basis of the Combined Score, we prepare a list of the top-18 drugs with the highest binding affinity for 5 viral proteins present in SARS-CoV-2. Subsequently, this list may be used for the creation of new useful drugs. Shrimon Mukherjee, Madhusudan Ghosh, Partha Basuchowdhuri |
SDM | 2 |