VLDB 2026 Research / reviewers in the wild / expert
Ghanshyam Verma
dblp:232/4684
· DBLP profile ↗
5ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0002-1394-6386ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Question answering and dialogue systems · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Energy systems and smart grids · 100% |
Topics — the 2 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › question generation
question-answer pair generation |
1.0 | 1 | 2026 | EvalQAG: A Framework for Automatic Complex QA Generation and a Benchmark QA Dataset for Policy Documents · AAAI 2026 |
Natural language and speech › Question answering and dialogue systems
question generation |
1.0 | 1 | 2026 | EvalQAG: A Framework for Automatic Complex QA Generation and a Benchmark QA Dataset for Policy Documents · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
structured prompting · 2.0retrieval-augmented generation · 2.0large language model · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EvalQAG: A Framework for Automatic Complex QA Generation and a Benchmark QA Dataset for Policy DocumentsabstractAccelerating research in renewable energy policy is critical for addressing climate change and enabling informed decision-making. Question answering (QA) over public policy documents presents unique challenges due to their legal structure, conditional dependencies, and domain-specific vocabulary. In this paper, we introduce EvalQAG, a framework for generating high-quality QA pairs from renewable energy policy documents. EvalQAG combines structured prompts, retrieval-augmented inputs, and multi-stage evaluation using large language models (LLMs) to support accurate and diverse QA generation. Using this framework, we construct REPolicyQA, a domain-specific QA dataset comprising approximately 160,000 QA pairs from over 1,000 U.S. renewable energy policy documents. The dataset covers five policy-relevant question types: Yes/No, Yes/No with Conditions, Factual, Legal Obligation, and Descriptive, which capture a wide range of reasoning patterns grounded in regulatory texts. We evaluate multiple QA models and uncover significant performance gaps, particularly in legal reasoning and conditional inference, highlighting major shortcomings in current systems. Our results establish EvalQAG as a generalizable QA generation pipeline for policy texts and position REPolicyQA as a new benchmark for advancing QA research in policy and regulatory domains. We believe this work can foster impactful research in the renewable energy sector, particularly by enabling more robust and explainable QA systems for legal and condition-heavy regulatory documents. Kirtan Brijeshbhai Soni, Krish Rupapara, Arpit Rana, Ghanshyam Verma, Paul Buitelaar |
AAAI | 4 |
| 2025 | Empowering Recommender Systems using Automatically Generated Knowledge Graphs and Reinforcement LearningabstractPersonalized recommender systems play a crucial role in direct marketing, particularly in financial services, where delivering relevant content can enhance customer engagement and promote informed decision-making. This study explores interpretable knowledge graph (KG)-based recommender systems by proposing two distinct approaches for personalized article recommendations within a multinational financial services firm. The first approach leverages Reinforcement Learning (RL) to traverse a KG constructed from both structured (tabular) and unstructured (textual) data, enabling interpretability through Path Directed Reasoning (PDR). The second approach employs the XGBoost algorithm, with post-hoc explainability techniques such as SHAP and ELI5 to enhance transparency. By integrating machine learning with automatically generated KGs, our methods not only improve recommendation accuracy but also provide interpretable insights, facilitating more informed decision-making in customer relationship management. Ghanshyam Verma, Simanta Sarkar, Devishree Pillai, John P. McCrae, János A Perge, Shovon Sengupta, Paul Buitelaar |
LDK | 1 |
| 2024 | Enabling personalised disease diagnosis by combining a patient's time-specific gene expression profile with a biomedical knowledge baseabstractBACKGROUND: Recent developments in the domain of biomedical knowledge bases (KBs) open up new ways to exploit biomedical knowledge that is available in the form of KBs. Significant work has been done in the direction of biomedical KB creation and KB completion, specifically, those having gene-disease associations and other related entities. However, the use of such biomedical KBs in combination with patients' temporal clinical data still largely remains unexplored, but has the potential to immensely benefit medical diagnostic decision support systems. RESULTS: We propose two new algorithms, LOADDx and SCADDx, to combine a patient's gene expression data with gene-disease association and other related information available in the form of a KB, to assist personalized disease diagnosis. We have tested both of the algorithms on two KBs and on four real-world gene expression datasets of respiratory viral infection caused by Influenza-like viruses of 19 subtypes. We also compare the performance of proposed algorithms with that of five existing state-of-the-art machine learning algorithms (k-NN, Random Forest, XGBoost, Linear SVM, and SVM with RBF Kernel) using two validation approaches: LOOCV and a single internal validation set. Both SCADDx and LOADDx outperform the existing algorithms when evaluated with both validation approaches. SCADDx is able to detect infections with up to 100% accuracy in the cases of Datasets 2 and 3. Overall, SCADDx and LOADDx are able to detect an infection within 72 h of infection with 91.38% and 92.66% average accuracy respectively considering all four datasets, whereas XGBoost, which performed best among the existing machine learning algorithms, can detect the infection with only 86.43% accuracy on an average. CONCLUSIONS: We demonstrate how our novel idea of using the most and least differentially expressed genes in combination with a KB can enable identification of the diseases that a patient is most likely to have at a particular time, from a KB with thousands of diseases. Moreover, the proposed algorithms can provide a short ranked list of the most likely diseases for each patient along with their most affected genes, and other entities linked with them in the KB, which can support health care professionals in their decision-making. Ghanshyam Verma, Dietrich Rebholz-Schuhmann, Michael G. Madden |
BMC Bioinform. | 1 |
| 2019 | Ranked MSD: A New Feature Ranking and Feature Selection Approach for Biomarker Identification
Ghanshyam Verma, Alokkumar Jha, Dietrich Rebholz-Schuhmann, Michael G. Madden |
CD-MAKE | 1 |
| 2018 | Deep Convolution Neural Network Model to Predict Relapse in Breast CancerabstractA mishap in anti-cancer drug distribution is critical in breast cancer patients due to poor prediction model to identify the treatment regime in ER+ve and ER-ve (Estrogen Receptor (ER)) patients. The traditional method for the prediction depends on the change in expression across the normal-disease pair. However, it certainly misses the multidimensional aspect and underlying cause of relapse, such as various mutations, drug dosage side effects, methylation, etc. In this paper, we have developed a multi-layer neural network model to classify multidimensional genomics data into their similar annotation group. Further, we used this multi-layer cancer genomics perceptron for annotating differentially expressed genes (DEGs) to predict relapse based on ER status in breast cancer. This approach provides multivariate identification of genes, not just by differential expression, but, cause-effect of disease status due to drug overdosage and genomics-driven drug balancing method. The multi-layered neural network model, where each layer defines the relationship of similar databases with multidimensional knowledge. We illustrate that the use of multilayer knowledge graph with gene expression data for training the deep convolution neural network stratify the patient relapse and drug dosage along with underlying molecular properties. Alokkumar Jha, Ghanshyam Verma, Yasar Khan, Qaiser Mehmood 0001, Dietrich Rebholz-Schuhmann, Ratnesh Sahay |
ICMLA | 2 |