VLDB 2026 Research / reviewers in the wild / expert
Gollam Rabby
dblp:271/5642
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-1212-0101ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AIssistant: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
Sasi Kiran Gaddipati, Farhana Keya, Gollam Rabby, Sören Auer |
PAKDD (4) | 3 |
| 2026 | Iterative hypothesis generation for scientific discovery with Monte Carlo self-refining treesabstractScientific hypothesis generation is central to discovery, enabling researchers to propose ideas, design experiments, and validate knowledge. Yet, producing hypotheses that are both novel and empirically grounded remains challenging. Traditional methods rely heavily on human intuition, while automated approaches often lack scientific rigor and meaningful validation. This work presents the Monte Carlo Self-Refine Tree (MC-NEST), a general-purpose framework that automates hypothesis generation by integrating Monte Carlo Tree Search (MCTS) with adaptive sampling and iterative self-evaluation. MC-NEST treats hypothesis generation as a structured search problem, dynamically balancing exploration and refinement through strategy selection. We evaluate MC-NEST on a benchmark spanning biomedicine, social science, and computer science. Experimental results show that MC-NEST consistently outperforms state-of-the-art prompt-based baselines across four human-annotated criteria: novelty, clarity, significance, and verifiability. Specifically, it achieves average scores of 2.65, 2.74, and 2.80 in social science, computer science, and biomedicine, respectively, outperforming baseline scores of 2.36, 2.51, and 2.52. These findings highlight MC-NEST’s ability to generate interpretable, scientifically meaningful hypotheses across domains. Furthermore, MC-NEST is designed to support human-AI collaboration through its interpretable tree structure, enabling future integration where domain experts can guide exploration and validate hypotheses transparently and reproducibly. By integrating strategic search, self-refinement, and adaptive validation, MC-NEST offers a scalable and effective path toward responsible, automated scientific discovery. Gollam Rabby, Diyana Muhammed, Prasenjit Mitra 0001, Sören Auer |
Inf. Sci. | 1 |
| 2026 | SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence EmbeddingsabstractAbstract Generating context-aware, high-quality, and innovative scientific ideas remains a central challenge in AI-supported research. We introduce SCI-IDEA , a two-stage framework combining large language models (LLMs) with a specialised Aha-Moment Detection module for iterative idea refinement. The first stage extracts structured facets: objectives, methodology, evaluation, and future work from previous publications, compressing each paper to $${\sim }$$ 200 tokens to enable scalable processing of researcher profiles that would otherwise exceed frontier LLM context windows. The second stage integrates these facets with a token-level embedding approach to identify novelty and context gaps, indicating candidate ideas as Aha moments when they surpass predefined thresholds for both novelty and surprise. We evaluate SCI-IDEA across 100 computer-science researcher profiles, 4 LLMs (GPT-4o, GPT−4.5, DeepSeek-32B, DeepSeek-70B), three embedding strategies, and 5 prompting configurations via a hybrid protocol of automated LLM-as-judge scoring with GPT−4.1 and a human evaluation with 15 PhD-level domain experts. These two evaluation signals diverge substantially: LLM-based scores exceed expert ratings by 3–4 points on a 10-point scale, with near-zero inter-rater correlation ( $$r = 0.02$$ –0.17, $$p > 0.05$$ ), indicating that automated scores reflect relative architectural comparisons rather than human-verified measures of absolute idea quality. Within the LLM-evaluated setting, SCI-IDEA with token-level embeddings achieves a mean quality score of 6.97, a modest but constant improvement over the LLM-only baseline ( $$\Delta = +0.20$$ points; $$p < 0.05$$ ) and driven primarily by improved feasibility scores ( $$+0.21$$ points, $$p < 0.001$$ ). Expert evaluators independently confirmed this trend while assigning lower absolute scores. These findings suggest an architectural advantage of facet-based context modelling and token-level novelty detection, though their practical importance deserves validation through longitudinal studies with larger domain experts. This work is validated in the computer science domain, but its application to other fields requires further investigation. Farhana Keya, Gollam Rabby, Sören Auer, Sahar Vahdati, Prasenjit Mitra 0001, Mohamad Yaser Jaradeh |
Mach. Learn. | 2 |
| 2025 | EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion RecognitionabstractEffective human-AI interaction relies on AI's ability to accurately perceive and interpret human emotions. Current benchmarks for vision and vision-language models are severely limited, offering a narrow emotional spectrum that overlooks nuanced states (e.g., bitterness, intoxication) and fails to distinguish subtle differences between related feelings (e.g., shame vs. embarrassment). Existing datasets also often use uncontrolled imagery with occluded faces and lack demographic diversity, risking significant bias. To address these critical gaps, we introduce EmoNet Face, a comprehensive benchmark suite. EmoNet Face features: (1) A novel 40-category emotion taxonomy, meticulously derived from foundational research to capture finer details of human emotional experiences. (2) Three large-scale, AI-generated datasets (EmoNet HQ, Binary, and Big) with explicit, full-face expressions and controlled demographic balance across ethnicity, age, and gender. (3) Rigorous, multi-expert annotations for training and high-fidelity evaluation. (4) We build Empathic Insight Face, a model achieving human-expert-level performance on our benchmark. The publicly released EmoNet Face suite—taxonomy, datasets, and model—provides a robust foundation for developing and evaluating AI systems with a deeper understanding of human emotions. Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby, Maurice Kraus, Felix Friedrich, Huu Nguyen, Krishna Kalyan, Kourosh Nadi, Kristian Kersting, Sören Auer |
NeurIPS | 3 |
| 2023 | Pruning and re-ranking the frequent patterns in knowledge graph profiling using machine learning
Gollam Rabby, Farhana Keya, Vojtech Svátek, Blerina Spahiu |
LDK | 1 |
| 2023 | Multi-class classification of COVID-19 documents using machine learning algorithms
Gollam Rabby, Petr Berka |
J. Intell. Inf. Syst. | 1 |
| 2020 | Ontologies Supporting Research-Related Information Foraging Using Knowledge Graphs: Literature Survey and Holistic Model Mapping
Viet Bach Nguyen, Vojtech Svátek, Gollam Rabby, Óscar Corcho |
EKAW | 3 |