EDBT 2026 Demo / reviewers in the wild / expert
Sikun Guo
dblp:317/4714
· DBLP profile ↗
5ranked-venue papers in the field
3as first author
5since 2021 · last 2026
0000-0002-4764-3359ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 4 (2 first)Big Data, Cloud & Distributed Data Systems · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated biomedical hypothesis generation with time-aware hypergraph contrastive learningabstractAbstract Research in scientific domains now generates more than a million articles annually, overwhelming researchers and hindering discovery. This surge has sparked interest in biomedical hypothesis generation (HG), which aims to uncover implicit patterns among biomedical concepts. Most existing methods focus on pairwise link prediction, overlooking the complex, multi-concept relationships underlying many breakthroughs. We introduce HyHG , a temporal Hy pergraph contrastive learning framework for biomedical H ypothesis G eneration, which redefines hypotheses as hyperedges—sets of co-mentioned concepts in an article. By representing articles as hyperedges and organizing them into a temporal hypergraph, HyHG captures the evolution of scientific ideas over time. A transformer-based architecture learns from historical hyperedge sequences to predict future hyperedges—sets of concepts likely to co-occur in the future literature. To distinguish genuine hypotheses from misleading ones, HyHG employs a time-anchored contrastive loss and hard negative sampling based on minimal edits to real hyperedges. We demonstrate that HyHG achieves state-of-the-art performance on three biomedical datasets. Our code and data are available at: https://github.com/amir-hassan25/Temporal-Hypergraph-Contrastive-Learning. Amir Hassan Shariatmadari, Sikun Guo, Nathan C. Sheffield, Aidong Zhang 0001, Kishlay Jha |
Knowl. Inf. Syst. | 2 |
| 2025 | HyHG: A Temporal Hypergraph Contrastive Learning Framework for Biomedical Hypothesis GenerationabstractBiomedical research now generates more than a million articles annually, overwhelming researchers and hindering discovery. This surge has sparked interest in biomedical hypothesis generation (HG), which aims to uncover implicit patterns among biomedical concepts. Most existing methods focus on pairwise link prediction, overlooking the complex, multi-concept relationships underlying many breakthroughs. We introduce HyHG, a temporal Hypergraph contrastive learning framework for biomedical Hypothesis Generation, which redefines hypotheses as hyperedges–sets of co-mentioned concepts in an article. By representing articles as hyperedges and organizing them into a temporal hypergraph, HyHG captures the evolution of scientific ideas over time. A transformer-based architecture learns from historical hyperedge sequences to predict future hyperedges–sets of concepts likely to co-occur in future literature. To distinguish genuine hypotheses from misleading ones, HyHG employs a timeanchored contrastive loss and hard negative sampling based on minimal edits to real hyperedges. We demonstrate state-of-the-art performance on three biomedical datasets. Our code and data are available at: https://github.com/amirhassan25/Temporal-Hypergraph-Contrastive-Learning. Amir Hassan Shariatmadari, Sikun Guo, Nathan C. Sheffield, Aidong Zhang 0001, Kishlay Jha |
ICDM | 2 |
| 2025 | IdeaBench: Benchmarking Large Language Models for Research Idea GenerationabstractLarge Language Models (LLMs) have revolutionized interactions between human and artificial intelligence (AI) systems, demonstrating state-of-the-art performance across various domains, including scientific discovery and hypothesis generation. However, the absence of a comprehensive and systematic evaluation framework for LLM-driven research idea generation hinders a rigorous understanding of their strengths and limitations. To address this gap, we propose IdeaBench, a benchmark system that provides a structured dataset and evaluation framework for standardizing the assessment of research idea generation by LLMs. Our dataset comprises titles and abstracts from 2,374 influential papers across eight research domains, along with their 29,408 referenced works, creating a context-rich environment that mirrors human researchers' ideation processes. By profiling LLMs as domain-specific researchers and grounding them in similar contextual constraints, we directly leverage the models' knowledge learned from the pre-training stage to generate new research ideas. To systematically evaluate LLMs' research ideation capability and approximate human assessment, we propose a reference-based metric that aligns with human judgment to quantify idea quality with the assistance of LLMs. Through this evaluation, we find that while LLMs excel at generating novel ideas, they may struggle with generating feasible ideas. IdeaBench serves as a critical resource for benchmarking and comparing LLMs, ultimately advancing research on AI's role in automating scientific discovery. Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang, Myles Kim, Corey M. Williams, Stefan Bekiranov, Aidong Zhang 0001 |
KDD (2) | 1 |
| 2025 | Optimizing External and Internal Knowledge of Foundation Models for Scientific DiscoveryabstractIn the emerging landscape of AI-driven scientific discovery, foundation models hold significant promise for enhancing research ideation and overall scientific advancement. This paper explores a future where foundation models should be able to effectively utilize both external and internal knowledge sources to maximize their role in scientific discovery. The core challenge lies in optimizing two knowledge types: external knowledge, drawn from diverse data sources, and internal knowledge, the parametric understanding acquired during training. We propose a dual-framework solution for this optimization, including X-augmented generation and in-context X learning. X-augmented generation approaches, such as retrieval-augmented generation, knowledge graph-augmented generation, and third-party tool integration, enhance external knowledge processing. In-context X learning methods, including in-context adversarial learning and in-context reinforcement learning, improve models’ internal knowledge adaptation and utility for scientific tasks. We aim to inspire the research community by proposing a bold pathway toward leveraging foundation models as active participants in scientific discovery, tackling the inherent complexity of optimizing vast, multimodal knowledge sources. By addressing this challenge, we envision a future where foundation models catalyze breakthroughs across disciplines, ultimately leading to a more dynamic, collaborative, and insight-driven scientific process. Sikun Guo, Guangzhi Xiong, Aidong Zhang 0001 |
SDM | 1 |
| 2024 | Embracing Foundation Models for Advancing Scientific DiscoveryabstractMachine learning foundation models, particularly large language models (LLMs) such as GPT-4o, have revolutionized traditional applications in computer vision and natural language processing, marking a significant shift in recent years. Building on these advancements, recent efforts have explored the potential of foundation models in hypothesis generation, highlighting their possibility in aiding human researchers in scientific discovery. In this paper, we envision a future where academia increasingly integrates foundation models to accelerate and enhance the process of scientific discovery. Motivated by potential application scenarios of foundation models in scientific research, our vision is anchored in a central question: How can we accelerate scientific discovery with the aid of foundation models? To address this overarching question, we raise two key challenges that need to be addressed: (1) how to effectively harness the parametric knowledge embedded in foundation models to propel scientific discovery? and (2) how to develop rigorous yet scalable methods to evaluate the effectiveness of foundation models in supporting scientific research? To tackle these two challenges, we propose our approaches, termed knowledge-grounded Chain-of-Idea (KG-CoI) hypothesis generation and IdeaBench - Benchmarking LLM hypothesis generators in a customizable manner. Through addressing these challenges, we outline our vision in hope to inspire new ideas and innovations in harnessing foundation models for advancing scientific discovery, paving the way for a new era of research collaboration between humans and artificial intelligence. Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Aidong Zhang 0001 |
IEEE Big Data | 1 |