Wanhao Liu

dblp:390/2445 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 54% Planning, search and constraint satisfaction · 18% Optimization for machine learning · 18%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
combinatorial optimization
0.912025
MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
hierarchical search
0.912025
MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model
0.912025
MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search · NeurIPS 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses · ICLR 2025
Natural language and speech › Language models and text generation › text generation
scientific hypothesis generation
0.912025
MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search · NeurIPS 2025
Computational science and engineering
computational chemistry
0.912025
MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses · ICLR 2025
Machine learning › Kernel, tree and ensemble methods
model ensemble
0.312025
MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
scientific discovery
0.312025
MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses · ICLR 2025

Methods — techniques the papers use, named apart from their topics

retrieval · 1.7multi-agent framework · 1.7reward landscape shaping · 0.9hierarchical search · 0.9ensemble · 0.9direct preference optimization · 0.9
YearPublicationVenuePosition
2025 MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
abstract
Scientific discovery contributes largely to the prosperity of human society, and recent progress shows that LLMs could potentially catalyst the process. However, it is still unclear whether LLMs can discover novel and valid hypotheses in chemistry. In this work, we investigate this main research question: whether LLMs can automatically discover novel and valid chemistry research hypotheses, given only a research question? With extensive discussions with chemistry experts, we adopt the assumption that a majority of chemistry hypotheses can be resulted from a research background question and several inspirations. With this key insight, we break the main question into three smaller fundamental questions. In brief, they are: (1) given a background question, whether LLMs can retrieve good inspirations; (2) with background and inspirations, whether LLMs can lead to hypothesis; and (3) whether LLMs can identify good hypotheses to rank them higher. To investigate these questions, we construct a benchmark consisting of 51 chemistry papers published in Nature or a similar level in 2024 (all papers are only available online since 2024). Every paper is divided by chemistry PhD students into three components: background, inspirations, and hypothesis. The goal is to rediscover the hypothesis given only the background and a large chemistry literature corpus consisting the ground truth inspiration papers, with LLMs trained with data up to 2023. We also develop an LLM-based multi-agent framework that leverages the assumption, consisting of three stages reflecting the more smaller questions. The proposed method can rediscover many hypotheses with very high similarity with the ground truth ones, covering the main innovations.
Zonglin Yang 0001, Wanhao Liu, Ben Gao, Tong Xie, Wanli Ouyang, Soujanya Poria, Erik Cambria, Dongzhan Zhou
ICLR2
2025 MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
abstract
Large language models (LLMs) have shown promise in automating scientific hypothesis generation, yet existing approaches primarily yield coarse-grained hypotheses lacking critical methodological and experimental details. We introduce and formally define the new task of fine-grained scientific hypothesis discovery, which entails generating detailed, experimentally actionable hypotheses from coarse initial research directions. We frame this as a combinatorial optimization problem and investigate the upper limits of LLMs' capacity to solve it when maximally leveraged. Specifically, we explore four foundational questions: (1) how to best harness an LLM's internal heuristics to formulate the fine-grained hypothesis it itself would judge as the most promising among all the possible hypotheses it might generate, based on its own internal scoring-thus defining a latent reward landscape over the hypothesis space; (2) whether such LLM-judged better hypotheses exhibit stronger alignment with ground-truth hypotheses; (3) whether shaping the reward landscape using an ensemble of diverse LLMs of similar capacity yields better outcomes than defining it with repeated instances of the strongest LLM among them; and (4) whether an ensemble of identical LLMs provides a more reliable reward landscape than a single LLM. To address these questions, we propose a hierarchical search method that incrementally proposes and integrates details into the hypothesis, progressing from general concepts to specific experimental configurations. We show that this hierarchical process smooths the reward landscape and enables more effective optimization. Empirical evaluations on a new benchmark of expert-annotated fine-grained hypotheses from recent literature show that our method consistently outperforms strong baselines.
Zonglin Yang 0001, Wanhao Liu, Ben Gao, Wei Li 0076, Tong Xie, Lidong Bing, Wanli Ouyang, Erik Cambria, Dongzhan Zhou
NeurIPS2