VLDB 2026 Research / reviewers in the wild / expert
Yanjun Shao
dblp:253/6512
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Knowledge representation and reasoning · 39% Multi-agent systems · 22% Deep learning architectures and training · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
LLM-based multi-agent planning |
1.0 | 1 | 2026 | SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data Science · ACL (1) 2026 |
Data mining
automated data science |
1.0 | 1 | 2026 | SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data Science · ACL (1) 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › qualitative reasoning
chemical reasoning |
0.9 | 1 | 2025 | ChemAgent: Self-updating Memories in Large Language Models Improves Chemical Reasoning · ICLR 2025 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.9 | 1 | 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation Models · ICLR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › case-based reasoning
memory-based reasoning |
0.9 | 1 | 2025 | ChemAgent: Self-updating Memories in Large Language Models Improves Chemical Reasoning · ICLR 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | ChemAgent: Self-updating Memories in Large Language Models Improves Chemical Reasoning · ICLR 2025 |
Bioinformatics and computational biology › sequence analysis › sequence modeling
DNA sequence modeling |
0.9 | 1 | 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation Models · ICLR 2025 |
Bioinformatics and computational biology
genomics |
0.9 | 1 | 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation Models · ICLR 2025 |
Program synthesis and code generation
code generation with language models |
0.9 | 1 | 2025 | OpenHands: An Open Platform for AI Software Developers as Generalist Agents · ICLR 2025 |
Bioinformatics and computational biology › molecular informatics
cheminformatics |
0.3 | 1 | 2025 | ChemAgent: Self-updating Memories in Large Language Models Improves Chemical Reasoning · ICLR 2025 |
Bioinformatics and computational biology
drug discovery |
0.3 | 1 | 2025 | ChemAgent: Self-updating Memories in Large Language Models Improves Chemical Reasoning · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 3.7hyperparameter tuning · 2.0ensemble selection · 2.0transformer · 1.7state space model · 1.7retrieval-augmented generation · 1.7large language model agent · 1.7gated convolution · 1.7dilated convolution · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data ScienceabstractLarge Language Models (LLMs) have enabled dynamic reasoning in automated data analytics, yet recent multi-agent systems remain limited by rigid, single-path workflows that restrict strategic exploration and often lead to suboptimal outcomes.To overcome these limitations, we propose SPIO (Sequential Plan Integration and Optimization), a framework that replaces rigid workflows with adaptive, multi-path planning across four core modules: data preprocessing, feature engineering, model selection, and hyperparameter tuning.In each module, specialized agents generate diverse candidate strategies, which are cascaded and refined by an optimization agent.SPIO offers two operating modes: SPIO-S for selecting a single optimal pipeline, and SPIO-E for ensembling top-k pipelines to maximize robustness.Extensive evaluations on Kaggle and OpenML benchmarks show that SPIO consistently outperforms state-of-the-art baselines, achieving an average performance gain of 5.6%.By explicitly exploring and integrating multiple solution paths, SPIO delivers a more flexible, accurate, and reliable foundation for automated data science.* denotes equal contribution.† denotes corresponding author(s). Wonduk Seo, Juhyeon Lee, Yanjun Shao, Qingshan Zhou, Yi Bu 0001 |
ACL (1) | 3 |
| 2025 | OpenHands: An Open Platform for AI Software Developers as Generalist AgentsabstractSoftware is one of the most powerful tools that we humans have at our disposal; it allows a skilled programmer to interact with the world in complex and profound ways. At the same time, thanks to improvements in large language models (LLMs), there has also been a rapid development in AI agents that interact with and effect change in their surrounding environments. In this paper, we introduce OpenHands, a platform for the development of powerful and flexible AI agents that interact with the world in similar ways to a human developer: by writing code, interacting with a command line, and browsing the web. We describe how the platform allows for the implementation of new agents, utilization of various LLMs, safe interaction with sandboxed environments for code execution, and incorporation of evaluation benchmarks. Based on our currently incorporated benchmarks, we perform an evaluation of agents over 13 challenging tasks, including software engineering (e.g., SWE-Bench) and web browsing (e.g., WebArena), amongst others. Released under the permissive MIT license, OpenHands is a community project spanning academia and industry with more than 2K contributions from over 186 contributors in less than six months of development, and will improve going forward. Xingyao Wang 0002, Boxuan Li, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Yueqi Song, Bowen Li 0002, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang 0002, Binyuan Hui, Junyang Lin |
ICLR | 16 |
| 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation ModelsabstractIn recent years, A variety of methods based on Transformer and state space model (SSM) architectures have been proposed, advancing foundational DNA language models.
However, there is a lack of comparison between these recent approaches and the classical architecture—convolutional networks (CNNs)—on foundation model benchmarks.
This raises the question: are CNNs truly being surpassed by these recent approaches based on transformer and SSM architectures? In this paper, we develop a simple but well-designed CNN-based method, termed ConvNova. ConvNova identifies and proposes three effective designs: 1) dilated convolutions, 2) gated convolutions, and 3) a dual-branch framework for gating mechanisms.
Through extensive empirical experiments, we demonstrate that ConvNova significantly outperforms recent methods on more than half of the tasks across several foundation model benchmarks. For example, in histone-related tasks, ConvNova exceeds the second-best method by an average of 5.8\%, while generally utilizing fewer parameters and enabling faster computation. In addition, the experiments observed findings that may be related to biological characteristics. This indicates that CNNs are still a strong competitor compared to Transformers and SSMs. We anticipate that this work will spark renewed interest in CNN-based methods for DNA foundation models. Yu Bo, Weian Mao, Yanjun Shao, Weiqiang Bai, Peng Ye 0006, Xinzhu Ma, Hao Chen 0041, Chunhua Shen |
ICLR | 3 |
| 2025 | ChemAgent: Self-updating Memories in Large Language Models Improves Chemical ReasoningabstractChemical reasoning usually involves complex, multi-step processes that demand precise calculations, where even minor errors can lead to cascading failures. Furthermore, large language models (LLMs) encounter difficulties handling domain-specific formulas, executing reasoning steps accurately, and integrating code ef- effectively when tackling chemical reasoning tasks. To address these challenges, we present ChemAgent, a novel framework designed to improve the performance of LLMs through a dynamic, self-updating library. This library is developed by decomposing chemical tasks into sub-tasks and compiling these sub-tasks into a structured collection that can be referenced for future queries. Then, when presented with a new problem, ChemAgent retrieves and refines pertinent information from the library, which we call memory, facilitating effective task decomposition and the generation of solutions. Our method designs three types of memory and a library-enhanced reasoning component, enabling LLMs to improve over time through experience. Experimental results on four chemical reasoning datasets from SciBench demonstrate that ChemAgent achieves performance gains of up to 46% (GPT-4), significantly outperforming existing methods. Our findings suggest substantial potential for future applications, including tasks such as drug discovery and materials science. Our code can be found at https://github.com/gersteinlab/ChemAgent. Xiangru Tang, Muyang Ye, Yanjun Shao, Xunjian Yin, Siru Ouyang, Wangchunshu Zhou, Pan Lu, Zhuosheng Zhang 0001, Yilun Zhao 0001, Arman Cohan, Mark Gerstein |
ICLR | 4 |