VLDB 2026 Research / reviewers in the wild / expert
Bohan Lyu 0001
dblp:355/3278-1
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0003-2479-6314ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 29% Generative modeling · 16% Probabilistic and Bayesian machine learning · 15% | |
| Software engineering, system software, and programming languages
1 paper |
Software maintenance and evolution · 50% Empirical software engineering · 50% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks |
1.0 | 1 | 2026 | Exploring Diffusion Models' Corruption Stage in Few-Shot Fine-tuning and Mitigating with Bayesian Neural Networks · KDD (1) 2026 |
Machine learning › Generative modeling
diffusion model |
1.0 | 1 | 2026 | Exploring Diffusion Models' Corruption Stage in Few-Shot Fine-tuning and Mitigating with Bayesian Neural Networks · KDD (1) 2026 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot adaptation |
1.0 | 1 | 2026 | Exploring Diffusion Models' Corruption Stage in Few-Shot Fine-tuning and Mitigating with Bayesian Neural Networks · KDD (1) 2026 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
1.0 | 1 | 2026 | Exploring Diffusion Models' Corruption Stage in Few-Shot Fine-tuning and Mitigating with Bayesian Neural Networks · KDD (1) 2026 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.9 | 1 | 2025 | Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation · ICML 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge management
knowledge internalization |
0.9 | 1 | 2025 | Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation · ICML 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
large language model benchmarking |
0.9 | 1 | 2025 | Surge: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors · EMNLP 2025 |
Computer vision › Vision and language
multimodal evaluation |
0.9 | 1 | 2025 | MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks · ICLR 2025 |
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models |
0.9 | 1 | 2025 | Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation · ICML 2025 |
Natural language and speech › Machine translation › machine translation evaluation
automatic evaluation metrics |
0.8 | 1 | 2024 | VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation · EMNLP 2024 |
Machine learning › Generative modeling
video generation |
0.8 | 1 | 2024 | VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation · EMNLP 2024 |
Machine learning › Generative modeling
image generation |
0.3 | 1 | 2026 | Exploring Diffusion Models' Corruption Stage in Few-Shot Fine-tuning and Mitigating with Bayesian Neural Networks · KDD (1) 2026 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation › personalization
subject-driven generation |
0.3 | 1 | 2026 | Exploring Diffusion Models' Corruption Stage in Few-Shot Fine-tuning and Mitigating with Bayesian Neural Networks · KDD (1) 2026 |
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
science question answering |
0.3 | 1 | 2025 | Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation · ICML 2025 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
surrogate model |
0.3 | 1 | 2025 | Surge: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors · EMNLP 2025 |
Empirical software engineering › mining software repositories
github |
0.3 | 1 | 2025 | Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub · ACL (1) 2025 |
Software maintenance and evolution
software ecosystems |
0.3 | 1 | 2025 | Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
large language model prompting · 2.6variational inference · 1.0regularization · 1.0bayesian neural network · 1.0tool-generated solutions · 0.9supervised fine-tuning · 0.9scaling law analysis · 0.9multimodal evaluation metrics · 0.9accuracy-based problem categorization · 0.9human preference learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Diffusion Models' Corruption Stage in Few-Shot Fine-tuning and Mitigating with Bayesian Neural NetworksabstractFew-shot fine-tuning of Diffusion Models (DMs) is a key advancement, significantly reducing training costs and enabling personalized AI applications. However, we explore the training dynamics of DMs and observe an unanticipated phenomenon: during the training process, image fidelity initially improves, then unexpectedly deteriorates with the emergence of noisy patterns, only to recover later with severe overfitting. We term the stage with generated noisy patterns as corruption stage. To understand this corruption stage, we begin by heuristically modeling the one-shot fine-tuning scenario, and then extend this modeling to more general cases. Through this modeling, we identify the primary cause of this corruption stage: a narrowed learning distribution inherent in the nature of few-shot fine-tuning. To tackle this, we apply Bayesian Neural Networks (BNNs) on DMs with variational inference to implicitly broaden the learned distribution, and present that the learning target of the BNNs can be naturally regarded as an expectation of the diffusion loss and a further regularization with the pretrained DMs. This approach is highly compatible with current few-shot fine-tuning methods in DMs and does not introduce any extra inference costs. Experimental results demonstrate that our method significantly mitigates corruption, and improves the fidelity, quality and diversity of the generated images in both object-driven and subject-driven generation tasks. Jiaru Zhang, Yang Hua 0001, Bohan Lyu 0001, Hao Wang 0022, Tao Song 0003, Haibing Guan |
KDD (1) | 4 |
| 2025 | Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHubabstractBohan Lyu, Xin Cong, Heyang Yu, Pan Yang, Cheng Qian, Zihe Wang, Yujia Qin, Yining Ye, Yaxi Lu, Chen Qian, Zhong Zhang, Yukun Yan, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Bohan Lyu 0001, Xin Cong, Heyang Yu, Pan Yang 0022, Cheng Qian 0008, Yujia Qin, Yining Ye, Yaxi Lu, Zhong Zhang 0004, Yukun Yan, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 1 |
| 2025 | Surge: On the Potential of Large Language Models as General-Purpose Surrogate Code ExecutorsabstractNeural surrogate models are powerful and efficient tools in data mining.Meanwhile, large language models (LLMs) have demonstrated remarkable capabilities in code-related tasks, such as generation and understanding.However, an equally important yet underexplored question is whether LLMs can serve as surrogate models for code execution prediction.To systematically investigate it, we introduce SURGE, a comprehensive benchmark with 1160 problems covering 8 key aspects: multilanguage programming tasks, competitionlevel programming problems, repository-level code analysis, high-cost scientific computing, time-complexity-intensive algorithms, buggy code analysis, programs dependent on specific compilers or execution environments, and formal mathematical proof verification.Through extensive analysis of 21 open-source and proprietary LLMs, we examine scaling laws, data efficiency, and predictive accuracy.Our findings reveal important insights about the feasibility of LLMs as efficient surrogates for computational processes. Bohan Lyu 0001, Siqiao Huang, Zichen Liang |
EMNLP | 1 |
| 2025 | MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World TasksabstractWe present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users.
Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal tasks, while enabling cost-effective and accurate model evaluation.
In particular, we collected 505 realistic tasks encompassing over 8,000 samples from 16 expert annotators to extensively cover the multimodal task space. Instead of unifying these problems into standard multi-choice questions (like MMMU, MM-Bench, and MMT-Bench), we embrace a wide range of output formats like numbers, phrases, code, \LaTeX, coordinates, JSON, free-form, etc. To accommodate these formats, we developed over 40 metrics to evaluate these tasks.
Unlike existing benchmarks, MEGA-Bench offers a fine-grained capability report across multiple dimensions (e.g., application, input type, output format, skill), allowing users to interact with and visualize model capabilities in depth. We evaluate a wide variety of frontier vision-language models on MEGA-Bench to understand their capabilities across these dimensions. Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang 0068, Yubo Wang 0019, Yuansheng Ni, Ziyan Jiang, Wang Zhu 0001, Bohan Lyu 0001, Dongfu Jiang, Hexiang Hu, Xiang Yue, Wenhu Chen |
ICLR | 10 |
| 2025 | Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage AdaptationabstractLarge Language Models (LLMs) demonstrate promising capabilities in solving scientific problems but often suffer from the issue of hallucination. While integrating LLMs with tools can mitigate this issue, models fine-tuned on tool usage become overreliant on them and incur unnecessary costs. Inspired by how human experts assess problem complexity before selecting solutions, we propose a novel two-component fine-tuning method, Adapting while Learning (AWL). In the first component World Knowledge Learning (WKL), LLMs internalize scientific knowledge by learning from tool-generated solutions. In the second component Tool Usage Adaptation (TUA), we categorize problems as easy or hard based on the model’s accuracy, and train it to maintain direct reasoning for easy problems while switching to tools for hard ones. We validate our method on 6 scientific benchmark datasets across climate science, epidemiology, physics, and other domains. Compared to the original instruct model (8B), models post-trained with AWL achieve 29.11% higher answer accuracy and 12.72% better tool usage accuracy, even surpassing state-of-the-art models including GPT-4o and Claude-3.5 on 4 custom-created datasets. Our code is open-source at https://github.com/Rose-STL-Lab/Adapting-While-Learning. Bohan Lyu 0001, Yadi Cao, Duncan Watson-Parris, Leon Bergen, Taylor Berg-Kirkpatrick, Rose Yu |
ICML | 1 |
| 2025 | Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving of InequalitiesabstractLLM-based formal proof assistants (e.g., in Lean) hold great promise for automating mathematical discovery. But beyond syntactic correctness, do these systems truly understand mathematical structure as humans do? We investigate this question in context of mathematical inequalities---specifically the prover's ability to recognize that the given problem simplifies by applying a known inequality such as AM/GM. Specifically, we are interested in their ability to do this in a {\em compositional setting} where multiple inequalities must be applied as part of a solution. We introduce \ineqcomp, a benchmark built from elementary inequalities through systematic transformations, including variable duplication, algebraic rewriting, and multi-step composition. Although these problems remain easy for humans, we find that most provers---including Goedel, STP, and Kimina-7B---struggle significantly. DeepSeek-Prover-V2-7B shows relative robustness, but still suffers a 20\% performance drop (pass@32). Even for DeepSeek-Prover-V2-671B model, the gap between compositional variants and seed problems exists, implying that simply scaling up the model size alone does not fully solve the compositional weakness. Strikingly, performance remains poor for all models even when formal proofs of the constituent parts are provided in context, revealing that the source of weakness is indeed in compositional reasoning. Our results expose a persisting gap between the generalization behavior of current AI provers and human mathematical intuition. All data and evaluation code can be found at \url{https://github.com/haoyuzhao123/LeanIneqComp}. Yihan Geng, Shange Tang, Bohan Lyu 0001, Hongzhou Lin, Chi Jin 0001, Sanjeev Arora |
NeurIPS | 5 |
| 2024 | VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video GenerationabstractXuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do, Yuansheng Ni, Bohan Lyu, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Yuchen Lin, Wenhu Chen. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Dongfu Jiang, Ge Zhang 0009, Max Ku, Achint Soni, Sherman Siu, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang 0068, Quy Duc Do, Yuansheng Ni, Bohan Lyu 0001, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Y. Lin, Wenhu Chen |
EMNLP | 14 |