VLDB 2026 Research / reviewers in the wild / expert
Han Weng
dblp:349/5252
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 71% Compilers and program optimization · 29% | |
| Artificial intelligence
1 paper |
Language models and text generation · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program synthesis and code generation
code generation with language models |
1.6 | 2 | 2025 | EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning · ICML 2025 EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization · NeurIPS 2024 |
Natural language and speech › Language models and text generation
LLM agents |
1.0 | 1 | 2026 | UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured Data · ACL (1) 2026 |
Compilers and program optimization
code efficiency optimization |
0.8 | 1 | 2024 | EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization · NeurIPS 2024 |
Program synthesis and code generation › code generation with language models
fine-tuning for code generation |
0.3 | 1 | 2025 | EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
self-correction · 2.0react · 2.0code generation · 2.0large language model fine-tuning · 0.9execution-based code selection · 0.9large language model · 0.8execution profiling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured DataabstractIn real-world business environments, data is stored in a variety of sources, including structured relational databases, semi-structured databases, and unstructured files.The ability to extract reasonable insights across these diverse sources is integral to data-driven decisionmaking.Existing benchmarks, however, are limited in assessing agents' capabilities across these diverse data types.To address this gap, we introduce UniDataBench, a multi-source benchmark designed to evaluate the performance of data analytics agents in handling diverse data sources.Specifically, UniDataBench is constructed based on real-life industry analysis reports, employing a pipeline to synthesize data that aligns with authentic analytical trends.It encompasses diverse datasets spanning relational databases, CSV files, and NoSQL stores to reflect real-world business settings, and provides a unified framework for evaluating how effectively agents can explore multiple data formats, extract insights, and generate meaningful summaries and recommendations.Based on UniDataBench, we propose a novel LLM-based agent named ReActInsight, an autonomous agent that performs end-to-end analysis over diverse data sources by automatically discovering cross-source linkages, decomposing goals, and generating robust, self-correcting code to extract actionable insights.Our benchmark and agent together provide a framework for facilitating the development of data analytics agents in real-world applications. Han Weng, Yuanfeng Song, Xiaoming Yin |
ACL (1) | 1 |
| 2025 | Towards Database-Free Text-to-SQL Evaluation: A Graph-Based Metric for Functional CorrectnessabstractExecution Accuracy and Exact Set Match are two predominant metrics for evaluating the functional correctness of SQL queries in modern Text-to-SQL tasks. However, both metrics have notable limitations: Exact Set Match fails when queries are functionally equivalent but syntactically different, while Execution Accuracy is prone to false positives due to inadequately prepared test databases, which can be costly to create, particularly in large-scale industrial applications. To overcome these challenges, we propose a novel graph-based metric, FuncEvalGMN, that effectively overcomes the deficiencies of the aforementioned metric designs. Our method utilizes a relational operator tree (ROT), referred to as RelNode, to extract rich semantic information from the logical execution plan of SQL queries, and embed it into a graph. We then train a graph neural network (GNN) to perform graph matching on pairs of SQL queries through graph contrastive learning. FuncEvalGMN offers two highly desired advantages: (i) it requires only the database schema to derive logical execution plans, eliminating the need for extensive test database preparation, and (ii) it demonstrates strong generalization capabilities on unseen datasets. These properties highlight FuncEvalGMN’s robustness as a reliable metric for assessing functional correctness across a wide range of Text-to-SQL applications. Longjie Cui, Han Weng, Yingxiang Yang, Xiaoming Yin, Jiajun Xie |
COLING | 3 |
| 2025 | EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuningabstractAs large language models (LLMs) play an increasingly important role in code generation, enhancing both correctness and efficiency has become crucial. Current methods primarily focus on correctness, often overlooking efficiency. To address this gap, we introduce SWIFTCODE to improve both aspects by fine-tuning LLMs on a high-quality dataset comprising correct and efficient code samples. Our methodology involves leveraging multiple LLMs to generate diverse candidate code solutions for various tasks across different programming languages. We then evaluate these solutions by directly measuring their execution time and memory usage through local execution. The code solution with the lowest execution time and memory consumption is selected as the final output for each task. Experimental results demonstrate significant improvements when fine-tuning with SWIFTCODE. For instance, Qwen2.5-Coder-7B-Instruct’s pass@1 score increases from 44.8% to 57.7%, while the average execution time for correct tasks decreases by 48.4%. SWIFTCODE offers a scalable and effective solution for advancing AI-driven code generation, benefiting both software development and computational problem-solving. Dong Huang 0005, Guangtao Zeng, Jianbo Dai, Meng Luo 0010, Han Weng, Yuhao Qing, Heming Cui, Zhijiang Guo, Jie Zhang 0050 |
ICML | 5 |
| 2024 | EffiLearner: Enhancing Efficiency of Generated Code via Self-OptimizationabstractLarge language models (LLMs) have shown remarkable progress in code generation, but their generated code often suffers from inefficiency, resulting in longer execution times and higher memory consumption. To address this issue, we propose EffiLearner, a self-optimization framework that utilizes execution overhead profiles to improve the efficiency of LLM-generated code. EffiLearner first generates code using an LLM, then executes it locally to capture execution time and memory usage profiles. These profiles are fed back to the LLM, which then revises the code to reduce overhead. To evaluate the effectiveness of EffiLearner, we conduct extensive experiments on EffiBench and two commonly used code generation benchmarks with 16 open-source and 6 closed-source models. Our evaluation results demonstrate that through iterative self-optimization, EffiLearner significantly enhances the efficiency of LLM-generated code. For example, the execution time (ET) of StarCoder2-15B for the EffiBench decreases from 0.93 (s) to 0.12 (s) which reduces 87.1\% execution time requirement compared with the initial code. The total memory usage (TMU) of StarCoder2-15B also decreases from 22.02 (Mb*s) to 2.03 (Mb*s), which decreases 90.8\% total memory consumption during the execution process. Dong Huang 0005, Jianbo Dai, Han Weng, Puzhen Wu, Yuhao Qing, Heming Cui, Zhijiang Guo, Jie Zhang 0050 |
NeurIPS | 3 |
| 2023 | A graph-based code representation method to improve code readability classification
Qing Mi, Han Weng, Qinghang Bao, Longjie Cui, Wei Ma 0008 |
Empir. Softw. Eng. | 3 |