VLDB 2026 Research / reviewers in the wild / expert
Wenxiang Hu
dblp:141/4590
· DBLP profile ↗
8ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 100% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 41% Program synthesis and code generation · 27% Empirical software engineering · 16% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
instruction tuning |
1.5 | 2 | 2024 | WizardCoder: Empowering Code Large Language Models with Evol-Instruct · ICLR 2024 WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning · ACL (1) 2024 |
Natural language and speech › Language models and text generation
code language models |
1.0 | 2 | 2025 | WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning · ACL (1) 2024 EpiCoder: Encompassing Diversity and Complexity in Code Generation · ICML 2025 |
Compilers and program optimization
code generation |
0.9 | 1 | 2025 | EpiCoder: Encompassing Diversity and Complexity in Code Generation · ICML 2025 |
Program synthesis and code generation › code generation with language models
repository-level code generation |
0.9 | 1 | 2025 | EpiCoder: Encompassing Diversity and Complexity in Code Generation · ICML 2025 |
Natural language and speech › Language models and text generation
code generation |
0.8 | 1 | 2024 | WizardCoder: Empowering Code Large Language Models with Evol-Instruct · ICLR 2024 |
Compilers and program optimization › deep learning compiler
deep learning compiler optimization |
0.4 | 1 | 2020 | Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks · OSDI 2020 |
Parallel and multicore computing › parallel scheduling
runtime scheduling |
0.4 | 1 | 2020 | Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks · OSDI 2020 |
Parallel and multicore computing
task scheduling |
0.4 | 1 | 2020 | Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks · OSDI 2020 |
Software maintenance and evolution
code recommendation |
0.2 | 1 | 2016 | Bing developer assistant: improving developer productivity by recommending sample code · SIGSOFT FSE 2016 |
Software testing
compiler testing |
0.2 | 1 | 2016 | An empirical comparison of compiler testing techniques · ICSE 2016 |
Empirical software engineering › software evaluation
empirical evaluation of testing techniques |
0.2 | 1 | 2016 | An empirical comparison of compiler testing techniques · ICSE 2016 |
Empirical software engineering
mining software repositories |
0.2 | 1 | 2016 | An empirical comparison of compiler testing techniques · ICSE 2016 |
Information retrieval › document retrieval › domain-specific retrieval
code search |
0.1 | 1 | 2016 | Bing developer assistant: improving developer productivity by recommending sample code · SIGSOFT FSE 2016 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 1.7feature tree synthesis · 1.7instruction tuning · 0.8instruction fine-tuning · 0.8instruction evolution · 0.8code mining · 0.5API usage mining · 0.5empirical comparison · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EpiCoder: Encompassing Diversity and Complexity in Code GenerationabstractExisting methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data. In this paper, we introduce a novel feature tree-based synthesis framework, which revolves around hierarchical code features derived from high-level abstractions of code. The feature tree is constructed from raw data and refined iteratively to increase the quantity and diversity of the extracted features, which captures and recognizes more complex patterns and relationships within the code. By adjusting the depth and breadth of the sampled subtrees, our framework provides precise control over the complexity of the generated code, enabling functionalities that range from function-level operations to multi-file scenarios. We fine-tuned widely-used base models to obtain EpiCoder series, achieving state-of-the-art performance on multiple benchmarks at both the function and file levels. In particular, empirical evidence indicates that our approach shows significant potential in the synthesizing of repository-level code data. Our code and data are publicly available. Yaoxiang Wang, Haoling Li, Xin Zhang 0099, Jie Wu 0001, Xiao Liu 0029, Wenxiang Hu, Zhongxin Guo, Yangyu Huang, Yujiu Yang 0001, Jinsong Su, Qi Chen 0009, Scarlett Li |
ICML | 6 |
| 2025 | An optimized hierarchical point cloud registration algorithm
Fuqun Zhao, Wenxiang Hu |
Multim. Syst. | 3 |
| 2024 | WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction TuningabstractZhaojian Yu, Xin Zhang, Ning Shang, Yangyu Huang, Can Xu, Yishujie Zhao, Wenxiang Hu, Qiufeng Yin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhaojian Yu, Xin Zhang 0099, Yangyu Huang, Yishujie Zhao, Wenxiang Hu, Qiufeng Yin |
ACL (1) | 7 |
| 2024 | WizardCoder: Empowering Code Large Language Models with Evol-InstructabstractCode Large Language Models (Code LLMs), such as StarCoder, have demonstrated remarkable performance in various code-related tasks. However, different from their counterparts in the general language modeling field, the technique of instruction fine-tuning remains relatively under-researched in this domain. In this paper, we present Code Evol-Instruct, a novel approach that adapts the Evol-Instruct method to the realm of code, enhancing Code LLMs to create novel models, WizardCoder. Through comprehensive experiments on five prominent code generation benchmarks, namely HumanEval, HumanEval+, MBPP, DS-1000, and MultiPL-E, our models showcase outstanding performance. They consistently outperform all other open-source Code LLMs by a significant margin. Remarkably, WizardCoder 15B even surpasses the well-known closed-source LLMs, including Anthropic's Claude and Google's Bard, on the HumanEval and HumanEval+ benchmarks. Additionally, WizardCoder 34B not only achieves a HumanEval score comparable to GPT3.5 (ChatGPT) but also surpasses it on the HumanEval+ benchmark. Furthermore, our preliminary exploration highlights the pivotal role of instruction complexity in achieving exceptional coding performance. Can Xu 0002, Pu Zhao 0004, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma 0004, Qingwei Lin, Daxin Jiang |
ICLR | 6 |
| 2020 | Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks
Lingxiao Ma, Zhi Yang 0001, Jilong Xue, Youshan Miao, Wenxiang Hu, Fan Yang 0024, Lidong Zhou |
OSDI | 7 |
| 2018 | Learning to Accelerate Symbolic Execution via Code TransformationabstractSymbolic execution is an effective but expensive technique for automated test generation. Over the years, a large number of refined symbolic execution techniques have been proposed to improve its efficiency. However, the symbolic execution efficiency problem remains, and largely limits the application of symbolic execution in practice. Orthogonal to refined symbolic execution, in this paper we propose to accelerate symbolic execution through semantic-preserving code transformation on the target programs. During the initial stage of this direction, we adopt a particular code transformation, compiler optimization, which is initially proposed to accelerate program concrete execution by transforming the source program into another semantic-preserving target program with increased efficiency (e.g., faster or smaller). However, compiler optimizations are mostly designed to accelerate program concrete execution rather than symbolic execution. Recent work also reported that unified settings on compiler optimizations that can accelerate symbolic execution for any program do not exist at all. Therefore, in this work we propose a machine-learning based approach to tuning compiler optimizations to accelerate symbolic execution, whose results may also aid further design of specific code transformations for symbolic execution. In particular, the proposed approach LEO separates source-code functions and libraries through our program-splitter, and predicts individual compiler optimization (i.e., whether a type of code transformation is chosen) separately through analyzing the performance of existing symbolic execution. Finally, LEO applies symbolic execution on the code transformed by compiler optimization (through our local-optimizer). We conduct an empirical study on GNU Coreutils programs using the KLEE symbolic execution engine. The results show that LEO significantly accelerates symbolic execution, outperforming the default KLEE configurations (i.e., turning on/off all compiler optimizations) in various settings, e.g., with the default training/testing time, LEO achieves the highest line coverage in 50/68 programs, and its average improvement rate on all programs is 46.48%/88.92% in terms of line coverage compared with turning on/off all compiler optimizations. Junjie Chen 0003, Wenxiang Hu, Lingming Zhang 0001, Dan Hao 0001, Sarfraz Khurshid, Lu Zhang 0023 |
ECOOP | 2 |
| 2016 | An empirical comparison of compiler testing techniquesabstractCompilers, as one of the most important infrastructure of today's digital world, are expected to be trustworthy. Different testing techniques are developed for testing compilers automatically. However, it is unknown so far how these testing techniques compared to each other in terms of testing effectiveness: how many bugs a testing technique can find within a time limit. Junjie Chen 0003, Wenxiang Hu, Dan Hao 0001, Yingfei Xiong 0001, Hongyu Zhang 0002, Lu Zhang 0023 |
ICSE | 2 |
| 2016 | Bing developer assistant: improving developer productivity by recommending sample codeabstractIn programming practice, developers often need sample code in order to learn how to solve a programming-related problem. For example, how to reuse an Application Programming Interface (API) of a large-scale software library and how to implement a certain functionality. We believe that previously written code can help developers understand how others addressed the similar problems and can help them write new programs. We develop a tool called Bing Developer Assistant (BDA), which improves developer productivity by recommending sample code mined from public software repositories (such as GitHub) and web pages (such as Stack Overflow). BDA can automatically mine code snippets that implement an API or answer a code search query. It has been implemented as a free-downloadable extension of Microsoft Visual Studio and has received more than 670K downloads since its initial release in December 2014. BDA is publicly available at: http://aka.ms/devassistant. Hongyu Zhang 0002, Anuj Jain, Gaurav Khandelwal, Chandrashekhar Kaushik, Scott Ge, Wenxiang Hu |
SIGSOFT FSE | 6 |