Wenxiang Hu

dblp:141/4590 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 100%
Software engineering, system software, and programming languages
4 papers
Compilers and program optimization · 41% Program synthesis and code generation · 27% Empirical software engineering · 16%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
instruction tuning
1.522024
WizardCoder: Empowering Code Large Language Models with Evol-Instruct · ICLR 2024
WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning · ACL (1) 2024
Natural language and speech › Language models and text generation
code language models
1.022025
WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning · ACL (1) 2024
EpiCoder: Encompassing Diversity and Complexity in Code Generation · ICML 2025
Compilers and program optimization
code generation
0.912025
EpiCoder: Encompassing Diversity and Complexity in Code Generation · ICML 2025
Program synthesis and code generation › code generation with language models
repository-level code generation
0.912025
EpiCoder: Encompassing Diversity and Complexity in Code Generation · ICML 2025
Natural language and speech › Language models and text generation
code generation
0.812024
WizardCoder: Empowering Code Large Language Models with Evol-Instruct · ICLR 2024
Compilers and program optimization › deep learning compiler
deep learning compiler optimization
0.412020
Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks · OSDI 2020
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.412020
Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks · OSDI 2020
Parallel and multicore computing
task scheduling
0.412020
Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks · OSDI 2020
Software maintenance and evolution
code recommendation
0.212016
Bing developer assistant: improving developer productivity by recommending sample code · SIGSOFT FSE 2016
Software testing
compiler testing
0.212016
An empirical comparison of compiler testing techniques · ICSE 2016
Empirical software engineering › software evaluation
empirical evaluation of testing techniques
0.212016
An empirical comparison of compiler testing techniques · ICSE 2016
Empirical software engineering
mining software repositories
0.212016
An empirical comparison of compiler testing techniques · ICSE 2016
Information retrieval › document retrieval › domain-specific retrieval
code search
0.112016
Bing developer assistant: improving developer productivity by recommending sample code · SIGSOFT FSE 2016

Methods — techniques the papers use, named apart from their topics

fine-tuning · 1.7feature tree synthesis · 1.7instruction tuning · 0.8instruction fine-tuning · 0.8instruction evolution · 0.8code mining · 0.5API usage mining · 0.5empirical comparison · 0.2
YearPublicationVenuePosition
2025 EpiCoder: Encompassing Diversity and Complexity in Code Generation
abstract
Existing methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data. In this paper, we introduce a novel feature tree-based synthesis framework, which revolves around hierarchical code features derived from high-level abstractions of code. The feature tree is constructed from raw data and refined iteratively to increase the quantity and diversity of the extracted features, which captures and recognizes more complex patterns and relationships within the code. By adjusting the depth and breadth of the sampled subtrees, our framework provides precise control over the complexity of the generated code, enabling functionalities that range from function-level operations to multi-file scenarios. We fine-tuned widely-used base models to obtain EpiCoder series, achieving state-of-the-art performance on multiple benchmarks at both the function and file levels. In particular, empirical evidence indicates that our approach shows significant potential in the synthesizing of repository-level code data. Our code and data are publicly available.
Yaoxiang Wang, Haoling Li, Xin Zhang 0099, Jie Wu 0001, Xiao Liu 0029, Wenxiang Hu, Zhongxin Guo, Yangyu Huang, Yujiu Yang 0001, Jinsong Su, Qi Chen 0009, Scarlett Li
ICML6
2025 An optimized hierarchical point cloud registration algorithm
Fuqun Zhao, Wenxiang Hu
Multim. Syst.3
2024 WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning
abstract
Zhaojian Yu, Xin Zhang, Ning Shang, Yangyu Huang, Can Xu, Yishujie Zhao, Wenxiang Hu, Qiufeng Yin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhaojian Yu, Xin Zhang 0099, Yangyu Huang, Yishujie Zhao, Wenxiang Hu, Qiufeng Yin
ACL (1)7
2024 WizardCoder: Empowering Code Large Language Models with Evol-Instruct
abstract
Code Large Language Models (Code LLMs), such as StarCoder, have demonstrated remarkable performance in various code-related tasks. However, different from their counterparts in the general language modeling field, the technique of instruction fine-tuning remains relatively under-researched in this domain. In this paper, we present Code Evol-Instruct, a novel approach that adapts the Evol-Instruct method to the realm of code, enhancing Code LLMs to create novel models, WizardCoder. Through comprehensive experiments on five prominent code generation benchmarks, namely HumanEval, HumanEval+, MBPP, DS-1000, and MultiPL-E, our models showcase outstanding performance. They consistently outperform all other open-source Code LLMs by a significant margin. Remarkably, WizardCoder 15B even surpasses the well-known closed-source LLMs, including Anthropic's Claude and Google's Bard, on the HumanEval and HumanEval+ benchmarks. Additionally, WizardCoder 34B not only achieves a HumanEval score comparable to GPT3.5 (ChatGPT) but also surpasses it on the HumanEval+ benchmark. Furthermore, our preliminary exploration highlights the pivotal role of instruction complexity in achieving exceptional coding performance.
Can Xu 0002, Pu Zhao 0004, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma 0004, Qingwei Lin, Daxin Jiang
ICLR6
2020 Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks
Lingxiao Ma, Zhi Yang 0001, Jilong Xue, Youshan Miao, Wenxiang Hu, Fan Yang 0024, Lidong Zhou
OSDI7
2018 Learning to Accelerate Symbolic Execution via Code Transformation
abstract
Symbolic execution is an effective but expensive technique for automated test generation. Over the years, a large number of refined symbolic execution techniques have been proposed to improve its efficiency. However, the symbolic execution efficiency problem remains, and largely limits the application of symbolic execution in practice. Orthogonal to refined symbolic execution, in this paper we propose to accelerate symbolic execution through semantic-preserving code transformation on the target programs. During the initial stage of this direction, we adopt a particular code transformation, compiler optimization, which is initially proposed to accelerate program concrete execution by transforming the source program into another semantic-preserving target program with increased efficiency (e.g., faster or smaller). However, compiler optimizations are mostly designed to accelerate program concrete execution rather than symbolic execution. Recent work also reported that unified settings on compiler optimizations that can accelerate symbolic execution for any program do not exist at all. Therefore, in this work we propose a machine-learning based approach to tuning compiler optimizations to accelerate symbolic execution, whose results may also aid further design of specific code transformations for symbolic execution. In particular, the proposed approach LEO separates source-code functions and libraries through our program-splitter, and predicts individual compiler optimization (i.e., whether a type of code transformation is chosen) separately through analyzing the performance of existing symbolic execution. Finally, LEO applies symbolic execution on the code transformed by compiler optimization (through our local-optimizer). We conduct an empirical study on GNU Coreutils programs using the KLEE symbolic execution engine. The results show that LEO significantly accelerates symbolic execution, outperforming the default KLEE configurations (i.e., turning on/off all compiler optimizations) in various settings, e.g., with the default training/testing time, LEO achieves the highest line coverage in 50/68 programs, and its average improvement rate on all programs is 46.48%/88.92% in terms of line coverage compared with turning on/off all compiler optimizations.
Junjie Chen 0003, Wenxiang Hu, Lingming Zhang 0001, Dan Hao 0001, Sarfraz Khurshid, Lu Zhang 0023
ECOOP2
2016 An empirical comparison of compiler testing techniques
abstract
Compilers, as one of the most important infrastructure of today's digital world, are expected to be trustworthy. Different testing techniques are developed for testing compilers automatically. However, it is unknown so far how these testing techniques compared to each other in terms of testing effectiveness: how many bugs a testing technique can find within a time limit.
Junjie Chen 0003, Wenxiang Hu, Dan Hao 0001, Yingfei Xiong 0001, Hongyu Zhang 0002, Lu Zhang 0023
ICSE2
2016 Bing developer assistant: improving developer productivity by recommending sample code
abstract
In programming practice, developers often need sample code in order to learn how to solve a programming-related problem. For example, how to reuse an Application Programming Interface (API) of a large-scale software library and how to implement a certain functionality. We believe that previously written code can help developers understand how others addressed the similar problems and can help them write new programs. We develop a tool called Bing Developer Assistant (BDA), which improves developer productivity by recommending sample code mined from public software repositories (such as GitHub) and web pages (such as Stack Overflow). BDA can automatically mine code snippets that implement an API or answer a code search query. It has been implemented as a free-downloadable extension of Microsoft Visual Studio and has received more than 670K downloads since its initial release in December 2014. BDA is publicly available at: http://aka.ms/devassistant.
Hongyu Zhang 0002, Anuj Jain, Gaurav Khandelwal, Chandrashekhar Kaushik, Scott Ge, Wenxiang Hu
SIGSOFT FSE6