Weiliang Will Zeng

dblp:324/5007 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 51% Compilers and program optimization · 34% Software testing · 15%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%
Artificial intelligence
1 paper
Graph learning · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
code generation with language models
0.912025
How efficient is LLM-generated code? A rigorous & high-standard benchmark · ICLR 2025
Parallel and multicore computing › task scheduling
DAG scheduling
0.712023
Neural DAG Scheduling via One-Shot Priority Sampling · ICLR 2023
Parallel and multicore computing
task scheduling
0.712023
Neural DAG Scheduling via One-Shot Priority Sampling · ICLR 2023
Compilers and program optimization
instruction scheduling
0.612022
Neural Topological Ordering for Computation Graphs · NeurIPS 2022
Software testing
test generation
0.312025
How efficient is LLM-generated code? A rigorous & high-standard benchmark · ICLR 2025
Machine learning › Graph learning › graph neural network
attention-based graph neural network
0.212022
Neural Topological Ordering for Computation Graphs · NeurIPS 2022
Machine learning › Graph learning
graph neural network
0.212022
Neural Topological Ordering for Computation Graphs · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

topoformer · 1.1encoder-decoder · 1.1attention-based graph neural network · 1.1rao-blackwellization · 0.9large language model · 0.9one-shot priority sampling · 0.7neural scheduling · 0.7
YearPublicationVenuePosition
2025 How efficient is LLM-generated code? A rigorous & high-standard benchmark
abstract
The emergence of large language models (LLMs) has significantly pushed the frontiers of program synthesis. Advancement of LLM-based program synthesis calls for a thorough evaluation of LLM-generated code. Most evaluation frameworks focus on the (functional) correctness of generated code; efficiency, as an important measure of code quality, has been overlooked in existing evaluations. In this work, we develop ENAMEL (EfficeNcy AutoMatic EvaLuator), a rigorous and high-standard benchmark for evaluating the capability of LLMs in generating efficient code. Firstly, we propose a new efficiency metric called eff@k, which generalizes the pass@k metric from correctness to efficiency and appropriately handles right-censored execution time. Furthermore, we derive an unbiased and variance-reduced estimator of eff@k via Rao–Blackwellization; we also provide a numerically stable implementation for the new estimator. Secondly, to set a high-standard for efficiency evaluation, we employ a human expert to design best algorithms and implementations as our reference solutions of efficiency, many of which are much more efficient than existing canonical solutions in HumanEval and HumanEval+. Moreover, to ensure a rigorous evaluation, we employ a human expert to curate strong test case generators to filter out wrong code and differentiate suboptimal algorithms. An extensive study across 30 popular LLMs using our benchmark ENAMEL shows that LLMs still fall short of generating expert-level efficient code. Using two subsets of our problem set, we demonstrate that such deficiency is because current LLMs struggle in designing advanced algorithms and are barely aware of implementation optimization.
Ruizhong Qiu, Weiliang Will Zeng, James Ezick, Christopher Lott, Hanghang Tong
ICLR2
2023 Neural DAG Scheduling via One-Shot Priority Sampling
Wonseok Jeon, Mukul Gagrani, Burak Bartan, Weiliang Will Zeng, Harris Teague, Piero Zappi, Christopher Lott
ICLR4
2022 Neural Topological Ordering for Computation Graphs
abstract
Recent works on machine learning for combinatorial optimization have shown that learning based approaches can outperform heuristic methods in terms of speed and performance. In this paper, we consider the problem of finding an optimal topological order on a directed acyclic graph (DAG) with focus on the memory minimization problem which arises in compilers. We propose an end-to-end machine learning based approach for topological ordering using an encoder-decoder framework. Our encoder is a novel attention based graph neural network architecture called \emph{Topoformer} which uses different topological transforms of a DAG for message passing. The node embeddings produced by the encoder are converted into node priorities which are used by the decoder to generate a probability distribution over topological orders. We train our model on a dataset of synthetically generated graphs called layered graphs. We show that our model outperforms, or is on-par, with several topological ordering baselines while being significantly faster on synthetic graphs with up to 2k nodes. We also train and test our model on a set of real-world computation graphs, showing performance improvements.
Mukul Gagrani, Corrado Rainone, Harris Teague, Wonseok Jeon, Roberto Bondesan, Herke van Hoof, Christopher Lott, Weiliang Will Zeng, Piero Zappi
NeurIPS9