VLDB 2026 Research / reviewers in the wild / expert
Zhenwen Li
dblp:254/2103
· DBLP profile ↗
7ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0002-1601-8403ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Information extraction and text analysis · 31% Efficient and distributed learning · 28% Knowledge representation and reasoning · 21% | |
| Software engineering, system software, and programming languages
2 papers |
Software testing · 54% Program synthesis and code generation · 36% Requirements engineering and software design · 11% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% |
Topics — the 11 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
semantic parsing |
1.0 | 2 | 2022 | Exploring the Secrets Behind the Learning Difficulty of Meaning Representations for Semantic Parsing · EMNLP 2022 Benchmarking Meaning Representations in Neural Semantic Parsing · EMNLP (1) 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic representation |
1.0 | 2 | 2022 | Exploring the Secrets Behind the Learning Difficulty of Meaning Representations for Semantic Parsing · EMNLP 2022 Benchmarking Meaning Representations in Neural Semantic Parsing · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.4 | 1 | 2020 | Composing Elementary Discourse Units in Abstractive Summarization · ACL 2020 |
Machine learning › Efficient and distributed learning › model compression
efficient architecture design |
0.4 | 1 | 2020 | AutoShrink: A Topology-Aware NAS for Discovering Efficient Neural Architecture · AAAI 2020 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2020 | AutoShrink: A Topology-Aware NAS for Discovering Efficient Neural Architecture · AAAI 2020 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.4 | 1 | 2020 | AutoShrink: A Topology-Aware NAS for Discovering Efficient Neural Architecture · AAAI 2020 |
Natural language and speech › Information extraction and text analysis › semantic parsing
neural semantic parsing |
0.4 | 1 | 2020 | Benchmarking Meaning Representations in Neural Semantic Parsing · EMNLP (1) 2020 |
Software testing
automated testing |
0.4 | 1 | 2020 | Clustering test steps in natural language toward automating test automation · ESEC/SIGSOFT FSE 2020 |
Software testing › test generation › automated test generation
test script generation |
0.4 | 1 | 2020 | Clustering test steps in natural language toward automating test automation · ESEC/SIGSOFT FSE 2020 |
Natural language and speech › Language models and text generation
code generation |
0.2 | 1 | 2024 | InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models · NeurIPS 2024 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2020 | AutoShrink: A Topology-Aware NAS for Discovering Efficient Neural Architecture · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
automatic evaluation metrics · 1.5natural language processing · 1.0incremental structural stability metric · 0.6constraint solving · 0.6word embedding · 0.4reinforcement learning · 0.4neural semantic parsing · 0.4neural architecture search · 0.4k-means clustering · 0.4hierarchical agglomerative clustering · 0.4execution engine · 0.4edge shrinking · 0.4directed acyclic graph · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language ModelsabstractLarge Language Models for code (code LLMs) have witnessed tremendous progress in recent years. With the rapid development of code LLMs, many popular evaluation benchmarks, such as HumanEval, DS-1000, and MBPP, have emerged to measure the performance of code LLMs with a particular focus on code generation tasks. However, they are insufficient to cover the full range of expected capabilities of code LLMs, which span beyond code generation to answering diverse coding-related questions. To fill this gap, we propose InfiBench, the first large-scale freeform question-answering (QA) benchmark for code to our knowledge, comprising 234 carefully selected high-quality Stack Overflow questions that span across 15 programming languages. InfiBench uses four types of model-free automatic metrics to evaluate response correctness where domain experts carefully concretize the criterion for each question. We conduct a systematic evaluation for over 100 latest code LLMs on InfiBench, leading to a series of novel and insightful findings. Our detailed analyses showcase potential directions for further advancement of code LLMs. InfiBench is fully open source at https://infi-coder.github.io/infibench and continuously expanding to foster more scientific and systematic practices for code LLM evaluation. Linyi Li 0001, Shijie Geng, Zhenwen Li, Yibo He, Hao Yu 0016, Ziyue Hua, Guanghan Ning, Tao Xie 0001, Hongxia Yang |
NeurIPS | 3 |
| 2022 | Exploring the Secrets Behind the Learning Difficulty of Meaning Representations for Semantic ParsingabstractPrevious research has shown that the design of Meaning Representation (MR) greatly influences model performance of a neural semantic parser.Therefore, designing a good MR is a long-term goal for semantic parsing.However, it is still an art as there is no quantitative indicator that can tell us which MR among a set of candidates may have the best final model performance.In practice, in order to select an MR, researchers often have to go through the whole training-testing process for all MR candidates, and the process often costs a lot.In this paper, we propose a data-aware metric called ISS (denoting incremental structural stability) of MRs, and demonstrate that ISS is highly correlated with model performance.The finding shows that ISS can be used as an indicator for designing MRs to avoid the costly training-testing process. Zhenwen Li, Qian Liu 0033, Jian-Guang Lou, Tao Xie 0001 |
EMNLP | 1 |
| 2022 | NL2Viz: natural language to visualization via constrained syntax-guided synthesisabstractRecent development in NL2CODE (Natural Language to Code) research allows end-users, especially novice programmers to create a concrete implementation of their ideas such as data visualization by providing natural language (NL) instructions. An NL2CODE system often fails to achieve its goal due to three major challenges: the user's words have contextual semantics, the user may not include all details needed for code generation, and the system results are imperfect and require further refinement. To address the aforementioned three challenges for NL to Visualization, we propose a new approach and its supporting tool named NL2VIZ with three salient features: (1) leveraging not only the user's NL input but also the data and program context that the NL query is upon, (2) using hard/soft constraints to reflect different confidence levels in the constraints retrieved from the user input and data/program context, and (3) providing support for result refinement and reuse. Zhengkai Wu, Vu Le 0002, Ashish Tiwari 0001, Sumit Gulwani, Arjun Radhakrishna, Ivan Radicek, Gustavo Soares, Xinyu Wang 0006, Zhenwen Li, Tao Xie 0001 |
ESEC/SIGSOFT FSE | 9 |
| 2020 | AutoShrink: A Topology-Aware NAS for Discovering Efficient Neural ArchitectureabstractResource is an important constraint when deploying Deep Neural Networks (DNNs) on mobile and edge devices. Existing works commonly adopt the cell-based search approach, which limits the flexibility of network patterns in learned cell structures. Moreover, due to the topology-agnostic nature of existing works, including both cell-based and node-based approaches, the search process is time consuming and the performance of found architecture may be sub-optimal. To address these problems, we propose AutoShrink, a topology-aware Neural Architecture Search (NAS) for searching efficient building blocks of neural architectures. Our method is node-based and thus can learn flexible network patterns in cell structures within a topological search space. Directed Acyclic Graphs (DAGs) are used to abstract DNN architectures and progressively optimize the cell structure through edge shrinking. As the search space intrinsically reduces as the edges are progressively shrunk, AutoShrink explores more flexible search space with even less search time. We evaluate AutoShrink on image classification and language tasks by crafting ShrinkCNN and ShrinkRNN models. ShrinkCNN is able to achieve up to 48% parameter reduction and save 34% Multiply-Accumulates (MACs) on ImageNet-1K with comparable accuracy of state-of-the-art (SOTA) models. Specifically, both ShrinkCNN and ShrinkRNN are crafted within 1.5 GPU hours, which is 7.2× and 6.7× faster than the crafting time of SOTA CNN and RNN models, respectively. Tunhou Zhang, Hsin-Pai Cheng, Zhenwen Li, Feng Yan 0001, Chengyu Huang 0001, Hai Li 0001, Yiran Chen 0001 |
AAAI | 3 |
| 2020 | Composing Elementary Discourse Units in Abstractive SummarizationabstractIn this paper, we argue that elementary discourse unit (EDU) is a more appropriate textual unit of content selection than the sentence unit in abstractive summarization.To well handle the problem of composing EDUs into an informative and fluent summary, we propose a novel summarization method that first designs an EDU selection model to extract and group informative EDUs and then an EDU fusion model to fuse the EDUs in each group into one sentence.We also design the reinforcement learning mechanism to use EDU fusion results to reward the EDU selection action, boosting the final summarization performance.Experiments on CNN/Daily Mail have demonstrated the effectiveness of our model. Zhenwen Li, Sujian Li |
ACL | 1 |
| 2020 | Benchmarking Meaning Representations in Neural Semantic ParsingabstractMeaning representation is an important component of semantic parsing.Although researchers have designed a lot of meaning representations, recent work focuses on only a few of them.Thus, the impact of meaning representation on semantic parsing is less understood.Furthermore, existing work's performance is often not comprehensively evaluated due to the lack of readily-available execution engines.Upon identifying these gaps, we propose UNIMER, a new unified benchmark on meaning representations, by integrating existing semantic parsing datasets, completing the missing logical forms, and implementing the missing execution engines.The resulting unified benchmark contains the complete enumeration of logical forms and execution engines over three datasets × four meaning representations.A thorough experimental study on UNIMER reveals that neural semantic parsing approaches exhibit notably different performance when they are trained to generate different meaning representations.Also, program alias and grammar rules heavily impact the performance of different meaning representations.Our benchmark, execution engines and implementation can be found on: https Qian Liu 0033, Jian-Guang Lou, Zhenwen Li, Xueqing Liu 0001, Tao Xie 0001, Ting Liu 0002 |
EMNLP (1) | 4 |
| 2020 | Clustering test steps in natural language toward automating test automationabstractFor large industrial applications, system test cases are still often described in natural language (NL), and their number can reach thousands. Test automation is to automatically execute the test cases. Achieving test automation typically requires substantial manual effort for creating executable test scripts from these NL test cases. In particular, given that each NL test case consists of a sequence of NL test steps, testers first implement a test API method for each test step and then write a test script for invoking these test API methods sequentially for test automation. Across different test cases, multiple test steps can share semantic similarities, supposedly mapped to the same API method. However, due to numerous test steps in various NL forms under manual inspection, testers may not realize those semantically similar test steps and thus waste effort to implement duplicate test API methods for them. To address this issue, in this paper, we propose a new approach based on natural language processing to cluster similar NL test steps together such that the test steps in each cluster can be mapped to the same test API method. Our approach includes domain-specific word embedding training along with measurement based on Relaxed Word Mover’sDistance to analyze the similarity of test steps. Our approach also includes a technique to combine hierarchical agglomerative clustering and K-means clustering post-refinement to derive high-quality and manually-adjustable clustering results. The evaluation results of our approach on a large industrial mobile app, WeChat, show that our approach can cluster the test steps with high accuracy, substantially reducing the number of clusters and thus reducing the downstream manual effort. In particular, compared with the baseline approach, our approach achieves 79.8% improvement on cluster quality, reducing 65.9% number of clusters, i.e., the number of test API methods to be implemented. Linyi Li 0001, Zhenwen Li, Guanghua He, Xia Zeng, Yuetang Deng, Tao Xie 0001 |
ESEC/SIGSOFT FSE | 2 |