EDBT 2026 Demo / reviewers in the wild / expert
Weiming Zhang 0004
dblp:20/612-4
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0001-4229-7063ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 33% Software testing · 33% Debugging and program repair · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Debugging and program repair
automated debugging |
0.9 | 1 | 2025 | NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debugging · EMNLP 2025 |
Program synthesis and code generation
code generation with language models |
0.9 | 1 | 2025 | DebateCoder: Towards Collective Intelligence of LLMs via Test Case Driven LLM Debate for Code Generation · ACL (1) 2025 |
Software testing
test generation |
0.9 | 1 | 2025 | DebateCoder: Towards Collective Intelligence of LLMs via Test Case Driven LLM Debate for Code Generation · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.7multi-agent debate · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DebateCoder: Towards Collective Intelligence of LLMs via Test Case Driven LLM Debate for Code GenerationabstractWith the impressive reasoning and text generation capabilities of large language models (LLMs), methods leveraging multiple LLMs to debate each other have garnered increasing attention. However, existing debate-based approaches remain limited in effectiveness in structured and detailed domains represented by code generation due to several reasons: 1) Reliance on different instances of the same LLM for debate, neglecting the potential benefits of integrating diverse models with varied internal knowledge for more comprehensive code generation, 2) under-utilization of test cases, and 3) reliance on third-party LLM moderators for result consolidation and decision-making, probably introducing hallucinations and judgment errors. To address these challenges, we propose DebateCoder to collect intelligence of LLMs via test case-driven debate for code generation. In DebateCoder, test cases serve as a medium for models to analyze code and identify bugs, while opposing models generate test cases to challenge each other’s code during the debate process. These test cases, along with their execution results, are elaborately leveraged to refine and enhance the code through a novel contrastive analysis process. Furthermore, DebateCoder leverages test case outcomes to assess code quality and determine convergence criteria. Unlike previous approaches, DebateCoder emphasizes the collaborative improvement of both models through competitive debate and interactive analysis. Abundant experimental results on two datasets demonstrate the effectiveness of DebateCoder. Jizheng Chen, Kounianhua Du, Xinyi Dai, Weiming Zhang 0004, Xihuai Wang, Yasheng Wang, Ruiming Tang, Weinan Zhang 0001, Yong Yu 0001 |
ACL (1) | 4 |
| 2025 | LLM4CD: Leveraging Large Language Models for Open-World Knowledge Augmented Cognitive DiagnosisabstractCognitive diagnosis (CD) plays a crucial role in intelligent education, evaluating students' comprehension of knowledge concepts based on their test histories. However, current CD methods often model students, exercises, and knowledge concepts solely on their ID relationships, neglecting the abundant semantic relationships present within the educational data space. Furthermore, contemporary intelligent tutoring systems (ITS) frequently involve the addition of new students and exercises, creating cold-start scenarios that ID-based methods find challenging to manage effectively. The advent of large language models (LLMs) offers the potential for overcoming this challenge with open-world knowledge. In this paper, we propose LLM4CD, which Leverages Large Language Models for open-world knowledge Augmented Cognitive Diagnosis. Our method utilizes the open-world knowledge of LLMs to construct cognitively expressive textual representations, which are then encoded to introduce rich semantic information into the CD task. Additionally, we propose an innovative bi-level encoder framework that models students' test histories through two levels of encoders: a macro-level cognitive text encoder and a micro-level knowledge state encoder. This approach substitutes traditional ID embeddings with semantic representations, enabling the model to accommodate new students and exercises with open-world knowledge and address the cold-start problem. Extensive experimental results demonstrate that LLM4CD consistently outperforms previous CD models on multiple real-world datasets, validating the effectiveness of leveraging LLMs to introduce rich semantic information into the CD task. Weiming Zhang 0004, Lingyue Fu, Qingyao Li, Kounianhua Du, Jianghao Lin, Jingwei Yu, Wei Xia 0001, Weinan Zhang 0001, Ruiming Tang, Yong Yu 0001 |
CIKM | 1 |
| 2025 | NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code DebuggingabstractWeiming Zhang, Qingyao Li, Xinyi Dai, Jizheng Chen, Kounianhua Du, Weiwen Liu, Yasheng Wang, Ruiming Tang, Yong Yu, Weinan Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weiming Zhang 0004, Qingyao Li, Xinyi Dai, Jizheng Chen, Kounianhua Du, Weiwen Liu, Yasheng Wang, Ruiming Tang, Yong Yu 0001, Weinan Zhang 0001 |
EMNLP | 1 |
| 2025 | CECAT: Certainty-Guaranteed Environment for Computerized Adaptive TestingabstractComputerized Adaptive Testing (CAT) is a fundamental issue in intelligent education, which aims to select a relatively small number of questions to assess student’s ability. However, existing approaches often suffer from a limited selection area within the question bank, as questions without ground truth answers in the practice log are considered invalid for selection. The selection area refers to the range of questions in the question bank that have answers. To expand the selection area, we proposed a framework called Certainty-Guaranteed Environment for Computerized Adaptive Testing (CECAT). First, Simulative Interaction Module (SIM) leverages questions and their answer results in practice log in training data, so as to predict the distribution of answer result for each question without answer result in practice log. Second, α-Certainty is designed to measure the statistical similarity of the predicted distribution of answer results obtained from SIM, so as to exclude the statistically dissimilar expanded answer results. Extensive experiments on three public datasets demonstrate that CECAT outperforms twelve strong baselines in terms of assessing the student’s ability. Our code is available at https://anonymous.4open.science/r/SimCAT-F31B/README.md. Jingwei Yu, Zhenyu Mu, Weiwen Liu, Ting Long, Weiming Zhang 0004, Weinan Zhang 0001, Yong Yu 0001 |
IJCNN | 5 |