VLDB 2026 Research / reviewers in the wild / expert
Daoguang Zan
dblp:305/5798
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2025
0009-0009-4269-8543ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsabstractRecent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities.
However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning. Bofei Gao, Feifan Song 0001, Zhe Yang 0013, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li 0039, Chenghao Ma, Liang Chen 0024, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang 0009, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu 0001, Baobao Chang |
ICLR | 13 |
| 2025 | Multi-SWE-bench: A Multilingual Benchmark for Issue ResolvingabstractThe task of issue resolving aims to modify a codebase to generate a patch that addresses a given issue. However, most existing benchmarks focus almost exclusively on Python, making them insufficient for evaluating Large Language Models (LLMs) across different programming languages. To bridge this gap, we introduce a multilingual issue-resolving benchmark, called Multi-SWE-bench, covering 8 languages of Python, Java, TypeScript, JavaScript, Go, Rust, C, and C++. In particular, this benchmark includes a total of 2,132 high-quality instances, carefully curated by 68 expert annotators, ensuring a reliable and accurate evaluation of LLMs on the issue-resolving task. Based on human-annotated results, the issues are further classified into three difficulty levels. We evaluate a series of state-of-the-art models on Multi-SWE-bench, utilizing both procedural and agent-based frameworks for issue resolving. Our experiments reveal three key findings: (1) Limited generalization across languages: While existing LLMs perform well on Python issues, their ability to generalize across other languages remains limited; (2) Performance aligned with human-annotated difficulty: LLM-based agents' performance closely aligns with human-assigned difficulty, with resolution rates decreasing as issue complexity rises; and (3) Performance drop on cross-file issues: The performance of current methods significantly deteriorates when handling cross-file issues. These findings highlight the limitations of current LLMs and underscore the need for more robust models capable of handling a broader range of programming languages and complex issue scenarios. Daoguang Zan, Zhirong Huang, Hanwu Chen, Shulin Xin, Linhao Zhang, Aoyan Li, Xiaojian Zhong, Yongsheng Xiao, Liangqiang Chen, Yuyu Zhang, Rui Long |
NeurIPS | 1 |
| 2025 | Private-library-oriented code generation with large language models
Daoguang Zan, Bei Chen 0008, Yongshun Gong, Junzhi Cao, Fengji Zhang, Bingchao Wu, Bei Guan, Yilong Yin, Yongji Wang 0002 |
Knowl. Based Syst. | 1 |
| 2024 | A GAN-Based Data Poisoning Framework Against Anomaly Detection in Vertical Federated LearningabstractIn vertical federated learning (VFL), commercial entities collaboratively train a model while preserving data privacy. However, a malicious participant's poisoning attack may degrade the performance of this collaborative model. The main challenge in achieving the poisoning attack is the absence of access to the server-side top model, leaving the malicious par-ticipant without a clear target model. To address this challenge, we introduce an innovative end-to-end poisoning framework P-GAN. Specifically, the malicious participant initially employs semi-supervised learning to train a surrogate target model. Subsequently, this participant employs a GAN-based method to produce adversarial perturbations to degrade the surrogate target model's performance. Finally, the generator is obtained and tailored for VFL poisoning. Besides, we develop an anomaly detection algorithm based on a deep auto-encoder (DAE), offering a robust defense mechanism to VFL scenarios. Through extensive experiments, we evaluate the efficacy of P-GAN and DAE, and further analyze the factors that influence their performance. Daoguang Zan, Wei Li 0326, Bei Guan, Yongji Wang 0002 |
ICC | 2 |
| 2024 | The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language ModelsabstractPre-trained Language models (PLMs) have been acknowledged to contain harmful information, such as social biases, which may cause negative social impacts or even bring catastrophic results in application. Previous works on this problem mainly focused on using black-box methods such as probing to detect and quantify social biases in PLMs by observing model outputs. As a result, previous debiasing methods mainly finetune or even pre-train PLMs on newly constructed anti-stereotypical datasets, which are high-cost. In this work, we try to unveil the mystery of social bias inside language models by introducing the concept of {\sc Social Bias Neurons}. Specifically, we propose {\sc Integrated Gap Gradients (IG$^2$)} to accurately pinpoint units (i.e., neurons) in a language model that can be attributed to undesirable behavior, such as social bias. By formalizing undesirable behavior as a distributional property of language, we employ sentiment-bearing prompts to elicit classes of sensitive words (demographics) correlated with such sentiments. Our IG$^2$ thus attributes the uneven distribution for different demographics to specific Social Bias Neurons, which track the trail of unwanted behavior inside PLM units to achieve interoperability. Moreover, derived from our interpretable technique, {\sc Bias Neuron Suppression (BNS)} is further proposed to mitigate social biases. By studying BERT, RoBERTa, and their attributable differences from debiased FairBERTa, IG$^2$ allows us to locate and suppress identified neurons, and further mitigate undesired behaviors. As measured by prior metrics from StereoSet, our model achieves a higher degree of fairness while maintaining language modeling ability with low cost\footnote{This work contains examples that potentially implicate stereotypes, associations, and other harms that could be offensive to individuals in certain social groups.}. Yan Liu 0002, Xiaokang Chen, Daoguang Zan, Min-Yen Kan, Tsung-Yi Ho |
ICLR | 5 |
| 2024 | FIA-TE: Feature Inference Attack on Decision Tree Ensembles in Vertical Federated LearningabstractVertical federated learning (VFL) enables multiple parties to collaboratively train a model while preserving privacy. However, recent studies have raised concerns about the susceptibility of VFL models, including those using logistic regression and neural networks, to feature inference attacks. Meanwhile, the non-differentiable characteristics of decision tree ensembles make conducting such attacks impractical. To address this challenge, we introduce a feature inference attack framework FIA-TE tailored for decision tree ensembles, including gradient boosted decision trees (GBDT) and random forest. Specifically, we distill the knowledge from trees into neural networks by leaf embedding and structure distillation to create a targeted model for the inference attack. We then employ a generative model based on the deconvolutional network for capturing correlation features and reconstructing the target features. Through extensive experiments on table and image data, we evaluate the effectiveness of our framework and provide an analysis of potential influencing factors. Daoguang Zan, Wei Li 0326, Bei Guan, Yongji Wang 0002 |
ICME | 2 |
| 2024 | GraphCoder: Enhancing Repository-Level Code Completion via Coarse-to-fine Retrieval Based on Code Context GraphabstractThe performance of repository-level code completion depends upon the effective leverage of both general and repository-specific knowledge. Despite the impressive capability of code LLMs in general code completion tasks, they often exhibit less satisfactory performance on repository-level completion due to the lack of repository-specific knowledge in these LLMs. To address this problem, we propose GraphCoder, a retrieval-augmented code completion framework that leverages LLMs' general code knowledge and the repository-specific knowledge via a graph-based retrieval-generation process. In particular, GraphCoder captures the context of completion target more accurately through code context graph (CCG) that consists of control-flow, data- and control-dependence between code statements, a more structured way to capture the completion target context than the sequence-based context used in existing retrieval-augmented approaches; based on CCG, GraphCoder further employs a coarse-to-fine retrieval process to locate context-similar code snippets with the completion target from the current repository. Experimental results demonstrate both the effectiveness and efficiency of GraphCoder: Compared to baseline retrieval-augmented methods, GraphCoder achieves higher exact match (EM) on average, with increases of +6.06 in code match and +6.23 in identifier match, while using less time and space. Wei Liu 0189, Ailun Yu, Daoguang Zan, Wei Zhang 0004, Haiyan Zhao 0001, Zhi Jin 0001, Qianxiang Wang |
ASE | 3 |
| 2023 | Large Language Models Meet NL2Code: A SurveyabstractDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Wang Yongji, Jian-Guang Lou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Daoguang Zan, Bei Chen 0008, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang 0002, Jian-Guang Lou |
ACL (1) | 1 |
| 2023 | RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationabstractFengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, Weizhu Chen. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Fengji Zhang, Bei Chen 0008, Jacky W. Keung, Jin Liu 0016, Daoguang Zan, Jian-Guang Lou, Weizhu Chen |
EMNLP | 6 |
| 2023 | CodeT: Code Generation with Generated Tests
Bei Chen 0008, Fengji Zhang, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, Weizhu Chen |
ICLR | 4 |
| 2023 | Hierarchical and Contrastive Representation Learning for Knowledge-Aware RecommendationabstractIncorporating knowledge graph into recommendation is an effective way to alleviate data sparsity. Most existing knowledge-aware methods usually perform recursive embedding propagation by enumerating graph neighbors. However, the number of nodes’ neighbors grows exponentially as the hop number increases, forcing the nodes to be aware of vast neighbors under this recursive propagation for distilling the high-order semantic relatedness. This may induce more harmful noise than useful information into recommendation, leading the learned node representations to be indistinguishable from each other, that is, the well-known over-smoothing issue. To relieve this issue, we propose a Hierarchical and CONtrastive representation learning framework for knowledge-aware recommendation named HiCON. Specifically, for avoiding the exponential expansion of neighbors, we propose a hierarchical message aggregation mechanism to interact separately with low-order neighbors and meta-path-constrained high-order neighbors. Moreover, we also perform cross-order contrastive learning to enforce the representations to be more discriminative. Extensive experiments on three datasets show the remarkable superiority of HiCON over state-of-the-art approaches. The code is available now1. Bingchao Wu, Yangyuxuan Kang, Daoguang Zan, Bei Guan, Yongji Wang 0002 |
ICME | 3 |
| 2023 | Uncovering and Quantifying Social Biases in Code GenerationabstractWith the popularity of automatic code generation tools, such as Copilot, the study of the potential hazards of these tools is gaining importance. In this work, we explore the social bias problem in pre-trained code generation models. We propose a new paradigm to construct code prompts and successfully uncover social biases in code generation models. To quantify the severity of social biases in generated code, we develop a dataset along with three metrics to evaluate the overall social bias and fine-grained unfairness across different demographics. Experimental results on three pre-trained code generation models (Codex, InCoder, and CodeGen) with varying sizes, reveal severe social biases. Moreover, we conduct analysis to provide useful insights for further choice of code generation models with low social bias. Yan Liu 0002, Xiaokang Chen, Yan Gao 0002, Fengji Zhang, Daoguang Zan, Jian-Guang Lou, Tsung-Yi Ho |
NeurIPS | 6 |
| 2022 | CERT: Continual Pre-training on Sketches for Library-oriented Code GenerationabstractCode generation is a longstanding challenge, aiming to generate a code snippet based on a natural language description. Usually, expensive text-code paired data is essential for training a code generation model. Recently, thanks to the success of pre-training techniques, large language models are trained on large unlabelled code corpora and perform well in generating code. In this paper, we investigate how to leverage an unlabelled code corpus to train a model for library-oriented code generation. Since it is a common practice for programmers to reuse third-party libraries, in which case the text-code paired data are harder to obtain due to the huge number of libraries. We observe that library-oriented code snippets are more likely to share similar code sketches. Hence, we present CERT with two steps: a sketcher generates the sketch, then a generator fills the details in the sketch. Both the sketcher and generator are continually pre-trained upon a base model using unlabelled data. Also, we carefully craft two benchmarks to evaluate library-oriented code generation named PandasEval and NumpyEval. Experimental results have shown the impressive performance of CERT. For example, it surpasses the base model by an absolute 15.67% improvement in terms of pass@1 on PandasEval. Our work is available at https://github.com/microsoft/PyCodeGPT. Daoguang Zan, Bei Chen 0008, Dejian Yang, Zeqi Lin, Bei Guan, Yongji Wang 0002, Weizhu Chen, Jian-Guang Lou |
IJCAI | 1 |
| 2022 | Complex Question Answering over Incomplete Knowledge Graph as N-ary Link PredictionabstractThe Question Answering over Knowledge Graph (KGQA) task seeks entities (answers) from the Knowledge Graph (KG) in order to answer natural language questions. In practice, KG is often incomplete, with numerous missing links and nodes. With such an incomplete KG, it is tricky to use the semantics inside the KG to get the golden answers, particularly for complex questions. Some current efforts concentrate on using external corpora to overcome KG sparsity; however, identifying and obtaining the corpora is challenging. Other types of work aim to leverage the pre-trained embeddings to resolve the issue but perform slightly worse on complex questions involving numerous triple facts in KG. To address the aforementioned problems, we present a framework CAPKGQA, which transforms Complex KGQA into an n-Ary link Prediction task capable of explicitly modeling complex questions. Furthermore, previous methods also suffer from incomplete KG throughout the candidate answer generation phase. Therefore, we devise an embedding-based retrieval strategy to extract more reliable candidate answers from incomplete KG. Extensive experiments reveal that our approach beats the state-of-the-art models on incomplete and complex KGQA tasks by a significant margin. Daoguang Zan, Kun Zhou 0002, Wei Wu 0014, Wayne Xin Zhao, Bingchao Wu, Bei Guan, Yongji Wang 0002 |
IJCNN | 1 |
| 2022 | S2QL: Retrieval Augmented Zero-Shot Question Answering over Knowledge Graph
Daoguang Zan, Yuanmeng Yan, Wei Wu 0014, Bei Guan, Yongji Wang 0002 |
PAKDD (3) | 1 |
| 2021 | Large-Scale Relation Learning for Question Answering over Knowledge Bases with Pre-trained Language ModelsabstractThe key challenge of question answering over knowledge bases (KBQA) is the inconsistency between the natural language questions and the reasoning paths in the knowledge base (KB).Recent graph-based KBQA methods are good at grasping the topological structure of the graph but often ignore the textual information carried by the nodes and edges.Meanwhile, pre-trained language models learn massive open-world knowledge from the large corpus, but it is in the natural language form and not structured.To bridge the gap between the natural language and the structured KB, we propose three relation learning tasks for BERTbased KBQA, including relation extraction, relation matching, and relation reasoning.By relation-augmented training, the model learns to align the natural language expressions to the relations in the KB as well as reason over the missing connections in the KB.Experiments on WebQSP show that our method consistently outperforms other baselines, especially when the KB is incomplete. Yuanmeng Yan, Daoguang Zan, Wei Wu 0014, Weiran Xu |
EMNLP (1) | 5 |