VLDB 2026 Research / reviewers in the wild / expert
Yangkai Du
dblp:260/9581
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
3 papers |
Program analysis · 70% Program synthesis and code generation · 16% Software maintenance and evolution · 14% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis › program comparison
code similarity |
0.9 | 1 | 2025 | Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting · AAAI 2025 |
Program analysis
source code analysis |
0.9 | 1 | 2025 | Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting · AAAI 2025 |
Software maintenance and evolution
code clone detection |
0.8 | 1 | 2024 | AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone Detection · AAAI 2024 |
Program analysis
program representation |
0.8 | 1 | 2024 | AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone Detection · AAAI 2024 |
Program analysis
binary analysis |
0.7 | 1 | 2023 | CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code · EMNLP 2023 |
Program analysis › program representation
control flow graph |
0.7 | 1 | 2023 | CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
self-supervised contrastive learning · 0.9code rewriting · 0.9knowledge transfer · 0.8contrastive learning · 0.8pseudo code · 0.7bidirectional instruction-level control flow graph · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code RewritingabstractLarge Language Models (LLMs) have demonstrated remarkable proficiency in generating code. However, the misuse of LLM-generated (synthetic) code has raised concerns in both educational and industrial contexts, underscoring the urgent need for synthetic code detectors. Existing methods for detecting synthetic content are primarily designed for general text and struggle with code due to the unique grammatical structure of programming languages and the presence of numerous ``low-entropy'' tokens. Building on this, our work proposes a novel zero-shot synthetic code detector based on the similarity between the original code and its LLM-rewritten variants. Our method is based on the observation that differences between LLM-rewritten and original code tend to be smaller when the original code is synthetic. We utilize self-supervised contrastive learning to train a code similarity model and evaluate our approach on two synthetic code detection benchmarks. Our results demonstrate a significant improvement over existing SOTA synthetic content detectors, delivering notable gains in both performance and robustness on the APPS and MBPP benchmarks. Yangkai Du, Tengfei Ma 0001, Lingfei Wu 0001, Xuhong Zhang 0002, Shouling Ji, Wenhai Wang |
AAAI | 2 |
| 2024 | AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone DetectionabstractCode Clone Detection, which aims to retrieve functionally similar programs from large code bases, has been attracting increasing attention. Modern software often involves a diverse range of programming languages. However, current code clone detection methods are generally limited to only a few popular programming languages due to insufficient annotated data as well as their own model design constraints. To address these issues, we present AdaCCD, a novel cross-lingual adaptation method that can detect cloned codes in a new language without annotations in that language. AdaCCD leverages language-agnostic code representations from pre-trained programming language models and propose an Adaptively Refined Contrastive Learning framework to transfer knowledge from resource-rich languages to resource-poor languages. We evaluate the cross-lingual adaptation results of AdaCCD by constructing a multilingual code clone detection benchmark consisting of 5 programming languages. AdaCCD achieves significant improvements over other baselines, and achieve comparable performance to supervised fine-tuning. Yangkai Du, Tengfei Ma 0001, Lingfei Wu 0001, Xuhong Zhang 0002, Shouling Ji |
AAAI | 1 |
| 2023 | CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo CodeabstractAutomatically generating function summaries for binaries is an extremely valuable but challenging task, since it involves translating the execution behavior and semantics of the low-level language (assembly code) into human-readable natural language.However, most current works on understanding assembly code are oriented towards generating function names, which involve numerous abbreviations that make them still confusing.To bridge this gap, we focus on generating complete summaries for binary functions, especially for stripped binary (no symbol table and debug information in reality).To fully exploit the semantics of assembly code, we present a control flow graph and pseudo code guided binary code summarization framework called CP-BCS.CP-BCS utilizes a bidirectional instruction-level control flow graph and pseudo code that incorporates expert knowledge to learn the comprehensive binary function execution behavior and logic semantics.We evaluate CP-BCS on 3 different binary optimization levels (O1, O2, and O3) for 3 different computer architectures (X86, X64, and ARM).The evaluation results demonstrate CP-BCS is superior and significantly improves the efficiency of reverse engineering. * Corresponding author.with limited high-level information, making it difficult to read and understand, as shown in Figure 1.Even an experienced reverse engineer needs to spend a significant amount of time determining the functionality of an assembly code snippet. Lingfei Wu 0001, Tengfei Ma 0001, Xuhong Zhang 0002, Yangkai Du, Peiyu Liu 0003, Shouling Ji, Wenhai Wang |
EMNLP | 5 |
| 2021 | Structured Self-Supervised Pretraining for Commonsense Knowledge Graph CompletionabstractAbstract To develop commonsense-grounded NLP applications, a comprehensive and accurate commonsense knowledge graph (CKG) is needed. It is time-consuming to manually construct CKGs and many research efforts have been devoted to the automatic construction of CKGs. Previous approaches focus on generating concepts that have direct and obvious relationships with existing concepts and lack an capability to generate unobvious concepts. In this work, we aim to bridge this gap. We propose a general graph-to-paths pretraining framework that leverages high-order structures in CKGs to capture high-order relationships between concepts. We instantiate this general framework to four special cases: long path, path-to-path, router, and graph-node-path. Experiments on two datasets demonstrate the effectiveness of our methods. The code will be released via the public GitHub repository. Jiayuan Huang, Yangkai Du, Shuting Tao, Pengtao Xie |
Trans. Assoc. Comput. Linguistics | 2 |