Yangkai Du

dblp:260/9581 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
3 papers
Program analysis · 70% Program synthesis and code generation · 16% Software maintenance and evolution · 14%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program analysis › program comparison
code similarity
0.912025
Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting · AAAI 2025
Program analysis
source code analysis
0.912025
Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting · AAAI 2025
Software maintenance and evolution
code clone detection
0.812024
AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone Detection · AAAI 2024
Program analysis
program representation
0.812024
AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone Detection · AAAI 2024
Program analysis
binary analysis
0.712023
CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code · EMNLP 2023
Program analysis › program representation
control flow graph
0.712023
CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

self-supervised contrastive learning · 0.9code rewriting · 0.9knowledge transfer · 0.8contrastive learning · 0.8pseudo code · 0.7bidirectional instruction-level control flow graph · 0.7
YearPublicationVenuePosition
2025 Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
abstract
Large Language Models (LLMs) have demonstrated remarkable proficiency in generating code. However, the misuse of LLM-generated (synthetic) code has raised concerns in both educational and industrial contexts, underscoring the urgent need for synthetic code detectors. Existing methods for detecting synthetic content are primarily designed for general text and struggle with code due to the unique grammatical structure of programming languages and the presence of numerous ``low-entropy'' tokens. Building on this, our work proposes a novel zero-shot synthetic code detector based on the similarity between the original code and its LLM-rewritten variants. Our method is based on the observation that differences between LLM-rewritten and original code tend to be smaller when the original code is synthetic. We utilize self-supervised contrastive learning to train a code similarity model and evaluate our approach on two synthetic code detection benchmarks. Our results demonstrate a significant improvement over existing SOTA synthetic content detectors, delivering notable gains in both performance and robustness on the APPS and MBPP benchmarks.
Yangkai Du, Tengfei Ma 0001, Lingfei Wu 0001, Xuhong Zhang 0002, Shouling Ji, Wenhai Wang
AAAI2
2024 AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone Detection
abstract
Code Clone Detection, which aims to retrieve functionally similar programs from large code bases, has been attracting increasing attention. Modern software often involves a diverse range of programming languages. However, current code clone detection methods are generally limited to only a few popular programming languages due to insufficient annotated data as well as their own model design constraints. To address these issues, we present AdaCCD, a novel cross-lingual adaptation method that can detect cloned codes in a new language without annotations in that language. AdaCCD leverages language-agnostic code representations from pre-trained programming language models and propose an Adaptively Refined Contrastive Learning framework to transfer knowledge from resource-rich languages to resource-poor languages. We evaluate the cross-lingual adaptation results of AdaCCD by constructing a multilingual code clone detection benchmark consisting of 5 programming languages. AdaCCD achieves significant improvements over other baselines, and achieve comparable performance to supervised fine-tuning.
Yangkai Du, Tengfei Ma 0001, Lingfei Wu 0001, Xuhong Zhang 0002, Shouling Ji
AAAI1
2023 CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code
abstract
Automatically generating function summaries for binaries is an extremely valuable but challenging task, since it involves translating the execution behavior and semantics of the low-level language (assembly code) into human-readable natural language.However, most current works on understanding assembly code are oriented towards generating function names, which involve numerous abbreviations that make them still confusing.To bridge this gap, we focus on generating complete summaries for binary functions, especially for stripped binary (no symbol table and debug information in reality).To fully exploit the semantics of assembly code, we present a control flow graph and pseudo code guided binary code summarization framework called CP-BCS.CP-BCS utilizes a bidirectional instruction-level control flow graph and pseudo code that incorporates expert knowledge to learn the comprehensive binary function execution behavior and logic semantics.We evaluate CP-BCS on 3 different binary optimization levels (O1, O2, and O3) for 3 different computer architectures (X86, X64, and ARM).The evaluation results demonstrate CP-BCS is superior and significantly improves the efficiency of reverse engineering. * Corresponding author.with limited high-level information, making it difficult to read and understand, as shown in Figure 1.Even an experienced reverse engineer needs to spend a significant amount of time determining the functionality of an assembly code snippet.
Lingfei Wu 0001, Tengfei Ma 0001, Xuhong Zhang 0002, Yangkai Du, Peiyu Liu 0003, Shouling Ji, Wenhai Wang
EMNLP5
2021 Structured Self-Supervised Pretraining for Commonsense Knowledge Graph Completion
abstract
Abstract To develop commonsense-grounded NLP applications, a comprehensive and accurate commonsense knowledge graph (CKG) is needed. It is time-consuming to manually construct CKGs and many research efforts have been devoted to the automatic construction of CKGs. Previous approaches focus on generating concepts that have direct and obvious relationships with existing concepts and lack an capability to generate unobvious concepts. In this work, we aim to bridge this gap. We propose a general graph-to-paths pretraining framework that leverages high-order structures in CKGs to capture high-order relationships between concepts. We instantiate this general framework to four special cases: long path, path-to-path, router, and graph-node-path. Experiments on two datasets demonstrate the effectiveness of our methods. The code will be released via the public GitHub repository.
Jiayuan Huang, Yangkai Du, Shuting Tao, Pengtao Xie
Trans. Assoc. Comput. Linguistics2