EDBT 2026 Demo / reviewers in the wild / expert
Yifan An
dblp:380/3739
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2026
0009-0000-0257-3628ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Software maintenance and evolution · 100% |
Topics — the 1 heaviest of 1, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software maintenance and evolution
code clone detection |
1.0 | 1 | 2026 | Scalable Large-Scale Multi-Granularity Code Clone Detection via Clustering Search and Pre-Trained Models · IEEE Trans. Software Eng. 2026 |
Methods — techniques the papers use, named apart from their topics
pre-trained model · 1.0entropy-based filtering · 1.0clustering search · 1.0IVF Flat · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable Large-Scale Multi-Granularity Code Clone Detection via Clustering Search and Pre-Trained ModelsabstractCode cloning is a common phenomenon in software development, which reduces developers’ programming efforts but also poses risks of defect inheritance. Clone detection locates exact or similar pieces of code within or between software systems. With the amount of source code increasing steadily, efficient and large-scale clone detection has become a necessity. Moreover, code clones may occur at various levels of code granularity, e.g., file, function, and block level, which pose more challenges for efficient clone detection. Although numerous methods have been proposed to detect code clones at different granularities, they often suffer from low detection efficiency, false positive results and are typically limited to identifying clones at a specific granularity. In this paper, we introduce an efficient clone detection, named MGCD, to detect code clones among large-scale codebases. Specifically, we embed function-level code into vectors using a pre-trained model and perform clustering search with the IVF Flat algorithm to identify clone candidates. These candidates are then filtered through an entropy-based method to enhance accuracy and avoid false positive results. Moreover, we leverage the information from function-level clone detection results to further conduct file and block level clone detection. We evaluate our approach on the BigCloneBench benchmark. Experimental results show that our approach only takes 0.23 ms to search clone results among 800,000 functions and achieves high precision and recall. Yifan An, Xiang Gao 0012, Hailong Sun 0001 |
IEEE Trans. Software Eng. | 1 |