Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yifan An

dblp:380/3739 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
0009-0000-0257-3628ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 100%

Topics — the 1 heaviest of 1, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution
code clone detection
1.012026
Scalable Large-Scale Multi-Granularity Code Clone Detection via Clustering Search and Pre-Trained Models · IEEE Trans. Software Eng. 2026

Methods — techniques the papers use, named apart from their topics

pre-trained model · 1.0entropy-based filtering · 1.0clustering search · 1.0IVF Flat · 1.0
YearPublicationVenuePosition
2026 Scalable Large-Scale Multi-Granularity Code Clone Detection via Clustering Search and Pre-Trained Models
abstract
Code cloning is a common phenomenon in software development, which reduces developers’ programming efforts but also poses risks of defect inheritance. Clone detection locates exact or similar pieces of code within or between software systems. With the amount of source code increasing steadily, efficient and large-scale clone detection has become a necessity. Moreover, code clones may occur at various levels of code granularity, e.g., file, function, and block level, which pose more challenges for efficient clone detection. Although numerous methods have been proposed to detect code clones at different granularities, they often suffer from low detection efficiency, false positive results and are typically limited to identifying clones at a specific granularity. In this paper, we introduce an efficient clone detection, named MGCD, to detect code clones among large-scale codebases. Specifically, we embed function-level code into vectors using a pre-trained model and perform clustering search with the IVF Flat algorithm to identify clone candidates. These candidates are then filtered through an entropy-based method to enhance accuracy and avoid false positive results. Moreover, we leverage the information from function-level clone detection results to further conduct file and block level clone detection. We evaluate our approach on the BigCloneBench benchmark. Experimental results show that our approach only takes 0.23 ms to search clone results among 800,000 functions and achieves high precision and recall.
Yifan An, Xiang Gao 0012, Hailong Sun 0001
IEEE Trans. Software Eng.1