Shuwen Bai

dblp:379/2362 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
0009-0007-6050-172XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Program analysis · 100%
Artificial intelligence
1 paper

Topics — the 1 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program analysis
code representation learning
0.812024
VarGAN: Adversarial Learning of Variable Semantic Representations · IEEE Trans. Software Eng. 2024

Methods — techniques the papers use, named apart from their topics

byte-pair encoding · 1.5adversarial training · 1.5
YearPublicationVenuePosition
2024 VarGAN: Adversarial Learning of Variable Semantic Representations
abstract
Variable names are of critical importance in code representation learning. However, due to diverse naming conventions, variables often receive arbitrary names, leading to long-tail, out-of-vocabulary (OOV), and other well-known problems. While the Byte-Pair Encoding (BPE) tokenizer has addressed the surface-level recognition of low-frequency tokens, it has not noticed the inadequate training of low-frequency identifiers by code representation models, resulting in an imbalanced distribution of rare and common identifiers. Consequently, code representation models struggle to effectively capture the semantics of low-frequency variable names. In this paper, we propose VarGAN, a novel method for variable name representations. VarGAN strengthens the training of low-frequency variables through adversarial training. Specifically, we regard the code representation model as a generator responsible for producing vectors from source code. Additionally, we employ a discriminator that detects whether the code input to the generator contains low-frequency variables. This adversarial setup regularizes the distribution of rare variables, making them overlap with their corresponding high-frequency counterparts in the vector space. Experimental results demonstrate that VarGAN empowers CodeBERT to generate code vectors that exhibit more uniform distribution for both low- and high-frequency identifiers. There is an improvement of 8% in similarity and relatedness scores compared to VarCLR in the IdBench benchmark. VarGAN is also validated in downstream tasks, where it exhibits enhanced capabilities in capturing token- and code-level semantics.
Yalan Lin, Chengcheng Wan 0001, Shuwen Bai, Xiaodong Gu 0002
IEEE Trans. Software Eng.3