Yanxian Huang

dblp:331/8233 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Agents in software engineering: survey, landscape, and vision
Yanlin Wang 0001, Wanjun Zhong, Yanxian Huang, Ensheng Shi, Min Yang 0002, Jiachi Chen, Hui Li 0057, Yuchi Ma, Qianxiang Wang, Zibin Zheng
Autom. Softw. Eng.3
2024 SparseCoder: Identifier-Aware Sparse Transformer for File- Level Code Summarization
abstract
Code summarization aims to generate natural language descriptions of source code, facilitating programmers to understand and maintain it rapidly. While previous code summarization efforts have predominantly focused on method-level, this paper studies file-level code summarization, which can assist programmers in understanding and maintaining large source code projects. Unlike method-level code summarization, file-level code summarization typically involves long source code within a single file, which makes it challenging for Transformer-based models to understand the code semantics for the maximum input length of these models is difficult to set to a large number that can handle long code input well, due to the quadratic scaling of computational complexity with the input sequence length. To address this challenge, we propose SparseCoder, an identifier-aware sparse transformer for effectively handling long code sequences. Specifically, the SparseCoder employs a sliding window mechanism for self-attention to model short-term dependencies and leverages the structure message of code to capture long-term dependencies among source code identifiers by introducing two types of sparse attention patterns named global and identifier attention. To evaluate the performance of SparseCoder, we construct a new dataset FILE-CS for file-level code summarization in Python. Experimental results show that our SparseCoder model achieves state-of-the-art performance compared with other pre-trained models, including full self-attention and sparse models. Additionally, our model has low memory overhead and achieves comparable performance with models using full self-attention mechanism. Furthermore, we verify the generality of SparseCoder on other code understanding tasks, i.e., code clone detection and code search, and results show that our model outperforms baseline models in both tasks, demonstrating that our model can generate better code representations for various downstream tasks. Our source code and experimental data are anonymously available at: https://github.com/DeepSoftwareAnalytics/SparseCoder.
Yanlin Wang 0001, Yanxian Huang, Daya Guo, Hongyu Zhang 0002, Zibin Zheng
SANER2
2022 Exploring Representation-level Augmentation for Code Search
abstract
Code search, which aims at retrieving the most relevant code fragment for a given natural language query, is a common activity in software development practice.Recently, contrastive learning is widely used in code search research, where many data augmentation approaches for source code (e.g., semantic-preserving program transformation) are proposed to learn better representations.However, these augmentations are at the raw-data level, which requires additional code analysis in the preprocessing stage and additional training costs in the training stage.In this paper, we explore augmentation methods that augment data (both code and query) at representation level which does not require additional data processing and training, and based on this we propose a general format of representationlevel augmentation that unifies existing methods.Then, we propose three new augmentation methods (linear extrapolation, binary interpolation, and Gaussian scaling) based on the general format.Furthermore, we theoretically analyze the advantages of the proposed augmentation methods over traditional contrastive learning methods on code search.We experimentally evaluate the proposed representationlevel augmentation methods with state-of-theart code search models on a large-scale public dataset consisting of six programming languages.The experimental results show that our approach can consistently boost the performance of the studied code search models.
Haochen Li 0009, Chunyan Miao, Cyril Leung, Yanxian Huang, Yuan Huang 0002, Hongyu Zhang 0002, Yanlin Wang 0001
EMNLP4