VLDB 2026 Research / reviewers in the wild / expert
Haiyang Yang
dblp:283/7179
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMUpdater: Automatic comment synchronization via edit model guided LLMs
Haiyang Yang, Qi Xie 0010, Shixiang Cai, Li Kuang, Yingjie Xia |
Empir. Softw. Eng. | 1 |
| 2025 | TACLR: A Scalable and Efficient Retrieval-based Method for Industrial Product Attribute Value IdentificationabstractYindu Su, Huike Zou, Lin Sun, Ting Zhang, Haiyang Yang, Chen Li Yu, David Lo, Qingheng Zhang, Shuguang Han, Jufeng Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yindu Su, Huike Zou, Ting Zhang 0011, Haiyang Yang, Chen Li Yu, David Lo 0001, Qingheng Zhang, Shuguang Han, Jufeng Chen |
ACL (1) | 5 |
| 2025 | APIMig: A Project-Level Cross-Multi-Version API Migration Framework Based on Evolution Knowledge GraphabstractAPI migration is essential for software maintenance due to the rapid evolution of third-party libraries where API elements may change continuously through updates. There are two main challenges for API migration at the project level, especially across multiple versions: 1) lack of specific library evolution knowledge across multi-version; 2) difficulty in identifying the chain of changes at the project level. This paper proposes a project-level cross-multi-version API migration framework APIMig. We first construct an API evolution knowledge graph (KG) to capture changes between adjacent library versions and then derive coherent cross-version API evolution knowledge by KG reasoning. Second, we design a chain exploration algorithm to track the chain of changes and aggregate the affected code segments. Finally, a large language model is employed in completing API migration by providing the API evolution knowledge and the chain of changes. We construct an evolution KG for the Lucene library from version 4.0.0 to 10.1.0 and evaluate our approach through project migration pairs that depend on different major versions. Our framework shows improvements over the baseline in migrating projects across 7 major versions, achieving average increases of 16.52% in CodeBLEU scores and 28.49% in VCEU scores in GPT-4o. Li Kuang, Qi Xie 0010, Haiyang Yang, HaoYue Kang, Yingjie Xia |
IJCAI | 3 |
| 2025 | DLCoG: A Novel Framework for Dual-Level Code Comment Generation Based on Semantic Segmentation and In-Context LearningabstractIn large software projects with collaborative development, comprehensive code comments are crucial for code readability and maintainability. Code comments mainly include method comments and inline comments, where the former describes the functionality globally, and the latter describes the implementation details locally. Existing methods typically generate these two kinds of comments with specific locations independently, which results in weak correlations between comments and code context, as well as high model inference costs due to long token inputs. To address these issues, we define the combination of inline comments and method comments as Dual-Level Code Comments. We formulate the novel task of automatically generate dual-level code comments based on given code and propose an approach named DLCoG (DualLevel Code Comment Generation) to automate this task. First, a Semantic Segmentation and Identification multi-task model based on CodeBERT, termed Se2Iden (Semantic Segmentation and Identification model), is proposed to identify code segments requiring inline comments. Next, we retrieve similar samples to adopting the in-context learning paradigm, which can enhance the generation quality of large language models (LLMs) in specific domains. Finally, the LLM is guided to generate duallevel code comments using Chain-of-Thought (CoT) prompts that first produce inline comments, followed by method comments. We manually constructed a high-quality clean Java dataset consisting of*> based on open-source Java projects by (i) determining comments type and (ii) manually associating inline comments with their corresponding code. Then, we trained a multi-task learning model based on CodeBERT to automatically take the two steps needed, termed ICSA (Inline Comment Classification and Scope Association), thus to expand to a dataset containing 80k dual-level code comments. Experimental results on clean and extended datasets show that DLCoG outperforms all baselines by substantial margins. The contextual information provided by DLCoG can effectively improve the inline comments generated by LLM. Coordinated generation of dual-level comment also brings effective improvements to method comments, which is particularly significant when there are few contextual examples. Our work fills the long-standing gap in the dual-level code comment generation field, and can provide insights for future research in this direction. We provide open-source datasets and source code for future research. Haiyang Yang, Qingyang Yan, Weihuan Min, Zhao Wei, Li Kuang, Yingjie Xia |
ICPC | 2 |
| 2025 | ChatDL: An LLM-Based Defect Localization Approach for Software in IIoT Flexible ManufacturingabstractWith the rapid advancement of flexible manufacturing in the Industrial Internet of Things (IIoT), there has been a significant increase in the number of IIoT devices and application software aimed at meeting various needs. The software defects may lead to delays or crashes in flexible manufacturing system, thereby affecting the production schedule. Automated software defect localization based on code changes can significantly reduce development and maintenance time costs, thereby maintaining the competitive edge of flexible manufacturing in the IIoT. Current efforts in software defect localization are primarily based on deep learning models or information retrieval models. This article investigates the performance of large language models (LLMs) in software defect localization and optimizes localization accuracy by combining it with an information retrieval model. Our empirical study reveals that GPT, given a software defect description, is unable to determine whether specific code changes are relevant. The model is unable to provide accurate answers, which aligns with the generative nature of LLMs where responses are generated according to probability distributions. However, the combined framework of LLMs and information retrieval models proposed in this article outperforms the current state-of-the-art models on public datasets. We conclude that LLMs can enhance localization performance when used as side information in conjunction with existing information retrieval models. The effectiveness of the framework has been validated through experiments conducted on publicly available datasets and in practical applications within IIoT projects. This offers valuable insights into the application and development of LLMs for defect localization in the software development and maintenance processes in the IIoT flexible manufacturing. Haiyang Yang, Yulu Zhou, Li Kuang |
IEEE Internet Things J. | 1 |
| 2024 | ASKDetector: An AST-Semantic and Key Features Fusion based Code Comment Mismatch DetectorabstractCode comments are essential for programming comprehension. Nevertheless, developers often neglect to update comments after modifying the source code. Wrong code comments may lead to bugs in the maintenance process, thus affecting the reliability of the software. So, timely comment mismatch detection is crucial for software development and maintenance. However, existing works have the following two limitations: 1) the lack of use of code structural and sequential information, and 2) the ignorance of existing associations between code and comments. In this paper, we propose a new model called ASKDetector (AST-Semantic and Key features fusion based mismatch Detector). For the first limitation, we encode code with an attention-based preorder traversal abstract syntax tree sequence to obtain both order and structural information. And CodeBERT is utilized to capture contextual semantic features further. For the second one, we encode extracted association information between the code snippets and comments to reduce the semantic gap. The correlations between the encoders are learned through a fusion layer and a multi-layer perceptron. The experimental results prove that our detector outperforms the state-of-the-art model in evaluation metrics, where our F1 and accuracy exceed an average of 3.4%. Haiyang Yang, Hao Chen 0116, Zhirui Kuai, Shuyuan Tu, Li Kuang |
ICPC | 1 |
| 2024 | Generation, augmentation, and alignment: a pseudo-source domain based method for source-free domain adaptation
Yuntao Du 0001, Haiyang Yang, Mingcai Chen, Hongtao Luo, Juan Jiang, Yi Xin 0003, Chong-Jun Wang |
Mach. Learn. | 2 |
| 2023 | HumanBench: Towards General Human-Centric Perception with Projector Assisted PretrainingabstractHuman-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this path from the aspects of both benchmark and pretraining methods. Specifically, we propose a HumanBench based on existing datasets to comprehensively evaluate on the common ground the generalization abilities of different pretraining methods on 19 datasets from 6 diverse downstream tasks, including person ReID, pose estimation, human parsing, pedestrian attribute recognition, pedestrian detection, and crowd counting. To learn both coarse-grained and fine-grained knowledge in human bodies, we further propose a Projector AssisTed Hierarchical pretraining method (PATH) to learn diverse knowledge at different granularity levels. Comprehensive evaluations on HumanBench show that our PATH achieves new state-of-the-art results on 17 downstream datasets and on-par results on the other 2 datasets. The code will be publicly at https://github.com/OpenGVLab/HumanBench. Shixiang Tang, Qingsong Xie, Meilin Chen, Yizhou Wang 0007, Yuanzheng Ci, Lei Bai 0001, Feng Zhu 0006, Haiyang Yang, Rui Zhao 0001, Wanli Ouyang |
CVPR | 9 |
| 2023 | Cycle-consistent Masked AutoEncoder for Unsupervised Domain Generalization
Haiyang Yang, Shixiang Tang, Feng Zhu 0006, Yizhou Wang 0007, Meilin Chen, Lei Bai 0001, Rui Zhao 0001, Wanli Ouyang |
ICLR | 1 |
| 2022 | Retrieve-Guided Commit Message Generation with Semantic Similarity And DisparityabstractHigh quality commit messages are important for program understanding and maintenance, which describes the content of code changes. Neural-based methods are most popular ways to generate commit messages, but they neglect retrieval results including retrieval diffs and retrieval messages. The existing combination models of neural-based methods and retrievalbased methods have two major limitations: a) Only use retrieval diffs but ignore retrieval messages. b) Seldom consider the similarity and disparity between the retrieval results and given diff. To address the above two issues, we propose a retrieveguided method named ReGenSD to generate commit messages, which consists of three steps. Firstly, we apply a similaritybased IR technique to get retrieval diff and retrieval messages. Secondly, we introduce a selective mechanism to decide whether to use retrieve-guided model based on lexical similarity between retrieval diff and given diff. Lastly, for retrieve-guided model, we design a novel seq2seq network with Bi-LSTM that takes given diff, retrieval diff and retrieval message as input. We introduce a relation gate in encoder to leverage retrieval message adaptively based on semantic similarity, and a difference vector in decoder to refine the utilization of retrieval message based on semantic disparity. Experimental results on an open source dataset demonstrate that retrieval messages guidance can facilitate commit message generation task. Besides, ablation experiments prove the effectiveness of our proposed mechanisms on adjusting the use of retrieval results. Haiyang Yang, Li Kuang |
APSEC | 3 |
| 2022 | Domain Invariant Masked Autoencoders for Self-supervised Learning from Multi-domains
Haiyang Yang, Shixiang Tang, Meilin Chen, Yizhou Wang 0007, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Wanli Ouyang |
ECCV (31) | 1 |
| 2022 | Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector LearningabstractUnsupervised pretraining methods for object detection aim to learn object discrimination and localization ability from large amounts of images. Typically, recent works design pretext tasks that supervise the detector to predict the defined object priors. They normally leverage heuristic methods to produce object priors, \emph{e.g.,} selective search, which separates the prior generation and detector learning and leads to sub-optimal solutions. In this work, we propose a novel object detection pretraining framework that could generate object priors and learn detectors jointly by generating accurate object priors from the model itself. Specifically, region priors are extracted by attention maps from the encoder, which highlights foregrounds. Instance priors are the selected high-quality output bounding boxes of the detection decoder. By assuming objects as instances in the foreground, we can generate object priors with both region and instance priors. Moreover, our object priors are jointly refined along with the detector optimization. With better object priors as supervision, the model could achieve better detection capability, which in turn promotes the object priors generation. Our method improves the competitive approaches by \textbf{+1.3 AP}, \textbf{+1.7 AP} in 1\% and 10\% COCO low-data regimes object detection. Yizhou Wang 0007, Meilin Chen, Shixiang Tang, Feng Zhu 0006, Haiyang Yang, Lei Bai 0001, Rui Zhao 0001, Yunfeng Yan, Donglian Qi, Wanli Ouyang |
NeurIPS | 5 |
| 2022 | KGAT: An Enhanced Graph-Based Model for Text Classification
Xin Wang 0064, Haiyang Yang, Xingpeng Zhang, Kan Ji, Yuhong Wu, Huayi Zhan |
NLPCC (1) | 3 |
| 2022 | InCo: Intermediate Prototype Contrast for Unsupervised Domain Adaptation
Yuntao Du 0001, Hongtao Luo, Haiyang Yang, Juan Jiang, Chong-Jun Wang |
ECML/PKDD (1) | 3 |
| 2022 | CGMBL: Combining GAN and Method Name for Bug LocalizationabstractDevelopers often need to locate buggy code files in the software quality maintenance process. Bug localization aims to automatically identify potentially buggy source code files from the project codes for developers based on the bug reports. Up to now, researchers have proposed many methods to advance this task. However, the early studies only focus on the accuracy of capturing text features or the efficiency of calculating relevance scores, which do not consider the semantic gap between bug reports in natural language and codes in programming language. In this paper, we propose a novel adversarial learning model to bridge the semantic gap. Due to the different characteristics of natural language and programming language, we propose two different representation models for bug reports and code files respectively, and regards the two representation models as the generators. Then we construct adversarial learning by adding a discriminator to distinguish the source of representations so that the model can learn the public features of different texts. In addition, method name is the summary of the code function, and the relevant method name often appears in the bug report. We consider the method name information according to whether the method name appears in the report. Our model can dynamically integrate the information to improve the model effect. We evaluate our model on three open-source java project datasets and compare it with four state-of-the-art methods. The experimental results show that our model outperforms the baseline models and has a significant improvement in evaluation metrics. Besides, we conduct ablation experiments to explain each module’s contribution to the model. Haiyang Yang, Zilun Yan, Li Kuang |
QRS | 2 |