VLDB 2026 Research / reviewers in the wild / expert
Gong Chen 0007
dblp:49/4553-7
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-3700-7268ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-view adaptive contrastive learning for information retrieval based fault localization
Chunying Zhou, Xiaoyuan Xie, Gong Chen 0007, Bing Li 0010 |
Autom. Softw. Eng. | 3 |
| 2026 | HypeAssign: Hypergraph contrastive learning for issue assignment
Chunying Zhou, Gong Chen 0007, Xiaoyuan Xie |
Empir. Softw. Eng. | 2 |
| 2026 | PreMulBVD: A pretraining-based multi-modal binary vulnerability detection framework
Chenliang Xing, Xiaoyuan Xie, Qi Xin 0001, Gong Chen 0007 |
J. Syst. Softw. | 4 |
| 2026 | IssueCourier: Multi-Relational Heterogeneous Temporal Graph Neural Network for Open-Source Issue AssignmentabstractIssue assignment plays a critical role in open-source software (OSS) maintenance, which involves recommending the most suitable developers to address the reported issues. Given the high volume of issue reports in large-scale projects, manually assigning issues is tedious and costly. Previous studies have proposed automated issue assignment approaches that primarily focus on modeling issue report textual information, developers’ expertise, or interactions between issues and developers based on historical issue-fixing records. However, these approaches often suffer from performance limitations due to the presence of incorrect and missing labels in OSS datasets, as well as the long tail of developer contributions and the changes in developer activity as the project evolves. To address these challenges, we propose IssueCourier, a novel Multi-Relational Heterogeneous Temporal Graph Neural Network approach for issue assignment. Specifically, we formalize five key relationships among issues, developers, and source code files to construct a heterogeneous graph. Then, we further adopt a temporal slicing technique that partitions the graph into a sequence of time-based subgraphs to learn stage-specific patterns. Furthermore, we provide a benchmark dataset with relabeled ground truth to address the problem of incorrect and missing labels in existing OSS datasets. Finally, to evaluate the performance of IssueCourier, we conduct extensive experiments on our benchmark dataset. The results show that IssueCourier can improve over the best baseline up to 45.49% in top-1 and 31.97% in MRR. Chunying Zhou, Xiaoyuan Xie, Gong Chen 0007, Bing Li 0010 |
IEEE Trans. Software Eng. | 3 |
| 2025 | Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code SearchabstractCode search is a vital activity in software engineering, focused on identifying and retrieving the correct code snippets based on a query provided in natural language. Approaches based on deep learning techniques have been increasingly adopted for this task, enhancing the initial representations of both code and its natural language descriptions. Despite this progress, there remains an unexplored gap in ensuring consistency between the representation spaces of code and its descriptions. Furthermore, existing methods have not fully leveraged the potential relevance between code snippets and their descriptions, presenting a challenge in discerning fine-grained semantic distinctions among similar code snippets. To address these challenges, we introduce a multi-task hedging contrastive Learning framework for Code Search, referred to as HedgeCode. HedgeCode is structured around two primary training phases. The first phase, known as the representation alignment stage, proposes a hedging contrastive learning approach. This method aims to detect subtle differences between code and natural language text, thereby aligning their representation spaces by identifying relevance. The subsequent phase involves multi-task joint learning, wherein the previously trained model serves as the encoder. This stage optimizes the model through a combination of supervised and self-supervised contrastive learning tasks. Our framework's effectiveness is demonstrated through its performance on the CodeSearchNet benchmark, showcasing HedgeCode's ability to address the mentioned limitations in code search tasks. Gong Chen 0007, Xiaoyuan Xie, Daniel Tang, Qi Xin 0001 |
ICSE | 1 |
| 2025 | Revisit the Intuition of Mutation-Based Fault Localization in Real-world ProgramsabstractMutation-based fault localization (MBFL) is an automated fault localization method that has been extensively studied in recent years.The intuition behind MBFL is based on the assumption that mutation operations can correct faults in a program.However, this assumption has only been experimented and validated on simulated datasets, and whether it truly holds in the real world has never been investigated.Fault types in simulated datasets are simple and differ significantly from the complex and diverse faults found in realworld programs.Therefore, to investigate whether MBFL works in the real world, it is necessary to validate its intuition in the real world.The goal of this study is to analyze whether the intuition of MBFL still holds in the real world.We quantified the MBFL intuition by establishing an algorithm, which eliminated the interference of factors unrelated to MBFL itself, allowing us to directly validate the intuition of MBFL.Based on this algorithm, we conducted extensive experiments on both real-world programs and programs in simulated datasets.The results revealed an interesting trend: due to the complexity of faults in real-world programs compared to those in simulated datasets, MBFL's intuition probably cannot hold in the real world.This indicates that MBFL's intuition is difficult to hold in the real world.Consequently, we focused on analyzing the real-world faulty versions and summarized a set of mutation operators that perform better in the real world by studying the types and effects of each mutant, providing guidance for the application of MBFL. Chenliang Xing, Gong Chen 0007, Qi Xin 0001, Xiaoyuan Xie |
Internetware | 2 |
| 2025 | RFMC-CS: a representation fusion based multi-view momentum contrastive learning framework for code search
Gong Chen 0007, Xiaoyuan Xie |
Autom. Softw. Eng. | 1 |
| 2024 | FMCS: Improving Code Search by Multi-Modal Representation Fusion and Momentum Contrastive LearningabstractCode search is a critical task in software engineering, which is to search relevant codes from the codebase based on the natural language query. Although existing code search methods based on multi-modal contrast learning have achieved advanced performance, these methods still have limitations in the representation learning of multi-modal data and do not sufficiently explore the role of functionally equivalent code pairs in representation learning. To address these limitations, we propose a code search framework based on multi-modal representation fusion and momentum contrastive learning, named FMCS. We effectively retain the semantic and structural information of the code by multi-modal representation fusion. We further learn the correlation between the relevant samples by the momentum contrastive learning between samples. The experimental results on the CodeSearchNet benchmark show the effectiveness of FMCS. Gong Chen 0007, Xiaoyuan Xie |
QRS | 2 |
| 2023 | ML-KGCL: Multi-level Knowledge Graph Contrastive Learning for Recommendation
Gong Chen 0007, Xiaoyuan Xie |
DASFAA (2) | 1 |