VLDB 2026 Research / reviewers in the wild / expert
Fuxiang Chen
dblp:167/7868
· DBLP profile ↗
16ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 2 first-author · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TR-GAN: Data-Augmentation-Aware Transformer-Rectification-Based Generative Adversarial Networks for Long-Term Cloud Workload ForecastingabstractMaximum utilisation of minimal amount of resources is pivotal for achieving a sustainable operation in large-scale Cloud Data Centres. Prediction driven resource provisioning in Cloud Data Centres is a potential approach to execute Cloud workloads in a sustianable way. Traditional prediction models often struggle to deliver accurate predictions under dynamic and heterogeneous cloud workloads, as capturing long-range dependencies and sudden workload spikes is often challenging in Cloud environments. In addition, recent time series models such as the Adversarial Error Correction Generative Adversarial Network (AEC-GAN) characterise shortcomings when applied to cloud workload datasets, particularly whilst managing volatility and learning irregular patterns. To address such challenges, this paper proposes a novel prediction model using Data Augment Aware Transformer Rectification-based Generative Adversarial Networks (TR-GAN), which incorporates a continuous and conditional learning Transformer block in the GAN's generator module to serve both as a data distribution moderator and as a data augmentation generator, ultimately to deliver accurate predictions. TR-GAN is the first GAN-based model tailored for long-term cloud workload forecasting that explicitly couples data augmentation with sequence rectification. Unlike discriminative forecasters, Its generative formulation allows it to model the intrinsic variability and uncertainty of cloud workloads,generate context-aware synthetic data to improve generalization under sparse or irregular patterns, and iteratively refine predictions through an adversarial learning process, thereby reducing error accumulation in long-horizon forecasts. The prediction performance of the proposed model is evaluated with two widely-used cloud workload datasets, namely the Google clusters and Alibaba traces. Experimental results demonstrate that the proposed TR-GAN model can deliver a prediction improvement of around 15% than notable state-of-the-art models, including Informer, Autoformer and AEC-GAN, for long forecasting horizons. Zekun Sun, Fuxiang Chen, Yao Lu 0021, John Panneerselvam, Lu Liu 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2025 | MPKT: Multi-Perspective Knowledge TracingabstractKnowledge tracing (KT) is an essential technique for predicting students' future performance based on the analysis of their previous learning activities. The exploration of question relevance has been shown to have a significant impact on predicting student performance. However, existing studies rely on basic attention mechanisms when investigating relevance, without fully utilizing the relationships between questions, concepts, and interactions. Additionally, current methods often neglect the number of repetitions on specific concepts when modeling forgetting behavior. This paper introduces a new knowledge tracing model that integrates multiple features into the attention mechanism to track and predict students' mastery of questions more accurately. Specifically, we explore question correlation from two perspectives: co-occurrence and answer consistency. Then, a forgetting feature is introduced to simulate the natural process of students gradually forgetting previously learned knowledge over time, considering the effects of repeated learning of the same concepts and the intervals between learning sessions. Experimental results demonstrate that our model surpasses existing popular and state-of-the-art knowledge tracing models on multiple metrics and robustness, effectively enhancing the quality of educational instruction and aids in the realization of personalized learning. Hongyun Wang, Longcheng Li, Lu Liu 0001, Zixuan Han, Fuxiang Chen |
HPCC | 6 |
| 2025 | AdvFusion: Adapter-based Knowledge Transfer for Code Summarization on Code Language ModelsabstractProgramming languages can benefit from one another by utilizing a pre-trained model for software engineering tasks such as code summarization and method name prediction. While full fine-tuning of Code Language Models (Code-LMs) has been explored for multilingual knowledge transfer, research on Parameter Efficient Fine-Tuning (PEFT) for this purpose is lim-ited. AdapterFusion, a PEFT architecture, aims to enhance task performance by leveraging information from multiple languages but primarily focuses on the target language. To address this, we propose AdvFusion, a novel PEFT-based approach that effectively learns from other languages before adapting to the target task. Evaluated on code summarization and method name prediction, AdvFusion outperforms AdapterFusion by up to 1.7 points and surpasses LoRA with gains of 1.99, 1.26, and 2.16 for Ruby, JavaScript, and Go, respectively. We open-source our scripts for replication purposes11https://github.com/ist1373/AdvFusion. Iman Saberi, Amirreza Esmaeili, Fatemeh Hendijani Fard, Fuxiang Chen |
SANER | 4 |
| 2025 | Correction to: Utilization of pre-trained language models for adapter-based knowledge transfer in software engineering
Iman Saberi, Fatemeh Hendijani Fard, Fuxiang Chen |
Empir. Softw. Eng. | 3 |
| 2025 | Reliable Indoor Localization in Multibuilding Environments: Leveraging Environment-Invariant and Position-Related FeaturesabstractReceived Signal Strength Indicator (RSSI)-based indoor localization offers a cost-effective solution for autonomous mobile robot navigation in 3D indoor environments, including cross-floor and multi-building structures. However, localization accuracy is fundamentally constrained by the low sampling density and unstable measurement of RSSI data. So far, existing methods neglect cross-environment RSSI coherence (e.g., repeated signal patterns in geometrically similar areas), resulting in unreliable fingerprint databases. What’s more, most approaches fail to model the spatial hierarchy of buildings, floors, and coordinates, which leads to lower accuracy in indoor positioning model predictions. To address these issues, we propose EP-3DLoc, a novel 3D indoor localization framework that combines an Environment-Invariant feature-based Data Completion (EIC) method with a Position-Related feature-based Localization (PRL) method. The EIC enhances data quality by filling in sparse RSSI data using environment-invariant features, which are recurring RSSI patterns found in similar environmental structures. The PRL module combines multi-scale RSSI signal processing (raw data and image-like data) with a multi-task network that analyzes location relationships, enhancing localization accuracy in 3D environments. Experimental results on public datasets (TUT2018, UTSIndoorLoc, and UJIIndoorLoc) have demonstrated that EP-3DLoc achieves state-of-the-art performance on indoor localization in multi-building environments. Further testing on the self-constructed dataset HZAUIndoorLoc have revealed that EP-3DLoc not only outperforms existing methods in localization accuracy but also maintains low energy consumption and strong resistance to interference. The dataset HZAUIndoorLoc is available at https://github.com/Hanzoe/HZAUIndoorLoc-Dataset. Wenhan Long, Xinlong Wen, Hao Liu 0056, Songquan Li, Fuxiang Chen, Lu Liu 0001, Rongbo Zhu |
IEEE Internet Things J. | 7 |
| 2025 | A graph regularized overlapping community discovery framework with three-way decisions
Xiaoyang Zou, Jinxin Cao, Hengrong Ju, Weiping Ding 0001, Lu Liu 0001, Fuxiang Chen, Di Jin 0001 |
Inf. Sci. | 6 |
| 2024 | A semi-supervised GCN-based community detection algorithmabstractCommunity detection reveals the unique characteristics and relationships of in-network members, differentiated from out-of-community members and plays a pivotal role in network analysis. In recent years, deep learning techniques have made great strides in the application of community detection, especially on label sampling models, which train graph convolutional networks by constructing a balanced training set through structural centre localization and neighbourhood node expansion. However, such algorithms are limited on centre selection in community structure. To address this, this paper introduces a novel community detection algorithm based on peak density adaptive iterative segmentation, or LDACN in short. First, labels are assigned by adaptively selecting the centre node of the community structure, which results in a more even distribution of labels in the network. Subsequently, more accurate community segmentation is achieved by taking advantage of GCN's combined ability in capturing the connectivity relationships between the nodes and their intrinsic characteristics. Experimental results on synthetic and real-world network datasets show that our algorithm improves the effectiveness of community segmentation compared to the state-of-the-art algorithms. Shu-Han Shi, Lu Liu 0001, Zixuan Han, Fuxiang Chen |
ISPA | 6 |
| 2024 | Studying Versioning in Stack OverflowabstractIn Stack Overflow (SO), a post consists of multiple components: title, question, answers, question tags, and comments. Developers can create any of these components and make changes, which we call 'edits'. Edits are an important aspect of QA websites to ensure the quality and correctness of the texts. We performed multiple analyses on the revision history of 23 million SO posts from 2008 to 2023, and we gain a more comprehensive understanding of developers' content maintenance behaviors which lay the foundation for further research. Fuxiang Chen, Mijung Kim, Fatemeh Hendijani Fard |
ASE | 2 |
| 2024 | Utilization of pre-trained language models for adapter-based knowledge transfer in software engineering
Iman Saberi, Fatemeh Hendijani Fard, Fuxiang Chen |
Empir. Softw. Eng. | 3 |
| 2023 | The repertoire of copy number alteration signatures in human cancerabstractCopy number alterations (CNAs) are a predominant source of genetic alterations in human cancer and play an important role in cancer progression. However comprehensive understanding of the mutational processes and signatures of CNA is still lacking. Here we developed a mechanism-agnostic method to categorize CNA based on various fragment properties, which reflect the consequences of mutagenic processes and can be extracted from different types of data, including whole genome sequencing (WGS) and single nucleotide polymorphism (SNP) array. The 14 signatures of CNA have been extracted from 2778 pan-cancer analysis of whole genomes WGS samples, and further validated with 10 851 the cancer genome atlas SNP array dataset. Novel patterns of CNA have been revealed through this study. The activities of some CNA signatures consistently predict cancer patients' prognosis. This study provides a repertoire for understanding the signatures of CNA in cancer, with potential implications for cancer prognosis, evolution and etiology. Ziyu Tao, Shixiang Wang, Chenxu Wu, Wei Ning, Guangshuai Wang, Kaixuan Diao, Fuxiang Chen, Xue-Song Liu |
Briefings Bioinform. | 11 |
| 2022 | On the transferability of pre-trained language models for low-resource programming languagesabstractA recent study by Ahmed and Devanbu reported that using a corpus of code written in multilingual datasets to fine-tune multilingual Pre-trained Language Models (PLMs) achieves higher performance as opposed to using a corpus of code written in just one programming language. However, no analysis was made with respect to fine-tuning monolingual PLMs. Furthermore, some programming languages are inherently different and code written in one language usually cannot be interchanged with the others, i.e., Ruby and Java code possess very different structure. To better understand how monolingual and multilingual PLMs affect different programming languages, we investigate 1) the performance of PLMs on Ruby for two popular Software Engineering tasks: Code Summarization and Code Search, 2) the strategy (to select programming languages) that works well on fine-tuning multilingual PLMs for Ruby, and 3) the performance of the fine-tuned PLMs on Ruby given different code lengths. Fuxiang Chen, Fatemeh Hendijani Fard, David Lo 0001, Timofey Bryksin |
ICPC | 1 |
| 2022 | An exploratory study on code attention in BERTabstractMany recent models in software engineering introduced deep neural models based on the Transformer architecture or use transformer-based Pre-trained Language Models (PLM) trained on code. Although these models achieve the state of the arts results in many downstream tasks such as code summarization and bug detection, they are based on Transformer and PLM, which are mainly studied in the Natural Language Processing (NLP) field. The current studies rely on the reasoning and practices from NLP for these models in code, despite the differences between natural languages and programming languages. There is also limited literature on explaining how code is modeled. Rishab Sharma, Fuxiang Chen, Fatemeh Hendijani Fard, David Lo 0001 |
ICPC | 2 |
| 2022 | LAMNER: code comment generation using character language model and named entity recognitionabstractCode comment generation is the task of generating a high-level natural language description for a given code method/function. Although researchers have been studying multiple ways to generate code comments automatically, previous work mainly considers representing a code token in its entirety semantics form only (e.g., a language model is used to learn the semantics of a code token), and additional code properties such as the tree structure of a code are included as an auxiliary input to the model. There are two limitations: 1) Learning the code token in its entirety form may not be able to capture information succinctly in source code, and 2) The code token does not contain additional syntactic information, inherently important in programming languages. Rishab Sharma, Fuxiang Chen, Fatemeh Hendijani Fard |
ICPC | 2 |
| 2019 | Paraphrase Diversification Using Counterfactual DebiasingabstractThe problem of generating a set of diverse paraphrase sentences while (1) not compromising the original meaning of the original sentence, and (2) imposing diversity in various semantic aspects, such as a lexical or syntactic structure, is examined. Existing work on paraphrase generation has focused more on the former, and the latter was trained as a fixed style transfer, such as transferring from positive to negative sentiments, even at the cost of losing semantics. In this work, we consider style transfer as a means of imposing diversity, with a paraphrasing correctness constraint that the target sentence must remain a paraphrase of the original sentence. However, our goal is to maximize the diversity for a set of k generated paraphrases, denoted as the diversified paraphrase (DP) problem. Our key contribution is deciding the style guidance at generation towards the direction of increasing the diversity of output with respect to those generated previously. As pre-materializing training data for all style decisions is impractical, we train with biased data, but with debiasing guidance. Compared to state-of-the-art methods, our proposed model can generate more diverse and yet semantically consistent paraphrase sentences. That is, our model, trained with the MSCOCO dataset, achieves the highest embedding scores, .94/.95/.86, similar to state-of-the-art results, but with a lower mBLEU score (more diverse) by 8.73%. Sunghyun Park 0005, Seung-won Hwang, Fuxiang Chen, Jaegul Choo, Jung-Woo Ha 0001, Sunghun Kim 0001, Jinyeong Yim |
AAAI | 3 |
| 2019 | NL2pSQL: Generating Pseudo-SQL Queries from Under-Specified Natural Language QuestionsabstractFuxiang Chen, Seung-won Hwang, Jaegul Choo, Jung-Woo Ha, Sunghun Kim. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Fuxiang Chen, Seung-won Hwang, Jaegul Choo, Jung-Woo Ha 0001, Sunghun Kim 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2015 | Crowd debuggingabstractResearch shows that, in general, many people turn to QA sites to solicit answers to their problems. We observe in Stack Overflow a huge number of recurring questions, 1,632,590, despite mechanisms having been put into place to prevent these recurring questions. Recurring questions imply developers are facing similar issues in their source code. However, limitations exist in the QA sites. Developers need to visit them frequently and/or should be familiar with all the content to take advantage of the crowd's knowledge. Due to the large and rapid growth of QA data, it is difficult, if not impossible for developers to catch up. To address these limitations, we propose mining the QA site, Stack Overflow, to leverage the huge mass of crowd knowledge to help developers debug their code. Our approach reveals 189 warnings and 171 (90.5%) of them are confirmed by developers from eight high-quality and well-maintained projects. Developers appreciate these findings because the crowd provides solutions and comprehensive explanations to the issues. We compared the confirmed bugs with three popular static analysis tools (FindBugs, JLint and PMD). Of the 171 bugs identified by our approach, only FindBugs detected six of them whereas JLint and PMD detected none. Fuxiang Chen, Sunghun Kim 0001 |
ESEC/SIGSOFT FSE | 1 |