VLDB 2026 Research / reviewers in the wild / expert
Yichen Li 0006
dblp:27/2248-6
· DBLP profile ↗
25ranked-venue papers
13as first author
25since 2021 · last 2026
0009-0009-8630-2504ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 10 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedCD: Towards Consolidated Distillation for Heterogeneous Federated LearningabstractKnowledge Distillation (KD) serves as an effective approach to addressing heterogeneity issues in Federated Learning (FL), leveraging additional datasets to align local and global models better. There are two primary distillation paradigms: feature-based distillation, which utilizes intermediate-layer features of the network, and logit-based distillation, which employs the final layer's logit outputs. However, existing studies often select distillation methods based on intuitive and empirical evidence when facing different heterogeneous settings, neglecting the intrinsic relationship between distillation paradigms and heterogeneity. This oversight may result in suboptimal federated knowledge distillation performance under heterogeneous conditions. In this paper, we propose the Consolidated Distillation for Heterogeneous Federated Learning - FedCD that balances knowledge representations from both feature-based and logit-based distillation to enhance performance. Specifically, to address the misalignment between knowledge conveyed by features and logits, we aggregate features from different layers via cross-layer attention to preserve semantic knowledge, followed by distribution modeling using Gaussian Mixture Models. This process strengthens knowledge distillation by constraining the transformation of different network layers' features under a consolidated distribution, thereby mitigating impacts from both data and model heterogeneity. Extensive experiments demonstrate that FedCD outperforms state-of-the-art methods by over 10.72% and validate the effectiveness of our approach. Yichen Li 0006, Huifa Li, Xinlin Zhuang, Haochen Xue, Haozhao Wang, Muhammad Imran Razzak |
AAAI | 1 |
| 2026 | Data-Centric Sequential Recommendation with Relation-Augmented GenerationabstractData-Centric Sequential Recommendation (DaCSR) has emerged as a promising technique that enhances dataset quality to better capture user preferences without increasing training complexity. However, mining item relations to improve data quality remains challenging due to the intricate nature of interaction sequences. Existing methods predominantly either: 1) optimize models to learn such item relations from fixed datasets at significant training cost, or 2) employ generative models to adaptively learn only interaction patterns, which lack interpretability and cannot guarantee effective data quality enhancement. In this paper, we pioneer a relation-guided dataset augmentation and regeneration framework for sequential recommendation called \textbf{RaSR}. This framework can significantly improve model performance on original datasets while maintaining training efficiency without modifying the model architecture. Specifically, we first preprocess user interactions to construct standardized sequential data and extract semantic representations via a Large Language Model (LLM). We then build a multi-relation graph with manually predefined metrics and semantic representations to generate augmented datasets. Finally, a relation-aware generator can produce regenerated datasets with both the multi-relation graph and the augmented dataset. To verify the effectiveness of RaSR, we conduct experiments on various backbone models and datasets, and achieve significant performance improvement compared to training the model only on the original dataset. Yichen Li 0006, Yichen Tan, Yijing Shan, Haozhao Wang, Rui Zhang 0003, Muhammad Imran Razzak, Ruixuan Li 0001 |
AAAI | 1 |
| 2026 | Unbiased Rectification for Sequential Recommender Systems Under Fake OrdersabstractFake orders pose increasing threats to sequential recommender systems by misleading recommendation results through artificially manipulated interactions, including click farming, context-irrelevant substitutions, and sequential perturbations. Unlike injecting carefully designed fake users to influence recommendation performance, fake orders embedded within genuine user sequences aim to disrupt user preferences and mislead recommendation results, thereby manipulating exposure rates of specific items to gain competitive advantages. To protect users' authentic interest preferences and eliminate misleading information, this paper aims to perform precise and efficient rectification on compromised sequential recommender systems while avoiding the enormous computational and time costs of retraining existing models. Specifically, we identify that fake orders are not absolutely harmful—in certain cases, partial fake orders can even have a data augmentation effect. Based on this insight, we propose Dual-view Identification and Targeted Rectification (DITaR), which primarily identifies harmful samples to achieve unbiased rectification of the system. The core idea of this method is to obtain differentiated representations from collaborative and semantic views for precise detection, and then filters detected suspicious fake orders to select truly harmful ones for targeted rectification with gradient ascent. This ensures that useful information in fake orders is not removed while preventing bias residue. Moreover, it maintains the original data volume and sequence structure, thus protecting system performance and trustworthiness to achieve optimal unbiased rectification. Extensive experiments on three datasets demonstrate that DITaR achieves superior performance compared to state-of-the-art methods in terms of recommendation quality, computational efficiency, and system robustness. Qiyu Qin, Yichen Li 0006, Haozhao Wang, Cheng Wang 0025, Rui Zhang 0003, Ruixuan Li 0001 |
AAAI | 2 |
| 2026 | UMPIRE: Unveiling LLM-generated Posts via Redundant ExpressionsabstractXiaoquan Yi, Haixing Wu, Haozhao Wang, Yichen Li, Yuhua Li, Rui Zhang, Ruixuan Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoquan Yi, Haixing Wu, Haozhao Wang, Yichen Li 0006, Yuhua Li 0003, Rui Zhang 0003, Ruixuan Li 0001 |
ACL (1) | 4 |
| 2026 | Rhythm of Opinion: Interpretable Hawkes-Graph Networks for Hierarchical Opinion Propagation
Yulong Li 0002, Zhixiang Lu, Peixin Guo, Simin Lai, Haochen Xue, Xiwei Liu, Yichen Li 0006, Zhaodong Wu, Mian Zhou, Muhammad Imran Razzak, Qingxia Li, Jionglong Su |
WWW | 8 |
| 2026 | FedCHG: Graph autoencoder enhanced federated learning for cross-Domain heterogeneous graph
Jiyuan He, Yichen Li 0006, Wenchao Xu 0001, Haozhao Wang, Yining Qi, Hongwei Lu, Ruixuan Li 0001 |
Expert Syst. Appl. | 3 |
| 2025 | FedSSI: Rehearsal-Free Continual Federated Learning with Synergistic Synaptic IntelligenceabstractContinual Federated Learning (CFL) allows distributed devices to collaboratively learn novel concepts from continuously shifting training data while avoiding \textit{knowledge forgetting} of previously seen tasks. To tackle this challenge, most current CFL approaches rely on extensive rehearsal of previous data. Despite effectiveness, rehearsal comes at a cost to memory, and it may also violate data privacy. Considering these, we seek to apply regularization techniques to CFL by considering their cost-efficient properties that do not require sample caching or rehearsal. Specifically, we first apply traditional regularization techniques to CFL and observe that existing regularization techniques, especially synaptic intelligence, can achieve promising results under homogeneous data distribution but fail when the data is heterogeneous. Based on this observation, we propose a simple yet effective regularization algorithm for CFL named \textbf{FedSSI}, which tailors the synaptic intelligence for the CFL with heterogeneous data settings. FedSSI can not only reduce computational overhead without rehearsal but also address the data heterogeneity issue. Extensive experiments show that FedSSI achieves superior performance compared to state-of-the-art methods. Yichen Li 0006, Haozhao Wang, Yining Qi, Tianzhe Xiao, Ruixuan Li 0001 |
ICML | 1 |
| 2025 | FedRE: Robust and Effective Federated Learning with Privacy PreferenceabstractDespite Federated Learning (FL) employing gradient aggregation at the server for distributed training to prevent the privacy leakage of raw data, private information can still be divulged through the analysis of uploaded gradients from clients. Substantial efforts have been made to integrate local differential privacy (LDP) into the system to achieve a strict privacy guarantee. However, existing methods fail to take practical issues into account by merely perturbing each sample with the same mechanism while each client may have their own privacy preferences on privacy-sensitive information (PSI), which is not uniformly distributed across the raw data. In such a case, excessive privacy protection from private-insensitive information can additionally introduce unnecessary noise, which may degrade the model performance. In this work, we study the PSI within data and develop FedRE, that can simultaneously achieve robustness and effectiveness benefits with LDP protection. More specifically, we first define PSI with regard to the privacy preferences of each client. Then, we optimize the LDP by allocating less privacy budget to gradients with higher PSI in a layer-wise manner, thus providing a stricter privacy guarantee for PSI. Furthermore, to mitigate the performance degradation caused by LDP, we design a parameter aggregation mechanism based on the distribution of the perturbed information. We conducted experiments with text tamper detection on T-SROIE and DocTamper datasets, and FedRE achieves competitive performance compared to state-of-the-art methods. Tianzhe Xiao, Yichen Li 0006, Yu Zhou 0053, Yining Qi, Yi Liu 0087, Wei Wang 0395, Haozhao Wang, Yi Wang 0004, Ruixuan Li 0001 |
ICMR | 2 |
| 2025 | Efficient Knowledge Transfer in Federated Recommendation for Joint Venture EcosystemabstractThe current Federated Recommendation System (FedRS) focuses on personalized recommendation services and assumes clients are personalized IoT devices (e.g., Mobile phones). In this paper, we deeply dive into new but practical FedRS applications within the joint venture ecosystem. Subsidiaries engage as participants with their users and items. However, in such a situation, merely exchanging item embedding is insufficient, as user bases always exhibit both overlaps and exclusive segments, demonstrating the complexity of user information. Meanwhile, directly uploading user information is a violation of privacy and unacceptable. To tackle the above challenges, we propose an efficient and privacy-enhanced federated recommendation for the joint venture ecosystem (FR-JVE) that each client transfers more common knowledge from other clients with a distilled user's \textit{rating preference} from the local dataset. More specifically, we first transform the local data into a new format and apply model inversion techniques to distill the rating preference with frozen user gradients before the federated training. Then, a bridge function is employed on each client side to align the local rating preference and aggregated global preference in a privacy-friendly manner. Finally, each client matches similar users to make a better prediction for overlapped users. From a theoretical perspective, we analyze how effectively FR-JVE can guarantee user privacy. Empirically, we show that FR-JVE achieves superior performance compared to state-of-the-art methods. Yichen Li 0006, Yijing Shan, Yi Liu 0087, Haozhao Wang, Cheng Wang 0025, Wei Wang 0395, Yi Wang 0004, Ruixuan Li 0001 |
NeurIPS | 1 |
| 2025 | Resource-Constrained Federated Continual Learning: What Does Matter?abstractFederated Continual Learning (FCL) aims to enable sequential privacy-preserving model training on streams of incoming data that vary in edge devices by preserving previous knowledge while adapting to new data. Current FCL literature focuses on restricted data privacy and access to previously seen data while imposing no constraints on the training overhead. This is unreasonable for FCL applications in real-world scenarios, where edge devices are primarily constrained by resources such as storage, computational budget, and label rate. We revisit this problem with a large-scale benchmark and analyze the performance of state-of-the-art FCL approaches under different resource-constrained settings. Various typical FCL techniques and six datasets in two incremental learning scenarios (Class-IL and Domain-IL) are involved in our experiments. Through extensive experiments amounting to a total of over 1,000+ GPU hours, we find that, under limited resource-constrained settings, existing FCL approaches, with no exception, fail to achieve the expected performance. Our conclusions are consistent in the sensitivity analysis. This suggests that most existing FCL methods are particularly too resource-dependent for real-world deployment. Moreover, we study the performance of typical FCL techniques with resource constraints and shed light on future research directions in FCL. Yichen Li 0006, Jiahua Dong 0001, Haozhao Wang, Yining Qi, Rui Zhang 0003, Ruixuan Li 0001 |
NeurIPS | 1 |
| 2025 | Feature Distillation is the Better Choice for Model-Heterogeneous Federated LearningabstractModel-Heterogeneous Federated Learning (Hetero-FL) has attracted growing attention for its ability to aggregate knowledge from heterogeneous models while keeping private data locally. To better aggregate knowledge from clients, ensemble distillation, as a widely used and effective technique, is often employed after global aggregation to enhance the performance of the global model. However, simply combining Hetero-FL and ensemble distillation does not always yield promising results and can make the training process unstable. The reason is that existing methods primarily focus on logit distillation, which, while being model-agnostic with softmax predictions, fails to compensate for the knowledge bias arising from heterogeneous models.
To tackle this challenge, we propose a stable and efficient Feature Distillation for model-heterogeneous Federated learning, dubbed FedFD, that can incorporate aligned feature information via orthogonal projection to integrate knowledge from heterogeneous models better. Specifically, a new feature-based ensemble federated knowledge distillation paradigm is proposed. The global model on the server needs to maintain a projection layer for each client-side model architecture to align the features separately. Orthogonal techniques are employed to re-parameterize the projection layer to mitigate knowledge bias from heterogeneous models and thus maximize the distilled knowledge. Extensive experiments show that FedFD achieves superior performance compared to state-of-the-art methods. Yichen Li 0006, Xiuying Wang 0015, Wenchao Xu 0001, Haozhao Wang, Yining Qi, Jiahua Dong 0001, Ruixuan Li 0001 |
NeurIPS | 1 |
| 2025 | Enhancing Privacy in Multimodal Federated Learning with Information TheoryabstractMultimodal federated learning (MMFL) has gained increasing popularity due to its ability to leverage the correlation between various modalities, meanwhile preserving data privacy for different clients. However, recent studies show that correlation between modalities increase the vulnerability of federated learning against Gradient Inversion Attack (GIA). The complicated situation of MMFL privacy preserving can be summarized as follows: 1) different modality transmits different amounts of information, thus requires various protection strength; 2) correlation between modalities should be taken into account. This paper introduces an information theory perspective to analyze the leaked privacy in process of MMFL, and tries to propose a more reasonable protection method \textbf{Sec-MMFL} based on assessing different information leakage possibilities of each modality by conditional mutual information and adjust the corresponding protection strength. Moreover, we use mutual information to reduce the cross-modality information leakage in MMFL. Experiments have proven that our method can bring more balanced and comprehensive protection at an acceptable cost. Tianzhe Xiao, Yichen Li 0006, Yining Qi, Yi Liu 0087, Wei Wang 0395, Haozhao Wang, Yi Wang 0004, Ruixuan Li 0001 |
NeurIPS | 2 |
| 2025 | Personalized Federated Recommendation for Cold-Start Users via Adaptive Knowledge FusionabstractFederated Recommendation System (FRS) usually offers recommendation services for users while keeping their data locally to ensure privacy. Currently, most FRS literature assumes that fixed users participate in federated training with personal IoT devices (e.g., mobile phones and PC). However, users may join incrementally, and retraining the entire FRS for each new participating user is unfeasible due to the high training costs and the limited global knowledge contribution from a small number of new users. To guarantee the quality service for these new users, we take a dive into the federated recommendation for cold-start users, a novel scenario where the new participating users can directly obtain a promising recommendation without comprehensive training with all participating users by leveraging both transferred knowledge from the converged warm clients and the knowledge learned from the local data. Yichen Li 0006, Yijing Shan, Yi Liu 0087, Haozhao Wang, Wei Wang 0395, Yi Wang 0004, Ruixuan Li 0001 |
WWW | 1 |
| 2025 | Re-Fed+: A Better Replay Strategy for Federated Incremental LearningabstractFederated learning (FL) has emerged as a significant distributed machine learning paradigm. It allows the training of a global model through user collaboration without the necessity of sharing their original data. Traditional FL generally assumes that each client's data remains fixed or static. However, in real-world scenarios, data typically arrives incrementally, leading to a dynamically expanding data domain. In this study, we examine catastrophic forgetting within Federated Incremental Learning (FIL) and focus on the training resources, where edge clients may not have sufficient storage to keep all data or computational budget to implement complex algorithms designed for the server-based environment. We propose a general and low-cost framework for FIL named Re-Fed+, which is designed to help clients cache important samples for replay. Specifically, when a new task arrives, each client initially caches selected previous samples based on their global and local significance. The client then trains the local model using both the cached samples and the new task samples. From a theoretical perspective, we analyze how effectively Re-Fed+ can identify significant samples for replay to alleviate the catastrophic forgetting issue. Empirically, we show that Re-Fed+ achieves competitive performance compared to state-of-the-art methods. Yichen Li 0006, Haozhao Wang, Yining Qi, Wei Liu 0144, Ruixuan Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Better Knowledge Enhancement for Privacy-Preserving Cross-Project Defect PredictionabstractABSTRACT Cross‐project defect prediction (CPDP) poses a nontrivial challenge to construct a reliable defect predictor by leveraging data from other projects, particularly when data owners are concerned about data privacy. In recent years, federated learning (FL) has become an emerging paradigm to guarantee privacy information by collaborative training a global model among multiple parties without sharing raw data. While the direct application of FL to the CPDP task offers a promising solution to address privacy concerns, the data heterogeneity arising from proprietary projects across different companies or organizations will bring troubles for model training. In this paper, we study the privacy‐preserving CPDP with data heterogeneity under the FL framework. To address this problem, we propose a novel knowledge enhancement approach named FedDP with two simple but effective solutions: 1. local heterogeneity awareness and 2. global knowledge distillation. Specifically, we employ open‐source project data as the distillation dataset and optimize the global model with the heterogeneity‐aware local model ensemble via knowledge distillation. Experimental results on 19 projects from two datasets demonstrate that our method significantly outperforms baselines. Yichen Li 0006, Haozhao Wang, Lei Zhao 0001 |
J. Softw. Evol. Process. | 2 |
| 2024 | Towards Efficient Replay in Federated Incremental LearningabstractIn Federated Learning (FL), the data in each client is typically assumed fixed or static. However, data often comes in an incremental manner in real-world applications, where the data domain may increase dynamically. In this work, we study catastrophic forgetting with data heterogeneity in Federated Incremental Learning (FIL) scenarios where edge clients may lack enough storage space to retain full data. We propose to employ a simple, generic frame-work for FIL named Re-Fed, which can coordinate each client to cache important samples for replay. More specifically, when a new task arrives, each client first caches selected previous samples based on their global and local im-portance. Then, the client trains the local model with both the cached samples and the samples from the new task. The-oretically, we analyze the ability of Re-Fed to discover important samples for replay thus alleviating the catastrophic forgetting problem. Moreover, we empirically show that Re-Fed achieves competitive performance compared to state-of-the-art methods. Yichen Li 0006, Qunwei Li, Haozhao Wang, Ruixuan Li 0001, Leon Wenliang Zhong |
CVPR | 1 |
| 2024 | Personalized Federated Domain-Incremental Learning Based on Adaptive Knowledge Matching
Yichen Li 0006, Wenchao Xu 0001, Haozhao Wang, Yining Qi, Jingcai Guo, Ruixuan Li 0001 |
ECCV (46) | 1 |
| 2024 | FedTA: Unsupervised Federated Prototype Learning with Temperature AdaptationabstractFederated Learning (FL) has emerged as a foundational paradigm that enables collaborative training of deep neural networks across distributed clients while ensuring data privacy. However, most of the existing research primarily concentrates on supervised federated learning tailored for specific downstream tasks. In this paper, we concentrate on unsupervised federated learning, where clients collaborate to identify data patterns, structures, or representations. Contrastive learning methods have proven highly effective for unsupervised learning in server-based environments. However, a straightforward adaptation of the contrastive learning approach for this setting falls short. Upon analyzing client behaviors during training, we identified an unstable training process stemming from NonIID issues, which can result in diminished performance. To address this problem, we propose a novel Temperature-Adaption based unsupervised Federated prototype learning approach, termed FedTA. This method incorporates two straightforward yet potent solutions: 1. Prototype Enhancement and 2. Temperature-Adaption. Experimental results show that our proposed approach surpasses the state-of-the-art, achieving a performance accuracy improvement of up to 5.3%. Juan Zhao 0010, Xiaoquan Yi, Ruixuan Li 0001, Yuhua Li 0003, Haozhao Wang, Yichen Li 0006, Zhiying Deng |
HPCC | 6 |
| 2024 | FedCDA: Federated Learning with Cross-rounds Divergence-aware AggregationabstractIn Federated Learning (FL), model aggregation is pivotal. It involves a global server iteratively aggregating client local trained models in successive rounds without accessing private data. Traditional methods typically aggregate the local models from the current round alone. However, due to the statistical heterogeneity across clients, the local models from different clients may be greatly diverse, making the obtained global model incapable of maintaining the specific knowledge of each local model. In this paper, we introduce a novel method, FedCDA, which selectively aggregates cross-round local models, decreasing discrepancies between the global model and local models.
The principle behind FedCDA is that due to the different global model parameters received in different rounds and the non-convexity of deep neural networks, the local models from each client may converge to different local optima across rounds. Therefore, for each client, we select a local model from its several recent local models obtained in multiple rounds, where the local model is selected by minimizing its divergence from the local models of other clients. This ensures the aggregated global model remains close to all selected local models to maintain their data knowledge. Extensive experiments conducted on various models and datasets reveal our approach outperforms state-of-the-art aggregation methods. Haozhao Wang, Yichen Li 0006, Yuan Xu 0033, Ruixuan Li 0001, Tianwei Zhang 0004 |
ICLR | 3 |
| 2024 | SR-FDIL: Synergistic Replay for Federated Domain-Incremental LearningabstractFederated Learning (FL) is to allow multiple clients to collaboratively train a model while keeping their data locally. However, existing FL approaches typically assume that the data in each client is static and fixed, which cannot account for incremental data with domain shift, leading to catastrophic forgetting on previous domains, particularly when clients are common edge devices that may lack enough storage to retain full samples of each domain. To tackle this challenge, we proposeFederatedDomain-IncrementalLearning viaSynergisticReplay (SR-FDIL), which alleviates catastrophic forgetting by coordinating all clients to cache samples and replay them. More specifically, when new data arrives, each client selects the cached samples based not only on their importance in the local dataset but also on their correlation with the global dataset. Moreover, to achieve a balance between learning new data and memorizing old data, we propose a novel client selection mechanism by jointly considering the importance of both old and new data. We conducted extensive experiments on several datasets of which the results demonstrate that SR-FDIL outperforms state-of-the-art methods by up to 4.05% in terms of average accuracy of all domains. Yichen Li 0006, Wenchao Xu 0001, Yining Qi, Haozhao Wang, Ruixuan Li 0001, Song Guo 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Hypernetwork-driven centralized contrastive learning for federated graph classification
Jianian Zhu, Yichen Li 0006, Haozhao Wang, Yining Qi, Ruixuan Li 0001 |
World Wide Web (WWW) | 2 |
| 2023 | DaFKD: Domain-aware Federated Knowledge DistillationabstractFederated Distillation (FD) has recently attracted increasing attention for its efficiency in aggregating multiple diverse local models trained from statistically heterogeneous data of distributed clients. Existing FD methods generally treat these models equally by merely computing the average of their output soft predictions for some given input distillation sample, which does not take the diversity across all local models into account, thus leading to degraded performance of the aggregated model, especially when some local models learn little knowledge about the sample. In this paper, we propose a new perspective that treats the local data in each client as a specific domain and design a novel domain knowledge aware federated distillation method, dubbed DaFKD, that can discern the importance of each model to the distillation sample, and thus is able to optimize the ensemble of soft predictions from diverse models. Specifically, we employ a domain discriminator for each client, which is trained to identify the correlation factor between the sample and the corresponding domain. Then, to facilitate the training of the domain discriminator while saving communication costs, we propose sharing its partial parameters with the classification model. Extensive experiments on various datasets and settings show that the proposed method can improve the model accuracy by up to 6.02% compared to state-of-the-art baselines. Haozhao Wang, Yichen Li 0006, Wenchao Xu 0001, Ruixuan Li 0001, Yufeng Zhan, Zhigang Zeng |
CVPR | 2 |
| 2023 | An Efficient Design Smell Detection Approach with Inter-class Relation (S)abstractCode smell indicates the potential designed problems and quality of source code affecting the software maintenance and readability.Hence, detecting code smells in a timely and effective manner can provide guides for developers in refactoring.Existing methods generally treat these code smells from different granularities equally by merely extracting tokens-based or abstract syntax tree(AST)-based code representation, which does not take the diversity of code smells into account, especially when fewer researches concern design smells.To tackle this challenge, we propose Design Smell Detection through Inter-class Relation, which leverages the corresponding design smells features for code smell detection.More specifically, we employ AST-tokens instead of traditional word-tokens or AST to obtain code syntax information from the deep dimension.Meanwhile, we analyze the common structural feature of design smells and propose the interclass relation among different class files contained in the same package.Moreover, to verify the effectiveness of our proposed method, we carry out extensive experiments with various settings on our new dataset and the results demonstrate that our method outperforms state-of-the-art methods by up to 31% in terms of F1-measure of all code smells.The code is available at: https://github.com/xzb777/designSmellUML Yichen Li 0006 |
SEKE | 2 |
| 2022 | Multi-Label Code Smell Detection with Hybrid Model based on Deep LearningabstractCode smell is an indicator of potential problems in a software design that have a negative impact on readability and maintainability.Hence, it is essential for developers to make out the code smell to get tips on code maintenance in time.Fortunately, many approaches like metric-based, heuristic-based, machine-learning based and deep-learning based have been proposed to detect code smells.However, existing methods, using the simple code representation to describe different code smells unilaterally, cannot efficiently extract enough rich information from source code.What is more, one code snippet often has several code smells at the same time and there is a lack of multilabel code smell detection based on deep learning.In this paper, we propose a hybrid model with multi-level code representation to further optimize the code smell detection.First, we parse the code into the abstract syntax tree(AST) with control and data flow edges and the graph convolution network is applied to get the prediction at the syntactic and semantic level.Then we use the bidirectional long-short term memory network with attention mechanism to analyze the code tokens at the token-level in the meanwhile.Finally we get the fusion prediction result of the models.Experimental results show that our model can perform outstanding not only in single code smell detection but also in multi-label code smell detection. Yichen Li 0006 |
SEKE | 1 |
| 2022 | Hybrid Model with Multi-Level Code Representation for Multi-Label Code Smell Detection (077)abstractCode smell is an indicator of potential problems in a software design that have a negative impact on readability and maintainability. Hence, detecting code smells in a timely and effective manner can provide guides for developers in refactoring. Fortunately, many approaches like metric-based, heuristic-based, machine-learning-based and deep-learning-based have been proposed to detect code smells. However, existing methods, using the simple code representation to describe different code smells unilaterally, cannot efficiently extract enough rich information from source code. In addition, one code snippet often has several code smells at the same time and there is a lack of multi-label code smell detection based on deep learning. In this paper, we present a large-scale dataset for the multi-label code smell detection task since there is still no publicly sufficient dataset for this task. The release of this dataset would push forward the research in this field. Based on it, we propose a hybrid model with multi-level code representation to further optimize the code smell detection. First, we parse the code into the abstract syntax tree (AST) with control and data flow edges and the graph convolution network is applied to get the prediction at the syntactic and semantic level. Then we use the bidirectional long-short term memory network with attention mechanism to analyze the code tokens at the token-level in the meanwhile. Finally, we get the fusion prediction result of the models. Experimental results illustrate that our proposed model outperforms the state-of-the-art methods not only in single code smell detection but also in multi-label code smell detection. Yichen Li 0006, An Liu 0002, Lei Zhao 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |