Yanxin Yang

dblp:116/7287 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Enhance Language Model-based Repair for Memory-related Vulnerabilities via Knowledge-and Semantic-guided Analysis
abstract
Memory-related vulnerabilities often result in system crashes and performance drops, imposing significant risks for embedded systems. Despite the potential of Language Models (LMs) in program repair, existing LM-based approaches struggle with these vulnerabilities due to two primary limitations: i) LMs do not possess adequate domain knowledge concerning program analysis and the characteristics of memory-related vulnerabilities, and ii) LMs face constraints in managing contexts as the size of programs increases. To address this issue, we introduce MVRepair, a novel lightweight Language Model (ℓLM)-driven framework built upon a domain-specific knowledge library that is developed through the examination of 7,935 real-world memory-related vulnerabilities. By using our proposed knowledge-based analysis strategy and semantic-guided segmentation mechanism, MVRepair can substantially enhance the LM’s ability to repair programs with memory-related vulnerabilities. Comprehensive experimental results on 8,118 real-world memory-related vulnerabilities demonstrate that, compared with state-of-the-art LM-based approaches, MVRepair yields improvements of a minimum of 23.8% in EM, 31.9% in BLEU-4, and 16.7% in CodeBLEU.
Hao Shen 0002, Ming Hu 0003, Yanxin Yang, Xiaofei Xie, Mingsong Chen 0001
DATE3
2026 AdaFedRec: Adaptive Heterogeneous Federated Recommender Systems Across Multi-Device Users
Zhenkai Li, Ming Hu 0003, Chentao Jia, Yining Sun, Zhufeng Lu, Yanxin Yang, Xiaofei Xie, Mingsong Chen 0001
ICDE7
2026 S${}^{2}$2FL: Toward Efficient and Accurate Heterogeneous Split Federated Learning
abstract
Along with the prosperity of Artificial Intelligence (AI) and Internet of Things (IoT) techniques, Split Federated Learning (SFL) is becoming popular in designing Artificial Intelligence of Things (AIoT) applications, since it enables knowledge sharing among resource-constrained devices without compromising data privacy. By offloading a portion of the full model onto cloud servers, SFL can not only enable AIoT devices to accommodate large models beyond their capabilities, but also reduce their overall local training efforts. However, due to various inherent data and device heterogeneity issues, existing SFL methods greatly suffer from low inference performance and slow convergence, especially when the network bandwidth of devices is limited. To address these problems, this paper presents a novel SFL approach named Sliding Split Federated Learning (S2FL). Unlike traditional SFL methods that train the same portion of models on each device, S2FL maintains different model portions on heterogeneous AIoT devices adaptively according to their current computing capability and network bandwidth based on our proposed adaptive model sliding split method, which can balance the training time between devices and mitigate the notorious straggler problem caused by devices with weak computation power and long network transmission time. Meanwhile, based on our proposed data balance-aware training mechanism, S2FL enables the training of the server model portion on the balanced local data of grouped devices, thus alleviating the degradation of inference accuracy caused by data heterogeneity. Comprehensive experimental results highlight the superiority of S2FL over conventional SFL methods, where S2FL can achieve up to 11.58% inference accuracy improvement and 3.82× training acceleration.
Dengke Yan, Ming Hu 0003, Xiaofei Xie, Yanxin Yang, Mingsong Chen 0001
IEEE Trans. Computers4
2025 FilterFL: Knowledge Filtering-based Data-Free Backdoor Defense for Federated Learning
abstract
Due to the lack of data auditing techniques for untrusted clients, Federated Learning (FL) is vulnerable to backdoor attacks.Although various methods have been proposed to protect FL against backdoor attacks, they still exhibit poor defense performance in extreme data heterogeneity scenarios.Worse still, these methods strongly rely on additional datasets, violating the privacy protection requirements of FL.To overcome the above shortcomings, this paper proposes a novel data-free backdoor defense approach for FL, named FilterFL, which strives to prevent uploaded client models with backdoor knowledge from participating in the aggregation operation in each FL communication round.Based on our knowledge extraction and backdoor filtering schemes using two well-designed Conditional Generative Adversarial Networks (CGANs), FilterFL extracts incremental knowledge learned by a newly updated global model and filters its backdoor components, which can be used to generate one sample that reflects backdoor knowledge for each category.If an uploaded local model can confidently classify a generated sample into its target category, the knowledge contributed by the model will be excluded from the aggregation.In this way, FilterFL can effectively defend against backdoor attacks without using any additional auxiliary data.Comprehensive experiments on well-known datasets demonstrate that, compared with state-of-the-art methods, our approach achieves the best defense performance within various data heterogeneity scenarios.
Yanxin Yang, Ming Hu 0003, Xiaofei Xie, Yihao Huang 0001, Mingsong Chen 0001
CCS1
2025 MMDFL: Multi-Model-based Decentralized Federated Learning for Resource-Constrained AIoT Systems
abstract
Along with the prosperity of Artificial Intelligence (AI) techniques, more and more Artificial Intelligence of Things (AIoT) applications adopt Federated Learning (FL) to enable collaborative learning without compromising the privacy of devices. Since existing centralized FL methods suffer from the problems of single-point-offailure and communication bottleneck caused by the parameter server, we are witnessing an increasing use of Decentralized Federated Learning (DFL), which is based on Peer-to-Peer (P2P) communication without using a global model. However, DFL still faces three major challenges, i.e., limited computing power and network bandwidth of resource-constrained devices, non-Independent and Identically Distributed (non-IID) device data, and all-neighbor-dependent knowledge aggregation operations, all of which greatly suppress the learning potential of existing DFL methods. To address these problems, this paper presents an efficient DFL framework named MMDFL based on our proposed multi-model-based learning and knowledge aggregation mechanism. Specifically, MMDFL adopts multiple traveler models, which perform local training individually along their traversed devices, accelerating and maximizing knowledge learning and sharing among devices. Moreover, based on our proposed device selection strategy, MMDFL enables each traveler to adaptively explore its next best neighboring device to further enhance the DFL training performance, taking into account issues of data heterogeneity, limited resources and catastrophic forgetting phenomenon. Experimental results from simulation and a real testbed show that, compared with state-of-the-art DFL methods, MMDFL can not only significantly reduce the communication overhead but also achieve better overall classification performance for both IID and non-IID scenarios.
Dengke Yan, Yanxin Yang, Ming Hu 0003, Xin Fu 0001, Mingsong Chen 0001
DAC2
2025 Rising from Ashes: Generalized Federated Learning via Dynamic Parameter Reset
abstract
Although Federated Learning (FL) is promising in privacy-preserving collaborative model training, it faces low inference performance due to heterogeneous data among clients. Due to heterogeneous data in each client, FL training easily learns the specific overfitting features. Existing FL methods adopt the coarse-grained average aggregation strategy, which causes the global model to easily get stuck in local optima, resulting in low generalization of the global model. Specifically, this paper presents a novel FL framework named FedPhoenix to address this issue, which stochastically resets partial parameters to destroy some features of the global model in each round to guide the FL training to learn multiple generalized features for inference rather than specific overfitting features. Experimental results on various well-known datasets demonstrate that compared to SOTA FL methods, FedPhoenix can achieve up to 20.73\% accuracy improvement.
Ming Hu 0003, Yanxin Yang, Xiaofei Xie, Zekai Chen 0004, Chenyu Song, Mingsong Chen 0001
NeurIPS3
2024 AdaptiveFL: Adaptive Heterogeneous Federated Learning for Resource-Constrained AIoT Systems
abstract
Although Federated Learning (FL) is promising to enable collaborative learning among Artificial Intelligence of Things (AIoT) devices, it suffers from the problem of low classification performance due to various heterogeneity factors (e.g., computing capacity, memory size) of devices and uncertain operating environments. To address these issues, this paper introduces an effective FL approach named AdaptiveFL based on a novel fine-grained width-wise model pruning mechanism, which can generate various heterogeneous local models for heterogeneous AIoT devices. By using our proposed reinforcement learning-based device selection strategy, AdaptiveFL can adaptively dispatch suitable heterogeneous models to corresponding AIoT devices based on their available resources for local training. Experimental results show that, compared to state-of-the-art methods, AdaptiveFL can achieve up to 8.94% inference improvements for both IID and non-IID scenarios.
Chentao Jia, Ming Hu 0003, Zekai Chen 0004, Yanxin Yang, Xiaofei Xie, Yang Liu 0003, Mingsong Chen 0001
DAC4
2024 SampDetox: Black-box Backdoor Defense via Perturbation-based Sample Detoxification
abstract
The advancement of Machine Learning has enabled the widespread deployment of Machine Learning as a Service (MLaaS) applications. However, the untrustworthy nature of third-party ML services poses backdoor threats. Existing defenses in MLaaS are limited by their reliance on training samples or white-box model analysis, highlighting the need for a black-box backdoor purification method. In our paper, we attempt to use diffusion models for purification by introducing noise in a forward diffusion process to destroy backdoors and recover clean samples through a reverse generative process. However, since a higher noise also destroys the semantics of the original samples, it still results in a low restoration performance. To investigate the effectiveness of noise in eliminating different types of backdoors, we conducted a preliminary study, which demonstrates that backdoors with low visibility can be easily destroyed by lightweight noise and those with high visibility need to be destroyed by high noise but can be easily detected. Based on the study, we propose SampDetox, which strategically combines lightweight and high noise. SampDetox applies weak noise to eliminate low-visibility backdoors and compares the structural similarity between the recovered and original samples to localize high-visibility backdoors. Intensive noise is then applied to these localized areas, destroying the high-visibility backdoors while preserving global semantic information. As a result, detoxified samples can be used for inference, even by poisoned models. Comprehensive experiments demonstrate the effectiveness of SampDetox in defending against various state-of-the-art backdoor attacks.
Yanxin Yang, Chentao Jia, Dengke Yan, Ming Hu 0003, Tianlin Li, Xiaofei Xie, Xian Wei, Mingsong Chen 0001
NeurIPS1
2022 Local offset point cloud transformer based implicit surface reconstruction
abstract
Abstract Implicit neural representations, such as MLP, can well recover the topology of watertight object. However, MLP fails to recover geometric details of watertight object and complicated topology due to dealing with point cloud in a point‐wise manner. In this paper, we propose a point cloud transformer called local offset point cloud transformer (LOPCT) as a feature fusion module. Before using MLP to learn the implicit function, the input point cloud is first fed into the local offset transformer, which adaptively learns the dependency of the local point cloud and obtains the enhanced features of each point. The feature‐enhanced point cloud is then fed into the MLP to recover the geometric details and sharp features of watertight object and complex topology. Extensive reconstruction experiments of watertight object and complex topology demonstrate that our method achieves comparable or better results than others in terms of recovering sharp features and geometric details. In addition, experiments on watertight objects demonstrate the robustness of our method in terms of average result.
Yanxin Yang, Sanguo Zhang
Comput. Graph. Forum1
2012 A Comparative Study of Two Independent Component Analysis Using Reference Signal Methods
Jian-Xun Mi, Yanxin Yang
ICIC (3)2