Xiaoyu Yi 0003

dblp:210/5029-3 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-4755-6172ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Persistent Clean-Label Backdoor Attacks on Semisupervised Social Graph Node Classification
abstract
Semisupervised social graph node classification (SSGNC) attempts to deduce node-related information of social graph with limited labeled training samples. It is primarily deployed in large-scale graph processing, e.g., malicious client detection, knowledge graph, and recommender system. However, in this article, we identify that the SSGNC model is also extremely sensitive to backdoor attacks. We present a novel persistent clean-label backdoor attack (PerCBA) on SSGNC, which selectively poisons unmarked training nodes before learning to compel the trained model to misclassify trigger-embedded inputs into malicious class. Specifically, PerCBA employs a style-agnostic trigger generator with adjustable perturbation strategy to produce perturbed triggers. These triggers are pasted onto a small subset of unmarked nodes ($< \, 4\%$), enabling the adversary to covertly poison the training graph and implant backdoors into the model without modifying labels. Additionally, to ensure SSGNC robustness when confronted with homogenous threats, we present a testing sample filtering-based defense strategy for PerCBA. It employs feature distribution to identify poisoned nodes and applies Gaussian blur and thresholding to remove the trigger fraction, thereby restoring suspicious data to clean states. Extensive experiments on SOTA SSGNC models and datasets indicate that PerCBA performs high attack success rates (maxima 96.25%) while remaining evasive, and the defense method can effectively mitigate attacks and purify backdoored models.
Xiao Yang 0016, Gaolei Li, Xinzheng Feng, Xiaoyu Yi 0003, Jianhua Li 0001
IEEE Trans. Comput. Soc. Syst.4
2025 3D-MGW: A Memory-Efficient Grouped Watermark for Multi-Object 3D Gaussian Splatting
abstract
Multi-object 3D Gaussian Splatting (3DGS) technology aims to efficiently synthesize complex 3D scenes from images while allowing users to manipulate objects through textual prompts. Training multi-object 3DGS models requires substantial computational resources, making it necessary to protect the generated 3D objects from unauthorized reproduction, modification, and distribution. Existing watermarking solutions suffer from high memory consumption and require additional time overhead. Moreover, they cannot precisely localize watermarks to specific objects, making it difficult to trace individual contributions when multiple creators collaborate on a same multi-object scene. To address these challenges, we propose 3D-MGW, a novel memory-efficient grouped watermark for multi-object 3DGS. Within 3D-MGW, background scenes are reconstructed from images, while diffusion models are incorporated to guide highquality 3DGS synthesis from prompts. To eliminate additional training overhead, watermark embedding is integrated within the 3DGS training process rather than implementing it separately. Subsequently, a grouped Gaussian strategy is introduced to enable granular, high-capacity multi-object watermark. Additionally, a Gaussian compression module is proposed to eliminate redundant primitives, reducing the storage footprint of Gaussian models. Through comprehensive experiments, our 3D-MGW demonstrates exceptional performance, achieving 95% watermark extraction accuracy under 64-bit watermarks while reducing storage utilization by up to 61%, highlighting its substantial potential for multi-object 3DGS applications.
Hui Su, Gaolei Li, Wenkai Huang 0003, Xiaoyu Yi 0003, Jianhua Li 0001
ICPADS5
2024 MKPL: Multi-dimensional Knowledge-embedded Prompt Learning for Few-shot Malware Family Recognition
abstract
Large language models (LLMs) bring great potential for next-generation malware family recognition with their capacity to understand complex code semantics by integrating multi-dimensional data features. However, existing fine-tuning methods still rely on well-labelled datasets and powerful computation resources, which is particularly challenging when the variety and amount of malware grow in real-time. To more effectively recognize unknown malware varieties based on LLMs, a novel multi-dimensional knowledge-embedded prompt learning (MKPL) framework is proposed, in which prompts are generated through two main steps: 1) cross-linguistic prompt paraphrasing (CPP) for embedding multi-dimensional knowledge into templates, and 2) prompt scoring for selecting the most effective prompt templates. Moreover, to reduce feature loss during prompt tuning, a sampling-infer-concatenation pipeline is designated to process these long API malware sequences. Specifically, a single-sentence template can be upgraded to a multi-sentence template by integrating statistic features into CPP, which is essential to improve the robustness of recognition results. Comprehensive experiments across eight malware families in few-shot scenarios demonstrate the proposed method’s superior performance in all metrics.
Shuilin Li, Gaolei Li, Xiaoyu Yi 0003, Jianhua Li 0001, Mianxiong Dong, Kaoru Ota
HPCC4
2024 LateBA: Latent Backdoor Attack on Deep Bug Search via Infrequent Execution Codes
abstract
Backdoor attacks can mislead deep bug search models by exploring model-sensitive assembly code, which can change alerts to benign results and cause buggy binaries to enter production environments. But assembly instructions have strict constraints and dependencies, and these additional model-sensitive assembly codes destroy semantics and syntax and are easily detected by dynamic analysis or context-based detection. To escape from the dynamic analysis-based detection, we propose a novel latent backdoor attack (LateBA) scheme based on the locality principle of program execution, which only poisons a few of infrequent execution codes, minimizing the effects on the original code logic. In LateBA, a progressive seed mutating strategy is designated to change the American Fuzzy Lop (AFL)-based path search tool to pay more attention to infrequent execution codes. With this strategy, the optimal range to positions in the whole program is determined. Subsequently, triggers are target model-sensitive assembly instructions, and try to minimize the variables that have been called in the context instructions in the trigger. Finally, we employ code semantic feature comparisons to select precise trigger injection positions within these ranges. The selection criteria of the trigger injection position is whether the corresponding code segments in this position have a data dependency relationship with other code segments. We evaluate the performance of LateBA over 7 deep bug search tasks. The results demonstrate the attack success rate of the proposed LateBA is considerable and competitive against the baselines.
Xiaoyu Yi 0003, Gaolei Li, Wenkai Huang 0003, Xi Lin 0003, Jianhua Li 0001, Yuchen Liu 0001
Internetware1
2024 HSESR: Hierarchical Software Execution State Representation for Ultralow-Latency Threat Alerting Over Internet of Things
abstract
To reduce attack risks in Internet of Things (IoT), many security vendors conduct software security analysis on IoT devices all the time. However, how to build an ultralow-latency threat alerting strategy using software vulnerability information still faces challenges. First, existing terminal threat detection methods for IoT systems relying on Indicators of Compromise (IoC) threat intelligence can only cover limited software vulnerabilities so the alert validity rate is still very low. Second, most users lack security knowledge and cannot proactively distinguish high-risk vulnerabilities, resulting in untimely reporting. In this article, a novel hierarchical software execution state representation (HSESR) scheme is proposed for ultralow latency threat alerting over IoT systems based on Beyond 5G. In HSESR, function call graphs are recorded and delivered to edge servers for swiftly identifying suspicious threat behaviors based on deep graph representation, while corresponding instruction sequences are delivered to the cloud data center for further matching the vulnerability information via recurrent semantic representation. To improve the effectiveness of HSESR, the graph representation is also actively encapsulated into the corresponding semantic representation, together acting as an implicit threat behavior signature, which is essential to associate with a security patch. Moreover, to accelerate the detection of suspicious behaviors, we also propose a deep reinforcement learning-based graph searching (DRL-GS) strategy to crop the huge function call graph of the entire software to timely report high-risk threat behaviors with minimized resource consumption. By instancing 1-day attacks on a simulated beyond 5G IoT system, the performance of HSESR is trustfully competitive against existing baselines, and the efficiency of threat detection was increased by 21.63%.
Xiaoyu Yi 0003, Gaolei Li, Bei Chen 0004, Xi Lin 0003, Yuchen Liu 0001, Jianhua Li 0001
IEEE Internet Things J.1
2023 SemSBA: Semantic-perturbed Stealthy Backdoor Attack on Federated Semi-supervised Learning
abstract
Federated semi-supervised learning (FSSL) has been perceived as a promising approach that leverages semi-supervised learning and federated learning (FL) to provide powerful privacy preservation while reducing the burden on human supervision. However, due to the lack of strict participant identification and the significant proportion of unlabeled samples, FSSL is more susceptible to covert backdoor attacks than traditional machine learning. To validate this speculation, a novel semantic-perturbed stealthy backdoor attack (SemSBA) scheme is proposed for FSSL-based systems. In SemSBA, we select original natural semantic features in the unlabeled training samples as backdoor triggers and then generate poisoned samples by adding adversarial perturbations that move them across the model decision boundary. With SemSBA, the adversary can trigger the hidden backdoor in the victim model during the inference stage without any deliberate modifications on testing samples. To further improve the strength and robustness of the attack, a pseudo label steering enhancement strategy is also designed to perturb the weakly-augmented version of unlabeled samples to induce target pseudo label allocations. Additionally, to improve the attack success rate, we amplify the weight of the local backdoored model during FSSL’s model aggregation process to manipulate the game between benign clients and malicious clients. Extensive experiments based on two benchmark datasets demonstrate that the proposed SemSBA scheme can achieve comparable stealthiness against existing attacks.
Yingrui Tong, Jun Feng 0007, Gaolei Li, Xi Lin 0003, Chengcheng Zhao, Xiaoyu Yi 0003, Jianhua Li 0001
ICPADS6