EDBT 2026 Demo / reviewers in the wild / expert
Yuxia Sun
dblp:132/5449
· DBLP profile ↗
15ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0002-5959-0629ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Computer networks · 4 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UAP4MA: Leveraging Multi-Agent Bandits to Generate Universal Adversarial Perturbations for Malware AttributionabstractAdvanced Persistent Threat (APT) malware group attribution is crucial for cybersecurity, yet current models remain vulnerable to adversarial attacks. Recent approaches for generating universal adversarial perturbations (UAPs) in the malware problem space have exposed substantial security risks, as a single UAP can mislead classification models across a wide range of malware. However, these methods are limited by suboptimal attack effectiveness and efficiency and rely heavily on confidence scores, restricting their practicality. To address these limitations, we propose UAP4MA, the first decision-based, problem-space, black-box UAP generation method tailored specifically for APT malware attribution models. UAP4MA employs a multi-agent Multi-Armed Bandit (MAB) framework, where agents collaboratively construct a UAP by applying functionality-preserving transformations chosen from 13 distinct types, leveraging Optimized Initialization, Collaborative Multi-Agents, and Dynamic Agent Strategies to enhance exploration and exploitation. Extensive experiments on our newly released AMG43 dataset and the APTMalware dataset demonstrate UAP4MA's outstanding attack performance, achieving over four times the fooling rate (FR), double the attack success rate (ASR), and an 80% reduction in training time compared to state-of-the-art methods. These results underscore UAP4MA's effectiveness and efficiency, establishing it as a powerful and practical approach for challenging APT attribution models. Yuxia Sun, Hengfeng Hu, Yepang Liu 0001, Haocheng Liang, Zhiquan Liu 0001, Jianfeng Ma 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | FedAPA: Server-side Gradient-Based Adaptive Personalized Aggregation for Federated Learning on Heterogeneous DataabstractPersonalized federated learning (PFL) tailors models to clients' unique data distributions while preserving privacy. However, existing aggregation-weight-based PFL methods often struggle with heterogeneous data, facing challenges in accuracy, computational efficiency, and communication overhead. We propose FedAPA, a novel PFL method featuring a server-side, gradient-based adaptive aggregation strategy to generate personalized models, by updating aggregation weights based on gradients of client-parameter changes with respect to the aggregation weights in a centralized manner. FedAPA guarantees theoretical convergence and achieves superior accuracy and computational efficiency compared to 10 PFL competitors across three datasets, with competitive communication overhead. The code and full proofs are available at: https://github.com/Yuxia-Sun/FL_FedAPA. Yuxia Sun, Aoxiang Sun, Siyi Pan, Zhixiao Fu, Jingcai Guo |
IJCAI | 1 |
| 2025 | WasmGuard: Enhancing Web Security through Robust Raw-Binary Detection of WebAssembly MalwareabstractWebAssembly (Wasm), a binary instruction format designed for efficient cross-platform execution, has rapidly become a foundational web standard, widely adopted in browsers, client-side, and server-side applications. However, its growing popularity has led to an increase in Wasm-targeted malware, including cryptojackers and obfuscated malicious scripts, which pose significant threats to web security. In spite of progress in deep learning based detection methods for Wasm malware, such as MINOS, these approaches face substantial performance degradation in adversarial environments. In our experiments, MINOS's detection accuracy dropped to 49.90% under adversarial attacks, revealing critical vulnerabilities. To address this, we introduce WasmGuard, a robust malware detection framework tailored for Wasm. WasmGuard employs FGSM-based adversarial training with prior-based initialization for perturbation bytes in customized sections, coupled with a novel adversarial contrastive learning objective. Using our large-scale dataset, WasmMal-15K (publicly available at https://github.com/Yuxia-Sun/WasmMal GitHub), WasmGuard outperforms six competing methods, achieving up to 99.20% Robust Accuracy and 99.93% Standard Accuracy under PGD-50 adversarial attacks, while maintaining low training overhead. Additionally, we have released WebChecker, a WasmGuard-powered browser plugin, providing real-time protection against malicious Wasm files, at https://github.com/Yuxia-Sun/WasmGuard. Yuxia Sun, Huihong Chen, Zhixiao Fu, Wenjian Lv, Zitao Liu 0001, Haolin Liu 0001 |
WWW | 1 |
| 2025 | $MGAP^{3}$MGAP3: Malware Group Attribution Based on PerceiverIO and Polytype Pre-TrainingabstractThe escalating prevalence of Advanced Persistent Threat (APT) malware demands more effective methods to accurately attribute malware to specific APT groups. Traditional manual attribution processes are labor-intensive and error-prone, while existing automated methods are hampered by small dataset sizes, inadequate representation learning, and poor noise reduction during preprocessing. To address these challenges, we introduce the AMG25 dataset, which expands the pool of malware samples labeled with APT group affiliations. Concurrently, we propose the MGAP3model (Malware Group Attribution based on PerceiverIO and Polytype Pre-training), which enhances attribution performance by incorporating hierarchical pre-training for disassembled codes and leveraging multi-view statistical features, all within a unified PerceiverIO architecture. This model adeptly captures complex program structures and interactions cross multiple code granularities, through a series of innovative polytype pre-training tasks. Additionally, we have developed a novel noise filtering technique that focuses on user-defined function codes, substantially reducing overfitting and boosting performance. Furthermore, a streamlined version of the model, MGAP3-Lite, has been developed to accelerate training while preserving robust performance. Extensive experiments have validated the effectiveness of our models and underscored the importance of the proposed pre-training technique. Yuxia Sun, Aoxiang Sun, Saiqin Long, Zhetao Li |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Personalized Federated Learning for Green Industrial IoTabstractIn recent years, federated learning (FL) has gained increasing attention in industrial Internet-of-Things (IIoT) domains due to its privacy-preserving advantages. However, prior works commonly adopt a one-size-fits-all strategy for FL computation resource management and reward allocation, disregarding the time-varying participant states across different FL training rounds. Consequently, these methods fail to ensure the sustainability and active participation of IIoT devices in realistic FL deployments. To bridge this gap, we propose a personalized FL methodology for green IIoT systems powered by renewable energy sources. We first establish an incentive model along with its preference parameter-solving scheme to accurately characterize the incentive preferences of individual FL participants. Subsequently, a personalized participant scheduling approach is developed to accommodate dynamic resource usage patterns and diverse incentive preferences among FL participants. Our technique integrates empirical insights into conventional proximal policy optimization methods to accelerate policy learning within reinforcement learning frameworks. Experimental results on an FL prototype system show that our methodology improves the FL model accuracy by 25.92% compared with representative baseline algorithms. Kun Cao 0001, Yangguang Cui, Rui Xu 0013, Yuxia Sun, Zhiquan Liu 0001, Chaohong Tan |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | FedLFP: Communication-Efficient Personalized Federated Learning on Non-IID Data in Mobile Edge Computing EnvironmentsabstractMobile Edge Computing (MEC) facilitates computing and storage at edge nodes near user devices, reducing latency and optimizing bandwidth. Federated Learning (FL) complements MEC by enabling privacy-preserving collaborative model training across edge nodes without sharing raw data. However, in MEC environments, FL faces challenges such as communication inefficiency and data heterogeneity (Non-IID), which degrade model performance and hinder convergence. To address these issues, we propose FedLFP, a communication-efficient personalized federated learning approach using label-free prototypes for Non-IID data in MEC. FedLFP employs three key strategies: (1) a Label-Free Prototype strategy to reduce communication costs and mitigate privacy risks, (2) a centroid prototype and combined clustering weight strategy to improve global prototype quality by considering data quantity and confidence levels, and (3) a multifaceted weighted contrastive learning strategy to enhance local representation learning and global alignment. We evaluated FedLFP on Android malware recognition using the KronoDroid dataset and standard image classification tasks, with eight configurations representing practical Non-IID settings. Experimental results show that FedLFP consistently outperforms thirteen state-of-the-art FL methods in accuracy, communication and computational efficiency. Additionally, we provide theoretical guarantees for the convergence of FedLFP under Non-IID conditions. Yuxia Sun, Siyi Pan, Aoxiang Sun, Zhixiao Fu, Saiqin Long, Zhetao Li |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Multimodal Dual-Embedding Networks for Malware Open-Set RecognitionabstractMalware open-set recognition (MOSR) is an emerging research domain that aims at jointly classifying malware samples from known families and detecting the ones from novel unknown families, respectively. Existing works mostly rely on a well-trained classifier considering the predicted probabilities of each known family with a threshold-based detection to achieve the MOSR. However, our observation reveals that the feature distributions of malware samples are extremely similar to each other even between known and unknown families. Thus, the obtained classifier may produce overly high probabilities of testing unknown samples toward known families and degrade the model performance. In this article, we propose the multi\modal dual-embedding networks, dubbed MDENet, to take advantage of comprehensive malware features from different modalities to enhance the diversity of malware feature space, which is more representative and discriminative for down-stream recognition. Concretely, we first generate a malware image for each observed sample based on their numeric features using our proposed numeric encoder with a re- designed multiscale CNN structure, which can better explore their statistical and spatial correlations. Besides, we propose to organize tokenized malware features into a sentence for each sample considering its behaviors and dynamics, and utilize language models as the textual encoder to transform it into a representable and computable textual vector. Such parallel multimodal encoders can fuse the above two components to enhance the feature diversity. Last, to further guarantee the open-set recognition (OSR), we dually embed the fused multimodal representation into one primary space and an associated sub-space, i.e., discriminative and exclusive spaces, with contrastive sampling and -bounded enclosing sphere regularizations, which resort to classification and detection, respectively. Moreover, we also enrich our previously proposed large-scaled malware dataset MAL-100 with multimodal characteristics and contribute an improved version dubbed MAL-100+. Experimental results on the widely used malware dataset Mailing and the proposed MAL-100+ demonstrate the effectiveness of our method. Jingcai Guo, Yuanyuan Xu 0004, Wenchao Xu 0001, Yufeng Zhan, Yuxia Sun, Song Guo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Energy inefficiency diagnosis for Android applications: a literature review
Yuxia Sun, Jiefeng Fang, Yanjia Chen, Yepang Liu 0001, Song Guo 0001, Xinkai Chen, Ziyuan Tan |
Frontiers Comput. Sci. | 1 |
| 2023 | Conservative Novelty Synthesizing Network for Malware Recognition in an Open-Set ScenarioabstractWe study the challenging task of malware recognition on both known and novel unknown malware families, called malware open-set recognition (MOSR). Previous works usually assume the malware families are known to the classifier in a close-set scenario, i.e., testing families are the subset or at most identical to training families. However, novel unknown malware families frequently emerge in real-world applications, and as such, require recognizing malware instances in an open-set scenario, i.e., some unknown families are also included in the test set, which has been rarely and nonthoroughly investigated in the cyber-security domain. One practical solution for MOSR may consider jointly classifying known and detecting unknown malware families by a single classifier (e.g., neural network) from the variance of the predicted probability distribution on known families. However, conventional well-trained classifiers usually tend to obtain overly high recognition probabilities in the outputs, especially when the instance feature distributions are similar to each other, e.g., unknown versus known malware families, and thus, dramatically degrade the recognition on novel unknown malware families. To address the problem and construct an applicable MOSR system, we propose a novel model that can conservatively synthesize malware instances to mimic unknown malware families and support a more robust training of the classifier. More specifically, we build upon the generative adversarial networks to explore and obtain marginal malware instances that are close to known families while falling into mimical unknown ones to guide the classifier to lower and flatten the recognition probabilities of unknown families and relatively raise that of known ones to rectify the performance of classification and detection. A cooperative training scheme involving the classification, synthesizing and rectification are further constructed to facilitate the training and jointly improve the model performance. Moreover, we also build a new large-scale malware dataset, named MAL-100, to fill the gap of lacking a large open-set malware benchmark dataset. Experimental results on two widely used malware datasets and our MAL-100 demonstrate the effectiveness of our model compared with other representative methods. Jingcai Guo, Song Guo 0001, Shiheng Ma, Yuxia Sun, Yuanyuan Xu 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Question answering model based on machine reading comprehension with knowledge enhancement and answer verificationabstractSummary Deep learning has led to important breakthroughs in natural language processing and obtained the state‐of‐the‐art results on machine reading comprehension. However, it is essential to consider the entity recognition and the detection of unanswerable questions for accuracy improvement. A novel question answering model is proposed with knowledge enhancement and answer verification to promote the performance of reading comprehension. With knowledge enhancement, the proposed model is able to recognize entities from the passage and detect word boundary precisely. To deal with unanswerable questions, the answerability of questions is evaluated based on the textual entailment. Empirical studies suggest that the proposed model has better ability of reading comprehension than others, with improvement on question answering tasks. Zimin (Max) Yang, Yuxia Sun, Qingxuan Kuang |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Disclosing and Locating Concurrency Bugs of Interrupt-Driven IoT ProgramsabstractThe Internet of Things (IoT) is envisioned as a distributed network formed by many end devices, e.g., the motes of wireless sensor network (WSN). These important IoT end devices enable ubiquitous sensing of environments and provide reliable services for mission-critical applications. However, programs running on WSN devices are typically interrupt-driven and prone to interrupt-induced concurrency bugs, which are primarily caused by erroneous interleavings among interrupt procedure instances (IPIs) (namely, executions of interrupt processing logic). In this paper, we use a set of dynamic bug patterns to characterize the concurrency bugs due to buggy access-interleavings among IPIs to shared resources, including shared memory locations and shared communication channels. By matching the above bug patterns, a dynamic analysis approach called disclosing and locating concurrency bugs of interrupt-driven IoT programs based on dynamic bug patterns (Daemon) is proposed to automatically detect and locate concurrency bugs in WSN programs. A GUI tool of Daemon is developed. As the empirical studies exhibit, the tool can discover concurrency bugs effectively and locate the buggy source lines visually. Yuxia Sun, Shing-Chi Cheung, Song Guo 0001 |
IEEE Internet Things J. | 1 |
| 2019 | Analyzing and Disentangling Interleaved Interrupt-Driven IoT ProgramsabstractIn the Internet of Things (IoT) community, wireless sensor network (WSN) is a key technique to enable ubiquitous sensing of environments and provide reliable services to applications. WSN programs, typically interrupt-driven, implement the functionalities via the collaboration of interrupt procedure instances (IPIs, namely executions of interrupt processing logic). However, due to the complicated concurrency model of WSN programs, the IPIs are interleaved intricately and the program behaviors are hard to predicate from the source codes. Thus, to improve the software quality of WSN programs, it is significant to disentangle the interleaved executions and develop various IPI-based program analysis techniques, including offline and online ones. As the common foundation of those techniques, a generic efficient and real-time algorithm to identify IPIs is urgently desired. However, the existing instance-identification approach cannot satisfy the desires. In this paper, we first formally define the concept of IPI. Next, we propose a generic IPI-identification algorithm, and prove its correctness, real-time, and efficiency. We also conduct comparison experiments to illustrate that our algorithm is more efficient than the existing one in terms of both time and space. As the theoretical analyses and empirical studies exhibit, our algorithm provides the groundwork for IPI-based analyses of WSN programs in IoT environment. Yuxia Sun, Song Guo 0001, Shing-Chi Cheung, Yong Tang 0001 |
IEEE Internet Things J. | 1 |
| 2017 | A Hybrid Collaborative Filtering Model with Deep Structure for Recommender SystemsabstractCollaborative filtering (CF) is a widely used approach in recommender systems to solve many real-world problems. Traditional CF-based methods employ the user-item matrix which encodes the individual preferences of users for items for learning to make recommendation. In real applications, the rating matrix is usually very sparse, causing CF-based methods to degrade significantly in recommendation performance. In this case, some improved CF methods utilize the increasing amount of side information to address the data sparsity problem as well as the cold start problem. However, the learned latent factors may not be effective due to the sparse nature of the user-item matrix and the side information. To address this problem, we utilize advances of learning effective representations in deep learning, and propose a hybrid model which jointly performs deep users and items’ latent factors learning from side information and collaborative filtering from the rating matrix. Extensive experimental results on three real-world datasets show that our hybrid model outperforms other methods in effectively utilizing side information and achieves performance improvement. Zhonghuo Wu, Yuxia Sun, Lingfeng Yuan, Fangxi Zhang |
AAAI | 4 |
| 2017 | Modeling the Impacts of WiFi Signals on Energy Consumption of Smartphones
Yuxia Sun, Junxian Chen, Yong Tang 0001 |
CollaborateCom | 1 |
| 2017 | Cryptanalysis on a Secret-Sharing Based Conditional Proxy Re-Encryption Scheme
Yuxia Sun |
Mob. Networks Appl. | 1 |