Siyang Yu

dblp:147/1420 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An efficient parallel DeepFM for recommendation systems based on spark
Qi Lai, Zhibang Yang, Siyang Yu, Zhuo Tang, Mingxing Duan
J. Parallel Distributed Comput.4
2025 A Cross-Silo Vulnerability Federated Learning Approach Based on Content Chunking
abstract
The proliferation of vulnerable code poses a significant threat to software system security and user privacy. Given the inefficiency inherent in manual vulnerability analysis, there has been a pronounced surge of interest in automating vulnerability management using machine learning techniques. However, the scarcity of publicly accessible and large-scale datasets in the vulnerability domain impedes the advancement of automated methodologies. The advent of federated learning has introduced the potential utilization of private data for learning, while ensuring privacy and security within this paradigm presents a novel challenge. To solve this problem, we introduce a new approach called vulnerability solution with abstract syntax tree (AST), SOEHash, and clustering (V-ASC). We first obtain the AST of the vulnerability code to obtain the underlying pattern of the vulnerability. To protect data privacy as well as to extract vector features of the, we use the SOEHash algorithm to process the AST. Finally, to speed up the process of similarity comparison between vectors, we use an unsupervised clustering algorithm to transform the set of vectors into individual vulnerability clusters. Experiments on a recent vulnerability code dataset validate the effectiveness and efficiency of V-ASC.
Weisheng Zhang, Jiapeng Zhang 0001, Siyang Yu, Mingxing Duan, Kenli Li 0001
IEEE Internet Things J.3
2025 A novel shilling attack on black-box recommendation systems for multiple targets
Shuangyu Liu, Siyang Yu, Zhibang Yang, Mingxing Duan, Xiangke Liao
Neural Comput. Appl.2
2025 Towards Accurate Truth Discovery With Privacy-Preserving Over Crowdsourced Data Streams
abstract
Truth discovery endeavors to extract valuable information from multi-source data through weighted aggregation. Some studies have integrated differential privacy techniques into traditional truth discovery algorithms to protect data privacy. However, due to the neglect of outliers and limitations in budget allocation, these schemes still need improvement in the accuracy of discovery results. To solve these challenges, we propose a privacy-preserving scheme called PriPTD to achieve secure and accurate truth discovery services over crowdsourced data streams. Instead of assuming that worker weights are always stable between two neighboring timestamps, we delve deeper to consider outliers where worker weights change rapidly. Accordingly, we develop an outlier-aware weight estimation method with a time series model to capture and handle these outliers. Furthermore, to ensure data utility under a limited budget, we devise a weight-aware budget allocation algorithm. Its core idea is that timestamps with higher importance consume a larger proportion of the remaining budget. Additionally, we design a noise-aware error adjustment approach to mitigate the adverse effects of introduced noise on accuracy. Theoretical analysis and extensive experiments validate our scheme. Final comparative experiments against existing works confirm that our scheme achieves more accurate truth discovery while preserving privacy.
Zhimao Gong, Zhibang Yang, Shenghong Yang, Siyang Yu, Kenli Li 0001, Mingxing Duan
IEEE Trans. Knowl. Data Eng.4
2024 Conformity-aware adoption maximization in competitive social networks
abstract
Influence maximization (IM) problem is an extensively studied problem in social networks. It aims to find a small set of users in the social network to initiate the diffusion process and maximize the expected influence spread. Existing works on conformity-aware IM focus on the interaction between influence and conformity in a single-influence setting and ignore the role of conformity in a competitive and multiple-influence setting. This paper proposes a conformity-aware independent cascade (C-IC) model that considers the competition among multiple influences as well as the role of conformity in a user’s decision-making. It is proved that the adoption of an influence under the C-IC model is monotone and submodular. Meanwhile, we formulate two adoption maximization (AM) problems, O-AM and S-AM, which are both NP-hard. Because estimating the adoption through diffusion simulations is very time-consuming, we propose a reverse adoption estimation (RAE) method based on a reverse multiple influence sampling (RMIS) technology for the C-IC model and integrate it into the D-SSA-fix (Nguyenet al., 2018) framework, DSSA for short, to compute a solution with approximation guarantee. To further boost the performance, we present a fast one-hop adoption estimation (OAE) method and develop a heuristic algorithm based on OAE, called GOAE. Extensive experiments on eight real-world social networks show that the C-IC model is superior to a non-conformity diffusion model and that RAE+DSSA and GOAE are efficient and effective. In most cases, GOAE finds comparable solutions to RAE+DSSA and CELF with less time and memory overhead. GOAE is five to six orders of magnitude faster than CELF and RAE+DSSA is up to three orders of magnitude faster than CELF on NetHEPT. GOAE runs up to four to five orders of magnitude faster than RAE+DSSA with at most two orders of magnitude less memory usage. GOAE is more scalable than RAE+DSSA in terms of the number of seeds and the size of the social network.
Yikun Hu 0001, Siyang Yu, Xu Zhou 0001, Keqin Li 0001
Neurocomputing3
2024 MC-Net: Realistic Sample Generation for Black-Box Attacks
abstract
One area of current research on adversarial attacks is how to generate plausible adversarial examples when only a small number of datasets are available. Current adversarial attack algorithms used to attack these black-box systems face a number of challenges, such as difficulty in training convergence, ambiguous sample images, substitute models collapse, unsatisfactory attack success rates, high query cost, and low defense capability improvement of target models. As a result, constructing plausible adversarial situations in a few known real-world sample circumstances remains difficult. As a solution to the aforementioned issues, this study introduces MC-Net, a novel multi-stage and multi-class balanced generating method based on a limited number of samples to generate realistic adversarial examples. Firstly, a multi-task learning approach is used to train the GAN by fully utilizing the small samples, ensuring that the size of the generated dataset for each category is balanced. In addition, we design a weight-balancing strategy to ensure faster convergence of each sub-network. Then, in the second stage, the generated samples of different categories are used to train a substitute model, and the distillation method is adopted to learn the output distribution of the target model. Finally, adversarial examples are constructed on the generated samples to complete the attack on the target models. Extensive experiments have proven that MC-Net has the following advantages: 1) The substitute model converges quickly using limited samples and queries; 2) High attack success rates can be obtained with a few queries; and 3) The constructed adversarial examples significantly improve the target model’s defense. Furthermore, we only utilize a few queries for the Microsoft Azure online model to obtain a satisfactory result. Our code can be found at https://github.com/jiaokailun/A-fast.
Mingxing Duan, Kailun Jiao, Siyang Yu, Zhibang Yang, Bin Xiao 0001, Kenli Li 0001
IEEE Trans. Inf. Forensics Secur.3
2024 Deep Reinforcement Learning-Based Multi-Layer Cascaded Resilient Recovery for Cyber-Physical Systems
abstract
Cyber-physical systems (CPSs) are intricate systems integrating both physical and computational components. When these components fail due to malfunction or cyber-attack, potentially leading to significant damage or even collapse of the network topology. Cyber resilience, defined as the capability of a network to restore its function and structure after component failures, is crucial for ensuring that CPSs can sustain their operational capabilities in the face of complex disturbances. Recently, CPS resilience has garnered increasing attention, leading to the development of various resilience recovery methods. However, most existing studies address network and physical layer resilience in isolation, which hampers the ability to implement adaptive resilience recovery decisions across different systems. To overcome these limitations, we propose a multi-layered cascaded resilient recovery framework grounded in deep reinforcement learning. Initially, we synthesize the complex interactions between the information and physical layers in CPS resilience recovery from a global perspective, modeling the interrelations within CPSs. Subsequently, we introduce a hybrid resilient recovery strategy, encompassing both horizontal and vertical resilient recovery. The correlation matrix is used to partition the system into horizontal and vertical resilience slices. The resilient recovery strategy is subsequently modeled as an optimization problem using these slices. Following this, the Deep Recurrent Q-learning (DRQL) algorithm is introduced to implement the resilient recovery strategy in CPSs. While DRQL exhibits strong adaptability, it may lead to the sparse selection of critical samples, thereby hindering the learning process and convergence on essential experiences. To address this issue, we further develop the RR-DRQL algorithm, designed to identify the optimal CPS resilient recovery strategy. The RR-DRQL algorithm is rigorously proven to converge to the optimal solution through extensive theoretical analysis. Comprehensive experiments demonstrate that the RR-DRQL algorithm surpasses existing resilience recovery methods by 3.8%–25% regarding resilient policy recovery performance across realistic scenarios and various simulation platforms.
Kai Zhong 0004, Zhibang Yang, Siyang Yu, Kenli Li 0001
IEEE Trans. Serv. Comput.3
2023 A parallel game model-based intrusion response system for cross-layer security in industrial internet of things
abstract
Summary With the rise of industrialization, the importance of the industrial Internet of Things (IIoT) has increased significantly, and with it comes a variety of security threats. Therefore, the security of these networks is critical. Industrial Response Systems (IRSs), as the last line of security, plays an important role in the security system of the Industrial Internet of Things. In this paper, a new IRS model based on the non‐cooperative game is proposed. First, by combining the Partially Observable Markov Decision Process (POMDP) model with the stochastic game model based on the expanded attack tree, our model could effectively perceive the changes at each node. Second, our model incorporates the alarms of intrusion detection system (IDS) and the physical quantities of sensors in Industrial Cyber‐Physical System (ICPS) into the quantization system so that the model can respond to intruders more accurately and comprehensively. Finally, we develop this model based on multiprocessors to speed up the solution process, and adopt an approximation algorithm to reduce the number of iterations of the POMDP
Siyang Yu, Fan Wu 0016, Baoding Chen, Ronghui Cao, Zhibang Yang, Keqin Li 0001
Concurr. Comput. Pract. Exp.1
2023 A data balancing approach based on generative adversarial network
Lixiang Yuan, Siyang Yu, Zhibang Yang, Mingxing Duan, Kenli Li 0001
Future Gener. Comput. Syst.2
2022 Community search over large semantic-based attribute graphs
Peiying Lin, Siyang Yu, Xu Zhou 0001, Peng Peng 0001, Kenli Li 0001, Xiangke Liao
World Wide Web2
2021 Work in Progress: Path-based Graph Partition for Parallel Hardware-accelerated Functional Verification
abstract
Functional verification of large scale circuit design is a basic problem in Very Large Scale Integrated (VLSI) design. With the increasing scale of the circuit, it is urgent to divide the whole large scale circuit into some smaller sub-circuits so as to perform parallel functional verification on multiple hardware processors. The partition problem of hardware-accelerated functional verification can be regarded as a graph partition problem. However, unlike the traditional graph partition requirements for minimum cutting, the hardware-accelerated functional verification partition needs to reduce the simulation depth and improve the parallelism of the simulation. Therefore, partition for hardware-accelerated functional verification is a problem combined with graph partitioning and schedule. While the traditional schedule algorithms have high complexity and cannot handle large scale Directied Acyclic Graph (DAG) scheduling. To tackle the parallelism, depth, and cut edge problem, we design a new method, called path-metis. Path-metis combines the scheduling idea, such as the critical path information and task priority of the DAG, into the traditional multilevel partitioning method. Our preliminary experiments on real circuits show the effectiveness of the method, and the simulation depth can be reduced by about 11.35% on average compared with metis only with 27.58% cut size increasing.
Peiying Lin, Kenli Li 0001, Cen Chen 0002, Siyang Yu
RTAS5
2021 Efficient Parallel Secure Outsourcing of Modular Exponentiation to Cloud for IoT Applications
abstract
Modular exponentiation, an operation widely utilized in cryptographic protocols to transfer text and other forms of data, can also be applied to Internet-of-Things (IoT) devices with high security requirements. However, due to the high resource consumption of modular exponentiation, IoT devices can face the problem of resource insufficient. Fortunately, the secure outsourcing scheme offers a new solution for resource-constrained devices. In this article, we apply a parallel secure outsourcing scheme to provide the possibility for modular exponentiation operation, which is used in the IoT devices. After that, the task of modular exponentiation is decomposed and we introduce the scheme in more detail. In addition, based on this scheme, we designed an extension scheme for RSA, providing enhanced security for IoT devices. Finally, the analysis of experimental results based on 512-4096 b of data indicates the superiority in scalability and time consumption over the previous schemes.
Qilin Hu, Mingxing Duan, Zhibang Yang, Siyang Yu, Bin Xiao 0001
IEEE Internet Things J.4
2019 An Automated Learning-Based Procedure for Large-scale Vehicle Dynamics Modeling on Baidu Apollo Platform
abstract
In the autonomous driving industry, vehicle dynamic models are important to control-in-the-loop simulations. For current commercial self-driving simulators, vehicle dynamic models are expressed explicitly by sophisticated analytical equations, which are accurate but difficult to build and expensive to scale to fleets of vehicles of different brands. In this paper, we introduce a highly automated learning-based vehicle dynamic modeling procedure, which has been deployed on Baidu Apollo self-driving platform, to support cross-vehicle data-driven applications on a large scale. Compared with our previous analytical models, the end-to-end learning-based dynamic models can achieve high accuracy with significantly reduced re-development effort.
Kecheng Xu, Xiangquan Xiao, Siyang Yu, Jiangtao Hu, Jinghao Miao, Jingao Wang
IROS5