Su Yao

dblp:198/8323 · DBLP profile ↗
← Back
35ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0001-5165-2787ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 21 · 4 first-author · 18 since 2021Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Transferable Graph Condensation from the Causal Perspective
abstract
The increasing scale of graph datasets has significantly improved the performance of graph representation learning methods, but it has also introduced substantial training challenges. Graph dataset condensation techniques have emerged to compress large datasets into smaller yet information-rich datasets, while maintaining similar test performance. However, these methods strictly require downstream applications to match the original dataset and task, which often fails in cross-task and cross-domain scenarios. To address these challenges, we propose a novel causal-invariance-based and transferable graph dataset condensation method, named TGCC, providing effective and transferable condensed datasets. Specifically, to preserve domain-invariant knowledge, we first extract domain causal-invariant features from the spatial domain of the graph using causal interventions. Then, to fully capture the structural and feature information of the original graph, we perform enhanced condensation operations. Finally, through spectral-domain Enhanced contrastive learning, we inject the causal-invariant features into the condensed graph, ensuring that the compressed graph retains the causal information of the original graph. Experimental results on five public datasets and our novel FinReport dataset demonstrate that TGCC achieves up to a 13.41% improvement in cross-task and cross-domain complex scenarios compared to existing methods, and achieves state-of-the-art performance on 5 out of 6 datasets in the single dataset and task scenario.
Huaming Du, Su Yao, Yiying Wang, Yueyang Zhou, Jinshi Zhang, Yu Zhao 0019, Guisong Liu, Hegui Zhang, Carl Yang 0001, Gang Kou
AAAI3
2026 KANC: An Interpretable Network Performance Prediction Model based on KAN and GNN
Yizhong Hu, Jianfeng Guan, Su Yao, Yang Liu 0038
ICC5
2026 FENIX: Enabling In-Network DNN Inference with FPGA-Enhanced Programmable Switches
Tong Li 0014, Yinchao Zhang, Xiangsheng Zeng, Su Yao, Ke Xu 0002
NSDI6
2026 Forewarned is Forearmed: A Responsive Congestion Control with Non-intrusive Uplink Dynamics Capture
Yiying Lin, Shenghui Wei, Enhuan Dong, Kang Chen 0001, Tong Li 0014, Yinchao Zhang, Renjie Xie, Su Yao, Ke Xu 0002, Changqiao Xu
SIGCOMM9
2026 A hypergraph-based model for tumor prognosis using local and global information fusion on H&E-stained histology images
Yanfen Cui, Zhenhui Li, Xiuming Zhang, Su Yao, Dacheng Yang, Zhishun Liu, Shiwei Luo, Guangjun Yang, Lixu Yan, Xiangtian Zhao, Yingqiu Huo, Jiahui Ma, Wenfeng He, Tao Tan 0002, Anant Madabhushi, Jinglei Tang, Zaiyi Liu, Cheng Lu 0001
Medical Image Anal.6
2026 OCEAN: Optional Capability-Based En Route Acknowledgement in Network Layer
abstract
High security and low latency are important in mission-critical data transmission, such as the end-to-end transmission in Industrial IoT (IIoT). However, existing schemes often struggle to simultaneously meet these demanding requirements due to hardware limitations and the lack of a packet lossless forwarding protocol in the network layer data plane. To address this challenge, we propose OCEAN (Optional Capability-based En route Acknowledgement in Network layer). OCEAN includes (1) an in-network caching hardware, which is a programmable Application Specific Integrated Circuit (ASIC) integrated with a Field Programmable Gate Array (FPGA), and (2) a packet lossless forwarding protocol in the network layer data plane. In OCEAN, each packet was generated by an authorized end device, while each en route node verifies the packet, and caches it until receiving the acknowledgment from the next en route node. It incurs negligible latency to packet forwarding when there is no packet loss while retransmitting the packet at the en route node after a short timeout, which reduces the packet forwarding latency. Besides that, the per-packet verification guarantees that the adversary could not subvert the forwarding protocol. Our simulation in the BMv2 environment confirms its functionality, and the hardware implementation demonstrates that it can process packets at line rate with a total processing latency ranging from 2519 ns to 6160 ns, which is negligible in end-to-end transmission.
Su Yao, Songtao Fu, Qi Li 0002, Zhuotao Liu, Yinchao Zhang, Ke Xu 0002
IEEE Trans. Dependable Secur. Comput.1
2026 SeqFedEDT: Accelerating Sequential Federated Learning on non-IID Data via Element-Wise Decoupled Training
abstract
Sequential federated learning (SFL) trains models collaboratively across clients in a chain manner. This order shows communication efficiency compared to traditional FL in a parallel manner with a star topology. However, SFL can fail to produce stable training results when clients have significant statistical heterogeneity among their local data distributions. To address these challenges, we propose a novel element-wise model decoupling framework namedSeqFedEDTthat accelerates SFL training by separating model parameters of each client into a shared subset for global knowledge collaboration and a personalized subset for migrating data heterogeneity. We explore three types of parameter contribution scoring metrics based on gradient, Fisher information, and parameter importance (PI) for personalized parameter selection. In addition, we propose a quantile-based thresholding mechanism to separate shared and personalized subsets and explore the best performance quantile selection in numerical studies. Extensive experiments demonstrate thatSeqFedEDToutperforms eight state-of-the-art methods across diverse datasets and heterogeneity scenarios. All code and results are available athttps://github.com/tian0920/SeqFedEDT.
Tian Du, Xingyan Chen, Yaling Liu, Su Yao, Gang Kou, Fuzhen Zhuang, Changqiao Xu, Gabriel-Miro Muntean
IEEE Trans. Mob. Comput.5
2026 OMEGA: A Comprehensive Cloud-Edge-Device Authentication and Key Agreement Scheme for Collaborative Multi-Factory Manufacturing in IIoT
abstract
Cloud Manufacturing (CMfg) has revolutionized traditional manufacturing by enabling resource sharing across factory boundaries. As industry increasingly adopts cloud-edge-device collaborative ecosystems, effective authentication and key agreement (AKA) mechanisms play a critical role in safeguarding security across these complex multi-factory environments. Existing schemes, however, focus primarily on isolated binary security relationships (e.g., device-edge, device-cloud, or edge-cloud pairs individually), neglecting the integrated cloud-edge-device collaborative security demands inherent in multi-factory environments, while often relying on trusted authorities and high-overhead cryptographic mechanisms. This leads to redundant authentication processes, increased latency, and security vulnerabilities as devices must separately establish connections with each entity. To bridge these gaps, this paper introduces OMEGA: a comprehensive Cloud-Edge-Device Authentication and Key Agreement scheme for collaborative multi-factory manufacturing that pioneers an integrated security architecture. OMEGA’s distinctive advantage lies in its ability to establish all necessary secure connections (device-edge, device-cloud, and edge-cloud) through a single authentication request from the smart manufacturing device (SMD), significantly reducing authentication overhead. By leveraging lightweight hash operation, OMEGA creates a cohesive security fabric that enables SMDs to concurrently access specialized capabilities from multiple clouds while leveraging edge computing for time-sensitive operations. Security analysis using both Real-Or-Random (ROR) model and ProVerif formal verification tool demonstrates OMEGA achieves robust security while performance evaluation confirms its superior efficiency in industrial environments.
Kexian Liu, Jianfeng Guan, Su Yao, Ilsun You, Hongke Zhang
IEEE Trans. Netw. Serv. Manag.3
2025 Decentralized Offloading for AI Inference with Heterogeneous Models in Mobile Cellular Networks
abstract
Enabling on-device intelligence that supports AI inference, has become a main trend in mobile application development. To enhance intelligence on resource-limited smartphones in mobile cellular networks, edge computing offloads tasks to powerful servers. However, current AI inference models exhibit significant heterogeneity across various platforms, posing challenges to edge offloading implementation. The complexity of managing these heterogeneous models is exacerbated by the super-large scale of smartphones and the need for realtime information exchange. These challenges, stemming from prohibitive communication overheads, render current offloading approaches economically undesirable. This paper proposes LOSA, a decentralized task-offloading scheme for edge-assisted AI inference systems with three main components. First, a performance estimator to measure the gains in offloading diverse AI model execution to the edge. Second, a workload forecaster that predicts the future workload based on historical traces. Last, a scheduler determines the optimal place to execute the model by solving a stochastic optimization problem with local model performance and workload prediction. Simulation results show how LOSA outperforms the current solutions while maintaining low communication overhead.
Yiying Lin, Su Yao, Xingyan Chen
HPCC2
2025 PIPE: Identity-Aware Privacy-Enhanced Source and Path Verification for Strengthened Network Accountability
abstract
Network-layer security threats have become increasingly sophisticated, exposing significant vulnerabilities in the current Internet architecture. Despite various proposed solutions, the field faces fundamental challenges in balancing user privacy with network security and achieving practical deployment. This paper presents a novel approach called PIPE (Privacy-preserving Identity and Path Enhancement), which leverages a distributed infrastructure of Key Distribution Servers (KDS) to integrate Decentralized Identity (DID) with source and path verification. Our solution binds user identity, address, path, and data while maintaining privacy through encryption and per-hop address transformation. By embedding DID information in address and implementing encrypted path verification, we achieve enhanced network accountability without compromising privacy. Experimental results demonstrate PIPE’s practicality and advantages over existing approaches in terms of deployment flexibility and security guarantees. Our work contributes to the evolution of secure network architectures by balancing accountability requirements with privacy protection while ensuring practical deployability.
Jianfeng Guan, Kexian Liu, Su Yao, Ye Qin, Songtao Fu, Ke Xu 0002
ICNP3
2025 Optimizing Distributed LLM Serving through Request Scheduling and Key-Value Cache Sharing
abstract
The widespread deployment of Large Language Models (LLMs) is often constrained by the significant computational and memory demands of the inference process. A critical bottleneck in distributed serving systems arises from the redundant processing of requests that share common prefixes, such as system prompts or few-shot examples. Traditional contentagnostic load balancers fail to exploit these redundancies, leading to inefficient resource utilization and increased latency. This paper introduces a dynamic, prefix-aware request scheduling system designed to optimize distributed LLM serving. Our approach intelligently routes incoming requests to specific GPU workers by analyzing prompt content and matching it with the resident Key-Value (KV) caches across the cluster. By colocating requests with shared prefixes, our system maximizes KV cache reuse, minimizes expensive prefill computations, and enables more efficient batched attention operations at the worker level. We implemented and evaluated this scheduler on a 12 -node GPU cluster using a real-world chatbot workload. The results demonstrate the profound impact of content-aware scheduling: our system increased the aggregate prefill throughput by over 144 % and reduced the median Time-to-First-Token by over 40 % compared to a conventional Round-Robin policy. These performance gains are a direct result of an 82 % relative increase in the prefix cache hit rate, validating our approach as a highly effective and cost-efficient strategy for enhancing the throughput and responsiveness of large-scale LLM services.
Hongye Jiang, Su Yao, Cui Ting, Changqiao Xu
ICPADS3
2025 Malva: A Jitter-Aware Online Pruning Framework for DNN Inference Tasks
abstract
In fields like autonomous driving, strict constraints are imposed on the computing latency of deep neural network (DNN) inference tasks on edge servers. However, it is typical for edge servers to execute multiple tasks in parallel to serve multiple users, causing severe latency jitter due to resource competition, which seriously affects timeliness. Existing works ignore the computing jitter and regard computing latency as a deterministic value, failing to meet the timeliness requirement. To address this issue, we propose Malva, a framework for finegrained online pruning for DNN tasks, allowing flexible pruning at runtime based on jitter conditions. Specifically, Malva first partitions the DNN model into blocks and applies early exiting and pruning methods to create block variants. Then, the Malva scheduler flexibly selects the variant to be executed or exits early according to urgency and jitter conditions. Moreover, we propose a novel urgency-aware prediction strategy to estimate the accuracy impact of variants with incomplete pathway information during scheduling. Stress testing shows Malva can strictly maintain a zero deadline miss rate and significantly increase the stress required to cause the first deadline miss while still outperforming state-of-the-art methods in accuracy.
Ziyan Fu 0001, Yongheng Deng, Yingjun Wu, Zhibo Wang 0001, Su Yao, Yaoxue Zhang, Ju Ren 0001
IWQoS6
2025 Pegasus: A Universal Framework for Scalable Deep Learning Inference on the Dataplane
abstract
The paradigm of Intelligent DataPlane (IDP) embeds deep learning (DL) models on the network dataplane to enable intelligent traffic analysis at line-speed. However, the current use of the match-action table (MAT) abstraction on the dataplane is misaligned with DL inference, leading to several key limitations, including accuracy degradation, limited scale, and lack of generality. This paper proposes Pegasus to address these limitations. Pegasus translates DL operations into three dataplane-oriented primitives to achieve generality: Partition, Map, and SumReduce. Specifically, Partition "divides" high-dimensional features into multiple low-dimensional vectors, making them more suitable for the dataplane; Map "conquers" computations on the low-dimensional vectors in parallel with the technique of Fuzzy Matching, while SumReduce "combines" the computation results. Additionally, Pegasus employs Primitive Fusion to merge computations, improving scalability. Finally, Pegasus adopts full-precision weights with fixed-point activations to improve accuracy. Our implementation on a P4 switch demonstrates that Pegasus can effectively support various types of DL models, including Multi-Layer Perceptron (MLP), Recurrent Neural Network (RNN), Convolutional Neural Network (CNN), and AutoEncoder models on the dataplane. Meanwhile, Pegasus outperforms state-of-the-art approaches with an average accuracy improvement of up to 22.8%, along with up to 248× larger model size and 212× larger input scale.
Yinchao Zhang, Su Yao, Kang Chen 0001, Tong Li 0014, Zhuotao Liu, Yi Zhao 0011, Lexuan Zhang, Qi Li 0002, Ke Xu 0002
SIGCOMM2
2025 Blockchain-based cross-domain IoT data sharing: A lightweight, secure edge-assisted approach
Kexian Liu, Jianfeng Guan, Su Yao, Hongke Zhang
Comput. Networks3
2025 SecFFT: Safeguarding Federated Fine-Tuning for Large Vision Language Models Against Covert Backdoor Attacks in IoRT Networks
abstract
As the large vision language models (LVLMs) and embodied intelligent robotic networks continue to advance at a remarkable pace, particularly in applications spanning smart cities, power grids, factories, and transportation, visual perception and understanding have emerged as foundational elements to overcoming performance limitations in such intelligent systems. However, since general pretrained models are not well-suited to specific tasks, federated fine-tuning (FFT) has gained attention as a promising technique for enhancing the performance of vision-based perception models by leveraging data and computational power distributed across nodes. The rise of advanced persistent threats has revealed significant vulnerabilities in existing defense mechanisms, which struggle to mitigate sophisticated backdoor attacks toward FFT for LVLMs. To address these challenges, this article proposes the SecFFT method, which tackles both the stealthiness and complexity of backdoor strategies. The approach incorporates instantaneous attack behavior detection based on frequency-domain distribution consistency and introduces a long-term secure aggregation mechanism aimed at identifying hidden attack intentions. These strategies effectively limit the feasibility of adversaries attempting to bypass defense measures by concealing their behaviors. Experiments conducted on public datasets demonstrate that SecFFT significantly improves defense success rates, model performance, and detection accuracy, particularly in response to highly covert, multiround backdoor attacks.
Zan Zhou 0001, Changqiao Xu, Sizhe Huang, Su Yao
IEEE Internet Things J.7
2025 LEOEdge: A Satellite-Ground Cooperation Platform for the AI Inference in Large LEO Constellation
abstract
With the rapid growth of low earth orbit (LEO) satellites, enabling LEO AI inference becomes a fast-increasing trend. However, due to resource heterogeneity, scheduling complexity, and fast movement, how to decide the place of executing each AI inference task is nontrivial in LEO systems. In this paper, we propose LEOEdge, an edge-assisted AI inference system for LEO satellites. We first introduce the adaptive modeling technologies that automatically generate the model for each satellite according to its computation resources. We then propose a layered scheduling optimization scheme to schedule the AI inference task in a distributed manner. LEOEdge also designs a seamless data transmission scheme to avoid transmission failure due to the LEO satellite movement. We conduct a series of simulation tests to validate the performance of the proposed LEOEdge, in terms of the neural network searching efficiency, average time execution latency, and delivery latency.
Su Yao, Yiying Lin, Ke Xu 0002, Mingwei Xu 0001, Changqiao Xu, Hongke Zhang
IEEE J. Sel. Areas Commun.1
2025 Secure Fault Localization in Path Aware Networking
abstract
Secure data forwarding is critical for users to meet their requirements. In this paper, we propose D3 (Demon Detector in Data Plane), a source-driven, secure fault localization mechanism, which empowers the source to localize faulty link in Path Aware Networking, thus circumventing faulty link to guarantee secure data forwarding. D3 utilizes the source to instruct the on-path routers, thus empowering it to detect whether the on-path routers forward the packet as expected. Compared with existing schemes that are difficult to be deployed in practice due to the heavy storage, computation, and communication overhead, D3 offloads most of the on-path router's storage and computation overhead, thus dramatically improving the deployment efficiency. Particularly, the length of the additional packet header in D3 is 2-5 times less than the state-of-the-art mechanisms, thus having a low communication overhead. Besides that, the destination in D3 could keep stateless processing, thus having backward compatibility and eliminating the opportunity for DoS attacks toward a stateful destination. The BMv2 and Barefoot Tofino hardware evaluations show that D3 could achieve high fault localization accuracy and process the packet at line rate.
Songtao Fu, Qi Li 0002, Xiaoliang Wang 0004, Su Yao, Xuewei Feng, Xinle Du, Kao Wan, Ke Xu 0002
IEEE Trans. Dependable Secur. Comput.4
2025 Multi-Agent Reinforcement Learning for Task Offloading in Crowd-Edge Computing
abstract
The Crowd-edge (CE) computing paradigm facilitates the utilization of the computational resources through simultaneously relying the edge computing and the collaboration among various mobile devices (MDs). Most existing works, focusing on offloading tasks from device to edge servers by centralized solutions, are unable to distribute tasks to massive MDs in CE. Meanwhile, designing a decentralized task offloading solution enabling task subscribers to individually make offloading decisions can be challenging given the randomness of crowd resource provisioning and limited knowledge of global status variations. In this paper, we propose a decentralized crowd-edge task offloading solution that enables users to optimally offload tasks to the CE in a distributed manner. Specifically, we formulate the corresponding problem as a stochastic optimization with partially observable status. By observing network and process delays at the crowd side, we further reform the optimization forms and provide a novel approximation policy, enabling users to optimize their offloading strategy based on local observations without interaction with each other. We then solve this task offloading problem by developing a Mixed Multi-Agent Proxy Policy Optimization algorithm (mixed MAPPO). Extensive testing, including numerical and system-level simulations, was conducted to validate the performance of the proposed algorithm in terms of task delay (including the processing delay and transmission delay), load rate, and resource utilization.
Su Yao, Ju Ren 0001, Weiqiang Wang 0002, Ke Xu 0002, Mingwei Xu 0001, Hongke Zhang
IEEE Trans. Mob. Comput.1
2024 Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
abstract
The common toxicity and societal bias in contents generated by large language models (LLMs) necessitate strategies to reduce harm.Present solutions often demand whitebox access to the model or substantial training, which is impractical for cutting-edge commercial LLMs.Moreover, prevailing prompting methods depend on external tool feedback and fail to simultaneously lessen toxicity and bias.Motivated by social psychology principles, we propose a novel strategy named perspective-taking prompting (PET) that inspires LLMs to integrate diverse human perspectives and self-regulate their responses.This self-correction mechanism can significantly diminish toxicity (up to 89%) and bias (up to 73%) in LLMs' responses.Rigorous evaluations and ablation studies are conducted on two commercial LLMs (ChatGPT and GLM) and three open-source LLMs, revealing PET's superiority in producing less harmful responses, outperforming five strong baselines."Words kill, words give life; they're either poison or fruit-you choose."~Proverbs 18:21 (MSG)
Rongwu Xu, Zi'an Zhou, Tianwei Zhang 0004, Zehan Qi, Su Yao, Ke Xu 0002, Wei Xu 0039, Han Qiu 0001
EMNLP5
2024 DKGAuth: Blockchain-Assisted Distributed Key Generation and Authentication for Cross-Domain Intelligent IoT
abstract
The widespread adoption of intelligent Internet of Things (IoT) has sparked increased efforts to foster extensive data interaction and collaboration across diverse fields, leading to a trust crisis in cross-domain scenarios. Moreover, cross-domain collaboration increases the complexity of key management, especially in resource-constrained IoT environments where high computational costs are impractical. This situation poses risks of key leakage and inefficient key updates. This paper introduces DKGAuth, a blockchain-based method for distributed key generation and authentication tailored for resource-constrained cross-domain intelligent IoT systems. Initially, we propose a lightweight cross-domain authentication architecture based on blockchain to address the trust crisis effectively among different domains in the intelligent IoT. Secondly, building upon this architecture, we introduce a distributed key generation method that revolutionizes the key infrastructure to address key management concerns. Additionally, we design an algorithm to combine key factors, minimizing costs associated with both key generation and updates. Finally, we establish a simulation environment to assess the computational, storage, and read/write overheads of our approach. In the same configuration, compared to other solutions, the efficiency of key updates improves by 83% when updated 100 times.
Kexian Liu, Jianfeng Guan, Su Yao, Hongke Zhang
IEEE Internet Things J.3
2024 FedPAGE: Pruning Adaptively Toward Global Efficiency of Heterogeneous Federated Learning
abstract
When workers are heterogeneous in computing and transmission capabilities, the global efficiency of federated learning suffers from the straggler issue, i.e., the slowest worker drags down the overall training process. We propose a novel and efficient federated learning framework named FedPAGE, where workers perform distributed pruning adaptively towards global efficiency, i.e., fast training and high accuracy. For fast training, we develop a pruning rate learning approach generating an adaptive pruning rate for each worker, making the overall update time approximate to the fastest worker’s update time, i.e., no stragglers. For high accuracy, we find that structural similarity between sub-models is essential to global model accuracy in the distributed pruning, and thus propose the CIG_X pruning scheme to ensure maximum similarity. Meanwhile, we adopt the sparse training and design model aggregating of different size sub-models to cope with distributed pruning. We prove the convergence of FedPAGE and demonstrate the effectiveness of FedPAGE on image classification and natural language inference tasks. Compared with the state-of-the-art, FedPAGE achieves higher accuracy with the same speedup ratio.
Guangmeng Zhou, Qi Li 0002, Yang Liu 0038, Yi Zhao 0011, Qi Tan 0003, Su Yao, Ke Xu 0002
IEEE/ACM Trans. Netw.6
2023 Differentially Private Learning with Per-Sample Adaptive Clipping
abstract
Privacy in AI remains a topic that draws attention from researchers and the general public in recent years. As one way to implement privacy-preserving AI, differentially private learning is a framework that enables AI models to use differential privacy (DP). To achieve DP in the learning process, existing algorithms typically limit the magnitude of gradients with a constant clipping, which requires carefully tuned due to its significant impact on model performance. As a solution to this issue, latest works NSGD and Auto-S innovatively propose to use normalization instead of clipping to avoid hyperparameter tuning. However, normalization-based approaches like NSGD and Auto-S rely on a monotonic weight function, which imposes excessive weight on small gradient samples and introduces extra deviation to the update. In this paper, we propose a Differentially Private Per-Sample Adaptive Clipping (DP-PSAC) algorithm based on a non-monotonic adaptive weight function, which guarantees privacy without the typical hyperparameter tuning process of using a constant clipping while significantly reducing the deviation between the update and true batch-averaged gradient. We provide a rigorous theoretical convergence analysis and show that with convergence rate at the same order, the proposed algorithm achieves a lower non-vanishing bound, which is maintained over training iterations, compared with NSGD/Auto-S. In addition, through extensive experimental evaluation, we show that DP-PSAC outperforms or matches the state-of-the-art methods on multiple main-stream vision and language tasks.
Shuheng Shen, Su Yao, Ke Xu 0002
AAAI3
2023 Neuron Pruning-Based Federated Learning for Communication-Efficient Distributed Training
Jianfeng Guan, Su Yao
ICA3PP (4)3
2023 MASK: Practical Source and Path Verification Based on Multi-AS-Key
abstract
The source and path verification in Path-Aware Networking considers the two critical issues: (1) end hosts could verify that the network follows their forwarding decisions, and (2) both on-path routers and destination host could authenticate the source of packets and filter the malicious traffic. Unfortunately, the state-of-the-art mechanisms require heavy communication overhead in the network and computation overhead in the router; moreover, it is difficult to meet the dynamic requirements of the end host. We propose a user-driven mechanism, source and path verification based on Multi-AS-Key (MASK). MASK decreases the communication overhead by a short additional packet header and reduces the computation overhead by separating the control and data plane in terms of the cryptographic operation. Furthermore, it utilizes the stateful user to instruct the stateless routers to process the packet with a user-driven policy, thus satisfying the user’s requirements such as detecting the packet drop and replay attack. With the plausible design, the communication overhead for realistic path lengths is 1/2 to 1/10 compared with the state-of-the-art mechanisms. We implement MASK in the BMv2 environment and commodity Barefoot Tofino programmable switch, testify that MASK introduces significantly less overhead than the state-of-the-art mechanisms, and demonstrate that MASK could achieve the verification in the programmable switch at line rate.
Songtao Fu, Qi Li 0002, Xiaoliang Wang 0004, Su Yao, Yangfei Guo, Xinle Du, Ke Xu 0002
IEEE/ACM Trans. Netw.5
2022 D3: Lightweight Secure Fault Localization in Edge Cloud
abstract
In pursuit of high-performance applications, the cloud is moving out of the data center and towards the edge. Secure data forwarding is critical for the users between the edge and the remote cloud. In this paper, we propose D3 (Demon Detector in Data Plane), a lightweight, secure fault localization mechanism, which can enable the users in the edge cloud to localize faulty links and thus avoid the faulty links to guarantee secure data forwarding along the path to the remote cloud. D3 utilizes the user to instruct the transit routers, thus empowering the user to detect whether the transit routers forward the packet as expected. Compared with existing schemes that are difficult to be deployed in practice due to the incurred heavy storage, computation, and communication overhead, D3 offloads most of the transit router’s storage and computation overhead, thus dramatically improving the deployment efficiency. Particularly, the length of the additional packet header in D3 is 2-5 times less than the state-of-the-art mechanisms, and the extra control packet overhead is ten times less while keeping a little constant storage overhead in the data plane. The evaluations in BMv2 and Barefoot Tofino hardware show that D3 could achieve high fault localization accuracy and efficiency.
Songtao Fu, Qi Li 0002, Xiaoliang Wang 0004, Su Yao, Xuewei Feng, Xinle Du, Kao Wan, Ke Xu 0002
ICDCS4
2022 Blockchain-Empowered Collaborative Task Offloading for Cloud-Edge-Device Computing
abstract
How to enable high-performance task offloading and preserve the trust between participants is imperative yet nontrivial to the Cloud-Edge-Device (CED) computing, mainly because the resources are geo-distributed and operated by different parties. Also, the CED participants are highly dynamic and heterogeneous in resource provision and may conflict in interest. This paper proposes BlockChain-empowered CED (BC-CED), a blockchain-empowered collaborative task offloading for CED computing. In BC-CED, blockchain plays a central role in the main functionality of CED, including task offloading, brokerage of resource usage, and incentives. We distinguish the BC-CED from the existing solutions by modifying the blockchain consensus process, enabling the participants to reach an agreement via solving the task offloading problem. For this purpose, we formulate the offloading problem by considering the computation capabilities of candidate nodes and the network performance. BC-CED allows each participant to apply reinforcement learning-based methods to solve this problem and compete for the right of block output by comparing the offloading policy performance and accepting the best policy as the offloading scheme within the next period. We also propose a truthful incentive mechanism to encourage resource contributions in BC-CED and force them to be honest. Extensive tests by implementing our solutions in a commercialized blockchain platform have shown how BC-CED achieves a superior performance in task offloading and blockchain maintenance.
Su Yao, Qiang Qu 0001, Ke Xu 0002, Mingwei Xu 0001
IEEE J. Sel. Areas Commun.1
2022 Multi-layer pseudo-supervision for histopathology tissue semantic segmentation using patch-level classification labels
abstract
Tissue-level semantic segmentation is a vital step in computational pathology. Fully-supervised models have already achieved outstanding performance with dense pixel-level annotations. However, drawing such labels on the giga-pixel whole slide images is extremely expensive and time-consuming. In this paper, we use only patch-level classification labels to achieve tissue semantic segmentation on histopathology images, finally reducing the annotation efforts. We propose a two-step model including a classification and a segmentation phases. In the classification phase, we propose a CAM-based model to generate pseudo masks by patch-level labels. In the segmentation phase, we achieve tissue semantic segmentation by our propose Multi-Layer Pseudo-Supervision. Several technical novelties have been proposed to reduce the information gap between pixel-level and patch-level annotations. As a part of this paper, we introduce a new weakly-supervised semantic segmentation (WSSS) dataset for lung adenocarcinoma (LUAD-HistoSeg). We conduct several experiments to evaluate our proposed model on two datasets. Our proposed model outperforms five state-of-the-art WSSS approaches. Note that we can achieve comparable quantitative and qualitative results with the fully-supervised model, with only around a 2% gap for MIoU and FwIoU. By comparing with manual labeling on a randomly sampled 100 patches dataset, patch-level labeling can greatly reduce the annotation time from hours to minutes. The source code and the released datasets are available at: https://github.com/ChuHan89/WSSS-Tissue.
Chu Han, Jiatai Lin, Jinhai Mai, Yi Wang 0031, Qingling Zhang 0006, Bingchao Zhao, Xin Chen 0058, Xipeng Pan, Zhenwei Shi 0002, Zeyan Xu, Su Yao, Lixu Yan, Xiaomei Huang, Changhong Liang, Guoqiang Han 0002, Zaiyi Liu
Medical Image Anal.11
2021 Adaptive Convolutional Neural Network Structure for Network Traffic Classification
abstract
Network traffic classification has been highly concerned by academia and industry for decades. In recent years, deep learning has attracted many scholars to use it in network traffic classification due to its excellent performance in the fields of computer vision and natural language processing. However, the performance of the neural network depends on its structure in the same dataset. When looking for the neural network to classify network traffic, it is necessary to constantly adjust the structure of the neural network to achieve better results, which is very time-consuming and experience-dependent. To solve the above problem, this paper proposes an Adaptive Convolutional Neural Network Structure for Network Traffic Classification (ACNNS-NTC) algorithm. The proposed algorithm first pre-processes the network traffic data used for training and testing, and then uses particle swarm optimization algorithm to optimize the network structure of the convolutional neural network, to generate convolutional neural network structure for network traffic classification, and verify the classification results. Experimental results show that the accuracies of the ACNNS-NTC algorithm on public datasets (ISCX-IDS2012, USTC-TFC2016, CIC-IDS2017) are above 99%. At the same time, the generated convolutional neural network has a more succinct structure and fewer model parameters compared with the existing methods.
Zhuang Han, Jianfeng Guan, Yanan Yao, Su Yao
ICPADS4
2021 MASK: Practical Source and Path Verification based on Multi-AS-Key
abstract
The source and path verification in path-aware Internet consider the two critical issues: (1) end hosts could verify that their forwarding decisions followed by the network, (2) both intermediate routers and destination host could authenticate the source of packets and filter the malicious traffic. Unfortunately, the current verification mechanism requires validation operations in each router on the path in an inter-domain environment, thus requiring high communication and computation overhead, reducing its usefulness; besides, it is also difficult to meet the dynamic requirements of the end host. Ideally, the verification should be secure and provide the customized capability to meet the end host’s requirements. We propose a new mechanism called source and path verification based on Multi-AS-Key (MASK). Instead of each packet verified and marked at each router on the path, MASK improves the verification by empowering the end hosts to instruct the routers to achieve the verification, thus decreasing the router’s overhead while ensuring security performance to meet the end host’s requirements. With the plausible design, the communication overhead for realistic path lengths is 3–8 times smaller than the state-of-the-art mechanisms. The computation overhead in the routers is 2-5 times smaller. We implement our design in the BMv2 environment and commodity Barefoot Tofino programmable switch, demonstrating that MASK introduces significantly less overhead than the existing mechanisms.
Songtao Fu, Ke Xu 0002, Qi Li 0002, Xiaoliang Wang 0004, Su Yao, Yangfei Guo, Xinle Du
IWQoS5
2021 BSLA: Blockchain-assisted Secure and Lightweight Authentication for SGIN
Jianfeng Guan, Su Yao, Tianhong Zhang, Xiaokang Su, Chuanqing Li
Comput. Commun.3
2020 Stochastic Cost Minimization Mechanism Based on Identifier Network for IoT Security
abstract
An identifier network (IN), as one of the promising network architectures to solve the IP dual properties problems, has been applied in many areas, including Internet of Things (IoT) for smart cities scenario. The separation mechanisms of an access/core network and identifier/location can benefit IoT in terms of trust management, access control, and privacy protection. The core network in IN is independent of the access network by introducing two namespaces, which makes the core network difficult to be attacked but easy for trust management. However, access network, such as access wireless sensor network (WSN), is facing serious trust and security challenge. Therefore, this article addresses this problem by using network address shuffling. We present an optimization framework of defense cost for IoT security and formulate it as a stochastic cost optimization problem by considering the impacts of network address shuffling control, network autoimmunity control, and defense cost. To improve its generality, we adopt a Lyapunov optimization theory and transform the formulated optimization problem into a queue stability problem, and further decompose the queue stability problem into two subproblems to solve the initial optimization problem. Finally, a novel stochastic cost minimization mechanism (SCMM) consisting of two algorithms for the derived subproblems is proposed. Through logical theoretical analyses, it is proved that our proposed mechanism can achieve the optimized results while maintaining the security of each access WSN and guaranteeing the limited network resources. The extensive simulation results verify that the tradeoff between defense strategy and limited network resources can be well tackled by the balancing factor.
Su Yao, Jianfeng Guan, Yang Liu 0038
IEEE Internet Things J.1
2020 COMP: Online Control Mechanism for Profit Maximization in Privacy- Preserving Crowdsensing
abstract
As a novel sensing paradigm, crowdsensing has gained great attention due to large-scale user participation, low cost and wide data source, replacing traditional sensor based sensing in intelligent transportation, environmental monitoring, urban public management, etc. In crowdsensing, however, user privacy leakage is a common but fatal problem, where the participants in crowdsensing might not provide their data if their sensing data expose their personal private information or even lead to malicious attacks. Additionally, it is still challenging for the platform to consider the randomness of sensing task arrival, the dynamic participation of participants and the complexity of task allocation. To this end, an online control mechanism is presented to maximize the profit of platform while guaranteeing system stability and providing personalized location privacy protection. By exploiting Lyapunov optimization theory, we transform the optimization problem into a queue stability problem, decomposing it into three subproblems further. Through rigorous theoretical analysis, we prove that our time-averaged profit is approximately optimal. We also carry out extensive simulations to verify the superiority of our proposed mechanism.
Yang Liu 0038, Tong Feng, Mugen Peng, Zhongbai Jiang, Jianfeng Guan, Su Yao
IEEE J. Sel. Areas Commun.7
2020 Triple U-net: Hematoxylin-aware nuclei segmentation with progressive dense feature aggregation
Bingchao Zhao, Xin Chen 0058, Zhiwen Yu 0002, Su Yao, Lixu Yan, Zaiyi Liu, Changhong Liang, Chu Han
Medical Image Anal.5
2017 GBC-based caching function group selection algorithm for SINET
Jianfeng Guan, Zhiwei Yan, Su Yao, Changqiao Xu, Hongke Zhang
J. Netw. Comput. Appl.3
2016 The Cache Location Selection Based on Group Betweenness Centrality Maximization
Jianfeng Guan, Zhiwei Yan, Su Yao, Changqiao Xu, Hongke Zhang
QSHINE3