Linshan Jiang

dblp:183/1884 · DBLP profile ↗
← Back
41ranked-venue papers
6as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 4 first-author · 10 since 2021Systems, architecture and hardware · 10 · 6 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 'They've Stolen My GPL-Licensed Model!': Toward Standardized and Transparent Model Licensing
abstract
As model parameter sizes scale into the billions and training consumes zettaFLOPs of computation, the reuse of Machine Learning (ML) assets and collaborative development have become increasingly prevalent in the ML community. These ML assets, including models, datasets, and software, may originate from various sources and be published under different licenses, which govern the use and distribution of licensed works and their derivatives. However, commonly chosen licenses, such as GPL and Apache, are software-specific and are not clearly defined or bounded in the context of model publishing. Meanwhile, the reused assets may also be under free-content licenses and model licenses, which pose a potential risk of license noncompliance and rights infringement within the model production workflow. In this paper, we address these challenges along two lines: 1) For ML workflow compliance, we propose ModelGo (MG) Analyzer, a tool that incorporates a vocabulary for ML workflow management and encoded license rules, enabling ontological reasoning to analyze rights granting and compliance issues. 2) For standardized model publishing, we introduce ModelGo Licenses, a set of modell-specific licenses that provide flexible options to meet the diverse needs of the ML community. MG Analyzer is built on Turtle language and Notation3 reasoning engine, envisioned as a first step toward Linked Open Data for ML workflow management. We have also encoded our proposed model licenses into rules and demonstrated the effects of GPL and other commonly used licenses in model publishing, along with the flexibility advantages of our licenses, through comparisons and experiments.
Moming Duan, Rui Zhao 0009, Linshan Jiang, Nigel Shadbolt, Bingsheng He
WWW3
2026 MASI: Memory-Adaptive Inference Framework for Spiking Neural Networks on Edge Devices
abstract
The rapid development of the Internet of Things (IoT) applications necessitates resource-efficient computing paradigms that can unify heterogeneous sensing modalities. Spiking Neural Networks (SNNs) meet this need with their event-driven and energy-efficient processing nature. However, deploying SNNs on mobile and embedded platforms is hindered by strict and fluctuating memory budgets. While prior work explores lightweight model design and system-level memory management, these methods either sacrifice accuracy or incur high runtime overhead due to timestep-dependent dynamics. To tackle these challenges, we propose a memory-adaptive framework MASI that enables efficient on-device SNN inference by combining (1) a fine-grained memory-adaptive layer slicing strategy, (2) a timestep-agnostic scheduler that maximizes memory utilization with minimal fragmentation, and (3) a timestep-aware early-exit mechanism that reduces redundant calculations. Evaluated on diverse workloads and edge devices, MASI can dynamically adapt to runtime memory availability, approximately reducing memory usage by 20.67% and inference latency by 58.53% on average with negligible accuracy loss compared to other feasible on-device implementations under memory constraints.
Di Yu 0001, Helin Zheng, Changze Lv, Xin Du 0002, Linshan Jiang, Xiang Liu 0017, Gang Pan 0001, Shuiguang Deng
WWW5
2025 Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference
abstract
Deep learning models are increasingly utilized on resource-constrained edge devices for real-time data analytics. Recently, Vision Transformer and their variants have shown exceptional performance in various computer vision tasks. However, their substantial computational requirements and low inference latency create significant challenges for deploying such models on resource-constrained edge devices. To address this issue, we propose a novel framework, ED-ViT, which is designed to efficiently split and execute complex Vision Transformers across multiple edge devices. Our approach involves partitioning Vision Transformer models into several sub-models, while each dedicated to handling a specific subset of data classes. To further reduce computational overhead and inference latency, we introduce a class-wise pruning technique that decreases the size of each sub-model. Through extensive experiments conducted on five datasets using three model architectures and actual implementation on edge devices, we demonstrate that our method significantly cuts down inference latency on edge devices and achieves a reduction in model size by up to 28.9 times and 34.1 times, respectively, while maintaining test accuracy comparable to the original Vision Transformer. Additionally, we compare ED-ViT with two state-of-the-art methods that deploy CNN and SNN models on edge devices, evaluating metrics such as accuracy, inference time, and overall model size. Our comprehensive evaluation underscores the effectiveness of the proposed ED-ViT framework.
Xiang Liu 0017, Yijun Song, Xia Li 0005, Huiying Lan, Linshan Jiang, Jialin Li 0001
ICDCS7
2025 FedWCM: Unleashing the Potential of Momentum-based Federated Learning in Long-Tailed Scenarios
abstract
Federated Learning (FL) enables decentralized model training while preserving data privacy. Despite its benefits, FL faces challenges with non-identically distributed (non-IID) data, especially in long-tailed scenarios with imbalanced class samples. Momentum-based FL methods, often used to accelerate FL convergence, struggle with these distributions, resulting in biased models and making FL hard to converge. To understand this challenge, we conduct extensive investigations into this phenomenon, accompanied by a layer-wise analysis of neural network behavior. Based on these insights, we propose FedWCM, a method that dynamically adjusts momentum using global and per-round data to correct directional biases introduced by long-tailed distributions. Extensive experiments show that FedWCM resolves non-convergence issues and outperforms existing methods, enhancing FL’s efficiency and effectiveness in handling client heterogeneity and data imbalance.
Tianle Li, Yongzhi Huang 0002, Linshan Jiang, Qipeng Xie, Chang Liu 0093, Wenfeng Du, Lu Wang 0002, Kaishun Wu
ICPP3
2025 Exploiting Label Skewness for Spiking Neural Networks in Federated Learning
abstract
The energy efficiency of deep spiking neural networks (SNNs) aligns with the constraints of resource-limited edge devices, positioning SNNs as a promising foundation for intelligent applications leveraging the extensive data collected by these devices. To safeguard data privacy, federated learning (FL) facilitates collaborative SNN-based model training by leveraging data distributed across edge devices without transmitting local data to a central server. However, existing FL approaches encounter challenges in handling label-skewed data across devices, inducing drift in the local SNN model and consequently impairing the performance of the global SNN model. To tackle these problems, we propose a novel framework called FedLEC, which incorporates intra-client label weight calibration to balance the learning intensity across local labels and inter-client knowledge distillation to mitigate local SNN model bias caused by label absence. Extensive experiments with three different structured SNNs across five datasets (i.e., three non-neuromorphic and two neuromorphic datasets) demonstrate the efficiency of FedLEC. Compared to seven state-of-the-art FL algorithms, FedLEC achieves an average accuracy improvement of approximately 11.59% for the global SNN model under various label skew distribution settings.
Di Yu 0001, Xin Du 0002, Linshan Jiang, Huijing Zhang, Shuiguang Deng
IJCAI3
2025 ECC-SNN: Cost-Effective Edge-Cloud Collaboration for Spiking Neural Networks
abstract
Most edge-cloud collaboration frameworks rely on the substantial computational and storage capabilities of cloud-based artificial neural networks (ANNs). However, this reliance results in significant communication overhead between edge devices and the cloud, as well as high computational energy consumption, especially when applied to resource-constrained edge devices. To address these challenges, we propose ECC-SNN, a novel edge-cloud collaboration framework that incorporates energy-efficient spiking neural networks (SNNs) to offload more computational workload from the cloud to the edge, thereby improving cost-effectiveness and reducing reliance on the cloud. ECC-SNN employs a joint training approach that integrates ANN and SNN models, enabling edge devices to leverage knowledge from cloud models for enhanced performance while reducing energy consumption and processing latency. Furthermore, ECC-SNN features an on-device incremental learning algorithm that enables edge models to continuously adapt to dynamic environments, reducing the communication overhead and resource consumption associated with frequent cloud update requests. Extensive experimental results on four datasets demonstrate that ECC-SNN improves accuracy by 4.15%, reduces average energy consumption by 79.4%, and lowers average processing latency by 39.1%.
Di Yu 0001, Changze Lv, Xin Du 0002, Linshan Jiang, Wentao Tong, Xiaoqing Zheng, Shuiguang Deng
IJCAI4
2025 Cost-Effective On-Device Sequential Recommendation with Spiking Neural Networks
abstract
On-device sequential recommendation (SR) systems are designed to make local inferences using real-time features, thereby alleviating the communication burden on server-based recommenders when handling concurrent requests from millions of users. However, the resource constraints of edge devices, including limited memory and computational capacity, pose significant challenges to deploying efficient SR models. Inspired by the energy-efficient and sparse computing properties of deep Spiking Neural Networks (SNNs), we propose a cost-effective on-device SR model named SSR, which encodes dense embedding representations into sparse spike-wise representations and integrates novel spiking filter modules to extract temporal patterns and critical features from item sequences, optimizing computational and memory efficiency without sacrificing recommendation accuracy. Extensive experiments on real-world datasets demonstrate the superiority of SSR. Compared to other SR baselines, SSR achieves comparable recommendation performance while reducing energy consumption by an average of 59.43%. In addition, SSR significantly lowers memory usage, making it particularly well-suited for deployment on resource-constrained edge devices.
Di Yu 0001, Changze Lv, Xin Du 0002, Linshan Jiang, Qing Yin, Wentao Tong, Xiaoqing Zheng, Shuiguang Deng
IJCAI4
2025 One-shot Federated Learning Methods: A Practical Guide
abstract
One-shot Federated Learning (OFL) is a distributed machine learning paradigm that constrains client-server communication to a single round, addressing privacy and communication overhead issues associated with multiple rounds of data exchange in traditional Federated Learning (FL). OFL demonstrates the practical potential for integration with future approaches that require collaborative training models, such as large language models (LLMs). However, current OFL methods face two major challenges: data heterogeneity and model heterogeneity, which result in subpar performance compared to conventional FL methods. Worse still, despite numerous studies addressing these limitations, a comprehensive summary is still lacking. To address these gaps, this paper presents a systematic analysis of the challenges faced by OFL and thoroughly reviews the current methods. We also offer an innovative categorization method and analyze the trade-offs of various techniques. Additionally, we discuss the most promising future directions and the technologies that should be integrated into the OFL field. This work aims to provide guidance and insights for future research.
Xiang Liu 0017, Zhenheng Tang, Xia Li 0005, Yijun Song, Sijie Ji, Bo Han 0003, Linshan Jiang, Jialin Li 0001
IJCAI8
2025 HARMONY: A Privacy-preserving and Sensor-agnostic Tele-monitoring system
abstract
Global aging necessitates tele-monitoring systems to provide real-time tracking and timely assistance for older adults living independently. While pervasive wireless devices (e.g., CSI, IMU, UWB) enable cost-effective, non-intrusive monitoring, existing systems lack flexibility, limiting their adaptability to different environments. In this work, we posit that the motion dynamics of human movement are invariant across sensing modalities, inspiring the design of HARMONY—a privacy-preserving, sensor-agnostic system that supports multi-modal inputs and diverse tele-monitoring tasks. HARMONY incorporates Modality-agnostic Data Processing to uniformly encrypt multi-modal signals and Task-specific Activity Recognition for seamless tasks adaptation. A novel Encrypted-processing Engine then significantly accelerates computations on encrypted data by optimizing matrix and convolution operations. Evaluations across five different sensing modalities show that HARMONY consistently achieves high accuracy while delivering 3.5 × to 130 × speedups over state-of-the-art baselines. Our results demonstrate that HARMONY is a practical, scalable, and privacy-centric prototype for next-generation remote healthcare.
Qipeng Xie, Weizheng Wang 0001, Yongzhi Huang 0002, Linshan Jiang, Jiafei Wu, Shuxin Zhong, Lu Wang 0002, Kaishun Wu
IJCAI5
2025 GrainSegNet: Towards Effective and Efficient Material Microstructure Segmentation
Ruoyu Tian, Xuecheng Zhang, Yingyao Wang, Chaojie Gu, Yuefei Zhang, Linshan Jiang
INDIN7
2025 FedWMSAM: Fast and Flat Federated Learning via Weighted Momentum and Sharpness-Aware Minimization
abstract
In federated learning (FL), models must \emph{converge quickly} under tight communication budgets while \emph{generalizing} across non-IID client distributions. These twin requirements have naturally led to two widely used techniques: client/server \emph{momentum} to accelerate progress, and \emph{sharpness-aware minimization} (SAM) to prefer flat solutions. However, simply combining momentum and SAM leaves two structural issues unresolved in non-IID FL. We identify and formalize two failure modes: \emph{local–global curvature misalignment} (local SAM directions need not reflect the global loss geometry) and \emph{momentum-echo oscillation} (late-stage instability caused by accumulated momentum). To our knowledge, these failure modes have not been jointly articulated and addressed in the FL literature. We propose \textbf{FedWMSAM} to address both failure modes. First, we construct a momentum-guided global perturbation from server-aggregated momentum to align clients' SAM directions with the global descent geometry, enabling a \emph{single-backprop} SAM approximation that preserves efficiency. Second, we couple momentum and SAM via a cosine-similarity adaptive rule, yielding an early-momentum, late-SAM two-phase training schedule. We provide a non-IID convergence bound that \emph{explicitly models the perturbation-induced variance} $\sigma_\rho^2=\sigma^2+(L\rho)^2$ and its dependence on $(S,K,R,N)$ on the theory side. We conduct extensive experiments on multiple datasets and model architectures, and the results validate the effectiveness, adaptability, and robustness of our method, demonstrating its superiority in addressing the optimization challenges of Federated Learning. Our code is available at \url{https://github.com/Li-Tian-Le/NeurlPS_FedWMSAM}.
Tianle Li, Yongzhi Huang 0002, Linshan Jiang, Chang Liu 0093, Qipeng Xie, Wenfeng Du, Lu Wang 0002, Kaishun Wu
NeurIPS3
2025 QoE-Optimized MultiPath Scheduling for Video Services in Large-Scale Peer-to-Peer CDNs
abstract
Video content providers such as Douyin implement Peer-to-Peer Content Delivery Networks (PCDNs) to reduce the costs associated with Content Delivery Networks (CDNs) while still maintaining optimal user-perceived quality of experience (QoE). PCDNs rely on the remaining resources of edge devices, such as edge access devices and hosts, to store and distribute data with a Multiple-Server-to-One-Client (MS2OC) communication pattern. MS2OC parallel transmission pattern suffers from severe data out-of-order issues. PCDNs offer significant cost savings by using multiple low-cost edge devices. However, due to its unique characteristics, including pull-based streaming transmission, many heterogeneous paths, and large receiving buffers, directly applying existing schedulers designed for Multipath TCP (MPTCP) to PCDN fails to meet the two goals of high aggregate bandwidth and low end-to-end delivery latency. To tackle this issue, we provide a detailed overview of Douyin’s self-developed PCDN video transmission system and introduce the first QoE-enhanced packet-level scheduler for PCDN systems, named Pscheduler. Pscheduler evaluates path quality with a congestion-control-decoupled algorithm and employs our proposed path-pick-packet method for data distribution, ensuring a smooth video playback experience. Additionally, we propose a redundant transmission algorithm to enhance task download speeds for segmented video transmission. Our extensive online A/B tests, involving 100,000 Douyin users generating tens of millions of video data points, demonstrate that Pscheduler achieves an average improvement of 60% in goodput, a 20% reduction in data delivery waiting time, and a 30% reduction in rebuffering rates. Furthermore, we conducted simulation experiments that further validate the effectiveness of Pscheduler, confirming its improvements in performance metrics under various network conditions.
Dehui Wei, Jiao Zhang 0002, Xiang Liu 0017, Zhichen Xue, Tao Huang 0005, Linshan Jiang, Jialin Li 0001
IEEE J. Sel. Areas Commun.7
2025 RoboCam: Model-Based Robotic Visual Sensing for Precise Inspection of Mesh Screens
abstract
The 3D-printed mesh screen with dense penetrating pores is a new structure for massive manufacturing of molded pulp package products. However, some of the pores may be clogged by the printing material powder during the printing process. Such defects negatively affect the quality of the pulp packages produced using the mesh screen mold. To pinpoint the defects, we design a model-based robotic visual sensing system, called RoboCam, which uses a robotic arm to carry a high-resolution camera for full inspection of a mold consisting of joined mesh screens. To inspect the entire mold, RoboCam plans the camera poses to capture multiple images of the mold and render synthesized images as references for identifying the clogged pores. In particular, we propose novel designs to rectify the inherent run-time pose errors of the robotic system for ensuring the reference quality and to accelerate the reference rendering for reducing inspection latency. Extensive evaluation shows that RoboCam’s design outperforms various baselines, including three existing computer vision and convolution neural network-based inspection systems. RoboCam achieves a recall rate of 94.95% within 528 seconds latency for inspecting an entire mold with 13,000 designed pores.
Duc Van Le, Linshan Jiang, Zhuoran Chen, Xiaohua Peng, Daren Ho, Jianmin Zheng, Rui Tan 0001
ACM Trans. Sens. Networks3
2025 SNN-IoT: Efficient Partitioning and Enabling of Deep Spiking Neural Networks in IoT Services
abstract
Spiking Neural Networks (SNNs), due to their inherent biological plausibility and energy-saving characteristics, naturally align with the requirements of IoT services. However, current SNNs require a multi-layer structure to achieve effective applications across various fields. The multi-layer deep SNNs with massive model parameters demand computational resources, rendering them incompatible with resource-constrained IoT devices. To address this problem, in this work, a deep SNN partitioning framework called SNN-IoT is proposed to run complex SNN models on IoT devices. The SNN-IoT first partitions a full deep SNN model into smaller sub-models, leveraging the event-driven sparsity of SNNs and channel-level firing patterns to distribute filters with lower levels of spike activity onto devices with more constrained resources. The SNN model partitioning and deployment is formulated as an optimization problem and is solved using a greedy search assignment mechanism. Furthermore, a channel-wise pruning method exploits the varying degrees of channel activity, effectively reducing each sub-model's size and computational load without compromising performance. Extensive experiments conducted on four non-neuromorphic and two neuromorphic datasets have demonstrated that the SNN-IoT framework not only efficiently partitions deep SNNs and enables their deployment on IoT devices but also significantly reduces the inference latency and energy consumption for IoT services. The experiment uses 9 Raspberry Pi-4B as the IoT devices, and results show that SNN-IoT may reduce the average latency and energy consumption by about 60.7% and 49.9%, respectively, while maintaining the inference accuracy.
Xin Du 0002, Wentao Tong, Linshan Jiang, Di Yu 0001, Zhiliang Wu, Qiang Duan 0002, Shuiguang Deng
IEEE Trans. Serv. Comput.3
2024 High Performance Computing Framework for Variable Selection on Genome-wide Association Studies
abstract
Variable selection for genome-wide association studies (GWAS) has been a major research focus for decades. With the exponential growth of biological and biomedical data in the era of big data, scientists are confronted with the challenge of extracting meaningful information from vast datasets while managing the inherent heterogeneity in bioinformatics. To date, there are no highly effective tools that support high-dimensional datasets and achieve robust variable selection performance, all while accounting for the non-i.i.d. features and structured relatedness among explanatory and response variables.To address these challenges, we introduce the first high-performance computing framework for variable selection in GWAS. Our framework integrates various state-of-the-art methods, allowing researchers to easily combine different techniques and fully explore their potential. Additionally, our approach employs novel optimization strategies to solve the problem efficiently, even for high-dimensional data with sparse characteristics. By processing the data holistically, the framework delivers comprehensive analysis and accurate linkage mapping associations. Designed for ease of use, the framework is implemented in Python and offers seamless deployment, making it accessible to a wide range of researchers.
Xiang Liu 0017, Jing Diao, Mengyao Zheng, Jihe Li, Dehui Wei, Qipeng Xie, Xia Li 0005, Linshan Jiang
BIBM9
2024 Novel Truncated-rank Graph-structured and Tree-guided Sparse Linear Mixed Models for Variable Selection on Genome-wide Association Studies
abstract
Variable selection for genome-wide association studies is a key focus for bioinformatics researchers in high-performance computing. The rapid growth of biological and biomedical data demands has led to high-dimensional, heterogeneous datasets characterized by non-i.i.d. properties and numerous response variables, often resulting in false negatives or positives in recovered results. Traditional methods, when nal̈ively applied, yield suboptimal performance due to confounding factors. To account for the complex interdependencies in heterogeneous data and enhance the practical outcomes of genome-wide association studies, we introduce two methods, TGsLMM and TTsLMM, which balance effects between response and explanatory variables for subpopulation inference. Our unified framework performs sparse variable selection using graph-structured or tree-guided structures in a low-rank linear mixed model. Additionally, we extend our approach to high-dimensional datasets and adaptively select the covariance structure for genomic data. Extensive experiments on synthetic and three real-world datasets emphasize the robustness and effectiveness of our proposed methods, achieving the highest ROC area compared to baselines and superior results for future potential.
Xiang Liu 0017, Jing Diao, Mengyao Zheng, Jihe Li, Yongyi Xie, Kang Lai, Xiao Geng, Yijun Song, Linshan Jiang
BIBM10
2024 FusionFrame: A Fusion Dataflow Scheduling Framework for DNN Accelerators via Analytical Modeling
Liutao Zheng, Huiying Lan, Xiang Liu 0017, Linshan Jiang, Xuehai Zhou
ICA3PP (6)4
2024 LiteCrypt: Enhancing IoMT Security with Optimized HE and Lightweight Dual-Authorization
abstract
The integration of 5G/6G networks with intelligent healthcare systems has enabled early disease detection through patient data monitoring. However, the Internet of Medical Things (IoMT) and remote healthcare services introduce significant privacy and security risks. In this paper, we propose LiteCrypt, which addresses these challenges by introducing an optimized Homomorphic Convolutional Neural Networks (HCNN) structure for secure inference and a lightweight Threshold Signature Scheme (TSS) based dual-authorization mechanism. To enhance the practicality of Homomorphic Encryption (HE)-based secure inference in telemedicine applications, LiteCrypt presents an optimized HCNN framework that ensures efficient and adaptable operations across multiple datasets. A high-performance GPU-accelerated HE engine is developed to address the computational demands of HE operations, enabling real-time processing of encrypted patient data. Besides, LiteCrypt introduces a novel TSS-based dual-authorization protocol, requiring consent from both the patient and the hospital to access patient data, thereby mitigating unauthorized access risks. The system adapts to a flexible 2-out-of-3 authorization scheme for emergencies, ensuring timely data retrieval while maintaining security. To overcome the initial challenge of prolonged computation time due to compute-intensive operations, In LiteCrypt, we utilized the lightweight TSS protocol, based on Oblivious Transfer (OT), which is designed for resource-constrained IoMT devices, reducing computation time from 11.9 to 0.11 seconds. Empirical validation demonstrates LiteCrypt’s superior performance, achieving a 233-fold increase in processing speed, a $96 \%$ reduction in encrypted message size, and a 28-fold speed increase using GPUs.
Qipeng Xie, Weizheng Wang 0001, Yongzhi Huang 0002, Mengyao Zheng, Shuai Shang, Linshan Jiang, Salabat Khan, Kaishun Wu
ICPADS6
2024 Practical Hybrid Gradient Compression for Federated Learning Systems
Sixu Hu, Linshan Jiang, Bingsheng He
IJCAI2
2024 EC-SNN: Splitting Deep Spiking Neural Networks for Edge Devices
Di Yu 0001, Xin Du 0002, Linshan Jiang, Wentao Tong, Shuiguang Deng
IJCAI3
2024 Poster Abstract: Threshold Cryptography-based Authentication Protocol for Remote Healthcare
abstract
With the advancement of the Internet of Medical Things (IoMT) and cryptographic technologies, remote healthcare services have become more widespread, presenting new challenges for patient privacy and data security. Conventional security mechanisms, such as centralized authentication and key distribution systems, are susceptible to single points of failure and significant management burdens, potentially leading to compromised authentication centers and internal security threats. In response, this study presents a threshold signature algorithm, it uses Distributed Key Generation (DKG) that distributes private keys without the need for a trusted key distributor, requiring the cooperative signature of at least two nodes for authentication. This approach not only circumvents the risk of single points of failure but also enhances the system’s robustness and efficiency. The experimental results validate its prospective utility in safeguarding remote healthcare data.
Qipeng Xie, Linshan Jiang, Siyang Jiang, Salabat Khan, Weizheng Wang 0001, Kaishun Wu
IPSN3
2024 FedLPA: One-shot Federated Learning with Layer-Wise Posterior Aggregation
abstract
Efficiently aggregating trained neural networks from local clients into a global model on a server is a widely researched topic in federated learning. Recently, motivated by diminishing privacy concerns, mitigating potential attacks, and reducing communication overhead, one-shot federated learning (i.e., limiting client-server communication into a single round) has gained popularity among researchers. However, the one-shot aggregation performances are sensitively affected by the non-identical training data distribution, which exhibits high statistical heterogeneity in some real-world scenarios. To address this issue, we propose a novel one-shot aggregation method with layer-wise posterior aggregation, named FedLPA. FedLPA aggregates local models to obtain a more accurate global model without requiring extra auxiliary datasets or exposing any private label information, e.g., label distributions. To effectively capture the statistics maintained in the biased local datasets in the practical non-IID scenario, we efficiently infer the posteriors of each layer in each local model using layer-wise Laplace approximation and aggregate them to train the global parameters. Extensive experimental results demonstrate that FedLPA significantly improves learning performance over state-of-the-art methods across several metrics.
Xiang Liu 0017, Liangxi Liu, Feiyang Ye 0004, Yunheng Shen, Xia Li 0005, Linshan Jiang, Jialin Li 0001
NeurIPS6
2024 Efficiency Optimization Techniques in Privacy-Preserving Federated Learning With Homomorphic Encryption: A Brief Survey
abstract
Federated learning (FL) offers distributed machine learning on edge devices. However, the FL model raises privacy concerns. Various techniques, such as homomorphic encryption (HE), differential privacy, and multiparty cooperation, are used to address the privacy issues of the FL model. Among them, HE ensures greater security and privacy since end-to-end encryption maintains data privacy throughout the computation process. Compared with other privacy-preserving techniques, HE does not require the establishment of a trusted environment or protocol among multiple parties and does not involve any artificial noise that can impair system performance. Unfortunately, it suffers from efficiency overhead when applied to privacy-preserving FL (PPFL). Some existing surveys on PPFL discuss the generic construction and organization of PPFL from the perspective of practical HE deployment in PPFL. However, none of them covers the efficiency optimization of HE when applied to PPFL. This article conducts a comprehensive review of the efficiency optimization of HE when applied to PPFL. First, we review general optimization strategies and discuss their limitations when applied directly to HE-based PPFL. Second, an overview of algorithmic, hardware, and hybrid optimizations is provided, along with a discussion of their adaptation. Additionally, we provide a detailed taxonomy of optimizations. Finally, we suggest future HE-based PPFL research directions.
Qipeng Xie, Siyang Jiang, Linshan Jiang, Yongzhi Huang 0002, Salabat Khan, Wangchen Dai, Zhe Liu 0001, Kaishun Wu
IEEE Internet Things J.3
2024 OFL-W3: A One-shot Federated Learning System on Web 3.0
abstract
Federated Learning (FL) addresses the challenges posed by data silos, which arise from privacy, security regulations, and ownership concerns. Despite these barriers, FL enables these isolated data repositories to participate in collaborative learning without compromising privacy or security. Concurrently, the advancement of blockchain technology and decentralized applications (DApps) within Web 3.0 heralds a new era of transformative possibilities in web development. As such, incorporating FL into Web 3.0 paves the path for overcoming the limitations of data silos through collaborative learning. However, given the transaction speed constraints of core blockchains such as Ethereum (ETH) and the latency in smart contracts, employing one-shot FL, which minimizes client-server interactions in traditional FL to a single exchange, is considered more apt for Web 3.0 environments. This paper presents a practical one-shot FL system for Web 3.0, termed OFL-W3. OFL-W3 capitalizes on blockchain technology by utilizing smart contracts for managing transactions. Meanwhile, OFL-W3 utilizes the Inter-Planetary File System (IPFS) coupled with Flask communication, to facilitate backend server operations to use existing one-shot FL algorithms. With the integration of the incentive mechanism, OFL-W3 showcases an effective implementation of one-shot FL on Web 3.0, offering valuable insights and future directions for AI combined with Web 3.0 studies.
Linshan Jiang, Moming Duan, Bingsheng He, Peishen Yan, Yang Hua 0001, Tao Song 0003
Proc. VLDB Endow.1
2023 Poster Abstract: CNN-guardian: Secure Neural Network Inference Acceleration on Edge GPU
abstract
The rapid development of AI applications powered by deep learning in edge devices boosts the opportunity for real-time health monitoring. To address the potential privacy concern in the inference phase, homomorphic encryption (HE) is an alternative solution that encrypts inference data without exposing raw data and has several distinct advantages, (i.e., single-round communication, lightweight bandwidth consumption, and non-interactive computation). However, the computational overhead on the current HE-based privacy-preserving inference necessitates a substantial amount of time, which is not feasible for some real-time applications on edge devices. To address this issue, we propose CNN-guardian, a unified and compact neural network structure for real-time inference in HE-based inference on edge GPU. CNN-guardian designs a HE-friendly neural network and GPU engine that optimizes HE operations to accelerate the inference in the HE domain.
Qipeng Xie, Hao Yang 0062, Linshan Jiang, Siyang Jiang, Shiyu Shen 0001, Salabat Khan, Zhe Liu 0001, Kaishun Wu
SenSys3
2023 Correction to "Privacy-Preserving Blockchain-Based Federated Learning for IoT Devices"
abstract
In[1], on page 1824,Fig. 3should be as follows:
Yang Zhao 0017, Jun Zhao 0007, Linshan Jiang, Rui Tan 0001, Dusit Niyato, Zengxiang Li, Lingjuan Lyu
IEEE Internet Things J.3
2022 PriMask: Cascadable and Collusion-Resilient Data Masking for Mobile Cloud Inference
abstract
Mobile cloud offloading is indispensable for inference tasks based on large-scale deep models. However, transmitting privacy-rich inference data to the cloud incurs concerns. This paper presents the design of a system called PriMask, in which the mobile device uses a secret small-scale neural network called MaskNet to mask the data before transmission. PriMask significantly weakens the cloud's capability to recover the data or extract certain private attributes. The MaskNet is cascadable in that the mobile can opt in to or out of its use seamlessly without any modifications to the cloud's inference service. Moreover, the mobiles use different MaskNets, such that the collusion between the cloud and some mobiles does not weaken the protection for other mobiles. We devise a split adversarial learning method to train a neural network that generates a new MaskNet quickly (within two seconds) at run time. We apply PriMask to three mobile sensing applications with diverse modalities and complexities, i.e., human activity recognition, urban environment crowdsensing, and driver behavior recognition. Results show PriMask's effectiveness in all the three applications.
Linshan Jiang, Qun Song 0001, Rui Tan 0001, Mo Li 0001
SenSys1
2022 Attack-aware Synchronization-free Data Timestamping in LoRaWAN
abstract
Low-power wide-area network technologies such as long-range wide-area network (LoRaWAN) are promising for collecting low-rate monitoring data from geographically distributed sensors, in which timestamping the sensor data is a critical system function. This article considers a synchronization-free approach to timestamping LoRaWAN uplink data based on signal arrival time at the gateway, which well matches LoRaWAN’s one-hop star topology and releases bandwidth from transmitting timestamps and synchronizing end devices’ clocks at all times. However, we show that this approach is susceptible to a frame delay attack consisting of malicious frame collision and delayed replay. Real experiments show that the attack can affect the end devices in large areas up to about 50,000, m 2 . In a broader sense, the attack threatens any system functions requiring timely deliveries of LoRaWAN frames. To address this threat, we propose a LoRaTS gateway design that integrates a commodity LoRaWAN gateway and a low-power software-defined radio receiver to track the inherent frequency biases of the end devices. Based on an analytic model of LoRa’s chirp spread spectrum modulation, we develop signal processing algorithms to estimate the frequency biases with high accuracy beyond that achieved by LoRa’s default demodulation. The accurate frequency bias tracking capability enables the detection of the attack that introduces additional frequency biases. We also investigate and implement a more crafty attack that uses advanced radio apparatuses to eliminate the frequency biases. To address this crafty attack, we propose a pseudorandom interval hopping scheme to enhance our frequency bias tracking approach. Extensive experiments show the effectiveness of our approach in deployments with real affecting factors such as temperature variations.
Chaojie Gu, Linshan Jiang, Rui Tan 0001, Mo Li 0001, Jun Huang 0001
ACM Trans. Sens. Networks2
2021 An Electromagnetic Covert Channel based on Neural Network Architecture
abstract
Outsourcing the design of deep neural networks may incur cybersecurity threats from the hostile designers. This paper studies a new covert channel attack that leaks the inference results over the air through a hostile design of the neural network architecture and the computing device's electromagnetic radiation when executing the neural network. Specifically, the hostile neural network consists of a series of binary models that correspond to all classes and are executed sequentially. The execution terminates once any binary model given the input is positive about its responsible class. We describe an approach to generate such binary models by pruning a benign neural network that is trained using the standard method to deal with all the classes. Compared with the benign neural network, the hostile one has similar memory usage and negligible classification accuracy drop, but distinct inference times for the samples of different classes. As a result, the hostile neural network's classification result can be eavesdropped by measuring the duration of the electromagnetic radiation emanated from the computing device. As neural networks are stored and transmitted as data files, this covert channel attack is more stealthy to the anti-malware than other code-based attacks. We implement the described attack on two edge computing devices that execute the hostile neural network on CPU or GPU. Evaluation shows 100% empirical accuracy in eavesdropping the inference results.
Chaojie Gu, Rui Tan 0001, Linshan Jiang
ICPADS4
2021 Privacy-Preserving Blockchain-Based Federated Learning for IoT Devices
abstract
Home appliance manufacturers strive to obtain feedback from users to improve their products and services to build a smart home system. To help manufacturers develop a smart home system, we design a federated learning (FL) system leveraging a reputation mechanism to assist home appliance manufacturers to train a machine learning model based on customers’ data. Then, manufacturers can predict customers’ requirements and consumption behaviors in the future. The working flow of the system includes two stages: in the first stage, customers train the initial model provided by the manufacturer using both the mobile phone and the mobile-edge computing (MEC) server. Customers collect data from various home appliances using phones, and then they download and train the initial model with their local data. After deriving local models, customers sign on their models and send them to the blockchain. In case customers or manufacturers are malicious, we use the blockchain to replace the centralized aggregator in the traditional FL system. Since records on the blockchain are untampered, malicious customers or manufacturers’ activities are traceable. In the second stage, manufacturers select customers or organizations as miners for calculating the averaged model using received models from customers. By the end of the crowdsourcing task, one of the miners, who is selected as the temporary leader, uploads the model to the blockchain. To protect customers’ privacy and improve the test accuracy, we enforce differential privacy (DP) on the extracted features and propose a new normalization technique. We experimentally demonstrate that our normalization technique outperforms batch normalization when features are under DP protection. In addition, to attract more customers to participate in the crowdsourcing FL task, we design an incentive mechanism to award participants.
Yang Zhao 0017, Jun Zhao 0007, Linshan Jiang, Rui Tan 0001, Dusit Niyato, Zengxiang Li, Lingjuan Lyu
IEEE Internet Things J.3
2021 On Lightweight Privacy-preserving Collaborative Learning for Internet of Things by Independent Random Projections
abstract
The Internet of Things (IoT) will be a main data generation infrastructure for achieving better system intelligence. This article considers the design and implementation of a practical privacy-preserving collaborative learning scheme, in which a curious learning coordinator trains a better machine learning model based on the data samples contributed by a number of IoT objects, while the confidentiality of the raw forms of the training data is protected against the coordinator. Existing distributed machine learning and data encryption approaches incur significant computation and communication overhead, rendering them ill-suited for resource-constrained IoT objects. We study an approach that applies independent random projection at each IoT object to obfuscate data and trains a deep neural network at the coordinator based on the projected data from the IoT objects. This approach introduces light computation overhead to the IoT objects and moves most workload to the coordinator that can have sufficient computing resources. Although the independent projections performed by the IoT objects address the potential collusion between the curious coordinator and some compromised IoT objects, they significantly increase the complexity of the projected data. In this article, we leverage the superior learning capability of deep learning in capturing sophisticated patterns to maintain good learning performance. Extensive comparative evaluation shows that this approach outperforms other lightweight approaches that apply additive noisification for differential privacy and/or support vector machines for learning in the applications with light to moderate data pattern complexities.
Linshan Jiang, Rui Tan 0001, Xin Lou 0005, Guosheng Lin
ACM Trans. Internet Things1
2020 Attack-Aware Data Timestamping in Low-Power Synchronization-Free LoRaWAN
abstract
Low-power wide-area network technologies such as LoRaWAN are promising for collecting low-rate monitoring data from geographically distributed sensors, in which timestamping the sensor data is a critical system function. This paper considers a synchronization-free approach to timestamping LoRaWAN uplink data based on signal arrival time at the gateway, which well matches LoRaWAN’s one-hop star topology and releases bandwidth from transmitting timestamps and synchronizing end devices’ clocks at all times. However, we show that this approach is susceptible to a frame delay attack consisting of malicious frame collision and delayed replay. Real experiments show that the attack can affect the end devices in large areas up to about 50, 000 m2. In a broader sense, the attack threatens any system functions requiring timely deliveries of LoRaWAN frames. To address this threat, we propose a LoRaTS gateway design that integrates a commodity LoRaWAN gateway and a low-power software-defined radio receiver to track the inherent frequency biases of the end devices. Based on an analytic model of LoRa’s chirp spread spectrum modulation, we develop signal processing algorithms to estimate the frequency biases with high accuracy beyond that achieved by LoRa’s default demodulation. The accurate frequency bias tracking capability enables the detection of the attack that introduces additional frequency biases. Extensive experiments show the effectiveness of our approach.
Chaojie Gu, Linshan Jiang, Rui Tan 0001, Mo Li 0001, Jun Huang 0001
ICDCS2
2020 Lightweight and Unobtrusive Data Obfuscation at IoT Edge for Remote Inference
abstract
Executing deep neural networks for inference on the server-class or cloud backend based on the data generated at the edge of the Internet of Things is desirable due primarily to the limited compute power of the edge devices and the need to protect the confidentiality of the inference neural networks. However, such a remote inference scheme incurs concerns regarding the privacy of the inference data transmitted by the edge devices to the curious backend. This article presents a lightweight and unobtrusive approach to obfuscate the inference data at the edge devices. It is lightweight in that the edge device only needs to execute a small-scale neural network; it is unobtrusive in that the edge device does not need to indicate whether obfuscation is applied. Extensive evaluation by three case studies of free-spoken digit recognition, handwritten digit recognition, and American sign language recognition shows that our approach effectively protects the confidentiality of the raw forms of the inference data while effectively preserving backend's inference accuracy.
Dixing Xu, Mengyao Zheng, Linshan Jiang, Chaojie Gu, Rui Tan 0001, Peng Cheng 0001
IEEE Internet Things J.3
2020 Resilience Bounds of Network Clock Synchronization with Fault Correction
abstract
Naturally occurring disturbances and malicious attacks can lead to faults in synchronizing the clocks of two network nodes. In this article, we investigate the fundamental resilience bounds of network clock synchronization for a system of N nodes against the peer-to-peer synchronization faults. Our analysis is based on practical synchronization algorithms with time complexity down to O ( N 3 ) that attempt to correct the faults by checking the consistency among the following three types of data: (1) the estimated faults, (2) the estimated clock offsets among the nodes, and (3) the measured clock offsets from the potentially faulty peer-to-peer synchronization sessions. Our analysis gives the following three major results. First, the maximum number of faults that can be corrected by the algorithms has a tight bound of ⌊ N /2 ⌋ − 1 when every node pair performs a synchronization session. Second, by converting the fault resilience problem to a graph-theoretic edge connectivity problem and applying Menger’s theorem, we develop an algorithm to compute the tight bound when not every node pair performs a synchronization session. Third, the number of synchronization sessions to achieve the capability of correcting any K faults has a lower bound of ⌈ N (2 K +1) / 2 ⌉ ; we also develop an algorithm to schedule the synchronization sessions to approach the lower bound. The above results provide basic understanding and useful guidelines to the design of resilient clock synchronization systems. For instance, our results suggest that, the four-node network achieves the highest degree of resilience that is defined as the ratio of the maximum number of correctable faults to the number of synchronization sessions. Therefore, by organizing a large-scale clock synchronization system into a hierarchy of multiple tiers with each consisting of four-node synchronization groups, we can achieve satisfactory and understood resilience against faults with reduced synchronization sessions.
Linshan Jiang, Rui Tan 0001, Arvind Easwaran
ACM Trans. Sens. Networks1
2019 LoRa-Based Localization: Opportunities and Challenges
Chaojie Gu, Linshan Jiang, Rui Tan 0001
EWSN2
2019 Differentially Private Collaborative Learning for the IoT Edge
Linshan Jiang, Xin Lou 0005, Rui Tan 0001, Jun Zhao 0007
EWSN1
2018 An Insurance-based Incentive Mechanism for Mobile Crowdsourcing to Improve System Security
abstract
In a crowdsourcing system, security is a critical issue which affects the participation willingness of users. To motivate users' participation, most of existing work provide additional reward to compensate their loss due to security issues. However, more efficient way is to motivate the users to arm with higher security capability, to reduce the infection probability from the attackers and malicious software. In this paper, we propose an insurance-based incentive framework to motivate the users to upgrade to a higher security level. The framework can be formed as a Stackelberg game, where crowdsourcing platform is the leader and the users are followers. Through backward induction, we found that a Nash Equilibrium exists in the Stackelberg game. Simulation result shows that the proposed mechanism can enhance both social welfare, platform utility and users' utility in the crowdsourcing system.
Linshan Jiang, Jin Zhang 0001
CSCWD2
2018 Resilience Bounds of Sensing-Based Network Clock Synchronization
abstract
Recent studies exploited external periodic synchronous signals to synchronize a pair of network nodes to address a threat of delaying the communications between the nodes. However, the sensing-based synchronization may yield faults due to nonmalicious signal and sensor noises. This paper considers a system of N nodes that will fuse their peer-to-peer synchronization results to correct the faults. Our analysis gives the lower bound of the number of faults that the system can tolerate when N is up to 12. If the number of faults is no greater than the lower bound, the faults can be identified and corrected. We also prove that the system cannot tolerate more than N - 2 faults. Our results can guide the design of resilient sensing-based clock synchronization systems.
Rui Tan 0001, Linshan Jiang, Arvind Easwaran, Jothi Prasanna Shanmuga Sundaram
ICPADS2
2018 An Insurance-Based Framework Against Security Threat in Mobile Crowdsourcing Systems
abstract
Mobile crowdsourcing is a popular computing paradigm that enables smart devices to measure and collect various sensing data. When the users participate in the sensing platform, they may face various attacks and fall in a non-secure environment. Under the malicious attack, the sensing data which should be transmitted to the platform may suffer from data loss, which reduces the users' utility and platform's utility. The data loss can also lead to the reduction of the users' participatory motivations because they are not able to earn expected money due to data loss. Therefore, the total sensing quality and the social welfare will be affected eventually. To solve this problem, in this paper, we propose a novel insurance-based framework to compensate the data loss due to security threat in mobile crowdsourcing systems. This framework can motivate the users with high-security levels to participate in the crowdsourcing system, thus improves the platform's utility. We formulate our framework as a Stackelberg game, where the platform is a leader and the users are followers. The theoretical analysis shows that the Nash Equilibrium exists in our framework and it can maximize the platform's revenue while considering the users' participatory willingness. Simulation results show that our framework achieves more participators, more platform's utility and social welfare, compared with existing mechanism.
Linshan Jiang, Jin Zhang 0001
ICPADS2
2016 Many-to-many matching for combinatorial spectrum trading
abstract
Dynamic spectrum access (DAS) is an efficient way to redistribute spare channels among users. Conventionally, dynamic spectrum access is conducted through (double) spectrum auction, where a third-party auctioneer collects bids from buyers and sellers, and determines the spectrum allocation. Rather than placing bids only on individual channels, combinatorial spectrum auction allows buyers to express their valuations for different combinations of channels. However, auction mechanisms are generally vulnerable to the collusion between the auctioneer and buyers or sellers. Furthermore, to find the optimal allocation in combinatorial auction is usually NP-hard. In this paper, we propose to leverage a many-to-many matching framework to realize combinatorial spectrum trading. Unlike traditional many-to-many matching problem, spectrum matching is more challenging, because spectrum allocation is interference-limited rather than quota-limited. To deal with this problem, we propose a novel matching algorithm, which takes buyers' interference relationship into consideration. We theoretically prove that the matching result is individual rational, strong pairwise stable and is a subgame-perfect Nash equilibrium of the corresponding spectrum bargaining game. Simulation results show that the proposed algorithm can converge to a stable matching within a few iterations.
Linshan Jiang, Haofan Cai, Yanjiao Chen, Jin Zhang 0001, Baochun Li
ICC1
2016 Spectrum Matching
abstract
Dynamic spectrum access (DSA) redistributes spectrum from service providers with spare channels to those in need for them. Existing works on such spectrum exchange mainly focus on double auctions, where an auctioneer centrally enforces a certain spectrum allocation policy. In this paper, we take a different and new perspective, proposing to use matching as an alternative tool to realize DSA in a distributed way for a free market, which consists of only buyers and sellers, but no trustworthy third-party authority. Compared with conventional many-to-one matching problems, the spectrum matching problem is distinctively challenging due to the interference bound between buyers: the same channel can be reused by an unlimited number of non-interfering buyers, but must be exclusively occupied by only one of interfering buyers. In this paper, we firstly formulate the spectrum matching problem as a many-to-one matching with peer effects, i.e., a buyer's utility is affected by other buyers who are matched to the same seller. We then present a two-stage distributed algorithm that converges to an interference-free and Nash-stable matching result. Simulations show that the proposed distributed matching algorithm can achieve 90% of the social welfare from the optimal matching result.
Yanjiao Chen, Linshan Jiang, Haofan Cai, Jin Zhang 0001, Baochun Li
ICDCS2