Xin Du 0002

dblp:18/4833-2 · DBLP profile ↗
← Back
35ranked-venue papers
5as first author
33since 2021 · last 2026
0000-0003-1108-0248ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TPipe: Efficient Spiking Transformer Training with Time Parallelism and Asynchronous Pipeline
Yubing Bao, Zhihui Lu 0002, Qiang Duan 0002, Changze Lv, Xin Du 0002, Zeyi Deng, Jingqi Feng, Sen Liu 0002, Yang Chen 0001, Xin Wang 0002
INFOCOM5
2026 MASI: Memory-Adaptive Inference Framework for Spiking Neural Networks on Edge Devices
abstract
The rapid development of the Internet of Things (IoT) applications necessitates resource-efficient computing paradigms that can unify heterogeneous sensing modalities. Spiking Neural Networks (SNNs) meet this need with their event-driven and energy-efficient processing nature. However, deploying SNNs on mobile and embedded platforms is hindered by strict and fluctuating memory budgets. While prior work explores lightweight model design and system-level memory management, these methods either sacrifice accuracy or incur high runtime overhead due to timestep-dependent dynamics. To tackle these challenges, we propose a memory-adaptive framework MASI that enables efficient on-device SNN inference by combining (1) a fine-grained memory-adaptive layer slicing strategy, (2) a timestep-agnostic scheduler that maximizes memory utilization with minimal fragmentation, and (3) a timestep-aware early-exit mechanism that reduces redundant calculations. Evaluated on diverse workloads and edge devices, MASI can dynamically adapt to runtime memory availability, approximately reducing memory usage by 20.67% and inference latency by 58.53% on average with negligible accuracy loss compared to other feasible on-device implementations under memory constraints.
Di Yu 0001, Helin Zheng, Changze Lv, Xin Du 0002, Linshan Jiang, Xiang Liu 0017, Gang Pan 0001, Shuiguang Deng
WWW4
2026 DepAsync: An Asynchronous SNN Accelerator Based on Core-Dependency
abstract
Spiking Neural Networks (SNNs) are widely used in brain-inspired computing and neuroscience research. Several many-core accelerators have been built to improve the running speed and energy efficiency of SNNs. However, current accelerators generally need explicit synchronization among all cores after each timestep of SNNs, which poses a challenge to overall efficiency. This paper proposes DepAsync, an asynchronous architecture that eliminates inter-core synchronization, facilitating fast and energy-efficient SNN inference with commendable scalability. The main idea is to exploit the dependency of neuromorphic cores predetermined at compile time. We design a DepAsync scheduler for each core to trace the running state of its dependencies and control the core to safely forward to the next timestep without waiting for other cores to complete their tasks. This approach prevents the necessity for global synchronization, allowing DepAsync to minimize core waiting time facing inherent core and time imbalance in SNN workloads. The comprehensive evaluations using five SNN workloads show that DepAsync achieves 2.47x speedup and 1.55x energy efficiency compared to the state-of-the-art synchronization architectures.
Zhuo Chen 0044, De Ma, Xiaofei Jin, Qinghui Xing, Ouwen Jin, Xin Du 0002, Shuibing He, Gang Pan 0001
IEEE Trans. Computers6
2025 Efficient Joint Communication and Computation Placement for Large-scale SNN Simulation on Supercomputers
abstract
Spiking Neural Network (SNN) simulation involves emulating the activation and firing of spiking neurons on hardware platforms. This is a highly time-sensitive task, requiring the simulation of billions of neurons and their intercommunication within a few milliseconds. Each neuron performs a complex, interdependent multi-stage communication and computation task. We consider the task placement of SNN on supercomputers to accelerate SNN simulation. Existing task placement methods for SNN simulations have two major limitations. First, they lack the capability to handle large-scale SNNs with billions of neurons. Second, they focus primarily on optimizing communication delay, while neglecting multi-stage computation delays in SNN simulations. In this paper, we formalize the SNN Joint Multi-stage Communication and Computation Placement (SJCCP) problem. We demonstrate that SJCCP can be solved using an approximation algorithm with an approximation ratio of $O\left( {{k^2}\sqrt {\log n\log k} } \right)$, where n is the number of voxels in the SNN and k is the number of GPUs. To further reduce the time complexity of solving SJCCP in practice, we propose a novel efficient framework, FastSJP, tailored for large-scale SNN placement. Then we apply the FastSJP framework to a human brain simulation that runs a large-scale SNN model derived from authentic biological data on a supercomputer equipped with 1024 GPUs. Experimental results verify that our framework notably reduces time overhead, ranging from 17.31% to 28.45%, compared to state-of-the-art methods. Leveraging the computational power of the supercomputer, FastSJP maximizes the problem size and processing performance, significantly advancing the development of brain-inspired intelligence.
Yubing Bao, Zhihui Lu 0002, Xin Du 0002, Qiang Duan 0002, Jirui Yang, Jin Zhao 0001, Geyong Min, Yang Chen 0001, Shijing Hu 0001, Xin Wang 0002
ICDCS3
2025 MetricEmbedding: Accelerate Metric Nearness by Tropical Inner Product
abstract
The Metric Nearness Problem involves restoring a non-metric matrix to its closest metric-compliant form, addressing issues such as noise, missing values, and data inconsistencies. Ensuring metric properties, particularly the $O(N^3)$ triangle inequality constraints, presents significant computational challenges, especially in large-scale scenarios where traditional methods suffer from high time and space complexity. We propose a novel solution based on the tropical inner product (max-plus operation), which we prove satisfies the triangle inequality for non-negative real matrices. By transforming the problem into a continuous optimization task, our method directly minimizes the distance to the target matrix. This approach not only restores metric properties but also generates metric-preserving embeddings, enabling real-time updates and reducing computational and storage overhead for downstream tasks. Experimental results demonstrate that our method achieves up to 60$\times$ speed improvements over state-of-the-art approaches, and efficiently scales from $1e4 \times 1e4$ to $1e5 \times 1e5$ matrices with significantly lower memory usage.
Muyang Cao, Jiajun Yu, Xin Du 0002, Gang Pan 0001, Wei Wang 0011
ICML3
2025 Dendritic Localized Learning: Toward Biologically Plausible Algorithm
abstract
Backpropagation is the foundational algorithm for training neural networks and a key driver of deep learning’s success. However, its biological plausibility has been challenged due to three primary limitations: weight symmetry, reliance on global error signals, and the dual-phase nature of training, as highlighted by the existing literature. Although various alternative learning approaches have been proposed to address these issues, most either fail to satisfy all three criteria simultaneously or yield suboptimal results. Inspired by the dynamics and plasticity of pyramidal neurons, we propose Dendritic Localized Learning (DLL), a novel learning algorithm designed to overcome these challenges. Extensive empirical experiments demonstrate that DLL satisfies all three criteria of biological plausibility while achieving state-of-the-art performance among algorithms that meet these requirements. Furthermore, DLL exhibits strong generalization across a range of architectures, including MLPs, CNNs, and RNNs. These results, benchmarked against existing biologically plausible learning algorithms, offer valuable empirical insights for future research. We hope this study can inspire the development of new biologically plausible algorithms for training multilayer networks and advancing progress in both neuroscience and machine learning. Our code is available at https://github.com/Lvchangze/Dendritic-Localized-Learning.
Changze Lv, Zhenghua Wang, Zhibo Xu, Di Yu 0001, Xin Du 0002, Xiaoqing Zheng, Xuanjing Huang 0001
ICML8
2025 BMapper: A Scalable and Efficient Framework for Brain Simulations Acceleration on Supercomputers
abstract
Brain simulation is an inherently highly parallel and time-sensitive task, requiring the simulation of billions of neurons and their interactions within just a few milliseconds. With the growing availability of brain data from biological research, more realistic and detailed simulations are becoming feasible. However, this also poses unprecedented challenges for parallel computing due to the extreme sparsity and heterogeneity of the emerging workloads. Efficient deployment of such workloads on modern HPC systems is critical to overcoming these challenges. We propose BMapper, a deployment framework that enables efficient parallel execution of brain simulations on supercomputers. BMapper comprises three synergistic components: BPartitioning, which introduces a novel multi-dimensional hybrid partitioning strategy to balance workloads across GPUs and reduce inter-GPU spike traffic; BPlacement, which applies deterministic spectral partitioning to minimize inter-server communication; and BRelaying, which identifies lightly loaded GPUs to assist the top-k heavily loaded ones by relaying spike traffic. These components work together to balance loads and minimize communication overhead, enabling high-speed simulation of large-scale brain models. BMapper has been deployed to simulate up to 10 billion neurons on a 1000-GPU supercomputer, achieving 25.15%–47.48% faster execution than state-of-the-art methods.
Yubing Bao, Zhihui Lu 0002, Qiang Duan 0002, Xin Du 0002, Yandan Tan, Yang Chen 0001, Yang Xu 0010
ICPP4
2025 Exploiting Label Skewness for Spiking Neural Networks in Federated Learning
abstract
The energy efficiency of deep spiking neural networks (SNNs) aligns with the constraints of resource-limited edge devices, positioning SNNs as a promising foundation for intelligent applications leveraging the extensive data collected by these devices. To safeguard data privacy, federated learning (FL) facilitates collaborative SNN-based model training by leveraging data distributed across edge devices without transmitting local data to a central server. However, existing FL approaches encounter challenges in handling label-skewed data across devices, inducing drift in the local SNN model and consequently impairing the performance of the global SNN model. To tackle these problems, we propose a novel framework called FedLEC, which incorporates intra-client label weight calibration to balance the learning intensity across local labels and inter-client knowledge distillation to mitigate local SNN model bias caused by label absence. Extensive experiments with three different structured SNNs across five datasets (i.e., three non-neuromorphic and two neuromorphic datasets) demonstrate the efficiency of FedLEC. Compared to seven state-of-the-art FL algorithms, FedLEC achieves an average accuracy improvement of approximately 11.59% for the global SNN model under various label skew distribution settings.
Di Yu 0001, Xin Du 0002, Linshan Jiang, Huijing Zhang, Shuiguang Deng
IJCAI2
2025 ECC-SNN: Cost-Effective Edge-Cloud Collaboration for Spiking Neural Networks
abstract
Most edge-cloud collaboration frameworks rely on the substantial computational and storage capabilities of cloud-based artificial neural networks (ANNs). However, this reliance results in significant communication overhead between edge devices and the cloud, as well as high computational energy consumption, especially when applied to resource-constrained edge devices. To address these challenges, we propose ECC-SNN, a novel edge-cloud collaboration framework that incorporates energy-efficient spiking neural networks (SNNs) to offload more computational workload from the cloud to the edge, thereby improving cost-effectiveness and reducing reliance on the cloud. ECC-SNN employs a joint training approach that integrates ANN and SNN models, enabling edge devices to leverage knowledge from cloud models for enhanced performance while reducing energy consumption and processing latency. Furthermore, ECC-SNN features an on-device incremental learning algorithm that enables edge models to continuously adapt to dynamic environments, reducing the communication overhead and resource consumption associated with frequent cloud update requests. Extensive experimental results on four datasets demonstrate that ECC-SNN improves accuracy by 4.15%, reduces average energy consumption by 79.4%, and lowers average processing latency by 39.1%.
Di Yu 0001, Changze Lv, Xin Du 0002, Linshan Jiang, Wentao Tong, Xiaoqing Zheng, Shuiguang Deng
IJCAI3
2025 Cost-Effective On-Device Sequential Recommendation with Spiking Neural Networks
abstract
On-device sequential recommendation (SR) systems are designed to make local inferences using real-time features, thereby alleviating the communication burden on server-based recommenders when handling concurrent requests from millions of users. However, the resource constraints of edge devices, including limited memory and computational capacity, pose significant challenges to deploying efficient SR models. Inspired by the energy-efficient and sparse computing properties of deep Spiking Neural Networks (SNNs), we propose a cost-effective on-device SR model named SSR, which encodes dense embedding representations into sparse spike-wise representations and integrates novel spiking filter modules to extract temporal patterns and critical features from item sequences, optimizing computational and memory efficiency without sacrificing recommendation accuracy. Extensive experiments on real-world datasets demonstrate the superiority of SSR. Compared to other SR baselines, SSR achieves comparable recommendation performance while reducing energy consumption by an average of 59.43%. In addition, SSR significantly lowers memory usage, making it particularly well-suited for deployment on resource-constrained edge devices.
Di Yu 0001, Changze Lv, Xin Du 0002, Linshan Jiang, Qing Yin, Wentao Tong, Xiaoqing Zheng, Shuiguang Deng
IJCAI3
2025 Universal Backdoor Defense via Label Consistency in Vertical Federated Learning
abstract
Backdoor attacks in vertical federated learning (VFL) are particularly concerning as they can covertly compromise VFL decision-making, posing a severe threat to critical applications of VFL. Existing defense mechanisms typically involve either label obfuscation during training or model pruning during inference. However, the inherent limitations on the defender's access to the global model and complete training data in VFL environments fundamentally constrain the effectiveness of these conventional methods. To address these limitations, we propose the Universal Backdoor Defense (UBD) framework. UBD leverages Label Consistent Clustering (LCC) to synthesize plausible latent triggers associated with the backdoor class. This synthesized information is then utilized for mitigating backdoor threats through Linear Probing (LP), guided by a constraint on Batch Normalization (BN) statistics. Positioned within a unified VFL backdoor defense paradigm, UBD offers a generalized framework for both detection and mitigation that critically does not necessitate access to the entire model or dataset. Extensive experiments across multiple datasets rigorously demonstrate the efficacy of the UBD framework, achieving state-of-the-art performance against diverse backdoor attack types in VFL, including both dirty-label and clean-label variants.
Peng Chen 0030, Haolong Xiang, Xin Du 0002, Xiaolong Xu 0001, Xuhao Jiang, Zhihui Lu 0002, Jirui Yang, Qiang Duan 0002, Wan-Chun Dou
IJCAI3
2025 Backdoor Attack on Vertical Federated Graph Neural Network Learning
abstract
Federated Graph Neural Network (FedGNN) integrate federated learning (FL) with graph neural networks (GNNs) to enable privacy-preserving training on distributed graph data. Vertical Federated Graph Neural Network (VFGNN), a key branch of FedGNN, handles scenarios where data features and labels are distributed among participants. Despite the robust privacy-preserving design of VFGNN, we have found that it still faces the risk of backdoor attacks, even in situations where labels are inaccessible. This paper proposes BVG, a novel backdoor attack method that leverages multi-hop triggers and backdoor retention, requiring only four target-class nodes to execute effective attacks. Experimental results demonstrate that BVG achieves nearly 100% attack success rates across three commonly used datasets and three GNN models, with minimal impact on the main task accuracy. We also evaluated various defense methods, and the BVG method maintained high attack effectiveness even under existing defenses. This finding highlights the need for advanced defense mechanisms to counter sophisticated backdoor attacks in practical VFGNN applications.
Jirui Yang, Peng Chen 0030, Zhihui Lu 0002, Jianping Zeng 0002, Qiang Duan 0002, Xin Du 0002, Ruijun Deng
IJCAI6
2025 SeMi: When Imbalanced Semi-Supervised Learning Meets Mining Hard Examples
abstract
Semi-Supervised Learning (SSL) can leverage abundant unlabeled data to boost model performance. However, the class-imbalanced data distribution in real-world scenarios poses great challenges to SSL, resulting in performance degradation. Existing class-imbalanced semi-supervised learning (CISSL) methods mainly focus on rebalancing datasets but ignore the potential of using hard examples to enhance performance, making it difficult to fully harness the power of unlabeled data even with sophisticated algorithms. To address this issue, we propose a method that enhances the performance of Imbalanced Semi-Supervised Learning by Mining Hard Examples (SeMi). This method distinguishes the entropy differences among logits of hard and easy examples, thereby identifying hard examples and increasing the utility of unlabeled data, better addressing the imbalance problem in CISSL. In addition, we maintain a class-balanced memory bank with confidence decay for storing high-confidence embeddings to enhance the pseudo-labels' reliability. Although our method is simple, it is effective and seamlessly integrates with existing approaches. We perform comprehensive experiments on standard CISSL benchmarks and experimentally demonstrate that our proposed SeMi outperforms existing state-of-the-art methods on multiple benchmarks, especially in reversed scenarios, where our best result shows approximately a 54.8% improvement over the baseline methods. Our code is available at https://github.com/pywin/SeMi.
Yin Wang 0004, Hao Lu 0009, Zhen Qin 0004, Hailiang Zhao, Guanjie Cheng, Xin Du 0002, Ge Su, Li Kuang, MengChu Zhou, Shuiguang Deng
ACM Multimedia7
2025 A Data Replication Placement Strategy for the Distributed Storage System in Cloud-Edge-Terminal Orchestrated Computing Environments
abstract
Cloud-edge-terminal orchestrated computing, as an expansion of cloud computing, has sunk resources to the edge nodes and terminal equipment, which can provide high-quality services for delay-sensitive applications and reduce the cost of network communication. Due to the high volume of data generated by Internet of Things (IoT) devices and the limited storage capacities of edge nodes, a significant number of terminal devices are now being considered for utilization as storage nodes. However, because of the heterogeneous storage capacity and reliability of these hardware devices and the different data requirements of user services, the performance and storage reliability of applications deployed in cloud-edge-terminal orchestrated computing environments have become urgent problems to be solved. Especially, for a distributed storage system in these environments, it is required to ensure reliable storage of the generated data and its’ replications. In this paper, we first implement a distributed storage system and construct a data replication placement model. Then, based on the constructed model, we formulate the data replication placement problem and design a data replication placement strategy called DRPS to solve it. The DRPS covers a ranks-based replication storage node selection algorithm and a greedy load balancing algorithm, which can select appropriate hardware devices for different data requirements of services and is implemented in the data storage system to store replications and balance loads. We design extensive experiments to verify the effectiveness of DRPS. The results indicate that the proposed strategy outperforms other state-of-the-art algorithms in terms of system delay reduction by 39.9%, an increase of 43.3% in the replication numbers, a 27.5% improvement in memory utilization, and a reduction of unreliability rate by 82.0%.
Peng Chen 0030, Mengke Zheng, Xin Du 0002, Muhammad Bilal 0003, Zhihui Lu 0002, Qiang Duan 0002, Xiaolong Xu 0001
IEEE Internet Things J.3
2025 Mapping Large-Scale Spiking Neural Network on Arbitrary Meshed Neuromorphic Hardware
abstract
Neuromorphic hardware systems—designed as 2D-mesh structures with parallel neurosynaptic cores—have proven highly efficient at executing large-scale spiking neural networks (SNNs). A critical challenge, however, lies in mapping neurons efficiently to these cores. While existing approaches work well with regular, fully functional mesh structures, they falter in real-world scenarios where hardware has irregular shapes or non-functional cores caused by defects or resource fragmentation. To address these limitations, we propose a novel mapping method based on an innovative space-filling curve: the Adaptive Locality-Preserving (ALP) curve. Using a unique divide-and-conquer construction algorithm, the ALP curve ensures adaptability to meshes of any shape while maintaining crucial locality properties—essential for efficient mapping. Our method demonstrates exceptional computational efficiency, making it ideal for large-scale deployments. These distinctive characteristics enable our approach to handle complex scenarios that challenge conventional methods. Experimental results show that our method matches state-of-the-art solutions in regular-shape mapping while achieving significant improvements in irregular scenarios, reducing communication overhead by up to 57.1%.
Ouwen Jin, Qinghui Xing, Zhuo Chen 0044, Ming Zhang 0018, De Ma, Ying Li 0001, Xin Du 0002, Shuibing He, Shuiguang Deng, Gang Pan 0001
IEEE Trans. Parallel Distributed Syst.7
2025 SNN-IoT: Efficient Partitioning and Enabling of Deep Spiking Neural Networks in IoT Services
abstract
Spiking Neural Networks (SNNs), due to their inherent biological plausibility and energy-saving characteristics, naturally align with the requirements of IoT services. However, current SNNs require a multi-layer structure to achieve effective applications across various fields. The multi-layer deep SNNs with massive model parameters demand computational resources, rendering them incompatible with resource-constrained IoT devices. To address this problem, in this work, a deep SNN partitioning framework called SNN-IoT is proposed to run complex SNN models on IoT devices. The SNN-IoT first partitions a full deep SNN model into smaller sub-models, leveraging the event-driven sparsity of SNNs and channel-level firing patterns to distribute filters with lower levels of spike activity onto devices with more constrained resources. The SNN model partitioning and deployment is formulated as an optimization problem and is solved using a greedy search assignment mechanism. Furthermore, a channel-wise pruning method exploits the varying degrees of channel activity, effectively reducing each sub-model's size and computational load without compromising performance. Extensive experiments conducted on four non-neuromorphic and two neuromorphic datasets have demonstrated that the SNN-IoT framework not only efficiently partitions deep SNNs and enables their deployment on IoT devices but also significantly reduces the inference latency and energy consumption for IoT services. The experiment uses 9 Raspberry Pi-4B as the IoT devices, and results show that SNN-IoT may reduce the average latency and energy consumption by about 60.7% and 49.9%, respectively, while maintaining the inference accuracy.
Xin Du 0002, Wentao Tong, Linshan Jiang, Di Yu 0001, Zhiliang Wu, Qiang Duan 0002, Shuiguang Deng
IEEE Trans. Serv. Comput.1
2024 EC-SNN: Splitting Deep Spiking Neural Networks for Edge Devices
Di Yu 0001, Xin Du 0002, Linshan Jiang, Wentao Tong, Shuiguang Deng
IJCAI2
2024 A Hierarchical Neural Task Scheduling Algorithm in the Operating System of Neuromorphic Computers
Pan Lv, Xin Du 0002, Ouwen Jin, Shuiguang Deng
KSEM (4)3
2024 Mitigating critical nodes in brain simulations via edge removal
Yubing Bao, Xin Du 0002, Zhihui Lu 0002, Jirui Yang, Shih-Chia Huang, Jianfeng Feng, Qibao Zheng
Comput. Networks2
2024 Universal adversarial backdoor attacks to fool vertical federated learning
Peng Chen 0030, Xin Du 0002, Zhihui Lu 0002, Hongfeng Chai
Comput. Secur.2
2024 CoLLaRS : A cloud-edge-terminal collaborative lifelong learning framework for AIoT
Shijing Hu 0001, Junxiong Lin, Zhihui Lu 0002, Xin Du 0002, Qiang Duan 0002, Shih-Chia Huang
Future Gener. Comput. Syst.4
2024 A balanced and reliable data replica placement scheme based on reinforcement learning in edge-cloud environments
Mengke Zheng, Xin Du 0002, Zhihui Lu 0002, Qiang Duan 0002
Future Gener. Comput. Syst.2
2024 Towards transferable adversarial attacks on vision transformers for image classification
Xu Guo 0004, Peng Chen 0030, Zhihui Lu 0002, Hongfeng Chai, Xin Du 0002
J. Syst. Archit.5
2024 HRCM: A Hierarchical Regularizing Mechanism for Sparse and Imbalanced Communication in Whole Human Brain Simulations
abstract
Brain simulation is one of the most important measures to understand how information is represented and processed in the brain, which usually needs to be realized in supercomputers with a large number of interconnected graphical processing units (GPUs). For the whole human brain simulation, tens of thousands of GPUs are utilized to simulate tens of billions of neurons and tens of trillions of synapses for the living brain to reveal functional connectivity patterns. However, as an application of the irregular spares communication problem on a large-scale system, the sparse and imbalanced communication patterns of the human brain make it particularly challenging to design a communication system for supporting large-scale brain simulations. To face this challenge, this paper proposes a hierarchical regularized communication mechanism, HRCM. The HRCM maintains a hierarchical virtual communication topology (HVCT) with a merge-forward algorithm that exploits the sparsity of neuron interactions to regularize inter-process communications in brain simulations. HRCM also provides a neuron-level partition scheme for assigning neurons to simulation processes to balance the communication load while improving resource utilization. In HRCM, neuron partition is formulated as a k-way graph partition problem and solved efficiently by the proposed hybrid multi-constraint greedy (HMCG) algorithm. HRCM performs finer-grained neuron-level communication control while leveraging voxel-level control as the basis, thus being more effective in balancing inter-process traffic in large-scale simulations. The hierarchical characteristics of the finer-grained communication control are considered by the problem formulation and algorithm design in HRCM. HRCM has been implemented in human brain simulations at the scale of up to 86 billion neurons running on 10000 GPUs. Results obtained from extensive simulation experiments verify the effectiveness of HRCM in significantly reducing communication delay, increasing resource usage, and shortening simulation time for large-scale human brain models.
Xin Du 0002, Minglong Wang, Zhihui Lu 0002, Qiang Duan 0002, Yuhao Liu 0008, Jianfeng Feng, Huarui Wang
IEEE Trans. Parallel Distributed Syst.1
2023 HSFL: Efficient and Privacy-Preserving Offloading for Split and Federated Learning in IoT Services
abstract
Distributed machine learning methods like Federated Learning (FL) and Split Learning (SL) meet the growing demands of processing large-scale datasets under privacy restrictions. Recently, FL and SL are combined in hybrid SLFL (SFL) frameworks to exploit both methods’ advantages to facilitate ubiquitous intelligence in the Internet of Things (IoT), for example, smart finance. Despite its significant impact on the performance and costs of SFL, model decomposition that splits an ML model into the client-server pair has not been sufficiently studied, especially for SFL in a large-scale dynamic IoT environment. In this paper, we propose a new SFL framework HSFL with a lightweight model decomposition method to offload a part of model training to the edge server. Specifically, we develop a method for estimating the training latency of HSFL and designed a metric for measuring privacy leakage in HSFL, based on which we formulate model decomposition in HSFL as an optimization problem with privacy protection as a constraint. Then, we transform the formulated problem into a contextual bandit problem and design an efficient algorithm to solve it. We have conducted thorough evaluations of the proposed HSFL framework through extensive experiments on a prototype testbed and a simulation platform. The experimental results validate the superiority of HSFL over the state-of-the-art benchmarks in terms of training latency, efficiency, scalability, and privacy protection.
Ruijun Deng, Xin Du 0002, Zhihui Lu 0002, Qiang Duan 0002, Shih-Chia Huang, Jie Wu 0003
ICWS2
2023 Fidan: a predictive service demand model for assisting nursing home health-care robots
abstract
While population aging has sharply increased the demand for nursing staff, it has also increased the workload of nursing staff.Although some nursing homes use robots to perform part of the work, such robots are the type of robots that perform set tasks.The requirements in actual application scenarios often change, so robots that perform set tasks cannot effectively reduce the workload of nursing staff.In order to provide practical help to nursing staff in nursing homes, we innovatively combine the LightGBM algorithm with the machine learning interpretation framework SHAP (Shapley Additive exPlanations) and use comprehensive data analysis methods to propose a service demand prediction model Fidan (Forecast service demand model).This model analyzes and predicts the demand for elderly services in nursing homes based on relevant health management data (including physiological and sleep data), ward round data, and nursing service data collected by IoT devices.We optimise the model parameters based on Grid Search during the training process.The experimental results show that the Fidan model has an accuracy rate of 86.61% in predicting the demand for elderly services.
Feng Zhou 0014, Xin Du 0002, Zhihui Lu 0002, Shih-Chia Huang
Connect. Sci.2
2023 A Blockchain-Assisted Intelligent Edge Cooperation System for IoT Environments With Multi-Infrastructure Providers
abstract
While edge computing has the potential to offer low-latency services and overcome the limitations of traditional cloud computing, it presents new challenges in terms of trust, security, and privacy (TSP) in Internet of Things environments. Cooperative edge computing (CEC) has emerged as a solution to address these challenges through resource sharing among edge nodes. However, for multi-infrastructure providers, incentive and trust mechanisms among edge nodes are crucial technical issues that must be addressed alongside system latency and reliability to meet performance requirements. In this article, we propose a blockchain-assisted intelligent edge cooperation system (BIECS) to systematically solve these issues. By leveraging blockchain technology, we construct trust among edge nodes and employ an incentive mechanism for resource sharing among multi-infrastructure providers. We formulate the system performance optimization as a multiobjective joint optimization problem and solve it efficiently through a two-stage strategy for selecting edge nodes. We first design an improved long short term memory (LSTM) model for resource prediction and then select edge nodes for executing offloaded tasks and handling the corresponding blockchain process related to each task execution. To evaluate the performance of BIECS, we implement the system based on Hyperledger Fabric and design extensive experiments. Our proposed system achieves better performance in terms of system delay, throughput, and resource utilization compared to state-of-the-art schemes for edge cooperation.
Xin Du 0002, Xuzhao Chen, Zhihui Lu 0002, Qiang Duan 0002, Jie Wu 0003, Patrick C. K. Hung
IEEE Internet Things J.1
2023 An Adaptive Mechanism for Dynamically Collaborative Computing Power and Task Scheduling in Edge Environment
abstract
Edge computing can provide high bandwidth and low-latency service for big data tasks by leveraging the edge side’s computing, storage, and network resources. With the development of microservice and docker technology, service providers can flexibly and dynamically cache microservice at the edge side to respond efficiently with limited resources. Automatically caching needed services on the nearest edge nodes and dynamically scheduling users’ requests can realize that computing power and software services flow with the users to provide continuous services. However, achieving the goal needs to overcome many challenges, such as the significant fluctuation of user devices’ requests at the edge side and the lack of collaboration among edge nodes. In this article, dynamic computing power scheduling and collaborative task scheduling among edge nodes are comprehensively developed. The problem is considered a multiobjective optimization problem, including sequentially minimizing the deadline missing rate of requests and the average task completion time. We propose an adaptive mechanism for dynamically collaborative computing power and task scheduling (ADCS) in the edge environment to solve this problem. It adopts the greedy decision method to schedule computing tasks to meet their deadline requirements. At the same time, it uses the best-fit method to adjust the computing resources according to the changes of users’ requests. The simulation results show that ADCS can decrease the deadline missing rate and reduce the average completion time. Compared with DSR and CoDSR, the deadline missing rate is reduced by 59.91% and 19.95%, respectively. The average completion time is decreased by 37.87% and 6.71%.
Yangchuan Xu, Lulu Chen, Zhihui Lu 0002, Xin Du 0002, Jie Wu 0003, Patrick C. K. Hung
IEEE Internet Things J.4
2022 Regularizing Sparse and Imbalanced Communications for Voxel-based Brain Simulations on Supercomputers
abstract
Inter-process communications form a performance bottleneck for large-scale brain simulations. The sparse and imbalanced communication patterns of human brain make it particularly challenging to design a communication system for supporting large-scale brain simulations. In this paper, we tackle the communication challenges posed by large-scale brain simulations with sparse and imbalanced communication patterns. We design a virtual communication topology with a merge and forward algorithm that exploits the sparsity to regularize inter-process communications. To balance the communication loads of different processes, we formulate voxel partition in brain simulations as a k-way graph partition problem and propose a constrained deterministic greedy algorithm to solve the problem effectively. We conducted extensive simulation experiments for evaluating the performance of the proposed communication scheme and found that the proposed method may significantly reduce communication overheads and shorten simulation time for large-scale brain models.
Yuhao Liu 0008, Xin Du 0002, Zhihui Lu 0002, Qiang Duan 0002, Jianfeng Feng, Minglong Wang, Jie Wu 0003
ICPP2
2022 BIECS: A Blockchain-based Intelligent Edge Cooperation System for Latency-Sensitive Services
abstract
Although the emerging edge computing paradigm offers a promising approach to overcoming some limitations of conventional cloud computing, the heterogeneous edge nodes with highly diverse system capacities bring new challenges to service provisioning especially for latency-sensitive services. Cooperative edge computing (CEC) has been proposed for facing such challenges through resource sharing among edge nodes. However, some technical issues must be fully addressed to make CEC effective, among which incentive and trust mechanisms and performance optimization are crucial for latency-sensitive service provision. In this paper, we design a novel blockchain-based intelligent edge cooperation system named BIECS to tackle these challenges systematically. BIECS provides incentive to edge nodes for resource sharing and enables trust among cooperative nodes upon a distributed platform leveraging the blockchain technology. In order to optimize system performance for meeting the requirements of latency-sensitive services, we propose a two-stage strategy for node selection in BIECS that chooses the most appropriate edge nodes for executing offloaded tasks and recording related transactions in the blockchain. We also implemented a prototype of BIECS based on Hyperledger Fabric and conducted extensive experiments for evaluating the performance of BIECS. The obtained experiment results verify that the proposed BIECS achieves better performance in system delay and throughput compared to the state-of-the-art methods for edge cooperation.
Xin Du 0002, Xuzhao Chen, Zhihui Lu 0002, Qiang Duan 0002, Jie Wu 0003
ICWS1
2022 Coordinate-based efficient indexing mechanism for intelligent IoT systems in heterogeneous edge computing
Songtao Tang, Xin Du 0002, Zhihui Lu 0002, Keke Gai, Jie Wu 0003, Patrick C. K. Hung, Kim-Kwang Raymond Choo
J. Parallel Distributed Comput.2
2022 EVFL: An explainable vertical federated learning for data-oriented Artificial Intelligence systems
Peng Chen 0030, Xin Du 0002, Zhihui Lu 0002, Jie Wu 0003, Patrick C. K. Hung
J. Syst. Archit.2
2022 Improved LSTM-Based Time-Series Anomaly Detection in Rail Transit Operation Environments
abstract
Anomaly detection is crucial to the reliability and safety of rail transit systems. The rapid development of Internet of Things (IoT) and cloud technologies together with recent advances in machine learning offered various cloud-based data-driven approaches to automatic anomaly detection. However, the challenges introduced by the different types of equipment in rail transit systems with highly diverse data distributions and the lack of labeled anomaly data have not been sufficiently addressed. In this article, we attempt to cope with such challenges by proposing an improved long short term memory (LSTM)-based time-series anomaly detection scheme. The key elements of the proposed scheme include an improved LSTM model that may achieve more accurate time-series prediction for various rail transit devices and a method for determining an appropriate error threshold for detecting anomalies based on the prediction errors. In order to further enhance anomaly detection performance, we also propose a pruning algorithm for reducing the number of false anomalies. Our method does not rely on scarce anomaly labels but dynamically determines a threshold of prediction errors to identify anomalies; therefore, it overcomes the challenge of the extremely uneven distribution of rail transit data. We conducted extensive experiments in a real metro operation environment for performance evaluation. The experiment results prove the effectiveness of the proposed scheme and show a superior performance of the scheme compared to existing anomaly detection methods.
Xin Du 0002, Zhihui Lu 0002, Qiang Duan 0002, Jie Wu 0003
IEEE Trans. Ind. Informatics2
2020 A Novel Data Placement Strategy for Data-Sharing Scientific Workflows in Heterogeneous Edge-Cloud Computing Environments
abstract
The deployment of datasets in the heterogeneous edge-cloud computing paradigm has received increasing attention in state-of-the-art research. However, due to their large sizes and the existence of private scientific datasets, finding an optimal data placement strategy that can minimize data transmission as well as improve performance, remains a persistent problem. In this study, the advantages of both edge and cloud computing are combined to construct a data placement model that works for multiple scientific workflows. Apparently, the most difficult research challenge is to provide a data placement strategy to consider shared datasets, both within individual and among multiple workflows, across various geographically distributed environments. According to the constructed model, not only the storage capacity of edge micro-datacenters, but also the data transfer between multiple clouds across regions must be considered. To address this issue, we considered the characteristics of this model and identified the factors that are causing the transmission delay. The authors propose using a discrete particle swarm optimization algorithm with differential evolution (DE-DPSO) to distribute dataset during workflow execution. Based on this, a new data placement strategy named DE-DPSO-DPS is proposed. DE-DPSO-DPS is evaluated using several experiments designed in simulated heterogeneous edge-cloud computing environments. The results demonstrate that our data placement strategy can effectively reduce the data transmission time and achieve superior performance as compared to traditional strategies for data-sharing scientific workflows.
Xin Du 0002, Songtao Tang, Zhihui Lu 0002, Jie Wu 0003, Keke Gai, Patrick C. K. Hung
ICWS1
2020 ORHRC: Optimized Recommendations of Heterogeneous Resource Configurations in Cloud-Fog Orchestrated Computing Environments
abstract
The cloud-fog orchestrated computing environments devolve computing tasks from the cloud center to the fog nodes, providing more heterogeneous configurations for the operation of workloads. Compared to the conventional cloud computing environment, the physical conditions at the fog nodes in the cloud-fog orchestrated computing environments are more complex and changeable. Therefore, the configurations that the fog nodes provide are heterogeneous and varying. This requires the configuration selection model to adapt to changeable configurations. The previous configuration selection models are applied to the limited and fixed configurations in the conventional cloud environment, but not to the complex cloud-fog orchestrated computing environments. To address this problem, we propose Optimized Recommendations of Heterogeneous Resource Configurations(ORHRC), a model that provides users with a reliable cloud configuration recommendation service. ORHRC uses the matrix factorization algorithm and neural network to build a recommendation model, which combines the operating characteristics of workloads as the explicit ratings and implicit feedback, to give configuration recommendations. Comprehensive experiments on a real-world dataset demonstrate that the hit rate of configurations of ORHRC is 24% higher than Micky and 15% higher than Selecta.
Ai Xiao, Zhihui Lu 0002, Xin Du 0002, Jie Wu 0003, Patrick C. K. Hung
ICWS3