Zhaolong Jian

dblp:282/4086 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-1543-3207ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Computer networks · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AEDS: An Affinity-Driven Efficient DRL-Based Task Scheduling Framework for Edge Computing
abstract
Edge computing is a promising paradigm that deploys computing resources at the network edge to provide services. Many existing solutions leverage deep reinforcement learning (DRL) to optimize task scheduling, yet they often rely on global scheduling approaches. However, such solutions result in an excessively large decision search space, reducing task scheduling efficiency in complex environments. Additionally, the cold start problem impedes the generation of optimal scheduling strategies. To address these challenges, we propose AEDS, a DRL-based task scheduling framework designed to enhance scheduling efficiency. AEDS optimizes the decision-making process from three aspects: (1)Decision Space Reduction.AEDS incorporates a novel affinity matching mechanism that identifies the most suitable edge cluster based on task characteristics, thereby significantly narrowing the decision search space. (2)Decision Process Optimization.AEDS adopts a hybrid strategy combining offline pre-training and online fine-tuning to address the cold start problem. Offline pre-training with historical task data ensures effective initial scheduling, while online fine-tuning periodically updates the DRL model to enhance long-term adaptability to dynamic system changes. (3)Decision Strategy Calibration.AEDS proposes a task migration solution to adapt to real-time workload variations dynamically. It utilizes triple queues to assess server overload and dynamically calibrates the scheduling strategy through task migration within interconnected clusters. Comprehensive experimental results validate the efficacy of AEDS. Compared with existing frameworks, AEDS reduces task latency by$28.23\%$and enhances task completion rate by$10.28\%$. Furthermore, by effectively narrowing the decision scope, AEDS accelerates the decision-making process by a remarkable$88.06\%$
Zhaolong Jian, Xueshuo Xie, Qiankun Dong, Mulin Li, Tao Li 0022
IEEE Trans. Mob. Comput.2
2025 Cochain: Architectural Support Mechanism for Blockchain-Based Task Scheduling
Yaozheng Fang, Yibing Jiang, Xueshuo Xie, Zhaolong Jian, Tao Li 0022, Zhiguo Wan, Grace Guiling Wang
APPT4
2025 Col-TEEs: Secure and Efficient Collaborative Inference Framework in Heterogeneous TEEs
abstract
Deep Neural Network (DNN) inference is a key enabler of edge intelligence, but its high computational demands and the imperative to protect the intellectual property of pretrained models present substantial challenges. Although recent research has started using Trusted Execution Environments (TEEs) for secure DNN inference, their limited memory and computational resources make it highly challenging to find an optimal balance between performance and security. To address this challenge, we propose Col-TEEs, a secure and efficient collaborative inference framework for heterogeneous TEEs. Col-TEEs strategically partitions the DNN inference process into device-side TEE and cloud-side TEE execution, aiming to maximize inference efficiency while ensuring model and data security. We first propose an optimal secure range determination method, which defines a secure interval by simulating model extraction and input reconstruction attacks, ensuring the confidentiality of pre-trained models and the privacy of user inputs. Furthermore, within this secure range, Col-TEEs employs an adaptive partitioning method that comprehensively evaluates the characteristics of heterogeneous TEE platforms. By combining a precise inference latency prediction model with heuristic algorithms, Col-TEEs dynamically selects the optimal partition point to balance security and inference performance. Additionally, Col-TEEs incorporates an integrity verification mechanism to secure intermediate data during network transmission. Extensive experiments demonstrate that Col-TEEs effectively balances performance and security, achieving an average$\mathbf{2. 9 8}$-fold increase in inference speed and a 63.54 % reduction in power consumption. Compared to collaborative inference without TEE protection, Col-TEEs introduces only a 4.65 % additional performance overhead.
Zhaolong Jian, Xueshuo Xie, Mulin Li
ICPADS3
2025 Hybrid-Granularity Parallelism Support for Fast Transaction Processing in Blockchain-Based Federated Learning
abstract
Blockchain-based Federated Learning (BCFL) is widely recognized as a promising solution for collaboratively training machine learning models while maintaining system security. Since blockchain systems are transaction-driven, the efficiency of transaction processing is directly related to the performance and availability of the BCFL system. Previous research has primarily focused on optimizing storage mechanisms or integrating Trusted Execution Environment (TEE) to reduce transaction processing pressure. However, the performance of BCFL remains constrained by slow transaction processing. This critical bottleneck arises from scalar instruction operations in transaction execution engines and the inherent serial transaction processing mechanism. In this paper, we propose a novel hybrid-granularity parallelism architecture, HGP, to greatly accelerate transaction processing in BCFL systems. HGP achieves this through three major innovations: (1) a suite of extended vector instructions, which reduces the instruction number and execution latency by enabling vectorized data I/O and computation using very long instruction word (VLIW) techniques, (2) the scalable transaction grouping method that generates parallelizable transaction groups through transaction signature verification and read-write conflict detection, and (3) the multi-EVM (Ethereum Virtual Machine) parallel processing mechanism that processes a group of transactions using multiple execution engine threads, and maintains global consistency through group scheduling. Through these optimizations, HGP accelerates the transaction processing with both data-level and thread-level parallelism. We evaluate HGP by executing BCFL tasks over classic ResNet18, MobileNet, and SqueezeNet. The experimental results demonstrate that HGP achieves up to a$3.8 \times$improvement in CPU utilization and a$1.6 \times$improvement in memory utilization. Furthermore, HGP significantly speeds up the transaction processing performance of three critical tasks by up to$24.5 \times, 12.4 \times$, and$2.8 \times$, respectively.
Mulin Li, Zhaolong Jian, Xueshuo Xie, Wajdy Othman
IPDPS2
2025 SmartZone: Runtime Support for Secure and Efficient On-Device Inference on ARM TrustZone
abstract
On-device inference is a burgeoning paradigm that performs model inference locally on end devices, allowing private data to remain local. ARM TrustZone as a widely supported trusted execution environment has been applied to provide confidentiality protection for on-device inference. However, with the rise of large-scale models like large language models (LLMs), TrustZone-based on-device inference faces challenges in migration difficulties and inefficient execution. The rudimentary TEE OS on TrustZone lacks both the inference runtime needed for building models and the parallel support necessary to accelerate inference. Moreover, the limited secure memory resources on end devices further constrain the model size and degrade performance. In this paper, we propose SmartZone to provide runtime support for secure and efficient on-device inference on TrustZone. SmartZone consists three main components: (1) a trusted inference-oriented operator set, providing the underlying mechanisms adapted to the TrustZone’s execution mode for trusted inference of DNN models and LLMs. (2) the proactive multi-threading parallel support, which increases the number of CPU cores in the secure state via cross-world thread collaboration to achieve parallelism, and (3) the on-demand secure memory management method, which statically allocates the appropriate secure memory size based on pre-execution resource analysis. We implement a prototype of SmartZone on the Raspberry Pi 3B+ board and evaluate it on four well-known DNN models and llama2 LLM. Extensive experimental results show that SmartZone provides end-to-end protection for on-device inference while maintaining excellent performance. Compared to the origin trusted inference, SmartZone accelerates the inference speed by up to 4.26× and reduces energy consumption by 65.81%.
Zhaolong Jian, Qiankun Dong, Longkai Cheng, Xueshuo Xie, Tao Li 0022
IEEE Trans. Computers1
2024 Memory-Efficient and Secure DNN Inference on TrustZone-enabled Consumer IoT Devices
abstract
Edge intelligence enables resource-demanding Deep Neural Network (DNN) inference without transferring original data, addressing concerns about data privacy in consumer Inter-net of Things (IoT) devices. For privacy-sensitive applications, deploying models in hardware-isolated trusted execution environments (TEEs) becomes essential. However, the limited secure memory in TEEs poses challenges for deploying DNN inference, and alternative techniques like model partitioning and offloading introduce performance degradation and security issues. In this paper, we present a novel approach for advanced model deployment in TrustZone that ensures comprehensive privacy preservation during model inference. We design a memory-efficient management method to support memory-demanding inference in TEEs. By adjusting the memory priority, we effectively mitigate memory leakage risks and memory overlap conflicts, resulting in 32 lines of code alterations in the trusted operating system. Additionally, we leverage two tiny libraries: S-Tinylib (2,538 LoCs), a tiny deep learning library, and Tinylibm (827 LoCs), a tiny math library, to support efficient inference in TEEs. We implemented a prototype on Raspberry Pi 3B+ and evaluated it using three well-known lightweight DNN models. The experimental results demonstrate that our design significantly improves inference speed by 3.13 times and reduces power consumption by over 66.5% compared to non-memory optimization method in TEEs.
Xueshuo Xie, Haoxu Wang, Zhaolong Jian, Tao Li 0022, Wei Wang 0012, Grace Guiling Wang
INFOCOM3
2024 DRS: A deep reinforcement learning enhanced Kubernetes scheduler for microservice-based system
abstract
Summary Recently, Kubernetes is widely used to manage and schedule the resources of microservices in cloud‐native distributed applications, as the most famous container orchestration framework. However, Kubernetes preferentially schedules microservices to nodes with rich and balanced CPU and memory resources on a single node. The native scheduler of Kubernetes, called Kube‐scheduler, may cause resource fragmentation and decrease resource utilization. In this paper, we propose a deep reinforcement learning enhanced Kubernetes scheduler named DRS. We initially frame the Kubernetes scheduling problem as a Markov decision process with intricately designed state , action , and reward structures in an effort to increase resource usage and decrease load imbalance. Then, we design and implement DRS mointor to perceive six parameters concerning resource utilization and create a thorough picture of all available resources globally. Finally, DRS can automatically learn the scheduling policy through interaction with the Kubernetes cluster, without relying on expert knowledge about workload and cluster status. We implement a prototype of DRS in a Kubernetes cluster with five nodes and evaluate its performance. Experimental results highlight that DRS overcomes the shortcomings of Kube‐scheduler and achieves the expected scheduling target with three workloads. With only 3.27% CPU overhead and 0.648% communication delay, DRS outperforms Kube‐scheduler by 27.29% in terms of resource utilization and reduces load imbalance by 2.90 times on average.
Zhaolong Jian, Xueshuo Xie, Yaozheng Fang, Yibing Jiang, Ye Lu 0004, Ankan Dash, Tao Li 0022, Grace Guiling Wang
Softw. Pract. Exp.1
2023 TSC-VEE: A TrustZone-Based Smart Contract Virtual Execution Environment
abstract
TrustZone as a trusted execution environment (TEE) has been proven to preserve the confidentiality of blockchain transactions supported by smart contracts. Despite some academic effort, TrustZone can only support limited languages for now. The lack of the corresponding execution environment for smart contracts seriously hinders blockchain applications from directly running on TrustZone. In this paper, we design the first virtual execution environment named TSC-VEE for performing Solidity smart contracts on TrustZone, to the best of our knowledge. TSC-VEE can be decomposed into fourfold: (1) an instruction set adapted to the isolation and world switching mechanism of TrustZone. (2) a runtime memory management mechanism that provides a pair of instructions with the corresponding processing mechanism to allocate and release the work memory. (3) a hybrid granularity resource analysis algorithm which computes and records the value of maximum stack height and static gas cost through bytecode pre-execution, avoiding runtime overflow and invalid computations. (4) a cross-isolation-environment prefetching approach that supports loading and storing the storage data from the normal world into the secure world on TrustZone before execution, thus avoiding switching the world state frequently at runtime. Extensive experimental results show that TSC-VEE can perform smart contracts correctly and efficiently on TrustZone. Compared with the most commonly used Ethereum client—Geth, TSC-VEE achieves execution performance improvements by$9.29\times$. We also implement the Ethereum virtual machine—evmoneon TrustZone. TSC-VEE can reduce the latency by 12.63% with our optimization techniques, and decrease the work memory footprint by 22.95% on average when executing various scale contracts.
Zhaolong Jian, Ye Lu 0004, Youyang Qiao, Yaozheng Fang, Xueshuo Xie, Dayi Yang, Tao Li 0022
IEEE Trans. Parallel Distributed Syst.1
2022 ATOM: Architectural Support and Optimization Mechanism for Smart Contract Fast Update and Execution in Blockchain-Based IoT
abstract
Blockchain-based Internet of Things (BC-IoT) brings the advantages of blockchain into traditional IoT systems. In BC-IoT, the smart contract has been widely used for automatic, trusted, and decentralized applications. Smart contracts require frequent adjust and fast update due to various reasons, such as inevitable code bugs, changes of applications, or security requirements. However, previous smart contract architecture and updating mechanism are low speed and cause high overhead, because they are based on recompilation and redeployment in BC-IoT. Meanwhile, smart contract execution is so time consuming due to contract instruction dispatching and operand loading in the stack-based Ethereum virtual machine (EVM). To address these issues, we propose a new smart contract architecture and optimization mechanism for BC-IoTs, ATOM, which provides architectural supports to update contract economically and fast executing in instructionwise for the first time, to the best of our knowledge. We design a compact Application-oriented Instruction (AoI) set to describe application operations. We can construct the bytecode of smart contract from application by directly assembling templates prebuilt upon the AoIs rather than by compilation. We also present an optimized mechanism for AoI execution to enable access addressable storage place rather than the indirect access through stack. We perform ATOM on a BC-IoT testbed based on private Ethereum and Hyperledger Burrow. The experimental results highlight that ATOM is more efficient than state-of-the-art approaches. ATOM can reduce update latency by 62.7%, ledger size by 70%, and gas usage by 90% on average, respectively. Compared with the traditional smart contract architecture, ATOM can improve EVM Memory access efficiency significantly by up to$10\times $and achieve improvement of execution efficiency with up to$1.6\times $.
Tao Li 0022, Yaozheng Fang, Zhaolong Jian, Xueshuo Xie, Ye Lu 0004, Grace Guiling Wang
IEEE Internet Things J.3
2022 SmartVM: A Smart Contract Virtual Machine for Fast On-Chain DNN Computations
abstract
Blockchain-based artificial intelligence (BC-AI) has been applied for protecting deep neural network (DNN) data from being tampered with, which is expected to further boost trusted distributed AI applications in many fields. However, due to smart contract execution environment architectural defects, it is challenging for previous BC-AI systems to support computing-intensive tasks on-chain performing such as DNN convolution operations. They have to offload computations and a large amount of data from blockchain to off-chain platforms to execute smart contracts as native code. This failure to take advantage of data locality has become one of the major critical performance bottlenecks in BC-AI system. To this end, in this article, we propose SmartVM with optimization methods to support on-chain DNN inference for BC-AI system. The key idea is to design and optimize the computing mechanism and storage structure of smart contract execution environment according to the characteristics of DNN such as high computational parallelism and large data volume. We decompose SmartVM into three components: 1) a compact DNN-oriented instruction set to describe computations in a short number of instructions to reduce interpretation time. 2) a memory management mechanism to make SmartVM memory dynamic free/allocated according to the size of DNN feature maps. 3) a block-based weight prefetching and parallel computing method to organize each layer's computing and weights prefetching in a pipelined manner. We perform the typical image classification in a private Ethereum blockchain testbed to evaluate SmartVM performance. Experimental results highlight that SmartVM can support DNN inference on-chain with roughly the same efficiency against the native code execution. Compared with the traditional off-chain computing, SmartVM can speed up the overall execution by70×,16×,11×, and12×over LeNet5, AlexNet, ResNet18, and MobileNet, respectively. The memory footprint can be reduced by84%,90.8%,94.3%, and93.7%over the above four models, while offering the same level model accuracy. This article sheds light on the design space of the smart contract virtual machine for DNN computation and is promising to further boost BC-AI applications.
Tao Li 0022, Yaozheng Fang, Ye Lu 0004, Jinni Yang, Zhaolong Jian, Zhiguo Wan, Yusen Li
IEEE Trans. Parallel Distributed Syst.5
2021 WIP: Sysnif: Constructing Workflow from Interleaved Logs in Intelligent IoT System
abstract
The massive smart devices in intelligent IoT can be broken due to malicious attacks and system failures. As a nonintrusive method, workflows mined from system logs facilitate administrators to quickly locate and diagnose anomalies in time. System logs are usually interleaved since there are lots of concurrent and asynchronous operations and executions on large scale IoT devices. Consequently, it is so challenging to construct an adaptive workflow from these logs and realize the real-time anomaly detection. To meet this challenge, in this paper, we propose a two-stage workflow construction approach named Sysnif, which includes offline construction and online adjustment. First, the window-based dependence computing method is employed to obtain the context of execution paths. Second, a weight-greedy algorithm is designed to denoise the interleaved system logs effectively. Third, in order to match system mechanism variation, the online micro-iteration adjusting algorithm is presented to update the workflow model. Experiment results highlight that Sysnif can outperform state-of-the-art methods, such as Logsed, on dataset of OpenStack logs by 22.4% on recall, meanwhile maintaining the same precision roughly. Sysnif can achieve an average precision and recall of 93.8% and 94.7%, respectively.
Zongming Jin, Xueshuo Xie, Yaozheng Fang, Zhaolong Jian, Ye Lu 0004, Guangying Li
WOWMOM4
2021 Fast Policy Interpretation and Dynamic Conflict Resolution for Blockchain-Based IoT System
abstract
Although the blockchain‐based Internet of Things (BC‐IoT) has been applied in many fields, it still faces many security attacks due to lacking policy‐based security management (PbSM). Previous PbSM is usually time‐consuming, which is difficult to integrate into BC‐IoT directly. The high‐latency policy conflict resolving in traditional PbSM cannot meet the BC‐IoT’s low‐latency requirement. Moreover, the conflict resolution rate is low as the PbSM usually neglects the runtime information. Therefore, it is challenging that achieving an efficient PbSM for BC‐IoT and overcomes both time and resource consumption. To address the problem, we propose a novel PbSM for BC‐IoT named FPICR to realize fast policy interpretation and dynamic conflict resolution efficiently. We first present policy templates based on system log to interpret policy in high speed in BC‐IoT. Benefiting from matching the characteristics of the system processing, FPICR supports interpreting a policy into the smart contract directly without complex content parsing. We then propose a weighted directed policy graph (WDPG) to evaluate the importance of the deployed policies more accurately. To improve the policy conflict resolution rate, we implement the resolution algorithm through reconstructing the WDPG. Taking the traits of these properties, FPICR thus can also remove the redundant data to compress storage space by the WDPG. Experiment results highlight that FPICR outperforms the baseline in all measure metrics. Especially, compared with the state‐of‐the‐art method, the speedup of interpretation in FPICR is about up to 2.1×. The conflict resolution rate in FPICR can be improved by 6.2% on average and achieve up to 96.1%.
Yaozheng Fang, Zhaolong Jian, Zongming Jin, Xueshuo Xie, Ye Lu 0004, Tao Li 0022
Wirel. Commun. Mob. Comput.2