Zhihui Lu 0002

dblp:51/4748 · also ZhiHui Lu 0002, ZhiHui Lv 0002 · DBLP profile ↗
← Back
91ranked-venue papers
12as first author
49since 2021 · last 2026
0000-0001-5706-7503ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 5 first-author · 11 since 2021Computer networks · 14 · 13 since 2021Software engineering, systems software and programming languages · 13 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 8 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Security and privacy · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
YearPublicationVenuePosition
2026 InfoDecom: Decomposing Information for Defending Against Privacy Leakage in Split Inference
abstract
Split inference (SI) enables users to access deep learning (DL) services without directly transmitting raw data. However, recent studies reveal that data reconstruction attacks (DRAs) can recover the original inputs from the smashed data sent from the client to the server, leading to significant privacy leakage. While various defenses have been proposed, they often result in substantial utility degradation, particularly when the client-side model is shallow. We identify a key cause of this trade-off: existing defenses apply excessive perturbation to redundant information in the smashed data. To address this issue in computer vision tasks, we propose InfoDecom, a defense framework that first decomposes and removes redundant information and then injects noise calibrated to provide theoretically guaranteed privacy. Experiments demonstrate that InfoDecom achieves a superior utility-privacy trade-off compared to existing baselines.
Ruijun Deng, Zhihui Lu 0002, Qiang Duan 0002
AAAI2
2026 TPipe: Efficient Spiking Transformer Training with Time Parallelism and Asynchronous Pipeline
Yubing Bao, Zhihui Lu 0002, Qiang Duan 0002, Changze Lv, Xin Du 0002, Zeyi Deng, Jingqi Feng, Sen Liu 0002, Yang Chen 0001, Xin Wang 0002
INFOCOM2
2026 RenHaze: A Coarse-to-Fine Rendering Framework for Improving Robustness to Haze
abstract
Large-scale datasets centered on images have driven advancements in deep learning-based computer vision applications. While there is an abundance of datasets containing images depicting favorable weather scenes, datasets featuring images of adverse weather conditions, especially the presence of haze, are scarce due to challenges in their collection. In response, we leverage the advantages of deep learning techniques to introduce a novel approach for facilitating the rendering of realistic and diverse hazy images, named RenHaze. To be specific, RenHaze adopts a denseness parameter \(\omega\) to control the haze level of output images and consists of five subnets, including a content exploitation (CE) subnet, a depth exploitation (DE) subnet, a haze exploitation (HE) subnet, an image generation (IG) subnet, and an image discernment (ID) subnet. The CE, DE, and HE subnets are responsible for extracting features from the source clear image, depth image, and reference hazy image, respectively, and then providing them for the IG subnet. The IG subnet is used to perform image translation in a coarse-to-fine manner, while the ID subnet is employed to discern the realism of the rendered image and provide feedback to the IG subnet for generating the desired output. Extensive experiments demonstrate the superiority of the proposed model over competing IG methods in terms of the realism and diversity of synthesized hazy images, as well as its effectiveness in boosting the performance of computer vision tasks such as object detection and semantic segmentation in real-world hazy environments.
Trung-Hieu Le, Shih-Chia Huang, Quoc-Viet Hoang, Zhihui Lu 0002
ACM Trans. Intell. Syst. Technol.4
2025 Efficient Joint Communication and Computation Placement for Large-scale SNN Simulation on Supercomputers
abstract
Spiking Neural Network (SNN) simulation involves emulating the activation and firing of spiking neurons on hardware platforms. This is a highly time-sensitive task, requiring the simulation of billions of neurons and their intercommunication within a few milliseconds. Each neuron performs a complex, interdependent multi-stage communication and computation task. We consider the task placement of SNN on supercomputers to accelerate SNN simulation. Existing task placement methods for SNN simulations have two major limitations. First, they lack the capability to handle large-scale SNNs with billions of neurons. Second, they focus primarily on optimizing communication delay, while neglecting multi-stage computation delays in SNN simulations. In this paper, we formalize the SNN Joint Multi-stage Communication and Computation Placement (SJCCP) problem. We demonstrate that SJCCP can be solved using an approximation algorithm with an approximation ratio of $O\left( {{k^2}\sqrt {\log n\log k} } \right)$, where n is the number of voxels in the SNN and k is the number of GPUs. To further reduce the time complexity of solving SJCCP in practice, we propose a novel efficient framework, FastSJP, tailored for large-scale SNN placement. Then we apply the FastSJP framework to a human brain simulation that runs a large-scale SNN model derived from authentic biological data on a supercomputer equipped with 1024 GPUs. Experimental results verify that our framework notably reduces time overhead, ranging from 17.31% to 28.45%, compared to state-of-the-art methods. Leveraging the computational power of the supercomputer, FastSJP maximizes the problem size and processing performance, significantly advancing the development of brain-inspired intelligence.
Yubing Bao, Zhihui Lu 0002, Xin Du 0002, Qiang Duan 0002, Jirui Yang, Jin Zhao 0001, Geyong Min, Yang Chen 0001, Shijing Hu 0001, Xin Wang 0002
ICDCS2
2025 BMapper: A Scalable and Efficient Framework for Brain Simulations Acceleration on Supercomputers
abstract
Brain simulation is an inherently highly parallel and time-sensitive task, requiring the simulation of billions of neurons and their interactions within just a few milliseconds. With the growing availability of brain data from biological research, more realistic and detailed simulations are becoming feasible. However, this also poses unprecedented challenges for parallel computing due to the extreme sparsity and heterogeneity of the emerging workloads. Efficient deployment of such workloads on modern HPC systems is critical to overcoming these challenges. We propose BMapper, a deployment framework that enables efficient parallel execution of brain simulations on supercomputers. BMapper comprises three synergistic components: BPartitioning, which introduces a novel multi-dimensional hybrid partitioning strategy to balance workloads across GPUs and reduce inter-GPU spike traffic; BPlacement, which applies deterministic spectral partitioning to minimize inter-server communication; and BRelaying, which identifies lightly loaded GPUs to assist the top-k heavily loaded ones by relaying spike traffic. These components work together to balance loads and minimize communication overhead, enabling high-speed simulation of large-scale brain models. BMapper has been deployed to simulate up to 10 billion neurons on a 1000-GPU supercomputer, achieving 25.15%–47.48% faster execution than state-of-the-art methods.
Yubing Bao, Zhihui Lu 0002, Qiang Duan 0002, Xin Du 0002, Yandan Tan, Yang Chen 0001, Yang Xu 0010
ICPP2
2025 PreFabric: Eliminating Conflicts for High-Throughput Permissioned Blockchains
abstract
Permissioned blockchains have found widespread adoption across diverse scenarios, ensuring data authenticity and integrity. However, transaction conflicts, as an inherent performance challenge in permissioned blockchains, can significantly decrease system throughput and thus degrade its Quality of Service (QoS) under substantial transaction contention. Existing approaches mitigate conflicts typically by either aborting or blocking transactions in advance, encountering two main issues: (i) resource wastage due to transaction failure and (ii) performance degradation, particularly under large block sizes or high transaction contention. In this paper, we propose PreFabric, a novel permissioned blockchain framework that guarantees high throughput by resolving the transaction conflict problem. We first conduct a comprehensive analysis of the transaction scenarios preceding simulation execution of the endorsing phase in the blockchain system to identify potential conflict-causing situations. Then, we devise an key-locking method to prevent transaction conflicts and propose concurrency control strategies based on dependency analysis, encompassing a transaction merging mechanism, an key-renaming mechanism and concurrent validating mechanisms, to improve system throughput. The experimental results demonstrate the superior performance of our method over state-of-the-art methods, with 2.1× higher effective throughput and 0.48× lower latency.
Junxiong Lin, Zhihui Lu 0002, Yiguang Zhang, Ruijun Deng, Qiang Duan 0002, Hengqi Guo, Xu Guo 0004, Baoqi Huang
ICWS2
2025 UIFV: Data Reconstruction Attack in Vertical Federated Learning
abstract
Vertical Federated Learning (VFL) enables collaborative machine learning without the need for participants to share their raw private data. However, recent studies have uncovered privacy risks, where adversaries might reconstruct sensitive features through data leakage during the learning process. Al-though existing data reconstruction methods are effective to some extent, they exhibit limitations in VFL scenarios, as initiating an attack requires meeting more stringent conditions. To gain a comprehensive understanding of the risks of data reconstruction in VFL, this paper proposes a unified framework, the Unified InverNet Framework in VFL (UIFV), for data reconstruction under realistic black-box threat models. Within the UIFV framework, we consider four attack scenarios, strictly adhering to VFL protocols to maintain confidentiality. Experiments on four datasets show that our methods significantly outperform state-of-the-art techniques in terms of applicability and attack precision. Our work reveals severe privacy vulnerabilities within VFL systems that pose real threats to practical VFL applications, thus confirming the necessity of further enhancing privacy protection in the VFL architecture. Overall, this paper provides a thorough analysis of the risks of data reconstruction in VFL and offers important guidance to enhance the security of VFL deployments.
Jirui Yang, Peng Chen 0030, Zhihui Lu 0002, Qiang Duan 0002, Yubing Bao
ICWS3
2025 Dynamic Model and Node Selection for Collaborative Inference of Large/Small Models in Vehicular Networks
abstract
Collaborative inference between large cloud-hosted models and small edge-deployed models offers a promising solution for balancing the accuracy and efficiency of ML-based applications in vehicular networks. Selecting the appropriate models and their hosting nodes for performing various inference tasks plays a crucial role in collaborative inference in vehicular networks. However, existing solutions, primarily based on deep reinforcement learning (DRL), suffer critical limitations, including delayed and suboptimal decisions on model and node selection in dynamic environments. To address these challenges, we propose a dynamic model and node selection strategy for a collaborative inference framework, grounded in active inference theory. Our strategy dynamically aligns task requirements with model capabilities and node capacities by considering factors such as vehicular mobility, latency constraints, task complexity, and model accuracy. Additionally, when significant drops in inference accuracy are detected, we fine-tune and update the models deployed on both the edge and cloud, ensuring reliable and up-to-date inference. By leveraging active inference to minimize free energy through Bayesian belief updates, our framework reduces average latency by 23.2%, lowers task failure rates by 67%, and achieves superior load balancing compared to existing methods. It also demonstrates robust dynamic performance with a 5.1% failure rate under 200% traffic surges, and its hybrid update strategy maintains 85.4% accuracy after 72 hours, effectively addressing the complex and dynamic conditions of vehicular networks.
Mengke Zheng, Zhihui Lu 0002, Qiang Duan 0002, Baoqi Huang, Shijing Hu 0001
ICWS2
2025 Universal Backdoor Defense via Label Consistency in Vertical Federated Learning
abstract
Backdoor attacks in vertical federated learning (VFL) are particularly concerning as they can covertly compromise VFL decision-making, posing a severe threat to critical applications of VFL. Existing defense mechanisms typically involve either label obfuscation during training or model pruning during inference. However, the inherent limitations on the defender's access to the global model and complete training data in VFL environments fundamentally constrain the effectiveness of these conventional methods. To address these limitations, we propose the Universal Backdoor Defense (UBD) framework. UBD leverages Label Consistent Clustering (LCC) to synthesize plausible latent triggers associated with the backdoor class. This synthesized information is then utilized for mitigating backdoor threats through Linear Probing (LP), guided by a constraint on Batch Normalization (BN) statistics. Positioned within a unified VFL backdoor defense paradigm, UBD offers a generalized framework for both detection and mitigation that critically does not necessitate access to the entire model or dataset. Extensive experiments across multiple datasets rigorously demonstrate the efficacy of the UBD framework, achieving state-of-the-art performance against diverse backdoor attack types in VFL, including both dirty-label and clean-label variants.
Peng Chen 0030, Haolong Xiang, Xin Du 0002, Xiaolong Xu 0001, Xuhao Jiang, Zhihui Lu 0002, Jirui Yang, Qiang Duan 0002, Wan-Chun Dou
IJCAI6
2025 Backdoor Attack on Vertical Federated Graph Neural Network Learning
abstract
Federated Graph Neural Network (FedGNN) integrate federated learning (FL) with graph neural networks (GNNs) to enable privacy-preserving training on distributed graph data. Vertical Federated Graph Neural Network (VFGNN), a key branch of FedGNN, handles scenarios where data features and labels are distributed among participants. Despite the robust privacy-preserving design of VFGNN, we have found that it still faces the risk of backdoor attacks, even in situations where labels are inaccessible. This paper proposes BVG, a novel backdoor attack method that leverages multi-hop triggers and backdoor retention, requiring only four target-class nodes to execute effective attacks. Experimental results demonstrate that BVG achieves nearly 100% attack success rates across three commonly used datasets and three GNN models, with minimal impact on the main task accuracy. We also evaluated various defense methods, and the BVG method maintained high attack effectiveness even under existing defenses. This finding highlights the need for advanced defense mechanisms to counter sophisticated backdoor attacks in practical VFGNN applications.
Jirui Yang, Peng Chen 0030, Zhihui Lu 0002, Jianping Zeng 0002, Qiang Duan 0002, Xin Du 0002, Ruijun Deng
IJCAI3
2025 GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
abstract
Speculative decoding accelerates inference in large language models (LLMs) by generating multiple draft tokens simultaneously. However, existing methods often struggle with token misalignment between the training and decoding phases, limiting their performance. To address this, we propose GRIFFIN, a novel framework that incorporates a token-alignable training strategy and a token-alignable draft model to mitigate misalignment. The training strategy employs a loss masking mechanism to exclude highly misaligned tokens during training, preventing them from negatively impacting the draft model's optimization. The token-alignable draft model introduces input tokens to correct inconsistencies in generated features. Experiments on LLaMA, Vicuna, Qwen and Mixtral models demonstrate that GRIFFIN achieves an average acceptance length improvement of over 8\% and a speedup ratio exceeding 7\%, outperforming current speculative decoding state-of-the-art methods. Our code and GRIFFIN's draft models will be released publicly in https://github.com/hsj576/GRIFFIN.
Shijing Hu 0001, Xingyu Xie, Zhihui Lu 0002, Kim-Chuan Toh, Pan Zhou 0002
NeurIPS4
2025 Friend Discovery Scheme with Privacy Protection in Mobile Social Networks
Chaoliang Li, Xuwei Zhu, Zhihui Lu 0002, Hua Huang 0006
WASA (1)4
2025 Quantum Machine Learning: Hybrid System of Quantum and Classical Computing
Lian Peng, Meikang Qiu, Zhihui Lu 0002
WASA (1)4
2025 A Data Replication Placement Strategy for the Distributed Storage System in Cloud-Edge-Terminal Orchestrated Computing Environments
abstract
Cloud-edge-terminal orchestrated computing, as an expansion of cloud computing, has sunk resources to the edge nodes and terminal equipment, which can provide high-quality services for delay-sensitive applications and reduce the cost of network communication. Due to the high volume of data generated by Internet of Things (IoT) devices and the limited storage capacities of edge nodes, a significant number of terminal devices are now being considered for utilization as storage nodes. However, because of the heterogeneous storage capacity and reliability of these hardware devices and the different data requirements of user services, the performance and storage reliability of applications deployed in cloud-edge-terminal orchestrated computing environments have become urgent problems to be solved. Especially, for a distributed storage system in these environments, it is required to ensure reliable storage of the generated data and its’ replications. In this paper, we first implement a distributed storage system and construct a data replication placement model. Then, based on the constructed model, we formulate the data replication placement problem and design a data replication placement strategy called DRPS to solve it. The DRPS covers a ranks-based replication storage node selection algorithm and a greedy load balancing algorithm, which can select appropriate hardware devices for different data requirements of services and is implemented in the data storage system to store replications and balance loads. We design extensive experiments to verify the effectiveness of DRPS. The results indicate that the proposed strategy outperforms other state-of-the-art algorithms in terms of system delay reduction by 39.9%, an increase of 43.3% in the replication numbers, a 27.5% improvement in memory utilization, and a reduction of unreliability rate by 82.0%.
Peng Chen 0030, Mengke Zheng, Xin Du 0002, Muhammad Bilal 0003, Zhihui Lu 0002, Qiang Duan 0002, Xiaolong Xu 0001
IEEE Internet Things J.5
2025 Distributed DRL-Based Integrated Sensing, Communication, and Computation in Cooperative UAV-Enabled Intelligent Transportation Systems
abstract
The integration of sensing, communication, and computation (ISCC) is a critical technology that will support various emerging wireless services in future 6G networks. The unmanned aerial vehicles (UAVs) equipped with edge servers can be used as an aerial service platform in intelligent transportation systems (ITSs) to offer ISCC services to vehicles. This article studies an aerial UAV network comprising a central UAV and secondary UAVs to realize sensing of the global ITS environment and data fusion computation through collaborative UAVs. To enhance the service performance of ISCC, we maximize the success rate of ISCC services and the energy efficiency of UAVs by jointly optimizing bandwidth allocation, power allocation, and computing capacity control while ensuring the sensing and data processing latency requirements. Leveraging the network architecture and collaboration requirements of UAVs, we propose the multi-UAV collaborative Air-ISCC (MCAI) algorithm based on the asynchronous advantage actor-critic algorithm, which obtains the optimal ISCC service policy by co-training a deep reinforcement learning model with multiple UAVs. Sufficient experimental results show that MCAI enhances energy efficiency by 10.51% to 80.12% compared with the baselines. Moreover, MCAI exhibits good scalability, strengthening its feasibility in real scenarios.
Peng Hou 0003, Yi Huang 0020, Hongbin Zhu, Zhihui Lu 0002, Shih-Chia Huang, Yang Yang 0001, Hongfeng Chai
IEEE Internet Things J.4
2025 The Effect of Domain Terms on Password Security
abstract
The predominant authentication method still relies on usernames and passwords. To enhance memorability, domain terms may have been opted to include as part of passwords. However, there is little analysis of the extent to which such practice affects password security, so there is a lack of guidance on how users use domain terms on websites with different domain characteristics. To address the problem, we propose a novel approach to analyze the security effect of using domain terms in passwords. The methodology primarily consists of three stages. First, we utilize Web crawlers to harvest domain vocabularies, subsequently leveraging the TextRank algorithm to rank their importance. Second, we propose an algorithm for constructing a simulated domain-specific password dataset by replacing password elements with domain terms. Third, password guessing experiments are done on the dataset using PCFG (Probabilistic Context-Free Grammar) and the Markov model to evaluate the impact of domain terms on password security. The experimental results indicate that, for systems without clear domain, 20% domain terms replacement in the test set can reduce the cracking rate by up to 5.45%. In contrast, for domain-specific systems, 20% domain terms replacement in the training set can increase the cracking rate by 6.45%. These findings provide practical guidance on the application of domain knowledge in password creation for different types of systems. In summary, this study offers a novel perspective for exploring the security implications of passwords influenced by specific domains.
Yubing Bao, Jianping Zeng 0002, Jirui Yang, Ruining Yang, Zhihui Lu 0002
ACM Trans. Priv. Secur.5
2025 Amalgamating Knowledge for Object Detection in Rainy Weather Conditions
abstract
In recent years, object detection has significantly advanced by using deep learning, especially convolutional neural networks. Most of the existing methods have focused on detecting objects under favorable weather conditions and achieved impressive results. However, object detection in the presence of rain remains a crucial challenge owing to the visibility limitation. In this article, we introduce an amalgamating knowledge network (AK-Net) to deal with the problem of detecting objects hampered by rain. The proposed AK-Net obtains performance improvement by associating object detection with visibility enhancement, and it is composed of five subnetworks: rain streak removal (RSR) subnetwork, raindrop removal (RDR) subnetwork, foggy rain removal (FRR) subnetwork, feature transmission (FT) subnetwork, and object detection (OD) subnetwork. Our approach is flexible; it can adopt different object detection models to construct the OD subnetwork for the final inference of objects. The RSR, RDR, and FRR subnetworks are responsible for producing clean features from rain streak, raindrop, and foggy rain images, respectively, and offer them to the OD subnetwork through the FT subnetwork for efficient object prediction. Experimental results indicate that the mean average precision (mAP) achieved by our proposed AK-Net was up to 19.58% and 26.91% higher than those produced using competitive methods on published iRain and RID datasets, respectively, while preserving the fast-running time of the baseline detector.
Trung-Hieu Le, Shih-Chia Huang, Quoc-Viet Hoang, Zdenek Lokaj, Zhihui Lu 0002
ACM Trans. Intell. Syst. Technol.5
2024 FinDS2: A Novel Data Synthesis System for Fintech Product Risks
abstract
This paper starts from the application scenarios and pain points in the financial industry, focusing on the synthesis of financial technology (Fintech) product risk data. By integrating existing research in the field of data synthesis, we propose a novel Fintech Data Synthesis System (FinDS2). To implement this system, we have designed and developed a data synthesis product that leverages cloud computing and micro-services. The product includes a cloud-native data lake, an intelligent algorithm engine, and a SAAS tool engine. It collects data on regulatory rules and penalties in the finance industry and offers highly configurable and automated features. By tailoring the synthesis process to Fintech application scenarios, our proposed FinDS2 generates high-quality synthetic data efficiently. Through practical validation, we demonstrate that the synthesized data closely resembles the original data, effectively meeting the data synthesis needs for managing Fintech product risk.
Xiaozheng Du, Xu Guo 0004, Feng Zhou 0014, Mingyu Gu, Zhihui Lu 0002
CSCloud5
2024 An Adversarial Attack Method Against Financial Fraud Detection Model Beta Wavelet Graph Neural Network via Node Injection
abstract
Financial fraud detection plays a crucial role in maintaining financial security and risk control. Many types of financial data, such as transaction networks and entity relationship networks, can be represented as graph-structured data. Graph neural networks, exemplified by the Beta Wavelet Graph Neural Network (BWGNN), are instrumental in financial fraud detection. However, current research on adversarial attacks against graph neural networks primarily focuses on vanilla GNNs, with limited exploration into adversarial attacks targeting models like BWGNN used in financial fraud detection. This paper presents effective adversarial attack methods tailored to BWGNN. Leveraging node injection as an adversarial attack method, we construct surrogate models that closely resemble the structure of BWGNN, significantly enhancing the attack performance. Additionally, by incorporating dropout layers after the input layer of the surrogate model, we further enhance the attack effectiveness. This paper reveals the adversarial vulnerabilities of financial fraud detection models represented by BWGNN, which holds significant implications for enhancing the security of fraud detection models applied in critical financial security domains.
Hengqi Guo, Xiaozheng Du, Jirui Yang, Zhihui Lu 0002
CSCloud5
2024 FINSEC: An Efficient Microservices-Based Detection Framework for Financial AI Model Security
abstract
Artificial intelligence technology, such as fraud detection and biometrics, has recently been widely used in financial security. However, the related machine learning models may have algorithmic risks, and attackers can use loopholes in the models themselves to circumvent censorship or even steal private data. Our research team has designed and implemented a cloud service-based algorithmic risk detection platform for fintech products using advanced AI technologies to address this challenge. The platform can assess the algorithmic risk of machine learning models used for regulation in several common scenarios in the financial sector and provide early warnings of potential risk factors. Our platform, built upon cloud services, boasts high performance and embraces the principle of loose coupling. Our research aims to furnish the FinTech industry with a pragmatic tool for model algorithm risk detection.
Peng Chen 0030, Zhihui Lu 0002, Xiaozheng Du
CSCloud4
2024 Efficiency Meets Security: A Delegated Transport Authentication Protocol for the Cloud
abstract
This paper presents a novel authentication method for proxy transmission via an intermediate server, focusing on low overhead for cloud gaming and online services. It improves data speed by optimizing authentication and boosts security with an efficient encryption mechanism, reducing resource usage. Performance evaluations reveal lower latency and higher throughput than traditional methods, effectively countering network attacks. Solutions to address deployment challenges such as network complexity and hardware adaptability include adaptive network optimization and regular security audits. Future initiatives aim to integrate 5G, IoT, and edge computing technologies with AI-driven enhancements for security and responsiveness. Demonstrations in theory and practice have shown this invention superior, promising enhanced data transmission for service providers and users and offering significant market and socio-economic potential.
Zhihui Lu 0002, Yu Liu 0135, Lingyu Hou
CSCloud4
2024 Revisiting Learned Index with Byte-addressable Persistent Storage
abstract
Byte-addressable Persistent Storage (BPS), such as persistent memory and CXL-enabled SSDs, has become an extension of main memory. This opens up new possibilities for indexes that operate and persist data directly on the memory bus. Recent learned indexes exploit data distribution and have shown great potential for some workloads. Despite some work proposed for integrating learned indexes into BPS, they are mainly based on Intel’s first-generation persistent memory. The current design suffers from the following problems: 1) Excessive storage line accesses due to large node in learned indexes; 2) Inefficient concurrency control due to volatile cache; 3) Write amplification due to mismatch access granularity.
Rui Zhang 0112, Sicheng Liang, Shangyi Sun, Shaonan Ma, Chengying Huan, Lulu Chen, Zhihui Lu 0002, Yang Xu 0010, Ming Yan 0009, Jie Wu 0003
ICPP8
2024 TuneChain: An Online Configuration Auto-Tuning Approach for Permissioned Blockchain Systems
abstract
The increasing prevalence of blockchain technology has drawn significant attention to the need for effective Quality of Service (QoS) management in blockchain service provision. In this context, the online tuning of system configurations is pivotal for automatic blockchain services to meet QoS requirements. Past studies on configuration tuning have primarily focused on system adaptability to hardware and network environments, overlooking the dynamic nature of the highly diverse workloads, thus resulting in suboptimal system performance. This paper presents TuneChain, an online configuration auto-tuning approach for permissioned blockchain systems, which addresses the limitations of current methods, particularly in handling dynamic workloads while minimizing tuning costs. TuneChain leverages a Conflict Emergency Mechanism (CF-EM) to mitigate the impact of transaction conflicts on effective throughput and employs the Proximal Policy Optimization (PPO) algorithm coupled with a multi-instance mechanism to offer adaptive configuration recommendations tailored to diverse workloads. Additionally, TuneChain incorporates a Tuning Causal Model (TCModel) based on expert knowledge to guide decision-making in configuration tuning, thereby reducing unnecessary exploration and improving efficiency. Extensive evaluations demonstrate that TuneChain outperforms state-of-the-art approaches to configuration tuning in adapting to dynamic workloads, showcasing its efficacy in enhancing blockchain service performance.
Junxiong Lin, Ruijun Deng, Zhihui Lu 0002, Yiguang Zhang, Qiang Duan 0002
ICWS3
2024 Mitigating critical nodes in brain simulations via edge removal
Yubing Bao, Xin Du 0002, Zhihui Lu 0002, Jirui Yang, Shih-Chia Huang, Jianfeng Feng, Qibao Zheng
Comput. Networks3
2024 Universal adversarial backdoor attacks to fool vertical federated learning
Peng Chen 0030, Xin Du 0002, Zhihui Lu 0002, Hongfeng Chai
Comput. Secur.3
2024 CoLLaRS : A cloud-edge-terminal collaborative lifelong learning framework for AIoT
Shijing Hu 0001, Junxiong Lin, Zhihui Lu 0002, Xin Du 0002, Qiang Duan 0002, Shih-Chia Huang
Future Gener. Comput. Syst.3
2024 PBRL-TChain: A performance-enhanced permissioned blockchain for time-critical applications based on reinforcement learning
Yiguang Zhang, Junxiong Lin, Zhihui Lu 0002, Qiang Duan 0002, Shih-Chia Huang
Future Gener. Comput. Syst.3
2024 A balanced and reliable data replica placement scheme based on reinforcement learning in edge-cloud environments
Mengke Zheng, Xin Du 0002, Zhihui Lu 0002, Qiang Duan 0002
Future Gener. Comput. Syst.3
2024 Distributed DRL-Based Intelligent Over-the-Air Computation in Unmanned Aerial Vehicle Swarm-Assisted Intelligent Transportation System
abstract
Unmanned aerial vehicle (UAV)-based edge computing has been widely applied in intelligent transportation systems (ITSs) owing to its ease of deployment and high mobility. In this article, we study intelligent over-the-air computation (AirComp) in UAV swarm-assisted ITS. To develop a holistic service framework for UAV swarm, we consider the heterogeneity of Internet of Things Devices (IoTDs) and UAVs. We model the 3-D deployment of UAVs, service configuration, bandwidth allocation, the control of computing capacity, and transmission power as a joint optimization problem. To tackle this complex problem, we first propose a dual time-scale architecture based on deep reinforcement learning (DRL). This architecture enables UAVs to achieve seamless coverage of IoTDs on larger time scales, while collaborative UAVs dynamically provide services on smaller time scales. Next, we propose an intelligent AirComp algorithm D2IAC based on distributed DRL to obtain the optimal UAV deployment and dynamic service policies on different time scales. The D2IAC algorithm consists of three subalgorithms, i.e., TD3-based UAV deployment (TBUD), UAV services configuration (USC), and REINFORCE-based dynamic service (RBDS). Sufficient experimental results show that the proposed algorithm can achieve 3-D deployment of UAVs with coverage improvement from 9% to 36% compared to clustering, center layout, and random algorithms. Regarding dynamic services, compared with the deep deterministic policy gradient algorithm, greedy, fixed, and random strategies, the service durations of UAV swarm are improved by 32.95%–93.72% and the resource utilization is improved by 36.19%–49.61%.
Peng Hou 0003, Yi Huang 0020, Hongbin Zhu, Zhihui Lu 0002, Shih-Chia Huang, Yang Yang 0001, Hongfeng Chai
IEEE Internet Things J.4
2024 Towards transferable adversarial attacks on vision transformers for image classification
Xu Guo 0004, Peng Chen 0030, Zhihui Lu 0002, Hongfeng Chai, Xin Du 0002
J. Syst. Archit.3
2024 Multidomain Object Detection Framework Using Feature Domain Knowledge Distillation
abstract
Object detection techniques have been widely studied, utilized in various works, and have exhibited robust performance on images with sufficient luminance. However, these approaches typically struggle to extract valuable features from low-luminance images, which often exhibit blurriness and dim appearence, leading to detection failures. To overcome this issue, we introduce an innovative unsupervised feature domain knowledge distillation (KD) framework. The proposed framework enhances the generalization capability of neural networks across both low- and high-luminance domains without incurring additional computational costs during testing. This improvement is made possible through the integration of generative adversarial networks and our proposed unsupervised KD process. Furthermore, we introduce a region-based multiscale discriminator designed to discern feature domain discrepancies at the object level rather than from the global context. This bolsters the joint learning process of object detection and feature domain distillation tasks. Both qualitative and quantitative assessments shown that the proposed method, empowered by the region-based multiscale discriminator and the unsupervised feature domain distillation process, can effectively extract beneficial features from low-luminance images, outperforming other state-of-the-art approaches in both low- and sufficient-luminance domains.
Da-Wei Jaw, Shih-Chia Huang, Zhihui Lu 0002, Benjamin C. M. Fung, Sy-Yen Kuo
IEEE Trans. Cybern.3
2024 HRCM: A Hierarchical Regularizing Mechanism for Sparse and Imbalanced Communication in Whole Human Brain Simulations
abstract
Brain simulation is one of the most important measures to understand how information is represented and processed in the brain, which usually needs to be realized in supercomputers with a large number of interconnected graphical processing units (GPUs). For the whole human brain simulation, tens of thousands of GPUs are utilized to simulate tens of billions of neurons and tens of trillions of synapses for the living brain to reveal functional connectivity patterns. However, as an application of the irregular spares communication problem on a large-scale system, the sparse and imbalanced communication patterns of the human brain make it particularly challenging to design a communication system for supporting large-scale brain simulations. To face this challenge, this paper proposes a hierarchical regularized communication mechanism, HRCM. The HRCM maintains a hierarchical virtual communication topology (HVCT) with a merge-forward algorithm that exploits the sparsity of neuron interactions to regularize inter-process communications in brain simulations. HRCM also provides a neuron-level partition scheme for assigning neurons to simulation processes to balance the communication load while improving resource utilization. In HRCM, neuron partition is formulated as a k-way graph partition problem and solved efficiently by the proposed hybrid multi-constraint greedy (HMCG) algorithm. HRCM performs finer-grained neuron-level communication control while leveraging voxel-level control as the basis, thus being more effective in balancing inter-process traffic in large-scale simulations. The hierarchical characteristics of the finer-grained communication control are considered by the problem formulation and algorithm design in HRCM. HRCM has been implemented in human brain simulations at the scale of up to 86 billion neurons running on 10000 GPUs. Results obtained from extensive simulation experiments verify the effectiveness of HRCM in significantly reducing communication delay, increasing resource usage, and shortening simulation time for large-scale human brain models.
Xin Du 0002, Minglong Wang, Zhihui Lu 0002, Qiang Duan 0002, Yuhao Liu 0008, Jianfeng Feng, Huarui Wang
IEEE Trans. Parallel Distributed Syst.3
2023 A Practical Clean-Label Backdoor Attack with Limited Information in Vertical Federated Learning
abstract
Vertical Federated Learning (VFL) facilitates collaboration on model training among multiple parties, each owning partitioned features of the distributed dataset. Although backdoor attacks have been found as one of the main threats to FL security, research on backdoor attacks in VFL is still in the infant stage. Existing methods for VFL backdoor attacks rely on predicting sample pseudo-labels using approaches such as label inference, which require substantial additional information not readily available in practical FL scenarios. To evaluate the practical vulnerability of VFL to backdoor attacks, we present a target-efficient clean backdoor (TECB) attack for VFL. The TECB approach consists of two phases – i) Clean Backdoor Poisoning (CBP) and Target Gradient Alignment (TGA). In the CBP phase, the adversary trains a backdoor trigger and poisons the model during VFL training. The poisoned model is further fine-tuned in the TGA phase to enhance its efficacy in complex multi-classification tasks. Compared to the existing methods, the proposed TECB achieves a highly effective backdoor attack with very limited information about the target class samples, which is more practical in typical VFL settings. Experimental results verify the superior performance of TECB, achieving above 97% attack success rate (ASR) on three widely used datasets (CIFAR10, CIFAR100, and CINIC-10) with only 0.1% of target labels known, which outperforms the state-of-the-art attack methods. This study uncovers the potential backdoor risks in VFL, enabling the development of secure VFL applications in areas like finance, healthcare, and beyond. Source code is available at: https://github.com/13thDayOLunarMay/TECB-attack
Peng Chen 0030, Jirui Yang, Junxiong Lin, Zhihui Lu 0002, Qiang Duan 0002, Hongfeng Chai
ICDM4
2023 HSFL: Efficient and Privacy-Preserving Offloading for Split and Federated Learning in IoT Services
abstract
Distributed machine learning methods like Federated Learning (FL) and Split Learning (SL) meet the growing demands of processing large-scale datasets under privacy restrictions. Recently, FL and SL are combined in hybrid SLFL (SFL) frameworks to exploit both methods’ advantages to facilitate ubiquitous intelligence in the Internet of Things (IoT), for example, smart finance. Despite its significant impact on the performance and costs of SFL, model decomposition that splits an ML model into the client-server pair has not been sufficiently studied, especially for SFL in a large-scale dynamic IoT environment. In this paper, we propose a new SFL framework HSFL with a lightweight model decomposition method to offload a part of model training to the edge server. Specifically, we develop a method for estimating the training latency of HSFL and designed a metric for measuring privacy leakage in HSFL, based on which we formulate model decomposition in HSFL as an optimization problem with privacy protection as a constraint. Then, we transform the formulated problem into a contextual bandit problem and design an efficient algorithm to solve it. We have conducted thorough evaluations of the proposed HSFL framework through extensive experiments on a prototype testbed and a simulation platform. The experimental results validate the superiority of HSFL over the state-of-the-art benchmarks in terms of training latency, efficiency, scalability, and privacy protection.
Ruijun Deng, Xin Du 0002, Zhihui Lu 0002, Qiang Duan 0002, Shih-Chia Huang, Jie Wu 0003
ICWS3
2023 Joint computation offloading and resource allocation based on deep reinforcement learning in C-V2X edge computing
Peng Hou 0003, Zhihui Lu 0002, Bo Li 0025, Zongshan Wang
Appl. Intell.3
2023 Fidan: a predictive service demand model for assisting nursing home health-care robots
abstract
While population aging has sharply increased the demand for nursing staff, it has also increased the workload of nursing staff.Although some nursing homes use robots to perform part of the work, such robots are the type of robots that perform set tasks.The requirements in actual application scenarios often change, so robots that perform set tasks cannot effectively reduce the workload of nursing staff.In order to provide practical help to nursing staff in nursing homes, we innovatively combine the LightGBM algorithm with the machine learning interpretation framework SHAP (Shapley Additive exPlanations) and use comprehensive data analysis methods to propose a service demand prediction model Fidan (Forecast service demand model).This model analyzes and predicts the demand for elderly services in nursing homes based on relevant health management data (including physiological and sleep data), ward round data, and nursing service data collected by IoT devices.We optimise the model parameters based on Grid Search during the training process.The experimental results show that the Fidan model has an accuracy rate of 86.61% in predicting the demand for elderly services.
Feng Zhou 0014, Xin Du 0002, Zhihui Lu 0002, Shih-Chia Huang
Connect. Sci.4
2023 A Blockchain-Assisted Intelligent Edge Cooperation System for IoT Environments With Multi-Infrastructure Providers
abstract
While edge computing has the potential to offer low-latency services and overcome the limitations of traditional cloud computing, it presents new challenges in terms of trust, security, and privacy (TSP) in Internet of Things environments. Cooperative edge computing (CEC) has emerged as a solution to address these challenges through resource sharing among edge nodes. However, for multi-infrastructure providers, incentive and trust mechanisms among edge nodes are crucial technical issues that must be addressed alongside system latency and reliability to meet performance requirements. In this article, we propose a blockchain-assisted intelligent edge cooperation system (BIECS) to systematically solve these issues. By leveraging blockchain technology, we construct trust among edge nodes and employ an incentive mechanism for resource sharing among multi-infrastructure providers. We formulate the system performance optimization as a multiobjective joint optimization problem and solve it efficiently through a two-stage strategy for selecting edge nodes. We first design an improved long short term memory (LSTM) model for resource prediction and then select edge nodes for executing offloaded tasks and handling the corresponding blockchain process related to each task execution. To evaluate the performance of BIECS, we implement the system based on Hyperledger Fabric and design extensive experiments. Our proposed system achieves better performance in terms of system delay, throughput, and resource utilization compared to state-of-the-art schemes for edge cooperation.
Xin Du 0002, Xuzhao Chen, Zhihui Lu 0002, Qiang Duan 0002, Jie Wu 0003, Patrick C. K. Hung
IEEE Internet Things J.3
2023 Federated Deep Reinforcement Learning-Based Intelligent Dynamic Services in UAV-Assisted MEC
abstract
Unmanned aerial vehicles (UAVs)-assisted multiaccess edge computing (MEC) has emerged as a promising solution in B5G/6G networks. The high flexibility and seamless connectivity of UAVs make them well suited for providing enhanced communications coverage and efficient computing support. Particularly, in situations where ground facilities may be compromised or communication is unreliable. In this article, we study joint dynamic service switching and resource allocation for multiple UAVs in MEC network. We consider the heterogeneity of tasks and UAVs and model the dynamic service process of UAVs as a sequential decision problem based on the Markovian decision process. To enable dynamic and intelligent UAV service, we first propose a centralized dynamic service algorithm DDPG-based centralized (DDBC) based on deep reinforcement learning. However, given the training difficulties of the centralized algorithm, we propose a more promising distributed learning algorithm FLBF, which combines federated learning. We conduct extensive simulations to evaluate the effectiveness and advantages of the proposed algorithms. Our results show that DDBC and FLBF significantly reduce the system cost by 17.99%–35.72% and 12.30%–31.26%, respectively, compared to the comparative algorithms. Furthermore, FLBF can effectively improve the convergence speed with guaranteed learning performance, indicating its suitability for model training in UAV-assisted MEC networks.
Peng Hou 0003, Zongshan Wang, Sen Liu 0002, Zhihui Lu 0002
IEEE Internet Things J.5
2023 AsyFed: Accelerated Federated Learning With Asynchronous Communication Mechanism
abstract
As a new distributed machine learning (ML) framework for privacy protection, federated learning (FL) enables substantial Internet of Things (IoT) devices (e.g., mobile phones, tablets, etc.) to participate in collaborative training of an ML model. FL can protect the data privacy of IoT devices without exposing their raw data. However, the diversity of IoT devices may degrade the overall training process due to the straggler issue. To tackle this problem, we propose a gear-based asynchronous FL (AsyFed) architecture. It adds a gear layer between the clients and the FL server as a mediator to store the model parameters. The key insight is that we group these clients with similar training abilities into the same gear. The clients within the same gear conduct synchronous training. These gears then communicate with the global FL server asynchronously. Besides, we propose a T-step mechanism to reduce the weight from the slow gear when they are communicating with the FL server. The extensive experiment evaluations indicate that AsyFed outperforms FedAvg (baseline synchronous FL scheme) and some state-of-the-art asynchronous FL methods in terms of training accuracy or speed under different data distributions. The only negligible overhead is that we leverage the extra layer (gear layer) to preserve part of the model parameters.
Zhixin Li 0003, Chunpu Huang, Keke Gai, Zhihui Lu 0002, Jie Wu 0003, Lulu Chen, Yangchuan Xu, Kim-Kwang Raymond Choo
IEEE Internet Things J.4
2023 An Adaptive Mechanism for Dynamically Collaborative Computing Power and Task Scheduling in Edge Environment
abstract
Edge computing can provide high bandwidth and low-latency service for big data tasks by leveraging the edge side’s computing, storage, and network resources. With the development of microservice and docker technology, service providers can flexibly and dynamically cache microservice at the edge side to respond efficiently with limited resources. Automatically caching needed services on the nearest edge nodes and dynamically scheduling users’ requests can realize that computing power and software services flow with the users to provide continuous services. However, achieving the goal needs to overcome many challenges, such as the significant fluctuation of user devices’ requests at the edge side and the lack of collaboration among edge nodes. In this article, dynamic computing power scheduling and collaborative task scheduling among edge nodes are comprehensively developed. The problem is considered a multiobjective optimization problem, including sequentially minimizing the deadline missing rate of requests and the average task completion time. We propose an adaptive mechanism for dynamically collaborative computing power and task scheduling (ADCS) in the edge environment to solve this problem. It adopts the greedy decision method to schedule computing tasks to meet their deadline requirements. At the same time, it uses the best-fit method to adjust the computing resources according to the changes of users’ requests. The simulation results show that ADCS can decrease the deadline missing rate and reduce the average completion time. Compared with DSR and CoDSR, the deadline missing rate is reduced by 59.91% and 19.95%, respectively. The average completion time is decreased by 37.87% and 6.71%.
Yangchuan Xu, Lulu Chen, Zhihui Lu 0002, Xin Du 0002, Jie Wu 0003, Patrick C. K. Hung
IEEE Internet Things J.3
2023 BESIFL: Blockchain-Empowered Secure and Incentive Federated Learning Paradigm in IoT
abstract
Federated learning (FL) offers a promising approach to efficient machine learning with privacy protection in distributed environments, such as Internet of Things (IoT) and mobile-edge computing (MEC). The effectiveness of FL relies on a group of participant nodes that contribute their data and computing capacities to the collaborative training of a global model. Therefore, preventing malicious nodes from adversely affecting the model training while incentivizing credible nodes to contribute to the learning process plays a crucial role in enhancing FL security and performance. Seeking to contribute to the literature, we propose a blockchain-empowered secure and incentive FL (BESIFL) paradigm in this article. Specifically, BESIFL leverages blockchain to achieve a fully decentralized FL system, where effective mechanisms for malicious node detections and incentive management are fully integrated in a unified framework. The experimental results show that the proposed BESIFL is effective in improving FL performance through its protection against malicious nodes, incentive management, and selection of credible nodes.
Zhihui Lu 0002, Keke Gai, Qiang Duan 0002, Junxiong Lin, Jie Wu 0003, Kim-Kwang Raymond Choo
IEEE Internet Things J.2
2022 Regularizing Sparse and Imbalanced Communications for Voxel-based Brain Simulations on Supercomputers
abstract
Inter-process communications form a performance bottleneck for large-scale brain simulations. The sparse and imbalanced communication patterns of human brain make it particularly challenging to design a communication system for supporting large-scale brain simulations. In this paper, we tackle the communication challenges posed by large-scale brain simulations with sparse and imbalanced communication patterns. We design a virtual communication topology with a merge and forward algorithm that exploits the sparsity to regularize inter-process communications. To balance the communication loads of different processes, we formulate voxel partition in brain simulations as a k-way graph partition problem and propose a constrained deterministic greedy algorithm to solve the problem effectively. We conducted extensive simulation experiments for evaluating the performance of the proposed communication scheme and found that the proposed method may significantly reduce communication overheads and shorten simulation time for large-scale brain models.
Yuhao Liu 0008, Xin Du 0002, Zhihui Lu 0002, Qiang Duan 0002, Jianfeng Feng, Minglong Wang, Jie Wu 0003
ICPP3
2022 BIECS: A Blockchain-based Intelligent Edge Cooperation System for Latency-Sensitive Services
abstract
Although the emerging edge computing paradigm offers a promising approach to overcoming some limitations of conventional cloud computing, the heterogeneous edge nodes with highly diverse system capacities bring new challenges to service provisioning especially for latency-sensitive services. Cooperative edge computing (CEC) has been proposed for facing such challenges through resource sharing among edge nodes. However, some technical issues must be fully addressed to make CEC effective, among which incentive and trust mechanisms and performance optimization are crucial for latency-sensitive service provision. In this paper, we design a novel blockchain-based intelligent edge cooperation system named BIECS to tackle these challenges systematically. BIECS provides incentive to edge nodes for resource sharing and enables trust among cooperative nodes upon a distributed platform leveraging the blockchain technology. In order to optimize system performance for meeting the requirements of latency-sensitive services, we propose a two-stage strategy for node selection in BIECS that chooses the most appropriate edge nodes for executing offloaded tasks and recording related transactions in the blockchain. We also implemented a prototype of BIECS based on Hyperledger Fabric and conducted extensive experiments for evaluating the performance of BIECS. The obtained experiment results verify that the proposed BIECS achieves better performance in system delay and throughput compared to the state-of-the-art methods for edge cooperation.
Xin Du 0002, Xuzhao Chen, Zhihui Lu 0002, Qiang Duan 0002, Jie Wu 0003
ICWS3
2022 Coordinate-based efficient indexing mechanism for intelligent IoT systems in heterogeneous edge computing
Songtao Tang, Xin Du 0002, Zhihui Lu 0002, Keke Gai, Jie Wu 0003, Patrick C. K. Hung, Kim-Kwang Raymond Choo
J. Parallel Distributed Comput.3
2022 EVFL: An explainable vertical federated learning for data-oriented Artificial Intelligence systems
Peng Chen 0030, Xin Du 0002, Zhihui Lu 0002, Jie Wu 0003, Patrick C. K. Hung
J. Syst. Archit.3
2022 A Resource Recommendation Model for Heterogeneous Workloads in Fog-Based Smart Factory Environment
abstract
The wide deployment of advanced robots with industrial IoT (IIoT) technologies in smart factories generates a large volume of data during production and a wide variety of data processing workloads are launched to maintain productivity and safety of smart manufacture. The emerging fog computing paradigm offers a promising solution to enhancing data processing performance in a smart factory environment while on the other hand brings in new challenges to resource management, which call for a more effective approach for recommending resource configurations to heterogeneous workloads. In this paper, we propose an Optimized Recommendations of Heterogeneous Resource Configurations (ORHRC) model that employs machine learning techniques to provide resource configuration recommendations for the heterogeneous workloads in a fog computing-based smart factory environment. ORHRC learns a recommendation model by leveraging the operating characteristics and execution time of workloads on fog servers with different configurations. We also design a decision model in ORHRC to further improve prediction accuracy and reduce operational overheads. Experiment results show that ORHRC outperforms the state of art configuration recommendation methods in terms of average prediction accuracy.Note to Practitioners—The various data processing workloads in a smart factory environment need to be processed by the computational resources with optimal configurations for meeting their performance requirements. In this paper, we employ machine learning technologies for enabling automatic recommendation of resource configurations to heterogeneous workloads. Specifically, we develop an Optimized Recommendations of Heterogeneous Resource Configurations (ORHRC) model that can identify the optimal resource configurations for various workloads. We also conducted extensive experiments that verify the effectiveness of the proposed ORHRC model.
Lulu Chen, Zhihui Lu 0002, Ai Xiao, Qiang Duan 0002, Jie Wu 0003, Patrick C. K. Hung
IEEE Trans Autom. Sci. Eng.2
2022 Improved LSTM-Based Time-Series Anomaly Detection in Rail Transit Operation Environments
abstract
Anomaly detection is crucial to the reliability and safety of rail transit systems. The rapid development of Internet of Things (IoT) and cloud technologies together with recent advances in machine learning offered various cloud-based data-driven approaches to automatic anomaly detection. However, the challenges introduced by the different types of equipment in rail transit systems with highly diverse data distributions and the lack of labeled anomaly data have not been sufficiently addressed. In this article, we attempt to cope with such challenges by proposing an improved long short term memory (LSTM)-based time-series anomaly detection scheme. The key elements of the proposed scheme include an improved LSTM model that may achieve more accurate time-series prediction for various rail transit devices and a method for determining an appropriate error threshold for detecting anomalies based on the prediction errors. In order to further enhance anomaly detection performance, we also propose a pruning algorithm for reducing the number of false anomalies. Our method does not rely on scarce anomaly labels but dynamically determines a threshold of prediction errors to identify anomalies; therefore, it overcomes the challenge of the extremely uneven distribution of rail transit data. We conducted extensive experiments in a real metro operation environment for performance evaluation. The experiment results prove the effectiveness of the proposed scheme and show a superior performance of the scheme compared to existing anomaly detection methods.
Xin Du 0002, Zhihui Lu 0002, Qiang Duan 0002, Jie Wu 0003
IEEE Trans. Ind. Informatics3
2021 A blockchain-based evidential and secure bulk-commodity supervisory system
abstract
In recent years, the commodities industry has grown rapidly under the stimulus of domestic demand and the expansion of cross-border trade. It has also been combined with the rapid development of e-commerce technology in the same period to form a flexible and efficient e-commerce system for bulk commodities. However, the hasty combination of both has inspired a lack of effective regulatory measures in the bulk industry, leading to constant industry chaos. Among them, the problem of lagging evidence in regulatory platforms is particularly prominent. Based on this, we design a blockchain-based evidential and secure bulk-commodity supervisory system (abbr. BeBus). Setting different privacy protection policies for each participant in the system, the solution ensures effective forensics and tamper-proof evidence to meet the needs of the bulk business scenario.
Junxiong Lin, Zhihui Lu 0002, Jie Wu 0003, Houhao Ye, Wenbing Huang 0004, Xuzhao Chen
ICSS3
2021 IoT Microservice Deployment in Edge-Cloud Hybrid Environment Using Reinforcement Learning
abstract
The edge-cloud hybrid environment requires complex deployment strategies to enable the smart Internet-of-Things (IoT) system. However, current service deployment strategies use simple, generalized heuristics and ignore the heterogeneous characteristics in the edge-cloud hybrid environment. In this article, we devise a method to find a microservice-based service deployment strategy that can reduce the average waiting time of IoT devices in the hybrid environment. For this purpose, we first propose a microservice-based deployment problem (MSDP) based on the heterogeneous and dynamic characteristics in the edge-cloud hybrid environment, including heterogeneity of edge server capacities, dynamic geographical information of IoT devices, and changing device preference for applications and complex application structures. We then propose a multiple buffer deep deterministic policy gradient (MB_DDPG) to provide more preferable service deployment solutions. Our algorithm leverages reinforcement learning and neural network to learn a deployment strategy without any human instruction. Therefore, the service provider can make full use of limited resources to improve the Quality of Service (QoS). Finally, we implement MB_DDPG based on real-world data sets and some synthetic data, and we also implement another two algorithms, genetic algorithm and random algorithm, as a contrast. The experimental results demonstrate that MB_DDPG is able to learn a preferable strategy which, in terms of average waiting time, outperforms genetic algorithm and the random algorithm by 32% and 44%, respectively.
Lulu Chen, Yangchuan Xu, Zhihui Lu 0002, Jie Wu 0003, Keke Gai, Patrick C. K. Hung, Meikang Qiu
IEEE Internet Things J.3
2020 A Novel Data Placement Strategy for Data-Sharing Scientific Workflows in Heterogeneous Edge-Cloud Computing Environments
abstract
The deployment of datasets in the heterogeneous edge-cloud computing paradigm has received increasing attention in state-of-the-art research. However, due to their large sizes and the existence of private scientific datasets, finding an optimal data placement strategy that can minimize data transmission as well as improve performance, remains a persistent problem. In this study, the advantages of both edge and cloud computing are combined to construct a data placement model that works for multiple scientific workflows. Apparently, the most difficult research challenge is to provide a data placement strategy to consider shared datasets, both within individual and among multiple workflows, across various geographically distributed environments. According to the constructed model, not only the storage capacity of edge micro-datacenters, but also the data transfer between multiple clouds across regions must be considered. To address this issue, we considered the characteristics of this model and identified the factors that are causing the transmission delay. The authors propose using a discrete particle swarm optimization algorithm with differential evolution (DE-DPSO) to distribute dataset during workflow execution. Based on this, a new data placement strategy named DE-DPSO-DPS is proposed. DE-DPSO-DPS is evaluated using several experiments designed in simulated heterogeneous edge-cloud computing environments. The results demonstrate that our data placement strategy can effectively reduce the data transmission time and achieve superior performance as compared to traditional strategies for data-sharing scientific workflows.
Xin Du 0002, Songtao Tang, Zhihui Lu 0002, Jie Wu 0003, Keke Gai, Patrick C. K. Hung
ICWS3
2020 ORHRC: Optimized Recommendations of Heterogeneous Resource Configurations in Cloud-Fog Orchestrated Computing Environments
abstract
The cloud-fog orchestrated computing environments devolve computing tasks from the cloud center to the fog nodes, providing more heterogeneous configurations for the operation of workloads. Compared to the conventional cloud computing environment, the physical conditions at the fog nodes in the cloud-fog orchestrated computing environments are more complex and changeable. Therefore, the configurations that the fog nodes provide are heterogeneous and varying. This requires the configuration selection model to adapt to changeable configurations. The previous configuration selection models are applied to the limited and fixed configurations in the conventional cloud environment, but not to the complex cloud-fog orchestrated computing environments. To address this problem, we propose Optimized Recommendations of Heterogeneous Resource Configurations(ORHRC), a model that provides users with a reliable cloud configuration recommendation service. ORHRC uses the matrix factorization algorithm and neural network to build a recommendation model, which combines the operating characteristics of workloads as the explicit ratings and implicit feedback, to give configuration recommendations. Comprehensive experiments on a real-world dataset demonstrate that the hit rate of configurations of ORHRC is 24% higher than Micky and 15% higher than Selecta.
Ai Xiao, Zhihui Lu 0002, Xin Du 0002, Jie Wu 0003, Patrick C. K. Hung
ICWS2
2020 Distributed gas concentration prediction with intelligent edge devices in coal mine
Yiwen Zhang 0001, Haishuai Guo, Zhihui Lu 0002, Lu Zhan, Patrick C. K. Hung
Eng. Appl. Artif. Intell.3
2020 Swarm Decision Table and Ensemble Search Methods in Fog Computing Environment: Case of Day-Ahead Prediction of Building Energy Demands Using IoT Sensors
abstract
Building energy demand prediction (BEDP) concerns sensing the environment using the Internet of Things (IoT), making seamless decisions and responding and controlling certain devices automatically, intelligently, and quickly. Typically, the BEDP application can be empowered by fog computing where the sensed data are processed at the edge nodes rather than in a central cloud. The challenge is that in this decentralized IoT environment, the machine learning algorithm implemented at the fog node must learn a model from the incoming data accurately and fast. Which type of incremental learning algorithms, combined with traditional or swarm types of stochastic feature selection methods, are more suitable for BEDP? In this article, this topic is investigated in detail by introducing a new incremental learning model, the swarm decision table (SDT) in comparison with the classical decision tree. The simulation experiments using an empirical energy consumption data set that represent a typical IoT-connected BEDP scenario are tested, and the SDT shows superior results in terms of accuracy and time, demonstrating it as a suitable machine learning candidate in a fog computing environment.
Tengyue Li, Simon Fong 0001, Xuqi Li, Zhihui Lu 0002, Amir Hossein Gandomi
IEEE Internet Things J.4
2020 BPS: A reliable and efficient pub/sub communication model with blockchain-enhanced paradigm in multi-tenant edge cloud
Yibo Huang 0005, Rui Zhang 0112, Zhihui Lu 0002, Yiming Zhang 0018, Jie Wu 0003, Lu Zhan, Patrick C. K. Hung
J. Parallel Distributed Comput.3
2020 ARVMEC: Adaptive Recommendation of Virtual Machines for IoT in Edge-Cloud Environment
Junnan Li 0003, Zhihui Lu 0002, Jie Wu 0003, Patrick C. K. Hung, Abdulhameed Alelaiwi
J. Parallel Distributed Comput.3
2020 Image Haze Removal Using Airlight White Correction, Local Light Filter, and Aerial Perspective Prior
abstract
Light is scattered and absorbed when travelling through atmosphere particles, leading to visibility attenuation for images captured, especially in hazy scenes. In addition, hazy images may suffer from color distortion caused by haze or sandstorm, resulting in a poor visual quality. In order to effectively enhance visibility and correct possible color casts for such images, we propose a new image dehazing algorithm based on an improved haze optical model, which consists of three modules: airlight white correction (AWC), local light filter (LLF), and aerial perspective prior (APP). In the proposed algorithm, the AWC module detects and corrects possible color cast, the LLF module downplays non-hazy bright pixels (e.g., headlight and white objects) for more accurate airlight estimation, and the APP module uses the minimum/maximum channel and their difference for scene transmission estimation. The experimental results demonstrate that the proposed method outperforms other state-of-the-art dehazing methods in three ways: 1) our results have better visual quality; 2) our method performs the best in terms of color restoration; and 3) our method is very efficient at removing haze and color casts.
Yan-Tsung Peng, Zhihui Lu 0002, Fan-Chieh Cheng, Yalun Zheng, Shih-Chia Huang
IEEE Trans. Circuits Syst. Video Technol.2
2020 BoR: Toward High-Performance Permissioned Blockchain in RDMA-Enabled Network
abstract
Known as a distributed ledger, blockchain is becoming prevalent due to its decentralization, traceability and tamper resistance. Particularly, permissioned blockchain such as Hyperledger Fabric shows great application prospects as the infrastructure of IoT security, credit management, etc. Many cloud platforms like AWS, Azure, Oracle and IBM cloud currently provide blockchain as a service, in which tenants can quickly build permissioned blockchain and run smart contract based applications. However, the transactions throughput and scalability in the permissioned blockchain are not ideal, despite many optimization efforts in consensus protocol and parallel chain. Existing solutions still reveals some limitations like excessive CPU scheduling, inefficient block broadcast and high latency of initial blocks synchronization when new nodes join blockchain network. Inspired by the emerging RDMA (Remote Direct Memory Access) network, we propose BoR, an RDMA-based permissioned blockchain framework. By offloading the block transfer transaction into RDMA NICs, it can increase block broadcast speed and reduce block sync delay. We exploit the RDMA primitives to redesign the block synchronization protocol and accelerate DPoS (Delegated Proof of Stake) consensus process for higher throughput and lower latency in kernel-bypass manner. As demonstrated in our evaluation with different workloads, BoR with lower CPU utilization significantly outperforms the state-of-the-art EoS blockchain.
Yibo Huang 0005, Zhihui Lu 0002, Xin Zhou 0009, Jie Wu 0003, Qifeng Tang, Patrick C. K. Hung
IEEE Trans. Serv. Comput.3
2019 Children Privacy Identification System in LINE Chatbot for Smart Toys
abstract
Children's privacy concerns about smart toys are becoming more and more critical in the toy industry. Parents and guardians continue to strive to protect their children from unnecessary privacy risks such as collection, and unconsented use of or access to their children's information. However, there is still no standardized privacy framework, which focuses on smart toys in this paradigm; making it difficult to determine possible privacy violation in for example determining whether a phrase shared with a smart toy is sensitive or not. To overcome this challenge, we build a privacy identification system through Chatbot technology. We call this system a Children Privacy Identification (CPI) system. To develop CPI system, we divide our research works into two parts: (1) Collect the phrase from the smart toys; and (2) Explore privacy Identification based on Personally Identifiable Information (PII) and Children's Online Privacy Protection Act (COPPA). For illustration, we integrate the CPI system in LINE Chatbot. The result shows that people feel more comfortable in talking to LINE Chatbot with privacy protection.
Pei-Chun Lin, Benjamin Yankson, Zhihui Lu 0002, Patrick C. K. Hung
CLOUD3
2019 Data Privacy Protection for Edge Computing of Smart City in a DIKW Architecture
Yucong Duan, Zhihui Lu 0002, Zhangbing Zhou, Xiaobing Sun 0001, Jie Wu 0003
Eng. Appl. Artif. Intell.2
2019 A self-adaptive approach to service deployment under mobile edge computing for autonomous driving
Zhihui Lu 0002, Bing Li 0010, Bo Hang, Jie Wu 0003, Xiaohua Xuan
Eng. Appl. Artif. Intell.2
2019 A general AI-defined attention network for predicting CDN performance
Junnan Li 0003, Zhihui Lu 0002, Jie Wu 0003, Shalin Huang, Meikang Qiu
Future Gener. Comput. Syst.2
2019 QaMeC: A QoS-driven IoVs application optimizing deployment scheme in multimedia edge clouds
Zhihui Lu 0002, Patrick C. K. Hung, Shih-Chia Huang, Zhenfang Wang
Future Gener. Comput. Syst.2
2019 Bigdata logs analysis based on seq2seq networks for cognitive Internet of Things
Pin Wu, Zhihui Lu 0002, Zhidan Lei, Xiaoqiang Li 0002, Meikang Qiu, Patrick C. K. Hung
Future Gener. Comput. Syst.2
2019 RDMA-driven MongoDB: An approach of RDMA enhanced NoSQL paradigm for large-Scale data processing
Yibo Huang 0005, Zhihui Lu 0002, Ming Yan 0009, Jie Wu 0003, Patrick C. K. Hung, Qifeng Tang
Inf. Sci.3
2019 Efficiently querying large process model repositories in smart city cloud workflow systems based on quantitative ordering relations
Hua Huang 0006, Zhihui Lu 0002, Rong Peng, Zaiwen Feng, Xiaohua Xuan, Patrick C. K. Hung, Shih-Chia Huang
Inf. Sci.2
2019 All-Or-Nothing data protection for ubiquitous communication: Challenges and perspectives
Han Qiu 0001, Katarzyna Kapusta, Zhihui Lu 0002, Meikang Qiu, Gérard Memmi
Inf. Sci.3
2018 SERAC3: Smart and economical resource allocation for big data clusters in community clouds
Junnan Li 0003, Zhihui Lu 0002, Wei Zhang 0085, Jie Wu 0003, Bo Li 0025, Patrick C. K. Hung
Future Gener. Comput. Syst.2
2018 Automating smart recommendation from natural language API descriptions via representation learning
Zhihui Lu 0002, Bing Li 0010, Bo Hang
Future Gener. Comput. Syst.2
2018 IoTDeM: An IoT Big Data-oriented MapReduce performance prediction extended model in multiple edge clouds
Zhihui Lu 0002, Nini Wang, Jie Wu 0003, Meikang Qiu
J. Parallel Distributed Comput.1
2018 A data-driven approach of performance evaluation for cache server groups in content delivery network
Zhihui Lu 0002, Wei Zhang 0085, Jie Wu 0003, Shalin Huang, Patrick C. K. Hung
J. Parallel Distributed Comput.2
2018 Smart-toy-edge-computing-oriented data exchange based on blockchain
Zhihui Lu 0002, Jie Wu 0003
J. Syst. Archit.2
2018 Toy-IoT-Oriented data-driven CDN performance evaluation model with deep learning
Wei Zhang 0085, Zhihui Lu 0002, Jie Wu 0003, Huanying Zou, Shalin Huang
J. Syst. Archit.2
2017 LTSS: Load-Adaptive Traffic Steering and Forwarding for Security Services in Multi-Tenant Cloud Datacenters
Xuekai Du, Zhihui Lu 0002, Qiang Duan 0002, Jie Wu 0003, Chengrong Wu
J. Comput. Sci. Technol.2
2017 InSTechAH: Cost-effectively autoscaling smart computing hadoop cluster in private cloud
Zhihui Lu 0002, Jie Wu 0003, Patrick C. K. Hung
J. Syst. Archit.1
2017 Multi-policy-aware MapReduce resource allocation and scheduling for smart computing cluster
Zhihui Lu 0002, Nini Wang, Jie Wu 0003, Patrick C. K. Hung
J. Syst. Archit.2
2016 Analysis of Big Data Platform with OpenStack and Hadoop
Zhihui Lu 0002, Nini Wang, Jie Wu 0003, Shalin Huang
APSCC2
2016 Comparison and Improvement of Hadoop MapReduce Performance Prediction Models in the Private Cloud
Nini Wang, Zhihui Lu 0002, Jie Wu 0003
APSCC3
2015 A Novel Reactive-Predictive Hybrid Resource Provision Method in Cloud Datacenter
Guorui Sun, Zhihui Lu 0002, Jie Wu 0003, Patrick C. K. Hung
APSCC2
2015 CPFirewall: A Novel Parallel Firewall Scheme for FWaaS in the Cloud Environment
Zhenfang Wang, Zhihui Lu 0002, Jie Wu 0003, Kang Fan
APSCC2
2014 Implementing a novel load-aware auto scale scheme for private cloud resource management platform
abstract
Resources dynamical allocation and management is always an important feature in cloud computing. Auto Scale allows users to scale their cloud resources capacity according to elastic loads timely, which has been widely used in mature public cloud. For private cloud, there are some different features from public cloud. It is more flexible to use Auto Scale technique to provide QoS guarantees and ensure system health. In this paper, we design a novel Auto Load-aware Scale scheme for private cloud environment. We describe scale in and scale out strategy based on prediction algorithm. We implement our scheme on OpenStack platform. Both simulation and experiments are carried out to evaluate our work. The experiments show that our scheme has better performance in resource utilization while providing high SLA levels.
Zhihui Lu 0002, Jie Wu 0003, Shiyong Zhang, YiPing Zhong
NOMS2
2013 PROSE: Proactive, Selective CDN Participation for P2P Streaming
Zhihui Lu 0002, Lijiang Chen, Jie Wu 0003, Da Deng 0002, Sijia Huang, Yi Huang 0020
J. Comput. Sci. Technol.1
2012 CPDID: A Novel CDN-P2P Dynamic Interactive Delivery Scheme for Live Streaming
abstract
Although many streaming application providers have relied on CDN services, there are several barriers to making CDN a more common service: expensive construction cost and fixed service mode. The rise of cloud computing requires "CDN as a Service" with a more open service mode, which requires content services available on-demand and be utilized in an open and loosely-coupled fashion. In most of the current streaming systems, CDN provides serve requesters with a passive and static state, there is no dynamic interaction with P2P systems, so the total CDN-P2P-Hybrid efficiency is not high. In this paper, we present CPDID: a CDN-P2P Dynamic Interactive Delivery scheme for Live Streaming. We explore CPDID architecture based on REST interface and JSON message format. And then we propose an identifying and selecting 'upload amplification nodes' algorithm to more efficiently utilize CDN resource. Our experimental results show that CPDID achieves at least 10-25% performance improvement compared with the existing native CDN-P2P-Hybrid schemes. At last, we analyze the prospective research direction and propose our future work.
Zhihui Lu 0002, Jie Wu 0003, Yi Huang 0020, Lijiang Chen, Da Deng 0002
ICPADS1
2012 WS-RM: Resource Management Using Web Service
abstract
We have been participated in the early research work of using Web service for system management and its standardization work. In this paper, we first review the concept of using Web service for management and compare two main industry standardization efforts - WSDM and WS-Management. We present WS-RM, our resource management model, its implementation, the testbed and management scenarios. Then, we present several selected gap items that are identified between realworld management requirements and current standard as well as our proposed solutions. In the end, we discuss some future work into cloud computing management and using Web service for management integration.
Zhihui Lu 0002
SERVICES2
2011 Scalable and Reliable Live Streaming Service through Coordinating CDN and P2P
abstract
In order to fully utilize CDN network edge nearby transmission capability and P2P scalable end-to-end transmission capability, while overcoming CDN's limited service capacity and P2P's dynamics, we need to combine the CDN and P2P technologies together. In recent years, some researchers have begun to focus on CDN-P2P-hybrid architecture and ISP-friendly P2P content delivery technology. In this paper, we firstly make an analysis on main problems of CDN-P2P-hybrid technology, and we compare some current existing solutions. And then we propose a novel scalable and reliable live streaming service scheme through coordinating CDN and P2P. In this scheme, different overlay hybrid methods, we design a smoother CDN and P2P overlay hybrid approach. After that we design a new peer node buffer structure for consuming media block from both CDN and P2P sources. And then we propose a novel algorithm for identifying and serving super node. The simulation experiment results show that our CDN-P2P-hybrid based live streaming scheme has better performance and reliability than pure P2P. Finally, we make a conclusion and analyze future research directions.
Zhihui Lu 0002, Xiao Hong Gao, Sijia Huang, Yi Huang 0020
ICPADS1
2011 Apply WS-Management to Manage Real Resources: MAPE Use Case Study
abstract
Web Services standard-based system resource management is a new direction. In this paper, in order to verify the WS-Management standard whether to satisfy the realistic system resource management requirements, we use IBM MAPE categories to find more WS-Management related use cases. We design three typical use cases based on MAPE principles. Finally, we make a conclusion and propose our next-step work.
Zhihui Lu 0002, Jie Wu 0003, Weiming Fu
ICWS1
2010 A Novel Cloud-Oriented WS-Management-Based Resource Management Model
abstract
Cloud computing environment requires a more open and loosely-coupled service and resource management model. Web Services for Management specification (WS-Management), as an initiative of DMTF organization, can help to manage IT resources cross multiple domains in cloud environment. In this paper, we propose a novel WS-Management-based Cloud-oriented resource management model. We describe the main components of this model. And then, we discuss our management model verification experimental scheme focusing on DASH resource. Finally, we present conclusion and future work.
Zhihui Lu 0002, Jie Wu 0003, Weiming Fu
ICWS1
2010 CPH-VoD: A Novel CDN-P2P-Hybrid Architecture Based VoD Scheme
Zhihui Lu 0002, Jie Wu 0003, Lijiang Chen, Sijia Huang, Yi Huang 0020
WISE1
2009 MWS-MCS: A Novel Multi-agent-assisted and WS-Management-based Composite Service Management Scheme
abstract
From the analysis of some hard drawbacks faced with system service and composite service management today, a novel multi-agent-assisted and WS-management-based composite service management scheme: MWS-MCS is introduced in this paper. Firstly, we propose model architecture of this scheme. The prototype system had proved the feasibility of this design scheme. At last, we conclude this paper and analyze the prospective research challenges.
Zhihui Lu 0002, Jie Wu 0003, YiPing Zhong
ICWS1
2009 Research on WS-Management-based System and Network Resource Management Middleware Model
abstract
Nowadays, system and network resource management software should deal with more and more heterogeneous specific interfaces of different resource. This is a tightly-coupled management model. Recent years, Web Services have become the major technology and architecture of SOA for enterprise applications. Web Services provides a loosely-coupled management model. In this paper, based on the Web Services-based management protocol-WSManagement, we propose the distributed System and Network Resource Management Middleware Model. In this model, every managed IT resource provides the manageability interfaces via WS-Management specification. Furthermore, we utilize WSManagement Java implementation prototype-wiseman and WMI management interface to carry out the scheme implementation and test case work of the novel model, and then analyze the experiment results. At last, we analyze the prospective research direction and challenges in this field.
Zhihui Lu 0002, Jie Wu 0003, Shiyong Zhang, YiPing Zhong
ICWS1
2008 MultiPeerCast: A Tree-Mesh-Hybrid P2P Live Streaming Scheme Design and Implementation Based on PeerCast
abstract
In this paper, we firstly analyze the mechanism of PeerCast as a P2P live streaming solution. An then, based on the shortcoming analysis of tree-based PeerCast, we propose a improved tree-mesh-hybrid scheme-MultiPeerCast, including Multi-thread media transmission, multiple-to-one overlay network reconstruct, optimized buffer design, media retrieving style changing from push mode to pull mode. As part of experiment work, we discuss an improved P2P live streaming prototype system implementation based on MultiPeerCast. The experiment results verify MultiPeerCast scheme is more stable and efficient than PeerCast. At last, from our research experiences and related survey, we analyze the prospective research direction and challenges in this field.
Zhihui Lu 0002, Jie Wu 0003, Shiyong Zhang, YiPing Zhong
HPCC1
2004 Research on Service Model of Content Delivery Grid
Zhihui Lu 0002, Shiyong Zhang, YiPing Zhong
APWeb1