Gang Liu 0038

dblp:37/2109-38 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0003-0971-714XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Computer networks · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 ST-GCN and Reinforcement Learning-Assisted Dynamic Multistrategy Task Offloading in Edge-IoT Vehicular Networks
abstract
The Internet of Things (IoT) enables intelligent transportation services by connecting vehicles with roadside infrastructure and generating time-sensitive data. To support low-latency processing, edge-IoT vehicular networks deploy distributed edge servers near mobile users. However, high vehicular mobility and heterogeneous edge resources make it difficult for existing approaches to effectively exploit spatio-temporal mobility patterns and to support real-time offloading decisions. To address these challenges, this paper proposes TPADO, a Trajectory Prediction-Aware Dynamic Offloading framework that integrates a Spatio-Temporal Graph Convolutional Network (ST-GCN) with a multi-agent decision mechanism based on Proximal Policy Optimization (PPO). TPADO employs ST-GCN to perform high-fidelity trajectory prediction by explicitly modeling the graph structure of vehicular networks, thereby enabling proactive candidate-node selection and mobility-aware result delivery. Based on the predicted mobility information, we further design a hierarchical multi-strategy offloading framework, where a DRL-based policy layer adaptively selects offloading strategies, and a rule layer performs fine-grained node assignment and task partitioning. Extensive simulation results demonstrate that TPADO achieves the best overall performance among the compared methods. Compared with the centralized DQN baseline, it reduces global average latency by 6.4% and system saturation by 3.01 percentage points, while also delivering higher throughput and task success rate. These results validate the effectiveness and generalizability of the proposed framework.
Chuang Li 0004, Gang Liu 0038, Yanhua Wen, Junyan Hu, Qingyu Shi 0001, Zhao Tong 0001
IEEE Internet Things J.3
2026 MC-ORAM: A Concurrent ORAM Scheme for Multi-User Shared Storage
abstract
The expansion of cloud-based shared storage increases data privacy concerns. While data encryption technologies can safeguard data content, they cannot prevent the leakage of data access patterns. By re-encrypting data and changing storage location after each access, Oblivious Random Access Machine (ORAM) can effectively avoid information leakage from memory access patterns. While ORAM was initially designed for single-user applications, most existing multi-user ORAM solutions have drawbacks, such as dependence on a trusted proxy, high client storage overhead, and low throughput. To address the issues of multi-user ORAM systems, this paper explores the design of proxyless ORAM solutions in shared storage scenarios and proposes MC-ORAM, a new multi-user oblivious data storage framework. It ensures data consistency through client collaboration in a proxyless architecture, achieves higher throughput by differentiating request processing based on privacy protection requirements, and uses recursion for optimization. We implemented MC-ORAM and analyzed its performance using a variety of indicators, as well as conducting comparative evaluations against alternative schemes. The results show that, on average, MC-ORAM reduces response latency by 18.1% and improves throughput by 23.8% compared to TaoStore.
Chuang Li 0004, Duo Hu, Gang Liu 0038, Yanhua Wen, Zhuo Tang
IEEE Trans. Computers3
2026 Performance Optimization of Split Federated Learning in Heterogeneous Edge Computing Environments
abstract
Clients in federated learning (FL) may exhibit varying computing capabilities, leading to prolonged training latency when deploying complex deep neural networks. To address this challenge, split federated learning (SFL) presents an approach that offloads the main computational workload from resource-constrained devices to a server, while enabling parallel training. However, there are two significant limitations of existing SFL frameworks: The adoption of a uniform cut layer strategy fails to take into account the heterogeneous among clients; it fails to effectively utilize server-side resources to improve training efficiency. This article presents a framework, i.e., heterogeneous split federated learning, which considers personalized cut layer selection and server resource configuration to accelerate SFL in heterogeneous edge computing environments. By splitting the global model into two components for each client, our framework jointly optimizes both client-side workload, batch size control, and server resource configuration strategy, while considering device heterogeneity. Specifically, we develop an alternating iterative scheduling algorithm to obtain an approximate scheme for the cut layer, batch sizes, and server resource configuration to alleviate the impact of device heterogeneity. The experimental results illustrate that HSFL outperforms the compared methods, achieving performance improvements of up to 3.9%$\sim$32.2% across two datasets under various data distribution scenarios, which demonstrates the effectiveness of the proposed strategies.
Junyan Hu, Yuansheng Liang, Yanping Chen 0006, Gang Liu 0038, Weiwei Chen 0004, Lixin Duan
IEEE Trans. Ind. Informatics4
2026 VM-ORAM: A Novel High-Performance ORAM Architecture for Efficient Data Integrity Verification in Industrial Cloud
abstract
With the rapid surge in industrial data, cloud computing has been integrated into Industrial Internet of Things (IIoT) systems to store, compute, and share massive data. In this process, privacy and data integrity are core concerns. Oblivious RAM (i.e., ORAM) is a technology widely applied to defend against cloud storage access pattern attacks. However, most existing ORAM systems do not consider the integration of data integrity verification technology. Although there are integrity verification systems integrated into conventional Path ORAM and Ring ORAM, they are not suitable for existing new ORAM systems. And the existing data integrity verification ORAM system still has the problem of excessive performance overhead. To address these challenges, this article proposes a novel high-performance data integrity ORAM system, VM-ORAM. Optimizes the ORAM integrity verification process by integrating dynamic scheduling and multipath eviction strategies, thereby minimizing performance loss. The comprehensive analysis and experimental results of this article show that VM-ORAM system not only defends against data tampering attacks, but also maintains high performance of the system.
Chuang Li 0004, Gang Liu 0038, Changyao Tan, Limei Liu, Wenhua Ye, Anthony T. Chronopoulos
IEEE Trans. Ind. Informatics3
2026 Cost-Optimized Periodic DAG-Structured Task Offloading in Multi-User MEC Systems Using Reinforcement Learning
abstract
Reinforcement Learning (RL) has emerged as a promising solution for task offloading due to its adaptability to dynamic environments and ability to reduce online computational overhead. Thereby, this article explores RL for optimizing periodic Directed Acyclic Graph (DAG) task offloading in multi-user Mobile Edge Computing (MEC) systems, aiming to minimize overall costs, including user device energy consumption and server computational charges. A key contribution of this work is the explicit modeling of user competition for limited edge resources, where concurrent access leads to dynamic contention, significantly affecting offloading latency and energy usage. However, this optimization task faces two main challenges: the high dimensionality of task states and the large action space, both of which increase learning complexity. To address this, we propose a dynamic and distributed Proximal Policy Optimization (PPO)-based offloading framework. An encoder is employed to map DAG node features and structural information into a lower-dimensional representation, reducing computational overhead and improving learning efficiency. Additionally, we incorporate behavioral cloning to imitate greedy policies as the PPO agent’s initial behavior, effectively narrowing the action space and accelerating convergence. By combining representation learning and imitation-based initialization, our method enables the PPO agent to quickly adapt to environmental dynamics, leveraging both prior knowledge and real-time feedback to make informed offloading decisions. Simulation results confirm that our approach achieves rapid convergence and outperforms existing baselines in cost reduction, demonstrating its effectiveness for periodic task offloading in MEC scenarios. The source code and implementation details are available at: https://github.com/xiaolutihua/GAT/tree/master .
Yan Wang 0022, Gang Liu 0038, Keqin Li 0001
ACM Trans. Internet Techn.3
2025 A High-Quality Data-Driven Incentive Mechanism for Multi-Platform Mobile Crowdsensing
abstract
With the widespread adoption of large-scale mobile smart devices, Mobile Crowdsensing is evolving from a single-platform paradigm toward a multi-platform architecture. In such multi-platform environments, ensuring task completion while maintaining a balanced user borrowing-to-lending (BL) ratio across platforms has emerged as a critical challenge, particularly in enhancing the quality of sensing data. To address this issue, we propose a Multi-stage High-quality Data-driven Incentive Mechanism (MHDIM). The mechanism comprises three sequential stages. In the first stage, we design the ITA algorithm, which jointly considers user data quality and task urgency to allocate tasks among intra-platform users. Each platform then uploads information about its idle users and unassigned tasks to a cross-platform management entity. In the second stage, we propose the CMSIP algorithm, which leverages user quality and the BL ratio of each platform to assist the cross-platform entity in assigning tasks with shared Points of Interest. In the third stage, the CMDIP algorithm is introduced to handle cross-platform task allocation involving different points of interest, taking into account user data quality, the spatial distance between users and tasks, and the platform-level BL ratio. Experimental results demonstrate that, compared with existing mainstream incentive mechanisms, the proposed MHDIM significantly outperforms in terms of system utility, task completion rate, and data quality within multi-platform settings.
Minghe Zhang, Lihong Chen, Gang Liu 0038
TrustCom4
2025 DC-ORAM: An ORAM Scheme Based on Dynamic Compression of Data Blocks and Position Map
abstract
Oblivious RAM (ORAM) is an efficient cryptographic primitive that prevents leakage of memory access patterns. It has been referenced by modern secure processors and plays an important role in memory security protection. Although the most advanced ORAM has made great progress in performance optimization, the access overhead (i.e., data blocks) and on-chip (i.e., PosMap) storage overhead is still too high, which will lead to problems such as low system performance. To overcome the above challenges, in this paper, we propose a DC-ORAM system, which reduces the data access overhead and on-chip PosMap storage overhead by using dynamic compression technology. Specifically, we use byte stream redundancy compression technology to compress data blocks on the ORAM tree. And in PosMap, a high-bit multiplexing strategy is used to achieve data compression for binary high-bit repeated data of leaf labels (or path labels). By introducing the above compression technology, in this work, compared with conventional Path ORAM, the compression rate of the ORAM tree is$52.9\%$, and the compression rate of PosMap is$40.0\%$. In terms of performance, compared to conventional Path ORAM, our proposed DC-ORAM system reduces the average latency by$33.6\%$. In addition, we apply the compression technology proposed in this work to the Ring ORAM system. By comparison, it is found that with the same compression ratio as Path ORAM, our design can still reduce latency by an average of$21.5\%$.
Chuang Li 0004, Changyao Tan, Gang Liu 0038, Yanhua Wen, Yan Wang 0022, Kenli Li 0001
IEEE Trans. Computers3
2025 Enabling Large Scale LoRa Parallel Decoding With High-Dimensional and High-Accuracy Features
abstract
LoRaWAN is a prominent technology for Low Power Wide Area Networks (LPWAN). However, the increasing network size has introduced a significant challenge: packet collisions resulting from concurrent transmissions in LoRaWAN. Previous studies either overlooked the issue by examining limited features or tackled it with intricate receivers employing up to eight antennas. To achieve a more favorable balance between implementation cost and system performance, we introduce$\text{Hi}^{2}\text{LoRa}$—a solution utilizing highly dimensional and accurate features for LoRa concurrent decoding, implemented with only two receiving antennas. The feature dimensions are expanded through an exploration of various hardware imperfections and inherent channel state information specific to each transceiver pair. To enhance feature accuracy, low pass filters and BiLSTM networks are applied to capture and learn their temporal patterns. Additionally, an efficient collision suppression strategy is introduced to mitigate feature corruption from concurrently transmitted packets. Extensive real-world testbed evaluations demonstrate that the achievable concurrency in$\text{Hi}^{2}\text{LoRa}$approaches that of state-of-the-art approaches with significantly higher complexity (e.g., utilizing eight antennas) or exceeds prior work by a factor of 2.7 with comparable complexity (e.g., using two antennas).
Weiwei Chen 0004, Xianjin Xia, Shuai Wang 0008, Tian He 0001, Shuai Wang 0021, Gang Liu 0038, Caishi Huang
IEEE Trans. Mob. Comput.6
2025 HM-ORAM: A Lightweight Crash-consistent ORAM Framework on Hybrid Memory System
abstract
Byte-addressable non-volatile memory (NVM) is a promising alternative technology for main memory, allowing the processor to access persistent data in the main memory directly. Systems with emerging NVM as the main memory still suffer from information leakage, performance, and endurance challenges. Fortunately, Oblivious RAM (ORAM) provides provable secure address obfuscation and encryption and can be used to build a secure memory system. However, it is challenging to support crash consistency with the increasing use of NVM as persistent memory. Current state-of-the-art PS-ORAM systems are designed based on pure NVM memory systems, which suffer from low performance. Different from the previous PS-ORAM system, in this work, we consider a hybrid memory system (both NVM and DRAM) and aim to build a crash-consistent ORAM framework with low overhead. We propose the HM-ORAM framework, which includes a novel ORAM tree-level partition scheme, a lightweight data persistency architecture, and a secure access protocol. It takes full advantage of both hybrid memory technologies to achieve the design requirements to support crash consistency and persistence of ORAM systems with minimal write overhead. Without compromising security, HM-ORAM can outperform an NVM-based ORAM (i.e., NVM-ORAM) system by 1.23× and 1.39× in non-recursive and recursive implementations.
Gang Liu 0038, Kenli Li 0001, Rujia Wang
ACM Trans. Storage1
2024 Automated Optical Accelerator Search Toward Superior Acceleration Efficiency, Inference Robustness, and Development Speed
abstract
Remarkable breakthroughs but daunting complexities of deep learning have aroused widespread interest in dedicated deep neural network (DNN) acceleration hardware, among which optical accelerators (OAs) are particularly promising thanks to their unprecedentedly high-performance-per-watt. However, the development of OAs is much slower than that of electrical accelerators due to threefold challenges. First, the OA design space is ample and discrete, making it tough for OA optimization; Second, the ecosystem that facilitates OA development is still in its infancy. Techniques to support OA design remain less explored, limiting both the achievable performance and the innovative development of OAs; and Third, OAs are highly sensitive to fabrication-induced process variations and thermal fluctuations (i.e., PTVs), which degrades OAs’ inference robustness and even renders them unusable in practice. In this article, we develop AutOAS, the first-of-its-kind framework for Automated Optical Accelerator Search, in order to jointly boost acceleration efficiency, inference robustness, and development speed. Our AutOAS comprises four enabling components: 1) a holistic OA search space, which takes full consideration of OAs’ micro-architectures (e.g., the type, shape and size of core functional units for data computation and data access), dataflow choices, DNN-to-accelerator mapping methods, memory hierarchy and PTV mitigation techniques; 2) a PTV Regulator, which can emulate the impact of PTVs on OAs’ inference accuracy based on given PTV profiles, and enables energy-efficient PTV mitigation on OAs; 3) an O-Performance Predictor, which enables accurate yet efficient predictions of an OA’s energy, throughput (latency) and chip area according to the DNN model and OA architecture parameters; and 4) two O-Search Engines (i.e., a differentiable search engine and an evolutionary search engine), which can automatically explore the large design space of OAs and identify the optimal accelerators to maximize the acceleration targets. Based on 10 DNN models widely applied in both computer vision and sequence modeling tasks, extensive experiments and ablation studies validate the effectiveness of our PTV Regulator, O-Performance Predictor, and O-Search Engines, as well as the superior performance of AutOAS-generated OAs.
Mengquan Li, Kenli Li 0001, Chao Wu 0006, Gang Liu 0038, Mingfeng Lan, Yunchuan Qin, Zhuo Tang, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 RFL-APIA: A Comprehensive Framework for Mitigating Poisoning Attacks and Promoting Model Aggregation in IIoT Federated Learning
abstract
With the development of industrial Internet of Things (IIoT), federated learning (FL) is important for protecting sensitive data from various Internet of Things devices (i.e., clients in FL). Despite FL's privacy benefits, attackers (e.g., untrusted clients) can still compromise the performance of the global model through model poisoning attacks. Unfortunately, two key challenges hinder effective detection and impact the performance of the global model in FL: first, accurately identifying malicious models to defend against attacks, and second, efficiently aggregating local models after detecting malicious clients. To address these challenges, we propose an improved FL system based on fuzzy rules, termed RFL-APIA. Compared to the conventional FL system, we have designed two novel components, federated learning generalized depth detection (FedGDD) and Fedsv-Weighted, to enhance performance and mitigate model poisoning attacks. Specifically, FedGDD introduces variance reduction by examining the relationship between local and global model gradients, thereby mitigating interference in nonindependent and identical distributed settings. It further implements an adaptive penalty factor-based scoring system, leveraging variations in local model updates for precise identification and mitigation of attacks. Based on FedGDD's output, the Fedsv-Weighted mechanism dynamically updates the global model's aggregation weights by considering local models' contributions, thus improving model aggregation. Extensive experiments demonstrate that RFL-APIA effectively prevents model poisoning attacks during training, ensuring model security, and guaranteeing a certain level of accuracy and convergence for the global model.
Chuang Li 0004, Aoli He, Gang Liu 0038, Yanhua Wen, Anthony T. Chronopoulos, Aristotelis Giannakos
IEEE Trans. Ind. Informatics3
2024 Towards Intelligent Adaptive Edge Caching Using Deep Reinforcement Learning
abstract
The tremendous expansion of edge data traffic poses great challenges to network bandwidth and service responsiveness for mobile computing. Edge caching has emerged as a promising method to alleviate these issues by storing a portion of data at the network edge. However, existing caching approaches suffer from either poor caching efficiency with low content-hit ratio or unintelligence of caching policies lacking self-adjustability. In this paper, we propose ICE, a novel Intelligent Edge Caching scheme using a deep reinforcement learning (DRL) method to capture specific valuable information from the requested data. With the benefit of our proposed popularity model based on Newton's law of cooling, ICE fully takes into account the popularity of the contents to be cached and leverages the formulated Markov decision model to decide whether or not the contents should be cached. Moreover, to further improve the caching efficiency, we propose a novel distributed multi-node caching framework, named DCCC, assisted by a multi-tiered caching hierarchy. Comprehensive experiments show that the single-node ICE scheme greatly improves the cache hit rate and contents exchanging time in comparison with both DRL-based and legacy approaches, and our distributed multi-node caching scheme DCCC further significantly improves the overall utilization of caching space.
Ting Wang 0001, Yuxiang Deng, Mingsong Chen 0001, Gang Liu 0038, Jieming Di, Keqin Li 0001
IEEE Trans. Mob. Comput.5
2023 Optimal Trading Mechanism Based on Differential Privacy Protection and Stackelberg Game in Big Data Market
abstract
Big data has become a fundamental resource and a commodity in economic activities, thus, it is necessary to build a market model capable of supporting efficient data trading. However, two major challenges remain. First, researches have considered constructing data trading mechanisms, while few of them are based on the method of measuring data value in multiple dimensions. Second, a data market involved an intermediary trading platform (i.e., a third party) which is honest but curious, results may be obtained due the the leakage of private information. In this article, we design TM-OUE, a data trading mechanism based on Optimized Unary Encoding that enables reasonable trading mechanism and protects the privacy of data trading. First of all, we combine qualitative and quantitative methods to measure the value of data in multiple dimensions and formulate data trading model between the data provider and data users. Then, we utilize an Optimized Unary Encoding (OUE) protocol to protect the privacy of the data trading mechanism. Based on the above steps, we develop a two-stage single leader multi-follower Stackelberg game to jointly maximize profits of the data provider and data users. Experimental results demonstrate that TM-OUE can offer appropriate price for data and maximize benefits both for data providers and data users, which guarantees fair data trades while protecting privacy.
Chuang Li 0004, Aoli He, Yanhua Wen, Gang Liu 0038, Anthony T. Chronopoulos
IEEE Trans. Serv. Comput.4
2022 PS-ORAM: efficient crash consistency support for oblivious RAM on NVM
abstract
Oblivious RAM (ORAM) is a provable secure primitive to prevent access pattern leakage on the memory bus. By randomly remapping the data blocks and accessing redundant blocks, ORAM prevents access pattern leakage through ob-fuscation. Byte-addressable non-volatile memory (NVM) is considered as the candidate for main memory due to its better scalability, competitive performance, and persistent data store. While there is much prior work focusing on improving ORAM's performance on the conventional DRAM-based memory system, when the memory technology shifts to use NVM, ensuring an efficient crash-consistent ORAM is needed for security, correctness, and performance. Directly using traditional software-based crash consistency support for ORAM system is not only expensive but also insecure.
Gang Liu 0038, Kenli Li 0001, Rujia Wang
ISCA1
2022 Towards an energy-efficient Data Center Network based on deep reinforcement learning
Yang Wang 0019, Ting Wang 0001, Gang Liu 0038
Comput. Networks4
2022 A Many-to-Many Demand and Response Hybrid Game Method for Cloud Environments
abstract
In this article, we design a service mechanism for profits optimization between multiple cloud providers and multiple cloud customers (many-to-many). We explore this problem from the perspective of game theory but take a different approach compared with existing cloud resource pricing game methods. First, we regard the relationships among multiple cloud customers as an evolutionary game, and formulate the competitions among the multiple cloud providers as a noncooperative game. Eventually, we form a hybrid game model in which the strategy of each customer and each cloud provider is affected not only by the other side but also by customers or cloud providers other than themselves. Second, based on the hybrid game model, we simulate the bargaining process between cloud providers and customers by controlling supply and demand allocation, and try to ultimately achieve a balanced supply and demand state, i.e., a win-win situation. For each cloud customer and provider, we design a utility function. A customer’s utility involves net profits and the cloud providers’ bidding strategies, and a cloud provider’s utility involves net profits and the cloud customers’ demand strategies. Both sides attempt to maximize their own profits under the influences of each other. We prove that our proposed strategies enable each of the two games to converge to their own equilibrium. Finally, the strategies of cloud customers and providers can be implemented through an iterative proximal algorithm ($\mathcal {IPA}$) and a distributed iterative algorithm ($\mathcal {DIA}$). The experimental results validate our methods and show that the proposed method can benefit both multiple cloud providers and customers.
Gang Liu 0038, Anthony T. Chronopoulos, Chubo Liu, Zhuo Tang
IEEE Trans. Cloud Comput.1
2021 ICE: Intelligent Caching at the Edge
abstract
The unprecedented growth of mobile data traffic brings unique challenges for network bandwidth and server resources to meet the diverse QoE (Quality of Experience). Caching becomes a promising way to alleviate these issues by storing a subset of data at the network edge, for which caching policy becomes critical. To this end, various caching schemes have been put forward, however, these schemes are either not intelligent lacking the ability of self-learning and self-decision-making, or inefficient with low data hit rate. Based on these observations, in this paper, we propose a novel Intelligent Caching framework at the Edge, named ICE, via deep reinforcement learning to capture certain valued information of the requested data. Notably, in our approach, the popularity of the data to be cached will be explored and considered. A Markov decision model is further developed to determine whether the data should be cached. The evaluation shows that ICE greatly improves the hit rate in comparison with the state-of-the-art approaches, and reduces the energy consumption for data transmission. Furthermore, based on ICE, the users' QoE is greatly improved. In conclusion, both theoretical analysis and experimental results prove the effectiveness and high performance of ICE compared with conventional strategies.
Ting Wang 0001, Mingsong Chen 0001, Gang Liu 0038, Jieming Di, Shui Yu 0001
GLOBECOM4
2020 Game theory-based optimization of distributed idle computing resources in cloud environments
Gang Liu 0038, Guanghua Tan, Kenli Li 0001, Anthony T. Chronopoulos
Theor. Comput. Sci.1