VLDB 2026 Research / reviewers in the wild / expert
Zicong Hong
dblp:239/5523
· DBLP profile ↗
50ranked-venue papers
8as first author
45since 2021 · last 2026
0000-0001-5689-382XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 18 · 3 first-author · 15 since 2021Systems, architecture and hardware · 17 · 4 first-author · 15 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 7 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Multimodal Serving via Module MultiplexingabstractMultimodal learning enables models to process and reason over diverse information sources, unlocking human-like perceptual and cognitive capabilities. As such models gain adoption, efficiently serving them on GPUs has become increasingly important. However, the modular architecture of multimodal models poses significant challenges to existing unimodal serving systems, which treat models as monolithic and overlook inter-module heterogeneity. This results in severe GPU underutilization. To address this, we propose Eevee, a multimodal serving system based on a new scheduling paradigm we call module multiplexing. Unlike prior approaches that execute all modules sequentially with uniform batch sizes, Eevee schedules modality-specific modules concurrently on the same GPU with independently tuned batching and resource allocation. This design enables fine-grained GPU sharing, boosting intra-GPU parallelism and improving request-level throughput. We implement a prototype of Eevee and evaluate it on several representative multimodal models (e.g., CLIP, BLIP, LLaVA, InternVL). Our results show that Eevee significantly outperforms state-of-the-art serving systems in both throughput and GPU utilization. Zicong Hong, Yuyan Chen, Peng Li 0017, Wuhui Chen, Song Guo 0001 |
EuroSys | 1 |
| 2026 | Crimson: Collaborative Parameter Updates for Efficient Pipeline Training of Large Language ModelsabstractLarge language models (LLMs) have driven significant progress in natural language processing, yet their training and fine-tuning remain limited by memory constraints, particularly the substantial memory footprints of optimizer states. Existing solutions address this challenge by offloading optimizer states and update tasks to the CPU, but this often leads to increased GPU idleness due to the CPU's limited computational capabilities, especially in pipeline parallelism. Yapeng Jiang, Wuhui Chen, Ganhong Huang, Yuzhou Huang, Zicong Hong, Song Guo 0001, Yue Yu 0001 |
EuroSys | 5 |
| 2026 | FlashServe: Adaptive Kernel Provision for Quantized LLM Serving
Yuyan Chen, Junyuan Liang, Wuhui Chen, Zicong Hong, Song Guo 0001, Ruiyan Zhuang, Yi Quan |
ICDCS | 5 |
| 2026 | Director: Accelerating Distributed MoE Serving via Online Proactive Expert PlacementabstractExpert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs. Existing works for optimizing expert placement focus on leveraging past requests' expert activation patterns. However, they demonstrate deficiencies facing diverse and rapidly changing request patterns, calling for an online, proactive approach. Implementing such an approach requires addressing several challenges: the uncertainty associated with incoming requests' expert activation, the cost of expert migration, and the NP-hard complexity in optimization. Therefore, we present Director, a new distributed MoE serving system that minimizes end-to-end latency via prediction-driven, online expert placement. Director uses either a lightweight cascaded predictor or a low-bit quantized replica for expert activation patterns of incoming requests. An online migration module then enacts the changes with near-zero downtime by executing migrations in compute-bound phases, keeping disruption bounded. At its core, a relaxation-based expert placement optimizer operates under capacity constraints, runs in polynomial time, and achieves a (1+ ϵ) approximation ratio. Finally, we implement a prototype and demonstrate, through extensive experiments, a reduction in end-to-end latency of 11 ~ 55% for popular MoE models (e.g., Mistral, DeepSeek and Qwen) compared to existing work. Qianli Liu, Kaibin Guo, Zicong Hong, Peng Li 0017, Fahao Chen, Song Guo 0001 |
INFOCOM | 3 |
| 2026 | PPAI: Enabling Personalized LLM Agent Interoperability for Collaborative Edge Intelligence
Zile Wang, Qianli Liu, Kaibin Guo, Zicong Hong, Song Guo 0001 |
INFOCOM | 6 |
| 2026 | Fate: Fasss sEsdge Inference of Mixture-of-Experts Models via Cross-Layer GateabstractWith the rapid growth and rising complexity of web content, edge-deployed LLMs have become essential for enhancing users' online experiences. However, sparsely-activated Mixture-of-Experts (MoE) models, which are well-suited for edge scenarios, face significant memory bottleneck challenges. Offload-based methods have been proposed to mitigate the problem, but they face difficulties with expert prediction. To promote the application of MoE models in edge scenarios, we propose Fate, an offloading system designed for MoE models to enable efficient inference in resource-constrained environments. The key insight behind Fate is that gate inputs from adjacent layers can be effectively used for expert prefetching, achieving high prediction accuracy. Furthermore, Fate employs a shallow-favoring expert caching strategy that increases the expert hit rate to 99%. Additionally, Fate integrates tailored quantization strategies for cache optimization and I/O efficiency. Experimental results show that, compared to baselines, Fate achieves up to 1.34×-5.07× prefill speedup and 1.26×-4.41× decoding speedup, while maintaining inference quality. Zhiyuan Fang, Xingfan Yu, Yuegui Huang, Zicong Hong, Yufeng Lyu, Wuhui Chen, Yue Yu 0001, Fan Yu 0004 |
WWW | 4 |
| 2026 | Scaling Blockchain via Dynamic ShardingabstractSharding is considered a promising solution for scaling blockchain systems. However, most existing sharding systems have not considered the dynamics of the environment when making a sharding strategy, including the change of pending transactions, the leaving and joining of participants, and malicious attacks, which could cause performance instability and security issues. To address it, in this paper, we propose an intelligent and efficient dynamic sharding technology to advance the blockchain system performance and security. We first propose a formal and general evaluation framework for blockchain sharding in a dynamic environment, and conclude an optimization target for the system performance and security. To achieve a long-term benefit for the optimization target, a deep reinforcement learning (DRL)-based sharding approach has been proposed to intelligently make optimal sharding strategies. Next, we propose an adaptive resharding protocol to efficiently reduce the overhead introduced by dynamic sharding. Our experimental results illustrate that our proposed dynamic sharding in a simulation testbed can achieve 2.8 times transactions per second compared to traditional static sharding systems, and guarantee high security in a dynamic environment. Zicong Hong, Xiaoyu Qiu, Wuhui Chen, Yufeng Zhan, Song Guo 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | AFedLF: Adaptive Layer Freezing of Foundation Models in Heterogeneous Federated LearningabstractThe rise of pre-trained foundation models (FMs) has popularized the trend of fine-tuning FMs to fit downstream tasks, while Federated Learning (FL) has become the de-facto approach for training distributed data with privacy-preservation. However, fine-tuning FMs in FL faces overwhelming overheads due to its bulky nature. While freezing parameters in FM have the potential to accelerate FL training, existing freezing strategies statically freeze parameters on specified or already converged layers, incur severe accuracy degradation, and resource-inefficiency in heterogeneous environments. In this paper, we propose AFedLF, an adaptive freezing framework for FM in FL, to accelerate its wall-clock time for convergence without losing its final accuracy. However, this poses great challenges, as different freezing strategies lead to different accuracy gains and time overheads, while unfreezing more layers may bring marginal accuracy gains but significant time overheads. To address this challenge, AFedLF mathematically establishes a correlation between the freezing strategy and the accuracy gain and time overhead, and allocates adaptive freezing strategies to clients, based on our insight that unfreezing more layers on devices with strong computation and communication capabilities helps improve resource efficiency. Besides, AFedLF incorporates our well-designed intermediate result caching scheme with constant approximation ratios utilizing the limited storage capacity on mobile devices to cache intermediate results to skip forward propagation, further saving wall-clock time. Finally, we implemented AFedLF using an open-source FL benchmark, and extensive trace-driven experimental results showed that AFedLF accelerates wall-clock time by up to 6.1× compared to state-of-the-art solutions, without sacrificing accuracy. Yue Zeng 0002, Jie Zhang 0076, Song Guo 0001, Zhihao Qu, Zicong Hong, Bin Tang 0002, Junlong Zhou, Jiaying Yu |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch PipelineabstractMixture of Experts (MoE), with its distinctive sparse structure, enables the scaling of language models up to trillions of parameters without significantly increasing computational costs. However, the substantial parameter size presents a challenge for inference, as the expansion in GPU memory cannot keep pace with the growth in parameters. Although offloading techniques utilise memory from the CPU and disk and parallelise the I/O and computation for efficiency, the computation for each expert in MoE models is often less than the I/O, resulting in numerous bubbles in the pipeline. Zhiyuan Fang, Yuegui Huang, Zicong Hong, Yufeng Lyu, Wuhui Chen, Yue Yu 0001, Fan Yu 0004, Zibin Zheng |
ASPLOS (2) | 3 |
| 2025 | Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
Qianli Liu, Zicong Hong, Peng Li 0017, Fahao Chen, Song Guo 0001 |
INFOCOM | 2 |
| 2025 | D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM ServingabstractThe mixture of experts (MoE) model is a sparse variant of large language models (LLMs), designed to hold a better balance between intelligent capability and computational overhead. Despite its benefits, MoE is still too expensive to deploy on resource-constrained edge devices, especially with the demands of on-device inference services. Recent research efforts often apply model compression techniques, such as quantization, pruning and merging, to restrict MoE complexity. Unfortunately, due to their predefined static model optimization strategies, they cannot always achieve the desired quality-overhead trade-off when handling multiple requests, finally degrading the on-device quality of service. These limitations motivate us to propose the D2MoE, an algorithm-system co-design framework that matches diverse task requirements by dynamically allocating the most proper bit-width to each expert. Specifically, inspired by the nested structure of matryoshka dolls, we propose the matryoshka weight quantization (MWQ) to progressively compress expert weights in a bit-nested manner and reduce the required runtime memory. On top of it, we further optimize the I/O-computation pipeline and design a heuristic scheduling algorithm following our hottest-expert-bit-first (HEBF) principle, which maximizes the expert parallelism between I/O and computation queue under constrained memory budgets, thus significantly reducing the idle temporal bubbles waiting for the experts to load. Evaluations on real edge devices show that D2MoE improves the overall inference throughput by up to 1.39× and reduces the peak memory footprint by up to 53% over the latest on-device inference frameworks, while still preserving comparable serving accuracy as its INT8 counterparts. Qihua Zhou, Zicong Hong, Song Guo 0001 |
MobiCom | 3 |
| 2025 | DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert PruningabstractDespite the significant breakthrough of Mixture-of-Experts (MoE), the increasing scale of these MoE models presents huge memory and storage challenges. Existing MoE pruning methods, which involve reducing parameter size with a uniform sparsity across all layers, often lead to suboptimal outcomes and performance degradation due to varying expert redundancy in different MoE layers. To address this, we propose a non-uniform pruning strategy, dubbed Differentiable Expert Pruning (DiEP), which adaptively adjusts pruning rates at the layer level while jointly learning inter-layer importance, effectively capturing the varying redundancy across different MoE layers. By transforming the global discrete search space into a continuous one, our method handles exponentially growing non-uniform expert combinations, enabling adaptive gradient-based pruning. Extensive experiments on five advanced MoE models demonstrate the efficacy of our method across various NLP tasks. Notably, \textbf{DiEP} retains around 92\% of original performance on Mixtral 8$\times$7B with only half the experts, outperforming other pruning methods by up to 7.1% on the challenging MMLU dataset. Sikai Bai, Haoxi Li, Jie Zhang 0076, Zicong Hong, Song Guo 0001 |
NeurIPS | 4 |
| 2025 | Obscura: Concealing Recomputation Overhead in Training of Large Language Models with Bubble-filling Pipeline Transformation
Yuzhou Huang, Yapeng Jiang, Zicong Hong, Wuhui Chen, Bin Wang 0034, Weixi Zhu, Yue Yu 0001, Zibin Zheng |
USENIX ATC | 3 |
| 2025 | Pistis: A Decentralized Knowledge Graph Platform Enabling Ownership-Preserving SPARQL QueryingabstractDecentralized Knowledge Graph (DKG) platforms allow the sharing of knowledge with multiple owners. While data owners can share their data with others by encrypting their data before sharing it, this naïve approach prevents data encrypted by different owners from being queried together, as it compromises query verifiability, an essential DKG platform feature. We propose Pistis, the first DKG platform capable of preserving ownership while also enabling verifiable SPARQL queries. Two novel techniques facilitate this: owner-managed end-to-end encryption and collaborative query verification. In Pistis, data owners thus encrypt their data individually and collaborate to construct an authenticated data structure (ADS) with a global key by means of secret sharing and secure multi-party computation. Then, by indexing KG data as ciphertext over the ADS, Pistis offers a cryptographic scheme called VO-SPARQL that facilitates verifiable queries on encrypted KG data with multiple owners. Pistis provides succinct proofs for two-stage SPARQL queries, including subgraph queries based on the ADS and aggregation on encrypted intermediate results based on a key-aggregate cryptographic primitive. A theoretical analysis and an empirical study provide detailed insight into the performance of Pistis while offering provable security. Enyuan Zhou, Song Guo 0001, Zicong Hong, Christian S. Jensen, Yang Xiao 0014, Jinwen Liang, Dalin Zhang 0001 |
Proc. VLDB Endow. | 3 |
| 2025 | Qora: Neural-Enhanced Interference-Aware Resource Provisioning for Serverless ComputingabstractServerless is an emerging cloud paradigm that offers fine-grained resource sharing through serverless functions. However, this resource sharing can cause interference, leading to performance degradation and QoS violations. Existing white box-based approaches for serverless resource provision often demand extensive expert knowledge, which is challenging to obtain due to the complexity of interference sources. This paper proposes Qora, a neural-enhanced interference-aware resource provisioning system for serverless computing. We model the resource provisioning of serverless functions as a novel combinatorial optimization problem, wherein the constraints on the queries per second are derived from neural network performance model. By leveraging neural networks to model the nonlinear performance fluctuations under various interference sources, our approach better captures the real-world behavior of serverless functions. To solve the formulated problem efficiently, rather than adopting commercial optimizer solvers like Gurobi, we propose a two-stage-VNS algorithm that searches discrete variables more efficiently and supports Sigmoid activations, avoiding introducing redundant discrete variables. Unlike pure machine learning methods lacking theoretical optimal guarantees, our approach is rigorously proven globally optimal based on optimization theory. We implement Qora on Kubernetes as a serverless system automating resource provisioning. Experimental results demonstrate that Qora reduces the QoS violation rate by 98% while reducing up to 35% resource costs compared with the state-of-the-arts. Note to Practitioners—From the perspective of cloud service providers, this paper considers the automatic resource provisioning for serverless functions. To improve hardware utilization, cloud providers tend to co-locate serverless functions on the same server. However, co-located functions compete for shared resources (memory bandwidth, L3 cache, etc.), which causes interference and leads to performance degradation and QoS violations. We use neural networks to build the performance models of interference-prone serverless functions and form the resource allocation optimization problem with neural network performance models as constraints. Compared to white box modeling methods, our neural network modeling adapts to complex and variable interference. Compared to deep reinforcement learning methods, our combinatorial optimization methods have stronger interpretability. In order to solve this optimization problem efficiently, we design the two-stage-VNS solution algorithm. We implement Qora on Kubernetes as a serverless system, which can automatically allocate computing resources. Experiments with small-scale real clusters and large-scale simulations demonstrate the effectiveness of Qora. Ruifeng Ma, Yufeng Zhan, Chuge Wu, Zicong Hong, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Collaborative Neural Architecture Search for Personalized Federated LearningabstractPersonalized federated learning (pFL) is a promising approach to train customized models for multiple clients over heterogeneous data distributions. However, existing works on pFL often rely on the optimization of model parameters and ignore the personalization demand on neural network architecture, which can greatly affect the model performance in practice. Therefore, generating personalized models with different neural architectures for different clients is a key issue in implementing pFL in a heterogeneous environment. Motivated by Neural Architecture Search (NAS), a model architecture searching methodology, this paper aims to automate the model design in a collaborative manner while achieving good training performance for each client. Specifically, we reconstruct the centralized searching of NAS into the distributed scheme called Personalized Architecture Search (PAS), where differentiable architecture fine-tuning is achieved via gradient-descent optimization, thus making each client obtain the most appropriate model. Furthermore, to aggregate knowledge from heterogeneous neural architectures, a knowledge distillation-based training framework is proposed to achieve a good trade-off between generalization and personalization in federated learning. Extensive experiments demonstrate that our architecture-level personalization method achieves higher accuracy under the non-iid settings, while not aggravating model complexity over state-of-the-art benchmarks. Yi Liu 0057, Song Guo 0001, Jie Zhang 0076, Zicong Hong, Yufeng Zhan, Qihua Zhou |
IEEE Trans. Computers | 4 |
| 2025 | SecPQ: Secure Prediction Queries on Encrypted Outsourced DatabasesabstractPrediction queries have revolutionized data search by integrating machine learning models and traditional data processing operations for advanced analytics. However, existing prediction query frameworks for outsourced databases face a critical security vulnerability: data flows are processed in plaintext on semi-honest servers, making them susceptible to data breaches. The main challenge in achieving secure prediction queries is that machine learning inference and data processing operations are distinct functionalities, while most current cryptographic frameworks support only a single type of operation on specific encrypted data. To bridge this crucial gap, we propose$\mathsf {SecPQ}$, the first framework tailored for secure prediction queries. Our approach unifies decision tree pipelines and data processing operations, such as selection, projection, and equality-joining, through equality matching on encrypted outsourced data. This enables the design of secure prediction queries with decision tree pipelines operating on encrypted data. We provide formal security definitions and proofs for$\mathsf {SecPQ}$. To further optimize the efficiency of secure prediction queries, we leverage order-preserving encryption to construct$\mathsf {SecPQ}_{{{ope}}}$, which offers improved query efficiency at the expense of weaker security properties compared with$\mathsf {SecPQ}$. Extensive experimental evaluations on billions of records demonstrate the feasibility and effectiveness of both$\mathsf {SecPQ}$and$\mathsf {SecPQ}_{{{ope}}}$. Jinwen Liang, Song Guo 0001, Zicong Hong, Enyuan Zhou, Chuan Zhang 0003, Bin Xiao 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Model Decomposition and Reassembly for Purified Knowledge Transfer in Personalized Federated LearningabstractPersonalized federated learning (pFL) is to collaboratively train non-identical machine learning models for different clients to adapt to their heterogeneously distributed datasets. State-of-the-art pFL approaches pay much attention on exploiting clients’ inter-similarities to facilitate the collaborative learning process, meanwhile, can barely escape from the irrelevant knowledge pooling that is inevitable during the aggregation phase, and thus hindering the optimization convergence and degrading the personalization performance. To tackle such conflicts between facilitating collaboration and promoting personalization, we propose a novel pFL framework, dubbed pFedC, to first decompose the global aggregated knowledge into several compositional branches, and then selectively reassemble the relevant branches for supporting conflicts-aware collaboration among contradictory clients. Specifically, by reconstructing each local model into a shared feature extractor and multiple decomposed task-specific classifiers, the training on each client transforms into a mutually reinforced and relatively independent multi-task learning process, which provides a new perspective for pFL. Besides, we conduct a purified knowledge aggregation mechanism via quantifying the combination weights for each client to capture clients’ common prior, as well as mitigate potential conflicts from the divergent knowledge caused by the heterogeneous data. Extensive experiments over various models and datasets demonstrate the effectiveness and superior performance of the proposed algorithm. Jie Zhang 0076, Song Guo 0001, Xiaosong Ma, Wenchao Xu 0001, Qihua Zhou, Jingcai Guo, Zicong Hong, Jun Shan |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token CondensationabstractMixture-of-Experts (MoE) is an emerging technique for scaling large models with sparse activation. MoE models are typically trained in a distributed manner with anexpert parallelismscheme, where experts in each MoE layer are distributed across multiple GPUs. However, the default expert parallelism suffers from the heavy network burden due to the all-to-all intermediate data exchange among GPUs before and after the expert run. Some existing works have proposed to reduce intermediate data exchanges by transferring experts to reduce the network loads, however, which would decrease parallelism level of expert execution and make computation inefficient. The weaknesses of existing works motivate us to explore whether it is possible to reduce inter-GPU traffic while maintaining a high degree of expert parallelism. This paper gives a positive response by presentingLuffy, a communication-efficient distributed MoE training system with two new techniques. First,Luffymigrates sequences among GPUs to hide heavy token pulling paths within GPUs and avoid copying experts over GPUs. Second, we propose token condensation that identifies similar tokens and then eliminates redundant transmissions. We implementLuffybased on PyTorch and evaluate its performance on a testbed of 16 V100 GPUs.Luffysystem can achieve a speedup of up to$2.73\times $compared to state-of-the-art MoE training systems. Fahao Chen, Peng Li 0017, Zicong Hong, Zhou Su 0001, Song Guo 0001 |
IEEE Trans. Netw. | 3 |
| 2025 | Sparrow: Expediting Smart Contract Execution for Blockchain Sharding via Inter-Shard CachingabstractSharding is a promising solution to scale blockchain by separating the system into multiple shards to process transactions in parallel. However, due to state separation and shard isolation, it is still challenging to efficiently support smart contracts on a blockchain sharding system where smart contracts can interact with each other, involving states maintained by multiple shards. Specifically, existing sharding systems adopt a costly multi-step collaboration mechanism to execute smart contracts, resulting in long latency and low throughput. This article proposesSparrow, a blockchain sharding protocol achieving one-step execution for smart contracts. To break shard isolation, inspired by non-local hotspot data caching in traditional databases, we propose a new idea ofinter-shard caching, allowing a shard to prefetch and cache frequently accessed contract states of other shards. The miner can thus use the inter-shard cache to pre-execute a pending transaction, retrieve all its contract invocations, and commit it to multiple shards in one step. Particularly, we first propose a speculative dispersal cache synchronisation mechanism for efficient and secure cache synchronization across shards in Byzantine environments. Then, we propose a multi-branch exploration mechanism to solve the rollback problem during the optimistic one-step execution of contract invocations with dependencies. We also present a series of conflict resolution mechanisms to decrease the rollback caused by inherent transaction conflicts. We implement prototypes forSparrowand existing sharding systems, and the evaluation shows thatSparrowimproves the throughput by$2.44\times$and reduces the transaction latency by 30% compared with the existing sharding systems. Junyuan Liang, Peiyuan Yao, Wuhui Chen, Zicong Hong, Ting Cai 0002, Zibin Zheng |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationabstractServerless ML inference is an emerging cloud computing paradigm for low-cost, easy-to-manage inference services. In serverless ML inference, each call is executed in a container; however, the cold start of containers results in long inference delays. Unfortunately, most existing works do not work well because they still need to load models into containers from scratch, which is the bottleneck based on our observations. Therefore, this paper proposes a low-latency serverless ML inference system called Optimus via a new container management scheme. Our key insight is that the loading of a new model can be significantly accelerated when using an existing model with a similar structure in a warm but idle container. We thus develop a novel idea of inter-function model transformation for serverless ML inference, which delves into models within containers at a finer granularity of operations, designs a set of in-container meta-operators for both CNN and transformer model transformation, and develops an efficient scheduling algorithm with linear complexity for a low-cost transformation strategy. Our evaluations on thousands of models show that Optimus reduces inference latency by 24.00% ~ 47.56% in both simulated and real-world workloads compared to state-of-the-art work. Zicong Hong, Song Guo 0001, Sifu Luo, Wuhui Chen, Roger Wattenhofer, Yue Yu 0001 |
EuroSys | 1 |
| 2024 | Porygon: Scaling Blockchain via 3D ParallelismabstractRecently, stateless blockchains have been proposed to alleviate the storage overhead for nodes. A stateless blockchain achieves storage-consensus parallelism, where storage workloads are offloaded from on-chain consensus, enabling more resource-constraint nodes to participate in the consensus. However, existing stateless blockchains still suffer from limited throughput. In this paper, we present Porygon, a novel stateless blockchain with three-dimensional (3D) parallelism. First, Porygon separates the storage and consensus of transactions as the stateless blockchain, achieving the storage-consensus parallelism. This first-dimensional parallelism divides the processing of transactions into several stages and scales the network by supporting more nodes in the system. Based on such a design, we then propose a pipeline mechanism to achieve second-dimensional inter-block parallelism, where relevant stages of processing transactions are pipelined efficiently, thereby reducing transaction latency. Finally, Porygon presents a sharding mechanism to achieve third-dimensional inner-block parallelism. By sharding the executions of transactions of a block and adopting a lightweight cross-shard coordination mechanism, Porygon can effectively execute both intra-shard and cross-shard transactions, consequently achieving outstanding transaction throughput. We evaluate the performance of Porygon by extensive experiments on an implemented prototype and large-scale simulations. Compared with existing blockchains, Porygon boosts throughput by up to 20x, reduces network usage by more than 50%, and simultaneously requires only 5MB of storage consumption per node. Wuhui Chen, Ding Xia, Zhongteng Cai, Hongning Dai, Zicong Hong, Junyuan Liang, Zibin Zheng |
ICDE | 6 |
| 2024 | Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language ModelsabstractFine-tuning the learnable prompt for a pre-trained vision-language model (VLM), such as CLIP, has demonstrated exceptional efficiency in adapting to a broad range of downstream tasks. Existing prompt tuning methods for VLMs do not distinguish spurious features introduced by biased training data from invariant features, and employ a uniform alignment process when adapting to unseen target domains. This can impair the cross-modal feature alignment when the testing data significantly deviate from the distribution of the training data, resulting in a poor out-of-distribution (OOD) generalization performance. In this paper, we reveal that the prompt tuning failure in such OOD scenarios can be attribute to the undesired alignment between the textual and the spurious feature. As a solution, we propose **CoOPood**, a fine-grained prompt tuning method that can discern the causal features and deliberately align the text modality with the invariant feature. Specifically, we design two independent contrastive phases using two lightweight projection layers during the alignment, each with different objectives: 1) pulling the text embedding closer to invariant image embedding and 2) pushing text embedding away from spurious image embedding. We have illustrated that **CoOPood** can serve as a general framework for VLMs and can be seamlessly integrated with existing prompt tuning methods. Extensive experiments on various OOD datasets demonstrate the performance superiority over state-of-the-art methods. Jie Zhang 0076, Xiaosong Ma, Song Guo 0001, Peng Li 0017, Wenchao Xu 0001, Xueyang Tang, Zicong Hong |
ICML | 7 |
| 2024 | OTAS: An Elastic Transformer Serving System via Token AdaptationabstractTransformer model empowered architectures have become a pillar of cloud services that keeps reshaping our society. However, the dynamic query loads and heterogeneous user requirements severely challenge current transformer serving systems, which rely on pre-training multiple variants of a foundation model, i.e., with different sizes, to accommodate varying service demands. Unfortunately, such a mechanism is unsuitable for large transformer models due to the additional training costs and excessive I/O delay. In this paper, we introduce OTAS, the first elastic serving system specially tailored for transformer models by exploring lightweight token management. We develop a novel idea called token adaptation that adds prompting tokens to improve accuracy and removes redundant tokens to accelerate inference. To cope with fluctuating query loads and diverse user requests, we enhance OTAS with application-aware selective batching and online token adaptation. OTAS first batches incoming queries with similar service-level objectives to improve the ingress throughput. Then, to strike a tradeoff between the overhead of token increment and the potentials for accuracy improvement, OTAS adaptively adjusts the token execution strategy by solving an optimization problem. We implement and evaluate a prototype of OTAS with multiple datasets, which show that OTAS improves the system utility by at least 18.2%. Wenchao Xu 0001, Zicong Hong, Song Guo 0001, Haozhao Wang, Jie Zhang 0076, Deze Zeng |
INFOCOM | 3 |
| 2024 | Front-running Attack in Sharded Blockchains and Fair Cross-shard Consensus
Wuhui Chen, Sifu Luo, Tiantian Gong, Zicong Hong, Aniket Kate |
NDSS | 5 |
| 2024 | Towards Safe Concept Transfer of Multi-Modal Diffusion via Causal Representation EditingabstractRecent advancements in vision-language-to-image (VL2I) diffusion generation have made significant progress. While generating images from broad vision-language inputs holds promise, it also raises concerns about potential misuse, such as copying artistic styles without permission, which could have legal and social consequences. Therefore, it's crucial to establish governance frameworks to ensure ethical and copyright integrity, especially with widely used diffusion models. To address these issues, researchers have explored various approaches, such as dataset filtering, adversarial perturbations, machine unlearning, and inference-time refusals. However, these methods often lack either scalability or effectiveness. In response, we propose a new framework called causal representation editing (CRE), which extends representation editing from large language models (LLMs) to diffusion-based models. CRE enhances the efficiency and flexibility of safe content generation by intervening at diffusion timesteps causally linked to unsafe concepts. This allows for precise removal of harmful content while preserving acceptable content quality, demonstrating superior effectiveness, precision and scalability compared to existing methods. CRE can handle complex scenarios, including incomplete or blurred representations of unsafe concepts, offering a promising solution to challenges in managing harmful content generation in diffusion-based models. Peiran Dong, Song Guo 0001, Jie Zhang 0076, Zicong Hong |
NeurIPS | 6 |
| 2024 | Poisoning Attack on Federated Knowledge Graph EmbeddingabstractFederated Knowledge Graph Embedding (FKGE) is an emerging collaborative learning technique for deriving expressive representations (i.e., embeddings) from client-maintained distributed knowledge graphs (KGs). However, poisoning attacks in FKGE, which lead to biased decisions by downstream applications, remain unexplored. This paper is the first work to systematize the risks of FKGE poisoning attacks, from which we develop a novel framework for poisoning attacks that force the victim client to predict specific false facts. Unlike centralized KGEs, FKGE maintains KGs locally, making direct injection of poisoned data challenging. Instead, attackers must create poisoned data without access to the victim's KG and inject it indirectly through FKGE aggregation. Specifically, to create poisoned data, the attacker first infers the targeted relations in the victim's local KG via a new KG component inference attack. Then, to accurately mislead the victim's embeddings via aggregation, the attacker locally trains a shadow model using the poisoned data and uses an optimized dynamic poisoning scheme to adjust the model and generate progressive poisoned updates. Our experimental results demonstrate the attack's effectiveness, achieving a remarkable success rate on various KGE models (e.g., 100% on TransE with WN18RR) while keeping the original task's performance nearly unchanged. Enyuan Zhou, Song Guo 0001, Zhixiu Ma, Zicong Hong, Tao Guo 0004, Peiran Dong |
WWW | 4 |
| 2024 | Efficient Execution of Arbitrarily Complex Cross-Shard Contracts for Blockchain ShardingabstractSharding is a promising solution to enhance the scalability of blockchain. However, previous sharding systems adopt the lock-based cross-shard protocol to exclusively handle one-shot cross-shard transactions, leading to low-efficiency executions and unavailable calls when handling complex cross-shard contracts that introduce multi-shot cross-shard transactions to invoke multiple contracts managed by different shards.In this paper, we aim to enable efficient execution of arbitrarily complex cross-shard contracts in blockchain sharding systems. First, we perform a calling-flow analysis on Ethereum contracts with more than 180 million real-world transactions and find that about 30% transactions invoke complex contracts. Then, motivated by the properties of these complex contracts, we propose an off-chain execution model, called ShardCon, to achieve efficient executions for complex cross-shard contracts by decoupling the contract execution from the cross-shard consensus. Next, we introduce a cross-shard contract execution engine and a contract-driven deployment rule to the overheads introduced by off-chain executions. Moreover, to adapt to the multi-chain property of a sharding system, we introduce an off-chain state atomic commit protocol. Finally, we implement a prototype and evaluate it with concrete cross-shard contracts, showing that ShardCon can achieve more than 10x increase in throughput and 2x decrease in confirmation latency than the state-of-the-art sharding systems. Wuhui Chen, Zicong Hong, Gang Xiao 0003, Linlin Du, Zibin Zheng |
IEEE Trans. Computers | 3 |
| 2024 | AoI-Aware Service Provisioning in Edge Computing for Digital Twin Network Slicing RequestsabstractDigital twins are poised to enter our lives with Industry 4.0. The Digital Twin Network (DTN) paradigm is projected to deliver upon the promise of efficient collaboration among digital twins to enable complicated and systematic services across many domains, through depicting an overall picture of a group of physical objects. To achieve timely data processing of digital twins, Mobile Edge Computing (MEC) shifts the computational power towards the network edge, and network slicing is well-suited to bundle heterogeneous physical resources to build logical networks based on edge servers for accommodating DTNs. In light of this, in this paper we investigate DTN slicing-enabled service provisioning in MEC, where each DTN slice consists of one master digital twin and a set of worker digital twins, and each worker digital twin is synchronized through collecting data from a respective object periodically. The master digital twin aggregates the processed data from worker digital twins to model the DTN continuously for user query services, whilst meeting delay requirements of users. We capture the utility gain of a DTN slicing request based on the DTN model quality at its master digital twin that is impacted by the Age of Information (AoI), and we focus on two novel optimization problems: the utility maximization problem for a single DTN slicing request, and the dynamic utility maximization problem for multiple DTN slicing requests. We propose an approximation algorithm for the former, and an online algorithm with a provable competitive ratio for the latter. We also evaluate the performance of the proposed algorithms through simulations. Experimental results demonstrate that the proposed algorithms are promising, outperforming their counterparts by at least 10.2%. Jing Li 0093, Song Guo 0001, Weifa Liang, Jianping Wang 0001, Quan Chen 0003, Zicong Hong, Zichuan Xu, Wenzheng Xu, Bin Xiao 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Chiron: A Robustness-Aware Incentive Scheme for Edge Learning via Hierarchical Reinforcement LearningabstractOver the past few years, edge learning has achieved significant success in mobile edge networks. Few works have designed incentive mechanism that motivates edge nodes to participate in edge learning. However, most existing works only consider myopic optimization and assume that all edge nodes are honest, which lacks long-term sustainability and the final performance assurance. In this paper, we propose Chiron, an incentive-driven Byzantine-resistant long-term mechanism based on hierarchical reinforcement learning (HRL). First, our optimization goal includes both learning-algorithm performance criteria (i.e., global accuracy) and systematical criteria (i.e., resource consumption), which aim to improve the edge learning performance under a given resource budget. Second, we propose a three-layer HRL architecture to handle long-term optimization, short-term optimization, and byzantine resistance, respectively. Finally, we conduct experiments on various edge learning tasks to demonstrate the superiority of the proposed approach. Specifically, our system can successfully exclude malicious nodes and lazy nodes out of the edge learning participation and achieves 14.96% higher accuracy and 12.66% higher total utility than the state-of-the-art methods under the same budget limit. Yi Liu 0057, Song Guo 0001, Yufeng Zhan, Leijie Wu, Zicong Hong, Qihua Zhou |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Long-Term Adaptive VCG Auction Mechanism for Sustainable Federated Learning With Periodical Client ShiftingabstractFederated Learning (FL) system needs to incentivize clients since they may be reluctant to participate in the resource consuming process. Existing incentive mechanisms fail to construct a sustainable environment for the long-term development of FL system: 1) They seldom focus on system economic properties (e.g., social welfare, individual rationality, and incentive compatibility) to guarantee client attraction. 2) Current online auction modeling methods divide the whole continual process into multiple independent rounds and solve them one-by-one, which breaks the correlation between each round. Besides, the inherent characteristics of FL system (model-agnostic and privacy-sensitive) also prevent it from the optimal strategy by precise mathematical analysis. 3) Current system modelings ignore the practical problem of periodical client shifting, which cannot adaptively update its strategy to handle system dynamics. To overcome the above challenges, this paper proposes a long-term adaptive Vickrey–Clarke–Groves (VCG) auction mechanism for FL system, which incorporate a multi-branch deep reinforcement learning (DRL) algorithm. First, VCG auction is the only one that can simultaneously guarantee all crucial economic properties. Second, we extend the economic properties to long-term forms and apply the experience-driven DRL algorithm to directly obtain long-term optimal strategy, without any prior system knowledge. Third, we reconstruct a multi-branch DRL network to accommodate periodical client shifting by adaptive decision head switching for different time periods. Finally, we theoretically prove he extended economic properties (i.e., IC) and conduct extensive experiments on several real-world datasets. Compared with state-of-the-art approaches, the long-term social welfare of FL system increases by 36% with a 37% reduction in payment. Besides, the multi-branch network can adaptively handle periodical client shifting on the timeline. Leijie Wu, Song Guo 0001, Zicong Hong, Yi Liu 0057, Wenchao Xu 0001, Yufeng Zhan |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | MoltDB: Accelerating Blockchain via Ancient State SegregationabstractBlockchain store states in Log-Structured Merge (LSM) tree-based database. Due to blockchain traceability, the growing ancient states are inevitably stored in the databases. Unfortunately, by default, this process mixescurrentandancientstates in the data layout, increasing unnecessary disk I/O access and slowing transaction execution. This paper proposes MoltDB, a scalable LSM-based database for efficient transaction execution through a novel idea ofancient state segregation, i.e., to segregate current and ancient states in the data layout. However, the frequently generated and uncertainly accessed characteristics of ancient states make the segregation challenging. Thus, we develop an “extract-compact” mechanism to batch extraction process for frequently generated ancient states and the LSM compaction process to relieve additional disk I/O overhead. Moreover, we design an adaptive LSM-based storage for the uncertainly accessed ancient states extracted for on-demand access. We implement MoltDB as a database engine compatible with many mainstream blockchains and integrate it into Ethereum for evaluation. Experimental results show that MoltDB achieves 1.3 × transaction throughput and 30% disk I/O latency savings over the state-of-the-art works. Junyuan Liang, Wuhui Chen, Zicong Hong, Haogang Zhu, Wangjie Qiu, Zibin Zheng |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | Prophet: Conflict-Free Sharding Blockchain via Byzantine-Tolerant Deterministic OrderingabstractSharding scales throughput by splitting blockchain nodes into parallel groups. However, different shards’ independent and random scheduling for cross-shard transactions results in numerous conflicts and aborts, since cross-shard transactions from different shards may access the same account. A deterministic ordering can eliminate conflicts by determining a global order for transactions before processing, as proved in the database field. Unfortunately, due to the intertwining of the Byzantine environment and information isolation among shards, there is no trusted party able to predetermine such an order for cross-shard transactions. To tackle this challenge, this paper proposes Prophet, a conflict-free sharding blockchain based on Byzantine-tolerant deterministic ordering. It first depends on untrusted self-organizing coalitions of nodes from different shards to pre-execute cross-shard transactions for prerequisite information about ordering. It then determines a trusted global order based on stateless ordering and post-verification for pre-executed results, through shard cooperation. Following the order, the shards thus orderly execute and commit transactions without conflicts. Prophet orchestrates the pre-execution, ordering, and execution processes in the sharding consensus for minimal overhead. We rigorously prove the determinism and serializability of transactions under the Byzantine and sharded environment. An evaluation of our prototype shows that Prophet improves the throughput by 3.11× and achieves nearly no aborts on 1 million Ethereum transactions compared with state-of-the-art sharding. Zicong Hong, Song Guo 0001, Enyuan Zhou, Wuhui Chen, Jinwen Liang, Jie Zhang 0076, Albert Y. Zomaya |
INFOCOM | 1 |
| 2023 | GriDB: Scaling Blockchain Database via Sharding and Off-Chain Cross-Shard MechanismabstractBlockchain databases have attracted widespread attention but suffer from poor scalability due to underlying non-scalable blockchains. While blockchain sharding is necessary for a scalable blockchain database, it poses a new challenge named on-chain cross-shard database services. Each cross-shard database service (e.g., cross-shard queries or inter-shard load balancing) involves massive cross-shard data exchanges, while the existing cross-shard mechanisms need to process each cross-shard data exchange via the consensus of all nodes in the related shards (i.e., on-chain) to resist a Byzantine environment of blockchain, which eliminates sharding benefits. To tackle the challenge, this paper presents GriDB, the first scalable blockchain database, by designing a novel off-chain cross-shard mechanism for efficient cross-shard database services. Borrowing the idea of off-chain payments, GriDB delegates massive cross-shard data exchange to a few nodes, each of which is randomly picked from a different shard. Considering the Byzantine environment, the untrusted delegates cooperate to generate succinct proof for cross-shard data exchanges, while the consensus is only responsible for the low-cost proof verification. However, different from payments, the database services' verification has more requirements (e.g., completeness, correctness, freshness, and availability); thus, we introduce several new authenticated data structures (ADS). Particularly, we utilize consensus to extend the threat model and reduce the complexity of traditional accumulator-based ADS for verifiable cross-shard queries with a rich set of relational operators. Moreover, we study the necessity of inter-shard load balancing for a scalable blockchain database and design an off-chain and live approach for both efficiency and availability during balancing. An evaluation of our prototype shows the performance of GriDB in terms of scalability in workloads with queries and updates. Zicong Hong, Song Guo 0001, Enyuan Zhou, Wuhui Chen, Huawei Huang, Albert Y. Zomaya |
Proc. VLDB Endow. | 1 |
| 2023 | VeriDKG: A Verifiable SPARQL Query Engine for Decentralized Knowledge GraphsabstractThe ability to decentralize knowledge graphs (KG) is important to exploit the full potential of the Semantic Web and realize the Web 3.0 vision. However, decentralization also renders KGs more prone to attacks with adverse effects on data integrity and query verifiability. While existing studies focus on ensuring data integrity, how to ensure query verifiability - thus guarding against incorrect, incomplete, or outdated query results - remains unsolved. We propose VeriDKG, the first SPARQL query engine for decentralized knowledge graphs (DKG) that offers both data integrity and query verifiability guarantees. The core of VeriDKG is the RGB-Trie, a new blockchain-maintained authenticated data structure (ADS) facilitating correctness proofs for SPARQL query results. VeriDKG enables verifiability of subqueries by gathering global index information on subgraphs using the RGB-Trie, which is implemented as a new variant of the Merkle prefix tree with an RGB color model. To enable verifiability of the final query result, the RGB-Trie is integrated with a cryptographic accumulator to support verifiable aggregation operations. A rigorous analysis of query verifiability in VeriDKG is presented, along with evidence from an extensive experimental study demonstrating its state-of-the-art query performance on the largeRDFbench benchmark. Enyuan Zhou, Song Guo 0001, Zicong Hong, Christian S. Jensen, Yang Xiao 0014, Dalin Zhang 0001, Jinwen Liang, Qingqi Pei |
Proc. VLDB Endow. | 3 |
| 2023 | MSTDB: A Hybrid Storage-Empowered Scalable Semantic Blockchain DatabaseabstractBlockchain has been regarded as a trusted carrier for distributed data storage. With large volumes of valuable data stored on blockchain, data query has become a major requirement. However, the existing blockchains do not provide efficient query functionality because of their deep-rooted chain structure. Blockchain database is a new direction that constructs index on top of blockchain to provide rich query functionalities. The existing works are either insecure because the query process separates from the blockchain consensus, or inscalable because all the data needs to be stored in the block. In this paper, we propose a novel semantic blockchain database called MSTDB. We design a hybrid on/off chain blockchain storage architecture in which the majority of blockchain storage is offloaded to the off-chain storage and a novel index structure named Merkle Semantic Trie (MST) is designed to be a secure and semantic bridge between on- and off-chain. Based on MST, MSTDB provides a variety of semantic query functions including multi-keyword query, range query, Top-K query, and cross-chain query. To improve the performance further, we design some index compression and query preprocessing techniques for MSTDB. Extensive experiments demonstrate the effectiveness and efficiency of our blockchain database. Enyuan Zhou, Zicong Hong, Yang Xiao 0014, Dongxiao Zhao, Qingqi Pei, Song Guo 0001, Rajendra Akerkar |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Benzene: Scaling Blockchain With Cooperation-Based ShardingabstractSharding has been considered as a prominent approach to enhance the limited performance of blockchain. However, most sharding systems leverage a non-cooperative design, which lowers the fault tolerance resilience due to the decreased mining power as the consensus execution is limited to each separated shard. To this end, we present Benzene, a novel sharding system that enhances the performance by cooperation-based sharding while defending the per-shard security. First, we establish a double-chain architecture for function decoupling. This architecture separates transaction-recording functions from consensus-execution functions, thereby enabling the cross-shard cooperation during consensus execution while preserving the concurrency nature of sharding. Second, we design a cross-shard block verification mechanism leveraging Trusted Execution Environment (TEE), via which miners can verify blocks from other shards during the cooperation process with the minimized overheads. Finally, we design a voting-based consensus protocol for cross-shard cooperation. Transactions in each shard are confirmed by all shards that simultaneously cast votes, consequently achieving an enhanced fault tolerance and lowering the confirmation latency. We implement Benzene and conduct both prototype experiments and large-scale simulations to evaluate the performance of Benzene. Results show that Benzene achieves superior performance than existing sharding/non-sharding blockchain protocols. In particular, Benzene achieves a linearly-improved throughput with the increased number of shards (e.g., 32,370 transactions per second with 50 shards) and maintains a lower confirmation latency than Bitcoin (with more than 50 shards). Meanwhile, Benzene maintains a fixed fault tolerance at 1/3 even with the increased number of shards. Zhongteng Cai, Junyuan Liang, Wuhui Chen, Zicong Hong, Hongning Dai, Zibin Zheng |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | SocialChain: Decoupling Social Data and Applications to Return Your Data OwnershipabstractSocial data produced from widely emerged social media activities are expected to promote information dissemination and engagement, or even make business intelligence more powerful. However, the recent increase in social media incidents of illegal surveillance and data breaches raises questions about the current data ownership model, in which centralized applications collect and control large amounts of user data. In this paper, we present SocialChain, which is a decentralized social data storage and sharing system based on blockchain that decouples user data and social applications to return data ownership to the user. We adopt Personal Data Store to extend off-chain storage for the social data, set up an identity establishment mechanism that can support WebID-based authentication functions using a unique identity assignment (i.e., WebID) as well as certificateless cryptography, and design a general framework that leverages smart contracts to help securely store and share social data in an automated manner. We develop a software prototype based on Ethereum and conduct case studies to test the effects of the adopted techniques on the performance. Experimental results show that SocialChain can provide easy-to-use interfaces while introducing relatively low latency, cost, and overhead and that it can support real-world social media applications. Ting Cai 0002, Zicong Hong, Wuhui Chen, Zibin Zheng, Yang Yu 0027 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Proactive look-ahead control of transaction flows for high-throughput payment channel networkabstractBlockchain technology has gained popularity owing to the success of cryptocurrencies such as Bitcoin and Ethereum. Nonetheless, the scalability challenge largely limits its applications in many real-world scenarios. Off-chain payment channel networks (PCNs) have recently emerged as a promising solution by conducting payments through off-chain channels. However, the throughput of current PCNs does not yet meet the growing demands of large-scale systems because: 1) most PCN systems only focus on maximizing the instantaneous throughput while failing to consider network dynamics in a long-term perspective; 2) transactions are re-actively routed in PCNs, in which intermediate nodes only passively forward every incoming transaction. These limitations of existing PCNs inevitably lead to channel imbalance and the failure of routing subsequent transactions. To address these challenges, we propose a novel proactive look-ahead algorithm (PLAC) that controls transaction flows from a long-term perspective and proactively prevents channel imbalance. In particular, we first conduct a measurement study on two real-world PCNs to explore their characteristics in terms of transaction distribution and topology. On that basis, we propose PLAC based on deep reinforcement learning (DRL), which directly learns the system dynamics from historical interactions of PCNs and aims at maximizing the long-term throughput. Furthermore, we develop a novel graph convolutional network-based model for PLAC, which extracts the inter-dependency between PCN nodes to consequently boost the performance. Extensive evaluations on real-world datasets show that PLAC improves state-of-the-art PCN routing schemes w.r.t the long-term throughput from 6.6% to 34.9%. Wuhui Chen, Xiaoyu Qiu, Zicong Hong, Zibin Zheng, Hongning Dai |
SoCC | 3 |
| 2022 | Cycle: Sustainable Off-Chain Payment Channel Network with Asynchronous RebalancingabstractPayment channel network (PCN) is a promising off-chain technology for blockchain scalability, but it suffers from poor sustainability in practice. In other words, due to the imbalanced transfer in channels, the balance in one direction of channels gradually becomes exhausted until the PCN is rebalanced via a consensus-based rebalancing protocol, during which the involved channels must be suspended. This paper presents Cycle, the first off-chain protocol for a sustainable PCN. It not only keeps the PCN at a balanced level consistently but also avoids the channel freeze incurred by the rebalancing protocol, leading to minimum failed payments and sustained PCN service, respectively. Cycle achieves these benefits based on a novel idea of asynchronous rebalancing. During the normal off-chain running, the participants share the information about their payments and asynchronously rebalance the PCN following the principle that payments along circular channels can cancel each other out. To guarantee security, the protocol resolves the disputes resulting from network latency or malicious participants by a message mechanism for synchronization and a smart contract for arbitration. Moreover, to address the privacy concern during the information sharing, a truncated Laplace mechanism is designed to achieve differential privacy. Finally, we provide a proof-of-concept implementation in Ethereum, over which a real data-based simulation shows that Cycle satisfies 31% more payments than the state-of-the-art technique. Zicong Hong, Song Guo 0001, Rui Zhang 0080, Peng Li 0017, Yufeng Zhan, Wuhui Chen |
DSN | 1 |
| 2022 | Sustainable Federated Learning with Long-term Online VCG Auction MechanismabstractFederated learning (FL) clients may be reluctant to participate in the energy-consuming FL unless they are incentivized. Existing incentive mechanisms seldom consider the economic properties, e.g., social welfare, individual rationality and incentive compatibility, which significantly limits the sustainability of FL to attract more clients. The Vickrey–Clarke–Groves (VCG) auction is an ideal mechanism for simultaneously guaranteeing all crucial economic properties to maximize social welfare. However, VCG auction cannot be applied directly to FL scenarios due to the following challenges: 1) It requires precise analytical derivation of the optimal strategy, which is unavailable due to the inherent model-unknown and privacy-sensitive characteristics of FL. 2) Current auction modeling decomposes the entire process into multiple independent rounds and solves them one-by-one, which breaks the successive correlation between rounds in the long-term training process of FL. To overcome these challenges, this paper presents a long-term online VCG auction mechanism for FL that employs an experience-driven deep reinforcement learning algorithm to obtain the optimal strategy. Besides, we extend long-term forms of the crucial economic properties for the successive FL process. Furthermore, knowledge transfer is applied to reduce the excessive training overhead arising from the VCG payment rules. By exploiting the environmental similarity among sub-auctions, we develop the strategy sharing to significantly cut the training time by half. Finally, we theoretically prove the extended economic properties and conduct extensive experiments on multiple real-world datasets. Compared with state-of-the-art approaches, the long-term social welfare of FL increases by 36% with a 37% reduction in payment. Leijie Wu, Song Guo 0001, Yi Liu 0057, Zicong Hong, Yufeng Zhan, Wenchao Xu 0001 |
ICDCS | 4 |
| 2022 | Scaling Blockchain via Layered ShardingabstractAs a promising solution to blockchain scalability, sharding divides blockchain nodes into small groups called shards, splitting the workload. Existing works for sharding, however, are limited by cross-shard transactions, since they need to split each cross-shard transaction into multiple sub-transactions, each of which costs a consensus round to commit. In this paper, we introduce PYRAMID, a novel sharding system based on the idea of layered sharding. In PYRAMID, the nodes with better hardware are allowed to participate in multiple shards and store the blockchains of these shards thus they can validate and execute the cross-shard transactions without splitting. Next, to commit the cross-shard transactions with consistency among the related shards, we design a cooperative cross-shard consensus based on collective signature-based inter-shard collaboration. Furthermore, we present an optimization framework to compute an optimal layered sharding strategy maximizing the transaction throughput with the constraint of system security and node resource. Finally, we implement a prototype for PYRAMID based on Ethereum and the experimental results reveal the efficiency of PYRAMID in terms of performance and scalability, especially in workloads with a high percentage of cross-shard transactions. PYRAMID improves the throughput by up to 3.2 times compared with the state-of-the-art works and achieves about 3821 transaction per seconds for 20 shards. Zicong Hong, Song Guo 0001, Peng Li 0017 |
IEEE J. Sel. Areas Commun. | 1 |
| 2021 | Incentive-Driven Long-term Optimization for Edge Learning by Hierarchical Reinforcement MechanismabstractEdge Learning is an emerging distributed machine learning in mobile edge network. Limited works have designed mechanisms to incentivize edge nodes to participate in edge learning. However, their mechanisms only consider myopia optimization on resource consumption, which results in the lack of learning algorithm performance guarantee and longterm sustainability. In this paper, we propose Chiron, an incentive-driven long-term mechanism for edge learning based on hierarchical deep reinforcement learning. First, our optimization goal combines learning-algorithms metric (i.e., model accuracy) with system metric (i.e., learning time, and resource consumption), which can improve edge learning quality under a fixed training budget. Second, we present a two-layer H-DRL design with exterior and inner agents to achieve both long-term and short-term optimization for edge learning, respectively. Finally, experiments on three different real-world datasets are conducted to demonstrate the superiority of our proposed approach. In particular, compared with the state-of-the-art methods under the same budget constraint, the final global model accuracy and time efficiency can be increased by 6.5 % and 39 %, respectively. Our implementation is available at https://github.com/Joey61Liuyi/Chiron. Yi Liu 0057, Leijie Wu, Yufeng Zhan, Song Guo 0001, Zicong Hong |
ICDCS | 5 |
| 2021 | Pyramid: A Layered Sharding Blockchain SystemabstractSharding can significantly improve the blockchain scalability, by dividing nodes into small groups called shards that can handle transactions in parallel. However, all existing sharding systems adopt complete sharding, i.e., shards are isolated. It raises additional overhead to guarantee the atomicity and consistency of cross-shard transactions and seriously degrades the sharding performance. In this paper, we present Pyramid, the first layered sharding blockchain system, in which some shards can store the full records of multiple shards thus the cross-shard transactions can be processed and validated in these shards internally. When committing cross-shard transactions, to achieve consistency among the related shards, a layered sharding consensus based on the collaboration among several shards is presented. Compared with complete sharding in which each cross-shard transaction is split into multiple sub-transactions and cost multiple consensus rounds to commit, the layered sharding consensus can commit cross-shard transactions in one round. Furthermore, the security, scalability, and performance of layered sharding with different sharding structures are theoretically analyzed. Finally, we implement a prototype for Pyramid and its evaluation results illustrate that compared with the state-of-the-art complete sharding systems, Pyramid can improve the transaction throughput by 2.95 times in a system with 17 shards and 3500 nodes. Zicong Hong, Song Guo 0001, Peng Li 0017, Wuhui Chen |
INFOCOM | 1 |
| 2021 | Hierarchical Pricing Mechanism With Financial Stability for Decentralized Crowdsourcing: A Smart Contract ApproachabstractSoftware crowdsourcing is an emerging approach to software engineering with great potential for the subdivision and assignment of large-scale tasks. However, because of the centralization of the traditional crowdsourcing platform, information disclosure and nontransparent accounting may be difficult to avoid. To address this issue, we first introduce a novel blockchain-enabled crowdsourcing platform that integrates the functions of task assignment and resource lending via two dedicated smart contracts. Second, to ensure financial stability in the blockchain-enabled market and to match the difficulty of the received tasks with the ability of the workers, we design a dynamic, hierarchical pricing mechanism based on economic modeling methods and heterogeneous agent theory. With this mechanism, the market is divided dynamically into multiple levels according to the remuneration of the customers' offer and the market value of the workers' resources. Additional constraints are proposed to avoid possible malicious trading behavior from workers in the resource lending process. We prove theoretically the rationality of our model and demonstrate the dynamics of the model. We show that the market price and demand can be convergent and test the cost of executing the two smart contracts. Finally, extensive experimental results demonstrate the correctness and feasibility of the platform and confirm that the hierarchical pricing mechanism can maintain the stability of the market. Weikun Zhang, Zicong Hong, Wuhui Chen |
IEEE Internet Things J. | 2 |
| 2020 | SkyChain: A Deep Reinforcement Learning-Empowered Dynamic Blockchain Sharding SystemabstractTo overcome the limitations on the scalability of current blockchain systems, sharding is widely considered as a promising solution that divides the network into multiple disjoint groups processing transactions in parallel to improve throughput while decreasing the overhead of communication, computation, and storage. However, most existing blockchain sharding systems adopt a static sharding policy that cannot efficiently deal with the dynamic environment in the blockchain system, i.e., joining and leaving of nodes, and malicious attack. This paper presents SkyChain, a novel dynamic sharding-based blockchain framework to achieve a good balance between performance and security without compromising scalability under the dynamic environment. We first propose an adaptive ledger protocol to guarantee that the ledgers can merge or split efficiently based on the dynamic sharding policy. Then, to optimize the sharding policy under dynamic environment with high dimensional system states, a deep reinforcement learning-based sharding approach has been proposed, the goals of which include: 1) building a framework to evaluate the blockchain sharding systems from the aspects of performance and security; 2) adjusting the re-sharding interval, shard number and block size to maintain a long-term balance of the system’s performance and security. Experimental results show that SkyChain can effectively improve the performance and security of the sharding system without compromising scalability under the dynamic environment in the blockchain system. Zicong Hong, Xiaoyu Qiu, Yufeng Zhan, Song Guo 0001, Wuhui Chen |
ICPP | 2 |
| 2020 | Smart Contract-based Hierarchical Auction Mechanism for Edge Computing in Blockchain-empowered IoTabstractEdge computing is a promising paradigm to expand the capability of Internet of Things (IoT) devices by computation offloading. To establish a distributed ledger to provide a secure and trusted environment for the resource allocation between edge servers and IoT devices, the emerging blockchain technology has attracted a lot of attention recently. However, in practice, edge resource allocation in IoT devices often involves multi-layer structures, which poses a challenge due to information incompleteness among different layers. Moreover, how to design a suitable and efficient blockchain framework for hierarchical resource allocation markets is a critical issue. In this paper, we apply blockchain to propose a secure and efficient hierarchical resource allocation framework for edge computing. First, we study the edge computing resource allocation problem in the hierarchical market of IoT devices, in which the IoT devices beyond the coverage of Access Points can participate in the resource allocation through middlemen. To solve the problem, a smart contract-based hierarchical auction mechanism is developed. The edge computing resources allocated in the top market can be continually reallocated to the sub-markets based on the mechanism, which then leads an efficient solution that maximizes the social welfare of the whole participants. Moreover, the mechanism is implemented as a smart contract in the blockchain, which enforces the rule of the hierarchical auction in a non-deniable and automated manner. Finally, the extensive simulations demonstrate the correctness and performance of the proposed mechanism. Zetao Yang, Zicong Hong, Shenghui Li, Wuhui Chen |
WoWMoM | 3 |
| 2019 | Cooperative and Distributed Computation Offloading for Blockchain-Empowered Industrial Internet of ThingsabstractOffloading computation-intensive blockchain mining tasks to the edge servers (ESs) is a promising solution for blockchain-empowered Industrial Internet of Things (IIoT) because the computing capabilities in IIoT are usually limited, whereas the blockchain mining tasks are computationally intensive. However, the computation offloading solutions for data processing tasks and for blockchain mining tasks have been studied separately. Moreover, most of the existing solutions for offloading assume that all IIoT devices can directly connect to the ESs or cloud data centers. To address these issues, in this paper, we propose a multihop cooperative and distributed computation offloading algorithm that considers the data processing tasks and the mining tasks together for blockchain-empowered IIoT. First, we study the multihop computation offloading problem for both the data processing tasks and the mining tasks to minimize the economic cost of IIoT devices. Second, we formulate the offloading problem as a potential game in which the IIoT devices can make their decisions autonomously and prove the existence of Nash equilibrium (NE) for the game. Third, we design an efficient distributed algorithm based on exchanging messages between IIoT devices to achieve the NE with low computational complexity. Lastly, our experimental results demonstrate that our distributed algorithm scales well as the number of IIoT devices increases and has the minimum system cost compared with other approaches. Wuhui Chen, Zhen Zhang 0022, Zicong Hong, Chuan Chen 0001, Jiajing Wu, Sabita Maharjan, Zibin Zheng, Yan Zhang 0002 |
IEEE Internet Things J. | 3 |
| 2019 | Joint Computation Offloading and Coin Loaning for Blockchain-Empowered Mobile-Edge ComputingabstractThe blockchain-empowered mobile-edge computing (MEC) is a promising solution for enhancing the computation capabilities of mobile equipments (MEs) to process computation-intensive tasks such as the real-time data processing tasks and mining tasks. However, because of the “cold start” and “long return” problems, efficient computation offloading cannot be achieved in blockchain-empowered MEC because the MEs do not always have enough coins to afford the offloading service cost. In this article, we study the joint computation-offloading and coin-loaning problem for blockchain-empowered MEC to minimize the total cost of all MEs. We introduce the banks that can provide loan services to the MEs to address the above two issues. We formulate the problem as a noncooperative game to model the competitions between the myopic MEs. By using a potential game method, we prove the existence of a pure-strategy Nash equilibrium (NE) and design a distributed algorithm to achieve the NE point with low computational complexity. We also provide an upper bound on the price of anarchy of the game by theoretical proof. Besides, two smart contracts are designed to automatically perform the computing resource trading and coin loaning processes. Lastly, our simulation results show that our proposed algorithm can significantly reduce the total cost of all MEs, has better performance compared with other solutions, and scales well as the number of MEs increases. Moreover, the financial cost for executing the two smart contracts on the Ethereum network is low. Zhen Zhang 0022, Zicong Hong, Wuhui Chen, Zibin Zheng, Xu Chen 0004 |
IEEE Internet Things J. | 2 |
| 2019 | Multi-Hop Cooperative Computation Offloading for Industrial IoT-Edge-Cloud Computing EnvironmentsabstractThe concept of the industrial Internet of things (IIoT) is being widely applied to service provisioning in many domains, including smart healthcare, intelligent transportation, autopilot, and the smart grid. However, because of the IIoT devices' limited onboard resources, supporting resource-intensive applications, such as 3D sensing, navigation, AI processing, and big-data analytics, remains a challenging task. In this paper, we study the multi-hop computation-offloading problem for the IIoT-edge-cloud computing model and adopt a game-theoretic approach to achieving Quality of service (QoS)-aware computation offloading in a distributed manner. First, we study the computation-offloading and communication-routing problems with the goal of minimizing each task's computation time and energy consumption, formulating the joint problem as a potential game in which the IIoT devices determine their computation-offloading strategies. Second, we apply a free-bound mechanism that can ensure a finite improvement path to a Nash equilibrium. Third, we propose a multi-hop cooperative-messaging mechanism and develop two QoS-aware distributed algorithms that can achieve the Nash equilibrium. Our simulation results show that our algorithms offer a stable performance gain for IIoT in various scenarios and scale well as the device size increases. Zicong Hong, Wuhui Chen, Huawei Huang, Song Guo 0001, Zibin Zheng |
IEEE Trans. Parallel Distributed Syst. | 1 |