Jin Dong 0004

dblp:31/507-4 · DBLP profile ↗
← Back
23ranked-venue papers
0as first author
23since 2021 · last 2026
0009-0006-9708-0220ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Security and privacy · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FedSEA-LLaMA: A Secure, Efficient and Adaptive Federated Splitting Framework for Large Language Models
abstract
Private data holds promise for improving LLMs due to its high quality, but its scattered distribution across data silos and the high computational demands of LLMs limit their deployment in federated environments. To address this, the transformer-based federated split models are proposed, which offload most model parameters to the server (or distributed clients) while retaining only a small portion on the client to ensure data privacy. Despite this design, they still face three challenges: 1) Peer-to-peer key encryption struggles to secure transmitted vectors effectively; 2) The auto-regressive nature of LLMs means that federated split learning can only train and infer sequentially, causing high communication overhead; 3) Fixed partition points lack adaptability to downstream tasks. In this paper, we introduce FedSEA-LLaMA, a Secure, Efficient, and Adaptive Federated splitting framework based on LLaMA2. First, we inject Gaussian noise into forward-pass hidden states to enable secure end-to-end vector transmission. Second, we employ attention-mask compression and KV cache collaboration to reduce communication costs, accelerating training and inference. Third, we allow users to dynamically adjust the partition points for input/output blocks based on specific task requirements. Experiments on natural language understanding, summarization, and conversational QA tasks show that FedSEA-LLaMA maintains performance comparable to centralized LLaMA2 and achieves up to 8× speedups in training and inference. Further analysis of privacy attacks and different partition points also demonstrates the effectiveness of FedSEA-LLaMA in security and adaptability.
Zishuai Zhang 0001, Hainan Zhang 0001, Qinnan Zhang, Jin Dong 0004, Yongxin Tong, Zhiming Zheng 0001
AAAI5
2026 FluxZK: Scalable and Efficient Zero-Knowledge Proof Computation via GPU Acceleration
abstract
Zero-knowledge succinct non-interactive arguments of knowledge (zkSNARKs) are a key technology to privacy-preserving applications today. The complexity of proof generation, however, heavily constrains throughput in latency-sensitive environments. The computational burden primarily stems from two fundamental algorithms: Multi-Scalar Multiplication (MSM) and the Number Theoretic Transform (NTT). We propose a series of optimizations for these two kernels, including computation-transfer pipelining, load balancing, and memory access fusion, achieving 1.97 × to 2.16 × proof generation speedup over a state-of-the-art open source GPU acceleration library. Our design also supports out-of-core computation, enabling the generation of large-scale ZKP proofs.
Xinwei Qiang, Liukun Yu, Zhengyi Li 0002, Shixuan Sun, Jingwen Leng, Chen Chen 0067, Jiaping Gui, Zhenzhe Zheng 0001, Jin Dong 0004, Minyi Guo
HPDC10
2026 Alzo: Auto-Tuning with Reinforcement Learning for DAG-based Blockchains
abstract
As critical infrastructure for Web 3.0, DAG-based blockchains promise high throughput for DeFi, IoT, and DApps. However, realizing this potential is challenging, as system performance is dictated by a multitude of interdependent parameters across network, node, and consensus layers. Manual configuration fails to adapt to dynamic workloads, leading to suboptimal performance. We introduce Alzo, a novel auto-tuner that employs hierarchical reinforcement learning (HRL) to navigate this complex configuration space. By decomposing the DAG blockchain's workflow into distinct stages, Alzo's HRL policy learns from stage-level performance metrics to control critical parameters governing consensus, execution, and graph topology in real-time. Furthermore, we employ a shadow-control loop to ensure the safety of all parameter adjustments. Our experiments show that Alzo significantly outperforms other configurations, achieving higher throughput and lower latency under variable workloads with minimal overhead.
Qiuyu Ding, Rongkai Zhang 0005, Qinnan Zhang, Jieyi Long, Mingchao Wan, Jin Dong 0004
WWW8
2026 CodeBC: A more secure large language model for smart contract code generation in blockchain
Lingxiang Wang, Hainan Zhang 0001, Qinnan Zhang, Hongwei Zheng 0003, Jin Dong 0004, Zhiming Zheng 0001
Neurocomputing6
2026 FedPDM: Representation enhanced federated learning with privacy preserving diffusion models
Fuzhen Zhuang, Yiqi Tong, Xiao Zhang 0015, Zhaojun Hu, Jiejie Zhao, Jin Dong 0004
Knowl. Based Syst.7
2026 ENClose: Encrypted Nonlinear Closed-Loop Control Over Fully Homomorphic Encryption
abstract
This work proposes an encrypted controller framework for closed-loop control systems with nonlinear dynamics over fully homomorphic encryption (FHE). Unlike differential privacy and output masking, FHE is a cryptographic primitive that provides assumption-based confidentiality guarantees under standard hardness assumptions. We observe that existing encrypted control frameworks remain largely limited to linear open-loop systems, primarily due to two key challenges: rapid ciphertext noise accumulation in feedback loops and the substantial computational overhead of nonlinear operations. In control systems, feedback is essential for real-time error correction, while nonlinear characteristics are critical for accurately modelling complex system behaviours. To address these challenges, we propose ENClose, a novel encrypted control framework that enables low-latency execution of both feedback control and nonlinear function evaluation. Specifically, ENClose introduces a low-latency homomorphic nonlinear computation framework that accelerates functional bootstrapping (FBS) by combining function segmentation with tree-based encrypted selection. This framework not only mitigates noise accumulation in encrypted feedback loops but also significantly improves the efficiency of FBS under high-precision settings, meeting the computational demands of dynamic control systems. Experimental results show that ENClose achieves a 3× to 20× speedup over state-of-the-art encrypted controllers. We validate ENClose through realworld applications, including multi-vehicle formation, spring–mass–damper control, and anomaly recovery, where the results demonstrate high-precision tracking and successful reconvergence after anomalies.
Song Bian 0001, Yuexiang Jin, Dong Zhao 0004, Yunhao Fu, Haowen Pan, Yi Chen 0012, Bo Zhang 0142, Changrui Ren, Jin Dong 0004, Zhenyu Guan 0002
IEEE Trans. Inf. Forensics Secur.10
2026 HKT-SmartAudit: Distilling Lightweight Models for Smart Contract Auditing
abstract
The rapid growth of blockchain technology has driven the widespread adoption of smart contracts; however, their inherent vulnerabilities have led to significant financial losses. Traditional auditing methods, while essential, struggle to keep pace with the increasing complexity and scale of smart contracts. Large language models (LLMs) offer promising capabilities for automating vulnerability detection, but their adoption is often limited by high computational costs. Although prior work has explored leveraging large models through agents or workflows, relatively little attention has been given to improving the performance of smaller, fine-tuned models—a critical factor for achieving both efficiency and data privacy. In this paper, we introduce HKT-SmartAudit, a framework for developing lightweight models optimized for smart contract auditing. It features a multi-stage knowledge distillation pipeline that integrates classical distillation, external domain knowledge, and reward-guided learning to transfer high-quality insights from large teacher models. A single-task learning strategy is employed to train compact student models that maintain high accuracy and robustness while significantly reducing computational overhead. Experimental results show that our distilled models outperform both commercial tools and larger models in detecting complex vulnerabilities and logical flaws, offering a practical, secure, and scalable solution for smart contract auditing. The source code is available in the GitHub repository1.
Jing Sun 0002, Zijian Zhang 0001, Xianhao Zhang, Meng Li 0006, Yuqiang Sun 0001, Daoyuan Wu, Yang Liu 0003, Chunmiao Li, Mingchao Wan, Jin Dong 0004
IEEE Trans. Inf. Forensics Secur.12
2026 Nexus: A Novel Transaction Processing Framework for Permissioned Blockchain
abstract
The transaction execution layer is a key determinant of throughput in permissioned blockchains. While recent Shared Memory Pools (SMP)-based approaches improve throughput by enabling all consensus nodes to participate in transaction packaging, they face two fundamental limitations. First, the performance bottleneck shifts from the consensus layer to the transaction execution layer as transaction number confirmed in a round increases. Second, these approaches are vulnerable to “transaction duplication” attacks where malicious clients can simultaneously send the same transaction to multiple consensus nodes, thereby decreasing the number of valid transactions in block proposals. To address these limitations, this paper introducesNexus, a novel blockchain transaction processing framework with high scalability.Nexusleverages the idle computational resources of full nodes to enable transaction execution in parallel with the consensus. Moreover,Nexusallows each node to handle only a fraction of the total transactions and share execution results with others. This approach reduces overall transaction execution time, increases throughput, and decreases latency. Lastly,Nexusintroduces a transaction partitioning mechanism that effectively addresses the “transaction duplication” attack and achieves load balancing between clients and consensus nodes. Our implementation ofNexusdemonstrates significant improvements: throughput increases by 4x to 15x, and latency is reduced by 50% to 70%.
Shengjie Guan, Rongkai Zhang 0005, Qiuyu Ding, Mingxuan Song, Jieyi Long, Mingchao Wan, Taifu Yuan, Jin Dong 0004
IEEE Trans. Parallel Distributed Syst.9
2026 EquiLink Bridge: A Semi-Custodial Approach to Cross-Chain Transactions via TEE
abstract
The rising demand for blockchain interoperability is accelerating advancements in cross-chain bridge technologies, which are crucial for a seamless information transfer in multi-blockchain ecosystems. Existing blockchain bridges are typically classified into two categories: custodial and non-custodial. Custodial bridges use a trusted third party for easier and faster transactions but depend on custodian trust, while non-custodial bridges enhance transparency and control with smart contracts but increased complexity and latency. Currently, no bridge design successfully combines the benefits of both while avoiding their drawbacks. This paper presents EquiLink, a semi-custodial bridge that combines the benefits of both custodial and non-custodial methods. EquiLink employs a smart contract, known as the EquiLink Service, to initiate cross-chain transfers. It then uses the EquiLink Network, a system composed of remote-attested Trusted Execution Environments (TEEs), to verify and issue these transfers between two blockchains. Any eligible participants validated through remote attestation can join the EquiLink Network and contribute to the bridge’s functionality. Additionally, participants are regulated by an economic model, providing an extra layer of security through economic incentives. This semi-custodial bridge enhances transparency and control for users. Meanwhile, it mitigates the risks associated with centralized custody and decentralization. In the evaluation, EquiLink is resilient against both replay and physical attacks. Additionally, it operates efficiently, reducing transaction costs by 14.1% and latency by 18.9%
Tingda Shen, Yebo Feng, Jin Dong 0004, Konglin Zhu, Lei Jiao 0002, Lin Zhang 0013
IEEE Trans. Serv. Comput.3
2026 Scheduling Training-Inference Co-Location in Demand Response for Sustainable Edge AI
abstract
In the pursuit of data privacy and reduced latency, the adoption of edge intelligence has surged. Meanwhile, the enormous increase in AI has resulted in significant energy consumption. Edge intelligence plays a crucial role in Energy Demand Response (EDR). However, existing edge intelligence falls short of meeting the demands of co-locating training and inference tasks while satisfying EDR. Specifically, the intertwinement between balancing energy consumption, system delay and model accuracy, and uncertain future inputs adds to the challenge of designing an online sustainable system for co-located training and inference tasks. To address these challenges, we propose a novel two-timescale system for co-locating training and inference EDR. Our approach satisfies EDR by strategically planning training schedules on macro-timescales and migrating inference requests between heterogeneous edges on micro-timescales while minimizing long-term cost. We introduce a novel online polynomial time algorithm that first breaks down the problem into two subproblems, which are subsequently solved using an online-learning-based fractional algorithm and a randomized roun ding algorithm, respectively. Rigorous analysis demonstrates that our approach achieves both sublinear dynamic regret and sublinear dynamic fit. Extensive trace-driven evaluations validate the practical superiority of our approach over multiple existing methods, highlighting its effectiveness in real-world scenarios.
Konglin Zhu, Siyuan Wei, Xuan'er Wu, Lei Jiao 0002, Jin Dong 0004, Lin Zhang 0013
IEEE Trans. Serv. Comput.5
2025 ORCA: Mitigating Over-Reliance for Multi-Task Dwell Time Prediction with Causal Decoupling
abstract
Dwell time (DT) is a critical post-click metric for evaluating user preference in recommender systems, complementing the traditional click-through rate (CTR). Although multi-task learning is widely adopted to jointly optimize DT and CTR, we observe that multi-task models systematically collapse their DT predictions to the shortest and longest bins, under-predicting the moderate durations. We attribute this moderate-duration bin under-representation to over-reliance on the CTR-DT spurious correlation, and propose ORCA to address it with causal-decoupling. Specifically, ORCA explicitly models and subtracts CTR's negative transfer while preserving its positive transfer. We further introduce (i) feature-level counterfactual intervention, and (ii) a task-interaction module with instance inverse-weighting, weakening CTR-mediated effect and restoring direct DT semantics. ORCA is model-agnostic and easy to deploy. Experiments show an average 10.6% lift in DT metrics without harming CTR. Code is available at https://github.com/Chrissie-Law/ORCA-Mitigating-Over-Reliance-for-Multi-Task-Dwell-Time-Prediction-with-Causal-Decoupling.
Huishi Luo, Fuzhen Zhuang, Yongchun Zhu, Yiqing Wu, Bo Kang, Ruobing Xie, Feng Xia 0006, Deqing Wang 0001, Jin Dong 0004
CIKM9
2025 H2Tune: Federated Foundation Model Fine-Tuning with Hybrid Heterogeneity
abstract
Different from existing federated fine-tuning (FFT) methods for foundation models, hybrid heterogeneous federated fine-tuning (HHFFT) is an under-explored scenario where clients exhibit double heterogeneity in model architectures and downstream tasks. This hybrid heterogeneity introduces two significant challenges: 1) heterogeneous matrix aggregation, where clients adopt different large-scale foundation models based on their task requirements and resource limitations, leading to dimensional mismatches during LoRA parameter aggregation; and 2) multi-task knowledge interference, where local shared parameters, trained with both task-shared and task-specific knowledge, cannot ensure only task-shared knowledge is transferred between clients. To address these challenges, we propose H2Tune, a federated foundation model fine-tuning with hybrid heterogeneity. Our framework H2Tune consists of three key components: (i) sparsified triple matrix decomposition to align hidden dimensions across clients through constructing rank-consistent middle matrices, with adaptive sparsification based on client resources; (ii) relation-guided matrix layer alignment to handle heterogeneous layer structures and representation capabilities; and (iii) alternating task-knowledge disentanglement mechanism to decouple shared and specific knowledge of local model parameters through alternating optimization. Theoretical analysis proves a convergence rate of O(1/√T). Extensive experiments show our method achieves up to 15.4% accuracy improvement compared to state-of-the-art baselines.
Yiqi Tong, Zhaojun Hu, Fuzhen Zhuang, Xiao Zhang 0015, Jin Dong 0004
ECAI8
2025 Know Your Account: Double Graph Inference-Based Account De-Anonymization on Ethereum
abstract
The scaled Web 3.0 digital economy, represented by decentralized finance (DeFi), has sparked increasing interest in the past few years, which usually relies on blockchain for token transfer and diverse transaction logic. However, illegal behaviors, such as financial fraud, hacker attacks, and money laundering, are rampant in the blockchain ecosystem and seriously threaten its integrity and security. In this paper, we propose a novel double graph-based Ethereum account de-anonymization inference method, dubbed DBG4ETH, which aims to capture the behavioral patterns of accounts comprehensively and has more robust analytical and judgment capabilities for current complex and continuously generated transaction behaviors. Specifically, we first construct a global static graph to build complex interactions between the various account nodes for all transaction data. Then, we also construct a local dynamic graph to learn about the gradual evolution of transactions over different periods. Different graphs focus on information from different perspectives, and features of global and local, static and dynamic transaction graphs are available through DBG4ETH. In addition, we propose an adaptive confidence calibration method to predict the results by feeding the calibrated weighted prediction values into the classifier. Experimental results show that DBG4ETH achieves state-of-the-art results in the account identification task, improving the F1-score by at least 3.75% and up to 40.52% compared to processing each graph type individually and outperforming similar account identity inference methods by 5.23 % to 12.91 %.
Shuyi Miao, Wangjie Qiu, Hongwei Zheng 0003, Qinnan Zhang, Xiaofan Tu, Xunan Liu, Yang Liu 0003, Jin Dong 0004, Zhiming Zheng 0001
ICDE8
2025 FEZE: Alignment-Flexible Zero-Shot Vertical Federated Learning
abstract
Different from existing vertical federated learning (VFL), zero-shot VFL (ZVFL) is an under-explored scenario where test classes are absent from partial parties' training sets. In extreme cases, some test classes even have no training samples for all parties. Traditionally, existing zero-shot methods require abundant seen-class samples for effective knowledge transfer to recognize unseen classes. However, both the limited aligned samples and different seen classes pose several unique challenges to ZVFL. The primary challenge lies in the seen-to-unseen transfer insufficiency, as the scarcity of aligned samples and diverse seen-class distributions across parties severely limits the model's capability to learn discriminative features that can generalize to unseen classes. Moreover, the multi-party bias inconsistency arises as different parties tend to be biased towards their own seen classes during prediction, leading to skewed classification results at the active party. To address these challenges, we propose FEZE, an alignment-flexible zero-shot vertical federated learning framework. Specifically, we introduce a relation learning network to capture class-feature relationships between class labels and feature representations across heterogeneous feature spaces, enabling unseen class recognition through relationship inference. Additionally, we design a meta-relation learning mechanism that leverages diverse class-feature patterns to tackle the insufficient feature generalization from limited seen-class samples. Finally, we propose an alignment-flexible adaptive aggregation strategy that achieves adaptively aggregation based on inconsistent prediction spaces with arbitrary number of aligned samples. Theoretical analysis proves that FEZE can achieve a convergence rate of O(1/T). In the most challenging zero-shot scenario without aligned samples, FEZE surpasses state-of-the-art baselines by an average of 7.47% across three datasets.
Yiqi Tong, Yiyang Duan, Fuzhen Zhuang, Xiao Zhang 0015, Zhaojun Hu, Jin Dong 0004
KDD (2)7
2025 Engorgio: An Arbitrary-Precision Unbounded-Size Hybrid Encrypted Database via Quantized Fully Homomorphic Encryption
Song Bian 0001, Haowen Pan, Zhou Zhang 0016, Yunhao Fu, Jiafeng Hua, Bo Zhang 0142, Yier Jin, Jin Dong 0004, Zhenyu Guan 0002
USENIX Security Symposium10
2025 Exploring AIoT Blockchain Transaction Semantic Detection and Incentive Mechanism With Evolutionary Game Toward Web 3.0 Ecosystem
abstract
In the Web 3.0 ecosystem, blockchain and Artificial Intelligence of Things (AIoT) construct the infrastructure, where blockchain transaction semantic detection (BTSD) aims to enhance blockchain security by identifying illegal transactions through distributed miners executing AI algorithms. However, the computational cost of performing semantic detection discourages miners from participating without adequate incentives. Existing studies focus on algorithmic aspects of BTSD, which generally ignore the critical issue of incentive mechanism. To fill this gap, we propose the first incentive-based BTSD framework in the transaction pool phase, emphasizing how incentives affect the behavior of miners and users. We use evolutionary game theory to model miner-user interactions and define three key scenarios to simulate the impact of reward decay and penalty factors on system dynamics. Our results demonstrate that adjusting these parameters significantly influences the number of miners engaging in semantic detection and users initiating legitimate transactions. Under certain conditions, a well-designed incentive mechanism can lead to an Evolutionary Stable Strategy (ESS), thereby achieving systemic stability. This study introduces a novel incentive mechanism for BTSD during the transaction pool phase and validates its effectiveness through both theoretical insights and numerical solutions to enhance blockchain security.
Qinnan Zhang, Zishuai Zhang 0001, Yiran Chen 0026, Misha Xu, Zehui Xiong, Jiequ Ji, Wangjie Qiu, Hongwei Zheng 0003, Jianming Zhu 0002, Jin Dong 0004, Zhiming Zheng 0001
IEEE Internet Things J.10
2025 MCHEAS: Optimizing Large-Parameter NTT Over Multicluster In-Situ FHE Accelerating System
abstract
Fully Homomorphic encryption (FHE) enables high-level security but with a heavy computation workload, necessitating software-hardware co-design for aggressive acceleration. Recent works on specialized accelerators for HE evaluation have made significant progress in supporting lightweight RNS-CKKS applications, especially those with high-density in-memory computing techniques. To fulfill higher computational demands for more general applications, this article proposes multicluster HE accelerating system (MCHEAS), an accelerating system comprising multiple in-situ HE processing accelerators, each functioning as a cluster to perform large-parameter RNS-CKKS evaluation collaboratively. MCHEAS features optimization strategies including the synchronous, preemptive swap, square-diagonal, and odd-even index separation. Using these strategies to compile the computation and transmission of number theoretic transform (NTT) coefficients, the method optimizes the intercluster data swaps, a major bottleneck in NTT computations. Evaluations show that under 1 GHz, with different intercluster data transfer bandwidths, our approach accelerates NTT computations by 26.40% to 51.75%. MCHEAS also improves computing unit utilization by 10.30% to 33.97%, with a maximum peak utilization rate of up to 99.62%. MCHEAS achieves 17.63% to 34.67% speedups for HE operations involving NTT, and 15.12% to 30.62% speedups for demonstrated applications, while enhancing the computing units’ utilization by 5.18% to 21.87% during application execution. Furthermore, we compare MCHEAS with SOTA designs under a specific intercluster data transfer bandwidth, achieving up to$81.45\times $their area efficiencies in applications.
Zhenyu Guan 0002, Luchang Lei, Hongyang Jia, Yi Chen 0012, Bo Zhang 0142, Changrui Ren, Jin Dong 0004, Song Bian 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2025 Advanced Smart Contract Vulnerability Detection via LLM-Powered Multi-Agent Systems
abstract
Blockchain’s inherent immutability, while transformative, creates critical security risks in smart contracts, where undetected vulnerabilities can result in irreversible financial losses. Current auditing tools and approaches often address specific vulnerability types, yet there is a need for a comprehensive solution that can detect a wide range of vulnerabilities with high accuracy. We propose LLM-SmartAudit, a novel framework that leverages Large Language Models (LLMs) to automate smart contract vulnerability detection and analysis. Using a multi-agent conversational architecture with a buffer-of-thought mechanism, LLM-SmartAudit maintains a dynamic record of insights generated throughout the audit process. This enables a collaborative system of specialized agents to iteratively refine their assessments, enhancing the accuracy and depth of vulnerability detection. To evaluate its effectiveness, LLM-SmartAudit was tested on three datasets: a benchmark for common vulnerabilities, a real-world project corpus, and a CVE dataset. It outperformed existing tools with 98% accuracy on common vulnerabilities and demonstrates higher accuracy in real-world scenarios. Additionally, it successfully identifies 12 out of 13 CVEs, surpassing other LLM-based methods. These results demonstrate the effectiveness of multi-agent collaboration in automated smart contract auditing, offering a scalable, adaptive, and highly efficient solution for blockchain security analysis.
Jing Sun 0002, Yuqiang Sun 0001, Ye Liu 0012, Daoyuan Wu, Zijian Zhang 0001, Xianhao Zhang, Meng Li 0006, Yang Liu 0003, Chunmiao Li, Mingchao Wan, Jin Dong 0004, Liehuang Zhu
IEEE Trans. Software Eng.12
2024 ESC-NTT: An Elastic, Seamless and Compact Architecture for Multi-Parameter NTT Acceleration
abstract
Fully homomorphic encryption (FHE) and post-quantum cryptography (PQC) heavily rely on number theoretic transform (NTT) to accelerate polynomial multiplication, However, most existing NTT accelerators lack flexibility when the underlying modulus and polynomial lengths change. Current designs often store twiddle factors in on-chip storage, facing a noticeable drawback when frequent parameter changes occur, leading to a potential 50% decrease in computation speed due to the input bandwidth limitations. To address this challenge, we propose ESC-NTT, a fully-pipelined and flexible architecture for handling NTTs with varying parameters. ESC-NTT, a complete custom architecture, continuously performs$N$-point (inverse) NTT, negacyclic NTT (NCN), and inverse NCN (INCN) without introducing bubbles during modulus and NTT length switches. Additionally, we introduce a twiddle factor generator (TFG) module to replace on-chip factor storage and save 68.7% twiddle factors' bandwidth compared to inputting every factor. In the experiment, ESC-NTT is implemented on a Xilinx Alveo U280 FPGA and synthesized in a 28 nm CMOS technology. In the case of frequent modulus switching and same on-chip storage, the calculation speed of ESC-NTT is 1.05× to 241.39× that of existing FHE accelerators when performing 4096-point NTT.
Zhenyu Guan 0002, Luchang Lei, Hongyang Jia, Yi Chen 0012, Bo Zhang 0142, Jin Dong 0004, Song Bian 0001
DATE9
2024 TierFlow: A Pipelined Layered BFT Consensus Protocol for Large-Scale Blockchain
abstract
As the coverage of permissioned blockchains expands and the number of participating replicas increases, a scalable and efficient Byzantine Fault Tolerant (BFT) protocol is essential for large-scale blockchain. Unfortunately, previous BFT consensus protocols rely on a single leader to drive the protocol, which becomes a bottleneck for system scalability when the number of replicas exceeds a certain threshold. Although some proposals suggest hierarchically grouping nodes into different layers to alleviate verification pressure on the single leader. However, existing solutions only support serial execution between layers, causing performance and latency bottlenecks. To address these issues, we propose TierFlow, the first layered consensus protocol that supports pipelined execution, maintaining high throughput in scenarios with a large-scale deployment of replicas. TierFlow innovatively addresses the serial execution bottleneck in layered consensus by decoupling inter-layer consensus. To eliminate redundant phases, we use a pre-proof method to advance the next round of verification, and utilize delayed verification to merge similar verification workflows. We implement TierFlow and compare it with advanced BFT protocols such as HotStuff and Fast-HotStuff. We conduct extensive experiments with over 100 replicas, demonstrating that TierFlow achieves throughput 14x higher than Fast-HotStuff in large-scale application scenarios, with the performance disparity widening as scale increases.
Yongkang Yu, Jinchun He, Xinwei Xu, Qinnan Zhang, Wangjie Qiu, Hongwei Zheng 0003, Jin Dong 0004
TrustCom8
2024 A comprehensive survey of federated transfer learning: challenges, methods and applications
abstract
Abstract Federated learning (FL) is a novel distributed machine learning paradigm that enables participants to collaboratively train a centralized model with privacy preservation by eliminating the requirement of data sharing. In practice, FL often involves multiple participants and requires the third party to aggregate global information to guide the update of the target participant. Therefore, many FL methods do not work well due to the training and test data of each participant may not be sampled from the same feature space and the same underlying distribution. Meanwhile, the differences in their local devices (system heterogeneity), the continuous influx of online data (incremental data), and labeled data scarcity may further influence the performance of these methods. To solve this problem, federated transfer learning (FTL), which integrates transfer learning (TL) into FL, has attracted the attention of numerous researchers. However, since FL enables a continuous share of knowledge among participants with each communication round while not allowing local data to be accessed by other participants, FTL faces many unique challenges that are not present in TL. In this survey, we focus on categorizing and reviewing the current progress on federated transfer learning, and outlining corresponding solutions and applications. Furthermore, the common setting of FTL scenarios, available datasets, and significant related research are summarized in this survey.
Fuzhen Zhuang, Xiao Zhang 0015, Yiqi Tong, Jin Dong 0004
Frontiers Comput. Sci.5
2024 FedSQ: A Secure System for Federated Vector Similarity Queries
abstract
Vector databases have emerged as crucial tools for managing and retrieving representation embeddings of unstructured data. Given the explosive growth of data, vector data is often distributed and stored across multiple organizations. However, privacy concerns and regulations like GDPR present new challenges in collaborative and secure queries, also known as federated queries, over those vector data distributed across various data owners. Although existing research has attempted to enable such query services for low-dimensional data, such as relational and spatial data, these solutions can be inefficient in answering vector similarity queries involving high-dimensional data. Therefore, we are motivated to develop a new prototype system called FedSQ that (1) ensures privacy protection across data owners and (2) balances query efficiency and result accuracy when processing federated vector similarity queries. To achieve these goals, FedSQ utilizes advanced secure multi-party computation techniques to prevent information leakage during query processing and incorporates indexing and sampling based optimizations to strike a proper performance balance.
Zeqi Zhu, Zeheng Fan, Yuxiang Zeng, Yexuan Shi, Yi Xu 0013, Mengmeng Zhou, Jin Dong 0004
Proc. VLDB Endow.7
2024 Phantasm: Adaptive Scalable Mining Toward Stable BlockDAG
abstract
Blockchain technology builds an immutable and append-only ledger in peer-to-peer networks, which attracts attention from various fields. However, traditional chain-based blockchain systems typically have the problem of low throughput, leading to unsatisfactory performance. Among the proposed solutions, introducing a structure of the Directed Acyclic Graph (DAG) into the blockchain reaches a high transaction throughput. Such an approach enables blocks to refer to more than one previous block, thus processing blocks in parallel with better performance. However, existing DAG-based blockchain schemes do not establish a deterministic rule for block reference priority. Adversaries can initiate a splitting attack to select block references to affect DAG topology, making the consensus unstable. In this paper, we propose a more stable consensus protocol named Phantasm, aiming to stabilize the ordering result in the consensus protocol. The referred blocks can be decided after computing a solution to the block puzzle and the difficulty of this solution affects the number of block references. We design two strategies to guide the honest nodes to select references so that they can resist the splitting attacks to stabilize the ordering. Theoretical analysis and simulation experiments show that Phantasm is more stable than the classic DAG-based blockchain consensus protocol Phantom regarding the ordering results.
Zijian Zhang 0001, Kaiyu Feng, Mingchao Wan, Meng Li 0006, Jin Dong 0004, Liehuang Zhu
IEEE Trans. Serv. Comput.6