EDBT 2026 Demo / reviewers in the wild / expert
Jinbao Zhu
dblp:211/9680 · also Jin Bao Zhu
· DBLP profile ↗
17ranked-venue papers
12as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Theory of computation · 4 · 4 first-author · 4 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Security and privacy · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Secure Aggregation with Top-K Sparsification in Decentralized Federated LearningabstractSecure aggregation is a vital component for mitigating gradient leakage in federated learning, but its communication cost conventionally scales with the gradient dimension. This becomes prohibitive for large models and even more pronounced in decentralized federated learning with limited bandwidth and unreliable nodes. Top-K gradient sparsification is an effective approach to reduce communication by transmitting only a few entries of the full gradient, while maintaining competitive model accuracy. Nevertheless, the top-K entries selected by each user are unpredictable and vary across users, which poses a challenge for efficient sparse secure aggregation. This paper studies information-theoretic secure aggregation with top-K sparsification in decentralized federated learning under user dropouts and user collusion. We propose a communication-efficient sparse secure aggregation scheme that offloads dimension-dependent overhead to an offline phase and protects private gradients using random masks and permutations. Experimental results demonstrate that our scheme preserves accuracy comparable to full-gradient aggregation even with only 1% gradient sparsification, while substantially reducing the communication cost. Hengxuan Tang, Jinbao Zhu, Xiaohu Tang 0004 |
ISIT | 2 |
| 2025 | Generalized Lagrange Coded Computing: A Flexible Computation-Communication Tradeoff for Resilient, Secure, and Private ComputationabstractWe consider the problem of evaluating arbitrary multivariate polynomials over a massive dataset containing multiple inputs, on a distributed computing system with a leader node and multiple worker nodes. Generalized Lagrange Coded Computing (GLCC) codes are proposed to simultaneously provide resiliency against stragglers who do not return computation results in time, security against adversarial workers who deliberately modify results for their benefit, and information-theoretic privacy of the dataset amidst possible collusion of workers. GLCC codes are constructed by first partitioning the dataset into multiple groups, then encoding the dataset using carefully designed interpolating polynomials, and sharing multiple encoded data points to each worker, such that interference computation results across groups can be eliminated at the leader. Particularly, GLCC codes include the state-of-the-art Lagrange Coded Computing (LCC) codes as a special case, and exhibit a more flexible tradeoff between communication and computation overheads in optimizing system efficiency. Furthermore, we apply GLCC to distributed training of machine learning models, and demonstrate that GLCC codes achieve a speedup of up to$2.5-3.9\times $over LCC codes in training time, across experiments for training image classifiers on different datasets, model architectures, and straggler patterns. Jinbao Zhu, Hengxuan Tang, Yijia Chang |
IEEE Trans. Commun. | 1 |
| 2025 | Secure Embedding Aggregation for Cross-Silo Federated Representation LearningabstractRepresentation learning plays a pivotal role in modern applications by enabling high-quality embeddings that support various downstream tasks such as recommendation, clustering, and personalized services. In federated representation learning (FRL), a central server collaborates withNclients, each holding private data, to jointly learn representations of entities (e.g., users in a social network). However, existing embedding aggregation protocols often fall short in either ensuring privacy protections or fully leveraging aggregation opportunities, leaving sensitive data exposed or vulnerable to collusion. To address these challenges, we propose SecEA, a secure embedding aggregation protocol that fully exploits all potential aggregation opportunities across all entities among clients while providing provable privacy guarantees. SecEA defends both local entities and their embeddings—ensuring computational security against a curious server and statistical privacy against up toTN/2 colluding clients. Comprehensive experiments on various representation learning tasks in cross-silo scenarios demonstrate that SecEA incurs a negligible performance loss (within 5%) compared to protocols with weaker or no privacy guarantees, and its additional computational latency significantly diminishes when training deeper models on larger datasets. A parallel mechanism is also included, which helps further improve the efficiency linearly. These results underscore that SecEA not only provides full privacy protections for both entity and embedding, but also preserves the utility of the learned representations. Jiaxiang Tang, Jinbao Zhu, Kai Zhang 0039, Lichao Sun 0001, Changyu Dong |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | On the Communication-Computation Tradeoff for Symmetric Private Linear ComputationabstractWe consider the problem of symmetric private linear computation (SPLC) over a replicated storage system with colluding and straggler constraints. The SPLC problem allows the user to privately compute a linear combination of multiple files from a set of replicated servers, even in the presence of straggler servers that can bottleneck the entire computation. It is guaranteed that a certain number of colluding servers learn nothing about the coefficients of the linear combination, and the user must not learn any information about the files other than the desired linear computation. Unlike previous private computation literature that mainly focused on decreasing download cost from servers, we aim to establish a flexible tradeoff between communication costs and computational complexities. In particular, we propose a novel SPLC scheme under the assumption of a fixed number of stragglers. Additionally, by generalizing this SPLC scheme, we construct an adaptive SPLC scheme capable of tolerating the presence of a varying number of stragglers, even if their identities and numbers are unknown in advance. Compared to the SPLC scheme with a fixed number of stragglers, the adaptive SPLC scheme achieves a lower communication cost based on the actual number of stragglers, albeit at the cost of increased computational complexities. Both types of SPLC schemes achieve flexible performance tradeoffs and can be employed to optimize system efficiency in practice. Jinbao Zhu, Xiaohu Tang 0004 |
IEEE Trans. Inf. Theory | 1 |
| 2025 | A General Coding Framework for Adaptive Private Information RetrievalabstractThe problem ofT-colluding private information retrieval (PIR) enables the user to retrieve one out ofMfiles from a distributed storage system withNservers without revealing anything about the index of the desired file to any group of up toTcolluding servers. In the considered storage system, theMfiles are stored across theNdistributed servers in anX-secureK-coded manner such that any group of up toXcolluding servers learns nothing about the files; the storage overhead at each server is reduced by a factor of 1/Kcompared to the total size of the files; and the files can be reconstructed from anyK+Xservers. However, in practical scenarios, when the user retrieves the desired file from the distributed system, some servers may respond to the user very slowly or not respond at all. These servers are referred to asstragglers, and particularly their identities and numbers are unknown in advance and may change over time. This paper considers the adaptive PIR problem that can be capable of tolerating the presence of a varying number of stragglers. We propose a general coding method for designing adaptive PIR schemes by introducing the concept of afeasible PIR coding framework. We demonstrate that anyfeasible PIR coding frameworkover a finite field Fqwith sizeqcan be used to construct an adaptive PIR scheme that achieves a retrieval rate of 1 −K+X+T−1/N−Ssimultaneously for all numbers of stragglers 0 ≤S≤N−(K+X+T) over the same finite field. Additionally, we provide an implementation of thefeasible PIR coding framework, ensuring that the adaptive PIR scheme operates over any finite field Fqwith sizeq≥N+max{K,N−(K+X+T−1)}. Jinbao Zhu, Xiaohu Tang 0004 |
IEEE Trans. Inf. Theory | 1 |
| 2024 | Cross-Domain Few-Shot Hyperspectral Image Classification with Bias Diminishing and Domain BridgingabstractCross-domain few-shot hyperspectral image classification involves acquiring knowledge from a substantial set of labeled samples in the source domain and subsequently applying this knowledge to tasks within target domains characterized by a limited number of labeled samples. Few-shot learning combined with domain adaptation is currently a popular method for the above problem, but how to learn effective representations in the few-shot learning stage and the domain adaptation stage respectively is still a huge challenge. Therefore, we propose a new hyperspectral image classification method that adds the gradually vanishing bridge(GVB) module in the domain adaptation stage to improve domain-invariant representation learning, and adds the bias diminishing(BD) module in the few-shot learning stage to improve metric-based representation learning. The experimental results demonstrate that the proposed method outperforms the cited state-of-the-art methods on four public HSI datasets. Jiachen Bei, Guo Cao, Jinbao Zhu, Yingchun Han |
IGARSS | 3 |
| 2024 | Fusion of Dilated Convolution in CNN and Transformer Networks for Hyperspectral Image ClassificationabstractAlthough existing Convolutional Neural Networks (CNNs) have exhibited commendable performance in hyperspectral image classification, they often emphasize local features. In recent years, transformers have garnered interest for capturing global features in hyperspectral images. This paper introduces a hyperspectral image classification framework, Dilated Convolution in CNN and Transformer Networks (DCCTnet), which employs various types of convolutions for local feature extraction and integrates dilated convolutions into multi-head self-attention mechanisms. Initially, multi-stage dilated convolutions are applied for multi-level feature extraction from the image. Following this, a grouped convolution is employed for feature fusion. To integrate CNN and transformers seamlessly, a pioneering Multi-Scale Multi-Head Self-Attention mechanism (MS-MHSA) is proposed, incorporating multiple dilated convolutions in MHSA to capture local-global multi-scale hyperspectral features. This mechanism is seamlessly integrated with the CNN branch, harnessing the strengths of both CNN and transformers. Through extensive experiments on two standard datasets, our proposed method demonstrates higher classification accuracy compared to other state-of-the-art networks. Jinbao Zhu, Guo Cao, Jiachen Bei, Yingchun Han |
IGARSS | 1 |
| 2024 | Private Multiple Linear Computation: A Flexible Communication-Computation TradeoffabstractWe consider the problem of private multiple linear computation (PMLC) over a replicated storage system with colluding and unresponsive constraints. In this scenario, the user wishes to privately compute$P$linear combinations of$M$files from a set of$N$replicated servers without revealing any information about the coefficients of these linear combinations to any$T$colluding servers, in the presence of$S$unresponsive servers that do not provide any information in response to user queries. Our focus is on more general performance metrics where the communication and computational overheads incurred by the user are not neglected. Additionally, the communication and computational overheads for servers are also taken into consideration. Unlike most previous literature that primarily focused on download cost from servers as a performance metric, we propose a novel PMLC scheme to establish a flexible tradeoff between communication costs and computational complexities. Jinbao Zhu, Lanping Li, Xiaohu Tang 0004, Ping Deng 0003 |
ISIT | 1 |
| 2024 | A Novel Design of a Unilateral Nuclear Magnetic Resonance Sensor for Soil Moisture Detection Based on a Simplified Analytical ModelabstractSoil moisture (SM) is a key state variable in terrestrial systems because it controls the exchange of water and energy between the continental surface and the atmosphere. Nuclear magnetic resonance (NMR) technology is widely used for the analysis of porous media in SM due to its unique sensitivity to hydrogen protons. Unlike traditional laboratory NMR systems, unilateral NMR (UNMR) systems allow for the placement of detection targets outside the sensor space, thereby enabling in situ detection capabilities. However, in the existing designs of UNMR sensors, the magnetic field location is typically determined after the sensor has been designed. In this study, a novel UNMR sensor design scheme is proposed based on a simplified analytical model (SAM) to solve this problem. In contrast to conventional practices, this scheme places a primary emphasis on the identification of detection positions as its initial step, followed by the computation of magnet structure parameters. Concurrently, the mechanical design of the proposed UNMR sensor offers a more adaptable approach to regulation. The scheme consists of two components: magnetic field calculation and optimization of structural parameters. Notably, the proposed model exhibits a remarkable enhancement in calculation efficiency, surpassing the baseline by more than 70 times within a single iteration, compared with the traditional analytical model (TAM). The goodness of fit between the measured magnetic field distribution and the optimized results surpasses 0.99, thereby providing additional evidence of the sensor’s effectiveness. In addition, the sensor’s performance is demonstrated through measurements conducted on samples with varying SM content. Tingting Lin 0001, Hualiang Wang, Zhengping Li, Jinbao Zhu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Secure Embedding Aggregation for Federated Representation LearningabstractWe consider a federated representation learning framework, where with the assistance of a central server, a group of N distributed clients train collaboratively over their private data, for the representations (or embeddings) of a set of entities (e.g., users in a social network). Under this framework, for the key step of aggregating local embeddings trained privately at the clients, we develop a secure embedding aggregation protocol named SecEA, which leverages all potential aggregation opportunities among all the clients, while providing privacy guarantees for the set of local entities and corresponding embeddings simultaneously at each client, against a curious server and up to T < N/2 colluding clients. Jiaxiang Tang, Jinbao Zhu, Lichao Sun 0001 |
ISIT | 2 |
| 2023 | Information-Theoretically Private Matrix Multiplication From MDS-Coded StorageabstractWe study two problems of private matrix multiplication, over a distributed computing system consisting of a master node, and multiple servers that collectively store a family of public matrices using Maximum-Distance-Separable (MDS) codes. In the first problem of Private and Secure Matrix Multiplication (PSMM) from colluding servers, the master intends to compute the product of its confidential matrix$\mathbf {A}$with a target matrix stored on the servers, without revealing any information about$\mathbf {A}$and the index of target matrix to some colluding servers. In the second problem of Fully Private Matrix Multiplication (FPMM) from colluding servers, the matrix$\mathbf {A}$is also selected from another family of public matrices stored at the servers in MDS form. In this case, the indices of the two target matrices should both be kept private from colluding servers. We develop novel strategies for the two PSMM and FPMM problems, which simultaneously guarantee information-theoretic data/index privacy and computation correctness. We compare the proposed PSMM strategy with a previous PSMM strategy with a weaker privacy guarantee (non-colluding servers), and demonstrate substantial improvements over the previous strategy in terms of communication and computation overheads. Moreover, compared with a baseline FPMM strategy that uses the idea of Private Information Retrieval (PIR) to directly retrieve the desired matrix multiplication, the proposed FPMM strategy significantly reduces storage overhead, but slightly incurs large communication and computation overheads. Jinbao Zhu, Jie Li 0019 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Generalized Lagrange Coded Computing: A Flexible Computation-Communication TradeoffabstractWe consider the problem of evaluating arbitrary multivariate polynomials over a massive dataset, in a distributed computing system with a master node and multiple worker nodes. Generalized Lagrange Coded Computing (GLCC) codes are proposed to provide robustness against stragglers who do not return computation results in time, adversarial workers who deliberately modify results for their benefit, and information-theoretic security of the dataset amidst possible collusion of workers. GLCC codes are constructed by first partitioning the dataset into multiple groups, and then encoding the dataset using carefully designed interpolation polynomials, such that interference computation results across groups can be eliminated at the master. Particularly, GLCC codes include the state-of-the-art Lagrange Coded Computing (LCC) codes as a special case, and achieve a more flexible tradeoff between communication and computation overheads in optimizing system efficiency. Jinbao Zhu |
ISIT | 1 |
| 2022 | Multi-User Blind Symmetric Private Information Retrieval From Coded ServersabstractThe problem of Multi-user Blind$X$-secure$T$-colluding Symmetric Private Information Retrieval from Maximum Distance Separable (MDS) coded storage system with$B$Byzantine and$U$unresponsive servers (U-B-MDS-MB-XTSPIR) is studied in this paper. Specifically, a database consisting of multiple files, each labeled by$M$indices, is stored at the distributed system with$N$servers according to$(N,K+X)$MDS codes over$\mathbb {F}_{q}$such that any group of up to$X$colluding servers learn nothing about the data files. There are$M$users, in which each user$m,m=1,\ldots,M$privately selects an index$\theta _{m}$and wishes to jointly retrieve the file specified by the$M$users’ indices$(\theta _{1},\ldots,\theta _{M})$from the storage system, while keeping its index$\theta _{m}$private from any$T_{m}$colluding servers, where there exists$B$Byzantine servers that can send arbitrary responses maliciously to confuse the users retrieving the desired file and$U$unresponsive servers that will not respond any message at all. In addition, each user must not learn information about the other users’ indices and the database more than the desired file. An U-B-MDS-MB-XTSPIR scheme is constructed based on Lagrange encoding. The scheme achieves a retrieval rate of$1-\frac {K+X+T_{1}+\ldots +T_{M}+2B-1}{N-U}$with secrecy rate$\frac {K+X+T_{1}+\ldots +T_{M}-1}{ N-(K+X+T_{1}+\ldots +T_{M}+2B+U-1)}$on the finite field of size$q\geq N+\max \{K, N-(K+X+T_{1}+\ldots +T_{M}+2B+U-1)\}$for any number of files. Jinbao Zhu, Qifa Yan, Xiaohu Tang 0004 |
IEEE J. Sel. Areas Commun. | 1 |
| 2022 | Symmetric Private Polynomial Computation From Lagrange EncodingabstractThe problem of$X$-secure$T$-colluding symmetric Private Polynomial Computation (PPC) from coded storage system with$B$Byzantine and$U$unresponsive servers is studied in this paper. Specifically, a dataset consisting of$M$files is stored across$N$distributed servers according to$(N,K+X)$Maximum Distance Separable (MDS) codes such that any group of up to$X$colluding servers can not learn anything about the data files. A user wishes to privately evaluate one out of a set of candidate polynomial functions over the$M$files from the system, while guaranteeing that any$T$colluding servers can not learn anything about the identity of the desired function and the user can not learn anything about the$M$data files more than the desired polynomial function evaluations, in the presence of$B$Byzantine servers that can send arbitrary responses maliciously to confuse the user and$U$unresponsive servers that will not respond any information at all. A novel symmetric PPC scheme using Lagrange encoding is proposed. This scheme achieves a PPC rate of$1-\frac {G(K+X-1)+T+2B}{N-U}$with secrecy rate$\frac {G(K+X-1)+T}{N-(G(K+X-1)+T+2B+U)}$and finite field size$N+\max \{K,N-(G(K+X-1)+T+2B+U)\}$, where$G$is the maximum degree over all the candidate polynomial functions. Moreover, to further measure the efficiency of PPC schemes, upload cost, query complexity, server computation complexity and decoding complexity required to implement the scheme are analyzed. Remarkably, the PPC setup studied in this paper generalizes all the previous MDS coded PPC setups and the degraded schemes strictly outperform the best known schemes in terms of (asymptotical) PPC rate, which is the main concern of the PPC schemes. Jinbao Zhu, Qifa Yan, Xiaohu Tang 0004 |
IEEE Trans. Inf. Theory | 1 |
| 2021 | Improved Constructions for Secure Multi-Party Batch Matrix MultiplicationabstractThis paper investigates the problem of Secure Multi-party Batch Matrix Multiplication (SMBMM), where a user aims to compute the pairwise products$\mathbf {A}\divideontimes \mathbf {B}\triangleq (\mathbf {A}^{(1)}\mathbf {B}^{(1)},\ldots,\mathbf {A}^{(M)}\mathbf {B}^{(M)})$of two batch of massive matrices$\mathbf {A}$and$\mathbf {B}$that are generated from two sources, through$N$honest but curious servers which share some common randomness. The matrices$\mathbf {A}$(resp.$\mathbf {B}$) must be kept secure from any subset of up to$X_{\mathbf {A}}$(resp.$X_{\mathbf {B}}$) servers even if they collude, and the user must not obtain any information about$(\mathbf {A},\mathbf {B})$beyond the products$\mathbf {A}\divideontimes \mathbf {B}$. A novel computation strategy for single secure matrix multiplication problem (i.e., the case$M=1$) is first proposed, and then is generalized to the strategy for SMBMM by means of cross subspace alignment. The SMBMM strategy focuses on the tradeoff between recovery threshold (the number of successful computing servers that the user needs to wait for), system cost (upload cost, the amount of common randomness, and download cost) and system complexity (encoding, computing, and decoding complexities). Notably, compared with the known result by Chenet al., the strategy for the degraded case$X= X_{\mathbf {A}}=X_{\mathbf {B}}$achieves better recovery threshold, amount of common randomness, download cost and decoding complexity when$X$is less than some parameter threshold, while the performance with respect to other measures remain identical. Jinbao Zhu, Qifa Yan, Xiaohu Tang 0004 |
IEEE Trans. Commun. | 1 |
| 2021 | Capacity-Achieving Private Information Retrieval Schemes From Uncoded Storage Constrained Servers With Low Sub-PacketizationabstractThis paper investigates reducing sub-packetization of capacity-achieving schemes for uncoded Storage Constrained Private Information Retrieval (SC-PIR) systems. In the SC-PIR system, a user aims to download one out of K files from N servers while revealing nothing about the identity of the requested file to any individual server, in which the K files are stored at the N servers in an uncoded form and each server can store up to μK equivalent files, where μ is the normalized storage capacity of each server. We first prove that there exists a capacity-achieving SC-PIR scheme for a given storage design if and only if all the packets are stored exactly at M\triangleq μN servers for μ such that M=μN ∈ {2,3,...,N}. Then, the optimal sub-packetization for capacity-achieving linear SC-PIR schemes is characterized as the solution to an optimization problem, which is typically hard to solve since it involves non-continuous indicator functions. Moreover, a new notion of array called Storage Design Array (SDA) is introduced for the SC-PIR system. With any given SDA, an associated capacity-achieving SC-PIR scheme is constructed. Next, the SC-PIR schemes that have equal-size packets are investigated. Furthermore, the optimal equal-size sub-packetization among all capacity-achieving linear SC-PIR schemes characterized by Woolsey et al. is proved to be \frac N(M-1)gcd(N,M), which is achieved by a construction of SDA. Finally, by allowing unequal size of packets, a greedy SDA construction is proposed, where the sub-packetization of the associated SC-PIR scheme is upper bounded by \frac N(M-1)gcd(N,M). Among all capacity-achieving linear SC-PIR schemes, the sub-packetization is optimal when min{M,N-M}|N or M=N, and within a multiplicative gap \frac min{M,N-M}gcd(N,M) of the optimal one in general. In particular, for the special case N=d·M±1 where the positive integer d ≥ 2, we propose another SDA construction to obtain lower sub-packetization. Jinbao Zhu, Qifa Yan, Xiaohu Tang 0004, Ying Miao 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2020 | A New Capacity-Achieving Private Information Retrieval Scheme With (Almost) Optimal File Length for Coded ServersabstractIn a distributed storage system, private information retrieval (PIR) guarantees that a user retrieves one file from the system without revealing any information about the identity of its interested file to any individual server. In this paper, we investigate an (N, K, M) coded server model of PIR, where each of M files is distributed to N servers in the form of (N, K) maximum distance separable (MDS) code for some N > K and M > 1. As a result, we propose a new capacity-achieving (N, K, M) coded linear PIR scheme such that it can be implemented with file length (K(N-K)/(gcd(N,K)), which is much smaller than the previous best result K(N/(gcd(N,K)))M-1. Notably, among all the capacity-achieving coded linear PIR schemes, we show that the file length is optimal if M > ⌊K/(gcd(N,K)) - K/(N-K)⌋ + 1 or min(K, N - K)|N, and within a multiplicative gap (min(K,N-K))/(gcd(N,K) ) of a lower bound on the minimum file length otherwise. Jinbao Zhu, Qifa Yan, Xiaohu Tang 0004 |
IEEE Trans. Inf. Forensics Secur. | 1 |