VLDB 2026 Research / reviewers in the wild / expert
Di Zhang 0011
dblp:80/3482-11
· DBLP profile ↗
19ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-4875-1319ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EECF: An edge-end collaborative framework with optimized lightweight model
Dewen Qiao, Zhenyan Wang, Jiamiao Liu, Xuetao Chen, Di Zhang 0011, Maolan Zhang |
Expert Syst. Appl. | 5 |
| 2026 | Updatable Multi-Party Private Set Intersection for Real-Time Collaborative Threat Intelligence
Ze Jiang, Biwen Chen, Zhongming Wang, Di Zhang 0011, Xiaoguo Li, Tao Xiang 0001, Xiaofeng Liao 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Approximate statistic query via sampling for IoT data trading and sharing
Fan Yang 0064, Di Zhang 0011, Junqing Le, Jiali Wu |
Comput. Networks | 2 |
| 2025 | DLA-FCIL: Federated Class-Incremental Learning for Dynamic Data with Forgetting Compensation and Auxiliary GeneratorsabstractFederated Class-Incremental Learning (FCIL) enables dynamic model updates, but suffers from local and global catastrophic forgetting due to limited client storage and cross-client class non-i.i.d. issues. To address catastrophic forgetting, we propose a multi-scale FCIL scheme, DLA-FCIL, which incorporates a Double-Loss-assisted forgetting compensation mechanism and the Auxiliary generators based on data characteristics. Specifically, to mitigate local catastrophic forgetting, we incorporate an auxiliary generator on local clients for knowledge replay, augmenting the training datasets with generated samples to form hybrid datasets. Then, to fully leverage the hybrid datasets, we design a double-loss forgetting compensation mechanism. This mechanism includes a gradient-weighted compensation loss that normalizes forgetting rates across old class knowledge, and a semantic-transition compensation loss that extracts the semantic relationships between old and new classes, preventing abrupt shifts in semantic consistency during the class transitions. Besides, to effectively alleviate the catastrophic forgetting problem caused by global class imbalance, the trained auxiliary generator is sent to a proxy server with minimal communication cost to build an i.i.d. dataset, enabling the development of an optimal global model. Finally, comparison experiments on CIFAR100, ImageNet-Subset, and Tiny-ImageNet datasets demonstrate that DLA-FCIL consistently outperforms other FCIL baselines by approximately 3–15% in test accuracy. Junqing Le, Di Zhang 0011, Xiaofeng Liao 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2025 | EMAFL: Evolutionary Momentum Auxiliary Adaptive Accelerating Federated LearningabstractThe utilization of federated learning (FL) has witnessed notable advancements in the domain of edge computing (EC). However, limited edge resources and heterogeneous devices restrict the accelerated training of the FL model. To address this issue, we introduce the biological evolutionary mechanism and momentum gradient descent (MGD) update approach into FL, called the EMAFL scheme, aiming to achieve accelerated model training and maximize resource utilization, simultaneously. Specifically, we first update the local model with particle swarm optimization (PSO) for each device and perform MGD on the updated local model. Next, by a toy example, we illustrate the necessity of adopting the different number of local iterations for heterogeneous devices in a resource-limited environment. Analytical convergence of the EMAFL scheme, premised on a delineated resource budget is subsequently explored. This yields a mathematical delineation correlating the quantity of local iterations for heterogeneous devices with the optimal model parameters. Predicated on the prior theoretical examinations, an adaptive control algorithm is devised to ascertain the local iteration count pertinent to each device following every communication round. Finally, through a lot of experiments compared with the benchmarks, the advantages of EMAFL in model accuracy, resource consumption, and Non-IID issues are verified. Dewen Qiao, Songtao Guo, Xuetao Chen, Pengzhan Zhou, Di Zhang 0011 |
IEEE Internet Things J. | 5 |
| 2025 | ECBSecDA: Edge Computing and Blockchain-Based Security Data Aggregation Scheme in a Smart GridabstractThe smart grid, which integrates with many advanced technologies and applications, brings great convenience for the data interaction between users and service providers. It is inevitable to collect a large amount of fine-grained, real-time data of users. However, these data transmissions have a huge communication burden and security issue in the smart grid. Therefore, this paper proposes an Edge Computing and Blockchain-based Security Data Aggregation scheme (ECBSecDA) for secure and fast data aggregation. This scheme is a three-layer data aggregation framework including user layer, edge computing layer and blockchain layer, and each layer uses Paillier cryptosystem to encrypt and then aggregate data. In user layer, we improve group signature by sharing group parameters for reducing communication overhead. In edge computing layer, we use group signature tracking algorithm to identify the attackers’ identity and make them re-register. Then, we design an edge node diversion algorithm to deal with big data flow issues for the aggregated data from the user layer, achieving low latency communication. Finally, the blockchain layer further aggregates and manages the received data from edge computing layer. This paper analyzes the security of our scheme, which ensures the privacy of users’ data. We conducted some comparisons with relevant work, and experimental results show that our scheme has lower computation cost and communication overhead. Aijuan Wang, Jingyue Ma, Di Zhang 0011, Ximeng Liu |
IEEE Internet Things J. | 5 |
| 2025 | LLMBD: Backdoor defense via large language model paraphrasing and data voting in NLPabstractWith the rapid development of natural language processing (NLP), backdoor attacks have emerged as a significant security threat. These attacks inject malicious triggers into NLP models, causing them to produce adversarial output while remaining functional under normal input. To eliminate backdoors, existing data-driven defense methods typically transform backdoored samples into normal samples. However, these defenses lack the scalability to adapt effectively to various backdoor attacks. To address this challenge, we propose LLMBD, a novel data-driven backdoor defense method that leverages large language models (LLMs) for paraphrasing. Specifically, LLMBD uses large language models with optimized prompts to paraphrase the input text, eliminating potential backdoors while maintaining semantic integrity and textual fluency. During the training and inference phase, we apply grouping and major voting mechanisms to bypass residual backdoors in the paraphrased dataset. Finally, we validate the robustness and defense effectiveness of LLMBD through comprehensive model evaluations. Experimental results on datasets including SST-2, IMDB, and HSOL under various backdoor attack types (BadNets, AddSent, Synbkd, Stylebkd) show that LLMBD significantly outperforms existing methods such as RAP, STRIP, ParaFuzz, and TextGuard. On the SST-2, HSOL, and IMDb datasets, LLMBD achieves an average ASR drop of 0.278, with the average CACC maintained at 0.897. LLMBD exhibits superior robustness, generalization, and performance preservation without modifications to the backdoored model, providing an efficient and model-agnostic defense strategy against diverse backdoor threats. Fei Ouyang, Di Zhang 0011, Chunlong Xie, Hao Wang 0003, Tao Xiang 0001 |
Knowl. Based Syst. | 2 |
| 2025 | NAAFL: A Non-Authoritative Anarchic Federated Learning for Defending Against Malicious AttacksabstractThe centralized server in traditional federated learning (FL) is authoritative (i.e. decisive control), which may cause immeasurable damage to the system's security in the event of decision failure or attack. To weaken the authority of the central server, existing studies have proposed blockchain-based federated learning (BFL) approaches. However, existing BFL still suffers from high resource overhead and difficulty in resisting high malicious ratio (more than 50%) attacks. To address the above challenge, this paper proposes an efficient and secure non-authoritative (i.e. highly decentralized) anarchic (i.e. distributed self-governance) federated learning framework which is named NAAFL. During the local process of NAAFL, an area credit-based screening mechanism for participating devices is proposed to ensure that participating devices are always highly trusted devices with higher total credit values. Then, to effectively exclude a high percentage of malicious training gradients, a multi-device validation voting mechanism based on historical information is designed to construct the global gradient. Subsequently, to weaken the central server authority and reduce the resource overhead while guaranteeing security, a secure and low-consumption consensus mechanism based on the federation chain is proposed, and the overhead is further reduced by a momentum acceleration algorithm. Finally, the theoretical analysis and experimental simulation are conducted on the proposed NAAFL. The results further show that the proposed NAAFL outperforms existing studies and can defend against attacks with up to 80% malicious ratio, which exceeds the common threshold (50%) of existing studies. Meanwhile, the overhead of NAAFL is reduced by about 77.51% compared to BFL. Ruihong Xiu, Junqing Le, Di Zhang 0011, Qingguo Lü, Tao Xiang 0001, Xiaofeng Liao 0001 |
IEEE Trans. Sustain. Comput. | 3 |
| 2024 | NLPSweep: A comprehensive defense scheme for mitigating NLP backdoor attacks
Tao Xiang 0001, Fei Ouyang, Di Zhang 0011, Chunlong Xie, Hao Wang 0227 |
Inf. Sci. | 3 |
| 2024 | Secure Redactable Blockchain With Dynamic SupportabstractBlockchain is extensively applied to many fields as an immutable distributed ledger. However, the immutability contradicts regulations such as the GDPR ruling “the right to be forgotten” of data. Besides, numerous emerging blockchain-based applications call for elastic data management. To erase some data, redactable blockchains are proposed for breaking the immutability in a controlled way. Unfortunately, the prior solutions may suffer from poor security and centralized control of the redaction privilege. They cannot support dynamic nodes, where the departure of participators will result in a single point of failure. This paper proposes a noveldynamic and decentralizedattribute-basedchameleonhash (DACH) to make blockchain history mutable, achieving asecurely anddynamicallyredactable blockchain (SDR-chain) in a decentralized setting. We first propose the formal definition, security models, and concrete construction of our DACH. Meanwhile, we design a delegation algorithm of DACH to support a dynamically changing committee, where participators can freely and securely leave and join the network. Then, the transactions of the SDR-chain are redacted by computing DACH collisions. The security is analyzed in the random oracle model. Finally, theoretical analysis and experimental evaluation demonstrate that our SDR-chain is superior to the prior solutions in terms of security and functionality. Di Zhang 0011, Junqing Le, Tao Xiang 0001, Xiaofeng Liao 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Secure and Efficient Continuous Learning Model for Traffic Flow PredictionabstractHigh-performance traffic flow prediction models provide reliable future road information and optimize traffic navigation systems. However, the traffic data used for model learning contains lots of private information, and the existing privacy-preserving strategies always reduce the accuracy of prediction models. Besides, an effective traffic flow prediction model needs to be continuously and rapidly updated to adapt to dynamic changes in the traffic environment. Thus, we propose a Secure and Efficient Continuous Learning Model (SE-CLM) based on broad learning, spatial correlation, and adaptive sampling processing techniques to realize accurate and efficient traffic flow prediction under strong privacy protection. Specifically, SE-CLM is constructed on the broad network architecture to enable fast and continuous model training. This model is trained on a cloud server by combining the spatial correlation of traffic flows, to achieve accurate traffic flow prediction. Besides, an adaptive sampling strategy is designed to further improve the prediction accuracy of the model under the protection with differential privacy (DP), where the budget allocation for DP is optimized by adaptively sampling traffic flows with different timestamps for noise perturbation processing. Furthermore, the experimental simulations are conducted in real vehicular mobility datasets. The experimental results show that the designed spatial-based SE-CLM achieve more accurate and efficient traffic flow prediction than those of the other existing schemes. The adaptive sampling strategy not only significantly reduces the DP-noise added in traffic flows but also a 20% reduction in communication volume compared to other strategies. Finally, the security analysis also verifies that SE-CLM satisfies w-event ε-DP. Junqing Le, Di Zhang 0011, Fan Yang 0064, Tao Xiang 0001, Xiaofeng Liao 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | Towards Privacy-Preserving and Practical Data Trading for Aggregate StatisticabstractData trading is an effective way for commercial companies to obtain massive personal data to develop their data-driven businesses. However, when data owners may want to sell their data without revealing privacy, data consumers also face the dilemma of high purchase costs due to purchasing too much invalid data. Therefore, there is an urgent need for a data trading scheme that can protect personal privacy and save expenses simultaneously. In this paper, we design a privACy-preserving and praCtical aggrEgateStatiStic trading scheme (named as ACCESS). Technically, we focus on the group-level pricing strategy to make ACCESS easier to implement. The differential privacy technique is applied to protect the data owners' privacy, and the sampling algorithm is adopted to reduce the data consumers' costs. Specifically, to provide a maximum tolerant privacy loss guarantee for the data owners, we design a decision algorithm to detect whether a conflict occurs between the consumer-specified accuracy level and the maximum tolerable privacy loss budget. Besides, to minimize the purchase cost for the data brokers, we develop a sampling-based aggregation method consisting of two sampling algorithms (called as BUSA and BKSA, respectively). BUSA enables reducing purchase costs with no additional background knowledge. Once the data broker knows the data boundary, BKSA can significantly reduce the amount of data that needs to be purchased, thereby the purchase cost is reduced. Rigorous theoretical analysis and extensive experiments (over four real-world and public datasets) further demonstrate the practicability of ACCESS. Fan Yang 0064, Xiaofeng Liao 0001, Nankun Mu, Di Zhang 0011 |
IEEE Trans. Sustain. Comput. | 5 |
| 2023 | Secure and Efficient Data Deduplication in JointCloud StorageabstractData deduplication can efficiently eliminate data redundancies in cloud storage and reduce the bandwidth requirement of users. However, most previous schemes depending on the help of a trusted key server (KS) are vulnerable and limited because they suffer from revealing information, poor resistance to attacks, great computational overhead, etc. In particular, if the trusted KS fails, the whole system stops working, i.e., single-point-of-failure. In this article, we propose aSecure andEfficient dataDeduplication scheme (named SED) in a JointCloud storage system which provides the global services via collaboration with various clouds. SED also supports dynamic data update and sharing without the help of the trusted KS. Moreover, SED can overcome the single-point-of-failure that commonly occurs in the classic cloud storage system. According to the theoretical analyses, our SED ensures the semantic security in the random oracle model and has strong anti-attack ability such as the brute-force attack resistance and the collusion attack resistance. Besides, SED can effectively eliminate data redundancies with low computational complexity and communication and storage overhead. The efficiency and functionality of SED improves the usability in client-side. Finally, the comparing results show that the performance of our scheme is superior to that of the existing schemes Di Zhang 0011, Junqing Le, Nankun Mu, Jiahui Wu 0001, Xiaofeng Liao 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Privacy-Preserving Federated Learning With Malicious Clients and Honest-but-Curious ServersabstractFederated learning (FL) enables multiple clients to jointly train a global learning model while keeping their training data locally, thereby protecting clients’ privacy. However, there still exist some security issues in FL, e.g., the honest-but-curious servers may mine privacy from clients’ model updates, and the malicious clients may launch poisoning attacks to disturb or break global model training. Moreover, most previous works focus on the security issues of FL in the presence of only honest-but-curious servers or only malicious clients. In this paper, we consider a stronger and more practical threat model in FL, where the honest-but-curious servers and malicious clients coexist, named as the non-fully trusted model. In the non-fully trusted FL, privacy protection schemes for honest-but-curious servers are executed to ensure that all model updates are indistinguishable, which makes malicious model updates difficult to detect. Toward this end, we present an Adaptive Privacy-Preserving FL (Ada-PPFL) scheme with Differential Privacy (DP) as the underlying technology, to simultaneously protect clients’ privacy and eliminate the adverse effects of malicious clients on model training. Specifically, we propose an adaptive DP strategy to achieve strong client-level privacy protection while minimizing the impact on the prediction accuracy of the global model. In addition, we introduce DPAD, an algorithm specifically designed to precisely detect malicious model updates, even in cases where the updates are protected by DP measures. Finally, the theoretical analysis and experimental results further illustrate that the proposed Ada-PPFL enables client-level privacy protection with 35% DP-noise savings, and maintains similar prediction accuracy to models without malicious attacks. Junqing Le, Di Zhang 0011, Long Jiao, Kai Zeng 0001, Xiaofeng Liao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | SecEDMO: Enabling Efficient Data Mining with Strong Privacy Protection in Cloud ComputingabstractFrequent itemsets mining and association rules mining are among the top used algorithms in the area of data mining. Secure outsourcing of data mining tasks to the third-party cloud is an effective option for data owners. However, due to the untrust cloud and the distrust between data owners, the traditional algorithms which only work over plaintext should be re-considered to take security and privacy concerns into account. For example, each data owner may not be willing to disclose their own private data to others during the cooperative data mining process. The previous solutions are either not sufficiently secure or not efficient. Therefore, we propose aSecure andEfficientDataMiningOutsourcing (SecEDMO) scheme for secure outsourcing of frequent itemsets mining and association rules mining over the joint database (i.e., database aggregated from multiple data owners) in the paradigm of cloud computing. Based on our customized lightweight symmetric homomorphic encryption algorithm and a secure comparison algorithm, SecEDMO can ensure strong privacy protection and low data mining latency simultaneously. Moreover, the well-designed virtual transaction insertion algorithm can hide the information of the original database while still preserving the cloud’s ability to perform data mining over the obfuscated data. By evaluation of a numerical experiment and theoretical comparisons, the correctness, security, and efficiency of SecEDMO are confirmed. Jiahui Wu 0001, Nankun Mu, Junqing Le, Di Zhang 0011, Xiaofeng Liao 0001 |
IEEE Trans. Cloud Comput. | 5 |
| 2021 | Exploring the redaction mechanisms of mutable blockchains: A comprehensive surveyabstractBlockchain technology has attracted tremendous interest from both industry and academia. It is typically used to record a public history of transactions (e.g., payment/smart contract data), but storing nonpayment/contract data in transactions has been common. The ability to store data unrelated to payment/contract such as illicit data on blockchain may be abused for malicious purposes. For example, one may use blockchain to store the data related to child pornography and copyright violations, which are publicly visible and immutable. Moreover, an immutable blockchain is not suitable for all blockchain-based applications. So far, numerous redaction mechanisms for the mutable blockchain have been developed. In this paper, we aim at conducting a comprehensive survey that reviews and analyzes the state-of-the-art redaction mechanisms. We start by giving a general presentation of blockchain and summarize the typical methods of inserting data in blockchain. Next, we discuss the challenges of designing the redaction mechanism and propose a list of evaluation criteria. Then, redaction mechanisms of the existing mutable blockchains are systemically reviewed and analyzed based on our evaluation criteria. The analyses include algorithmic overviews, performance limitations, and security vulnerabilities. Finally, the comparisons and analyses provide new insights into these mechanisms. This survey will provide developers and researchers a comprehensive view and facilitate the design of future mutable blockchains. Di Zhang 0011, Junqing Le, Tao Xiang 0001, Xiaofeng Liao 0001 |
Int. J. Intell. Syst. | 1 |
| 2020 | Anonymous Privacy Preservation Based on m-Signature and Fuzzy Processing for Real-Time Data ReleaseabstractThe real-time data generated from various smart devices will be released and shared for public to obtain numerous benefits. However, it will lead to individual privacy leakage because of data mining or analysis. Currently, many existing privacy protection models either fail to be directly applied in real-time data release or are unsatisfactory in terms of data utility and privacy protection. Toward this end, based on m-signature and fuzzy processing, an anonymous privacy protection model, named PMF, is proposed in this paper. Specifically, for the proposed model there are five advantages: 1) PMF defines m-signature for making each bucket with at least m different sensitive values instead of generating any counterfeit tuples, which can not only resist h-difference attack but also improve practical value; 2) the buckets satisfying m-signature are variable over time, and this flexibility of m-signature can improve the efficiency of dynamic update; 3) PMF can effectively insert, delete, and modify real-time data for release; 4) PMF applies fuzzy processing to handle the tuples in the candidate list, which strikes a good balance between the utility of released data and privacy protection; and 5) PMF adopts greedy heuristic algorithm to process update operations, which greatly reduces the information loss of released data. Furthermore, PMF is obviously more secure than the existing models in the real-time data release. Finally, the results of the comparison experiments on real-world and synthetic datasets illustrate that PMF is superior to the existing models in terms of data utility and efficiency. Junqing Le, Di Zhang 0011, Nankun Mu, Xiaofeng Liao 0001, Fan Yang 0064 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | An Anonymous Off-Blockchain Micropayments Scheme for Cryptocurrencies in the Real WorldabstractBlockchain, as a secured, trusted, and decentralized architecture, is used to create secure and tamper-proof payment schemes, which can serve economies and societies without trusted parties. However, the transparency and traceability of blockchain severely restrict the anonymity of participants in the real world, which will cause participants' privacy leakage. Toward this end, in this paper, an anonymous off-blockchain micropayments scheme (AOM) is proposed for cryptocurrencies in the real world. In AOM, a payee receives micropayments from an “honest-but-curious” intermediary T by solving puzzles which are generated based on the standard RSA assumption. Meanwhile, T also receives micropayments from the payers by payee's solutions and T will randomly select the inputs of the merging transaction Tmer. In order to improve service efficiency of T and resist denial of service attack, one of the outputs of Tmeris paid for T as a service fee. Besides, AOM simultaneously ensures the correctness and fairness of transactions. Finally, from the analyses of property and security, AOM has strong unlinkability, ability for anti-attacks and unforgeability. Di Zhang 0011, Junqing Le, Nankun Mu, Xiaofeng Liao 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2018 | A fast and efficient approach to color-image encryption based on compressive sensing and fractional Fourier transform
Di Zhang 0011, Xiaofeng Liao 0001, Bo Yang 0025, Yushu Zhang 0001 |
Multim. Tools Appl. | 1 |