VLDB 2026 Research / reviewers in the wild / expert
Fan Zhang 0044
dblp:21/3626-44
· DBLP profile ↗
26ranked-venue papers
0as first author
23since 2021 · last 2026
0000-0001-7456-8377ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Security and privacy · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Resilient UAV Swarm with Fast Connectivity Recovery and Extensive CoverageabstractTo address partial node failures in unmanned aerial vehicle swarms, self-healing communication techniques are commonly employed to restore backbone connectivity while preserving area coverage. However, existing heuristic methods struggle to scale under large-scale failures and dynamic conditions, while learning-based approaches often suffer from spatial collapse, resulting in significant coverage loss. To overcome these limitations, we propose a resilient self-healing framework that enables rapid connectivity recovery and wide-area coverage through a divide-and-conquer strategy. First, we introduce a buffered dynamic virtual force expansion mechanism that categorizes pairwise distances into repulsive, neutral, and attractive zones, allowing nodes to disperse appropriately while preserving communication links and maintaining safety buffers. Subsequently, we design a multipartite graph convolution module to reason over subnetwork-level interactions and facilitate cross-subnetwork reconnection with global structural awareness. Finally, we develop an adaptive fusion strategy that combines both outputs with time-aware weighting to generate the final motion decisions. Experimental results in both random and uniform deployment scenarios demonstrate that our approach outperforms many state-of-the-art methods in terms of connectivity restoration speed and communication coverage. Yabin Peng, Chenyu Zhou 0004, Hainan Cui, Tong Duan, Fan Zhang 0044, Shaoxun Liu |
AAAI | 6 |
| 2026 | FHPSAC: FPGA-based High-Parallelism SAC AcceleratorabstractReinforcement learning (RL) enables autonomous decision-making in applications such as robotics and control, and Soft Actor-Critic (SAC) is a leading model-free algorithm for continuous tasks. However, SAC’s small-batch training updates and fine-grained computation lead to heavy scheduling and kernel-launch overheads on GPU, limiting efficiency. In this work, we present FHPSAC (FPGA-based High-Parallelism SAC Accelerator), the first FPGA-accelerated architecture dedicated to SAC training. First, we propose a hardware–software co-designed on-chip memory hierarchy to statically partition and allocate SAC’s training data for conflict-free parallel access. Second, we build a high-parallelism accelerator with a tensor core for GEMM (General Matrix Multiply) and a lightweight unit for irregular elementwise/reduction kernels. Finally, we implement the full system on a Xilinx XCVU9P FPGA and demonstrate significant speedup with low power. Experimental results show that FHPSAC obtains 5.33–14.85× speedup compared with the Intel Xeon Gold 6130 CPU, while outperforming an NVIDIA A100-SXM4 GPU by 2.90–10.43× in training latency with an average power of 40.17 W. FHPSAC substantially reduces SAC training latency, providing a computational foundation for large-scale SAC deployments. Jiabin Xu, Wang Fan, Xuegong Zhou, Wei Cao 0002, Fengzhe Zhang, Fan Zhang 0044, Xinsheng Yu 0001 |
FCCM | 7 |
| 2026 | DADSA: Dual-Side Adaptive Deep Safety Alignment for Large Language Models
Kunlin Li, Yabin Peng, Chenyu Zhou 0001, Fan Zhang 0044, Jiangtao Ma, Yaqiong Qiao, Wei Huang 0035 |
Inf. Process. Manag. | 4 |
| 2026 | EADOD: Ensemble adversarial defense via orthogonal distillation
Xinyuan Miao, Mingqi Qiao, Wei Huang 0035, Jiayu Du, Fan Zhang 0044, Guangjiao Zhou |
Knowl. Based Syst. | 5 |
| 2026 | Time-Aware Cybersecurity Knowledge Graph Reasoning Method for Vulnerability AnalysisabstractIn the digital age, software security is essential for the stability of information systems and data protection, yet increasing complexity in software systems has made vulnerabilities a significant cybersecurity threat, leading to data breaches, system crashes, and service disruptions. Traditional vulnerability assessments usually analyze vulnerabilities in isolation, ignoring their time relations and the risk of attackers exploiting multiple vulnerabilities simultaneously, known as co-exploitation. This paper proposes an innovative time-aware cybersecurity knowledge graph (TCG) reasoning method called TCGFormer, which is designed to address these challenges. TCGFormer comprises four modules: (1) an entity encoding module that adjusts attention based on positional information and interaction frequency, (2) a novel attention mechanism for encoding relational topology graphs, (3) a joint sequence encoding module for extracting temporal representations and node relations from historical interactions, and (4) a parameter learning module for predicting entities and relations. Extensive experiments on three public temporal datasets demonstrate that TCGFormer significantly outperforms existing baseline methods, and validation on a cybersecurity knowledge graph dataset—including NVD, CVE details, CWE database, and EDB—further confirms its efficacy in identifying co-exploitation behaviors. Kunlin Li, Fan Zhang 0044, Jiangtao Ma, Yaqiong Qiao |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | An Agile Deployment System for Password Recovery on FPGAabstractHardware-based acceleration of password recovery remains a pressing challenge, as CPUs and GPUs struggle to efficiently process modern cryptographic primitives. Although FPGAs offer superior performance-per-watt, their widespread adoption is limited by long development cycles, manual optimization, and the absence of an end-to-end deployment framework that jointly accelerates password generation and verification. To address this gap, we propose the first agile, end-to-end FPGA-based password recovery system that unifies deep-learning-driven password generation and cryptographic verification within a single deployment workflow. The framework consists of: (1) a customized Neural Processing Unit (NPU) that accelerates GAN-based password generation models such as PassGAN; (2) an automated, template-based accelerator generator for verification kernels, built on reusable Chisel hardware primitives; and (3) a multi-objective Design Space Exploration (DSE) engine that co-optimizes kernel-level parameters (e.g., loop unrolling) and system-level parallelism to determine globally optimal FPGA configurations. We deploy the system on a heterogeneous platform combining a Zynq MPSoC with dual Virtex UltraScale+ FPGAs. Experimental results show that the NPU outperforms an NVIDIA Tesla V100 by 82.16% in PassGAN inference throughput. The full system achieves 1.90× higher end-to-end throughput and 2.32× better energy efficiency than GPU-based implementations, and delivers an average 32.58% speedup over state-of-the-art FPGA-only verification designs. These results demonstrate the practicality and scalability of our architecture for real-world password recovery workflows. Liming Deng, Guowei Zhu, Xitian Fan, Guangwei Xie, Mingqian Sun, Xuegong Zhou, Wei Cao 0002, Fan Zhang 0044, Xinsheng Yu 0001 |
IEEE Trans. Computers | 8 |
| 2025 | Hadamard Transform Based Backdoor Attack
Qiulong Yang, Jiayu Du, Xinyuan Miao, Fan Zhang 0044 |
PRCV (3) | 4 |
| 2025 | S3Det: a fast object detector for remote sensing images based on artificial to spiking neural network conversionabstractArtificial neural networks (ANNs) have made great strides in the field of remote sensing image object detection. However, low detection efficiency and high power consumption have always been significant bottlenecks in remote sensing. Spiking neural networks (SNNs) process information in the form of sparse spikes, creating the advantage of high energy efficiency for computer vision tasks. However, most studies have focused on simple classification tasks, and only a few researchers have applied SNNs to object detection in natural images. In this study, we consider the parsimonious nature of biological brains and propose a fast ANN-to-SNN conversion method for remote sensing image detection. We establish a fast sparse model for pulse sequence perception based on group sparse features and conduct transform-domain sparse resampling of the original images to enable fast perception of image features and encoded pulse sequences. In addition, to meet accuracy requirements in relevant remote sensing scenarios, we theoretically analyze the transformation error and propose channel self-decaying weighted normalization (CSWN) to eliminate neuron overactivation. We propose S3Det, a remote sensing image object detection model. Our experiments, based on a large publicly available remote sensing dataset, show that S3Det achieves an accuracy performance similar to that of the ANN. Meanwhile, our transformed network is only 24.32% as sparse as the benchmark and consumes only 1.46 W, which is 1/122 of the original algorithm’s power consumption. Fan Zhang 0044, Guangwei Xie, Yanzhao Gao, Xiaofeng Qi, Mingqian Sun |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2025 | AHCA: Agile Design Framework for Hashcat Acceleration Based on FPGAabstractThis article presents AHCA, an agile design framework for Field Programmable Gate Array (FPGA)-based Hashcat acceleration that automates the generation of optimized register transfer level (RTL) code. Our approach is centered on a proposed automated design method using a parameterized domain-specific template (DST) and a specific hardware operator library. The framework analyzes an algorithm’s graph to extract key hardware operators and their interconnection network. To support diverse user inputs, we introduce an innovative operator matching strategy using subgraph isomorphism, which maps algorithms to our operator library. This matched information, combined with design space exploration (DSE), is used to configure the DST and generate the final RTL code, avoiding redundancy for previously implemented algorithms. Compared to state-of-the-art high-level synthesis (HLS) tools, AHCA demonstrates a maximum performance enhancement of 797×, a Look-Up table (LUT) efficiency improvement of up to 105×, and an energy efficiency gain of up to 676×. When deployed on an FPGA for password cracking, the AHCA-generated hardware achieves a 63.95× enhancement in energy efficiency over CPUs and a 4.71× improvement over GPUs. Liming Deng, Guowei Zhu, Xitian Fan, Wei Cao 0002, Xuegong Zhou, Fan Zhang 0044, Shaobo Yang |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2025 | PTME: A Regular Expression Matching Engine Based on Speculation and Enumerative Computation on FPGAabstractFast regular expression matching is an essential task for deep packet inspection. In previous works, the regular expression matching engine on FPGA struggled to achieve an ideal balance between resource consumption and throughput. Speculation and enumerative computation exploits the statistical properties of deterministic finite automata, allowing for more efficient pattern matching. Existing related designs mostly revolve around vector instructions and multiple processors/cores or SIMD instruction sets, with a lack of implementation on FPGA platforms. We design a parallelized two-character matching engine on FPGA for efficiently fast filtering off fields with no pattern features. We transform the state transitions with sequential dependencies to the existing problem of elements in one set, enabling the proposed design to achieve high throughput with low resource consumption and support dynamic updates. Results show that compared with the traditional DFA matching, with a maximum resource consumption of 25% for on-chip FFs (74323/1045440) and LUTs (123902/522720), there is an improvement in throughput of 8.08–229.96× speedup and 87.61–99.56% speed-up(percentage improvement) for normal traffic, and 11.73–39.59× speedup and 91.47–97.47% speed-up(percentage improvement) for traffic with high-frequency match hits. Compared with the state-of-the-art similar implementation, our circuit on a single FPGA chip is superior to existing multi-core designs. Mingqian Sun, Guangwei Xie, Fan Zhang 0044, Wei Guo 0018, Xitian Fan, Jiayu Du |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2025 | FPGA-Based Large-Scale Sorting with Optimized Bandwidth UtilizationabstractFast sorting of large-scale data is an essential task for data centers. In previous works, the existing computational model of sorting kernel still results in lower bandwidth utilization on the external memory bus. And the execution of merge operations in merge sort circuit on FPGAs depends on control commands from the host CPU. In this case, the merge sort circuit is not fully offloaded to hardware layer for acceleration, resulting in a performance loss. We design an on-chip merge sort controller to efficiently command the merge sort process. The proposed controller has the ability to schedule multiple on-chip computing kernels simultaneously in a more efficient mode, thus ensuring that the circuit has a better bandwidth utilization. Meanwhile, fundamental factors affecting the performance of merge sort are studied and analyzed, and we propose a high-performance merge sort architecture. Results show that using the proposed controller-centered architecture, an overall improvement of 20%–30% in sorting throughput can be achieved. Compared with the state-of-the-art previous merge sorting implementation on FPGA, our circuit can achieve 1.22/1.46 \(\times\) speedup. Mingqian Sun, Guangwei Xie, Fan Zhang 0044, Wei Guo 0018, Xitian Fan, Jiayu Du |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2025 | DVHetero: A Framework for Designing and Validating Heterogeneous SoC with RISC-V Processor and CGRAabstractCGRA, as a coprocessor in SoCs, has been widely studied. However, there is limited research on how to efficiently debug and verify SoCs composed of CGRAs and processors during the design process. To address this gap, we introduce DVHetero. DVHetero incorporates a simulation and validation framework, SoCDiff, which enables comprehensive SoC simulation, debugging, and rapid error localization. Using this verification framework, we successfully implemented and validated the entire SoC. The SoC includes a Chisel-based CGRA generator and provides a pipelined CGRA architecture template. The CGRA is tightly integrated with the RISC-V processor, allowing for efficient DMA-based data transfer and MMIO support within the SoC. The pipelined CGRA architecture generated by DVHetero shows a 1.27× improvement in area efficiency and a 10.54× increase in mapping speed compared to the state-of-the-art CGRA framework, HierCGRA. Additionally, compared to state-of-the-art CGRA-SoC systems FDRA, DVHetero demonstrates a 1.67× increase in execution speed and a 4.34× improvement in area efficiency. Guowei Zhu, Liming Deng, Kaisen Zhang, Wang Fan, Boyin Jin, Wei Cao 0002, Fengzhe Zhang, Xuegong Zhou, Fan Zhang 0044, Xinsheng Yu 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 9 |
| 2024 | Ensemble Adversarial Defense via Integration of Multiple Dispersed Low Curvature ModelsabstractThe integration of an ensemble of deep learning models has been extensively explored to enhance defense against adversarial attacks. The diversity among sub-models increases the attack cost required to deceive the majority of the ensemble, thereby improving the adversarial robustness. While existing approaches mainly center on increasing diversity in feature representations or dispersion of first-order gradients with respect to input, the limited correlation between these diversity metrics and adversarial robustness constrains the performance of ensemble adversarial defense. In this work, we aim to enhance ensemble diversity by reducing attack transferability. We identify second-order gradients, which depict the loss curvature, as a key factor in adversarial robustness. Computing the Hessian matrix involved in second-order gradients is computationally expensive. To address this, we approximate the Hessian-vector product using differential approximation. Given that low curvature provides better robustness, our ensemble model was designed to consider the influence of curvature among different sub-models. We introduce a novel regularizer to train multiple more-diverse low-curvature network models. Extensive experiments across various datasets demonstrate that our ensemble model exhibits superior robustness against a range of attacks, underscoring the effectiveness of our approach. Kaikang Zhao, Xi Chen 0112, Wei Huang 0035, Liuxin Ding, Xianglong Kong, Fan Zhang 0044 |
IJCNN | 6 |
| 2024 | Diversity supporting robustness: Enhancing adversarial robustness via differentiated ensemble predictions
Xi Chen 0112, Wei Huang 0035, Ziwen Peng, Wei Guo 0018, Fan Zhang 0044 |
Comput. Secur. | 5 |
| 2024 | A simple framework to enhance the adversarial robustness of deep learning-based intrusion detection system
Xinwei Yuan, Wei Huang 0035, Hongliang Ye, Xianglong Kong, Fan Zhang 0044 |
Comput. Secur. | 6 |
| 2024 | Formal verification of robustness and resilience of learning-enabled state estimation systems
Wei Huang 0035, Gaojie Jin, Youcheng Sun, Fan Zhang 0044, Xiaowei Huang 0001 |
Neurocomputing | 6 |
| 2024 | Correction: Adversarial defence by learning differentiated feature representation in deep ensemble
Xi Chen 0112, Wei Huang 0035, Wei Guo 0018, Fan Zhang 0044, Jiayu Du, Zhizhong Zhou |
Mach. Vis. Appl. | 4 |
| 2024 | Adversarial defence by learning differentiated feature representation in deep ensemble
Xi Chen 0112, Huang Wei, Wei Guo 0018, Fan Zhang 0044, Jiayu Du, Zhizhong Zhou |
Mach. Vis. Appl. | 4 |
| 2024 | Boosting Multimode Ruling in DHR Architecture With Metamorphic RelationsabstractABSTRACT The DHR architecture provides a revolutionary security defense structure for cyberspace. The multimode ruling in DHR is expected to alleviate the oracle problem, which still suffers from the existence of common model vulnerability. In this work, we design a test segmentation method to transform multimode ruling to a metamorphic testing problem. The text test input that causes inconsistency of heterogeneous executors is converted to a condition set, and we extract subsets of conditions based on its syntax tree. The original test can exploit a specific vulnerability, the follow‐up tests are composed by different subsets of conditions within the original test. We collect the execution matrix for the follow‐up tests to analyse the impact of each subset of conditions on ruling decision. Metamorphic relations are extracted based on the localization of independent condition, that is, the subsets of conditions that can impact ruling decision independently. The executors in an inconsistent ruling should be examined with metamorphic testing methods, rather than traditional majority voting mechanism. The proposed test segmentation and improved multimode ruling methods are evaluated on two DHR‐based cases, SQL injection in cyber‐range system and deserialization attack in ‐ project. The experimental results show that our test segmentation can help to locate malicious expressions and the metamorphic testing‐based multimode ruling can generate more correct results than majority voting mechanism with an average 15.8% performance loss. Ruosi Li, Xianglong Kong, Wei Guo 0018, Jingdong Guo, Hongfa Li, Fan Zhang 0044 |
Softw. Test. Verification Reliab. | 6 |
| 2023 | Unified Accelerator for Attention and Convolution in Inference Based on FPGAabstractMany models combining Transformers with convolutional neural networks (CNNs) for computer vision tasks have achieved state-of-the-art results. However, due to the different computation patterns between attention and convolution, using a dedicated Transformer or CNN accelerator will inevitably reduce the computing efficiency of the other. To overcome this problem, we propose a unified architecture for attention and convolution on FPGA. We reduce runtime overhead by offloading part of self-attention computations offline before inference. Furthermore, we present a unified mapping method according to the computing characteristics of attention-based and convolution-based models. This accelerator implements multi-head attention in Transformer, independent ResNet-50 and hybrid blocks of attention and con-volution in BoTNet-50 at 200MHz on Xilinx Virtex Ultrascale+ XCVU37P. Experimental results show that the solution is nearly 3.62 times more energy-efficient than the NVIDIA V100 GPU, and the computational efficiency is 11.86% and 28.29% higher than the state-of-the-art Transformer and ResNet-50 accelerators, respectively. Fan Zhang 0044, Xitian Fan, Jianliang Shen, Wei Guo 0018, Wei Cao 0002 |
ISCAS | 2 |
| 2023 | GAMBD: Generating adversarial malware against MalConv
Wei Guo 0018, Fan Zhang 0044, Jiayu Du |
Comput. Secur. | 3 |
| 2023 | SFTN: Fast object detection for aerial imagesabstractAbstract The task of remote sensing image object detection in low latency scenes is of great research significance. To address the problem that the current high‐precision object detection algorithm based on a feature pyramid network is slow due to a large number of parameters and complicated computation, a fast remote sensing image object detection method based on a Single‐scale Feature Transformation Network (SFTN) is proposed. Firstly, based on the single‐scale remote sensing image features, the new channel features are quickly generated by a linear transformation of the original features and convolution kernel clustering optimization using cosine similarity; secondly, in order to obtain multi‐scale receptive fields, a parallel residual hole convolution module is designed to cover multi‐category remote sensing object scales on the feature map; finally, angle variables are introduced and optimized using angle similarity to effectively improve the object orientation accuracy. The experimental results on different datasets show that the method in this paper improves the detection speed rapidly while ensuring the accuracy of remote sensing image object detection, which is better than many remote sensing image object detection methods. The results demonstrate the reliability and robustness of the method. Fan Zhang 0044, Wei Guo 0018, Mingqian Sun |
IET Image Process. | 2 |
| 2023 | A Combined Usage of NLP Libraries Towards Analyzing Software DocumentsabstractSoftware documents are commonly processed by natural language processing (NLP) libraries to extract information. The libraries provide similar functional APIs to achieve NLP tasks, numerous toolkits result in a problem of selection. In this work, we propose a method to combine the strengths of different NLP libraries to avoid the subjective selection of a specific NLP library. The combined usage is conducted through two steps, i.e. document-level selection of primary NLP library and sentence-level overwriting. The primary NLP library is determined according to the overlap degree of the results. The highest overlap degree indicated the most effective NLP library on a specific NLP task. Through sentence-level overwriting, the possible fine-gained improvements from other libraries are extracted to overwrite the outputs of primary library. We evaluate the combined method with six widely used NLP libraries and 200 documents from three different sources. The results show that the combined method can generally outperform all the studied NLP libraries in terms of accuracy. The finding means that our combined method can be used instead of individual NLP library for more effective results. Xianglong Kong, Hangyi Zhuo, Zhechun Gu, Xinyun Cheng, Fan Zhang 0044 |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2020 | MDLB: a metadata dynamic load balancing mechanism based on reinforcement learningabstractWith the growing amount of information and data, object-oriented storage systems have been widely used in many applications, including the Google File System, Amazon S3, Hadoop Distributed File System, and Ceph, in which load balancing of metadata plays an important role in improving the input/output performance of the entire system. Unbalanced load on the metadata server leads to a serious bottleneck problem for system performance. However, most existing metadata load balancing strategies, which are based on subtree segmentation or hashing, lack good dynamics and adaptability. In this study, we propose a metadata dynamic load balancing (MDLB) mechanism based on reinforcement learning (RL). We learn that the Q_learning algorithm and our RL-based strategy consist of three modules, i.e., the policy selection network, load balancing network, and parameter update network. Experimental results show that the proposed MDLB algorithm can adjust the load dynamically according to the performance of the metadata servers, and that it has good adaptability in the case of sudden change of data volume. Zhaoqi Wu, Fan Zhang 0044, Wei Guo 0018, Guangwei Xie |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2020 | Erratum to: MDLB: a metadata dynamic load balancing mechanism based on reinforcement learningabstractUnfortunately the corresponding author’s ORCID was incorrect. It should be: Fan ZHANG, https://orcid.org/0000-0001-7456-8377 Zhaoqi Wu, Fan Zhang 0044, Wei Guo 0018, Guangwei Xie |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2019 | A Reliable Physical Unclonable Function Based on Differential Charging CapacitorsabstractPhysical Unclonable Function (PUF) is an emerging security primitive for cryptography applications. However, achieving a very high reliability against the environmental variations remains a main challenge in PUF design and a key barrier for its commercialization. This paper presents a new PUF design based on the charging of a symmetric MOS capacitor pair by constant current with cross-coupled positive feedback inverters. The proposed weak PUF features high raw response reliability against variations in power supply and temperature without power-up reset noise and other issues due to the power-down and up of an array of cells. Extensive Monte-Carlo simulations have been performed using a standard 110nm CMOS process technology. The simulated results show an almost ideal uniqueness of 50.03% and superior reliability of 97.70% over a temperature range from 0 °C to 80 °C, and 96.20% with the supply voltage varies from 1.2 V to 1.8 V. The response bit can be generated at a rate of 27.78 Mbps with an average power consumption of 20.86 μW at 1.5V, and the energy consumption is only 750 fJ/bit. Wei Guo 0018, Chip-Hong Chang, Yuan Cao 0003, Shaojun Wei, Shouyi Yin, Chenchen Deng, Leibo Liu, Fan Zhang 0044 |
ISCAS | 10 |