EDBT 2026 Demo / reviewers in the wild / expert
Shengyu Fan
dblp:243/1367
· DBLP profile ↗
23ranked-venue papers
3as first author
23since 2021 · last 2026
0009-0009-9160-8540ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 2 first-author · 17 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Falcon: Algorithm-Hardware Co-Design for Efficient Fully Homomorphic Encryption AcceleratorabstractFully homomorphic encryption (FHE) enables computation on encrypted data without compromising privacy, positioning it as a promising solution for secure cloud computing. However, its substantial computational overhead impedes practical deployment, prompting the development of dedicated hardware accelerators. In practice, when deploying cryptographic algorithm optimizations on FHE accelerators, hardware constraints typically such as limited memory capacity, often lead to a disparity between theoretical algorithmic advantage and achievable hardware efficiency. Liang Kong 0005, Xianglong Deng, Guang Fan 0001, Shengyu Fan, Yilan Zhu, Geng Yang 0001, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005 |
ASPLOS (2) | 4 |
| 2026 | An Efficient and Scalable Hardware Architecture for Number Theoretic Transform on FPGA with Design AutomationabstractFully Homomorphic Encryption (FHE) has become a promising approach to protecting data privacy in emerging application scenarios. Unfortunately, FHE suffers from significant processing speed degradation compared to plaintext computation, with one of the primary bottlenecks being the time-consuming Number Theoretic Transform (NTT). Therefore, accelerating NTT to accommodate various FHE parameters is crucial to advancing FHE towards practical use. With highly reconfigurable and performant logical fabrics, Field Programmable Gate Arrays (FPGAs) have exhibited great potential in NTT acceleration. By decomposing large-point NTT with strong data dependency into independent and simple small-point NTTs, the emerging Ten-step NTT (TNTT) algorithms intuitively enable higher parallelism and thereby have the potential to explore better performance compared to traditional algorithms. However, our quantitative analysis reveals that TNTT exhibits significant performance degradation as parallelism increases due to additional varying-size transpositions and Hadamard products. This paper proposes AutoNest, an efficient and scalable hardware architecture, along with an accelerator auto-generation framework for TNTT. The proposed hardware architecture maximizes performance by 1) adopting a 2D block decomposition dataflow to address critical path delays in transpose logic, thereby improving clock frequency. 2) integrating algorithm-level costfree twiddle factor fusion to reduce the number of modular multiplications in Hadamard products, thereby allowing higher parallelism on chip. Moreover, we also deliver an accelerator generation framework conducting automated design space exploration to elaborate a performant TNTT architecture under the target FPGAs' resource budget for user-defined FHE parameters. Experimental results on the AMD-Xilinx U280 FPGA demonstrate that NTT accelerators generated by AutoNest achieve an average speedup of$2.31 \times$compared to prior designs. Yilan Zhu, Geng Yang 0001, Xingyu Tian, Dilshan Kumarathunga, Liang Kong 0005, Xianglong Deng, Shengyu Fan, Guang Fan 0001, Guiming Shi, Bo Zhang 0098, Yisong Chang, Shoumeng Yan, Zhenman Fang, Mingzhe Zhang 0005 |
HPCA | 7 |
| 2026 | HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration
Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Geng Yang 0001, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng, Meng Li 0004, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005 |
ISCA | 9 |
| 2026 | MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access Optimization
Xianglong Deng, Guang Fan 0001, Shengyu Fan, Mingzhe Zhang 0005 |
ISCA | 7 |
| 2026 | A Bitwidth-Flexible Modular Multiplier with Shift-Free Accumulation for Efficient NTT Acceleration in FHE
Shengyu Fan, Xianglong Deng, Rui Hou 0001, Mingzhe Zhang 0005 |
ISCAS | 3 |
| 2026 | TensorFHE+: Fully Homomorphic Encryption Acceleration Based on Linear AlgebraabstractFully Homomorphic Encryption (FHE) enables encrypted data processing on untrusted cloud servers, crucial for privacy-sensitive applications. Despite its potential, performance overheads (about 10, 000× slower) limit adoption. ASIC accelerators outperform GPUs/FPGAs by optimizing specific operations but rely on costly 7nm processes and large on-chip memory, hindering cost-effective deployment. Balancing efficiency with manufacturing constraints remains critical. This paper presents TensorFHE+, a GPU-optimized FHE acceleration framework leveraging Tensor Cores to accelerate Number Theoretic Transform (NTT) operations. Key innovations include: 1) Decomposing CKKS kernels into vector/matrix operations for hardware utilization; 2) Vectorized modulo arithmetic; 3) Data layout optimization for memory efficiency. Evaluated on NVIDIA A100, TensorFHE+ outperforms TensorFHE [1] by 1.44× in average (up to 1.69× on ResNet-20) and surpasses prior GPU implementations [2], [3]. The design also demonstrates compatibility with commercial linear algebra accelerators, enabling efficient FHE deployment. Yintai Sun, Shengyu Fan, Zhenhua Yin, Xinkai Song, Xing Hu 0001, Zidong Du, Qi Guo 0001, Weizhi Xu 0001, Rui Hou 0001, Dan Meng 0002, Song Bian 0001, Mingzhe Zhang 0005 |
IEEE Trans. Computers | 2 |
| 2025 | The Future of Fully Homomorphic Encryption System: From a Storage I/O Perspective
Erci Xu, Shengyu Fan, Xianglong Deng, Guiming Shi, Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Shoumeng Yan, Mingzhe Zhang 0005 |
APPT | 4 |
| 2025 | WPC: Weight Plaintext Compression for CNN Inference based on RNS-CKKSabstractConvolutional neural network (CNN) inference based on RNS-CKKS enables secure processing on encrypted data but introduces significant weight size overhead. Weight plaintext, weight in RNS-CKKS format, can reach tens to hundreds of gigabytes. Existing compression methods either add high computational cost or yield low compression rates. In this work, we propose WPC, Weight Plaintext Compression, to compress weight plaintext for RNS-CKKS-based CNN inference. We observe that the transformation from the weight in CNN models to the weight plaintext in RNS-CKKS format involves an operation akin to the Discrete Fourier Transform, which shifts data between the time and frequency domains while retaining redundant information from periodic and discrete data. Based on this observation, we first introduce the Periodic Transmit Theorem, which states that periodic patterns can be preserved during the transformation process, thereby enabling compression. We then propose Channel Innermost Packing Scheme and Rotation Padding to rearrange the weight data into periodic patterns for compression. Results show that WPC achieves 1.25 to 2.18 times speedup on an A100 GPU and 46.08 to 139.11 times compression rate. Guiming Shi, Shengyu Fan, Xianglong Deng, Liang Kong 0005, Jingwei Cai, Shuwen Deng, Mingzhe Zhang 0005, Kaisheng Ma |
CCS | 3 |
| 2025 | Corrosion Hammer: A Self-Activated Bit-Flip Attack to the Processing-In-Memory AcceleratorabstractIn this paper, taking ReRAM-based PIM accelerators as an example, we present a novel attack framework called Corrosion Hammer, which builds based on the Bit Flip Attack (BFA).Unlike previous BFA methods that require explicit memory fault injection techniques, such as Row Hammer, to modify sensitive bits in the victim Neural Network model, Corrosion Hammer implants the trojan during the hardware-software co-design phase and flips sensitive bits using read disturbance, which is a common noise in ReRAM caused by normal read operations.Furthermore, we explore the impact of inputs on the activation time consumption of the trojan and propose a method to expedite activation using normal input.Our experimental results demonstrate that Corrosion Hammer achieves an extremely covert trojan implantation and activation method, with an adversarial attack success rate of 92.46%.Additionally, using a specially designed method, Trojan activation is 61.98× faster compared to activation in an undisturbed normal operation state.It provides a way to significantly speed up the Trojan activation. Mengxin Zheng, Shengyu Fan, Qian Lou, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005 |
CF | 3 |
| 2025 | WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA CoresabstractThe application of Fully Homomorphic Encryption (FHE) is rapidly gaining traction as a means to maintain data confidentiality while performing computations on encrypted data. Given the accessibility and computational power, GPUs hold promise for significantly accelerating FHE operations. However, existing GPU-based acceleration solutions face several formidable challenges, notably the extensive occurrence of pipeline stalls induced by memory access and suboptimal harnessing of GPU hardware. This paper presents WarpDrive, a comprehensive framework for GPU-based FHE acceleration. Through sophisticated computation decomposition and fine-grained memory access design, WarpDrive significantly reduces the number of instructions by $\mathbf{7 3 \%}$ and pipeline stalls by $\mathbf{8 6 \%}$ compared to the state-of-the-art solution. Additionally, WarpDrive features a framework that supports the concurrent utilization of CUDA Cores and Tensor Cores within the NTT operation, for the first time, achieving performance that surpasses that of any single type of processing unit. Furthermore, we fully exploit the intra-ciphertext parallelism to elevate both computation and memory utilization, achieving up to $2.12 \times$ improvements without the need for ciphertext batching. Experimental results demonstrate that our optimizations highly enhance the performance of homomorphic operations. On an NVIDIA A100 GPU, WarpDrive achieves a throughput of 1218 KOPS for NTT and 305 KOPS for homomorphic multiplication, outperforming the state-of-the-art GPU solution (TensorFHE) by factors of $13.4 \times$ and $3.5 \times$, respectively. For the specific FHE workload, even under a much smaller batch size, our approach achieves $2.8 \times$ the performance of TensorFHE. Guang Fan 0001, Mingzhe Zhang 0005, Fangyu Zheng, Shengyu Fan, Xianglong Deng, Wenxu Tang, Liang Kong 0005, Shoumeng Yan |
HPCA | 4 |
| 2025 | Analysis of Bit-Flip Attacks on Encrypted Neural NetworksabstractWith the swift progression of artificial intelligence and deep learning, neural networks have achieved remarkable success in domains such as image recognition, natural language processing, and autonomous driving. However, the proliferation of data scales and the extensive deployment of computational resources have engendered significant privacy concerns for users. In scenarios involving personal sensitive data, the safeguarding of privacy is of utmost importance. Homomorphic encryption technology, particularly the CKKS scheme, is capable of performing computations with minimal computational error while preserving data privacy, and it has been extensively utilized in encrypted neural networks. This paper studies Bit-Flip Attacks (BFAs) on encrypted neural networks under the RNS-CKKS scheme. We empirically analyze the effects of bit flips at different memory locations—covering ciphertext data, model weights, and evaluation keys—and report their observable outcomes (silent misclassification, irregular yet decodable outputs, or computation aborts). Under a realistic threat model where the adversary cannot precisely target bytes nor observe model predictions, BFAs can corrupt results but do not leak additional information. Our findings indicate that key corruption often produces conspicuous anomalies that offer detection potential for defenders. Yilan Zhu, Rui Hou 0001, Dan Meng 0002, Shengyu Fan, Mingzhe Zhang 0005 |
ICPADS | 5 |
| 2025 | FAST: An FHE Accelerator for Scalable-parallelism with Tunable-bitabstractFully Homomorphic Encryption (FHE) enables direct computation on encrypted data, providing substantial security advantages in cloud-based modern society.However, FHE suffers from significant computational overhead compared to plaintext computation, hindering its adoption in real-world applications.While many accelerators have been designed to address performance bottlenecks, most do not fully leverage cryptographic optimization technologies, leaving room for further performance enhancements.In this work, we propose FAST, an FHE accelerator incorporating recent cryptographic optimizations, including hoisting technology and the gadget decomposition key-switching method (named KLSS method).We analyze ciphertext level consumption throughout application execution and observe that workload requirements vary significantly with different ciphertext levels for both hybrid and KLSS key-switching methods.Additionally, we note the differing computational precision requirements for these key-switching methods.Based on these observations, we designed a versatile framework that supports multiple key-switching methods during a single application execution and integrates hoisting technology. Shengyu Fan, Xianglong Deng, Liang Kong 0005, Guiming Shi, Guang Fan 0001, Dan Meng 0002, Rui Hou 0001, Mingzhe Zhang 0005 |
ISCA | 1 |
| 2025 | Neo: Towards Efficient Fully Homomorphic Encryption Acceleration using Tensor CoreabstractFully Homomorphic Encryption (FHE) is an emerging cryptographic technique for privacy-preserving computation, which enables computations on the encrypted data.Nonetheless, the massive computational demands of FHE prevent its further application to real-world workloads.To tackle this problem, several studies focus on the ASIC-based acceleration for FHE.However, the rapid evolution of FHE algorithms poses challenges to the generality of ASIC accelerator design.By contrast, a number of works rely on GPGPUs for FHE accelerations, due to the high parallelism and flexibility provided by GPGPUs.In this work, we propose a GPGPU-based acceleration solution that supports the Cheon-Kim-Kim-Song (CKKS) scheme by further exploiting Tensor Core(TCU) capabilities.In our study, we * Both author contributed equally to this research. Xianglong Deng, Shengyu Fan, Dan Meng 0002, Rui Hou 0001, Mingzhe Zhang 0005 |
ISCA | 4 |
| 2025 | HAWK: Fully Homomorphic Encryption Accelerator with Fixed-Word Key Decomposition Switching
Liang Kong 0005, Shengyu Fan, Xianglong Deng, Guang Fan 0001, Guiming Shi, Yilan Zhu, Geng Yang 0001, Shoumeng Yan, Mingzhe Zhang 0005 |
MICRO | 2 |
| 2025 | LP-HENN: fully homomorphic encryption accelerator with high energy efficiencyabstractAbstract Fully homomorphic encryption (FHE) enables direct computation on encrypted data without decryption, ensuring data privacy in cloud computing scenarios and preventing the leakage of sensitive information. However, the computational overhead of HE typically exceeds that of plaintext computation by 4 to 5 orders of magnitude, while energy consumption is 5 to 6 orders of magnitude higher. These substantial performance and energy overheads significantly hinder the widespread adoption of FHE. This paper proposed LP-HENN, a novel low-power and energy-efficient FHE accelerator architecture that leverages a RISC-V vector coprocessor and ReRAM crossbar arrays. LP-HENN targets power-constrained application scenarios such as edge devices, aiming to provide highly energy-efficient acceleration support for FHE applications. LP-HENN leverages the collaborative work of the vector processor and ReRAM crossbars, employing optimization strategies to achieve full pipelining and minimize memory access. Furthermore, this paper proposed a parameter selection model for early-stage architecture design, which achieves an optimal balance between performance and energy consumption through the collaborative optimization of multiple parameters. Experimental results show that, for an FHE-based convolutional neural network (HE-CNN) inference application, LP-HENN achieves a 31.82Ã- and 11920.56Ã- improvement in performance and energy efficiency, respectively, compared to CPU. Compared to FxHENN, the state-of-the-art FPGA-based FHE accelerator with high energy efficiency for edge devices, LP-HENN achieves a 2.36Ã- and 10.04Ã- improvement in performance and energy efficiency, respectively. The energy efficiency of LP-HENN is comparable to that of F1, the state-of-the-art ASIC FHE accelerator, while featuring a low power design suitable for edge computing. Zhuoyu Tian, Shengyu Fan, Xianglong Deng, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005 |
Cybersecur. | 3 |
| 2025 | Corrosion Hammer: a self-activated bit-flip attack to the processing-in-memory acceleratorabstractAbstract The Resistive Random-Access-Memory (ReRAM) crossbar-based Processing-In-Memory (PIM) accelerator shows great promise in accelerating neural networks (NNs). This technique boasts low energy consumption and exceptional performance in multiplication and accumulations (MAC) operations, making ReRAM-based PIM accelerators an ideal solution for intelligent computing in wearable and low-power mobile devices. However, security concerns related to PIM have not been adequately addressed. In this paper, we present a new attack framework called SolutionName for ReRAM-based PIM accelerators. SolutionName builds upon the Bit Flip Attack (BFA), a weight modification attack that manipulates the NN function by flipping specific bits in the deployed quantized NN model. Unlike previous BFA methods that require explicit memory fault injection techniques, such as Row Hammer, to modify sensitive bits in the victim NN, SolutionName implants the trojan during the hardware-software co-design phase and flips sensitive bits using read disturbance. Read disturbance is a common noise in ReRAM caused by normal read operations. This approach enables the trojan to be activated quietly during normal use, eliminating the need for explicit attacks. Furthermore, we explore the impact of inputs on the activation time of the trojan and propose a method to expedite activation using normal input. Our experimental results demonstrate that SolutionName achieves an extremely covert trojan implantation and activation method, with an adversarial attack success rate of 94.38%. Additionally, with a specially designed method, the trojan activation can be accelerated on average by 61.98 $$\times $$ × , providing controllable activation. Mengxin Zheng, Shengyu Fan, Qian Lou, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005 |
Cybersecur. | 3 |
| 2024 | Efficient Selection Based on Integrated Information for Dialogue State TrackingabstractDialogue State Tracking (DST) is a critical component in Task-Oriented Dialogue (TOD) systems, responsible for generating the dialogue state at each turn. Current approaches often struggle with complicated conversational contexts, primarily due to issues of information redundancy and insufficiency, which adversely affects accuracy. To address these challenges, we propose a novel approach termed Selection based on Integrated Information (SII). This method comprises three key components: an Information Integrator, which distills core information from the dialogue; an Information Selector, which identifies the most pertinent core information for each slot; and a State Predictor, which executes predictions based on the selected information. By focusing on selected information, SII demonstrates enhanced performance, achieving joint goal accuracies of 55.44% and 59.89% on the MultiWOZ2.0 and MultiWOZ2.1 datasets, respectively. Hongyun Du, Jikun Dong, Shengyu Fan, Shengjie Jia, Feiyue Diao, Jiran Zhu, Hui Yu 0010, Weizhi Xu 0001 |
IJCNN | 3 |
| 2024 | Trinity: A General Purpose FHE AcceleratorabstractFully Homomorphic Encryption (FHE) is crucial for privacy-preserving computing, which allows direct computation on encrypted data. While various FHE schemes have been proposed, none of them efficiently support both arithmetic FHE and logic FHE simultaneously. To address this issue, researchers explore the combination of different FHE schemes within a single application and propose algorithms for the conversion between them. Unfortunately, all prior ASIC-based FHE accelerators are designed to support a single FHE scheme, and none of them supports the acceleration for FHE scheme conversion. This necessitates FHE acceleration systems to integrate multiple accelerators for different schemes, leading to increased system complexity and hindering performance enhancement. In this paper, we present the first multi-modal FHE accelerator based on a unified architecture, which efficiently supports CKKS, TFHE, and their conversion scheme within a single accelerator. To achieve this goal, we first analyze the theoretical foundations of the aforementioned schemes and highlight their composition from a finite number of arithmetic kernels. Then, we investigate the challenges for efficiently supporting these kernels within a unified architecture, which include 1) concurrent support for NTT and FFT, 2) maintaining high hardware utilization across various polynomial lengths, and 3) ensuring consistent performance across diverse arithmetic kernels. To tackle these challenges, we propose a novel FHE accelerator named Trinity, which in-corporates algorithm optimizations, hardware component reuse, and dynamic workload scheduling to enhance the acceleration of CKKS, TFHE, and their conversion scheme. By adaptive select the proper allocation of components for NTT and MAC, Trinity maintains high utilization across NTTs with various polynomial lengths and imbalanced arithmetic workloads. The experiment results show that, for the pure CKKS and TFHE workloads, the performance of our Trinity outperforms the state-of-the- art accelerator for CKKS (SHARP) and TFHE (Morphling) by 1.49 x and 4.23 x, respectively. Moreover, Trinity achieves 919.3 x performance improvement for the FHE-conversion scheme over the CPU-based implementation. Notably, despite the performance improvement, the hardware overhead of Trinity is only 85 % of the summed circuit areas of SHARP and Morphling. Xianglong Deng, Shengyu Fan, Zhicheng Hu, Zhuoyu Tian, Jiangrui Yu, Dingyuan Cao 0002, Dan Meng 0002, Rui Hou 0001, Meng Li 0004, Qian Lou, Mingzhe Zhang 0005 |
MICRO | 2 |
| 2023 | TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPUabstractIn the cloud computing era, privacy protection is becoming pervasive in a broad range of applications (e.g., machine learning, data mining, etc). Fully Homomorphic Encryption (FHE) is considered the perfect solution as it enables privacy-preserved computation on untrusted servers. Unfortunately, the prohibitive performance overhead blocks the wide adoption of FHE (about 10, 000× slower than the normal computation). As heterogeneous architectures have gained remarkable success in several fields, achieving high performance for FHE with specifically designed accelerators seems to be a natural choice. Until now, most FHE accelerators have focused on efficiently implementing one FHE operation at a time based on ASIC and with significantly higher performance than GPU and FPGA. However, recent state-of-the-art FHE accelerators rely on an expensive and large on-chip storage and a high-end manufacturing process (i.e., 7nm), which increase the cost of FHE adoption.In this paper, we propose TensorFHE, an FHE acceleration solution based on GPGPU for real applications on encrypted data. TensorFHE utilizes Tensor Core Units (TCUs) to boost the computation of Number Theoretic Transform (NTT), which is the part of FHE with highest time-cost. Moreover, TensorFHE focuses on performing as many FHE operations as possible in a certain time period rather than reducing the latency of one operation. Based on such an idea, TensorFHE introduces operation-level batching to fully utilize the data parallelism in GPGPU. We experimentally prove that it is possible to achieve comparable performance with GPGPU as with state-of-the-art ASIC accelerators. TensorFHE performs 913 KOPS and 88 KOPS for NTT and HMULT (key FHE kernels) within NVIDIA A100 GPGPU, which is 2.61× faster than state-of-the-art FHE implementation on GPGPU; Moreover, TensorFHE provides comparable performance to the ASIC FHE accelerators, which makes it even 2.9× faster than the F1+ with a specific workload. Such a pure software acceleration based on commercial hardware with high performance can open up usage of state-of-the-art FHE algorithms for a broad set of applications in real systems. Shengyu Fan, Weizhi Xu 0001, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005 |
HPCA | 1 |
| 2023 | Poseidon: Practical Homomorphic Encryption AcceleratorabstractWith the development of the important solution for privacy computing, the explosion of data size and computing intensity in Fully Homomorphic Encryption (FHE) has brought enormous challenges to the hardware design. In this paper, we propose a practical FHE accelerator - "Poseidon", which focuses on improving the hardware resource and bandwidth consumption. Poseidon supports complex FHE operations like Bootstrapping, Keyswitch, Rotation and so on, under limited FPGA resources. It refines these operations by abstracting five key operators: Modular Addition (MA), Modular Multiplication (MM), Number Theoretic Transformation (NTT), Automorphsim and Shared Barret Reduction (SBT). These operators are combined and reused to implement higher-level FHE operations. To utilize the FPGA resources more efficiently and improve the parallelism, we adopt the radix-based NTT algorithm and propose HFAuto, an optimized automorphism implementation suitable for FPGA. Then, we design the hardware accelerator based on the optimized key operators and HBM to maximize computational efficiency. We evaluate Poseidon with four domain-specific FHE benchmarks on Xilinx Alveo U280 FPGA. Empirical results show that the efficient reuse of the operator cores and on-chip storage enables superior performance compared with the state-of-the-art GPU, FPGA and accelerator ASICs. We highlight the following results: (1) up to 370× speedup over CPU for the basic operations of FHE; (2) up to 1300×/52× speedup over CPU and the FPGA solution for the key operators; (3) up to 10.6×/8.7× speedup over GPU and the ASIC solution for the FHE benchmark. Yinghao Yang 0001, Huaizhi Zhang, Shengyu Fan, Mingzhe Zhang 0005, Xiaowei Li 0001 |
HPCA | 3 |
| 2023 | Tell me your position: Distantly supervised biomedical entity relation extraction using entity position marker
Jiran Zhu, Jikun Dong, Hongyun Du, Yanfang Geng, Shengyu Fan, Hui Yu 0010, Zengzhen Shao, Yaping Yang, Weizhi Xu 0001 |
Neural Networks | 5 |
| 2023 | Accelerating Convolutional Neural Network by Exploiting Sparsity on GPUsabstractThe convolutional neural network (CNN) is an important deep learning method, which is widely used in many fields. However, it is very time consuming to implement the CNN where convolution usually takes most of the time. There are many zero values in feature maps and filters, which leads to redundant calculations and memory accesses if dense methods are used to compute convolution. Many works recently have made use of sparsity to skip the calculations for zero values to reduce the inference time of the CNN. On the graphics processing unit platform, current works cannot fully exploit the sparsity of the feature map and achieve satisfactory performance. Therefore, we design a new parallel strategy to transform the feature map into a new storage format to avoid the redundant computation of zero values on graphics processing units. Also considering the sparsity in the feature map, we propose a fused storage format to combine the convolution operation with the following pooling operation, to further improve the performance. We carry out experiments with mainstream CNN models and achieve better performance compared with cuDNN and cuSPARSE. For VGG-19, ResNet-50, DenseNet-121, and RegNetX-16GF, 1.97×, 2.23×, 2.74×, and 1.58× speedups respectively are obtained over cuDNN. The speedups over cuSPARSE respectively are 2.10×, 1.83×, 2.35×, and 1.35× when only using the first method. Weizhi Xu 0001, Yintai Sun, Shengyu Fan, Hui Yu 0010, Xin Fu 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | Multi-attention deep neural network fusing character and word embedding for clinical and biomedical concept extraction
Shengyu Fan, Hui Yu 0010, Xiaoya Cai, Yanfang Geng, Guangzhen Li, Weizhi Xu 0001, Yaping Yang |
Inf. Sci. | 1 |