VLDB 2026 Research / reviewers in the wild / expert
Geng Yang 0001
dblp:97/3523-1
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-6921-1007ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Falcon: Algorithm-Hardware Co-Design for Efficient Fully Homomorphic Encryption AcceleratorabstractFully homomorphic encryption (FHE) enables computation on encrypted data without compromising privacy, positioning it as a promising solution for secure cloud computing. However, its substantial computational overhead impedes practical deployment, prompting the development of dedicated hardware accelerators. In practice, when deploying cryptographic algorithm optimizations on FHE accelerators, hardware constraints typically such as limited memory capacity, often lead to a disparity between theoretical algorithmic advantage and achievable hardware efficiency. Liang Kong 0005, Xianglong Deng, Guang Fan 0001, Shengyu Fan, Yilan Zhu, Geng Yang 0001, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005 |
ASPLOS (2) | 7 |
| 2026 | An Efficient and Scalable Hardware Architecture for Number Theoretic Transform on FPGA with Design AutomationabstractFully Homomorphic Encryption (FHE) has become a promising approach to protecting data privacy in emerging application scenarios. Unfortunately, FHE suffers from significant processing speed degradation compared to plaintext computation, with one of the primary bottlenecks being the time-consuming Number Theoretic Transform (NTT). Therefore, accelerating NTT to accommodate various FHE parameters is crucial to advancing FHE towards practical use. With highly reconfigurable and performant logical fabrics, Field Programmable Gate Arrays (FPGAs) have exhibited great potential in NTT acceleration. By decomposing large-point NTT with strong data dependency into independent and simple small-point NTTs, the emerging Ten-step NTT (TNTT) algorithms intuitively enable higher parallelism and thereby have the potential to explore better performance compared to traditional algorithms. However, our quantitative analysis reveals that TNTT exhibits significant performance degradation as parallelism increases due to additional varying-size transpositions and Hadamard products. This paper proposes AutoNest, an efficient and scalable hardware architecture, along with an accelerator auto-generation framework for TNTT. The proposed hardware architecture maximizes performance by 1) adopting a 2D block decomposition dataflow to address critical path delays in transpose logic, thereby improving clock frequency. 2) integrating algorithm-level costfree twiddle factor fusion to reduce the number of modular multiplications in Hadamard products, thereby allowing higher parallelism on chip. Moreover, we also deliver an accelerator generation framework conducting automated design space exploration to elaborate a performant TNTT architecture under the target FPGAs' resource budget for user-defined FHE parameters. Experimental results on the AMD-Xilinx U280 FPGA demonstrate that NTT accelerators generated by AutoNest achieve an average speedup of$2.31 \times$compared to prior designs. Yilan Zhu, Geng Yang 0001, Xingyu Tian, Dilshan Kumarathunga, Liang Kong 0005, Xianglong Deng, Shengyu Fan, Guang Fan 0001, Guiming Shi, Bo Zhang 0098, Yisong Chang, Shoumeng Yan, Zhenman Fang, Mingzhe Zhang 0005 |
HPCA | 2 |
| 2026 | HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration
Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Geng Yang 0001, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng, Meng Li 0004, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005 |
ISCA | 8 |
| 2025 | HAWK: Fully Homomorphic Encryption Accelerator with Fixed-Word Key Decomposition Switching
Liang Kong 0005, Shengyu Fan, Xianglong Deng, Guang Fan 0001, Guiming Shi, Yilan Zhu, Geng Yang 0001, Shoumeng Yan, Mingzhe Zhang 0005 |
MICRO | 8 |
| 2024 | E4SA: An Ultra-Efficient Systolic Array Architecture for 4-Bit Convolutional Neural NetworksabstractMany studies have demonstrated that 4-bit precision quantization can achieve comparable accuracy to floating-point DNNs, sparking significant interest in efficiently accelerating compressed DNNs, especially 4-bit convolutions, on edge devices. However, we observe that conventional systolic array (SA) architectures designed for DNNs cannot fully exploit the advantages of high DSP computational density offered by 4-bit DSP packing. Although state-of-the-art FPGA-based SA architectures (e.g., AutoSA) exhibit flexibility in accommodating 4-bit DSP packing, they suffer from resource consumption and data supply latency issues, especially when adapting to various convolution spatial sizes. This work introduces a customizable and ultra-efficient SA architectural template for 4-bit convolution, called E4SA. First, we propose a fine-grained row-temporal weight stationary dataflow that aligns with the specific requirements of 4-bit DSP full packing (4bF packing). Based on this, we design a cost-effective SA unit (SAU) composed of 4bF-packing-based processing elements (PEs) to enhance computational efficiency. This includes column-shared packed-data splitters and shift-register-based feature-map/weight fetchers to ensure continuous data supply, all of which are locally interconnected via more cost-effective registers. In addition, we develop a two-level hierarchy SA that decomposes the original large SA into parallel 4×4 SAU sets, which not only allows multiple PEs in the same column to share data splitting and reorganization logic and thus reducing the LUT overhead, but also maintains near-theoretical latency across various convolutional spatial sizes. Experimental results demonstrate that E4SA achieves up to 576.6 GOPS with 13.8× higher GOPS/DSP efficiency and 51.6× higher GOPS/kLUTs efficiency compared to 4-bit AutoSA-based design. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Junrong Zhang 0002, Weiying Xie, Yunsong Li 0001 |
FPGA | 1 |
| 2024 | SA4: A Comprehensive Analysis and Optimization of Systolic Array Architecture for 4-bit ConvolutionsabstractMany studies have demonstrated that 4-bit precision quantization can maintain accuracy levels comparable to those of floating-point deep neural networks (DNNs). Thus, it has sparked a keen interest in the efficient acceleration of such compressed DNNs, especially 4-bit convolutions, on edge devices. However, we observe that conventional systolic array (SA) architectures, widely adopted for DNN acceleration, fail to fully exploit the high computational density benefits of 4 -bit DSP packing. In this paper, we conduct the first comprehensive analysis of the integration of modern DSP packing techniques (specifically, 4-bit fully DSP packing) into the 4-bit systolic array design for convolutions. First, we introduce a row-temporal weight stationary 4-bit SA dataflow that complements the loop execution order inherent in 4-bit fully DSP packing in conventional SAs, which is called BaseSA. Next, we analyze the performance and resource efficiency of BaseSA, and identify two inefficiencies in the integration: 1) excessive LUT resource utilization that constraints the overall SA size, and 2) large latency gap to the theoretical optimum, due to various stalls in data supplies. To overcome these obstacles, we propose SA4: an HLS-based, customizable, and ultra-efficient hierarchical $\underline{\text { SA}}$ architecture optimized for 4 -bit convolutions. The core unit in SA4 is a delicately designed cost-effective SA unit (SAU), which 1) replaces the costly buffer-based data suppliers for activations and weights with shift-register-based ones, 2) replaces LUT-intensive FIFO connections between SA PEs (processing elements) with registers, and 3) replaces the finite state machines (FSM) and data unpacking logic inside each PE with a global FSM inside each SAU and a data splitter shared by a column of PEs. While such an SAU can only support a small spatial size for an SA due to its delicate design, we further scale it out using an array of SAUs. Experimental results show that our proposed SA4 achieves 1153.2 GOPS on the AMD-Xilinx Ultra96-V2 FPGA, with a $13.8 \times$ increase in GOPS/DSP efficiency and a $49 \times$ increase in GOPS/kLUTs efficiency compared to a straightforward SA and 4-bit DSP packing integration. Our SA4 project is open sourced here: https://github.com/Michaela1224/SA4. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Junrong Zhang 0002, Weiying Xie, Yunsong Li 0001 |
FPL | 1 |
| 2024 | SDA: Low-Bit Stable Diffusion Acceleration on Edge FPGAsabstractThis paper introduces SDA, the first effort to adapt the expensive stable diffusion (SD) model for edge FPGA deployment. First, we apply quantization-aware training to quantize its weights to 4 -bit and activations to 8 -bit ($W 4 A 8$) with a negligible accuracy loss. Based on that, we propose a high-performance hybrid systolic array (hybridSA) architecture that natively executes convolution and attention operators across varying quantization bit-widths (e.g., $W 4 A 8$ and all 8 -bit $Q K^{T} V$ in attention). To improve computational efficiency, hybridSA integrates diverse DSP packing techniques into hybrid weightstationary and output-stationary dataflows that are optimized for convolution and attention. It also supports flexible dataflow transitions to address the distinct demands of its output sequence by subsequent nonlinear operators. Moreover, we observe that nonlinear operators become the new performance bottleneck after the acceleration of convolution and attention, and offload them onto the FPGA as well. To reduce the latency of each nonlinear operator, we pipeline its own execution at a fine granularity. To minimize the resource utilization of nonlinear operators, we carefully balance their execution with hybridSA in a coarse-grained pipeline. Experimental results demonstrate that our low-bit ($W 4 A 8$) SDA accelerator on the embedded AMDXilinx ZCU102 FPGA achieves a speedup of $97.3 \times$ (which takes about $\mathrm{2 . 1}$ minutes for one SD inference), compared to the original SD-v1.5 model on the ARM Cortex-A53 CPU (which takes about 3.5 hours for one SD inference). Our SDA project is open sourced here: https://github.com/Michaela1224/SDA_code. Geng Yang 0001, Yanyue Xie, Zhong Jia Xue, Sung-En Chang, Yanyu Li, Peiyan Dong, Jie Lei 0001, Weiying Xie, Yanzhi Wang 0001, Xue Lin 0001, Zhenman Fang |
FPL | 1 |
| 2024 | Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR ClassificationabstractIn multimodal land cover classification (MLCC), a common challenge is the redundancy in data distribution, where task-irrelevant information from multiple modalities can hinder the effective integration of their unique features. To tackle this, we introduce the Multimodal Informative Vit (MIVit), a system with an innovative information aggregate-distributing mechanism. This approach redefines redundancy levels and integrates performance-aware elements into the fused representation, facilitating the learning of semantics in both forward and backward directions. MIVit stands out by significantly reducing redundancy in the empirical distribution of each modality’s separate and fused features. It employs oriented attention fusion (OAF) for extracting shallow local shape features across modalities in horizontal and vertical dimensions, and a Transformer feature extractor for extracting deep global features through long-range attention. We also propose an information aggregation constraint (IAC) based on mutual information, designed to remove redundant information and preserve complementary information within embedded features. Additionally, the information distribution flow (IDF) in MIVit enhances performance-awareness by distributing global classification information across different modalities’ feature maps. This architecture also addresses missing modality challenges with lightweight independent modality classifiers, reducing the computational load typically associated with Transformers. Our results show that MIVit’s bidirectional aggregate-distributing mechanism between modalities is highly effective, achieving an average overall accuracy of 95.56% across three multimodal datasets. This performance surpasses current state-of-the-art methods in MLCC. The code for MIVit is accessible at https://github.com/icey-zhang/MIViT. Jie Lei 0001, Weiying Xie, Geng Yang 0001, Daixun Li, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | HyBNN: Quantifying and Optimizing Hardware Efficiency of Binary Neural NetworksabstractBinary neural network (BNN), where both the weight and the activation values are represented with one bit, provides an attractive alternative to deploy highly efficient deep learning inference on resource-constrained edge devices. However, our investigation reveals that, to achieve satisfactory accuracy gains, state-of-the-art (SOTA) BNNs, such as FracBNN and ReActNet, usually have to incorporate various auxiliary floating-point components and increase the model size, which in turn degrades the hardware performance efficiency. In this article, we aim to quantify such hardware inefficiency in SOTA BNNs and further mitigate it with negligible accuracy loss. First, we observe that the auxiliary floating-point (AFP) components consume an average of 93% DSPs, 46% LUTs, and 62% FFs, among the entire BNN accelerator resource utilization. To mitigate such overhead, we propose a novel algorithm-hardware co-design, called FuseBNN , to fuse those AFP operators without hurting the accuracy. On average, FuseBNN reduces AFP resource utilization to 59% DSPs, 13% LUTs, and 16% FFs. Second, SOTA BNNs often use the compact MobileNetV1 as the backbone network but have to replace the lightweight 3 × 3 depth-wise convolution (DWC) with the 3 × 3 standard convolution (SC, e.g., in ReActNet and our ReActNet-adapted BaseBNN) or even more complex fractional 3 × 3 SC (e.g., in FracBNN) to bridge the accuracy gap. As a result, the model parameter size is significantly increased and becomes 2.25× larger than that of the 4-bit direct quantization with the original DWC (4-Bit-Net); the number of multiply-accumulate operations is also significantly increased so that the overall LUT resource usage of BaseBNN is almost the same as that of 4-Bit-Net. To address this issue, we propose HyBNN , where we binarize depth-wise separation convolution (DSC) blocks for the first time to decrease the model size and incorporate 4-bit DSC blocks to compensate for the accuracy loss. For the ship detection task in synthetic aperture radar imagery on the AMD-Xilinx ZCU102 FPGA, HyBNN achieves a detection accuracy of 94.8% and a detection speed of 615 frames per second (FPS), which is 6.8× faster than FuseBNN+ (94.9% accuracy) and 2.7× faster than 4-Bit-Net (95.9% accuracy). For image classification on the CIFAR-10 dataset on the AMD-Xilinx Ultra96-V2 FPGA, HyBNN achieves 1.5× speedup and 0.7% better accuracy over SOTA FracBNN. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Yunsong Li 0001, Weiying Xie |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2023 | HyBNN: Quantifying and Optimizing Hardware Efficiency of Binary Neural NetworksabstractBinary neural network (BNN) has recently presented a promising opportunity for deep learning inferences on resource-constrained edge devices. Using extreme data precision, i.e., 1-bit weight and 1-bit activation, BNN not only significantly reduces the network memory footprint, but also trades massive multiply-accumulate operations for much cheaper logical XNOR and population count operations. However, our investigation reveals that, to achieve satisfactory accuracy gains, state-of-the-art (SOTA) BNNs, such as FracBNN [4] and ReActNet [1], usually have to incorporate various auxiliary floating-point ($AFP$) components and increase the model size, which in turn degrades the hardware performance efficiency. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Yunsong Li 0001, Weiying Xie |
FCCM | 1 |
| 2023 | Guided Hybrid Quantization for Object Detection in Remote Sensing Imagery via One-to-One Self-TeachingabstractDeep convolutional neural networks (CNNs) have improved remote sensing image analysis, but their high computational demands may limit their deployment on low-end devices with limited resources, such as intelligent satellites and unmanned aerial vehicles. Considering the computation complexity, we propose a Guided Hybrid Quantization with One-to-one Self-Teaching (GHOST) framework. More concretely, we first design a structure called guided quantization self-distillation (GQSD), an innovative idea for realizing a lightweight model through the synergy of quantization and distillation. The training process of the quantization model is guided by its full-precision model, which is time-saving and cost-saving without preparing a huge pre-trained model in advance. Second, we put forward a hybrid quantization (HQ) module that automatically acquires the optimal bit-width by imposing a threshold constraint on the distribution distance between the center point and samples in the weight search space, aiming to retain more shallow detail information that is advantageous for small object detection. Third, to improve information transformation, we propose a one-to-one self-teaching (OST) module to give the student network the ability to self-judgment. A switch control machine (SCM) builds a bridge between the student and teacher networks in the same location to help the teacher reduce wrong guidance and impart vital knowledge about objects without vast background information to the student. This distillation method allows a model to learn from itself and gain substantial improvement without any additional supervision. Extensive experiments on a multimodal dataset (VEDAI) and single-modality datasets (DOTA, NWPU, and DIOR) show that object detection based on GHOST outperforms the existing detectors. The tiny parameters (<9.7 MB) and Bit-Operations (BOPs) (<2158 G) compared with any remote sensing-based, lightweight, or distillation-based algorithms demonstrate the superiority in the lightweight design domain. Our code and model will be released at https://github.com/icey-zhang/GHOST. Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Geng Yang 0001, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Algorithm/Hardware Codesign for Real-Time On-Satellite CNN-Based Ship Detection in SAR ImageryabstractRecently, the convolutional neural network (CNN)-based approach for on-satellite ship detection in synthetic aperture radar (SAR) images has received increasing attention since it does not rely on predefined imagery features and distributions that are required in conventional detection methods. To achieve high detection accuracy, most of the existing CNN-based methods leverage complex off-the-shelf CNN models for optical imagery. Unfortunately, this usually leads to expensive computational cost, which is hard to process in real time using resource-constrained devices deployed in the harsh satellite environment. In this article, we propose OSCAR-RT, the first end-to-end algorithm/hardware codesign framework for real-time on-satellite CNN-based SAR ship detection, which can simultaneously produce an accurate and hardware-friendly CNN model and an ultraefficient field-programmable gate array (FPGA)-based hardware accelerator that can be deployed on satellites. With the real-time on-satellite processing speed in mind, we start from a state-of-the-art compact CNN model for optical imagery. To eliminate the sharp decrease in the detection accuracy for SAR imagery, we analyze the discrepancy between the SAR domain and optical domain and propose to adapt the model by adjusting the output feature size to better detect relatively smaller objects in SAR imagery. To improve the detection speed, we propose to develop a fully pipelined interlayer streaming accelerator architecture, where all the layers of the CNN model can be concurrently processed using on-chip FPGA resources. To achieve this architecture, we first propose a hardware-guided, progressive, and structural pruning strategy, which is guided by our modeled hardware metrics and applies state-of-the-art coarse-grained and fine-grained filter pruning as well as mixed-precision quantization techniques. Moreover, to improve the reusability and portability of the hardware accelerator design, we develop a library of highly optimized CNN components in high-level synthesis, together with their performance and resource models. Finally, we map the pruned CNN model onto these hardware library components in a fully pipelined interlayer streaming fashion, by adjusting their parallelism factors to balance the execution of each layer and fit into the resource constraint. Experimental results using the adapted MobileNetV1, MobileNetV2, and SqueezeNet models on the widely used SAR ship detection dataset (SSDD) demonstrate the effectiveness of OSCAR-RT; for the MobileNetV1 model, it achieves an average precision of 94%, a detection speed of 652 frames/s on the Xilinx VC709 FPGA evaluation board while consuming about 5.8-W power. Geng Yang 0001, Jie Lei 0001, Weiying Xie, Zhenman Fang, Yunsong Li 0001, Xin Zhang 0092 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Rank-Aware Generative Adversarial Network for Hyperspectral Band SelectionabstractTraditional clustering-based band selection (BS) methods treat each band as individuals, and selection is conducted by enlarging the difference between clusters, which leads to the loss of band interaction and information saliency evaluation. In this article, we propose a BS method named rank-aware generative adversarial network (R-GAN) to address these problems. First, centralized reference feature extraction (FE) with GAN aids R-GAN to combine interpretability and interband relevance. Then, the reference feature is refined with the saliency estimation provided by the rank-aware strategy. According to data characteristics, there are two versions of rank computation including tensor and matrix. Finally, the structural similarity index measurement (SSIM) maps the saliency to the original data space to obtain the final BS result. Extensive comparison experiments with popular existing BS approaches on five hyperspectral images (HSIs) datasets show that the proposed R-GAN can address spectral saliency effectively and select more informative band subsets, which outperforms other competitors for both detection and classification tasks. For example, on the SD-1 dataset, the ten bands selected by R-GAN achieve 0.982 ± 0.003 with an improvement of 13.7% in the area under the curve (AUC) value of anomaly detection performance. The peaked accuracy surpasses the baseline by 0.46% for the classification on the PaviaU dataset. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001, Geng Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |