EDBT 2026 Demo / reviewers in the wild / expert
Huanlin Luo
dblp:257/8789
· DBLP profile ↗
5ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-2802-2725ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Energy-Efficient FPGA-Based Vision Transformer Accelerator via Software-Hardware Co-DesignabstractThe ViT models have a large number of parameters and intensive matrix computations, making it challenging to deploy on resource-constrained FPGAs for acceleration. In this paper, we propose an end-to-end quantized ViT accelerator that adopts a multi-kernel architecture and a time-multiplexed scheduling strategy. We implement customized hardware for key operations in ViTs, enabling efficient resource utilization and memory-friendly features. Experiments show that compared to Edge-MoE, our accelerator achieves an 8.64× improvement in energy efficiency. Meanwhile, hardware resource consumption is significantly reduced, making it more suitable for resource-constrained FPGAs. Jiacheng Cao, Huanlin Luo |
FCCM | 4 |
| 2025 | EQViTA: an End-To-End Quantized Vision Transformer Accelerator Implemented on Resource-Constrained FPGAsabstractVision Transformer (ViT) has achieved great success in computer vision tasks, and FPGA-based ViT inference acceleration has recently gained widespread attention. However, the massive number of parameters and intensive matrix computations make it challenging to accelerate ViT models on resource-constrained FPGAs. To address this challenge, prior works have explored ViT quantization and approximate implementations of non-linear operations, but significant hardware resource consumption and potential performance optimization opportunities remain. In this paper, we propose an end-to-end quantized ViT accelerator, EQViTA. Its multi-kernel architecture and time-multiplexed scheduling strategy enable efficient resource utilization and memory-friendly features. We customize designs for the key operators of ViTs. First, we compress the model using INT4 quantization and implement the dataflow of the quantized model through resource-optimized dequantization. Second, we eliminate the convolution hardware module by using the convolution-to-linear mapping, thereby reducing resource usage. Finally, we achieve low-cost and highly parallel acceleration of self-attention and linear transformations through efficient exponential approximation for softmax and an adaptive linear transformation engine. Experiments on the Xilinx ZCU106 FPGA show that compared with state-of-the-art works, EQViTA achieves$1.03 \times$to$8.64 \times$improvements in energy efficiency and$1.14 \times$to$18.6 \times$improvements in normalized throughput. Meanwhile, EQViTA significantly reduces LUT, FF, BRAM, and DSP resource consumption, making it more suitable for resource-constrained FPGAs. Compared to the full-precision DeiT-Tiny model implemented with PyTorch, the INT4 quantized model deployed on EQViTA exhibits a 5.93 % accuracy drop. Jiacheng Cao, Huanlin Luo, Jian Wang 0036, Jinmei Lai 0001 |
FPL | 4 |
| 2022 | An Effective Test Method for Block RAMs in Heterogeneous FPGAs Based on a Novel Partial Bitstream Relocation TechniqueabstractBlock RAMs (BRAMs) play an important role in modern heterogenous FPGAs, hence how to test them comprehensively and effectively becomes a major concern. On-chip Partial Bitstream Relocation (PBR) technique based on FPGA Dynamic Partial Reconfiguration (DPR) can decrease the time spent on configuring modules in FPGA while reducing the memory resources overhead for storing partial bitstreams of the reconfigurable modules. The previous PBR technique is difficult to be combined with BRAM test directly, because they are somehow tedious, unsuitable for large-scale design or limited to specific devices. Besides, the problem exists for BRAM testing is that fault model is still incomplete and testing algorithms need to be improved to achieve higher fault coverage. An Effective BRAM test method based on a novel PBR technique is proposed in this paper. Our test method establishes a complete fault model for BRAM and improves the testing algorithms for faults in BRAM ECC circuits and intra-word coupling faults in SRAM cells. On-board experiments are carried out with Xilinx xc7vx690t device, and 14 BRAM configurations are used to fully test BRAMs. In conjunction with the proposed PBR technique, the number of configurations can be reduced to 10, which leads to a 35.7% time saving. Changpeng Sun, Huanlin Luo, Jiafeng Liu, Jian Wang 0036, Jinmei Lai 0001, Gang Qu 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | 3-D Auxiliary Classifier GAN for Hyperspectral Anomaly Detection via Weakly Supervised LearningabstractHyperspectral anomaly detection (AD) is important in Earth observation and remote sensing. However, the low spatial resolution of hyperspectral images, insufficient samples and lack of prior information limit the detection accuracy. To solve these problems, in this paper, we propose an auxiliary classifier generative adversarial network model based on a three-dimensional (3D) convolutional neural network named 3D AC-GAN. Firstly, the model is based on a 3D convolutional neural network design, with 3D tensors as samples. The network maintains valuable image spatial spectrum joint features to achieve good detection results. It can also generate sufficient samples to achieve dataset augmentation, solving the overfitting problem in GAN training. Secondly, we train the model with a weakly supervised method. The label of the samples is obtained through the coarse scanning method. Then, the AC-GAN is trained with the bootstrapping method to mitigate the impact of noise labels. The experimental results show that our proposed algorithm outperforms state-of-the-art AD algorithms. Huanlin Luo, Haowen Zhu, Shengyang Liu, Xinzhong Zhu, Jinmei Lai 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | DSSNet: A Simple Dilated Semantic Segmentation Network for Hyperspectral Imagery ClassificationabstractDeep learning-based methods have presented a promising performance in the task of hyperspectral imagery classification (HSIC). However, recent methods usually are considered HSIC as a patchwise image classification problem and addressed it by giving a single label to the patch surrounding a pixel. In this letter, we propose a new semantic segmentation network that can directly label each pixel in an end-to-end manner. Compared with patchwise models, our method can significantly improve training effectiveness and reduce some manual parameters. Another challenge in HSIC is that the spatial resolution of hyperspectral imagery is relatively low; in that case, the pooling operation may result in resolution and coverage loss. To address this issue, we introduce dilated convolution to our model and construct a dilated semantic segmentation network (DSSNet). Different from some existing works, DSSNet is specially designed for HSIC without complicated architecture, and no pretrained models are required. The joint spatial-spectral information can be extracted via an end-to-end manner and, thus, avoid various preprocessing or postprocessing operations. Experiments on two public data sets have demonstrated the effectiveness of our improvements compared with some of the latest deep learning-based HSIC models. Bin Pan, Zhenwei Shi 0001, Huanlin Luo, Xianchao Lan |
IEEE Geosci. Remote. Sens. Lett. | 5 |