Han Jiao 0003

dblp:38/10222-3 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HFRWKV: A High-Performance Fully On-Chip Hardware Accelerator for RWKV
abstract
RWKV achieves Transformer-level performance with linear memory complexity for long sequences. However, its inherent sequentiality limits hardware parallelism and triggers memory-bound bottlenecks via frequent off-chip weight accesses. To address these challenges, we propose HFRWKV, an FPGA-based hardware accelerator for RWKV. Within the matrix operation module, we propose a novel Δ-PoT hardware-friendly hybrid-precision quantization strategy, which enhances performance while maintaining acceptable accuracy. For the complex operations including exponentiation and division, we introduce a method featuring reusable architectures combined with lookup tables or piecewise linear approximation, which is algorithmically refined to effectively balance precision and hardware resource consumption. We adopt a fully on-chip computing system integrating parallel matrix-vector array and efficient pipelining, employing computation reordering and chunked double buffering to eliminate data bottlenecks and maximize throughput. We implement HFRWKV on the Alveo U50 and U280 platform. Experimental results show that compared to a CPU, a throughput improvement of 63.48× and an energy efficiency improvement of 139.17×. Compared to GPUs, achieves a throughput improvement of 32.33× and an energy efficiency improvement of 171.36×.
Zhenghao Zeng, Han Jiao 0003, Yihua Huang 0005
FPGA3
2026 RRAM: A Reconfigurable Request-Rate-Aware Multi-DNN Accelerator Based on FPGA
abstract
The latency-sensitive image recognition systems in edge data centers are facing the workloads of multiple deep neural networks (Multi-DNN) and multi-user dynamic requests. Most of the existing solutions are based on GPUs or ASICs, whose fixed hardware architecture has limited low-latency capabilities when dealing with varying request rates in random dynamic scenarios. The dynamic reconfigurability of Field-Programmable Gate Arrays (FPGAs) can enable the hardware architecture to adapt in accordance with fluctuating request rates. Therefore, this study proposes a Request-Rate-Aware Multi-DNN accelerator framework (RRAM) based on FPGA, thereby improving latency performance. The RRAM includes a reconfigurable, computation-memory-interleaved multi-core accelerator architecture, a dynamic bandwidth adaptor for stable data transfers, and a performance analytic model for random request scenarios that enables an end-to-end latency evaluation. To efficiently explore the design space, a workload-aware resource boundary search strategy is also introduced, which eliminates a majority of invalid solutions. Implemented on the U250 platform, RRAM achieves better Service-Level Agreement (SLA) satisfaction rates compared to the fixed-architecture baseline accelerators under varying request rates, and improves the Average Normalized Turnaround Time (ANTT) by 1.3x to 6.1x. Furthermore, the average power consumption of RRAM is only 22.8% and 44.9% of the GPUs of RTX3090 and A100, while the latency is improved by up to 18.2x and 4.6x.
Han Jiao 0003, Wenjin Huang, Kailing Zhou, Dinghua Xu, Zhiyong Pang, Yihua Huang 0005
IEEE Internet Things J.1
2026 Relationship-Experts Transformer for Image Captioning
abstract
Image captioning is a cross-modal text generation task aimed at understanding the relationships among various objects in an image. Therefore, accurately expressing object–object relations remains a key bottleneck for transformer-based image captioning. Prior methods usually inject semantic and geometric relations once and keep them fixed while only updating visual features, creating a mismatch—evolving visuals vs. frozen relations—that weakens relational guidance and leads to feature entanglement. We propose the Relationship-Experts Transformer (RET), which treats semantic and geometric relations as learnable experts that guide object visual features (students) and co-evolve with them. In RET, we first design the Relationship-Guided Feature Aggregation (RGFA) module, which is analogous to experts-guided student learning, specifically utilizing the relationship kernel (the expert’s knowledge brain) to guide the learning of the object visual features (students). Secondly, we develop the Experts Knowledge Updating (EKU) module, which continuously iterates expert knowledge during training to enhance the expert’s guiding ability over the student. Finally, we design the Student Knowledge Selector (SKS) module to adaptively select object visual features enhanced with different relations under the guidance of semantic and geometric experts to generate descriptive texts embodying semantic and geometric knowledge. Experiments on the MSCOCO dataset demonstrate that our model achieves state-of-the-art performance. All codes are available at https://github.com/songchuanle-1/RET .
Chuanle Song, Wenjin Huang, Han Jiao 0003, Yihua Huang 0005
ACM Trans. Multim. Comput. Commun. Appl.3
2025 An Efficient FPGA-Based Hardware Accelerator of Fully Quantized Mamba-2
abstract
The Mamba-2 model introduces a State Space Duality (SSD) mechanism, based on original State Space Models (SSMs), that accelerates training and improves accuracy. However, efficient hardware acceleration for Mamba-2 faces challenges. Numerous element-wise operations fail to fully utilize GPU tensor cores, diminishing inference efficiency. Furthermore, research on full-quantization strategies for Mamba-2 is lacking. To address this, we propose a hybrid-precision full-quantization strategy, Hfqmamba2, balancing performance and hardware resource usage. Applying this strategy to Mamba-2 models of various sizes shows that accuracy loss remains within an acceptable range. We also propose an efficient FPGA-based hardware accelerator for Mamba-2. Given the distinct data flow characteristics of the two RMSNorm (Root Mean Square Normalization) layers in the hardware implementation, we introduce a reconfigurable hardware architecture based on a segmented quantization strategy, improving efficiency and flexibility by using a segmented lookup table to approximate the inverse square root operation. For the selective SSM layer operations, we design an intra-layer computation pipeline to enhance processing efficiency. Through design space exploration, we configure two versions of the hardware accelerator and evaluate their performance on the Alveo U50 platform. Experimental results show that both configurations achieve 99.63 % bandwidth utilization. Compared to the CPU, the hardware accelerator achieves a 114.05× speedup and a 282.75× improvement in energy efficiency. It also outperforms the PyTorch implementation on the GPU, achieving a 29.81 ×speedup and a 297.87 ×improvement in energy efficiency. Additionally, the hardware implementation shows a 1.94 ×speedup and a 35.89 ×improvement in energy efficiency over the official CUDA-accelerated version.
Kailing Zhou, Han Jiao 0003, Wenjin Huang, Yihua Huang 0005
FCCM2
2025 EViL: An Efficient Vision-LSTM Accelerator Based on FPGA
abstract
Recently, the Vision-LSTM (ViL) model, built upon Extended Long Short-Term Memory (xLSTM) building blocks, has attracted widespread attention due to its excellent performance and its linear complexity. Due to the unified compute architecture of GPUs, efficiently deploying the ViL layer on GPU is challenging because of its complex dataflow, thereby necessitating custom hardware accelerators. Moreover, significant differences in the activation distribution and quantization sensitivity among modules, making most existing quantization methods for CNNs not suitable for ViL layers. To address these challenges, we propose an intra-layer mixed-precision quantization method, termed ClipQuant, for the ViL layer. By introducing variable quantization range parameters$\alpha$and scaling parameters$\beta$, assessing the quantization sensitivity of each module, and imposing suitable quantization parameter constraints, we achieve near-lossless quantization of the ViL layer (accuracy loss$1.10 \times \sim 1.49 \times$performance improvement and a$36.95 \times \sim 45.50 \times$improvement in energy efficiency.
Zexuan Deng, Han Jiao 0003, Wenjin Huang, Yihua Huang 0005
FPL2
2025 An FPGA-based Quantization and Acceleration Framework for Multi-DNN
abstract
Deep neural networks occupy an irreplaceable place in both modern consumer and industrial sectors. The advent of INFerence-as-a-Service (INFasS) by cloud providers facilitates end-users in various domains. Numerous studies have tailored accelerators to boost the efficiency of multi-DNN inference towards these scenarios. However, current research seldom employs quantization techniques to improve the performance of multi-DNN accelerator. FPGAs are characterized by high parallelism and flexible reconfigurability, with their heterogeneous resources enabling support of diverse precision levels. In this context, we propose a framework for quantizing and accelerating multi-DNN on FPGA, including 1. a hardware-aware inter-layer mixed-precision quantization algorithm tailored for multi-DNN inference, 2. an FPGA architecture that facilitates ultra-low bit-width mixed-precision quantization, and 3. a low-overhead scheduling algorithm. Our co-design improves the throughput by up to 2.19x and reduces the response time by up to 28% compared to a unified-precision baseline. In addition, our design achieves similar DSP efficiency and better energy efficiency compared to related work.
Han Jiao 0003, Wenjin Huang, Yihua Huang 0005
ISCAS2
2025 HCG: Streaming DCNN Accelerator With a Hybrid Computational Granularity Scheme on FPGA
abstract
With the growth of field-programmable gate array (FPGA) hardware resources, streaming DCNN accelerators leverage interconvolutional-layer parallelism to enhance throughput. In existing streaming accelerators, convolution nodes typically adopt layer- or column-based tiling methods, where the tiled input feature map (Ifmap) encompasses all input channels. This approach facilitates the comprehensive calculation of the output feature map (Ofmap) and maximizes interlayer parallelism. The computational granularity, defined in this study as the calculated rows or columns of Ofmap based on each tiled Ifmap data, significantly influences on-chip Ifmap storage and off-chip weight bandwidth (BW). The uniform application of computational granularity across all nodes inevitably impacts the memory-BW tradeoff. This article introduces a novel streaming accelerator with a hybrid computational granularity (HCG) scheme. Each node employs an independently optimized computational granularity, enabling a more flexible memory-BW tradeoff and more effective utilization of FPGA resources. However, this hybrid scheme can introduce pipeline bubbles and increase system pipeline complexity and control logic. To address these challenges, this article theoretically analyzes the impact of computational granularity on individual computing nodes and the overall system, aiming to establish a seamless system pipeline without pipeline bubbles and simplify system design. Furthermore, the article develops a hardware overhead model and employs a heuristic algorithm to optimize computational granularity for each computing node, achieving optimal memory-BW tradeoff and higher throughput. Finally, the effectiveness of the proposed design and optimization methodology is validated through the implementation of a 3-TOPS ResNet-18 accelerator on the Alveo U250 development board under BW constraints of 25, 20, and 15 GB/s. Additionally, accelerators for 4-TOPS VGG-16, 4-TOPS ResNet-34, 5-TOPS ResNet-50, 3-TOPS MobileNetV1, 4-TOPS ConvNeXt-T, and 4-TOPS ResNeXt-50 are implemented, surpassing the performance of most existing works.
Wenjin Huang, Conghui Luo, Baoze Zhao, Han Jiao 0003, Yihua Huang 0005
IEEE Trans. Neural Networks Learn. Syst.4
2025 HIN: Hierarchical Interaction Network for Image Captioning
abstract
The purpose of the image captioning task is to understand the content of an image and generate corresponding descriptive text. Traditional approaches to image captioning typically generate descriptive text by extracting different types of visual features from an image and performing feature interactions. However, these methods often fail to fully exploit the interactions between different types of visual features, leading to suboptimal feature integration. To address this limitation, we propose a novel Hierarchical Interaction Network (HIN) , designed to continuously extract and interact with different types of visual features to perform more effective multilevel feature interactions. Our HIN consists of three key modules: firstly, we design the Cross-Type Feature Alignment (CTFA) encoder, which aligns different types of visual features by three global features, so that the subsequent modules can effectively carry out the Hierarchical Interaction (HI) ; secondly, the HI module, which utilizes different types of multilevel features output from the encoder to carry out feature interactions and information mining, so as to generate fully mined multilevel features. The Bottom-up Gated Attention Fusion (BGAF) decoder is finally designed to perform the multilevel decoding of the features mined by our HI module, further enhancing the feature interaction capabilities of our HIN. Moreover, additional experiments on the MS-COCO dataset show that our model achieves new state-of-the-art performance. All codes are available at https://github.com/songchuanle-1/HIN .
Chuanle Song, Wei Zhou 0042, Han Jiao 0003, Wenjin Huang, Yihua Huang 0005
ACM Trans. Multim. Comput. Commun. Appl.3