Huidong Ji

dblp:385/9544 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0005-7202-1681ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Approx-L: An Error-Balanced Approximate Floating-Point Divider with Multi-Level Linear Compensation
Shangshang Yao, Huidong Ji, Zuoning Chen
ACM Great Lakes Symposium on VLSI2
2025 A Computation and Energy Efficient Hardware Architecture for SSL Acceleration
abstract
In Computer Vision (CV), the deployment of Convolutional Neural Networks (CNNs) is often hindered by their substantial computational requirements and large labeled datasets. Self-supervised learning (SSL) serves as an effective approach to reducing the reliance on labeled data with the option of augmentation methods to infer and train CNNs. Excluding irrelevant features accelerates learning and improves optimization. We propose a Field-Programmable Gate Array (FPGA)-based hardware accelerator architecture tailored for SSL framework, leveraging its parallelism and reconfigurability to expedite block matching, optimize sparse convolutions, and manage data reuse, significantly improving resource and energy efficiency. The implementation and evaluation of our work on Xilinx ZCU102 FPGA working at 200 MHz confirm that the similarity finding part's FPGA accelerations with a low hardware overhead generates a latency of 0.0106 seconds, surpassing GPU and CPU, and in the sparse CNN's FPGA acceleration part, with the processing of VGG16 and ResNet50, compared with the related FPGA-based works, our design claims a maximum of 3.08× throughput improvement and 1.5× in energy efficiency.
Huidong Ji, Sheng Li 0019, Chen Ding 0010, Jiawei Xu 0001, Qitao Tan, Jun Liu 0075, Ao Li 0004, Xulong Tang, Lirong Zheng 0001, Geng Yuan, Zhuo Zou
ASP-DAC1
2025 DuoQ: A DSP Utilization-aware and Outlier-free Quantization for FPGA-based LLMs Acceleration
abstract
Quantization enables efficient deployment of large language models (LLMs) on FPGAs, but its presence of outliers affects the accuracy of the quantized model. Existing methods mainly deal with outliers through channel-wise or tokenwise isolation and encoding, which leads to expensive dynamic quantization. To address this problem, we introduce DuoQ, an FPGA-oriented algorithm-hardware co-design framework. DuoQ effectively eliminates outliers through learnable equivalent transformations and low-semantic token awareness in the quantization scheme part, facilitating per-tensor quantization with 4 -bits. We co-design the quantization algorithm and hardware architecture. Specifically, DuoQ accelerates end-to-end LLM through a novel DSP-aware PE unit design and encoder design. In addition, two types of post-processing units assist in the realization of nonlinear functions and dynamic token awareness. Experimental results show that compared with platforms with different architectures, DuoQ’s computational efficiency and energy efficiency are improved by up to $8.8 \times$ and $23.45 \times$. In addition, DuoQ has achieved accuracy improvements compared to other outlieraware software and hardware works.
Zhuoquan Yu, Huidong Ji, Junfu Wu, Xiaoze Yan, Lirong Zheng 0001, Zhuo Zou
DAC2
2025 Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
abstract
Zeroth-order (ZO) optimization is an emerging deep neural network (DNN) training paradigm that offers computational simplicity and memory savings. However, this seemingly promising approach faces a significant and long-ignored challenge. ZO requires generating a substantial number of Gaussian random numbers, which poses significant difficulties and even makes it infeasible for hardware platforms, such as FPGAs and ASICs. In this paper, we identify this critical issue, which arises from the mismatch between algorithm and hardware designers. To address this issue, we proposed PeZO, a perturbation-efficient ZO framework. Specifically, we design random number reuse strategies to significantly reduce the demand for random number generation and introduce a hardware-friendly adaptive scaling method to replace the costly Gaussian distribution with a uniform distribution. Our experiments show that PeZO reduces the required LUTs and FFs for random number generation by 48.6% and 12.7%, and saves at maximum 86% power consumption, all without compromising training performance, making ZO optimization feasible for on-device training. To the best of our knowledge, we are the first to explore the potential of on-device ZO optimization, providing valuable insights for future research.
Qitao Tan, Sung-En Chang, Huidong Ji, Chence Yang, Ci Zhang, Jun Liu 0075, Zheng Zhan 0001, Zhenman Fang, Zhuo Zou, Yanzhi Wang 0001, Jin Lu 0001, Geng Yuan
ICCAD4
2025 SAIndust: A Self-Aware Heterogeneous Computing Framework for Industrial Internet of Things
abstract
Distributed collaborative automation and resource scheduling are important for improving the productivity of intelligent manufacturing in the Industrial Internet of Things (IIoT). However, current efforts at the edge layer, where a large number of operations converge and device interactions are concentrated, are inadequate in dealing with the resulting computational heterogeneity and dynamic changes in the operating environment. To address these issues, we propose a self-aware heterogeneous computing framework (SAIndust). First, we design and implement a fine-grained heterogeneous resource virtualization technology based on Kubernetes, which pools computing resources and implements circulation to improve resource utilization. Then, we design a self-aware method that drives distributed system state update and scheduling, which is an autonomic optimization framework for real-time scheduling. Finally, we build a physical prototype platform and develop a practical plug-and-play deployment and evaluation tools. Experiments with deep learning applications with different resource intensities show that its 1.54% and 1.85% GPU virtualization overheads and standard deviation of resource allocation can achieve good virtualization performance and high fidelity. On the other hand, while achieving a 56.9% reduction in the average age of information and only a 25.6% increase in the average CPU cost, SAIndust can reduce the resource saturation by an average of 8.71% and achieve a maximum throughput increase of 5.12× compared to related methods in medium-scale to ultra-large-scale edge clusters.
Zhuoquan Yu, Jichao Leng, Huidong Ji, Lirong Zheng 0001, Zhuo Zou
IEEE Internet Things J.4