Hanqiu Chen

dblp:330/9889 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 7 first-author · 9 since 2021
YearPublicationVenuePosition
2026 ForgeBench: A Machine Learning Benchmark Suite and Auto-Generation Framework for Next-Generation HLS Tools
abstract
High-Level Synthesis (HLS) has gained traction in hardware design, yet two key limitations prevent it from becoming mainstream. ❶ First , existing HLS benchmarks are outdated and unrepresentative of modern machine learning (ML) workloads. Legacy suites such as MachSuite [1] and Rosetta [2] rely on outdated algorithms, while Rodinia [3] and Poly-Bench [4] were designed for GPUs and software compilers, respectively, and thus offer limited relevance to HLS evaluation. Even recent efforts like HLSyn [5] exclude modern architectures such as Large Language Models (LLMs). Meanwhile, the research community increasingly evaluates HLS tools using advanced ML models through ad-hoc, private implementations, making fair and reproducible comparisons impossible without a standardized, open-source benchmark. ❷ Second , current HLS tools are accelerator-oriented , targeting simple standalone functions and generating one-off accelerators via static dataflow analysis. However, modern ML hardware design is architecture-oriented , requiring general-purpose architectures with flexible dataflow control, multi-application mapping, and complex design hierarchies featuring shared module extraction and reuse. This mismatch between tools and design practice limits HLS productivity. To address these limitations, we propose ForgeBench , an ML-centric HLS benchmark suite and generation framework for next-generation HLS tool design and evaluation. Our key contributions are: (1) an open-source, extensible HLS design generation framework supporting easy integration of new designs and applications; (2) over 10,000 high-quality, diverse ML-representative HLS benchmarks that expose limitations of modern HLS frameworks on ML workloads; and (3) an architecture-oriented benchmark suite featuring pairs of HLS designs with manually implemented module-level reuse and mappings, serving as a baseline for evaluating future modular HLS tools. We demonstrate the diversity of our generated designs across latency and resource usage, and show the potential for module reuse within HLS designs. ForgeBench is open-sourced at https://github.com/hchen799/ForgeBench .
Andy Wanna, Hanqiu Chen, Cong Hao
FCCM2
2026 COMETS: Cost-effective Multi-node Efficient Training System with Memory Pooling and Sharing
Hanqiu Chen, Shao-Peng Yang 0001, Mohammadreza Soltaniyeh, Shuyi Pei, Bryan S. Kim, Cong Hao
ICS1
2024 ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture Model
abstract
Compute Express Link (CXL) emerges as a solution for wide gap between computational speed and data communication rates among host and multiple devices. It fosters a unified and coherent memory space between host and CXL storage devices such as such as Solid-state drive (SSD) for memory expansion, with a corresponding DRAM implemented as the device cache. However, this introduces challenges such as substantial cache miss penalties, sub-optimal caching due to data access granularity mismatch between the DRAM "cache" and SSD "memory", and inefficient hardware cache management. To address these issues, we propose a novel solution, named ICGMM, which optimizes caching and eviction directly on hardware, employing a Gaussian Mixture Model (GMM)-based approach. We prototype our solution on an FPGA board, which demonstrates a noteworthy improvement compared to the classic Least Recently Used (LRU) cache strategy. We observe a decrease in the cache miss rate ranging from 0.32% to 6.14%, leading to a substantial 16.23% to 39.14% reduction in the average SSD access latency. Furthermore, when compared to the state-of-the-art Long Short-Term Memory (LSTM)-based cache policies, our GMM algorithm on FPGA showcases an impressive latency reduction of over 10,000 times. Remarkably, this is achieved while demanding much fewer hardware resources.
Hanqiu Chen, Yitu Wang, Vitorio Cargnini, Mohammadreza Soltaniyeh, Gongjin Sun, Pradeep Subedi, Yiran Chen 0001, Cong Hao
DAC1
2024 Residual-INR: Communication Efficient On-Device Learning Using Implicit Neural Representation
abstract
Edge computing is a distributed computing paradigm that collects and processes data at or near the source of data generation. The on-device learning at edge relies on device-to-device wireless communication to facilitate real-time data sharing and collaborative decision-making among multiple devices. This significantly improves the adaptability of the edge computing system to the changing environments. However, as the scale of the edge computing system is getting larger, communication among devices is becoming the bottleneck because of the limited bandwidth of wireless communication leads to large data transfer latency. To reduce the amount of device-to-device data transmission and accelerate on-device learning, in this paper, we propose Residual-INR, a fog computing-based communication-efficient on-device learning framework by utilizing implicit neural representation (INR) to compress images/videos into neural network weights. Residual-INR enhances data transfer efficiency by collecting JPEG images from edge devices, compressing them into INR format at the fog node, and redistributing them for on-device learning. By using a smaller INR for full image encoding and a separate object INR for high-quality object region reconstruction through residual encoding, our technique can reduce the encoding redundancy while maintaining the object quality. Residual-INR is a promising solution for edge on-device learning because it reduces data transmission by up to 5.16 × across a network of 10 edge devices. It also facilitates CPU-free accelerated on-device learning, achieving up to 2.9 × speedup without sacrificing accuracy. Our code is available at: https://github.com/sharc-lab/Residual-INR.
Hanqiu Chen, Xuebin Yao, Pradeep Subedi, Cong Hao
ICCAD1
2024 Survey of Machine Learning for Software-assisted Hardware Design Verification: Past, Present, and Prospect
abstract
With the ever-increasing hardware design complexity comes the realization that efforts required for hardware verification increase at an even faster rate. Driven by the push from the desired verification productivity boost and the pull from leap-ahead capabilities of machine learning (ML), recent years have witnessed the emergence of exploiting ML-based techniques to improve the efficiency of hardware verification. In this article, we present a panoramic view of how ML-based techniques are embraced in hardware design verification, from formal verification to simulation-based verification, from academia to industry, and from current progress to future prospects. We envision that the adoption of ML-based techniques will pave the road for more scalable, more intelligent, and more productive hardware verification.
Nan Wu 0009, Hanqiu Chen, Steve Dai, Cong Hao, Cunxi Yu, Yuan Xie 0001
ACM Trans. Design Autom. Electr. Syst.4
2023 DGNN-Booster: A Generic FPGA Accelerator Framework For Dynamic Graph Neural Network Inference
abstract
Dynamic Graph Neural Networks (DGNNs) are becoming increasingly popular due to their effectiveness in analyzing and predicting the evolution of complex interconnected graph-based systems. However, hardware deployment of DGNNs still remains a challenge. First, DGNNs do not fully utilize hardware resources because temporal data dependencies cause low hardware parallelism. Additionally, there is currently a lack of generic DGNN hardware accelerator frameworks, and existing GNN accelerator frameworks have limited ability to handle dynamic graphs with changing topologies and node features. To address the aforementioned challenges, in this paper, we propose DGNN-Booster, which is a novel Field-Programmable Gate Array (FPGA) accelerator framework for real-time DGNN inference using High-Level Synthesis (HLS). It includes two different FPGA accelerator designs with different dataflows that can support the most widely used DGNNs. We showcase the effectiveness of our designs by implementing and evaluating two representative DGNN models on ZCU102 board and measuring the end-to-end performance. The experiment results demonstrate that DGNN-Booster can achieve a speedup of up to 5.6× compared to the CPU baseline (6226R), 8.4× compared to the GPU baseline (A6000) and 2.1× compared to the FPGA baseline without applying optimizations proposed in this paper. Moreover, DGNN-Booster can achieve over 100× and over 1000× runtime energy efficiency than the CPU and GPU baseline respectively. Our implementation code and on-board measurements are publicly available at https://github.com/sharc-lab/DGNN-Booster.
Hanqiu Chen, Cong Hao
FCCM1
2023 Hardware/Software Co-design for Machine Learning Accelerators
abstract
This abstract highlights challenges in machine learning accelerator design and proposes solutions through software/hardware co-design techniques. To optimize single object detection, we introduce Mask-Net, a lightweight network that eliminates redundant computation. To address hardware limitations in Dynamic Graph Neural Networks (DGNNs), we present DGNN-Booster, a graph-agnostic FPGA accelerator. Our designs are open-source, generic, and applicable to real-world scenarios.
Hanqiu Chen, Cong Hao
FCCM1
2023 Rapid-INR: Storage Efficient CPU-Free DNN Training Using Implicit Neural Representation
abstract
Implicit Neural Representation (INR) is an innovative approach for representing complex shapes or objects without explicitly defining their geometry or surface structure. Instead, INR represents objects as continuous functions. Previous research has demonstrated the effectiveness of using neural networks as INR for image compression, showcasing comparable performance to traditional methods such as JPEG. However, INR holds potential for various applications beyond image compression. This paper introduces Rapid-INR, a novel approach that utilizes INR for encoding and compressing images, thereby accelerating neural network training in computer vision tasks. Our methodology involves storing the whole dataset directly in INR format on a GPU, mitigating the significant data communication overhead between the CPU and GPU during training. Additionally, the decoding process from INR to RGB format is highly parallelized and executed on-the-fly. To further enhance compression, we propose iterative and dynamic pruning, as well as layer-wise quantization, building upon previous work. We evaluate our framework on the image classification task, utilizing the ResNet-18 backbone network and three commonly used datasets with varying image sizes. Rapid-INR reduces memory consumption to only 5% of the original dataset size and achieves a maximum 6× speedup over the PyTorch training pipeline, as well as a maximum 1.2× speedup over the DALI training pipeline, with only a marginal decrease in accuracy. Importantly, Rapid-INR can be readily applied to other computer vision tasks and backbone networks with reasonable engineering efforts. Our implementation code is publicly available at https://github.com/sharc-lab/Rapid-INR.
Hanqiu Chen, Stephen B. R. Fitzmeyer, Cong Hao
ICCAD1
2022 Mask-Net: A Hardware-efficient Object Detection Network with Masked Region Proposals
abstract
Object detection on embedded systems is challenging because it is hard to achieve real-time inference with low energy consumption and limited hardware resources. Another challenge is to find hardware-friendly methods to avoid redundant computation. To address these challenges, in this work, we propose Mask-Net, a hardware-efficient object detection network with masked region proposals in regular shapes. First, we propose a hardware-friendly region proposal method to avoid redundant computation as much as possible and as early as possible, with slight or no accuracy loss. Second, we demonstrate that our method is generalizable by applying it to several detection backbones including SkyNet, ResNet-18 and UltraNet. Our method performs well in different scenarios, including DAC-SDC dataset, UAV123 dataset and OTB100 dataset. We choose SkyNet as our base model to design an accelerator and verify our design on Xilinx ZCU106 FPGA. We observe a speedup of 1.3× and about 30% energy consumption reduction when the FPGA runs at different frequencies from 124 MHz to 214 MHz with only a slight accuracy loss. We also conduct a design space exploration and demonstrate that our accelerator can achieve a theoretical speedup of 1.76× with masked region proposals. This is achieved by optimally allocating DSPs to different parts of the accelerator to balance the computations before and after the mask.
Hanqiu Chen, Cong Hao
ASAP1