VLDB 2026 Research / reviewers in the wild / expert
Chaofang Ma
dblp:351/6344
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0009-2082-4916ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HERO: Hardware-Efficient RL-based Optimization Framework for NeRF Quantization
Yipu Zhang 0002, Chaofang Ma, Jinming Ge, Jiang Xu 0001, Wei Zhang 0012 |
ASP-DAC | 2 |
| 2026 | Grin: HyperGNN Training Framework for Efficient Edge Inference via Hypergraph RestructuringabstractHypergraph neural networks (HyperGNNs) have garnered increasing attention for their ability to model high-order relationships in various domains. However, the extremely sparse connections inherent to hypergraphs result in numerous off-chip memory accesses, posing a long-latency inference issue on edge devices. Existing hardware accelerators focus solely on exploiting the limited data reuse opportunities in hypergraphs to mitigate this issue, without addressing the underlying cause: the sparsity of the hypergraph structures themselves.To address the fundamental limitation, this paper proposes Grin, a general HyperGNN training framework. It is designed to restructure hypergraphs for enhancing inference efficiency on edge devices regardless of hardware architectures while improving model performance. Specifically, hyperedge pruning within Grin is utilized to eliminate redundant computation workloads, effectively lowering overall off-chip memory accesses. Moreover, Grin redefines the objective of traditional data augmentation by incorporating hardware efficiency alongside model accuracy. This shift enables significantly increased data reuse in the remaining computation workloads, thereby ensuring model performance and further reducing off-chip memory accesses. Experiments demonstrate that, with increased model accuracy, deploying Grin-optimized hypergraphs on the state-of-the-art (SOTA) accelerator achieves an average inference speedup of 1.41× compared to the original hypergraphs on the same accelerator, while reducing off-chip memory accesses by 27.60%. Furthermore, this deployment achieves a 14.82× speedup over the SOTA GPU-based system. Chaofang Ma, Jiang Xu 0001, Wei Zhang 0012 |
DATE | 1 |
| 2026 | DRACO: A Hardware-Efficient Robot Rigid Body Dynamics Accelerator with Precision-Aware Quantization FrameworkabstractRigid Body Dynamics (RBD) computation is a critical component of robotic control, often dominating system runtime due to its algorithmic complexity and high parallelism demands. CPUs suffer from limited parallelism and cache-unfriendly access patterns, while GPUs incur prohibitive memory-access latency and per-task response time, making them unsuitable for real-time control. Both platforms also consume excessive power for edge deployment. FPGAs offer superior latency, energy efficiency, and customizable hardware-level parallelism, emerging as promising targets for RBD acceleration. However, existing FPGA designs still face critical limitations. First, the intensive use of multiply-accumulate operations leads to high Digital Signal Processing (DSP) slices consumptionespecially for high degrees-of-freedom (DOF) robots-resulting in limited scalability. Second, RBD functions include mass matrix inversion function, which is inefficient on FPGA due to reciprocal operations falling on the longest latency path, severely limiting performance. Third, mismatched processing rates across modules introduce idle cycles, resulting in poor DSP utilization. To address these issues, we propose DRACO, a hardwareefficient and high-performance RBD accelerator based on FPGA, introducing three key innovations. First, we propose a precisionaware quantization framework that reduces DSP demand by up to$4 \times$while preserving motion accuracy. This is also the first study to systematically evaluate quantization impact on robot control and motion for hardware acceleration. Second, we leverage a hardware-efficient division deferring optimization in mass matrix inversion algorithm, which decouples reciprocal operations from the longest latency path to improve the performance. Finally, we present an inter-module DSP reuse methodology to improve DSP utilization and save DSP usage. Experiment results show that DRACO achieves up to$8 \times$throughput improvement and$7.4 \times$latency reduction over state-of-the-art (SOTA) RBD accelerators across various robot types, demonstrating its effectiveness and scalability for high-DOF robotic systems. Yipu Zhang 0002, Linfeng Du, Chaofang Ma, Jiang Xu 0001, Wei Zhang 0012 |
HPCA | 5 |
| 2025 | Robin: RWKV Accelerator using Block Circulant Matrices based on FPGAabstractRecent advancements in linear-attention models, such as RWKV, have opened up new possibilities for efficient sequence processing by reducing the computational overhead of traditional Transformer architectures. Field-programmable gate arrays (FPGAs) offer a compelling solution for deep learning applications by providing customizable hardware architectures that enhance computational efficiency and flexibility. However, deploying these models on FPGAs introduces several challenges. Previous FPGA deployment workflows tend to focus on general machine learning tasks, lacking sufficient integration between software and hardware optimizations. Besides, FPGAs are constrained by limited on-chip and off-chip memory, posing significant challenges for weight storage. Moreover, the predominance of linear operations on the GPU runtime leads to significant computational bottlenecks. These obstacles necessitate innovative solutions to bridge the performance gap between FPGAs and GPUs while preserving model accuracyTo overcome these challenges, we introduce Robin, a fine-grained FPGA accelerator workflow that integrates both algorithm-level and hardware-level optimization. Robin leverages a weight compression technique based on Partial Block Circulant Matrices (PBCM), which effectively reduces storage demands while maintaining accuracy. Based on PBCM, our design employs a configurable circulant computing core that fully exploits the bit-width efficiency of DSP48E resources through two DSP packaging strategies to support both circulant and standard matrix operations. The combined end-to-end software-hardware co-design enables Robin to achieve up to a 3.09× increase in throughput and a 7.31× boost in energy efficiency compared to high-end Tesla A100 GPU implementations, making it a compelling solution for deploying RWKV models on FPGAs. Shangkun Li, Chuyi Dai, Chaofang Ma, Wei Zhanh |
ICCAD | 4 |
| 2025 | ScanNow: A Scan Window-Based Sparse Matrix Multiplication Accelerator DesignabstractSparse matrix-matrix multiplication (SpMM) is a prevailing kernel in scientific and artificial intelligence applications. However, the irregular memory access behaviors caused by diverse sparse patterns in SpMM lead to a significant performance bottleneck for traditional computing platforms, driving the demand for dedicated hardware accelerators. Unfortunately, the existing hardware accelerators are not flexible enough to handle diverse SpMM workloads due to their rigid task dispatch schemes, limiting the overall performance.To address the challenge, this paper introduces ScanNow, a scan window-based SpMM accelerator design. ScanNow first presents a scan window-based task dispatch scheme designed to minimize the time-consuming irregular memory access behaviors by exploiting potential data reuse opportunities. The proposed task dispatch scheme effectively and dynamically balances input and output data reuse for row-wise dataflow according to real-time sparse patterns. Moreover, a tailored hardware architecture, featuring multiple memory components and customized schedulers, is proposed to accommodate the task dispatch scheme and improve computation parallelism. Experiments demonstrate that ScanNow achieves an average speedup of 1.67x across a wide range of SpMM workloads compared to the state-of-the-art hardware accelerator. Chaofang Ma, Hanwei Fan |
ICCAD | 1 |
| 2025 | FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization AccelerationabstractLegalization is a critical yet time-consuming step in very large-scale integration (VLSI) design, tasked with iteratively relocating standard cells to eliminate overlaps while resolving design rule violations. This process is repeatedly invoked during VLSI physical design. However, increasing spatial constraints and complex design rules impose significant challenges on existing CPU- and GPU-based legalizers, including suboptimal task assignment, inefficient algorithm, and long hardware idle time caused by processing tasks with irregular computational patterns in parallel. Linfeng Du, Yipu Zhang 0002, Chaofang Ma, Hanwei Fan, Jiang Xu 0001, Wei Zhang 0012 |
ICPP | 5 |
| 2023 | Online Reliability Evaluation Design: Select Reliable CRPs for Arbiter PUF and Its VariantsabstractPhysical Unclonable Function (PUF) is a hardware security primitive with broad application prospects. Variants of the arbiter PUF have been proposed to resist modeling attacks. However, their low reliability issue limits their applications. To solve the low reliability issue, this paper proposes an Online Reliability Evaluation (ORE) design for the arbiter PUF and its variants. Moreover, a corresponding machine learning method to select reliable Challenge Response Pairs (CRPs) for applications is proposed. Based on the ORE design, a small number of CRPs and their reliability levels are collected during the enrollment phase. Then they are trained to build reliability models for predicting the responses and reliability levels of other challenges. Since the ORE design does not change the security structures of the arbiter PUF and its variants, the resistance to modeling attacks of PUF designs equipped with it is maintained. Compared to the previous work that tests 100,000 times per CRP, our design is time-saving in the enrollment phase since each CRP is only tested three times for training reliability models. The proposed design is implemented under the 40nm process. Experimental results on real chips show that all the CRPs selected by our reliability models are indeed reliable for applications, verifying the effectiveness of our method. Chaofang Ma, Jianan Mu, Jing Ye 0001, Yuan Cao 0003, Huawei Li 0001, Xiaowei Li 0001 |
ETS | 1 |