Chengyu Ma

dblp:32/7220 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MoEA: A Mixed-Precision Edge Accelerator for CNN-MSA Models with Fine-Tuning Support
Qiwei Dang, Chengyu Ma, Zhiwang Huo, Guoming Yang, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
ASP-DAC2
2026 Constructions of Maximally Recoverable Codes with Two Global Parities and Length q + 1
Jie Hao 0001, Chengyu Ma, Kenneth W. Shum
ISIT2
2026 Hierarchical-ISA Supporting Row-Wise Operands for Efficient DNN Computation
abstract
Deep neural networks (DNNs) have become a cornerstone in advancing artificial intelligence, but their complexity often leads to inefficient hardware utilization due to varying structure characteristics and excessive memory accesses. Domain-specific architectures (DSAs) offer a solution by optimizing data locality through data stationary, tiling, and layer fusion, which minimize memory access and energy consumption while boosting performance. However, current approaches lack flexibility for efficient memory management at the appropriate granularity, causing misaligned accesses and decreasing reuse potential for variable-sized tiles. To this end, we propose a hierarchical Instruction Set Architecture (hierarchical-ISA) combining a RISC-V ISA and a flexible CISC-style macro-ISA (mISA). Unlike byte-level RISC-V, mISA employs row-wise tiles as the fundamental operand, enabling efficient data reuse across adjacent iterations as well as residual connections. This mISA approach simplifies DNN programming, enhances data partitioning and manipulation efficiency, and enables a hardware-software co-designed Remapping mechanism that facilitates data reuse without physical data movement. Experiments show 31.8%–72.0% reductions in off-chip memory access across MobileNet, ResNet, Swin Transformer, MobileViT, along with speedups of 2.9× to 7.4× compared to previous DNN accelerators. We also conduct comparisons under the roofline model with NVIDIA RTX A6000 and Intel Core i7-10700K. The results show that our arithmetic intensity reaches up to 26.0× that of i7-10700K and 22.6× that of A6000.
Zhiwang Huo, Wenzhe Zhao 0001, Qiwei Dang, Chengyu Ma, Guoming Yang, Gelin Fu, Tian Xia 0008, Pengju Ren
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2026 FP2: A 2-bit Floating-Point Format for Edge-AI Inference and Fine-Tuning
abstract
The increasing scale of Deep Neural Networks (DNNs) has made 2-bit quantization crucial for mitigating memory bottlenecks on edge devices. Low-bitwidth floating-point formats, offering larger dynamic ranges and avoiding quantization steps, have emerged as promising alternatives to fixed-point quantization. However, constructing viable floating-point representations with fewer than 3 bits remains challenging, as conventional formats require at least one sign bit, one exponent bit, and one mantissa bit. We address this challenge by introducing a novel data compression method that uses a 4-bit encoding space to represent two floating-point values, achieving an effective storage density of 2 bits per value. Depending on the bit width of the exponent and mantissa, we propose two different 2-bit floating-point encodings:fp2-e1m0andfp2-e0m1. Based onfp2, we introduce two computing architectures that simplify floating-point multiply-accumulate (MAC) operations into bitwise addition and logic operations, reducing floating-point computation by factors of$2\times $and$4\times $. As a result,fp2offers a practical solution for efficient inference using floating-point arithmetic on resource-constrained edge devices. Moreover, we analyze the error characteristics of thefp2data format from three perspectives. To validate the effectiveness of thefp2format, we conduct experiments on ResNet18/50 and ConvNeXt-Tiny using the CIFAR-10 and ImageNet-1K datasets. Compared tofp4, our approach reduces model size by 47%, with accuracy loss is less than 2 percentage points. Notably, on CIFAR-10, some results are close to those offp32. In contrast, when evaluated under 2-bit GPTQ,fp2demonstrates significant advantages over the baseline method on the LLAMA model. For hardware evaluation, we implement our design at the RTL level and evaluate it on both FPGA and ASIC platforms. Compared to computation architectures based onfp4, ourfp4$\times $fp2processing element (PE) array reduces area by 15% and power consumption by 8%. Furthermore, ourfp2$\times $fp2PE array achieves a remarkable 78% reduction in both area and power consumption.
Qiwei Dang, Chengyu Ma, Haiduo Huang, Gelin Fu, Zhiwang Huo, Guoming Yang, Pengchen Zong, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 DeepCell: Self-Supervised Multiview Fusion for Circuit Representation Learning
abstract
We introduce DeepCell, a novel circuit representation learning framework that effectively integrates multiview information from both And-Inverter Graphs (AIGs) and Post-Mapping (PM) netlists. At its core, DeepCell employs a self-supervised Mask Circuit Modeling (MCM) strategy, inspired by masked language modeling, to fuse complementary circuit representations from different design stages into unified and rich embeddings. To our knowledge, DeepCell is the first framework explicitly designed for PM netlist representation learning, setting new benchmarks in both predictive accuracy and reconstruction quality. We demonstrate the practical efficacy of DeepCell by applying it to critical EDA tasks such as functional Engineering Change Orders (ECO) and technology mapping. Extensive experimental results show that DeepCell significantly surpasses state-of-the-art open-source EDA tools in efficiency and performance. The code is available at https://github.com/cure-lab/DeepCell.
Zhengyuan Shi, Chengyu Ma, Lingfeng Zhou, Hongyang Pan, Fan Yang 0001, Zhufei Chu, Qiang Xu 0001
ICCAD2
2009 Certificate revocation release policies
abstract
Public key infrastructure provides a promising foundation for verifying the authenticity of communicating parties and transferring trust over the Internet. The key issue in public key infrastructure is how to process certificate revocations. Previous research in this area has concentrated on the tradeoffs that can be made among different revocation options. No rigorous efforts have been made to understand the probability distribution of certificate revocation requests based on real empirical data. In this study, we first collect real data from VeriSign and suggest a functional form for the probability density function of certificate revocation requests. Exponential distribution function is chosen as it adequately approximates the real data. We then provide an economic model based on which a certificate authority can choose the optimal Certificate Revocation List (CRL) release interval considering the intrinsic properties among different types of certificate services. To conclude we draw some insights by comparing the performance of four different CRL strategies.
Giri Kumar Tayi, Chengyu Ma, Yingjiu Li
J. Comput. Secur.3
2006 On the Release of CRLs in Public Key Infrastructure
Chengyu Ma, Yingjiu Li
USENIX Security Symposium1