VLDB 2026 Research / reviewers in the wild / expert
Yufei Ma 0002
dblp:32/333-2
· DBLP profile ↗
34ranked-venue papers
13as first author
20since 2021 · last 2026
0000-0002-2670-524XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 13 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RV-CIM: Energy-Delay Optimized Mapping and Architecture Co-Design for a RISC-V Multi-core SoC with Configurable DCIM Cluster
Ninghui Shang, Ying Liu 0069, Zecheng Zhou, Jiyong Hu, Zhiqiang Guo, Zhiyuan Chen 0009, Guoxiang Li, Yufei Ma 0002, Le Ye |
APPT | 10 |
| 2026 | A Full-Stack Framework for GNN Acceleration via Partition-Compiler-Architecture Co-DesignabstractGraph Neural Networks (GNNs) have achieved remarkable success across domains such as recommendation and scientific computing, yet their practical deployment remains constrained by high execution cost. The diversity of GNN model structures and the sparsity of real-world graphs pose two fundamental challenges for hardware acceleration: supporting heterogeneous operator patterns and achieving high resource utilization under irregular data access. Existing accelerators often address only one aspect, either targeting specific models with hardwired pipelines or applying general architectures with limited efficiency. To address these challenges, we propose SWITCHBLADE, a full-stack framework for GNN acceleration through the coordinated design of partitioning, compilation, and architecture. SWITCHBLADE addresses these challenges through three key components. First, a phase-based intermediate representation unifies diverse GNN models by abstracting computation stages for model-independent code generation. Second, a fine-grained graph partitioner enhances data locality and reduces memory traffic by adapting to graph topology and model semantics. Third, the hardware architecture supports stream-level parallelism and decoupled execution to exploit cross-shard and inter-phase concurrency. Evaluation on representative models and datasets shows that SWITCHBLADE achieves up to 1.85× speedup and 19.03× energy savings over an NVIDIA V100 GPU, while outperforming state-of-the-art GNN accelerators across diverse full-graph workloads, demonstrating both high efficiency and broad model generality. Yangjie Zhou 0001, Shuwen Lu, Cong Guo 0003, Jingwen Leng, Yufei Ma 0002, Yun Liang 0001, Minyi Guo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | Quartet: A Digital Compute-in-Memory Versatile AI Accelerator With Heterogeneous Tensor Engines and Off-Chip-Less DataflowabstractAlthough most AI core operations can be formulated as matrix multiplications (MMs), their characteristics are quite different. In some cases, one input may be a constant weight matrix while the other is a dynamic feature matrix, i.e., FWMM, or both inputs may be dynamic features as FFMM. Furthermore, the data involved can also be sparse, leading to variations like SpFWMM and SpFFMM. To address this challenge, this paper investigates a versatile accelerator architecture for AI algorithms based on heterogeneous tensor engines. For MM operators with varying characteristics, this paper proposes four-quadrant heterogeneous tensor engines to handle FWMM, SpFWMM, FFMM, and SpFFMM, respectively. These four tensor engines are comprised of SRAM-based single address digital compute-in-memory (CIM) array, SRAM-based multi-address digital CIM array, systolic array, and multi-SIMD array, respectively. In addition, to improve the AI execution efficiency, this paper proposes a dual-level multi-issue mechanism to achieve inter-operator and inter-block parallelization along with an off-chip-less dataflow enabled by on-chip unified memory pool. Thanks to the integration of aforementioned innovations, this paper develops a versatile AI acceleration chip Quartet, which achieves exceptional energy efficiency. Specifically, for graph convolutional network on PubMed, it demonstrates a$19.56\times $and$3.47\times $improvement in energy efficiency compared to similar works, ReDCIM and TensorCIM, respectively. Yikan Qiu, Guoxiang Li, Meng Wu 0005, Yifan Jia 0009, Le Ye, Yufei Ma 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | Accelerating Diffusion Transformer via Increment-Calibrated Caching with Channel-Aware Singular Value DecompositionabstractDiffusion transformer (DiT) models have achieved remarkable success in image generation, thanks for their exceptional generative capabilities and scalability. Nonetheless, the iterative nature of diffusion models (DMs) results in high computation complexity, posing challenges for deployment. Although existing cache-based acceleration methods try to utilize the inherent temporal similarity to skip redundant computations of DiT, the lack of correction may induce potential quality degradation. In this paper, we propose increment-calibrated caching, a training-free method for DiT acceleration, where the calibration parameters are generated from the pre-trained model itself with low-rank approximation. To deal with the possible correction failure arising from outlier activations, we introduce channel-aware Singular Value Decomposition (SVD), which further strengthens the calibration effect. Experimental results show that our method always achieve better performance than existing naive caching methods with a similar computation resource budget. When compared with 35-step DDIM, our method eliminates more than 45% computation and improves IS by 12 at the cost of less than 0.06 FID increase. Zhiyuan Chen 0009, Yifan Jia 0009, Le Ye, Yufei Ma 0002 |
CVPR | 5 |
| 2025 | 3D-SubG: A 3D Stacked Hybrid Processing Near/In-Memory Accelerator for Subgraph GNNsabstractSubgraph Graph Neural Networks (GNNs) are emerging as a promising approach to enhance GNN expressiveness, but their more complex graph structures with numerous independent and irregular subgraphs pose significant hardware deployment challenges. In this work, we propose 3D-SubG, a 3D stacked hybrid processing-near/in-memory accelerator for subgraph GNNs. With hybrid bonding packaging technology, a logic die is 3D stacked with a DRAM die for highly parallel memory accesses. The logic die employs digital SRAM-based processing-in-memory (PIM) macros to boost computation density and minimize data transfer. We further propose a bit-level non-zero gathering method to exploit graph sparsity for PIM, a workloadbalanced mapping strategy for subgraph allocation onto different logic-to-DRAM blocks, and a distributed global pooling approach to reduce inter-block data movements. Experimental results show that 3D-SubG achieves average improvements of $146.11 \times$ in performance, $934.18 \times$ in area efficiency, and $1171.80 \times$ in energy efficiency compared to RTX 3090Ti. Guoxiang Li, Runnan Xu, Ruohang Xu, Yikan Qiu, Renati Tuerhong, Muhan Zhang, Le Ye, Yufei Ma 0002 |
DAC | 8 |
| 2025 | 3D-TokSIM: Stacking 3D Memory with Token-Stationary Compute-in-Memory for Speculative LLM InferenceabstractThe LLM decoding process poses a significant challenge for memory bandwidth due to its autoregressive nature. Prior 2D memory solutions fail to overcome this memory bottleneck due to limited memory-to-logic bandwidth. In this work, we propose 3D-TokSIM, a cross-stack solution by stacking 3D memory on logic die with a specially designed token-stationary compute-in-memory (CIM) to efficiently accelerate speculative decoding. Our CIM is developed with novel token-stationary dataflow to reduce data movement on logic die to save power and balance computation and memory access. To further reduce the buffer requirements, we perform architecture exploration and allocate notable CIM resources for achieving higher decoding parallelism. Compared to RTX 3090 GPU, 3D-TokSIM achieves 15.1 $\times$ throughput and $324 \times$ energy efficiency improvements on speculative Llama2-7B decoding. Boya Lv, Meng Wu 0005, Fengyun Yan, Yufei Ma 0002, Ru Huang 0001, Le Ye |
DAC | 6 |
| 2025 | 3D-MoE: Accelerating Multi-Expert Activated LLMs on 3D In/Near-Memory Computing Architecture via Hybrid ParallelismabstractThe Mixture-of-Expert (MoE) architecture constitutes a significant advancement in the scalability of large- language models (LLMs), achieving better performance compared to traditional dense models with the same size of activated parameters. However, the MoE models introduce two main challenges for their edge deployment: (1) increased memory storage requirements due to the sparse parameter composition and (2) aggravated yon Neumann bottleneck caused by the inherent unpredictability of dynamic expert routing. Recent trends in state-of-the-art MoE models suggest that scaling up both activated and the total number of experts leads to better model performance. However, this multi-expert activated LLM architecture further exacerbates the aforementioned issues. To address these challenges, this paper presents a 3D stacked near-DRAM computing architecture based on available hybrid bonding and mini-TSV techniques, thereby increasing both storage capacity with the same die area and memory bandwidth for MoE models. We also employ 6T-SRAM-based Digital Computing-in-Memory (DCIM) macros as the primary computing components and weight buffers to optimize the effective throughput of LLM inference. Furthermore, we propose a hybrid parallel computation scheme with Greedy-Swap Expert Grouping (GSEG) strategy. It balances the weight-parameter placement between tile groups in 3D-MoE architecture to enhance hardware utilization and overall performance. Through comprehensive benchmarking, the proposed accelerator demonstrates 143.10× and 68.89× speedup in token throughput, and 870.98× and 47.11× improvements in energy efficiency compared to CPU and GPU, respectively, confirming its advantages for multi-expert activated LLM inference. Xinyu Qu, Runnan Xu, Yufei Ma 0002 |
ICCAD | 4 |
| 2024 | An In-Memory Computing Accelerator with Reconfigurable Dataflow for Multi-Scale Vision Transformer with Hybrid TopologyabstractTransformer models equipped with multi-head attention (MHA) mechanism have demonstrated promise in computer vision (CV) tasks, i.e., vision transformers (ViTs). Nevertheless, the lack of inductive bias in ViTs leads to substantial computational and storage requirements, hindering their deployment on resource-constrained edge devices. To this end, multi-scale hybrid models are proposed to take the advantages of both transformers and convolutional neural networks (CNNs). However, existing domain-specific architectures focus on the optimization of either convolution or MHA at the expense of flexibility. In this work, an in-memory computing (IMC) accelerator is proposed to efficiently accelerate ViTs with hybrid MHA and convolution topology by introducing pipeline reordering. SRAM-based digital IMC macro is utilized to mitigate memory access bottleneck, while avoiding analog non-ideality. The reconfigurable processing engines and interconnections are investigated to enable the adaptable mapping of both convolution and MHA. Under typical workloads, experimental results exhibit that our proposed IMC architecture delivers 2.20× to 2.52× speedup and 40.6% to 74.8% energy reduction compared with the baseline design. Zhiyuan Chen 0009, Yufei Ma 0002, Yifan Jia 0009, Guoxiang Li, Meng Wu 0005, Le Ye, Ru Huang 0001 |
DAC | 2 |
| 2024 | AIG-CIM: A Scalable Chiplet Module with Tri-Gear Heterogeneous Compute-in-Memory for Diffusion AccelerationabstractThe emergence of Diffusion models has gained significant attention in the field of Artificial Intelligence Generated Content. While Diffusion demonstrates impressive image generation capability, it faces hardware deployment challenges due to its unique model architecture and computation requirement. In this paper, we present a hardware accelerator design, i.e. AIG-CIM, which incorporates tri-gear heterogeneous digital compute-in-memory to address the flexible data reuse demands in Diffusion models. Our framework offers a collaborative design methodology for large generative models from the computational circuit-level to the multi-chip-module system-level. We implemented and evaluated the AIG-CIM accelerator using TSMC 22nm technology. For several Diffusion inferences, scalable AIG-CIM chiplets achieve 21.3× latency reduction, up to 231.2× throughput improvement and three orders of magnitude energy efficiency improvement compared to RTX 3090 GPU. Yiqi Jing, Meng Wu 0005, Yufei Ma 0002, Ru Huang 0001, Le Ye |
DAC | 5 |
| 2024 | Sparsity-Aware In-Memory Neuromorphic Computing Unit With Configurable Topology of Hybrid Spiking and Artificial Neural NetworkabstractSpiking neural networks (SNNs) have shown great potential in achieving high energy efficiency and low power consumption compared to artificial neural networks (ANNs). However, there remains a significant accuracy gap between SNNs and ANNs. To address this issue, we present an in-memory neuromorphic computing (IMNC) chip that supports hybrid spiking/artificial neural networks (S/ANNs) and sparsity-aware data flows. With the IMNC chip, we aim to improve inference accuracy while simultaneously achieving high energy efficiency through optimization at the algorithm, architecture, and circuit levels. First, at the algorithm level, we note that SNNs extract temporal features from input spikes using time-domain convolution operations. Based on this insight, we efficiently utilize leaky integrate (LI) neurons to hybridize SNNs and ANNs, thereby improving accuracy while maintaining highly sparse operations. Second, at the architecture level, we design a sparsity-aware architecture that supports a hybrid S/ANN topology with varying sparsity. Finally, at the circuit level, we propose a ring-based in-memory computing (IMC) macro, whose energy consumption is inversely proportional to the input sparsity, making it ideal for performing energy-efficient multiplication and accumulation (MAC) operations in both SNNs and ANNs. We evaluate the proposed hybrid S/ANNs on various classification tasks and demonstrate their stronger classification and generalization ability compared with pure SNNs. Notably, our IMNC chip, fabricated using 22 nm CMOS technology, achieves impressive measured accuracy rates of over 95% for voice activity detection (VAD) and ECG anomaly detection. Additionally, our IMNC chip demonstrates superior dynamic energy efficiency of 0.43 pJ per synaptic operation, outperforming related works. Ying Liu 0069, Zhiyuan Chen 0009, Zhixuan Wang, Ru Huang 0001, Le Ye, Yufei Ma 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2024 | DCIM-GCN: Digital Computing-in-Memory Accelerator for Graph Convolutional NetworkabstractGraph convolutional network (GCN) has gained great success in a diverse range of intelligent tasks. However, the hardware performance of GCNs is often bounded by random and non-continuous memory accesses due to the sparse graph data, which incur high latency and high power consumption. The emerging computing-in-memory (CIM) architecture significantly reduces the overhead of data movements, which is suitable for memory-intensive GCN acceleration. Existing analog-based CIM solutions require a large amount of analog-to-digital (AD) and digital-to-analog (DA) conversions, which dominate the overall area and power consumption. Furthermore, the analog non-ideality can degrade accuracy and reliability of CIM. To address these challenges, this work proposes a digital CIM accelerator based on SRAM, called DCIM-GCN, to accelerate GCN algorithm. DCIM-GCN introduces innovations on three levels: circuit, architecture, and algorithm. At the circuit level, digital CIM is proposed with SRAM sub-arrays to eliminate the power and area expensive AD/DA converters. Furthermore, we have incorporated the multi-address feature into the digital CIM, thereby leveraging its ability to efficiently process sparse matrix multiplication. At the architecture level, the sparsity-aware computation engine takes advantage of sparsity in GCNs and leverages CIM to minimize memory accesses and data movements. Finally, at the algorithm level, the balance mapping algorithm tackles workload imbalance issues, while the vertex reorder algorithm reduces idle states for aggregation engines, resulting in increased hardware utilization. Our DCIM-GCN achieves 1.89$\times$and 2.42$\times$speedup and 4.58$\times$and 9.46$\times$energy efficiency improvement on average over other CIM-based graph accelerators, e.g., PASGCN and PIM-GCN, respectively. Yufei Ma 0002, Yikan Qiu, Guoxiang Li, Meng Wu 0005, Le Ye, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | RIMAC: An Array-Level ADC/DAC-Free ReRAM-Based in-Memory DNN Processor with Analog Cache and ComputationabstractBy directly computing in analog domain, processing-in-memory (PIM) is emerging as a promising alternative to overcome the memory bottleneck of traditional von-Neuman architecture, especially for deep neural networks (DNNs). However, the data outside PIM macros in most existing PIM accelerators are stored and operated as digital signals that require massive expensive digital-to-analog (D/A) and analog-to-digital (A/D) converters. In this work, an array-level ADC/DAC-free ReRAM-based in-memory DNN processor named RIMAC is proposed, which accelerates various DNNs in pure analog-domain with analog cache and analog computation modules to eliminate the expensive D/A and A/D conversions. Our experiment result shows the peak energy efficiency is improved by about 34.8×, 97.6×, 10.7×, and 14.0× compared to PRIME, ISAAC, Lattice, and 21'DAC for various DNNs on ImageNet, respectively. Meng Wu 0005, Yufei Ma 0002, Le Ye, Ru Huang 0001 |
ASP-DAC | 3 |
| 2023 | A Model-Specific End-to-End Design Methodology for Resource-Constrained TinyML HardwareabstractTiny machine learning (TinyML) becomes appealing as it enables machine learning on resource-constrained devices with ultra low energy and small form factor. In this paper, a model-specific end-to-end design methodology is presented for TinyML hardware design. First, we introduce an end-to-end system evaluation method using Roofline models, which considering both AI and other general-purpose computing to guide the architecture design choices. Second, to improve the efficiency of AI computation, we develop an enhanced design space exploration framework, TinyScale, to enable optimal low-voltage operation for energy-efficient TinyML. Finally, we present a use case driven design selection method to search the optimal hardware design across a set of application use cases. Our model-specific design methodology is evaluated on both TSMC 22nm and 55nm technology for MLPerf Tiny benchmark and a keyword spotting (KWS) SoC design. With the help of our end-to-end design methodology, an optimal TinyML hardware can be automatically explored with significant energy and EDP improvements for a diverse of TinyML use cases. Yanchi Dong, Kaixuan Du, Yiqi Jing, Qijun Wang, Pixian Zhan, Fengyun Yan, Yufei Ma 0002, Yun Liang 0001, Le Ye, Ru Huang 0001 |
DAC | 9 |
| 2023 | DCIM-3DRec: A 3D Reconstruction Accelerator with Digital Computing-in-Memory and Octree-Based SchedulerabstractLearning-based 3D reconstruction has evolved rapidly with promising quality, while it requires high-performance hardware for interactive applications. In this work, a reconstruction accelerator called DCIM-3DRec is presented which leverages digital computing-in-memory (DCIM) design to facilitate learning-based reconstruction deployment on realtime and low-power edge platforms. The DCIM-3DRec is designed with the following features: a reconfigurable DCIM macro array for high data reuse and macro utilization, and an Octree-based subdivision scheduler for efficient management of 3D space prediction. The DCIM-3DRec accelerator is implemented and evaluated in TSMC 55 nm technology, with a DCIM macro efficiency of 19.4 TOPS/W at INT8. Overall, the DCIM-3DRec accelerator achieves 23× performance gain and four orders of magnitude energy efficiency improvement compared to a Nvidia RTX3090 GPU. Yiqi Jing, Meng Wu 0005, Fengyun Yan, Yufei Ma 0002, Le Ye |
ISLPED | 7 |
| 2023 | Research progress on low-power artificial intelligence of things (AIoT) chip design
Le Ye, Zhixuan Wang, Yufei Ma 0002, Linxiao Shen, Yihan Zhang 0002, Meng Wu 0005, Ying Liu 0069, Yiqi Jing, Hao Zhang 0119, Ru Huang 0001 |
Sci. China Inf. Sci. | 4 |
| 2023 | An 82-nW 0.53-pJ/SOP Clock-Free Spiking Neural Network With 40-μs Latency for AIoT Wake-Up Functions Using a Multilevel-Event-Driven Bionic Architecture and Computing-in-Memory TechniqueabstractThis article presents a clock-free spiking neural network (SNN) intelligent inference engine (IIE) for artificial intelligence of things (AIoT) sensor nodes, which often operate in random-sparse-event (RSE) scenarios. The IIE drastically reduces the system’s long-term average (LTA) power consumption, improves energy efficiency, and achieves microsecond level inference latency. Three techniques are proposed: 1) A clock-free SNN architecture without clock tree, frame generator, and arbiter, is driven by the output spikes, which are encoded with level-crossing (LC) sampling method; the circuit activity is completely related to event activity and spike rates, dramatically reducing the overall power consumption and latency. 2) The bioinspired leaky-integrate-fire (LIF) neurons directly extract the time-domain information from asynchronous spikes, reducing the network size and number of operations. 3) The computing-in-memory (CIM) and mixed-signal synapse-neuron circuits are employed to increase the SNN parallelism and avoid weight movements, thus improving the energy efficiency and response speed. The measured LTA power is bounded at 82 nW while the event-driven chip is on call and waiting for events; the energy efficiency is 0.53 pJ per synapse operation (SOP), only 1/3 that of state-of-the-art methods at 4bit weights even with 180 nm technology. We demonstrate electrocardiogram (ECG) recognition as a typical AIoT application, and the power consumption is less than 350 nW. The measured accuracy of abnormal ECG detection is 90.5%. Moreover, the latency is only$40 \mu \text{s}$to realize real-time NN inference. This work provides an effective solution for AIoT nodes that require both ultralow power and fast response. Ying Liu 0069, Yufei Ma 0002, Zhixuan Wang, Linxiao Shen, Jiayoon Ru, Ru Huang 0001, Le Ye |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | DCIM-GCN: Digital Computing-in-Memory to Efficiently Accelerate Graph Convolutional NetworksabstractComputing-in-memory (CIM) is emerging as a promising architecture to accelerate graph convolutional networks (GCNs) normally bounded by redundant and irregular memory transactions. Current analog based CIM requires frequent analog and digital conversions (AD/DA) that dominate the overall area and power consumption. Furthermore, the analog non-ideality degrades the accuracy and reliability of CIM. In this work, an SRAM based digital CIM system is proposed to accelerate memory intensive GCNs, namely DCIM-GCN, which covers innovations from CIM circuit level eliminating costly AD/DA converters to architecture level addressing irregularity and sparsity of graph data. DCIM-GCN achieves 2.07X, 1.76X, and 1.89× speedup and 29.98×, 1.29×, and 3.73× energy efficiency improvement on average over CIM based PIMGCN, TARe, and PIM-GCN, respectively. Yikan Qiu, Yufei Ma 0002, Meng Wu 0005, Le Ye, Ru Huang 0001 |
ICCAD | 2 |
| 2022 | Hybrid Stochastic-Binary Computing for Low-Latency and High-Precision Inference of CNNsabstractThe appealing property of low area, low power, and high bit error tolerance has made Stochastic Computing (SC) a promising alternative to conventional binary arithmetic for many computation intensive tasks, e.g., convolutional neural networks (CNNs). However, current SC-based CNN accelerators suffer from the intrinsic computation error and exponentially growing latency. In this work, we optimize both the architecture of SC multiply-and-accumulate (MAC) unit and the overall acceleration strategy of CNN accelerator to favor SC. A low-complexity bit-stream-extending method is proposed to suppress the computation error of SC and ensure the trained fix-point model can be deployed into SC-based hardware without fine-tuning. Besides, distribution-determined partition scheme is developed to design hybrid stochastic-binary computing (SBC) MAC unit which boosts the processing of bit streams at a minimum overhead. For the overall accelerator, the SBC-based MAC array is extended to reuse hardware resources and improve throughput, since the judiciously chosen loop unrolling strategy can better benefit SC operations. The proposed CNN accelerator with extended SBC-MAC array is synthesized and validated using TSMC 28nm CMOS on several representative CNNs, targeted at ImageNet dataset. Compared with precise binary implementation, our proposed design gains 44% area reduction and 50% power saving but induces only 4% additional computation latency and 0.5% accuracy degradation. Zhiyuan Chen 0009, Yufei Ma 0002, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | A Flexible and Efficient FPGA Accelerator for Various Large-Scale and Lightweight CNNsabstractTo enable efficient deployment of convolutional neural networks (CNNs) on embedded platforms for different computer vision applications, several convolution variants have been introduced, such as depthwise convolution (DWCV), transposed convolution (TPCV), and dilated convolution (DLCV). To address the utilization degradation issue occurred in a general convolution engine for these emerging operators, a highly flexible and reconfigurable hardware accelerator is proposed to efficiently support various CNN-based vision tasks. Firstly, to avoid workload imbalance of TPCV, a zero transfer and skipping (ZTS) method is proposed to reorganize the computation process. To eliminate the redundant zero calculations of TPCV and DLCV, a sparsity-alike processing (SAP) method is proposed based on weight-oriented dataflow. Secondly, the DWCV or pooling layers are configured to be directly executed after standard convolutions without external memory accesses. Furthermore, a programmable execution schedule is introduced to gain better flexibility. Finally, the proposed accelerator is evaluated on Intel Arria 10 SoC FPGA. Experimental results show state-of-the-art performance on both large-scale and lightweight CNNs for image segmentation or classification. Specifically, the accelerator can achieve a processing speed up to 339.9 FPS and computational efficiency up to 0.58 GOPS/DSP, which is$3.3\times $better than the prior art evaluated on the same network. Yufei Ma 0002, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | SWIFT: Small-World-based Structural Pruning to Accelerate DNN Inference on FPGAabstractState-of-the-art DNN pruning approaches achieved high sparsity. However, these methods usually do not consider the intrinsic graph property of DNNs, leading to an irregular pruned network. Consequently, hardware accelerators cannot directly benefit from such pruning, suffering additional cost on indexing, control and data paths. Inspired by the observation that the brain and real-world networks follow a Small-World model, we propose a graph-based progressive structural pruning technique, SWIFT, that integrates local clusters and global sparsity in DNNs to benefit the dataflow and workload balance of the accelerators. In particular, we propose an output stationary FPGA architecture to accelerate DNN inference and integrate it with the structural sparsity by SWIFT, so that the communication and computation of clustered zero weights are eliminated. In addition, a full mesh data router is designed to adaptively direct inputs into corresponding processing elements (PEs) for different layer configurations and skipping zero operations. The proposed SWIFT is evaluated with multiple DNNs on different datasets. It achieves sparsity ratio up to 76% for CIFAR-10, 83% for CIFAR-100, 76% for the SVHN datasets. Moreover, our proposed SWIFT FPGA accelerator achieves up to 4.4× improvement in throughput for different dense networks with a marginal hardware overhead. Yufei Ma 0002, Yu Cao 0001, Le Ye, Ru Huang 0001 |
FPGA | 1 |
| 2020 | In-Memory Computing: The Next-Generation AI Computing ParadigmabstractTo overcome the memory bottleneck of von-Neuman architecture, various memory-centric computing techniques are emerging to reduce the latency and energy consumption caused by data communication. The great success of artificial intelligence (AI) algorithms, which involve a large number of computations and data movements, has motivated and accelerated the recent researches of in-memory computing (IMC) techniques to significantly reduce or even diminish the accesses of off-chip data, where memory is not only storing data but can also directly output computation results. For example, the multiply-and-accumulate (MAC) operations in deep learning algorithms can be realized by accessing the memory using the input activations. This paper will investigate the recent trends of IMC from techniques (SRAM, flash, RRAM and other types of non-volatile memory) to architecture and to applications, which will serve as a guide to the future advances on computing in-memory (CIM). Yufei Ma 0002, Yuan Du, Jun Lin 0001, Zhongfeng Wang 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | Optimizing Stochastic Computing for Low Latency Inference of Convolutional Neural NetworksabstractThe appealing property of low area, low power, flexible precision, and high bit error tolerance has made Stochastic Computing (SC) a promising alternative to conventional binary arithmetic for many computation intensive tasks, e.g., convolutional neural networks (CNNs). However, to relieve the intrinsic fluctuation noise in SC, long bit stream is normally required in SC-based CNN accelerators to achieve satisfactory accuracy, which leads to extortionate latency. Although the bit parallel structure of a SC multiplier has been proposed to reduce latency, the resulting extra overhead still considerably degrade the overall efficiency of SC. In this paper, we optimize both the micro-architecture of SC multiply-and-accumulate (MAC) unit and the overall acceleration scheme of CNN accelerator to favor SC. An optimized and scalable SC-MAC unit, which fully utilizes the property of low-discrepancy bit stream, is proposed with adjustable parameters to reduce the latency with minor area increase. For the overall accelerator, the parallel dimensions of SC-based MAC array are extended to reuse hardware resources and improve throughput, since the judiciously chosen loop unrolling strategy can better benefit SC operations. The proposed CNN accelerator with extended SC-MAC array is synthesized and demonstrated using TSMC 28nm CMOS on several representative CNNs, which gains 2× performance speedup, 2.8× energy savings and 15% area reduction compared to state-of-the-art SC based CNN accelerator. Zhiyuan Chen 0009, Yufei Ma 0002, Zhongfeng Wang 0001 |
ICCAD | 2 |
| 2020 | Automatic Compilation of Diverse CNNs Onto High-Performance FPGA AcceleratorsabstractA broad range of applications are increasingly benefiting from the rapid and flourishing development of convolutional neural networks (CNNs). The FPGA-based CNN inference accelerator is gaining popularity due to its high-performance and low-power as well as FPGA's conventional advantage of reconfigurability and flexibility. Without a general compiler to automate the implementation, however, significant efforts and expertise are still required to customize the design for each CNN model. In this paper, we present an register-transfer level (RTL)-level CNN compiler that automatically generates customized FPGA hardware for the inference tasks of various CNNs, in order to enable high-level fast prototyping of CNNs from software to FPGA and still keep the benefits of low-level hardware optimization. First, a general-purpose library of RTL modules is developed to model different operations at each layer. The integration and dataflow of physical modules are predefined in the top-level system template and reconfigured during compilation for a given CNN algorithm. The runtime control of layer-by-layer sequential computation is managed by the proposed execution schedule so that even highly irregular and complex network topology, e.g., GoogLeNet and ResNet, can be compiled. The proposed methodology is demonstrated with various CNN algorithms, e.g., NiN, VGG, GoogLeNet, and ResNet, on two standalone Intel FPGAs, Arria 10, and Stratix 10, achieving end-to-end inference throughputs of 969 GOPS and 1604 GOPS, respectively, with batch size of one. Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Performance Modeling for CNN Inference Accelerators on FPGAabstractThe recently reported successes of convolutional neural networks (CNNs) in many areas have generated wide interest in the development of field-programmable gate array (FPGA)-based accelerators. To achieve high performance and energy efficiency, an FPGA-based accelerator must fully utilize the limited computation resources and minimize the data communication and memory access, both of which are impacted and constrained by a variety of design parameters, e.g., the degree and dimension of parallelism, the size of on-chip buffers, the bandwidth of the external memory, and many more. The large design space of the accelerator makes it impractical to search for the optimal design in the implementation phase. To address this problem, a performance model is described to estimate the performance and resource utilization of an FPGA implementation. By this means, the performance bottleneck and design bound can be identified and the optimal design option can be explored early in the design phase. The proposed performance model is validated using a variety of CNN algorithms comparing the results with on-board test results on two different FPGAs. Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Automatic Compiler Based FPGA Accelerator for CNN TrainingabstractTraining of convolutional neural networks (CNNs) on embedded platforms to support on-device learning is earning vital importance in recent days. Designing flexible training hardware is much more challenging than inference hardware, due to design complexity and large computation/memory requirement. In this work, we present an automatic compiler based FPGA accelerator with 16-bit fixed-point precision for complete CNN training, including Forward Pass (FP), Backward Pass (BP) and Weight Update (WU). We implemented an optimized RTL library to perform training-specific tasks and developed an RTL compiler to automatically generate FPGA-synthesizable RTL based on user-defined constraints. We present a new cyclic weight storage/access scheme for on-chip BRAM and off-chip DRAM to efficiently implement non-transpose and transpose operations during FP and BP phases, respectively. Representative CNNs for CIFAR-10 dataset are implemented and trained on Intel Stratix 10 GX FPGA using proposed hardware architecture, demonstrating up to 479 GOPS performance. Shreyas K. Venkataramanaiah, Yufei Ma 0002, Shihui Yin, Eriko Nurvitadhi, Aravind Dasu, Yu Cao 0001, Jae-sun Seo |
FPL | 2 |
| 2018 | Algorithm-hardware co-design of single shot detector for fast object detection on FPGAsabstractThe rapid improvement in computation capability has made convolutional neural networks (CNNs) a great success in recent years on image classification tasks, which has also prospered the development of objection detection algorithms with significantly improved accuracy. However, during the deployment phase, many applications demand low latency processing of one image with strict power consumption requirement, which reduces the efficiency of GPU and other general-purpose platform, bringing opportunities for specific acceleration hardware, e.g. FPGA, by customizing the digital circuit specific for the inference algorithm. Therefore, this work proposes to customize the detection algorithm, e.g. SSD, to benefit its hardware implementation with low data precision at the cost of marginal accuracy degradation. The proposed FPGA-based deep learning inference accelerator is demonstrated on two Intel FPGAs for SSD algorithm achieving up to 2.18 TOPS throughput and up to 3.3× superior energy-efficiency compared to GPU. Yufei Ma 0002, Tu Zheng, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo |
ICCAD | 1 |
| 2018 | ALAMO: FPGA acceleration of deep learning algorithms with a modularized RTL compiler
Yufei Ma 0002, Naveen Suda, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo |
Integr. | 1 |
| 2018 | Optimizing the Convolution Operation to Accelerate Deep Neural Networks on FPGA
Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Optimizing Loop Operation and Dataflow in FPGA Acceleration of Deep Convolutional Neural Networks
Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo |
FPGA | 1 |
| 2017 | An automatic RTL compiler for high-throughput FPGA implementation of diverse deep convolutional neural networksabstractConvolutional neural networks (CNNs) are rapidly evolving and being applied to a broad range of applications. Given a specific application, an increasing challenge is to search the appropriate CNN algorithm and efficiently map it to the target hardware. The FPGA-based accelerator has the advantage of reconfigurability and flexibility, and has achieved high-performance and low-power. Without a general compiler to automate the implementation, however, significant efforts and expertise are still required to customize the design for each CNN model. In this work, we present an RTL-level CNN compiler that automatically generates customized FPGA hardware for the inference tasks of various CNNs, in order to enable high-level fast prototyping of CNNs from software to FPGA and still keep the benefits of low-level hardware optimization. First, a general-purpose library of RTL modules is developed to model different operations at each layer. The implementation of each module is optimized at the RTL level. Given a CNN algorithm, its structure is abstracted to a directed acyclic graph (DAG) and then complied with RTL modules in the library. The integration and dataflow of physical modules are predefined in the top-level system template and reconfigured during compilation. The runtime control of layer-by-layer sequential computation is managed by the proposed execution schedule so that even highly irregular and complex network topology, e.g. ResNet, can be compiled. The proposed methodology is demonstrated with end-to-end FPGA implementations of various CNN algorithms (e.g. NiN, VGG-16, ResNet-50, and ResNet-152) on two standalone Intel FPGAs, Stratix V and Arria 10. The performance and overhead of the automated compilation are evaluated. The compiled FPGA accelerators exhibit superior performance compared to state-of-the-art automation-based works by >2× for various CNNs. Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo |
FPL | 1 |
| 2017 | End-to-end scalable FPGA accelerator for deep residual networksabstractThis work presents an efficient hardware accelerator design of deep residual learning algorithms, which have shown superior image recognition accuracy (>90% top-5 accuracy on ImageNet database). Two key objectives of the acceleration strategy are to (1) maximize resource utilization and minimize data movements, and (2) employ scalable and reusable computing primitives to optimize physical design under hardware constraints. Furthermore, we present techniques for efficient integration and communication of these primitives in deep residual convolutional neural networks (CNNs) that exhibit complex, non-uniform layer connections. The proposed hardware accelerator efficiently implements state-of-the-art ResNet-50/152 algorithms on Arria-10 FPGA, demonstrating 285.1/315.5 GOPS of throughput and 27.2/71.7 ms of latency, respectively. Yufei Ma 0002, Minkyu Kim 0001, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo |
ISCAS | 1 |
| 2016 | Throughput-Optimized OpenCL-based FPGA Accelerator for Large-Scale Convolutional Neural NetworksabstractConvolutional Neural Networks (CNNs) have gained popularity in many computer vision applications such as image classification, face detection, and video analysis, because of their ability to train and classify with high accuracy. Due to multiple convolution and fully-connected layers that are compute-/memory-intensive, it is difficult to perform real-time classification with low power consumption on today?s computing systems. FPGAs have been widely explored as hardware accelerators for CNNs because of their reconfigurability and energy efficiency, as well as fast turn-around-time, especially with high-level synthesis methodologies. Previous FPGA-based CNN accelerators, however, typically implemented generic accelerators agnostic to the CNN configuration, where the reconfigurable capabilities of FPGAs are not fully leveraged to maximize the overall system throughput. In this work, we present a systematic design space exploration methodology to maximize the throughput of an OpenCL-based FPGA accelerator for a given CNN model, considering the FPGA resource constraints such as on-chip memory, registers, computational resources and external memory bandwidth. The proposed methodology is demonstrated by optimizing two representative large-scale CNNs, AlexNet and VGG, on two Altera Stratix-V FPGA platforms, DE5-Net and P395-D8 boards, which have different hardware resources. We achieve a peak performance of 136.5 GOPS for convolution operation, and 117.8 GOPS for the entire VGG network that performs ImageNet classification on P395-D8 board. Naveen Suda, Vikas Chandra, Ganesh Dasika, Abinash Mohanty, Yufei Ma 0002, Sarma B. K. Vrudhula, Jae-sun Seo, Yu Cao 0001 |
FPGA | 5 |
| 2016 | Scalable and modularized RTL compilation of Convolutional Neural Networks onto FPGAabstractDespite its popularity, deploying Convolutional Neural Networks (CNNs) on a portable system is still challenging due to large data volume, intensive computation and frequent memory access. Although previous FPGA acceleration schemes generated by high-level synthesis tools (i.e., HLS, OpenCL) have allowed for fast design optimization, hardware inefficiency still exists when allocating FPGA resources to maximize parallelism and throughput. A direct hardware-level design (i.e., RTL) can improve the efficiency and achieve greater acceleration. However, this requires an in-depth understanding of both the algorithm structure and the FPGA system architecture. In this work, we present a scalable solution that integrates the flexibility of high-level synthesis and the finer level optimization of an RTL implementation. The cornerstone is a compiler that analyzes the CNN structure and parameters, and automatically generates a set of modular and scalable computing primitives that can accelerate various deep learning algorithms. Integrating these modules together for end-to-end CNN implementations, this work quantitatively analyzes the complier's design strategy to optimize the throughput of a given CNN model with the FPGA resource constraints. The proposed methodology is demonstrated on Altera Stratix-V GXA7 FPGA for AlexNet and NIN CNN models, achieving 114.5 GOPS and 117.3 GOPS, respectively. This represents a 1.9× improvement in throughput when compared to the OpenCL-based design. The results illustrate the promise of the automatic compiler solution for modularized and scalable hardware acceleration of deep learning. Yufei Ma 0002, Naveen Suda, Yu Cao 0001, Jae-sun Seo, Sarma B. K. Vrudhula |
FPL | 1 |
| 2015 | Energy-efficient reconstruction of compressively sensed bioelectrical signals with stochastic computing circuitsabstractCompressive sensing (CS) allows acquiring sparse signals at sub-Nyquist rate, offering an energy-efficient solution to data acquisition. This is especially important to reduce communication data for mobile medical applications. However, reconstructing the signal from CS is usually left off-line due to the complex computations. In this paper, we integrate two key technologies to enable on-line energy-efficient CS signal reconstruction. These are (1) the use of Bayesian CS Belief Propagation (CS-BP) as the algorithm basis and (2) the novel design of stochastic computing (SC) circuits to efficiently map CS-BP algorithm. The overall signal reconstruction system is implemented with digital SC circuits in 65nm CMOS and recovers compressively sensed electrocardiography (ECG) and electromyography (EMG) signals with 11X to 8X data compression factor. Compared to a conventional binary design, post-layout simulation results show that the proposed stochastic design performs reconstruction with 5X energy-delay product improvement and 2X area reduction. Yufei Ma 0002, Minkyu Kim 0001, Yu Cao 0001, Jae-sun Seo, Sarma B. K. Vrudhula |
ICCD | 1 |