EDBT 2026 Demo / reviewers in the wild / expert
Fengbin Tu
dblp:161/0248
· DBLP profile ↗
49ranked-venue papers
6as first author
36since 2021 · last 2026
0000-0003-2228-8829ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 5 first-author · 34 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HR-DCIM: High-Reliability Floating-Point Digital CIM Architecture With Unified Low-Cost Iterative Error CorrectionabstractDigital computing-in-memory (CIM) is a promising computing paradigm for the neural network (NN) acceleration. However, during the actual deployment process of digital CIM chips, we find that existing digital CIM designs face severe computing reliability issues, which are crucial for real product development but remain underexplored. Therefore, this work pioneers a systematic computing reliability analysis for digital CIM across off-memory and in-memory levels. We find that both the off-memory floating-point (FP) exponent alignment and the in-memory random cell bit-flip errors impair digital CIM's computing reliability, causing significant truncation and bit-flip accuracy loss. Critically, existing reliability solutions are incompatible with the unique multi-row accumulation structure of digital CIM, which either severely damage digital CIM's performance or result in prohibitive overhead. To address above challenges, we propose HR-DCIM: a highreliability FP digital CIM architecture featuring unified lowcost iterative error correction. Specifically, for the off-memory reliability, we propose an exponent-mantissa joint-alignment mechanism to repurpose inherent invalid bits of aligned mantissas as compensation bits to reduce alignment truncation loss, without damaging digital CIM's performance. Then, for the in-memory reliability, we propose a remainder aliasing-based unified multiply-accumulation (MAC) error correction mechanism to correct possible MAC errors caused by various cell error cases with low-cost iteration. Experimental results show that the proposed techniques enable digital CIM to maintain high performance and efficiency across various operating voltage conditions without significant accuracy loss. Yiqi Wang 0005, Zhiheng Yue, Zihan Wu 0006, Huiming Han, Shaojun Wei, Yang Hu 0001, Fengbin Tu, Shouyi Yin |
HPCA | 8 |
| 2026 | VAR-Turbo: Unlocking the Potential of Visual Autoregressive Models Through Dual RedundancyabstractImage synthesis task has recently drawn enormous attention from both the academia and the industry due to the recent advancements of generative models, which now can generate photorealistic images with conditional words from human. Among the generative models for image synthesis task, Visual AutoRegressive (VAR) model is a promising avenue due to its strong scalability. Nevertheless, its exorbitant computational cost poses a formidable obstacle to widespread adoption. To this end, we propose a dedicated software/hardware co-design framework dubbed VAR-Turbo for unlocking the potential of VAR models. Specifically, in the software level, we propose a Draft-Free Parallel Decoding scheme by exploiting the Image Redundancy, which can decrease the sample steps by$>80 \%$, and a combination of Token Aggregation and Dynamic Bypass that capitalizes on the Model Redundancy introduced by the generative Transformer to reduce the computational load by$>60 \%$. In the hardware level, we propose a dedicated accelerator featuring 1) A Unified Attention Core and 2) Radix Sort Core, which can support the aforementioned algorithm pipeline seamlessly and efficiently. Under the collaborative design and synergy of software and hardware, VAR-Turbo achieves averagely 5047.4x, 210.3x, 6.1x, 3.8x speedups and 24818.2x, 423.5x, 6.0x, 7.8x energy-efficiency improvements over Xeon 8168 CPU, Nvidia V100, ViTCoD and AdapTiV, while maintaining the generation quality. Xujiang Xiang, Fengbin Tu |
HPCA | 2 |
| 2026 | Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit ComputationabstractLarge Language Models (LLMs) have transformed society, but their computational and energy needs hinder efficient inference. The memory wall, the growing processor-memory speed disparity, remains a critical bottleneck for LLM. While Process-in-Memory (PIM) architectures address this challenge by co-locating computation with memory, achieving 5-20 × higher bandwidth than GPUs, existing scalable PIM solutions face critical trade-offs in flexibility, capacity, and efficiency when handling LLMs' dynamic memory-compute patterns and operator diversity. DRAM-PIM suffers from inter-bank communication overhead despite its vector parallelism. SRAM-PIM offers sub10ns latency for matrix operation but is constrained by limited capacity. This work introduces CompAir, a scalable PIM architecture that integrates DRAM-PIM and SRAM-PIM with hybrid bonding, enabling efficient linear computations while unlocking multi-granularity data pathways. We further develop CompAirNoC, an advanced network-on-chip (NoC) with an embedded arithmetic logic unit that performs non-linear operations during data movement. Such a design offloads the centralized communication bottleneck in the channel level to distributed banks, simultaneously reducing communication overhead and area cost for scalability. Finally, we develop a hierarchical Instruction Set Architecture that ensures both flexibility and programmability of the hybrid PIM. Experiments show CompAir delivers 1.83-7.98× faster prefill and 1.95 - 6.28 × faster decoding versus state-ofthe-art PIM designs, with 3.52 × lower energy than GPU-PIM hybrids. This work presents the first systematic exploration of hybrid DRAM-PIM and SRAM-PIM architectures with innetwork computation, paving the way towards a scalable PIM system for LLM inference. Songchen Ma, Huanyu Qu, Jia Chen 0032, Junfeng Lin, Fengbin Tu |
ISCA | 7 |
| 2026 | DiTPA: A DiT-Based Action Planner Accelerator Exploiting Action-Denoising-Multimodality Redundancy for Embodied Artificial IntelligenceabstractRecent advances in multimodal vision-languageaction (VLA) models have endowed embodied artificial intelligence (embodied AI) systems with remarkable perception, reasoning, and planning capabilities. Among these VLA models, diffusion transformers (DiTs) have become the backbone for action planning due to their strong and continuous generation capability. However, multimodal DiT-based action planners typically need to generate hundreds of actions to complete a single task, and each action requires about 10-50 denoising steps. This results in extremely low action frequencies, preventing real-time deployment in embodied AI applications. In this work, we systematically analyze the inference and data distribution characteristics of DiT-based action planners and observe significant computational redundancy across action, denoising and multimodality. Motivated by these findings, we propose DiTPA, a softwarehardware co-designed DiT-based action planner accelerator to fully exploit these three levels of redundancy. We first introduce a DiTPA software framework, consisting of (1) an orientationconditioned action prediction mechanism to reuse actions with minimal orientation variation (action redundancy), (2) an alternating denoising with feature reuse technique that replaces lowimpact iterations with low-cost residual computation for noise updates (denoising redundancy), and (3) a calibrated multimodal approximate computing strategy that eliminates redundant multimodal operations based on modality lifespan and attention sparsity (multimodality redundancy). At the hardware level, the DiTPA accelerator supports this redundancy-aware framework through an action predictor, a reconfigurable processing element (PE) array, and a multimodal scheduler. Together, these innovations convert high computational redundancy into substantial performance and energy efficiency gains. Owing to softwarehardware co-design, DiTPA obtains average action frequency of 217.65Hz and task execution time of 1.73s on the LIBERO-Long benchmark, with only 1.05 W power consumption. It achieves 386.93×, 13.22×, 9.54× speedups and 2356.77 ×, 8.71 ×, 11.59 × energy-efficiency improvements over NVIDIA A40, EXION and Ditto, while maintaining the task success rate. The code is opensourced in https://github.com/fengbintu/ISCA2026-DiTPA. Longke Yan, Jiancong Li, Yongkun Wu, Fengbin Tu |
ISCA | 5 |
| 2026 | ERSRP: A 55nm 46.28 MPixels/(s·mm2) 104 FPS Efficient Real-Time Super-Resolution Processor with Layer-Fused Lightweight EngineabstractThis paper presents ERSRP, a 55nm Edge Real-Time Super-Resolution Processor that achieves 104 FPS at FHD resolution with a peak area efficiency of 46.28 MPixels/(s·mm2), 8.86× higher than the state-of-the-art. The processor is designed through a software-hardware co-optimization approach, addressing the low utilization, high memory demand, and workload imbalance challenges inherent in lightweight SR networks. At the algorithmic level, an Ultra-Lightweight Super-Resolution (ULSR) model is proposed that integrates depth-wise and point-wise separable blocks with a pixel-shuffle mechanism to achieve high-quality reconstruction (37.18 dB PSNR and 0.9581 SSIM on Set5) with only 5.62K parameters. At the hardware level, the ERSRP introduces a Lightweight Accelerated Engine (LAE) sup-porting a Layer Parallel Computing Scheme (LPCS) to improve lightweight operator throughput by 48.9%. A Point-wise Layer Fused Scheme (PLFS) further enhances utilization by 4.95× through inter-core and intra-core fusion without intermediate memory. Fabricated in a 55nm UMC CMOS process, the ERSRP achieves a throughput of 215.7 MPixels/s and supports 104 FPS real-time SR at FHD. Gaoxiang Wu, Liang Chang 0002, Jingke Wang, Zhicheng Hu, Xin Zhao 0044, Fengbin Tu, Jun Zhou 0017 |
ISCAS | 8 |
| 2026 | MRCIM: A Many-Core Reconfigurable Computing-in-Memory Processor Combining CPU and Tensor Modes for NN AccelerationabstractMany-core architecture is a promising architecture to accelerate increasingly larger neural networks (NNs). Most many-core architectures couple a standalone CPU core and a tensor core together as a compute node. However, the existing architectures suffer from inefficiency at the architecture, data flow, and control flow levels: The standalone scalar CPU core with deep out-of-order pipeline and low data parallelism per instruction incurs high hardware overhead and low throughput; Fixed proportions of CPU and tensor cores execute computations alternately in each cluster, leading to core under-utilization under diverse workloads; The MIMD parallelism strategy causes redundant instruction cache (I-Cache) accesses, which increases power consumption. To tackle the above limitations, we propose MRCIM, a many-core reconfigurable computing-in-memory (CIM) processor with reconfigurable cores featuring both CPU and tensor modes. 1) We design a reconfigurable CPU core by reusing the CIM-based tensor core’s inherent memory and computing logic to simplify the pipeline logic and improve the data parallelism of conventional CPU. 2) We propose interleaved workload execution (IWE) and adaptive workload mapping (AWM) scheduling strategies, which dynamically adjust the proportion of CPU core and tensor core in a cluster, making them work in parallel with high utilization. 3) We propose a hybrid MIMD/SIMD control flow to bypass unnecessary I-Cache accesses by instruction forwarding and sharing, thereby reducing power consumption. Experimental results show MRCIM achieves 166.48x~446.67x speedup and 96.76x~309.01x energy saving over Intel i9-13900k CPU, 12.62x~27.62x speedup and 5.49x~17.82x energy saving over NVIDIA RTX 4090 GPU. Compared with state-of-the-art NN processor architectures, our MRCIM achieves average 6.84x, 7.51x, and 3.66x speedup and average 4.57x, 3.03x, and 3.11x energy saving over Simba, LUT-ICC, and MAICC. Yiqi Wang 0005, Zihan Wu 0006, Huiming Han, Shaojun Wei, Yang Hu 0001, Chao Li 0009, Fengbin Tu, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | PAMA: Large-Scale GNN Acceleration with Pre-Aggregation in Multi-Node ArchitectureabstractGraph Neural Networks (GNNs) have demonstrated exceptional performance in real-world applications, which often involve large-scale graphs with billions of vertices and numerous features per vertex. Large-scale workload requires multi-node systems to enhance computing power and memory capacity. However, accelerating large-scale GNNs on multi-node systems faces two key challenges. (1) Graph irregularity and high-dimensional features lead to excessive redundant inter-node communication. (2) Computational dependency in GNN results in waiting issues and underutilization of computing resources in accelerator nodes. To address the challenges, this work proposes PAMA, a pre-aggregation-based multi-node architecture for GNN acceleration. For challenge (1), we propose a pre-aggregation approach to avoid redundant feature transmissions, which is facilitated by a complementary communication scheme. For challenge (2), a batched staggered aggregation-transformation pipeline dataflow is proposed to alleviate the waiting issues. Additionally, a reconfigurable computing core that dynamically adapts to different workloads is designed to further improve computing resource utilization. The evaluation results show that PAMA achieves a$9.5-16 \times$speedup over the baseline multi-node system. Fengbin Tu, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ASAP | 2 |
| 2025 | SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit SynthesisabstractDigital Computing-in-Memory (DCIM) is an innovative technology that integrates multiply-accumulation (MAC) logic directly into memory arrays to enhance the performance of modern AI computing. However, the need for customized memory cells and logic components currently necessitates significant manual effort in DCIM design. Existing tools for facilitating DCIM macro designs struggle to optimize subcircuit synthesis to meet user-defined performance criteria, thereby limiting the potential system-level acceleration that DCIM can offer. To address these challenges and enable the agile design of DCIM macros with optimal architectures, we present SynDCIM - a performance-aware DCIM compiler that employs multi-spec-oriented subcircuit synthesis. SynDCIM features an automated performance-to-layout generation process that aligns with user-defined performance expectations. This is supported by a scalable subcircuit library and a multi-spec-oriented searching algorithm for effective subcircuit synthesis. The effectiveness of SynDCIM is demonstrated through extensive experiments and validated with a test chip fabricated in a 40nm CMOS process. Testing results reveal that designs generated by SynDCIM exhibit competitive performance when compared to state-of-the-art manually designed DCIM macros. Kunming Shao, Fengshi Tian, Jiakun Zheng, Jia Chen 0032, Jingyu He, Hui Wu 0010, Jinbo Chen 0002, Xihao Guan, Fengbin Tu, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 11 |
| 2025 | ER-DCIM: Error-Resilient Digital CIM Architecture with Run-Time MAC-Cell Error CorrectionabstractDigital computing-in-memory (CIM) is an emerging solution to break through the limitations of memory wall by integrating digital logic into SRAM, which is able to achieve high area and energy efficiency with no accuracy loss. Digital CIM’s SRAM cells are still prone to errors like conventional SRAM due to noise and variation, especially under low-voltage operation for high energy efficiency. The SRAM cell errors cause multiplyaccumulation (MAC) result errors, which may seriously damage the neural network inference accuracy. However, traditional SRAM’s error correcting code (ECC) that corrects errors in one read-out row is incompatible with digital CIM, which reads out multiple rows simultaneously for computation. Detecting and correcting computational MAC errors and SRAM cell errors (i.e., MAC-cell errors) in digital CIM remain largely unexplored.To address digital CIM’s unique MAC-cell error resilience needs, we propose ER-DCIM, an error-resilient digital CIM with run-time MAC-cell error correction to guarantee computation correctness. The proposed residue code-based MAC error correction mechanism is the first to correct additive errors in the MAC result in real time during DCIM computation. Then, we propose a progressive cell error correction mechanism to correct underlying cell error in a timely manner, avoiding performance loss due to stalling computation. Further, we design a mode switcher to repurpose redundant error-resilient logic reserved for low-voltage mode to improve performance in high-voltage mode. Experimental results show that the proposed techniques enable digital CIM to maintain high throughput and energy efficiency without accuracy loss in both low-voltage and high-voltage modes. Yiqi Wang 0005, Zihan Wu 0006, Shaojun Wei, Yang Hu 0001, Fengbin Tu, Shouyi Yin |
HPCA | 6 |
| 2025 | CoXplorer: Multi-Staged Co-Exploration Framework for AI Model Compression and Accelerator DesignabstractThe rapid evolution of artificial intelligence (AI) algorithms demands efficient computing chips, positioning algorithm-hardware co-design as a crucial optimization strategy. However, automating the co-design process remains challenging due to the lack of a unified exploration framework for both algorithmic and hardware domains, as existing tools - hardware design space exploration (DSE) and compression neural architecture search (Compression NAS) - operate independently, relying entirely on manual collaboration. This paper presents CoXplorer, a co-exploration framework that connects model-compression optimization space and architecture design space. We make three key contributions: (1) a multi-staged co-design space decomposition method that enables systematic exploration of compression-hardware design choices with reduced complexity, (2) an AC-Copilot toolchain enhanced with multi-grained performance modeling driven by hardware simulation-compilation hierarchical cooperation to fulfill various evaluation requirements of co-exploration, enabling balanced simulation accuracy-efficiency trade-offs, and (3) a co-exploration workflow with hierarchical and bottleneck-guided search for harmonizing optimization objectives of both model and hardware design spaces, resulting in improved search efficiency. We validate the CoXplorer on two edge chips, which achieve 53.7% throughput and 45.8% energy efficiency improvements for the CNN acceleration, and 7.5× speedup with 9.9× energy efficiency boost for the Transformer acceleration. A case study on large language model acceleration shows CoXplorer’s extensibility to emerging workloads, enhancing LLAMA2-7B inference throughput from 6.75 to 25.46 tokens/s via co-optimization with compression and near-memory computing architecture. Songchen Ma, Yonghao Tan, Pingcheng Dong, Di Pang, Yu Liu 0007, Luhong Liang, Kwang-Ting Cheng, Fengbin Tu |
ICCAD | 11 |
| 2025 | Rethinking Control Flow in Spatial Architectures: Insights Into Control Flow Plane DesignabstractSpatial architecture is a high-performance paradigm that employs control flow graphs and data flow graphs as computation model, and producer/consumer models as execution model. However, existing spatial architectures struggle with control flow handling challenges. Upon thoroughly characterizing their PE execution models, we observe that they lack autonomous, peer-to-peer, and temporally loosely-coupled control flow handling capability. This degrades its performance in intensive control programs. To tackle the existing control flow handling challenges, Marionette, a spatial architecture with an explicit-designed control flow plane, is proposed. We elaborately develop a full stack of Marionette architecture, from ISA, compiler, simulator to RTL. Marionette's flexible Control Flow Plane enables autonomous, peer-to-peer, and temporally loosely-coupled control flow management. Its Proactive PE Configuration ensures computation-overlapped and timely configuration to promote Branch Divergence handling capability. Besides, Marionette's Agile PE Assignment improves pipeline performance of imperfect loops. Compared to state-of-the-art spatial architectures, the experimental results demonstrate that Marionette outperforms Softbrain, TIA, REVEL, and RipTide by geomean 2.88$\mathbf{\times}$, 3.38$\mathbf{\times}$, 1.55$\mathbf{\times}$, and 2.66$\mathbf{\times}$in a variety of challenging intensive control programs. Jinyi Deng, Xinru Tang, Linyun Zhang, Fengbin Tu, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
IEEE Trans. Computers | 6 |
| 2025 | Exploiting the Memory-Compute-Coupling Feature for CIM Accelerator Design OptimizationabstractSRAM computing-in-memory (CIM) accelerators have evolved as a promising solution to the memory wall problem in neural network (NN) models. By integrating memory and compute resources in each macro, CIM accelerators offer massive in-situ computing parallelism and large memory capacity, enabling spatial mapping with layer fusion and potentially keeping layers stationary in CIM. However, CIM’s memory-compute coupling (MCC) feature poses challenges in designing CIM accelerators. From an architecture aspect, designers must balance CIM’s memory and compute resources by optimizing the macro’s memory-compute ratio (MCR) configuration across diverse scenarios. From a mapping aspect, conventional mappings, which allocate each macro exclusively to each layer, face two major problems: a layer-fusion dilemma (the accelerator suffers from excessive memory access due to layer replications or performance degradation due to load imbalance) and a layer-eviction issue (storing layers stationary in CIM is usually infeasible due to limited CIM capacity). To address these challenges, this paper introduces MCC-DSE, an MCC-aware Design Space Exploration framework for architecture-mapping co-optimization of CIM accelerators. We also propose a three-axis CIM division mapping, which interleaves multiple layers in each macro to concurrently optimize memory access and performance during layer fusion as well as reserves a part of CIM memory in each macro for layer pinning. Compared to baseline architecture and mapping, MCC-DSE shows a 1.4x 8.3x EDP reduction across various workloads and chip areas. Moreover, MCC-DSE provides insights into CIM accelerator optimization, such as selecting optimal MCR and configuring CIM dynamically for different scenarios. Yongkun Wu, Jia Chen 0032, Zhenhua Zhu 0002, Jingyu He, Pingcheng Dong, Yonghao Tan, Xin Zhao 0044, Liang Chang 0002, Yu Wang 0002, Fengbin Tu, Chi-Ying Tsui, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 11 |
| 2024 | AdaP-CIM: Compute-in-Memory Based Neural Network Accelerator Using Adaptive PositabstractThis study proposes two novel approaches to address memory wall issues in AI accelerator designs for large neural networks. The first approach introduces a new format called adaptive Posit (AdaP) with two exponent encoding schemes that dynamically extend the dynamic range of its representation at run time with minimal hardware overhead. The second approach proposes using compute-in-memory (CIM) with speculative input alignment (SAU) to implement the AdaP multiply-and-accumulate (MAC) computation, significantly reducing the delay, area, and power consumption for the max exponent computation. The proposed approaches outperform state-of-the-art quantization methods and achieve significant energy and area efficiency improvements. Jingyu He, Fengbin Tu, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 2 |
| 2024 | Exploiting Similarity Opportunities of Emerging Vision AI Models on Hybrid Bonding ArchitectureabstractWhile extensive research has focused on optimizing performance and efficiency in vision-based AI accelerators, an unexplored phenomenon, Clustering Similarity Effect, presents a significant opportunity for further improvement. This effect reveals that clusters of neighboring data points exhibit similar values, enabling the potential to skip redundant computations.To fully capitalize on the potential of the Clustering Similarity Effect (CSE), this work integrates hybrid bonding DRAM technology. We conduct a comprehensive analysis of the associated design considerations and integration overhead. Leveraging these insights, we propose a novel CSE-aware architecture specifically tailored for hybrid bonding memory. This architecture facilitates similarity detection and adapts to the inherent data characteristics associated with CSE.Compared with state-of-the-art 2D/2.5D AI accelerators, the hybrid bonding baseline demonstrates an average energy efficiency improvement of $2.89 \times \sim 14.28 \times$ and an area efficiency improvement of $2.67 \times \sim 7.68 \times$. Incorporating the similarity optimizations further enhances energy efficiency and area efficiency improvement to $5.69 \times \sim 28.13 \times$ and $3.82 \times \sim 10.98 \times$, respectively. Zhiheng Yue, Huizheng Wang, Jiahao Fang, Jinyi Deng, Guangyang Lu, Fengbin Tu, Yubin Qin, Yang Wang 0089, Chao Li 0009, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ISCA | 6 |
| 2024 | Multi-Issue Butterfly Architecture for Sparse Convex Quadratic ProgrammingabstractConvex quadratic optimization solvers are extensively utilized in various domains; however, achieving optimal performance in diverse situations remains a significant challenge due to the sparse nature of objective and constraint matrices. General-purpose architectures struggle with hardware utilization when performing critical sparse matrix operations, such as factorization and multiplication. To address this issue, we introduce a pipelined spatial architecture, Multi-Issue Butterfly (MIB), which supports all primitive scalar, vector, and matrix operations required by the Alternating Direction Method of Multipliers (ADMM) based solver algorithm. The proposed architecture features a butterfly computational network with innovative working modes for each node, controlled by runtime instructions. We developed a companion scheduling method for matrix operations based on their sparsity patterns. For factorization, an elimination tree guides the network instructions reordering to avoid data hazards caused by computation dependencies. For matrix-vector multiplication, data prefetching resolves structural hazards caused by read and write conflicts to register files. Instructions without hazards are issued simultaneously to increase pipeline throughput and function unit utilization. We evaluate the proposed architecture using FPGA prototypes, representing the first fully FPGA-based generic QP solver. Our assessment includes extensive performance and efficiency bench-marks across 100 QP problems from five application domains. Compared to the same algorithm variation running on CPU backends, our prototype achieves a geometric mean of$30.5\times$end-to-end speedup,$127.0 \times$greater energy efficiency, and$16.5\times$less runtime jitter. In comparison to GPU backends, the prototype attains a geometric mean of$4.3\times$faster end-to-end speedup,$21.7\times$higher energy efficiency, and$33.4\times$less runtime jitter. Maolin Wang 0002, Ian McInerney, Bartolomeo Stellato, Fengbin Tu, Stephen P. Boyd, Hayden Kwok-Hay So, Kwang-Ting Cheng |
MICRO | 4 |
| 2024 | SWG: an architecture for sparse weight gradient computation
Fengbin Tu, Shaojun Wei, Shouyi Yin |
Sci. China Inf. Sci. | 2 |
| 2024 | DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network InferenceabstractTo accelerate the inference of deep neural networks (DNNs), quantization with low-bitwidth numbers is actively researched. A prominent challenge is to quantize the DNN models into low-bitwidth numbers without significant accuracy degradation, especially at very low bitwidths (< 8 bits). This work targets an adaptive data representation with variablelength encoding called DyBit. DyBit can dynamically adjust the precision and range of separate bit-fields to be adapted to the DNN weights/activations distribution. We also propose a hardware-aware quantization framework with a mixed-precision accelerator to trade-off the inference accuracy and speedup. Experimental results demonstrate that the ImageNet inference accuracy via DyBit is 1.97% higher than the state-of-the-art at 4-bit quantization, and the proposed framework can achieve up to 8.1× speedup compared with the original ResNet-50 model. Jiajun Zhou 0004, Jiajun Wu 0006, Yizhao Gao 0002, Yuhao Ding, Chaofan Tao, Fengbin Tu, Kwang-Ting Cheng, Hayden Kwok-Hay So, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | HDSuper: High-Quality and High Computational Utilization Edge Super-Resolution Accelerator With Hardware-Algorithm Co-Design TechniquesabstractSuper-resolution (SR) techniques have been employed to construct high-definition images from low-quality images. Various neural networks have demonstrated excellent image-reconstruction quality in SR accelerators. However, deploying SR networks on edge devices is limited by resources and power consumption induced by significant algorithm parameters, computation complexity, and external memory accesses. This work explores the hardware algorithm co-design techniques to provide an end-to-end platform with a lightweight super-resolution network (LSR) and an efficient, high-quality SR accelerator HDSuper. For algorithm design, the improved depth-wise separable convolution and pixelshuffle layers are developed to reduce network size and computation complexity by considering the hardware constraints. Also, the improved channel attention (CA) blocks enhance the image reconstruction quality. For hardware accelerator design, we design a unified computing core (UCC) combined with an efficient flattening-and-allocation (F-A) mapping strategy to support various operators with high computational utilization. In addition, we design the patch computing scheme to reduce the external memory access of the hardware architecture. Based on the evaluation, the proposed algorithm achieves high-quality image reconstruction with$37.44dB$PSNR. Finally, the FPGA demonstration and ASIC layout under UMC 55nm are achieved with low power consumption ($2.08 W$and$152 mW$) under the lowest hardware resources compared to the state-of-the-art works. Xin Zhao 0044, Liang Chang 0002, Dongqi Fan, Zhicheng Hu, Ting Yue, Fengbin Tu, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | AutoDCIM: An Automated Digital CIM CompilerabstractDigital Computing-in-Memory (DCIM) is an emerging architecture that integrates digital logic into memory for efficient AI computing. However, current DCIM designs heavily rely on manual efforts. This increases DCIM design time and limits the optimization space, making it challenging to satisfy the user specifications of diverse AI applications. This paper presents AutoDCIM, the first automated DCIM compiler. Au-toDCIM takes the user specifications as inputs and generates a DCIM macro architecture with an optimized layout. AutoDCIM’s template-based generation balances handcrafted cell design and agile macro development. AutoDCIM’s layout exploration loop analyzes diverse DCIM array partitioning schemes to satisfy user specifications. The auto-generated DCIM macros present competitive efficiency results in comparison with state-of-the-art silicon-verified DCIM macros. Jia Chen 0032, Fengbin Tu, Kunming Shao, Fengshi Tian, Xiao Huo, Chi-Ying Tsui, Kwang-Ting Cheng |
DAC | 2 |
| 2023 | PIM-HLS: An Automatic Hardware Generation Tool for Heterogeneous Processing-In-Memory-based Neural Network AcceleratorsabstractProcessing-in-memory (PIM) architectures have shown great abilities for neural network (NN) acceleration on edge devices that demand low latency under severe area constraints. Heterogeneous PIM architectures with different PIM implementation approaches such as RRAM-based PIM and SRAM-based PIM can further improve the performance. However, the automatic generation of heterogeneous PIM architectures faces the following two unresolved problems. First, existing work has not considered the design for heterogeneous PIM-based NN accelerators with multiple memory technologies. Second, for PIM with insufficient memory on edge devices, it is challenging to find the optimal runtime weight scheduling strategy in an O(L!) optimization space for the NN with L layers.In this paper, we propose PIM-HLS, an automatic hardware generation tool for heterogeneous PIM-based NN accelerators. Aiming at the problems above, we first point out that heterogeneous PIM can improve the performance under severe area constraints. Then we optimize the architectures for each NN layer by taking the advantage of different memory technologies. We also define the optimization problem of runtime weight scheduling and mapping for the first time, and propose a dynamic-programming-based weight scheduling algorithm to reduce the optimization space to O(L2). We implement PIM-HLS to automatically generate the hardware code and the instructions. Results show that we achieve an averagely 5.9× speedup with 72.8% less area compared with state-of-the-art PIM designs. Zhenhua Zhu 0002, Guohao Dai 0001, Fengbin Tu, Hanbo Sun, Kwang-Ting Cheng, Huazhong Yang, Yu Wang 0002 |
DAC | 4 |
| 2023 | ECSSD: Hardware/Data Layout Co-Designed In-Storage-Computing Architecture for Extreme ClassificationabstractWith the rapid growth of classification scale in deep learning systems, the final classification layer becomes extreme classification with a memory footprint exceeding the main memory capacity of the CPU or GPU. The emerging in-storage-computing technique offers an opportunity on account of the fact that SSD has enough storage capacity for the parameters of extreme classification. However, the limited performance of naive in-storage-computing schemes is insufficient to support the heavy workload of extreme classification. Siqi Li 0013, Fengbin Tu, Liu Liu 0017, Jilan Lin, Zheng Wang 0075, Yangwook Kang, Yufei Ding 0001, Yuan Xie 0001 |
ISCA | 2 |
| 2023 | Towards Efficient Control Flow Handling in Spatial Architecture via Architecting the Control Flow PlaneabstractSpatial architecture is a high-performance architecture that uses control flow graphs and data flow graphs as the computational model and producer/consumer models as the execution models. However, existing spatial architectures suffer from control flow handling challenges. Upon categorizing their PE execution models, we find that they lack autonomous, peer-to-peer, and temporally loosely-coupled control flow handling capability. This leads to limited performance in intensive control programs. Jinyi Deng, Xinru Tang, Linyun Zhang, Boxiao Han, Hongjun He, Fengbin Tu, Leibo Liu, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
MICRO | 8 |
| 2023 | SPG: Structure-Private Graph Database via SqueezePIRabstractMany relational data in our daily life are represented as graphs, making graph application an important workload. Because of the large scale of graph datasets, moving graph data to the cloud becomes a popular option. To keep the confidential and private graph secure from an untrusted cloud server, many cryptographic techniques are leveraged to hide the content of the data. However, protecting only the data content is not enough for a graph database. Because the structural information of the graph can be revealed through the database accessing track. In this work, we study the graph neural network (GNN), an important graph workload to mine information from a graph database. We find that the server is able to infer which node is processing during the edge retrieving phase and also learn its neighbor indices during GNN's aggregation phase. This leads to the leakage of the information of graph structure data. In this work, we present SPG, a structure-private graph database with SqueezePIR. Our SPG is built on top of Private Information Retrieval (PIR), which securely hides which nodes/neighbors are accessed. In addition, we propose SqueezePIR, a compression technique to overcome the computation overhead of PIR. Based on our evaluation, our SqueezePIR achieves 11.85× speedup on average with less than 2% accuracy loss when compared to the state-of-the-art FastPIR protocol. Ling Liang 0003, Jilan Lin, Zheng Qu 0002, Ishtiyaque Ahmad, Fengbin Tu, Trinabh Gupta, Yufei Ding 0001, Yuan Xie 0001 |
Proc. VLDB Endow. | 5 |
| 2023 | SDP: Co-Designing Algorithm, Dataflow, and Architecture for In-SRAM Sparse NN AccelerationabstractProcessing-in-memory (PIM) is a promising architecture for neural network (NN) acceleration. Most previous PIMs are based on analog computing, so their accuracy and memory cell array utilization are limited by analog deviation and ADC overhead. Digital PIM is an emerging type of PIM architecture that integrates digital logic in memory cells, which can make full utilization of the cell array without accuracy loss. However, digital PIM’s rigid crossbar architecture and full array activation raise new challenges in sparse NN acceleration. Conventional unstructured or structured sparsity cannot perform well on both the weight and input side of digital PIM. We take the opportunities from digital PIM’s bit-serial processing and in-memory customization, to tackle the above challenges by the co-designing sparse algorithm, multiplication dataflow, and PIM architecture. At the algorithm level, we propose double-broadcast hybrid-grained pruning to exploit weight sparsity with better accuracy and efficiency balance. At the dataflow level, we propose a bit-serial Booth in-SRAM multiplication dataflow for stable acceleration from the input side. At the architecture level, we design a sparse digital PIM (SDP) accelerator with customized SRAM-PIM macros to support the proposed techniques. SDP achieves$3.59\times $,$8.15\times $,$3.11\times $area efficiency, and$6.95\times $,$29.44\times $,$39.40\times $energy savings, over state-of-the-art sparse NN architectures SIGMA, SRE, and Bit Prudent. Fengbin Tu, Yiqi Wang 0005, Ling Liang 0003, Yufei Ding 0001, Leibo Liu, Shaojun Wei, Shouyi Yin, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | SPCIM: Sparsity-Balanced Practical CIM Accelerator With Optimized Spatial-Temporal Multi-Macro UtilizationabstractCompute-in-memory (CIM) is a promising technique that reduces data movement in neural network (NN) acceleration. To achieve higher efficiency, some recent CIM accelerators exploit NN sparsity based on CIM’s small-grained operation unit (OU) feature. However, new problems arise in a practical multi-macro accelerator: The mismatch between workload parallelism and CIM macro organization causes spatial under-utilization; The multiple macros’ different computation time leads to temporal under-utilization. To solve the under-utilization problems, we propose a Sparsity-balanced Practical CIM accelerator (SPCIM), including optimized dataflow and hardware architecture design. For the CIM dataflow design, we first propose a reconfigurable cluster topology for CIM macro organization. Then we regularize weight sparsity in the OU-height pattern and reorder the weight matrix based on the sparsity ratio. The cluster topology can be reshaped to match workload parallelism for higher spatial utilization. Each CIM cluster’s workload is dynamically rebalanced for higher temporal utilization. Our hardware architecture supports the proposed dataflow with a spatial input dispatcher and a temporal workload allocator. Experimental results show that, compared with the baseline sparse CIM accelerator that suffers from spatial and temporal under-utilization, SPCIM achieves$2.94\times $speedup and$2.86\times $energy saving. The proposed sparsity-balanced dataflow and architecture are generic and scalable, which can be applied to other CIM accelerators. We strengthen two state-of-the-art CIM accelerators with the SPCIM techniques, improving their energy efficiency by$1.92\times $and$5.59\times $, respectively. Yiqi Wang 0005, Fengbin Tu, Leibo Liu, Shaojun Wei, Yuan Xie 0001, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Reconfigurability, Why It Matters in AI Tasks Processing: A Survey of Reconfigurable AI ChipsabstractNowadays, artificial intelligence (AI) technologies, especially deep neural networks (DNNs), play an vital role in solving many problems in both academia and industry. In order to simultaneously meet the demand of performance, energy efficiency and flexibility in DNN processing, various reconfigurable AI chips have been proposed in the past several years. They are based on FPGA or CGRA platforms and have domain-specific reconfigurability to customize the computing units and data paths for different DNN tasks without re-produce the chips. This paper surveys typical reconfigurable AI chips from three reconfiguration hierarchies: processing element level, processing element array level, and chip level. Each reconfiguration hierarchy covers a set of important optimization techniques for DNN computation which are frequently adopted in real life. This paper lists the reconfigurable AI chip works in chronological order, discusses the hardware development process for each optimization techniques, and analyzes the necessity of reconfigurability in AI tasks processing. The trends of each reconfiguration hierarchy and insights about the cooperation of techniques from different hierarchies are also proposed. Shaojun Wei, Xinhan Lin, Fengbin Tu, Yang Wang 0089, Leibo Liu, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | STAR: An STGCN ARchitecture for Skeleton-Based Human Action RecognitionabstractSkeleton-based human action cognition (HAR) has drawn increasing attention recently. As an emerging approach for skeleton-based HAR tasks, Spatial-Temporal Graph Convolution Network (STGCN) achieves remarkable performance by fully exploiting the skeleton topology information via graph convolution. Unfortunately, existing GCN accelerators lose efficiency when processing STGCN models due to two limitations. (1) At the dataflow level, the hardware parallelism of GCN accelerators cannot match the computation parallelism of STGCN models, leading to computing resource under-utilization. (2) At the computation level, GCN accelerators fail to exploit the inherent temporal redundancy in STGCN models. To overcome the limitations, this paper proposes STAR, an STGCN architecture for skeleton-based human action recognition. STAR is designed based on the characteristics of different computation phases in STGCN. For limitation (1), a spatial-temporal dimension consistent (STDC) dataflow is proposed to fully exploit the data reuse opportunities in all the different dimensions of STGCN. For limitation (2), we propose a node-wise exponent sharing scheme and a temporal-structured redundancy elimination mechanism, to exploit the inherent temporal redundancy specially introduced by STGCN. To further address the under-utilization induced by redundancy elimination, we design a dynamic data scheduler to manage the feature data storage and schedule the features and weights for valid computation in real time. STAR achieves$4.48\times $,$5.98\times $,$2.54\times $, and$103.88\times $energy savings on average over the HyGCN, AWB-GCN, TPU, and Jetson TX2 GPU. Fengbin Tu, Mengqi Niu, Zhiheng Yue, Leibo Liu, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | DOTA: detect and omit weak attentions for scalable transformer accelerationabstractTransformer Neural Networks have demonstrated leading performance in many applications spanning over language understanding, image processing, and generative modeling. Despite the impressive performance, long-sequence Transformer processing is expensive due to quadratic computation complexity and memory consumption of self-attention. In this paper, we present DOTA, an algorithm-architecture co-design that effectively addresses the challenges of scalable Transformer inference. Based on the insight that not all connections in an attention graph are equally important, we propose to jointly optimize a lightweight Detector with the Transformer model to accurately detect and omit weak connections during runtime. Furthermore, we design a specialized system architecture for end-to-end Transformer acceleration using the proposed attention detection mechanism. Experiments on a wide range of benchmarks demonstrate the superior performance of DOTA over other solutions. In summary, DOTA achieves 152.6x and 4.5x performance speedup and orders of magnitude energy-efficiency improvements over GPU and customized hardware, respectively. Zheng Qu 0002, Liu Liu 0017, Fengbin Tu, Zhaodong Chen 0001, Yufei Ding 0001, Yuan Xie 0001 |
ASPLOS | 3 |
| 2022 | Alleviating datapath conflicts and design centralization in graph analytics accelerationabstractPrevious graph analytics accelerators have achieved great improvement on throughput by alleviating irregular off-chip memory accesses. However, on-chip side datapath conflicts and design centralization have become the critical issues hindering further throughput improvement. In this paper, a general solution, Multiple-stage Decentralized Propagation network (MDP-network), is proposed to address these issues, inspired by the key idea of trading latency for throughput. Besides, a novel High throughput Graph analytics accelerator, HiGraph, is proposed by deploying MDP-network to address each issue in practice. The experiment shows that compared with state-of-the-art accelerator, HiGraph achieves up to 2.2× speedup (1.5× on average) as well as better scalability. Haiyang Lin, Mingyu Yan, Mo Zou, Fengbin Tu, Xiaochun Ye, Dongrui Fan, Yuan Xie 0001 |
DAC | 5 |
| 2022 | Accelerating Spatiotemporal Supervised Training of Large-Scale Spiking Neural Networks on GPUabstractSpiking neural networks (SNNs) have great potential to achieve brain-like intelligence, however, it suffers low accuracy of conventional synaptic plasticity rules and low training efficiency on GPUs. Recently, the emerging backpropagation through time (BPTT) inspired learning algorithms bring new opportunities to boost the accuracy of SNNs, while training on GPUs still remains inefficient due to the complex spatiotemporal dynamics and huge memory consumption, which restricts the model exploration for SNNs and prevents the advance of neuromorphic computing. In this work, we build a framework to solve the inefficiency of BPTT-based SNN training on modern GPUs. To reduce the memory consumption, we optimize the dataflow by saving CONV/FC results only in the forward pass and recomputing other intermediate results in the backward pass. Then, we customize kernel functions to accelerate the neural dynamics for all training stages. Finally, we provide a Pytorch interface to make our framework easy-to-deploy in real systems. Compared to vanilla Pytorch implementation, our framework can achieve up to 2.13 x end-to-end speedup and consume only 0.41 x peak memory on the CIFAR10 dataset. Moreover, for the distributed training on the large ImageNet dataset, we can achieve up to 1.81 x end-to-end speedup and consume only 0.38 x peak memory. Ling Liang 0003, Zhaodong Chen 0001, Lei Deng 0003, Fengbin Tu, Guoqi Li 0002, Yuan Xie 0001 |
DATE | 4 |
| 2022 | INSPIRE: in-storage private information retrieval via protocol and architecture co-designabstractPrivate Information Retrieval (PIR) plays a vital role in secure, database-centric applications. However, existing PIR protocols explore a massive working space containing hundreds of GiBs of query and database data. As a consequence, PIR performance is severely bounded by storage communication, making it far from practical for real-world deployment. Jilan Lin, Ling Liang 0003, Zheng Qu 0002, Ishtiyaque Ahmad, Liu Liu 0017, Fengbin Tu, Trinabh Gupta, Yufei Ding 0001, Yuan Xie 0001 |
ISCA | 6 |
| 2022 | Dynamic Sparse Attention for Scalable Transformer AccelerationabstractTransformers are the mainstream of NLP applications and are becoming increasingly popular in other domains such as Computer Vision. Despite the improvements in model quality, the enormous computation costs make Transformers difficult at deployment, especially when the sequence length is large in emerging applications. Processing attention mechanism as the essential component of Transformer is the bottleneck of execution due to the quadratic complexity. Prior art explores sparse patterns in attention to support long sequence modeling, but those pieces of work are on static or fixed patterns. We demonstrate that the sparse patterns are dynamic, depending on input sequences. Thus, we propose the Dynamic Sparse Attention (DSA) that can efficiently exploit dynamic sparse patterns in attention. Compared with other methods, our approach can achieve better trade-offs between accuracy and model complexity. Moving forward, we identify challenges and provide solutions to implement DSA on existing hardware (GPUs) and specialized hardware in order to achieve practical speedup and efficiency improvements for Transformer execution. Liu Liu 0017, Zheng Qu 0002, Zhaodong Chen 0001, Fengbin Tu, Yufei Ding 0001, Yuan Xie 0001 |
IEEE Trans. Computers | 4 |
| 2022 | H2Learn: High-Efficiency Learning Accelerator for High-Accuracy Spiking Neural NetworksabstractAlthough spiking neural networks (SNNs) take benefits from the bioplausible neural modeling, the low accuracy under the common local synaptic plasticity learning rules limits their application in many practical tasks. Recently, an emerging SNN supervised learning algorithm inspired by backpropagation through time (BPTT) from the domain of artificial neural networks (ANNs) has successfully boosted the accuracy of SNNs, and helped improve the practicability of SNNs. However, current general-purpose processors suffer from low efficiency when performing BPTT for SNNs due to the ANN-tailored optimization. On the other hand, current neuromorphic chips cannot support BPTT because they mainly adopt local synaptic plasticity rules for simplified implementation. In this work, we propose H2Learn, a novel architecture that can achieve high efficiency for BPTT-based SNN learning, which ensures high accuracy of SNNs. At the beginning, we characterized the behaviors of BPTT-based SNN learning. Benefited from the binary spike-based computation in the forward pass and weight update, we first design look-up table (LUT)-based processing elements in the forward engine and weight update engine to make accumulations implicit and to fuse the computations of multiple input points. Second, benefited from the rich sparsity in the backward pass, we design a dual-sparsity-aware backward engine, which exploits both input and output sparsity. Finally, we apply a pipeline optimization between different engines to build an end-to-end solution for the BPTT-based SNN learning. Compared with the modern NVIDIA V100 GPU, H2Learn achieves$7.38\times $area saving,$5.74-10.20\times $speedup, and$5.25-7.12\times $energy saving on several benchmark datasets. Ling Liang 0003, Zheng Qu 0002, Zhaodong Chen 0001, Fengbin Tu, Yujie Wu 0002, Lei Deng 0003, Guoqi Li 0002, Peng Li 0001, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | GQNA: Generic Quantized DNN Accelerator With Weight-Repetition-Aware Activation AggregatingabstractQuantization is a prominent approach to compress model sizes of deep neural networks (DNNs), which clusters high-precision weights into a smaller set of quantization levels and represents high-precision weights by low-precision indexes. To achieve the same accuracy, nonuniform quantized DNNs (NUQ-DNNs) with unequal quantization intervals need lower index precision than uniform quantized DNNs (UQ-DNNs) with equal intervals, achieving smaller model sizes. Hence, deploying NUQ-DNNs on accelerators costs less on- and off-chip memory accesses than UQ-DNNs, which are more valuable for edge devices. However, accelerating NUQ-DNNs is nontrivial, since weight indexes cannot be directly used for computations. Previous NUQ-DNN accelerators adopt standard convolutions by decoding weight indexes into actual-weights multiplied with activations, causing abundant look-up overhead and redundant computations. In this work, we propose a weight-repetition-aware activation aggregating (WPAA) convolution approach to accelerate inference of variable-precision NUQ- and UQ-DNNs. By merging convolutions of multiple kernels, WPAA requires no look-up operation and removes redundant computations. Based on WPAA, we design a generic quantized DNN accelerator (GQNA). Furthermore, we propose a layer-adaptive kernel-reordering merging scheme to off-line adjust merging order of kernels for minimizing energy consumption of GQNA. Implemented under TSMC 28-nm technology, GQNA achieves 31.9 and 32.6 TOPS/W energy efficiency for 1-b UQ- and NUQ-VGG-16, respectively. Jianxun Yang, Fengbin Tu, Yiqi Wang 0005, Leibo Liu, Shaojun Wei, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | ADROIT: An Adaptive Dynamic Refresh Optimization Framework for DRAM Energy Saving In DNN TrainingabstractTo achieve high accuracy, DNN training usually consumes and generates myriads of data, which requires a large DRAM for efficient processing. The refresh power consumption in large DRAM has become a severe problem. Previous refresh energy saving methods have drawbacks on usability, flexibility or training supporting. We propose ADROIT, an adaptive dynamic refresh optimization framework for various DNNs and processing platforms. ADROIT dynamically adjusts the refresh rates for different types of data according to runtime loss feedback in DNN training. Data idle time, lifetime and size are taken into consideration to reduce the search space of refresh rate and remove most refresh operations. Experimental results show that ADROIT can reduce the refresh energy and total DRAM energy in DNN training by up to 98.9% and 24.7% respectively, while maintaining the accuracy. Moreover, ADROIT can automatically apply to different DNNs and hardware platforms without tedious manual configuration. Xinhan Lin, Fengbin Tu, Leibo Liu, Shaojun Wei, Shouyi Yin |
DAC | 3 |
| 2021 | Brain-Inspired Computing: Adventure from Beyond CMOS Technologies to Beyond von Neumann Architectures ICCAD Special Session PaperabstractThe goal of this special session paper is to introduce and discuss different breakthrough technologies as well as novel architectures and how they together may reshape the future of Artificial Intelligent. Our aim is to provide a comprehensive overview on the latest advances in brain-inspired computing and how the latter can be realized when emerging technologies, using beyond-CMOS devices, are coupled with novel computing paradigms that go beyond von Neumann architectures. Different emerging technologies like Ferroelectric Field-Effect Transistor (FeFET), Phase Change Memory (PCM), and Resistive RAM (ReRAM) are discussed, demonstrating their promising capability in building neuromorphic computing architectures that are inspired by nature. In addition, this special session paper discusses various novel concepts such as Logic-in-Memory (LIM), Processing-in-Memory (PIM), and Spiking Neural Networks (SNNs) towards exploring the far-reaching consequences of beyond von Neumann computing on accelerating deep learning. Finally, the latest trends in brain-inspired computing are summarized into algorithm, technology, and application-driven innovations towards comparing different PIM architectures. Hussam Amrouch, Jian-Jia Chen, Kaushik Roy 0001, Yuan Xie 0001, Indranil Chakraborty, Wenqin Huangfu, Ling Liang 0003, Fengbin Tu, Cheng Wang 0036, Mikail Yayla |
ICCAD | 8 |
| 2020 | STC: Significance-aware Transform-based Codec Framework for External Memory Access ReductionabstractDeep convolutional neural networks (DCNNs), with extensive computation, require considerable external memory bandwidth and storage for intermediate feature maps. External memory accesses for feature maps become a significant energy bottleneck for DCNN accelerators. Many works have been done on quantizing feature maps into low precision to decrease the costs for computation and storage. There is an opportunity that the large amount of correlation among channels in feature maps can be exploited to further reduce external memory access. Towards this end, we propose a novel compression framework called Significance-aware Transform-based Codec (STC). In its compression process, significance-aware transform is introduced to obtain low-correlated feature maps in an orthogonal space, as the intrinsic representations of original feature maps. The transformed feature maps are quantized and encoded to compress external data transmission. For the next layer computation, the data will be reloaded with STC's reconstruction process. The STC framework can be supported with a small set of extensions to current DCNN accelerators. We implement STC extensions to the baseline TPU architecture for hardware evaluation. The strengthened TPU achieves average reduction of 2.57x in external memory access, 1.95x~2.78x improvement of system-level energy efficiency, with a negligible accuracy loss of only 0.5%. Fengbin Tu, Man Shi, Yang Wang 0089, Leibo Liu, Shaojun Wei, Shouyi Yin |
DAC | 2 |
| 2020 | DUET: Boosting Deep Neural Network Efficiency on Dual-Module ArchitectureabstractDeep Neural Networks (DNNs) have been driving the mainstream of Machine Learning applications. However, deploying DNNs on modern hardware with stringent latency requirements and energy constraints is challenging because of the compute-intensive and memory-intensive execution patterns of various DNN models. We propose an algorithm-architecture co-design to boost DNN execution efficiency. Leveraging the noise resilience of nonlinear activation functions in DNNs, we propose dual-module processing that uses approximate modules learned from original DNN layers to compute insensitive activations. Therefore, we can save expensive computations and data accesses of unnecessary sensitive activations. We then design an Executor-Speculator dual-module architecture with support for balance execution and memory access reduction. With acceptable model inference quality degradation, our accelerator design can achieve 2.24x speedup and 1.97x energy efficiency improvement for compute-bound Convolutional Neural Networks (CNNs) and memory-bound Recurrent Neural Networks (RNNs). Liu Liu 0017, Zheng Qu 0002, Lei Deng 0003, Fengbin Tu, Shuangchen Li, Xing Hu 0001, Yufei Ding 0001, Yuan Xie 0001 |
MICRO | 4 |
| 2019 | A High Throughput Acceleration for Hybrid Neural Networks With Efficient Resource Management on FPGAabstractDeep learning is the amazing technology which has promoted the development of artificial intelligence and achieved many amazing successes in intelligent fields. Convolution-based layers (CLs), fully connected layers (FLs) and recurrent layers (RLs) are three types of layers in classic neural networks. Most intelligent tasks are implemented by the hybrid neural networks (hybrid-NNs), which are commonly composed of different layer-blocks (LBs) of CLs, FLs, and RLs. Because the CLs require the most computation in hybrid-NNs, many field-programmable gate array (FPGA)-based accelerators focus on CLs acceleration and have demonstrated great performance. However, the CLs accelerators lead to an underutilization of FPGA resources in the acceleration of the whole hybrid-NN. To fully exploit the logic resources and the memory bandwidth in the acceleration of CLs/FLs/RLs, we propose an FPGA resource efficient mapping mechanism for hybrid-NNs. The mechanism first improves the utilization of DSPs by integrating multiple small bit-width operations on one DSP. Then the LB-level spatial mapping is used to exploit the complementary features between different neural networks in the hybrid-NN. We evaluate the mapping mechanism by implementing four hybrid-NNs on Xilinx Virtex7 690T FPGA. The proposed mechanism achieves a peak performance of 1805.8 giga operations per second (GOPs). With the analysis on resource utilization and throughput, the proposed method exploits more computing power in FPGA and achieves up to $4.13 \times$ higher throughput than the state-of-the-art acceleration. Shouyi Yin, Shibin Tang, Xinhan Lin, Fengbin Tu, Leibo Liu, Shaojun Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | Reconfigurable Architecture for Neural Approximation in Multimedia ComputingabstractDue to inherent error resiliency, many high performance multimedia applications can be approximated by multilayer perceptrons (MLPs), with little quality loss. An MLP accelerator can be designed to improve the power efficiency of multimedia systems. However, previous MLP accelerators' fixed computational pattern lowers the performance when the MLP topology varies for different applications. In this paper, we propose a scheduling framework to guide mapping MLPs onto limited hardware resources. The scheduling framework adjusts the computational patterns for various MLP topologies, obtaining 30% higher performance than the conventional scheduling. We implement a reconfigurable neural architecture (RNA) to support different patterns in the framework and further improve the performance and efficiency. RNA achieves a speedup of 572× on the approximable part, whole application speedup of 7.9× and energy savings of 6.3×, with little quality loss on the benchmarks. Fengbin Tu, Shouyi Yin, Leibo Liu, Shaojun Wei |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Parana: A Parallel Neural Architecture Considering Thermal Problem of 3D Stacked MemoryabstractRecent advances in deep learning (DL) have stimulated increasing interests in neural networks (NN). From the perspective of operation type and network architecture, deep neural networks can be categorized into full convolution-based neural network (ConvNet), recurrent neural network (RNN), and fully-connected neural network (FCNet). Different types of neural networks are usually cascaded and combined as a hybrid neural network (Hybrid-NN) to complete real-life cognitive tasks. Such hybrid-NN implementation is memory-intensive with large number of memory accesses, hence the performance of hybrid-NN is often limited by the insufficient memory bandwidth. A “3D + 2.5D” integration system, which integrates a high-bandwidth 3D stacked DRAM side-by-side with a highly-parallel neural processing unit (NPU) on a silicon interposer, overcomes the bandwidth bottleneck in hybrid-NN acceleration. However, intensive concurrent 3D DRAM accesses produced by the NPU lead to a serious thermal problem in 3D DRAM. In this paper, we propose a neural processor calledParanafor hybrid-NN acceleration in consideration of thermal problem of 3D DRAM. Parana solves the thermal problem of 3D memory by optimizing both the total number of memory accesses and memory accessing behaviors. For memory accessing behaviors, Parana balances the memory bandwidth by spatial division mapping hybrid-NN onto computing resources, which efficiently avoids that masses of memory accesses are issued in a short time period. To reduce the total number of memory accesses, we design a new NPU architecture and propose a memory-oriented tiling and scheduling mechanism to exploit the maximum utilization of on-chip buffer. Experimental results show that Parana reduces the peak temperature by up to 54.72$^\circ$C and the steady temperature by up to 32.27$^\circ$C over state-of-the-art accelerators with 3D memory without performance degradation. Shouyi Yin, Shibin Tang, Xinhan Lin, Fengbin Tu, Leibo Liu, Jishen Zhao, Cong Xu 0002, Shuangchen Li, Yuan Xie 0001, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | LCP: a layer clusters paralleling mapping method for accelerating inception and residual networks on FPGAabstractDeep convolutional neural networks (DCNNs) have been widely used in various AI applications. Inception and Residual are two promising structures adopted in many important modern DCNN models, including AlphaGo Zero's model. These structures allow considerably increasing the depth and width of the network to improve accuracy, without increasing the computational budget or the difficulty of convergence. Various accelerators for DCNNs have been proposed based on FPGA platform because it has advantages of high performance, good power efficiency, and fast development round, etc. However, previous FPGA mapping methods cannot fully adapt to the different data localities among layers and other characteristics of Inception and Residual, which leads to a under-utilization of FPGA resources. We propose LCP, a Layer Clusters Paralleling mapping method to classify the layers into clusters based on their differences of parameters and data localities, and then accelerate them in different partitions of FPGA. We evaluate our mapping method by implementing Inception/Residual modules from GoogLeNet [8] and ResNet-50 [4] on Xilinx VC709 (Virtex 690T) FPGA. The results show that the proposed method fully utilizes resources and achieves up to 4.03× performance than the baseline and 2.00× performance than the state-of-the-art methods. Xinhan Lin, Shouyi Yin, Fengbin Tu, Leibo Liu, Shaojun Wei |
DAC | 3 |
| 2018 | RANA: Towards Efficient Neural Acceleration with Refresh-Optimized Embedded DRAMabstractThe growing size of convolutional neural networks (CNNs) requires large amounts of on-chip storage. In many CNN accelerators, their limited on-chip memory capacity causes massive off-chip memory access and leads to very high system energy consumption. Embedded DRAM (eDRAM), with higher density than SRAM, can be used to improve on-chip buffer capacity and reduce off-chip access. However, eDRAM requires periodic refresh to maintain data retention, which costs much energy consumption. Refresh is unnecessary if the data's lifetime in eDRAM is shorter than the eDRAM's retention time. Based on this principle, we propose a Retention-Aware Neural Acceleration (RANA) framework for CNN accelerators to save total system energy consumption with refresh-optimized eDRAM. The RANA framework includes three levels of techniques: a retention-aware training method, a hybrid computation pattern and a refresh-optimized eDRAM controller. At the training level, CNN's error resilience is exploited in training to improve eDRAM's tolerable retention time. At the scheduling level, RANA assigns each CNN layer with a computation pattern that consumes the lowest energy. At the architecture level, a refresh-optimized eDRAM controller is proposed to alleviate unnecessary refresh operations. We implement an evaluation platform to verify RANA. Owing to the RANA framework, 99.7% eDRAM refresh operations can be removed with negligible performance and accuracy loss. Compared with the conventional SRAM-based CNN accelerator, an eDRAM-based CNN accelerator strengthened by RANA can save 41.7% off-chip memory access and 66.2% system energy consumption, with the same area cost. Fengbin Tu, Shouyi Yin, Leibo Liu, Shaojun Wei |
ISCA | 1 |
| 2018 | Bit-width Adaptive Accelerator Design for Convolution Neural NetworkabstractConvolutional neural networks (CNNs) have achieved great success in many applications. Recently, various FPGA-based accelerators have been proposed to improve the performance of CNNs. However, current most FPGA-based methods only use the same bit-width selection for all CNN layers which lead to very low resource utilization and difficulty in further performance improvement. In this paper, we propose a bit-width adaptive accelerator design approach which can adapt to the CNN layers with various bit-width requirements in a same network. We construct multiple different bit-width convolutional processors to compute the CNN layers in parallel way. We partition the FPGA DSP resources and use our optimization approach to find the optimal resource allocation. On a Xilinx Virtex-7 FPGA, our design approach achieves higher throughput than the state-of-the-art FPGA-based CNN accelerators from 5.48× to 7.25× and by 6.20× on average, when we evaluate the convolutional layers of AlexNet and deeper VGG CNNs. Jianxin Guo, Shouyi Yin, Fengbin Tu, Shibin Tang, Leibo Liu, Shaojun Wei |
ISCAS | 4 |
| 2018 | An Energy Efficient JPEG Encoder with Neural Network Based Approximation and Near-Threshold ComputingabstractJPEG compression is an important part in low-power multimedia applications. This paper proposes an approach that leverages the error resilience of JPEG for different energy budgets. We select and train neural networks to approximate DCT and quantization code regions in JPEG. Then we design an architecture called reconfigurable neural unit (RNU) to accelerate trained neural networks which replace original codes. In addition, some architecture innovations are proposed to make our JPEG encoder works efficiently in near-threshold voltage region. This JPEG encoder synthesized with a 40nm CMOS technology, is able to operate at 40MHz for a 0.6V supply voltage. Results show up to 5.0 × energy reduction with 2.5 × performance degradation when compared to using a 1.0V nominal supply voltage. Shouyi Yin, Fengbin Tu, Leibo Liu, Shaojun Wei |
ISCAS | 3 |
| 2018 | GNA: Reconfigurable and Efficient Architecture for Generative Network AccelerationabstractGenerative networks have become ubiquitous in image generation applications like image super-resolution, image to image translation, and text to image synthesis. They are usually composed of convolutional (CONV) layers, convolution-based residual blocks, and deconvolutional (DeCONV) layers. Previous works on neural network acceleration focus too much on optimizing CONV layers computation such as data-reuse or parallel computation, but have low processing element (PE) utilization in computing residual blocks and DeCONV layers: residual blocks require very high memory bandwidth when performing elementwise additions on residual paths; DeCONV layers have imbalanced operation counts for different outputs. In this paper, we propose a dual convolution mapping method for CONV and DeCONV layers to make full use of the available PE resources. A cross-layer scheduling method is also proposed to avoid extra off-chip memory access in residual block processing. Precision-adaptive PEs and buffer bandwidth reconfiguration are used to support flexible bitwidths for both inputs and weights in deep neural networks. We implement a generative network accelerator (GNA) based on intra-PE processing, inter-PE processing, and cross-layer scheduling techniques. Owing to the proposed optimization techniques, GNA achieves energy efficiency of 2.05 TOPS/W with 61% higher PE utilization than traditional methods in generative network acceleration. Jiale Yan, Shouyi Yin, Fengbin Tu, Leibo Liu, Shaojun Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Deep Convolutional Neural Network Architecture With Reconfigurable Computation PatternsabstractDeep convolutional neural networks (DCNNs) have been successfully used in many computer vision tasks. Previous works on DCNN acceleration usually use a fixed computation pattern for diverse DCNN models, leading to imbalance between power efficiency and performance. We solve this problem by designing a DCNN acceleration architecture called deep neural architecture (DNA), with reconfigurable computation patterns for different models. The computation pattern comprises a data reuse pattern and a convolution mapping method. For massive and different layer sizes, DNA reconfigures its data paths to support a hybrid data reuse pattern, which reduces total energy consumption by 5.9~8.4 times over conventional methods. For various convolution parameters, DNA reconfigures its computing resources to support a highly scalable convolution mapping method, which obtains 93% computing resource utilization on modern DCNNs. Finally, a layer-based scheduling framework is proposed to balance DNA's power efficiency and performance for different DCNNs. DNA is implemented in the area of 16 mm2at 65 nm. On the benchmarks, it achieves 194.4 GOPS at 200 MHz and consumes only 479 mW. The system-level power efficiency is 152.9 GOPS/W (considering DRAM access power), which outperforms the state-of-the-art designs by one to two orders. Fengbin Tu, Shouyi Yin, Shibin Tang, Leibo Liu, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | RNA: a reconfigurable architecture for hardware neural acceleration
Fengbin Tu, Shouyi Yin, Leibo Liu, Shaojun Wei |
DATE | 1 |
| 2015 | Neural approximating architecture targeting multiple application domainsabstractApproximate computing emerges as a promising technique for high energy efficiency. Multi-layer perceptron (MLP) models can be used to approximate many modern applications, with little quality loss. However, the various MLP topologies limits the hardwares performance in all cases. In this paper, a scheduling framework is proposed to guide mapping MLPs onto limited hardware resources with high performance. We then design a reconfigurable neural architecture (RNA) to support the proposed scheduling framework. RNA can be reconfigured to accelerate different MLP topologies, and achieves higher performance than other MLP accelerators. Fengbin Tu, Shouyi Yin, Leibo Liu, Shaojun Wei |
ISCAS | 1 |