EDBT 2026 Demo / reviewers in the wild / expert
Xiankui Xiong
dblp:207/4010
· DBLP profile ↗
20ranked-venue papers
0as first author
20since 2021 · last 2026
0009-0009-5194-6174ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 18 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeepPiC: xPU-PIM Cluster Architecture with Adaptive Resource-Aware Task Orchestration for DeepSeek-Style MoE InferenceabstractThe success of DeepSeek has driven demand for deploying high-performance inference clusters. However, due to its Transformer-based autoregressive structure, DeepSeek remains severely bandwidth-bound, limiting the scalability of traditional xPU (e.g., GPU/TPU). While DRAM-based processing-inmemory (PIM) offers a promising solution to overcome memory bottlenecks, its use in inference clusters for DeepSeek remains underexplored due to three challenges: (1) non-trivial inter-device communication overhead; (2) the need for expert parallelism in the mixture-of-experts (MoE) module; and (3) lack of efficient task offloading to PIM. To this end, we propose DeepPiC, a novel xPU-PIM cluster architecture designed for DeepSeek-style models with multi-latent attention (MLA) and MoE modules. DeepPiC introduces a heterogeneous xPU+HBM-PIM device to accelerate low arithmetic intensity operations. It can seamlessly replace conventional xPU devices without any modification to clusterlevel interconnect topology. However, DeepPiC cannot fully realize its performance potential under static scheduling, which fails to adapt to shifting compute and memory demands driven by multidimensional variability (model heterogeneity, cluster-scale volatility, runtime dynamics). This induces inter-device communication overhead and intra-device underutilization. Thus, we propose Adaptive Resource-Aware Task Orchestration (ARTO), a two-phase strategy that decouples global model partitioning from local task assignment by dynamically coordinating (1) crossdevice parallelism optimization and (2) intra-device xPU/PIM mapping. Evaluated on DeepSeek V3-671B using H20-, A100-, and $\mathbf{H 2 0 0}$-Cluster ($\mathbf{H 2 0}$ serves as a compute-limited alternative to high-end GPUs), DeepPiC (H20+HBM-PIM) achieves up to $\mathbf{3} \times \mathbf{, 2} \times$ and $\mathbf{1. 3} \times$ speedup over $\mathbf{H 2 0}$-, A100-, and $\mathbf{H 2 0 0}$-Cluster at small batch sizes, while maintaining $\mathbf{7 4 \%}$ and $\mathbf{5 4 \%}$ of A100and $\mathbf{H 2 0 0}$-Cluster performance at large batch sizes. These results demonstrate that DeepPiC enables low-end xPU to approach or even exceed premium ones by fundamentally overcoming memory bottlenecks via adaptive scheduling that orchestrates PIM and xPU heterogeneous resources. Manni Li, Zijian Huang 0017, Wending Zhao, Yinyin Lin, Chengchen Wang, Haidong Tian, Xiankui Xiong |
ASP-DAC | 9 |
| 2026 | MPiCO: Memory-Pool-Based XPU-PIM Cluster over Optical I/O with Load-Imbalance-Aware Assignment and Execution-Site-Matching Mapping Strategies for MoE InferenceabstractWe first propose MPiCO, a memory-pool-based XPU–PIM cluster over Optical I/O, together with Load-Imbalance-Aware Assignment (LIAA) and Execution-Site-Matching Mapping (ESMM) strategies. Confining processing-in-memory (PIM) to a small set of HBMs in a hybrid HBM–DDR pool, MPiCO cuts PIM cost and offsets the resulting performance loss by eliminating inter-XPU communication overhead. LIAA resolves MoE load imbalance via dynamic assignment of warm experts to XPU/PIM, and ESMM avoids PIM-induced bandwidth loss by aligning address mapping: interleaved for XPU, PIM-friendly mapping dedicated to PIM-dies. On DeepSeek-V3 671B, MPiCO with LIAA and ESMM achieves a 2.4 × speedup and 3.5 × higher energy efficiency over H20-Electric I/O (EIO) cluster, 3 × lower PIM cost than H20-EIO with local PIM, and a 1.8 × speedup over a state-of-the-art MoE platform. Yinyin Lin, Chengchen Wang, Haidong Tian, Xiankui Xiong |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | BLOOM: Bit-Slice Framework for DNN Acceleration with Mixed-PrecisionabstractDeep neural networks (DNNs) have revolutionized numerous AI applications, but their vast model sizes and limited hardware resources present significant deployment challenges. Model quantization offers a promising solution to bridge the gap between DNN size and hardware capacity. While INT8 quantization has been widely used, recent research has pushed for even lower precision, such as INT4. However, the presence of outliers-values with unusually large magnitudes-limits the effectiveness of current quantization techniques. Previous compression-based acceleration methods that incorporate outlieraware encoding introduce complex logic. A critical issue we have identified is that serialization and deserialization dominate the encoding/decoding time in these compression workflows, leading to substantial performance penalties during workflow execution. To address this challenge, we introduce a novel computing approach and a compatible architecture design named “BLOOM”. BLOOM leverages the strengths of the “bit-slicing” method, effectively combining structured mixed-precision and bit-level sparsity with adaptive dataflow techniques. The key insight of BLOOM is that outliers require higher precision, while normal values can be processed at lower precision. By interleaving 4-bit values, we efficiently exploit the inherent sparsity in the highprecision components. As a result, the BLOOM-based accelerator outperforms the existing outlier-aware accelerators by an average $1.2 \sim 4.0 \times$ speedup and $24.6 \% \sim 71.3 \%$ energy reduction, respectively, without model accuracy loss. Fangxin Liu, Ning Yang 0012, Zongwu Wang, Xuanpeng Zhu, Haidong Yao, Xiankui Xiong, Li Jiang 0002, Haibing Guan |
DAC | 6 |
| 2025 | OPS: Outlier-Aware Precision-Slice Framework for LLM AccelerationabstractLarge language models (LLMs) have transformed numerous AI applications, with on-device deployment becoming increasingly important for reducing cloud computing costs and protecting user privacy. However, the astronomical model size and limited hardware resources pose significant deployment challenges. Model quantization is a promising approach to mitigate this gap, but the presence of outliers in LLMs reduces its effectiveness. Previous efforts addressed this issue by employing compression-based encoding for mixed-precision quantization. These approaches struggle to balance model accuracy with hard-ware efficiency due to their value-wise outlier granularity and complex encoding/decoding hardware logic. To address this, we propose OPS (Outlier-aware Precision-Slicing), an acceleration framework that exploits massive sparsity in the higher-order part of LLMs by splitting 16-bit values into a 4-bit/12-bit format. Crucially, OPS introduces an early bird mechanism that leverages the high-order 4-bit computation to predict the importance of the full calculation result. This mechanism enables efficient computational skips by continuing execution only for important computations and using preset values for less significant ones. This scheme can be efficiently integrated with existing hardware accelerators like systolic arrays without complex encoding/decoding. As a result, OPS outperforms state-of-the-art outlier-aware accelerators, achieving a 1.3 − 4.3× performance boost with minimal model accuracy loss. This approach enables more efficient on-device LLM deployment, effectively balancing computational efficiency and model accuracy. Fangxin Liu, Ning Yang 0012, Zongwu Wang, Xuanpeng Zhu, Haidong Yao, Xiankui Xiong, Li Jiang 0002 |
DATE | 6 |
| 2025 | LsCMM-H: A TCO-Optimized Hybrid CXL Memory Expansion Architecture with Log StructureabstractIn the era of big data, the demand for memory capacity in modern computing systems is surging. The CXL-SSD, NAND Flash-based memory expander using emerging Compute Express Link (CXL), has become a promising solution for efficient memory expansion. However, the memory-expansion scenario poses severe performance and endurance challenges for CXL-SSDs, and existing works fail to fully address them due to the usage of traditional SSDs as back-end media. To optimize these aspects, we propose LsCMM-H, a Total-Cost-of-Ownership (TCO) -efficient CXL-SSD architecture with Zoned Namespace (ZNS) SSDs as back-end media for better latency and lifetime. LsCMM-H employs hardware-software co-designed log management, low-overhead data-tiering-based garbage collection mechanism, and read acceleration to leverage the benefits of ZNS. Based on our evaluation, LsCMM-H reduces tail latency by 49.9%, improves throughput by 41.9%, endurance by 280.5%, and saves TCO by 72.4% compared to vanilla CXL-SSD. The additional comparison also demonstrates the superiority of our proposed log structure. Xiangrui Zhang, Sirui Peng, Zhiwang Guo, Haidong Tian, Xiankui Xiong, Xiaoyong Xue, Xiaoyang Zeng |
ICCAD | 6 |
| 2025 | CPSnB: Compressing and Processing Spatial Similarity near Memory Bank for DNNsabstractNear memory bank processing (NMBP) architecture only benefits memory-bound operations of DNNs in terms of energy consumption. Drawing on the insight that data compression can reduce the compute density of operators, transforming compute-bound operations into memory-bound operations, We propose CPSnB, a NMBP architecture combined with preserving numerical jump-spatial similarity compression (PNJ-SSC) method. CPSnB provides a tiling strategy for optimizing operators of different DNN models. Compared to the systolic host-side accelerator and existing dense and sparse NMBP, CPSnB significantly reduces energy consumption. Analysis of the experimental results indicates that a 60% compression ratio of activation can enhance the versatility of CPSnB in processing DNN operators to 22.3 times. Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong |
ISCAS | 9 |
| 2025 | APCPU: Adaptive-Pooling Compression Processing Unit for Energy-Efficient DNNs ProcessingabstractIntegrating compression in the multiply-and-accumulate (MAC) path can significantly improve the energy efficiency of DNN operators. However, existing unstructured sparse compression (USSC) methods struggle to effectively compress activations with low sparsity. Computing core processing USSC face challenges such as load imbalance and complex index control circuit design. Based on insights into local spatial correlation, a block-wise adaptive-pooling compression (APC) method is proposed to achieve a high compression ratio for activations. Furthermore, this paper proposes an APCPU to integrate APC into the MAC path with minimal overhead, facilitating highly energy-efficient sparse processing of DNN operators. Leveraging a hybrid data flow design to achieve load balancing results in speedups of 1.25× to 1.33×. The experiment results show that the APCPU achieves energy savings of 1.35× and 1.27× compared to JPZ-PU, and 2.63× and 2.71× compared to CSC-PU when evaluated on AlexNet and Bert. Wang Wang, Wending Zhao, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong |
ISCAS | 9 |
| 2025 | GPOS: A General and Precise Offloading Strategy for High Generality of DNN Acceleration by OCP and NDP Co-OptimizingabstractThe arithmetic intensity (ArI) of different DNNs can be opposite. This challenges the generality of single acceleration architectures, including both dedicated on-chip processing (OCP) and near-data processing (NDP). Neither architecture can simultaneously achieve optimal energy efficiency and performance for operators with opposite ArI. It is relatively straightforward to think of combining the respective advantages of OCP and NDP. However, few publications have addressed their real-time co-optimization, primarily due to the lack of a quantifiable offloading method. Here, we propose GPOS, a general and precise offloading strategy that supports high generality of DNN acceleration. GPOS comprehensively considers the complex interactions between OCP and NDP, including hardware configurations, dataflow (DF), DNN model, and interdie data movements (DMs). Three quantifiable indicators—ArI, execution cost (Ex-cost), and DM-cost—are employed to precisely evaluate the impacts of these interactions on energy and latency. GPOS adopts a four-step flow with progressive refinement: each of the first three steps focuses on a single indicator at the operator level, while the final step performs context-based calibration to address operator interdependencies and avoid offsetting NDP benefits. Narrowing down offloading candidates in step 1 and step 3 significantly accelerates real-time quantitative analysis. Optimized mapping techniques and NDP-input stationary DF are proposed to reduce Ex-cost and extend operator types supported by NDP. Next, for the first time, sparsity—one of the most popular methods for energy optimization that can alter data reuse or ArI—is quantitatively investigated for its impacts on offloading using GPOS. Our evaluations include representative DNNs, including GPT-2, Bert, RNN, CNN, and MLP. GPOS achieves the minimum energy and latency for each benchmark, with geometric mean speedups of 49.0% and 94.1%, and geometric mean energy savings of 45.8% and 89.2% over All-OCP and All-NDP, respectively. GPOS also reduces offloading analysis latency by a geometric mean of 92.7% compared to the evaluation that traverses each operator and its relative combinations. On average, sparsity further improves performance and energy efficiency by increasing the number of operators offloaded to NDP. However, for DNNs where all operators exhibit either very high or very low ArI, the number of offloaded operators remains unchanged, even after sparsity is applied. Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2024 | FullSparse: A Sparse-Aware GEMM Accelerator with Online Sparsity PredictionabstractLeveraging sparsity optimizes storage and computation for resource-constrained devices in Deep Learning Neural Networks (DNNs). While neural networks naturally incorporate sparsity through operations like ReLU and quantization, diverse sparsity levels (0.2% to 99%) pose challenges for the design of computational units. In this paper, we provide an energy-efficient GEMM accelerator named FullSparse which is designed for diverse applications, accommodating varying sparsity levels in matrix multiplication (0.2% to 99%). This paper introduces three features for nuanced sparsity support: multi-sparsity control, predictive result sparsity, and a multi-sparsity-compatible PE array. Experimental evaluations affirm that our implementation while ensuring adaptability to sparsity, exhibits superior computational power comparable to the existing designs. Jiangnan Yu, Fan Yang 0001, Yuxuan Qiao, Xiankui Xiong, Haidong Yao, Yecheng Zhang |
CF | 6 |
| 2024 | ARCTIC: Agile and Robust Compute-In-Memory Compiler with Parameterized INT/FP Precision and Built-In Self TestabstractDigital Compute-in-Memory (DCIM) architectures are playing an increasingly vital role in artificial intelligence (AI) applications due to their significant energy efficiency enhancement. Coupling memory and computing logic in DCIM requires extensive customization of custom cells and layouts, thus increasing design complexity and implementation effort. To adapt to the swiftly evolving AI algorithms, DCIM compiler for agile customization is required. Previous DCIM compilers accelerate the customization process but only focus on integer computation. Moreover, with technology node scaling down, design-for-test circuits are critical for robust chip design, while previous built-in-self-test (BIST) schemes for traditional memory fail to offer support for DCIM. This paper presents ARCTIC, an agile and robust DCIM compiler supporting parameterized integer/floating-point formats with corresponding BIST circuits. To support variable precision formats (including integer and floating-point), ARCTIC applies adaptive topology and layout optimization schemes for optimal performance. The compiler is also equipped with DCIM-friendly MarchCIM BIST circuits for efficient post-silicon tests with negligible area overhead. The energy efficiency of the generated DCIM macros remains competent with the state-of-the-art counterparts. Haozhe Zhu, Siqi He, Chengchen Wang, Xiankui Xiong, Haidong Tian, Xiaoyang Zeng, Chixiao Chen |
DATE | 6 |
| 2024 | Multi-Weather Degradation-Aware Transformer for Image RestorationabstractRestoring images under different adverse weather conditions with a single model is practical in many applications. Most existing weather restoration approaches are only able to handle a specific type of degradation, which is often insufficient in real-world scenarios where the weather type is unknown. In this paper, we propose a holistic solution to solve multiple weather degradations using a single model. Specifically, we build a weather-type aware Transformer, an efficient architecture that can restore images degraded by different adverse weathers with the same set of parameters. For model training, we first use contrastive loss to train an auxiliary hypernetwork capable of extracting content-independent, distortion-aware feature embeddings. Guided by these weather-dependent features, the image restoration Transformer can adaptively modulate its parameters using hypernetworks and feature-wise linear modulation blocks, conducting both local and global operations adaptively for images with different degradations. Qualitative and quantitative results on the multi-weather benchmark demonstrate that our model achieves significant improvements compared with previous state-of-the-arts, with even less computational cost. Ruoxi Zhu, Minfeng Wu, Xiankui Xiong, Xuanpeng Zhu, Yibo Fan |
ICASSP | 3 |
| 2024 | FSMM: An Efficient Matrix Multiplication Accelerator Supporting Flexible SparsityabstractSparse matrix multiplication is a critical operation in deep learning. However, matrix sparsity leads to irregular data flow, which would degrade the efficiency of matrix multiplication. Traditional accelerators, equipped with additional hardware units to address this issue, often experience the issue of low hardware utilization. Furthermore, N : M structured sparsity and corresponding hardware architectures face challenges such as accuracy degradation, limited flexibility, and restricted applicability. In this paper, we propose a Flexible Sparse Matrix Multiplication Accelerator (FSMM), which can improve the efficiency of sparse matrix multiplication through both algorithmic-level and hardware-level optimizations. At the algorithmic level, we propose the matrix-matrix multiplication with block-level outer production and fine-grained matrix reordering algorithm. The algorithm balances the sparsity of each column of a matrix block, which improves matrix compression, balances the load, and speeds up computation. This algorithm reduces storage by 8.2% ~ 85.9%. At the hardware-level, we introduce a flexible architecture for matrix multiplication. It selects the most suitable data path to complete matrix multiplication based on the sparsity of the reordered matrix. FSMM achieves a speedup of 1.90× ~ 16.18× over Systolic Array and 1.70× ~ 2.87× over the existing TSTC approach. Yuxuan Qiao, Fan Yang 0001, Yecheng Zhang, Xiankui Xiong, Haidong Yao |
ICCAD | 4 |
| 2024 | SFFTNet: Sparse Feature Fusion Transformer Network for Image DeblurringabstractThe U-Net structure, with its an encoder-decoder architecture, has been widely adopted by many deep learning methods for image deblurring. Most methods concentrate on the design of encoder and decoder block and use skip connections to connect them. However, this simple skip connection strategy is insufficient to fully exploit the correlation of multi-scale features, which can result in a potential loss of deblurring performance. To address this issue, we design an effective Sparse Feature Fusion Transformer capable of integrating multi-scale features to replace the skip connections in the U-Net framework. Specifically, We propose a cross-attention mechanism with a learnable top-k selection operator to adaptively preserve highly correlated cross-attention values for feature fusion. This approach ensures that the fused multi-scale features effectively integrate contextual information, resulting in high-quality image deblurring. Additionally, we introduce the position encoding generator scheme to ensure that our deblurring network can be applied to images of any size. Comprehensive experimental results demonstrate that our proposed method outperforms the the state-of-the-art methods. Faxing Lei, Ming-e Jing, Xiankui Xiong, Xuanpeng Zhu, Yibo Fan |
ISCAS | 5 |
| 2024 | LauWS: Local Adaptive Unstructured Weight Sparsity of Load Balance for DNN in Near-Data ProcessingabstractMemory wall issue has become the overwhelming bottleneck of future systems due to the explosive parameter growth and low computing density large language model (LLM). Near-data processing (NDP) could alleviate data traffic and energy consumption, but the storage demand of LLM is still enormous. Weight sparsity is helpful for reducing data capacity. Unstructured sparsity sacrifices less accuracy compared to structured one, but the random non-zero values distribution in NDP leads to load imbalance among parallel processing units. Here we propose LauWS which is seamlessly combined into various prior arts of sparsity. LauWS follows the local characteristics of feature distribution in weight matrix for various models, preserving even tiny features and discarding non-feature values as far as possible region by region. That is the key for LauWS achieving a trade-off between high prune ratio (PR) and less accuracy loss (AL). Evaluations are carried out based on a GDDR6-based bank-NDP system. The typical optimization compared to the no-prune includes 38% speedup at 0.8PR with no AL for MLP, 22.7% speedup at 0.5PR with no AL for GPT-2, 23.6% speedup at 0.5PR with the lowest perplexity for OPT-125m. Wang Wang, Manni Li, Yinyin Lin, Guhyun Kim, Yosub Song, Chengchen Wang, Xiankui Xiong |
ISCAS | 10 |
| 2023 | Graph Representation Learning for Microarchitecture Design Space ExplorationabstractDesign optimization of modern microprocessors is a complex task due to the exponential growth of the design space. This work presents GRL-DSE, an automatic microarchitecture search framework based on graph embeddings. GRL-DSE uses graph representation learning to build a compact and continuous embedding space. Multi-objective Bayesian optimization using an ensemble surrogate model conducts microarchitecture design space exploration in the graph embedding space to efficiently and holistically optimize performance-power-area (PPA) objectives. Experimental studies on RISC-V BOOM show that GRLDSE outperforms previous techniques by 74.59% on Pareto front quality and outperforms manual designs in terms of PPA. Xiaoling Yi, Jialin Lu, Xiankui Xiong, Dong Xu 0015, Fan Yang 0001 |
DAC | 3 |
| 2023 | TPNoC: An Efficient Topology Reconfigurable NoC GeneratorabstractWith the core count increasing in Chip to support various data-intensive workloads, Network-on-chip (NoC) has become the better solution for addressing on-chip interconnection. Various data-intensive workloads have different traffic patterns that require NoC with different topologies and microarchitectures. On the one hand, topology type selection has a great influence on the final performance, area, and energy. However, it is difficult to change the topology type in the traditional NoC RTL design process once it is determined. On the other hand, NoC platforms have many tunable micro-architecture design parameters, which require careful design space exploration to trade off performance advantages and overhead. Designing and validating each microarchitecture of NoCs to account for various trade-offs will greatly exacerbate the design cost issue. Jiangnan Yu, Fan Yang 0001, Xiaoling Yi, Chixiao Chen, Jun Tao 0001, Dong Xu 0015, Xiankui Xiong |
ACM Great Lakes Symposium on VLSI | 7 |
| 2023 | Luminance-Preserving Visible and Near-Infrared Image Fusion Network with Edge GuidanceabstractNear-infrared (NIR) images and visible (VIS) images can provide mutually complementary information for each other, thus the fusion of the two modalities can create images of high quality even in adverse conditions. However, the luminance of NIR and VIS images may be inconsistent in some regions, resulting in color distortion and unrealistic appearance in the fused images. The existing methods perform poorly at luminance retention. Aiming at the problem and based on deep learning framework, we propose an edge-guided method which can be applied to the image fusion network. Edge maps are utilized as prior knowledge of images to boost the performance of the neural network. Additionally, we propose a luminance-preserving loss function combined with max-edge loss to further improve the image quality. Experimental results show the superiority of our method. Ruoxi Zhu, Yi Ling, Xiankui Xiong, Dong Xu 0015, Xuanpeng Zhu, Yibo Fan |
ICIP | 3 |
| 2022 | A 11.6μ W Computing-on-Memory-Boundary Keyword Spotting Processor with Joint MFCC-CNN Ternary QuantizationabstractThis paper presents an ultra-low-power keyword spotting processor using an algorithm-architecture co-design approach. Joint MFCC-CNN ternary weight quantization is proposed to reduce power consumption. The Mel filter and the DCT module are merged into one matrix multiplication. The merged coefficients and the weights of the rest NN classifier are ternary-quantized, causing less than 3% accuracy loss but 39 × energy efficiency improvement. Moreover, a Computing-on-Memory-Boundary macro is adopted to store the quantized coefficients and weights, and perform matrix multiplications. Compared to the existing computing-in-memory technology, the proposed technique can reduce power consumption due to the higher utilization ratio. To verify the proposed techniques, a keyword spotting processor prototype is designed with 28nm CMOS technology. Simulation results show that the prototype achieves power consumption of 11.6μ W under a power supply of 0.72V and a clock frequency of 250KHz. Xinru Jia, Haozhe Zhu, Yunzheng Wang, Jinshan Zhang 0006, Xiankui Xiong, Dong Xu 0015, Chixiao Chen, Qi Liu 0010 |
ISCAS | 6 |
| 2022 | An Automated Compiler for RISC-V Based DNN AcceleratorabstractMultifarious hardware accelerators are developed for the widely used Deep neural networks (DNN). Nowadays the SoCs composed of a general processor and a coupled accelerator are becoming prevalent. Compared to the specialized DNN accelerator for one specific DNN, this kind of coupled architecture is programmable and supports diverse DNNs. However, for the low-level programming interface of the co-processor-like accelerator and the multi-hierarchy memory structure, programming for the DNN accelerator is not easy work. Meanwhile, there are a couple of tensor compilers that deploy the DNN on various hardware. In this work, we combine the flexibility of the tensor compiler and the high efficiency of the hardware accelerator by proposing an automated compiler that can compile tensor programs and generate high-performance programs for programmable DNN accelerators. Our compiler is based on TVM [1] and target at Rocket Chip Coprocessor (RoCC) [2]. The compiler is flexible and supports many kinds of RISC-V instructions. The programmer can define the hardware constraints in the proposed compiler which makes the generated code more efficient. Our compiler can lower the program with the ping-pong strategy and the generated code can achieve 26% speed up compared to the baseline. Wuzhen Xie, Xiaoling Yi, Ruiyao Pu, Xiankui Xiong, Haidong Yao, Chixiao Chen, Jun Tao 0001, Fan Yang 0001 |
ISCAS | 6 |
| 2022 | NNASIM: An Efficient Event-Driven Simulator for DNN Accelerators with Accurate Timing and Area ModelsabstractIn this paper, we propose NNASIM, an efficient timing and area accurate event-driven simulator for custom DNN accelerators. NNASIM is a highly-modular and highly parameterized modeling framework. We build accurate timing and area models for common accelerator modules like GEMM, ALU array, and crossbar using ASIC synthesis flows. These models are fed into the event-driven simulator for fast simulation. NNASIM is integrated with a RISC-V simulator. This approach guarantees the functional correctness of the accelerator simulation at the instruction level. The experimental results show that our model evaluates the performance and area of DNN accelerators with less than 0.76% and 2.83% error, respectively, compared to RTL implementations. NNASIM allows designers to model the performance and area of the accelerator at a high level, and thus enables the systematic microarchitecture design space exploration of the custom accelerators. Index Terms accelerators. Xiaoling Yi, Jiangnan Yu, Xiankui Xiong, Dong Xu 0015, Chixiao Chen, Jun Tao 0001, Fan Yang 0001 |
ISCAS | 4 |