Tian Xia 0008

dblp:90/4765-8 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0002-2520-3731ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 3 first-author · 17 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Double Rounding: Nearly Lossless Adaptive Bit Switching for QAT
abstract
Model quantization is widely applied for compressing and accelerating deep neural networks (DNNs). However, conventional Quantization-Aware Training (QAT) focuses on training DNNs with uniform bit-width. The bit-width settings vary across different hardware and transmission demands, which induces considerable training and storage costs. Hence, the scheme of one-shot joint training multiple precisions is proposed to address this issue. Previous works either store a larger FP32 model to switch between different precision models for higher accuracy or store a smaller INT8 model but compromise accuracy due to using shared quantization parameters. In this paper, we introduce the Double Rounding quantization method, which fully utilizes the quantized representation range to accomplish nearly lossless bit-switching while reducing storage by using the highest integer precision instead of full precision. Furthermore, we observe a competitive interference among different precisions during one-shot joint training, primarily due to inconsistent gradients of quantization scales during backward propagation. To tackle this problem, we propose an Adaptive Learning Rate Scaling (AdaScale) technique that dynamically adapts learning rates for various precisions to optimize the training process. Additionally, we extend our Double Rounding to one-shot mixed precision training and develop a Hessian-Aware Stochastic Bit-switching (HessBit) strategy. Experimental results on the ImageNet-1K classification demonstrate that our methods have enough advantages to state-of-the-art one-shot joint QAT in both multi-precision and mixed-precision. We validate the feasibility of our method on detection and segmentation tasks, as well as on LLMs task.
Haiduo Huang, Tian Xia 0008, Pengju Ren
AAAI3
2026 PartialNet: Compute Less, Perform Better
abstract
Achieving a balance between low parameter count, reduced FLOPs, and high accuracy and throughput remains a central challenge in neural network design. To address this, we propose the partial channel mechanism (PCM), which leverages the inherent redundancy in feature map channels. PCM divides feature map channels into multiple groups, each processed by distinct operations such as convolution, attention, pooling, or identity mapping. Building on this, we introduce partial attention convolution (PATConv), a novel module that efficiently fuses convolution and visual attention within a unified framework. Our results demonstrate that PATConv can fully replace both standard convolution and visual attention modules, leading to significant reductions in parameters and FLOPs. Furthermore, PATConv enables three efficient visual attention variants: Partial Channel Attention, Partial Spatial Attention, and Partial Self-Attention. To further optimize the allocation of channel splits, we propose dynamic {partial convolution (DPConv), which adaptively learns the optimal split ratio for each layer, achieving a better trade-off between speed and accuracy. By integrating PATConv and DPConv, we develop a new hybrid network family, PartialNet, which achieves superior top-1 accuracy and inference speed on ImageNet-1K, and demonstrates strong performance on COCO detection and segmentation tasks.
Haiduo Huang, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
AAAI2
2026 MoEA: A Mixed-Precision Edge Accelerator for CNN-MSA Models with Fine-Tuning Support
Qiwei Dang, Chengyu Ma, Zhiwang Huo, Guoming Yang, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
ASP-DAC5
2026 Hierarchical-ISA Supporting Row-Wise Operands for Efficient DNN Computation
abstract
Deep neural networks (DNNs) have become a cornerstone in advancing artificial intelligence, but their complexity often leads to inefficient hardware utilization due to varying structure characteristics and excessive memory accesses. Domain-specific architectures (DSAs) offer a solution by optimizing data locality through data stationary, tiling, and layer fusion, which minimize memory access and energy consumption while boosting performance. However, current approaches lack flexibility for efficient memory management at the appropriate granularity, causing misaligned accesses and decreasing reuse potential for variable-sized tiles. To this end, we propose a hierarchical Instruction Set Architecture (hierarchical-ISA) combining a RISC-V ISA and a flexible CISC-style macro-ISA (mISA). Unlike byte-level RISC-V, mISA employs row-wise tiles as the fundamental operand, enabling efficient data reuse across adjacent iterations as well as residual connections. This mISA approach simplifies DNN programming, enhances data partitioning and manipulation efficiency, and enables a hardware-software co-designed Remapping mechanism that facilitates data reuse without physical data movement. Experiments show 31.8%–72.0% reductions in off-chip memory access across MobileNet, ResNet, Swin Transformer, MobileViT, along with speedups of 2.9× to 7.4× compared to previous DNN accelerators. We also conduct comparisons under the roofline model with NVIDIA RTX A6000 and Intel Core i7-10700K. The results show that our arithmetic intensity reaches up to 26.0× that of i7-10700K and 22.6× that of A6000.
Zhiwang Huo, Wenzhe Zhao 0001, Qiwei Dang, Chengyu Ma, Guoming Yang, Gelin Fu, Tian Xia 0008, Pengju Ren
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2026 SparseMap: A Sparse Tensor Accelerator Framework Based on Evolution Strategy
abstract
The growing demand for sparse tensor algebra (SpTA) in machine learning and big data has driven the development of various sparse tensor accelerators. However, most existing manually designed accelerators are limited to specific scenarios, and it’s time-consuming and challenging to adjust a large number of design factors when scenarios change. Therefore, automating the design of SpTA accelerators is crucial. Nevertheless, previous works focus solely on eithermapping(i.e., tiling communication and computation in space and time) orsparse strategy(i.e., bypassing zero elements for efficiency), leading to suboptimal designs due to the lack of consideration of both. A unified framework that jointly optimizes both is urgently needed. However, integrating mapping and sparse strategies leads to a combinatorial explosion in the design space(e.g., as large asO(1041) for the workloadP32×64×Q64×48=Z32×48). This vast search space renders most conventional optimization methods (e.g., particle swarm optimization, reinforcement learning and Monte Carlo tree search) inefficient. To address this challenge, we propose an evolution strategy-based sparse tensor accelerator optimization framework, called SparseMap. SparseMap constructing a more essential-factor design space with the consideration of both mapping and sparse strategy. We introduce a series of enhancements to genetic encoding and evolutionary operators, enabling SparseMap to efficiently explore the vast and diverse design space. We quantitatively compare SparseMap with prior works and classical optimization methods, demonstrating that SparseMap consistently finds superior solutions.
Boran Zhao, Haiming Zhai, Zihang Yuan, Hetian Liu, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2026 FP2: A 2-bit Floating-Point Format for Edge-AI Inference and Fine-Tuning
abstract
The increasing scale of Deep Neural Networks (DNNs) has made 2-bit quantization crucial for mitigating memory bottlenecks on edge devices. Low-bitwidth floating-point formats, offering larger dynamic ranges and avoiding quantization steps, have emerged as promising alternatives to fixed-point quantization. However, constructing viable floating-point representations with fewer than 3 bits remains challenging, as conventional formats require at least one sign bit, one exponent bit, and one mantissa bit. We address this challenge by introducing a novel data compression method that uses a 4-bit encoding space to represent two floating-point values, achieving an effective storage density of 2 bits per value. Depending on the bit width of the exponent and mantissa, we propose two different 2-bit floating-point encodings:fp2-e1m0andfp2-e0m1. Based onfp2, we introduce two computing architectures that simplify floating-point multiply-accumulate (MAC) operations into bitwise addition and logic operations, reducing floating-point computation by factors of$2\times $and$4\times $. As a result,fp2offers a practical solution for efficient inference using floating-point arithmetic on resource-constrained edge devices. Moreover, we analyze the error characteristics of thefp2data format from three perspectives. To validate the effectiveness of thefp2format, we conduct experiments on ResNet18/50 and ConvNeXt-Tiny using the CIFAR-10 and ImageNet-1K datasets. Compared tofp4, our approach reduces model size by 47%, with accuracy loss is less than 2 percentage points. Notably, on CIFAR-10, some results are close to those offp32. In contrast, when evaluated under 2-bit GPTQ,fp2demonstrates significant advantages over the baseline method on the LLAMA model. For hardware evaluation, we implement our design at the RTL level and evaluate it on both FPGA and ASIC platforms. Compared to computation architectures based onfp4, ourfp4$\times $fp2processing element (PE) array reduces area by 15% and power consumption by 8%. Furthermore, ourfp2$\times $fp2PE array achieves a remarkable 78% reduction in both area and power consumption.
Qiwei Dang, Chengyu Ma, Haiduo Huang, Gelin Fu, Zhiwang Huo, Guoming Yang, Pengchen Zong, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
IEEE Trans. Circuits Syst. I Regul. Pap.8
2026 AdapSNE: Adaptive Fireworks-Optimized and Entropy-Guided Dataset Sampling for Edge DNN Training
abstract
Training deep neural networks (DNNs) on edge devices faces challenges due to the large-scale datasets required, which are costly for edge devices, especially in large language model (LLM) tasks. To address this, a DNN-free method called Near-Memory Sampling (NMS) has been introduced. NMS reduces dimensionality and performs exemplar sampling in the reduced space, avoiding architectural bias and improving generalization. However, NMS has two limitations: 1) The mismatch between the search method and the non-monotonic property of the perplexity error function leads to the emergence of outliers; 2) Key parameter (i.e., target perplexity) is selected empirically, introducing arbitrariness and leading to uneven sampling. These two issues lead torepresentative biasof exemplars, resulting in degraded accuracy. To overcome these, we propose AdapSNE, which integrates the Fireworks Algorithm (FWA) for efficient non-monotonic search to avoid outliers and uses entropy-guided optimization for uniform sampling, ensuring representative training samples. To reduce the cost of iterative computations, we design an accelerator with custom dataflow and time-multiplexing mechanisms. Experimental results show that AdapSNE outperforms state-of-the-art methods, including both DNN-based (DQAS) and DNN-free (NMS) approaches, across small-scale image datasets, large-scale datasets, and the MMLU benchmark for LLM tasks.
Boran Zhao, Hetian Liu, Zihang Yuan, Li Zhu 0003, Lina Xie, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 A Groupwise Add-Multiply-Shift-Accumulate Datapath for Efficient DNN Accelerators
abstract
Multiply–accumulate(MAC) units account for a large fraction of the power and area in modern deep neural network (DNN) accelerators. Although low-bitwidth quantization reduces hardware overhead, the high cost of multipliers remains a fundamental bottleneck in modern accelerator datapaths. This article proposes add–multiply–shift–accumulate (AMC), a groupwise arithmetic datapath that reduces multiplier count by sharing base multiplications across groups of neighboring weights and generating residual products using lightweight shift operation. To support efficient deployment, we design a compact residual encoding and buffer organization that allows AMC arrays to be constructed with minimal decoding and control overheads. While AMC can be directly applied to existing quantized models, we further introduce a lightweight residual-aware fine-tuning (RAF) procedure to increase AMC compatibility. We implement AMC-based accelerators in SystemVerilog and synthesize them in TSMC 28-nm CMOS technology across operating frequencies from 500MHz to 1GHz. At the compute unit level, AMC reduces arithmetic area by 39.5%–62.8% and dynamic power by 32.2%–60.3% compared with optimized baseline multipliers. When integrated into CNN and Vision Transformer accelerators, AMC achieves$1.34\times $–$18.90\times $higher area efficiency and up to$10.16\times $higher energy efficiency than prior designs while preserving baseline inference accuracy.
Zhiwang Huo, Wenzhe Zhao 0001, Yuanchang Gong, Tian Xia 0008, Zheng Wang 0001, Pengju Ren
IEEE Trans. Very Large Scale Integr. Syst.4
2025 Magellan: A High-Performance Loop-Guided Prefetcher for Indirect Memory Access
abstract
Graph analytics and sparse linear algebra applications heavily rely on indirect memory access (IMA).IMAs are characterized by poor temporal and spatial locality, which causes frequent high-latency DRAM accesses.While dedicated hardware prefetchers for IMA have been explored, they target narrow access patterns and tend to introduce significant hardware complexity.Software prefetching offers a promising alternative, leveraging compiler analysis to prefetch indirection patterns.However, existing software prefetchers struggle with sparse applications due to limited loop iterations and complex IMA patterns across nested loops.We propose Magellan, a novel loop-guided software prefetcher designed to detect and schedule IMA prefetches efficiently.Magellan introduces two key innovations: (1) extracting dependence graphs across loop levels to detect complex IMA patterns and (2) capturing inner-outer loop semantics to prefetch for both current and future iterations.We evaluate Magellan on 14 memory-intensive benchmarks using real-world datasets from social networks and web graphs.Compared to the best existing IMA software prefetcher, Magellan reduces cache misses by 25% and dynamic instruction counts by 14% on average.This results in a 1.14× average speedup, with performance gains of up to 1.41×.
Gelin Fu, Tian Xia 0008, Mingzhuo Yin, Prashant J. Nair, Mieszko Lis, Pengju Ren
ISCA2
2025 GeGS-PCR: Fast and Robust Color 3D Point Cloud Registration with Two-Stage Geometric-3DGS Fusion
abstract
We address the challenge of point cloud registration using color information, where traditional methods relying solely on geometric features often struggle in low-overlap and incomplete scenarios. To overcome these limitations, we propose GeGS-PCR, a novel two-stage method that combines geometric, color, and Gaussian information for robust registration. Our approach incorporates a dedicated color encoder that enhances color features by extracting multi-level geometric and color data from the original point cloud. We introduce the Geometric-3DGS module, which encodes the local neighborhood information of colored superpoints to ensure a globally invariant geometric-color context. Leveraging LORA optimization, we maintain high performance while preserving the expressiveness of 3DGS. Additionally, fast differentiable rendering is utilized to refine the registration process, leading to improved convergence. To further enhance performance, we propose a joint photometric loss that exploits both geometric and color features. This enables strong performance in challenging conditions with extremely low point cloud overlap. We validate our method by colorizing the Kitti dataset as ColorKitti and testing on both Color3DMatch and Color3DLoMatch datasets. Our method achieves state-of-the-art performance with Registration Recall at 99.9%, Relative Rotation Error as low as 0.013, and Relative Translation Error as low as 0.024, improving precision by at least a factor of 2.
Haiduo Huang, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
NeurIPS3
2024 Differential-Matching Prefetcher for Indirect Memory Access
abstract
Indirect memory access is a critical bottleneck for modern CPUs, especially for graph analysis and sparse linear algebra applications, where the values of one data array are used to generate the fetching addresses of another array. It often causes irregular data accesses with poor temporal and spatial locality that are difficult to be captured by conventional hardware prefetchers. For many complex workloads, such indirect access patterns may have different types and are nested in a multiplelevel form. Moreover, branch mispredictions would further disturb their patterns, making them even harder to detect. As a result, existing hardware prefetchers are unable to fully prefetch complex indirect patterns. This paper proposes DMP, a low-cost hardware prefetcher to improve the memory latency in several representative irregular workloads. DMP targets four types of indirect memory access patterns including single, range, multi-level, and multi-way indirect access. DMP uses differential matching to identify an indirect access pattern in pair with its corresponding index stream. Then DMP uses a flexible prefetching mechanism to dynamically adapt the prefetching degree to maintain prefetching coverage. We evaluate the performance, energy consumption, and transistor cost of DMP among various algorithms from GAP, NAS, and HPCG benchmarks. DMP improves performance by 1.8 × (up to 5.6 ×) on average against state-of-the-art hardware prefetchers and 1.2 × (up to 2.3 ×) speedup against state-of-the-art compiler-based prefetcher Prodigy. Besides, the proposed design is optimized to take only 0.9KB of storage, making it feasible to be integrated into current CPU designs.
Gelin Fu, Tian Xia 0008, Zhongpei Luo, Wenzhe Zhao 0001, Pengju Ren
HPCA2
2023 PrSpMV: An Efficient Predictable Kernel for SpMV
abstract
Sparse Matrix-Vector Multiplication (SpMV) has been widely applied in scientific computation, industry simulation, and intelligent computation domains, which is the critical algorithm in all these applications. Due to the poor data locality, low cache usage, and extremely irregular branch patterns caused by the highly sparse and random distributions, SpMV optimization has become one of the most challenging problems for modern high-performance processors. In this paper, we study the bottlenecks of SpMV on current out-of-order CPUs and propose a novel SpMV kernel named PrSpMV to improve its performance by pursuing high predictability. Specifically, we improve the memory access regularity and locality by creating serialized access patterns so that the data prefetching efficiency and cache usage are optimized. We also improve pipeline efficiency by creating regular branch patterns to make branch prediction more accurate. Experiment results show that using the above optimization approaches, PrSpMV can eliminate nearly all branch mispredictions. Moreover, it can also significantly reduce the average L2 cache miss rate from 57% to 20% via efficiently leveraging hardware prefetchers. By using PrSpMV, stride prefetcher can be boosted with 1.31× speedup and dedicated irregular prefetcher can be improved with 1.40× speedup. Meanwhile, on commercial high-end Intel processors, it achieves 1.32× speedup against some state-of-the-art SpMV kernels.
Gelin Fu, Tian Xia 0008, Shaoru Qu, Zhongpei Luo, Pengyu Cheng, Runfan Guo, Yitong Ding, Pengju Ren
ICCD2
2023 TAQ: Top-K Attention-Aware Quantization for Vision Transformers
abstract
Model quantization can reduce the memory footprint of the neural network and improve the computing efficiency. However, the sparse attention in Transformer models is difficult to quantize, the main challenge is that changing the order of attention values and shifting attention regions might lead to incorrect prediction results. To address this problem, we propose quantization method, termed TAQ, which uses the proposed TOP-K attention-aware loss to search the quantization parameters. Further, we combine the sequential and parallel quantization methods to optimize the procedure. We evaluate the generalization ability of TAQ on various vision Transformer variants, and its performance on image classification and object detection tasks. TAQ makes the TOP-K attention ranking more consistent before and after quantization, and significantly reduces the attention shifting rate, compared with PTQ4ViT, TAQ improves the performance by 0.66 and 0.45, respectively on ImageNet and COCO, achieves the state-of-the-art performance.
Lili Shi, Haiduo Huang, Bowei Song, Meng Tan, Wenzhe Zhao 0001, Tian Xia 0008, Pengju Ren
ICIP6
2023 REMAP: A Spatiotemporal CNN Accelerator Optimization Methodology and Toolkit Thereof
abstract
Designing convolutional neural network (CNN) accelerators is getting more difficult owing to the fast-increasing types of CNN models. Some approaches use constant dataflow and microarchitecture that have lower design complexity. However, these accelerators are difficult to adapt with the highly-diverse CNN models and often suffer from low process element utilization. Some other accelerators resort to reconfigurable devices, such as field-programmable gate array (FPGA) and coarse-grained reconfigurable array to support flexible dataflows in order to fit diverse CNN layers. However, layer-by-layer processing may require more energy for frequent reconfiguration and off-chip DDR access. In this work, we introduce a reconfigurable pipeline accelerator (RPA) that can reduce the latency and DDR access by pipelining the compuptation of CNN layers. Although there have been several researches that try to speedup the design process by automatically exploring subset of the accelerator design space, identifying an available automated design tool that can effectively find the complete and optimal design scheme remains a problem, especially for the novel RPA architecture type. Unfortunately, comprehensive exploration of the whole design space faces an excessive large searching space. To tackle this problem, we propose REMAP, a toolkit for designing CNN accelerators based on the Monte Carlo tree search (MCTS) method. To efficiently search the huge design space, we propose several methods to improve searching efficiency. Evaluations show that REMAP significantly outperforms some state-of-the-art approaches; compared with GAMMA, it achieves an average speed increase of$14.75\times $, and an energy reduction of 45.45%; it also achieves a speed increase of$32.6\times $against ConfuciuX on MobileNetV2 and ResNet50. We also show an FPGA accelerator implementation which is based on REMAP’s search result, and it achieves high performance in real-time CNN tasks. This indicates that REMAP can provide high-quality design exploration with valuable insights and useful architecture design guidances.
Boran Zhao, Tian Xia 0008, Haiming Zhai, Fulun Ma, Hanzhi Chang, Wenzhe Zhao 0001, Pengju Ren
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 An Energy-and-Area-Efficient CNN Accelerator for Universal Powers-of-Two Quantization
abstract
CNN model computation on edge devices is tightly restricted to the limited resource and power budgets, which motivates the low-bit quantization technology to compress CNN models into 4-bit or lower format to reduce the model size and increase hardware efficiency. Most current low-bit quantization methods use uniform quantization that maps weight and activation values onto evenly-distributed levels, which usually results in accuracy loss due to distribution mismatch. Meanwhile, some non-uniform quantization methods propose specialized representation that can better match various distribution shapes but are usually difficult to be efficiently accelerated on hardware. In order to achieve low-bit quantization with high accuracy and hardware efficiency, this paper proposes Universal Power-of-Two (UPoT), a novel low-bit quantization method that represents values as the addition of multiple power-of-two values selected from a series of subsets. By updating the subset contents, UPoT can provide adaptive quantization levels for various distributions. For each CNN model layer, UPoT automatically searches for the optimized distribution that minimizes the quantization error. Moreover, we design an efficient accelerator system with specifically optimized power-of-two multipliers and requantization units. Evaluations show that the proposed architecture can provide high-performance CNN inference with reduced circuit area and energy, and outperforms several mainstream CNN accelerators with higher ($8\times $–$65\times $) area efficiency and ($2\times $–$19\times $) energy efficiency. Further experiments of 4/3/2-bit quantization on ResNet18/50, MobileNet_V2 and EfficientNet models show that our UPoT can achieve high model accuracy which greatly outperform other state-of-the-art low-bit quantization methods by 0.3%–6%. The results indicate that our approach provides a highly-efficient accelerator for low-bit CNN model quantization with low hardware overheads and good model accuracy.
Tian Xia 0008, Boran Zhao, Gelin Fu, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 Optimizing FPGA-Based DNN Accelerator With Shared Exponential Floating-Point Format
abstract
In recent years, low-precision fixed-point computation has become a widely used technique for neural network inference on FPGAs. However, this approach has some limitations, as certain neural networks are difficult to quantify using fixed-point arithmetic, such as those involved in super-resolution scaling, image denoising, and other scenarios that lack sufficient conditions for fine-tuning. Furthermore, deploying a floating-point precision neural network directly on an FPGA would lead to significant hardware overhead and low computational efficiency. To address this issue, this paper proposes an FPGA-friendly floating-point data format that achieves the same storage density as int8 without sacrificing inference accuracy or requiring fine-tuning. Additionally, this paper presents an FPGA-based neural network accelerator that is compatible with the proposed format, utilizing DSP resources to increase the number of DSP cascading from 7 to 16, and solving the back-to-back accumulation issue of floating-point numbers. This design achieves comparable resource consumption and execution efficiency to those of 8-bit fixed-point accelerators. Experimental results demonstrate that the accelerator proposed in this study achieves the same accuracy as the native floating point on multiple neural networks without fine-tuning, and remains high computing performance. When deployed on the Xilinx ZU9P, the performance achieves 4.072 TFlops at 250 MHz, which outperforms the previous works, including the Xilinx official DPU.
Wenzhe Zhao 0001, Qiwei Dang, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 A Comprehensive Performance Model of Sparse Matrix-Vector Multiplication to Guide Kernel Optimization
abstract
Sparse Matrix-Vector Multiplication (SpMV) is important in scientific and industrial applications and remains a well-known challenge for modern CPUs due to high sparsity and irregularity. Many researchers try to improve SpMV performance by designing dedicated data formats and computation patterns. However, out-of-order superscalar CPUs have complex micro-architectures where exist complicated interactions and restrictions among software and hardware factors. It is hard to systematically study the effectiveness of optimization methods on the overall performance, as its benefits may be undermined by other factors. In this paper, we thoroughly study the execution of SpMV on modern CPUs and propose a comprehensive performance model to reveal the critical factors and their relationships. Specifically, we first study the coding characteristics of SpMV kernels to identify key factors worthy of attention. Then we model the execution of SpMV as two overlapped parts: CPU pipeline and memory latency. Both are carefully modeled with related hardware and software factors. We also model SIMD performance with the usage of specific SIMD instructions and vector registers. Experiments show that our model matches the actual execution of real-world processors. Guided by the model, we propose SpV8, a novel SpMV kernel that optimizes critical factors to improve computation efficiency and memory bandwidth. Experiments on Intel/AMD x86 and ARM AArch64 platforms show that SpV8 outperforms several state-of-the-art approaches with large margins, achieving average$3.4\times$over Intel Math Kernel Library and$1.4\times$over the best existing approach. Such results indicate that the proposed model is capable of valuable guidance for efficient SpMV optimizations.
Tian Xia 0008, Gelin Fu, Zhongpei Luo, Lucheng Zhang, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Parallel Distributed Syst.1
2023 HIPU: A Hybrid Intelligent Processing Unit With Fine-Grained ISA for Real-Time Deep Neural Network Inference Applications
abstract
Neural network algorithms have shown superior performance over conventional algorithms, leading to the designation and deployment of dedicated accelerators in practical scenarios. Coarse-grained accelerators achieve high performance but can support only a limited number of predesigned operators, which cannot cover the flexible operators emerging in modern neural network algorithms. Therefore, fine-grained accelerators, such as instruction set architecture (ISA)-based accelerators, have become a hot research topic due to their sufficient flexibility to cover the unpredefined operators. The main challenges for fine-grained accelerators include the undesired long delays of single-image inference when performing multibatch inference, as well as the difficulty of meeting real-time constraints when processing multiple tasks simultaneously. This article proposes a hybrid intelligent processing unit (HIPU) to address the aforementioned problems. Specifically, we design a novel conversion-free data format, expanding the single-instruction multiple-data (SIMD) instruction set and optimizing the microarchitecture design to improve the performance. We also arrange the inference schedule to guarantee scalability on multicores. The experimental results show that the proposed accelerator maintains high multiply–accumulation (MAC) utilization for all common operators and achieves high performance with 4–$7\times $speedup against NVIDIA RTX2080Ti GPU. Finally, the proposed accelerator is manufactured using TSMC 28-nm technology, achieving 1 GHz for each core, with a peak performance of 13 TOPS.
Wenzhe Zhao 0001, Guoming Yang, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Very Large Scale Integr. Syst.3
2022 MI2D: Accelerating Matrix Inversion with 2-Dimensional Tile Manipulations
abstract
Matrix inversion is critical in mathematics and scientific applications. Large-scale dense matrix inversion is especially challenging for modern computers due to its heavy dependency of matrix elements and the poor temporal data locality. In this paper, we propose a novel accelerator termed MI2D, which converts matrix inversion into regular matrix multiplications using 2-dimensional cross-tile operations and novel algorithms for efficient data reuse and computations. Our evaluations show that MI2D can be easily integrated with existing matrix engines in modern high-end CPU and NPU, and effectively improves matrix inversion with 2.7× speedup against Intel Skylake CPU, and 24× against NVIDIA RTX 2080 Ti.
Lingfeng Chen, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
ACM Great Lakes Symposium on VLSI2
2021 SpV8: Pursuing Optimal Vectorization and Regular Computation Pattern in SpMV
abstract
Sparse Matrix-Vector Multiplication (SpMV) plays an important role in many scientific and industry applications, and remains a well-known challenge due to the high sparsity and irregularity. Most existing researches on SpMV try to pursue high vectorization efficiency. However, such approaches may suffer from non-negligible speculation penalty due to their irregular computation patterns. In this paper, we propose SpV8, a novel approach that optimizes both speculation and vectorization in SpMV. Specifically, SpV8 analyzes data distribution in different matrices and row panels, and accordingly applies optimization method that achieves the maximal vectorization with regular computation patterns. We evaluate SpV8 on Intel Xeon CPU and compare with multiple state-of-art SpMV algorithms using 71 sparse matrices. The results show that SpV8 achieves up to 10× speedup (average 2.8×) against the standard MKL SpMV routine, and up to 2.4× speedup (average 1.4×) against the best existing approach. Moreover, SpMV features very low preprocessing overhead in all compared approaches, which indicates SpV8 is highly-applicable in real-world applications.
Tian Xia 0008, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
DAC2
2021 CAQ: Context-Aware Quantization via Reinforcement Learning
abstract
Model quantization is a crucial step for porting Deep Neural Networks (DNNs) on embedded devices to meet the limited computation and storage resources requirement. Traditional methods usually obtain the scaling factor and quantize the weights based on the information of single layer. However, our analysis indicate that these selection methods of scaling factor overlook the differences and dependencies among layers, leading to large truncation errors or zeroing errors, which is the main reason for the performance degradation. To this end, we propose a Context-Aware Quantization (CAQ) scheme, which formalizes the model quantization as a global optimization problem and leverages reinforcement learning to search for the optimal scaling factors based on the entire model. Further, we adopt shift-based scaling factors to narrow the search space to improve the search efficiency, additionally, it reduces the computational complexity during the inference phase, and also provides a simpler and more robust activation calibration solution. We extensively test our scheme on a wide range of Neural Networks, including ResNet 50/101/152, InceptionV3 and MobileNetV2 on ImageNet, the entire search process only takes about 1 hour on a single GeForce RTX 2080 Ti. Compared with the existed methods, Our scheme can get a better performance, which could maintain the post-quantization accuracy loss less than 0.25%, while reducing memory footprint by 5%-8% and multiply accumulate (MAC) operations by 2%-4%. Besides, we further show that the CAQ can be applied on other tasks, such as object detection and segmentation.
Zhijun Tu, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren, Nanning Zheng 0001
IJCNN3
2021 Joint Critics Mechanism: A Universal Framework for Multi-targets Visual Navigation
abstract
Regarding to target-driven visual navigation problem, training a universal value function or policy function approximator is considered to be a fairly difficult task, because there may exits potential conflicts among different targets. When modeling navigation as a goal-conditional reinforcement learning problem, the algorithm can only support a relatively small number of goals, which limits the universality of the reinforcement learning based methods. In this work, we proposed a framework for multi-targets visual navigation, termed Joint Critics Mechanism, to better train the universal policy function approximator. Recognizing that target-specific network has better convergence, we use the target-specific value network to estimate the advantage of the target-universal policy network for better convergence. In this way, we avoid complexity of learning competitive targets and achieve a better convergence with a larger number of targets. For evaluation, we conduct experiments in realistic simulation environments and the results prove the rationality and effectiveness of our proposed framework.
Youzhuo Wang, Wenzhe Zhao 0001, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren
IJCNN3
2021 PIT: Processing-In-Transmission With Fine-Grained Data Manipulation Networks
abstract
In the domain of data parallel computation, most works focus on data flow optimization inside the PE array and favorable memory hierarchy to pursue the maximum parallelism and efficiency, while the importance of data contents has been overlooked for a long time. As we observe, for structured data, insights on the contents (i.e., their values and locations within a structured form) can greatly benefit the computation performance, as fine-grained data manipulation can be performed. In this paper, we claim that by providing a flexible and adaptive data path, an efficient architecture with capability of fine-grained data manipulation can be built. Specifically, we design SOM, a portable and highly-adaptive data transmission network, with the capability of operand sorting, non-blocking self-route ordering and multicasting. Based on SOM, we propose the processing-in-transmission architecture (PITA), which extends the traditional SIMD architecture to perform some fundamental data processing during its transmission, by embedding multiple levels of SOM networks on the data path. We evaluate the performance of PITA in two irregular computation problems. We first map the matrix inversion task onto PITA and show considerable performance gain can be achieved, resulting in 3x-20x speedup against Intel MKL, and 20x-40x against cuBLAS. Then we evaluate our PITA on sparse CNNs. The results indicate that PITA can greatly improve computation efficiency and reduce memory bandwidth pressure. We achieved 2x-9x speedup against several state-of-art accelerators on sparse CNN, where nearly 100 percent PE efficiency is maintained under high sparsity. We believe the concept of PIT is a promising computing paradigm that can enlarge the capability of traditional parallel architecture.
Pengchen Zong, Tian Xia 0008, Jianming Tong, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Computers2
2020 COCOA: Content-Oriented Configurable Architecture Based on Highly-Adaptive Data Transmission Networks
abstract
In domain of parallel computation, most works focus on optimizing PE organization or memory hierarchy to pursue the maximum efficiency, while the importance of data contents has been overlooked for a long time. Actually for structured data, insights on data contents (i.e. values and locations within a structured form) can greatly benefit the computation performance, as fine-grained data manipulation can be performed. In this paper, we claim that by providing a flexible and adaptive data path, an efficient architecture with capability of fine-grained data manipulation can be built. Specifically, we propose COCOA, a novel content-oriented configurable architecture, which integrates multi-functional data reorganization networks in traditional computing scheme to handle the contents of data during the transmission path, so that they can be processed more efficiently. We evaluate COCOA on various problems: complex matrix algorithm (matrix inversion) and sparse DNN. The results indicates that COCOA is versatile enough to achieve high computation efficiency in both cases.
Tian Xia 0008, Pengchen Zong, Jianming Tong, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
ACM Great Lakes Symposium on VLSI1
2020 Exploring Better Speculation and Data Locality in Sparse Matrix-Vector Multiplication on Intel Xeon
abstract
Sparse Matrix-Vector Multiplication (SpMV) is a fundamental workload of numerous applications. However, for today's high-end superscalar CPUs, such as Intel Xeon series, it is usually difficult to efficiently perform SpMV due to the irregular, matrix-dependent data access and computation pattern. While many researches focus on optimizing the memory bandwidth bound by improving data locality, this work dives into the execution of SpMV computation on Intel Xeon CPU and reveals that the bad-speculation penalty is significant in many sparse matrices and too expensive to be ignored. We study and characterize sparsity structure types that are more vulnerable to the cache miss penalty or the bad speculation penalty, respectively. Based on this insight, we proposed a fast preprocessing method, which divides the matrix into sub-matrices and determines the critical performance bound of sub-matrices according to the data distribution characteristics. On each submatrix, a combination of dedicated row reordering strategies is performed to efficiently alleviate its key performance bounds: bad speculation, cache miss, or both. Our matrix representation is based on standard Compressed Sparse Row (CSR) format, and can be easily adapted to existing SpMV libraries. Our approach is evaluated on Intel Xeon Gold 6146 Processor with a wide-range of matrices from the SuiteSparse benchmarks. The results demonstrate that the proposed approach achieves an average 1.8× speedup (up to 2.5×) on multi-threaded MKL Sparse Routines, with a quite low pre-processing cost. Additionally, when used in conjunction with MKL's original optimization method, our approach can further prompt the speedup, to average 3.6 × (up to 8.3 ×), This result indicates that our method can serve as a fast and wide-spectrum optimization method which is compatible with existing routines.
Tian Xia 0008, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
ICCD2