Tian Zhi

dblp:134/7486 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
8since 2021 · last 2025
0009-0003-3449-0474ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CISGraph: A Contribution-Driven Accelerator for Pairwise Streaming Graph Analytics
abstract
Recent research observed that pairwise query is practical enough in real-world streaming graph analytics. Given a pair of distinct vertices, existing approaches coalesce or prune vertex activations to decrease computations. However, they still suffer from severe invalid computations because they ignore contribution variations in graph updates, hindering performance improvement. In this work, we propose to enhance pairwise analytics by taking updates contributions into account. We first identify that graph updates from one batch have a distinct impact on query results and experience obvious diverse computation overheads. We then introduce CISGraph, a novel Contribution-driven pairwise accelerator with valuable updates Identification and Scheduling. Specifically, inspired by triangle inequality, CIS-Graph categorizes graph updates into three levels according to contributions, prioritizes valuable updates, delays possible-valuable updates, and drops useless updates to eliminate wasteful computations. As far as we know, CISGraph is the first hardware accelerator that supports efficient pairwise queries on streaming graphs. Experimental results show that CISGraph substantially outperforms state-of-the-art streaming graph processing systems by 25 x on average in response time.
Songyu Feng, Mo Zou, Tian Zhi, Zidong Du
DATE3
2024 Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
abstract
Deploying advanced large language models on edge devices, such as smartphones and robotics, is a growing trend that enhances user data privacy and network connectivity resilience while preserving intelligent capabilities. However, such a task exhibits single-batch computing with incredibly low arithmetic intensity, which poses the significant challenges of huge memory footprint and bandwidth demands on limited edge resources. To address these issues, we introduce Cambricon-LLM, a chiplet-based hybrid architecture with NPU and a dedicated NAND flash chip to enable efficient on-device inference of 70B LLMs. Such a hybrid architecture utilizes both the high computing capability of NPU and the data capacity of the NAND flash chip, with the proposed hardware-tiling strategy that minimizes the data movement overhead between NPU and NAND flash chip. Specifically, the NAND flash chip, enhanced by our innovative in-flash computing and on-die ECC techniques, excels at performing precise lightweight on-die processing. Simultaneously, the NPU collaborates with the flash chip for matrix operations and handles special function computations beyond the flash's on-die processing capabilities. Overall, Cambricon-LLM enables the on-device inference of 70B LLMs at a speed of 3.44 token/s, and 7B LLMs at a speed of 36.34 token/s, which is over 22× to 45× faster than existing flash-offloading technologies, showing the potentiality of deploying powerful LLMs in edge devices.
Zhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai, Ziyuan Nan, Xinkai Song, Yifan Hao 0001, Jie Zhang 0048, Tian Zhi, Yongwei Zhao 0001, Zidong Du, Xing Hu 0001, Qi Guo 0001, Tianshi Chen 0002
MICRO10
2023 Cambricon-R: A Fully Fused Accelerator for Real-Time Learning of Neural Scene Representation
abstract
Neural scene representation (NSR) initiates a new methodology of encoding a 3D scene with neural networks by learning from dozens of photos taken from different camera positions. NSR not only achieves significant improvement in the quality of novel view synthesis and 3D reconstruction but also reduces the camera cost from the expensive laser cameras to the cheap color cameras on the shelf. However, performing 3D scene encoding using NSR is far from real-time due to the extremely low hardware utilization (only utilization of hardware peak performance), which greatly limits its applications in real-time AR/VR interactions
Xinkai Song, Yuanbo Wen 0001, Xing Hu 0001, Tianbo Liu 0006, Haoxuan Zhou, Husheng Han, Tian Zhi, Zidong Du, Wei Li 0008, Rui Zhang 0040, Chen Zhang 0001, Lin Gao 0004, Qi Guo 0001, Tianshi Chen 0002
MICRO7
2023 Hardware Acceleration for SLAM in Mobile Systems
Zhe Fan, Yifan Hao 0001, Tian Zhi, Qi Guo 0001, Zidong Du
J. Comput. Sci. Technol.3
2023 DyPipe: A Holistic Approach to Accelerating Dynamic Neural Networks with Dynamic Pipelining
Yimin Zhuang, Xing Hu 0001, Xiaobing Chen, Tian Zhi
J. Comput. Sci. Technol.4
2022 Tetris: A Heuristic Static Memory Management Framework for Uniform Memory Multicore Neural Network Accelerators
Xiaobing Chen, Hao Qi 0004, Shaohui Peng, Yimin Zhuang, Tian Zhi, Yunji Chen
J. Comput. Sci. Technol.5
2022 Cambricon-G: A Polyvalent Energy-Efficient Accelerator for Dynamic Graph Neural Networks
abstract
Graph neural networks (GNNs), which extend traditional neural networks for processing graph-structured data, have been widely used in many fields. The GNN computation mainly consists of theedge processingto generate messages by combining the edge/vertex features and thevertex processingto update the vertex features with aggregated messages. In addition to nontrivial vector operations in the edge processing, huge random accesses and neural network operations in the vertex processing, the graph topology of GNNs may also vary during the computation (i.e., dynamic GNNs). The above characteristics pose significant challenges on existing architectures. In this article, we propose a novel accelerator named CAMBRICON-G for efficient processing of both dynamic and static GNNs. The key of CAMBRICON-G is to abstract the irregular computation of a broad range of GNN variants to the process of regularly tiledadjacent cuboid(which extends the traditional adjacent matrix of graph by adding the dimension of vertex features). The intuition is that the adjacent cuboid facilitates exploitation of both data locality and parallelism by offeringmultidimensional multilevel tiling(including spatial and temporal tiling) opportunities. To perform themultidimensional spatial tiling, the CAMBRICON-G architecture mainly consists of the cuboid engine (CE) and hybrid on-chip memory. The CE has multiple vertex processing units (VPUs) working in a coordinated manner to efficiently process the sparse data and dynamically update the graph topology with dedicated instructions. The hybrid on-chip memory contains the topology-aware cache and multiple scratchpad memory to reduce off-chip memory access. To perform themultidimensional temporal tiling, an easy-to-use programming model is provided to flexibly explore different tiling options for large graphs. Experimental results show that compared against Nvidia P100 GPU, the performance and energy efficiency can be improved by$7.14\times $and$20.18\times $, respectively, on various GNNs, which validates both the versatility and energy efficiency of CAMBRICON-G.
Xinkai Song, Tian Zhi, Zhe Fan, Wei Li 0008, Xing Hu 0001, Zidong Du, Qi Guo 0001, Yunji Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Space-address decoupled scratchpad memory management for neural network accelerators
abstract
Summary Deep neural networks have been demonstrated to be useful in varieties of intelligent tasks, and various specialized NN accelerators have been proposed recently to improve the hardware efficiency, which are typically equipped with software‐managed scratchpad memory (SPM) for high performance and energy efficiency. However, traditional SPM management techniques cause memory fragmentation for NN accelerators, and thus lead to low utilization of precious SPM. The main reason is that traditional techniques are originally designed for managing fixed‐length registers rather than variable‐length memory blocks. In this article, we propose a novel SPM management approach for NN accelerators. The basic intuition is that NN computation/memory behaviors are predictable and relatively regular compared with traditional applications, and thus most information can be determined at compile time. In addition, by exploiting the variable‐length feature of SPM, we propose to divide the allocation process into two passes: the space assignment and the address assignment pass, which are simultaneously (and implicitly) performed in traditional one‐pass allocation techniques. Experimental results on the memory requests of a representative NN accelerator demonstrate that the proposed approach can significantly reduce the memory consumption by 30% at most compared with state‐of‐the‐art SPM management techniques, and the memory usage is only 2% larger than that of the theoretical optimal allocation.
Shiyan Sun, Xunyu Chen, Tian Zhi, Qi Guo 0001, Yunji Chen
Concurr. Comput. Pract. Exp.4
2020 DWM: A Decomposable Winograd Method for Convolution Acceleration
abstract
Winograd's minimal filtering algorithm has been widely used in Convolutional Neural Networks (CNNs) to reduce the number of multiplications for faster processing. However, it is only effective on convolutions with kernel size as 3x3 and stride as 1, because it suffers from significantly increased FLOPs and numerical accuracy problem for kernel size larger than 3x3 and fails on convolution with stride larger than 1. In this paper, we propose a novel Decomposable Winograd Method (DWM), which breaks through the limitation of original Winograd's minimal filtering algorithm to a wide and general convolutions. DWM decomposes kernels with large size or large stride to several small kernels with stride as 1 for further applying Winograd method, so that DWM can reduce the number of multiplications while keeping the numerical accuracy. It enables the fast exploring of larger kernel size and larger stride value in CNNs for high performance and accuracy and even the potential for new CNNs. Comparing against the original Winograd, the proposed DWM is able to support all kinds of convolutions with a speedup of ∼2, without affecting the numerical accuracy.
Xishan Zhang, Rui Zhang 0040, Tian Zhi, Deyuan He, Jiaming Guo, Chang Liu 0021, Qi Guo 0001, Zidong Du, Shaoli Liu, Tianshi Chen 0002, Yunji Chen
AAAI4
2020 Fixed-Point Back-Propagation Training
abstract
Recent emerged quantization technique (i.e., using low bit-width fixed-point data instead of high bit-width floating-point data) has been applied to inference of deep neural networks for fast and efficient execution. However, directly applying quantization in training can cause significant accuracy loss, thus remaining an open challenge. In this paper, we propose a novel training approach, which applies a layer-wise precision-adaptive quantization in deep neural networks. The new training approach leverages our key insight that the degradation of training accuracy is attributed to the dramatic change of data distribution. Therefore, by keeping the data distribution stable through a layer-wise precision-adaptive quantization, we are able to directly train deep neural networks using low bit-width fixed-point data and achieve guaranteed accuracy, without changing hyper parameters. Experimental results on a wide variety of network architectures (e.g., convolution and recurrent networks) and applications (e.g., image classification, object detection, segmentation and machine translation) show that the proposed approach can train these neural networks with negligible accuracy losses (-1.40%-1.3%, 0.02% on average), and speed up training by 252% on a state-of-the-art Intel CPU.
Xishan Zhang, Shaoli Liu, Rui Zhang 0040, Chang Liu 0021, Shiyi Zhou, Jiaming Guo, Qi Guo 0001, Zidong Du, Tian Zhi, Yunji Chen
CVPR10
2020 ALT: Optimizing Tensor Compilation in Deep Learning Compilers with Active Learning
abstract
Deep learning compilers serve as the central role of scheduling neural network execution. State-of-the-art method of tensor compilation in deep learning compilers requires a long time tuning, which greatly hinders the model's deployment. In this paper, we propose ALT, an active learning tuning method for tensor computation compilation. ALT leverages a sampling strategy based on active learning to find more informative samples to be labeled. The sampling strategy is performed by an active learning exploration module which mainly consists of an uncertainty predictor, which predicts uncertainty of unseen samples, and a score predictor which evaluates the current performance of the whole method. We design a novel ping-pang way of iteration between the score predictor and the uncertainty predictor. Experiments on real workloads show that ALT helps to achieve 1.93 × - 2.49 × time reduction to obtain the optimal schedule compared to state-of-the-art deep learning compilers. When tuning under the same time budget, the end to end inference time of a set of neural networks can be improved by 1.04× - 1.07×.
Tian Zhi, Zidong Du, Qi Guo 0001, Ninghui Sun, Yunji Chen
ICCD2
2020 Self-Aware Neural Network Systems: A Survey and New Perspective
abstract
Neural network (NN) processors are specially designed to handle deep learning tasks by utilizing multilayer artificial NNs. They have been demonstrated to be useful in broad application fields such as image recognition, speech processing, machine translation, and scientific computing. Meanwhile, innovative self-aware techniques, whereby a system can dynamically react based on continuously sensed information from the execution environment, have attracted attention from both academia and industry. Actually, various self-aware techniques have been applied to NN systems to significantly improve the computational speed and energy efficiency. This article surveys state-of-the-art self-aware NN systems (SaNNSs), which can be achieved at different layers, that is, the architectural layer, the physical layer, and the circuit layer. At the architectural layer, SaNNS can be characterized from a data-centric perspective where different data properties (i.e., data value, data precision, dataflow, and data distribution) are exploited. At the physical layer, various parameters of physical implementation are considered. At the circuit layer, different logics and devices can be used for high efficiency. In fact, the self-awareness of existing SaNNS is still in a preliminary form. We propose a comprehensive SaNNS from a new perspective, that is, the model layer, to exploit more opportunities for high efficiency. The proposed system is called as MinMaxNN, which features model switching and elastic sparsity based on monitored information from the execution environment. The model switching mechanism implies that models (i.e., min and max model) dynamically switch given different inputs for both efficiency and accuracy. The elastic sparsity mechanism indicates that the sparsity of NNs can be dynamically adjusted in each layer for efficiency. The experimental results show that compared with traditional SaNNS, MinMaxNN can achieve 5.64× and 19.66% performance improvement and energy reduction, respectively, without notable loss of accuracy and negative effects on developers' productivity.
Zidong Du, Qi Guo 0001, Yongwei Zhao 0001, Tian Zhi, Yunji Chen, Zhiwei Xu 0002
Proc. IEEE4
2020 Addressing Irregularity in Sparse Neural Networks Through a Cooperative Software/Hardware Approach
abstract
Neural networks have become the dominant algorithms rapidly as they achieve state-of-the-art performance in a broad range of applications such as image recognition, speech recognition, and natural language processing. However, neural networks keep moving toward deeper and larger architectures, posing a great challenge to hardware systems due to the huge amount of data and computations. Although sparsity has emerged as an effective solution for reducing the intensity of computation and memory accesses directly, irregularity caused by sparsity (including sparse synapses and neurons) prevents accelerators from completely leveraging the benefits, i.e., it also introduces costly indexing module in accelerators. In this article, we propose a cooperative software/hardware approach to address the irregularity of sparse neural networks efficiently. Initially, we observe the local convergence, namely larger weights tend to gather into small clusters during training. Based on that key observation, we propose a software-based coarse-grained pruning technique to reduce the irregularity of sparse synapses drastically. The coarse-grained pruning technique, together with local quantization, significantly reduces the size of indexes and improves the network compression ratio. We further design a multi-core hardware accelerator, Cambricon-SE, to address the remaining irregularity of sparse synapses and neurons efficiently. The novel accelerator have three key features: 1) selector modulesto filter unnecessary synapses and neurons, 2) compress/decompress modules for exploiting the sparsity in data transmission (which is rarely studied in previous work), and 3) a multi-core architecture with elevated throughput to meet the real-time processing requirement. Compared against a state-of-the-art sparse neural network accelerator, our accelerator is 1.20x and 2.72x better in terms of performance and energy efficiency, respectively. Moreover, for real-time video analysis tasks, Cambricon-SE can process 1080p video at the speed of 76.59 fps.
Tian Zhi, Xuda Zhou, Zidong Du, Qi Guo 0001, Shaoli Liu, Bingrui Wang, Yuanbo Wen 0001, Chao Wang 0003, Xuehai Zhou, Ling Li 0001, Tianshi Chen 0002, Ninghui Sun, Yunji Chen
IEEE Trans. Computers2
2020 Machine Learning Computers With Fractal von Neumann Architecture
abstract
Machine learning techniques are pervasive tools for emerging commercial applications and many dedicated machine learning computers on different scales have been deployed in embedded devices, servers, and data centers. Currently, most machine learning computer architectures still focus on optimizing performance and energy efficiency instead of programming productivity. However, with the fast development in silicon technology, programming productivity, including programming itself and software stack development, becomes the vital reason instead of performance and power efficiency that hinders the application of machine learning computers. In this article, we propose Cambricon-F, which is a series of homogeneous, sequential, multi-layer, layer-similar, and machine learning computers with same ISA. A Cambricon-F machine has a fractal von Neumann architecture to iteratively manage its components: it is with von Neumann architecture and its processing components (sub-nodes) are still Cambricon-F machines with von Neumann architecture and the same ISA. Since different Cambricon-F instances with different scales can share the same software stack on their common ISA, Cambricon-Fs can significantly improve the programming productivity. Moreover, we address four major challenges in Cambricon-F architecture design, which allow Cambricon-F to achieve a high efficiency. We implement two Cambricon-F instances at different scales, i.e., Cambricon-F100 and Cambricon-F1. Compared to GPU based machines (DGX-1 and 1080Ti), Cambricon-F instances achieve 2.82x, 5.14x better performance, 8.37x, 11.39x better efficiency on average, with 74.5, 93.8 percent smaller area costs, respectively. We further propose Cambricon-FR, which enhances the Cambricon-F machine learning computers to flexibly and efficiently support all the fractal operations with a reconfigurable fractal instruction set architecture. Compared to the Cambricon-F instances, Cambricon-FR machines achieve 1.96x, 2.49x better performance on average. Most importantly, Cambricon-FR computers are able to save the code length with a factor of 5.83, thus significantly improving the programming productivity.
Yongwei Zhao 0001, Zhe Fan, Zidong Du, Tian Zhi, Ling Li 0001, Qi Guo 0001, Shaoli Liu, Zhiwei Xu 0002, Tianshi Chen 0002, Yunji Chen
IEEE Trans. Computers4
2019 TDSNN: From Deep Neural Networks to Deep Spike Neural Networks with Temporal-Coding
abstract
Continuous-valued deep convolutional networks (DNNs) can be converted into accurate rate-coding based spike neural networks (SNNs). However, the substantial computational and energy costs, which is caused by multiple spikes, limit their use in mobile and embedded applications. And recent works have shown that the newly emerged temporal-coding based SNNs converted from DNNs can reduce the computational load effectively. In this paper, we propose a novel method to convert DNNs to temporal-coding SNNs, called TDSNN. Combined with the characteristic of the leaky integrate-andfire (LIF) neural model, we put forward a new coding principle Reverse Coding and design a novel Ticking Neuron mechanism. According to our evaluation, our proposed method achieves 42% total operations reduction on average in large networks comparing with DNNs with no more than 0.5% accuracy loss. The evaluation shows that TDSNN may prove to be one of the key enablers to make the adoption of SNNs widespread.
Lei Zhang 0008, Shengyuan Zhou, Tian Zhi, Zidong Du, Yunji Chen
AAAI3
2019 Partition and Scheduling Algorithms for Neural Network Accelerators
Xiaobing Chen, Shaohui Peng, Luyang Jin, Yimin Zhuang, Jin Song, Weijian Du, Shaoli Liu, Tian Zhi
APPT8
2019 ZhuQue: A Neural Network Programming Model Based on Labeled Data Layout
Weijian Du, Linyang Wu, Xiaobing Chen, Yimin Zhuang, Tian Zhi
APPT5
2019 Compiling Optimization for Neural Network Accelerators
Jin Song, Yimin Zhuang, Xiaobing Chen, Tian Zhi, Shaoli Liu
APPT4
2019 Deep Fusion: A Software Scheduling Method for Memory Access Optimization
Yimin Zhuang, Shaohui Peng, Xiaobing Chen, Shengyuan Zhou, Tian Zhi, Wei Li 0008, Shaoli Liu
NPC5
2018 Leveraging Subgraph Extraction for Performance Portable Programming Frameworks on DL Accelerators
Huiying Lan, Tian Zhi
NPC3
2017 TuNao: A High-Performance and Energy-Efficient Reconfigurable Accelerator for Graph Processing
abstract
Large-scale graph processing is now a crucial task of many commercial applications, and it is conventionally supported by general-purpose processors. These processors are designed to flexibly support highly diverse workloads with classic techniques such as on-chip cache and dynamic pipelining. Yet, it is difficult for the on-chip cache to exploit irregular data locality in large-scale graph processing, even though there are a few high-degree vertices that are frequently accessed in real-world graphs, it is not efficient to perform regular arithmetic operations via sophisticated dynamic pipelining. In short, general-purpose processors could not be the ideal platforms to graph processing. In this paper, we design a reconfigurable graph processing accelerator, with the purpose of providing an energy-efficient and flexible hardware platform for large-scale graph processing. This accelerator features two main components, i.e., the on-chip storage to exploit the data locality of graph processing, and the reconfigurable functional units to adapt to diversified operations in different graph processing tasks. On a total of 36 practical graph processing tasks, we demonstrate that, on average, our accelerator design achieves 1.58x and 25.56x better performance and energy efficiency, respectively, than the GPU baseline.
Jinhong Zhou, Shaoli Liu, Qi Guo 0001, Xuda Zhou, Tian Zhi, Dao-Fu Liu, Chao Wang 0003, Xuehai Zhou, Yunji Chen, Tianshi Chen 0002
CCGrid5
2017 A survey of neural network accelerators
Tian Zhi, Tianshi Chen 0002
Frontiers Comput. Sci.3