VLDB 2026 Research / reviewers in the wild / expert
Shaoli Liu
dblp:85/9414
· DBLP profile ↗
40ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 3 first-authorArtificial intelligence and machine learning · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
20 papers |
Hardware accelerators and domain-specific architectures · 69% Processor architecture and microarchitecture · 13% Interconnection networks and networks-on-chip · 11% | |
| Artificial intelligence
14 papers |
Efficient and distributed learning · 48% Image recognition and object detection · 17% Transfer learning and domain adaptation · 9% |
Topics — the 30 heaviest of 56, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
2.8 | 10 | 2020 | Addressing Irregularity in Sparse Neural Networks Through a Cooperative Software/Hardware Approach · IEEE Trans. Computers 2020 Addressing Sparsity in Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 An Instruction Set Architecture for Machine Learning · ACM Trans. Comput. Syst. 2018 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
2.1 | 7 | 2020 | ParaML: A Polyvalent Multicore Accelerator for Machine Learning · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 DWM: A Decomposable Winograd Method for Convolution Acceleration · AAAI 2020 An Instruction Set Architecture for Machine Learning · ACM Trans. Comput. Syst. 2018 |
Machine learning › Efficient and distributed learning
model compression |
1.8 | 4 | 2022 | Rethinking the Importance of Quantization Bias, Toward Full Low-Bit Training · IEEE Trans. Image Process. 2022 Fixed-Point Back-Propagation Training · CVPR 2020 DWM: A Decomposable Winograd Method for Convolution Acceleration · AAAI 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
sparse neural network acceleration |
1.4 | 4 | 2020 | Addressing Irregularity in Sparse Neural Networks Through a Cooperative Software/Hardware Approach · IEEE Trans. Computers 2020 Addressing Sparsity in Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 Cambricon-S: Addressing Irregularity in Sparse Neural Networks through A Cooperative Software/Hardware Approach · MICRO 2018 |
Machine learning › Efficient and distributed learning
low-precision training |
1.0 | 2 | 2022 | Rethinking the Importance of Quantization Bias, Toward Full Low-Bit Training · IEEE Trans. Image Process. 2022 Fixed-Point Back-Propagation Training · CVPR 2020 |
Machine learning › Efficient and distributed learning › model compression
quantization |
1.0 | 2 | 2022 | Rethinking the Importance of Quantization Bias, Toward Full Low-Bit Training · IEEE Trans. Image Process. 2022 Fixed-Point Back-Propagation Training · CVPR 2020 |
Computer vision › Image recognition and object detection
object detection |
1.0 | 2 | 2021 | Distilling Object Detectors with Feature Richness · NeurIPS 2021 Domain-Specific Suppression for Adaptive Object Detection · CVPR 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
convolution acceleration |
0.9 | 2 | 2021 | A Decomposable Winograd Method for N-D Convolution Acceleration in Video Analysis · Int. J. Comput. Vis. 2021 DWM: A Decomposable Winograd Method for Convolution Acceleration · AAAI 2020 |
Hardware accelerators and domain-specific architectures › many-core accelerator
multi-core accelerator |
0.9 | 2 | 2020 | ParaML: A Polyvalent Multicore Accelerator for Machine Learning · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 Addressing Irregularity in Sparse Neural Networks Through a Cooperative Software/Hardware Approach · IEEE Trans. Computers 2020 |
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
multi-view 3d object detection |
0.8 | 1 | 2024 | RecurrentBEV: A Long-Term Temporal Fusion Framework for Multi-view 3D Detection · ECCV (72) 2024 |
Natural language and speech › Language models and text generation › text generation
story generation |
0.8 | 1 | 2024 | Ex3: Automatic Novel Writing by Extracting, Excelsior and Expanding · ACL (1) 2024 |
Computer vision › Video understanding and tracking › temporal modeling
temporal fusion |
0.8 | 1 | 2024 | RecurrentBEV: A Long-Term Temporal Fusion Framework for Multi-view 3D Detection · ECCV (72) 2024 |
Processor architecture and microarchitecture
instruction set architecture |
0.7 | 3 | 2020 | An Instruction Set Architecture for Machine Learning · ACM Trans. Comput. Syst. 2018 Cambricon: An Instruction Set Architecture for Neural Networks · ISCA 2016 Machine Learning Computers With Fractal von Neumann Architecture · IEEE Trans. Computers 2020 |
Processor architecture and microarchitecture › instruction set architecture › instruction set customization
domain-specific ISA |
0.6 | 2 | 2018 | An Instruction Set Architecture for Machine Learning · ACM Trans. Comput. Syst. 2018 Cambricon: An Instruction Set Architecture for Neural Networks · ISCA 2016 |
Machine learning › Efficient and distributed learning › communication compression
gradient quantization |
0.6 | 1 | 2022 | Rethinking the Importance of Quantization Bias, Toward Full Low-Bit Training · IEEE Trans. Image Process. 2022 |
Computer vision › Image recognition and object detection › object detection
domain adaptive object detection |
0.5 | 1 | 2021 | Domain-Specific Suppression for Adaptive Object Detection · CVPR 2021 |
Machine learning › Transfer learning and domain adaptation
domain-invariant representation learning |
0.5 | 1 | 2021 | Domain-Specific Suppression for Adaptive Object Detection · CVPR 2021 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.5 | 1 | 2021 | Distilling Object Detectors with Feature Richness · NeurIPS 2021 |
Computer vision › Image recognition and object detection › object detection
knowledge distillation for detection |
0.5 | 1 | 2021 | Distilling Object Detectors with Feature Richness · NeurIPS 2021 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.5 | 1 | 2021 | Domain-Specific Suppression for Adaptive Object Detection · CVPR 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › convolution optimization
winograd convolution |
0.5 | 1 | 2021 | A Decomposable Winograd Method for N-D Convolution Acceleration in Video Analysis · Int. J. Comput. Vis. 2021 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.5 | 2 | 2020 | Cambricon-S: Addressing Irregularity in Sparse Neural Networks through A Cooperative Software/Hardware Approach · MICRO 2018 Addressing Irregularity in Sparse Neural Networks Through a Cooperative Software/Hardware Approach · IEEE Trans. Computers 2020 |
Memory systems
on-chip memory |
0.4 | 2 | 2017 | An Accelerator for High Efficient Vision Processing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 DaDianNao: A Neural Network Supercomputer · IEEE Trans. Computers 2017 |
Compilers and program optimization
machine learning compiler |
0.3 | 1 | 2018 | An Instruction Set Architecture for Machine Learning · ACM Trans. Comput. Syst. 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.3 | 1 | 2017 | An Accelerator for High Efficient Vision Processing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Interconnection networks and networks-on-chip › die-to-die interconnect
inter-chip communication |
0.3 | 1 | 2017 | DaDianNao: A Neural Network Supercomputer · IEEE Trans. Computers 2017 |
Interconnection networks and networks-on-chip › on-chip communication
network-on-chip communication |
0.3 | 1 | 2017 | Stealth-ACK: stealth transmissions of NoC acknowledgements · Sci. China Inf. Sci. 2017 |
Interconnection networks and networks-on-chip
optical interconnection networks |
0.3 | 1 | 2017 | DaDianNao: A Neural Network Supercomputer · IEEE Trans. Computers 2017 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.3 | 3 | 2021 | Domain-Specific Suppression for Adaptive Object Detection · CVPR 2021 A Small-Footprint Accelerator for Large-Scale Neural Networks · ACM Trans. Comput. Syst. 2015 DaDianNao: A Machine-Learning Supercomputer · MICRO 2014 |
Processor architecture and microarchitecture › instruction set architecture › instruction set style
load-store architecture |
0.2 | 1 | 2016 | Cambricon: An Instruction Set Architecture for Neural Networks · ISCA 2016 |
Methods — techniques the papers use, named apart from their topics
local quantization · 1.5coarse-grained pruning · 1.5indexing module · 1.0compressed synapse storage · 1.0winograd method · 1.0n-d convolution · 1.0fractal architecture · 0.8recurrent fusion · 0.8extract-expand pipeline · 0.8bird's-eye-view representation · 0.8LLM-based generation · 0.8stochastic rounding · 0.6adaptive piecewise quantization · 0.6gradient manipulation · 0.5domain-specific suppression · 0.5simulation · 0.5winograd minimal filtering · 0.4computational primitive analysis · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PE Loss: Perception-Enhanced Distortion-Oriented Loss for Image RestorationabstractImage restoration is the inverse problem of recovering high-quality images from knowledge of degraded images; it includes image super-resolution, image denoising, image deblurring, etc. The objective of image restoration methods is to minimize the error defined by the loss function between the network output image and the corresponding ground truth. Distortion-oriented loss functions are fundamental and commonly used in image restoration. However, these functions treat all pixels as equally important without differentiating between sharp and blurred edge areas, which does not match human visual perception. As a result, these methods can produce accurate but blurred images. To address this issue and achieve both accurate and perceptually satisfactory results, we propose a novel perception-enhanced distortion-oriented loss (PE loss) for image restoration, inspired by the Mach band effect. This effect demonstrates that sharp edges are perceived as having better quality than blurred edges by the human visual system. Our approach includes designing a blur factor map that detects blurred pixels and penalizes them by amplifying their error. The PE loss is a simple yet effective plug-and-play method, and we apply it to state-of-the-art networks. Extensive quantitative and qualitative experiments show that our method can restore images with sharp edges and high perceptual quality. Xishan Zhang, Shaoli Liu |
Comput. Vis. Media | 5 |
| 2026 | Generation-Augmented Intelligent Industrial Defect Detection Model Under Extremely Low-Sample ConditionsabstractDeep-learning-based defect detection models require large and diverse datasets, yet obtaining sufficient defect samples in industrial environments is challenging. Existing sample-generation-based methods alleviate data scarcity, but still depend on a certain amount of training data and struggle to enforce the strict structural specifications of industrial products, limiting their applicability. To address these issues, we propose a sample-generation-based defect detection model. The core of our approach is a defect generation network that employs multiscale progressive learning to extract multiresolution features and enables augmentation from a single sample. To satisfy structural constraints without suppressing diversity, we introduce a structural attention mechanism that guides the generative process rather than imposing direct mask-based restrictions. Additionally, we design a simple yet effective foreground—background reconstruction loss that better preserves structural details compared with conventional reconstruction losses. Our method requires only a single sample to initiate data augmentation. As a result, collecting a small number of representative defect samples is sufficient to significantly enhance detection performance, and the low data requirement allows broad applicability across diverse industrial defect detection tasks. Experimental results demonstrate that our method outperforms existing models, with improved sample quality and detection performance, achieving up to a 28% and 20% increase in [email protected] on the DeepPCB and NEU-DET datasets, respectively. Our method proves even more effective when the sample size is limited. Zehua Jian, Shaoli Liu, Jianhua Liu 0005, Jiachun Huang |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | TopoPIS: Topology-constrained pipe instance segmentation via adaptive curvature convolution
Shaoli Liu |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Ex3: Automatic Novel Writing by Extracting, Excelsior and ExpandingabstractHuang Lei, Jiaming Guo, Guanhua He, Xishan Zhang, Rui Zhang, Shaohui Peng, Shaoli Liu, Tianshi Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Huang Lei, Jiaming Guo, Guanhua He, Xishan Zhang, Rui Zhang 0040, Shaohui Peng, Shaoli Liu, Tianshi Chen 0002 |
ACL (1) | 7 |
| 2024 | RecurrentBEV: A Long-Term Temporal Fusion Framework for Multi-view 3D Detection
Xishan Zhang, Rui Zhang 0040, Guanhua He, Shaoli Liu |
ECCV (72) | 6 |
| 2024 | PCSGAN: A Perceptual Constrained Generative Model for Railway Defect Sample Expansion From a Single ImageabstractMany deep learning based railway defect detection methods have been proposed in recent years. They have greatly improved the efficiency and accuracy of defects detection. However, detection of railway defects remains challenging because of the limited number and types of defect samples, thus, general deep learning methods cannot be applied. In this paper, we designed a “Perceptually Constrained Single Image Generative Adversarial Network” (PCSGAN) to expand the number of railway defect image samples. PCSGAN uses a pyramidal structure to learn the internal features of a single image. In addition, we proposed a masking and a perceptual reconstruction loss mechanism to impose specific positional and structural constraints on the images. We tested the method using railway defects images and compared it to other single image generation models. The experiment results show that the images generated by PCSGAN take into account railway prior knowledge, generate railway structure which satisfied the constraints imposed by railway infrastructure designs, and also provide new information. High image realism and lowest Single Image FID were obtained, and the effectiveness of PCSGAN in the defect detection task were also validated. Sen He 0005, Zehua Jian, Shaoli Liu, Jianhua Liu 0005 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Latent Representation Self-Supervised Pose Network for Accurate Monocular Pipe Pose EstimationabstractAccurate pipe pose estimation plays a pivotal role in the automatic assembly of pipelines. Recently, data-driven deep neural networks have been proven capable of estimating pose. Nonetheless, a large number of labeled datasets are required during the training process. One effective solution is to estimate pose using self-supervised learning. However, existing algorithms are difficult to deal with textureless objects (like pipes), and they avoid the occlusion problem. To this end, in this article, we propose a latent representation self-supervised pose network (LSPN) for accurate monocular pipe pose estimation. We train our network with synthetic RGB (Red, Green, Blue) data, where only a few labeled samples are used to establish the latent pose space, whereas a large number of structured unlabeled samples are used to learn latent pose representation in self-supervised learning. Experiments demonstrate that LSPN achieves excellent performance on real data and is robust to different environments, such as illumination changes and self-occlusion. Shaoli Liu, Jianhua Liu 0005, Wenxiong Zhang |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Rethinking the Importance of Quantization Bias, Toward Full Low-Bit TrainingabstractQuantization is a promising technique to reduce the computation and storage costs of DNNs. Low-bit ( ≤ 8 bits) precision training remains an open problem due to the difficulty of gradient quantization. In this paper, we find two long-standing misunderstandings of the bias of gradient quantization noise. First, the large bias of gradient quantization noise, instead of the variance, is the key factor of training accuracy loss. Second, the widely used stochastic rounding cannot solve the training crash problem caused by the gradient quantization bias in practice. Moreover, we find that the asymmetric distribution of gradients causes a large bias of gradient quantization noise. Based on our findings, we propose a novel adaptive piecewise quantization method to effectively limit the bias of gradient quantization noise. Accordingly, we propose a new data format, Piecewise Fixed Point (PWF), to present data after quantization. We apply our method to different applications including image classification, machine translation, optical character recognition, and text classification. We achieve approximately 1.9 ∼ 3.5× speedup compared with full precision training with an accuracy loss of less than 0.5%. To the best of our knowledge, this is the first work to quantize gradients of all layers to 8 bits in both large-scale CNN and RNN training with negligible accuracy loss. Chang Liu 0021, Xishan Zhang, Rui Zhang 0040, Ling Li 0001, Shiyi Zhou, Zidong Du, Shaoli Liu, Tianshi Chen 0002 |
IEEE Trans. Image Process. | 9 |
| 2021 | Domain-Specific Suppression for Adaptive Object DetectionabstractDomain adaptation methods face performance degradation in object detection, as the complexity of tasks require more about the transferability of the model. We propose a new perspective on how CNN models gain the transferability, viewing the weights of a model as a series of motion patterns. The directions of weights, and the gradients, can be divided into domain-specific and domain-invariant parts, and the goal of domain adaptation is to concentrate on the domain-invariant direction while eliminating the disturbance from domain-specific one. Current UDA object detection methods view the two directions as a whole while optimizing, which will cause domain-invariant direction mismatch even if the output features are perfectly aligned. In this paper, we propose the domain-specific suppression, an exemplary and generalizable constraint to the original convolution gradients in backpropagation to detach the two parts of directions and suppress the domain-specific one. We further validate our theoretical analysis and methods on several domain adaptive object detection tasks, including weather, camera configuration, and synthetic to real-world adaptation. Our experiment results show significant advance over the state-of-the-art methods in the UDA object detection field, performing a promotion of 10.2 ∼ 12.2% mAP on all these domain adaptation scenarios. Rui Zhang 0040, Yangyang Xia, Xishan Zhang, Shaoli Liu |
CVPR | 7 |
| 2021 | Distilling Object Detectors with Feature RichnessabstractIn recent years, large-scale deep models have achieved great success, but the huge computational complexity and massive storage requirements make it a great challenge to deploy them in resource-limited devices. As a model compression and acceleration method, knowledge distillation effectively improves the performance of small models by transferring the dark knowledge from the teacher detector. However, most of the existing distillation-based detection methods mainly imitating features near bounding boxes, which suffer from two limitations. First, they ignore the beneficial features outside the bounding boxes. Second, these methods imitate some features which are mistakenly regarded as the background by the teacher detector. To address the above issues, we propose a novel Feature-Richness Score (FRS) method to choose important features that improve generalized detectability during distilling. The proposed method effectively retrieves the important features outside the bounding boxes and removes the detrimental features within the bounding boxes. Extensive experiments show that our methods achieve excellent performance on both anchor-based and anchor-free detectors. For example, RetinaNet with ResNet-50 achieves 39.7% in mAP on the COCO2017 dataset, which even surpasses the ResNet-101 based teacher detector 38.9% by 0.8%. Our implementation is available at https://github.com/duzhixing/FRS. Zhixing Du, Rui Zhang 0040, Xishan Zhang, Shaoli Liu, Tianshi Chen 0002, Yunji Chen |
NeurIPS | 5 |
| 2021 | A Decomposable Winograd Method for N-D Convolution Acceleration in Video Analysis
Rui Zhang 0040, Xishan Zhang, Xianzhuo Wang, Pengwei Jin, Shaoli Liu, Ling Li 0001, Yunji Chen |
Int. J. Comput. Vis. | 7 |
| 2020 | DWM: A Decomposable Winograd Method for Convolution AccelerationabstractWinograd's minimal filtering algorithm has been widely used in Convolutional Neural Networks (CNNs) to reduce the number of multiplications for faster processing. However, it is only effective on convolutions with kernel size as 3x3 and stride as 1, because it suffers from significantly increased FLOPs and numerical accuracy problem for kernel size larger than 3x3 and fails on convolution with stride larger than 1. In this paper, we propose a novel Decomposable Winograd Method (DWM), which breaks through the limitation of original Winograd's minimal filtering algorithm to a wide and general convolutions. DWM decomposes kernels with large size or large stride to several small kernels with stride as 1 for further applying Winograd method, so that DWM can reduce the number of multiplications while keeping the numerical accuracy. It enables the fast exploring of larger kernel size and larger stride value in CNNs for high performance and accuracy and even the potential for new CNNs. Comparing against the original Winograd, the proposed DWM is able to support all kinds of convolutions with a speedup of ∼2, without affecting the numerical accuracy. Xishan Zhang, Rui Zhang 0040, Tian Zhi, Deyuan He, Jiaming Guo, Chang Liu 0021, Qi Guo 0001, Zidong Du, Shaoli Liu, Tianshi Chen 0002, Yunji Chen |
AAAI | 10 |
| 2020 | Fixed-Point Back-Propagation TrainingabstractRecent emerged quantization technique (i.e., using low bit-width fixed-point data instead of high bit-width floating-point data) has been applied to inference of deep neural networks for fast and efficient execution. However, directly applying quantization in training can cause significant accuracy loss, thus remaining an open challenge. In this paper, we propose a novel training approach, which applies a layer-wise precision-adaptive quantization in deep neural networks. The new training approach leverages our key insight that the degradation of training accuracy is attributed to the dramatic change of data distribution. Therefore, by keeping the data distribution stable through a layer-wise precision-adaptive quantization, we are able to directly train deep neural networks using low bit-width fixed-point data and achieve guaranteed accuracy, without changing hyper parameters. Experimental results on a wide variety of network architectures (e.g., convolution and recurrent networks) and applications (e.g., image classification, object detection, segmentation and machine translation) show that the proposed approach can train these neural networks with negligible accuracy losses (-1.40%-1.3%, 0.02% on average), and speed up training by 252% on a state-of-the-art Intel CPU. Xishan Zhang, Shaoli Liu, Rui Zhang 0040, Chang Liu 0021, Shiyi Zhou, Jiaming Guo, Qi Guo 0001, Zidong Du, Tian Zhi, Yunji Chen |
CVPR | 2 |
| 2020 | Addressing Irregularity in Sparse Neural Networks Through a Cooperative Software/Hardware ApproachabstractNeural networks have become the dominant algorithms rapidly as they achieve state-of-the-art performance in a broad range of applications such as image recognition, speech recognition, and natural language processing. However, neural networks keep moving toward deeper and larger architectures, posing a great challenge to hardware systems due to the huge amount of data and computations. Although sparsity has emerged as an effective solution for reducing the intensity of computation and memory accesses directly, irregularity caused by sparsity (including sparse synapses and neurons) prevents accelerators from completely leveraging the benefits, i.e., it also introduces costly indexing module in accelerators. In this article, we propose a cooperative software/hardware approach to address the irregularity of sparse neural networks efficiently. Initially, we observe the local convergence, namely larger weights tend to gather into small clusters during training. Based on that key observation, we propose a software-based coarse-grained pruning technique to reduce the irregularity of sparse synapses drastically. The coarse-grained pruning technique, together with local quantization, significantly reduces the size of indexes and improves the network compression ratio. We further design a multi-core hardware accelerator, Cambricon-SE, to address the remaining irregularity of sparse synapses and neurons efficiently. The novel accelerator have three key features: 1) selector modulesto filter unnecessary synapses and neurons, 2) compress/decompress modules for exploiting the sparsity in data transmission (which is rarely studied in previous work), and 3) a multi-core architecture with elevated throughput to meet the real-time processing requirement. Compared against a state-of-the-art sparse neural network accelerator, our accelerator is 1.20x and 2.72x better in terms of performance and energy efficiency, respectively. Moreover, for real-time video analysis tasks, Cambricon-SE can process 1080p video at the speed of 76.59 fps. Tian Zhi, Xuda Zhou, Zidong Du, Qi Guo 0001, Shaoli Liu, Bingrui Wang, Yuanbo Wen 0001, Chao Wang 0003, Xuehai Zhou, Ling Li 0001, Tianshi Chen 0002, Ninghui Sun, Yunji Chen |
IEEE Trans. Computers | 6 |
| 2020 | Machine Learning Computers With Fractal von Neumann ArchitectureabstractMachine learning techniques are pervasive tools for emerging commercial applications and many dedicated machine learning computers on different scales have been deployed in embedded devices, servers, and data centers. Currently, most machine learning computer architectures still focus on optimizing performance and energy efficiency instead of programming productivity. However, with the fast development in silicon technology, programming productivity, including programming itself and software stack development, becomes the vital reason instead of performance and power efficiency that hinders the application of machine learning computers. In this article, we propose Cambricon-F, which is a series of homogeneous, sequential, multi-layer, layer-similar, and machine learning computers with same ISA. A Cambricon-F machine has a fractal von Neumann architecture to iteratively manage its components: it is with von Neumann architecture and its processing components (sub-nodes) are still Cambricon-F machines with von Neumann architecture and the same ISA. Since different Cambricon-F instances with different scales can share the same software stack on their common ISA, Cambricon-Fs can significantly improve the programming productivity. Moreover, we address four major challenges in Cambricon-F architecture design, which allow Cambricon-F to achieve a high efficiency. We implement two Cambricon-F instances at different scales, i.e., Cambricon-F100 and Cambricon-F1. Compared to GPU based machines (DGX-1 and 1080Ti), Cambricon-F instances achieve 2.82x, 5.14x better performance, 8.37x, 11.39x better efficiency on average, with 74.5, 93.8 percent smaller area costs, respectively. We further propose Cambricon-FR, which enhances the Cambricon-F machine learning computers to flexibly and efficiently support all the fractal operations with a reconfigurable fractal instruction set architecture. Compared to the Cambricon-F instances, Cambricon-FR machines achieve 1.96x, 2.49x better performance on average. Most importantly, Cambricon-FR computers are able to save the code length with a factor of 5.83, thus significantly improving the programming productivity. Yongwei Zhao 0001, Zhe Fan, Zidong Du, Tian Zhi, Ling Li 0001, Qi Guo 0001, Shaoli Liu, Zhiwei Xu 0002, Tianshi Chen 0002, Yunji Chen |
IEEE Trans. Computers | 7 |
| 2020 | ParaML: A Polyvalent Multicore Accelerator for Machine LearningabstractIn recent years, machine learning (ML) techniques are proven to be powerful tools in various emerging applications. Traditionally, ML techniques are processed on general-purpose CPUs and GPUs, but their energy efficiencies are limited due to their excessive support for flexibility. As an efficient alternative to CPUs/GPUs, hardware accelerators are still limited as they often accommodate only a single ML technique (family). However, different problems may require different ML techniques, which implies that such accelerators may achieve poor learning accuracy or even be ineffective. In this paper, we present a polyvalent accelerator architecture integrated with multiple processing cores, called ParaML, which accommodates ten representative ML techniques, including k-means, k-nearest neighbors (k-NN), naive Bayes (NB), support vector machine (SVM), linear regression (LR), classification tree (CT), deep neural network (DNN), learning vector quantization (LVQ), parzen window (PW), and principal component analysis (PCA). Benefited from our thorough analysis on computational primitives and locality properties of different ML techniques, the single-core ParaML can perform up to 1056 GOP/s (e.g., additions and multiplications) in an area of 3.51 mm2and consumes 596 mW only, estimated by ICC and PrimeTime PX with postsynthesis netlist, respectively. Compared with the NVIDIA K20M GPU (28-nm process), the single-core ParaML (65-nm process) is 1.21× faster, and can reduce the energy by 137.93×. We also compare the single-core ParaML with other accelerators. Compared with PRINS, single-core ParaML achieves 72.09× and 2.57× energy benefit for k-NN and k-means, respectively, and speeds up each query in k-NN by 44.76×. Compared with EIE, the single-core ParaML achieves 5.02× speedup and 4.97× energy benefit with 11.62× less area when evaluating with dense DNN. Compared with TPU, the single-core ParaML achieves 2.45× better power efficiency (5647 Gop/W versus 2300 Gop/W) with 321.36× less area. Compared to the single-core version, the 8-core ParaML will further improve the speedup up to 3.98× with an area of 13.44 mm2and a power of 2036 mW. Shengyuan Zhou, Qi Guo 0001, Zidong Du, Dao-Fu Liu, Tianshi Chen 0002, Ling Li 0001, Shaoli Liu, Jinhong Zhou, Olivier Temam, Xiaobing Feng 0002, Xuehai Zhou, Yunji Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2019 | Partition and Scheduling Algorithms for Neural Network Accelerators
Xiaobing Chen, Shaohui Peng, Luyang Jin, Yimin Zhuang, Jin Song, Weijian Du, Shaoli Liu, Tian Zhi |
APPT | 7 |
| 2019 | Compiling Optimization for Neural Network Accelerators
Jin Song, Yimin Zhuang, Xiaobing Chen, Tian Zhi, Shaoli Liu |
APPT | 5 |
| 2019 | Cambricon-F: machine learning computers with fractal von neumann architectureabstractMachine learning techniques are pervasive tools for emerging commercial applications and many dedicated machine learning computers on different scales have been deployed in embedded devices, servers, and data centers. Currently, most machine learning computer architectures still focus on optimizing performance and energy efficiency instead of programming productivity. However, with the fast development in silicon technology, programming productivity, including programming itself and software stack development, becomes the vital reason instead of performance and power efficiency that hinders the application of machine learning computers. Yongwei Zhao 0001, Zidong Du, Qi Guo 0001, Shaoli Liu, Ling Li 0001, Zhiwei Xu 0002, Tianshi Chen 0002, Yunji Chen |
ISCA | 4 |
| 2019 | Deep Fusion: A Software Scheduling Method for Memory Access Optimization
Yimin Zhuang, Shaohui Peng, Xiaobing Chen, Shengyuan Zhou, Tian Zhi, Wei Li 0008, Shaoli Liu |
NPC | 7 |
| 2019 | Addressing Sparsity in Deep Neural NetworksabstractNeural networks (NNs) have been demonstrated to be useful in a broad range of applications, such as image recognition, automatic translation, and advertisement recommendation. State-of-the-art NNs are known to be both computationally and memory intensive, due to the ever-increasing deep structure, i.e., multiple layers with massive neurons and connections (i.e., synapses). Sparse NNs have emerged as an effective solution to reduce the amount of computation and memory required. Though existing NN accelerators are able to efficiently process dense and regular networks, they cannot benefit from the reduction of synaptic weights. In this paper, we propose a novel accelerator, Cambricon-X, to exploit the sparsity and irregularity of NN models for increased efficiency. The proposed accelerator features a processing element (PE)-based architecture consisting of multiple PEs. An indexing module efficiently selects and transfers needed neurons to connected PEs with reduced bandwidth requirement, while each PE stores irregular and compressed synapses for local computation in an asynchronous fashion. With 16 PEs, our accelerator is able to achieve at most 544 GOP/s in a small form factor (6.38 mm2and 954 mW at 65 nm). Experimental results over a number of representative sparse networks show that our accelerator achieves, on average, $7.23\times$ speedup and $6.43\times$ energy saving against the state-of-the-art NN accelerator. We further investigate possibilities of leveraging activation sparsity and multi-issue controller, which improve the efficiency of Cambricon-X. To ease the burden of programmers, we also propose a high efficient library-based programming environment for our accelerator. Xuda Zhou, Zidong Du, Shijin Zhang, Lei Zhang 0008, Huiying Lan, Shaoli Liu, Ling Li 0001, Qi Guo 0001, Tianshi Chen 0002, Yunji Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2018 | Cambricon-S: Addressing Irregularity in Sparse Neural Networks through A Cooperative Software/Hardware ApproachabstractNeural networks have become the dominant algorithms rapidly as they achieve state-of-the-art performance in a broad range of applications such as image recognition, speech recognition and natural language processing. However, neural networks keep moving towards deeper and larger architectures, posing a great challenge to the huge amount of data and computations. Although sparsity has emerged as an effective solution for reducing the intensity of computation and memory accesses directly, irregularity caused by sparsity (including sparse synapses and neurons) prevents accelerators from completely leveraging the benefits; it also introduces costly indexing module in accelerators. In this paper, we propose a cooperative software/hardware approach to address the irregularity of sparse neural networks efficiently. Initially, we observe the local convergence, namely larger weights tend to gather into small clusters during training. Based on that key observation, we propose a software-based coarse-grained pruning technique to reduce the irregularity of sparse synapses drastically. The coarse-grained pruning technique, together with local quantization, significantly reduces the size of indexes and improves the network compression ratio. We further design a hardware accelerator, Cambricon-S, to address the remaining irregularity of sparse synapses and neurons efficiently. The novel accelerator features a selector module to filter unnecessary synapses and neurons. Compared with a state-of-the-art sparse neural network accelerator, our accelerator is 1.71× and 1.37× better in terms of performance and energy efficiency, respectively. Xuda Zhou, Zidong Du, Qi Guo 0001, Shaoli Liu, Chengsi Liu, Chao Wang 0003, Xuehai Zhou, Ling Li 0001, Tianshi Chen 0002, Yunji Chen |
MICRO | 4 |
| 2018 | BenchIP: Benchmarking Intelligence Processors
Jinhua Tao, Zidong Du, Qi Guo 0001, Huiying Lan, Lei Zhang 0008, Shengyuan Zhou, Lingjie Xu, Shan Tang, Allen Rush, Willian Chen, Shaoli Liu, Yunji Chen, Tianshi Chen 0002 |
J. Comput. Sci. Technol. | 13 |
| 2018 | An Instruction Set Architecture for Machine LearningabstractMachine Learning (ML) are a family of models for learning from the data to improve performance on a certain task. ML techniques, especially recent renewed neural networks (deep neural networks), have proven to be efficient for a broad range of applications. ML techniques are conventionally executed on general-purpose processors (such as CPU and GPGPU), which usually are not energy efficient, since they invest excessive hardware resources to flexibly support various workloads. Consequently, application-specific hardware accelerators have been proposed recently to improve energy efficiency. However, such accelerators were designed for a small set of ML techniques sharing similar computational patterns, and they adopt complex and informative instructions (control signals) directly corresponding to high-level functional blocks of an ML technique (such as layers in neural networks) or even an ML as a whole. Although straightforward and easy to implement for a limited set of similar ML techniques, the lack of agility in the instruction set prevents such accelerator designs from supporting a variety of different ML techniques with sufficient flexibility and efficiency. In this article, we first propose a novel domain-specific Instruction Set Architecture (ISA) for NN accelerators, called Cambricon, which is a load-store architecture that integrates scalar, vector, matrix, logical, data transfer, and control instructions, based on a comprehensive analysis of existing NN techniques. We then extend the application scope of Cambricon from NN to ML techniques. We also propose an assembly language, an assembler, and runtime to support programming with Cambricon, especially targeting large-scale ML problems. Our evaluation over a total of 16 representative yet distinct ML techniques have demonstrated that Cambricon exhibits strong descriptive capacity over a broad range of ML techniques and provides higher code density than general-purpose ISAs such as x86, MIPS, and GPGPU. Compared to the latest state-of-the-art NN accelerator design DaDianNao [7] (which can only accommodate three types of NN techniques), our Cambricon-based accelerator prototype implemented in TSMC 65nm technology incurs only negligible latency/power/area overheads, with a versatile coverage of 10 different NN benchmarks and 7 other ML benchmarks. Compared to the recent prevalent ML accelerator PuDianNao, our Cambricon-based accelerator is able to support all the ML techniques as well as the 10 NNs but with only approximate 5.1% performance loss. Yunji Chen, Huiying Lan, Zidong Du, Shaoli Liu, Jinhua Tao, Qi Guo 0001, Ling Li 0001, Yuan Xie 0001, Tianshi Chen 0002 |
ACM Trans. Comput. Syst. | 4 |
| 2017 | TuNao: A High-Performance and Energy-Efficient Reconfigurable Accelerator for Graph ProcessingabstractLarge-scale graph processing is now a crucial task of many commercial applications, and it is conventionally supported by general-purpose processors. These processors are designed to flexibly support highly diverse workloads with classic techniques such as on-chip cache and dynamic pipelining. Yet, it is difficult for the on-chip cache to exploit irregular data locality in large-scale graph processing, even though there are a few high-degree vertices that are frequently accessed in real-world graphs, it is not efficient to perform regular arithmetic operations via sophisticated dynamic pipelining. In short, general-purpose processors could not be the ideal platforms to graph processing. In this paper, we design a reconfigurable graph processing accelerator, with the purpose of providing an energy-efficient and flexible hardware platform for large-scale graph processing. This accelerator features two main components, i.e., the on-chip storage to exploit the data locality of graph processing, and the reconfigurable functional units to adapt to diversified operations in different graph processing tasks. On a total of 36 practical graph processing tasks, we demonstrate that, on average, our accelerator design achieves 1.58x and 25.56x better performance and energy efficiency, respectively, than the GPU baseline. Jinhong Zhou, Shaoli Liu, Qi Guo 0001, Xuda Zhou, Tian Zhi, Dao-Fu Liu, Chao Wang 0003, Xuehai Zhou, Yunji Chen, Tianshi Chen 0002 |
CCGrid | 2 |
| 2017 | Stealth-ACK: stealth transmissions of NoC acknowledgements
Jinhua Tao, Shaoli Liu, Tianshi Chen 0002, Rui Mao 0001 |
Sci. China Inf. Sci. | 3 |
| 2017 | DaDianNao: A Neural Network SupercomputerabstractMany companies are deploying services largely based on machine-learning algorithms for sophisticated processing of large amounts of data, either for consumers or industry. The state-of-the-art and most popular such machine-learning algorithms are Convolutional and Deep Neural Networks (CNNs and DNNs), which are known to be computationally and memory intensive. A number of neural network accelerators have been recently proposed which can offer high computational capacity/area ratio, but which remain hampered by memory accesses. However, unlike the memory wall faced by processors on general-purpose workloads, the CNNs and DNNs memory footprint, while large, is not beyond the capability of the on-chip storage of a multi-chip system. This property, combined with the CNN/DNN algorithmic characteristics, can lead to high internal bandwidth and low external communications, which can in turn enable high-degree parallelism at a reasonable area cost. In this article, we introduce a custom multi-chip machine-learning architecture along those lines, and evaluate performance by integrating electrical and optical inter-chip interconnects separately. We show that, on a subset of the largest known neural network layers, it is possible to achieve a speedup of 656.63× over a GPU, and reduce the energy by 184.05× on average for a 64-chip system. We implement the node down to the place and route at 28 nm, containing a combination of custom storage and computational units, with electrical inter-chip interconnects. Shaoli Liu, Ling Li 0001, Shijin Zhang, Tianshi Chen 0002, Zhiwei Xu 0002, Olivier Temam, Yunji Chen |
IEEE Trans. Computers | 2 |
| 2017 | An Accelerator for High Efficient Vision ProcessingabstractIn recent years, neural network accelerators have been shown to achieve both high energy efficiency and high performance for a broad application scope within the important category of recognition and mining applications. Still, both the energy efficiency and performance of such accelerators remain limited by memory accesses. In this paper, we focus on image applications, arguably the most important category among recognition and mining applications. The neural networks which are state-of-the-art for these applications are convolutional neural networks (CNNs), and they have an important property: weights are shared among many neurons, considerably reducing the neural network memory footprint. This property allows to entirely map a CNN within an SRAM, eliminating all DRAM accesses for weights. By further hoisting this accelerator next to the image sensor, it is possible to eliminate all remaining DRAM accesses, i.e., for inputs and outputs. In this paper, we propose such a CNN accelerator, placed next to a CMOS or CCD sensor. The absence of DRAM accesses combined with a careful exploitation of the specific data access patterns within CNNs allows us to design an accelerator which is highly energy-efficient. We present a single-core implementation down to the layout at 65 nm, with a modest footprint of 5.94mm$^{\boldsymbol {2}}$and consuming only 336mW, but still about$\boldsymbol {30\times }$faster than high-end GPUs. For visual processing with higher resolution and frame-rate requirements, we further present a multicore implementation with elevated performance. Zidong Du, Shaoli Liu, Robert Fasthuber, Tianshi Chen 0002, Paolo Ienne, Ling Li 0001, Qi Guo 0001, Xiaobing Feng 0002, Yunji Chen, Olivier Temam |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Cambricon: An Instruction Set Architecture for Neural NetworksabstractNeural Networks (NN) are a family of models for a broad range of emerging machine learning and pattern recondition applications. NN techniques are conventionally executed on general-purpose processors (such as CPU and GPGPU), which are usually not energy-efficient since they invest excessive hardware resources to flexibly support various workloads. Consequently, application-specific hardware accelerators for neural networks have been proposed recently to improve the energy-efficiency. However, such accelerators were designed for a small set of NN techniques sharing similar computational patterns, and they adopt complex and informative instructions (control signals) directly corresponding to high-level functional blocks of an NN (such as layers), or even an NN as a whole. Although straightforward and easy-to-implement for a limited set of similar NN techniques, the lack of agility in the instruction set prevents such accelerator designs from supporting a variety of different NN techniques with sufficient flexibility and efficiency. In this paper, we propose a novel domain-specific Instruction Set Architecture (ISA) for NN accelerators, called Cambricon, which is a load-store architecture that integrates scalar, vector, matrix, logical, data transfer, and control instructions, based on a comprehensive analysis of existing NN techniques. Our evaluation over a total of ten representative yet distinct NN techniques have demonstrated that Cambricon exhibits strong descriptive capacity over a broad range of NN techniques, and provides higher code density than general-purpose ISAs such as ×86, MIPS, and GPGPU. Compared to the latest state-of-the-art NN accelerator design DaDianNao [5] (which can only accommodate 3 types of NN techniques), our Cambricon-based accelerator prototype implemented in TSMC 65nm technology incurs only negligible latency/power/area overheads, with a versatile coverage of 10 different NN benchmarks. Shaoli Liu, Zidong Du, Jinhua Tao, Yuan Xie 0001, Yunji Chen, Tianshi Chen 0002 |
ISCA | 1 |
| 2016 | Cambricon-X: An accelerator for sparse neural networksabstractNeural networks (NNs) have been demonstrated to be useful in a broad range of applications such as image recognition, automatic translation and advertisement recommendation. State-of-the-art NNs are known to be both computationally and memory intensive, due to the ever-increasing deep structure, i.e., multiple layers with massive neurons and connections (i.e., synapses). Sparse neural networks have emerged as an effective solution to reduce the amount of computation and memory required. Though existing NN accelerators are able to efficiently process dense and regular networks, they cannot benefit from the reduction of synaptic weights. In this paper, we propose a novel accelerator, Cambricon-X, to exploit the sparsity and irregularity of NN models for increased efficiency. The proposed accelerator features a PE-based architecture consisting of multiple Processing Elements (PE). An Indexing Module (IM) efficiently selects and transfers needed neurons to connected PEs with reduced bandwidth requirement, while each PE stores irregular and compressed synapses for local computation in an asynchronous fashion. With 16 PEs, our accelerator is able to achieve at most 544 GOP/s in a small form factor (6.38 mm2and 954 mW at 65 nm). Experimental results over a number of representative sparse networks show that our accelerator achieves, on average, 7.23x speedup and 6.43x energy saving against the state-of-the-art NN accelerator. Shijin Zhang, Zidong Du, Lei Zhang 0008, Huiying Lan, Shaoli Liu, Ling Li 0001, Qi Guo 0001, Tianshi Chen 0002, Yunji Chen |
MICRO | 5 |
| 2016 | IMR: High-Performance Low-Cost Multi-Ring NoCsabstractA ring topology is a common solution of network-on-chip (NoC) in industry, but is frequently criticized to have poor scalability. In this paper, we present a novel type of multi-ring NoC called isolated multi-ring (IMR), which can even support chip multiprocessors (CMPs) with 1,024 cores. In IMR, any pair of cores are connected via at least one isolated ring, so that each packet can reach the destination without transferring from one ring to another. Therefore, IMR no longer needs expensive routers as mesh, which not only enhances the network performance but also reduces hardware overheads. We utilize simulated evolution to design optimized IMR topologies. We compare these IMR topologies against nine representative NoCs (e.g., traditional mesh, multi mesh, low-cost mesh, Express-virtual-channels mesh (EVC), torus ring, and hierarchical ring). We observe from experiments that IMR significantly outperforms its competitors in both saturation throughput and latency across all scenarios considered. For example, in a 16 × 16 CMP, IMR improves the saturation throughput of a state-of-the-art mesh (EVC) by 265.29 percent on average, and reduces the average packet latency on SPLASH-2 application traces by 71.58 percent, while consuming 5.08 percent less area and 9.76 percent less power. In a 32 × 32 CMP, IMR averagely improves the saturation throughput of EVC by 191.58 percent, and averagely reduces the packet latency on SPLASH-2 application traces by 23.09 percent, while consuming 2.86 percent less area and 10.81 percent less power. Shaoli Liu, Tianshi Chen 0002, Ling Li 0001, Xiaoxue Feng, Zhiwei Xu 0002, Haibo Chen 0001, Fred Chong, Yunji Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Stable Matching Scheduler for Single-ISA Heterogeneous Multi-core Processors
Lei Wang 0015, Shaoli Liu, Longbing Zhang, Junhua Xiao |
APPT | 2 |
| 2015 | PuDianNao: A Polyvalent Machine Learning AcceleratorabstractMachine Learning (ML) techniques are pervasive tools in various emerging commercial applications, but have to be accommodated by powerful computer systems to process very large data. Although general-purpose CPUs and GPUs have provided straightforward solutions, their energy-efficiencies are limited due to their excessive supports for flexibility. Hardware accelerators may achieve better energy-efficiencies, but each accelerator often accommodates only a single ML technique (family). According to the famous No-Free-Lunch theorem in the ML domain, however, an ML technique performs well on a dataset may perform poorly on another dataset, which implies that such accelerator may sometimes lead to poor learning accuracy. Even if regardless of the learning accuracy, such accelerator can still become inapplicable simply because the concrete ML task is altered, or the user chooses another ML technique. Dao-Fu Liu, Tianshi Chen 0002, Shaoli Liu, Jinhong Zhou, Shengyuan Zhou, Olivier Temam, Xiaobing Feng 0002, Xuehai Zhou, Yunji Chen |
ASPLOS | 3 |
| 2015 | A Small-Footprint Accelerator for Large-Scale Neural NetworksabstractMachine-learning tasks are becoming pervasive in a broad range of domains, and in a broad range of systems (from embedded systems to data centers). At the same time, a small set of machine-learning algorithms (especially Convolutional and Deep Neural Networks, i.e., CNNs and DNNs) are proving to be state-of-the-art across many applications. As architectures evolve toward heterogeneous multicores composed of a mix of cores and accelerators, a machine-learning accelerator can achieve the rare combination of efficiency (due to the small number of target algorithms) and broad application scope. Until now, most machine-learning accelerator designs have been focusing on efficiently implementing the computational part of the algorithms. However, recent state-of-the-art CNNs and DNNs are characterized by their large size. In this study, we design an accelerator for large-scale CNNs and DNNs, with a special emphasis on the impact of memory on accelerator design, performance, and energy. We show that it is possible to design an accelerator with a high throughput, capable of performing 452 GOP/s (key NN operations such as synaptic weight multiplications and neurons outputs additions) in a small footprint of 3.02mm2 and 485mW; compared to a 128-bit 2GHz SIMD processor, the accelerator is 117.87 × faster, and it can reduce the total energy by 21.08 ×. The accelerator characteristics are obtained after layout at 65nm. Such a high throughput in a small footprint can open up the usage of state-of-the-art machine-learning algorithms in a broad set of systems and for a broad set of applications. Tianshi Chen 0002, Shijin Zhang, Shaoli Liu, Zidong Du, Dongsheng Wang 0002, Chengyong Wu, Ninghui Sun, Yunji Chen, Olivier Temam |
ACM Trans. Comput. Syst. | 3 |
| 2015 | FreeRider: Non-Local Adaptive Network-on-Chip Routing with Packet-Carried Propagation of Congestion InformationabstractNon-local adaptive routing techniques, which utilize statuses of both local and distant links to make routing decisions, have recently been shown to be effective solutions for promoting the performance of Network-on-Chip (NoC). The essence of non-local adaptive routing was an additional network dedicated to propagate congestion information of distant links on the NoC. While the dedicated Congestion Propagation Network (CPN) helps routers to make promising routing decisions, it incurs additional wiring and power costs and becomes an unnecessary decoration when the load of NoC is light. Moreover, the CPN has to be extended if one would utilize more sophisticated congestion information to enhance the performance of NoC, bringing in even larger wiring and power costs. This paper proposes an innovative non-local adaptive routing technique called FreeRider, which does not use a dedicated CPN but instead leverages free bits in head flits of existing packets to carry and propagate rich congestion information without introducing additional wires or flits. In order to balance the network load, FreeRider adopts a novel three-stage strategy of output link selection, which adequately utilizes the propagated information to make routing decisions. Experimental results on both synthetic traffic patterns and application traces show that FreeRider achieves better throughput, shorter latency, and smaller power consumption than a state-of-the-art adaptive routing technique with dedicated CPN. Shaoli Liu, Tianshi Chen 0002, Ling Li 0001, Xi Li 0003, Mingzhe Zhang 0005, Chao Wang 0003, Haibo Meng, Xuehai Zhou, Yunji Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | DaDianNao: A Machine-Learning SupercomputerabstractMany companies are deploying services, either for consumers or industry, which are largely based on machine-learning algorithms for sophisticated processing of large amounts of data. The state-of-the-art and most popular such machine-learning algorithms are Convolutional and Deep Neural Networks (CNNs and DNNs), which are known to be both computationally and memory intensive. A number of neural network accelerators have been recently proposed which can offer high computational capacity/area ratio, but which remain hampered by memory accesses. However, unlike the memory wall faced by processors on general-purpose workloads, the CNNs and DNNs memory footprint, while large, is not beyond the capability of the on chip storage of a multi-chip system. This property, combined with the CNN/DNN algorithmic characteristics, can lead to high internal bandwidth and low external communications, which can in turn enable high-degree parallelism at a reasonable area cost. In this article, we introduce a custom multi-chip machine-learning architecture along those lines. We show that, on a subset of the largest known neural network layers, it is possible to achieve a speedup of 450.65x over a GPU, and reduce the energy by 150.31x on average for a 64-chip system. We implement the node down to the place and route at 28nm, containing a combination of custom storage and computational units, with industry-grade interconnects. Yunji Chen, Shaoli Liu, Shijin Zhang, Liqiang He, Ling Li 0001, Tianshi Chen 0002, Zhiwei Xu 0002, Ninghui Sun, Olivier Temam |
MICRO | 3 |
| 2014 | Archipelago: A Floorplan Optimized for Concurrent Multiple Applications on Network-on-ChipabstractIn practice, a chip multiprocessor (CMP) with many nodes often runs multiple applications simultaneously and nodes allocated to different applications seldom communicate. Leveraging the non-uniformity of communication, it is possible to optimize the Network-on-Chip (NoC) performance for the case of single chip running multiple applications through reducing the physical links lengths between nodes which tend to be allocated to the same application. In this paper, we propose a novel, nearly zero-cost floor plan method called Archipelago which rearranges the lengths of physical links by flipping hard IP of nodes. With communication monitoring and dynamic thread scheduling, Archipelago significantly reduces the average latency of NoC for concurrent multiple small- and medium-size applications. Experimental results show that Archipelago averagely reduces the packet latency by 19.10% and maximum to 24.4% for three-node applications on 4 × 4 mesh NoC. Shaoli Liu, Lei Wang 0015, Longbing Zhang, Meng Wen |
NAS | 2 |
| 2014 | Auxiliary stream for optimizing memory access of video decoders
Shaoli Liu, Ling Li 0001, Yunji Chen, Weiwu Hu |
Sci. China Inf. Sci. | 1 |
| 2013 | Motion Estimation Without Integer-Pel SearchabstractThe typical motion estimation (ME) consists of three main steps, including spatial-temporal prediction, integer-pel search, and fractional-pel search. The integer-pel search, which seeks the best matched integer-pel position within a search window, is considered to be crucial for video encoding. It occupies over 50% of the overall encoding time (when adopting the full search scheme) for software encoders, and introduces remarkable area cost, memory traffic, and power consumption to hardware encoders. In this paper, we find that video sequences (especially high-resolution videos) can often be encoded effectively and efficiently even without integer-pel search. Such counter-intuitive phenomenon is not only because that spatial-temporal prediction and fractional-pel search are accurate enough for the ME of many blocks. In fact, we observe that when the predicted motion vector is biased from the optimal motion vector (mainly for boundary blocks of irregularly moving objects), it is also hard for integer-pel search to reduce the final rate-distortion cost: the deviation of reference position could be alleviated with the fractional-pel interpolation and rate-distortion optimization techniques (e.g., adaptive macroblock mode). Considering the decreasing proportion of boundary blocks caused by the increasing resolution of videos, integer-pel search may be rather cost-ineffective in the era of high-resolution. Experimental results on 36 typical sequences of different resolutions encoded with x264, which is a widely-used video encoder, comply with our analysis well. For 1080p sequences, removing the integer-pel search saves 57.9% of the overall H.264 encoding time on average (compared to the original x264 with full integer-pel search using default parameters), while the resultant performance loss is negligible: the bit-rate is increased by only 0.18%, while the peak signal-to-noise ratio is decreased by only 0.01 dB per frame averagely. Ling Li 0001, Shaoli Liu, Yunji Chen, Tianshi Chen 0002 |
IEEE Trans. Image Process. | 2 |
| 2011 | Video Encoding without Integer-Pel Motion EstimationabstractMotion estimation (ME) consists of three main steps, including spatial-temporal prediction, integer-pel ME and fractional-pel ME. However, we find that video sequences (especially high resolution sequences) can be encoded efficiently even without integer-pel ME. Shaoli Liu, Ling Li 0001, Yunji Chen, Tianshi Chen 0002 |
DCC | 1 |