EDBT 2026 Demo / reviewers in the wild / expert
Bosheng Liu
dblp:124/4033
· DBLP profile ↗
18ranked-venue papers
8as first author
10since 2021 · last 2024
0000-0001-7123-7425ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Quantization-aware Optimization Approach for CNNs Inference on CPUsabstractData movements through the memory hierarchy are a fundamental bottleneck in the majority of convolutional neural network (CNN) deployments on CPUs. Loop-level optimization and hybrid bitwidth quantization are two representative optimization approaches for memory access reduction. However, they were carried out independently because of the significantly increased complexity of design space exploration. We present QAOpt, a quantization-aware optimization approach that can reduce the high complexity when combining both for CNN deployments on CPUs. We develop a bitwidth-sensitive quantization strategy that can perform the trade-off between model accuracy and data movements when deploying both loop-level optimization and mixed precision quantization. Also, we provide a quantization-aware pruning process that can reduce the design space for high efficiency. Evaluation results demonstrate that our work can achieve better energy efficiency under acceptable accuracy loss. Jiasong Chen, Zeming Xie, Weipeng Liang, Bosheng Liu, Xin Zheng 0001, Jigang Wu, Xiaoming Xiong |
ASPDAC | 4 |
| 2024 | Accelerating Frequency-domain Convolutional Neural Networks Inference using FPGAsabstractLow-end field programmable gate arrays (FPGAs) are difficult to deploy typical convolutional neural networks (C- NNs) owing to the limited hardware resources and the increasing model computational complexity. Fast Fourier transform (FFT) is a promising solution for saving both computation and memory footprint by convolving in the frequency domain. However, few FPGA accelerators can take full advantage at the computation level, because of the distinct element-wise complex calculation in the frequency domain. In this work, we present an FPGA-based 8-bit inference accelerator (called FAF) that packs frequency-domain calculations into digital signal processing (DSP) blocks to fully utilize DSPs for performance boost. We then provide a mapping dataflow to maximize the reduction of redundant packing operations by frequency-domain data reuse. Evaluations based on representative CNN benchmarks show that our work can achieve 1.5-6.9× better power efficiency compared with representative FPGA baselines. Bosheng Liu, Yongqi Xu, Jigang Wu, Xiaoming Chen 0003, Peng Liu 0045, Qingguo Zhou, Yinhe Han 0001 |
ISCAS | 2 |
| 2024 | A Complementary Resistive Switch-Based Balanced Ternary LogicabstractMemristors offer advantages in terms of high speed, high integration density, and non-volatility, making them a promising option for efficient logic applications. Recent works have explored the design methodology for ternary logic in memristor-based computing-in-memory (CIM) systems. However, existing methods require a large number of devices and are susceptible to noise interference. To address these issues, this work proposes a reliable in-memory computing paradigm for balanced ternary logic based on complementary resistive switch (CRS), which can be considered as two anti-serially connected memristors. Six balanced ternary logic gates are designed based on the proposed method, which support parallel operations when integrated into the CRS crossbar array. To demonstrate the efficiency of the proposed method, a 1-tri full adder is designed by using the proposed logic gates. The feasibility of the design is verified by Cadence Virtuoso using the Voltage Threshold Adaptive Memristor (VTEAM) model. The Monte Carlo simulation of the full adder verifies the reliability of the proposed method. Compared to existing methods, both the operation steps and area overhead are reduced using the proposed approach. Zhijian Peng, Peng Liu 0045, Lian Yao, Zhiqiang You, Bosheng Liu, Jigang Wu |
ITC-Asia | 5 |
| 2024 | Share-Aware Joint Model Deployment and Task Offloading for Multi-Task InferenceabstractIn vehicular edge computing, efficient strategies for model deployment and task offloading offer tremendous potential to reduce response time for machine learning inference. However, existing works do not pay much attention to that there are shared structures among different types of inference tasks. This limits the improvement in response time. This paper aims to fill this gap by investigating a share-aware joint model deployment and task offloading problem for multi-task inference in vehicular edge computing. We formulate the problem with an objective to minimize the total response time of all inference requests, under constraints of per task response time, per roadside unit storage capacity, etc. We prove that the formulated problem is NP-hard. To solve the problem, a time period aware algorithm, called TPA, is proposed with guaranteed approximation ratio. In TPA, an iterative approach is designed to solve the problem of maximizing system throughput during a certain time period. Then, the certain time period approximates to the minimum time period of completing all requests. The algorithms are evaluated in the environment comprising two CPUs, two GPUs, state-of-the-art multi-task learning models and the dataset of Google cluster-usage trace. Simulation results derived from this environment show that, the proposed TPA outperforms the state-of-the-art methods for all cases, in terms of the total response time of all requests. For example, TPA can significantly reduce the total response time by at least$73.72\%$for different numbers of RSUs considered, compared with state-of-the-art methods. Yalan Wu, Jigang Wu, Long Chen 0006, Bosheng Liu, Mianyang Yao, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Accelerating Convolutional Neural Networks in Frequency Domain via Kernel-Sharing ApproachabstractConvolutional neural networks (CNNs) are typically computationally heavy. Fast algorithms such as fast Fourier transforms (FFTs), are promising in significantly reducing computation complexity by replacing convolutions with frequency-domain element-wise multiplication. However, the increased high memory access overhead of complex weights counteracts the computing benefit, because frequency-domain convolutions not only pad weights to the same size as input maps, but also have no sharable complex kernel weights. In this work, we propose an FFT-based kernel-sharing technique called FS-Conv to reduce memory access. Based on FS-Conv, we derive the sharable complex weights in frequency-domain convolutions, which has never been solved. FS-Conv includes a hybrid padding approach, which utilizes the inherent periodic characteristic of FFT transformation to provide sharable complex weights for different blocks of complex input maps. We in addition build a frequency-domain inference accelerator (called Yixin) that can utilize the sharable complex weights for CNN accelerations. Evaluation results demonstrate the significant performance and energy efficiency benefits compared with the state-of-the-art baseline. Bosheng Liu, Hongyi Liang, Jigang Wu, Xiaoming Chen 0003, Peng Liu 0045, Yinhe Han 0001 |
ASP-DAC | 1 |
| 2023 | Two-Level Scheduling Algorithms for Deep Neural Network Inference in Vehicular NetworksabstractIn vehicular networks, task scheduling at the microarchitecture-level and network-level offers tremendous potential to improve the quality of computing services for deep neural network (DNN) inference. However, existing task scheduling works only focus on either one of the two levels, which results in inefficient utilization of computing resources. This paper aims to fill this gap by formulating a two-level scheduling problem for DNN inference tasks in a vehicular network, with an objective of minimizing total weighted sum of response time and energy consumption for all tasks under the following constraints: per task response time, per vehicle energy consumption, per vehicle storage capacity. We first formulate the problem and prove that it is NP-hard. A group transformation based algorithm, called GTA, is proposed. GTA makes scheduling decisions at the network-level using the group transformation based approach, and at the microarchitecture-level using a greedy strategy. In addition, an algorithm, denoted as DRL, is proposed to decrease total weighted sum of response time and energy consumption for all tasks. DRL trains two models with deep reinforcement learning to achieve two-level scheduling. The proposed algorithms are evaluated on a platform consisting of a desktop, Raspberry Pi, Eyeriss, OSM, SUMO, NS-3. Simulation results show that DRL outperforms the state-of-the-art methods for all cases, while the proposed GTA outperforms the state-of-the-art methods for most cases, in terms of total weighted sum of response time and energy consumption. Compared with four baseline algorithms, GTA and DRL reduce the total weighted sum of response time and energy consumption by 41.49% and 62.38%, on average respectively, for different numbers of tasks. Yalan Wu, Jigang Wu, Mianyang Yao, Bosheng Liu, Long Chen 0006, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Frequency-Domain Inference Acceleration for Convolutional Neural Networks Using ReRAMsabstractConvolutional neural networks (CNNs) (including 2D and 3D convolutions) are popular in video analysis tasks such as action recognition and activity understanding. Fast algorithms such as fast Fourier transforms (FFTs) are promising in significantly reducing computation complexity by transforming convolution into frequency domain. In frequency space, conventional spatial convolutions are replaced with simpler element-wise complex multiplications. Conventional application-specific-integrated-circuit (ASIC) based frequency-domain accelerators can achieve effective performance boost but come at the cost of significant energy consumption, owing to the hierarchical memory organization. We propose a frequency-domain resistive random access memory (ReRAM) based inference accelerator called FDA that can process element-wise complex multiplication in memory for both 2D and 3D CNNs. Each ReRAM-based frequency-domain process element (PE) with two ReRAM cells can perform an element-wise complex multiplication in two continuous execution cycles. We then provide a flexible dataflow to alleviate the redundant data movements by frequency-domain data reuse and inherent symmetrical characteristic for both 2D and 3D convolutions. Evaluation results based on representative both 2D and 3D CNN benchmarks demonstrate that FDA outperforms state-of-the-art baselines with better performance and energy efficiency. Bosheng Liu, Zhuoshen Jiang, Yalan Wu, Jigang Wu, Xiaoming Chen 0003, Peng Liu 0045, Qingguo Zhou, Yinhe Han 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Search-Free Inference Acceleration for Sparse Convolutional Neural NetworksabstractSparse convolution neural networks (CNNs) are promising in reducing both memory usage and computational complexity while still preserving high inference accuracy. State-of-the-art sparse CNN accelerators can deliver high throughput by skipping zero weights and/or activations. To operate on only nonzero weights and activations, sparse accelerators typically search pairs of nonzero weights and activations for multiplication-accumulation (MAC) operations. However, the conventional search operation results in a severe limitation in the processing element (PE) array scale because of the enormous demands of internal interconnection and memory bandwidth. In this article, we first provide a design principle to free the search process of sparse CNN accelerations. Specifically, the indexes of the static compressed weights access the dynamic activations directly to avoid the search process for MAC operations. We then develop two search-free inference accelerators, called Swan and Swan-flexible, for sparse CNN accelerations. Swan supports search-free sparse convolution accelerations for interconnection and bandwidth saving. Compared with Swan, Swan-flexible not only has the search-free capability but also comprises a configurable architecture for optimum throughput. We formulate a mathematical optimization problem by combining the configurable characterization with the compressive dataflow to optimize the overall throughput. Evaluations based on a place-and-route process show that the proposed designs, in a compact factor of 4096 PEs, achieve 1.5–$2.7\times $higher speedup and 6.0–$13.6\times $better energy efficiency than representative accelerator baselines with the same PE array scale. Bosheng Liu, Xiaoming Chen 0003, Yinhe Han 0001, Jigang Wu, Liang Chang 0003, Peng Liu 0045 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | F3D: Accelerating 3D Convolutional Neural Networks in Frequency Space Using ReRAMabstract3D convolutional neural networks (CNNs) are widely deployed in video analysis. Fast algorithms such as fast Fourier transforms (FFTs) are gaining popularity in reducing computation complexity for their superior capability of replacing convolutions with simpler element-wise multiplications. Conventional frequency-domain dedicated accelerators employ memory hierarchy organization for high throughput but at the expensive costs of a significant amount of data movements and energy consumptions. This paper presents F3D, a processingin-memory frequency-domain accelerator using resistive random access memory (ReRAM). F3D supports frequency-domain complex number multiplications directly in ReRAM-based crossbar architecture. We alleviate the overheads of redundant data movements in ReRAM-based complex number multiplications by data reuse and the inherent symmetry of inputs in the frequency space. Evaluation results demonstrate that F3D outperforms state-of-the-art accelerators with significant improvements in performance and energy efficiency. Bosheng Liu, Zhuoshen Jiang, Jigang Wu, Xiaoming Chen 0003, Yinhe Han 0001, Peng Liu 0045 |
DAC | 1 |
| 2021 | Fault Modeling and Efficient Testing of Memristor-Based MemoryabstractMemristor-based memory technology is one of the emerging memory technologies, which is a potential candidate to replace traditional memories. Efficient test solutions are required to enable the quality and reliability of such products. In previous works, fault models are caused by open, short and bridge defects and parametric variations during the fabrication. However, these fault models cannot describe the bridge defects that cause the state of the faulty cell to an undefined state. In this paper, we analyze the different effects of bridge defects and aggregate their faulty behavior into new fault models, undefined coupling fault and dynamic undefined coupling fault. In addition, an enhanced March algorithm is designed to detect all the modeled faults. In one resistor crossbar with$N$memristors, the enhanced March algorithm requires$8N$write and$7N$read operations with negligible hardware overhead. To reduce the test time, a March RC algorithm is proposed based on read operations with new reference currents, which requires$4N+2$write and$6N$read operations. Analytical results show that the proposed test algorithms can detect all the modeled faults outperforming all the previous methods. Subsequently, a Design-for-Testability scheme is proposed to implement March RC algorithm with a little area overhead. Peng Liu 0045, Zhiqiang You, Jigang Wu, Bosheng Liu, Yinhe Han 0001, Krishnendu Chakrabarty |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2020 | Search-free Accelerator for Sparse Convolutional Neural NetworksabstractSparsification is an efficient solution to reduce the demand of on-chip memory space for deep convolutional neural networks (CNNs). Most of state-of-the-art CNN accelerators can deliver high throughput for sparse CNNs by searching pairs of nonzero weights and activations, and then sending them to processing elements (PEs) for multiplication-accumulation (MAC) operations. However, their PE scales are difficult to be increased for superior and efficient computing because of the significant internal interconnect and memory bandwidth consumption. To deal with this dilemma, we propose a sparsity-aware architecture, called Swan, which frees the search process for sparse CNNs under limited interconnect and bandwidth resources. The architecture comprises two parts: a MAC unit that can free the search operation for the sparsity-aware MAC calculation, and a systolic compressive dataflow that well suits the MAC architecture and greatly reuses inputs for interconnect and bandwidth saving. With the proposed architecture, only one column of the PEs needs to load/store data while all PEs can operate in full scale. Evaluation results based on a place-and-route process show that the proposed design, in a compact factor of 4096 PEs, 4.9TOP/s peak performance, and 2.97W power running at 600MHz, achieves 1.5-2.1× speedup and 6.0-9.1× higher energy efficiency than state-of-the-art CNN accelerators with the same PE scale. Bosheng Liu, Xiaoming Chen 0003, Yinhe Han 0001, Ying Wang 0001, Xiaowei Li 0001 |
ASP-DAC | 1 |
| 2020 | Swallow: A Versatile Accelerator for Sparse Neural NetworksabstractSparse neural networks (SNNs) are emerging as a promising technique for resource-limited intelligent embedded systems because of the compact model size and the un compromised accuracy. Recently, most of the dedicated neural network accelerators are beginning to exploit the sparsity of neural network models for performance boost and energy saving. However, existing sparsity-aware accelerators fail to support both sparse weights and activations in neural networks or support them at the same time for both convolutional (Conv) layers and fully connected (FC) layers, which dominate the computational time of neural networks. In this article, we propose a novel sparsity-aware accelerator architecture, called Swallow, to sufficiently improve the inference performance by eliminating ineffectual weights and activations of neural networks. Swallow comprises: 1) a 2-D systolic architecture that fully utilizes the sparsity of both weights and activations in both Conv and FC layers and 2) a sparsity-aware dataflow which is optimized to reuse both weights and activations and to achieve high processing element (PE) utilization by sparse matrix multiplication tiling. Comprehensive evaluations based on a place-and-route process show that Swallow, with 614 GOP/s peak performance and 1.26-W power, outperforms a state-of-the-art sparsity-aware accelerator Cambricon-X by 1.32× in term of energy efficiency. Bosheng Liu, Xiaoming Chen 0003, Yinhe Han 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Simulate-the-hardware: training accurate binarized neural networks for low-precision neural acceleratorsabstractThis work investigates how to effectively train binarized neural networks (BNNs) for the specialized low-precision neural accelerators. When mapping BNNs onto the specialized neural accelerators that adopt fixed-point feature data representation and binary parameters, due to the operation overflow caused by short fixed-point coding, the BNN inference results from the deep learning frameworks on CPU/GPU will be inconsistent with those from the accelerators. This issue leads to a large deviation between the training environment and the inference implementation, and causes potential model accuracy losses when deployed on the accelerators. Therefore, we present a series of methods to contain the overflow phenomenon, and enable typical deep learning frameworks like Tensorflow to effectively train BNNs that could work with high accuracy and convergence speed on the specialized neural accelerators. Ying Wang 0001, Bosheng Liu, Yinhe Han 0001, Xiaowei Li 0001 |
ASP-DAC | 3 |
| 2019 | Addressing the issue of processing element under-utilization in general-purpose systolic deep learning acceleratorsabstractAs an energy-efficient hardware solution for deep neural network (DNN) inference, systolic accelerators are particularly popular in both embedded and datacenter computing scenarios. Despite their excellent performance and energy efficiency, however, systolic DNN accelerators are naturally facing a resource under-utilization problem - not all DNN models can well match the fixed processing elements (PEs) in a systolic array implementation, because typical DNN models vary significantly from applications to applications. Consequently, state-of-the-art hardware solutions are not expected to deliver the nominal (peak) performance and energy efficiency as claimed because of resource under-utilization. To deal with this dilemma, this study proposes a novel systolic DNN accelerator with a flexible computation mapping and dataflow scheme. By providing three types of parallelism and dynamically switching among them: channel-direction mapping, planar mapping, and hybrid, our accelerator offers the adaptability to match various DNN models to the fixed hardware resources, and thus, enables flexibly exploiting PE provision and data reuse for a wide range of DNN models to achieve optimal performance and energy efficiency. Bosheng Liu, Xiaoming Chen 0003, Ying Wang 0001, Yinhe Han 0001, Xiaowei Li 0001 |
ASP-DAC | 1 |
| 2019 | Merging Everything (ME): A Unified FPGA Architecture Based on Logic-in-Memory TechniquesabstractNo abstract available. Xiaoming Chen 0003, Longxiang Yin, Bosheng Liu, Yinhe Han 0001 |
DAC | 3 |
| 2019 | ACG-Engine: An Inference Accelerator for Content Generative Neural NetworksabstractThe technological breakthrough in Generative Adversarial Networks (GAN) has propelled the advancement of content generative applications such as AI-based paintings, style transfer, and music composition. However, in contrast to previous deep learning models for prediction and categorization, generative networks generally rely on instance normalization (IN) layer for better feature distribution, which performs significantly better than batch normalization(BN) in image style-transfer, image to image translation, etc. Unlike batch or group normalization that can be fused into convolutional layers and ignored during the network inference stage, an instance normalization layer induces intensive computation and memory access. However, prior deep learning accelerator designs for traditional Neural Network and Generative Adversarial Networks mostly focus on the acceleration of convolution and deconvolution layer but lack of support for IN operations, which could become a performance bottleneck on edge devices with insufficient computational power. To address this problem, we propose an inference accelerator for content generation (ACG-Engine) aimed to support the fundamental operations of generative networks, including convolution layers, deconvolution layers, specifically instance normalization layer. We performed a hardware-aware mathematical transformation of the IN operation for less computation complexity and memory-friendliness, so that it can be efficiently mapped to the classic 2D processing element array. Owing to the proposed optimization techniques, ACG-Engine achieves 4.56X speedup and improve power efficiency up to 29X compared to prior baseline acceleration scheme in generative network acceleration. In addition, ACG-Engine can achieve performance comparable to the classic CNN-specific accelerators with negligible power consumption and area overhead. Ying Wang 0001, Bosheng Liu, Yinhe Han 0001 |
ICCAD | 5 |
| 2019 | Accelerating DNN-based 3D point cloud processing for mobile computing
Bosheng Liu, Xiaoming Chen 0003, Yinhe Han 0001, Xiaowei Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2016 | C-brain: a deep learning accelerator that tames the diversity of CNNs through adaptive data-level parallelizationabstractConvolutional neural networks (CNN) accelerators have been proposed as an efficient hardware solution for deep learning based applications, which are known to be both compute-and-memory intensive. Although the most advanced CNN accelerators can deliver high computational throughput, the performance is highly unstable. Once changed to accommodate a new network with different parameters like layers and kernel size, the fixed hardware structure, may no longer well match the data flows. Consequently, the accelerator will fail to deliver high performance due to the underutilization of either logic resource or memory bandwidth. To overcome this problem, we proposed a novel deep learning accelerator, which offers multiple types of data-level parallelism: inter-kernel, intra-kernel and hybrid. Our design can adaptively switch among the three types of parallelism and the corresponding data tiling schemes to dynamically match different networks or even different layers of a single network. No matter how we change the hardware configurations or network types, the proposed network mapping strategy ensures the optimal performance and energy-efficiency. Compared with previous state-of-the-art NN accelerators, it is possible to achieve a speedup of 4.0x-8.3x for some layers of the well-known large scale CNNs. For the whole phase of network forward-propagation, our design achieves 28.04% PE energy saving, 90.3% on-chip memory energy saving on average. Lili Song, Ying Wang 0001, Yinhe Han 0001, Xin Zhao 0044, Bosheng Liu, Xiaowei Li 0001 |
DAC | 5 |