Baoting Li

dblp:235/0785 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0008-4807-3731ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Low-Error Approximate Logarithmic Multiplier with Symmetric LUT for Efficient DNN Training
Baoting Li, Tai Yu, Xuchong Zhang, Hongbin Sun 0001
ISCAS2
2025 Integrating models of real aboveground scene and underground geological structures at an open pit mine
abstract
Background As information technology has advanced and been popularized, open pit mining has rapidly developed toward integration and digitization. The three-dimensional reconstruction technology has been successfully applied to geological reconstruction and modeling of surface scenes in open pit mines. However, an integrated modeling method for surface and underground mine sites has not been reported. Methods In this study, we propose an integrated modeling method for open pit mines that fuses a real scene on the surface with an underground geological model. Based on oblique photography, a real-scene model was established on the surface. Based on the surface-stitching method proposed, the upper and lower surfaces and sides of the model were constructed in stages to construct a complete underground three-dimensional geological model, and the aboveground and underground models were registered together to build an integrated open pit mine model. Results The oblique photography method used reconstructed a surface model of an open pit mine using a real scene. The surface-stitching algorithm proposed was compared with the ball-pivoting and Poisson algorithms, and the integrity of the reconstructed model was markedly superior to that of the other two reconstruction methods. In addition, the surface-stitching algorithm was applied to the reconstruction of different formation models and showed good stability and reconstruction efficiency. Finally, the aboveground and underground models were accurately fitted after registration to form an integrated model. Conclusions The proposed method can efficiently establish an integrated open pit model. Based on the integrated model, an open pit auxiliary planning system was designed and realized. It supports the functions of mining planning and output calculation, assists users in mining planning and operation management, and improves production efficiency and management levels.
Biao Dong, Wenjun Tan, Weichao Chang, Baoting Li, Yanliang Guo, Quanxing Hu, Guangwei Liu
Virtual Real. Intell. Hardw.4
2024 An Efficient Sparse-Aware Summation Optimization Strategy for DNN Accelerator
abstract
Due to the various applications and high sparsity of deep neural network (DNN), a lot of sparse-aware DNN accelerators have been proposed to exploit the sparsity in DNN. Furthermore, it is essential to optimize for the accumulations and inter-channel aggregations in DNN to reduce memory overhead and improve performance of DNN accelerator. However, the uncertain number and location of non-zero element in DNN pose critical challenges for optimizing such accelerators and this inspires us to explore an efficient spare-aware summation optimization strategy for DNN accelerator. In this paper, we leverage the strategy that trading higher cost memory storage/access for lower cost computation to propose a random index based sparse-aware adder tree (RAT), which achieves a better trade-off among performance, hardware resource overhead and adaptability. Synthesis and simulation results demonstrate that, compared with reference design, the proposed design achieves 1.71× and 1.52× the normalized area efficiency and energy efficiency improvement on ResNet18, respectively.
Danqing Zhang, Baoting Li, Xuchong Zhang, Hongbin Sun 0001
ISCAS2
2024 DQ-STP: An Efficient Sparse On-Device Training Processor Based on Low-Rank Decomposition and Quantization for DNN
abstract
Due to the bottleneck problems such as scenario-varying application, significant data communication overhead and privacy protection between off-line training and on-line inference, intelligent edge devices capable of adaptively fine-tuning the deep neural network (DNN) models for specific tasks have become the most urgent need. However, the computational cost is intolerable for ordinary on-device training (ODT), which inspires us to explore an efficient ODT processor, named DQ-STP. In this paper, we leverage a series of optimization techniques using software-hardware co-design. On the one hand, the proposed design incorporates SVD-based low-rank decomposition,$2^{n}$quantization and ACBN algorithm on the software side. This unifies the sparse computing mode of convolutional layers and enhancing weight sparsity. On the other hand, the proposed design effectively leverages data sparsity on the hardware side through four techniques: 1) The flag compressed sparse row is proposed to compress input feature maps and gradient maps. 2) A unified processing element (PE) array comprising shifters and adders is proposed to expedite forward and error propagation steps. 3) The PE arrays for error propagation and weight gradients generation are separated to enhance throughput. 4) A sparse alignment strategy is proposed to further enhance PE utilization. Through these software and hardware co-optimization, the proposed DQ-STP achieves an area efficiency and peak energy efficiency of 41.2 GOPS/mm2 and 90.63 TOPS/W. In comparison to state-of-the-art reference designs, the proposed DQ-STP demonstrates a$2.19\times $improvement in normalized area efficiency and a$1.85\times $enhancement in energy efficiency.
Baoting Li, Danqing Zhang, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 Multi-region Quality Assessment Based on Spatial-Temporal Community Detection from Computed Tomography Images
Tao Wen 0001, Tongze Xu, Baoting Li, Zhenning Wu
ADMA (4)4
2023 A Fine-Grained Verification Method for Blockchain Data Based on Merkle Path Sharding
Liang Wen, Zhiqiong Wang, Tingyu Cui, Caiyun Shi, Baoting Li, Zhongming Yao
ADMA (4)5
2023 ACBN: Approximate Calculated Batch Normalization for Efficient DNN On-Device Training Processor
abstract
Batch normalization (BN) has been established as a very effective component in deep learning, largely helping accelerate the convergence of deep neural network (DNN) training. Nevertheless, its hardware architecture has not received much attention in the field of DNN on-device training processors. Several previous designs incur either high off-chip memory traffic or high circuit complexity, and hence have deficiencies in terms of hardware efficiency and performance. This article proposes approximately calculated BN (ACBN) to achieve a much better tradeoff between hardware efficiency and performance for DNN on-device training processors. The accuracy and convergence rate of the proposed ACBN have been extensively evaluated using four typical DNN models. Compared with the state-of-the-art reference design, the hardware simulation results show the proposed ACBN can at least reduce floating point operations by 22.2% and save external memory access by 33.3% on average. Moreover, the proposed ACBN introduces 63.6% data sparsity for the backward propagation of BN layers of VGG16 on average. To the best of our knowledge, we are the first to introduce data sparsity for the backward propagation of BN layers. The ACBN module is implemented on Zynq UltraScale+ ZCU102 system-on-chip (SoC) field-programmable gate array (FPGA), and the results show that the implementation of ACBN hardware module saves 33.9% look-up table (LUT), 49.4% flip-flop (FF), 75% digital signal processor (DSP), and reduces the power by 12.4% compared with the reference design while achieving better performance.
Baoting Li, Fujie Luo, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2021 Dynamic Dataflow Scheduling and Computation Mapping Techniques for Efficient Depthwise Separable Convolution Acceleration
abstract
Depthwise separable convolution (DSC) has become one of the essential structures for lightweight convolutional neural networks. Nevertheless, its hardware architecture has not received much attention. Several previous hardware designs incur either high off-chip memory traffic or large on-chip memory usage, and hence have deficiency in terms of hardware efficiency as well as performance. This paper proposes two efficient dynamic design techniques, i.e. adaptive row-based dataflow scheduling and adaptive computation mapping, to achieve a much better trade-off between hardware efficiency and performance for DSC-based lightweight CNN accelerator. The effectiveness and efficiency of the proposed dynamic design techniques have been extensively evaluated using six DSC-based lightweight CNNs. Compared with the reference architectures, the simulation results show the proposed architectural techniques can at least reduce on-chip buffer size by 50.4% and improve the performance of convolution calculation by 1.18× while maintaining the minimum off-chip memory traffic. MobileNetV2 is implemented on Zynq UltraScale+ ZCU102 SoC FPGA, and the results show the proposed accelerator can achieve 381.7 frames per second (fps), which is 1.43× of the reference design, and it can save about 36.3% on-chip buffer size compared with the reference design, while maintaining the same off-chip memory traffic.
Baoting Li, Xuchong Zhang, Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 Designing Efficient Shortcut Architecture for Improving the Accuracy of Fully Quantized Neural Networks Accelerator
abstract
Network quantization is an effective solution to compress Deep Neural Networks (DNN) that can be accelerated with custom circuit. However, existing quantization methods suffer from significant loss in accuracy. In this paper, we propose an efficient shortcut architecture to enhance the representational capability of DNN between different convolution layers. We further implement the shortcut hardware architecture to effectively improve the accuracy of fully quantized neural networks accelerator. The experimental results show that our shortcut architecture can obviously improve network accuracy while increasing very few hardware resources ( 0.11 × and 0.17 × for LUT and FF respectively) compared with the whole accelerator.
Baoting Li, Longjun Liu, Yanming Jin, Hongbin Sun 0001, Nanning Zheng 0001
ASP-DAC1