Shengli Lu

dblp:48/8123 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0001-5769-8671ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 CTSNet: Collaborative temporal-spatial net with dual-branch cross-attention for dynamic IR drop prediction
Hongshun Zheng, Hongxiang Liu, Shengli Lu
Integr.5
2025 Low-latency and energy-efficient FPGA accelerator for sparse neural networks in edge LiDAR-based 3D object detection
Jinyi Li, Hengrui Hu, Shengli Lu
J. Supercomput.5
2025 Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
abstract
Sparse triangular solve (SpTRSV) is widely used in various domains. Numerous studies have been conducted using CPUs, GPUs, and specific hardware accelerators, where dataflows can be categorized into coarse and fine granularity. Coarse dataflows offer good spatial locality but suffer from low parallelism, while fine dataflows provide high parallelism but disrupt the spatial structure, leading to increased nodes and poor data reuse. This article proposes a novel hardware accelerator for SpTRSV or SpTRSV-like directed acyclic graphs (DAGs). The accelerator implements a medium granularity dataflow through hardware-software codesign and achieves both excellent spatial locality and high parallelism. In addition, a partial sum caching mechanism is introduced to reduce the blocking frequency of processing elements (PEs), and a reordering algorithm of intranode edges’ computation is developed to enhance data reuse. Experimental results on 245 benchmarks with node counts reaching up to 85392 demonstrate that this work achieves average performance improvements of$7.0\times $(up to$27.8\times $) over CPUs and$5.8\times $(up to$98.8\times $) over GPUs. Compared with the state-of-the-art technique (DPU-v2), this work shows a$2.5\times $(up to$5.9\times $) average performance improvement and$1.7\times $(up to$4.1\times $) average energy efficiency enhancement.
Shengli Lu
IEEE Trans. Very Large Scale Integr. Syst.3
2021 HFPQ: deep neural network compression by hardware-friendly pruning-quantization
YingBo Fan, Wei Pang 0003, Shengli Lu
Appl. Intell.3
2021 Few-shot fine-grained classification with Spatial Attentive Comparison
Xiaoqian Ruan, Guosheng Lin, Cheng Long 0001, Shengli Lu
Knowl. Based Syst.4
2015 A Ripple Control Dual-Mode Single-Inductor Dual-Output Buck Converter With Fast Transient Response
abstract
A novel integrated single-inductor dual-output buck converter based on the ripple control technique is presented, which can provide two independent output voltages (1.2 and 1.8 V) only using one inductor. The peak current common-mode and ripples compare differential-mode are adopted to improve the system transient response. The converter can automatically switch between pulsewidth modulation and pulse-skip modulation to improve the light load conversion efficiency. The proposed converter has been fabricated in a 0.18 μm 1P6M CMOS process. The experimental results show the transient response time is about 10 μs and the cross-regulation is less than 0.05 mV/mA while the load current suddenly changes 200 mA. In general, the conversion efficiency at light load is above 86% and the peak efficiency at the full load can achieve 93.5%.
Caixia Han, Shengli Lu
IEEE Trans. Very Large Scale Integr. Syst.5
2007 Multi-Cycle Measuring and Data Fusion for Large End Face Run-Out of Tapered Roller
abstract
The large end face run-out is a kind of shape and position tolerance, and is an important index in deciding the class of the tapered roller. The auto-measuring and sorting equipments for the tapered roller introduced in this article is an intelligent measuring equipment that orientates in pursuing measurement and choosing of every tapered roller on grinding spot. According to professional standard about the tapered roller, the large end face run-out (notated as SDW) must be measured dynamically during revolving on normal V-shape table. In order to eliminate the disturbance caused by electromagnetism interference and machine vibration, this equipment adopts multi-cycle measuring under single sensor, making better use of the regeneration of the sampling signal and based on analysis of the main factors leading to the large end face run-out and signal characteristics. On the other hand, two kinds of signal-characteristics-based data association threshold filtering algorithm and an improved weight arithmetic mean data fusion algorithm are applied on data processing software embedded in it. This equipment has been proved by practical application and has obtained a Chinese patent.
Shengli Lu, Weichang Liang
ISDA1