EDBT 2026 Demo / reviewers in the wild / expert
Tianshuo Lu
dblp:407/2205
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0005-9301-4130ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Configurable Streaming Accelerator for LUT-Based Super-Resolution on FPGA
Xuzhuo Hu, Jianyang Ding, Huachen Zhang, Tianshuo Lu, ZhiLei Chai |
APPT | 5 |
| 2026 | RACP: An Efficient RISC-V Domain-Specific Processor for Arbitrary-Size Kernel CNNs
Jianyang Ding, Tianshuo Lu, Huachen Zhang, Xuzhuo Hu, ZhiLei Chai |
ISCAS | 3 |
| 2026 | An efficient RISC-V processor with customized instruction set for sparse DNN acceleration on embedded system
Jianyang Ding, Huachen Zhang, Tianshuo Lu, ZhiLei Chai |
J. Syst. Archit. | 4 |
| 2025 | Hybrid-SANet: Hybrid Self-attention Transformer for Efficient Image Super-Resolution
Jianyang Ding, Huachen Zhang, Nachuan Zhang, Tianshuo Lu, ZhiLei Chai |
CGI (3) | 5 |
| 2025 | EVO-QNN: Efficient Mixed-Precision Quantization Inference on RISC-V-Based Edge DeviceabstractMixed-Precision Quantized Neural Network (MPQNN) helps balance inference precision and efficiency under resource constraints, while most of them lack high-energy-efficiency hardware acceleration solutions. To address these challenges, we propose a SW/HW co-design framework termed EVO-QNN for low-energy and low-latency inference. Specifically, EVO-QNN manages bit-widths of operators and extends SIMD instructions based on customized RISC-V core. Experimental results demonstrate that our framework can achieve performance improvement ranging from 1.23x to 1.58x, with only a 1.02% and 6.61% increasing in area and power consumption for 2–8 bit convolution operators. Tianshuo Lu, Jianyang Ding, Huachen Zhang, ZhiLei Chai |
FCCM | 1 |
| 2025 | RV-ESMC: Efficient Sparse Matrix Convolution Processor based on RISC-V Custom instructions for Edge PlatformsabstractAs the demand for deep neural network (DNN) inference on edge platforms grows, deploying compute-intensive DNNs on resource-constrained devices remains challenging. This paper proposes a novel sparse convolution acceleration processor, RV-ESMC, based on RISC-V architecture, with custom instructions to enable efficient edge DNN inference. RV-ESMC provides flexibility by supporting inline assembly calls in C programming. Experimental results indicate that RV-ESMC can reduce execution time by over 70% in DNNs with convolution operations compared to conventional instruction sets. The functionality of RV-ESMC is validated on an FPGA platform and its performance is comprehensively evaluated based on a 55nm CMOS process. The results show that RV-ESMC can achieve a peak energy efficiency of 675 GOPS/W. Huachen Zhang, Jianyang Ding, Tianshuo Lu, ZhiLei Chai |
FCCM | 4 |
| 2025 | Optimizing Sparse Matrix Convolution on RISC-V Core: Custom Instructions for Embedded SystemabstractWith the increasing demand for deep neural network (DNN) inference tasks on embedded platforms, deploying compute-intensive DNNs on resource-constrained embedded platforms faces challenges. While sparsification technology offers a potential solution, its implementation on edge platforms still faces difficulties. In this article, we propose a novel sparse convolution acceleration processor based on RISC-V architecture, and design specialized custom instructions to enable efficient edge DNN inference. To this end, we mainly address three technical issues. In response to numerical characteristics of sparse convolution, the designed processor can implement a hardware-friendly architecture that transforms convolutions into sparse matrix multiplication. Additionally, it employs a column-major and element-level parallel strategy to optimize load imbalance issues present in the Gustavson algorithm, thereby enhancing sparse matrix computations. To further improve computational efficiency, our work is designed by incorporating efficient execution units that reduce instruction execution overheads while minimizing memory access frequency. Compared to traditional accelerators, our work supports custom instruction formats in the C programming language, offering superior flexibility. Extensive experimental results indicate that our work can reduce execution time over 70% when running most DNNs with convolution operations compared to conventional instruction sets. Moreover, the functionality of our work is validated on an FPGA platform, and its performance is comprehensively evaluated based on a 55 nm CMOS process. The results show that our work can achieve a peak energy efficiency of 675 GOPS/W in most network inference tasks, demonstrating exceptional computational performance and energy efficiency. Huachen Zhang, Jianyang Ding, Tianshuo Lu, ZhiLei Chai |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | ISRLUT: Integer-Only FHD Image Super-Resolution Based on Neural Lookup Table and Near-Memory ComputingabstractWhile Deep Neural Networks (DNNs) have achieved remarkable progress in Image Super-Resolution (SR) task, they face significant challenges for edge processing FHD images. Complex DNN operators lead to high hardware resource consumption and latency. Computational inefficiency of FPU increases energy consumption, while DDR access overhead and on-chip memory overflow further constrain real-time capabilities. To address this, we propose ISRLUT, a novel accelerator architecture focused on integer-only inference and near-memory computing. Its core contributions include: (1) Fusion of Neural LUT arithmetic with reconfigurable compute units, transforming unified LUT operators from DNN operators and enhancing hardware utilization; (2) An integer-only inference and parallel architecture, eliminating floating-point dependencies and significantly reducing energy consumption; (3) An innovative internal operator memory management scheme coupled with Tile-based Buffer Overlap and Private Cache Mechanism. We deploy ISRLUT on FPGA and ASIC platforms. Experiments demonstrate that ISRLUT achieves efficient performance: For 4 \(\times\) upscaling, it requires only 36.9 KB of storage and achieves a PSNR of 30.21 dB on Set5. Hardware implementation using a 55 nm ASIC consumes merely 0.0337 W power, delivers an energy efficiency of 7278.6 Mpixels/s/W, and achieves a real-time frame rate of 118 FPS for 4 \(\times\) FHD processing, validating its superiority in energy efficiency and hardware utilization. Tianshuo Lu, Jianyang Ding, Huachen Zhang, ZhiLei Chai |
ACM Trans. Reconfigurable Technol. Syst. | 1 |