VLDB 2026 Research / reviewers in the wild / expert
Jianyang Ding
dblp:255/1369
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Configurable Streaming Accelerator for LUT-Based Super-Resolution on FPGA
Xuzhuo Hu, Jianyang Ding, Huachen Zhang, Tianshuo Lu, ZhiLei Chai |
APPT | 2 |
| 2026 | RACP: An Efficient RISC-V Domain-Specific Processor for Arbitrary-Size Kernel CNNs
Jianyang Ding, Tianshuo Lu, Huachen Zhang, Xuzhuo Hu, ZhiLei Chai |
ISCAS | 2 |
| 2026 | An efficient RISC-V processor with customized instruction set for sparse DNN acceleration on embedded system
Jianyang Ding, Huachen Zhang, Tianshuo Lu, ZhiLei Chai |
J. Syst. Archit. | 2 |
| 2025 | Hybrid-SANet: Hybrid Self-attention Transformer for Efficient Image Super-Resolution
Jianyang Ding, Huachen Zhang, Nachuan Zhang, Tianshuo Lu, ZhiLei Chai |
CGI (3) | 2 |
| 2025 | EVO-QNN: Efficient Mixed-Precision Quantization Inference on RISC-V-Based Edge DeviceabstractMixed-Precision Quantized Neural Network (MPQNN) helps balance inference precision and efficiency under resource constraints, while most of them lack high-energy-efficiency hardware acceleration solutions. To address these challenges, we propose a SW/HW co-design framework termed EVO-QNN for low-energy and low-latency inference. Specifically, EVO-QNN manages bit-widths of operators and extends SIMD instructions based on customized RISC-V core. Experimental results demonstrate that our framework can achieve performance improvement ranging from 1.23x to 1.58x, with only a 1.02% and 6.61% increasing in area and power consumption for 2–8 bit convolution operators. Tianshuo Lu, Jianyang Ding, Huachen Zhang, ZhiLei Chai |
FCCM | 2 |
| 2025 | RV-ESMC: Efficient Sparse Matrix Convolution Processor based on RISC-V Custom instructions for Edge PlatformsabstractAs the demand for deep neural network (DNN) inference on edge platforms grows, deploying compute-intensive DNNs on resource-constrained devices remains challenging. This paper proposes a novel sparse convolution acceleration processor, RV-ESMC, based on RISC-V architecture, with custom instructions to enable efficient edge DNN inference. RV-ESMC provides flexibility by supporting inline assembly calls in C programming. Experimental results indicate that RV-ESMC can reduce execution time by over 70% in DNNs with convolution operations compared to conventional instruction sets. The functionality of RV-ESMC is validated on an FPGA platform and its performance is comprehensively evaluated based on a 55nm CMOS process. The results show that RV-ESMC can achieve a peak energy efficiency of 675 GOPS/W. Huachen Zhang, Jianyang Ding, Tianshuo Lu, ZhiLei Chai |
FCCM | 2 |
| 2025 | Lightweight Transformer with Enhanced Inverted Residual Blocks for Bird Sound RecognitionabstractBird sound recognition relies on capturing unique characteristics of avian vocalizations to achieve accurate species identification. Recent advances in deep learning have significantly improved classification accuracy and Transformer-based models stand out due to their superior ability to model long-range dependencies. However, most of them suffer from high computational complexity, posing challenges for deployment on edge devices in the wild. To address this issue, this paper proposes a lightweight Transformer-based model incorporating enhanced Inverted Residual Blocks (IRB). To this end, we first replace original Multilayer Perceptron (MLP) modules with IRB. Additionally, we propose a stage-wise incremental strategy for setting expansion factors to reduce redundancy. At the same time, we incorporate residual connections before and after depthwise convolutions to maintain model performance. Furthermore, we refine attention configurations throughout various phases of the network. We strategically minimize redundant attention in the initial stages, while intensifying its application in the crucial stages. This approach achieves an optimal balance between computation and accuracy. Experimental results on the Bird-CLEF2023, DCASE2020, and Birdsdata demonstrate that our proposed model can achieve approximately a 5-fold reduction in parameters, an 85% decrease in computational load, and a 2.7-fold increase in inference speed on Jetson AGX Xavier, while preserving high accuracy. Xiangyu Cheng, Shengnan Fan, Jianyang Ding, ZhiLei Chai |
IJCNN | 3 |
| 2025 | Optimizing Sparse Matrix Convolution on RISC-V Core: Custom Instructions for Embedded SystemabstractWith the increasing demand for deep neural network (DNN) inference tasks on embedded platforms, deploying compute-intensive DNNs on resource-constrained embedded platforms faces challenges. While sparsification technology offers a potential solution, its implementation on edge platforms still faces difficulties. In this article, we propose a novel sparse convolution acceleration processor based on RISC-V architecture, and design specialized custom instructions to enable efficient edge DNN inference. To this end, we mainly address three technical issues. In response to numerical characteristics of sparse convolution, the designed processor can implement a hardware-friendly architecture that transforms convolutions into sparse matrix multiplication. Additionally, it employs a column-major and element-level parallel strategy to optimize load imbalance issues present in the Gustavson algorithm, thereby enhancing sparse matrix computations. To further improve computational efficiency, our work is designed by incorporating efficient execution units that reduce instruction execution overheads while minimizing memory access frequency. Compared to traditional accelerators, our work supports custom instruction formats in the C programming language, offering superior flexibility. Extensive experimental results indicate that our work can reduce execution time over 70% when running most DNNs with convolution operations compared to conventional instruction sets. Moreover, the functionality of our work is validated on an FPGA platform, and its performance is comprehensively evaluated based on a 55 nm CMOS process. The results show that our work can achieve a peak energy efficiency of 675 GOPS/W in most network inference tasks, demonstrating exceptional computational performance and energy efficiency. Huachen Zhang, Jianyang Ding, Tianshuo Lu, ZhiLei Chai |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2025 | ISRLUT: Integer-Only FHD Image Super-Resolution Based on Neural Lookup Table and Near-Memory ComputingabstractWhile Deep Neural Networks (DNNs) have achieved remarkable progress in Image Super-Resolution (SR) task, they face significant challenges for edge processing FHD images. Complex DNN operators lead to high hardware resource consumption and latency. Computational inefficiency of FPU increases energy consumption, while DDR access overhead and on-chip memory overflow further constrain real-time capabilities. To address this, we propose ISRLUT, a novel accelerator architecture focused on integer-only inference and near-memory computing. Its core contributions include: (1) Fusion of Neural LUT arithmetic with reconfigurable compute units, transforming unified LUT operators from DNN operators and enhancing hardware utilization; (2) An integer-only inference and parallel architecture, eliminating floating-point dependencies and significantly reducing energy consumption; (3) An innovative internal operator memory management scheme coupled with Tile-based Buffer Overlap and Private Cache Mechanism. We deploy ISRLUT on FPGA and ASIC platforms. Experiments demonstrate that ISRLUT achieves efficient performance: For 4 \(\times\) upscaling, it requires only 36.9 KB of storage and achieves a PSNR of 30.21 dB on Set5. Hardware implementation using a 55 nm ASIC consumes merely 0.0337 W power, delivers an energy efficiency of 7278.6 Mpixels/s/W, and achieves a real-time frame rate of 118 FPS for 4 \(\times\) FHD processing, validating its superiority in energy efficiency and hardware utilization. Tianshuo Lu, Jianyang Ding, Huachen Zhang, ZhiLei Chai |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | Multimodal Fusion-GMM based Gesture Recognition for Smart Home by WiFi SensingabstractWith the thriving development of communication systems and the ubiquity of commodity WiFi devices, wireless intelligent sensing has attracted increasingly attention owing to its importance in Human-Computer Interaction (HCI) where hand gesture recognition has been widely studied these years. However, most of gesture recognition approaches have a limited performance owing to the background noises. In this paper, we present a multimodal fusion-Gaussian mixed model (GMM) based gesture recognition scheme by exploiting channel state information (CSI) of WiFi signals. To this end, we firstly design two theoretical underpinnings, including a sensing model and a recognition model. Firstly, the sensing model is established to investigate the impact of gestures on the propagation properties of WiFi signals and quantify the correlation between the CSI dynamics and various gestures. Secondly, the recognition model is used to exploit physical gesture-induced signal changes to infer potential gestures information. Then, we implement the proposed scheme on a set of WiFi devices and evaluate it in both laboratory and corridor environments. The experiment results show that the proposed scheme can achieve average recognition accuracies of 96% and 94% in these two scenarios, respectively. This shows promise for future ubiquitous hands-free gesture-based interaction with mobile devices. Jianyang Ding, Yong Wang 0013, Hongyan Si, Jiannan Ma, Jingwen He, Shaozhong Fu |
VTC Spring | 1 |
| 2022 | Three-Dimensional Indoor Localization and Tracking for Mobile Target Based on WiFi SensingabstractWith the prevalence of commodity WiFi devices and development of the Internet of Things (IoT), the usage of WiFi has been extended from communication to context aware. Based on this, indoor localization has attracted increasing attention in the academic community, without the need for additional sensors and any active engagement from the users. However, the localization performance is vulnerable to the background noises due to relying on signals changes. To address this issue, in this article, we propose a passive 3-D indoor localization with a radio map for the mobile target by exploiting channel state information (CSI) of WiFi signals, realizing human–computer interaction (HCI). To this end, 3-D space is first divided into multiple independent regions and we construct a spatial radio map by traversing all the subspaces using CSI measurements. Next, to fully characterize the profiles of the locations of the mobile target, we reconstruct CSI time series to form a CSI tensor through integrating WiFi transmission links and a CANDECAMP/PARAFAC (CP) decomposition method is applied to this tensor for obtaining representative features. Then, the features-location data set is optimized by combining$t$-distributed stochastic neighbor embedding (t-SNE) and information theory method to reconstruct a fine-grained fingerprint map for improving system performance. Finally, a recurrent neural network (RNN) model is introduced to learn the features data set optimized and then build a nonlinear correlation between input and output for realizing the purpose of accurate indoor localization. The proposed scheme is implemented on a set of commodity WiFi devices and evaluated in indoor scenarios. Based on real-world CSI data, our experimental results confirm the effectiveness of the proposed scheme in terms of localization accuracy and robustness against the noises. Jianyang Ding, Yong Wang 0013, Hongyan Si, Jiwei Xing |
IEEE Internet Things J. | 1 |