EDBT 2026 Demo / reviewers in the wild / expert
Yiren Zhu
dblp:305/6293
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiveHPS-Lite: A Lightweight LiDAR-based Motion Capture System for Edge ApplicationsabstractRecent advances in LiDAR-based 3D human motion capture have demonstrated significant potential for large-scale applications in unconstrained environments. However, achieving real-time performance remains challenging, particularly under the computational constraints of edge devices where deploying large deep learning models is often impractical. To address these limitations, we propose LiveHPS-Lite, a lightweight single-LiDAR-based human motion capture system, offering enhanced computational efficiency with competitive performance. In particular, we introduce a novel architecture by streamlining backbone components across all processing stages in the LiveHPS++ framework and replacing inconsistent sequential modules with parallelizable minGRUs. We implement the proposed architecture on NVIDIA Jetson Xavier NX with TensorRT acceleration, achieving real-time performance on edge. Comprehensive evaluations on benchmark datasets show that LiveHPS-Lite achieves comparable or superior accuracy while significantly reducing computational complexity. The proposed LiveHPS-Lite achieves up to $6.71 \times$ faster inference speed compared to the-state-of-the-art solutions, delivering real-time performance even on a computationally limited edge device. This work contributes a practical solution for deploying high-performance 3D human pose estimation models in real-world applications. Yiren Zhu, Junsheng Zhou, Yiming Ren 0001, Hanshu Hezi, Yuexin Ma |
ASP-DAC | 1 |
| 2025 | A Neural Rendering Coprocessor With Optimized Ray Representation and MarchingabstractNeural rendering, a transformative approach for 3-D scene reconstruction and rendering, has advanced rapidly in recent years. This article introduces an energy-efficient neural rendering coprocessor that implements the popular and widely used instant neural graphics primitive (Instant-NGP) algorithm. In particular, we address the challenges of limited resources for deploying Instant-NGP on edge by proposing a dedicated architecture, which incorporates three main innovations: 1) we optimize occupancy grid queries in the ray marching module by partitioning the grid and decoupling the query process from sampling point generation, which improves both efficiency and memory usage; 2) we introduce a bilinked list-based ray switching strategy, which ensures continuous pipeline utilization to overcome the inefficiencies caused by sequential processing; and 3) we optimize the hash encoding process by incorporating quantization-aware training (QAT), enabling the hash table to fit into on-chip memory, thereby improving performance on resource-constrained devices. To demonstrate the effectiveness of our architecture, we design and fabricate a proof-of-concept chip using 40-nm CMOS technology and develop a testing system to evaluate its performance. Measurement results validate the advantages of the proposed design, showing that our chip achieves superior energy efficiency compared to both server and edge graphics processing units (GPUs), as well as other state-of-the-art neural rendering chip designs. Zhechen Yuan, Binzhe Yuan, Chaolin Rao, Yiren Zhu, Yunxiang He, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Reconfigurable and Energy-Efficient Architecture for Deploying Multi-Layer RNNs on FPGAabstractRecurrent Neural Networks (RNNs) are extensively applied in sequence prediction tasks such as sentiment analysis, and machine translation. Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) are popular recurrent layers for mitigating gradient vanishing challenges. This paper introduces a reconfigurable hardware architecture that supports LSTM, GRU, and Fully Connected (FC) layers to accommodate the diverse structures of RNNs with minimal overhead. The proposed three-mode architecture utilizes dynamically reconfigurable components, allowing seamless mode switching among the three types of layers. Besides, Dynamic Compression (DC) for intermediate results is introduced to minimize precision loss, and layer decomposition as well as processing element grouping techniques are used to improve the processing efficiency. To validate the proposed architecture, a proof-of-concept prototype system using Intel Arria10 FPGA is built. Evaluation results demonstrate a 12% improvement over the state-of-the-art accelerator in energy efficiency when configured as GRU with a layer size set to$512{\times }$512, accompanied by a 93% reduction in block RAM and an 81% reduction in DSP resources. Xiangyu Zhang 0002, Yiren Zhu, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Half-Bridge-Active-Clamp Converter with High Step-down Capabilities for More Electric Aircraft ApplicationsabstractA Half-bridge-active-clamp (HBAC) converter is used as an alternative to replace Dual active bridge (DAB) converter for 540V/28V on-board grid in more electric aircraft applications. The HBAC converter provides an opportunity to reduce the turn-ratio of the high frequency transformer compared to DAB converter with reduced current stress on the low voltage side benefit from a current-fed LV bridge. HBAC converter shows a much better performance both on efficiency and volume compared to conventional DAB converter. In addition, a Model Predictive Control (MPC) method for LV output current regulation is proposed based on HBAC converter with the aim of achieving fast dynamic performance. Yiren Zhu, Tao Yang 0020, Serhiy Bozhko, Pat Wheeler |
IECON | 1 |