Shengzhe Lyu

dblp:381/4582 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0004-7331-7700ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 50% Reconfigurable computing and FPGAs · 50%
Computer networks
1 paper
Physical-layer communications · 87% Cellular and mobile networks · 13%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%
Artificial intelligence
2 papers
Efficient and distributed learning · 54% 3D vision · 46%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA accelerator
2.022026
SwiftChannel: Algorithm-Hardware Co-Design for Deep Learning-Based 5G Channel Estimation · IEEE Trans. Mob. Comput. 2026
ViM-Q: Energy Efficient Algorithm-Hardware Co-Design for Dynamically Quantized Vision Mamba Models · FPGA 2026
Physical-layer communications
channel estimation
1.012026
SwiftChannel: Algorithm-Hardware Co-Design for Deep Learning-Based 5G Channel Estimation · IEEE Trans. Mob. Comput. 2026
Physical-layer communications › channel estimation
deep learning-based channel estimation
1.012026
SwiftChannel: Algorithm-Hardware Co-Design for Deep Learning-Based 5G Channel Estimation · IEEE Trans. Mob. Comput. 2026
Hardware accelerators and domain-specific architectures
algorithm-hardware co-design
1.012026
ViM-Q: Energy Efficient Algorithm-Hardware Co-Design for Dynamically Quantized Vision Mamba Models · FPGA 2026
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.012026
ViM-Q: Energy Efficient Algorithm-Hardware Co-Design for Dynamically Quantized Vision Mamba Models · FPGA 2026
Wearable and physiological sensing › radio frequency sensing
mmwave radar sensing
0.912025
Argus: Multi-View Egocentric Human Mesh Reconstruction Based on Stripped-Down Wearable mmWave Add-on · SenSys 2025
Machine learning › Efficient and distributed learning
model quantization
0.312026
ViM-Q: Energy Efficient Algorithm-Hardware Co-Design for Dynamically Quantized Vision Mamba Models · FPGA 2026
Computer vision › 3D vision
human mesh recovery
0.312025
Argus: Multi-View Egocentric Human Mesh Reconstruction Based on Stripped-Down Wearable mmWave Add-on · SenSys 2025

Methods — techniques the papers use, named apart from their topics

quantization-aware training · 2.0knowledge distillation · 2.0high-level synthesis · 2.0convolutional neural network · 2.0attention mechanism · 2.0associative scan · 2.0multi-view sensing · 1.7mmwave radar · 1.7deep neural network · 1.7state-space model · 1.0state space model · 1.0
YearPublicationVenuePosition
2026 ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA
abstract
Vision Mamba (ViM) models offer a compelling efficiency advantage over Transformers by leveraging the linear complexity of State Space Models (SSMs), yet efficiently deploying them on FPGAs remains challenging. Linear layers struggle with dynamic activation outliers that render static quantization ineffective, while uniform quantization fails to capture the weight distribution at low bit-widths. Furthermore, while associative scan accelerates SSMs on GPUs, its memory access patterns are misaligned with the streaming dataflow required by FPGAs. To address these challenges, we present ViM-Q1, a scalable algorithm-hardware co-design for end-to-end ViM inference on the edge. We introduce a hardware-aware quantization scheme combining dynamic per-token activation quantization and per-channel smoothing to mitigate outliers, alongside a custom 4-bit per-block Additive Power-of-Two (APoT) weight quantization. The models are deployed on a runtime-parameterizable FPGA accelerator featuring a linear engine employing a Lookup-Table (LUT) unit to replace multiplications with shift-add operations, and a fine-grained pipelined SSM engine that parallelizes the state dimension while preserving sequential recurrence. Crucially, the hardware supports runtime configuration, adapting to diverse dimensions and input resolutions across the ViM family. Implemented on an AMD ZCU102 FPGA, ViM-Q achieves an average 4.96× speedup and 59.8× energy efficiency gain over a quantized NVIDIA RTX 3090 GPU baseline for low-batch inference on ViM-tiny. This co-design shows a viable path for deploying ViM models on resource-constrained edge devices.
Shengzhe Lyu, Yuhan She, Patrick S. Y. Hung, Ray C. C. Cheung, Weitao Xu
FCCM1
2026 ViM-Q: Energy Efficient Algorithm-Hardware Co-Design for Dynamically Quantized Vision Mamba Models
abstract
State-space models (SSMs), such as Mamba, provide an efficient alternative to Transformers for vision tasks by replacing their quadratic-cost self-attention with linear complexity state update. However, efficiently deploying Vision Mamba (ViM) models on FPGA platforms is challenging, as the latency is dominated by two key components: linear layers and the selective SSM. For the linear layers, highly dynamic activation outliers across tokens render conventional static quantization techniques ineffective. Meanwhile, while the associative scan algorithm is effective in accelerating SSM on GPUs, its data access pattern is fundamentally mismatched with FPGA architectures when mapping the model's inherently sequential recurrence, creating a critical dataflow bottleneck.
Shengzhe Lyu, Yuhan She, Patrick S. Y. Hung, Ray C. C. Cheung, Weitao Xu
FPGA1
2026 SwiftChannel: Algorithm-Hardware Co-Design for Deep Learning-Based 5G Channel Estimation
abstract
Channel estimation is crucial in 5G communication networks for optimizing transmission parameters and ensuring reliable, high-speed communication. However, the use of multiple-input and multiple-output (MIMO) and millimeter-wave (mmWave) in 5G networks presents challenges in achieving accurate estimation under strict latency requirements on resource-limited hardware platforms. To address these challenges, we proposeSwiftChannel, an algorithm-hardware co-design framework that integrates a hardware-friendly deep learning-based channel estimator with a dedicated accelerator. Our approach employs a convolutional neural network enhanced with a parameter-free attention mechanism, which effectively reconstructs full-resolution spatial-frequency domain channel matrices from low-resolution least squares (LS) estimates. We further develop a multi-stage model compression pipeline combining knowledge distillation, convolution re-parameterization, and quantization-aware training, resulting in substantial model size reduction with negligible accuracy loss. The hardware accelerator, implementing the compressed model and the LS estimator on FPGA platforms using High-level Synthesis (HLS), features a fine-grained pipeline architecture and optimized dataflow strategies. Tested on a Zynq UltraScale+ RFSoC, the accelerator achieves sub-millisecond latency, providing up to 24x speed-up and over 33x improvement in energy efficiency compared to GPU-based solutions. Extensive evaluations demonstrate that the proposed design generalizes not only across various noise levels and user mobilities, but also to a variety of unseen channel profiles, outperforming state-of-the-art baselines. By unifying algorithmic innovation with hardware-aware design, our work presents a future-proof channel estimation solution for 5G MIMO systems. The source codes for the dataset synthesis, deep learning algorithm, and HLS-based FPGA design are accessible via GitHub.
Shengzhe Lyu, Yuhan She, Di Duan, Tao Ni 0003, Yu Hin Chan, Chengwen Luo 0001, Ray C. C. Cheung, Weitao Xu
IEEE Trans. Mob. Comput.1
2025 Argus: Multi-View Egocentric Human Mesh Reconstruction Based on Stripped-Down Wearable mmWave Add-on
abstract
In this paper, we propose Argus, a wearable add-on system based on stripped-down (i.e., compact, lightweight, low-power, limited-capability) mmWave radars. It is the first to achieve egocentric human mesh reconstruction in a multi-view manner. Compared with conventional frontal-view mmWave sensing solutions, it addresses several pain points, such as restricted sensing range, occlusion, and the multipath effect caused by surroundings. To overcome the limited capabilities of the stripped-down mmWave radars (with only one transmit antenna and three receive antennas), we tackle three main challenges and propose a holistic solution, including tailored hardware design, sophisticated signal processing, and a deep neural network optimized for high-dimensional complex point clouds. Extensive evaluation shows that Argus achieves performance comparable to traditional solutions based on high-capability mmWave radars, with an average vertex error of 6.5 cm, solely using stripped-down radars deployed in a multi-view configuration. It presents robustness and practicality across conditions, such as with unseen users and different host devices.
Di Duan, Shengzhe Lyu, Mu Yuan, Hongfei Xue, Tianxing Li 0001, Weitao Xu, Kaishun Wu, Guoliang Xing
SenSys2