EDBT 2026 Demo / reviewers in the wild / expert
Qinzhe Zhi
dblp:389/7233
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0001-4866-7213ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DayPQ: Dynamic Layerwise Pruning and Quantization for LLM Inference AccelerationabstractThe deployment of the prevailing large language models (LLMs) on resource-constrained hardware faces critical challenges due to their massive computational and memory requirements. Pruning and quantization (PQ), as two predominant techniques for LLM acceleration, face the challenge of fully preserving model capabilities while achieving significant acceleration in practice. Static PQ irreversibly discards potentially critical parameters and degrades precision across diverse inputs, which must be compensated through fine-tuning. Dynamic PQ preserves accuracy but incurs prohibitive hardware overhead and additional latency from runtime adaptations. Achieving low-latency and overhead-efficient dynamic PQ requires the co-design of adaptive compression tightly integrated with hardware. In this work, we present DayPQ, an algorithm-hardware co-optimization technique for efficient LLM inference: self-adaptive norm-based pruning and bias-shifting FP8 quantization at the algorithm level, synergized with a multiplication-free shift–add PE array and switchable dataflow at the hardware level, achieving dynamic acceleration with the negligible accuracy loss and zero additional latency. DayPQ enables adaptive PQ by intrinsically exploiting normalization operations from the model’s inference process instead of extra operations. Simultaneously, it establishes dynamic orthogonality between sparsity and quantization, where these dual compression mechanisms synergistically enhance the redundancy reduction beyond mere additive effects of independent operations. Our evaluation shows that DayPQ obtains an average$3.13\times $higher energy efficiency and$2.03\times $performance boost than three state-of-the-art (SOTA) accelerators. Yongchao Xu, Ziying Zhuo, Qinzhe Zhi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Local-GS: An Order-Independent Gaussian Splatting Training Accelerator Exploiting Splat Localityabstract3D Gaussian Splatting has emerged as the SOTA approach for 3D representation and view synthesis. While Gaussian Splatting has demonstrated impressive capability and rendering quality on desktop GPUs, achieving on-demand training on resource-constrained edge devices is still challenging. In this work, we identified the training bottleneck from a few perspectives including algorithm splat locality and the limited memory and hardware under-utilization. To address these problems, we present Local-GS, a 3D Gaussian Splatting training accelerator with order-independent rendering to break the depth-wise data dependency between overlapping Gaussians. We further incorporate a parallel pixel intersection test unit to schedule thread workload based on Gaussian splat locality and improve hardware utilization. A set of unified training-rendering cores are designed to achieve efficient splat-level parallel rendering and gradient propagation. Our Local-GS is implemented in 7 nm and is evaluated by several real-world 3D scenes. Compared to edge Jetson NX GPU, Local-GS achieve 26.9-53 $\times$ training speedup and three orders of magnitude efficiency boost. Qinzhe Zhi, Yiqi Jing, Le Ye, Ru Huang 0001 |
DAC | 2 |
| 2025 | GenSoC: A Multi-Agent-Assisted SoC Generation Methodology Leveraging Open-Source HardwareabstractThe complexity and heterogeneity of system-on-chip (SoC) architecture keep rapidly growing and require prolonged design cycle and cost. Recent advancements in large language models (LLMs) have opened up new avenues for agile design. In this work, we present an LLM-based multi-agent assisted SoC design methodology, which utilizes LLM agents to automatically and intelligently select, integrate, and verify SoC design. We first constructed a comprehensive IP library by retrieving existing open-source IPs as foundation design resource. By leveraging collaborative LLM multi-agent, each agent is pre-configured with unique guidelines and toolsets for SoC design steps including IP selection, SoC integration, and verification. Our methodology has been applied on two SoC design cases. The generated SoCs can achieve notable up to 27.18× and 29.67× energy efficiency improvements respectively compared to SoCs generated by existing open-source platforms. Peiran Yan, Qinzhe Zhi, Lifeng Liu |
ISLPED | 2 |
| 2024 | SPARK: An Efficient Hybrid Acceleration Architecture with Run-Time Sparsity-Aware Scheduling for TinyML LearningabstractCurrently most TinyML devices only focus on inference, as training requires much more hardware resources. In this paper, we introduce SPARK, an efficient hybrid acceleration architecture with run-time sparsity-aware scheduling for TinyML learning. Besides a standalone accelerator, an in-pipeline acceleration unit is integrated within the CPU pipeline to support simultaneous forward and backward propagation. To better utilize sparsity and improve hardware utilization, a sparsity-aware acceleration scheduler is implemented to schedule the workload between two acceleration units. A unified memory system is also constructed to support transposable data fetch, reducing memory access. We implement SPARK using TSMC 22nm technology and evaluate different TinyML tasks. Compared with the baseline accelerator, SPARK achieves 4.1× performance improvement in average with only 2.27% area overhead. SPARK also outperforms off-shelf edge devices in performance by 9.4× with 446.0× higher efficiency. Qinzhe Zhi, Yanchi Dong, Le Ye |
DAC | 2 |