EDBT 2026 Demo / reviewers in the wild / expert
Ziyang Shen
dblp:254/2557
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0004-4180-6171ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Heterogeneous Decision Spiking Transformer Accelerator with Locality-dependent KV Product Cache and Compute Pattern Reconfigurable Engine
Ziyang Shen, Zhipeng Liao, Sitan Shen, Chaoming Fang, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 1 |
| 2026 | Breaking the I/O Bottleneck: I/O Coordination Optimization for Efficient Large-Scale LLM Fine-TuningabstractLarge Language Models (LLMs) with tens or even hundreds of billions of parameters have become the foundation of modern AI applications. However, fine-tuning such massive models is severely constrained by the limited GPU memory. Existing memory-saving systems, such as ZeRO-based offloading in DeepSpeed, reduce GPU memory usage but inevitably incur substantial I/O overhead, especially when model states reside on slow storage devices, such as NVMe SSDs. As a result, the memory bottleneck in large-scale fine-tuning is transformed into an I/O bottleneck. Although prior systems have employed strategies like parameter prefetching and partial asynchronous execution, they remain limited by synchronous I/O-communication dependencies and the lack of fine-grained read/write I/O scheduling. To address these limitations, we propose IOC, an I/O Coordination Optimization framework that maximizes pipeline parallelism across different phases of LLM fine-tuning. IOC introduces three key mechanisms: (1) An All-Gather prefetching technique based on an I/O state hash table, which completely decouples All-Gather prefetching from parameter I/O, achieving continuous overlap among I/O, communication, and computation; (2) The parameter update phase is refactored into an asynchronous pipeline with explicit I/O isolation, where the optimizer state write-back is executed in a semi-asynchronous manner, thereby mitigating read/write contention and reducing synchronization stalls; (3) Multi-disk parallelism is leveraged by introducing an additional disk to further relieve I/O contention and defer synchronization waits to the latest possible time point. Experimental results demonstrate that IOC significantly accelerates LLM finetuning while preserving low memory consumption. The end-toend fine-tuning time on the Llama-70B model is reduced by $\mathbf{2 1. 5 \%}$ and 34.3% in single-disk and multi-disk configurations compared to the baseline. Ziyang Shen, Hongchao Du, Kaihuan Lin, Yin Lin, Qiao Li 0001, Chun Jason Xue |
ISPASS | 1 |
| 2025 | An Area-Efficient and Bit-Width Configurable Carry-Save Adder Tree for Spiking TransformersabstractSpiking transformers have been successfully applied to multiple applications with comparable accuracy with native transformers. Achieving a high energy efficiency with spiking transformers requires dedicated hardware design, especially a specialized matrix multiplication engine optimized for spike input. In this paper, we propose a carry-save adder (CSA) array with an improved energy and area efficiency for spiking transformer matrix multiplication computation. A two-staged CSA structure is proposed to support maximum logic reuse between 1b self-attention mode and 8b linear mode. Besides, a cubic-mesh architecture is proposed to organize CSA trees to reuse weight in different timesteps. Compared to a baseline accumulation design, the proposed architecture achieved a 2.7x area reduction and a 3.75x power reduction with the same throughput, showing that the optimized computation array has great potential to be applied in digital neuromorphic accelerators. Chaoming Fang, Ziyang Shen, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 2 |
| 2023 | NBSSN: A Neuromorphic Binary Single-Spike Neural Network for Efficient Edge IntelligenceabstractNeuromorphic computing approaches such as Spiking Neural Networks (SNN) have been increasingly adopted in bio-signal processing and interpretation due to its intrinsic neurodynamic attribute. Nevertheless, reconciling performance and power efficiency in SNN implementation is still a bottleneck. Single-spike neural coding scheme, which is an extremely sparse coding scheme, provides a solution to bridge the gap. In this work, a neuromorphic architecture, using binary single spike neural signals, is proposed with both algorithm and hardware implementation. A sparsity-aware spatial-temporal back-propagation training method is proposed together with a single-spike coding scheme. Also, a novel neuromorphic accelerator is co-designed with algorithmic optimization and implemented in 40nm CMOS process. Experimental results show that the proposed processor reaches an accuracy of 94.61% on the MNIST dataset, 93.59% on the N-MNIST dataset, and 93.27% on the ECG dataset, respectively, while consumes$0.173\mu\mathrm{J}$per ECG classification task and 0.16mm2on-chip area. The overall power consumption is reduced by 91.68% compared to the state-of-the-art systems. Ziyang Shen, Fengshi Tian, Chaoming Fang, Xiaoyong Xue, Jie Yang 0033, Mohamad Sawan |
ISCAS | 1 |
| 2022 | A Compact Online-Learning Spiking Neuromorphic Biosignal ProcessorabstractReal-time biosignal processing on wearable devices has attracted worldwide attention for its potential in healthcare applications. However, the requirement of low-area, low-power and high adaptability to different patients challenge conventional algorithms and hardware platforms. In this design, a compact online learning neuromorphic hardware architecture with ultralow power consumption designed explicitly for biosignal processing is proposed. A trace-based Spiking-Timing-Dependent-Plasticity (STDP) algorithm is applied to realize hardware-friendly online learning of a single-layer excitatory-inhibitory spiking neural network. Several techniques, including event-driven architecture and a fully optimized iterative computation approach, are adopted to minimize the hardware utilization and power consumption for the hardware implementation of online learning. Experiment results show that the proposed design reaches the accuracy of 87.36% and 83% for the Mixed National Institute of Standards and Technology database (MNIST) and ECG classification. The hardware architecture is implemented on a Zynq-7020 FPGA. Implementation results show that the Look-Up Table (LUT) and Flip Flops (FF) utilization reduced by 14.87 and 7.34 times, respectively, and the power consumption reduced by 21.69% compared to state of the art. Chaoming Fang, Ziyang Shen, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 2 |