EDBT 2026 Demo / reviewers in the wild / expert
Liuyang Zhang
dblp:161/9058
· DBLP profile ↗
11ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Layer-wised Mixed-Precision CIM Accelerator with Bit-level Sparsity-aware ADCs for NAS-Optimized CNNsabstractExploring multiple precisions as well as sparsities for a computingin-memory (CIM) based convolutional accelerators is challenging. To further improve energy efficiency with minimal accuracy loss, this paper develops a neural architecture search (NAS) method to identify precision for each layer of the CNN and further leverages bit-level sparsity. The results indicate that following this approach, ResNet-18 and VGG-16 not only maintain their accuracy but also implement layer-wised mixed-precision effectively. Furthermore, there is a substantial enhancement in the bit-level sparsity of weights within each layer, with an average bit-level sparsity exceeding 90% per bit, thus providing broader possibilities for hardware-level sparsity optimization. In terms of hardware design, a mixed-precision (2/4/8-bit) readout circuit as well as a bit-level sparsity-aware Analog-to-Digital Converter (ADC) are both proposed to reduce system power consumption. Based on bit-level sparsity mixed-precision CNNs benchmarks, post-layout simulation results in 28nm reveal that the proposed accelerator achieves up to 245.72 TOPS/W energy efficiency, which shows about 2.52 -- 6.57× improvement compared to the state-of-the-art SRAM-based CIM accelerators. Haoxiang Zhou, Zikun Wei, Dingbang Liu, Liuyang Zhang, Chenchen Ding, Jiaqi Yang 0009, Wei Mao 0002, Hao Yu 0001 |
ASP-DAC | 4 |
| 2025 | SHWCIM: A Scalable Heterogeneous Workload Computing-in-Memory ArchitectureabstractThis study introduces HWCIM, a SRAM-based Computing-In-Memory core, and SHWCIM, a CIM-capable Coarse-Grained Reconfigurable Architecture, to enhance re-source utilization, multi-functionality, and on-chip memory size in SRAM-based CIM designs. Evaluated using the SMIC 55nm process, HWCIM achieves 1.6x lower power, 2.8x higher energy efficiency, and up to 4.1x smaller area compared to previous CIM and CGRA works. Additionally, SHWCIM delivers an average 105.9x speedup over existing CGRAs and consumes 2-5x less energy than the Nvidia A40 GPU on realistic workloads. Zhibiao Xue, Liuyang Zhang |
DATE | 4 |
| 2025 | A 28-nm 135.19 TOPS/W Bootstrapped-SRAM Compute-in-Memory Accelerator With Layer-Wise Precision and SparsityabstractArtificial intelligence (AI) edge devices demand high energy efficiency as well as inference accuracy. SRAM-based compute-in-memory (CIM) accelerators have great potential for power reduction but still need to exploit higher throughput and better linearity performance. To meet edge-AI computing demands by CIM works, it is crucial to optimize algorithms and parameters for specific circuit systems to achieve hardware acceleration. This work firstly employs neural network search (NAS) method to find out the layer-wise optimized precisions and sparsities for convolutional neural networks (CNNs). Then, a 144-Kb charge-domain signed mixed-precision (2/4/8-bit) CIM accelerator employing bootstrapped SRAM cells with 9-transistors and 1-capacitor (9T1C) structure is proposed that incorporates a bit-level sparsity-aware analog-to-digital converter (ADC). This work not only achieves highly linear parallel accumulation operations to meet AI computing demands but also implements a hardware and software co-optimization system tailored to specific data characteristics. The design is verified on NAS-optimized networks VGG-16 and ResNet-18 using Cifar-10 dataset, which could achieve an equivalent accuracy at 4-bit of 68.68% while maintaining a high energy efficiency at 2-bit of 135.19TOPS/W by measurements. Wei Mao 0002, Dingbang Liu, Haoxiang Zhou, Fuyi Li, Kai Li 0024, Qiuping Wu, Jiaqi Yang 0009, Liuyang Zhang, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2025 | Terahertz High-Contrast Imaging for Delamination Detection in Ultralayered Composites Based on Local SymmetryabstractTerahertz (THz) time-domain spectroscopy gains its popularity in internal defect detection of nonpolar dielectrics, and emerges as a promising inspection technique for delamination defects in composite materials due to its submillimeter level resolution and high penetrability. However, the contrast of THz images decreases significantly with an increasing number of layers due to dramatic signal attenuation and severe multiple reflections, which degrades the temporal and spatial resolution and subsequently hinders the accuracy of defect characterization. In this article, we propose a novel THz defect characterization framework for inspecting ultralayered glass fiber-reinforced polymer composites. The framework reconstructs the laminate structure by utilizing the local symmetry characteristic of the reflected THz pulses and the point cloud density. Then, THz images are extracted from both the reflected time-domain signal and local symmetry function. Pixel-level clustering of THz images is performed for accurate defect characterization. The defect intersection over union for 25 delamination defects in a 31-layer composite can reach more than 0.87. Our proposed strategy offers a promising route for accurate delamination assessment in ultralayered composites, and can be extended to automatic THz characterization in industrial applications. Yuqing Cui, Yafei Xu, Donghai Han, Liuyang Zhang, Ruqiang Yan 0001, Xuefeng Chen 0002 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | LAMPS: A Layer-wised Mixed-Precision-and-Sparsity Accelerator for NAS-Optimized CNNs on FPGAabstractThe increasing model size and computation load of convolutional neural networks (CNN) pose a grand challenge to deploy CNN models on edge computing devices. To further improve performance without significant accuracy loss, this paper developed a neural architecture search (NAS) method to achieve a layer-wise mixed-precision-and-sparsity (LAMPS) CNN. However, this optimization cannot be fully utilized and directly mapped to existing AI accelerators due to the irregu- lar computation of sparse and multi-precision data. To tackle this challenge, this work proposed a LAMPS vector systolic accelerator and demonstrated state-of-the-art results. Experi- mental results show that the LAMPS accelerator on Xilinx ZCU102 achieves an average performance of 756.83 GOPS and 470.25 GOPS when accelerating the NAS-optimized VGG16 and Resnet18, respectively, leading to 1.3-6.0x speed-up over the state- of-the-art accelerators on FPGA. Shuxin Yang, Chenchen Ding, Mingqiang Huang, Kai Li 0024, Chenghao Li 0010, Zikun Wei, Sixiao Huang, Jingyao Dong, Liuyang Zhang, Hao Yu 0001 |
FCCM | 9 |
| 2023 | Graph attention U-Net to fuse multi-sensor signals for long-tailed distribution fault diagnosis
Yuangui Yang, Tianfu Li, Chuang Sun 0001, Liuyang Zhang, Ruqiang Yan 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | Hybrid energy-efficient scheduling measures for flexible job-shop problem with variable machining speeds
Zhenzhen Wei, Wenzhu Liao, Liuyang Zhang |
Expert Syst. Appl. | 3 |
| 2020 | High-Density, Low-Power Voltage-Control Spin Orbit Torque Memory with Synchronous Two-Step Write and Symmetric Read TechniquesabstractVoltage-control spin orbit torque (VC-SOT) magnetic tunnel junction (MTJ) has the potential to achieve high-speed and low-power spintronic memory, owing to the adaptive voltage modulated energy barrier of the MTJ. However, the three-terminal device structure needs two access transistors (one for write operation and the other one for read operation) and thus occupies larger bit-cell area compared to two terminal MTJs. A feasible method to reduce area overhead is to stack multiple VC-SOT MTJs on a common antiferromagnetic strip to share the write access transistors. In this structure, high density can be achieved. However, write and read operations face problems and the design space is not sure given a strip length. In this paper, we propose a synchronous two-step multi-bit write and symmetric read method by exploiting the selective VC-SOT driven MTJ switching mechanism. Then hybrid circuits are designed and evaluated based a physics-based VC-SOT MTJ model and a 40nm CMOS design-kit to show the feasibility and performance of our method. Our work enables high-density, low-power, high-speed voltage-control SOT memory. Wang Kang 0001, Liuyang Zhang, He Zhang 0011, Brajesh Kumar Kaushik, Weisheng Zhao 0001 |
DATE | 3 |
| 2016 | Quantitative evaluation of reliability and performance for STT-MRAMabstractDue to its non-volatility, high access speed, ultra low power consumption and unlimited writing/reading cycles, STT-MRAM (Spin Transfer Torque Magnetic Random Access Memory) has emerged as the most promising candidate for the next generation universal memory. However, the process of commercialization of STT-MRAM is hampered by its poor reliability. Generally, these reliability issues are caused by the PVT (Process Variations, Voltage, and Temperature) of both MTJ (Magnetic Tunneling Junction) and transistor. Mitigation and alleviating the impacts of the intrinsic properties and PVT on STT-MRAM is a challenging work. This paper discusses the errors occurring in STT-MRAM resulting from its poor reliability, and analyzes the causes of such errors. To obtain a quantitative assessment of PVT impact on STT-MRAM reliability, we investigate three aspects: writing/reading operation error rate, power consumption and access delay of a single cell. This study is carried out on Cadence platform for 45 nm technology node and the PMA (Perpendicular Magnetic Anisotropy) MTJ model used in the investigation comes from SP INLIB. These quantitative information would be helpful for designing reliability enhancing strategies of STT-MRAM. Liuyang Zhang, Aida Todri, Wang Kang 0001, Youguang Zhang, Lionel Torres, Yuanqing Cheng, Weisheng Zhao 0001 |
ISCAS | 1 |
| 2016 | Further studies on impulsive consensus of multi-agent nonlinear systems with control gain error
Tiedong Ma, Liuyang Zhang |
Neurocomputing | 2 |
| 2016 | Impulsive consensus of multi-agent nonlinear systems with control gain error
Tiedong Ma, Liuyang Zhang, Feiya Zhao, Yixiang Xu |
Neurocomputing | 2 |