Yu Liu 0113

dblp:97/2274-113 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0001-9789-0959ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Multiplication-Free Floating-Point CIM Architecture with a 5-Bit Approximate Squaring Circuit
Xin Li 0099, Juntao Ge, Miao Long, Fugui Jiang, Chenglong Duan, Yang Yang 0025, Yu Liu 0113, Zhi-Ting Lin
ISCAS8
2026 A Floating-Point CIM Macro Featuring Asymmetric Exponent Encoding and Adaptive Mantissa Truncation for High-Efficiency AI Edge Computing
Zhi-Ting Lin, Rongtao Li, Yu Liu 0113, Xin Li 0099, Xiulong Wu
ISCAS6
2026 An offset-compensation capacitor-coupled DRAM sense amplifier with symmetric sensing and high PVT stability
Chenghu Dai, Jing Lv, Yongqi Qin, Chunyu Peng, Xin Li 0099, Yu Liu 0113, Xiulong Wu, Zhi-Ting Lin
Integr.7
2026 Analysis and Design of Memory Testing Algorithm for Computing-in-Memory Using MBIST
abstract
Computing-in-memory (CIM), as a novel computing architecture for the future, effectively overcomes the bottlenecks in the von Neumann architecture. The CIM architecture embeds logic into the memory array to reduce the data transfer between the processor and memory. However, embedding logic into the memory array increases the test complexity. In this study, we offer a comprehensive examination of the challenges associated with CIM and introduce a novel March-like test algorithm, named March CC, tailored for CIM chips. Computational elements are added to the read/write operation sequences, combining the tests in memory mode and computing mode into one step, which significantly improves the test efficiency. In comparison to the traditional March C− test algorithm, the proposed March CC test algorithm, with a complexity of only 10 N , enhances the fault coverage from 66.7% to 79.8% for six common single-cell fault (SCF) models and nine common double-cell fault (DCF) models. Furthermore, the March CC algorithm demonstrates good compatibility and is applicable to various memory configurations, such as SRAM, RRAM, and MRAM CIM architectures.
Zhi-Ting Lin, Siyan Li, Qiushi Feng, Changxin Yue, Yuanyang Wang, Yunlong Liu 0006, Yu Liu 0113, Licai Hao, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Xiulong Wu
ACM J. Emerg. Technol. Comput. Syst.9
2026 A 28 nm 1.3 TFLOPS/mm2 Floating-Point SRAM-Based CIM Macro With Asynchronous Normalization and Parallel Sorting Alignment for AI-Edge Chip
abstract
State-of-the-art AI edge devices require floating-point (FP) multiply-accumulate (MAC) operations with high-energy efficiency and inference accuracy. FP computing in-memory (FP-CIM) has a broader range of applications compared to integer CIM. However, FP-CIM can incur greater power, delay, and area overheads than integer CIM due to the inherent complexity of FP computational flow. In this article, we introduce a new method for asynchronous exponent normalization and parallel mantissa alignment. This approach allows us to add exponents and find the maximum sum simultaneously. We also replace the traditional subtraction and shifting for mantissa alignment with a cross-structure maximum-finding method, enabling FP-CIM to be achieved with lower delay, area, and power overheads. The macro is designed in TSMC 28 nm process, with a memory size of 6 Kb, a layout area of 0.067 mm2, and an area efficiency of 1.3 TFLOPS/mm2. Simulation results show that the macro computational frequency and energy efficiency can reach 150 MHz and 12.8 TFLOPS/W, respectively, at 900 mV, while performing FP- MAC operations.
Zhi-Ting Lin, Miao Long, Yang Yang 0025, Lintao Chen, Yu Liu 0113, Xin Li 0099, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.8
2025 A Floating-Point SRAM-based CIM Macro with Asynchronous Normalization and Parallel Sorting Alignment
abstract
Floating-point computing-in-memory (FP-CIM) has a broader range of applications compared to integer CIM. However, FP-CIM can incur greater power, delay, and area overheads than integer CIM due to a more complex computational flow. In this paper, we introduce a new method for asynchronous exponent normalization and parallel mantissa alignment. This approach allows us to add exponents and find the maximum sum simultaneously. We also replace the traditional subtraction and shifting for mantissa alignment with a time-cycle lookup method, enabling FP-CIM to be achieved with lower delay, area, and power overheads. The macro is designed in the TSMC 28nm process, with a memory size of 6Kb, a layout area of 0.067mm2, and an area efficiency of 1.3TFLOPS/mm2. Simulation results show that the macro computational frequency and energy efficiency can reach 150MHz and 12.8TFLOPS/W, respectively at 900mV.
Zhi-Ting Lin, Dongcheng Wang, Rongtao Li, Shichen Yu, Yu Liu 0113, Xin Li 0099, Xiulong Wu
ISCAS8
2025 TSCIM: A 28nm Transposed Stochastic CIM Macro for On-Chip Training and Inference
abstract
This work introduces a novel Transposed Stochastic Computing-in-Memory (TSCIM) macro designed to enhance the efficiency of on-chip training and inference. The macro incorporates a novel stochastic quantization strategy and utilizes a transposed separated wordline SRAM to enable multi-bit signed MAC operations. Furthermore, a stochastic adder tree is utilized to minimize area and power consumption overhead. The design includes a 4Kb SRAM CIM macro implemented in 28 nm CMOS technology. Simulation results show that the power consumption of the stochastic accumulation circuit (SAC) is reduced by 63.6%, while the area overhead is decreased by a factor of 7.73 compared to designs using full adder (FA) adder trees. Additionally, the computation latency is decreased by 16× compared to traditional stochastic circuits. The TSCIM macro can achieve a peak energy efficiency of 63.02 TOPS/W and an area efficiency of 15.54 TOPS/mm2.
Yu Liu 0113, Yang Lou, Kangkang Mao, Xin Li 0099, Chenghu Dai, Xiulong Wu, Zhi-Ting Lin
ISCAS1
2025 A 28-nm 9T1C SRAM-Based CIM Macro With Hierarchical Capacitance Weighting and Two-Step Capacitive Comparison ADCs for CNNs
abstract
In the realm of charge-domain computing-in-memory (CIM) macros, reducing the area of capacitor ladder and analog-to-digital converter (ADC) while maintaining high throughput remains a significant challenge. This brief introduces an adjustable-weight CIM macro designed to enhance both energy efficiency and area efficiency for convolutional neural networks (CNNs). The proposed architecture uses: 1) a customized 9T1C bit cell for sensing margin improvement and bidirectional decoupled read ports; 2) a hierarchical capacitance weighting (HCW) structure that achieves a weight accumulation of 1/2/4 bits with less capacitance area and weighting time; and 3) a two-step capacitive comparison ADCs (TC-ADCs) readout scheme to improve area efficiency and throughput. The proposed 8-kb static random address memory (SRAM) CIM macro is implemented using 28-nm CMOS technology. It can achieve an energy efficiency of 224.4 TOPS/W and an area efficiency of 21.894 TOPS/mm2, and the accuracies on MNIST, CIFAR-10, and CIFAR-100 datasets are 99.67%, 89.13%, and 67.58% with a 4-b input and 4-b weight.
Zhi-Ting Lin, Runru Yu, Miao Long, Yu Liu 0113, Jianxing Zhou, Qingchuan Zhu, Yue Zhao 0029, Lintao Chen, Chunyu Peng, Qiang Zhao 0007, Xin Li 0099, Chenghu Dai, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.5
2024 Domain-aware double attention network for zero-shot sketch-based image retrieval with similarity loss
Ming Zhu 0016, Nian Wang 0001, Feiyang Gu, Yu Liu 0113, Xin Li 0099
Vis. Comput.5
2022 TSDLPP: A Novel Two-Stage Deep Learning Framework For Prognosis Prediction Based on Whole Slide Histopathological Images
abstract
Recently, digital pathology image-based prognosis prediction has become a hot topic in healthcare research to make early decisions on therapy and improve the treatment quality of patients. Therefore, there has been a recent surge of interest in designing deep learning method solving the problem of prognosis prediction with digital pathology images. However, whole slide histopathological images (WSIs) based prognosis prediction is still a challenge due to the large size of pathological images, the heterogeneity of tumors and the high cost of region of interests (ROIs) labeling. In this study, we design a novel two-stage deep learning framework for prognosis prediction (TSDLPP) based on WSIs. Our proposed framework consists of two-stage paradigms: 1) training tissue decomposition network (TDNet) to divide WSIs into cancerous and non-cancerous regions, 2) integrating general prognosis-related densely connected CNN (GPR-DCCNN) and morphology-specific prognosis-related densely connected CNNs (MSPR-DCCNNs) to extract different level features of pathological images. In the end, we apply TSDLPP to the prognosis prediction of breast cancer using The Cancer Genome Atlas (TCGA) datasets. Experiment results demonstrate that TSDLPP obtains superior performance of prognosis prediction compared with the existing state-of-arts methods.
Yu Liu 0113, Ao Li 0001, Jiangshu Liu, Gang Meng
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 DeepPhos: prediction of protein phosphorylation sites with deep learning
abstract
MOTIVATION: Phosphorylation is the most studied post-translational modification, which is crucial for multiple biological processes. Recently, many efforts have been taken to develop computational predictors for phosphorylation site prediction, but most of them are based on feature selection and discriminative classification. Thus, it is useful to develop a novel and highly accurate predictor that can unveil intricate patterns automatically for protein phosphorylation sites. RESULTS: In this study we present DeepPhos, a novel deep learning architecture for prediction of protein phosphorylation. Unlike multi-layer convolutional neural networks, DeepPhos consists of densely connected convolutional neuron network blocks which can capture multiple representations of sequences to make final phosphorylation prediction by intra block concatenation layers and inter block concatenation layers. DeepPhos can also be used for kinase-specific prediction varying from group, family, subfamily and individual kinase level. The experimental results demonstrated that DeepPhos outperforms competitive predictors in general and kinase-specific phosphorylation site prediction. AVAILABILITY AND IMPLEMENTATION: The source code of DeepPhos is publicly deposited at https://github.com/USTCHIlab/DeepPhos. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fenglin Luo, Yu Liu 0113, Xing-Ming Zhao, Ao Li 0001
Bioinform.3