VLDB 2026 Research / reviewers in the wild / expert
Xin Li 0099
dblp:09/1365-99
· DBLP profile ↗
15ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-0125-5254ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Multiplication-Free Floating-Point CIM Architecture with a 5-Bit Approximate Squaring Circuit
Xin Li 0099, Juntao Ge, Miao Long, Fugui Jiang, Chenglong Duan, Yang Yang 0025, Yu Liu 0113, Zhi-Ting Lin |
ISCAS | 1 |
| 2026 | A Floating-Point CIM Macro Featuring Asymmetric Exponent Encoding and Adaptive Mantissa Truncation for High-Efficiency AI Edge Computing
Zhi-Ting Lin, Rongtao Li, Yu Liu 0113, Xin Li 0099, Xiulong Wu |
ISCAS | 7 |
| 2026 | An offset-compensation capacitor-coupled DRAM sense amplifier with symmetric sensing and high PVT stability
Chenghu Dai, Jing Lv, Yongqi Qin, Chunyu Peng, Xin Li 0099, Yu Liu 0113, Xiulong Wu, Zhi-Ting Lin |
Integr. | 6 |
| 2026 | A 28 nm 1.3 TFLOPS/mm2 Floating-Point SRAM-Based CIM Macro With Asynchronous Normalization and Parallel Sorting Alignment for AI-Edge ChipabstractState-of-the-art AI edge devices require floating-point (FP) multiply-accumulate (MAC) operations with high-energy efficiency and inference accuracy. FP computing in-memory (FP-CIM) has a broader range of applications compared to integer CIM. However, FP-CIM can incur greater power, delay, and area overheads than integer CIM due to the inherent complexity of FP computational flow. In this article, we introduce a new method for asynchronous exponent normalization and parallel mantissa alignment. This approach allows us to add exponents and find the maximum sum simultaneously. We also replace the traditional subtraction and shifting for mantissa alignment with a cross-structure maximum-finding method, enabling FP-CIM to be achieved with lower delay, area, and power overheads. The macro is designed in TSMC 28 nm process, with a memory size of 6 Kb, a layout area of 0.067 mm2, and an area efficiency of 1.3 TFLOPS/mm2. Simulation results show that the macro computational frequency and energy efficiency can reach 150 MHz and 12.8 TFLOPS/W, respectively, at 900 mV, while performing FP- MAC operations. Zhi-Ting Lin, Miao Long, Yang Yang 0025, Lintao Chen, Yu Liu 0113, Xin Li 0099, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2025 | A Floating-Point SRAM-based CIM Macro with Asynchronous Normalization and Parallel Sorting AlignmentabstractFloating-point computing-in-memory (FP-CIM) has a broader range of applications compared to integer CIM. However, FP-CIM can incur greater power, delay, and area overheads than integer CIM due to a more complex computational flow. In this paper, we introduce a new method for asynchronous exponent normalization and parallel mantissa alignment. This approach allows us to add exponents and find the maximum sum simultaneously. We also replace the traditional subtraction and shifting for mantissa alignment with a time-cycle lookup method, enabling FP-CIM to be achieved with lower delay, area, and power overheads. The macro is designed in the TSMC 28nm process, with a memory size of 6Kb, a layout area of 0.067mm2, and an area efficiency of 1.3TFLOPS/mm2. Simulation results show that the macro computational frequency and energy efficiency can reach 150MHz and 12.8TFLOPS/W, respectively at 900mV. Zhi-Ting Lin, Dongcheng Wang, Rongtao Li, Shichen Yu, Yu Liu 0113, Xin Li 0099, Xiulong Wu |
ISCAS | 9 |
| 2025 | TSCIM: A 28nm Transposed Stochastic CIM Macro for On-Chip Training and InferenceabstractThis work introduces a novel Transposed Stochastic Computing-in-Memory (TSCIM) macro designed to enhance the efficiency of on-chip training and inference. The macro incorporates a novel stochastic quantization strategy and utilizes a transposed separated wordline SRAM to enable multi-bit signed MAC operations. Furthermore, a stochastic adder tree is utilized to minimize area and power consumption overhead. The design includes a 4Kb SRAM CIM macro implemented in 28 nm CMOS technology. Simulation results show that the power consumption of the stochastic accumulation circuit (SAC) is reduced by 63.6%, while the area overhead is decreased by a factor of 7.73 compared to designs using full adder (FA) adder trees. Additionally, the computation latency is decreased by 16× compared to traditional stochastic circuits. The TSCIM macro can achieve a peak energy efficiency of 63.02 TOPS/W and an area efficiency of 15.54 TOPS/mm2. Yu Liu 0113, Yang Lou, Kangkang Mao, Xin Li 0099, Chenghu Dai, Xiulong Wu, Zhi-Ting Lin |
ISCAS | 4 |
| 2025 | MTJ based temperature compensated beta multiplier Voltage ReferenceabstractThis article mainly explores the interaction between CMOS devices and Magnetic Tunnel Junction (MTJ) devices in terms of temperature characteristics, aiming to achieve a CMOS beta multiplier circuit that combines low power consumption and wide temperature adaptability, making it a stable reference voltage source. The proposed design utilizes the Tunneling Magneto Resistance (TMR) effect of MTJ to compensate for the performance mismatch caused by temperature changes in CMOS. This design adopts TSMC 28nm CMOS craft, which can generate a reference voltage with a linearity of 0.57%/V, a temperature coefficient of 43.6ppm/°C, and a stable voltage of 441.6mV at a minimum supply voltage of 0.6V and a temperature range of 10~110 °C. The noise of this voltage at a frequency of 10Hz is 69.7uV/sqrt (Hz), the power suppression ratio is -62.1dB, and the power consumption is 4.092nW. Yongliang Zhou, Yingxue Sun, Jingxue Zhong, Chengxing Dai, Weizhe Tan, Chunyu Peng, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu |
ISCAS | 7 |
| 2025 | MTJ based Temperature-Adaptive VCO (TAVCO) for Compensating CP-PLL Frequency DriftabstractThe Charge Pump Phase-Locked Loop (CP-PLL) is a commonly utilized component in contemporary mixed-signal electronic systems. It is widely employed for clock generation, synchronization, and frequency synthesis in both digital and wireless functionalities. However, the frequency accuracy of oscillators can be adversely affected by variations in frequency across a broad temperature range. To address this issue, the voltage-controlled oscillator designed in this study employs a four-stage differential delay structure, chosen for its simple circuit architecture, favorable control linearity, and low noise characteristics. This research integrates the temperature behaviors of Complementary Metal-Oxide-Semiconductor (CMOS) and Magnetic Tunnel Junction (MTJ) technologies, utilizing 28nm CMOS technology to enhance the frequency stability of ring oscillators effectively. Simulation results indicate that frequency drift is reduced by 92% within the temperature range of -80°C to 125°C. Yongliang Zhou, Jingxue Zhong, Yingxue Sun, Chengxing Dai, Weizhe Tan, Chunyu Peng, Wenjuan Lu, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu |
ISCAS | 8 |
| 2025 | Full-Array Boolean Logic CIM Macro With Self-Recycling 10T-SRAM Cell for AES SystemsabstractComputing in memory (CIM), which alleviates the need to transfer a large amount of data between processor and memory, significantly reducing latency and energy consumption, is a promising new computing architecture for addressing the von Neumann bottleneck problem. This article proposes a CIM array structure composed of self-recycling 10T static random access memory (SRAM) cells, which can realize orthogonal data writing, and multiple Boolean logical operations for the entire array. The self-recycling and full-array activation characteristics are extremely suitable for accelerating diverse data processing algorithms such as the Advanced Encryption Standard (AES). A 4-kb SRAM is implemented in 55-nm CMOS technology to verify the effectiveness of the design. Compared with other state-of-the-art architectures, the throughput and the operating frequency of the proposed CIM macro are increased to 843 GOPS/kb ($2.64\times $) and 823.7 MHz ($2.6\times $), respectively. The energy efficiency reaches 246.9 TOPS/W. When applied to the AES, the energy consumption is 35.77% less than the digital CIM architecture that is not self-recycling. Xin Li 0099, Lintao Chen, Yang Lou, Baofa Wu, Jiajun Long, Yongliang Zhou, Chunyu Peng, Xiulong Wu, Zhi-Ting Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | A 28-nm 9T1C SRAM-Based CIM Macro With Hierarchical Capacitance Weighting and Two-Step Capacitive Comparison ADCs for CNNsabstractIn the realm of charge-domain computing-in-memory (CIM) macros, reducing the area of capacitor ladder and analog-to-digital converter (ADC) while maintaining high throughput remains a significant challenge. This brief introduces an adjustable-weight CIM macro designed to enhance both energy efficiency and area efficiency for convolutional neural networks (CNNs). The proposed architecture uses: 1) a customized 9T1C bit cell for sensing margin improvement and bidirectional decoupled read ports; 2) a hierarchical capacitance weighting (HCW) structure that achieves a weight accumulation of 1/2/4 bits with less capacitance area and weighting time; and 3) a two-step capacitive comparison ADCs (TC-ADCs) readout scheme to improve area efficiency and throughput. The proposed 8-kb static random address memory (SRAM) CIM macro is implemented using 28-nm CMOS technology. It can achieve an energy efficiency of 224.4 TOPS/W and an area efficiency of 21.894 TOPS/mm2, and the accuracies on MNIST, CIFAR-10, and CIFAR-100 datasets are 99.67%, 89.13%, and 67.58% with a 4-b input and 4-b weight. Zhi-Ting Lin, Runru Yu, Miao Long, Yu Liu 0113, Jianxing Zhou, Qingchuan Zhu, Yue Zhao 0029, Lintao Chen, Chunyu Peng, Qiang Zhao 0007, Xin Li 0099, Chenghu Dai, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 13 |
| 2025 | A High-Performance and High-Robustness Triple-Node-Upset Tolerant Latch Based on Redundant-Node HardeningabstractIn response to the issues of high cost, large overhead, and limited node fault tolerance in current latch hardening techniques, this article proposes a latch circuit resistant to triple-node-upset (TNU) based on redundant-node hardening technology. This latch comprises eight 1P2N modules interlocked, with its output isolated by two levels of C-elements (CEs), achieving tolerance to TNU. The performance of the redundant-node reinforcement TNU tolerant latch (RNRTTL) was simulated and verified using CMOS 65 nm technology. The simulation results indicate that the RNRTTL circuit has a D-Q delay of 14.14 ps, static power consumption of$4.03~\mu $w, an area of$32.87~\mu $m2, and an area-static power-D–Q delay-product (APDP) of 1873, respectively. Compared to the triple-node upset tolerant latches TTLL, TNU-latch, TNURL, and HLTNURL reported in the current literature, the proposed latch demonstrates an average reduction of 219.9%, 164.9%, 150.7%, and 2464.8% in D-Q delay, static power consumption, area, and APDP, respectively, indicating that the RNRTTL latch has superior comprehensive performance; furthermore, a series of 2000 Monte Carlo (MC) simulations on the node group$\langle $Q, X0, X$8\rangle $reveal that the proposed latch circuit possesses good stability, making it suitable for harsh radiation environments. Qiang Zhao 0007, Qingyi Liu, Licai Hao, Xin Li 0099, Shengyue Zhang, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | A Timing-Shared Adaptive Sensing Methodology for Low-Voltage SRAMabstractLowing static random access memory (SRAM) supply voltage could highly improve energy efficiency, yet energy efficiency still not attains optimal point due to the constraint of the weakest bit-cell, especially in low-voltage SRAM. Adaptive sensing methodology is proposed for the challenge and consists of four elements: Switch unit is built to implement cross-sensing operation, Timing-Shared Decoupling Latch Sense Amplifier (TS-DLSA) allows rapid and successive sensing, the judging module is utilized to trigger the FLAG signal which is for adaptive timing controller to cut off word line (WL). Energy efficiency was obtained by compressing the activation delay of WL compared to global timing scheme. The proposed adaptive sensing methodology was performed with TSMC 28-nm CMOS process, and evaluated in SRAM array of 128x128, 256x256, 512x512, and 1024x1024. Monte Carlo simulation results are formed to confirm that, compared to the global timing scheme, the proposed adaptive sensing methodology has reduced WL activation delay by 75.1%∼41.8% and read operation energy overhead by 77.5%∼30.9% from 0.6V to 1.2V. Compared to the current-latched sense amplifier with a footswitch (FS-CLSA) with proposed adaptive sensing methodology, TS-DLSA with proposed scheme have reduced the WL activation by 4.1%∼10.5% and read operation energy overhead by 15.4%∼23.8% from 256x256 to 2048x2048 at 0.6V. The more cells mounted on the BL, the higher energy revenue gains. Yongliang Zhou, Saiai Wu, Wenjuan Lu, Chunyu Peng, Xin Li 0099, Xiulong Wu |
ISCAS | 8 |
| 2024 | Domain-aware double attention network for zero-shot sketch-based image retrieval with similarity loss
Ming Zhu 0016, Nian Wang 0001, Feiyang Gu, Yu Liu 0113, Xin Li 0099 |
Vis. Comput. | 6 |
| 2021 | Two-Stage Difference-Based Estimation Method for Timing Skew in TI-ADCsabstractA two-stage difference-based estimation method is presented to extract the timing skew of each sub-ADC in time interleaved analog to digital converters (TIADCs). Compared with the previous techniques, such as the correlation-based method, the main advantage is that it greatly reduces the number of multipliers and does not require digital filters. Simulation results of a four-channel TI-ADC behavioral model and experimental results from a commercial 12-bit 3.6-GS/s TI- ADC demonstrate the effectiveness and superiority of the proposed estimation method. Xin Li 0099, Christian Vogel 0001, Jianhui Wu 0001 |
ISCAS | 1 |
| 2020 | Active Noise Shaping SAR ADC Based on ISDM with the 5MHz BandwidthabstractThis paper reports a hybrid noise shaping (NS) successive-approximation registers (SAR) ADC based on incremental sigma-delta modulator (ISDM). The combination of ISDM and NS-SAR ADC achieves the high signal-to-noise ratio (SNR) with a few extra timing and hardware penalty. The reuse of integrator for both ISDM and NS releases the demands of high-resolution multi-input comparator replaced by only a 2-input one, which also helps suppress the noise from comparator and clock to the top plate of capacitor digital-to-analog converter (CDAC) relatively. With the embedded ISDM, the gain of amplifier used in the finite impulse response (FIR) filter for NS is relaxed remarkably, and wider bandwidth (BW) as well as lower oversampling rate (OSR) is realized compared to those traditional NS-SAR ADCs. The proposed hybrid ADC is designed in 40nm CMOS process with a 1.1V supply. It achieves the Signal-to-Noise Distortion Ratio (SNDR) and the Spurious Free Dynamic Range (SFDR) of 82.4dB and 97.1dBc respectively at 80MS/s, with a signal bandwidth (BW) of 5MHz and a total power consumption of 883uW. Xin Li 0099, Chenggang Yan 0002, Jianhui Wu 0001 |
ISCAS | 2 |