EDBT 2026 Demo / reviewers in the wild / expert
Ming Liu 0022
dblp:20/2039-22
· DBLP profile ↗
24ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-0937-7547ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An RRAM-based Multi-Timescale Spiking Processor with Reconfigurable Neurons
Jinhao Liang, Fangduo Zhu, Siyuan Ouyang, Jingsong Zhang, Xumeng Zhang, Qi Liu 0010, Ming Liu 0022 |
ISCAS | 9 |
| 2026 | A 13-GS/s 9-bit Time-Interleaved Pipelined-SAR ADC With Common-Mode Regulated Floating-Inverter-Amplifier and Rapid-Tracking Bootstrapped SwitchabstractThis article presents a 13GS/s 9-bit 8-channel time-interleaved (TI) Pipelined-SAR (Pipe-SAR) ADC. A common-mode regulated floating-inverter-amplifier (CMR-FIA) is proposed to overcome the common-mode voltage variation due to the charge leakage through the parasitic capacitor, thereby eliminating the need for common-mode feedback (CMFB) circuitry embedded in the body of FIA. By combining an adaptively biased (AB) technique, the proposed FIA facilitates the high-speed and robust Pipe-SAR ADCs. In addition, a rapid-tracking bootstrapped signal generation is introduced to achieve high-linearity with short sampling time in an ultra-high speed sampling network. The ADC prototype is fabricated is a 28nm-CMOS process, the achieved spurious-free dynamic range (SFDR) and signal-to-noise and distortion ratio (SNDR) at the Nyquist input are 56.4dB and 41.75dB, respectively. Consuming 97mW at 13GS/s, it yields a Schreier figure of merit ($\text{FoM}_{\mathrm {S}}$) of 150dB. With the proposed CMR-FIA, the ADC’s SNDR variation is within 2.48dB across the input common-mode range of 0.3V to 0.8V. Ji Guo, Danfeng Zhai, Wenning Jiang, Qi Liu 0010, Ming Liu 0022 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2026 | A 1024-Ch 583-nW/Ch Spike-Sorting SoC With Sparsity-Aware Spike Detection Scratchpad and Ultra-Low-Leakage Dual-Voltage 5T-SRAM for 16K-Template ClusteringabstractThis paper presents an energy-efficient spike-sorting system-on-chip (SoC) designed for closed-loop brain-computer interfaces of massive probing channels. The design first incorporates a sparsity/similarity-aware spike detection scratchpad, leveraging a bit-wise differential encoder and zero-friendly read-out circuits, reducing the dynamic power consumption of spike detection by 77.7%. To mitigate static power dissipation, it also introduces an ultra-low-leakage dual-voltage 5T-SRAM array with level-shifter embedded sense amplifiers, achieving an 82.2% leakage power reduction of neural signal buffering by applying half$V_{DD}$on SRAM cells. Additionally, a memory hierarchy architecture combining on-chip SRAM and off-chip FeRAM, along with a firing-rate-based Osort for cluster template management, minimizes off-chip memory access to only 9.7% with a latency of$11.7\mu $s for 1024-channel spike sorting. A silicon prototype is fabricated in 28-nm CMOS technology, which achieves a power consumption of 583nW/channel and an area consumption of 0.0012mm2/channel. The chip supports real-time spike sorting with up to 16K templates,$21.3\times $greater than the state-of-the-art spike-sorting processor. Hao Jiang 0024, Zexing Chen, Jiajun Lu, Siqi He, Liangjian Lyu, Jiamin Xu, Shiwei Liu 0002, Yingping Chen, Chixiao Chen, Qi Liu 0010, Ming Liu 0022 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 12 |
| 2025 | An Energy-Efficient High-Utilization Hardware Architecture for Attention Mechanism in Transformer using Balanced Systolic Array and Multi-Row Interleaved Operation OrderingabstractTransformer-based neural networks have achieved remarkable performance. Designing energy-efficient and high-speed accelerators for the attention mechanism, which dominates the energy and latency in Transformers, has become increasingly significant. Existing attention accelerators commonly use algorithm-hardware co-design to achieve higher energy efficiency and speed. However, deeply customized algorithms make these accelerators dependent on a particular application. Therefore, optimizing hardware architecture is crucial for achieving general-purpose acceleration. We observe two limitations in the hardware architecture of existing attention accelerators. First, the widely used input stationary, weight stationary, and output stationary systolic arrays (SAs) can’t balance data reuse, register saving, and utilization, which hinders to build more energy-efficient and faster SA-based accelerators. Second, layer-by-layer operation ordering introduces high SRAM access overhead of intermediate results. To address the first limitation, we propose the “Balanced Systolic Array”, which improves energy efficiency by 40% compared to conventional systolic arrays and achieves a utilization rate of 99.5%. To address the second limitation, we propose “Multi-Row Interleaved” operation ordering, which reduces the SRAM energy by 31.7% By integrating two techniques, the proposed attention accelerator achieves a 39% improvement in energy efficiency and a 38% enhancement in throughput×energy efficiency compared to previous works. Haiyang Zhou, Hongyang Hu, Jinshan Yue, Hanghang Gao, Yuanlu Xie, Xiaoxin Xu, Chunmeng Dou, Ming Liu 0022 |
DAC | 8 |
| 2025 | EIGEN: Enabling Efficient 3DIC Interconnect with Heterogeneous Dual-Layer Network-on-Active-InterposerabstractChiplet-based 3DICs have emerged as modular solutions for large-scale, high-performance computing systems. However, unlike monolithic Network-on-Chips (NoCs), 3DICs using active interposers encounter difficulties in managing heterogeneous and high-load network traffic patterns. The emerging field of Networks-on-Active-Interposer (NOAI), independently designed from top dies’ interconnects, is desired to address these challenges by supporting flexible topologies, low-latency memory access, and reduced traffic congestion. To satisfy the requirements, we propose a heterogeneous dual-layer interconnect architecture, EIGEN, for chiplet-interposer systems, along with a reinforcement learning (RL)-based routing framework, which can provide efficient and flexible communication for chiplet-based 3DICs. EIGEN features an application-aware switch-programable interconnection layer (AspLayer) and a dynamic packet-routing interconnection layer (DynLayer) on the interposer, respectively. To optimize inter-chiplet data communication, we also develop an RL-based path routing framework tailored to this dual-layer architecture. The RL framework’s state space includes topology, network, and memory metrics, while the RL reward is set as the product of average packet latency and link utilization. The effectiveness of EIGEN is demonstrated through different applications where CPU, GPU, and AI accelerator chiplets are integrated via NOAI. Simulation results show that the proposed EIGEN achieves up to $\mathbf{6 7. 1 6 \%}$ latency reduction, $\mathbf{5 3. 8 9 \%}$ hops reduction and $\mathbf{1 1. 2 1 \%}$ runtime reduction compared to the state-of-the-art (SOTA) chiplets interconnect architectures. Furthermore, sensitivity analysis of network scaling shows that the EIGEN architecture and framework exhibit strong scalability, with latency reduction of $\mathbf{2 0. 9 6 \%}$ to $\mathbf{5 2. 7 7 \%}$ and hop reduction of $\mathbf{1 9. 7 2 \%}$ to $\mathbf{4 3. 4 1 \%}$ as the system scales from $4 \times 4$ to $32 \times 32$ chiplets. Siyao Jia, Bo Jiao 0003, Haozhe Zhu, Chixiao Chen, Qi Liu 0010, Ming Liu 0022 |
HPCA | 6 |
| 2025 | A monolithic 3D IGZO-RRAM-SRAM-integrated architecture for robust and efficient compute-in-memory enabling equivalent-ideal device metrics
Shengzhe Yan, Zhaori Cong, Zhuoyu Dai, Zeyu Guo 0002, Zhihang Qian, Xufan Li, Chuanke Chen, Nianduan Lu, Chunmeng Dou, Guanhua Yang, Xiaoxin Xu, Di Geng, Jinshan Yue, Ling Li 0013, Ming Liu 0022 |
Sci. China Inf. Sci. | 18 |
| 2024 | FPIA: Communication-Aware Multi-Chiplet Integration With Field-Programmable Interconnect Fabric on Reusable Silicon InterposerabstractSilicon interposer re-usage is drawing attention for cost-effective multi-chiplet integrated systems. To address the communication awareness of inter/off-chiplet interconnect, the paper proposes a field-programmable interconnect fabric and develops its corresponding automatic physical integration tool. The tile-based fabric consists of turnout, cross-over boxes and parallel tracks. It features micro-bump-wise connecting flexibility and hardware efficiency. The automation flow performs chiplet location optimization and efficient bump-to-bump routing, supporting multi-lane bus interconnect and miscellaneous external ports. The methodology is validated by 9 different integration scenarios, where the routability is guaranteed when the local resource utilization ratio approaches 94.5%. The data’s maximum interconnect latency is 2.2 ns and the energy consumption is 1.18 pJ/bit at a bitrate of 1 Gbps. The latency consumes$16.5\times \sim ~53.4\times $fewer clock cycles than the state-of-the-art network-on-package-based reusable interposer architectures. Bo Jiao 0003, Haozhe Zhu, Jundong Zhu, Dexin Wen, Lingli Wang, Jun Tao 0001, Chixiao Chen, Yinhe Han 0001, Qi Liu 0010, Ninghui Sun, Ming Liu 0022 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 15 |
| 2023 | CNN Accelerator at the Edge With Adaptive Zero Skipping and Sparsity-Driven Data FlowabstractAn energy-efficient convolutional neural network (CNN) accelerator is proposed for low-power inference on edge devices. An adaptive zero skipping technique is proposed to dynamically skip the zeros in either activations or weights, depending on which has the higher sparsity. The characteristic of non-zero data aggregation is explored to enhance the effectiveness of adaptive zero skipping in performance boosting. To mitigate the load imbalance issue after zero skipping, a sparsity-driven data flow and low-complexity dynamic task allocation are employed for different convolution layers. Facilitated further by a two-stage distiller, the proposed accelerator achieves$5.42\times $,$3.41\times $, and$3.42\times $performance boosting for VGG16, AlexNet, and Mobilenet-v1, respectively, compared to the baseline. Implemented in a 55-nm low power CMOS technology, the proposed accelerator achieves an effective energy efficiency of 2.41 TOPS/W, 2.35 TOPS/W, and 0.64 TOPS/W for VGG16, AlexNet, and Mobilenet-v1, respectively, at 100 MHz and 1.08 V supply voltage. Ming Liu 0022, Changchun Zhou 0001, Siyuan Qiu, Yifan He 0002, Hailong Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Sparsity-Aware Clamping Readout Scheme for High Parallelism and Low Power Nonvolatile Computing-in-Memory Based on Resistive MemoryabstractThe input parallelism of resistive memory (RRAM) based nonvolatile computing-in-memory (nvCIM) structure is limited by the signal margin as well as the readout precision. In this work, we propose a sparsity-aware clamping (SAC) scheme and its circuit implementation for nvCIM by co-design of circuit and algorithm. It can adaptively tune the quantized range and resolution of the readout circuit according to the degree of sparsity in neural network models. As a result, the SAC scheme can effectively increase the input parallelism of nvCIMs without incurring degradation on the signal margin or increasing the hardware cost for analogue readout. A case study on processing a multi-layer perceptron (MLP) model with the proposed nvCIM structure shows that the SAC scheme can improve the throughput by 2 times and increase the energy efficiency by 25.35% with negligible inference accuracy loss. Linfang Wang, Wang Ye, Junjie An, Chunmeng Dou, Qi Liu 0010, Meng-Fan Chang, Ming Liu 0022 |
ISCAS | 7 |
| 2021 | Investigation of weight updating modes on oxide-based resistive switching memory synapse towards neuromorphic computing applications
Qingting Ding, Tiancheng Gong, Jie Yu 0027, Xiaoxin Xu, Hangbing Lv, Feng Zhang 0014, Ming Liu 0022 |
Sci. China Inf. Sci. | 11 |
| 2021 | Simulations of single event effects on the ferroelectric capacitor-based non-volatile SRAM design
Jianjian Wang, Jinshun Bi, Bo Li 0051, Sandip Majumdar, Lanlong Ji, Ming Liu 0022, Zhangang Zhang |
Sci. China Inf. Sci. | 9 |
| 2020 | Design of a Single-Stage Wireless Charger with 92.3%-Peak-Efficiency for Portable Devices ApplicationsabstractThis summary presents a fully-integrated wireless charger to achieve high efficiency with low cost and volume. The charger realizes power rectification, voltage regulation and CCCV charging in one power stage only. A bootstrapping technique is also designed for on-chip integration of the bootstrap capacitors. A chip prototype was fabricated in a standard 0.35μm CMOS process with a die area of 8mm2. The charger achieves peak efficiency of 92.3% and 91.4% when the charging currents are 1A and 1.5A, respectively. Lin Cheng 0001, Xinyuan Ge, Wai Chiu Ng, Wing-Hung Ki, Tsz Fai Kwok, Chi-Ying Tsui, Ming Liu 0022 |
ASP-DAC | 8 |
| 2020 | A Low Power 4T2C nvSRAM With Dynamic Current Compensation Operation SchemeabstractThis study proposed a novel nonvolatile static random access memory (nvSRAM) cell with two ferroelectric capacitors (FeCAPs) embedded inside a 4T SRAM cell, i.e., 4T2C, for minimal area penalty and full logic compatibility. The FeCAP with 10-nm-thick Hf0.5Zr0.5O2film shows excellent ferroelectricity (Pr = 15 μC/cm2) and good memory characteristics (cycles 1011). The 4T2C nvSRAM is capable of storing and restoring previous memory states for nonvolatile data storage. To compensate the leakage current in the dynamic nodes of 4T load less SRAM, we propose a dynamic current compensation operation scheme by exploiting the polarization-dependent leakage current of FeCAP. Outstanding characteristics were achieved in this nvSRAM cell: 1) elimination of the dc path; 2) ultralow store and restore power consumption; and 3) high area efficiency. Tiancheng Gong, Qingting Ding, Yuling Zhao, Xiaoyong Xue, Hangbing Lv, Ming Liu 0022 |
IEEE Trans. Very Large Scale Integr. Syst. | 12 |
| 2019 | Total ionizing dose effects on graphene-based charge-trapping memory
Jinshun Bi, Sandip Majumdar, Bo Li 0051, Jing Liu 0035, Yannan Xu, Ming Liu 0022 |
Sci. China Inf. Sci. | 7 |
| 2018 | Low-Noise High-Linearity 56Gb/s PAM-4 Optical Receiver in 45nm SOI CMOSabstractA 56Gb/s PAM-4 linear optical receiver with low noise and high linearity is presented. The fully integrated receiver comprises a transimpedance amplifier (TIA), a variable gain amplifier (VGA), an output buffer, auxiliary analog loops and on-chip bias circuitry. As will be shown, low noise and high linearity often contradict each other, thus both the TIA and VGA implement novel gain control techniques for linear operation while realizing low noise design, making them favorable for PAM-4 signal amplification. Designed and implemented in 45nm SOI CMOS technology, the receiver accomplishes state-of-the-art input-referred noise current of 1.8μArms, 74.4dB transimpedance gain and 23GHz bandwidth while consuming 37mW. The dynamic range achieved is 29dB, enabling large input overload of 0.8mA for PAM-4 compliant signaling. Dan Li 0011, Yiqun Liu 0011, Ming Liu 0022, Li Geng |
ISCAS | 4 |
| 2017 | Total ionizing dose effects and annealing behaviors of HfO2-based MOS capacitor
Yannan Xu, Jinshun Bi, Gaobo Xu, Bo Li 0051, Ming Liu 0022 |
Sci. China Inf. Sci. | 6 |
| 2015 | A 1G-cell floating-gate NOR flash memory in 65 nm technology with 100 ns random access time
Zongliang Huo, Ming Liu 0022, Liyang Pan |
Sci. China Inf. Sci. | 5 |
| 2011 | 3D Integration of CMOL Structures for FPGA ApplicationsabstractIn this paper, a novel 3D CMOS nanohybrid technology, 3D CMOL, is introduced to establish FPGA chips. By combining two leading technologies, hybrid CMOS/nanoelectronic circuit (CMOL) and 3D integration, 3D CMOL can provide a feasible and more efficient fabrication/assembly process than the existing 2D CMOL. Furthermore, 3D CMOL FPGA implements circuits in three dimensions so that it can increase the density of the nanodevices and achieve higher performance compared to 2D CMOL and field programmable nanowire interconnect (FPNI). This paper presents the architecture optimization, 3D integration, defect tolerance, and performance evaluation of 3D CMOL FPGA. It is expected that this technology can lead to technology breakthroughs towards the development of the next FPGA generation. Zine-Eddine Abid, Ming Liu 0022, Wei Wang 0003 |
IEEE Trans. Computers | 2 |
| 2011 | FPGA Based on Integration of CMOS and RRAMabstractIn this paper, a novel CMOS-nano hybrid reconfigurable field-prgrammable gate array architecture (rFPGA) is introduced based on resistive memory (RRAM) devices. Different from the existing crossbar-based CMOS-nano architectures, rFPGA consists of mainly 1T1R RRAM structures that can be fabricated by using a CMOS-compatible process. These devices can efficiently establish FPGA block memories. More importantly, novel RRAM routing switches are developed to replace the CMOS routing switches to achieve significant density enhancement and power reduction. The simulation results demonstrate that 2-D and 3-D rFPGAs provide at least 2$\times$–3$\times$overall improvement in terms of area with 20% lower power consumption, compared with the corresponding CMOS FPGAs. Sansiri Tanachutiwat, Ming Liu 0022, Wei Wang 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | A low power single ended input differential output low noise amplifier for L1/L2 bandabstractA 1.2~1.6 GHz low power single-ended input to differential output low noise amplifier for Global Navigation Satellite System (GNSS, such as GPS, Galileo, China Compass, etc.) is proposed in the paper. A new current-reused technique with mismatch adjustment is adopted to achieve low power, low noise factor, high voltage gain and low mismatch simultaneously. The LNA is fabricated in TSMC 0.18 μm RF CMOS process, with only 2 mA current consumption under 1.8 V supply. The LNA exhibits a differential voltage gain of 27-30 dB, a noise figure of 2.4-3.0 dB and gain mismatch less than 0.5 dB. Yonghui Ji, Ming Liu 0022, Shibing Long, Zhaoan Yu, Manhong Zhang |
ISCAS | 2 |
| 2010 | Formation and annihilation of Cu conductive filament in the nonpolar resistive switching Cu/ZrO2: Cu/Pt ReRAMabstractWe report a ZrO2-based resistive memory composed of a thin Cu doped ZrO2layer sandwiched between Pt bottom and Cu top electrode. The Cu/ZrO2:Cu/Pt shows excellent nonpolar resistive switching behaviors, such as free-electroforming, high ON/OFF resistance ratio (106), fast Set/Reset speed (50 ns/100 ns), and reliable data retention (>10 years). The temperature-dependent switching characteristics show that a metallic filamentary channel is responsible for the low resistance state. Further analysis reveals that the physical origin of this metallic filament is the nanoscale Cu conductive filament. On this basis, we propose that the set process and the reset process stem from the electrochemical reactions in the filament, in which a thermal effect is greatly involved. Ming Liu 0022, Qi Liu 0010, Shibing Long, Weihua Guan |
ISCAS | 1 |
| 2008 | Analyzing mixed carbon nanotube bundles: A current density studyabstractCarbon nanotube (CNT) bundle is a promising candidate for future on-chip interconnects and electro-thermal applications due to its superior electrical and thermal properties. A realistic CNT bundle is a mixture of single-wall/multi-wall CNTs due to the nature of bottom-up fabrication process. In order to utilize CNT bundle, its current density performance needs to be analyzed considering both types of CNTs. In this paper, compact current density models are introduced for CNTs and mixed CNT bundles considering the influence of diameter and other structural factors. The current density performance estimation based on these models match the recent measurement results. Considering realistic CNT densities and geometries, the mixed bundle shows various levels of improvement over the copper wire counterpart in terms of current density. Liwei Shang, Ming Liu 0022, Sansiri Tanachutiwat, Wei Wang 0003 |
ISCAS | 2 |
| 2007 | Fault Tolerance Circuit for AM-OLEDabstractA novel circuit employing p-type low-temperature poly-Si thin-film transistors is introduced for active matrix-organic light-emitting diode (AM-OLED) circuits to automatically detect short defects and switch to a spare OLED. This design maintains the luminance of the OLED pixel without changing the driving current in the event of defects. Experimental results show that not only is fault tolerance capability obtained during operation, but also a significant amount (around 90%) of power consumption is saved compared with the standard driving circuits. Dayong Li, Ming Liu 0022, Wei Wang 0003 |
ISCAS | 2 |
| 2006 | Hybrid Nanoelectronics: Future of Computer Technology
Wei Wang 0003, Ming Liu 0022, Andrew Hsu |
J. Comput. Sci. Technol. | 2 |