VLDB 2026 Research / reviewers in the wild / expert
Yimao Cai
dblp:50/9704
· DBLP profile ↗
38ranked-venue papers
1as first author
29since 2021 · last 2026
0000-0002-6854-8211ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 19 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHIP-MAP: A Collaborative Optimization Framework for Macro Placement Using Large Language ModelsabstractAs integrated circuits continue to grow in both scale and complexity, macro placement plays a critical role in physical design, directly affecting chip-level performance, power, and area (PPA). Traditional macro placement methods, such as simulated annealing, analytical optimization, and reinforcement learning, face limitations including slow convergence, heavy dependence on large datasets, and over-reliance on intermediate PPA indicators rather than final PPA. Large language models (LLMs) offer strong generative power and semantic reasoning that can potentially automate macro layout tasks while addressing the aforementioned problems in traditional methods, but their limited understanding of layout rules and lack of iterative, feedback-driven refinement make direct application challenging. To address this, we propose CHIP-MAP, a macro placement framework based on multi-agent collaboration and feedback-driven optimization. Furthermore, we introduce two innovative tools: the Module Link Weight Analyzer (MWA) and the Standard Cell Usability Score (SCUS), which are designed to guide fine-grained layout refinement. We evaluate CHIP-MAP on five benchmarks ranging from low-power cores to large multi-core processors implemented at 130nm and 45nm technology nodes. Results show that it achieves up to 1.5% area reduction and an average repair of 61.6% of total negative slack (TNS), while also reducing wirelength and improving timing. Yiming Du, Renye Yan, Yunfan Yang, Frank Qu, Jiajun Tan, ZhiYu Zheng, Yiming Gan, Ling Liang 0003, Zongwei Wang 0001, Yimao Cai |
DATE | 10 |
| 2026 | GMaC: NvCIM Architecture for Parallel Point-based Point Cloud Acceleration via Geometric Mapping and Address-Index Computation
Zongwei Wang 0001, Ling Liang 0003, Yimao Cai |
DATE | 4 |
| 2026 | SONIC: Smart Optimization for Neural-Integrated CMP with Timing-Aware FillsabstractDummy fill insertion is essential for CMP uniformity but remains challenging due to the nonlinear CMP process, the large optimization space, and timing degradation caused by parasitic coupling. We propose SONIC, a differentiable CMP-driven dummy fill optimization framework that employs a neural CMP simulator to directly optimize planarization objectives using gradient-based methods. SONIC further integrates a timing-aware fill insertion strategy to mitigate coupling capacitance near critical nets. Experimental results demonstrate that SONIC achieves competitive planarization quality with up to 1830× runtime speedup over a full-chip CMP simulator. Compared with the state-of-the-art model-based method, SONIC reduces height variation, line deviation, and outliers by up to 86.16%, 90.10%, and 51.61%, respectively, while achieving a 77.67% runtime reduction and lowering coupling capacitance by 13.05%. Jiajun Tan, Yiming Du, Yiming Gan, Ling Lang 0002, Yibo Lin, Zongwei Wang 0001, Yimao Cai |
DATE | 9 |
| 2026 | DIAMoND: Dynamic Inference for Adaptive Edge MOE with Heterogeneous In-NAND and Near-DRAM Compute Architecture
Tianyang Luo, Shuzhang Zhong, Dongxue Zhao, Renjie Wei, Meng Li 0004, Guangyu Sun 0003, Zongwei Wang 0001, Yimao Cai |
ISCA | 11 |
| 2026 | A Sparsity-Aware Reconfigurable Sensing-Quantization Circuit for RRAM-Based Analog Compute-in-Memory
Zhuoya Chen, Hao Ding 0011, Yunfan Yang, Haisu Zhang, Jinshan Li, Shigeng Zhao, Yunyi Fu, Zongwei Wang 0001, Yimao Cai |
ISCAS | 9 |
| 2026 | An IZO-Based 2T0C Compute-in-Memory Array with Adaptive Read Voltage Boosting for Energy-Efficient Edge AI
Hao Ding 0011, Jiye Li, Xiantong Qiu, Yunfan Yang, Gaoqi Yang, Shengdong Zhang, Zongwei Wang 0001, Yimao Cai |
ISCAS | 10 |
| 2026 | A Near-Sensor Image Compression Architecture with RRAM-based Hyperdimensional Encoder for Smart Vision
Haoyang Gu, Zongwei Wang 0001, Jingshan Li, Zezhi Chen, Ling Liang 0003, Wengao Lu, Yimao Cai |
ISCAS | 8 |
| 2026 | RRAM-based CAM for Energy-Efficient In-Memory Text Compression System
Lianliang Wu, Hao Ai, Ao Shi, Haokai Guan, Kexun Li, Kefan Tao, Yulin Feng, Zongwei Wang 0001, Yimao Cai, Peng Huang 0004 |
ISCAS | 13 |
| 2026 | A 1.8-ns, 45fJ/bit Time-Domain Sensing Scheme with Offset-Cancelled Resistance-to-Time Converter and 10-5 BER for Digital RRAM Compute-in-Memory
Shigeng Zhao, Hao Ding 0011, Yunfan Yang, Jinshan Li, Zhuoya Chen, Xing Zhang 0002, Zongwei Wang 0001, Yimao Cai |
ISCAS | 8 |
| 2026 | Efficient LoRA-Based Weight Update Write-Back in 3D NAND Flash for Large Language Models
Dongxue Zhao, Tianyang Luo, Ling Liang 0003, Zongwei Wang 0001, Yimao Cai |
ISCAS | 5 |
| 2026 | OmniGuard: Two-Level Protection Framework for RRAM-Based Accelerator With High Efficiency and FlexibilityabstractRRAM-based Deep Neural Network (DNN) accelerators have gained widespread usage in edge devices. However, the security vulnerabilities of RRAM-based accelerators hinder their real application. Current research on safeguarding RRAM-based accelerators predominantly relies on a single-level protection approach. This has resulted in the restriction of its protection scope, the rigidity and lack of generality in the protection method, or has had an impact on the computational efficiency of the system. As a result, it encounters substantial challenges in attaining comprehensive optimization across multiple dimensions, such as universality, the scope of protection, and security-related overheads. In this paper, we develop specific analyses on accelerators and attacks and build graph-based representations. Based on these, we partition the RRAM-based accelerators’ security into two levels: on-chip security and off-chip security. Furthermore, we propose a two-level protection framework for RRAM-based accelerators, which is calledOmniGuard. At the on-chip security level,OmniGuardproposes a bit-grained shuffle to achieve protection while using lightweight Benes Networks to maintain the CIM capability. At the off-chip level,OmniGuardproposes an RRAM-based AES engine to introduce the AES algorithm into the accelerator with significant acceleration and minimal overhead. Evaluation results demonstrate thatOmniGuardprovides powerful and flexible protection while achieving 1.33×∼4.38× speedup and 1.26×∼2.65× power savings, with only 5% energy overhead and 3% area overhead. Ling Liang 0003, Yunfan Yang, Jinlong Lin, Meng Li 0004, Zongwei Wang 0001, Yimao Cai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | REF-CIM: A 40-nm Non-Ideality Tolerant and Energy Efficient RRAM Compute-in-Memory Macro With Configurable Precision for Edge AI
Hao Ding 0011, Yunfan Yang, Zongwei Wang 0001, Jinshan Li, Lin Bao, Ling Liang 0003, Yimao Cai |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment
Renye Yan, Jikang Cheng, Yaozhong Gan, Shikun Sun, Yunfan Yang, Ling Liang 0003, Jinlong Lin, Yeshuang Zhu, Jie Zhou 0001, Junliang Xing, Yimao Cai, Ru Huang 0001 |
ICCV | 13 |
| 2025 | SA-CIM: A 28nm 16Mb RRAM-based Sparsity-Aware Compute-In-Memory Macro for Edge AI Algorithm ProcessingabstractCompute-in-memory (CIM) for edge devices is usually constrained by on-chip resources, including on-chip memory and physical chip size, which hinders the deployment of more complex neural networks. By leveraging the sparsity of neural networks, the overall memory requirements and energy consumption can be reduced. However, Existing sparsity-aware architectures cannot achieve high energy efficiency due to off-chip sparsity control. This work proposes:1) Hybrid sparsity regulation strategy. The sparsity encoding and alignment circuit is designed and implemented, realizing on-chip sparsity detecting and encoding. 2) Sparsity-aware compute-in-memory (CIM) array based on RRAMs. The in-situ deployment of unstructured sparsity is implemented inside the CIM array, and the CIM array and sparsity are tightly coupled by sparsity read/write. This work demonstrates the design and evaluation of SA-CIM: a sparsity-aware CIM macro with 16Mb RRAM with fine-grained sparsity detecting and encoding capacity, achieving energy efficiency of 22.7TOP/W@8b/8b. Hao Ding 0011, Zongwei Wang 0001, Jinshan Li, Shigeng Zhao, Heting Gao, Junbo Ao, Ling Liang 0003, Yimao Cai, Ru Huang 0001 |
ISCAS | 8 |
| 2025 | HRC-CIM: Hybrid RRAM-Capacitor Cell based Compute-in-Memory with High Linearity, Parallelism and Energy EfficiencyabstractRRAM-based Compute-in-memory (CIM) has emerged as a promising computing paradigm for artificial intelligence (AI) algorithms. However, the low on/off ratio and high on-current have been the major challenges to enhance the accuracy, parallelism, and energy efficiency. In this paper, we propose a novel Hybrid RRAM-Capacitor (HRC) cell based CIM macro to address these issues. The proposed HRC cell achieves a high on-off ratio with sub-100nA on-current and eliminates direct current path during computation, which significantly enhances both parallelism and energy efficiency. The write-verify scheme for RRAM programming is optimized for HRC cell array, and is further supported by a quantization result calibration technique using a dummy column to ensure high linearity and accuracy in analog domain multiply-and-accumulate (MAC) operations. A HRC-CIM macro has been designed and demonstrated using 28nm technology node, enabling block-level parallelism across 64 rows with 4-bit input per row, and delivering an energy efficiency of up to 40.40 TOPS/W @8b-IN/8b-W. Jinshan Li, Zongwei Wang 0001, Hao Ding 0011, Yunfan Yang, Shigeng Zhao, Shengyu Bao, Ruiqing Xie, Zhuoya Chen, Yimao Cai, Ru Huang 0001 |
ISCAS | 10 |
| 2025 | A Generic Circuit Platform of Covering All Rational Numbers for Determnistic Stochastic Computing in the Time DimensionabstractThe key advantages of stochastic computing (SC) are its simple arithmetic circuits and high tolerance for bit errors. Deterministic SC (DSC) improves SC by using deterministic bitstreams, resulting in accurate results without fluctuation. Recently, a platform that conducts DSC in the time domain has been proposed. This platform enhances the feasibility of DSC and brings it closer to becoming a mainstream computational paradigm. However, its circuit is unable to represent all rational numbers. This work improves the platform by enabling the representation of all rational numbers for computation. Furthermore, it reduces hardware costs and power consumption, paving the way for new SC applications. Xiangye Wei, Junjian Ma, Liming Xiu, Yimao Cai |
ISCAS | 4 |
| 2025 | Polynomial geometric transformation based on IGZO charge trapping RAM array for machine vision calibration
Lin Bao, Haisu Zhang, Zongwei Wang 0001, Linbo Shan, Cuimei Wang, Yimao Cai, Shanguo Huang |
Sci. China Inf. Sci. | 6 |
| 2025 | Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE AcceleratorabstractFully homomorphic encryption over torus (TFHE) enables the execution of arbitrary functions on encrypted data through programmable bootstrapping (PBS). However, performing all operations on ciphertext during PBS results in high computational and memory requirements, limiting the deployment of PBS in real-world scenarios. Previous TFHE accelerator designs have attempted to improve performance by employing specific dataflow and functional units, but these techniques may require large off-chip bandwidth or on-chip storage when scaling up computation capacity. Additionally, the design of specialized functional units may limit the utilization of computation units when facing dynamic secure parameter settings. To address these challenges and further improve PBS throughput in TFHE, we propose Matrix , an ASIC-based architecture that balances off-chip bandwidth and on-chip storage according to the execution flow of PBS. In Matrix , we utilize a unified special-prime-based processing element (PE) that achieves high utilization with minimal resource overhead. Furthermore, we propose a hybrid PBS dataflow that can efficiently reduce computation complexity and memory requirements. Compared to state-of-the-art TFHE accelerators, Matrix achieves 1.43 × -5.66 × throughput improvement for PBS. For ZAMA Deep-NN benchmark, we achieve 525.60× and 68.06× speedup compared to CPU and GPU, respectively. 1 Ling Liang 0003, Fahong Zhang 0004, Zhirui Li, Xin Fan 0009, Dimin Niu, Meng Li 0004, Zhiyong Li 0016, Zongwei Wang 0001, Hongzhong Zheng, Yimao Cai, Yuan Xie 0001 |
ACM Trans. Archit. Code Optim. | 12 |
| 2024 | Autoencoder Reconstruction Model for Long-Horizon ExplorationabstractConventional reinforcement learning (RL) algorithms often necessitate millions of environment interactions to ascertain an efficacious policy. In stark contrast, humans, leveraging their curiosity mechanisms, can develop proficient policies with minimal effort. Drawing inspiration from this observation, we introduce the Autoencoder Reconstruction Model(ARM), a curiosity-driven RL model that significantly reduces interactions while enhancing policy effectiveness. ARM employs an autoencoder module, utilizing a deep neural network to learn feature representations from the environment. ARM utilizes its Curiosity Measurement Module to motivate RL agents for effective exploration, particularly in environments with sparse rewards. ARM also introduces an innovative mechanism to balance the exploration-exploitation dilemma. Theoretical analyses reveal that the reward shaping introduced by the ARM aligns with the potential-based reward shaping paradigm, thereby preserving the optimality of reinforcement learning. We will release the source code and trained models to facilitate further studies in this research direction. Renye Yan, Yaozhong Gan, Yunfan Yang, Zhaoke Yu, Zongxi Liu, Ling Liang 0003, Yimao Cai |
IJCNN | 9 |
| 2024 | An isolated symmetrical 2T2R cell enabling high precision and high density for RRAM-based in-memory computing
Yaotian Ling, Zongwei Wang 0001, Lin Bao, Shengyu Bao, Yimao Cai, Ru Huang 0001 |
Sci. China Inf. Sci. | 7 |
| 2024 | Investigation and mitigation of Mott neuronal oscillation fluctuation in spiking neural network
Lindong Wu, Zongwei Wang 0001, Lin Bao, Linbo Shan, Zhizhen Yu, Yunfan Yang, Shuangjie Zhang, Guandong Bai, Cuimei Wang, John Robertson, Yuan Wang 0001, Yimao Cai, Ru Huang 0001 |
Sci. China Inf. Sci. | 12 |
| 2024 | A Perspective of Using Frequency-Mixing as Entropy in Random Number Generation for Portable Hardware Cybersecurity IPabstractTrue random number generator (TRNG) is a crucial component in security. In typical TRNGs, entropy comes directly from device noises. In this work, an improved method of using frequency-mixing as means for enriching entropy is implemented. A group of electromagnetic waves are mixed to create an irregular waveform that is then sampled to generate a random bitstream. Some part of the bitstream is fed back to the system for influencing the future frequencies of the sourcing waves, making it a chaotic system. The circuit-level support for this TRNG is the TAF-DPS (Time-Average-Frequency Direct Period Synthesis) technology. It can be digitally implemented, making the TRNG a portable IP. The merits of this TRNG include no need of special device, no post-processing, free of bias, programmable throughput, and hard-to-recognize spectrum. Those features make the TRNG suitable for a large array of applications, particularly for security in cyberspace. This TRNG is validated by a silicon chip on a 180 nm process, also on a FPGA. Xiangye Wei, Liming Xiu, Yimao Cai |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Mnemonic Dictionary Learning for Intrinsic Motivation in Reinforcement LearningabstractReinforcement learning for hard-exploration tasks remains challenging due to the long-term dependence and sparse-and-delay rewards in complex environments. In these challenging tasks, intrinsic motivation has become a dominant paradigm to enable the agent to explore the environment when no external reward feedback is available. In this work, inspired by studies from the human memory mechanism, we present a mnemonic dictionary learning (MDL) model for intrinsic motivation in reinforcement learning. The MDL model leverages sparse dictionary learning to incremental abstract the exploration histories into a compact memory-like dictionary, providing an excellent intrinsic motivation model. This mnemonic dictionary model not only drives the agent to explore novel stats in the environments indicated by the memory reconstruction error but also helps the agent to remember the key states and structure of the environments using its learned bases and reconstruction coefficients. The proposed MDL model can serve as a generative module for existing exploration methods. Extensive experimental results on typical sparse-reward tasks demonstrate its effectiveness and applicability over several competing algorithms. We will release the source code and trained models to facilitate further studies in this research direction. Renye Yan, Yuan Zhan, Pin Tao, Zongwei Wang 0001, Yimao Cai, Junliang Xing |
IJCNN | 6 |
| 2022 | Facial Action Unit Detection by Exploring the Weak Relationships Between AU Labels
Mengke Tian, Hengliang Zhu, Yimao Cai, Pengrong Lin, Yingzhuo Huang, Xiaochen Xie |
CollaborateCom (2) | 4 |
| 2022 | PIMulator-NN: An Event-Driven, Cross-Level Simulation Framework for Processing-In-Memory-Based Neural Network AcceleratorsabstractProcessing-in-memory (PIM) architecture has been proposed to accelerate state-of-the-art neuro-inspired algorithms, such as deep neural networks. In this article, we present PIMulator-NN, an event-driven, cross-level simulation framework for PIM-based neural network accelerators. By employing an event-driven simulation mechanism, PIMulator-NN is able to model architecture details and capture design details of the architecture. Moreover, we integrate the main-stream circuit-level simulation framework with PIMulator-NN to accurately simulate the area, latency, and energy consumption of analog computation units. To demonstrate the usage of PIMulator-NN, we implement several PIM designs with PIMulator-NN and perform detailed simulation. The simulation results show that memory access and interconnects make considerable impacts on system-level performance and energy. Note that such results are hard to be captured by conventional performance model-based estimations. We found some anti common-sense results while modeling the architecture details with PIMulator-NN. With several architecture templates, PIMulator-NN provides the users with a platform to build up their PIM architecture quickly. PIMulator-NN is able to capture the impacts of different design choices (e.g., dataflow, interconnect, data parallelism, etc.), and this could enable users to explore their design space efficiently. Qilin Zheng, Yijin Guan, Zongwei Wang 0001, Yimao Cai, Yiran Chen 0001, Guangyu Sun 0003, Ru Huang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | A High Accuracy Multiple-Command Speech Recognition ASIC Based on Configurable One-Dimension Convolutional Neural NetworkabstractSpeech command interaction has drawn much attention in smart application market. Many of previous chips achieve an ultra-low power consumption at the cost of a certain accuracy loss, and they are designed only for the fixed speech command recognition tasks, which is inflexible and restrains further development. Here, we demonstrate a configurable speech command recognition ASIC with an ultra-high accuracy fabricated by the TSMC commercial 180-nm CMOS technology. In this chip, Mel-Frequency Cepstrum Coefficients (MFCCs) are used as speech features and a One-Dimension Convolutional Neural Network (1-D CNN) is adopted for the speech feature recognition, which simplifies the design of network and the storage method of memory. Moreover, the configurable 1-D CNN layer of the network ensures the diversity and flexibility of the commands. The measurement results indicate that the chip achieves a 95.6% accuracy on Google Speech Command Database (GSCD) when working at 16 MHz and keeping a reasonable power consumption as 26.4 mW. Moreover, the chip supports max 30 speech commands at a time, which is better than the state-of-the-art chips. Lindong Wu, Zongwei Wang 0001, Yimao Cai, Ru Huang 0001 |
ISCAS | 5 |
| 2021 | In-memory computing with emerging nonvolatile memory devices
Caidie Cheng, Pek Jun Tiw, Yimao Cai, Xiaoqin Yan, Yuchao Yang 0001, Ru Huang 0001 |
Sci. China Inf. Sci. | 3 |
| 2021 | Recent progress of integrated circuits and optoelectronic chips
Yue Hao 0001, Genquan Han, Jincheng Zhang 0001, Xiaohua Ma 0001, Zhangming Zhu, Yanan Han, Ling Yang 0003, Jiangyi Shi, Wei Zhang 0343, Biao Pan, Yangqi Huang, Qi Liu 0010, Yimao Cai, Xin Ou, Tiangui You, Huaqiang Wu, Bin Gao 0006, Guoping Guo, Yonghua Chen, Xiangfei Chen, Chunlai Xue, Lixia Zhao, Xihua Zou, Lianshan Yan |
Sci. China Inf. Sci. | 21 |
| 2021 | Optimization Schemes for In-Memory Linear Regression Circuit With Memristor ArraysabstractRecently, an in-memory analog circuit based on crosspoint memristor arrays was reported, which enables solving linear regression problems in one step and can be used to train many other machine learning algorithms. To explore its potential for computing accelerator applications, it is of fundamental importance to improve the computing speed of the circuit,i.e., the circuit response towards correct outputs. In this work, we comprehensively studied the transfer function of this circuit, resulting in a quadratic eigenvalue problem that describes the distribution of poles. The minimal real part of non-zero eigenvalues defines the dominant pole, which in turn dominates the response time. Simulations for multiple linear regression solutions with different datasets evidence that, the computing time does not necessarily increase with problem size. The dominant pole is related to parameters in the circuit, including feedback conductance, and gain bandwidth products of operational amplifiers. By optimizing these parameters synergistically, the dominant pole shifts to higher frequencies and the computing speed is consequently optimized. Our results provide a guideline for design and optimization of in-memory machine learning accelerators with analog memristor arrays. Also, issues including power consumption, impact of noise and variation of sources and memristors are investigated to offer a comprehensive evaluation of the circuit performance. Zhong Sun, Shengyu Bao, Yimao Cai, Daniele Ielmini, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | Lattice: An ADC/DAC-less ReRAM-based Processing-In-Memory Architecture for Accelerating Deep Convolution Neural NetworksabstractNonvolatile Processing-In-Memory (NVPIM) has demonstrated its great potential in accelerating Deep Convolution Neural Networks (DCNN). However, most of existing NVPIM designs require costly analog-digital conversions and often rely on excessive data copies or writes to achieve performance speedup. In this paper, we propose a new NVPIM architecture, namely, Lattice, which calculates the partial sum of the dot products between the feature map and weights of network layers in a CMOS peripheral circuit to eliminate the analog-digital conversions. Lattice also naturally offers an efficient data mapping scheme to align the data of the feature maps and the weights and hence, avoiding the excessive data copies or writes in the previous NVPIM designs. Finally, we develop a zero-flag encoding scheme to save the energy of processing zero-values in sparse DCNNs. Our experimental results show that Lattice improves the system energy efficiency by 4× ~ 13.22× compared to three state-of-the-art NVPIM designs: ISAAC, PipeLayer, and FloatPIM. Qilin Zheng, Zongwei Wang 0001, Zishun Feng, Bonan Yan, Yimao Cai, Ru Huang 0001, Yiran Chen 0001, Chia-Lin Yang, Hai Li 0001 |
DAC | 5 |
| 2020 | MobiLattice: A Depth-wise DCNN Accelerator with Hybrid Digital/Analog Nonvolatile Processing-In-Memory BlockabstractNonvolatile Processing-In-Memory (NVPIM) architecture is a promising technology to enable energy-efficient inference of Deep Convolutional Neural Networks (DCNNs). One major advantage of NVPIM is that the vector dot-product operations can be completed efficiently by analog computing inside a Nonvolatile Memory (NVM) crossbar. However, its inference efficiency is severely downgraded when processing depth-wise convolution layers, which have been widely employed in many lightweight DCNNs. One major challenge is that the cell utilization is extreme low when mapping the depth-wise convolution layer to a crossbar. To overcome this problem, we propose a novel hybrid mode NVPIM architecture, namely, MobiLattice. With moderate hardware overhead, MobiLattice enables both analog and digital mode operations on NVM crossbars. While conventional convolution layers are computed efficiently using the analog mode, the computation efficiency of depth-wise convolution layers are substantially improved using the digital mode by mitigating the redundant memory space in the NVM crossbars. Experimental results show that, compared to prior approaches where only the analog mode is supported by the NVPIM architecture, MobiLattice can speedup the processing of typical depth-wise DCNNs by 2 ~ 5× on average and up to 30× by combining with some extreme quantization schemes. Qilin Zheng, Zongwei Wang 0001, Guangyu Sun 0003, Yimao Cai, Ru Huang 0001, Yiran Chen 0001, Hai Li 0001 |
ICCAD | 5 |
| 2019 | Enhance the Robustness to Time Dependent Variability of ReRAM-Based Neuromorphic Computing Systems with Regularization and 2R SynapseabstractTime Dependent Variability (TDV) is one of the major concerns in implementing a Neuromorphic Computing System (NCS) with Resistive Random Access Memory (ReRAM). In this work, we propose a variation-distribution aware training algorithm to enhance the robustness of NCS to TDV without incurring extra hardware overhead by leveraging algorithm-level regularization and hardware-level 2R synapse structure. Simulation results on image recognition tasks show that our method improves the system accuracy by up to ∼4% and ∼10% under the worst-case TDV condition for MNIST and CIFAR-10, respectively. Detailed analysis also shows that our method allows the NCS to use synapses with higher resistance than conventional design for the same accuracy requirement, introducing potential energy saving. Qilin Zheng, Zongwei Wang 0001, Yimao Cai, Ru Huang 0001, Bing Li 0017, Yiran Chen 0001, Hai Li 0001 |
ISCAS | 4 |
| 2019 | Investigation of NbOx-based volatile switching device with self-rectifying characteristics
Yichen Fang, Zongwei Wang 0001, Caidie Cheng, Zhizhen Yu, Yuchao Yang 0001, Yimao Cai, Ru Huang 0001 |
Sci. China Inf. Sci. | 7 |
| 2018 | Integration of biocompatible organic resistive memory and photoresistor for wearable image sensing application
Yichen Fang, Zongwei Wang 0001, Yuchao Yang 0001, Jintong Xu, Yimao Cai, Ru Huang 0001 |
Sci. China Inf. Sci. | 7 |
| 2014 | Resistive switching in organic memory devices for flexible applicationsabstractThe organic resistance memories show great potentials for future flexible applications. In this paper the main challenges and typical recent progress of the organic resistance memory devices are discussed. A kind of single-component polymer resistance memory device based on polychloro-paraxylylene (parylene-C) is focused, with excellent chemical stability and high CMOS process compatibility as well as further reduction of operation current, which is promising for future information storage in flexible systems. Ru Huang 0001, Yimao Cai, Yefan Liu, Wenliang Bai, Yongbian Kuang, Yangyuan Wang |
ISCAS | 2 |
| 2011 | Design and Implementation of a Peripheral Bus Based on a New Kind of Reconfigurable SystemabstractReconfigurable system-on-a-chip(SoC) is an important trend of embedded system. It is not only to achieve a higher performance but also flexible enough. In this paper a new kind of reconfigurable system using SoP(System on a Package) technology is presented and a new kind of peripheral bus which is used to form a whole system architecture is proposed based on the reconfigurable system. Using this peripheral bus, we can form a new embedded system easily by changing different intellectual property(IP) cores. We can make the whole system much more smaller, higher levels of integration, lower costs and lower power by using this chip and the peripheral bus. Compare with the earlier system, the new system using the reconfigurable chip which is of the same function is much smaller and lighter. All these are very suitable for small satellites and consumer electronics. Yimao Cai, Yuanfu Zhao, Lidong Lan |
DASC | 1 |
| 2011 | Editor's note
Ru Huang 0001, Runsheng Wang, Yimao Cai |
Sci. China Inf. Sci. | 4 |
| 2008 | Novel vertical channel double gate structures for high density and low power flash memory applications
Ru Huang 0001, FaLong Zhou, Yimao Cai, DaKe Wu, Xing Zhang 0002 |
Sci. China Ser. F Inf. Sci. | 3 |