Liu Liu 0017

dblp:74/7037-17 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-0792-8146ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 4 first-author · 13 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author
YearPublicationVenuePosition
2026 STARC: Selective Token Access with Remapping and Clustering for Efficient LLM Decoding on PIM Systems
abstract
Serving large language models (LLMs) places significant pressure on memory systems due to frequent accesses and growing key–value (KV) caches as context lengths increase. Processing-in-memory (PIM) architectures offer high internal bandwidth and near-data compute parallelism, but current designs target dense attention and perform poorly under the irregular access patterns of dynamic KV cache sparsity. To mitigate this limitation, we propose STARC, a sparsity-optimized data mapping scheme for efficient LLM decoding on PIM. STARC clusters semantically similar KV pairs and co-locates them contiguously within PIM banks, enabling retrieval at cluster granularity by matching queries against precomputed centroids. This bridges the gap between fine-grained sparse attention and row-level PIM operations, improving utilization while minimizing overhead. On a simulated HBM-PIM system, under constrained KV budgets, STARC achieves up to 78% and 65% reductions in attention-layer latency and energy over token-wise sparsity methods, and up to 93% and 92% reductions relative to full attention, while preserving model accuracy.
Zehao Fan, Yunzhen Liu, Garrett Gagnon, Yayue Hou, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu 0017
ASPLOS (2)8
2026 TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
abstract
LLM inference is increasingly limited by memory bandwidth, and the bottleneck worsens at long context as the KV cache grows. CXL memory adds capacity to offload weights and KV, but its link and device-side DDR bandwidth are far below HBM, so decoding stalls once traffic shifts to the CXL tier. Many CXL controllers are starting to add genericlosslesscompression, yet applying commodity codecs directly to standard word-major LLM tensors is largely ineffective, especially for token-major KV streams. We propose TRACE (Traffic-Reduced Architecture for Compression and Elasticity), which preserves the unmodified CXL.mem interface but changes the device-internal representation. It stores tensors in a channel-major, disaggregated bit-plane layout, and applies a KV-specific transform before compression, converting mixed-field words into low-entropy plane streams that commodity codecs can compress. The same substrate enables precision-proportional fetch by reading only the required bit-planes. Across public LLMs, TRACE reduces BF16 weight footprint by 25.2% and BF16 KV footprint by 46.9% losslessly, with per-layer KV ratios peaking at 2.69×. In tracedriven system modeling, once KV spills to CXL, GPT-OSS-120B-MXFP4 improves throughput at 128k tokens from 16.28 to 68.99 tok/s (4.24×). DRAMSim3 shows up to 40.3% lower DRAM access energy under plane-aligned fetch. A 7nm SystemVerilog implementation sustains 256 GB/s device bandwidth. Relative to a CXL controller with generic inline lossless compression, TRACE only adds 7.2% area, 4.7% power, and 6.0% load-to-use latency at 2 GHz and 0.7V.
Rui Xie 0006, Asad Ul Haq, Yunhua Fang, Linsen Ma, Zirak Burzin Engineer, Liu Liu 0017, Tong Zhang 0002
IEEE Trans. Computers6
2025 NORA: Noise-Optimized Rescaling of LLMs on Analog Compute-in-Memory Accelerators
abstract
Large Language Models (LLMs) have become critical in AI applications, yet current digital AI accelerators suffer from significant energy inefficiencies due to frequent data movement. Analog compute-in-memory (CIM) accelerators offer a potential solution for improving energy efficiency but introduce non-idealities that can degrade LLM accuracy. While analog CIM has been extensively studied for traditional deep neural networks, its impact on LLMs remains unexplored, particularly concerning the large influence of Analog CIM non-idealities. In this paper, we conduct a sensitivity analysis on the effects of analog-induced noise on LLM accuracy. We find that while LLMs demonstrate robustness to weight-related noise, they are highly sensitive to quantization noise and additive Gaussian noise. Based on these insights, we propose a noise-optimized rescaling method to mitigate LLM accuracy loss by shifting the non-ideality burden from the sensitive input/output to the more resilient weight. Through rescaling, we can implement the OPT-6.7b model on simulated analog CIM hardware with less than 1% accuracy loss from the floating-point baseline, compared to a much higher loss of around 30% without rescaling.
Yayue Hou, Hsinyu Tsai, Kaoutar El Maghraoui, Tayfun Gokmen, Geoffrey W. Burr, Liu Liu 0017
DATE6
2025 SAGE: Saliency-Aware Grouping for Efficient Mapping of LLMs on Analog Compute-in-Memory
abstract
Large Language Models (LLMs) demand high memory bandwidth and computational efficiency, posing significant challenges for deployment on traditional digital accelerators. Analog Compute-in-Memory (ACIM) architectures offer an attractive alternative by co-locating storage and computation to reduce data movement. However, executing LLMs on ACIM systems remains challenging due to hardware non-idealities and the unique statistical properties of LLM inputs and outputs in FC layers. In particular, long-tailed data distributions containing large-amplitude "salient values" degrade analog signal quality under quantization and system noise. In this work, we propose SAGE (Saliency-Aware Grouping for Efficient Mapping), a training-free strategy that improves noise resilience by reordering weight and input channels of FC layers based on statistical characteristics of LLMs. We identify kurtosis as a key factor affecting analog robustness and develop a saliency-aware mapping method that reduces output kurtosis to enhance the signal-to-noise ratio. We further introduce a reconfigurable tile design that supports mixed-precision execution and maximizes array utilization across layers. Evaluations on multiple LLMs and benchmarks show that SAGE significantly improves inference accuracy and energy efficiency based on ACIM simulation without requiring retraining.
Yayue Hou, Garrett Gagnon, Hsinyu Tsai, Kaoutar El Maghraoui, Geoffrey W. Burr, Liu Liu 0017
ICCAD7
2025 BitWeaver: Read-Time Truncation in Memory
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in generating contextually relevant responses to prompts, but their inference performance is often constrained by a severe memory bottleneck in the self-attention stage.This bottleneck, which is inherently memory-bound, has led to extensive research into strategies for reducing the size of the Key-Value (KV) cache.Many existing approaches employ quantization to lower the data precision and reduce data volume.However, these methods are constrained by memory technology, which requires the data read from memory to match the data written into it.As a result, such strategies must apply reductions at write-time, limiting either model performance or achievable speedup.We identify that enabling precision scaling at read-time -after data has been stored in memory -offers a unique opportunity to simultaneously reduce memory traffic and retain model accuracy.To this end, we develop a read-time precision-scaling mechanism and introduce BitWeaver, a hardware-enabled solution for in-memory truncation.BitWeaver dynamically reduces data precision during memory reads, achieving up to a 3× increase in memory throughput and execution speedups of up to 80% compared to baseline.Additionally, BitWeaver enhances sparse KV cache strategies by improving the efficiency of state-of-the-art sparsity techniques.
Garrett Gagnon, Srikanth Malla, Yangwook Kang, Liu Liu 0017
ICS4
2023 ECSSD: Hardware/Data Layout Co-Designed In-Storage-Computing Architecture for Extreme Classification
abstract
With the rapid growth of classification scale in deep learning systems, the final classification layer becomes extreme classification with a memory footprint exceeding the main memory capacity of the CPU or GPU. The emerging in-storage-computing technique offers an opportunity on account of the fact that SSD has enough storage capacity for the parameters of extreme classification. However, the limited performance of naive in-storage-computing schemes is insufficient to support the heavy workload of extreme classification.
Siqi Li 0013, Fengbin Tu, Liu Liu 0017, Jilan Lin, Zheng Wang 0075, Yangwook Kang, Yufei Ding 0001, Yuan Xie 0001
ISCA3
2023 Dynamic N: M Fine-Grained Structured Sparse Attention Mechanism
abstract
Transformers are becoming the mainstream solutions for various tasks like NLP and Computer vision. Despite their success, the high complexity of the attention mechanism hinders them from being applied to latency-sensitive tasks. One opportunity to accelerate the attention mechanism is leveraging the sparsity in the attention weight matrix. However, due to the dilemma between "dynamic" and "fine-grained", previous studies fail to achieve speedup on GPUs under moderate sequence lengths. They also require costly retraining to recover accuracy. In this paper, we present DFSS, the first GPU-friendly dynamic fine-grained pruning mechanism, to address this dilemma. DFSS dynamically prunes the full attention score matrix to N:M fine-grained structured sparse pattern. Our key insight is that on the dynamic side, N:M sparsity is friendly to pruning and encoding the sparse matrix on GPU. On the fine-grained side, it always preserves the dominant entries in each row. We develop a dynamic sampled dense-dense matrix multiplication kernel, first of its kind, that multiplies the query and key matrices, prunes the result, and encodes the compressed sparse matrix without overhead. Compared with previous studies, DFSS achieves speedup in arbitrary sequence lengths. It only takes a few fine-tuning epochs to reach on-par accuracy with full attention mechanism. We provide both theoretical and empirical evidence to demonstrate DFSS is a good approximation of the full attention mechanism. We evaluate the 1:2 and 2:4 sparsity under different settings and achieve 1.38 ~ 1.86× speedups over the full-attention on A100 GPU. On tasks from various domains with sequence lengths from 384 to 4096, its accuracy is on par with the full attention after only a couple of finetuning epochs from the dense pre-trained model.
Zhaodong Chen 0001, Zheng Qu 0002, Yuying Quan, Liu Liu 0017, Yufei Ding 0001, Yuan Xie 0001
PPoPP4
2022 DOTA: detect and omit weak attentions for scalable transformer acceleration
abstract
Transformer Neural Networks have demonstrated leading performance in many applications spanning over language understanding, image processing, and generative modeling. Despite the impressive performance, long-sequence Transformer processing is expensive due to quadratic computation complexity and memory consumption of self-attention. In this paper, we present DOTA, an algorithm-architecture co-design that effectively addresses the challenges of scalable Transformer inference. Based on the insight that not all connections in an attention graph are equally important, we propose to jointly optimize a lightweight Detector with the Transformer model to accurately detect and omit weak connections during runtime. Furthermore, we design a specialized system architecture for end-to-end Transformer acceleration using the proposed attention detection mechanism. Experiments on a wide range of benchmarks demonstrate the superior performance of DOTA over other solutions. In summary, DOTA achieves 152.6x and 4.5x performance speedup and orders of magnitude energy-efficiency improvements over GPU and customized hardware, respectively.
Zheng Qu 0002, Liu Liu 0017, Fengbin Tu, Zhaodong Chen 0001, Yufei Ding 0001, Yuan Xie 0001
ASPLOS2
2022 A one-for-all and o(v log(v ))-cost solution for parallel merge style operations on sorted key-value arrays
abstract
The processing of sorted key-value arrays using a “merge style operation (MSO)” is a very basic and important problem in domains like scientific computing, deep learning, database, graph analysis, sorting, set-operation etc. MSOs dominate the execution time in some important applications like SpGEMM and graph mining. For example, sparse vector addition as an MSO takes up to 98% execution time in SpGEMM in our experiment. For this reason, accelerating MSOs on CPU, GPU, and accelerators using parallel execution has been extensively studied but the solutions in prior work have three major limitations. (1) They treat different MSOs as isolated problems using incompatible methods and an unified solution is still lacking. (2) They do not have the flexibility to support variable key/value sizes and value calculations in the runtime given a fixed hardware design. (3) They require a quadratic hardware cost (O(V2)) for given parallelism V in most cases.
Bangyan Wang, Lei Deng 0003, Fei Sun 0002, Guohao Dai 0001, Liu Liu 0017, Yu Wang 0002, Yuan Xie 0001
ASPLOS5
2022 INSPIRE: in-storage private information retrieval via protocol and architecture co-design
abstract
Private Information Retrieval (PIR) plays a vital role in secure, database-centric applications. However, existing PIR protocols explore a massive working space containing hundreds of GiBs of query and database data. As a consequence, PIR performance is severely bounded by storage communication, making it far from practical for real-world deployment.
Jilan Lin, Ling Liang 0003, Zheng Qu 0002, Ishtiyaque Ahmad, Liu Liu 0017, Fengbin Tu, Trinabh Gupta, Yufei Ding 0001, Yuan Xie 0001
ISCA5
2022 Dynamic Sparse Attention for Scalable Transformer Acceleration
abstract
Transformers are the mainstream of NLP applications and are becoming increasingly popular in other domains such as Computer Vision. Despite the improvements in model quality, the enormous computation costs make Transformers difficult at deployment, especially when the sequence length is large in emerging applications. Processing attention mechanism as the essential component of Transformer is the bottleneck of execution due to the quadratic complexity. Prior art explores sparse patterns in attention to support long sequence modeling, but those pieces of work are on static or fixed patterns. We demonstrate that the sparse patterns are dynamic, depending on input sequences. Thus, we propose the Dynamic Sparse Attention (DSA) that can efficiently exploit dynamic sparse patterns in attention. Compared with other methods, our approach can achieve better trade-offs between accuracy and model complexity. Moving forward, we identify challenges and provide solutions to implement DSA on existing hardware (GPUs) and specialized hardware in order to achieve practical speedup and efficiency improvements for Transformer execution.
Liu Liu 0017, Zheng Qu 0002, Zhaodong Chen 0001, Fengbin Tu, Yufei Ding 0001, Yuan Xie 0001
IEEE Trans. Computers1
2021 ENMC: Extreme Near-Memory Classification via Approximate Screening
abstract
Extreme classification (XC) is the essential component of large-scale Deep Learning Systems for a wide range of application domains, including image recognition, language modeling, and recommendation. As classification categories keep scaling in real-world applications, the classifier’s parameters could reach several thousands of Gigabytes, way exceed the on-chip memory capacity. With the advent of near-memory processing (NMP) architectures, offloading the XC component onto NMP units could alleviate the memory-intensive problem. However, naive NMP design with limited area and power budget cannot afford the computational complexity of full classification. To tackle the problem, we first propose a novel screening method to reduce the computation and memory consumption by efficiently approximating the classification output and identifying a small portion of key candidates that require accurate results. Then, we design a new extreme-classification-tailored NMP architecture, namely ENMC, to support both screening and candidates-only classification. Overall, our approximate screening method achieves 7.3 × speedup over the CPU baseline, and ENMC further improves the performance by 7.4 × and demonstrates 2.7 × speedup compared with the state-of-the-art NMP baseline.
Liu Liu 0017, Jilan Lin, Zheng Qu 0002, Yufei Ding 0001, Yuan Xie 0001
MICRO1
2021 Efficient tensor core-based GPU kernels for structured sparsity under reduced precision
abstract
The success of DNN comes at the expense of excessive memory/computation cost, which can be addressed by exploiting reduced precision and sparsity jointly. Existing sparse GPU kernels, however, fail to achieve practical speedup over cuBLASHgemm under half-precision. Those for fine-grained sparsity suffer from low data reuse, and others for coarse-grained sparsity are limited by the wrestling between kernel performance and model quality under different grain sizes. We propose column-vector-sparse-encoding that has a smaller grain size under the same reuse rate compared with block sparsity. Column-vector-sparse-encoding can be applied to both SpMM & SDDMM, two major sparse DNN operations. We also introduce the Tensor-Core-based 1D Octet Tiling that has efficient memory access and computation patterns under small grain size. Based on these, we design SpMM and SDDMM kernels and achieve 1.71-7.19x speedup over cuSPARSE. Practical speedup is achieved over cuBLASHgemm under >70% and >90% sparsity with 4x1 grain size and half-precision.
Zhaodong Chen 0001, Zheng Qu 0002, Liu Liu 0017, Yufei Ding 0001, Yuan Xie 0001
SC3
2020 INVITED: Computation on Sparse Neural Networks and its Implications for Future Hardware
abstract
Neural network models are widely used in solving many challenging problems, such as computer vision, personalized recommendation, and natural language processing. Those models are very computationally intensive and reach the hardware limit of the existing server and IoT devices. Thus, finding better model architectures with much less amount of computation while maximally preserving the accuracy is a popular research topic. Among various mechanisms that aim to reduce the computation complexity, identifying the zero values in the model weights and in the activations to avoid computing them is a promising direction. In this paper, we summarize the current status of the research on the computation of sparse neural networks, from the perspective of the sparse algorithms, the software frameworks, and the hardware accelerations. We observe that the search for the sparse structure can be a general methodology for high-quality model explorations, in addition to a strategy for high-efficiency model execution. We discuss the model accuracy influenced by the number of weight parameters and the structure of the model. The corresponding models are called to be located in the weight dominated and structure dominated regions, respectively. We show that for practically complicated problems, it is more beneficial to search large and sparse models in the weight dominated region. In order to achieve the goal, new approaches are required to search for proper sparse structures, and new sparse training hardware needs to be developed to facilitate fast iterations of sparse models.
Fei Sun 0002, Minghai Qin, Tianyun Zhang, Liu Liu 0017, Yen-Kuang Chen, Yuan Xie 0001
DAC4
2020 Boosting Deep Neural Network Efficiency with Dual-Module Inference
abstract
Using deep neural networks (DNNs) in machine learning tasks is promising in delivering high-quality results but challenging to meet stringent latency requirements and energy constraints because of the memory-bound and the compute-bound execution pattern of DNNs. We propose a big-little dual-module inference to dynamically skip unnecessary memory accesses and computations to accelerate DNN inference. Leveraging the noise-resilient feature of nonlinear activation functions, we propose to use a lightweight little module that approximates the original DNN layer, termed as the big module, to compute activations of the insensitive region that are more noise-resilient. Hence, the expensive memory accesses and computations of the big module can be reduced as the results are only calculated in the sensitive region. For memory-bound models such as recurrent neural networks (RNNs), our method can reduce the overall memory accesses by 40% on average and achieve 1.54x to 1.75x speedup on a commodity CPU-based server platform with a negligible impact on model quality. In addition, our method can reduce the operations of the compute-bound models such as convolutional neural networks (CNNs) by 3.02x, with only a 0.5% accuracy drop.
Liu Liu 0017, Lei Deng 0003, Zhaodong Chen 0001, Shuangchen Li, Yihua Yang, Yufei Ding 0001, Yuan Xie 0001
ICML1
2020 DUET: Boosting Deep Neural Network Efficiency on Dual-Module Architecture
abstract
Deep Neural Networks (DNNs) have been driving the mainstream of Machine Learning applications. However, deploying DNNs on modern hardware with stringent latency requirements and energy constraints is challenging because of the compute-intensive and memory-intensive execution patterns of various DNN models. We propose an algorithm-architecture co-design to boost DNN execution efficiency. Leveraging the noise resilience of nonlinear activation functions in DNNs, we propose dual-module processing that uses approximate modules learned from original DNN layers to compute insensitive activations. Therefore, we can save expensive computations and data accesses of unnecessary sensitive activations. We then design an Executor-Speculator dual-module architecture with support for balance execution and memory access reduction. With acceptable model inference quality degradation, our accelerator design can achieve 2.24x speedup and 1.97x energy efficiency improvement for compute-bound Convolutional Neural Networks (CNNs) and memory-bound Recurrent Neural Networks (RNNs).
Liu Liu 0017, Zheng Qu 0002, Lei Deng 0003, Fengbin Tu, Shuangchen Li, Xing Hu 0001, Yufei Ding 0001, Yuan Xie 0001
MICRO1
2020 SemiMap: A Semi-Folded Convolution Mapping for Speed-Overhead Balance on Crossbars
abstract
Crossbar architecture has been widely used in neural network (NN) accelerators, involving conventional and emerging devices. It performs well on the fully connected layer through efficient vector-matrix multiplication. Whereas, the advantages degrade on the convolutional layer with huge data reuse, since the execution speed and resource overhead are imbalanced when using existing fully unfolded or fully folded mapping strategy. To address this issue, we propose a novel semi-folded mapping (SemiMap) framework for implementing the convolution on crossbars. It simultaneously folds the physical resources along the row dimension of feature maps (FMs) and unfolds them along the column dimension. The former reduces the resource overhead, and the latter maintains the parallelism. An FM slicing scheme is further proposed to enable the processing of large-size image. Via our mapping framework, a row-by-row streaming pipeline for intraimage dataflow and periodical pipeline for interimage dataflow are easy to be obtained. To validate the idea, we build a many-crossbar architecture with several designs to guarantee the overall functionality and performance. Based on the measurement data of a fabricated chip, a mapping compiler and a cycle-accurate simulator are developed for the hardware simulation of large-scale networks. We evaluate the proposed SemiMap on various convolutional NNs across different network scale. ${>} 35 {\times }$ resource saving and several hundred times cycle reduction are demonstrated compared to the existing fully unfolded and fully folded strategies, respectively. This paper jumps out of the current extreme mapping schemes, and provides a balanced solution on how to efficiently deploy the computational graphs with data reuse on many-crossbar architecture.
Lei Deng 0003, Yuan Xie 0001, Ling Liang 0003, Guanrui Wang, Liang Chang 0002, Xing Hu 0001, Liu Liu 0017, Jing Pei, Guoqi Li 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2019 Dynamic Sparse Graph for Efficient Deep Learning
Liu Liu 0017, Lei Deng 0003, Xing Hu 0001, Maohua Zhu, Guoqi Li 0002, Yufei Ding 0001, Yuan Xie 0001
ICLR (Poster)1
2019 L1-Norm Batch Normalization for Efficient Training of Deep Neural Networks
abstract
Batch normalization (BN) has recently become a standard component for accelerating and improving the training of deep neural networks (DNNs). However, BN brings in additional calculations, consumes more memory, and significantly slows down the training iteration. Furthermore, the nonlinear square and sqrt operations in the normalization process impede low bit-width quantization techniques, which draw much attention to the deep learning hardware community. In this paper, we propose an$L1$-norm BN (L1BN) with only linear operations in both forward and backward propagations during training. L1BN is approximately equivalent to the conventional$L2$-norm BN (L2BN) by multiplying a scaling factor that equals$({\pi }/{2})^{1/2}$. Experiments on various convolutional neural networks and generative adversarial networks reveal that L1BN can maintain the same performance and convergence rate as L2BN but with higher computational efficiency. In real application-specified integrated circuit synthesis with reduced resources, L1BN achieves 25% speedup and 37% energy saving compared to the original L2BN. Our hardware-friendly normalization method not only surpasses L2BN in speed but also simplifies the design of deep learning accelerators. Last but not least, L1BN promises a fully quantized training of DNNs, which empowers future artificial intelligence applications on mobile devices with transfer and continual learning capability.
Guoqi Li 0002, Lei Deng 0003, Liu Liu 0017, Yuan Xie 0001, Luping Shi
IEEE Trans. Neural Networks Learn. Syst.4
2017 Building energy-efficient multi-level cell STT-RAM caches with data compression
abstract
Spin-transfer torque magnetic random access memory (STT-RAM) technology has emerged as a potential replacement of SRAM in cache design, especially for building large-scale and energy-efficient last level caches. Compared with single-level cell (SLC), multi-level cell (MLC) STT-RAM is expected to double cache capacity and increase system performance. However, the two-step read/write access schemes incur considerable energy consumption and performance degradation. In this paper, we propose two techniques using data compression to optimize MLC STT-RAM cache design. The first technique tries to compress a cache line and fit it into only the soft-bit region of the cells, so that reading or writing this cache line takes only one step which is fast and energy-efficient. We introduce a second technique to increase the cache capacity by enabling the left hard-bit region to store another compressed cache line, which can improve the system performance for memory intensive workloads. The experimental results show that, compared with a conventional MLC STT-RAM last level cache design, our overhead minimized technique reduces the dynamic energy consumption by 38.2% on average with the same system performance, and our capacity augmented technique boosts the system performance by 6.1% with 19.2% dynamic energy saving on average, across the evaluated multi-programmed benchmarks.
Liu Liu 0017, Ping Chi, Shuangchen Li, Yuanqing Cheng, Yuan Xie 0001
ASP-DAC1
2016 Leveraging 3D Technologies for Hardware Security: Opportunities and Challenges
abstract
3D die stacking and 2.5D interposer design are promising technologies to improve integration density, performance and cost. Current approaches face serious issues in dealing with emerging security challenges such as side channel attacks, hardware trojans, secure IC manufacturing and IP piracy. By utilizing intrinsic characteristics of 2.5D and 3D technologies, we propose novel opportunities in designing secure systems. We present: (i) a 3D architecture for shielding side-channel information; (ii) split fabrication using active interposers; (iii) circuit camouflage on monolithic 3D IC, and (iv) 3D IC-based security processing-in-memory (PIM). Advantages and challenges of these designs are discussed, showing that the new designs can improve existing countermeasures against security threats and further provide new security features.
Peng Gu 0008, Shuangchen Li, Dylan C. Stow, Russell Barnes, Liu Liu 0017, Yuan Xie 0001, Eren Kursun
ACM Great Lakes Symposium on VLSI5
2016 NVSim-CAM: a circuit-level simulator for emerging nonvolatile memory based content-addressable memory
abstract
Ternary Content-Addressable Memory (TCAM) is widely used in networking routers, fully associative caches, search engines, etc. While the conventional SRAM-based TCAM suffers from the poor scalability, the emerging nonvolatile memories (NVM, i.e., MRAM, PCM, and ReRAM) bring evolution for the TCAM design. It effectively reduces the cell size, and makes significant energy reduction and scalability improvement. New applications such as associative processors/accelerators are facilitated by the emergence of the nonvolatile TCAM (nvTCAM). However, nvTCAM design is challenging. In addition to the emerging device's uncertainty, the nvTCAM cell structure is so diverse that it results in a design space too large to explore manually. To tackle these challenges, we propose a circuit-level model and develop a simulation tool, NVSim-CAM, which helps researchers to make early design decisions, and to evaluate device/circuit innovations. The tool is validated by HSPICE simulations and data from fabricated chips. We also present a case study to illustrate how NVSim-CAM benefits the nvTCAM design. In the case study, we propose a novel 3D vertical ReRAM based TCAM cell, the 3DvTCAM. We project the advantages/disadvantages and explore the design space for the proposed cell with NVSim-CAM.
Shuangchen Li, Liu Liu 0017, Peng Gu 0008, Cong Xu 0002, Yuan Xie 0001
ICCAD2