Frank Liu 0001

dblp:18/2008 · also Frank Y. Liu, Frank Ying Liu · DBLP profile ↗
← Back
64ranked-venue papers
10as first author
14since 2021 · last 2025
0000-0001-6615-0739ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 52 · 10 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-Gather
abstract
Graph processing requires irregular, fine-grained random access patterns incompatible with contemporary off-chip memory architecture, leading to inefficient data access. This inefficiency makes graph processing an extremely memory-bound application. Because of this, existing graph processing accelerators typically employ a graph tiling-based or processing-in-memory (PIM) approach to relieve the memory bottleneck. In the tiling-based approach, a graph is split into chunks that fit within the on-chip cache to maximize data reuse. In the PIM approach, arithmetic units are placed within memory to perform operations such as reduction or atomic addition. However, both approaches have several limitations, especially when implemented on current memory standards (i.e., DDR). Because the access granularity provided by DDR is much larger than that of the graph vertex property data, much of the bandwidth and cache capacity are wasted. PIM is meant to alleviate such issues, but it is difficult to use in conjunction with the tiling-based approach, resulting in a significant disadvantage. Furthermore, placing arithmetic units inside a memory chip is expensive, thereby supporting multiple types of operation is thought to be impractical. To address the above limitations, we present Piccolo, an end-to-end efficient graph processing accelerator with fine-grained in-memory random scatter-gather. Instead of placing expensive arithmetic units in off-chip memory, Piccolo focuses on reducing the off-chip traffic with non-arithmetic function-in-memory of random scatter-gather. To fully benefit from in-memory scatter-gather, Piccolo redesigns the cache and miss-handling architecture (MHA) of the accelerator such that it can enjoy both the advantage of tiling and in-memory operations. Piccolo achieves a maximum speedup of 3.28 × and a geometric mean speedup of 1.62 ×, along with up to 59.7% reduction in energy consumption across various and extensive benchmarks.
Changmin Shin 0002, Jaeyong Song 0002, Hongsun Jang, Dogeun Kim, Jun Sung, Taehee Kwon 0002, Jae Hyung Ju, Frank Liu 0001, YeonKyu Choi, Jinho Lee 0001
HPCA8
2024 Semi-supervised Learning of Dynamical Systems with Neural Ordinary Differential Equations: A Teacher-Student Model Approach
abstract
Modeling dynamical systems is crucial for a wide range of tasks, but it remains challenging due to complex nonlinear dynamics, limited observations, or lack of prior knowledge. Recently, data-driven approaches such as Neural Ordinary Differential Equations (NODE) have shown promising results by leveraging the expressive power of neural networks to model unknown dynamics. However, these approaches often suffer from limited labeled training data, leading to poor generalization and suboptimal predictions. On the other hand, semi-supervised algorithms can utilize abundant unlabeled data and have demonstrated good performance in classification and regression tasks. We propose TS-NODE, the first semi-supervised approach to modeling dynamical systems with NODE. TS-NODE explores cheaply generated synthetic pseudo rollouts to broaden exploration in the state space and to tackle the challenges brought by lack of ground-truth system data under a teacher-student model. TS-NODE employs an unified optimization framework that corrects the teacher model based on the student's feedback while mitigating the potential false system dynamics present in pseudo rollouts. TS-NODE demonstrates significant performance improvements over a baseline Neural ODE model on multiple dynamical system modeling tasks.
Yu Wang 0167, Yuxuan Yin, Karthik Somayaji Nanjangud Suryanarayana, Ján Drgona, Malachi Schram, Mahantesh Halappanavar, Frank Liu 0001, Peng Li 0001
AAAI7
2023 AutoNF: Automated Architecture Optimization of Normalizing Flows with Unconstrained Continuous Relaxation Admitting Optimal Discrete Solution
abstract
Normalizing flows (NF) build upon invertible neural networks and have wide applications in probabilistic modeling. Currently, building a powerful yet computationally efficient flow model relies on empirical fine-tuning over a large design space. While introducing neural architecture search (NAS) to NF is desirable, the invertibility constraint of NF brings new challenges to existing NAS methods whose application is limited to unstructured neural networks. Developing efficient NAS methods specifically for NF remains an open problem. We present AutoNF, the first automated NF architectural optimization framework. First, we present a new mixture distribution formulation that allows efficient differentiable architecture search of flow models without violating the invertibility constraint. Second, under the new formulation, we convert the original NP-hard combinatorial NF architectural optimization problem to an unconstrained continuous relaxation admitting the discrete optimal architectural solution, circumventing the loss of optimality due to binarization in architectural optimization. We evaluate AutoNF with various density estimation datasets and show its superior performance-cost trade-offs over a set of existing hand-crafted baselines.
Yu Wang 0167, Ján Drgona, Jiaxin Zhang 0005, Karthik Somayaji Nanjangud Suryanarayana, Malachi Schram, Frank Liu 0001, Peng Li 0001
AAAI6
2023 FPGA Acceleration of GCN in Light of the Symmetry of Graph Adjacency Matrix
abstract
Graph Convolutional Neural Networks (GCNs) are widely used to process large-scale graph data. Different from deep neural networks (DNNs), GCNs are sparse, irregular, and unstructured, posing unique challenges to hardware acceleration with regular processing elements (PEs). In particular, the adja-cency matrix of a GCN is extremely sparse, leading to frequent but irregular memory access, low spatial/temporal data locality and poor data reuse. Furthermore, a realistic graph usually consists of unstructured data (e.g., unbalanced distributions), creating significantly different processing times and imbalanced workload for each node in GCN acceleration. To overcome these challenges, we propose an end-to-end hardware-software co-design to accelerate GCNs on resource-constrained FPGAs with the features including: (1) A custom dataflow that leverages symmetry along the diagonal of the adjacency matrix to accelerate feature aggregation for undirected graphs. We utilize either the upper or the lower triangular matrix of the adjacency matrix to perform aggregation in GCN to improve data reuse. (2) Unified compute cores for both aggregation and transform phases, with full support to the symmetry-based dataflow. These cores can be dynamically reconfigured to the systolic mode for transformation or as individual accumulators for aggregation in GCN processing. (3) Preprocessing of the graph in software to rearrange the edges and features to match the custom dataflow. This step improves the regularity in memory access and data reuse in the aggregation phase. Moreover, we quantize the GCN precision from FP32 to INT8 to reduce the memory footprint without losing the inference accuracy. We implement our accelerator design in Intel Stratix10 MX FPGA board with HBM2, and demonstrate$1.3\times-110.5\times$improvement in end-to-end GCN latency as compared to the state-of the-art FPGA implementations, on the graph datasets of Cora, Pubmed, Citeseer and Reddit.
Gopikrishnan Raveendran Nair, Han-Sok Suh, Mahantesh Halappanavar, Frank Liu 0001, Jae-sun Seo, Yu Cao 0001
DATE4
2023 Disentangling Learning Representations with Density Estimation
Eric C. Yeats, Frank Liu 0001, Hai Li 0001
ICLR2
2023 Accelerating Scientific Simulations with Bi-Fidelity Weighted Transfer Learning
abstract
High-fidelity modeling is an essential design tool for many engineering applications. However, for complex systems, computational cost can be a limiting factor. Analyzing parameter sensitivity, uncertainty quantification, and design optimization require many model evaluations. Surrogate models are often used to develop the relationship between model parameters and quantities of interest. However, in the case of complex systems, surrogate models require several degrees of freedom and, thus, a large number of data points to determine the correct dependencies. For many applications, this may be prohibitively expensive. The reduction of computational requirements can be achieved by leveraging low-fidelity models. Low-fidelity models represent the system at a coarser resolution with the advantage of computational efficiency. Therefore, a bi-fidelity modeling paradigm, which augments the accuracy of a low-fidelity model in a computationally efficient manner by invoking limited runs of a high-fidelity model, can be leveraged to sufficiently balance the accuracy and computational requirements. In this work, a bi-fidelity weighted transfer learning method using neural networks was applied to a computational fluid dynamics heat transfer modeling problem. The transfer learning advantage was investigated as a function of hyperparameters. Our main finding is that the use of a bi-fidelity modeling paradigm achieves accuracy close to that of a high-fidelity Gaussian process model while significantly reducing computational cost. The bi-fidelity model achieves comparable performance with 90 high-fidelity samples-that is, 60% less than the samples needed to achieve similar accuracy without the use of bi-fidelity modeling,
Katarzyna Borowiec, Dan Lu 0001, Vikas Chandan, Samrat Chatterjee, Pradeep Ramuhalli, Ramakrishna Tipireddy, Mahantesh Halappanavar, Frank Liu 0001
ICMLA8
2023 A 3D Implementation of Convolutional Neural Network for Fast Inference
abstract
Low latency inference has many applications in edge machine learning. In this paper, we present a run-time configurable convolutional neural network (CNN) inference ASIC design for low-latency edge machine learning. By implementing a 5-stage pipelined CNN inference model in a 3D ASIC technology, we demonstrate that the model distributed on two dies utilizing face-to-face (F2F) 3D integration achieves superior performance. Our experimental results show that the design based on 3D integration achieves 43% better energy-delay product when compared to the traditional 2D technology.
Narasinga Rao Miniskar, Pruek Vanna-Iampikul, Aaron R. Young, Sung Kyu Lim, Frank Liu 0001, Jieun Yoo, Corrinne Mills, Farah Fahim, Jeffrey S. Vetter
ISCAS5
2022 Gradient-Based Novelty Detection Boosted by Self-Supervised Binary Classification
abstract
Novelty detection aims to automatically identify out-of-distribution (OOD) data, without any prior knowledge of them. It is a critical step in data monitoring, behavior analysis and other applications, helping enable continual learning in the field. Conventional methods of OOD detection perform multi-variate analysis on an ensemble of data or features, and usually resort to the supervision with OOD data to improve the accuracy. In reality, such supervision is impractical as one cannot anticipate the anomalous data. In this paper, we propose a novel, self-supervised approach that does not rely on any pre-defined OOD data: (1) The new method evaluates the Mahalanobis distance of the gradients between the in-distribution and OOD data. (2) It is assisted by a self-supervised binary classifier to guide the label selection to generate the gradients, and maximize the Mahalanobis distance. In the evaluation with multiple datasets, such as CIFAR-10, CIFAR-100, SVHN and TinyImageNet, the proposed approach consistently outperforms state-of-the-art supervised and unsupervised methods in the area under the receiver operating characteristic (AUROC) and area under the precision-recall curve (AUPR) metrics. We further demonstrate that this detector is able to accurately learn one OOD class in continual learning.
Jingbo Sun 0003, Li Yang 0009, Jiaxin Zhang 0005, Frank Liu 0001, Mahantesh Halappanavar, Deliang Fan, Yu Cao 0001
AAAI4
2022 NashAE: Disentangling Representations Through Adversarial Covariance Minimization
Eric C. Yeats, Frank Liu 0001, David Womble, Hai Li 0001
ECCV (27)2
2022 Ultra Low Latency Machine Learning for Scientific Edge Applications
abstract
In this paper, we present an FPGA design of an extremely low latency scientific machine learning application at the edge. Real-time prediction of errant high-energy particle beams at scientific facilities such as Spallation Neutron Source (SNS) is crucial to avoid damages to the equipment. Machine learning techniques are becoming increasingly effective to detect subtle signatures of the errant beams in the noisy sensor signals. However, to minimize potential damage done by errant beam, real-time errant beam detection has to be completed with extremely low latency, usually less than 1 microsecond. By stream processing the input features and employing out-of-order execution of decision nodes among the decision trees, we demonstrate that our highly efficient FPGA implementation can achieve 60 nanoseconds of computing latency for complex random forest models with 10,000 input features.
Narasinga Rao Miniskar, Aaron R. Young, Frank Liu 0001, Willem Blokland, Anthony M. Cabrera, Jeffrey S. Vetter
FPL3
2022 IRIS-BLAS: Towards a Performance Portable and Heterogeneous BLAS Library
abstract
This paper presents IRIS-BLAS, a novel heterogeneous and performance portable BLAS library. IRIS-BLAS is built on top of the IRIS runtime and multiple vendor and open-source BLAS libraries. It can transparently use all the architectures/devices available in a heterogeneous system, using the appropriate BLAS library based on the task mapping at run time. Thus, IRIS-BLAS is portable across a broad spectrum of architectures and BLAS libraries, alleviating the worry of application developers about modifying the application source code. Even though the emphasis is on portability, IRIS-BLAS provides competitive or even better performance than other state-of-the-art references. Moreover, IRIS-BLAS offers new features such as efficiently using extremely heterogeneous systems composed of multiple GPUs from different hardware vendors.
Narasinga Rao Miniskar, Mohammad Alaul Haque Monil, Pedro Valero-Lara, Frank Liu 0001, Jeffrey S. Vetter
HIPC4
2021 Evolutionary NAS in Light of Model Stability for Accurate Continual Learning
abstract
Continual learning, the capability to learn new knowledge from streaming data without forgetting the previous knowledge, is a critical requirement for dynamic learning systems, especially for emerging edge devices such as self-driving cars and drones. However, continual learning is still facing the catastrophic forgetting problem. Previous work illustrate that model performance on continual learning is not only related to the learning algorithms but also strongly dependent on the inherited model, i.e., the model where continual learning starts. The better stability of the inherited model, the less catastrophic forgetting and thus, the inherited model should be elaborately selected. Inspired by this finding, we develop an evolutionary neural architecture search (ENAS) algorithm that emphasizes the Stability of the inherited model, namely ENAS-S. ENAS-S aims to find optimal architectures for accurate continual learning on edge devices. On CIFAR-10 and CIFAR-100, we present that ENAS-S achieves competitive architectures with lower catastrophic forgetting and smaller model size when learning from a data stream, as compared with handcrafted DNNs.
Xiaocong Du, Zheng Li 0020, Jingbo Sun 0003, Frank Liu 0001, Yu Cao 0001
IJCNN4
2021 A Memory Efficient Lock-Free Circular Queue
abstract
Hardware queues are import in many applications, such as data transfer, synchronization of concurrent modules with the need of mutual exclusion constructs. State of the art bounded (of a fixed size) lock free circular queues are implemented either by read/write atomic operations, or barrier conditions, or by separating dequeue and enqueue operations. However, these queues always require an unused element at all the times to safeguard the front and rear pointers of the queue, so as to avoid data race conditions, which leads to the waste of memory. The waste of memory is especially disadvantageous in applications such as I/O data transfer, and image transfer between processing filters, when large element size is needed, We propose a lock- free solution of the bounded circular queue through read/write atomic operations, but without the need of an extra element in the queue. The proposed solution is implemented and verified in both Verilog and 'C' languages. We also demonstrate its effectiveness by comparing its area and delay metrics with the implementations of other existing designs of queue.
Narasinga Rao Miniskar, Frank Liu 0001, Jeffrey S. Vetter
ISCAS2
2021 On the Stochastic Stability of Deep Markov Models
abstract
Deep Markov models (DMM) are generative models which are scalable and expressive generalization of Markov models for representation, learning, and inference problems. However, the fundamental stochastic stability guarantees of such models have not been thoroughly investigated. In this paper, we present a novel stability analysis method and provide sufficient conditions of DMM's stochastic stability. The proposed stability analysis is based on the contraction of probabilistic maps modeled by deep neural networks. We make connections between the spectral properties of neural network's weights and different types of used activation function on the stability and overall dynamic behavior of DMMs with Gaussian distributions. Based on the theory, we propose a few practical methods for designing constrained DMMs with guaranteed stability. We empirically substantiate our theoretical results via intuitive numerical experiments using the proposed stability constraints.
Ján Drgona, Sayak Mukherjee, Jiaxin Zhang 0005, Frank Liu 0001, Mahantesh Halappanavar
NeurIPS4
2020 Scaled Population Arithmetic for Efficient Stochastic Computing
abstract
We propose a new Scaled Population (SP) based arithmetic computation approach that achieves considerable improvements over existing stochastic computing (SC) techniques. First, SP arithmetic introduces scaling operations that significantly reduce the numerical errors as compared to SC. Experiments show accuracy improvements of a single multiplication and addition operation by 6.3× and 4.0×, respectively. Secondly, SP arithmetic erases the inherent serialization associated with stochastic computing, thereby significantly improves the computational delays. We design each of the operations of SP arithmetic to take O(1) gate delays, and eliminate the need of serially iterating over the bits of the population vector. Our SP approach improves the area, delay and power compared with conventional stochastic computing on an FPGA-based implementation. We also apply our SP scheme on a handwritten digit recognition application (MNIST), improving the recognition accuracy by 32.79% compared to SC.
Sunil P. Khatri, Jiang Hu 0001, Frank Liu 0001
ASP-DAC4
2020 Deffe: a data-efficient framework for performance characterization in domain-specific computing
abstract
As the computer architecture community moves toward the end of traditional device scaling, domain-specific architectures are becoming more pervasive. Given the number of diverse workloads and emerging heterogeneous architectures, exploration of this design space is a constrained optimization problem in a high-dimensional parameter space. In this respect, predicting workload performance both accurately and efficiently is a critical task for this exploration. In this paper, we present Deffe: a framework to estimate workload performance across varying architectural configurations. Deffe uses machine learning to improve the performance of this design space exploration. By casting the work of performance prediction itself as transfer learning tasks, the modelling component of Deffe can leverage the learned knowledge on one workload and "transfer" it to a new workload. Our extensive experimental results on a contemporary architecture toolchain (RISC-V and GEM5) and infrastructure show that the method can achieve superior testing accuracy with an effective reduction of 32-80× in terms of the amount of required training data. The overall run-time can be reduced from 400 hours to 5 hours when executed over 24 CPU cores. The infrastructure component of Deffe is based on scalable and easy-to-use open-source software components.
Frank Liu 0001, Narasinga Rao Miniskar, Dwaipayan Chakraborty, Jeffrey S. Vetter
CF1
2020 Online Knowledge Acquisition with the Selective Inherited Model
abstract
Continual learning, which updates machine learning models according to streaming data, is increasingly needed in the dynamic systems. Such a scenario requires both the preservation of previous knowledge, as well as the adaptation to new observations, with high computational and memory efficiency at the edge. Previous approaches attempt to learn the knowledge class by class from scratch, using either regularization based or memory replay-based methods. However, they still suffer from severe accuracy drop, a.k.a catastrophic forgetting, during this incremental process. Moreover, as the entire model is involved in each updating, their computation cost is too expensive for edge computing. In this work, we propose a novel brain- inspired paradigm named acquisitive learning (AL). Different from previous approaches that focus only on model adaptation, AL emphasizes the importance of both knowledge inheritance and acquisition: the model is first pre-trained and selected in the cloud (the selective inherited model) and then adapted to new knowledge (the acquisition). The quality of the inherited model is monitored by the landscape of the loss function, while the acquisition is realized by segmented training. The combination of both steps reduces accuracy drop by >10× on the CIFAR- 100 dataset. Furthermore, AL benefits edge computing with 5× reduction in latency per training image on FPGA prototype and 150× reduction in training FLOPs.
Xiaocong Du, Shreyas K. Venkataramanaiah, Zheng Li 0020, Jae-sun Seo, Frank Liu 0001, Yu Cao 0001
IJCNN5
2019 A Memory-Efficient Markov Decision Process Computation Framework Using BDD-based Sampling Representation
abstract
Although Markov Decision Process (MDP) has wide applications in autonomous systems as a core model in Reinforcement Learning, a key bottleneck is the large memory utilization of the state transition probability matrices. This is particularly problematic for computational platforms with limited memory, or for Bayesian MDP, which requires dozens of such matrices. To mitigate this difficulty, we propose a highly memory-efficient representation for probability matrices using Binary Decision Diagram (BDD) based sampling, and develop a corresponding (Bayesian/classical) MDP solver on a CPU-GPU platform. Simulation results indicate our approach reduces memory by one and two orders of magnitude for Bayesian/classical MDP, respectively.
Sunil P. Khatri, Jiang Hu 0001, Frank Liu 0001
DAC4
2019 Single-Net Continual Learning with Progressive Segmented Training
abstract
There is an increasing need of continual learning in dynamic systems, such as the self-driving vehicle, the surveillance drone, and the robotic system. Such a system requires learning from the data stream, training the model to preserve previous information and adapt to a new task, and generating a single-headed vector for future inference. Different from previous approaches with dynamic structures, this work focuses on a single network and model segmentation to prevent catastrophic forgetting. Leveraging the redundant capacity of a single network, model parameters for each task are separated into two groups: one important group which is frozen to preserve current knowledge, and secondary group to be saved (not pruned) for a future learning. A fixed-size memory containing a small amount of previously seen data is further adopted to assist the training. Without additional regularization, the simple yet effective approach of Progressive Segmented Training (PST) successfully incorporates multiple tasks and achieves the state-of-the-art accuracy in the single-head evaluation on CIFAR-10 and CIFAR-100 datasets. Moreover, the segmented training significantly improves computation efficiency in continual learning at the edge.
Xiaocong Du, Gouranga Charan, Frank Liu 0001, Yu Cao 0001
ICMLA3
2019 An Efficient Graph Compressor Based on Adaptive Prefix Encoding
abstract
In this paper we introduce APEC, a graph compression/decompression framework. A key component of APEC is adaptive prefix code, a novel variable-length coding scheme which can adapt to varying characteristics of different vertices in the graph data. APEC also encompasses many software optimization techniques including compressed vertex indexing, bit counting and parallelization. The net outcome is that APEC not only achieves up to 20% improvement on compression ratio, which is equivalent to 2.28 bits/edge, but also as much as 9x faster in compression and up to 20x faster in decompression compared to the existing frameworks. Moreover, APEC is capable of random accessing compressed data and performing compression on extremely large graph datasets.
Jinho Lee 0001, Frank Liu 0001
SSDBM2
2018 Optimization of Genomics Analysis Pipeline for Scalable Performance in a Cloud Environment
Carlos H. A. Costa, Claudia Misale, Frank Liu 0001, Marcio Silva, Hubertus Franke, Paul Crumley, Bruce D'Amora
BIBM3
2015 Efficient Transient Analysis of Power Delivery Network With Clock/Power Gating by Sparse Approximation
abstract
Transient analysis of large-scale power delivery network (PDN) is a critical task to ensure the functional correctness and desired performance of today's integrated circuits (ICs), especially if significant transient noises are induced by clock and/or power gating due to the utilization of extensive power management. In this paper, we propose an efficient algorithm for PDN transient analysis based on sparse approximation. The key idea is to exploit the fact that the transient response caused by clock/power gating is often localized and the voltages at many other “inactive” nodes are almost unchanged, thereby rendering a unique sparse structure. By taking advantage of the underlying sparsity of the solution structure, a modified conjugate gradient algorithm is developed and tuned to efficiently solve the PDN analysis problem with low computational cost. Our numerical experiments based on standard benchmarks demonstrate that the proposed transient analysis with sparse approximation offers up to 2.2× runtime speedup over other traditional methods, while simultaneously achieving similar accuracy.
Hengliang Zhu, Yuanzhe Wang, Frank Liu 0001, Xin Li 0001, Xuan Zeng 0001, Peter Feldmann
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 A Time-Unrolling Method to Compute Sensitivity of Dynamic Systems
abstract
Sensitivities of the dynamic system responses with respect to the system parameters are highly valuable, with broad applications such as system tuning and uncertainty quantification. Compared to the direct methods, adjoint methods are much more efficient when the number of parameters is large. In this paper, we present a time-unrolling method to compute adjoint sensitivities. Instead of explicitly constructing the adjoint system, which quite often is nontrivial, our time-unrolling method implicitly retrace the response trajectory by utilizing the fitting polynomial of the integration methods. This paper provides theoretical foundation of the method as well as experimental demonstrations of its effectiveness.
Frank Liu 0001, Peter Feldmann
DAC1
2014 IFM: A Scalable High Resolution Flood Modeling Framework
Swati Singhal, Sandhya Aneja, Frank Liu 0001, Lucas Correia Villa Real, Thomas George
Euro-Par3
2012 Dynamic river network simulation at large scale
abstract
Fully dynamic modeling of large scale river networks is still a challenge. In this paper we describe SPRINT, an inter-disciplinary collaborative effort between computer engineering and hydroscience to address the computational aspect of this challenge. Although algorithmic details differ, SPRINT draws many design considerations from SPICE, one of the most fundamental EDA tools. Experimental results demonstrate that SPRINT is capable of simulating large river basins at over 100x faster than real time.
Frank Liu 0001, Ben R. Hodges
DAC1
2012 2012 TAU power grid simulation contest: Benchmark suite and results
abstract
Although power grid analysis has been an active research area for a number of years, increasing chip size has exposed new challenges in this traditional topic. The simulation of these large scale networks is becoming a dominant step in the design verification flow and it often requires the very largest computer available to the design team. To spur academic research in this vital verification step, the IBM Austin Research Laboratory, with support from the ACM TAU Workshop, has successfully organized two annual TAU Power Grid Simulation Contests, and over twenty university teams across the world have participated. For 2012, the contest is focused on dynamic analysis and parallel implementation.
Zhuo Li 0001, Raju Balasubramanian, Frank Liu 0001, Sani R. Nassif
ICCAD3
2011 An efficient mask optimization method based on homotopy continuation technique
abstract
In sub-wavelength lithography, traditional resolution enhancement techniques (e.g., OPC) cannot guarantee the optimality of the mask. In this paper, we present a novel inverse lithography method to solve the mask optimization problem. Recognizing that when formulated on a pixel-by-pixel basis with partially coherent optical models, the problem is a large-scale nonlinear optimization problem, we cast the optimization flow into a homotopy framework and apply an efficient numerical continuation technique. Compared to earlier pixel-based inverse lithography methods, our homotopy approach is not only more efficient, but also capable of naturally addressing the mask manufactureability problem. Experiment results in a state-of-the-art lithography environment show that our method generates high fidelity wafer images, and is 100× faster than previously reported inverse lithography method.
Frank Liu 0001, Xiaokang Shi
DATE1
2011 2011 TAU power grid simulation contest: Benchmark suite and results
abstract
Benchmark suite is an immensely useful tool in performing research since it allows for rapid and clear comparison between different approaches to solving CAD problems. Technology scaling with decrease in supply voltage, increase in power density and frequency will continue to impose strong challenges in designing of robust power delivery networks. An accurate analysis of power delivery networks has become an absolute necessity. A critical issue in power grid analysis is the large size of the power grid network. At the 45-nm technology node, the typical size of the power grid network is in the range of hundreds of million nodes. In this paper, we review the TAU 2011 Power Grid Simulation Contest. This contest was held to seek new efficient methods for solving very large power grid networks. Accuracy, run-time and memory were used as metrics to evaluate the solutions and consequently, prizes were awarded to the top three teams. The benchmarks in [1] are expanded to include larger networks that were created from real industry designs. These are made public along with the score from various teams that participated in the contest. These new benchmarks would aid in furthering academic research to address the increasing demands in the analysis of very large power grid networks.
Zhuo Li 0001, Raju Balasubramanian, Frank Liu 0001, Sani R. Nassif
ICCAD3
2011 Pure nodal analysis for efficient on-chip interconnect model order reduction
abstract
This paper described a model-order reduction (MOR) method based on a novel pure-nodal analysis formulation (PNA) which permits the use of symmetric, positive-definite Cholesky solvers for all circuit topologies. Moreover, frequently occurring special cases, e.g., inductor-resistor tree structures result in particular types of matrices that are solved by an even faster linear time algorithm. The model order reduction algorithms also uses symmetric-Lanczos iteration and non- standard inner-products for generating the Krylov subspace basis. Its efficiency is supported by a wide range of industrial examples.
Frank Liu 0001, Peter Feldmann
ISCAS1
2011 Virtual Probe: A Statistical Framework for Low-Cost Silicon Characterization of Nanoscale Integrated Circuits
abstract
In this paper, we propose a new technique, referred to as virtual probe (VP), to efficiently measure, characterize, and monitor spatially-correlated inter-die and/or intra-die variations in nanoscale manufacturing process. VP exploits recent breakthroughs in compressed sensing to accurately predict spatial variations from an exceptionally small set of measurement data, thereby reducing the cost of silicon characterization. By exploring the underlying sparse pattern in spatial frequency domain, VP achieves substantially lower sampling frequency than the well-known Nyquist rate. In addition, VP is formulated as a linear programming problem and, therefore, can be solved both robustly and efficiently. Our industrial measurement data demonstrate the superior accuracy of VP over several traditional methods, including 2-D interpolation, Kriging prediction, and k-LSE estimation.
Wangyang Zhang, Xin Li 0001, Frank Liu 0001, Emrah Acar, Rob A. Rutenbar, R. D. (Shawn) Blanton
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2011 Statistical Modeling and Simulation of Threshold Variation Under Random Dopant Fluctuations and Line-Edge Roughness
abstract
The threshold voltage (Vth) of a nanoscale transistor is severely affected by random dopant fluctuations and line-edge roughness. The analysis of these effects usually requires atomistic simulations which are too expensive in computation for statistical design. In this work, we develop an efficient SPICE simulation method and statistical variation model that accurately predict threshold variation as a function of dopant fluctuations and gate length change caused by lithography and the etching process. By understanding the physical principles of atomistic simulations, we: 1) identify the appropriate method to divide a nonuniform gate into slices in order to map those fluctuations into the device model; 2) extract the variation ofVthfrom the strong-inversion region instead of the leakage current, benefiting from the linearity of the saturation current with respect toVth; 3) propose a compact model ofVthvariation that is scalable with gate size and the amount of dopant and gate length fluctuations; and 4) investigate the interaction with non-rectangular gate (NRG) and reverse narrow width effect (RNWE). The proposed SPICE simulation method is validated with atomistic simulation results. Given the post-lithography gate geometry, this approach correctly models the variation of device output current in all operating regions. Based on the new results, we further project the amount ofVthvariation at advanced technology nodes, helping shed light on the challenges of future robust circuit design.
Yun Ye 0001, Frank Liu 0001, Min Chen 0024, Sani R. Nassif, Yu Cao 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Physical design techniques for optimizing RTA-induced variations
abstract
At 65nm and below, Rapid Thermal Annealing (RTA) makes a significant contribution to manufacturing process variations, degrading the parametric yield. RTA-induced variability strongly depends on circuit layout patterns, particularly the distribution of the density of the Shallow Trench Isolation (STI) regions. In this work, we investigate a two-step approach to reduce the impact of RTA-induced variations. We first solve a floorplanning problem that aims to reduce the RTA variations by evening out the STI density distribution. Next, we insert dummy polysilicon fills to further improve the uniformity of the STI density. Experimental results show that our floorplanner can reduce the global RTA variations by 39% and the local variations by 29% on average with low overhead compared to a traditional floorplanner, and the proposed dummy fill algorithm can further reduce the RTA variations to negligible amounts. Moreover, when inserting dummy fills, for the layouts obtained by our floorplanner, on average, 24% fewer dummy polysilicon fills are inserted, as compared to the results from a traditional floorplanner.
Yaoguang Wei, Jiang Hu 0001, Frank Liu 0001, Sachin S. Sapatnekar
ASP-DAC3
2010 Multi-Wafer Virtual Probe: Minimum-cost variation characterization by exploring wafer-to-wafer correlation
abstract
In this paper, we propose a new technique, referred to as Multi-Wafer Virtual Probe (MVP) to efficiently model wafer-level spatial variations for nanoscale integrated circuits. Towards this goal, a novel Bayesian inference is derived to extract a shared model template to explore the wafer-to-wafer correlation information within the same lot. In addition, a robust regression algorithm is proposed to automatically detect and remove outliers (i.e., abnormal measurement data with large error) so that they do not bias the modeling results. The proposed MVP method is extensively tested for silicon measurement data collected from 200 wafers at an advanced technology node. Our experimental results demonstrate that MVP offers superior accuracy over other traditional approaches such as VP and EM, if a limited number of measurement data are available.
Wangyang Zhang, Xin Li 0001, Emrah Acar, Frank Liu 0001, Rob A. Rutenbar
ICCAD4
2010 Modeling and Analysis of the Nonrectangular Gate Effect for Postlithography Circuit Simulation
abstract
For nanoscale CMOS devices, gate roughness has severe impact on the deviceI-Vcharacteristics, particularly in the subthreshold region. In particular, the nonrectangular gate (NRG) geometries are caused by subwavelength lithography and have relatively low spatial frequency. In this paper, we present an analytical approach to model NRG effects onI-Vcharacteristics. To predict the change ofI-Vcharacteristics due to the NRG effect, the proposed model converts the postlithography gate profile into an equivalent gate length (Le) , which is a function of the gate bias voltage but independent of the drain bias voltage. We demonstrate the accuracy of this approach by comparing it to TCAD simulation results for 65-nm technology. The newLemodel is readily integrated into standard transistor models in traditional circuit simulation tools, such as SPICE, for both dc and transient analyses. We further develop a generic procedure to systematically extract theLevalue from the postlithography gate profile. The interaction with the narrow-width effect is also efficiently incorporated into the proposed algorithm. TCAD verification demonstrates that the proposedLemodel is simple for implementation, scalable with both transistor geometries and bias conditions, and also continuous across all the operation regions.
Ritu Singhal, Asha Balijepalli, Anupama R. Subramaniam, Chi-Chao Wang, Frank Liu 0001, Sani R. Nassif, Yu Cao 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2010 The Impact of NBTI Effect on Combinational Circuit: Modeling, Simulation, and Analysis
abstract
Negative-bias-temperature instability (NBTI) has become the primary limiting factor of circuit life time. In this paper, we develop a hierarchical framework for analyzing the impact of NBTI on the performance of logic circuits under various operation conditions, such as the supply voltage, temperature, and node switching activity. Given a circuit topology and input switching activity, we propose an efficient method to predict the degradation of circuit speed over a long period of time. The effectiveness of our method is comprehensively demonstrated with the International Symposium on Circuits and Systems (ISCAS) benchmarks and a 65-nm industrial design. Furthermore, we extract the following key design insights for reliable circuit design under NBTI effect, including: 1) During dynamic operation, NBTI-induced degradation is relatively insensitive to supply voltage, but strongly dependent on temperature; 2) There is an optimum supply voltage that leads to the minimum of circuit performance degradation; circuit degradation rate actually goes up if supply voltage is lower than the optimum value; 3) Circuit performance degradation due to NBTI is highly sensitive to input vectors. The difference in delay degradation is up to 5× for various static and dynamic operations. Finally, we examine the interaction between NBTI effect, and process and design uncertainty in realistic conditions.
Wenping Wang 0004, Shengqi Yang, Sarvesh Bhardwaj, Sarma B. K. Vrudhula, Frank Liu 0001, Yu Cao 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2009 Predicting variability in nanoscale lithography processes
abstract
As lithography process nodes shrink to sub-wavelength levels generating acceptable layout patterns becomes a challenging problem. Traditionally, complex convolution based lithography simulations are used to estimate areas of high variability. These methods are slow and infeasible for large scale full chip analysis. This work proposes a solution to this problem by using machine learning techniques to identify layout areas that are more prone to variability. A novel target layout representation is proposed, and the latest support vector machine (SVM) algorithms are used to detect variability within standard cells and between cells in a simulated full chip layout.
Dragoljub Gagi Drmanac, Frank Liu 0001, Li-C. Wang
DAC2
2009 Variability analysis under layout pattern-dependent rapid-thermal annealing process
abstract
Rapid-Thermal Annealing (RTA) with radiation heating is recently adopted in nanoscale CMOS fabrication in order to achieve ultra-shallow junction with maximum dopant activation rate. However, recent results report the systematic shift of threshold voltage (Vth) and increased Vth variation due to RTA process [1--2]. The exact amount of variations depends on layout pattern density, RTA heating temperature (T) and effective annealing time. In this work, we develop joint thermal/TCAD simulation and compact modeling tools to analyze performance variability under various layout pattern densities and RTA conditions. With the new simulation capability, we recognize two major variation mechanisms under RTA: the change of effective channel length (Leff) induced by lateral dopant diffusion, and the fluctuation of equivalent oxide thickness (EOT) due to incomplete dopant activation. We perform device simulations to quantify transistor performance shift due to Leff and EOT variations. Moreover, we propose a suite of compact models that bridge the underlying RTA process with device parameter change for efficient design optimization. The new tools are validated with published silicon data at 45nm and 65nm nodes. They will facilitate physical designers to predict and mitigate circuit performance variability due to the layout-dependent RTA process.
Yun Ye 0001, Frank Liu 0001, Min Chen 0024, Yu Cao 0001
DAC2
2009 Modeling of layout-dependent stress effect in CMOS design
abstract
Strain technology has been successfully integrated into CMOS fabrication to improve carrier transport properties since 90 nm node. Due to the non-uniform stress distribution in the channel, the enhancement in carrier mobility, velocity, and threshold voltage shift strongly depend on circuit layout, leading to systematic performance variations among transistors. A compact stress model that physically captures this behavior is essential to bridge the process technology with design optimization. In this paper, starting from the first principle, a new layout-dependent stress model is proposed as a function of layout, temperature, and other device parameters. Furthermore, a method of layout decomposition is developed to partition the layout into a set of simple patterns for efficient model extraction. These solutions significantly reduce the complexity in stress modeling and simulation. They are comprehensively validated by TCAD simulation and published Si-data, including the state-of-the-art strain technologies and the STI stress effect. By embedding them into circuit analysis, the interaction between layout and circuit performance is well benchmarked at 45 nm node.
Chi-Chao Wang, Frank Liu 0001, Min Chen 0024, Yu Cao 0001
ICCAD3
2009 Finite-Point-Based Transistor Model: A New Approach to Fast Circuit Simulation
abstract
In this paper, a new approach of transistor modeling is developed for fast statistical circuit simulation in the presence of variations. For both the I-V and C-V characteristics of a transistor, finite data points are identified based on their physical meanings and their importance in circuit operation. The impact of process and design variations is embedded into these key points using analytical expressions. During the simulation, the entire I -V and C -V curves are interpolated from these points with simple polynomial formulas. This novel approach significantly enhances the simulation speed with sufficient accuracy. The model is implemented in Verilog-A to support generic circuit simulators. The accuracy and convergence of the proposed model are comprehensively evaluated through a set of benchmark circuits, including nand, a pass-gate, latches, AOI, ring oscillators, and an adder. Compared to SPICE simulations with the BSIM models, the simulation time can be reduced by 7 times in transient analysis and more than 9 times in Monte-Carlo simulations.
Min Chen 0024, Frank Liu 0001, Yu Cao 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2008 Statistical modeling and simulation of threshold variation under dopant fluctuations and line-edge roughness
abstract
The threshold voltage (Vth) of a nanoscale transistor is severely affected by random dopant fluctuations and line-edge roughness. The analysis of these effects usually requires atomistic simulations that are too expensive in computation for statistical circuit design. In this work, we develop an efficient SPICE simulation method and statistical transistor model that accurately predict threshold variation as a function of dopant fluctuations and gate length change caused by sub-wavelength lithography and the gate etching process. By understanding the physical principles of atomistic simulations, we (a) identify the appropriate method to divide a non-uniform gate into slices in order to map those fluctuations into the device model; (b) extract the variation of V th from the strong-inversion region instead of the leakage current, benefiting from the linearity of the saturation current with respect to Vth; and (c) propose a compact model of Vth variation that is scalable with gate size and the amount of dopant and gate length fluctuations. The proposed SPICE simulation method is fully validated against atomistic simulation results. Given the post-lithography gate geometry, this approach correctly models the variation of device output current in all operating regions. Based on the new results, we further project the amount of V th variation at advanced technology nodes, helping shed light on the challenges of future robust circuit design.
Yun Ye 0001, Frank Liu 0001, Sani R. Nassif, Yu Cao 0001
DAC2
2007 A New Methodology for Interconnect Parasitics Extraction Considering Photo-Lithography Effects
abstract
Even with the wide adaptation of resolution enhancement techniques in sub-wavelength lithography, the geometry of the fabricated interconnect is still quite different from the drawn one. Existing layout parasitic extraction (LPE) tools assume perfect geometry, thus introducing significant error in the extracted parasitic models, which in turn cases significant error in timing verification and signal integrity analysis. Our simulation shows that the RC parasitics extracted from perfect GDS-II geometry can be as much as 20% different from those extracted from the post litho/etching simulation geometry. This paper presents a new LPE methodology and related fast algorithms for interconnect parasitic extraction under photolithographic effects. Our methodology is compatible with the existing design flow. Experimental results show that the proposed methods are accurate and efficient.
Nancy Y. Zhou, Zhuo Li 0001, Weiping Shi, Frank Liu 0001
ASP-DAC5
2007 A General Framework for Spatial Correlation Modeling in VLSI Design
abstract
Many characteristics of VLSI designs, such as process variations, demonstrate strong spatial correlations. Accurately modeling of these correlated behaviors is crucial for many timing and power analyses to be valid. This paper proposes a new spatial model with a long-range trend component, a smooth correlation component, as well as a truly random component. The efficient method to construct such a spatial model is based on the Generalized Least Square fitting and the structured correlation functions, which are actually the generalization of the popular Pelgrom mismatch models. Experimental results on industrial benchmarks show that the method is not only highly effective for variability modeling, but can also be used for other spatially distributed characteristics such as IR drops and on-chip temperature distributions.
Frank Liu 0001
DAC1
2007 Modeling and Analysis of Non-Rectangular Gate for Post-Lithography Circuit Simulation
abstract
In the nano regime it has become increasingly important to consider the impact of non-rectangular gate (NRG) shape caused due to sub-wavelength lithography. NRG dramatically increases the leakage current and requires geometry dependent transistor models for post-litho circuit simulation. In this paper, we propose a coherent modeling approach for non-rectangular gates based on equivalent gate length (Le). A gate-voltage dependent model of Le is developed which is scalable with design conditions, continuous across weak and strong inversion regions, accurate for both leakage and saturation current, and compatible with standard circuit analysis tools. We systematically verify this approach with 65nm TCAD simulations. A generic CAD algorithm is further proposed to predict the value of Le under various non-rectangular geometries. The interaction with the narrow-width effect is efficiently convolved in this method. Depending on the gate geometry, the leakage current can vary more than 15X at 65nm technology node. Our analytical method well captures this effect. Finally, we extrapolate the impact of NRG effect on future technology generations. The proposed model can be easily extracted from TCAD tools or direct silicon data. It bridges the gap between lithography, simulation, and circuit analysis for measuring transistor performance under increasingly severe NRG effect.
Ritu Singhal, Asha Balijepalli, Anupama R. Subramaniam, Frank Liu 0001, Sani R. Nassif, Yu Cao 0001
DAC4
2007 The Impact of NBTI on the Performance of Combinational and Sequential Circuits
abstract
Negative-bias-temperature-instability (NBTI) has become the primary limiting factor of circuit lifetime. In this work, we develop a general framework for analyzing the impact of NBTI on the performance of a circuit, based on various circuit parameters such as the supply voltage, temperature, and node switching activity of the signals etc. We propose an efficient method to predict the degradation of circuit performance based on circuit topology and the switching activity of the signals over long periods of time. We demonstrate our results on ISCAS benchmarks and a 65nm industrial design. The framework is used to provide key design insights for designing reliable circuits. The key design insights that we obtain are: (1) degradation due to NBTI is most sensitive on the input patterns and the duty cycle; the difference in the delay degradation can be up to 5X for various static and dynamic conditions, (2) during dynamic operation, NBTI-induced degradation is relatively insensitive to supply voltage, but strongly dependent on temperature; (3) NBTI has marginal impact on the clock signal.
Wenping Wang 0004, Shengqi Yang, Sarvesh Bhardwaj, Rakesh Vattikonda, Sarma B. K. Vrudhula, Frank Liu 0001, Yu Cao 0001
DAC6
2007 Fast statistical circuit analysis with finite-point based transistor model
Min Chen 0024, Frank Liu 0001, Yu Cao 0001
DATE3
2007 Efficient computation of current flow in signal wires for reliability analysis
abstract
Electromigration(EM) andself-heatingare critical reliability concerns for metal wires in high performance designs. EM reliability rules for a VLSI technology are typically expressed in terms of average, root-mean-square and peak current limits for each metal layer in the technology. To ensure EM reliability of a design, current flowing through each wire segment in the design should not violate the EM reliability rules. In this work, we present closed-from analytical models for efficient computation of average, root-mean-square and peak currents through any element in an arbitrary RC tree. The proposed models are validated against SPICE simulations for several RC nets extracted from an industrial ASIC design. The results show that the models exhibit very good accuracy with a mean error of only 3.1% in root-mean-square and 0.2% in average current estimation.
Kanak Agarwal 0001, Frank Liu 0001
ICCAD2
2007 An efficient method for statistical circuit simulation
abstract
The dynamic behavior of a VLSI circuit can be described by a system of differential-algebraic equations. When some circuit elements are affected by process variations, the dynamic behavior of the circuit will deviate from its nominal trajectory. Monte-Carlo-type random sampling methods are widely used to estimate the trajectory deviation. However they can be quite time-consuming when the dimension of the parameter space is large. This paper offers an alternative solution by casting the problem into the theoretic frame work of non-linear non-Gaussian filtering. To estimate the mean and variance of the time-dependent circuit trajectory, we develop a method based on unscented transformation, which is an efficient Bayesian analysis sampling technique. Theoretically the method has linear runtime complexity. Experimental results show that compared to traditional Monte-Carlo methods, the new method can achieve over 10times speedup with less than 2% error.
Frank Liu 0001
ICCAD1
2007 Integrated Placement and Skew Optimization for Rotary Clocking
abstract
The clock distribution network is a key component of any synchronous VLSI design. High power dissipation and pressure volume temperature-induced variations in clock skew have started playing an increasingly important role in limiting the performance of the clock network. Rotary clocking is a novel technique which employs unterminated rings formed by differential transmission lines to save power and reduce skew variability. Despite its appealing advantages, rotary clocking requires flip-flop locations to match predesigned clock skew on rotary clock rings. This requirement poses a difficult chicken-and-egg problem which prevents its wide application. In this paper, we propose an integrated placement and skew scheduling methodology to break this hurdle, making rotary clocking compatible with practical design flows. A network flow based flip-flop assignment algorithm and a cost-driven skew optimization algorithm are developed. We also present an integer linear programming formulation that minimizes maximum capacitance loaded at any of the rotary rings, thereby maximizing the operating frequency. Experimental results on benchmark circuits show that our method can reduce the tapping cost (measured as the total length of the wire segments connecting the rotary rings to the clock sinks) for rotary clocking by 33%-53%
Ganesh Venkataraman, Jiang Hu 0001, Frank Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2007 Fast Variational Interconnect Delay and Slew Computation Using Quadratic Models
abstract
Interconnects constitute a dominant source of circuit delay for modern chip designs. The variations of critical dimensions in modern VLSI technologies lead to variability in interconnect performance that must be fully accounted for in timing verification. However, handling a multitude of inter-die/intra-die variations and assessing their impacts on circuit performance can dramatically complicate the timing analysis. In this paper, a practical interconnect delay and slew analysis technique is presented to facilitate efficient evaluation of wire performance variability. By harnessing a collection of computationally efficient procedures and closed-form formulas, process variations are directly mapped into the variability of the output delay and slew. An efficient method based on sensitivity analysis is implemented to calculate driving point models under variations for gate-level timing analysis. The proposed adjoint technique not only provides statistical performance variations of the interconnect network under analysis, but also produces delay and slew expressions parameterized in the underlying process variations in a quadratic parametric form. As such, it can be harnessed to enable statistical timing analysis while considering important statistical correlations. Our experimental results have indicated that the presented analysis is accurate regardless of location of sink nodes and it is also robust over a wide range of process variations.
Xiaoji Ye, Frank Liu 0001, Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2006 A practical method to estimate interconnect responses to variabilities
abstract
Variabilities in metal interconnect structures can affect circuit timing performance or even cause function failure in VLSI designs. This paper proposes a method to estimate the difference between the nominal and perturbed circuit waveforms by calculating the moments in frequency-domain via efficient iterative method. The algorithm can be used to accurately reproduce the differential waveforms, or to provide efficient early estimates on the timing impact of the variabilities for RC networks
Frank Liu 0001
DATE1
2006 Integrated placement and skew optimization for rotary clocking
abstract
The clock distribution network is a key component on any synchronous VLSI design. As technology moves into the nanometer era, innovative clocking techniques are required to solve the power dissipation and variability issues. Rotary clocking is a novel technique which employs unterminated rings formed by differential transmission lines to save power and reduce skew variability. Despite its appealing advantages, rotary clocking requires latch locations to match pre-designed clock skew on rotary clock rings. This requirement is a difficult chicken-and-egg problem which prevents its wide application. In this work, we proposed an integrated placement and skew scheduling methodology to break this hurdle, making rotary clocking compatible with practical design flows. A network flow based latch assignment algorithm and a cost-driven skew optimization algorithm are developed. Experiments show that our method can generate chip placements which satisfy the unique requirements of rotary clocks, without sacrificing design quality. By enabling concurrent clock network and placement design, our method can also be applied in other clocking methodologies as well
Ganesh Venkataraman, Jiang Hu 0001, Frank Liu 0001, Cliff C. N. Sze
DATE3
2006 Practical variation-aware interconnect delay and slew analysis for statistical timing verification
abstract
Interconnects constitute a dominant source of circuit delay for modern chip designs. The variations of critical dimensions in modern VLSI technologies lead to variability in interconnect performance that must be fully accounted for in timing verification. However, handling a multitude of inter-die/intra-die variations and assessing their impacts on circuit performance can dramatically complicate the timing analysis. In this paper, a practical interconnect delay and slew analysis technique is presented to facilitate efficient evaluation of wire performance variability. By harnessing a collection of computationally efficient procedures and closed-form formulas, process and input signal variations are directly mapped into the variability of the output delay and slew. Since our approach produces delay and slew expressions parameterized in the underlying process variations, it can be harnessed to enable statistical timing analysis while considering important statistical correlations. Our experimental results have indicated that the presented analysis is accurate regardless of location of sink nodes and it is also robust over a wide range of process variations.
Xiaoji Ye, Peng Li 0001, Frank Liu 0001
ICCAD3
2005 A noise-driven effective capacitance method with fast embedded noise rule calculation for functional noise analysis
abstract
We present a noise-driven effective capacitance method for estimating the combined propagation noise and crosstalk noise. Gate propagation noise rules are efficiently calculated inside the Ceff procedure to determine a linear Thevenin model of the victim driver. A voltage-dependent current source model [2, 6] of the driver, along with a load capacitor is analyzed to generate the gate output waveform, from which noise rules are directly extracted. This method removes potential errors introduced in traditional look-up table or fitted-equation based noise rules. The linear driver Thevenin model can then be employed to analyze the propagation noise, while the same Thevenin resistance can be used to analyze the crosstalk noise. The combined coupling and propagation noise can then be estimated using superposition. In this work, we extend the popular timing-driven effective capacitance method into the noise domain. Similar to the effective capacitance method in timing analysis, this technique can successfully separate the nonlinear driver analysis from the linear interconnect analysis. In addition, the linear driver model can significantly ease the task of finding the worst-case peak alignment among all the victim and aggressor noise sources. Experimental results on both RC and RLC nets from industry designs show both accuracy and efficiency compared to SPICE results.
Haihua Su, David Widiger, Chandramouli V. Kashyap, Frank Liu 0001, Byron Krauter
DAC4
2005 Modeling Interconnect Variability Using Efficient Parametric Model Order Reduction
abstract
Assessing IC manufacturing process fluctuations and their impacts on IC interconnect performance has become unavoidable for modern DSM designs. However, the construction of parametric interconnect models is often hampered by the rapid increase in computational cost and model complexity. In this paper we present an efficient yet accurate parametric model order reduction algorithm for addressing the variability of IC interconnect performance. The efficiency of the approach lies in a novel combination of low-rank matrix approximation and multi-parameter moment matching. The complexity of the proposed parametric model order reduction is as low as that of a standard Krylov subspace method when applied to a nominal system. Under the projection-based framework, our algorithm also preserves the passivity of the resulting parametric models.
Peng Li 0001, Frank Liu 0001, Xin Li 0001, Lawrence T. Pileggi, Sani R. Nassif
DATE2
2004 Variational delay metrics for interconnect timing analysis
abstract
In this paper we develop an approach to model interconnect delay under process variability for timing analysis and physical design optimization. The technique allows for closed-form computation of interconnect delay probability density functions (PDFs) given variations in relevant process parameters such as linewidth, metal thickness, and dielectric thickness. We express the resistance and capacitance of a line as a linear function of random variables and then use these to compute circuit moments. Finally, these variability-aware moments are used in known closed-form delay metrics to compute interconnect delay PDFs. We compare the approach to SPICE based Monte Carlo simulations and report an error in mean and standard deviation of delay of 1% and 4% on average, respectively.
Kanak Agarwal 0001, Dennis Sylvester, David T. Blaauw, Frank Liu 0001, Sani R. Nassif, Sarma B. K. Vrudhula
DAC4
2004 Sparse and efficient reduced order modeling of linear subcircuits with large number of terminals
abstract
In the process of designing state-of-the art VLSI circuit we often encounter large but highly structured linear subcircuits with large number of terminals. Classical examples are power supply networks, clock distribution networks, large data buses, etc. Various applications would benefit from efficient high level models of such networks. Unfortunately the existing model-order-reduction algorithms are not adapted to handle more than a few tens of terminals. This talk introduces RecMOR, an algorithm for the computation of reduced order models of structured linear circuits with numerous I/O ports. The algorithm exploits certain regularities of the subcircuit response that are typical in numerous applications of interest. When these regularities are present, the normally dense matrix-transfer function of the subcircuit contains sub-blocks that in some sense are significantly low rank and can be compactly modeled by the recently introduced SVDMOR algorithm. The new RecMOR algorithm decomposes the large matrix-transfer function recursively, and applies SVDMOR compression adaptively to the sub-blocks of the transfer function. The result is a reduced order model that is sparse, efficient, and directly usable as an efficient substitute of the subcircuit in circuit simulations. The method is illustrated on several circuit examples.
Peter Feldmann, Frank Liu 0001
ICCAD2
2004 Closed-form delay and slew metrics made easy
abstract
For optimizations like physical synthesis and static timing analysis, efficient interconnect delay and slew computation is critical. Since one cannot often afford to run asymptotic waveform evaluation (Pillage and Rohrer, 1990), constant time solutions are required. This work presents the first complete solution to closed-form formulas for both delay and also for slew. Our metrics are derived from matching circuit moments to the lognormal distribution. From a single table, one can easily implement the metrics for delay and slew for both step and ramp inputs. Experiments validate the effectiveness of the metrics for nets from a real industrial design.
Charles J. Alpert, Frank Liu 0001, Chandramouli V. Kashyap, Anirudh Devgan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2004 Closed-form expressions for extending step delay and slew metrics to ramp inputs for RC trees
abstract
Recent years have seen significant research in finding closed form expressions for the delay of an RC circuit that improves upon the Elmore delay model. However, several of these formulae assume a step excitation, leaving it to the reader to find a suitable extension to ramp-we always refer to saturated ramps in this paper-inputs. The few works that do consider ramp inputs do not present a closed-form formula that works for a wide range of possible input slews. We propose the PERI (probability distribution function extension for ramp inputs) technique, that extends delay metrics for step inputs to the more general and realistic non-step (such as a ramp) inputs. Although there has been little work done in finding good slew (which is also referred as signal transition time) metrics, we also show how one can extend a slew metric for step inputs to the non-step case. We validate the efficacy of our approach through experimental results from several hundred RC dominated nets extracted from an industry application specific integrated circuit design.
Chandramouli V. Kashyap, Charles J. Alpert, Frank Liu 0001, Anirudh Devgan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 A delay metric for RC circuits based on the Weibull distribution
abstract
Physical synthesis optimizations require fast and accurate analysis of RC networks. Elmore first proposed matching circuit moments to a probability density function (PDF), which led to widespread adoption of his simple and fast metric. The more recently proposed PRIMO and H-gamma metrics match the circuit moments to the PDF of a Gamma statistical distribution. We instead propose to match the circuit moments to a Weibull distribution and derive a new delay metric called Weibull-based delay (WED). The primary advantages of WED over PRIMO and H-gamma are its efficiency and ease of implementation. Experiments show that WED is robust and has satisfactory accuracy at both near- and far-end nodes.
Frank Liu 0001, Chandramouli V. Kashyap, Charles J. Alpert
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 Delay and slew metrics using the lognormal distribution
abstract
For optimizations like physical synthesis and static timing analysis, efficient interconnect delay and slew computation is critical. Since one cannot often afford to run AWE[12], constant time solutions are required. This work presents the first complete solution to closed form formulae for both delay and slew. Our metrics are derived from matching circuit moments to the lognormal distribution. From a single table, one can easily implement the metrics for delay and slew for both step and ramp inputs. Experiments validate the effectiveness of the metrics for nets from a real industrial design.
Charles J. Alpert, Frank Liu 0001, Chandramouli V. Kashyap, Anirudh Devgan
DAC2
2003 A Heuristic to Determine Low Leakage Sleep State Vectors for CMOS Combinational Circuits
Rahul M. Rao, Frank Liu 0001, Jeffrey L. Burns, Richard B. Brown
ICCAD2
2003 Full chip leakage estimation considering power supply and temperature variations
abstract
Leakage power is emerging as a key design challenge in current and future CMOS designs. Since leakage is critically dependent on operating temperature and power supply, we present a full chip leakage estimation technique which accurately accounts for power supply and temperature variations. State of the art techniques are used to compute the thermal and power supply profile of the entire chip. Closed-form models are presented which relate leakage to temperature and VDD variations. These models coupled with the thermal and VDD profile are used to generate an accurate full chip leakage estimation technique considering environmental variations. The results of this approach are demonstrated on large-scale industrial designs.
Haihua Su, Frank Liu 0001, Anirudh Devgan, Emrah Acar, Sani R. Nassif
ISLPED2
2003 Closed form expressions for extending step delay and slew metrics to ramp inputs
abstract
Recent years have seen significant research in finding closed form expressions for the delay of an RC circuit that improves upon the Elmore delay model. However, several of these formulae assume a step excitation, leaving it to the reader to find a suitable extension to ramp inputs (we always assume a saturated ramp in this paper). The few works that do consider ramp inputs do not present a closed-form formula that works for a wide range of possible input slews. We propose the PERI (Probability distribution function Extension for Ramp Inputs) technique, that extends delay metrics for step inputs to the more general and realistic non-step inputs. Although there has been little work done in finding good slew - which is also referred as signal transition time - metrics, we also show how one can extend a slew metric for step inputs to the non-step case. We validate the efficacy of our approach through experimental results from several hundred RC dominated nets extracted from an industry ASIC design.
Chandramouli V. Kashyap, Charles J. Alpert, Frank Liu 0001, Anirudh Devgan
ISPD3
2002 A delay metric for RC circuits based on the Weibull distribution
abstract
Physical design optimizations such as placement, interconnect synthesis, floorplanning, and routing require fast and accurate analysis of RC networks. Because of its simple close form and fast evaluation, the Elmore delay metric has been widely adopted. The recently proposed delay metrics PRIMO and H-gamma match the first three circuit moments to the probability density function of a gamma statistical distribution. Although these methods demonstrate impressive accuracy compared to other delay metrics, their implementations tend to be challenging. As an alternative to matching to the gamma distribution, we propose to match the first two circuit moments to a Weibull distribution. The result is a new delay metric called Weibull based Delay (WED). The primary advantages of WED over PRIMO and H-gamma are its efficiency and ease of implementation. Experiments show that WED is robust and has satisfactory accuracies at both near- and far-end nodes.
Frank Liu 0001, Chandramouli V. Kashyap, Charles J. Alpert
ICCAD1