Nancy Y. Zhou

dblp:45/11470 · also Nancy Ying Zhou · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Electronic design automation · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
physical design
0.432012
$O(mn)$ Time Algorithm for Optimal Buffer Insertion of Nets With $m$ Sinks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Guiding a physical design closure system to produce easier-to-route designs with more predictable timing · DAC 2012
Wire Sizing for Non-Tree Topology · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Electronic design automation › physical design
interconnect optimization
0.222012
$O(mn)$ Time Algorithm for Optimal Buffer Insertion of Nets With $m$ Sinks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Wire Sizing for Non-Tree Topology · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Electronic design automation › physical design
buffer insertion
0.112012
$O(mn)$ Time Algorithm for Optimal Buffer Insertion of Nets With $m$ Sinks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Electronic design automation › design optimization
buffer minimization
0.112012
$O(mn)$ Time Algorithm for Optimal Buffer Insertion of Nets With $m$ Sinks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Electronic design automation › design methodology
design closure
0.112012
Guiding a physical design closure system to produce easier-to-route designs with more predictable timing · DAC 2012
Electronic design automation › physical design › placement
routability-driven placement
0.112012
Guiding a physical design closure system to produce easier-to-route designs with more predictable timing · DAC 2012
Electronic design automation
boundary element method
0.112007
Fast Capacitance Extraction in Multilayer, Conformal and Embedded Dielectric using Hybrid Boundary Element Method · DAC 2007
Electronic design automation › physical design › parasitic extraction
capacitance extraction
0.112007
Fast Capacitance Extraction in Multilayer, Conformal and Embedded Dielectric using Hybrid Boundary Element Method · DAC 2007
Electronic design automation › physical design
clock network synthesis
0.112007
Wire Sizing for Non-Tree Topology · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Electronic design automation › physical design › parasitic extraction
interconnect parasitic extraction
0.112007
Fast Capacitance Extraction in Multilayer, Conformal and Embedded Dielectric using Hybrid Boundary Element Method · DAC 2007
Electronic design automation › physical design › interconnect optimization
wire sizing
0.112007
Wire Sizing for Non-Tree Topology · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Electronic design automation › physical design
timing optimization
0.012012
Guiding a physical design closure system to produce easier-to-route designs with more predictable timing · DAC 2012

Methods — techniques the papers use, named apart from their topics

steiner wire model · 0.1physical synthesis · 0.1linked list data structure · 0.1dynamic programming · 0.1multilayer green's function · 0.1equivalent charge method · 0.1elmore delay model · 0.1RC network decomposition · 0.1
YearPublicationVenuePosition
2013 Clock power minimization using structured latch templates and decision tree induction
abstract
This work proposes a novel latch placement methodology by computing optimized placement templates with significantly lower local clock tree capacitance at a one-time cost per standard cell library. By directly minimizing local clock tree capacitance, overall chip power is reduced. The proposed methodology first generates optimized placement solutions for a wide range of input configurations. Then, a redundancy removal approach using set-theoretic annotation is proposed demonstrating it is possible to remove over 99% of the templates with no information loss. Finally, a decision tree induction algorithm with novel impurity metric enables extremely fast template selection during the clock optimization stage of a modern physical design flow. The proposed approach reduces the local clock tree capacitance by 20-30% on average roughly equating to between a 1 and 4 watt reduction in total dynamic power on a 100-watt 22-nm microprocessor. Additionally, because of a priori generation, template selection during physical design is extremely fast.
Samuel I. Ward, Natarajan Viswanathan, Nancy Y. Zhou, Cliff C. N. Sze, Zhuo Li 0001, Charles J. Alpert, David Z. Pan
ICCAD3
2012 Guiding a physical design closure system to produce easier-to-route designs with more predictable timing
abstract
Physical synthesis has emerged as one of the most important tools in design closure, which starts with the logic synthesis step and generates a new optimized netlist and its layout for the final signoff process. As stated in [1], "it is a wrapper around traditional place and route, whereby synthesis-based optimization are interwoven with placement and routing." A traditional physical synthesis tool generally focuses on design closure with Steiner wire model. It optimizes timing/area/power with the assumption that each net can be routed with optimal Steiner tree. However, advanced design rules, more IP and hierarchical design styles for super-large billion-gate designs, serious buffering problems from interconnect scaling and metal layer stacks make routing a much more challenging problem [2]. This paper discusses a series of techniques that may relieve this problem, and guide the physical design closure system to produce not only easier to route designs, but also better timing quality. Open challenges are also overviewed at the end.
Zhuo Li 0001, Charles J. Alpert, Gi-Joon Nam, Cliff C. N. Sze, Natarajan Viswanathan, Nancy Y. Zhou
DAC6
2012 $O(mn)$ Time Algorithm for Optimal Buffer Insertion of Nets With $m$ Sinks
abstract
Buffer insertion is an effective technique to reduce interconnect delay. In this paper, we give a simple$O(mn)$time algorithm for optimal buffer insertion, where$m$is the number of sinks and$n$is the number of buffer positions. When$m$is small, our algorithm is a significant improvement over the recent$O(n\log^{2}n)$time algorithm by Shi and Li, and the$O(n^{2})$time algorithm of van Ginneken. For$b$buffer types, our algorithms runs in$O(b^{2}n+bmn)$time, an improvement of the recent$O(bn^{2})$algorithm by Li and Shi. The improvement is made possible by an innovative linked list that can perform addition of a wire, addition of a buffer in amortized$O(1)$time, and smart design of pointers. We then present the extension of our algorithm for the buffer cost minimization problem, which improves the previous best algorithm. On industrial test cases, the new algorithms is faster than previous best algorithms by an order of magnitude.
Zhuo Li 0001, Nancy Y. Zhou, Weiping Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 Ultra-fast interconnect driven cell cloning for minimizing critical path delay
abstract
In a complete physical synthesis flow, optimization transforms, that can improve the timing on critical paths that are already well-optimized by a series of powerful transforms (timing driven placement, buffering and gate sizing) are invaluable. Finding such a transform is quite challenging, to say nothing of efficiency. This work explores innovative cloning (gate duplication) techniques to improve timing-closure in a physical synthesis environment.
Zhuo Li 0001, David A. Papa, Charles J. Alpert, Shiyan Hu 0001, Weiping Shi, Cliff C. N. Sze, Nancy Y. Zhou
ISPD7
2008 SRAM methodology for yield and power efficiency: per-element selectable supplies and memory reconfiguration schemes
abstract
We present a novel power-aware yield enhancement design methodology and reconfiguration scheme for deep submicron SRAM designs. We show that with the continued trend of raising array supply to counter process variations, it is more effective to use a per-element selectable virtual power-supply scenario as opposed to single array supply with traditional redundancy schemes. The element can be a bank, a sub-array, or an independent row/column, and the element's virtual supply value is determined based on fail bitmaps. The technique can also be used in conjunction with traditional redundancy schemes to further improve the efficiency. The supply and redundancy assignments can be obtained by relying on memory reconfiguration algorithms. For this, we propose a greedy yet accurate algorithm that runs in O(nlogn) as opposed to average case O(n2) traditional algorithms. The methodology leads to significant power savings ranging from 20% to 50% for 65nm technology. We expect the savings to increase in future technologies as leakage powers dominate. To the best of our knowledge, this is the first time such a methodology is applied to SRAM designs.
Rouwaida Kanj, Rajiv V. Joshi, Zhuo Li 0001, Jente B. Kuang, Hung C. Ngo, Nancy Y. Zhou, Weiping Shi, Sani R. Nassif
ISLPED6
2007 A New Methodology for Interconnect Parasitics Extraction Considering Photo-Lithography Effects
abstract
Even with the wide adaptation of resolution enhancement techniques in sub-wavelength lithography, the geometry of the fabricated interconnect is still quite different from the drawn one. Existing layout parasitic extraction (LPE) tools assume perfect geometry, thus introducing significant error in the extracted parasitic models, which in turn cases significant error in timing verification and signal integrity analysis. Our simulation shows that the RC parasitics extracted from perfect GDS-II geometry can be as much as 20% different from those extracted from the post litho/etching simulation geometry. This paper presents a new LPE methodology and related fast algorithms for interconnect parasitic extraction under photolithographic effects. Our methodology is compatible with the existing design flow. Experimental results show that the proposed methods are accurate and efficient.
Nancy Y. Zhou, Zhuo Li 0001, Weiping Shi, Frank Liu 0001
ASP-DAC1
2007 Fast Capacitance Extraction in Multilayer, Conformal and Embedded Dielectric using Hybrid Boundary Element Method
abstract
In modern VLSI circuits, metal conductors are separated by multiple planar, conformal or embedded dielectric media. Previous algorithms based on Boundary Element Method (BEM) are inefficient to extract interconnect capacitance due to the complex dielectric structures. In this paper, we present a new algorithm that combines multilayer Green's function with the equivalent charge method to efficiently deal with the complex dielectrics. The multilayer Green's function is efficient to model layered dielectric media, while the equivalent charge method is powerful to model non-planar complex dielectric. Our method can also model ground plane and reflective boundary wall. From experimental results, the new method is significantly faster than previous methods in realistic conditions, i.e., 70X speedup and 99% memory saving compared with FastCap and 2X speedup and 80% memory saving compared with PHiCap for complex dielectric structure with similar accuracy.
Nancy Y. Zhou, Zhuo Li 0001, Weiping Shi
DAC1
2007 Wire Sizing for Non-Tree Topology
abstract
Most existing methods for interconnect wire sizing are designed for RC trees. With the increasing popularity of the non-tree topology in clock networks and multiple link networks, wire sizing for non-tree networks becomes an important problem. In this paper, we propose the first systematic method to size the wires of general non-tree RC networks. Our method consists of three steps: 1) decompose a non-tree RC network into a tree RC network such that the Elmore delay at every sink remains unchanged; 2) size wires of the tree; and 3) merge the wires back to the original non-tree network. All three steps can be implemented in low-order polynomial time. Using this method, previous wiresizing techniques for tree topology for various objectives, such as minimizing the maximum delay, minimizing the total area or power, and reducing skew variability under process variations, can be applied to non-tree topologies. For certain types of networks, such as the tree+link network, our method gives the optimal solution, provided the tree wire sizing is optimal. Compared with the previous best wire-sizing method for non-tree circuits we can achieve 2% to 17% Elmore delay reduction with 14% to 30% total wire area reduction. Compared with unsized minimum width networks, our delay is 25% less and the skew is 34% less, under SPICE simulation. For the tree+link network, we can achieve significant delay reduction and zero skew in nominal case, while get up to 66% skew variation reduction.
Zhuo Li 0001, Nancy Y. Zhou, Weiping Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2