Hyein Lee 0001

dblp:29/6859-1 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0001-8037-2207ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Electronic design automation · 92% Energy-efficient computing · 4% Integrated circuit design · 4%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
physical design
1.762020
Heuristic Methods for Fine-Grain Exploitation of FDSOI · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
PROBE: A Placement, Routing, Back-End-of-Line Measurement Utility · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
MILP-Based Optimization of 2-D Block Masks for Timing-Aware Dummy Segment Removal in Self-Aligned Multiple Patterning Layouts · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Electronic design automation › physical design
routing
0.422018
PROBE: A Placement, Routing, Back-End-of-Line Measurement Utility · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Vertical M1 Routing-Aware Detailed Placement for Congestion and Wirelength Reduction in Sub-10nm Nodes · DAC 2017
Electronic design automation › physical design › placement
detailed placement
0.312017
Vertical M1 Routing-Aware Detailed Placement for Congestion and Wirelength Reduction in Sub-10nm Nodes · DAC 2017
Electronic design automation › physical design › lithography
multiple patterning lithography
0.312017
MILP-Based Optimization of 2-D Block Masks for Timing-Aware Dummy Segment Removal in Self-Aligned Multiple Patterning Layouts · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Electronic design automation › physical design › routing
detailed routing
0.212015
Evaluation of BEOL design rule impacts using an optimal ILP-based detailed router · DAC 2015
Electronic design automation › physical design
clock network synthesis
0.212013
Smart non-default routing for clock power reduction · DAC 2013
Energy-efficient computing › dynamic power reduction
clock power reduction
0.212013
Smart non-default routing for clock power reduction · DAC 2013
Electronic design automation › design methodology
design flow optimization
0.112020
Heuristic Methods for Fine-Grain Exploitation of FDSOI · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Integrated circuit design
low-power circuit design
0.112020
Heuristic Methods for Fine-Grain Exploitation of FDSOI · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020

Methods — techniques the papers use, named apart from their topics

integer linear programming · 0.7mixed integer linear programming · 0.6sensitivity-based heuristic optimization · 0.4placement · 0.3markov transition matrix · 0.3distributed optimization · 0.3smart non-default routing · 0.2
YearPublicationVenuePosition
2020 Heuristic Methods for Fine-Grain Exploitation of FDSOI
abstract
Fully depleted silicon on insulator (FDSOI) is attractive for its low cost and low power; the mixed-Vt and body-bias levers that it affords expand the performance-power solution space. However, in FDSOI, different-Vt (i.e., low Vt and regular Vt) devices must be isolated from each other, which makes realization of fine-grained mixed-Vt/body-biasing in layout extremely challenging. In this article, we study heuristic methods aimed at exploitation of fine-grained mixed-Vt in FDSOI implementation. We propose a novel speed domain partitioning (SDP) problem formulation that comprehends the spatial contiguity restrictions arising from flip-well structure of low Vt regions in popular 28-nm commercial FDSOI offerings. We explore a wide space of implementation flows that include an integer linear programming (ILP)-based approach, and a heuristic (sensitivity-based) optimization. Our experimental studies have been performed across multiple commercial enablements. We observe that outcomes are library- and design-dependent. For implementations using generic library options, up to 20% speed improvement with 54% low Vt region is seen for one out of four testcases studied. For implementations using “rich” library options, up to 7% speed improvement with 26% low Vt region is achieved. We provide a discussion that summarizes root-cause, intrinsic difficulties of fine-grained exploitation of mixed-Vt. Finally, we suggest a “decision tree” to help assess a design's amenability to fine-grained mixed-Vt implementation, and to help guide design flow selection for better design QoR.
Hamed Fatemi, Andrew B. Kahng, Hyein Lee 0001, José Pineda de Gyvez
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 Enhancing sensitivity-based power reduction for an industry IC design context
Hamed Fatemi, Andrew B. Kahng, Hyein Lee 0001, Jiajia Li 0002, José Pineda de Gyvez
Integr.3
2018 PROBE: A Placement, Routing, Back-End-of-Line Measurement Utility
abstract
In advanced technology nodes, correctly choosing among available back-end-of-line (BEOL) stack options is important to meet stringent design quality of results requirements. However, it is nontrivial to evaluate BEOL stack options since the routing outcomes highly depend on the input design (e.g., netlist, placement, etc.). In this paper, we propose a systematic framework to measure routing capacity of a BEOL stack as well as inherent capability of routers. Based on our experimental results, we observe consistent results across mesh-like placement and placements from various placers. Also, our proposed framework enables new insights into important questions regarding BEOL stack options. Using our framework, we further study the relation between the routing hotspot size and routing failure empirically. Lastly, we present an analytical study based on exponentiation of a Markov transition matrix about the impact of design size on routing failure.
Alex Kahng, Andrew B. Kahng, Hyein Lee 0001, Jiajia Li 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Vertical M1 Routing-Aware Detailed Placement for Congestion and Wirelength Reduction in Sub-10nm Nodes
abstract
Aggressive pitch scaling in sub-10nm nodes has introduced complex design rules which make routing extremely challenging. Cell architectures have also been changed to meet the design rules. For example, metal layers below M1 are used to gain additional routing resources. New cell architectures wherein inter-row M1 routing is allowed force consideration of vertical alignment of cells. In this work, we propose a mixed-integer linear programming (MILP)-based, detailed placement optimization to maximize direct vertical M1 routing utilization for congestion and wirelength reduction.
Peter Debacker, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Praveen Raghavan, Lutong Wang
DAC4
2017 Optimal multi-row detailed placement for yield and model-hardware correlation improvements in sub-10nm VLSI
abstract
In sub-10nm, nodes, a change or step in diffusion height between adjacent standard cells causes yield loss as well as a form of model-hardware miscorrelation called neighbor diffusion effect (NDE). Cell libraries must inevitably have multiple diffusion heights (numbers of fins in PFETs and NFETs) in order to enable flexible exploration of the power-performance envelope for design. However, this brings step-induced risks of NDE, for which guardbanding is costly, as well as yield loss. Special filler cells can protect against harmful NDE effects, but are costly in terms of area. In this work, we develop dynamic programming-based single-row and double-row detailed placement optimizations that optimally minimize the impacts of NDE. Our algorithms support a richer set of cell movements than in previous works - i.e., flipping, relocating and reordering within the original row; we also consider cell displacement and flipping costs. Importantly, to our knowledge, our dynamic programming-based optimal detailed placement algorithm is the first to handle multiple rows with multiple-height cells that can be reordered. We further develop a timing-aware approach, which is capable of recovering (or, improving) the worst negative slack (WNS) by creating additional diffusion steps around timing-critical cells.
Changho Han, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Lutong Wang, Bangqi Xu
ICCAD4
2017 MILP-Based Optimization of 2-D Block Masks for Timing-Aware Dummy Segment Removal in Self-Aligned Multiple Patterning Layouts
abstract
Self-aligned multiple patterning, due to its low overlay error, has emerged as the leading option for 1-D gridded back-end-of-line (BEOL) in sub-14-nm nodes. To form actual routing patterns from a uniform “sea of wires,” cut masks are needed for line-end cutting or realization of space between routing segments. The line-end cutting results in nonfunctional (i.e., dummy fill) patterns that change wire capacitance, and hence design timing and power. Therefore, to remove such dummy fill patterns, extra 2-D block masks are used. However, 2-D block masks cannot remove arbitrary dummy fill patterns, due to design rule constraints on the block mask shapes. In this paper, we address the timing-aware optimization of 2-D block mask layouts under various sets of mask rules that are derived from mask patterning technology options (e.g., 193i and 193d) for foundry 7-/5-nm (N7/N5) BEOL. Our central contribution is a mixed integer linear programming (MILP) optimization that minimizes timing impact due to dummy metal segments while satisfying block mask rules and metal density constraints. We also propose a distributed optimization flow to improve the scalability. With our optimizer, we recover up to 84% of the worst negative slack impact from dummy segments, with up to 64% dummy removal rate. We further extend our MILP to a co-optimization of cut and block masks. This paper gives new insights into fundamental limits of benefit from emerging cut and block mask technology options.
Peter Debacker, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Praveen Raghavan, Lutong Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2016 Measuring progress and value of IC implementation technology
abstract
Over the past decade, “Moore's Law” has become increasingly well-understood as being a law of “value scaling”: success of new electronics- and semiconductor-based products depends on improved cost-efficiency, utility, and value. Design Automation (DA) provides fundamental tools and methodologies that glue together disparate technological advances - across architectures, circuits, process and integration - into actually-realized product and system benefits. However, it is still unknown how to measure and credit “progress” of DA with respect to realized product- and system-level benefits. Thus, it is also challenging for industry, academia and funding entities to envision, and define, R&D objectives and prioritizations. In this paper, we contend that “measuring progress and value” is tractable at the level of EDA technology and R&D efforts. We describe four example assessments of progress and value of IC implementation technology: (i) assessment of progress of EDA tools (e.g., for P&R and STA) across multiple releases; (ii) establishing “upper bounds” on future progress (e.g., for 3DIC layout); (iii) robust rank-ordering of alternative design enablements (e.g., including routing tools and interconnect stack options); and (iv) lowering barriers to measuring progress and value of academic research results in “more real-world” contexts.
Andrew B. Kahng, Hyein Lee 0001, Jiajia Li 0002
ICCAD2
2015 Evaluation of BEOL design rule impacts using an optimal ILP-based detailed router
abstract
Continued technology scaling with more pervasive use of multi-patterning has led to complex design rules and increased difficulty of maintaining high layout densities. Intuitively, emerging constraints such as unidirectional patterning or increased via spacing will decrease achievable density of the final place-and-route solution, worsening die area and product cost. However, no methodology exists for accurate assessment of design rules' impact on physical chip implementation. At the same time, this is a crucial need for early development of BEOL process technologies, particularly with FinFET or future vertical-device architectures where cell footprints can become much smaller than in bulk planar CMOS technologies. In this work, we study impacts of patterning technology choices and associated design rules on physical implementation density, with respect to cost-optimal design rule-correct detailed routing. A key contribution is an Integer Linear Programming (ILP) based optimal router (OptRouter) which considers complex design rules that arise in sub-20nm process technologies. Using OptRouter, we assess wirelength and via count impacts of various design rules (implicitly, patterning technology choices) by analyzing optimal routing solutions of clips (i.e., switchbox instances) extracted from post-detailed route layouts in an advanced technology.
Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001
DAC3
2015 Scalable Detailed Placement Legalization for Complex Sub-14nm Constraints
abstract
Technology scaling to 10nm and below introduces complex intra-row and inter-row constraints in standard-cell detailed placement. Examples of such constraints are found in rules for drain-drain abutment, minimum implant region area and width, oxide diffusion (OD) notching and jogging, etc. Typically, these rules are too complex for the normal global-detailed placement flow to fully consider. On the other hand, guardbanding the library cell design so that arbitrary cell placement adjacencies are all “correct by construction” has increasingly high area cost. This motivates the introduction of a final legalization phase for standard-cell placement tools in advanced (particularly 10nm and 7nm) foundry nodes. In this work, we develop a mixed integer-linear programming (MILP)-based placer, called DFPlacer, for final-phase design rule violation (DRV) fixing. DFPlacer finds (near-)DRV-free solutions considering various complex layout constraints including minimum implant width, drain-drain abutment, and oxide diffusion jogs. To overcome the runtime limitation of MILP-based approaches, we implement a distributable optimization strategy based on partitioning of the block layout into windows of cells that can be independently legalized. Using layouts in an abstracted 7nm library, we find that DFPlacer fixes 99% of DRVs on average with minimal impacts on area and timing. We also study an area-DRV tradeoff between two types of standard-cell library strategies, namely, with and without dummy poly gates.
Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001
ICCAD3
2014 Minimum implant area-aware gate sizing and placement
abstract
With reduction of minimum feature size, the minimum implant area (MinIA) constraint is emerging as a new challenge for the physical implementation flow in sub-22nm technology. In particular, the MinIA constraint induces a new problem formulation wherein gate sizing and V_t-swapping must now be linked closely with detailed placement changes. To solve this new problem, we propose heuristic methods that fix MinIA violations and reduce power with gate sizing while minimizing placement perturbation to avoid creating extra timing violations. Compared to recent versions of commercial P&R tools, our methodologies achieve significant reductions (up to 100%) in the number of MinIA violations under timing/power constraints.
Andrew B. Kahng, Hyein Lee 0001
ACM Great Lakes Symposium on VLSI2
2014 Horizontal benchmark extension for improved assessment of physical CAD research
abstract
The rapid growth in complexity and diversity of IC designs, design flows and methodologies has resulted in a benchmark-centric culture for evaluation of performance and scalability in physicaldesign algorithm research. Landmark papers in the literature present vertical benchmarks that can be used across multiple design flow stages; artificial benchmarks with characteristics that mimic those of real designs; artificial benchmarks with known optimal solutions; as well as benchmark suites created by major companies from internal designs and/or open-source RTL. However, to our knowledge, there has been no work on horizontal benchmark creation, i.e., the creation of benchmarks that enable maximal, comprehensive assessments across commercial and academic tools at one or more specific design stages. Typically, the creation of horizontal benchmarks is limited by mismatches in data models, netlist formats, technology files, library granularity, etc. across different tools, technologies, and benchmark suites. In this paper, we describe methodology and robust infrastructure for horizontal benchmark extension" that permits maximal leverage of benchmark suites and technologies in "apples-to-apples" assessment of both industry and academic optimizers. We demonstrate horizontal benchmark extensions, and the assessments that are thus enabled, in two well-studied domains: place-and-route (four combinations of academic placers/routers, and two commercial P&R tools) and gate sizing (two academic sizers, and three commercial tools). We also point out several issues and precepts for horizontal benchmark enablement.
Andrew B. Kahng, Hyein Lee 0001, Jiajia Li 0002
ACM Great Lakes Symposium on VLSI2
2013 Smart non-default routing for clock power reduction
abstract
At advanced process nodes, non-default routing rules (NDRs) are integral to clock network synthesis methodologies. NDRs apply wider wire widths and spacings to address electromigration constraints, and to reduce parasitic and delay variations. However, wider wires result in larger driven capacitance and dynamic power. In this work, we quantify the potential for capacitance and power reduction through the application of "smart" NDR (SNDR) that substitute narrower-width NDRs on selected clock network segments, while maintaining skew, slew, delay and EM reliability criteria. We propose a practical methodology to apply smart NDRs in standard clock tree synthesis flows. Our studies with a 32/28nm library and open-source benchmarks confirm substantial (average of 9.2%) clock wire capacitance reduction and an average of 4.9% clock switching power savings over the current fixed-NDR methodology, without loss of QoR in the clock distribution.
Andrew B. Kahng, Seokhyeong Kang, Hyein Lee 0001
DAC3
2013 High-performance gate sizing with a signoff timer
abstract
Process and device scaling in late-CMOS technologies highlight leakage power as a critical challenge for the semiconductor industry. Careful gate sizing and Vth-swapping can reduce leakage, but prior optimizations based on convex or dynamic programming (i) are often based on unrealistic assumptions about circuit delay and slew propagation, (ii) fail to handle practical design rules such as transition time or load upper bounds, and (iii) do not scale well to input complexities when full extracted parasitics are available. Seeing substantial opportunities for improvement, we present a multithreaded, stochastic optimization (Trident2.0) for gate sizing and Vthassignment to minimize leakage power subject to capacitance, slew and timing constraints. Scalability and high performance of Trident2.0 are validated on ISPD-2013 Gate Sizing Contest benchmarks.
Andrew B. Kahng, Seokhyeong Kang, Hyein Lee 0001, Igor L. Markov, Pankit Thapar
ICCAD3