EDBT 2026 Demo / reviewers in the wild / expert
Kwangsoo Han
dblp:145/7573
· DBLP profile ↗
14ranked-venue papers
5as first author
1since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Electronic design automation · 95% Integrated circuit design · 2% Energy-efficient computing · 2% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
physical design |
1.7 | 6 | 2020 | Optimal Generalized H-Tree Topology and Buffering for High-Performance and Low-Power Clock Distribution · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 MILP-Based Optimization of 2-D Block Masks for Timing-Aware Dummy Segment Removal in Self-Aligned Multiple Patterning Layouts · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 Benchmarking of Mask Fracturing Heuristics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Electronic design automation › physical design › clock network synthesis
clock network optimization |
0.4 | 1 | 2020 | Optimal Generalized H-Tree Topology and Buffering for High-Performance and Low-Power Clock Distribution · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Electronic design automation › physical design › clock network synthesis
clock tree synthesis |
0.4 | 1 | 2020 | Optimal Generalized H-Tree Topology and Buffering for High-Performance and Low-Power Clock Distribution · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Electronic design automation › physical design › placement
detailed placement |
0.3 | 1 | 2017 | Vertical M1 Routing-Aware Detailed Placement for Congestion and Wirelength Reduction in Sub-10nm Nodes · DAC 2017 |
Electronic design automation › physical design › optical proximity correction
inverse lithography technology |
0.3 | 1 | 2017 | Benchmarking of Mask Fracturing Heuristics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Electronic design automation › physical design
mask optimization |
0.3 | 1 | 2017 | Benchmarking of Mask Fracturing Heuristics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Electronic design automation › physical design › lithography
multiple patterning lithography |
0.3 | 1 | 2017 | MILP-Based Optimization of 2-D Block Masks for Timing-Aware Dummy Segment Removal in Self-Aligned Multiple Patterning Layouts · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Electronic design automation › physical design
buffer insertion |
0.2 | 1 | 2015 | A global-local optimization framework for simultaneous multi-mode multi-corner clock skew variation reduction · DAC 2015 |
Electronic design automation › physical design
clock network synthesis |
0.2 | 1 | 2015 | A global-local optimization framework for simultaneous multi-mode multi-corner clock skew variation reduction · DAC 2015 |
Electronic design automation › physical design › clock network synthesis
clock skew optimization |
0.2 | 1 | 2015 | A global-local optimization framework for simultaneous multi-mode multi-corner clock skew variation reduction · DAC 2015 |
Electronic design automation › physical design › routing
detailed routing |
0.2 | 1 | 2015 | Evaluation of BEOL design rule impacts using an optimal ILP-based detailed router · DAC 2015 |
Energy-efficient computing › dynamic power reduction
clock power reduction |
0.1 | 1 | 2020 | Optimal Generalized H-Tree Topology and Buffering for High-Performance and Low-Power Clock Distribution · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Integrated circuit design
low-power circuit design |
0.1 | 1 | 2020 | Optimal Generalized H-Tree Topology and Buffering for High-Performance and Low-Power Clock Distribution · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Electronic design automation › physical design
routing |
0.1 | 1 | 2017 | Vertical M1 Routing-Aware Detailed Placement for Congestion and Wirelength Reduction in Sub-10nm Nodes · DAC 2017 |
Methods — techniques the papers use, named apart from their topics
linear programming · 0.7mixed integer linear programming · 0.6integer linear programming · 0.5k-means clustering · 0.4dynamic programming · 0.4distributed optimization · 0.3branch-and-price · 0.3benchmark generation · 0.3machine learning-based latency prediction · 0.2buffer sizing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | 2024 ICCAD CAD Contest Problem A: Reinforcement Logic Optimization for a General Cost FunctionabstractTraditionally, logic synthesis/optimization metric would be majorly determined by PPA (power, performance, area). However, as the technology node shrinks and the design process becomes extremely complicated, iterative optimization flow and local re-synthesis might be invoked to optimize for more variant purposes. It is necessary to have a methodology which is not just a simple cost-function-based algorithm but also can optimize and legalize a design according to a more complex cost estimator. Chung-Han Chou, Chih-Jen Hsu, Chi-An Wu, Kuan-Hua Tu, Kwangsoo Han, Zhuo Li 0001 |
ICCAD | 5 |
| 2020 | Optimal Generalized H-Tree Topology and Buffering for High-Performance and Low-Power Clock DistributionabstractClock power, skew and maximum latency are three key metrics for clock distribution in low-power and high-performance designs. An H-tree offers minimum clock skew and good robustness against variations, but at the cost of large wirelength and clock power. On the other hand, a “fishbone” clock network with spine-ribs structures has smaller wirelength, latency and clock power, but larger skew, as compared to an H-tree. No previous work enables systematic exploration of the regime between H-tree and spine to achieve an optimal tradeoff among clock power, skew, and latency. In this paper, we study the concept of a generalized H-tree (GH-tree)-a topologically balanced tree with an arbitrary sequence of branching factors-and propose a dynamic programming-based method to determine optimal clock power, skew, and latency, in the space of GH-tree solutions. Our method co-optimizes clock tree topology and buffering along branches according to fitted electrical models. We further propose a balanced K-means clustering and a linear programming (LP)-guided buffer placement approach to embed the GH-tree with respect to a given sink placement. We validate our solutions in commercial clock tree synthesis (CTS) tool flows, in a commercial foundry's 28LP technology. The results show up to 30% clock power reduction while achieving similar skew and maximum latency as CTS solutions from recent versions of leading commercial place-and-route tools. Our proposed approach also achieves up to 56% clock power reduction while achieving similar skew and maximum latency as compared to CTS solutions from a state-of-the-art academic tool. Kwangsoo Han, Andrew B. Kahng, Jiajia Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Prim-Dijkstra Revisited: Achieving Superior Timing-driven Routing TreesabstractThe Prim-Dijkstra (PD ) construction [1] was first presented over 20 years ago as a way to efficiently trade off between shortest-path and minimum-wirelength routing trees. This approach has stood the test of time, having been integrated into leading semiconductor design methodologies and electronic design automation tools. PD optimizes the conflicting objectives of wirelength (WL) and source-sink pathlength (PL) by blending the classic Prim and Dijkstra spanning tree algorithms. However, as this work shows, PD can sometimes demonstrate significant suboptimality for both WL and PL. This quality degradation can be especially costly for advanced nodes because (i) wire delays form a much larger component of total stage delay, i.e., timing-driven routing is critical, and (ii) modern designs are severely power-constrained (e.g., mobile, IoT), which makes low-capacitance wiring important. Consequently, achieving a good timing and power tradeoff for routing is required to build a market-leading product[2]. This work introduces a new problem formulation that incorporates the total detour cost in the objective function to optimize the detour to every sink in the tree, not just the worst detour. We then propose a new PD-II construction which directly improves upon the original PD construction by repairing the tree to simultaneously reduce both WL and PL. The PD-II approach achieves improvement for both objectives, making it a clear win over PD, for virtually zero additional runtime cost. PD-II is a spanning tree algorithm (which is useful for seeding global routing); however, since Steiner trees are needed for timing estimation, this work also includes a post-processing algorithm called DAS to convert PD-II trees into balanced Steiner trees. Experimental results demonstrate that this construction outperforms the recent state-of-the-art academic tool, SALT [36], for high-fanout nets, achieving up to 36.46% PL improvement with similar WL on average for 20K nets of size ≥ 32 terminals from DAC 2012 contest benchmark designs [37]. Charles J. Alpert, Wing-Kai Chow, Kwangsoo Han, Andrew B. Kahng, Zhuo Li 0001, Derong Liu 0002, Sriram Venkatesh |
ISPD | 3 |
| 2017 | Vertical M1 Routing-Aware Detailed Placement for Congestion and Wirelength Reduction in Sub-10nm NodesabstractAggressive pitch scaling in sub-10nm nodes has introduced complex design rules which make routing extremely challenging. Cell architectures have also been changed to meet the design rules. For example, metal layers below M1 are used to gain additional routing resources. New cell architectures wherein inter-row M1 routing is allowed force consideration of vertical alignment of cells. In this work, we propose a mixed-integer linear programming (MILP)-based, detailed placement optimization to maximize direct vertical M1 routing utilization for congestion and wirelength reduction. Peter Debacker, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Praveen Raghavan, Lutong Wang |
DAC | 2 |
| 2017 | Optimal multi-row detailed placement for yield and model-hardware correlation improvements in sub-10nm VLSIabstractIn sub-10nm, nodes, a change or step in diffusion height between adjacent standard cells causes yield loss as well as a form of model-hardware miscorrelation called neighbor diffusion effect (NDE). Cell libraries must inevitably have multiple diffusion heights (numbers of fins in PFETs and NFETs) in order to enable flexible exploration of the power-performance envelope for design. However, this brings step-induced risks of NDE, for which guardbanding is costly, as well as yield loss. Special filler cells can protect against harmful NDE effects, but are costly in terms of area. In this work, we develop dynamic programming-based single-row and double-row detailed placement optimizations that optimally minimize the impacts of NDE. Our algorithms support a richer set of cell movements than in previous works - i.e., flipping, relocating and reordering within the original row; we also consider cell displacement and flipping costs. Importantly, to our knowledge, our dynamic programming-based optimal detailed placement algorithm is the first to handle multiple rows with multiple-height cells that can be reordered. We further develop a timing-aware approach, which is capable of recovering (or, improving) the worst negative slack (WNS) by creating additional diffusion steps around timing-critical cells. Changho Han, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Lutong Wang, Bangqi Xu |
ICCAD | 2 |
| 2017 | Benchmarking of Mask Fracturing HeuristicsabstractAggressive resolution enhancement techniques such as inverse lithography (ILT) often lead to complex, nonrectilinear mask shapes which make mask writing extremely slow and expensive. To reduce shot count of complex mask shapes, mask writers allow overlapping shots, due to which the problem of fracturing mask shapes with minimum shot count is NP-hard. The need to account for e-beam proximity effect makes mask fracturing even more challenging. Although a number of fracturing heuristics have been proposed, there has been no systematic study to analyze the quality of their solutions. In this paper, we first propose a method to generate tight upper and lower bounds for actual ILT mask shapes by formulating mask fracturing as an integer linear program and solving it using branch and price. Since the integer program requires significant computational resources to compute reasonable bounds, we propose a new method to generate benchmarks with known optimal solutions, that can be used to evaluate the suboptimality of mask fracturing heuristics. To make the generated benchmark shapes realistic, we further propose a novel automated benchmark generation method that takes any ILT shape as input and returns a benchmark shape which looks similar to the input shape and for which the optimal fracturing solution is known. Using these methods, we compare the suboptimality of four mask fracturing heuristics. Our results show that even a state-of-the-art prototype (version of) capability within a commercial EDA tool for e-beam mask shot decomposition can be suboptimal by as much as 2.6× for real ILT shapes and by 6.0× for generated benchmarks. Tuck-Boon Chan, Puneet Gupta 0001, Kwangsoo Han, Abde Ali Kagalwalla, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | MILP-Based Optimization of 2-D Block Masks for Timing-Aware Dummy Segment Removal in Self-Aligned Multiple Patterning LayoutsabstractSelf-aligned multiple patterning, due to its low overlay error, has emerged as the leading option for 1-D gridded back-end-of-line (BEOL) in sub-14-nm nodes. To form actual routing patterns from a uniform “sea of wires,” cut masks are needed for line-end cutting or realization of space between routing segments. The line-end cutting results in nonfunctional (i.e., dummy fill) patterns that change wire capacitance, and hence design timing and power. Therefore, to remove such dummy fill patterns, extra 2-D block masks are used. However, 2-D block masks cannot remove arbitrary dummy fill patterns, due to design rule constraints on the block mask shapes. In this paper, we address the timing-aware optimization of 2-D block mask layouts under various sets of mask rules that are derived from mask patterning technology options (e.g., 193i and 193d) for foundry 7-/5-nm (N7/N5) BEOL. Our central contribution is a mixed integer linear programming (MILP) optimization that minimizes timing impact due to dummy metal segments while satisfying block mask rules and metal density constraints. We also propose a distributed optimization flow to improve the scalability. With our optimizer, we recover up to 84% of the worst negative slack impact from dummy segments, with up to 64% dummy removal rate. We further extend our MILP to a co-optimization of cut and block masks. This paper gives new insights into fundamental limits of benefit from emerging cut and block mask technology options. Peter Debacker, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Praveen Raghavan, Lutong Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Delay uncertainty and signal criticality driven routing channel optimization for advanced DRAM productsabstractSignal delay uncertainty induced by crosstalk is a critical challenge to the physical design of long interconnect channels in DRAM products at the 2× and 1× technology nodes. Due to severe cost challenges in a high-volume, commodity market, layout resources including channel width, buffers, and number of metal routing layers are extremely scarce. We describe a new channel optimizer that reduces crosstalk-induced delay uncertainty, weighted by signal criticality and aware of signal activity correlations (e.g., to reduce delay uncertainty by mutual shielding). Instead of the typical signal net permutation strategy, we apply (pessimistic) timing-driven swizzling to minimize the delay uncertainty cost function. Contributions of this work include (1) an accurate and efficient analytical crosstalk delay calculator, (2) scalability up to hundreds of signals and tracks in the routing channel through use of greedy and decomposition strategies as well as a pair-swapping approach, and (3) experimental studies that demonstrate up to 24% reduction of the worst-case criticalityweighted delay uncertainty (or, 34ps of absolute delay uncertainty reduction) compared with the typical signal permutation approach. Samyoung Bang, Kwangsoo Han, Andrew B. Kahng, Mulong Luo |
ASP-DAC | 2 |
| 2016 | Improved performance of 3DIC implementations through inherent awareness of mix-and-match die stacking
Kwangsoo Han, Andrew B. Kahng, Jiajia Li 0002 |
DATE | 1 |
| 2015 | Evaluation of BEOL design rule impacts using an optimal ILP-based detailed routerabstractContinued technology scaling with more pervasive use of multi-patterning has led to complex design rules and increased difficulty of maintaining high layout densities. Intuitively, emerging constraints such as unidirectional patterning or increased via spacing will decrease achievable density of the final place-and-route solution, worsening die area and product cost. However, no methodology exists for accurate assessment of design rules' impact on physical chip implementation. At the same time, this is a crucial need for early development of BEOL process technologies, particularly with FinFET or future vertical-device architectures where cell footprints can become much smaller than in bulk planar CMOS technologies. In this work, we study impacts of patterning technology choices and associated design rules on physical implementation density, with respect to cost-optimal design rule-correct detailed routing. A key contribution is an Integer Linear Programming (ILP) based optimal router (OptRouter) which considers complex design rules that arise in sub-20nm process technologies. Using OptRouter, we assess wirelength and via count impacts of various design rules (implicitly, patterning technology choices) by analyzing optimal routing solutions of clips (i.e., switchbox instances) extracted from post-detailed route layouts in an advanced technology. Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001 |
DAC | 1 |
| 2015 | A global-local optimization framework for simultaneous multi-mode multi-corner clock skew variation reductionabstractAs combinations of signoff corners grow in modern SoCs, minimization of clock skew variation across corners is important. Large skew variation can cause difficulties in multi-corner timing closure because fixing violations at one corner can lead to violations at other corners. Such "ping-pong" effects lead to significant power and area overheads and time to signoff. We propose a novel framework encompassing both global and local clock network optimizations to minimize the sum of skew variations across different PVT corners between all sequentially adjacent sink pairs. The global optimization uses linear programming to guide buffer insertion, buffer removal and routing detours. The local optimization is based on machine learning-based predictors of latency change; these are used for iterative optimization with tree surgery, buffer sizing and buffer displacement operators. Our optimization achieves up to 22% total skew variation reduction across multiple testcases implemented in foundry 28nm technology, as compared to a best-practices CTS solution using a leading commercial tool. Kwangsoo Han, Jiajia Li 0002, Andrew B. Kahng, Siddhartha Nath, Jongpil Lee |
DAC | 1 |
| 2015 | Scalable Detailed Placement Legalization for Complex Sub-14nm ConstraintsabstractTechnology scaling to 10nm and below introduces complex intra-row and inter-row constraints in standard-cell detailed placement. Examples of such constraints are found in rules for drain-drain abutment, minimum implant region area and width, oxide diffusion (OD) notching and jogging, etc. Typically, these rules are too complex for the normal global-detailed placement flow to fully consider. On the other hand, guardbanding the library cell design so that arbitrary cell placement adjacencies are all “correct by construction” has increasingly high area cost. This motivates the introduction of a final legalization phase for standard-cell placement tools in advanced (particularly 10nm and 7nm) foundry nodes. In this work, we develop a mixed integer-linear programming (MILP)-based placer, called DFPlacer, for final-phase design rule violation (DRV) fixing. DFPlacer finds (near-)DRV-free solutions considering various complex layout constraints including minimum implant width, drain-drain abutment, and oxide diffusion jogs. To overcome the runtime limitation of MILP-based approaches, we implement a distributable optimization strategy based on partitioning of the block layout into windows of cells that can be independently legalized. Using layouts in an abstracted 7nm library, we find that DFPlacer fixes 99% of DRVs on average with minimal impacts on area and timing. We also study an area-DRV tradeoff between two types of standard-cell library strategies, namely, with and without dummy poly gates. Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001 |
ICCAD | 1 |
| 2014 | OCV-aware top-level clock tree optimizationabstractThe clock trees of high-performance synchronous circuits have many clock logic cells (e.g., clock gating cells, multiplexers and dividers) in order to achieve aggressive clock gating and required performance across a wide range of operating modes and conditions. As a result, clock tree structures have become very complex and difficult to optimize with automatic clock tree synthesis (CTS) tools. In advanced process nodes, CTS becomes even more challenging due to on-chip variation (OCV) effects. In this paper, we present a new CTS methodology that optimizes clock logic cell placements and buffer insertions in the top level of a clock tree. We formulate the top-level clock tree optimization problem as a linear program that minimizes a weighted sum of timing slacks, clock uncertainty and wirelength. Experimental results in a commercial 28nm FDSOI technology show that our method can improve post-CTS worst negative slack across all modes/corners by up to 320ps compared to a leading commercial provider's CTS flow. Tuck-Boon Chan, Kwangsoo Han, Andrew B. Kahng, Jae-Gon Lee, Siddhartha Nath |
ACM Great Lakes Symposium on VLSI | 2 |
| 2014 | Benchmarking of mask fracturing heuristicsabstractAggressive resolution enhancement techniques such as inverse lithography (ILT) often lead to complex, non-rectilinear mask shapes which make mask writing extremely slow and expensive. To reduce shot count of complex mask shapes, mask writers allow overlapping shots, due to which the problem of fracturing mask shapes with minimum shot count is NP-hard. The need to correct for e-beam proximity effect makes mask fracturing even more challenging. Although a number of fracturing heuristics have been proposed, there has been no systematic study to analyze the quality of their solutions. In this work, we propose a new method to generate benchmarks with known optimal solutions that can be used to evaluate the suboptimality of mask fracturing heuristics. We also propose a method to generate tight upper and lower bounds for actual ILT mask shapes by formulating mask fracturing as an integer linear program and solving it using branch and price. Our results show that a state-of-the-art prototype [version of] capability within a commercial EDA tool for e-beam mask shot decomposition can be suboptimal by as much as 3.7× for generated benchmarks, and by as much as 3.6× for actual ILT shapes. Tuck-Boon Chan, Puneet Gupta 0001, Kwangsoo Han, Abde Ali Kagalwalla, Andrew B. Kahng, Emile Sahouria |
ICCAD | 3 |