Derong Liu 0002

dblp:181/2755-2 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
0since 2021 · last 2019
0000-0001-8949-6919ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 6 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Electronic design automation · 75% Energy-efficient computing · 12% Interconnection networks and networks-on-chip · 8%

Topics — the 21 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
physical design
2.172019
Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
TILA-S: Timing-Driven Incremental Layer Assignment Avoiding Slew Violations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
OPERON: optical-electrical power-efficient route synthesis for on-chip signals · DAC 2018
Electronic design automation › physical design › routing › multilayer routing
layer assignment
0.932018
TILA-S: Timing-Driven Incremental Layer Assignment Avoiding Slew Violations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Incremental Layer Assignment Driven by an External Signoff Timing Engine · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Incremental layer assignment for critical path timing · DAC 2016
Electronic design automation › physical design
routing
0.832019
Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Streak: Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups · DAC 2017
Incremental Layer Assignment Driven by an External Signoff Timing Engine · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Electronic design automation › physical design › routing › VLSI routing
bus routing
0.722019
Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Streak: Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups · DAC 2017
Electronic design automation › analog circuit synthesis
topology synthesis
0.722019
Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Streak: Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups · DAC 2017
Electronic design automation › physical design
gate sizing
0.522016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
OSFA: a new paradigm of gate-sizing for power/performance optimizations under multiple operating conditions · DAC 2015
Interconnection networks and networks-on-chip
on-chip interconnect
0.422018
OPERON: optical-electrical power-efficient route synthesis for on-chip signals · DAC 2018
Streak: Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups · DAC 2017
Interconnection networks and networks-on-chip
optical network-on-chip
0.312018
OPERON: optical-electrical power-efficient route synthesis for on-chip signals · DAC 2018
Energy-efficient computing › energy-efficient communication
power-aware routing
0.312018
OPERON: optical-electrical power-efficient route synthesis for on-chip signals · DAC 2018
Electronic design automation › physical design › routing
timing-driven routing
0.312018
TILA-S: Timing-Driven Incremental Layer Assignment Avoiding Slew Violations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Energy-efficient computing › power management
supply voltage optimization
0.322016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
OSFA: a new paradigm of gate-sizing for power/performance optimizations under multiple operating conditions · DAC 2015
Electronic design automation
timing analysis
0.312017
Incremental Layer Assignment Driven by an External Signoff Timing Engine · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Electronic design automation › timing analysis
timing sign-off
0.312017
Incremental Layer Assignment Driven by an External Signoff Timing Engine · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Hardware reliability and fault tolerance
aging
0.212016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Electronic design automation
logic synthesis
0.212016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Hardware reliability and fault tolerance › aging › transistor aging
negative bias temperature instability
0.212016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Energy-efficient computing
power management
0.212016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Electronic design automation › physical design
timing optimization
0.212016
Incremental layer assignment for critical path timing · DAC 2016
Energy-efficient computing
low-power design
0.212015
OSFA: a new paradigm of gate-sizing for power/performance optimizations under multiple operating conditions · DAC 2015
Electronic design automation › physical design › routing › routability
routability optimization
0.112019
Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Electronic design automation › physical design › routing
global routing
0.112017
Incremental Layer Assignment Driven by an External Signoff Timing Engine · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017

Methods — techniques the papers use, named apart from their topics

wire synthesis · 0.4topology generation · 0.4post-refinement · 0.4bottom-up clustering · 0.4multiprocessing · 0.3min-cost max-flow · 0.3min-cost flow · 0.3lagrangian relaxation · 0.3hyper net formulation · 0.3greedy algorithm · 0.3
YearPublicationVenuePosition
2019 Hardware-software co-design of slimmed optical neural networks
abstract
Optical neural network (ONN) is a neuromorphic computing hardware based on optical components. Since its first on-chip experimental demonstration, it has attracted more and more research interests due to the advantages of ultra-high speed inference with low power consumption. In this work, we design a novel slimmed architecture for realizing optical neural network considering both its software and hardware implementations. Different from the originally proposed ONN architecture based on singular value decomposition which results in two implementation-expensive unitary matrices, we show a more area-efficient architecture which uses a sparse tree network block, a single unitary block and a diagonal block for each neural network layer. In the experiments, we demonstrate that by leveraging the training engine, we are able to find a comparable accuracy to that of the previous architecture, which brings about the flexibility of using the slimmed implementation. The area cost in terms of the Mach-Zehnder interferometers, the core optical components of ONN, is 15%-38% less for various sizes of optical neural networks.
Zheng Zhao 0003, Derong Liu 0002, Meng Li 0004, Zhoufeng Ying, Biying Xu, Bei Yu 0001, Ray T. Chen, David Z. Pan
ASP-DAC2
2019 Exploiting Wavelength Division Multiplexing for Optical Logic Synthesis
abstract
Photonic integrated circuit (PIC), as a promising alternative to traditional CMOS circuit, has demonstrated the potential to accomplish on-chip optical interconnects and computations in ultra-high speed and/or low power consumption. Wavelength division multiplexing (WDM) is widely used in optical communication for enabling multiple signals being processed and transferred independently. In this work, we apply WDM to optical logic PIC synthesis to reduce the PIC area.
Zheng Zhao 0003, Derong Liu 0002, Zhoufeng Ying, Biying Xu, Chenghao Feng, Ray T. Chen, David Z. Pan
DATE2
2019 Device Layer-Aware Analytical Placement for Analog Circuits
abstract
The layouts of analog/mixed-signal (AMS) integrated circuits (ICs) are dramatically different from their digital counterparts. AMS circuit layouts usually include a variety of devices, including transistors, capacitors, resistors, and inductors. A complicated AMS IC system with hierarchical structure may also consist of pre-laid out subcircuits. Different types of devices can occupy different manufacturing layers. Therefore, during the layout stage, the devices require co-optimization to achieve high circuit performance. Leveraging the fact that some devices can be built by mutually exclusive layers, they can be carefully designed to overlap each other to effectively reduce the total area and wirelength without degrading the circuit performance. In this paper, we propose an analytical framework to tackle the device layer-aware analog placement problem. Experimental results show that on average the proposed techniques can reduce the total area and half-perimeter wirelength by 9% and 23%, respectively. To verify the routability of the placement results, we also develop an analog global router, which demonstrates that the device layer-aware placement can achieve 18% shorter wirelength during global routing.
Biying Xu, Shaolan Li, Chak-Wa Pui, Derong Liu 0002, Linxiao Shen, Yibo Lin, Nan Sun 0001, David Z. Pan
ISPD4
2019 Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups
abstract
As very large scale integration technology scales to deep submicron, design for interconnections becomes increasingly challenging. The traditional bus routing follows a sequential bit-by-bit order, and few works explicitly target interbit regularity for signal groups via multilayer topology selection. To overcome these limitations, we present Streak, an efficient framework that combines topology generation and wire synthesis with a global view of optimization and constrained metal layer track resource allocation. In the framework, an identification stage decomposes binding groups into a set of representative objects; with the generated backbones, equivalent topologies are accompanied by the bits in every object; then a formulation guides the routing considering wire congestion and design regularity. Furthermore, a bottom-up clustering methodology based on layer prediction targets to enhance the routability; a post-refinement stage is developed to match the source-to-sink distance deviation among bits in one group. Experimental results using industrial benchmarks demonstrate the effectiveness of the proposed technique.
Derong Liu 0002, Bei Yu 0001, Vinicius S. Livramento, Salim Chowdhury, Duo Ding, Huy Vo, Akshay Sharma, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 OPERON: optical-electrical power-efficient route synthesis for on-chip signals
abstract
As VLSI technology scales to deep sub-micron, optical interconnect becomes an attractive alternative for on-chip communication. The traditional optical routing works mainly optimize the path loss, and few works explicitly exploit the optical-electrical co-design of on-chip interconnects. To overcome these limitations, we present an efficient framework that directs the hybrid optical and electrical routes with a global view of power optimization. In this framework, on-chip signal bits are processed as hyper nets; the combination of optical and electrical routes are designed for hyper nets; then a formulation is given to find the appropriate solution of each hyper net and follows a speed-up algorithm; a min-cost max-flow network is utilized to reduce the consumed optical waveguides. Experimental results demonstrate the effectiveness of the proposed framework.
Derong Liu 0002, Zheng Zhao 0003, Zheng Wang 0036, Zhoufeng Ying, Ray T. Chen, David Z. Pan
DAC1
2018 Prim-Dijkstra Revisited: Achieving Superior Timing-driven Routing Trees
abstract
The Prim-Dijkstra (PD ) construction [1] was first presented over 20 years ago as a way to efficiently trade off between shortest-path and minimum-wirelength routing trees. This approach has stood the test of time, having been integrated into leading semiconductor design methodologies and electronic design automation tools. PD optimizes the conflicting objectives of wirelength (WL) and source-sink pathlength (PL) by blending the classic Prim and Dijkstra spanning tree algorithms. However, as this work shows, PD can sometimes demonstrate significant suboptimality for both WL and PL. This quality degradation can be especially costly for advanced nodes because (i) wire delays form a much larger component of total stage delay, i.e., timing-driven routing is critical, and (ii) modern designs are severely power-constrained (e.g., mobile, IoT), which makes low-capacitance wiring important. Consequently, achieving a good timing and power tradeoff for routing is required to build a market-leading product[2]. This work introduces a new problem formulation that incorporates the total detour cost in the objective function to optimize the detour to every sink in the tree, not just the worst detour. We then propose a new PD-II construction which directly improves upon the original PD construction by repairing the tree to simultaneously reduce both WL and PL. The PD-II approach achieves improvement for both objectives, making it a clear win over PD, for virtually zero additional runtime cost. PD-II is a spanning tree algorithm (which is useful for seeding global routing); however, since Steiner trees are needed for timing estimation, this work also includes a post-processing algorithm called DAS to convert PD-II trees into balanced Steiner trees. Experimental results demonstrate that this construction outperforms the recent state-of-the-art academic tool, SALT [36], for high-fanout nets, achieving up to 36.46% PL improvement with similar WL on average for 20K nets of size ≥ 32 terminals from DAC 2012 contest benchmark designs [37].
Charles J. Alpert, Wing-Kai Chow, Kwangsoo Han, Andrew B. Kahng, Zhuo Li 0001, Derong Liu 0002, Sriram Venkatesh
ISPD6
2018 TILA-S: Timing-Driven Incremental Layer Assignment Avoiding Slew Violations
abstract
As very large scale integration technology scales to deep submicrometer and beyond, interconnect delay greatly limits the circuit performance. The traditional 2-D global routing and subsequent net by net assignment of available empty tracks on various layers lacks a global view for timing optimization. To overcome the limitation, this paper presents a timing driven incremental layer assignment tool, to reassign layers among routing segments of critical nets and noncritical nets. Lagrangian relaxation techniques are proposed to iteratively provide consistent layer/via assignments. Modeling via min-cost flow for layer shuffling avoids using integer programming and yet guarantees integer solutions via uni-modular property of the inherent model. In addition, multiprocessing of K × K partitions of the whole chip provides runtime speed up. Furthermore, a slew targeted optimization is presented to reduce the number of violations incrementally through iteration-based Lagrangian relaxation, followed by a post greedy algorithm to fix local violations. Certain parameters introduced in the models provide tradeoff between timing optimization and via count. Experimental results in both ISPD 2008 and industry benchmark suites demonstrate the effectiveness of the proposed incremental algorithms.
Derong Liu 0002, Bei Yu 0001, Salim Chowdhury, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2017 Streak: Synergistic Topology Generation and Route Synthesis for On-Chip Performance-Critical Signal Groups
abstract
As VLSI technology scales to deep sub-micron, design for interconnections becomes increasingly challenging. The traditional bus routing follows a sequential bit-by-bit order, and few works explicitly target inter-bit regularity for signal groups via multilayer topology selection. To overcome these limitations, we present Streak, an efficient framework that combines topology generation and wire synthesis with a global view of optimization and constrained metal layer track resource allocation. In the framework, an identification stage decomposes binding groups into a set of representative objects; with the generated backbones, equivalent topologies are accompanied by the bits in every object; then a formulation guides the routing considering wire congestion and design regularity. Experimental results using industrial benchmarks demonstrate the effectiveness of the proposed technique.
Derong Liu 0002, Vinicius S. Livramento, Salim Chowdhury, Duo Ding, Huy Vo, Akshay Sharma, David Z. Pan
DAC1
2017 Incremental Layer Assignment Driven by an External Signoff Timing Engine
abstract
Modern technologies provide wide and thick metal layers that must be wisely used to reduce the delay of critical interconnections. After global routing, incremental layer assignment can improve the circuit timing by properly selecting critical interconnect segments to be routed in the faster (but very limited) wires on upper layers. Existing techniques based on net-by-net iterative improvement may get stuck at locally-optimal solutions depending on net ordering. Recent techniques rule out such drawback through the simultaneous iterative improvement of all nets, but they unfortunately rely on objective functions that may guide the optimization off critical paths. As opposed to all reported techniques, which rely on simplified, overly pessimistic timing models, this paper proposes the decoupling of incremental layer assignment from the timing analysis and the exploitation of flow conservation conditions so as to enable the use of an external signoff timing engine. The novel technique was experimentally compared with two state-of-the art works, leading to 50% less timing violations under total negative slack metric and 35% less timing violations under worst negative slack metric with similar overhead in number of vias.
Vinicius S. Livramento, Derong Liu 0002, Salim Chowdhury, Bei Yu 0001, David Z. Pan, José Luís Güntzel, Luiz Cláudio Villar dos Santos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Incremental Layer Assignment for Timing Optimization
abstract
With VLSI technology nodes scaling into the nanometer regime, interconnect delay plays an increasingly critical role in timing. For layer assignment, most works deal with via counts or total net delays, ignoring critical paths of each net and resulting in potential timing issues. In this article, we propose an incremental layer assignment framework targeting delay optimization in timing the critical path of each net. A set of novel techniques are presented: self-adaptive quadruple partition based on K × K division benefits the runtime; semidefinite programming is utilized for each partition; and the sequential mapping algorithm guarantees integer solutions while satisfying edge capacities; additionally, concurrent mapping offers a global view of assignment and post delay optimization reduces the path timing violations. The effectiveness of our work is verified by ISPD’08 benchmarks.
Derong Liu 0002, Bei Yu 0001, Salim Chowdhury, David Z. Pan
ACM Trans. Design Autom. Electr. Syst.1
2016 Incremental layer assignment for critical path timing
abstract
With VLSI technology nodes scaling into nanometer regime, interconnect delay plays an increasingly critical role in timing. For layer assignment, most works deal with via counts or total net delays, ignoring critical paths of each net and resulting in potential timing issues. In this paper we propose an incremental layer assignment framework targeting at delay optimization for critical path of each net. A set of novel techniques are presented: self-adaptive quadruple partition based on KxK division benefits the run-time; semidefinite programming is utilized for each partition; post mapping algorithm guarantees integer solutions while satisfying edge capacities. The effectiveness of our work is verified by ISPD'08 benchmarks.
Derong Liu 0002, Bei Yu 0001, Salim Chowdhury, David Z. Pan
DAC1
2016 OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions
abstract
Modern systems-on-a-chip and microprocessors, e.g., those in smart phones and laptops, typically have multiple operating conditions, such as video streaming, Web browsing, standby, and so on. They will have different performance targets and run under different supply voltages. Gate sizing (with threshold voltage assignment) is a fundamental step for power/performance optimization. However, conventional gate sizing algorithms only consider one scenario, e.g., the performance-critical operating condition, which may be over-design for other operating conditions. In addition, reliability has become a prime concern in nanometer designs, and gate sizing has been employed to mitigate aging. However: 1) previous aging-affected delay models do not take into account more than one operating condition to estimate the aging impact and 2) earlier aging aware gate sizing algorithms only consider one operating condition at a time. In this paper, we present a new paradigm of aging aware gate sizing, one-size-fits-all (OSFA), which performs power/performance optimizations across multiple operating conditions. The existing delay model for negative bias temperature instability (NBTI) is extended to take into account multiple operating conditions, and incorporated into our OSFA framework. Based on OSFA, we also adjust the supply voltage targeting overall power optimization. A speed-up heuristic is proposed to scale our OSFA design space exploration methodology for higher number of operating conditions. Experimental results on industry-strength benchmarks demonstrate that: 1) compared with conventional approach OSFA could provide an average 6.1% reduction in power without performance loss; 2) NBTI-aware OSFA framework can provide significant improvement in comparison with guard-band based traditional NBTI-aware gate sizing approach; and 3) percentage savings compared to conventional methodology increases with the number of operating conditions.
Subhendu Roy, Derong Liu 0002, Jagmohan Singh, Junhyung Um, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 OSFA: a new paradigm of gate-sizing for power/performance optimizations under multiple operating conditions
abstract
Modern SoCs and microprocessors, e.g., those in smart phones and laptops, typically have multiple operating conditions, such as video streaming, web browsing, standby, and so on. They will have different performance targets and run under different supply voltages. Gate sizing (with threshold voltage assignment) is a fundamental step for power/performance optimization. However, conventional gate sizing algorithms only consider one scenario, e.g., the performance-critical operating condition, which may be over-design for other operating conditions. In this paper, we present a new paradigm of gate sizing, OSFA (One-Size-Fits-All), which performs power/performance optimizations across multiple operating conditions. Based on OSFA, we also adjust the supply voltage targeting overall power optimization. Experimental results on industry-strength benchmarks demonstrate that compared with conventional approach OSFA could provide an average 6.1% reduction in power without performance loss.
Subhendu Roy, Derong Liu 0002, Junhyung Um, David Z. Pan
DAC2
2015 TILA: Timing-Driven Incremental Layer Assignment
abstract
As VLSI technology scales to deep submicron and beyond, interconnect delay greatly limits the circuit performance. The traditional 2D global routing and subsequent net by net assignment of available empty tracks on various layers lacks a global view for timing optimization. To overcome the limitation, this paper presents a timing driven incremental layer assignment tool, TILA, to reassign layers among routing segments of critical nets and non-critical nets. Lagrangian relaxation techniques are proposed to iteratively provide consistent layer/via assignments. Modeling via min-cost flow for layer shuffling avoids using integer programming and yet guarantees integer solutions via uni-modular property of the inherent model. In addition, multiprocessing of K × K partitions of the whole chip provides run time speed up. Certain parameters introduced in the models provide trade-off between timing optimization and via count. Experimental results in both ISPD'08 and industry benchmark suites demonstrate the effectiveness of the proposed incremental algorithms.
Bei Yu 0001, Derong Liu 0002, Salim Chowdhury, David Z. Pan
ICCAD2