Wai-Chung Tang

dblp:85/6096 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Electronic design automation · 74% Reconfigurable computing and FPGAs · 26%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA physical design
0.212014
A scalable routability-driven analytical placer with global router integration for FPGAs (abstract only) · FPGA 2014
Electronic design automation › physical design
placement
0.212014
A scalable routability-driven analytical placer with global router integration for FPGAs (abstract only) · FPGA 2014
Electronic design automation › physical design
clock routing
0.112012
Postgrid Clock Routing for High Performance Microprocessor Designs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Electronic design automation
physical design
0.112012
Postgrid Clock Routing for High Performance Microprocessor Designs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Reconfigurable computing and FPGAs
FPGA design flow
0.112007
How Much Can Logic Perturbation Help from Netlist to Final Routing for FPGAs · DAC 2007
Electronic design automation
logic synthesis
0.112007
How Much Can Logic Perturbation Help from Netlist to Final Routing for FPGAs · DAC 2007
Electronic design automation › logic synthesis › logic restructuring
rewiring
0.112007
How Much Can Logic Perturbation Help from Netlist to Final Routing for FPGAs · DAC 2007
Electronic design automation › physical design › routing
global routing
0.112014
A scalable routability-driven analytical placer with global router integration for FPGAs (abstract only) · FPGA 2014
Electronic design automation › physical design
clock network synthesis
0.012012
Postgrid Clock Routing for High Performance Microprocessor Designs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Electronic design automation › physical design
placement and routing
0.012007
How Much Can Logic Perturbation Help from Netlist to Final Routing for FPGAs · DAC 2007
Electronic design automation › physical design › routing › routability
routability optimization
0.012007
How Much Can Logic Perturbation Help from Netlist to Final Routing for FPGAs · DAC 2007

Methods — techniques the papers use, named apart from their topics

conjugate gradient · 0.2bound2bound net model · 0.2partition-based path expansion · 0.1rewiring · 0.1logic perturbation · 0.1
YearPublicationVenuePosition
2014 A scalable routability-driven analytical placer with global router integration for FPGAs (abstract only)
abstract
As the sizes of modern circuits become bigger and bigger, implementing those large circuits into FPGA becomes arduous. The state-of-the-art academic FPGA place-and-route tool, VPR, has good quality but needs around a whole day to complete a placement when the input circuit contains millions of lookup tables, excluding the runtime for routing. To expedite the placement process, we propose a routability-driven placement algorithm for FPGA that adopts techniques used in ASIC global placer. Our placer follows the lower-bound-and-upper-bound iterative optimization process in ASIC placers like Ripple. In the lower-bound computation, the total HPWL, modeled using the Bound2Bound net model, is minimized using the conjugate gradient method. In the upper-bound computation, an almost-legalized result is produced by spreading cells linearly in the placement area. Those positions are then served as fixed-point anchors and fed into the next lower-bound computation. Furthermore, global routing will be performed in the upper-bound computation to estimate the routing segment usage, as a mean to consider congestion in placement. We tested our approach using 20 MCNC benchmarks and 4 large benchmarks for performance and scalability. Experimental results show that based on the island-style architecture which VPR is most optimized for, our approach can obtain a placement result 8x faster than VPR with 2% more in channel width, or 3x faster with 1% more in channel width when congestion is being considered. Our approach is even 14x faster than VPR in placing large benchmarks with over 10,000 lookup tables, with only 7% more in channel width.
Ka-Chun Lam, Wai-Chung Tang, Evangeline F. Y. Young
FPGA2
2013 Mountain-mover: An intuitive logic shifting heuristic for improving timing slack violating paths
abstract
Based on a simple intuitive notion, in this paper, we propose an efficient post-placement improvement scheme. Based on the given timing slack distribution of a circuit, a corresponding “slack mountain map” can be visualized with peaks representing most violating (negative) slacks and valleys representing non-critical (positive) slacks respectively. Guided by this map, violating paths are eliminated or improved when slack mountains are flattened by applying a local logic perturbation technique (rewiring) iteratively to shift logic resources from critical to non-critical areas. Due to the locality property of the rewiring technique, to better avoid being stuck at local minimums, instead of running rewiring operations from the peak top towards lower areas, we do this local logic shifting starting from “sea areas” (most non-critical) towards peak (most critical) areas. At the end, as the slack map is more flattened, a circuit with slack violations more evenly distributed can be yielded. Comparing to the recent work [1], our experimental results demonstrate that this scheme can obtain a better or comparable delay reduction but with CPU time one order of magnitude smaller.
Wai-Chung Tang, Yu-Liang Wu, Cliff C. N. Sze, Charles J. Alpert
ASP-DAC2
2012 ECO timing optimization with negotiation-based re-routing and logic re-structuring using spare cells
abstract
To maintain a lower re-masking cost, Engineering Change Order (ECO) using pre-placed spare cells for buffer insertion and gate sizing has been shown to be practical for fixing timing violating paths (ECO paths). However, in the previously known best scheme DCP [1, 2], re-routings are done with each path optimized according to its surrounding available spare cells without considering potential exchanges with neighboring active cells, and spare cell arbitration between competing ECO paths are less addressed. Besides, the extra flexibility for allowing logic restructuring was not exploited. In this work, we develop a framework harnessing the following more flexible strategies to make the usage of spare cells for ECO timing optimization more powerful: (1) a negotiation based re-routing scheme yielding a more global view in solving resource competition arbitration; (2) an extended gate sizing operation to allow exchanges of active gates with spare gates of different function types through equivalent logic re-structuring. Our experiments upon MCNC and ITC benchmarks with highly injected timing violations show that compared to DCP, our newly proposed framework can cut down the average total negative slack (TNS) by 50% and reduce the number of unsolved ECO paths by 31%.
Wai-Chung Tang, Yi Diao, Yu-Liang Wu
ASP-DAC2
2012 Almost every wire is removable: A modeling and solution for removing any circuit wire
abstract
Rewiring is a flexible and useful logic transformation technique through which a target wire can be removed by adding its alternative logics without changing the circuit functionality. In today's deep sub-micron era, circuit wires have become a dominating factor in most EDA processes and there are situations where removing a certain set of (perhaps extremely unwanted) wires is very useful. However, it has been experimentally suggested that the rewiring rate (percentage of original circuit wires being removable by rewiring) is only 30 to 40% for optimized circuits in the past. In this paper, we propose a generalized error cancellation modeling and flow to show that theoretically almost every circuit wire is removable under this flow. In the Flow graph Error Cancellation based Rewiring (FECR) scheme we propose here, a rewiring rate of 95% of even optimized circuits is obtainable under this scheme, affirming the basic claim of this paper. To our knowledge, this is the first known rewiring scheme being able to achieve this near complete rewiring rate. Consequently, this wire-removal process can now be considered as a powerful atomic and universal operation for logic transformations, as virtually every circuit node can also be removed through repetitions of this rewiring process. Besides, this modeling can also serve as a general framework containing many other rewiring techniques as its special cases.
Tak-Kei Lam, Wai-Chung Tang, Yu-Liang Wu
DATE3
2012 WRIP: logic restructuring techniques for wirelength-driven incremental placement
abstract
This paper presents WRIP - a Wirelength-driven Rewiring-based Incremental Placement which effectively reduces wirelength of the optimized placement of industrial large-scale standard cell designs. WRIP uses a powerful logic synthesis technique called logic rewiring which restructures the local circuits while preserving the logic functionality and reduces the wirelength under an accurate estimation of the half perimeter wirelength (HPWL) metric. We integrated WRIP into an industrial EDA tool and tested it upon several real designs with hundreds of thousands of movable objects. Tested on circuits which has been fully optimized by the state-of-the-art industrial placement tool, our experiments showed that on average WRIP reduces wirelength by 2.25% after placement and 2.45% after global routing in HPWL and Steiner WL model respectively. The runtime of WRIP is only about half an hour for the largest tested ASIC circuit. This is the first attempt to fully integrate powerful logic synthesis into industrial placement tools with real-life effectiveness and efficiency.
Wai-Chung Tang, Yu-Liang Wu, Cliff C. N. Sze, Charles J. Alpert
ACM Great Lakes Symposium on VLSI2
2012 Postgrid Clock Routing for High Performance Microprocessor Designs
abstract
Designing a high-quality clock network is very important in very large-scale integrated designs today, as it is the clock network that synchronizes all the elements of a chip, and it is also a major source of power dissipation of a system. Early study by Pham in 2006 shows that about 18.1% of the total clock capacitance was due to this postgrid clock routing (i.e., lower mesh wires plus clock twig wires). In this paper, we proposed a partition-based path expansion algorithm to solve this postgrid clock routing problem effectively. Experimental results on industrial test cases show that our algorithm can improve over the latest work by Shelar on this problem significantly by reducing the wire capacitance by 24.6% and the wirelength by 23.6%.
Haitong Tian, Wai-Chung Tang, Evangeline F. Y. Young, Cliff C. N. Sze
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2012 ECR: A Powerful and Low-Complexity Error Cancellation Rewiring Scheme
abstract
Rewiring is known to be a class of logic restructuring technique that is at least equally powerful in flexibility compared to other logic transformation techniques. Especially it is wiring sensitive and is particularly useful for interconnect-based circuit synthesis processes. One of the most well-studied rewiring techniques is the ATPG-based Redundancy Addition and Removal (RAR) technique which adds a redundant alternative wire to make an originally irredundant target wire become redundant and thus removable. In this article, we propose a new Error-Cancellation-based Rewiring scheme (ECR) which can also identify non-RAR-based rewiring operations with high efficiency. In ECR scheme, it is not necessary for alternative wires to be redundant. Based on the notion of error cancellation, we analyze and reformulate the rewiring problem, and a more generalized rewiring scheme is developed to detect more rewiring cases which are not obtainable by existing schemes while it still maintains a low runtime complexity. Comparing with the most recent non-RAR rewiring tool IRRA, the total number of alternative wires found by our approach is about doubled (202%) while the CPU time used is just slightly more (8%) upon benchmarks preoptimized by ABC’s rewriting. Our experimental results also suggest that the ECR engine is more powerful than IRRA in FPGA technology mapping.
Tak-Kei Lam, Wai-Chung Tang, Yu-Liang Wu
ACM Trans. Design Autom. Electr. Syst.2
2011 On applying erroneous clock gating conditions to further cut down power
abstract
All of today's known clock gating techniques only disable clocks on valid (”correct”) clock gating conditions, like idle states or observability don't cares (ODC), whose applying will not change the circuit functionality. In this paper, we explore a technique that allows shutting down certain clocks during invalid cycles, which if applied alone will certainly cause erroneous results. However, the erroneous results will be corrected either during the current or later stages by injecting other clock gating conditions to cancel out each other's error effects before they reach the primary outputs. Under this model, conditions across multiple flip-flop stages can also be analyzed to locate easily correctable erroneous clock gating conditions. Experimental results show that by using this error cancellation technique, a total power (including dynamic and leakage power) cut of up to 23% and in average of around 6% could be stably achieved, no matter with or without applying Power Compiler (which brought a power cut of 4% in average) together. The results indicate that the power saving conditions found by this new technique were nearly orthogonal (independent) to what can be done by the popular commercial power optimization tool. The idea of these new multi-stage logic error cancellation operations can potentially be applied to other sequential logic synthesis problems as well.
Tak-Kei Lam, Wai-Chung Tang, Yu-Liang Wu
ASP-DAC3
2011 Grid-to-ports clock routing for high performance microprocessor designs
abstract
Clock distribution in VLSI designs is of crucial importance and it is also a major source of power dissipation of a system. For today's high performance microprocessors, clock signals are usually distributed by a global clock grid covering the whole chip, followed by post-grid routing that connects clock loads to the clock grid. Early study [2] shows that about 18.1% of the total clock capacitance dissipation was due to this post-grid clock routing (i.e., lower mesh wires plus clock twig wires). This post-grid clock routing problem is thus an important one but not many previous works have addressed it. In this paper, we try to solve this problem of connecting clock ports to the clock grid through reserved tracks on multiple metal layers, with delay and slew constraints. Note that a set of routing tracks are reserved for this grid-to-ports clock wires in practice because of the conventional modular design style of high-performance microprocessors. We propose a new expansion algorithm based on the heap data structure to solve the problem effectively. Experimental results on industrial test cases show that our algorithm can improve over the latest work on this problem [1] significantly by reducing the capacitance by 24.6% and the wire length by 23.6%. We also validate our results using hspice simulation. Finally, our approach is very efficient and for larger test cases with about 2000 ports, the runtime is in seconds.
Haitong Tian, Wai-Chung Tang, Evangeline F. Y. Young, Cliff C. N. Sze
ISPD2
2010 Logic synthesis for low power using clock gating and rewiring
abstract
Traditionally, clock gating for power saving is mainly done at Register Transistor Level (RTL), while in a lower logical level some synthesis techniques, e.g. Observability Don't Care (ODC) can also be used to provide more power savings. In this paper, we propose an effective logic level ODC-based clock gating scheme that aims to reduce the intra-module dynamic power of sequential circuits. It is accompanied with a rewiring-based pruning scheme to trim down the incurred area overhead. Switching activity is put into account in the optimization processes. Extensive experimental results obtained by using ModelSim and PowerCompiler on the ISCAS-89 benchmarks showed that without rewired area trimming, an average of 40% on clock power and 12% on total dynamic power can be saved with a total cell area overhead of 6%. When rewiring was applied to trim down the area overhead, a similar clock power saving of 40% and an appealing 17% of total dynamic power saving can be achieved with area similar (-1%) to the original non ODC-clock-gated circuits.
Tak-Kei Lam, Steve Yang, Wai-Chung Tang, Yu-Liang Wu
ACM Great Lakes Symposium on VLSI3
2007 How Much Can Logic Perturbation Help from Netlist to Final Routing for FPGAs
abstract
One unique property of an FPGA chip is that any logic perturbation inside its Look-Up-Tables (LUTs) is totally area/delay-free. Amongst others, this free LUT-internal resource perturbation can also be used to trade for critical LUT-external logic/wire removals for EDA improvements, an extra flexibility ignored before. Using rewiring technique for such logic perturbations, we show that significant cut-downs upon already excellent results from the state-of-the-art DAOmap mappings and the TVPR place-and-route can still be obtained. This logic perturbation operation can further reduce the number of LUTs by up to 33.7% (avg. 10%) without delay penalty and also reduce critical path delay by up to 31.7% (avg. 11%) without disturbing placement or sacrificing area in the final routing. For delay reduction, under proper rewiring strategy, the CPU time used by rewiring is only 5% of the total run time consumed by TVPR's placement and routing. This idea of perturbing logic between the free LUT-internal and critical LUT-external circuit resources is simple and proved to be powerful. The encouraging results suggest a new technique for an optimization domain less explored for FPGA design flow.
Catherine L. Zhou, Wai-Chung Tang, Wing-Hang Lo, Yu-Liang Wu
DAC2
2007 Further Improve Excellent Graph-Based FPGA Technology Mapping by Rewiring
abstract
FPGA technology mapping is conventionally solved without altering the circuit by modeling the circuit as a direct acyclic graph for the ease of applying graph algorithms. Clearly there is room for further improvement on even optimal technology mapping results if logic perturbation can be applied. In this paper, we propose logic-aware minimization methods to further reduce both depth and area for the purely-graph-based depth-optimal FPGA mapped results. For area minimization, the proposed method perturbs the subject circuit using rewiring technique and incrementally reduce the mapping area. Improving the outstanding technology mapping algorithm DAOMap, the method can further reduce area by 10.9%. An area reduction of 13.4% is achieved with synthesis results from BDS-pga. A logic level reduction scheme is also proposed and it further reduces LUT depth for half of the circuits tested without area penalty. A combination of logic level reduction and area minimization techniques can improve both the LUT depth and area by 11.3% and 6.1%, compared to results of FlowMap and FlowSYN respectively.
Wai-Chung Tang, Wing-Hang Lo, Yu-Liang Wu
ISCAS1