EDBT 2026 Demo / reviewers in the wild / expert
Wuxi Li
dblp:149/4660
· DBLP profile ↗
28ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-9887-5109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 8 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LegoMap: Optimization for High-Throughput Transformer Computing on AI Engine-Based FPGAs
Hailiang Hu, Haodong Chang, Donghao Fang, Zhenrui Wang, Wuxi Li, Rongjian Liang, Bo Yuan 0001, Jiang Hu 0001 |
FCCM | 5 |
| 2026 | A New Approach to Performance-Driven Analog IC PlacementabstractA major obstacle in analog design automation is that circuit performance is sensitive to layout, yet accurately capturing this impact within layout tools is very expensive. To address this challenge, we propose a performance-driven analog IC placement approach, called VPlace, guided by machine learning. Our approach leverages a novel application of the VQ-VAE technique to improve robustness during the placement stage, in conjunction with a recent machine learning-based macromodeling method. We further demonstrate that data preparation strategies, which directly affect the efficiency of investigating the solution space, play an important role in determining both the accuracy of machine learning models and the resulting circuit performance. Experimental results show that VPlace achieves 22%-26% and 10%-16% performance improvements over an open-source analog layout tool and a prior machine learning–based performance-driven analog placement technique, respectively. Donghao Fang, Hailiang Hu, Wuxi Li, Jiang Hu 0001 |
ISPD | 3 |
| 2025 | Global Placement Exploiting Soft 2D RegularityabstractCell placement is a step of paramount importance in chip physical design and requests relentless effort for continuous improvement. Recently, designs with two-dimensional (2D) processing element arrays have become popular primarily due to their deep neural network hardware applications. The 2D array regularity is similar to but different from the regularity of conventional datapath designs. To exploit the 2D array regularity, this work develops a new global placement technique, Placement of Arrays with SOft Regularity (PASOR), built upon RePlAce, the state-of-the-art placement framework. Experimental results from various designs show that the proposed approach can reduce global routing wirelength by 11% and 6% compared to RePlAce and a previous work on datapath driven placement, respectively. Donghao Fang, Boyang Zhang 0007, Hailiang Hu, Wuxi Li, Bo Yuan 0001, Jiang Hu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2024 | SysMix: Mixed-Size Placement for Systolic-Array-Based Hierarchical DesignsabstractSystolic array designs are gaining popularity due to their applications in hardware acceleration for ML computing, such as CNNs and transformers. Increasingly large ML models necessitate very high circuit energy-efficiency, which is highly correlated with minimizing placement wirelength in chip physical design. However, existing placement techniques are mostly general purpose and overlook unique properties of systolic array designs. We propose a mixed-size placement approach, called SysMix, which is tailored for systolic arrays and leverage their partial regularity in hierarchical design methodologies. Experimental results from multiple CNN designs show that SysMix achieves 53% wirelength reduction and 15X speedup compared to a commercial placer and a state-of-the-art academic placer. Donghao Fang, Hailiang Hu, Wuxi Li, Bo Yuan 0001, Jiang Hu 0001 |
ICCAD | 3 |
| 2024 | HeteroExcept: A CPU-GPU Heterogeneous Algorithm to Accelerate Exception-aware Static Timing AnalysisabstractStatic timing analysis (STA) for large-scale modern circuits requires extensive handling of false paths, multi-cycle paths, and other types of path exceptions. Despite the linear nature of timing propagation, we show that exception-aware STA is NP-hard and thus requires a long runtime to solve using conventional CPU-based methods. To overcome this runtime challenge, we propose a general CPU-GPU heterogeneous algorithm, HeteroExcept, that can handle common types of path exceptions and efficiently generate an accurate path report. Our algorithm targets runtime efficiency at the scale of thousands of exception rules and millions of circuit elements. To further improve the performance, we optimize our GPU implementation by introducing a cost-effective data exchange strategy between CPU and GPU. Experimental results demonstrate up to 6.84× and 12.93× speed-up compared to industrial timers, PrimeTime and OpenSTA. Zizheng Guo 0001, Zuodong Zhang, Wuxi Li, Tsung-Wei Huang, Xizhe Shi, Yufan Du, Yibo Lin, Runsheng Wang, Ru Huang 0001 |
ICCAD | 3 |
| 2024 | Calibration-Based Differentiable Timing Optimization in Non-linear Global PlacementabstractPlacement plays a crucial role in the timing closure of integrated circuit (IC) physical design. This paper presents an efficient and effective calibration-based differentiable timing-driven global placement engine. Our key innovation is a calibration technique that approximates a precise but expensive reference timer, such as a signoff timer, using a lightweight simple timer. This calibrated simple timer inherently accounts for intricate timing exceptions and common path pessimism removal (CPPR) prevalent in industry designs. Extending this calibrated simple timer into a differentiable timing engine enables ultrafast yet accurate timing optimization in non-linear global placement. Experimental results on various industry designs demonstrate the superiority of the proposed framework over the latest AMD Vivado and traditional net-weighting methods across key metrics including maximum clock frequency, wirelength, routability, and overall back-end runtime. Wuxi Li, Yuji Kukimoto, Grégory Servel, Ismail Bustany, Mehrdad E. Dehkordi |
ISPD | 1 |
| 2023 | General-Purpose Gate-Level Simulation with Partition-Agnostic ParallelismabstractGate-level simulation with delay annotation is a both critical and time-consuming task in the circuit design flow. It is highly nontrivial to parallelize a simulation process, especially on designs with arbitrary general-purpose sequential elements such as latches, gated clocks, and scan chains. Current works on parallelizing gate-level simulation are fundamentally incompatible with these design elements and are highly reliant on circuit partitioning to achieve the best performance. In this paper, we propose a general-purpose gate-level simulation engine with partition-agnostic parallelism. We propose a general sequential behavior encoding technique and a fast event scheduling algorithm for general-purpose simulation tasks. Experimental results have shown up to 30× speed-up over commercial simulation engines. Zizheng Guo 0001, Zuodong Zhang, Xun Jiang 0002, Wuxi Li, Yibo Lin, Runsheng Wang, Ru Huang 0001 |
DAC | 4 |
| 2023 | Systolic Array Placement on FPGAsabstractSystolic array designs have regained popularity in recent years, particularly for their applications in accelerating CNN (Convolutional Neural Network) computing in hardware, including on FPGAs. However, existing FPGA layout techniques are primarily designed for general-purpose applications and have not fully leveraged the regularity of systolic arrays to enhance solution quality. This paper presents a new algorithmic approach for systolic array placement on FPGAs. Our approach enables 23% – 25% wirelength reduction for CNN circuits compared to an industrial tool and state-of-the-art academic methods. Moreover, it usually leads to significantly reduced routing resource utilization, accelerated placement runtime and improved timing performance. Hailiang Hu, Donghao Fang, Wuxi Li, Bo Yuan 0001, Jiang Hu 0001 |
ICCAD | 3 |
| 2023 | RapidStream 2.0: Automated Parallel Implementation of Latency-Insensitive FPGA Designs Through Partial ReconfigurationabstractField-programmable gate arrays (FPGAs) require a much longer compilation cycle than conventional computing platforms such as CPUs. In this article, we shorten the overall compilation time by co-optimizing the HLS compilation (C-to-RTL) and the back-end physical implementation (RTL-to-bitstream). We propose a split compilation approach based on the pipelining flexibility at the HLS level, which allows us to partition designs for parallel placement and routing. We outline a number of technical challenges and address them by breaking the conventional boundaries between different stages of the traditional FPGA tool flow and reorganizing them to achieve a fast end-to-end compilation. Our research produces RapidStream, a parallelized and physical-integrated compilation framework that takes in a latency-insensitive program in C/C++ and generates a fully placed and routed implementation. We present two approaches. The first approach (RapidStream 1.0) resolves inter-partition routing conflicts at the end when separate partitions are stitched together. When tested on the Xilinx U250 FPGA with a set of realistic HLS designs, RapidStream achieves a 5 to 7× reduction in compile time and up to 1.3× increase in frequency when compared with a commercial off-the-shelf toolchain. In addition, we provide preliminary results using a customized open-source router to reduce the compile time up to an order of magnitude in cases with lower performance requirements. The second approach (RapidStream 2.0) prevents routing conflicts using virtual pins. Testing on Xilinx U280 FPGA, we observed 5 to 7× compile time reduction and 1.3× frequency increase. Licheng Guo, Pongstorn Maidee, Chris Lavin, Eddie Hung, Wuxi Li, Jason Lau, Weikang Qiao, Yuze Chi, Linghao Song, Yuanlong Xiao, Alireza Kaviani, Zhiru Zhang, Jason Cong |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2022 | Global Placement Exploiting Soft 2D RegularityabstractCell placement is such a critical step for chip physical design that it needs many kinds of efforts for improvement. Recently, designs with 2D processing element arrays have become popular primarily due to their deep neural network computing applications. The 2D array regularity is similar to but different from the regularity of conventional datapath designs. To exploit the 2D array regularity, this work develops a new global placement technique built upon RePlAce, the latest state-of-the-art placement framework. Experimental results from various designs show that the proposed technique can reduce half-perimeter wirelength and Steiner tree wirelength by about $6%$ and $12%$, respectively. Donghao Fang, Boyang Zhang 0007, Hailiang Hu, Wuxi Li, Bo Yuan 0001, Jiang Hu 0001 |
ISPD | 4 |
| 2022 | elfPlace: Electrostatics-Based Placement for Large-Scale Heterogeneous FPGAsabstractelfPlaceis a flat nonlinear placement algorithm for large-scale heterogeneous field-programmable gate arrays (FPGAs). We adopt the analogy between placement and electrostatic systems initially proposed byePlaceand extend it to tackle heterogeneous blocks in FPGA designs. To achieve satisfiable solution quality with fast and robust numerical convergence, an augmented Lagrangian formulation together with a preconditioning technique and a normalized subgradient-based multiplier updating scheme are proposed. Besides pure-wirelength minimization, we also propose a unified instance area adjustment scheme to simultaneously optimize routability, pin density, and downstream clustering compatibility. We further propose run-to-run deterministic GPU acceleration techniques to speedup the global placement. Our experiments on the ISPD 2016 benchmark suite show thatelfPlaceoutperforms four state-of-the-art FPGA placersUTPlaceF,RippleFPGA,GPlace3.0, andUTPlaceF-DLby 13.5%, 10.2%, 8.8%, and 7.0%, respectively, in routed wirelength with competitive runtime. Yibai Meng, Wuxi Li, Yibo Lin, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | DREAMPlace: Deep Learning Toolkit-Enabled GPU Acceleration for Modern VLSI PlacementabstractPlacement for very large-scale integrated (VLSI) circuits is one of the most important steps for design closure. We propose a novel GPU-accelerated placement framework DREAMPlace, by casting the analytical placement problem equivalently to training a neural network. Implemented on top of a widely adopted deep learning toolkit PyTorch, with customized key kernels for wirelength and density computations, DREAMPlace can achieve around 40× speedup in global placement without quality degradation compared to the state-of-the-art multithreaded placer RePlAce. We believe this work shall open up new directions for revisiting classical EDA problems with advancements in AI hardware and software. Yibo Lin, Zixuan Jiang, Jiaqi Gu 0002, Wuxi Li, Shounak Dhar, Haoxing Ren, Brucek Khailany, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | High-Definition Routing Congestion Prediction for Large-Scale FPGAsabstractTo speed up the FPGA placement and routing closure, we propose a novel approach to predict the routing congestion map for large-scale FPGA designs at the placement stage. After reformulating the problem into an image translation task, our proposed approach leverages recent advancement in generative adversarial learning to address the task. Particularly, state-of-the-art generative adversarial networks for high-resolution image translation are used along with well-engineered features extracted from the placement stage. Unlike available approaches, our novel framework demonstrates a capability of handling large-scale FPGA designs. With its superior accuracy, our proposed approach can be incorporated into the placement engine to provide congestion prediction resulting in up to 7% reduction in routed wirelength for the most congested design in ISPD 2016 benchmark. Mohamed Baker Alawieh, Wuxi Li, Yibo Lin, Love Singhal, Mahesh A. Iyer, David Z. Pan |
ASP-DAC | 2 |
| 2020 | S3DET: Detecting System Symmetry Constraints for Analog Circuits with Graph SimilarityabstractSymmetry and matching between critical building blocks have a significant impact on analog system performance. However, there is limited research on generating system level symmetry constraints. In this paper, we propose a novel method of detecting system symmetry constraints for analog circuits with graph similarity. Leveraging spectral graph analysis and graph centrality, the proposed algorithm can be applied to circuits and systems of large scale and different architectures. To the best of our knowledge, this is the first work in detecting system level symmetry constraints for analog and mixed-signal (AMS) circuits. Experimental results show that the proposed method can achieve high accuracy of 88.3% with low false alarm rate of less than 1.1% in largescale AMS designs. Wuxi Li, Keren Zhu 0001, Biying Xu, Yibo Lin, Linxiao Shen, Xiyuan Tang, Nan Sun 0001, David Z. Pan |
ASP-DAC | 2 |
| 2020 | FLOPS: EFficient On-Chip Learning for OPtical Neural Networks Through Stochastic Zeroth-Order OptimizationabstractOptical neural networks (ONNs) have attracted extensive attention due to its ultra-high execution speed and low energy consumption. The traditional software-based ONN training, however, suffers the problems of expensive hardware mapping and inaccurate variation modeling while the current on-chip training methods fail to leverage the self-learning capability of ONNs due to algorithmic inefficiency and poor variation- robustness. In this work, we propose an on-chip learning method to resolve the aforementioned problems that impede ONNs' full potential for ultra-fast forward acceleration. We directly optimize optical components using stochastic zeroth-order optimization on-chip, avoiding the traditional high-overhead back-propagation, matrix decomposition, or in situ devicelevel intensity measurements. Experimental results demonstrate that the proposed on-chip learning framework provides an efficient solution to train integrated ONNs with 3~4× fewer ONN forward, higher inference accuracy, and better variation-robustness than previous works. Jiaqi Gu 0002, Zheng Zhao 0003, Chenghao Feng, Wuxi Li, Ray T. Chen, David Z. Pan |
DAC | 4 |
| 2020 | ABCDPlace: Accelerated Batch-Based Concurrent Detailed Placement on Multithreaded CPUs and GPUsabstractPlacement is an important step in modern verylarge-scale integrated (VLSI) designs. Detailed placement is a placement refining procedure intensively called throughout the design flow, thus its efficiency has a vital impact on design closure. However, since most detailed placement techniques are inherently greedy and sequential, they are generally difficult to parallelize. In this article, we present a concurrent detailed placement framework, ABCDPlace, exploiting multithreading and graphic processing unit (GPU) acceleration. We propose batch-based concurrent algorithms for widely adopted sequential detailed placement techniques, such as independent set matching, global swap, and local reordering. The experimental results demonstrate that ABCDPlace can achieve 2× -5× faster runtime than sequential implementations with multithreaded CPU and over 10× with GPU on ISPD 2005 contest benchmarks without quality degradation. On larger industrial benchmarks, we show more than 16× speedup with GPU over the state-of-the-art sequential detailed placer. ABCDPlace finishes the detailed placement of a 10-million-cell industrial design in 1 min. Yibo Lin, Wuxi Li, Jiaqi Gu 0002, Haoxing Ren, Brucek Khailany, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | DREAMPlace: Deep Learning Toolkit-Enabled GPU Acceleration for Modern VLSI PlacementabstractPlacement for very-large-scale integrated (VLSI) circuits is one of the most important steps for design closure. This paper proposes a novel GPU-accelerated placement framework DREAMPlace, by casting the analytical placement problem equivalently to training a neural network. Implemented on top of a widely-adopted deep learning toolkit PyTorch, with customized key kernels for wirelength and density computations, DREAMPlace can achieve over 30× speedup in global placement without quality degradation compared to the state-of-the-art multi-threaded placer RePlAce. We believe this work shall open up new directions for revisiting classical EDA problems with advancement in AI hardware and software. Yibo Lin, Shounak Dhar, Wuxi Li, Haoxing Ren, Brucek Khailany, David Z. Pan |
DAC | 3 |
| 2019 | Simultaneous Placement and Clock Tree Construction for Modern FPGAsabstractModern field-programmable gate array (FPGA) devices often contain complex clocking architectures to achieve high-performance and flexible clock networks. The physical structure of these clock networks, however, are pre-manufactured, unadjustable, and with only limited routing resources. Most conventional FPGA placement algorithms rarely consider clock feasibility, and therefore lead to clock routing failures. Some recent works adopt simplified clock routing models (e.g., the bounding box model) to force clock legality during placement, which, however, can often overestimate clock routing demands and results in unnecessary placement quality degradation. To address these limitations, in this paper, we propose a generic FPGA placement framework that can simultaneously optimize placement quality and ensure clock feasibility by explicit clock tree construction. We demonstrate the effectiveness and efficiency of the proposed approach using the ISPD 2017 Clock-Aware Placement Contest benchmark suite. Compared with other state-of-the-art clock legalization algorithms, the proposed approach can achieve the best routed wirelength with competitive runtime. Wuxi Li, Mehrdad E. Dehkordi, David Z. Pan |
FPGA | 1 |
| 2019 | elfPlace: Electrostatics-based Placement for Large-Scale Heterogeneous FPGAsabstractelfplace is a flat nonlinear placement algorithm for large-scale heterogeneous field-programmable gate arrays (FPGAs). We adopt the analogy between placement and electrostatic systems initially proposed by ePlace and extend it to tackle heterogeneous blocks in FPGA designs. To achieve satisfiable solution quality with fast and robust numerical convergence, an augmented Lagrangian formulation together with a preconditioning technique and a normalized subgradient-based multiplier updating scheme are proposed. Besides pure-wirelength minimization, we also propose a unified instance area adjustment scheme to simultaneously optimize routability, pin density, and downstream clustering compatibility. Our experiments on ISPD 2016 benchmark suite show that elfPlace outperforms four state-of-the-art FPGA placers UTPlaceF, RippleFPGA, GPlace3.0, and UTPlaceF-DL by 13.6%, 11.3%, 8.9%, and 7.1%, respectively, in routed wirelength with competitive runtime. Wuxi Li, Yibo Lin, David Z. Pan |
ICCAD | 1 |
| 2019 | A New Paradigm for FPGA Placement Without Explicit PackingabstractPlacement and packing are two important but separated optimization steps in a conventional field programmable gate array (FPGA) implementation flow. A packing engine clusters logic elements, like lookup tables and flip-flops, into configurable logic blocks, while a placement engine determines their physical locations in FPGA layouts. This paper presents a new paradigm for FPGA placement without an explicit packing stage. In the proposed framework, the solution spaces of placement and packing are simultaneously explored in a smooth and elegant way. Our experiments on ISPD 2016 and 2017 benchmark suites demonstrate the effectiveness of the proposed framework. Wuxi Li, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | A Practical Split Manufacturing Framework for Trojan Prevention via Simultaneous Wire Lifting and Cell InsertionabstractTrojans and backdoors inserted by untrusted foundries have become serious threats to hardware security. Split manufacturing is proposed to hide important circuit structures and prevent Trojan insertion by fabricating partial interconnections in trusted foundries. Existing split manufacturing frameworks, however, usually lack security guarantee and suffer from poor scalability. It is observed that inserting dummy cells and wires can have high potential on overcoming the security and scalability problems of existing methods, but it is not compatible with current security definition. In this paper, we focus on answering the questions on how to define the notion of security and how to realize the required security level effectively and efficiently when the insertion of dummy cells and wires is considered. We first generalize existing security criterion by modeling the split manufacturing process as a graph problem. Then, a sufficient condition is derived for the proposed security criterion to avoid the computationally intensive operations in traditional methods. To further enhance the scalability of the framework, we propose a secure-by-construction split manufacturing flow. For the first time, a novel mixed-integer linear programming (MILP) formulation is proposed to simultaneously consider cell and wire insertion together with wire lifting. A Lagrangian relaxation algorithm with a minimum-cost flow transformation technique is employed to solve the MILP formulation efficiently. With extensive experiments, our framework demonstrates significantly better efficiency, overhead reduction and security guarantee compared with the previous state-of-the-art. Meng Li 0004, Bei Yu 0001, Yibo Lin, Wuxi Li, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | A practical split manufacturing framework for Trojan prevention via simultaneous wire lifting and cell insertionabstractTrojans and backdoors inserted by untrusted foundries have become serious threats to hardware security. Split manufacturing is proposed to prevent Trojan insertion proactively. Existing methods depend on wire lifting to hide partial circuit interconnections, which usually suffer from large overhead and lack of security guarantee. In this paper, we propose a novel split manufacturing framework that not only guarantees to achieve the required security level but also allows for a drastic reduction of the introduced overhead. In our framework, insertion of dummy circuit cells and wires is considered simultaneously with wire lifting. To support cell and wire insertion, we propose a new security criterion, and further derive its sufficient condition to avoid computation intensive operations in traditional methods. Then, for the first time, a novel mixed integer linear programming formulation is proposed to simultaneously consider cell and wire insertion together with wire lifting, which significantly enlarges the design space to guarantee the realization of the sufficient condition under the security requirements and overhead constraints. With extensive experimental results, our framework demonstrates much better efficiency, overhead reduction, and security guarantee compared with existing methods. Meng Li 0004, Bei Yu 0001, Yibo Lin, Wuxi Li, David Z. Pan |
ASP-DAC | 5 |
| 2018 | UTPlaceF: A Routability-Driven FPGA Placer With Physical and Congestion Aware PackingabstractField programmable gate array (FPGA) packing and placement without routability consideration could lead to unroutable results for high-utilization designs. Conventional FPGA packing and placement approaches are shown to have severe difficulties to yield good routability. In this paper, we propose an FPGA packing and placement engine called UTPlaceF that simultaneously optimizes wirelength and routability. A novel physical and congestion aware packing algorithm and a hierarchical detailed placement technique are proposed. UTPlaceF outperforms state-of-the-art FPGA placers simultaneously in runtime and solution quality on International Symposium on Physical Design (ISPD) 2016 benchmark suite. Compared with the top three winners of ISPD'16 FPGA placement contest, UTPlaceF can deliver 6.2%, 11.6%, and 29.1% better routed wirelength with shorter runtime. Wuxi Li, Shounak Dhar, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | UTPlaceF 2.0: A High-Performance Clock-Aware FPGA Placement EngineabstractModern field-programmable gate array (FPGA) devices contain complex clock architectures on top of configurable logics. Unlike application specific integrated circuits (ASICs), the physical structure of clock networks in an FPGA is pre-manufactured and cannot be adjusted to different applications. Furthermore, clock routing resources are typically limited for high-utilization designs. Consequently, clock architectures impose extra clock constraints and further complicate physical implementation tasks such as placement. Traditional ASIC placement techniques only optimize conventional design metrics such as wirelength, routability, power, and timing without clock legality consideration. It is imperative to have new techniques to honor clock constraints during placement for FPGAs. In this article, we propose a high-performance FPGA placement engine, UTPlaceF 2.0, that optimizes wirelength and routability while honoring complex clock constraints. Our proposed approaches consist of an iterative minimum-cost-flow-based cell assignment as well as a clock-aware packing for producing clock-legal yet high-quality placement solutions. UTPlaceF 2.0 won first place in the ISPD’17 clock-aware FPGA placement contest organized by Xilinx, outperforming the second- and the third-place winners by 4.0% and 10.0%, respectively, in routed wirelength with competitive runtime, on a set of industry benchmarks. Wuxi Li, Yibo Lin, Meng Li 0004, Shounak Dhar, David Z. Pan |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2017 | UTPlaceF 3.0: A parallelization framework for modern FPGA global placement: (Invited paper)abstractGlobal placement is a major runtime bottleneck of modern FPGA physical synthesis. As the FPGA capacity grows rapidly, new innovative global placement approaches are in great demand for more efficient circuit mapping and prototyping. In this paper, we propose a parallelization framework for modern FPGA global placement, UTPlaceF 3.0. Two major techniques are presented to boost the performance of a state-of-the-art quadratic placer with only small quality degradation: 1) placement-driven block-Jacobi preconditioning and 2) parallelized incremental placement correction. Experimental results show that UTPlaceF 3.0 can take full advantages of modern multi-core CPUs and achieves more than 5X speedup over sequential implementation with competitive placement quality. Wuxi Li, Meng Li 0004, David Z. Pan |
ICCAD | 1 |
| 2017 | Placement mitigation techniques for power grid electromigrationabstractIn advanced technology nodes, power grid metal wires are prone to electromigration (EM) failures due to small wire sizes and high unidirectional current densities. Power grid EM failures usually happen around weak power grid connections delivering current to high power-consuming regions. Previously, power grid EM was mostly addressed at the post-routing stage, which may be too late for a large number of EM violations in modern designs. In this paper, we propose a new set of incremental placement techniques to mitigate power grid EM, including cell move, single row placement, and single tile placement. Experimental results demonstrate the proposed placement techniques can effectively reduce EM violations with negligible wirelength and placement density impacts. Wei Ye 0008, Yibo Lin, Wuxi Li, Yiwei Fu, Yongsheng Sun, Canhui Zhan, David Z. Pan |
ISLPED | 4 |
| 2016 | UTPlaceF: a routability-driven FPGA placer with physical and congestion aware packingabstractFPGA packing and placement without routability consideration could lead to unroutable results for high-utilization designs. Conventional FPGA packing and placement approaches are shown to have severe difficulties to yield good routability. In this paper, we propose a FPGA packing and placement engine called UTPlaceF that simultaneously optimizes wirelength and routablity. A novel physical and congestion aware packing alogrithm and several congestion aware detailed placement techniques are proposed. Compared with the top 3 winners of ISPD'16 FPGA placement contest, UTPlaceF can achieve 3.3%, 7.7% and 28.3% better routed wirelength with similar or shorter runtime. Wuxi Li, Shounak Dhar, David Z. Pan |
ICCAD | 1 |
| 2014 | Live demonstration: An optimization software and a design case of a novel dual band wireless power and data transmission systemabstractBiomedical implanted electronic devices utilize inductively coupled coils for power and data transmission. To achieve efficient power transmission and high data rate, a dual band telemetry system has been proposed where different carrier frequencies can be used for power and data transmission, respectively. In this system, the physical parameters and geometrical structure of the coils must be carefully optimized to avoid the strong power interference to data transmission, which diminishes the advantages of multiple carrier frequencies. In this demonstration, an optimization procedure is introduced about the optimization of two pairs of coils based on an overlapping structure which minimizes power-to-signal interference in the data transmission. We applied the optimization procedure to a practical design case, and the results showed that our optimization procedure can achieve both optimal power transmission efficiency and high signal to interference ratio. Xiyan Li, Wuxi Li, Guoxing Wang |
ISCAS | 3 |