EDBT 2026 Demo / reviewers in the wild / expert
Zhili Xiong
dblp:242/5718
· DBLP profile ↗
6ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiffRouter: A Differentiable Routing Framework for UltraScale FPGAsabstractRouting has become a major bottleneck in large-scale FPGA design flows, where traditional routers rely on sequential net-by-net processing and rip-up-reroute heuristics. This inherent sequentialism limits parallelism, leads to suboptimal utilization of routing resources, and significantly increases runtime on modern FPGAs. This paper introduces DiffRouter, the first differentiable, gradient-based FPGA routing framework with GPU acceleration. By lifting routing into a continuous optimization space, DiffRouter enables highly concurrent routing and improved global optimality before projecting solutions back onto the discrete fabric. The framework consists of three key stages: (1) Resource-pruned preprocessing, which extracts compact routing variables per net to reduce memory footprint and problem complexity; (2) Differentiable global routing via gradient descent, which formulates routing as a Lagrangian-relaxed optimization that minimizes wirelength while iteratively enforcing connectivity and congestion constraints as relaxed terms; and (3) Postprocessing, which maps the continuous wire distributions to a valid UltraScale routing solution and completes solutions by detailed routing. Experiments on the FPGA 2024 Routing Contest benchmarks show that DiffRouter outperforms state-of-the-art parallel routers, achieving over 5× speedup over RWRoute and 2× over Vivado, and running 18% faster than the state-of-the-art academic FPGA router Potter while delivering competitive critical path wirelength. Xiaohan Gao, Zhili Xiong, David Z. Pan |
FCCM | 2 |
| 2026 | GrandPlan: Differentiable, Simultaneous Top-Level Floorplanning and Partition-Level Cell Placement for Large-Scale IP-CoresabstractTop-level floorplanning is a critical step in industrial physical design, where the die is partitioned into exactly abutted regions with carefully allocated areas to enable efficient hierarchical place-and-route and achieve desired power, performance, and area (PPA) trade-offs. In current practice, however, floorplanning remains largely manual and sub-optimal, as designers rely on RTL hierarchy with limited physical guidance; commercial tools cannot feasibly perform flat optimization at IP-core scale. As a result, late-stage routability-driven partition resizing often triggers cascading boundary changes, disrupting neighboring partitions and significantly increasing turnaround time and engineering cost. To address this challenge, we present GrandPlan, a GPU-accelerated, differentiable, end-to-end framework that co-optimizes top-level floorplanning and partition-level cell placement within a single automated loop. Leveraging custom CUDA kernels, GrandPlan generates clean, rectilinear partition boundaries while concurrently placing macros and standard cells. The framework consists of three tightly coupled stages: (1) flat IP-core placement with differentiable grouping objectives, (2) boundary refinement via simulated annealing under area and routability constraints, and (3) routability-aware fence-region placement. Experiments on eight large-scale industrial IP-cores (up to 25M cells) show that GrandPlan reduces total wirelength by up to 14% and cross-partition (feedthrough) wirelength by 27% on average compared to human-expert-crafted baselines, with an average runtime of only 1.2 hours. Zhili Xiong, Yi-Chen Lu, David Z. Pan, Haoxing Ren |
ISPD | 1 |
| 2025 | Differentiable Timing-Driven FPGA Placement with Smooth Optimization and ML-Based Delay CalibrationabstractPlacement is a critical stage in FPGA physical design, determining instance locations within available device resources. FPGA timing-driven placement aims to minimize routed wirelength while optimizing timing metrics such as total negative slack (TNS) and worst negative slack (WNS). In this work, we propose a differentiable timing-driven FPGA placement framework, enabling effective optimization convergence. We introduce a novel timing preconditioning method that smooths the timing objective landscape, along with a general multiplier scheduling scheme to effectively balance multiple objectives, including wirelength, WNS, and TNS. To further improve timing estimation during placement, we develop an XGBoost-based net delay prediction model to calibrate the timing model. Experiments on the ISPD 2016 contest benchmarks demonstrate that our placer achieves a 4% reduction in critical path delay, a 19% improvement in half-perimeter wirelength, and maintains a similar total place-and-route runtime (0.99 ×) relative to a state-of-the-art GPU-accelerated timing-driven FPGA placer. Yu-Kang Lin, Zhili Xiong, David Z. Pan |
ICCAD | 2 |
| 2024 | A Data-Driven, Congestion-Aware and Open-Source Timing-Driven FPGA Placer Accelerated by GPUsabstractPlacement plays a pivotal role in the modern FPGA physical design flow to determine the locations of the design instances among the available FPGA device resources, impacting routability and performance. Due to the lack of open-source accurate timing models for high-performance FPGAs, academic placement research has focused primarily on wirelength opti- mization rather than timing optimizations. This work presents an open-source timing-driven FPGA placer accelerated on GPU that employs a congestion-aware and data-driven timing model with timing optimizations at global placement and legalization. The placement objective incorporates an additional term to optimize the timing arcs in the lagrangian formulation. While packing and legalizing look-up tables (LUTs) and flip-flops (FFs), we emphasize timing-critical nets to remain within the Slice, minimizing overall path delay. On the ISPD'2016 contest benchmarks employing an AMD-Xilinx UltraScale architecture, our placer is 3× faster than the commercial AMD Vivado with similar critical path delay (×1.02) and 40% faster routing runtime. Zhili Xiong, Rachel Selina Rajarathnam, David Z. Pan |
FCCM | 1 |
| 2020 | Cross-layer congestion control of wireless sensor networks based on fuzzy sliding mode control
Shaocheng Qu, Liang Zhao 0014, Zhili Xiong |
Neural Comput. Appl. | 3 |
| 2019 | Differential Privacy with Variant-Noise for Gaussian Processes Classification
Zhili Xiong, Longyuan Li, Junchi Yan, Hao He 0007, Yaohui Jin |
PRICAI (3) | 1 |