VLDB 2026 Research / reviewers in the wild / expert
Jing Mai
dblp:293/1199
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-3787-1534ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TDM Signal Grouping and Package Pin Assignment for 2.5D Multi-FPGA Systems with Lookahead PlacementabstractLarge-scale multi-FPGA systems are widely used in modern emulation systems. As a critical part of the multi-FPGA system design flow, TDM signal grouping and package pin assignment directly impact the final placement and routing in the FPGA physical implementation. Poor pin assignments cause severe congestion and timing degradation at the logic-element level, while existing approaches lack accurate congestion modeling during system-level partitioning. This paper presents Chimew, a novel pin assignment methodology that leverages placement prototyping to predict logic-element-level congestion before physical implementation precisely. The proposed method co-optimizes signal grouping and pin placement through iterative refinement guided by congestion-aware cost functions derived from fast global placement. Experimental results demonstrate a 28% congestion reduction and up to 2.87ns less worst negative slack (WNS) compared to industrial tools while achieving a 100% success rate across diverse multi-FPGA benchmarks. Runzhe Tao, Jing Mai, Xun Jiang 0002, Cuiliu Yang, Haoyu Jie, Kan Huang, Richard Y. Sun, Yibo Lin |
FPGA | 3 |
| 2026 | LEGALM 2.0: A Versatile Augmented Lagrangian Method-Based Methodology for Mixed-Cell-Height LegalizationabstractLegalization is a crucial step in VLSI physical design, ensuring design rule compliance while minimizing disruptions to global placement. With the rise of multi-row-height cells in advanced nodes, mixed-cell-height legalization poses significant challenges due to complex cell shapes and design constraints. In this work, we present LEGALM 2.0, a versatile legalization methodology that efficiently handles routability and hybrid region constraints. Our approach introduces a linearized augmented Lagrangian formulation, a scanline-based initial legalization algorithm, and a connectivity-based local optimum escape strategy to enhance convergence. Additionally, we propose a block gradient descent method and a GPU-optimized triplefold partitioning strategy for improved parallelism. Experimental results show that LEGALM 2.0 outperforms state-of-the-art legalizers, achieving 6-36% better quality scores on ICCAD2017 benchmarks and 1.61-4.30W speedup on large-scale designs. For hybrid region constraints, it reduces displacement by 25% and wirelength perturbation by 21%, demonstrating its effectiveness in modern physical design. Jing Mai, Chunyuan Zhao, Zuodong Zhang, Zhixiong Di, Runsheng Wang, Yibo Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | RUPlace: Optimizing Routability via Unified Placement and Routing FormulationabstractPlacement plays a critical role in VLSI physical design, particularly in optimizing routability. With continuous advancements in semiconductor manufacturing technology, increased integration, and growing design complexity, managing routing congestion during placement has become increasingly challenging. Despite the widespread techniques to improve routability, these methods often lack theoretical guidance or sever the intrinsic connection between placement and routing optimization. In this paper, we propose RUPlace, an ADMMbased placer for unified optimization of placement and routing. Leveraging Wasserstein distance and bilevel optimization, our approach provides a unified framework for congestion optimization by alternately running global routing and incremental placement. Furthermore, we introduce a simple yet effective model for cell inflation-based global placement, where convex programming is employed to determine the optimal inflation ratio. Experimental results on a diverse set of open-source industrial benchmarks from CircuitNet and Chipyard demonstrate that our method achieves superior congestion reduction compared to widely used tools such as OpenROAD, Xplace 2.0, and DREAMPlace 4.1, while maintaining competitive wirelength and runtime. Jing Mai, Zuodong Zhang, Yibo Lin |
DAC | 2 |
| 2025 | LEGALM: Efficient Legalization for Mixed-Cell-Height Circuits with Linearized Augmented Lagrangian MethodabstractAdvanced technologies increasingly adopt mixed-cell-height circuits due to their superior power efficiency, compact area usage, enhanced routability, and improved performance. However, the complex constraints of modern circuit design, including routing challenges and fence region constraints, increase the difficulty of mixed-cell-height legalization. In this paper, we introduce LEGALM, a state-of-the-art mixed-cell-height legalizer that can address routability and fence region constraints more efficiently. We propose an augmented Lagrangian formulation coupled with a block gradient descent method that offers a novel analytical perspective on the mixed-cell-height legalization problem. To further enhance efficiency, we develop a series of GPU-accelerated kernels and a triplefold partitioning technique with minor quality overhead. Experimental results on ICCAD-2017 and modified ISPD-2015 benchmarks show that our approach significantly outperforms current state-of-the-art legalization algorithms in both quality and efficiency. Jing Mai, Chunyuan Zhao, Zuodong Zhang, Zhixiong Di, Yibo Lin, Runsheng Wang, Ru Huang 0001 |
ISPD | 1 |
| 2025 | A Robust FPGA Router With Optimization of High-Fanout Nets and Intra-CLB ConnectionsabstractRouting is the most time-consuming step in the implementation flow of field programmable gate array (FPGA) designs. With the advance in transistor scaling and system integration, hardware resources in FPGA devices are growing in a larger quantity and diversity. The routing architecture is designed to be more complicated for mapping RTL designs to FPGA devices correctly, which brings significant challenges for current FPGA routing algorithms. The key challenges for routing algorithms lie in large solution space and heavy congestion, especially for high-fanout nets (HFNets). We propose a partition-based algorithm to accelerate the routing of HFNets, which decomposes the global routing guide to shrink the search space for connecting each sink. Meanwhile, the congestions existing inside configurable logic block (CLB) are hard to handle by traditional sequential negotiation-based algorithms, because the industrial routing architecture is quite complex. We propose a concurrent intra-CLB rerouting algorithm to effectively resolve routing congestion inside a CLB tile induced by connections between intra-CLB logic pins, e.g., logic elements and switch boxes. Experimental results on modified ISPD2016 benchmarks demonstrate that our framework can achieve 100% routability in 9.8% less wirelength and$11\times $less runtime, while the state-of-the-art VTR 8.0 routing algorithm fails at 7 of 12 benchmarks. Xun Jiang 0002, Jing Mai, Zhixiong Di, Yibo Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | MORPH: More Robust ASIC Placement for Hybrid Region Constraint ManagementabstractModern ASIC placement tools encompass three categories of region constraints: default regions, fence regions, and guide regions. Region constraints pose significant challenges to existing placement algorithms, compromising the versatility and robustness required for diverse placement workloads. In this work, we propose MORPH, a more robust ASIC placer designed for hybrid region constraints. We integrate hybrid region constraints into a unified multi-electrostatic formulation that features a shared electrostatics model and a binary-lifting-based region pruning algorithm. We develop a more robust nonlinear placement framework that includes second-order information and a hybrid-region-aware legalization algorithm to address convergence issues. Experiments on the ISPD 2015 benchmark suite demonstrate 5.6-14.3% HPWL improvement and 10--24% overflow reduction compared to state-of-the-art region-aware placers. Further experiments on the ISPD 2015 benchmark suite and its variants show that the proposed techniques can achieve over 30% HPWL improvement and up to a twofold reduction in overflow with more stable convergence. Jing Mai, Zuodong Zhang, Yibo Lin, Runsheng Wang, Ru Huang 0001 |
ICCAD | 1 |
| 2024 | Multielectrostatic FPGA Placement Considering SLICEL-SLICEM Heterogeneity, Clock Feasibility, and Timing OptimizationabstractWhen modern FPGA architecture becomes increasingly complicated, modern FPGA placement is a mixed optimization problem with multiple objectives, including wirelength, routability, timing closure, and clock feasibility. Typical FPGA devices nowadays consist of heterogeneous SLICEs like SLICEL and SLICEM. The resources of a SLICE can be configured to {LUT, FF, distributed RAM, SHIFT, CARRY}. Besides such heterogeneity, advanced FPGA architectures also bring complicated constraints like timing, clock routing, carry chain alignment, etc. The above heterogeneity and constraints impose increasing challenges to FPGA placement algorithms. In this work, we propose a multielectrostatic FPGA placer considering the aforementioned SLICEL–SLICEM heterogeneity under timing, clock routing and carry chain alignment constraints. We first propose an effective SLICEL–SLICEM heterogeneity model with a novel electrostatic-based density formulation. We also design a dynamically adjusted preconditioning and carry chain alignment technique to stabilize the optimization convergence. We then propose a timing-driven net weighting scheme to incorporate timing optimization. Finally, we put forward a nested Lagrangian relaxation-based placement framework to incorporate the optimization objectives of wirelength, routability, timing, and clock feasibility. Experimental results on both academic and industrial benchmarks demonstrate that our placer outperforms the state-of-the-art placers in quality and efficiency. Jing Mai, Zhixiong Di, Yibo Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | LEAPS: Topological-Layout-Adaptable Multi-Die FPGA Placement for Super Long Line MinimizationabstractMulti-die FPGAs are crucial components in modern computing systems, particularly for high-performance applications such as artificial intelligence and data centers. Super long lines (SLLs) provide interconnections between super logic regions (SLRs) for a multi-die FPGA on a silicon interposer. They have significantly higher delay compared to regular interconnects, which need to be minimized. With the increase in design complexity, the growth of SLLs gives rise to challenges in timing and power closure. Existing placement algorithms focus on optimizing the number of SLLs but often face limitations due to specific topologies of SLRs. Furthermore, they fall short of achieving continuous optimization of SLLs throughout the entire placement process. This highlights the necessity for more advanced and adaptable solutions. In this paper, we propose LEAPS, a comprehensive, systematic, and adaptable multi-die FPGA placement algorithm for SLL minimization. Our contributions are threefold: 1) proposing a high-performance global placement algorithm for multi-die FPGAs that optimizes the number of SLLs while addressing other essential design constraints such as wirelength, routability, and clock routing; 2) introducing a versatile method for more complex SLR topologies of multi-die FPGAs, surpassing the limitations of existing approaches; and 3) executing continuous optimization of SLL counts across the whole placement stages, including global placement (GP), legalization (LG), and detailed placement (DP). Experimental results demonstrate the effectiveness of LEAPS in reducing SLLs and enhancing circuit performance. Compared with the most recent state-of-the-art (SOTA) method, LEAPS achieves an average reduction of 43.08% in SLL counts and 9.99% in HPWL while exhibiting a notable 34.34$\times$improvement in runtime. Zhixiong Di, Runzhe Tao, Jing Mai, Yibo Lin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | MacroRank: Ranking Macro Placement Solutions Leveraging Translation EquivariancyabstractModern large-scale designs make extensive use of heterogeneous macros, which can significantly affect routability. Predicting the final routing quality in the early macro placement stage can filter out poor solutions and speed up design closure. By observing that routing is correlated with the relative positions between instances, we propose MacroRank, a macro placement ranking framework leveraging translation equivariance and a Learning to Rank technique. The framework is able to learn the relative order of macro placement solutions and rank them based on routing quality metrics like wirelength, number of vias, and number of shorts. The experimental results show that compared with the most recent baseline, our framework can improve the Kendall rank correlation coefficient by 49.5% and the average performance of top-30 prediction by 8.1%, 2.3%, and 10.6% on wirelength, vias, and shorts, respectively. Jing Mai, Xiaohan Gao, Muhan Zhang, Yibo Lin |
ASP-DAC | 2 |
| 2023 | A Robust FPGA Router with Concurrent Intra-CLB ReroutingabstractRouting is the most time-consuming step in the FPGA design flow with increasingly complicated FPGA architectures and design scales. The growing complexity of connections between logic pins inside CLBs of FPGAs challenges the efficiency and quality of FPGA routers. Existing negotiation-based rip-up and reroute schemes will result in a large number of iterations when generating paths inside CLBs. In this work, we propose a robust routing framework for FPGAs with complex connections between logic elements and switch boxes. We propose a concurrent intra-CLB rerouting algorithm that can effectively resolve routing congestion inside a CLB tile. Experimental results on modified ISPD 2016 benchmarks demonstrate that our framework can achieve 100% routability in less wirelength and runtime, while the state-of-the-art VTR 8.0 routing algorithm fails at 4 of 12 benchmarks. Jing Mai, Zhixiong Di, Yibo Lin |
ASP-DAC | 2 |
| 2022 | Multi-electrostatic FPGA placement considering SLICEL-SLICEM heterogeneity and clock feasibilityabstractModern field-programmable gate arrays (FPGAs) contain heterogeneous resources, including CLB, DSP, BRAM, IO, etc. A Configurable Logic Block (CLB) slice is further categorized to SLICEL and SLICEM, which can be configured as specific combinations of instances in {LUT, FF, distributed RAM, SHIFT, CARRY}. Such kind of heterogeneity challenges the existing FPGA placement algorithms. Meanwhile, limited clock routing resources also lead to complicated clock constraints, causing difficulties in achieving clock feasible placement solutions. In this work, we propose a heterogeneous FPGA placement framework considering SLICEL-SLICEM heterogeneity and clock feasibility based on a multi-electrostatic formulation. We support a comprehensive set of the aforementioned instance types with a uniform algorithm for wirelength, routability, and clock optimization. Experimental results on both academic and industrial benchmarks demonstrate that we outperform the state-of-the-art placers in both quality and efficiency. Jing Mai, Yibai Meng, Zhixiong Di, Yibo Lin |
DAC | 1 |
| 2021 | Ultrafast CPU/GPU Kernels for Density Accumulation in PlacementabstractDensity accumulation is a widely-used primitive operation in physical design, especially for placement. Iterative invocation in the optimization flow makes it one of the runtime bottlenecks. Accelerating density accumulation is challenging due to data dependency and workload imbalance. In this paper, we propose efficient CPU/GPU kernels for density accumulation by decomposing the problem into two phases: constant-time density collection for each instance and a linear-time prefix sum. We develop CPU and GPU dedicated implementations, and demonstrate promising efficiency benefits on tasks from large-scale placement problems. Zizheng Guo 0001, Jing Mai, Yibo Lin |
DAC | 2 |
| 2021 | Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Peking UniversityabstractShi et al. (2018) proposed a highly parallel polynomial filtering eigensolver for the computation of planetary normal modes. As a challenge at the Student Cluster Competition in The International Conference for High Performance Computing, Networking, Storage and Analysis (SC19), we reproduce the computational efficiency of the polynomial filtering eigensolver on our Intel Xeon machine. We present the weak scalability, scaling of runtime with model size (in a fixed interval) and the strong scalability results in this report. Yihua Cheng, Zejia Fan, Jing Mai, Yifan Wu 0005, Pengcheng Xu 0005, Yuxuan Yan, Zhenxin Fu, Yun Liang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |