Shixiong Kai

dblp:233/5903 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
14since 2021 · last 2025
0009-0006-1563-3179ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 PCBAgent: An Agent-based Framework for High-Density Printed Circuit Board Placement
abstract
Recently, printed circuit board (PCB) placement has emerged as a significant challenge since the scale of PCB designs has rapidly enlarged. Furthermore, the presence of various types of constraints with differing tolerance priorities hampers the automation of PCB layout design, necessitating substantial manual effort. To address this problem, we introduce a novel agent-based framework that automatically generates PCB layouts meeting industrial constraints through user interactions. This framework includes two main agents: a reinforcement learning (RL)-based agent for layout inference and fine-tuning, and a large language model (LLM)-based agent for interactive optimization. Experimental results on 17 industrial tasks show that our framework outperforms other state-of-the-art methods.
Lin Chen 0029, Ran Chen 0001, Shoubo Hu, Xufeng Yao, Zhentao Tang, Shixiong Kai, Mingxuan Yuan, Jianye Hao, Bei Yu 0001, Jiang Xu 0001
ASP-DAC6
2025 FTAFP: A Feedthrough-Aware Floorplanner for Hierarchical Design of Large-Scale SoCs
abstract
Floorplanning is a critical step in the physical design of digital integrated circuits (ICs). As circuit complexity grows, the hierarchical design paradigm of large-scale systems on chips (SoCs) is gradually emerging, introducing new optimization challenges, particularly with feedthrough. Feedthrough is a through-module connection, yet it would require additional buffers and ports inside the module for data transmission. Excessive feedthroughs will inevitably hinder the routability within reusable modules, causing congestion and timing problems. However, few works have addressed the challenges of feedthrough modeling and optimization.
Kanglin Tian, Jianwang Zhai, Shixiong Kai, Bei Yu 0001
ASP-DAC5
2025 ReMaP: Macro Placement by Recursively Prototyping and Periphery-Guided Relocating
abstract
We introduce the ReMaP framework, which generates expert-quality macro placements through recursively prototyping and periphery-guided relocating. A key innovation is ABPlace, an angle-based analytical method that arranges macros along an ellipse to facilitate a rough distribution near the periphery, while optimizing dataflow, minimizing overlap, and ensuring convergence. Based on the results of ABPlace, an efficient heuristic is proposed to position macros along the chip’s periphery, mirroring practices often employed by experts. Our framework outperforms three leading macro placers in both WNS and TNS across eight test cases, achieving improvements up to 34.15% in WNS and 65.39% in TNS, as tested on the popular OpenROAD-flow-scripts infrastructure. Additionally, our parameter autotuning method further improves timing by 8.75%.
Yunqi Shi, Xi Lin 0001, Shixiong Kai, Ke Xue 0001, Mingxuan Yuan, Chao Qian 0001, Zhi-Hua Zhou
DAC4
2025 Timing-Driven Global Placement by Efficient Critical Path Extraction
abstract
Timing optimization during the global placement of integrated circuits has been a significant focus for decades, yet it remains a complex, unresolved issue. Recent analytical methods typically use pin-level timing information to adjust net weights, which is fast and simple but neglects the path-based nature of the timing graph. The existing path-based methods, however, cannot balance the accuracy and efficiency due to the exponential growth of number of critical paths. In this work, we propose a GPU-accelerated timing-driven global placement framework, integrating accurate path-level information into the efficient DREAMPlace infrastructure. It optimizes the fine-grained pin-to-pin attraction objective and is facilitated by efficient critical path extraction. We also design a quadratic distance loss function specifically to align with the RC timing model. Experimental results demonstrate that our method significantly outperforms the current leading timing-driven placers, achieving an average improvement of 40.5% in total negative slack (TNS) and 8.3% in worst negative slack (WNS), as well as an improvement in half-perimeter wirelength (HPWL).
Yunqi Shi, Shixiong Kai, Xi Lin 0001, Ke Xue 0001, Mingxuan Yuan, Chao Qian 0001
DATE3
2025 LaMPlace: Learning to Optimize Cross-Stage Metrics in Macro Placement
abstract
Machine learning techniques have shown great potential in enhancing macro placement, a critical stage in modern chip design. However, existing methods primarily focus on *online* optimization of *intermediate surrogate metrics* that are available at the current placement stage, rather than directly targeting the *cross-stage metrics*---such as the timing performance---that measure the final chip quality. This is mainly because of the high computational costs associated with performing post-placement stages for evaluating such metrics, making the *online* optimization impractical. Consequently, these optimizations struggle to align with actual performance improvements and can even lead to severe manufacturing issues. To bridge this gap, we propose **LaMPlace**, which **L**earns **a** **M**ask for optimizing cross-stage metrics in macro placement. Specifically, LaMPlace trains a predictor on *offline* data to estimate these *cross-stage metrics* and then leverages the predictor to quickly generate a mask, i.e., a pixel-level feature map that quantifies the impact of placing a macro in each chip grid location on the design metrics. This mask essentially acts as a fast evaluator, enabling placement decisions based on *cross-stage metrics* rather than *intermediate surrogate metrics*. Experiments on commonly used benchmarks demonstrate that LaMPlace significantly improves the chip quality across several key design metrics, achieving an average improvement of 9.6\%, notably 43.0\% and 30.4\% in terms of WNS and TNS, respectively, which are two crucial cross-stage metrics that reflect the final chip quality in terms of the timing performance.
Zijie Geng, Jie Wang 0005, Ziyan Liu 0001, Zhentao Tang, Shixiong Kai, Mingxuan Yuan, Jianye Hao, Feng Wu 0001
ICLR6
2025 CORE: Collaborative Optimization with Reinforcement Learning and Evolutionary Algorithm for Floorplanning
abstract
Floorplanning is the initial step in the physical design process of Electronic Design Automation (EDA), directly influencing subsequent placement, routing, and final power of the chip. However, the solution space in floorplanning is vast, and current algorithms often struggle to explore it sufficiently, making them prone to getting trapped in local optima. To achieve efficient floorplanning, we propose **CORE**, a general and effective solution optimization framework that synergizes Evolutionary Algorithms (EAs) and Reinforcement Learning (RL) for high-quality layout search and optimization. Specifically, we propose the Clustering-based Diversified Evolutionary Search that directly perturbs layouts and evolves them based on novelty and performance. Additionally, we model the floorplanning problem as a sequential decision problem with B*-Tree representation and employ RL for efficient learning. To efficiently coordinate EAs and RL, we propose the reinforcement-driven mechanism and evolution-guided mechanism. The former accelerates population evolution through RL, while the latter guides RL learning through EAs. The experimental results on the MCNC and GSRC benchmarks demonstrate that CORE outperforms other strong baselines in terms of wirelength and area utilization metrics, achieving a 12.9\% improvement in wirelength. CORE represents the first evolutionary reinforcement learning (ERL) algorithm for floorplanning, surpassing existing RL-based methods. The code is available at https://github.com/yeshenpy/CORE.
Pengyi Li 0001, Shixiong Kai, Jianye Hao, Ruizhe Zhong, Hongyao Tang, Zhentao Tang, Mingxuan Yuan, Junchi Yan
NeurIPS2
2025 Benchmarking End-To-End Performance of AI-Based Chip Placement Algorithms
abstract
Chip placement is a critical step in the Electronic Design Automation (EDA) workflow, which aims to arrange chip modules on the canvas to optimize the performance, power, and area (PPA) metrics of final designs.Recent advances show great potential of AI-based algorithms in chip placement.However, due to the lengthy EDA workflow, evaluations of these algorithms often focus on intermediate surrogate metrics, which are computationally efficient but often misalign with the final end-to-end performance (i.e., the final design PPA).To address this challenge, we propose to build ChiPBench, a comprehensive benchmark specifically designed to evaluate the effectiveness of AI-based algorithms in final design PPA metrics.Specifically, we generate a diverse evaluation dataset from $20$ circuits across various domains, such as CPUs, GPUs, and NPUs.We then evaluate six state-of-the-art AI-based chip placement algorithms on the dataset and conduct a thorough analysis of their placement behavior.Extensive experiments show that AI-based chip placement algorithms produce unsatisfactory final PPA results, highlighting the significant influence of often-overlooked factors like regularity and dataflow.We believe ChiPBench will effectively bridge the gap between academia and industry.
Zijie Geng, Zhaojie Tu, Jie Wang 0005, Yuxi Qian, Zhexuan Xu, Ziyan Liu 0001, Zhentao Tang, Shixiong Kai, Mingxuan Yuan, Jianye Hao, Bin Li 0025, Feng Wu 0001
NeurIPS10
2024 PreRoutGNN for Timing Prediction with Order Preserving Partition: Global Circuit Pre-training, Local Delay Learning and Attentional Cell Modeling
abstract
Pre-routing timing prediction has been recently studied for evaluating the quality of a candidate cell placement in chip design. It involves directly estimating the timing metrics for both pin-level (slack, slew) and edge-level (net delay, cell delay), without time-consuming routing. However, it often suffers from signal decay and error accumulation due to the long timing paths in large-scale industrial circuits. To address these challenges, we propose a two-stage approach. First, we propose global circuit training to pre-train a graph auto-encoder that learns the global graph embedding from circuit netlist. Second, we use a novel node updating scheme for message passing on GCN, following the topological sorting sequence of the learned graph embedding and circuit graph. This scheme residually models the local time delay between two adjacent pins in the updating sequence, and extracts the lookup table information inside each cell via a new attention mechanism. To handle large-scale circuits efficiently, we introduce an order preserving partition scheme that reduces memory consumption while maintaining the topological dependencies. Experiments on 21 real world circuits achieve a new SOTA R2 of 0.93 for slack prediction, which is significantly surpasses 0.59 by previous SOTA method. Code will be available at: https://github.com/Thinklab-SJTU/EDA-AI.
Ruizhe Zhong, Junjie Ye 0002, Zhentao Tang, Shixiong Kai, Mingxuan Yuan, Jianye Hao, Junchi Yan
AAAI4
2024 Multi-Agent Trajectory Prediction with Scalable Diffusion Transformer
abstract
Accurate prediction of multi-agent spatiotemporal systems is critical to various real-world applications, such as autonomous driving, sports, and multiplayer games.Unfortunately, modeling multiagent trajectories is challenging due to its complicated, interactive, and multi-modal nature.Recently, diffusion models have achieved great success in modeling multi-modal distribution and trajectory generation, showing promising ability in resolving this problem.Motivated by this, in this paper, we propose a novel multi-agent trajectory prediction framework, dubbed Scalable Diffusion Transformer (SDT), which is naturally designed to learn the complicated distribution and implicit interactions among agents.We evaluate SDT on a set of real-world benchmark datasets and compare it with representative baseline methods, which demonstrates the better multi-agent trajectory prediction ability of SDT in terms of accuracy and diversity.
Shenyu Zhang 0001, Shixiong Kai, Chang Chen 0015, Yuzheng Zhuang, Zhengbang Zhu, Minghuan Liu, Weinan Zhang 0001
DAI2
2024 JigsawPlanner: Jigsaw-like Floorplanner for Eliminating Whitespace and Overlap among Complex Rectilinear Modules
abstract
As an early step in physical design, floorplanning plays a pivotal role in determining the performance upper bounds for downstream tasks and greatly impacts the PPA (power, performance, area) of the system. Efforts in floorplanning usually simplify modules as rectangles; however, assumptions of rectangular modules are not necessary for modern floorplan designs but could restrict floorplan solutions, typically resulting in lower chip area utilization with whitespace or overlaps among modules. In this paper, we challenge the widely accepted fixed-outline floorplanning problem setting which could lead to an inherent trade-off between whitespace and overlaps. We introduce JigsawPlanner, a novel and flexible Jigsaw-like floorPlanner that facilitates floorplanning to handle complex-shaped rectilinear modules. Given a global floorplan solution derived from an analytical method, we obtain the central positions of each module. Subsequently, we respectively partition the chip and modules into multiple grids and submodules, and assign these submodules to the grids using hierarchical Jonker-Volgenant algorithms and Cellular Automata. Empirical evaluations on public datasets show that JigsawPlanner can effectively eliminate white-space and overlaps simultaneously and significantly reduce the Half Perimeter Wire Length (HPWL) by an average of 10.50% compared to state-of-the-art baselines.
Xingbo Du, Ruizhe Zhong, Shixiong Kai, Zhentao Tang, Jianye Hao, Mingxuan Yuan, Junchi Yan
ICCAD3
2024 Reinforcement Learning Policy as Macro Regulator Rather than Macro Placer
abstract
In modern chip design, placement aims at placing millions of circuit modules, which is an essential step that significantly influences power, performance, and area (PPA) metrics. Recently, reinforcement learning (RL) has emerged as a promising technique for improving placement quality, especially macro placement. However, current RL-based placement methods suffer from long training times, low generalization ability, and inability to guarantee PPA results. A key issue lies in the problem formulation, i.e., using RL to place from scratch, which results in limits useful information and inaccurate rewards during the training process. In this work, we propose an approach that utilizes RL for the refinement stage, which allows the RL policy to learn how to adjust existing placement layouts, thereby receiving sufficient information for the policy to act and obtain relatively dense and precise rewards. Additionally, we introduce the concept of regularity during training, which is considered an important metric in the chip design industry but is often overlooked in current RL placement methods. We evaluate our approach on the ISPD 2005 and ICCAD 2015 benchmark, comparing the global half-perimeter wirelength and regularity of our proposed method against several competitive approaches. Besides, we test the PPA performance using commercial software, showing that RL as a regulator can achieve significant PPA improvements. Our RL regulator can fine-tune placements from any method and enhance their quality. Our work opens up new possibilities for the application of RL in placement, providing a more effective and efficient approach to optimizing chip design. Our code is available at \url{https://github.com/lamda-bbo/macro-regulator}.
Ke Xue 0001, Ruo-Tong Chen, Xi Lin 0001, Yunqi Shi, Shixiong Kai, Chao Qian 0001
NeurIPS5
2024 FlexPlanner: Flexible 3D Floorplanning via Deep Reinforcement Learning in Hybrid Action Space with Multi-Modality Representation
abstract
In the Integrated Circuit (IC) design flow, floorplanning (FP) determines the position and shape of each block. Serving as a prototype for downstream tasks, it is critical and establishes the upper bound of the final PPA (Power, Performance, Area). However, with the emergence of 3D IC with stacked layers, existing methods are not flexible enough to handle the versatile constraints. Besides, they typically face difficulties in aligning the cross-die modules in 3D ICs due to their heuristic representations, which could potentially result in severe data transfer failures. To address these issues, we propose FlexPlanner, a flexible learning-based method in hybrid action space with multi-modality representation to simultaneously handle position, aspect ratio, and alignment of blocks. To our best knowledge, FlexPlanner is the first learning-based approach to discard heuristic-based search in the 3D FP task. Thus, the solution space is not limited by the heuristic floorplanning representation, allowing for significant improvements in both wirelength and alignment scores. Specifically, FlexPlanner models 3D FP based on multi-modalities, including vision, graph, and sequence. To address the non-trivial heuristic-dependent issue, we design a sophisticated policy network with hybrid action space and asynchronous layer decision mechanism that allow for determining the versatile properties of each block. Experiments on public benchmarks MCNC and GSRC show the effectiveness. We significantly improve the alignment score from 0.474 to 0.940 and achieve an average reduction of 16% in wirelength. Moreover, our method also demonstrates zero-shot transferability on unseen circuits.
Ruizhe Zhong, Xingbo Du, Shixiong Kai, Zhentao Tang, Jianye Hao, Mingxuan Yuan, Junchi Yan
NeurIPS3
2023 RITA: Boost Driving Simulators with Realistic Interactive Traffic Flow
abstract
High-quality traffic flow generation is the core module in building simulators for autonomous driving. However, the majority of available simulators are incapable of replicating traffic patterns that accurately reflect the various features of real-world data while also simulating human-like reactive responses to the tested autopilot driving strategies. Taking one step forward to addressing such a problem, we propose Realistic Interactive TrAffic flow (RITA) as an integrated component of existing driving simulators to provide high-quality traffic flow for the evaluation and optimization of the tested driving strategies. RITA is developed with consideration of three key features, i.e., fidelity, diversity, and controllability, and consists of two core modules called RITABackend and RITAKit. RITABackend is built to support vehicle-wise control and provide traffic generation models from real-world datasets, while RITAKit is developed with easy-to-use interfaces for controllable traffic generation via RITABackend. We demonstrate RITA’s capacity to create diversified and high-fidelity traffic simulations in several highly interactive highway scenarios. The experimental findings demonstrate that our produced RITA traffic flows exhibit all three key features, hence enhancing the completeness of driving strategy evaluation. Moreover, we showcase the possibility for further improvement of baseline strategies through online fine-tuning with RITA traffic flows.
Zhengbang Zhu, Shenyu Zhang 0001, Yuzheng Zhuang, Yuecheng Liu, Minghuan Liu, Ziqing Gong, Shixiong Kai, Qiang Gu, Bin Wang 0034, Siyuan Cheng 0012, Xinyu Wang 0001, Jianye Hao, Yong Yu 0001
DAI7
2023 TOFU: A Two-Step Floorplan Refinement Framework for Whitespace Reduction
abstract
Floorplanning, as an early step in physical design, will greatly affect the PPA of the later stages. To achieve better performance while main-taining relatively the same chip size, the utilization of the generated floorplan needs to be high and constraints related to design rules, routability, power should be honored. In this paper, we propose a two-step framework, called TOFU, for floorplan whitespace reduction with fixed-outline and soft/pre- placed/hard modules modeled. Whitespace is first reduced by iteratively refining the locations of modules. Then the modules near whitespace will be changed into rectilinear shapes to further improve the utilization. To ensure the legality and quality of the intermediate floorplan during the refinement process, a constraint graph-based legalizer with a novel constraint graph construction method is proposed. Experimental results show that the whitespace of the initial floorplans generated by Corblivar [1] can be reduced by about 70% on average and up to 90% in several cases. Moreover, the resulting wirelength is also 3% shorter due to a higher utilization.
Shixiong Kai, Chak-Wa Pui, Shougao Jiang, Bin Wang 0034, Yu Huang 0005, Jianye Hao
DATE1
2020 A Multi-Task Reinforcement Learning Approach for Navigating Unsignalized Intersections
abstract
Navigating through unsignalized intersections is one of the most challenging problems in urban environments for autonomous vehicles. Existing methods need to train specific policy models to deal with different tasks including going straight, turning left and turning right. In this paper we formulate intersection navigation as a multi-task reinforcement learning problem and propose a unified learning framework for all three navigation tasks at the intersections. We propose to represent multiple tasks with a unified four-dimensional vector, which elements mean a common sub-task and three specific target sub-tasks respectively. Meanwhile, we design a vectorized reward function combining with deep Q-networks (DQN) to learn to handle multiple intersection navigation tasks concurrently. We train the agent to navigate through intersections by adjusting the speed of the ego vehicle under given route. Experimental results in both simulation and realworld vehicle test demonstrate that the proposed multi-task DQN algorithm outperforms baselines for all three navigation tasks in several different intersection scenarios.
Shixiong Kai, Bin Wang 0034, Jianye Hao, Wulong Liu
IV1
2018 Cooperative transportation control of multiple mobile manipulators through distributed optimization
Shixiong Kai
Sci. China Inf. Sci.2