EDBT 2026 Demo / reviewers in the wild / expert
Hailong You
dblp:42/3308
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0003-3427-5320ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 12 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Efficient Optimization Framework for Netlist Partitioning by Co-optimizing Moves and Replications
Qiwang Chen, Hailong You, Shunyang Bi, Zehong Wei |
ISCAS | 2 |
| 2026 | ParSCo: Performance-Driven Partitioning and Scheduling Co-optimization Framework for Processor-based EmulationabstractAs the scale and complexity of designs increase, functional verification becomes a critical part of the very-large-scale integration (VLSI) design flow. However, existing processor-based emulation systems suffer from inefficiencies due to the misalignment objective between partitioning and scheduling, which are traditionally treated as separate and independent stages during compilation. To address this issue, we propose ParSCo , a partitioning and scheduling co-optimization framework that explicitly aligns the objectives of both stages by jointly considering cut minimization and topological order balancing (TOB) under multiple constraints. To integrate these objectives and constraints into our framework, we incorporate them into all partitioning and scheduling stages and further develop a set of novel techniques, including TOB-aware coarsening with multiple constraints , global growing initial partitioning with fixed nodes , TopoRefinement , and partitioning-aware scheduling , which collectively enhance the co-optimization process in emulation compilation. Furthermore, we establish theorems that reduce the time complexity of gain calculation and update to O (1), significantly improving the computational efficiency of the whole process. Furthermore, we evaluate the proposed method on the public and open-source chip design benchmarks, which have up to nearly 10 million cells. ParSCo significantly extends ideas and algorithms that first appeared in our previous work TopoOrderPart and achieves a 15% improvement. Extensive experimental results demonstrate the effectiveness of ParSCo , achieving an average improvement of 22.5% in time step reduction, 72% enhancement in TOB, and 55% acceleration in CPU time compared to the state-of-the-art (SOTA) two-stage partitioning and scheduling approach. Shunyang Bi, Hailong You, Cong Li 0023, Richard Sun |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2026 | Accurate Analytic Equation Generation for Compact Modeling with Physics-Assisted Kolmogorov-Arnold NetworksabstractThis article proposes a method to generate accurate and concise analytic equations for device compact modeling using Physics-Assisted Kolmogorov–Arnold Networks (PKAN). The equations are directly extracted from the trained neural network architecture. PKAN uses variable activation functions informed by prior physical knowledge to model device behaviors. Similarity constraints map these trained activation functions to mathematical symbols. Sparsification techniques simplify the network structure, producing concise and explicit equations. This article also presents four approaches for physics-assisted device modeling using PKAN: (1) generating entire continuous equations without human intervention, (2) applying correlation factors to existing models without requiring knowledge of internal physical mechanisms, (3) revising specific parts of existing models, and (4) automatically extending existing models. Experimental results show that PKAN demonstrates significant accuracy improvements, achieving error reductions of 91.8%, 91.5%, 66.2%, and 83.7% for corresponding experiments, respectively. These findings demonstrate PKAN’s potential for various device modeling applications. By combining the precision of neural networks with the clarity of symbolic representation, PKAN offers a powerful tool for device modeling applications. Zhengguang Tang, Zhenhai Cui, Cong Li 0023, Handing Wang, Hailong You |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2025 | A blockchain-based resource sharing incentivization mechanism for multi-to-multi in compute first networking
Zixu Zhang, Chenhao Ren, Hailong You |
Comput. Networks | 4 |
| 2024 | A High Performance Detailed Router Based on Integer Programming with Adaptive Route GuidesabstractDetailed routing is a crucial and time-consuming stage for ASIC design. As the number and complexity of design rules increase, it is challenging to achieve high solution quality and fast speed at the same time in detailed routing. In this work, a high performance detailed routing algorithm named IPAG with integer programming (IP) is proposed. The IP formulation uses the selection of candidate routes as decision variables. High quality candidate routes are generated by queue-based rip-up and reroute with adaptive global route guidance. A design rule checking engine which can simultaneously process nets with multiple routes is designed, to efficiently construct penalty parameters in the IP formulation. Experimental results on ISPD 2018 detailed routing benchmark show that IPAG achieves better solution quality in shorter or comparable runtime, as compared to the state-of-the-art academic detailed router. Zhongdong Qi, Shizhe Hu, Qi Peng 0003, Hailong You, Zhangming Zhu |
ASPDAC | 4 |
| 2024 | An Efficient Hypergraph Partitioner under Inter - Block Interconnection ConstraintsabstractMulti-FPGA systems are increasingly employed for very large scale integration circuit emulation and prototyping. Due to limited I/O resources, each FPGA often only has direct physical connections to a few other FPGAs. Therefore, if signals between FPGAs originate from a source FPGA and flow toward a target FPGA not directly connected to the source FPGA, intermediate FPGAs will be used as hops in the signal path. These FPGA-hops increase signal delays and the number of physical lines used in signal multiplexing between FPGAs, degrading system performance. To address these issues, researchers proposed partitioners that guarantees zero hop, but they lead to a considerable cut-size. In this paper, building on previous research, we introduce a new candidate block propagation theorem and optimize the initial partition process based on its corollary. Additionally, we also present a method for correcting violations during uncoarsening to improve the solver capability. Results of experiments demonstrate that our proposed method significantly reduces the cut size by 96% while retaining comparable running times. Benzheng Li, Hailong You, Shunyang Bi |
DATE | 2 |
| 2024 | TopoOrderPart: a Multi-level Scheduling-Driven Partitioning Framework for Processor-Based EmulationabstractIn a compilation flow of processor-based emulation (PBE), partitioning involves dividing a large netlist into smaller pieces and assigning them to different processors. Furthermore, the scheduling process must adhere to the levels of the netlist, which are determined by topological ordering, and the logic gates in the same level can be emulated in parallel. However, during the netlist partitioning stage, assigning most gates at the same level to one processor would undermine the benefits of parallelization in scheduling, leading to overall performance degradation. This paper proposes the TopoOrderPart, the first scheduling-driven partitioning framework for simultaneous balancing topological order and minimizing the cut size, which holds significant value in reducing time steps of scheduling. In particular, the topological order balancing and cut size are considered throughout the multilevel paradigm, and balance-aware coarsening achieves balancing between clusters in the early stage, with super-far root growing initial partitioning obtaining the better partition by selecting those root nodes in distant relationship within the connection space and two novel TopoRefine algorithms further enhancing the solution. Experimental results show TopoOrderPart can improve 69% topological order balancing and 0.53× run time while maintaining comparable cut size, compared to the state-of-the-art partitioner. Shunyang Bi, Hailong You, Cong Li 0023, Richard Sun |
ICCAD | 3 |
| 2024 | MaPart: An Efficient Multi-FPGA System-Aware Hypergraph Partitioning FrameworkabstractMulti-FPGA systems (MFSs) are increasingly important in addressing VLSI circuit emulation and prototyping. However, the limitations of I/O resources between FPGAs have driven the usage of TDM and FPGA-hop technologies, which complicate the partitioning problem. Consequently, designing a suitable partitioning process for MFS has emerged as a critical research question affecting overall system performance. This paper proposes MaPart, a novel hypergraph partitioning framework, which aims to minimize the maximum path delay in MFS. MaPart combines binary search with a non-hop partitioner, TopoPart+, to minimize the maximum hop count during the partitioning process. Compared to previous non-hop partitioner, TopoPart+ provides enhanced problem-solving capabilities and achieves a remarkable 96% reduction in cut-size. Furthermore, the framework incorporates two successive local refinement algorithms that optimize the time-division multiplexing ratio, reduce total hop count, and alleviate congestion on critical paths. Additionally, MaPart includes a system-level router based on layered graphs, enabling flexible control of the hop count based on the timing criticality of each path. Experimental results demonstrate that the proposed framework achieves a significant 37% reduction in delay compared to baseline algorithms when evaluated using publicly available benchmarks. Benzheng Li, Shunyang Bi, Hailong You, Zhongdong Qi, Richard Sun |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | ASSURER: A PPA-friendly Security Closure Framework for Physical DesignabstractHardware security is emerging in the very large scale integration (VLSI). The seminal threats, like hardware Trojan insertion, probing attacks, and fault injection, are hard to detect and almost impossible to fix at post-design stage. The optimal solution is to prevent them at the physical design stage. Usually, defending against them may cause a lot of power, performance, and area (PPA) loss. In this paper, we propose a PPA-friendly physical layout security closure framework ASSURER. Reward-directed placement refinement and multi-threshold partition algorithm are proposed to assure Trojan threats are empty. Cleaning up probing attacks is established on a patch-based ECO routing flow. Evaluated on the ISPD'22 benchmarks, ASSURER can clean out the Trojan threat with no leakage power increase when shrinking the physical layout area. When not shrinking, ASSURER only increases 14% total power. Compared with the work of first place in the ISPD2022 Contest, ASSURE reduced 53% additional total power consumption, and probing vulnerability can be reduced by 97.6% under the premise of timing closure. We believe this work shall open up a new perspective for preventing Trojan insertion and probing attacks. Hailong You, Zhengguang Tang, Benzheng Li, Cong Li 0023, Xiaojue Zhang |
ASP-DAC | 2 |
| 2023 | Machine Learning Based Framework for Fast Resource Estimation of RTL Designs Targeting FPGAsabstractField-programmable gate arrays (FPGAs) have grown to be an important platform for integrated circuit design and hardware emulation. However, with the dramatic increase in design scale, it has become a key challenge to partition very large scale integration into multi-FPGA systems. Fast estimation of FPGA on-chip resource usage for individual sub-circuit blocks early in the circuit design flow will provide an essential basis for reasonable circuit partition. It will also help FPGA designers to tune the circuits in hardware description language. In this article, we propose a framework for fast estimation of the on-chip resources consumed by register transfer level (RTL) designs with machine learning methods. We extensively collect RTL designs as a dataset, extract features from the result of a parser tool and analyze their roles, and train a targeted three-stage ensemble learning model. A 5,513× speedup is achieved while having 27% relative absolute error. Although the effect is sufficient to support RTL circuit partition, we discuss how the estimation quality continues to be improved. Benzheng Li, Hailong You, Zhongdong Qi |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2022 | Effective and Efficient Detailed Routing with Adaptive Rip-up Scheme and Pin Access RefinementabstractDetailed routing is one of the most complex and time-consuming stages of VLSI design process. Due to the rapidly growing problem scale and increasing number of design rules in advanced technology nodes, a feasible routing result can only be achieved after many rounds of rip-up and reroute (R&R) iterations, which takes a significantly long runtime. In this paper, we propose several effective and efficient techniques to handle the design rule violations in detailed routing. An adaptive rip-up scheme with two strategies of different effort is designed, which can speed up the R&R phase with comparable solution quality. To cope with the pin access challenge with complex design rule constraints, approaches to refine the pin connections are proposed. Besides, some specific design rules are handled in a post-processing manner efficiently. Experiment result shows that the number of design rule violations can be reduced by 69% with 28% lower runtime on average, after integrating these techniques in Dr. CU 2.0. Zhongdong Qi, Jingchong Zhang, Gengjie Chen, Hailong You |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | High quality hypergraph partitioning for logic emulation
Benzheng Li, Zhongdong Qi, Zhengguang Tang, Xiyi He, Hailong You |
Integr. | 5 |
| 2021 | Placement for Wafer-Scale Deep Learning AcceleratorabstractTo meet the growing demand from deep learning applications for computing resources, accelerators by ASIC are necessary. A wafer-scale engine (WSE) is recently proposed [1], which is able to simultaneously accelerate multiple layers from a neural network (NN). However, without a high-quality placement that properly maps NN layers onto the WSE, the acceleration efficiency cannot be achieved. Here, the WSE placement resembles the traditional ASIC floor plan problem of placing blocks onto a chip region, but they are fundamentally different. Since the slowest layer determines the compute time of the whole NN on WSE, a layer with a heavier workload needs more computing resources. Besides, locations of layers and protocol adapter cost of internal 10 connections will influence inter-layer communication overhead. In this paper, we propose GigaPlacer to handle this new challenge. A binary-search-based framework is developed to obtain a minimum compute time of the NN. Two dynamic-programming-based algorithms with different optimizing strategies are integrated to produce legal placement. The distance and adapter cost between connected layers will be further minimized by some refinements. Compared with the first place of the ISPD2020 Contest, GigaPlacer reduces the contest metric by up to 6.89% and on average 2.09%, while runs 7.23X faster. Benzheng Li, Qi Du, Dingcheng Liu, Jingchong Zhang, Gengjie Chen, Hailong You |
ASP-DAC | 6 |