VLDB 2026 Research / reviewers in the wild / expert
Kaixiang Zhu
dblp:353/2870
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Novel Multi-Corner Delay Padding using Path Relationship Analysis and Dual DecompositionabstractMulti-corner timing analysis is essential for ensuring the robustness of circuits under variations in process, voltage, and temperature (PVT). Along with clock skew scheduling, delay padding is used to address hold violations. However, applying padding consistently in multiple corners is challenging due to conflicting constraints and the prevalence of “ping-pong” effects. This paper presents a novel methodology that uses dual decomposition to tackle this challenge. The problem is divided into a set of network flow problems, one for each corner. These problems are coupled through shared delay variables. Coordinating these subproblems using Lagrange multipliers ensures consistent padding assignments across corners. Additionally, traditional padding methods often struggle with physical feasibility. The incorporation of path relationship analysis is proposed to identify viable, physically feasible padding locations. Experimental results on industrial benchmarks demonstrate that the proposed method efficiently identifies feasible padding solutions and achieves the minimum clock period that satisfies the setup and hold time constraints for all corners. Compared to the single worst-case corner baseline, the optimized clock period is reduced by up to 9%, highlighting the effectiveness of our approach. Kaixiang Zhu, Lingli Wang, Wai-Shing Luk |
ASP-DAC | 1 |
| 2026 | A Collaborative Framework for Multi-Level Multi-Objective Design Space ExplorationabstractHigh-level synthesis (HLS) tools have drawn considerable attention in recent years because they can automatically generate hardware description code from high-level semantics under compiler-controlled configurations. However, the time-consuming design process, the inherent trade-offs among design objectives, and the often suboptimal quality of RTL produced by HLS have meant that prior studies rarely scale to or investigate the downstream stages beyond HLS.In this paper, we present COLA, an end to end design space exploration (DSE) framework that effectively automates the adaptive tuning of compiler transformation sequences and logic synthesis directives. First, we introduce MOEBO, a holistic Bayesian optimization method that builds multiple local surrogate models within trust regions while maintaining a global surrogate to correct local search bias and align decisions across regions. We further design a cooperative acquisition maximization scheme that coordinates these surrogates to propose diverse and promising candidates in parallel. Additionally, we employ reinforcement learning (RL) techniques to optimize logic synthesis by exploring the design space more effectively, improving the quality of the generated RTL and minimizing the area-delay product. The RL model dynamically adapts the logic synthesis directives to achieve better optimization outcomes over traditional methods. Experimental results show that, our framework achieves a substantial speedup across diverse accelerators for varying kernel granularities with a better trade-off between area and performance. Kaixiang Zhu, Yuping Bai, Yunfei Dai, Lingli Wang |
DATE | 2 |
| 2026 | An End-to-End Compilation Flow with Reinforcement Learning-Guided Logic Synthesis
Kaixiang Zhu, Lingli Wang |
ISCAS | 2 |
| 2026 | LOFMPL: An Open-source Logic Optimization Framework with MFFC-based Hypergraph Partition and Reinforcement Learning for Large CircuitsabstractAs the size of a circuit increases, previous reinforcement learning (RL) approaches struggle to effectively explore the logic optimization sequences of large-scale Boolean networks due to the long runtime overhead with poor optimization results. This article proposes LOFMPL: an open-source logic optimization framework with Maximum Fanout-Free Cone (MFFC) based hypergraph partitioning and reinforcement learning. The novel two-stage MFFC-based hypergraph partitioning can divide the circuit into highly independent subnetworks, which can be explored by an enhanced parallel RL-based design space exploration engine with an improved objective function. The experiment is conducted based on more than 150 benchmarks with logic optimization and ASIC technology mapping tasks and compared with other ML-based and greedy methods. The different partitioning algorithms are also compared for the subsequent logic optimization. Experimental results demonstrate that the proposed partitioning algorithm significantly enhances optimization quality without greatly increasing partitioning time, outperforming the KaHypar algorithm. Additionally, for the logic optimization task, the proposed method achieves a node-level-product improvement of 13% over the RLG synthesis exploration technique, 3% over the ESE reinforcement learning framework, 14% over the Boils synthesis method, and 7% over the DRiLLS synthesis method, while delivering greater reductions in node count compared with the Bulls-Eye optimization technique. For the ASIC technology mapping task, the proposed method achieves an area-delay-product improvement of 23% over the LSOracle framework, 9% over the Boils synthesis method, and 5% over the DRiLLS synthesis method. Hence, LOFMPL can achieve better results within the same runtime constraints compared with state-of-the-art works. Kaixiang Zhu, Zhen Li 0059, Jide Zhang, Wai-Shing Luk, Lingli Wang |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2025 | Yield-driven Clock Skew Scheduling Based on Generalized Extreme Value DistributionabstractClock skew scheduling is a cost-effective technique for enhancing the synchronous digital VLSI systems. The technique solely requires adjusting the clock skew to meet the timing constraints of the signal paths in order to increase the clock frequency or yield. In the past, Gaussian distributions were commonly assumed in the probability density function (PDF) modeling of maximum and minimum path delays under process variations. However, this assumption may not be appropriate due to the notable asymmetry of the actual path delay distributions. In this paper, we suggest the generalized extreme value (GEV) distribution as a potential alternative. Furthermore, we evaluate maximum likelihood estimation, linear-moments, and the method of moments (MoM) for parameter estimation. Experimental results show that the GEV distribution can more accurately approximate the cumulative distribution function (CDF) of the benchmark circuit path delays, resulting in an average improvement of 40% in the Kolmogorov-Smirnov (KS) statistic. Furthermore, yield-driven clock skew scheduling based on the GEV distribution produces superior timing yield outcomes compared to that based on the Gaussian distribution, with an improvement in timing yield up to 33% and average 8%. Kaixiang Zhu, Wai-Shing Luk, Lingli Wang |
ASP-DAC | 1 |
| 2025 | GEF: A GNN-Based Evaluation Framework for FPGA Routing ArchitectureabstractThe routing architecture significantly impacts the performance of modern FPGAs, motivating extensive research into its design space exploration (DSE). However, DSE efficiency is hindered by non-generalizable parametrization methods and considerable runtime overhead of FPGA architecture evaluation tools. In this paper, we propose GEF, a GNN-based FPGA Evaluation Framework that predicts routability and area-delay product (ADP) across various routing architectures. In GEF, we introduce Intra-Tile Graph, a novel intermediate representation (IR) that encodes global routing patterns in a compact form, serving as the input to predictors. The Routability Predictor (Rou-P) integrates Self-Attention Pooling (SAGPool), while the ADP Predictor (ADP-P) benefits from intermediate supervision through auxiliary node-level labels. Experimental results demonstrate the high accuracy of GEF, with Rou-P achieving 94.56% and ADP-P 94.57 %, respectively. We also conduct ablation studies, which further validate that GEF achieves substantial enhancements through efficient architecture modeling and timingaware analysis. Finally, a case study on routing architecture exploration with the incorporation of GEF is presented, which achieves a$\mathbf{1 5} \boldsymbol{\times}$speedup and enhanced improvements. Our codes are are available from https://github.com/RapidFlex/GEF. Yuanqi Wang, Yunfei Dai, Kaixiang Zhu, Huizhen Kuang, Eric Ren, Xifan Tang, Weijun Qin, Lingli Wang |
FPL | 4 |
| 2025 | DynVec: An End-to-End Framework for Efficient Vector-Dataflow ExecutionabstractHigh-performance computing (HPC) and hardware acceleration increasingly rely on dataflow architectures to achieve scalable parallelism and efficiency. High-level synthesis (HLS) facilitates accelerator design from high-level programs, but conventional tools often require intrusive source-level modifications and struggle to optimize irregular workloads. Dynamically scheduled HLS frameworks offer a promising direction for addressing control flow divergence and memory irregularity by generating dataflow accelerators. However, they lack compile-time parallelism optimizations such as vectorization and incur significant hardware overhead. Moreover, modern compilers can generate vectorized code using memory access and computational patterns. Nevertheless, in programs with irregular control flow or data-dependent behavior, such patterns are unknown until runtime, limiting the effectiveness of static vectorization strategies.To address these challenges, we propose DynVec, a unified vector-dataflow framework that integrates dynamic scheduling and vectorization to exploit runtime parallelism beyond conventional models. We address the vectorization of irregular kernels through an MLIR-based context-aware vectorizer that effectively identifies vectorizable operations and, through dataflow scheduling, generates a vector-dataflow execution graph that explicitly models control flow constructs, data and control interfaces, and memory operations. DynVec encapsulates high-level elastic units designed with built-in vectorization support, allowing customizable and adaptive execution behavior. Our compiler preserves the structural hierarchy of the kernel by combining vector and scalar operations in a bottom-up, type-safe manner. Experiments show that our approach achieves significant speedup compared to state-of-the-art HLS implementations across various regular and irregular applications. Moreover, compared to hybrid accelerators that separately support dynamic parallelism and vectorization, DynVec delivers superior performance. Xianfeng Cao, Kaixiang Zhu, Wenbo Yin, Lingli Wang |
ICCAD | 3 |
| 2025 | RLUT: A Reduced LUT Architecture with Fine-Grained Scalability and Its Automatic Design Flow for Large Frequent FunctionsabstractAs technology scaling exacerbates interconnect resistance in advanced nodes, FPGA architectures demand enhanced programmable logic blocks (PLBs) to minimize global metal routing. However, it is expensive to raise the functionality of LUTs due to exponential area growth with the number of inputs, resulting in poor scalability. Moreover, LUTs are redundant since practical functions in real-world benchmarks only account for an extremely small proportion of all the functions. For example, only 16,424 out of more than 100 trillion NPN classes of 6-input functions are used in the mapped netlists of the VTR8 and KOIOS benchmarks. Therefore, we propose a reduced LUT architecture, named RLUT, to efficiently implement most of the frequent functions. The compact structure of the MUX tree in LUTs is preserved and reduced, while the reduced programmable bits are connected to the MUX tree according to the bit assignment generated automatically by the proposed algorithms. Results of evaluations by a full EDA flow show that, compared with the modified Stratix10 baseline, the proposed 8-input PLB with 75 SRAM bits, named Dual-RLUT6, reduces the maximum logic levels significantly by 20.85%, while the critical path delay is improved by 10.11% at the cost of 4.65% area overhead. Moucheng Yang, Chengyu Zeng, Kaixiang Zhu, Lingli Wang |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2021 | Developing an Online Examination Timetabling System Using Artificial Bee Colony Algorithm in Higher Education
Kaixiang Zhu, Lily D. Li, Michael M. Li |
BROADNETS | 1 |