VLDB 2026 Research / reviewers in the wild / expert
Mingyang Kou
dblp:276/2023
· DBLP profile ↗
12ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-9562-1365ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CERT: A Curved Escape Routing Framework for High-Speed Differential Pairs in Dense BGA PackagesabstractWith the rapid advancement of 5G communication, AI accelerators, and high-performance computing (HPC) systems, the data rate of BGA differential signals has surpassed 56Gbps. Traditional 45° routing causes impedance discontinuity and EMI issues, while also triggering design rule violations (DRVs) with route-keepout areas in dense BGA packages. In contrast, curved routing can provide superior signal integrity (SI), but relies on labor-intensive manual design. To address the above issues, this paper proposes CERT, a novel smooth (tangent-continuous) curved escape routing framework for high-speed differential pairs in dense BGA packages. A tailored hexagonal graph model with polyline-to-arc conversion rules is proposed, which guarantees the tangent continuity of curved wires for differential pairs and enables the direct reuse of traditional polyline escape routing algorithms. Experiments show CERT completes routing of the industrial benchmark in 2s (vs. 1 week for manual routing). To the best of our knowledge, this is the first automatic curved escape routing work for differential signals with route-keepout area adaptability. Weiqing Ji, Boxuan Xu, Hongli Dai, Chaojie Liu, Mingyang Kou, Hailong Yao 0002 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | Integrated Track Assignment and Detailed Routing for Enhanced Triple Patterning LithographyabstractAs semiconductor manufacturing advances toward smaller technology nodes, triple-patterning lithography (TPL) has become indispensable. Existing TPL-aware routers address manufacturability constraints too late, leading to numerous stitches, conflicts, and mask density imbalances. To overcome this, we propose a "shift-TPL-left" strategy that, for the first time, integrates TPL awareness into the track assignment (TA) stage and tightly coordinates it with detailed routing (DR). This approach guides the TPL-aware detailed routing from a more macroscopic level, resulting in faster convergence and fewer DRC violations. Experimental results demonstrate that our method achieves DRC clean in 80% of cases on the ISPD’18 dataset, outperforms the state-of-the-art TPL-aware routing method by 7 × in mask balance score, and achieves a 6 × speedup in runtime. Chengkai Wang, Weiqing Ji, Mingyang Kou, Nengyong Zhu, Hailong Yao 0002 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | GPS: GNN-Based Two-Stage Pre-Scheduling Loop Mapping Method on CGRAsabstractCoarse-grained reconfigurable architecture (CGRA) has emerged as a promising solution for accelerating computationally intensive applications, particularly in the field of artificial intelligence. One of the primary challenges for CGRA compilers is generating effective mapping results for complex applications within a limited time-frame. This paper presents an enhanced pre-scheduling method that integrates Integer Linear Programming (ILP) and Graph Neural Networks (GNN), along with a corresponding two-stage mapping approach. This combination significantly reduces the search space and accelerates the solution process for mapping problems. Experimental results demonstrate performance improvements ranging from $29.4 \%$ to $406.7 \%$, along with compilation time reductions of up to $1106.8 \times$ compared to existing compilation techniques, as well as excellent scalability. Mingyang Kou, Weiqing Ji, Shouyi Yin, Hailong Yao 0002 |
DAC | 1 |
| 2025 | Mr.TPL: A Method for Multi-Pin Net Router in Triple Patterning LithographyabstractTriple patterning lithography (TPL) has been recognized as one of the most promising solutions to print critical features in advanced technology nodes. A critical challenge within TPL is the effective assignment of the layout to masks. Recently, various layout decomposition methods and TPL-aware routing methods have been proposed to consider TPL. However, these methods typically result in numerous conflicts and stitches, and are mainly designed for 2-pin nets. This paper proposes a multipin net routing method in triple patterning lithography, called Mr.TPL. Experimental results demonstrate that Mr.TPL reduces color conflicts by 81.17%, decreases stitches by 76.89%, and achieves up to $5.4 \times$ speed improvement compared to the state-of-the-art TPL-aware routing method. Chengkai Wang, Weiqing Ji, Mingyang Kou, Zhiyang Chen 0006, Nengyong Zhu, Hailong Yao 0002 |
DAC | 3 |
| 2025 | VAER: Via-Aware Escape Routing for Chiplet Interconnection
Haochang Tian, Weiqing Ji, Mingyang Kou, Chengkai Wang, Hailong Yao 0002 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | NaviMap: Partial Order-Guided Neural Architecture via Deep Q-Networks for Efficient CGRA MappingabstractCoarse-Grained Reconfigurable Architectures (CGRAs) have emerged as promising solutions for energyefficient computing in edge devices and datacenter accelerators. While offering substantial performance benefits, their adoption is hindered by the NP-hard loop mapping problem during compilation. In this paper, we present NaviMap, a neuralsymbolic framework combining partial order-aware graph embeddings with deep reinforcement learning. Experimental results demonstrate that NaviMap achieves a$1.95 \times$speedup in solving CGRA mapping problems compared with state-of-the-art methods, while producing mappings with equivalent or superior performance. Mingyang Kou, Jun Zeng 0001, Xinyu Peng, Weiqing Ji, Hailong Yao 0002 |
ICCD | 1 |
| 2023 | TAEM 2.0: A Faster Transfer-Aware Effective Loop Mapping for Heterogeneous Resources on CGRAabstractCoarse-grained reconfigurable architectures (CGRAs) are energy-efficient and processing-flexible platforms to perform parallel computation. CGRAs combine the advantages of flexibility of general-purpose processors (GPPs) and energy efficiency of application-specific integrated circuits (ASICs). During the compilation process, the CGRA compiler needs to convert the high-level language codes into a data flow graph, and then map it onto CGRA to generate instruction flow and configuration context. The instruction mapping schemes of the CGRA compiler have a great impact on the efficiency and energy consumption of CGRAs. Furthermore, the quality of the instruction mapping schemes of the CGRA compiler highly depends on how the compiler maps data dependencies using different CGRA resources. This article proposes an enhanced transfer-aware loop mapping method, TAEM 2.0, based on state-of-the-art TAEM algorithm. Based on a parallel iterative IBBMCX algorithm and comprehensive CGRA resources analysis strategy, this method efficiently processes the complex situations of utilizing all those heterogeneous resources on CGRA and significantly accelerates the compilation process. Experimental results show TAEM 2.0 can accelerate the compilation process by$4.40\times $while generating the same or better mapping results on CGRA, when compared to the state-of-art mapping technique. Mingyang Kou, Jiangyuan Gu, Hailong Yao 0002, Shaojun Wei, Shouyi Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | GEML: GNN-based efficient mapping method for large loop applications on CGRAabstractCoarse-grained reconfigurable architecture (CGRA) is an emerging hardware architecture, with reconfigurable Processing Elements (PEs) for executing operations efficiently and flexibly. One major challenge for current CGRA compilers is the scalability issue for large loop applications, where valid loop mapping results cannot be obtained in an acceptable time. This paper proposes an enhanced loop mapping method based on Graph Neural Network (GNN), which effectively addresses the scalability issue and generates valid loop mapping results for large applications. Experimental results show that the proposed method enhances the compilation time by 10.8x on average over existing methods, with even better loop mapping solutions. Mingyang Kou, Jun Zeng 0001, Boxiao Han, Jiangyuan Gu, Hailong Yao 0002 |
DAC | 1 |
| 2022 | KunlunTVM: A Compilation Framework for Kunlun Chip Supporting Both Training and InferenceabstractWith the rapid development of deep learning, training big neural network models demands huge amount of computing power.Therefore, many accelerators are designed to meet the performance requirements. Recently, series of Kunlun chips have been released, which claim comparable performance over GPUs. However, there lacks an end-to-end compiler to support both training and inference on Kunlun chip,leaving large performance optimization space to be explored. This paper presents KunlunTVM, the first end-to-end compiler based on TVM, supporting both training and inference tasks on Kunlun Chip. Experimental results show that KunlunTVM achieves up to 5x training performance improvement over the existing framework PaddlePaddle supporting Kunlun chip. It is noteworthy that the proposed methods are general and extensible for the TVM framework targeting different backends. Jun Zeng 0001, Mingyang Kou, Hailong Yao 0002 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | NeuroSchedule: A Novel Effective GNN-based Scheduling Method for High-level SynthesisabstractHigh-level synthesis (HLS) is widely used for transferring behavior-level specifications into circuit-level implementations. As a critical step in HLS, scheduling arranges the execution order of operations for enhanced performance. However, existing scheduling methods suffer from either exponential runtime or poor quality of solutions. This paper proposes an efficient and effective GNN-based scheduling method called NeuroSchedule, with both fast runtime and enhanced solution quality. Major features are as follows: (1) The learning problem for HLS scheduling is formulated for the first time, and a new machine learning framework is proposed. (2) Pre-training models are adopted to further enhance the scalability for various scheduling problems with different settings. Experimental results show that NeuroSchedule obtains near-optimal solutions while achieving more than 50,000x improvement in runtime compared with the ILP-based scheduling method. At the same time, NeuroSchedule improves the scheduling results by 6.10% on average compared with state-of-the-art entropy-directed method. To the best of our knowledge, this is the first GNN-based scheduling method for HLS. Jun Zeng 0001, Mingyang Kou, Hailong Yao 0002 |
NeurIPS | 2 |
| 2021 | Splitter-Aware Multiterminal Routing With Length-Matching Constraint for RSFQ CircuitsabstractAided by the advancement of super-conductive materials, rapid single flux quantum (RSFQ) digital circuits are emerging as a promising complement or even replacement of the traditional CMOS digital integrated circuits. RSFQ digital circuits typically work at a low temperature of around 4.2 K, i.e., around −268.95 °C. Nevertheless, the operating frequency of RSFQ digital circuits reaches up to 770 GHz, which is orders of magnitudes faster than contemporary CMOS digital circuits. The high operating frequency causes critical design challenges especially for the clock networks and data path signals, where relative skew on wires need to be observed for achieving the correct functionality. Therefore, for designing a timing-variability-aware SFQ layout, it is necessary to match the PTL delays that are proportional to their respective lengths. And the matching of PTL delays should be carried out by extensions in PTL lengths. To meet the above-mentioned critical timing requirements, it is necessary to incorporate length-matching constraints into a routing problem, which is transformed from the timing requirements of matching the PTL delays during the logical synthesis stage. However, existing routing algorithms are inherently limited by preallocated splitters (SPLs), which complicates the subsequent routing stage under length-matching constraints. In this article, in order to effectively address the length-matching constraints, we reallocate SPLs to fully utilize routing resources. We propose the first multiterminal routing algorithm for RSFQ circuits, which integrates SPL reallocation into the routing stage and achieves 100% routing completion in the tested benchmarks. Compared with the state-of-the-art method, the proposed multiterminal routing algorithm reduces the required area by 17% and the runtime by 7%. Mingyang Kou, Pei-Yi Cheng, Jun Zeng 0001, Tsung-Yi Ho, Kazuyoshi Takagi, Hailong Yao 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | TAEM: Fast Transfer-Aware Effective Loop Mapping for Heterogeneous Resources on CGRAabstractCoarse-grained reconfigurable architecture (CGRA) is an energy-efficient and processing-flexible parallel computing architecture. Efficiency of CGRA highly depends on how to map data dependencies using different CGRA resources. Previous works investigated different strategies for transferring data dependencies, using registers, processing elements (PEs) and memory. However, these works do not consider all those resources in CGRA and take a long time during compilation period. This paper proposes a Transfer-Aware Effective loop Mapping (TAEM) method for CGRA, which can efficiently utilize all those heterogeneous resources on CGRA and significantly accelerate the compilation time. Experimental results show that TAEM is able to reduce the compilation time by 11.1x over the state-of-the-art technique RAMP, while keeping the same or better performance of loop mapping results. Mingyang Kou, Jiangyuan Gu, Shaojun Wei, Hailong Yao 0002, Shouyi Yin |
DAC | 1 |