VLDB 2026 Research / reviewers in the wild / expert
Jiangyuan Gu
dblp:184/7513
· DBLP profile ↗
16ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-1190-7524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hetero-ChipletSim: Bridging Chiplet, Interconnect and Packaging Heterogeneity in Multi-Chiplet System SimulationabstractWith the end of Moore’s Law, multi-chiplet systems have emerged as a promising solution featuring heterogeneity across chiplets, interconnects and packaging. Existing simulators lack support for such multi-level heterogeneity, making accurate architectural exploration difficult. We propose Hetero-ChipletSim (HCS), a simulation methodology that directly integrates heterogeneous chiplet models while incorporating die-to-die(D2D) interconnect and packaging effects, enabling fast and accurate evaluation of multi-chiplet systems. Sensitivity analysis provides insights into design trade-offs under heterogeneous integration. Xuguang Yuan, Jiangyuan Gu, Qidie Wu, Yang Hu 0001, Shaojun Wei, Shouyi Yin |
DATE | 2 |
| 2025 | DIAG: A Refined Four-layer Agile Hardware Developing Flow for Generating Flexible Reconfigurable ArchitecturesabstractRapid evolution in application algorithms, exemplified by advancements in artificial intelligence, wireless communication, and sciencific computing, necessitates a focus on developing energy-efficient, highly-flexible parallel computing architectures. This urgency is further amplified by the need for agile hardware development techniques to mitigate design complexity and reduce costs. Among emerging agile hardware development techniques, generative HDL stands out due to its straightforward grammatical structure and compatibility with hardware design thinking, yet it remains underutilized. In response, this paper introduces a novel four-layer agile developing flow, termed DIAG, innovatively leveraging unique Plugin-Service technology. The DIAG framework is applied to an extensible reconfigurable architecture generator, enabling the generation of diverse CGRA designs suitable for accelerating task computations across multiple application domains. Our comprehensive experiments on the CGRA generator design validate the efficiency of the DIAG flow and underscore generative-HDL's significant potential for complex, large-scale hardware development. Haojia Hui, Jiangyuan Gu, Xunbo Hu, Shaojun Wei, Shouyi Yin |
ASP-DAC | 2 |
| 2025 | Computing Efficiency Improvement for Multi-PEA CGRA with Built-in Control DesignabstractThe growing demands of modern applications, such as AI, graph computing, and big data processing, are driving the increase in algorithmic scale and computational workload.As a result, Multi-PEA CGRA has been a popular choice because of its high computing power.However, such kinds of Architecture are confronted with control problem due to the large amount of PEA required management on Architecture.To address this challenge, this paper propose a built-in control design(Control Element, CE) for multi-PEA CGRA to improve computing efficiency.In this paper, this paper has compared the execution time, and power consumption with and without CE.Experiments demonstrate that our CE can reduce 97.3% execution time, in which the proportion of PEA work time is 76.2% and at least 79.7% power consumption. Jiangyuan Gu, Xunbo Hu, Zidi Qin, Shaojun Wei, Shouyi Yin |
CF | 1 |
| 2025 | P2P-Chiplet: Partition and Placement Co-Optimization for Multi-Chiplet ArchitectureabstractThe rising cost and complexity of cutting-edge process nodes have impeded large monolithic System-on-Chip to follow Moore’s Law, forcing chip designers to embrace Multi-Chiplet architectures. Multi-chiplet designs achieve cost reduction while maintaining near-monolithic performance by disaggregating a large die into smaller chiplets and integrating them through advanced packaging. The payback of this Disaggregation-Integration paradigm critically depends on the efficacy of chiplet Partition and Placement framework. However, existing frameworks fail to harness the potential merits offered by Partition-Placement Co-Optimization. Serving as an input provider for placement, partition phase typically adjusts block-to-die assignments to guide subsequent placement. This sequential dependency implies an inherent Partition-Placement (P2P) Inconsistency problem: solutions optimal solely in partition or placement may finally cause an inferior solution. Hence, this paper proposes P2P-Chiplet, a Partition-Placement Co-Optimization framework for multi-chiplet designs. Firstly, an optimized ACG structure, named as HeteroACG, is introduced to aggregate topological partition and physical placement optimization spaces. Then, the sequential partition-placement flow is decomposed into interleaved fine-grained epochs and an alternating progressive optimization strategy is employed to preserve P2P Consistency and bring better co-optimized solutions. Finally, experimental results show that, compared with existing chiplet partition-placement frameworks, our proposed P2P Chiplet notably mitigates potential performance bottlenecks while effectively reducing costs within acceptable overhead. Qidie Wu, Jiangyuan Gu, Xuguang Yuan, Shaojun Wei, Shouyi Yin |
ICCAD | 2 |
| 2023 | RMP-MEM: A HW/SW Reconfigurable Multi-Port Memory Architecture for Multi-PEA Oriented CGRAabstractCoarse-Grained Reconfigurable Architecture (CGRA), especially the one with multiple parallelized Processing Element Arrays (PEA), possesses flexible programmability and high parallel computational efficiency, which relies upon an efficient memory architecture to deliver the corresponding computing power. Multi-PEA oriented CGRA allows for mapping various applications and thus demands a flexible memory to adapt to the ever-changing workloads, whose parallel access also requires an efficient multi-port memory. However, the existing memory designs for CGRA are hard to satisfy those requirements since conventional rigid memories fail to provide the desired flexibility due to fixed structure, and traditional multi-port designs are impractical due to large overhead. Therefore, this paper proposes a hardware/software (HW/SW) hybrid reconfigurable multi-port memory architecture (RMP-MEM) with an instructive analysis for the multi-PEA oriented CGRA. RMP-MEM supports adaptive memory partition and programmer-defined access modes to adapt the different features of memory accesses. Also, RMP-MEM achieves an efficient multi-port implementation by a partially shared mechanism. Furthermore, the microarchitecture of RMP-MEM is optimized multi-directionally, resulting in a significant performance gain. The experimental results indicate that RMP-MEM reduces the parallel access latency by 81.1% and exhibits 28.3% energy efficiency improvement compared to prior designs. Qidie Wu, Jiangyuan Gu, Youxu Lin, Boxiao Han, Hongjun He, Yang Hu 0001, Leibo Liu, Shaojun Wei, Shouyi Yin |
DAC | 2 |
| 2023 | TAEM 2.0: A Faster Transfer-Aware Effective Loop Mapping for Heterogeneous Resources on CGRAabstractCoarse-grained reconfigurable architectures (CGRAs) are energy-efficient and processing-flexible platforms to perform parallel computation. CGRAs combine the advantages of flexibility of general-purpose processors (GPPs) and energy efficiency of application-specific integrated circuits (ASICs). During the compilation process, the CGRA compiler needs to convert the high-level language codes into a data flow graph, and then map it onto CGRA to generate instruction flow and configuration context. The instruction mapping schemes of the CGRA compiler have a great impact on the efficiency and energy consumption of CGRAs. Furthermore, the quality of the instruction mapping schemes of the CGRA compiler highly depends on how the compiler maps data dependencies using different CGRA resources. This article proposes an enhanced transfer-aware loop mapping method, TAEM 2.0, based on state-of-the-art TAEM algorithm. Based on a parallel iterative IBBMCX algorithm and comprehensive CGRA resources analysis strategy, this method efficiently processes the complex situations of utilizing all those heterogeneous resources on CGRA and significantly accelerates the compilation process. Experimental results show TAEM 2.0 can accelerate the compilation process by$4.40\times $while generating the same or better mapping results on CGRA, when compared to the state-of-art mapping technique. Mingyang Kou, Jiangyuan Gu, Hailong Yao 0002, Shaojun Wei, Shouyi Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Mixed-granularity parallel coarse-grained reconfigurable architectureabstractCoarse-Grained Reconfigurable Architecture (CGRA) is a high-performance computing architecture. However, existing CGRA silicon utilization is low due to the lack of fine-grained parallelism inside Processing Element (PE) and general coarse-grained parallel approach on PE array. No fine-grained parallelism in PE not only leads to low silicon utilization of PE, but also makes the mapping loose and irregular. No generalized parallel method for the mapping cause low PE utilization on CGRA. Our goal is to design an execution model and a Mixed-granularity Parallel CGRA (MP-CGRA), which is capable to fine-grained parallelize operators excution in PEs and parallelize data transmission in channels, leading to a compact mapping. A coarse-grained general parallel method is proposed to vectorize the compact mapping. Evaluated with Machsuite, MP-CGRA achieves an improvement of 104.65% silicon utilization on PE array and a 91.40% performance per area improvement compared with baseline-CGRA. Jinyi Deng, Linyun Zhang, Kexiang Deng, Shibin Tang, Jiangyuan Gu, Boxiao Han, Leibo Liu, Shaojun Wei, Shouyi Yin |
DAC | 7 |
| 2022 | GEML: GNN-based efficient mapping method for large loop applications on CGRAabstractCoarse-grained reconfigurable architecture (CGRA) is an emerging hardware architecture, with reconfigurable Processing Elements (PEs) for executing operations efficiently and flexibly. One major challenge for current CGRA compilers is the scalability issue for large loop applications, where valid loop mapping results cannot be obtained in an acceptable time. This paper proposes an enhanced loop mapping method based on Graph Neural Network (GNN), which effectively addresses the scalability issue and generates valid loop mapping results for large applications. Experimental results show that the proposed method enhances the compilation time by 10.8x on average over existing methods, with even better loop mapping solutions. Mingyang Kou, Jun Zeng 0001, Boxiao Han, Jiangyuan Gu, Hailong Yao 0002 |
DAC | 5 |
| 2021 | Combining Memory Partitioning and Subtask Generation for Parallel Data Access on CGRAsabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are attractive reconfigurable platforms with the advantages of high performance and power efficiency. In a CGRA based computing system, the computations are often mapped onto the CGRA with parallel memory accesses. To fully exploit the on-chip memory bandwidth, memory partitioning algorithms are widely used to reduce access conflicts. CGRAs have a fixed storage fabric and limited size memory due to the severe area constraints. Previous memory partitioning algorithms assumed that data could be completely transferred into the target memory. However, in practice, we often encounter situations where on-chip storage is insufficient to store the complete data. In order to perform the computation of these applications in the memory-limited CGRA, we first develop a memory partitioning strategy with continual placement, which can also avoid data preprocessing, and then divide the kernel into multiple subtasks that suit the size of the target memory. Experimental results show that, compared to the state-of-the-art method, our approach achieves a 43.2% reduction in data preparation time and an 18.5% improvement in overall performance. If the subtask generation scheme is adopted, our approach can achieve a 14.4% overall performance improvement while reducing memory requirements by 99.7%. Jiangyuan Gu, Shouyi Yin, Leibo Liu, Shaojun Wei |
ASP-DAC | 2 |
| 2021 | A Multiple-Precision Multiply and Accumulation Design with Multiply-Add Merged Strategy for AI AcceleratingabstractMultiply and accumulations(MAC) are fundamental operations for domain-specific accelerator with AI applications ranging from filtering to convolutional neural networks(CNN). This paper proposes an energy-efficient MAC design, supporting a wide range of bit-width, for both signed and unsigned operands. Firstly, based on the classic Booth algorithm, we propose the Booth algorithm to propose a multiply-add merged strategy. The design can not only support both signed and unsigned operations but also eliminate the delay, area and power overheads from the adder of traditional MAC units. Then a multiply-add merged design method for flexible bit-width adjustment is proposed using the fusion strategy. In addition, treating the addend as a partial product makes the operation easy to pipeline and balanced. The comprehensive improvement in delay, area and power can meet various requirements from different applications and hardware design. By using the proposed method, we have synthesized MAC units for several operation modes using a SMIC 40-nm library. Comparison with other MAC designs shows that the proposed design method can achieve up to 24.1% and 28.2% PDP and ADP improvement for bit-width fixed MAC designs, and 28.43% ~ 38.16% for bit-width adjustable ones. When pipelined, the design has decreased the latency by more than 13%. The improvement in power and area is up to 8.0% and 8.1% respectively. Jiangyuan Gu, Shouyi Yin, Leibo Liu, Shaojun Wei |
ASP-DAC | 2 |
| 2020 | TAEM: Fast Transfer-Aware Effective Loop Mapping for Heterogeneous Resources on CGRAabstractCoarse-grained reconfigurable architecture (CGRA) is an energy-efficient and processing-flexible parallel computing architecture. Efficiency of CGRA highly depends on how to map data dependencies using different CGRA resources. Previous works investigated different strategies for transferring data dependencies, using registers, processing elements (PEs) and memory. However, these works do not consider all those resources in CGRA and take a long time during compilation period. This paper proposes a Transfer-Aware Effective loop Mapping (TAEM) method for CGRA, which can efficiently utilize all those heterogeneous resources on CGRA and significantly accelerate the compilation time. Experimental results show that TAEM is able to reduce the compilation time by 11.1x over the state-of-the-art technique RAMP, while keeping the same or better performance of loop mapping results. Mingyang Kou, Jiangyuan Gu, Shaojun Wei, Hailong Yao 0002, Shouyi Yin |
DAC | 2 |
| 2018 | Stress-Aware Loops Mapping on CGRAs with Dynamic Multi-Map ReconfigurationabstractWith VLSI process technology scaling into nano-scale, the increasingly serious aging issues (e.g., NBTI and HCI aging effects) have brought a significant threat to system reliability. Coarse-grained reconfigurable architectures (CGRAs) exhibit the feature to reconfigure and execute different mapping schemes (Maps) dynamically, compensating for each other to mitigate aging issues effectively. In this paper, a two-stage stress-aware loops mapping algorithm is first proposed for the CGRA-mapped designs by jointing the intra-kernel and inter-kernel stress optimizations. With pipelining techniques, the intra-kernel stress optimization employs the stress-aware force-directed and effective MCC (Maximal Compatibility Class) methods to optimize operations' placement and mapping distribution on processing elements (PEs), which helps to avoid overmany operations to be mapped on the same PEs and reduce the accumulated stresses. By leveraging the dynamic reconfiguration feature, the inter-kernel stress optimization develops a multi-map scheduling method to reconfigure a set of ordered maps on CGRA dynamically, which diversifies the PEs' usage and compensates for the stresses on different PEs among them. Experimental results show that our approach can reduce the maximum stress by 82.0% for NBTI and 70.4% for HCI, and improve the aging efficiency by 6.01X and MTTF by 3.16X averagely, while keeping the optimized performance. Jiangyuan Gu, Shouyi Yin, Leibo Liu, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Energy-aware loops mapping on multi-vdd CGRAs without performance degradationabstractCoarse Grained Reconfigurable Architectures (C-GRAs) have been paid an increasing attention due to their inherent advantages of high performance and energy efficiency. As we know, multi-Vddtechnique is popularly used to reduce energy consumption, and modulo scheduling is one of widely-used pipeline techniques to improve performance. To achieve both high performance and energy-efficiency simultaneously, this paper proposes an energy-aware mapping algorithm integrating multi-Vddassignment into the scheduling and mapping procedures of loop applications. Also, an energy-aware FDS (eFDS) algorithm and a rapid MCC searching method based on compatibility concept are successfully adopted to solve the bi-objective optimization problem. The experimental results show that the proposed approach brings 18.7% energy reduction and 1.44X energy-efficiency improvement while keeping optimized performance. Jiangyuan Gu, Shouyi Yin, Leibo Liu, Shaojun Wei |
ASP-DAC | 1 |
| 2017 | Stress-Aware Loops Mapping on CGRAs with Considering NBTI Aging EffectabstractWith the process scaling into nano-scale VLSI technology, the increasingly serious aging issues (e.g. NBTI aging effect) bring a significant threat to system reliability. Coarsegrained reconfigurable architectures (CGRAs) exhibit the feature to reconfigure different mapping schemes (Maps) dynamically during loops execution, which can mitigate the aging issues on CGRAs effectively. In this paper, we propose a stress-aware loops mapping algorithm by jointing intra-kernel and inter-kernel stress optimizations strategies in the early phase of CGRA-mapped designs. With the pipelining technique, a stress-aware force-directed method is introduced in the intra-kernel optimization, avoiding many operations to be mapped on some certain PEs and reducing the stresses accumulated on them. By leveraging the dynamic reconfiguration, a multi-map scheduling method is proposed in the inter-kernel stress optimization to find a set of ordered maps to reconfigure dynamically, which diversifies PE usages and compensates for the accumulated stresses on different PEs among them. Experimental results show our proposed approach enlarges the maximum stress reduction up to 78.9% and improves the MTTF by 340.3% on average while keeping the optimized performance. Jiangyuan Gu, Shouyi Yin, Shaojun Wei |
DAC | 1 |
| 2017 | Conflict-Free Loop Mapping for Coarse-Grained Reconfigurable Architecture with Multi-Bank MemoryabstractCoarse-grained reconfigurable architecture (CGRA) is a promising architecture with high performance, high power-efficiency and attraction of flexibility. The computation-intensive parts of an application (e.g., loops) are often mapped on CGRA for acceleration. Due to the high parallel data access demands, the architecture with multi-bank memory is proposed to improve parallelism. For CGRA with multi-bank memory, a joint solution, which simultaneously considers the memory partitioning and modulo scheduling, is proposed to achieve a valid mapping with better performance. In this solution, the modulo scheduling and operator scheduling are used to achieve a valid loop mapping and a valid data placement without any memory access conflicts. By avoiding the pipelining stalls caused by conflicts, the performance of loop mapping is greatly improved. The experimental results on benchmarks of the Livermore, Polybench and Mediabench show that our approach can improve the performance of loops on CGRA to 1.89×, 1.49× and 1.37× compared with REGIMap, HTDM and REGIMap with memory partitioning, at cost of an acceptable increase in compilation time. Shouyi Yin, Xianqing Yao, Dajiang Liu, Jiangyuan Gu, Leibo Liu, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2016 | Joint Modulo Scheduling and Vdd Assignment for Loop Mapping on Dual- Vdd CGRAsabstractCoarse-grained reconfigurable architecture (CGRA) is becoming an increasingly attractive platform because of its high performance and power (or energy) efficiency. To reduce energy consumption, the dual-Vddtechnique has been employed in CGRAs, and the modulo scheduling technique is widely used to improve performance of applications. To achieve both high performance and energy-efficiency simultaneously, this paper formulates the solution as a biobjective optimization problem of energy consumption and initiation interval of loop pipelines on CGRAs, and proposes a joint modulo scheduling and dual-Vddassignment approach. The experimental results show that the proposed approach can bring a significant energy reduction of 24.8% and kernel energy efficiency acceleration of 1.41× on average, while the performance is maintained. Shouyi Yin, Jiangyuan Gu, Dajiang Liu, Leibo Liu, Shaojun Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |