Mengke Ge

dblp:219/4222 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0001-7888-9370ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Thermal-Aware Scheduling for DNN Inference on 3D Logic-to-DRAM Process-Near-Memory Architecture
Shiji Ke, Conghui Li, Mengke Ge, Song Chen 0001, Yi Kang
HPDC3
2026 Throughput Maximization for Transformer Inference on Processing Near-Memory Architectures
abstract
The advent of Transformers has revolutionized fields such as computer vision and natural language processing. However, their memory-intensive nature creates significant hurdles for conventional computing platforms such as CPUs and GPUs. Processing near-memory (PNM) architecture has arisen as a promising solution to mitigate the memory wall problem. However, efficiently deploying Transformer models on PNM architecture remains a cutting-edge challenge. To address the practical demands of cloud and edge computing, we propose a novel mapping framework called Energon, which aims to facilitate high-throughput inference of encoderbased Transformers on PNM-based neural network (NN) accelerators, catering to both non-latency-sensitive and latencybounded scenarios. Firstly, Energon introduces a novel pipeline parallelism strategy based on an XY-aligned layout, which offers an enhanced flexibility in pipeline layout compared to existing pipeline parallelism approaches, while adapting to the finegrained partitioning scheme tailored for Transformers to achieve efficient mass parallelism. Secondly, Energon formulates the mapping optimization problems using dynamic programming and integer linear programming, respectively, to jointly optimize network partitioning and pipeline layout construction for a globally optimal mapping solution. Experimental results demonstrate that Energon significantly improves the inference throughput of encoder-based Transformers on PNM accelerators, outperforming state-of-the-art mapping frameworks by 1.1× to 2.3×. Under user-defined latency bounds, it enhances the inference throughput by an average of 43% and up to 123%.
Mengke Ge, Yingjian Zhong, Song Chen 0001, Yi Kang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems
abstract
The advent of Transformers has revolutionized computer vision, offering a powerful alternative to convolutional neural networks (CNNs), especially with the local attention mechanism that excels at capturing local structures within the input and achieve state-of-the-art performance. Processing in-memory (PIM) architecture offers extensive parallelism, low data movement costs, and scalable memory bandwidth, making it a promising solution to accelerate Transformer with memory-intensive operations. However, the crucial issue lies in efficiently deploying an entire model onto resource-limited PIM system while parallelizing each transformer block with potentially many computational branches based on local-attention mechanisms. We present Allspark, which focuses on workload orchestration for visual Transformers on PIM systems, aiming at minimizing inference latency. Firstly, to fully utilize the massive parallelism of PIM, Allspark employs a fine-grained partitioning scheme for computational branches, and formats a systematic layout and interleaved dataflow with maximized data locality and reduced data movement. Secondly, Allspark formulates the scheduling of the complete model on a resource-limited distributed PIM system as an integer linear programming (ILP) problem. Thirdly, as local-global data interactions exhibit complex yet regular dependencies, Allspark provides a two-stage placement method, which simplifies the challenging placement of computational branches on the PIM system into the structured layout and greedy-based binding, to minimize NoC communication costs. Extensive experiments on 3D-stacked DRAM-based PIM systems show that Allspark brings$1.2\times$$\sim$$24.0\times$inference speedup for various visual Transformers over baselines. Compared to Nvidia V100 GPU, Allspark-enriched PIM system yields average speedups of$2.3\times$and energy savings of$20\times$$\sim$$55\times$.
Mengke Ge, Junpeng Wang 0002, Binhan Chen, Yingjian Zhong, Haitao Du, Song Chen 0001, Yi Kang
IEEE Trans. Computers1
2024 NicePIM: Design Space Exploration for Processing-In-Memory DNN Accelerators With 3-D Stacked-DRAM
abstract
With the widespread use of deep neural networks (DNNs) in intelligent systems, DNN accelerators with high performance and energy efficiency are greatly demanded. As one of the feasible processing-in-memory (PIM) architectures, 3D-stacked-DRAM-based PIM (DRAM-PIM) architecture enables large-capacity memory and low-cost memory access, which is a promising solution for DNN accelerators with better performance and energy efficiency. However, the low-access-cost characteristics of stacked DRAM and the distributed manner of memory access and data storing require us to rebalance the hardware design and DNN mapping. In this paper, we propose NicePIM to efficiently explore the design space of hardware architecture and DNN mapping of DRAM-PIM-based DNN inference accelerators, which consists of three key components: PIM-Tuner, PIM-Mapper and Data-Scheduler. PIM-Tuner optimizes the hardware configurations leveraging a DNN model for classifying area-compliant PIM-node designs and a deep kernel learning model for identifying better hardware parameters. PIM-Mapper explores a variety of DNN mapping configurations, including parallelism between branches of DNN, DNN layer partitioning, DRAM capacity allocation and data layout pattern in DRAM to generate high-hardware-utilization DNN mapping schemes for various hardware configurations. The Data-Scheduler employs an integer-linear-programming-based data scheduling algorithm to alleviate the inter-PIM-node communication overhead of data-sharing brought by DNN layer partitioning. Experimental results demonstrate that NicePIM can optimize hardware configurations for DRAM-PIM systems effectively and can generate high-quality DNN mapping schemes with latency and energy cost reduced by 37% and 28% on average respectively compared to the baseline method.
Junpeng Wang 0002, Mengke Ge, Bo Ding 0004, Qi Xu 0004, Song Chen 0001, Yi Kang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Synthesizing Brain-network-inspired Interconnections for Large-scale Network-on-chips
abstract
Brain network is a large-scale complex network with scale-free, small-world, and modularity properties, which largely supports this high-efficiency massive system. In this article, we propose to synthesize brain-network-inspired interconnections for large-scale network-on-chips. First, we propose a method to generate brain-network-inspired topologies with limited scale-free and power-law small-world properties, which have a low total link length and extremely low average hop count approximately proportional to the logarithm of the network size. In addition, given the large-scale applications, considering the modularity of the brain-network-inspired topologies, we present an application mapping method, including task mapping and deterministic deadlock-free routing, to minimize the power consumption and hop count. Finally, a cycle-accurate simulator BookSim2 is used to validate the architecture performance with different synthetic traffic patterns and large-scale test cases, including real-world communication networks for the graph processing application. Experiments show that, compared with other topologies and methods, the brain-network-inspired network-on-chips (NoCs) generated by the proposed method present significantly lower average hop count and lower average latency. Especially in graph processing applications with a power-law and tightly coupled inter-core communication, the brain-network-inspired NoC has up to 70% lower average hop count and 75% lower average latency than mesh-based NoCs.
Mengke Ge, Xiaobing Ni, Qi Xu 0004, Song Chen 0001, Jinglei Huang, Yi Kang, Feng Wu 0001
ACM Trans. Design Autom. Electr. Syst.1
2020 Synthesizing A Generalized Brain-inspired Interconnection Network for Large-scale Network-on-chip Systems
abstract
Brain network is a large-scale complex network with scale-free, small-world, and modularity properties, which to a large extent supports this high-efficiency massively parallel computing system known in the world. In this paper, we propose a three-stage method to synthesize a brain-inspired interconnection network for large-scale network-on-chip systems, which minimizes communication hop count, dynamic power consumption, and energy-delay-product. Topology generation, core assignment, and routing path allocation are executed in these three stages, respectively. Experimental results show that our synthesis method can construct large-scale brain-inspired NoC systems with higher communication efficiency and superior performance compared to the state-of-the-art.
Mengke Ge, Qi Xu 0004, Huajie Ruan, Xiaobing Ni, Song Chen 0001, Yi Kang
ACM Great Lakes Symposium on VLSI1
2020 Generalized Fault-Tolerance Topology Generation for Application-Specific Network-on-Chips
abstract
The network-on-chips (NoCs)-based communication architecture is a promising candidate for addressing communication bottlenecks in many-core processors and neural network processors. In this article, we consider the generalized fault-tolerance topology generation problem, where the link (physical channel) or switch failures can happen, for application-specific NoCs (ASNoCs). With a user-defined maximum number of faults, K, we propose an integer linear programming (ILP)-based method to generate ASNoC topologies, which can tolerate at most K faults in switches or links. Given the communication requirements between cores and their floorplan, we first propose a convex-cost-flow-based method to solve a core mapping (CM) problem for building connections between the cores and switches. Second, an ILP-based method is proposed to solve the routing path allocation (PA) problem, where K+1 switch-disjoint routing paths are allocated for every communication flow between the cores. Finally, to reduce switch sizes, we propose to share the switch ports for the connections between the cores and switches and formulate the port sharing problem as a clique-partitioning problem, which is solved by iteratively finding a set of the maximum cliques. Additionally, we propose an ILP-based method to simultaneously solve the CM and routing PA problems when only physical link failures are considered. The experimental results show that the power consumption of fault-tolerance topologies increases almost linearly with K because of the routing path redundancy for fault tolerance. When both switch faults and link faults are considered, port sharing can reduce the average power consumption of fault-tolerance topologies with K = 1, K = 2, and K = 3 by 18.08%, 28.88%, and 34.20%, respectively. When considering only the physical link faults, the experimental results show that compared to the fault-tolerant topology generation (FTTG) algorithm, the proposed method reduces power consumption and hop count by 10.58% and 6.25%, respectively; compared to the de Bruijn Digraph (DBG)-based method, the proposed method reduces power consumption and hop count by 21.72% and 9.35%, respectively.
Song Chen 0001, Mengke Ge, Jinglei Huang, Qi Xu 0004, Feng Wu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2