VLDB 2026 Research / reviewers in the wild / expert
Mingyu Wang 0003
dblp:65/8491-3
· DBLP profile ↗
17ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-4006-8870ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 2 first-author · 15 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NBCache: An Efficient and Scalable Non-Blocking Cache for Coherent Multi-Chiplet Systems
Zhirong Ye, Tao Lu 0012, Zhaolin Li, Zhiyi Yu, Mingyu Wang 0003 |
ASP-DAC | 7 |
| 2026 | LRM-GPU: Alleviating Synchronization Overhead for Multi-Chiplet GPU Architecture
Baiqing Zhong, Zhirong Ye, Haiqiu Huang, Zhaolin Li, Zhiyi Yu, Mingyu Wang 0003 |
HPCA | 8 |
| 2026 | PipeIMC: A Pipelined In-SRAM Computing Architecture
Yikai Cui, Renhao Fan, Weike Li, Mingyu Wang 0003, Zhaolin Li |
ISCA | 5 |
| 2026 | MAX-SM: High-Utilization Dynamic SM Partitioning for Heterogeneous Workloads on Multitasking Chiplet-Based GPUs
Mingyu Wang 0003, Tao Lu 0012, Baiqing Zhong, Zhaolin Li, Zhiyi Yu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | MagiCache: A Virtual In-Cache Computing EngineabstractThe rise of data-parallel applications poses a significant challenge to the energy consumption of computing architectures.In-cache computation is a promising solution for achieving high parallelism and energy efficiency because it can eliminate data movement between the cache and the processor.Existing in-cache computing architectures transform a portion of cache arrays into computing arrays, with all rows of these arrays serving as computing lines.The remaining cache arrays are used as cachelines to store the data required by computing arrays or processors.However, in these array-level in-cache computing architectures, only a few computing lines in each computing array are active at runtime while the others are idle, which incurs severe cache capacity loss and space underutilization.In addition, bursty memory accesses of data-parallel applications also cause significant in-cache data movement latency.To address these problems, we propose MagiCache, a virtual in-cache computing engine.First, we design a novel cacheline-level in-cache computing architecture in which each cache array can configure some rows as computing lines and the other rows as cachelines with negligible overhead.Second, a virtual engine is further designed on this novel architecture to dynamically allocate different rows of each array as computing lines or cachelines based on runtime computation and storage requirements, thus realizing efficient cacheline-level space management.Finally, we present an instruction chaining technique to overlap the bursty access latency by enabling asynchronous execution of computing arrays.Evaluation results show that MagiCache achieves a 1.19x-1.61xspeedup over the state-of-the-art in-cache computing architectures with 6.5 KB of additional storage.Our cacheline-level space * These authors contributed equally to this work. Renhao Fan, Yikai Cui, Weike Li, Mingyu Wang 0003, Zhaolin Li |
ISCA | 4 |
| 2025 | C3ache: Towards Hierarchical Cache-Centric Computing for Sparse Matrix Multiplication on GPGPUs
Mingyu Wang 0003, Baiqing Zhong, Haiqiu Huang, Guangjie Cao, Zhiyi Yu |
MICRO | 2 |
| 2025 | CINOC: Computing in Network-On-Chip With Tiled Many-Core Architectures for Large-Scale General Matrix MultiplicationsabstractLarge-scale general matrix multiplications (LMMs) are the key bottlenecks in various computation domains such as Transformer applications. However, it is a challenge to perform LMMs efficiently on traditional multi/many-core processor systems due to the large amount of memory access and the tight dependence of data transmission. By analyzing the aforementioned problems, we propose a computing in network-on-chip paradigm to perform LMMs by mitigating the performance losses caused by limited on-chip cache resources and memory bandwidth. Specifically, we propose a co-design of computable network-on-chip and the last-level cache method in tiled many-core architectures, which can reconstruct the redundant cache capacity as computable input buffer to balance the demands of computing, storage, and communication for the running LMM applications. Furthermore, a data-aware thread execution mechanism is also proposed to maximize the computational efficiency of thread streams in computable network. At the software level, memory-friendly matrix partitioning strategy, hybrid routing method and programming model are designed to bridge the gap between application demands and mismatched hardware/software interfaces. Experimental evaluations demonstrate that this proposed work achieves a computational latency reduction of 45% compared to the state-of-the-art GPU architecture, and the inference performance is improved by$2\times $of the GPT network. Mingyu Wang 0003, Jiahua Yan, Tao Lu 0012, Zhiyi Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | CCacheSim: A Circuit-Architecture Cross-Level Simulation Framework for SRAM-Based in-Cache Computing System EvaluationabstractSRAM-based Compute-In-Memory (CIM) circuits have demonstrated significant performance and energy efficiency advantages. Although numerous frameworks or tools have emerged for simulating CIM-based systems, most frameworks are tailored for specific DNN accelerators and rarely consider SRAM-CIM solutions in general processor systems because it is difficult to establish an effective mechanism to build the CIM data path to integrate the SRAM-CIM module that is tightly coupled with the cache hierarchy into the system. To address this problem, we propose a circuit-architecture cross-level simulation framework named CCacheSim for in-cache computing system. CCacheSim integrates the simulation of SRAM-CIM circuit timing and energy consumption characteristics, providing circuit-level accuracy evaluation support for in-cache computing system simulations. For the circuit level, the SRAM-CIM model can automatically generate a circuit-level netlist and conduct accurate simulation through corresponding configurations, thus balancing accuracy and agility for early design exploration. For the architectural level, to efficiently support the portable integration of SRAM-CIM module to in-cache computing system, a configurable hardware programming interface is implemented within the cache model to manage the interaction of the control stream between processor and cache for CIM tasks. Moreover, a request queue based access mechanism is proposed to ensure the completeness of the operands required by CIM tasks. To validate the proposed framework, CCacheSim is implemented to simulate varying configurations of in-cache computing systems and SRAM-CIM modules. The results prove that CCacheSim can conduct accurate performance and energy consumption evaluation for given processor architecture with given CIM module. CCacheSim can support flexible and effective design space exploration for in-cache computing system. Baiqing Zhong, Mingyu Wang 0003, Yicong Zhang, Zhiyi Yu |
ICCD | 2 |
| 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsabstractGeneral-purpose graphics processing unit (GPGPU), widely recognized as an exceptional computing platform for de-ploying emerging parallel applications, requires strict adherence to atomicity and memory consistency models for shared variable synchronization. This is crucial to ensure deterministic execution and leverage the performance advantages of the GPGPU single-instruction -multiple-threads architecture. However, the escalating demand for shared variable updates across thread blocks, notably in applications like deep neural networks and graph analysis, significantly exacerbates the serialization overhead of atomic operations due to the von Neumann bottleneck. Additionally, the overhead introduced by memory fences supporting the memory consistency model further complicates this fine-grained synchronization requirement. To address these challenges, this paper proposes Atomic Cache, facilitating an In-Cache computing hardware-software co-design for GPGPUs. At the software level, we propose relaxed memory consistency based on non-ordering commutativity to alleviate the execution of in-cache atomic operations, thereby mitigating the performance overhead of memory fences. At the hardware level, we present the In-Situ Store Atomic Cache Macro, which empowers the Atomic Cache to efficiently execute atomic logic and arithmetic operations within the cache array. This innovation alleviates the von Neumann bottleneck associated with serialized execution of atomic operations. The experimental evaluation results demonstrate that the Atomic Cache can save more than 60% of memory access energy while incurring only 9.42% chip area overhead. Furthermore, it not only delivers an average speedup ratio of 2.59 × and an IPC performance improvement of 1.48× for RISC-V GPGPUs, but also achieves an average speedup ratio of 1.31 × and an IPC performance improvement of 39.92% when compared to state-of-the-art designs employing local atomic buffers. Yicong Zhang, Mingyu Wang 0003, Wangguang Wang, Yangzhan Mai, Haiqiu Huang, Zhiyi Yu |
MICRO | 2 |
| 2023 | CRAFT: Common Router Architecture for Throughput Optimization
Jiahua Yan, Mingyu Wang 0003, Zhiyi Yu |
ICA3PP (3) | 2 |
| 2023 | LWSDP: Locality-Aware Warp Scheduling and Dynamic Data Prefetching Co-design in the Per-SM Private Cache of GPGPUsabstractGeneral Purpose Graphics Processing Units (GPG-PUs) employ frequent context switching to mask the long-latency of memory operations. However, GPGPUs still suffer from stagnation due to the incomplete overlapping of memory operations. To alleviate this stagnation and enhance Memory-Level Parallelism (MLP), it is crucial to overlap and minimize memory operations. This paper conducts a comprehensive analysis of data locality in GPGPUs and proposes an approach called Locality-Aware Warp Scheduling and Dynamic Data Prefetching (LWSDP) Co-design in the Per-SM Private Cache of GPGPUs, which effectively utilizes data locality to improve MLP. In addition to employing a coordinated scheduler and dynamic data prefetching, we incorporate Prefetching Requests Admitted Cache Access Re-execution (PRA-CAR) to mitigate the adverse impact of excessive prefetching memory requests on memory saturation. Experimental results demonstrate that LWSDP achieves an average 33.02% performance improvement and an average 28.16% miss rate reduction compared to the previous schedulers on data locality-sensitive kernels. Wangguang Wang, Mingyu Wang 0003, Yicong Zhang, Yukun Wei, Zhiyi Yu |
ICPADS | 2 |
| 2023 | A Scalable Deadlock-Free Static Routing Algorithm for Chiplet-Based SystemsabstractThe utilization of the Chiplet methodology can accelerate VLSI system development and provide better flexibility. Building interconnection networks across multiple Chiplets and ensuring high-performance deadlock-free routing in systems with diverse irregular topologies is a challenging task.To avoid the reordering introduced by adaptive routing algorithms, a scalable static deadlock-free routing algorithm specifically designed for Chiplets is proposed. This approach capitalizes on static routing, a feature that ensures consistent message order due to fixed paths, fundamentally averting reordering issues. By adaptively configuring the state of the local router, it is possible to proactively initiate turns or exit detour loops, thereby effectively preventing deadlocks. Furthermore, by employing state configuration, the routing remains deadlock-free even when there are variations in the number of Chiplets, scale, or internal topology. This showcases the system’s scalability.Due to limited wiring resources, we chose classical routing algorithms up*/down* to compare, and the results showed a significant latency advantages, with a saturation injection rate approximately 1.5 to 2 times higher. Mingyu Wang 0003, Yicong Zhang, Tao Lu 0012, Zhiyi Yu |
ICPADS | 2 |
| 2023 | A 1.97 TFLOPS/W Configurable SRAM-Based Floating-Point Computation-in-Memory Macro for Energy-Efficient AI ChipsabstractFloating-point (FP) computation-in-memory (CIM) technology is increasingly demanded by low-power neural network training. In this work, we propose an energy-efficient configurable SRAM-based FP CIM macro. A mantissa parallel alignment method is proposed to improve calculation speed and accuracy in FP multiply-accumulation (MAC) operations. The separated mantissa CIM and exponent CIM are designed to enable pipelining of exponent and mantissa operations to increase computation throughput. Furthermore, the macro can be flexibly set to BF16 or FP32 precision by configuring accumulators. The proposed FP CIM macro is analyzed in 40 nm CMOS technology, and the estimated area is 0.48 mm2, The simulation results show that the macro achieves a frequency of 294 MHz in 1.1 V. In BF16 mode, the macro can achieve a peak throughput of 56.5 GFLOPS and an energy efficiency of 1.97 TFLOPS/W while the peak throughput and energy efficiency are 16 GFLOPS and 0.62 TFLOPS/W in FP32 mode. Yangzhan Mai, Mingyu Wang 0003, Chuanghao Zhang, Baiqing Zhong, Zhiyi Yu |
ISCAS | 2 |
| 2023 | MAICC : A Lightweight Many-core Architecture with In-Cache Computing for Multi-DNN Parallel InferenceabstractThe growing complexity and diversity of neural networks in the fields of autonomous driving and intelligent robots have facilitated the research of many-core architectures, which can offer sufficient programming flexibility to simultaneously support multi-DNN parallel inference with different network structures and sizes compared to domain-specific architectures. However, due to the tight constraints of area and power consumption, many-core architectures typically use lightweight scalar cores without vector units and are almost unable to meet the high-performance computing needs of multi-DNN parallel inference. To solve the above problem, we design an area- and energy-efficient many-core architecture by integrating large amounts of lightweight processor cores with RV32IMA ISA. The architecture leverages the emerging SRAM-based computing-in-memory technology to implement vector instruction extensions by reusing memory cells in the data cache instead of conventional logic circuits. Thus, the data cache in each core can be reconfigured as the memory part and the computing part with the latter tightly coupled with the core pipeline, enabling parallel execution of the basic RISC-V instructions and the extended multi-cycle vector instructions. Furthermore, a corresponding execution framework is proposed to effectively map DNN models onto the many-core architecture by using intra-layer and inter-layer pipelining, which potentially supports multi-DNN parallel inference. Experimental results show that the proposed MAICC architecture obtains a 4.3 × throughput and 31.6 × energy efficiency over CPU (Intel i9-13900k). MAICC also achieves a 1.8 × energy efficiency over GPU (RTX 4090) with only 4MB on-chip memory and 28 mm2 area. Renhao Fan, Yikai Cui, Qilin Chen, Mingyu Wang 0003, Youhui Zhang, Zhaolin Li |
MICRO | 4 |
| 2023 | TensorCache: Reconstructing Memory Architecture With SRAM-Based In-Cache Computing for Efficient Tensor Computations in GPGPUsabstractGeneral purpose graphics processing units (GPGPUs) have emerged as a convincing and pivotal computing platform for deep learning applications. However, the fundamental tensor computations for neural networks on GPGPUs are still restricted by the von Neumann bottleneck. The memory bandwidth and energy consumption of moving a large amount of neural network data between the memory hierarchy and computational units of GPGPUs dominate the overall computational cost. To address these challenges, this article proposes TensorCache to reconstruct memory architecture with static random-access memory (SRAM)-based In-Cache Computing for efficient tensor computations in GPGPUs. It provides an innovative digital SRAM processing-in-memory (PIM) solution by transforming the cache array into large-scale PIM units, effectively mitigating the significant performance and energy consumption losses caused by data movement. To enable efficient hardware-software co-design for TensorCache, a decoupled architecture-based SRAM-PIM macro (SPM) is introduced at the hardware level, supporting in-memory bit-parallel comparison (IMBC) and near-memory radix-4 booth encoder (NRBE) for efficient mixed-precision floating-point (FP) tensor computations. At the software level, a programming model leveraging the GPGPU’s flexible programmability is proposed to bridge the gap between application demands and mismatched hardware/software interfaces. Experimental evaluations demonstrate that TensorCache achieves up to$38.59\times $speedup and$16.26\times $throughput enhancement compared to GPU CUDA Cores. Furthermore, it attains an acceleration of up to$1.78\times $and$3.87\times $throughput improvement compared to GPU Tensor Cores, while saving power consumption in tensor computations by over 90% with a mere 21% chip area overhead. Yicong Zhang, Mingyu Wang 0003, Yangzhan Mai, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | ERA-LSTM: An Efficient ReRAM-Based Architecture for Long Short-Term MemoryabstractProcessing-in-memory (PIM) architecture based on resistive random access memory (ReRAM) crossbars is a promising solution to the memory bottleneck that long short-term memory (LSTM) faces. Based on the dataflow analysis of the LSTM computing paradigm, this article proposes to adopt the ReRAM-based analog approximate computing to conduct the LSTM-specific element-wise computation. Combined with the dot-product computation implemented with ReRAM crossbars, a new LSTM processing tile is designed to significantly reduce the demand for analog-to-digital converters (ADCs), which is the major part of power consumption of existing designs. Next, we elaborate on a mapping scheme to efficiently deploy large-scale LSTM onto multiple processing tiles. Finally, an architecture enhancement is proposed to support crossbar-friendly LSTM pruning to further improve efficiency. This overall design, named ERA-LSTM, is presented. Our evaluation shows that it can outperform two state-of-the-art FPGA-based LSTM accelerators by 103.6 and 35.9 times, respectively; compared with a state-of-the-art ReRAM-based LSTM accelerator with digital element-wise computation, it is 6.1 times more efficient. Moreover, our experiments demonstrate that the impact of hardware constraints and approximation errors on the inference accuracy can be effectively reduced by the proposed fine-tuning scheme and by optimizing the design of the approximator. Jianhui Han, Mingyu Wang 0003, Zhaolin Li, Youhui Zhang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | A Spatial and Temporal Locality-Aware Adaptive Cache Design With Network Optimization for Tiled Many-Core ArchitecturesabstractThe spatial locality and the temporal locality of workloads are the root causes for cache designs to overcome the memory wall problem. However, the real memory access behavior for each of these applications can be very different. It gives the opportunities to explore further performance improvement due to different cache organization requirements. To address this issue, a spatial and temporal locality-aware adaptive cache is proposed, which dynamically partitions the private last level cache bank as prefetch region or victim region at runtime to explore the locality characteristics. The prefetch region speculates the data blocks in subsequent addresses to exploit the spatial locality, while the victim region collects the evicted data blocks from the upper memory hierarchy to exploit the temporal locality. Fast data prefetch with prioritized dynamic buffer management and adaptive burst-aware routing is realized in the proposed hybrid burst-support network-on-chip (HBNoC). By combining the adaptive cache partition with HBNoC, the off-chip misses and the on-chip network usage are greatly reduced. Experimental results demonstrate that the proposed adaptive cache design reduces up to 25% off-chip misses and improves 11.3% performance on average compared with the prior design, respectively. Mingyu Wang 0003, Zhaolin Li |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |