EDBT 2026 Demo / reviewers in the wild / expert
Yangzhan Mai
dblp:352/9936
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2024
0009-0008-3895-4340ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 71% GPUs and heterogeneous computing · 18% Parallel and multicore computing · 5% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache design |
0.8 | 1 | 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic Operations · MICRO 2024 |
GPUs and heterogeneous computing
GPU computing |
0.8 | 1 | 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic Operations · MICRO 2024 |
Memory systems › processing-in-memory › computing-in-memory › in-SRAM computing
in-cache computing |
0.8 | 1 | 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic Operations · MICRO 2024 |
Memory systems › memory consistency
memory consistency model |
0.8 | 1 | 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic Operations · MICRO 2024 |
Memory systems › memory consistency › memory consistency model
weak memory model |
0.8 | 1 | 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic Operations · MICRO 2024 |
Processor architecture and microarchitecture
atomic operations |
0.2 | 1 | 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic Operations · MICRO 2024 |
Parallel and multicore computing › synchronization
fine-grain synchronization |
0.2 | 1 | 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic Operations · MICRO 2024 |
Methods — techniques the papers use, named apart from their topics
in-situ store atomic cache macro · 0.8hardware-software co-design · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsabstractGeneral-purpose graphics processing unit (GPGPU), widely recognized as an exceptional computing platform for de-ploying emerging parallel applications, requires strict adherence to atomicity and memory consistency models for shared variable synchronization. This is crucial to ensure deterministic execution and leverage the performance advantages of the GPGPU single-instruction -multiple-threads architecture. However, the escalating demand for shared variable updates across thread blocks, notably in applications like deep neural networks and graph analysis, significantly exacerbates the serialization overhead of atomic operations due to the von Neumann bottleneck. Additionally, the overhead introduced by memory fences supporting the memory consistency model further complicates this fine-grained synchronization requirement. To address these challenges, this paper proposes Atomic Cache, facilitating an In-Cache computing hardware-software co-design for GPGPUs. At the software level, we propose relaxed memory consistency based on non-ordering commutativity to alleviate the execution of in-cache atomic operations, thereby mitigating the performance overhead of memory fences. At the hardware level, we present the In-Situ Store Atomic Cache Macro, which empowers the Atomic Cache to efficiently execute atomic logic and arithmetic operations within the cache array. This innovation alleviates the von Neumann bottleneck associated with serialized execution of atomic operations. The experimental evaluation results demonstrate that the Atomic Cache can save more than 60% of memory access energy while incurring only 9.42% chip area overhead. Furthermore, it not only delivers an average speedup ratio of 2.59 × and an IPC performance improvement of 1.48× for RISC-V GPGPUs, but also achieves an average speedup ratio of 1.31 × and an IPC performance improvement of 39.92% when compared to state-of-the-art designs employing local atomic buffers. Yicong Zhang, Mingyu Wang 0003, Wangguang Wang, Yangzhan Mai, Haiqiu Huang, Zhiyi Yu |
MICRO | 4 |
| 2023 | A 1.97 TFLOPS/W Configurable SRAM-Based Floating-Point Computation-in-Memory Macro for Energy-Efficient AI ChipsabstractFloating-point (FP) computation-in-memory (CIM) technology is increasingly demanded by low-power neural network training. In this work, we propose an energy-efficient configurable SRAM-based FP CIM macro. A mantissa parallel alignment method is proposed to improve calculation speed and accuracy in FP multiply-accumulation (MAC) operations. The separated mantissa CIM and exponent CIM are designed to enable pipelining of exponent and mantissa operations to increase computation throughput. Furthermore, the macro can be flexibly set to BF16 or FP32 precision by configuring accumulators. The proposed FP CIM macro is analyzed in 40 nm CMOS technology, and the estimated area is 0.48 mm2, The simulation results show that the macro achieves a frequency of 294 MHz in 1.1 V. In BF16 mode, the macro can achieve a peak throughput of 56.5 GFLOPS and an energy efficiency of 1.97 TFLOPS/W while the peak throughput and energy efficiency are 16 GFLOPS and 0.62 TFLOPS/W in FP32 mode. Yangzhan Mai, Mingyu Wang 0003, Chuanghao Zhang, Baiqing Zhong, Zhiyi Yu |
ISCAS | 1 |
| 2023 | TensorCache: Reconstructing Memory Architecture With SRAM-Based In-Cache Computing for Efficient Tensor Computations in GPGPUsabstractGeneral purpose graphics processing units (GPGPUs) have emerged as a convincing and pivotal computing platform for deep learning applications. However, the fundamental tensor computations for neural networks on GPGPUs are still restricted by the von Neumann bottleneck. The memory bandwidth and energy consumption of moving a large amount of neural network data between the memory hierarchy and computational units of GPGPUs dominate the overall computational cost. To address these challenges, this article proposes TensorCache to reconstruct memory architecture with static random-access memory (SRAM)-based In-Cache Computing for efficient tensor computations in GPGPUs. It provides an innovative digital SRAM processing-in-memory (PIM) solution by transforming the cache array into large-scale PIM units, effectively mitigating the significant performance and energy consumption losses caused by data movement. To enable efficient hardware-software co-design for TensorCache, a decoupled architecture-based SRAM-PIM macro (SPM) is introduced at the hardware level, supporting in-memory bit-parallel comparison (IMBC) and near-memory radix-4 booth encoder (NRBE) for efficient mixed-precision floating-point (FP) tensor computations. At the software level, a programming model leveraging the GPGPU’s flexible programmability is proposed to bridge the gap between application demands and mismatched hardware/software interfaces. Experimental evaluations demonstrate that TensorCache achieves up to$38.59\times $speedup and$16.26\times $throughput enhancement compared to GPU CUDA Cores. Furthermore, it attains an acceleration of up to$1.78\times $and$3.87\times $throughput improvement compared to GPU Tensor Cores, while saving power consumption in tensor computations by over 90% with a mere 21% chip area overhead. Yicong Zhang, Mingyu Wang 0003, Yangzhan Mai, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |