Baiqing Zhong

dblp:352/8734 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0001-7455-815XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 LRM-GPU: Alleviating Synchronization Overhead for Multi-Chiplet GPU Architecture
Baiqing Zhong, Zhirong Ye, Haiqiu Huang, Zhaolin Li, Zhiyi Yu, Mingyu Wang 0003
HPCA1
2026 MAX-SM: High-Utilization Dynamic SM Partitioning for Heterogeneous Workloads on Multitasking Chiplet-Based GPUs
Mingyu Wang 0003, Tao Lu 0012, Baiqing Zhong, Zhaolin Li, Zhiyi Yu
IEEE Trans. Parallel Distributed Syst.4
2025 C3ache: Towards Hierarchical Cache-Centric Computing for Sparse Matrix Multiplication on GPGPUs
Mingyu Wang 0003, Baiqing Zhong, Haiqiu Huang, Guangjie Cao, Zhiyi Yu
MICRO3
2024 CCacheSim: A Circuit-Architecture Cross-Level Simulation Framework for SRAM-Based in-Cache Computing System Evaluation
abstract
SRAM-based Compute-In-Memory (CIM) circuits have demonstrated significant performance and energy efficiency advantages. Although numerous frameworks or tools have emerged for simulating CIM-based systems, most frameworks are tailored for specific DNN accelerators and rarely consider SRAM-CIM solutions in general processor systems because it is difficult to establish an effective mechanism to build the CIM data path to integrate the SRAM-CIM module that is tightly coupled with the cache hierarchy into the system. To address this problem, we propose a circuit-architecture cross-level simulation framework named CCacheSim for in-cache computing system. CCacheSim integrates the simulation of SRAM-CIM circuit timing and energy consumption characteristics, providing circuit-level accuracy evaluation support for in-cache computing system simulations. For the circuit level, the SRAM-CIM model can automatically generate a circuit-level netlist and conduct accurate simulation through corresponding configurations, thus balancing accuracy and agility for early design exploration. For the architectural level, to efficiently support the portable integration of SRAM-CIM module to in-cache computing system, a configurable hardware programming interface is implemented within the cache model to manage the interaction of the control stream between processor and cache for CIM tasks. Moreover, a request queue based access mechanism is proposed to ensure the completeness of the operands required by CIM tasks. To validate the proposed framework, CCacheSim is implemented to simulate varying configurations of in-cache computing systems and SRAM-CIM modules. The results prove that CCacheSim can conduct accurate performance and energy consumption evaluation for given processor architecture with given CIM module. CCacheSim can support flexible and effective design space exploration for in-cache computing system.
Baiqing Zhong, Mingyu Wang 0003, Yicong Zhang, Zhiyi Yu
ICCD1
2023 A 1.97 TFLOPS/W Configurable SRAM-Based Floating-Point Computation-in-Memory Macro for Energy-Efficient AI Chips
abstract
Floating-point (FP) computation-in-memory (CIM) technology is increasingly demanded by low-power neural network training. In this work, we propose an energy-efficient configurable SRAM-based FP CIM macro. A mantissa parallel alignment method is proposed to improve calculation speed and accuracy in FP multiply-accumulation (MAC) operations. The separated mantissa CIM and exponent CIM are designed to enable pipelining of exponent and mantissa operations to increase computation throughput. Furthermore, the macro can be flexibly set to BF16 or FP32 precision by configuring accumulators. The proposed FP CIM macro is analyzed in 40 nm CMOS technology, and the estimated area is 0.48 mm2, The simulation results show that the macro achieves a frequency of 294 MHz in 1.1 V. In BF16 mode, the macro can achieve a peak throughput of 56.5 GFLOPS and an energy efficiency of 1.97 TFLOPS/W while the peak throughput and energy efficiency are 16 GFLOPS and 0.62 TFLOPS/W in FP32 mode.
Yangzhan Mai, Mingyu Wang 0003, Chuanghao Zhang, Baiqing Zhong, Zhiyi Yu
ISCAS4