EDBT 2026 Demo / reviewers in the wild / expert
Yuang Ma
dblp:380/5782
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0000-7885-3562ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SLaB: Sparse-Lowrank-Binary Decomposition for Efficient Large Language Models
Yuang Ma, Yi Kang |
ISCAS | 2 |
| 2026 | VFE-CIM: An Algorithm-Hardware Co-Designed Computing-in-Memory Accelerator for Efficient Voxel Feature Encoding in Large-Scale Point Clouds
Yuang Ma, Shiyu Fan, Song Chen 0001, Yi Kang |
ISCAS | 1 |
| 2026 | I2Rec: Enabling Intra and Inter Batch Reuse in Recommendation Systems with PIM ArchitectureabstractThe Deep Learning Recommendation Model (DLRM), one of the most popular recommendation system models, faces a performance bottleneck due to its memory-bound embedding layers. In recent years, processing-in-memory (PIM) has emerged as a solution to address the “memory wall” problem. Numerous PIM-based works have been published aiming to enhance DLRM performance by exploiting data locality. However, existing methods have yet to fully capitalize on locality. To better resolve the locality issue of the embedding layer and boost the performance of DLRM, we propose I2Rec, an architecture based on PIM that can further explore the locality in a DLRM system. I2Rec employs both intra-batch and inter-batch reuse strategies, releasing the potential of inter-batch reuse. As the embedding table size grows, I2Rec can uncover more reuse opportunities so the locality can be utilized more efficiently. Compared with spatial locality methods, I2Rec avoids a long preprocessing flow and achieves better locality exploration. Experimental results show that I2Rec achieves a 1.28× speedup and reduces memory accesses by 27% compared with intra-batch reuse alone under the same cache size and up to 1.41× speedup with a little extra overhead. Additionally, I2Rec outperforms state-of-the-art spatial locality algorithms, reducing memory traffic to 40% and achieving a 2.40× improvement in performance. Shiyu Fan, Yuang Ma, Yi Kang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | HNM-CIM: An Algorithm-Hardware Co-designed SRAM-based CIM for Transformer Acceleration Exploiting Hybrid N:M SparsityabstractSRAM-based computing-in-memory (CIM) is an efficient technology for computing neural networks where matrix operations are dominated. However, leveraging sparsity in CIM presents challenges due to the crossbar architecture, which complicates the avoidance of zero element calculations. Previous CIM designs have demonstrated that sparsity can improve energy efficiency, but these approaches often lead to non-negligible accuracy loss or substantial hardware overhead. To address this challenge, we propose a hybrid N:M CIM (HNM-CIM), an algorithm-architecture co-design framework for accelerating Transformers. At the algorithm level, we propose a hybrid N:M pruning (HNMP), a method that combines structured and unstructured sparsity. This approach maintains regularity while preserving the random distributions of sparsity, thereby enhancing model sparsity with negligible accuracy loss and ensuring CIM compatibility. At the hardware level, we introduce a hybrid N:M sparse digital CIM (HNM-CIM) to support HNMP, which can accelerate Transformers with hybrid N:M sparsity patterns. Experimental results show that HNMP can reduce Transformer models by about 3.1× on model size with negligible accuracy loss. Compared with state-of-the-art references, HNM-CIM yields about 2.46× speed up and 1.43× area savings. Yuang Ma, Yulong Meng, Zihao Xuan, Song Chen 0001, Yi Kang |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2025 | SysCIM: A Heterogeneous Chip Architecture for High-Efficiency CNN Training at EdgeabstractNeural network training is notoriously computationally intensive and time-consuming. Quantization technology is promising to improve training efficiency by using lower data bitwidths to reduce storage and computing requirements. Currently, state-of-the-art quantization training algorithms have a negligible loss of accuracy, which requires dedicated quantization circuits for dynamic quantization of large amounts of data. In addition, the matrix transposition problem during neural network training gradually becomes a challenge as the network size increases. To address this problem, we propose a quantized training architecture which is a heterogeneous architecture consisting of a computing-in-memory (CIM) macro and a systolic array. First, the CIM macro realizes efficient transpose matrix multiplication through flexible data path control, which handles the need for transpose operation of the weight matrix in neural network training. Second, the systolic array utilizes two different data flows in the forward (FW) and backward (BW) propagation for the transpose matrix multiplication of the activation matrix in neural network training and provides higher computational throughput. Then, we design efficient dedicated quantization circuits for quantization algorithms to support efficient quantization training. Experimental results show that the area and power consumption of the two specialized quantization circuits are reduced by a factor of 1.35 and 5.4, on average, compared to floating-point computing circuits. The architecture achieves 4.05 tera operations per second per wat (TOPS/W) energy efficiency @ INT8 convolutional neural network (CNN) training at the 28-nm process. Compared to a state of the art (SOTA) quantization training architecture, SysCIM shows$1.8\times $energy efficiency. Shuai Wang 0040, Yuang Ma, Yi Kang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | AFT-CIM: An Energy Efficient ADC-Free Transpose Computing-in-Memory Macro for MAC OperationsabstractComputing-in-memory (CIM) based on SRAM is a promising technique to implement energy-efficient matrix computing in artificial intelligence (AI) edge devices. The ability to support both inference and training on a single macro is desired for AI edge devices, while most existing SRAM-based CIM macros only support inference. In this paper, we propose an energy-efficient and ADC-free transpose CIM (AFT-CIM) macro that can support both inference and training on a single macro. First, we introduce a computational circuit with switchable row or column inputs, which not only maintains input flexibility but also simplifies the circuits. Furthermore, we propose an orthogonal path adder tree (OPAT) that achieves flexible switching between two computation paths: row-wise accumulation and column-wise accumulation. The different combinations of input and accumulation directions enable different data computation paths on the proposed AFT-CIM macro. Different computations required in DNN training, such as matrix multiplication or transpose matrix multiplication, can be realized by flexibly controlling the data flow of the AFT-CIM macro. A 32Kb SRAM CIM macro is designed using 28 nm CMOS technology. The circuit-level evaluation shows that the power consumption of the OPAT circuit is reduced by 1.7× and the area overhead is reduced by 1.3× compared to the design using two separate adder trees. The peak energy efficiency of the CIM macro reaches 62.9 TOPS/W. Compared to SOTA transpose CIM macros, AFT-CIM shows 2.3× to 3.9× energy efficiency. Shuai Wang 0040, Yuang Ma, Yi Kang |
ISCAS | 2 |