EDBT 2026 Demo / reviewers in the wild / expert
Xin Ju 0005
dblp:227/4699-5
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0006-0875-196XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Precision boundary modeling for area-efficient Block Floating Point accumulation
Xin Ju 0005, Yasong Cao, Zhongdi Luo, Jianchao Yang, Jingkui Yang, Dong Chen 0015, Mei Wen |
J. Syst. Archit. | 3 |
| 2026 | C-CIM: A Multi-Mode Convolution-Capable SRAM-CIMabstractSRAM is widely used in computing-in-memory (CIM) neural network accelerators because of its relatively mature technology and good compatibility with complementary metal oxide semiconductor logic process. Digital SRAM-CIM is favored by researchers because of its stability and accuracy. However, the current digital SRAM-CIM macro only supports the weight-stationary dataflows, which means the repeated movement of graph data. Some special deep neural network layers, such as depth-wise, make the utilization of computing resources inside CIM low. To overcome these problems, we propose C-CIM, which can switch between input-stationary and weight-stationary dataflows and support matrix multiplication as well as convolution operations with multiple mainstream convolution kernel sizes (1×1, 3×3, 5×5 and 7×7). The C-CIM achieves an average performance of 27.31TOPS/W@8b at a frequency of 1GHz. Experimental results show that our proposed SRAM-CIM successfully outperforms baseline in terms of performance optimization, achieving up to 7.6× performance speedup and up to 86.84% reduction in activation relocation. Renyu Yang, Xin Ju 0005, Mei Wen, Jinjin Deng, Junzhong Shen, Tianyu Wang 0003, Zhaoyan Shen, Zili Shao |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | WinAcc: Window-based Acceleration of Neural Networks Using Block Floating PointabstractDeep Neural Networks (DNNs) impose significant computational demands, necessitating optimizations for computational and energy efficiencies. Per-vector scaling, which applies a scaling factor to blocks of elements using narrow integer types, effectively reduces storage and computational overhead. However, the frequent occurrence of floating-point accumulations between vectors limits further improvements in energy efficiency. State-of-the-art accelerators address this challenge by grouping and summing vector products based on their exponent differences, thereby reducing the overhead associated with intra-group shifting and accumulation. Nevertheless, this approach increases the complexity of register usage and grouping logic, leading to limited energy benefits and hardware efficiency. In this context, we introduce WinAcc, a novel algorithm and architecture co-designed solution that utilizes a low-cost accumu-lator to handle the majority of data in DNNs, offering low area overhead and high energy efficiency gains. Our key insight is that the data of DNNs follows a Laplace-like distribution, which enables the use of a customized data format with a narrow dynamic range to encode most of the data. This allows for the design of a low-cost accumulator with narrow shifters and adders, significantly reducing reliance on floating-point accumulator and consequently improving energy efficiency. Compared with state-of-the-art architecture Bucket, WinAcc achieves 33.95% energy reduction across seven representative DNNs and reduces area by 9.5% while maintaining superior model performance. Xin Ju 0005, Mei Wen, Yasong Cao, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008 |
DATE | 1 |
| 2025 | SmartBlock: Adaptive Block Floating Point Quantization for Efficient DNN AccelerationabstractDeep Neural Networks (DNNs) have achieved remarkable success as model sizes continue to grow, driving the need for optimizations in both computational and energy efficiency. Block Floating Point (BFP) quantization has emerged as an effective model compression technique, offering a favorable trade-off between model accuracy and hardware cost. However, the frequent use of floating-point (FP) accumulation across BFP blocks remains a significant bottleneck, limiting further improvements in energy efficiency. State-of-the-art (SotA) accelerators mitigate this issue by introducing low-overhead accumulators with a narrower dynamic range ahead of the FP accumulator to handle a small range of values. While this approach reduces the activation of power-hungry alignment and format conversion units, it increases the complexity of the processing elements (PEs), thereby limiting the overall energy savings. Xin Ju 0005, Jingkui Yang, Mei Wen, Minjin Tang, Zhaoyun Chen, Yang Shi 0008 |
ICPP | 1 |
| 2025 | Super Microscaling: Enhancing Precision and Hardware Efficiency in Deep Learning QuantizationabstractThe Microscaling (MX) data format, an state-of-the-art quantization technique for deep learning tensor operations, suffers from precision loss at low bit-widths and potential hardware overhead. To overcome these limitations, we propose the Super Microscaling (SMX) data format and its accompanying hardware engine. SMX ensures higher accuracy by optimizing the conversion from Floating Point (FP) to MX, particularly in low-bit width scenarios. We also present spatial-temporal reuse strategy and two-level dequantization hardware architecture, which boosts efficiency in terms of hardware area and energy consumption. Experimental results demonstrate that SMX outperforms MX, achieving average accuracy improvements of 55.86% for BERT, 40.89% for ResNet50, and 37.51% for VGG, while reducing hardware area and energy consumption by 22.74% and 25.46%, respectively. Zhuang Cao, Xin Ju 0005, Mei Wen, Yang Guo 0003 |
ISCAS | 3 |
| 2025 | CAMO: A High-Performance CIM-based Lightweight CNN Accelerator for Mobile DevicesabstractDigital Compute-in-Memory (CIM) macros revolutionize the Von Neumann architecture by significantly reducing data movements between CPU and memory. However, when dealing with lightweight CNNs with various convolution types, existing GEMM (general matrix multiplication)-oriented solutions suffer from underutilization and large activation traffic, leading to unsatisfied energy and area efficiencies. In this context, we propose CAMO, in which the key contributions are: (1) A novel convolution mapping mechanism suitable for CIM macros, that maximizes data reuse and reduces activation traffic. (2) A convolution-capable CIM macro, that also supports small-scale GEMM. (3) A CIM-based architecture that supports multiple computing modes. The experimental results show that CAMO achieves up to 31.49× performance speedup and 77.5% activation traffic reduction compared to the baseline architecture. Xin Ju 0005, Renyu Yang, Mei Wen, Junzhong Shen, Tianyu Wang 0009, Zhaoyan Shen, Zili Shao |
ISCAS | 1 |
| 2025 | A quantized network processor for mixed-precision deep learning models based on enhanced Microscaling format
Mei Wen, Xin Ju 0005, Zhuang Cao, Yang Guo 0003 |
J. Syst. Archit. | 3 |
| 2024 | Enhancing the PE Utilization for Multi-Precision Systolic Array via Optimizing Computation LatencyabstractSystolic array (SA) architectures are widely recognized as the optimal choice for Convolutional Neural Networks (CNNs). However, existing SAs suffer from reduced computational efficiency when confronted with an inadequate workload scale. Furthermore, this issue becomes even more pronounced in the accelerators that support multiple precisions. In this paper, by analyzing the under-utilization of processing element (PE), we propose a SA accelerator that optimizes computation latency for multi-precision scenarios. Considering dynamic changes in data precision, we incorporate a switching strategy to further enhance computational efficiency. Experimental results demonstrate that our proposed design only incurs a 1.203% increase in area compared to the classic approach, while achieving an average performance improvement of 20% on small-scale CNN models. Mei Wen, Xin Ju 0005, Junzhong Shen, Yang Guo 0003 |
ISCAS | 3 |
| 2024 | ABS: Accumulation Bit-Width Scaling Method for Designing Low-Precision Tensor CoreabstractA big gap exists between deep neural network (DNN) applications’ computational demand and the computing power of DNN accelerators. Low-precision floating-point (LP-FP) computation is one of the important means to improve the performance of DNN training and inference. However, the high-precision accumulators are typically applied to summating the dot products during general matrix multiplication (GEMM) in tensor cores (TCs). As the precision of data decreases, the accumulator becomes the main consumer of multiply-accumulate’s (MAC’s) area and power. Reducing the accumulators’ bit-width is of significant importance for improving the area- and energy-efficiency of TCs. There are two main challenges: 1) theoretical support on the floating-point (FP) formats with the lowest bit-width of TC’s accumulators and 2) how to integrate the LP-FP TC in the framework of DNN training and inference to evaluate its benefits. In this article, we propose accumulation bit-width scaling (ABS), a novel ABS method, to guide the design of LP-FP TCs. We 1) implement this method by constructing a novel variance retention ratio (VRR) model to predict the FP format with the minimum bit-width for TC’s accumulator; 2) provide a generator of DNN accelerator based on a systolic-array (SA) TC, supporting many low-precision configurations; and 3) design an LP-FP DNN executing framework that supports software-simulation mode and hardware-accelerator mode to run LP-FP DNN tasks. The experimental results show that the LP-FP TC guided by our ABS method has a maximum reduction of 76.47% and 75.60% in area and power consumption, respectively, compared with the advanced TCs. Yasong Cao, Mei Wen, Zhongdi Luo, Xin Ju 0005, Haolan Huang, Junzhong Shen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |