EDBT 2026 Demo / reviewers in the wild / expert
Tianshuo Bai
dblp:349/1978
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-3966-1442ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Operator-Circuit Co-design Digital SOT-MRAM Computing-in-Memory Accelerator with Double Bit Density and Full-Utilized Bandwidth/ThroughputabstractComputing-in-Memory (CIM) demonstrates exceptional performance on edge AI applications, owing to its in-situ computation capability with minimal data transfer consumption. However, volatile CIMs suffer from inevitable data retention power overhead, while non-volatile MRAM-CIMs still necessitate periodic weight updates constrained by limited memory space, diminishing the intrinsic advantage of CIMs. In this work, we propose a digital SOT-MRAM CIM accelerator with circuit-architecture-operator cross-layer design, achieving double bit density and full utilization of both data transmission bandwidth and computing throughput, thereby satisfying the stringent hardware demands for edge AI applications. Firstly, we propose a refined 2T-1MTJ non-complementary memory cell with an XOR-integrated pre-charged sense amplifier (X-SA), which significantly promotes the storage density and consumes only 6.284 fJ per read-based XOR operation. Then, we devise a channel-flatten data mapping (CFDM) scheme and an operator-aware residual fusion (OARF) structure to full utilize the storage and computing resources. Furthermore, an operator fusion method towards non-linear layers is proposed, achieving an 89.84% size reduction in non-binary parameters. System-level simulations at 40nm demonstrate that our work achieves 284.25 TOPS/W energy efficiency and 5.41 TOPS/mm2area efficiency with an accuracy of 98.72% (87.78%) on MNIST (CIFAR-10) dataset. Tianshuo Bai, Jingcheng Gu, Lehao Tan, Wente Yi, Haolin Ge, Zhenyu Xue, He Zhang 0011, Na Lei, Biao Pan |
DATE | 1 |
| 2025 | HyIMC: Analog-Digital Hybrid In-Memory Computing SoC for High-Quality Low-Latency Speech EnhancementabstractIn-memory computing (IMC) holds significant promise for accelerating deep learning-based speech enhancement (DL-SE). However, existing IMC architectures face challenges in simultaneously achieving high precision, energy efficiency, and the necessary parallelism for DL-SE's inherent temporal dependencies. This paper introduces HyIMC, a novel hybrid analog-digital IMC architecture designed to address these limitations. HyIMC features: 1) a hybrid analog-digital design optimized for DL-SE algorithms; 2) a schedule controller that efficiently manages recurrent dataflow within skip connections; and 3) non-key dimension shrinkage, a model compression technique that preserves accuracy. Implemented on a 40nm eFlash-based IMC SoC prototype, HyIMC achieves 160 TOPS/W energy efficiency, compresses the DL-SE model size by ~600%, improves the feature of merit by ~1200%, and enhances perceptual evaluation of speech quality by ~120%. Wanru Mao, Guangyao Wang, Tianshuo Bai, Jingcheng Gu, Xitong Yang, Aifei Zhang, Xiaohang Wei, Wang Kang 0001 |
DATE | 4 |
| 2025 | PAR-CIM: A Precise/Approximate Reconfigurable Digital CIM Macro with 0.35-4b Fractional Mixed-Bitwidth QuantizationabstractDigital computing-in-memory (DCIM) enables efficient deep neural networks (DNNs) acceleration but faces limitations in resource overhead, energy efficiency, and architectural flexibility. Existing approximate or reconfigurable DCIM solutions tackle these issues partially without achieving a holistic balance. To address this, we propose PAR-CIM, a highly energy-efficient reconfigurable CIM macro that integrates precise and approximate paradigm. First, we introduce layer/gate-level approximate computation (LGAC) into the adder tree (AT) of the DCIM core, achieving full operation with only 0.35× the area of traditional implementations. Then, we develop a 0.35-4b fractional mixed-bitwidth quantization (FMBQ) algorithm, combining second-order Taylor sensitivity analysis with DoReFa-Net. This is complemented by a high-precision low-approximation (HPLA) mapping scheme to enhance energy efficiency. Additionally, a multi-bit reconfigurable computation mode (MBRM) strategy further improves architectural flexibility and enables the implementation of the proposed design. Under 40nm technology, PAR-CIM achieves 3048 TOPS/W at 1b/1b operations. With FMBQ, ResNet18 and our custom V-FuseMBA trained on CIFAR-10 achieve over 86.61% compression with accuracy loss under 0.74%, reaching classification accuracies of 93.67% and 92.86%, respectively. Zhenyu Xue, Wente Yi, Tianshuo Bai, Lehao Tan, Jingcheng Gu, Weijie Ding, Wang Kang 0001, Biao Pan |
ICCAD | 4 |
| 2025 | TongueBCI: An Interaction Method Based on EEG Signals from Tongue Movement Direction
Dingming Tan, Zifeng Ni, Baiqiao Zhang, Chao Zhou 0012, Tianshuo Bai, Juan Liu 0008, Xiangxian Li, Yulong Bian |
ICXR | 5 |
| 2024 | An End-to-End In-Memory Computing System Based on a 40-nm eFlash-Based IMC SoC: Circuits, Toolchains, and Systems Co-Design FrameworkabstractDespite its promising potential for Artificial Intelligence (AI) applications, current In-Memory Computing (IMC) technology faces a variety of challenges before mass production. One of the major challenges we face is the absence of efficient toolchains for deploying canonical networks on IMC chips. To address this issue, we propose a co-designed framework that integrates circuit, toolchain, and system elements specifically for IMC. More specifically, our framework consists of several key techniques to improve the key performance including (a) an 8-bit hardware-friendly Quantization-Aware Training (QAT) approach to quantify the deep learning network from floating-point data to fixed-point data, (b) a novel operator optimization technique to increase the computing precision when running the algorithm models on the IMC chips, and (c) an efficient mapping strategy based on the Integer Linear Programming (ILP) approach to improve the computation resource utilization of the IMC array. We assess our method on our 40nm eFlash-based IMC SoC chip with voice recognition, speech noise reduction, and person detection tasks. Our experimental results show an accuracy over 94.60% in a quiet environment and 87.27% in a white noise environment and a false recognition rate below 1 time per 24 hours for voice recognition, a 21.53% improvement for the Perceptual Evaluation of Speech Quality (PESQ) for noise reduction, and a 97.80% accuracy in person detection. Tianshuo Bai, Wanru Mao, Guangyao Wang, Aifei Zhang, Shihang Fu, Shuaikai Liu, Jianchao Hu, Xitong Yang, Biao Pan, Wei W. Xing, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |