VLDB 2026 Research / reviewers in the wild / expert
Yichuan Bai
dblp:57/9544
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-8895-1780ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-DeviceabstractEdge deployment of low-batch large language models (LLMs) faces critical memory bandwidth bottlenecks when executing memory-intensive general matrix-vector multiplications (GEMV) operations. While digital processing-in-memory (PIM) architectures promise to accelerate GEMV operations, existing PIM-equipped edge devices still suffer from three key limitations: limited bandwidth improvement, component under-utilization in mixed workloads, and low compute capacity of computing units (CUs). In this paper, we propose CD-PIM to address these challenges through three key innovations. First, we introduce a high-bandwidth compute-efficient mode (HBCEM) that enhances bandwidth by dividing each bank into four pseudo-banks through segmented global bitlines. Second, we propose a low-batch interleaving mode (LBIM) to improve component utilization by overlapping GEMV operations with GEMM operations. Third, we design a compute-efficient CU that performs enhanced GEMV operations in a pipelined manner by serially feeding weight data into the computing core. Forth, we adopt a column-wise mapping for the key-cache matrix and row-wise mapping for the value-cache matrix, which fully utilizes CU resources. Our evaluation shows that compared to a GPU-only baseline and state-of-the-art PIM designs, our CD-PIM achieves 11.42× and 4.25× speedup on average within a single batch in HBCEM mode, respectively. Moreover, for low-batch sizes, the CD-PIM achieves an average speedup of 1.12× in LBIM compared to HBCEM. Chao Fang 0005, Xiaoyong Song, Anying Jiang, Yichuan Bai |
DATE | 6 |
| 2026 | A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP SystemsabstractMixture-of-Experts (MoE) models facilitate edge deployment by decoupling model capacity from active computation, yet their large memory footprint drives the need for GPU systems with near-data processing (NDP) capabilities that offload experts to dedicated processing units. However, deploying MoE models on such edge-based GPU-NDP systems faces three critical challenges: 1) severe load imbalance across NDP units due to non-uniform expert selection and expert parallelism, 2) insufficient GPU utilization during expert computation within NDP units, and 3) extensive data pre-profiling necessitated by unpredictable expert activation patterns for pre-fetching. To address these challenges, this paper proposes an efficient inference framework featuring three key optimizations. First, the underexplored tensor parallelism in MoE inference is exploited to partition and compute large expert parameters across multiple NDP units simultaneously towards edge low-batch scenarios. Second, a load-balancing-aware scheduling algorithm distributes expert computations across NDP units and GPU to maximize resource utilization. Third, a dataset-free pre-fetching strategy proactively loads frequently accessed experts to minimize activation delays. Experimental results show that our framework enables GPU-NDP systems to achieve 2.41× on average and up to 2.56× speedup in end-to-end latency compared to state-of-the-art approaches, significantly enhancing MoE inference efficiency in resource-constrained environments. Chao Fang 0005, Yichuan Bai, Yuan Du |
DATE | 6 |
| 2026 | A Compilation Framework for Domain Specific Multi-Chiplet Systems with Double-Layer Genetic Algorithms
Yichuan Bai, Yuan Du |
ISCAS | 2 |
| 2025 | FBQuant: FeedBack Quantization for Large Language ModelsabstractDeploying Large Language Models (LLMs) on edge devices is increasingly important, as it eliminates reliance on network connections, reduces expensive API calls, and enhances user privacy. However, on-device deployment is challenging due to the limited computational resources of edge devices. In particular, the key bottleneck stems from memory bandwidth constraints related to weight loading. Weight-only quantization effectively reduces memory access, yet often induces significant accuracy degradation. Recent efforts to incorporate sub-branches have shown promise for mitigating quantization errors, but these methods either lack robust optimization strategies or rely on suboptimal objectives. To address these gaps, we propose FeedBack Quantization (FBQuant), a novel approach inspired by negative feedback mechanisms in automatic control. FBQuant inherently ensures that the reconstructed weights remain bounded by the quantization process, thereby reducing the risk of overfitting. To further offset the additional latency introduced by sub-branches, we develop an efficient CUDA kernel that decreases 60% of extra inference time. Comprehensive experiments demonstrate the efficiency and effectiveness of FBQuant across various LLMs. Notably, for 3-bit Llama2-7B, FBQuant improves zero-shot accuracy by 1.2%. Yijiang Liu, Hengyu Fang, Liulu He, Rongyu Zhang, Yichuan Bai, Yuan Du |
IJCAI | 5 |
| 2025 | BE-NPU: A Bandwidth-Efficient Neural Processing Unit With Adaptive Processing Schemes for Reduced Off-Chip Bandwidth DemandabstractExisting neural processing units (NPUs) mainly focus on the optimized multiply-accumulate (MAC) arrays for efficient inference of convolutional neural networks (CNNs). However, off-chip data transmission usually keeps NPUs waiting during CNN inference, causing up to 38.4GB/s off-chip bandwidth (OCB) demand for mobile AI devices. And none of the previous benchmarks quantitatively evaluate the bandwidth efficiency of different NPU architectures. In addition, CNNs exhibit distinct characteristics of off-chip data transmission when applied to different fields, and it has become a challenging task for NPUs to support different CNNs efficiently with reasonable OCB demand. To address the aforementioned issues, this paper proposes the Bandwidth-Peak Performance Ratio for n percentages of ideal frame rate (BPPR-n%) to demonstrate the normalized OCB demand of different NPU architectures. A bandwidth-efficient NPU (BE-NPU) is introduced with adaptive processing schemes to reduce the OCB demand during inference of different CNNs. The adaptive processing schemes include both instruction-level and thread-level schemes. For the instruction-level scheme, decoupled execute/access is introduced into depth-first (DF) and layer-first (LF) schemes to improve the concurrency between NPU calculation (CAL) and direct memory access (DMA) instructions. For the thread-level scheme, DF and LF threads are hybridly processed to further improve overall NPU efficiency. Compared with state-of-the-art works, BE-NPU achieves 48.1%~80.6% reduction of BPPR-80% and 67.0%~95.1% reduction of BPPR-95%. The proposed architecture is synthesized with TSMC 28nm technology node. BE-NPU utilizes 14.3% additional logic gates compared with baseline implementation. Yichuan Bai, Yaqing Li, Yuan Du |
IEEE Trans. Computers | 1 |
| 2025 | A Hybrid-Structured Lossless Compression-Decompression Engine for Intermediate Feature Maps in Vision Neural NetworksabstractWith the continuous evolution of vision neural networks, the off-chip transmission and storage of intermediate feature maps have become a major bottleneck during inference, especially in resource-constrained edge devices. Lossless compression of the intermediate feature maps exhibits the plug-and-play feature without requiring additional evaluation or retraining. However, previous lossless compression hardware engines face challenges in the tradeoff between hardware complexity and compression ratio. To address this issue, this article proposes a hybrid-structured lossless compression and decompression engine of intermediate feature maps in vision neural networks. The proposed work combines multidimensional prediction, run-length encoding (RLE), and extended encoding (EE). In the predictive stage, delta and contextual prediction are employed to enhance sparsity for compression efficiency. In the encoding stage, RLE minimizes redundancy between adjacent bytes, and EE is used to select various compression methods to decrease bit-level redundancy dynamically. The average compression ratio of widely used vision neural networks is 48.42%, which is better than that of Huffman coding. Compared with the state-of-the-art works, it improves the average compression ratio by 20.90%. For hardware implementation, hardware reuse and data reuse are employed to achieve 17.95% and 21.18% reduction of the gate count and the power, respectively. Experimental results show that we achieve a throughput-per-area of 1.10 [(bits/cycle)/K GCs] and a throughput-per-power of 1.19 [(bits/cycle)/mW] in the 28-nm process node. Junyong Hua, Yichuan Bai, Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | Optoelectronic Computing Evaluation and Deployment Platform Based on a 256-MAC Silicon Photonic ChipabstractThe deceleration of Moore's Law has led to increasing difficulties in advancing the computational speed and power efficiency of Complementary-Metal-Oxide-Semiconductor (CMOS) chips. As a solution to this challenge, optical computing emerges as a promising technology, boasting low energy consumption, high processing speed, and extensive bandwidth. Yet, a critical obstacle remains: the absence of a co-simulation platform that incorporates both photonic chips and peripheral electrical circuits. This paper addresses this gap by introducing a hybrid optoelectronic computing evaluation and deployment platform utilizing Simulink tools. Based on the measured data from the silicon optical computing chip, we have deployed an image filtering algorithm and a convolutional neural network onto this platform. The optical computing chip achieves an accuracy of 86.4% on the ImageNet image dataset. Through evaluation, we have identified the most substantial impacts on calculation results. To achieve an image classification accuracy of 80%, the signal-to-noise ratio (SNR) of the low-speed DAC must be a minimum of 52 dB. These findings provide crucial insights into the optimization of optical computing systems. Likai Li, Yichuan Bai, Shengping Liu, Sunan He, Yaqing Li, Yuan Du |
ISCAS | 2 |
| 2024 | Low-Latency PAE: Permutation-Based Address Encryption Hardware Engine for IoT Real-Time Memory ProtectionabstractIn Internet of Things (IoT) endpoint devices, some data or address ciphers are used for real-time memory protection to mitigate some side-channel attacks against memories. To better meet the requirements of real-time memory protection, this article proposes a hardware engine of permutation-based address encryption (PAE) to implement memory address encryption with flexible width adaptation, low latency, and low hardware overhead. When evaluated with TSMC’s 40-nm standard CMOS technology, PAE features lightweight characteristics with a gate count of 0.589 KGates, which is only 0.37% of advanced encryption standard (AES) and 33.50% of address cipher Galois field encryption (GF-Enc). The security of PAE in memory protection is quantitatively proven through both logic cryptanalysis and side-channel attacks. The results show that PAE performs effective mitigation in some side-channel attacks and provides better security than other address ciphers in resisting the brute-force attack, chosen-plaintext attack, and the differential attack. A RISC-V system with PAE and AES is deployed on an field-programmable gate array platform to analyze the impact on performance. The evaluation data show that PAE has no impact on system throughput in the case analysis, while AES reduces system throughput by 89.47%. Xuewen He, Yichuan Bai, Zhongfeng Wang 0001, Yuan Du |
IEEE Internet Things J. | 2 |
| 2024 | A Compilation Framework for SRAM Computing-in-Memory Systems With Optimized Weight Mapping and Error CorrectionabstractDeploying convolution-based algorithms into SRAM computing-in-memory (CIM) systems faces various challenges, such as operator incompatibility and intrinsic non-ideal error. This paper proposes a compilation framework to address this issue. Efficient weight mapping strategies are introduced to improve the utilization of SRAM-CIM macro. The intrinsic non-ideal errors of SRAM-CIM macro are also taken into consideration, and two efficient error correction schemes are proposed, which include calibration of computation voltage linear error (CCVLE) and the mitigation of analog-to-digital quantization error (MAQE). In addition, bit-width flexibility and signed-unsigned reconfigurability are also supported to facilitate the deployment of various convolution-based algorithms. ResNet18, finite impulse response (FIR) filtering, and Gaussian image filtering are deployed into a multi-macro SRAM-CIM system. These algorithms serve as deployment representatives of convolutional neural network (CNN), digital signal processing (DSP), and digital image processing (DIP), respectively. The results show that the introduced weight mapping strategies improve the macro utilization by 63.29% and 21.10% for two types of frequently used convolution layers compared to the traditional strategy. Moreover, the proposed error correction schemes achieve similar algorithm accuracy to the floating-point results, and the deployment result of ResNet18 achieves 66.3%~70.1% top-1 classification accuracy evaluated on the ImageNet dataset with different throughput tradeoffs. Yichuan Bai, Yaqing Li, Heng Zhang 0024, Aojie Jiang, Yuan Du |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | An Efficient High-Throughput Structured-Light Depth EngineabstractIn this article, an efficient high-throughput depth engine is proposed to generate high-quality 3-D depth maps for speckle-pattern structured-light depth cameras. A dynamic-binarization (DB) method is introduced with a significant reduction of computational complexity in contrast to the sum-of-absolute-distance (SAD) method. The depth map evaluation shows good robustness compared with other window-based correlation methods. Parallel architecture and reuse of intermediate results are employed for efficient hardware implementation. Our design is verified on a field-programmable gate array (FPGA) and implemented in the SMIC 55-nm CMOS technology, achieving a frame rate of 1731.77 fps ($640\times480$) with an area efficiency of 3.75 fps/KGE. The proposed engine shows a$2.71\times $promotion of area efficiency in contrast to the SAD-based implementation. In addition, the subpixel estimation algorithm deployed in postprocessing is optimized for efficient hardware implementation, reducing the gate count by 69.2% without significant performance loss. Yichuan Bai, Mingzhe Jiang, Qingyu Zhu, Yuan Du, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |