VLDB 2026 Research / reviewers in the wild / expert
Xin Si
dblp:216/3686
· DBLP profile ↗
32ranked-venue papers
1as first author
30since 2021 · last 2026
0000-0002-4993-0087ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 1 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-explorationabstractAs an emerging type of AI computing accelerator, SRAM Computing-In-Memory (CIM) accelerators feature high energy efficiency and throughput. However, various CIM designs and under-explored mapping strategies impede the full exploration of compute and storage balancing in SRAM-CIM accelerator, potentially leading to significant performance degradation. To address this issue, we propose CIM-Tuner, an automatic tool for hardware balancing and optimal mapping strategy under area constraint via hardware-mapping co-exploration. It ensures universality across various CIM designs through a matrix abstraction of CIM macros and a generalized accelerator template. For efficient mapping with different hardware configurations, it employs fine-grained two-level strategies comprising accelerator-level scheduling and macro-level tiling. Compared to prior CIM mapping, CIM-Tuner’s extended strategy space achieves 1.58× higher energy efficiency and 2.11× higher throughput. Applied to SOTA CIM accelerators with identical area budget, CIM-Tuner also delivers comparable improvements. The simulation accuracy is silicon-verified and CIM-Tuner tool is open-sourced at https://github.com/champloo2878/CIM-Tuner.git. Jinwu Chen, He Wang 0028, Zhe Jiang 0004, Jun Yang 0006, Xin Si, Zhenhua Zhu 0002 |
DATE | 6 |
| 2026 | A Highly-Scalable and Full-Connected SOT P-Bit Ising Annealer for Combinatorial Optimization
Huanghui Wang, Yantong Di, Zeying Ding, Jiongzhe Su, Jingchao Zhang, Xin Si, Bo Liu 0019, Hao Cai 0001 |
ISCAS | 8 |
| 2026 | Taming polarized fitting: BLINEX-Pcomp with asymmetric risk penalty for robust Pcomp classification
Xin Si, Yingjie Tian 0001, Panos M. Pardalos |
Neural Networks | 2 |
| 2026 | An LSGQ-FFS Framework for Adaptive Optimization of Hybrid INT-CIM ArchitectureabstractHybrid computing-in-memory (CIM) has recently gained significant attention due to its ability to leverage the strengths of both digital CIM (DCIM) and analog CIM (ACIM). The multibit fusion (MF) scheme enhances energy efficiency by fusing low-bit results, which typically require multiple read-out cycles, into a single-cycle read out. However, the relationship between hybrid INT-CIM circuit design and network performance based on the MF scheme has not yet been systematically explored. In addition, we investigate how different MF configurations affect the performance of various neural networks. To address this gap, we first propose a less-significant group quantization (LSGQ) model, which defines and explores the design space of hybrid INT-CIM. Second, we develop a FastFuse-Search (FFS) algorithm, which optimizes configurations for different networks to strike a better balance between model accuracy and energy efficiency. Based on the experimental results, some key considerations on hybrid CIM design are derived. FFS yields a$1.72\times $energy-efficiency boost with negligible accuracy loss. Finally, we fabricate a 28-nm hybrid INT-CIM test chip, achieving 59.74 TOPS/W and 0.96 TOPS/mm2, with performance metrics of 23.21 perplexity for GPT-2, 68.69% accuracy for ResNet18, and 80.53% accuracy for ViT. Shaochen Li, Xi Chen 0107, Yujia Xiong, Lingyi Kong, He Wang 0028, Tianhui Jiao, Yan Yan 0030, Xin Si |
IEEE Trans. Very Large Scale Integr. Syst. | 12 |
| 2025 | TRIFP-DCIM: A Toggle-Rate-Immune Floating-point Digital Compute-in-Memory Design with Adaptive-Asymmetric Compute-TreeabstractFloating-point compute-in-memory (FP-CIM) is regarded as an attractive approach to enhancing the energy efficiency of complex neural networks. Digital domain compute mechanism has been widely utilized in CIM designs owing to its high robustness to PVT variations. However, the energy consumption of digital CIM is significantly influenced by the toggle rate of the compute-tree. This work proposes a toggle-rate-immune floating-point digital CIM (TRIFP-DCIM) design with 34.03% compute energy reduction on average. Combined with the TRIFP-DCIM design, a toggle-rate gathering method is employed in the neural network training/inference process with almost no accuracy loss. Experiment results show that the TRIFP-DCIM can achieve 14.51--36.83 TFLOPS/W @BF16 in 28nm technology process. Tianhui Jiao, Shaochen Li, Zhican Zhang, Xi Chen 0107, Xin Si |
ASP-DAC | 8 |
| 2025 | OutlierCIM: Outlier-Aware Digital CIM-Based LLM Accelerator with Hybrid-Strategy Quantization and Unified FP-INT ComputationabstractActivation outliers in Large Language Models (LLMs), which exhibit large magnitudes but small quantities, significantly affect model performance and pose challenges for the acceleration of LLMs. To address this bottleneck, researchers have proposed several co-design frameworks with outlier-aware algorithms and dedicated hardware. However, they face challenges balancing model accuracy with hardware efficiency when accelerating LLMs in a low bit-width manner. To this end, we propose OutlierCIM, the first algorithm and hardware codesign framework for the compute-in-memory (CIM) accelerator with outlier-aware quantization algorithm. The key contributions of OutlierCIM are 1) an outlier-clustered tiling strategy that regulates memory access and reduces inefficient workloads which are both introduced by outliers, 2) a hybrid-strategy quantization and a reconfigurable double-bit CIM macro array that overcome the low storage utilization and high latency of outlier-based LLM quantization, and 3) a quantization factor post-processing strategy and a dedicated quantizer that efficiently unify the multiplication and accumulation of outlier-caused FP-INT workloads. Implemented in a 28 nm CMOS technology, OutlierCIM occupies an area of $2.25 \mathrm{~mm}^{2}$. When evaluated at comprehensive benchmarks, OutlierCIM achieves up to $4.54 \times$ energy efficiency improvement and $3.91 \times$ speedup compared to the state-of-the-art outlier-aware accelerators. Zihan Zou, Shikuang Chen, Chen Zhang 0001, Xin Si, Hao Cai 0001, Bo Liu 0019 |
DAC | 7 |
| 2025 | H3D-LLM: Heterogeneous 3D Chiplet Design for LLM Inference with Dynamic Task Scheduling and Memory-Aware OrchestrationabstractThe exponential growth of Large Language Model (LLM) intensifies hardware demands for energy-efficient, low-latency architectures with scalable memory bandwidth. While 3D chiplet integration addresses conventional systems’ memory wall limitations, three critical challenges persist: asymmetric compression constraints from divergent sparsity-precision requirements across attention/projection layers, tier-level load imbalance from static resource allocation in dynamic computation patterns, and coupling-induced signal degradation in high-density TSV networks, especially under LLM-phase-specific traffic with spatiotemporal burstiness. To address these, we present H3D-LLM, a vertically heterogeneous architecture combining analog/digital Computing-in-Memory (CIM) and Neural Processing Unit (NPU) chiplets through three innovations. First, a Sparse-Aware Dynamic Execution Framework (SADEF) with Precision-Adaptive Quantization Mechanism (PAQM) enables hardware-aware compression via layer-wise unstructured sparsity detection and INT4/8-FP/BF16 mixed precision. Second, a 3D Spatio-Temporal Interleaved Parallelism (3D-STIP) with Semantic-Aware Tiered Storage (SATS) eliminates resource imbalance and improves memory efficiency through dynamic sub-batch partitioning and Key-Value (KV) cache aware management. Third, a Phase-Adaptive TSV Management (PATM) scheme with dynamic encoding and cluster-based allocation enhances inter-connect efficiency and signal integrity through runtime-aware partitioning and phase-specific dataflow scheduling. Evaluations on Llama-7B demonstrate that H3D-LLM achieves 12.3× higher energy efficiency and 8.4× faster inference than the A800 GPU, while its TSV strategy increases eye height by 12% and reduces bit error rate by up to 60× compared to naïve 3D accelerators. Hui Kou, Chenjie Xia, Liyi Li 0004, Hao Cai 0001, Xin Si, Bo Liu 0019 |
ICCAD | 6 |
| 2025 | SparCIM: A Heterogeneous CIM-Based Accelerator for Large Language Models with Contextual and Unstructured Bit SparsityabstractTransformer-based Large Language Models (LLMs) exhibit high computation and memory demands, especially on resource-constrained edge devices. Sparsification techniques like contextual sparsity have shown potential in reducing these demands by selectively pruning model components, but their integration with efficient hardware remains challenging. Firstly, sparsity prediction introduces extra energy and latency cost. Secondly, prefilling and decoding stages in LLM inference cause imbalanced workload. Thirdly, Computing-in-Memory (CIM) suffers from low utilization due to unstructured bit sparsity. To address the above challenges, this paper proposes SparCIM, a heterogeneous CIM accelerator optimized for decoder-only LLMs with contextual and unstructured bit sparsity, leveraging both Analog-Domain and Digital-Domain CIM architectures. The main contributions are: (1) a Heterogeneous CIM Architecture (HCA) which utilizes Analog CIM for sparse prediction of attention heads and Feed Forward Network parameters, while Digital CIM processes essential computations based on contextual sparsity, enhancing energy efficiency; (2) a Dynamic Workload Allocator (DWA) which configures DCIM memory and computing modes dynamically to reduce external memory access and token generation latency; (3) a Bit-level Butterfly Sparsity Alignment Unit (BSAU) which leverages bit-level sparsity, improving CIM utilization and throughput. The proposed accelerator is implemented with an industry 28-nm technology, achieving 1.27-5.98× higher energy efficiency compared to state-of-the-art transformer accelerators with minimal accuracy loss. Xingyu Xu 0008, Zihan Zou, Xin Si, Bo Liu 0019 |
ICCAD | 6 |
| 2025 | A L0 Framework with Anisotropic Sparsity and Fairness for 3D Denoising
Jinfeng Qian, Xin Si, Yong Zhao 0004, Jingliang Zhang |
ICIC (1) | 3 |
| 2025 | H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM InferenceabstractLow-batch large language model (LLM) inference has been extensively applied to edge-side generative tasks, such as personal chat helper, virtual assistant, reception bot, private edge server, etc.To efficiently handle both prefill and decoding stages in LLM inference, near-memory processing (NMP) enabled heterogeneous computation paradigm has been proposed.However, existing NMP designs typically embed processing engines into DRAM dies, resulting in limited computation capacity, which in turn restricts their ability to accelerate edge-side low-batch LLM inference.To tackle this problem, we propose H 2 -LLM, a Hybrid-bondingbased Heterogeneous accelerator for edge-side low-batch LLM inference.To balance the trade-off between computation capacity and bandwidth intrinsic to hybrid-bonding technology, we propose * Co-corresponding authors. Cong Li 0008, Yihan Yin, Xintong Wu, Jingchen Zhu, Zhutianya Gao, Dimin Niu, Qiang Wu 0012, Xin Si, Yuan Xie 0001, Chen Zhang 0001, Guangyu Sun 0003 |
ISCA | 8 |
| 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIMabstractSRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computational precision.However, the pursuit of higher performance necessitates more complex circuit designs and increased operating frequencies, which exacerbate IR-drop issues.Severe IR-drop can significantly degrade chip performance and even threaten reliability.Conventional circuit-level IR-drop mitigation methods, such as back-end optimizations, are resource-intensive and often compromise power, performance, and area (PPA).To address these challenges, we propose AIM, comprehensive software and hardware co-design for architecture-level IR-drop mitigation in high-performance PIM.Initially, leveraging the bit-serial and in-situ dataflow processing properties of PIM, we introduce R tog and HR, which establish a direct correlation between PIM workloads and IR-drop.Building on this foundation, we propose LHR and WDS, enabling extensive exploration of architecture-level IR-drop mitigation while maintaining computational accuracy through software optimization.Subsequently, we develop IR-Booster, a dynamic adjustment mechanism that integrates software-level HR information with hardwarebased IR-drop monitoring to adapt the V-f pairs of the PIM macro, achieving enhanced energy efficiency and performance.Finally, we propose the HR-aware task mapping method, bridging software and hardware designs to achieve optimal improvement.Post-layout simulation results on a 7nm 256-TOPS PIM chip demonstrate that AIM achieves up to 69.2% IR-drop mitigation, resulting in 2.29× energy efficiency improvement and 1.152× speedup. Yuanpeng Zhang 0002, Xing Hu 0010, Xi Chen 0107, Zhihang Yuan, Cong Li 0008, Jingchen Zhu, Xin Si, Wei Gao 0058, Qiang Wu 0012, Runsheng Wang, Guangyu Sun 0003 |
ISCA | 9 |
| 2025 | Variation-Adaptive Negative Bitline and Skip Bitline Pre-charge Scheme for Low-Power SRAMabstractThe negative bitline (NBL) write assist circuitry is widely adopted in static random access memory (SRAM) for its significant improvement in write yield and acceptable cost. Due to PVT variations and the trade-offs between yield and energy in different applications, the adaptive configuration of NBL enable timing and negative voltage level is increasingly essential, while existing NBL schemes lack a comprehensive methodology to adjust NBL options effectively. This paper presents a PVT variation-adaptive NBL (VA-NBL) write assist scheme and skip bitline pre-charge (SBP) circuitry for low-power SRAM. The VA-NBL scheme establishes a mapping relationship between NBL options and PVT conditions and employs on-chip PVT sensors to adaptively control NBL options, ensuring the lowest power consumption and satisfying target yield under PVT variations. Applied in a 64kb SRAM macro, we achieved different optimal NBL options under various PVT conditions, resulting in a maximum write energy power savings of 47%, while introducing negligible area overhead on dynamic voltage frequency scaling (DVFS) systems that already integrate PVT sensors. Additionally, the SBP circuitry for hierarchical bitline architecture avoids unnecessary pre-charge for write operations and further yields a 53% reduction in write energy. Lishuo Deng, Changwei Yan, Xin Si, Weiwei Shan |
ISCAS | 4 |
| 2025 | Modeling of Less-Significant Group Quantization for Hybrid CIM ArchitectureabstractHybrid computing-in-memory (CIM) has gained growing interest in recent times due to its ability to combine the strengths of both digital CIM (DCIM) and analog CIM (ACIM). The Multi-Bit Fusion (MF) scheme enhances energy efficiency by fusing low-bit results, which would typically require multiple readout cycles, into a single-cycle readout. However, the relationship between hybrid INT-CIM circuit design and network performance based on the MF scheme has yet to be systematically explored. To fill this gap, a less-significant group quantization (LSGQ) model is proposed, defining and exploring the hybrid INT-CIM design space. Experimental results demonstrate up to 1.83x improvement in energy efficiency with minimal impact on performance. A 28nm hybrid INT-CIM test chip is fabricated, achieving 52.16 TOPS/W, 0.96 TOPS/mm2, with respective performance metrics of 48.58 perplexity for GPT-2, 75.51% accuracy for ResNet18, and 81.5% accuracy for ViT. Shaochen Li, Xi Chen 0107, Lingyi Kong, He Wang 0028, Yi Yang 0001, Xin Si |
ISCAS | 7 |
| 2025 | A 22-nm 64-kB lightning-like hybrid computing-in-memory macro with a compressed adder tree and analog-storage quantizers for transformer and CNNs
An Guo 0001, Xi Chen 0107, Fangyuan Dong, Jinwu Chen, Zhihang Yuan, Xing Hu 0010, Guangyu Sun 0003, Arindam Basu, Jun Yang 0006, Xin Si |
Sci. China Inf. Sci. | 11 |
| 2025 | Expansion of the memory pyramid in the era of large models: compute-intensive compute-in-memory and memory-intensive compute-in-memory
Zhican Zhang, Yi Yang 0001, Zhaoyang Zhang 0008, Jinwu Chen, Xin Si, Jun Yang 0006 |
Sci. China Inf. Sci. | 9 |
| 2025 | A Robust 3D Mesh Segmentation Algorithm With Anisotropic Sparse EmbeddingabstractABSTRACT 3D mesh segmentation, as a very challenging problem in computer graphics, has attracted considerable interest. The most popular methods in recent years are data‐driven methods. However, such methods require a large amount of accurately labeled data, which is difficult to obtain. In this article, we propose a novel mesh segmentation algorithm based on anisotropic sparse embedding. We first over‐segment the input mesh and get a collection of patches. Then these patches are embedded into a latent space via an anisotropic ‐regularized optimization problem. In the new space, the patches that belong to the same part of the mesh will be closer, while those belonging to different parts will be farther. Finally, we can easily generate the segmentation result by clustering. Various experimental results on the PSB and COSEG datasets show that our algorithm is able to get perception‐aware results and is superior to the state‐of‐the‐art algorithms. In addition, the proposed algorithm can robustly deal with meshes with different poses, different triangulations, noises, missing regions, or missing parts. Yong Zhao 0004, Xin Si, Jingliang Zhang |
Comput. Animat. Virtual Worlds | 4 |
| 2025 | Joint-Learning: A Robust Segmentation Method for 3D Point Clouds Under Label NoiseabstractABSTRACT Most of point cloud segmentation methods are based on clean datasets and are easily affected by label noise. We present a novel method called Joint‐learning, which is the first attempt to apply a dual‐network framework to point cloud segmentation with noisy labels. Two networks are trained simultaneously, and each network selects clean samples to update its peer network. The communication between two networks is able to exchange the knowledge they learned, possessing good robustness and generalization ability. Subsequently, adaptive sample selection is proposed to maximize the learning capacity. When the accuracies of both networks are no longer improving, the selection rate is reduced, which results in cleaner selected samples. To further reduce the impact of noisy labels, for unselected samples, we provide a joint label correction algorithm to rectify their labels via two networks' predictions. We conduct various experiments on S3DIS and ScanNet‐v2 datasets under different types and rates of noises. Both quantitative and qualitative results verify the reasonableness and effectiveness of the proposed method. By comparison, our method is substantially superior to the state‐of‐the‐art methods and achieves the best results in all noise settings. The average performance improvement is more than 7.43%, with a maximum of 11.42%. Tingyun Miao, Yong Zhao 0004, Xin Si, Jingliang Zhang |
Comput. Animat. Virtual Worlds | 5 |
| 2025 | An On-Chip-Training Keyword-Spotting Chip Using Interleaved Pipeline and Computation-in-Memory Cluster in 28-nm CMOSabstractTo improve the precision of keyword spotting (KWS) for individual users on edge devices, we propose an on-chip-training KWS (OCT-KWS) chip for private data protection while also achieving ultralow -power inference. Our main contributions are: 1) identity interchange and interleaved pipeline methods during backpropagation (BP), enabling the pipelined execution of operations that traditionally had to be performed sequentially, reducing cache requirements for loss values by 95.8%; 2) all-digital isolated-bitline (BL)-based computation-in-memory (CIM) macro, eliminating ineffective computations caused by glitches, achieving 2.03$\times$higher energy efficiency; and 3) multisize CIM cluster-based BP data flow, designing each CIM macro collaboratively to achieve all-time full utilization, reducing 47.2% of output feature map (Ofmap) access. Fabricated in 28-nm CMOS and enhanced with a refined library characterization methodology, this chip achieves both the highest training energy efficiency of 101.5 TOPS/W and the lowest inference energy of 9.9nJ/decision among current KWS chips. By retraining a three-class depthwise-separable convolutional neural network (DSCNN), detection accuracy on the private dataset increases from 80.8% to 98.9%. Junyi Qian, Peng Cao 0002, Xin Si, Weiwei Shan |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | A Hybrid Domain and Pipelined Analog Computing Chain for MVM ComputationabstractIn this article, a stream-architecture and pipelined hybrid computing chain is presented to process matrix-vector multiplication (MVM). In each stage of the computing chain, a primary multiply-accumulate (MAC) stage consisting of charge, time, and digital domain processing units makes signed or unsigned$8\times 1\times 8$bit MAC operations and MSB quantization. Based on the stream architecture, the length of the computing chain can be configured to fit different MVM applications. In the charge-domain MAC unit, a double-plate sampling and weighted capacitor array with writing yield and efficiency enhanced 7T bitcell and three-step weighting scheme is implemented. To utilize the speed and resolution advantages of time-domain computing, a high linearity voltage-to-time converter (VTC) followed by a dynamic tristate delay chain is proposed to transfer and store MAC values from the charge domain in the time domain. To realize fast analog readout, a folding type and distributed time-to-digital converter (TDC) is proposed. To fully eliminate the offset and variation in the distributed TDC, a specific residue readout timing and back-end calibration scheme are applied. In the digital domain, a double-input and double-clock dynamic D flip-flop is built to realize partial sum transmission and accumulation in a single cycle with low energy and area consumption. Post-simulation results show that this computing chain can achieve 20.89–40.72-TOPS/W energy efficiency and 4.498-TOPS/mm2 throughput. Tianzhu Xiong, Yuyang Ye 0001, Xin Si, Jun Yang 0006 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | FDCA: Fine-grained Digital-CIM based CNN Accelerator with Hybrid Quantization and Weight-Stationary DataflowabstractDigital-Compute-in-memory (DCIM) has demonstrated significant energy and area efficiency in convolutional neural network (CNN) accelerators, particularly for high precision applications. However, to mitigate parasitic effects on word and bit lines, most DCIMs employ fine-grained multiply-accumulate operations, which introduces new challenges and opportunities but has not been widely explored. This paper proposes FDCA: a fine-grained digital-CIM based CNN accelerator with hybrid quantization and weight-stationary dataflow, in which the key contributions are :1) a hybrid quantization approach for CNNs leveraging hessian trace and approximation is utilized. This method incorporates the ratio of computation time and storage time into quantization, achieving high efficiency while maintaining accuracy; 2) a Cartesian Genetic Programming based approximate shift and accumulate with error compensation is proposed, where an approximate adder tree is generated to compensate for errors introduced by DCIM; 3) an optimized weight-stationary dataflow is used to improve the utilization of CIM and eliminate dataflow stalls. The experimental results demonstrate that under 28-nm process, when running VGG16 and ResNet50 on CIFAR100, the proposed FDCA achieves 17.1TOPS/W and 18.79TOPS/W with only a slight decrease in accuracy by 0.71% and 0.98%, respectively. Compared to previous works, this work achieves 1.76× and 1.28× better in energy efficiency with less accuracy loss. Bo Liu 0019, Qingwen Wei, Yang Zhang 0132, Xingyu Xu 0008, Zihan Zou, Xinxiang Huang, Xin Si, Hao Cai 0001 |
DAC | 7 |
| 2024 | Complementary Series-connected STT-MTJ for Time-based Computing-in-MemoryabstractComputing-in-memory (CIM) based on spin transfer torque magnetic random access memory (STT-MRAM) is promised to be an effective way to overcome the "memory wall" bottleneck. In this work, we proposed a novel complementary series-connected magnetic tunnel junction (STT-MTJ) structure for time-based Computing-in-Memory (CST-CIM). The bit-cell with four transistors and one MTJ is utilized to establish a series-connected structure to improve the limited resistance of MTJ, which can be applied for high-linearity and sufficient-margin multiply-and-accumulate (MAC) operation. In addition, for peripheral computing circuit, a customized successive-approximation-register time-to-digital converter (SAR-TDC) is used for high energy efficiency and low latency. To optimize the multi-bit MAC operation, we proposed a novel hardware friendly signed binary weight mapping strategy, which can provide the computing flexibility with 1-8bit quantization. Simulation result shows the proposed CST-CIM architecture can achieve low computation latency of 5ns and peak energy efficiency of 106.7 TOPS/W. Bo Liu 0019, Xin Si, Hao Cai 0001 |
ISCAS | 3 |
| 2024 | ROTA-I/O: Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and RobustnessabstractIn safety-critical systems, timing accuracy is the key to achieving precise I/O control. To meet such strict timing requirements, dedicated hardware assistance has recently been investigated and developed. However, these solutions are often fragile, due to unforeseen timing defects. In this paper, we propose a robust and timing-accurate I/O co-processor, which manages I/O tasks using Execution Time Servers (ETSs) and a two-level scheduler. The ETSs limit the impact of timing defects between tasks, and the scheduler prioritises ETSs based on their importance, offering a robust and configurable scheduling infrastructure. Based on the hardware design, we present an ETS-based timing-accurate I/O schedule, with the ETS parameters configured to further enhance robustness against timing defects. Experiments show the proposed I/O control method outperforms the state-of-the-art method in terms of timing accuracy and robustness without introducing significant overhead. Zhe Jiang 0004, Shuai Zhao 0004, Xin Si, Gang Chen 0023, Nan Guan |
RTSS | 4 |
| 2024 | A 22-nm 264-GOPS/mm2 6T SRAM and Proportional Current Compute Cell-Based Computing-in-Memory Macro for CNNsabstractWith the rise of artificial intelligence and big data applications, the general-purpose Von Neumann architecture is no longer capable of fulfilling the requirements of these application scenarios. The large amount of parallelizable and repeatable multiply-and-accumulate (MAC) operations in deep neural networks provide the possibility for the emergence of storage-computing integrated architectures. Current-based computation and quantization are employed to circumvent signal margin limitations on the power supply voltage of the computing unit, thereby facilitating low-power design. The proposed design is a computing-in-memory (CIM) circuit based on current sampling accumulation and applies a current-sensing analog-to-digital converter design that exhibits reduced sensitivity to parasitic capacitance compared to voltage-based analog-to-digital converters. Its power consumption is proportional to the input current, achieving higher area efficiency and energy efficiency gains. The design of the CIM circuit based on the current sampling in the 22-nm FDSOI process is fabricated with an area efficiency of 264 GOPS/mm2. The peak energy efficiency is 20.81 TOPS/W, and the inference accuracy reaches 92.11% when employed to VGG-16 under CIFAR-10 dataset. Feiran Liu, Anran Yin, Bo Wang 0023, Zhongyuan Feng, Xiang Li 0147, Tianzhu Xiong, Xin Si |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2024 | A 28 nm 16-kb Sign-Extension-Less Digital-Compute-in-Memory Macro With Extension-Friendly Compute Units and Accuracy-Adjustable Adder-TreeabstractConventional digital-domain SRAM compute-in-memory (CIM) faces challenges in handling multiply-and-accumulate (MAC) operations with signed values, either in serial data feeding mode or extra sign-bit processing. The proposed CIM macro has the following features: 1) a sign-extension-less array multiplication circuit structure that eliminates the need for converting partial sums into 2’s complement, which removes the constraints related to handling specific symbol bits; 2) developing a circuit that avoids signed bit extension shift and accumulate, resulting in reduced area cost; and 3) integrating an adder structure that provides adjustable accuracy, thereby enhancing network adaptability as compared to traditional approximation techniques. A fabricated 28 nm 16-kb sign-extension-less DCIM was tested with the highest MAC speed with 5.6 ns (Signed 8 b IN&W 23 b Out) and achieved the best energy efficiency with 40.15 TOPS/W over a wide range of network adaptability. Xin Si, Fangyuan Dong, Shengnan He, Anran Yin, Xiang Li 0147 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | From macro to microarchitecture: reviews and trends of SRAM-based compute-in-memory circuits
Zhaoyang Zhang 0008, Jinwu Chen, Xi Chen 0107, An Guo 0001, Bo Wang 0023, Tianzhu Xiong, Yuyao Kong, Xingyu Pu, Shengnan He, Xin Si, Jun Yang 0006 |
Sci. China Inf. Sci. | 10 |
| 2022 | Enabling High-Quality Uncertainty Quantification in a PIM Designed for Bayesian Neural NetworkabstractUncertainty quantification measures the prediction uncertainty of a neural network facing out-of-training-distribution samples. Bayesian Neural Networks (BNNs) can provide high-quality uncertainty quantification by introducing specific noise to the weights during inference. To accelerate BNN inference, ReRAM processing-in-memory (PIM) architecture is a competitive solution to provide both high-efficient computing and in-situ noise generation at the same time. However, there normally exists a huge gap between the generated noise in PIM hardware and that required by a BNN model. We demonstrate that the quality of uncertainty quantification is substantially degraded due to this gap. To solve this problem, we propose a holistic framework called W2W-PIM. We first introduce an efficient method to generate noise in ReRAM PIM design according to the demand of a BNN model. In addition, the PIM architecture is carefully modified to enable the noise generation and evaluate uncertainty quality. Moreover, a calibration unit is further introduced to reduce the noise gap caused by imperfection of the noise model. Comprehensive evaluation results demonstrate that W2W-PIM framework can achieve high-quality uncertainty quantification and high energy-efficiency at the same time. Bingzhe Wu, Guangyu Sun 0003, Zhe Zhang 0006, Zhihang Yuan, Runsheng Wang, Ru Huang 0001, Dimin Niu, Hongzhong Zheng, Zhichao Lu, Meng-Fan Chang, Tianchan Guan, Xin Si |
HPCA | 14 |
| 2022 | ShareFloat CIM: A Compute-In-Memory Architecture with Floating-Point Multiply-and-Accumulate OperationsabstractCompute-in-memory (CIM) has been widely explored to overcome “Von-Neumann bottleneck” for its high throughput and energy efficiency. However, recent compute-in-memory works can only support integer (INT)-type multiply-and-accumulate (MAC) operations. Floating point MACs (FP-MAC) are highly required to achieve both high performance training and high accuracy inference. In this paper, we proposed a ShareFloat CIM architecture which can support FP-MAC operations. Neural networks with ShareFloat MAC can achieve almost the same accuracy as that with FP64 MAC. A 28nm 64Kb ShareFloat CIM macro was further implemented with an energy efficiency of 18.8 TFLOPS/W and 73.11% accuracy when applied to a VGG-16 network with ShareFloat MAC and CIFAR-100 dataset. An Guo 0001, Yongliang Zhou, Bo Wang 0023, Tianzhu Xiong, Xin Si, Jun Yang 0006 |
ISCAS | 7 |
| 2022 | SNNIM: A 10T-SRAM based Spiking-Neural-Network-In-Memory architecture with capacitance computationabstractSpiking-Neural-Networks (SNN) have natural advantages in high-speed signal processing and big data operation. However, due to the complex implementation of synaptic arrays, SNN based accelerators may face low area utilization and high energy consumption. Computing-In-Memory (CIM) shows great potential in performing intensive and high energy efficient computations. In this work, we proposed a JOT-SRAM based Spiking-Neural-Network-In-Memory architecture (SNNIM) with 28nm CMOS technology node. A compact JOT-SRAM bit-cell was developed to realize signed 5bit synapses arrays and configurable bias arrays (SYBIA). The soma array based standard 8T-SRAM (SMTA) stores the soma membrane voltage and the threshold value. A capacitance computation scheme (CCA) between them was proposed to support various SNN operations. The proposed SNNIM achieved energy efficiency of 25.18 TSyOPSI. And the proposed SNNIM achieved 1.79+× better array efficiency compared with previous works. Bo Wang 0023, Xiang Li 0147, Anran Yin, Zhongyuan Feng, Yuyao Kong, Tianzhu Xiong, Haiming Hsu, Yongliang Zhou, An Guo 0001, Jun Yang 0006, Xin Si |
ISCAS | 14 |
| 2022 | VCCIM: a voltage coupling based computing-in-memory architecture in 28 nm for edge AI applications
An Guo 0001, Xi Chen 0107, Xin Si |
CCF Trans. High Perform. Comput. | 4 |
| 2021 | A 40nm 1Mb 35.6 TOPS/W MLC NOR-Flash Based Computation-in-Memory Structure for Machine LearningabstractComputation-in-memory (CIM) is a feasible method to overcome "Von-Neumann bottleneck" with high throughput and energy efficiency. In this paper, we proposed a 1Mb Multi-Level (MLC) NOR Flash based CIM (MLFlash- CIM) structure with 40nm technology node. A multi-bit readout circuit was proposed to realize adaptive quantization, which comprises a current interface circuit, a multi-level analog shift amplifier (AS-Amp) and an 8-bit SAR-ADC. When applied to a modified VGG-16 Network with 16 layers, the proposed MLFlash-CIM can achieve 92.73% inference accuracy under CIFAR-10 dataset. This CIM structure also achieved a peak throughput of 3.277 TOPS and an energy efficiency of 35.6 TOPS/W with 4-bit multiplication and accumulation (MAC) operations. Sitao Zeng, Zhiguo Zhu, Zhaolong Qin, Chen Wang 0130, Jingjing Li 0001, Sanfeng Zhang 0001, Yajuan He, Chunmeng Dou, Xin Si, Meng-Fan Chang, Qiang Li 0021 |
ISCAS | 10 |
| 2019 | A Half-Select Disturb-Free 11T SRAM Cell With Built-In Write/Read-Assist Scheme for Ultralow-Voltage OperationsabstractThis paper presents a half-select disturb-free 11T static random access memory (SRAM) cell for ultralow-voltage operations. The proposed SRAM cell is well suited for bit-interleaving architecture, which helps to improve the soft-error immunity with error correction coding. The read static noise margin (RSNM) and the write margin (WM) are significantly improved due to its built-in write/read-assist scheme. The experimental results in a 40-nm standard CMOS technology indicate that at a 0.5-V supply voltage, RSNM of the proposed SRAM cell is$19.8\times $and$0.96\times $as that of 6T and 8T SRAM cells with min-area, respectively. It achieves$11.84\times $and$9.56\times $higher WM correspondingly. As a result, a lower minimum operation voltage is obtained. In addition, its leakage power consumption is reduced by 53.3% and 44.5% when compared with 6T and 8T SRAM cell with min-area, respectively. Yajuan He, Jiubai Zhang, Xiaoqing Wu, Xin Si, Shaowei Zhen, Bo Zhang 0027 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Parallelizing SRAM arrays with customized bit-cell for binary neural networksabstractRecent advances in deep neural networks (DNNs) have shown Binary Neural Networks (BNNs) are able to provide a reasonable accuracy on various image datasets with a significant reduction in computation and memory cost. In this paper, we explore two BNNs: hybrid BNN (HBNN) and XNOR-BNN, where the weights are binarized to +1/-1 while the neuron activations are binarized to 1/0 and +1/-1 respectively. Two SRAM bit cell designs are proposed, namely, 6T SRAM for HBNN and customized 8T SRAM for XNOR-BNN. In our design, the high-precision multiply-and-accumulate (MAC) is replaced by bitwise multiplication for HBNN or XNOR for XNOR-BNN plus bit-counting operations. To parallelize the weighted sum operation, we activate multiple word lines in the SRAM array simultaneously and digitize the analog voltage developed along the bit line by a multi-level sense amplifier (MLSA). In order to partition the large matrices in DNNs, we investigate the impact of sensing bit-levels of MLSA on the accuracy degradation for different sub-array sizes and propose using the nonlinear quantization technique to mitigate the accuracy degradation. With 64×64 sub-array size and 3-bit MLSA, HBNN and XNOR-BNN architectures can minimize the accuracy degradation to 2.37% and 0.88%, respectively, for an inspired VGG-16 network on the CIFAR-10 dataset. Design space exploration of SRAM based synaptic architectures with the conventional row-by-row access scheme and our proposed parallel access scheme are also performed, showing significant benefits in the area, latency and energy-efficiency. Finally, we have successfully taped-out and validated the proposed HBNN and XNOR-BNN designs in TSMC 65 nm process with measured silicon data, achieving energy-efficiency >100 TOPS/W for HBNN and >50 TOPS/W for XNOR-BNN. Rui Liu 0005, Xiaochen Peng, Xiaoyu Sun 0001, Win-San Khwa, Xin Si, Jia-Jing Chen, Jia-Fang Li, Meng-Fan Chang, Shimeng Yu |
DAC | 5 |