EDBT 2026 Demo / reviewers in the wild / expert
Shan Shen
dblp:15/4616
· DBLP profile ↗
20ranked-venue papers
12as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 10 first-author · 14 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mask-based Meta-Learning for Stuck-at Faults Tolerance in ReRAM Computing SystemsabstractReRAM crossbar-based computing-in-memory (CIM) systems offer computational efficiency but suffer from significant accuracy degradation under stuck-at faults (SAFs). Conventional approaches like retraining-based methods fail to effectively generalize across diverse SAFs ratios. To address this challenge, we propose the Mask-based Meta Learning (MML) framework, leveraging meta-learning’s multi-task generalization capability to achieve robust performance of varying SAFs scenarios. Within the MML framework, a SAFs mask-based task formulation is used to create meta-learning tasks based on SAFs masks rather than dataset. Second, we define a new meta-learning objective by integrating different SAFs masks into the meta loss. Finally, we develop a SAFs-sensitivity guided weight importance search algorithm and dynamically expands the adjustment range of crucial weights using the ReRAM array’s redundant cells to further enhance model performance. Experimental results show that our method demonstrates superior robustness and generalization performance over the state-of-the-art robust training approaches. Moreover, our method can maintain high accuracy across various SAFs ratios while the accuracy of other approaches illustrates large fluctuations. Zhan Shen, Shan Shen, Zhen Mei 0001, Daying Sun |
ASP-DAC | 4 |
| 2026 | Effective RC Reduction via Graph Sparsification for Accurate Post-Simulation of Mixed-Signal ICs
Shan Shen, Wenjian Yu |
ASP-DAC | 3 |
| 2026 | OpenACM: An Open-Source SRAM-Based Approximate CiM CompilerabstractThe rise of data-intensive AI workloads has exacerbated the "memory wall" bottleneck. Digital Compute-in-Memory (DCiM) using SRAM offers a scalable solution, but its vast design space makes manual design impractical, creating a need for automated compilers. A key opportunity lies in approximate computing, which leverages the error tolerance of AI applications for significant energy savings. However, existing DCiM compilers focus on exact arithmetic, failing to exploit this optimization. This paper introduces OpenACM, the first open-source, accuracy-aware compiler for SRAM-based approximate DCiM architectures. OpenACM bridges the gap between application error tolerance and hardware automation. Its key contribution is an integrated library of accuracy-configurable multipliers (exact, tunable approximate, and logarithmic), enabling designers to make fine-grained accuracy-energy trade-offs. The compiler automates the generation of the DCiM architecture, integrating a transistor-level customizable SRAM macro with variation-aware characterization into a complete, open-source physical design flow based on OpenROAD and the FreePDK45 library. This ensures full reproducibility and accessibility, removing dependencies on proprietary tools. Experimental results on representative convolutional neural networks (CNNs) demonstrate that OpenACM achieves energy savings of up to 64% with negligible loss in application accuracy. The framework is available on OpenACM:URL. JunHao Ma, Xingyang Li, Yule Sheng, Bochang Wang, Yiheng Wu, Shan Shen, Daying Sun |
DATE | 9 |
| 2026 | Error expectation-driven design and energy optimization of approximate multipliers
Yanghui Wu, Daying Sun, Shan Shen, Xiong Cheng |
Integr. | 4 |
| 2026 | Efficient Parallel ILU Factorization and Forward/Backward Substitution with Application to Large-Scale Nonlinear Circuit SimulationabstractEfficient techniques are proposed for parallel incomplete LU (ILU) factorization and forward/backward substitution, for the sparse matrices with the same sparsity pattern. These parallel algorithms are then used as a preconditioner for the generalized minimal residual (GMRES) algorithm to obtain an ILU-GMRES solver for large-scale circuit simulation. The novelty of the parallel ILU and substitution algorithms includes a subtree-based task scheduling scheme, a nested dissection-based approach for generating task queues, and the task packing and reverse-order execution techniques for forward/backward substitution. Experiments on 43 matrices dumped from circuit simulation show that the 8-thread parallel ILU with threshold (ILUT) factorization and forward/backward substitution with the proposed techniques achieve 4.2 \(\times\) and 3.3 \(\times\) parallel speedups on average, respectively. The proposed parallel ILUT-GMRES solver runs 5.3 \(\times\) , on average, faster than PARDISO on these benchmarks. When integrated into Ngspice, it enables up to 2.3 \(\times\) and 1.7 \(\times\) faster execution of a step of Newton-Raphson iteration than the commercial parallel HSPICE, for the DC analysis and time integration stages, respectively. Jiawen Cheng, Shan Shen, Zhenya Zhou, Wenjian Yu |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | Deep Learning Inspired Capacitance Extraction TechniquesabstractWith the advancement of integrated circuit (IC), the process technology becomes more complicated and the design margin shrinks. Thus, the parasitic extraction is more demanded during IC design. In this invited paper, we survey the research progress on IC capacitance extraction, especially the usage of deep-learning technologies in relevant problems. Firstly, a method based on graph neural network (GNN) for predicting the parasitic capacitances in the pre-layout design stage is presented. It exhibits potential benefit for the optimization of SRAM design. Then, the deep-learning-inspired methods for post-layout capacitance extraction are presented, including CNN-Cap, NAS-Cap and GNN-Cap, etc. They can revamp the accuracy drawback of layout parasitic extraction (LPE) method and the efficiency drawback of 3-D capacitance field solver. Lastly, we briefly review the deep-learning technique for improving the accuracy of the random walk based 3-D capacitance solver for the structures under the advanced process technology. Wenjian Yu, Shan Shen, Dingcheng Yang, Haoyuan Li 0004, Jiechen Huang, Chunyan Pei |
ASP-DAC | 2 |
| 2025 | Few-shot Learning on AMS Circuits and Its Application to Parasitic Capacitance PredictionabstractGraph representation learning is a powerful method to extract features from graph-structured data, such as analog/mixed-signal (AMS) circuits. However, training deep learning models for AMS designs is severely limited by the scarcity of integrated circuit design data. In this work, we present CircuitGPS, a few-shot learning method for parasitic effect prediction in AMS circuits. The circuit netlist is represented as a heterogeneous graph, with the coupling capacitance modeled as a link. CircuitGPS is pre-trained on link prediction and fine-tuned on edge regression. The proposed method starts with a small-hop sampling technique that converts a link or a node into a subgraph. Then, the subgraph embeddings are learned with a hybrid graph Transformer. Additionally, CircuitGPS integrates a low-cost positional encoding that summarizes the positional and structural information of the sampled subgraph. CircuitGPS improves the accuracy of coupling existence by at least 20% and reduces the MAE of capacitance estimation by at least 0.067 compared to existing methods. Our method demonstrates strong inherent scalability, enabling direct application to diverse AMS circuit designs through zero-shot learning. Furthermore, the ablation studies provide valuable insights into graph models for representation learning. Shan Shen, Hector Rodriguez Rodriguez, Wenjian Yu |
DAC | 1 |
| 2025 | Transferable Parasitic Estimation via Graph Contrastive Learning and Label Rebalancing in AMS CircuitsabstractGraph representation learning on Analog-Mixed Signal (AMS) circuits is crucial for various downstream tasks, e.g., parasitic estimation. However, the scarcity of design data, the unbalanced distribution of labels, and the inherent diversity of circuit implementations pose significant challenges to learning robust and transferable circuit representations. To address these limitations, we propose CircuitGCL, a novel graph contrastive learning framework that integrates representation scattering and label rebalancing to enhance transferability across heterogeneous circuit graphs. CircuitGCL employs a self-supervised strategy to learn topology-invariant node embeddings through hyperspherical representation scattering, eliminating dependency on large-scale data. Simultaneously, balanced mean squared error (BMSE) and balanced softmax cross-entropy (BSCE) losses are introduced to mitigate label distribution disparities between circuits, enabling robust and transferable parasitic estimation. Evaluated on parasitic capacitance estimation (edge-level task) and ground capacitance classification (node-level task) across TSMC 28nm AMS designs, CircuitGCL outperforms all state-of-the-art (SOTA) methods, with the R2improvement of 33.64% ~ 44.20% for edge regression and F1-score gain of 0.9× ~ 2.1× for node classification. Our code is available at https://github.com/ShenShan123/CircuitGCL. Shan Shen, Shenglu Hua, Jiawei Liu 0006, Jianwang Zhai, Chuan Shi 0001, Wenjian Yu |
ICCAD | 1 |
| 2025 | OpenYield: An Open-Source SRAM Yield Analysis and Optimization Benchmark SuiteabstractStatic Random-Access Memory (SRAM) yield analysis is essential for semiconductor innovation, yet research progress faces a critical challenge: the large gap between simplified academic models and the complexities observed in practice. The lack of open, higher-fidelity benchmarks has hindered reproducibility and transferability, as promising academic techniques often fail to carry over to more realistic settings. We present OpenYield, an open-source ecosystem that aims to narrow this gap through three contributions: (i) An SRAM circuit generator that explicitly incorporates second-order effects (interconnect/line parasitics, inter-cell leakage coupling, and peripheralcircuit variations) that are commonly omitted in academic studies. (ii) A standardized evaluation platform with a simple interface and baseline yield-analysis implementations to enable fair comparisons and reproducible research on these higherfidelity circuits. (iii) An optimization platform for transistor-level sizing under these models, supporting reproducible studies of robustness/efficiency trade-offs. OpenYield aims to foster more reproducible and transferable progress in SRAM-yield research. The framework is publicly available at OpenYield:URL. Shan Shen, Xingyang Li, Zhuohua Liu, Junhao Ma, Yiheng Wu, Yuquan Sun, Wei W. Xing |
ICCD | 1 |
| 2024 | Deep-Learning-Based Pre-Layout Parasitic Capacitance Prediction on SRAM DesignsabstractTo achieve higher system energy efficiency, SRAM in SoCs is often customized. The parasitic effects cause notable discrepancies between pre-layout and post-layout circuit simulations, leading to difficulty in converging design parameters and excessive design iterations. Is it possible to well predict the parasitics based on the pre-layout circuit, so as to perform parasitic-aware pre-layout simulation? In this work, we propose a deep-learning-based 2-stage model to accurately predict these parasitics in pre-layout stages. The model combines a Graph Neural Network (GNN) classifier and Multi-Layer Perceptron (MLP) regressors, effectively managing class imbalance of the net parasitics in SRAM circuits. We also employ Focal Loss to mitigate the impact of abundant internal net samples and integrate subcircuit information into the graph to abstract the hierarchical structure of schematics. Experiments on 4 real SRAM designs show that our approach not only surpasses the state-of-the-art model in parasitic prediction by a maximum of 19X reduction of error but also significantly boosts the simulation process by up to 598X speedup. Shan Shen, Dingcheng Yang, Chunyan Pei, Bei Yu 0001, Wenjian Yu |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | Ultra8T: A sub-threshold 8T SRAM with leakage detection
Shan Shen, Yongliang Zhou, Wenjian Yu |
Integr. | 1 |
| 2023 | Parallel Incomplete LU Factorization Based Iterative Solver for Fixed-Structure Linear Equations in Circuit SimulationabstractA series of fixed-structure sparse linear equations are solved in a circuit simulation process. We propose a parallel incomplete LU (ILU) preconditioned GMRES solver for those equations. A new subtree-based scheduling algorithm for ILU factorization and forward/backward substitution is adopted to overcome the load-balancing and data locality problem of the conventional levelization-based scheduling. Experimental results show that the proposed scheduling algorithm can achieve up to 2.6X speedup for ILU factorization and 3.1X speedup for forward/backward substitution compared to the levelization-based scheduling. The proposed ILU-GMRES solver achieves around 4X parallel speedup with 8 threads, which is up to 2.1X faster than that based on the levelization-based scheme. The proposed parallel solver also shows remarkable advantage over existing methods (including HSPICE) on transient simulation of linear and nonlinear circuits. Shan Shen, Wenjian Yu |
ASP-DAC | 4 |
| 2023 | A Timing Yield Model for SRAM Cells at Sub/Near-Threshold Voltages Based on a Compact Drain Current ModelabstractSub/near-threshold static random-access memory (SRAM) design is crucial for addressing the memory bottleneck in power-constrained applications. However, the high integration density and reliability under process variations demand an accurate estimation of extremely small failure probabilities. To capture such a “rare event” in memory circuits, the time and storage overhead of conventional simulations based on the Monte Carlo (MC) analysis cannot be tolerated. On the other hand, classic analytical methods predicting failure probabilities from a physical expression become inaccurate in the sub/near-threshold voltage domain due to the hypothetical distribution or the oversimplified drain current ($I_{ds}$) model for nanoscale devices. This work first proposes a simple but efficient empirical$I_{ds}$model to describe the drain-induced barrier lowering (DIBL) effect. Based on that, the probability density functions of the interest metrics in SRAM are derived. Two analytical models are then put forward to evaluate SRAM dynamic stabilities, including the access time failure and the write failure. The proposed models can be extended easily to different types of SRAM with different read/write-assist circuits. The models are validated against MC simulations across different operating voltages and temperatures. The average relative errors at 0.5-V$V_{\mathrm{ DD}}$are only 8.8% for the access-time failure model and 10.4% for the write failure model. The size of the required sample data set is$43.6\times $smaller than that of the state-of-the-art method. Shan Shen, Peng Cao 0002, Longxing Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | A Design of Timing Speculation SRAM-Based L1 Caches With PVT Autotracking Under Near-Threshold VoltagesabstractTo improve the performance of SRAM in caches under near-threshold voltages, several timing speculation techniques, such as the cross-sensing SRAM (CS-SRAM), are proposed. Meanwhile, for a given process, voltage, and temperature (PVT) condition, CS-SRAM has an optimal bitline discharging time (TBL) to achieve the lowest average access latency. However, existing timing speculation caches do not track the variations of different PVT conditions to adjust the access timing to the optimal TBL point, on which the system possesses the lowest average memory access time. In this article, we propose a design of CS-SRAM-based L1 caches with a PVT autotracking mechanism, namely TS-PULP, which adjusts both the TBL and the frequency of the system clock to the optimal points. To quantify the improvement of our approach, a cycle-accurate RTL model of CS-SRAM and a field-programmable gate array (FPGA) prototype of the proposed L1 caches with the open-source system on chip (SoC) platform PULP have also been implemented. According to the evaluation results from RTL simulations and the FPGA prototype, the proposed caches can achieve a similar performance (over 80%) of the original standard cell memory (SCM)-based PULP design under TSMC 28 nm 0.5 V and 25 °C with only about 30% chip area. In addition, we introduced the figure of merit (FOM) of million instructions per second (MIPS), area, and energy (MAE) to comprehensively evaluate different approaches. The proposed scheme, TS-PULP, achieves the best FOM of MAE among four architectures (PULP with SCM, PULP with SRAM, TS-cache, and TS-PULP) with different cache sizes. Qingde Lin, Ke Tan 0004, Tianxiang Shao, Shan Shen, Jun Yang 0006 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2020 | TYMER: A Yield-based Performance Model for Timing-speculation SRAMabstractIn low power designs, timing-speculative techniques are proposed to boost the SRAM frequency and throughput. This paper proposes TYMER, a unified yield-based performance model for timing-speculative SRAM. In TYMER, the first sub-model evaluates access-time yield at different worldline enable time for a general 6T SRAM under low supply voltages, while the second one uses the yield results to estimate optimal sensing time and the overall read latency for speculative SRAM. TYMER is not only compared with simulation results but also the measurements from 28nm fabricated speculative SRAM chips. Both cases show precise evaluation results under different operating conditions. Shan Shen, Liang Pang 0002, Tianxiang Shao, Xiao Shi 0001, Longxing Shi |
DAC | 1 |
| 2020 | Modeling and Designing of a PVT Auto-tracking Timing-speculative SRAMabstractIn the low supply voltage region, the performance of 6T cell SRAM degrades seriously, which takes more time to achieve the sufficient voltage difference on bitlines. Timing- speculative techniques are proposed to boost the SRAM frequency and the throughput with speculatively reading data in an aggressive timing and correcting timing failures in one or more extended cycles. However, the throughput gains of timing- speculative SRAM are affected by the process, voltage and temperature (PVT) variations, which causes the timing design of speculative SRAM to be either too aggressive or too conservative. This paper first proposes a statistical model to abstract the characteristics of speculative SRAM and shows the presence of an optimal sensing time that maximizes the overall throughput. Then, with the guidance of the performance model, a PVT auto-tracking speculative SRAM is designed and fabricated, which can dynamically self-tune the bitline sensing to the optimal time as the working condition changes. According to the measurement results, the maximum throughput gain of the proposed 28nm SRAM is 1.62X compared to the baseline at 0.6V VDD. Shan Shen, Tianxiang Shao, Jun Yang 0006, Longxing Shi |
DATE | 1 |
| 2020 | TS Cache: A Fast Cache With Timing-Speculation Mechanism Under Low Supply VoltagesabstractTo mitigate the ever-worsening “power wall” problem, more and more applications need to expand their working voltage to the wide-voltage range including the nearthreshold region. However, the read delay distribution of the static random access memory (SRAM) cells under the nearthreshold voltage shows a more serious long-tail characteristic than that under the nominal voltage due to the process fluctuation. Such degradation of SRAM delay makes the SRAM-based cache a performance bottleneck of systems as well. To avoid unreliable data reading, circuit-level studies use larger/more transistors in a bitcell by sacrificing chip area and the static power of cache arrays. Architectural studies propose the auxiliary error correction or block disabling/remapping methods in fault-tolerant caches, which worsen both the hit latency and energy efficiency due to the complex accessing logic. This article proposes a timing-speculation (TS) cache to boost the cache frequency and improve energy efficiency under low supply voltages. In the TS cache, the voltage differences of bitlines (BLs) are continuously evaluated twice by a sense amplifier (SA), and the access timing error can be detected much earlier than that in prior methods. According to the measurement results from the fabricated chips, the TS L1 cache aggressively increases its frequency to 1.62× and 1.92× compared with the conventional scheme at 0.5- and 0.6-V supply voltages, respectively. Shan Shen, Tianxiang Shao, Xiaojing Shang, Jun Yang 0006, Longxing Shi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Detecting the phase behavior on cache performance using the reuse distance vectors
Shan Shen, Longxing Shi |
J. Syst. Archit. | 1 |
| 2008 | Detection of Infarct Lesions From Single MRI Modality Using Inconsistency Between Voxel Intensity and Spatial Location - A 3-D Automatic ApproachabstractDetection of infarct lesions using traditional segmentation methods is always problematic due to intensity similarity between lesions and normal tissues, so that multispectral MRI modalities were often employed for this purpose. However, the high costs of MRI scan and the severity of patient conditions restrict the collection of multiple images. Therefore, in this paper, a new 3-D automatic lesion detection approach was proposed, which required only a single type of anatomical MRI scan. It was developed on a theory that, when lesions were present, the voxel-intensity-based segmentation and the spatial-location-based tissue distribution should be inconsistent in the regions of lesions. The degree of this inconsistency was calculated, which indicated the likelihood of tissue abnormality. Lesions were identified when the inconsistency exceeded a defined threshold. In this approach, the intensity-based segmentation was implemented by the conventional fuzzy c-mean (FCM) algorithm, while the spatial location of tissues was provided by prior tissue probability maps. The use of simulated MRI lesions allowed us to quantitatively evaluate the performance of the proposed method, as the size and location of lesions were prespecified. The results showed that our method effectively detected lesions with 40-80% signal reduction compared to normal tissues (similarity index > 0.7). The capability of the proposed method in practice was also demonstrated on real infarct lesions from 15 stroke patients, where the lesions detected were in broad agreement with true lesions. Furthermore, a comparison to a statistical segmentation approach presented in the literature suggested that our 3-D lesion detection approach was more reliable. Future work will focus on adapting the current method to multiple sclerosis lesion detection. Shan Shen, André J. Szameitat, Annette Sterr |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2005 | MRI fuzzy segmentation of brain tissue using neighborhood attraction with neural-network optimizationabstractImage segmentation is an indispensable process in the visualization of human tissues, particularly during clinical analysis of magnetic resonance (MR) images. Unfortunately, MR images always contain a significant amount of noise caused by operator performance, equipment, and the environment, which can lead to serious inaccuracies with segmentation. A robust segmentation technique based on an extension to the traditional fuzzy c-means (FCM) clustering algorithm is proposed in this paper. A neighborhood attraction, which is dependent on the relative location and features of neighboring pixels, is shown to improve the segmentation performance dramatically. The degree of attraction is optimized by a neural-network model. Simulated and real brain MR images with different noise levels are segmented to demonstrate the superiority of the proposed technique compared to other FCM-based methods. This segmentation method is a key component of an MR image-based classification system for brain tumors, currently being developed. Index Terms-Improved fuzzy c-means clustering (IFCM), magnetic resonance imaging (MRI), neighborhood attraction, segmentation. Shan Shen, William Sandham, Malcolm Granat, Annette Sterr |
IEEE Trans. Inf. Technol. Biomed. | 1 |