EDBT 2026 Demo / reviewers in the wild / expert
Shaodi Wang
dblp:142/0321
· DBLP profile ↗
12ranked-venue papers
5as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 4GS/s 8b Time-Interleaved SAR ADC with LSB-Repeating-Based Background Offset Calibration and Adaptive-Average Residue Estimator
Yunsong Tao, Xiyu He, Xiaoge Zhu, Anqiang Guo, Long Kong, Shaodi Wang, Yi Zhong 0002, Lu Jie 0001, Nan Sun 0001 |
ISCAS | 9 |
| 2024 | Spatial-Frequency Characteristics of EEG Associated With the Mental Stress in Human-Machine SystemsabstractAccurate assessment of user mental stress in human-machine system plays a crucial role in ensuring task performance and system safety. However, the underlying neural mechanisms of stress in human-machine tasks and assessment methods based on physiological indicators remain fundamental challenges. In this paper, we employ a virtual unmanned aerial vehicle (UAV) control experiment to explore the reorganization of functional brain network patterns under stress conditions. The results indicate enhanced functional connectivity in the frontal theta band and central beta band, as well as reduced functional connectivity in the left parieto-occipital alpha band, which is associated with increased mental stress. Evaluation of network metrics reveals that decreased global efficiency in the theta and beta bands is linked to elevated stress levels. Subsequently, inspired by the frequency-specific patterns in the stress brain network, a cross-band graph convolutional network (CBGCN) model is constructed for mental stress brain state recognition. The proposed method captures the spatial-frequency topological relationships of cross-band brain networks through multiple branches, with the aim of integrating complex dynamic patterns hidden in the brain network and learning discriminative cognitive features. Experimental results demonstrate that the neuro-inspired CBGCN model improves classification performance and enhances model interpretability. The study suggests that the proposed approach provides a potentially viable solution for recognizing stress states in human-machine system by using EEG signals. Qunli Yao, Heng Gu, Shaodi Wang, Xiaoli Li 0002 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Design of Ultracompact Content Addressable Memory Exploiting 1T-1MTJ CellabstractContent addressable memories (CAMs) are a promising category of computing-in-memory (CiM) elements that can perform highly parallel and efficient search operations for routers, pattern matching, and other data-intensive applications. Various magnetic tunnel junction (MTJ)-based CAM designs have been proposed to realize zero standby power and high-performance search. However, due to the relatively small tunnel magneto-resistance (TMR) ratio, MTJ-based CAMs require extra transistors and differential MTJ branches to distinguish between the parallel and anti-parallel resistance states, resulting in significant area and energy overhead. In this article, we propose a device-circuit co-design approach for an ultracompact CAM design by only exploiting a 1T-1MTJ structure in each cell. We propose a 2-step search scheme to enable the parallel in-memory search operation across the proposed CAM array and demonstrate the sufficient sensing margin of the array in a successful search operation. Evaluation results suggest that our proposed 1T-1MTJ-based CAM design improves$179\times /301\times $area efficiency compared with the state-of-the-art 15T-4MTJ/20T-6MTJ CAM design. Application benchmarking on hyperdimensional computing (HDC) inference shows a$54.6\times /12.8\times $speedup compared with GPU/20T-6MTJ CAM-based approaches. Cheng Zhuo, Kai Ni 0004, Mohsen Imani, Yuxuan Luo 0001, Shaodi Wang, Deming Zhang, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | SSM-CIM: An Efficient CIM Macro Featuring Single-Step Multi-bit MAC Computation for CNN Edge InferenceabstractCompute-in-memory (CIM) is a promising approach to solving the memory-wall problem existing in traditional computing architectures. In this paper, we introduce SSM-CIM, a charge-domain, static random-access memory (SRAM)-based CIM macro designed for area-energy-efficient convolutional neural network (CNN) inference. SSM-CIM utilizes an original sign-magnitude data encoding method for both inputs and weights. By codesigning four adjacent SRAM computing cells and employing a 3-bit digital-to-analog converter (DAC), SSM-CIM performs accurate 4-bit multiply-and-accumulate (MAC) computation in a single step, eliminating the peripheral digital shift-and-add circuits. To digitize the MAC computing results, a dedicated multi-reference assisted SAR ADC is designed by reusing the reference voltages from the DAC, which offers significant power and area savings. In addition, analog computing errors and quantization errors are analyzed to ensure the multi-bit computing accuracy of SSM-CIM. SSM-CIM is implemented and evaluated using 28-nm global foundry process. The post-layout simulation results validate the excellent computing linearity and accuracy of SSM-CIM. Benefitting from the compact layout design and fully parallel computing flow, the$144\times 256$macro achieves a peak throughput of 2.3 TOPS, an area efficiency of 10.2 TOPS/mm2, and an energy efficiency of 205.4 TOPS/W with 4-bit weights and 4-bit inputs. Heng Zhang 0024, Sunan He, Xinjie Guo, Shaodi Wang, Yuan Du |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2017 | Hybrid VC-MTJ/CMOS non-volatile stochastic logic for efficient computingabstractIn this paper, we propose a non-volatile stochastic computing (SC) scheme using voltage-controlled magnetic tunnel junction (VC-MTJ) and negative differential resistance (NDR). The proposed design includes a VC-MTJ based true stochastic bit stream generator and VC-MTJ and NDR based stochastic adder, multiplier, register, which are experimentally demonstrated using 60nm VC-MTJ and CMOS NDR connected on die. These components are then used to realize FIR filter and AdaBoost (machine-learning algorithm). 3X–37X energy advantage is shown for the proposed SC compared with CMOS binary arithmetic ASIC and SC designs. Shaodi Wang, Saptadeep Pal, Tianmu Li, Andrew Pan, Cecile Grezes, Pedram Khalili Amiri, Kang L. Wang, Puneet Gupta 0001 |
DATE | 1 |
| 2017 | Assessing Benefits of a Buried Interconnect Layer in Digital DesignsabstractIn sub-15 nm technology nodes, local metal layers have witnessed extremely high congestion leading to pin-access-limited designs, and hence affecting the chip area and related performance. In this paper, we assess the benefits of adding a buried interconnect layer below the device layers for the purpose of reducing cell area, improving pin access, and reducing chip area. After adding the buried layer to a projected 7 nm standard cell library, results show ~9%-13% chip area reduction and 126% pin access improvement. This shows that buried interconnect, as an integration primitive, is very promising as an alternative method to density scaling. Liheng Zhu, Yasmine Badr, Shaodi Wang, Subramanian S. Iyer, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | A Word Line Pulse Circuit Technique for Reliable Magnetoelectric Random Access MemoryabstractA word line pulse (WLP) circuit scheme is proposed toward the implementation of magnetoelectric random access memory (MeRAM). The circuit improves the write error rate (WER) and cell area efficiency by generating a better write pulse compared to conventional bitline pulse (BLP) techniques in terms of the pulse slew rate and amplitude. For the voltage-controlled magnetic anisotropy-induced precessional switching of the magnetic tunnel junction (MTJ), the write pulse shape has a large impact on the switching probability. Typically, a square shape pulse results in higher switching probability compared to that of a triangular shape pulse with long rise and falling edges, since the square shape pulse causes a more stable precessional trajectory of the free layer magnetization by providing a relatively constant in-plane-dominant effective field. Compared to the BLP scheme, the WLP can generate a better square shape pulse by eliminating discharge paths under the pulse condition, using the gain of the access transistor, and effectively diminishing the capacitive loading which needs to be driven. A macrospin compact model of voltage-controlled MTJ shows that the WLP can improve WER by${10}^{7}$times and allow MeRAM to have four-time improvement in area efficiency of driver circuits compared to the BLP. Hochul Lee, Shaodi Wang, Farbod Ebrahimi, Puneet Gupta 0001, Pedram Khalili Amiri, Kang L. Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | MTJ variation monitor-assisted adaptive MRAM writeabstractSpin-transfer torque random access memory (STT-RAM) and magnetoelectric random access memory (MeRAM) are promising non-volatile memory technologies. But STT-RAM and Me RAM both suffer from high write error rate due to thermal fluctuation of magnetization. Temperature and wafer-level process variation significantly exacerbate these problems. In this paper, we propose a design that adaptively selects optimized write pulse for STT-RAM and MeRAM to overcome ambient process and temperature variation. To enable the adaptive write, we design specific MTJ-based variation monitor, which precisely senses process and temperature variation. The monitor is over 10X faster, 5X more energy-efficient, and 20X smaller compared with conventional thermal monitors of similar accuracy. With adaptive write, the write latency of STT-RAM and MeRAM cache are reduced by up to 17% and 59% respectively, and application run time is improved by up to 41%. Shaodi Wang, Hochul Lee, Cecile Grezes, Pedram Khalili Amiri, Kang L. Wang, Puneet Gupta 0001 |
DAC | 1 |
| 2016 | MEMRES: A Fast Memory System Reliability SimulatorabstractWith scaling technology, emerging nonvolatile devices, and data-intensive applications, memory faults have become a major reliability concern for computing systems. With various hardware and software approaches proposed to address this issue, a comprehensive evaluation is required to understand the effectiveness of these solutions. Considering the complex nature of various memory faults as well as interactions between various correction mechanisms, we propose MEMRES, a fast main memory system reliability simulator. It enables memory fault simulation with error-correcting code (ECC) algorithms and modern memory reliability management, including memory page retirement, mirroring, scrubbing, and hardware sparing. MEMRES is computationally efficient in obtaining memory failure probabilities in the presence of multiple failure mechanisms and complex correction scheme, allowing the optimization of memory system reliability, the prediction of emerging memory reliability, and designing a reliability enhancement technique. The accuracy of MEMRES is verified by an existing analytical model and an existing memory fault simulator. We performed a case study on spin-transfer torque random access memory (STT-RAM)-based main memory, and the results indicate that in-memory ECC can significantly mitigate the write error rate of STT-RAM, demonstrating the capability of handling emerging memory system. Shaodi Wang, Henry Chaohong Hu, Hongzhong Zheng, Puneet Gupta 0001 |
IEEE Trans. Reliab. | 1 |
| 2016 | An Evaluation Framework for Nanotransfer Printing-Based Feature-Level Heterogeneous Integration in VLSI CircuitsabstractWe develop an evaluation framework to assess the potential benefits of feature-level heterogeneous integration (HGI) in nanoscale VLSI circuits. We study, for the first time, the impact of HGI on circuit delay, layout area, and power by comparing the integration of 15-nm InGaAs and Ge FinFETs via nanotransfer printing with the baseline Si-only FinFET technology. To properly account for the performance, power, and area tradeoffs, we perform comprehensive evaluations, including synthesis, placement, and routing of digital circuit benchmarks. We show the circuits designed with an HGI exhibit lower delay and power due to improved device performance at the cost of larger area induced by misalignment errors. We also demonstrate that the HGI misalignment area penalties can be drastically reduced using posttransfer fin trimming. Our findings provide substantial motivation for industry to explore HGI as a technology route for the post-Si era. Greg Leung, Shaodi Wang, Andrew Pan, Puneet Gupta 0001, Chi On Chui |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | PROCEED: A Pareto Optimization-Based Circuit-Level Evaluator for Emerging DevicesabstractEvaluation of novel devices in the context of circuits is crucial to identifying and maximizing their value. We propose a new framework, Pareto optimization-based circuit-level evaluator for emerging device (PROCEED), that uses comprehensive performance, power, and area metrics for accurate device-circuit coevaluation through optimization of digital circuit benchmarks. The PROCEED assesses technology suitability over a wide operating region (megahertz to gigahertz) by leveraging available circuit knobs (threshold voltage assignment, power management, sizing, and so on). It improves the benchmark accuracy by 3x to 115x compared with the existing methods while offering orders of magnitude improvements in runtime over full physical design implementation flows. To illustrate the PROCEED's capabilities, we deploy it to assess emerging technologies, including novel tunneling field-effect transistors, compared with conventional silicon CMOS. As a further illustration, we extend PROCEED to evaluate future heterogeneous integration of varied devices onto the same silicon substrate. Shaodi Wang, Andrew Pan, Chi On Chui, Puneet Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | PROCEED: A pareto optimization-based circuit-level evaluator for emerging devicesabstractEvaluation of novel devices in a circuit context is crucial to identifying and maximizing their value. We propose a new framework, PROCEED, and metrics for accurate device-circuit co-evaluation through proper optimization of digital circuit benchmarks. PROCEED assesses technology suitability over a wide operating region (MHz to GHz) by leveraging available circuit knobs (Vtassignment, power management, sizing, etc.) and improves accuracy by 3X to 115X compared to existing methods while offering orders of magnitude improvements in runtime over full physical design implementation flows. To illustrate PROCEED's capabilities, we deploy it to assess novel tunneling transistors (TFETs) compared to conventional CMOS. Shaodi Wang, Andrew Pan, Chi On Chui, Puneet Gupta 0001 |
ASP-DAC | 1 |