EDBT 2026 Demo / reviewers in the wild / expert
Xingsheng Wang
dblp:60/8001
· DBLP profile ↗
14ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0001-8335-2033ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPEAR: Spike-Aware Point Pruning and Fused FPS-KNN Sampling for Spiking Point Transformer Acceleration Co-Design
Yilong Fang, Pinfeng Jiang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ISCAS | 7 |
| 2026 | DyMamba: dynamic Mamba for microscopy image semantic segmentationabstractMOTIVATION: Segmentation of cell bodies and organelles in microscopy images is critical for biological research, particularly in scenarios with multiple regions of interest where spatial continuity is essential. The Mamba architecture, derived from State Space Models (SSMs), has recently gained attention for efficiently modeling long-range dependencies in sequences, achieving excellent results in both natural and medical image segmentation. However, in vision tasks, current Mamba scanning strategies mainly focus on raster-scanning and local-scanning, which introduce spatial discontinuities, severely affecting the effectiveness of segmentation at the pixel level, especially in dense segmentation tasks. RESULTS: In this article, we propose DyMamba, a Mamba-based model featuring a dynamic scanning strategy that adaptively plans scanning paths based on local features and complexity. In addition, to address the challenges of detail prediction and small object detection, we introduce a local aware module that performs pixel-level regional processing on images. DyMamba achieves robust segmentation across diverse microscopy image types, including cell-, organelle- and tissue-scale images. Experiments on six datasets and multiple scanning strategies demonstrate the excellent performance of our method in segmenting microscopy images, achieving an average improvement of 6.9% in mDice and 4.3% in mIoU over state-of-the-art methods across all datasets. AVAILABILITY: The code is released at https://github.com/cbqBit/dymamba. Buqing Cai, Xingsheng Wang, Zhuo Jia, Fa Zhang 0001, Bin Hu 0001 |
Bioinform. | 2 |
| 2025 | SPARTA: Spike-Aware Token Skipping Co-Optimization with Heterogeneous ReRAM-CIM Architecture for Spiking Transformer AccelerationabstractSpiking neural networks (SNNs) have emerged as a promising paradigm for energy-efficient neural computation, offering advantages over artificial neural networks (ANNs) by using sparse spike activations. Among SNN models, spiking transformers have shown great potential in achieving high accuracy while maintaining low energy consumption, making them ideal for resource-constrained applications. However, a significant challenge in accelerating Spiking Transformers lies in managing the inherent unstructured sparsity in spike activations. This sparsity introduces substantial hardware overhead, limiting the efficiency of existing accelerators.To address these challenges, we propose SPARTA, an algorithm-hardware co-optimization framework designed specifically for spiking transformers. SPARTA utilizes a spike-aware dynamic token skipping algorithm, which applies reinforcement learning to selectively skip less informative tokens, achieving structured token-level sparsity in both spatial and temporal domains. Additionally, we introduce the spike-aware token prediction algorithm to predict and eliminate inactive tokens, further improving efficiency. Meanwhile, we present a dedicated heterogeneous ReRAM-based Compute-in-Memory (CIM) hardware architecture tailored to support token-level sparsity, which integrates a ReRAM analog CIM engine for linear layer and a token-spike fusion engine for optimized token routing and spike attention. Experimental results show that SPARTA achieves up to 543.1× and 10.2× speedup with 308.0× and 5.2× energy efficiency improvement compared to GPU and the state-of-the-art SNN accelerator "COMPASS", while preserving high model accuracy. Pinfeng Jiang, Yilong Fang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ICCAD | 7 |
| 2025 | A Novel Memristor-Based Majority Logic and Efficient Approximate Full Adder for Image ProcessingabstractThe memristor, an emerging non-volatile memory device, is well-suited for in-memory computing (IMC) due to its capability to simultaneously store data and perform logic operations. Majority (MAJ) logic is a type of expressive Boolean logic that performs well in logic synthesis. This paper proposes a novel MAJ logic implementation scheme to improve the latency of n-bit full adder (FA) from 6n+1 to 5n+1. Additionally, an approximate computing approach is introduced, yielding an n-bit approximate adder with a significantly reduced delay of 2n+1, at the cost of precision. Both exact and approximate adders were verified experimentally, and the error quality metrics of the approximate design were thoroughly evaluated. Furthermore, hybrid configurations combining exact and approximate adders were applied to image processing tasks, where the peak signal-to-noise ratio (PSNR) for most hybrid designs remained within an acceptable range (>30dB). Zhouchao Gan, Yifeng Xiong, Fan Yang 0148, Xiangshui Miao, Xingsheng Wang |
ISCAS | 6 |
| 2025 | ReBA: A Hybrid Sparse Reconfigurable Butterfly Accelerator for Solving Partial Differential Equations via Hardware and Algorithm Co-DesignabstractPartial differential equations (PDEs) are widely used in many scientific and engineering fields. Traditional numerical methods for solving PDEs often struggle with complex equations and require extensive computation. Recent advances in neural operators, particularly the Galerkin Transformer (GT-FNO), offer faster solutions but present unique challenges for hardware due to complex neural network operations such as attention, Fourier transforms, and convolution. To address these challenges, we propose ReBA, a dedicated hardware and algorithm co-design framework specified for GT-FNO. Our design implements hybrid sparsity schemes for efficient workload balance, significantly reducing hardware overhead while maintaining high accuracy. ReBA introduces a reconfigurable dual-engine architecture, featuring an adaptable butterfly processing element (BPE), which efficiently supports FFT and other DNN operations through flexible BPE allocation, offering high parallelism and low-latency computation. Experiments validate that ReBA delivers significant speedups of 34.57×, 2.70×, and 1.82 51.26× over conventional CPUs, GPUs, and prior state-of-the-art accelerators in solving PDEs. Pinfeng Jiang, Yilong Fang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ISCAS | 7 |
| 2025 | Optimizing hardware-software co-design based on non-ideality in memristor crossbars for in-memory computing
Pinfeng Jiang, Danzhe Song, Menghua Huang, Fan Yang 0148, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 8 |
| 2025 | ISARA: An Island-Style Systolic Array Reconfigurable Accelerator Based on Memristors for Deep Neural NetworksabstractThe demand for edge artificial intelligence (AI) is significant, particularly in revolutionary technological areas such as the Internet of Things, autonomous driving, and industrial control. However, reliable and high-performance edge AI is still constrained by computing hardware, and improving the performance and reliability of edge AI accelerators remains a key focus for researchers. This work proposes a memristor/resistive random access memory (RRAM)-based island-style systolic array reconfigurable accelerator (ISARA) that meets the reliability and performance requirements of edge AI. Inspired by the island-style architecture of FPGAs, this work proposes a flexible-tile architecture based on RRAM processing element (PE) islands, optimizing the data flow within the systolic array. The design of network-on-chip reduces data processing latency. In addition, to enhance computational efficiency, this work incorporates a bit-fusion scheme within the flexible tile, which reduces analog-to-digital converter (ADC) power consumption and addresses the conductance variation of RRAM. To date, only a few works have completed the entire process from simulation, design, and fabrication to hardware testing. This work fully realizes the design and validation of a new accelerator based on RRAM chips, demonstrating the reliability of RRAM-based systolic array accelerators for the first time. After deploying algorithms, the hardware accelerator achieved recognition rates comparable to software. Compared to similar works, ISARA’s computational efficiency exceeds theirs and has flexible reconfigurability. The same deep neural network (DNN) models are adopted for evaluation and compared to other accelerators, and ISARA’s processing latency is reduced by 200 times. Fan Yang 0148, Pinfeng Jiang, Xiangshui Miao, Xingsheng Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | Multi-order Differential Neural Network for TCAD Simulation of the Semiconductor DevicesabstractTechnology Computer Aided Design (TCAD) is a crucial step in the design and manufacturing of semiconductor devices. It involves solving physical equations that describe the behavior of semiconductor devices to predict various device parameters. Traditional TCAD methods, such as finite volume and finite element methods, discretize relevant physical equations to achieve numerical simulations of devices, significantly burdening the computation resources. For the first time, this paper proposes a novel method for TCAD simulation based on Physics-Informed Neural Networks (PINNs). We proposed multi-order differential neural network (MDNN), an improved Radial Basis Function Neural Network (RBFNN) model. By training MDNN, it achieves the couple solution of the Poisson equation and drift-diffusion equation under steady-state conditions, without the need for a pre-existing dataset. To the best of our knowledge, this marks the first instance of an ML-TCAD simulation that does not require any pre-existing data. For an example of PN junction diode, this method effectively simulates the basic physical characteristics of the device, with a self-consistent solution error of less than 1×10−5. Zifei Cai, Anaoxue Huang, Yifeng Xiong, Dejiang Mu, Xiangshui Miao, Xingsheng Wang |
DAC | 6 |
| 2024 | Efficient implementation of majority-inverter graph logic and arithmetic functions with memristor arrays
Zhouchao Gan, Yinghao Ma, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 6 |
| 2023 | A comprehensive study of device variability of sub-5 nm nanosheet transistors and interplay with quantum confinement variation
Haowen Luo, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 4 |
| 2023 | Modeling and physical mechanism analysis of the effect of a polycrystalline-ferroelectric gate on FE-FinFETs
Chengxu Wang, Zichong Zhang, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 6 |
| 2013 | SRAM device and cell co-design considerations in a 14nm SOI FinFET technologyabstractWe report a systematic study on the impact of process and statistical variability on SRAM design in a 14nm SOI FinFET technology node. A comprehensive statistical compact modelling strategy is developed for the early delivery of reliable PDK model, which enables TCAD-based transistor-cell co-design and path finding during the early phase of a technology node. Binjie Cheng, Xingsheng Wang, Andrew R. Brown, Jente B. Kuang, Dave Reid, Campbell Millar, Sani R. Nassif, A. Asenov |
ISCAS | 2 |
| 2011 | New reliability mechanisms in memory design for sub-22nm technologiesabstractThe TRAMS (Terascale Reliable Adaptive MEMORY Systems) project addresses in an evolutionary way the ultimate CMOS scaling technologies and paves the way for revolutionary, most promising beyond-CMOS technologies. In this abstract we show the significant variability levels of future 18 and 13nm device bulk-CMOS technologies as well as its dramatic effect on the yield of memory cells, and what kind of circuit solution would be required to maintain the current yield level. Later, we discuss the impact of errors at the system level, and different approaches at system level to adapt the heterogeneous systems to user's requirements. Nivard Aymerich, A. Asenov, Andrew R. Brown, Ramon Canal, Binjie Cheng, Joan Figueras, Antonio González 0001, Enric Herrero, S. Markov, Miguel Corbalan, Peyman Pouyan, Tanausú Ramírez, Antonio Rubio 0001, Elena I. Vatajelu, Xavier Vera, Xingsheng Wang, Paul Zuber |
IOLTS | 16 |
| 2010 | Capturing intrinsic parameter fluctuations using the PSP compact modelabstractStatistical variability (SV) presents increasing challenges to CMOS scaling and integration at nanometer scales. It is essential that SV information is accurately captured by compact models in order to facilitate reliable variability aware design. Using statistical compact model parameter extraction for the new industry standard compact model PSP, we investigate the accuracy of standard statistical parameter generation strategies in statistical circuit simulations. Results indicate that the typical use of uncorrelated normal distribution of the statistical compact model parameters may introduce considerable errors in the statistical circuit simulations. Binjie Cheng, Daryoosh Dideban, Negin Moezi, Campbell Millar, Gareth Roy, Xingsheng Wang, Scott Roy, A. Asenov |
DATE | 6 |