EDBT 2026 Demo / reviewers in the wild / expert
Hongyang Shang
dblp:377/0016
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0007-6276-1947ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CADC: Crossbar-Aware Dendritic Convolution for Efficient In-memory ComputingabstractConvolutional neural networks (CNNs) are computationally intensive and often accelerated using crossbar-based in-memory computing (IMC) architectures. However, large convolutional layers must be partitioned across multiple crossbars, generating numerous partial sums (psums) that require additional buffer, transfer, and accumulation, thus introducing significant system-level overhead. Inspired by dendritic computing principles from neuroscience, we propose crossbar-aware dendritic convolution (CADC), a novel approach that dramatically increases sparsity in psums by embedding a nonlinear dendritic function (zeroing negative values) directly within crossbar computations. Experimental results demonstrate that CADC significantly reduces psums, eliminating 80% in LeNet-5 on MNIST, 54% in ResNet-18 on CIFAR-10, 66% in VGG-16 on CIFAR-100, and up to 88% in spiking neural networks (SNN) on the DVS Gesture dataset. The induced sparsity from CADC provides two key benefits: (1) enabling zero-compression and zero-skipping, thus reducing buffer and transfer overhead by 29.3%, and accumulation overhead by 47.9%; (2) minimizing ADC quantization noise accumulation, resulting in small accuracy degradation-only 0.01% for LeNet-5, 0.1% for ResNet-18, 0.5% for VGG-16, and 0.9% for SNN. Compared to vanilla convolution (vConv), CADC exhibits accuracy changes ranging from +0.11 % to + 0. 1 9 % for LeNet-5, - 0. 0 4 % t - 0. 2 7 % for ResNet-18, +0. 9 9 % to + 1. 6 0 % for VGG-16, and - 0. 5 7 % to + 1. 3 2 % for SNN, across crossbar sizes from $64 \times 64$ to $256 \times 256$. Ultimately, a SRAM-based IMC implementation of CADC achieves 2.15 TOPS and 40.8 TOPS/W for ResNet-18 (4/2/4b), realizing a $11 \times-18 \times$ speedup and $1.9 \times-22.9 \times$ improvement in energy efficiency compared to existing IMC accelerators. Ye Ke, Hongyang Shang, Arindam Basu |
ASP-DAC | 4 |
| 2026 | Near-Memory Architecture for Threshold-Ordinal Surface-Based Corner Detection of Event CamerasabstractEvent-based Cameras (EBCs) are widely utilized in surveillance and autonomous driving applications due to their high speed and low power consumption. Corners are essential low-level features in event-driven computer vision, and novel algorithms utilizing event-based representations, such as Threshold-Ordinal Surface (TOS), have been developed for corner detection. However, the implementation of these algorithms on resource-constrained edge devices is hindered by significant latency, undermining the advantages of EBCs. To address this challenge, a near-memory architecture for efficient TOS updates (NM-TOS) is proposed. This architecture employs a read-write decoupled 8T SRAM cell and optimizes patch update speed through pipelining. Hardware-software co-optimized peripheral circuits and dynamic voltage and frequency scaling (DVFS) enable power and latency reductions. Compared to traditional digital implementations, our architecture reduces latency/energy by 24.7×/1.2× at Vdd=1.2 V or 1.93×/6.6× at Vdd=0.6 V based on 65nm CMOS process. Monte Carlo simulations confirm robust circuit operation, demonstrating zero bit error rate at operating voltages above 0.62 V, with only 0.2% at 0.61 V and 2.5% at 0.6 V. Corner detection evaluation using precision-recall area under curve (AUC) metrics reveals minor AUC reductions of 0.027 and 0.015 at 0.6 V for two popular EBC datasets. Hongyang Shang, An Guo 0001, Ye Ke, Arindam Basu |
DATE | 1 |
| 2026 | SRAM-Based Compute-in-Memory Accelerator for Linear-decay Spiking Neural Networks
Hongyang Shang, Yahan Yang, Arindam Basu |
ISCAS | 1 |
| 2026 | 3D Stack In-Sensor-Computing (3DS-ISC): Accelerating Time-Surface Construction for Neuromorphic Event CamerasabstractThis work proposes a 3D Stack In-Sensor-Computing (3DS-ISC) architecture for efficient event-based vision processing. A real-time normalization method using an exponential decay function is introduced to construct the time-surface, reducing hardware usage while preserving temporal information. The circuit design utilizes the leakage characterization of Dynamic Random Access Memory(DRAM) for timestamp normalization. Custom interdigitated metal-oxide-metal capacitor (MOMCAP) is used to store the charge and low leakage switch (LL switch) is used to extend the effective charge storage time. The 3DS-ISC architecture integrates sensing, memory, and computation to overcome the memory wall problem, reducing power, latency, and reducing area by$69\times $,$2.2\times $and$1.9\times $, respectively, compared with its 2D counterpart. Moreover, compared to works using a 16-bit SRAM to store timestamps, the ISC analog array can reduce power consumption by three orders of magnitude. In real computer vision (CV) tasks, we applied the spatial-temporal correlation filter (STCF) for denoise, and 3D-ISC achieved almost equivalent accuracy compared to the digital implementation using high precision timestamps. As for the image classification, time-surface constructed by 3D-ISC is used as the input of GoogleNet, achieving 99% on N-MNIST, 85% on N-Caltech101, 78% on CIFAR10-DVS, and 97% on DVS128 Gesture, comparable with state-of-the-art results on each dataset. Additionally, the 3D-ISC method is also applied to image reconstruction using the DAVIS240C dataset, achieving the highest average SSIM (0.62) among three methods. This work establishes a foundation for real-time, resource-efficient event-based processing and points to future integration of advanced computational circuits for broader applications. Hongyang Shang, Ye Ke, Arindam Basu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | High Energy-efficiency and Low latency In-Memory Computing using Analog Accumulator and In-Memory ADC with shared ReferencesabstractThis article proposes a $256 \times 128$ in-memory computing array using reconfigurable in-memory analog-to-digital conversion with shared references for high area efficiency (area overhead of 3% is $9 X$ better than traditional). A dual-8T SRAM bitcell is used to achieve read-write decoupling and store ternary weights. Read World Line Under Drive enabled Cascode helps to minimize current variations producing high linearity. Multi-bit input is handled with low latency and high energyefficiency by using bit-slicing (BS) with near-memory chargesharing based binary weighted accumulator (CHA). Using noise resilient training, we show software comparable performance for a MLP on MNIST, VGG-8 on CIFAR-10, and graph attention network on Cora, with respective accuracy reductions of only $0.1 \%, 0.8 \%$ and 0.5% due to non-idealities. The proposed macro demonstrates high energy/area efficiency (1146 TOPS/W, 27 TOPS $/ \mathbf{m m}^{\mathbf{2}}$ at $\mathbf{1} \boldsymbol{/} \mathbf{2} \boldsymbol{/} \mathbf{1 b}$) in $\mathbf{6 5 ~ n m}$ CMOS. It increases throughput (by $1.9 X$) and linearity (by $23 X$) compared to input pulse-width modulation by using BS and CHA. Compared to conventional BS with digital accumulation after ADC, this method has $1.7 X / 6.6 X$ better energy-efficiency/throughput by reducing ADC operations. Zhengnan Fu, Hongyang Shang, Arindam Basu |
DAC | 4 |
| 2025 | Topkima-Former: Low-Energy, Low-Latency Inference for Transformers Using Top-k In-Memory ADCabstractTransformer has emerged as a leading architecture in neural language processing (NLP) and computer vision (CV). However, the extensive use of nonlinear operations, like softmax, poses a performance bottleneck during transformer inference and comprises up to 40% of the total latency. Hence, we propose innovations at the circuit, algorithm and architecture levels to accelerate the transformer. At the circuit level, we propose Topkima—combining top-kactivation selection with in-memory ADC (IMA) to implement efficient softmax without any sorting overhead. Only theklargest activations are sent to softmax calculation block, reducing the huge computational cost of softmax. At the algorithmic level, a modified training scheme utilizes top-kactivations only during the forward pass, combined with a sub-top-kmethod to address the crossbar size limitation by aggregating each sub-top-kvalues as global top-k. At the architecture level, we introduce a fine pipeline for efficiently scheduling data flows and an improved scale-free technique for removing scaling cost. The combined system, dubbed Topkima-Former, enhances$1.8\times -84\times $speedup and$1.2\times -36\times $energy efficiency (EE) over prior In-memory computing (IMC) accelerators. Compared to a conventional softmax macro and a digital top-k(Dtopk) softmax macro, our proposed Topkima softmax macro achieves about$15\times $and$8\times $faster speed respectively. Experimental evaluations demonstrate minimal (0.42% to 1.60%) accuracy loss for different models in both vision and NLP tasks. Xiaoqi Peng, Hongyang Shang, Ye Ke, Xiaofeng Yang 0004, Arindam Basu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |