EDBT 2026 Demo / reviewers in the wild / expert
Xilong Kang
dblp:398/9942
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0002-3853-2398ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D-TANoC: Thermal-Aware 3D LLM Accelerator with Hierarchical NoC and Operator-Aware Dataflow Mapping Strategy
Xingyu Xu 0008, Dengke Liu, Zihan Zou, Xilong Kang, Hui Kou, Hao Cai 0001, Bo Liu 0019 |
ISCAS | 5 |
| 2026 | Low Bit-Width LLM Acceleration via Symmetric Lookup Format and Compute-in-Decoding Paradigm
Zihan Zou, Jiaming Lin, Xinming Yan, Shikuang Chen, Chen Zhang 0025, Xilong Kang, Hao Cai 0001, Bo Liu 0019 |
IEEE Trans. Computers | 8 |
| 2025 | SUArch: Accelerating Layer-wise N: M Sparse Pattern with a Unified Architecture for Deep-learning Edge DeviceabstractDeep neural networks are of the essence for user applications on edge devices. However, the computation and memory-intensive nature of deep neural networks conflicts with the resource-constrained devices. Moreover, the heterogeneity across different models imposes new challenges on deployment on edge devices. To boost the capabilities of edge devices, we propose SUArch, which innovates on three fronts: 1) a layer-wise N:M sparsity aware training approach to strike a balance between accuracy and training cost; 2) a sparsity alignment unit based on the butterfly network to maximize hardware utilization and eliminate extra overhead; 3) a mode-heterogenous processing element array to effectively accomplish the unified support for Convolution Neural Network and Transformer. The experimental results demonstrate that when running convolution-based and attention-based models under an industrial 28-nm process, the proposed SUArch realizes an energy efficiency of 52.1 TOPS/W. Compared to state-of-the-art architecture, SUArch achieves an energy efficiency improvement of 2.07× while accuracy loss is within 0.7%. Xilong Kang, Qingwen Wei, Ningyuan Li 0004, Xingyu Xu 0008, Hao Cai 0001, Bo Liu 0019 |
ASP-DAC | 1 |
| 2025 | A Layer-wise N: M Sparsity Aware Transformer Accelerator leveraging Temporal Locality with Butterfly NetworkabstractIn recent years, Transformer-based neural networks have driven significant advancements across various fields. However, the attention mechanism results in high memory access and computational demands, making deployment on edge devices challenging. To address these issues, we propose a layer-wise N:M sparsity-aware Transformer accelerator with two key innovations: 1) a layer-wise N:M sparsity-based parallel dimension compression strategy that effectively leverages the temporal locality of feature values, and 2) a butterfly network-based scattering unit that ensures correct accumulation of partial products after multiplication. Experimental results demonstrate that the proposed accelerator achieves up to a 3.86× improvement in runtime cycles. When running Deit-s on Imagenet-1k dataset, the accelerator delivers an energy efficiency of up to 34.4 TOPS/W at 68.9% sparsity under an industrial 28-nm technology, with only a 0.39% accuracy loss, representing a 1.46~1.48× improvement over prior work. Qinfan Wang, Xilong Kang, Qingwen Wei, Cai Hao, Bo Liu 0019 |
ISCAS | 3 |