EDBT 2026 Demo / reviewers in the wild / expert
Shucheng Du
dblp:318/4292
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 68% Hardware accelerators and domain-specific architectures · 29% Electronic design automation · 4% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › processing-in-memory
computing-in-memory |
1.6 | 2 | 2025 | Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-Memory · DAC 2025 CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts · Sci. China Inf. Sci. 2024 |
Memory systems › processing-in-memory › computing-in-memory
in-SRAM computing |
0.9 | 1 | 2025 | Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-Memory · DAC 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-Memory · DAC 2025 |
Memory systems
processing-in-memory |
0.9 | 1 | 2025 | Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-Memory · DAC 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
vision transformer accelerator |
0.9 | 1 | 2025 | Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-Memory · DAC 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.8 | 1 | 2024 | CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts · Sci. China Inf. Sci. 2024 |
Memory systems › processing-in-memory
ReRAM-based processing-in-memory |
0.8 | 1 | 2024 | CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts · Sci. China Inf. Sci. 2024 |
Electronic design automation
hardware/software co-design |
0.2 | 1 | 2024 | CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts · Sci. China Inf. Sci. 2024 |
Methods — techniques the papers use, named apart from their topics
neural architecture search · 1.5computing-in-memory · 1.5reconfigurable CIM macro · 0.9hybrid SRAM-RRAM architecture · 0.9fusion scheduling · 0.9decoupled chunk attention · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-MemoryabstractVision Transformers (ViTs) are new foundation models for vision applications. Edge-deploying ViTs to realize energy-saving, low-latency, and high-performance dense predictions have wide applications, such as autonomous driving and surveillance image analysis. However, the quadratic complexity of the self-attention mechanism renders ViTs slow and resource-intensive, particularly for pixel-level dense predictions that involve long contexts. Additionally, the pyramid-like architecture of modern ViT variants leads to an unbalanced workload, further reducing hardware utilization and decreasing the throughput of conventional edge devices. To this end, we propose an algorithm-hardware co-optimized edge ViT accelerator tailored for efficient dense predictions. At the algorithm level, we propose a decoupled chunk attention (DCA) mechanism implemented in a pipelined manner to reduce off-chip memory access, thereby enabling efficient dense predictions within limited on-chip memory. At the architecture level, we introduce a hybrid architecture that combines SRAM-based computing-in-memory (CIM) and nonvolatile RRAM storage to eliminate extensive off-chip memory access, with a fusion scheduling to balance workloads and minimize intermediate on-chip memory access. At the circuit level, a bit/element two-way-reconfigurable CIM macro is proposed to improve hardware utilization across pyramidal ViT blocks with varied matrix sizes. The experimental results on object detection, semantic segmentation, and depth estimation tasks demonstrate that our design can efficiently process patch lengths up to 16384 with a speedup of 18.5×-217.1×, a reduction in memory accesses of 1.7×-7.4×, and an improvement in energy efficiency of 1.8×, under less than 1% performance degradation. Yi Li 0049, Zijian Ye, Xiangqu Fu, Songqi Wang, Shucheng Du, Ning Lin, Dashan Shang, Jinshan Yue, Xiaojuan Qi 0001, Feng Zhang 0014 |
DAC | 5 |
| 2024 | CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-expertsabstractAbstract Artificial intelligence (AI) has experienced substantial advancements recently, notably with the advent of large-scale language models (LLMs) employing mixture-of-experts (MoE) techniques, exhibiting human-like cognitive skills. As a promising hardware solution for edge MoE implementations, the computing-in-memory (CIM) architecture collocates memory and computing within a single device, significantly reducing the data movement and the associated energy consumption. However, due to diverse edge application scenarios and constraints, determining the optimal network structures for MoE, such as the expert’s location, quantity, and dimension on CIM systems remains elusive. To this end, we introduce a software-hardware co-designed neural architecture search (NAS) framework, C IM-based M oE N AS (CMN), focusing on identifying a high-performing MoE structure under specific hardware constraints. The results of the NYUD-v2 dataset segmentation on the RRAM (SRAM) CIM system reveal that CMN can discover optimized MoE configurations under energy, latency, and performance constraints, achieving 29.67 × ( 43.10 ×) energy savings, 175.44 ×( 109.89 ×) speedup, and 12.24 × smaller model size compared to the baseline MoE-enabled Visual Transformer, respectively. This co-design opens up an avenue toward high-performance MoE deployments in edge CIM systems. Shihao Han, Sishuo Liu, Shucheng Du, Mingzi Li, Zijian Ye, Xiaoxin Xu, Dashan Shang |
Sci. China Inf. Sci. | 3 |
| 2024 | Erratum to: CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts
Shihao Han, Sishuo Liu, Shucheng Du, Mingzi Li, Zijian Ye, Xiaoxin Xu, Dashan Shang |
Sci. China Inf. Sci. | 3 |