EDBT 2026 Demo / reviewers in the wild / expert
Wei Zhang 0062
dblp:10/4661-62
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0004-4590-6559ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | μ-ORCA: Optimizing Acceleration for Microsecond-Scale Deep Neural Network Inference on ACAPabstractHeterogeneous reconfigurable platforms with tensor cores, such as AMD ACAP, are increasingly adopted for deep neural network (DNN) inference due to their high throughput and flexibility. However, their suitability for microsecond-scale inference on small problem sizes remains underexplored. In jet-tagging applications in high-energy physics, inefficient on-chip communication and large inter-layer latency prevent existing frameworks from meeting the 1-μ s latency budget. Moreover, hardware overheads such as synchronization and VLIW processor prologue are often overlooked, making it infeasible to optimize accelerators correctly. To address these problems, we propose µ-ORCA, a customized heterogeneous accelerator framework for ultra-low-latency model inference. µ-ORCA enables direct inter-layer communication between DNN layers on the AIE array, instead of using shared memory tiles or FPGA fabric. Moreover, a 512-bit/cycle cascade connection is applied instead of a 32-bit/cycle DMA connection. µ-ORCA also provides an overhead-aware performance model that adapts to different NN layer sizes, and conducts design space exploration to optimize end-to-end latency. µ-ORCA supports MLP and DeepSets models with non-MM kernels, including bias, ReLU, and global aggregation on AIE. We evaluate µ-ORCA on the AMD ACAP VEK280 platform. Experimental results show that µ-ORCA achieves average latency reduction of > 1.70 × and > 1.83 × compared with different state-of-the-art ACAP frameworks, and achieves 0.93 μ s latency for a 6-layer real-world DeepSets model, satisfying the latency budget. We open source µ-ORCA at https://github.com/arc-research-lab/u-ORCA. Shixin Ji, Jinming Zhuang, Zhuoping Yang, Xingzhen Chen, Wei Zhang 0062, Peipei Zhou 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | Towards Accelerator Customization in Real-time Safety-critical Systems
Shixin Ji, Xingzhen Chen, Wei Zhang 0062, Zhuoping Yang, Jinming Zhuang, Sarah Schultz, Yukai Song, Jingtong Hu, Alex K. Jones, Zheng Dong 0002, Peipei Zhou 0001 |
FPGA | 3 |
| 2025 | ART: Customizing Accelerators for DNN-Enabled Real-Time Safety-Critical Systems
Shixin Ji, Xingzhen Chen, Jinming Zhuang, Wei Zhang 0062, Zhuoping Yang, Sarah Schultz, Yukai Song, Jingtong Hu, Alex K. Jones, Zheng Dong 0002, Peipei Zhou 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | DERCA: DetERministic Cycle-Level Accelerator on Reconfigurable Platforms in DNN-Enabled Real-Time Safety-Critical SystemsabstractDeep neural network (DNN) models are increasingly deployed in real-time, safety-critical systems such as autonomous vehicles, driving the need for specialized AI accelerators. However, most existing accelerators support only non-preemptive execution or limited preemptive scheduling at the coarse granularity of DNN layers. This restriction leads to frequent priority inversion due to the scarcity of preemption points, resulting in unpredictable execution behavior and, ultimately, system failure. To address these limitations and improve the real-time performance of AI accelerators, we propose DERCA, a novel accelerator architecture that supports fine-grained, intra-layer flexible preemptive scheduling with cycle-level determinism. DERCA incorporates an on-chip Earliest Deadline First (EDF) scheduler to reduce both scheduling latency and variance, along with a customized dataflow design that enables intralayer preemption points (PPs) while minimizing the overhead associated with preemption. Leveraging the limited preemptive task model, we perform a comprehensive predictability analysis of DERCA, enabling formal schedulability analysis and optimized placement of preemption points within the constraints of limited preemptive scheduling. We implement DERCA on the AMD ACAP VCK190 reconfigurable platform. Experimental results show that DERCA outperforms state-of-the-art designs using non-preemptive and layer-wise preemptive dataflows, with less than 5 % overhead in worst-case execution time (WCET) and only 6% additional resource utilization. DERCA is open-sourced on GitHub: https://github.com/arc-research-lab/DERCA Shixin Ji, Zhuoping Yang, Xingzhen Chen, Wei Zhang 0062, Jinming Zhuang, Alex K. Jones, Zheng Dong 0002, Peipei Zhou 0001 |
RTSS | 4 |
| 2024 | Reducing Smart Phone Environmental Footprints with In-Memory ProcessingabstractSmart phones have revolutionized the availability of computing to the consumer. Recently, smart phones have been aggressively integrating artificial intelligence (AI) capabilities into their devices. The custom designed processors for the latest phones integrate incredibly capable and energy efficient graphics processors (GPUs) and tensor processors (TPUs) to accommodate this emerging AI workload and on-device inference. Unfortunately, smart phones are far from sustainable and have a substantial carbon footprint that continues to be dominated by environmental impacts from their manufacture and far less so by the energy required to power their operation. In this paper we explore the possibility of reversing the trend to increase the dedicated silicon dedicated to emerging application workloads in the phone. Instead we consider how in-memory processing using the DRAM already present in the phone could be used in place of dedicated GPU/TPU devices for AI inference. We explore the potential savings in embodied carbon that could be possible with this tradeoff and provide some analysis of the potential of in-memory computing to compete with these accelerators. While it may not be possible to achieve the same throughput, we suggest that the responsiveness to the user may be sufficient using in-memory computing, while both the embodied and operational carbon footprints could be improved. Our approach can save circa $10-15 \mathrm{~kg} \mathrm{CO}_{2}$. Zhuoping Yang, Wei Zhang 0062, Shixin Ji, Peipei Zhou 0001, Alex K. Jones |
CODES+ISSS | 2 |
| 2011 | Context, Computation, and Optimal ROC Performance in Hierarchical Models
Lo-Bin Chang, Ya Jin, Wei Zhang 0062, Eran Borenstein, Stuart Geman |
Int. J. Comput. Vis. | 3 |