VLDB 2026 Research / reviewers in the wild / expert
Xingzhen Chen
dblp:140/7209
· DBLP profile ↗
10ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-4865-3708ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DORA: Dataflow-Instruction Orchestration Architecture for DNN AccelerationabstractModern DNN workloads are increasingly diverse in operation types, tensor shapes, and execution dependencies, making it difficult for customized accelerators to sustain high hardware efficiency across models. We propose DORA, an instruction-based overlay architecture as well as a compilation framework that explicitly describes dataflow via a proposed ISA, enabling fine-grained control of data movement, computation, and synchronization at the layer level. Xingzhen Chen, Zhuoping Yang, Jinming Zhuang, Shixin Ji, Sarah Schultz, Zheng Dong 0002, Weisong Shi, Peipei Zhou 0001 |
FCCM | 1 |
| 2026 | DORA: Dataflow-Instruction Orchestration Architecture for DNN AccelerationabstractAs deep neural networks develop significantly more diverse and complex, achieving high performance and efficiency on complicated DNN models faces pressing challenges. Modern DNN workloads are increasingly diverse in operation types, tensor shapes, and execution dependencies, making it difficult to sustain high hardware efficiency across models. In addition, a generic accelerator often incurs substantial overhead when executing diverse workloads. Xingzhen Chen, Zhuoping Yang, Jinming Zhuang, Shixin Ji, Sarah Schultz, Zheng Dong 0002, Weisong Shi, Peipei Zhou 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2026 | μ-ORCA: Optimizing Acceleration for Microsecond-Scale Deep Neural Network Inference on ACAPabstractHeterogeneous reconfigurable platforms with tensor cores, such as AMD ACAP, are increasingly adopted for deep neural network (DNN) inference due to their high throughput and flexibility. However, their suitability for microsecond-scale inference on small problem sizes remains underexplored. In jet-tagging applications in high-energy physics, inefficient on-chip communication and large inter-layer latency prevent existing frameworks from meeting the 1-μ s latency budget. Moreover, hardware overheads such as synchronization and VLIW processor prologue are often overlooked, making it infeasible to optimize accelerators correctly. To address these problems, we propose µ-ORCA, a customized heterogeneous accelerator framework for ultra-low-latency model inference. µ-ORCA enables direct inter-layer communication between DNN layers on the AIE array, instead of using shared memory tiles or FPGA fabric. Moreover, a 512-bit/cycle cascade connection is applied instead of a 32-bit/cycle DMA connection. µ-ORCA also provides an overhead-aware performance model that adapts to different NN layer sizes, and conducts design space exploration to optimize end-to-end latency. µ-ORCA supports MLP and DeepSets models with non-MM kernels, including bias, ReLU, and global aggregation on AIE. We evaluate µ-ORCA on the AMD ACAP VEK280 platform. Experimental results show that µ-ORCA achieves average latency reduction of > 1.70 × and > 1.83 × compared with different state-of-the-art ACAP frameworks, and achieves 0.93 μ s latency for a 6-layer real-world DeepSets model, satisfying the latency budget. We open source µ-ORCA at https://github.com/arc-research-lab/u-ORCA. Shixin Ji, Jinming Zhuang, Zhuoping Yang, Xingzhen Chen, Wei Zhang 0062, Peipei Zhou 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Towards Accelerator Customization in Real-time Safety-critical Systems
Shixin Ji, Xingzhen Chen, Wei Zhang 0062, Zhuoping Yang, Jinming Zhuang, Sarah Schultz, Yukai Song, Jingtong Hu, Alex K. Jones, Zheng Dong 0002, Peipei Zhou 0001 |
FPGA | 2 |
| 2025 | ART: Customizing Accelerators for DNN-Enabled Real-Time Safety-Critical Systems
Shixin Ji, Xingzhen Chen, Jinming Zhuang, Wei Zhang 0062, Zhuoping Yang, Sarah Schultz, Yukai Song, Jingtong Hu, Alex K. Jones, Zheng Dong 0002, Peipei Zhou 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | DERCA: DetERministic Cycle-Level Accelerator on Reconfigurable Platforms in DNN-Enabled Real-Time Safety-Critical SystemsabstractDeep neural network (DNN) models are increasingly deployed in real-time, safety-critical systems such as autonomous vehicles, driving the need for specialized AI accelerators. However, most existing accelerators support only non-preemptive execution or limited preemptive scheduling at the coarse granularity of DNN layers. This restriction leads to frequent priority inversion due to the scarcity of preemption points, resulting in unpredictable execution behavior and, ultimately, system failure. To address these limitations and improve the real-time performance of AI accelerators, we propose DERCA, a novel accelerator architecture that supports fine-grained, intra-layer flexible preemptive scheduling with cycle-level determinism. DERCA incorporates an on-chip Earliest Deadline First (EDF) scheduler to reduce both scheduling latency and variance, along with a customized dataflow design that enables intralayer preemption points (PPs) while minimizing the overhead associated with preemption. Leveraging the limited preemptive task model, we perform a comprehensive predictability analysis of DERCA, enabling formal schedulability analysis and optimized placement of preemption points within the constraints of limited preemptive scheduling. We implement DERCA on the AMD ACAP VCK190 reconfigurable platform. Experimental results show that DERCA outperforms state-of-the-art designs using non-preemptive and layer-wise preemptive dataflows, with less than 5 % overhead in worst-case execution time (WCET) and only 6% additional resource utilization. DERCA is open-sourced on GitHub: https://github.com/arc-research-lab/DERCA Shixin Ji, Zhuoping Yang, Xingzhen Chen, Wei Zhang 0062, Jinming Zhuang, Alex K. Jones, Zheng Dong 0002, Peipei Zhou 0001 |
RTSS | 3 |
| 2025 | AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationabstractGPUs are critical for compute-intensive applications, yet emerging workloads such as recommender systems, graph analytics, and data analytics often exceed GPU memory capacity. Existing solutions allow GPUs to use CPU DRAM or SSDs as external memory, and the GPU-centric approach enables GPU threads to directly issue NVMe requests, further avoiding CPU intervention. However, current GPU-centric approaches adopt synchronous I/O, forcing threads to stall during long communication delays. Zhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones, Peipei Zhou 0001 |
SC | 3 |
| 2024 | Challenges and Opportunities to Enable Large-Scale Computing via Heterogeneous ChipletsabstractFast-evolving artificial intelligence (AI) algorithms such as large language models have been driving the ever-increasing computing demands in today’s data centers. Heterogeneous computing with domain-specific architectures (DSAs) brings many opportunities when scaling up and scaling out the computing system. In particular, heterogeneous chiplet architecture is favored to keep scaling up and scaling out the system as well as to reduce the design complexity and the cost stemming from the traditional monolithic chip design. However, how to interconnect computing resources and orchestrate heterogeneous chiplets is the key to success. In this paper, we first discuss the diversity and evolving demands of different AI workloads. We discuss how chiplet brings better cost efficiency and shorter time to market. Then we discuss the challenges in establishing chiplet interface standards, packaging, and security issues. We further discuss the software programming challenges in chiplet systems. Zhuoping Yang, Shixin Ji, Xingzhen Chen, Jinming Zhuang, Dharmesh Jani, Peipei Zhou 0001 |
ASPDAC | 3 |
| 2022 | INFless: a native serverless system for low-latency, high-throughput inferenceabstractModern websites increasingly rely on machine learning (ML) to improve their business efficiency. Developing and maintaining ML services incurs high costs for developers. Although serverless systems are a promising solution to reduce costs, we find that the current general purpose serverless systems cannot meet the low latency, high throughput demands of ML services. Laiping Zhao, Xingzhen Chen, Keqiu Li |
ASPLOS | 7 |
| 2014 | BigOP: Generating Comprehensive Big Data Workloads as a Benchmarking Framework
Yuqing Zhu 0001, Jianfeng Zhan, Chuliang Weng, Raghunath Othayoth Nambiar, Jinchao Zhang 0001, Xingzhen Chen, Lei Wang 0004 |
DASFAA (2) | 6 |