EDBT 2026 Demo / reviewers in the wild / expert
Dingcheng Jiang
dblp:380/5596
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0004-4379-694XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XY-Serve: End-to-End Versatile Production Serving for Dynamic LLM WorkloadsabstractMeeting growing demands for low latency and cost efficiency in production-grade large language model (LLM) serving systems requires integrating advanced optimization techniques. However, dynamic and unpredictable input-output lengths of LLM, compounded by these optimizations, exacerbate the issues of workload variability, making it difficult to maintain high efficiency on AI accelerators, especially DSAs with tile-based programming models. To address this challenge, we introduce XY-Serve, a versatile, Ascend NPU native, end-to-end production LLM-serving system. The core idea is an abstraction mechanism that smooths out the workload variability by decomposing computations into unified, hardware-friendly, fine-grained meta primitives. Then, kernels can efficiently execute without concerning the irregularity of workload. After this abstraction mechanism, for Attention, we propose a meta-kernel that computes the basic pattern of GEMM-Softmax-GEMM with architectural-aware tile sizes. For Linear, we introduce a virtual padding scheme that adapts to dynamic shape changes while using highly efficient GEMM primitives with assorted fixed tile sizes. XY-Serve sits harmoniously with vLLM. Experimental results show up to 95% end-to-end throughput improvement compared with current publicly available baselines on Ascend NPUs. We also set a new performance record for Linear (average 14.6% faster) and Attention (average 21.5% faster) kernels relative to existing libraries. Lastly, we demonstrate the generality of our technologies on GPU platform. Mingcong Song, Xinru Tang, Fengfan Hou, Yipeng Ma, Runqiu Xiao, Hongjie Si, Dingcheng Jiang, Shouyi Yin, Yang Hu 0001, Guoping Long |
ASPLOS (1) | 9 |
| 2026 | MoEntwine: Unleashing the Potential of Wafer-Scale Chips for Large-Scale Expert Parallel InferenceabstractAs large language models (LLMs) continue to scale up, mixture-of-experts (MoE) has become a common technology in SOTA models. MoE models rely on expert parallelism (EP) to alleviate memory bottleneck, which introduces all-to-all communication to dispatch and combine tokens across devices. However, in widely-adopted GPU clusters, high-overhead crossnode communication makes all-to-all expensive, hindering the adoption of EP. Recently, wafer-scale chips (WSCs) have emerged as a platform integrating numerous devices on a wafer-sized interposer. WSCs provide a unified high-performance network connecting all devices, presenting a promising potential for hosting MoE models. Yet, their network is restricted to a mesh topology, causing imbalanced communication pressure and performance loss. Moreover, the lack of on-wafer disk leads to high-overhead expert migration on the critical path. To fully unleash this potential, we first propose Entwined Ring Mapping (ER-Mapping), which co-designs the mapping of attention and MoE layers to balance communication pressure and achieve better performance. We find that under ER-Mapping, the distribution of cold and hot links in the attention and MoE layers is complementary. Therefore, to hide the migration overhead, we propose the Non-invasive Balancer (NI-Balancer), which splits a complete expert migration into multiple steps and alternately utilizes the cold links of both layers. Evaluation shows ER-Mapping achieves communication reduction up to 62 %. NIBalancer further delivers 54 % and 22 % improvements in MoE computation and communication, respectively. Compared with the SOTA NVL72 supernode, the WSC platform delivers an average 39 % higher per-device MoE performance owing to its scalability to larger EP. Xinru Tang, Jingxiang Hou, Dingcheng Jiang, Taiquan Wei, Jinyi Deng, Huizheng Wang, Qize Yang, Haoran Shang, Chao Li 0009, Yang Hu 0001, Shouyi Yin |
HPCA | 3 |
| 2026 | TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsabstractLarge language models (LLMs) demand significant memory and computation resources. Wafer-scale chips (WSCs) provide high computation power and die-to-die (D2D) bandwidth but face a unique trade-off between on-chip memory and compute resources due to limited wafer area. Therefore, tensor parallelism strategies for wafer should leverage communication advantages while maintaining memory efficiency to maximize WSC performance. However, existing approaches fail to address these challenges. To address these challenges, we propose the tensor stream partition paradigm (TSPP), which reveals an opportunity to leverage WSCs' abundant communication bandwidth to alleviate stringent on-chip memory constraints. However, the 2D mesh topology of WSCs lacks long-distance and flexible interconnects, leading to three challenges: 1) severe tail latency, 2) prohibitive D2D traffic contention, and 3) intractable search time for optimal design. We present TEMP, a framework for LLM training on WSCs that leverages topology-aware tensor-stream partition, trafficconscious mapping, and dual-level wafer solving to overcome hardware constraints and parallelism challenges. These integrated approaches optimize memory efficiency and throughput, unlocking TSPP's full potential on WSCs. Evaluations show TEMP achieves$1.7 \times$average throughput improvement over state-of-the-art LLM training systems across various models. Huizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang, Qize Yang, Jingxiang Hou, Chao Li 0009, Jinyi Deng, Yang Hu 0001, Shouyi Yin |
HPCA | 4 |
| 2026 | FACE: Fully Overlapped PD Scheduling and Multi-Level Architecture Co-Exploration on WaferabstractThe rapid expansion of large language models (LLMs) parameter scales imposes unprecedented demands on compute, memory, and communication resources for inference deployment. Wafer-scale chips, leveraging advanced packaging technologies, deliver high-density integration of compute and memory with high die-to-die (D2D) communication bandwidth, providing a compelling architectural approach to satisfy these resource requirements. However, its unprecedented chip area introduces significant architectural design complexities. Waferscale chips feature a multi-level architecture spanning the wafer, die, and core levels, involving numerous critical design parameters and trade-offs, which still lack systematic understanding and exploration. Moreover, this poses major challenges for LLM serving scheduling. Existing methods, largely adapted from GPUbased systems, fail to fully leverage the advantages of waferscale chips and mitigate their limitations, making it difficult to efficiently translate massive hardware resources into actual performance gains. To address these challenges, we introduce FACE, a coexploration framework for jointly optimizing multi-level architecture and serving scheduling. We first establish a flexible and extensible wafer-scale hardware template to systematically explore the optimal architecture and micro-architecture parameters. Leveraging the fine-grained control and high interconnect bandwidth of wafer-scale chips, FACE implements an LLM scheduling strategy that achieves fully overlapped prefill-decode execution and efficient KV cache management, maximizing hardware resource utilization to improve LLM service quality. Our evaluation demonstrates that FACE can achieve an average overall performance improvement of 3.68 × across various LLM models and datasets compared to the state-of-the-art (SOTA) LLM serving system on wafer-scale chips. Dehao Kong, Dingcheng Jiang, Jinyi Deng, Yang Hu 0001, Shouyi Yin |
HPCA | 4 |
| 2025 | Cramming a Data Center into One Cabinet, a Co-Exploration of Computing and Hardware Architecture of Waferscale ChipabstractThe rapid advancements in large language models (LLMs) have significantly increased hardware demands.Wafer-scale chips, which integrate numerous compute units on an entire wafer, offer a highdensity computing solution for data centers and can extend Moore's Law at system level.However, current wafer-scale data center architectures face inefficiencies, such as uncoordinated resource allocation and lack of co-optimization for system area, preventing optimal integration density and performance within given cost and physical constraints.We propose a co-exploration approach of computing and hardware architectures to bridge this gap.We first develop an optimized wafer-scale single-cabinet data center model, integrating configurable on-chip memory dies and employing a vertically stacked hardware architecture.Based on this model, we introduce Titan, an automated exploration framework for intra-chip and inter-chip architecture design and optimization.Based on the architecture features of wafer-scale systems with optimal integration density, Titan establishes parameter dependencies to co-design the computing and hardware architectures.To reduce the design cycle for wafer-scale systems, Titan introduces vertical area constraints and pre-checks physical limits by integrating a series of reliability prediction models.It also integrates hardware Xingmao Yu, Dingcheng Jiang, Jinyi Deng, Chao Li 0009, Shouyi Yin, Yang Hu 0001 |
ISCA | 2 |
| 2025 | An Energy-Efficient, High-Frame-Rate, and Reconfigurable EKF-SLAM Processor With Full Acceleration for Autonomous Mobile RobotsabstractIn many intelligent edge applications involving Autonomous Mobile Robots (AMRs), efficient and real-time localization and mapping is a fundamental issue. Extended Kalman Filter Simultaneous Localization and Mapping (EKFSLAM) algorithm is a classic and successful solution to realize localization and mapping, while it is computationally intensive and poses a challenge for real-time tasks in small and micro robots. To address this issue, this work proposes an energy-efficient, highframe-rate, and reconfigurable EKF-SLAM processor. Firstly, a heterogeneous dual-core architecture is proposed to enable full acceleration of both matrix operations and nonlinear calculations in EKF-SLAM at the hardware architecture level. Secondly, a Reconfigurable Matrix Accelerator (RMA) and Reconfigurable Nonlinear Accelerator (RNA) are proposed to maximize data reuse and support diverse nonlinear functions at the data flow level. Thirdly, a data property-aware strategy is proposed at the data property level, which exploits matrix symmetry, sparsity, and dependency to reduce storage significantly and eliminate redundant computations. FPGA validation results show that the proposed design can achieve a frame rate of 774 fps and an energy efficiency of 0.66 mJ/frame, when performing mapping processes involving 60 landmarks at 100 MHz. Bingqiang Liu, Yequan Zhao, Minjie Bao, Zhendong Fan, Dingcheng Jiang, Zixuan Shen, Yulong Tan, Zaisheng He, Dengke Xu, Ke Wang 0028, Chao Wang 0096, Lining Sun |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | Live Demonstration: A Reconfigurable, Energy-efficient and High-frame-rate EKF-SLAM Accelerator Based SoC Design for Autonomous Mobile Robot ApplicationsabstractThis demonstration shows a Extend Kalman Filter-Simultaneous Localization And Mapping (EKF-SLAM) accelerator based System On Chip (SoC) design for Autonomous Mobile Robots (AMR). The AMR platform consists of a multi-sensor system with a wheel encoder and LiDAR, and a ZYNQ-7000 FPGA based SoC featuring an EKF-SLAM hardware accelerator. This AMR system achieves real-time SLAM with significant energy efficient improvement against the state-of-the-art designs. Dingcheng Jiang, Bingqiang Liu, Ao Hu, Yequan Zhao, Minjie Bao, Zhendong Fan, Zixuan Shen, Ke Wang 0028, Chao Wang 0096 |
ISCAS | 1 |