EDBT 2026 Demo / reviewers in the wild / expert
Zhantong Zhu
dblp:401/6284
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0002-4216-7581ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chat-A2: An LLM-aided Design Space Exploration Framework for High-Performance CPU DesignabstractMulti-objective design space exploration (DSE) for complex high-performance CPUs presents a significant challenge due to extensive parameter ranges and the vastness of the design space. Prior works either overlook microarchitectural information during DSE or necessitate intricate analysis for power, performance and area (PPA) evaluations. In this work, we present a novel DSE methodology for CPUs, which leverages large language models (LLMs) as an assistive tool to accelerate and automate the DSE flow. We develop an LLM-aided architecture-oriented DSE framework, i.e., Chat-A2, with satisfactory effectiveness and flexibility. Chat- $\mathbf{A}^{\mathbf{2}}$ is able to quickly explore 98% of the Pareto hyper-volume covered by the real Pareto optimal set without utilizing pre-sampled datasets or complex, customized analytical mechanisms. Experiments on an open-source high-performance RISC-V XiangShan CPU conclude that Chat-A ${ }^{\mathbf{2}}$ can obtain up to 19% normalized performance improvement while achieving 21.3% normalized area reduction compared to the default architecture optimized by professional human architects. Zhantong Zhu, Kangbo Bai |
ASP-DAC | 1 |
| 2026 | ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM TrainingabstractIn large-scale distributed LLM training, communication between devices becomes the key performance bottleneck. Chiplet technology can integrate multiple dies into a package to scale-up node performance with higher bandwidth. Meanwhile, optical interconnect (OI) technology offers long-reach, high-bandwidth links, making it well suited for scale-out networks. The combination of these two technologies has the potential to overcome communication bottlenecks within and across packages. In this work, we present ChipLight, a cross-layer multi-objective design and optimization method for training clusters leveraging chiplet and OI. We first abstract an architecture model for such complex clusters, co-optimizing chiplet architecture, training parallel strategy, and OI network topology. Based on such models, we tailor the design space exploration flow by combining both black-box and white-box methodologies. Evaluated by our experimental results, ChipLight achieves significantly improved training efficiency and provides valuable design insights for the development of future training clusters. Kangbo Bai, Zhantong Zhu |
DATE | 2 |
| 2026 | ACES: A Chiplet Architecture with Resource Partition and Dynamic Scheduling for Agentic LLMsabstractAgentic LLM is an emerging working paradigm leveraging large language models (LLMs) as assistant agents for complex, multi-step tasks. However, the operations of LLM agents also create unique workload characteristics with highly dynamic resource demands. In this work, we propose ACES, a co-designed solution leveraging scalable chiplet architecture together with dynamic workload scheduling for agentic LLMs. At hardware level, the chiplet architecture is designed to support a zoned fabric with flexible Swing Zones. Upon this, at software level, a conversation-centric dynamic scheduling approach is adopted, which includes topology-aware homing, proactive caching, and adaptive resource conversion to accommodate to data locality and resource balance. We evaluate our architecture using Llama-3 8B models on representative agentic-RAG-derived tasks. The system evaluations demonstrate that our design achieves 2.33× throughput improvement and an average 58% conversation latency reduction compared to state-of-the-art DistServe-like chiplet baselines, showcasing superior performance and scalability. Hongou Li, Zhantong Zhu |
DATE | 3 |
| 2025 | Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUsabstractWith the rapid advent of generative models, efficiently deploying these models on specialized hardware has become critical. Tensor Processing Units (TPUs) are designed to accelerate AI workloads, but their high power consumption neces-sitates innovations for improving efficiency. Compute-in-memory (CIM) has emerged as a promising paradigm with superior area and energy efficiency. In this work, we present a TPU architecture that integrates digital CIM to replace conventional digital systolic arrays in matrix multiply units (MXUs). We first establish a CIM-based TPU architecture model and simulator to evaluate the benefits of CIM for diverse generative model inference. Building upon the observed design insights, we further explore various CIM-based TPU architectural design choices. Up to 44.2% and 33.8% performance improvement for large language model and diffusion transformer inference, and 27.3 × reduction in MXU energy consumption can be achieved with different design choices, compared to the baseline TPUv4i architecture. Zhantong Zhu, Hongou Li, Wenjie Ren, Meng Wu 0005, Le Ye, Ru Huang 0001 |
DATE | 1 |
| 2025 | LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-OptimizationabstractLLM inference on mobile devices faces extraneous challenges due to limited memory bandwidth and computational resources. To address these issues, speculative inference and processing-in-memory (PIM) techniques have been explored at the algorithmic and hardware levels. However, speculative inference results in more compute-intensive GEMM operations, creating new design trade-offs for existing GEMV-accelerated PIM architectures. Furthermore, there exists a significant amount of redundant draft tokens in tree-based speculative inference, necessitating efficient token management schemes to minimize energy consumption. In this work, we present LP-Spec, an architecture-dataflow co-design leveraging hybrid LPDDR5 performance-enhanced PIM architecture with draft token pruning and dynamic workload scheduling to accelerate LLM speculative inference. A near-data memory controller is proposed to enable data reallocation between DRAM and PIM banks. Furthermore, a data allocation unit based on the hardware-aware draft token pruner is developed to minimize energy consumption and fully exploit parallel execution opportunities. Compared to end-to-end LLM inference on other mobile solutions such as mobile NPUs or GEMV-accelerated PIMs, our LP-Spec achieves 13.21× , 7.56 ×, and 99.87× improvements in performance, energy efficiency, and energy-delay-product (EDP). Compared with prior AttAcc PIM and RTX 3090 GPU, LP-Spec can obtain 12.83× and 415.31× EDP reduction benefits. Zhantong Zhu, Yandong He |
ICCAD | 2 |