EDBT 2026 Demo / reviewers in the wild / expert
Dehao Kong
dblp:255/3403
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WaferSim: A Simulation Infrastructure for LLM Service on Wafer-Scale Chips
Dehao Kong, Jiamu Fu, Jian Weng 0002, Jinyi Deng, Yang Hu 0001, Shouyi Yin |
APPT | 3 |
| 2026 | Characterizing Cloud-Native LLM Inference at Bytedance and Exposing Optimization Challenges and Opportunities for Future AI AcceleratorsabstractAs a major provider of LLM inference services, ByteDance has continuously explored diverse accelerator options to meet the rapidly growing inference demands of various heterogeneous LLM scenarios with higher cost-effectiveness, thereby enabling LLMs to serve more people worldwide. However, during this process, we have found that the complexity and opacity of cloud scenarios and corresponding cloud accelerators make it difficult for academia and many innovative chip startups to fully understand the real demands and challenges of these scenarios, which in turn severely restricts innovation and application potential in this field. To bridge this gap, we first present and analyze the data and characteristics of the ByteDance Doubao LLM app across multiple dimensions, helping the community understand real-world cloud scenarios, and detail the challenges and opportunities we have identified. Second, we propose and plan to open-source our multi-level evaluation framework, XPU-Perf, which includes benchmarks spanning instructions, operators, and models. This framework improves interpretability and trustworthiness, and helps promising new accelerator architectures gain wider adoption and development. Finally, we present comparative results of four typical accelerators, summarize their shortcomings and challenges, conduct in-depth analysis, and highlight numerous architectural and scheduling innovation opportunities we have observed. Jingwei Cai, Dehao Kong, Hantao Huang, Zishan Jiang, Zixuan Ma, Qingyu Guo, Guiming Shi, Mingyu Gao 0001, Kaisheng Ma, Minghui Yu |
HPCA | 2 |
| 2026 | FACE: Fully Overlapped PD Scheduling and Multi-Level Architecture Co-Exploration on WaferabstractThe rapid expansion of large language models (LLMs) parameter scales imposes unprecedented demands on compute, memory, and communication resources for inference deployment. Wafer-scale chips, leveraging advanced packaging technologies, deliver high-density integration of compute and memory with high die-to-die (D2D) communication bandwidth, providing a compelling architectural approach to satisfy these resource requirements. However, its unprecedented chip area introduces significant architectural design complexities. Waferscale chips feature a multi-level architecture spanning the wafer, die, and core levels, involving numerous critical design parameters and trade-offs, which still lack systematic understanding and exploration. Moreover, this poses major challenges for LLM serving scheduling. Existing methods, largely adapted from GPUbased systems, fail to fully leverage the advantages of waferscale chips and mitigate their limitations, making it difficult to efficiently translate massive hardware resources into actual performance gains. To address these challenges, we introduce FACE, a coexploration framework for jointly optimizing multi-level architecture and serving scheduling. We first establish a flexible and extensible wafer-scale hardware template to systematically explore the optimal architecture and micro-architecture parameters. Leveraging the fine-grained control and high interconnect bandwidth of wafer-scale chips, FACE implements an LLM scheduling strategy that achieves fully overlapped prefill-decode execution and efficient KV cache management, maximizing hardware resource utilization to improve LLM service quality. Our evaluation demonstrates that FACE can achieve an average overall performance improvement of 3.68 × across various LLM models and datasets compared to the state-of-the-art (SOTA) LLM serving system on wafer-scale chips. Dehao Kong, Dingcheng Jiang, Jinyi Deng, Yang Hu 0001, Shouyi Yin |
HPCA | 2 |
| 2025 | RAM-Wafer: RL-Based Automatic Mapping Framework for Large-Scale AI Training on Wafer-Scale ComputingabstractWafer-scale computing, with its high integration density and die-to-die bandwidth, offers a promising solution to the exponentially growing computational demands of large AI models. However, mapping large-scale AI training workloads onto wafer-scale architectures poses unique challenges compared to traditional GPU or AI accelerator clusters-namely, limited on-chip memory, non-uniform collective communication, and an exponentially large, sparsely populated search space. To address these challenges, we introduce RAM-Wafer, an innovative reinforcement learning (RL)-based automatic mapping framework designed specifically for wafer-scale computing. Built by extending the production-level AI compiler framework OpenXLA, our compiler-based end-to-end mapping solution incorporates an accurate and fast performance model that accounts for the distinctive constraints of wafer-scale systems. Moreover, our RL-based mapping method efficiently explores the vast search space to identify near-optimal mapping solutions in a fraction of the time required by conventional methods. Extensive experimental results demonstrate that RAM-Wafer outperforms manual expert baselines by 22 % and 9.5 % on Dojo and Waferscale GPU platforms, respectively, and achieves improvements of 15 % and 6.5 % compared to a genetic algorithm (GA). Additionally, RAM-Wafer reduces search time by$28 \times$, cutting mapping time from 2.3 hours to just 5 minutes. Dehao Kong, Xufeng He, Shaopeng Zhai, Yang Hu 0001, Shouyi Yin |
ICCD | 2 |
| 2025 | WSC-LLM: Efficient LLM Service and Architecture Co-exploration for Wafer-scale ChipsabstractThe deployment of large language models (LLMs) imposes significant demands on computing, memory, and communication resources.Wafer-scale technology enables the high-density integration of multiple single-die chips with high-speed Die-to-Die (D2D) interconnections, presenting a promising solution to meet these demands arising from LLMs.However, given the limited wafer area, a trade-off needs to be made among computing, storage, and communication resources.Maximizing the benefits and minimizing the drawbacks of wafer-scale technology is crucial for enhancing the performance of LLM service systems, which poses challenges to both architecture and scheduling.Unfortunately, existing methods cannot effectively address these challenges.To bridge the gap, we propose WSC-LLM, an architecture and scheduling co-exploration framework.We first define a highly configurable general hardware template designed to explore optimal architectural parameters for wafer-scale chips.Based on it, we Dehao Kong, Jingxiang Hou, Chao Li 0009, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ISCA | 2 |
| 2024 | Fuzzy Boundary-Guided Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) is a challenging task that identifies camouflaged objects from highly similar backgrounds. Existing methods typically treat the whole object equally while neglecting the indistinguishable regions that require more attention than other regions. In this paper, we propose a Fuzzy Boundary-Guided Network (FBG-Net) for camouflaged object detection, which mimics the human behavior that pays more attention to these low-confidence regions when observing objects. Specifically, we devise two main building blocks: (1) Mixed Semantics Aggregation Module (MSAM) to integrate boundary and texture features cumulatively in the high-to-low scales, and (2) Fuzzy Boundary-Guided Module (FBGM) to locate and enhance the low-confidence regions under the guidance of fuzzy boundary. Extensive experiments demonstrate the effectiveness of FBG-Net with superior performance to existing state-of-the-art methods. Code is available at https://github.com/YAOSL98/FBG-Net. Qi Jia 0001, Shuilian Yao, Youcan Xu, Yu Liu 0012, Dehao Kong, Longin Jan Latecki |
ICME | 5 |
| 2019 | Operating Modes Analysis and Control Strategy of Single-Stage Isolated Modular Multilevel Converter (I-MMC) for Medium Voltage AC/DC GridabstractFor the medium-voltage (MV) AC/DC grid integrating with renewable sources, this paper introduces a single-stage isolated modular multilevel converter (I-MMC) topology. Compared with the conventional modular multilevel converter with DC transformer, I-MMC has the advantages of smaller size and simpler control strategy, which can realize single-stage power conversion among MVAC, MVDC and LVDC ports. Its basic circuit working process is given and three representative working modes of the I-MMC are discussed in detail. Coordinative control strategy based on double modulation freedoms is proposed for different modes of I-MMC. Moreover, an (10kV AC, ±10kV DC, 760V DC) I-MMC system is established in MATLAB/SIMULINK, and a scale-down experimental platform is built to verify the proposed control strategy. Dehao Kong, Hong Ying, Zhongchen Pei |
IECON | 1 |