EDBT 2026 Demo / reviewers in the wild / expert
Bo Zhao 0019
dblp:94/4810-19
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-0768-3444ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bat: Efficient Generative Recommender Serving with Bipartite AttentionabstractGenerative Recommenders (GRs) have recently emerged as promising alternatives to traditional Deep Learning Recommendation Models (DLRMs). Despite their potential, GRs remain computationally expensive in inference, exhibiting compute-bound characteristics similar to the prefill stage of Large Language Model (LLM) inference. Prefix caching can reduce redundant computation by reusing previously constructed KV caches. However, the unique properties of GRs, i.e., highly personalized user profiles and real-time item retrieval, make cache reuse across queries challenging, resulting in limited computational savings. Jie Sun 0017, Shaohang Wang, Zimo Zhang, Peng Sun 0006, Bo Zhao 0019, Bingsheng He, Fei Wu 0001, Zeke Wang |
ASPLOS (2) | 7 |
| 2026 | CHARM: Chiplet Heterogeneity-Aware Runtime Mapping SystemabstractPublisher Copyright: © 2026 Copyright held by the owner/author(s) Alessandro Fogli, Bo Zhao 0019, Peter R. Pietzuch, Jana Giceva |
EuroSys | 2 |
| 2026 | SHARP: Shared State Reduction for Efficient Matching of Sequential Patterns
Matthias Weidlich 0001, Bo Zhao 0019 |
Proc. VLDB Endow. | 4 |
| 2025 | Scalpel: High Performance Contention-Aware Task Co-Scheduling for Shared Cache HierarchyabstractFor scientific computing applications that consist of many loosely coupled tasks, efficient scheduling is critical to achieve high performance and good quality of service (QoS). One of the challenges for co-running tasks is the frequent contention for shared cache hierarchy of multi-core processors. Such contention significantly increases cache miss rate and therefore, results in performance deterioration for computational tasks. This paper presents Scalpel, a contention-aware task grouping and co-scheduling approach for efficient task scheduling on shared cache hierarchy. Scalpel utilizes the shared cache access features of tasks to group them in a heuristic way, which reduces the contention within groups by achieving equal shared cache locality, while maintaining load balancing between groups. Based thereon, it proposes a two-level scheduling strategy to schedule groups to processors and assign tasks to available cores in a timely manner, while considering the impact of task scheduling on shared cache locality to minimize task execution time. Experiments show that Scalpel reduces the shared cache miss rate by up to 2.14× and optimizes the execution time by up to 1.53× for scientific computing benchmarks, compared to several baseline approaches. Song Liu 0007, Zengyuan Zhang, Xinhe Wan, Bo Zhao 0019, Weiguo Wu |
IEEE Trans. Computers | 5 |
| 2024 | Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor CollectionsabstractDeep learning (DL) jobs use multi-dimensional parallelism, i.e., combining data, model, and pipeline parallelism, to use large GPU clusters efficiently. Long-running jobs may experience changes to their GPU allocation: (i) resource elasticity during training adds or removes GPUs; (ii) hardware maintenance may require redeployment on different GPUs; and (iii) GPU failures force jobs to run with fewer devices. Current DL frameworks tie jobs to a set of GPUs and thus lack support for these scenarios. In particular, they cannot change the multi-dimensional parallelism of an already-running job in an efficient and model-independent way. Marcel Wagenländer, Bo Zhao 0019, Luo Mai, Peter R. Pietzuch |
SOSP | 3 |
| 2024 | OLAP on Modern Chiplet-Based ProcessorsabstractChiplet-based CPUs, which combine multiple independent dies on a single package, allow hardware to scale to higher CPU core counts at the cost of more memory heterogeneity and performance variability. This introduces challenges when existing query engines are deployed on chiplet-based CPUs, as current designs make assumptions about uniform memory access, cache locality and consistent core performance, e.g., leading to ineffective CPU utilization. In this paper, we analyse the performance impact when query engines ignore chiplet-specific properties. We demonstrate that a naïve deployment can result in a significant degradation of query processing efficiency, exhibiting non-linear scaling even within a single CPU socket domain. Based on comprehensive experiments, we explore approaches to deploy query engines on chiplet-based CPUs with improved performance: we show that distributing processing tasks according to a chiplet-aware strategy achieves higher resource utilization and scalability, yielding an up to 7× speedup compared to hardware-oblivious approaches. Alessandro Fogli, Bo Zhao 0019, Peter R. Pietzuch, Maximilian Bandle, Jana Giceva |
Proc. VLDB Endow. | 2 |
| 2023 | MSRL: Distributed Reinforcement Learning with Dataflow Fragments
Huanzhou Zhu, Bo Zhao 0019, Yaodong Yang 0001, Peter R. Pietzuch |
USENIX ATC | 2 |
| 2023 | TurboStencil: You only compute once for stencil computation
Song Liu 0007, Xinhe Wan, Zengyuan Zhang, Bo Zhao 0019, Weiguo Wu |
Future Gener. Comput. Syst. | 4 |
| 2021 | EIRES: Efficient Integration of Remote Data in Event Stream ProcessingabstractTo support reactive and predictive applications, complex event processing (CEP) systems detect patterns in event streams based on predefined queries. To determine the events that constitute a query match, their payload data may need to be assessed together with data from remote sources. Such dependencies are problematic, since waiting for remote data to be fetched interrupts the processing of the stream. Yet, without event selection based on remote data, the query state to maintain may grow exponentially. In either case, the performance of the CEP system degrades drastically. Bo Zhao 0019, Han van der Aa, Thanh Tam Nguyen, Nguyen Quoc Viet Hung, Matthias Weidlich 0001 |
SIGMOD Conference | 1 |
| 2020 | Load Shedding for Complex Event Processing: Input-based and State-based TechniquesabstractComplex event processing (CEP) systems that evaluate queries over streams of events may face unpredictable input rates and query selectivities. During short peak times, exhaustive processing is then no longer reasonable, or even infeasible, and systems shall resort to best-effort query evaluation and strive for optimal result quality while staying within a latency bound. In traditional data stream processing, this is achieved by load shedding that discards some stream elements without processing them based on their estimated utility for the query result. We argue that such input-based load shedding is not always suitable for CEP queries. It assumes that the utility of each individual element of a stream can be assessed in isolation. For CEP queries, however, this utility may be highly dynamic: Depending on the presence of partial matches, the impact of discarding a single event can vary drastically. In this work, we therefore complement input-based load shedding with a state-based technique that discards partial matches. We introduce a hybrid model that combines both input-based and state-based shedding to achieve high result quality under constrained resources. Our experiments indicate that such hybrid shedding improves the recall by up to 14× for synthetic data and 11.4× for real-world data, compared to baseline approaches. Bo Zhao 0019, Nguyen Quoc Viet Hung, Matthias Weidlich 0001 |
ICDE | 1 |
| 2018 | Complex Event Processing under Constrained Resources by State-Based Load SheddingabstractComplex event processing (CEP) systems evaluate queries over event streams for low-latency detection of user-specified event patterns. They need to process streams of growing volume and velocity, while the heterogeneity of event sources yields unpredictable input rates. Against this background, models and algorithms for the optimisation of CEP systems have been proposed in the literature. However, when input rates grow by orders of magnitude during short peak times, exhaustive real-time processing of event streams becomes infeasible. CEP systems shall therefore resort to best-effort query evaluation, which maximises the accuracy of pattern detection while staying within a predefined latency bound. For traditional data stream processing, this is achieved by load shedding that drops some input data without processing it, guided by the estimated importance of particular data entities for the processing accuracy. In this work, we argue that such input-based load shedding is not suited for CEP queries in all situations. Unlike for relational stream processing, where the impact of shedding is assessed based on the operator selectivity, the importance of an event for a CEP query is highly dynamic and largely depends on the state of query processing. Depending on the presence of particular partial matches, the impact of dropping a single event can vary drastically. Hence, this PhD project is devoted to state-based load shedding that, instead of dropping input events, discards partial matches to realise best-effort processing under constrained resources. In this paper, we describe the addressed problem in detail, sketch our envisioned solution for state-based load shedding, and present preliminary experimental results that indicate the general feasibility of our approach. Bo Zhao 0019 |
ICDE | 1 |
| 2015 | Beyond Data Parallelism: Identifying Parallel Tasks in Sequential Programs
Zhen Li 0005, Bo Zhao 0019, Ali Jannesari, Felix Wolf 0001 |
ICA3PP (4) | 2 |