EDBT 2026 Demo / reviewers in the wild / expert
Qinwen Shi
dblp:234/8525
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BTune: Bottleneck-Centric Configuration Parameters Tuning for Mixed WorkloadsabstractMany applications rely heavily on configuration tuning to improve throughput and reduce latency. The existing state-of-the-art methods still have some limitations. First, application-centric tuning makes it difficult to generalize to dynamic or mixed application scenarios. Second, they exhibit limited efficacy in high-dimensional parameter space exploration and are susceptible to local optima entrapment. To overcome these issues, this paper introduces BTune, a novel bottleneck-centric configuration parameters tuning method. By dynamically identifying bottlenecks and adjusting relevant parameters, it avoids the inefficiency of blindly tuning all parameters. In addition, inspired by mixture of expert, we present a modular neural architecture that decomposes tuning tasks into specialized subspaces. Each expert module focuses on a subset of parameters relevant to a specific resource bottleneck, while a gating mechanism dynamically routes tuning decisions based on observed system state. Experimental results show that, compared with OPPerTune, BTune not only demonstrates accelerated discovery of near-optimal parameter configurations in singleapplication scenarios, but also achieves much higher optimization efficacy in mixed-application scenarios. Qinwen Shi, Yuwan Li, Yubo Deng, Yuanchao Xu 0003 |
ICPADS | 1 |
| 2024 | Leaf: Learning-based Stream-level Fair Scheduling for Deep Learning AcceleratorsabstractDeep learning accelerators, as the cornerstone of machine learning systems, expedite neural network training and inference. With their computational power escalating annually, multi-tasking becomes imperative to harness their full potential. However, similar to other parallel processing systems, deep learning accelerators confront the problem of performance fluctuations which could result in unpredictable kernel latency, suboptimal resource utilization, and exacerbated tail latency. This paper identifies the unfairness in stream-level scheduling as the root cause of these performance fluctuations. To mitigate this issue, we introduce Leaf, an innovative learning-based stream-level fair scheduling method that dynamically learns and adapts scheduling policies by leveraging feature vectors extracted from enqueued kernels and accelerator status. To tackle the issue of scalability and latency constraints inherent in stream-level scheduling, we devise a scalable scheduling framework and an online scheduler switching mechanism for Leaf. Preliminary implementation on a commercial grade deep learning accelerator demonstrates that Leaf can significantly reduce kernel latency variation by $10 \sim 20$ times, sustain high and stable resource utilization, and markedly decrease workload runtime by mitigating tail latency, outperforming the accelerator’s native scheduler. Bojun Cao, Mengjuan Gao, Qinwen Shi, Yuanchao Xu 0002 |
ICPADS | 4 |
| 2024 | vLFS: Learning-based Scheduling for AI Chip VirtualizationabstractThe virtualization of AI chips ensures security in multi-user scenarios by partitioning the AI chip into multiple logically isolated instances. While each instance is independent in terms of computing resources, they share the same task scheduler, resulting in implicit competition for scheduler time slices which could in turn degrade overall performance. To address the problem, we propose vLFS, a deep reinforcement learning-based scheduling algorithm that opens up a new multidimensional optimization space for AI chip virtualization. vLFS schedules tasks from different instances by combining task load characteristics and the runtime status of the underlying AI chip, while also considering collaborative scheduling between the host and the device, thereby significantly improving AI chip utilization. We implement vLFS on a real-world system and conduct extensive comparisons against heuristic scheduling methods. Mengjuan Gao, Qinwen Shi, Bojun Cao, Yuanchao Xu 0002 |
ISPA | 3 |
| 2023 | An Empirical Study of Memory Pool Based Allocation and Reuse in CUDA Graph
Ruyi Qian, Mengjuan Gao, Qinwen Shi, Yuanchao Xu 0002 |
ICA3PP (5) | 3 |
| 2023 | EagerReuse: An Efficient Memory Reuse Approach for Complex Computational GraphabstractMemory reuse is a promising approach for deep neural network (DNN) to reduce memory consumption because it does not introduce any additional runtime overhead. We observe that existing memory reuse algorithms consider only the effect of an individual data feature (either tensor size or tensor lifetime) on memory reuse and ignore the relative position relationship (RPR) among tensors. As computational graphs grow slightly more complex, the mining of memory reuse becomes insufficient. To address this issue, we propose a new memory reuse algorithm—EagerReuse, which can exploit more memory reuse opportunities by analyzing RPR among tensors and reusing them as quickly as possible. We evaluated the algorithms with inference models in TensorFlow Model Garden, and the results show that the EagerReuse outperforms the state-of-the-art algorithms in three out of seven cases. For more complex computational graphs, EagerReuse can achieve better memory usage with slightly higher but acceptable overhead. Ruyi Qian, Bojun Cao, Mengjuan Gao, Qinwen Shi, Yuanchao Xu 0002, Qirun Huo, Keni Qiu |
ICPADS | 4 |