EDBT 2026 Demo / reviewers in the wild / expert
Yuanchao Xu 0002
dblp:78/10373-2
· DBLP profile ↗
14ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-4165-9138ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Parallel Serving for Vision-Language Models via Modal Decoupling and SchedulingabstractVision-Language Models (VLMs) have demonstrated strong performance in tasks such as image captioning and visual question answering. Under mixed workloads, however, the differing inference pipelines for text-only and multimodal requests create heterogeneity that existing serving systems fail to optimize—leading to high latency and poor fairness. We propose Duet-Infer, a modality-aware serving framework that enhances single-GPU serving efficiency for VLMs through three key contributions: (i) parallel computation enabled by preprocessing parallelism and decoupled vision-language execution, (ii) a shared memory manager that eliminates weight redundancy and supports efficient encoder cache sharing, and (iii) a fairness-aware scheduler that reduces delays for multimodal requests without penalizing text-only ones. Implemented within vLLM and evaluated on realistic workloads, DuetInfer reduces P99 TTFT by up to 33.7% and end-to-end latency by up to 20%. Yijia Yang, Yubo Deng, Yuanchao Xu 0002, Keni Qiu |
DATE | 4 |
| 2024 | Leaf: Learning-based Stream-level Fair Scheduling for Deep Learning AcceleratorsabstractDeep learning accelerators, as the cornerstone of machine learning systems, expedite neural network training and inference. With their computational power escalating annually, multi-tasking becomes imperative to harness their full potential. However, similar to other parallel processing systems, deep learning accelerators confront the problem of performance fluctuations which could result in unpredictable kernel latency, suboptimal resource utilization, and exacerbated tail latency. This paper identifies the unfairness in stream-level scheduling as the root cause of these performance fluctuations. To mitigate this issue, we introduce Leaf, an innovative learning-based stream-level fair scheduling method that dynamically learns and adapts scheduling policies by leveraging feature vectors extracted from enqueued kernels and accelerator status. To tackle the issue of scalability and latency constraints inherent in stream-level scheduling, we devise a scalable scheduling framework and an online scheduler switching mechanism for Leaf. Preliminary implementation on a commercial grade deep learning accelerator demonstrates that Leaf can significantly reduce kernel latency variation by $10 \sim 20$ times, sustain high and stable resource utilization, and markedly decrease workload runtime by mitigating tail latency, outperforming the accelerator’s native scheduler. Bojun Cao, Mengjuan Gao, Qinwen Shi, Yuanchao Xu 0002 |
ICPADS | 5 |
| 2024 | vLFS: Learning-based Scheduling for AI Chip VirtualizationabstractThe virtualization of AI chips ensures security in multi-user scenarios by partitioning the AI chip into multiple logically isolated instances. While each instance is independent in terms of computing resources, they share the same task scheduler, resulting in implicit competition for scheduler time slices which could in turn degrade overall performance. To address the problem, we propose vLFS, a deep reinforcement learning-based scheduling algorithm that opens up a new multidimensional optimization space for AI chip virtualization. vLFS schedules tasks from different instances by combining task load characteristics and the runtime status of the underlying AI chip, while also considering collaborative scheduling between the host and the device, thereby significantly improving AI chip utilization. We implement vLFS on a real-world system and conduct extensive comparisons against heuristic scheduling methods. Mengjuan Gao, Qinwen Shi, Bojun Cao, Yuanchao Xu 0002 |
ISPA | 5 |
| 2023 | An Empirical Study of Memory Pool Based Allocation and Reuse in CUDA Graph
Ruyi Qian, Mengjuan Gao, Qinwen Shi, Yuanchao Xu 0002 |
ICA3PP (5) | 4 |
| 2023 | EagerReuse: An Efficient Memory Reuse Approach for Complex Computational GraphabstractMemory reuse is a promising approach for deep neural network (DNN) to reduce memory consumption because it does not introduce any additional runtime overhead. We observe that existing memory reuse algorithms consider only the effect of an individual data feature (either tensor size or tensor lifetime) on memory reuse and ignore the relative position relationship (RPR) among tensors. As computational graphs grow slightly more complex, the mining of memory reuse becomes insufficient. To address this issue, we propose a new memory reuse algorithm—EagerReuse, which can exploit more memory reuse opportunities by analyzing RPR among tensors and reusing them as quickly as possible. We evaluated the algorithms with inference models in TensorFlow Model Garden, and the results show that the EagerReuse outperforms the state-of-the-art algorithms in three out of seven cases. For more complex computational graphs, EagerReuse can achieve better memory usage with slightly higher but acceptable overhead. Ruyi Qian, Bojun Cao, Mengjuan Gao, Qinwen Shi, Yuanchao Xu 0002, Qirun Huo, Keni Qiu |
ICPADS | 6 |
| 2020 | FCDM: A Methodology Based on Sensor Pattern Noise Fingerprinting for Fast Confidence Detection to Adversarial AttacksabstractDeep neural networks (DNNs) have shown phenomenal success in many real-world applications. However, a concerning weakness of DNNs is their vulnerability to adversarial attacks. Although there exist some methods to detect adversarial attacks, they often suffer from high computational cost and constraints on certain types of attacks, and ignore external features that could aid during attack detection. In this article, we propose fast confidence detection method (FCDM), an innovative method for fast confidence detection of adversarial attacks based on measuring the integrity of sensor pattern noise fingerprinting embedded in input examples. We note that the existing adversarial detectors are often designed as a binary classifier to differentiate clean or adversarial examples. However, the detection of adversarial examples can be much more complicated than such a scenario. Our key insight is that the confidence level of detecting an input sample as an adversarial example is a more useful info for the system to properly take an action to resist potential attacks. The experimental results show that FCDM is capable to give a confidence distribution model of the most popular adversarial attacks. And, using the confidence distribution model, FCDM can quickly determine the confidence level of the input sample. Based on different properties of the confidence distribution models associated with these adversarial attacks, FCDM can provide early attack warning including even the possible attack types of the adversarial attack examples. FCDM also has the following advantages: 1) it is effective for both a white-box attack and black-box attack; 2) it do not depend on the class of adversarial attacks and can be used as both known attack defense and unknown attack defense; and 3) it does not need to know the details of the DNN model and does not affect the functionality of the DNN. Since fast confidence detection method (FCDM) is a computationally heavy task, we propose an FPGA-based accelerator based on a series of optimization techniques, such as the quantization, data reuse and operation replacement, etc. We implement our method on an FPGA platform and achieve a system clock frequency of 279 MHz with a power consumption of the only 0.7626 W. Moreover, in the real system performance test, we obtain a high efficiency of 29.740 IPS/W and a low latency of just 44.1 ms with very marginal accuracy loss. Yazhu Lan, Kent W. Nixon, Qingli Guo, Guohe Zhang, Yuanchao Xu 0002, Hai Li 0001, Yiran Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | Fast Confidence Detection: One Hot Way to Detect Adversarial Attacks via Sensor Pattern Noise FingerprintingabstractDeep Neural Networks (DNNs) have shown phenomenal success in a wide range of real-world applications. However, a concerning weakness of DNNs is that they are vulnerable to adversarial attacks. Although there exist methods to detect adversarial attacks, they often suffer constraints on specific attack types and provide limited information to downstream systems. We specifically note that existing adversarial detectors are often binary classifiers, which differentiate clean or adversarial examples. However, detection of adversarial examples is much more complicated than such a scenario. Our key insight is that the confidence probability of detecting an input sample as an adversarial example will be more useful for the system to properly take action to resist potential attacks. In this work, we propose an innovative method for fast confidence detection of adversarial attacks based on integrity of sensor pattern noise embedded in input examples. Experimental results show that our proposed method is capable of providing a confidence distribution model of most of popular adversarial attacks. Furthermore, our presented method can provide early attack warning with even the attack types based on different properties of the confidence distribution models. Since fast confidence detection is a computationally heavy task, we propose an FPGA-Based hardware architecture based on a series of optimization techniques, such as incremental multi-level quantization and etc. We realize our proposed method on an FPGA platform and achieve a high efficiency of 29.740 IPS/W with a power consumption of only 0.7626W. Yazhu Lan, Qingli Guo, Guohe Zhang, Yuanchao Xu 0002, Kent W. Nixon, Hai Li 0001, Yiran Chen 0001 |
FPGA | 4 |
| 2018 | A peripheral circuit reuse structure integrated with a retimed data flow for low power RRAM crossbar-based CNNabstractConvolutional computations implemented in RRAM crossbar-based Computing System (RCS) demonstrate the outstanding advantages of high performance and low power. However, current designs are energy-unbalanced among the three parts of RRAM crossbar computation, peripheral circuits and memory accesses, and the latter two factors can significantly limit the potential gains of RCS. Addressing the problem of high power overhead of peripheral circuits in RCS, this paper proposes a Peripheral Circuit Unit (PeriCU)-Reuse scheme to meet power budgets in energy constrained embedded systems. The underlying idea is to put the expensive ADCs/DACs onto spotlight and arrange multiple convolution layers to be sequentially served by the same PeriCU. In the solution, the first step is to determine the number of PeriCUs which are organized by cycle frames. Inside a cycle frame, the layers are computed in parallel inter-PeriCUs while sequentially intra-PeriCU. Furthermore, a layer retiming technique is exploited to further improve the energy of RCS by assigning two adjacent layers within the same PeriCU so as to bypass the energy consuming memory accesses. The experiments of five convolutional applications validate that the PeriCU-Reuse scheme integrated with the retiming technique can efficiently meet variable power budgets, and further reduce energy consumption efficiently. Keni Qiu, Weiwen Chen, Yuanchao Xu 0002, Lixue Xia, Yu Wang 0002, Zili Shao |
DATE | 3 |
| 2018 | Efficient energy management by exploiting retention state for self-powered nonvolatile processors
Keni Qiu, Zhiyao Gong, Dongqin Zhou, Weiwen Chen, Yuanchao Xu 0002, Yongpan Liu |
J. Syst. Archit. | 5 |
| 2017 | i-BEP: A non-redundant and high-concurrency memory persistency modelabstractByte-addressable, non-volatile memory (NVM) technologies enable fast persistent updates but incur potential data inconsistency upon a failure. Recent proposals present several persistency models to guarantee data consistency. However, they fail to express the minimal persist ordering as a result of inducing unnecessary ordering constraints. In this paper, we propose i-BEP, a non-redundant high concurrency memory persistency model, which expresses epoch dependency via persist directed acyclic graph instead of program order. Additionally, we propose two techniques, background persist and deferred eviction, to enhance the performance of i-BEP. We demonstrate that i-BEP can improve the performance by 15% for typical data structures on average over buffered epoch persistency (BEP) model. Yuanchao Xu 0002, Zeyi Hou, Junfeng Yan, Hu Wan 0001 |
DATE | 1 |
| 2017 | Retention state-enabled and progress-driven energy management for self-powered nonvolatile processorsabstractEnergy harvesting instead of battery is a better power source for wearable devices due to many advantages such as long operation time without maintenance and comfort to users. However, harvested energy is naturally unstable and program execution will be interrupted frequently. To solve this problem, nonvolatile processor (NVP) has been proposed because it can back up volatile state before the system energy is depleted. However, this backup process also introduces non-negligible energy and area overhead. To improve the performance of NVP, retention state has been proposed recently which can enable a system to retain the volatile data to wait for power resumption instead of saving data immediately. The goal of this paper is to forward program execution as much as possible by exploiting retention state. Specifically, two objectives are achieved. The first objective is to minimize power failures of the system if there is a great probability to get power resumption during retention state. The second objective of this paper is to achieve maximum computation efficiency if it is unlikely to avoid power failure. Compared to the instant backup scheme, evaluation results report that power failure can be reduced by 81.6% and computation efficiency can be increased by 2.5x by the proposed retention state-aware energy management strategy. Zhiyao Gong, Keni Qiu, Dongqin Zhou, Weiwen Chen, Yuanchao Xu 0002, Yongpan Liu |
RTCSA | 5 |
| 2017 | Data re-allocation enabled cache locking for embedded systems
Chun Jason Xue, Keni Qiu, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Mengying Zhao |
J. Syst. Archit. | 5 |
| 2016 | Refresh-aware loop scheduling for high performance low power volatile STT-RAMabstractThe highlighted advantages of low leakage power, high storage density and immunity to electronic magnetic radiation make STT-RAM a promising candidate to build cache, SPM or main memory in embedded systems. However, write operations on STT-RAM have considerably longer latency and higher energy consumption than conventional SRAM. To solve this problem, researchers have proposed to relax STT-RAM's non-volatility and to have it work in a fast and low power mode. Under this volatile mode, refresh operations are needed to guarantee data correctness if their lifespan is larger than the retention time. It is observed that this refresh overhead is significant for data in stencil loops with the characteristic of constant read and write dependencies. This paper proposes a loop scheduling technique which can traverse loops in a new direction such that data lifespan can be greatly shortened. Therefore, overall refresh overhead can be efficiently mitigated so as to improve performance and reduce power consumption. The experimental results indicate that access latency and dynamic energy can be improved by 21.4~96.0% and 22.0~95.5% respectively by the proposed scheduling scheme. Keni Qiu, Junpeng Luo, Zhiyao Gong, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Tao Li 0006, Chun Jason Xue |
ICCD | 6 |
| 2016 | Reducing Synchronization Cost for Single-Level Store in Mobile Systems
Yuanchao Xu 0002, Hu Wan 0001, Keni Qiu, Tao Li 0006, Weigong Zhang |
J. Comput. Sci. Technol. | 1 |