Siwei Ye

dblp:279/6847 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 RT-VirtIO: Towards the Real-Time Performance of VirtIO in a Two-Tier Computing Architecture
abstract
With the popularity of virtualization technology, ensuring reliable I/O operations with timing constraints in virtual environments becomes increasingly critical. Timing-predictable virtual I/O enhances the responsiveness and efficiency of virtualized systems, facilitating their seamless integration into time-critical applications such as industrial automation and robotics. Its significance lies in meeting rigorous performance standards, minimizing latency, and consistently delivering predictable I/O performance. As a result, virtual machines can effectively support mission-critical and time-sensitive workloads. However, due to the complicated system architecture, the I/O operations in virtualization face competition from tasks within the same virtual machine and those in different virtual machines who share the same host machine. This study presents RT-VirtIO, a practical approach to provide predictable real-time I/O operations. RT-VirtIO addresses the challenges associated with lengthy data paths and complex resource management. Through early-stage characterization, this study identifies key factors contributing to poor I/O real-time performance and then builds an analytical model and a learning-based data-driven model to predict the tail I/O latency. Leveraging these two models, RT-VirtIO effectively captures these dynamics, enabling the development of a general and applicable optimization framework. Experimental results demonstrate that RT-VirtIO significantly improves real-time performance in virtual environments (by 20.07% ~ 30.90%) without necessitating hardware modifications, which exhibit promising applicability across a broader range of scenarios.
Siwei Ye, Minqing Sun, Huifeng Zhu, Yier Jin, An Zou
DATE1
2025 HARD: Hardening Real-Time Scheduling and Analysis for Accelerator Enabled Computing
abstract
Despite the advancements in supporting artificial intelligence, accelerator-enabled computing architectures still struggle to meet strict timing constraints due to the complex interactions between CPU cores and accelerators. Although various scheduling and response-time analysis techniques have been developed, a significant gap remains between the conservative hard real-time schedulability (i.e., worst-case response times) and the average measured schedulability on real systems. This pessimism significantly limits the deployment of hard real-time tasks on accelerator-enabled computing platforms. To address this, we propose HARD, a real-time scheduling approach that integrates scheduling strategies, response time analysis, and practical scheduler designs for general accelerator-enabled computing platforms. Benefiting the subtask level segmented characteristics that are ignored by classic schedulers, the proposed HARD can significantly improve the theoretically guaranteed hard real-time schedulability. Extensive experiments on off-the-shelf Intel CPUs and NVIDIA GPUs show that HARD outperforms state-of-the-art scheduling and analysis approaches, delivering a 11.3% improvement in hard real-time schedulability and a remarkable 45.1 % reduction in pessimism.
Yinchen Ni, Tianrui Ma, Jintao Chen 0001, Chongye Yang, Siwei Ye, Yuankai Xu, Yier Jin, An Zou
RTAS5
2025 FALCON: FPGA Accelerated Real-Time Intelligent Controller for Autonomous Systems
abstract
The growing complexity and stringent real-time demands of autonomous systems, such as self-driving cars and drones, have driven the adoption of intelligent control methods based on deep neural networks (DNNs). While these methods offer improved control performance over traditional modelbased approaches, they also pose significant computational challenges, particularly for resource-constrained platforms. FieldProgrammable Gate Arrays (FPGAs) offer an attractive solution due to their energy efficiency and customizable architecture. In this work, we propose FALCON, an innovative approach for designing real-time intelligent controllers for autonomous systems using FPGA accelerators. Our approach begins with designing DNN-based intelligent controllers with varying levels of complexity and accuracy on the FPGA platform. Then, a performance function is proposed to capture the interplay among controller complexity, computational behavior, physical system characteristics, and overall control performance. Based on this performance function, we develop an algorithm-hardware codesign framework to determine the optimal control complexity, hardware configuration, and resource allocation. Finally, a case study on the co-design of intelligent controllers and FPGAbased overlay processors, together with a hardware-in-the-loop simulator, is conducted to demonstrate the advantages of the proposed methods. Compared to benchmarking controllers on other platforms, FALCON's optimized intelligent controller using FPGA accelerators shows competitive control performance with superior real-time capability and power efficiency. FALCON's optimization reduces the worst-case response time (WCRT) by up to 46.52%, improves the control performance by$1.93 \times$compared to the default setup. For performance per power efficiency, FALCON achieves a$3.67 \times$improvement compared to the DNN intelligent controller on TX2 and a remarkable$30.78 \times$improvement compared to traditional MPC on CPU.
Siwei Ye, Jintao Chen 0001, Yehan Ma, An Zou
RTSS1
2024 FastTuning: Enabling Fast and Efficient Hyper-Parameter Tuning With Partitioning and Parallelism of Search Space
abstract
Hyper-parameter tuning (HPT) for deep learning (DL) models is prohibitively expensive. Sequential model-based optimization (SMBO) emerges as the state-of-the-art (SOTA) approach to automatically optimize HPT performance due to its heuristic advantages. Unfortunately, focusing on algorithm optimization rather than a large-scale parallel HPT system, existing SMBO-based approaches still cannot effectively remove their strong sequential nature, posing two performance problems: (1)extremely low tuning speedand (2)sub-optimal model quality. In this paper, we propose FastTuning, a fast, scalable, and generic system aiming at parallelly accelerating SMBO-based HPT for large DL/ML models. The key is to partition the highly complex search space into multiple smaller sub-spaces, each of which is assigned to and optimized by a different tuning worker in parallel. However, determining the right level of resource allocation to strike a balance between quality and cost remains a challenge. To address this, we further propose NIMBLE, a dynamic scheduling strategy that is specially designed for FastTuning, including (1) Dynamic Elimination Algorithm, (2) Sub-space Re-division, and (3) Posterior Information Sharing. Finally, we incorporate 6 SOTAs (i.e., 3 tuning algorithms and 3 parallel tuning tools) into FastTuning. Experimental results, on ResNet18, VGG19, ResNet50, and ResNet152, show that FastTuning can consistently offer much faster tuning speed (up to$80\times$) with better accuracy (up to 4.7% improvement), thereby enabling the application of automatic HPT to real-life DL models.
Xiaqing Li, Qi Guo 0001, Guangyan Zhang, Siwei Ye, Guanhua He, Yiheng Yao, Rui Zhang 0040, Yifan Hao 0001, Zidong Du
IEEE Trans. Parallel Distributed Syst.4