Sichao Chen

dblp:228/3937 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Performance Prediction of Concurrent DNN Training Tasks in GPU Spatial Sharing Environments
abstract
GPU sharing is commonly employed in GPU clusters to improve utilization, with spatial sharing being one of the most widely adopted techniques. However, spatial sharing can lead to resource interference, making task execution times difficult to predict. Predictable execution times for each task are crucial in GPU cluster management and task scheduling. In this article, we propose a performance predictor for multi-DNN training tasks in GPU spatial sharing environments. We first conduct experiments on spatial sharing for multiple DNN workloads on a single GPU, demonstrating that concurrent execution of multiple tasks improves overall performance and GPU resource utilization compared to serial execution. By analyzing warp stall reasons collected during task execution, we investigate the interference for computation and memory resources under MPS on GPUs. Finally, we design a performance predictor that predicts the execution time of a target DNN training task when it runs concurrently with other tasks under GPU spatial sharing via MPS. The predictor is capable of predicting the execution time of each task for previously unseen combinations of DNN training tasks. Extensive evaluations on modern GPUs show that compared to other baseline methods, our approach exhibits higher prediction accuracy, as well as improved stability and robustness. Experiments on multiple GPU architectures, as well as at higher concurrency levels, further demonstrate that our method possesses strong generalization and scalability. We also conducted a performance analysis under diverse workload pattern and a case study to validate the practical applicability of our predictor in real scheduling environments.
Sichao Chen, Desheng Wang 0002, Weizhe Zhang, Meng Hao 0002, Yu-Chu Tian
ACM Trans. Archit. Code Optim.1
2025 ServerlessLego: An Elastic Serverless Framework Assembling Model Building Blocks to Provide SLO-Aware Inference Services
abstract
Inference of large language models (LLMs) is common in cloud environments. As the elastic resource management capabilities and the flexible pay-as-you-go billing model offered by serverless, LLM inference services are increasingly migrated to serverless platforms. However, the increasing size of LLMs in recent years has introduced a new cold start issue for serverless frameworks, which in turn impacts their scalability under dynamic workloads. To address these issues, we propose ServerlessLego, an elastic serverless computing framework. ServerlessLego partitions LLMs into layers, then groups and deploys them to different instances, and loads these groups in parallel. These instances perform a subscription-based pipeline. To address dynamically request loads, ServerlessLego models the incoming request patterns and the inference time of running requests, providing an SLO-Aware instance scheduling. Experiments show that ServerlessLego reduces the cold start time of serverless frameworks by 58.15 % and improves throughput by 43.39 % compared to the baseline for dynamic workloads. Moreover, ServerlessLego can horizontally schedule instance based on request SLOs and arrival rates.
Desheng Wang 0002, Weizhe Zhang, Sichao Chen, Yuming Feng 0002
ICPADS4
2025 HyDLR: Load-Aware Dynamic Rescheduling for Deep Learning Hybrid Deployment
abstract
Resource contention, driven by traffic surges from online services, presents a significant challenge in hybrid clusters where latency-sensitive and best-effort deep learning tasks are colocated. To address this, we propose HyDLR, a dynamic, loadaware hybrid deployment scheduling method that dynamically reallocates offline tasks to ensure Quality of Service (QoS) for online services while enhancing overall resource utilization. The bursty nature and stringent QoS demands of online tasks, coupled with the fluctuating resource footprints of offline tasks, can lead to severe resource pressure on nodes and undermine system stability. HyDLR first designs a load-aware rescheduling policy that dynamically identifies resource hotspots by monitoring metrics such as CPU satisfaction degree, memory, and GPU memory utilization. It then leverages eviction and task migration to optimize workload distribution. Furthermore, a two-stage filtering algorithm, guided by a multi-objective optimization model, targets system-wide load balancing and minimal rescheduling overhead. By incorporating a dynamically adjusted priority queue and a cost-feedback mechanism, HyDLR improves scheduling efficiency without compromising stability. Experimental results demonstrate that HyDLR significantly reduces the frequency of task migrations while achieving a well-balanced system load. The rate of cascading rescheduling events is kept below 3%, demonstrating superior performance over existing approaches. This work offers an effective solution for resource management in complex, hybrid deployment scenarios, laying a foundation for more efficient data center scheduling and demonstrating strong potential for practical adoption.
Desheng Wang 0002, Shuo Si, Sichao Chen, Weizhe Zhang
ICPADS4
2025 DynGPU: A Dynamic GPU Sharing Framework for Enhanced Resource Utilization and Task Scheduling in Concurrent DNN Training
abstract
Training deep neural networks (DNNs) is a common task in GPU clusters. However, in practical cluster environments, multiple concurrent DNN training tasks often fail to fully leverage GPU resources, resulting in suboptimal GPU utilization. Furthermore, existing GPU sharing frameworks primarily rely on static scheduling and frequently overlook task deadlines, leading to task delays and inefficient scheduling. To address these issues, we propose a dynamic GPU sharing framework (DynGPU) that intercepts GPU kernel executions to perform resource scheduling in multi-task environments. DynGPU incorporates a dynamic task priority adjustment mechanism that adapts task priorities in real time based on task progress, historical data, and remaining time to deadlines. By guaranteeing resources for high-priority tasks while maximizing resource allocation for low-priority tasks, DynGPU reduces resource contention and improves system throughput, enabling more timely task completions. Experiments show that, compared to dedicated GPU execution, DynGPU can reserve up to 97.5 % of throughput for high-priority tasks. Compared to state-of-the-art baselines, DynGPU achieves up to an 8.4 % improvement in task completion time.
Zhiji Yu, Desheng Wang 0002, Weizhe Zhang, Sichao Chen, Meng Hao 0002, Yu-Chu Tian
ICPADS4
2025 COFFA: A Co-Design Framework for Fused-Grained Reconfigurable Architecture Towards Efficient Irregular Loop Handling
abstract
Coarse-Grained Reconfigurable Architecture (CGRA) emerges as a competitive accelerator due to its high flexibility and energy efficiency. However, most CGRAs are effective for computation-intensive applications with regular loops but struggle with irregular loops containing control flows. These loops introduce fine-grained logic operations and are costly to execute by coarse-grained arithmetic units in CGRA. Efficiently handling such logic operations necessitates incorporating Boolean algebra optimization, which can improve logic density and reduce logic depth. Unfortunately, no previous research has incorporated it into the compilation flow to support irregular loops efficiently.We proposeCOFFA, an open-source framework for heterogeneous architecture with a RISC-V CPU and a fused-grained reconfigurable accelerator, which integrates coarse-grained arithmetic and fine-grained logic units, along with flexible IO units and distributed interconnects. As a software/hardware co-design framework,COFFAhas a powerful compiler that extracts and optimizes fine-grained logic operations from irregular loops, performs coarse-grained arithmetic and memory optimizations, and offloads the loops to the accelerator.Across various challenging benchmarks with irregular loops,COFFAachieves significant performance and energy efficiency improvements over an in-order, an out-of-order RISC-V CPUs, and a recent FPGA, respectively. Moreover, compared with the state-of-the-art CGRAUE-CGRAandHycube,COFFAcan achieve 2.5× and 3.5× performance gains, respectively.
Yuan Dai, Xuchen Gao, Yunhui Qiu, Jingyuan Li 0003, Yuhang Cao, Yiqing Mao, Sichao Chen, Wenbo Yin, Wai-Shing Luk, Lingli Wang
IEEE Trans. Computers7
2024 An extended trust and distrust network-based dual fuzzy recommendation model and its application based on user-generated content
Sichao Chen, Shengjia Zhou
Expert Syst. Appl.1
2024 HierCGRA: A Novel Framework for Large-scale CGRA with Hierarchical Modeling and Automated Design Space Exploration
abstract
Coarse-grained reconfigurable arrays (CGRAs) are promising design choices in computation-intensive domains, since they can strike a balance between energy efficiency and flexibility. A typical CGRA comprises processing elements (PEs) that can execute operations in applications and interconnections between them. Nevertheless, most CGRAs suffer from the ineffectiveness of supporting flexible architecture design and solving large-scale mapping problems. To address these challenges, we introduce HierCGRA, a novel framework that integrates hierarchical CGRA modeling, Chisel-based Verilog generation, LLVM-based data flow graph (DFG) generation, DFG mapping, and design space exploration (DSE). With the graph homomorphism (GH) mapping algorithm, HierCGRA achieves a faster mapping speed and higher PE utilization rate compared with the existing state-of-the-art CGRA frameworks. The proposed hierarchical mapping strategy achieves 41× speedup on average compared with the ILP mapping algorithm in CGRA-ME. Furthermore, the automated DSE based on Bayesian optimization achieves a significant performance improvement by the heterogeneity of PEs and interconnections. With these features, HierCGRA enables the agile development for large-scale CGRA and accelerates the process of finding a better CGRA architecture.
Sichao Chen, Su Zheng, Guowei Zhu, Jingyuan Li 0003, Yazhou Yan, Yuan Dai, Wenbo Yin, Lingli Wang
ACM Trans. Reconfigurable Technol. Syst.1
2024 FDRA: A Framework for a Dynamically Reconfigurable Accelerator Supporting Multi-Level Parallelism
abstract
Coarse-grained reconfigurable architectures (CGRAs) have emerged as promising accelerators due to their high flexibility and energy efficiency. However, existing open source works often lack integration of CGRAs with CPU systems and corresponding toolchains. Moreover, there is rare support for the accelerator instruction pipelining to overlap data communication, computation, and configuration across multiple tasks. In this article, we propose FDRA, an open source exploration framework for a heterogeneous system-on-chip (SoC) with a RISC-V processor and a dynamically reconfigurable accelerator (DRA) supporting loop, instruction, and task levels of parallelism. FDRA encompasses parameterized SoC modeling, Verilog generation, source-to-source application code transformation using frontend and DRA compilers, SoC simulation, and FPGA prototyping. FDRA incorporates the extraction of periodic accumulative operators and multi-dimensional linear load/store operators from nested loops. The DRA enables accessing the shared L2 cache with virtual addresses and supports direct memory access with arbitrary start addresses and data lengths. Integrated into the RISC-V Rocket SoC, our DRA achieves a remarkable 55× acceleration for loop kernels and improves energy efficiency by 29×. Compared to state-of-the-art RISC-V vector units, our DRA demonstrates a 2.9× speed improvement and 3.5× greater energy efficiency. In contrast to previous CGRA+RISC-V SoCs, our SoC achieves a minimum speedup of 5.2×.
Yunhui Qiu, Yiqing Mao, Xuchen Gao, Sichao Chen, Wenbo Yin, Lingli Wang
ACM Trans. Reconfigurable Technol. Syst.4
2021 MSDF: A General Open-Domain Multi-skill Dialog Framework
Yu Zhao 0043, Xinshuo Hu, Yunxin Li, Baotian Hu, Dongfang Li 0002, Sichao Chen, Xiaolong Wang 0001
NLPCC (2)6