VLDB 2026 Research / reviewers in the wild / expert
Chengsi Gao
dblp:260/8370
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2023
0000-0002-3479-9020ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Layer-Puzzle: Allocating and Scheduling Multi-task on Multi-core NPUs by Using Layer HeterogeneityabstractIn this work, we propose Layer-Puzzle, a multi-task allocation and scheduling framework for multi-core NPUs. Based on the proposed latency-prediction model and dynamic parallelization scheme, Layer-Puzzle can generate near-optimal results for each layer under given hardware resources and traffic congestion levels. As an online scheduler, Layer-Puzzle performs a QoS-aware and dynamic scheduling method that picks the superior version from the previously compiled results and co-runs the selected tasks to improve system performance. Our experiments on MLPerf show that Layer-Puzzle can achieve up to 1.61X, 1.53X, and 1.95X improvement in ANTT, STP, and PE utilization, respectively. Chengsi Gao, Ying Wang 0001, Cheng Liu 0008, Mengdi Wang 0004, Yinhe Han 0001, Lei Zhang 0008 |
DATE | 1 |
| 2023 | IVP: An Intelligent Video Processing Architecture for Video StreamingabstractRecently, video processing tasks, such as video enhancement and analysis, have received increasing attention from both academics and industries. However, the current video processing procedure on edge decouples the decoding phase and the subsequent video processing tasks, missing the opportunity to accelerate the procedure by orchestrating video decoding and enhancement stages. Thus, we propose an intelligent video processing workflow and architecture(IVP) for cloud-edge video streaming. For edge devices that receive compressed videos, IVP can perform direct DNN-based video enhancement, e.g., super-resolution and frame-interpolation. By leveraging the metadata motion vectors and residuals extracted from the encoded video, our architecture will significantly eliminate unnecessary frame pixels being processed by the DNNs and improve execution efficiency. The proposed IVP and workflow are proved to reduce up to 90% of the processing latency while producing accurate and high-quality videos. Furthermore, we observe a significant portion of similar optical flow in time domain of continuous videos, which can be used to reduce the computation overhead. Thus, to utilize such temporal similarity of optical flow, the proposed IVP is upgraded to be capable of reusing previous computation results, which further improves energy efficiency of the whole system. Chengsi Gao, Ying Wang 0001, Yinhe Han 0001, Lei Zhang 0008 |
IEEE Trans. Computers | 1 |
| 2023 | A Framework for Neural Network Architecture and Compile Co-optimizationabstractThe efficiency of deep neural network (DNN) solutions on real hardware devices are mainly decided by the DNN architecture and the compiler-level scheduling strategy on the hardware. When we try to fully exploit the underlying hardware and obtain the optimal tradeoff between DNN accuracy and runtime performance, we discovered that the two optimization goals of DNN architecture and scheduling policy are intimately related to each other. However, current hardware-aware Neural Architecture Search (NAS) methods primarily focus on the DNN architecture search process, ignoring the effects of various compiler-level scheduling strategies (e.g., graph-level optimization, loop transformations, parallelization, etc.) on network candidates being evaluated in the search process. As a result, they may overlook the true-optimal DNN implementations on hardware, which can only be discovered by trying-out different combinations of scheduling strategies and DNN architectures. This work proposes a NAS framework (CHaNAS) that searches for not only the network architecture but also the dedicated compiler-level scheduling policy, as the optimal co-design solution on the target hardware. We propose to use a block-based pre-scheduling methodology to reduce the co-design search space and enable the automatic generation of the optimal co-design, including the network architecture and the tensor programs that practice the scheduling policy. Further, we introduce a new search objective function based on the generalization gap to prevent the selection of architectures that are prone to overfitting. We evaluate CHaNAS on Imagenet on different hardware back-ends against the state-of-the-art hardware-aware search method based on the MobileNet-v3 search space. Experimental results show that the co-design solutions obtained by ChaNAS show up to 1.6×, 1.9×, and 1.7×, 24 performance boost on NVIDIA P100 GPU, Intel Xeon 8163 CPU, and Samsung Note 10 Mobile, respectively, over the baselines of the same-level accuracy. Ying Wang 0001, Chengsi Gao, Cheng Liu 0008, Lei Zhang 0008 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | Olympus: Reaching Memory-Optimality on DNN ProcessorsabstractIn DNN processors, main memory consumes much more energy than arithmetic operations. Therefore, many memory-oriented network scheduling (MONS) techniques are introduced to exploit on-chip data reuse opportunities and reduce accesses to memory. However, to derive the theoretical lower bound of memory overhead for DNNs is still a significant challenge, which also sheds light on how to reach memory-level optimality by means of network scheduling. Prior work on MONS mainly focused on disparate optimization techniques or missed some of the data reusing opportunities in diverse network models, thus their results are likely to deviate from the true memory-optimality that can be achieved in processors. This paper introduces Olympus, which comprehensively considers the entire memory-level DNN scheduling space, formally analyzes the true memory-optimality and also how to reach the memory-optimal schedules for an arbitrary DNN running on a DNN processor. The key idea behind Olympus is to derive a true memory lower-bound regarding both the intra-layer and inter-layer reuse opportunities, which has not been simultaneously explored by prior works. Evaluation on SOTA DNN processors of different architectures shows that Olympus can guarantee the minimum off-chip memory access, and it reduces 12.3-85.6% DRAM access and saves 7.4-70.3% energy on the latest network models. Xuyi Cai, Ying Wang 0001, Kaijie Tu, Chengsi Gao, Lei Zhang 0008 |
IEEE Trans. Computers | 4 |
| 2022 | Amphis: Managing Reconfigurable Processor Architectures With Generative Adversarial LearningabstractDynamic resources management in reconfigurable processors often manifests as a hard online decision-making task, which should yield premier solutions that must meet Quality-of-Service (QoS) requirements while maximizing the system’s efficiency. Most prior works rely on a hard-to-train predictor to model the complicated relationships between processor configurations and performance. To decide the proper resource allocation, the predictor needs to tentatively evaluate a group of possible configurations, and then decide the best configuration for the workload. This tedious process has an expensive runtime overhead for resource configuration in processors. Besides, prior works focus on improving the prediction accuracy, however, higher performance prediction cannot guarantee a good system outcome. Inspired by recent advances in adversarial learning, we present a generative adversarial network (GAN)-based framework, Amphis, which can directly generate the on-demand processor configuration for any scheduled-in application. By evaluating Amphis on a reconfigurable processor with 18 different workloads, our results demonstrate that the GAN-based method provides tremendous overhead reduction (up to 90%) compared to the SOTA prediction-based method WNNM while providing higher resource utilization. Ying Wang 0001, Chengsi Gao, Yinhe Han 0001, Lei Zhang 0008 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | An Intelligent Video Processing Architecture for Edge-cloud Video StreamingabstractThis work proposes an intelligent video processing architecture for bandwidth-efficient edge-cloud video streaming. On receiving the bandwidth-saving low-quality video streaming in compressed format, the proposed architecture can perform direct DNN-based video enhancement, e.g., super-resolution and motion-compensated frame interpolation (MCFI), on streams. By utilizing the metadata motion vectors and residuals extracted from the encoded video, our workflow will significantly eliminate the unnecessary pixels being processed by the video-enhancing DNNs, and greatly promote the execution efficiency. The evaluation results on popular datasets show that our architecture can reduce the edge-side processing latency of video-enhancing DNNs by 90% compared to the traditional flow while producing accurate and high-quality videos on edge. Chengsi Gao, Ying Wang 0001, Lei Zhang 0008 |
DAC | 1 |
| 2021 | Tenet: A Neural Network Model Extraction Attack in Multi-core ArchitectureabstractAs neural networks (NNs) are being widely deployed in many cloud-oriented systems for safety-critical tasks, the privacy and security of NNs become significant concerns to users in the cloud platform that shares the computation infrastructure such as memory resource. In this work, we observed that the memory timing channel in the shared memory of cloud multi-core architecture poses the risk of network model information leakage. Based on the observation, we propose a learning-based method to steal the model architecture of the NNs by exploiting the memory timing channel without any high-level privilege or physical access. We first trained an end-to-end measurement network offline to learn the relation between memory timing information and NNs model architecture. Then, we performed an online attack and reconstructed the target model using the prediction from the measurement network. We evaluated the proposed attack method on a multi-core architecture simulator. The experimental results show that our learning-based attack method can reconstruct the target model with high accuracy and improve the adversarial attack success rate by 42.4%. Chengsi Gao, Bing Li 0017, Ying Wang 0001, Lei Zhang 0008 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2021 | CHaNAS: coordinated search for network architecture and scheduling policyabstractAutomatically design an efficient DNN solution for a given deep learning task on the target hardware mainly decided by the neural network architecture and the schedule mapping strategy, where the two goals are closely coupled with each other to fully exploit the advantages of the underlying hardware. Prior hardware-aware Neural Architecture Search (NAS) methods mostly ignore the impacts of different scheduling policies (e.g., graph-level optimization, loop transformations, parallelization, etc.) on network candidates being evaluated in the search process. Thus, they may miss the true-optimal architecture that can only be discovered by trying-out different scheduling policies. This work proposes a NAS framework (CHaNAS) that searches for not only the network architecture but also the dedicated scheduling policy, as the optimal co-design solution on target hardware that fully exploits the advantages of the underlying hardware. We propose to use a block-based pre-scheduling methodology to reduce the co-design search space, and enable the automatic generation of the optimal co-design, including the network architecture and the tensor programs that practice the scheduling policy. We evaluate CHaNAS on Imagenet on different hardware back-ends against the state-of-the-art hardware-aware search method MobileNet-v3. Experimental results show that the co-design solutions obtained by ChaNAS show up to 1.6x, 1.9x, and 1.7x performance boost on NVIDIA P100 GPU, Intel Xeon 8163 CPU, and Samsung Note 10 Mobile, respectively, over the baselines of the same-level accuracy. Ying Wang 0001, Gangliang Lin, Chengsi Gao, Cheng Liu 0008, Lei Zhang 0008 |
LCTES | 4 |