EDBT 2026 Demo / reviewers in the wild / expert
Yunfei Song
dblp:192/5416
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning CompilationabstractWith the rapid development of deep learning models and hardware support for dense computing, the deep learning (DL) workload characteristics changed significantly from a few hot spots on compute-intensive operations to a broad range of operations scattered across the models. Accelerating a few compute-intensive operations using the expert-tuned implementation of primitives doesn't fully exploit the performance potential of AI hardware. Various efforts have been made to compile a full deep neural network (DNN) graph. One of the biggest challenges is to achieve high-performance tensor compilation by generating expert-level performance code for the dense compute-intensive operations and applying compilation optimization at the scope of DNN computation graph across multiple compute-intensive operations. We present oneDNN Graph Compiler, a tensor compiler that employs a hybrid approach of using techniques from both compiler optimization and expert-tuned kernels for high-performance code generation of the deep neural network graph. oneDNN Graph Compiler addresses unique optimization challenges in the deep learning domain, such as low-precision computation, aggressive fusion of graph operations, optimization for static tensor shapes and memory layout, constant weight optimization, and memory buffer reuse. Experimental results demonstrate significant performance gains over existing tensor compiler and primitives library for performance-critical DNN computation graphs and end-to-end models on Intel® Xeon® Scalable Processors. Zhennan Qin, Yijie Mei, Jingze Cui, Yunfei Song, Ciyong Chen, Longsheng Du, Xianhang Cheng, Baihui Jin, Jason Ye, Eric Lin, Dan Lavery |
CGO | 5 |
| 2024 | Distributed Rendering for Cloud Gaming in Cloud-Edge-End Cooperation NetworksabstractCurrently, the typical framework for cloud gaming involves the use of cloud servers or a combination of cloud servers and edge servers to provide services. At the same time, there is an increasing number of intelligent user devices with computing resources, and the computational capabilities of user devices are also becoming stronger. However, existing frameworks often treat user devices as terminals capable of network connectivity and display, which leads to a waste of computational resources. Therefore, this paper proposes a cloud-edge-end collaborative cloud gaming distributed rendering framework that incorporates user devices into the cloud gaming service network. The framework not only inherits the advantages of traditional frameworks, but also takes advantage of the computing power of user devices, enhances the usability of the framework, makes the utilization of computing resources higher, and makes the framework's work more flexible. Additionally, this paper designs a server operational cost optimization algorithm based on machine learning and metaheuristic algorithms. We compare and evaluate the proposed framework and algorithm with the mainstream framework and algorithm, and the experimental results prove its effectiveness Yipei He, Yongqiang Gao, Yunfei Song |
CSCWD | 3 |
| 2024 | Joint Task Offloading and Resource Allocation for NOMA-Based Vehicular NetworksabstractVehicle Edge Computing (VEC) is a critical technology that can achieve low latency and energy consumption for Telematics. However, with the high-speed mobility of new energy-electric vehicles and their cross-regional nature, performing high-quality service of vehicle tasks on the VEC model is still challenging. Considering the electric vehicle range problem, this paper plans to jointly optimize task offloading, task result forwarding and computational resource allocation (OOFR) within the maximum tolerable delay of vehicle tasks to minimize vehicle tasks' delay and energy consumption. The non-orthogonal multiple access (NOMA) technology is used with roadside units (RSUs) to achieve multiplexing of limited spectrum resources. This allows multiple vehicle users to perform task transmission simultaneously, thus reducing vehicle task transmission delay. In addition, we propose a cooperative game approach based on NOMA for task grouping to reduce the signal interference of vehicle task transmission. Finally, a deep reinforcement learning method is proposed for task offloading decision selection. A simulation platform is built to compare with MEC, COMO and MADDPG methods, combined with simulation results, showing that the superiority of our proposed scheme is verified. Yunfei Song, Yongqiang Gao, Yipei He |
CSCWD | 1 |
| 2023 | A framework for deep neural network multiuser authorization based on channel pruningabstractSummary Various deep neural network (DNN) model watermarks have been proposed by researchers to verify copyrights for deep neural networks DNN. However, most DNN watermarking methods cannot prevent attackers from stealing and using the model. Unlike many existing approaches, this paper uses a channel pruning algorithm to protect DNN models, which verifies DNN models copyrights but also prevents the illegal use of DNN models. In this work, the pruning threshold or pruning rate is used as the secret key of a DNN model. After the secret key is distributed to multiple users, they prune the DNN model with the secret key, and the pruned and fine‐tuned model is provided to the users. The users can verify ownership of the model according to the pruning accuracy and fine‐tuning accuracy. If the secret key is incorrect, the accuracy of the model after fine‐tuning will be very low, and users will be unable to use the reasoning function of the fine‐tuned model. Based on the CIFAR‐10 and CIFAR‐100 datasets, we conducted experiments on five popular DNN models. The experimental results show that we can authorize multiple users by pruning very few channels in the convolution layers of the DNN model. Linna Wang, Yunfei Song, Yujia Zhu, Daoxun Xia, Guoquan Han |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | Deep neural network watermarking based on a reversible image hiding network
Linna Wang, Yunfei Song, Daoxun Xia |
Pattern Anal. Appl. | 2 |
| 2021 | FDA$^3$: Federated Defense Against Adversarial Attacks for Cloud-Based IIoT ApplicationsabstractAlong with the proliferation of artificial intelligence and Internet of things (IoT) techniques, various kinds of adversarial attacks are increasingly emerging to fool deep neural networks (DNNs) used by industrial IoT (IIoT) applications. Due to biased training data or vulnerable underlying models, imperceptible modifications on inputs made by adversarial attacks may result in devastating consequences. Although existing methods are promising in defending such malicious attacks, most of them can only deal with limited existing attack types, which makes the deployment of large-scale IIoT devices a great challenge. To address this problem, in this article, we present an effective federated defense approach named FDA3that can aggregate defense knowledge against adversarial examples from different sources. Inspired by federated learning, our proposed cloud-based architecture enables the sharing of defense capabilities against different attacks among IIoT devices. Comprehensive experimental results show that the generated DNNs by our approach can not only resist more malicious attacks than existing attack-specific adversarial training methods, but also prevent IIoT applications from new attacks. Yunfei Song, Tian Liu 0005, Tongquan Wei, Xiangfeng Wang 0001, Zhe Tao, Mingsong Chen 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Fault-tolerant routing algorithm based on disjoint paths in 3-ary n-cube networks with structure faults
Weibei Fan, Zhijie Han 0001, Yunfei Song, Ruchuan Wang 0001 |
J. Supercomput. | 4 |