Penghao Sun

dblp:189/2278 · DBLP profile ↗
← Back
31ranked-venue papers
15as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 8 first-author · 6 since 2021Systems, architecture and hardware · 8 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 One sketch for all super hosts: Large-scale network monitoring with UnivCar
Jiqiang Xia, Jianjin Zhao, Penghao Sun, Jianhua Peng
Future Gener. Comput. Syst.5
2026 Kirin: Efficient In-Storage Learned Compaction for LSM-Trees via System-Algorithm Co-Design
Shengan Zheng, Penghao Sun, Jin Pu, Kaijiang Deng, Bowen Zhang 0012, Weihan Kong, Yifan Hua, Linpeng Huang
Proc. VLDB Endow.3
2026 Shiro: Efficient and Accurate In-Storage Data Lifetime Separation for nand Flash SSDs
abstract
The log-structured nature of NAND flash storage necessitates garbage collection in SSDs. Garbage collection (GC) is a major source of runtime write amplification (WA), leading to faster device wear out and interference with host I/Os. The key to mitigating this problem is separating data by lifetime so that data in the same flash block are invalidated within temporal proximity. For higher lifetime prediction accuracy and adaptibility, prior works proposed using machine learning algorithms for data separation. However, existing learning-based solutions perform data lifetime prediction at the host side, leading to several drawbacks. First, host-side prediction does not have knowledge of the internal data movement inside the SSD during GC, and thus fails to leverage the opportunity to further separate GC writes, resulting in suboptimal WA reduction in the long term. Second, performing prediction at the host significantly prolongs the I/O critical path and consumes host resources that could otherwise be used for serving user applications. We present Shiro, a holistic FTL design that performs instorage data separation for both user writes and GC writes for maximal long-term WA reduction. For user writes, Shiro uses a sequence model to accurately predict data lifetime by learning lifetime distribution from long historical access patterns. For GC writes, Shiro incorporates a reinforcement learning-assisted page migration strategy that takes direct feedback from longterm WA to further improve data separation efficacy. To address the challenges posed by performing fine-grained and real-time machine learning decisions inside the resource-constrained SSD, we propose a suite of enabling techniques to keep computation and storage overhead low. Extensive evaluation of Shiro on real-world traces shows that Shiro can deliver 29 WA compared with conventional FTL and state-of-the-art instorage data separation schemes. Furthermore, thanks to lower data migration overhead during GC, Shiro achieves significantly higher steady-state I/O performance.
Penghao Sun, Shengan Zheng, Litong You, Wanru Zhang, Ruoyan Ma, Feng Zhu 0024, Linpeng Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 TxISC: Transactional File Processing in Computational SSDs
abstract
Computational SSDs implement the in-storage computing (ISC) paradigm and benefit applications by taking over I/O-intensive tasks from the host. Existing works have proposed various frameworks aiming at easy access to ISC functionalities, and among them generic frameworks with file-based abstractions offer better usability. However, since intermediate output by ISC tasks may leave files in a dirty state, concurrent access to and the integrity of file data should be properly managed, which has not been fully addressed. In this paper, we present TxISC, a generic ISC framework that coordinates the host kernel and device firmware to offer a versatile file-based programming model. Under the hood, TxISC turns each invocation of an ISC task into a transaction with full ACID guarantee, fully covering concurrency control and data protection. TxISC implements transactions at low cost by leveraging the out-of-place write characteristic of NAND flash. Evaluation on full-stack hardware shows that transactions incur almost no runtime performance penalty compared with existing ISC architectures. Application case studies demonstrate that the programming model of TxISC can be used to offload complex logic and deliver significant speedup over host-only solutions.
Penghao Sun, Shengan Zheng, Kaijiang Deng, Jin Pu, Maojun Yuan, Feng Zhu 0024, Linpeng Huang
DATE1
2025 CSGC: Collaborative File System Garbage Collection with Computational Storage
Jin Pu, Shengan Zheng, Penghao Sun, Linpeng Huang
Euro-Par (2)3
2025 CoDDoS: Detecting and mitigating diverse DDoS attacks with programmable switches
Jiqiang Xia, Le Tian 0002, Yuxiang Hu 0004, Ziyong Li, Penghao Sun, Jianhua Peng
Comput. Commun.5
2024 MEFusion: Unsupervised Mutual Enhancement for Multimodal Image Fusion
abstract
Image fusion aims to extract valuable information from each modality to create a fused image. Currently, state-of-the-art image fusion approaches tend to initially decompose each modality into distinct yet complementary features, and transfer beneficial information through carefully hand-crafted or learned fusion rules to the target. Nevertheless, previous approaches treat each modality in isolation before fusion, potentially under-utilising the complementary information available across modalities. In this word, we introduce a novel method called MEFusion that pioneers cross-modality mutual enhancement before feature decomposition. By harnessing the individual strengths of each modality, MEFusion elevates the overall quality and comprehensiveness of the fusion outcome. To facilitate a bidirectional enhancement for each feature across modalities, we have designed a pluggable co-attention mechanism that seamlessly integrates into a lightweight dual-path transformer. Furthermore, to enrich the details of each modality, we propose an unsupervised cross-modality mutual enhancement loss, which overcomes the limitations of requiring paired training data for enhancement tasks. Extensive experiments conducted on several benchmark datasets demonstrate the superiority of our proposed MEFusion method in terms of traditional fusion metrics and perceptual quality improvement of fused images.
Yushe Cao, Siwen Jiao, Penghao Sun, Baoyun Peng, Dian-xi Shi, Yuanchun Shi
ECAI3
2024 Enabling efficient routing for traffic engineering in SDN with Deep Reinforcement Learning
Xinglong Pei, Penghao Sun, Yuxiang Hu 0004, Dan Li 0007, Le Tian 0002
Comput. Networks2
2024 Multi-resource interleaving for task scheduling in cloud-edge system by deep reinforcement learning
Xinglong Pei, Penghao Sun, Yuxiang Hu 0004, Dan Li 0007, Le Tian 0002, Ziyong Li
Future Gener. Comput. Syst.2
2024 SuperGuardian: Superspreader removal for cardinality estimation in data streaming
Jie Lu 0006, Hongchang Chen, Penghao Sun, Tao Hu 0002, Zhen Zhang 0049, Quan Ren
Inf. Syst.3
2024 Realizing the Carbon-Aware Service Provision in ICT System
abstract
The ever-growing carbon emission of information infrastructure accounts for a significant proportion of the global carbon emissions. Existing studies reduce carbon consumption mainly by improving power efficiency on specific facilities or energy source structures. However, these methods do not jointly consider the impact of computation and network resource distribution on carbon emission. In this paper, we propose a data-driven scheme named EcoNet using reinforcement learning to reduce carbon emissions by jointly scheduling computation and network resources. We dynamically monitor the status of the computation and network facilities using cloud-edge collaboration and software-defined networking. Based on the collected status information, we formulate the resource scheduling problem as an optimization problem, which comprehensively considers the carbon emission, electricity price, and quality of service. The problem has high computation complexity, and we solve the problem with the proposed EcoNet to achieve efficient scheduling and near-optimal performance based on the collected network status information. The evaluation results show that EcoNet can maintain good Quality of Service and save at least 17% of the overall cost considering the electricity bills and carbon emissions.
Penghao Sun, Julong Lan, Yuxiang Hu 0001, Zehua Guo 0001, Jiangxing Wu 0001
IEEE Trans. Netw. Serv. Manag.1
2023 Learning-based Data Separation for Write Amplification Reduction in Solid State Drives
abstract
Garbage collection in SSDs causes write amplification. The key to mitigating this problem is separating data by lifetime. Prior works proposed using machine learning to accurately predict data lifetime but prediction is performed at the host side, burdening the host storage stack. We present PHFTL, a practical, holistic FTL design with device-side learning-based data separation. The machine learning model in PHFTL accurately and adaptively predicts the lifetime of every written page. A suite of enabling techniques are introduced to keep computation and storage overhead low. Extensive evaluation of PHFTL demonstrates superiority over state-of-the-art and feasibility on real hardware.
Penghao Sun, Litong You, Shengan Zheng, Wanru Zhang, Ruoyan Ma, Guanzhong Wang, Feng Zhu 0024, Linpeng Huang
DAC1
2023 OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory
abstract
In this paper, we present OpenEmbedding, a distributed parameter server system for deep learning recommendation models (DLRM) workloads. In order to support rapid growth in the number of features and the model size (Terabytes are common) of DLRM workloads, OpenEmbedding takes advantage of emerging persistent memory (PMem) to address scalability and reliability issues in training DLRMs. Compared to DRAM, PMem can have much lower per-GB cost, higher density, and non-volatility, while with slightly low access performance to DRAM. OpenEmbedding uses DRAM as cache and PMem as storage for the sparse features and develops a simple but effective pipeline processing approach to optimize the access latency of the sparse features in PMem. For reliability, we develop a lightweight synchronous checkpointing scheme that is specially co-designed with the pipelined cache to reduce the run-time overhead of checkpointing. Our evaluations on a real-world industry workload consisting of billions of parameters demonstrate 1) the effectiveness of our PMem-aware optimizations, 2) checkpointing mechanism with near-zero run-time overhead to the training performance and 3) fast recovery with up to 3.97× speedup compared to the state-of-the-art. OpenEmbedding has been deployed in hundreds of scenarios in industry within 4Paradigm, and is open-sourced1.
Cheng Chen 0008, Jun Yang 0022, Mian Lu, Zhao Zheng, Bingsheng He, Weng-Fai Wong, Liang You, Penghao Sun, Yuping Zhao, Fenghua Hu, Andy Rudoff
ICDE10
2023 Virtual self-adaptive bitmap for online cardinality estimation
Jie Lu 0006, Hongchang Chen, Tao Hu 0002, Penghao Sun, Zhen Zhang 0049
Inf. Syst.5
2022 Defensive deception framework against reconnaissance attacks in the cloud with deep reinforcement learning
Huanruo Li, Shumin Huo, Penghao Sun
Sci. China Inf. Sci.5
2022 An optimal defensive deception framework for the container-based cloud with deep reinforcement learning
abstract
Abstract Defensive deception is emerging to reveal stealthy attackers by presenting intentionally falsified information. To implement it in the increasing dynamic and complex cloud, major concerns remain about the establishment of precise adversarial model and the adaptive decoy placement strategy. However, existing studies do not fulfil both issues because of (1) the insufficiency on extracting potential threats in virtualisation technique, (2) the inadequate learning on the agility of target environment, and (3) the lack of measurement for placement strategy. In this study, an optimal defensive deception framework is proposed for the container based‐cloud. The System Risk Graph (SRG) is formalised to depict an updatable adversarial model with the automatic orchestration platform. Afterwards, a Deep Reinforcement Learning (DRL) model is trained based on SRG. The well‐trained DRL agent generates optimal placement strategies for the orchestration platform to distribute decoys and deceptive routings. Lastly, the coefficient of deception, , is defined to evaluate the effectiveness of placement strategy. Simulation results show that the proposed method increases by 30.22%, and increase the detection ratio on the random walker attacker and persistent attacker by 30.69% and 51.10%, respectively.
Huanruo Li, Penghao Sun, Shumin Huo
IET Inf. Secur.3
2022 Enabling Scalable Routing in Software-Defined Networks With Deep Reinforcement Learning on Critical Nodes
abstract
Traditional routing schemes usually use fixed models for routing policies and thus are not good at handling complicated and dynamic traffic, leading to performance degradation (e.g., poor quality of service). Emerging Deep Reinforcement Learning (DRL) coupled with Software-Defined Networking (SDN) provides new opportunities to improve network performance with automatic traffic analysis and policy generation. However, existing DRL-based routing solutions usually rely on all node information to make routing decisions for the network and hence are both hard to converge in large networks and vulnerable to topology changes. In this paper, we propose ScaleDeep, a scalable DRL-based routing scheme for SDN, which improves the routing performance and is resilient to topology changes. Essentially, ScaleDeep takes advantage of partial control on network nodes and DRL. We select a set of critical nodes from a network as driver nodes, which can simulate the entire network operation, based on the control theory. By observing the traffic variation on the driver nodes, DRL dynamically adjusts some link weights for a weighted shortest path algorithm to change the routing paths and improve the routing performance. Limiting the control on driver nodes improves the convergence ability of DRL and reduces the dependency of the DRL agent on the fixed network topology. To validate the performance of ScaleDeep, we conduct packet-level simulations on different topologies. The results show that ScaleDeep outperforms existing DRL-based schemes by reducing the average flow completion time by up to 36% and exhibiting better robustness against minor topology changes.
Penghao Sun, Zehua Guo 0001, Junfei Li, Yang Xu 0010, Julong Lan, Yuxiang Hu 0001
IEEE/ACM Trans. Netw.1
2021 OrderSketch: An Unbiased and Fast Sketch for Frequency Estimation of Data Streams
Jie Lu 0006, Hongchang Chen, Penghao Sun, Tao Hu 0002, Zhen Zhang 0049
Comput. Networks3
2021 ScaleDRL: A Scalable Deep Reinforcement Learning Approach for Traffic Engineering in SDN with Pinning Control
Penghao Sun, Zehua Guo 0001, Julong Lan, Junfei Li, Yuxiang Hu 0001, Thar Baker
Comput. Networks1
2020 QOS-Aware Flow Control for Power-Efficient Data Center Networks with Deep Reinforcement Learning
abstract
Reducing the power consumption and maintaining the Flow Completion Time (FCT) for the Quality of Service (QoS) of applications in Data Center Networks (DCNs) are two major concerns for data center operators. However, existing works either fail in guaranteeing the QoS due to the neglect of the FCT constraints or achieve a less satisfying power efficiency. In this paper, we propose SmartFCT, which employs Software-Defined Networking (SDN) coupled with the Deep Reinforcement Learning (DRL) to improve the power efficiency of DCNs and guarantee the FCT. The DRL agent can generate a dynamic policy to consolidate traffic flows into fewer active switches in the DCN for power efficiency, and the policy also leaves different margins in different active links and switches to avoid FCT violation of unexpected short bursts of flows. Simulation results show that with similar FCT guarantee, SmartFCT can save 8% more of the power consumption compared to the state-of-the-art solutions.
Penghao Sun, Zehua Guo 0001, Sen Liu 0002, Julong Lan, Yuxiang Hu 0001
ICASSP1
2020 Improving the Scalability of Deep Reinforcement Learning-Based Routing with Control on Partial Nodes
abstract
Machine Learning (ML)-based routing optimization has been proposed to optimize the performance of flow routing for future networks, such as Software-Defined Networks (SDNs). However, existing studies are either hard to converge for large networks or vulnerable to topology changes. In this paper, we propose SINET, a scalable and intelligent network control framework for routing optimization. To improve the robustness and scalability, SINET selects several critical routing nodes to be directly controlled by a Deep Reinforcement Learning (DRL) agent, which dynamically generates routing policy to optimize network performance. Simulation results show that SINET can reduce the average flow completion time by at least 32% for a network with 82 nodes and exhibit better robustness against minor topology changes, compared to other DRL-based schemes.
Penghao Sun, Julong Lan, Zehua Guo 0001, Yang Xu 0010, Yuxiang Hu 0001
ICASSP1
2020 DeepMigration: Flow Migration for NFV with Graph-based Deep Reinforcement Learning
abstract
Network Function Virtualization (NFV) enables flexible deployment of network services as applications. Network operators expect to use a limited number of Network Function (NF) instances to handle the fluctuating traffic load and provide network services. However, it is a big challenge to guarantee the Quality of Service (QoS) under the unpredictable network traffic while minimizing the processing resources. One typical solution is to realize NF scale-out, scale-in and load balancing by elastically migrating the related traffic flows with SoftwareDefined Networking (SDN). However, it is difficult to optimally migrate flows since many real-time statuses of NF instances should be considered to make accurate decisions. In this paper, we propose DeepMigration to solve the problem by efficiently and dynamically migrating traffic flows among different NF instances. DeepMigration is a Deep Reinforcement Learning (DRL)-based solution coupled with Graph Neural Network (GNN). By taking advantages of the graph-based relationship deduction ability from our customized GNN and the self-evolution ability from the experience training of DRL, DeepMigration can accurately model the cost (e.g., migration latency) and the benefit (e.g., reducing the number of NF instances) of flow migration among different NF instances and generate dynamic and effective flow migration policies to improve the QoS. Experiment results show that DeepMigration requires less migration cost and saves up to 71.6{%} of the computation time than existing solutions.
Penghao Sun, Julong Lan, Zehua Guo 0001, Di Zhang 0002, Xianfu Chen, Yuxiang Hu 0001, Zhi Liu 0002
ICC1
2020 DeepWeave: Accelerating Job Completion Time with Deep Reinforcement Learning-based Coflow Scheduling
abstract
To improve the processing efficiency of jobs in distributed computing, the concept of coflow is proposed. A coflow is a collection of flows that are semantically correlated in a multi-stage computation task. A job consists of multiple coflows and can be usually formulated as a Directed-Acyclic Graph (DAG). A proper scheduling of coflows can significantly reduce the completion time of jobs in distributed computing. However, this scheduling problem is proved to be NP-hard. Different from existing schemes that use hand-crafted heuristic algorithms to solve this problem, in this paper, we propose a Deep Reinforcement Learning (DRL) framework named DeepWeave to generate coflow scheduling policies. To improve the inter-coflow scheduling ability in the job DAG, DeepWeave employs a Graph Neural Network (GNN) to process the DAG information. DeepWeave learns from the history workload trace to train the neural networks of the DRL agent and encodes the scheduling policy in the neural networks, which make coflow scheduling decisions without expert knowledge or a pre-assumed model. The proposed scheme is evaluated with a simulator using real-life traces. Simulation results show that DeepWeave completes jobs at least 1.7X faster than the state-of-the-art solutions.
Penghao Sun, Zehua Guo 0001, Junfei Li, Julong Lan, Yuxiang Hu 0001
IJCAI1
2020 Traffic modeling and optimization in datacenters with graph neural network
Junfei Li, Penghao Sun
Comput. Networks2
2020 SmartFCT: Improving power-efficiency for data center networks with deep reinforcement learning
Penghao Sun, Zehua Guo 0001, Sen Liu 0002, Julong Lan, Yuxiang Hu 0001
Comput. Networks1
2020 MARVEL: Enabling controller load balancing in software-defined networks with multi-agent reinforcement learning
Penghao Sun, Zehua Guo 0001, Gang Wang 0014, Julong Lan, Yuxiang Hu 0001
Comput. Networks1
2020 Efficient flow migration for NFV with Graph-aware deep reinforcement learning
Penghao Sun, Julong Lan, Junfei Li, Zehua Guo 0001, Tao Hu 0002
Comput. Networks1
2020 FTLink: Efficient and flexible link fault tolerance scheme for data plane in Software-Defined Networking
Tao Hu 0002, Peng Yi 0003, Julong Lan, Yuxiang Hu 0004, Penghao Sun
Future Gener. Comput. Syst.5
2019 ACST: Audit-based compromised switch tolerance for enhancing data plane robustness in software-defined networking
Tao Hu 0002, Peng Yi 0003, Julong Lan, Yuxiang Hu 0004, Penghao Sun
Comput. Networks5
2019 TIDE: Time-relevant deep reinforcement learning for routing optimization
Penghao Sun, Yuxiang Hu 0004, Julong Lan, Le Tian 0002, Min Chen 0003
Future Gener. Comput. Syst.1
2017 RFC: Range feature code for TCAM-based packet classification
Penghao Sun, Julong Lan
Comput. Networks1