VLDB 2026 Research / reviewers in the wild / expert
Heng Yu 0005
dblp:29/1429-5
· DBLP profile ↗
17ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0001-6907-9958ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 4 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DDoS Detection at the Scale of One Hundred Tbps
Yunming Xiao, Xijun Luo, Youliang Jiang, Aike Wang, Heng Yu 0005, Jiahao Cao 0001, Yong Jiang 0001, Jilong Wang 0001, Mingwei Xu 0001, Congcong Miao |
NSDI | 7 |
| 2026 | XFir: Accelerating New-Flow Setup on Host Servers of a Large Cloud NetworkabstractIn today's cloud networks, host servers widely deploy Data Processing Units (DPUs) as network accelerators under the "Sep-Path" paradigm. However, as server capabilities scale with increasing CPU cores and network bandwidth, the software slow path (executed on a DPU's CPU) has become a critical bottleneck for workloads with high new-flow rates. Meanwhile, new-flow setup logic on host servers must continuously evolve to meet diverse and changing customer demands, making flexibility a key requirement alongside performance. To address this gap, we present XFir, the first hardware-accelerated new-flow setup system for cloud host servers that delivers high CPS throughput while preserving sufficient flexibility. XFir leverages a next-generation DPU equipped with a Cloud Network co-Processor (CNP) to execute the host server's new-flow setup logic. XFir redesigns the host-server flow-setup datapath and table layout, optimizes LPM lookups, and introduces CPU-CNP collaboration mechanisms to further improve performance and reliability. Our evaluation shows that XFir achieves over 776K new-flow CPS on a single host server with 11.7μs slow-path latency. Compared to prior work (Fornax), XFir achieves 4.8x CPS and reduces latency by 69.2%. Moreover, XFir is cost-effective to deploy, requiring only a single DPU per host. Overall, XFir improves new-flow throughput while maintaining development flexibility at low financial cost. Shihan Lin, Shunqiao Jiang, Chao Pei, Jian Zhao 0006, Wenjun Wu 0001, Lijun Zhuang, Qingmin Liu, Heng Yu 0005, Yibo Huang 0005, Yifei Zhu 0001, Yunming Xiao, Ang Chen 0001, Linghe Kong, Congcong Miao |
SIGCOMM | 11 |
| 2026 | Dorado: Scaling SmartNIC Session Tables on Commodity DDRs
Heng Yu 0005, Jiajun Liang, Baozeng Zhang, Guozhi Lin, Xinyi Zhang 0004, Jian Zhao 0006, Ziyue Zhai, Chao Pei, Jilong Wang 0001, Gaogang Xie, Ang Chen 0001, Congcong Miao |
SIGCOMM | 1 |
| 2025 | Fornax: A Hardware-Centric Session Management in Large Public Cloud NetworkabstractSmartNIC is increasingly utilized to accelerate cloud network components. The effectiveness and correctness of hardware acceleration heavily rely on its management mechanism. Unfortunately, traditional management mechanisms adopt software-centric architecture, which treats flow as the basic management unit and completely relies on one-way commands to manage the flow table, making it challenging to support various cloud network scenarios while managing extremely large tables. In this paper, we advocate for a radical new mechanism to shift the management paradigm from software-centric architecture to hardware-centric architecture, which adopts session as the basic management unit and designs two-way protocols to facilitate the management process. We propose and implement a first-of-its-kind system, called Fornax, a novel management architecture for large public cloud networks. At the core of Fornax is leveraging a session-empowered hardware engine to provide various management capabilities. Besides, Fornax utilizes a light-weight software manager to enhance system scalability, and hardware-driven management protocols to improve resource efficiency. Our testbed evaluations demonstrate that Fornax can reduce the software storage usage by 80% and CPU usage by 77% with little hardware resource overhead. Our large-scale production results show that Fornax can manage up to 16M session entries while significantly reducing the resource overhead by over 79%. Heng Yu 0005, Jian Zhao 0006, Guozhi Lin, Baozeng Zhang, Yunpeng Guan, Jiajun Liang, Chao Pei, Yachen Wang, Xin Jin 0008, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 1 |
| 2023 | Serpens: A High Performance FaaS Platform for Network FunctionsabstractMore and more enterprises deploy applications on Function-as-a-Service (FaaS) platforms to improve resource efficiency and save monetary costs. Network Functions (NFs) suffer from staggered peaks of traffic patterns and could benefit from fine-grained resource multiplexing in FaaS platform. However, naively exploring existing FaaS platforms to support NFs can introduce significant performance overheads in three aspects, including slow instance startup, remote state access for NFs, and costly packet delivery between NFs. To address these problems, we propose${\sf Serpens}$, a high performance FaaS platform for NFs. First,${\sf Serpens}$proposes a reusable NF runtime design to slash instance startup overhead. Second,${\sf Serpens}$designs a novel state management mechanism to support local state access. Third,${\sf Serpens}$introduces an advanced service chaining approach to avoid extra packet delivery. Besides,${\sf Serpens}$designs an NF scaling mechanism to minimize performance fluctuation. We have implemented a prototype of${\sf Serpens}$and conducted comprehensive experiments. Compared with the NFs and Service Function Chains (SFCs) that run on existing FaaS platforms,${\sf Serpens}$can improve the throughput by more than 10× and reduce the latency by more than 90%. Heng Yu 0005, Han Zhang 0009, Junxian Shen, Yantao Geng, Jilong Wang 0001, Congcong Miao, Mingwei Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Scorpius: Proactive Code Preparation to Accelerate Function StartupabstractMassive enterprises deploy their applications on public clouds to relieve infrastructure management burden. However, applications are faced with highly fluctuating workloads, while clouds provision exclusive resources at coarse time granularity, resulting in severely low resource efficiency. Function-as-a-Service (FaaS) platform enables fine-grained resource multiplexing, which has the potential to improve efficiency. However, FaaS platforms could consume several seconds to start functions and the long startup latency can severely hurt the performance of applications. In this paper, we measure the FaaS platforms and find that most startup latency is occupied by code preparation. To reduce the code preparation latency with little resource overhead, we propose Scorpius, a FaaS platform that proactively prepares code based on the historical data of functions. It combines two optimization categories: (1) To reduce the code size, Scorpius proposes to proactively prepare partial libraries over servers and run functions on the server with most library sharing. (2) To advance the start time, Scorpius proposes to predict the function overload with a simple model and proactively scale code to more servers. We have implemented a prototype of Scorpius and conducted extensive experiments. Evaluation results demonstrate that compared with state-of-the-art methods, Scorpius can reduce the code preparation latency by 87.6% with only 9.3% storage overhead. Heng Yu 0005, Junxian Shen, Han Zhang 0009, Jilong Wang 0001, Congcong Miao, Mingwei Xu 0001 |
IWQoS | 1 |
| 2022 | Detecting Ephemeral Optical Events with OpTel
Congcong Miao, Minggang Chen, Arpit Gupta, Zili Meng, Lianjin Ye, Jingyu Xiao, Zekun He, Xulong Luo, Jilong Wang 0001, Heng Yu 0005 |
NSDI | 11 |
| 2021 | Predicting Crowd Flows via Pyramid Dilated Deeper Spatial-temporal NetworkabstractPredicting crowd flows is crucial for urban planning, traffic management and public safety. However, predicting crowd flows is not trivial because of three challenges: 1) highly heterogeneous mobility data collected by various services; 2) complex spatio-temporal correlations of crowd flows, including multi-scale spatial correlations along with non-linear temporal correlations. 3) diversity in long-term temporal patterns. To tackle these challenges, we proposed an end-to-end architecture, called pyramid dilated spatial-temporal network (PDSTN), to effectively learn spatial-temporal representations of crowd flows with a novel attention mechanism. Specifically, PDSTN employs the ConvLSTM structure to identify complex features that capture spatial-temporal correlations simultaneously, and then stacks multiple ConvLSTM units for deeper feature extraction. For further improving the spatial learning ability, a pyramid dilated residual network is introduced by adopting several dilated residual ConvLSTM networks to extract multi-scale spatial information. In addition, a novel attention mechanism, which considers both long-term periodicity and the shift in periodicity, is designed to study diverse temporal patterns. Extensive experiments were conducted on three highly heterogeneous real-world mobility datasets to illustrate the effectiveness of PDSTN beyond the state-of-the-art methods. Moreover, PDSTN provides intuitive interpretation into the prediction. Congcong Miao, Jiajun Fu, Jilong Wang 0001, Heng Yu 0005, Botao Yao, Anqi Zhong, Zekun He |
WSDM | 4 |
| 2021 | Octans: Optimal Placement of Service Function Chains in Many-Core SystemsabstractNetwork Function Virtualization (NFV) offers service delivery flexibility and reduces overall costs by running service function chains (SFCs) on commodity servers with many cores. Existing solutions for placing SFCs in one server treat all CPU cores as equal and allocate isolated CPU cores to network functions (NFs). However, advanced servers often adopt Non-Uniform Memory Access (NUMA) architecture to improve the scalability of many-core systems. CPU cores are grouped into nodes, incurring performance degradation due to cross-node memory access and intra-node resource contention. Our evaluation shows that randomly selecting cores to place NFs in an SFC could suffer from 39.2 percent lower throughput comparing to an optimal placement solution. In this article, we propose Octans, an NFV orchestrator to achieve maximum aggregate throughput of all SFCs in many-core systems. Octans first formulates the optimization problem as a Non-Linear Integer Programming (NLIP) Model. Then we identify the key factor for problem solving as evaluating the throughput drop of an NF caused by other NFs in the same SFC or different SFCs, i.e., performance drop index, and propose a formal and accurate prediction model based on system level performance metrics. Finally, we propose two online algorithms to quickly find near-optimal placement solutions for one-time and incremental deployment. Extensive evaluation on a prototype implementation shows that Octans significantly improves the aggregate throughput comparing to two state-of-the-art placement solutions by 27.1 ~ 45.2 percent for one-time deployment and by 20.9 ~ 38.1 percent for incremental deployment, with very low prediction errors. Moreover, Octans could quickly find a near-optimal placement solution with tiny optimality gap. Heng Yu 0005, Zhilong Zheng, Junxian Shen, Congcong Miao, Chen Sun 0005, Hongxin Hu, Jun Bi, Jilong Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Serpens: A High-Performance Serverless Platform for NFVabstractMany enterprises run Network Function Virtualization (NFV) services on public clouds to relieve management burdens and reduce costs. However, NFV operators still face the burden of choosing the right types of virtual machines (VMs) for various network functions (NFs), as well as the cost of renting VMs at a granularity of months or years while many VMs remain idle during valley hours. A recent computing model named serverless computing automatically executes user-defined functions on requests arrival, and charges users based on the number of processed requests. For NFV operators, serverless computing has the potential of completely relieving NF management burden and significantly reducing costs. Nevertheless, naively exploring existing serverless platforms for NFV introduces significant performance overheads in three aspects, including high remote state access latency, long NF launching time, and high packet delivery latency between NFs. To address these problems, we propose Serpens, a high-performance serverless platform for NFV. Firstly, Serpens designs a novel state management mechanism to support local state access. Secondly, Serpens proposes an efficient NF execution model to provide fast NF launching and avoid extra packet delivery. We have implemented a prototype of Serpens. Evaluation results demonstrate that Serpens could significantly improve performance for NFs and service function chains (SFCs) comparing to existing serverless platforms. Junxian Shen, Heng Yu 0005, Zhilong Zheng, Chen Sun 0005, Mingwei Xu 0001, Jilong Wang 0001 |
IWQoS | 2 |
| 2019 | Buffet: Enabling Multi-Tenant Network FunctionsabstractMany enterprises outsource traffic processing to third- party Network Function (NF) service providers to relieve management burden and reduce cost. NF providers have to process packets from multiple tenants simultaneously. However, most existing software based NFs are designed for one single tenant without internal state isolation mechanisms. These NFs cannot be securely shared across multiple tenants. Existing solutions that support multitenancy are either inefficient or ad-hoc for specific NFs. In this paper, we propose Buffet, a general and efficient framework that enables multitenancy for a wide range of NFs. First, Buffet introduces a general programming abstraction for various NFs to relieve NF developers from considering isolation details. Second, Buffet proposes dynamic tenant-level affinity to achieve high performance and resource efficiency. Finally, Buffet exploits SmartNIC offloading to eliminate host CPU overhead. We have implemented a prototype of Buffet. Evaluation results demonstrate that Buffet can effectively enable multitenancy for a wide range of NFs with high performance and resource efficiency. Heng Yu 0005, Junxian Shen, Chen Sun 0005, Zhilong Zheng, Jilong Wang 0001 |
GLOBECOM | 1 |
| 2019 | Octans: Optimal Placement of Service Function Chains in Many-Core SystemsabstractNetwork Function Virtualization (NFV) has the potential to offer service delivery flexibility and reduce overall costs by running service function chains (SFCs) on commodity servers with many cores. Existing solutions for placing SFCs in one server treat all CPU cores as equal and allocate isolated CPU cores to different network functions (NFs). However, advanced servers often adopt Non-Uniform Memory Access (NUMA) architecture to improve the scalability of many-core systems. CPU cores are grouped into nodes, incurring performance bottleneck due to cross-node memory access and intra-node resource contention. Our evaluation shows that randomly selecting cores to place NFs in an SFC could suffer from 39.2% lower throughput comparing to an optimal placement solution. In this paper, we propose Octans, an NFV orchestrator to achieve maximum aggregate throughput of all SFCs in many-core systems. Octans first formulates the optimization problem as a Non-Linear Integer Programming (NLIP) model. Then we identify the key factor for problem solving as evaluating the throughput drop of an NF caused by other NFs in the same SFC or different SFCs, i.e. performance drop index, and propose a formal and precise prediction model based on system level performance metrics. Finally, we propose an efficient heuristic algorithm to quickly find near-optimal placement solutions. We have implemented a prototype of Octans. Extensive evaluation shows that Octans significantly improves the aggregate throughput comparing to two state-of the-art placement mechanisms by 26.7%~51.8%, with very low prediction errors of SFC performance (an average deviation of 2.6%). Moreover, Octans could quickly find a near-optimal placement solution with tiny optimality gap (1.2%~3.5%). Zhilong Zheng, Jun Bi, Heng Yu 0005, Haiping Wang 0002, Chen Sun 0005, Hongxin Hu |
INFOCOM | 3 |
| 2018 | GEN: A GPU-Accelerated Elastic Framework for NFVabstractNetwork Function Virtualization (NFV) has the potential to enhance service delivery flexibility and reduce overall costs by provisioning software-based service function chains (SFCs) on commodity hardware. However, we observe that existing CPU-based SFC solutions cannot achieve both high performance and high elasticity simultaneously. To address such a critical challenge, we seek beyond CPU and exploit the capability of Graphics Processing Unit (GPU) to support NFV. We propose GEN, a GPU-based high performance and elastic framework for NFV. As opposed to pipeline-based SFCs in existing GPU-based NFV systems, GEN proposes to support RTC-based SFCs to improve processing performance. Meanwhile, GEN offers great elasticity of network function (NF) scaling up and down by allocating a different number of fine-grained GPU threads to an NF during runtime. We have implemented a prototype of GEN. Preliminary evaluation results demonstrate that GEN improves performance with RTC-based SFCs, and supports adaptive, precise, and fast NF scaling for NFV. Zhilong Zheng, Jun Bi, Chen Sun 0005, Heng Yu 0005, Hongxin Hu, Zili Meng, Shuhe Wang, Kai Gao 0001 |
APNet | 4 |
| 2018 | Grus: Enabling Latency SLOs for GPU-Accelerated NFV SystemsabstractGraphics Processing Unit (GPU) has been recently exploited as a hardware accelerator to improve the performance of Network Function Virtualization (NFV). However, GPU-accelerated NFV systems suffer from significant latency variation when multiple network functions (NFs) are co-located in the same machine, which prevents operators from supporting latency Service Level Objectives (SLOs). Existing research efforts to address this problem can only guarantee a limited number of SLOs with very low resource utilization efficiency. In this paper, we present the Grus framework to support latency SLOs in GPU-accelerated NFV systems. Grus thoroughly analyzes the sources of latency variation and proposes three design principles: (1) dynamic batch size setting is needed to bound packet batching latency in CPU; (2) a reordering mechanism for data transfer over PCI-E is required to guarantee the stalling time; and (3) maximizing concurrency in GPU is necessary to avoid NF execution waiting time. Guided by the principles, Grus consists of two logical layers including an infrastructure layer and a scheduling layer. The infrastructure layer is equipped with an in-CPU Reorder-able Worker Pool that could adjust batching size and packet transfer order, and in-GPU Controllable Concurrent Executors to provide maximized concurrency. The scheduling layer runs a heuristic algorithm to perform accurate and fast scheduling to guarantee SLOs based on our prediction models. We have implemented a prototype of Grus. Extensive evaluations demonstrate that Grus can significantly reduce latency variation and satisfy 4.5 × more SLO terms than state-of-the-art solutions. Zhilong Zheng, Jun Bi, Haiping Wang 0002, Chen Sun 0005, Heng Yu 0005, Hongxin Hu, Kai Gao 0001 |
ICNP | 5 |
| 2018 | SmartPartNet: Part-Informed Person Detection for Body-Worn SmartphonesabstractWe are interested in the development of image-based person detection algorithms for wearable computing using commodity smartphones. We focus on the use of smartphones as a wearable device because it is a practical means of augmenting human sensing for applications such as navigation for the blind or assisting social interaction. We identify two unique features of developing a vision-based person detector for body-worn smartphones: (1) the detector must take into account the strong bias in the size of people in the images taken with a wearable device and (2) the detector must consider the low image quality due to dim lighting and rapid ego-motion which leads to motion blur. In order to account for the unique distribution over the visibility of body parts when using a wearable camera, we propose a part-based person detector specialized for chestmounted smartphones. We perform extensive ablative analysis on the usefulness of part information, providing several insights regarding the design of the optimal person detector across different application domains. To account for the frequent occurrence of motion blur in our target domain, we introduce a data augmentation technique to generate synthetic motion-blurred images during training. In addition to addressing the aforementioned features, the final detector must also run in real-time using only smartphone resources. We leverage recent progress in deep neural networks for mobile devices and show that our proposed person detector, SmartPartNet, obtains performance similar to state-of-the-art pedestrian detection networks, while being 3X smaller and 5X faster. Heng Yu 0005, Eshed Ohn-Bar, Donghyun Yoo, Kris Makoto Kitani |
WACV | 1 |
| 2018 | Left Atrial Appendage Segmentation Using Fully Convolutional Neural Networks and Modified Three-Dimensional Conditional Random FieldsabstractThrombosis has become a global disease threatening human health. The left atrial appendage (LAA) is a major source of thrombosis in patients with atrial fibrillation (AF). Positive correlation exists between LAA volume and AF risk. LAA morphology has been suggested to influence thromboembolic risk in AF patients and to help predict thromboembolic events in low-risk patient groups. Automatic segmentation of LAA can greatly help physicians diagnose AF. In consideration of the large anatomical variations of the LAA, we proposed a robust method for automatic LAA segmentation on computed tomographic angiography (CTA) data using fully convolutional neural networks with three-dimensional (3-D) conditional random fields (CRFs). After manual localization of ROI of LAA, we adopted the FCN in natural image segmentation and transferred their learned models by fine-tuning the networks to segment each 2-D LAA slice. Subsequently, we used a modified dense 3-D CRF that accounts for the 3-D spatial information and larger contextual information to refine the segmentations of all slices. Our method was evaluated on 150 sets of CTA data using five-fold cross validation. Compared with manual annotation, we obtained a mean dice overlap of and a mean volume overlap of with a computation time of less than 40 s per volume. Experimental results demonstrated the robustness of our method in dealing with large anatomical variations and computational efficiency for adoption in a daily clinical routine.). Cheng Jin 0007, Jianjiang Feng, Heng Yu 0005, Jiang Liu 0014, Jiwen Lu, Jie Zhou 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2017 | NFP: Enabling Network Function Parallelism in NFVabstractSoftware-based sequential service chains in Network Function Virtualization (NFV) could introduce significant performance overhead. Current acceleration efforts for NFV mainly target on optimizing each component of the sequential service chain. However, based on the statistics from real world enterprise networks, we observe that 53.8% network function (NF) pairs can work in parallel. In particular, 41.5% NF pairs can be parallelized without causing extra resource overhead. In this paper, we present NFP, a high performance framework, that innovatively enables network function parallelism to improve NFV performance. NFP consists of three logical components. First, NFP provides a policy specification scheme for operators to intuitively describe sequential or parallel NF chaining intents. Second, NFP orchestrator intelligently identifies NF dependency and automatically compiles the policies into high performance service graphs. Third, NFP infrastructure performs light-weight packet copying, distributed parallel packet delivery, and load-balanced merging of packet copies to support NF parallelism. We implement an NFP prototype based on DPDK in Linux containers. Our evaluation results show that NFP achieves significant latency reduction for real world service chains. Chen Sun 0005, Jun Bi, Zhilong Zheng, Heng Yu 0005, Hongxin Hu |
SIGCOMM | 4 |