VLDB 2026 Research / reviewers in the wild / expert
Yibo Huang 0005
dblp:246/5079 · also Bobo Huang
· DBLP profile ↗
12ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-9215-4298ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Skyline: A Cloud Centric Internet Monitoring Engine
Shixian Guo, Yangyang Bai, Kefei Liu 0004, Zhenyang Zhong, Sisi Wen, Yongbin Dong, Anjian Chen, Jiale Feng, Lingpei Meng, Siwan Chen, Juntao Zhong, Chaoran Hu, Yibo Huang 0005, Yiming Qiu 0001 |
NSDI | 24 |
| 2026 | XFir: Accelerating New-Flow Setup on Host Servers of a Large Cloud NetworkabstractIn today's cloud networks, host servers widely deploy Data Processing Units (DPUs) as network accelerators under the "Sep-Path" paradigm. However, as server capabilities scale with increasing CPU cores and network bandwidth, the software slow path (executed on a DPU's CPU) has become a critical bottleneck for workloads with high new-flow rates. Meanwhile, new-flow setup logic on host servers must continuously evolve to meet diverse and changing customer demands, making flexibility a key requirement alongside performance. To address this gap, we present XFir, the first hardware-accelerated new-flow setup system for cloud host servers that delivers high CPS throughput while preserving sufficient flexibility. XFir leverages a next-generation DPU equipped with a Cloud Network co-Processor (CNP) to execute the host server's new-flow setup logic. XFir redesigns the host-server flow-setup datapath and table layout, optimizes LPM lookups, and introduces CPU-CNP collaboration mechanisms to further improve performance and reliability. Our evaluation shows that XFir achieves over 776K new-flow CPS on a single host server with 11.7μs slow-path latency. Compared to prior work (Fornax), XFir achieves 4.8x CPS and reduces latency by 69.2%. Moreover, XFir is cost-effective to deploy, requiring only a single DPU per host. Overall, XFir improves new-flow throughput while maintaining development flexibility at low financial cost. Shihan Lin, Shunqiao Jiang, Chao Pei, Jian Zhao 0006, Wenjun Wu 0001, Lijun Zhuang, Qingmin Liu, Heng Yu 0005, Yibo Huang 0005, Yifei Zhu 0001, Yunming Xiao, Ang Chen 0001, Linghe Kong, Congcong Miao |
SIGCOMM | 13 |
| 2026 | DistDPU: A Disaggregated DPU Architecture for High-Performance and Cost-Efficient AI CloudsabstractAI training and inference are driving cloud networks toward terabit-per-second (Tbps) bandwidth per server, challenging the scalability and efficiency of today's cloud network architectures. A prevalent design scales bandwidth by stacking monolithic Data Processing Units (DPUs), but this approach tightly couples control and data plane resources, leading to excessive cost, power consumption, and operational complexity. We identify a fundamental control-data plane divergence in AI clouds: while data plane bandwidth demand grows rapidly, control plane demand remains largely flat due to the dominance of elephant flows. As a result, monolithic DPUs become systematically over-provisioned when used as bandwidth scaling primitives. Lizhou Gao, Yuanyi Zhu, Chao Pei, Chuhao Chen 0001, Zijian Li 0003, Jian Zhao 0006, Dongbo Gu, Hongchen Ren, Jiyuan Chen, Yunpeng Guan, Jianye Yuan, Yibo Huang 0005, Yang Xu 0010 |
SIGCOMM | 17 |
| 2025 | Exposing RDMA NIC Resources for Software-Defined SchedulingabstractPeer Reviewed Yibo Huang 0005, Yiming Qiu 0001, Yunming Xiao, Archit Bhatnagar, Sylvia Ratnasamy, Ang Chen 0001 |
APNet | 1 |
| 2025 | Remote Direct Code ExecutionabstractWe propose remote direct code execution (RDX), which elevates the power of RDMA from memory access to code execution. We target runtime extension frameworks such as Wasm filters, BPF programs, and UDF functions, where RDX enables an agentless architecture that unlocks capabilities such as fast extension injection, update consistency guarantees, and minimal resource contention. We outline the roadmap for RDX around a new CodeFlow abstraction, encompassing programming remote extensions, exposing management stubs, remotely validating and JIT compiling code, seamlessly linking code to local context, managing remote extension state, and synchronizing code to targets. The case studies and initial results demonstrate the feasibility of RDX and its potential to spark the next wave of RDMA innovations. Yibo Huang 0005, Yiming Qiu 0001, Daqian Ding, Patrick Tser Jern Kon, Yiwen Zhang 0008, Yuzhou Mao, Archit Bhatnagar, Mosharaf Chowdhury, Srini Devadas, Jiarong Xing, Ang Chen 0001 |
HotNets | 1 |
| 2023 | PFtree: Optimizing Persistent Adaptive Radix Tree for PM Systems on eADR Platform
Rui Zhang 0112, Shangyi Sun, Lulu Chen, Yibo Huang 0005, Ming Yan 0009, Jie Wu 0003 |
DASFAA (1) | 5 |
| 2023 | Simplifying Cloud Management with Cloudless ComputingabstractCloud computing has transformed the IT industry, but managing cloud infrastructures remains a difficult task. We make a case for putting today's management practices, known as "Infrastructure-as-Code," on a firmer ground via a principled design. We call this end goal Cloudless Computing: it aims to simplify cloud infrastructure management tasks by supporting them "as-a-service," analogous to serverless computing that relieves users of the burden of managing server instances. By assisting tenants with these tasks, cloud resources will be presented to their users more readily without the undue burden of complex control. We describe the research problems by examining the typical lifecycle of today's cloud infrastructure management, and identify places where a cloudless approach will advance the state of the art. Yiming Qiu 0001, Patrick Tser Jern Kon, Jiarong Xing, Yibo Huang 0005, Xinyu Wang 0006, Peng Huang 0005, Mosharaf Chowdhury, Ang Chen 0001 |
HotNets | 4 |
| 2023 | Remote Direct Memory Introspection
Jiarong Xing, Yibo Huang 0005, Danyang Zhuo, Srini Devadas, Ang Chen 0001 |
USENIX Security Symposium | 3 |
| 2022 | An ultra-low latency and compatible PCIe interconnect for rack-scale communicationabstractEmerging network-attached resource disaggregation architecture requires ultra-low latency rack-scale communication. However, current hardware offloading (e.g., RDMA) and user-space (e.g., mTCP) communication schemes still rely on heavily layered protocol stacks which requires the translation between PCIe bus and network protocol, or complex connection/memory resource management within RNICs, inevitably bringing latency overhead. Yibo Huang 0005, Ming Yan 0009, Cunming Liang, Yang Xu 0010, Wenxiong Zou, Yiming Zhang 0018, Rui Zhang 0112, Chunpu Huang, Jie Wu 0003 |
CoNEXT | 1 |
| 2020 | BPS: A reliable and efficient pub/sub communication model with blockchain-enhanced paradigm in multi-tenant edge cloud
Yibo Huang 0005, Rui Zhang 0112, Zhihui Lu 0002, Yiming Zhang 0018, Jie Wu 0003, Lu Zhan, Patrick C. K. Hung |
J. Parallel Distributed Comput. | 1 |
| 2020 | BoR: Toward High-Performance Permissioned Blockchain in RDMA-Enabled NetworkabstractKnown as a distributed ledger, blockchain is becoming prevalent due to its decentralization, traceability and tamper resistance. Particularly, permissioned blockchain such as Hyperledger Fabric shows great application prospects as the infrastructure of IoT security, credit management, etc. Many cloud platforms like AWS, Azure, Oracle and IBM cloud currently provide blockchain as a service, in which tenants can quickly build permissioned blockchain and run smart contract based applications. However, the transactions throughput and scalability in the permissioned blockchain are not ideal, despite many optimization efforts in consensus protocol and parallel chain. Existing solutions still reveals some limitations like excessive CPU scheduling, inefficient block broadcast and high latency of initial blocks synchronization when new nodes join blockchain network. Inspired by the emerging RDMA (Remote Direct Memory Access) network, we propose BoR, an RDMA-based permissioned blockchain framework. By offloading the block transfer transaction into RDMA NICs, it can increase block broadcast speed and reduce block sync delay. We exploit the RDMA primitives to redesign the block synchronization protocol and accelerate DPoS (Delegated Proof of Stake) consensus process for higher throughput and lower latency in kernel-bypass manner. As demonstrated in our evaluation with different workloads, BoR with lower CPU utilization significantly outperforms the state-of-the-art EoS blockchain. Yibo Huang 0005, Zhihui Lu 0002, Xin Zhou 0009, Jie Wu 0003, Qifeng Tang, Patrick C. K. Hung |
IEEE Trans. Serv. Comput. | 1 |
| 2019 | RDMA-driven MongoDB: An approach of RDMA enhanced NoSQL paradigm for large-Scale data processing
Yibo Huang 0005, Zhihui Lu 0002, Ming Yan 0009, Jie Wu 0003, Patrick C. K. Hung, Qifeng Tang |
Inf. Sci. | 1 |