Yunming Xiao

dblp:235/7025 · DBLP profile ↗
← Back
26ranked-venue papers
12as first author
23since 2021 · last 2026
0000-0002-4913-4881ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 6 first-author · 17 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RAG on the Decentralized Web: An Empirical Study of Akash and Golem
abstract
Decentralized compute markets claim to be an alternative to the cloud, yet the systems community lacks hard evidence on whether they can run modern workloads end to end. We put this claim to the test by deploying a Retrieval-Augmented Generation (RAG) pipeline on the Akash and Golem networks, two popular decentralized computing markets. Our analysis reveals a striking paradox: supply is abundant, with over 10,000 CPU cores and 50 TiB of memory available, yet utilization remains minimal (about 0.3% on Akash). When mapped carefully, several RAG pipeline stages execute correctly and up to 3 × cheaper than AWS; others degrade under node churn and network bottlenecks. We also observe that decentralization in practice is already hybrid, with centralized gateways and offchain services masking blockchain complexity. Taken together, our findings suggest that decentralized markets are neither hype nor drop-in cloud replacements, but an underutilized substrate whose viability depends on better scheduling, reliability mechanisms, and trust primitives.
Suting Chen, Matteo Varvello, Aleksandar Kuzmanovic, Yunming Xiao
APNet4
2026 Secure Vickrey Auctions for Online Advertising
Archit Bhatnagar, Yunming Xiao, Ang Chen 0001, Amrita Roy Chowdhury 0001
NSDI2
2026 Cost-effective and Reliable Global Internet Peering with Programmable Switches
Congcong Miao, Zhiyi Yao, Jianchao Lv, Jinglin Wang, Shihan Lin, Xinyi Zhang 0004, Yunming Xiao, Jiwu Bu, Yachen Wang, Xianneng Zou, Yong Jiang 0001, Marco Canini, Gaogang Xie
NSDI7
2026 DDoS Detection at the Scale of One Hundred Tbps
Yunming Xiao, Xijun Luo, Youliang Jiang, Aike Wang, Heng Yu 0005, Jiahao Cao 0001, Yong Jiang 0001, Jilong Wang 0001, Mingwei Xu 0001, Congcong Miao
NSDI1
2026 XFir: Accelerating New-Flow Setup on Host Servers of a Large Cloud Network
abstract
In today's cloud networks, host servers widely deploy Data Processing Units (DPUs) as network accelerators under the "Sep-Path" paradigm. However, as server capabilities scale with increasing CPU cores and network bandwidth, the software slow path (executed on a DPU's CPU) has become a critical bottleneck for workloads with high new-flow rates. Meanwhile, new-flow setup logic on host servers must continuously evolve to meet diverse and changing customer demands, making flexibility a key requirement alongside performance. To address this gap, we present XFir, the first hardware-accelerated new-flow setup system for cloud host servers that delivers high CPS throughput while preserving sufficient flexibility. XFir leverages a next-generation DPU equipped with a Cloud Network co-Processor (CNP) to execute the host server's new-flow setup logic. XFir redesigns the host-server flow-setup datapath and table layout, optimizes LPM lookups, and introduces CPU-CNP collaboration mechanisms to further improve performance and reliability. Our evaluation shows that XFir achieves over 776K new-flow CPS on a single host server with 11.7μs slow-path latency. Compared to prior work (Fornax), XFir achieves 4.8x CPS and reduces latency by 69.2%. Moreover, XFir is cost-effective to deploy, requiring only a single DPU per host. Overall, XFir improves new-flow throughput while maintaining development flexibility at low financial cost.
Shihan Lin, Shunqiao Jiang, Chao Pei, Jian Zhao 0006, Wenjun Wu 0001, Lijun Zhuang, Qingmin Liu, Heng Yu 0005, Yibo Huang 0005, Yifei Zhu 0001, Yunming Xiao, Ang Chen 0001, Linghe Kong, Congcong Miao
SIGCOMM15
2026 Horizon: A Hyper-Edge Observability Engine for Live Streaming Networks
abstract
Live streaming services power mainstream real-time interactions on top of dedicated live streaming networks (LiveNets). Yet making LiveNets reliable at scale is challenging: failures arise on the userfacing delivery path and within streaming protocol and application logic, so operators need both continuous runtime monitoring to detect and localize incidents quickly and proactive preflight testing to exercise changes under representative environments and sustained playback behavior. Meeting these goals hinges on the right vantage point: the observability workflow must traverse the same network paths and delivery stacks as users while remaining controllable and non-intrusive. We present Horizon, which leverages near-user, provider-managed hyper-edge devices and orchestrates them into a shared fleet that supports both always-on monitoring and customizable, scenario-driven validation. Horizon has been deployed in production for over three years; in 2025, it identified 2,000+ major network incidents using 100,000+ hyper-edge agents.
Daqian Ding, Shixian Guo, Zhendong Xie, Aifang Xu, Changqian Wang, Kefei Liu 0004, Jialin Li 0001, Yunming Xiao, Heming Cui, Yiming Qiu 0001
SIGCOMM10
2026 CubeTrace: Microscopic Network Tracing for Heterogeneous Cloud Gateways
Yunming Xiao, Yinchao Yang, Jiaqi Zheng 0001, Xuqian Li, Dongbo Gu, Jun Zhang 0014, Miantao Wan, Chao Pei, Chen Tian 0001, Mingwei Xu 0001, Ang Chen 0001, Congcong Miao
SIGCOMM1
2025 Exposing RDMA NIC Resources for Software-Defined Scheduling
abstract
Peer Reviewed
Yibo Huang 0005, Yiming Qiu 0001, Yunming Xiao, Archit Bhatnagar, Sylvia Ratnasamy, Ang Chen 0001
APNet3
2025 SQUiD: Synthesizing Relational Databases from Unstructured Text
abstract
Relational databases are central to modern data management, yet most data exists in unstructured forms like text documents.To bridge this gap, we leverage large language models (LLMs) to automatically synthesize a relational database by generating its schema and populating its tables from raw text.We introduce SQUiD, a novel neurosymbolic framework that decomposes this task into four stages, each with specialized techniques.Our experiments show that SQUiD consistently outperforms baselines across diverse datasets.Our code and datasets are publicly available at
Mushtari Sadia, Zhenning Yang, Yunming Xiao, Ang Chen 0001, Amrita Roy Chowdhury 0001
EMNLP3
2025 PreAcher: Secure and Practical Password Pre-Authentication by Content Delivery Networks
Shihan Lin, Suting Chen, Yunming Xiao, Yanqi Gu, Aleksandar Kuzmanovic, Xiaowei Yang 0001
NSDI3
2025 Unlocking ECMP Programmability for Precise Traffic Control
Yunming Xiao, Weizhen Dang, Xiang Li 0223, Zekun He, Jilong Wang 0001, Aleksandar Kuzmanovic, Ang Chen 0001, Congcong Miao
NSDI2
2025 Enabling Anonymous Online Streaming Analytics at the Network Edge
abstract
In recent years, content hyper-giants have increasingly deployed server infrastructure and services close to end-users within “eyeball” networks. Still, online streaming analytics has largely remained unaffected by this trend. This is despite the fact that most of the “big data” is received in real-time and is most valuable at the time of arrival. The inability to process data at the network edge is caused by a common setting where user profiles, necessary for analytics, are stored deep in the data center backends. This setting also carries privacy concerns as such user profiles are individually identifiable, yet the users are almost blind to what data is associated with their identities and how the data is analyzed. In this article, we revise this arrangement, and plant encrypted semantic cookies at the user end. By redesigning the cookie content without altering existing protocols, semantic cookies enable the capture and pre-processing of user data at edge ISPs or CDNs while preserving user anonymity. Additionally, lightweight cryptographic algorithms like partially homomorphic encryption can protect web providers’ proprietary data from CDNs. We present Snatch , a QUIC-based streaming analytics prototype that achieves up to 200x faster user analytics, with common-case improvements of 10-30x.
Yunming Xiao, Yanqi Gu, Sen Lin 0009, Aleksandar Kuzmanovic
ACM Trans. Comput. Syst.1
2024 Snatch: Online Streaming Analytics at the Network Edge
abstract
In recent years, we have witnessed a growing trend of content hyper-giants deploying server infrastructure and services close to end-users, in "eyeball" networks. Still, one of the services that remained largely unaffected by this trend is online streaming analytics. This is despite the fact that most of the "big data" is received in real time and is most valuable at the time of arrival. The inability to process requests at the network edge is caused by a common setting where user profiles, necessary for analytics, are stored deep in the data center back-ends. This setting also carries privacy concerns as such user profiles are individually identifiable, yet the users are almost blind to what data is associated with their identities and how the data is analyzed. In this paper, we revise this arrangement, and plant encrypted semantic cookies at the user end. Without altering any of the existing protocols, this enables capturing and analytically pre-processing user requests soon after they are generated, at edge ISPs or content providers' off-nets. In addition, it ensures user anonymity perseverance during the analytics. We design and implement Snatch, a QUIC-based streaming analytics prototype, and demonstrate that it speeds up user analytics by up to 200x, and by 10-30x in the common case.
Yunming Xiao, Sen Lin 0009, Aleksandar Kuzmanovic
EuroSys1
2024 MegaTE: Extending WAN Traffic Engineering to Millions of Endpoints in Virtualized Cloud
abstract
In today's virtualized cloud, containers and virtual machines (VMs) are prevailing methods to deploy applications with different tenant requirements. However, these requirements are at odds with the resource allocation capabilities of conventional networking stacks in wide-area networks (WANs). In particular, existing WAN traffic engineering (TE) systems at the granularity of aggregated traffic flows are not designed to cater to each individual flow. In this paper, we advocate for a radical new approach to extend TE systems to involve millions of virtual instance endpoints. We propose and implement a first-of-its-kind system, called MegaTE, to satisfy the needs of each fine-grained traffic flow at the virtual instance level. At the core of the MegaTE system is the paradigm shift from the top-down centralized control to the bottom-up asynchronous query in the TE control loop, combined with eBPF-based segment routing on the data plane and TE optimization contraction on the control plane. We evaluate MegaTE using flow-level simulations with production traffic traces. Our results show that MegaTE supports 20× more endpoints with the similar algorithm run time compared to prior work. MegaTE has been adopted by large-scale public cloud providers. Notably, Tencent rolled out MegaTE in its cloud WAN since December 2022. Our production analysis shows that MegaTE reduces the packet latency of real-time applications by up to 51%.
Congcong Miao, Zhizhen Zhong, Yunming Xiao, Senkuo Zhang, Yinan Jiang, Zizhuo Bai, Chaodong Lu, Jingyi Geng, Zekun He, Yachen Wang, Xianneng Zou, Chuanchuan Yang
SIGCOMM3
2024 Conspirator: SmartNIC-Aided Control Plane for Distributed ML Workloads
Yunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao, Puneet Sharma 0001, Aleksandar Kuzmanovic
USENIX ATC1
2024 Mitigating Poor Data Quality Impact with Federated Unlearning for Human-Centric Metaverse
abstract
Federated Learning (FL), which has been employed to train machine learning models on the data with a distributed manner, could enhance the immersive user experience for the human-centric metaverse. However, it’s challenging to train machine learning models accurately and promptly with FL for the human-centric metaverse due to massive data communication and user unreliability. User experience could be negatively affected by using low-quality machine learning models for human-centric metaverse, e.g., it cannot scrutinize and arrive at decisions accurately and timely. To resolve this pressing issue, we propose MetaFul a federated unlearning solution which reduces the negative influences of low-quality data with no data transmission by removing low-quality training models at the server side. To be specific, MetaFul includes three main components. (i) Low-throughput federated learning (LT-FL) addresses the issue of large model transmission in FL by decreasing the dimension and the number of transmitted model parameters. (ii) Loss-based model quality assessment (LM-QA) utilizes the model loss generated in LT-FL to estimate user data quality. (iii) Non-communicative federated unlearning (NC-FUL) revokes the low-quality data impact on the FL model with careful designed federated unlearning at the server side. Both LM-QA and NC-FUL have no communications with clients. Finally, extensive evaluations are conducted to show MetaFul could improve the model accuracy by at least 2.5% and decrease the user perception time by at least 19.3% in human-centric metaverse compared to benchmarks.
Pengfei Wang 0013, Zongzheng Wei, Heng Qi, Shaohua Wan 0001, Yunming Xiao, Geng Sun 0001, Qiang Zhang 0008
IEEE J. Sel. Areas Commun.5
2023 TENSOR: Lightweight BGP Non-Stop Routing
abstract
As the solitary inter-domain protocol, BGP plays an important role in today's Internet. Its failures threaten network stability and will usually result in large-scale packet losses. Thus, the non-stop routing (NSR) capability that protects inter-domain connectivity from being disrupted by various failures, is critical to any Autonomous System (AS) operator. Replicating the BGP and underlying TCP connection status is key to realizing NSR. But existing NSR solutions, which heavily rely on OS kernel modifications, have become impractical due to providers' adoption of virtualized network gateways for better scalability and manageability.
Congcong Miao, Yunming Xiao, Marco Canini, Ruiqiang Dai, Shengli Zheng, Jilong Wang 0001, Jiwu Bu, Aleksandar Kuzmanovic, Yachen Wang
SIGCOMM2
2023 Demo: PDNS: A Fully Privacy-Preserving DNS
abstract
The Domain Name System (DNS) is a key component of Internet-based communication and its privacy has been neglected for years. Recently, DNS over HTTPS has improved the situation by fixing the issue of in-path middleboxes. Further progress has been made with proxy-based solutions such as Oblivious DoH, which separate a user's identity from their DNS queries. However, these solutions rely on non-collusion between DNS resolvers and proxy networks. This paper instead proposes PDNS, a new DNS extension that uses Private Information Retrieval to allow DNS resolvers to operate on blind queries, thereby eliminating any privacy leaks.
Yunming Xiao, Chenkai Weng, Ruijie Yu, Peizhi Liu, Matteo Varvello, Aleksandar Kuzmanovic
SIGCOMM1
2023 Decoding the Kodi Ecosystem
abstract
Free and open-source media centers are experiencing a boom in popularity for the convenience they offer users seeking to remotely consume digital content. Kodi is today’s most popular home media center, with millions of users worldwide. Kodi’s popularity derives from its ability to centralize the sheer amount of media content available on the Web, both free and copyrighted . Researchers have been hinting at potential security concerns around Kodi, due to add-ons injecting unwanted content as well as user settings linked with security holes. Motivated by these observations, this article conducts the first comprehensive analysis of the Kodi ecosystem: 15,000 Kodi users from 104 countries, 11,000 unique add-ons, and data collected over 9 months. Our work makes three important contributions. Our first contribution is that we build “crawling” software ( de-Kodi ) which can automatically install a Kodi add-on, explore its menu, and locate (video) content. This is challenging for two main reasons. First, Kodi largely relies on visual information and user input which intrinsically complicates automation. Second, the potential sheer size of this ecosystem (i.e., the number of available add-ons) requires a highly scalable crawling solution. Our second contribution is that we develop a solution to discover Kodi add-ons. Our solution combines Web crawling of popular websites where Kodi add-ons are published (LazyKodi and GitHub) and SafeKodi , a Kodi add-on we have developed which leverages the help of Kodi users to learn which add-ons are used in the wild and, in return, offers information about how safe these add-ons are, e.g., do they track user activity or contact sketchy URLs/IP addresses. Our third contribution is a classifier to passively detect Kodi traffic and add-on usage in the wild. Our analysis of the Kodi ecosystem reveals the following findings. We find that most installed add-ons are unofficial but safe to use. Still, 78% of the users have installed at least one unsafe add-on, and even worse, such add-ons are among the most popular. In response to the information offered by SafeKodi, one-third of the users reacted by disabling some of their add-ons. However, the majority of users ignored our warnings for several months attracted by the content such unsafe add-ons have to offer. Last but not least, we show that Kodi’s auto-update, a feature active for 97.6% of SafeKodi users, makes Kodi users easily identifiable by their ISPs. While passively identifying which Kodi add-on is in use is, as expected, much harder, we also find that many unofficial add-ons do not use HTTPS yet, making their passive detection straightforward. 1
Yunming Xiao, Matteo Varvello, Marc Anthony Warrior, Aleksandar Kuzmanovic
ACM Trans. Web1
2022 Blockchain Mining: Optimal Resource Allocation
abstract
Having enabled numerous applications, blockchains have attracted not only much attention, in the past decade, but also huge amount of resources: talent, capital, energy, etc. Focusing on the mining side of the market, in this paper, we aim at understanding how to efficiently use the resources mining and staking pools attract. We start with developing predictions about factors that increase the efficient allocation of pools' resources. We then test our predictions based on a general model for optimal resource allocation that we develop, as well as data we collected on pools' actual resource allocations. We find that pools can increase resource efficiency by mining for more blockchains as well as by increasing the frequency of resource re-allocation. Further, we enroll to mining pools as a miner to understand and comment on how pools can encourage their miners to increase the efficiency of their allocation. While our empirical investigation mostly focuses on the BTC family, we show that our theory and results are general and applicable to the Ethereum family as well as other proof-of-work (PoW) and proof-of-stake (PoS) chains.
Yunming Xiao, Sarit Markovich, Aleksandar Kuzmanovic
AFT1
2022 FIAT: frictionless authentication of IoT traffic
abstract
Home IoT (Internet of Things) deployments are vulnerable to local adversaries, compromising a LAN, and remote adversaries, compromising either the accounts associated with IoT devices or third-party devices like mobile phones used to control the IoT. There is, however, a fundamental difference between an attacker and a legitimate IoT user: the physical interaction with the device (e.g., via a mobile app) used to operate the IoT. Such physical interactions can be used to build frictionless authentications. However, their integration with IoT requires each vendor to independently adopt them, which is both complex and expensive. We instead design and build FIAT, the first third-party mechanism to automatically authorize IoT traffic by learning recurring traffic and validating human actions behind unpredictable traffic. FIAT does not require modification of the IoT devices or apps, as it operates passively on network traffic. Our evaluation shows that FIAT achieves high accuracy with minimal impact on the user experience.
Yunming Xiao, Matteo Varvello
CoNEXT1
2022 Blockchain-Enhanced Federated Learning Market With Social Internet of Things
abstract
The machine learning performance usually could be improved by training with massive data. However, requesters can only select a subset of devices with limited training data to execute federated learning (FL) tasks as a result of their limited budgets in today’s IoT scenario. To resolve this pressing issue, we devise a blockchain-enhanced FL market (BFL) to$(i)$make data in computationally bounded devices available for training with social Internet of things,$(ii)$maximize the amount of training data with given budgets for an FL task, and$(iii)$decentralize the FL market with blockchain. To achieve these goals, we firstly propose a trust-enhanced collaborative learning strategy (TCL) and a quality-oriented task allocation algorithm (QTA), where TCL enables training data sharing among trusted devices with social Internet of things, and QTA allocates suitable devices to execute FL tasks while maximizing the training quality with fixed budgets. Then, we devise an encrypted model training scheme (EMT) based on a simple but countervailable differential privacy methodology to prevent attacks from malicious devices. In addition, we also propose a contribution-driven delegated proof of stake (DPoS) consensus mechanism to guarantee the fairness of reward distribution in the block generation process. Finally, extensive evaluations are conducted to verify the proposed BFL could improve the total utility of requesters and average accuracy of FL models significantly.
Pengfei Wang 0013, Yian Zhao, Mohammad S. Obaidat, Zongzheng Wei, Heng Qi, Chi Lin 0001, Yunming Xiao, Qiang Zhang 0008
IEEE J. Sel. Areas Commun.7
2021 FIAT: frictionless authentication of IoT traffic
abstract
The average US household currently hosts more than 10 Internet of Things (IoT) devices [2]. Many research papers [5, 8] have demonstrated critical security concerns of the IoT, often due to lack of best practices like partial usage of HTTPS, or old ciphers. Even when best security practices are implemented, the IoT is still vulnerable to many attacks. Intruders can penetrate the home WiFi and directly control some IoT devices. They can compromise the account associated with an IoT device, mostly relying on username and password, or of third-party services like IFTTT [4]. They can also compromise the devices where IoT apps run, i.e., mostly mobile phones [1].
Yunming Xiao, Matteo Varvello
CoNEXT1
2020 De-Kodi: Understanding the Kodi Ecosystem
abstract
Free and open source media centers are currently experiencing a boom in popularity for the convenience and flexibility they offer users seeking to remotely consume digital content. This newfound fame is matched by increasing notoriety—for their potential to serve as hubs for illegal content—and a presumably ever-increasing network footprint. It is fair to say that a complex ecosystem has developed around Kodi, composed of millions of users, thousands of “add-ons”—Kodi extensions from 3rd-party developers—and content providers. Motivated by these observations, this paper conducts the first analysis of the Kodi ecosystem. Our approach is to build “crawling” software around Kodi which can automatically install an addon, explore its menu, and locate (video) content. This is challenging for many reasons. First, Kodi largely relies on visual information and user input which intrinsically complicates automation. Second, no central aggregators for Kodi addons exist. Third, the potential sheer size of this ecosystem requires a highly scalable crawling solution. We address these challenges with de-Kodi, a full fledged crawling system capable of discovering and crawling large cross-sections of Kodi’s decentralized ecosystem. With de-Kodi, we discovered and tested over 9,000 distinct Kodi addons. Our results demonstrate de-Kodi, which we make available to the general public, to be an essential asset in studying one of the largest multimedia platforms in the world. Our work further serves as the first ever transparent and repeatable analysis of the Kodi ecosystem at large.
Marc Anthony Warrior, Yunming Xiao, Matteo Varvello, Aleksandar Kuzmanovic
WWW2
2019 Close spatial arrangement of mutants favors and disfavors fixation
abstract
Cooperation is ubiquitous across all levels of biological systems ranging from microbial communities to human societies. It, however, seemingly contradicts the evolutionary theory, since cooperators are exploited by free-riders and thus are disfavored by natural selection. Many studies based on evolutionary game theory have tried to solve the puzzle and figure out the reason why cooperation exists and how it emerges. Network reciprocity is one of the mechanisms to promote cooperation, where nodes refer to individuals and links refer to social relationships. The spatial arrangement of mutant individuals, which refers to the clustering of mutants, plays a key role in network reciprocity. Besides, many other mechanisms supporting cooperation suggest that the clustering of mutants plays an important role in the expansion of mutants. However, the clustering of mutants and the game dynamics are typically coupled. It is still unclear how the clustering of mutants alone alters the evolutionary dynamics. To this end, we employ a minimal model with frequency independent fitness on a circle. It disentangles the clustering of mutants from game dynamics. The distance between two mutants on the circle is adopted as a natural indicator for the clustering of mutants or assortment. We find that the assortment is an amplifier of the selection for the connected mutants compared with the separated ones. Nevertheless, as mutants are separated, the more dispersed mutants are, the greater the chance of invasion is. It gives rise to the non-monotonic effect of clustering, which is counterintuitive. On the other hand, we find that less assortative mutants speed up fixation. Our model shows that the clustering of mutants plays a non-trivial role in fixation, which has emerged even if the game interaction is absent.
Yunming Xiao, Bin Wu 0016
PLoS Comput. Biol.1
2018 Common Knowledge Based Transfer Learning for Traffic Classification
abstract
Deep neural networks have been used for traffic classification, and promising results are obtained. However, most previous work confined to one specific classification task and restricted the classifiers potential performance and applications. As the traffic flows can be labeled from different perspectives, the performance of the classifier might be improved by exploring more meaningful latent features. For this purpose, we adopted a multi-output DNN model that simultaneously learns different traffic classification tasks. The common knowledge of traffic is exploited by the synergy among the tasks and boosts the individual performances of the tasks. Experiments show that this structure has the potential to meet new future demands and achieve the classification with advanced speeds and fair accuracies. Yet, due to the heavy training cost, the neural networks, though achieving good performance, are hard to implement in the real environment. We further show that few-shot learning could be a viable approach.
Yunming Xiao, Haifeng Sun 0001, Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001
LCN1