VLDB 2026 Research / reviewers in the wild / expert
Yiting Xia
dblp:87/10380
· DBLP profile ↗
23ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0003-2981-1459ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 5 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Matryoshka: Realizing Hyperscale Data Center Network Design for the AI Era
Yan Cai 0018, Jialong Li 0006, Kutalmis Akpinar, Hany Morsy, Sunil Khaunte, Yiting Xia, Ying Zhang 0022 |
NSDI | 8 |
| 2026 | SyncWise: Error-Aware Time Synchronization for Reconfigurable Data Center Networks
Yiming Lei 0002, Jialong Li 0006, Zhengqing Liu, Raj Joshi, Yiting Xia |
NSDI | 5 |
| 2026 | OpenOptics: Enabling Open Research and Implementation of Optical Data Center Networks
Yiming Lei 0002, Federico De Marchi 0002, Jialong Li 0006, Raj Joshi, Shu-Ting Wang, Balakrishnan Chandrasekaran 0002, Yiting Xia |
NSDI | 8 |
| 2026 | Unlocking Diversity of Fast-Switched Optical Data Center Networks With Unified RoutingabstractOptical data center networks (DCNs) are emerging as a promising solution for cloud infrastructure in the post-Moore’s Law era, particularly with the advent of “fast-switched” optical architectures capable of circuit reconfiguration at microsecond or even nanosecond scales. However, frequent reconfiguration of optical circuits introduces a unique challenge: in-flight packets risk loss during these transitions, hindering the deployment of many mature optical hardware designs due to the lack of suitable routing solutions. In this paper, we presentUnifiedRouting forOptical networks (URO), a general routing framework designed to support fast-switched optical DCNs across various hardware architectures. URO combines theoretical modeling of this novel routing problem with practical implementation on programmable switches, enabling precise, time-based packet transmission. Our prototype on Intel Tofino2 switches achieves a minimum circuit duration of$\mathrm {2~\mu \text {s} }$, ensuring end-to-end, loss-free application performance. Large-scale simulations using production DCN traffic validate URO’s generality across different hardware configurations, demonstrating its effectiveness and efficient system resource utilization. Jialong Li 0006, Federico De Marchi 0002, Yiming Lei 0002, Raj Joshi, Balakrishnan Chandrasekaran 0002, Yiting Xia |
IEEE Trans. Netw. | 6 |
| 2026 | Optimizing Mixture-of-Experts Inference Time via Model Deployment and Communication SchedulingabstractAs machine learning models scale in size and complexity, their computational requirements become a significant barrier. Mixture-of-Experts (MoE) models alleviate this issue by selectively activating relevant experts. Despite this, MoE models are hindered by high communication overhead from all-to-all operations, low GPU utilization, and complications from heterogeneous GPU environments. This paper presents Comet, which optimizes both model deployment and all-to-all communication scheduling to address these challenges in MoE inference. Comet achieves minimal communication times by strategically ordering token transmissions in all-to-all communications. It improves GPU utilization by colocating experts from different models on the same device, avoiding the limitations of all-to-all communication. We analyze Comet’s optimization strategies theoretically across four common GPU cluster settings: exclusive vs. colocated models on GPUs, and homogeneous vs. heterogeneous GPUs. Comet provides optimal solutions for three cases, and for the remaining NP-hard scenario, it offers a polynomial-time sub-optimal solution with only a 1.09× degradation from the optimal, as shown in the simulation results. Comet is the first approach to minimize MoE inference time via optimal model deployment and communication scheduling across various scenarios. Evaluations demonstrate that Comet significantly accelerates inference, achieving speedups of up to 2.63× in homogeneous clusters and 2.91× in heterogeneous environments. Moreover, Comet enhances GPU utilization by up to 2.38× compared to existing methods. Jialong Li 0006, Shreyansh Tripathi, Lakshay Rastogi, Yiming Lei 0002, Rui Pan 0003, Yiting Xia |
IEEE Trans. Netw. | 6 |
| 2025 | ClipMind: A Framework for Auditing Short-Format Video Recommendations Using Multimodal AI ModelsabstractWe are witnessing a significant shift in social media platforms; we are transitioning from chronological social media feeds to feeds that are driven by AI recommendation systems. While the main goal of AI recommendation systems is to suggest engaging content to users, there are also some associated risks: AI recommendation systems can promote extreme content, causing negative consequences like online polarization and user radicalization. Overall, there is a pressing need to design powerful techniques that allow us to audit AI recommendation systems. Motivated by this, our work introduces ClipMind, a scalable and generalizable framework using advanced AI models to audit these recommendation algorithms on short-format video platforms like TikTok and YouTube Shorts. We demonstrate the merits of our framework by collecting social media feeds from TikTok. Our analysis shows that TikTok’s recommendation algorithm increasingly recommends similar videos when a user expresses interest in mainstream topics like Food and Beauty Care. On the other hand, by investigating niche interests (War and Mental Health), we find no evidence of informational rabbit holes of extreme content on TikTok. Our work contributes to efforts that leverage AI for social good, as our framework can be used by several interested stakeholders, including users, social media platforms, regulators, and researchers, to understand and audit video-based algorithmic recommendations. Aoyu Gong, Sepehr Mousavi, Yiting Xia, Savvas Zannettou |
ICWSM | 3 |
| 2025 | Unlocking Superior Performance in Reconfigurable Data Center Networks with Credit-Based TransportabstractThe large-scale, end-to-end implementation of microsecond-switched reconfigurable data center networks (RDCNs), coupled with innovative routing and topology designs that provide continuous routes abstracting away frequent topology changes, demonstrates promise as a viable alternative to Clos networks in the post-Moore's Law era. However, the gap remains in transport performance, with current transport solutions falling short of unlocking their full performance. In this paper, we introduce Flare, a novel credit-based transport protocol that ensures reliable traffic delivery, low latency, and leverages the rapidly reconfiguring circuits of the RDCN to opportunistically route traffic over short paths, maximizing throughput. In simulations, Flare enables RDCNs to outperform Clos networks, achieving up to 1.15× higher throughput even under adversarial traffic. Additionally, it delivers up to 2× and 1.5× higher throughput than NDP and ExpressPass, and up to 10×, 15×, and 3.5× shorter flow completion time (FCT) than ExpressPass, TDTCP, and Bolt. Our testbed implementation further demonstrates the feasibility of Flare's mechanisms with programmable switches and DPDK. Federico De Marchi 0002, Jialong Li 0006, Ying Zhang 0022, Wei Bai 0001, Yiting Xia |
SIGCOMM | 5 |
| 2024 | Occam: A Programming System for Reliable Network ManagementabstractThe complexity of large networks makes their management a daunting task. State-of-the-art network management tools use workflow systems for automation, but they do not adequately address the substantial challenges in operation reliability. This paper presents Occam, a programming system that simplifies the development of reliable network management tasks. We leverage the fact that most modern network management systems are backed with a source-of-truth database, and thus customize database techniques to the context of network management. Occam exposes an easy-to-use programming model for network operators to express the key management logic, while shielding them from reliability concerns, such as operational conflicts and task atomicity. Instead, the Occam runtime provides these reliability guardrails automatically. Our evaluation demonstrates Occam's effectiveness in simplifying management tasks, minimizing network vulnerable time and assisting with failure recovery. Jiarong Xing, Kuo-Feng Hsu, Yiting Xia, Yan Cai 0018, Ying Zhang 0022, Ang Chen 0001 |
EuroSys | 3 |
| 2024 | Uniform-Cost Multi-Path Routing for Reconfigurable Data Center NetworksabstractReconfigurable data center networks (RDCNs) are arising as a promising data center network (DCN) design in the post-Moore's law era. However, the constantly reconfigured network topology in RDCNs invalidates the assumption of using hop count as the cost metric for routing, e.g., the status quo Equal-Cost Multi-Path routing (ECMP) in traditional DCNs. Unfortunately, existing routing solutions in RDCNs stick to the old assumption and deliver suboptimal performance either high in latency or low in bandwidth efficiency. In this paper, we redefine the cost metric for RDCN routing with uniform cost to unify the effects of topology disruption and hop count on latency and bandwidth efficiency. We propose Uniform-Cost Multi-Path routing (UCMP), an ECMP equivalent for RDCNs, where minimizing uniform cost leads flows of various sizes to the right balance between latency and bandwidth efficiency. Our simulation shows that UCMP achieves 53% to 98% lower flow completion time (FCT) and 1.55× bandwidth efficiency compared to the state-of-the-art RDCN routing strategy, and our testbed implementation demonstrates sustainable switch resource usage of UCMP as RDCNs scale. Jialong Li 0006, Haotian Gong, Federico De Marchi 0002, Aoyu Gong, Yiming Lei 0002, Wei Bai 0001, Yiting Xia |
SIGCOMM | 7 |
| 2022 | "What color are the fish's scales?" Exploring parents' and children's natural interactions with a child-friendly virtual agent during storybook readingabstractWith increasing integration of AI-powered agents into educational technologies available to families with young children, the landscape of how caregivers and children interact together or separately with these technologies is underexplored. Understanding the nature of these interactions could critically inform the design of educational technologies to facilitate children’s learning, research methods for use in future studies involving adult-child dyad technology use, and policy decisions regarding the use of educational technology with young children. In this study, we explored the natural interactions among parent, child, and a child-friendly virtual rabbit character named Floppy. Floppy is a virtual agent in a Smart Speaker app that models adult dialogic reading and conversational strategies for use with young children (ages 4-6 years). Over a span of four to six weeks, 18 parent-child dyads read The Rainbow Fish, a classic children’s book by Marcus Pfister, during 24 at-home, remote sessions with Floppy. Of the 189 conversations generated during this time, 125 were initiated by a prompt spoken by Floppy. Though there were some variations among the dyads, across all conversations, parent-driven interactions made up 63% of the conversations, followed by child-driven conversations at 15.3%, Floppy-driven at 14.3%, and Floppy-and-parent-driven at 7.4%. A select few parents were more comfortable having their children interact directly with Floppy, whereas the majority of the parents would direct children’s attention back to themselves or help children understand the questions by repeating or reformulating Floppy’s prompts. More than half of the parents reported that their children formed emotional connections with the virtual character. These findings point to a need to clearly define the role of virtual agents, even ones with limited AI, in this type of triadic interaction. Grace C. Lin, Ilana Schoenfeld, Meredith M. Thompson, Yiting Xia, Cigdem Uz Bilgin, Kathryn Leech |
IDC | 4 |
| 2022 | Hop-On Hop-Off Routing: A Fast Tour across the Optical Data Center Network for Latency-Sensitive FlowsabstractOptical data center networks show promise to serve as the next-generation cloud infrastructure especially with their cost and power benefits. The need to set up dedicated optical circuits between endpoints before they can exchange data, however, delays latency-sensitive (“mice”) flows. We find the state-of-the-art solution to reducing flow latency produces sub-optimal paths. To address this issue, we leverage programmable switches to realize Hop-On Hop-Off (HOHO) routing, where mice flows are forwarded along the minimal-latency paths. We prove the optimality and robustness of our algorithm and sketch an implementation on programmable switches. In our packet-level simulations, HOHO routing reduces the flow-completion times for mice flows by up to 35% and the average path length by 15% compared to the state-of-the-art solution. Jialong Li 0006, Yiming Lei 0002, Federico De Marchi 0002, Raj Joshi, Balakrishnan Chandrasekaran 0002, Yiting Xia |
APNet | 6 |
| 2022 | Efficient flow scheduling in distributed deep learning training with echelon formationabstractThis paper discusses why flow scheduling does not apply to distributed deep learning training and presents EchelonFlow, the first network abstraction to bridge the gap. EchelonFlow deviates from the common belief that semantically related flows should finish at the same time. We reached the key observation, after extensive workflow analysis of diverse training paradigms, that distributed training jobs observe strict computation patterns, which may consume data at different times. We devise a generic method to model the drastically different computation patterns across training paradigms, and formulate EchelonFlow to regulate flow finish times accordingly. Case studies of mainstream training paradigms under EchelonFlow demonstrate the expressiveness of the abstraction, and our system sketch suggests the feasibility of an EchelonFlow scheduling system. Rui Pan 0003, Yiming Lei 0002, Jialong Li 0006, Binhang Yuan, Yiting Xia |
HotNets | 6 |
| 2021 | A Social Network Under Social Distancing: Risk-Driven Backbone Management During COVID-19 and Beyond
Yiting Xia, Ying Zhang 0022, Zhizhen Zhong, Guanqing Yan, Chiunlin Lim, Satyajeet Ahuja, Soshant Bali, Alexander Nikolaidis, Kimia Ghobadi, Manya Ghobadi |
NSDI | 1 |
| 2021 | Capacity-efficient and uncertainty-resilient backbone network planning with hoseabstractThis paper presents Facebook's design and operational experience of a Hose-based backbone network planning system. This initial adoption of the Hose model in network planning is driven by the capacity and demand uncertainty pressure of backbone expansion. Since the Hose model abstracts the aggregated traffic demand per site, peak traffic flows at different times can be multiplexed to save capacity and buffer traffic spikes. Our core design involves heuristic algorithms to select Hose-compliant traffic matrices and cross-layer optimization between the optical and IP networks. We evaluate the system performance in production and share insights from years of production experience. Hose-based network planning can save 17.4% capacity and drops 75% less traffic under fiber cuts. As the first study of Hose in network planning, our work has the potential to inspire follow-up research. Satyajeet Ahuja, Vinayak Dangui, Soshant Bali, Abishek Gopalan, Petr Lapukhov, Yiting Xia, Ying Zhang 0022 |
SIGCOMM | 8 |
| 2021 | ARROW: restoration-aware traffic engineeringabstractFiber cut events reduce the capacity of wide-area networks (WANs) by several Tbps. In this paper, we revive the lost capacity by reconfiguring the wavelengths from cut fibers into healthy fibers. We highlight two challenges that made prior solutions impractical and propose a system called Arrow to address them. First, our measurements show that contrary to common belief, in most cases, the lost capacity is only partially restorable. This poses a cross-layer challenge from the Traffic Engineering (TE) perspective that has not been considered before: “Which IP links should be restored and by how much to best match the TE objective?” To address this challenge, Arrow's restoration-aware TE system takes a set of partial restoration candidates (that we call LotteryTickets) as input and proactively finds the best restoration plan. Second, prior work has not considered the reconfiguration latency of amplifiers. However, in practical settings, amplifiers add tens of minutes of reconfiguration delay. To enable fast and practical restoration, Arrow leverages optical noise loading and bypasses amplifier reconfiguration altogether. We evaluate Arrow using large-scale simulations and a testbed. Our testbed demonstrates Arrow's end-to-end restoration latency is eight seconds. Our large-scale simulations compare Arrow to the state-of-the-art TE schemes and show it can support 2.0x--2.4x more demand without compromising 99.99% availability. Zhizhen Zhong, Manya Ghobadi, Alaa Khaddaj, Jonathan Leach, Yiting Xia, Ying Zhang 0022 |
SIGCOMM | 5 |
| 2020 | Weaver: Efficient Coflow Scheduling in Heterogeneous Parallel NetworksabstractLeveraging application-level requirements expressed in Coflows has been shown to improve application-level communication efficiency. However, most existing works assume all application traffic is serviced by one monolithic network. This over-simplified assumption is no longer sufficient in a modern, evolving data center which operates on multiple generations of network fabrics, an architecture that we define as Heterogeneous Parallel Networks (HPNs). In this paper, we present the first scheduler, called Weaver, that addresses the Coflow management problem in HPNs. To exploit HPNs fully, achieving high communication efficiency for applications is crucial, yet it is also challenging because of the complex traffic patterns in applications and the heterogeneous bandwidth distribution in HPNs. Weaver addresses these challenges at two levels. At the microscopic level, for each application, Weaver leverages an efficient algorithm to exploit the distributed bandwidth in HPNs, which we proved to be within a constant factor of the optimal. At the macroscopic level involving multiple applications, Weaver can adopt a range of application traffic scheduling policies as desired by the system operator. Under realistic traffic, Weaver enables HPNs to service Coflows as efficiently as a monolithic network with equivalent aggregated capacity. Xin Sunny Huang, Yiting Xia, T. S. Eugene Ng |
IPDPS | 2 |
| 2018 | Republic: Data Multicast Meets Hybrid Rack-Level Interconnections in Data CenterabstractData multicast is a crucial data transfer pattern in distributed big-data processing. However, due to the lack of network and system level support, data processing relies on unicast-based application layer multicast. In recent years, there has been a surge in interest in using various emerging circuit switching technologies to build data centers having hybrid packet-circuit switched rack-level interconnections, i.e., hybrid data centers. These physical layer innovations fundamentally change the inter-rack communication capability, especially the capability of multicast communication. We propose Republic, a complete system that addresses the challenging issues in achieving high-performance data multicast in hybrid data centers. Republic abstracts the underlying network complexity as a data multicast service and provides a unified Republic API for data center applications requesting data multicast. Republic is implemented and deployed in a hybrid data center testbed. Testbed evaluation shows that Republic can improve data multicast in Apache Spark machine learning applications by as much as 4.0x. Xiaoye Sun, Yiting Xia, Simbarashe Dzinamarira, Xin Sunny Huang, Dingming Wu 0002, T. S. Eugene Ng |
ICNP | 2 |
| 2018 | Masking failures from application performance in data center networks with shareable backupabstractShareable backup is an economical and effective way to mask failures from application performance. A small number of backup switches are shared network-wide for repairing failures on demand so that the network quickly recovers to its full capacity without applications noticing the failures. This approach avoids complications and ineffectiveness of rerouting. We propose ShareBackup as a prototype architecture to realize this concept and present the detailed design. We implement ShareBackup on a hardware testbed. Its failure recovery takes merely 0.73ms, causing no disruption to routing; and it accelerates Spark and Tez jobs by up to 4.1X under failures. Large-scale simulations with real data center traffic and failure model show that ShareBackup reduces the percentage of job flows prolonged by failures from 47.2% to as little as 0.78%. In all our experiments, the results for ShareBackup have little difference from the no-failure case. Dingming Wu 0002, Yiting Xia, Xiaoye Sun, Xin Sunny Huang, Simbarashe Dzinamarira, T. S. Eugene Ng |
SIGCOMM | 2 |
| 2017 | Stop Rerouting!: Enabling ShareBackup for Failure Recovery in Data Center NetworksabstractThis paper introduces sharable backup as a novel solution to failure recovery in data center networks. It allows the entire network to share a small pool of backup devices. This proposal is grounded in three key observations. First, the traditional rerouting-based failure recovery is ineffective, because bandwidth loss from failures degrades application performance drastically. Therefore, failed devices should be replaced to restore bandwidth. Second, failures in data centers are rare but destructive [11], so it is desirable to seek cost-effective backup options. Third, the emergence of configurable data center network architectures promises feasibility of bringing backup devices online dynamically. We design the ShareBackup prototype architecture to realize this idea. Compared to rerouting-based solutions, ShareBackup provides more bandwidth with short path length at low cost. Yiting Xia, Xin Sunny Huang, T. S. Eugene Ng |
HotNets | 1 |
| 2017 | A Tale of Two Topologies: Exploring Convertible Data Center Network Architectures with Flat-treeabstractThis paper promotes convertible data center network architectures, which can dynamically change the network topology to combine the benefits of multiple architectures. We propose the flat-tree prototype architecture as the first step to realize this concept. Flat-tree can be implemented as a Clos network and later be converted to approximate random graphs of different sizes, thus achieving both Clos-like implementation simplicity and random-graph-like transmission performance. We present the detailed design for the network architecture and the control system. Simulations using real data center traffic traces show that flat-tree is able to optimize various workloads with different topology options. We implement an example flat-tree network on a 20-switch 24-server testbed. The traffic reaches the maximal throughput in 2.5s after a topology change, proving the feasibility of converting topology at run time. The network core bandwidth is increased by 27.6% just by converting the topology from Clos to approximate random graph. This improvement can be translated into acceleration of applications as we observe reduced communication time in Spark and Hadoop jobs. Yiting Xia, Xiaoye Sun, Simbarashe Dzinamarira, Dingming Wu 0002, Xin Sunny Huang, T. S. Eugene Ng |
SIGCOMM | 1 |
| 2016 | Flat-tree: A Convertible Data Center Network Architecture from Clos to Random GraphabstractClos networks are easy to implement, whereas random graphs have good performance. We propose flat-tree, a convertible data center network architecture, to combine the best of both worlds. Flat-tree can change the network topology dynamically, so the data center can be implemented as a Clos network and be converted to approximate random graphs of different sizes. To serve the heterogeneous workloads in data centers, flat-tree can organize the network as functionally separate zones each having a different topology. Workloads are placed into suitable zones that best optimize the performance. Simulation results demonstrate that flat-tree has similar performance to random graphs. Yiting Xia, T. S. Eugene Ng |
HotNets | 1 |
| 2015 | Blast: Accelerating high-performance data analytics applications by optical multicastabstractMulticast data dissemination is the performance bottleneck for high-performance data analytics applications in cluster computing, because terabytes of data need to be distributed routinely from a single data source to hundreds of computing servers. The state-of-the-art solutions for delivering these massive data sets all rely on application-layer overlays, which suffer from inherent performance limitations. This paper presents Blast, a system for accelerating data analytics applications by optical multicast. Blast leverages passive optical power splitting to duplicate data at line rate on a physical-layer broadcast medium separate from the packet-switched network core. We implement Blast on a small-scale hardware testbed. Multicast transmission can start 33ms after an application issues the request, resulting in a very small control overhead. We evaluate Blast's performance at the scale of thousands of servers through simulation. Using only a 10Gbps optical uplink per rack, Blast achieves upto 102× better performance than the state-of-the-art solutions even when they are used over a non-blocking core network with a 400Gbps uplink per rack. Yiting Xia, T. S. Eugene Ng, Xiaoye Sun |
INFOCOM | 1 |
| 2011 | RPIM: Inferring BGP Routing Policies in ISP NetworksabstractBGP dictates routing between autonomous systems with rich policy mechanisms in today's Internet. Operators translate high-level policy objectives into low-level router configurations without a comprehensive understanding of the actual effects on the network behavior, leaving the routing management an error-prone and time-consuming procedure. A fundamental question is: how to verify the intended routing principles against the actual routing effects of an ISP? In this paper, we develop a Routing Policy Inference Model (RPIM) as the first step towards addressing this fundamental issue. RPIM extracts various policy patterns from the BGP routing tables and translates them into high-level policy objectives of the ISP using a grouping and matching technique. Our work bridges the gap between the high-level policy objectives and the actual routing effects, which provides network operators with a novel approach to verify their policy design principles, thus facilitating the routing management. We evaluate our approach by extensive simulations using the Internet AS-level topology from CAIDA and the real routing data from the Abilene network. Simulation results show that RPIM achieves over 78.94% average inference accuracy in our suggested optimal threshold range. We also verify RPIM on several operating ISPs by the registered policies in an Internet Routing Registry (IRR). A representative case study on AS3292 demonstrates that RPIM effectively infers high-level policy objectives from routing data. Jingping Bi, Yiting Xia, Chengchen Hu |
GLOBECOM | 3 |