EDBT 2026 Demo / reviewers in the wild / expert
Minlan Yu
dblp:89/6345
· DBLP profile ↗
91ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0002-2381-0212ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 69 · 7 first-author · 23 since 2021Systems, architecture and hardware · 7 · 4 since 2021Software engineering, systems software and programming languages · 6 · 4 since 2021Security and privacy · 4 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Connecting 100K+ GPUs: Building the Communication Stack for Large-Scale LLM TrainingabstractThe arrival of 100K+ GPU clusters marks a new frontier in AI infrastructure. Standard communication stack meets new challenges as physical topologies span multiple datacenter buildings, introducing high bandwidth-delay product links where latency increases by up to 30× compared to intra-rack traffic. Furthermore, the transition toward Mixture-of-Experts architectures generating bursty all-to-all patterns that create transient congestion hotspots. These constraints, combined with an operational environment where hardware failures shift from anomalies to frequent occurrences, renders traditionally lightweight operations like initialization and resource management challenging. Hongyi Zeng, Min Si, Pavan Balaji, Yongzhou Chen, Ching-Hsiang Chu, Adithya Gangidi, Prashanth Kannan, Bingzhe Liu, Saif Hasan, Deep Shah, Ashmitha Jeevaraj Shetty, Gregory R. Steinbrecher, Srikanth Sundaresan, Yulun Wang, Yexin Wu, Mingran Yang, Kenny Yu, Minlan Yu, Cen Zhao, Shengbao Zheng, Wesley Bland, Denis Boyda, Suman Gumudavelli, Subodh Iyengar, Cristian Lumezanu, Rui Miao 0001, Venkat Ramesh, Jingliang Ren, Maxim Samoylov, Jan Seidel, Qiye Tan, Xinfeng Xie, Yimeng Zhao, Shuqiang Zhang, Art Zhu |
SIGCOMM | 19 |
| 2025 | OctoCache: Caching Voxels for Accelerating 3D Occupancy Mapping in Autonomous Systemsabstract3D mapping systems are crucial for creating digital representations of physical environments, widely used in autonomous robot navigation, 3D visualization, and AR/VR. This paper focuses on OctoMap, a leading 3D mapping framework using an octree-based structure for spatial efficiency. However, OctoMap's performance is limited by slow updates due to costly memory accesses. We introduce OctoCache, a software system that accelerates OctoMap through (1) optimized cache memory access, (2) refined voxel ordering, and (3) workflow parallelization. OctoCache achieves speedups of 45.63%~88.01% in 3D environment construction tasks compared to standard OctoMap. Deployed in UAV navigation scenarios, OctoCache demonstrates up to 3.02× speedup and reduces mission completion time by up to 28%. These results highlight OctoCache's potential to enhance 3D mapping efficiency in autonomous navigation, advancing robotics and environmental modeling. Peiqing Chen, Minghao Li 0003, Zishen Wan, Yu-Shun Hsiao, Minlan Yu, Vijay Janapa Reddi, Zaoxing Liu |
ASPLOS (2) | 5 |
| 2025 | Your network doesn't end at the NIC: A case for unifying the inter-host and intra-host networks in (AI) datacentersabstractModern ML workloads increasingly rely on direct communication between host devices—such as GPUs, NVMe SSDs, and DRAM—spanning intra-host and inter-host networks. However, today's intra-host network lacks hardware-level primitives for routing across heterogeneous interconnects, hindering efficient use of alternative paths and leading to sub-optimal performance under failures or congestion. Furthermore, the inter-host network treats the NIC as the endpoint, with intra-host interconnects like PCIe running oblivious to inter-host network protocols. This prevents leveraging multiple paths for communication between host devices across different servers. To address these limitations, we propose expanding the datacenter network layer to encompass the intra-host network, making intra-host devices first-class network endpoints. Our scheme envisions hardware-level routing and forwarding across multiple intra-host interconnects and makes intra-host devices visible to the inter-host network. This unified approach provides a principled foundation for robust, efficient peer-to-peer communication between storage and compute hardware devices in AI datacenters. Raj Joshi, Saksham Agarwal, ChonLam Lao, Minlan Yu |
HotNets | 4 |
| 2025 | Don't stop me Now: Embedding based Scheduling for LLMSabstractEfficient scheduling is crucial for interactive Large Language Model (LLM) applications, where low request completion time directly impacts user engagement. Size-based scheduling algorithms like Shortest Remaining Process Time (SRPT) aim to reduce average request completion time by leveraging known or estimated request sizes and allowing preemption by incoming jobs with shorter service times. However, two main challenges arise when applying size-based scheduling to LLM systems. First, accurately predicting output lengths from prompts is challenging and often resource-intensive, making it impractical for many systems. As a result, the state-of-the-art LLM systems default to first-come, first-served scheduling, which can lead to head-of-line blocking and reduced system efficiency. Second, preemption introduces extra memory overhead to LLM systems as they must maintain intermediate states for unfinished (preempted) requests.
In this paper, we propose TRAIL, a method to obtain output predictions from the target LLM itself. After generating each output token, we recycle the embedding of its internal structure as input for a lightweight classifier that predicts the remaining length for each running request. Using these predictions, we propose a prediction-based SRPT variant with limited preemption designed to account for memory overhead in LLM systems. This variant allows preemption early in request execution when memory consumption is low but restricts preemption as requests approach completion to optimize resource utilization. On the theoretical side, we derive a closed-form formula for this SRPT variant in an M/G/1 queue model, which demonstrates its potential value. In our system, we implement this preemption policy alongside our embedding-based prediction method. Our refined predictions from layer embeddings achieve 2.66x lower mean absolute error compared to BERT predictions from sequence prompts. TRAIL achieves 1.66x to 2.01x lower mean latency on the Alpaca dataset and 1.76x to 24.07x lower mean time to the first token compared to the state-of-the-art serving system. Rana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang, Minlan Yu, Michael Mitzenmacher |
ICLR | 5 |
| 2025 | Fast Inference for Augmented Large Language ModelsabstractAugmented Large Language Models (LLMs) enhance standalone LLMs by integrating external data sources through API calls. In interactive applications, efficient scheduling is crucial for maintaining low request completion times, directly impacting user engagement. However, these augmentations introduce new scheduling challenges: the size of augmented requests (in tokens) no longer correlates proportionally with execution time, making traditional size-based scheduling algorithms like Shortest Job First less effective. Additionally, requests may require different handling during API calls, which must be incorporated into scheduling.
This paper presents MARS, a novel inference framework that optimizes augmented LLM latency by explicitly incorporating system- and application-level considerations into scheduling. MARS introduces a predictive, memory-aware scheduling approach that integrates API handling and request prioritization to minimize completion time. We implement MARS on top of vLLM and evaluate its performance against baseline LLM inference systems, demonstrating improvements in end-to-end latency by 27%-85% and reductions in TTFT by 4%-96% compared to the existing augmented-LLM system, with even greater gains over vLLM. Our implementation is available online. Rana Shahout, Cong Liang 0005, Shiji Xin, Qianru Lao, Yong Cui 0001, Minlan Yu, Michael Mitzenmacher |
NeurIPS | 6 |
| 2025 | Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and Storage
Hamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu, Ramesh Govindan, Neal Cardwell, Nandita Dukkipati, Meng-Jung Tsai, Chris DeForeest, Yuxue Jin, Charles J. Carver, Jan Kopanski, Liqun Cheng, Amin Vahdat |
NSDI | 4 |
| 2025 | eTran: Extensible Kernel Transport with eBPF
Zhongjie Chen, Qingkai Meng 0001, ChonLam Lao, Fengyuan Ren, Minlan Yu, Yang Zhou 0008 |
NSDI | 6 |
| 2025 | Minder: Faulty Machine Detection for Large-scale Distributed Model Training
Yangtao Deng, Zhuo Jiang, Xingjian Zhang 0009, Zhang Zhang 0003, Zuquan Song, Gaohong Liu, Fuliang Li, Shuguang Wang, Haibin Lin, Jianxi Ye, Minlan Yu |
NSDI | 15 |
| 2025 | Decouple and Decompose: Scaling Resource Allocation with DeDe
Zhiying Xu, Minlan Yu, Francis Y. Yan |
OSDI | 2 |
| 2025 | HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM InferenceabstractDisaggregated Large Language Model (LLM) inference decouples the compute-intensive prefill stage from the memory-intensive decode stage, allowing low-end, compute-focused GPUs for prefill and high-end, memory-rich GPUs for decode, which reduces cost while maintaining high throughput. However, transmitting Key-Value (KV) data between the two stages can be a bottleneck, especially for long prompts. Additionally, the computational overhead in the two stages is key for optimizing Job Completion Time (JCT), and KV data size can become prohibitive for long prompts and sequences. Existing KV quantization methods can alleviate transmission and memory bottlenecks, but they introduce significant dequantization overhead, exacerbating the computation time. Zeyu Zhang 0005, Haiying Shen, Shay Vargaftik, Ran Ben-Basat, Michael Mitzenmacher, Minlan Yu |
SIGCOMM | 6 |
| 2025 | Intent-Driven Network Management with Multi-Agent LLMs: The Confucius FrameworkabstractAdvancements in Large Language Models (LLMs) are significantly transforming network management practices. In this paper, we present our experience developing Confucius, a multi-agent framework for network management at Meta. We model network management workflows as directed acyclic graphs (DAGs) to aid planning. Our framework integrates LLMs with existing management tools to achieve seamless operational integration, employs retrieval-augmented generation (RAG) to improve long-term memory, and establishes a set of primitives to systematically support human/model interaction. To ensure the accuracy of critical network operations, Confucius closely integrates with existing network validation methods and incorporates its own validation framework to prevent regressions. Remarkably, Confucius is a production-ready LLM development framework that has been operational for two years, with over 60 applications onboarded. To our knowledge, this is the first report on employing multi-agent LLMs for hyper-scale networks. Zhaodong Wang, Samuel Lin, Guanqing Yan, Soudeh Ghorbani, Minlan Yu, Jiawei Zhou 0012, Nathan Hu, Lopa Baruah, Sam Peters, Srikanth Kamath, Jerry Yang, Ying Zhang 0022 |
SIGCOMM | 5 |
| 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingabstractReliability is essential for ensuring efficiency in LLM training. However, many real-world reliability issues remain difficult to resolve, resulting in wasted resources and degraded model performance. Unfortunately, today's collective communication libraries operate as black boxes, hiding critical information needed for effective root cause analysis. Yangtao Deng, Qinlong Wang, Xiaoyun Zhi, Zhuo Jiang, Haohan Xu, Zuquan Song, Gaohong Liu, Shuguang Wang, Wencong Xiao, Jianxi Ye, Minlan Yu, Hong Xu 0001 |
SOSP | 15 |
| 2025 | Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
Weiqi Feng, Yangrui Chen, Yanghua Peng, Haibin Lin, Minlan Yu |
USENIX ATC | 6 |
| 2024 | SmartNIC Security Isolation in the Cloud with S-NICabstractModern smart NICs provide little isolation between the network functions belonging to different tenants. These NICs also do not protect network functions from the datacenter-provided management OS which runs on the smart NIC. We describe concrete attacks which allow a network function's state to leak to (or be modified by) another network function or the management OS. We then introduce S-NIC, a new hardware design for smart NICs that provides strong isolation guarantees. S-NIC pervasively virtualizes hardware accelerators, enforces single-owner semantics for each line in on-NIC cache and RAM, and provides dedicated bus bandwidth for each network function. Using this design, we eliminate side channels involving shared hardware state, and give each network function the illusion of having a private smart NIC. We show how these virtual NICs can be integrated with preexisting datacenter technologies for virtual LANs and trusted host-level computations like SGX enclaves. The overall result is that S-NIC enables strongly-isolated, NIC-accelerated datacenter applications; in these applications, network functions and host-level code receive hardware-guaranteed isolation from other applications and the datacenter provider. Yang Zhou 0008, Mark Wilkening, James W. Mickens, Minlan Yu |
EuroSys | 4 |
| 2024 | BitMatcher: Bit-level Counter Adjustment for SketchesabstractSketch has been widely used in the field of large-scale data stream processing. However, common fixed-counter algorithms such as Count-Min Sketch have to allocate larger counters, which wastes a lot of memory due to the high skewness of real-world data streams. To reduce memory usage, we propose to dynamically adjust the counter size that matches the distribution of the data stream. We introduce BitMatcher, a fast global-adjusting algorithm that automatically adjusts the counter to the appropriate size to match the data stream. During stream processing, BitMatcher identifies items hashed into a bucket based on isolated fingerprints. If it overflows, BitMatcher changes the flag bits in the bucket and dynamically increases or shrinks the size of some counters in a fine-grained manner. BitMatcher can also relocate a cold item in the bucket with the idea of cuckoo hashing to preserve the potential hot item while achieving global load balancing. Through the above way of dealing with overflow caused by skewed data, BitMatcher precisely manipulates allocated bits and maximizes memory utilization. The experiments show that BitMatcher has high throughput and can outperform SOTA by up to 4 orders of magnitude in terms of accuracy. We also deployed BitMatcher on several platforms, showing its software and hardware scalability. Qilong Shi, Chengjun Jia, Wenjun Li 0004, Zaoxing Liu, Tong Yang 0003, Jianan Ji, Gaogang Xie, Weizhe Zhang, Minlan Yu |
ICDE | 9 |
| 2024 | DINT: Fast In-Kernel Distributed Transactions with eBPF
Yang Zhou 0008, Xingyu Xiang, Matthew Kiley, Sowmya Dharanipragada, Minlan Yu |
NSDI | 5 |
| 2024 | THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
Minghao Li 0003, Ran Ben-Basat, Shay Vargaftik, ChonLam Lao, Kevin Xu, Michael Mitzenmacher, Minlan Yu |
NSDI | 7 |
| 2023 | Electrode: Accelerating Distributed Protocols with eBPF
Yang Zhou 0008, Zezhou Wang, Sowmya Dharanipragada, Minlan Yu |
NSDI | 4 |
| 2023 | Scalable Distributed Massive MIMO Baseband Processing
Junzhi Gong, Anuj Kalia, Minlan Yu |
NSDI | 3 |
| 2023 | Rearchitecting the TCP Stack for I/O-Offloaded Content Delivery
Deondre Martin Ng, Junzhi Gong, Youngjin Kwon, Minlan Yu, KyoungSoo Park |
NSDI | 5 |
| 2023 | Practical Intent-driven Routing Configuration Synthesis
Sivaramakrishnan Ramanathan, Ying Zhang 0022, Mohab Gawish, Yogesh Mundada, Zhaodong Wang, Sangki Yun, Eric Lippert, Walid Taha, Minlan Yu, Jelena Mirkovic |
NSDI | 9 |
| 2023 | Direct Telemetry AccessabstractFine-grained network telemetry is becoming a modern datacenter standard and is the basis of essential applications such as congestion control, load balancing, and advanced troubleshooting. As network size increases and telemetry gets more fine-grained, there is a tremendous growth in the amount of data needed to be reported from switches to collectors to enable network-wide view. As a consequence, it is progressively hard to scale data collection systems. Jonatan Langlet, Ran Ben-Basat, Gabriele Oliaro, Michael Mitzenmacher, Minlan Yu, Gianni Antichi |
SIGCOMM | 5 |
| 2023 | Teal: Learning-Accelerated Optimization of WAN Traffic EngineeringabstractThe rapid expansion of global cloud wide-area networks (WANs) has posed a challenge for commercial optimization engines to efficiently solve network traffic engineering (TE) problems at scale. Existing acceleration strategies decompose TE optimization into concurrent subproblems but realize limited parallelism due to an inherent tradeoff between run time and allocation performance. Zhiying Xu, Francis Y. Yan, Rachee Singh, Justin T. Chiu, Alexander M. Rush, Minlan Yu |
SIGCOMM | 6 |
| 2023 | Optimal Oblivious Routing With Concave Objectives for Structured NetworksabstractOblivious routing distributes traffic from sources to destinations following predefined routes with rules independent of traffic demands. While finding optimal oblivious routing with a concave objective is intractable for general topologies, we show that it is tractable for structured topologies often used in datacenter networks. To achieve this, we apply graph automorphism and prove the existence of the optimal automorphism-invariant solution. This result reduces the search space to targeting the optimal automorphism-invariant solution. We design an iterative algorithm to obtain such a solution by alternating between convex optimization and a linear program. The convex optimization finds an automorphism-invariant solution based on representative variables and constraints, making the problem tractable. The linear program generates adversarial demands to ensure the final result satisfies all possible demands. Since the construction of the representative variables and constraints are combinatorial problems, we design polynomial-time algorithms for the construction. We evaluate the iterative algorithm in terms of throughput performance, scalability, and generality over three potential applications. The algorithm i) improves the throughput up to 87.5% for partially deployed FatTree and achieves up to$2.55\times $throughput gain for DRing over heuristic algorithms, ii) scales for three considered topologies with a thousand switches, iii) applies to a general structured topology with non-uniform link capacity and server distribution. Kanatip Chitavisutthivong, Sucha Supittayapornpong, Pooria Namyar, Mingyang Zhang 0005, Minlan Yu, Ramesh Govindan |
IEEE/ACM Trans. Netw. | 5 |
| 2022 | Xatu: boosting existing DDoS detection systems using auxiliary signalsabstractTraditional DDoS attack detection monitors volumetric traffic features to detect attack onset. To reduce false positives, such detection is often conservative---raising an alert only after a sustained period of observed anomalous behavior. However, contemporary attacks tend to be short, which combined with a long detection delay means that most of the attack still reaches and impacts the victim. We propose Xatu, a system that utilizes auxiliary signals to improve the accuracy and timeliness of existing DDoS detection systems. We explore two types of auxiliary signals, attack preparation signals and the history of prior attacks. These signals can be easily mined from existing traffic monitoring systems in many ISP networks. To leverage these auxiliary signals for attack detection, we propose a multi-timescale LSTM model, which derives both long-term and short-term patterns from diverse auxiliary signals. We then leverage survival analysis to quickly detect attacks when they occur while minimizing false positives and thus scrubbing costs. We evaluate Xatu on traffic from a large ISP, using commercial defense alert data to label prevalent attack events. Xatu would help the commercial defense scrub up to 44.1% additional anomalous traffic and would reduce its median detection delay by 9.5 minutes.1 Zhiying Xu, Sivaramakrishnan Ramanathan, Alexander M. Rush, Jelena Mirkovic, Minlan Yu |
CoNEXT | 5 |
| 2022 | Optimal Oblivious Routing for Structured NetworksabstractOblivious routing distributes traffic from sources to destinations following predefined routes with rules independent of traffic demands. While finding optimal oblivious routing is intractable for general topologies, we show that it is tractable for structured topologies often used in datacenter networks. To achieve this, we apply graph automorphism and prove the existence of the optimal automorphism-invariant solution. This result reduces the search space to targeting the optimal automorphism-invariant solution. We design an iterative algorithm to obtain such a solution by alternating between two linear programs. The first program finds an automorphism-invariant solution based on representative variables and constraints, making the problem tractable. The second program generates adversarial demands to ensure the final result satisfies all possible demands. Since, the construction of the representative variables and constraints are combinatorial problems, we design polynomial-time algorithms for the construction. We evaluate proposed iterative algorithm in terms of throughput performance, scalability, and generality over three potential applications. The algorithm i) improves the throughput up to 87.5% over a heuristic algorithm for partially deployed FatTree, ii) scales for FatClique with a thousand switches, iii) is applicable to a general structured topology with non-uniform link capacity and server distribution. Sucha Supittayapornpong, Pooria Namyar, Mingyang Zhang 0005, Minlan Yu, Ramesh Govindan |
INFOCOM | 4 |
| 2022 | Evolvable Network Telemetry at Facebook
Yang Zhou 0008, Ying Zhang 0022, Minlan Yu, Dexter Cao, Yu-Wei Eric Sung, Starsky H. Y. Wong |
NSDI | 3 |
| 2022 | Carbink: Fault-Tolerant Far Memory
Yang Zhou 0008, Hassan M. G. Wassel, Sihang Liu 0001, James W. Mickens, Minlan Yu, Chris Kennelly, David E. Culler, Henry M. Levy, Amin Vahdat |
OSDI | 6 |
| 2022 | SwitchV: automated SDN switch validation with P4 modelsabstractIncreasing demand on computer networks continuously pushes manufacturers to incorporate novel features and capabilities into their switches at an ever-accelerating pace. However, the traditional approach to switch development relies on informal specifications and handcrafted tests to ensure reliability, which are tedious and slow to maintain and update, effectively putting feature velocity at odds with reliability. Kinan Dak Albab, Jonathan DiLorenzo, Stefan Heule, Ali Kheradmand, Steffen Smolka, Konstantin Weitz, Muhammad Timarzi, Minlan Yu |
SIGCOMM | 9 |
| 2022 | Hashing Design in Modern Networks: Challenges and Mitigation Techniques
Yunhong Xu, Keqiang He, Rui Wang 0025, Minlan Yu, Nick G. Duffield, Hassan M. G. Wassel, Shidong Zhang, Leonid B. Poutievski, Junlan Zhou, Amin Vahdat |
USENIX ATC | 4 |
| 2021 | Zero-CPU Collection with Direct Telemetry AccessabstractProgrammable switches are driving a massive increase in fine-grained measurements. This puts significant pressure on telemetry collectors that have to process reports from many switches. Past research acknowledged this problem by either improving collectors' stack performance or by limiting the amount of data sent from switches. In this paper, we take a different and radical approach: switches are responsible for directly inserting queryable telemetry data into the collectors' memory, bypassing their CPU, and thereby improving their collection scalability. We propose to use a method we call direct telemetry access, where switches jointly write telemetry reports directly into the same collector's memory region, without coordination. Our solution, DART, is probabilistic, trading memory redundancy and query success probability for CPU resources at collectors. We prototype DART using commodity hardware such as P4 switches and RDMA NICs and show that we get high query success rates with a reasonable memory overhead. For example, we can collect INT path tracing information on a fat tree topology without a collector's CPU involvement while achieving 99.9% query success probability and using just 300 bytes per flow. Jonatan Langlet, Ran Ben-Basat, Sivaramakrishnan Ramanathan, Gabriele Oliaro, Michael Mitzenmacher, Minlan Yu, Gianni Antichi |
HotNets | 6 |
| 2021 | A throughput-centric view of the performance of datacenter topologiesabstractWhile prior work has explored many proposed datacenter designs, only two designs, Clos-based and expander-based, are generally considered practical because they can scale using commodity switching chips. Prior work has used two different metrics, bisection bandwidth and throughput, for evaluating these topologies at scale. Little is known, theoretically or practically, how these metrics relate to each other. Exploiting characteristics of these topologies, we prove an upper bound on their throughput, then show that this upper bound better estimates worst-case throughput than all previously proposed throughput estimators and scales better than most of them. Using this upper bound, we show that for expander-based topologies, unlike Clos, beyond a certain size of the network, no topology can have full throughput, even if it has full bisection bandwidth; in fact, even relatively small expander-based topologies fail to achieve full throughput. We conclude by showing that using throughput to evaluate datacenter performance instead of bisection bandwidth can alter conclusions in prior work about datacenter cost, manageability, and reliability. Pooria Namyar, Sucha Supittayapornpong, Mingyang Zhang 0005, Minlan Yu, Ramesh Govindan |
SIGCOMM | 4 |
| 2021 | Aquila: a practically usable verification system for production-scale programmable data planesabstractThis paper presents Aquila, the first practically usable verification system for Alibaba's production-scale programmable data planes. Aquila addresses four challenges in building a practically usable verification: (1) specification complexity; (2) verification scalability; (3) bug localization; and (4) verifier self validation. Specifically, first, Aquila proposes a high-level language that facilitates easy expression of specifications, reducing lines of specification codes by tenfold compared to the state-of-the-art. Second, Aquila constructs a sequential encoding algorithm to circumvent the exponential growth of states associated with the upscaling of data plane programs to production level. Third, Aquila adopts an automatic and accurate bug localization approach that can narrow down suspects based on reported violations and pinpoint the culprit by simulating a fix for each suspect. Fourth and finally, Aquila can perform self validation based on refinement proof, which involves the construction of an alternative representation and subsequent equivalence checking. To this date, Aquila has been used in the verification of our production-scale programmable edge networks for over half a year, and it has successfully prevented many potential failures resulting from data plane bugs. Bingchuan Tian, Mengqi Liu 0001, Ennan Zhai, Yu Zhou 0008, Mengjing Ma, Xionglie Wei, Hongqiang Harry Liu, Ming Zhang 0005, Chen Tian 0001, Minlan Yu |
SIGCOMM | 16 |
| 2021 | Jaqen: A High-Performance Switch-Native Approach for Detecting and Mitigating Volumetric DDoS Attacks with Programmable Switches
Zaoxing Liu, Hun Namkung, Georgios Nikolaidis, Jeongkeun Lee, Changhoon Kim, Xin Jin 0008, Vladimir Braverman, Minlan Yu, Vyas Sekar |
USENIX Security Symposium | 8 |
| 2020 | Detecting routing loops in the data planeabstractRouting loops can harm network operation. Existing loop detection mechanisms, including mirroring packets, storing state on switches, or encoding the path onto packets, impose significant overheads on either the switches or the network. Jan Kucera 0004, Ran Ben-Basat, Mário Kuka, Gianni Antichi, Minlan Yu, Michael Mitzenmacher |
CoNEXT | 5 |
| 2020 | Challenging the Stateless Quo of Programmable SwitchesabstractProgrammable switches based on the Protocol Independent Switch Architecture (PISA) have greatly enhanced the flexibility of today's networks by allowing new packet protocols to be deployed without any hardware changes. They have also been instrumental in enabling a new computing paradigm in which parts of an application's logic run within the network core (in-network computing). Nadeen Gebara, Alberto Lerner, Mingran Yang, Minlan Yu, Paolo Costa, Manya Ghobadi |
HotNets | 4 |
| 2020 | Quantifying the Impact of Blocklisting in the Age of Address ReuseabstractBlocklists, consisting of known malicious IP addresses, can be used as a simple method to block malicious traffic. However, blocklists can potentially lead to unjust blocking of legitimate users due to IP address reuse, where more users could be blocked than intended. IP addresses can be reused either at the same time (Network Address Translation) or over time (dynamic addressing). We propose two new techniques to identify reused addresses. We built a crawler using the BitTorrent Distributed Hash Table to detect NATed addresses and use the RIPE Atlas measurement logs to detect dynamically allocated address spaces. We then analyze 151 publicly available IPv4 blocklists to show the implications of reused addresses and find that 53-60% of blocklists contain reused addresses having about 30.6K-45.1K listings of reused addresses. We also find that reused addresses can potentially affect as many as 78 legitimate users for as many as 44 days. Sivaramakrishnan Ramanathan, Anushah Hossain, Jelena Mirkovic, Minlan Yu, Sadia Afroz 0001 |
Internet Measurement Conference | 4 |
| 2020 | BLAG: Improving the Accuracy of Blacklists
Sivaramakrishnan Ramanathan, Jelena Mirkovic, Minlan Yu |
NDSS | 3 |
| 2020 | Routing Oblivious Measurement Analytics
Ran Ben-Basat, Gil Einziger, Shir Landau Feibish, Danny Raz, Minlan Yu |
Networking | 6 |
| 2020 | Enabling Premium Service for Streaming Video in Cellular Networks
Ramesh Govindan, Ajay Mahimkar, N. K. Shankaranarayanan, Jia Wang 0001, Minlan Yu |
Networking | 6 |
| 2020 | Poster: CO2: Collaborative Packet Classification for Network Functions with Overselection
Yunhong Xu, Hao Wu 0023, Nick G. Duffield, Bin Liu 0001, Minlan Yu |
Networking | 5 |
| 2020 | Sundial: Fault-tolerant Clock Synchronization for Datacenters
Gautam Kumar 0001, Hema Hariharan, Hassan M. G. Wassel, Peter Hochschild, Dave Platt, Simon L. Sabato, Minlan Yu, Nandita Dukkipati, Prashant Chandra, Amin Vahdat |
OSDI | 8 |
| 2020 | PINT: Probabilistic In-band Network TelemetryabstractCommodity network devices support adding in-band telemetry measurements into data packets, enabling a wide range of applications, including network troubleshooting, congestion control, and path tracing. However, including such information on packets adds significant overhead that impacts both flow completion times and application-level performance. Ran Ben-Basat, Sivaramakrishnan Ramanathan, Gianni Antichi, Minlan Yu, Michael Mitzenmacher |
SIGCOMM | 5 |
| 2020 | Scouts: Improving the Diagnosis Process Through Domain-customized Incident RoutingabstractIncident routing is critical for maintaining service level objectives in the cloud: the time-to-diagnosis can increase by 10x due to mis-routings. Properly routing incidents is challenging because of the complexity of today's data center (DC) applications and their dependencies. For instance, an application running on a VM might rely on a functioning host-server, remote-storage service, and virtual and physical network components. It is hard for any one team, rule-based system, or even machine learning solution to fully learn the complexity and solve the incident routing problem. We propose a different approach using per-team Scouts. Each teams' Scout acts as its gate-keeper --- it routes relevant incidents to the team and routes-away unrelated ones. We solve the problem through a collection of these Scouts. Our PhyNet Scout alone --- currently deployed in production --- reduces the time-to-mitigation of 65% of mis-routed incidents in our dataset. Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri, Vincent Liu 0001, Ricardo Bianchini, Ramaswamy Aditya, Xiaohang Wang 0008, Henry Lee, David A. Maltz, Minlan Yu, Behnaz Arzani |
SIGCOMM | 11 |
| 2020 | Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICsabstractProgrammable data plane has been moving towards deployments in data centers as mainstream vendors of switching ASICs enable programmability in their newly launched products, such as Broadcom's Trident-4, Intel/Barefoot's Tofino, and Cisco's Silicon One. However, current data plane programs are written in low-level, chip-specific languages (e.g., P4 and NPL) and thus tightly coupled to the chip-specific architecture. As a result, it is arduous and error-prone to develop, maintain, and composite data plane programs in production networks. This paper presents Lyra, the first cross-platform, high-level language & compiler system that aids the programmers in programming data planes efficiently. Lyra offers a one-big-pipeline abstraction that allows programmers to use simple statements to express their intent, without laboriously taking care of the details in hardware; Lyra also proposes a set of synthesis and optimization techniques to automatically compile this "big-pipeline" program into multiple pieces of runnable chip-specific code that can be launched directly on the individual programmable switches of the target network. We built and evaluated Lyra. Lyra not only generates runnable real-world programs (in both P4 and NPL), but also uses up to 87.5% fewer hardware resources and up to 78% fewer lines of code than human-written programs. Ennan Zhai, Hongqiang Harry Liu, Rui Miao 0001, Yu Zhou 0008, Bingchuan Tian, Chen Sun 0005, Dennis Cai, Ming Zhang 0005, Minlan Yu |
SIGCOMM | 10 |
| 2020 | Microscope: Queue-based Performance Diagnosis for Network FunctionsabstractBy moving monolithic network appliances to software running on commodity hardware, network function virtualization allows flexible resource sharing among network functions and achieves scalability with low cost. However, due to resource contention, network functions can suffer from performance problems that are hard to diagnose. In particular, when many flows traverse a complex topology of NF instances, it is hard to pinpoint root causes for a flow experiencing performance issues such as low throughput or high latency. Simply maintaining resource counters at individual NFs is not sufficient since the effect of resource contention can propagate across NFs and over time. In this paper, we introduce Microscope, a performance diagnosis tool, for network functions that leverages queuing information at NFs to identify the root causes (i.e., resources, NFs, traffic patterns of flows etc.). Our evaluation on realistic NF chains and traffic shows that we can correctly capture root causes behind 89.7% of performance impairments, up to 2.5 times more than the state-of-the-art tools with low overhead. Junzhi Gong, Muhammad Bilal Anwer, Aman Shaikh, Minlan Yu |
SIGCOMM | 5 |
| 2020 | Cheetah: Accelerating Database Queries with Switch PruningabstractModern database systems are growing increasingly distributed and struggle to reduce query completion time with a large volume of data. In this paper, we leverage programmable switches in the network to partially offload query computation to the switch. While switches provide high performance, they have resource and programming constraints that make implementing diverse queries difficult. To fit in these constraints, we introduce the concept of data pruning -- filtering out entries that are guaranteed not to affect output. The database system then runs the same query but on the pruned data, which significantly reduces processing time. We propose pruning algorithms for a variety of queries. We implement our system, Cheetah, on a Barefoot Tofino switch and Spark. Our evaluation on multiple workloads shows 40 - 200% improvement in the query completion time compared to Spark. Muhammad Tirmazi, Ran Ben-Basat, Minlan Yu |
SIGMOD Conference | 4 |
| 2019 | DETER: Deterministic TCP Replay for Performance Diagnosis
Rui Miao 0001, Mohammad Alizadeh, Minlan Yu |
NSDI | 4 |
| 2019 | HPCC: high precision congestion controlabstractCongestion control (CC) is the key to achieving ultra-low latency, high bandwidth and network stability in high-speed networks. From years of experience operating large-scale and high-speed RDMA networks, we find the existing high-speed CC schemes have inherent limitations for reaching these goals. In this paper, we present HPCC (High Precision Congestion Control), a new high-speed CC mechanism which achieves the three goals simultaneously. HPCC leverages in-network telemetry (INT) to obtain precise link load information and controls traffic precisely. By addressing challenges such as delayed INT information during congestion and overreac-tion to INT information, HPCC can quickly converge to utilize free bandwidth while avoiding congestion, and can maintain near-zero in-network queues for ultra-low latency. HPCC is also fair and easy to deploy in hardware. We implement HPCC with commodity programmable NICs and switches. In our evaluation, compared to DCQCN and TIMELY, HPCC shortens flow completion times by up to 95%, causing little congestion even under large-scale incasts. Rui Miao 0001, Hongqiang Harry Liu, Lingbo Tang, Zheng Cao 0003, Ming Zhang 0005, Frank Kelly, Mohammad Alizadeh, Minlan Yu |
SIGCOMM | 11 |
| 2019 | Risk based planning of network changes in evolving data centersabstractData center networks evolve as they serve customer traffic. When applying network changes, operators risk impacting customer traffic because the network operates at reduced capacity and is more vulnerable to failures and traffic variations. The impact on customer traffic ultimately translates to operator cost (e.g., refunds to customers). However, planning a network change while minimizing the risks is challenging as we need to adapt to a variety of traffic dynamics and cost functions while scaling to large networks and large changes. Today, operators often use plans that maximize the residual capacity (MRC), which often incurs a high cost under different traffic dynamics. Instead, we propose Janus, which searches the large planning space by leveraging the high degree of symmetry in data center networks. Our evaluation on large Clos networks and Facebook traffic traces shows that Janus generates plans in real-time only needing 33~71% of the cost of MRC planners while adapting to a variety of settings. Omid Alipourfard, Jérémie Koenig, Christopher Harshaw, Amin Vahdat, Minlan Yu |
SOSP | 6 |
| 2018 | SENSS Against Volumetric DDoS AttacksabstractVolumetric distributed denial-of-service (DDoS) attacks can bring any network to a halt. Because of their distributed nature and high volume, the victim often cannot handle these attacks alone and needs help from upstream ISPs. Today's Internet has no automated mechanism for victims to ask ISPs for help in attack handling and ISPs themselves do not offer such services. We propose SENSS, a security service for collaborative mitigation of volumetric DDoS attacks. SENSS enables the victim of an attack to request attack monitoring and filtering on demand, and to pay for the services rendered. Requests can be sent both to the immediate and to remote ISPs, in an automated and secure manner, and can be authenticated by these ISPs, without having prior trust with the victim. Simple and generic SENSS APIs enable victims to build custom detection and mitigation approaches against a variety of DDoS attacks. SENSS is deployable with today's infrastructure, and it has strong economic incentives both for ISPs and for the attack victims. It is also very effective in sparse deployment, offering full protection to direct customers of early adopters, and considerable protection to remote victims when deployed strategically. Deployment on the largest 1% of ISPs protects not just direct customers of these ISPs, but everyone on the Internet, from 90% of volumetric DDoS attacks. Sivaramakrishnan Ramanathan, Jelena Mirkovic, Minlan Yu, Ying Zhang 0022 |
ACSAC | 3 |
| 2018 | Wide-area analytics with multiple resourcesabstractRunning data-parallel jobs across geo-distributed sites has emerged as a promising direction due to the growing need for geo-distributed cluster deployment. A key difference between geo-distributed and intra-cluster jobs is the heterogeneous (and often constrained) nature of compute and network resources across the sites. We propose Tetrium, a system for multi-resource allocation in geo-distributed clusters, that jointly considers both compute and network resources for task placement and job scheduling. Tetrium significantly reduces job response time, while incorporating several other performance goals with simple control knobs. Our EC2 deployment and trace-driven simulations suggest that Tetrium improves the average job response time by up to 78% compared to existing data-locality-based solutions, and up to 55% compared to Iridium, the recently proposed geo-distributed analytics system. Chien-Chun Hung, Ganesh Ananthanarayanan, Leana Golubchik, Minlan Yu, Mingyang Zhang 0005 |
EuroSys | 4 |
| 2018 | Decoupling Algorithms and Optimizations in Network FunctionsabstractNetwork function virtualization promises a path to rapid innovation in networks. However, due to the complexity of developing these functions, innovations have been slow. Designing a network function is a daunting task that requires combining packet processing optimizations with the network function logic. It is not possible to ignore packet processing optimizations either: an optimized pipeline can have 3 times better performance than an unoptimized pipeline. In this paper, we introduce NFMorph, a framework wherein the network function logic is decoupled from the packet processing optimizations. Developers would specify the packet processing algorithm in a high level language. The runtime then identifies the best set of optimizations on the packet processing algorithm based on the domain knowledge specified by operators and optimization templates for common NF primitives. NFMorph can also justin-time reoptimize based on the workload and environment constraints. Omid Alipourfard, Minlan Yu |
HotNets | 2 |
| 2018 | Cold Filter: A Meta-Framework for Faster and More Accurate Stream ProcessingabstractApproximate stream processing algorithms, such as Count-Min sketch, Space-Saving, etc., support numerous applications in databases, storage systems, networking, and other domains. However, the unbalanced distribution in real data streams poses great challenges to existing algorithms. To enhance these algorithms, we propose a meta-framework, called Cold Filter (CF), that enables faster and more accurate stream processing. Yang Zhou 0008, Tong Yang 0003, Jie Jiang 0008, Bin Cui 0001, Minlan Yu, Xiaoming Li 0001, Steve Uhlig |
SIGMOD Conference | 5 |
| 2017 | Stream Aggregation Through Order SamplingabstractThis paper introduces a new single-pass reservoir weighted-sampling stream aggregation algorithm, Priority-Based Aggregation (PBA). While order sampling is a powerful and efficient method for weighted sampling from a stream of uniquely keyed items, there is no current algorithm that realizes the benefits of order sampling in the context of stream aggregation over non-unique keys. A naive approach to order sample regardless of key then aggregate the results is hopelessly inefficient. In distinction, our proposed algorithm uses a single persistent random variable across the lifetime of each key in the cache, and maintains unbiased estimates of the key aggregates that can be queried at any point in the stream. The basic approach can be supplemented with a Sample and Hold pre-sampling stage with a sampling rate adaptation controlled by PBA. This approach represents a considerable reduction in computational complexity compared with the state of the art in adapting Sample and Hold to operate with a fixed cache size. Concerning statistical properties, we prove that PBA provides unbiased estimates of the true aggregates. We analyze the computational complexity of PBA and its variants, and provide a detailed evaluation of its accuracy on synthetic and trace data. Weighted relative error is reduced by 40% to 65% at sampling rates of 5% to 17%, relative to Adaptive Sample and Hold; there is also substantial improvement for rank queries. Nick G. Duffield, Yunhong Xu, Liangzhen Xia, Nesreen K. Ahmed, Minlan Yu |
CIKM | 5 |
| 2017 | CherryPick: Adaptively Unearthing the Best Cloud Configurations for Big Data Analytics
Omid Alipourfard, Hongqiang Harry Liu, Jianshu Chen, Shivaram Venkataraman, Minlan Yu, Ming Zhang 0005 |
NSDI | 5 |
| 2017 | Enabling Wide-Spread Communications on Optical Fabric with MegaSwitch
Li Chen 0008, Kai Chen 0005, Zhonghua Zhu, Minlan Yu, George Porter, Chunming Qiao |
NSDI | 4 |
| 2017 | SilkRoad: Making Stateful Layer-4 Load Balancing Fast and Cheap Using Switching ASICsabstractIn this paper, we show that up to hundreds of software load balancer (SLB) servers can be replaced by a single modern switching ASIC, potentially reducing the cost of load balancing by over two orders of magnitude. Today, large data centers typically employ hundreds or thousands of servers to load-balance incoming traffic over application servers. These software load balancers (SLBs) map packets destined to a service (with a virtual IP address, or VIP), to a pool of servers tasked with providing the service (with multiple direct IP addresses, or DIPs). An SLB is stateful, it must always map a connection to the same server, even if the pool of servers changes and/or if the load is spread differently across the pool. This property is called per-connection consistency or PCC. The challenge is that the load balancer must keep track of millions of connections simultaneously. Rui Miao 0001, Hongyi Zeng, Changhoon Kim, Jeongkeun Lee, Minlan Yu |
SIGCOMM | 5 |
| 2016 | LossRadar: Fast Detection of Lost Packets in Data Center NetworksabstractPacket losses are common in data center networks, may be caused by a variety of reasons (e.g., congestion, blackhole), and have significant impacts on application performance and network operations. Thus, it is important to provide fast detection of packet losses independent of their root causes. We also need to capture both the locations and packet header information of the lost packets to help diagnose and mitigate these losses. Unfortunately, existing monitoring tools that are generic in capturing all types of network events often fall short in capturing losses fast with enough details and low overhead. Due to the importance of loss in data centers, we propose a specific monitoring system designed for loss detection. We propose LossRadar, a system that can capture individual lost packets and their detailed information in the entire network on a fine time scale. Our extensive evaluation on prototypes and simulations demonstrates that LossRadar is easy to implement in hardware switches, achieves low memory and bandwidth overhead, while providing detailed information about individual lost packets. We also build a loss analysis tool that demonstrates the usefulness of LossRadar with a few example applications. Rui Miao 0001, Changhoon Kim, Minlan Yu |
CoNEXT | 4 |
| 2016 | FlowRadar: A Better NetFlow for Data Centers
Rui Miao 0001, Changhoon Kim, Minlan Yu |
NSDI | 4 |
| 2016 | Trumpet: Timely and Precise Triggers in Data CentersabstractAs data centers grow larger and strive to provide tight performance and availability SLAs, their monitoring infrastructure must move from passive systems that provide aggregated inputs to human operators, to active systems that enable programmed control. In this paper, we propose Trumpet, an event monitoring system that leverages CPU resources and end-host programmability, to monitor every packet and report events at millisecond timescales. Trumpet users can express many *network-wide events*, and the system efficiently detects these events using *triggers* at end-hosts. Using careful design, Trumpet can evaluate triggers by inspecting every packet at full line rate even on future generations of NICs, scale to thousands of triggers per end-host while bounding packet processing delay to a few microseconds, and report events to a controller within 10 milliseconds, even in the presence of attacks. We demonstrate these properties using an implementation of Trumpet, and also show that it allows operators to describe new network events such as detecting correlated bursts and loss, identifying the root cause of transient congestion, and detecting short-term anomalies at the scale of a data center tenant. Masoud Moshref, Minlan Yu, Ramesh Govindan, Amin Vahdat |
SIGCOMM | 2 |
| 2016 | Guest Editors' Introduction: Special Issue on Management of Softwarized NetworksabstractThere is currently a strong interest in both industry and academia in the softwarization of telecommunication networks and cloud computing infrastructures. This evolution is enabled by three paradigms. First, Software-Defined Networking (SDN) allows network control to be separated from the forwarding plane and allows for a flexible management of the network resources. Second, Network Virtualization (NV) brings virtualization concepts to the network, similar to cloud computing, which was enabled by virtualization of servers. Third, Network Function Virtualization (NFV) focuses on virtualization of software-based network functions. Instead of installing and managing dedicated hardware devices for these functions, they are implemented as software components and deployed on commodity hardware infrastructures, in most cases operated by a network operator or cloud infrastructure provider. Service Function Chaining (SFC) consists of building services using virtual network functions (VNFs). These three paradigms are synergetic and reinforce each other when used together. Several initial SDN, NV and NFV deployments are already operational in providers’ networks. Filip De Turck, Prosper Chemouil, Raouf Boutaba, Minlan Yu, Christian Esteve Rothenberg, Kohei Shiomoto |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2015 | Scheduling jobs across geo-distributed datacentersabstractWith growing data volumes generated and stored across geo-distributed datacenters, it is becoming increasingly inefficient to aggregate all data required for computation at a single datacenter. Instead, a recent trend is to distribute computation to take advantage of data locality, thus reducing the resource (e.g., bandwidth) costs while improving performance. In this trend, new challenges are emerging in job scheduling, which requires coordination among the datacenters as each job runs across geo-distributed sites. In this paper, we propose novel job scheduling algorithms that coordinate job scheduling across datacenters with low overhead, while achieving near-optimal performance. Our extensive simulation study with realistic job traces shows that the proposed scheduling algorithms result in up to 50% improvement in average job completion time over the Shortest Remaining Processing Time (SRPT) based approaches. Chien-Chun Hung, Leana Golubchik, Minlan Yu |
SoCC | 3 |
| 2015 | SCREAM: sketch resource allocation for software-defined measurementabstractSoftware-defined networks can enable a variety of concurrent, dynamically instantiated, measurement tasks, that provide fine-grain visibility into network traffic. Recently, there have been many proposals for using sketches for network measurement. However, sketches in hardware switches use constrained resources such as SRAM memory, and the accuracy of measurement tasks is a function of the resources devoted to them on each switch. This paper presents SCREAM, a system for allocating resources to sketch-based measurement tasks that ensures a user-specified minimum accuracy. SCREAM estimates the instantaneous accuracy of tasks so as to dynamically adapt the allocated resources for each task. Thus, by finding the right amount of resources for each task on each switch and correctly merging sketches at the controller, SCREAM can multiplex resources among network-wide measurement tasks. Simulations with three measurement tasks (heavy hitter, hierarchical heavy hitter, and super source/destination detection) show that SCREAM can support more measurement tasks with higher accuracy than existing approaches. Masoud Moshref, Minlan Yu, Ramesh Govindan, Amin Vahdat |
CoNEXT | 2 |
| 2015 | Re-evaluating Measurement Algorithms in SoftwareabstractWith the advancement of multicore servers, there is a new trend of moving network functions to software servers. Measurement is critical to most network functions as it not only helps the operators understand the network usage and detect anomalies, but also produces feedback to the control loop in management tasks such as load balancing and traffic engineering. Traditional researches on measurement algorithms mainly focus on reducing the memory usage leveraging the fact that measurement can sustain bounded inaccuracy. In this study, we re-evaluate these algorithms on software servers in order to understand their tradeoffs of accuracy and performance. We observe that simple hash tables work better than more advanced measurement algorithms for a variety of measurement scenarios. This is because with better cache design in modern servers and the skewness in the access patterns of measurement tasks, the memory usage of measurement tasks is largely irrelevant to the packet processing performance. Omid Alipourfard, Masoud Moshref, Minlan Yu |
HotNets | 3 |
| 2015 | The Dark Menace: Characterizing Network-based Attacks in the CloudabstractAs the cloud computing market continues to grow, the cloud platform is becoming an attractive target for attackers to disrupt services and steal data, and to compromise resources to launch attacks. In this paper, using three months of NetFlow data in 2013 from a large cloud provider, we present the first large-scale characterization of inbound attacks towards the cloud and outbound attacks from the cloud. We investigate nine types of attacks ranging from network-level attacks such as DDoS to application-level attacks such as SQL injection and spam. Our analysis covers the complexity, intensity, duration, and distribution of these attacks, highlighting the key challenges in defending against attacks in the cloud. By characterizing the diversity of cloud attacks, we aim to motivate the research community towards improving future security solutions for cloud systems. Rui Miao 0001, Rahul Potharaju, Minlan Yu, Navendu Jain |
Internet Measurement Conference | 3 |
| 2015 | Rapier: Integrating routing and scheduling for coflow-aware data center networksabstractIn the data flow models of today's data center applications such as MapReduce, Spark and Dryad, multiple flows can comprise a coflow group semantically. Only completing all flows in a coflow is meaningful to an application. To optimize application performance, routing and scheduling must be jointly considered at the level of a coflow rather than individual flows. However, prior solutions have significant limitation: they only consider scheduling, which is insufficient. To this end, we present Rapier, a coflow-aware network optimization framework that seamlessly integrates routing and scheduling for better application performance. Using a small-scale testbed implementation and large-scale simulations, we demonstrate that Rapier significantly reduces the average coflow completion time (CCT) by up to 79.30% compared to the state-of-the-art scheduling-only solution, and it is readily implementable with existing commodity switches. Yangming Zhao, Kai Chen 0005, Wei Bai 0001, Minlan Yu, Chen Tian 0001, Yanhui Geng, Yiming Zhang 0003, Dan Li 0001, Sheng Wang 0006 |
INFOCOM | 4 |
| 2015 | Hopper: Decentralized Speculation-aware Cluster Scheduling at ScaleabstractAs clusters continue to grow in size and complexity, providing scalable and predictable performance is an increasingly important challenge. A crucial roadblock to achieving predictable performance is stragglers, i.e., tasks that take significantly longer than expected to run. At this point, speculative execution has been widely adopted to mitigate the impact of stragglers. However, speculation mechanisms are designed and operated independently of job scheduling when, in fact, scheduling a speculative copy of a task has a direct impact on the resources available for other jobs. In this work, we present Hopper, a job scheduler that is speculation-aware, i.e., that integrates the tradeoffs associated with speculation into job scheduling decisions. We implement both centralized and decentralized prototypes of the Hopper scheduler and show that 50% (66%) improvements over state-of-the-art centralized (decentralized) schedulers and speculation strategies can be achieved through the coordination of scheduling and speculation. Xiaoqi Ren, Ganesh Ananthanarayanan, Adam Wierman, Minlan Yu |
SIGCOMM | 4 |
| 2015 | Condor: Better Topologies Through Declarative DesignabstractThe design space for large, multipath datacenter networks is large and complex, and no one design fits all purposes. Network architects must trade off many criteria to design cost-effective, reliable, and maintainable networks, and typically cannot explore much of the design space. We present Condor, our approach to enabling a rapid, efficient design cycle. Condor allows architects to express their requirements as constraints via a Topology Description Language (TDL), rather than having to directly specify network structures. Condor then uses constraint-based synthesis to rapidly generate candidate topologies, which can be analyzed against multiple criteria. We show that TDL supports concise descriptions of topologies such as fat-trees, BCube, and DCell; that we can generate known and novel variants of fat-trees with simple changes to a TDL file; and that we can synthesize large topologies in tens of seconds. We also show that Condor supports the daunting task of designing multi-phase network expansions that can be carried out on live networks. Brandon Schlinker, Radhika Niranjan Mysore, Jeffrey C. Mogul, Amin Vahdat, Minlan Yu, Ethan Katz-Bassett, Michael Rubin |
SIGCOMM | 6 |
| 2015 | Joint VM placement and topology optimization for traffic scalability in dynamic datacenter networks
Yangming Zhao, Yifan Huang 0001, Kai Chen 0005, Minlan Yu, Sheng Wang 0006, Dongsheng Li 0001 |
Comput. Networks | 4 |
| 2014 | Programmable measurement architectureabstractMeasurement is at least half of network management. Many data centers require huge capital investments to build larger networks with higher link speeds; yet provide surprisingly little visibility into the network and traffic. Switch vendors often treat measurement as a second-class citizen, devoting most resources to control functions. Operators have limited control over what (not) to measure, and thus have to integrate incomplete measurement data from individual devices. Inspired by software-defined networking, we propose to design and build a new programmable measurement architecture that bridges the gap between operator's measurement requirements and device capabilities. We allow operators to flexibly program queries about the network state they need, provide generic and efficient primitives at many devices (hosts, switches, and reconfigurable devices), and automatically match the queries with the primitives. Our solutions have gained significant interests from both production data center operators and programmable switch vendors. Minlan Yu |
ANCS | 1 |
| 2014 | Tango: Simplifying SDN Control with Automatic Switch Property Inference, Abstraction, and OptimizationabstractA major benefit of software-defined networking (SDN) over traditional networking is simpler and easier control of network devices. The diversity of SDN switch implementation properties, which include both diverse switch hardware capabilities and diverse control-plane software behaviors, however, can make it difficult to understand and/or to control the switches in an SDN network. In this paper, we present Tango, a novel framework to explore the issues of understanding and optimization of SDN control, in the presence of switch diversity. The basic idea of Tango is novel, simple, and yet quite powerful. In particular, different from all previous SDN control systems, which either ignore switch diversity or depend on that switches can and will report diverse switch implementation properties, Tango introduces a novel, proactive probing engine that infers key switch capabilities and behaviors, according to a well-structured set of Tango patterns, where a Tango pattern consists of a sequence of standard OpenFlow commands and a corresponding data traffic pattern. Utilizing the inference results from Tango patterns and additional application API hints, Tango conducts automatic switch control optimization, despite diverse switch capabilities and behaviors. Evaluating Tango on both hardware switches and emulated software switches, we show that Tango can infer flow table sizes, which are key switch implementation properties, within less than 5% of actual values, despite diverse switch caching algorithms, using a probing algorithm that is asymptotically optimal in terms of probing overhead. We demonstrate cases where routing and scheduling optimizations based on Tango improves the rule installation time by up to 70% in our hardware switch testbed. Aggelos Lazaris, Daniel Tahara, Xin Huang 0008, Li Erran Li, Andreas Voellmy, Yang Richard Yang, Minlan Yu |
CoNEXT | 7 |
| 2014 | DIBS: just-in-time congestion mitigation for data centersabstractData centers must support a range of workloads with differing demands. Although existing approaches handle routine traffic smoothly, intense hotspots--even if ephemeral--cause excessive packet loss and severely degrade performance. This loss occurs even though congestion is typically highly localized, with spare buffer capacity at nearby switches. In this paper, we argue that switches should share buffer capacity to effectively handle this spot congestion without the monetary hit of deploying large buffers at individual switches. Specifically, we present detour-induced buffer sharing (DIBS), a mechanism that achieves a near lossless network without requiring additional buffers at individual switches. Using DIBS, a congested switch detours packets randomly to neighboring switches to avoid dropping the packets. We implement DIBS in hardware, on software routers in a testbed, and in simulation, and we demonstrate that it reduces the 99th percentile of delay-sensitive query completion time by up to 85%, with very little impact on other traffic. Kyriakos Zarifis, Rui Miao 0001, Matt Calder, Ethan Katz-Bassett, Minlan Yu, Jitendra Padhye |
EuroSys | 5 |
| 2014 | GRASS: Trimming Stragglers in Approximation Analytics
Ganesh Ananthanarayanan, Chien-Chun Hung, Xiaoqi Ren, Ion Stoica, Adam Wierman, Minlan Yu |
NSDI | 6 |
| 2014 | Enforcing Network-Wide Policies in the Presence of Dynamic Middlebox Actions using FlowTags
Seyed Kaveh Fayaz, Luis Chiang, Vyas Sekar, Minlan Yu, Jeffrey C. Mogul |
NSDI | 4 |
| 2014 | The Need for End-to-End Evaluation of Cloud Availability
Zi Hu, Calvin Ardi, Ethan Katz-Bassett, Harsha V. Madhyastha, John S. Heidemann, Minlan Yu |
PAM | 7 |
| 2014 | SENSS: observe and control your own traffic in the internetabstractWe propose a new software-defined security service -- SENSS -- that enables a victim network to request services from remote ISPs for traffic that carries source IPs or destination IPs from this network's address space. These services range from statistics gathering, to filtering or quality of service guarantees, to route reports or modifications. The SENSS service has very simple, yet powerful, interfaces. This enables it to handle a variety of data plane and control plane attacks, while being easily implementable in today's ISP. Through extensive evaluations on realistic traffic traces and Internet topology, we show how SENSS can be used to quickly, safely and effectively mitigate a variety of large-scale attacks that are largely unhandled today. Abdulla Alwabel, Minlan Yu, Ying Zhang 0022, Jelena Mirkovic |
SIGCOMM | 2 |
| 2014 | NIMBUS: cloud-scale attack detection and mitigationabstractNo abstract available. Rui Miao 0001, Minlan Yu, Navendu Jain |
SIGCOMM | 2 |
| 2014 | Flow-level state transition as a new switch primitive for SDNabstractNo abstract available. Masoud Moshref, Apoorv Bhargava, Adhip Gupta, Minlan Yu, Ramesh Govindan |
SIGCOMM | 4 |
| 2014 | DREAM: dynamic resource allocation for software-defined measurementabstractSoftware-defined networks can enable a variety of concurrent, dynamically instantiated, measurement tasks, that provide fine-grain visibility into network traffic. Recently, there have been many proposals to configure TCAM counters in hardware switches to monitor traffic. However, the TCAM memory at switches is fundamentally limited and the accuracy of the measurement tasks is a function of the resources devoted to them on each switch. This paper describes an adaptive measurement framework, called DREAM, that dynamically adjusts the resources devoted to each measurement task, while ensuring a user-specified level of accuracy. Since the trade-off between resource usage and accuracy can depend upon the type of tasks, their parameters, and traffic characteristics, DREAM does not assume an a priori characterization of this trade-off, but instead dynamically searches for a resource allocation that is sufficient to achieve a desired level of accuracy. A prototype implementation and simulations with three network-wide measurement tasks (heavy hitter, hierarchical heavy hitter and change detection) and diverse traffic show that DREAM can support more concurrent tasks with higher accuracy than several other alternatives. Masoud Moshref, Minlan Yu, Ramesh Govindan, Amin Vahdat |
SIGCOMM | 2 |
| 2013 | Scalable Rule Management for Data Centers
Masoud Moshref, Minlan Yu, Abhishek B. Sharma, Ramesh Govindan |
NSDI | 2 |
| 2013 | Software Defined Traffic Measurement with OpenSketch
Minlan Yu, Lavanya Jose, Rui Miao 0001 |
NSDI | 1 |
| 2013 | Don't drop, detour!abstractToday's data centers must support a range of workloads with different demands. While existing approaches handle routine traffic smoothly, ephemeral but intense hotspots cause excessive packet loss and severely degrade performance. This loss occurs even though the congestion is typically highly localized, with spare buffer capacity available at nearby switches. Matt Calder, Rui Miao 0001, Kyriakos Zarifis, Ethan Katz-Bassett, Minlan Yu, Jitendra Padhye |
SIGCOMM | 5 |
| 2013 | SIMPLE-fying middlebox policy enforcement using SDNabstractNetworks today rely on middleboxes to provide critical performance, security, and policy compliance capabilities. Achieving these benefits and ensuring that the traffic is directed through the desired sequence of middleboxes requires significant manual effort and operator expertise. In this respect, Software-Defined Networking (SDN) offers a promising alternative. Middleboxes, however, introduce new aspects (e.g., policy composition, resource management, packet modifications) that fall outside the purvey of traditional L2/L3 functions that SDN supports (e.g., access control or routing). Zafar Ayyub Qazi, Cheng-Chun Tu, Luis Chiang, Rui Miao 0001, Vyas Sekar, Minlan Yu |
SIGCOMM | 6 |
| 2012 | Tradeoffs in CDN designs for throughput oriented trafficabstractInternet delivery infrastructures are traditionally optimized for low-latency traffic, such as the Web traffic. However, in recent years we are witnessing a massive growth of throughput-oriented applications, such as video streaming. These applications introduce new tradeoffs and design choices for content delivery networks (CDNs). In this paper, we focus on understanding two key design choices: (1) What is the impact of the number of CDN's peering points and server locations on its aggregate throughput and operating costs? (2) How much can ISP-CDNs benefit from using path selection to maximize its aggregate throughput compared to other CDNs who only have control at the edge? Answering these questions is challenging because content distribution involves a complex ecosystem consisting of many parties (clients, CDNs, ISPs) and depends on various settings which differ across places and over time. We introduce a simple model to illustrate and quantify the essential tradeoffs in CDN designs. Using extensive analysis over a variety of network topologies (with varying numbers of CDN peering points and server locations), operating cost models, and client video streaming traces, we observe that: (1) Doubling the number of peering points roughly doubles the aggregate throughput over a wide range of values and network topologies. In contrast, optimal path selection improves the CDN aggregate throughput by less than 70\%, and in many cases by as little as a few percents. (2) Keeping the number of peering points constant, but reducing the number of location (data centers) at which the CDN is deployed can significantly reduce operating costs. Minlan Yu, Wenjie Jiang 0001, Haoyuan Li 0001, Ion Stoica |
CoNEXT | 1 |
| 2012 | Latency Equalization as a New Network Service PrimitiveabstractMultiparty interactive network applications such as teleconferencing, network gaming, and online trading are gaining popularity. In addition to end-to-end latency bounds, these applications require that the delay difference among multiple clients of the service is minimized for a good interactive experience. We propose a Latency EQualization (LEQ) service, which equalizes the perceived latency for all clients participating in an interactive network application. To effectively implement the proposed LEQ service, network support is essential. The LEQ architecture uses a few routers in the network as hubs to redirect packets of interactive applications along paths with similar end-to-end delay. We first formulate the hub selection problem, prove its NP-hardness, and provide a greedy algorithm to solve it. Through extensive simulations, we show that our LEQ architecture significantly reduces delay difference under different optimization criteria that allow or do not allow compromising the per-user end-to-end delay. Our LEQ service is incrementally deployable in today's networks, requiring just software modifications to edge routers. Minlan Yu, Marina Thottan, Li Erran Li |
IEEE/ACM Trans. Netw. | 1 |
| 2011 | Profiling Network Performance for Multi-tier Data Center Applications
Minlan Yu, Albert G. Greenberg, David A. Maltz, Jennifer Rexford, Srikanth Kandula, Changhoon Kim |
NSDI | 1 |
| 2010 | CloudPolice: taking access control out of the networkabstractCloud computing environments impose new challenges on access control techniques due to multi-tenancy, the growing scale and dynamicity of hosts within the cloud infrastructure, and the increasing diversity of cloud network architectures. The majority of existing access control techniques were originally designed for enterprise environments that do not share these challenges and, as such, are poorly suited for cloud environments. In this paper, we argue that it is both sufficient and advantageous to implement access control only within the hypervisors at the end-hosts. We thus propose Cloud-Police, a system that implements a hypervisor-based access control mechanism. We argue that, not only can CloudPolice support more sophisticated access control policies, it can do so in a manner that is simpler, more scalable and more robust than existing network-based techniques. Lucian Popa 0002, Minlan Yu, Steven Y. Ko, Sylvia Ratnasamy, Ion Stoica |
HotNets | 2 |
| 2010 | Scalable flow-based networking with DIFANEabstractIdeally, enterprise administrators could specify fine-grain policies that drive how the underlying switches forward, drop, and measure traffic. However, existing techniques for flow-based networking rely too heavily on centralized controller software that installs rules reactively, based on the first packet of each flow. In this paper, we propose DIFANE, a scalable and efficient solution that keeps all traffic in the data plane by selectively directing packets through intermediate switches that store the necessary rules. DIFANE relegates the controller to the simpler task of partitioning these rules over the switches. DIFANE can be readily implemented with commodity switch hardware, since all data-plane functions can be expressed in terms of wildcard rules that perform simple actions on matching packets. Experiments with our prototype on Click-based OpenFlow switches show that DIFANE scales to larger networks with richer policies. Minlan Yu, Jennifer Rexford, Michael J. Freedman, Jia Wang 0001 |
SIGCOMM | 1 |
| 2009 | Virtually eliminating router bugsabstractSoftware bugs in routers lead to network outages, security vulnerabilities, and other unexpected behavior. Rather than simply crashing the router, bugs can violate protocol semantics, rendering traditional failure detection and recovery techniques ineffective. Handling router bugs is an increasingly important problem as new applications demand higher availability, and networks become better at dealing with traditional failures. In this paper, we tailor software and data diversity (SDD) to the unique properties of routing protocols, so as to avoid buggy behavior at run time. Our bug-tolerant router executes multiple diverse instances of routing software, and uses voting to determine the output to publish to the forwarding table, or to advertise to neighbors. We design and implement a router hypervisor that makes this parallelism transparent to other routers, handles fault detection and booting of new router instances, and performs voting in the presence of routing-protocol dynamics, without needing to modify software of the diverse instances. Experiments with BGP message traces and open-source software running on our Linux-based router hypervisor demonstrate that our solution scales to large networks and efficiently masks buggy behavior. Eric Keller, Minlan Yu, Matthew Caesar 0001, Jennifer Rexford |
CoNEXT | 2 |
| 2009 | BUFFALO: bloom filter forwarding architecture for large organizationsabstractIn enterprise and data center networks, the scalability of the data plane becomes increasingly challenging as forwarding tables and link speeds grow. Simply building switches with larger amounts of faster memory is not appealing, since high-speed memory is both expensive and power hungry. Implementing hash tables in SRAM is not appealing either because it requires significant overprovisioning to ensure that all forwarding table entries fit. Instead, we propose the BUFFALO architecture, which uses a small SRAM to store one Bloom filter of the addresses associated with each outgoing link. We provide a practical switch design leveraging flat addresses and shortest-path routing. BUFFALO gracefully handles false positives without reducing the packet-forwarding rate, while guaranteeing that packets reach their destinations with bounded stretch with high probability. We tune the sizes of Bloom filters to minimize false positives for a given memory size. We also handle routing changes and dynamically adjust Bloom filter sizes using counting Bloom filters in slow memory. Our extensive analysis, simulation, and prototype implementation in kernel-level Click show that BUFFALO significantly reduces memory cost, increases the scalability of the data plane, and improves packet-forwarding performance. Minlan Yu, Alex Fabrikant, Jennifer Rexford |
CoNEXT | 1 |