VLDB 2026 Research / reviewers in the wild / expert
Mosharaf Chowdhury
dblp:42/1518 · also N. M. Mosharaf Kabir Chowdhury
· DBLP profile ↗
59ranked-venue papers
11as first author
28since 2021 · last 2026
0000-0003-0884-6740ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 32 · 10 first-author · 9 since 2021Software engineering, systems software and programming languages · 11 · 7 since 2021Systems, architecture and hardware · 10 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action MemoryabstractAutonomous web agents powered by large language models (LLMs) show strong potential for performing goal-oriented tasks such as information retrieval, report generation, and online transactions.These agents mark a key step toward practical embodied reasoning in open web environments.However, existing approaches remain limited in reasoning depth and efficiency: vanilla linear methods fail at multistep reasoning and lack effective backtracking, while other search strategies are coarse-grained and computationally costly.We introduce Branch-and-Browse, a fine-grained web agent framework that unifies structured reasoningacting, contextual memory, and efficient execution.It (i) employs explicit subtask management with tree-structured exploration for controllable multi-branch reasoning, (ii) bootstraps exploration through efficient web state replay with background reasoning, and (iii) leverages a page action memory to share explored actions within and across sessions.On the WebArena benchmark, Branch-and-Browse achieves a task success rate of 35.8% and reduces execution time by up to 40.4% relative to state-of-the-art methods.These results demonstrate that Branch-and-Browse is a reliable and efficient framework for LLM-based web agents.Code is available at https://github.com/ SymbioticLab/Branch-and-Browse. Shiqi He, Yue Cui 0001, Yaliang Li, Bolin Ding, Mosharaf Chowdhury |
ACL (1) | 6 |
| 2026 | TetriServe: Efficiently Serving Mixed DiT Workloads
Runyu Lu, Shiqi He, Wenxuan Tan, Shenggui Li, Jeff J. Ma, Ang Chen 0001, Mosharaf Chowdhury |
ASPLOS (2) | 8 |
| 2025 | DPack: Efficiency-Oriented Privacy Budget SchedulingabstractMachine learning (ML) models can leak information about users, and differential privacy (DP) provides a rigorous way to bound that leakage under a given budget. This DP budget can be regarded as a new type of computing resource in workloads of multiple ML models training on user data. Once it is used, the DP budget is forever consumed. Therefore, it is crucial to allocate it most efficiently to train as many models as possible. This paper presents a scheduler for the privacy resources that optimizes for efficiency. We formulate privacy scheduling as a new type of multidimensional knapsack problem, called privacy knapsack, which maximizes DP budget efficiency. We show that privacy knapsack is NP-hard, hence practical algorithms are necessarily approximate. We develop an approximation algorithm for privacy knapsack, DPack, and evaluate it on microbenchmarks and on a new, synthetic private-ML workload we developed from the Alibaba ML cluster trace. We show that DPack: (1) often approaches the efficiency-optimal schedule, (2) consistently schedules more tasks compared to a state-of-the-art privacy scheduling algorithm that focused on fairness instead of efficiency (1.3-1.7× in Alibaba, 1.0-2.6X in microbenchmarks), but (3) sacrifices some level of fairness for efficiency. Using DPack, DP ML operators should be able to train more models on the same amount of user data while offering the same privacy guarantee to their users. Pierre Tholoniat, Kelly Kostopoulou, Mosharaf Chowdhury, Asaf Cidon, Roxana Geambasu, Mathias Lécuyer |
EuroSys | 3 |
| 2025 | Remote Direct Code ExecutionabstractWe propose remote direct code execution (RDX), which elevates the power of RDMA from memory access to code execution. We target runtime extension frameworks such as Wasm filters, BPF programs, and UDF functions, where RDX enables an agentless architecture that unlocks capabilities such as fast extension injection, update consistency guarantees, and minimal resource contention. We outline the roadmap for RDX around a new CodeFlow abstraction, encompassing programming remote extensions, exposing management stubs, remotely validating and JIT compiling code, seamlessly linking code to local context, managing remote extension state, and synchronizing code to targets. The case studies and initial results demonstrate the feasibility of RDX and its potential to spark the next wave of RDMA innovations. Yibo Huang 0005, Yiming Qiu 0001, Daqian Ding, Patrick Tser Jern Kon, Yiwen Zhang 0008, Yuzhou Mao, Archit Bhatnagar, Mosharaf Chowdhury, Srini Devadas, Jiarong Xing, Ang Chen 0001 |
HotNets | 8 |
| 2025 | The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and OptimizationabstractAs the adoption of Generative AI in real-world services grow explosively, energy has emerged as a critical bottleneck resource. However, energy remains a metric that is often overlooked, under-explored, or poorly understood in the context of building ML systems. We present the ML.ENERGY Benchmark, a benchmark suite and tool for measuring inference energy consumption under realistic service environments, and the corresponding ML.ENERGY Leaderboard, which have served as a valuable resource for those hoping to understand and optimize the energy consumption of their generative AI services. In this paper, we explain four key design principles for benchmarking ML energy we have acquired over time, and then describe how they are implemented in the ML.ENERGY Benchmark. We then highlight results from the early 2025 iteration of the benchmark, including energy measurements of 40 widely used model architectures across 6 different tasks, case studies of how ML design choices impact energy consumption, and how automated optimization recommendations can lead to significant (sometimes more than 40%) energy savings without changing what is being computed by the model. The ML.ENERGY Benchmark is open-source and can be easily extended to various customized models and application scenarios. Jae-Won Chung, Jeff J. Ma, Oh Jun Kweon, Yuxuan Xia, Zhiyu Wu, Mosharaf Chowdhury |
NeurIPS | 8 |
| 2024 | INFA-FinOps for Cloud Data IntegrationabstractOver the past decade, businesses have migrated to the cloud for its simplicity, elasticity, and resilience. Cloud ecosystems offer a variety of computing and storage options, enabling customers to choose configurations that maximize productivity. However, determining the right configuration to minimize cost while maximizing performance is challenging, as workloads vary and cloud offerings constantly evolve. Many businesses are overwhelmed with choice overload and often end up making suboptimal choices that lead to inflated cloud spending and/or poor performance.In this paper, we describe INFA-FinOps, an automated system that helps Informatica customers strike a balance between cost efficiency and meeting SLAs for Informatica Advanced Data Integration (aka CDI-E) workloads. We first describe common workload patterns observed in CDI-E customers and show how INFA-FinOps selects optimal cloud resources and configurations for each workload, adjusting them as workloads and cloud ecosystems change. It also makes recommendations for actions that require user review or input. Finally, we present performance benchmarks on various enterprise use cases and conclude with lessons learned and potential future enhancements. Atam Prakash Agrawal, Anant Mittal, Shivangi Srivastava, Michael Brevard, Valentin Moskovich, Mosharaf Chowdhury |
IEEE Big Data | 6 |
| 2024 | IaC-Eval: A Code Generation Benchmark for Cloud Infrastructure-as-Code ProgramsabstractInfrastructure-as-Code (IaC), an important component of cloud computing, allows the definition of cloud infrastructure in high-level programs. However, developing IaC programs is challenging, complicated by factors that include the burgeoning complexity of the cloud ecosystem (e.g., diversity of cloud services and workloads), and the relative scarcity of IaC-specific code examples and public repositories. While large language models (LLMs) have shown promise in general code generation and could potentially aid in IaC development, no benchmarks currently exist for evaluating their ability to generate IaC code. We present IaC-Eval, a first step in this research direction. IaC-Eval's dataset includes 458 human-curated scenarios covering a wide range of popular AWS services, at varying difficulty levels. Each scenario mainly comprises a natural language IaC problem description and an infrastructure intent specification. The former is fed as user input to the LLM, while the latter is a general notion used to verify if the generated IaC program conforms to the user's intent; by making explicit the problem's requirements that can encompass various cloud services, resources and internal infrastructure details. Our in-depth evaluation shows that contemporary LLMs perform poorly on IaC-Eval, with the top-performing model, GPT-4, obtaining a pass@1 accuracy of 19.36%. In contrast, it scores 86.6% on EvalPlus, a popular Python code generation benchmark, highlighting a need for advancements in this domain. We open-source the IaC-Eval dataset and evaluation framework at https://github.com/autoiac-project/iac-eval to enable future research on LLM-based IaC code generation. Patrick Tser Jern Kon, Yiming Qiu 0001, Weijun Fan, Owen Park, George Elengikal, Yuxin Kang, Ang Chen 0001, Mosharaf Chowdhury, Myungjin Lee, Xinyu Wang 0006 |
NeurIPS | 12 |
| 2024 | Vulcan: Automatic Query Planning for Live ML Analytics
Yiwen Zhang 0008, Xumiao Zhang, Ganesh Ananthanarayanan, Anand Padmanabha Iyer, Yuanchao Shu, Paramvir Bahl, Z. Morley Mao, Mosharaf Chowdhury |
NSDI | 8 |
| 2024 | Managing Memory Tiers with CXL in Virtualized Environments
Yuhong Zhong, Daniel S. Berger, Carl A. Waldspurger, Ryan Wee, Ishwar Agarwal, Rajat Agarwal, Frank Hady, Karthik Kumar, Mark D. Hill, Mosharaf Chowdhury, Asaf Cidon |
OSDI | 10 |
| 2024 | Reducing Energy Bloat in Large Model TrainingabstractTraining large AI models on numerous GPUs consumes a massive amount of energy, making power delivery one of the largest limiting factors in building and operating datacenters for AI workloads. However, we observe that not all energy consumed during training directly contributes to end-to-end throughput; a significant portion can be removed without slowing down training. We call this portion energy bloat. Jae-Won Chung, Yile Gu, Insu Jang, Luoxi Meng, Nikhil Bansal 0001, Mosharaf Chowdhury |
SOSP | 6 |
| 2024 | Fed-ensemble: Ensemble Models in Federated Learning for Improved Generalization and Uncertainty QuantificationabstractThe increase in the computational power of edge devices has opened up the possibility of processing some of the data at the edge and distributing model learning. This paradigm is often called federated learning (FL), where edge devices exploit their local computational resources to train models collaboratively. Though FL has seen recent success, it is unclear how to characterize uncertainties in FL predictions. In this paper, we proposeFed-ensemble: a simple approach that brings model ensembling to FL. Instead of aggregating local models to update a single global model,Fed-ensembleuses random permutations to update a group of$K$models and then obtains predictions through model averaging.Fed-ensemblecan be readily utilized within established FL methods and does not impose a computational overhead compared with single-model methods. Empirical results show that our model has superior performance over several FL algorithms on a wide range of data sets and excels in heterogeneous settings often encountered in FL applications. Also, by carefully choosing client-dependent weights in the inference stage,Fed-ensemblebecomes personalized and yields even better performance. Theoretically, we show that predictions on new data from all$K$models belong to the same predictive posterior distribution under a neural tangent kernel regime. This result, in turn, sheds light on the generalization advantages of model averaging and justifies the uncertainty quantification capability. We also illustrate thatFed-ensemblehas an elegant Bayesian interpretation.Note to Practitioners—provides an algorithm that extracts a set of$K$solutions without imposing any additional communication overhead in FL. Given multiple solutions,Fed-ensemblecan be exploited to personalize inference as well as quantify uncertainty. Such capabilities may be beneficial within multiple practical systems that require uncertainty-aware decision-making. Further,Fed-ensemblemay be useful for model validation and hypothesis testing. Naichen Shi, Fan Lai 0001, Raed Kontar, Mosharaf Chowdhury |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Pyxis: Scheduling Mixed Tasks in Disaggregated DatacentersabstractDisaggregating compute from storage is an emerging trend in cloud computing. Effectively utilizing resources in both compute and storage pool is the key to high performance. The state-of-the-art scheduler provides optimal scheduling decisions for workloads with homogeneous tasks. However, cloud applications often generate a mix of tasks with diverse compute and IO characteristics, resulting in sub-optimal performance for existing solutions. We present Pyxis, a system that provides optimal scheduling decisions for mixed workloads in disaggregated datacenters with theoretical guarantees. Pyxis is capable of maximizing overall throughput while meeting latency SLOs. Pyxis decouples the scheduling of different tasks. Our insight is that the optimal solution has an “all-or-nothing” structure that can be captured by a singleturning pointin the spectrum of tasks. Based on task characteristics, the turning point partitions the tasks either all to storage nodes or all to compute nodes (none to storage nodes). We theoretically prove that the optimal solution has such a structure, and design an online algorithm with sub-second convergence. We implement a prototype of Pyxis. Experiments on CloudLab with various synthetic and application workloads show that Pyxis improves the throughput by 3–21× over the state-of-the-art solution. Chao Jin 0007, Mosharaf Chowdhury, Zhenming Liu, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryabstractThe increasing demand for memory in hyperscale applications has led to memory becoming a large portion of the overall datacenter spend. The emergence of coherent interfaces like CXL enables main memory expansion and offers an efficient solution to this problem. In such systems, the main memory can constitute different memory technologies with varied characteristics. In this paper, we characterize memory usage patterns of a wide range of datacenter applications across the server fleet of Meta. We, therefore, demonstrate the opportunities to offload colder pages to slower memory tiers for these applications. Without efficient memory management, however, such systems can significantly degrade performance. Hasan Al Maruf, Hao Wang 0011, Abhishek Dhanotia, Johannes Weiner, Niket Agarwal, Pallab Bhattacharya, Chris Petersen 0002, Mosharaf Chowdhury, Shobhit O. Kanaujia, Prakash Chauhan |
ASPLOS (3) | 8 |
| 2023 | Auxo: Efficient Federated Learning via Scalable Client ClusteringabstractFederated learning (FL) is an emerging machine learning (ML) paradigm that enables heterogeneous edge devices to collaboratively train ML models without revealing their raw data to a logically centralized server. However, beyond the heterogeneous device capacity, FL participants often exhibit differences in their data distributions, which are not independent and identically distributed (Non-IID). Many existing works present point solutions to address issues like slow convergence, low final accuracy, and bias in FL, all stemming from client heterogeneity. Fan Lai 0001, Yinwei Dai, Aditya Akella, Harsha V. Madhyastha, Mosharaf Chowdhury |
SoCC | 6 |
| 2023 | Egeria: Efficient DNN Training with Knowledge-Guided Layer FreezingabstractTraining deep neural networks (DNNs) is time-consuming. While most existing solutions try to overlap/schedule computation and communication for efficient training, this paper goes one step further by skipping computing and communication through DNN layer freezing. Our key insight is that the training progress of internal DNN layers differs significantly, and front layers often become well-trained much earlier than deep layers. To explore this, we first introduce the notion of training plasticity to quantify the training progress of internal DNN layers. Then we design Egeria, a knowledge-guided DNN training system that employs semantic knowledge from a reference model to accurately evaluate individual layers' training plasticity and safely freeze the converged ones, saving their corresponding backward computation and communication. Our reference model is generated on the fly using quantization techniques and runs forward operations asynchronously on available CPUs to minimize the overhead. In addition, Egeria caches the intermediate outputs of the frozen layers with prefetching to further skip the forward computation. Our implementation and testbed experiments with popular vision and language models show that Egeria achieves 19%-43% training speedup w.r.t. the state-of-the-art without sacrificing accuracy. Decang Sun, Kai Chen 0005, Fan Lai 0001, Mosharaf Chowdhury |
EuroSys | 5 |
| 2023 | Simplifying Cloud Management with Cloudless ComputingabstractCloud computing has transformed the IT industry, but managing cloud infrastructures remains a difficult task. We make a case for putting today's management practices, known as "Infrastructure-as-Code," on a firmer ground via a principled design. We call this end goal Cloudless Computing: it aims to simplify cloud infrastructure management tasks by supporting them "as-a-service," analogous to serverless computing that relieves users of the burden of managing server instances. By assisting tenants with these tasks, cloud resources will be presented to their users more readily without the undue burden of complex control. We describe the research problems by examining the typical lifecycle of today's cloud infrastructure management, and identify places where a cloudless approach will advance the state of the art. Yiming Qiu 0001, Patrick Tser Jern Kon, Jiarong Xing, Yibo Huang 0005, Xinyu Wang 0006, Peng Huang 0005, Mosharaf Chowdhury, Ang Chen 0001 |
HotNets | 8 |
| 2023 | ModelKeeper: Accelerating DNN Training via Automated Training Warmup
Fan Lai 0001, Yinwei Dai, Harsha V. Madhyastha, Mosharaf Chowdhury |
NSDI | 4 |
| 2023 | Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training
Jae-Won Chung, Mosharaf Chowdhury |
NSDI | 3 |
| 2023 | AdaEmbed: Adaptive Embedding for Large-Scale Recommendation Models
Fan Lai 0001, Wei Zhang 0044, William Tsai, Xiaohan Wei, Yuxi Hu 0001, Sabin Devkota, Jongsoo Park, Zeliang Chen, Ellie Wen, Paul Rivera, Chun-cheng Jason Chen, Mosharaf Chowdhury |
OSDI | 16 |
| 2023 | Oobleck: Resilient Distributed Training of Large Models Using Pipeline TemplatesabstractOobleck enables resilient distributed training of large DNN models with guaranteed fault tolerance. It takes a planning-execution co-design approach, where it first generates a set of heterogeneous pipeline templates and instantiates at least f + 1 logically equivalent pipeline replicas to tolerate any f simultaneous failures. During execution, it relies on already-replicated model states across the replicas to provide fast recovery. Oobleck provably guarantees that some combination of the initially created pipeline templates can be used to cover all available resources after f or fewer simultaneous failures, thereby avoiding resource idling at all times. Evaluation on large DNN models with billions of parameters shows that Oobleck provides consistently high throughput, and it outperforms state-of-the-art fault tolerance solutions like Bamboo and Varuna by up to 13.9×. Insu Jang, Zhenning Yang, Zhen Zhang 0063, Xin Jin 0008, Mosharaf Chowdhury |
SOSP | 5 |
| 2022 | Hydra : Resilient and Highly Available Remote Memory
Youngmoon Lee, Hasan Al Maruf, Mosharaf Chowdhury, Asaf Cidon, Kang G. Shin |
FAST | 3 |
| 2022 | FedScale: Benchmarking Model and System Performance of Federated Learning at ScaleabstractWe present FedScale, a federated learning (FL) benchmarking suite with realistic datasets and a scalable runtime to enable reproducible FL research. FedScale datasets encompass a wide range of critical FL tasks, ranging from image classification and object detection to language modeling and speech recognition. Each dataset comes with a unified evaluation protocol using real-world data splits and evaluation metrics. To reproduce realistic FL behavior, FedScale contains a scalable and extensible runtime. It provides high-level APIs to implement FL algorithms, deploy them at scale across diverse hardware and software backends, and evaluate them at scale, all with minimal developer efforts. We combine the two to perform systematic benchmarking experiments and highlight potential opportunities for heterogeneity-aware co-optimizations in FL. FedScale is open-source and actively maintained by contributors from different institutions at http://fedscale.ai. We welcome feedback and contributions from the community. Fan Lai 0001, Yinwei Dai, Sanjay Sri Vallabh Singapuram, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury |
ICML | 7 |
| 2022 | Justitia: Software Multi-Tenancy in Hardware Kernel-Bypass Networks
Yiwen Zhang 0008, Brent E. Stephens, Mosharaf Chowdhury |
NSDI | 4 |
| 2022 | Aequitas: admission control for performance-critical RPCs in datacentersabstractWith the increasing popularity of disaggregated storage and microservice architectures, high fan-out and fan-in Remote Procedure Calls (RPCs) now generate most of the traffic in modern datacenters. While the network plays a crucial role in RPC performance, traditional traffic classification categories cannot sufficiently capture their importance due to wide variations in RPC characteristics. As a result, meeting service-level objectives (SLOs), especially for performance-critical (PC) RPCs, remains challenging. Yiwen Zhang 0008, Gautam Kumar 0001, Nandita Dukkipati, Xian Wu 0001, Priyaranjan Jha, Mosharaf Chowdhury, Amin Vahdat |
SIGCOMM | 6 |
| 2022 | CDI-E: An Elastic Cloud Service for Data EngineeringabstractWe live in the gilded age of data-driven computing. With public clouds offering virtually unlimited amounts of compute and storage, enterprises collecting data about every aspect of their businesses, and advances in analytics and machine learning technologies, data driven decision making is now timely, cost-effective, and therefore, pervasive. Alas, only a handful of power users can wield today's powerful data engineering tools. For one thing, most solutions require knowledge of specific programming interfaces or libraries. Furthermore, running them requires complex configurations and knowledge of the underlying cloud for cost-effectiveness. We decided that a fundamental redesign is in order to democratize data engineering for the masses at cloud scale. The result is Informatica Cloud Data Integration - Elastic (CDI-E). Since the early 1990s, Informatica has been a pioneer and industry leader in building no-code data engineering tools. Non-experts can express complex data engineering tasks using a graphical user interface (GUI). Informatica CDI-E is built to incorporate the simplicity of GUI in the design layer with an elastic and highly scalable run time to handle data in any format without little to no user input using automated optimizations. Users upload their data to the cloud in any format and can immediately use them in conjunction with their data management and analytic tools of choice using CDI-E GUI. Implementation began in the Spring of 2017, and Informatica CDI-E has been generally available since the Summer of 2019. Today, CDI-E is used in production by a growing number of small and large enterprises to make sense of data in arbitrary formats. In this paper, we describe the architecture of Informatica CDI-E and its novel no-code data engineering interface. The paper highlights some of the key features of CDI-E: simplicity without loss in productivity and extreme elasticity. It concludes with lessons we learned and an outlook of the future. Prakash C. Das, Shivangi Srivastava, Valentin Moskovich, Anmol Chaturvedi, Anant Mittal, Yongqin Xiao, Mosharaf Chowdhury |
Proc. VLDB Endow. | 7 |
| 2021 | Ship Compute or Ship Data? Why Not Both?
Jingfeng Wu, Xin Jin 0008, Mosharaf Chowdhury |
NSDI | 4 |
| 2021 | Oort: Efficient Federated Learning via Guided Participant Selection
Fan Lai 0001, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury |
OSDI | 4 |
| 2021 | Programmable packet scheduling with a single queueabstractProgrammable packet scheduling enables scheduling algorithms to be programmed into the data plane without changing the hardware. Existing proposals either have no hardware implementations for switch ASICs or require multiple strict-priority queues. Zhuolong Yu, Chuheng Hu, Jingfeng Wu, Xiao Sun 0004, Vladimir Braverman, Mosharaf Chowdhury, Zhenhua Liu 0002, Xin Jin 0008 |
SIGCOMM | 6 |
| 2020 | AlloX: compute allocation in hybrid clustersabstractModern deep learning frameworks support a variety of hardware, including CPU, GPU, and other accelerators, to perform computation. In this paper, we study how to schedule jobs over such interchangeable resources - each with a different rate of computation - to optimize performance while providing fairness among users in a shared cluster. We demonstrate theoretically and empirically that existing solutions and their straightforward modifications perform poorly in the presence of interchangeable resources, which motivates the design and implementation of AlloX. At its core, AlloX transforms the scheduling problem into a min-cost bipartite matching problem and provides dynamic fair allocation over time. We theoretically prove its optimality in an ideal, offline setting and show empirically that it works well in the online scenario by incorporating with Kubernetes. Evaluations on a small-scale CPU-GPU hybrid cluster and large-scale simulations highlight that AlloX can reduce the average job completion time significantly (by up to 95% when the system load is high) while providing fairness and preventing starvation. Tan N. Le, Xiao Sun 0004, Mosharaf Chowdhury, Zhenhua Liu 0002 |
EuroSys | 3 |
| 2020 | Sol: Fast Distributed Computation Over Slow Networks
Fan Lai 0001, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury |
NSDI | 5 |
| 2020 | Near-Optimal Latency Versus Cost Tradeoffs in Geo-Distributed Storage
Muhammed Uluyol, Anthony Huang, Ayush Goel, Mosharaf Chowdhury, Harsha V. Madhyastha |
NSDI | 4 |
| 2020 | NetLock: Fast, Centralized Lock Management Using Programmable SwitchesabstractLock managers are widely used by distributed systems. Traditional centralized lock managers can easily support policies between multiple users using global knowledge, but they suffer from low performance. In contrast, emerging decentralized approaches are faster but cannot provide flexible policy support. Furthermore, performance in both cases is limited by the server capability. Zhuolong Yu, Yiwen Zhang 0008, Vladimir Braverman, Mosharaf Chowdhury, Xin Jin 0008 |
SIGCOMM | 4 |
| 2020 | Effectively Prefetching Remote Memory with Leap
Hasan Al Maruf, Mosharaf Chowdhury |
USENIX ATC | 2 |
| 2019 | Tiresias: A GPU Cluster Manager for Distributed Deep Learning
Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin, Yibo Zhu 0001, Myeongjae Jeon, Junjie Qian, Hongqiang Harry Liu, Chuanxiong Guo |
NSDI | 2 |
| 2019 | Near Optimal Coflow Scheduling in NetworksabstractThe coflow scheduling problem has emerged as a popular abstraction in the last few years to study data communication problems within a data center[6]. In this basic framework, each coflow has a set of communication demands and the goal is to schedule many coflows in a manner that minimizes the total weighted completion time. A coflow is said to complete when all its communication needs are met. This problem has been extremely well studied for the case of complete bipartite graphs that model a data center with full bisection bandwidth and several approximation algorithms and effective heuristics have been proposed recently[1,2,29]. In this work, we study a slightly different model of coflow scheduling in general graphs (to capture traffic between data centers [15,29]) and develop practical and efficient approximation algorithms for it. Our main result is a randomized 2 approximation algorithm for the single path and free path model, significantly improving prior work. In addition, we demonstrate via extensive experiments that the algorithm is practical, easy to implement and performs well in practice. Mosharaf Chowdhury, Samir Khuller, Manish Purohit, Sheng Yang 0005 |
SPAA | 1 |
| 2018 | Pas de deux: Shape the Circuits, and Shape the Apps too!abstractDespite continued efforts toward building high bandwidth, low cost datacenter networks with reconfigurable optical fabrics, the impact of optical networks on datacenter applications has received little attention. Given the constraints of optical networks and the semantics of datacenter applications, we believe the network-application intersection to be the next innovation hotspot. In this paper, we specifically focus on data-parallel applications for two primary reasons: they are a natural fit to exploit high bandwidth optical fabrics, and they often form structured communication patterns or coflows. Hong Zhang 0025, Kai Chen 0005, Mosharaf Chowdhury |
APNet | 3 |
| 2018 | Mitigating the Latency-Accuracy Trade-off in Mobile Data Analytics SystemsabstractAn increasing amount of mobile analytics is performed on data that is procured in a real-time fashion to make real-time decisions. Such tasks include simple reporting on streams to sophisticated model building. However, the practicality of these analyses are impeded in several domains because they are faced with a fundamental trade-off between data collection latency and analysis accuracy. In this paper, we first study this trade-off in the context of a specific domain, Cellular Radio Access Networks (RAN). We find that the trade-off can be resolved using two broad, general techniques: intelligent data grouping and task formulations that leverage domain characteristics. Based on this, we present CellScope, a system that applies a domain specific formulation and application of Multi-task Learning (MTL) to RAN performance analysis. It uses three techniques: feature engineering to transform raw data into effective features, a PCA inspired similarity metric to group data from geographically nearby base stations sharing performance commonalities, and a hybrid online-offline model for efficient model updates. Our evaluation shows that CellScope's accuracy improvements over direct application of ML range from 2.5× to 4.4× while reducing the model update overhead by up to 4.8×. We have also used CellScope to analyze an LTE network of over 2 million subscribers, where it reduced troubleshooting efforts by several magnitudes. We then apply the underlying techniques in CellScope to another domain specific problem, mobile phone energy bug diagnosis, and show that the techniques are general. Anand Padmanabha Iyer, Li Erran Li, Mosharaf Chowdhury, Ion Stoica |
MobiCom | 3 |
| 2018 | Dynamic Query Re-Planning using QOOP
Kshiteej Mahajan, Mosharaf Chowdhury, Aditya Akella, Shuchi Chawla 0001 |
OSDI | 2 |
| 2018 | Distributed Lock Management with RDMA: Decentralization without StarvationabstractLock managers are a crucial component of modern distributed systems. However, with the increasing availability of fast RDMA-enabled networks, traditional lock managers can no longer keep up with the latency and throughput requirements of modern systems. Centralized lock managers can ensure fairness and prevent starvation using global knowledge of the system, but are themselves single points of contention and failure. Consequently, they fall short in leveraging the full potential of RDMA networks. On the other hand, decentralized (RDMA-based) lock managers either completely sacrifice global knowledge to achieve higher throughput at the risk of starvation and higher tail latencies, or they resort to costly communications in order to maintain global knowledge, which can result in significantly lower throughput. Dong Young Yoon, Mosharaf Chowdhury, Barzan Mozafari |
SIGMOD Conference | 2 |
| 2017 | No!: Not Another Deep Learning FrameworkabstractIn recent years, deep learning has pervaded many areas of computing due to the confluence of an explosive growth of large-scale computing capabilities, availability of datasets, and advances in learning techniques. While this rapid growth has resulted in diverse deep learning frameworks, it has also led to inefficiencies for both the users and developers of these frameworks. Specifically, adopting useful techniques across frameworks -- both to perform learning tasks and to optimize performance -- involves significant repetitions and reinventions. Peifeng Yu, Mosharaf Chowdhury |
HotOS | 3 |
| 2017 | Efficient Memory Disaggregation with Infiniswap
Juncheng Gu, Youngmoon Lee, Yiwen Zhang 0008, Mosharaf Chowdhury, Kang G. Shin |
NSDI | 4 |
| 2017 | Resilient Datacenter Load Balancing in the WildabstractProduction datacenters operate under various uncertainties such as traffic dynamics, topology asymmetry, and failures. Therefore, datacenter load balancing schemes must be resilient to these uncertainties; i.e., they should accurately sense path conditions and timely react to mitigate the fallouts. Despite significant efforts, prior solutions have important drawbacks. On the one hand, solutions such as Presto and DRB are oblivious to path conditions and blindly reroute at fixed granularity. On the other hand, solutions such as CONGA and CLOVE can sense congestion, but they can only reroute when flowlets emerge; thus, they cannot always react timely to uncertainties. To make things worse, these solutions fail to detect/handle failures such as blackholes and random packet drops, which greatly degrades their performance. Hong Zhang 0025, Junxue Zhang 0001, Wei Bai 0001, Kai Chen 0005, Mosharaf Chowdhury |
SIGCOMM | 5 |
| 2016 | HUG: Multi-Resource Fairness for Correlated and Elastic Demands
Mosharaf Chowdhury, Zhenhua Liu 0002, Ali Ghodsi 0002, Ion Stoica |
NSDI | 1 |
| 2016 | Altruistic Scheduling in Multi-Resource Clusters
Robert Grandl, Mosharaf Chowdhury, Aditya Akella, Ganesh Ananthanarayanan |
OSDI | 2 |
| 2016 | EC-Cache: Load-Balanced, Low-Latency Cluster Caching with Online Erasure Coding
K. V. Rashmi, Mosharaf Chowdhury, Jack Kosaian, Ion Stoica, Kannan Ramchandran |
OSDI | 2 |
| 2016 | CODA: Toward Automatically Identifying and Scheduling Coflows in the DarkabstractLeveraging application-level requirements using coflows has recently been shown to improve application-level communication performance in data-parallel clusters. However, existing coflow-based solutions rely on modifying applications to extract coflows, making them inapplicable to many practical scenarios. Hong Zhang 0025, Li Chen 0008, Bairen Yi, Kai Chen 0005, Mosharaf Chowdhury, Yanhui Geng |
SIGCOMM | 5 |
| 2015 | Efficient Coflow Scheduling Without Prior KnowledgeabstractInter-coflow scheduling improves application-level communication performance in data-parallel clusters. However, existing efficient schedulers require a priori coflow information and ignore cluster dynamics like pipelining, task failures, and speculative executions, which limit their applicability. Schedulers without prior knowledge compromise on performance to avoid head-of-line blocking. In this paper, we present Aalo that strikes a balance and efficiently schedules coflows without prior knowledge. Mosharaf Chowdhury, Ion Stoica |
SIGCOMM | 1 |
| 2014 | Efficient coflow scheduling with VarysabstractCommunication in data-parallel applications often involves a collection of parallel flows. Traditional techniques to optimize flow-level metrics do not perform well in optimizing such collections, because the network is largely agnostic to application-level requirements. The recently proposed coflow abstraction bridges this gap and creates new opportunities for network scheduling. In this paper, we address inter-coflow scheduling for two different objectives: decreasing communication time of data-intensive jobs and guaranteeing predictable communication time. We introduce the concurrent open shop scheduling with coupled resources problem, analyze its complexity, and propose effective heuristics to optimize either objective. We present Varys, a system that enables data-intensive frameworks to use coflows and the proposed algorithms while maintaining high network utilization and guaranteeing starvation freedom. EC2 deployments and trace-driven simulations show that communication stages complete up to 3.16X faster on average and up to 2X more coflows meet their deadlines using Varys in comparison to per-flow mechanisms. Moreover, Varys outperforms non-preemptive coflow schedulers by more than 5X. Mosharaf Chowdhury, Yuan Zhong 0001, Ion Stoica |
SIGCOMM | 1 |
| 2013 | Leveraging endpoint flexibility in data-intensive clustersabstractMany applications do not constrain the destinations of their network transfers. New opportunities emerge when such transfers contribute a large amount of network bytes. By choosing the endpoints to avoid congested links, completion times of these transfers as well as that of others without similar flexibility can be improved. In this paper, we focus on leveraging the flexibility in replica placement during writes to cluster file systems (CFSes), which account for almost half of all cross-rack traffic in data-intensive clusters. The replicas of a CFS write can be placed in any subset of machines as long as they are in multiple fault domains and ensure a balanced use of storage throughout the cluster. Mosharaf Chowdhury, Srikanth Kandula, Ion Stoica |
SIGCOMM | 1 |
| 2012 | Coflow: a networking abstraction for cluster applicationsabstractCluster computing applications -- frameworks like MapReduce and user-facing applications like search platforms -- have application-level requirements and higher-level abstractions to express them. However, there exists no networking abstraction that can take advantage of the rich semantics readily available from these data parallel applications. Mosharaf Chowdhury, Ion Stoica |
HotNets | 1 |
| 2012 | Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing
Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauly, Michael J. Franklin, Scott Shenker, Ion Stoica |
NSDI | 2 |
| 2012 | Surviving failures in bandwidth-constrained datacentersabstractDatacenter networks have been designed to tolerate failures of network equipment and provide sufficient bandwidth. In practice, however, failures and maintenance of networking and power equipment often make tens to thousands of servers unavailable, and network congestion can increase service latency. Unfortunately, there exists an inherent tradeoff between achieving high fault tolerance and reducing bandwidth usage in network core; spreading servers across fault domains improves fault tolerance, but requires additional bandwidth, while deploying servers together reduces bandwidth usage, but also decreases fault tolerance. We present a detailed analysis of a large-scale Web application and its communication patterns. Based on that, we propose and evaluate a novel optimization framework that achieves both high fault tolerance and significantly reduces bandwidth usage in the network core by exploiting the skewness in the observed communication patterns. Peter Bodík, Ishai Menache, Mosharaf Chowdhury, Pradeepkumar Mani, David A. Maltz, Ion Stoica |
SIGCOMM | 3 |
| 2012 | FairCloud: sharing the network in cloud computingabstractThe network, similar to CPU and memory, is a critical and shared resource in the cloud. However, unlike other resources, it is neither shared proportionally to payment, nor do cloud providers offer minimum guarantees on network bandwidth. The reason networks are more difficult to share is because the network allocation of a virtual machine (VM) X depends not only on the VMs running on the same machine with X, but also on the other VMs that X communicates with and the cross-traffic on each link used by X. In this paper, we start from the above requirements--payment proportionality and minimum guarantees--and show that the network-specific challenges lead to fundamental tradeoffs when sharing cloud networks. We then propose a set of properties to explicitly express these tradeoffs. Finally, we present three allocation policies that allow us to navigate the tradeoff space. We evaluate their characteristics through simulation and testbed experiments to show that they can provide minimum guarantees and achieve better proportionality than existing solutions. Lucian Popa 0002, Gautam Kumar 0001, Mosharaf Chowdhury, Arvind Krishnamurthy, Sylvia Ratnasamy, Ion Stoica |
SIGCOMM | 3 |
| 2012 | ViNEYard: Virtual Network Embedding Algorithms With Coordinated Node and Link MappingabstractNetwork virtualization allows multiple heterogeneous virtual networks (VNs) to coexist on a shared infrastructure. Efficient mapping of virtual nodes and virtual links of a VN request onto substrate network resources, also known as the VN embedding problem, is the first step toward enabling such multiplicity. Since this problem is known to beNP-hard, previous research focused on designing heuristic-based algorithms that had clear separation between the node mapping and the link mapping phases. In this paper, we present ViNEYard-a collection of VN embedding algorithms that leverage better coordination between the two phases. We formulate the VN embedding problem as a mixed integer program through substrate network augmentation. We then relax the integer constraints to obtain a linear program and devise two online VN embedding algorithms D-ViNE and R-ViNE using deterministic and randomized rounding techniques, respectively. We also present a generalized window-based VN embedding algorithm (WiNE) to evaluate the effect of lookahead on VN embedding. Our simulation experiments on a large mix of VN requests show that the proposed algorithms increase the acceptance ratio and the revenue while decreasing the cost incurred by the substrate network in the long run. Mosharaf Chowdhury, Muntasir Raihan Rahman, Raouf Boutaba |
IEEE/ACM Trans. Netw. | 1 |
| 2011 | Managing data transfers in computer clusters with orchestraabstractCluster computing applications like MapReduce and Dryad transfer massive amounts of data between their computation stages. These transfers can have a significant impact on job performance, accounting for more than 50% of job completion times. Despite this impact, there has been relatively little work on optimizing the performance of these data transfers, with networking researchers traditionally focusing on per-flow traffic management. We address this limitation by proposing a global management architecture and a set of algorithms that (1) improve the transfer times of common communication patterns, such as broadcast and shuffle, and (2) allow scheduling policies at the transfer level, such as prioritizing a transfer over other transfers. Using a prototype implementation, we show that our solution improves broadcast completion times by up to 4.5X compared to the status quo in Hadoop. We also show that transfer-level scheduling can reduce the completion time of high-priority transfers by 1.7X. Mosharaf Chowdhury, Matei Zaharia, Justin Ma, Michael I. Jordan, Ion Stoica |
SIGCOMM | 1 |
| 2010 | Topology-Awareness and Reoptimization Mechanism for Virtual Network Embedding
Nabeel Farooq Butt, Mosharaf Chowdhury, Raouf Boutaba |
Networking | 2 |
| 2010 | A survey of network virtualization
Mosharaf Chowdhury, Raouf Boutaba |
Comput. Networks | 1 |
| 2009 | iMark: An identity management framework for network virtualization environmentabstractNetwork virtualization has been propounded as an open and flexible future internetworking paradigm that allows multiple virtual networks (VNs) to co-exist on a shared physical substrate. Each VN in a network virtualization environment (NVE) is free to implement its own naming, addressing, routing, and transport mechanisms. While such flexibility allows fast and easy deployment of diversified applications and services, ensuring end-to-end communication and universal connectivity poses a daunting challenge. This paper advocates that effective and efficient management of heterogeneous identifier spaces is the key to solving the problem of end-to-end connectivity in an NVE. We propose iMark, an identity management framework based on a global identity space, which enables end hosts to communicate with each other within and outside of their own networks through a set of controllers, adapters, and well-placed mappings without sacrificing the autonomy of the concerned VNs. We describe the procedures that manipulate these mappings between different identifier spaces and provide performance evaluation of the proposed framework. Mosharaf Chowdhury, Fida-E. Zaheer, Raouf Boutaba |
Integrated Network Management | 1 |
| 2009 | Virtual Network Embedding with Coordinated Node and Link MappingabstractRecently network virtualization has been proposed as a promising way to overcome the current ossification of the Internet by allowing multiple heterogeneous virtual networks (VNs) to coexist on a shared infrastructure. A major challenge in this respect is the VN embedding problem that deals with efficient mapping of virtual nodes and virtual links onto the substrate network resources. Since this problem is known to be NP-hard, previous research focused on designing heuristic-based algorithms which had clear separation between the node mapping and the link mapping phases. This paper proposes VN embedding algorithms with better coordination between the two phases. We formulate the VN embedding problem as a mixed integer program through substrate network augmentation. We then relax the integer constraints to obtain a linear program, and devise two VN embedding algorithms D-ViNE and R-ViNE using deterministic and randomized rounding techniques, respectively. Simulation experiments show that the proposed algorithms increase the acceptance ratio and the revenue while decreasing the cost incurred by the substrate network in the long run. Mosharaf Chowdhury, Muntasir Raihan Rahman, Raouf Boutaba |
INFOCOM | 1 |