VLDB 2026 Research / reviewers in the wild / expert
Puneet Sharma 0001
dblp:13/5262-1
· DBLP profile ↗
68ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-4594-8164ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 50 · 6 first-author · 11 since 2021Systems, architecture and hardware · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Not A DPU in Name Only! Unleashing RDMA-capable DPUs in Multi-Tenant Serverless Clouds with NADINO
Shixiong Qi, Songyu Zhang, K. K. Ramakrishnan, Diman Zad Tootaghaj, Hardik Soni 0001, Puneet Sharma 0001 |
EuroSys | 6 |
| 2026 | AdaGen: Workload-Adaptive Cluster Scheduler for Latency-Optimal LLM Inference ServingabstractThe inference workloads of Large Language Models (LLMs) pose significant latency and cost challenges due to increasing model sizes and demand for real-time responses. Existing cluster schedulers for multi-instance LLM serving primarily focus on load balancing to optimize memory usage, which is insufficient for workloads with diverse request characteristics. In such cases, the compute layout — the arrangement of tokens across iterations within each instance—plays a crucial role in determining latency. We propose AdaGen, a workload-adaptive cluster scheduler that minimizes latency and thus maximizes SLO attainment by optimizing compute layouts across instances. AdaGen employs a multi-step scheduling strategy: it first classifies requests based on prefill and decode lengths, then balances load, and finally performs selective distributed execution across instances. Each step incrementally refines the scheduling based on the compute layouts derived from the decision of the previous step. To avoid the overhead of actual execution to generate the layouts, AdaGen introduces a novel simulation-based estimator. Extensive experiments using production workloads show that AdaGen achieves up to 3.6× higher SLO attainment and 2× better cost-efficiency compared to the existing systems, while ensuring scalability. Sudipta Saha Shubha, Ayush Goel, Diman Zad Tootaghaj, Khaled Diab 0001, Hardik Soni 0001, K. K. Ramakrishnan, Puneet Sharma 0001, Haiying Shen |
EuroSys | 7 |
| 2026 | Scaling Attention Beyond GPUs for LLM InferenceabstractScaling inference for large language models is increasingly constrained by limited GPU memory, primarily due to the expanding intermediate states (KV caches) required for long-context generation and multi-user workloads. Once the KV cache exceeds the capacity of high-bandwidth memory, it must be offloaded to host memory and reloaded on demand, a workflow severely bottlenecked by the CPU–GPU interconnect, typically PCIe. Existing approaches exploiting offload KV caches to CPU memory and selectively reload partial segments for attention computation often underutilize CPU compute resources and suffer from accuracy degradation. We present Beyond, a drop-in runtime that integrates a smart offloading scheme to selectively identify and retain salient KV entries across continuous decoding sessions, together with a hybrid CPU–GPU attention mechanism for scalable inference. Beyond executes dense attention over recent KV entries stored in GPU memory while performing parallel, per-head sparse attention on salient contextual KV entries residing in CPU memory. The outputs are fused efficiently through a log-sum-exp scheme. During the bandwidth-constrained decoding phase, oversized KV caches are processed cooperatively by the aggregated CPU and GPU memory bandwidth, with only minimal PCIe data movement. Experiments across diverse models and workloads demonstrate that Beyond improves scalability, supports longer sequences and larger batch sizes, and outperforms existing sparse attention baselines in both efficiency and accuracy—all on commodity GPU hardware. Weishu Deng, Peiran Du, Lingfeng Xiang, Chen Zhong 0002, Faraz Ahmed, Lianjie Cao, Puneet Sharma 0001, Song Jiang 0001, Hui Lu 0001, Jia Rao |
HPDC | 9 |
| 2026 | Griffin: Coherency-Aware Task Scheduling and Memory Allocation for CXL InterconnectsabstractCXL is an emerging interconnect that has the potential to efficiently realize memory disaggregation. This is because CXL enables the expansion of memory beyond individual hosts, and supports coherent memory sharing among multiple hosts. However, CXL introduces several performance overheads due to the cache coherency protocol for memory sharing, as well as placement constraints for shared data, which, if ignored, can lead to correctness issues. This paper presents the first analysis of the impact of CXL memory sharing and shows that the overheads of hardware-based coherency in CXL interconnects are substantial. We then propose Griffin, a new coherency-aware task and memory allocator for CXL disaggregated memory systems. Griffin introduces new abstractions and algorithms that allow it to prioritize which data is allocated remotely and to which memory node, to efficiently reduce the coherence overheads associated with both the amount of shared data and the load on CXL coherence resources. Our simulation results show that Griffin reduces the total memory time by up to 4.29 × compared to a standard baseline and 1.71 × compared to an advanced baseline. Suyeon Lee, Khaled Diab 0001, Diman Zad Tootaghaj, Lianjie Cao, Puneet Sharma 0001, Ada Gavrilovska |
ICS | 5 |
| 2026 | Beyond Monoliths: Enabling Flexible and Composable AI Systems via Memory DisaggregationabstractModern state-of-the-art AI systems are increasingly built as monolithic supernodes integrating large numbers of specialized accelerators with proprietary high-bandwidth interconnects. These systems provision compute, memory, and networking resources in fixed ratios at design time. As AI workloads evolve, their resource demands increasingly diverge from these static configurations, leading to underutilization, limited scalability, and high operational cost. Composable systems based on disaggregated resources offer a more flexible alternative by allowing memory and compute capacity to be scaled independently without replicating an entire supernode. Divya Kiran Kadiyala, Lianjie Cao, Jinsun Yoo, Puneet Sharma 0001, Samantika Sury, Alexandros Daglis |
SIGCOMM | 4 |
| 2026 | CCSwitch: A Scalable Data Plane for Non-Blocking In-Network Collective CommunicationabstractCollective communication operations in AI and HPC workloads generate heavy network traffic. Offloading these operations to network switches reduces latency, but performing arithmetic and replication at line rate is difficult, especially as port counts and link speeds grow. Existing in-network approaches rely on accumulation buffers that not only limit throughput but also require complex state management to handle stragglers and congestion. We present CCSwitch, a modular switching fabric built from 4×4 non-blocking Collective Engines (CEs). Each CE combines spatial and temporal parallelism to perform reductions without accumulation buffers. CEs compose into k-ary n-tree topologies, scaling to 32- and 256-port switches while preserving non-blocking throughput. Source routing and flit-level synchronization keep per-switch state minimal. Our FPGA implementation shows that CCSwitch's quaternary-tree reduction fabric uses up to 23% fewer LUTs and 12–30% fewer flip-flops than a comparable Clos-based design at equal throughput. Enabling the full feature set—source routing, replication, and time-multiplexed VCs—uses 1.4–1.8× more LUTs than the circuit-switched baseline, well below the 3–5× overhead typical of packet-switched NoC routers, while supporting concurrent collectives on shared links. Sumukh Pinge, Hardik Soni 0001, Bob Lantz, Khaled Diab 0001, Lianjie Cao, Tajana Rosing, Puneet Sharma 0001 |
SIGCOMM | 7 |
| 2026 | DynamoServe: A Distributed Tiered Memory System for Multi-tenant LLM ServingabstractThe rapid adoption of large language models (LLMs) has increased the need for efficient multi-tenant inference systems that maximize GPU utilization. However, existing frameworks struggle to scale due to the high memory demands of model weights and key-value (KV) caches. We present DynamoServe, a multi-tenant LLM serving framework that addresses these challenges through three key innovations: (1) leveraging stranded GPU memory to offload model weights and KV caches, (2) mitigating resource fragmentation in multi-workload environments, and (3) improving memory locality through coordinated data placement and demand-driven weight migration across GPUs. Together, these techniques enable high-throughput, low-latency inference. Experiments on state-of-the-art models show that DynamoServe significantly improves memory efficiency without sacrificing latency. Diman Zad Tootaghaj, Khaled Diab 0001, Bob Lantz, Hanjiang Wu, K. K. Ramakrishnan, Md Ashfaqur Rahaman, Ryan Stutsman, Puneet Sharma 0001, Tushar Krishna |
SIGCOMM | 8 |
| 2025 | Self-Clocked Round-Robin Packet Scheduling
Erfan Sharafzadeh, Raymond Matson, Jean Tourrilhes, Puneet Sharma 0001, Soudeh Ghorbani |
NSDI | 4 |
| 2025 | Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA FabricsabstractServerless computing offers resource efficiency but suffers from a heavyweight data plane. We present Palladium, a DPU-offloaded serverless data plane enabling distributed zero-copy communication. Palladium uses two-sided RDMA and cross-processor shared memory to mitigate limitations of wimpy DPU cores. Its DPU-enabled network engine (DNE) isolates RDMA resources and manages flows across tenants. By converting HTTP/TCP to RDMA at ingress, Palladium reduces protocol overhead on the critical path. Shixiong Qi, Songyu Zhang, K. K. Ramakrishnan, Diman Zad Tootaghaj, Hardik Soni 0001, Puneet Sharma 0001 |
SIGCOMM | 6 |
| 2025 | Can Hardware Outsmart Software in Tiered Memory Management? A CMM-H Case StudyabstractWith the advent of Compute Express Link (CXL), hardware-managed memory tiering has become a reality. In this paper, we investigate Samsung's CXL Memory Module-Hybrid (CMM-H), a CXL Type 3 device integrating DRAM and NAND flash managed by an FPGA-based controller and providing byte-addressable memory interface via the cxl.mem protocol. We perform a detailed evaluation of CMM-H and compare its performance with OS-level and block-level tiering solutions. Our results highlight the performance benefits of CMM-H for cache-hit scenarios and identify key limitations for cache-miss situations, offering insights into the trade-offs involved in adopting hardware-managed memory tiering in emerging CXL-based systems. Lingfeng Xiang, Lianjie Cao, Faraz Ahmed, Jia Rao, Hui Lu 0001, Puneet Sharma 0001 |
SYSTOR | 8 |
| 2025 | Efficient and Lightweight Model-Predictive IoT Energy ManagementabstractMulti-sensor IoT devices often rely on renewable energy and batteries to support diverse field deployments. A device’s sensors use significant energy, so careful energy management is needed to maximize its operating time while meeting application requirements. Model Predictive Control (MPC), a successful energy management technique for other application domains, requires significant computational resources usually not available in typical IoT sensing devices. We develop and evaluate a low-complexity approximation to MPC for IoT device operations powered by renewable (solar) energy, where the approximation is guided by the charging characteristics of real batteries. The complex MPC optimization problem is solved by decomposing it into a time-dependent energy allocation problem and a task-dependent sensor scheduling problem, each of which is solvable with cubic time complexity. We utilize a novel combination of an incremental max-min fair allocation method and a recursive dynamic programming-like procedure. This low-complexity predictive optimization step is integrated with a very simple parabola-based adaptive solar prediction to provide a full system solution, termed Predictive EneRgy Management for IoT (PERMIT) , that can be implemented in lightweight IoT devices. PERMIT is evaluated through experiments on a solar-powered Raspberry Pi device with multiple sensors. We also complement our evaluations with simulations based on real device data traces. Results show that PERMIT’s low-complexity algorithm approximates the exact MPC solution very closely. PERMIT also performs significantly better than Signpost, a comparable IoT energy management solution. Elizabeth Liri, K. K. Ramakrishnan, Koushik Kar, Geoff Lyon, Puneet Sharma 0001 |
ACM Trans. Internet Things | 5 |
| 2024 | Accelerating Containerized Machine Learning WorkloadsabstractTo facilitate various Machine Learning (ML) training and inference tasks, enterprises tend to build large and expensive clusters and share them among different teams for diverse ML workloads. Virtualized platforms (containers/VMs) and schedulers are typically deployed to allow such access, manage heterogeneous resources and schedule ML jobs in these clusters. However, allocating resource budgets for different ML jobs to achieve best performance and cluster resource efficiency remains a significant challenge. This work proposes Nearchus to accelerate distributed ML training while ensuring high resource efficiency by using adaptive resource allocation. Nearchus automatically identifies potential performance bottlenecks for running jobs and re-allocates resources to provide optimized run-time performance with high resource efficiency. Nearchus’s resource configuration significantly improves the training speed of individual jobs up to 71.4%–129.1% against state-of-the-art resource schedulers, and reduces job completion and queuing time by 35.6% and 67.8%, respectively. Ali Tariq, Lianjie Cao, Faraz Ahmed, Eric Rozner, Puneet Sharma 0001 |
NOMS | 5 |
| 2024 | Conspirator: SmartNIC-Aided Control Plane for Distributed ML Workloads
Yunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao, Puneet Sharma 0001, Aleksandar Kuzmanovic |
USENIX ATC | 5 |
| 2023 | When Caching Systems Meet Emerging Storage Devices: A Case StudyabstractBlock-layer caching systems improve the I/O performance by using hybrid storage devices; the advent of fast, byte-addressable storage enables caching systems to further leverage new storage tiers (e.g., with persistent memory as the cache device and SSD as the backend device) to achieve better caching performance. However, the new storage devices also challenge the design and implementation of existing block-based caching systems. This paper conducts a comprehensive performance study of a popular caching system, Open CAS, and identifies new, unrevealed software bottlenecks. Our observations and root cause analysis cast light on optimizing the software stack of caching systems to incorporate emerging storage technologies. Lianjie Cao, Faraz Ahmed, Hui Lu 0001, Puneet Sharma 0001 |
HotStorage | 5 |
| 2022 | Slice-Tune: a system for high performance DNN autotuningabstractAutotuning DNN models prior to their deployment is an essential but time-consuming task. Using expensive (and power-hungry) GPU and TPU accelerators efficiently is also key. Since DNNs do not always use a GPU fully, spatial multiplexing of multiple models can provide just the right amount of GPU resources for each DNN. We find that a DNN model tuned with the maximum GPU resources has higher inference latency if less GPU resources are available at inference time. We present methods to tune a DNN model, so that we provide the right amount of accelerator resources during tuning. Thus, even when a wide range of GPU resources are available at inference time, the tuned model achieves low inference latency. Further, existing autotuning frameworks take a long time to tune a model due to inefficient utilization of the client and server-side CPU and GPU. Our system, Slice-Tune, improves several autotuning frameworks to efficiently use system resources by re-thinking the partitioning of tasks between the client and server (where models are profiled on the server GPU), in a Kubernetes environment. We increase parallelism during tuning by sharding the tuning model across multiple tuning application instances, providing concurrent tuning of different operators of a model. We also scale server instances to achieve better GPU multiplexing. Slice-Tune reduces DNN autotuning time in a single GPU and in GPU clusters. Slice-Tune decreases DNN autotuning time by up to 75%, and increase autotuning throughput by a factor of 5, across 3 different autotuning frameworks (TVM, Ansor, and Chameleon). Aditya Dhakal, K. K. Ramakrishnan, Sameer G. Kulkarni, Puneet Sharma 0001, Junguk Cho |
Middleware | 4 |
| 2022 | Metered Boot: Trusted Framework for Application Usage Rights Management in Virtualized EcosystemsabstractThe adoption of virtualization and cloud computing technologies have revolutionized how services and applications can be developed, deployed, and operated to achieve better elasticity, flexibility, and scalability. Multiple stakeholders can be involved for providing online services; each of them plays one or more roles (i.e., service operator, application vendor, and infrastructure provider) to create a customized operating model based on the business requirements. The operating model changes from one business to another, and it may even change at different stages of the same business. A trusted relationship among stakeholders for secure information exchange is the key to enable such flexibility. However, traditional usage compliance methods (e.g., in-person audit, dynamic licensing, and subscription) lack explicit trust among involved parties and the flexibility and scalability to support dynamic sizing of services and applications with low overhead. In this work, we argue the need for a new trust framework to manage application usage rights and propose Metered Boot to provide trusted, capacity/usage-based usage rights management for services and applications deployed in virtualized environments. Metered Boot decouples application workload instantiation for service operators, usage rights governance for application vendors, and resource provisioning for infrastructure providers. We leverage cryptoprocessors (e.g., Trusted Platform Module (TPM)) on commodity servers to generate trusted proofs which are managed by efficient cryptographic construction, Merkle hash tree, for usage rights compliance. We integrated our framework with OpenStack and demonstrate that Metered Boot is able to achieve high scalability and low overhead for instantiating virtual network functions (VNFs). Arun Raghuramu, Lianjie Cao, Puneet Sharma 0001, Joon-Myung Kang, Chen-Nee Chuah, Vinay Saxena |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2022 | Voyager: Revisiting Available Bandwidth Estimation With a New Class of Methods - Decreasing- Chirp-Train MethodsabstractThe available bandwidth (ABW) of a network path is a crucial metric for various applications, such as traffic engineering, congestion control, multimedia streaming, and path selection in software-defined wide-area networks (SDWAN). In recent years, a new class of measurement methods have been proposed to estimate the available bandwidth, decreasing-chirp-train methods. However, the performance and limitations of this new class of methods are neither well studied nor fairly compared beyond simulation studies. In this work, we implement Voyager, a modular framework that allows us to conduct a fair and thorough comparison of how a variety of modern bandwidth estimation methods perform under different network paths and traffic conditions. We shed light on the characteristics and limitations of the internal algorithms of various methods, and propose two new methods. We investigate the impact of various bottleneck types and traffic types, and we explore the performance of these methods on high speed links where interrupt coalescence can cause measurement noise. We finally test Voyager on long-distance Internet links and with live traffic and report our findings. Chang Liu 0156, Jean Tourrilhes, Chen-Nee Chuah, Puneet Sharma 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2021 | Invenio: Communication Affinity Computation for Low-Latency MicroservicesabstractMicroservices enable rapid service deployment and scaling. Integrating poorly-understood microservice components into Service Function Chains (SFCs) or graphs limits a provider's control over service delivery latency, however. Orchestration frameworks currently instantiate and place myriads of microservice components without knowing the impact of placement decisions on latency. Amit Sheoran, Sonia Fahmy, Puneet Sharma 0001, Navin Modi |
ANCS | 3 |
| 2021 | Co-locating containerized workload using service mesh telemetryabstractThe cloud-native architecture and container-based technologies are revolutionizing how online services and applications are designed, developed, and managed by offering better elasticity and flexibility to developers and operators. However, the increasing adoption of microservice and serverless designs makes application workload more decomposed and transient at a larger scale. Most existing container orchestration systems still manage application workload based on simple system-level resource usage and policies manually created by operators, leading to ineffective application-agnostic scheduling and extra management burden for operators. Lianjie Cao, Puneet Sharma 0001 |
CoNEXT | 2 |
| 2021 | Epinoia: Intent Checker for Stateful NetworksabstractIntent-Based Networking (IBN) has been increasingly deployed in production enterprise networks. Automated network configuration in IBN lets operators focus on intents- i.e., the end to end business objectives-rather than spelling out details of the configurations that implement these objectives. Automation brings its own concerns as the administrators cannot rely on traditional network troubleshooting tools. This situation is further exacerbated in the case of stateful Network Functions (NFs) whose packet processing behavior depends on previously observed traffic patterns. To ensure that the network configuration and state derived from network automation matches the administrator’s specified intent, we propose, Epinoia, a network intent checker for stateful networks. Epinoia relies on a unified model for NFs by leveraging the causal precedence relationships that exist between NF packet I/Os and states. Scalability of Epinoia is achieved by decomposing intents into sub-checking tasks and maintaining a causality graph between checked invariants. Epinoia checks for network-wide intent violations incrementally to reduce overhead in the event of network changes. Our evaluation results using real-world network topologies show that Epinoia can perform comprehensive checking within a few seconds per network with intent updates. Huazhe Wang, Puneet Sharma 0001, Faraz Ahmed, Joon-Myung Kang, Chen Qian 0001, Mihalis Yannakakis |
ICCCN | 2 |
| 2021 | Accurate Available Bandwidth Measurement with Packet Batching Mitigation for High Speed NetworksabstractMeasuring the Available Bandwidth (ABW) is an important function for traffic engineering, and in software-defined metro and wide-area network (SD-WAN) applications. Because network speeds are increasing, it is timely to re-visit the effectiveness of ABW measurement again. A significant challenge arises because of Interrupt Coalescence (IC), that network interface drivers use to mitigate the overhead when processing packets at high speed, but introduce packet batching. IC distorts receiver timing and decreases the ABW estimation. This effect is further exacerbated with software-based forwarding platforms that exploit network function virtualization (NFV) and the lower-cost and flexibility that NFV offers, and with the increased use of poll-mode packet processing popularized by the Data Plane Development Kit (DPDK) library. We examine the effectiveness of the ABW estimation with the popular probe rate models (PRM) such as PathChirp and PathCos++, and show that there is a need to improve upon them. We propose a modular packet batching mitigation that can be adopted to improve both the increasing PRM models like PathChirp and decreasing models like PathCos++. Our mitigation techniques improve the accuracy of ABW estimation substantially when packet batching occurs either at the receiver due to IC, DPDK based processing or intermediate NFV-based forwarding nodes. We also show that our technique helps improve estimation significantly in the presence of cross-traffic. Vincent Tran, Jean Tourrilhes, K. K. Ramakrishnan, Puneet Sharma 0001 |
LANMAN | 4 |
| 2020 | Homa: An Efficient Topology and Route Management Approach in SD-WAN OverlaysabstractThis paper presents an efficient topology and route management approach in Software-Defined Wide Area Networks (SD-WAN). Traditional WANs suffer from low utilization and lack of global view of the network. Therefore, during failures, topology/service/traffic changes, or new policy requirements, the system does not always converge to the global optimal state. Using Software Defined Networking architectures in WANs provides the opportunity to design WANs with higher fault tolerance, scalability, and manageability. We exploit the correlation matrix derived from monitoring system between the virtual links to infer the underlying route topology and propose a route update approach that minimizes the total route update cost on all flows. We formulate the problem as an integer linear programming optimization problem and provide a centralized control approach that minimizes the total cost while satisfying the quality of service (QoS) on all flows. Experimental results on real network topologies demonstrate the effectiveness of the proposed approach in terms of disruption cost and average disrupted flows. Diman Zad Tootaghaj, Faraz Ahmed, Puneet Sharma 0001, Mihalis Yannakakis |
INFOCOM | 3 |
| 2019 | Data-driven Resource Allocation in Virtualized Environments
Lianjie Cao, Sonia Fahmy, Puneet Sharma 0001 |
IM | 3 |
| 2018 | Data-driven resource flexing for network functions visualizationabstractResource flexing is the notion of allocating resources on-demand as workload changes. This is a key advantage of Virtualized Network Functions (VNFs) over their non-virtualized counterparts. However, it is difficult to balance the timeliness and resource efficiency when making resource flexing decisions due to unpredictable workloads and complex VNF processing logic. Lianjie Cao, Sonia Fahmy, Puneet Sharma 0001, Shandian Zhe |
ANCS | 3 |
| 2018 | We don't need no licensing serverabstractCloudification of edge to core infrastructure has led to new and rich application and service deployment and operational models. These ecosystems have complex relationships between the application vendors, infrastructure operators and application users. Traditional licensing and compliance enforcement methods such as those based on in person audits and dynamic issuing of license keys inhibit the resource provisioning and consumption flexibility offered by cloudified services due to scalability and management overheads. In this work, we argue the need for a trusted framework for application usage rights compliance. This new architecture named "Metered Boot" provides a way to realize trusted, capacity/usage based rights compliance for service deployments that allows decoupling of usage rights governed by application vendors from the resource provisioning by the infrastructure provider. We have built a Metered Boot prototype for a particular usecase of NFV usage rights compliance. Puneet Sharma 0001, Vinay Saxena, Arun Raghuramu, Chen-Nee Chuah |
HotNets | 1 |
| 2017 | Software defined network inference with evolutionary optimal observation matrices
Mehdi Malboubi, Yanlei Gong, Zijun Yang, Xiong Wang 0001, Chen-Nee Chuah, Puneet Sharma 0001 |
Comput. Networks | 6 |
| 2016 | Decentralizing Network Inference Problems With Multiple-Description Fusion Estimation (MDFE)abstractNetwork inference (or tomography) problems, such as traffic matrix estimation or completion and link loss inference, have been studied rigorously in different networking applications. These problems are often posed as under-determined linear inverse (UDLI) problems and solved in a centralized manner, where all the measurements are collected at a central node, which then applies a variety of inference techniques to estimate the attributes of interest. This paper proposes a novel framework for decentralizing these large-scale under-determined network inference problems by intelligently partitioning it into smaller subproblems and solving them independently and in parallel. The resulting estimates, referred to as multiple descriptions, can then be fused together to compute the global estimate. We apply this Multiple Description and Fusion Estimation (MDFE) framework to three classical problems: traffic matrix estimation, traffic matrix completion, and loss inference. Using real topologies and traces, we demonstrate how MDFE can speed up computation while maintaining (even improving) the estimation accuracy and how it enhances robustness against noise and failures. We also show that our MDFE framework is compatible with a variety of existing inference techniques used to solve the UDLI problems. Mehdi Malboubi, Cuong Vu, Chen-Nee Chuah, Puneet Sharma 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2015 | Software Defined Network Inference with Passive/Active Evolutionary-Optimal pRobing (SNIPER)abstractA key requirement for network management is the accurate and reliable monitoring of relevant network characteristics. In today's large-scale networks, this is a challenging task due to the hard constraints of network measurement resources. This paper proposes a new framework, SNIPER, which leverages the flexibility provided by Software-Defined Networking (SDN) to design the optimal observation or measurement matrix that can leads to the best achievable estimation accuracy using Matrix Completion (MC) techniques. To cope with the complexity of designing large-scale optimal observation matrices, we use the Evolutionary Optimization Algorithms (EOA) which directly target the ultimate estimation accuracy as the optimization objective function. We evaluate the performance of SNIPER using both synthetic and real network measurement traces from different network topologies and by considering two main applications including per-flow size and delay estimations. Our results show that SNIPER can be applied to a variety of network performance measurements under hard resource constraints. For example, by measuring 8.8\% of per-flow path delays in Harvard network, congested paths can be detected with probability 0.94. To demonstrate the feasibility of our framework, we also have implemented a prototype of SNIPER in Mininet. Mehdi Malboubi, Yanlei Gong, Xiong Wang 0001, Chen-Nee Chuah, Puneet Sharma 0001 |
ICCCN | 5 |
| 2015 | Network Policy Whiteboarding and CompositionabstractWe present Policy Graph Abstraction (PGA) that graphically expresses network policies and service chain requirements, just as simple as drawing whiteboard diagrams. Different users independently draw policy graphs that can constrain each other. PGA graph clearly captures user intents and invariants and thus facilitates automatic composition of overlapping policies into a coherent policy. Jeongkeun Lee, Joon-Myung Kang, Chaithan Prakash, Yoshio Turner, Aditya Akella, Charles Clark, Yadi Ma, Puneet Sharma 0001, Ying Zhang 0022 |
SIGCOMM | 8 |
| 2015 | PGA: Using Graphs to Express and Automatically Reconcile Network PoliciesabstractSoftware Defined Networking (SDN) and cloud automation enable a large number of diverse parties (network operators, application admins, tenants/end-users) and control programs (SDN Apps, network services) to generate network policies independently and dynamically. Yet existing policy abstractions and frameworks do not support natural expression and automatic composition of high-level policies from diverse sources. We tackle the open problem of automatic, correct and fast composition of multiple independently specified network policies. We first develop a high-level Policy Graph Abstraction (PGA) that allows network policies to be expressed simply and independently, and leverage the graph structure to detect and resolve policy conflicts efficiently. Besides supporting ACL policies, PGA also models and composes service chaining policies, i.e., the sequence of middleboxes to be traversed, by merging multiple service chain requirements into conflict-free composed chains. Our system validation using a large enterprise network policy dataset demonstrates practical composition times even for very large inputs, with only sub-millisecond runtime latencies. Chaithan Prakash, Jeongkeun Lee, Yoshio Turner, Joon-Myung Kang, Aditya Akella, Sujata Banerjee, Charles Clark, Yadi Ma, Puneet Sharma 0001, Ying Zhang 0022 |
SIGCOMM | 9 |
| 2014 | Democratic Resolution of Resource Conflicts Between SDN Control ProgramsabstractResource conflicts are inevitable on any shared infrastructure. In Software-Defined Networks (SDNs), different controller modules with diverse objectives may be installed on the SDN controller. Each module independently generates resource requests that may conflict with the objectives of a different module. For example, a controller module for maintaining high availability may want resource allocations that require too much core network bandwidth and thus conflict with another module that aims to minimize core bandwidth usage. In such a situation, it is imperative to identify and install resource allocations that achieve network wide global objectives that may not be known to individual modules, e.g., high availability with acceptable bandwidth usage. This problem has received only limited attention, with most prior work focused on detecting, avoiding, and resolving rule-level conflicts in the context of OpenFlow. Alvin AuYoung, Yadi Ma, Sujata Banerjee, Jeongkeun Lee, Puneet Sharma 0001, Yoshio Turner, Jeffrey C. Mogul |
CoNEXT | 5 |
| 2014 | Intelligent SDN based traffic (de)Aggregation and Measurement Paradigm (iSTAMP)abstractFine-grained traffic flow measurement, which provides useful information for network management tasks and security analysis, can be challenging to obtain due to monitoring resource constraints. The alternate approach of inferring flow statistics from partial measurement data has to be robust against dynamic temporal/spatial fluctuations of network traffic. In this paper, we propose an intelligent Traffic (de)Aggregation and Measurement Paradigm (iSTAMP), which partitions TCAM entries of switches/routers into two parts to: 1) optimally aggregate part of incoming flows for aggregate measurements, and 2) de-aggregate and directly measure the most informative flows for per-flow measurements. iSTAMP then processes these aggregate and per-flow measurements to effectively estimate network flows using a variety of optimization techniques. With the advent of Software-Defined-Networking (SDN), such real-time rule (re)configuration can be achieved via OpenFlow or other similar SDN APIs. We first show how to design the optimal aggregation matrix for minimizing the flow-size estimation error. Moreover, we propose a method for designing an efficient-compressive flow aggregation matrix under hard resource constraints of limited TCAM sizes. In addition, we propose an intelligent Multi-Armed Bandit based algorithm to adaptively sample the most “rewarding” flows, whose accurate measurements have the highest impact on the overall flow measurement and estimation performance. We evaluate the performance of iSTAMP using real traffic traces from a variety of network environments and by considering two applications: traffic matrix estimation and heavy hitter detection. Also, we have implemented a prototype of iSTAMP and demonstrated its feasibility and effectiveness in Mininet environment. Mehdi Malboubi, Chen-Nee Chuah, Puneet Sharma 0001 |
INFOCOM | 4 |
| 2014 | Application-driven bandwidth guarantees in datacentersabstractProviding bandwidth guarantees to specific applications is becoming increasingly important as applications compete for shared cloud network resources. We present CloudMirror, a solution that provides bandwidth guarantees to cloud applications based on a new network abstraction and workload placement algorithm. An effective network abstraction should enable applications to easily and accurately specify their requirements, while simultaneously enabling the infrastructure to provision resources efficiently for deployed applications. Prior research has approached the bandwidth guarantee specification by using abstractions that resemble physical network topologies. We present a contrasting approach of deriving a network abstraction based on application communication structure, called Tenant Application Graph or TAG. CloudMirror also incorporates a new workload placement algorithm that efficiently meets bandwidth requirements specified by TAGs while factoring in high availability considerations. Extensive simulations using real application traces and datacenter topologies show that CloudMirror can handle 40% more bandwidth demand than the state of the art (e.g., the Oktopus system), while improving high availability from 20% to 70%. Jeongkeun Lee, Yoshio Turner, Myungjin Lee, Lucian Popa 0002, Sujata Banerjee, Joon-Myung Kang, Puneet Sharma 0001 |
SIGCOMM | 7 |
| 2014 | Streaming Solutions for Fine-Grained Network Traffic Measurements and AnalysisabstractOnline network traffic measurements and analysis is critical for detecting and preventing any real-time anomalies in the network. We propose, implement, and evaluate an online, adaptive measurement platform, which utilizes real-time traffic analysis results to refine subsequent traffic measurements. Central to our solution is the concept of Multi-Resolution Tiling (MRT), a heuristic approach that performs sequential analysis of traffic data to zoom into traffic subregions of interest. However, MRT is sensitive to transient traffic spikes. In this paper, we propose three novel traffic streaming algorithms that overcome the limitations of MRT and can cater to varying degrees of computational and storage budgets, detection latency, and accuracy of query response. We evaluate our streaming algorithms on a highly parallel and programmable hardware as well as a traditional software-based platforms. The algorithms demonstrate significant accuracy improvement over MRT in detecting anomalies consisting of synthetic hard-to-track elephant flows and global icebergs. Our proposed algorithms maintain the worst-case complexities of the MRT while incurring only a moderate increase in average resource utilization. Nicholas Hosein, Soheil Ghiasi, Chen-Nee Chuah, Puneet Sharma 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2013 | Compressive sensing network inference with multiple-description fusion estimationabstractWe have previously introduced Multiple Description Fusion Estimation (MDFE) framework that partitions a large-scale Under-Determined Linear Inverse (UDLI) problem into smaller sub-problems that can be solved independently and in parallel. The resulting estimates, referred to as multiple descriptions, can then be fused together to compute the global estimate [1]. In this paper, we extend MDFE framework to make it compatible with Compressive Sensing (CS) network inference, where the attributes of interests (i.e. unknowns) are fluctuating rapidly over time and/or space. For this purpose, we propose a new clustering based technique to intelligently divide a large-scale compressive sensing problem into smaller sub-problems where observations between sub-spaces contain redundancy. We apply this new framework, referred to as Compressive Sensing MDFE (CS-MDFE), to three classical inference problems in networking: traffic matrix estimation, traffic matrix completion, and loss inference. Using real topologies and traces, we demonstrate how CS-MDFE can improve the estimation accuracy and speed up computation time, and how it enhances robustness against noise and failures. We also show that this framework is compatible with different CS inference techniques. Mehdi Malboubi, Cuong Vu, Chen-Nee Chuah, Puneet Sharma 0001 |
GLOBECOM | 4 |
| 2013 | Corybantic: towards the modular composition of SDN control programsabstractSoftware-Defined Networking (SDN) promises to enable vigorous innovation, through separation of the control plane from the data plane, and to enable novel forms of network management, through a controller that uses a global view to make globally-valid decisions. The design of SDN controllers creates novel challenges; much previous work has focused on making them scalable, reliable, and efficient. Jeffrey C. Mogul, Alvin AuYoung, Sujata Banerjee, Lucian Popa 0002, Jeongkeun Lee, Jayaram Mudigonda, Puneet Sharma 0001, Yoshio Turner |
HotNets | 7 |
| 2013 | Enhancing network management frameworks with SDN-like control
Puneet Sharma 0001, Sujata Banerjee, Sébastien Tandel, Renato Aguiar, Raphael Amorim, David Pinheiro |
IM | 1 |
| 2013 | Pegasus: Precision hunting for icebergs and anomalies in network flowsabstractAccurate online network monitoring is crucial for detecting attacks, faults, and anomalies, and determining traffic properties across the network. With high bandwidth links and consequently increasing traffic volumes, it is difficult to collect and analyze detailed flow records in an online manner. Traditional solutions that decouple data collection from analysis resort to sampling and sketching to handle large monitoring traffic volumes. We propose a new system, Pegasus, to leverage commercially available co-located compute and storage devices near routers and switches. Pegasus adaptively manages data transfers between monitors and aggregators based on traffic patterns and user queries. We use Pegasus to detect global icebergs or global heavy-hitters. Icebergs are flows with a common property that contribute a significant fraction of network traffic. For example, DDoS attack detection is an iceberg detection problem with a common destination IP. Other applications include identification of “top talkers,” top destinations, and detection of worms and port scans. Experiments with Abilene traces, sFlow traces from an enterprise network, and deployment of Pegasus as a live monitoring service on PlanetLab show that our system is accurate and scales well with increasing traffic and number of monitors. Sriharsha Gangam, Puneet Sharma 0001, Sonia Fahmy |
INFOCOM | 2 |
| 2013 | Decentralizing network inference problems with Multiple-Description Fusion Estimation (MDFE)abstractTwo forms of network inference (or tomography) problems have been studied rigorously: (a) traffic matrix estimation or completion based on link-level traffic measurements, and (b) link-level loss or delay inference based on end-to-end measurements. These problems are often posed as underdetermined linear inverse (UDLI) problems and solved in a centralized manner, where all the measurements are collected at a central node, which then applies a variety of inference techniques to estimate the attributes of interest. This paper proposes a novel framework for decentralizing these large-scale UDLI network inference problems by intelligently partitioning it into smaller sub-problems and solving them independently and in parallel. The resulting estimates, referred to as multiple descriptions, can then be fused together to compute the global estimate. We apply this Multiple Description and Fusion Estimation (MDFE) framework to three classical problems: traffic matrix estimation, traffic matrix completion, and loss inference. Using real topologies and traces, we demonstrate how MDFE can speed up computation time while maintaining (even improving) the estimation accuracy and how it enhances robustness against noise and failures. We also show that our MDFE framework is compatible with a variety of existing inference techniques used to solve the UDLI problems. Mehdi Malboubi, Cuong Vu, Chen-Nee Chuah, Puneet Sharma 0001 |
INFOCOM | 4 |
| 2012 | NEEM: Network energy efficiency managerabstractThe energy consumed by networks is growing and while it is not the dominant contributor to IT energy spend, the absolute power consumption numbers are staggeringly large. As servers and cooling within data centers and enterprises becomes more energy-efficient, it is critical that we address the energy management for networking devices now and make these devices energy efficient as well as energy proportional. Network device, topology and route control based on configurations, traffic demands and performance requirements can be leveraged for reducing network energy consumption. In this paper we present NEEM (Network Energy Efficiency Manager), a network energy management solution for automating the policy driven analysis for energy efficiency and implementing the network changes that can result in significant savings by adapting to the network traffic dynamics and requirements. Our initial results on an enterprise network, based on offline analysis show that NEEM can save over 30% of network energy consumption. NEEM is a vendor-neutral, standards-based solution. Puneet Sharma 0001, Sujata Banerjee, Deniz Demir, Srikanth Natarajan, Swamy Mandavilli |
NOMS | 1 |
| 2011 | DevoFlow: scaling flow management for high-performance networksabstractOpenFlow is a great concept, but its original design imposes excessive overheads. It can simplify network and traffic management in enterprise and data center environments, because it enables flow-level control over Ethernet switching and provides global visibility of the flows in the network. However, such fine-grained control and visibility comes with costs: the switch-implementation costs of involving the switch's control-plane too often and the distributed-system costs of involving the OpenFlow controller too frequently, both on flow setups and especially for statistics-gathering. Andrew R. Curtis, Jeffrey C. Mogul, Jean Tourrilhes, Praveen Yalagandula, Puneet Sharma 0001, Sujata Banerjee |
SIGCOMM | 5 |
| 2010 | Network Integrated Transparent TCP AcceleratorabstractNetwork device vendors have recently opened up the processing capabilities on their hardware platform to support third-party applications. In this paper, we explore the requirements and overheads associated with co-locating middlebox functionality on such computing resources on networking hardware. In particular, we use an example of TCP acceleration proxy (CHART) that improves throughput over networks with delay and loss. The CHART system, developed by HP and its partners provides enhanced TCP/IP performance and service quality guarantees by deploying performance accelerating proxies, which enables legacy clients to benefit by high-performance network service. Use of the TCP proxy, however, requires manual configuration on the clients changing http proxy and/or routing table settings. Can we remove the need to configure end-hosts by inserting a transparent TCP proxy in the path, without losing performance? To address this question, we implement the accelerator on HP's x86-based processing blade designed to integrate network applications within switch architecture as well as on low-end home routers with OpenWRT. We describe the implementation detail such as flow redirection for transparency and new mechanisms required for easy insertion of proxies in the network path. We also evaluate its performance on HP's experimental testbed in terms of throughput and additional processing overhead. Jeongkeun Lee, Puneet Sharma 0001, Jean Tourrilhes, Rick McGeer, Jack Brassil, Andy C. Bavier |
AINA | 2 |
| 2010 | Leveraging Correlations between Capacity and Available Bandwidth to Scale Network MonitoringabstractRecently, there has been a tremendous growth in the number of installed distributed computing platforms such as those for content distribution networks, cloud computing infrastructures, and distributed data centers. Such distributed platforms need a scalable end-to-end (e2e) network monitoring component to provide Quality of Service (QoS) guarantees to the services and improve the overall performance. An important challenge for a network monitoring infrastructure is the periodicity of the measurements as this aspect trades off the monitoring overheads with staleness of the results. In the Network Genome project, we explore the relationships between different e2e network metrics with the aim of leveraging such relationships for reducing monitoring costs while maintaining measurement accuracy. We perform our analysis using long range network measurements from PlanetLab, where we have been collecting e2e network data (route, number of hops, capacity bandwidth and available bandwidth) as part of the S3 system since January 2006. In this paper, we focus on the correlation between the Capacity and Available Bandwidth metrics between host pairs in the PlanetLab testbed. Our analysis shows that the ranking of hosts with respect to their Capacity to/from a set of nodes is a good indicator of the ranking of hosts with respect to their Available Bandwidth to/from the same set of nodes. Praveen Yalagandula, Sung-Ju Lee 0001, Puneet Sharma 0001, Sujata Banerjee |
GLOBECOM | 3 |
| 2010 | DevoFlow: cost-effective flow management for high performance enterprise networksabstractThe OpenFlow framework enables flow-level control over Ethernet switching, as well as centralized visibility of the flows in the network. OpenFlow's coupling of these features comes with costs, however: the distributed-system costs of involving the OpenFlow controller on flow setups, and the switch-implementation costs of involving the switch's control plane too often. Jeffrey C. Mogul, Jean Tourrilhes, Praveen Yalagandula, Puneet Sharma 0001, Andrew R. Curtis, Sujata Banerjee |
HotNets | 4 |
| 2010 | Understanding the Effectiveness of a Co-Located Wireless Channel Monitoring Surrogate SystemabstractIn Wireless Local Area Networks (WLANs), channel management is important in achieving reliable data communications and satisfying QoS requirements. The key aspects of wireless channel management are monitoring the channel quality and adapting quickly to the network conditions by switching to a better channel. We propose a wireless channel monitoring system with co-located monitoring surrogates. Our system works on multi-radio Access Points (APs) where a co-located surrogate radio monitors the condition of various channels while the master radio serves the clients for data communication. Although we have designed our system for generic WLANs, we believe it will be most useful for IEEE 802.11n networks where there are a large number of channels and dynamic frequency selection is required. Our system enables intelligent, fast channel adaptation, reduces service disruption time, and consequently helps realize the performance potential of 802.11n. We present our multi-radio co-located wireless channel monitoring surrogate system and evaluate its effectiveness on our IEEE 802.11n network testbed. We also perform case studies to demonstrate the benefit our system brings compared against the existing schemes. Jeongkeun Lee, Sung-Ju Lee 0001, Puneet Sharma 0001, Sungjoon Choi 0001 |
ICC | 3 |
| 2010 | ElasticTree: Saving Energy in Data Center Networks
Brandon Heller, Srinivasan Seetharaman, Priya Mahadevan, Yiannis Yiakoumis, Puneet Sharma 0001, Sujata Banerjee, Nick McKeown |
NSDI | 5 |
| 2010 | No more middlebox: integrate processing into networkabstractTraditionally, in-network services like firewall, proxy, cache, and transcoders have been provided by dedicated hardware middleboxes. A recent trend has been to remove the middleboxes by deploying the network services into switch/router-integrated computing modules or separate server/blade machines. In this abstract, by using a web Ad-insertion application as an example, we demonstrate our in-network processing (INP) framework that orchestrates various computing resources and network devices and enables seamless and efficient deployments of network services. Jeongkeun Lee, Jean Tourrilhes, Puneet Sharma 0001, Sujata Banerjee |
SIGCOMM | 3 |
| 2009 | The Case for Service OverlaysabstractThe Internet was designed as a packet-switched network in the 1960's and 1970's, with the explicit intent of sacrificing quality-of-service guarantees for an individual application in order to optimize channel usage and provide optimal median service for all applications. This approach was successful, since the application mix of the Internet heretofore has been dominated by applications with low quality-of-service needs: primarily bulk data transfer and low-bandwidth text- based interactive applications. As the Internet absorbs other networks and applications with strong quality-of-service requirements (television, voice over IP) , this tradeoff changes. We are faced with the problem of introducing the quality- of-service guarantees of circuit-switching into packet-switched networks. Fundamental change to the lower layers of the Internet stack have proven infeasible; even strongly-motivated, well-designed modifications which made transition a first-class design consideration have had difficult introductions. ATM and IPv6 are two recent examples. One effective transition strategy for new networking techniques has been the use of overlays. In this paper, we introduce the concept of a service overlay network, to offer circuit-switched behavior on legacy IP networks, and establish the requirements on the underlying IP network to make this strategy effective. Jack Brassil, Rick McGeer, Puneet Sharma 0001, Praveen Yalagandula, Brian L. Mark, Stephen Schwab |
ICCCN | 3 |
| 2009 | Supporting application network flows with multiple QoS constraintsabstractThere is a growing need to support real-time applications over the Internet. Real-time interactive applications often have multiple quality-of-service (QoS) requirements which are application specific. Traditional provisioning of QoS in the Internet through IP routing - Intserv or Diffserv - faces many technical challenges, and is also deterred by the huge deployment issues. As an alternative, application providers often build their own application-specific overlay networks to meet their QoS requirements. In this paper, we present a unified framework which can serve diverse applications with multiple QoS constraints. Our scalable flow route management architecture, called MCQoS, employs a hybrid approach using a path vector protocol to disseminate aggregated path information combined with on-demand path discovery to find paths that match the diverse QoS requirements. It uses a distributed algorithm to dynamically adapt to an alternate path when the current path fails to satisfy the required QoS constraints. We do large-scale simulation and analysis to show that our approach is both efficient and scalable, and that it substantially outperforms the state of the art protocols in accuracy. Our simulation results show that MCQoS can reduce the false negative percentage to less than 1% compared with 5-10% in other approaches, and eliminates false positives, whereas other schemes have false positive rates of 10-20% with minimal increase in protocol overhead. Finally, we implemented and deployed our system on the Planetlab testbed for evaluation in a real network environment. Amit Mondal, Puneet Sharma 0001, Sujata Banerjee, Aleksandar Kuzmanovic |
IWQoS | 2 |
| 2009 | A Power Benchmarking Framework for Network Devices
Priya Mahadevan, Puneet Sharma 0001, Sujata Banerjee, Parthasarathy Ranganathan |
Networking | 2 |
| 2009 | NodeWiz: Fault-tolerant grid information service
Sujoy Basu, Lauro Beltrão Costa, Francisco Vilar Brasileiro, Sujata Banerjee, Puneet Sharma 0001, Sung-Ju Lee 0001 |
Peer-to-Peer Netw. Appl. | 5 |
| 2008 | API Design Challenges for Open Router Platforms on Proprietary Hardware
Jeffrey C. Mogul, Praveen Yalagandula, Jean Tourrilhes, Rick McGeer, Sujata Banerjee, Tim Connors, Puneet Sharma 0001 |
HotNets | 7 |
| 2008 | Bandwidth-Aware Routing in Overlay NetworksabstractIn the absence of end-to-end quality of service (QoS), overlay routing has been used as an alternative to the default best effort Internet routing. Using end-to-end network measurement, the problematic parts of the path can be bypassed, resulting in improving the resiliency and robustness to failures. Studies have shown that overlay paths can give better latency, loss rate, and TCP throughput. Overlay routing also offers flexibility as different routes can be used based on application needs. There have been very few proposals of using bandwidth as the main metric of interest, which is of great concern in media applications. We introduce our scheme BARON (Bandwidth-Aware Routing in Overlay Networks) that utilizes capacity between the end hosts to identify viable overlay paths and measures available bandwidth to select the best route. We propose our path selection approaches, and using the measurements between 174 PlanetLab nodes and over 13,189 paths, we evaluate the usefulness of overlay routes in terms of bandwidth gain. Our results show that among 658,526 overlay paths, 25% have larger bandwidth than their native IP routes, and over 86% of (source, destination) pairs have at least one overlay route with larger bandwidth than the default IP routes. We also present the effectiveness of BARON in preserving the bandwidth requirement over time for a few selected Internet paths. Sung-Ju Lee 0001, Sujata Banerjee, Puneet Sharma 0001, Praveen Yalagandula, Sujoy Basu |
INFOCOM | 3 |
| 2008 | Minerva: Learning to Infer Network Path PropertiesabstractKnowledge of the network path properties such as latency, hop count, loss and bandwidth is key to the performance of overlay networks, grids and P2P applications. Network operators also use these metrics for managing and diagnosing problems in their networks. However, the size of the Internet makes the task of measuring these metrics immensely difficult. A more scalable approach of inference and estimation of these metrics based on partial measurements has been recently adopted. Current inference approaches do not adapt to different network topologies and the evolution of the network over time. In this paper, we propose a novel learning based approach, called Minerva, for the inferencing of inter-node properties. Minerva uses partial measurements to create signature-like profiles for the participating nodes. These signatures are later used as input to a trained Bayesian network module to estimate the different network properties. We have built a system based on our approach and present performance results from real network measurements obtained from the Planet-Lab testbed. The sensitivity of the system to different parameters including training set, measurement overhead, and size of network have also been studied in this paper. Rita H. Wouhaybi, Puneet Sharma 0001, Sujata Banerjee, Andrew T. Campbell |
INFOCOM | 2 |
| 2008 | QoS-guaranteed path selection algorithm for service compositionabstractService overlay networking is an emerging approach, which employs overlay nodes to provide advanced services by dynamically composing it from basic services available on overlay nodes. Advanced service request from users can have different and multiple quality-of-service (QoS) requirements and findi Puneet Sharma 0001, Sujata Banerjee |
QSHINE | 2 |
| 2007 | Aggregating Bandwidth for Multihomed Mobile Collaborative CommunitiesabstractMultihomed, mobile wireless computing and communication devices can spontaneously form communities to logically combine and share the bandwidth of each other's wide-area communication links using inverse multiplexing. But, membership in such a community can be highly dynamic, as devices and their associated WWAN links randomly join and leave the community. We identify the issues and trade-offs faced in designing a decentralized inverse multiplexing system in this challenging setting and determine precisely how heterogeneous WWAN links should be characterized and when they should be added to, or deleted from, the shared pool. We then propose methods of choosing the appropriate channels on which to assign newly arriving application flows. Using video traffic as a motivating example, we demonstrate how significant performance gains can be realized by adapting allocation of the shared WWAN channels to specific application requirements. Our simulation and experimentation results show that collaborative bandwidth aggregation systems are, indeed, a practical and compelling means of achieving high-speed Internet access for groups of wireless computing devices beyond the reach of public or private access points Puneet Sharma 0001, Sung-Ju Lee 0001, Jack Brassil, Kang G. Shin |
IEEE Trans. Mob. Comput. | 1 |
| 2006 | Implementation and Evolution of Packet Striping for Media Streaming Over Multiple Burst-Loss ChannelsabstractModern mobile devices are multi-homed with WLAN and WWAN communication interfaces. In a community of nodes with such multi-homed devices-locally inter-connected via high-speed WLAN but each globally connected to larger networks via low-speed WWAN, striping high-volume traffic from remote large networks over a bundle of low speed WWAN links can overcome the bandwidth mismatch problem between WLAN and WWAN. In our previous work, we showed that a packet striping system for such multi-homed devices-a mapping of delay-sensitive packets by an intermediate gateway to multiple channels using combination of retransmissions (ARQ) and forward error corrections (FEC)-can dramatically enhance the overall performance. In this paper, we improve upon a previous algorithm in two respects. First, by introducing two-tier dynamic programming tables to memoize computed solutions, packet striping decisions translate to simple table lookup operations given stationary network statistics. Doing so drastically reduces striping operation complexity. Second, new weighting functions are introduced into the hybrid ARQ/FEC algorithm to drive the long-term striping system evolution away from pathological local minima that are far from the global optimum. Results show the new algorithm performs efficiently and gives improved performance by avoiding local minima compared to the previous algorithm Gene Cheung, Puneet Sharma 0001, Sung-Ju Lee 0001 |
ICME | 2 |
| 2006 | Distributed Querying of Internet Distance InformationabstractAbstract — Estimation of network proximity among nodes is an important building block in several applications like service selection and composition, multicast tree formation, and overlay construction. Recently, scalable techniques have been proposed to estimate inter-node latencies, including network coordinate systems like GNP and Vivaldi. However, existing mechanisms for querying such information do not scale well to a very large number of nodes, when one wants to accurately find a set of nodes globally closest to a given node. In this paper we are concerned with distributing the position data among a set of infrastructure nodes, and propose ways of partitioning and querying this data. The trade-offs between accuracy and overhead in this distributed infrastructure are explored. We evaluate our solution through simulations with real and synthetic network measurement data. I. Rodrigo Fonseca, Puneet Sharma 0001, Sujata Banerjee, Sung-Ju Lee 0001, Sujoy Basu |
INFOCOM | 2 |
| 2006 | SmartSeer: Using a DHT to Process Continuous Queries Over Peer-to-Peer NetworksabstractAbstract — As the academic world moves away from physical journals and proceedings towards online document repositories, the ability to efficiently locate work of interest among the torrent of newly-generated papers will become increasingly important. To aid in this endeavor, we designed SmartSeer, a system that allows users to register personalized continuous queries over the CiteSeer database of technical documents. Users are then alerted whenever papers that match their queries are put online. SmartSeer has two main design requirements. First, to allow effective information retrieval, it should support rich continuous queries (as opposed to simple keyword searches). Second, to make effective use of donated infrastructure, it should be capable of running on a loosely maintained group of unreliable machines spread across multiple organizations (as opposed to assuming a reliable and tightly coupled distributed system). Existing work on distributed continuous query systems fails at least one of these requirements. Our design for SmartSeer is based on Distributed Hash Tables (DHTs), and thereby leverages previous work on DHT-based query systems. A prototype of SmartSeer has been implemented and evaluated on Planetlab. Though we evaluate our design only for the SmartSeer application, we believe it also provides useful insights into other distributed and rich continuous query systems (web alerts, news alerts etc). I. Jayanthkumar Kannan, Beverly Yang, Scott Shenker, Puneet Sharma 0001, Sujata Banerjee, Sujoy Basu, Sung-Ju Lee 0001 |
INFOCOM | 4 |
| 2006 | QoS-Guaranteed Path Selection Algorithm for Service CompositionabstractIn this paper, QoS-guaranteed path selection algorithm for service composition is described. A heuristic algorithm, K-closest pruning (KCP), is used to solve the problem of multi-constraint service path selection for SON in polynomial time. The main feature of this algorithm is that the path selected by this algorithm meets all the QoS requirements specified by the user/application. The main idea in this approach is to leverage network proximity information to reduce search space, i.e. to reduce the number of qualified overlay nodes/links, for service path selection for a given request. This paper demonstrates the two distinct objectives reuse and load balancing and their use in modifying the path selection of KCP algorithm to achieve the objectives with affecting performance of KCP algorithm Puneet Sharma 0001, Sujata Banerjee |
IWQoS | 2 |
| 2006 | Improving aggregated channel performance through decentralized channel monitoring
Puneet Sharma 0001, Jack Brassil, Sung-Ju Lee 0001, Kang G. Shin |
Comput. Networks | 1 |
| 2005 | NodeWiz: peer-to-peer resource discovery for gridsabstractEfficient resource discovery based on dynamic attributes such as CPU utilization and available bandwidth is a crucial problem in the deployment of computing grids. Existing solutions are either centralized or unable to answer advanced resource queries (e.g., range queries) efficiently. We present the design of NodeWiz, a grid information service (CIS) that allows multi-attribute range queries to be performed efficiently in a distributed manner. This is obtained by aggregating the directory services of individual organizations in a peer-to-peer information service. Sujoy Basu, Sujata Banerjee, Puneet Sharma 0001, Sung-Ju Lee 0001 |
CCGRID | 3 |
| 2005 | Distributed communication paradigm for wireless community networksabstractDistributed computing has been widely embraced as a cost-effective means of performing compute-intensive tasks by pooling the computational resources of collaborating systems. We envision the emergence of an analogous approach to communication resource sharing which we call distributed communication. Distributed communication enables sharing a set of relatively low-speed WAN channels emanating from communities of multi-homed devices interconnected with a highspeed wireless LAN. We envisage opportunities to aggregate cellular links from spontaneously formed ad hoc groups of mobile devices, as well as broadband access links (e.g., DSL) from neighboring residences. But, will individuals be willing to share bandwidth as easily as they share bits? A prototype system that we have constructed convinces us that the technical challenges of distributed communication can be overcome. And there appears to be no other means of satisfying the growing demand for access bandwidth as quickly and as cheaply. Puneet Sharma 0001, Sung-Ju Lee 0001, Jack Brassil |
ICC | 1 |
| 2005 | Striping Delay-Sensitive Packets Over Multiple Bursty Wireless ChannelsabstractMulti-homed mobile devices have multiple wireless communication interfaces, each connecting to the Internet via a low speed and bursty WAN link such as a cellular link. We propose a packet striping system for such multi-homed devices — a mapping of packets by agateway to multiple channels, such that the overall performance is enhanced. We model and analyze the striping of delay-sensitive packets over multiple burst-loss channels. We derive the expected packet loss ratio when FEC (Forward Error Correction) and retransmissions are applied for error protection over multiple channels. We next model and analyze the case when the channels are bandwidth-limited. We develop a dynamic programming based algorithm that solves the optimal striping problem for the ARQ, the FEC, and the hybrid FEC/ARQ case. Gene Cheung, Puneet Sharma 0001, Sung-Ju Lee 0001 |
ICME | 2 |
| 2005 | Distributed querying of Internet distance informationabstractEstimation of network proximity among nodes is an important building block in several applications like service selection and composition, multicast tree formation, and overlay construction. Recently, scalable techniques have been proposed to estimate inter-node latencies, including network coordinate systems like GNP and Vivaldi. However, existing mechanisms for querying such information do not scale well to a very large number of nodes, when one wants to accurately find a set of nodes globally closest to a given node. In this paper we are concerned with distributing the position data among a set of infrastructure nodes, and propose ways of partitioning and querying this data. The trade-offs between accuracy and overhead in this distributed infrastructure are explored. We evaluate our solution through simulations with real and synthetic network measurement data. Rodrigo Fonseca, Puneet Sharma 0001, Sujata Banerjee, Sung-Ju Lee 0001, Sujoy Basu |
INFOCOM | 2 |
| 2005 | Striping Delay-sensitive Packets over Multiple Burst-loss Channels with Random DelaysabstractMulti-homed mobile devices have multiple wireless communication interfaces, each connecting to the Internet via a long range but low speed and bursty WAN link such as a cellular link. We propose a packet striping system for such multi-homed devices - a mapping of delay-sensitive packets by an intermediate gateway to multiple channels, such that the overall performance is enhanced. In particular, we model and analyze the striping of delay-sensitive packets over multiple burst-loss channels with random delays. We first derive the expected packet loss ratio when forward error correction (FEC) is applied for error protection over multiple channels. We next model and analyze the case when the channels are bandwidth-limited with shifted-gamma-distributed transmission delays. We develop a dynamic programming-based algorithm that solves the optimal striping problem for the ARQ, the FEC, and the hybrid FEC/ARQ case. Gene Cheung, Puneet Sharma 0001, Sung-Ju Lee 0001 |
ISM | 2 |
| 2004 | Handheld Routers: Intelligent Bandwidth Aggregation for Mobile Collaborative CommunitiesabstractMulti-homed, mobile wireless computing and communication devices can spontaneously form communities to logically combine and share the bandwidth of each other's wide-area communication links using inverse multiplexing. But membership in such a community can be highly dynamic, as devices and their associated WAN links randomly join and leave the community. We identify the issues and tradeoffs faced in designing a decentralized inverse multiplexing system in this challenging setting, and determine precisely how heterogeneous WAN links should be characterized, and when they should be added to, or deleted from, the shared pool. We then propose methods of choosing the appropriate channels on which to assign newly-arriving application flows. Using video traffic as a motivating example, we demonstrate how significant performance gains can be realized by adapting allocation of the shared WAN channels to specific application requirements. Our simulation and experimentation results show that collaborative bandwidth aggregation systems are, indeed, a practical and compelling means of achieving high-speed Internet access for groups of wireless computing devices beyond the reach of public or private access points. Puneet Sharma 0001, Sung-Ju Lee 0001, Jack Brassil, Kang G. Shin |
BROADNETS | 1 |
| 2004 | Distributed Channel Monitoring for Wireless Bandwidth Aggregation
Puneet Sharma 0001, Sung-Ju Lee 0001, Jack Brassil, Kang G. Shin |
NETWORKING | 1 |