VLDB 2026 Research / reviewers in the wild / expert
Sambit Sahu
dblp:76/157
· DBLP profile ↗
48ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 22Artificial intelligence and machine learning · 11 · 6 since 2021Systems, architecture and hardware · 7 · 2 first-authorDatabases, data management, data science and information retrieval · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 64% Deep learning architectures and training · 23% Planning, search and constraint satisfaction · 10% | |
| Human-computer interaction and pervasive computing
3 papers |
Ubiquitous computing and smart environments · 41% Health and well-being technologies · 29% User interface design and tools · 16% | |
| Computer networks
9 papers |
Network measurement and analytics · 26% Internet architecture and protocols · 25% Content delivery and video streaming · 19% | |
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Cloud and datacenter computing · 76% Distributed systems · 16% Memory systems · 7% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Smart cities and intelligent transportation · 92% Energy systems and smart grids · 8% |
Topics — the 30 heaviest of 60, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
model routing |
1.0 | 1 | 2026 | Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection · ACL (1) 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
skill estimation |
1.0 | 1 | 2026 | Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection · ACL (1) 2026 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.9 | 1 | 2025 | RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization · ICLR 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 1 | 2025 | Dense Backpropagation Improves Training for Sparse Mixture-of-Experts · NeurIPS 2025 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.9 | 1 | 2025 | RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization · ICLR 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization · ICLR 2025 |
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models |
0.9 | 1 | 2025 | T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning · NeurIPS 2025 |
Health and well-being technologies
behavior change |
0.3 | 2 | 2013 | The dubuque electricity portal: evaluation of a city-scale residential electricity consumption feedback system · CHI 2013 The dubuque water portal: evaluation of the uptake, use and impact of residential water consumption feedback · CHI 2012 |
Natural language and speech › Question answering and dialogue systems
conversational agents |
0.3 | 1 | 2025 | T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › transformer › transformer training
transformer pretraining |
0.3 | 1 | 2025 | Dense Backpropagation Improves Training for Sparse Mixture-of-Experts · NeurIPS 2025 |
Smart cities and intelligent transportation › urban mobility
urban mobility analysis |
0.2 | 1 | 2016 | Singapore in Motion: Insights on Public Transport Service Level Through Farecard and Mobile Data Analytics · KDD 2016 |
Ubiquitous computing and smart environments › smart buildings
energy consumption feedback |
0.2 | 1 | 2013 | The dubuque electricity portal: evaluation of a city-scale residential electricity consumption feedback system · CHI 2013 |
User interface design and tools
feedback systems |
0.2 | 1 | 2013 | The dubuque electricity portal: evaluation of a city-scale residential electricity consumption feedback system · CHI 2013 |
Collaborative and social computing › social computing
social comparison |
0.1 | 1 | 2012 | The dubuque water portal: evaluation of the uptake, use and impact of residential water consumption feedback · CHI 2012 |
Cloud and datacenter computing
cloud bursting |
0.1 | 1 | 2012 | Seagull: Intelligent Cloud Bursting for Enterprise Applications · USENIX ATC 2012 |
Smart cities and intelligent transportation
demand prediction |
0.1 | 1 | 2011 | Multi-granular demand forecasting in SmarterWater · UbiComp 2011 |
Data mining
pattern mining |
0.1 | 1 | 2011 | Activity analysis based on low sample rate smart meters · KDD 2011 |
Cloud and datacenter computing › elastic computing
cloud elasticity |
0.1 | 1 | 2011 | Kingfisher: Cost-aware elasticity in the cloud · INFOCOM 2011 |
Network performance modeling › performance prediction
latency prediction |
0.1 | 1 | 2010 | On suitability of Euclidean embedding for host-based network coordinate systems · IEEE/ACM Trans. Netw. 2010 |
Network measurement and analytics
network coordinate system |
0.1 | 1 | 2010 | On suitability of Euclidean embedding for host-based network coordinate systems · IEEE/ACM Trans. Netw. 2010 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.1 | 1 | 2010 | Change Management in Enterprise IT Systems: Process Modeling and Capacity-optimal Scheduling · INFOCOM 2010 |
Cloud and datacenter computing
resource management |
0.1 | 3 | 2012 | Seagull: Intelligent Cloud Bursting for Enterprise Applications · USENIX ATC 2012 Kingfisher: Cost-aware elasticity in the cloud · INFOCOM 2011 Impact of Load Sharing on Provisioning Services with Consistency Requirements · INFOCOM 2006 |
Distributed systems
peer-to-peer systems |
0.1 | 2 | 2004 | A lightweight, robust P2P system to handle flash crowds · IEEE J. Sel. Areas Commun. 2004 A Lightweight, Robust P2P System to Handle Flash Crowds · ICNP 2002 |
Cloud and datacenter computing › cloud service management
service provisioning |
0.1 | 1 | 2006 | Impact of Load Sharing on Provisioning Services with Consistency Requirements · INFOCOM 2006 |
Network measurement and analytics › traffic characterization
application traffic characterization |
0.1 | 1 | 2005 | Measurement-based Characterization of a Collection of On-line Games (Awarded Best Student Paper!) · Internet Measurement Conference 2005 |
Content delivery and video streaming
peer-to-peer content distribution |
0.1 | 1 | 2005 | Can unstructured P2P protocols survive flash crowds? · IEEE/ACM Trans. Netw. 2005 |
Internet architecture and protocols
multicast |
0.0 | 1 | 2004 | Scalability of Reliable Group Communication Using Overlays · INFOCOM 2004 |
Content delivery and video streaming
overlay multicast |
0.0 | 1 | 2004 | Scalability of Reliable Group Communication Using Overlays · INFOCOM 2004 |
Internet architecture and protocols
peer-to-peer networks |
0.0 | 1 | 2004 | A lightweight, robust P2P system to handle flash crowds · IEEE J. Sel. Areas Commun. 2004 |
Transport protocols and congestion control
rate control |
0.0 | 1 | 2004 | Scalability of Reliable Group Communication Using Overlays · INFOCOM 2004 |
Methods — techniques the papers use, named apart from their topics
synthetic data generation · 1.0LLM annotation · 1.0unified objective · 0.9top-k routing · 0.9large language model · 0.9exponential moving average · 0.9direct preference optimization · 0.9dense gradient approximation · 0.9caching · 0.9statistical framework · 0.4household behavior modeling · 0.4fixture characteristic modeling · 0.4survey · 0.3interviews · 0.3field study · 0.3mobile geolocation data analytics · 0.2joint telco-and-farecard learning · 0.2farecard data analytics · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert SelectionabstractTianyi Niu, Justin Chen, Genta Indra Winata, Shi-Xiong Zhang, Supriyo Chakraborty, Sambit Sahu, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianyi Niu, Justin Chih-Yao Chen, Genta Indra Winata, Supriyo Chakraborty, Sambit Sahu, Yue Zhang 0004, Elias Stengel-Eskin, Mohit Bansal |
ACL (1) | 6 |
| 2025 | SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation
Berkcan Kapusuzoglu, Supriyo Chakraborty, Renkun Ni, Stephen Rawls, Sambit Sahu |
IEEE Big Data | 5 |
| 2025 | RainbowPO: A Unified Framework for Combining Improvements in Preference OptimizationabstractRecently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully aligned models with human preferences, there is a lack of understanding regarding the contributions of their additional components. Moreover, fair and consistent comparisons are scarce, making it difficult to discern which components genuinely enhance downstream performance. In this work, we propose RainbowPO, a unified framework that demystifies the effectiveness of existing DPO methods by categorizing their key components into seven broad directions. We integrate these components into a single cohesive objective, enhancing the performance of each individual element. Through extensive experiments, we demonstrate that RainbowPO outperforms existing DPO variants. Additionally, we provide insights to guide researchers in developing new DPO methods and assist practitioners in their implementations. Hanyang Zhao, Genta Indra Winata, David D. Yao, Wenpin Tang, Sambit Sahu |
ICLR | 7 |
| 2025 | T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic PlanningabstractLarge Language Models (LLMs) have demonstrated impressive capabilities as intelligent agents capable of solving complex problems. However, effective planning in scenarios involving dependencies between API or tool calls-particularly in multi-turn conversations-remains a significant challenge. To address this, we introduce T1, a tool-augmented, multi-domain, multi-turn conversational dataset specifically designed to capture and manage inter-tool dependencies across diverse domains. T1 enables rigorous evaluation of agents' ability to coordinate tool use across nine distinct domains (4 single domain and 5 multi-domain) with the help of an integrated caching mechanism for both short- and long-term memory, while supporting dynamic replanning-such as deciding whether to recompute or reuse cached results. Beyond facilitating research on tool use and planning, T1 also serves as a benchmark for evaluating the performance of open-weight and proprietary large language models. We present results powered by T1-Agent highlighting their ability to plan and reason in complex, tool-dependent scenarios. Amartya Chakraborty, Paresh Dashore, Nadia Bathaee, Anmol Jain, Sambit Sahu, Milind R. Naphade, Genta Indra Winata |
NeurIPS | 7 |
| 2025 | Dense Backpropagation Improves Training for Sparse Mixture-of-ExpertsabstractMixture of Experts (MoE) pretraining is more scalable than dense Transformer pretraining, because MoEs learn to route inputs to a sparse set of their feedforward parameters. However, this means that MoEs only receive a sparse backward update, leading to training instability and suboptimal performance. We present a lightweight approximation method that gives the MoE router a dense gradient update while continuing to sparsely activate its parameters. Our method, which we refer to as Default MoE, substitutes missing expert activations with default outputs consisting of an exponential moving average of expert outputs previously seen over the course of training. This allows the router to receive signals from every expert for each token, leading to significant improvements in training performance. Our Default MoE outperforms standard TopK routing in a variety of settings without requiring significant computational overhead. Ashwinee Panda, Vatsal Baherwani, Zain Sarwar, Benjamin Thérien, Sambit Sahu, Tom Goldstein, Supriyo Chakraborty |
NeurIPS | 5 |
| 2025 | Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A SurveyabstractPreference tuning is a crucial process for aligning deep generative models with human preferences. This survey offers a thorough overview of recent advancements in preference tuning and the integration of human feedback. The paper is organized into three main sections: 1) introduction and preliminaries: an introduction to reinforcement learning frameworks, preference tuning tasks, models, and datasets across various modalities: language, speech, and vision, as well as different policy approaches, 2) in-depth exploration of each preference tuning approach: a detailed analysis of the methods used in preference tuning, and 3) applications, discussion, and future directions: an exploration of the applications of preference tuning in downstream tasks, including evaluation methods for different modalities, and an outlook on future research directions. Our objective is to present the latest methodologies in preference tuning and model alignment, enhancing the understanding of this field for researchers and practitioners. We hope to encourage further engagement and innovation in this area. Additionally, we provide a GitHub link https://github.com/hanyang1999/Preference-Tuning-with-Human-Feedback. Genta Indra Winata, Hanyang Zhao, Wenpin Tang, David D. Yao, Sambit Sahu |
J. Artif. Intell. Res. | 7 |
| 2016 | Singapore in Motion: Insights on Public Transport Service Level Through Farecard and Mobile Data AnalyticsabstractGiven the changing dynamics of mobility patterns and rapid growth of cities, transport agencies seek to respond more rapidly to needs of the public with the goal of offering an effective and competitive public transport system. A more data-centric approach for transport planning is part of the evolution of this process. In particular, the vast penetration of mobile phones provides an opportunity to monitor and derive insights on transport usage. Real time and historical analyses of such data can give a detailed understanding of mobility patterns of people and also suggest improvements to current transit systems. On its own, however, mobile geolocation data has a number of limitations. We thus propose a joint telco-and-farecard-based learning approach to understanding urban mobility. The approach enhances telecommunications data by leveraging it jointly with other sources of real-time data. The approach is illustrated on the First- and last-mile problem as well as route choice estimation within a densely-connected train network. Hasan Poonawala, Vinay Kolar, Sebastien Blandin, Laura Wynter, Sambit Sahu |
KDD | 5 |
| 2014 | cAppCloud: Contextual Personalized Application CloudabstractExtending ubiquitous data access to personalized applications stack, we propose and build a cloud based intelligent personalized application leveraging cloud to allow users access to their personalized applications from anywhere on any device. Our system is smart and intelligent in the sense that it adapts itself based on (i) available bandwidth between the user location and the cloud system, (ii) user profile that is automatically sensed and learned, (iii) user locations among other context profiles. First, we develop a framework for application virtualization that decouples an application from any specific platform or device and enables access to personalized application access at per user level. Our solution is context aware and learns the usage context based on the location and user profile and leverages this to minimize the download bandwidth requirement. Next we provide an implementation of this framework for applications on Windows platform leveraging Amazon S3 cloud storage-although Capp Cloud can be implemented for any platform with any other cloud systems. Given the era of Internet of Things (IoT) and various Cloud Enabled Intelligent Applications, we feel that Capp Cloud can meet various key requirements to facilitate the interesting scenarios. Sambit Sahu |
CloudCom | 2 |
| 2014 | Mobile App Connecting People Based on Personality Detection and Image Perception AnalysisabstractWe propose a personality detection algorithm based on human perceptions of images and created a mobile social application named Date gram to test the efficiency of our algorithm and connect people who share similar image perceptions. We detect user's Big Five personality based on their perceptual opinions on several images. Users in Date gram are matched to play visual-sentiment games, through which their connections based on personality similarity are further strengthened. Datagram utilizes techniques of image sentiment detection to analyze user's personality. Given the collected personal sentiment perception data, we proposed personality-mapping algorithm to categorize users into "Big Five" personality. After the Adjective-Noun-Pairs (ANPs) are extracted from each image, a personalized image-profile that consists of multiple pairs of ANPs and corresponding user decisions is generated. This profile is treated as the input of our algorithm, and a modified Latent Dirichlet Allocation (LDA) model is applied to extract personality-related features. "Big Five" personality values are calculated by a word-personality mapping algorithm where the technique of Linguistic Inquiry and Word Count (LIWC) is used. To evaluate our proposed algorithm, we did user studies over visual sentiments as well as personality detections. The results demonstrate that sentimental opinions from different persons vary significantly towards the same image, people would prefer to connect with those sharing similar perceptions of images and similar personalities with them. Ruichi Yu, Shuguan Yang, Guifan Li, Chao Qian 0005, Sambit Sahu, Ching-Yung Lin |
ISM | 5 |
| 2014 | Cost-Aware Cloud Bursting for Enterprise ApplicationsabstractThe high cost of provisioning resources to meet peak application demands has led to the widespread adoption of pay-as-you-go cloud computing services to handle workload fluctuations. Some enterprises with existing IT infrastructure employ a hybrid cloud model where the enterprise uses its own private resources for the majority of its computing, but then “bursts” into the cloud when local resources are insufficient. However, current commercial tools rely heavily on the system administrator’s knowledge to answer key questions such as when a cloud burst is needed and which applications must be moved to the cloud. In this article, we describe Seagull, a system designed to facilitate cloud bursting by determining which applications should be transitioned into the cloud and automating the movement process at the proper time. Seagull optimizes the bursting of applications using an optimization algorithm as well as a more efficient but approximate greedy heuristic. Seagull also optimizes the overhead of deploying applications into the cloud using an intelligent precopying mechanism that proactively replicates virtualized applications, lowering the bursting time from hours to minutes. Our evaluation shows over 100% improvement compared to naïve solutions but produces more expensive solutions compared to ILP. However, the scalability of our greedy algorithm is dramatically better as the number of VMs increase. Our evaluation illustrates scenarios where our prototype can reduce cloud costs by more than 45% when bursting to the cloud, and that the incremental cost added by precopying applications is offset by a burst time reduction of nearly 95%. Tian Guo 0001, Upendra Sharma, Prashant J. Shenoy, Timothy Wood 0001, Sambit Sahu |
ACM Trans. Internet Techn. | 5 |
| 2013 | The dubuque electricity portal: evaluation of a city-scale residential electricity consumption feedback systemabstractThis paper describes the Dubuque Electricity Portal, a city-scale system aimed at supporting voluntary reductions of electricity consumption. The Portal provided each household with fine-grained feedback on its electricity use, as well as using incentives, comparisons, and goal setting to encourage conservation. Logs, a survey and interviews were used to evaluate the user experience of the Portal during a 20-week pilot with 765 volunteer households. Although the volunteers had already made a wide range of changes to conserve electricity prior to the pilot, those who used the Portal decreased their electricity use by about 3.7%. They also reported increased understanding of their usage, and reported taking an array of actions - both changing their behavior and their electricity infrastructure. The paper discusses the experience of the system's users, and describes challenges for the design of ECF systems, including balancing accessibility and security, a preference for time-based visualizations, and the advisability of multiple modes of feedback, incentives and information presentation. Thomas Erickson, Ming Li 0009, Younghun Kim, Ajay Deshpande, Sambit Sahu, Tian Chao, Noi Sukaviriya, Milind R. Naphade |
CHI | 5 |
| 2013 | Heat pump detection from coarse grained smart meter data with positive and unlabeled learningabstractRecent advances in smart metering technology enable utility companies to have access to tremendous amount of smart meter data, from which the utility companies are eager to gain more insight about their customers. In this paper, we aim to detect electric heat pumps from coarse grained smart meter data for a heat pump marketing campaign. However, appliance detection is a challenging task, especially given a very low granularity and partial labeled even unlabeled data. Traditional methods install either a high granularity smart meter or sensors at every appliance, which is either too expensive or requires technical expertise. We propose a novel approach to detect heat pumps that utilizes low granularity smart meter data, prior sales data and weather data. In particular, motivated by the characteristics of heat pump consumption pattern, we extract novel features that are highly relevant to heat pump usage from smart meter data and weather data. Under the constraint that only a subset of heat pump users are available, we formalize the problem into a positive and unlabeled data classification and apply biased Support Vector Machine (BSVM) to our extracted features. Our empirical study on a real-world data set demonstrates the effectiveness of our method. Furthermore, our method has been deployed in a real-life setting where the partner electric company runs a targeted campaign for 292,496 customers. Based on the initial feedback, our detection algorithm can successfully detect substantial number of non-heat pump users who were identified heat pump users with the prior algorithm the company had used. Hongliang Fei, Younghun Kim, Sambit Sahu, Milind R. Naphade, Sanjay K. Mamidipalli, John Hutchinson |
KDD | 3 |
| 2012 | The dubuque water portal: evaluation of the uptake, use and impact of residential water consumption feedbackabstractThe Dubuque Water Portal is a system aimed at supporting voluntary reductions of water consumption that is intended to be deployed city-wide. It provides each household with fine-grained, near real time feedback on their water consumption, as well as using techniques like social comparison, weekly games, and news and chat to encourage water conservation. This study used logs, a survey and interviews to evaluate a 15-week pilot with 303 households. It describes the Portal's design, and discusses its adoption, use and impacts. The system resulted in a 6.6% decrease in water consumption, and the paper employs qualitative methods to look at the ways in which the Portal was (or wasn't) effective in supporting its users and enabling them to reduce their consumption. The paper concludes with a discussion of design implications for residential feedback systems, and possible engagement models. Thomas Erickson, Mark Podlaseck, Sambit Sahu, Tian Chao, Milind R. Naphade |
CHI | 3 |
| 2012 | Designing standard service offerings for reducing IT service management costabstractA significant portion of a typical Enterprise IT services budget is spent on steady state management operations known as ”operational expenses”. Recently ”Standard Services Offerings” based approaches are proposed to reduce such cost where only a few set of versions and configurations are used to build an Enterprise IT services platform. A key challenge in these approaches is the choice of appropriate base set of configurations that can be used as seed configurations. Our paper addresses the above problem by studying several methods towards choosing the right set of application stack with configurations that minimizes the cost of building the target services environment from this seed set of offerings. We present four different methods and use model driven simulation studies to compare the performance of these approaches. We validate our key findings by analyzing our approaches for a large Enterprise customer IT environment. Sanghwan Lee 0002, Sambit Sahu, Rajeev Puri |
NOMS | 2 |
| 2012 | Seagull: Intelligent Cloud Bursting for Enterprise Applications
Tian Guo 0001, Upendra Sharma, Timothy Wood 0001, Sambit Sahu, Prashant J. Shenoy |
USENIX ATC | 4 |
| 2011 | CloudNaaS: a cloud networking platform for enterprise applicationsabstractEnterprises today face several challenges when hosting line-of-business applications in the cloud. Central to many of these challenges is the limited support for control over cloud network functions, such as, the ability to ensure security, performance guarantees or isolation, and to flexibly interpose middleboxes in application deployments. In this paper, we present the design and implementation of a novel cloud networking system called CloudNaaS. Customers can leverage CloudNaaS to deploy applications augmented with a rich and extensible set of network functions such as virtual network isolation, custom addressing, service differentiation, and flexible interposition of various middleboxes. CloudNaaS primitives are directly implemented within the cloud infrastructure itself using high-speed programmable network elements, making CloudNaaS highly efficient. We evaluate an OpenFlow-based prototype of CloudNaaS and find that it can be used to instantiate a variety of network functions in the cloud, and that its performance is robust even in the face of large numbers of provisioned services and link/device failures. Theophilus Benson, Aditya Akella, Anees Shaikh, Sambit Sahu |
SoCC | 4 |
| 2011 | Trip analyzer through smartphone appsabstractBroad usage of Smartphones and mobile apps enables both individual trip summary and regional travel demand analysis. In this paper, we describe a trip analysis system, as part of smarter transit service. This trip analysis system consists of mobile apps and a centralized analyzer. It identifies the travel mode and purpose of the trips sensed by mobile devices, provides trip summaries and insights to mobile subscribers, and generates meaningful patterns to support traffic operation planning and transit system design. It is developed and deployed to the Smartphones of the volunteers in Dubuque, IA, to serve both the volunteers and the transit agencies. Preliminary evaluation has demonstrated the applicability of the design. Ming Li 0009, Sambit Sahu, Milind R. Naphade |
GIS | 3 |
| 2011 | Efficient Server Consolidation Considering Intra-Cluster TrafficabstractDue to the advancement of virtual machine technologies, large server farms or data centers consolidate their servers into a set of more powerful machines to increase the flexibility, agility, and adaptability to the fluctuating resource demands. Furthermore, managing a small number of powerful machines rather than a large number of low-performing machines decreases the management costs. In such a consolidation procedure, the network traffic as well as the usage of CPU and memory should be taken into account because the network bandwidth is also a critical resource. Since the network traffic of two servers or applications can be eliminated when they are consolidated into the same machine, traditional binpacking algorithms for the consolidation may not work efficiently. In this paper, we propose a cluster based consolidation algorithm, which reduces the maximum resource usage ratio of the machines. Through extensive simulation analysis, we show that the network traffic based clustering approach achieves a near optimal maximum resource usage, that is, the usage is close to the simple lower bound that only considers CPU and memory usage. Sanghwan Lee 0002, Sambit Sahu |
GLOBECOM | 2 |
| 2011 | Multi-granular demand forecasting in SmarterWaterabstractIn this paper we describe the multi-resolution water consumption prediction based on the SmarterWater system. This prediction service provides household consumption projection and regional demand forecasting for both short-term and med-term. Water consumption prediction, together with the other functions in the SmarterWater service, has been deployed to Dubuque, IA. Consumption behavior change after accessing the service has been observed. Ming Li 0009, Sambit Sahu, Milind R. Naphade, Feng Chen 0001 |
UbiComp | 3 |
| 2011 | Mesh Router Placement Exploiting Obstacles for Mitigating InterferenceabstractThe performance (such as throughput) of wireless mesh networks (WMNs) are affected by many factors such as routing protocols and interferences. To improve the throughput, several studies have proposed efficient routing protocols to find better paths and new link state metrics to accurately represent the wireless link qualities. Most of these works assume that the locations of the mesh routers are pre-defined. In this paper, we propose a mesh router placement scheme based on genetic algorithm accounting for dynamic location for all the mesh routers, and indoor obstacles that can minimize the interferences among the mesh routers. Through extensive simulations, we find that the existence of obstacles such as walls can be exploited to decrease the interferences to improve the performance of the indoor wireless mesh networks. We show that the proposed scheme improves the performance by 30-40% compared to the random selection scheme. Furthermore, the existence of walls improves the performance by as much as 50% compared to the cases without walls. Sanghwan Lee 0002, Youngjun Lee, Min Sun Jeong, Sambit Sahu |
ICCCN | 4 |
| 2011 | A Cost-Aware Elasticity Provisioning System for the CloudabstractIn this paper we present Kingfisher, a cost-aware system that provides efficient support for elasticity in the cloud by (i) leveraging multiple mechanisms to reduce the time to transition to new configurations, and (ii) optimizing the selection of a virtual server configuration that minimizes the cost. We have implemented a prototype of Kingfisher and have evaluated its efficacy on a laboratory cloud platform. Our experiments with varying application workloads demonstrate that Kingfisher is able to (i) decrease the cost of virtual server resources by as much as 24% compared to the current cost-unaware approach, (ii) reduce by an order of magnitude the time to transition to a new configuration through multiple elasticity mechanisms in the cloud, and (iii), illustrate the opportunity for design alternatives which trade-off the cost of server resources with the time required to scale the application. Upendra Sharma, Prashant J. Shenoy, Sambit Sahu, Anees Shaikh |
ICDCS | 3 |
| 2011 | Kingfisher: Cost-aware elasticity in the cloudabstractIn this paper we present Kingfisher, a cost-aware system that provides efficient support for elasticity in the cloud by (i) leveraging multiple mechanisms to reduce the time to transition to new configurations, and (ii) optimizing the selection of a virtual server configuration that minimizes the cost. We have implemented a prototype of Kingfisher and have evaluated its efficacy on a laboratory cloud platform. Our experiments with varying application workloads demonstrate that Kingfisher is able to (i) decrease the cost of virtual server resources by as much as 24% compared to the current cost-unaware approach, (ii) reduce by an order of magnitude the time to transition to a new configuration through multiple elasticity mechanisms in the cloud, and (iii), illustrate the opportunity for further design alternatives which trade-off the cost of server resources with the time required to scale the application. Upendra Sharma, Prashant J. Shenoy, Sambit Sahu, Anees Shaikh |
INFOCOM | 3 |
| 2011 | Activity analysis based on low sample rate smart metersabstractActivity analysis disaggregates utility consumption from smart meters into specific usage that associates with human activities. It can not only help residents better manage their consumption for sustainable lifestyle, but also allow utility managers to devise conservation programs. Existing research efforts on disaggregating consumption focus on analyzing consumption features with high sample rates (mainly between 1 Hz ~ 1MHz). However, many smart meter deployments support sample rates at most 1/900 Hz, which challenges activity analysis with occurrences of parallel activities, difficulty of aligning events, and lack of consumption features. We propose a novel statistical framework for disaggregation on coarse granular smart meter readings by modeling fixture characteristics, household behavior, and activity correlations. This framework has been implemented into two approaches for different application scenarios, and has been deployed to serve over 300 pilot households in Dubuque, IA. Interesting activity-level consumption patterns have been identified, and the evaluation on both real and synthetic datasets has shown high accuracy on discovering washer and shower. Feng Chen 0001, Bingsheng Wang, Sambit Sahu, Milind R. Naphade, Chang-Tien Lu |
KDD | 4 |
| 2011 | Smarter Water Management: A Challenge for Spatio-Temporal Network Databases
KwangSoo Yang, Shashi Shekhar 0001, Sambit Sahu, Milind R. Naphade |
SSTD | 4 |
| 2010 | Regional behavior change detection via local spatial scanabstractRegional human behavior change refers to the scenarios that people in a certain area exhibit significant behavior deviation from their neighbors and their own past. This regional pattern usually reveals underlying changes of living environment, such as regional development, immigration, disease breakout; or uncovers demographic information from special events, for instance, start/end of school holidays, or religious holidays. Statistically significant behavior changes contain both temporal and spatial characteristics. In this paper, we propose local spatial scan statistic to identify regional behavior changes. To accelerate local search, spatial index is modified to provide data-driven clusters and scalable data access. Base on the restricted spatial index, we provide both exact and approximated approaches to compute local spatial scan. Simulation analysis and case studies on water bills of 15K households validated the efficiency and effectiveness of these approaches on identifying regional behavior changes. Feng Chen 0001, Sambit Sahu, Milind R. Naphade |
GIS | 3 |
| 2010 | Core Tree Optimization in Hybrid Peer to Peer Real Time Broadcasting SystemabstractPeer to peer based live multimedia streaming systems over the Internet have been gaining popularity these days because they can be easily deployed without router based IP multicast capabilities. However, the peers frequently join and leave the system so that some peers may experience service disruptions, which degrade the overall quality and reputation of the services. In this situation, a hybrid approach is considered as a viable alternative. In the hybrid approach, the service providers deploy multiple proxies all over the Internet and make them constitute a core tree in the peer to peer system so that the proxies behave as stable nodes that are always on. Thus the system quality can increase. In this paper, we propose a centralized core tree construction scheme based on various optimization algorithms. The scheme focuses on constructing a core tree that maximizes the overall quality of the system. Because the real time multimedia streaming systems have higher QoS requirements than the content distribution such as Akamai, we take into account not only the locations of the proxies but also the topology formed by the proxies. The general framework can obtain a tree for arbitrary performance metrics not only for a specific metric such as minimum delay. Through simulations over diverse scenarios, we show that our optimization based approaches can provide a better core tree than some greedy heuristic approaches. Sanghwan Lee 0002, Sambit Sahu |
GLOBECOM | 2 |
| 2010 | Optimized Hybrid Overlay Design for Real Time BroadcastingabstractLive multimedia streaming over the Internet has steadily gained popularity over the past decade primarily fueled by the growth in the available network bandwidth and rich multimedia applications. Various approaches to support live multimedia streaming can be broadly categorized into two alternatives, namely, IP multicast based and overlay based. While IP multicast based solution is dependent on the availability of IP multicast enabled network, peer-to-peer approach does not require any changes at the the network layer. However in a peer-to-peer based approach, peers frequently join and leave the system which often results in service disruptions, hence degradation of the overall quality and reputation of the services. In this situation, a hybrid approach is considered as a viable alternative where the service providers deploy multiple proxies all over the Internet and make them constitute a core tree in the peer-to-peer system. The proxies behave as stable nodes that are always in the system in spite of the frequent join and leave of the ordinary peers. In the hybrid approach, the main issue is the construction of the core tree that maximizes the overall quality of the system. The locations of the proxies as well as the topology formed by the proxies affect the performance due to the real time requirements in a live streaming. In this paper, we propose a centralized core tree construction scheme based on various optimization algorithms. Our proposed method provides a general framework to obtain a tree for arbitrary performance metrics. Furthermore, we show that our optimization based approaches can provide a better core tree than some greedy heuristic approaches such that the average delay from the source to each proxy decreases by 20-40%. Sanghwan Lee 0002, Sambit Sahu |
ICCCN | 2 |
| 2010 | Change Management in Enterprise IT Systems: Process Modeling and Capacity-optimal SchedulingabstractWe provide a formal model for the Change Management process for Enterprise IT systems, and develop change scheduling algorithms that seek to attain the "change capacity" of the system. The change management process handles critical updates in the system that often use overlapping sets of servers, resulting in scheduling conflicts between the corresponding change classes. Furthermore, applications are typically associated with certain permissible downtime windows, which impose constraints on the timing of the change executions. Scheduling of changes for such systems represent a complex dynamic optimization question. In a limiting fluid regime, where changes are assumed nonatomic, we develop a scheduling policy that provably attains the change capacity of the system. We then propose and evaluate an atomic approximation of the optimal fluid scheduling policy, which is well suited for application to a real change management system. Simulation results demonstrate that the expected change execution delay and the capacity attained by the approximate policy is close to the best attainable values, when unavoidable capacity losses due to fragmentation effects are taken into account and is significantly better than a randomized scheduling policy. Praveen Kumar Muthuswamy, Koushik Kar, Sambit Sahu, Prashant Pradhan, Saswati Sarkar |
INFOCOM | 3 |
| 2010 | Network distance based coordinate systems for P2P multimedia streamingabstractNetwork distance estimation through Euclidean embedding of Internet hosts has been extensively studied as a scalable approach. In this method, each Internet host is assigned a computed co-ordinate where the distance between any two arbitrary hosts is approximated by their Euclidean distance. In this paper, we investigate whether such co-ordinate based method can be leveraged in designing efficient overlay based routing path for peer-to-peer streaming. Specifically we explore the usage of co-ordinate based estimation with well known Delaunay Triangulation (DT) based routing path design for peer-to-peer streaming. We devise algorithms for this combined approach that we refer as e-DT. Using real measurement data sets, we show that e-DT improves the overlay based routing significantly over purely DT based approach. The comparison with other well known overlay based routing approaches indicates that e-DT is able to construct similar overlay paths with smaller average forwarding degree requirements at each peering node. In addition, e-DT can support both point-to-multi point and multi point to multi point streaming requirements. Sanghwan Lee 0002, Sambit Sahu |
NOMS | 2 |
| 2010 | Characterizing Online GamesabstractOnline games are a rapidly growing Internet application. In order to run a successful online game, game companies and game infrastructure providers must properly manage game workloads and content so that they can maximize player satisfaction while minimizing their own costs. Toward this end, this paper provides a comprehensive, long-term analysis of several popular online games and their players using one of the richest data sets available for online games. Chris Chambers, Wu-chang Feng, Sambit Sahu, Debanjan Saha, David Brandt |
IEEE/ACM Trans. Netw. | 3 |
| 2010 | On suitability of Euclidean embedding for host-based network coordinate systems
Sanghwan Lee 0002, Zhi-Li Zhang, Sambit Sahu, Debanjan Saha |
IEEE/ACM Trans. Netw. | 3 |
| 2007 | PDA: A Tool for Automated Problem Determination
Hai Huang 0002, Raymond B. Jennings III, Yaoping Ruan, Ramendra K. Sahoo, Sambit Sahu, Anees Shaikh |
LISA | 5 |
| 2007 | Fundamental Effects of Clustering on the Euclidean Embedding of Internet Hosts
Sanghwan Lee 0002, Zhi-Li Zhang, Sambit Sahu, Debanjan Saha, Mukund Srinivasan |
Networking | 3 |
| 2007 | Improving the resilience of content distribution networks to large scale distributed denial of service attacks
Kang-Won Lee 0002, Suresh Chari, Anees Shaikh, Sambit Sahu, Pau-Chen Cheng |
Comput. Networks | 4 |
| 2006 | Design and Evaluation of a Network Distance Based Planning ServiceabstractIn this paper, we design and evaluate the prototype of a network planning service utilizing the coordinate based embedding of network hosts. The kernel of our prototype consists of a scalable network distance embedding method, a core set of services built on top of this embedding, and a generic set of APIs exposed to the applications for utilizing these services. The implemented service core consists of four generic services that we argue to be common to a wide range of applications requiring management of services and monitoring of network distance among Internet hosts. The proposed service does not require any support from the end-hosts. We evaluate the implemented service core using real Internet data and demonstrate its efficacy. Our experience provides several key insights for the design and management of a suitable embedding scheme. Sanghwan Lee 0002, Sambit Sahu, Debanjan Saha |
GLOBECOM | 2 |
| 2006 | Impact of Load Sharing on Provisioning Services with Consistency Requirements
Daniel A. M. Villela, Vishal Misra, Dan Rubenstein, Sambit Sahu |
INFOCOM | 4 |
| 2006 | An observation-based approach towards self-managing web servers
Abhishek Chandra, Prashant Pradhan, Renu Tewari, Sambit Sahu, Prashant J. Shenoy |
Comput. Commun. | 4 |
| 2005 | Last mile problem in overlay designabstractPerformance of overlay networks is dependent on last-mile connections, since they require that data traverse these last-mile bottlenecks at each forwarding step. This requires several times more upstream bandwidth than downstream, further exaggerating the asymmetry between down-stream and upstream bandwidth in last-mile technologies. This imbalance can cause packet queuing at the outgoing network interface of forwarding nodes, increasing latency and causing packet losses. We describe a model of a last-mile constrained overlay network and formulate and use it to solve a simplified latency- and bandwidth-bounded overlay construction problem. We observe that queueing delay may be a significant component of the end-to-end delay and approaches ignoring this may potentially result in an overlay network violating the delay and/or loss bounds. We observe that allowing a small amount of loss, it is possible to support a significantly large number of nodes. For a given end to end delay and loss bound we identify feasible degree (fan out) of each nodes. Our study sheds insights which provide engineering guidelines for designing overlays accounting for last mile problem in the Internet. Parijat Dube, Zhen Liu 0001, Sambit Sahu, Jeremy Silber |
GLOBECOM | 3 |
| 2005 | Protecting content distribution networks from denial of service attacksabstractIn this paper, we develop two mechanisms to detect DoS attacks against CDN-hosted Web sites and CDN infrastructure servers. First, we propose a novel request routing algorithm which allows CDN servers to effectively distinguish attacks from legitimate requests. Our scheme, based on a keyed hash function, significantly improves the resilience of servers to DoS attacks. Second, we introduce several site allocation algorithms based on binary codes which insure that an attack on one hosted Web site has a limited impact on other hosted sites. Our scheme guarantees that a specified minimum number of servers remain available for non-victimized sites. Together, the proposed schemes significantly improve the resilience of CDN-hosted Web sites, and complement other work on countering distributed DoS attacks. Kang-Won Lee 0002, Suresh Chari, Anees Shaikh, Sambit Sahu, Pau-Chen Cheng |
ICC | 4 |
| 2005 | Measurement-based Characterization of a Collection of On-line Games (Awarded Best Student Paper!)
Chris Chambers, Wu-chang Feng, Sambit Sahu, Debanjan Saha |
Internet Measurement Conference | 3 |
| 2005 | Can unstructured P2P protocols survive flash crowds?abstractToday's Internet periodically suffers from hot spots, a.k.a., flash crowds. A hot spot is typically triggered by an unanticipated news event that triggers an unanticipated surge of users that request data objects from a particular site, temporarily overwhelming the site's delivery capabilities. During this time, the large majority of users that attempt to get these objects face the frustrating experience of not being able to retrieve the content they want while still being able to communicate effectively with all other parts of the network. In this paper, we examine whether simple, undirected peer-to-peer search protocols can be used as a backup to deliver content whose popularity suddenly spikes. We model a simple, representative, undirected peer-to-peer search protocol in which clients cache only those objects they have explicitly requested. Because the object that becomes hot initially has limited popularity, the number of cache points, were they to remain fixed, would be insufficient to handle the level of demand during the flash crowd. However, as searches complete, more copies of the object become available. We analyze this natural scaling phenomenon and show that during the flash crowd, copies are distributed to requesting clients at a fast enough rate such that these simple protocols can indeed be used to scalably retrieve content that suddenly becomes "hot". Dan Rubenstein, Sambit Sahu |
IEEE/ACM Trans. Netw. | 2 |
| 2004 | Augmenting overlay trees for failure resiliencyabstractOverlay trees typically use directed trees as efficient structures for disseminating information, but their single-path structure means that just one node failure results in the disconnection of all descendants, possibly a significant portion of the graph. The addition of extra "backup" links to a directed tree can provide alternate data paths that significantly reduce the number of nodes disconnected when some set of nodes are removed from the graph. We investigate several deterministic and randomized algorithms for adding such backup links to a directed tree and analyze the connectedness of the resulting graphs when nodes in the network fail with some random probability. We present closed-form approximations and simulation measurements for the connectivity of these augmented trees in networks ranging from hundreds to hundreds of thousands of nodes. We also identify and measure the costs of adding backup links, using simulations and real-world measurements from overlays constructed using PlanetLab latency data. We find that, with node failure rates up to 10%, deterministic backup link selection policies offer comparable resiliency to random backup links with significantly lower overhead and resource usage. Jeremy Silber, Sambit Sahu, Jatinder Singh, Zhen Liu 0001 |
GLOBECOM | 2 |
| 2004 | Scalability of Reliable Group Communication Using OverlaysabstractThis study provides some new insights into the scalability of reliable group communication mechanisms using overlays. These mechanisms use individual TCP connections for packet transfers between end-systems. End-systems store incoming packets and forward them to downstream nodes using different unicast TCP connections. In this paper we assume that buffers in end-systems are large enough for the transfers. It is shown that the throughput of the reliable overlay group communication scales in the sense that for all multicast tree sizes and topologies, the group throughput is strictly positive under natural conditions. This is in contrast with the IP supported multicast paradigm where reliable protocols have vanishing throughput when the group size tends to infinity. The scalability of packet delay and buffer occupancy is then investigated. In the absence of additional control, the occupancy of the buffer and the latency in the end-systems explodes with time. It is then shown that proactive rate throttle mechanism implemented at the source leads to finite packet latency and buffer occupancy in any end-system of the network provided certain moment conditions are satisfied by cross traffic in the routers. François Baccelli, Augustin Chaintreau, Zhen Liu 0001, Anton Riabov, Sambit Sahu |
INFOCOM | 5 |
| 2004 | A lightweight, robust P2P system to handle flash crowdsabstractAn Internet flash crowd (also known as hot spots) is a phenomenon that results from a sudden, unpredicted increase in an on-line object's popularity. Currently, there is no efficient means within the Internet to deliver Web objects scalably under hot spot conditions to all clients that desire the object. We present peer-to-peer (P2P) randomized overlays to obviate flash-crowd symptoms (PROOFS), a simple, lightweight, P2P approach that uses randomized overlay construction and randomized, scoped searches to locate and deliver objects efficiently under heavy demand to all users that desire them. We evaluate PROOFS' robustness in environments in which clients join and leave the P2P network, as well as in environments in which clients are not always fully cooperative. Through a mix of simulation and prototype experimentation in the Internet, we show that randomized approaches like PROOFS should effectively relieve flash crowd symptoms in dynamic, limited-participation environments. Angelos Stavrou, Dan Rubenstein, Sambit Sahu |
IEEE J. Sel. Areas Commun. | 3 |
| 2002 | A Lightweight, Robust P2P System to Handle Flash CrowdsabstractInternet flash crowds (a.k.a. hot spots) are a phenomenon that result from a sudden, unpredicted increase in an on-line object's popularity. Currently, there is no efficient means within the Internet to scalably deliver Web objects under hot spot conditions to all clients that desire the object. We present PROOFS: a simple, lightweight, peer-to-peer (P2P) approach that uses randomized overlay construction and randomized, scoped searches to efficiently locate and deliver objects under heavy demand to all users that desire them. We evaluate PROOFS' robustness in environments in which clients join and leave the P2P network as well as in environments in which clients are not always fully cooperative. Through a mix of simulation and prototype experimentation in the Internet, we show that randomized approaches like PROOFS should effectively relieve flash crowd symptoms in dynamic, limited-participation environments. Angelos Stavrou, Dan Rubenstein, Sambit Sahu |
ICNP | 3 |
| 2002 | On the sensitivity of cooperative caching performance to workload and network characteristicsabstractA rich body of literature exists on several aspects of cooperative caching [1, 2, 3, 4, 5], including object placement and replacement algorithms [1], mechanisms for reducing the overhead of cooperation [2, 3], and the performance impact of cooperation [3, 4, 5]. However, while several studies have focused on quantifying the performance benefit of cooperative caching, their conclusions on the effectiveness of such cooperation vary significantly. The source of this apparent disagreement lies mainly in their different assumptions about workload and network characteristics, and about the degree of cooperation among caches.To more comprehensively evaluate the practical benefit of cooperative caching, we explore the sensitivity of the benefit of cooperation to workload characteristics such as object popularity distribution, temporal locality, one time referencing behavior, and to network characteristics such as latencies between clients, proxies, and servers. Furthermore, we identify a critical workload characteristic, which we call average access density, and show that it has a crucial impact on the effectiveness of cooperative caching.In this extended abstract, we report on a few important results selected from our extensive study reported in [6]. In particular, assuming an LFU-based cache management policy, we arrive at the following conclusions. First, cooperative caching is only effective when the average access density (defined as the ratio of the number of requests to the number of distinct objects in a time window) is relatively high. Second, the effectiveness of cooperative caching decreases as the skew in object popularity increases. Higher skew means that only a small number of objects are most frequently accessed reducing the benefit of larger caches, and therefore of cooperation. Khalil Amiri, Sambit Sahu, Chitra Venkatramani |
SIGMETRICS | 3 |
| 2000 | On achievable service differentiation with token bucket marking for TCPabstractThe Differentiated services (diffserv) architecture has been proposed as a scalable solution for providing service differentiation among flows without any per-flow buffer management inside the core of the network. It has been advocated that it is feasible to provide service differentiation among a set of flows by choosing an appropriate “marking profile” for each flow. In this paper, we examine (i) whether it is possible to provide service differentiation among a set of TCP flows by choosing appropriate marking profiles for each flow, (ii) under what circumstances, the marking profiles are able to influence the service that a TCP flow receives, and, (iii) how to choose a correct profile to achieve a given service level. We derive a simple, and yet accurate, analytical model for determining the achieved rate of a TCP flow when edge-routers use “token bucket” packet marking and core-routers use active queue management for preferential packet dropping. From our study, we observe three important results: (i) the achieved rate is not proportional to the assured rate, (ii) it is not always possible to achieve the assured rate and, (iii) there exist ranges of values of the achieved rate for which token bucket parameters have no influence. We find that it is not easy to regulate the service level achieved by a TCP flow by solely setting the profile parameters. In addition, we derive conditions that determine when the bucket size influences the achieved rate, and rates that can be achieved and those that cannot. Our study provides insight for choosing appropriate token bucket parameters for the achievable rates. Sambit Sahu, Philippe Nain, Christophe Diot, Victor Firoiu, Don Towsley |
SIGMETRICS | 1 |
| 2000 | Traffic models and admission control for variable bit rate continuous media transmission with deterministic service
Sambit Sahu, Victor Firoiu, Don Towsley, James F. Kurose |
Perform. Evaluation | 1 |