VLDB 2026 Research / reviewers in the wild / expert
Diwakar Krishnamurthy
dblp:36/3062
· DBLP profile ↗
60ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0002-6098-4801ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 20 · 8 since 2021Software engineering, systems software and programming languages · 13 · 2 first-authorHuman-computer interaction and ubiquitous computing · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Systems, architecture and hardware · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Detecting Intent in XR for Nonspeaking Autistic TypersabstractNonspeaking autistic individuals often build expressive communication through spelling on physical letterboards, but skilled one-on-one support is difficult to scale. XR letterboards could expand access, yet commodity near-hand press recognizers are brittle under the motor strategies common in nonspeaking typers (e.g., brief pokes, glides, folded-finger taps), leading to missed interactions that disrupt fluency and trust. We present SAGE (Spatial Anticipation of Gestural Events), a lightweight intent-inference layer that recovers intended selections even when no device press is registered. SAGE first fits a per-user, per-key spatial prior from historical device-registered taps, then fuses this prior at runtime with fingertip proximity and wrist-anchored pointing, using temporal peak detection to infer both press timing and letter identity. We evaluate SAGE on a HoloLens 2 AR letterboard dataset from four nonspeaking participants (244,620 hand-tracking samples) and compare against the headset recognizer and two physics-based replay baselines. On known taps, SAGE improves attribution quality (F1=0.91 vs. best baseline 0.52); on missed attempts, it recovers 71% of unregistered taps with correct letter identity for all recovered events. We contribute SAGE, the first dataset of nonspeaking XR typing trajectories at this granularity (released anonymized), and evidence that intent-aware modelling can make XR spelling more forgiving to motor diversity. Pratishtha, Ahmadreza Nazari, Kenzy Hamed, Lorans Alabood, Vikram Jaswal, Diwakar Krishnamurthy |
AVI | 6 |
| 2026 | Super-Resolution Meets Compression in Live VV: Light on Bits, Rich on QualityabstractVolumetric video enables fully immersive free-viewpoint experiences but imposes bandwidth and computation costs that prevent real-time deployment on everyday devices. Existing codecs achieve strong compression but struggle to meet live-streaming latency budgets, while super-resolution methods often depend on GPU acceleration or costly training. We present Super-VV, an end-to-end streaming framework that is light on bits yet rich on quality. Super-VV combines three key components: (i) a quality-preserving downsampler that prunes redundant geometry while retaining perceptual detail, (ii) a Fast Encoder/Decoder that extends Draco with tile-parallel and in-memory optimizations, and (iii) a lightweight, CPU-based upsampler that uses k-nearest-neighbors interpolation to restore dense geometry and color without GPUs or training. Together, these modules sustain real-time throughput of up to 73 FPS on commodity CPUs, while reducing bandwidth by 98% compared to raw transmission and maintaining compelling visual quality (PSNR of 42.75 dB) compared to the original volumetric videos. Sepehr Ganji, Amir Allahveran, Mea Wang, Diwakar Krishnamurthy |
MMSys | 4 |
| 2026 | Personalized Adaptive Virtual Object Placement in AR for Nonspeaking Autistic Users Using Behavioral CloningabstractNonspeaking autistic individuals (“nonspeakers”) represent about one-third of the autistic population, yet most lack access to an effective alternative to speech. This lack of effective communication significantly limits their access to educational, social, and employment opportunities. Some nonspeakers have learned to spell words and sentences by pointing to letters on a physical letterboard held in their field of view by a trained human assistant. While effective, this method relies on the assistant for positioning the letterboard, limiting user autonomy and privacy. We report here a system we developed that uses Behavioral Cloning (BC) to automatically and adaptively position a virtual letterboard in Augmented Reality (AR). By observing finger, palm, head, and physical letterboard poses during real-life interactions between a nonspeaker and their assistant, we train a BC Machine Learning (ML) model that can adapt the placement of a virtual letterboard for that user. Results from 11 experiments (3 emulated scenarios and 8 nonspeaking autistic participants) show that our approach can accurately replicate the actions of the human assistant of any given user, outperforming a non-ML baseline personalized placement policy in both positional and rotational accuracies. Further, our novel BC formulation overcomes traditional data-efficiency limitations, allowing us to achieve high accuracy with a modest training effort. This work represents a foundational step toward enabling more autonomous and private communication for nonspeakers. Ahmadreza Nazari, Lorans Alabood, Kaylyn B. Feeley, Vikram Jaswal, Diwakar Krishnamurthy |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2026 | vEdge: Flow-Based Network Slicing for Smart Cities in Edge Cloud Environments
Fekri Saleh, Abraham O. Fapojuwo, Diwakar Krishnamurthy |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2025 | Grab-and-Release Spelling in XR: A Feasibility Study for Nonspeaking Autistic People Using Video-Passthrough DevicesabstractThis paper investigates the feasibility of using video-passthrough Extended Reality (XR) devices to support communication in nonspeaking autistic individuals.Prior XR research with this population has relied on expensive augmented reality headsets and limited interactions to near-hand tapping.We present LetterBox, a novel application developed for video-passthrough XR headsets such as the Meta Quest series, which enables spelling through a "grabsnap-release" interaction.The app supports three immersion levels and includes a dynamic pass-through window tracking caregiver presence.We conducted a study with 19 participants across four Lorans Alabood, Ahmadreza Nazari, Travis Dow, Souad Alabood, Vikram Jaswal, Diwakar Krishnamurthy |
Conference on Designing Interactive Systems | 6 |
| 2025 | AR-Based Embodied Avatar Assistance for Nonspeaking Autistic People? Design and Feasibility StudyabstractMany nonspeaking autistic individuals rely on Communication and Regulation Partners (CRPs) to develop spelling-based communication using physical letterboards, but this support is often geographically inaccessible.We developed a remote presence system using Augmented Reality (AR) to enable immersive, collaborative spelling instruction.The system features holographic letterboards and fully embodied avatars with real-time head and hand tracking, allowing remote interaction between students and CRPs.In a study with 18 nonspeaking autistic participants, 15 (83%) successfully completed avatar-supported sessions.Interaction was higher, and participants reported a preference for the avatar condition over voice-only support.These findings demonstrate the feasibility of avatar-based AR telepresence for remote communication training.The system provides a demonstration of AR-supported interaction designed with nonspeaking autistic individuals-an underrepresented group in HCI-and offers design insights for inclusive telepresence technologies that address geographic and accessibility barriers. Travis Dow, Pratishtha, Lorans Alabood, Vikram Jaswal, Diwakar Krishnamurthy |
Conference on Designing Interactive Systems | 5 |
| 2025 | Enhancing Network Slice Identification in Beyond 5G: A Comparative Study of Machine Learning ApproachesabstractNetwork Slice Identification (NSI) is crucial for managing Quality of Service (QoS) in Beyond 5G (B5G) networks, particularly for smart city applications. This study explores and compares supervised, unsupervised, and semi-supervised learning techniques for NSI, addressing the challenges of limited labelled data in production environments. We use a publicly available 5G network dataset to model and perform comparisons among supervised, unsupervised, and semi-supervised learning approaches. Our methodology involves feature selection, dimensionality reduction using t-SNE, and addressing class imbalance through undersampling. We evaluate model performance using accuracy and Silhouette Score. Our results show a Random Forest Classifier achieves 100% accuracy with supervised learning. The unsupervised K-Means clustering model, optimized with both t-SNE and undersampling, achieves a mean accuracy of 92.83%. Semi-supervised learning using a self-training method and being trained on only 10% of the training data points performs comparably to the supervised models. Importantly, we demonstrate the robustness check using test data perturbation injecting additional variability in data simulating unknown 5G network fluctuations. This comparative analysis provides insights into the trade-offs between different learning approaches for NSI in B5G networks, offering practical solutions for scenarios with different conditions of labelled data. Sujesh Padhi, Diwakar Krishnamurthy, Abraham O. Fapojuwo |
NOMS | 2 |
| 2025 | SLA-Driven Performance Prediction for Co-Hosted DNNsabstractDeep Neural Networks (DNNs) are complex and versatile Machine Learning (ML) algorithms, essential to many systems. A single system can run multiple DNNs simultaneously, each handling different tasks to meet diverse user requests. Organizational DNN deployments must adhere to strict response time targets defined by Service Level Agreements (SLAs), as failure to meet these targets can incur significant costs. Therefore, accurately predicting performance under various workloads is crucial for optimizing resource allocation and avoiding SLA violations. Despite much research on DNNs, there is limited literature on predicting their performance across varying workload-resource configurations, with existing studies often neglecting important variables or lacking accuracy. Furthermore, predicting the performance of co-hosted DNNs that share resources is an understudied area. This paper addresses this gap by examining the response time behavior of co-hosted DNNs under various resource, workload, and DNN-related factors. We propose an ML-based performance modeling strategy to predict the likelihood of meeting predefined SLA-driven response time targets in co-hosted deployments. Our study identifies key factors related to the host system and the DNNs' isolated performance, enabling effective prediction without extensive historical data collection. This approach helps mitigate potential SLA violations proactively. Our model achieves an SLA compliance prediction accuracy of 90.2% and an f-score of 98.0%, demonstrating strong generalization to unseen workloads and resource configurations. Sarah Shah, Yasaman Amannejad, Diwakar Krishnamurthy |
NOMS | 3 |
| 2025 | Enabling Distance-Aware Real-Time Volumetric Video StreamingabstractLive, real-time volumetric video streaming enables immersive remote communication by transmitting 3D representations of participants, but faces significant bandwidth and computational challenges on consumer hardware. We introduce a distance-aware volumetric video streaming method that prunes point cloud data in real time based on both real-world and virtual viewing distances before any conventional compression or tiling is applied. Our prototype, RealityStream, uses a single Kinect v1 for capture, referencing a precomputed VMAF-driven distance table to guide adaptive downsampling for transmission. This effectively reduces bandwidth demand by up to 65% while maintaining acceptable visual fidelity. Implemented on a commodity laptop and streamed to a standalone VR headset, RealityStream is shown to be practical, achieving end-to-end latencies of around 75 ms. These findings highlight that taking into account both real and virtual distance is a low-complexity approach for efficient, real-time volumetric video capture and streaming. RealityStream will benefit existing volumetric streaming platforms, making real-time 3D communication more efficient over networks. Kyle Jorgensen, Mea Wang, Diwakar Krishnamurthy |
NOSSDAV | 3 |
| 2025 | Compression and Transmission of 8K Stereoscopic VR Using VAE-GAN Latents and Standard EncodersabstractDespite the heightened popularity of Virtual Re-ality (VR), streaming high-resolution stereoscopic VR remains a challenge. This is primarily due to significant bandwidth demands of high-definition VR content. While advanced deep neural networks (DNN s) have demonstrated the potential to outperform standard codecs, their integration into real-world transmission frameworks is complex and not directly compatible with current encoding standards. To bridge this gap, this paper proposes a novel technique for compressing and transmitting 8K stereoscopic scenes using Variational Autoencoder (VAE) GAN latents represented as 3-channel RGB scenes that can be transmitted via standard encoders. The proposed method reduces bandwidth requirements by 45.1 % across different 8K scenes while maintaining visual quality, highlighting the effectiveness of the approach. This study also investigates the impact of varying patch-sizes of input frames for model training and evaluate its influence on client-side reconstructions. We then explore various transmission configurations of latent frames. Our findings suggest that while residual transmission offers limited benefits for 3-channel latent frame compression, raw transmission consistently yields better results, particularly for texture-heavy scenes. To the best of our knowledge, this is the first such transmission study on 8K stereoscopic scenes for cloud-based VR, providing valuable insights for optimizing high-resolution VR streaming systems. Code: github.com/sampreetucalgary07/8K-VR-compression. Sampreet Vaidya, Hatem Abou-Zeid, Diwakar Krishnamurthy |
WCNC | 3 |
| 2025 | eSlice: Elastic Inter-Slice Resource Allocation for Smart City ApplicationsabstractNetwork slicing is a fundamental enabler for the advancement of fifth generation (5G) and beyond 5G (B5G) networks, offering customized service-level agreements (SLAs) for distinct slices such as enhanced mobile broadband (eMBB), massive machine-type communications (mMTC), and ultra-reliable low-latency communication (URLLC). However, smart city applications often require multiple slices concurrently, posing significant challenges in resource allocation, service isolation, and maintaining performance guarantees. This paper presents eSlice, an elastic inter-slice resource allocation mechanism specifically designed to address the dynamic requirements of smart city applications. eSlice organizes applications into hierarchical slices, leveraging cloud-native resource scaling to dynamically adapt to real-time demands. It integrates two novel algorithms: the Proactive eSlice Allocation Algorithm (PeSAA), which ensures the fair distribution of resources across the substrate network, and the Reactive eSlice Allocation Algorithm (ReSAA), which employs Multi-Agent Reinforcement Learning (MARL) to dynamically coordinate, reallocate, and recover unused resources as network conditions evolve. Experimental results demonstrate that eSlice significantly outperforms existing methods, achieving 94.3% resource utilization in simulation-based experiments under constrained urban-scale scenarios, providing a robust solution for dynamic resource management in 5G-enabled smart city networks. Fekri Saleh, Abraham O. Fapojuwo, Diwakar Krishnamurthy |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | Budget Aware Performance Test Selection for MicroservicesabstractThe microservice architecture is being increasingly used to build complex applications. This software architecture offers scalability, modularity, agility to development processes. However, ensuring optimal performance prior to deployment emerges as a significant challenge, especially within the fast-paced environments of Continuous Integration/Continuous Deployment (CI/CD) pipelines. Traditional performance testing processes are often reliant on synthetic scenarios and lengthy testing processes, and can be challenging to adopt in such an environment where testing needs to be both realistic and quick. Addressing this need, our paper proposes and implements an innovative framework that leverages real-world usage traces to identify and execute a small yet essential set of performance tests. This approach aims to seamlessly integrate with CI/CD workflows, offering developers quick feedback on performance issues and scalability constraints that may arise from changes to one or more microservices. Through a series of empirical evaluations, we compare the ability of different techniques in identifying a set of performance tests that can capture historically observed system behaviour and be executed within a specified time budget. In our paper, we successfully replicated the response time distribution of a 24-hour test on our custom testbench within just a five-minute test, achieving a relative percent error of only 0.96% and 1.40% at the 95thand 99thpercentile, respectively. This considerable decrease in time and resources necessary for load testing and response time modeling demonstrates the efficiency and effectiveness of our approach. Quinn Cooper, Diwakar Krishnamurthy, Yasaman Amannejad |
CLOUD | 2 |
| 2024 | A Comparative Analysis of Generative Adversarial Networks for Generating Cloud WorkloadsabstractWhile Generative Adversarial Networks (GANs) have been highly successful in areas such as image generation, their efficacy in generating time series data, specifically for cloud workload applications, is not yet very well-established. Several GAN architectures have been proposed for time series generation, however there is a lack of comprehensive comparative analysis among these models for different real-world datasets in cloud workload domain. Additionally, prior research has not thoroughly explored the performance of models in relation to dataset attributes, including length of the data sequences, their seasonality and stationarity. This paper bridges this gap by focusing on cloud work-load time series data. We compare TimeGAN, RGAN, TTS- GAN, and V-GAN architectures using three real-world trace datasets-Alibaba 2017, Alibaba 2018, and Azure-to evaluate their performance when applied to these datasets with diverse characteristics. We intend this study to be an empirical guide for practitioners and researchers to choose the most appropriate GAN model based on the unique characteristics of their time series data. In this paper, we introduced a way to employ existing statistical measures to preprocess and characterize the datasets from varying standpoints. Then we used these datasets to assess the quality of these models' outputs qualitatively and quantitatively with respect to diversity, fidelity, and usability, for each kind of the input data. Our findings revealed the capabilities and limitations of each model, with regards to data characteristics such as sequence length, seasonality and stationarity. Based on our results, TimeGAN and TTS-GAN emerged as top-performing models in general across different datasets and sequence lengths. TimeGAN showed superiority with capturing short term tem-poral dynamics, while TTS-GAN outperformed in capturing long term dependencies. The transformer-based architecture employed in TTS-GAN makes it adept for handling highly seasonal data across both short and long sequence lengths. Conversely, TimeGAN demonstrated superior performance in accurately capturing highly seasonal data over shorter periods. Niloofar Sharifisadr, Diwakar Krishnamurthy, Yasaman Amannejad |
CLOUD | 2 |
| 2024 | From Letterboards to Holograms: Advancing Assistive Technology for Nonspeaking Autistic Individuals with the HoloBoardabstractAbout one-third of autistic individuals are nonspeaking, i.e., they cannot use speech to convey their thoughts reliably. Many in this population communicate via spelling, a process in which they point to letters on a letterboard held upright in their field of view by a trained Communication and Regulation Partner (CRP). This paper focuses on transitioning such individuals to more independent, digital spelling that requires less support from the CRP, a goal most nonspeakers we consulted with desire. To enable this transition, we followed an approach that mimics an environment familiar to the nonspeaker and that harnesses the skills they already possess from physical letterboard training. Using this approach, we developed HoloBoard, a system that allows a nonspeaker, their CRP, and others, e.g., researchers, to share a common Augmented Reality (AR) environment containing a virtual letterboard. We configured the system to offer a brief (less than 10 minutes, on average) training module with graduated spelling tasks on the virtual letterboard. In a study involving 23 participants, 16 completed the entire module. These participants were able to spell words on the virtual letterboard without the CRP holding that board, an outcome we had not expected. When offered the opportunity to continue interacting with the virtual letterboard after the training module, 14 performed more complicated tasks than we had anticipated, spelling full sentences, or even offering feedback on the HoloBoard using solely the virtual board. Furthermore, five of these participants used the system solo, i.e., with the CRP and researchers absent from the virtual environment. These results suggest that training with the HoloBoard can lay the foundation for more independent communication, providing new social and educational opportunities for this marginalized population. Lorans Alabood, Travis Dow, Kaylyn B. Feeley, Vikram Jaswal, Diwakar Krishnamurthy |
CHI | 5 |
| 2024 | Optimization of B5G Network Slicing for Smart Cities Applications Using the MGBFS AlgorithmabstractThis paper introduces the matroid-based modified greedy breadth-first search (MGBFS) algorithm, a pioneering approach tailored for optimizing network slicing for smart city applications. It adeptly allocates virtual network functions and virtual links, considering constraints while optimizing for cost and latency considerations. This approach extends beyond the boundaries of the access network, offering a comprehensive solution spanning the edge, transport, and core network domains. Additionally, we propose a dual integer linear programming (ILP) optimization problem to jointly minimize the embedding cost and latency of the network. An extensive comparison of the MGBFS algorithm with an existing baseline model and ILP confirms the superiority of our proposed method. Joyeeta Rani Barai, Abraham O. Fapojuwo, Diwakar Krishnamurthy |
ICC | 3 |
| 2024 | TYLE: Tile-based Dynamic Quality Enhancement for 360-degree Video StreamingabstractIn recent years, live streaming of 360° videos has increased in popularity due to the emergence of virtual and mixed reality (VR/MR) applications. The high quality and bird-eye view characteristics of VR/MR pose real-time challenges for 360° video streaming. The tile-based 360° video streaming has been created based on Dynamic Adaptive Streaming over HTTP (DASH), mainly focusing on adaptation algorithms but ignoring the impact of coding properties and variations in video content on streaming quality. In this paper, we take on the challenge in a new direction through the exploration of the potential of CRF rate control. Our deep quality inspection of CRF transcoded videos lead to the proposal of a dynamic quality ladder for a more continuous quality provisioning in contrast to the conventional discrete quality levels. Our quality selection is based on visual quality and CRF rate control rather than just the resolution and bitrate in conventional DASH. We propose TYLE, a Tile-based dynamic quality enhancement for 360° video streaming. Without increasing the bandwidth demand, TYLE serves the content in the Field-of-View (FoV) at the highest quality level and content in the near-FoV region with improved visual quality compare to conventional DASH. TYLE maintains smoother and stabler playback, especially under the condition challenged by bandwidth under-provisioning and high motion-activity videos. Sonali Keshava Murthy Naik, Mea Wang, Diwakar Krishnamurthy |
IPCCC | 3 |
| 2024 | Personalizing an AR-based Communication System for Nonspeaking Autistic UsersabstractNonspeaking autistic individuals ("nonspeakers") represent about one-third of the autistic population, and most are never provided with an effective alternative to speech, hindering their educational, employment, and social opportunities. Some individuals have learned to spell words and sentences by pointing to letters on a physical letterboard held vertically in their field of view by a trained human assistant. While this method is effective, nonspeakers have expressed to us a desire to transition towards a more independent communication method that relies less on a human assistant, which would provide them with more autonomy and privacy. Augmented Reality (AR) based communication systems have the potential to address this objective. For example, an AR-based communication system can lessen the reliance on a human assistant by employing a virtual letterboard that is automatically and adaptively placed in a personalized manner that considers a given user’s unique motor skills and movement patterns. In this paper, we explore the use of Behavioural Cloning (BC) to derive such a personalized placement policy. Specifically, we observe finger, palm, head, and physical letterboard poses during real-life interactions between a nonspeaker and their assistant. These observations are then used to train a BC Machine Learning (ML) model that can adapt the placement of a virtual letterboard for that user within an AR environment. Results show that our approach can accurately replicate the actions of the human assistant of any given user, outperforming a non-ML baseline personalized placement policy in both positional and rotational accuracies. This work represents a foundational step toward enabling more autonomous and private communications for nonspeakers, thereby opening up new opportunities for them. Ahmadreza Nazari, Lorans Alabood, Kaylyn B. Feeley, Vikram Jaswal, Diwakar Krishnamurthy |
IUI | 5 |
| 2024 | Evaluating Gaze Interactions within AR for Nonspeaking Autistic UsersabstractNonspeaking autistic individuals often face significant inclusion barriers in various aspects of life, mainly due to a lack of effective communication means. Specialized computer software, particularly delivered via Augmented Reality (AR), offers a promising and accessible way to improve their ability to engage with the world. While research has explored near-hand interactions within AR for this population, gaze-based interactions remain unexamined. Given the fine motor skill requirements and potential for fatigue associated with near-hand interactions, there is a pressing need to investigate the potential of gaze interactions as a more accessible option. This paper presents a study investigating the feasibility of eye gaze interactions within an AR environment for nonspeaking autistic individuals. We utilized the HoloLens 2 to create an eye gaze-based interactive system, enabling users to select targets either by fixating their gaze for a fixed period or by gazing at a target and triggering selection with a physical button (referred to as a ‘clicker’). We developed a system called HoloGaze that allows a caregiver to join an AR session to train an autistic individual in gaze-based interactions as appropriate. Using HoloGaze, we conducted a study involving 14 nonspeaking autistic participants. The study had several phases, including tolerance testing, calibration, gaze training, and interacting with a complex interface: a virtual letterboard. All but one participant were able to wear the device and complete the system’s default eye calibration; 10 participants completed all training phases that required them to select targets using gaze only or gaze-click. Interestingly, the 7 users who chose to continue to the testing phase with gaze-click were much more successful than those who chose to continue with gaze alone. We also report on challenges and improvements needed for future gaze-based interactive AR systems for this population. Our findings pave the way for new opportunities for specialized AR solutions tailored to the needs of this under-served and under-researched population. Ahmadreza Nazari, Lorans Alabood, Molly Kay Rathbun, Vikram Jaswal, Diwakar Krishnamurthy |
VRST | 5 |
| 2023 | Interactive AR Applications for Nonspeaking Autistic People? - A Usability StudyabstractAbout one-third of autistic people are nonspeaking, and most are never provided access to an effective alternative to speech. Thoughtfully designed AR applications could provide members of this population with structured learning opportunities, including training on skills that underlie alternative forms of communication. A fundamental step toward creating such opportunities, however, is to investigate nonspeaking autistic people’s ability to tolerate a head-mounted AR device and to interact with virtual objects. We present the first study to examine the usability of an interactive AR-based application by this population. We recruited 17 nonspeaking autistic subjects to play a HoloLens 2 game we developed that involved holographic animations and buttons. Almost all subjects tolerated the device long enough to begin the game, and most completed increasingly challenging tasks that involved pressing holographic buttons. Based on the results, we discuss best practice design and process recommendations. Our findings contradict prevailing assumptions about nonspeaking autistic people and thus open up exciting possibilities for AR-based solutions for this understudied and underserved population. Ahmadreza Nazari, Ali Shahidi, Kate M. Kaufman, Julia E. Bondi, Lorans Alabood, Vikram Jaswal, Diwakar Krishnamurthy, Mea Wang |
CHI | 7 |
| 2023 | Predicting the Performance of DNNs to Support Efficient Resource AllocationabstractNumerous organizations are adopting sophisticated Machine Learning (ML) algorithms for their operations. To ensure the optimal performance of ML systems, organizations require insights into the response time of such systems under realistic user workloads. However, despite the widespread adoption of ML models, research on predicting the response time of a system serving an ML model under varying resources and user workloads is limited. In this paper, we address this gap by proposing a modeling approach to predict response times of multiple well-known Deep Neural Networks (DNNs) under simultaneously varying resource settings and user workloads. We join a classifier and a regressor to identify the optimal resource setting for meeting a DNN's response time target, and to predict the response time under the allocated resource setting. Our technique enables performance modeling without the need to collect extensive data during system operation, thus empowering pre-deployment predictions. The results demonstrate that our approach can generalize to unseen resource and workload scenarios, guaranteeing accurate predictions of compliance with response time targets 98.05% of the time and offering response time predictions with a mean prediction error of 9.10%. Sarah Shah, Yasaman Amannejad, Diwakar Krishnamurthy |
CNSM | 3 |
| 2023 | Transfer Learning for Online Prediction of Virtual Reality Cloud Gaming TrafficabstractCloud-based Virtual Reality (VR) gaming is gaining popularity to provide immersive experiences without requiring bulky hardware. However, managing network resources for these games is crucial to prevent subpar user experience and unnecessary costs. Predicting VR traffic patterns, such as video frame sizes can enable proactive network resource allocation and lead to improved quality of service (QoS). To this end, in this paper we first evaluate the efficacy of various Machine Learning (ML) models for predicting gaming traffic frame sizes using data collected from a real-world cloud-based VR game testbed. We then investigate the effectiveness of transfer learning (TL) in predicting frame size traffic patterns across different games and network conditions using an online learning method that we propose. The findings show that using the TL approach for online learning prediction can reduce overall traffic prediction error by up to 54%. Overall, this paper contributes to the understanding of cloud-based VR traffic patterns and can be of interest to developers, practitioners, and researchers interested in optimizing the performance of such systems. Sampreet Vaidya, Hatem Abou-Zeid, Diwakar Krishnamurthy |
GLOBECOM | 3 |
| 2023 | AR-Based Educational Software for Nonspeaking Autistic People - A Feasibility StudyabstractApproximately one-third of individuals with autism are nonspeaking: They cannot communicate effectively using speech. Some traditional accounts suggest that these individuals cannot talk because they lack the symbolic capacity for language. And yet, recent studies have shown that these individuals’ cognitive abilities are vastly underestimated by standardized tests, and that difficulties with motor skills and movement contribute to their difficulty with speech. One consequence of the traditional accounts of nonspeaking autism is that life skills (rather than academic content) tend to be emphasized in schooling. Without access to meaningful academic content, their educational and vocational opportunities are significantly limited. Recent studies have proposed the use of head-mounted Augmented Reality (AR) applications as a means of providing engaging, customizable, and age-appropriate content to this population. Specifically, such applications can address the unique sensory and motor needs of nonspeaking autistic students, e.g., allow them to move freely around the room as they interact with lessons in the application. This paper describes the design and evaluation of the first AR application aimed to facilitate tailored educational experiences for nonspeaking autistic students. After extensive consultations with nonspeaking people, parents, and professionals, we developed our application to run on HoloLens 2 offering lessons and multiple-choice comprehension and spelling questions. We conducted a study involving five nonspeaking autistic participants and two specialized educators. Through a design critique process and an iterative design refinement approach, we show that most of our participants successfully interacted with the application and completed different types of lesson tasks. Based on quantitative data from the study sessions and qualitative feedback from participants and educators, we provide recommendations for UI and UX design that will promote the development and use of such software for this under-served and under-researched population. Ali Shahidi, Lorans Alabood, Kate M. Kaufman, Vikram Jaswal, Diwakar Krishnamurthy, Mea Wang |
ISMAR | 5 |
| 2023 | Antifreeze: High-Quality Adaptive Live Streaming with Real-time TranscoderabstractThe demand for real-time video streaming is increasing due to emerging live and interactive applications like virtual conferencing/collaboration and augmented/mixed reality. Real-time video transcoding and streaming face challenges, as inefficient transcoding can cause delays and hinder viewer QoE. In this paper, we propose Antifreeze, a complete end-to-end solution for real-time transcoding and streaming. Antifreeze includes a transcoding-aware adaptation algorithm that considers visual quality, bandwidth, buffer dynamics, and transcoding time to maximize client QoE. By dynamically adapting through on-the-fly transcoding, Antifreeze provides a personalized streaming experience. Results demonstrate that Antifreeze reduces playback stalls and improves visual quality in live video streaming sessions across different bandwidth profiles. Asif Ali Mehmuda, Reza Hedayati Majdabadi, Mea Wang, Diwakar Krishnamurthy |
LCN | 4 |
| 2021 | Oasis: Performance Matching IoT System EmulationabstractInternet of Things (IoT) and its applications are proliferating. Performance and scalability are key aspects of an IoT system. A scalable IoT system can gracefully accommodate a growth in the number of IoT devices while still meeting performance requirements. This motivates the need for tools that allow the performance and scalability of an IoT system to be evaluated prior to deployment. Implementing the actual system at full scale and then evaluating it through measurements is ideal from the point of view of realism. However, such an approach can be expensive and inflexible. Emulating an IoT deployment on commodity hardware is an attractive alternative since it allows various system design alternatives to be evaluated in a flexible way without the need for expensive full scale implementations of the alternatives. Recently, many such IoT emulation platforms have been proposed. However, none of these platforms are appropriate for performance and scalability testing since they do not provide a mechanism to match the performance characteristics of the IoT devices being emulated. We present Oasis, a system that addresses this limitation. Oasis uses Docker containers to emulate IoT devices. It leverages Docker's resource allocation and network emulation capabilities to match the performance characteristics of an emulated device to that of its native counterpart. We also show that the ability to accurately match performance improves scalability as well. Navid Alipour, Mea Wang, Diwakar Krishnamurthy |
CLOUD | 3 |
| 2021 | Fast and Efficient Performance Tuning of MicroservicesabstractThe microservice architecture is being increasingly adopted. Microservices often rely on containerization technology, facilitating agile development and permitting flexible deployment on cloud platforms. Many microservice applications are interactive. Consequently, there is a need for pre-deployment performance tuning techniques to ensure that an application will meet its end user response time requirements post-deployment. Additionally, the tuning process should be efficient, i.e., allocate just enough resources to minimize costs in cloud-based deployments. Furthermore, the tuning process needs to be fast to facilitate agile deployments. We design and evaluate a technique called MOAT (Microservice Application Performance Tuner) that embodies these requiremenis. MOAT conducts iterative performance tests to determine resource allocations for the individual microservices in an application for any given workload. It exploits a novel optimization technique that identifies resource allocations while requiring only a limited number of performance tests to explore the tuning space. Validation using an experimental system shows that MOAT outperforms a competing approach based on Bayesian optimization in terms of both solution speed and resource allocation efficiency. Vahid MirzaEbrahim Mostofi, Diwakar Krishnamurthy, Martin F. Arlitt |
CLOUD | 2 |
| 2020 | RAD: Detecting Performance Anomalies in Cloud-based Web ServicesabstractWeb services hosted on public cloud platforms are often subjected to performance anomalies. Runtime detection of such anomalies is crucial for operations in cloud data centers. With ever-increasing data center size, complexities in software applications and dynamic traffic workload patterns, automatically detecting performance anomalies is a challenging task. In this paper, we propose RAD, a lightweight runtime anomaly detection technique that does not require application level instrumentation and can be easily implemented for detecting anomalies in multi-tier cloud-based Web services. In particular, we focus on anomalies that are difficult to detect by simply monitoring system level metrics alone, such as anomalies that are caused by contention from within a service and also those caused by shared resource contention by other services running on the cloud. RAD continuously monitors service resource metrics and uses a queuing network model to detect performance anomalies at runtime. Additionally, RAD uses historical data and implements a statistical methodology to diagnose the root cause of an anomaly. We evaluate RAD on a private cloud and also on the EC2 public cloud platform to show that RAD incurs extremely low levels of performance overhead on the service and is effective for detecting anomalies in both multi-tier monolithic services and microservices. Joydeep Mukherjee, Alexandru Baluta, Marin Litoiu, Diwakar Krishnamurthy |
CLOUD | 4 |
| 2020 | CONTRAST: Container-based Transcoding for Interactive Video StreamingabstractInteractive video streaming applications are becoming increasingly popular. To maintain the Quality of Experience (QoE) of an end user, interactive streaming platforms need to transcode a video stream, i.e., adapt the quality of the video content, to match the network conditions between the platform and the user as well as the device capabilities of the end user. Modern video codecs such as High Efficiency Video Coding (HEVC) require significant computational resources for transcoding operations. Consequently, there is a need for systems that can perform transcoding quickly at runtime to sustain the real-time performance required for interactive streaming while at the same time using just the right amount of computational resources for the transcoding operations. This paper addresses this need by designing and implementing CONTRAST, a Container- based Distributed Transcoding Framework for Interactive Video Streaming. For any given stream and transcoding resolution, CONTRAST exploits a profiling technique to automatically determine the degree of parallelism, Le., the number of processing cores, demanded by the transcoding process to sustain the stream’s frame rate. It then launches Docker containers configured with the required number of cores to perform the transcoding. Experiments using a set of realistic video streams show that CONTRAST is able to sustain the frame rate requirements for interactive streams in a more resource efficient manner compared to baseline techniques that do not consider the degree of parallelism. To the best of our knowledge, our paper is the first to establish best practices for implementing transcoding platforms for interactive streaming videos encoded using a modem video codec. Sajad Sameti, Mea Wang, Diwakar Krishnamurthy |
NOMS | 3 |
| 2020 | Contention Aware Web of Things Emulation TestbedabstractSince the advent of the Web, new Web benchmarking tools have frequently been introduced to keep up with evolving workloads and environments. The introduction of Web of Things (WoT) marks the beginning of another important paradigm that requires new benchmarking tools and testbeds. Such a WoT benchmarking testbed can enable the comparison of different WoT application configurations and workload scenarios under assumptions regarding WoT application resource demands and WoT device network characteristics. The powerful computational capabilities of modern commodity multicore servers along with the limited resource consumption footprints of WoT devices suggest the feasibility of a benchmarking testbed that can emulate the application behaviour of a large number of WoT devices on just a single multicore server. However, to obtain test results that reflect the true performance of the system being emulated, care must be exercised to detect and consider the impact of testbed bottlenecks on performance results. For example, if too many WoT devices are emulated then performance metrics obtained from a test run, e.g., WoT device response times, would only reflect contention among emulated devices for shared multicore server resources instead of providing a true indication of the performance of the WoT system being emulated. We develop a testbed that helps a user emulate a system consisting of multiple WoT devices on a single multicore server by exploiting Docker containers. Furthermore, we devise a novel mechanism for the user to check whether shared resource contention in the testbed has impacted the integrity of test results. Our solution allows for careful scaling of experiments and enables resource efficient evaluation of a wide range of WoT systems, architectures, application characteristics, workload scenarios, and network conditions. Raoufehsadat Hashemian, Niklas Carlsson, Diwakar Krishnamurthy, Martin F. Arlitt |
ICPE | 3 |
| 2020 | PRIMA: Subscriber-Driven Interference Mitigation for Cloud ServicesabstractNetwork services, e.g., video streaming services, are increasingly being deployed on public cloud platforms. Such services often employ horizontal scaling where a group of resource instances, e.g., virtual machines (VMs), handle incoming workload. The response time of such services is often affected by interference, i.e., contention among resource instances belonging to multiple cloud subscribers for shared cloud resources. Most commercial cloud platforms do not support built-in mechanisms to detect interference and mitigate its impact. Consequently, subscribers of such platforms, i.e., network service providers, need to deploy their own mechanisms to ensure a specified end user response time target is continuously met even in the face of fluctuations in workload and interference. This paper describes PRIMA, our implementation of such a mechanism. PRIMA uses automated and controlled performance tests to build models that capture the joint impact of workload and interference on the response time of each resource instance employed by a service. It adapts the system to changing workload and interference conditions by using these models at runtime to control the number of instances in the system and the distribution of load among these instances. Unlike existing subscriber-oriented interference mitigation techniques in literature, PRIMA guarantees that a subscriber-specified response time threshold is satisfied at every resource instance assigned to a service. Furthermore, in contrast to these approaches PRIMA can help a subscriber avoid using more instances than necessary by automatically selecting at runtime the least number of instances required for handling the observed workload and interference. We experimentally validate the effectiveness of PRIMA in both private and public cloud environments. Results show that PRIMA outperforms competing approaches proposed by us and others, including those that are commonly used in practice. They also reveal that PRIMA can automatically calibrate its models at runtime to account for any model prediction errors. Joydeep Mukherjee, Diwakar Krishnamurthy |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2019 | Fast and Lightweight Execution Time Predictions for Spark ApplicationsabstractUsers and operators of cloud-based Spark clusters often require quick insights on how the execution time of an application is likely to be impacted by the resources allocated to the application, e.g., the number of Spark executor cores assigned, and the size of the data to be processed. Existing techniques typically require extensive prior executions of the application under various resource allocation settings and data sizes to obtain an accurate model. In this paper, we explore the accuracy of a model with less prior executions of the application. Such a model can be useful for situations where quick predictions are required and little cluster resources are available for building a model. We use logs from two executions of an application with small sample data and different resource settings and explore the accuracy of the predictions for other resource allocation settings and input data sizes. Yasaman Amannejad, Sarah Shah, Diwakar Krishnamurthy, Mea Wang |
CLOUD | 3 |
| 2019 | Quick Execution Time Predictions for Spark ApplicationsabstractThe Apache Spark cluster computing platform is being increasingly used to develop big data analytics applications. There are many scenarios that require quick estimates of the execution time of any given Spark application. For example, users and operators of a Spark cluster often require quick insights on how the execution time of an application is likely to be impacted by the resources allocated to the application, e.g., the number of Spark executor cores assigned, and the size of the data to be processed. Job schedulers can benefit from fast estimates at runtime that would allow them to quickly conFigure a Spark application for a desired execution time using the least amount of resources. While others have developed models to predict the execution time of Spark applications, such models typically require extensive prior executions of applications under various resource allocation settings and data sizes. Consequently, these techniques are not suited for situations where quick predictions are required and very little cluster resources are available for the experimentation needed to build a model. This paper proposes an alternative approach called PERIDOT that addresses this limitation. The approach involves executing a given application under a fixed resource allocation setting with two different-sized, small subsets of its input data. It analyzes logs from these two executions to estimate the dependencies between internal stages in the application. Information on these dependencies combined with knowledge of Spark's data partitioning mechanisms is used to derive an analytic model that can predict execution times for other resource allocation settings and input data sizes. We show that deriving a model using just these two reference executions allows PERIDOT to accurately predict the performance of a variety of Spark applications spanning text analytics, linear algebra, machine learning and Spark SQL. In contrast, we show that a state-of-the-art machine learning based execution time prediction algorithm performs poorly when presented with such limited training data. Sarah Shah, Yasaman Amannejad, Diwakar Krishnamurthy, Mea Wang |
CNSM | 3 |
| 2019 | Container-based Real-time Video TranscodingabstractWith the ever growing popularity of video services, maintaining high Quality of Experience (QoE) of the end users with heterogeneous devices and network conditions is becoming more challenging. Each user requires content that matches with their device capability and network conditions. This motivates the need for flexible video transcoding, which enables changing the properties of videos on-the-fly to fit different users. However, the transcoding process is compute intensive especially when handling modern video coding standards such as High Efficiency Video Coding (HEVC) and supporting emerging applications such as live broadcasts. Consequently, there is a need for lightweight and resource-efficient systems that can perform transcoding quickly at real-time to sustain desired user QoE requirements. We design and implement a container-based video transcoding system to address this need. We experimentally show that our system can meet real-time transcoding while using less computational resources than a native transcoding approach. Our work also identifies container and transcoder parameters that can impact the overall performance of the proposed system. Sajad Sameti, Mea Wang, Diwakar Krishnamurthy |
LCN | 3 |
| 2018 | Stride: Distributed Video Transcoding in SparkabstractOn one hand, since the introduction of UHD (ultra-high definition) videos, e.g., 4K and 8K videos, it is becoming more resource and time intensive to transcode videos. On the other hand, the increasing demand for video streaming implies more videos need to be transcoded. These two facts motivate the need for techniques to speedup coding and transcoding time. In this paper, we propose Stride, the first distributed video transcoding system that leverages the Apache Spark big data platform. The design of Stride is transcoder agnostic, meaning it can adopt any transcoder implementation (e.g., FFMPEG) without any modification. We provide an experimental characterization of the impact of video transcoding and Spark configuration parameters to identify the optimal settings. We also compare Stride with competing approaches. Our results show that Stride achieves 3.27 times speedup when the computing power (i.e., the number of vCPUs in a cloud) is increased by a factor of 4, which is significantly higher than the other alternatives we explore. In particular, Spark's dynamic task scheduler allows Stride to reduce transcoding time by 19.86% compared to an implementation without Spark. Our benchmark study suggests that Stride can support transcoding from 4K to 1080p (full HD) at a rate matching the video bitrate using approximately only 24 virtual cores. Sajad Sameti, Mea Wang, Diwakar Krishnamurthy |
IPCCC | 3 |
| 2018 | Subscriber-Driven Cloud Interference Mitigation for Network ServicesabstractNetwork services, e.g., video streaming services, are increasingly being deployed on public cloud platforms. Such services often employ horizontal scaling where a group of resource instances, e.g., virtual machines (VMs), handle the incoming workload. The response time of such services is often affected by interference, i.e., contention among resource instances belonging to multiple cloud subscribers for shared cloud resources. Most commercial cloud platforms do not support built-in mechanisms to detect interference and mitigate its impact. This paper outlines a solution called PRIMA that subscribers of such platforms, i.e., network service operators, can deploy to ensure a specified end user response time target is met even in the face of fluctuations in workload and interference. PRIMA uses automated and controlled performance tests to build models that capture the joint impact of workload and interference on the response time of each resource instance employed by a service. PRIMA adapts the system to changing workload and interference conditions by using these models at runtime to control the number of instances in the system and the distribution of load among these instances. Unlike existing subscriber-oriented interference mitigation techniques in literature, PRIMA provides an explicit mechanism to guarantee that the specified response time threshold is met at every resource instance assigned to a service. Furthermore, in contrast to these approaches PRIMA can help an operator avoid using more instances than necessary for handling the observed workload and interference. Joydeep Mukherjee, Diwakar Krishnamurthy |
IWQoS | 2 |
| 2017 | IRIS: Iterative and Intelligent Experiment SelectionabstractBenchmarking is a widely-used technique to quantify the performance of software systems. However, the design and implementation of a benchmarking study can face several challenges. In particular, the time required to perform a benchmarking study can quickly spiral out of control, owing to the number of distinct variables to systematically examine. In this paper, we propose IRIS, an IteRative and Intelligent Experiment Selection methodology, to maximize the information gain while minimizing the duration of the benchmarking process. IRIS selects the region to place the next experiment point based on the variability of both dependent, i.e., response, and independent variables in that region. It aims to identify a performance function that minimizes the response variable prediction error for a constant and limited experimentation budget. We evaluate IRIS for a wide selection of experimental, simulated and synthetic systems with one, two and three independent variables. Considering a limited experimentation budget, the results show IRIS is able to reduce the performance function prediction error up to 4.3 times compared to equal distance experiment point selection. Moreover, we show that the error reduction can further improve through system-specific parameter tuning. Analysis of the error distributions obtained with IRIS reveals that the technique is particularly effective in regions where the response variable is sensitive to changes in the independent variables. Raoufehsadat Hashemian, Niklas Carlsson, Diwakar Krishnamurthy, Martin F. Arlitt |
ICPE | 3 |
| 2017 | Subscriber-Driven Interference Detection for Cloud-Based Web ServicesabstractWeb services are now increasingly being hosted on public cloud infrastructure as a service platforms such as the Amazon Web service elastic compute cloud (EC2). However, previous studies have shown that the virtualized infrastructure used in public clouds can introduce contention among virtual machines (VMs) for shared physical host resources eventually leading to performance problems. Subscribers in a public cloud platform typically do not have access to metrics that can directly quantify the adverse impact of such inter-VM interference on Web service response times. We present a software probe based system to address this limitation. The probe is a lightweight application that runs on each Web service VM that needs to be monitored. We periodically measure the probe's response time on a monitored VM. We then compare this response time with the probe's previously recorded baseline no-interference response time when it executes in isolation on a VM of the same type. Statistically significant increase in the probe's response time from the baseline is used to detect interference. The probe also indicates the type of contention at the physical host that causes the interference. This information can be exploited by a subscriber to mitigate the problem. Results show that our approach is quite effective over two different cloud platforms and a wide variety of workload scenarios. In particular, results indicate that Web service instances hosted on EC2 suffer from interference. Our probe was able to detect 93% of performance degradations triggered by such interference. In all these cases, the probe imposed an average overhead of only 3%-4% on the mean response time of the Web service being monitored. Joydeep Mukherjee, Diwakar Krishnamurthy, Mea Wang |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2016 | Predicting Web service response time percentilesabstractPredicting Web service response time percentiles is often an important aspect of service level management exercises. Existing techniques can be very time consuming since they involve the manual construction of complex analytic or simulation models. To address this problem, we propose Prospective, a fully automated and data-driven approach for predicting Web service response time percentiles. Prospective relies on historical response time data collected from a Web service. Given a specification for workload expected at the Web service over a planning horizon, Prospective uses this historical data to offer predictions for response time percentiles of interest. At the core of Prospective is a lightweight simulator that uses collaborative filtering to estimate response time behaviour of the service based on behaviour observed historically. Results show that Prospective is able to predict various response time percentiles of interest with high accuracy for a wide variety of workloads. Yasaman Amannejad, Diwakar Krishnamurthy, Behrouz Homayoun Far |
CNSM | 2 |
| 2015 | Detecting performance interference in cloud-based web servicesabstractWeb services have increasingly begun to rely on public cloud platforms. The virtualization technologies employed by public clouds can however trigger contention between virtual machines (VMs) for shared physical machine (PM) resources thereby leading to performance problems for the Web service. Past studies have exploited PM level performance metrics such as Clock Cycles Per Instruction to detect such platform induced performance interference. Unfortunately, public cloud customers do not have access to such metrics. They can typically only access VM-level metrics and application level metrics such as transaction response times and such metrics alone are often not useful for detecting inter-VM contention. This poses a difficult challenge to Web service operators for detecting and managing platform induced performance interference issues inside the cloud. We propose a machine learning based interference detection technique to address this problem. The technique applies collaborative filtering to predict whether a given transaction being processed by a Web service is suffering adversely from interference. The results can then be used by a management controller to trigger remedial actions, e.g., reporting problems to the system manager or switching cloud providers. Results using a realistic Web benchmark show that the approach is effective. The most effective variant of our approach is able to detect about 96% of performance interference events with almost no false alarms. Yasaman Amannejad, Diwakar Krishnamurthy, Behrouz Homayoun Far |
IM | 2 |
| 2015 | Managing Performance Interference in Cloud-Based Web ServicesabstractWeb services have increasingly begun to rely on public cloud platforms. The virtualization technologies employed by public clouds can, however, trigger contention between virtual machines (VMs) for shared physical machine resources, thereby leading to performance problems for Web services. Past studies have exploited physical-machine-level performance metrics such as clock cycles per instruction to detect such platform-induced performance interference. Unfortunately, public cloud customers do not have access to such metrics. They can only typically access VM-level metrics and application-level metrics such as transaction response times, and such metrics alone are often not useful for detecting inter-VM contention. This poses a difficult challenge to Web service operators for detecting and mitigating platform-induced performance interference issues inside the cloud. We propose a machine-learning-based interference detection technique to address this problem. The technique applies collaborative filtering to predict whether a given transaction being processed by a Web service is adversely suffering from interference. The results can be then used by a management controller to trigger remedial actions, e.g., reporting problems to the system manager or switching cloud providers. Results using a realistic Web benchmark show that the approach is effective. The most effective variant of our approach is able to detect about 96% of performance interference events with almost no false alarms. Furthermore, we show that a load redistribution technique that exploits the information from our detection technique is able to more effectively mitigate the interference than techniques that are interference agnostic. Yasaman Amannejad, Diwakar Krishnamurthy, Behrouz Homayoun Far |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2015 | Resource Contention Detection in Virtualized EnvironmentsabstractPublic and private cloud computing environments employ virtualization methods to consolidate application workloads onto shared servers. Modern servers typically have one or more sockets each with one or more computing cores, a multi-level caching hierarchy, a memory subsystem, and an interconnect to the memory of other sockets. While resource management methods may manage application performance by controlling the sharing of processing time and input-output rates, there is generally no management of contention for virtualization kernel resources or for the memory hierarchy and subsystems. Yet such contention can have a significant impact on application performance. Hardware platform specific counters have been proposed for detecting such contention. We show that such counters alone are not always sufficient for detecting contention. We propose a software probe based approach for detecting contention for shared platform resources and demonstrate its effectiveness. We show that the probe imposes low overhead and is remarkably effective at detecting performance degradations due to inter-VM interference over a wide variety of workload scenarios and on two different server architectures. The probe successfully detected virtualization-induced software bottleneck and memory contention on both server architectures. Our approach supports the management of workload placement on shared servers and pools of shared servers. Joydeep Mukherjee, Diwakar Krishnamurthy, Jerome A. Rolia |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2014 | Using Web Mining to Support Low Cost Historical Vehicle Traffic Analytics
Charanjeet Kaur, Diwakar Krishnamurthy, Behrouz Homayoun Far |
SEKE | 2 |
| 2014 | Characterizing the scalability of a Web application on a multi-core serverabstractSUMMARY The advent of multi‒core technology motivates new studies to understand how efficiently Web servers utilize such hardware. This paper presents a detailed performance study of a Web server application deployed on a modern eight‒core server. Our study shows that default Web server configurations result in poor scalability with increasing core counts. We study two different types of workloads, namely, a workload with intense TCP/IP related OS activity and the SPECweb2009 Support workload with more application‒level processing. We observe that the scaling behaviour is markedly different for these workloads, mainly because of the difference in the performance of static and dynamic requests. While static requests perform poorly when moving from using one socket to both sockets in the system, the converse is true for dynamic requests. We show that, contrary to what was suggested by previous work, Web server scalability improvement policies need to be adapted based on the type of workload experienced by the server. The results of our experiments reveal that with workload‒specific Web server configuration strategies, a multi‒core server can be utilized up to 80% while still serving requests without significant queuing delays; utilizations beyond 90% are also possible, while still serving requests with ‘acceptable’ response times. Copyright © 2014 John Wiley & Sons, Ltd. Raoufehsadat Hashemian, Diwakar Krishnamurthy, Martin F. Arlitt, Niklas Carlsson |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Resource contention detection and management for consolidated workloads
Joydeep Mukherjee, Diwakar Krishnamurthy, Jerome A. Rolia, Chris Hyser |
IM | 2 |
| 2013 | RPO: Runtime web server optimization under simultaneous multithreading
Samira Musabbir, Diwakar Krishnamurthy, Giuliano Casale |
IM | 2 |
| 2013 | Improving the scalability of a multi-core web serverabstractImproving the performance and scalability of Web servers enhances user experiences and reduces the costs of providing Web-based services. The advent of Multi-core technology motivates new studies to understand how efficiently Web servers utilize such hardware. This paper presents a detailed performance study of a Web server application deployed on a modern 2 socket, 4-cores per socket server. Our study show that default, "out-of-the-box" Web server configurations can cause the system to scale poorly with increasing core counts. We study two different types of workloads, namely a workload that imposes intense TCP/IP related OS activity and the SPECweb2009 Support workload, which incurs more application-level processing. We observe that the scaling behaviour is markedly different for these two types of workloads, mainly due to the difference in the performance characteristics of static and dynamic requests. The results of our experiments reveal that with workload-specific Web server configuration strategies a modern Multi-core server can be utilized up to 80% while still serving requests without significant queuing delays; utilizations beyond 90% are also possible, while still serving requests with acceptable response times. Raoufehsadat Hashemian, Diwakar Krishnamurthy, Martin F. Arlitt, Niklas Carlsson |
ICPE | 2 |
| 2013 | Performance models of storage contention in cloud environmentsabstractWe propose simple models to predict the performance degradation of disk requests due to storage device contention in consolidated virtualized environments. Model parameters can be deduced from measurements obtained inside Virtual Machines (VMs) from a system where a single VM accesses a remote storage server. The parameterized model can then be used to predict the effect of storage contention when multiple VMs are consolidated on the same server. We first propose a trace-driven approach that evaluates a queueing network with fair share scheduling using simulation. The model parameters consider Virtual Machine Monitor level disk access optimizations and rely on a calibration technique. We further present a measurement-based approach that allows a distinct characterization of read/write performance attributes. In particular, we define simple linear prediction models for I/O request mean response times, throughputs and read/write mixes, as well as a simulation model for predicting response time distributions. We found our models to be effective in predicting such quantities across a range of synthetic and emulated application workloads. Stephan Kraft, Giuliano Casale, Diwakar Krishnamurthy, Des Greer, Peter Kilpatrick |
Softw. Syst. Model. | 3 |
| 2012 | Overcoming Web Server Benchmarking Challenges in the Multi-core EraabstractWeb-based services are used by many organizations to support their customers and employees. An important consideration in developing such services is ensuring the Quality of Service (QoS) that users experience is acceptable. Recent years have seen a shift toward deploying Web service son multi-core hardware. Leveraging the performance benefits of multi-core hardware is a non-trivial task. In particular, systematic Web server benchmarking techniques are needed so organizations can verify their ability to meet customer QoS objectives while effectively utilizing such hardware. However, our recent experiences suggest that the multi-core era imposes significant challenges to Web server benchmarking. In particular, due to limitations of current hardware monitoring tools, we found that a large number of experiments are needed to detect complex bottlenecks that can arise in a multi-core system due to contention for shared resources such as cache hierarchy, memory controllers and processor inter-connects. Furthermore, multiple load generator instances are needed to adequately stress multi-core hardware. This leads to practical challenges in validating and managing the test results. This paper describes the automation strategies we employed to overcome these challenges. We make our test harness available for other researchers and practitioners working on similar studies. Raoufehsadat Hashemian, Diwakar Krishnamurthy, Martin F. Arlitt |
ICST | 2 |
| 2012 | Web workload generation challenges - an empirical investigationabstractSUMMARY Workload generators are widely used for testing the performance of Web‐based systems. Typically, these tools are also used to collect measurements such as throughput and end‐user response times that are often used to characterize the QoS provided by a system to its users. However, our study finds that Web workload generation is more difficult than it seems. In examining the popular RUBiS client generator, we found that reported response times could be grossly inaccurate, and that the generated workloads were less realistic than expected, causing server scalability to be incorrectly estimated. Using experimentation, we demonstrate how the Java virtual machine and the Java network library are the root causes of these issues. Our work serves as an example of how to verify the behavior of a Web workload generator. Copyright © 2011 John Wiley & Sons, Ltd. Raoufehsadat Hashemian, Diwakar Krishnamurthy, Martin F. Arlitt |
Softw. Pract. Exp. | 2 |
| 2012 | BURN: Enabling Workload Burstiness in Customized Service BenchmarksabstractWe introduce BURN, a methodology to create customized benchmarks for testing multitier applications under time-varying resource usage conditions. Starting from a set of preexisting test workloads, BURN finds a policy that interleaves their execution to stress the multitier application and generate controlled burstiness in resource consumption. This is useful to study, in a controlled way, the robustness of software services to sudden changes in the workload characteristics and in the usage levels of the resources. The problem is tackled by a model-based technique which first generates Markov models to describe resource consumption patterns of each test workload. Then, a policy is generated using an optimization program which sets as constraints a target request mix and user-specified levels of burstiness at the different resources in the system. Burstiness is quantified using a novel metric called overdemand, which describes in a natural way the tendency of a workload to keep a resource congested for long periods of time and across multiple requests. A case study based on a three-tier application testbed shows that our method is able to control and predict burstiness for session service demands at a fine-grained scale. Furthermore, experiments demonstrate that for any given request mix our approach can expose latency and throughput degradations not found with nonbursty workloads having the same request mix. Giuliano Casale, Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia |
IEEE Trans. Software Eng. | 3 |
| 2012 | DEC: Service Demand Estimation with ConfidenceabstractWe present a new technique for predicting the resource demand requirements of services implemented by multitier systems. Accurate demand estimates are essential to ensure the efficient provisioning of services in an increasingly service-oriented world. The demand estimation technique proposed in this paper has several advantages compared with regression-based demand estimation techniques, which many practitioners employ today. In contrast to regression, it does not suffer from the problem of multicollinearity, it provides more reliable aggregate resource demand and confidence interval predictions, and it offers a measurement-based validation test. The technique can be used to support system sizing and capacity planning exercises, costing and pricing exercises, and to predict the impact of changes to a service upon different service customers. Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia, Stephen Dawson |
IEEE Trans. Software Eng. | 2 |
| 2011 | MODE: Mix Driven On-line Resource Demand Estimation
Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia |
CNSM | 2 |
| 2011 | A trace-based service level planning framework for enterprise application clouds
Anas Youssef, Diwakar Krishnamurthy |
CNSM | 2 |
| 2011 | IO performance prediction in consolidated virtualized environmentsabstractWe propose a trace-driven approach to predict the performance degradation of disk request response times due to storage device contention in consolidated virtualized environments. Our performance model evaluates a queueing network with fair share scheduling using trace-driven simulation. The model parameters can be deduced from measurements obtained inside Virtual Machines (VMs) from a system where a single VM accesses a remote storage server. The parameterized model can then be used to predict the effect of storage contention when multiple VMs are consolidated on the same virtualized server. The model parameter estimation relies on a search technique that tries to estimate the splitting and merging of blocks at the the Virtual Machine Monitor (VMM) level in the case of multiple competing VMs. Simulation experiments based on traces of the Postmark and FFSB disk benchmarks show that our model is able to accurately predict the impact of workload consolidation on VM disk IO response times. Stephan Kraft, Giuliano Casale, Diwakar Krishnamurthy, Des Greer, Peter Kilpatrick |
ICPE | 3 |
| 2011 | Towards automated HPC scheduler configuration tuningabstractAbstract High performance computing (HPC) systems allow researchers and businesses to harness large amounts of computing power needed for solving complex problems. In such systems a job scheduler prioritizes the execution of jobs belonging to users of the system in a manner that allows the system to satisfy performance objectives for various groups of users while simultaneously making efficient use of available resources. Typically, system administrators have the responsibility of manually configuring or tuning the job scheduler such that the performance objectives of user groups as well as system‐level performance objectives are met. Modern job schedulers used in production systems are quite complex. Through detailed trace‐driven simulations, we show that manually tuning the configuration of production schedulers in an environment characterized by multiple performance objectives is very challenging and may not be feasible. To alleviate this problem, this paper describes a toolset that can help a system administrator to automatically configure a scheduler such that the performance objectives for various classes of users in the system as well as other system‐level performance objectives can be satisfied. A unique aspect of this work that differentiates it from the existing work on scheduler tuning is that it has been implemented to work with a widely used production scheduler. Furthermore, in contrast to the existing work it considers the challenging real‐world problem of delivering different levels of performance to different classes of users. System administrators can exploit the toolset to react quickly to changes in performance objectives and workload conditions. Case studies using synthetic and real HPC workloads demonstrate the effectiveness of the technique. Copyright © 2011 John Wiley & Sons, Ltd. Diwakar Krishnamurthy, Mehrnoush Alemzadeh, Mahmood Moussavi |
Concurr. Comput. Pract. Exp. | 1 |
| 2011 | WAM - The Weighted Average Method for Predicting the Performance of Systems with Bursts of Customer SessionsabstractPredictive performance models are important tools that support system sizing, capacity planning, and systems management exercises. We introduce the Weighted Average Method (WAM) to improve the accuracy of analytic predictive performance models for systems with bursts of concurrent customers. WAM considers the customer population distribution at a system to reflect the impact of bursts. The WAM approach is robust with respect to distribution functions, including heavy-tail-like distributions, for workload parameters. We demonstrate the effectiveness of WAM using a case study involving a multitier TPC-W benchmark system. To demonstrate the utility of WAM with multiple performance modeling approaches, we developed both Queuing Network Models and Layered Queuing Models for the system. Results indicate that WAM improves prediction accuracy for bursty workloads for QNMs and LQMs by 10 and 12 percent, respectively, with respect to a Markov Chain approach reported in the literature. Diwakar Krishnamurthy, Jerome A. Rolia |
IEEE Trans. Software Eng. | 1 |
| 2009 | Automatic Stress Testing of Multi-tier Systems by Dynamic Bottleneck Switch Generation
Giuliano Casale, Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia |
Middleware | 3 |
| 2007 | Towards Autonomic Provisioning of Wireless Grid ServicesabstractThe wireless grid paradigm has been proposed recently to support seamless collaborations among devices. Realizing the promise of the wireless grid involves providing a set of middleware services that would allow devices to autonomously share their resources. However, existing wireless grid middleware provide only rudimentary support in this regard. This paper describes our initial experiences in realizing autonomic resource management capabilities for wireless grids. Specifically, we present protocols that support self-configuration mechanisms such as resource discovery and brokering. The proposed protocols leverage similar techniques in the wired grid domain while adding new functionality to handle the dynamic nature of wireless grids. We demonstrate the protocols by using them to build a grid of Bluetooth devices. We also used Bluetooth grid to compare the performance of our service discovery protocol with other similar protocols. Results indicate that our discovery protocol strikes a balance between the compared protocols in terms of resource discovery time and expressiveness of resource description. This allows the protocol to support sophisticated resource brokering and scheduling techniques without incurring prohibitive overheads. Furthermore, in contrast to the existing protocols our protocol supports service advertisements. This facilitates more efficient handling of dynamic situations such as devices leaving and joining the grid. Mohamed El-Darieby, Diwakar Krishnamurthy |
Integrated Network Management | 2 |
| 2006 | Replay: A Model-Based Service for Supporting Transparent Cluster Analysis ToolsabstractGrid computing environments typically federate heterogeneous resource clusters belonging to several organizations. To fully realize the promise of a grid environment, it is necessary to support tools that help obtain insights into the behaviour of individual clusters. This paper describes a cluster service called Replay that simplifies the development and maintenance of such tools. The service provides a common model-based interface for obtaining current and historical information about a cluster. Replay can manage multiple views which allows tools to obtain information about an existing cluster as well as information that shows how a cluster might have behaved under alternate configurations and workloads. The model-based interface allows tools to be ported to different clusters with little effort. Furthermore, Replay uses different mechanisms to manage information that typically changes infrequently and information that can change in a more dynamic, continuous manner allowing it to handle information more efficiently than existing services that provide similar functionality. The paper presents a job analysis tool to illustrate the utility of the service Diwakar Krishnamurthy, Cameron Kiddle, Jerome A. Rolia, Rob Simmonds |
CLUSTER | 1 |
| 2006 | A Synthetic Workload Generation Technique for Stress Testing Session-Based SystemsabstractEnterprise applications are often business critical but lack effective synthetic workload generation techniques to evaluate performance. These workloads are characterized by sessions of interdependent requests that often cause and exploit dynamically generated responses. Interrequest dependencies must be reflected in synthetic workloads for these systems to exercise application functions correctly. This poses significant challenges for automating the construction of representative synthetic workloads and manipulating workload characteristics for sensitivity analyses. This paper presents a technique to overcome these problems. Given request logs for a system under study, the technique automatically creates a synthetic workload that has specified characteristics and maintains the correct interrequest dependencies. The technique is demonstrated through a case study involving a TPC-W e-commerce system. Results show that incorrect performance results can be obtained by neglecting interrequest dependencies, thereby highlighting the value of our technique. The study also exploits our technique to investigate the impact of several workload characteristics on system performance. Results establish that high variability in the distributions of session length, session idle times, and request service times can cause increased contention among sessions, leading to poor system responsiveness. To the best of our knowledge, these are the first results of this kind for a session-based system. We believe our technique is of value for studies where fine control over workload is essential Diwakar Krishnamurthy, Jerome A. Rolia, Shikharesh Majumdar |
IEEE Trans. Software Eng. | 1 |
| 2001 | Characterizing the scalability of a large web-based shopping systemabstractThis article presents an analysis of five days of workload data from a large Web-based shopping system. The multitier environment of this Web-based shopping system includes Web servers, application servers, database servers, and an assortment of load-balancing and firewall appliances. We characterize user requests and sessions and determine their impact on system performance scalability. The purpose of our study is to assess scalability and support capacity planning exercises for the multitier system. We find that horizontal scalability is not always an adequate mechanism for supporting increased workloads and that personalization and robots can have a significant impact on system scalability. Martin F. Arlitt, Diwakar Krishnamurthy, Jerome A. Rolia |
ACM Trans. Internet Techn. | 2 |