EDBT 2026 Demo / reviewers in the wild / expert
Ramana Rao Kompella
dblp:98/2327 · also Ramana Kompella
· DBLP profile ↗
80ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-7559-8997ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 35 · 9 first-author · 3 since 2021Systems, architecture and hardware · 16 · 5 since 2021Artificial intelligence and machine learning · 14 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Software engineering, systems software and programming languages · 8 · 2 since 2021Security and privacy · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Classifying Implementations of Cryptographic Primitives and Protocols that Use Post-Quantum AlgorithmsabstractClassification techniques can be used to analyze system behaviors, network protocols, and cryptographic primitives based on identifiable traits. While useful for defense, such classification can also be leveraged by attackers to infer system configurations, detect vulnerabilities, and tailor attacks such as denial-of-service, key recovery, or downgrade attacks. In this paper, we study the feasibility of classifying post-quantum (PQ) algorithms by analyzing implementations of key exchange and digital signatures, their use within secure protocols, and their integration into SNARK generation libraries. Unlike traditional cryptography, PQ algorithms have larger memory requirements and variable computational costs. Our research examines two post-quantum cryptography libraries, liboqs and CIRCL, evaluating TLS, SSH, QUIC, OpenVPN, and OpenID Connect (OIDC) across Windows, Ubuntu, and macOS. We also analyze pysnark and lattice_zksnark for SNARK generation and verification on Ubuntu. Experimental results show that (1) classical and PQ key exchange and signature algorithms can be distinguished with accuracies of 98% and 100%; (2) specific PQ algorithms can be identified with 97% accuracy for key exchange and 86% for signatures; (3) implementations of the same algorithm in liboqs and CIRCL are distinguishable with up to 100% accuracy; and (4) within CIRCL, PQ and hybrid key exchange implementations can be distinguished with 97% accuracy. For secure protocols, we can determine whether key exchange is classical or PQ and identify the PQ algorithm used. SNARK generation and verification in pysnark and lattice_zksnark are distinguishable with 100% accuracy. We demonstrate real-world applicability by identifying PQ-enabled TLS domains in the Tranco dataset and integrating our methods into QUARTZ, an open-source risk and threat analyzer by Cisco. Tushin Mallick, Ashish Kundu, Ramana Rao Kompella, Cristina Nita-Rotaru |
SACMAT | 3 |
| 2026 | PokéLLMon: A Grounding and Reasoning Benchmark for Large Language Models in Pokémon BattlesabstractDeveloping grounding techniques for LLMs poses two requirements for interactive environments, i.e., (i) the presence of rich knowledge beyond the scope of existing LLMs and (ii) the complexity of tasks that require strategic reasoning. Existing environments fail to meet both requirements due to their simplicity or reliance on commonsense knowledge already encoded in LLMs for interaction. In this article, we present PokéLLMon, a new benchmark enriched with fictional game knowledge and characterized by the intense, dynamic, and adversarial gameplay of Pokémon battles, setting new challenges for the development of grounding and reasoning techniques in interactive environments. Empirical evaluations demonstrate that existing LLMs lack game knowledge and struggle in Pokémon battles. We investigate grounding techniques that leverage feedback and game knowledge, and provide a thorough analysis of reasoning methods from a new perspective of action consistency. Additionally, we introduce higher-level reasoning challenges when playing against human players. The implementation of our benchmark is released at: https://github.com/git-disl/PokeLLMon . Sihao Hu, Tiansheng Huang, Gaowen Liu, Ramana Rao Kompella, Ling Liu 0001 |
ACM Trans. Internet Techn. | 4 |
| 2025 | Targeted Forgetting of Image Subgroups in CLIP ModelsabstractFoundation models (FMs) such as CLIP have demonstrated impressive zero-shot performance across various tasks by leveraging large-scale, unsupervised pre-training. However, they often inherit harmful or unwanted knowledge from noisy internet-sourced datasets, compromising their reliability in real-world applications. Existing model unlearning methods either rely on access to pre-trained datasets or focus on coarse-grained unlearning (e.g., entire classes), leaving a critical gap for fine-grained unlearning. In this paper, we address the challenging scenario of selectively forgetting specific portions of knowledge within a class—without access to pre-trained data—while preserving the model’s overall performance. We propose a novel three-stage approach that progressively unlearns targeted knowledge while mitigating over-forgetting. It consists of (1) a forgetting stage to fine-tune the CLIP on samples to be forgotten, (2) a reminding stage to restore performance on retained samples, and (3) a restoring stage to recover zero-shot capabilities using model souping. Additionally, we introduce knowledge distillation to handle the distribution disparity between forgetting/retaining samples and unseen pre-trained data. Extensive experiments on CIFAR-10, ImageNet-1K, and style datasets demonstrate that our approach effectively unlearns specific subgroups while maintaining strong zero-shot performance on semantically similar subgroups and other categories, significantly outperforming baseline unlearning methods, which lose effectiveness under the CLIP unlearning setting. Zeliang Zhang 0001, Gaowen Liu, Charles Fleming, Ramana Rao Kompella, Chenliang Xu |
CVPR | 4 |
| 2025 | Towards Training Robustness Against Dynamic Errors in Quantum Machine LearningabstractQuantum machine learning, crucial in the noisy intermediate-scale quantum (NISQ) era, confronts challenges in error mitigation. Current noise-aware training (NAT) methods often assume static error rates in quantum neural networks (QNNs), overlooking the dynamic nature of quantum noise. Our work highlights how error rates fluctuate over time and across different qubits, affecting QNN performance even when overall error rates are similar. We introduce a novel NAT strategy that dynamically adjusts to standard and fatal error conditions, incorporating a low-complexity search method to identify fatal errors during optimization. This strategy significantly improves robustness, maintaining competitive performance with leading NAT methods across varying error scenarios. Shijin Duan, Gaowen Liu, Charles Fleming, Ramana Rao Kompella, Xiaolin Xu 0001, Shaolei Ren |
DAC | 4 |
| 2025 | Quantum-Resistant Security: PQC Readiness and Research Challenges (Invited)abstractWhat is your PQC-readiness” - in this paper, we have explored this problem, some of the key challenges to address it and how ciphersuite dependency graphs alongwith observability plays a vital role in addressing some of those challenges. Quantum computing capabilities are evolving fast and on course to develop a cryptographically relevant quantum computer (CRQC). With that several widely used classical cryptography protocols such as RSA are poised to be broken in the next few years. NIST has announced three cryptography protocols as part of the first batch of post-quantum cryptography standards. However, implementation and adoption of quantum-resistant cryptography is a hard problem given the complexities of today’s internet and computing stack. Our work has led to development of algorithms and a system Quartz (Quantum Risk and Threat Analyzer) for observability for quantum vulnerabilities for cryptography suites, where they are used, and analyzing their risks. In that context, we used observability, and the concept of ciphersuite dependency graphs in order to determine use of quantum-unsafe cryptography, and its influence on cryptography supply chain, and generation of cryptography bill of materials (CBOM). Ashish Kundu, Ramana Rao Kompella |
DAC | 2 |
| 2025 | A First-order Generative Bilevel Optimization Framework for Diffusion ModelsabstractDiffusion models, which iteratively denoise data samples to synthesize high-quality outputs, have achieved empirical success across domains. However, optimizing these models for downstream tasks often involves nested bilevel structures, such as tuning hyperparameters for fine-tuning tasks or noise schedules in training dynamics, where traditional bilevel methods fail due to the infinite-dimensional probability space and prohibitive sampling costs. We formalize this challenge as a generative bilevel optimization problem and address two key scenarios: (1) fine-tuning pre-trained models via an inference-only lower-level solver paired with a sample-efficient gradient estimator for the upper level, and (2) training diffusion model from scratch with noise schedule optimization by reparameterizing the lower-level problem and designing a computationally tractable gradient estimator. Our first-order bilevel framework overcomes the incompatibility of conventional bilevel methods with diffusion processes, offering theoretical grounding and computational practicality. Experiments demonstrate that our method outperforms existing fine-tuning and hyperparameter search baselines. Quan Xiao, Hui Yuan 0002, A F M Saif, Gaowen Liu, Ramana Rao Kompella, Mengdi Wang 0001, Tianyi Chen 0002 |
ICML | 5 |
| 2025 | SwitchQNet: Optimizing Distributed Quantum Computing for Quantum Data Centers with Switch NetworksabstractDistributed Quantum Computing (DQC) provides a scalable architecture by interconnecting multiple quantum processor units (QPUs).Among various DQC implementations, quantum data centers (QDCs) -where QPUs in different racks are connected through reconfigurable optical switch networks -are becoming feasible in the near term.However, the latency of cross-rack communications and dynamic switch reconfigurations poses unique challenges to communications in QDCs, significantly increasing the overall latency, thereby also reducing the overall fidelity.In this paper, we address these challenges by introducing a novel compiler that optimizes scheduling of communications across the program and network layers.Our evaluation shows that it reduces the overall latency by 8.02× over prior approaches with a small overhead and can be integrated with quantum error correction (QEC) to facilitate fault-tolerant quantum computing (FTQC).We have open-sourced our codes at https://zenodo.org/records/15377656. Hezi Zhang, Haotian Hu, Keyi Yin, Hassan Shapourian, Jiapeng Zhao, Ramana Rao Kompella, Reza Nejabati, Yufei Ding 0001 |
ISCA | 7 |
| 2025 | Quantized-ViT Efficient Training via Fisher Matrix Regularization
Yuzhang Shang, Gaowen Liu, Ramana Rao Kompella, Yan Yan 0002 |
MMM (3) | 3 |
| 2025 | Orientation-anchored Hyper-Gaussian for 4D Reconstruction from Casual VideosabstractWe present Orientation-anchored Gaussian Splatting (OriGS), a novel framework for high-quality 4D reconstruction from casually captured monocular videos.
While recent advances extend 3D Gaussian Splatting to dynamic scenes via various motion anchors, such as graph nodes or spline control points, they often rely on low-rank assumptions and fall short in modeling complex, region-specific deformations inherent to unconstrained dynamics.
OriGS addresses this by introducing a hyperdimensional representation grounded in scene orientation.
We first estimate a Global Orientation Field that propagates principal forward directions across space and time, serving as stable structural guidance for dynamic modeling.
Built upon this, we propose Orientation-aware Hyper-Gaussian, a unified formulation that embeds time, space, geometry, and orientation into a coherent probabilistic state.
This enables inferring region-specific deformation through principled conditioned slicing, adaptively capturing diverse local dynamics in alignment with global motion intent.
Experiments demonstrate the superior reconstruction fidelity of OriGS over mainstream methods in challenging real-world dynamic scenes. Junyi Wu 0002, Jiachen Tao, Haoxuan Wang 0002, Gaowen Liu, Ramana Rao Kompella, Yan Yan 0002 |
NeurIPS | 5 |
| 2025 | Layer-Wise Security Framework and Analysis for the Quantum InternetabstractWith its significant security potential, the quantum internet is poised to revolutionize technologies like cryptography and communications. Although it boasts enhanced security over traditional networks, the quantum internet still encounters unique security challenges essential for safeguarding its Confidentiality, Integrity, and Availability (CIA). This study explores these challenges by analyzing the vulnerabilities and the corresponding mitigation strategies across different layers of the quantum internet, including physical, link, network, and application layers. We assess the severity of potential attacks, evaluate the expected effectiveness of mitigation strategies, and identify vulnerabilities within diverse network configurations, integrating both classical and quantum approaches. Our research highlights the dynamic nature of these security issues and emphasizes the necessity for adaptive security measures. The findings underline the need for ongoing research into the security dimension of the quantum internet to ensure its robustness, encourage its adoption, and maximize its impact on society. Zebo Yang, Ali Ghubaish, Raj Jain, Ala I. Al-Fuqaha, Aiman Erbad, Ramana Rao Kompella, Hassan Shapourian, Reza Nejabati |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | Comp-Diff: A Unified Pruning and Distillation Framework for Compressing Diffusion ModelsabstractRecently, generative models such as diffusion models (DMs) have gained prominence in various applications, and there is a growing demand for their deployment on resource-constrained devices. Model pruning provides an effective solution by reducing the model redundancy without significantly impacting performance. However, most existing model pruning methods are designed for classification models and often lead to substantial performance degradation when applied to generative models. To address this issue, we propose Comp-Diff, a novel two-stage framework of pruning and knowledge distillation tailored for diffusion models. In the pruning stage, we propose a new structured content-aware pruning (CaP) method within Comp-Diff to identify and preserve informative units (filters/channels) that actually contribute to the generative capability of the model. Specifically, we introduce input perturbations to the pre-trained model and measure each unit’s importance score using gradients induced by these perturbations. Units with higher importance scores are considered more informative and are retained to maintain the model’s generative power. In the fine-tuning stage of Comp-Diff, we propose the distribution-aware knowledge distillation (DaKD) method, which effectively transfers fine-grained knowledge from the original model to the pruned one on both attention and noise distribution levels. In addition, DaKD includes an adversarial loss to improve the quality and diversity of generated outputs. To verify and evaluate our method, we apply the proposed Comp-Diff on three representative tasks: unconditional image generation, conditional image generation, and text-to-image generation. Extensive experiments on both multi-step and one-step diffusion models demonstrate that the proposed framework consistently yields compact models and outperforms existing pruning techniques by a large margin. Wei Xiang 0001, Kang Han, Gaowen Liu, Ramana Rao Kompella |
IEEE Trans. Multim. | 5 |
| 2024 | Large Language Models Can Learn Temporal ReasoningabstractWhile large language models (LLMs) have demonstrated remarkable reasoning capabilities, they are not without their flaws and inaccuracies.Recent studies have introduced various methods to mitigate these limitations.Temporal reasoning (TR), in particular, presents a significant challenge for LLMs due to its reliance on diverse temporal concepts and intricate temporal logic.In this paper, we propose TG-LLM, a novel framework towards languagebased TR.Instead of reasoning over the original context, we adopt a latent representation, temporal graph (TG) that enhances the learning of TR.A synthetic dataset (TGQA), which is fully controllable and requires minimal supervision, is constructed for fine-tuning LLMs on this text-to-TG translation task.We confirmed in experiments that the capability of TG translation learned on our dataset can be transferred to other TR tasks and benchmarks.On top of that, we teach LLM to perform deliberate reasoning over the TGs via Chain-of-Thought (CoT) bootstrapping and graph data augmentation.We observed that those strategies, which maintain a balance between usefulness and diversity, bring more reliable CoTs and final results than the vanilla CoT distillation. 1 * Equal contribution. 1 Code and data are available at https://github.com/ xiongsiheng/TG-LLM.Once upon a time in the quaint town of Weston, a baby boy named John Thompson was brought into the world in the year 1921.Growing up, he had a vibrant spirit and an adventurous soul.…Step 1: Text-to-Temporal Graph Translation True or false: event (John Thompson owned Pearl Network) was longer in duration than event (Sophia Parker was married to John Thompson) ?The duration for each event can be calculated as follows:(John Thompson owned Pearl Network) starts at 1942, ends at 1967 Siheng Xiong, Ali Payani, Ramana Rao Kompella, Faramarz Fekri |
ACL (1) | 3 |
| 2024 | OnePerc: A Randomness-aware Compiler for Photonic Quantum ComputingabstractThe photonic platform holds great promise for quantum computing. Nevertheless, the intrinsic probabilistic characteristic of its native fusion operations introduces substantial randomness into the computing process, posing significant challenges to achieving scalability and efficiency in program execution. In this paper, we introduce a randomness-aware compilation framework designed to concurrently achieve scalability and efficiency. Our approach leverages an innovative combination of offline and online optimization passes, with a novel intermediate representation serving as a crucial bridge between them. Through a comprehensive evaluation, we demonstrate that this framework significantly outperforms the most efficient baseline compiler in a scalable manner, opening up new possibilities for realizing scalable photonic quantum computing. Hezi Zhang, Jixuan Ruan, Hassan Shapourian, Ramana Rao Kompella, Yufei Ding 0001 |
ASPLOS (3) | 4 |
| 2024 | Enhancing Large Language Models through Transforming Reasoning Problems into Classification TasksabstractIn this paper, we introduce a novel approach for enhancing the reasoning capabilities of large language models (LLMs) for constraint satisfaction problems (CSPs), by converting reasoning problems into classification tasks. Our method leverages the LLM’s ability to decide when to call a function from a set of logical-linguistic primitives, each of which can interact with a local “scratchpad” memory and logical inference engine. Invocation of these primitives in the correct order writes the constraints to the scratchpad memory and enables the logical engine to verifiably solve the problem. We additionally propose a formal framework for exploring the “linguistic” hardness of CSP reasoning-problems for LLMs. Our experimental results demonstrate that under our proposed method, tasks with significant computational hardness can be converted to a form that is easier for LLMs to solve and yields a 40% improvement over baselines. This opens up new avenues for future research into hybrid cognitive models that integrate symbolic and neural approaches. Tarun Raheja, Raunak Sinha, Advit Deepak, Will Healy, Jayanth Srinivasa, Myungjin Lee, Ramana Rao Kompella |
LREC/COLING | 7 |
| 2024 | Riemannian Multinomial Logistics Regression for SPD Neural NetworksabstractDeep neural networks for learning Symmetric Positive Definite (SPD) matrices are gaining increasing attention in machine learning. Despite the significant progress, most existing SPD networks use traditional Euclidean classifiers on an approximated space rather than intrinsic classifiers that accurately capture the geometry of SPD manifolds. In-spired by Hyperbolic Neural Networks (HNNs), we propose Riemannian Multinomial Logistics Regression (RMLR) for the classification layers in SPD networks. We introduce a unified framework for building Riemannian classifiers under the metrics pulled back from the Euclidean space, and showcase our framework under the parameterized Log-Euclidean Metric (LEM) and Log-Cholesky Metric (LCM). Besides, our framework offers a novel intrinsic explanation for the most popular LogEig classifier in existing SPD networks. The effectiveness of our method is demonstrated in three applications: radar recognition, human action recognition, and electroencephalography (EEG) classification. The code is available at https://github.com/GitZH-Chen/SPDMLR.git. Ziheng Chen 0001, Yue Song 0002, Gaowen Liu, Ramana Rao Kompella, Xiaojun Wu 0001, Nicu Sebe |
CVPR | 4 |
| 2024 | Efficient Multitask Dense Predictor via BinarizationabstractMulti-task learning for dense prediction has emerged as a pivotal area in computer vision, enabling simultaneous processing of diverse yet interrelated pixel-wise prediction tasks. However, the substantial computational demands of state-of-the-art (SoTA) models often limit their widespread deployment. This paper addresses this challenge by introducing network binarization to compress resource-intensive multi-task dense predictors. Specifically, our goal is to significantly accelerate multi-task dense prediction models via Binary Neural Networks (BNNs) while maintaining and even improving model performance at the same time. To reach this goal, we propose a Binary Multi-task Dense Predictor, Bi -MTPD, and several variants of Bi -MTPD, in which a multi-task dense predictor is constructed via specified binarized modules. Our systematical analysis of this predictor reveals that performance drop from binarization is primarily caused by severe information degradation. To address this issue, we introduce a deep information bottleneck layer that enforces representations for downstream tasks satisfying Gaussian distribution in forward propagation. Moreover, we introduce a knowledge distillation mechanism to correct the direction of information flow in backward propagation. Intriguingly, one variant of Bi -MTPD outperforms full-precision (FP) multi-task dense prediction SoTAs, ARTC [2] (CNN-based) and InvPT [50] (ViT-Based). This result indicates that Bi -MTPD is not merely a naive trade-off between performance and efficiency, but is rather a benefit of the redundant information flow thanks to the multi-task architecture. Code is available at BiMTDP. Yuzhang Shang, Dan Xu 0002, Gaowen Liu, Ramana Rao Kompella, Yan Yan 0002 |
CVPR | 4 |
| 2024 | Enhancing Post-Training Quantization Calibration Through Contrastive LearningabstractPost-training quantization (PTQ) converts a pre-trained full-precision (FP) model into a quantized model in a training-free manner. Determining suitable quantization parameters, such as scaling factors and zero points, is the primary strategy for mitigating the impact of quantization noise (calibration) and restoring the performance of the quantized models. However, the existing activation calibration methods have never considered information degradation between pre- (FP) and post-quantized activations. In this study, we introduce a well-defined distributional metric from information theory, mutual information, into PTQ calibration. We aim to calibrate the quantized activations by maximizing the mutual information between the pre- and post-quantized activations. To realize this goal, we establish a contrastive learning (CL) framework for the calibration, where the quantization parameters are optimized through a self-supervised proxy task. Specifically, by leveraging CL during the PTQ calibration, we can benefit from pulling the positive pairs of quantized and FP activations collected from the same input samples, while pushing negative pairs from different samples. Thanks to the ingeniously designed critic function, we avoid the unwanted but of tenencountered collision solution in CL, especially in calibration scenarios where the amount of calibration data is limited. Additionally, we provide a theoretical guarantee that minimizing our designed loss is equivalent to maximizing the desired mutual information. Consequently, the quantized activations retain more information, which ultimately enhances the performance of the quantized network. Experimental results show that our method can effectively serve as an add-on module to existing SoTA PTQ methods. Yuzhang Shang, Gaowen Liu, Ramana Rao Kompella, Yan Yan 0002 |
CVPR | 3 |
| 2024 | A Method for Bilevel Optimization with Convex Lower-Level ProblemabstractGradient-based bilevel optimization methods have been applied to a wide range of applications including hyper-parameter optimization, meta-learning, and model pruning. However, it is known that the bilevel optimization problem is difficult to solve, and the finite-time guarantee has only been established for simpler bilevel problems with a strongly-convex lower-level problem. In this work, we propose an iterative bilevel optimization method that sequentially solves simple approximate problems of the original problem. Despite the lack of strong convexity in the lower level, we show that the proposed method converges to an ϵ-stationary-point with an iteration complexity of $\mathcal{O}\left( {{\varepsilon ^{ - 1}}} \right)$. Experiments have verified the effectiveness of the method. Santiago Paternain, Gaowen Liu, Ramana Rao Kompella, Tianyi Chen 0002 |
ICASSP | 4 |
| 2024 | LightPure: Realtime Adversarial Image Purification for Mobile Devices Using Diffusion ModelsabstractAutonomous mobile systems increasingly rely on deep neural networks for perception and decision-making. While effective, these systems are vulnerable to adversarial machine learning attacks where small perturbations in the input could significantly impact the outcome of the system. Common countermeasures include leveraging adversarial training and/or data or network transformation. Although widely used, the main drawback of these countermeasures is that they require full and invasive access to the classifiers, which are typically proprietary. Additionally, the cost of training or retraining is often prohibitively expensive for large models. To tackle this, purification models have recently been proposed. The aim is to incorporate a "purification" layer before classification, thereby eliminating the necessity to modify the classifier. Despite their effectiveness, state-of-the-art purification methods are compute-intensive, rendering them unsuitable for mobile systems where resources are constrained and large latency is not desired. Hossein Khalili, Vincent Li, Brandan Bright, Ali Payani, Ramana Rao Kompella, Nader Sehatbakhsh |
MobiCom | 6 |
| 2024 | Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit DifferenceabstractAs Large Language Models (LLMs) demonstrate extensive capability in learning from documents, LLM unlearning becomes an increasingly important research area to address concerns of LLMs in terms of privacy, copyright, etc. A conventional LLM unlearning task typically involves two goals: (1) The target LLM should forget the knowledge in the specified forget documents; and (2) it should retain the other knowledge that the LLM possesses, for which we assume access to a small number of retain documents. To achieve both goals, a mainstream class of LLM unlearning methods introduces an optimization framework with a combination of two objectives – maximizing the prediction loss on the forget documents while minimizing that on the retain documents, which suffers from two challenges, degenerated output and catastrophic forgetting. In this paper, we propose a novel unlearning framework called Unlearning from Logit Difference (ULD), which introduces an assistant LLM that aims to achieve the opposite of the unlearning goals: remembering the forget documents and forgetting the retain knowledge. ULD then derives the unlearned LLM by computing the logit difference between the target and the assistant LLMs. We show that such reversed objectives would naturally resolve both aforementioned challenges while significantly improving the training efficiency. Extensive experiments demonstrate that our method efficiently achieves the intended forgetting while preserving the LLM’s overall capabilities, reducing training time by more than threefold. Notably, our method loses 0% of model utility on the ToFU benchmark, whereas baseline methods may sacrifice 17% of utility on average to achieve comparable forget quality. Jiabao Ji, Yujian Liu, Yang Zhang 0001, Gaowen Liu, Ramana Rao Kompella, Sijia Liu 0001, Shiyu Chang |
NeurIPS | 5 |
| 2024 | From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion ModelsabstractWhile state-of-the-art diffusion models (DMs) excel in image generation, concerns regarding their security persist. Earlier research highlighted DMs' vulnerability to data poisoning attacks, but these studies placed stricter requirements than conventional methods like 'BadNets' in image classification. This is because the art necessitates modifications to the diffusion training and sampling procedures. Unlike the prior work, we investigate whether BadNets-like data poisoning methods can directly degrade the generation by DMs. In other words, if only the training dataset is contaminated (without manipulating the diffusion process), how will this affect the performance of learned DMs? In this setting, we uncover bilateral data poisoning effects that not only serve an adversarial purpose (compromising the functionality of DMs) but also offer a defensive advantage (which can be leveraged for defense in classification tasks against poisoning attacks). We show that a BadNets-like data poisoning attack remains effective in DMs for producing incorrect images (misaligned with the intended text conditions). Meanwhile, poisoned DMs exhibit an increased ratio of triggers, a phenomenon we refer to as 'trigger amplification', among the generated images. This insight can be then used to enhance the detection of poisoned training data. In addition, even under a low poisoning ratio, studying the poisoning effects of DMs is also valuable for designing robust image classifiers against such attacks. Last but not least, we establish a meaningful linkage between data poisoning and the phenomenon of data replications by exploring DMs' inherent data memorization tendencies. Code is available at https://github.com/OPTML-Group/BiBadDiff. Zhuoshi Pan, Yuguang Yao, Gaowen Liu, Bingquan Shen, H. Vicky Zhao, Ramana Rao Kompella, Sijia Liu 0001 |
NeurIPS | 6 |
| 2024 | UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion ModelsabstractThe technological advancements in diffusion models (DMs) have demonstrated unprecedented capabilities in text-to-image generation and are widely used in diverse applications. However, they have also raised significant societal concerns, such as the generation of harmful content and copyright disputes. Machine unlearning (MU) has emerged as a promising solution, capable of removing undesired generative capabilities from DMs. However, existing MU evaluation systems present several key challenges that can result in incomplete and inaccurate assessments. To address these issues, we propose UnlearnCanvas, a comprehensive high-resolution stylized image dataset that facilitates the evaluation of the unlearning of artistic styles and associated objects. This dataset enables the establishment of a standardized, automated evaluation framework with 7 quantitative metrics assessing various aspects of the unlearning performance for DMs. Through extensive experiments, we benchmark 9 state-of-the-art MU methods for DMs, revealing novel insights into their strengths, weaknesses, and underlying mechanisms. Additionally, we explore challenging unlearning scenarios for DMs to evaluate worst-case performance against adversarial prompts, the unlearning of finer-scale concepts, and sequential unlearning. We hope that this study can pave the way for developing more effective, accurate, and robust DM unlearning methods, ensuring safer and more ethical applications of DMs in the future. The dataset, benchmark, and codes are publicly available at this link. Chongyu Fan, Yuguang Yao, Jinghan Jia, Jiancheng Liu, Gaoyuan Zhang, Gaowen Liu, Ramana Rao Kompella, Xiaoming Liu 0002, Sijia Liu 0001 |
NeurIPS | 9 |
| 2024 | Adaptive Deep Neural Network Inference Optimization with EENetabstractWell-trained deep neural networks (DNNs) treat all test samples equally during prediction. Adaptive DNN inference with early exiting leverages the observation that some test examples can be easier to predict than others. This paper presents EENet, a novel early-exiting scheduling framework for multi-exit DNN models. Instead of having every sample go through all DNN layers during prediction, EENet learns an early exit scheduler, which can intelligently terminate the inference earlier for certain predictions, which the model has high confidence of early exit. As opposed to previous early-exiting solutions with heuristics-based methods, our EENet framework optimizes an early-exiting policy to maximize model accuracy while satisfying the given per-sample average inference budget. Extensive experiments are conducted on four computer vision datasets (CIFAR-10, CIFAR-100, ImageNet, Cityscapes) and two NLP datasets (SST-2, AgNews). The results demonstrate that the adaptive inference by EENet can outperform the representative existing early exit techniques. We also perform a detailed visualization analysis of the comparison results to interpret the benefits of EENet. Fatih Ilhan, Ka-Ho Chow 0001, Sihao Hu, Tiansheng Huang, Selim F. Tekin, Wenqi Wei 0001, Yanzhao Wu 0001, Myungjin Lee, Ramana Rao Kompella, Hugo Latapie, Gaowen Liu, Ling Liu 0001 |
WACV | 9 |
| 2023 | Flame: Simplifying Topology Extension in Federated LearningabstractDistributed machine learning approaches, including a broad class of federated learning (FL) techniques, present a number of benefits when deploying machine learning applications over widely distributed infrastructures. The benefits are highly dependent on the details of the underlying machine learning topology, which specifies the functionality executed by the participating nodes, their dependencies and interconnections. Current systems lack the flexibility and extensibility necessary to customize the topology of a machine learning deployment. We present Flame, a new system that provides flexibility of the topology configuration of distributed FL applications around the specifics of a particular deployment context, and is easily extensible to support new FL architectures. Flame achieves this via a new high-level abstraction Topology Abstraction Graphs (TAGs). TAGs decouple the ML application logic from the underlying deployment details, making it possible to specialize the application deployment with reduced development effort. Flame is released as an open source project, and its flexibility and extensibility support a variety of topologies and mechanisms, and can facilitate the development of new FL methodologies. Harshit Daga, Jaemin Shin 0005, Dhruv Garg, Ada Gavrilovska, Myungjin Lee, Ramana Rao Kompella |
SoCC | 6 |
| 2023 | Causal-DFQ: Causality Guided Data-free Network QuantizationabstractModel quantization, which aims to compress deep neural networks and accelerate inference speed, has greatly facilitated the development of cumbersome models on mobile and edge devices. There is a common assumption in quantization methods from prior works that training data is available. In practice, however, this assumption cannot always be fulfilled due to reasons of privacy and security, rendering these methods inapplicable in real-life situations. Thus, data-free network quantization has recently received significant attention in neural network compression. Causal reasoning provides an intuitive way to model causal relationships to eliminate data-driven correlations, making causality an essential component of analyzing data-free problems. However, causal formulations of data-free quantization are inadequate in the literature. To bridge this gap, we construct a causal graph to model the data generation and discrepancy reduction between the pre-trained and quantized models. Inspired by the causal understanding, we propose the Causality-guided Data-free Network Quantization method, Causal-DFQ, to eliminate the reliance on data via approaching an equilibrium of causality-driven intervened distributions. Specifically, we design a content-style-decoupled generator, synthesizing images conditioned on the relevant and irrelevant factors; then we propose a discrepancy reduction loss to align the intervened distributions of the pre-trained and quantized models. It is worth noting that our work is the first attempt towards introducing causality to data-free quantization problem. Extensive experiments demonstrate the efficacy of Causal-DFQ. The code is available at Causal-DFQ. Yuzhang Shang, Bingxin Xu, Gaowen Liu, Ramana Rao Kompella, Yan Yan 0002 |
ICCV | 4 |
| 2023 | Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity ModelingabstractGraph neural networks (GNNs) have found extensive applications in learning from graph data. However, real-world graphs often possess diverse structures and comprise nodes and edges of varying types. To bolster the generalization capacity of GNNs, it has become customary to augment training graph structures through techniques like graph augmentations and large-scale pre-training on a wider array of graphs. Balancing this diversity while avoiding increased computational costs and the notorious trainability issues of GNNs is crucial. This study introduces the concept of Mixture-of-Experts (MoE) to GNNs, with the aim of augmenting their capacity to adapt to a diverse range of training graph structures, without incurring explosive computational overhead. The proposed Graph Mixture of Experts (GMoE) model empowers individual nodes in the graph to dynamically and adaptively select more general information aggregation experts. These experts are trained to capture distinct subgroups of graph structures and to incorporate information with varying hop sizes, where those with larger hop sizes specialize in gathering information over longer distances. The effectiveness of GMoE is validated through a series of experiments on a diverse set of tasks, including graph, node, and link prediction, using the OGB benchmark. Notably, it enhances ROC-AUC by $1.81\%$ in ogbg-molhiv and by $1.40\%$ in ogbg-molbbbp, when compared to the non-MoE baselines. Our code is publicly available at https://github.com/VITA-Group/Graph-Mixture-of-Experts. Haotao Wang, Ziyu Jiang, Yuning You, Yan Han 0001, Gaowen Liu, Jayanth Srinivasa, Ramana Rao Kompella, Zhangyang Wang |
NeurIPS | 7 |
| 2018 | Fault Localization in Large-Scale Network Policy DeploymentabstractThe recent advances in network management automation and Software-Defined Networking (SDN) facilitate network policy management tasks. At the same time, these new technologies create a new mode of failure in the management cycle itself. Network policies are presented in an abstract model at a centralized controller and deployed as low-level rules across network devices. Thus, any software and hardware element in that cycle can be a potential cause of underlying network problems. In this paper, we present and solve a network policy fault localization problem that arises in operating policy management frameworks for a production network. We formulate our problem via risk modeling and propose a greedy algorithm that quickly localizes faulty policy objects in the network policy. We then design and develop SCOUT-a fully-automated system that produces faulty policy objects and further pinpoints physical-level failures which made the objects faulty. Evaluation results using a real testbed and extensive simulations demonstrate that SCOUT detects faulty objects with small false positives and false negatives. Praveen Tammana, Chandra Nagarajan, Pavan Mamillapalli, Ramana Rao Kompella, Myungjin Lee |
ICDCS | 4 |
| 2015 | vHaul: Towards Optimal Scheduling of Live Multi-VM Migration for Multi-tier ApplicationsabstractLive virtual machine (VM) migration enables seamless movement of an online server from one location to another to achieve failure recovery, load balancing, and system maintenance. Beyond single VM migration, a multi-tier application involves a group of correlated VMs and its live migration will require careful scheduling of the migrations of the member VMs. Our observations from extensive experiments using a variety of multi-tier applications suggest that, in a dedicated data center with dedicated migration links, different migration strategies result in distinct performance impacts on a multi-tier application. The root cause of the problem is the inter-dependence between functional components of a multitier application. We leverage these observations in vHaul, a system that coordinates multi-VM migration to approximate the optimal scheduling. Our evaluation of a vHaul prototype on Xen suggests that vHaul yields the optimal multi-VM live migration schedules. Further, our application-level evaluation using Apache Olio, a web 2.0 cloud application, shows that the optimal migration schedule produced by vHaul outperforms the worst-case schedule by 43% in application throughput. Moreover, the optimal schedule significantly reduces service latency during migration by up to 70%. Hui Lu 0001, Cong Xu 0010, Ramana Rao Kompella, Dongyan Xu |
CLOUD | 4 |
| 2015 | vFair: latency-aware fair storage scheduling via per-IO cost-based differentiationabstractIn virtualized data centers, multiple VMs are consolidated to access a shared storage system. Effective storage resource management, however, turns out to be challenging, as VM workloads exhibit various IO patterns and diverse loads. To multiplex the underlying hardware resources among VMs, providing fairness and isolation while maintaining high resource utilization becomes imperative for effective storage resource management. Existing schedulers such as Linux CFQ or SFQ can provide some fairness, but it has been observed that synchronous IO tends to lose fair shares significantly when competing with aggressive VMs. Hui Lu 0001, Brendan Saltaformaggio, Ramana Rao Kompella, Dongyan Xu |
SoCC | 3 |
| 2015 | Inferring the Network Latency Requirements of Cloud Tenants
Jeffrey C. Mogul, Ramana Rao Kompella |
HotOS | 2 |
| 2015 | vRead: Efficient Data Access for Hadoop in Virtualized CloudsabstractWith its unlimited scalability and on-demand access to computation and storage, a virtualized cloud platform is the perfect match for big data systems such as Hadoop. However, virtualization introduces a significant amount of overhead to I/O intensive applications due to device virtualization and VMs or I/O threads scheduling delay. In particular, device virtualization causes significant CPU overhead as I/O data needs to be moved across several protection boundaries. We observe that such overhead especially affects the I/O performance of the Hadoop distributed file system (HDFS). In fact, data read from an HDFS datanode VM must go through virtual devices multiple times --- incurring non-negligible virtualization overhead --- even though both client VM and datanode VM may be running on the same machine. In this paper, we propose vRead, a programmable framework which connects I/O flows from HDFS applications directly to their data. vRead enables direct "reads" to the disk images of datanode VMs from the hypervisor. By doing so, vRead can significantly avoid device virtualization overhead, resulting in improved I/O throughput as well as CPU savings for Hadoop workloads and other applications relying on HDFS. Cong Xu 0010, Brendan Saltaformaggio, Sahan Gamage, Ramana Rao Kompella, Dongyan Xu |
Middleware | 4 |
| 2015 | A flow measurement architecture to preserve application structure
Myungjin Lee, Mohammad Y. Hajjat, Ramana Rao Kompella, Sanjay G. Rao |
Comput. Networks | 3 |
| 2014 | ElastiCon: an elastic distributed sdn controllerabstractSoftware Defined Networking (SDN) has become a popular paradigm for centralized control in many modern networking scenarios such as data centers and cloud. For large data centers hosting many hundreds of thousands of servers, there are few thousands of switches that need to be managed in a centralized fashion, which cannot be done using a single controller node. Previous works have proposed distributed controller architectures to address scalability issues. A key limitation of these works, however, is that the mapping between a switch and a controller is statically configured, which may result in uneven load distribution among the controllers as traffic conditions change dynamically. To address this problem, we propose ElastiCon, an elastic distributed controller architecture in which the controller pool is dynamically grown or shrunk according to traffic conditions. To address the load imbalance caused due to spatial and temporal variations in the traffic conditions, ElastiCon automatically balances the load across controllers thus ensuring good performance at all times irrespective of the traffic dynamics. We propose a novel switch migration protocol for enabling such load shifting, which conforms with the Openflow standard. We further design the algorithms for controller load balancing and elasticity. We also build a prototype of ElastiCon and evaluate it extensively to demonstrate the efficacy of our design. Advait Abhay Dixit, Fang Hao, Sarit Mukherjee, T. V. Lakshman, Ramana Rao Kompella |
ANCS | 5 |
| 2014 | vPipe: Piped I/O Offloading for Efficient Data Movement in Virtualized CloudsabstractVirtualization introduces a significant amount of overhead for I/O intensive applications running inside virtual machines (VMs). Such overhead is caused by two main sources: (1) device virtualization and (2) VM scheduling. Device virtualization causes significant CPU overhead as I/O data need to be moved across several protection boundaries. VM scheduling introduces delays to the overall I/O processing path due to the wait time of VMs' virtual CPUs in the run queue. We observe that such overhead particularly affects many applications involving piped I/O data movements, such as web servers, streaming servers, big data analytics, and storage, because the data has to be transferred first into the application from the source I/O device and then back to the sink I/O device, incurring the virtualization overhead twice. In this paper, we propose vPipe, a programmable framework to mitigate this problem for a wide range of applications running in virtualized clouds. vPipe enables direct "piping" of application I/O data from source to sink devices, either files or TCP sockets, at virtual machine monitor (VMM) level. By doing so, vPipe can avoid both device virtualization overhead and VM scheduling delays, resulting in improved I/O throughput and application performance as well as significant CPU savings. Sahan Gamage, Cong Xu 0010, Ramana Rao Kompella, Dongyan Xu |
SoCC | 3 |
| 2014 | Graph sample and hold: a framework for big-graph analyticsabstractSampling is a standard approach in big-graph analytics; the goal is to efficiently estimate the graph properties by consulting a sample of the whole population. A perfect sample is assumed to mirror every property of the whole population. Unfortunately, such a perfect sample is hard to collect in complex populations such as graphs (e.g. web graphs, social networks), where an underlying network connects the units of the population. Therefore, a good sample will be representative in the sense that graph properties of interest can be estimated with a known degree of accuracy. Nesreen K. Ahmed, Nick G. Duffield, Jennifer Neville, Ramana Rao Kompella |
KDD | 4 |
| 2014 | FineComb: Measuring Microscopic Latency and Loss in the Presence of ReorderingabstractModern stock trading and cluster applications require microsecond latencies and almost no losses in data centers. This paper introduces an algorithm called FineComb that can obtain fine-grain end-to-end loss and latency measurements between edge routers in these networks. Such a mechanism can allow managers to distinguish between latencies and loss singularities caused by servers and those caused by the network. Compared to prior work, such as Lossy Difference Aggregator (LDA), which focused on switch-level latency measurements, the requirement of end-to-end latency measurements introduces the challenge of reordering that occurs commonly in IP networks due to churn. The problem is even more acute in switches across data center networks that employ multipath routing algorithms to exploit the inherent path diversity. Without proper care, a loss estimation algorithm can confound loss and reordering; furthermore, any attempt to aggregate delay estimates in the presence of reordering results in severe errors. FineComb deals with these problems using order-agnostic packet digests and a simple new idea we call stash recovery. Our evaluation demonstrates that FineComb is orders of magnitude more accurate than LDA in loss and delay estimates in the presence of reordering. Myungjin Lee, Sharon Goldberg, Ramana Rao Kompella, George Varghese |
IEEE/ACM Trans. Netw. | 3 |
| 2013 | On the impact of packet spraying in data center networksabstractModern data center networks are commonly organized in multi-rooted tree topologies. They typically rely on equal-cost multipath to split flows across multiple paths, which can lead to significant load imbalance. Splitting individual flows can provide better load balance, but is not preferred because of potential packet reordering that conventional wisdom suggests may negatively interact with TCP congestion control. In this paper, we revisit this “myth” in the context of data center networks which have regular topologies such as multi-rooted trees. We argue that due to symmetry, the multiple equal-cost paths between two hosts are composed of links that exhibit similar queuing properties. As a result, TCP is able to tolerate the induced packet reordering and maintain a single estimate of RTT. We validate the efficacy of random packet spraying (RPS) using a data center testbed comprising real hardware switches. We also reveal the adverse impact on the performance of RPS when the symmetry is disturbed (e.g., during link failures) and suggest solutions to mitigate this effect. Advait Abhay Dixit, Pawan Prakash, Y. Charlie Hu, Ramana Rao Kompella |
INFOCOM | 4 |
| 2013 | CobWeb: In-network cobbling of web traffic
Hitesh Khandelwal, Fang Hao, Sarit Mukherjee, Ramana Rao Kompella, T. V. Lakshman |
Networking | 4 |
| 2013 | PhishLive: A View of Phishing and Malware Attacks from an Edge Router
Lianjie Cao, Thibaut Probst, Ramana Rao Kompella |
PAM | 3 |
| 2013 | vTurbo: Accelerating Virtual Machine I/O Processing Using Designated Turbo-Sliced Core
Cong Xu 0010, Sahan Gamage, Hui Lu 0001, Ramana Rao Kompella, Dongyan Xu |
USENIX ATC | 4 |
| 2013 | Network Sampling: From Static to Streaming GraphsabstractNetwork sampling is integral to the analysis of social, information, and biological networks. Since many real-world networks are massive in size, continuously evolving, and/or distributed in nature, the network structure is often sampled in order to facilitate study. For these reasons, a more thorough and complete understanding of network sampling is critical to support the field of network science. In this paper, we outline a framework for the general problem of network sampling by highlighting the different objectives, population and units of interest, and classes of network sampling methods. In addition, we propose a spectrum of computational models for network sampling methods, ranging from the traditionally studied model based on the assumption of a static domain to a more challenging model that is appropriate for streaming domains. We design a family of sampling methods based on the concept of graph induction that generalize across the full spectrum of computational models (from static to streaming) while efficiently preserving many of the topological properties of the input graphs. Furthermore, we demonstrate how traditional static sampling algorithms can be modified for graph streams for each of the three main classes of sampling methods: node, edge, and topology-based sampling. Experimental results indicate that our proposed family of sampling methods more accurately preserve the underlying properties of the graph in both static and streaming domains. Finally, we study the impact of network sampling algorithms on the parameter estimation and performance evaluation of relational classification algorithms. Nesreen K. Ahmed, Jennifer Neville, Ramana Rao Kompella |
ACM Trans. Knowl. Discov. Data | 3 |
| 2013 | Protocol Responsibility Offloading to Improve TCP Throughput in Virtualized EnvironmentsabstractVirtualization is a key technology that powers cloud computing platforms such as Amazon EC2. Virtual machine (VM) consolidation, where multiple VMs share a physical host, has seen rapid adoption in practice, with increasingly large numbers of VMs per machine and per CPU core. Our investigations, however, suggest that the increasing degree of VM consolidation has serious negative effects on the VMs’ TCP performance. As multiple VMs share a given CPU, the scheduling latencies, which can be in the order of tens of milliseconds, substantially increase the typically submillisecond round-trip times (RTTs) for TCP connections in a datacenter, causing significant degradation in throughput. In this article, we propose a lightweight solution, called vPRO, that (a) offloads the VM’s TCP congestion control function to the driver domain to improve TCP transmit performance; and (b) offloads TCP acknowledgment functionality to the driver domain to improve the TCP receive performance. Our evaluation of a vPRO prototype on Xen suggests that vPRO substantially improves TCP receive and transmit throughputs with minimal per-packet CPU overhead. We further show that the higher TCP throughput leads to improvement in application-level performance, via experiments with Apache Olio, a Web 2.0 cloud application, and Intel MPI benchmark. Sahan Gamage, Ramana Rao Kompella, Dongyan Xu, Ardalan Kangarlou |
ACM Trans. Comput. Syst. | 2 |
| 2013 | High-Fidelity Per-Flow Delay Measurements With Reference Latency InterpolationabstractNew applications such as soft real-time data center applications, algorithmic trading, and high-performance computing require extremely low latency (in microseconds) from networks. Network operators today lack sufficient fine-grain measurement tools to detect, localize, and repair delay spikes that cause application service level agreement (SLA) violations. A recently proposed solution called LDA provides a scalable way to obtain latency, but only provides aggregate measurements. However, debugging application-specific problems requires per-flow measurements since different flows may exhibit significantly different characteristics even when they are traversing the same link. To enable fine-grained per-flow measurements in routers, we propose a new scalable architecture called reference latency interpolation (RLI) that is based on our observation that packets potentially belonging to different flows that are closely spaced to each other exhibit similar delay properties. In our evaluation using simulations over real traces, we show that while having small overhead, RLI achieves a median relative error of 12% and one to two orders of magnitude higher accuracy than previous per-flow measurement solutions. We also observe RLI achieves as high accuracy as LDA in aggregate latency estimation, and RLI outperforms LDA in standard deviation estimation. Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella |
IEEE/ACM Trans. Netw. | 3 |
| 2012 | On the performance projectability of MapReduceabstractA key challenge faced by users of public clouds today is how to request for the right amount of resources in the production datacenter that satisfies a target performance for a given cloud application. An obvious approach is to develop a performance model for a class of applications such as MapReduce. However, several recent studies have shown that even for the class of well-studied MapReduce jobs, their running times can be seriously affected by numerous external factors ranging from dozen or so configuration parameters, to the physical machine characteristics (CPU, memory, disk, and network bandwidth), to implementation deficiencies such as Java, garbage collection. These factors make direct performance modeling extremely difficult. In this paper, we propose a more practical systematic methodology to solve this problem. Our approach develops a projection model, based on insights into performance bottlenecks of MapReduce jobs and their scaling properties, and parameterized with component running times based on profiling on small clusters with sampled inputs. Evaluation results show our projection model can predict job running times with 2.7% of accuracy when scaling to 32 nodes. Di Xie, Y. Charlie Hu, Ramana Rao Kompella |
CloudCom | 3 |
| 2012 | vSlicer: latency-aware virtual machine scheduling via differentiated-frequency CPU slicingabstractRecent advances in virtualization technologies have made it feasible to host multiple virtual machines (VMs) in the same physical host and even the same CPU core, with fair share of the physical resources among the VMs. However, as more VMs share the same core/CPU, the CPU access latency experienced by each VM increases substantially, which translates into longer I/O processing latency perceived by I/O-bound applications. To mitigate such impact while retaining the benefit of CPU sharing, we introduce a new class of VMs called latency-sensitive VMs (LSVMs), which achieve better performance for I/O-bound applications while maintaining the same resource share (and thus cost) as other CPU-sharing VMs. LSVMs are enabled by vSlicer, a hypervisor-level technique that schedules each LSVM more frequently but with a smaller micro time slice. vSlicer enables more timely processing of I/O events by LSVMs, without violating the CPU share fairness among all sharing VMs. Our evaluation of a vSlicer prototype in Xen shows that vSlicer substantially reduces network packet round-trip times and jitter and improves application-level performance. For example, vSlicer doubles both the connection rate and request processing throughput of an Apache web server; reduces a VoIP server's upstream jitter by 62%; and shortens the execution times of Intel MPI benchmark programs by half or more. Cong Xu 0010, Sahan Gamage, Pawan N. Rao, Ardalan Kangarlou, Ramana Rao Kompella, Dongyan Xu |
HPDC | 5 |
| 2012 | Network Sampling Designs for Relational Classification
Nesreen K. Ahmed, Jennifer Neville, Ramana Rao Kompella |
ICWSM | 3 |
| 2012 | MAPLE: a scalable architecture for maintaining packet latency measurementsabstractLatency has become an important metric for network monitoring since the emergence of new latency-sensitive applications (e.g., algorithmic trading and high-performance computing). To satisfy the need, researchers have proposed new architectures such as LDA and RLI that can provide fine-grained latency measurements. However, these architectures are fundamentally ossified in their design as they are designed to provide only a specific pre-configured aggregate measurement---either average latency across all packets (LDA) or per-flow latency measurements (RLI). Network operators, however, need latency measurements at both finer (e.g., packet) as well as flexible (e.g., flow subsets) levels of granularity. To bridge this gap, we propose an architecture called MAPLE that essentially stores packet-level latencies in routers and allows network operators to query the latency of arbitrary traffic sub-populations. MAPLE is built using scalable data structures with small storage needs (uses only 12.8 bits/packet), and uses a novel mechanism to reduce the query bandwidth significantly (by a factor of 17 compared to the naive method of sending packet queries individually). Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella |
Internet Measurement Conference | 3 |
| 2012 | Cuckoo sampling: Robust collection of flow aggregates under a fixed memory budgetabstractCollecting per-flow aggregates in high-speed links is challenging and usually requires traffic sampling to handle peak rates and extreme traffic mixes. Static selection of sampling rates is problematic, since worst-case resource usage is orders of magnitude higher than the average. To address this issue, adaptive schemes have been proposed in the last few years that periodically adjust packet sampling rates to network conditions. However, such proposals rely on complex algorithms and data structures of costly maintenance. As a consequence, adaptive sampling is still not widely implemented in routers. Josep Sanjuàs-Cuxart, Pere Barlet-Ros, Nick G. Duffield, Ramana Rao Kompella |
INFOCOM | 4 |
| 2012 | The TCP Outcast Problem: Exposing Unfairness in Data Center Networks
Pawan Prakash, Advait Abhay Dixit, Y. Charlie Hu, Ramana Rao Kompella |
NSDI | 4 |
| 2012 | The only constant is change: incorporating time-varying network reservations in data centersabstractIn multi-tenant datacenters, jobs of different tenants compete for the shared datacenter network and can suffer poor performance and high cost from varying, unpredictable network performance. Recently, several virtual network abstractions have been proposed to provide explicit APIs for tenant jobs to specify and reserve virtual clusters (VC) with both explicit VMs and required network bandwidth between the VMs. However, all of the existing proposals reserve a fixed bandwidth throughout the entire execution of a job. Di Xie, Ning Ding 0004, Y. Charlie Hu, Ramana Rao Kompella |
SIGCOMM | 4 |
| 2012 | On the efficacy of fine-grained traffic splitting protocols in data center networksabstractCurrent multipath routing techniques split traffic at a per-flow level because, according to conventional wisdom, forwarding packets of a TCP flow along different paths leads to packet reordering which is detrimental to TCP. In this paper, we revisit this "myth" in the context of cloud data center networks which have regular topologies such as multi-rooted trees. We argue that due to the symmetry in the multiple equal-cost paths in such networks, simply spraying packets of a given flow among all equal-cost paths, leads to balanced queues across multiple paths, and consequently little packet reordering. Using a testbed comprising of NetFPGA switches, we show how cloud applications benefit from better network utilization in data centers. Advait Abhay Dixit, Pawan Prakash, Ramana Rao Kompella, Y. Charlie Hu |
SIGMETRICS | 3 |
| 2012 | A scalable architecture for maintaining packet latency measurementsabstractLatency has become an important metric for network monitoring since the emergence of new latency-sensitive applications (e.g., algorithmic trading and high-performance computing). In this paper, to provide latency measurements at both finer (e.g., packet) as well as flexible (e.g., flow subsets) levels of granularity, we propose an architecture called MAPLE that essentially stores packet-level latencies in routers and allows network operators to query the latency of arbitrary traffic sub-populations. MAPLE is built using a scalable data structure called SVBF with small storage needs. Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella |
SIGMETRICS | 3 |
| 2012 | Router Support for Fine-Grained Latency MeasurementsabstractAn increasing number of datacenter network applications, including automated trading and high-performance computing, have stringent end-to-end latency requirements where even microsecond variations may be intolerable. The resulting fine-grained measurement demands cannot be met effectively by existing technologies, such as SNMP, NetFlow, or active probing. We propose instrumenting routers with a hash-based primitive that we call a Lossy Difference Aggregator (LDA) to measure latencies down to tens of microseconds even in the presence of packet loss. Because LDA does not modify or encapsulate the packet, it can be deployed incrementally without changes along the forwarding path. When compared to Poisson-spaced active probing with similar overheads, our LDA mechanism delivers orders of magnitude smaller relative error; active probing requires 50-60 times as much bandwidth to deliver similar levels of accuracy. Although ubiquitous deployment is ultimately desired, it may be hard to achieve in the shorter term; we discuss a partial deployment architecture called mPlane using LDAs for intrarouter measurements and localized segment measurements for interrouter measurements. Ramana Rao Kompella, Kirill Levchenko, Alex C. Snoeren, George Varghese |
IEEE/ACM Trans. Netw. | 1 |
| 2012 | Opportunistic Flow-Level Latency Estimation Using Consistent NetFlowabstractThe inherent measurement support in routers (SNMP counters or NetFlow) is not sufficient to diagnose performance problems in IP networks, especially for flow-specific problems where the aggregate behavior within a router appears normal. Tomographic approaches to detect the location of such problems are not feasible in such cases as active probes can only catch aggregate characteristics. To address this problem, in this paper, we propose a Consistent NetFlow (CNF) architecture for measuring per-flow delay measurements within routers. CNF utilizes the existing NetFlow architecture that already reports the first and last timestamps per flow, and it proposes hash-based sampling to ensure that two adjacent routers record the same flows. We devise a novel Multiflow estimator that approximates the intermediate delay samples from other background flows to significantly improve the per-flow latency estimates compared to the naive estimator that only uses actual flow samples. In our experiments using real backbone traces and realistic delay models, we show that the Multiflow estimator is accurate with a median relative error of less than 20% for flows of size greater than 100 packets. We also show that Multiflow estimator performs two to three times better than a prior approach based on trajectory sampling at an equivalent packet sampling rate. Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella |
IEEE/ACM Trans. Netw. | 3 |
| 2011 | Opportunistic flooding to improve TCP transmit performance in virtualized cloudsabstractVirtualization is a key technology that powers cloud computing platforms such as Amazon EC2. Virtual machine (VM) consolidation, where multiple VMs share a physical host, has seen rapid adoption in practice with increasingly large number of VMs per machine and per CPU core. Our investigations, however, suggest that the increasing degree of VM consolidation has serious negative effects on the VMs' TCP transport performance. As multiple VMs share a given CPU, the scheduling latencies, which can be in the order of tens of milliseconds, substantially increase the typically sub-millisecond round-trip times (RTTs) for TCP connections in a datacenter, causing significant degradation in throughput. In this paper, we propose a light-weight solution called vFlood that (a) allows a TCP sender VM to opportunistically flood the driver domain in the same host, and (b) offloads the VM's TCP congestion control function to the driver domain in order to mask the effects of VM consolidation. Our evaluation of a vFlood prototype on Xen suggests that vFlood substantially improves TCP transmit throughput with minimal per-packet CPU overhead. Further, our application-level evaluation using Apache Olio, a web 2.0 cloud application, indicates a 33% improvement in the number of operations per second. Sahan Gamage, Ardalan Kangarlou, Ramana Rao Kompella, Dongyan Xu |
SoCC | 3 |
| 2011 | Sketching the delay: tracking temporally uncorrelated flow-level latenciesabstractPacket delay is a crucial performance metric for real-time, network-based applications. Obtaining per-flow delay measurements is particularly important to network operators, but is computationally challenging in high-speed links. Recently, passive delay measurement techniques have been proposed that outperform traditional active probing in terms of accuracy and network overhead. However, such techniques rely on the empirical observation that packet delays across different flows are temporally correlated, an assumption that is not met in presence of traffic prioritization, load balancing policies, or due to intricacies of the switch fabric. Josep Sanjuàs-Cuxart, Pere Barlet-Ros, Nick G. Duffield, Ramana Rao Kompella |
Internet Measurement Conference | 4 |
| 2011 | Scheduling in mapreduce-like systems for fast completion timeabstractLarge-scale data processing needs of enterprises today are primarily met with distributed and parallel computing in data centers. MapReduce has emerged as an important programming model for these environments. Since today's data centers run many MapReduce jobs in parallel, it is important to find a good scheduling algorithm that can optimize the completion times of these jobs. While several recent papers focused on optimizing the scheduler, there exists very little theoretical understanding of the scheduling problem in the context of MapReduce. In this paper, we seek to address this problem by first presenting a simplified abstraction of the MapReduce scheduling problem, and then formulate the scheduling problem as an optimization problem.We devise various online and offline algorithms to arrive at a good ordering of jobs to minimize the overall job completion times. Since optimal solutions are hard to compute (NP-hard), we propose approximation algorithms that work within a factor of 3 of the optimal. Using simulations, we also compare our online algorithm with standard scheduling strategies such as FIFO, Shortest Job First and show that our algorithm consistently outperforms these across different job distributions. Hyunseok Chang, Murali S. Kodialam, Ramana Rao Kompella, T. V. Lakshman, Myungjin Lee, Sarit Mukherjee |
INFOCOM | 3 |
| 2011 | RelSamp: Preserving application structure in sampled flow measurementsabstractThe Internet has significantly evolved in the number and variety of applications. Network operators need mechanisms to constantly monitor and study these applications. Given modern applications routinely consist of several flows, potentially to many different destinations, existing measurement approaches such as Sampled NetFlow sample only a few flows per application session. To address this issue, in this paper, we introduce RelSamp architecture that implements the notion of related sampling where flows that are part of the same application session are given higher probability. In our evaluation using real traces, we show that RelSamp achieves 5-10x more flows per application session compared to Sampled NetFlow for the same effective number of sampled packets. We also show that behavioral and statistical classification approaches such as BLINC, SVM and C4.5 achieve up to 50% better classification accuracy compared to Sampled NetFlow, while not breaking existing management tasks such as volume estimation. Myungjin Lee, Mohammad Y. Hajjat, Ramana Rao Kompella, Sanjay G. Rao |
INFOCOM | 3 |
| 2011 | On the efficacy of fine-grained traffic splitting protocolsin data center networksabstractMulti-rooted tree topologies are commonly used to construct high-bandwidth data center network fabrics. In these networks, switches typically rely on equal-cost multipath (ECMP) routing techniques to split traffic across multiple paths, such that packets within a flow traverse the same end-to-end path. Unfortunately, since ECMP splits traffic based on flow-granularity, it can cause load imbalance across paths resulting in poor utilization of network resources. More fine-grained traffic splitting techniques are typically not preferred because they can cause packet reordering that can, according to conventional wisdom, lead to severe TCP throughput degradation. In this work, we revisit this fact in the context of regular data center topologies such as fat-tree architectures. We argue that packet-level traffic splitting, where packets of a flow are sprayed through all available paths, would lead to a better load-balanced network, which in turn leads to significantly more balanced queues and much higher throughput compared to ECMP. Advait Abhay Dixit, Pawan Prakash, Ramana Rao Kompella |
SIGCOMM | 3 |
| 2011 | Fine-grained latency and loss measurements in the presence of reorderingabstractModern trading and cluster applications require microsecond latencies and almost no losses in data centers. This paper introduces an algorithm called FineComb that can estimate fine-grain end-to-end loss and latency measurements between edge routers in these data center networks. Such a mechanism can allow managers to distinguish between latencies and loss singularities caused by servers and those caused by the network. Compared to prior work, such as Lossy Difference Aggregator (LDA), that focused on switch-level latency measurements, the requirement of end-to-end latency measurements introduces the challenge of reordering that occurs commonly in IP networks due to churn. The problem is even more acute in switches across data center networks that employ multipath routing algorithms to exploit the inherent path diversity. Without proper care, a loss estimation algorithm can confound loss and reordering; further, any attempt to aggregate delay estimates in the presence of reordering results in severe errors. FineComb deals with these problems using order-agnostic packet digests and a simple new idea we call stash recovery. Our evaluation demonstrates that FineComb can provide orders of magnitude better accuracy in loss and delay estimates in the presence of reordering compared to LDA. Myungjin Lee, Sharon Goldberg, Ramana Rao Kompella, George Varghese |
SIGMETRICS | 3 |
| 2010 | Two Samples are Enough: Opportunistic Flow-level Latency Estimation using NetFlowabstractThe inherent support in routers (SNMP counters or NetFlow) is not sufficient to diagnose performance problems in IP networks, especially for flow-specific problems and hence, the aggregate behavior within a router appears normal. To address this problem, in this paper, we propose a Consistent NetFlow (CNF) architecture for measuring per-flow performance measurements within routers. CNF utilizes NetFlow architecture that already reports the first and last timestamps per-flow, and hash-based sampling for ensuring that two routers record same flows. We devise a novel Multiflow estimator that approximates the intermediate delay samples from other background flows to improve the per-flow latency estimates significantly compared to the naive estimator that only uses actual flow samples. In our experiments using real backbone traces and realistic delay models, we show that Multiflow estimator is accurate with a median relative error of less than 20% for flows of size greater than 100 packets. We also show that prior approach based on trajectory sampling performs about 2-3× worse. Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella |
INFOCOM | 3 |
| 2010 | PhishNet: Predictive Blacklisting to Detect Phishing AttacksabstractPhishing has been easy and effective way for trickery and deception on the Internet. While solutions such as URL blacklisting have been effective to some degree, their reliance on exact match with the blacklisted entries makes it easy for attackers to evade. We start with the observation that attackers often employ simple modifications (e.g., changing top level domain) to URLs. Our system, PhishNet, exploits this observation using two components. In the first component, we propose five heuristics to enumerate simple combinations of known phishing sites to discover new phishing URLs. The second component consists of an approximate matching algorithm that dissects a URL into multiple components that are matched individually against entries in the blacklist. In our evaluation with real-time blacklist feeds, we discovered around 18,000 new phishing URLs from a set of 6,000 new blacklist entries. We also show that our approximate matching algorithm leads to very few false positives (3%) and negatives (5%). Pawan Prakash, Ramana Rao Kompella |
INFOCOM | 3 |
| 2010 | dFault: Fault Localization in Large-Scale Peer-to-Peer Systems
Pawan Prakash, Ramana Rao Kompella, Venugopalan Ramasubramanian, Ranveer Chandra |
Middleware | 2 |
| 2010 | vSnoop: Improving TCP Throughput in Virtualized Environments via Acknowledgement OffloadabstractVirtual machine (VM) consolidation has become a common practice in clouds, Grids, and datacenters. While this practice leads to higher CPU utilization, we observe its negative impact on the TCP throughput of the consolidated VMs: As more VMs share the same core/CPU, the CPU scheduling latency for each VM increases significantly. Such increase leads to slower progress of TCP transmissions to the VMs. To address this problem, we propose an approach called vSnoop, where the driver domain of a host acknowledges TCP packets on behalf of the guest VMs - whenever it is safe to do so. Our evaluation of a Xen-based prototype indicates that vSnoop constantly achieves TCP throughput improvement for VMs (of orders of magnitude in some scenarios). We further show that the higher TCP throughput leads to improvement in application- level performance, via experiments with a two-tier online auction application and two suites of MPI benchmarks. Ardalan Kangarlou, Sahan Gamage, Ramana Rao Kompella, Dongyan Xu |
SC | 3 |
| 2010 | Not all microseconds are equal: fine-grained per-flow measurements with reference latency interpolationabstractNew applications such as algorithmic trading and high-performance computing require extremely low latency (in microseconds). Network operators today lack sufficient fine-grain measurement tools to detect, localize and repair performance anomalies and delay spikes that cause application SLA violations. A recently proposed solution called LDA provides a scalable way to obtain latency, but only provides aggregate measurements. However, debugging application-specific problems requires per-flow measurements, since different flows may exhibit significantly different characteristics even when they are traversing the same link. To enable fine-grained per-flow measurements in routers, we propose a new scalable architecture called reference latency interpolation (RLI) that is based on our observation that packets potentially belonging to different flows that are closely spaced to each other exhibit similar delay properties. In our evaluation using simulations over real traces, we show that RLI achieves a median relative error of 12% and one to two orders of magnitude higher accuracy than previous per-flow measurement solutions with small overhead. Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella |
SIGCOMM | 3 |
| 2010 | Shedding Light on Enterprise Network Failures Using SpotlightabstractFault localization in enterprise networks is extremely challenging. A recent approach called Sherlock makes some headway into this problem by using an inference algorithm over a multi-tier probabilistic dependency graph that relates fault symptoms with possible root causes (e.g., routers, servers). A key limitation of Sherlock is its scalability because of the use of complicated inference algorithms based on Bayesian networks. We present a fault localization system called Spotlight that essentially uses two basic ideas. First, it compresses a multi-tier dependency graph into a bipartite graph with direct probabilistic edges between root causes and symptoms. Second, it runs a novel weighted greedy minimum set cover algorithm to provide fast inference. Through extensive simulations with real service dependency graphs and enterprise network topologies reported previously in literature, we show that Spotlight is about 100× faster than Sherlock in typical settings, with comparable accuracy in diagnosis. Dipu John, Pawan Prakash, Ramana Rao Kompella, Ranveer Chandra |
SRDS | 3 |
| 2010 | CLAMP: Efficient class-based sampling for flexible flow monitoring
Mohit Saxena, Ramana Rao Kompella |
Comput. Networks | 2 |
| 2010 | Fault Localization via Risk ModelingabstractInternet backbone networks are under constant flux in order to keep up with demand and offer new features. The pace of change in technology often outstrips the pace of introduction of associated fault monitoring capabilities that are built into today's IP protocols and routers. Moreover, some of these new technologies cross networking layers, raising the potential for unanticipated interactions and service disruptions, which the individual layers' built-in monitoring capabilities may not detect. In these instances, operators typically employ higher layer monitoring techniques such as end-to-end liveness probing to detect lower or cross-layer failures, but lack tools to precisely determine where a detected failure may have occurred. In this paper, we evaluate the effectiveness of using risk modeling to translate high-level failure notifications into lower layer root causes in two specific scenarios in a tier-1 ISP. We show that a simple greedy heuristic works with accuracy exceeding 80 percent for many failure scenarios in simulation, while delivering extremely high precision (greater than 80 percent). We report our operational experience using risk modeling to isolate optical component and MPLS control plane failures in an ISP backbone. Ramana Rao Kompella, Jennifer Yates, Albert G. Greenberg, Alex C. Snoeren |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2010 | Covenant: An architecture for cooperative scheduling in 802.11 wireless networksabstractWireless networks based on 802.11a/b/g protocols have gained wide-spread acceptance in both enterprise as well as home networks. However, these devices lack native support for many advanced features such as service differentiation, etc., that are required in specific application domains. In this paper, we propose Covenant, a software based cooperative scheduling framework to provide a rich set of features for applications that require nodes to cooperate with each other to satisfy system-wide objectives. We propose a novel 2 1/2-stage pipeline architecture as an efficient mechanism to implement cooperative scheduling among multiple nodes. We demonstrate how Covenant can be easily implemented in software, thus requiring absolutely no hardware or firmware changes to the already widely installed base of 802.11a/b/g based wireless devices. We also evaluate, using a real Linux based test-bed with Covenant drivers, the efficacy of the approach on two different scheduling disciplines: proportional priority and strict priority. We demonstrate that these scheduling disciplines are effective in providing service guarantees to multimedia applications even in the presence of other competing traffic. Ishwar Ramani, Ramana Rao Kompella, Sriram Ramabhadran, Alex C. Snoeren |
IEEE Trans. Wirel. Commun. | 2 |
| 2009 | A Framework for Efficient Class-Based SchedulingabstractWith an increasing requirement for network monitoring tools to classify traffic and track security threats, newer and efficient ways are needed for collecting traffic statistics and monitoring of network flows. However, traditional solutions based on random packet sampling treat all flows as equal and therefore, do not provide the flexibility required for these applications. In this paper, we propose a novel architecture called CLAMP that provides an efficient framework to implement size-based sampling. At the heart of CLAMP is a novel data structure called composite bloom filter (CBF) that consists of a set of bloom filters that work together to encapsulate various class definitions. In comparison to previous approaches that implement simple size-based sampling, our architecture requires substantially lower memory (upto 80x) and results in higher flow coverage (upto 8x more flows) under specific configurations. Mohit Saxena, Ramana Rao Kompella |
INFOCOM | 2 |
| 2009 | Every microsecond counts: tracking fine-grain latencies with a lossy difference aggregatorabstractMany network applications have stringent end-to-end latency requirements, including VoIP and interactive video conferencing, automated trading, and high-performance computing---where even microsecond variations may be intolerable. The resulting fine-grain measurement demands cannot be met effectively by existing technologies, such as SNMP, NetFlow, or active probing. We propose instrumenting routers with a hash-based primitive that we call a Lossy Difference Aggregator (LDA) to measure latencies down to tens of microseconds and losses as infrequent as one in a million.Such measurement can be viewed abstractly as what we refer to as a coordinated streaming problem, which is fundamentally harder than standard streaming problems due to the need to coordinate values between nodes. We describe a compact data structure that efficiently computes the average and standard deviation of latency and loss rate in a coordinated streaming environment. Our theoretical results translate to an efficient hardware implementation at 40 Gbps using less than 1% of a typical 65-nm 400-MHz networking ASIC. When compared to Poisson-spaced active probing with similar overheads, our LDA mechanism delivers orders of magnitude smaller relative error; active probing requires 50--60 times as much bandwidth to deliver similar levels of accuracy. Ramana Rao Kompella, Kirill Levchenko, Alex C. Snoeren, George Varghese |
SIGCOMM | 1 |
| 2008 | cSamp: A System for Network-Wide Flow Monitoring
Vyas Sekar, Michael K. Reiter, Walter Willinger, Hui Zhang 0001, Ramana Rao Kompella, David G. Andersen |
NSDI | 5 |
| 2008 | Designing packet buffers for router linecards
Sundar Iyer, Ramana Rao Kompella, Nick McKeown |
IEEE/ACM Trans. Netw. | 2 |
| 2007 | Detection and Localization of Network Black HolesabstractInternet backbone networks are under constant flux, struggling to keep up with increasing demand. The pace of technology change often outstrips the deployment of associated fault monitoring capabilities that are built into today's IP protocols and routers. Moreover, some of these new technologies cross networking layers, raising the potential for unanticipated interactions and service disruptions that the built-in monitoring systems cannot detect. In such instances, failures may cause data packets to be silently dropped inside the network without triggering any alarms or responses (e.g., the failure is not routed around). So-called "silent failures" or "black holes" represent a critical threat to today's rapidly evolving networks. In this paper, we present a simple and effective method to detect and diagnose such silent failures. Our method uses active measurement between edge routers to raise alarms whenever end-to-end connectivity is disrupted, regardless of the cause. These alarms feed localization agents that employ spatial correlation techniques to isolate the root-cause of failure. Using data from two real systems deployed on sections of a tier-I ISP network, we successfully detect and localize three known black holes. Further, we present simulation results demonstrating that our system accurately and precisely (both greater than 80% according to our metrics) localizes a variety of failures classes. Ramana Rao Kompella, Jennifer Yates, Albert G. Greenberg, Alex C. Snoeren |
INFOCOM | 1 |
| 2007 | On scalable attack detection in the network
Ramana Rao Kompella, Sumeet Singh, George Varghese |
IEEE/ACM Trans. Netw. | 1 |
| 2005 | The Power of Slicing in Internet Flow Measurement
Ramana Rao Kompella, Cristian Estan |
Internet Measurement Conference | 1 |
| 2005 | IP Fault Localization Via Risk Modeling
Ramana Rao Kompella, Jennifer Yates, Albert G. Greenberg, Alex C. Snoeren |
NSDI | 1 |
| 2004 | On scalable attack detection in the networkabstractCurrent intrusion detection and prevention systems seek to detect a wide class of network intrusions (e.g., DoS attacks, worms, port scans)at network vantage points. Unfortunately, all the IDS systems we know of keep per-connection or per-flow state. Thus it is hardly surprising that IDS systems (other than signature detection mechanisms) have not scaled to multi-gigabit speeds. By contrast, note that both router lookups and fair queuing have scaled to high speeds using aggregation via prefix lookups or DiffServ. Thus in this paper, we initiate research into the question as to whether one can detect attacks without keeping per-flow state. We will show that such aggregation, while making fast implementations possible, immediately cause two problems. First, aggregation can cause behavioral aliasing where, for example, good behaviors can aggregate to look like bad behaviors. Second, aggregated schemes are susceptible to spoofing by which the intruder sends attacks that have appropriate aggregate behavior. We examine a wide variety of DoS attacks and show that several categories (bandwidth based, claim-and-hold, host scanning) can be scalably detected. By contrast, it appears that stealthy port-scanning cannot be scalably detected without keeping per-flow state. Ramana Rao Kompella, Sumeet Singh, George Varghese |
Internet Measurement Conference | 1 |
| 2004 | Reduced state fair queuing for edge and core routersabstractDespite many years of research, fair queuing still faces a number of implementation challenges in high speed routers. In particular, in spite of proposals such as DiffServ, the state needs for even simple schedulers are still large for heavily channelized core routers and for edge routers. An earlier proposal, Stochastic Fair Queuing, reduces state but at the expense of added unfairness between certain flows. Another earlier scheme, Core Stateless Fair Queuing, requires header changes and does not address the state needs of edge routers. By contrast, our paper proposes a randomization technique that removes the need to store deficit counters per flow in Deficit Round Robin and its variants. Even without the counters, we show, using both analysis and simulation, that randomized technique preserves throughput fairness properties of DRR. This randomization technique introduced in this paper can be used to considerably reduce the state requirements of high speed schedulers in edge and core routers, making hardware designs feasible. The randomization idea in this paper can also be applied to other round robin schedulers as well as potentially in entirely different scenarios wherever deficits need to be tracked over time explicitly. Ramana Rao Kompella, George Varghese |
NOSSDAV | 1 |
| 2003 | Practical lazy scheduling in sensor networksabstractExperience has shown that the power consumption of sensors and other wireless computational devices is often dominated by their communication patterns. We present a practical realization of lazy packet scheduling that attempts to minimize the total transmission energy in a broadcast network by dynamically adjusting each node's transmission power and rate on a per-packet basis. Lazy packet scheduling leverages the fact that many channel coding schemes are more efficient at lower transmission rates; that is, the energy required to send a fixed amount of data can be reduced by transmitting the data at a lower bit rate and transmission powe.The optimal per-packet transmission rate in a multi-node network is governed in practice by the available bit rates of the given transceiver(s), the nodes' delay tolerance, and the offered load at every node contending for the shared broadcast channel. We propose an extension to the traditional CSMACA MAC scheme called L-CSMACA that allows individual nodes to continually estimate the current demand for a broadcast channel and adjust their transmission schedules accordingly. Our simulation results show that L-CSMACA can provide improved energy efficiency in a single-hop, broadcast network (20--25% with more than 10 nodes, and up to 99% for four nodes with a standard power function) for both Poisson and bursty arrivals with only minor degradation the capacity of the channe. Ramana Rao Kompella, Alex C. Snoeren |
SenSys | 1 |