EDBT 2026 Demo / reviewers in the wild / expert
Hani Jamjoom
dblp:45/1350
· DBLP profile ↗
43ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0001-6143-1064ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 5 since 2021Computer networks · 13 · 5 first-authorSecurity and privacy · 9 · 7 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GPU Travelling: Efficient Confidential Collaborative Training with TEE-Enabled GPUsabstractConfidential collaborative machine learning (ML) enables multiple mutually distrusted data holders to jointly train an ML model while preserving the confidentiality of their private datasets due to regulatory or competitive reasons. However, existing works need frequent data and model exchanges during training via slower conventional links. They face increasing challenges due to the exponentially growing sizes of models and datasets in modern training workloads like large language models (LLMs), resulting in prohibitively high communication costs. In this paper, we propose a novel mechanism called GPU Travelling that leverages recently emerged confidential GPUs. With our rigorous design, the GPU can securely travel to the specific data holder to load the dataset directly into the GPU's protected memory and then return for training, eliminating the need for data transmission while ensuring confidentiality up to a data-centre level. We developed a prototype using Intel TDX and NVIDIA H100 and evaluated its performance on llm.c, a CUDA-based LLM training project, and demonstrated the performance and feasibility while maintaining strong security guarantees. The results showed at least 4x speed improvement when transmitting a 512 MiB dataset chunk versus conventional transmission. Shixuan Zhao 0002, Zhongshu Gu, Salman Ahmed 0001, Enriquillo Valdez, Hani Jamjoom, Zhiqiang Lin 0001 |
CCS | 5 |
| 2025 | Rex: Closing the language-verifier gap with safe and usable kernel extensions
Jinghao Jia, Ruowen Qin, Milo Craun, Egor Lukiyanov, Ayush Bansal, Minh Phan, Michael V. Le, Hubertus Franke, Hani Jamjoom, Tianyin Xu, Dan Williams 0001 |
USENIX ATC | 9 |
| 2024 | Crossing Shifted Moats: Replacing Old Bridges with New Tunnels to Confidential ContainersabstractThe Confidential Containers (CoCo) project, as an open-source community initiative, inherits the system architecture of Kata Containers while integrating confidential computing to protect cloud-native container workloads. However, there exists a misalignment in the threat model and trusted computing base (TCB) between Kata Containers and confidential computing. The shifted trust boundaries could potentially expose a range of vulnerabilities, particularly in scenarios where a malicious actor on the host gains access to the CoCo's unprotected control interface. This paper conducts a thorough examination of CoCo's system architecture, exploring the attack surface resulting from the discord in trust boundaries. We have assessed all API endpoints of CoCo's control interface, categorizing them based on their security properties. Drawing from these insights, we have developed a bifurcation approach to splitting CoCo's control interface. This involves establishing an owner-side controller and minimizing the capabilities of the existing host-side controller. Under this framework, the host-side controller is exclusively responsible for allocating and recycling compute resources, while dedicated workload owners can directly manage their containers through alternative secure tunnels. This approach ensures seamless integration with cloud-native orchestration layers and aligns CoCo with the threat model of confidential computing. By doing so, it effectively prevents untrusted hosts from accessing confidential data and interfering with the execution of workloads within protected domains. Enriquillo Valdez, Salman Ahmed 0001, Zhongshu Gu, Christophe de Dinechin, Pau-Chen Cheng, Hani Jamjoom |
CCS | 6 |
| 2024 | DeTA: Minimizing Data Leaks in Federated Learning via Decentralized and Trustworthy AggregationabstractFederated learning (FL) relies on a central authority to oversee and aggregate model updates contributed by multiple participating parties in the training process. This centralization of sensitive model updates naturally raises concerns about the trustworthiness of the central aggregation server, as well as the potential risks associated with server failures or breaches, which could result in loss and leaks of model updates. Moreover, recent attacks have demonstrated that, by obtaining the leaked model updates, malicious actors can even reconstruct substantial amounts of private data belonging to training participants. This underscores the critical necessity to rethink the existing FL system architecture to mitigate emerging attacks in the evolving threat landscape. One straightforward approach is to fortify the central aggregator with confidential computing (CC), which offers hardware-assisted protection for runtime computation and can be remotely verified for execution integrity. However, a growing number of security vulnerabilities have surfaced in tandem with the adoption of CC, indicating that depending solely on this singular defense may not provide the requisite resilience to thwart data leaks. Pau-Chen Cheng, Kevin Eykholt, Zhongshu Gu, Hani Jamjoom, K. R. Jayaram, Enriquillo Valdez, Ashish Verma 0001 |
EuroSys | 4 |
| 2024 | GNNIC: Finding Long-Lost Sibling Functions with Abstract Similarity
Qiushi Wu, Zhongshu Gu, Hani Jamjoom, Kangjie Lu |
NDSS | 3 |
| 2024 | Fast (Trapless) Kernel Probes Everywhere
Jinghao Jia, Michael V. Le, Salman Ahmed 0001, Dan Williams 0001, Hani Jamjoom, Tianyin Xu |
USENIX ATC | 5 |
| 2024 | SeaK: Rethinking the Design of a Secure Allocator for OS Kernel
Zicheng Wang 0010, Yicheng Guang, Yueqi Chen 0001, Zhenpeng Lin, Michael V. Le, Dang K. Le, Dan Williams 0001, Xinyu Xing 0001, Zhongshu Gu, Hani Jamjoom |
USENIX Security Symposium | 10 |
| 2023 | Securing Container-based Clouds with Syscall-aware SchedulingabstractContainer-based clouds—in which containers are the basic unit of isolation—face security concerns because, unlike Virtual Machines, containers directly interface with the underlying highly privileged kernel through the wide and vulnerable system call interface. Regardless of whether a container itself requires dangerous system calls, a compromised or malicious container sharing the host (a bad neighbor) can compromise the host kernel using a vulnerable syscall, thereby compromising all other containers sharing the host. Michael V. Le, Salman Ahmed 0001, Dan Williams 0001, Hani Jamjoom |
AsiaCCS | 4 |
| 2023 | Not All Data are Created Equal: Data and Pointer Prioritization for Scalable Protection Against Data-Oriented Attacks
Salman Ahmed 0001, Hans Liljestrand, Hani Jamjoom, Matthew Hicks, N. Asokan, Danfeng Yao |
USENIX Security Symposium | 3 |
| 2021 | Glitching Demystified: Analyzing Control-flow-based Glitching Attacks and DefensesabstractHardware fault injection, or glitching, attacks can compromise the security of devices even when no software vulnerabilities exist. Attempts to analyze the hardware effects of glitching are subject to the Heisenberg effect and there is typically a disconnect between what people “think” is possible and what is actually possible with respect to these attacks. In this work, we attempt to provide some clarity to the impacts of attacks and defenses for control-flow modification through glitching. First, we introduce a glitching emulation framework, which provides a scalable playground to test the effects of bit flips on specific instruction set architectures (ISAs) (i.e., the fault tolerance of the instruction encoding). Next, we examine real glitching experiments using the ChipWhisperer, a popular microcontroller using open-source glitching hardware. These real-world experiments provide novel insights into how glitching attacks are realized and might be defended against in practice. Finally, we present GLITCHRESISTOR, an open-source, software-based glitching defense tool that can automatically insert glitching defenses into any existing source code, in an architecture-independent way. We evaluated GLITCHRESISTOR, which integrates numerous software-only defenses against powerful and real-world glitching attacks. Our findings indicate that software-only defenses can be implemented with acceptable run-time and size overheads, while completely mitigating some single-glitch attacks, minimizing the likelihood of a successful multi-glitch attack (i.e., a success rate of 0.000306%), and detecting failed glitching attempts at a high rate (between 79.2% and 100%). Chad Spensky, Aravind Machiry, Nathan Burow, Hamed Okhravi, Rick Housley, Zhongshu Gu, Hani Jamjoom, Christopher Krügel, Giovanni Vigna |
DSN | 7 |
| 2021 | Confidential computing for OpenPOWERabstractThis paper presents Protected Execution Facility (PEF), a virtual machine-based Trusted Execution Environment (TEE) for confidential computing on Power ISA. PEF enables protected secure virtual machines (SVMs). Like other TEEs, PEF verifies the SVM prior to execution. PEF utilizes a Trusted Platform Module (TPM), secure boot, and trusted boot as well as newly introduced architectural changes for Power ISA systems. Exploiting these architectural changes requires new firmware, the Protected Execution Ultravisor. PEF is supported in the latest version of the POWER9 chip. PEF demonstrates that access control for isolation and cryptography for confidentiality is an effective approach to confidential computing. We particularly focus on how our design (i) balances between access control and cryptography, (ii) maximizes the use of existing security components, and (iii) simplifies the management of the SVM life cycle. Finally, we evaluate the performance of SVMs in comparison to normal virtual machines on OpenPOWER systems. Guerney D. H. Hunt, Ramachandra Pai, Michael V. Le, Hani Jamjoom, Sukadev Bhattiprolu, Rick Boivie, Laurent Dufour, Brad Frey, Mohit Kapur, Kenneth A. Goldman, Ryan Grimm, Janani Janakirman, John M. Ludden, Paul Mackerras, Cathy May, Elaine R. Palmer, Bharata Bhasker Rao, Lawrence Roy, William A. Starke, Jeffrey Stuecheli, Enriquillo Valdez, Wendel Voigt |
EuroSys | 4 |
| 2019 | Houdini's Escape: Breaking the Resource Rein of Linux Control GroupsabstractLinux Control Groups, i.e., cgroups, are the key building blocks to enable operating-system-level containerization. The cgroups mechanism partitions processes into hierarchical groups and applies different controllers to manage system resources, including CPU, memory, block I/O, etc. Newly spawned child processes automatically copy cgroups attributes from their parents to enforce resource control. Unfortunately, inherited cgroups confinement via process creation does not always guarantee consistent and fair resource accounting. In this paper, we devise a set of exploiting strategies to generate out-of-band</>workloads via de-associating processes from their original process groups. The system resources consumed by such workloads will not be charged to the appropriate cgroups. To further demonstrate the feasibility, we present five case studies within Docker containers to demonstrate how to break the resource rein of cgroups in realistic scenarios. Even worse, by exploiting those cgroups' insufficiencies in a multi-tenant container environment, an adversarial container is able to greatly amplify the amount of consumed resources, significantly slow-down other containers on the same host, and gain extra unfair advantages on the system resources. We conduct extensive experiments on both a local testbed and an Amazon EC2 cloud dedicated server. The experimental results demonstrate that a container can consume system resources (e.g., CPU) as much as $200\times$ of its limit, and reduce both computing and I/O performance of particular workloads in other co-resident containers by 95%. Xing Gao 0001, Zhongshu Gu, Zhengfa Li, Hani Jamjoom, Cong Wang 0006 |
CCS | 4 |
| 2019 | Reaching Data Confidentiality and Model Accountability on the CalTrainabstractDistributed collaborative learning (DCL) paradigms enable building joint machine learning models from distrusted multi-party participants. Data confidentiality is guaranteed by retaining private training data on each participant's local infrastructure. However, this approach makes today's DCL design fundamentally vulnerable to data poisoning and backdoor attacks. It limits DCL's model accountability, which is key to backtracking problematic training data instances and their responsible contributors. In this paper, we introduce CALTRAIN, a centralized collaborative learning system that simultaneously achieves data confidentiality and model accountability. CALTRAIN enforces isolated computation via secure enclaves on centrally aggregated training data to guarantee data confidentiality. To support building accountable learning models, we securely maintain the links between training instances and their contributors. Our evaluation shows that the models generated by CALTRAIN can achieve the same prediction accuracy when compared to the models trained in non-protected environments. We also demonstrate that when malicious training participants tend to implant backdoors during model training, CALTRAIN can accurately and precisely discover the poisoned or mislabeled training data that lead to the runtime mispredictions. Zhongshu Gu, Hani Jamjoom, Dong Su, Heqing Huang 0001, Jialong Zhang 0001, Tengfei Ma 0001, Dimitrios E. Pendarakis, Ian M. Molloy |
DSN | 2 |
| 2016 | CRONets: Cloud-Routed Overlay NetworksabstractOverlay networking and ISP-assisted tunneling are effective solutions to overcome problematic BGP routes and bypass troublesome autonomous systems. Despite their demonstrated effectiveness, overlay support is not broadly available. In this paper, we propose Cloud-Routed Overlay Networks (CRONets), whereby users can readily build their own overlays using nodes from global and well-provisioned cloud providers like IBM Softlayer or Amazon EC2. While previous studies have demonstrated the benefits of overlay networks with the high-speed experimental Internet2 backbone, we are the first to evaluate the improvements in a realistic -- cloud -- setting. We conduct a large-scale experiment where we observe 6,600 Internet paths. The results show that CRONets improve the throughput for 78% of the default Internet paths with a median and average improvement factors of 1.67 and 3.27 times respectively, at a tenth of the cost of leasing private lines of comparable performance. We also performed a longitudinal measurement, and demonstrate that the performance gains are consistent over time with only a small number of overlay nodes needed to be deployed. However, given the size and dynamic nature of the Internet routing system (e.g., due to congestion and failures), selecting the proper path is still a challenging problem. To address it, we propose a novel solution based on the newly-introduced MPTCP extensions. Our experiments show that MPTCP can achieve the maximum observed throughput across the different overlay paths. Chris X. Cai, Franck Le, Xin Sun 0002, Geoffrey G. Xie, Hani Jamjoom, Roy H. Campbell |
ICDCS | 5 |
| 2016 | Gremlin: Systematic Resilience Testing of MicroservicesabstractModern Internet applications are being disaggregated into a microservice-based architecture, with services being updated and deployed hundreds of times a day. The accelerated software life cycle and heterogeneity of language runtimes in a single application necessitates a new approach for testing the resiliency of these applications in production infrastructures. We present Gremlin, a framework for systematically testing the failure-handling capabilities of microservices. Gremlin is based on the observation that microservices are loosely coupled and thus rely on standard message-exchange patterns over the network. Gremlin allows the operator to easily design tests and executes them by manipulating inter-service messages at the network layer. We show how to use Gremlin to express common failure scenarios and how developers of an enterprise application were able to discover previously unknown bugs in their failure-handling code without modifying the application. Victor Heorhiadi, Shriram Rajagopalan, Hani Jamjoom, Michael K. Reiter, Vyas Sekar |
ICDCS | 3 |
| 2016 | Version Traveler: Fast and Memory-Efficient Version Switching in Graph Processing Systems
Xiaoen Ju, Dan Williams 0001, Hani Jamjoom, Kang G. Shin |
USENIX ATC | 3 |
| 2016 | Enabling Efficient Hypervisor-as-a-Service Clouds with Eemeral VirtualizationabstractWhen considering a hypervisor, cloud providers must balance conflicting requirements for simple, secure code bases with more complex, feature-filled offerings. This paper introduces Dichotomy, a new two-layer cloud architecture in which the roles of the hypervisor are split. The cloud provider runs a lean hyperplexor that has the sole task of multiplexing hardware and running more substantial hypervisors (called featurevisors) that implement features. Cloud users choose featurevisors from a selection of lightly-modified hypervisors potentially offered by third-parties in an "as-a-service" model for each VM. Rather than running the featurevisor directly on the hyperplexor using nested virtualization, Dichotomy uses a new virtualization technique called eemeral virtualization which efficiently (and repeatedly) transfers control of a VM between the hyperplexor and featurevisor using memory mapping techniques. Nesting overhead is only incurred when the VM is accessed by the featurevisor. We have implemented Dichotomy in KVM/QEMU and demonstrate average switching times of 80 ms, two to three orders of magnitude faster than live VM migration. We show that, for the featurevisor applications we evaluated, VMs hosted in Dichotomy deliver up to 12% better performance than those hosted on nested hypervisors, and continue to show benefit even when the featurevisor applications run as often as every 2.5~seconds. Dan Williams 0001, Yaohui Hu, Umesh Deshpande, Piush K. Sinha, Nilton Bila, Kartik Gopalan, Hani Jamjoom |
VEE | 7 |
| 2015 | Flux: multi-surface computing in AndroidabstractWith the continued proliferation of mobile devices, apps will increasingly become multi-surface, running seamlessly across multiple user devices (e.g., phone, tablet, etc.). Yet general systems support for multi-surface app is limited to (1) screencasting, which relies on a single master device's computing power and battery life or (2) cloud backing, which is unsuitable in the face of disconnected operation or untrusted cloud providers. We present an alternative approach: Flux, an Android-based system that enables any app to become multi-surface through app migration. Flux overcomes device heterogeneity and residual dependencies through two key mechanisms. Selective Record/Adaptive Replay records just those device-agnostic app calls that lead to the generation of app-specific device-dependent state in system services and replays them on the target. Checkpoint/Restore in Android (CRIA) transitions an app into a state in which device-specific information can be safely discarded before checkpointing and restoring the app. Our implementation of Flux can migrate many popular, unmodified Android apps---including those with extensive device interactions like 3D accelerated graphics---across heterogeneous devices and is fast enough for interactive use. Alexander Van't Hof, Hani Jamjoom, Jason Nieh, Dan Williams 0001 |
EuroSys | 2 |
| 2014 | TideWatch: Fingerprinting the cyclicality of big data workloadsabstractIntrinsic to “big data” processing workloads (e.g., iterative MapReduce, Pregel, etc.) are cyclical resource utilization patterns that are highly synchronized across different resource types as well as the workers in a cluster. In Infrastructure as a Service settings, cloud providers do not exploit this characteristic to better manage VMs because they view VMs as “black boxes.” We present TideWatch, a system that automatically identifies cyclicality and similarity in running VMs. TideWatch predicts period lengths of most VMs in Hadoop workloads within 9% of actual iteration boundaries and successfully classifies up to 95% of running VMs as participating in the appropriate Hadoop cluster. Furthermore, we show how TideWatch can be used to improve the timing of VM migrations, reducing both migration time and network impact by over 50% when compared to a random approach. Dan Williams 0001, Shuai Zheng 0002, Xiangliang Zhang 0001, Hani Jamjoom |
INFOCOM | 4 |
| 2013 | Pico replication: a high availability framework for middleboxesabstractMiddleboxes are being rearchitected to be service oriented, composable, extensible, and elastic. Yet system-level support for high availability (HA) continues to introduce significant performance overhead. In this paper, we propose Pico Replication (PR), a system-level framework for middleboxes that exploits their flow-centric structure to achieve low overhead, fully customizable HA. Unlike generic (virtual machine level) techniques, PR operates at the flow level. Individual flows can be checkpointed at very high frequencies while the middlebox continues to process other flows. Furthermore, each flow can have its own checkpoint frequency, output buffer and target for backup, enabling rich and diverse policies that balance---per-flow---performance and utilization. PR leverages OpenFlow to provide near instant flow-level failure recovery, by dynamically rerouting a flow's packets to its replication target. We have implemented PR and a flow-based HA policy. In controlled experiments, PR sustains checkpoint frequencies of 1000Hz, an order of magnitude improvement over current VM replication solutions. As a result, PR drastically reduces the overhead on end-to-end latency from 280% to 15.5% and throughput overhead from 99.5% to 3.2%. Shriram Rajagopalan, Dan Williams 0001, Hani Jamjoom |
SoCC | 3 |
| 2013 | Mizan: a system for dynamic load balancing in large-scale graph processingabstractPregel [23] was recently introduced as a scalable graph mining system that can provide significant performance improvements over traditional MapReduce implementations. Existing implementations focus primarily on graph partitioning as a preprocessing step to balance computation across compute nodes. In this paper, we examine the runtime characteristics of a Pregel system. We show that graph partitioning alone is insufficient for minimizing end-to-end computation. Especially where data is very large or the runtime behavior of the algorithm is unknown, an adaptive approach is needed. To this end, we introduce Mizan, a Pregel system that achieves efficient load balancing to better adapt to changes in computing needs. Unlike known implementations of Pregel, Mizan does not assume any a priori knowledge of the structure of the graph or behavior of the algorithm. Instead, it monitors the runtime characteristics of the system. Mizan then performs efficient fine-grained vertex migration to balance computation and communication. We have fully implemented Mizan; using extensive evaluation we show that---especially for highly-dynamic workloads---Mizan provides up to 84% improvement over techniques leveraging static graph pre-partitioning. Zuhair Khayyat, Karim Awara, Amani AlOnazi, Hani Jamjoom, Dan Williams 0001, Panos Kalnis |
EuroSys | 4 |
| 2013 | Escape Capsule: Explicit State Is Robust and Scalable
Shriram Rajagopalan, Dan Williams 0001, Hani Jamjoom, Andy Warfield |
HotOS | 3 |
| 2013 | What to discover before migrating to the cloud
Niyu Ge, Hani Jamjoom, Ea-Ee Jan, Lakshminarayanan Renganarayana, Xiaolan Zhang 0001 |
IM | 3 |
| 2013 | Split/Merge: System Support for Elastic Execution in Virtual Middleboxes
Shriram Rajagopalan, Dan Williams 0001, Hani Jamjoom, Andy Warfield |
NSDI | 3 |
| 2013 | To 4, 000 compute nodes and beyond: network-aware vertex placement in large-scale graph processing systemsabstractThe explosive growth of "big data" is giving rise to a new breed of large scale graph systems, such as Pregel. This poster describes our ongoing work in characterizing and minimizing the communication cost of Bulk Synchronous Parallel (BSP) graph mining systems, like Pregel, when scaling to 4,096 compute nodes. Existing implementations generally assume a fixed communication cost. This is sufficient in small deployments as the BSP programming model (i.e., overlapping computation and communication) masks small variations in the underlying network. In large scale deployments, such variations can dominate the overall runtime characteristics. In this poster, we first quantify the impact of network communication on the total compute time of a Pregel system. We then propose an efficient vertex placement strategy that subsamples highly connected vertices and applies the Reverse Cuthill-McKee (RCM) algorithm to efficiently partition the input graph and place partitions closer to each other based on their expected communication patterns. We finally describe a vertex replication strategy to further reduce communication overhead. Karim Awara, Hani Jamjoom, Panos Kalnis |
SIGCOMM | 2 |
| 2012 | The Xen-Blanket: virtualize once, run everywhereabstractCurrent Infrastructure as a Service (IaaS) clouds operate in isolation from each other. Slight variations in the virtual machine (VM) abstractions or underlying hypervisor services prevent unified access and control across clouds. While standardization efforts aim to address these issues, they will take years to be agreed upon and adopted, if ever. Instead of standardization, which is by definition provider-centric, we advocate a user-centric approach that gives users an unprecedented level of control over the virtualization layer. We introduce the Xen-Blanket, a thin, immediately deployable virtualization layer that can homogenize today's diverse cloud infrastructures. We have deployed the Xen-Blanket across Amazon's EC2, an enterprise cloud, and a private setup at Cornell University. We show that a user-centric approach to homogenize clouds can achieve similar performance to a paravirtualized environment while enabling previously impossible tasks like cross-provider live migration. The Xen-Blanket also allows users to exploit resource management opportunities like oversubscription, and ultimately can reduce costs for users. Dan Williams 0001, Hani Jamjoom, Hakim Weatherspoon |
EuroSys | 2 |
| 2012 | Virtual machine migration in an over-committed cloudabstractWhile early emphasis of Infrastructure as a Service (IaaS) clouds was on providing resource elasticity to end users, providers are increasingly interested in over-committing their resources to maximize the utilization and returns of their capital investments. In principle, over-committing resources hedges that users — on average — only need a small portion of their leased resources. When such hedge fails (i.e., resource demand far exceeds available physical capacity), providers must mitigate this provider-induced overload, typically by migrating virtual machines (VMs) to underutilized physical machines. Recent works on VM placement and migration assume the availability of target physical machines [1], [2]. However, in an over-committed cloud data center, this is not the case. VM migration can even trigger cascading overloads if performed haphazardly. In this paper, we design a new VM migration algorithm (called Scattered) that minimizes VM migrations in over-committed data centers. Compared to a traditional implementation, our algorithm can balance host utilization across all time epochs. Using real-world data traces from an enterprise cloud, we show that our migration algorithm reduces the risk of overload, minimizes the number of needed migrations, and has minimal impact on communication cost between VMs. Xiangliang Zhang 0001, Zon-Yin Shae, Shuai Zheng 0002, Hani Jamjoom |
NOMS | 4 |
| 2011 | Analysis and Modeling of Social Influence in High Performance Computing Workloads
Shuai Zheng 0002, Zon-Yin Shae, Xiangliang Zhang 0001, Hani Jamjoom, Liana L. Fong |
Euro-Par (1) | 4 |
| 2011 | Application-aware virtual machine migration in data centersabstractWhile virtual machine (VM) migration is allowing data centers to rebalance workloads across physical machines, the promise of a maximally utilized infrastructure is yet to be realized. Part of the challenge is due to the inherent dependencies between VMs comprising a multi-tier application, which introduce complex load interactions between the underlying physical servers. For example, simply moving an overloaded VM to a (random) underloaded physical machine can inadvertently overload the network. We introduce AppAware-a novel, computationally efficient scheme for incorporating (1) inter-VM dependencies and (2) the underlying network topology into VM migration decisions. Using simulations, we show that our proposed method decreases network traffic by up to 81%compared to a well known alternative VM migration method that is not application-aware. Vivek Shrivastava, Petros Zerfos, Hani Jamjoom, Yew-Huey Liu, Suman Banerjee 0001 |
INFOCOM | 4 |
| 2011 | Overdriver: handling memory overload in an oversubscribed cloudabstractWith the intense competition between cloud providers, oversubscription is increasingly important to maintain profitability. Oversubscribing physical resources is not without consequences: it increases the likelihood of overload. Memory overload is particularly damaging. Contrary to traditional views, we analyze current data center logs and realistic Web workloads to show that overload is largely transient: up to 88.1% of overloads last for less than 2 minutes. Regarding overload as a continuum that includes both transient and sustained overloads of various durations points us to consider mitigation approaches also as a continuum, complete with tradeoffs with respect to application performance and data center overhead. In particular, heavyweight techniques, like VM migration, are better suited to sustained overloads, whereas lightweight approaches, like network memory, are better suited to transient overloads. We present Overdriver, a system that adaptively takes advantage of these tradeoffs, mitigating all overloads within 8% of well-provisioned performance. Furthermore, under reasonable oversubscription ratios, where transient overload constitutes the vast majority of overloads, Overdriver requires 15% of the excess space and generates a factor of four less network traffic than a migration-only approach. Dan Williams 0001, Hani Jamjoom, Yew-Huey Liu, Hakim Weatherspoon |
VEE | 2 |
| 2010 | A service composition framework for market-oriented high performance computing cloudabstractDespite the success High Performance Computing (HPC) across a number of application domains, the adoption of HPC resources and applications is still limited, primarily due to its high capital cost, system complexity, application availability, and service delivery model. Recently, several research efforts have shown that the emerging Cloud Com-puting service model can improve on-demand access to HPC capacity as utility. This paper introduces a framework for on-demand composing and deploying available HPC applica-tions as services on HPC clouds. The composition is enabled by an ontology that describes dependencies and relationships among HPC software and resources. Tran Vu Pham, Hani Jamjoom, Kirk E. Jordan, Zon-Yin Shae |
HPDC | 2 |
| 2009 | Rule-Based Problem Classification in IT Service ManagementabstractProblem management is a critical and expensive element for delivering IT service management and touches various levels of managed IT infrastructure. While problem management has been mostly reactive, recent work is studying how to leverage large problem ticket information from similar IT infrastructures to probatively predict the onset of problems. Because of the sheer size and complexity of problem tickets, supervised learning algorithms have been the method of choice for problem ticket classification, relying on labeled (or pre-classified) tickets from one managed infrastructure to automatically create signatures for similar infrastructures. However, where there are insufficient preclassified data, leveraging human expertise to develop classification rules can be more efficient. In this paper, we describe a rule-based crowdsourcing approach, where experts can author classification rules and a social networking-based platform (called xPad) is used to socialize and execute these rules by large practitioner communities. Using real data sets from several large IT delivery centers, we demonstrate that this approach balances between two key criteria: accuracy and cost effectiveness. Yixin Diao, Hani Jamjoom, David Loewenstern |
IEEE CLOUD | 2 |
| 2009 | iPoG: fast interactive proximity querying on graphsabstractGiven an author-conference graph, how do we answer proximity queries (e.g., what are the most related conferences for John Smith?); how can we tailor the search result if the user provides additional yes/no type of feedback (e.g., what are the most related conferences for John Smith given that he does not like ICML?)? Given the potential computational complexity, we mainly devote ourselves to addressing the computational issues in this paper by proposing an efficient solution (referred to as iPoG-B) for bipartite graphs. Our experimental results show that the proposed fast solution (iPoGB) achieves significant speedup, while leading to the same ranking result. Hanghang Tong, Huiming Qu, Hani Jamjoom, Christos Faloutsos |
CIKM | 3 |
| 2008 | Measuring Proximity on Graphs with Side InformationabstractThis paper studies how to incorporate side information (such as users' feedback) in measuring node proximity on large graphs. Our method (ProSIN) is motivated by the well-studied random walk with restart (RWR). The basic idea behind ProSIN is to leverage side information to refine the graph structure so that the random walk is biased towards/away from some specific zones on the graph. Our case studies demonstrate that ProSIN is well-suited in a variety of applications, including neighborhood search, center-piece subgraphs, and image caption. Given the potential computational complexity of ProSIN, we also propose a fast algorithm (Fast-ProSIN) that exploits the smoothness of the graph structures with/without side information. Our experimental evaluation shows that fast-ProSIN achieves significant speedups (up to 49x) over straightforward implementations. Hanghang Tong, Huiming Qu, Hani Jamjoom |
ICDM | 3 |
| 2008 | Enterprise mashups and Web 2.0 for management
Hani Jamjoom, Nikos Anerousis |
NOMS | 1 |
| 2007 | Service Assurance Process Re-Engineering Using Location-aware Infrastructure IntelligenceabstractThe continuous introduction of converged services such as VoIP and Video-On-Demand has created many operational challenges for service providers. In this paper, we describe how to use location-aware technologies, not only to integrate disparate management applications, but also to transform the underlying process to use geographical views as the focal point of management operations. Based on an engagement with a large Cable provider, we have designed and implemented 3i - Integrated Infrastructure Intelligence - to address key issues in the service assurance process. 3i is highly componentized and provides an intuitive way for creating role-based views through dynamic scoping, event aggregation and status projection, and location-driven active probing. We analyze the current service assurance process and compare it with the improved process after introducing 3i. Overall, the re-engineered process offers base execution improvements in alarm collection, problem drill-down and reporting, as well as complexity improvements throughout. 3i is a fully implemented tool and has demonstrated capabilities beyond its original intended scope as a decision support tool in planning and marketing functions. Hani Jamjoom, Nikos Anerousis, Raymond B. Jennings III, Debanjan Saha |
Integrated Network Management | 1 |
| 2007 | Networkmd: topology inference and failure diagnosis in the last mileabstractHealth monitoring, automated failure localization and diagnosis have all become critical to service providers of large distribution networks (e.g., digital cable and fiber-to-the-home), due to the increases in scale and complexity of their offered services. Existing automated failure diagnosis solutions typically assume complete knowledge of network topology, which in practice is rarely available. The solution presented in this paper - Network Management and Diagnosis (NetworkMD) - is an automated failure diagnosis system that can infer failure groups based on historical failure data, and optionally geographical information. The inferred failure groups mirror missing topologies, and can be used to localize failures, diagnose root causes of problems, and detect misconfiguration in known topologies. NetworkMD uses an unsupervised learning algorithm based on non-negative matrix factorization (NMF) to infer failure groups. Using cable network as the primary example, we demonstrate the effectiveness of NetworkMD in both simulated settings and real environment using data collected from a commercial network serving hundreds of thousands of customers via thousands of intermediate network devices. Yun Mao, Hani Jamjoom, Shu Tao, Jonathan M. Smith |
Internet Measurement Conference | 2 |
| 2006 | Failure diagnosis with incomplete information in cable networksabstractCable network has become one of the most popular ways of high-speed Internet access for homes and small businesses. For broadband cable providers, managing their large-scale network infrastructure is highly challenging because these networks are geographically dispersed and contain a large number of devices. A single administrative area typically serves hundreds of thousands of end-customers (or cable modems) with thousands of intermediate distribution devices that operate on different protocol layers. Such devices include routers, Cable Modem Termination Systems (CMTS's), fiber nodes, repeaters, etc. It is, thus, critical for the providers to monitor the health of their infrastructure and perform quick failure diagnosis. However, there are many challenges in failure diagnosis. In this paper, we focus on two challenges that are motivated by input from a large U.S. cable provider: missing device status information and incomplete topology. Yun Mao, Hani Jamjoom, Shu Tao |
CoNEXT | 2 |
| 2006 | On the role and controllability of persistent clients in traffic aggregates
Hani Jamjoom, Kang G. Shin |
IEEE/ACM Trans. Netw. | 1 |
| 2004 | The Impact of Concurrency Gains on the Analysis and Control of Multi-threaded Internet ServicesabstractWith the proliferation of Internet services, many solutions have emerged to provide quality-of-service (QoS) guarantees when the demands for the hosted services exceed the server's capacity. We take an analytical approach to answering key questions in the design and performance of application-level QoS techniques, especially those that are based on the multi-threading or multi-processing abstraction. Key to our analysis is the integration of the effects of concurrency into the interactions between multi-threaded services. To this end, we extend traditional time-sharing models to develop the multi-threaded round-robin (MTRR) servers, a more accurate model of operation of typical multi-threaded Internet services. For this model, we first develop powerful, yet computationally-efficient, mathematical relationships that describe the performance (in terms of throughput and response time) of multi-threaded services. We then apply optimization techniques to derive the optimal allocation of threads given specific QoS objective functions. Using realistic workloads on a typical Web server, we show the efficacy and accuracy of the proposed new methodology. Hani Jamjoom, Chun-Ting Chou, Kang G. Shin |
INFOCOM | 1 |
| 2004 | Resynchronization and controllability of bursty service requestsabstractThere is an increasing prevalence of interactive Web sessions in the Internet. These are mostly short-lived TCP connections that are delay-sensitive and have transfer times dominated by TCP backoffs, if any, during connection establishment. Unfortunately, arrivals of such connections at a server tend to be bursty, and can trigger multiple retransmissions, resulting in long average client-perceived delays. Traditional traffic control mechanisms, such as token bucket filters, are designed to complement admission control mechanisms, by regulating throughput, bounding service times, and protecting systems from overload. However, they cannot control connection-establishment delays, and thus, do not provide effective control of client-perceived delays. We first present the surprising discovery of a resynchronization property of retransmitted requests that exacerbates client-perceived delays when traditional control mechanisms are used. Then, we introduce a novel, multistage filtering scheme called Abacus Filters (AFs) that limits the client-perceived delay while maximizing server throughput even in the case of bursty connection arrivals. Analysis of delay-control properties of various filtering mechanisms is presented, along with a detailed performance evaluation. AFs are shown to provide tight delay control and better complement traditional admission control policies. Hani Jamjoom, Padmanabhan Pillai, Kang G. Shin |
IEEE/ACM Trans. Netw. | 1 |
| 2003 | Persistent dropping: an efficient control of traffic aggregatesabstractFlash crowd events (FCEs) present a real threat to the stability of routers and end-servers. Such events are characterized by a large and sustained spike in client arrival rates, usually to the point of service failure. Traditional rate-based drop policies, such as Random Early Drop (RED), become ineffective in such situations since clients tend to be persistent, in the sense that they make multiple retransmission attempts before aborting their connection. As it is built into TCP's congestion control, this persistence is very widespread, making it a major stumbling block to providing responsive aggregate traffic controls. This paper focuses on analyzing and building a coherent model of the effects of client persistence on the controllability of aggregate traffic. Based on this model, we propose a new drop strategy called persistent dropping to regulate the arrival of SYN packets and achieves three important goals: (1) it allows routers and end-servers to quickly converge to their control targets without sacrificing fairness, (2) it minimizes the portion of client delay that is attributed to the applied controls, and (3) it is both easily implementable and computationally tractable. Using a real implementation of this controller in the Linux kernel, we demonstrate its efficacy, up to 60% delay reduction for drop probabilities less than 0.5. Hani Jamjoom, Kang G. Shin |
SIGCOMM | 1 |
| 2001 | Adaptive packet filtersabstractAdaptive packet filters (APFs) are motivated by the proliferation of distributed servers and the lack of quality-of-service (QoS) management solutions for them. APFs merge packet-filtering and server load monitoring into a novel load-sensitive packet-filtering abstraction for overload protection and QoS differentiation. They integrate well into network protocol stacks and firewalls, scale to large server farms while remaining completely transparent to the applications. Experimental results of our prototype implementation demonstrate APFs' efficacy in providing QoS differentiation and overload protection with minimal overheads. John Reumann, Hani Jamjoom, Kang G. Shin |
GLOBECOM | 2 |