H. Howie Huang

dblp:39/560 · DBLP profile ↗
← Back
79ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0001-8588-7680ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 7 first-author · 5 since 2021Security and privacy · 13 · 10 since 2021Databases, data management, data science and information retrieval · 12 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Computer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Generalizable Graph-based Reinforcement Learning Agents for Automated Cyber Defense
Isaiah J. King, Benjamin Bowman, H. Howie Huang
DSN3
2026 From Time and Place to Preference: LLM-Driven Geo-Temporal Context in Recommendations
abstract
Recommender systems focus on timestamps as numeric or cyclical values, often ignoring real-world context like seasons, holidays and events. We present a scalable framework that utilizes large language models (LLMs) to create geo-temporal embeddings from timestamps and coarse locations, capturing holidays, seasonal trends, and local/global events. We then introduce a geo-temporal embedding informativeness test as a lightweight diagnostic, demonstrating on MovieLens, LastFM, and a large-scale production dataset that these embeddings provide a predictive signal consistent with the outcomes of full model integrations. Geo-temporal embeddings were integrated into sequential models via feature fusion with metadata. Our findings underscore the importance of adaptive and hybrid strategies for improving recommendations. We also release a context-enriched MovieLens dataset. https://github.com/yejinjennyK/movielens-1m-geo-temporal-context https://huggingface.co/datasets/yejinjennyK/movielens-1m-geo-temporal-context.
Yejin Kim 0005, Shaghayegh Agah, Mayur Nankani, Maria Peifer, Feifei Peng, H. Howie Huang, Sardar Hamidian
SIGIR7
2025 Exploring the Efficacy of Multi-Agent Reinforcement Learning for Autonomous Cyber Defence: A CAGE Challenge 4 Perspective
abstract
As cyber threats become increasingly automated and sophisticated, novel solutions must be introduced to improve defence of enterprise networks. Deep Reinforcement Learning (DRL) has demonstrated potential in mitigating these advanced threats. Single DRL Agents have proven utility toward execution of autonomous cyber defence. Despite the success of employing single DRL Agents, this approach presents significant limitations, especially regarding scalability within large enterprise networks. An attractive alternative to the single agent approach is the use of Multi-Agent Reinforcement Learning (MARL). However, developing MARL agents is costly with few options for examining MARL cyber defence techniques against adversarial agents. This paper presents a MARL network security environment, the fourth iteration of the Cyber Autonomy Gym for Experimentation (CAGE) challenges. This challenge was specifically designed to test the efficacy of MARL algorithms in an enterprise network. Our work aims to evaluate the potential of MARL as a robust and scalable solution for autonomous network defence.
Mitchell Kiely, Metin Ahiskali, Etienne Borde, Benjamin Bowman, David Bowman, Dirk Van Bruggen, KC Cowan, Prithviraj Dasgupta, Erich Devendorf, Ben Edwards, Alex Fitts, Sunny Fugate, Ryan Gabrys, Wayne Gould, H. Howie Huang, Jules Jacobs, Ryan Kerr, Isaiah J. King, Li Li 0009, Luis Martinez, Christopher Moir, Craig Murphy, Olivia Naish, Claire Owens, Miranda Purchase, Ahmad Ridley, Adrian Taylor, Sara Farmer, William John Valentine, Yiyi Zhang 0002
AAAI15
2025 Demystifying optimized prompts in language models
abstract
Modern language models (LMs) are not robust to out-of-distribution inputs.Machine generated ("optimized") prompts can be used to modulate LM outputs and induce specific behaviors while appearing completely uninterpretable.In this work, we investigate the composition of optimized prompts, as well as the mechanisms by which LMs parse and build predictions from optimized prompts.We find that optimized prompts primarily consist of punctuation and noun tokens which are more rare in the training data.Internally, optimized prompts are clearly distinguishable from natural language counterparts based on sparse subsets of the model's activations.Across various families of instruction-tuned models, optimized prompts follow a similar path in how their representations form through the network. 1
Rimon Melamed, Lucas H. McCabe, H. Howie Huang
EMNLP3
2025 Revelio: Revealing Important Message Flows in Graph Neural Networks
abstract
Explainability is crucial for the deployment of Graph Neural Networks (GNNs) in real-world applications. Unfortunately, existing explanation methods primarily focus on identifying important graph components, such as nodes and edges, rather than providing insights into the fundamental message passing mechanisms of GNNs. This shortcoming impedes our understanding of how GNNs make predictions and limits their deployment in critical applications. In this paper, we introduce Revelio, a novel method to provide faithful explanations of message flows in GNNs. Revelio leverages a learning-based approach to quantify the importance of message flows, excelling in terms of faithfulness, compatibility, and efficiency. Our extensive experiments on both synthetic and real-world datasets demonstrate the superiority of Revelio through quantitative and qualitative assessments.
Isaiah J. King, H. Howie Huang
ICDE3
2025 Trail: A Knowledge Graph-Based Approach for Attributing Advanced Persistent Threats
abstract
Open-source intelligence exchanges provide a rich repository of indicators of compromise (IOCs). These IOCs are used to build detection signatures and blocklists in production cybersecurity environments as well as prior works. In this work, we investigate their utility for cyberattack attribution. To do this, we create a novel system called Trail that builds a knowledge graph of network-based IOC co-occurrences in cyberattacks, and their relations to other IOCs. After analyzing 4,500 cybersecurity events attributed to 22 different advanced persistent threats (APTs), the knowledge graph holds over 2.1 million nodes with 7.9 million edges. We analyze the knowledge graph this system produces using conventional machine learning, graph analytics, and a graph neural network to quantify the degree to which APTs leave identifiable clues in their IOCs. Using the Trail method to enrich the IOC feature space, IOCs can individually be attributed to the APT that generated them with 45% accuracy. When attributing groups of IOCs that made up cyberattacks, indirect resource reuse alone accurately attributed 82% of samples. When we used both graph topology and feature analysis and analyzed events with a graph neural network, attribution accuracy increased to 84%. Finally, we conducted a 6-month study of new cyber events our models had never seen. We found that our models continue to achieve similar accuracy on real-world data to what was observed experimentally, so long as the database is no more than 1 month out of date.
Isaiah J. King, Ramiro Ramirez, Benjamin Bowman, H. Howie Huang
ICDE4
2025 Predicting Movie Hits Before They Happen with LLMs
abstract
Addressing the cold-start issue in content recommendation remains a critical ongoing challenge.In this work, we focus on tackling the cold-start problem for movies on a large entertainment platform.Our primary goal is to forecast the popularity of cold-start movies using Large Language Models (LLMs) leveraging movie metadata.This method could be integrated into retrieval systems within the personalization pipeline or could be adopted as a tool for editorial teams to ensure fair promotion of potentially overlooked movies that may be missed by traditional or algorithmic solutions.Our study validates the effectiveness of this approach compared to established baselines and those we developed.
Shaghayegh Agah, Yejin Kim 0005, Mayur Nankani, Kevin Foley, H. Howie Huang, Sardar Hamidian
UMAP6
2025 SynCSE: syntax graph-based contrastive learning of sentence embeddings
Yejin Kim 0005, Dongsuk Oh, H. Howie Huang
Expert Syst. Appl.3
2024 Prov2vec: Learning Provenance Graph Representation for Anomaly Detection in Computer Systems
abstract
Modern cyber attackers use advanced zero-day exploits, highly targeted spear phishing, and other social engineering techniques to gain access, and also use evasion techniques to maintain a prolonged presence within the victim network while working gradually towards the objective. To minimize damage, detecting these Advanced Persistent Threats as early in the campaign as possible is crucial. This paper proposes, Prov2vec, a system for the continuous monitoring of enterprise host’s behavior to detect attackers’ activities. It leverages the data provenance graph built using system event logs to get complete visibility into the execution state of an enterprise host and the causal relationship between system entities. It proposes a novel provenance graph kernel to obtain the canonical representation of the system behavior, which is compared against its historical behaviors and that of other hosts to detect the deviation from the norm. These representations are used in several machine learning models to evaluate their ability to capture the underlying behavior of an endpoint host. We have empirically demonstrated that the provenance graph kernel produces a much more compact representation compared to existing methods while improving prediction ability.
Bibek Bhattarai, H. Howie Huang
ARES2
2024 Fine-grained Graph-based Anomaly Detection on Vehicle Controller Area Networks
abstract
Electronic components in vehicles communicate with one another by broadcasting messages over the controller area network (CAN) bus. The CAN message protocol is notoriously insecure, lacking both encryption and authentication for performance reasons. Vehicle manufacturers instead opt for "security through obscurity" and try to keep the meanings of CAN messages industry secrets. This approach has led to the discovery of several alarming, and unaddressed vulnerabilities. For this reason, it is imperative to develop a security monitoring system for the CAN bus. However, any such intrusion detection system is limited by severe memory constraints–in-vehicle ECUs rarely have more than 1MB of RAM. In this work, we explore the potential for lightweight graph kernel-based intrusion detection systems that work in conjunction with byte analysis of individual messages. Our approach extends the state-of-the-art in this field, which only classifies batches of messages as malicious or benign, rather than performing fine-grained anomaly detection. We analyze the precedence graph formed by CAN message ordering in conjunction with the bytes those messages contain to create a high-performance, low-memory anomaly detector. Our analysis revealed that this approach can detect a wide variety of attack types in both moving and stationary vehicles. We demonstrated that our method performs more precisely than prior works in the same field while requiring less than 100KB of memory.
Isaiah J. King, Benjamin Bowman, H. Howie Huang
IEEE Big Data3
2024 JITSPMM: Just-in-Time Instruction Generation for Accelerated Sparse Matrix-Matrix Multiplication
abstract
Achieving high performance for Sparse Matrix-Matrix Multiplication (SpMM) has received increasing research attention, especially on multi-core CPUs, due to the large input data size in applications such as graph neural networks (GNNs). Most existing solutions for SpMM computation follow the ahead-of-time (AOT) compilation approach, which compiles a program entirely before it is executed. AOT compilation for SpMM faces three key limitations: unnecessary memory access, additional branch overhead, and redundant instructions. These limitations stem from the fact that crucial information pertaining to SpMM is not known until runtime. In this paper, we propose JITSpMM, a just-in-time (JIT) assembly code generation framework to accelerated SpMM computation on multi-core CPUs with SIMD extensions. First, JITSpMM integrates the JIT assembly code generation technique into three widely-used workload division methods for SpMM to achieve balanced workload distribution among CPU threads. Next, with the availability of runtime information, JITSpMM employs a novel technique, coarse-grain column merging, to maximize instruction-level parallelism by unrolling the performance-critical loop. Furthermore, JITSpMM intelligently allocates registers to cache frequently accessed data to minimizing memory accesses, and employs selected SIMD instructions to enhance arithmetic throughput. We conduct a performance evaluation of JITSpMM and compare it two AOT baselines. The first involves existing SpMM implementations compiled using the Intel icc compiler with auto-vectorization. The second utilizes the highly-optimized SpMM routine provided by Intel MKL. Our results show that JITSpMM provides an average improvement of 3.8× and 1.4×, respectively.
Thomas B. Rolinger, H. Howie Huang
CGO3
2024 Improving Content Recommendation: Knowledge Graph-Based Semantic Contrastive Learning for Diversity and Cold-Start Users
abstract
Addressing the challenges related to data sparsity, cold-start problems, and diversity in recommendation systems is both crucial and demanding. Many current solutions leverage knowledge graphs to tackle these issues by combining both item-based and user-item collaborative signals. A common trend in these approaches focuses on improving ranking performance at the cost of escalating model complexity, reducing diversity, and complicating the task. It is essential to provide recommendations that are both personalized and diverse, rather than solely relying on achieving high rank-based performance, such as Click-through rate, Recall, etc. In this paper, we propose a hybrid multi-task learning approach, training on user-item and item-item interactions. We apply item-based contrastive learning on descriptive text, sampling positive and negative pairs based on item metadata. Our approach allows the model to better understand the relationships between entities within the knowledge graph by utilizing semantic information from text. It leads to more accurate, relevant, and diverse user recommendations and a benefit that extends even to cold-start users who have few interactions with items. We perform extensive experiments on two widely used datasets to validate the effectiveness of our approach. Our findings demonstrate that jointly training user-item interactions and item-based signals using synopsis text is highly effective. Furthermore, our results provide evidence that item-based contrastive learning enhances the quality of entity embeddings, as indicated by metrics such as uniformity and alignment.
Scott Rome, Kevin Foley, Mayur Nankani, Rimon Melamed, Abhay Yadav, Maria Peifer, Sardar Hamidian, H. Howie Huang
LREC/COLING10
2024 Prompts have evil twins
abstract
We discover that many natural-language prompts can be replaced by corresponding prompts that are unintelligible to humans but that provably elicit similar behavior in language models. We call these prompts “evil twins” because they are obfuscated and uninterpretable (evil), but at the same time mimic the functionality of the original natural-language prompts (twins). Remarkably, evil twins transfer between models. We find these prompts by solving a maximum-likelihood problem which has applications of independent interest.
Rimon Melamed, Lucas H. McCabe, Tanay Wakhare, H. Howie Huang, Enric Boix-Adserà
EMNLP5
2024 Maui: Black-Box Edge Privacy Attack on Graph Neural Networks
abstract
Graphs are ubiquitous data structures with nodes representing objects and edges representing relationships between them. Graph Neural Networks (GNNs) have recently been proposed to study graph-structured data, but unfortunately, are susceptible to privacy leakage. This issue becomes more urgent as GNNs gain wide deployment in many real-world settings including social network analysis, bioinformatics, and cybersecurity. In this paper, we propose the first link inference attack that can compromise user data under the most difficult security settings, which we call Maui. We demonstrate that private edge information can be inferred by a malicious user with a black-box approach. Extensive experiments on six real-world datasets show our attacks conduct effective link inference attacks in various scopes. Our attack achieves significant performance improvements over the current state-of-the-art. When targeting 2-layer Graph Convolution Networks, for inferring edges of a single node, our attack outperforms the best existing method by 12.0%, increasing from 83.7% to 95.7%; when inferring edges of the entire graph, our attack achieves a 19.6% improvement, from 67.7% to 87.3%. Our results underscore the need for countermeasures against privacy attacks in GNNs, as they can reveal rich information about graph structures.
Isaiah J. King, H. Howie Huang
Proc. Priv. Enhancing Technol.3
2023 EdgeTorrent: Real-time Temporal Graph Representations for Intrusion Detection
abstract
Anomaly-based intrusion detection aims to learn the normal behaviors of a system and detect activity that deviates from it. One of the best ways to represent the behavior of a computer network is through provenance graphs: dynamic networks of entity interactions over time. When provenance graphs deviate from their normal behaviors, it could be indicative of a malicious actor attempting to compromise the network. However, efficiently characterizing the normal behavior of large temporal graphs is challenging. To do this, we propose EdgeTorrent, an end-to-end anomaly-based intrusion detection system for provenance graph analysis. EdgeTorrent leverages a novel high-performance message passing neural network for graph embedding over a stream of edges to capture both temporal and topological changes in the system. These embeddings are then processed by a novel adversarially trained sequence analyzer that alerts when a series of graph embeddings changes in an unexpected way. EdgeTorrent preserves temporal ordering during message passing, and its streaming-focused design allows users to conduct out-of-core inference on billion-edge graphs, faster than real-time. We show that our method outperforms state-of-the-art graph-kernel approaches on several host monitoring data sets; notably, it is the first intrusion detection system to perfectly classify the StreamSpot data set. Additionally, we show it is the best-performing method on a real-world, billion-edge data set encompassing 11 days of benign and attack data.
Isaiah J. King, Xiaokui Shu, Jiyong Jang, Kevin Eykholt, Taesung Lee, H. Howie Huang
RAID6
2023 Euler: Detecting Network Lateral Movement via Scalable Temporal Link Prediction
abstract
Lateral movement is a key stage of system compromise used by advanced persistent threats. Detecting it is no simple task. When network host logs are abstracted into discrete temporal graphs, the problem can be reframed as anomalous edge detection in an evolving network. Research in modern deep graph learning techniques has produced many creative and complicated models for this task. However, as is the case in many machine learning fields, the generality of models is of paramount importance for accuracy and scalability during training and inference. In this article, we propose a formalized approach to this problem with a framework we call Euler . It consists of a model-agnostic graph neural network stacked upon a model-agnostic sequence encoding layer such as a recurrent neural network. Models built according to the Euler framework can easily distribute their graph convolutional layers across multiple machines for large performance improvements. Additionally, we demonstrate that Euler -based models are as good, or better, than every state-of-the-art approach to anomalous link detection and prediction that we tested. As anomaly-based intrusion detection systems, our models efficiently identified anomalous connections between entities with high precision and outperformed all other unsupervised techniques for anomalous lateral movement detection. Additionally, we show that as a piece of a larger anomaly detection pipeline, Euler models perform well enough for use in real-world systems. With more advanced, yet still lightweight, alerting mechanisms ingesting the embeddings produced by Euler models, precision is boosted from 0.243, to 0.986 on real-world network traffic.
Isaiah J. King, H. Howie Huang
ACM Trans. Priv. Secur.2
2022 SteinerLog: Prize Collecting the Audit Logs for Threat Hunting on Enterprise Network
abstract
Advanced cyberattacks are carried out in multiple stages, where each stage performs a specific task corresponding to the campaign. While these steps are designed to blend in with benign activities, they leave their activity footprints across multiple logs on the machines inside the victim environment. The majority of these footprints when looked at in isolation seem benign to the activity monitors. Existing threat hunting systems require a significant amount of human effort to correlate these events in order to detect and reconstruct an attack campaign. This paper introduces SteinerLog, an end-to-end system to automate the task of correlating the alerts to detect ongoing attack campaigns within an enterprise network. SteinerLog takes the alerts generated by mature intelligence-based and anomaly-based alerting systems and uses causal analysis to extract the group of events that are most likely to represent the attackers' activities. It performs hierarchical graph traversal to perform cross-host attacker activity correlation, which includes detecting the compromised entities, reconstructing the attackers' steps, and abstracting them into easy-to-understand attack graphs. The experiments show that it is able to detect APT campaigns in real-time and scale to an enterprise system with hundreds of workstations.
Bibek Bhattarai, H. Howie Huang
AsiaCCS2
2022 Graggle: A Graph-based Approach to Document Clustering
abstract
Document recommendation systems have traditionally relied upon high-dimensional vector representations that scale poorly in corpora with diverse vocabularies. Existing graph-based approaches focus on the metadata of documents and, unfortunately, ignore the content of the papers. In this work, we have designed and implemented a new system we call Graggle, which builds a graph to model a corpus. Nodes are papers, and edges represent significant words shared between them. We then leverage modern graph learning techniques to turn this graph into a highly efficient tool for dimensionality reduction. Documents are represented as low-dimensional vector embeddings generated with a graph autoencoder. Our experiments show that this approach outperforms traditional document vector-based and text autoencoding approaches on labeled data. Additionally, we have applied this technique to a repository of unlabeled research documents about the novel coronavirus to demonstrate its effectiveness as a real-world tool.
Isaiah J. King, H. Howie Huang
IEEE Big Data2
2022 Don't Judge a Language Model by Its Last Layer: Contrastive Learning with Layer-Wise Attention Pooling
abstract
Recent pre-trained language models (PLMs) achieved great success on many natural language processing tasks through learning linguistic features and contextualized sentence representation. Since attributes captured in stacked layers of PLMs are not clearly identified, straightforward approaches such as embedding the last layer are commonly preferred to derive sentence representations from PLMs. This paper introduces the attention-based pooling strategy, which enables the model to preserve layer-wise signals captured in each layer and learn digested linguistic features for downstream tasks. The contrastive learning objective can adapt the layer-wise attention pooling to both unsupervised and supervised manners. It results in regularizing the anisotropic space of pre-trained embeddings and being more uniform. We evaluate our model on standard semantic textual similarity (STS) and semantic search tasks. As a result, our method improved the performance of the base contrastive learned BERT_{base} and variants.
Dongsuk Oh, Yejin Kim 0005, Hodong Lee, H. Howie Huang, Heuiseok Lim
COLING4
2022 Illuminati: Towards Explaining Graph Neural Networks for Cybersecurity Analysis
abstract
Graph neural networks (GNNs) have been utilized to create multi-layer graph models for a number of cybersecurity applications from fraud detection to software vulnerability analysis. Unfortunately, like traditional neural networks, GNNs also suffer from a lack of transparency, that is, it is challenging to interpret the model predictions. Prior works focused on specific factor explanations for a GNN model. In this work, we have designed and implemented Illuminati, a comprehensive and accurate explanation framework for cybersecurity applications using GNN models. Given a graph and a pre-trained GNN model, Illuminati is able to identify the important nodes, edges, and attributes that are contributing to the prediction while requiring no prior knowledge of GNN models. We evaluate Illuminati in two cybersecurity applications, i.e., code vulnerability detection and smart contract vulnerability detection. The experiments show that Illuminati achieves more accurate explanation results than state-of-the-art methods, specifically, 87.6% of subgraphs identified by Illuminati are able to retain their original prediction, an improvement of 10.3% over others at 77.3%. Furthermore, the explanation of Illuminati can be easily understood by the domain experts, suggesting the significant usefulness for the development of cybersecurity applications.
Yuede Ji, H. Howie Huang
EuroS&P3
2022 TLPGNN: A Lightweight Two-Level Parallelism Paradigm for Graph Neural Network Computation on GPU
abstract
Graph Neural Networks (GNNs) are an emerging class of deep learning models on graphs, with many successful applications, such as, recommendation systems, drug discovery, and social network analysis. The GNN computation includes both regular neural network operations and general graph convolution operations, which take the majority of the total computation time. Though several recent works have been proposed to accelerate the computation for GNNs, they face the limitations of heavy pre-processing, low efficient atomic operations, and unnecessary kernel launches. In this paper, we design TLPGNN, a lightweight two-level parallelism paradigm for GNN computation. First, we conduct a systematic analysis on the hardware resource usage of GNN workloads to deeply understand the specialties of GNN workloads. With the insightful observations, we then divide the GNN computation into two levels, i.e., vertex parallelism for the first level and feature par- allelism for the second. Next, we employ a novel hybrid dynamic workload assignment to address the imbalanced workload distribution. Furthermore, we fuse the kernels to reduce the number of kernel launches and cache the frequently accessed data into registers to avoid unnecessary memory traffics. Together, TLPGNN is able to significantly outperform existing GNN computation systems, such as DGL, GNNAdivsor, and FeatGraph, by 5.6×, 7.7×, and 3.3×, respectively, on the average.
Yuede Ji, H. Howie Huang
HPDC3
2022 NestedGNN: Detecting Malicious Network Activity with Nested Graph Neural Networks
abstract
Network attacks are dramatically increasing over the years. A graph can accurately model the network activities. Therefore, graph-based techniques are frequently used to detect network threats. Motivated by the strong representation of graph neural networks (GNNs), many GNN-based techniques have been proposed for various security problems, such as network threat detection, malware detection, insider threat detection, and fraud detection. Most GNNs work on the classical attributed graph structure, while we observe that a nested graph structure is a more accurate representation for modelling enterprise network, where the communications between hosts form a graph, while the local activities of each host, e.g., local event graph, form an inner graph. Observing no existing GNNs can directly learn on such a nested graph, in this paper, we designed NestedGNN, the first graph neural network for nested graphs. NestedGNN consists of three layers, i.e., inner GNN layers, nested graph layers, and outer GNN layers. We successfully applied it to compromised host detection. NestedGNN can significantly improve the performance over traditional methods on a publicly available cybersecurity dataset.
Yuede Ji, H. Howie Huang
ICC2
2022 Mnemonic: A Parallel Subgraph Matching System for Streaming Graphs
abstract
Finding patterns in large highly connected datasets is critical for value discovery in business development and scientific research. This work focuses on the problem of subgraph matching on streaming graphs, which provides utility in a myriad of real-world applications ranging from social network analysis to cybersecurity. Each application poses a different set of control parameters, including the restrictions for a match, type of data stream, and search granularity. The problem-driven design of existing subgraph matching systems makes them challenging to apply for different problem domains. This paper presents Mnemonic, a programmable system that provides a high-level API and democratizes the development of a wide variety of subgraph matching solutions. Importantly, Mnemonic also delivers key data management capabilities and optimizations to support real-time processing on long-running, high-velocity multi-relational graph streams. The experiments demonstrate the versatility of Mnemonic, as it outperforms several state-of-the-art systems by up to two orders of magnitude.
Bibek Bhattarai, H. Howie Huang
IPDPS2
2022 Euler: Detecting Network Lateral Movement via Scalable Temporal Graph Link Prediction
Isaiah J. King, H. Howie Huang
NDSS2
2022 Pikachu: Temporal Walk Based Dynamic Graph Embedding for Network Anomaly Detection
abstract
Enterprise networks evolve constantly over time. In addition to the network topology, the order of information flow is crucial to detect cyber-threats in a constantly evolving network. Majority of the existing technique uses static snapshot to learn from dynamic network. However, using static snapshots is not sufficient as it largely ignores highly granular temporal information and leads to information loss due to approximation of aggregation granularity. In this work, we propose PIKACHU, a sophisticated, unsupervised, temporal walk-based dynamic network embedding technique that can capture both network topology as well as highly granular temporal information. PIKACHU learns the appropriate and meaningful representation by preserving the temporal order of nodes. This is important information to detect Advanced Persistent Threat (APT) as temporal order helps to understand the lateral movement of the attacker. Experiments on two open-source datasets: LANL and OpTC datasets demonstrated the effectiveness in detecting network anomalies. PIKACHU achieves True Positive Rate (TPR) of 95.1% in LANL and 98.7% on OpTC dataset. Furthermore, in the LANL dataset, it achieves a 4.65% reduction in False Positive Rate (FPR) despite similar area under ROC curve (AUC). In the OpTC dataset 16% improvement in AUC was obtained in comparison to the other state-of-the-art approaches.
Ramesh Paudel, H. Howie Huang
NOMS2
2021 Vestige: Identifying Binary Code Provenance for Vulnerability Detection
Yuede Ji, Lei Cui 0003, H. Howie Huang
ACNS (2)3
2021 BugGraph: Differentiating Source-Binary Code Similarity with Graph Triplet-Loss Network
abstract
Binary code similarity detection, which answers whether two pieces of binary code are similar, has been used in a number of applications,such as vulnerability detection and automatic patching. Existing approaches face two hurdles in their efforts to achieve high accuracy and coverage: (1) the problem of source-binary code similarity detection, where the target code to be analyzed is in the binary format while the comparing code (with ground truth) is in source code format. Meanwhile, the source code is compiled to the comparing binary code with either a random or fixed configuration (e.g.,architecture, compiler family, compiler version, and optimization level), which significantly increases the difficulty of code similarity detection; and (2) the existence of different degrees of code similarity. Less similar code is known to be more, if not equally, important in various applications such as binary vulnerability study. To address these challenges, we design BugGraph, which performs source-binary code similarity detection in two steps. First, BugGraph identifies the compilation provenance of the target binary and compiles the comparing source code to a binary with the same provenance.Second, BugGraph utilizes a new graph triplet-loss network on the attributed control flow graph to produce a similarity ranking. The experiments on four real-world datasets show that BugGraph achieves 90% and 75% true positive rate for syntax equivalent and similar code, respectively, an improvement of 16% and 24% overstate-of-the-art methods. Moreover, BugGraph is able to identify 140 vulnerabilities in six commercial firmware.
Yuede Ji, Lei Cui 0003, H. Howie Huang
AsiaCCS3
2021 Automatic Generation of High-Performance Inference Kernels for Graph Neural Networks on Multi-Core Systems
abstract
Graph neural networks are powerful in learning from high-dimensional graph-structured data, for which a number of frameworks such as DGL and Pytorch-geometrics have been developed to facilitate the construction, training, and deployment of such models. Unfortunately, existing systems underperform when inferring on huge graph data on multi-core CPUs. Furthermore, traditional graph processing systems are struggling with complexity issues due to their low-level programming interfaces. In this paper, we present a new compiler-based software framework Gin optimized for graph neural network inference, which offers a user-friendly interface, via an intuitive programming model, for defining graph neural network models. Gin builds high-level dataflow graphs as intermediate representations, which are transformed into highly efficient codes and then compiled into binary inference kernels. Our evaluation shows that Gin significantly accelerates the inference on billion-edge graphs, beating three state-of-the-art solutions i.e., DGL, Tensorflow, and Pytorch-geometrics, by 31.44×on average, with much higher CPU and memory bandwidth utilization. In addition, Gin is able to achieve considerable speedup (up to 7.6×) over traditional graph processing system Ligra.
H. Howie Huang
ICPP2
2021 Optimizing Job Reliability Through Contention-Free, Distributed Checkpoint Scheduling
abstract
A datacenter that consists of hundreds or thousands of servers can provide virtualized environments to a large number of cloud applications and jobs that value the requirement of reliability very differently. Checkpointing a virtual machine (VM) is a proven technique to improve reliability. However, existing checkpoint scheduling techniques for enhancing reliability of distributed systems fails to achieve satisfactory results, either because they tend to offer the same, fixed reliability to all jobs, or because their solutions are tied up to specific applications and rely on centralized checkpoint control mechanisms. In this work, we first show that reliability can be significantly improved through contention-free scheduling of checkpoints. Then, inspired by the Carrier Sense Multiple Access (CSMA) protocol in wireless congestion control, we propose a novel framework for distributed and contention-free scheduling of VM checkpointing to provide reliability as a transparent, elastic service. We quantify reliability in closed form by studying system stationary behaviours, and maximize job reliability through utility optimization. Our design is validated via a proof-of-concept prototype that leverages readily available implementations in Xen hypervisors. The proposed checkpoint scheduling is shown to significantly reduce checkpointing interference and improve reliability by as much as one order of magnitude over contention-oblivious checkpoint schemes.
Yu Xiang 0003, Hang Liu 0001, Tian Lan 0001, H. Howie Huang, Suresh Subramaniam 0001
IEEE Trans. Netw. Serv. Manag.4
2020 VGRAPH: A Robust Vulnerable Code Clone Detection System Using Code Property Triplets
abstract
Software vulnerabilities are a common attack vector for cyber adversaries. This problem has been exacerbated by the wealth of open-source software projects, as code is often copy-pasted to new locations. This causes a serious problem when a new security vulnerability is discovered in a particular software project, as it may potentially affect many others. Discovering vulnerable code reuse in source code is known as vulnerable code clone detection. This is a very challenging problem as the cloned code has the potential to be modified, sometimes significantly, from the original code, while still retaining the underlying vulnerability. Existing vulnerable clone detection techniques are either too strict, missing vulnerabilities when they have subtle modifications, or are too narrow, applicable only to a small number of vulnerability types. In this work we present VGRAPH, a technique for identifying vulnerable code clones, which is more robust to code modification, while still remaining generic to all vulnerability types. VGRAPHs are representations of vulnerable source code comprising three graph-based components representing code property relationships extracted from the contextual code, the vulnerable code, and the patched code. We develop a matching algorithm utilizing these three graph-based components which is able to identify vulnerable code clones with a precision of 98% and recall of 97%. Even for highly modified code clones, we are able to identify over 100 more vulnerable clones than the best performing comparison work ReDeBug. When we apply our technique to several versions of popular software packages (e.g., FFMpeg, OpenSSL), we are able to identify 10 vulnerabilities which were silently patched and are not listed in the National Vulnerability Database.
Benjamin Bowman, H. Howie Huang
EuroS&P2
2020 Aquila: Adaptive Parallel Computation of Graph Connectivity Queries
abstract
Graph connectivity algorithms answer whether two nodes in a graph are connected under specific conditions, which are beneficial to a number of applications, such as pattern recognition and cybersecurity. Unfortunately, existing graph computing frameworks support only a small number of connectivity algorithms and achieve low computation parallelism. In this paper, we have designed an adaptive parallel computation framework, Aqila, that covers a wide range of different highly optimized graph connectivity algorithms. Given a graph, Aqila first transforms the query if it can be answered with partial computation. During the computation, Aqila is able to greatly reduce the workload by up to 98%. Furthermore, Aqila identifies the irregular tasks in the connectivity algorithms and applies different parallel strategies for different tasks. As a result, Aqila significantly outperforms existing systems such as Multistep, Galois, Ligra, GraphChi, X-Stream, DFS, and Boost, by average 13x, 53x, 264x, 364x, 1,369x, 45x, and 255x, respectively.
Yuede Ji, H. Howie Huang
HPDC2
2020 Detecting Lateral Movement in Enterprise Computer Networks with Unsupervised Graph AI
Benjamin Bowman, Craig Laprade, Yuede Ji, H. Howie Huang
RAID4
2020 GraphOne: A Data Store for Real-time Analytics on Evolving Graphs
abstract
There is a growing need to perform a diverse set of real-time analytics (batch and stream analytics) on evolving graphs to deliver the values of big data to users. The key requirement from such applications is to have a data store to support their diverse data access efficiently, while concurrently ingesting fine-grained updates at a high velocity. Unfortunately, current graph systems, either graph databases or analytics engines, are not designed to achieve high performance for both operations; rather, they excel in one area that keeps a private data store in a specialized way to favor their operations only. To address this challenge, we have designed and developed G raph O ne , a graph data store that abstracts the graph data store away from the specialized systems to solve the fundamental research problems associated with the data store design. It combines two complementary graph storage formats (edge list and adjacency list) and uses dual versioning to decouple graph computations from updates. Importantly, it presents a new data abstraction, GraphView , to enable data access at two different granularities of data ingestions (called data visibility ) for concurrent execution of diverse classes of real-time graph analytics with only a small data duplication. Experimental results show that G raph O ne is able to deliver 11.40× and 5.36× average speedup in ingestion rate against LLAMA and Stinger, the two state-of-the-art dynamic graph systems, respectively. Further, they achieve an average speedup of 8.75× and 4.14× against LLAMA and 12.80× and 3.18× against Stinger for BFS and PageRank analytics (batch version), respectively. G raph O ne also gains over 2,000× speedup against Kickstarter, a state-of-the-art stream analytics engine in ingesting the streaming edges and performing streaming BFS when treating first half as a base snapshot and rest as streaming edge in a synthetic graph. G raph O ne also achieves an ingestion rate of two to three orders of magnitude higher than graph databases. Finally, we demonstrate that it is possible to run concurrent stream analytics from the same data store.
H. Howie Huang
ACM Trans. Storage2
2019 GraphOne: A Data Store for Real-time Analytics on Evolving Graphs
H. Howie Huang
FAST2
2019 CECI: Compact Embedding Cluster Index for Scalable Subgraph Matching
abstract
Subgraph matching finds all distinct isomorphic embeddings of a query graph on a data graph. For large graphs, current solutions face the scalability challenge due to expensive joins, excessive false candidates, and workload imbalance. In this paper, we propose a novel framework for subgraph listing based on Compact Embedding Cluster Index (\idx), which divides the data graph into multiple embedding clusters for parallel processing. The \sub has three unique techniques: utilizing the BFS-based filtering and reverse-BFS-based refinement to prune the unpromising candidates early on, replacing the edge verification with set intersection to speed up the candidate verification, and using search cardinality based cost estimation for detecting and dividing large embedding clusters in advance. The experiments performed on several real and synthetic datasets show that the \sub outperforms state-of-the-art solutions on average by 20.4× for listing all embeddings and by 2.6× for enumerating the first 1,024 embeddings.
Bibek Bhattarai, Hang Liu 0001, H. Howie Huang
SIGMOD Conference3
2019 SIMD-X: Programming and Processing of Graph Algorithms on GPUs
Hang Liu 0001, H. Howie Huang
USENIX ATC2
2018 TriCore: parallel triangle counting on GPUs
Hang Liu 0001, H. Howie Huang
SC3
2018 iSpan: parallel identification of strongly connected components with spanning trees
Yuede Ji, Hang Liu 0001, H. Howie Huang
SC3
2018 A Novel ReRAM-Based Processing-in-Memory Architecture for Graph Traversal
abstract
Graph algorithms such as graph traversal have been gaining ever-increasing importance in the era of big data. However, graph processing on traditional architectures issues many random and irregular memory accesses, leading to a huge number of data movements and the consumption of very large amounts of energy. To minimize the waste of memory bandwidth, we investigate utilizing processing-in-memory (PIM), combined with non-volatile metal-oxide resistive random access memory (ReRAM), to improve both computation and I/O performance. We propose a new ReRAM-based processing-in-memory architecture called RPBFS, in which graph data can be persistently stored and processed in place. We study the problem of graph traversal, and we design an efficient graph traversal algorithm in RPBFS. Benefiting from low data movement overhead and high bank-level parallel computation, RPBFS shows a significant performance improvement compared with both the CPU-based and the GPU-based BFS implementations. On a suite of real-world graphs, our architecture yields a speedup in graph traversal performance of up to 33.8×, and achieves a reduction in energy over conventional systems of up to 142.8×.
Zhaoyan Shen, Duo Liu 0002, Zili Shao, H. Howie Huang, Tao Li 0006
ACM Trans. Storage5
2017 Graphene: Fine-Grained IO Management for Graph Computing
Hang Liu 0001, H. Howie Huang
FAST2
2017 Falcon: Scaling IO Performance in Multi-SSD Volumes
H. Howie Huang
USENIX ATC2
2017 Elastic Reliability Optimization Through Peer-to-Peer Checkpointing in Cloud Computing
abstract
Modern day data centers coordinate hundreds of thousands of heterogeneous tasks and aim at delivering highly reliable cloud computing services. Although offering equal reliability to all users benefits everyone at the same time, users may find such an approach either inadequate or too expensive to fit their individual requirements, which may vary dramatically. In this paper, we propose a novel method for providing elastic reliability optimization in cloud computing. Our scheme makes use of peer-to-peer checkpointing and allows user reliability levels to be jointly optimized based on an assessment of their individual requirements and total available resources in the data center. We show that the joint optimization can be efficiently solved by a distributed algorithm using dual decomposition. The solution improves resource utilization and presents an additional source of revenue to data center operators. Our validation results suggest a significant improvement of reliability over existing schemes.
Juzi Zhao, Yu Xiang 0003, Tian Lan 0001, H. Howie Huang, Suresh Subramaniam 0001
IEEE Trans. Parallel Distributed Syst.4
2017 An Adaptive IO Prefetching Approach for Virtualized Data Centers
abstract
Cloud and data center applications often make heavy use of virtualized servers, where flash-based solid-state drives (SSDs) have become popular alternatives over hard drives for data-intensive applications. Traditional data prefetching focuses on applications running on bare metal systems using hard drives. In contrast, virtualized systems using SSDs present different challenges for data prefetching. Most existing prefetching techniques, if applied unchanged in such environments, are likely to either fail to fully utilize SSDs, interfere with virtual machine I/O requests, or cause too much overhead if run in every virtualized instance.In this work, we demonstrate that data prefetching, when run in a virtualization-friendly manner can provide significant performance benefits for a wide range of data-intensive applications. We have designed and developed VIO-prefetching , consisting of accurate prediction of application needs in runtime and adaptive feedback-directed prefetching that scales with application needs, while being considerate to underlying storage devices and host systems. We have implemented a real system in Linux and evaluated it on different storage devices with the virtualization layer.Our comprehensive study provides insights of VIO-prefetching’s behavior at various virtualization system configurations, e.g., the number of VMs, in-guest processes, application types, etc. The proposed method improves virtual I/O performance up to 43 percent with the average of 14 percent for 1 to 12 VMs while running various applications on a Xen virtualization system.
Ron Chi-Lung Chiang, Ahsen J. Uppal, H. Howie Huang
IEEE Trans. Serv. Comput.3
2016 Performance Analysis of GPU-Based Convolutional Neural Networks
abstract
As one of the most important deep learning models, convolutional neural networks (CNNs) have achieved great successes in a number of applications such as image classification, speech recognition and nature language understanding. Training CNNs on large data sets is computationally expensive, leading to a flurry of research and development of open-source parallel implementations on GPUs. However, few studies have been performed to evaluate the performance characteristics of those implementations. In this paper, we conduct a comprehensive comparison of these implementations over a wide range of parameter configurations, investigate potential performance bottlenecks and point out a number of opportunities for further optimization.
Xiaqing Li, Guangyan Zhang, H. Howie Huang, Zhufan Wang
ICPP3
2016 G-store: high-performance graph store for trillion-edge processing
abstract
High-performance graph processing brings great benefits to a wide range of scientific applications, e.g., biology networks, recommendation systems, and social networks, where such graphs have grown to terabytes of data with billions of vertices and trillions of edges. Subsequently, storage performance plays a critical role in designing a high-performance computer system for graph analytics. In this paper, we present G-Store, a new graph store that incorporates three techniques to accelerate the I/O and computation of graph algorithms. First, G-Store develops a space-efficient tile format for graph data, which takes advantage of the symmetry present in graphs as well as a new smallest number of bits representation. Second, G-Store utilizes tile-based physical grouping on disks so that multi-core CPUs can achieve high cache and memory performance and fully utilize the throughput from an array of solid-state disks. Third, G-Store employs a novel slide-cache-rewind strategy to pipeline graph I/O and computing. With a modest amount of memory, G-Store utilizes a proactive caching strategy in the system so that all fetched graph data are fully utilized before evicted from memory. We evaluate G-Store on a number of graphs against two state-of-the-art graph engines and show that G-Store achieves 2 to 8× saving in storage and outperforms both by 2 to 32×. G-Store is able to run different algorithms on trillion-edge graphs within tens of minutes, setting a new milestone in semi-external graph processing system.
H. Howie Huang
SC2
2016 iBFS: Concurrent Breadth-First Search on GPUs
abstract
Breadth-First Search (BFS) is a key graph algorithm with many important applications. In this work, we focus on a special class of graph traversal algorithm - concurrent BFS - where multiple breadth-first traversals are performed simultaneously on the same graph. We have designed and developed a new approach called iBFS that is able to run i concurrent BFSes from i distinct source vertices, very efficiently on Graphics Processing Units (GPUs). iBFS consists of three novel designs. First, iBFS develops a single GPU kernel for joint traversal of concurrent BFS to take advantage of shared frontiers across different instances. Second, outdegree-based GroupBy rules enables iBFS to selectively run a group of BFS instances which further maximizes the frontier sharing within such a group. Third, iBFS brings additional performance benefit by utilizing highly optimized bitwise operations on GPUs, which allows a single GPU thread to inspect a vertex for concurrent BFS instances. The evaluation on a wide spectrum of graph benchmarks shows that iBFS on one GPU runs up to 30x faster than executing BFS instances sequentially, and on 112 GPUs achieves near linear speedup with the maximum performance of 57,267 billion traversed edges per second (TEPS).
Hang Liu 0001, H. Howie Huang
SIGMOD Conference2
2016 On Soft Error Reliability of Virtualization Infrastructure
abstract
Hardware errors are no longer exceptions in modern cloud data centers. Although virtualization provides software failure isolation among different virtual machines (VM), the virtualization infrastructure including the hypervisor and privileged VMs remains vulnerable to hardware errors. What makes matters worse is that such errors are unlikely bounded by the virtualization boundary and may lead to loss of work in multiple guest VMs due to unexpected and/or mishandled failures. To understand reliability implication of hardware errors in virtualized systems, in this paper we develop a simulation-based framework that enables a comprehensive fault injection study on the hypervisor with a wide range of configurations. Our analysis shows that, in current systems, many hardware errors can propagate through various paths for an extended time before an observed failure (e.g., whole system crash). We further discuss the challenges of designing error tolerance techniques for the hypervisor.
H. Howie Huang
IEEE Trans. Computers2
2015 DualVisor: Redundant Hypervisor Execution for Achieving Hardware Error Resilience in Datacenters
abstract
Virtualization technology as the foundation of cloud computing provides many benefits in cost, security, and management, but all of them rely on the reliability of the underlying virtualization software - the hypervisor (or virtual machine monitor). Cloud data centers are built upon 10Ks to 100Ks commodity servers. Hardware errors in these large scale computer systems are not rare events. When hardware errors occur during the hypervisor execution, they may cause failures or data corruptions in co-located VMs, undermining the whole system reliability. In this paper, we propose DualVisor, that uses a software redundancy based fault tolerance technique to protect the hypervisor from hardware errors. DualVisor replicates hypervisor executions and data structures for error detection and recovery. In this work, we first study the need for a hardware error-resilient hypervisor. Then, we discuss the design considerations in detail. We implement a prototype in the hypervisor to demonstrate the feasibility and evaluate the performance overhead. Our preliminary results show that the performance overhead of DualVisor is fairly small (less than 6%) for tested applications.
H. Howie Huang
CCGRID2
2015 IOrchestra: supporting high-performance data-intensive applications in the cloud via collaborative virtualization
abstract
Multi-tier data-intensive applications are widely deployed in virtualized data centers for high scalability and reliability. As the response time is vital for user satisfaction, this requires achieving good performance at each tier of the applications in order to minimize the overall latency. However, in such virtualized environments, each tier (e.g., application, database, web) is likely to be hosted by different virtual machines (VMs) on multiple physical servers, where a guest VM is unaware of changes outside its domain, and the hypervisor also does not know the configuration and runtime status of a guest VM. As a result, isolated virtualization domains lend themselves to performance unpredictability and variance. In this paper, we propose IOrchestra, a holistic collaborative virtualization framework, which bridges the semantic gaps of I/O stacks and system information across multiple VMs, improves virtual I/O performance through collaboration from guest domains, and increases resource utilization in data centers. We present several case studies to demonstrate that IOrchestra is able to address numerous drawbacks of the current practice and improve the I/O latency of various distributed cloud applications by up to 31%.
Ron Chi-Lung Chiang, H. Howie Huang, Timothy Wood 0001, Changbin Liu, Oliver Spatscheck
SC2
2015 Enterprise: breadth-first graph traversal on GPUs
abstract
The Breadth-First Search (BFS) algorithm serves as the foundation for many graph-processing applications and analytics workloads. While Graphics Processing Unit (GPU) offers massive parallelism, achieving high-performance BFS on GPUs entails efficient scheduling of a large number of GPU threads and effective utilization of GPU memory hierarchy. In this paper, we present Enterprise, a new GPU-based BFS system that combines three techniques to remove potential performance bottlenecks: (1) streamlined GPU threads scheduling through constructing a frontier queue without contention from concurrent threads, yet containing no duplicated frontiers and optimized for both top-down and bottom-up BFS. (2) GPU workload balancing that classifies the frontiers based on different out-degrees to utilize the full spectrum of GPU parallel granularity, which significantly increases thread-level parallelism; and (3) GPU based BFS direction optimization quantifies the effect of hub vertices on direction-switching and selectively caches a small set of critical hub vertices in the limited GPU shared memory to reduce expensive random data accesses. We have evaluated Enterprise on a large variety of graphs with different GPU devices. Enterprise achieves up to 76 billion traversed edges per second (TEPS) on a single NVIDIA Kepler K40, and up to 122 billion TEPS on two GPUs that ranks No. 45 in the Graph 500 on November 2014. Enterprise is also very energy-efficient as No. 1 in the GreenGraph 500 (small data category), delivering 446 million TEPS per watt.
Hang Liu 0001, H. Howie Huang
SC2
2015 Swiper: Exploiting Virtual Machine Vulnerability in Third-Party Clouds with Competition for I/O Resources
abstract
The emerging paradigm of cloud computing, e.g., Amazon Elastic Compute Cloud (EC2), promises a highly flexible yet robust environment for large-scale applications. Ideally, while multiple virtual machines (VM) share the same physical resources (e.g., CPUs, caches, DRAM, and I/O devices), each application should be allocated to an independently managed VM and isolated from one another. Unfortunately, the absence of physical isolation inevitably opens doors to a number of security threats. In this paper, we demonstrate in EC2 a new type of security vulnerability caused by competition between virtual I/O workloads-i.e., by leveraging the competition for shared resources, an adversary could intentionally slow down the execution of a targeted application in a VM that shares the same hardware. In particular, we focus on I/O resources such as hard-drive throughput and/or network bandwidth-which are critical for data-intensive applications. We design and implement Swiper, a framework which uses a carefully designed workload to incur significant delays on the targeted application and VM with minimum cost (i.e., resource consumption). We conduct a comprehensive set of experiments in EC2, which clearly demonstrates that Swiper is capable of significantly slowing down various server applications while consuming a small amount of resources.
Ron Chi-Lung Chiang, Sundaresan Rajasekaran, Nan Zhang 0004, H. Howie Huang
IEEE Trans. Parallel Distributed Syst.4
2015 Exploring Data-Level Error Tolerance in High-Performance Solid-State Drives
abstract
Flash storage systems have exhibited great benefits over magnetic hard drives such as low input-output (I-O) latency, and high throughput. However, NAND flash based Solid-State Drives (SSDs) are inherently prone to soft errors from various sources, e.g., wear-out, program and read disturbance, and hot-electron injections. To address this issue, flash devices employ different error-correction codes (ECC) to detect and correct soft errors. Using ECC induces non-trivial overhead costs in terms of flash area, performance, and energy consumption. In this work, we evaluate the feasibility of reducing the need for strong ECC while maintaining the correct execution of the applications. Specifically, we explore data-level error tolerance in various data-centric applications, and study the system implications for designing a low-cost yet high performance flash storage system, SoftFlash. We explore three key aspects of enabling SoftFlash. First, we design an error modeling framework that can be used in runtime for monitoring and estimating the error rates of real-world flash devices. Our experiments show that the error rate of SSDs can be modeled with reasonable accuracy (13%) using parameters accessible from operating systems. Second, we carry out extensive fault-injection experiments on a wide range of applications including multimedia, scientific computation, and cloud computing to understand the requirements and characteristics of data level error tolerance. We find that the data from these applications show high error resiliency, and can produce acceptable results even with high error rates. Third, we conduct a case study to show the benefits of leveraging data-level error tolerance in flash devices. Our results show that, for many data-centric applications, the proposed SoftFlash system can achieve acceptable results (or better in certain cases), with more than a 40% performance improvement, and a third of the energy consumption.
H. Howie Huang
IEEE Trans. Reliab.2
2014 UniCache: Hypervisor Managed Data Storage in RAM and Flash
abstract
Application and OS-level caches are crucial for hiding I/O latency and improving application performance. However, caches are designed to greedily consume memory, which can cause memory-hogging problems in a virtualized data centers since the hypervisor cannot tell for what a virtual machine uses its memory. A group of virtual machines may contain a wide range of caches: database query pools, memcached key-value stores, disk caches, etc., each of which would like as much memory as possible. The relative importance of these caches can vary significantly, yet system administrators currently have no easy way to dynamically manage the resources assigned to a range of virtual machine data caches in a unified way. To improve this situation, we have developed UniCache, a system that provides a hypervisor managed volatile data store that can cache data either in hypervisor controlled main memory (hot data) or on Flash based storage (cold data). We propose a two-level cache management system that uses a combination of recency information, object size, and a prediction of the cost to recover an object to guide its eviction algorithm. We have built a prototype of UniCache using Xen, and have evaluated its effectiveness in a shared environment where multiple virtual machines compete for storage resources.
Jinho Hwang, Wei Zhang 0052, Ron Chi-Lung Chiang, Timothy Wood 0001, H. Howie Huang
IEEE CLOUD5
2014 Big data machine learning and graph analytics: Current state and future challenges
abstract
Big data machine learning and graph analytics have been widely used in industry, academia and government. Continuous advance in this area is critical to business success, scientific discovery, as well as cybersecurity. In this paper, we present some current projects and propose that next-generation computing systems for big data machine learning and graph analytics need innovative designs in both hardware and software that provide a good match between big data algorithms and the underlying computing and storage resources.
H. Howie Huang, Hang Liu 0001
IEEE BigData1
2014 Xentry: Hypervisor-Level Soft Error Detection
abstract
Cloud data centers leverage virtualization to share commodity hardware resources, where virtual machines (VMs) achieve fault isolation by containing VM failures within the virtualization boundary. However, hypervisor failure induced by soft errors will most likely affect multiple, if not all, VMs on a single physical host. Existing fault detection techniques are not well equipped to handle such hypervisor failures. In this paper, we propose a new soft error detection framework, Xentry (a sentry on soft error for Xen), that focuses on limiting error propagation within and from the hypervisor. In particular, we have designed a VM transition detection technique to identify incorrect control flow before VM execution resumes, and a runtime detection technique to shorten detection latency. This framework requires no hardware modification and has been implemented in the Xen hypervisor. The experiment results show that Xentry incurs very small performance overhead and detects over 99% of the injected faults.
Ron Chi-Lung Chiang, H. Howie Huang
ICPP3
2014 Mortar: filling the gaps in data center memory
abstract
Data center servers are typically overprovisioned, leaving spare memory and CPU capacity idle to handle unpredictable workload bursts by the virtual machines running on them. While this allows for fast hotspot mitigation, it is also wasteful. Unfortunately, making use of spare capacity without impacting active applications is particularly difficult for memory since it typically must be allocated in coarse chunks over long timescales. In this work we propose re- purposing the poorly utilized memory in a data center to store a volatile data store that is managed by the hypervisor. We present two uses for our Mortar framework: as a cache for prefetching disk blocks, and as an application-level distributed cache that follows the memcached protocol. Both prototypes use the framework to ask the hypervisor to store useful, but recoverable data within its free memory pool. This allows the hypervisor to control eviction policies and prioritize access to the cache. We demonstrate the benefits of our prototypes using realistic web applications and disk benchmarks, as well as memory traces gathered from live servers in our university's IT department. By expanding and contracting the data store size based on the free memory available, Mortar improves average response time of a web application by up to 35% compared to a fixed size memcached deployment, and improves overall video streaming performance by 45% through prefetching.
Jinho Hwang, Ahsen J. Uppal, Timothy Wood 0001, H. Howie Huang
VEE4
2014 Exploring Dynamic Redundancy to Resuscitate Faulty PCM Blocks
abstract
DRAM technology challenges have increased the necessity to adapt to the emerging memory technologies like Phase-Change Memory (PCM or PRAM). While such emerging technologies provide benefits like storage density, nonvolatility, and low energy consumption, they are constrained by limited write endurance that becomes more pronounced with process variation. In this article, we explore a novel PRAM-based main memory system which resuscitates a group of faulty pages in a cost-effective manner to significantly extend the PCM main memory lifetime while minimizing the performance impact. In particular, we explore three different dimensions of dynamic redundancy levels and group sizes, and design low-cost hardware and software support for our proposed schemes. We aim to have minimal hardware modifications (that have less than 1% on-chip and off-chip area overheads). Also, our schemes can improve the PRAM lifetime by up to 105× (times) over a chip with no error correction capabilities, and outperform prior schemes such as DRM and ECP at a small fraction of the hardware cost. The performance overhead resulting from our scheme is less than 8% on average across 21 applications from SPEC2006, Splash-2, and PARSEC benchmark suites.
Jie Chen 0020, Guru Venkataramani, H. Howie Huang
ACM J. Emerg. Technol. Comput. Syst.3
2014 TRACON: Interference-Aware Schedulingfor Data-Intensive Applicationsin Virtualized Environments
abstract
Large-scale data centers leverage virtualization technology to achieve excellent resource utilization, scalability, and high availability. Ideally, the performance of an application running inside a virtual machine (VM) shall be independent of co-located applications and VMs that share the physical machine. However, adverse interference effects exist and are especially severe for data-intensive applications in such virtualized environments. In this work, we present TRACON, a novel Task and Resource Allocation CONtrol framework that mitigates the interference effects from concurrent data-intensive applications and greatly improves the overall system performance. TRACON utilizes modeling and control techniques from statistical machine learning and consists of three major components: the interference prediction model that infers application performance from resource consumption observed from different VMs, the interference-aware scheduler that is designed to utilize the model for effective resource management, and the task and resource monitor that collects application characteristics at the runtime for model adaption. We implement and validate TRACON with a variety of cloud applications. The evaluation results show that TRACON can achieve up to 25 percent improvement on application throughput on virtualized servers.
Ron Chi-Lung Chiang, H. Howie Huang
IEEE Trans. Parallel Distributed Syst.2
2013 DUAL: Reliability-Aware Power Management in Data Centers
abstract
A virtualized data center hosts users and applications within a large number of virtual machines (VM) to achieve easy provisioning and high utilization of physical resources. Energy efficiency and reliability are two primary concerns for operating a data center. Power saving techniques, such as dynamic voltage and frequency scaling (DVFS), are often employed to reduce the supply voltages of the CPUs in runtime when the computer system utilization is low. However, DVFS can potentially decrease the system reliability - the processors at low voltages are more likely to encounter soft errors that may result in VM or system crashes. In this work, we propose a data center management framework, DUAL, which consists of the new virtual machine power and reliability analysis tools. The framework is designed to balance the dual needs of a data center: reducing energy consumption and providing high reliability. The evaluations show that DUAL can help maintain the desired reliability and significantly reduce power consumption, which in turn will lower the overall operational cost of a data center.
Kayo Teramoto, Allan Morales, H. Howie Huang
CCGRID4
2013 Mortar: filling the gaps in data center memory
abstract
Data center servers are typically overprovisioned, leaving spare memory and CPU capacity idle to handle unpredictable workload bursts by the virtual machines running on them [1, 2, 3]. While this allows for fast hotspot mitigation, it is also wasteful. Unfortunately, making use of spare capacity without impacting active applications is particularly difficult for memory since it typically must be allocated in coarse chunks over long timescales [4, 5, 6, 7]. In this work we propose repurposing the poorly utilized memory in a data center to store a volatile data store that is managed by the hypervisor. We present two uses for our Mortar framework: as a cache for prefetching disk blocks [8, 9, 10], and as an application-level distributed cache that follows the memcached protocol [11, 12]. Both prototypes use the framework to ask the hypervisor to store useful, but recoverable data within its free memory pool. This allows the hypervisor to control eviction policies and prioritize access to the cache.
Jinho Hwang, Ahsen J. Uppal, Timothy Wood 0001, H. Howie Huang
SoCC4
2013 GPU-accelerated scalable solver for banded linear systems
abstract
Solving a banded linear system efficiently is important to many scientific and engineering applications. Current solvers achieve good scalability only on the linear systems that can be partitioned into independent subsystems. In this paper, we present a GPU based, scalable Bi-Conjugate Gradient Stabilized solver that can be used to solve a wide range of banded linear systems. We utilize a row-oriented matrix decomposition method to divide the banded linear system into several correlated sub-linear systems and solve them on multiple GPUs collaboratively. We design a number of GPU and MPI optimizations to speedup inter-GPU and inter-machine communications. We evaluate the solver on Poisson equation and advection diffusion equation as well as several other banded linear systems. The solver achieves a speedup of more than 21 times running from 6 to 192 GPUs on the XSEDE's Keeneland supercomputer and because of small communication overhead, can scale upto 32 GPUs on Amazon EC2 with relatively slow ethernet network.
Hang Liu 0001, Jung Hee Seo, Rajat Mittal 0002, H. Howie Huang
CLUSTER4
2013 Achieving high job execution reliability using underutilized resources in a computational economy
Woochul Kang, H. Howie Huang, Andrew S. Grimshaw
Future Gener. Comput. Syst.2
2012 Understanding the effects of hypervisor I/O scheduling for virtual machine performance interference
abstract
In virtualized environments, the customers who purchase virtual machines (VMs) from a third-party cloud would expect that their VMs run in an isolated manner. However, the performance of a VM can be negatively affected by co-resident VMs. In this paper, we propose vExplorer, a distributed VM I/O performance measurement and analysis framework, where one can use a set of representative I/O operations to identify the I/O scheduling characteristics within a hypervisor; and potentially leverage this knowledge to carry out I/O based performance attacks to slow down the execution of the target VMs. We evaluate our prototype on both Xen and VMware platforms with four server benchmarks and show that vExplorer is practical and effective. We also conduct similar tests on Amazon's EC2 platform and successfully slow down the performance of target VMs.
Ziye Yang, Haifeng Fang, Yingjun Wu, Chunqi Li, H. Howie Huang
CloudCom6
2012 RePRAM: Re-cycling PRAM faulty blocks for extended lifetime
abstract
As main memory systems begin to face the scaling challenges from DRAM technology, future computer systems need to adapt to the emerging memory technologies like Phase-Change Memory (PCM or PRAM). While these newer technologies offer advantages such as storage density, non-volatility, and low energy consumption, they are constrained by limited write endurance that becomes more pronounced with process variation. In this paper, we propose a novel PRAM-based main memory system, RePRAM (Recycling PRAM), which leverages a group of faulty pages and recycles them in a managed way to significantly extend the PRAM lifetime while minimizing the performance impact. In particular, we explore two different dimensions of dynamic redundancy levels and group sizes, and design low-cost hardware and software support for RePRAM. Our proposed scheme involves minimal hardware modifications (that have less than 1% on-chip and off-chip area overheads). Also, our schemes can improve the PRAM lifetime by up to 43× (times) over a chip with no error correction capabilities, and outperform prior schemes such as DRM and ECP at a small fraction of the hardware cost. The performance overhead resulting from our scheme is less than 7% on average across 21 applications from SPEC2006, Splash-2, and PARSEC benchmark suites.
Jie Chen 0020, Guru Venkataramani, H. Howie Huang
DSN3
2012 Providing reliability as an elastic service in cloud computing
abstract
Modern day data centers coordinate hundreds of thousands of heterogeneous tasks and aim at delivering highly reliable cloud computing services. Although offering equal reliability to all users benefits everyone at the same time, users may find such an approach either too inadequate or too expensive to fit their individual requirements, which may vary dramatically. In this paper, we propose a novel method for providing reliability as an elastic and on-demand service. Our scheme makes use of peer-to-peer checkpointing and allows user reliability levels to be jointly optimized based on an assessment of their individual requirements and total available resources in the data center. We show that the joint optimization can be efficiently solved by a distributed algorithm using dual decomposition. The solution improves resource utilization and presents an additional source of revenue to data center operators. Our validation results suggest a significant improvement of reliability over existing schemes.
Nakharin Limrungsi, Juzi Zhao, Yu Xiang 0003, Tian Lan 0001, H. Howie Huang, Suresh Subramaniam 0001
ICC5
2012 Flashy prefetching for high-performance flash drives
abstract
While hard drives hold on to the capacity advantage, flash-based solid-state drives (SSD) with high bandwidth and low latency have become good alternatives for I/O-intensive applications. Traditional data prefetching has been primarily designed to improve I/O performance on hard drives. The same techniques, if applied unchanged on flash drives, are likely to either fail to fully utilize SSDs, or interfere with application I/O requests, both of which could result in undesirable application performance. In this work, we demonstrate that data prefetching, when effectively harnessing the high performance of SSDs, can provide significant performance benefits for a wide range of data-intensive applications. The new technique, flashy prefetching, consists of accurate prediction of application needs in runtime and adaptive feedback-directed prefetching that scales with application needs, while being considerate to underlying storage devices. We have implemented a real system in Linux and evaluated it on four different SSDs. The results show 65-70% prefetching accuracy and an average 20% speedup on LFS, web search engine traces, BLAST, and TPC-H like benchmarks across various storage drives.
Ahsen J. Uppal, Ron Chi-Lung Chiang, H. Howie Huang
MSST3
2012 Just-in-Time Analytics on Large File Systems
abstract
As file systems reach the petabytes scale, users and administrators are increasingly interested in acquiring high-level analytical information for file management and analysis. Two particularly important tasks are the processing of aggregate and top-k queries which, unfortunately, cannot be quickly answered by hierarchical file systems such as ext3 and NTFS. Existing preprocessing-based solutions, e.g., file system crawling and index building, consume a significant amount of time and space (for generating and maintaining the indexes) which in many cases cannot be justified by the infrequent usage of such solutions. In this paper, we advocate that user interests can often be sufficiently satisfied by approximate-i.e., statistically accurate-answers. We develop Glance, a just-in-time sampling-based system which, after consuming a small number of disk accesses, is capable of producing extremely accurate answers for a broad class of aggregate and top-k queries over a file system without the requirement of any prior knowledge. We use a number of real-world file systems to demonstrate the efficiency, accuracy, and scalability of Glance.
H. Howie Huang, Nan Zhang 0004, Wei Wang 0082, Gautam Das 0001, Alex Szalay
IEEE Trans. Computers1
2011 rPRAM: Exploring Redundancy Techniques to Improve Lifetime of PCM-based Main Memory
abstract
Future main memory systems will confront the scaling challenges posed by DRAM technology and should adapt themselves to use the emerging memory technologies like Phase Change Memory (PCM, or PRAM). PCM offers advantages such as storage density, non-volatility, and lower energy consumption. However, they are constrained by limited write endurance and reduced performance. In this paper, we propose a novel PCM-based main memory system, rPRAM, that explores advanced redundancy techniques to resuscitate faulty PCM pages and reuse these pages to store data. Our preliminary experiments show that rPRAM has the potential to extend the lifetime of PCM based memory commensurate with the existing schemes like ECP, while incurring only a negligible fraction of hardware cost compared to ECP.
Jie Chen 0020, Zachary Winter, Guru Venkataramani, H. Howie Huang
PACT4
2011 Just-in-Time Analytics on Large File Systems
H. Howie Huang, Nan Zhang 0004, Wei Wang 0082, Gautam Das 0001, Alex Szalay
FAST1
2011 Performance modeling and analysis of flash-based storage devices
abstract
Flash-based solid-state drives (SSDs) will become key components in future storage systems. An accurate performance model will not only help understand the state-of-the-art of SSDs, but also provide the research tools for exploring the design space of such storage systems. Although over the years many performance models were developed for hard drives, the architectural differences between two device families prevent these models from being effective for SSDs. The hard drive performance models cannot account for several unique characteristics of SSDs, e.g., low latency, slow update, and expensive block-level erase. In this paper, we utilize the black-box modeling approach to analyze and evaluate SSD performance, including latency, bandwidth, and throughput, as it requires minimal a priori information about the storage devices. We construct the black-box models, using both synthetic workloads and real-world traces, on three SSDs, as well as an SSD RAID. We find that, while the black-box approach may produce less desirable performance predictions for hard disks, a black-box SSD model with a comprehensive set of workload characteristics can produce accurate predictions for latency, bandwidth, and throughput with small errors.
H. Howie Huang, Alex Szalay, Andreas Terzis
MSST1
2011 TRACON: interference-aware scheduling for data-intensive applications in virtualized environments
abstract
Large-scale data centers leverage virtualization technology to achieve excellent resource utilization, scalability, and high availability. Ideally, the performance of an application running inside a virtual machine (VM) shall be independent of co-located applications and VMs that share the physical machine. However, adverse interference effects exist and are especially severe for data-intensive applications in such virtualized environments. In this work, we present TRACON, a novel Task and Resource Allocation CONtrol framework that mitigates the interference effects from concurrent dataintensive applications and greatly improves the overall system performance. TRACON utilizes modeling and control techniques from statistical machine learning and consists of three major components: the interference prediction model that infers application performance from resource consumption observed from different VMs, the interference-aware scheduler that is designed to utilize the model for effective resource management, and the task and resource monitor that collects application characteristics at the runtime for model adaption. We simulate TRACON with a wide variety of data-intensive applications including bioinformatics, data mining, video processing, email and web servers, etc. The evaluation results show that TRACON can achieve up to 50% improvement on application runtime, and up to 80% on I/O throughput for data-intensive applications in virtualized data centers.
Ron Chi-Lung Chiang, H. Howie Huang
SC2
2011 Design, implementation and evaluation of a virtual storage system
abstract
Abstract Large organizations always have a strong demand for storage from data‐intensive applications and instruments. In this paper, we present the design, implementation, and evaluation of a new virtual storage system, Storage@desk, which can aggregate a large number of distributed machines within an organization to provide storage services with quality of service guarantees. Because storage virtualization is the prominent goal, Storage@desk provides clients with the abstraction of a hard drive by utilizing the Internet SCSI protocol. As such, data access to new storage services is transparent so that clients do not need to modify any existing applications nor change their current practices. Storage@desk replicates data and employs version‐based journaling for high availability. It utilizes a market‐based model for resource management and a feedback controller for automated performance control. We have developed a prototype of Storage@desk that implements the core components. Copyright © 2010 John Wiley & Sons, Ltd.
H. Howie Huang, Andrew S. Grimshaw
Concurr. Comput. Pract. Exp.1
2010 VMGuard: An Integrity Monitoring System for Management Virtual Machines
abstract
A cloud computing provider can dynamically allocate virtual machines (VM) based on the needs of the customers, while maintaining the privileged access to the Management Virtual Machine that directly manages the hardware and supports the guest VMs. The customers must trust the cloud providers to protect the confidentiality and integrity of their applications and data. However, as the VMs from different customers are running on the same host, an attack to the management virtual machine will easily lead to the compromise of the guest VMs. Therefore, it is critical for a cloud computing system to ensure the trustworthiness of management VMs. To this end, we propose VMGuard, an integrity monitoring and detecting system for management virtual machines in a distributed environment. VMGuard utilizes a special VM, Guard Domain, which runs on each physical node to monitor the co-resident management VMs. The integrity measurements collected by the Guard Domains are sent to the VMGuard server for safe store and independent analysis. The experimental evaluation of a Xen-based prototype shows that VMGuard can quickly detect the root kit attacks while the performance overhead is low.
Haifeng Fang, Yiqiang Zhao, Hongyong Zang, H. Howie Huang, Yuzhong Sun, Zhiyong Liu 0002
ICPADS4
2010 PicFS: The Privacy-Enhancing Image-Based Collaborative File System
abstract
Cloud computing makes available a vast amount of computation and storage resources in the pay-as-you-go manner. However, the users of cloud storage have to trust the providers to ensure the data privacy and confidentiality. In this paper, we present the Privacy-enhancing Image-based Collaborative File System (PicFS), a network file system that steganographically encodes itself into images and provides anonymous uploads and downloads from a media sharing website. PicFS provides plausible deniability by preventing traffic and image analysis by any third party from revealing the existence of PicFS or compromising its data. Because all accesses are anonymized, users of PicFS are dissociated from their data, which protects users against being compelled to release their keys. For further security and ease of use, we develop a method for automatically generating a large set of non-suspicious images to serve as input to the system. Our prototype leverages a number of existing technologies, including the F5 algorithm for steganography, Quick-Flickr for Flickr API access, Tor for anonymization, and FUSE-J for user-level file system calls. We show that the PicFS is indeed practical as the prototype demonstrates satisfactory performance in the real-world environment.
Chris Sosa, Blake C. Sutton, H. Howie Huang
ICPADS3
2010 Black-Box Performance Modeling for Solid-State Drives
abstract
Flash-based Solid-State Drives (SSDs) have become a promising alternative to magnetic Hard Disk Drives (HDDs) thanks to the large improvements in performance, power consumption, and shock resistance. An accurate SSD performance model will provide the important research tools for exploring the design space of flash-based storage systems. While many HDD performance models have been developed, architectural differences prevent these models from being effective for SSDs, mostly because their designs cannot accurately account for many unique SSD characteristics (e.g., low latencies, slow updates, and expensive erases). In this paper, we utilize the black-box modeling technique to analyze and evaluate SSD performance, including latency, bandwidth, and throughput. Such an approach is appealing because it requires minimal a priori information about SSDs. We construct and evaluate our models on three commercial SSDs. Although this approach may lead to less accurate predictions for HDDs, we find that a black-box model with a comprehensive set of workload characteristics can achieve the mean relative errors of 20%, 13%, and 6% for latency, bandwidth, and throughput predictions, respectively.
H. Howie Huang
MASCOTS2
2010 A control-theoretic approach to automated local policy enforcement in computational grids
H. Howie Huang
Future Gener. Comput. Syst.1
2008 Analyzing the feasibility of building a new mass storage system on distributed resources
abstract
Abstract The average PC now contains a large and increasing amount of storage with an ever greater amount left unused. We believe there is an opportunity for organizations to harness the vast unused storage capacity on their PCs to create a very large, low‐cost, shared storage system. What is needed is the proper storage system architecture and software to exploit and manage the unused portions of existing PC storage devices across an organization and make it reliably accessible to users and applications. We call our vision of such a storage system Storage@desk (SD). This paper describes our first step towards the realization of SD—a study of machine and storage characteristics and usage in a model organization. We studied 729 PCs in an academic institution for 91 days, monitoring the configuration, load and usage of the major machine subsystems, i.e. disk, memory, CPU and network. To further analyze the availability characteristics of storage in an SD system, we performed a trace‐driven simulation of some basic storage allocation strategies. This paper presents the results of our data collection efforts, our analysis of the data, our simulation results and our conclusion that an SD system is indeed feasible and holds promise as a cost‐effective way to create massive storage systems. Copyright © 2007 John Wiley & Sons, Ltd.
H. Howie Huang, John F. Karpovich, Andrew S. Grimshaw
Concurr. Comput. Pract. Exp.1
2007 You Can't Always Get What You Want: Achieving Differentiated Service Levels with Pricing Agents in a Storage Grid
abstract
We have designed a new storage grid called Storage@desk to harness unused storage available on desktop machines and turn it into a useful resource for clients. Given the complexity of managing clientspecific QoS requirements, and the dynamism inherent in supply and demand for resources, even a highly experienced system administrator cannot effectively manage resource allocation. In this paper, we present a market-based resource allocation model where pricing agents help resource providers adjust the prices as demand fluctuates. With derivative-following pricing, an agent requires no knowledge of competitors or consumers, which reduces communication overheads and avoids bottlenecks in the system. Individual clients need a variety of service levels and are in competition in scarce resources. Under the budget constraints, the consumers can't always get what they want. The budgets serve as an incentive for the consumers to react to the price signals. We simulate our model using real world trace data and the results show that, using this model, the system allows the consumers to achieve QoS goals under sufficient budgets and degrade in accordance with relative budget amounts.
H. Howie Huang, Andrew S. Grimshaw, John F. Karpovich
Web Intelligence1
2006 The Cost of Transparency: Grid-Based File Access on the Avaki Data Grid
H. Howie Huang, Andrew S. Grimshaw
ISPA1