VLDB 2026 Research / reviewers in the wild / expert
Ming Zhao 0002
dblp:z/MingZhao2
· DBLP profile ↗
55ranked-venue papers
9as first author
21since 2021 · last 2025
0000-0001-9531-4464ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 8 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Smart Blue Light Pole-Based Real-Time Crowd Counting for Smart CampusesabstractCrowd counting on a smart campus can provide valuable insights into pedestrian behaviors and is critical to making decisions about campus designs and improvements. Conventional image or video based crowd-counting approaches have serious privacy risks, while cloud-based solutions also face challenges in providing real-time responses. This paper proposes Height-Aware Human Classifier for Crowd Counting (HAWC-CC), a real-time, cost-effective, and privacy-protected smart blue light pole-based crowd-counting framework for smart campuses, using Light Detection and Range (LiDAR) sensors and computing devices installed on the blue light poles to collect traffic data and analyze patterns in real time while protecting data privacy. HAWC-CC is built upon two novel methods. The first is an adaptive clustering approach that dynamically adjusts to point cloud structure and density for accurate and efficient clustering. The second method is a Height-Aware Human Classifier (HAWC) that projects 3D point clouds into height-augmented multiple 2D views and uses a lightweight CNN to detect humans from the projected views. HAWC-CC and its quantized version are then implemented on Nvidia Jetson and Coral TPU for real-world deployment on the blue light poles. A comprehensive 3D LiDAR dataset is collected and curated for single-person detection and crowd counting. An extensive evaluation shows that HAWC-CC significantly outperforms the representative related works (PointNet-Cc, AutoEncoder-Cc, and OC-SVM-CC) by up to 86.61% in Mean Absolute Error (MAE) and 90.44% in Mean Squared Error (MSE). Additionally, HAWC-CC is the only framework that meets the real-time requirements for LiDAR-based crowd counting, processing one LiDAR sample in 17.42 ms. HAWC demonstrates the highest robustness to limited training data, achieving a notable accuracy of 90.29% with only 0.1% of the training data. HAWC-CC outperforms SOTA RGB image-based solutions in high-density settings, achieving 97.64% accuracy. Kaiqi Zhao 0002, Krishna Gundu, Zohair Zaidi, Ming Zhao 0002 |
ICDCS | 5 |
| 2025 | Waterfall: Fast Network Flow Rules Checking and Conflict ResolutionabstractSoftware Defined Networking (SDN) enables a centralized manageable framework to control network devices and their policies using device-specific flow rules. When administrators deploy flow rules to support business policies, the network controller checks them against existing rules to detect conflicts and ensure consistency, security, and functionality in the data plane. Existing offline conflict detection methods are not scalable due to state explosion and often lead to networking chaos due to inefficiency. This paper presents Waterfall, designed to minimize the number of flow rule-checking operations. We propose a novel Equivalence Class (EC) creation and prioritization technique that simplifies conflict detection by organizing rules with similar patterns and processing them accordingly. Analogous to a multi-stage waterfall, our algorithm optimizes downstream stages by reducing unnecessary comparisons, ensuring efficient conflict detection. Our comprehensive evaluation demonstrates Waterfall’s effectiveness through significant reductions in computation time ($O(mKH)$, where m is the number of matched flow-rules which is far less than the total number of flow-rules, K is the number of attributes (headers) in flow rules, H is the number of hash functions in Bloom filter for attribute matching), making it ideal for real-time flow rule checking and conflict resolution in SDN environments. In our evaluation, Waterfall achieved a remarkable 1.3X improvement in conflict detection and 4.4X improvement for conflict resolution over the state-of-the-art solution for the Stanford topology which is a popular topology to represent real-world networking scenarios. We also evaluate the scalability of the solution using a synthetic dataset containing 15K flow rules that have three virtual network functions. Our solution achieved a$90.53~\mu $s conflict detection and resolution time for the large synthetic dataset. This lightweight approach promises substantial benefits for real-time flow rule checking in SDN environments. Neha Vadnere, Dijiang Huang, Abdulhakim Sabur, Jim Luo, Ming Zhao 0002 |
IEEE Trans. Netw. | 5 |
| 2025 | The Past, Present, and Future of Storage Technologies (Part 1 of 2)
Geoffrey H. Kuenning, Youjip Won, Ming Zhao 0002, Erez Zadok |
ACM Trans. Storage | 3 |
| 2025 | The Past, Present, and Future of Storage Technologies (part 2 of 2)abstractThe Past, Present, and Future of Storage Technologies (part 2 of 2)Any good research project begins with a "literature review"-a process of looking for papers relevant to a topic of interest, reviewing them to identify those more useful while discarding the rest, then looking for more papers, and repeating this process over and over until you feel you have reached some saturation point.That is the point when you're coming up against the same papers you have seen already (which we like to call the "transitive closure" point).This review stage is fairly time-consuming but also critical; in fact, many research papers get rejected for neglecting to cite important related work.Similarly, practitioners may run into roadblocks that they could have avoided if they had been aware of all the existing literature.So, are you a new graduate student or an employee at a company who is interested in innovating in a given storage technology?Do you wish you could find a single publication that would summarize (almost) everything there is to know about a specific technology?If so, this Special Issue (both parts) is hopefully for you because, unlike conference papers, our journal articles have no page limits and thus allow authors to discuss any technology with as much detail as needed and include a comprehensive bibliography for those interested in more.Incidentally, we found out that survey papers tend to get well cited and received.We believe this is because they become a "go to" source on a topic, thus saving the readers from having to read and cite many other papers.Authors of survey papers may see their articles cited over and over.In 2023, TOS's Editor-in-Chief ( EiC ) reached out to a few senior people to brainstorm ideas for special issues.Because putting together a special issue is a huge task, he recruited several Associate Editors (AEs) to help.We settled on an idea particularly suitable for journals: survey papers.And we decided to focus on the bottom of the storage stack: storage technologies and media.We also debated whether we should focus on futuristic technologies, current, or past ones.In the end, we opted to include everything: all storage technologies, regardless of their age or maturity, can teach us something useful.Normally, TOS authors submit their full manuscript for review.However, because this special issue's survey nature would likely mean longer papers, we wanted to provide better direction to authors.So, we posted a CFP asking the prospective authors to submit a short one-page abstract.We provided guidance to prospective authors as to what makes a good survey paper.Specifically, authors would have to survey many related papers in their chosen area, so as to make their survey the most comprehensive paper on the topic to date.Secondly, it was not enough to just summarize past papers; authors also needed to provide insight into why and how a given technology evolved over the years, and where it might go in the future. Geoffrey H. Kuenning, Youjip Won, Ming Zhao 0002, Erez Zadok |
ACM Trans. Storage | 3 |
| 2025 | WALSH: Write-Aggregating Log-Structured Hashing for Hybrid MemoryabstractPersistent memory (PM) brings important opportunities for improving data storage including the widely used hash tables. However, PM is not friendly to small writes, which causes existing PM hashes to suffer from high hardware write amplification. Hybrid memory offers the performance and concurrency of DRAM and the durability and capacity of PM, but existing hybrid memory hashes cannot deliver high performance, low DRAM footprint, and fast recovery at the same time. This paper proposes WALSH, a flat hash with novel log-structured separate chaining designs to optimize the performance while ensuring low DRAM footprint and fast recovery. To address the overhead of hash resizing and garbage collection (GC), WALSH further proposes partial resizing/GC mechanisms and a 4-phase protocol for concurrent hash operations. As a result, WALSH is the first flat index for hybrid memory with embedded write aggregation ability. A comprehensive evaluation shows that WALSH substantially outperforms state-of-the-art hybrid memory hashes; e.g., its insert throughput is up to 2.4X that of related works while saving more than 87% of DRAM. WALSH also provides efficient recovery; e.g., it can recover a dataset with 1 billion objects in just a few seconds. Yongfeng Wang, Zhiguang Chen 0001, Yutong Lu, Ming Zhao 0002 |
ACM Trans. Storage | 5 |
| 2025 | Introduction to the Special Section on MSST 2024
Ming Zhao 0002, Benjamin C. Reed |
ACM Trans. Storage | 1 |
| 2024 | Self-Supervised Quantization-Aware Knowledge DistillationabstractQuantization-aware training (QAT) and Knowledge Distillation (KD) are combined to achieve competitive performance in creating low-bit deep learning models. However, existing works applying KD to QAT require tedious hyper-parameter tuning to balance the weights of different loss terms, assume the availability of labeled training data, and require complex, computationally intensive training procedures for good performance. To address these limitations, this paper proposes a novel Self-Supervised Quantization-Aware Knowledge Distillation (SQAKD) framework. SQAKD first unifies the forward and backward dynamics of various quantization functions, making it flexible for incorporating various QAT works. Then it formulates QAT as a co-optimization problem that simultaneously minimizes the KL-Loss between the full-precision and low-bit models for KD and the discretization error for quantization, without supervision from labels. A comprehensive evaluation shows that SQAKD substantially outperforms the state-of-the-art QAT and KD works for a variety of model architectures. Our code is at: https://github.com/kaiqi123/SQAKD.git. Kaiqi Zhao 0002, Ming Zhao 0002 |
AISTATS | 2 |
| 2024 | Can ZNS SSDs be Better Storage Devices for Persistent Cache?abstractBlock-based regular SSDs have been widely used as storage backends for persistent cache systems due to their explicitly lower cost and persistence compared to DRAM. However, the caching workloads are both write- and update-intensive. It incurs a large amount of device-level write amplification (WA) in the internal garbage collection (GC), which can lead to SSD lifespan and potential performance issues. Zoned Namespace SSDs (ZNS SSDs) offer a new interface for modern SSDs to overcome the limitations of regular SSDs in some use cases. As ZNS SSDs need much lower internal over-provisioning, they can offer a larger capacity compared with regular SSDs. Considering these two advantages of ZNS SSDs, we aim to explore three possible schemes to adapt the existing persistent cache system on ZNS SSDs and analyze their benefits and limitations. We conduct comprehensive evaluations to further illustrate the tradeoffs of each scheme. Based on our research and investigation, we conclude that ZNS SSDs exhibit promising results as better storage backends for persistent cache. Further, the co-design between cache management and zone management can potentially enhance the cache efficiency and performance. Chongzhuo Yang, Zhang Cao 0002, Ming Zhao 0002, Zhichao Cao 0002 |
HotStorage | 4 |
| 2023 | AdaCache: A Disaggregated Cache System with Adaptive Block Size for Cloud Block StorageabstractNVMe SSD caching has demonstrated impressive capabilities in solving cloud block storage's I/O bottleneck and enhancing application performance in public, private, and hybrid cloud environments. However, traditional host-side caching solutions have several serious limitations. First, the cache cannot be shared across hosts, leading to low cache utilization. Second, the commonly-used fix-sized cache block allocation mechanism is unable to provide good cache performance with low memory overhead for diverse cloud workloads with vastly different I/O patterns. This paper presents AdaCache, a novel userspace disaggregated cache system that utilizes adaptive cache block allocation for cloud block storage. First, AdaCache proposes an innovative adaptive cache block allocation scheme that allocates cache blocks based on the request size to achieve both good cache performance and low memory overhead. Second, AdaCache proposes a group-based cache organization that stores cache blocks into groups to solve the fragmentation problem brought by variable-sized cache blocks. Third, AdaCache designs a two-level cache replacement policy that replaces cache blocks in both single blocks and groups to improve the hit ratio. Experimental results with real-world traces show that AdaCache can substantially improve I/O performance and reduce storage access caused by cache miss with a much lower memory usage compared to traditional fix-sized cache systems. Runyu Jin, Ni Fan, Devasena Inupakutika, Bridget Davis, Ming Zhao 0002 |
CLOUD | 6 |
| 2023 | Automatic Attention Pruning: Improving and Automating Model Pruning using AttentionsabstractPruning is a promising approach to compress deep learning models in order to deploy them on resource-constrained edge devices. However, many existing pruning solutions are based on unstructured pruning, which yields models that cannot efficiently run on commodity hardware; and they often require users to manually explore and tune the pruning process, which is time-consuming and often leads to sub-optimal results. To address these limitations, this paper presents Automatic Attention Pruning (AAP), an adaptive, attention-based, structured pruning approach to automatically generate small, accurate, and hardware-efficient models that meet user objectives. First, it proposes iterative structured pruning using activation-based attention maps to effectively identify and prune unimportant filters. Then, it proposes adaptive pruning policies for automatically meeting the pruning objectives of accuracy-critical, memory-constrained, and latency-sensitive tasks. A comprehensive evaluation shows that AAP substantially outperforms the state-of-the-art structured pruning works for a variety of model architectures. Our code is at: https://github.com/kaiqi123/Automatic-Attention-Pruning.git. Kaiqi Zhao 0002, Animesh Jain, Ming Zhao 0002 |
AISTATS | 3 |
| 2023 | A Contrastive Knowledge Transfer Framework for Model Compression and Transfer LearningabstractKnowledge Transfer (KT) achieves competitive performance and is widely used for image classification tasks in model compression and transfer learning. Existing KT works transfer the information from a large model ("teacher") to train a small model ("student") by minimizing the difference of their conditionally independent output distributions. However, these works overlook the high-dimension structural knowledge from the intermediate representations of the teacher, which leads to limited effectiveness, and they are motivated by various heuristic intuitions, which makes it difficult to generalize. This paper proposes a novel Contrastive Knowledge Transfer Framework (CKTF), which enables the transfer of sufficient structural knowledge from the teacher to the student by optimizing multiple contrastive objectives across the intermediate representations between them. Also, CKTF provides a generalized agreement to existing KT techniques and increases their performance significantly by deriving them as specific cases of CKTF. The extensive evaluation shows that CKTF consistently outperforms the existing KT works by 0.04% to 11.59% in model compression and by 0.4% to 4.75% in transfer learning on various models and datasets. Kaiqi Zhao 0002, Ming Zhao 0002 |
ICASSP | 3 |
| 2023 | Semantic Privacy-Preserving for Video Surveillance Services on the EdgeabstractIntelligent Video surveillance systems, leveraging edge computing, have become increasingly prevalent in various facilities, providing advanced monitoring and management capabilities. However, these systems can inadvertently compromise personally identifiable information, such as human images, leading to privacy violations. We introduced a semantic privacy-preserving video surveillance service on the edge to address this critical issue. Unlike traditional centralized models, the solution operates as a decentralized machine learning framework within the video surveillance infrastructure at the edge. Its primary focus is protecting private information extracted from captured video streaming data. This research integrates cutting-edge machine learning techniques, including scene graph generation and semantic communication approaches, by enabling edge nodes to exchange parameters for training, referencing, and safeguarding data privacy and ownership. These innovations collectively contribute to the protection of human privacy. The performance evaluation confirms that the solution is an efficient and effective privacy protection platform, offering a significant advancement over conventional centralized solutions. Alexander Y. C. Huang, Dijiang Huang, Ming Zhao 0002 |
SEC | 4 |
| 2023 | Poster: Self-Supervised Quantization-Aware Knowledge DistillationabstractQuantization-aware training (QAT) achieves competitive performance and is widely used for image classification tasks in model compression. Existing QAT works start with a pre-trained full-precision model and perform quantization during retraining. However, these works require supervision from the ground-truth labels whereas sufficient labeled data are infeasible in real-world environments. Also, they suffer from accuracy loss due to reduced precision, and no algorithm consistently achieves the best or the worst performance on every model architecture. To address the aforementioned limitations, this paper proposes a novel Self-Supervised Quantization-Aware Knowledge Distillation framework (SQAKD). SQAKD unifies the forward and backward dynamics of various quantization functions, making it flexible for incorporating the various QAT works. With the full-precision model as the teacher and the low-bit model as the student, SQAKD reframes QAT as a co-optimization problem that simultaneously minimizes the KL-Loss (i.e., the Kullback-Leibler divergence loss between the teacher's and student's penultimate outputs) and the discretization error (i.e., the difference between the full-precision weights/activations and their quantized counterparts). This optimization is achieved in a self-supervised manner without labeled data. The evaluation shows that SQAKD significantly improves the performance of various state-of-the-art QAT works (e.g., PACT, LSQ, DoReFa, and EWGS). SQAKD establishes stronger baselines and does not require extensive labeled training data, potentially making state-of-the-art QAT research more accessible. Kaiqi Zhao 0002, Ming Zhao 0002 |
SEC | 2 |
| 2023 | GPU-enabled Function-as-a-Service for Machine Learning InferenceabstractFunction-as-a-Service (FaaS) is emerging as an important cloud computing service model as it can improve the scalability and usability of a wide range of applications, especially Machine-Learning (ML) inference tasks that require scalable resources and complex software configurations. These inference tasks heavily rely on GPUs to achieve high performance; however, support for GPUs is currently lacking in the existing FaaS solutions. The unique event-triggered and short-lived nature of functions poses new challenges to enabling GPUs on FaaS, which must consider the overhead of transferring data (e.g., ML model parameters and inputs/outputs) between GPU and host memory. This paper proposes a novel GPU-enabled FaaS solution that enables ML inference functions to efficiently utilize GPUs to accelerate their computations. First, it extends existing FaaS frameworks such as OpenFaaS to support the scheduling and execution of functions across GPUs in a FaaS cluster. Second, it provides caching of ML models in GPU memory to improve the performance of model inference functions and global management of GPU memories to improve cache utilization. Third, it offers co-designed GPU function scheduling and cache management to optimize the performance of ML inference functions. Specifically, the paper proposes locality-aware scheduling, which maximizes the utilization of both GPU memory for cache hits and GPU cores for parallel processing. A thorough evaluation based on real-world traces and ML models shows that the proposed GPU-enabled FaaS works well for ML inference tasks, and the proposed locality-aware scheduler achieves a speedup of 48x compared to the default, load balancing only schedulers. Ming Zhao 0002, Kritshekhar Jha, Sungho Hong |
IPDPS | 1 |
| 2022 | Exploring Edge Machine Learning-based Stress Prediction using Wearable DevicesabstractStress is a central factor in our daily lives, impacting performance, decisions, well-being, and our interactions with others. With the development of IoT technology, smart wearable devices can handle diverse operations, including networking and recording biometric signals. The enhanced data processing capability of wearables has also allowed for increased stress awareness among users. Edge computing on such devices enables real-time feedback which can provide an opportunity to prevent severe consequences that might result if stress is left unaddressed. Edge computing can also strengthen privacy by implementing stress prediction on local devices without transferring personal information to the public cloud.This paper presents a framework for real-time stress prediction, specifically for police training cadets, using wearable devices and machine learning with support from cloud computing. We developed an application for Fitbit and the user's accompanying smartphone to collect heart rate fluctuations and corresponding stress levels entered by users and a cloud backend for storing data and training models. Real-world data for this study was collected from police cadets during a police academy training program. Machine learning classifiers for stress prediction were built using this data through classic machine learning models and neural networks. To analyze efficiency across different environments, the models were optimized using model compression and other relevant techniques and tested on cloud and edge environments. Evaluation using real data and real devices showed that the highest accuracy came from XGBoost and Tensorflow neural network models, and on-edge stress prediction models produced lower latency results than in-cloud prediction. Sang-Hun Sim, Tara Paranjpe, Nicole Roberts, Ming Zhao 0002 |
ICMLA | 4 |
| 2022 | Knowledge Distillation via Module Replacing for Automatic Speech Recognition with Recurrent Neural Network Transducer
Kaiqi Zhao 0002, Animesh Jain, Nathan Susanj, Athanasios Mouchtaris, Lokesh Gupta, Ming Zhao 0002 |
INTERSPEECH | 7 |
| 2022 | Performance Evaluation on CXL-enabled Hybrid Memory PoolabstractThe emerging cache coherent Compute Express Link (CXL) interconnect provides a practical way to disaggregate cloud memory resources from monolithic servers into memory pools with DRAM-level access latency. While DRAM-only memory pool improves the resource utilization and reduces the Total Cost of Ownership (TCO) for cloud providers, we investigate the possibility of applying cheaper SSDs to a memory pooling system to further reduce the cost of cloud servers without sacrificing the application's performance. In this study, we build a simulated CXL-enabled DRAM-SSD hybrid memory pool based on Linux and commodity hardware, and conduct performance evaluation by running representative cloud workloads which cover deep learning training, database, data analytics and video processing on the testbed. The evaluation results show that a hybrid memory pool can potentially reduce memory cost while maintaining the same level of application performance for computation-intensive applications. For example, with memory overcommit ratio of 2, the performance degradation of training ResNet50 on ImageNet dataset is only 2.68%. Runyu Jin, Bridget Davis, Devasena Inupakutika, Ming Zhao 0002 |
NAS | 5 |
| 2021 | Characterizing Loop Acceleration in Heterogeneous ComputingabstractComputation intensive applications usually consist of multiple nested or flattened loops. These loops are the main building blocks of the applications and embody a specific type of execution pattern. In order to reduce the running time of the loops, developers need to analyze the loops in the code and try to parallelize them on hardware accelerators, such as GPUs, TPUs, and FPGAs, which are increasingly available in the cloud. Unfortunately, the lack of understanding of loop characteristics and the ability of hardware accelerators in handling these types of loops prevents developers from choosing the right platform to develop their applications in the cloud. Also, developing and optimizing code for a specific accelerator is a time-consuming effort. To address these issues, this paper studies the effectiveness of different processors in accelerating common patterns of loops. It identifies five important types of loops that commonly exist in real-world applications, and presents Loopy, the implementations of these loops optimized for different architectures. Using Loopy, the paper also evaluates different hardware in accelerating the loop patterns. The result reveals the architectural differences among different accelerators with regard to different loop patterns. It also provides insights for the developers to choose the right accelerators for their applications. The current version of Loopy supports both FPGAs and GPUs, which are the most versatile and available accelerators. Saman Biookaghazadeh, Fengbo Ren, Ming Zhao 0002 |
CLOUD | 3 |
| 2021 | Learning Cache Replacement with CACHEUS
Liana V. Rodriguez, Farzana Beente Yusuf, Steven Lyons, Eysler Paz, Raju Rangaswami, Jason Liu 0001, Ming Zhao 0002, Giri Narasimhan |
FAST | 7 |
| 2021 | SecureFL: Privacy Preserving Federated Learning with SGX and TrustZone
Eugene N. Kuznetsov, Ming Zhao 0002 |
SEC | 3 |
| 2021 | Toward Multi-FPGA Acceleration of the Neural NetworksabstractHigh-throughput and low-latency Convolutional Neural Network (CNN) inference is increasingly important for many cloud- and edge-computing applications. FPGA-based acceleration of CNN inference has demonstrated various benefits compared to other high-performance devices such as GPGPUs. Current FPGA CNN-acceleration solutions are based on a single FPGA design, which are limited by the available resources on an FPGA. In addition, they can only accelerate conventional 2D neural networks. To address these limitations, we present a generic multi-FPGA solution, written in OpenCL, which can accelerate more complex CNNs (e.g., C3D CNN) and achieve a near linear speedup with respect to the available single-FPGA solutions. The design is built upon the Intel Deep Learning Accelerator architecture, with three extensions. First, it includes updates for better area efficiency (up to 25%) and higher performance (up to 24%). Second, it supports 3D convolutions for more challenging applications such as video learning. Third, it supports multi-FPGA communication for higher inference throughput. The results show that utilizing multiple FPGAs can linearly increase the overall bandwidth while maintaining the same end-to-end latency. In addition, the design can outperform other FPGA 2D accelerators by up to 8.4 times and 3D accelerators by up to 1.7 times. Saman Biookaghazadeh, Pravin Kumar Ravi, Ming Zhao 0002 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2020 | Pacon: Improving Scalability and Efficiency of Metadata Service through Partial ConsistencyabstractTraditional distributed file systems (DFS) use centralized service to manage metadata. Many studies based on this centralized architecture enhanced metadata processing capability by scaling the metadata server cluster, which is however still difficult to keep up with the growing number of clients and the increasingly metadata-intensive applications. Some solutions abandoned the centralized metadata service and improved scalability by embedding a private metadata service in an HPC application, but these solutions are suitable for only some specific applications and the absence of global namespace makes data sharing and management difficult. This paper addresses the shortcomings of existing studies by optimizing the consistency model of client- side metadata cache for the HPC scenario using a novel partial consistency model. It provides the application with strong consistency guarantee for only its workspace, thus improving metadata scalability without adding hardware or sacrificing the versatility and manageability of DFSes. In addition, the paper proposes batch permission management to reduce path traversal overhead, thereby improving metadata processing efficiency. The result is a library (Pacon) that allows existing DFSes to achieve partial consistency for scalable and efficient metadata management. The paper also presents a comprehensive evaluation using intensive benchmarks and representative application. For example, in file creation, Pacon improves the performance of BeeGFS by more than 76.4 times, and outperforms the state-of-the-art metadata management solution (IndexFS) by more than 4.6 times. Yutong Lu, Zhiguang Chen 0001, Ming Zhao 0002 |
IPDPS | 4 |
| 2019 | An Efficient and Flexible Metadata Management Layer for Local File SystemsabstractThe efficiency of metadata processing affects the file system performance significantly. There are two bottlenecks in metadata management in existing local file systems: 1) Path lookup is costly because it causes a lot of disk I/Os, which makes metadata operations inefficient. 2) Existing file systems have deep I/O stack in metadata management, resulting in additional processing overhead. To solve these two bottlenecks, we decoupled data and metadata management and proposed a metadata management layer for local file systems. First, we separated the metadata based on their locations in the namespace tree and aggregated the metadata into fixed-size metadata buckets (MDBs). This design fully utilizes the metadata locality and improves the efficiency of disk I/O in the path lookup. Second, we customized an efficient MDB storage system on the raw storage device. This design simplifies the file system I/O stack in the metadata management and allows metadata lookup to be completed with constant time complexity. Finally, this metadata management layer gives users the flexibility to choose metadata storage devices. We implemented a prototype called Otter. Our evaluation demonstrated that Otter outperforms native EXT4, XFS, Btrfs, BetrFS and TableFS in many metadata operations. For instance, Otter has 1.2 times to 9.6 times performance improvement over other tested file systems in file opening. Hongbo Li 0007, Yutong Lu, Zhiguang Chen 0001, Ming Zhao 0002 |
ICCD | 5 |
| 2019 | vPFS+: Managing I/O Performance for Diverse HPC ApplicationsabstractHigh-performance computing (HPC) systems are increasingly shared by a variety of data-and metadata-intensive parallel applications. However, existing parallel file systems employed for HPC storage management are unable to differentiate the I/O requests from concurrent applications and meet their different performance requirements. Previous work, vPFS, provided a solution to this problem by virtualizing a parallel file system and enabling proportional-share bandwidth allocation to the applications; but it cannot handle the increasingly diverse applications in today's HPC environments, including those that have different sizes of I/Os and those that are metadata-intensive. This paper presents vPFS+ which builds upon the virtualization framework provided by vPFS but addresses its limitations in supporting diverse HPC applications. First, a new proportional-share I/O scheduler, SFQ(D)+, is created to allow applications with various I/O sizes and issue rates to share the storage with good application-level fairness and system-level utilization. Second, vPFS+ extends the scheduling to also include metadata I/Os and provides performance isolation to metadata-intensive applications. vPFS+ is prototyped on PVFS2, a widely used open-source parallel file system, and evaluated using a comprehensive set of representative HPC benchmarks and applications (IOR, NPB BTIO, WRF, and multi-md-test). The results confirm that the new SFQ(D)+ scheduler can provide significantly better performance isolation to applications with small, bursty I/Os than the traditional SFQ(D) scheduler (3.35 times better) and the native PVFS2 (8.25 times better) while still making efficient use of the storage. The results also show that vPFS+ can deliver near-perfect proportional sharing (>95% of the target sharing ratio) to metadata-intensive applications. Ming Zhao 0002, Yiqi Xu |
MSST | 1 |
| 2019 | SmartDedup: Optimizing Deduplication for Resource-constrained Devices
Runyu Jin, Ming Zhao 0002 |
USENIX ATC | 3 |
| 2018 | RTVirt: enabling time-sensitive computing on virtualized systems through cross-layer CPU schedulingabstractVirtualization enables flexible application delivery and efficient resource consolidation, and is pervasively used to build various virtualized systems including public and private cloud computing systems. Many applications can benefit from computing on virtualized systems, including those that are time sensitive, but it is still challenging for existing virtualized systems to deliver application-desired timeliness. In particular, the lack of awareness between VM host- and guest-level schedulers presents a serious hurdle to achieving strong timeliness guarantees on virtualized systems. This paper presents RTVirt, a new solution to time-sensitive computing on virtualized systems through cross-layer scheduling. It allows the two levels of schedulers on a virtualized system to communicate key scheduling information and coordinate on the scheduling decisions. It enables optimal multiprocessor schedulers to support virtualized time-sensitive applications with strong timeliness guarantees and efficient resource utilization. RTVirt is prototyped on a widely used virtualization framework (Xen) and evaluated with diverse workloads. The results show that it can meet application deadlines (99%) or tail latency requirements (99.9th percentile) nearly perfectly; it can handle large numbers of applications and dynamic changes in their timeliness requirements; and it substantially outperforms the existing solutions in both timeliness and resource utilization. Ming Zhao 0002, Jorge Cabrera 0002 |
EuroSys | 1 |
| 2018 | Driving Cache Replacement with ML-based LeCaR
Giuseppe Vietri, Liana V. Rodriguez, Wendy A. Martinez, Steven Lyons, Jason Liu 0001, Raju Rangaswami, Ming Zhao 0002, Giri Narasimhan |
HotStorage | 7 |
| 2018 | Cross-Layer Optimization for Virtual Machine Resource ManagementabstractVirtualized systems (e.g., public and private clouds) are playing an increasingly vital role to support the computing of applications from different domains. Existing resource management solutions in such systems typically treat virtual machines (VMs) as black boxes, which presents a hurdle to achieving application-desired Quality of Service (QoS). This paper advocates the cooperation between VM host- and guest-layer schedulers for optimizing the resource management and application performance. It presents an approach to such cross-layer optimization by enabling the host-layer scheduler to feedback resource allocation decisions and adapt guest-layer application configurations. As case studies, the proposed approach is applied to virtualized databases and map services which have challenging dynamic and complex resource demands as well as sophisticated configurations. Specifically, for databases, the proposed approach adapts query executions by tuning the cost model parameters according to the available storage bandwidth and memory capacity. For map services, it adapts the quality of returned map imagery in order to meet the response time target as the workload intensity and available network bandwidth change over time. A prototype of the proposed approach is implemented on Xen and Hyper-V VMs, and evaluated using a TPC-H based database workload and a TerraFly-based map service workload. The results show that with the proposed host-to-guest application adaptation, the TPC-H workload improves its performance by 33.5%, and the TerraFly workload improves the map imagery quality by 40% and always meets its response time target, compared to the schemes without adaptation. Ming Zhao 0002, Lixi Wang, Yun Lv, Jing Xu 0012 |
IC2E | 1 |
| 2018 | DA Placement: A Dual-Aware Data Placement in a Deduplicated and Erasure-Coded Storage System
Mingzhu Deng, Ming Zhao 0002, Fang Liu 0002, Zhiguang Chen 0001, Nong Xiao 0001 |
ICA3PP (1) | 2 |
| 2018 | Improving the Performance and Endurance of Encrypted Non-Volatile Main Memory through Deduplicating WritesabstractNon-volatile memory (NVM) technologies are considered as promising candidates of the next-generation main memory. However, the non-volatility of NVMs leads to new security vulnerabilities. For example, it is not difficult to access sensitive data stored on stolen NVMs. Memory encryption can be employed to mitigate the security vulnerabilities, but it increases the number of bits written to NVMs due to the diffusion property and thereby aggravates the NVM wear-out induced by writes. To address these security and endurance challenges, this paper proposes DeWrite, a secure and deduplication-aware scheme to enhance the performance and endurance of encrypted NVMs based on a new in-line deduplication technique and the synergistic integrations of deduplication and memory encryption. Specifically, it performs low-latency in-line deduplication to exploit the abundant cache-line-level duplications leveraging the intrinsic read/write asymmetry of NVMs and light-weight hashing. It also opportunistically parallelizes the operations of deduplication and encryption and allows them to co-locate the metadata for high time and space efficiency. DeWrite was implemented on the gem5 with NVMain and evaluated using 20 applications from SPEC CPU2006 and PARSEC. Extensive experimental results demonstrate that DeWrite reduces on average 54% writes to encrypted NVMs, and speeds up memory writes and reads of encrypted NVMs by 4.2 × and 3.1 ×, respectively. Meanwhile, DeWrite improves the system IPC by 82% and reduces 40% of energy consumption on average. Pengfei Zuo, Yu Hua 0001, Ming Zhao 0002, Wen Zhou 0030, Yuncheng Guo |
MICRO | 3 |
| 2017 | Kaleido: Enabling Efficient Scientific Data Processing on Big-Data SystemsabstractBig-Data systems are increasingly important for solving the data-driven problems in many science domains including geosciences. However, existing big- data systems cannot support the efficient processing of self-describing data formats such as NetCDF which are commonly used by scientific communities for data distribution and sharing. This limitation presents a serious hurdle to the further adoption of big-data systems by science domains. This paper presents Kaleido, a solution to this problem by enabling big- data systems to efficiently store and process scientific data. Specifically, it enables Hadoop to directly store NetCDF data on HDFS, and process them in MapReduce using convenient APIs. It also enables Hive to support queries on NetCDF data, transparent to the users. Moreover, it employs optimizations tailored to scientific data, particularly dimension-aware layout which allows efficient execution of subset queries targeting any dimension of the multi- dimensional data. The paper presents a comprehensive evaluation of Kaleido using representative queries on a typical geoscientific dataset. The results show that Kaleido achieves substantial speedup and space saving compared to existing solutions for storing and processing NetCDF data on Hadoop, and it also substantially outperforms the state-of-the-art solutions for supporting subset queries on scientific data. Saman Biookaghazadeh, Shujia Zhou, Ming Zhao 0002 |
NAS | 3 |
| 2017 | TCASM: An asynchronous shared memory interface for high-performance application composition
Douglas Otstott, Latchesar Ionkov, Michael Lang 0003, Ming Zhao 0002 |
Parallel Comput. | 4 |
| 2016 | CloudCache: On-demand Flash Cache Management for Cloud Computing
Dulcardo Arteaga, Jorge Cabrera 0002, Jing Xu 0012, Swaminathan Sundararaman, Ming Zhao 0002 |
FAST | 5 |
| 2016 | CacheDedup: In-line Deduplication for Flash Caching
Wenji Li, Gregory Jean-Baptise, Juan Riveros, Giri Narasimhan, Tony Zhang, Ming Zhao 0002 |
FAST | 6 |
| 2016 | IBIS: Interposed Big-data I/O SchedulerabstractBig-data systems are increasingly shared by diverse, data-intensive applications from different domains. However, existing systems lack the support for I/O management, and the performance of big-data applications degrades in unpredictable ways when they contend for I/Os. To address this challenge, this paper proposes IBIS, an Interposed Big-data I/O Scheduler, to provide I/O performance differentiation for competing applications in a shared big-data system. IBIS transparently intercepts, isolates, and schedules an application's different phases of I/Os via an I/O interposition layer on every datanode of the big-data system. It provides a new proportional-share I/O scheduler, SFQ(D2), to allow applications to share the I/O service of each datanode with good fairness and resource utilization. It enables the distributed I/O schedulers to coordinate with one another and to achieve proportional sharing of the big-data system's total I/O service in a scalable manner. Finally, it supports the shared use of big-data resources by diverse frameworks and manages the I/Os from different types of big-data workloads (e.g., batch jobs vs. queries) across these frameworks. The prototype of IBIS is implemented in Hadoop/YARN, a widely used big-data system. Experiments based on a variety of representative applications (WordCount, TeraSort, Facebook, TPC-H) show that IBIS achieves good total-service proportional sharing with low overhead in both application performance and resource usages. IBIS is also shown to support various performance policies: it can deliver stronger performance isolation than native Hadoop/YARN (99% better for WordCount and 15% better for TPC-H queries) with good resource utilization; and it can also achieve perfect proportional slowdown with better application performance (30% better than native Hadoop). Yiqi Xu, Ming Zhao 0002 |
HPDC | 2 |
| 2015 | Enabling scientific data storage and processing on big-data systemsabstractBig-data systems are increasingly important for solving the data-driven problems in many science domains including geosciences. However, existing big-data systems cannot support the self-describing data formats such as NetCDF which are commonly used by scientific communities for data distribution and sharing. This limitation presents a serious hurdle to the further adoption of big-data systems by science domains and prevents scientific users from leveraging these systems to improve their productivity. This paper presents a solution to this problem by enabling big-data systems to directly store and process scientific data. Specifically, it enables Hadoop to efficiently store NetCDF data on HDFS and process them in MapReduce using convenient APIs. It also enables Hive to support standard queries on NetCDF data, transparently to users. The paper also presents an evaluation of the proposed solution using several representative queries on a typical geoscientific dataset. The results show that the proposed approach achieves substantial speedup (up to 20 times) and space saving (83% reduction), compared to the traditional approach which has to convert NetCDF data to CSV format for Hadoop and Hive to use them. Saman Biookaghazadeh, Yiqi Xu, Shujia Zhou, Ming Zhao 0002 |
IEEE BigData | 4 |
| 2014 | Enabling composite applications through an asynchronous shared memory interfaceabstractIn this work we address the growing need for mechanisms for intranode application composition. We provide a novel shared memory interface that allows composite applications, two or more coupled applications, to share internal data structures without blocking. This allows independent progress of the applications such that they can proceed in a parallel, overlapped fashion. Composite applications using in-node shared memory can reduce the amount of data to be communicated between nodes, allowing data reduction or analytics to be performed locally and in parallel. To validate our approach we implemented our solution in Linux and used two proxy-applications to demonstrate how applications can be coupled and compare the performance to a traditional solution. We also compared the impact of composite applications to the performance of their unmodified versions. Our solution incurs small overhead in HPC Linux environments and significantly outperforms preexisting approaches. Douglas Otstott, Noah Evans, Latchesar Ionkov, Ming Zhao 0002, Michael Lang 0003 |
IEEE BigData | 4 |
| 2014 | BigCache for big-data systemsabstractBig-data systems are increasingly used in many disciplines for important tasks such as knowledge discovery and decision making by processing large volumes of data. Big-data systems rely on hard-disk drive (HDD) based storage to provide the necessary capacity. However, as big-data applications grow rapidly more diverse and demanding, HDD storage becomes insufficient to satisfy their performance requirements. Emerging solid-state drives (SSDs) promise great IO performance that can be exploited by big-data applications, but they still face serious limitations in capacity, cost, and endurance and therefore must be strategically incorporated into big-data systems. This paper presents BigCache, an SSD-based distributed caching layer for big-data systems. It is designed to be seamlessly integrated with existing big-data systems and transparently accelerate IOs for diverse big-data applications. The management of the distributed SSD caches in BigCache is coordinated with the job management of big-data systems in order to support cache-locality-driven job scheduling. BigCache is prototyped in Hadoop to provide caching upon HDFS for MapReduce applications. It is evaluated using typical MapReduce applications, and the results show that BigCache reduces the runtime of WordCount by 38% and the runtime of TeraSort by 52%. The results also show that BigCache is able to achieve significant speedup by caching only partial input for the benchmarks, owing to its ability to cache partial input and its replacement policy that recognizes application access patterns. Michel Angelo Roger, Yiqi Xu, Ming Zhao 0002 |
IEEE BigData | 3 |
| 2014 | Client-side Flash Caching for Cloud SystemsabstractAs the size of cloud systems and the number of hosted VMs rapidly grow, the scalability of shared VM storage systems becomes a serious issue. Client-side flash-based caching has the potential to improve the performance of cloud VM storage by employing flash storage available on the client-side of the storage system to exploit the locality inherent in VM IOs. However, because of the limited capacity and durability of flash storage, it is important to determine the proper size and configuration of the flash caches used in cloud systems. This paper provides answers to the key design questions of cloud flash caching based on dm-cache, a block-level caching solution customized for cloud environments, and a large amount of long-term traces collected from real-world public and private clouds. The study first validates that cloud workloads have good cacheability and dm-cache-based flash caching incurs low overhead with respect to commodity flash devices. It further reveals that write-back caching substantially outperforms write-through caching in typical cloud environments due to the reduction of server IO load. It also shows that there is a tradeoff on making a flash cache persistent across client restarts which saves hours of cache warm-up time but incurs considerable overhead from committing every metadata update persistently. Finally, to reduce the data loss risk from using write-back caching, the paper proposes a new cache-optimized RAID technique, which minimizes the RAID overhead by introducing redundancy of cache dirty data only, and shows to be significantly faster than traditional RAID and write-through caching. Dulcardo Arteaga, Ming Zhao 0002 |
SYSTOR | 2 |
| 2013 | Write policies for host-side flash caches
Ricardo Koller, Leonardo Mármol, Raju Rangaswami, Swaminathan Sundararaman, Nisha Talagala, Ming Zhao 0002 |
FAST | 6 |
| 2013 | IBIS: interposed big-data I/O scheduler
Yiqi Xu, Adrian Suarez, Ming Zhao 0002 |
HPDC | 3 |
| 2013 | Massive GIS Database System with Autonomic Resource ManagementabstractGIS application hosts are becoming more and more complicated. Thus, their management is more time consuming, and reliability decreases with the complexity of GIS applications increasing. We have designed, implemented, and evaluated, a virtualized whole Large Scale Distributed Spatial Data Visualization System for optimizing maintainability and performance when handling large amount of GIS data. We employ the virtual machines (VMs) technique, load balance cluster techniques, and autonomic resource management to improve the system's performance. The proposed system was prototyped on TerraFly [1], a production web map service, and evaluated using actual TerraFly workloads. The results show that the virtual TerraFly system has both good performance and much better maintainability. Our experiments show that the proposed Virtual TerraFly Geo-database system has doubled the reliability, and saved 20-30% computing resources cost compared to current static peak-load physical machine node allocations. Yun Lu 0001, Ming Zhao 0002, Guangqiang Zhao, Lixi Wang, Naphtali Rishe |
ICMLA (2) | 2 |
| 2013 | On the design and implementation of a simulator for parallel file system researchabstractDue to the popularity and importance of Parallel File Systems (PFSs) in modern High Performance Computing (HPC) centers, PFS designs and I/O optimizations are active research topics. However, the research process is often time-consuming and faces cost and complexity challenges in deploying experiments in real HPC systems. This paper describes PFSsim, a trace-driven simulator of distributed storage systems that allows the evaluation of PFS designs, I/O schedulers, network structures, and workloads. PFSsim differentiates itself from related work in that it provides a powerful platform featuring a modular design with high flexibility in the modeling of subsystems including the network, clients, data servers and I/O schedulers. It does so by designing the simulator to capture abstractions found in common PFSs. PFSsim also exposes script-based interfaces for detailed configurations. Experiments and validation against real systems considering sub-modules and the entire simulator show that PFSsim is capable of simulating a representative PFS (PVFS2) and of modeling different I/O scheduler algorithms with good fidelity. In addition, the simulation speed is also shown to be acceptable. Yonggang Liu 0004, Renato J. O. Figueiredo, Yiqi Xu, Ming Zhao 0002 |
MSST | 4 |
| 2012 | vPFS: Bandwidth virtualization of parallel storage systemsabstractExisting parallel file systems are unable to differentiate I/Os requests from concurrent applications and meet per-application bandwidth requirements. This limitation prevents applications from meeting their desired Quality of Service (QoS) as high-performance computing (HPC) systems continue to scale up. This paper presents vPFS, a new solution to address this challenge through a bandwidth virtualization layer for parallel file systems. vPFS employs user-level parallel file system proxies to interpose requests between native clients and servers and to schedule parallel I/Os from different applications based on configurable bandwidth management policies. vPFS is designed to be generic enough to support various scheduling algorithms and parallel file systems. Its utility and performance are studied with a prototype which virtualizes PVFS2, a widely used parallel file system. Enhanced proportional sharing schedulers are enabled based on the unique characteristics (parallel striped I/Os) and requirement (high throughput) of parallel storage systems. The enhancements include new threshold- and layout-driven scheduling synchronization schemes which reduce global communication overhead while delivering total-service fairness. An experimental evaluation using typical HPC benchmarks (IOR, NPB BTIO) shows that the throughput overhead of vPFS is small (;96% of target sharing ratio) for competing applications with diverse I/O patterns. Yiqi Xu, Dulcardo Arteaga, Ming Zhao 0002, Yonggang Liu 0004, Renato J. O. Figueiredo, Seetharami Seelam |
MSST | 3 |
| 2012 | Modeling virtualized applications using machine learning techniquesabstractWith the growing adoption of virtualized datacenters and cloud hosting services, the allocation and sizing of resources such as CPU, memory, and I/O bandwidth for virtual machines (VMs) is becoming increasingly important. Accurate performance modeling of an application would help users in better VM sizing, thus reducing costs. It can also benefit cloud service providers who can offer a new charging model based on the VMs' performance instead of their configured sizes. In this paper, we present techniques to model the performance of a VM-hosted application as a function of the resources allocated to the VM and the resource contention it experiences. To address this multi-dimensional modeling problem, we propose and refine the use of two machine learning techniques: artificial neural network (ANN) and support vector machine (SVM). We evaluate these modeling techniques using five virtualized applications from the RUBiS and Filebench suite of benchmarks and demonstrate that their median and 90th percentile prediction errors are within 4.36% and 29.17% respectively. These results are substantially better than regression based approaches as well as direct applications of machine learning techniques without our refinements. We also present a simple and effective approach to VM sizing and empirically demonstrate that it can deliver optimal results for 65% of the sizing problems that we studied and produces close-to-optimal sizes for the remaining 35%. Sajib Kundu, Raju Rangaswami, Ajay Gulati, Ming Zhao 0002, Kaushik Dutta |
VEE | 4 |
| 2011 | Fuzzy Modeling Based Resource Management for Virtualized Database SystemsabstractThe hosting of databases on virtual machines (VMs) has great potential to improve the efficiency of resource utilization and the ease of deployment of database systems. This paper considers the problem of on-demand allocation of resources to a VM running a database serving dynamic and complex query workloads while meeting QoS (Quality of Service) requirements. An autonomic resource-management approach is proposed to address this problem. It uses adaptive fuzzy modeling to capture the behavior of a VM hosting a database with dynamically changing workloads and to predict its multi-type resource needs. A prototype of the proposed approach is implemented on Xen-based VMs and evaluated using workloads based on TPC-H and RUBiS. The results demonstrate that CPU and disk I/O bandwidth can be efficiently allocated to database VMs serving workloads with dynamically changing intensity and composition while meeting QoS targets. For TPC-H-based experiments, the resulting throughput is within 89.5-100% of what would be obtained using resource allocation based on peak loads, For RUBiS, the response time target (set based on the performance under peak-load-based allocation) is met for 97% of the time. Moreover, substantial resources are saved (about 62.6% of CPU and 76.5% of disk I/O bandwidth) in comparison to peak-load-based allocation. Lixi Wang, Jing Xu 0012, Ming Zhao 0002, Yi-Cheng Tu, José A. B. Fortes |
MASCOTS | 3 |
| 2011 | Towards Scalable Application Checkpointing with Parallel File System DelegationabstractThe ever-increasing scale of modern high-performance computing (HPC) systems presents a variety of challenges to the parallel file system (PFS) based storage in these systems. The scalability of application check pointing is a particularly important challenge because it is critical to the reliability of computing and it often dominates the I/Os in a HPC system. When a large number of parallel processes simultaneously perform check pointing, the PFS metadata servers can become a serious bottleneck due to the large volume of concurrent metadata operations. This paper specifically addresses this PFS metadata management issue in order to support scalable application check pointing in large HPC systems. It proposes a new technique named PFS-delegation which delegates the management of the PFS storage space used for check pointing to applications, thereby relieving the load of metadata operations on the PFS during their check pointing. This proposed technique is prototyped on PVFS2, a widely used PFS implementation, and evaluated on a HPC cluster using a representative parallel I/O benchmark, IOR. Experiments with up to 128 parallel processes show that the PFS-delegation based check pointing is significantly faster than the traditional shared-file and file-per-process based check pointing methods (7% and 10% speedup when the underlying PVFS2 uses a centralized metadata server, 22% and 31% speedup when using distributed metadata servers). The results also demonstrate that the PFS-delegation based check pointing substantially reduces the total number of metadata operations handled by the metadata servers during the check pointing. Dulcardo Arteaga, Ming Zhao 0002 |
NAS | 2 |
| 2010 | Application performance modeling in a virtualized environmentabstractPerformance models provide the ability to predict application performance for a given set of hardware resources and are used for capacity planning and resource management. Traditional performance models assume the availability of dedicated hardware for the application. With growing application deployment on virtualized hardware, hardware resources are increasingly shared across multiple virtual machines. In this paper, we build performance models for applications in virtualized environments. We identify a key set of virtualization architecture independent parameters that influence application performance for a diverse and representative set of applications. We explore several conventional modeling techniques and evaluate their effectiveness in modeling application performance in a virtualized environment. We propose an iterative model training technique based on artificial neural networks which is found to be accurate across a range of applications. The proposed approach is implemented as a prototype in Xen-based virtual machine environments and evaluated for accuracy, sensitivity to the training process, and overhead. Median modeling error in the range 1.16-6.65% across a diverse application set and low modeling overhead suggest the suitability of our approach in production virtualized environments. Sajib Kundu, Raju Rangaswami, Kaushik Dutta, Ming Zhao 0002 |
HPCA | 4 |
| 2009 | Cooperative Autonomic Management in Dynamic Distributed Systems
Jing Xu 0012, Ming Zhao 0002, José A. B. Fortes |
SSS | 2 |
| 2007 | A user-level secure grid file systemabstractA grid-wide distributed file system provides convenient data access interfaces that facilitate fine-grained cross-domain data sharing and collaboration. However, existing widely-adopted distributed file systems do not meet the security requirements for grid systems. This paper presents a Secure Grid File System (SGFS) which supports GSI-based authentication and access control, end-to-end message privacy, and integrity. It employs user-level virtualization of NFS to provide transparent grid data access leveraging existing, unmodified clients and servers. It supports user and application-tailored security customization per SGFS session, and leverages secure management services to control and configure the sessions. The system conforms to the GSI grid security infrastructure and allows for seamless integration with other grid middleware. A SGFS prototype is evaluated with both file system benchmarks and typical applications, which demonstrates that it can achieve strong security with an acceptable overhead, and substantially outperform native NFS in wide-area environments by using disk caching. Ming Zhao 0002, Renato J. O. Figueiredo |
SC | 1 |
| 2006 | Application-Tailored Cache Consistency for Wide-Area File SystemsabstractThe inability to perform optimizations based on application-specific information presents a hurdle to the deployment of pervasive LAN file systems across WAN environments. This paper proposes a novel approach addressing this problem through application-tailored caching and consistency in widearea file systems. It leverages widely available Network File System (NFS) deployments without any modifications to kernels nor applications, and employs middleware to dynamically establish Grid-wide Virtual File System (GVFS) sessions with application-tailored cache consistency. Two consistency models are discussed in this paper: a relaxed model based on invalidation polling, and a stronger model based on delegation and callback. Experimental evaluation based on microbenchmarks and scientific applications show that with application-tailored cache consistency, GVFS is able to both improve application runtimes and reduce server load significantly, compared to kernel-level NFS in WAN. Ming Zhao 0002, Renato J. O. Figueiredo |
ICDCS | 1 |
| 2005 | On the Use of Virtualization and Service Technologies to Enable Grid-Computing
Andréa M. Matsunaga, Maurício O. Tsugawa, Ming Zhao 0002, Vivekananthan Sanjeepan, Sumalatha Adabala, Renato J. O. Figueiredo, Herman Lam, José A. B. Fortes |
Euro-Par | 3 |
| 2005 | Supporting application-tailored grid file system sessions with WSRF-based servicesabstractThis paper presents novel service-based grid data management middleware that leverages standards defined by WSRF specifications to create and manage dynamic grid file system sessions. A unique aspect of the service is that the sessions it creates can be customized to address application data transfer needs. Application-tailored configurations enable selection of both performance-related features (block-based partial file transfers and/or whole-file transfers, cache parameters and consistency models) and reliability features (file system copy-on-write checkpointing to aid recovery of client-side failures; replication, autonomous failure detection and data access redirection for server-side failures). These enhancements, in addition to cross-domain user identity mapping and encrypted communication, are implemented via user level proxies managed by the service, requiring no changes to existing kernels. Sessions established using the service is mounted as distributed file systems and can be used transparently by unmodified binary applications. The paper analyzes the use of the service to support virtual machine based grid systems and workflow execution, and also reports on the performance and reliability of service managed wide-area file system sessions with experiments based on scientific applications (NanoMOS/Matlab, CHID, GAUSS and SPECseis). Ming Zhao 0002, Vineet Chadha, Renato J. O. Figueiredo |
HPDC | 1 |
| 2005 | From virtualized resources to virtual computing grids: the In-VIGO system
Sumalatha Adabala, Vineet Chadha, Puneet Chawla, Renato J. O. Figueiredo, José A. B. Fortes, Ivan Krsul, Andréa M. Matsunaga, Maurício O. Tsugawa, Jian Zhang 0005, Ming Zhao 0002 |
Future Gener. Comput. Syst. | 10 |
| 2004 | Distributed File System Support for Virtual Machines in Grid Computing
Ming Zhao 0002, Jian Zhang 0005, Renato J. O. Figueiredo |
HPDC | 1 |