VLDB 2026 Research / reviewers in the wild / expert
Jun Wang 0001
dblp:w/JunWang1
· DBLP profile ↗
119ranked-venue papers
21as first author
22since 2021 · last 2026
0000-0002-0926-4761ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 90 · 19 first-author · 12 since 2021Computer networks · 11 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-authorArtificial intelligence and machine learning · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?abstractLarge vision-language models (LVLMs) have achieved remarkable advancements in multimodal reasoning tasks. However, their widespread accessibility raises critical concerns about potential copyright infringement. Will LVLMs accurately recognize and comply with copyright regulations when encountering copyrighted content (i.e., user input, retrieved documents) in the context? Failure to comply with copyright regulations may lead to serious legal and ethical consequences, particularly when LVLMs generate responses based on copyrighted materials (e.g., retrieved book experts, news reports). In this paper, we present a comprehensive evaluation of various LVLMs, examining how they handle copyrighted content – such as book excerpts, news articles, music lyrics, and code documentation when they are presented as visual inputs. To systematically measure copyright compliance, we introduce a large-scale benchmark dataset comprising 50,000 multimodal query-content pairs designed to evaluate how effectively LVLMs handle queries that could lead to copyright infringement. Given that real-world copyrighted content may or may not include a copyright notice, the dataset includes query-content pairs in two distinct scenarios: with and without a copyright notice. For the former, we extensively cover four types of copyright notices to account for different cases. Our evaluation reveals that even state-of-the-art closed-source LVLMs exhibit significant deficiencies in recognizing and respecting the copyrighted content, even when presented with the copyright notice. To solve this limitation, we introduce a novel tool-augmented defense framework for copyright compliance, which reduces infringement risks in all scenarios. Our findings underscore the importance of developing copyright-aware LVLMs to ensure the responsible and lawful use of copyrighted content. Naen Xu, Jinghuai Zhang, Changjiang Li, Hengyu An, Chunyi Zhou 0001, Jun Wang 0001, Yuyuan Li 0001, Tianyu Du, Shouling Ji |
AAAI | 6 |
| 2026 | Classifier Enhancement Using Extended Context and Domain Experts for Semantic SegmentationabstractPrevalent semantic segmentation methods generally adopt a vanilla classifier to categorize each pixel into specific classes. Although such a classifier learns global information from the training data, this information is represented by a set of fixed parameters (weights and biases). However, each image has a different class distribution, which prevents the classifier from addressing the unique characteristics of individual images. At the dataset level, class imbalance leads to segmentation results being biased towards majority classes, limiting the model's effectiveness in identifying and segmenting minority class regions. In this paper, we propose an Extended Context-Aware Classifier (ECAC) that dynamically adjusts the classifier using global (dataset-level) and local (image-level) contextual information. Specifically, we leverage a memory bank to learn dataset-level contextual information of each class, incorporating the class-specific contextual information from the current image to improve the classifier for precise pixel labeling. Additionally, a teacher-student network paradigm is adopted, where the domain expert (teacher network) dynamically adjusts contextual information with ground truth and transfers knowledge to the student network. Comprehensive experiments illustrate that the proposed ECAC can achieve state-of-the-art performance across several datasets, including ADE20K, COCO-Stuff10K, and Pascal-Context. Huadong Tang, Youpeng Zhao 0002, Min Xu 0001, Jun Wang 0001, Qiang Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | Introduction to the Special Issue on Reliable Infrastructure and Edge Analytics for IoT
Wenqi Wei 0001, Balaji Palanisamy, Jun Wang 0001 |
ACM Trans. Internet Techn. | 3 |
| 2025 | MeRino: Entropy-Driven Design for Generative Language Models on IoT DevicesabstractGenerative Large Language Models (LLMs) stand as a revolutionary advancement in the modern era of artificial intelligence (AI). However, scaling down LLMs for resource-constrained hardware, such as Internet-of-Things (IoT) devices requires non-trivial efforts and domain knowledge. In this paper, we propose a novel information-entropy framework for designing mobile-friendly generative language models. The whole design procedure involves solving a mathematical programming (MP) problem, which can be done on the CPU within minutes, making it nearly zero-cost. We evaluate our designed models, termed MeRino, across fourteen NLP downstream tasks, showing their competitive performance against the state-of-the-art autoregressive transformer models under the mobile setting. Notably, MeRino achieves similar or better performance on both language modeling and zero-shot learning tasks, compared to the 350M parameter OPT while being 4.9x faster on NVIDIA Jetson Nano with 5.5x reduction in model size. Youpeng Zhao 0002, Huadong Tang, Qiang Wu 0001, Jun Wang 0001 |
AAAI | 5 |
| 2025 | AIRES: Accelerating Out-of-Core GCNs via Algorithm-System Co-DesignabstractGraph convolutional networks (GCNs) are fundamental in various scientific applications, ranging from biomedical protein-protein interactions (PPI) to large-scale recommendation systems. An essential component for modeling graph structures in GCNs is sparse general matrix-matrix multiplication (SpGEMM). As the size of graph data continues to scale up, SpGEMMs are often conducted in an out-of-core fashion due to limited GPU memory space in resource-constrained systems. Albeit recent efforts that aim to alleviate the memory constraints of out-of-core SpGEMM through either GPU feature caching, hybrid CPU-GPU memory layout, or performing the computation in sparse format, current systems suffer from both high I/O latency and GPU under-utilization issues. In this paper, we first identify the problems of existing systems, where sparse format data alignment and memory allocation are the main performance bottlenecks, and propose AIRES, a novel algorithm-system co-design solution to accelerate out-of-core SpGEMM computation for GCNs. Specifically, from the algorithm angle, AIRES proposes to alleviate the data alignment issues on the block level for matrices in sparse formats and develops a tiling algorithm to facilitate row block-wise alignment. On the system level, AIRES employs a three-phase dynamic scheduling that features a dual-way data transfer strategy utilizing a tiered memory system: integrating GPU memory, GPU Direct Storage (GDS), and host memory to reduce I/O latency and improve throughput. Evaluations show that AIRES significantly outperforms the state-of-the-art methods, achieving up to${1. 8} \times$lower latency in real-world graph processing benchmarks. Shakya Jayakody, Youpeng Zhao 0002, Jun Wang 0001 |
ASAP | 3 |
| 2025 | MHDiff: Memory- and Hardware-Efficient Diffusion Acceleration via Focal Pixel Aware QuantizationabstractDiffusion models have demonstrated superior performance in image generation tasks, thus becoming the mainstream model for generative visual tasks. Diffusion models need to execute multiple timesteps sequentially, resulting in a dramatic increase in workload. Existing accelerators leverage the data similarity between adjacent timesteps and perform mixed-precision differential quantization to accelerate diffusion models. However, merging differential values with raw inputs in each layer of each timestep to ensure computational correctness requires significant memory access for loading raw inputs, which creates a heavy memory burden. Moreover, mixed-precision computations may lead to low hardware utilization if not well designed. Unlike these works, we propose MHDiff, a tailored framework that identifies the focal pixels at the first layer and finetunes them to fit all layers, then represents focal pixels with high-precision while using low-precision for others, thereby accelerating diffusion models while minimizing memory burden. To improve hardware utilization, MHDiff employs a packing module that merges low-precision values into high-precision values to create full high-precision matrices and designs a processing element (PE) array to efficiently process the packed matrices. Extensive experiment results demonstrate that MHDiff can achieve satisfactory performance with negligible quality loss. Chunyu Qi, Xuhang Wang, Yuanzheng Yao, Naifeng Jing, Chen Zhang 0001, Jun Wang 0001, Zhihui Fu, Xiaoyao Liang, Zhuoran Song |
DAC | 7 |
| 2025 | SynGPU: Synergizing CUDA and Bit-Serial Tensor Cores for Vision Transformer Acceleration on GPUabstractVision Transformers (ViTs) have demonstrated remarkable performance in computer vision tasks by effectively extracting global features. However, their self-attention mechanism suffers from quadratic time and memory complexity as image resolution or video duration increases, leading to inefficiency on GPUs. To accelerate ViTs, existing works mainly focus on pruning tokens based on value-level sparsity. However, they miss the chance to achieve peak performance as they overlook the bit-level sparsity. Instead, we propose Inter-token Bit-sparsity Awareness (IBA) algorithm to accelerate ViTs by exploring bit-sparsity from similar tokens. Next, we implement IBA on GPUs that synergize CUDA and Tensor Cores by addressing two issues: firstly, the bandwidth congestion of the Register File hinders the parallel ability of CUDA and Tensor Cores. Secondly, due to the varying exponent of floating-point vectors, it is hard to accelerate bitsparse matrix multiplication and accumulation (MMA) in Tensor Core through fixed-point-based bit-level circuits. Therefore, we present SynGPU, an algorithm-hardware co-design framework, to accelerate ViTs. SynGPU enhances data reuse by a novel data mapping to enable full parallelism of CUDA and Tensor Cores. Moreover, it introduces Bit-Serial Tensor Core (BSTC) that supports fixed- and floating-point MMA by combining the fixedpoint Bit-Serial Dot Product (BSDP) and exponent alignment techniques. Extensive experiments show that SynGPU achieves an average of $2.15 \times \sim 3.95 \times$ speedup and $2.49 \times \sim 3.81 \times$ compute density over A100 GPU. Yuanzheng Yao, Chen Zhang 0001, Chunyu Qi, Jun Wang 0001, Zhihui Fu, Naifeng Jing, Xiaoyao Liang, Zhuoran Song |
DAC | 5 |
| 2025 | OLearning: A Geo-Distributed System for Device-Cloud Collaborative Computing
Zhihui Fu, Xiangmou Qu, Ruiguang Pei, Jun Wang 0001 |
DASFAA (6) | 5 |
| 2025 | SimDC: A High-Fidelity Device Simulation Platform for Device-Cloud Collaborative ComputingabstractThe advent of edge intelligence and escalating concerns for data privacy protection have sparked a surge of interest in device-cloud collaborative computing. Large-scale device deployments to validate prototype solutions are often prohibitively expensive and practically challenging, resulting in a pronounced demand for simulation tools that can emulate real-world scenarios. However, existing simulators predominantly rely solely on high-performance servers to emulate edge computing devices, overlooking (1) the discrepancies between virtual computing units and actual heterogeneous computing devices and (2) the simulation of device behaviors in real-world environments. In this paper, we propose a high-fidelity device simulation platform, called SimDC, which uses a hybrid heterogeneous resource and integrates high-performance servers and physical mobile phones. Utilizing this platform, developers can simulate numerous devices for functional testing cost-effectively and capture precise operational responses from varied real devices. To simulate real behaviors of heterogeneous devices, we offer a configurable device behavior traffic controller that dispatches results on devices to the cloud using a user-defined operation strategy. Comprehensive experiments on the public dataset show the effectiveness of our simulation platform and its great potential for application.1 Ruiguang Pei, Dan Peng, Zhihui Fu, Jun Wang 0001 |
ICDCS | 7 |
| 2024 | ALISE: Accelerating Large Language Model Serving with Speculative SchedulingabstractLarge Language Models (LLMs) represent a revolutionary advancement in the contemporary landscape of artificial general intelligence (AGI). As exemplified by ChatGPT, LLM-based applications necessitate minimal response latency and maximal throughput for inference serving. However, due to the unpredictability of LLM execution, the first-come-first-serve (FCFS) scheduling policy employed by current LLM serving systems suffers from head-of-line (HoL) blocking issues and long job response times. Youpeng Zhao 0002, Jun Wang 0001 |
ICCAD | 2 |
| 2024 | DFD: Distilling the Feature Disparity Differently for DetectorsabstractKnowledge distillation is a widely adopted model compression technique that has been successfully applied to object detection. In feature distillation, it is common practice for the student model to imitate the feature responses of the teacher model, with the underlying objective of improving its own abilities by reducing the disparity with the teacher. However, it is crucial to recognize that the disparities between the student and teacher are inconsistent, highlighting their varying abilities. In this paper, we explore the inconsistency in the disparity between teacher and student feature maps and analyze their impact on the efficiency of the distillation. We find that regions with varying degrees of difference should be treated separately, with different distillation constraints applied accordingly. We introduce our distillation method called Disparity Feature Distillation(DFD). The core idea behind DFD is to apply different treatments to regions with varying learning difficulties, simultaneously incorporating leniency and strictness. It enables the student to better assimilate the teacher’s knowledge. Through extensive experiments, we demonstrate the effectiveness of our proposed DFD in achieving significant improvements. For instance, when applied to detectors based on ResNet50 such as RetinaNet, FasterRCNN, and RepPoints, our method enhances their mAP from 37.4%, 38.4%, 38.6% to 41.7%, 42.4%, 42.7%, respectively. Our approach also demonstrates substantial improvements on YOLO and ViT-based models. The code is available at https://github.com/luckin99/DFD. Jinmin Li, Jun Wang 0001, Shaoming Wang, Chun Yuan 0003, Rizen Guo |
ICML | 5 |
| 2024 | ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV CachingabstractThe Transformer architecture has significantly advanced natural language processing (NLP) and has been foundational in developing large language models (LLMs) such as LLaMA and OPT, which have come to dominate a broad range of NLP tasks. Despite their superior accuracy, LLMs present unique challenges in practical inference, concerning the compute and memory-intensive nature. Thanks to the autoregressive characteristic of LLM inference, KV caching for the attention layers in Transformers can effectively accelerate LLM inference by substituting quadratic-complexity computation with linear-complexity memory accesses. Yet, this approach requires increasing memory as demand grows for processing longer sequences. The overhead leads to reduced throughput due to I/O bottlenecks and even out-of-memory errors, particularly on resource-constrained systems like a single commodity GPU. In this paper, we propose ALISA, a novel algorithm-system co-design solution to address the challenges imposed by KV caching. On the algorithm level, ALISA prioritizes tokens that are most important in generating a new token via a Sparse Window Attention (SWA) algorithm. SWA introduces high sparsity in attention layers and reduces the memory footprint of KV caching at negligible accuracy loss. On the system level, ALISA employs three-phase token-level dynamical scheduling and optimizes the trade-off between caching and recomputation, thus maximizing the overall performance in resource-constrained systems. In a single GPU-CPU system, we demonstrate that under varying workloads, ALISA improves the throughput of baseline systems such as FlexGen and vLLM by up to $3 \times$ and $1.9 \times$, respectively. Youpeng Zhao 0002, Di Wu 0016, Jun Wang 0001 |
ISCA | 3 |
| 2024 | Correction to: AA-forecast: anomaly-aware forecast for extreme events
Ashkan Farhangi, Jiang Bian 0003, Arthur Huang, Haoyi Xiong, Jun Wang 0001, Zhishan Guo |
Data Min. Knowl. Discov. | 5 |
| 2023 | Exploring Architecture, Dataflow, and Sparsity for GCN Accelerators: A Holistic FrameworkabstractRecent years have seen an increasing number of Graph Convolutional Network (GCN) models employed in various real-world applications. However, designing efficient architectures for GCN acceleration remains challenging due to the varied sparsity across graph datasets. Despite significant efforts, very few of the existing works have considered a holistic view of the entire GCN accelerator design, and therefore, the dynamic interactions between architecture, dataflow (i.e., data reuse and parallelization strategies), and compression format are not well studied Lingxiang Yin, Jun Wang 0001, Hao Zheng 0005 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | HCPerf: Driving Performance-Directed Hierarchical Coordination for Autonomous VehiclesabstractThe rapid development of autonomous driving poses new research challenges to the on-vehicle computing system. In particular, the execution time of autonomous driving tasks highly depends on the specific driving environment. For instance, the execution time of configurable sensor fusion increases significantly as the scene becomes complex, which leads to end-to-end deadline misses from sensing to control and may cause accidents. Thus, a framework that can effectively utilize the system resources to guarantee the end-to-end deadlines of autonomous driving tasks as well as effectively prioritize the responsiveness and throughput of the control commands is crucial for autonomous driving. In this paper, we propose HCPerf, a performance-directed hierarchical coordination framework that intelligently coordinates the autonomous driving tasks with high execution time variation and complex dependencies according to the driving performance in real-time. Specifically, HCPerf mainly consists of two coordinators. The internal coordinator intelligently schedules the tasks according to the driving performance of the vehicle in order to help them meet the end-to-end deadlines while well prioritizing the responsiveness and throughput of the control commands. At the same time, the external coordinator dynamically tunes the rates of tasks according to the schedulability in order to efficiently utilize the system resource. We conduct extensive experiments on both simulation and hardware testbeds with the representative autonomous driving application. The results show that HCPerf can effectively improve the driving performance by 7.69%-45.94% in different driving scenarios. Jialiang Ma, Li Li 0064, Zejiang Wang, Jun Wang 0001, Cheng-Zhong Xu 0001 |
ICDCS | 4 |
| 2023 | Parameter-Efficient Vision Transformer with Linear AttentionabstractRecent advances in vision transformers (ViTs) have achieved outstanding performance in visual recognition tasks, including image classification and detection. ViTs can learn global representations with their self-attention mechanism, but they are usually heavy-weight and unsuitable for resource-constrained devices. In this paper, we propose a novel linear feature attention (LFA) module to reduce computation costs for vision transformers and combine efficient mobile CNN modules to form a parameter-efficient and high-performance CNN-ViT hybrid model, called LightFormer, which can serve as a general-purpose backbone to learn both global and local representation. Comprehensive experiments demonstrate that LightFormer achieves competitive performance across different visual recognition tasks. On the ImageNet-1K dataset, LightFormer achieves top-1 accuracy of 78.5% with 5.5 million parameters. Our model also performs well when transferred to object detection and semantic segmentation tasks. On the MS COCO dataset, LightFormer attains mAP of 33.2 within the YOLOv3 framework, and on the Cityscapes dataset, with only a simple all-MLP decoder, LightFormer achieves mIoU of 78.5 and FPS of 15.3, surpassing state-of-the-art lightweight segmentation networks. Youpeng Zhao 0002, Huadong Tang, Yingying Jiang 0001, Yong A, Qiang Wu 0001, Jun Wang 0001 |
ICIP | 6 |
| 2023 | AA-forecast: anomaly-aware forecast for extreme events
Ashkan Farhangi, Jiang Bian 0003, Arthur Huang, Haoyi Xiong, Jun Wang 0001, Zhishan Guo |
Data Min. Knowl. Discov. | 5 |
| 2022 | HARMONY: Heterogeneity-Aware Hierarchical Management for Federated Learning SystemabstractFederated learning (FL) enables multiple devices to collaboratively train a shared model while preserving data privacy. However, despite its emerging applications in many areas, real-world deployment of on-device FL is challenging due to wildly diverse training capability and data distribution across heterogeneous edge devices, which highly impact both model performance and training efficiency. This paper proposes Harmony, a high-performance FL framework with heterogeneity-aware hierarchical management of training devices and training data. Unlike previous work that mainly focuses on heterogeneity in either training capability or data distribution, Harmony adopts a hierarchical structure to jointly handle both heterogeneities in a unified manner. Specifically, the two core components of Harmony are a global coordinator hosted by the central server and a local coordinator deployed on each participating device. Without accessing the raw data, the global coordinator first selects the participants, and then further reorganizes their training samples based on the accurate estimation of the runtime training capability and data distribution of each device. The local coordinator keeps monitoring the local training status and conducts efficient training with guidance from the global coordinator. We conduct extensive experiments to evaluate Harmony using both hardware and simulation testbeds on representative datasets. The experimental results show that Harmony improves the accuracy performance by 1.67% - 27.62%. In addition, Harmony effectively accelerates the training process up to $3.29\times$ and $1.84\times$ on average, and saves energy up to 88.41% and 28.04% on average. Chunlin Tian, Li Li 0064, Jun Wang 0001, Cheng-Zhong Xu 0001 |
MICRO | 4 |
| 2022 | Machine Learning in Real-Time Internet of Things (IoT) Systems: A SurveyabstractOver the last decade, machine learning (ML) and deep learning (DL) algorithms have significantly evolved and been employed in diverse applications, such as computer vision, natural language processing, automated speech recognition, etc. Real-time safety-critical embedded and Internet of Things (IoT) systems, such as autonomous driving systems, UAVs, drones, security robots, etc., heavily rely on ML/DL-based technologies, accelerated with the improvement of hardware technologies. The cost of a deadline (required time constraint) missed by ML/DL algorithms would be catastrophic in these safety-critical systems. However, ML/DL algorithm-based applications have more concerns about accuracy than strict time requirements. Accordingly, researchers from the real-time systems (RTSs) community address the strict timing requirements of ML/DL technologies to include in RTSs. This article will rigorously explore the state-of-the-art results emphasizing the strengths and weaknesses in ML/DL-based scheduling techniques, accuracy versus execution time tradeoff policies of ML algorithms, and security and privacy of learning-based algorithms in real-time IoT systems. Jiang Bian 0003, Abdullah Al Arafat, Haoyi Xiong, Jing Li 0025, Li Li 0064, Hongyang Chen 0001, Jun Wang 0001, Dejing Dou, Zhishan Guo |
IEEE Internet Things J. | 7 |
| 2022 | Mixed-Criticality Scheduling Upon Permitted Failure Probability and Dynamic PriorityabstractMany safety-critical real-time systems are considered certified when they meet failure probability requirements with respect to the maximum permitted incidences of failure per hour. In this article, the mixed-criticality task model with multiple worst case execution time (WCET) estimations is extended to incorporate such system-level certification restrictions. A new parameter is added to each task, characterizing the distribution of WCET estimations—the likelihood of all jobs of a task finishing their executions within the less pessimistic WCET estimates. Efficient algorithms are derived for scheduling mixed-criticality systems represented using this model for both uniprocessor and multiprocessor platforms for independent tasks. Furthermore, a 0/1 covariance matrix is introduced to represent the failure dependency between tasks. An efficient algorithm is proposed to schedule such failure-dependent tasks. Experimental analyses show our new model and algorithm outperform current state-of-the-art mixed-criticality scheduling algorithms. Zhishan Guo, Sudharsan Vaidhun, Luca Satinelli, Samsil Arefin, Jun Wang 0001, Kecheng Yang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Enhancing Proportional IO Sharing on Containerized Big Data File SystemsabstractBig Data platforms recently employ resource management systems, such as YARN, Mesos, and Google Borg, to provision computational resources. These systems adopt containerization to share the computing resources in a multi-tenant setting with low performance overhead and interference. However, it may be observed that tenants often interfere with each other on the underlying Big Data File Systems (BDFS), e.g., Hadoop File System, which have been widely deployed as a persistent layer in current data centers. A solution with systematic generality is to containerize BDFS itself to isolate and allocate its IO sources to multiple tenants. To this end, we conduct analysis on the ineffectiveness of proportionally sharing BDFS IO resource via containerization. This ineffectiveness is due to the scheduler of containerization in “pseudo-starvation” status, in which most of IO requests are backlogged in BDFS rather than in containerization scheduler. Without enough backlogged IO requests, existing schedulers might have to maximize device utilization rather than enforce proportional sharing policy. To resolve this ineffectiveness issue, we develop a cross-layer system calledBDFS-Container, which containerizes BDFS at the Linux block IO level. Central to BDFS-Container, we propose and design a proactive IOPS throttling-based mechanism namedIOPS Regulator, which achieves a trade-off between maximizing IO utilization and accurately proportional IO sharing. The evaluation results show that our method can improve proportionally sharing BDFS IO resources by 74.4 percent on average. Dan Huang 0001, Jun Wang 0001, Qing Liu 0002, Nong Xiao 0001, Huafeng Wu, Jiangling Yin |
IEEE Trans. Computers | 2 |
| 2021 | Overlapping Communication With Computation in Parameter Server for Scalable DL TrainingabstractScalability of distributed deep learning (DL) training with parameter server (PS) architecture is often communication constrained in large clusters. There are recent efforts that use a layer by layer strategy to overlap gradient communication with backward computation so as to reduce the impact of communication constraint on the scalability. However, the approaches could bring significant overhead in gradient communication. Meanwhile, they cannot be effectively applied to the overlap between parameter communication and forward computation. In this article, we propose and develop iPart, a novel approach that partitions communication and computation in various partition sizes to overlap gradient communication with backward computation and parameter communication with forward computation. iPart formulates the partitioning decision as an optimization problem and solves it based on a greedy algorithm to derive communication and computation partitions. We implement iPart in the open-source DL framework BigDL and perform evaluations with various DL workloads. Experimental results show that iPart improves the scalability of a cluster of 72 nodes by up to 94 percent over the default PS and 52 percent over the layer by layer strategy. Aidi Pi, Xiaobo Zhou 0002, Jun Wang 0001, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Disperse Access Considered Energy Inefficiency in Intel Optane DC Persistent Memory ServersabstractThe Intel Optane DC Persistent Memory Module (AEP), which is the first commercial available Non-Volatile Memory (NVM) product, offers comparable performance with DRAM while providing larger capacities and data persistence. Existing researches that substitute NVM with DRAM or hybridize them are either emulator-based or focused on how to improve the energy efficiency for writes. Unfortunately, the energy efficiency of the real AEP system is less explored. Based on real AEP, we observe that even though eliminating the DRAM-like refresh energy consumptions, AEP consumes significant different energy at different performance levels. Specifically, requests with time intervals (dispersed) underperform in both performance and energy efficiency when compared with the case of requests without time intervals (compact). This disparity and parallelism exploitation potentials motivate us to propose Sprint-AEP, an energy-efficiency-oriented scheduling method for AEP-equipped servers. Sprint-AEP fully activates adequate AEPs to serve most of the requests by deferring the write requests and prefetching the hottest data. The remaining AEPs will stay in idle mode with a low idle power to save energy. Besides, we also utilize the read parallelism to accelerate the sync and prefetching processes. Compared with energy-unaware AEP usages, our experimental results show that Sprint-AEP saves up to 26% energy with little performance degradation. Daping Li, Jiguang Wan 0001, Jun Wang 0001, Jian Zhou 0004, Kai Lu 0002, Fei Wu 0005, Changsheng Xie 0001 |
ICDCS | 3 |
| 2020 | Lelantus: Fine-Granularity Copy-On-Write Operations for Secure Non-Volatile MemoriesabstractBulk operations, such as Copy-on-Write (CoW), have been heavily used in most operating systems. In particular, CoW brings in significant savings in memory space and improvement in performance. CoW mainly relies on the fact that many allocated virtual pages are not written immediately (if ever written). Thus, assigning them to a shared physical page can eliminate much of the copy/initialization overheads in addition to improving the memory space efficiency. By prohibiting writes to the shared page, and merely copying the page content to a new physical page at the first write, CoW achieves significant performance and memory space advantages. Unfortunately, with the limited write bandwidth and slow writes of emerging Non-Volatile Memories (NVMs), such bulk writes can throttle the memory system. Moreover, it can add significant delays on the first write access to each page due to the need to copy or initialize a new page. Ideally, we need to enable CoW at fine-granularity, and hence only the updated cache blocks within the page need to be copied. To do this, we propose Lelantus, a novel approach that leverages secure memory metadata to allow fine-granularity CoW operations. Lelantus relies on a novel hardware-software co-design to allow tracking updated blocks of copied pages and hence delay the copy of the rest of the blocks until written. The impact of Lelantus becomes more significant when huge pages are deployed, e.g., 2MB or 1GB, as expected with emerging NVMs. Jian Zhou 0004, Amro Awad, Jun Wang 0001 |
ISCA | 3 |
| 2020 | Multi-layer Coordination for High-Performance Energy-Efficient Federated LearningabstractFederated Learning is designed for multiple mobile devices to collaboratively train an artificial intelligence model while preserving data privacy. Instead of collecting the raw training data from mobile devices to the cloud, Federated Learning coordinates a group of devices to train a shared model in a distributed manner with the training data located on the devices. However, in order to effectively deploy Federated Learning on resource-constrained mobile devices, several critical issues including convergence rate, scalability and energy efficiency should be well addressed. In this paper, we propose MCFL, a multi-layer online coordination framework for high-performance energy efficient federated learning. MCFL consists of two layers: a macro-layer on the central server and a micro-layer on each participating device. In each training round, the macro coordinator performs two tasks, namely, selecting the right devices to participate, and estimating a time limit, such that the overall training time is significantly reduced while still guaranteeing the model accuracy. Unlike existing systems, MCFL removes the restriction that participating devices must be connected to power sources, thus allowing more timely and ubiquitous training. This clearly requires on-device training to be highly energy-efficient. To this end, the micro coordinator determines optimal schedules for hardware resources in order to meet the time limit set by the macro coordinator with the least amount of energy consumption. Tested on real devices as well as simulation testbed, MCFL has shown to be able to effectively balance the convergence rate, model accuracy and energy efficiency. Compared with existing systems, MCFL can achieve a speedup up to 8.66× and reduce energy consumption by up to 76.5% during the training process. Li Li 0064, Jun Wang 0001, Cheng-Zhong Xu 0001 |
IWQoS | 2 |
| 2020 | NMTLAT: A New robust mobile Multi-Target Localization and Tracking Scheme in marine search and rescue wireless sensor networks under Byzantine attack
Jiangfeng Xian, Huafeng Wu, Xiaojun Mei, Yuanyuan Zhang 0015, Huixing Chen, Jun Wang 0001 |
Comput. Commun. | 6 |
| 2020 | ODDS: Optimizing Data-Locality Access for Scientific Data AnalysisabstractWhereas traditional scientific applications are computationally intensive, recent applications require more data-intensive analysis and visualization to extract knowledge from the explosive growth of scientific information and simulation data. As the computational power and size of compute clusters continue to increase, the I/O read rates and associated network for these data-intensive applications have been unable to keep pace. These applications suffer from long I/O latency due to the movement of “big data” from the network/parallel file system, which results in a serious performance bottleneck. To address this problem, we proposed a novel approach called “ODDS” to optimize data-locality access in scientific data analysis and visualization. ODDS leverages a distributed file system (DFS) to provide scalable data access for scientific analysis. Through exploiting the information of underlying data distribution in DFS, ODDS employs a novel data-locality scheduler to transform a compute-centric mapping into a data-centric one and enables each computational process to access the needed data from a local or nearby storage node. ODDS is suitable for parallel applications with dynamic process-to-data scheduling and for applications with static process-to-data assignment. To demonstrate the efficacy of our methods, we present and evaluate ODDS in the context of two state-of-the-art, scientific-analysis applications-mpiBLAST and ParaView-along with the Hadoop distributed file system (HDFS) across a wide variety of computing platform settings. In comparison to existing deployments using NFS, PVFS, or Lustre as the underlying storage systems, ODDS can greatly reduce the I/O cost and double overall performance. Jun Wang 0001, Dezhi Han, Jiangling Yin, Xiaobo Zhou 0002, Changjun Jiang 0002 |
IEEE Trans. Cloud Comput. | 1 |
| 2019 | Work-in-Progress: A Deep Learning Strategy for I/O Scheduling in Storage SystemsabstractUnder the big data era, there is a crucial need to improve the performance of storage systems for data-intensive applications. Data-intensive applications tend to behave in a predictable manner, which can be exploited for improving the performance of the storage system. At the storage level, we propose a deep recurrent neural network that learns the patterns of I/O requests and predicts the upcoming ones, such that memory contents can be pre-loaded at the right time to prevent cache/memory misses. Preliminary experimental results, on two real-world I/O logs of storage systems (from financial and web search), are reported-they partially demonstrate the effectiveness of the proposed method. Ashkan Farhangi, Jiang Bian 0003, Jun Wang 0001, Zhishan Guo |
RTSS | 3 |
| 2019 | SmartPC: Hierarchical Pace Control in Real-Time Federated Learning SystemabstractFederated Learning is a technique for learning AI models through the collaboration of a large number of resourceconstrained mobile devices, while preserving data privacy. Instead of aggregating the training data from devices, Federated Learning uses multiple rounds of parameter aggregation to train a model, wherein the participating devices are coordinated to incrementally update a shared model with their own parameters locally learned. To efficiently deploy Federated Learning system over mobile devices, several critical issues including realtimeliness and energy efficiency should be well addressed. This paper proposes SmartPC, a hierarchical online pace control framework for Federated Learning that balances the training time and model accuracy in an energy-efficient manner. SmartPC consists of two layers of pace control: global and local. Prior to every training round, the global controller first oversees the status (e.g., connectivity, availability, and energy/resource remained) of every participating device, then selects qualified devices and assigns them a well-estimated virtual deadline for task completion. Within such virtual deadline, a statistically significant proportion (e.g., 60%) of the devices are expected to complete one round of their local training and model updates, while the overall progress of multi-round training procedure is kept up adaptively. On each device, a local pace controller then dynamically adjusts device settings such as CPU frequency so that the learning task is able to meet the deadline with the least amount of energy consumption. We performed extensive experiments to evaluate SmartPC on both Android smartphones and simulation platforms using well-known datasets. The experiment results show that SmartPC reduces up to 32:8% energy consumption on mobile devices and achieves a speedup of 2.27 in training time without model accuracy degradation. Li Li 0064, Haoyi Xiong, Zhishan Guo, Jun Wang 0001, Cheng-Zhong Xu 0001 |
RTSS | 4 |
| 2019 | RangingNet: A convolutional deep neural network based ranging model for wireless sensor networks (WSN)
Huafeng Wu, Weijun Wang 0006, Jun Wang 0001, Prasant Mohapatra |
Comput. Commun. | 3 |
| 2019 | Efficient target detection in maritime search and rescue wireless sensor network using data fusion
Huafeng Wu, Jiangfeng Xian, Xiaojun Mei, Yuanyuan Zhang 0015, Jun Wang 0001, Junkuo Cao, Prasant Mohapatra |
Comput. Commun. | 5 |
| 2019 | An Efficient and Safe Road Condition Monitoring Authentication Scheme Based on Fog ComputingabstractIn recent years, with the development of intelligent vehicles and wireless sensor network technology, the research on road safety has attracted much attention in vehicular ad-hoc networks (VANETs). By sensing events on the road, vehicles can broadcast information to inform others of traffic jams or accidents. However, the mobile vehicle network has a large transmission delay, which makes real-time content transmission impossible. In this paper, a new certificateless aggregate signcryption scheme (CLASC) is proposed by using a fog computing framework that supports mobility, low latency, and location awareness. It is combined with online/offline encryption (OOE) technology, which reduces many time-consuming operations and improves the security of vehicle users and the efficiency of message authentication. In addition, the scheme has the characteristics of mutual authentication, anonymity, untraceability, and nondeniability. Based on the difficulty of the discrete logarithm problem (DLP) and the computational Diffie-Hellman (CDH) problem, the scheme is further proved to be unforgeability and confidentiality under the random oracle model. The simulation results show that compared with the existing schemes, this scheme can not only ensure the security requirements of the system but also achieve higher efficiency in computing and communication. Mingming Cui, Dezhi Han, Jun Wang 0001 |
IEEE Internet Things J. | 3 |
| 2019 | ApproxSSD: Data Layout Aware Sampling on an Array of SSDsabstractExecution of analytic frameworks on sample data sets is the current trend in response to increasing data size and demand for real-time analysis. Additionally, high-performance, energy-efficient Solid-State Drive (SSD) arrays are the primary storage subsystem for parallel data analysis systems. To exploit the benefits of SSD arrays when executing sample data set analytics, several key areas must be considered. First, due to logical to physical address translation, random data choice in data sampling jobs can cause unbalanced workloads among SSDs in the array. Second, after the data choice, existing task schedulers in data analysis frameworks can introduce non-negligible resource contentions resulting from the suboptimal Input/Output (I/O). The performance of SSDs is unpredictable because of their varying maintenance costs at runtime, which renders them hard to be managed by the scheduler. With the trend towards sample set data analytics and the use of SSDs, it is increasingly important to ensure balanced workloads and minimize resource contention. Without addressing these areas, sample-set data analytics on SSDs will continue to suffer from performance inefficiencies. In this paper, we propose ApproxSSD to perform on-disk layout-aware data sampling on SSD arrays. This proposed framework leverages data selection and task scheduling to improve the performance of many applications. ApproxSSD decouples I/O from the computation in task execution. This avoids potential I/O contentions and suboptimal workload balances. We have developed an open-source prototype system of ApproxSSD in Scala at Github. Our evaluation shows that ApproxSSD can achieve up to 2.7 times speed up at 10 percent sampling ratio under an example sampling workload when compared to Spark, while simultaneously maintaining high output accuracy. Jian Zhou 0004, Huafeng Wu, Jun Wang 0001 |
IEEE Trans. Computers | 3 |
| 2019 | Harnessing Data Movement in Virtual Clusters for In-Situ ExecutionabstractAs a result of increasing data volume and velocity, Big Data science at exascale has shifted towards the in-situ paradigm, where large scale simulations run concurrently alongside data analytics. With in-situ, data generated from simulations can be processed while still in memory, thereby avoiding the slow storage bottleneck. However, running simulations and analytics together on shared resources will likely result in substantial contention if left unmanaged, as demonstrated in this work, leading to much reduced efficiency of simulations and analytics. Recently, virtualization technologies such as Linux containers have been widely applied to data centers and physical clusters to provide highly efficient and elastic resource provisioning for consolidated workloads including scientific simulations and data analytics. In this paper, we investigate to facilitate network traffic manipulation and reduce mutual interference on the network for in-situ applications in virtual clusters. In order to dynamically allocate the network bandwidth when it is needed, we adopt SARIMA-based techniques to analyze and predict MPI traffic issued from simulations. Although this can be an effective technique, the naïve usage of network virtualization can lead to performance degradation for bursty asynchronous transmissions within an MPI job. We analyze and resolve this performance degradation in virtual clusters. Dan Huang 0001, Qing Liu 0002, Scott Klasky, Jun Wang 0001, Jong Choi 0001, Jeremy Logan, Norbert Podhorszki |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | An I/O Efficient Distributed Approximation Framework Using Cluster SamplingabstractIn this paper, we present an I/O efficient distributed approximation framework to support approximations on arbitrary sub-datasets of a large dataset. Due to the prohibitive storage overhead of caching offline samples for each sub-dataset, existing offline sample-based systems provide high accuracy results for only a limited number of sub-datasets, such as the popular ones. On the other hand, current online sample-based approximation systems, which generate samples at runtime, do not take into account the uneven storage distribution of a sub-dataset. They work well for uniform distribution of a sub-dataset while suffer low I/O efficiency and poor estimation accuracy on unevenly distributed sub-datasets. To address the problem, we develop a distribution aware method called CLAP (cluster sampling based approximation). Our idea is to collect the occurrences of a sub-dataset at each logical partition of a dataset (storage distribution) in the distributed system, and make good use of such information to enable I/O efficient online sampling. There are three thrusts in CLAP. First, we develop a probabilistic map to reduce the exponential number of recorded sub-datasets to a linear one. Second, we apply the cluster sampling with unequal probability theory to implement a distribution-aware method for efficient online sampling for a single or multiple sub-datasets. Third, we enrich CLAP support with more complex approximations such as ratio and regression using bootstrap based estimation beyond the simple aggragation approxiamtions. Forth, we add an option in CLAP to allow users specifying a target error bound when submitting an approximation job. Fifth, we quantitatively derive the optimal sampling unit size in a distributed file system by associating it with approximation costs and accuracy. We have implemented CLAP into Hadoop as an example system and open sourced it on GitHub. Our comprehensive experimental results show that CLAP can achieve a speedup by up to 20× over the precise execution. Xuhong Zhang 0002, Jun Wang 0001, Shouling Ji, Jiangling Yin, Rui Wang 0030, Xiaobo Zhou 0002, Changjun Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | A Correlation-Aware Page-Level FTL to Exploit Semantic Links in WorkloadsabstractNAND Flash based Solid State Disks (SSDs) are gaining tremendous popularity in today's storage market due to their unique erase-before-write feature. The Flash Translation Layer (FTL) in the SSDs redirects the incoming writes to a free physical address and manages a logical to physical address mapping table. However, this induces significant performance degradation to the SSDs. One of the main reasons is that current cache management in FTLs is mainly optimized for the temporal or spatial locality. However, because of multiple levels of data buffers in the whole storage architecture, the locality of internal disk I/O is relatively low. What's more, the increasing capacity of SSD not only generates large mapping tables, but also imposes high pressure on the efficiency of page-level address mapping. To overcome this limitation, we propose Correlation-Aware Page-level FTL, a.k.a CPFTL, which exploits I/O correlations in the workloads. In CPFTL, we develop a correlation-aware mapping table based on the correlation in read operations. We then build a correlation prediction table to support fast mapping entry lookup in the correlation-aware mapping table. Finally, we split read and write caches and build a skew-aware dirty entry index to improve the cache hit ratio and reduce the garbage collection overhead. Our emulator and prototype are open-sourced at: https://github.com/janzhou/SSD-Emulator. The experimental results show that CPFTL can reduce the average response time by 63.4 percent for read dominant workloads and 32.9 percent for transaction workloads. Jian Zhou 0004, Dezhi Han, Jun Wang 0001, Xiaobo Zhou 0002, Changjun Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | ArchSampler: Architecture-Aware Memory Sampling Library for In-Memory ApplicationsabstractWith the explosive rate of data growth, the limited scalability of the DRAM technology defies the performance potentials for in-memory applications. Fortunately, emerging non-volatile memory (NVM) technologies, such as Phase-Change Memory (PCM) and Memristor, are promising candidates for replacing DRAM. Emerging NVMs are very dense, hence promise large capacities. Additionally, NVMs are non-volatile, thus enable persistent applications and byte-addressable files. Both density and persistency are key enablers for in-memory applications. On the other side, emerging NVMs are slower than DRAM, thus optimizing for locality and avoiding contentions are key aspects to unlock the NVM performance. In this paper, we study the impact of memory contentions and architecture-oblivious implementations on the performance of sampling based in-memory approximation. Sampling has become an imperative technique used to accelerate big data processing, especially in today's emerging in-memory computing. However, we observe multiple times slow-down for nave and default implementations of in-memory data sampling. Accordingly, we propose ArchSampler, an architecture-aware sampling library. The main idea is to exploits the free choice of data samples to dynamically select which bank as a host to serve memory requests. Hence, ArchSampler enables efficient and high performing sampling through employing its knowledge of the NVM architectural details to maximize data locality and avoiding interthread contentions. Our evaluation shows that ArchSampler can achieve up to 1.62 speed up (1.20 on average) for different in-memory applications. Jian Zhou 0004, Jun Wang 0001 |
ICCD | 2 |
| 2018 | Performance Evaluation and Analysis for MPI-Based Data Movement in Virtual Switch NetworkabstractVirtualization technologies have been widely deployed in data centers and private clusters to provide highly efficient and elastic resource provisioning. Further, virtualization has been extended to the network layer, known as network virtualization. For example, independent virtual switches have become the primary provider of network services for various virtual machines, such as VMware, Xen and Docker. This approach allows the physical network to be decoupled from the overlying virtual switch networks. However, network virutalization introduces performance degradation and scalability bottleneck to communication-intensive frameworks, such as MPI. We quantify and analyze the performance degradation involved with collective communications as well as bursty asynchronous transmission (BAT) in vswitch network environments. Our experiments illustrate that the performance of MPI communication can be degraded up to 5× in the virtual environment. Dan Huang 0001, Jun Wang 0001, Dezhi Han |
NAS | 2 |
| 2018 | Missing data recovery using reconstruction in ocean wireless sensor networks
Huafeng Wu, Jiangfeng Xian, Jun Wang 0001, Siddhi Khandge, Prasant Mohapatra |
Comput. Commun. | 3 |
| 2018 | Prediction based opportunistic routing for maritime search and rescue wireless sensor network
Huafeng Wu, Jun Wang 0001, Raghavendra Rao Ananta, Vamsee Reddy Kommareddy, Rui Wang 0030, Prasant Mohapatra |
J. Parallel Distributed Comput. | 2 |
| 2018 | Speed Up Big Data Analytics by Unveiling the Storage Distribution of Sub-DatasetsabstractIn this paper, we study the problem of sub-dataset analysis over distributed file systems, e.g., the Hadoop file system. Our experiments show that the sub-datasets distribution over HDFS blocks, which is hidden by HDFS, can often cause corresponding analyses to suffer from a seriously imbalanced or inefficient parallel execution. Specifically, the content clustering of sub-datasets results in some computational nodes carrying out much more workload than others; furthermore, it leads to inefficient sampling of sub-datasets, as analysis programs will often read large amounts of irrelevant data. We conduct a comprehensive analysis on how imbalanced computing patterns and inefficient sampling occur. We then propose a storage distribution aware method to optimize sub-dataset analysis over distributed storage systems referred to as DataNet. First, we propose an efficient algorithm to obtain the meta-data of sub-dataset distributions. Second, we design an elastic storage structure called ElasticMap based on the HashMap and BloomFilter techniques to store the meta-data. Third, we employ distribution-aware algorithms for sub-dataset applications to achieve balanced and efficient parallel execution. Our proposed method can benefit different sub-dataset analyses with various computational requirements. Experiments are conducted on PRObEs Marmot 128-node cluster testbed and the results show the performance benefits of DataNet. Jun Wang 0001, Xuhong Zhang 0002, Jiangling Yin, Huafeng Wu, Dezhi Han |
IEEE Trans. Big Data | 1 |
| 2018 | Achieving Load Balance for Parallel Data Access on Distributed File SystemsabstractThe distributed file system, HDFS, is widely deployed as the bedrock for many parallel big data analysis. However, when running multiple parallel applications over the shared file system, the data requests from different processes/executors will unfortunately be served in a surprisingly imbalanced fashion on the distributed storage servers. These imbalanced access patterns among storage nodes are caused because a). unlike conventional parallel file system using striping policies to evenly distribute data among storage nodes, data-intensive file system such as HDFS store each data unit, referred to as chunk file, with several copies based on a relative random policy, which can result in an uneven data distribution among storage nodes; b). based on the data retrieval policy in HDFS, the more data a storage node contains, the higher probability the storage node could be selected to serve the data. Therefore, on the nodes serving multiple chunk files, the data requests from different processes/executors will compete for shared resources such as hard disk head and networkbandwidth, resulting in a degraded I/O performance. In this paper, we first conduct a complete analysis on how remote and imbalanced read/write patterns occur and how they are affected by the size of the cluster. We then propose novel methods, referred to as Opass, to optimize parallel data reads, as well as to reduce the imbalance of parallel writes on distributed file systems. Our proposed methods can benefit parallel data-intensive analysis with various parallel data access strategies. Opass adopts new matching-based algorithms to match processes to data so as to compute the maximum degree of data locality and balanced data access. Furthermore, to reduce the imbalance of parallel writes, Opass employs a heatmap for monitoring the I/O statuses of storage nodes and performs HM-LRU policy to select a local optimal storage node for serving write requests. Experiments are conducted on PRObE's Marmot 128-node cluster testbed and the results from both benchmark and well-known parallel applications show the performance benefits and scalability of Opass. Dan Huang 0001, Dezhi Han, Jun Wang 0001, Jiangling Yin, Xunchao Chen, Xuhong Zhang 0002, Jian Zhou 0004, Mao Ye 0008 |
IEEE Trans. Computers | 3 |
| 2018 | G-SD: Achieving Fast Reverse Lookup using Scalable Declustering Layout in Large-Scale File SystemsabstractWith the increasing popularity of cloud computing, current data centers contain petabytes of data in their datacenters. This requires thousands or tens of thousands of storage nodes at a single site. Node failure in these datacenters is normal instead of a rare situation. As a result, data reliability is a great concern. In order to achieve high reliability, data recovery or node reconstruction is a must. Although extensive research works have investigated how to sustain high performance and high reliability in case of node failure at large scale, a reverse lookup problem, namely finding the list of objects for the failed node is not well-addressed. As the first step of failure recovery, this process has a direct impact to the data recovery/node reconstruction. While existing solutions use metadata traversal or data distribution reversing methods for reverse lookup, which are either time consuming or expensive, the deterministic block placement schemes can achieve fast and efficient reverse lookup easily. However, they are designed for centralized, small-scale storage architectures such as RAID etc. Due to their lacking of scalability, they cannot be directly applied in large-scale storage systems. In this paper, we propose Group-Shifted Declustering (G-SD), a deterministic data layout for multi-way replication. G-SD addresses the scalability issue of our previous Shifted Declustering layout and supports fast and efficient reverse lookup. Our mathematical proofs demonstrate that G-SD is a scalable layout that maintains a high level of data availability. We implement a prototype of G-SD and its reverse lookup function on two open source file systems: Ceph and HDFS. Large scale experiments on the Marmot cluster demonstrate that the average speed of G-SD reverse lookup is more than 5× faster than the reverse lookup speed of existing schemes. Jun Wang 0001, Dezhi Han, Junyao Zhang 0007, Jiangling Yin |
IEEE Trans. Cloud Comput. | 1 |
| 2018 | A new rule-based power-aware job scheduler for supercomputers
Jun Wang 0001, Dezhi Han |
J. Supercomput. | 1 |
| 2018 | Workload Scheduling for Massive Storage Systems with Arbitrary Renewable SupplyabstractAs datacenters grow in scale, increasing energy costs and carbon emissions have led data centers to seek renewable energy, such as wind and solar energy. However, tackling the challenges associated with the intermittency and variability of renewable energy is difficult. This paper proposes a scheme called GreenMatch, which deploys an SSD cache to match green energy supplies with a time-shifting workload schedule while maintaining low latency for online data-intensive services. With the SSD cache, the process for a latency-sensitive request to access a disk is divided into two stages: a low-energy/low-latency online stage and a high-energy/high-latency off-line stage. As the process in the latter stage is off-line, it offers opportunities for time-shifting workload scheduling in response to variations of green energy supplies. We also allocate an HDD cache to guarantee data availability when renewable energy is inadequate. Furthermore, we design a novel replacement policy called Inactive P-disk First for the HDD cache to avoid inactive disk accesses. The experimental results show that GreenMatch can make full use of renewable energy while minimizing the negative impacts of intermittency and variability on performance and availability. Daping Li, Xiaoyang Qu, Jiguang Wan 0001, Jun Wang 0001, Xiaozhao Zhuang, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | DFS-container: achieving containerized block I/O for distributed file systemsabstractToday BigData systems commonly use resource management systems such as TORQUE, Mesos, and Google Borg to share the physical resources among users or applications. Enabled by virtualization, users can run their applications on the same node with low mutual interference. Container-based virtualizations (e.g., Docker and Linux Containers) offer a lightweight virtualization layer, which promises a near-native performance and is adopted by some Big-Data resource sharing platforms such as Mesos. Nevertheless, using containers to consolidate the I/O resources of shared storage systems is still at an early stage, especially in a distributed file system (DFS) such as Hadoop File System (HDFS). To overcome this issue, we propose a distributed middleware system, DFS-Container, by further containerizing DFS. We also evaluate and analyze the unfairness of using containers to proportionally allocate the I/O resource of DFS. Based on these analyses and evaluations, we propose and implement a new mechanism, IOPS-Regulator, which improve the fairness of proportional allocation by 74.4% on average. Dan Huang 0001, Jun Wang 0001, Qing Liu 0001, Xuhong Zhang 0002, Xunchao Chen, Jian Zhou 0004 |
SoCC | 2 |
| 2017 | SideIO: A Side I/O system framework for hybrid scientific workflow
Jun Wang 0001, Dan Huang 0001, Huafeng Wu, Jiangling Yin, Xuhong Zhang 0002, Xunchao Chen |
J. Parallel Distributed Comput. | 1 |
| 2017 | A new reliability model in replication-based big data storage systems
Jun Wang 0001, Huafeng Wu |
J. Parallel Distributed Comput. | 1 |
| 2017 | Deister: A light-weight autonomous block management in data-intensive file systems using deterministic declustering distribution
Jun Wang 0001, Xuhong Zhang 0002, Junyao Zhang 0007, Jiangling Yin, Dezhi Han, Dan Huang 0001 |
J. Parallel Distributed Comput. | 1 |
| 2017 | A reliable and energy-efficient storage system with erasure coding cacheabstractIn modern energy-saving replication storage systems, a primary group of disks is always powered up to serve incoming requests while other disks are often spun down to save energy during slack periods. However, since new writes cannot be immediately synchronized into all disks, system reliability is degraded. In this paper, we develop a high-reliability and energy-efficient replication storage system, named RERAID, based on RAID10. RERAID employs part of the free space in the primary disk group and uses erasure coding to construct a code cache at the front end to absorb new writes. Since code cache supports failure recovery of two or more disks by using erasure coding, RERAID guarantees a reliability comparable with that of the RAID10 storage system. In addition, we develop an algorithm, called erasure coding write (ECW), to buffer many small random writes into a few large writes, which are then written to the code cache in a parallel fashion sequentially to improve the write performance. Experimental results show that RERAID significantly improves write performance and saves more energy than existing solutions. Jiguang Wan 0001, Daping Li, Xiaoyang Qu, Jun Wang 0001, Changsheng Xie 0001 |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2017 | Energy-Aware Adaptive Restore Schemes for MLC STT-RAM CacheabstractFor the sake of higher cell density while achieving near-zero standby power, recent research progress in Magnetic Tunneling Junction (MTJ) devices has leveraged Multi-Level Cell (MLC) configurations of Spin-Transfer Torque Random Access Memory (STT-RAM). However, in orderto mitigate the write disturbance in an MLC strategy, data stored in the soft bit must be restored back immediately after the hard bit switching is completed. Furthermore, as the result of MTJ feature size scaling, the soft bit can be expected to become disturbed by the read sensing current, thus requiring an immediate restore operation to ensure the data reliability. In this paper, we design and analyze a novel Adaptive Restore Scheme for Write Disturbance (ARS-WD) and Read Disturbance (ARS-RD), respectively. ARS-WD alleviates restoration overhead by intentionally overwriting soft bit lines which are less likely to be read. ARS-RD, on the other hand, aggregates the potential writes and restore the soft bit line at the time of its eviction from higher level cache. Both of these two schemes are based on a lightweight forecasting approach for the future read behavior of the cache block. Our experimental results show substantial reduction in soft bit line restore operations, delivering 17.9 percent decrease in overall energy consumption and 9.4 percent increase in IPC, while incurring negligible capacity overhead. Moreover, ARS promotes advantages of MLC to provide a preferable L2 design alternative in terms of energy, area and latency product compared to SLC STT-RAM alternatives. Xunchao Chen, Navid Khoshavi, Ronald F. DeMara, Jun Wang 0001, Dan Huang 0001, Wujie Wen, Yiran Chen 0001 |
IEEE Trans. Computers | 4 |
| 2016 | Survey of data intensive computing technologies application to to security log data managementabstractData intensive computing research and technology developments offer the potential of providing significant improvements in several security log management challenges. Approaches to address the complexity, timeliness, expense, diversity, and noise issues have been identified. These improvements are motivated by the increasingly important role of analytics. Machine learning and expert systems that incorporate attack patterns are providing greater detection insights. Finding actionable indicators requires the analysis to combine security event log data with other network data such and access control lists, making the big-data problem even bigger. Automation of threat intelligence is recognized as not complete with limited adoption of standards. With limited progress in anomaly signature detection, movement towards using expert systems has been identified as the path forward. Techniques focus on matching behaviors of attackers to patterns of abnormal activity in the network. The need to stream, parse, and analyze large volumes of small, semi-structured data files can be feasibly addressed through a variety of techniques identified by researchers. This report highlights research in key areas, including protection of the data, performance of the systems and network bandwidth utilization. Anne M. Tall, Jun Wang 0001, Dezhi Han |
BDCAT | 2 |
| 2016 | Accelerating I/O Performance of SVM on HDFSabstractHadoop distributed file system (HDFS) is a major distributed file system for commodity clusters and cloud computing. Its extensive scalability and replica fault tolerance scheme makes it well suited for data-intensive application. Due to the tremendous growth of data, many computation-centric applications also become data-intensive. However, they are not optimal on HDFS, which leaves plenty of space for performance optimization. In this paper we ported an MPI-SVM solver, originally developed for HPC environment to the HDFS. We specifically improved the data pre-processing part that requires large amount of I/O operations by a deterministic scheduling method. Our improvement showed a balanced read pattern on each node. The time ratio between the longest process and the shortest process has been reduced by 60%. Also the average read time has significantly reduced by 78%. The data served on each node also showed a small variance in comparison with the originally ported SVM algorithm. We believe that our design avoids the overhead introduced by remote I/O operations, which will be beneficial to many algorithms when coping with large scale of data. Mao Ye 0008, Jun Wang 0001, Jiangling Yin, Xuhong Zhang 0002 |
CLUSTER | 2 |
| 2016 | AOS: adaptive overwrite scheme for energy-efficient MLC STT-RAM cacheabstractSpin-Transfer Torque Random Access Memory (STT-RAM) has been identified as an advantageous candidate for on-chip memory technology due to its high density and ultra low leakage power. Recent research progress in Magnetic Tunneling Junction (MTJ) devices has developed Multi-Level Cell (MLC) STT-RAM to further enhance cell density. To avoid the write disturbance in MLC strategy, data stored in the soft bit must be restored back immediately after the hard bit switching is completed. However, frequent restores are not only unnecessary, but also introduce a significant energy consumption overhead. In this paper, we propose an Adaptive Overwrite Scheme (AOS) which alleviates restoration overhead by intentionally overwriting selected soft bits based on RRD (Read Reuse Distance). Our experimental results show 54.6% reduction in soft bit restoration, delivering 10.8% decrease in overall energy consumption. Moreover, AOS promotes MLC to be a preferable L2 design alternative in terms of energy, area and latency product. Xunchao Chen, Navid Khoshavi, Jian Zhou 0004, Dan Huang 0001, Ronald F. DeMara, Jun Wang 0001, Wujie Wen, Yiran Chen 0001 |
DAC | 6 |
| 2016 | GreenMatch: Renewable-Aware Workload Scheduling for Massive Storage SystemsabstractAs datacenters grow in scale, increasing energy costs and carbon emissions have led data centers to seek renewable energy, such as wind and solar energy. However, tackling the challenges associated with the intermittent nature and variability of renewable energy is substantial. This paper proposes a scheme called GreenMatch, which deploys an SSD-cache to match green energy supplies with a time-shifting workload schedule while maintaining low latency for online data-intensive services. With the SSD-cache, the process for a latency-sensitive request to access a disk is divided into two stages: a low-energy low-latency online stage and a high-energy high-latency off-line stage. As the process in the latter stage is off-line, it offers opportunities for time-shifting workload scheduling in response to variations of green energy supplies. We also allocate an HDD-cache to guarantee data availability when renewable energy is non-adequate. Furthermore, we design a novel replacement policy called Inactive Disk First for the HDD-cache to avoid inactive disk accesses. The experimental results show that GreenMatch can make full use of renewable energy while minimizing the negative impact of intermittency and variability on performance and availability. Xiaoyang Qu, Jiguang Wan 0001, Jun Wang 0001, Liqiong Liu, Changsheng Xie 0001 |
IPDPS | 3 |
| 2016 | DataNet: A Data Distribution-Aware Method for Sub-Dataset Analysis on Distributed File SystemsabstractIn this paper, we study the problem of sub-dataset analysis over distributed file systems, e.g, the Hadoop file system. Our experiments show that the sub-datasets' distribution over HDFS blocks can often cause the corresponding analysis to suffer from a seriously imbalanced parallel execution. This is because the locality of individual sub-datasets is hidden by the Hadoop file system and the content clustering of sub-datasets results in some computational nodes carrying out much more workload than others. We conduct a comprehensive analysis on how the imbalanced computing patterns occur and their sensitivity to the size of a cluster. We then propose a novel method to optimize sub-dataset analysis over distributed storage systems referred to as DataNet. DataNet aims to achieve distribution-aware and workload-balanced computing and consists of the following three parts. Firstly, we propose an efficient algorithm with linear complexity to obtain the meta-data of sub-dataset distributions. Secondly, we design an elastic storage structure called ElasticMap based on the HashMap and BloomFilter techniques to store the meta-data. Thirdly, we employ a distribution-aware algorithm for sub-dataset applications to achieve a workload-balance in parallel-execution. Our proposed method can benefit different sub-dataset analyses with various computational requirements. Experiments are conducted on PRObEs Marmot 128-node cluster testbed and the results show the performance benefits of DataNet. Jun Wang 0001, Jiangling Yin, Jian Zhou 0004, Xuhong Zhang 0002 |
IPDPS | 1 |
| 2016 | TEES: A novel multiple criteria optimization scheme for temperature-constrained energy efficient storage
Jian Zhou 0004, Jun Wang 0001, Fei Wu 0005, Changsheng Xie 0001 |
J. Parallel Distributed Comput. | 2 |
| 2016 | A reliable power management scheme for consistent hashing based distributed key value storage systemsabstractDistributed key value storage systems are among the most important types of distributed storage systems currently deployed in data centers. Nowadays, enterprise data centers are facing growing pressure in reducing their power consumption. In this paper, we propose GreenCHT, a reliable power management scheme for consistent hashing based distributed key value storage systems. It consists of a multi-tier replication scheme, a reliable distributed log store, and a predictive power mode scheduler (PMS). Instead of randomly placing replicas of each object on a number of nodes in the consistent hash ring, we arrange the replicas of objects on nonoverlapping tiers of nodes in the ring. This allows the system to fall in various power modes by powering down subsets of servers while not violating data availability. The predictive PMS predicts workloads and adapts to load fluctuation. It cooperates with the multi-tier replication strategy to provide power proportionality for the system. To ensure that the reliability of the system is maintained when replicas are powered down, we distribute the writes to standby replicas to active servers, which ensures failure tolerance of the system. GreenCHT is implemented based on Sheepdog, a distributed key value storage system that uses consistent hashing as an underlying distributed hash table. By replaying 12 typical real workload traces collected from Microsoft, the evaluation results show that GreenCHT can provide significant power savings while maintaining a desired performance. We observe that GreenCHT can reduce power consumption by up to 35%–61%. Jiguang Wan 0001, Jun Wang 0001, Changsheng Xie 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2016 | Sapprox: Enabling Efficient and Accurate Approximations on Sub-datasets with Distribution-aware Online SamplingabstractIn this paper, we aim to enable both efficient and accurate approximations on arbitrary sub-datasets of a large dataset. Due to the prohibitive storage overhead of caching offline samples for each sub-dataset, existing offline sample based systems provide high accuracy results for only a limited number of sub-datasets, such as the popular ones. On the other hand, current online sample based approximation systems, which generate samples at runtime, do not take into account the uneven storage distribution of a sub-dataset. They work well for uniform distribution of a sub-dataset while suffer low sampling efficiency and poor estimation accuracy on unevenly distributed sub-datasets. To address the problem, we develop a distribution aware method calledSapprox. Our idea is to collect the occurrences of a sub-dataset at each logical partition of a dataset (storage distribution) in the distributed system, and make good use of such information to facilitate online sampling. There are three thrusts in Sapprox. First, we develop a probabilistic map to reduce the exponential number of recorded sub-datasets to a linear one. Second, we apply thecluster sampling with unequal probability theoryto implement a distribution-aware sampling method for efficient online sub-dataset sampling. Third, we quantitatively derive the optimal sampling unit size in a distributed file system by associating it with approximation costs and accuracy. We have implemented Sapprox into Hadoop ecosystem as an example system and open sourced it on GitHub. Our comprehensive experimental results show that Sapprox can achieve a speedup by up to 20× over the precise execution. Xuhong Zhang 0002, Jun Wang 0001, Jiangling Yin, Shouling Ji |
Proc. VLDB Endow. | 2 |
| 2016 | MAR: A Novel Power Management for CMP Systems in Data-Intensive EnvironmentabstractEmerging data-intensive applications are creating non-uniform CPU and I/O workloads which impose the requirement to consider both CPU and I/O effects in the power management strategies. Current approaches focus on scaling down the CPU frequency based on CPU busy/idle ratio without taking I/O into consideration. Therefore, they do not fully exploit the opportunities in power conservation. In this paper, we propose a novel power management scheme called model-free, adaptive, rule-based (MAR) in multiprocessor systems to minimize the CPU power consumption subject to performance constraints. By introducing new I/O wait status, MAR is able to accurately describe the relationship between core frequencies, performance and power consumption. Moreover, we adopt a model-free control method to filter out the I/O wait status from the traditional CPU busy/idle model in order to achieve fast responsiveness to burst situations and take full advantage of power saving. Our extensive experiments on a physical testbed demonstrate that, for SPEC benchmarks and data-intensive (TPC-C) benchmarks, an MAR prototype system achieves 95.8-97.8 percent accuracy of the ideal power saving strategy calculated offline. Compared with baseline solutions, MAR is able to save 12.3-16.1 percent more power while maintain a comparable performance loss of about 0.78-1.08 percent. In addition, more simulation results indicate that our design achieved 3.35-14.2 percent more power saving efficiency and 4.2-10.7 percent less performance loss under various CMP configurations as compared with various baseline approaches such as LAST, Relax, PID and MPC. Pengju Shang, Junyao Zhang 0007, Qingdong Wang, Jun Wang 0001 |
IEEE Trans. Computers | 6 |
| 2015 | Optimize Parallel Data Access in Big Data ProcessingabstractRecent years the Hadoop Distributed File System(HDFS) has been deployed as the bedrock for many parallel big data processing systems, such as graph processing systems, MPI-based parallel programs and scala/java-based Spark frameworks, which can efficiently support iterative and interactive data analysis in memory. The first part of my dissertation mainly focuses on studying parallel data accession distributed file systems, e.g, HDFS. Since the distributed I/O resources and global data distribution are often not taken into consideration, the data requests from parallel processes/executors will unfortunately be served in a remoter imbalanced fashion on the storage servers. In order to address these problems, we develop I/O middleware systems and matching-based algorithms to map parallel data requests to storage servers such that local and balanced data access can be achieved. The last part of my dissertation presents our plans to improve the performance of interactive data access in big data analysis. Specifically, most interactive analysis programs will scan through the entire data set regardless of which data is actually required. We plan to develop a content-aware method to quickly access required data without this laborious scanning process. Jiangling Yin, Jun Wang 0001 |
CCGRID | 2 |
| 2015 | VH-DSI: Speeding up Data Visualization via a Heterogeneous Distributed Storage InfrastructureabstractVisualizing and analyzing large-scale datasets are both critical and challenging, as they require substantial resources for data processing and storage. While the speed of supercomputers continues to set higher standard, the I/O systems have not kept in pace, resulting in a significant performance bottleneck. To alleviate the I/O bottleneck for scientific visualization applications, we propose a Visualization via a Heterogeneous Distributed Storage Infrastructure (VH-DSI) solution to improve I/O speed and accelerate overall visualization performance. VH-DSI replaces the traditional parallel file system with a distributed file system to support visualization applications. A new scheduling algorithm HeterSche is proposed in VH-DSI to assign computing tasks to data nodes with the consideration of cluster heterogeneity and data locality. VH-DSI also includes a design to support POSIX-IO for distributed file system. The performance evaluation has shown that the proposed VH-DSI solution can achieve significant performance improvement for visualization applications. Compared to the traditional visualization, the VH-DSI solution reduces the response time by at least 5 times. The HeterSche scheduling algorithm is capable to speed up visualization compared to other scheduling algorithms especially for large scale datasets. Juniarto Samsudin, Haixiang Shi, Jun Wang 0001 |
ICPADS | 4 |
| 2015 | Opass: Analysis and Optimization of Parallel Data Access on Distributed File SystemsabstractIn this paper, we study parallel data access on distributed file systems, e.g, the Hadoop file system. Our experiments show that parallel data read requests are often served data remotely and in an imbalanced fashion. This results in a serious disk access and data transfer contention on certain cluster/storage nodes. We conduct a complete analysis on how remote and imbalanced read patterns occur and how they are affected by the size of the cluster. We then propose a novel method to Optimize Parallel Data Access on Distributed File Systems referred to as Opass. The goal of Opass is to reduce remote parallel data accesses and achieve a higher balance of data read requests between cluster nodes. To achieve this goal, we represent the data read requests that are issued by parallel applications to cluster nodes as a graph data structure where edges weights encode the demands of data locality and load capacity. Then we propose new matching-based algorithms to match processes to data based on the configurations of the graph data structure so as to compute the maximum degree of data locality and balanced access. Our proposed method can benefit parallel data-intensive analysis with various parallel data access strategies. Experiments are conducted on PRObEs Marmot 128-node cluster tested and the results from both benchmark and well-known parallel applications show the performance benefits and scalability of Opass. Jiangling Yin, Jun Wang 0001, Jian Zhou 0004, Tyler Lukasiewicz, Dan Huang 0001, Junyao Zhang 0007 |
IPDPS | 2 |
| 2015 | GreenCHT: A power-proportional replication scheme for consistent hashing based key value storage systemsabstractDistributed key value storage systems are widely used by many popular networking corporations. Nevertheless, server power consumption has become a growing concern for key value storage system designers since the power consumption of servers contributes substantially to a data center's power bills. In this paper, we propose GreenCHT, a power-proportional replication scheme for consistent hashing based key value storage systems. GreenCHT consists of a power-aware replication strategy — multi-tier replication strategy and a centralized power control service — predictive power-mode scheduler. The multitier replication provides power-proportionality and ensures data availability, reliability, consistency, as well as fault-tolerance of the whole system. The predictive power-mode scheduler component predicts workloads and exploits load fluctuation to schedule nodes to be powered-up and powered-down. GreenCHT is implemented based on Sheepdog, a distributed key value system that uses consistent hashing as an underlying distributed hash table. By replicating twelve real workload traces collected from Microsoft, the evaluation results show that GreenCHT can provide significant power savings while maintaining an acceptable performance. We observed that GreenCHT can reduce power consumption by up to 35%–61%. Jiguang Wan 0001, Jun Wang 0001, Changsheng Xie 0001 |
MSST | 3 |
| 2015 | Message from the program co-chairsabstractWe would first like to express our many thanks to Resit Sendag, this year's General Chair, as well as the rest of the research community for the opportunity to chair the program for the 10thIEEE International Conference on Networking, Architecture, and Storage (NAS 2015). Jun Wang 0001, R. Iris Bahar |
NAS | 1 |
| 2015 | Achieving up to zero communication delay in BSP-based graph processing via vertex categorizationabstractThe Bulk Synchronous Parallel (BSP) model, which divides a graphing algorithm into multiple supersteps, has become extremely popular in distributed graph processing systems. However, the high number of network messages exchanged in each superstep of the graph algorithm will create a long period of time. We refer to this as a communication delay. Furthermore, the BSP's global synchronization barrier does not allow computation in the next superstrep to be scheduled during this communication delay. This communication delay makes up a large percentage of the overall processing time of a superstep. While most recent research has focused on reducing number of network messages, but communication delay is still a deterministic factor for overall performance. In this paper, we add a runtime communication and computation scheduler into current graph BSP implementations. This scheduler will move some computation from the next superstep to the communication phase in the current superstep to mitigate the communication delay. Finally, we prototyped our system, Zebra, on Apache Hama, which is an open source clone of the classic Google Pregel. By running a set of graph algorithms on an in-house cluster, our evaluation shows that our system could completely eliminate the communication delay in the best case and can achieve average 2X speedup over Hama. Xuhong Zhang 0002, Xunchao Chen, Jun Wang 0001, Tyler Lukasiewicz, Dezhi Han |
NAS | 4 |
| 2015 | On the Cooling of Energy Efficient StorageabstractEnergy consumption has become an important issue in storage systems. Existing energy control solutions emphasize power consumption without considering re- liability degradation that results from overburden of those long standing disks. In this paper, we develop a novel multiple criteria optimization scheme based on Fuzzy Decision Making theory, for the Cool Energy Efficient Storage System called CEES. CEES aims to enforce a temperature constraint as well as performance requirements while also keeping energy consumption to a minimum. This is achieved by aggregating all the decision criteria, such as I/O performance, power consumption, temperature and frequency of disk-status transition. We first calculate the satisfaction degree of each criteria. Then, we use the weighted averaging satisfaction degree to determine the system control sequence. The experimental results show that CEES is able to reduce disk temperature by 20–30% as compared with existing control methods, while obtaining comparable performance and power consumption. Jian Zhou 0004, Jun Wang 0001, Fei Wu 0005, Changsheng Xie 0001, Dezhi Han |
NAS | 2 |
| 2015 | PERP: Attacking the balance among energy, performance and recovery in storage systems
Junyao Zhang 0007, Qingdong Wang, Jiangling Yin, Jian Zhou 0004, Jun Wang 0001 |
J. Parallel Distributed Comput. | 5 |
| 2015 | ThinRAID: Thinning Down RAID Array for Energy ConservationabstractThe current power managements in RAID array are mostly designed to conserve energy by spinning down partial disks of standard RAID architecture. However, spinning down several disks not only decreases disk parallelism, but also creates new problems, for example, partial chunks of the stripe cannot be accessed directly or multiple chunks of the same stripe are stored on the same disk, which affect spatial locality. We refer these problems as stripe degradation, which results in further performance degradation. To avoid such problems, this paper proposes a new RAID storage architecture called ThinRAID, which uses a subset of disks to build a capacity-adaptive RAID array based on the volume of the data set. Also, the other non-essential disks are spun down to save energy. When the workload is projected to become heavier based on our forecast model, data are migrated to disks that have recently transitioned from standby to active. Furthermore, we also propose a novel data reorganization algorithm that can minimize data migration. We have implemented ThinRAID in the Linux kernel and evaluated its performance and energy efficiency by replaying seven representative traces. Experimental results show that ThinRAID can save 15-27 percent on energy on average over conventional RAID, with minimum performance degradation. In comparison to PARAID, ThinRAID achieves up to 62 percent performance improvement. Jiguang Wan 0001, Xiaoyang Qu, Jun Wang 0001, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2014 | SLAM: scalable locality-aware middleware for I/O in scientific analysis and visualizationabstractWhereas traditional scientific applications are computationally intensive, recent applications require more data-intensive analysis and visualization. As the computational power and size of compute clusters continue to increase, the I/O read rates and associated network cost for these data-intensive applications create a serious performance bottleneck when faced with the massive data sets of today's "big data" era. Jiangling Yin, Jun Wang 0001, Wu-chun Feng, Xuhong Zhang 0002, Junyao Zhang 0007 |
HPDC | 2 |
| 2014 | ScalScheduling: A Scalable Scheduling Architecture for MPI-based interactive analysis programsabstractIn today's large scale clusters, running tasks with high degrees of parallelism allows interactive data visualization/analysis to complete in seconds. However, conventional, centralized scheduling poses significant challenges for these interactive applications. As the amount of data to be processed grows, it becomes too heavy to move across the network. Thus, data processing tasks should be scheduled such that the amount of transferred data is minimized, i.e., realizing data locality computation. To implement this, a scheduler process should collect and analyze data distribution metadata prior to making scheduling decisions, which usually causes milliseconds or seconds of latency. Such scheduling delay is unacceptable for interactive data applications. In this paper, we present a Scalable Scheduling Architecture for conventional interactive data programs and refer to it as ScalScheduling. ScalScheduling is proposed to reduce task scheduling latency, while ensuring the worker processes achieve a high degree of data locality computation and load balance in heterogeneous environments. In our proposed architecture, each worker process uses a novel Modulo-based priority method to schedule its local tasks independently. Multiple scheduler processes are employed according to the number of worker processes to resolve the issue of concurrent requests and assign remote tasks with respect to load balance. We perform experiments using thousands of parallel processes, and the experimental results show the benefits of our proposed scheduling architecture as well as its potential for future oversize task scheduling problems on large-scale clusters. Jiangling Yin, Andrew Foran, Xuhong Zhang 0002, Jun Wang 0001 |
ICCCN | 4 |
| 2014 | Reliability Analysis on Shifted and Random Declustering Block Layouts in Scale-Out Storage ArchitecturesabstractReliability is a critical metric in the design and development of scale-out data storage clusters. A general multiway replication-based declustering scheme has been widely used in enterprise large-scale storage systems to improve the I/O parallelism. Unfortunately, given an increasing number of node failures, how often a cluster starts losing data when being scaled-out is not well investigated. In this paper, we studied the reliability of multi-way declustering layouts by developing an extended model, more specifically abstracting the Continuous Time Markov chain to an ordinary differentiate equation group, and analyzing their potential parallel recovery possibilities. Our comprehensive simulation results on Mat lab and SHARPE show that the shifted declustering layout outperforms the random declustering layout in a multi-way replication scale-out architecture, in terms of data loss probability and system reliability by up to 63% and 85% respectively. Our study on both 5-year and 10-year system reliability equipped with various recovery bandwidth settings shows that, the shifted declustering layout surpasses the random declustering layout in both cases by consuming up to 5.2% and 11% less recovery bandwidth. Jun Wang 0001, Jiangling Yin, Huijun Zhu, Yuanyuan Yang 0001 |
NAS | 1 |
| 2014 | SDAFT: A novel scalable data access framework for parallel BLAST
Jiangling Yin, Junyao Zhang 0007, Jun Wang 0001, Wu-chun Feng |
Parallel Comput. | 3 |
| 2013 | DL-MPI: Enabling data locality computation for MPI-based data-intensive applicationsabstractCurrently, most scientific applications based on MPI adopt a compute-centric architecture. Needed data is accessed by MPI processes running on different nodes through a shared file system. Unfortunately, the explosive growth of scientific data undermines the high performance of MPI-based applications, especially in the execution environment of commodity clusters. In this paper, we present a novel approach to enable data locality computation for MPI-based data-intensive applications and refer to it as DL-MPI. DL-MPI allows MPI-based programs to obtain data distribution information for compute nodes through a novel data locality API. In addition, the problem of allocating data processing tasks to parallel processes is formulated as an integer optimization problem with the objectives of achieving data locality computation and optimal parallel execution time. For heterogeneous runtime environments, we propose a scheduling algorithm based on probability to dynamically schedule tasks to processes by evaluating the unprocessed local data and the computing ability of each compute node. We demonstrate the functionality of our methods through the implementation of scientific data processing programs as well as the incorporation of DL-MPI with existing HPC applications. Jiangling Yin, Andrew Foran, Jun Wang 0001 |
IEEE BigData | 3 |
| 2013 | Supporting HPC Analytics Applications with Access Patterns Using Data Restructuring and Data-Centric Scheduling Techniques in MapReduceabstractCurrent High Performance Computing (HPC) applications have seen an explosive growth in the size of data in recent years. Many application scientists have initiated efforts to integrate data-intensive computing into computational-intensive HPC facilities, particularly for data analytics. We have observed several scientific applications which must migrate their data from an HPC storage system to a data-intensive one for analytics. There is a gap between the data semantics of HPC storage and data-intensive system, hence, once migrated, the data must be further refined and reorganized. This reorganization must be performed before existing data-intensive tools such as MapReduce can be used to analyze data. This reorganization requires at least two complete scans through the data set and then at least one MapReduce program to prepare the data before analyzing it. Running multiple MapReduce phases causes significant overhead for the application, in the form of excessive I/O operations. That is for every MapReduce phase, a distributed read and write operation on the file system must be performed. Our contribution is to develop a MapReduce-based framework for HPC analytics to eliminate the multiple scans and also reduce the number of data preprocessing MapReduce programs. We also implement a data-centric scheduler to further improve the performance of HPC analytics MapReduce programs by maintaining the data locality. We have added additional expressiveness to the MapReduce language to allow application scientists to specify the logical semantics of their data such that 1) the data can be analyzed without running multiple data preprocessing MapReduce programs, and 2) the data can be simultaneously reorganized as it is migrated to the data-intensive file system. Using our augmented Map-Reduce system, MapReduce with Access Patterns (MRAP), we have demonstrated up to 33 percent throughput improvement in one real application, and up to 70 percent in an I/O kernel of another application. Our results for scheduling show up to 49 percent improvement for an I/O kernel of a prevalent HPC analysis application. Saba Sehrish, Grant Mackey, Pengju Shang, Jun Wang 0001, John Bent |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2012 | Generic Parallel Programming for Massive Remote Sensing Data ProcessingabstractRemote Sensing (RS) data processing is characterized by massive remote sensing images and increasing amount of algorithms of higher complexity. Parallel programming for data-intensive applications like massive remote sensing image processing on parallel systems is bound to be especially trivial and challenging. We propose a C++ template mechanism enabled generic parallel programming skeleton for these remote sensing applications in high performance clusters. It provides both programming templates for distributed RS data and generic parallel skeletons for RS algorithms. Through one-side communication primitives provided by MPI, the distributed RS data template could provide a global view of the big RS data whose sliced data blocks are scattered among the distributed memory of cluster nodes. Moreover, by data serialization and RMA (Remote Memory Access), the data templates could also offer a simple and effective way to distribute and communicate massive remote sensing data with complex data structures. Furthermore, the generic parallel skeletons implement the recurring patterns of computation, performance optimization and pass the user-defined sequential functions as parameters of templates for type genericity. With the implemented skeletons, Developers without extensive parallel computing technologies can implement efficient parallel remote sensing programs without concerning for parallel computing details. Through experiments on remote sensing applications, we confirmed that our templates were productive and efficient. Yan Ma 0001, Lizhe Wang 0001, Dingsheng Liu, Peng Liu 0024, Jun Wang 0001, Jie Tao 0001 |
CLUSTER | 5 |
| 2012 | Parallel Processing of Massive EEG Data with MapReduceabstractAnalysis of neural signals like electroencephalogram (EEG) is one of the key technologies in detecting and diagnosing various brain disorders. As neural signals are non-stationary and non-linear in nature, it is almost impossible to understand their true physical dynamics until the recent advent of the Ensemble Empirical Mode Decomposition (EEMD) algorithm. The neural signal processing with EEMD is highly compute-intensive due to the high complexity of the EEMD algorithm. It is also data intensive because 1) EEG signals contain massive data sets 2) EEMD has to introduce a large number of trials in processing to ensure precision. The Map Reduce programming mode is a promising parallel computing paradigm for data intensive computing. To increase the efficiency and performance of the neural signal analysis, this research develops parallel EEMD neural signal processing with Map Reduce. In this paper, we implement the parallel EEMD with Hadoop in a modern cyber infrastructure. Test results and performance evaluation show that parallel EEMD can significantly improve the performance of neural signal processing. Lizhe Wang 0001, Dan Chen 0001, Rajiv Ranjan 0001, Samee Ullah Khan, Joanna Kolodziej, Jun Wang 0001 |
ICPADS | 6 |
| 2012 | NOHAA: A NOvel Framework for HPC Analytics over Windows AzureabstractHPC analytics has become increasingly vital to analyze the large volumes of data produced by sophisticated computing instruments. Meanwhile, with the successful development of cloud computing, more and more scientists are devoted to deploy HPC analytics in the ever-popular clouds, which poses new challenges mainly caused by different storage architectures, resource management mechanisms and programming APIs. Firstly, there exists a ``data semantics" gap between the way data are stored by Cloud platform and the way data will be accessed by the HPC Analytics. Secondly, data are mostly distributed across data nodes for in-house data-intensive clusters to achieve co-located computation and storage, however, it is challenging for the public clouds to mimic because their data are stored centrally. In this paper, we develop a new HPC analytics framework called NOHAA, to provide 1) a semantics-aware intelligent data upload interface and 2) a locality-aware hierarchical storage system in support of co-located computation and storage on Windows Azure. Our extensive real world experiments show that NOHAA significantly reduces the average data access time by up to 85% and accelerates the HPC analytics execution time by a factor of 2 to 7. Qiangju Xiao, Jun Wang 0001, Yan Ma 0001, Lizhe Wang 0001 |
ICPADS | 2 |
| 2012 | A new high-performance, energy-efficient replication storage system with reliability guaranteeabstractIn modern replication storage systems where data carries two or more multiple copies, a primary group of disks is always up to service incoming requests while other disks are often spun down to sleep states to save energy during slack periods. However, since new writes cannot be immediately synchronized onto all disks, system reliability is degraded. This paper develops PERAID, a new high-performance, energy-efficient replication storage system, which aims to improve both performance and energy efficiency without compromising reliability. It employs a parity software RAID as a virtual write buffer disk at the front end to absorb new writes. Since extra parity redundancy supplies two or more copies, PERAID guarantees comparable reliability with that of a replication storage system. In addition, PERAID offers better write performance compared to the replication system by avoiding the classical small-write problem in traditional parity RAID: buffering many small random writes into few large writes and writing to storage in a parallel fashion. By evaluating our PERAID prototype using two benchmarks and two real-life traces, we found that PERAID significantly improves write performance and saves more energy than existing solutions such as GRAID, eRAID. Jiguang Wan 0001, Jun Wang 0001, Changsheng Xie 0001 |
MSST | 3 |
| 2012 | Co-located Compute and Binary File Storage in Data-Intensive ComputingabstractWith the rapid development of computation capability, the massive increase in data volume has outmoded compute-intensive clusters for HPC analysis of large-scale data sets due to a huge amount of data transfer over network. Co-located compute and storage has been introduced in data-intensive clusters to avoid network bottleneck by launching the computation on nodes in which most of the input data reside. Chunk-based storage systems are typical examples, splitting data into blocks and randomly storing them across nodes. Records as the input data for the analysis are read from blocks. This method implicitly assumes that a single record resides on a single node and then data transfer can be avoided. However, this assumption does not always hold because there is a gap between records and blocks. The current solution overlooks the relationship between the computation unit as a record and the storage unit as a block. For situations with records belonging to one block, there would be no data transfer. But in practice, one record could consist of several blocks. This is especially true for binary files, which introduce extra data transfer due to preparing the input data before conducting the analysis. Blocks belonging to a single record are scattered randomly across the data nodes regardless of to the semantics of the records. To address these problems, we develop two solutions in this paper, one is to develop a Record-Based Block Distribution (RBBD) framework and the other is a data-centric scheduling using a Weighted Set Cover Scheduling (WSCS) to schedule the tasks. The Record-Based Block Distribution (RBBD) framework for data-intensive analytics aims to eliminate the gap between records and blocks and accomplishes zero data transfer among nodes. The Weighted Set Cover Scheduling (WSCS) is proposed to further improve the performance by optimizing the combination of nodes. Our experiments show that overlooking the record and block relationship can cause severe performance problems when a record is comprised of several blocks scattered in different nodes. Our proposed novel data storage strategy, Record-Based Block Distribution (RBBD), optimizes the block distribution according to the record and block relationship. By being combined with our novel scheduling Weighted Set Cover Scheduling (WSCS), we efficiently reduces extra data transfers, and eventually improves the performance of the chunk-based storage system. Using our RBBD framework and WSCS in chunk-based storage system, our extensive experiments show that the data transfer decreases by 36.4% (average) and the scheduling algorithm outperforms the random algorithm by 51%-62%; the deviation from the ideal solutions is no more than 6.8%. Qiangju Xiao, Pengju Shang, Jun Wang 0001 |
NAS | 3 |
| 2012 | TRAID: Exploiting Temporal Redundancy and Spatial Redundancy to Boost Transaction Processing Systems PerformanceabstractIn the past few years, more storage system applications have employed transaction processing techniques to ensure data integrity and consistency. Logging is one of the key requirements to ensure transaction Atomicity, Consistency, Isolation, Durability (ACID) properties and data recoverability in transaction processing systems (TPS). Recently, emerging complex I/O bound transactions have resulted in substantially more log content and higher log flushing latency. The latency will delay transaction commit and decrease the overall throughput of the TPS. On the other hand, RAID is widely used as the underlying storage system for Databases to guarantee system reliability and availability with high I/O performance. In this paper, we observe the overlap between the redundancies in the underlying RAID storage system and database logging system, and propose a novel reliable storage architecture called Transactional RAID (TRAID). TRAID deduplicates this overlap by only logging one compact version (XOR results) of recovery references for the updating data. It minimizes the amount of log content and thereby boosts the overall transaction processing performance. At the same time, TRAID guarantees the same RAID reliability, as well as recovery correctness and ACID semantics as current TPS setups. We experiment on two open-source database systems: Berkeley DB and PostgreSQL, with three different workloads: standard OLTP benchmark TPC-C, customized TPC-C with strong access locality, and customized TPC-C with write-intensive property. Then we test TRAID performance with "Group Commit” enabled. Finally, we evaluate the recovery efficiency of TRAID. Our extensive results demonstrate that for throughput, TRAID outperforms RAID by 43.24-69.5 percent for various workloads; it also saves on log space by 28.57-35.48 percent, and outperforms RAID by about 20 percent in throughput with "Group Commit” enabled. At last, we show that TRAID outperforms RAID from 28.7 to 35.7 percent during the recovery. Pengju Shang, Saba Sehrish, Jun Wang 0001 |
IEEE Trans. Computers | 3 |
| 2012 | Reduced Function Set Abstraction (RFSA) for MPI-IO
Saba Sehrish, Jun Wang 0001 |
J. Supercomput. | 2 |
| 2011 | VisIO: Enabling Interactive Visualization of Ultra-Scale, Time Series Data via High-Bandwidth Distributed I/O SystemsabstractPetascale simulations compute at resolutions ranging into billions of cells and write terabytes of data for visualization and analysis. Interactive visualization of this time series is a desired step before starting a new run. The I/O subsystem and associated network often are a significant impediment to interactive visualization of time-varying data, as they are not configured or provisioned to provide necessary I/O read rates. In this paper, we propose a new I/O library for visualization applications: VisIO. Visualization applications commonly use N-to-N reads within their parallel enabled readers which provides an incentive for a shared-nothing approach to I/O, similar to other data-intensive approaches such as Hadoop. However, unlike other data-intensive applications, visualization requires: (1) interactive performance for large data volumes, (2) compatibility with MPI and POSIX file system semantics for compatibility with existing infrastructure, and (3) use of existing file formats and their stipulated data partitioning rules. VisIO, provides a mechanism for using a non-POSIX distributed file system to provide linear scaling of I/O bandwidth. In addition, we introduce a novel scheduling algorithm that helps to co-locate visualization processes on nodes with the requested data. Testing using VisIO integrated into Para View was conducted using the Hadoop Distributed File System (HDFS) on TACC's Longhorn cluster. A representative dataset, VPIC, across 128 nodes showed a 64.4% read performance improvement compared to the provided Lustre installation. Also tested, was a dataset representing a global ocean salinity simulation that showed a 51.4% improvement in read performance over Lustre when using our VisIO system. VisIO, provides powerful high-performance I/O services to visualization applications, allowing for interactive performance with ultra-scale, time-series data. Christopher Mitchell, James P. Ahrens, Jun Wang 0001 |
IPDPS | 3 |
| 2011 | A Novel Power Management for CMP Systems in Data-Intensive EnvironmentabstractThe emerging data-intensive applications of today are comprised of non-uniform CPU and I/O intensive workloads, thus imposing a requirement to consider both CPU and I/O effects in the power management strategies. Only scaling down the processor's frequency based on its busy/idle ratio cannot fully exploit opportunities of saving power. Our experiments show that besides the busy and idle status, each processor may also have I/O wait phases waiting for I/O operations to complete. During this period, the completion time is decided by the I/O subsystem rather than the CPU thus scaling the processor to a lower frequency will not affect the performance but save more power. In addition, the CPU's reaction to the I/O operations may be significantly affected by several factors, such as I/O type (sync or unsync), instruction/job level parallelism, it cannot be accurately modeled via physics laws like mechanical or chemical systems. In this paper, we propose a novel power management scheme called MAR (modeless, adaptive, rule-based) in multiprocessor systems to minimize the CPU power consumption under performance constraints. By using richer feedback factors, e.g. the I/O wait, MAR is able to accurately describe the relationships among core frequencies, performance and power consumption. We adopt a modeless control model to reduce the complexity of system modeling. MAR is designed for CMP (Chip Multi Processor) systems by employing multi-input/multi-output (MIMO) theory and per core level DVFS (Dynamic Voltage and Frequency Scaling). Our extensive experiments on a physical test bed demonstrate that, for the SPEC benchmark and data-intensive (TPC-C) benchmark, the efficiency of MAR is 93.6-96.2\% accurate to the ideal power saving strategy calculated off-line. Compared with baseline solutions, MAR could save 22.5-32.5\% more power while keeping the comparable performance loss of about 1.8-2.9\%. In addition, simulation results show the efficiency of our design for various CMP configurations. Pengju Shang, Jun Wang 0001 |
IPDPS | 2 |
| 2011 | A Scalable Reverse Lookup Scheme Using Group-Based Shifted Declustering LayoutabstractRecent years have witnessed an increasing demand for super data clusters. The super data clusters have reached the petabyte-scale that can consist of thousands or tens of thousands storage nodes at a single site. For this architecture, reliability is becoming a great concern. In order to achieve a high reliability, data recovery and node reconstruction is a must. Although extensive research works have investigated how to sustain high performance and high reliability in case of node failures at large scale, a reverse lookup problem, namely finding the objects list for the failed node remains open. This is especially true for storage systems with high requirement of data integrity and availability, such as scientific research data clusters and etc. Existing solutions are either time consuming or expensive. Meanwhile, replication based block placement can be used to realize fast reverse lookup. However, they are designed for centralized, small-scale storage architectures. In this paper, we propose a fast and efficient reverse lookup scheme named Group-based Shifted Declustering (G-SD) layout that is able to locate the whole content of the failed node. G-SD extends our previous shifted declustering layout and applies to large-scale file systems. Our mathematical proofs and real-life experiments show that G-SD is a scalable reverse lookup scheme that is up to one order of magnitude faster than existing schemes. Junyao Zhang 0007, Pengju Shang, Jun Wang 0001 |
IPDPS | 3 |
| 2011 | Classified Power Capping with Distribution Trees in Cloud ComputingabstractPower management is becoming very important in data centers. Cloud computing is also one of the newest promising data center techniques which is appealing to many big companies. As cloud computing is different from current data centers in terms of power management due to a dynamic structure and property for its online service. Power budgeting, in terms of its important role in power management, provides powerful solutions for cloud computing with dynamic capabilities. To be specific, existing methods for data centers are based on power distribution units (PDU) divided by fixed locations on physical levels. However, it is not suitable for cloud with the dynamic property. We propose a power management design based at the logical level which uses a distribution tree with classified power capping by different service or workload types. By setting multiple trees, we can differentiate and analyze the effect of workload types and Service Level Agreements (SLAs) in terms of power characteristics. Zhengkai Wu, Junyao Zhang 0007, Christopher Giles, Jun Wang 0001 |
NAS | 4 |
| 2011 | A New Placement-Ideal Layout for Multiway Replication Storage SystemabstractTechnology trends are making sophisticated replication-based storage architectures become a standard commercial practice in today's computing. Existing solutions successfully developed optimal and near-optimal parallelism layouts such as declustered parity organizations at small-scale storage architectures. There are very few studies on multiway replication-based storage architectures that are significantly different from parity-based storage architectures. It is difficult to scale up to a large size because current placement-ideal solutions have a limited number of configurations. In this paper, we retrofit the desirable properties of optimal parallelism definitions in parity architectures for replication architectures, and propose a novel placement-ideal data layout-shifted declustering. Shifted declustering layout obtains optimal parallelism in a wide range of configurations, and obtains optimal high performance and load balancing in both fault-free and degraded mode. Our theoretical proofs and comprehensive simulation results show that shifted declustering is superior in performance, load balancing, and reliability to traditional layout schemes such as standard mirroring, chained declustering, group-rotational declustering, and existing parity layout schemes PRIME and RELPR. Pengju Shang, Jun Wang 0001, Huijun Zhu |
IEEE Trans. Computers | 2 |
| 2010 | MRAP: a novel MapReduce-based framework to support HPC analytics applications with access patternsabstractDue to the explosive growth in the size of scientific data sets, data-intensive computing is an emerging trend in computational science. Many application scientists are looking to integrate data-intensive computing into computational-intensive High Performance Computing facilities, particularly for data analytics. We have observed several scientific applications which must migrate their data from an HPC storage system to a data-intensive one. There is a gap between the data semantics of HPC storage and data-intensive system, hence, once migrated, the data must be further refined and reorganized. This reorganization requires at least two complete scans through the data set and then at least one MapReduce program to prepare the data before analyzing it. Running multiple MapReduce phases causes significant overhead for the application, in the form of excessive I/O operations. For every MapReduce application that must be run in order to complete the desired data analysis, a distributed read and write operation on the file system must be performed. Our contribution is to extend Map-Reduce to eliminate the multiple scans and also reduce the number of pre-processing MapReduce programs. We have added additional expressiveness to the MapReduce language to allow users to specify the logical semantics of their data such that 1) the data can be analyzed without running multiple data pre-processing MapReduce programs, and 2) the data can be simultaneously reorganized as it is migrated to the data-intensive file system. Using our augmented MapReduce system, MapReduce with Access Patterns (MRAP), we have demonstrated up to 33% throughput improvement in one real application, and up to 70% in an I/O kernel of another application. Saba Sehrish, Grant Mackey, Jun Wang 0001, John Bent |
HPDC | 3 |
| 2010 | Concentric Layout, a New Scientific Data Distribution Scheme in Hadoop File SystemabstractThe data generated by scientific simulation, sensor, monitor or optical telescope has increased with dramatic speed. In order to analyze the raw data fast and space efficiently, data pre-process operation is needed to achieve better performance in data analysis phase. Current research shows an increasing tread of adopting MapReduce framework for large scale data processing. However, the data access patterns which generally applied to scientific data set are not supported by current MapReduce framework directly. The gap between the requirement from analytics application and the property of MapReduce framework motivates us to provide support for these data access patterns in MapReduce framework. In our work, we studied the data access patterns in matrix files and proposed a new concentric data layout solution to facilitate matrix data access and analysis in MapReduce framework. Concentric data layout is a hierarchical data layout which maintains the dimensional property in large data sets. Contrary to the continuous data layout adopted in current Hadoop framework, concentric data layout stores the data from the same sub-matrix into one chunk, and then stores chunks symmetrically in a higher level. This matches well with the matrix like computation. The concentric data layout preprocesses the data beforehand, and optimizes the afterward run of MapReduce application. The experiments show that the concentric data layout improves the overall performance, reduces the execution time by about 38% when reading a 64 GB file. It also mitigates the unused data read overhead and increases the useful data efficiency by 32% on average. Pengju Shang, Saba Sehrish, Grant Mackey, Jun Wang 0001 |
NAS | 5 |
| 2010 | A Novel Weighted-Graph-Based Grouping Algorithm for Metadata PrefetchingabstractAlthough data prefetching algorithms have been extensively studied for years, there is no counterpart research done for metadata access performance. Existing data prefetching algorithms, either lack of emphasis on group prefetching, or bearing a high level of computational complexity, do not work well with metadata prefetching cases. Therefore, an efficient, accurate, and distributed metadata-oriented prefetching scheme is critical to leverage the overall performance in large distributed storage systems. In this paper, we present a novel weighted-graph-based prefetching technique, built on both direct and indirect successor relationship, to reap performance benefit from prefetching specifically for clustered metadata servers, an arrangement envisioned necessary for petabyte-scale distributed storage systems. Extensive trace-driven simulations show that by adopting our new metadata prefetching algorithm, the miss rate for metadata accesses on the client site can be effectively reduced, while the average response time of metadata operations can be dramatically cut by up to 67 percent, compared with legacy LRU caching algorithm and existing state-of-the-art prefetching algorithms. Jun Wang 0001, Hong Jiang 0001, Pengju Shang |
IEEE Trans. Computers | 2 |
| 2009 | Improving metadata management for small files in HDFSabstractScientific applications are adapting HDFS/MapReduce to perform large scale data analytics. One of the major challenges is that an overabundance of small files is common in these applications, and HDFS manages all its files through a single server, the Namenode. It is anticipated that small files can significantly impact the performance of Namenode. In this work we propose a mechanism to store small files in HDFS efficiently and improve the space utilization for metadata. Our scheme is based on the assumption that each client is assigned a quota in the file system, for both the space and number of files. In our approach, we utilize the compression method ‘harballing', provided by Hadoop, to better utilize the HDFS. We provide for new job functionality to allow for in-job archival of directories and files so that running MapReduce programs may complete without being killed by the JobTracker due to quota policies. This approach leads to better functionality of metadata operations and more efficient usage of the HDFS. Our analysis results show that we can reduce the metadata footprint in main memory by a factor of 42. Grant Mackey, Saba Sehrish, Jun Wang 0001 |
CLUSTER | 3 |
| 2009 | Overlapped checkpointing with hardware assistabstractWe present a new approach to handling the demanding I/O workload incurred during checkpoint writes encountered in High Performance Computing. Prior efforts to improve performance have been bound by issues such as hard drive limitations, and the network. Our research surpasses this limitation by providing a method to: (1) write checkpoint data to a high-speed, non-volatile buffer, and (2) asynchronously write this data to permanent storage while resuming computation. This removes the hard drive from the critical data path because our I/O node based buffers isolate the compute nodes from the storage servers. This solution is feasible because of industry declines in cost for high-capacity, non-volatile storage technologies. Testing was conducted using a standardized HPC benchmark on a test bed cluster at Los Alamos National Laboratory. Results show a definitive speedup factor for select workloads over writing directly to a typical global parallel file system; the Panasas ActiveScale File System. Christopher Mitchell, James Nunez, Jun Wang 0001 |
CLUSTER | 3 |
| 2009 | Smart read/write for MPI-IOabstractWe present a case for automating the selection of MPI-IO performance optimizations, with an ultimate goal to relieve the application programmer from these details, thereby improving their productivity. Programmers productivity has always been overlooked as compared to the performance optimizations in high performance computing community. In this paper we present RFSA, a reduced function set abstraction based on an existing parallel programming interface (MPI-IO) for I/O. MPI-IO provides high performance I/O function calls to the scientists/engineers writing parallel programs; who are required to use the most appropriate optimization of a specific function, hence limits the programmer productivity. Therefore, we propose a set of reduced functions with an automatic selection algorithm to decide what specific MPI-IO function to use. We implement a selection algorithm for I/O functions like read, write, etc. RFSA replaces 6 different flavors of read and write functions by one read and write function. By running different parallel I/O benchmarks on both medium-scale clusters and NERSC supercomputers, we show that RFSA functions impose minimal performance penalties. Saba Sehrish, Jun Wang 0001 |
IPDPS | 2 |
| 2009 | An advertisement-based peer-to-peer search algorithm
Jun Wang 0001, Hailong Cai |
J. Parallel Distributed Comput. | 1 |
| 2009 | A New Hierarchical Data Cache Architecture for iSCSI Storage ServerabstractWith the emergence of data-intensive applications, recent years have seen a fast-growing volume of I/O traffic propagated through the local I/O interconnect bus. This raises up a question for storage servers on how to resolve such a potential bottleneck. In this paper, we present a hierarchical data cache architecture called DCA to effectively slash local interconnect traffic and thus boost the storage server performance. A popular iSCSI storage server architecture is chosen as an example. DCA is composed of a read cache in NIC called NIC cache and a read/write unified cache in host memory called helper cache. The NIC cache services most portions of read requests without fetching data via the PCI bus, while the helper cache (1) supplies some portions of read requests per partial NIC cache hit, (2) directs cache placement for NIC cache, and (3) absorbs most transient writes locally. We develop a novel state-locality-aware cache placement algorithm called SLAP to improve the NIC cache hit ratio for mixed read and write workloads. To demonstrate the effectiveness of DCA, we develop a DCA prototype system and evaluate it with an open source iSCSI implementation under representative storage server workloads. Experimental results showed that DCA can boost iSCSI storage server throughput by up to 121 percent and reduce the PCI traffic by up to 74 percent compared with an iSCSI target without DCA. Jun Wang 0001, Xiaoyu Yao, Christopher Mitchell |
IEEE Trans. Computers | 1 |
| 2008 | Bridging the Gap Between Parallel File Systems and Local File Systems: A Case Study with PVFSabstractParallel I/O plays an increasingly important role in today's data intensive computing applications. While much attention has been paid to parallel read performance, most of this work has focused on the parallel file system, middleware, or application layers, ignoring the potential for improvement through more effective use of local storage. In this paper, we present the design and implementation of Segment-structuredOn-disk data Grouping and Prefetching (SOGP), a technique that leverages additional local storage to boost the local data read performance for parallel file systems, especially for those applications with partially overlapped access patterns. Parallel Virtual File System (PVFS) is chosen as an example. Our experiments show that an SOGP-enhanced PVFS prototype system can outperforma traditional Linux-Ext3-based PVFS for many applications and benchmarks, in some tests by as much as 230% in terms of I/O bandwidth. Jun Wang 0001 |
ICPP | 2 |
| 2008 | Shifted declustering: a placement-ideal layout scheme for multi-way replication storage architectureabstractRecent years have seen a growing interest in the deployment of sophisticated replication based storage architecture in data-intensive computing. Existing placement-ideal data layout solutions place an emphasis on declustered parity based storage. However, there exist major differences between parity and replication architectures, especially in data layouts. We retrofit the desirable properties of optimal parallelism in parity architectures for replication architectures, and propose a novel placement-ideal data layout ---- shifted declustering for replication based storage. Shifted declustering layout obtains optimal parallelism in a wide range of configurations, and obtains optimal high performance and load balancing in both fault-free and degraded modes. Our theoretical proofs and comprehensive simulation results show that shifted declustering is superiour in performance and load balancing to traditional replication layout schemes such as standard mirroring, chained declustering, group rotational declustering and existing parity layout schemes PRIME and RELPR in reference [4]. Huijun Zhu, Jun Wang 0001 |
ICS | 3 |
| 2008 | Energy-efficient high-performance storage systemabstractThis project is developing extended versions of RAIDs with low-power, and exploring novel energy-efficient disk array architectures using coding techniques for data intensive computing. Jun Wang 0001 |
IPDPS | 1 |
| 2008 | Exploiting In-Memory and On-Disk Redundancy to Conserve Energy in Storage SystemsabstractToday's storage systems place an imperative demand on energy efficiency. A storage system often places single-rotation- rate disks into standby mode by stopping them from spinning to conserve energy when the workload is not heavy. The major obstacle of this method is a high spin-up cost introduced by passively waking up the standby disk to service requests. In this paper, we propose a redundancy-based hierarchical I/O cache architecture called RIMAC to solve the problem. The idea of RIMAC is to enable data on the standby disk(s) to be recovered by accessing a two-level I/O cache and/or active disks if needed. In parity-based redundant disk arrays, RIMAC exploits parity redundancy to dynamically XOR-reconstruct the data being accessed toward standby disk(s) at both the cache and disk levels. By avoiding passive spin-ups, RIMAC can significantly improve both energy efficiency and performance. In RIMAC, we developed 1) two power-aware read request transformation schemes called Transformable Read in Cache (TRC) and Transformable Read on Disk (TRD), 2) a power-aware write request transformation policy for parity update, and 3) a second-chance parity cache replacement algorithm to favor the request transformation rate. We evaluated RIMAC by augmenting a validated storage system simulator, DiskSim, and tested three real-life server traces, including HP's cello99, UlUC's OLTP, and SPC's search engine. Comprehensive results indicate that RIMAC is able to reduce energy consumption by up to 18 percent and simultaneously improve the average response time by up to 34 percent compared with threshold-based power management schemes for single-rotation-rate disk- based RAIDs. Jun Wang 0001, Xiaoyu Yao, Huijun Zhu |
IEEE Trans. Computers | 1 |
| 2008 | eRAID: Conserving Energy in Conventional Disk-Based RAID SystemabstractRecently, high-energy consumption has become a serious concern for both storage servers and data centers. Recent research studies have utilized the short transition times of multispeed disks to decrease energy consumption. Manufacturing challenges and costs have so far prevented commercial deployment of multispeed disks. In this paper, we propose an energy saving policy, eRAID (energy-efficient RAID), for conventional disk-based mirrored and parity redundant disk array architectures. eRAID saves energy by spinning down partial or the entire mirror disk group with constraints of acceptable performance degradation. We first develop a multiconstraint energy-saving model for the RAID environment by considering both disk characteristics and workload features. Then, we develop a performance (response time and throughput) control scheme for eRAID based on the analytical model. Experimental results show that eRAID can save up to 32 percent energy while satisfying the predefined performance requirement. Jun Wang 0001, Huijun Zhu |
IEEE Trans. Computers | 1 |
| 2008 | HBA: Distributed Metadata Management for Large Cluster-Based Storage SystemsabstractAn efficient and distributed scheme for file mapping or file lookup is critical in decentralizing metadata management within a group of metadata servers. This paper presents a novel technique called Hierarchical Bloom Filter Arrays (HBA) to map filenames to the metadata servers holding their metadata. Two levels of probabilistic arrays, namely, the Bloom filter arrays with different levels of accuracies, are used on each metadata server. One array, with lower accuracy and representing the distribution of the entire metadata, trades accuracy for significantly reduced memory overhead, whereas the other array, with higher accuracy, caches partial distribution information and exploits the temporal locality of file access patterns. Both arrays are replicated to all metadata servers to support fast local lookups. We evaluate HBA through extensive trace-driven simulations and implementation in Linux. Simulation results show our HBA design to be highly effective and efficient in improving the performance and scalability of file systems in clusters with 1,000 to 10,000 nodes (or superclusters) and with the amount of data in the petabyte scale or higher. Our implementation indicates that HBA can reduce the metadata operation time of a single-metadata-server architecture by a factor of up to 43.9 when the system is configured with 16 metadata servers. Hong Jiang 0001, Jun Wang 0001, Feng Xian |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2007 | ASAP: An Advertisement-based Search Algorithm for Unstructured Peer-to-peer SystemsabstractMost of existing search algorithms for unstructured peer-to-peer (P2P) systems share one common approach: the requesting node sends out a query and the query message is repeatedly routed and forwarded to other peers in the overlay network. Due to multiple hops involved in query forwarding, the search may result in a long delay before it is answered. Furthermore, some incapable nodes may be easily overloaded when the query traffic becomes intensive or bursty. In this paper, we present a novel content-pushing, Advertisement-based search algorithm for unstructured P2P systems called ASAP. An advertisement (ad in brief) is a synopsis of contents a peer tends to share, and appropriately distributed and selectively cached by other peers in the system. In ASAP, nodes proactively advertise their contents by delivering ads, and selectively store interesting ads received from other peers. Upon a request, a node can locate the destination nodes by looking up its local ads repository, and thus obtain a one-hop search latency with modest search cost. Comprehensive experimental results show that, compared with traditional query-based search algorithms, ASAP achieves much better search efficiency, and maintains system load at a low level with small variances. In addition, ASAP works well under node churn. Jun Wang 0001, Hailong Cai |
ICPP | 2 |
| 2006 | Nexus: A Novel Weighted-Graph-Based Prefetching Algorithm for Metadata Servers in Petabyte-Scale Storage SystemsabstractAn efficient, accurate and distributed metadata-oriented prefetching scheme is critical to the overall performance in large distributed storage systems. In this paper, we present a novel weighted-graph-based prefetching technique, built on successor relationship, to gain performance benefit from prefetching specifically for clustered metadata servers, an arrangement envisioned necessary for petabyte-scale distributed storage systems. Extensive trace-driven simulations show that by adopting our new prefetching algorithm, the hit rate for metadata access on the client site can be increased by up to 13%, while the average response time of metadata operations can be reduced by up to 67%, compared with LRU and an existing state of the art prefetching algorithm. Hong Jiang 0001, Jun Wang 0001 |
CCGRID | 4 |
| 2006 | RIMAC: a novel redundancy-based hierarchical cache architecture for energy efficient, high performance storage systemsabstractEnergy efficiency becomes increasingly important in today's high-performance storage systems. It can be challenging to save energy and improve performance at the same time in conventional (i.e. single-rotation-rate) disk-based storage systems. Most existing solutions compromise performance for energy conservation. In this paper, we propose a redundancy-based, two-level I/O cache architecture called RIMAC to address this problem. The idea of RIMAC is to enable data on the standby disk to be recovered by accessing data in the two-level I/O cache or on currently active/idle disks. At both cache and disk levels, RIMAC dynamically transforms accesses toward standby disks by exploiting parity redundancy in parity-based redundant disk arrays. Because I/O requests that require physical accesses on standby disks involve long waiting time and high power consumption for disk spin-up (tens of seconds for SCSI disks), transforming those requests to accesses in a two-level, collaborative I/O cache or on active disks can significantly improve both energy efficiency and performance.In RIMAC, we developed i) two power-aware read request transformation schemes called Transformable Read in Cache (TRC) and Transformable Read on Disk (TRD), ii) a power-aware write request transformation policy for parity update and iii) a second-chance parity cache replacement algorithm to improve request transformation rate. We evaluated RIMAC by augmenting a validated storage system simulator, disksim. For several real-life server traces including HP's cello 99, TPC-D and SPC's search engine, RIMAC is shown to reduce energy consumption by up to 33% and simultaneously improve the average response time by up to 30%. Xiaoyu Yao, Jun Wang 0001 |
EuroSys | 2 |
| 2006 | Alliatrust: A Trustable Reputation Management Scheme for Unstructured P2P Systems
Jeffrey Gerard, Hailong Cai, Jun Wang 0001 |
GPC | 3 |
| 2006 | eRAID: A Queueing Model Based Energy Saving PolicyabstractRecently energy consumption becomes an ever critical concern for both low-end and high-end storage servers. In this paper, we propose an energy saving policy, eRAID (energy-efficient RAID), for mirrored redundant disk array systems. eRAID saves energy by spinning down partial or entire mirror disk group with controllable performance degradation. We first develop an energy-saving model for multi-disk environments by taking into account both disk characteristics and workload features. Then, we develop a queueing model based performance (response time and throughput) control scheme for eRAID. Experimental results show that eRAID can save up to 32% energy without violating predefined performance degradation constraints. Jun Wang 0001 |
MASCOTS | 2 |
| 2006 | A light-weight, collaborative temporary file system for clustered Web servers
Jun Wang 0001 |
J. Parallel Distributed Comput. | 1 |
| 2006 | Exploiting Geographical and Temporal Locality to Boost Search Efficiency in Peer-to-Peer SystemsabstractAs a hot research topic, many search algorithms have been presented and studied for unstructured peer-to-peer (P2P) systems during the past few years. Unfortunately, current approaches either cannot yield good lookup performance, or incur high search cost and system maintenance overhead. The poor search efficiency of these approaches may seriously limit the scalability of current unstructured P2P systems. In this paper, we propose to exploit two-dimensional locality to improve P2P system search efficiency. We present a locality-aware P2P system architecture called Foreseer, which explicitly exploits geographical locality and temporal locality by constructing a neighbor overlay and a friend overlay, respectively. Each peer in Foreseer maintains a small number of neighbors and friends along with their content filters used as distributed indices. By combining the advantages of distributed indices and the utilization of two-dimensional locality, our scheme significantly boosts P2P search efficiency while introducing only modest overhead. In addition, several alternative forwarding policies of Foreseer search algorithm are studied in depth on how to fully exploit the two-dimensional locality Hailong Cai, Jun Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2005 | Caching Routing Indices in Structured P2P OverlaysabstractBecause of the omnipresence of node dynamic activities, large scale P2P systems built on structured overlays suffer high maintenance overhead and compromised routing performance. In this paper, we study the characteristics of P2P node dynamic behaviors and present a novel routing indices caching scheme, called SORIC, which solves this problem by fully exploiting the round-trip pattern in node dynamic behaviors and heterogeneity among peers in the system. SORIC selectively caches routing indices of transient departed nodes in other relatively stable and capable nodes for two purposes. First, rejoin of the cached nodes is drastically simplified to O(1) complexity, thus cutting off a large portion of system maintenance overhead. Second, caching routing indices of departed nodes minimizes the negative effects of node departures and rejoins, and thus enables the system to sustain an uninterruptedly high quality routing service. Hailong Cai, Jun Wang 0001 |
ICPP | 2 |
| 2005 | Toward Effective NIC Caching: A Hierarchical Data Cache Architecture for iSCSI Storage ServersabstractIn this paper, we present a hierarchical data cache architecture called DCA to effectively slash local interconnect traffic and thus boost the storage server performance. DCA is composed of a read cache in NIC card called NIC cache and a read/write unify cache in host memory called Helper cache. NIC cache services most portion of read requests without fetching data via PCI bus, while Helper cache 1) supplies some portions of read requests given partial NIC cache hits; 2) directs cache placement for NIC cache and 3) absorbs most transient writes locally. We developed a novel State Locality Aware cache Placement algorithm called SLAP to improve NIC cache hit ratio for mixed read and write workloads. To demonstrate the effectiveness of DCA, we developed a DCA prototype system and evaluated it with open source iSCSI implementation under representative storage server workloads. Experimental results showed that DCA can boost iSCSI storage server throughput by up to 121% and slash the PCI traffic by up to 74% compared with an iSCSI target without DCA. Xiaoyu Yao, Jun Wang 0001 |
ICPP | 2 |
| 2005 | Achieving Stability and Fairness in Mobile Ad Hoc Networks
Ahmed M. Mahdy, Jitender S. Deogun, Jun Wang 0001 |
NETWORKING | 3 |
| 2005 | A novel state cache scheme in structured P2P systems
Hailong Cai, Jun Wang 0001, Jitender S. Deogun |
J. Parallel Distributed Comput. | 2 |
| 2004 | Hierarchical Bloom filter arrays (HBA): a novel, scalable metadata management system for large cluster-based storageabstractAn efficient and distributed scheme for file mapping or file lookup scheme is critical in decentralizing metadata management within a group of metadata servers. This work presents a technique called HBA (hierarchical Bloom filter arrays) to map file names to the servers holding their metadata. Two levels of probabilistic arrays, i.e., Bloom filter arrays, with different accuracies are used on each metadata server. One array, with lower accuracy and representing the distribution of the entire metadata, trades accuracy for significantly reduced memory overhead, while the other array, with higher accuracy, caches partial distribution information and exploits the temporal locality of file access patterns. Extensive trace-driven simulations have shown our HBA design to be highly effective and efficient in improving performance and scalability of file systems in clusters with 1,000 to 10,000 nodes (or superclusters). Hong Jiang 0001, Jun Wang 0001 |
CLUSTER | 3 |
| 2004 | A performance study on Internet Server Provider mail serversabstractThis work presents a comprehensive performance study on Internet Service Provider (ISP) mail server, which plays an important role in Internet-based distributed computing. By feeding the SPECmail2001 benchmark into an ISP mail server testbed set up by a commercial mail server software system - MDaemon 5.0.1, we study both networking and I/O performance by varying its user population from 200 to 10,000. The benchmark utilities are adopted for networking analysis while file system traces are collected for I/O measurement. Based on the benchmark study and offline trace analysis, we arrive at several important conclusions for ISP mail servers and give corresponding technical suggestions to improve the performance. First we observed that, in SMTP and POP sessions, the initial network connection setup step usually takes a very long time. Second, I/O latencies typically contribute to 40-55% of the total data transfer time in e-mail requests, especially in a server with large user population support. Third, A group of e-mail messages will easily make a remote recipient server become overloaded. Jun Wang 0001, Yiming Hu |
ISCC | 1 |
| 2004 | Foreseer: A Novel, Locality-Aware Peer-to-Peer System Architecture for Keyword Searches
Hailong Cai, Jun Wang 0001 |
Middleware | 2 |
| 2003 | CTFS: A New Light-Weight, Cooperative Temporary File System for Cluster-Based Web ServersabstractPrevious studies showed that I/O could become a major performance bottleneck in cluster-based Web servers. Adopting a large I/O buffer cache on separate server nodes is not a good performance-cost scheme and sometime infeasible because of the high price and poor reliability. Current native file systems do not work well for the poor performance. Specialized file systems suffer a poor portability problem. In this paper, we present a new light-weight, cooperative temporary file system (called CTFS) to boost I/O performance for cluster-based Web servers. CTFS has the following advantages: (a) consists of a peer-to-peer cooperative caching system using user-level communication technique to eliminate repeated file requests and conduct aggressive remote prefetch; (b) runs in the user space to achieve a good portability; and (c) organizes a group of files with good associated access locality together to form a cluster unit on disk and thereby providing a sustained high I/O performance without degradation. Comprehensive trace-driven simulation experiments show that CTFS achieves up to a 37% better entire system throughput and reduces up to 47% total disk I/O latency than those in asynchronous FFS for a 64 node cluster-based Web server. Jun Wang 0001 |
CLUSTER | 1 |
| 2003 | A Novel Reordering Write Buffer to Improve Write Performance of Log-Structured File SystemsabstractWe present a novel reordering write buffer which improves the performance of log-structured file systems (LFS). While UFS has a good write performance, high garbage-collection overhead degrades its performance under high disk space utilization. Previous research concentrated on how to improve the efficiency of the garbage collector after data is written to disk. We propose a new method that reduces the amount of work the garbage collector would do before data reaches disk. By classifying active and inactive data in memory into different segment buffers and then writing them to different disk segments, we force the disk segments to form a bimodal distribution. Most data blocks in active segments are quickly invalidated, while inactive segments remain mostly intact. Simulation results based on a wide range of both real-world and synthetic traces show that our method significantly reduces the garbage collection overhead, slashing the overall write cost of LFS by up to 53 percent, improving the write performance of LFS by up to 26 percent, and the overall system performance by up to 21 percent. Jun Wang 0001, Yiming Hu |
IEEE Trans. Computers | 1 |
| 2002 | WOLF - A Novel Reordering Write Buffer to Boost the Performance of Log-Structured File Systems
Jun Wang 0001, Yiming Hu |
FAST | 1 |
| 2002 | UCFS-A Novel User-Space, High Performance, Customized File System for Web Proxy ServersabstractWeb proxy caching servers play a key role in today's Web infrastructure. Previous studies have shown that disk I/O is one of the major performance bottlenecks of proxy servers. Most conventional file systems do not work well for proxy server workloads and have high overheads. This paper presents a novel, User-space, Customized File System, called UCFS, that can drastically improve the I/O performance of proxy servers. UCFS is a user-level software component of a proxy server which manages data on a raw disk or disk partition. Since the entire system runs in the user space, it is easy and inexpensive to implement. It also has good portability and maintainability. UCFS uses efficient in-memory meta-data tables to eliminate almost all I/O overhead of meta-data searches and updates. It also includes a novel file system called Cluster-structured File System (CFS). Similarly to the Log-structured File Systems (LFS), CFS uses large disk transfers to significantly improve disk write performance. However, CFS can also markedly improve file read operations and it does not generate garbage. Comprehensive simulation experiments using five representative real-world traces show that UCFS can significantly improve proxy server performance. For example, UCFS achieves 8-19 times better I/O performance than the state-of-the-art SQUID server running on a Unix Fast File System (FFS), 4-7.5 times better than SQUID on asynchronous FFS, and 3-9 times better than the improved SQUIDML. Jun Wang 0001, Yingwu Zhu, Yiming Hu |
IEEE Trans. Computers | 1 |