EDBT 2026 Demo / reviewers in the wild / expert
Jianfeng Zhan
dblp:12/2485
· DBLP profile ↗
72ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-3728-6837ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 since 2021Artificial intelligence and machine learning · 9 · 4 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Software engineering, systems software and programming languages · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Security and privacy · 2Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TimeMosaic: Temporal Heterogeneity Guided Time Series Forecasting via Adaptive Granularity Patch and Segment-wise DecodingabstractMultivariate time series forecasting is essential in domains such as finance, transportation, climate, and energy. However, existing patch-based methods typically adopt fixed-length segmentation, overlooking the heterogeneity of local temporal dynamics and the decoding heterogeneity of forecasting. Such designs lose details in information-dense regions, introduce redundancy in stable segments, and fail to capture the distinct complexities of short-term and long-term horizons. We propose TimeMosaic, a forecasting framework that aims to address temporal heterogeneity. TimeMosaic employs adaptive patch embedding to dynamically adjust granularity according to local information density, balancing motif reuse with structural clarity while preserving temporal continuity. In addition, it introduces segment-wise decoding that treats each prediction horizon as a related subtask and adapts to horizon-specific difficulty and information requirements, rather than applying a single uniform decoder. Extensive evaluations on benchmark datasets demonstrate that TimeMosaic delivers consistent improvements over existing methods, and our model trained on the large-scale corpus with 321 billion observations achieves performance competitive with state-of-the-art TSFMs. Kuiye Ding, Fanda Fan, Chunyi Hou, Zheya Wang, Lei Wang 0004, Zhengxin Yang, Jianfeng Zhan |
AAAI | 7 |
| 2025 | DualSG: A Dual-Stream Explicit Semantic-Guided Multivariate Time Series Forecasting FrameworkabstractMultivariate Time Series Forecasting plays a key role in many applications. Recent works have explored using Large Language Models for MTSF to take advantage of their reasoning abilities. However, many methods treat LLMs as end-to-end forecasters, which often leads to a loss of numerical precision and forces LLMs to handle patterns beyond their intended design. Alternatively, methods that attempt to align textual and time series modalities within latent space frequently encounter alignment difficulty. In this paper, we propose to treat LLMs not as standalone forecasters, but as semantic guidance modules within a dual-stream framework. We propose DualSG, a dual-stream framework that provides explicit semantic guidance, where LLMs act as Semantic Guides to refine rather than replace traditional predictions. As part of DualSG, we introduce Time Series Caption, an explicit prompt format that summarizes trend patterns in natural language and provides interpretable context for LLMs, rather than relying on implicit alignment between text and time series in the latent space. We also design a caption-guided fusion module that explicitly models inter-variable relationships while reducing noise and computation. Experiments on real-world datasets from diverse domains show that DualSG consistently outperforms 15 state-of-the-art baselines, demonstrating the value of explicitly combining numerical forecasting with semantic guidance. Kuiye Ding, Fanda Fan, Ruijie Jian, Luqi Gong, Yishan Jiang, Chunjie Luo, Jianfeng Zhan |
ACM Multimedia | 9 |
| 2025 | OpenClinicalAI: An open and dynamic model for Alzheimer's Disease diagnosis
Yunyou Huang, Xiaoshuang Liang, Jiyue Xie, Xiangjiang Lu, Xiuxia Miao, Fan Zhang 0047, Guoxin Kang, Suqin Tang, Jianfeng Zhan |
Expert Syst. Appl. | 11 |
| 2025 | BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector DatabasesabstractVector databases are designed to effectively store, organize, and retrieve high-dimensional vectors, enabling faster and more accurate querying and analysis. This study highlights that the performance of cutting-edge vector databases hinges on their proficiency in managing heterogeneous data embedding and handling compound queries. The former task revolves around converting varied data types into a cohesive vector format, while the latter involves processing multimodal or single-modal queries with precise constraints. The paper advocates for evaluating these dual tasks within an integrated benchmark framework. However, state-of-the-art vector database benchmarks overlook heterogeneous data embedding and compound queries, creating a gap in evaluating vector database performance. To address this gap, we introduce BigVectorBench, a benchmark suite designed to evaluate vector database performance. BigVectorBench contributes by defining and evaluating the embedding performance of heterogeneous data. Additionally, it abstracts compound queries, which are increasingly used in real-world applications, replacing unimodal vector searches. Our rigorous evaluations validate the two design decisions of BigVectorBench and identify performance bottlenecks of mainstream vector databases. Its source code and user manual are available from https://github.com/BenchCouncil/BigVectorBench. Guoxin Kang, Zhongxin Ge, Jingpei Hu, Xueya Zhang, Lei Wang 0004, Jianfeng Zhan |
Proc. VLDB Endow. | 6 |
| 2024 | IterLara: A Concise General-Purpose Algebraic Model
Hongxiao Li, Wanling Gao, Lei Wang 0004, Jianfeng Zhan |
COCOON (2) | 4 |
| 2024 | Deep Adaptive Graph Clustering via von Mises-Fisher DistributionsabstractGraph clustering has been a hot research topic and is widely used in many fields, such as community detection in social networks. Lots of works combining auto-encoder and graph neural networks have been applied to clustering tasks by utilizing node attributes and graph structure. These works usually assumed the inherent parameters (i.e., size and variance) of different clusters in the latent embedding space are homogeneous, and hence the assigned probability is monotonous over the Euclidean distance between node embeddings and centroids. Unfortunately, this assumption usually does not hold since the size and concentration of different clusters can be quite different, which limits the clustering accuracy. In addition, the node embeddings in deep graph clustering methods are usually L2 normalized so that it lies on the surface of a unit hyper-sphere. To solve this problem, we proposed D eep A daptive G raph C lustering via von Mises-Fisher distributions, namely DAGC. DAGC assumes the node embeddings H can be drawn from a von Mises-Fisher distribution and each cluster k is associated with cluster inherent parameters ρ k which includes cluster center μ and cluster cohesion degree κ. Then we adopt an EM-like approach (i.e., 𝒫( H | ρ ) and 𝒫( ρ | H ), respectively) to learn the embedding and cluster inherent parameters alternately. Specifically, with the node embeddings, we proposed to update the cluster centers in an attraction-repulsion manner to make the cluster centers more separable. And given the cluster inherent parameters, a likelihood-based loss is proposed to make node embeddings more concentrated around cluster centers. Thus, DAGC can simultaneously improve the intra-cluster compactness and inter-cluster heterogeneity. Finally, extensive experiments conducted on four benchmark datasets have demonstrated that the proposed DAGC consistently outperforms the state-of-the-art methods, especially on imbalanced datasets. Pengfei Wang 0008, Daqing Wu, Chong Chen 0002, Kunpeng Liu 0001, Yanjie Fu, Jianqiang Huang 0001, Yuanchun Zhou, Jianfeng Zhan, Xian-Sheng Hua 0001 |
ACM Trans. Web | 8 |
| 2023 | CMLCompiler: A Unified Compiler for Classical Machine LearningabstractClassical machine learning (CML) occupies nearly half of machine learning pipelines in production applications. Unfortunately, it fails to utilize the state-of-the-practice devices fully and performs poorly. Without a unified framework, the hybrid deployments of deep learning (DL) and CML also suffer from severe performance and portability issues. This paper presents the design of a unified compiler, called CMLCompiler, for CML inference. We propose two unified abstractions: operator representations and extended computational graphs. The CMLCompiler framework performs the conversion and graph optimization based on two unified abstractions, then outputs an optimized computational graph to DL compilers or frameworks. We implement CMLCompiler on TVM. The evaluation shows CMLCompiler's portability and superior performance. It achieves up to 4.38× speedup on CPU, 3.31× speedup on GPU, and 5.09× speedup on IoT devices, compared to the state-of-the-art solutions --- scikit-learn, intel sklearn, and hummingbird. Our performance of CML and DL mixed pipelines achieves up to 3.04x speedup compared with cross-framework implementations. The project documents and source code are available at https://www.computercouncil.org/cmlcompiler. Wanling Gao, Anzheng Li, Lei Wang 0004, Zihan Jiang 0006, Jianfeng Zhan |
ICS | 6 |
| 2023 | Hierarchical Masked 3D Diffusion Model for Video OutpaintingabstractVideo outpainting aims to complete missing areas at the edges of video frames adequately. Compared to image outpainting, it presents an additional challenge as the model should maintain the temporal consistency of the filled area. In this paper, we introduce a masked 3D diffusion model for video outpainting. We use the technique of mask modeling to train the 3D diffusion model. This allows us to use multiple guide frames to connect the results of multiple video clip inferences, thus ensuring temporal consistency and reducing jitter between adjacent frames. Meanwhile, we extract the global frames of the video as prompts and guide the model to obtain information other than the current video clip using cross-attention. We also introduce a hybrid coarse-to-fine inference pipeline to alleviate the artifact accumulation problem. The existing coarse-to-fine pipeline only uses the infilling strategy, which brings degradation because the time interval of the sparse frames is too large. Our pipeline benefits from bidirectional learning of the mask modeling and thus can employ a hybrid strategy of infilling and interpolation when generating sparse frames. Experiments show that our method achieves state-of-art results in video outpainting tasks. Fanda Fan, Chaoxu Guo, Litong Gong, Tiezheng Ge, Yuning Jiang 0001, Chunjie Luo, Jianfeng Zhan |
ACM Multimedia | 8 |
| 2023 | Scenario-Based AI Benchmark Evaluation of Distributed Cloud/Edge Computing SystemsabstractDistributed cloud/edge (DCE) platform has become popular in recent years. This paper proposes a new AI benchmark suite for assessing the performance of DCE platforms in machine learning (ML) and cognitive science applications. The benchmark suite is custom-designed to satisfy scenario-based performance requirements, namely the model training time, inference speed, model accuracy, job response time, quality of service, and system reliability. These metrics are substantiated by intensive experiments with real-life AI workloads. Our work is specially tailored for supporting massive AI multitasking across distributed resources in the networking environment. Our benchmark experiments were conducted on an AI-oriented AIRS cloud built at the Chinese University of Hong Kong, Shenzhen. We have tested a large number of ML/DL programs to narrow down the inclusion of ten representative AI kernel codes in the benchmark suite. Our benchmark results reveal the advantages of using the DCE systems cost-effectively in smart cities, healthcare, community surveillance, and transportation services. Our technical contributions are in the AIRS cloud architecture, benchmark design, testing, and distributed AI computing requirements. Our work will benefit computer system designers and AI application developers on clouds, edge, and mobile devices, that are supported by 5G mobile networks and AIoT resources. Tianshu Hao, Kai Hwang 0001, Jianfeng Zhan, Yuejin Li, Yong Cao 0001 |
IEEE Trans. Computers | 3 |
| 2022 | OLxPBench: Real-time, Semantically Consistent, and Domain-specific are Essential in Benchmarking, Designing, and Implementing HTAP SystemsabstractAs real-time analysis on the fresh data become in-creasingly compelling, more organizations deploy Hybrid Trans-actional/Analytical Processing (HTAP) systems to support real-time queries on data recently generated by online transaction processing. This paper argues that real-time queries, semantically consistent schema, and domain-specific workloads are essential in benchmarking, designing, and implementing HTAP systems. However, most state-of-the-art and state-of-the-practice bench-marks ignore those critical factors. Hence, at best, they are incommensurable and, at worst, misleading in benchmarking, designing, and implementing HTAP systems. This paper presents OLxPBench, a composite HTAP benchmark suite. OLxPBench proposes: (1) the abstraction of a hybrid transaction, performing a real-time query in-between an online transaction, to model widely-observed behavior pattern - making a quick decision while consulting real-time analysis; (2) a semantically consistent schema to express the relationships between OLTP and OLAP schema; (3) the combination of domain-specific and general benchmarks to characterize diverse application scenarios with varying resource demands. Our evaluations justify the three design decisions of OLxPBench and pinpoint the bottlenecks of two mainstream distributed HTAP DBMSs. International Open Benchmark Council (Bench Council) sets up the OLxP-Bench homepage at https://www.benchcouncil.org/olxpbench/. Its source code is available from https://github.com/BenchCouncil/olxpbench.git. Guoxin Kang, Lei Wang 0004, Wanling Gao, Fei Tang 0003, Jianfeng Zhan |
ICDE | 5 |
| 2022 | DeepCS: Training a deep learning model for cervical spondylosis recognition on small-labeled sensor data
Nana Wang 0001, Chunjie Luo, Yunyou Huang, Jianfeng Zhan |
Neurocomputing | 5 |
| 2021 | AIBench Scenario: Scenario-Distilling AI BenchmarkingabstractModern real-world application scenarios like Internet services consist of a diversity of AI and non-AI modules with huge code sizes and long and complicated execution paths, which raises serious benchmarking or evaluating challenges. Using AI components or micro benchmarks alone can lead to error-prone conclusions. This paper presents a methodology to attack the above challenge. We formalize a real-world application scenario as a Directed Acyclic Graph-based model and propose the rules to distill it into a permutation of essential AI and non-AI tasks, which we call a scenario benchmark. Together with seventeen industry partners, we extract nine typical scenario benchmarks. We design and implement an extensible, configurable, and flexible benchmark framework. We implement two Internet service AI scenario benchmarks based on the framework as proxies to two real-world application scenarios. We consider scenario, component, and micro benchmarks as three indispensable parts for evaluating. Our evaluation shows the advantage of our methodology against using component or micro AI benchmarks alone. The specifications, source code11Zenodo: https://doi.org/10.5281/zenodo.5158715 GitHub: https://github.com/BenchCouncil/aibench_scenario, testbed, and results are publicly available from https://www.benchcouncil.org/aibench/scenario/. Wanling Gao, Fei Tang 0003, Jianfeng Zhan, Lei Wang 0004, Zheng Cao 0003, Chuanxin Lan, Chunjie Luo, Xiaoli Liu 0002, Zihan Jiang 0006 |
PACT | 3 |
| 2021 | Ensemble Clustering-based Cervical Spondylosis Fine-classificationabstractState-of-the-art and state-of-the-practice clinical work fail to fine-classify early cervical spondylosis (CS) illness and hence can not provide personalized treatment. Surface electromyography (sEMG) is an essential physiological signal that can locate muscle problems associated with cervical spine dysfunction, and it is promising to provide sEMG-based CS fine-classification. However, the state-of-the-art clustering approaches on sEMG data have the following drawbacks: (1) the clustering results are unstable due to different algorithms or multiple values of a resolution parameter. (2) The samples of the same individual in the same state are clustered in different groups. This paper proposes an individual-centered three-layer ensemble clustering framework (in short, ITECF) to fine-classify cervical spondylosis (CS). ITECF mainly consists of three parts: base clusters generation, individual consensus, base cluster consensus. We define an individual consistency index named ICS to measure the consistency of samples of the same individual in the same partition and propose an ICS-based consensus to transform the sample-centered base clusters into individual-centered ones for partition selection. We evaluate our approaches against five clustering algorithms and five ensemble clustering frameworks on the three real-world sEMG data set using four metrics: Silhouette Coefficient (SC), Davies Bouldin score (DB), Calinski-Harabaz Index (CH), and ICS. The results show that ITECF outperforms the state-of-the-art algorithms. ITECF obtains eight fine-classifications of CS for the first time, each with different muscle pathological problems in muscle coordination, balance, and strength. And, this fine-classification has the potential for personalized treatment. Nana Wang 0001, Chunjie Luo, Yunyou Huang, Jianfeng Zhan |
BIBM | 4 |
| 2021 | AI-oriented Workload Allocation for Cloud-Edge ComputingabstractDifferent placement or collaboration policies in handling datasets and workloads across cloud, edge, and user-end may substantially affect a cloud-edge computing environment's overall performance. However, the common practice is to optimize the performance only on the edge layer, while ignoring the rest of the system. This paper calls attention to optimize AI-oriented workloads' performance across all components in cloud-edge architectures holistically. Our goal is to optimize AI-workload allocation in cloud clusters, edge servers, and end devices, achieving the minimum response time in latency-sensitive applications. This paper presents new workload allocation methods for AI workloads in cloud-edge computing systems. We have proposed two efficient allocation algorithms to reduce the end-to-end response time of single-workload and multi-jobs scenarios, respectively. We apply six edge AI workloads from a comprehensive edge computing benchmark - Edge AIBench for experiments. Besides, we conduct experiments in a real edge computing environment. Our experiment results demonstrate the high efficiency and effectiveness of our algorithms in real-life applications and datasets. Our multi-job allocation algorithm's end-to-end response time outperforms the other four baseline strategies by 33% to 63%. Tianshu Hao, Jianfeng Zhan, Kai Hwang 0001, Wanling Gao |
CCGRID | 2 |
| 2021 | HPC AI500 V2.0: The Methodology, Tools, and Metrics for Benchmarking HPC AI SystemsabstractRecent years witness a trend of applying large-scale distributed deep learning algorithms (HPC AI) in both business and scientific computing areas, whose goal is to speed up the training time to achieve a state-of-the-art quality. The HPC AI benchmarks accelerate the process. Unfortunately, benchmarking HPC AI systems at scale raises serious challenges. This paper presents a comprehensive HPC AI benchmarking methodology that achieves equivalence, representativeness, repeatability, and affordability. Among the nineteen AI workloads of AIBench Training–by far the most comprehensive AI benchmarks suite, we choose two representative and repeatable AI workloads in terms of both AI model and micro-architectural characteristics. The selected HPC AI benchmarks include both business and scientific computing: Image Classification and Extreme Weather Analytics. Finally, we propose three high levels of benchmarking and the corresponding rules to assure equivalence. To rank the performance of HPC AI systems, we present a new metric named Valid FLOPS, emphasizing both throughput performance and target quality. The evaluations show our methodology, benchmarks, and metrics can measure and rank the HPC AI systems in a simple, affordable and repeatable way. The specification, source code, datasets, and HPC AI500 ranking numbers are publicly available from https://www.benchcouncil.org/aibench/hpcai500/index.html. Zihan Jiang 0006, Wanling Gao, Fei Tang 0003, Lei Wang 0004, Xingwang Xiong, Chunjie Luo, Chuanxin Lan, Hongxiao Li, Jianfeng Zhan |
CLUSTER | 9 |
| 2021 | Finet: Using Fine-grained Batch Normalization to Train Light-weight Neural NetworksabstractTo build light-weight network, we propose a new normalization, Fine-grained Batch Normalization (FBN). Different from Batch Normalization (BN), which normalizes the final summation of the weighted inputs, FBN normalizes the intermediate state of the summation. We propose a novel lightweight network based on FBN, called Finet. At training time, the convolutional layer with FBN can be seen as an inverted bottleneck mechanism. FBN can be fused into convolution at inference time. After fusion, Finet uses the standard convolution with equal channel width, thus makes the inference more efficient. On ImageNet classification dataset, Finet achieves the state-of-art performance (65.706 % accuracy with 43M FLOPs, and 73.786% accuracy with 303M FLOPs), Moreover, experiments show that Finet is more efficient than other state-of-art lightweight networks. Chunjie Luo, Jianfeng Zhan, Lei Wang 0004, Wanling Gao |
IJCNN | 2 |
| 2021 | AIBench Training: Balanced Industry-Standard AI Training BenchmarkingabstractEarlier-stage evaluations of a new AI architecture/system need affordable AI benchmarks. Only using a few AI component benchmarks like MLPerf alone in the other stages may lead to misleading conclusions. Moreover, the learning dynamics are not well understood, and the benchmarks' shelf-life is short. This paper proposes a balanced benchmarking methodology. We use real-world benchmarks to cover the factors space that impacts the learning dynamics to the most considerable extent. After performing an exhaustive survey on Internet service AI domains, we identify and implement nineteen representative AI tasks with state-of-the-art models. For repeatable performance ranking (RPR subset) and workload characterization (WC subset), we keep two subsets to a minimum for affordability. We contribute by far the most comprehensive AI training benchmark suite. The evaluations show: (1) AIBench Training (v1.1) outperforms MLPerf Training (v0.7) in terms of diversity and representativeness of model complexity, computational cost, convergent rate, computation, and memory access patterns, and hotspot functions; (2) Against the AIBench full benchmarks, its RPR subset shortens the benchmarking cost by 64%, while maintaining the primary workload characteristics; (3) The performance ranking shows the single-purpose AI accelerator like TPU with the optimized TensorFlow framework performs better than that of GPUs while losing the latter's general support for various AI models. The specification, source code, and performance numbers are available from the AIBench homepage https://www.benchcouncil.org/aibench-training/index.html. Fei Tang 0003, Wanling Gao, Jianfeng Zhan, Chuanxin Lan, Lei Wang 0004, Chunjie Luo, Zheng Cao 0003, Xingwang Xiong, Zihan Jiang 0006, Tianshu Hao, Fanda Fan, Fan Zhang 0047, Yunyou Huang, Jianan Chen 0003, Mengjia Du, Chen Zheng 0001, Daoyi Zheng, Haoning Tang, Kunlin Zhan, Defei Kong, Chongkang Tan, Xinhui Tian, Yatao Li, Junchao Shao, Xiaoyu Wang 0002, Jiahui Dai, Hainan Ye |
ISPASS | 3 |
| 2019 | HybridTune: Spatio-Temporal Performance Data Correlation for Performance Diagnosis of Big Data Systems
Jiechao Cheng, Xiwen He, Lei Wang 0004, Jianfeng Zhan, Wanling Gao, Chunjie Luo |
J. Comput. Sci. Technol. | 5 |
| 2019 | Understanding Processors Design Decisions for Data Analytics in Homogeneous Data CentersabstractOur global economy increasingly depends on our ability to gather, analyze, link, and compare very large data sets. Keeping up with such big data poses challenges in terms of both computational performance and energy efficiency, and motivates different approaches to explore data center systems and architectures. To better understand the processor design decisions in context of data analytics in data centers, we conduct comprehensive evaluations using representative data analaytics workloads on representative conventional multi-core and many-core processors. After a comprehensive analysis of performance, power, energy efficiency and performance-cost efficiency, we have the following observations: contrasted with the conventional wisdom that uses wimpy many-core processors to improve energy-efficiency, the brawny multi-core processors with SMT (simultaneous multithreading) and dynamic overclocking technologies outperform the counterparts in terms of not only execution time, but also energy-efficiency for most of data analytics workloads in our experiments. Zhen Jia 0001, Wanling Gao, Yingjie Shi, Sally A. McKee, Zhenyan Ji, Jianfeng Zhan, Lei Wang 0004, Lixin Zhang 0002 |
IEEE Trans. Big Data | 6 |
| 2019 | Workload-Adaptive Configuration Tuning for Hierarchical Cloud SchedulersabstractCluster schedulers provide flexible resource sharing mechanism for best-effort cloud jobs, which occupy a majority in modern datacenters. Properly tuning a scheduler's configurations is the key to these jobs' performance because it decides how to allocate resources among them. Today's cloud scheduling systems usually rely on cluster operators to set the configuration and thus overlook the potential performance improvement through optimally configuring the scheduler according to the heterogeneous and dynamic cloud workloads. In this paper, we introduce AdaptiveConfig, a run-time configurator for cluster schedulers that automatically adapts to the changing workload and resource status in two steps. First, a comparison approach estimates jobs' performances under different configurations and diverse scheduling scenarios. The key idea here is to transform a scheduler's resource allocation mechanism and their variable influence factors (configurations, scheduling constraints, available resources, and workload status) into business rules and facts in a rule engine, thereby reasoning about these correlated factors in job performance comparison. Second, a workload-adaptive optimizer transforms the cluster-level searching of huge configuration space into an equivalent dynamic programming problem that can be efficiently solved at scale. We implement AdaptiveConfig on the popular YARN Capacity and Fair schedulers and demonstrate its effectiveness using real-world Facebook and Google workloads, i.e., successfully finding best configurations for most of scheduling scenarios and considerably reducing latencies by a factor of two with low optimization time. Rui Han 0001, Chi Harold Liu, Zan Zong, Lydia Y. Chen, Wending Liu, Jianfeng Zhan |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2018 | Data motifs: a lens towards fully understanding big data and AI workloadsabstractThe complexity and diversity of big data and AI workloads make understanding them difficult and challenging. This paper proposes a new approachto modelling and characterizing big data and AI workloads. We consider each big data and AI workload as a pipeline of one or more classes of units of computation performed on different initial or intermediate data inputs. Each class of unit of computation captures the common requirements while being reasonably divorced from individual implementations, and hence we call it a data motif. For the first time, among a wide variety of big data and AI workloads, we identify eight data motifs that take up most of the run time of those workloads, including Matrix, Sampling, Logic, Transform, Set, Graph, Sort and Statistic. We implement the eight data motifs on different software stacks as the micro benchmarks of an open-source big data and AI benchmark suite --- BigDataBench 4.0 (publicly available from http://prof.ict.ac.cn/BigDataBench), and perform comprehensive characterization of those data motifs from perspective of data sizes, types, sources, and patterns as a lens towards fully understanding big data and AI workloads. We believe the eight data motifs are promising abstractions and tools for not only big data and AI benchmarking, but also domain-specific hardware and software co-design. Wanling Gao, Jianfeng Zhan, Lei Wang 0004, Chunjie Luo, Daoyi Zheng, Fei Tang 0003, Biwei Xie, Chen Zheng 0001, Xiwen He, Hainan Ye |
PACT | 2 |
| 2018 | Deep Convolutional Neural Networks for Log Event Classification on Distributed Cluster SystemsabstractWith the widespread development of cloud computing, cluster systems are becoming increasingly complex, system logs is an universal and effective approach for automatic system management and troubleshooting. Log event classification as an effective preprocessing method for log analysis, which is helpful for system administrators to locate or predict components’ have errors or failures.In this paper, we design and implement an automatic log classification system based on deep CNN (Convolutional Neural Network) models, and take advantage of the feature engineering and learning algorithm to improve classification performance. First, in the feature engineering step, to address the problem of that the original unstructured event logs are unsuitable for numerical calculation in deep CNN models, we propose a novel and effective log preprocessing method, which include building categories dictionary libraries, filtering abundant information, generating numerical semantic feature vectors by calculating and combining the semantic similarity values for filtered log events. Additionally, in the learning step, we measure a series of deep CNN algorithms with varied hyper-parameter combinations by using standard evaluation metrics, and the results of our study reveal the advantages and potential capabilities of the proposed deep CNN models for log classification tasks on cluster systems. The optimal classification precision of our approach is 98.14%, which surpasses the popular traditional machine learning methods, and it can also be applied to other large-scale system logs with good accuracy. Just like the experiment results, different choices of learning algorithm do result in performance numbers varying, and subsequently careful feature engineering enables promoting performances, thus both of approaches contribute to best learning model finding. Jiechao Cheng, Jianfeng Zhan, Lei Wang 0004, Jinheng Li, Chunjie Luo |
IEEE BigData | 4 |
| 2018 | XOS: An Application-Defined Operating System for Datacenter ComputingabstractRapid growth of datacenter (DC) scale, urgency of cost control, increasing workload diversity, and huge software investment protection place unprecedented demands on the operating system (OS) efficiency, scalability, performance isolation, and backward-compatibility. The traditional OSes are not built to work with deep-hierarchy software stacks, large numbers of cores, tail latency guarantee, and increasingly rich variety of applications seen in modern DCs, and thus they struggle to meet the demands of such workloads. This paper presents XOS, an application-defined OS for modern DC servers. Our design moves resource management out of the OS kernel, supports customizable kernel subsystems in user space, and enables elastic partitioning of hardware resources. Specifically, XOS leverages modern hardware support for virtualization to move resource management functionality out of the conventional kernel and into user space, which lets applications achieve near bare-metal performance. We implement XOS on top of Linux to provide backward compatibility. XOS speeds up a set of DC workloads by up to 1.6× over our baseline Linux on a 24-core server, and outperforms the state-of-the-art Dune by up to 3.3× in terms of virtual memory management. In addition, XOS demonstrates good scalability and strong performance isolation. Chen Zheng 0001, Lei Wang 0004, Sally A. McKee, Lixin Zhang 0002, Hainan Ye, Jianfeng Zhan |
IEEE BigData | 6 |
| 2018 | CVR: efficient vectorization of SpMV on x86 processorsabstractSparse Matrix-vector Multiplication (SpMV) is an important computation kernel widely used in HPC and data centers. The irregularity of SpMV is a well-known challenge that limits SpMV’s parallelism with vectorization operations. Existing work achieves limited locality and vectorization efficiency with large preprocessing overheads. To address this issue, we present the Compressed Vectorization-oriented sparse Row (CVR), a novel SpMV representation targeting efficient vectorization. The CVR simultaneously processes multiple rows within the input matrix to increase cache efficiency and separates them into multiple SIMD lanes so as to take the advantage of vector processing units in modern processors. Our method is insensitive to the sparsity and irregularity of SpMV, and thus able to deal with various scale-free and HPC matrices. We implement and evaluate CVR on an Intel Knights Landing processor and compare it with five state-of-the-art approaches through using 58 scale-free and HPC sparse matrices. Experimental results show that CVR can achieve a speedup up to 1.70 × (1.33× on average) and a speedup up to 1.57× (1.10× on average) over the best existing approaches for scale-free and HPC sparse matrices, respectively. Moreover, CVR typically incurs the lowest preprocessing overhead compared with state-of-the-art approaches. Biwei Xie, Jianfeng Zhan, Xu Liu 0001, Wanling Gao, Zhen Jia 0001, Xiwen He, Lixin Zhang 0002 |
CGO | 2 |
| 2018 | Cosine Normalization: Using Cosine Similarity Instead of Dot Product in Neural Networks
Chunjie Luo, Jianfeng Zhan, Xiaohe Xue, Lei Wang 0004, Qiang Yang 0012 |
ICANN (1) | 2 |
| 2018 | AdaptiveConfig: Run-Time Configuration of Cluster Schedulers for Cloud Short-Running JobsabstractCluster schedulers provide flexible resource sharing mechanism for short-running jobs, which occupy a majority of cloud jobs. A scheduler's configuration decides how to allocate resources among jobs and hence it is crucial to their performances. Today's cloud platforms usually rely on cluster administrators to set this configuration, thus it is difficult to optimally configure the scheduler so as to minimize the latencies of heterogeneous and dynamically changing jobs in the cloud. In this paper, we introduce AdaptiveConfig, a run-time configurator for cluster schedulers that automatically adapts to the changing workload and resource status. This includes: (1) an estimator to calculate jobs' performances under different configurations and various scheduling scenarios. The key idea here is to transform a scheduler's resource allocation mechanisms and their variable influence factors (configuration parameters, scheduling constraints, available resources, and workload status) into business rules and facts in a rule engine, thereby reasoning about these correlated factors in job performance estimation. (2) A run-time optimizer that efficiently searches the configuration space to find the optimal configuration for the current workload. We implemented AdaptiveConfig on the popular YARN Capacity and Fair schedulers and demonstrate its effectiveness using workloads of Facebook jobs, i.e. considerably reducing latencies by 2.22 times (and up to 4.50 times) with low optimization overheads. Rui Han 0001, Zan Zong, Lydia Y. Chen, Jianfeng Zhan |
ICDCS | 5 |
| 2018 | Benchmarking Big Data Systems: A ReviewabstractWith the fast development of big data systems in recent years, a variety of open-source benchmarks have been built to evaluate and compare the workloads on these systems, and to promote their technology improvement. However, to date no comprehensive survey has been written on this topic. This paper attempts to fill the void by presenting a review of the state-of-the-art big data benchmarking efforts. The paper first gives an overview of popular open-source benchmarks from the point of view of big data systems. It then reviews the three important aspects of benchmarking - workload generation techniques, workload input data generation techniques, and metrics used to assess systems. For each aspect, the paper divides the surveyed benchmarks into different categories and describes some representative benchmarks, rather than all benchmarks listed, in each category, following the discussion of potential research directions to motivate future work in this area. Rui Han 0001, Lizy Kurian John, Jianfeng Zhan |
IEEE Trans. Serv. Comput. | 3 |
| 2017 | Towards memory and computation efficient graph processing on sparkabstractAlgorithms for large scale natural graph processing can be categorized into two types based on their value propagation behaviors: the unidirectional value propagation (UVP) algorithms and the bidirectional value propagation (BVP) algorithms. The behavior about how vertices interact with neighbors also differs between two algorithm types, which demands different system design choices. However, current distributed graph processing systems usually try to support both types in one general-purpose framework Such system design can not promise good performance and low resource consumption for both types. Especially, for UVP algorithms, current systems can not guarantee low memory footprint, computation efficiency and communication efficiency at the same time. In this paper, we propose a new graph processing engine on Spark, GraphV, which is specially designed for the unidirectional value propagation algorithms, and can satisfy all the above requirements for this type of algorithms. To retain the generalization for other algorithms, we also build a dual-engine framework by integrating GraphV with Spark's existing graph processing engine GraphX. The main design choices of GraphV include a cheap propagation-related partitioner, an one-step computation model, and a locality-aware local graph layout. According to the experiment results, GraphV is faster than GraphX by the factors of 1.2x-3.1x, with much less resource consumption. The source code of GraphV will be publicly available from http://prof.ict.ac.cn/GraphV. Xinhui Tian, Yuanqing Guo, Jianfeng Zhan, Lei Wang 0004 |
IEEE BigData | 3 |
| 2017 | Work-in-Progress: Maximizing Model Accuracy in Real-time and Iterative Machine LearningabstractAs iterative machine learning (ML) (e.g. neural network based supervised learning and k-means clustering) becomes more ubiquitous in our daily life, it is becoming increasingly important to complete model training quickly to support real-time decision making, while still achieving high model accuracy (e.g. low prediction errors) that is critical for profits of ML tasks. Motivated by the observation that the small proportions of accuracy-critical input data can contribute to large parts of model accuracy in many iterative ML applications, this paper introduces a system middleware to maximize model accuracy by spending the limited time budget on the most accuracy-related input data. To achieve this, our approach employs a fast method to divide the input data into multiple parts of similar points and represents each part with an aggregated data point. Using these points, it quickly estimates the correlations between different parts and model accuracy, thus allowing ML tasks to process the most accuracy-related parts first. We incorporate our approach with two popular supervised and unsupervised ML algorithms on Spark and demonstrate its benefits in providing high model accuracy under short deadlines. Rui Han 0001, Fan Zhang 0047, Lydia Y. Chen, Jianfeng Zhan |
RTSS | 4 |
| 2017 | CLAP: Component-Level Approximate Processing for Low Tail Latency and High Result Accuracy in Cloud Online ServicesabstractModern latency-critical online services such as search engines often process requests by consulting large input data spanning massive parallel components. Hence the tail latency of these components determines the service latency. To trade off result accuracy for tail latency reduction, existing techniques use the components responding before a specified deadline to produce approximate results. However, they skip a large proportion of components when load gets heavier, thus incurring large accuracy losses. In this paper, we propose CLAP to enable component-level approximate processing of requests for low tail latency and small accuracy losses. CLAP aggregates information of input data to create small aggregated data points. Using these points, CLAP reduces latency variance of parallel components and allows them to produce initial results quickly; CLAP also identifies the parts of input data most related to requests' result accuracies, thus first using these parts to improve the produced results to minimize accuracy losses. We evaluated CLAP using real services and datasets. The results show: (i) CLAP reduces tail latency by 6.46 times with accuracy losses of 2.2 percent compared to existing exact processing techniques; (ii) when using the same latency, CLAP reduces accuracy losses by 31.58 times compared to existing approximate processing techniques. Rui Han 0001, Siguang Huang, Jianfeng Zhan |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Understanding Big Data Analytics Workloads on Modern ProcessorsabstractBig data analytics workloads are very significant ones in modern data centers, and it is more and more important to characterize their representative workloads and understand their behaviors so as to improve the performance of data center computer systems. In this paper, we embark on a comprehensive study to understand the impacts and performance implications of the big data analytics workloads on the systems equipped with modern superscalar out-of-order processors. After investigating three most important application domains in Internet services in terms of page views and daily visitors, we choose 11 representative data analytics workloads and characterize their micro-architectural behaviors by using hardware performance counters. Our study reveals that the big data analytics workloads share many inherent characteristics, which place them in a different class from the traditional workloads and the scale-out services. To further understand the characteristics of big data analytics workloads, we perform correlation analysis to identify the most key factors that affect cycles per instruction (CPI). Also, we reveal that the increasing complexity of the big data software stacks will put higher pressures on the modern processor pipelines. Zhen Jia 0001, Jianfeng Zhan, Lei Wang 0004, Chunjie Luo, Wanling Gao, Rui Han 0001, Lixin Zhang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Auto-tuning Spark Big Data Workloads on POWER8: Prediction-Based Dynamic SMT ThreadingabstractMuch research work devotes to tuning big data analytics in modern data centers, since %the truth that even a small percentage of performance improvement immediately translates to huge cost savings because of the large scale. Simultaneous multithreading (SMT) receives great interest from data center communities, as it has the potential to boost performance of big data analytics by increasing the processor resources utilization. For example, the emerging processor architectures like POWER8 support up to 8-way multithreading. However, as different big data workloads have disparate architectural characteristics, how to identify the most efficient SMT configuration to achieve the best performance is challenging in terms of both complex application behaviors and processor architectures. In this paper, we specifically focus on auto-tuning SMT configuration for Spark-based big data workloads on POWE-R8. However, our methodology could be generalized and extended to other programming software stacks and other architectures. Zhen Jia 0001, Guancheng Chen, Jianfeng Zhan, Lixin Zhang 0002, Yonghua Lin, H. Peter Hofstee |
PACT | 4 |
| 2016 | BDTUne: Hierarchical correlation-based performance analysis and rule-based diagnosis for big data systemsabstractAlthough big data systems are in widespread use and there have much research efforts for improving big data systems performance, efficiently analysing and diagnosing performance bottlenecks over these massively distributed systems remain a major challenge. In this paper, we propose a hierarchical correlation-based analysis and rule-based diagnostic approach for big data systems. The key approaches lie in identifying performance bottlenecks, classifying root causes, analyzing performance according to multi-level performance metrics, and setting diagnostic rules for performance tuning. Based on this approach, we have implemented BDTune - a lightweight, extensible and transparent tool that can provide valuable insights into performance of big data applications with a very low overhead. We also report our experience on how to use BDTune to conduct performance analysis and performance bottlenecks diagnosis, and demonstrate BDTune can help users find the performance bottlenecks and provide optimization recommendations. Zhen Jia 0001, Lei Wang 0004, Jianfeng Zhan, Tianxu Yi |
IEEE BigData | 4 |
| 2016 | AccuracyTrader: Accuracy-Aware Approximate Processing for Low Tail Latency and High Result Accuracy in Cloud Online ServicesabstractModern latency-critical online services such as search engines often process requests by consulting large input data spanning massive parallel components. Hence the tail latency of these components determines the service latency. To trade off result accuracy for tail latency reduction, existing techniques use the components responding before a specified deadline to produce approximate results. However, they may skip a large proportion of components when load gets heavier, thus incurring large accuracy losses. This paper presents AccuracyTrader that produces approximate results with small accuracy losses while maintaining low tail latency. AccuracyTrader aggregates information of input data on each component to create a small synopsis, thus enabling all components producing initial results quickly using their synopses. AccuracyTrader also uses synopses to identify the parts of input data most related to arbitrary requests' result accuracy, thus first using these parts to improve the produced results in order to minimize accuracy losses. We evaluated AccuracyTrader using workloads in real services. The results show: (i) AccuracyTrader reduces tail latency by over 40 times with accuracy losses of less than 7% compared to existing exact processing techniques, (ii) when using the same latency, AccuracyTrader reduces accuracy losses by over 13 times comparing to existing approximate processing techniques. Rui Han 0001, Siguang Huang, Fei Tang 0003, Fu-Gui Chang, Jianfeng Zhan |
ICPP | 5 |
| 2016 | Characterization and architectural implications of big data workloadsabstractThe previous major efforts on big data benchmark either propose a large amount of workloads (e.g. a recent comprehensive big data benchmark suite—BigDataBench [4]), which impose cognitive difficulty on workload characterization and serious benchmarking cost; or only select a few workloads according to so-called popularity[1], which lead to partial or biased observations. Lei Wang 0004, Jianfeng Zhan, Zhen Jia 0001 |
ISPASS | 3 |
| 2016 | Pipelining image compositing in heterogeneous networking environmentsabstractAbstract Because of intensive inter‐node communications, image compositing has always been a bottleneck in parallel visualization systems. In a heterogeneous networking environment, the variation of link bandwidth and latency adds more uncertainty to the system performance. In this paper, we present a pipelining image compositing algorithm in heterogeneous networking environments, which is able to rearrange the direction of data flow of a compositing pipeline under strict ordering constraint. We introduce a novel directional image compositing operator that specifies not only the color and α channels of the output but also the direction of data flow when performing compositing. Based on this new operator, we thoroughly study the properties of image compositing pipelines in heterogeneous environments. We develop an optimization algorithm that could find the optimal pipeline from an exponentially large searching space in polynomial time. We conducted a comprehensive evaluation on the ns‐3 network simulator. Experimental results demonstrate the efficiency of our method. Copyright © 2016 John Wiley & Sons, Ltd. Dengming Zhu, Hong Qin 0001, Jianfeng Zhan, Jinzhu Gao |
Comput. Animat. Virtual Worlds | 5 |
| 2015 | Interference-Aware Component Scheduling for Reducing Tail Latency in Cloud Interactive ServicesabstractLarge-scale interactive services usually divide requests into multiple sub-requests and distribute them to a large number of server components for parallel execution. Hence the tail latency (i.e. The slowest component's latency) of these components determines the overall service latency. On a cloud platform, each component shares and competes node resources such as caches and I/O bandwidths with its co-located jobs, hence inevitably suffering from their performance interference. In this paper, we study the short-running jobs in a 12k-node Google cluster to illustrate the dynamic resource demands of these jobs, resulting in both individual components' latency variability over time and across different nodes and hence posing a major challenge to maintain low tail latency. Given this motivation, this paper introduces a dynamic and interference-aware scheduler for large-scale, parallel cloud services. At each scheduling interval, it collects workload and resource contention information of a running service, and predicts both the component latency on different nodes and the overall service performance. Based on the predicted performance, the scheduler identifies straggling components and conducts near-optimal component-node allocations to adapt to the changing workloads and performance interferences. We demonstrate that, using realistic workloads, the proposed approach achieves significant reductions in tail latency compared to the basic approach without scheduling. Rui Han 0001, Siguang Huang, Chenrong Shao, Shulin Zhan, Jianfeng Zhan, José Luis Vázquez-Poletti |
ICDCS | 6 |
| 2015 | PCS: Predictive Component-Level Scheduling for Reducing Tail Latency in Cloud Online ServicesabstractModern latency-critical online services often rely on composing results from a large number of server components. Hence the tail latency (e.g. The 99th percentile of response time), rather than the average, of these components determines the overall service performance. When hosted on a cloud environment, the components of a service typically co-locate with short batch jobs to increase machine utilizations, and share and contend resources such as caches and I/O bandwidths with them. The highly dynamic nature of batch jobs in terms of their workload types and input sizes causes continuously changing performance interference to individual components, hence leading to their latency variability and high tail latency. However, existing techniques either ignore such fine-grained component latency variability when managing service performance, or rely on executing redundant requests to reduce the tail latency, which adversely deteriorate the service performance when load gets heavier. In this paper, we propose PCS, a predictive and component-level scheduling framework to reduce tail latency for large-scale, parallel online services. It uses an analytical performance model to simultaneously predict the component latency and the overall service performance on different nodes. Based on the predicted performance, the scheduler identifies straggling components and conducts near-optimal component-node allocations to adapt to the changing performance interferences from batch jobs. We demonstrate that, using realistic workloads, the proposed scheduler reduces the component tail latency by an average of 67.05% and the average overall service latency by 64.16% compared with the state-of-the-art techniques on reducing tail latency. Rui Han 0001, Siguang Huang, Chenrong Shao, Shulin Zhan, Jianfeng Zhan, José Luis Vázquez-Poletti |
ICPP | 6 |
| 2015 | PowerTracer: Tracing Requests in Multi-Tier Services to Reduce Energy InefficiencyabstractAs energy has become one of the key operating costs in running a data center and power waste commonly exists, it is essential to reduce energy inefficiency inside data centers. In this paper, we develop an innovative framework, calledPowerTracer, for diagnosing energy inefficiency and saving power. Inside the framework, we first present a resource tracing method based on request tracing in multi-tier services of black boxes. Then, we propose a generalized methodology of applying a request tracing approach for energy inefficiency diagnosis and power saving in multi-tier service systems. With insights into service performance and resource consumption of individual requests, we develop (1) a bottleneck diagnosis tool that pinpoints the root causes of energy inefficiency, and (2) a power saving method that enables dynamic voltage and frequency scaling (DVFS) with online request tracing. We implement a prototype of PowerTracer, and conduct extensive experiments to validate its effectiveness. Our tool analyzes several state-of-the-practice and state-of-the-art DVFS control policies and uncovers existing energy inefficiencies. Meanwhile, the experimental results demonstrate that PowerTracer outperforms its peers in power saving. Jianfeng Zhan, Haining Wang 0001, Yunwei Gao, Chuliang Weng, Yong Qi 0001 |
IEEE Trans. Computers | 2 |
| 2015 | TSAC: Enforcing Isolation ofVirtual Machines in CloudsabstractVirtualization plays a vital role in building the infrastructure of clouds, and isolation is considered as one of its important features. However, we demonstrate with practical measurements that there exist two kinds of isolation problems in current virtualized systems, due to cache interference in a multi-core processor. That is, one virtual machine could degrade the performance or obtain the load information of another virtual machine, which running on a same physical machine. Then we present a time-sensitive contention management approach (TSAC) for allocating resources dynamically in the virtual machine monitor, in which virtual machines are controlled to share some physical resources (e.g., CPU or page color) in a dynamical manner, in order to enforce isolation between the virtual machines without sacrificing performance of the virtualized system. We have implemented a working prototype based on Xen, evaluated the implemented prototype with experiments, and experimental results show that TSAC could significantly improve isolation of virtualization. Specifically, compared to the default Xen, TSAC could improve the performance of the victim virtual machine by up to about 78 percent, and perform well in blocking its cache-based load information leakage. Chuliang Weng, Jianfeng Zhan, Yuan Luo 0003 |
IEEE Trans. Computers | 2 |
| 2014 | Digging deeper into cluster system logs for failure prediction and root cause diagnosisabstractAs the sizes of supercomputers and data centers grow towards exascale, failures become normal. System logs play a critical role in the increasingly complex tasks of automatic failure prediction and diagnosis. Many methods for failure prediction are based on analyzing event logs for large scale systems, but there is still neither a widely used one to predict failures based on both non-fatal and fatal events, nor a precise one that uses fine-grained information (such as failure type, node location, related application, and time of occurrence). A deeper and more precise log analysis technique is needed. We propose a three-step approach to draw out event dependencies and to identify failure-event generating processes. First, we cluster frequent event sequences into event groups based on common events. Then we infer causal dependencies between events in each event group. Finally, we extract failure rules based on the observation that events of the same event types, on the same nodes or from the same applications have similar operational behaviors. We use this rich information to improve failure prediction. Our approach semi-automates diagnosing the root causes of failure events, making it a valuable tool for system administrators. Xiaoyu Fu, Sally A. McKee, Jianfeng Zhan, Ninghui Sun |
CLUSTER | 4 |
| 2014 | BigOP: Generating Comprehensive Big Data Workloads as a Benchmarking Framework
Yuqing Zhu 0001, Jianfeng Zhan, Chuliang Weng, Raghunath Othayoth Nambiar, Jinchao Zhang 0001, Xingzhen Chen, Lei Wang 0004 |
DASFAA (2) | 2 |
| 2014 | BigDataBench: A big data benchmark suite from internet servicesabstractAs architecture, systems, and data management communities pay greater attention to innovative big data systems and architecture, the pressure of benchmarking and evaluating these systems rises. However, the complexity, diversity, frequently changed workloads, and rapid evolution of big data systems raise great challenges in big data benchmarking. Considering the broad use of big data systems, for the sake of fairness, big data benchmarks must include diversity of data and workloads, which is the prerequisite for evaluating big data systems and architecture. Most of the state-of-the-art big data benchmarking efforts target evaluating specific types of applications or system software stacks, and hence they are not qualified for serving the purposes mentioned above. This paper presents our joint research efforts on this issue with several industrial partners. Our big data benchmark suite-BigDataBench not only covers broad application scenarios, but also includes diverse and representative data sets. Currently, we choose 19 big data benchmarks from dimensions of application scenarios, operations/ algorithms, data types, data sources, software stacks, and application types, and they are comprehensive for fairly measuring and evaluating big data systems and architecture. BigDataBench is publicly available from the project home page http://prof.ict.ac.cn/BigDataBench. Also, we comprehensively characterize 19 big data workloads included in BigDataBench with varying data inputs. On a typical state-of-practice processor, Intel Xeon E5645, we have the following observations: First, in comparison with the traditional benchmarks: including PARSEC, HPCC, and SPECCPU, big data applications have very low operation intensity, which measures the ratio of the total number of instructions divided by the total byte number of memory accesses; Second, the volume of data input has non-negligible impact on micro-architecture characteristics, which may impose challenges for simulation-based big data architecture research; Last but not least, corroborating the observations in CloudSuite and DCBench (which use smaller data inputs), we find that the numbers of L1 instruction cache (L1I) misses per 1000 instructions (in short, MPKI) of the big data applications are higher than in the traditional benchmarks; also, we find that L3 caches are effective for the big data applications, corroborating the observation in DCBench. Lei Wang 0004, Jianfeng Zhan, Chunjie Luo, Yuqing Zhu 0001, Qiang Yang 0012, Yongqiang He, Wanling Gao, Zhen Jia 0001, Yingjie Shi, Chen Zheng 0001, Kent Zhan, Bizhu Qiu |
HPCA | 2 |
| 2014 | An Automatic Framework for Detecting and Characterizing Performance Degradation of Software SystemsabstractSoftware systems that run continuously over a long time have been frequently reported encountering gradual degradation issues. That is, as time progresses, software tends to exhibit degraded performance, deflated service capacity, or deteriorated QoS. Currently, the state-of-the-art approach of Mann-Kendall Test & Seasonal Kendall Test & Sen's Slope Estimator & Seasonal Sen's Slope Estimator (MKSK) detects and characterizes degradation via a combination of techniques in statistical trend analysis. Nevertheless, we pinpoint some drawbacks of MKSK in this paper: 1) MKSK cannot be automated for large scale software degradation analysis, 2) MKSK estimates the degradation trend of software in an oversimplified linear way, 3) MKSK is sensitive to noise, and 4) MKSK suffers from high computational complexity. To overcome all these limitations, we propose a more advanced approach called Modified Cox-Stuart Test & Iterative Hodrick-Prescott Filter (CSHP). The superiority of our CSHP approach over MKSK is validated through extensive Monte Carlo simulations, as well as a real performance dataset measured from 99 real-world web servers. Yong Qi 0001, Yangfan Zhou 0002, Pengfei Chen 0002, Jianfeng Zhan, Michael R. Lyu |
IEEE Trans. Reliab. | 5 |
| 2013 | A Relationship-Based VM Placement Framework of Cloud EnvironmentabstractManaging computation resources in a cost-effective way has become the core competence for a Cloud provider to win over the market because of the "pay-as-you-go" business model. Therefore, VM placement has become more and more important in the research and practices of VM management by determining at what condition and on which physical server a VM should be placed so that the SLA can be guaranteed and servers' utilization can be improved. Much existing work simply formulates the above issue to be a bin-packing problem, which does not take the VM relationships into account. However, the relationship information can greatly impact the SLA of the Cloud system and the resource utilization. Therefore, in this paper, we propose a relationship-based VM placement framework, SmartCRS, to optimize the VM placement procedure. SmartCRS reveals the relationships between VMs automatically. Then by using such information and based on a constraint library, it gives a proper VM placement plan. Finally, the plan is carried out by SmartCRS automatically or by Cloud administrators manually to improve the server utilization and guarantee the required SLA. Two case studies are conducted to demonstrate the effectiveness and efficiency of the proposed framework at the end of this paper. Xiaodong Zhang 0025, Ying Zhang 0012, Xing Chen 0002, Gang Huang 0001, Jianfeng Zhan |
COMPSAC | 6 |
| 2013 | Cost-Aware Cooperative Resource Provisioning for Heterogeneous Workloads in Data CentersabstractRecent cost analysis shows that the server cost still dominates the total cost of high-scale data centers or cloud systems. In this paper, we argue for a new twist on the classical resource provisioning problem: heterogeneous workloads are a fact of life in large-scale data centers, and current resource provisioning solutions do not act upon this heterogeneity. Our contributions are threefold: first, we propose a cooperative resource provisioning solution, and take advantage of differences of heterogeneous workloads so as to decrease their peak resources consumption under competitive conditions; second, for four typical heterogeneous workloads: parallel batch jobs, web servers, search engines, and MapReduce jobs, we build an agile system PhoenixCloud that enables cooperative resource provisioning; and third, we perform a comprehensive evaluation for both real and synthetic workload traces. Our experiments show that our solution could save the server cost aggressively with respect to the noncooperative solutions that are widely used in state-of-the-practice hosting data centers or cloud systems: for example, EC2, which leverages the statistical multiplexing technique, or RightScale, which roughly implements the elastic resource provisioning technique proposed in related state-of-the-art work. Jianfeng Zhan, Lei Wang 0004, Weisong Shi, Chuliang Weng, Xiutao Zang |
IEEE Trans. Computers | 1 |
| 2013 | Parallel Streamline Placement for 2D Flow FieldsabstractParallel streamline placement is still an open problem in flow visualization. In this paper, we propose an innovative method to place streamlines in parallel for 2D flow fields. This method is based on our proposed concept of local tracing areas (LTAs). An LTA is defined as a subdomain enclosed by streamlines and/or field borders, where the tracing of streamlines are localized. Given a flow field, it is initialized as an LTA, which is later recursively partitioned into hierarchical LTAs. Streamlines are placed within different LTAs simultaneously and independently. At the same time, to control the density of streamlines, each streamline is associated with an isolation zone and a saturation zone, both of which are center aligned with the streamline but have different widths. None of streamlines can trace into isolation zones of others. And new streamlines are only seeded within valid seeding areas (VSAs) that are enclosed by saturation zones and/or field borders. To implement the parallel strategy and the density control, a cell-based modeling is devised to describe isolation zones and LTAs as well as saturation zones and VSAs. With the help of these cell-based models, a heuristic seeding strategy is proposed to seed streamlines within irregular LTAs, and a cell-marking technique is used to control the seeding and tracing of streamlines. Test results show that the placement method can achieve highly parallel performance on shared memory systems without losing the quality of placements. Jianfeng Zhan, Beichen Liu, Jianguo Ning |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2012 | LogMaster: Mining Event Correlations in Logs of Large-Scale Cluster SystemsabstractThis paper presents a set of innovative algorithms and a system, named Log Master, for mining correlations of events that have multiple attributions, i.e., node ID, application ID, event type, and event severity, in logs of large-scale cloud and HPC systems. Different from traditional transactional data, e.g., supermarket purchases, system logs have their unique characteristics, and hence we propose several innovative approaches to mining their correlations. We parse logs into an n-ary sequence where each event is identified by an informative nine-tuple. We propose a set of enhanced apriori-like algorithms for improving sequence mining efficiency, we propose an innovative abstraction-event correlation graphs (ECGs) to represent event correlations, and present an ECGs-based algorithm for fast predicting events. The experimental results on three logs of production cloud and HPC systems, varying from 433490 entries to 4747963 entries, show that our method can predict failures with a high precision and an acceptable recall rates. Xiaoyu Fu, Jianfeng Zhan, Wei Zhou 0019, Zhen Jia 0001 |
SRDS | 3 |
| 2012 | CloudRank-D: benchmarking and ranking cloud computing systems for data processing applications
Chunjie Luo, Jianfeng Zhan, Zhen Jia 0001, Lei Wang 0004, Lixin Zhang 0002, Cheng-Zhong Xu 0001, Ninghui Sun |
Frontiers Comput. Sci. | 2 |
| 2012 | Performance analysis and optimization of MPI collective operations on multi-core clusters
Bibo Tu, Jianping Fan 0002, Jianfeng Zhan |
J. Supercomput. | 3 |
| 2012 | Precise, Scalable, and Online Request Tracing for Multitier Services of Black BoxesabstractAs more and more multitier services are developed from commercial off-the-shelf components or heterogeneous middleware without source code available, both developers and administrators need a request tracing tool to (1) exactly know how a user request of interest travels through services of black boxes and (2) obtain macrolevel user request behaviors of services without manually analyzing massive logs. This need is further exacerbated by IT system “agility,” which mandates the tracing tool to provide online performance data since offline approaches cannot reflect system changes in real time. Moreover, considering the large scale of deployed services, a pragmatic tracing approach should be scalable in terms of the cost in collecting and analyzing logs. In this paper, we introduce a precise, scalable, and online request tracing tool for multitier services of black boxes. Our contributions are threefold. First, we propose a precise request tracing algorithm for multitier services of black boxes, which only uses application-independent knowledge. Second, we present a microlevel abstraction, component activity graph, to represent causal paths of each request. On the basis of this abstraction, we use dominated causal path patterns to represent repeatedly executed causal paths that account for significant fractions, and we further present a derived performance metric of causal path patterns, latency percentages of components, to enable debugging performance-in-the-large. Third, we develop two mechanisms, tracing on demand and sampling, to significantly increase the system scalability. We implement a prototype of the proposed system, called PreciseTracer, and release it as open source code. In comparison with WAP5-a black-box tracing approach, PreciseTracer achieves higher tracing accuracy and faster response time. Our experimental results also show that PreciseTracer has low overhead, and still achieves high tracing accuracy even if an aggressive sampling policy is adopted, indicating that PreciseTracer is a promising tracing tool for large-scale production systems. Bo Sang, Jianfeng Zhan, Haining Wang 0001, Dongyan Xu, Lei Wang 0004, Zhen Jia 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | In Cloud, Can Scientific Communities Benefit from the Economies of Scale?abstractThe basic idea behind cloud computing is that resource providers offer elastic resources to end users. In this paper, we intend to answer one key question to the success of cloud computing: in cloud, can small-to-medium scale scientific communities benefit from the economies of scale? Our research contributions are threefold: first, we propose an innovative public cloud usage model for small-to-medium scale scientific communities to utilize elastic resources on a public cloud site while maintaining their flexible system controls, i.e., create, activate, suspend, resume, deactivate, and destroy their high-level management entities-service management layers without knowing the details of management. Second, we design and implement an innovative system-DawningCloud, at the core of which are lightweight service management layers running on top of a common management service framework. The common management service framework of DawningCloud not only facilitates building lightweight service management layers for heterogeneous workloads, but also makes their management tasks simple. Third, we evaluate the systems comprehensively using both emulation and real experiments. We found that for four traces of two typical scientific workloads: High-Throughput Computing (HTC) and Many-Task Computing (MTC), DawningCloud saves the resource consumption maximally by 59.5 and 72.6 percent for HTC and MTC service providers, respectively, and saves the total resource consumption maximally by 54 percent for the resource provider with respect to the previous two public cloud solutions. To this end, we conclude that small-to-medium scale scientific communities indeed can benefit from the economies of scale of public clouds with the support of the enabling system. Lei Wang 0004, Jianfeng Zhan, Weisong Shi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | A Runtime Fault Detection Method for HPC ClusterabstractAs the number of nodes keeps increasing, faults have become commonplace for HPC cluster. For fast recovery from faults, the fault detection method is necessary. Based on the usage patterns of HPC cluster, a automatic runtime fault detection mechanism is proposed in this paper: First, the normal activities for nodes in HPC cluster are modeled using runtime state by clustering analysis, Second, the fault detection process is implemented by comparing the current runtime state of nodes with normal activity models. A fault alarm is made immediately when the current runtime state deviates from the normal activity models. In the experiments, the faults are simulated by fault injection methods and the experimental results show that the runtime fault detection method in this paper can detect faults with high accuracy. Linping Wu, Hongbing Luo, Jianfeng Zhan, Dan Meng 0002 |
PDCAT | 3 |
| 2011 | Automatic performance debugging of SPMD-style parallel programs
Xu Liu 0001, Jianfeng Zhan, Kunlin Zhan, Weisong Shi, Dan Meng 0002, Lei Wang 0004 |
J. Parallel Distributed Comput. | 2 |
| 2010 | Online Event Correlations Analysis in System Logs of Large-Scale Cluster Systems
Wei Zhou 0019, Jianfeng Zhan, Dan Meng 0002 |
NPC | 2 |
| 2009 | Precise request tracing and performance debugging for multi-tier services of black boxesabstractAs more and more multi-tier services are developed from commercial components or heterogeneous middleware without the source code available, both developers and administrators need a precise request tracing tool to help understand and debug performance problems of large concurrent services of black boxes. Previous work fails to resolve this issue in several ways: they either accept the imprecision of probabilistic correlation methods, or rely on knowledge of protocols to isolate requests in pursuit of tracing accuracy. This paper introduces a tool named PreciseTracer to help debug performance problems of multi-tier services of black boxes. Our contributions are two-fold: first, we propose a precise request tracing algorithm for multi-tier services of black boxes, which only uses ap plication-independent knowledge; secondly, we present a component activity graph abstraction to represent causal paths of requests and facilitate end-to-end performance debugging. The low overhead and tolerance of noise make PreciseTracer a promising tracing tool for using on production systems. Jianfeng Zhan, Yong Li 0007, Lei Wang 0004, Dan Meng 0002, Bo Sang |
DSN | 2 |
| 2009 | Accurate Analytical Models for Message Passing on Multi-core ClustersabstractMemory hierarchy on multi-core clusters has two-fold characteristics: vertical memory hierarchy and horizontal memory hierarchy. Vertical memory hierarchy has been modeled by previous work (e.g. memory logP, lognP, log3P etc.) to analyze middlewarepsilas effects on point-to-point communication with different message sizes and message strides; Horizontal memory hierarchy has become more prominent due to distinct performance among three levels of communication in a multi-core cluster: intra-CMP, inter-CMP and inter-node, which should adequately be considered. Derived from lognP and log3P models, new analytical models mlognP and its reduction 2log{2,3}P are proposed to unitedly abstract memory hierarchy on multi-core clusters in vertical and horizontal levels. The results of performance evaluation show that it is indispensable to incorporate horizontal memory hierarchy into new models suitable for multi-core clusters, and 2log{2,3}P model can predict communication costs for message passing on multi-core clusters more accurately than log3P model. Bibo Tu, Jianping Fan 0002, Jianfeng Zhan |
PDP | 3 |
| 2008 | Multi-core aware optimization for MPI collectivesabstractMPI collective operations on multi-core clusters should be multi-core aware. In this paper, collective algorithms with hierarchical virtual topology focus on the performance difference among different communication levels on multi-core clusters, simply for intra-node and inter-node communication; Furthermore, to select befitting segment sizes for intra-node collective communication can cater to cache hierarchy in multi-core processors. Based on existing collective algorithms in MPICH2, above two techniques construct portable optimization methodology over MPICH2 for collective operations on multi-core clusters. Conforming to above optimization methodology, multi-core aware broadcast algorithm has been implemented and evaluated as a case study. The results of performance evaluation show that the multi-core aware optimization methodology over MPICH2 is efficient. Bibo Tu, Ming Zou, Jianfeng Zhan, Jianping Fan 0002 |
CLUSTER | 3 |
| 2008 | Design Techniques for the Scalability of Cluster Management Software on Dawning SupercomputersabstractCluster management software has faced more increased scalability challenge with ever enlarged cluster scale. Its good scalability rests with feasible design techniques focusing on hybrid software topologies with partitioning policy, non-blocking I/O multiplexing and message on demand. Design patterns are generic solutions to recurring software design problems, and above three important techniques are abstracted the design pattern of scalable cluster management software in this paper. According to this design pattern, some cluster management tools, such as job scheduling, MPI job launcher and so on, have been designed and applied on Dawning supercomputers. Some results of performance evaluation have shown that good scalability of cluster management software on Dawning supercomputers has benefited from this design pattern. Bibo Tu, Ming Zou, Jianfeng Zhan, Jianping Fan 0002 |
ISPA | 3 |
| 2008 | A Dynamic Provisioning Framework for Multi-tier Internet Applications in Virtualized Data CenterabstractWith the resurgence of virtualization technology, todaypsilas Internet data centers are shifting towards virtualized data centers. Internet applications tend to see dynamically varying workloads. To address the problem of performance management for multi-tier applications hosted in virtualized Internet data center, we propose a three-level automatic provisioning framework based on feedback control for multi-tier applications. Experiments demonstrate the effectiveness of our technique in SLA guarantees while obtaining improved resource utilization. Xu Liu 0001, Jianfeng Zhan |
PDCAT | 3 |
| 2008 | A Fast-Start, Fault-Tolerant MPI Launcher on Dawning SupercomputersabstractDaemon-based MPI launchers are the mainstream in nowadays, because they can startup processes rapidly. However, effective task management and fault tolerance become more important as the scale of supercomputers enlarges. A new fast-start and fault tolerant launcher, called SFLauncher, has been used to startup MPICH task on Dawning supercomputers. This paper details its features and implementation, with emphasis on scalability, self-organization algorithm and garbage reclamation. The results of performance evaluation on SFLauncher are also given. Xu Liu 0001, Bibo Tu, Jianfeng Zhan, Dan Meng 0002 |
PDCAT | 3 |
| 2007 | A layered design methodology of cluster system stackabstractThe application range of cluster has expanded beyond scientific computing, but the present cluster system software fails to provide a flexible architecture to promote code reuse and facilitate building cluster system software for different computing contexts, most of which are developed from scratch case by case, or integrated or packaged with “the best practice”. In this paper, we have proposed a layered design methodology to build cluster system stack with different layers concentrating on different functions, and developed common sets of core service as reusing framework for different computing context. Following this methodology, we have built Phoenix-a complete cluster system stack for both scientific and business computing, which is verified and deployed on Dawning 4000A super computer for scientific computing and other cluster systems for business computing. The qualitative evaluation and our practices show the design methodology of Phoenix has advantages over other methodologies. Jianfeng Zhan, Lei Wang 0004, Bibo Tu, Yu Wen 0001, Yuansheng Chen, Wei Zhou 0019, Dan Meng 0002, Ninghui Sun |
CLUSTER | 1 |
| 2007 | Grid Unit: A Self-Managing Building Block for Grid SystemabstractGrid system software is inherently complex, hard to build and maintain. In this paper, we propose a self- managing building block: Grid Unit, which facilitates constructing Grid system with higher availability and lower management overhead. We present an agent organization as autonomic management framework, and propose a self-recovering protocol to eliminate most of tough jobs from system administrator's routines. The system has been deployed on Dawning 4000A since 2004, the biggest node for China Grid system. We have done extensive experiments to evaluate Grid Unit, and the collected log data shows the availability of a Grid parallel process management service, built on the basis of Grid Unit, reaches 99.997%. Jianfeng Zhan, Lei Wang 0004, Ming Zou, Yulei Ding |
PDCAT | 1 |
| 2006 | A Failure-Aware Scheduling Strategy in Large-Scale Cluster SystemabstractAs the scale is expanding, node failure becomes a commonplace feature of large-scale cluster systems. As an important part of cluster operating system software, job scheduling takes charge with high efficient resource management and reasonable job scheduling. The function of job scheduling in cluster is divided into two sub-parts: job selection and node allocation. In this paper, we introduce a failure-aware scheduling strategy named LUNF (Longest Uptime Node First) node allocation policy using characterization of nodes' failure. Simulation results show that LUNF policy do better than random node allocation policy for the system performance. Linping Wu, Dan Meng 0002, Jianfeng Zhan, Lei Wang 0004, Bibo Tu |
CCGRID | 3 |
| 2006 | PhoenixG: A Unified Management Framework for Industrial Information GridabstractThe industrial information grid is a special kind of system, the users of which exclusively own geographically distributed computing resources for business service, and try to maintain the lowest total cost of ownership while guaranteeing quality of service. In this paper, we classify the industrial information grid as an extension to grid problem; develop a unified management framework for new management paradigm, which supports the distribution of administration labor and collaboration of system administrator at different locations; propose a self-organizing algorithm, which supports the initial establishment, daily management and exception processing of industrial information grid. Finally, we evaluate the performance of system management, and analyze the management overhead with this new management paradigm. Jianfeng Zhan, Gengpu Liu, Lei Wang 0004, Bibo Tu, Yang Li 0002, Yan Hao, Xuehai Hong, Dan Meng 0002, Ninghui Sun |
CCGRID | 1 |
| 2006 | An Integrated Adaptive Management System for Cluster-based Web ServicesabstractThe complexity of the cluster-based Web service challenges the traditional approaches, which fail to guarantee the reliability and real-time performance required. In this paper, we present an integrated adaptive management system (JAMS) for such service. The issues we discuss address to efficiently allocate resources and provide more effective QoS support under a wide range of load conditions. For the global resource level, we introduce spare instance and corresponding management strategy as a supplemental adaptive mechanism. The spare instances hosted on shared node afford better resource utilization and more effective QoS support in the case of overload or workload fluctuation. Further, it can relax the influence of the fault recovery from the hardware and software failure. For the local level, we apply a multipurpose linear-quadratic regulator (LQR) as basic adaptive element. The control scheme using reject time ratio as control input is able to provide guarantees for overload protection, resource control, Qos control, performance isolation, and effective management for spare instances. Results of experiments on both static and dynamic Web sites illustrate the efficiency and robustness of the multi-purpose LQR Dan Meng 0002, Jianfeng Zhan |
CLUSTER | 4 |
| 2006 | A proactive fault-detection mechanism in large-scale cluster systemsabstractTo improve the whole dependability of large-scale cluster systems, an online fault detection mechanism is proposed in this paper. This mechanism can detect the fault in time before node fails and enables the proactive fault management. The proposed mechanism is summarized as follows: first, the dynamic characteristics of cluster system running in normal activity are built using time series analysis methods. Second, the fault detection process is implemented by comparing the current running state of cluster system with normal running model. The fault alarm decision is made immediately when the current running state deviates the normal running model. The experiment results show that this mechanism can detect the fault in cluster system in good time. Linping Wu, Dan Meng 0002, Wen Gao 0001, Jianfeng Zhan |
IPDPS | 4 |
| 2006 | Easy and reliable cluster management: the self-management experience of Fire PhoenixabstractHigh-Performance clusters are rapidly becoming an important computing platform for both scientific and business applications. To fulfil the new demands and challenges, cluster system software is inevitably complex. Even for experienced administrators, the management of a cluster system is an exhausting job. This paper introduces Fire Phoenix, a scalable and self-managing cluster system software that supports both scientific and commercial applications. With the self-configuring and self-healing features, much of the machine configuration and error recovery can be done automatically. Our design has been proven effective in the operations of the Dawning 4000A supercomputer, which is the biggest cluster system in China Dan Meng 0002, Jianfeng Zhan, Lei Wang 0004, Linping Wu, Huang Wei |
IPDPS | 3 |
| 2006 | Design Patterns of Scalable Cluster System SoftwareabstractThe design pattern of cluster system software has an important influence on scalability of massive cluster system. The paper presents design patterns of scalable cluster system software, including scalable software topologies and optimized communication modes. These design patterns have been widely applied in Dawning series of supercomputers and some results of performance evaluation show their good scalability Bibo Tu, Ming Zou, Jianfeng Zhan, Lei Wang 0004, Jianping Fan 0002 |
PDCAT | 3 |
| 2006 | The Failure-rate Aware Scheduling Policies for Large-scale Cluster SystemsabstractWith the scale expanding, node failures become one of the important obstacles when using large-scale cluster systems. The traditional scheduling policies of cluster only took into account the factors such as jobs priority and node load with the node failure rate omitted. The function of job scheduling in cluster system can be divided into two sub-processes: job selection process and node allocation process. In this paper, we introduce several scheduling policies considering the node failure rate with which the more dependable nodes are selected during the node allocation process. In the end, we use the discrete event-driven simulation method to evaluate the policies and the simulation results show that the failure-rate aware scheduling policies do better than random node allocation policy for the system performance Linping Wu, Dan Meng 0002, Jianfeng Zhan, Bibo Tu |
PDCAT | 4 |
| 2005 | Adaptive Management of a Utility ComputingabstractThe complexity of the high performance Web-based application challenges the traditional approaches, which fail to guarantee the reliability and real-time performance required. In this paper, we have studied the adaptive mechanisms for managing such applications and explained them based on a prototype of an adaptive application management system (AMUS) in cluster. AMUS is composed of the SLA event-driven global resource manager, the server resource manager and the self-adapting application systems based on feedback control theory. The adoption of feedback control theory supports the application resource control in the case of the resource contention and the guarantee of the QoS performance in the changing environment Dan Meng 0002, Danjun Liu, Jianfeng Zhan |
CLUSTER | 5 |
| 2005 | Fire Phoenix Cluster Operating System Kernel and its EvaluationabstractFire Phoenix cluster operating system kernel (Phoenix kernel) is a minimum set of cluster core junctions with scalability and fault-tolerance support. In this paper, we define components of cluster operating system kernel, and introduce its internal mechanism for scalability and fault-tolerance support. Based on Phoenix kernel, user environments can be easily constructed according to users' needs. In addition, we evaluate Phoenix kernel from four different perspectives, such as fault-tolerance, scalability, performance impact on scientific computing, and easiness of constructing user environment. Our design has been proved in the practices of Dawning 4000A super server, which is the biggest cluster system for scientific computing in China Jianfeng Zhan, Ninghui Sun |
CLUSTER | 1 |