Zhendong Bei

dblp:139/4211 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0001-6875-5539ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 60% Performance modeling and evaluation · 20% High-performance computing · 14%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › configuration tuning
configuration auto-tuning
0.922022
OSC: An Online Self-Configuring Big Data Framework for Optimization of QoS · IEEE Trans. Computers 2022
Datasize-Aware High Dimensional Configurations Auto-Tuning of In-Memory Cluster Computing · ASPLOS 2018
Cloud and datacenter computing › big data platform
big data frameworks
0.612022
OSC: An Online Self-Configuring Big Data Framework for Optimization of QoS · IEEE Trans. Computers 2022
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.612022
OSC: An Online Self-Configuring Big Data Framework for Optimization of QoS · IEEE Trans. Computers 2022
High-performance computing
performance optimization
0.612022
OSC: An Online Self-Configuring Big Data Framework for Optimization of QoS · IEEE Trans. Computers 2022
Performance modeling and evaluation
benchmarking
0.312018
MIA: Metric Importance Analysis for Big Data Workload Characterization · IEEE Trans. Parallel Distributed Syst. 2018
Cloud and datacenter computing
cluster resource management and scheduling
0.312018
Datasize-Aware High Dimensional Configurations Auto-Tuning of In-Memory Cluster Computing · ASPLOS 2018
Performance modeling and evaluation
workload characterization
0.312018
MIA: Metric Importance Analysis for Big Data Workload Characterization · IEEE Trans. Parallel Distributed Syst. 2018
Cloud and datacenter computing
configuration tuning
0.212016
RFHOC: A Random-Forest Approach to Auto-Tuning Hadoop's Configuration · IEEE Trans. Parallel Distributed Syst. 2016
Parallel and multicore computing › data-parallel programming
mapreduce
0.212016
RFHOC: A Random-Forest Approach to Auto-Tuning Hadoop's Configuration · IEEE Trans. Parallel Distributed Syst. 2016
Performance modeling and evaluation
performance model construction
0.212016
RFHOC: A Random-Forest Approach to Auto-Tuning Hadoop's Configuration · IEEE Trans. Parallel Distributed Syst. 2016
Machine learning and data management
data management for machine learning
0.212022
OSC: An Online Self-Configuring Big Data Framework for Optimization of QoS · IEEE Trans. Computers 2022
High-performance computing › performance optimization
auto-tuning
0.112016
RFHOC: A Random-Forest Approach to Auto-Tuning Hadoop's Configuration · IEEE Trans. Parallel Distributed Syst. 2016

Methods — techniques the papers use, named apart from their topics

genetic algorithm · 1.7ensemble learning · 1.5incremental modeling · 1.1linkage clustering · 0.3kiviat plot · 0.3hierarchical modeling · 0.3random forest · 0.2
YearPublicationVenuePosition
2022 OSC: An Online Self-Configuring Big Data Framework for Optimization of QoS
abstract
Big-data frameworks such as MapReduce/Hadoop or Spark have many performance-critical configuration parameters which may interact with each other in a complex way. Their optimal values for an application on a given cluster are affected by not only the application itself but also its input data. This makes offline auto-configuration approaches hard to be used in practice because the input data of an application may change at each run. To address this issue, we propose an Online Self-Configuring (OSC) approach that automatically determines the optimal parameter values for a given application. OSC synergistically integrates three key techniques. First, OSC leveragesensemble learningto build a precise performance model for a given application. Second, it quantifies theimportanceof the parameters andinteraction intensitybetween them to accelerate the genetic algorithm for searching optimal configuration parameters. Third, OSC supports anincremental modelingapproach to achieve low overhead of the models for online needs. These techniques allow OSC to effectively learn the characteristics of an application and optimize its performance by automatically adjusting the configurations at runtime. Our implementation of OSC atop MapReduce/Hadoop 2.6 improves performance by 60 percent on average and up to 120 percent compared with the state-of-the-art approach. Lastly, the performance benefit of an application running on OSC generally increases along with its input data size.
Zhendong Bei, Nam Sung Kim, Kai Hwang 0001, Zhibin Yu 0001
IEEE Trans. Computers1
2021 Evaluation of residue-residue contact prediction methods: From retrospective to prospective
abstract
Sequence-based residue contact prediction plays a crucial role in protein structure reconstruction. In recent years, the combination of evolutionary coupling analysis (ECA) and deep learning (DL) techniques has made tremendous progress for residue contact prediction, thus a comprehensive assessment of current methods based on a large-scale benchmark data set is very needed. In this study, we evaluate 18 contact predictors on 610 non-redundant proteins and 32 CASP13 targets according to a wide range of perspectives. The results show that different methods have different application scenarios: (1) DL methods based on multi-categories of inputs and large training sets are the best choices for low-contact-density proteins such as the intrinsically disordered ones and proteins with shallow multi-sequence alignments (MSAs). (2) With at least 5L (L is sequence length) effective sequences in the MSA, all the methods show the best performance, and methods that rely only on MSA as input can reach comparable achievements as methods that adopt multi-source inputs. (3) For top L/5 and L/2 predictions, DL methods can predict more hydrophobic interactions while ECA methods predict more salt bridges and disulfide bonds. (4) ECA methods can detect more secondary structure interactions, while DL methods can accurately excavate more contact patterns and prune isolated false positives. In general, multi-input DL methods with large training sets dominate current approaches with the best overall performance. Despite the great success of current DL methods must be stated the fact that there is still much room left for further improvement: (1) With shallow MSAs, the performance will be greatly affected. (2) Current methods show lower precisions for inter-domain compared with intra-domain contact predictions, as well as very high imbalances in precisions between intra-domains. (3) Strong prediction similarities between DL methods indicating more feature types and diversified models need to be developed. (4) The runtime of most methods can be further optimized.
Zhendong Bei, Wenhui Xi, Min Hao 0002, Zhen Ju, Konda Mani Saravanan, Yanjie Wei
PLoS Comput. Biol.2
2018 Datasize-Aware High Dimensional Configurations Auto-Tuning of In-Memory Cluster Computing
abstract
In-Memory cluster Computing (IMC) frameworks (e.g., Spark) have become increasingly important because they typically achieve more than 10× speedups over the traditional On-Disk cluster Computing (ODC) frameworks for iterative and interactive applications. Like ODC, IMC frameworks typically run the same given programs repeatedly on a given cluster with similar input dataset size each time. It is challenging to build performance model for IMC program because: 1) the performance of IMC programs is more sensitive to the size of input dataset, which is known to be difficult to be incorporated into a performance model due to its complex effects on performance; 2) the number of performance-critical configuration parameters in IMC is much larger than ODC (more than 40 vs. around 10), the high dimensionality requires more sophisticated models to achieve high accuracy. To address this challenge, we propose DAC, a datasize-aware auto-tuning approach to efficiently identify the high dimensional configuration for a given IMC program to achieve optimal performance on a given cluster. DAC is a significant advance over the state-of-the-art because it can take the size of input dataset and 41 configuration parameters as the parameters of the performance model for a given IMC program, --- unprecedented in previous work. It is made possible by two key techniques: 1) Hierarchical Modeling (HM), which combines a number of individual sub-models in a hierarchical manner; 2) Genetic Algorithm (GA) is employed to search the optimal configuration. To evaluate DAC, we use six typical Spark programs, each with five different input dataset sizes. The evaluation results show that DAC improves the performance of six typical Spark programs, each with five different input dataset sizes compared to default configurations by a factor of 30.4x on average and up to 89x. We also report that the geometric mean speedups of DAC over configurations by default, expert, and RFHOC are 15.4x, 2.3x, and 1.5x, respectively.
Zhibin Yu 0001, Zhendong Bei, Xuehai Qian
ASPLOS2
2018 Configuring in-memory cluster computing using random forest
Zhendong Bei, Zhibin Yu 0001, Ni Luo, Chuntao Jiang, Cheng-Zhong Xu 0001, Shengzhong Feng
Future Gener. Comput. Syst.1
2018 MIA: Metric Importance Analysis for Big Data Workload Characterization
abstract
Data analytics is at the foundation of both high-quality products and services in modern economies and societies. Big data workloads run on complex large-scale computing clusters, which implies significant challenges for deeply understanding and characterizing overall system performance. In general, performance is affected by many factors at multiple layers in the system stack, hence it is challenging to identify the key metrics when understanding big data workload performance. In this paper, we propose a novel workload characterization methodology using ensemble learning, called Metric Importance Analysis (MIA), to quantify the respective importance of workload metrics. By focusing on the most important metrics, MIA reduces the complexity of the analysis without losing information. Moreover, we develop the MIA-based Kiviat Plot (MKP) and Benchmark Similarity Matrix (BSM) which provide more insightful information than the traditional linkage clustering based dendrogram to visualize program behavior (dis)similarity. To demonstrate the applicability of MIA, we use it to characterize three big data benchmark suites: HiBench, CloudRank-D and SZTS. The results show that MIA is able to characterize complex big data workloads in a simple, intuitive manner, and reveal interesting insights. Moreover, through a case study, we demonstrate that tuning the configuration parameters related to the important metrics found by MIA results in higher performance improvements than through tuning the parameters related to the less important ones.
Zhibin Yu 0001, Lieven Eeckhout, Zhendong Bei, Avi Mendelson, Cheng-Zhong Xu 0001
IEEE Trans. Parallel Distributed Syst.4
2016 RFHOC: A Random-Forest Approach to Auto-Tuning Hadoop's Configuration
abstract
Hadoop is a widely-used implementation framework of the MapReduce programming model for large-scale data processing. Hadoop performance however is significantly affected by the settings of the Hadoop configuration parameters. Unfortunately, manually tuning these parameters is very time-consuming, if at all practical. This paper proposes an approach, called RFHOC, to automatically tune the Hadoop configuration parameters for optimized performance for a given application running on a given cluster. RFHOC constructs two ensembles of performance models using a random-forest approach for the map and reduce stage respectively. Leveraging these models, RFHOC employs a genetic algorithm to automatically search the Hadoop configuration space. The evaluation of RFHOC using five typical Hadoop programs, each with five different input data sets, shows that it achieves a performance speedup by a factor of 2.11$\times$on average and up to 7.4$\times$over the recently proposed cost-based optimization (CBO) approach. In addition, RFHOC's performance benefit increases with input data set size.
Zhendong Bei, Zhibin Yu 0001, Cheng-Zhong Xu 0001, Lieven Eeckhout, Shengzhong Feng
IEEE Trans. Parallel Distributed Syst.1
2013 A characterization of big data benchmarks
abstract
Recently, big data has been evolved into a buzzword from academia to industry all over the world. Benchmarks are important tools for evaluating an IT system. However, benchmarking big data systems is much more challenging than ever before. First, big data systems are still in their infant stage and consequently they are not well understood. Second, big data systems are more complicated compared to previous systems such as a single node computing platform. While some researchers started to design benchmarks for big data systems, they do not consider the redundancy between their benchmarks. Moreover, they use artificial input data sets rather than real world data for their benchmarks. It is therefore unclear whether these benchmarks can be used to precisely evaluate the performance of big data systems. In this paper, we first analyze the redundancy among benchmarks from ICTBench, HiBench and typical workloads from real world applications: spatio-temporal data analysis for Shenzhen transportation system. Subsequently, we present an initial idea of a big data benchmark suite for spatio-temporal data. There are three findings in this work: (1) redundancy exists in these pioneering benchmark suites and some of them can be removed safely. (2) The workload behavior of trajectory data analysis applications is dramatically affected by their input data sets. (3) The benchmarks created for academic research cannot represent the cases of real world applications.
Zhibin Yu 0001, Zhendong Bei, Juanjuan Zhao 0001, Fan Zhang 0019, Yubin Zou, Ye Li 0002, Cheng-Zhong Xu 0001
IEEE BigData3