Mike Ji

dblp:210/8962 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 45% Energy-efficient computing · 26% Parallel and multicore computing · 19%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
1.132019
Heterogeneity Aware Workload Management in Distributed Sustainable Datacenters · IEEE Trans. Parallel Distributed Syst. 2019
Addressing Skewness in Iterative ML Jobs with Parameter Partition · INFOCOM 2019
Energy Efficiency Aware Task Assignment with DVFS in Heterogeneous Hadoop Clusters · IEEE Trans. Parallel Distributed Syst. 2018
Parallel and multicore computing › parallel scheduling
data-parallel job scheduling
0.412019
Addressing Skewness in Iterative ML Jobs with Parameter Partition · INFOCOM 2019
Energy-efficient computing
green datacenter
0.412019
Heterogeneity Aware Workload Management in Distributed Sustainable Datacenters · IEEE Trans. Parallel Distributed Syst. 2019
Energy-efficient computing › energy-aware scheduling
renewable energy-aware scheduling
0.412019
Heterogeneity Aware Workload Management in Distributed Sustainable Datacenters · IEEE Trans. Parallel Distributed Syst. 2019
Cloud and datacenter computing › resource management
workload management
0.412019
Heterogeneity Aware Workload Management in Distributed Sustainable Datacenters · IEEE Trans. Parallel Distributed Syst. 2019
Energy-efficient computing
energy-aware scheduling
0.312018
Energy Efficiency Aware Task Assignment with DVFS in Heterogeneous Hadoop Clusters · IEEE Trans. Parallel Distributed Syst. 2018
High-performance computing › cluster computing
heterogeneous clusters
0.312018
Energy Efficiency Aware Task Assignment with DVFS in Heterogeneous Hadoop Clusters · IEEE Trans. Parallel Distributed Syst. 2018
Parallel and multicore computing
task allocation
0.312018
Energy Efficiency Aware Task Assignment with DVFS in Heterogeneous Hadoop Clusters · IEEE Trans. Parallel Distributed Syst. 2018
Machine learning › Efficient and distributed learning
distributed training
0.112019
Addressing Skewness in Iterative ML Jobs with Parameter Partition · INFOCOM 2019
Cloud and datacenter computing › datacenter architecture
geo-distributed datacenters
0.112019
Heterogeneity Aware Workload Management in Distributed Sustainable Datacenters · IEEE Trans. Parallel Distributed Syst. 2019
Parallel and multicore computing › data-parallel programming
mapreduce
0.112018
Energy Efficiency Aware Task Assignment with DVFS in Heterogeneous Hadoop Clusters · IEEE Trans. Parallel Distributed Syst. 2018
Performance modeling and evaluation
performance diagnosis
0.112018
Profiling distributed systems in lightweight virtualized environments with logs and resource metrics · HPDC 2018

Methods — techniques the papers use, named apart from their topics

parameter partition · 0.8capacity model · 0.8nonlinear programming · 0.4constrained optimization · 0.4resource monitoring · 0.3log analysis · 0.3ant colony optimization · 0.3DVFS · 0.3
YearPublicationVenuePosition
2019 Addressing Skewness in Iterative ML Jobs with Parameter Partition
abstract
Computational skewness is a significant challenge in multi-tenant data-parallel clusters that introduce dynamic heterogeneity of machine capacity in distributed data processing. Previous efforts to addressing skewness mostly focus on batch jobs based on the assumption that processing time is linearly dependent on the size of partitioned data. However, they are illsuited for iterative machine learning (ML) jobs, which (1) exhibit a non-linear relationship between the size of partitioned parameters and processing time within each iteration, and (2) show an explicit binding relationship between input data and parameters for parameter update. In this paper, we present FlexPara, a parameter partition approach that leverages the non-linear relationship and provisions adaptive tasks to match the distinct machine capacity so as to address the skewness in iterative ML jobs on data-parallel clusters. FlexPara first predicts task processing time based on a capacity model designed for iterative ML jobs without the linear assumption. It then partitions parameters to parallel tasks through proactive parameter reassignment. Such reassignment can significantly reduce network transmission cost incurred by input data movement due to the binding relationship. We implement FlexPara in Spark and evaluate it with various ML jobs. Experimental results show that compared to hash partition, FlexPara speeds up the execution by up to 54% and 43% in private and NSF Chameleon clusters, respectively.
Wei Chen 0038, Xiaobo Zhou 0002, Sang-Yoon Chang, Mike Ji
INFOCOM5
2019 Heterogeneity Aware Workload Management in Distributed Sustainable Datacenters
abstract
The tremendous growth of cloud computing and large-scale data analytics highlight the importance of reducing datacenter power consumption and environmental impact of brown energy. While many Internet service operators have at least partially powered their datacenters by green energy, it is challenging to effectively utilize green energy due to the intermittency of renewable sources, such as solar or wind. We find that the geographical diversity of internet-scale services can be carefully scheduled to improve the efficiency of applying green energy in datacenters. In this paper, we propose a holistic heterogeneity-aware cloud workload management approach, sCloud, that aims to maximize the system goodput in distributed self-sustainable datacenters. sCloud adaptively places the transactional workload to distributed datacenters, allocates the available resource to heterogeneous workloads in each datacenter, and migrates batch jobs across datacenters, while taking into account the green power availability and QoS requirements. We formulate the transactional workload placement as a constrained optimization problem that can be solved by nonlinear programming. Then, we propose a batch job migration algorithm to further improve the system goodput when the green power supply varies widely at different locations. Finally, we extend sCloud by integrating a flexible batch job manager to dynamically control the job execution progress without violating the deadlines. We have implemented sCloud in a university cloud testbed with real-world weather conditions and workload traces. Experimental results demonstrate sCloud can achieve near-to-optimal system performance while being resilient to dynamic power availability. sCloud with the flexible batch job management approach outperforms a heterogeneity-oblivious approach by 37 percent in improving system goodput and 33 percent in reducing QoS violations.
Dazhao Cheng, Xiaobo Zhou 0002, Zhijun Ding, Yu Wang 0003, Mike Ji
IEEE Trans. Parallel Distributed Syst.5
2018 Profiling distributed systems in lightweight virtualized environments with logs and resource metrics
abstract
Understanding and troubleshooting distributed systems in the cloud is considered a very difficult problem because the execution of a single user request is distributed to multiple machines. Further, the multi-tenancy nature of cloud environments further introduces interference that causes performance issues. Most existing troubleshooting tools either focus on log analysis or intrusive tracing methods, leaving resource usage monitoring unexplored.
Aidi Pi, Wei Chen 0038, Xiaobo Zhou 0002, Mike Ji
HPDC4
2018 Energy Efficiency Aware Task Assignment with DVFS in Heterogeneous Hadoop Clusters
abstract
While Hadoop ecosystems become increasingly important for practitioners of large-scale data analysis, they also incur tremendous energy cost. This trend is driving up the need for designing energy-efficient Hadoop clusters in order to reduce the operational costs and the carbon emission associated with its energy consumption. However, despite extensive studies of the problem, existing approaches for energy efficiency have not fully considered the heterogeneity of both workload and machine hardware found in production environments. In this paper, we find that heterogeneity-oblivious task assignment approaches are detrimental to both performance and energy efficiency of Hadoop clusters. Our observation shows that even heterogeneity-aware techniques that aim to reduce the job completion time do not guarantee a reduction in energy consumption of heterogeneous machines. We propose a heterogeneity-aware task assignment approach, E-Ant, that aims to improve the overall energy consumption in a heterogeneous Hadoop cluster without sacrificing job performance. It adaptively schedules heterogeneous workloads on energy-efficient machines, without a priori knowledge of the workload properties. E-Ant employs an ant colony optimization approach that generates task assignment solutions based on the feedback of each task's energy consumption reported by Hadoop TaskTrackers in an agile way. Furthermore, we integrate DVFS technique with E-Ant to further improve the energy efficiency of heterogeneous Hadoop clusters. It relies on a DVFS controller to dynamically scale the CPU frequency of each slave machine in response to time-varying resource demands. Experimental results on a heterogeneous cluster with varying hardware capabilities show that E-Ant with DVFS improves the overall energy savings for a synthetic workload from Microsoft by 23 and 17 percent compared to Fair Scheduler and Tarazu, respectively.
Dazhao Cheng, Xiaobo Zhou 0002, Palden Lama, Mike Ji, Changjun Jiang 0002
IEEE Trans. Parallel Distributed Syst.4