Zhengxiong Hou

dblp:46/5393 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Container Workload Prediction Using Deep Domain Adaptation in Transfer Learning
Yunlan Wang, Tianhai Zhao, Jianhua Gu, Zhengxiong Hou, Chengwen Zhong
Euro-Par (1)6
2024 Stochastic Network Calculus Based Quality of Service Guarantee for Multi-class Traffic
abstract
With the rapid development of cloud computing technology, a variety of emerging network traffic types have higher quality of service(QoS) requirements. However, the traditional Internet’s best-effort service cannot meet the demands of cloud computing applications. Therefore, we propose a stochastic network calculus(SNC) based QoS guarantee mechanism consisting of two parts. The first part is MTACC, a multi-threshold adaptive admission control algorithm based on network calculus. MTACC classifies network traffic, calculates resource requirements, and introduces admission probabilities and demarcation parameters. As the network load reaches various thresholds, MTACC adjusts the demarcation parameter to modify the admission probabilities accordingly. The second part is TSRA, a two-stage resource allocation algorithm. In the first stage, basic resources are allocated to ensure the minimum resource requirements of the network traffic based on SNC. In the second stage, additional idle resources are allocated to improve the QoS of the network traffic based on the resource utility function. Finally, we simulate the QoS guarantee techniques using NS3. By comparing with Simple Sum and SCAC, we verify the effectiveness of our proposed QoS guarantee techniques. They provide robust performance guarantees for multi-class traffic such as delay-sensitive, bandwidth-sensitive, and packet loss-sensitive traffic.
Yunlan Wang, Tianhai Zhao, YongKuo Hu, Jianhua Gu, Zhengxiong Hou
IPCCC6
2024 Qualitative QoS-aware Scheduling of Moldable Parallel Jobs on HPC Clusters*
abstract
In service oriented high-performance computing (HPC) clusters, end users have various Quality of Service (QoS) requirements. Most of the existing research work focuses on quantitative QoS requirements, such as deadlines, for rigid jobs. While, in many cases, it is more convenient for users to qualitatively state QoS requirements (such as performance-sensitive) at the submission of their jobs. Almost all kinds of QoS requirements will be greatly impacted by job scheduling, which determine the degree of job parallelism, execution time and waiting time, etc. Most modern parallel applications are moldable in the sense that they can choose a resource allocation before execution. Traditional sequential job scheduling mechanism with fixed resource allocation appears to be an obstacle to improve QoS for end users. To address this issue, we propose a novel qualitative QoS-aware sub-queue simultaneous scheduling method for moldable parallel jobs (with variable resource allocation) on HPC Clusters. We first define the qualitative QoS models for end users, then present our sub-queue simultaneous scheduling method, including job sequencing and resource allocation algorithms for a set of moldable parallel jobs on multi-core clusters. Our method can efficiently sequence queuing jobs and allocate appropriate resources for simultaneously running some performance-sensitive jobs in a sub-queue rather than running them one by one. Experimental results demonstrate the effectiveness of our method to improving QoS for end users.
Zhengxiong Hou, Yubing Liu, Hong Shen 0001, Jianhua Gu
ISPA1
2024 Optimizing job scheduling by using broad learning to predict execution times on HPC clusters
Zhengxiong Hou, Hong Shen 0001, Qiying Feng, Zhiqi Lv, Xingshe Zhou 0001, Jianhua Gu
CCF Trans. High Perform. Comput.1
2022 Prediction of job characteristics for intelligent resource allocation in HPC systems: a survey and future directions
Zhengxiong Hou, Hong Shen 0001, Xingshe Zhou 0001, Jianhua Gu, Yunlan Wang, Tianhai Zhao
Frontiers Comput. Sci.1
2019 Machine Learning Based Performance Analysis and Prediction of Jobs on a HPC Cluster
abstract
There are a lot of middle-class or small-class high-performance computing clusters at universities and research institutes, etc. Large volumes of job logs have been accumulated after many years of operation. In this paper, on the basis of accumulated job logs on a high-performance computing cluster, we examine and analyze the job logs. Then, we study machine learning based performance analysis and prediction methods for parallel jobs. Various machine learning methods such as multivariate linear fitting, artificial neural network are used to build performance prediction models. We compare the errors of each model, and select the optimal prediction model for different users. The experimental results show that we can obtain reasonable prediction accuracy using the selected machine learning algorithms.
Zhengxiong Hou, Shuxin Zhao, Yunlan Wang, Jianhua Gu, Xingshe Zhou 0001
PDCAT1
2018 Prediction Method of Blasting Vibration by Optimized GEP Based on Spark
abstract
In order to minimize the damage of engineering blasting vibration, it is very important to accurately predict the blasting peak velocity. The GEP algorithm is used to analyze the relationship between blasting vibration and related parameters. Aiming to improve the time performance of GEP in dealing with large-scale engineering blasting data, the gene structure of GEP is adjusted to head, body and tail. Then the GEP algorithm is paralleled on Spark cluster. The optimized parallel GEP can greatly improves the global search efficiency. The experimental results show that the method can obviously reduce the running time and can improve the prediction accuracy in most cases.
Yunlan Wang, Tianhai Zhao, Zhengxiong Hou, Hussain Khanzada Muzammil
COMPSAC (2)4
2010 ASAAS: Application Software as a Service for High Performance Cloud Computing
abstract
Currently, SAAS (Software as a Service) solutions are usually provided for business, such as salesforce.com. Few work focus on the application software for high performance scientific computing. However, in the high performance cloud computing environment, traditional application software is not intrinsically service oriented. And the limitation of traditional software licenses is a bottleneck problem for large scale of dynamic users. To enable on-demand services for applications, we propose a solution: Application Software as a Service (ASAAS). It provides a web services portal, an on-demand software license service for the users. Application software is wrapped as web services on the basis of underlying computational resources. With a pay-for-use mode, there is no limitation for the licenses any more. The instant service rate, average job response time, and cost are analyzed for an evaluation. A case of implementation and the evaluation show that ASAAS can bring a much better effect than traditional mechanism.
Zhengxiong Hou, Xingshe Zhou 0001, Jianhua Gu, Yunlan Wang, Tianhai Zhao
HPCC1
2009 Experiences of On-Demand Execution for Large Scale Parameter Sweep Applications on OSG by Swift
abstract
Large scale parameter sweep application (PSA) is one of the main grid applications, which may have different characteristics and demands. In this paper, we describe how to use swift to enable the on-demand execution of large scale PSA on open science grid (OSG). The basic on-demand concept means providing appropriate grid resources for the application, which is decided by the characteristics and demands of the application. So we can get high reliability, efficiency, and scalability for large scale independent PSA jobs on OSG. The main on-demand policies include: trust based site selection and pre-selection; scheduling policy on-demand configuration; clustering for small jobs; adaptive execution and automatic data staging; divide and conquer for the scalability. Some usage examples of swift for executing large scale PSA are presented, such as dock, blast. The experimental results for the performance of different policies are presented, with a benchmarking workload size of 10,000 jobs.
Zhengxiong Hou, Michael Wilde, Mihael Hategan, Xingshe Zhou 0001, Ian T. Foster, Ben Clifford
HPCC1
2009 A Runtime Reputation Based Grid Resource Selection Algorithm on the Open Science Grid
abstract
The scheduling and execution for grid application is an important problem in the grid environment. To get the high reliability and efficiency, we propose a runtime reputation based grid resource selection algorithm. According to the accumulated raw score, the runtime reputation degree for a grid resource is quantified as an evaluating score in the runtime of an application. Instead of being dependent on the historical experiences, it is dynamically adaptive to the runtime availability, load, and performance of the grid resources. The execution framework on the grid is based on Globus Toolkit and Swift system. In a real production grid, Open Science Grid (OSG), a typical grid application with large scale independent jobs was experimented, which was based on BLAST application. The experimental results for the performance of different policies are presented, with a benchmarking workload size of 10,000 jobs. The runtime reputation and behavior statistics for the grid resources are also presented.
Zhengxiong Hou, Xingshe Zhou 0001, Michael Wilde, Jianhua Gu, Mihael Hategan
ICPADS1