Naunidh Singh

dblp:333/2227 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 91% Embedded and real-time systems · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › job scheduling › batch scheduling
backfilling
0.612022
DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing · IEEE Trans. Parallel Distributed Syst. 2022
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling
0.612022
DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing · IEEE Trans. Parallel Distributed Syst. 2022
Cloud and datacenter computing › job scheduling
reinforcement-learning-based scheduling
0.612022
DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing · IEEE Trans. Parallel Distributed Syst. 2022
Embedded and real-time systems › real-time scheduling
resource reservation
0.212022
DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing · IEEE Trans. Parallel Distributed Syst. 2022

Methods — techniques the papers use, named apart from their topics

hierarchical neural network · 0.6deep reinforcement learning · 0.6
YearPublicationVenuePosition
2022 DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing
abstract
Cluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. The experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%.
Yuping Fan, Boyang Li 0018, Dustin Favorite, Naunidh Singh, John T. Childers, Paul M. Rich, William E. Allcock, Michael E. Papka, Zhiling Lan
IEEE Trans. Parallel Distributed Syst.4