Yijia Zhang 0002

dblp:65/2747-2 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0001-6925-7777ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy Efficiency
abstract
Recent advancements in deep learning have significantly increased AI processors' energy consumption, which is becoming a critical factor limiting AI development. Dynamic Voltage and Frequency Scaling (DVFS) stands as a key method in power optimization. However, due to the latency of DVFS control in AI processors, previous works typically apply DVFS control at the granularity of a program's entire duration or sub-phases, rather than at the level of AI operators.
Yijia Zhang 0002, Fuchun Wei, Bingqiang Wang, Yanlin Liu, Zhiheng Hu, Xiaoxin Xu, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
ASPLOS (1)2
2022 HPC Data Center Participation in Demand Response: An Adaptive Policy With QoS Assurance
abstract
Demand response programs help stabilize the electricity grid by providing monetary stimulus to consumers if they regulate their power consumption following market requirements. Regulation service, a market that requires participants to regulate power by following a signal updated every few seconds, is particularly beneficial to HPC data centers since data centers are capable of increasing/decreasing power consumption owing to the flexibility in running workloads and the availability of power control mechanisms. While prior works have explored how data centers can provide regulation service reserves, Quality-of-Service (QoS) provisioning for the jobs running at the data centers has not been considered. In this work, we propose an Adaptive policy with QoS Assurance that enables data centers to participate in regulation service programs with assurance on job QoS. Our policy regulates data center power through job scheduling and server power capping. QoS assurance is achieved by applying a queueing-theoretic result to our job scheduling strategy. We evaluate our policy by experiments on a real cluster. Our results demonstrate that the proposed policy reduces electricity costs by 25-56% while providing QoS assurance. On the other hand, the baseline policies cannot meet QoS constraints in 9 of the 14 workload traces tested.
Yijia Zhang 0002, Daniel C. Wilson, Ioannis Paschalidis, Ayse K. Coskun
IEEE Trans. Sustain. Comput.1
2021 A Data Center Demand Response Policy for Real-World Workload Scenarios in HPC
abstract
Demand response programs offer an opportunity for large power consumers to save on electricity costs by modulating their power consumption in response to demand changes in the electricity grid. Multiple types of such programs exist; for example, regulation service programs enable a consumer to bid for a sustainable amount of power draw over a time period, along with a reserve amount they are able to provide at request of the electricity service provider. Data centers offer unique capabilities to participate in these programs since they have significant capacity to modify their power consumption through workload scheduling and CPU power limiting. This paper proposes a novel power management policy and a bidding policy that enable data centers to participate in regulation service programs under real-world constraints. The power management policy schedules computing jobs and applies server power-capping under both the constraints of power programs and the constraints of job Quality-of-Service (QoS). Simulations with workload traces from a real data center show that the proposed policies enable data centers to meet both the requirement of regulation service programs and the QoS requirement of jobs. We demonstrate that, by applying our policies, data centers can save their electricity costs by 10% while abiding by all the QoS constraints in a real-world scenario.
Yijia Zhang 0002, Daniel C. Wilson, Ioannis Paschalidis, Ayse K. Coskun
DATE1
2020 Quantifying the impact of network congestion on application performance and network metrics
abstract
In modern high-performance computing (HPC) systems, network congestion is an important factor that contributes to performance degradation. However, how network congestion impacts application performance is not fully understood. As Aries network, a recent HPC network architecture featuring a dragonfly topology, is equipped with network counters measuring packet transmission statistics on each router, these network metrics can potentially be utilized to understand network performance. In this work, by experiments on a large HPC system, we quantify the impact of network congestion on various applications' performance in terms of execution time, and we correlate application performance with network metrics. Our results demonstrate diverse impacts of network congestion: while applications with intensive MPI operations (such as HACC and MILC) suffer from more than 40% extension in their execution times under network congestion, applications with less intensive MPI operations (such as Graph500 and HPCG) are mostly not affected. We also demonstrate that a stall-to-flit ratio metric derived from Aries network counters is positively correlated with performance degradation and, thus, this metric can serve as an indicator of network congestion in HPC systems.
Yijia Zhang 0002, Taylor L. Groves, Brandon Cook 0001, Nicholas J. Wright, Ayse K. Coskun
CLUSTER1
2019 HPAS: An HPC Performance Anomaly Suite for Reproducing Performance Variations
abstract
Modern high performance computing (HPC) systems, including supercomputers, routinely suffer from substantial performance variations. The same application with the same input can have more than 100% performance variation, and such variations cause reduced efficiency and wasted resources. There have been recent studies on performance variability and on designing automated methods for diagnosing "anomalies" that cause performance variability. These studies either observe data collected from HPC systems, or they rely on synthetic reproduction of performance variability scenarios. However, there is no standardized way of creating performance variability inducing synthetic anomalies; so, researchers rely on designing ad-hoc methods for reproducing performance variability.
Emre Ates, Yijia Zhang 0002, Burak Aksar, Jim M. Brandt, Vitus J. Leung, Manuel Egele, Ayse K. Coskun
ICPP2
2019 Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning
abstract
As the size and complexity of high performance computing (HPC) systems grow in line with advancements in hardware and software technology, HPC systems increasingly suffer from performance variations due to shared resource contention as well as software- and hardware-related problems. Such performance variations can lead to failures and inefficiencies, which impact the cost and resilience of HPC systems. To minimize the impact of performance variations, one must quickly and accurately detect and diagnose the anomalies that cause the variations and take mitigating actions. However, it is difficult to identify anomalies based on the voluminous, high-dimensional, and noisy data collected by system monitoring infrastructures. This paper presents a novel machine learning based framework to automatically diagnose performance anomalies at runtime. Our framework leverages historical resource usage data to extract signatures of previously-observed anomalies. We first convert collected time series data into easy-to-compute statistical features. We then identify the features that are required to detect anomalies, and extract the signatures of these anomalies. At runtime, we use these signatures to diagnose anomalies with negligible overhead. We evaluate our framework using experiments on a real-world HPC supercomputer and demonstrate that our approach successfully identifies 98 percent of injected anomalies and consistently outperforms existing anomaly diagnosis techniques.
Ozan Tuncer, Emre Ates, Yijia Zhang 0002, Ata Turk, Jim M. Brandt, Vitus J. Leung, Manuel Egele, Ayse K. Coskun
IEEE Trans. Parallel Distributed Syst.3
2018 Level-Spread: A New Job Allocation Policy for Dragonfly Networks
abstract
The dragonfly network topology has attracted attention in recent years owing to its high radix and constant diameter. However, the influence of job allocation on communication time in dragonfly networks is not fully understood. Recent studies have shown that random allocation is better at balancing the network traffic, while compact allocation is better at harnessing the locality in dragonfly groups. Based on these observations, this paper introduces a novel allocation policy called Level-Spread for dragonfly networks. This policy spreads jobs within the smallest network level that a given job can fit in at the time of its allocation. In this way, it simultaneously harnesses node adjacency and balances link congestion. To evaluate the performance of Level-Spread, we run packet-level network simulations using a diverse set of application communication patterns, job sizes, and communication intensities. We also explore the impact of network properties such as the number of groups, number of routers per group, machine utilization level, and global link bandwidth. Level-Spread reduces the communication overhead by 16% on average (and up to 71%) compared to the state-of-the-art allocation policies.
Yijia Zhang 0002, Ozan Tuncer, Fulya Kaplan, Katzalin Olcoz, Vitus J. Leung, Ayse K. Coskun
IPDPS1