Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chao Chen 0024

dblp:66/3019-24 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0003-1960-4042ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Hardware reliability and fault tolerance · 29% Distributed systems · 15% GPUs and heterogeneous computing · 12%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 51% Program analysis · 49%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
1.022022
Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022
CARE: compiler-assisted recovery from soft failures · SC 2019
Compilers and program optimization
compiler-based fault tolerance
0.722022
Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022
CARE: compiler-assisted recovery from soft failures · SC 2019
Program analysis
loop analysis
0.712023
Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop Characteristics · Proc. ACM Program. Lang. 2023
Cloud and datacenter computing
job scheduling
0.712023
Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop Characteristics · Proc. ACM Program. Lang. 2023
Performance modeling and evaluation
workload characterization
0.712023
Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop Characteristics · Proc. ACM Program. Lang. 2023
Processor architecture and microarchitecture › instruction scheduling
compiler-based scheduling
0.612022
CASE: a compiler-assisted SchEduling framework for multi-GPU systems · PPoPP 2022
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU scheduling
0.612022
CASE: a compiler-assisted SchEduling framework for multi-GPU systems · PPoPP 2022
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.612022
CASE: a compiler-assisted SchEduling framework for multi-GPU systems · PPoPP 2022
Hardware reliability and fault tolerance › error recovery
transient fault recovery
0.612022
Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022
Hardware reliability and fault tolerance › soft errors
transient fault
0.412019
CARE: compiler-assisted recovery from soft failures · SC 2019
Hardware reliability and fault tolerance › soft errors
silent data corruption
0.312018
LADR: low-cost application-level detector for reducing silent output corruptions · HPDC 2018
Hardware reliability and fault tolerance
soft errors
0.312018
LADR: low-cost application-level detector for reducing silent output corruptions · HPDC 2018
Storage systems
crash recovery
0.212022
Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022
Hardware reliability and fault tolerance
transient errors
0.212022
Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022
High-performance computing
scientific computing systems
0.112018
LADR: low-cost application-level detector for reducing silent output corruptions · HPDC 2018

Methods — techniques the papers use, named apart from their topics

compiler instrumentation · 1.9machine learning · 1.3strength reduction · 1.1loop unrolling · 1.1code transformation · 1.1compiler-assisted recovery · 0.8checkpointing · 0.8throughput-oriented scheduling · 0.6application-level detection · 0.3
YearPublicationVenuePosition
2023 Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop Characteristics
abstract
Efficient management of shared resources is a critical problem in high-performance computing (HPC) environments. Existing workload management systems often promote non-sharing of resources among different co-executing applications to achieve performance isolation. Such schemes lead to poor resource utilization and suboptimal process throughput, adversely affecting user productivity. Tackling this problem in a scalable fashion is extremely challenging, since it requires the workload scheduler to possess an in-depth knowledge about various application resource requirements and runtime phases at fine granularities within individual applications. In this work, we show that applications’ resource requirements and execution phase behaviour can be captured in a scalable and lightweight manner at runtime by estimating important program artifacts termed as “ dynamic loop characteristics ”. Specifically, we propose a solution to the problem of efficient workload scheduling by designing a compiler and runtime cooperative framework that leverages novel loop-based compiler analysis for resource allocation . We present Beacons Framework , an end-to-end compiler and scheduling framework, that estimates dynamic loop characteristics, encapsulates them in compiler-instrumented beacons in an application, and broadcasts them during application runtime, for proactive workload scheduling. We focus on estimating four important loop characteristics : loop trip-count , loop timing , loop memory footprint , and loop data-reuse behaviour , through a combination of compiler analysis and machine learning. The novelty of the Beacons Framework also lies in its ability to tackle irregular loops that exhibit complex control flow with indeterminate loop bounds involving structure fields, aliased variables and function calls , which are highly prevalent in modern workloads. At the backend, Beacons Framework entails a proactive workload scheduler that leverages the runtime information to orchestrate aggressive process co-locations, for maximizing resource concurrency, without causing cache thrashing . Our results show that Beacons Framework can predict different loop characteristics with an accuracy of 85% to 95% on average, and the proactive scheduler obtains an average throughput improvement of 1.9x (up to 3.2x ) over the state-of-the-art schedulers on an Amazon Graviton2 machine on consolidated workloads involving 1000-10000 co-executing processes, across 51 benchmarks.
Girish Mururu, Sharjeel Khan, Bodhisatwa Chatterjee, Chao Chen 0024, Chris Porter, Ada Gavrilovska, Santosh Pande
Proc. ACM Program. Lang.4
2022 CASE: a compiler-assisted SchEduling framework for multi-GPU systems
abstract
Modern computing platforms tend to deploy multiple GPUs on a single node to boost performance. GPUs have large computing capacities and are an expensive resource. Increasing their utilization without causing performance degradation of individual workloads is an important and challenging problem. Although services such as NVIDIA's MPS allow multiple cooperative kernels to simultaneously run on a single device, they do not solve the co-execution problem for uncooperative, independent kernels on such a multi-GPU system. To tackle this problem, we propose CASE --- a fully automated compiler-assisted scheduling framework. During the compilation of an application, CASE constructs GPU tasks from CUDA programs and instruments the code with a probe before each one. At runtime, each probe conveys information about its task's resource requirements such as memory and the number of streaming multiprocessor (SMs) needed to a user-level scheduler. The scheduler then places each task onto a suitable device by employing a policy appropriate to the system. In our prototype, a throughput-oriented scheduling policy is implemented to evaluate our resource-aware scheduling framework. The Rodinia benchmark suite and the Darknet neural network framework were used in our evaluation. The results show that, as compared to existing state-of-the-art methods, CASE improves throughput by up to 2.5X for Rodinia, and up to 2.7X for Darknet on modern NVIDIA GPU platforms, mainly due to the fact that it improves the average system utilization by up to 3.36X and the job turnaround time by up to 4.9X. Meanwhile, it limits individual kernel performance degradation within 2.5%. CASE achieved peak system utilization of 78% for Rodinia and 80% for Darknet on a 4XV100 system.
Chao Chen 0024, Chris Porter, Santosh Pande
PPoPP1
2022 Near-Zero Downtime Recovery From Transient-Error-Induced Crashes
abstract
Due to the system scaling,transient errorscaused by external noise, e.g., heat fluxes and particle strikes, have become a growing concern for the current and upcoming exa-scale high-performance-computing (HPC) systems. Applications running on these systems are expected to experience transient errors more frequently than ever before, which will either lead them to generate incorrect outputs or cause them to crash. However, since such errors are still quite rare as compared to no-fault cases, desirable solutions call for low/no-overhead systems that do not compromise the performance under no-fault conditions and also allow very fast fault recovery to minimize downtime. In this article, we presentIterPro, a light-weight compiler-assisted resilience technique to quickly and accurately recover processes from transient-error-induced crashes. During the compilation of applications,IterProconstructs a set of recovery kernels for crash-prone instructions. These recovery kernels are executed to repair the corrupted process states on-the-fly upon occurrences of errors, enabling applications to continue their executions instead of being terminated. When constructing recovery kernels,IterProexploits side effects introduced by induction variable based code optimization techniques based on loop unrolling and strength reduction to improve its recovery capability. To this end, two new code transformation passes are introduced to expose the side effects for resilience purposes. We evaluatedIterProwith 4 scientific workloads as well as the NPB benchmarks suite. During their normal execution,IterProincurs almostzeroruntime overhead and a small, fixed27MBmemory overhead. Meanwhile,IterProcan recover on an average 83.55 percent of crash-causing errors within dozens of milliseconds with negligible downtime. We also evaluatedIterProwith parallel jobs running on 3072 cores and showed thatIterProcan successfully mask the impact of crash-causing errors by providing almost uninterrupted execution. Finally, we present our preliminary evaluation result for BLAS, which shows thatIterProis capable of recovering failures in libraries with a very high coverage rate of 83 percent and negligible overheads. With such an effective recovery mechanism,IterProcould tremendously mitigate the overheads and resource requirements of the resilience subsystem in future exa-scale systems.
Chao Chen 0024, Greg Eisenhauer, Santosh Pande
IEEE Trans. Parallel Distributed Syst.1
2019 CARE: compiler-assisted recovery from soft failures
abstract
As processors continue to boost the system performance with higher circuit density, shrinking process technology and near-threshold voltage (NTV) operations, they are projected to be more vulnerable to transient faults, which have become one of the major concerns for future extreme-scale HPC systems. Despite being relatively infrequent, crashes due to transient faults are incredibly disruptive, particularly for massively parallel jobs on supercomputers where they potentially kill the entire job, requiring an expensive rerun or restart from a checkpoint.
Chao Chen 0024, Greg Eisenhauer, Santosh Pande, Qiang Guan
SC1
2018 LADR: low-cost application-level detector for reducing silent output corruptions
abstract
Applications running on future high performance computing (HPC) systems are more likely to experience transient faults due to technology scaling trends with respect to higher circuit density, smaller transistor size and near-threshold voltage (NTV) operations. A transient fault could corrupt application state without warning, possibly leading to incorrect application output. Such errors are called silent data corruptions (SDCs).
Chao Chen 0024, Greg Eisenhauer, Matthew Wolf, Santosh Pande
HPDC1