EDBT 2026 Demo / reviewers in the wild / expert
Chao Chen 0024
dblp:66/3019-24
· DBLP profile ↗
5ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0003-1960-4042ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Hardware reliability and fault tolerance · 29% Distributed systems · 15% GPUs and heterogeneous computing · 12% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 51% Program analysis · 49% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
fault tolerance |
1.0 | 2 | 2022 | Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022 CARE: compiler-assisted recovery from soft failures · SC 2019 |
Compilers and program optimization
compiler-based fault tolerance |
0.7 | 2 | 2022 | Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022 CARE: compiler-assisted recovery from soft failures · SC 2019 |
Program analysis
loop analysis |
0.7 | 1 | 2023 | Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop Characteristics · Proc. ACM Program. Lang. 2023 |
Cloud and datacenter computing
job scheduling |
0.7 | 1 | 2023 | Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop Characteristics · Proc. ACM Program. Lang. 2023 |
Performance modeling and evaluation
workload characterization |
0.7 | 1 | 2023 | Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop Characteristics · Proc. ACM Program. Lang. 2023 |
Processor architecture and microarchitecture › instruction scheduling
compiler-based scheduling |
0.6 | 1 | 2022 | CASE: a compiler-assisted SchEduling framework for multi-GPU systems · PPoPP 2022 |
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU scheduling |
0.6 | 1 | 2022 | CASE: a compiler-assisted SchEduling framework for multi-GPU systems · PPoPP 2022 |
Parallel and multicore computing › parallel scheduling
runtime scheduling |
0.6 | 1 | 2022 | CASE: a compiler-assisted SchEduling framework for multi-GPU systems · PPoPP 2022 |
Hardware reliability and fault tolerance › error recovery
transient fault recovery |
0.6 | 1 | 2022 | Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022 |
Hardware reliability and fault tolerance › soft errors
transient fault |
0.4 | 1 | 2019 | CARE: compiler-assisted recovery from soft failures · SC 2019 |
Hardware reliability and fault tolerance › soft errors
silent data corruption |
0.3 | 1 | 2018 | LADR: low-cost application-level detector for reducing silent output corruptions · HPDC 2018 |
Hardware reliability and fault tolerance
soft errors |
0.3 | 1 | 2018 | LADR: low-cost application-level detector for reducing silent output corruptions · HPDC 2018 |
Storage systems
crash recovery |
0.2 | 1 | 2022 | Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022 |
Hardware reliability and fault tolerance
transient errors |
0.2 | 1 | 2022 | Near-Zero Downtime Recovery From Transient-Error-Induced Crashes · IEEE Trans. Parallel Distributed Syst. 2022 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2018 | LADR: low-cost application-level detector for reducing silent output corruptions · HPDC 2018 |
Methods — techniques the papers use, named apart from their topics
compiler instrumentation · 1.9machine learning · 1.3strength reduction · 1.1loop unrolling · 1.1code transformation · 1.1compiler-assisted recovery · 0.8checkpointing · 0.8throughput-oriented scheduling · 0.6application-level detection · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop CharacteristicsabstractEfficient management of shared resources is a critical problem in high-performance computing (HPC) environments. Existing workload management systems often promote non-sharing of resources among different co-executing applications to achieve performance isolation. Such schemes lead to poor resource utilization and suboptimal process throughput, adversely affecting user productivity. Tackling this problem in a scalable fashion is extremely challenging, since it requires the workload scheduler to possess an in-depth knowledge about various application resource requirements and runtime phases at fine granularities within individual applications. In this work, we show that applications’ resource requirements and execution phase behaviour can be captured in a scalable and lightweight manner at runtime by estimating important program artifacts termed as “ dynamic loop characteristics ”. Specifically, we propose a solution to the problem of efficient workload scheduling by designing a compiler and runtime cooperative framework that leverages novel loop-based compiler analysis for resource allocation . We present Beacons Framework , an end-to-end compiler and scheduling framework, that estimates dynamic loop characteristics, encapsulates them in compiler-instrumented beacons in an application, and broadcasts them during application runtime, for proactive workload scheduling. We focus on estimating four important loop characteristics : loop trip-count , loop timing , loop memory footprint , and loop data-reuse behaviour , through a combination of compiler analysis and machine learning. The novelty of the Beacons Framework also lies in its ability to tackle irregular loops that exhibit complex control flow with indeterminate loop bounds involving structure fields, aliased variables and function calls , which are highly prevalent in modern workloads. At the backend, Beacons Framework entails a proactive workload scheduler that leverages the runtime information to orchestrate aggressive process co-locations, for maximizing resource concurrency, without causing cache thrashing . Our results show that Beacons Framework can predict different loop characteristics with an accuracy of 85% to 95% on average, and the proactive scheduler obtains an average throughput improvement of 1.9x (up to 3.2x ) over the state-of-the-art schedulers on an Amazon Graviton2 machine on consolidated workloads involving 1000-10000 co-executing processes, across 51 benchmarks. Girish Mururu, Sharjeel Khan, Bodhisatwa Chatterjee, Chao Chen 0024, Chris Porter, Ada Gavrilovska, Santosh Pande |
Proc. ACM Program. Lang. | 4 |
| 2022 | CASE: a compiler-assisted SchEduling framework for multi-GPU systemsabstractModern computing platforms tend to deploy multiple GPUs on a single node to boost performance. GPUs have large computing capacities and are an expensive resource. Increasing their utilization without causing performance degradation of individual workloads is an important and challenging problem. Although services such as NVIDIA's MPS allow multiple cooperative kernels to simultaneously run on a single device, they do not solve the co-execution problem for uncooperative, independent kernels on such a multi-GPU system. To tackle this problem, we propose CASE --- a fully automated compiler-assisted scheduling framework. During the compilation of an application, CASE constructs GPU tasks from CUDA programs and instruments the code with a probe before each one. At runtime, each probe conveys information about its task's resource requirements such as memory and the number of streaming multiprocessor (SMs) needed to a user-level scheduler. The scheduler then places each task onto a suitable device by employing a policy appropriate to the system. In our prototype, a throughput-oriented scheduling policy is implemented to evaluate our resource-aware scheduling framework. The Rodinia benchmark suite and the Darknet neural network framework were used in our evaluation. The results show that, as compared to existing state-of-the-art methods, CASE improves throughput by up to 2.5X for Rodinia, and up to 2.7X for Darknet on modern NVIDIA GPU platforms, mainly due to the fact that it improves the average system utilization by up to 3.36X and the job turnaround time by up to 4.9X. Meanwhile, it limits individual kernel performance degradation within 2.5%. CASE achieved peak system utilization of 78% for Rodinia and 80% for Darknet on a 4XV100 system. Chao Chen 0024, Chris Porter, Santosh Pande |
PPoPP | 1 |
| 2022 | Near-Zero Downtime Recovery From Transient-Error-Induced CrashesabstractDue to the system scaling,transient errorscaused by external noise, e.g., heat fluxes and particle strikes, have become a growing concern for the current and upcoming exa-scale high-performance-computing (HPC) systems. Applications running on these systems are expected to experience transient errors more frequently than ever before, which will either lead them to generate incorrect outputs or cause them to crash. However, since such errors are still quite rare as compared to no-fault cases, desirable solutions call for low/no-overhead systems that do not compromise the performance under no-fault conditions and also allow very fast fault recovery to minimize downtime. In this article, we presentIterPro, a light-weight compiler-assisted resilience technique to quickly and accurately recover processes from transient-error-induced crashes. During the compilation of applications,IterProconstructs a set of recovery kernels for crash-prone instructions. These recovery kernels are executed to repair the corrupted process states on-the-fly upon occurrences of errors, enabling applications to continue their executions instead of being terminated. When constructing recovery kernels,IterProexploits side effects introduced by induction variable based code optimization techniques based on loop unrolling and strength reduction to improve its recovery capability. To this end, two new code transformation passes are introduced to expose the side effects for resilience purposes. We evaluatedIterProwith 4 scientific workloads as well as the NPB benchmarks suite. During their normal execution,IterProincurs almostzeroruntime overhead and a small, fixed27MBmemory overhead. Meanwhile,IterProcan recover on an average 83.55 percent of crash-causing errors within dozens of milliseconds with negligible downtime. We also evaluatedIterProwith parallel jobs running on 3072 cores and showed thatIterProcan successfully mask the impact of crash-causing errors by providing almost uninterrupted execution. Finally, we present our preliminary evaluation result for BLAS, which shows thatIterProis capable of recovering failures in libraries with a very high coverage rate of 83 percent and negligible overheads. With such an effective recovery mechanism,IterProcould tremendously mitigate the overheads and resource requirements of the resilience subsystem in future exa-scale systems. Chao Chen 0024, Greg Eisenhauer, Santosh Pande |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | CARE: compiler-assisted recovery from soft failuresabstractAs processors continue to boost the system performance with higher circuit density, shrinking process technology and near-threshold voltage (NTV) operations, they are projected to be more vulnerable to transient faults, which have become one of the major concerns for future extreme-scale HPC systems. Despite being relatively infrequent, crashes due to transient faults are incredibly disruptive, particularly for massively parallel jobs on supercomputers where they potentially kill the entire job, requiring an expensive rerun or restart from a checkpoint. Chao Chen 0024, Greg Eisenhauer, Santosh Pande, Qiang Guan |
SC | 1 |
| 2018 | LADR: low-cost application-level detector for reducing silent output corruptionsabstractApplications running on future high performance computing (HPC) systems are more likely to experience transient faults due to technology scaling trends with respect to higher circuit density, smaller transistor size and near-threshold voltage (NTV) operations. A transient fault could corrupt application state without warning, possibly leading to incorrect application output. Such errors are called silent data corruptions (SDCs). Chao Chen 0024, Greg Eisenhauer, Matthew Wolf, Santosh Pande |
HPDC | 1 |