Haochen Pan

dblp:255/7549 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
11since 2021 · last 2025
0009-0006-8992-5895ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 Dynostore: A Wide-Area Distribution System for the Management of Data Over Heterogeneous Storage
abstract
Data distribution across different facilities offers benefits such as enhanced resource utilization, increased resilience through replication, and improved performance by processing data near its source. However, managing such data is challenging due to heterogeneous access protocols, disparate authentication models, and the lack of a unified coordination framework. This paper presents DynoStore, a system that manages data across heterogeneous storage systems. At the core of DynoStore are data containers, an abstraction that provides standardized interfaces for seamless data management, irrespective of the underlying storage systems. Multiple data container connections create a cohesive wide-area storage network, ensuring resilience using erasure coding policies. Furthermore, a load-balancing algorithm ensures equitable and efficient utilization of storage resources. We evaluate DynoStore using benchmarks and realworld case studies, including the management of medical and satellite data across geographically distributed environments. Our results demonstrate a 10 % performance improvement compared to centralized cloud-hosted systems while maintaining competitive performance with state-of-the-art solutions such as Redis and IPFS. DynoStore also exhibits superior fault tolerance, withstanding more failures than traditional systems.
Dante D. Sánchez-Gallegos, José Luis González 0002, Maxime Gonthier, Valérie Hayot-Sasson, J. Gregory Pauloski, Haochen Pan, Kyle Chard, Jesús Carretero 0001, Ian T. Foster
CCGrid6
2025 Wrath: Workload Resilience Across Task Hierarchies in Task-Based Parallel Programming Frameworks
abstract
Failures in Task-based Parallel Programming (TBPP) can severely degrade performance and result in incomplete or incorrect outcomes. Existing failure-handling approaches, including reactive, proactive, and resilient methods such as retry and checkpointing mechanisms, often apply uniform retry mechanisms regardless of the root cause of failures, failing to account for the unique characteristics of TBPP frameworks such as heterogeneous resource availability and task-level failures. To address these limitations, we propose Wrath, a novel systematic approach that categorizes failures based on the unique layered structure of TBPP frameworks and defines specific responses to address failures at different layers. Wrath combines a distributed monitoring system and a resilient module to collaboratively address different types of failures in real time. The monitoring system captures execution and resource information, reports failures, and profiles tasks across different layers of TBPP frameworks. The resilient module then categorizes failures and responds with appropriate actions, such as hierarchically retrying failed tasks on suitable resources. Evaluations demonstrate that Wrath significantly improves TBPP robustness, tripling the task success rate and maintaining an application success rate of over 90 % for resolvable failures. Additionally, Wrath can reduce the time to failure by$20 \%-50 \%$, allowing tasks that are destined to fail to be identified and fail more quickly.
Zhuozhao Li, Valérie Hayot-Sasson, Haochen Pan, Maxime Gonthier, J. Gregory Pauloski, Ryan Chard, Kyle Chard, Ian T. Foster
CCGrid4
2025 D-Rex: Heterogeneity-Aware Reliability Framework and Adaptive Algorithms for Distributed Storage
abstract
The exponential growth of data necessitates distributed storage models, such as peer-to-peer systems and data federations.While distributed storage can reduce costs and increase reliability, the heterogeneity in storage capacity, I/O performance, and failure rates of storage resources makes their efficient use a challenge.Further, node failures are common and can lead to data unavailability and even data loss.
Maxime Gonthier, Dante D. Sánchez-Gallegos, Haochen Pan, Bogdan Nicolae, Hai Nguyen 0005, Valérie Hayot-Sasson, J. Gregory Pauloski, Jesús Carretero 0001, Kyle Chard, Ian T. Foster
ICS3
2025 Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
abstract
Achieving efficient task parallelism on many-core architectures is an important challenge. The widely used GNU OpenMP implementation of the popular OpenMP parallel programming model incurs high overhead for fine-grained, shortrunning tasks due to time spent on runtime synchronization. In this work, we introduce and analyze three key advances that collectively achieve significant performance gains. First, we introduce XQueue, a lock-less concurrent queue implementation to replace GNU's priority task queue and remove the global task lock. Second, we develop a scalable, efficient, and hybrid lock-free/lock-less distributed tree barrier to address the high hardware synchronization overhead from GNU's centralized barrier. Third, we develop two lock-less and NUMA-aware load balancing strategies. We evaluate our implementation using Barcelona OpenMP Task Suite (BOTS) benchmarks. We show that the use of XQueue and the distributed tree barrier can improve performance by up to$1522.8 \times$compared to the original GNU OpenMP. We further show that lock-less load balancing can improve performance by up to$4 \times$compared to GNU OpenMP using XQueue.
Maxime Gonthier, Poornima Nookala, Haochen Pan, Ian T. Foster, Ioan Raicu, Kyle Chard
IPDPS4
2024 An Empirical Investigation of Container Building Strategies and Warm Times to Reduce Cold Starts in Scientific Computing Serverless Functions
abstract
Serverless computing has revolutionized application development and deployment by abstracting infrastructure management, allowing developers to focus on writing code. To do so, serverless platforms dynamically create execution environments, often using containers. The cost to create and deploy these environments is known as "cold start" latency, and this cost can be particularly detrimental to scientific computing workloads characterized by sporadic and dynamic demands. We investigate methods to mitigate cold start issues in scientific computing applications by pre-installing Python packages in container images. Using data from Globus Compute and Binder, we empirically analyze cold start behavior and evaluate four strategies for building containers, including fully pre-built environments and dynamic, on-demand installations. Our results show that pre-installing all packages reduces initial cold start time but requires significant storage. Conversely, dynamic installation offers lower storage requirements but incurs repetitive delays. Additionally, we implemented a simulator and assessed the impact of different warm times, finding that moderate warm times significantly reduce cold starts without the excessive overhead of maintaining always-hot states.
André Bauer 0001, Maxime Gonthier, Haochen Pan, Ryan Chard, Daniel Grzenda, Martin Sträßer, J. Gregory Pauloski, Alok Kamatar, Matt Baughman, Nathaniel Hudson 0001, Ian T. Foster, Kyle Chard
e-Science3
2024 Diaspora: Resilience-Enabling Services for Real-Time Distributed Workflows
abstract
The need for real-time processing to enable automated decision making and experimental steering has driven a shift from high-performance computing workflows on a centralized system to a distributed approach that integrates remote data sources, edge devices, and diverse compute facilities. Under this paradigm, data can be processed close to the source where it is generated, thus reducing latency and bandwidth usage. System resilience is thus a key challenge, requiring distributed workflows to survive component failures and to meet stringent quality-of-service requirements, which results in the need to mitigate anomalies such as congestion and low availability of resources. To address these challenges, we propose Diaspora, a unified resilience framework that is inspired by event-driven communication patterns used in public clouds. Specifically, we propose an event fabric that extends across sites, facilities, and computations to provide timely, reliable, and accurate information about data, application, and resource status. On top of the event fabric, we build resilience-enabling services that combine QoS-aware data streaming, resilient data views, resilient compute and data resources, and anomaly detection and prediction, all of which collectively enhance workflow resilience for these scientific cases.
Bogdan Nicolae, Justin M. Wozniak, Tekin Bicer, Hai Nguyen 0005, Haochen Pan, Amal Gueroudji, Maxime Gonthier, Valérie Hayot-Sasson, Eliu A. Huerta, Kyle Chard, Ryan Chard, Matthieu Dorier, Nageswara S. V. Rao, Anees Al-Najjar, Alessandra Corsi, Ian T. Foster
e-Science6
2024 TaPS: A Performance Evaluation Suite for Task-based Execution Frameworks
abstract
Task-based execution frameworks, such as parallel programming libraries, computational workflow systems, and function-as-a-service platforms, enable the composition of distinct tasks into a single, unified application designed to achieve a computational goal and abstract the parallel and distributed execution of those tasks on arbitrary hardware. Research into these task executors has accelerated as computational sciences increasingly need to take advantage of parallel compute and/or heterogeneous hardware. However, the lack of evaluation standards makes it challenging to compare and contrast novel systems against existing implementations. Here, we introduce TaPS, the Task Performance Suite, to support continued research in distributed task executor frameworks. TaPS provides (1) a unified, modular interface for writing and evaluating applications using arbitrary execution frameworks and data management systems and (2) an initial set of reference synthetic and real-world science applications. We discuss how the design of TaPS supports the reliable evaluation of frameworks and demonstrate TaPS through a survey of benchmarks using the provided reference applications.
J. Gregory Pauloski, Valérie Hayot-Sasson, Maxime Gonthier, Nathaniel Hudson 0001, Haochen Pan, Ian T. Foster, Kyle Chard
e-Science5
2024 The globus compute dataset: An open function-as-a-service dataset from the edge to the cloud
André Bauer 0001, Haochen Pan, Ryan Chard, Yadu N. Babuji, Josh Bryan, Devesh Tiwari, Ian T. Foster, Kyle Chard
Future Gener. Comput. Syst.2
2022 Reliable Broadcast in Critical Applications: Asset Transfer and Smart Home
abstract
Asynchronous Byzantine reliable broadcast receives renewed attention recently, as it is fundamental to many fault-tolerant critical applications. This paper focuses on the Byzantine Reliable Broadcast protocol, which was first proposed by Bracha in 1987. Several recent protocols have improved the round and bit complexity of these algorithms. Motivated by practical network constraints in modern applications, this paper revisits the problem and reduces both complexity in communication and local computation. State-of-the-arts protocols are evaluated using the developed framework that simulates realistic bandwidth constraints. The evaluation demonstrates that our protocols, which use cryptographic hash functions and erasure coding in a novel way, have superior performance in critical applications such as asset transfer and smart home.
Yingjian Wu, Yicheng Shen, Haochen Pan, Lewis Tseng, Moayad Aloqaily
ICC3
2022 Cancellation in Systems: An Empirical Study of Task Cancellation Patterns and Failures
Utsav Sethi, Haochen Pan, Shan Lu 0001, Madan Musuvathi, Suman Nath
OSDI2
2021 Rabia: Simplifying State-Machine Replication Through Randomization
abstract
We introduce Rabia, a simple and high performance framework for implementing state-machine replication (SMR) within a datacenter. The main innovation of Rabia is in using randomization to simplify the design. Rabia provides the following two features: (i) It does not need any fail-over protocol and supports trivial auxiliary protocols like log compaction, snapshotting, and reconfiguration, components that are often considered the most challenging when developing SMR systems; and (ii) It provides high performance, up to 1.5x higher throughput than the closest competitor (i.e., EPaxos) in a favorable setup (same availability zone with three replicas) and is comparable with a larger number of replicas or when deployed in multiple availability zones.
Haochen Pan, Jesse Tuglu, Neo Zhou, Yicheng Shen, Xiong Zheng, Joseph Tassarotti, Lewis Tseng, Roberto Palmieri
SOSP1
2020 BBB: A Lightweight Approach to Evaluate Private Blockchains in Clouds
abstract
Evaluating Blockchain performance is not an easy task. It is difficult to compare different systems, since the evaluation is often incomprehensible and conducted in different environments with distinct workloads. Only a handful of prior tools were proposed, e.g., BLOCKBENCH and HFBench. Unfortunately, these tools have several limitations. We first identify these limitations. Second, motivated by our observations, we then present a benchmarking tool, Boston Blockchain Benchmarking (BBB). BBB is configurable, extensible, and easy-touse. In particular, BBB can be used to test Blockchain from a networking perspective, a feature that we have not observed in prior tools. Similar to BLOCKBENCH, we focus on the private Blockchain. Concretely, we integrate our tool with Mininet, and provide a simple mechanism to test how network properties (e.g., latency, bandwidth, package loss rate) affect the performance of the chosen Blockchain. We present our preliminary result of evaluating Ethereum. We stress that the architecture of BBB is general, and could be extended to other Blockchain systems. BBB is extremely lightweight and can be used on your laptop to test a small network. Such a feature allows quick evaluation of the Blockchain and speeds up innovation and development.
Haochen Pan, Xuheng Duan, Yingjian Wu, Lewis Tseng, Moayad Aloqaily, Azzedine Boukerche
GLOBECOM1
2020 CassandrEAS: Highly Available and Storage-Efficient Distributed Key-Value Store with Erasure Coding
abstract
In this work, we propose an erasure coding-based protocol that implements a key-value store with atomicity and near-optimal storage cost. Our protocol supports concurrent read and write operations while tolerating asynchronous communication and crash failures of any client and some fraction of servers. One novel feature is a tunable knob between the number of supported concurrent operations, availability, and storage cost. We implement our protocol into Cassandra, namely Cassan-drEAS (Cassandra + Erasure-coding Atomic Storage). Extensive evaluation using YCSB on Google Cloud Platform shows that CassandrEAS incurs moderate penalty on latency and throughput, yet saves significant amount of storage space.
Viveck R. Cadambe, Kishori M. Konwar, Muriel Médard, Haochen Pan, Lewis Tseng, Yingjian Wu
NCA4
2020 Reliable broadcast with trusted nodes: Energy reduction, resilience, and speed
Lewis Tseng, Yingjian Wu, Haochen Pan, Moayad Aloqaily, Azzedine Boukerche
Comput. Networks3
2019 Reliable Broadcast in Networks with Trusted Nodes
abstract
Broadcast is one of the fundamental primitives to enable large-scale networks such as sensor networks and IoT. There is a rich study on achieving reliable broadcast under various kind of failures. In this paper, we use the notion of trust to improve the performance of reliable broadcast. We focus on Certified Propagation Algorithm (CPA), one of the simple algorithms that does not rely on a cryptographic infrastructure and has a proven guarantee on resilience (number of node failures tolerated). Specifically, the paper has two main contributions: (i) A new algorithm Trust-CPA which integrates CPA with trusted nodes has been proposed and shown to increase the resilience from the original CPA, and (ii) A natural optimization problem related to Trust-CPA (i.e., finding the location to place trusted nodes to reduce the broadcast latency) has been proposed as well. We first show that it is NP-hard to find an exact answer and even NP-hard to find a good approximation. A greedy heuristic algorithm has been used and its efficacy has been examined using simulation. We show that our algorithm performs relatively well in geometric random graphs, an appropriate model for large- scale wireless sensor networks.
Lewis Tseng, Yingjian Wu, Haochen Pan, Moayad Aloqaily, Azzedine Boukerche
GLOBECOM3
2019 Distributed Causal Memory in the Presence of Byzantine Servers
abstract
We study distributed causal shared memory (or distributed read/write objects) in the client-server model over asynchronous message-passing networks in which some servers may suffer Byzantine failures. Since Ahamad et al. proposed causal memory in 1994, there have been abundant research on causal storage. Lately, there is a renewed interest in enforcing causal consistency in large-scale distributed storage systems (e.g., COPS, Eiger, Bolt-on). However, to the best of our knowledge, the fault-tolerance aspect of causal memory is not well studied, especially on the tight resilience bound. In our prior work, we showed that 2 f+1 servers is the tight bound to emulate crash-tolerant causal shared memory when up to f servers may crash. In this paper, we adopt a typical model considered in many prior works on Byzantine-tolerant storage algorithms and quorum systems. In the system, up to f servers may suffer Byzantine failures and any number of clients may crash. We constructively present an emulation algorithm for Byzantine causal memory using 3 f+1 servers. We also prove that 3 f+1 is necessary for tolerating up to f Byzantine servers. In other words, we show that 3 f+1 is a tight bound. For evaluation, we implement our algorithm in Golang and compare their performance with two state-of-the-art fault-tolerant algorithms that ensure atomicity in the Google Cloud Platform.
Lewis Tseng, Zezhi Wang, Haochen Pan
NCA4
2019 BBB: Make Benchmarking Blockchains Configurable and Extensible
abstract
Interest in Blockchain technology gradually developed after Bitcoin was introduced in 2008, and has grown exponentially as Blockchain can be applied in many scenarios that require coordination among parties that do not typically trust each other. Due to its popularity and wide applications, enormous number of Blockchain systems have been proposed in the past few years. One way to understand the performance is to build a benchmarking tool to evaluate different systems under different scenarios. A few research teams have build their benchmarking tools, such as BLOCKBENCH and HFBench. Unfortunately, there are several limitations of these tools. In this paper, we enumerate these limitations and challenges. We then present our tool - Boston Blockchain Benchmarking (BBB), which is more configurable and extensible than prior benchmarking tools. Moreover, BBB allows us to evaluate the impact of some common attacks against Blockchains. Finally, we present some preliminary results using BBB.
Xuheng Duan, Haochen Pan, Lewis Tseng, Yingjian Wu
PRDC2