EDBT 2026 Demo / reviewers in the wild / expert
Kevin Assogba
dblp:218/7041
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-0377-4576ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On Optimizing Checkpoint Restoration for HPC Applications: Leveraging Merkle Trees and Asynchronous I/OabstractEfficient checkpoint restoration is critical in high-performance computing (HPC) and AI applications, where slow recovery times disrupt workflows, waste resources, and hinder reproducibility. This work introduces a Merkle tree checkpoint restoration method to accelerate failure recovery and improve explainability. Our method integrates asynchronous I/O via the Liburing library to optimize scattered reads in HPC applications. Tested on the Polaris system at Argonne National Laboratory, it exhibits lower restoration time and memory consumption than state-of-the-art checkpoint restoration methods, reaching near-full efficiency with duplicated data. Our work advances scalable and efficient checkpointing solutions for HPC, ensuring reliable and fast failure recovery for large-scale simulations. Zackary Malkmus, Nigel Tan, Ian Lumsden, Kevin Assogba, M. Mustafa Rafique, Bogdan Nicolae, Michela Taufer |
HPDC | 4 |
| 2025 | Bayelemabaga: Creating Resources for Bambara NLPabstractAllahsera Auguste Tapo, Kevin Assogba, Christopher M Homan, M. Mustafa Rafique, Marcos Zampieri. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Allahsera Tapo, Kevin Assogba, Christopher Homan, M. Mustafa Rafique, Marcos Zampieri |
NAACL (Long Papers) | 2 |
| 2024 | Towards Affordable Reproducibility Using Scalable Capture and Comparison of Intermediate Multi-Run ResultsabstractEnsuring reproducibility in high-performance computing (HPC) applications is a significant challenge, particularly when nondeterministic execution can lead to untrustworthy results. Traditional methods that compare final results from multiple runs often fail because they provide sources of discrepancies only a posteriori and require substantial resources, making them impractical and unfeasible. This paper introduces an innovative method to address this issue by using scalable capture and comparing intermediate multi-run results. By capitalizing on intermediate checkpoints and hash-based techniques with user-defined error bounds, our method identifies divergences early in the execution paths. We employ Merkle trees for checkpoint data to reduce the I/O overhead associated with loading historical data. Our evaluations on the nondeterministic HACC cosmology simulation show that our method effectively captures differences above a predefined error bound and significantly reduces I/O overhead. Our solution provides a robust and scalable method for improving reproducibility, ensuring that scientific applications on HPC systems yield trustworthy and reliable results. Nigel Tan, Kevin Assogba, Walter J. Ashworth, Befikir Bogale, Franck Cappello, M. Mustafa Rafique, Michela Taufer, Bogdan Nicolae |
Middleware | 2 |
| 2023 | PredictDDL: Reusable Workload Performance Prediction for Distributed Deep LearningabstractAccurately predicting the training time of deep learning (DL) workloads is critical for optimizing the utilization of data centers and allocating the required cluster resources for completing critical model training tasks before a deadline. The state-of-the-art prediction models, e.g., Ernest and Cherrypick, treat DL workloads as black boxes, and require running the given DL job on a fraction of the dataset. Moreover, they require retraining their prediction models every time a change occurs in the given DL workload. This significantly limits the reusability of prediction models across DL workloads with different deep neural network (DNN) architectures. In this paper, we address this challenge and propose a novel approach where the prediction model is trained only once for a particular dataset type, e.g., ImageNet, thus completely avoiding tedious and costly retraining tasks for predicting the training time of new DL workloads. Our proposed approach, called PredictDDL, provides an end-to-end system for predicting the training time of DL models in distributed settings. PredictDDL leverages Graph HyperNetworks, a class of neural networks that takes computational graphs as input and produces vector representations of their DNNs. PredictDDL is the first prediction system that eliminates the need of retraining a performance prediction model for each new DL workload and maximizes the reuse of the prediction model by requiring running a DL workload only once for training the prediction model. Our extensive evaluation using representative workloads shows that PredictDDL achieves up to 9.8× lower average prediction error and 10.3× lower inference time compared to the state-of-the-art system, i.e., Ernest, on multiple DNN architectures. Kevin Assogba, Eduardo Lima, M. Mustafa Rafique, Minseok Kwon |
CLUSTER | 1 |
| 2023 | Building the I (Interoperability) of FAIR for Performance Reproducibility of Large-Scale Composable Workflows in RECUPabstractScientific computing communities increasingly run their experiments using complex data- and compute-intensive workflows that utilize distributed and heterogeneous architectures targeting numerical simulations and machine learning, often executed on the Department of Energy Leadership Computing Facilities (LCFs). We argue that a principled, systematic approach to implementing FAIR principles at scale, including fine-grained metadata extraction and organization, can help with the numerous challenges to performance reproducibility posed by such workflows. We extract workflow patterns, propose a set of tools to manage the entire life cycle of performance metadata, and aggregate them in an HPC-ready framework for reproducibility (RECUP). We describe the challenges in making these tools interoperable, preliminary work, and lessons learned from this experiment. Bogdan Nicolae, Tanzima Z. Islam, Robert B. Ross, Huub J. J. Van Dam, Kevin Assogba, Polina Shpilker, Mikhail Titov, Matteo Turilli, Tianle Wang 0001, Ozgur O. Kilic, Shantenu Jha, Line C. Pouchard |
e-Science | 5 |
| 2023 | Optimizing the Training of Co-Located Deep Learning Models Using Cache-Aware StaggeringabstractDespite significant advances, training deep learning models remains a time-consuming and resource-intensive task. One of the key challenges in this context is the ingestion of the training data, which involves non-trivial overheads: read the training data from a remote repository, apply augmentations and transformations, shuffle the training samples, and assemble them into mini-batches. Despite the introduction of abstractions such as data pipelines that aim to hide such overheads asynchronously, it is often the case that the data ingestion is slower than the training, causing a delay at each training iteration. This problem is further augmented when training multiple deep learning models simultaneously on powerful compute nodes that feature multiple GPUs. In this case, the training data is often reused across different training instances (e.g., in the case of multi-model or ensemble training) or even within the same training instance (e.g., data-parallel training). However, transparent caching solutions (e.g., OS-level POSIX caching) are not suitable to directly mitigate the competition between training instances that reuse the same training data. In this paper, we study the problem of how to minimize the makespan of running two training instances that reuse the same training data. The makespan is subject to a trade-off: if the training instances start at the same time, competition for I/O bandwidth slows down the data pipelines and increases the makespan. If one training instance is staggered, competition is reduced but the makespan increases. We aim to optimize this trade-off by proposing a performance model capable of predicting the makespan based on the staggering between the training instances, which can be used to find the optimal staggering that triggers just enough competition to make optimal use of transparent caching in order to minimize the makespan. Experiments with different combinations of learning models using the same training data demonstrate that (1) staggering is important to minimize the makespan; (2) our performance model is accurate and can predict the optimal staggering in advance based on calibration overhead. Kevin Assogba, Bogdan Nicolae, M. Mustafa Rafique |
HiPC | 1 |
| 2022 | On Realizing Efficient Deep Learning Using Serverless ComputingabstractServerless computing is gaining rapid popularity as it enables quick application deployment and seamless application scaling without managing complex computing resources. Re-cently, it has been explored for running data-intensive, e.g., deep learning (DL), workloads for improving application performance and reducing execution cost. However, serverless computing imposes resource-level constraints, specifically fixed memory allocation and short task timeouts, that lead to job failures. In this paper, we address these constraints and develop an effective runtime framework, DiSDeL, that improves the performance of DL jobs by leveraging data splitting techniques, and ensuring that an appropriate amount of memory is allocated to containers for storing application data and a suitable timeout is selected for each job based on its complexity in serverless deployments. We implement our approach using Apache OpenWhisk and TensorFlow platforms and evaluate it using representative DL workloads to show that it eliminates DL job failures and reduces action memory consumption and total training time by up to 44% and 46%, respectively as compared to a default serverless computing framework. Our evaluation also shows that DiSDeL achieves a performance improvement of up to 29% as compared to bare-metal TensorFlow environment in a multi-tenant setting. Kevin Assogba, Moiz Arif, M. Mustafa Rafique, Dimitrios S. Nikolopoulos |
CCGRID | 1 |
| 2022 | Exploiting CXL-based Memory for Distributed Deep LearningabstractDeep learning (DL) is being widely used to solve complex problems in scientific applications from diverse domains, such as weather forecasting, medical diagnostics, and fluid dynamics simulation. DL applications consume a large amount of data using large-scale high-performance computing (HPC) systems to train a given model. These workloads have large memory and storage requirements that typically go beyond the limited amount of main memory available on an HPC server. This significantly increases the overall training time as the input training data and model parameters are frequently swapped to slower storage tiers during the training process. In this paper, we use the latest advancements in the memory subsystem, specifically Compute Express Link (CXL), to provide additional memory and fast scratch space for DL workloads to reduce the overall training time while enabling DL jobs to efficiently train models using data that is much larger than the installed system memory. We propose a framework, called DeepMemoryDL, that manages the allocation of additional CXL-based memory, introduces a fast intermediate storage tier, and provides intelligent prefetching and caching mechanisms for DL workloads. We implement and integrate DeepMemoryDL with a popular DL platform, TensorFlow, to show that our approach reduces read and write latencies, improves the overall I/O throughput, and reduces the training time. Our evaluation shows a performance improvement of up to 34% and 27% compared to the default TensorFlow platform and CXL-based memory expansion approaches, respectively. Moiz Arif, Kevin Assogba, M. Mustafa Rafique, Sudharshan S. Vazhkudai |
ICPP | 2 |
| 2022 | Canary: Fault-Tolerant FaaS for Stateful Time-Sensitive ApplicationsabstractFunction-as-a-Service (FaaS) platforms have recently gained rapid popularity. Many stateful applications have been migrated to FaaS platforms due to their ease of deployment, scalability, and minimal management overhead. However, failures in FaaS have not been thoroughly investigated, thus making these desirable platforms unreliable for guaranteeing function execution and ensuring performance requirements. In this paper, we propose Canary, a highly resilient and fault-tolerant framework for FaaS that mitigates the impact of failures and reduces the overhead of function restart. Canary utilizes replicated container runtimes and application-level checkpoints to reduce application recovery time over FaaS platforms. Our evaluations using representative stateful FaaS applications show that Canary reduces the application recovery time and dollar cost by up to 83% and 12%, respectively over the default retry-based strategy. Moreover, it improves application availability with an additional average execution time and cost overhead of 14% and 8%, respectively, as compared to the ideal failure-free execution. Moiz Arif, Kevin Assogba, M. Mustafa Rafique |
SC | 2 |
| 2018 | Two-echelon location-routing optimization with time windows based on customer clustering
Yong Wang 0022, Kevin Assogba, Yong Liu 0028, Xiaolei Ma, Maozeng Xu, Yinhai Wang |
Expert Syst. Appl. | 2 |
| 2018 | Two-echelon logistics delivery and pickup network optimization based on integrated cooperation and transportation fleet sharing
Yong Wang 0022, Shouguo Peng, Chengcheng Xu 0001, Kevin Assogba, Haizhong Wang, Maozeng Xu, Yinhai Wang |
Expert Syst. Appl. | 4 |
| 2018 | Collaboration and transportation resource sharing in multiple centers vehicle routing optimization with delivery and pickup
Yong Wang 0022, Kevin Assogba, Yong Liu 0028, Maozeng Xu, Yinhai Wang |
Knowl. Based Syst. | 3 |