EDBT 2026 Demo / reviewers in the wild / expert
Seydou Ba
dblp:155/6599
· DBLP profile ↗
6ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0001-8837-3116ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Machine Learning Driven Auto-Tuning for Non-Uniform All-to-All CollectivesabstractNon-uniform all-to-all communication patterns present optimization challenges in parallel computing due to their irregular data distribution and dynamic behavior. While MPI_Alltoallv provides the standard interface for such exchanges, achieving optimal performance requires careful selection among multiple implementation variants and tuning of algorithm-specific parameters. This paper presents a data-driven autotuning framework that combines machine learning-based runtime prediction with a lookup-table mechanism for fast configuration selection. The ML model estimates the communication time of each algorithm configuration under a given system setup, allowing the framework to identify the optimal implementation and parameter set based on predicted performance. We validate our approach through comprehensive benchmarking of MPI_Alltoallv and two specialized algorithms across varying process counts, message sizes, and tunable parameters. Applied to a real MPI-based transitive closure application on the Fugaku supercomputer, our framework achieves up to$6.03 \times$reduction in communication time over the vendor implementation, providing a detailed understanding of non-uniform collective communication behavior and a practical framework for automatic performance optimization in HPC applications. Kunting Qi, Jens Domke, Seydou Ba, Venkatram Vishwanath, Michael E. Papka, Sidharth Kumar |
HiPC | 4 |
| 2025 | Parameterized Algorithms for Non-uniform All-to-allabstractMPI_Alltoallv generalizes the uniform all-to-all communication (MPI_Alltoall) by enabling the exchange of data-blocks of varied sizes among processes. This function plays a crucial role in facilitating many computational tasks, such as FFT calculations and graph mining operations. Popular MPI libraries, such as MPICH and OpenMPI, implement MPI_Alltoall using a combination of linear and logarithmic algorithms. However, MPI_Alltoallv typically relies only on variations of linear algorithms, missing the benefits of logarithmic approaches. Furthermore, current algorithms also overlook the intricacies of modern HPC system architectures, such as the significant performance gap between intra-node (local) and inter-node (global) communication. To address these problems, this paper presents two novel algorithms: Parameterized Logarithmic non-uniform All-to-all (ParLogNa) and Parameterized Linear nonuniform All-to-all (ParLinNa). ParLogNa is a tunable logarithmic time algorithm for non-uniform all-to-all, and ParLinNa is a hierarchical and tunable near-linear-time algorithm for non-uniform all-to-all. These algorithms efficiently address the trade-off between bandwidth maximization and latency minimization that existing implementations struggle to optimize. We show a performance improvement over the state-of-the-art implementations by factors of 42x and 138x on Polaris and Fugaku, respectively. Jens Domke, Seydou Ba, Sidharth Kumar |
HPDC | 3 |
| 2025 | Bine Trees: Enhancing Collective Operations by Optimizing Communication LocalityabstractCommunication locality plays a key role in the performance of collective operations on large HPC systems, especially on oversubscribed networks where groups of nodes are fully connected internally but sparsely linked through global connections. We present Bine (binomial negabinary) trees, a family of collective algorithms that improve communication locality. Bine trees maintain the generality of binomial trees and butterflies while cutting global-link traffic by up to \(33\%\). We implement eight Bine-based collectives and evaluate them on four large-scale supercomputers with Dragonfly, Dragonfly+, oversubscribed fat-tree, and torus topologies, achieving up to 5 × speedups and consistent reductions in global-link traffic across different vector sizes and node counts. Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato, Seydou Ba, Matteo Turisini, Jens Domke, Torsten Hoefler |
SC | 5 |
| 2019 | Monte Carlo Tree Search with Variable Simulation Periods for Continuously Running TasksabstractMonte Carlo Tree Search (MCTS) is widely used for planning in domains where the potential actions can be represented as a tree of sequential decisions. To efficiently select an action, MCTS usually needs to perform many simulations to build a reliable tree representation of the decision space. As such, a bottleneck to MCTS arises when enough simulations cannot be performed between action selections. This is particularly highlighted in continuously running tasks, for which the time available to perform simulations between actions tends to be limited due to the environment's state constantly changing. In this paper, we present an approach that extends the time available for Monte Carlo simulations when allowed. Our approach is to effectively balance the prospect of selecting the right action with the time that can be spared to perform MCTS simulations before the next action selection. For that, we considered the simulation time as a decision variable to be selected alongside an action. We extended the Hierarchical Optimistic Optimization applied to Tree (HOOT) method to adapt our approach to environments with a continuous decision space. We evaluated our approach on tasks with a continuous decision space using OpenAI gym's Pendulum and Continuous Mountain Car environments and on those with discrete action space using the arcade learning environment (ALE) platform. The evaluation results show that, with variable simulation times, the proposed approach outperforms the conventional MCTS in the evaluated continuous decision space tasks and improves the performance of MCTS in most of the ALE tasks. Seydou Ba, Takuya Hiraoka, Takashi Onishi, Toru Nakata, Yoshimasa Tsuruoka |
ICTAI | 1 |
| 2017 | Defragmentation Scheme Based on Exchanging Primary and Backup Paths in 1+1 Path Protected Elastic Optical NetworksabstractIn elastic optical networks (EONs), a major obstacle to using the spectrum resources efficiently is the spectrum fragmentation. In the literature, several defragmentation approaches have been presented. For 1+1 path protection, conventional defragmentation approaches consider designated primary and backup paths. This exposes the spectrum to fragmentations induced by the primary lightpaths, which are not to be disturbed in order to achieve hitless defragmentation. This paper proposes a defragmentation scheme using path exchanging in 1+1 path protected EONs. We exchange the path function of the 1+1 protection with the primary toggling to the backup state, while the backup becomes the primary. This allows both lightpaths to be reallocated during the defragmentation process, while they work as backup, offering hitless defragmentation. Considering path exchanging, we define a static spectrum reallocation optimization problem that minimizes the spectrum fragmentation while limiting the number of path exchanging and reallocation operations. We then formulate the problem as an integer linear programming (ILP) problem. We prove that a decision version of the defined static reallocation problem is NP-complete. We present a spectrum defragmentation process for dynamic traffic, and introduce a heuristic algorithm for the case that the ILP problem is not tractable. The simulation results show that the proposed scheme outperforms the conventional one and improves the total admissible traffic up to 10%. Seydou Ba, Bijoy Chand Chatterjee, Eiji Oki |
IEEE/ACM Trans. Netw. | 1 |
| 2016 | Computational time complexity of allocation problem for distributed servers in real-time applicationsabstractThis paper analyzes the computational time complexity of the allocation problem for data processing functions among multiple users and distributed servers in the distributed processing communication scheme for a real-time network application. In the distributed processing communication scheme, the application is processed on a data processing function in the distributed servers in order to minimize the delay time. We prove that the allocation problem for data processing functions among multiple users and distributed servers is an NP-complete problem. Seydou Ba, Akio Kawabata, Bijoy Chand Chatterjee, Eiji Oki |
APNOMS | 1 |