VLDB 2026 Research / reviewers in the wild / expert
Mohamed Wahib
dblp:10/6150
· DBLP profile ↗
71ranked-venue papers
10as first author
49since 2021 · last 2026
0000-0002-7165-2095ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 49 · 6 first-author · 35 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FRUGAL: Pushing GPU Applications beyond Memory LimitsabstractGPUs power modern scientific and AI applications, but their limited memory capacity restricts scalability. Buying GPUs with larger HBM is prohibitively expensive and still bounded by market limits. Existing solutions either exploit application-specific knowledge through out-of-core techniques, which lack generality, or rely on system-level page faulting, which is transparent but inefficient. We propose FRUGAL, an application-agnostic framework and methodology that reduces GPU memory footprint while sustaining high performance. FRUGAL formulates memory management as an optimization over an application’s execution graph, encompassing prefetching, kernel execution, and offloading. Using static analysis and profiling, FRUGAL applies a two-phase scheduling and migration strategy, solving an otherwise intractable optimization efficiently. Evaluations on Tiled Cholesky Decomposition, Tiled LU Decomposition, Tiny-CUDA-NN, and QuEST show that FRUGAL significantly reduces maximum GPU memory usage by 80.21%, 80.20%, 64.75% and 60.86% with only a geometric mean of 28.31% slowdown. FRUGAL allows applications to exceed hardware-imposed limits, and maintains strong performance scalability beyond existing GPU memory constraints, without additional hardware cost. Lingqi Zhang 0001, Jiajun Huang 0001, Chen Zhuang, Ivan R. Ivanov, Peng Chen 0035, Toshio Endo, Mohamed Wahib |
CGO | 8 |
| 2026 | EvoTADASHI: Genetic Programming for High-Performance Code Optimization
João E. Batista, Emil Vatai, Aleksandr Drozd, Mohamed Wahib |
EvoApplications (1) | 4 |
| 2026 | SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix MultiplicationabstractDistributed Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in high-performance computing and deep learning applications. The major performance bottleneck in distributed SpMM lies in substantial communication overhead, which limits both performance and scalability. In this paper, we identify two key sources of communication inefficiency in distributed SpMM: redundant data transfer due to sparsity unawareness, and suboptimal utilization of hierarchical network topology. To address these, we propose (1) a fine-grained, sparsity-aware communication strategy that reduces communication overhead by exploiting the sparsity pattern of the sparse matrix, and (2) a hierarchical communication strategy that maps the sparsity-aware strategy onto two-tier GPU network architectures, minimizing redundant data movement across slower inter-node links. We implement these optimizations in a comprehensive distributed SpMM framework, SHIRO. Extensive evaluations on real-world datasets show that SHIRO demonstrates strong scalability up to 128 GPUs, achieving geometric mean speedups of 221.5 ×, 56.0 ×, 23.4 ×, and 8.8 × in SpMM over four state-of-the-art baselines (CAGNET, SPA, BCL, and CoLa, respectively) at this scale. Chen Zhuang, Lingqi Zhang 0001, Benjamin Brock, Du Wu, Peng Chen 0035, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
ICS | 8 |
| 2026 | Looking for (Genomic) Needles in a Haystack: Sparsity-Driven Search for Identifying Correlated Genetic Mutations in Cancer
Ritvik Prabhu, Emil Vatai, Bernard Moussad, Emmanuel Jeannot, Ramu Anandakrishnan, Wu-chun Feng, Mohamed Wahib |
IPDPS | 7 |
| 2026 | Neural architecture search for generative adversarial networks with hybrid convolution
Yu Xue 0003, Yufeng Zou, Mohamed Wahib, Peng Chen 0035, Moncef Gabbouj |
Neurocomputing | 3 |
| 2026 | RT-RkNN: Reverse k Nearest Neighbor Queries as a Graphics Ray Casting Problem
Zhengyang Bai, Peng Chen 0035, Mohamed Wahib |
Proc. VLDB Endow. | 3 |
| 2026 | A Pairwise Comparison Relation-Assisted Multiobjective Evolutionary Neural Architecture Search Method With Multipopulation MechanismabstractNeural architecture search (NAS) has emerged as a powerful paradigm that enables researchers to automatically explore vast search spaces and discover efficient neural networks. However, NAS suffers from a critical bottleneck, i.e., the evaluation of numerous architectures during the search process demands substantial computing resources and time. In order to improve the efficiency of NAS, a series of methods have been proposed to reduce the evaluation time of neural architectures. However, they are not efficient enough and still only focus on the accuracy of architectures. Beyond classification accuracy, real-world applications increasingly demand more efficient and compact network architectures that balance multiple performance criteria. To address these challenges, we propose the SMEMNAS, a pairwise comparison relation-assisted multiobjective evolutionary algorithm (EA) based on a multipopulation (MP) mechanism. In the SMEMNAS, a surrogate model is constructed based on pairwise comparison relations to predict the accuracy ranking of architectures, rather than the absolute accuracy. Moreover, two populations cooperate with each other in the search process, i.e., a main population that guides the evolutionary process and a vice population that enhances search diversity. Our method aims to discover high-performance models that simultaneously optimize multiple objectives. We conduct comprehensive experiments on CIFAR-10, CIFAR-100, and ImageNet datasets to validate the effectiveness of our approach. With only a single GPU searching for 0.17 days, competitive architectures can be found by SMEMNAS, which achieves 78.91% accuracy with the MAdds of 570 M on the ImageNet. This work makes a significant advancement in the field of NAS. Yu Xue 0003, Pengcheng Jiang, Chenchen Zhu, MengChu Zhou, Mohamed Wahib, Moncef Gabbouj |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2025 | A Device-Side Execution Model for Multi-GPU Task GraphsabstractExecuting task graphs on multi-GPU systems presents challenges typically managed by CPU-side runtimes, which handle memory management, track dependencies, and balance load.However, the interplay of runtime components, CPUdriven kernel initialization, and dynamic task graph construction creates significant overhead.For static graphs, recent advancements have enabled GPU-side execution, demonstrating substantial performance gains in single-GPU scenarios.However, multi-GPU execution still lags behind in both usability and performance.In particular, no GPU-side solution exists for executing task graphs on multiple nodes.In this work, we introduce Mustard, a multi-GPU execution model that shifts execution of static task graphs entirely to the devices, drastically reducing overhead.Mustard offers a clean solution for executing CUDA graphs across multiple GPUs on multiple nodes without requiring modifications to GPU kernel code or the adoption of new runtime mechanisms or APIs.By transforming the task graph, Mustard enables precise tracking of task dependencies and load balancing directly on the GPU, eliminating the need for host CPU involvement.We evaluate our approach using generated graphs, as well as LU and Cholesky decomposition graphs.In a multi-node scenario with 64 GPUs, Mustard achieves an average 5.83× speedup over the linear algebra library SLATE.On a single node, compared to the best-performing baseline, Mustard delivers an average 1.66× speedup for LU and 1.29× for Cholesky. Ilyas Turimbetov, Mohamed Wahib, Didem Unat |
ICS | 2 |
| 2025 | Scaling Large-scale GNN Training to Thousands of Processors on CPU-based SupercomputersabstractGraph Convolutional Networks (GCNs), particularly for largescale graphs, are crucial across numerous domains.However, training distributed full-batch GCNs on large-scale graphs suffers from inefficient memory access patterns and high communication overhead.To address these challenges, we introduce SuperGCN, an efficient and scalable distributed GCN Chen Zhuang, Lingqi Zhang 0001, Du Wu, Peng Chen 0035, Jiajun Huang 0001, Xin Liu 0020, Rio Yokota, Nikoli Dryden, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
ICS | 11 |
| 2025 | SHF: Symmetrical Hierarchical Forest with Pretrained Vision Transformer Encoder for High-Resolution Medical SegmentationabstractThis paper presents a novel approach to addressing the long-sequence problem in high-resolution medical images for Vision Transformers (ViTs). Using smaller patches as tokens can enhance ViT performance, but quadratically increases computation and memory requirements. Therefore, the common practice for applying ViTs to high-resolution images is either to: (a) employ complex sub-quadratic attention schemes or (b) use large to medium-sized patches and rely on additional mechanisms within the model to capture the spatial hierarchy of details. We propose Symmetrical Hierarchical Forest (SHF), a lightweight approach that adaptively patches the input image to increase token information density and encode hierarchical spatial structures into the input embedding. We then apply a reverse depatching scheme to the output embeddings of the transformer encoder, eliminating the need for convolution-based decoders. Unlike previous methods that modify attention mechanisms \wahib{or use a complex hierarchy of interacting models}, SHF can be retrofitted to any ViT model to allow it to learn the hierarchical structure of details in high-resolution images without requiring architectural changes. Experimental results demonstrate significant gains in computational efficiency and performance: on the PAIP WSI dataset, we achieved a 3$\sim$32$\times$ speedup or a 2.95\% to 7.03\% increase in accuracy (measured by Dice score) at a $64K^2$ resolution with the same computational budget, compared to state-of-the-art production models. On the 3D medical datasets BTCV and KiTS, training was 6$\times$ faster, with accuracy gains of 6.93\% and 5.9\%, respectively, compared to models without SHF. Enzhi Zhang, Peng Chen 0035, Rui Zhong 0004, Du Wu, Jun Igarashi, Isaac Lyngaas, Xiao Wang 0004, Masaharu Munetomo, Mohamed Wahib |
NeurIPS | 9 |
| 2025 | A General and Scalable GCN Training Framework on CPU SupercomputersabstractGraph Convolutional Networks (GCNs) are widely used in various domains. However, training distributed full-batch GCNs on large-scale graphs poses challenges due to inefficient memory access patterns and high communication overhead. This paper presents a general and efficient GCN training framework on CPU supercomputers. It comprises a general aggregation kernel designed to optimize irregular memory access and a quantization method with label propagation to reduce communication overhead. Experimental results show that our method achieves a speedup of up to 4.1× compared with the SoTA implementations. Chen Zhuang, Peng Chen 0035, Xin Liu 0020, Rio Yokota, Nikoli Dryden, Lingqi Zhang 0001, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
PPoPP | 9 |
| 2025 | A Sample-Free Compilation Framework for Efficient Dynamic Tensor ComputationabstractDynamic-shape tensor computation poses challenges for shape-specific compilation due to variable input dimensions. Existing compilers rely on shape samples, incurring high tuning costs and performance degradation on unseen inputs. We present Helix, a dynamic tensor compilation framework with sample-free compilation and architecture-guided optimization to achieve both compilation efficiency and shape-general performance. To avoid shape sampling, Helix constructs shape-agnostic compilation by decomposing computations across architectural layers. A bidirectional strategy combines top-down abstraction to align tensor computations with architectural hierarchies, and bottom-up kernel construction to build efficient execution strategies from reusable, architecture-aligned micro-kernels. A hybrid analyzer ensures accuracy through profiling at lower architectural levels, and achieves scalability through architecture-informed modeling at higher levels and runtime. This hierarchical design eliminates shape-specific tuning and enables shape-adaptive execution. Evaluations conducted on x86 CPUs, ARM CPUs, and NVIDIA GPUs demonstrate that Helix reduces compilation time by 174 × over the existing compilers and delivers 2.26 × and 3.29 × execution speedups over vendor libraries and dynamic-shape compilers, respectively. Yangjie Zhou 0001, Weihao Cui, Zihan Liu 0002, Peng Chen 0035, Mohamed Wahib, Cong Guo 0003, Siyuan Feng 0007, Jintao Meng 0001, Haidong Lan, Jingwen Leng, Yun Lin 0001, Jin Song Dong 0001, Wenxi Zhu, Minwen Deng |
SC | 7 |
| 2025 | ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate DownscalingabstractSparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R2 scores in range of 0.98–0.99 against observation data. Xiao Wang 0004, Jong-Youl Choi, Takuya Kurihana, Isaac Lyngaas, Hong-Jun Yoon, Xi Xiao 0003, David Pugmire, Nasik Muhammad Nafi, Aristeidis Tsaris, Ashwin M. Aji, Maliha Hossain, Mohamed Wahib, Dali Wang, Peter E. Thornton, Prasanna Balaprakash, Moetasim Ashfaq, Dan Lu 0001 |
SC | 13 |
| 2025 | RAPTOR: Practical Numerical Profiling of Scientific ApplicationsabstractThe proliferation of low-precision units in modern high-performance architectures increasingly burdens domain scientists. Historically, the choice in HPC was easy: can we get away with 32 bit floating-point operations and lower bandwidth requirements, or is FP64 necessary? Driven by Artificial Intelligence, vendors introduce novel low-precision units for vector and tensor operations, and FP64 capabilities stagnate or are reduced. This forces scientists to re-evaluate their codes, but a trivial search-and-replace approach to go from FP64 to FP16 will not suffice. Faveo Hoerold, Ivan R. Ivanov, Akash Dhruv, William S. Moses, Anshu Dubey, Mohamed Wahib, Jens Domke |
SC | 6 |
| 2025 | Distributed Cross-Channel Hierarchical Aggregation for Foundation ModelsabstractVision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources—such as varying physical groundings or data acquisition systems—and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer. Aristeidis Tsaris, Isaac Lyngaas, John H. Lagergren, Mohamed Wahib, Larry M. York, Prasanna Balaprakash, Dan Lu 0001, Feiyi Wang, Xiao Wang 0004 |
SC | 4 |
| 2025 | Balanced and Elastic End-to-end Training of Dynamic LLMsabstractTo reduce the computational and memory overhead of Large Language Models, various approaches have been proposed. These include a) Mixture of Experts (MoEs), where token routing affects compute balance; b) gradual pruning of model parameters; c) dynamically freezing layers; d) dynamic sparse attention mechanisms; e) early exit of tokens as they pass through model layers; and f) Mixture of Depths (MoDs), where tokens bypass certain blocks. While these approaches are effective in reducing overall computation, they often introduce significant workload imbalance across workers. In many cases, this imbalance is severe enough to render the techniques impractical for large-scale distributed training, limiting their applicability to toy models due to poor efficiency. Mohamed Wahib, Muhammed Abdullah Soyturk, Didem Unat |
SC | 1 |
| 2025 | Noisy data-based attack: A new type of untargeted attack in Federated Learning and its countermeasures
Manh Cuong Dao, Phi-Le Nguyen, Hieu H. Pham 0001, Thanh-Hung Nguyen, Peng Chen 0035, Mohamed Wahib, Truong Thao Nguyen |
Future Gener. Comput. Syst. | 6 |
| 2025 | CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations
Johann Rudi, Youngjun Lee, Aidan H. Chadha, Mohamed Wahib, Klaus Weide, Jared O'Neal, Anshu Dubey |
Future Gener. Comput. Syst. | 4 |
| 2025 | Predictor-assisted evolutionary neural architecture search for spiking neural networks
Yu Xue 0003, Yiyu Tan, Peng Chen 0035, Mohamed Wahib |
Neurocomputing | 6 |
| 2025 | Enabling Visual Scene Recovery From Wi-Fi CSI for Occlusion-Free SurveillanceabstractWe introduce CSI-Inpainter, a novel approach for obstacle removal using Wi-Fi CSI. This method harnesses CSI data to reconstruct obscured visual elements, regardless of lighting conditions. Extensive empirical evaluation in both office and industrial settings demonstrates the effectiveness of CSI-Inpainter’s exceptional ability to identify and reconstruct occluded segments, outperforming traditional baselines and our received signal strength indicator (RSSI)-based work, RF-Inpainter in terms of visual quality. Our findings emphasize the superiority of CSI data over RSSI for providing richer visual information and underscore the critical role of optimal sensor placement and data fusion from multiple CSI sensors in enhancing the performance. CSI-Inpainter represents a significant advancement in obstacle removal for various applications like surveillance, offering new insights into the integration of wireless sensing and visual scene recovery, expanding the potential applications of Computer Vision in real-world environments. Cheng Chen 0068, Shoki Ohta, Takayuki Nishio, Mehdi Bennis, Jihong Park, Mohamed Wahib |
IEEE Internet Things J. | 6 |
| 2025 | YOLO-DKR: Differentiable architecture search based on kernel reusing for object detection
Yu Xue 0003, Chenhang Yao, Mohamed Wahib, Moncef Gabbouj |
Inf. Sci. | 3 |
| 2025 | Neural Architecture Search with Progressive Evaluation and Subpopulation PreservationabstractNeural architecture search (NAS) is an effective approach for automating the design of deep neural networks. Evolutionary computation (EC) is commonly used in NAS due to its global optimization capability. However, the evaluation phase of architecture candidates in EC-based NAS is compute-intensive, limiting its application for many real-world problems. To overcome this challenge, we propose a novel progressive evaluation strategy for the evaluation phase in convolutional neural network architecture search, in which the number of training epochs of network individuals is progressively increased. In addition, a subpopulation preservation strategy is proposed to preserve medium-size and large-size architectures to avoid prematurely discarding networks that may not perform well in the early stages but have the potential to excel with further optimization. Our proposed algorithm reduces the computational cost of the evaluation phase and promotes population diversity and fairness by preserving promising networks based on their distribution. We evaluate the proposed progressive evaluation and subpopulation preservation of NAS (PEPNAS) algorithm on the CIFAR10, CIFAR100, and ImageNet benchmark datasets, and compare it with 36 state-of-the-art algorithms, including manually designed networks, reinforcement learning (RL) algorithms, gradient-based algorithms, and other EC-based ones. The experimental results demonstrate that PEPNAS effectively identifies networks with competitive accuracy while also markedly improving the efficiency of the search process. For instance, PEPNAS discovers the architecture on CIFAR10 with a low-error rate of 2.38% using only 0.7 GPU days. We directly adopt the searched architecture for the image classification on the CIFAR100 and ImageNet datasets, which achieves the top 1 error rates of 16.46% and 26.25%, respectively. The code is available athttps://github.com/chajiajie/PEPNAS. Yu Xue 0003, Jiajie Zha, Danilo Pelusi, Peng Chen 0035, Tao Luo 0014, Liangli Zhen, Yan Wang 0015, Mohamed Wahib |
IEEE Trans. Evol. Comput. | 8 |
| 2025 | Vision transformer-based meta loss landscape exploration with actor-critic method
Enzhi Zhang, Rui Zhong 0004, Xingbang Du, Mohamed Wahib, Masaharu Munetomo |
J. Supercomput. | 4 |
| 2024 | Welcome Message from the IEEE Cluster 2024 Program ChairsabstractWe are thrilled to share the program for this year's IEEE Cluster conference, showcasing a diverse range of topics in cluster computing and emphasizing the field's ongoing significance. The program strikes a balance between established subjects like architectures, software environments, and scientific applications, and emerging areas such as data analytics and deep learning. Yutong Lu, Wu-chun Feng, Mohamed Wahib |
CLUSTER | 3 |
| 2024 | Surrogate-Assisted Evolutionary Neural Architecture Search with Isomorphic Training and Prediction
Pengcheng Jiang, Yu Xue 0003, Ferrante Neri, Mohamed Wahib |
ICIC (2) | 4 |
| 2024 | Real-time High-resolution X-Ray Computed TomographyabstractComputed Tomography (CT) serves as a key imaging technology that relies on computationally intensive filtering and back-projection algorithms for 3D image reconstruction. While conventional high-resolution image reconstruction (> 2K3) solutions provide quick results, they typically treat reconstruction as an offline workload to be performed remotely on large-scale HPC systems. The growing demand for post-construction AI-driven analytics and the need for real-time adjustments call for high-resolution reconstruction solutions that are feasible on local computing resources, i.e. a multi-GPU server at most. In this paper, we propose a novel approach that utilizes Tensor Cores to optimize image reconstruction without sacrificing precision. We also introduce a framework designed to enable real-time execution of end-to-end distributed image reconstruction in a multi-GPU environment. Evaluations conducted on a single Nvidia A100 and H100 GPU show performance improvements of 1.91 × and 2.15 × compared to highly optimized production libraries. Furthermore, our framework, when deployed on 8-card Nvidia A100 GPU system, demonstrates the ability to reconstruct real-world datasets into 20483 volumes (32 GB) in slightly more than one minute and 40963 volumes (256 GB) in 7 minutes. Du Wu, Peng Chen 0035, Xiao Wang 0004, Isaac Lyngaas, Takaaki Miyajima, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
ICS | 8 |
| 2024 | Progressive Neural Predictor with Score-Based SamplingabstractNeural architecture search (NAS) automates the design of neural networks, but faces high computational costs for evaluating the performance candidate architectures. Surrogate-assisted NAS methods use approximate computational models to get predictive estimation instead of real complete training, but also face the challenge of maintaining the balance between training cost and predictive effectiveness. In this paper, we propose a progressive neural predictor that uses score-based sampling (PNSS) to improve the performance of the surrogate model with limited training data. Different from existing algorithms that rely on initial sample selection, PNSS uses an online method to progressively select new samples of the surrogate model based on potential information from the previous search process. During the iterative process, the sampled scores are dynamically adjusted based on the prediction rankings in each round to keep track of good architectures, which gradually optimises the surrogate model. In this way, the processes of training the predictor and searching for architectures are jointly combined to improve the efficiency of sample utilization. In addition, the surrogate model with different degrees of training is assigned prediction confidence equal to the accuracy of the current stage. Experiments are conducted on NAS-Bench-101 and NAS-Bench-201 benchmarks. The experimental results show that the proposed PNSS algorithm outperforms the existing methods with limited training samples. In addition, visualisation of the search process and ablation study also shows the effectiveness of the progressive search. Yu Xue 0003, Ferrante Neri, Xiaoping Zhao, Mohamed Wahib |
IJCNN | 5 |
| 2024 | Validation Loss Landscape Exploration with Deep Q-LearningabstractOverfitting is a well-documented and studied issue in supervised learning. Human experts have been designing methods to reduce over-fitting by observing the validation knowledge, e.g., learning rate schedules, dropout, and adversarial training. We propose a validation-loss landscape exploration/exploitation method called VKI (Validation Knowledge Inheritance). We reformulate the traditional gradient optimization problem as a reinforcement learning task to explore the validation-loss landscape. In particular, by treating the gradient descent process as a Markov Decision Process (MDP), where the validation losses are treated as the costs, we train a Q-network as a controller to learn the future rewards and later use it to decrease the validation loss of workers. We conduct the experiments in two reduced gradient action spaces on the MNIST, CIFAR-10, and CIFAR-100 datasets with naive dense neural networks and ResNet-56. Our results show that VKI could rediscover hyperparameters schedule rules and improve the models’ training and generalization by only exploring/exploiting the validation loss landscape, e.g., increasing the learning rate accelerated training and penalty factor when overfitting. In comparison to other metaHPO (Hyper Parameter Optimization) methods, we empirically show that VKI can leverage its weights-loss regression of Q-Net to enable the transfer to the target dataset without heavy retraining but with light fine-tuning. Enzhi Zhang, Rui Zhong 0004, Masaharu Munetomo, Mohamed Wahib |
IJCNN | 4 |
| 2024 | autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm ArchitecturesabstractThis paper presents an open-source library that pushes the limits of performance portability for irregular General Matrix Multiplication (GEMM) on the widely-used Arm architectures. Our library, autoGEMM, is designed to support a wide range of Arm processors: from edge devices to HPCgrade CPUs. autoGEMM generates optimized kernels for various hardware configurations by auto-combining fragments of autogenerated micro-kernels that employ hand-written optimizations to maximize computational efficiency. We optimize the kernel pipeline by tuning the register reuse and the data load/store overlapping. In addition, we use a dynamic tiling scheme to generate balanced tile shapes. Finally, we position autoGEMM on top of the TVM framework where our dynamic tiling scheme prunes the search space for TVM to identify the optimal combination of parameters for code optimization. Evaluations on five different classes of Arm chips demonstrate the advantages of autoGEMM. For small matrices, autoGEMM achieves 98% of peak and up to 2.0x speedup over state-of-the-art libraries such as LIBXSMM and LibShalom. For irregular matrices (i.e. tall skinny and long rectangles), autoGEMM is 1.3-2.0x faster than widely-used libraries such as OpenBLAS and Eigen. autoGEMM is publicly available at: https://github.com/wudu98/autoGEMM. Du Wu, Jintao Meng 0001, Wenxi Zhu, Minwen Deng, Xiao Wang 0004, Tao Luo 0014, Mohamed Wahib, Yanjie Wei |
SC | 7 |
| 2024 | Adaptive Patching for High-resolution Image Segmentation with TransformersabstractAttention-based models are proliferating in the space of image analytics, including segmentation. The standard method of feeding images to transformer encoders is to divide the images into patches and then feed the patches to the model as a linear sequence of tokens. For high-resolution images, e.g. microscopic pathology images, the quadratic compute and memory cost prohibits the use of an attention-based model, if we are to use smaller patch sizes that are favorable in segmentation. The solution is to either use custom complex multi-resolution models or approximate attention schemes. We take inspiration from Adapative Mesh Refinement (AMR) methods in HPC by adaptively patching the images, as a pre-processing step, based on the image details to reduce the number of patches being fed to the model, by orders of magnitude. This method has a negligible overhead, and works seamlessly with any attention-based model, i.e. it is a pre-processing step that can be adopted by any attention-based model without friction. We demonstrate superior segmentation quality over SoTA segmentation models for real-world pathology datasets while gaining a geomean speedup of $6.9 \times$ for resolutions up to $64 K^{2}$, on up to 2,048 GPUs. Enzhi Zhang, Isaac Lyngaas, Peng Chen 0035, Xiao Wang 0004, Jun Igarashi, Yuankai Huo, Masaharu Munetomo, Mohamed Wahib |
SC | 8 |
| 2024 | Meta generative image and text data augmentation optimization
Enzhi Zhang, Bochen Dong, Mohamed Wahib, Rui Zhong 0004, Masaharu Munetomo |
J. Supercomput. | 3 |
| 2023 | Multi-GPU Communication Schemes for Iterative Solvers: When CPUs are Not in ChargeabstractThis paper proposes a fully autonomous execution model for multi-GPU applications that completely excludes the involvement of the CPU beyond the initial kernel launch. In a typical multi-GPU application, the host serves as the orchestrator of execution by directly launching kernels, issuing communication calls, and acting as a synchronizer for devices. We argue that this orchestration, or control flow path, causes undue overhead and can be delegated entirely to devices to improve performance in applications that require communication among peers. For the proposed CPU-free execution model, we leverage existing techniques such as persistent kernels, thread block specialization, device-side barriers, and device-initiated communication routines to write fully autonomous multi-GPU code and achieve significantly reduced communication overheads. We demonstrate our proposed model on two broadly used iterative solvers, 2D/3D Jacobi stencil and Conjugate Gradient(CG). Compared to the CPU-controlled baselines, the CPU-free model can improve 3D stencil communication latency by 58.8% and provide a 1.63x speedup for CG on 8 NVIDIA A100 GPUs. The project code is available at https://github.com/ParCoreLab/CPU-Free-model. Ismayil Ismayilov, Javid Baydamirli, Dogan Sagbili, Mohamed Wahib, Didem Unat |
ICS | 4 |
| 2023 | PERKS: a Locality-Optimized Execution Model for Iterative Memory-bound GPU ApplicationsabstractIterative memory-bound solvers commonly occur in HPC codes. Typical GPU implementations have a loop on the host side that invokes the GPU kernel as much as time/algorithm steps there are. The termination of each kernel implicitly acts the barrier required after advancing the solution every time step. We propose an execution model for running memory-bound iterative GPU kernels: PERsistent KernelS (PERKS). In this model, the time loop is moved inside persistent kernel, and device-wide barriers are used for synchronization. We then reduce the traffic to device memory by caching subset of the output in each time step in the unused registers and shared memory. PERKS can be generalized to any iterative solver: they largely independent of the solver's implementation. We explain the design principle of PERKS and demonstrate effectiveness of PERKS for a wide range of iterative 2D/3D stencil benchmarks (geomean speedup of 2.12x for 2D stencils and 1.24x for 3D stencils over state-of-art libraries), and a Krylov subspace conjugate gradient solver (geomean speedup of 4.86x in smaller SpMV datasets from SuiteSparse and 1.43x in larger SpMV datasets over a state-of-art library). All PERKS-based implementations available at: https://github.com/neozhang307/PERKS. Lingqi Zhang 0001, Mohamed Wahib, Peng Chen 0035, Jintao Meng 0001, Xiao Wang 0004, Toshio Endo, Satoshi Matsuoka |
ICS | 2 |
| 2023 | Revisiting Temporal Blocking Stencil OptimizationsabstractIterative stencils are used widely across the spectrum of High Performance Computing (HPC) applications. Many efforts have been put into optimizing stencil GPU kernels, given the prevalence of GPU-accelerated supercomputers. To improve the data locality, temporal blocking is an optimization that combines a batch of time steps to process them together. Under the observation that GPUs are evolving to resemble CPUs in some aspects, we revisit temporal blocking optimizations for GPUs. We explore how temporal blocking schemes can be adapted to the new features in the recent Nvidia GPUs, including large scratchpad memory, hardware prefetching, and device-wide synchronization. We propose a novel temporal blocking method, EBISU, which champions low device occupancy to drive aggressive deep temporal blocking on large tiles that are executed tile-by-tile. We compare EBISU with state-of-the-art temporal blocking libraries: STENCILGEN and AN5D. We also compare with state-of-the-art stencil auto-tuning tools that are equipped with temporal blocking optimizations: ARTEMIS and DRSTENCIL. Over a wide range of stencil benchmarks, EBISU achieves speedups up to 2.53x and a geometric mean speedup of 1.49x over the best state-of-the-art performance in each stencil benchmark. Lingqi Zhang 0001, Mohamed Wahib, Peng Chen 0035, Jintao Meng 0001, Xiao Wang 0004, Toshio Endo, Satoshi Matsuoka |
ICS | 2 |
| 2023 | KAKURENBO: Adaptively Hiding Samples in Deep Neural Network TrainingabstractThis paper proposes a method for hiding the least-important samples during the training of deep neural networks to increase efficiency, i.e., to reduce the cost of training. Using information about the loss and prediction confidence during training, we adaptively find samples to exclude in a given epoch based on their contribution to the overall learning process, without significantly degrading accuracy. We explore the converge properties when accounting for the reduction in the number of SGD updates. Empirical results on various large-scale datasets and models used directly in image classification and segmentation show that while the with-replacement importance sampling algorithm performs poorly on large datasets, our method can reduce total training time by up to 22\% impacting accuracy only by 0.4\% compared to the baseline. Truong Thao Nguyen, Balazs Gerofi, Edgar Josafat Martinez-Noriega, François Trahay, Mohamed Wahib |
NeurIPS | 5 |
| 2023 | Training Knowledge Inheritance Through Deep Q-NetabstractWhen training neural networks, the weights of the model are updated at each optimization step, and the older weights are discarded. In this paper, we propose a method called, Training Knowledge Inheritance (TKI), to use the knowledge about the progression of weight and loss data in reducing overfitting and improving the generalization in the later stages of training. We reformulate the traditional gradient optimization problem as a reinforcement learning task. In particular, by treating the trainable weight space as an environment, the learning rate as action, and the validation accuracies as the rewards, we train a Q-network (controller) to learn the discounted future validation accuracy and guide the later training of another network (worker). We conduct the experiments on the MNIST, CIFAR-10, and CIFAR-100 datasets with naive dense neural networks and ResNet-56. Our results show that TKI could rediscover learning rate schedule rules similar to previous works, including increasing, decaying, and cyclical repeating. Enzhi Zhang, Ruqin Wang, Mohamed Wahib, Rui Zhong 0004, Masaharu Munetomo |
SMC | 3 |
| 2023 | At the Locus of Performance: Quantifying the Effects of Copious 3D-Stacked Cache on HPC WorkloadsabstractOver the last three decades, innovations in the memory subsystem were primarily targeted at overcoming the data movement bottleneck. In this paper, we focus on a specific market trend in memory technology: 3D-stacked memory and caches. We investigate the impact of extending the on-chip memory capabilities in future HPC-focused processors, particularly by 3D-stacked SRAM. First, we propose a method oblivious to the memory subsystem to gauge the upper-bound in performance improvements when data movement costs are eliminated. Then, using the gem5 simulator, we model two variants of a hypothetical LARge Cache processor (LARC), fabricated in 1.5 nm and enriched with high-capacity 3D-stacked cache. With a volume of experiments involving a broad set of proxy-applications and benchmarks, we aim to reveal how HPC CPU performance will evolve, and conclude an average boost of 9.56× for cache-sensitive HPC applications, on a per-chip basis. Additionally, we exhaustively document our methodological exploration to motivate HPC centers to drive their own technological agenda through enhanced co-design. Jens Domke, Emil Vatai, Balazs Gerofi, Yuetsu Kodama, Mohamed Wahib, Artur Podobas, Sparsh Mittal, Miquel Pericàs, Lingqi Zhang 0001, Peng Chen 0035, Aleksandr Drozd, Satoshi Matsuoka |
ACM Trans. Archit. Code Optim. | 5 |
| 2023 | Simeuro: A Hybrid CPU-GPU Parallel Simulator for Neuromorphic Computing ChipsabstractWith the success of deep learning, there have been numerous efforts to build hardware for it. One approach that is gaining momentum is neuromorphic computing with spiking neural networks (SNNs), which are multiplication-free and open the possibility of using analog computing via novel technologies. However, to design effective and efficient hardware for such architectures, a fast and accurate software simulator is key. This article presents Simeuro, a fast and scalable system-level simulator for SNN models used in neuromorphic accelerators. The simulator uses spike-level details and configurable architectural constraints that are independent of the underlying hardware implementation. Simeuro supports a wide range of features including analog computing, novel memory (currently, RRAM is supported), and a full network-on-chip. The simulator can provide detailed simulation results such as routing statistics, energy consumption, delay, and accuracy of arbitrarily defined SNN architectures. Our simulator leverages a CPU-GPU hybrid environment to expedite the simulation by scaling out to multi-nodes equipped with multi-GPUs. We are able to conduct core simulations for a system-scale SNN chip of 20,000 neuromorphic cores on up to 512 A100 GPUs in a few minutes. Huaipeng Zhang, Nhut-Minh Ho, Dogukan Yigit Polat, Peng Chen 0035, Mohamed Wahib, Truong Thao Nguyen, Jintao Meng 0001, Rick Siow Mong Goh, Satoshi Matsuoka, Tao Luo 0014, Weng-Fai Wong |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | Why Globally Re-shuffle? Revisiting Data Shuffling in Large Scale Deep LearningabstractStochastic gradient descent (SGD) is the most prevalent algorithm for training Deep Neural Networks (DNN). SGD iterates the input data set in each training epoch processing data samples in a random access fashion. Because this puts enormous pressure on the I/O subsystem, the most common approach to distributed SGD in HPC environments is to replicate the entire dataset to node local SSDs. However, due to rapidly growing data set sizes this approach has become increasingly infeasible. Surprisingly, the questions of why and to what extent random access is required have not received a lot of attention in the literature from an empirical standpoint. In this paper, we revisit data shuffling in DL workloads to investigate the viability of partitioning the dataset among workers and performing only a partial distributed exchange of samples in each training epoch. Through extensive experiments on up to 2,048 GPUs of ABCI and 4,096 compute nodes of Fugaku, we demonstrate that in practice validation accuracy of global shuffling can be maintained when carefully tuning the partial distributed exchange. We provide a solution implemented in PyTorch that enables users to control the proposed data exchange scheme. Truong Thao Nguyen, François Trahay, Jens Domke, Aleksandr Drozd, Emil Vatai, Jianwei Liao 0001, Mohamed Wahib, Balazs Gerofi |
IPDPS | 7 |
| 2022 | Image Gradient Decomposition for Parallel and Memory-Efficient Ptychographic ReconstructionabstractPtychography is a popular microscopic imaging modality for many scientific discoveries and sets the record for highest image resolution. Unfortunately, the high image resolution for ptychographic reconstruction requires significant amount of memory and computations, forcing many applications to compromise their image resolution in exchange for a smaller memory footprint and a shorter reconstruction time. In this paper, we propose a novel image gradient decomposition method that significantly reduces the memory footprint for ptychographic reconstruction by tessellating image gradients and diffraction measurements into tiles. In addition, we propose a parallel image gradient decomposition method that enables asynchronous point-to-point communications and parallel pipelining with minimal overhead on a large number of GPUs. Our experiments on a Titanate material dataset (PbTiO3) with 16632 probe locations show that our Gradient Decomposition algorithm reduces memory footprint by 51 times. In addition, it achieves time-to-solution within 2.2 minutes by scaling to 4158 GPUs with a super-linear strong scaling efficiency at 364% compared to runtimes at 6 GPUs. This performance is 2.7 times more memory efficient, 9 times more scalable and 86 times faster than the state-of-the-art algorithm. Xiao Wang 0004, Aristeidis Tsaris, Debangshu Mukherjee, Mohamed Wahib, Peng Chen 0035, Mark Oxley, Olga Ovchinnikova, Jacob D. Hinkle |
SC | 4 |
| 2022 | Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep LearningabstractWe presentFastConv, a template-based code auto-generation open-source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming layers of convolutional neural networks. ARM CPUs cover a wide range of designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. Results show 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, ARM NN, and FeatherCNN on Kunpeng 920. Furthermore, performance portability experiments with various convolution shapes show that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920. CPU performance portability evaluation on VGG–16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively. Jintao Meng 0001, Chen Zhuang, Peng Chen 0035, Mohamed Wahib, Bertil Schmidt, Xiao Wang 0004, Haidong Lan, Dou Wu, Minwen Deng, Yanjie Wei, Shengzhong Feng |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | An Allreduce Algorithm and Network Co-design for Large-Scale Training of Distributed Deep LearningabstractDistributed training of Deep Neural Networks (DNNs) on High-Performance Computing (HPC) systems is becoming increasingly common. HPC systems dedicated entirely or mainly to Deep Learning (DL) workloads are becoming a reality. The collective communication overhead for calculating the average of weight gradients, e.g., an Allreduce operations, is one of the main factors limiting the scaling of data parallelism training. Several active efforts across different layers of the training stack including the training algorithms, parallelism strategy, communication algorithms, and system design have been proposed to cope with this communication challenge when scaling distributed training of DNNs. However, even with those methods, communication still becomes a bottleneck with the steady increase in model sizes, e.g.,100-10,000s MB, and the number of compute nodes, e.g., 1,000-10,000s of GPUs. In this work, we investigate the benefits of co-design of Allreduce algorithms and the network system. We propose to replace the Fat-tree network topology with a variant of Distributed Loop Network topology that guarantees a fixed routing paths length between any pairs of computing nodes for the communication pattern of halving-doubling Allreduce algorithm. We also propose a technique to eliminate/mitigate the network contention. Truong Thao Nguyen, Mohamed Wahib |
CCGRID | 2 |
| 2021 | An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural NetworksabstractDeep Neural Network (DNN) frameworks use distributed training to enable faster time to convergence and alleviate memory capacity limitations when training large models and/or using high dimension inputs. With the steady increase in datasets and model sizes, model/hybrid parallelism is deemed to have an important role in the future of distributed training of DNNs. We analyze the compute, communication, and memory requirements of Convolutional Neural Networks (CNNs) to understand the trade-offs between different parallelism approaches on performance and scalability. We leverage our model-driven analysis to be the basis for an oracle utility which can help in detecting the limitations and bottlenecks of different parallelism approaches at scale. We evaluate the oracle on six parallelization strategies, with four CNN models and multiple datasets (2D and 3D), on up to 1024 GPUs. The results demonstrate that the oracle has an average accuracy of about 86.74% when compared to empirical results, and as high as 97.57% for data parallelism. Albert Kahira, Truong Thao Nguyen, Leonardo Arturo Bautista-Gomez, Ryousei Takano, Rosa M. Badia, Mohamed Wahib |
HPDC | 6 |
| 2021 | Intra-page Cache Update in SLC-mode with Partial Programming in High Density SSDsabstractModern high density SSDs commonly designate a part of their capacity as a cache using an Single-level Cell (SLC)-mode region. Partial programming is then adopted for reducing space fragmentation in the SLC-mode pages, but it exacerbates program disturb including both in-page disturb and neighbouring page disturb. This paper proposes a partial programming scheme (called intra-page update) by updating hot, small size data inside a given page to minimize the negative impact induced by program disturb. Moreover, we introduce a novel data movement principle to separate hot and cold write data in the SLC-mode cache when updating the data or carrying out garbage collection. As a result, the hot updated data can be kept in the SLC-mode cache and the cold data will be flushed onto the high density SSD region. Simulation tests on several realistic disk traces show that our proposal improves bit error rate by 9.2%, and I/O performance by 9.3% on average, compared to state-of-the-art methods, without a noticeable decrease in total endurance. Jun Li 0062, Minjun Li, Zhigang Cai, François Trahay, Mohamed Wahib, Balazs Gerofi, Zhiming Liu 0001, Min Huang 0018, Jianwei Liao 0001 |
ICPP | 5 |
| 2021 | Performance portable back-projection algorithms on CPUs: agnostic data locality and vectorization optimizationsabstractComputed Tomography (CT) is a key 3D imaging technology that fundamentally relies on the compute-intense back-projection operation to generate 3D volumes. GPUs are typically used for back-projection in production CT devices. However, with the rise of power-constrained micro-CT devices, and also the emergence of CPUs comparable in performance to GPUs, back-projection for CPUs could become favorable. Unlike GPUs, extracting parallelism for back-projection algorithms on CPUs is complex given that parallelism and locality are not explicitly defined and controlled by the programmer, as is the case when using CUDA for instance. We propose a collection of novel back-projection algorithms that reduce the arithmetic computation, robustly enable vectorization, enforce a regular memory access pattern, and maximize the data locality. We also implement the novel algorithms as efficient back-projection kernels that are performance portable over a wide range of CPUs. Performance evaluation using a variety of CPUs from different vendors and generations demonstrates that our back-projection implementation achieves on average 5.2 times speedup over the multi-threaded implementation of the most widely used, and optimized, open library. With a state‐of‐the‐art CPU, we reach performance that rivals top-performing GPUs. Peng Chen 0035, Mohamed Wahib, Xiao Wang 0004, Shin'ichiro Takizawa, Takahiro Hirofuchi, Hirotaka Ogawa, Satoshi Matsuoka |
ICS | 2 |
| 2021 | Matrix Engines for High Performance Computing: A Paragon of Performance or Grasping at Straws?abstractMatrix engines or units, in different forms and affinities, are becoming a reality in modern processors; CPUs and otherwise. The current and dominant algorithmic approach to Deep Learning merits the commercial investments in these units, and deduced from the No. 1 benchmark in supercomputing, namely High Performance Linpack, one would expect an awakened enthusiasm by the HPC community, too. Hence, our goal is to identify the practical added benefits for HPC and machine learning applications by having access to matrix engines. For this purpose, we perform an in-depth survey of software stacks, proxy applications and benchmarks, and historical batch job records. We provide a cost-benefit analysis of matrix engines, both asymptotically and in conjunction with state-of-the-art processors. While our empirical data will temper the enthusiasm, we also outline opportunities to “misuse” these dense matrix-multiplication engines if they come for free. Jens Domke, Emil Vatai, Aleksandr Drozd, Peng Chen 0035, Yosuke Oyama, Lingqi Zhang 0001, Shweta Salaria, Daichi Mukunoki, Artur Podobas, Mohamed Wahib, Satoshi Matsuoka |
IPDPS | 10 |
| 2021 | Scalable FBP decomposition for cone-beam CT reconstructionabstractFiltered Back-Projection (FBP) is a fundamental compute intense algorithm used in tomographic image reconstruction. Cone-Beam Computed Tomography (CBCT) devices use a cone-shaped X-ray beam, in comparison to the parallel beam used in older CT generations. Distributed image reconstruction of cone-beam datasets typically relies on dividing batches of images into different nodes. This simple input decomposition, however, introduces limits on input/output sizes and scalability. Peng Chen 0035, Mohamed Wahib, Xiao Wang 0004, Takahiro Hirofuchi, Hirotaka Ogawa, Ander Biguri, Richard P. Boardman, Thomas Blumensath, Satoshi Matsuoka |
SC | 2 |
| 2021 | Efficient MPI-AllReduce for large-scale deep learning on GPU-clustersabstractSummary Training models on large‐scale GPUs‐accelerated clusters are becoming a commonplace due to the increase in complexity and size in deep learning models. One of the main challenges for distributed training is the collective communication overhead for large message sizes: up to hundreds of MB. In this paper, we propose two hierarchical distributed memory multileader AllReduce algorithms optimized for GPU‐accelerated clusters (named lr_lr and lr_rab ), in which GPUs inside a computing node perform an intra‐node communication phase to gather and store results of local reduced values to designated GPUs (known as node leaders). Node leaders then keep a role as an inter‐node communicator. Each leader exchanges one part of reduced values to the leaders of the other nodes in parallel. Hence, we are capable of significantly reducing the time for injecting data into the inter‐node network. We also overlap the inter‐node and intra‐node communication by implementing our proposal in a pipelined manner. We evaluate those algorithms on the discrete‐event simulation Simgrid. We show that our algorithms, lr_lr and lr_rab , can cut down the execution time of an AllReduce microbenchmark that uses the logical ring algorithm ( lr ) by up to 45% and 51%, respectively. With the pipelined implementation, our lr_lr_pipe achieves 15% performance improvement when compared with lr_lr . In addition, the simulation result also projects power savings for the network devices of up to 23% and 32%. Truong Thao Nguyen, Mohamed Wahib, Ryousei Takano |
Concurr. Comput. Pract. Exp. | 2 |
| 2021 | A computational-graph partitioning method for training memory-constrained DNNs
Fareed Qararyah, Mohamed Wahib, Doga Dikbayir, Mehmet Esat Belviranli, Didem Unat |
Parallel Comput. | 2 |
| 2020 | AN5D: automated stencil framework for high-degree temporal blocking on GPUsabstractStencil computation is one of the most widely-used compute patterns in high performance computing applications. Spatial and temporal blocking have been proposed to overcome the memory-bound nature of this type of computation by moving memory pressure from external memory to on-chip memory on GPUs. However, correctly implementing those optimizations while considering the complexity of the architecture and memory hierarchy of GPUs to achieve high performance is difficult. We propose AN5D, an automated stencil framework which is capable of automatically transforming and optimizing stencil patterns in a given C source code, and generating corresponding CUDA code. Parameter tuning in our framework is guided by our performance model. Our novel optimization strategy reduces shared memory and register pressure in comparison to existing implementations, allowing performance scaling up to a temporal blocking degree of 10. We achieve the highest performance reported so far for all evaluated stencil benchmarks on the state-of-the-art Tesla V100 GPU. Kazuaki Matsumura, Hamid Reza Zohouri, Mohamed Wahib, Toshio Endo, Satoshi Matsuoka |
CGO | 3 |
| 2020 | A Study of Single and Multi-device Synchronization Methods in Nvidia GPUsabstractGPUs are playing an increasingly important role in general-purpose computing. Many algorithms require synchronizations at different levels of granularity in a single GPU. Additionally, the emergence of dense GPU nodes also calls for multi-GPU synchronization. Nvidia's latest CUDA provides a variety of synchronization methods. Until now, there is no full understanding of the characteristics of those synchronization methods. This work explores important undocumented features and provides an in-depth analysis of the performance considerations and pitfalls of the state-of-art synchronization methods for Nvidia GPUs. The provided analysis would be useful when making design choices for applications, libraries, and frameworks running on single and/or multi-GPU environments. We provide a case study of the commonly used reduction operator to illustrate how the knowledge gained in our analysis can be useful. We also describe our micro-benchmarks and measurement methods. Lingqi Zhang 0001, Mohamed Wahib, Satoshi Matsuoka |
IPDPS | 2 |
| 2020 | Scaling distributed deep learning workloads beyond the memory capacity with KARMAabstractThe dedicated memory of hardware accelerators can be insufficient to store all weights and/or intermediate states of large deep learning models. Although model parallelism is a viable approach to reduce the memory pressure issue, significant modification of the source code and considerations for algorithms are required. An alternative solution is to use out-of-core methods instead of, or in addition to, data parallelism. We propose a performance model based on the concurrency analysis of out-of-core training behavior, and derive a strategy that combines layer swapping and redundant recomputing. We achieve an average of 1. 52x speedup in six different models over the state-of-the-art out-of-core methods. We also introduce the first method to solve the challenging problem of out-of-core multi-node training by carefully pipelining gradient exchanges and performing the parameter updates on the host. Our data parallel out-of-core solution can outperform complex hybrid model parallelism in training large models, e.g. Megatron-LM and Turning-NLG. Mohamed Wahib, Truong Thao Nguyen, Aleksandr Drozd, Jens Domke, Lingqi Zhang 0001, Ryousei Takano, Satoshi Matsuoka |
SC | 1 |
| 2019 | Topology-aware Sparse Allreduce for Large-scale Deep LearningabstractData parallelism is the dominant method used to scale-up deep learning (DL) training across multiple compute nodes. Collective communication of the local gradients between nodes is a critical bottleneck due to the significant increase in complexity and size of DL models. Researchers cope with this problem by one of the following solutions: a) optimizing the collective communication algorithm to account for the underlying network topology (topology-aware), and b) reducing the amount of transferred data. In the latter approach, sparse communication techniques communicate only the essential data, which helps to significantly cut down the communication cost. However, the diversity of message sizes, and unknown overlap of indicessets among compute nodes (i.e. irregular size), restricts their communication into the many-to-many scheme, which is ineffective in large-scale implementations. In this paper, we present allreduce algorithms that can exploit both sparse-communication and topology-aware techniques by mixing heterogeneous data formats (dense and sparse) with a trivial cost of computation. Truong Thao Nguyen, Mohamed Wahib, Ryousei Takano |
IPCCC | 2 |
| 2019 | Double-Precision FPUs in High-Performance Computing: An Embarrassment of Riches?abstractAmong the (uncontended) common wisdom in High-Performance Computing (HPC) is the applications' need for large amount of double-precision support in hardware. Hardware manufacturers, the TOP500 list, and (rarely revisited) legacy software have without doubt followed and contributed to this view. In this paper, we challenge that wisdom, and we do so by exhaustively comparing a large number of HPC proxy applications on two processors: Intel's Knights Landing (KNL) and Knights Mill (KNM). Although similar, the KNL and KNM architecturally deviate at one important point: the silicon area devoted to double-precision arithmetics. This fortunate discrepancy allows us to empirically quantify the performance impact in reducing the amount of hardware double-precision arithmetic. Our analysis shows that this common wisdom might not always be right. We find that the investigated HPC proxy applications do allow for a (significant) reduction in double-precision with little-to-no performance implications. With the advent of a failing of Moore's law, our results partially reinforce the view taken by modern industry (e.g., upcoming Fujitsu ARM64FX) to integrate hybrid-precision hardware units. Jens Domke, Kazuaki Matsumura, Mohamed Wahib, Keita Yashima, Toshiki Tsuchikawa, Yohei Tsuji, Artur Podobas, Satoshi Matsuoka |
IPDPS | 3 |
| 2019 | A versatile software systolic execution model for GPU memory-bound kernelsabstractThis paper proposes a versatile high-performance execution model, inspired by systolic arrays, for memory-bound regular kernels running on CUDA-enabled GPUs. We formulate a systolic model that shifts partial sums by CUDA warp primitives for the computation. We also employ register files as a cache resource in order to operate the entire model efficiently. We demonstrate the effectiveness and versatility of the proposed model for a wide variety of stencil kernels that appear commonly in HPC, and also convolution kernels (increasingly important in deep learning workloads). Our algorithm outperforms the top reported state-of-the-art stencil implementations, including implementations with sophisticated temporal and spatial blocking techniques, on the two latest Nvidia architectures: Tesla V100 and P100. For 2D convolution of general filter sizes and shapes, our algorithm is on average 2.5× faster than Nvidia's NPP on V100 and P100 GPUs. Peng Chen 0035, Mohamed Wahib, Shin'ichiro Takizawa, Ryousei Takano, Satoshi Matsuoka |
SC | 2 |
| 2019 | iFDK: a scalable framework for instant high-resolution image reconstructionabstractComputed Tomography (CT) is a widely used technology that requires compute-intense algorithms for image reconstruction. We propose a novel back-projection algorithm that reduces the projection computation cost to 1/6 of the standard algorithm. We also propose an efficient implementation that takes advantage of the heterogeneity of GPU-accelerated systems by overlapping the filtering and back-projection stages on CPUs and GPUs, respectively. Finally, we propose a distributed framework for high-resolution image reconstruction on state-of-the-art GPU-accelerated supercomputers. The framework relies on an elaborate interleave of MPI collective communication steps to achieve scalable communication. Evaluation on a single Tesla V100 GPU demonstrates that our back-projection kernel performs up to 1.6× faster than the standard FDK implementation. We also demonstrate the scalability and instantaneous CT capability of the distributed framework by using up to 2,048 V100 GPUs to solve 4K and 8K problems within 30 seconds and 2 minutes, respectively (including I/O). Peng Chen 0035, Mohamed Wahib, Shin'ichiro Takizawa, Ryousei Takano, Satoshi Matsuoka |
SC | 2 |
| 2018 | Efficient Algorithms for the Summed Area Tables Primitive on GPUsabstractTwo-dimensional Summed Area Tables (SAT) is a fundamental primitive used in image processing and machine learning applications. We present a collection of optimization methods for computing SAT on CUDA-enabled GPUs. Conventional approaches rely on computing the prefix sum in one dimension in parallel, transposing the matrix, then computing the prefix sum for the other dimension in parallel. Additionally, conventional methods use the scratchpad memory as cache. We propose a collection of algorithms that are scalable with respect to problem size. We use the register cache technique instead of the scratchpad memory and also employ a naive serial scan on the thread level for computing the prefix sum for one of the dimensions. Using a novel transpose-in-registers method we increase the inter-thread parallelism and outperform conventional SAT implementations. In addition, we significantly reduce both the communication between threads and the number of arithmetic instructions. On an Nvidia Pascal P100 GPU and Volta V100, our evaluations demonstrate that our implementations outperform state of the art libraries and yield up to 2.3x and 3.2x speedup over OpenCV and Nvidia NPP libraries, respectively. Peng Chen 0035, Mohamed Wahib, Shin'ichiro Takizawa, Ryousei Takano, Satoshi Matsuoka |
CLUSTER | 2 |
| 2017 | Numerical Optimization of ESA's Messenger Space Mission Benchmark
Martin Schlueter, Mohamed Wahib, Masaharu Munetomo |
EvoApplications (1) | 2 |
| 2016 | Daino: a high-level framework for parallel and efficient AMR on GPUsabstractAdaptive Mesh Refinement methods reduce computational requirements of problems by increasing resolution for only areas of interest. However, in practice, efficient AMR implementations are difficult considering that the mesh hierarchy management must be optimized for the underlying hardware. Architecture complexity of GPUs can render efficient AMR to be particularity challenging in GPU-accelerated supercomputers. This paper presents a compiler-based high-level framework that can automatically transform serial uniform mesh code annotated by the user into parallel adaptive mesh code optimized for GPU-accelerated supercomputers. We also present a method for empirical analysis of a uniform mesh to project an upper-bound on achievable speedup of a GPU-optimized AMR code. We show experimental results on three production applications. The speedups of code generated by our framework are comparable to hand-written AMR code while achieving good and weak scaling up to 1000 GPUs. Mohamed Wahib, Naoya Maruyama, Takayuki Aoki |
SC | 1 |
| 2015 | Automated GPU Kernel Transformations in Large-Scale Production Stencil ApplicationsabstractThis paper proposes an end-to-end framework for automatically transforming stencil-based CUDA programs to exploit inter-kernel data locality. The CUDA-to-CUDA transformation collectively replaces the user-written kernels by auto-generated kernels optimized for data reuse. The transformation is based on two basic operations, kernel fusion and fission, and relies on a series of automated steps: gathering metadata, generating graphs expressing dependencies and precedency constraints, searching for optimal kernel fissions/fusions, and generation of optimized code. The framework is modeled to provide the flexibility required for accommodating different applications, allowing the programmer to monitor and amend the intermediate results of different phases of the transformation. We demonstrate the practicality and effectiveness of automatic transformations in exploiting exposed data localities using a variety of real-world applications with large codebases that contain dozens of kernels and data arrays. Experimental results show that the proposed end-to-end automated approach, with minimum intervention from the user, improved performance of six applications with speedups ranging between 1.12x to 1.76x. Mohamed Wahib, Naoya Maruyama |
HPDC | 1 |
| 2014 | Scalable Kernel Fusion for Memory-Bound GPU ApplicationsabstractGPU implementations of HPC applications relying on finite difference methods can include tens of kernels that are memory-bound. Kernel fusion can improve performance by reducing data traffic to off-chip memory, kernels that share data arrays are fused to larger kernels where on-chip cache is used to hold the data reused by instructions originating from different kernels. The main challenges are a) searching for the optimal kernel fusions while constrained by data dependencies and kernels' precedences and b) effectively applying kernel fusion to achieve speedup. This paper introduces a problem definition and proposes a scalable method for searching the space of possible kernel fusions to identify optimal kernel fusions for large problems. The paper also proposes a codeless performance upper-bound projection model to achieve effective fusions. Results show that using the proposed scalable method for kernel fusion improved the performance of two real-world applications containing tens of kernels by 1.35x and 1.2x. Mohamed Wahib, Naoya Maruyama |
SC | 1 |
| 2013 | Highly optimized full GPU-acceleration of non-hydrostatic weather model SCALE-LESabstractSCALE-LES is a non-hydrostatic weather model developed at RIKEN, Japan. It is intended to be a global high-resolution model that would be scaled to exascale systems. This paper introduces the full GPU acceleration of all SCALE-LES modules. Moreover, the paper demonstrates the strategies to handle the unique challenges of accelerating SCALE-LES using GPU. The proposed acceleration is important for identifying the expectations and requirements of scaling SCALE-LES, and similar real world applications, into the exascale era. The GPU implementation includes the optimized GPU acceleration of SCALE-LES for a single GPU with both CUDA Fortran and OpenACC. It also includes scaling SCALE-LES for GPU-accelerated clusters. The results and analysis show how the optimization strategies affect the performance gain in SCALE-LES when moving from conventional CPU clusters towards GPU-powered clusters. Mohamed Wahib, Naoya Maruyama |
CLUSTER | 1 |
| 2011 | Advanced genetic algorithm to solve MINLP problems over GPUabstractIn this paper we propose a many-core implementation of evolutionary computation for GPGPU (General-Purpose Graphic Processing Unit) to solve non-convex Mixed Integer Non-Linear Programming (MINLP) and non-convex Non Linear Programming (NLP) problems using a stochastic algorithm. Stochastic algorithms being random in their behavior are difficult to implement over GPU like architectures. In this paper we not only succeed in implementation of a stochastic algorithm over GPU but show considerable speedups over CPU implementations. The stochastic algorithm considered for this paper is an adaptive resolution approach to genetic algorithm (arGA), developed by the authors of this paper. The technique uses the entropy measure of each variable to adjust the intensity of the genetic search around promising individuals. Performance is further improved by hybridization with adaptive resolution local search (arLS) operator. In this paper, we describe the challenges and design choices involved in parallelization of this algorithm to solve complex MINLPs over a commodity GPU using Compute Unified Device Architecture (CUDA) programming model. Results section shows several numerical tests and performance measurements obtained by running the algorithm over an nVidia Fermi GPU. We show that for difficult problems we can obtain a speedup of up to 20x with double precision and up to 42x with single precision. Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama |
IEEE Congress on Evolutionary Computation | 2 |
| 2011 | Optimization of parallel Genetic Algorithms for nVidia GPUsabstractLed by General Purpose computing over Graphical Processing Units (GPGPUs), the parallel computing area is witnessing a rapid change in dominant parallel systems. A major hurdle in this switch is the Single Instruction Multiple Thread (SIMT) architecture of GPUs which is usually not suitable for the design of legacy parallel algorithms. Genetic Algorithms (GAs) is no exception for that. GAs are commonly parallelized due to the high demanding computational needs. Given the performance of GPGPUs, the need to best exploit them to maximize computing efficiency for parallel GAs is demandingly growing. The goal of this paper is to shed light on the challenges parallel GAs designers/programmers will likely face while trying to achieve this, and to provide some practical advice on how to maximize GPGPU exploitation as a result. To that end, this paper provides a study on adapting legacy parallel GAs on GPGPU systems. The paper exposes the design challenges of nVidia's GPU architecture to the parallel GAs community by: discussing features of GPU, reviewing design issues in GPU relevant to parallel GAs, the design and introduction of new techniques to achieve an efficient implementation for parallel GAs and observing the effect of the pivotal points that both capitalize on the strengths of GPU and limit the deficiencies/overheads of GPUs. The paper demonstrates the performance of designed-for-GPGPU parallel GAs representing the entire spectrum of legacy parallel model of GAs over nVidia Tesla C1060 workstation showing a significant improvement in performance after optimizing and tuning the algorithms for GPU. Mohamed Wahib, Asim Munawar, Masaharu Munetomo, Kiyoshi Akama |
IEEE Congress on Evolutionary Computation | 1 |
| 2011 | A Framework for Cloud Embedded Web Services Utilized by Cloud ApplicationsabstractCloud computing is impacting the modern Internet computing and businesses in every aspect. One feature of clouds is the convenience of using the services offered by the cloud. Consequently, most cloud service providers use WS for users and developers to interface with the cloud. However, the current cloud WS are focused into core and fundamental modern computing functionalities. We anticipate as cloud developments tools mature and cloud applications become more popular, there will be an opportunity for designing and implementing applications/services to be embedded in the cloud for use by applications in the cloud. We propose a framework for WS deployment in the cloud to be usable by applications residing in the same cloud. The framework capitalizes on the cloud strong points to offer a higher value to the service consumer inside the cloud. The authoritative nature of clouds would enable more efficient models for WS publishing, indexing and description. Moreover, being hosted in the cloud, WScan build on the high scalability offered by the cloud with a much higher reliability. Finally, scheduling the instances using the WS in bundle with the WS instances could offer a LAN-like connectivity performance driving down the latency to the magnitude of lower microseconds. In this paper, we highlight the challenges and opportunities of cloud applications using cloud embedded Web services. We give a description of the different aspects by illustrating the different components, together with an end-to-end use case to show the applicability of the proposed system. Mohamed Wahib, Asim Munawar, Masaharu Munetomo, Kiyoshi Akama |
SERVICES | 1 |
| 2010 | A Bayesian Optimization Algorithm for De Novo ligand design based docking running over GPUabstractA principal fragment-based design approach is De Novo ligand design at which small-molecule structures from a database of existing compounds (or compounds that could be made) are docked into the protein binding site following a virtual synthesis scheme. New virtual structures can easily be constructed from combinatorial building blocks. Typically, tens of thousands of orientations are generated for each ligand candidate, therefore global optimization algorithms are usually employed to search the chemical space by generating new molecular structures through probing many different fragments in a combinatorial fashion. We propose using Bayesian Optimization Algorithm (BOA), a meta-heuristic algorithm, in searching the combination of pre-docked fragments through minimizing the energy of ligand-receptor docking. We further introduce the use of GPU (Graphical Processing Unit) to overcome the very long time required in evaluating each possible fragment combination. We show how the GPU utilization enables experimenting larger fragments and target receptors for more complex instances. The experiments resulted in regenerating three drug-like compounds defined in the ZINC database as well as finding a new compound. The Results show how the nVidia's Tesla C1060 GPU was utilized to accelerate the docking process by two orders of magnitude. Mohamed Wahib, Asim Munawar, Masaharu Munetomo, Kiyoshi Akama |
IEEE Congress on Evolutionary Computation | 1 |
| 2010 | The design, usage, and performance of GridUFO: A Grid based Unified Framework for Optimization
Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama |
Future Gener. Comput. Syst. | 2 |
| 2009 | Theoretical and Empirical Analysis of a GPU Based Parallel Bayesian Optimization AlgorithmabstractGeneral purpose computing over graphical processing units (GPGPUs) is a huge shift of paradigm in parallel computing that promises a dramatic increase in performance. But GPGPUs also bring an unprecedented level of complexity in algorithmic design and software development. In this paper we describe the challenges and design choices involved in parallelization of Bayesian optimization algorithm (BOA) to solve complex combinatorial optimization problems over nVidia commodity graphics hardware using compute unified device architecture (CUDA). BOA is a well-known multivariate estimation of distribution algorithm (EDA) that incorporates methods for learning Bayesian network (BN). It then uses BN to sample new promising solutions. Our implementation is fully compatible with modern commodity GPUs and therefore we call it gBOA (BOA on GPU). In the results section, we show several numerical tests and performance measurements obtained by running gBOA over an nVidia Tesla C1060 GPU. We show that in the best case we can obtain a speedup of up to 13x. Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama |
PDCAT | 2 |
| 2008 | Solving Large Instances of Capacitated Vehicle Routing Problem over Cell BEabstractThis paper presents a method to solve large instances of capacitated vehicle routing problem (CVRP) using cellular genetic algorithm (cGA) with local search (LS) over cell broadband engine (cell BE) architecture. We propose a unique parallelization model where computationally intensive local search (LS) runs on the available synergistic processing elements (SPEs) in parallel, while the power processor element (PPE) runs the cGA and acts as a controller for all the SPEs. We reproduce the results from earlier work in PPE only implementation of the algorithm, and we show a considerable reduction in execution time for parallel implementation over cell BE. Moreover, we extended it further to solve larger instances of CVRP (compared to the ones present in the CVRP literature), and got acceptable results in a reasonable amount of time. Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama |
HPCC | 2 |
| 2008 | A Survey: Genetic Algorithms and the Fast Evolving World of Parallel ComputingabstractThis paper gives a survey about the impact of modern parallel/distributed computing paradigms over parallel genetic algorithms (PGAs). Helping the GA community to feel more comfortable with the evolving parallel paradigms, and marking some areas of research for the high-performance computing (HPC) community is the major inspiration behind this survey. In the modern parallel computing paradigms we have considered only two major areas that have evolved very quickly during the past few years, namely, multicore computing and Grid computing. We discuss the challenges involved, and give potential solutions for these challenges. We also propose a hierarchical PGA suitable for Grid environment with multicore computational resources. Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama |
HPCC | 2 |
| 2007 | MHGrid: Towards an Ideal Optimization Environment for Global Optimization Problems Using Grid ComputingabstractThis paper introduces MHGrid, a framework that exploits meta-heuristics based search methods and grid computing to enable the transparent sharing of heterogeneous and dynamic resources offering a grid based global optimization framework. MHGrid allows a user to solve almost all kinds of global optimization problems in a black box manner with a minimal input from the user, it also allows the user to integrate his own solver into MHGrid. In this paper we will discuss the architecture and motivation of such a system. We will also discuss the challenges/complexities involved in constructing MHGrid. Mohamed Wahib, Asim Munawar, Masaharu Munetomo, Kiyoshi Akama |
PDCAT | 1 |