Didem Unat

dblp:96/2997 · also Didem Unat Erten · DBLP profile ↗
← Back
34ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-2351-0770ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 3 first-author · 13 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 UCTRACE: A Multi-Layer Profiling Tool for UCX-driven Communication
Emir Gencer, Mohammad Kefah Taha Issa, Ilyas Turimbetov, James D. Trotter, Didem Unat
IPDPS5
2026 Illuminating Multi-GPU Data Movement
abstract
GPUs have become the accelerators of choice for HPC and machine learning applications, thanks to their massive parallelism and high memory bandwidth. However, as GPU counts per node and across clusters continue to grow, inter-GPU communication has emerged as a major scalability bottleneck, requiring debugging and profiling tool support. In this talk, I will provide an overview of GPU-centric communication within and across compute nodes, highlighting vendor mechanisms and existing tool support for profiling multi- GPU communication. I will then discuss research challenges and opportunities in developing modern profiling tools. Finally, I will emphasize the need for more user-friendly profiling and analysis frameworks that illuminate communication data paths and help developers better understand and optimize data movement across networks. I will conclude by outlining ongoing efforts in our re- search group to address these challenges and advance the state of the art.
Didem Unat
ICPE1
2025 Uniconn: A Uniform High-Level Communication Library for Portable Multi-GPU Programming
abstract
Modern HPC and AI systems increasingly rely on multi-GPU clusters, where communication libraries such as MPI, NCCL/RCCL, and NVSHMEM enable data movement across GPUs. While these libraries are widely used in frameworks and solver packages, their distinct APIs, synchronization models, and integration mechanisms introduce programming complexity and limit portability. Performance also varies across workloads and system architectures, making it difficult to achieve consistent efficiency. These issues present a significant obstacle to writing portable, high-performance code for large-scale GPU systems. We present Uniconn, a unified, portable high-level C++ communication library that supports both point-to-point and collective operations across GPU clusters. Uniconn enables seamless switching between backends and APIs (host or device) with minimal or no changes to application code. We describe its design and core constructs, and evaluate its performance using network benchmarks, a Jacobi solver, and a Conjugate Gradient solver. Across three supercomputers, we compare Uniconn's overhead against CUDA/ROCm-aware MPI, NCCL/RCCL, and NVSHMEM on up to 64 GPUs. In most cases, Uniconn incurs negligible overhead, typically under 1 % for the Jacobi solver and under 2% for the Conjugate Gradient solver.
Dogan Sagbili, Sinan Ekmekçibasi, Khaled Z. Ibrahim, Tan Nguyen 0001, Didem Unat
CLUSTER5
2025 A Device-Side Execution Model for Multi-GPU Task Graphs
abstract
Executing task graphs on multi-GPU systems presents challenges typically managed by CPU-side runtimes, which handle memory management, track dependencies, and balance load.However, the interplay of runtime components, CPUdriven kernel initialization, and dynamic task graph construction creates significant overhead.For static graphs, recent advancements have enabled GPU-side execution, demonstrating substantial performance gains in single-GPU scenarios.However, multi-GPU execution still lags behind in both usability and performance.In particular, no GPU-side solution exists for executing task graphs on multiple nodes.In this work, we introduce Mustard, a multi-GPU execution model that shifts execution of static task graphs entirely to the devices, drastically reducing overhead.Mustard offers a clean solution for executing CUDA graphs across multiple GPUs on multiple nodes without requiring modifications to GPU kernel code or the adoption of new runtime mechanisms or APIs.By transforming the task graph, Mustard enables precise tracking of task dependencies and load balancing directly on the GPU, eliminating the need for host CPU involvement.We evaluate our approach using generated graphs, as well as LU and Cholesky decomposition graphs.In a multi-node scenario with 64 GPUs, Mustard achieves an average 5.83× speedup over the linear algebra library SLATE.On a single node, compared to the best-performing baseline, Mustard delivers an average 1.66× speedup for LU and 1.29× for Cholesky.
Ilyas Turimbetov, Mohamed Wahib, Didem Unat
ICS3
2025 CPU- and GPU-initiated Communication Strategies for Conjugate Gradient Methods on Large GPU Clusters
abstract
The Conjugate Gradient (CG) method is a key building block in numerous applications, yet its low computational intensity and sensitivity to communication overhead make it difficult to scale efficiently on multi-GPU systems. In light of recent advances in multi-GPU communication technologies, we revisit CG parallelization for large-scale GPU clusters.
James D. Trotter, Sinan Ekmekçibasi, Dogan Sagbili, Johannes Langguth, Xing Cai, Didem Unat
SC6
2025 Balanced and Elastic End-to-end Training of Dynamic LLMs
abstract
To reduce the computational and memory overhead of Large Language Models, various approaches have been proposed. These include a) Mixture of Experts (MoEs), where token routing affects compute balance; b) gradual pruning of model parameters; c) dynamically freezing layers; d) dynamic sparse attention mechanisms; e) early exit of tokens as they pass through model layers; and f) Mixture of Depths (MoDs), where tokens bypass certain blocks. While these approaches are effective in reducing overall computation, they often introduce significant workload imbalance across workers. In many cases, this imbalance is severe enough to render the techniques impractical for large-scale distributed training, limiting their applicability to toy models due to poor efficiency.
Mohamed Wahib, Muhammed Abdullah Soyturk, Didem Unat
SC3
2024 Snoopie: A Multi-GPU Communication Profiler and Visualizer
abstract
With data movement becoming one of the most expensive bottlenecks in computing, the need for profiling tools to analyze communication becomes crucial for effectively scaling multi-GPU applications. While existing profiling tools including first-party software by GPU vendors are robust and excel at capturing compute operations within a single GPU, support for monitoring GPU-GPU data transfers and calls issued by communication libraries is currently inadequate. To fill these gaps, we introduce Snoopie, an instrumentation-based multi-GPU communication profiling tool built on NVBit, capable of tracking peer-to-peer transfers and GPU-centric communication library calls. To increase programmer productivity, Snoopie can attribute data movement to the source code line and the data objects involved. It comes with multiple visualization modes at varying granularities, from a coarse view of the data movement in the system as a whole to specific instructions and addresses. Our case studies demonstrate Snoopie’s effectiveness in monitoring data movement, locating performance bugs in applications, and understanding concrete data transfers abstracted beneath communication libraries. The tool is publicly available at https://github.com/ParCoreLab/snoopie.
Mohammad Kefah Taha Issa, Muhammad Aditya Sasongko, Ilyas Turimbetov, Javid Baydamirli, Dogan Sagbili, Didem Unat
ICS6
2023 Multi-GPU Communication Schemes for Iterative Solvers: When CPUs are Not in Charge
abstract
This paper proposes a fully autonomous execution model for multi-GPU applications that completely excludes the involvement of the CPU beyond the initial kernel launch. In a typical multi-GPU application, the host serves as the orchestrator of execution by directly launching kernels, issuing communication calls, and acting as a synchronizer for devices. We argue that this orchestration, or control flow path, causes undue overhead and can be delegated entirely to devices to improve performance in applications that require communication among peers. For the proposed CPU-free execution model, we leverage existing techniques such as persistent kernels, thread block specialization, device-side barriers, and device-initiated communication routines to write fully autonomous multi-GPU code and achieve significantly reduced communication overheads. We demonstrate our proposed model on two broadly used iterative solvers, 2D/3D Jacobi stencil and Conjugate Gradient(CG). Compared to the CPU-controlled baselines, the CPU-free model can improve 3D stencil communication latency by 58.8% and provide a 1.63x speedup for CG on 8 NVIDIA A100 GPUs. The project code is available at https://github.com/ParCoreLab/CPU-Free-model.
Ismayil Ismayilov, Javid Baydamirli, Dogan Sagbili, Mohamed Wahib, Didem Unat
ICS5
2023 Bringing Order to Sparsity: A Sparse Matrix Reordering Study on Multicore CPUs
abstract
Many real-world computations involve sparse data structures in the form of sparse matrices. A common strategy for optimizing sparse matrix operations is to reorder a matrix to improve data locality. However, it's not always clear whether reordering will provide benefits over the unordered matrix, as its effectiveness depends on several factors, such as structural features of the matrix, the reordering algorithm and the hardware that is used. This paper aims to establish the relationship between matrix reordering algorithms and the performance of sparse matrix operations. We thoroughly evaluate six different matrix reordering algorithms on 490 matrices across eight multicore architectures, focusing on the commonly used sparse matrix-vector multiplication (SpMV) kernel. We find that reordering based on graph partitioning provides better SpMV performance than the alternatives for a large majority of matrices, and that the resulting performance is explained through a combination of data locality and load balancing concerns.
James D. Trotter, Sinan Ekmekçibasi, Johannes Langguth, Tugba Torun, Emre Düzakin, Aleksandar Ilic, Didem Unat
SC7
2023 Precise event sampling-based data locality tools for AMD multicore architectures
abstract
Summary We propose ComDetective, an inter‐thread communication analyzer, and ReuseTracker, a reuse distance analyzer, that leverage the hardware features in AMD processors to support low‐overhead profiling. Both tools employ the instruction‐based sampling (IBS) facility and debug registers in AMD processors to detect inter‐thread communication and data reuse. Different from prior arts, ComDetective differentiates the communication into true and false sharing, and ReuseTracker measures reuse distance in private and shared caches by also considering cache line invalidation with low overhead. Both tools can attribute the communications and reuses to source code lines. To our knowledge these tools are two of the few profiling tools designed specifically for AMD x86 architectures using IBS. Our tools are timely and relevant considering the rise in numbers of AMD processor based data centers and HPC systems. We perform experiments to evaluate the accuracy and overheads of the proposed tools on an AMD machine with two‐socket EPYC 7352 processors. ComDetective exhibits high accuracy while introducing 5.14 runtime and 1.4 memory overheads. ReuseTracker also displays high accuracy, which is 95%, with 11.76 runtime and 1.46 memory overheads. These overheads are much lower than the overheads of existing simulators and code instrumentation‐based tools. Lastly, we demonstrate the usage of the tools by having ComDetective and ReuseTracker facilitate the code refactoring of two data mining benchmarks to improve their performance by up to 29%.
Muhammad Aditya Sasongko, Milind Chabbi, Paul H. J. Kelly, Didem Unat
Concurr. Comput. Pract. Exp.4
2023 Precise Event Sampling on AMD Versus Intel: Quantitative and Qualitative Comparison
abstract
Precise event sampling is a profiling feature in commodity processors that can sample hardware events and accurately locate the instructions that trigger the events. This feature has been used in a large number of tools to detect application performance issues. Although precise event sampling is readily supported in modern multicore architectures, vendor supports exhibit great differences that affect their accuracy, stability, overhead, and functionality. This work presents the most comprehensive study to date on benchmarking the event sampling features of Intel PEBS and AMD IBS and performs in-depth analysis on key differences through series of microbenchmarks. Our qualitative and quantitative analysis shows that PEBS allows finer-grained and more accurate sampling of hardware events, while IBS offers richer set of information at each sample though it suffers from lower accuracy and stability. Moreover, OS signal delivery, which is a common method used by the profiling software, introduces significant time overhead to the original overhead incurred by the hardware mechanisms in both PEBS and IBS. We also found that both PEBS and IBS have bias in sampling events across multiple different locations in a code. Lastly, we demonstrate how our findings on microbenchmarks under different thread counts hold for a full-fledged profiling tool that runs on the state-of-the-art Intel and AMD machines. Overall our detailed comparisons serve as a great reference and provide invaluable information for hardware designers and profiling tool developers.
Muhammad Aditya Sasongko, Milind Chabbi, Paul H. J. Kelly, Didem Unat
IEEE Trans. Parallel Distributed Syst.4
2022 Mixed and Multi-Precision SpMV for GPUs with Row-wise Precision Selection
abstract
Sparse Matrix-Vector Multiplication (SpMV) is one of the key memory-bound kernels commonly used in industrial and scientific applications. To improve its data movement and benefit from higher compute rates, there are several efforts to utilize mixed precision on SpMV. Most of the prior-art focus on performing the entire SpMV in single-precision within a bigger context of an iterative solver (e.g., CG, GMRES). In this work, we are interested in a more fine-grained mixed-precision SpMV, where the level of precision is decided for each element in the matrix to be used in a single operation. We extend an existing entry-wise precision based approach by deciding precisions per row, motivated by the granularity of parallelism on a GPU where groups of threads process rows in CSR-based matrices. We propose mixed-precision CSR storage methods with row permutations and describe their greater efficiency and load-balancing compared to the existing method. We also consider a multi-precision case where single and double precision copies of the matrix are stored priorly and further extend our mixed-precision SpMV approach to comply with it. As such, we leverage a mixed-precision SpMV to obtain a multi-precision Jacobi method which is faster than yet almost as accurate as double-precision Jacobi implementation, and further evaluate a multi-precision Cardiac modeling algorithm. We demonstrate the effectiveness of the proposed SpMV methods on an extensive dataset of real-valued large sparse matrices from the SuiteSparse Matrix Collection using an NVIDIA V100 GPU.
Erhan Tezcan, Tugba Torun, Fahrican Kosar, Kamer Kaya, Didem Unat
SBAC-PAD5
2022 ReuseTracker: Fast Yet Accurate Multicore Reuse Distance Analyzer
abstract
One widely used metric that measures data locality is reuse distance —the number of unique memory locations that are accessed between two consecutive accesses to a particular memory location. State-of-the-art techniques that measure reuse distance in parallel applications rely on simulators or binary instrumentation tools that incur large performance and memory overheads. Moreover, the existing sampling-based tools are limited to measuring reuse distances of a single thread and discard interactions among threads in multi-threaded programs. In this work, we propose ReuseTracker —a fast and accurate reuse distance analyzer that leverages existing hardware features in commodity CPUs. ReuseTracker is designed for multi-threaded programs and takes cache-coherence effects into account. By utilizing hardware features like performance monitoring units and debug registers, ReuseTracker can accurately profile reuse distance in parallel applications with much lower overheads than existing tools. It introduces only 2.9× runtime and 2.8× memory overheads. Our tool achieves 92% accuracy when verified against a newly developed configurable benchmark that can generate a variety of different reuse distance patterns. We demonstrate the tool’s functionality with two use-case scenarios using PARSEC, Rodinia, and Synchrobench benchmark suites where ReuseTracker guides code refactoring in these benchmarks by detecting spatial reuses in shared caches that are also false sharing and successfully predicts whether some benchmarks in these suites can benefit from adjacent cache line prefetch optimization.
Muhammad Aditya Sasongko, Milind Chabbi, Mandana Bagheri-Marzijarani, Didem Unat
ACM Trans. Archit. Code Optim.4
2021 A computational-graph partitioning method for training memory-constrained DNNs
Fareed Qararyah, Mohamed Wahib, Doga Dikbayir, Mehmet Esat Belviranli, Didem Unat
Parallel Comput.5
2021 A Split Execution Model for SpTRSV
abstract
Sparse Triangular Solve (SpTRSV) is an important and extensively used kernel in scientific computing. Parallelism within SpTRSV depends upon matrix sparsity pattern and, in many cases, is non-uniform from one computational step to the next. In cases where the SpTRSV computational steps have contrasting parallelism characteristics- some steps are more parallel, others more sequential in nature, the performance of an SpTRSV algorithm may be limited by the contrasting parallelism characteristics. In this work, we propose a split-execution model for SpTRSV to automatically divide SpTRSV computation into two sub-SpTRSV systems and an SpMV, such that one of the sub-SpTRSVs has more parallelism than the other. Each sub-SpTRSV is then computed using different SpTRSV algorithms, which are possibly executed on different platforms (CPU or GPU). By analyzing the SpTRSV Directed Acyclic Graph (DAG) and matrix sparsity features, we use a heuristics-based approach to (i) automatically determine the suitability of an SpTRSV for split-execution, (ii) find the appropriate split-point, and (iii) execute SpTRSV in a split fashion using two SpTRSV algorithms while managing any required inter-platform communication. Experimental evaluation of the execution model on two CPU-GPU machines with a matrix dataset of 327 matrices from the SuiteSparse Matrix Collection shows that our approach correctly selects the fastest SpTRSV method (split or unsplit) for 88 percent of matrices on the Intel Xeon Gold (6148) + NVIDIA Tesla V100 and 83 percent on the Intel Core I7 + NVIDIA G1080 Ti platform achieving speedups up to 10x and 6.36x respectively.
Najeeb Ahmad, Buse Yilmaz, Didem Unat
IEEE Trans. Parallel Distributed Syst.3
2020 A Prediction Framework for Fast Sparse Triangular Solves
Najeeb Ahmad, Buse Yilmaz, Didem Unat
Euro-Par3
2020 Tiling-Based Programming Model for Structured Grids on GPU Clusters
abstract
Currently, more than 25% of supercomputers employ GPUs due to their massively parallel and power-efficient architectures. However, programming GPUs efficiently in a large scale system is a demanding task not only for computational scientists but also for programming experts as multi-GPU programming requires managing distinct address spaces, generating GPU-specific code and handling inter-device communication. To ease the programming effort, we propose a tiling-based high-level GPU programming model for structured grid problems. The model abstracts data decomposition, memory management and generation of GPU specific code, and hides all types of data transfer overheads. We demonstrate the effectiveness of the programming model on a heat simulation and a real-life cardiac modeling on a single GPU, on a single node with multiple-GPUs and multiple-nodes with multiple-GPUs. We also present performance comparisons under different hardware and software configurations. The results show that the programming model successfully overlaps communication and provides good speedup on 192 GPUs.
Burak Bastem, Didem Unat
HPC Asia2
2020 Adaptive Level Binning: A New Algorithm for Solving Sparse Triangular Systems
abstract
Sparse triangular solve (SpTRSV) is an important scientific kernel used in several applications such as preconditioners for Krylov methods. Parallelizing SpTRSV on multi-core systems is challenging since it exhibits limited parallelism due to computational dependencies and introduces high parallelization overhead due to finegrained and unbalanced nature of workloads. We propose a novel method, named Adaptive Level Binning (ALB), that addresses these challenges by eliminating redundant synchronization points and adapting the work granularity with an efficient load balancing strategy. Similar to the commonly used level-set methods for solving SpTRSV, ALB constructs level-sets of rows, where each level can be computed in parallel. Differently, ALB bins rows to levels adaptively and reduces redundant dependencies between rows. On an Intel® Xeon® Gold 6148 processor and NVIDIA® Tesla V100 GPU, ALB obtains 1.83x speedup on average and up to 5.28x speedup over Intel MKL and, over NVIDIA cuSPARSE, an average speedup of 2.80x and a maximum speedup of 39.40x for 29 matrices selected from Suite Sparse Matrix Collection.
Buse Yilmaz, Bugrra Sipahiogrlu, Najeeb Ahmad, Didem Unat
HPC Asia4
2019 ComDetective: a lightweight communication detection tool for threads
abstract
Inter-thread communication is a vital performance indicator in shared-memory systems. Prior works on identifying inter-thread communication employed hardware simulators or binary instrumentation and suffered from inaccuracy or high overheads---both space and time---making them impractical for production use. We propose ComDetective, which produces communication matrices that are accurate and introduces low runtime and low memory overheads, thus making it practical for production use.
Muhammad Aditya Sasongko, Milind Chabbi, Palwisha Akhtar, Didem Unat
SC4
2018 Runtime Determinacy Race Detection for OpenMP Tasks
Hassan Salehe Matar, Didem Unat
Euro-Par2
2018 Phase-Based Data Placement Scheme for Heterogeneous Memory Systems
abstract
Heterogeneous memory systems are equipped with two or more types of memories, which work in tandem to complement the capabilities of each other. The multiple memories can vary in latency, bandwidth and capacity characteristics across systems and they come in various configurations that can be managed by the programmer. This introduces an added programming complexity for the programmer. In this paper, we present a dynamic phase-based data placement scheme to assist the programmer in making decisions about program object allocations. We devise a cost model to assess the benefit of having an object in one type of memory over the other and apply the cost model at every application phase to capture the dynamic behaviour of an application. Our cost model takes into account the reference counts of objects and incurred transfer overhead when making a suggestion. In addition, objects can be transferred across memories asynchronously between phases to mask some of the transfer overhead. We test our cost model with a diverse set of applications from NAS Parallel and Rodinia benchmarks and perform experiments on Intel KNL, which is equipped with a high bandwidth memory (MCDRAM) and a high capacity memory (DDR). Our dynamic phase-based data placement performs better than initial placement and achieves comparable or better performance than cache mode of MCDRAM.
Mohammad Shakeel Laghari, Najeeb Ahmad, Didem Unat
SBAC-PAD3
2018 Phase asynchronous AMR execution for productive and performant astrophysical flows
Muhammed Nufail Farooqi, Tan Nguyen 0001, Weiqun Zhang, Ann S. Almgren, John Shalf, Didem Unat
SC6
2018 Fast multidimensional reduction and broadcast operations on GPU for machine learning
abstract
Summary Reduction and broadcast operations are commonly used in machine learning algorithms for different purposes. They widely appear in the calculation of the gradient values of a loss function, which are one of the core structures of neural networks. Both operations are implemented naively in many libraries usually for scalar reduction or broadcast; however, to our knowledge, there are no optimized multidimensional implementations available. This fact limits the performance of machine learning models requiring these operations to be performed on tensors. In this work, we address the problem and propose two new strategies that extend the existing implementations to perform on tensors. We introduce formal definitions of both operations using tensor notations, investigate their mathematical properties, and exploit these properties to provide an efficient solution for each. We implement our parallel strategies and test them on a CUDA enabled Tesla K40 m GPU accelerator. Our performant implementations achieve up to 75% of the peak device memory bandwidth on different tensor sizes and dimensions. Significant speedups against the implementations available in the Knet Deep Learning framework are also achieved for both operations.
Doga Dikbayir, Enis Berk Çoban, Ilker Kesen, Deniz Yuret, Didem Unat
Concurr. Comput. Pract. Exp.5
2018 BindMe: A thread binding library with advanced mapping algorithms
abstract
Summary Binding parallel tasks to cores according to a placement policy is one of the key aspects to achieve good performance in multicore machines because it can reduce on‐chip communication among parallel threads. Binding also prevents operating system from migrating threads, which improves data locality. However, there is no single mapping policy that works best among all different kinds of applications and platforms because each machine has a different topology and each application exhibits different communication pattern. Determining the best policy for a given application and machine requires extra programming effort. To relieve the programmer from that burden, we introduce BindMe, a thread binding library that assists programmer to bind threads to underlying hardware. BindMe incorporates state‐of‐the‐art mapping algorithms, which use communication pattern in an application to formulate an efficient task placement policy. We also introduce ChoiceMap, a communication aware mapping algorithm that respects mutual priorities of parallel tasks and performs a fair mapping by reducing communication volume among cores. We have tested BindMe and ChoiceMap with various applications from NAS parallel benchmark and Rodinia bechmark. Our results show that choosing a mapping policy that best suits the application behavior can increase its performance and no single policy gives the best performance across different applications.
Pirah Noor Soomro, Muhammad Aditya Sasongko, Didem Unat
Concurr. Comput. Pract. Exp.3
2018 Special issue on High performance computing conference (BASARIM-2017)
abstract
Numerical operations and calculations have been studied on parallel systems for many years for the purpose of achieving faster results and better performance. In recent years, these studies and their results have reached a certain level and the progress made in this regard has accelerated. The Turkish High Performance Computing Conference has been organized since 2009 for discussing challenges in high performance computing and identifying opportunities in the areas of cloud computing and big data processing for moving ahead. To this end, the integration of big data processing software stack and traditional parallel computing message passing protocols has been addressed in the keynote talk "HPC-enhanced IoT and Data-based Grid" by Geoffrey Charles Fox at the 5th Turkish High Performance Computing Conference (BASARIM-2017). With the objective of sharing and evaluating scientific research, experience, studies, and results related to high performance computing at this scientific event, this special issue presents the highlights of the program and selected high-quality papers from BASARIM-2017. Aktemur1 presents a new sparse matrix-vector multiplication (SpMV) implementation, named CSRLenGoto, which is based on complete loop unrolling. The method provides performance improvements especially for matrices whose row length is short. CSRLenGoto incurs an inexpensive preprocessing phase that is compensated in just a few SpMV repetitions. The author parallelized CSRLenGoto and integrated the operation into a state-of-the-art matrix partitioning approach as the kernel operation. The author observed up to 2.46× and on the average 1.29× speedup with respect to highly optimized Intel MKL's SpMV for matrices with short- or medium-length rows. Topcuoglu et al2 present a generic private information retrieval (PIR) scheme with parallel multi-exponentiations for multicore architectures. Without revealing any information as to which a data item is requested, PIR allows the data owners to share and/or retrieve data on remote repositories. Homomorphic cryptosystems are commonly exploited for PIR in the literature. Those approaches require not one but many modular exponentiations need to be computed and multiplied to obtain the desired result. The multi-exponentiation operation can be implemented by exponentiating the bases to their corresponding exponents one-by-one and the operation is extremely easy to parallelize. However, when the operation is considered as a whole, it can be computed in a more performance efficient way, but the combined multi-exponentiation is not straightforward to parallelize. Topcuoglu et al propose a generic tensor-based PIR scheme that is efficient and novel to parallelize multi-exponentiations on multicore processors with perfect load balance. The evaluation results demonstrate that the proposed load balancing methods make a parallel multi-exponentiation faster. Mumcuyan et al3 have developed a method for optimally bipartitioning sparse matrices with reordering and parallelism. Because communication between tasks can constitute one of the bottlenecks for scalability, an efficient task-to-processor assignment is crucial for performance. A main solution to this problem is to model the tasks as a hypergraph where the pins and nets represent the tasks and the communication among them, respectively. The hypergraph vertices are partitioned into a number of parts, which correspond to processors, in a way that the total number of vertices for each part is balanced and the amount of edges having endpoints in different parts is optimized to be minimal. Recently, for solving sparse matrix bipartitioning, a novel purely combinatorial approach has been proposed. The approach is based on branch-and-bound and can handle hypergraphs that cannot be optimally partitioned by using existing methods because of the problem's complexity. The work of Mumcuyan et al built on top of the previous study with three new ideas. (1) It applies matrix ordering methods to use more information in the earlier branches of the tree, (2) it leverages machine learning to select an ordering based on the matrix features, and (3) it parallelizes the search of an optimal bipartitioning. The results show that their techniques make the bipartitioning sparse matrices significantly faster. Dikbayir et al4 have developed two parallel and efficient implementations of broadcast and reduction operations for multidimensional arrays on GPU devices. Reduction and broadcast operations are commonly used in machine learning, especially when performing forward and backward passes in deep neural networks. Existing parallel methods usually developed for scalar reduction but with the increasing size of data and dimensionality, the need for implementations suitable for multidimensional tensors has emerged. The authors first set up a terminology for their methods and define both of operations mathematically in the high-dimensional space. They then analyze and evaluate the mathematical nature of the original algorithms in order to exploit any properties. For reduction, they take advantage of its associativity property to reduce multiple dimensions within a single kernel launch to minimize the amount of synchronizations needed in the process. For broadcast, they adapt data reuse in order to avoid replication of the input data in memory. Both of the implementations use coordinate calculation and index translation algorithms in order to map the GPU threads to the corresponding data elements of the multidimensional input tensors. The authors evaluate the algorithms implemented in CUDA on the NVIDIA K40 GPU accelerator. They compare the performance of the methods to the existing implementations in Knet, which is a deep learning framework. The new implementations are able to obtain up to 56x speed up over Knet while reaching %75 of the theoretical limit for memory bandwidth rate of K40 device for test scenarios involving large multidimensional tensors. Guler and Ozkasap5 address combining checkpointing techniques with primary-backup replication protocol to further improve efficiency in terms of client blocking time and overall system throughput. For this purpose, they propose an advanced primary-backup replication protocol, which minimizes the failover time by eliminating the recovery process in the event of rollback operation. The authors develop a software framework for a geographically replicated key-value store based on RocksDB and use the PlanetLab overlay network to execute the proposed primary-backup replication protocol. They conduct a thorough analysis of various checkpointing algorithms integrated with primary-backup replication. Using various metrics of interest including blocking time, checkpointing time, checkpoint size, failover time and throughput, and testing with realistic workloads, their findings indicate that the proposed primary-backup replication protocol, supported by Snappy-compressed-periodic-incremental-checkpointing technique, provides significant improvements in the system throughput and reduced blocking times compared to the traditional primary-backup replication protocol. Soomro et al6 propose BindMe, which is a thread binding library that supports advanced mapping algorithms. Current multicore machines contain a large number of cores and the number of cores is expected to increase in upcoming exascale multicore machines. Binding parallel tasks to cores according to a placement policy is one of the important performance boosting factors, because it can reduce on-chip communication among parallel threads. Binding also prevents operating system from migrating threads, which improves data locality. However, there is no single mapping policy that works best among all different kinds of applications and platforms because each machine has a different topology and each application exhibits different communication pattern. Determining the best policy for a given application and machine requires extra programming effort. The authors introduce the BindMe, a thread binding library that assists programmer to bind threads to underlying hardware. BindMe incorporates state-of-the-art mapping algorithms, which analyze communication pattern of an application to formulate a task placement policy. The authors also introduce ChoiceMap, a communication aware mapping algorithm that respects mutual priorities of parallel tasks and performs a fair mapping by reducing communication volume among cores. Their results show that choosing a mapping policy that best suits the application behavior can increase its performance and no single policy gives the best performance across different applications. Fisne and Ozsoy7 provide a real-time running software defined radio (SDR) with its full pipeline steps running on GPGPUs. Initially, they port the data preprocess step on to GPUs with Big Endian/Little Endian Conversion and Short/Complex Type Conversion. Second, FFT Process for Spectrum and Spectogram is ported on GPU. Signal Detection is achieved afterwards and now the wideband data is ready for down conversion to be reduced to narrowband data. For DDC operation, they have used a different technique then traditional approaches, where using FFT/IFFT blocks for filtering. After downsampling, narrowband data is processed for extracting the sound. FM demodulation and resampling steps are applied for this purpose. For all steps, the algorithms are parallelized both on CPU and GPU, design choices are given, performance results are listed, and analyses are made for reaching to real-time performance. The last demodulation steps are only implemented on CPU since these steps are sufficient enough for real-time requirements. Consequently, their work provides a design of a full running SDR implementation under real-time requirements. Muhtaroglu et al8 investigated several design choices for HPC services at different layers of the cloud computing architecture to simplify and broaden its use cases. They compared direct versus iterative parallel linear equation solvers for the platform-as-a-service layer. They observed that several matrix properties, identified before starting long-running solvers, can help HPC services automatically select the amount of computing resources per job, such that the job latency is minimized and the overall job throughput is maximized. They showed that, on top of the 2x-3x speedups gained from parallelization, one can achieve an additional 2x-3x speedup with careful selection of solver types and preconditioner combinations. They explored HPC application performance, load isolation, and deployment issues using application containers (Docker) while also comparing them to physical and virtual machines for the infrastructure-as-a-service layer. They found that Dockerized HPC can be setup and deployed much faster than physical or virtual HPC alternatives and its performance is comparable. Baeth and Aktas9 introduced a generic software architecture that can be integrated with existing social media software to enable users to track the dissemination of their data and generate special notifications by using complex event processing. Their solution utilizes social provenance data. The proposed architecture is designed to detect the candidate social media user accounts for copyright violations. They developed a prototype of this proposed architecture. The developed system has a set of extendible facade classes responsible for hiding the complexities of the utilized streaming and complex event processing engines. In turn, this approach makes their implementation highly decoupled. In addition, it adds an extra level of layers segregation by making it much easier to switch to different tools, libraries, and technologies. To facilitate testing of the software architecture, they developed a large-scale synthetic provenance dataset, discussed the details of the prototype implementation, and evaluated its performance. Their prototype performed well with the ingested large number of provenance workflows because the processing overhead is negligible. Deep learning has emerged as an effective solution to various text mining problems such as document classification and clustering, document summarization, web mining, and sentiment analysis. Karakus et al10 investigate several deep learning models for binary sentiment classification problem. They report a detailed comparison of the models in terms of accuracy and time performances. Two major deep learning architectures used in their study are Convolutional Neural Networks and Long Short-Term Memory. Karakus et al built several variants of these models by changing the number of layers, tuning the hyper-parameters, and combining models. They investigated the effect of using the pre-word embeddings with these models. Their experimental results have shown that the use of word embeddings with deep neural networks effectively yields performance improvements in terms of run-time and accuracy. Cloud computing provides scalable computing resources on demand. Challenges in cloud computing monitoring systems include detecting patterns that might lead to failure of the cloud system, detecting malfunctioning problems within the cloud platform after they occur, and issues related to the fact that existing monitoring solutions are tightly coupled to specific cloud platforms. To address these challenges, Aktas11 designed a hybrid cloud monitoring software architecture that can work as an add-on layer on top of existing cloud computing platforms. The proposed architecture is designed based on facade software design pattern and utilizes complex event processing concept in which data from various primitive metrics streams are processed to detect previously defined patterns. Prototype applications have been developed to demonstrate the architecture's usability. Performance tests were applied to prototype applications. Computation times required for the operation of the proposed architecture were found to be negligible.
Didem Unat, Mehmet S. Aktas
Concurr. Comput. Pract. Exp.1
2018 Output nondeterminism detection for programming models combining dataflow with shared memory
Hassan Salehe Matar, Erdal Mutlu, Serdar Tasiran, Didem Unat
Parallel Comput.4
2017 Nonintrusive AMR Asynchrony for Communication Optimization
Muhammed Nufail Farooqi, Didem Unat, Tan Nguyen 0001, Weiqun Zhang, Ann S. Almgren, John Shalf
Euro-Par2
2017 Overlapping Data Transfers with Computation on GPU with Tiles
abstract
GPUs are employed to accelerate scientific applications however they require much more programming effort from the programmers particularly because of the disjoint address spaces between the host and the device. OpenACC and OpenMP 4.0 provide directive based programming solutions to alleviate the programming burden however synchronous data movement can create a performance bottleneck in fully taking advantage of GPUs. We propose a tiling based programming model and its library that simplifies the development of GPU programs and overlaps the data movement with computation. The programming model decomposes the data and computation into tiles and treats them as the main data transfer and execution units, which enables pipelining the transfers to hide the transfer latency. Moreover, partitioning application data into tiles allows the programmer to still take advantage of GPU even though application data cannot fit into the device memory. The library leverages C++ lambda functions, OpenACC directives, CUDA streams and tiling API from TiDA to support both productivity and performance. We show the performance of the library on a data transfer-intensive and a compute-intensive kernels and compare its speedup against OpenACC and CUDA. The results indicate that the library can hide the transfer latency, handle the cases where there is no sufficient device memory, and achieves reasonable performance.
Burak Bastem, Didem Unat, Weiqun Zhang, Ann S. Almgren, John Shalf
ICPP2
2017 EmbedSanitizer: Runtime Race Detection Tool for 32-bit Embedded ARM
Hassan Salehe Matar, Serdar Tasiran, Didem Unat
RV3
2017 Object Placement for High Bandwidth Memory Augmented with High Capacity Memory
abstract
High bandwidth memory (HBM) is a new emerging technology that aims to improve the performance of bandwidth limited applications. Even though it provides high bandwidth, it must be augmented with DRAM to meet the memory capacity requirement of any applications. Due to this limitation, objects in an application should be optimally placed on the heterogeneous memory subsystems. In this study, we propose an object placement algorithm that places program objects to fast or slow memories in case the capacity of fast memory is insufficient to hold all the objects to increase the overall application performance. Our algorithm uses the reference counts and type of references (read or write) to make an initial placement of data. In addition, we perform various memory bandwidth benchmarks to be used in our placement algorithm on Intel Knights Landing (KNL) architecture. Not surprisingly high bandwidth memory sustains higher read bandwidth than write bandwidth, however, placing write-intensive data on HBM results in better overall performance because write-intensive data is punished by the DRAM speed more severely compared to read intensive data. Moreover, our benchmarks demonstrate that if a basic block makes references to both types of memories, it performs worse than if it makes references to only one type of memory in some cases. We test our proposed placement algorithm with 6 applications under various system configurations. By allocating objects according to our placement scheme, we are able to achieve a speedup of up to 2x.
Mohammad Shakeel Laghari, Didem Unat
SBAC-PAD2
2017 Trends in Data Locality Abstractions for HPC Systems
abstract
The cost of data movement has always been an important concern in high performance computing (HPC) systems. It has now become the dominant factor in terms of both energy consumption and performance. Support for expression of data locality has been explored in the past, but those efforts have had only modest success in being adopted in HPC applications for various reasons. them However, with the increasing complexity of the memory hierarchy and higher parallelism in emerging HPC systems, locality management has acquired a new urgency. Developers can no longer limit themselves to low-level solutions and ignore the potential for productivity and performance portability obtained by using locality abstractions. Fortunately, the trend emerging in recent literature on the topic alleviates many of the concerns that got in the way of their adoption by application developers. Data locality abstractions are available in the forms of libraries, data structures, languages and runtime systems; a common theme is increasing productivity without sacrificing performance. This paper examines these trends and identifies commonalities that can combine various locality concepts to develop a comprehensive approach to expressing and managing data locality on future large-scale high-performance computing systems.
Didem Unat, Anshu Dubey, Torsten Hoefler, John Shalf, Mark James Abraham, Mauro Bianco, Bradford L. Chamberlain, Romain Cledat, H. Carter Edwards, Hal Finkel, Karl Fürlinger, Frank Hannig, Emmanuel Jeannot, Amir Kamil, Jeff Keasler, Paul H. J. Kelly, Vitus J. Leung, Hatem Ltaief, Naoya Maruyama, Chris J. Newburn, Miquel Pericàs
IEEE Trans. Parallel Distributed Syst.1
2016 Perilla: metadata-based optimizations of an asynchronous runtime for adaptive mesh refinement
abstract
Hardware architecture is increasingly complex, urging the development of asynchronous runtime systems with advance resource and locality management supports. However, these supports may come at the cost of complicating the user interface while programming remains one of the major constraints to wide adoption of asynchronous runtimes in practice. In this paper, we propose a solution that leverages application metadata to enable challenging optimizations as well as to facilitate the task of transforming legacy code to an asynchronous representation. We develop Perilla, a task graph-based runtime system that requires only modest programming effort. Perilla utilizes metadata of an AMR software framework to enable various optimizations at the communication layer without complicating its API. Experimental results with different applications on up to 24K processor cores show that Perilla can realize up to 1.44x speedup over the synchronous code variant. The metadata enabled optimizations account for 25% to 100% of the performance improvement.
Tan Nguyen 0001, Didem Unat, Weiqun Zhang, Ann S. Almgren, Muhammed Nufail Farooqi, John Shalf
SC2
2011 Mint: realizing CUDA performance in 3D stencil methods with annotated C
abstract
We present Mint, a programming model that enables the non-expert to enjoy the performance benefits of hand coded CUDA without becoming entangled in the details. Mint targets stencil methods, which are an important class of scientific applications. We have implemented the Mint programming model with a source-to-source translator that generates optimized CUDA C from traditional C source. The translator relies on annotations to guide translation at a high level. The set of pragmas is small, and the model is compact and simple. Yet, Mint is able to deliver performance competitive with painstakingly hand-optimized CUDA. We show that, for a set of widely used stencil kernels, Mint realized 80% of the performance obtained from aggressively optimized CUDA on the 200 series NVIDIA GPUs. Our optimizations target three dimensional kernels, which present a daunting array of optimizations.
Didem Unat, Xing Cai, Scott B. Baden
ICS1
2009 An Adaptive Sub-sampling Method for In-memory Compression of Scientific Data
abstract
A current challenge in scientific computing is how to curb the growth of simulation datasets without losing valuable information. While wavelet based methods are popular, they require that data be decompressed before it can analyzed, for example, when identifying time-dependent structures in turbulent flows. We present adaptive coarsening, an adaptive subsampling compression strategy that enables the compressed data product to be directly manipulated in memory without requiring costly decompression.We demonstrate compression factors of up to 8 in turbulent flow simulations in three dimensions.Our compression strategy produces a non-progressive multiresolution representation, subdividing the dataset into fixed sized regions and compressing each region independently.
Didem Unat, Theodore Hromadka III, Scott B. Baden
DCC1