EDBT 2026 Demo / reviewers in the wild / expert
Nikoli Dryden
dblp:148/1273
· DBLP profile ↗
22ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-9965-3647ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoDL: A Framework for Studying Cross-Component Interference in Deep Learning Training PipelinesabstractDeep learning training on HPC systems exhibits significant step-time variability—data stalls can dominate 50–85% of training time [12, 31]—that limits efficiency. Existing benchmarks and profilers isolate individual components: I/O benchmarks replace compute with sleep, compute profilers ignore I/O, and GPU tracers operate on separate timescales. This isolation prevents understanding real variability, which stems from interactions among data loading, compute, and communication competing for shared CPU resources. We present CoDL, a framework that combines component-isolation benchmarking with multi-level tracing on a unified microsecond timeline, enabling diagnosis that reframes optimization from isolated component throughput to CPU–GPU scheduling coordination. On a single 4-GPU node, evaluation of ResNet-50 shows compute duration more than doubles (+127.4%) under active data loading and NCCL AllReduce degrades 2.8 ×, despite zero direct GPU-kernel overlap with data loading; the dominant mechanism is data-loading workers preempting GPU-launch threads on the CPU, inflating CUDA runtime API latency by 3.3 ×. Druva Dhakshinamoorthy, Ray A. O. Sinurat, Nikoli Dryden, Arnab Kumar Paul, Hariharan Devarajan |
HPDC | 3 |
| 2026 | HORATIO: Bridging Management and Analysis of Traces at ScaleabstractModern scientific and deep learning workloads on HPC systems rely on profiling and tracing across multiple software and hardware layers, generating diagnostic traces that often reach terabyte scale. Existing approaches manage these traces along two axes: trace formats and analysis tooling. Practitioners often convert raw traces into queryable formats, but doing so nearly doubles storage when raw files are retained for compatibility and requires upfront schema discovery that profiling tools cannot guarantee. A cleaner path is to make raw traces efficient in place, but this requires overcoming three limitations: lack of selective querying, analysis throughput bottlenecks, and lack of physical clustering. To address these limitations jointly, we developed Horatio, a raw trace management framework that indexes, analyzes, and physically clusters raw traces directly. Three findings emerge from our work. First, Horatio stores a lightweight RocksDB-backed auxiliary index alongside the raw trace, including gzip checkpoints, per-chunk bloom filters, and chunk-level statistics, delivering selective queries up to 75 × faster than naive Parquet at ∼ 1.01 × raw storage, with the highest cross-query mean throughput (232 M events/s) among state-of-the-art formats. Second, offloading event-level computation to a native C++ backend while keeping Dask for orchestration yields 80–83 × speedup over the original Dask-based DFAnalyzer. Third, lossless trace clustering that preserves the same input format yields a further 1.8–5.5 × end-to-end speedup across four AI and scientific workloads, and up to 230 × on h5bench where the preset aligns tightly with the cluster boundary, all with original layouts reconstructible on demand. Across five AI and scientific workloads, Horatio completes pipelines that DFAnalyzer cannot finish within 8 hours. On a 2.2 TB uncompressed trace, Horatio’s MPI-based mode scales to 16 × at 32 nodes on the full event set and its Dask-based path peaks at 4 × on a preset-filtered workload, both completing where DFAnalyzer hits OOM at every scale. Ray A. O. Sinurat, William Nixon, Haryadi S. Gunawi, Nikoli Dryden, Hariharan Devarajan |
SSDBM | 4 |
| 2026 | WADO: A Distributed WORM Storage Service for Asynchronous Data OperationsabstractAI-driven scientific workloads increasingly depend on data-intensive input pipelines, where deep learning frameworks must ingest and transform large datasets from hierarchical HPC storage. Existing system-centric data services improve movement and locality between the parallel file system (PFS), node-local storage, and memory. However, they do not directly optimize how input pipeline operations execute across scopes, stage overlap, and resource-specific parallelism. As scale grows, this gap causes worker stalls, contention, and poor hardware utilization. We present WADO, a distributed write-once-read-many (WORM) object-store runtime for data-centric workloads that closes this gap through three coordinated mechanisms: scope-centric processing, explicit pipeline decomposition, and interference-aware explicit parallelism. WADO dynamically maps operations to execution scopes, overlaps stages such as I/O, communication, and transformations, and applies contention-aware concurrency control to match hardware behavior at runtime. Our evaluation shows three main findings: (1) scope-centric processing preserves throughput under scale, improving mixed-operation throughput by up to 1.65 × ; (2) explicit pipeline decomposition converts serialized wait into overlapped progress, delivering up to 2.16 × higher sustained bandwidth; and (3) interference-aware explicit parallelism improves effective bandwidth by up to 4.4 × by avoiding oversubscription collapse. On Unet3D model training, these mechanisms translate to end-to-end gains, improving data loading performance by 4.1 × compared to baseline PyTorch on Lustre, and 1.51 × compared to DYAD, enabled by deeper pipelining, adaptive parallelism, and near-data transformation offloading. Karim Youssef, Hariharan Devarajan, Nikoli Dryden, Roger A. Pearce |
SSDBM | 3 |
| 2026 | STen: Productive and Efficient Sparsity in PyTorchabstractAs deep learning models grow, sparsity is becoming an increasingly critical component of deep neural networks, enabling improved performance and reduced storage. However, existing frameworks offer poor support for sparsity. Specialized sparsity engines focus exclusively on sparse inference, while general frameworks primarily focus on sparse tensors in classical formats and neglect the broader sparsification pipeline necessary for using sparse models, especially during training. Further, existing frameworks are not easily extensible: adding a new sparse tensor format or operator is challenging and time-consuming. To address this, we propose STen, a sparsity programming model and interface for PyTorch whose key design insight is the decoupling of sparsity layouts, operators, and sparsifiers into composable, first-class abstractions that users can independently define and combine. An automatic dispatch mechanism selects the best available sparse implementation and transparently falls back to dense execution, allowing STen to support virtually all sparsification methods while enabling rapid prototyping without sacrificing performance for optimized paths. We demonstrate the versatility of STen by expressing existing sparsification techniques within its abstraction, achieving a code size reduction of over 2×. Finally, we develop a novel, high-performance grouped n : m sparsity layout for CPU inference at moderate sparsity, accelerating end-to-end BERT BASE inference by up to 3.2×. STen brings high performance and ease of use, making sparsity readily accessible for existing PyTorch models. Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Timo Schneider, Saleh Ashkboos, Torsten Hoefler |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | Scaling Large-scale GNN Training to Thousands of Processors on CPU-based SupercomputersabstractGraph Convolutional Networks (GCNs), particularly for largescale graphs, are crucial across numerous domains.However, training distributed full-batch GCNs on large-scale graphs suffers from inefficient memory access patterns and high communication overhead.To address these challenges, we introduce SuperGCN, an efficient and scalable distributed GCN Chen Zhuang, Lingqi Zhang 0001, Du Wu, Peng Chen 0035, Jiajun Huang 0001, Xin Liu 0020, Rio Yokota, Nikoli Dryden, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
ICS | 8 |
| 2025 | Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual SparsityabstractAccelerating large language model (LLM) inference is critical for real-world deployments requiring high throughput and low latency. Contextual sparsity, where each token dynamically activates only a small subset of the model parameters, shows promise but does not scale to large batch sizes due to union of active neurons quickly approaching dense computation. We introduce Polar Sparsity, highlighting a key shift in sparsity importance from MLP to Attention layers as we scale batch size and sequence length. While MLP layers become more compute-efficient under batching, their sparsity vanishes. In contrast, attention becomes increasingly more expensive at scale, while their head sparsity remains stable and batch-invariant. We develop Selective Head Attention with hardware-efficient, sparsity-aware GPU kernels, delivering up to \(2.2\times\) end-to-end speedups for models like OPT, LLaMA-2 \& 3, Qwen, Mistral across various batch sizes and sequence lengths without compromising accuracy. To our knowledge, this is the first work to demonstrate that contextual sparsity can scale effectively to large batch sizes, delivering substantial inference acceleration with minimal changes, making Polar Sparsity practical for large-scale, high-throughput LLM deployment systems. Susav Shrestha, Bradley W. Settlemyer, Nikoli Dryden, Narasimha Reddy |
NeurIPS | 3 |
| 2025 | A General and Scalable GCN Training Framework on CPU SupercomputersabstractGraph Convolutional Networks (GCNs) are widely used in various domains. However, training distributed full-batch GCNs on large-scale graphs poses challenges due to inefficient memory access patterns and high communication overhead. This paper presents a general and efficient GCN training framework on CPU supercomputers. It comprises a general aggregation kernel designed to optimize irregular memory access and a quantization method with label propagation to reduce communication overhead. Experimental results show that our method achieves a speedup of up to 4.1× compared with the SoTA implementations. Chen Zhuang, Peng Chen 0035, Xin Liu 0020, Rio Yokota, Nikoli Dryden, Lingqi Zhang 0001, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
PPoPP | 5 |
| 2024 | Learning to Compose SuperWeights for Neural Parameter Allocation SearchabstractNeural parameter allocation search (NPAS) automates parameter sharing by obtaining weights for a network given an arbitrary, fixed parameter budget. Prior work has two major drawbacks we aim to address. First, there is a disconnect in the sharing pattern between the search and training steps, where weights are warped for layers of different sizes during the search to measure similarity, but not during training, resulting in reduced performance. To address this, we generate layer weights by learning to compose sets of SuperWeights, which represent a group of trainable parameters. These SuperWeights are created to be large enough so they can be used to represent any layer in the network, but small enough that they are computationally efficient. The second drawback we address is the method of measuring similarity between shared parameters. Whereas prior work compared the weights themselves, we argue this does not take into account the amount of conflict between the shared weights. Instead, we use gradient information to identify layers with shared weights that wish to diverge from each other. We demonstrate that our SuperWeight Networks consistently boost performance over the state-of-the-art on the ImageNet and CIFAR datasets in the NPAS setting. We further show that our approach can generate parameters for many network architectures using the same set of weights. This enables us to support tasks like efficient ensembling and anytime prediction, outperforming fully-parameterized ensembles with 17% fewer parameters1. Piotr Teterwak, Soren Nelson, Nikoli Dryden, Dina Bashkirova, Kate Saenko, Bryan A. Plummer |
WACV | 3 |
| 2022 | Neural Parameter Allocation Search
Bryan A. Plummer, Nikoli Dryden, Julius Frost, Torsten Hoefler, Kate Saenko |
ICLR | 2 |
| 2022 | A data-centric optimization framework for machine learningabstractRapid progress in deep learning is leading to a diverse set of quickly changing models, with a dramatically growing demand for compute. However, as frameworks specialize performance optimization to patterns in popular networks, they implicitly constrain novel and diverse models that drive progress in research. We empower deep learning researchers by defining a flexible and user-customizable pipeline for optimizing training of arbitrary deep neural networks, based on data movement minimization. The pipeline begins with standard networks in PyTorch or ONNX and transforms computation through progressive lowering. We define four levels of general-purpose transformations, from local intra-operator optimizations to global data movement reduction. These operate on a data-centric graph intermediate representation that expresses computation and data movement at all levels of abstraction, including expanding basic operators such as convolutions to their underlying computations. Central to the design is the interactive and introspectable nature of the pipeline. Every part is extensible through a Python API, and can be tuned interactively using a GUI. We demonstrate competitive performance or speedups on ten different networks, with interactive optimizations discovering new opportunities in EfficientNet. Oliver Rausch, Tal Ben-Nun, Nikoli Dryden, Andrei Ivanov, Shigang Li 0002, Torsten Hoefler |
ICS | 3 |
| 2022 | Motif Prediction with Graph Neural NetworksabstractLink prediction is one of the central problems in graph mining. However, recent studies highlight the importance of higher-order network analysis, where complex structures called motifs are the first-class citizens. We first show that existing link prediction schemes fail to effectively predict motifs. To alleviate this, we establish a general motif prediction problem and we propose several heuristics that assess the chances for a specified motif to appear. To make the scores realistic, our heuristics consider - among others - correlations between links, i.e., the potential impact of some arriving links on the appearance of other links in a given motif. Finally, for highest accuracy, we develop a graph neural network (GNN) architecture for motif prediction. Our architecture offers vertex features and sampling schemes that capture the rich structural properties of motifs. While our heuristics are fast and do not need any training, GNNs ensure highest accuracy of predicting motifs, both for dense (e.g., k-cliques) and for sparse ones (e.g., k-stars). We consistently outperform the best available competitor by more than 10% on average and up to 32% in area under the curve. Importantly, the advantages of our approach over schemes based on uncorrelated link prediction increase with the increasing motif size and complexity. We also successfully apply our architecture for predicting more arbitrary clusters and communities, illustrating its potential for graph mining beyond motif analysis. Maciej Besta, Raphael Grob, Cesare Miglioli, Nicola Bernold, Grzegorz Kwasniewski, Gabriel Gjini, Raghavendra Kanakagiri, Saleh Ashkboos, Lukas Gianinazzi, Nikoli Dryden, Torsten Hoefler |
KDD | 10 |
| 2022 | ENS-10: A Dataset For Post-Processing Ensemble Weather ForecastsabstractPost-processing ensemble prediction systems can improve the reliability of weather forecasting, especially for extreme event prediction. In recent years, different machine learning models have been developed to improve the quality of weather post-processing. However, these models require a comprehensive dataset of weather simulations to produce high-accuracy results, which comes at a high computational cost to generate. This paper introduces the ENS-10 dataset, consisting of ten ensemble members spanning 20 years (1998--2017). The ensemble members are generated by perturbing numerical weather simulations to capture the chaotic behavior of the Earth. To represent the three-dimensional state of the atmosphere, ENS-10 provides the most relevant atmospheric variables at 11 distinct pressure levels and the surface at \ang{0.5} resolution for forecast lead times T=0, 24, and 48 hours (two data points per week). We propose the ENS-10 prediction correction task for improving the forecast quality at a 48-hour lead time through ensemble post-processing. We provide a set of baselines and compare their skill at correcting the predictions of three important atmospheric variables. Moreover, we measure the baselines' skill at improving predictions of extreme weather events using our dataset. The ENS-10 dataset is available under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Saleh Ashkboos, Langwen Huang, Nikoli Dryden, Tal Ben-Nun, Peter D. Düben, Lukas Gianinazzi, Luca Kummer, Torsten Hoefler |
NeurIPS | 3 |
| 2022 | Spatial Mixture-of-ExpertsabstractMany data have an underlying dependence on spatial location; it may be weather on the Earth, a simulation on a mesh, or a registered image. Yet this feature is rarely taken advantage of, and violates common assumptions made by many neural network layers, such as translation equivariance. Further, many works that do incorporate locality fail to capture fine-grained structure. To address this, we introduce the Spatial Mixture-of-Experts (SMoE) layer, a sparsely-gated layer that learns spatial structure in the input domain and routes experts at a fine-grained level to utilize it. We also develop new techniques to train SMoEs, including a self-supervised routing loss and damping expert errors. Finally, we show strong results for SMoEs on numerous tasks, and set new state-of-the-art results for medium-range weather prediction and post-processing ensemble weather forecasts. Nikoli Dryden, Torsten Hoefler |
NeurIPS | 1 |
| 2021 | Clairvoyant prefetching for distributed machine learning I/OabstractI/O is emerging as a major bottleneck for machine learning training, especially in distributed environments. Indeed, at large scale, I/O takes as much as 85% of training time. Addressing this I/O bottleneck necessitates careful optimization, as optimal data ingestion pipelines differ between systems, and require a delicate balance between access to local storage, external filesystems, and remote nodes. We introduce NoPFS, a machine learning I/O middleware, which provides a scalable, flexible, and easy-to-use solution to the I/O bottleneck. NoPFS uses clairvoyance: Given the seed generating the random access pattern for training with SGD, it can exactly predict when and where a sample will be accessed. We combine this with an analysis of access patterns and a performance model to provide distributed caching policies that adapt to different datasets and storage hierarchies. NoPFS reduces I/O times and improves end-to-end training by up to 5.4× on the ImageNet-1k, ImageNet-22k, and CosmoFlow datasets. Nikoli Dryden, Roman Böhringer, Tal Ben-Nun, Torsten Hoefler |
SC | 1 |
| 2021 | Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networksabstractThe growing energy and performance costs of deep learning have driven the community to reduce the size of neural networks by selectively pruning components. Similarly to their biological counterparts, sparse networks generalize just as well, sometimes even better than, the original dense networks. Sparsity promises to reduce the memory footprint of regular networks to fit mobile devices, as well as shorten training time for ever growing networks. In this paper, we survey prior work on sparsity in deep learning and provide an extensive tutorial of sparsification for both inference and training. We describe approaches to remove and add elements of neural networks, different training strategies to achieve model sparsity, and mechanisms to exploit sparsity in practice. Our work distills ideas from more than 300 research papers and provides guidance to practitioners who wish to utilize sparsity today, as well as to researchers whose goal is to push the frontier forward. We include the necessary background on mathematical methods in sparsification, describe phenomena such as early structure adaptation, the intricate relations between sparsity and the training process, and show techniques for achieving acceleration on real hardware. We also define a metric of pruned parameter efficiency that could serve as a baseline for comparison of different sparse networks. We close by speculating on how sparsity can improve future workloads and outline major open problems in the field. Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, Alexandra Peste |
J. Mach. Learn. Res. | 4 |
| 2021 | Breaking (Global) Barriers in Parallel Stochastic Optimization With Wait-Avoiding Group AveragingabstractDeep learning at scale is dominated by communication time. Distributing samples across nodes usually yields the best performance, but poses scaling challenges due to global information dissemination and load imbalance across uneven sample lengths. State-of-the-art decentralized optimizers mitigate the problem, but require more iterations to achieve the same accuracy as their globally-communicating counterparts. We present Wait-Avoiding Group Model Averaging (WAGMA) SGD, a wait-avoiding stochastic optimizer that reduces global communication via subgroup weight exchange. The key insight is a combination of algorithmic changes to the averaging scheme and the use of a group allreduce operation. We prove the convergence of WAGMA-SGD, and empirically show that it retains convergence rates similar to Allreduce-SGD. For evaluation, we train ResNet-50 on ImageNet; Transformer for machine translation; and deep reinforcement learning for navigation at scale. Compared with state-of-the-art decentralized SGD variants, WAGMA-SGD significantly improves training throughput (e.g., 2.1× on 1,024 GPUs for reinforcement learning), and achieves the fastest time-to-solution (e.g., the highest score using the shortest training time for Transformer). Shigang Li 0002, Tal Ben-Nun, Giorgi Nadiradze, Salvatore Di Girolamo, Nikoli Dryden, Dan Alistarh, Torsten Hoefler |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs With Hybrid ParallelismabstractWe present scalable hybrid-parallel algorithms for training large-scale 3D convolutional neural networks. Deep learning-based emerging scientific workflows often require model training with large, high-dimensional samples, which can make training much more costly and even infeasible due to excessive memory usage. We solve these challenges by extensively applying hybrid parallelism throughout the end-to-end training pipeline, including both computations and I/O. Our hybrid-parallel algorithm extends the standard data parallelism with spatial parallelism, which partitions a single sample in the spatial domain, realizing strong scaling beyond the mini-batch dimension with a larger aggregated memory capacity. We evaluate our proposed training algorithms with two challenging 3D CNNs, CosmoFlow and 3D U-Net. Our comprehensive performance studies show that good weak and strong scaling can be achieved for both networks using up to 2K GPUs. More importantly, we enable training of CosmoFlow with much larger samples than previously possible, realizing an order-of-magnitude improvement in prediction accuracy. Yosuke Oyama, Naoya Maruyama, Nikoli Dryden, Erin McCarthy, Peter Harrington, Jan Balewski, Satoshi Matsuoka, Peter Nugent, Brian Van Essen |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Improving Strong-Scaling of CNN Training by Exploiting Finer-Grained ParallelismabstractScaling CNN training is necessary to keep up with growing datasets and reduce training time. We also see an emerging need to handle datasets with very large samples, where memory requirements for training are large. Existing training frameworks use a data-parallel approach that partitions samples within a mini-batch, but limits to scaling the minibatch size and memory consumption makes this untenable for large samples. We describe and implement new approaches to convolution, which parallelize using spatial decomposition or a combination of sample and spatial decomposition. This introduces many performance knobs for a network, so we develop a performance model for CNNs and present a method for using it to automatically determine efficient parallelization strategies. We evaluate our algorithms with microbenchmarks and image classification with ResNet-50. Our algorithms allow us to prototype a model for a mesh-tangling dataset, where sample sizes are very large. We show that our parallelization achieves excellent strong and weak scaling and enables training for previously unreachable datasets. Nikoli Dryden, Naoya Maruyama, Tom Benson, Tim Moon, Marc Snir, Brian Van Essen |
IPDPS | 1 |
| 2019 | Channel and filter parallelism for large-scale CNN trainingabstractAccelerating large-scale CNN training is needed to keep training times reasonable as datasets grow larger and models become more complex. Existing frameworks primarily scale using data-parallelism, but this is limited by the mini-batch size, which cannot grow arbitrarily. We introduce three algorithms that partition channel or filter data to exploit parallelism beyond the sample dimension. Further, they partition the parameters of convolutional layers, replacing global all reduces with segmented allreduces---smaller, concurrent allreduces among disjoint processor sets. These algorithms enable strong scaling, reduced communication overhead, and reduced memory pressure, enabling training of very wide CNNs. Nikoli Dryden, Naoya Maruyama, Tim Moon, Tom Benson, Marc Snir, Brian Van Essen |
SC | 1 |
| 2018 | Neural Network Based Silent Error DetectorabstractAs we move toward exascale platforms, silent data corruptions (SDC) are likely to occur more frequently. Such errors can lead to incorrect results. Attempts have been made to use generic algorithms to detect such errors. Such detectors have demonstrated high precision and recall for detecting errors, but only if they run immediately after an error has been injected. In this paper, we propose a neural network detector that can detect SDCs even multiple iterations after they were injected. We have evaluated our detector with 6 FLASH applications and 2 Mantevo mini-apps. Experiments show that our detector can detect more than 89% of SDCs with a false positive rate of less than 2%. Chen Wang 0004, Nikoli Dryden, Franck Cappello, Marc Snir |
CLUSTER | 2 |
| 2018 | A Lightweight Communication Runtime for Distributed Graph AnalyticsabstractDistributed-memory multi-core clusters enable in-memory processing of very large graphs with billions of nodes and edges. Recent distributed graph analytics systems have been built on top of MPI. However, communication in graph applications is very irregular, and each host exchanges different amounts of non-contiguous data with other hosts. MPI does not support such a communication pattern well, and it has limited ability to integrate communication with serialization, deserialization, and graph computation tasks. In this paper, we describe a lightweight communication runtime called LCI that supports a large number of threads on each host and avoids the semantic mismatches between the requirements of graph computations and the communication library in MPI. The implementation of LCI is informed by lessons learnt from two baseline MPI-based implementations. We have successfully integrated LCI with two state-of-the-art graph analytics systems - Gemini and Abelian. LCI improves the latency up to 3.5× for microbenchmarks compared to MPI solutions and improves the end-to-end performance of distributed graph algorithms by up to 2×. Hoang-Vu Dang, Roshan Dathathri, Gurbinder Gill, Alex Brooks, Nikoli Dryden, Andrew Lenharth, Loc Hoang, Keshav Pingali, Marc Snir |
IPDPS | 5 |
| 2018 | Gluon: a communication-optimizing substrate for distributed heterogeneous graph analyticsabstractThis paper introduces a new approach to building distributed-memory graph analytics systems that exploits heterogeneity in processor types (CPU and GPU), partitioning policies, and programming models. The key to this approach is Gluon, a communication-optimizing substrate. Roshan Dathathri, Gurbinder Gill, Loc Hoang, Hoang-Vu Dang, Alex Brooks, Nikoli Dryden, Marc Snir, Keshav Pingali |
PLDI | 6 |