VLDB 2026 Research / reviewers in the wild / expert
Armon Carigiet
dblp:273/4323
· DBLP profile ↗
3ranked-venue papers
0as first author
2since 2021 · last 2024
0009-0002-3555-767XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 71% GPUs and heterogeneous computing · 18% Parallel and multicore computing · 10% | |
| Theoretical computer science
2 papers |
Graph algorithms and graph theory · 54% Computational complexity · 46% | |
| Artificial intelligence
1 paper |
Graph learning · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › parallel numerical algorithms
communication-avoiding algorithms |
0.8 | 1 | 2024 | Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication · PPoPP 2024 |
GPUs and heterogeneous computing › GPU kernel
sparse-dense matrix multiplication |
0.8 | 1 | 2024 | Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication · PPoPP 2024 |
High-performance computing › sparse linear algebra
sparse matrix computation |
0.8 | 1 | 2024 | Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication · PPoPP 2024 |
High-performance computing › sparse linear algebra
sparse matrix multiplication |
0.8 | 1 | 2024 | Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication · PPoPP 2024 |
Computational complexity › algebraic complexity › matrix multiplication
sparse matrix multiplication |
0.8 | 1 | 2024 | Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication · PPoPP 2024 |
Machine learning › Graph learning › graph neural network › attention-based graph neural network
graph attention network |
0.7 | 1 | 2023 | High-Performance and Programmable Attentional Graph Neural Networks with Global Tensor Formulations · SC 2023 |
Machine learning › Graph learning
graph neural network |
0.7 | 1 | 2023 | High-Performance and Programmable Attentional Graph Neural Networks with Global Tensor Formulations · SC 2023 |
High-performance computing
performance optimization at scale |
0.7 | 1 | 2023 | High-Performance and Programmable Attentional Graph Neural Networks with Global Tensor Formulations · SC 2023 |
Parallel and multicore computing
parallel algorithms |
0.4 | 1 | 2020 | High-performance parallel graph coloring with strong guarantees on work, depth, and quality · SC 2020 |
Graph algorithms and graph theory
graph coloring |
0.4 | 1 | 2020 | High-performance parallel graph coloring with strong guarantees on work, depth, and quality · SC 2020 |
Graph algorithms and graph theory › graph coloring
parallel graph coloring |
0.4 | 1 | 2020 | High-performance parallel graph coloring with strong guarantees on work, depth, and quality · SC 2020 |
Methods — techniques the papers use, named apart from their topics
permutation · 1.5arrow matrix decomposition · 1.5tensor formulation · 1.3communication-minimizing routines · 1.3GraphBLAS · 1.3degeneracy ordering relaxation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix MultiplicationabstractWe propose a novel approach to iterated sparse matrix dense matrix multiplication, a fundamental computational kernel in scientific computing and graph neural network training. In cases where matrix sizes exceed the memory of a single compute node, data transfer becomes a bottleneck. An approach based on dense matrix multiplication algorithms leads to sub-optimal scalability and fails to exploit the sparsity in the problem. To address these challenges, we propose decomposing the sparse matrix into a small number of highly structured matrices called arrow matrices, which are connected by permutations. Our approach enables communication-avoiding multiplications, achieving a polynomial reduction in communication volume per iteration for matrices corresponding to planar graphs and other minor-excluded families of graphs. Our evaluation demonstrates that our approach outperforms a state-of-the-art method for sparse matrix multiplication on matrices with hundreds of millions of rows, offering near-linear strong and weak scaling. Lukas Gianinazzi, Alexandros Nikolaos Ziogas, Langwen Huang, Piotr Luczynski, Saleh Ashkboos, Florian Scheidl, Armon Carigiet, Chio Ge, Nabil Abubaker, Maciej Besta, Tal Ben-Nun, Torsten Hoefler |
PPoPP | 7 |
| 2023 | High-Performance and Programmable Attentional Graph Neural Networks with Global Tensor FormulationsabstractGraph attention models (A-GNNs), a type of Graph Neural Networks (GNNs), have been shown to be more powerful than simpler convolutional GNNs (C-GNNs). However, A-GNNs are more complex to program and difficult to scale. To address this, we develop a novel mathematical formulation, based on tensors that group all the feature vectors, targeting both training and inference of A-GNNs. The formulation enables straightforward adoption of communication-minimizing routines, it fosters optimizations such as vectorization, and it enables seamless integration with established linear algebra DSLs or libraries such as GraphBLAS. Our implementation uses a data redistribution scheme explicitly developed for sparse-dense tensor operations used heavily in GNNs, and fusing optimizations that further minimize memory usage and communication cost. We ensure theoretical asymptotic reductions in communicated data compared to the established message-passing GNN paradigm. Finally, we provide excellent scalability and speedups of even 4--5x over modern libraries such as Deep Graph Library. Maciej Besta, Pawel Renc, Robert Gerstenberger, Paolo Sylos Labini, Alexandros Nikolaos Ziogas, Tiancheng Chen, Lukas Gianinazzi, Florian Scheidl, Kalman Szenes, Armon Carigiet, Patrick Iff, Grzegorz Kwasniewski, Raghavendra Kanakagiri, Chio Ge, Sammy Jaeger, Jaroslaw Was, Flavio Vella, Torsten Hoefler |
SC | 10 |
| 2020 | High-performance parallel graph coloring with strong guarantees on work, depth, and qualityabstractWe develop the first parallel graph coloring heuristics with strong theoretical guarantees on work and depth and coloring quality. The key idea is to design a relaxation of the vertex degeneracy order, a well-known graph theory concept, and to color vertices in the order dictated by this relaxation. This introduces a tunable amount of parallelism into the degeneracy ordering that is otherwise hard to parallelize. This simple idea enables significant benefits in several key aspects of graph coloring. For example, one of our algorithms ensures polylogarithmic depth and a bound on the number of used colors that is superior to all other parallelizable schemes, while maintaining workefficiency. In addition to provable guarantees, the developed algorithms have competitive run-times for several real-world graphs, while almost always providing superior coloring quality. Our degeneracy ordering relaxation is of separate interest for algorithms outside the context of coloring. Maciej Besta, Armon Carigiet, Kacper Janda, Zur Vonarburg-Shmaria, Lukas Gianinazzi, Torsten Hoefler |
SC | 2 |