Saeed Soori

dblp:207/8337 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 67% Optimization for machine learning · 33%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 33% Parallel and multicore computing · 33% Memory systems · 33%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
0.612022
HyLo: A Hybrid Low-Rank Natural Gradient Descent Method · SC 2022
Machine learning › Efficient and distributed learning › model compression
low-rank approximation
0.612022
HyLo: A Hybrid Low-Rank Natural Gradient Descent Method · SC 2022
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
natural gradient descent
0.612022
HyLo: A Hybrid Low-Rank Natural Gradient Descent Method · SC 2022
Memory systems
cache
0.412020
MatRox: modular approach for improving data locality in hierarchical (Mat)rix App(Rox)imation · PPoPP 2020
Parallel and multicore computing › locality optimization
data locality optimization
0.412020
MatRox: modular approach for improving data locality in hierarchical (Mat)rix App(Rox)imation · PPoPP 2020
High-performance computing
performance optimization at scale
0.412020
MatRox: modular approach for improving data locality in hierarchical (Mat)rix App(Rox)imation · PPoPP 2020
Machine learning and data management
data management for machine learning
0.112020
MatRox: modular approach for improving data locality in hierarchical (Mat)rix App(Rox)imation · PPoPP 2020

Methods — techniques the papers use, named apart from their topics

structure analysis · 0.9storage format · 0.9code specialization · 0.9sherman-morrison-woodbury · 0.6kronecker factorization · 0.6fisher information matrix · 0.6
YearPublicationVenuePosition
2024 L-DATR: A Limited-Memory Distributed Asynchronous Trust-Region Method
abstract
A distributed approach is proposed in this work to solve large-scale optimization problems, called L-DATR, under the master/worker communication model. L-DATR is a distributed limited-memory trust-region method that allows worker nodes to perform asynchronous computations. Our method dynamically adjusts the step size and direction using trust-region strategies to improve stability and convergence. To our knowledge, this is the first implementation of a distributed trust-region limited memory quasi-Newton method with robust handling of asynchronous updates and non-uniform delays between nodes. Our method is communication-efficient because it communicates only vectors of the dimension of the decision variable. Our numerical experiments match our theoretical results and showcase significant stability improvements compared to state-of-the-art distributed algorithms.
Mohammad Jalali, Saeed Soori, Hadis Barati
ICIS2
2022 HyLo: A Hybrid Low-Rank Natural Gradient Descent Method
abstract
This work presents a Hybrid Low-Rank Natural Gradient Descent method, called HyLo, that accelerates the training time of deep neural networks. Natural gradient descent (NGD) requires computing the inverse of the Fisher information matrix (FIM), which is typically expensive at largescale. Kronecker factorization methods such as KFAC attempt to improve NGD's running time by approximating the FIM with Kronecker factors. However, the size of Kronecker factors increases quadratically as the model size grows. Instead, in HyLo, we use the Sherman-Morrison-Woodbury variant of NGD (SNGD) and propose a reformulation of SNGD to resolve its scalability issues. HyLo uses a computationally-efficient low-rank factorization to achieve superior timing for Fisher inverses. We evaluate HyLo on large models including ResNet-50, U-Net, and ResNet-32 on up to 64 GPUs. HyLo converges 1.4×-2.1× faster than the state-of-the-art distributed implementation of KFAC and reduces the computation and communication time up to 350 × and 10.7× on ResNet-50.
Baorun Mu, Saeed Soori, Bugra Can, Mert Gürbüzbalaban, Maryam Mehri Dehnavi
SC2
2020 DAve-QN: A Distributed Averaged Quasi-Newton Method with Local Superlinear Convergence Rate
abstract
In this paper, we consider distributed algorithms for solving the empirical risk minimization problem under the master/worker communication model. We develop a distributed asynchronous quasi-Newton algorithm that can achieve superlinear convergence. To our knowledge, this is the first distributed asynchronous algorithm with superlinear convergence guarantees. Our algorithm is communication-efficient in the sense that at every iteration the master node and workers communicate vectors of size $O(p)$, where $p$ is the dimension of the decision variable. The proposed method is based on a distributed asynchronous averaging scheme of decision vectors and gradients in a way to effectively capture the local Hessian information of the objective function. Our convergence theory supports asynchronous computations subject to both bounded delays and unbounded delays with a bounded time-average. Unlike in the majority of asynchronous optimization literature, we do not require choosing smaller stepsize when delays are huge. We provide numerical experiments that match our theoretical results and showcase significant improvement comparing to state-of-the-art distributed algorithms.
Saeed Soori, Konstantin Mishchenko, Aryan Mokhtari, Maryam Mehri Dehnavi, Mert Gürbüzbalaban
AISTATS1
2020 ASYNC: A Cloud Engine with Asynchrony and History for Distributed Machine Learning
abstract
ASYNC is a framework that supports the implementation of asynchrony and history for optimization methods on distributed computing platforms. The popularity of asynchronous optimization methods has increased in distributed machine learning. However, their applicability and practical experimentation on distributed systems are limited because current bulk-processing cloud engines do not provide a robust support for asynchrony and history. With introducing three main modules and bookkeeping system-specific and application parameters, ASYNC provides practitioners with a framework to implement asynchronous machine learning methods. To demonstrate ease-of-implementation in ASYNC, the synchronous and asynchronous variants of two well-known optimization methods, stochastic gradient descent and SAGA, are demonstrated in ASYNC.
Saeed Soori, Bugra Can, Mert Gürbüzbalaban, Maryam Mehri Dehnavi
IPDPS1
2020 MatRox: modular approach for improving data locality in hierarchical (Mat)rix App(Rox)imation
abstract
Hierarchical matrix approximations have gained significant traction in the machine learning and scientific community as they exploit available low-rank structures in kernel methods to compress the kernel matrix. The resulting compressed matrix, HMatrix, is used to reduce the computational complexity of operations such as HMatrix-matrix multiplications with tuneable accuracy in an evaluation phase. Existing implementations of HMatrix evaluations do not preserve locality and often lead to unbalanced parallel execution with high synchronization. Also, current solutions require the compression phase to re-execute if the kernel method or the required accuracy change. MatRox is a framework that uses novel structure analysis strategies with code specialization and a storage format to improve locality and create load-balanced parallel tasks for HMatrix-matrix multiplications. Modularization of the matrix compression phase enables the reuse of computations when there are changes to the input accuracy and the kernel function. The MatRox-generated code for matrix-matrix multiplication is 2.98X, 1.60X, and 5.98X faster than library implementations available in GOFMM, SMASH, and STRUMPACK respectively. Additionally, the ability to reuse portions of the compression computation for changes to the accuracy leads to up to 2.64X improvement with MatRox over five changes to accuracy using GOFMM.
Bangtian Liu, Kazem Cheshmi, Saeed Soori, Michelle Mills Strout, Maryam Mehri Dehnavi
PPoPP3
2018 Reducing Communication in Proximal Newton Methods for Sparse Least Squares Problems
abstract
Proximal Newton methods are iterative algorithms that solve l1-regularized least squares problems. Distributed-memory implementation of these methods have become popular since they enable the analysis of large-scale machine learning problems. However, the scalability of these methods is limited by the communication overhead on modern distributed architecture. We propose a stochastic variance-reduced proximal method along with iteration-overlapping and Hessian-reuse to find an efficient trade-off between computation complexity and data communication. The proposed RC-SFSITA algorithm reduces latency costs by a factor of k without altering bandwidth costs. RC-SFISTA is implemented on both MPI and Spark and compared to the state-of-the-art framework, ProxCoCoA. The performance of RC-SFISTA is evaluated on 1 to 512 nodes for multiple benchmarks and demonstrates speedups of up to 12× compared to ProxCoCoA with scaling properties that outperform the original algorithm.
Saeed Soori, Aditya Devarakonda, Zachary Blanco, James Demmel, Mert Gürbüzbalaban, Maryam Mehri Dehnavi
ICPP1