EDBT 2026 Demo / reviewers in the wild / expert
Mahdi Torabzadehkashi
dblp:224/1540
· DBLP profile ↗
5ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0002-7765-3064ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 61% Hardware accelerators and domain-specific architectures · 15% Memory systems · 15% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
computational storage |
0.9 | 2 | 2020 | Cost-effective, Energy-efficient, and Scalable Storage Computing for Large-scale AI Applications · ACM Trans. Storage 2020 Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020 |
Storage systems › computational storage
in-storage computing |
0.9 | 2 | 2020 | Cost-effective, Energy-efficient, and Scalable Storage Computing for Large-scale AI Applications · ACM Trans. Storage 2020 Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2020 | Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020 |
Memory systems › processing-in-memory
near-data processing |
0.4 | 1 | 2020 | Cost-effective, Energy-efficient, and Scalable Storage Computing for Large-scale AI Applications · ACM Trans. Storage 2020 |
Distributed systems › distributed machine learning
distributed training |
0.1 | 1 | 2020 | Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.1 | 1 | 2020 | Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020 |
Methods — techniques the papers use, named apart from their topics
distributed near-data processing · 0.9cost analysis · 0.9distributed training · 0.4batch size scheduling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Leveraging Computational Storage for Power-Efficient Distributed Data AnalyticsabstractThis article presents a family of computational storage drives (CSDs) and demonstrates their performance and power improvements due to in-storage processing (ISP) when running big data analytics applications. CSDs are an emerging class of solid state drives that are capable of running user code while minimizing data transfer time and energy. Applications that can benefit from in situ processing include distributed training, distributed inferencing, and databases. To achieve the full advantage of the proposed ISP architecture, we propose software solutions for workload balancing before and at runtime for training and inferencing applications. Other applications such as sharding-based databases can readily take advantage of our ISP structure without additional tooling. Experimental results on different capacity and form factors of CSDs show up to 3.1× speedup in processing while reducing the energy consumption and data transfer by up to 67% and 68%, respectively, compared to regular enterprise solid state drives. Ali Heydarigorji, Siavash Rezaei, Mahdi Torabzadehkashi, Hossein Bobarshad, Vladimir Castro Alves, Pai H. Chou |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Stannis: Low-Power Acceleration of DNN Training Using Computational Storage DevicesabstractComputational storage devices enable in-storage processing of data in place. These devices contain 64-bit application processors and hardware accelerators that can help improving performance and saving power by reducing or eliminating data movement between host computers and storage units. This paper proposes a framework, named Stannis, for distributed in-storage training of deep neural networks on clusters of computational storage devices. This in-storage processing style of training ensures that private data never leaves the storage while fully controlling the public sharing of data. The Stannis framework distributes the workload based on the processing power of each worker by determining the proper batch size for each node. Stannis also ensures the availability of input data for all nodes to avoid rank stall while maximizing the utilization and overall processing speed. Experimental results show up to 2.7x speedup and 69% reduction in energy consumption with no significant loss in accuracy. Ali Heydarigorji, Mahdi Torabzadehkashi, Siavash Rezaei, Hossein Bobarshad, Vladimir Castro Alves, Pai H. Chou |
DAC | 2 |
| 2020 | HyperTune: Dynamic Hyperparameter Tuning for Efficient Distribution of DNN Training Over Heterogeneous SystemsabstractDistributed training is a novel approach to accelerating training of Deep Neural Networks (DNN), but common training libraries fall short of addressing the distributed nature of heterogeneous processors or interruption by other workloads on the shared processing nodes. This paper describes distributed training of DNN on computational storage devices (CSD), which are NAND flash-based, high-capacity data storage with internal processing engines. A CSD-based distributed architecture incorporates the advantages of federated learning in terms of performance scalability, resiliency, and data privacy by eliminating the unnecessary data movement between the storage device and the host processor. The paper also describes Stannis, a DNN training framework that improves on the shortcomings of existing distributed training frameworks by dynamically tuning the training hyperparameters in heterogeneous systems to maintain the maximum overall processing speed in term of processed images per second and energy efficiency. Experimental results on image classification training benchmarks show up to 3.1x improvement in performance and 2.45x reduction in energy consumption when using Stannis plus CSD compare to the generic systems. Ali Heydarigorji, Siavash Rezaei, Mahdi Torabzadehkashi, Hossein Bobarshad, Vladimir Castro Alves, Pai H. Chou |
ICCAD | 3 |
| 2020 | Cost-effective, Energy-efficient, and Scalable Storage Computing for Large-scale AI ApplicationsabstractThe growing volume of data produced continuously in the Cloud and at the Edge poses significant challenges for large-scale AI applications to extract and learn useful information from the data in a timely and efficient way. The goal of this article is to explore the use of computational storage to address such challenges by distributed near-data processing. We describe Newport, a high-performance and energy-efficient computational storage developed for realizing the full potential of in-storage processing. To the best of our knowledge, Newport is the first commodity SSD that can be configured to run a server-like operating system, greatly minimizing the effort for creating and maintaining applications running inside the storage. We analyze the benefits of using Newport by running complex AI applications such as image similarity search and object tracking on a large visual dataset. The results demonstrate that data-intensive AI workloads can be efficiently parallelized and offloaded, even to a small set of Newport drives with significant performance gains and energy savings. In addition, we introduce a comprehensive taxonomy of existing computational storage solutions together with a realistic cost analysis for high-volume production, giving a good big picture of the economic feasibility of the computational storage technology. Jaeyoung Do, Victor da Cruz Ferreira, Hossein Bobarshad, Mahdi Torabzadehkashi, Siavash Rezaei, Ali Heydarigorji, Diego Fonseca Pereira de Souza, Brunno F. Goldstein, Leandro Santiago de Araújo, Min Soo Kim 0009, Priscila M. V. Lima, Felipe M. G. França, Vladimir Castro Alves |
ACM Trans. Storage | 4 |
| 2019 | Catalina: In-Storage Processing Acceleration for Scalable Big Data AnalyticsabstractCloud applications are increasingly playing a crucial role in big data analytics. New use cases such as autonomous cars and edge computing call for novel approaches mixing heterogeneous computing and machine learning. These applications typically process petabyte-scale datasets, therefore, requiring low-power and scalable storage providing low-latency and high-throughput data access. While data centers have been focusing on migrating from legacy HDDs and SATA SSDs by deploying high-throughput and low-latency NVMe SSDs, the data bottlenecks appear as capacity scales. One approach to tackle this problem is to enable processing to happen within the storage device -in-storage processing (ISP)- eliminating the need to move the data. In this paper, we investigated the deployment of storage units with embedded low-power application processors along with FPGA-based reconfigurable hardware accelerators to address both performance and energy efficiency. To this purpose, we developed a high-capacity solid-state drive (SSD) named Catalina equipped with a quad-core ARM A53 processor running a Linux operating system along with a highly efficient FPGA accelerator for running applications in-place. We evaluated our proposed approach on a case study application for a similarity search library called Faiss. Mahdi Torabzadehkashi, Siavash Rezaei, Ali Heydarigorji, Hossein Bobarshad, Vladimir Castro Alves, Nader Bagherzadeh |
PDP | 1 |