EDBT 2026 Demo / reviewers in the wild / expert
Mohammadreza Soltaniyeh
dblp:139/8299
· DBLP profile ↗
8ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-2608-7322ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COMETS: Cost-effective Multi-node Efficient Training System with Memory Pooling and Sharing
Hanqiu Chen, Shao-Peng Yang 0001, Mohammadreza Soltaniyeh, Shuyi Pei, Bryan S. Kim, Cong Hao |
ICS | 3 |
| 2025 | Revisiting Memory Hierarchies with CMM-H: Use Device-side Caching to Integrate DRAM and SSD for a Hybrid CXL MemoryabstractEmerging data-intensive applications increasingly demand large-scale, cost-effective, and high-performance memory solutions. Samsung's CXL Memory Module-Hybrid (CMM-H) uniquely integrates DDR DRAM and NAND flash storage within a single CXL-attached memory device for higher capacity while keeping still high performance. This paper provides a comprehensive exploration of the CMM-H module, detailing its architectural design, operational workflow, and caching mechanisms. We evaluate the performance of CMM-H through extensive experiments, highlighting its benefits and limitations. Our study contributes insights into hybrid CXL memory architectures and provides valuable guidance for future CXL memory design improvements. Mohammadreza Soltaniyeh, Gongjin Sun, Xuebin Yao, Amir Beygi, Ramdas Kachare, Dongwan Zhao, Hingkwan Huen, Senthil Murugesapandian, Caroline Kahn |
HotStorage | 1 |
| 2024 | ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture ModelabstractCompute Express Link (CXL) emerges as a solution for wide gap between computational speed and data communication rates among host and multiple devices. It fosters a unified and coherent memory space between host and CXL storage devices such as such as Solid-state drive (SSD) for memory expansion, with a corresponding DRAM implemented as the device cache. However, this introduces challenges such as substantial cache miss penalties, sub-optimal caching due to data access granularity mismatch between the DRAM "cache" and SSD "memory", and inefficient hardware cache management. To address these issues, we propose a novel solution, named ICGMM, which optimizes caching and eviction directly on hardware, employing a Gaussian Mixture Model (GMM)-based approach. We prototype our solution on an FPGA board, which demonstrates a noteworthy improvement compared to the classic Least Recently Used (LRU) cache strategy. We observe a decrease in the cache miss rate ranging from 0.32% to 6.14%, leading to a substantial 16.23% to 39.14% reduction in the average SSD access latency. Furthermore, when compared to the state-of-the-art Long Short-Term Memory (LSTM)-based cache policies, our GMM algorithm on FPGA showcases an impressive latency reduction of over 10,000 times. Remarkably, this is achieved while demanding much fewer hardware resources. Hanqiu Chen, Yitu Wang, Vitorio Cargnini, Mohammadreza Soltaniyeh, Gongjin Sun, Pradeep Subedi, Yiran Chen 0001, Cong Hao |
DAC | 4 |
| 2022 | Near-Storage Processing for Solid State Drive Based Recommendation Inference with SmartSSDs®abstractDeep learning-based recommendation systems are extensively deployed in numerous internet services, including social media, entertainment services, and search engines, to provide users with the most relevant and personalized content. Production scale deep learning models consist of large embedding tables with billions of parameters. DRAM-based recommendation systems incur a high infrastructure cost and limit the size of the deployed models. Recommendation systems based on solid-state drives (SSDs) are a promising alternative for DRAM-based systems. Systems based on SSDs can offer ample storage required for deep learning models with large embedding tables. This paper proposes SmartRec, an inference engine for deep learning-based recommendation systems that utilizes Samsung SmartSSD, an SSD with an on-board FPGA that can process data in-situ. We evaluate SmartRec with state-of-the-art recommendation models from Facebook and compare its performance and energy efficiency to a DRAM-based system on a CPU. We show SmartRec improves the energy efficiency of the recommendation inference task up to 10x in comparison to the baseline CPU implementation. In addition, we propose a novel application-specific caching system for SmartSSDs that allows the kernel on the FPGA to use its DRAM as a cache to minimize high latency SSD accesses. Finally, we demonstrate the scalability of our design by offloading the computation to multiple SmartSSDs to further improve performance. Mohammadreza Soltaniyeh, Veronica Lagrange Moutinho dos Reis, Matthew Bryson, Xuebin Yao, Richard P. Martin, Santosh Nagarakatte |
ICPE | 1 |
| 2022 | An Accelerator for Sparse Convolutional Neural Networks Leveraging Systolic General Matrix-matrix MultiplicationabstractThis article proposes a novel hardware accelerator for the inference task with sparse convolutional neural networks (CNNs) by building a hardware unit to perform Image to Column ( Im2Col ) transformation of the input feature map coupled with a systolic-array-based general matrix-matrix multiplication (GEMM) unit. Our design carefully overlaps the Im2Col transformation with the GEMM computation to maximize parallelism. We propose a novel design for the Im2Col unit that uses a set of distributed local memories connected by a ring network, which improves energy efficiency and latency by streaming the input feature map only once. The systolic-array-based GEMM unit in the accelerator can be dynamically configured as multiple GEMM units with square-shaped systolic arrays or as a single GEMM unit with a tall systolic array. This dynamic reconfigurability enables effective pipelining of Im2Col and GEMM operations and attains high processing element utilization for a wide range of CNNs. Further, our accelerator is sparsity aware, improving performance and energy efficiency by effectively mapping the sparse feature maps and weights to the processing elements, skipping ineffectual operations and unnecessary data movements involving zeros. Our prototype, SPOTS, is on average 2.16 \( \times \) , 1.74 \( \times \) , and 1.63 \( \times \) faster than Gemmini, Eyeriss, and Sparse-PE, which are prior hardware accelerators for dense and sparse CNNs, respectively. SPOTS is also 78 \( \times \) and 12 \( \times \) more energy-efficient when compared to CPU and GPU implementations, respectively. Mohammadreza Soltaniyeh, Richard P. Martin, Santosh Nagarakatte |
ACM Trans. Archit. Code Optim. | 1 |
| 2021 | Near-Storage Acceleration of Database Query Processing with SmartSSDsabstractSmart solid-state drives (SmartSSDs) with onboard FPGAs are becoming mainstream, providing opportunities for near-storage computation, which is appealing for increasing the performance of data-intensive workloads such as database query processing. This paper demonstrates the performance and energy improvements by offloading the filter and aggregation operations to the FPGA on a SmartSSD with real-system experiments. We make the observation that efficiently handling null entries in the data is important. Hence, we propose a novel design to manage the metadata to handle null entries. Our real system evaluation shows that offloading the query operations to the FPGA results in 9.6× improvement in performance while consuming 10.9× less energy compared to a conventional CPU-only query processing. Mohammadreza Soltaniyeh, Veronica Lagrange Moutinho dos Reis, Matthew Bryson, Richard P. Martin, Santosh Nagarakatte |
FCCM | 1 |
| 2018 | Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPsabstractAs shown in some prior studies, a significant percentage of data blocks accessed in parallel codes are private, and not keeping track of those blocks can improve the effectiveness of directory structures in Chip multiprocessors (CMPs). In this paper, we have two major contributions. First, we showed that compared to the classification of cache blocks at page granularity, data block classification (DBC) at subpage level helps to detect considerably more private data blocks. Based on this idea, we propose two different approaches for enhancing the effectiveness of directory caches in tiled CMPs. In the first approach, which is called quasi-dynamic subpage level DBC (QDBC), a data block is assumed to be private from the beginning of the program execution and stays private as long as the corresponding subpage is accessed by only one core. Our second approach, which is called dynamic subpage level DBC, turns a data block into private again after all blocks within the corresponding subpage are evicted from private cache hierarchy. Memory block classification at subpage level, however, may increase the frequency of the operating system involvement in updating the maintenance bits in page table entries. To overcome this, we propose, as a second contribution, a distributed table called as on-chip page table (o-CPT), which stores recently accessed page translations in the system. Our simulation results show that, compared to page level data classification, QDBC and DBC approaches relying on the o-CPT can detect significantly more private data blocks and considerably improve system performance. Mohammadreza Soltaniyeh, Ismail Kadayif, Ozcan Ozturk 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Quota setting router architecture for quality of service in GALS NoCabstractNetwork on Chip (NoC) is a new communication paradigm for emerging multi- and many-core architectures. Despite major benefits, like scalability and power efficiency, it suffers from lack of guaranteed bounded latency. Many contemporary applications, like multimedia and real-time applications, require such a guarantee. The growth of these applications in embedded systems emphasizes the need for guaranteed services in NoCs. Additionally, increasing numbers of cores in NoCs highlights the clock distribution issue. Globally asynchronous locally synchronous (GALS) NoC architectures propose to solve this issue through using asynchronous routers to connect synchronous blocks. This paper presents a novel approach for guaranteed service in a GALS NoC by using router with set port quota. We propose a novel router architecture which facilitates guaranteed latency for accessing shared media. Our simulations show up to 39% improvement in latency, with a negligible (up to 5%) power overhead. Kazem Cheshmi, Mohammadreza Soltaniyeh, Siamak Mohammadi, Jelena Trajkovic |
RSP | 2 |