Mohammadamin Ajdari

dblp:173/9805 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0003-1639-9303ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Characterization and Performance Analysis of Aging Effect on Enterprise HDDs
abstract
Hard Disk Drives (HDDs) are still a popular choice in data centers for cloud storage, due to lowest cost per terabyte and decent performance. Previous works have studied parameters related to disk failure and also proposed prediction models to predict the time that a disk fails. However, the impact of disk aging on performance has remained a big mystery. In this paper, we empirically study over 30 enterprise HDDs and analyze how different parameters related to aging (e.g., the amount of previous disk writes and power-on hours) affect the performance. We specifically provide six findings, summarized as four major insights. For example, we reveal that a new disk, after disk being used for over eight moths, shows up to $3.9 \times$ higher write latency. At the same time, no latency difference exists between disks older than two-years of age. Overall, a single used disk (with two-years of age) alongside three new disks in a RAID array worsen the I/O latency by 17%. Our findings can help storage architects and admins to (1) ensure high service-level-agreements for HDDbased storage systems, (2) detect used enterprise disks compared to fairly new disks even when disk S.M.A.R.T information are not available or have been altered deliberately to cause misleading information.
Ava Dezhban, Mohammadamin Ajdari, Hossein Asadi 0001
ISPASS2
2026 Do Not Get Fooled: Use True Device Utilization on High-Performance Storage Devices
abstract
Maximizing utilization of server components in data centers (through colocation of applications), while guaranteeing Quality of Service (QoS) for every application is of the major goals in enterprise environments. To ensure proper QoS, tools that measure a device utilization play a major role. Such tools report the capacity (i.e., number and type of applications) that may run on each device in real-time. We observe that the publicly-available tools to measure storage device utilization are very limited, and their utilization measurement logic is designed assuming traditional storage devices (with limited in-flight I/O requests). We reveal that the latest version of these tools such as Linux iostat have significant disparity with real utilization of emerging high-performance devices such as Persistent Memory or modern NVMe SSDs.In this paper, we first propose True Utilization metric for storage devices, which is the ratio of real-time I/O per Seconds (IOPS) to its saturation IOPS under same workload patterns. Second, we propose a framework that (a) explores different linear and non-linear models to predict saturation IOPS of each storage device under different workload patterns, (b) combines real-time IOPS measurements with our trained model outputs to estimate True Utilization accurately. We evaluate our proposed framework on different types of SSDs and persistent memories, and show it can accurately calculate True Utilization in real-time, and only have a negligible one-time training overhead.
Ali Sedaghatgoo, Mohammadamin Ajdari, Erfan Teymouri, Hossein Asadi 0001
ISPASS2
2023 Re-architecting I/O Caches for Emerging Fast Storage Devices
abstract
I/O caching has widely been used in enterprise storage systems to enhance the system performance with minimal cost. Using Solid-State Drives (SSDs) as an I/O caching layer on the top of arrays of Hard Disk Drives (HDDs) has been well studied in numerous studies. With emergence of ultra fast storage devices, recent studies suggest to use them as an I/O cache layer on top of mainstream SSDs in I/O intensive applications. Our detailed analysis shows despite significant potential of ultra-fast storage devices, existing I/O cache architectures may act as a major performance bottleneck in enterprise storage systems, which prevents to take advantage of the device full performance potentials.
Mohammadamin Ajdari, Pouria Peykani Sani, Amirhossein Moradi, Masoud Khanalizadeh Imani, Amir Hossein Bazkhanei, Hossein Asadi 0001
ASPLOS (3)1
2019 CIDR: A Cost-Effective In-Line Data Reduction System for Terabit-Per-Second Scale SSD Arrays
abstract
An SSD array, a storage system consisting of multiple SSDs per node, has become a design choice to implement a fast primary storage system, and modern storage architects now aim to achieve terabit-per-second scale performance with the next-generation SSD array. To reduce the storage cost and improve the device endurability, such SSD array must employ data reduction schemes (i.e., deduplication, compression), which provide high data reduction capability at minimum costs. However, existing data reduction schemes do not scale with the fast increasing performance of an SSD array, due to inhibitive amount of CPU resources (e.g., in software-based schemes) or low data reduction ratio (e.g., in SSD device wide deduplication) or being cost ineffective to address workload changes in datacenters (e.g., in ASIC-based acceleration). In this paper, we propose CIDR, a novel FPGA-based, cost-effective data reduction system for an SSD array to achieve the terabit-per-second scale storage performance. Our key ideas are as follows. First, we decouple data reduction related computing tasks from the unscalable host CPUs by offloading them to a scalable array of FPGA boards. Second, we employ a centralized, node-wide metadata management scheme to achieve an SSD array-wide, high data reduction. Third, our FPGA-based reconfiguration adapts to different workload patterns by dynamically balancing the amount of software and hardware tasks running on CPUs and FPGAs, respectively. For evaluation, we built our example CIDR prototype achieving up to 12.8 GB/s (0.1 Tbps) on one FPGA. CIDR outperforms the baseline for a write-only workload by up to 2.47x and a mixed read-write workload by an expected 3.2x, respectively. We showed CIDR's scalability to achieve Tbps-scale performance by measuring a two-FPGA CIDR and projecting the performance impacts for more FPGAs.
Mohammadamin Ajdari, Pyeongsu Park, Joonsung Kim 0001, Dongup Kwon, Jangwoo Kim
HPCA1
2019 FIDR: A Scalable Storage System for Fine-Grain Inline Data Reduction with Efficient Memory Handling
abstract
Storage systems play a critical role in modern servers which run highly data-intensive applications. To satisfy the high performance and capacity demands of such applications, storage systems now deploy an array of fast SSDs per server. To reduce the storage cost of employing many SSDs per server, storage systems actively perform inline data reduction (e.g., data deduplication, compression). Existing inline data reduction studies can achieve high performance and scalability by offloading computation-intensive data-reduction operations to dedicated hardware accelerators. However, such existing studies suffer from limited workload support and scalability. For example, they reduce only large data blocks, which incur many IO requests, leading to low data reduction rates, and their offloading overlooks memory-intensive operations, leading to the unoptimal scalability.
Mohammadamin Ajdari, Wonsik Lee, Pyeongsu Park, Joonsung Kim 0001, Jangwoo Kim
MICRO1
2018 DCS-ctrl: A Fast and Flexible Device-Control Mechanism for Device-Centric Server Architecture
abstract
Modern high-performance servers leverage a large number of emerging peripheral devices (e.g., data processing accelerators, non-volatile memory storage, high-bandwidth network cards) to meet ever-increasing performance demands of server applications. However, as such servers experience severe kernel overhead due to frequently invoked device operations (e.g., buffer management and data copy), server architects have proposed various hardware and software approaches to enable direct communications among the devices. Unfortunately, existing direct device-to-device (D2D) communication schemes still suffer from low performance and the lack of flexibility. First, software-based schemes depend on complicated kernel routines and necessitate multiple hardware-software and user-kernel boundary crossings, which significantly limit the performance improvement opportunities from direct D2D communications. On the other hand, hardware-based schemes require tight integration and custom-built devices, preventing architects from flexibly adding off-the-shelf devices. In this paper, we propose DCS-ctrl, a novel Hardware-based Device-Control (HDC) mechanism for Device-Centric Server (DCS) architecture to provide fast and CPU-efficient direct D2D communications among a large number of off-the-shelf peripheral devices. The key idea of DCS-ctrl is to implement a low-cost and flexible device-control mechanism on an independent FPGA device called HDC Engine. As HDC Engine manages all data and control transfers among devices at the hardware level, the server achieves high performance, scalability, and flexibility. First, optimizing both data and control paths at the hardware level minimizes the latency of inter-device communications. Second, implementing FPGA-based reconfigurable device controllers enables direct D2D communications among commodity devices and thus improves per-device flexibility. Third, merging heterogeneous device operations with intermediate data processing supports creates more opportunities for direct inter-device communications in server applications. Our DCS-ctrl prototype reduces the latency of software-based direct D2D communications by 42% and the CPU utilization by 52%.
Dongup Kwon, Jaehyung Ahn, Dongju Chae, Mohammadamin Ajdari, Suheon Bae, Youngsok Kim, Jangwoo Kim
ISCA4
2018 DiagSim: Systematically Diagnosing Simulators for Healthy Simulations
abstract
Simulators are the most popular and useful tool to study computer architecture and examine new ideas. However, modern simulators have become prohibitively complex (e.g., 200K+ lines of code) to fully understand and utilize. Users therefore end up analyzing and modifying only the modules of interest (e.g., branch predictor, register file) when performing simulations. Unfortunately, hidden details and inter-module interactions of simulators create discrepancies between the expected and actual module behaviors. Consequently, the effect of modifying the target module may be amplified or masked and the users get inaccurate insights from expensive simulations. In this article, we propose DiagSim, an efficient and systematic method to diagnose simulators. It ensures the target modules behave as expected to perform simulation in a healthy (i.e., accurate and correct) way. DiagSim is efficient in that it quickly pinpoints the modules showing discrepancies and guides the users to inspect the behavior without investigating the whole simulator. DiagSim is systematic in that it hierarchically tests the modules to guarantee the integrity of individual diagnosis and always provide reliable results. We construct DiagSim based on generic category-based diagnosis ideas to encourage easy expansion of the diagnosis. We diagnose three popular open source simulators and discover hidden details including implicitly reserved resources, un-documented latency factors, and hard-coded module parameter values. We observe that these factors have large performance impacts (up to 156%) and illustrate that our diagnosis can correctly detect and eliminate them.
Jae-Eon Jo, Gyu-hyeon Lee, Hanhwi Jang, Mohammadamin Ajdari, Jangwoo Kim
ACM Trans. Archit. Code Optim.5
2015 DCS: a fast and scalable device-centric server architecture
abstract
Conventional servers have achieved high performance by employing fast CPUs to run compute-intensive workloads, while making operating systems manage relatively slow I/O devices through memory accesses and interrupts. However, as the emerging workloads are becoming heavily data-intensive and the emerging devices (e.g., NVM storage, high-bandwidth NICs, and GPUs) come to enable low-latency and high-bandwidth device operations, the traditional host-centric server architectures fail to deliver high performance due to their inefficient device handling mechanisms. Furthermore, without resolving the architecture inefficiency, the performance loss will continue to increase as the emerging devices become faster.
Jaehyung Ahn, Dongup Kwon, Youngsok Kim, Mohammadamin Ajdari, Jangwoo Kim
MICRO4