Mansour Shafaei

dblp:98/9640 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 5 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Storage systems · 67% Performance modeling and evaluation · 29% Memory systems · 4%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › magnetic recording › shingled magnetic recording
drive-managed SMR
0.522017
Modeling Drive-Managed SMR Performance · ACM Trans. Storage 2017
Skylight - A Window on Shingled Disk Operation · ACM Trans. Storage 2015
Storage systems › magnetic recording
shingled magnetic recording
0.522017
Modeling Drive-Managed SMR Performance · ACM Trans. Storage 2017
Skylight - A Window on Shingled Disk Operation · ACM Trans. Storage 2015
Performance modeling and evaluation
storage performance modeling
0.312017
Modeling Drive-Managed SMR Performance · ACM Trans. Storage 2017
Performance modeling and evaluation
workload characterization
0.212015
Skylight - A Window on Shingled Disk Operation · ACM Trans. Storage 2015
Storage systems › magnetic storage
hard disk drive
0.222017
Modeling Drive-Managed SMR Performance · ACM Trans. Storage 2017
Skylight - A Window on Shingled Disk Operation · ACM Trans. Storage 2015
Memory systems › cache management › storage caching
persistent cache
0.112015
Skylight - A Window on Shingled Disk Operation · ACM Trans. Storage 2015

Methods — techniques the papers use, named apart from their topics

predictive simulation · 0.3black-box measurement · 0.3latency measurement · 0.2high-speed camera observation · 0.2
YearPublicationVenuePosition
2018 FSTL: A Framework to Design and Explore Shingled Magnetic Recording Translation Layers
abstract
We introduce FSTL, a Framework for Shingled Translation Layers: a toolkit for implementing host-side block translation layers for Shingled Magnetic Recording (SMR) drives. It provides a Linux kernel implementation of key translation mechanisms (write allocation, LBA translation, map persistence, and consistent copying) while allowing translation policy (e.g. layout, cleaning algorithms, crash recovery) to be implemented in a user-space controller which communicates through an ioctl-based API to the kernel data plane. Due to its use of a journaled write format, FSTL-based translation layers are able to handle synchronous and durable writes to random LBAs at high speed. We describe the architecture and implementation of FSTL, and present two FSTL-based translation layers implemented in 400 lines of Python each. Despite the simplicity of the controllers, experiments show our first translation layer performing 1.5x to 10x better than a drive-managed translation layer on trace replay experiments, with performance roughly comparable to the drive-managed device for real file system-based benchmarks; the second translation layer, based on a full-volume extent map, is shown to offer significantly better performance than prior work. Furthermore, we implement and evaluate three cleaning algorithms to demonstrate how FSTL-based translation layers may be readily modified, while still offering the robustness needed for long-duration benchmarks and use.
Mohammad Hossein Hajkazemi, Mania Abdi, Mansour Shafaei, Peter Desnoyers
MASCOTS3
2017 Virtual Guard: A Track-Based Translation Layer for Shingled Disks
Mansour Shafaei, Peter Desnoyers
HotStorage1
2017 Modeling Drive-Managed SMR Performance
abstract
Accurately modeling drive-managed Shingled Magnetic Recording (SMR) disks is a challenge, requiring an array of approaches including both existing disk modeling techniques as well as new techniques for inferring internal translation layer algorithms. In this work, we present the first predictive simulation model of a generally available drive-managed SMR disk. Despite the use of unknown proprietary algorithms in this device, our model that is derived from external measurements is able to predict mean latency within a few percent, and with an Root Mean Square (RMS) cumulative latency error of 25% or less for most workloads tested. These variations, although not small, are in most cases less than three times the drive-to-drive variation seen among seemingly identical drives.
Mansour Shafaei, Mohammad Hossein Hajkazemi, Peter Desnoyers, Abutalib Aghayev
ACM Trans. Storage1
2016 Write Amplification Reduction in Flash-Based SSDs Through Extent-Based Temperature Identification
Mansour Shafaei, Peter Desnoyers, Jim Fitzpatrick
HotStorage1
2016 Modeling SMR Drive Performance
abstract
No abstract available.
Mansour Shafaei, Mohammad Hossein Hajkazemi, Peter Desnoyers, Abutalib Aghayev
SIGMETRICS1
2015 Skylight - A Window on Shingled Disk Operation
abstract
We introduce Skylight, a novel methodology that combines software and hardware techniques to reverse engineer key properties of drive-managed Shingled Magnetic Recording (SMR) drives. The software part of Skylight measures the latency of controlled I/O operations to infer important properties of drive-managed SMR, including type, structure, and size of the persistent cache; type of cleaning algorithm; type of block mapping; and size of bands. The hardware part of Skylight tracks drive head movements during these tests, using a high-speed camera through an observation window drilled through the cover of the drive. These observations not only confirm inferences from measurements, but resolve ambiguities that arise from the use of latency measurements alone. We show the generality and efficacy of our techniques by running them on top of three emulated and two real SMR drives, discovering valuable performance-relevant details of the behavior of the real SMR drives.
Abutalib Aghayev, Mansour Shafaei, Peter Desnoyers
ACM Trans. Storage2
2014 HiTS: A High Throughput Memory Scheduling Scheme to Mitigate Denial-of-Service Attacks in Multi-core Systems
abstract
Sharing DRAM memory by multiple cores in a computer system potentially exposes the running threads on cores to denial-of-service (DoS) attacks. This issue is usually addressed by memory scheduling schemes that rotate the memory service among threads according to a certain ranking mechanism. These ranking-based schemes, however, often incur many memory banks' row-buffer conflicts which reduce the throughput of DRAM and the entire system. This paper proposes a new ranking-based memory scheduling scheme, called HiTS, to mitigate DoS attacks in multicore systems with the lowest performance degradation. HiTS achieves these by ranking threads according to each thread's memory usage/requirement. HiTS then enforces the ranking in a way that minimum performance overhead would occur and fairness is also balanced. The effectiveness of HiTS is evaluated by simulations with 18 different workloads running on 8- and 16-core machines. The simulation results show up to 15.8% improvements in terms of unfairness reduction and 24.1% in system throughput compared with the best existing scheduling scheme.
Mansour Shafaei, Yunsi Fei
SBAC-PAD1
2011 Numeral-Based Crosstalk Avoidance Coding to Reliable NoC Design
abstract
This paper proposes a Numeral-Based Crosstalk Avoidance Coding (NB-CAC) to protect communication channels of Network-on-Chips (NoCs) against crosstalk faults. The NB-CAC scheme produces code words without bit patterns '101' and '010' to eliminate harmful transition patterns from NoC channels. This is done by the use of a new numeral system proposed in the paper. Using the proposed numeral system, the NB-CAC scheme 1) can be utilized in NoC channels with any arbitrary width, and 2) can be implemented with low area, power, and timing overheads. VHDL and SPICE simulations have been carried out for a wide range of channel widths to evaluate delay, area, and power consumption of the NB-CAC codecs. Results of simulations reveal that the NB-CAC scheme completely removes crosstalk faults from NoC channel. In addition, the NB-CAC scheme provides reductions of 17.3% in area and 31.9% in power-delay product with respect to Fibonacci-based coding which has been recently proposed in literature.
Mansour Shafaei, Ahmad Patooghy, Seyed Ghassem Miremadi
DSD1
2010 Crosstalk modeling to predict channel delay in Network-on-Chips
abstract
Communication channels in Network-on-Chips (NoCs) are highly susceptible to crosstalk faults due to the use of nano-scale VLSI technologies in the fabrication of NoCs. Crosstalk faults cause variable timing delay in NoC channels based on the patterns of transitions appearing on the channels. This paper proposes an analytical model to estimate the timing delay of an NoC channel in the presence of crosstalk faults. The model calculates expected number of 4C, 3C, 2C, and 1C transition patterns to predict delay of a K-bit communication channel. The model is applicable for both non-protected channels and channels which are protected by crosstalk mitigation methods. Spice simulations are done in a wide range of working conditions to validate the proposed model. Delays extracted from the simulations are compared with those obtained from the model. Comparisons show that the proposed model accurately estimates the delay of NoC channels. In addition, the proposed model accelerates the evaluation phase of any crosstalk mitigation method by at least three orders of magnitude.
Ahmad Patooghy, Seyed Ghassem Miremadi, Mansour Shafaei
ICCD3
2010 FiRot: An Efficient Crosstalk Mitigation Method for Network-on-Chips
abstract
This paper proposes an efficient cross talk mitigation method for Network-on-Chips (NoCs). The proposed method investigates flits in each packet to minimize the number of harmful transition patterns appearing on the communication channels of NoC. To do this, the content of every flit is rotated with respect to the previously flit sent through the channel. Rotation is done to find a rotated version of the flit which minimizes the number of harmful transition patterns. A tag field is added into the rotated flit to enable the receiving side to recover the original flit. Maximum number of rotations is bounded by a fixed value to minimize the timing and power overheads of the proposed method. Evaluation of the proposed method is done in both analytical and simulation manners. VHDL-based simulations are carried out for several channel widths and several tag widths. Simulation results confirm that the proposed method effectively overcomes the cross talk problem while its timing and power overheads are negligible. Results of analytical evaluation are also in agreement with the simulation results.
Ahmad Patooghy, Mansour Shafaei, Seyed Ghassem Miremadi, Hajar Falahati, Somayyeh Taheri
PRDC2