Mahmood Fazlali

dblp:61/5947 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-1701-5562ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2024 Parallel Fractional Stochastic Gradient Descent With Adaptive Learning for Recommender Systems
abstract
The structural change toward the digital transformation of online sales elevates the importance of parallel processing techniques in recommender systems, particularly in the pandemic and post-pandemic era. Matrix factorization (MF) is a popular and scalable approach in collaborative filtering (CF) to predict user preferences in recommender systems. Researchers apply Stochastic Gradient Descent (SGD) as one of the most famous optimization techniques for MF. Paralleling SGD methods help address big data challenges due to the wide range of products and the sparsity in user ratings. However, these methods’ convergence rate and accuracy are affected by the dependency between the user and item latent factors, specifically in large-scale problems. Besides, the performance is sensitive to the applied learning rates. This article proposes a new parallel method to remove dependencies and boost speed-up by using fractional calculus to improve accuracy and convergence rate. We also apply adaptive learning rates to enhance the performance of our proposed method. The proposed method is based on Compute Unified Device Architecture (CUDA) platform. We evaluate the performance of our proposed method using real-world data and compare the results with the close baselines. The results show that our method can obtain high accuracy and convergence rate in addition to high parallelism.
Fatemeh Elahi, Mahmood Fazlali, Hadi Tabatabaee Malazi, Mehdi Elahi
IEEE Trans. Parallel Distributed Syst.2
2023 A fast MILP solver for high-level synthesis based on heuristic model reduction and enhanced branch and bound algorithm
Mina Mirhosseini, Mahmood Fazlali, Mohammad K. Fallah, Jeong-A Lee
J. Supercomput.2
2021 Parallel branch and bound algorithm for solving integer linear programming models derived from behavioral synthesis
Mohammad K. Fallah, Mahmood Fazlali
Parallel Comput.2
2021 Accelerating Louvain community detection algorithm on graphic processing unit
Mahmood Fazlali, Mehdi Hosseinzadeh 0001
J. Supercomput.2
2020 Scalable Parallel Genetic Algorithm For Solving Large Integer Linear Programming Models Derived From Behavioral Synthesis
abstract
Solving Integer Linear Programming (ILP) models generally lies in the category of NP-hard problems. Therefore, as the size of ILP models grows, the efficiency of exact algorithms for solving the models reduced significantly and for large models it is not possible to have the result. Genetic Algorithm (GA) is a metaheuristic method capable of adjusting and redesigning parameters and operations according to the characteristics of ILP models. Still GA has huge search space for large models and parallelization is a suitable technique to tackle this problem. This paper presents a scalable parallel GA to solve large ILP models derived from behavioral synthesis of digital circuits. We show that although models have non-binary variables, only binary variables are sufficient for coding chromosomes. We also use ”unknown” values for some genes to decrease the likelihood of inconsistency in the encoded constraints. Our experiments verify the efficiency and scalability of the proposed algorithm on multicore platforms. The proposed method outperforms IBM ILOG CPLEX 12.6 and MI-LXPM algorithm where the ILP models include 550 to 2258 int / binary decision variables. Also, the results indicate that the saturation point of using parallel processing elements for solving the large ILP models is at least 60.
Mohammad K. Fallah, Mina Mirhosseini, Mahmood Fazlali, Masoud Daneshtalab
PDP3
2019 High-speed GPU implementation of a secret sharing scheme based on cellular automata
Saeideh Kabirirad, Mahmood Fazlali, Ziba Eslami
J. Supercomput.2
2017 A hybrid bio-inspired learning algorithm for image segmentation using multilevel thresholding
Mohammad Mahdi Dehshibi, Mohamad Sourizaei, Mahmood Fazlali, Omid Talaee, Hossein Samadyar, Jamshid Shanbehzadeh
Multim. Tools Appl.3
2016 Metamorphic malware detection using opcode frequency rate and decision tree
abstract
Malware is defined as any type of malicious code that is the potent to harm a computer or a network. Modern malwares are accompanied with mutation characteristics, namely polymorphism and metamorphism. They let malwares to generate enormous number of variants. Rising number of metamorphic malwares entails hardship in analyzing them for signature extraction and database updates. In spite of the broad use of signature-based methods in the security products, they are not able detect the new unseen morphs of malware, and it is stemmed from changing the structure of malware as well as the signature in each infection. In this paper, a novel method is proposed in which the proportion of opcodes is used for detecting the new morphs. Decision trees are utilized for classification and detection of malware variants based on the rate of opcode frequencies. Three metrics for evaluating the proposed method are speed, efficiency and accuracy. It was observed in the course of experiments that speed and time complexity will not be challenging factors; because of the fast nature of extracting the frequencies of opcodes from source assembly file. Empirical validation reveals that the proposed method outperforms the entire commercial antivirus programs with a high level of efficiency and accuracy.
Mahmood Fazlali, Peyman Khodamoradi, Farhad Mardukhi, Masoud Nosrati, Mohammad Mahdi Dehshibi
Int. J. Inf. Secur. Priv.1
2014 Linear principal transformation: toward locating features in N-dimensional image space
Mohammad Mahdi Dehshibi, Mahmood Fazlali, Jamshid Shanbehzadeh
Multim. Tools Appl.2
2013 High-speed Binary Signed-Digit RNS adder with posibit and negabit encoding
abstract
Binary Signed-Digit Residue Number System (BSD-RNS) has been proposed in the literatures as an appropriate number system to perform the arithmetic operations in parallel. BSD-RNS addition is the basic operation and improving its performance results in efficient VLSI arithmetic circuits. Here, we present a new architecture for carry-free BSD-RNS addition utilizing a recently proposed posibit and negabit BSD representation. Compared to 2's complement BSD-RNS adder, the proposed architecture has 21% less delay. Besides, for a same delay (0.6ns), we obtain 48% less area and 28% less power than the most efficient existing BSD-RNS adder.
Somayeh Timarchi, Maryam Saremi, Mahmood Fazlali, Georgi Gaydadjiev
VLSI-SoC3
2013 Clustering Persian viseme using phoneme subspace for developing visual speech application
Mohammad Aghaahmadi, Mohammad Mahdi Dehshibi, Azam Bastanfard, Mahmood Fazlali
Multim. Tools Appl.4
2012 Efficient datapath merging for the overhead reduction of run-time reconfigurable systems
Mahmood Fazlali, Ali Zakerolhosseini, Georgi Gaydadjiev
J. Supercomput.1
2010 A unified addition structure for moduli set {2n-1, 2n, 2n+1} based on a novel RNS representation
abstract
Given that modulo 2n±1 are the most popular moduli in Residue Number Systems (RNS), a large variety of modulo 2n±1 adder designs have been proposed based on different number representations. However, in most of the cases, these encodings do not allow the implementation of a unified adder for all the moduli of the form 2n-1, 2n, and 2n+1. In this paper, we address the modular addition issue by introducing a new encoding, namely, the stored-unibit RNS. Moreover, we demonstrate how the proposed representation can be utilized to derive a unified design for the moduli set {2n-1,2n,2n+1}. Our approach enables a unified design for the moduli set adders, which opens the possibility to design reliable RNS processors with low hardware redundancy. Moreover, the proposed representation can be utilized in conjunction with any fast state of the art binary adder without requiring any extra hardware for end-around-carry addition.
Somayeh Timarchi, Mahmood Fazlali, Sorin Cotofana
ICCD2
2010 Efficient task scheduling for runtime reconfigurable systems
Mahmood Fazlali, Mojtaba Sabeghi, Ali Zakerolhosseini, Koen Bertels
J. Syst. Archit.1