Hadi Shahriar Shahhoseini

dblp:20/6614 · also Hadishahriar Shahhoseini · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-6042-0993ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorArtificial intelligence and machine learning · 3 · 1 since 2021Computer networks · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Approximate Reciprocal-based Divider
Ali Ghaderi, Nima Amirafshar, Hadi Shahriar Shahhoseini, Nima Taherinejad
ISCAS3
2026 Thermal-aware routing in three-dimensional network on chips with temperature prediction using principal component analysis and adaptive neuro-fuzzy inference system
Erfan Khedersolh, Majid Nezarat, Hadi Shahriar Shahhoseini, Mohammad Reza Mosavi
Eng. Appl. Artif. Intell.3
2025 PRIM: Hybrid Array-Compressor Multipliers with Carry Disregard and OR-based Approximation
abstract
This paper introduces an efficient new 4:1 compressor that uses carry disregard and OR-based approximation, leading to the development of 13 approximate unsigned multipliers. The proposed multipliers, 8-bit array-comPressor oR-based carry dIsregard Multipliers (PRIM8s) demonstrate significant improvements in area, power, delay, and Power-Delay-Product (PDP) by an average of 29%, 31%, 25%, and 47%, compared to the exact multiplier. In the approximate multiplier literature, with our hardware, we establish new Pareto fronts for most criteria. The effectiveness of the proposed multipliers for noise reduction is demonstrated in an image-processing application using a low-pass Gaussian filter. On average, PRIM8s reduce power consumption and improve speed by 32.36% and 19.01% compared to the exact multiplier, while also enhancing image quality, as indicated by a 0.14% increase in Structural Similarity Index Measure (SSIM).
Nima Amirafshar, Gulafshan, Hadi Shahriar Shahhoseini, Nima Taherinejad
ISCAS3
2025 A customized balanced-objective genetic algorithm for task scheduling in reconfigurable computing systems
Milad Gholamrezanejad, Hadi Shahriar Shahhoseini, Seyed Mehdi Mohtavipour
Knowl. Inf. Syst.2
2025 PRISA: A Potential Region-based Intelligent Search Algorithm for Dataflow Graph Mapping in Spatial CGRAs
abstract
Coarse-grained Reconfigurable Architectures (CGRAs) offer energy efficiency and programmability, making them integral to modern high-performance computing. However, complicated compilation when mapping the Dataflow Graph (DFG) to CGRA components leads to significant time overhead, area inefficiency, and routing challenges, especially for large-scale applications today. In this article, a fast and accurate DFG mapping approach called PRISA is proposed to reduce the compilation delay and obtain mapping solutions with higher qualities in a more reasonable time. This approach analytically identifies the potential and weak regions in the space of mapping solutions to guide the search algorithm and prevent ineffective examinations. Moreover, using the potential solutions and a novel sparse matrix permutation technique, we introduced Selective Initial Solution (SIS) to further improve the performance of PRISA without needing long-time optimizations. As the computational requirements in the PRISA are performed analytically, there is no additional time overhead. We conducted extensive experiments on the benchmark graphs of CGRA-ME and VPR-8 toolkits and obtained outstanding results compared to integer linear programming, evolutionary, and graph traversal approaches in terms of compilation time and mapping communication cost. Our approach could map the DFGs of VPR-8 toolkit with 3.64 maximal FIFO size requirement and 53.5 ms time overhead, on average.
Seyed Mehdi Mohtavipour, Hadi Shahriar Shahhoseini
ACM Trans. Reconfigurable Technol. Syst.2
2024 A Reconfigurable Approximate Computing RISC-V Platform for Fault-Tolerant Applications
abstract
The demand for energy-efficient and high-performance embedded systems drives the evolution of new hardware architectures, including concepts like approximate computing. This paper presents a novel reconfigurable embed-ded platform named “phoeniX”, using the standard RISC-V ISA, maximizing energy efficiency while maintaining acceptable application-level accuracy. The platform enables the integration of approximate circuits at the core level with diverse structures, accuracies, and timings without requiring modifications to the core, particularly in the control logic. The platform introduces novel control features, allowing configurable trade-offs between accuracy and energy consumption based on specific application requirements. To evaluate the effectiveness of the platform, experiments were conducted on a set of applications, such as image processing and Dhrystone benchmark. The core with its original execution engine, occupies 0.024mm2 of area, with average power consumption of 4.23m W at 1.1 V operating voltage, average energy-efficiency of 7.85pJ per operation at 620MHz frequency in 45nm CMOS technology. The configurable platform with a highly optimized 3-stage pipelined RV32I(E)M architecture, possesses a DMIPS/MHz of 1.89, and a CPI of 1.13, showcasing remarkable capabilities for an embedded processor.
Arvin Delavari, Faraz Ghoreishy, Hadi Shahriar Shahhoseini, Sattar Mirzakuchaki
DSD3
2024 ACE-CNN: Approximate Carry Disregard Multipliers for Energy-Efficient CNN-Based Image Classification
abstract
This paper presents the design and development of Signed Carry Disregard Multiplier (SCDM8), a family of signed approximate multipliers tailored for integration into Convolutional Neural Networks (CNNs). Extensive experiments were conducted on popular pre-trained CNN models, including VGG16, VGG19, ResNet101, ResNet152, MobileNetV2, InceptionV3, and ConvNeXt-T to evaluate the trade-off between accuracy and approximation. The results demonstrate that ACE-CNN outperforms other configurations, offering a favorable balance between accuracy and computational efficiency. In our experiments, when applied to VGG16, SCDM8 achieves an average reduction in power consumption of 35% with a marginal decrease in accuracy of only 1.5%. Similarly, when incorporated into ResNet152, SCDM8 yields an energy saving of 42% while sacrificing only 1.8% in accuracy. ACE-CNN provides the first approximate version of ConvNeXt which yields up to 72% energy improvement at the price of less than only 1.3% Top-1 accuracy. These results highlight the suitability of SCDM8 as an approximation method across various CNN models. Our analysis shows that the ACE-CNN outperforms state-of-the-art approaches in accuracy, energy efficiency, and computation precision for image classification tasks in CNNs. Our study investigated the resiliency of CNN models to approximate multipliers, revealing that ResNet101 demonstrated the highest resiliency with an average difference in the accuracy of 0.97%, whereas LeNet5 Inspired-CNN exhibited the lowest resiliency with an average difference of 2.92%. These findings aid in selecting energy-efficient approximate multipliers for CNN-based systems, and contribute to the development of energy-efficient deep learning systems by offering an effective approximation technique for multipliers in CNNs. The proposed SCDM8 family of approximate multipliers opens new avenues for efficient deep learning applications, enabling significant energy savings with virtually no loss in accuracy.
Salar Shakibhamedan, Nima Amirafshar, Ahmad Sedigh Baroughi, Hadi Shahriar Shahhoseini, Nima Taherinejad
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 On the security of 'an ultra-lightweight and secure scheme for communications of smart metres and neighbourhood gateways by utilisation of an ARM Cortex-M microcontroller'
abstract
Abstract In 2018, Abbasinezhad‐Mood and Nikooghadam (IEEE Transaction on Smart Grid, pp 6194–6205, 9(6), 2018) proposed an ultra‐lightweight secure scheme for neighbourhood area network () communications in smart grid. They have claimed that their protocol is secure against all known attacks in environment by providing informal security analysis besides a formal analysis which was done by using an automatic verification tool. However, by performing several attacks, this study shows that their scheme has serious security flaws. After performing each attack, lightweight countermeasures is proposed for securing their protocol against that attack.
Sonia Miri, Masoud Kaveh, Hadi Shahriar Shahhoseini, Mohammad Reza Mosavi, Saeed Aghapour
IET Inf. Secur.3
2023 Carry Disregard Approximate Multipliers
abstract
Several challenges in improving the performance of computing systems have given rise to emerging computing paradigms. One of these paradigms is approximate computing. Many applications require different levels of accuracy and are error-tolerance to a certain degree. Approximate computations can reduce the calculation complexities significantly and thus improve the performance. Here, we propose a methodology for designing approximate N-bit array multipliers based on carry disregarding. We evaluate and analyze the proposed multipliers both experimentally and theoretically. The proposed 8-bit multipliers, compared to the exact multiplier, reduce the critical path delay, power consumption, and area by 29%, 29%, and 30%, on average. Compared to the existing approximate array architectures in the literature, they have improved 14.3%, 22.8%, and 26.4%, respectively. Compared to the exact 16-bit multiplier, the proposed 16-bit multipliers have reduced the delay, power consumption, and area by 35%, 24%, and 23% on average. In an image processing application, we have also demonstrated the applicability of a wide range of proposed multipliers, which have Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) over 30 dB and 94%, respectively.
Nima Amirafshar, Ahmad Sadigh Baroughi, Hadi Shahriar Shahhoseini, Nima Taherinejad
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 An Approximate Carry Disregard Multiplier with Improved Mean Relative Error Distance and Probability of Correctness
abstract
Nowadays, a wide range of applications can tolerate certain computational errors. Hence, approximate computing has become one of the most attractive topics in computer architecture. Reducing accuracy in computations in a premeditated and appropriate manner reduces architectural complexities, and as a result, performance, power consumption, and area can improve significantly. This paper proposes a novel approximate multiplier design. The proposed design has been implemented using 45 nm CMOS technology and has been extensively evaluated. Compared to existing approximate architectures, the proposed approximate multiplier has higher accuracy. It also achieves better results in critical path delay, power consumption, and area up to 47.54 %, 75.24%, and 92.49%, respectively. Compared to the precise multipliers, our evaluations show that the critical path delay, power consumption, and area have been improved by 39%, 18%, and 6 %, respectively.
Nima Amirafshar, Ahmad Sadigh Baroughi, Hadi Shahriar Shahhoseini, Nima Taherinejad
DSD3
2022 AxE: An Approximate-Exact Multi-Processor System-on-Chip Platform
abstract
Due to the ever-increasing complexity of computing tasks, emerging computing paradigms that increase efficiency, such as approximate computing, are gaining momentum. However, so far, the majority of proposed solutions for hardware-based approximation have been application-specific and/or limited to smaller units of the computing system and require engineering effort for integration into the rest of the system. In this paper, we present Approximate and Exact Multi-Processor system-on-chip (AxE) platform. AxE is the first general-purpose approximate Multi-Processor System-on-Chip (MPSoC). AxE is a heterogeneous RISC-V platform with exact and approximate cores that allows exploring hardware approximation for any application and using software instructions. Using the full capacity of an entire MPSoC, especially a heterogeneous one such as AxE, is an increasingly challenging problem. Therefore, we also propose a task mapping method for running exact and approximable applications on AxE. That is a mixed task mapping, in which applications are viewed as a set of tasks that can be run independently on different processors with different capabilities (exact or approximate). We evaluated our proposed method on AxE and reached a 32% average execution speed-up and 21% energy consumption saving with an average of 99.3% accuracy on three mixed workloads. We also ran a sample image processing application, namely gray-scale filter, on AxE and will present its results.
Ahmad Sadigh Baroughi, Sini Huemer, Hadi Shahriar Shahhoseini, Nima Taherinejad
DSD3
2022 Energy efficient 3D network-on-chip based on approximate communication
Masoomeh Momeni, Hadi Shahriar Shahhoseini
Comput. Networks2
2020 A link-elimination partitioning approach for application graph mapping in reconfigurable computing systems
Seyed Mehdi Mohtavipour, Hadi Shahriar Shahhoseini
J. Supercomput.2
2018 Divisible Load Scheduling of Image Processing Applications on the Heterogeneous Star Network Using a new Genetic Algorithm
abstract
The divisible load scheduling of image processing applications on the heterogeneous star network is addressed in this paper. In our platform, processors and links have different speeds. Also the computation and communication overheads are considered. A new genetic algorithm for minimizing the processing time of low level image applications using divisible load theory is introduced. A closed form solution for the processing time and the image fractions that should be assigned to each processor are obtained. The optimum number of participating processors and the optimal sequence for load distribution with a new genetic algorithm are derived. The effect of different image and kernel sizes on processing time and speed up are investigated. Finally, to indicate the efficiency of our algorithm, several numerical experiments are presented.
Sahar Nikbakht Aali, Hadi Shahriar Shahhoseini, Nader Bagherzadeh
PDP2
2018 Hyper-chaotic Feeded GA (HFGA): a reversible optimization technique for robust and sensitive image encryption
Parisa Gholizadeh Pashakolaee, Hadi Shahriar Shahhoseini, Morteza Mollajafari
Multim. Tools Appl.2
2017 MP Mitigation in Urban Canyons using GPS-combined-GLONASS Weighted Vectorized Receiver
abstract
Multipath (MP) interference in urban canyons is one of the major sources of the positioning error. Among different methods for MP mitigation, vectorised receiver (VR) is a promising one in which channels can share their information, and as a result the stronger channels can aid the weaker or affected ones to be tracked more accurately. This study proposes a weighted VR (WVR) with the strategy of giving more weights to the observations with lower vulnerability to MP. To increase the number of immune satellites to MP, global positioning system, and global navigation satellite system integration in WVR has been discussed. The performance is also compared with conventional VR. The experimental results show that the proposed method both in a static mode under an exaggerated MP condition and in a drive through an urban canyon could find the position, respectively, with about 10 and 13% improvements.
Amir Tabatabaei, Mohammad Reza Mosavi, Hadi Shahriar Shahhoseini
IET Signal Process.3
2016 An efficient ACO-based algorithm for scheduling tasks onto dynamically reconfigurable hardware using TSP-likened construction graph
Morteza Mollajafari, Hadi Shahriar Shahhoseini
Appl. Intell.2
2014 Nonflat surface level pyramid: a high connectivity multidimensional interconnection network
Hadi Shahriar Shahhoseini, Ehsan Saleh Kandzi, Morteza Mollajafari
J. Supercomput.1
2013 Improving CompactMatrix phase in gang scheduling by changing transference condition and utilizing exchange
Hossein Amir, Hadi Shahriar Shahhoseini
J. Supercomput.2
2011 Security analysis and improvement of Smart Card-Based Authenticated Key Exchange protocol with CAPTCHAs for wireless mobile network
abstract
In 2010, Fan et al. proposed a Smart Card-Based Authenticated Key Exchange protocol with CAPTCHAs for wireless mobile networks. In this paper, it is shown that their proposed protocol is vulnerable to Key Compromise Impersonation (KCI) and ephemeral key compromise impersonation attacks while it does not provide the key confirmation attribute. Furthermore, it is inefficient due to number redundancy of rounds and computational costs. To overcome these weaknesses, two secure and efficient smart card-based Authenticated Key Exchange (AKE) protocols combined with CAPTCHAs are proposed for wireless mobile networks that provide many security attributes while they have a remarkable efficiency when compared with Fan et al.'s protocol in the terms of communication costs and computational complexities.
Maryam Saeed, Hadi Shahriar Shahhoseini
ISCC2
2011 Configuration Reusing in On-Line Task Scheduling for Reconfigurable Computing Systems
Maisam Mansub Bassiri, Hadi Shahriar Shahhoseini
J. Comput. Sci. Technol.2
2010 A New approach in on-line task scheduling for reconfigurable computing systems
abstract
Reconfiguration overhead is an important obstacle that limits the performance of on-line scheduling algorithms in reconfigurable computing systems and increases the overall execution time. Configuration reusing (task reusing) can decrease reconfiguration overhead considerably, particularly in periodic applications. In this paper, we present a new approach for on-line scheduling and placement in which configuration reusing is considered as a main characteristic in order to reduce reconfiguration overhead and decrease total execution time of the tasks. A large variety of experiments have been conducted on the proposed algorithm. Obtained results show considerable improvement in overall execution time of the tasks.
Maisam Mansub Bassiri, Hadi Shahriar Shahhoseini
ASAP2
2010 A New Metric for On-Line Scheduling and Placement in Reconfigurable Computing Systems
Maisam Mansub Bassiri, Hadi Shahriar Shahhoseini
ICA3PP (2)2
2010 APPMA - an Anti-phishing protocol with mutual Authentication
abstract
The phishing as an online identity theft is one of the fastest growing crimes in the Internet. Several counter-measures are proposed through the years, one of them is the Anti-phishing Authentication (APA) protocol that is based on SPEKE which is a Password Authenticated Key Exchange (PAKE) protocol. In this paper, it is shown that the APA protocol is vulnerable to password compromise impersonation, ephemeral key compromise impersonation and malicious server attacks. An improved anti-phishing protocol is also proposed that provides several security attributes including mutual authentication, forward secrecy, known session key security, no key control, Key confirmation, and resilience to Denning-Sacco, password compromise impersonation, Unknown Key Share (UKS), off-line dictionary, undetectable online dictionary, ephemeral key compromise impersonation, Key Compromise Impersonation (KCI), eavesdropping, message loss, message modification, message insertion and message replay attacks while it provides better efficiency when compared with the APA protocol.
Maryam Saeed, Hadi Shahriar Shahhoseini
ISCC2
2008 Performance modeling of partially reconfigurable computing systems
abstract
Reconfigurable systems have become very popular because of their impact in increasing performance improvement for executing a number of applications such as morphology, image compression, etc. These systems provide a general platform for executing applications. However, effectively using the full potential of these systems cars be challenging without the knowledge of the system's performance characteristics. A general analytical model is developed in this paper for reconfigurable computing systems based on queuing theory. The reconfigurable system is composed of a host processor coupled with a reconfigurable hardware such as an FPGA.
Foad Lotfifar, Hadi Shahriar Shahhoseini
AICCSA2
2006 The Best Irreducible Pentanomials For A Mastrovito GF Multiplier
abstract
There are three main irreducible polvnomials used,for daigning Galois Field nrultiplierr whicli are: all-one polvnon~ial (AOP), equallv spaced poivnomial (ESP) and irreducible trinomial. Althougli using these polvnomials causes low space and lime complexih, in the rntrltiplier archifechrre, the.v do not exist.fir many field degrees. It has been shown that an irreducible penranomial exists whenever an irreducible trinomial does not exist for a peld degree. Therefire, using pentanomia1.v in designing multipliers have practical importance. In this paper the Mashovito mtrltiplier using general irredtrcible pentano~nials is considered This can be used in impletnenting multipliers based on any h,pe of pentanomial. Then some special h,pes of pentanomials are introduced which reduce the number of XOR gates and/or de1a.v in the Mastrovito tnultiplier
Mohsen Bahramali, Hadi Shahriar Shahhoseini
AICCSA2
2006 Semi-Algorithmic Test Pattern Generation
abstract
Nowadays, the reliability and correctness of digital circuits has become increasingly important. In addition, digital systems design methodology has been changing to HDL-based design. Thus the test methods which are based on behavioral level design would be broadly applicable. The other factor primarily influencing the practical field of application of a specific test generation algorithm is the computational complexity of the algorithm. Random test pattern generation algorithms are much simple than the other types of algorithms. Hence, developing such algorithms is exceedingly beneficial. In this paper, an extension to random test pattern generation is proposed which makes it semi-algorithmic, so it is called SAT. SAT generates some parts of test vectors deterministically for the conditional nodes of the circuit’s CFG extracted from VDHL code. Simulation shows the quality of SAT generated test sets, according to path coverage, is better than ones produced by uniformly random test pattern generation.
Hadi Shahriar Shahhoseini, Babak Hosseini Kazerouni
AICCSA1
2006 A mesh-based routing protocol for wireless ad- sensor networks
abstract
Routing algorithms in wireless sensor networks have attracted much attention during recent years. Due to the limited resources of each sensor node, e.g. limited power supply and limited communication bandwidth, routing algorithm in such networks will be complex and affect the system performance and its lifetime. In this paper the area, which sensor laid out, is partitioned into some regions. Each region, which contains some sensors, is considered as a node of a Mesh. The nodes can communicate to their neighbor nodes through a number of virtual channels. Each virtual channel formed by connecting two sensors to adjacent regions that are located in Radio-range of each other. Simulation results show that the proposed algorithm (MBR) can be used to optimize the use of the network resources and reduce the power consumption of the network.
Foad Lotfifar, Hadi Shahriar Shahhoseini
IWCMC2