Ali Mahani 0001

dblp:119/0257 · also Ali K. Mahani, Ali Khayatzadeh Mahani · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0003-4916-202XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 8 since 2021Computer networks · 5 · 3 first-authorArtificial intelligence and machine learning · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HAWX: A Hardware-Aware FrameWork for Fast and Scalable ApproXimation of DNNs
Samira Nazari, Mohammad Saeed Almasi, Mahdi Taheri, Ali Azarpeyvand, Ali Mokhtari, Ali Mahani 0001, Christian Herglotz
DATE6
2026 An FPGA-Based SoC Architecture with a RISC-V Controller for Energy-Efficient Temporal-Coding Spiking Neural Networks
Mohammad Javad Sekonji, Ali Mahani 0001, Maryam Mirsadeghi, Mahdi Taheri
ISCAS2
2026 Robust DCNN: The impact of approximate multipliers in defending against adversarial attacks
Mohammad Javad Askarizadeh, Jorge Castro-Godínez, Ebrahim Farahmand, Ali Mahani 0001, Laura Cabrera Quiros, Carlos Salazar-García
Future Gener. Comput. Syst.4
2026 A RISC-V Accelerator for Sequence Decoding in Mobile DNA Sequencers
abstract
Modern nanopore sequencers generate raw signal data at high speed, demanding low-latency and energy-efficient basecalling pipelines to enable fully portable genomic analysis. In this work, we present a hardware accelerator for the Viterbi-based connectionist temporal classification (CTC) decoding stage of basecalling—a key bottleneck in translating neural network outputs into deoxyribonucleic acid (DNA) sequences. Our design is the first pipelined CTC Viterbi decoder architecture tailored for nanopore sequencing and is implemented on a Xilinx Virtex-7 (VC707) FPGA within a Linux-capable reduced instruction set computer-fifth generation (RISC-V) system-on-chip (SoC). The accelerator processes over 23 000 DNA bases per second at 100 MHz with about$4.3~\boldsymbol {\mu }$s per-sample latency and only 0.43-W overhead power. This corresponds to$\textbf {5.3}\times \mathbf {10^{4}}$bases/J ($19~\boldsymbol {\mu }$J/base) and yields approximately 7x end-to-end speedup over a CPU baseline, while reserving the baseline read-identity accuracy. For the same CTC task, the accelerator delivers 29x higher throughput than a recent FPGA beam-search decoder. These results demonstrate the viability of dedicated decoding accelerators for real time, on-device genomic processing in power-constrained environments.
Amin Savari, Ali Mahani 0001, Ebrahim Ghafar-Zadeh, Sebastian Magierowski
IEEE Trans. Very Large Scale Integr. Syst.2
2025 RL-Agent-based Early-Exit DNN Architecture Search Framework
abstract
This paper introduces a Reinforcement Learning (RL)-based framework for optimizing early-exit configurations in Deep Neural Networks (DNNs). By integrating RL with BranchyNet-inspired architectures, the framework dynamically determines optimal early exit placements and confidence thresholds, balancing inference time, energy consumption, and accuracy. Key contributions include an early-exit DNN architecture search, an RL-driven threshold optimization process during training, and a design-space exploration open-source framework. Experiments on models such as ResNet-18, VGG-16, and AlexNet, using benchmarks like CIFAR-10 and MNIST, reveal significant reductions in inference time (up to 69.7x) and power consumption while keeping accuracy drop within 1-2%. This work demonstrates that dynamic early-exit strategies can enhance DNN efficiency while maintaining performance, paving the way for resource-constrained applications.
Mahdi Taheri, Parth Patne, Natalia Cherezova, Ali Mahani 0001, Christian Herglotz, Maksim Jenihhin
DDECS4
2023 Design and Analysis of High Performance Heterogeneous Block-based Approximate Adders
abstract
Approximate computing is an emerging paradigm to improve the power and performance efficiency of error-resilient applications. As adders are one of the key components in almost all processing systems, a significant amount of research has been carried out toward designing approximate adders that can offer better efficiency than conventional designs; however, at the cost of some accuracy loss. In this article, we highlight a new class of energy-efficient approximate adders, namely, Heterogeneous Block-based Approximate Adders (HBAAs), and propose a generic configurable adder model that can be configured to represent a particular HBAA configuration. An HBAA, in general, is composed of heterogeneous sub-adder blocks of equal length, where each sub-adder can be an approximate sub-adder and have a different configuration. The sub-adders are mainly approximated through inexact logic and carry truncation. Compared to the existing design space, HBAAs provide additional design points that fall on the Pareto-front and offer a better quality-efficiency tradeoff in certain scenarios. Furthermore, to enable efficient design space exploration based on user-defined constraints, we propose an analytical model to efficiently evaluate the Probability Mass Function (PMF) of approximation error and other error metrics, such as Mean Error Distance (MED), Normalized Mean Error Distance (NMED), and Error Rate (ER) of HBAAs. The results show that HBAA configurations can provide around 15% reduction in area and up to 17% reduction in energy compared to state-of-the-art approximate adders.
Ebrahim Farahmand, Ali Mahani 0001, Muhammad Abdullah Hanif, Muhammad Shafique 0001
ACM Trans. Embed. Comput. Syst.2
2022 A Novel Fault-Tolerant Logic Style with Self-Checking Capability
abstract
We introduce a novel logic style with self-checking capability to enhance hardware reliability at logic level. The proposed logic cells have two-rail inputs/outputs, and the functionality for each rail of outputs enables construction of fault-tolerant configurable circuits. The AND and OR gates consist of 8 transistors based on CNFET technology, while the proposed XOR gate benefits from both CNFET and low-power MGDI technologies in its transistor arrangement. To demonstrate the feasibility of our new logic gates, we used an AES S-box implementation as the use case. The extensive simulation results using HSPICE indicate that the case-study circuit using on proposed gates has superior speed and power consumption compared to other implementations with error-detection capability.
Mahdi Taheri, Saeideh Sheikhpour, Ali Mahani 0001, Maksim Jenihhin
IOLTS3
2021 Accelerating Deep Convolutional Neural Network base on stochastic computing
Mohamad Hasani Sadi, Ali Mahani 0001
Integr.2
2020 A Novel Energy-Efficient Clustering Protocol Using Two-Stage Genetic Algorithm for Improving the Lifetime of Wireless Sensor Networks
abstract
Wireless sensor networks (WSNs) are beginning to be deployed at an accelerated pace, and they have attracted significant attention in a broad spectrum of applications. WSNs encompass a large number of sensor nodes enabling a base station (BS) to sense and transmit data over the area where WSN is spread. As most sensor nodes have a limited energy capacity and at the same time transmit critical information, enhancing the lifetime and the reliability of WSNs are essential factors in designing these networks. Among many approaches, clustering of sensor nodes has proved to be an effective method of reducing energy consumption and increasing lifetime of WSNs. In this paper, a new energy-efficient clustering protocol is implemented using a two-step Genetic Algorithm (GA). In the first step of GA, cluster heads (CHs) are selected, and in the second step, cluster members are chosen based on their distance to the selected CHs. Compared to other clustering protocols, the lifetime of WSNs in the proposed clustering is improved. This improvement is the consequence of the fact that this clustering considers energy efficient parameters in clustering protocol.
Ali Mahani 0001, Ebrahim Farahmand, Saeideh Sheikhpour, Nooshin Taheri-Chatrudi
Int. J. Comput. Intell. Appl.1
2020 A survey on fault injection methods of digital integrated circuits
Mohammad Eslami, Behnam Ghavami, Mohsen Raji, Ali Mahani 0001
Integr.4
2019 Gravitational search algorithm with both attractive and repulsive forces
Hamed Zandevakili, Esmat Rashedi, Ali Mahani 0001
Soft Comput.3
2018 A New ASIC Structure With Self-Repair Capability Using Field-Programmable Nanowire Interconnect Architecture
Hamed Zandevakili, Ali Mahani 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Availability Improvement Method for Repairable Systems Using Modified Ant Colony Optimisation
abstract
Reliability is deemed as an important issue in many broad applications, e.g., telecommunication systems, electric power systems, computation and parallel processing systems. In order to reach efficient performance and reliable/available design, we need to optimize the cost and reliability/availability of the proposed designs. The main contribution of this paper is to address the optimum model for repairable component with series-parallel structure using redundancy allocation problem (RAP). In this regard, two novel modified Ant Colony Optimization (ACO), are defined. In order to improve availability, ACOs employ visibility and pheromone update to optimize the RAP. The proposed methodology includes single objective to maximize the availability of case study system with two constraints-weight and cost of components. Hence, three meta-heuristic algorithms, i.e., ACO, Genetic algorithm (GA) and Particle Swarm Optimization (PSO) are applied to find the optimum structure. Finally, the simulation results of all the proposed meta-heuristic algorithms are compared. The comparison reveals that the modified ACO (MACO) provides the maximum availability among all the other meta-heuristic algorithms.
Ali Mahani 0001, Ebrahim Farahmand, Saeideh Sheikhpour, Nooshin Taheri-Chatrudi
Int. J. Comput. Intell. Appl.1
2017 Particle Swarm Optimization with Intelligent Mutation for Nonlinear Mixed-Integer Reliability-Redundancy Allocation
abstract
As improving system reliability in a basic system has been always one of the important concerns in reliability engineering; many studies have been developed in this regard. In this paper, a novel intelligent PSO (PSO-IM) is proposed. In suggested approach two different types of mutation operator, which controlled by Fuzzy controller, are applied to standard PSO for finding the best solution of reliability-redundancy allocation problems (RRAP). Also a heuristic inertia weight equation is introduced in our proposed PSO-IM. The main objective of our solution is to achieve maximum reliability in a basic system with acceptable redundancy allocation subject to cost, weight, volume. The proposed method (PSO-IM) significantly improves the search ability of basic PSO and finds the maximum reliability in comparison with previous works.
Saeideh Sheikhpour, Ali Mahani 0001
Int. J. Comput. Intell. Appl.2
2014 Channel assignment in multi-radio wireless mesh networks using an improved gravitational search algorithm
Mohammad Doraghinejad, Hossein Nezamabadi-pour, Ali Mahani 0001
J. Netw. Comput. Appl.3
2012 Two-stage uncertainty incorporating in optical core networks
abstract
In guaranteed-type applications the bandwidth planning and cost management are two important issues in deployment of optical core networks. A two-stage fuzzy-based approach is proposed for accommodating long-term demand uncertainties in dense wavelength division multiplexing optical networks. Here, the uncertainties are modelled using the Gaussian fuzzy membership functions. First, the forecasted part of demand matrix is introduced to Dijkstra shortest path-based routing algorithm. Then, the available wavelengths are assigned to demand uncertainties. Unlike existing algorithms, different bandwidth cost factors are assigned to the links of a lightpath according to network and links state information. The performance of proposed approach is evaluated on a typical optical link for uncertain traffic loads in Erlang mode. Simulation results show that the proposed fuzzy-based approach is up to 29% cost-effective for accommodating network demands in real-world applications comparing to existing algorithms.
Yousef Seifi Kavian, Ali Mahani 0001, Habib F. Rashvand
IET Commun.2
2011 Heavy-tail and voice over internet protocol traffic: queueing analysis for performance evaluation
abstract
Study the effects of concurrent voice connections on the performance metrics of communication network such as queue length, waiting time, packets service time and is very important. Mathematical analysis of such network especially with long-tail traffic will help us for a good capacity planning and also lead to an accurate admission control algorithms. In this study a mathematical model of a communication network supporting VoIP and back-ground traffic with long-tail service time is considered. Some problems of previous mathematical models are identified and a new queueing system is proposed in which specifically the coexisting of heavy-tail and voice flows is addressed. The long-tail service time is approximated via hyper-Erlang distribution and also to achieving an accurate performance model a Markov reward model is introduced. The available bandwidth for long-tail distribution varies according to the Markov chain, describing the utilisation factor of voice connection. Numerical results show a comparison between exponential and heavy-tail service time and finally the effects of concurrent voice connections on the service time of heavy-tailed back-ground packets is shown.
Ali Mahani 0001, Yousef Seifi Kavian, Majid Naderi, Habib F. Rashvand
IET Commun.1
2009 MAC-layer channel utilisation enhancements for wireless mesh networks
abstract
The authors focus on a wireless mesh network, that is, an ad hoc IEEE 802.11-based network whose nodes are either user devices or Access Points providing access to the mesh network or to the Internet. By relying on some work done within the IEEE 802.11s TG, the network nodes can use one control channel and one or more data channels, each on separate frequencies. Then, some problems related to channel access are identified and a MAC scheme is proposed that specifically addresses the problem of hidden terminals and the problem of coexisting control and data traffic on different frequency channels. An analytical model of the MAC scheme is presented and validated by using the Omnet++ simulator. Through the developed model, we show that our solution achieves very good performance both in regular and in very fragmented mesh topologies, and it significantly outperforms the standard 802.11 solution.
Ali Mahani 0001, Majid Naderi, Claudio Casetti, Carla Fabiana Chiasserini
IET Commun.1
2009 Wireless mesh networks channel reservation: modelling and delay analysis
abstract
In order to overcome the negotiation procedure bottleneck of the standard DCF in wireless mesh networks, the authors propose a new channel reservation function (CRF) that reduces the negotiation overhead of the DCF, which as a result reduces the overall transmission delay effectively without of any extra bandwidth consumption. Furthermore, the authors provide an analytical model for the proposed scheme for which the simulation results measure the amount that the new method can reduce the average total delay for both regular and fragmented mesh topologies demonstrating superiority of the new method over the classic 802.11 solution. Additionally, the authors extend the scheme to multichannel CRF upon which the proposed method can be used for multichannel applications.
Ali Mahani 0001, Habib F. Rashvand, Ebrahim Teimoury, Majid Naderi, Bahman Abolhassani
IET Commun.1