EDBT 2026 Demo / reviewers in the wild / expert
Mostafa E. Salehi
dblp:20/7718
· DBLP profile ↗
24ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0003-1733-6056ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient quantized transformer for atrial fibrillation detection in cross-domain datasets
Maedeh H. Toosi, Mahdi Mohammadi-nasab, Siamak Mohammadi, Mostafa E. Salehi |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | MLB-MAC: Multi-Level Binary MAC Array for Energy Efficient ML AcceleratorsabstractQuantization is a critical compression technique for optimizing deep neural networks (DNNs) on resource-constrained embedded devices. Efficient hardware utilization hinges on effective number representation. Integer representation, a widely adopted method, uses scaling factors and offsets to enhance network accuracy and simplify fractional bit selection through uniform quantization. On the other hand, non-uniform quantization is well-suited for DNNs with parameters following a normal distribution, helping reduce data width requirements. This paper introduces a novel non-uniform representation called MLB (Multi-Level Binary), which encompasses and extends integer representation. We propose an architecture that objectively compares these representations, demonstrating that MLB is a superset of integer representation in terms of accuracy. Our comprehensive analysis spans various data widths, parallel factorization, and DNN models. Our findings indicate that MLB outperforms integer representation in energy efficiency for lower bit widths (2–5 bits), whereas integer representation is more advantageous for higher bit widths (4–8 bits). Specifically, our work shows an average energy improvement of 1.1 to 2.3× and area saving up to 1.7× compared to integer multiply-accumulate (MAC) units, while preserving network accuracy. This research provides insights into the optimal choice of number representation based on bit-width requirements, highlighting the potential of MLB in enhancing the performance and efficiency of DNNs on embedded devices. Ali Ansarmohammadi, Reza Hojabr, Marzie Mastalizade, Najmeh Nazari, Mostafa E. Salehi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Inter-Layer Hybrid Quantization Scheme for Hardware Friendly Implementation of Embedded Deep Neural NetworksabstractCompression techniques have been widely deployed to amortize the model size and inference computations of Deep Neural Networks (DNNs), particularly for embedded systems. In this work, we propose an inter-layer approach that deploys a weight distribution aware quantization scheme (a hybrid of fixed-point and power-of-two) and multi-precision (3-bit and 4-bit) to better use heterogeneity in FPGA resources. Based on our evaluation, with similar hardware logic and memory resource usage, our proposed approach improved the throughput of the ResNet-18 network by 37% with negligible accuracy loss compared to the state-of-the-art on an embedded FPGA. Najmeh Nazari, Mostafa E. Salehi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2021 | ELC-ECG: Efficient LSTM Cell for ECG Classification Based on Quantized ArchitectureabstractLong Short-Term Memory (LSTM) is one of the most popular and effective Recurrent Neural Network (RNN) models used for sequence learning in applications such as ECG signal classification. Complex LSTMs could hardly be deployed on resource-limited bio-medical wearable devices due to the huge amount of computations and memory requirements. Binary LSTMs are introduced to cope with this problem. However, naive binarization leads to significant accuracy loss in ECG classification. In this paper, we propose an efficient LSTM cell along with a novel hardware architecture for ECG classification. By deploying 5-level binarized inputs and just 1- level binarization for weights, output, and in-memory cell activations, the delay of one LSTM cell operation is reduced 50x with about 0.004% accuracy loss in comparison with full precision design of ECG classification. Seyed Ahmad Mirsalari, Najmeh Nazari, Seyed Ali Ansarmohammadi, Sima Sinaei, Mostafa E. Salehi, Masoud Daneshtalab |
ISCAS | 5 |
| 2020 | MuBiNN: Multi-Level Binarized Recurrent Neural Network for EEG Signal ClassificationabstractRecurrent Neural Networks (RNN) are widely used for learning sequences in applications such as EEG classification. Complex RNNs could be hardly deployed on wearable devices due to their computation and memory-intensive processing patterns. Generally, reduction in precision leads much more efficiency and binarized RNNs are introduced as energy-efficient solutions. However, naive binarization methods lead to significant accuracy loss in EEG classification. In this paper, we propose a multi-level binarized LSTM, which significantly reduces computations whereas ensuring an accuracy pretty close to the full precision LSTM. Our method reduces the delay of the 3-bit LSTM cell operation 47× with less than 0.01% accuracy loss. Seyed Ahmad Mirsalari, Sima Sinaei, Mostafa E. Salehi, Masoud Daneshtalab |
ISCAS | 3 |
| 2020 | Multi-level Binarized LSTM in EEG Classification for Wearable DevicesabstractLong Short-Term Memory (LSTM) is widely used in various sequential applications. Complex LSTMs could be hardly deployed on wearable and resourced-limited devices due to the huge amount of computations and memory requirements. Binary LSTMs are introduced to cope with this problem, however, they lead to significant accuracy loss in some applications such as EEG classification which is essential to be deployed in wearable devices. In this paper, we propose an efficient multi-level binarized LSTM which has significantly reduced computations whereas ensuring an accuracy pretty close to full precision LSTM. By deploying 5-level binarized weights and inputs, our method reduces area and delay of MAC operation about $31\times and 27\times$ in 65nm technology, respectively with less than 0.01% accuracy loss. In contrast to many compute-intensive deep-learning approaches, the proposed algorithm is lightweight, and therefore, brings performance efficiency with accurate LSTM-based EEG classification to realtime wearable devices. Najmeh Nazari, Seyed Ahmad Mirsalari, Sima Sinaei, Mostafa E. Salehi, Masoud Daneshtalab |
PDP | 4 |
| 2020 | Aging-Aware Instruction-Level Statistical Dynamic Timing Analysis for Embedded ProcessorsabstractCMOS miniaturization and timing faults due to factors, such as aging, emphasize that embedded processor reliability is a major concern. Among the various aging mechanisms, negative bias temperature instability (NBTI) is encountered as the dominant factor. Techniques against NBTI are mostly based on aggressive Vdd scaling, decelerating aging at the expense of performance degradation. Traditionally, designers use conservative guard-bands to combat timing faults, leading to loss of efficiency. Some other reactive approaches use sensors, requiring hardware modification and large area and debug overheads. According to the literature, two opportunities exist to compensate for the performance loss: instruction timing slacks imposed by static timing analysis (STA) and application computational error resiliency. This article proposes an efficient estimation model for the instruction-level timing slack probability distribution function (PDF) and gives a dynamic approach for statistical timing analysis, which is used for dynamic frequency management to improve performance of both error-resilient and errorsensitive applications. To this aim, we introduce a metric called architecture timing-fault vulnerability factor, considering NBTI and Vdd effects. Simulation results show that the proposed timing slack PDF estimation model has an accuracy of about 94%, which can be used to increase throughput of error-resilient applications up to 3.2 times compared with when the traditional STA is used. Iraj Moghaddasi, Mostafa E. Salehi, Mehdi Kargahi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | TOT-Net: An Endeavor Toward Optimizing Ternary Neural NetworksabstractHigh computation demands and big memory resources are the major implementation challenges of Convolutional Neural Networks (CNNs) especially for low-power and resource-limited embedded devices. Many binarized neural networks are recently proposed to address these issues. Although they have significantly decreased computation and memory footprint, they have suffered from accuracy loss especially for large datasets. In this paper, we propose TOT-Net, a ternarized neural network with [-1, 0, 1] values for both weights and activation functions that has simultaneously achieved a higher level of accuracy and less computational load. In fact, first, TOT-Net introduces a simple bitwise logic for convolution computations to reduce the cost of multiply operations. To improve the accuracy, selecting proper activation function and learning rate are influential, but also difficult. As the second contribution, we propose a novel piece-wise activation function, and optimized learning rate for different datasets. Our findings first reveal that 0.01 is a preferable learning rate for the studied datasets. Third, by using an evolutionary optimization approach, we found novel piece-wise activation functions customized for TOT-Net. According to the experimental results, TOT-Net achieves 2.15%, 8.77%, and 5.7/5.52% better accuracy compared to XNOR-Net on CIFAR-10, CIFAR-100, and ImageNet top-5/top-1 datasets, respectively. Najmeh Nazari, Mohammad Loni, Mostafa E. Salehi, Masoud Daneshtalab, Mikael Sjödin |
DSD | 3 |
| 2019 | Instruction-Level NBTI Stress Estimation and Its Application in Runtime Aging Prediction for Embedded ProcessorsabstractLifetime reliability management of miniaturized CMOS devices continuously gets more importance with the shrinking of technology size. Neither of existing design-time solutions (like guard-banding) and runtime methods (like reactive monitoring) does efficiently address this issue; rather, proactive approaches, which use runtime aging prediction, are getting more promising to provide resiliency. Among various reliability threatening mechanisms in recent technologies, negative bias temperature instability is the dominant factor; it depends on multiple time-varying operational parameters, including temperature, supply voltage, and stress. This paper proposes an efficient instruction-level stress estimation model; accordingly, it introduces a runtime aging prediction approach for embedded processors, taking simultaneous impacts of the temperature, supply voltage, and stress variations. We propose instruction degradation factor and architecture degradation factor metrics, respectively, for fine-grained stress estimation and recurring runtime aging prediction. We also provide a simulation environment for model validation. Simulation results of several benchmarks show that the proposed stress estimation model has an accuracy of about 92%, indicating that the method is accurate enough, yet simple for runtime usage. Iraj Moghaddasi, Arash Fouman, Mostafa E. Salehi, Mehdi Kargahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Vulnerability Analysis of Adder Architectures Considering Design and Synthesis Constraints
Mostafa E. Salehi, Ali Azarpeyvand, Armin Hajaboutalebi Aboutalebi |
J. Electron. Test. | 1 |
| 2016 | Fast and accurate FPGA-based framework for processor architecture vulnerability analysis
Hoda Mahdiani, Saeed Safari, Mostafa E. Salehi |
Integr. | 3 |
| 2016 | Ultralow-Energy Variation-Aware Design: Adder Architecture StudyabstractPower consumption of digital systems is an important issue in nanoscale technologies and growth of process variation makes the problem more challenging. In this brief, we have analyzed the latency, energy consumption, and effects of process variation on different structures with respect to the design structure and logic depth to propose architectures with higher throughput, lower energy consumption, and smaller performance loss caused by process variation in application-specific integrated circuit design. We have exploited adders as different implementations of a processing unit, and propose architectural guidelines for finer technologies in subthreshold which are applicable to any other architecture. The results show that smaller computing building blocks have better energy efficiency and less performance degradation because of variation effects. In contrast, their computation throughput will be mid or less unless proper solutions, such as pipelined or parallel structures, are used. Therefore, our proposed solution to improve the throughput loss while reducing sensitivity to process variations is using simpler elements in deep pipelined designs or massively parallel structures. Hamed Dorosti, Ali Teymouri, Sied Mehdi Fakhraie, Mostafa E. Salehi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | A Fast Fault-Tolerant Architecture for Sauvola Local Image Thresholding Algorithm Using Stochastic ComputingabstractBinarization plays an important role in document image processing, particularly in degraded document images. Among all local image thresholding algorithms, Sauvola has excellent binarization performance for degraded document images. However, this algorithm is computationally intensive and sensitive to the noises from the internal computational circuits. In this paper, we present a stochastic implementation of Sauvola algorithm. Our experimental results show that the stochastic implementation of Sauvola needs much less time and area and can tolerate more faults, while consuming less power in comparison with its conventional implementation. M. Hassan Najafi, Mostafa E. Salehi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Voltage scaling and dark silicon in symmetric multicore processors
Hamid Nejatollahi, Mostafa E. Salehi |
J. Supercomput. | 2 |
| 2014 | An analytical method for reliability aware instruction set extension
Ali Azarpeyvand, Mostafa E. Salehi, Sied Mehdi Fakhraie |
J. Supercomput. | 2 |
| 2014 | Customized pipeline and instruction set architecture for embedded processing engines
Amir Yazdanbakhsh, Mostafa E. Salehi, Sied Mehdi Fakhraie |
J. Supercomput. | 2 |
| 2012 | CIVA: Custom instruction vulnerability analysis frameworkabstractThis paper describes a methodology for analyzing the vulnerability of custom instructions against the electronic faults, considering different operations and the custom instruction graph topology. Our approach enables designers to optionally constrain the operand types and also the custom functional unit structure to reach an acceptable vulnerability. We have developed a framework to evaluate the desired goal. The presented framework explores the effects of different operations and their dependencies on overall vulnerability of the custom functional units. Our experiments show that, in most cases, custom functional units with similar speed-ups in performance present different vulnerability to soft errors. Ali Azarpeyvand, Mostafa E. Salehi, Sied Mehdi Fakhraie |
DDECS | 2 |
| 2012 | Vulnerability Analysis for Custom InstructionsabstractToday circuits are becoming more vulnerable to electronic noises and reliable system design has emerged as a key challenge to embedded system design. Logic fault in terms of soft errors or transient faults are now a serious problem for embedded processors. Recent developments in customized embedded processors significantly focus on improving the performance and area of the processor by augmenting it with application specific custom functional units that implement custom instructions. This paper analyzes the effect of type, order, and bit-width of the operations of different custom instruction sub-graphs on the vulnerability of extensible processors. We have developed a framework for studying the effects of different operations and their dependencies on overall vulnerability of the custom functional units and our experiments show that, in most cases, similar custom functional units could have different vulnerabilities to soft errors. Our approach enables designers to optionally constrain the operand types and also the custom functional unit structure to reach an acceptable vulnerability. Ali Azarpeyvand, Mostafa E. Salehi, Sied Mehdi Fakhraie |
DSD | 2 |
| 2012 | Instruction set architectural guidelines for embedded packet-processing engines
Mostafa E. Salehi, Sied Mehdi Fakhraie, Amir Yazdanbakhsh |
J. Syst. Archit. | 1 |
| 2011 | Dynamic Voltage and Frequency Scheduling for Embedded Processors Considering Power/Performance TradeoffsabstractAn adaptive method to perform dynamic voltage and frequency scheduling (DVFS) for minimizing the energy consumption of microprocessor chips is presented. Instead of using a fixed update interval, the proposed DVFS system makes use of adaptive update intervals for optimal frequency and voltage scheduling. The optimization enables the system to rapidly track the workload changes so as to meet soft real-time deadlines. The technique, which can be realized with very simple hardware, is completely transparent to the application. The results of applying the method to some real application workloads demonstrate considerable power savings and fewer frequency updates compared to DVFS systems based on fixed update intervals. Mostafa E. Salehi, Mehrzad Samadi, Mehrdad Najibi, Ali Afzali-Kusha, Massoud Pedram, Sied Mehdi Fakhraie |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | Instruction reliability analysis for embedded processorsabstractAdvances in silicon technology and shrinking the feature size to nanometer scale make unreliability of nano devices the most important concern of fault-tolerant designs. Soft error analysis has been greatly aided by the concept of architectural vulnerability factor (AVF) and architecturally correct execution (ACE). In this work, we exploit the techniques of AVF analysis to introduce the instruction-level vulnerability metric for software reliability analysis. The proposed metric can be used to make judgments about the reliability of different programs on different processors with regard to architectural and compiler guidelines for improving the processor reliability. Ali Azarpeyvand, Mostafa E. Salehi, Farshad Firouzi, Amir Yazdanbakhsh, Sied Mehdi Fakhraie |
DDECS | 2 |
| 2010 | Architecture-Level Design Space Exploration of Super Scalar Microarchitecture for Network ApplicationsabstractIncreasing diversity in packet-processing applications and rapid increases in channel bandwidth lead to greater complexity in communication protocols. These factors result in larger computational loads for packet-processing engines that introduce high performance microprocessor designs as an important solution. This paper presents an exhaustive simulation for exploring the performance of instruction-level parallel super scalar processors executing packet-processing applications. Based on the simulation results, a design space exploration has been used to derive performance-efficient application-specific super scalar processor architecture based on MIPS instruction set architecture. Simple Scalar architecture toolset has been used for design space exploration and network applications have been investigated to guide the architecture exploration. The optimizations achieve up to 80% improvement in performance for representative packet-processing applications. Mostafa E. Salehi, Hamed Dorosti, Sied Mehdi Fakhraie |
DSD | 1 |
| 2009 | Quantitative analysis of packet-processing applications regarding architectural guidelines for network-processing-engine development
Mostafa E. Salehi, Sied Mehdi Fakhraie |
J. Syst. Archit. | 1 |
| 2006 | Dynamic voltage and frequency management based on variable update intervals for frequency settingabstractAn efficient adaptive method to perform dynamic voltage and frequency management (DVFM) for minimizing the energy consumption of microprocessor chips is presented. Instead of using a fixed update interval, the proposed DVFM system makes use of adaptive update intervals for optimal frequency and voltage scheduling. The optimization enables the system to rapidly track the workload changes so as to meet soft real-time deadlines. The method, which is based on introducing the concept of an effective deadline, utilizes the correlation between consecutive values of the workload. In practice because the frequency and voltage update rates are dynamically set based on variable update interval lengths, voltage fluctuations on the power network are also minimized. The technique, which may be implemented by simple hardware and is completely transparent from the application, leads to power savings of up to 60% for highly correlated workloads compared to DVFM systems based on fixed update intervals. Mehrdad Najibi, Mostafa E. Salehi, Ali Afzali-Kusha, Massoud Pedram, Sied Mehdi Fakhraie, Hossein Pedram |
ICCAD | 2 |