Nima Taherinejad

dblp:43/7291 · also Nima TaheriNejad · DBLP profile ↗
← Back
37ranked-venue papers
2as first author
27since 2021 · last 2026
0000-0002-1295-0332ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 1 first-author · 23 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Approximate Reciprocal-based Divider
Ali Ghaderi, Nima Amirafshar, Hadi Shahriar Shahhoseini, Nima Taherinejad
ISCAS4
2026 CryHop: Baby Cry Classification on Dunstan Dataset using Transfer Learning and Ensembling
Soheil Khooyooz, Loic Ricard, Jiayue Liu, Mostafa Haghi 0001, Nima Taherinejad
ISCAS5
2026 ARTS: An approximate reduced tree and segmentation-based multiplier
Mahla Salehi Sheikhali Kelayeh, Sahand Divsalar, Shaghayegh Vahdat, Nima Taherinejad
Future Gener. Comput. Syst.4
2026 LUTEA: LUT-Based Energy-Efficient Approximate Multiplier for More Sustainable Neural Network Applications
Mahla Salehi Sheikhali Kelayeh, Sahand Divsalar, Shaghayegh Vahdat, Nima Taherinejad
IEEE Trans. Computers4
2025 PRIM: Hybrid Array-Compressor Multipliers with Carry Disregard and OR-based Approximation
abstract
This paper introduces an efficient new 4:1 compressor that uses carry disregard and OR-based approximation, leading to the development of 13 approximate unsigned multipliers. The proposed multipliers, 8-bit array-comPressor oR-based carry dIsregard Multipliers (PRIM8s) demonstrate significant improvements in area, power, delay, and Power-Delay-Product (PDP) by an average of 29%, 31%, 25%, and 47%, compared to the exact multiplier. In the approximate multiplier literature, with our hardware, we establish new Pareto fronts for most criteria. The effectiveness of the proposed multipliers for noise reduction is demonstrated in an image-processing application using a low-pass Gaussian filter. On average, PRIM8s reduce power consumption and improve speed by 32.36% and 19.01% compared to the exact multiplier, while also enhancing image quality, as indicated by a 0.14% increase in Structural Similarity Index Measure (SSIM).
Nima Amirafshar, Gulafshan, Hadi Shahriar Shahhoseini, Nima Taherinejad
ISCAS4
2025 Dual-Path Cuffless PPG-Based Blood Pressure Estimation Using Conformer & Swin Transformer
abstract
This study introduces a novel dual-path deep learning framework using Photoplethysmogram (PPG) signals to address key challenges in continuous, non-invasive cuffless Blood Pressure (BP) monitoring. To this end, we introduce -for the first time- the use of two novel deep neural network architectures: Conformer-Transformer and 1D Swin Transformer. These architectures are adapted here to model both the morphological structure and rhythmic dynamics of PPG signals. This cross-domain transfer enables Arterial Blood Pressure (ABP) waveform reconstruction and significantly improves the accuracy and physiological consistency of Systolic Blood Pressure (SBP) and Diastolic Blood Pressure (DBP) estimation. Extensive experiments on two public datasets demonstrate that our methods consistently outperform mainstream baselines across multiple key metrics. Specifically, the Conformer-Transformer achieved the lowest Mean Absolute Error (MAE) of 2.979 mmHg for systolic and 1.603 mmHg for diastolic BP, improving upon previous studies by 9.6% and 8.4%, respectively, while delivering the best waveform reconstruction performance too. The Swin Transformer achieved a systolic MAE of 3.034 mmHg and a diastolic MAE of 1.714 mmHg. All experimental results conform to the British Hypertension Society (BHS) grade A and Association for the Advancement of Medical Instrumentation (AAMI) standards.
Caoyueshan Fan, Yiting Wei, Melanie Qiu, Mostafa Haghi 0001, Nima Taherinejad
IEEE J. Biomed. Health Informatics5
2024 RecogNoise: Machine-Learning-Based Recognition of Noisy Segments in Electrocardiogram Signals
abstract
Today, wearable technology is frequently used for continuous monitoring of physiological indicators in the health-care domain. However, mobile-health and wearable devices are generally used in ambulatory settings, hence vulnerable to noise. This interferes with the accuracy of Machine Learning (ML) models running on such systems and their decision-making procedures. To address this issue, we first need to identify the presence of noise. In this paper, we propose RecogNoise to detect noisy segments in Electrocardiography (ECG) recordings using heartbeat detection algorithms and ML. We evaluate our approach based on the MIT-BIH arrhythmia database and three types of noise, i.e., Electrode Motion (EM) , Baseline Wander (BW), and Muscle Artifact (MA), with different Signal to Noise Ratios (SNRs). We show that RecogNoise can detect noisy segments with an F1-score of 86.9% and an accuracy of 88.3%.
Amin Aminifar, Soheil Khooyooz, Anice Jahanjoo, Salar Shakibhamedan, Nima Taherinejad
ISCAS5
2024 High-Accuracy Stress Detection Using Wrist-Worn PPG Sensors
abstract
Stress has become a prevalent issue affecting individuals’ physical and mental well-being. Detecting stress is the first crucial step to managing it and preventing it from causing other health issues. In this paper, we present a new method to improve the performance of detecting stress, using a comfortable to wear sensor, namely Photoplethysmography (PPG), which is embedded virtually in all smartwatches. To this end, we use PPG sensor data from the publicly available wearable stress and affect detection dataset (WESAD). Using new denoising processes, segmentation methods, and key feature extract, we achieve 95.55% accuracy in detecting stress using the Support Vector Machine (SVM) algorithm. Simplifying the process alongside improved accuracy in this paper facilitates smartphone usage as a real-time stress detection, which we plan as future work.
Anice Jahanjoo, Nima Taherinejad, Amin Aminifar
ISCAS2
2024 Adaptive approximate computing in edge AI and IoT applications: A review
abstract
Recent advancements in hardware and software systems have been driven by the deployment of emerging smart health and mobility applications. These developments have modernized the traditional approaches by replacing conventional computing systems with cyber-physical and intelligent systems combining the Internet of Things (IoT) with Edge Artificial Intelligence. Despite the many advantages and opportunities of these systems within various application domains, the scarcity of energy, extensive computing needs, and limited communication must be considered when orchestrating their deployment. Inducing savings in these directions is central to the Approximate Computing (AxC) paradigm, in which the accuracy of some operations is traded off with energy, latency, and/or communication reductions. Unfortunately, the dynamics of the environments in which AxC-equipped IoT systems operate have been paid little attention. We bridge this gap by surveying adaptive AxC techniques applied to three emerging application domains, namely autonomous driving, smart sensing and wearables, and positioning, paying special attention to hardware acceleration. We discuss the challenges of such applications, how adaptive AxC can aid their deployment, and which savings it can bring based on traits of the data and devices involved. Insights arising thereof may serve as inspiration to researchers, engineers, and students active within the considered domains.
Hans Jakob Damsgaard, Antoine Grenier, Dewant Katare, Zain Taufique, Salar Shakibhamedan, Tiago Troccoli, Georgios Chatzitsompanis, Anil Kanduri, Aleksandr Ometov, Aaron Yi Ding, Nima Taherinejad, Georgios Karakonstantis, Roger F. Woods, Jari Nurmi
J. Syst. Archit.11
2024 Efficient Image Processing via Memristive-Based Approximate In-Memory Computing
abstract
Image processing algorithms continue to demand higher performance from computers. However, computer performance is not improving at the same rate as before. In response to the current challenges in enhancing computing performance, a wave of new technologies and computing paradigms is surfacing. Among these, memristors stand out as one of the most promising components due to their technological prospects and low power consumption. With efficient data storage capabilities and their ability to directly perform logical operations within the memory, they are well-suited for in-memory computation (IMC). Approximate computing emerges as another promising paradigm, offering improved performance metrics, notably speed. The tradeoff for this gain is the reduction of accuracy. In this article, we are using the stateful logic material implication (IMPLY) in the semi-serial topology and combine both the paradigms to further enhance the computational performance. We present three novel approximated adders that drastically improve speed and energy consumption with an normalized mean error distance (NMED) lower than 0.02 for most scenarios. We evaluated partially approximated Ripple carry adder (RCA) at the circuit-level and compared them to the State-of-the-Art (SoA). The proposed adders are applied in different image processing applications and the quality metrics are calculated. While maintaining acceptable quality, our approach achieves significant energy savings of 6%–38% and reduces the delay (number of computation cycles) by 5%–35%, demonstrating notable efficiency compared to exact calculations.
Fabian Seiler, Nima Taherinejad
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Accelerated Image Processing Through IMPLY-Based NoCarry Approximated Adders
abstract
As the demand for computational power increases drastically, traditional solutions to address those needs struggle to keep up. Consequently, there has been a proliferation of alternative computing paradigms aimed at tackling this disparity. Approximate Computing (AxC) has emerged as a modern way of improving speed, area efficiency, and energy consumption in error-resilient applications such as image processing or machine learning. The trade-off for these enhancements is the loss in accuracy. From a technology point of view, memristors have garnered significant attention due to their low power consumption and inherent non-volatility that makes them suitable for In-Memory Computation (IMC). Another computing paradigm that has risen to tackle the aforementioned disparity between the demand growth and performance improvement. In this work, we leverage a memristive stateful in-memory logic, namely Material Implication (IMPLY). We investigate advanced adder topologies within the context of AxC, aiming to combine the strengths of both of these novel computing paradigms. We present two approximated algorithms for each IMPLY based adder topology. When embedded in an Ripple Carry Adder (RCA), they reduce the number of steps by$6\%-54\%$and the energy consumption by$7\%-54\%$compared to the corresponding exact full adders. We compare our work to State-of-the-Art (SoA) approximations at circuit-level, which improves the speed and energy efficiency by up to$72\%$and$34\%$, while lowering the Normalized Median Error Distance (NMED) by up to$81\%$. We evaluate our adders in four common image processing applications, for which we introduce two new test datasets as well. When applied to image processing, our proposed adders can reduce the number of steps by up to$60\%$and the energy consumption by up to$57\%$, while also improving the quality metrics over the SoA in most cases.
Fabian Seiler, Nima Taherinejad
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 ACE-CNN: Approximate Carry Disregard Multipliers for Energy-Efficient CNN-Based Image Classification
abstract
This paper presents the design and development of Signed Carry Disregard Multiplier (SCDM8), a family of signed approximate multipliers tailored for integration into Convolutional Neural Networks (CNNs). Extensive experiments were conducted on popular pre-trained CNN models, including VGG16, VGG19, ResNet101, ResNet152, MobileNetV2, InceptionV3, and ConvNeXt-T to evaluate the trade-off between accuracy and approximation. The results demonstrate that ACE-CNN outperforms other configurations, offering a favorable balance between accuracy and computational efficiency. In our experiments, when applied to VGG16, SCDM8 achieves an average reduction in power consumption of 35% with a marginal decrease in accuracy of only 1.5%. Similarly, when incorporated into ResNet152, SCDM8 yields an energy saving of 42% while sacrificing only 1.8% in accuracy. ACE-CNN provides the first approximate version of ConvNeXt which yields up to 72% energy improvement at the price of less than only 1.3% Top-1 accuracy. These results highlight the suitability of SCDM8 as an approximation method across various CNN models. Our analysis shows that the ACE-CNN outperforms state-of-the-art approaches in accuracy, energy efficiency, and computation precision for image classification tasks in CNNs. Our study investigated the resiliency of CNN models to approximate multipliers, revealing that ResNet101 demonstrated the highest resiliency with an average difference in the accuracy of 0.97%, whereas LeNet5 Inspired-CNN exhibited the lowest resiliency with an average difference of 2.92%. These findings aid in selecting energy-efficient approximate multipliers for CNN-based systems, and contribute to the development of energy-efficient deep learning systems by offering an effective approximation technique for multipliers in CNNs. The proposed SCDM8 family of approximate multipliers opens new avenues for efficient deep learning applications, enabling significant energy savings with virtually no loss in accuracy.
Salar Shakibhamedan, Nima Amirafshar, Ahmad Sedigh Baroughi, Hadi Shahriar Shahhoseini, Nima Taherinejad
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Stochastic Computing for Reliable Memristive In-Memory Computation
abstract
In-Memory Computing (IMC) is a promising computing paradigm to accelerate Big Data applications. It reduces the data movement between memory and processing units, and provides massive parallelism. Memristive technology is one of the promising technologies for IMC. This emerging technology, however, is still in evolution, facing practical challenges. Memristive memories are prone to softerror while storing the data and during computations. The traditional binary encoding commonly used in memristive IMC is highly sensitive to soft-errors, which makes developing reliable memristive IMC more challenging. Stochastic Computing (SC) is a re-emerging computing paradigm that is highly robust against soft-errors as any bit flip leads to only a least significant bit error. In this work, we study SC as a solution to increase the reliability of memristive IMC. We investigate how and to what extent SC may address or improve the reliability issues of current memristive technology, and memristive IMC. We also evaluate the characteristics yielded by memristive stochastic IMC and compare them with those of the traditional reliability techniques.
Mohsen Riahi Alam, M. Hassan Najafi, Nima Taherinejad, Mohsen Imani, Lu Peng 0001
ACM Great Lakes Symposium on VLSI3
2023 Carry Disregard Approximate Multipliers
abstract
Several challenges in improving the performance of computing systems have given rise to emerging computing paradigms. One of these paradigms is approximate computing. Many applications require different levels of accuracy and are error-tolerance to a certain degree. Approximate computations can reduce the calculation complexities significantly and thus improve the performance. Here, we propose a methodology for designing approximate N-bit array multipliers based on carry disregarding. We evaluate and analyze the proposed multipliers both experimentally and theoretically. The proposed 8-bit multipliers, compared to the exact multiplier, reduce the critical path delay, power consumption, and area by 29%, 29%, and 30%, on average. Compared to the existing approximate array architectures in the literature, they have improved 14.3%, 22.8%, and 26.4%, respectively. Compared to the exact 16-bit multiplier, the proposed 16-bit multipliers have reduced the delay, power consumption, and area by 35%, 24%, and 23% on average. In an image processing application, we have also demonstrated the applicability of a wide range of proposed multipliers, which have Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) over 30 dB and 94%, respectively.
Nima Amirafshar, Ahmad Sadigh Baroughi, Hadi Shahriar Shahhoseini, Nima Taherinejad
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 An Approximate Carry Disregard Multiplier with Improved Mean Relative Error Distance and Probability of Correctness
abstract
Nowadays, a wide range of applications can tolerate certain computational errors. Hence, approximate computing has become one of the most attractive topics in computer architecture. Reducing accuracy in computations in a premeditated and appropriate manner reduces architectural complexities, and as a result, performance, power consumption, and area can improve significantly. This paper proposes a novel approximate multiplier design. The proposed design has been implemented using 45 nm CMOS technology and has been extensively evaluated. Compared to existing approximate architectures, the proposed approximate multiplier has higher accuracy. It also achieves better results in critical path delay, power consumption, and area up to 47.54 %, 75.24%, and 92.49%, respectively. Compared to the precise multipliers, our evaluations show that the critical path delay, power consumption, and area have been improved by 39%, 18%, and 6 %, respectively.
Nima Amirafshar, Ahmad Sadigh Baroughi, Hadi Shahriar Shahhoseini, Nima Taherinejad
DSD4
2022 AxE: An Approximate-Exact Multi-Processor System-on-Chip Platform
abstract
Due to the ever-increasing complexity of computing tasks, emerging computing paradigms that increase efficiency, such as approximate computing, are gaining momentum. However, so far, the majority of proposed solutions for hardware-based approximation have been application-specific and/or limited to smaller units of the computing system and require engineering effort for integration into the rest of the system. In this paper, we present Approximate and Exact Multi-Processor system-on-chip (AxE) platform. AxE is the first general-purpose approximate Multi-Processor System-on-Chip (MPSoC). AxE is a heterogeneous RISC-V platform with exact and approximate cores that allows exploring hardware approximation for any application and using software instructions. Using the full capacity of an entire MPSoC, especially a heterogeneous one such as AxE, is an increasingly challenging problem. Therefore, we also propose a task mapping method for running exact and approximable applications on AxE. That is a mixed task mapping, in which applications are viewed as a set of tasks that can be run independently on different processors with different capabilities (exact or approximate). We evaluated our proposed method on AxE and reached a 32% average execution speed-up and 21% energy consumption saving with an average of 99.3% accuracy on three mixed workloads. We also ran a sample image processing application, namely gray-scale filter, on AxE and will present its results.
Ahmad Sadigh Baroughi, Sini Huemer, Hadi Shahriar Shahhoseini, Nima Taherinejad
DSD4
2022 Sound Source Localization Using Stochastic Computing
abstract
Stochastic computing (SC) is an alternative computing paradigm that processes data in the form of long uniform bit-streams rather than conventional compact weighted binary numbers. SC is fault-tolerant and can compute on small, efficient circuits, promising advantages over conventional arithmetic for smaller computer chips. SC has been primarily used in scientific research, not in practical applications. Digital sound source localization (SSL) is a useful signal processing technique that locates speakers using multiple microphones in cell phones, laptops, and other voice-controlled devices. SC has not been integrated into SSL in practice or theory. In this work, for the first time to the best of our knowledge, we implement an SSL algorithm in the stochastic domain and develop a functional SC-based sound source localizer. The developed design can replace the conventional design of the algorithm. The practical part of this work shows that the proposed stochastic circuit does not rely on conventional analog-to-digital conversion and can process data in the form of pulse-width-modulated (PWM) signals. The proposed SC design consumes up to 39% less area than the conventional baseline design. The SC-based design can consume less power depending on the computational accuracy, for example, 6% less power consumption for 3-bit inputs. The presented stochastic circuit is not limited to SSL and is readily applicable to other practical applications such as radar ranging, wireless location, sonar direction finding, beamforming, and sensor calibration.
Peter Schober, Seyedeh Newsha Estiri, Sercan Aygün, Nima Taherinejad, M. Hassan Najafi
ICCAD4
2022 Approximate In-Memory Computing using Memristive IMPLY Logic and its Application to Image Processing
abstract
Approximate computing is a new way of performing calculations in digital systems. By applying this method, performance metrics, e.g., speed, are improved, but in return for this, the accuracy of the calculations is reduced. Memristors are electrical elements that can be used to perform logical calculations along with data storage. This makes memristors a good choice for In-Memory Computation (IMC). IMPLY logic is the first stateful logic proposed for memristive IMC. Approximate computing in memory, particularly using memristive stateful logic, has not been explored yet. In this paper, we combine these two concepts and propose a novel algorithm for serial IMPLY-based adders to implement an approximate full-adder. The proposed approximate full-adder was assessed in an image processing application, and image quality metrics like Peak Signal to Noise Ratio (PSNR) were calculated. In addition, different error quality metrics like Error Distance (ED) and Mean ED (MED) were assessed. Our study shows that the proposed method can achieve up to 40% improvement whereas maintaining the introduced error in an acceptable range (i.e., a PSNR above 32.4).
Seyed Erfan Fatemieh, Mohammad Reza Reshadinezhad, Nima Taherinejad
ISCAS3
2022 Sorting in Memristive Memory
abstract
Sorting data is needed in many application domains. Traditionally, the data is read from memory and sent to a general-purpose processor or application-specific hardware for sorting. The sorted data is then written back to the memory. Reading/writing data from/to memory and transferring data between memory and processing unit incur significant latency and energy overhead. In this work, we develop the first architectures for in-memory sorting of data to the best of our knowledge. We propose two architectures. The first architecture is applicable to the conventional format of representing data, i.e., weighted binary radix. The second architecture is proposed for developing unary processing systems, where data is encoded as uniform unary bit-streams. As we present, each of the two architectures has different advantages and disadvantages, making one or the other more suitable for a specific application. However, the common property of both is a significant reduction in the processing time compared to prior sorting designs. Our evaluations show on average 37 × and 138× energy reduction for binary and unary designs, respectively, compared to conventional CMOS off-memory sorting systems in a 45 nm technology. We designed a 3×3 and a 5×5 Median filter using the proposed sorting solutions, which we used for processing 64×64 pixel images. Our results show a reduction of 14× and 634× in energy and latency, respectively, with the proposed binary, and 5.6× and 152×10 3 in energy and latency with the proposed unary approach compared to those of the off-memory binary and unary designs for the 3 × 3 Median filtering system.
Mohsen Riahi Alam, M. Hassan Najafi, Nima Taherinejad
ACM J. Emerg. Technol. Comput. Syst.3
2022 Confidence-Enhanced Early Warning Score Based on Fuzzy Logic
abstract
Abstract Cardiovascular diseases are one of the world’s major causes of loss of life. The vital signs of a patient can indicate this up to 24 hours before such an incident happens. Healthcare professionals use Early Warning Score (EWS) as a common tool in healthcare facilities to indicate the health status of a patient. However, the chance of survival of an outpatient could be increased if a mobile EWS system would monitor them during their daily activities to be able to alert in case of danger. Because of limited healthcare professional supervision of this health condition assessment, a mobile EWS system needs to have an acceptable level of reliability - even if errors occur in the monitoring setup such as noisy signals and detached sensors. In earlier works, a data reliability validation technique has been presented that gives information about the trustfulness of the calculated EWS. In this paper, we propose an EWS system enhanced with the self-aware property confidence, which is based on fuzzy logic. In our experiments, we demonstrate that - under adverse monitoring circumstances (such as noisy signals, detached sensors, and non-nominal monitoring conditions) - our proposed Self-Aware Early Warning Score (SA-EWS) system provides a more reliable EWS than an EWS system without self-aware properties.
Maximilian Götzinger, Arman Anzanpour, Iman Azimi, Nima Taherinejad, Axel Jantsch, Amir-Mohammad Rahmani, Pasi Liljeberg
Mob. Networks Appl.4
2022 Detection and Removal of Motion Artifacts in PPG Signals
abstract
Abstract With the rise of wearable devices, which integrate myriad of health-care and fitness procedures into daily life, a reliable method for measuring various bio-signals in a daily setup is more desired than ever. Many of these physiological parameters, such as Heart rate (HR) and Respiratory Rate (RR), are extracted indirectly and using other signals such as Photoplethysmograph (PPG). Part of the reason is that in some cases, such as RR measurements, the devices which directly measure them are cumbersome to wear and thus, rather impractical. On the other hand, signals, such as PPG from which the RR can be extracted, are not very clean. This poses a challenge on reliable extraction of these metrics. The most important problem is that they are corrupted by motion artifacts. In this paper, we review the state of the art algorithms which are used to detect and filter motion artifacts in PPG signals and compare them in terms of their performance. The insight provided by this paper can help the scientists and engineers to obtain a better understanding of the field and be able to use the most suitable technique for their work, or come up with innovative solutions based on existing ones.
David Pollreisz, Nima Taherinejad
Mob. Networks Appl.2
2022 Mobile Health Technology: From Daily Care and Pandemics to their Energy Consumption and Environmental Impact
Nima Taherinejad, Paolo Perego, Amir-Mohammad Rahmani
Mob. Networks Appl.1
2022 High-Accuracy Multiply-Accumulate (MAC) Technique for Unary Stochastic Computing
abstract
Multiply-accumulate (MAC) operations are common in data processing and machine learning but costly in terms of hardware usage. Stochastic Computing (SC) is a promising approach for low-cost hardware design of complex arithmetic operations such as multiplication. Computing with deterministic unary bit-streams (defined as bit-streams with all 1s grouped together at the beginning or end of a bit-stream) has been recently suggested to improve the accuracy of SC. Conventionally, SC designs use multiplexer (MUX) units or OR gates to accumulate data in the stochastic domain. MUX-based addition suffers from scaling of data and OR-based addition from inaccuracy. This work proposes a novel technique for MAC operation on unary bit-streams that allows exact, non-scaled addition of multiplication results. By introducing a relative delay between the products, we control correlation between bit-streams and eliminate OR-based addition error. We evaluate the accuracy of the proposed technique compared to the state-of-the-art MAC designs. After quantization, the proposed technique demonstrates at least 37% and up to 100% decrease of the mean absolute error for uniformly distributed random input values, compared to traditional OR-based MAC designs. Further, we demonstrate that the proposed technique is practical and evaluate area, power and energy of three possible implementations.
Peter Schober, M. Hassan Najafi, Nima Taherinejad
IEEE Trans. Computers3
2021 MELODI: An Online Platform for Mass Education of Digital Design - HDL to Remote FPGA
abstract
Learning and teaching digital hardware design involves significant efforts on both sides, teachers and students. Hardware Description Languages (HDLs) simplify the design process, where the created designs can be tested either in simulators or using real hardware. For the latter, Field Programmable Gate Arrays (FPGAs) play a crucial role in facilitating and speeding up the prototyping process. Mass E-Learning of design, test, and prototyping Digital hardware (MELODI) is developed for scalable teaching of HDL to a large number of students. The primary goals are to a) minimize the requirements for students and b) reduce the resources required at the university. MELODI provides a complete HDL workflow, including real remote hardware prototyping on FPGAs without the need for any tool at the students’ side.
Friedrich Bauer, Felix Braun, Daniel Hauer, Axel Jantsch, Markus D. Kobelrausch, Martin Mosbeck, Nima Taherinejad, Philipp-Sebastian Vogt
FPL7
2021 A Fast Line Segment Detector Using Approximate Computing
abstract
The Line Segment Detector (LSD) algorithm is an underlying step of many image processing systems. Hence, its performance has a significant on the upper layers using the detected line segment for various purposed. In this paper, we propose a fast LSD algorithm. This method approximates several floating point operations, including the logarithmic Gamma function, by a series of lookup table searches. Due to the simplicity of such approximation (lookup table search) compared to the naïve implementation (calculation-based), this method is considerably faster. The proposed method has implications on reduction of the necessary efforts to implement and enhancement of the performance of the LSD hardware accelerators. Our experiments show that the proposed method reduces the run-time of the algorithm by 13% on average, with no considerable quality loss in the detection results. This improvement further propagates through other image processing algorithms using LSD.
Christoph Ossimitz, Nima Taherinejad
ISCAS2
2021 UBAR: User- and Battery-aware Resource Management for Smartphones
abstract
Smartphone users require high Battery Cycle Life (BCL) and high Quality of Experience (QoE) during their usage. These two objectives can be conflicting based on the user preference at run-time. Finding the best trade-off between QoE and BCL requires an intelligent resource management approach that considers and learns user preference at run-time. Current approaches focus on one of these two objectives and neglect the other, limiting their efficiency in meeting users’ needs. In this article, we present UBAR, User- and Battery-aware Resource management, which considers dynamic workload, user preference, and user plug-in/out pattern at run-time to provide a suitable trade-off between BCL and QoE. UBAR personalizes this trade-off by learning the user’s habits and using that to satisfy QoE, while considering battery temperature and State of Charge (SOC) pattern to maximize BCL. The evaluation results show that UBAR achieves 10% to 40% improvement compared to the existing state-of-the-art approaches.
Elham Shamsa, Alma Pröbstl, Nima Taherinejad, Anil Kanduri, Samarjit Chakraborty, Amir-Mohammad Rahmani, Pasi Liljeberg
ACM Trans. Embed. Comput. Syst.3
2021 SIXOR: Single-Cycle In-Memristor XOR
abstract
With the fast approach of the end of silicon scaling and existing problems, such as the Von-Neumann bottleneck, alternative computing paradigms are in demand. In-memory computation (IMC) is one of the most promising solutions, and memristive technology is one of the best platforms for that purpose. Many logic families have been proposed to enable memristive IMC, among which stateful logic family stands out due to its minimal power consumption and simplicity. In this work, to complement existing works, we propose the first stateful crossbar-compatible XOR atomic logic operation that requires only one cycle for its completion, which is two times faster than the current minimum required time for performing XOR (which is two cycles) using other atomic operations in comparable memristive stateful logic families. We show that, in an example case of an adder, by taking advantage of the proposed single-cycle in-memristor XOR (SIXOR), up to 4.5× speedup can be achieved compared to other SoA stateful adders. The gained speed-up scales up in more complex systems and calculations that use XOR.
Nima Taherinejad
IEEE Trans. Very Large Scale Integr. Syst.1
2020 Exact In-Memory Multiplication Based on Deterministic Stochastic Computing
abstract
Memristors offer the ability to both store and process data in memory, eliminating the overhead of data transfer between memory and processing unit. For data-intensive applications, developing efficient in-memory computing methods is under investigation. Stochastic computing (SC), a paradigm offering simple execution of complex operations, has been used for reliable and efficient multiplication of data in-memory. Current SC-based in-memory methods are incapable of producing accurate results. This work, to the best of our knowledge, develops the first accurate SC-based in-memory multiplier. For logical operations, we use Memristor-Aided Logic (MAGIC), and to generate bit-streams, we propose a novel method, which takes advantage of the intrinsic properties of memristors. The proposed design improves the speed and reduces the memory usage and energy consumption compared to the State-of-the-Art (SoA) accurate in-memory fixed-point and off-memory SC multipliers.
Mohsen Riahi Alam, M. Hassan Najafi, Nima Taherinejad
ISCAS3
2020 A Semiparallel Full-Adder in IMPLY Logic
abstract
Passive implementation of memristors has led to several innovative works in the field of electronics. Despite being primarily a candidate for memory applications, memristors have proven to be beneficial in several other circuits and applications as well. One of the use cases is the implementation of digital circuits such as adders. Among several logic implementations using memristors, IMPLY logic is one of the promising candidates and one of the first stateful logics proposed. In this logic, the result of a IMPLY b (a → b) is stored in b and it is always true (1) except for when a=0 and b=1. Given the intrinsic difference between IMPLY and Boolean logic, conventional operations such as binary addition need to be appropriated for implementation using IMPLY. This appropriation has two main constituents; the topology or structure of the adder and the algorithm performing the addition. In this paper, we proposed a new architecture for a digital full-adder, which is up to 41% faster than existing IMPLY-based serial designs while requiring up to 78% less area (memristors) compared to the existing parallel design. In addition to that, we present a review of the state-of-the-art in IMPLY-based adders, which appeared in the literature during the last five years and discuss their advantages or disadvantages. To be able to compare these designs in a generic condition, we define a Figure of Merit (FoM), in which both the number of memristors (nm) and the number of steps (ns) are equally important. However, we must bear in mind that m (which translates to area and consequently cost) and s (which represents the speed of the full-adder) carry different weights in designs with different constraints. IMPLY-based adders can be divided into two major micro-architectures; i) serial, in which each bit is processed after the other, ii) parallel, in which bits are added in parallel as long as possible. In the first category, relatively speaking, the number of memristors is low and the number of steps is high. This is quite the opposite for the second category. In serial adders, [1] is one of the very first efficient designs, which needs 3n+3 memristors and 29n steps for an n-bit addition. In [2], a different algorithm led to a smaller number of steps, i.e., 23n. These works were improved in terms of FoM, thanks to a new approach proposed in [3], where input memristors were used for storing the output result too. Thus, the new algorithm managed to reduce m and consequently increase FoM, without needing any structural changes. However, this approach, which was adopted by [4] as well, is not suitable for all applications, in particular when the respective input is needed for further operations and its value must not be lost. The parallel design proposed in [1] was for a long time uncontested until recently the authors in [4] proposed an approach, which reduced nmand thus improved FoM by applying a minimal change to the structure. Lastly, in our newer topology, namely semiparallel topology, by having two serial sections in parallel, we managed to considerably reduce the number of steps (compared to serial adders), without needing any additional memristors. The FoM of this design is even better than the parallel approach proposed in [4]. Compared to other serial works our design needs three extra switches, which is negligible compared to 2n and n switches required in the parallel approaches.
Shokat Ganjeheizadeh Rohani, Nima Taherinejad, David Radakovits
ISCAS2
2020 Self-aware Cyber-Physical Systems
abstract
In this article, we make the case for the new class of Self-aware Cyber-physical Systems. By bringing together the two established fields of cyber-physical systems and self-aware computing, we aim at creating systems with strongly increased yet managed autonomy, which is a main requirement for many emerging and future applications and technologies. Self-aware cyber-physical systems are situated in a physical environment and constrained in their resources, and they understand their own state and environment and, based on that understanding, are able to make decisions autonomously at runtime in a self-explanatory way. In an attempt to lay out a research agenda, we bring up and elaborate on five key challenges for future self-aware cyber-physical systems: (i) How can we build resource-sensitive yet self-aware systems? (ii) How to acknowledge situatedness and subjectivity? (iii) What are effective infrastructures for implementing self-awareness processes? (iv) How can we verify self-aware cyber-physical systems and, in particular, which guarantees can we give? (v) What novel development processes will be required to engineer self-aware cyber-physical systems? We review each of these challenges in some detail and emphasize that addressing all of them requires the system to make a comprehensive assessment of the situation and a continual introspection of its own state to sensibly balance diverse requirements, constraints, short-term and long-term objectives. Throughout, we draw on three examples of cyber-physical systems that may benefit from self-awareness: a multi-processor system-on-chip, a Mars rover, and an implanted insulin pump. These three very different systems nevertheless have similar characteristics: limited resources, complex unforeseeable environmental dynamics, high expectations on their reliability, and substantial levels of risk associated with malfunctioning. Using these examples, we discuss the potential role of self-awareness in both highly complex and rather more simple systems, and as a main conclusion we highlight the need for research on above listed topics.
Kirstie L. Bellman, Christopher Landauer, Nikil Dutt, Lukas Esterle, Andreas Herkersdorf, Axel Jantsch, Nima Taherinejad, Peter R. Lewis 0001, Marco Platzner, Kalle Tammemäe
ACM Trans. Cyber Phys. Syst.7
2020 A Semiparallel Full-Adder in IMPLY Logic
abstract
Passive implementation of memristors has led to several innovative works in the field of electronics. Despite being primarily a candidate for memory applications, memristors have proven to be beneficial in several other circuits and applications as well. One of the use cases is the implementation of digital circuits such as adders. Among several logic implementations using memristors, IMPLY logic is one of the promising candidates. In this brief, we present a new architecture for a digital full-adder, which is up to 41% faster than existing IMPLY-based serial designs while requiring up to 78% less area (memristors) compared to the existing parallel design.
Shokat Ganjeheizadeh Rohani, Nima Taherinejad, David Radakovits
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Applicability of Context-Aware Health Monitoring to Hydraulic Circuits
abstract
Monitoring is an important aspect of operation and maintenance in virtually every industrial system. However, the extent and methods of monitoring vastly vary in different systems, from fully automated to fully manual. One of the challenges of automated monitoring is the tediousness of, and the extent of engineering time and effort required to develop necessary models or machine learning algorithms for the units to be monitored. Model-free monitoring, on the other hand, can save resources and efforts substantially. However, more often than not they have a very limited scope and application. Such a system is needed, for example, to monitor entire Heating, Ventilation and Air Conditioning (HVAC) systems, consisting of different types of sensors such as temperature, pressure, humidity or flow sensors. Recently, we proposed the Context-Aware Health Monitoring (CAH) system for model-free monitoring of any injective-function black-box, and it was tested successfully on an AC motor. In this paper, we evaluate the CAH system for an entirely different industrial use-case, that is, a hydraulic circuit. The results show the potential for considerable benefits in monitoring HVAC systems. Moreover, in the light of applying CAH to different use-cases which may potentially need a different setup of parameters, we performed a sensitivity analysis on the values of different parameters in the system. The results show the robustness of CAH with regard to the values of these parameters.
Maximilian Götzinger, Edwin Willegger, Nima Taherinejad, Axel Jantsch, Thilo Sauter, Thomas Glatzl, P. Lilieberg
IECON3
2018 Enhancement of Classification of Small Data Sets Using Self-awareness - An Iris Flower Case-Study
abstract
In big-data (Deep) Neural Network (NN) algorithm is often used for classification. However, such a massive mine of data is not always available and a shortage of training data can significantly deteriorate the performance of NNs and other classifiers. Therefore, we propose a self-aware multiple classifier system suitable for “Small-Data” cases. This algorithm uses self-awareness to switch between classifiers to improve its performance. We tested the algorithm for the classification of iris flower species using the Iris standard database. Compared to NN, our algorithm showed up to 17% classification success rate improvement with up to 10 times smaller standard deviation.
Hedyeh A. Kholerdi, Nima Taherinejad, Axel Jantsch
ISCAS2
2017 Self-awareness in remote health monitoring systems using wearable electronics
abstract
In healthcare, effective monitoring of patients plays a key role in detecting health deterioration early enough. Many signs of deterioration exist as early as 24 hours prior having a serious impact on the health of a person. As hospitalization times have to be minimized, in-home or remote early warning systems can fill the gap by allowing in-home care while having the potentially problematic conditions and their signs under surveillance and control. This work presents a remote monitoring and diagnostic system that provides a holistic perspective of patients and their health conditions. We discuss how the concept of self-awareness can be used in various parts of the system such as information collection through wearable sensors, confidence assessment of the sensory data, the knowledge base of the patient's health situation, and automation of reasoning about the health situation. Our approach to self-awareness provides (i) situation awareness to consider the impact of variations such as sleeping, walking, running, and resting, (ii) system personalization by reflecting parameters such as age, body mass index, and gender, and (iii) the attention property of self-awareness to improve the energy efficiency and dependability of the system via adjusting the priorities of the sensory data collection. We evaluate the proposed method using a full system demonstration.
Arman Anzanpour, Iman Azimi, Maximilian Götzinger, Amir-Mohammad Rahmani, Nima Taherinejad, Pasi Liljeberg, Axel Jantsch, Nikil Dutt
DATE5
2017 SAMBA: A self-aware health monitoring architecture for distributed industrial systems
abstract
In the context of Industry 4.0, constantly evolving shop floors generate the need for a highly adaptive and autonomous automation system with lean maintenance, minimum downtime, maximum reliability, and resilience. Future Manufacturing Execution Systems (MESs) will be more complex and dynamic as well as distributed physically and logically. This makes it very difficult, if not impossible, for the conventional centralized architectures to effectively control these vibrant Cyber-Physical Production Systems (CPPSs). To address these issues, we propose Self-Aware health Monitoring and Bio-inspired coordination for distributed Automation systems (SAMBA), an architecture which tackles these challenges. SAMBA increases the ability of the system to intelligently adapt to rapidly changing environment and conditions of future CPPSs.
Lydia Siafara, Hedyeh A. Kholerdi, Aleksey Bratukhin, Nima Taherinejad, Alexander Wendt, Axel Jantsch, Albert Treytl, Thilo Sauter
IECON4
2016 Driver's drowsiness detection using an enhanced image processing technique inspired by the human visual system
abstract
Unfit drivers are the cause of tens of thousands of incidents on the roads which lead to injuries and deaths. Therefore, it is very important to take preventive measures against such incidents. One of the unfit driving conditions is driving while being drowsy. Using image processing techniques, drowsiness of the driver could be detected and hence such incidents could be prevented. In this work, inspired by how images are processed by the human visual system, an enhancement for driver's drowsiness detection is suggested. Furthermore, to improve the robustness of the drowsiness detection system, the mechanism for using energy levels in frames is changed. Lastly, a better decision making process is proposed. To measure the merit of the system, it is applied to a set of drivers' data. Test results show that using the proposed system, success rate of the drowsiness detection system is 90%.
Hedyeh A. Kholerdi, Nima Taherinejad, Reza Ghaderi, Yasser Baleghi 0001
Connect. Sci.2
2009 New robust and efficient ant colony algorithms: Using new interpretation of local updating process
Hossein Miar Naimi, Nima Taherinejad
Expert Syst. Appl.2