EDBT 2026 Demo / reviewers in the wild / expert
Sied Mehdi Fakhraie
dblp:f/SeidMehdiFakhraie · also Seid Mehdi Fakhraie
· DBLP profile ↗
41ranked-venue papers
1as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28Artificial intelligence and machine learning · 9 · 1 first-authorComputer networks · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Energy-efficient computing · 77% Electronic design automation · 23% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing
clock gating |
0.2 | 1 | 2015 | Power Efficient High-Level Synthesis by Centralized and Fine-Grained Clock Gating · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Energy-efficient computing
dynamic power reduction |
0.2 | 1 | 2015 | Power Efficient High-Level Synthesis by Centralized and Fine-Grained Clock Gating · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Electronic design automation
high-level synthesis |
0.2 | 1 | 2015 | Power Efficient High-Level Synthesis by Centralized and Fine-Grained Clock Gating · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Energy-efficient computing
power management |
0.2 | 1 | 2015 | Power Efficient High-Level Synthesis by Centralized and Fine-Grained Clock Gating · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.1 | 1 | 2015 | Power Efficient High-Level Synthesis by Centralized and Fine-Grained Clock Gating · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Lifetime improvement by exploiting aggressive voltage scaling during runtime of error-resilient applications
Farzaneh Nakhaee, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Sied Mehdi Fakhraie, Hamed Dorosti |
Integr. | 5 |
| 2016 | Reliability aware throughput management of chip multi-processor architecture via thread migration
Fatemeh Pouyan, Ali Azarpeyvand, Saeed Safari, Sied Mehdi Fakhraie |
J. Supercomput. | 4 |
| 2016 | Ultralow-Energy Variation-Aware Design: Adder Architecture StudyabstractPower consumption of digital systems is an important issue in nanoscale technologies and growth of process variation makes the problem more challenging. In this brief, we have analyzed the latency, energy consumption, and effects of process variation on different structures with respect to the design structure and logic depth to propose architectures with higher throughput, lower energy consumption, and smaller performance loss caused by process variation in application-specific integrated circuit design. We have exploited adders as different implementations of a processing unit, and propose architectural guidelines for finer technologies in subthreshold which are applicable to any other architecture. The results show that smaller computing building blocks have better energy efficiency and less performance degradation because of variation effects. In contrast, their computation throughput will be mid or less unless proper solutions, such as pipelined or parallel structures, are used. Therefore, our proposed solution to improve the throughput loss while reducing sensitivity to process variations is using simpler elements in deep pipelined designs or massively parallel structures. Hamed Dorosti, Ali Teymouri, Sied Mehdi Fakhraie, Mostafa E. Salehi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Multiplierless filter-bank based multicarrier system by using canonical signed digit representationabstractABSTRACT The filter‐bank based multicarrier (FBMC) system is a candidate for designing the physical layer of a cognitive radio because of its spectral efficiency and the spectral containment. The main drawback of such a system compared with orthogonal frequency division multiplexing systems is its high computation complexity, because each subcarrier is shaped by a non‐rectangular prototype filter. Although poly‐phase decomposition is suggested to decrease the sampling rate of filtering, the number of filtering operations (multiplications) is dramatically increased by the growing number of subcarriers and the amount of the desired spectral containment. Hence, hardware implementation of the FBMC system faces challenges such as high electrical power consumption and large silicon area occupation. In order to reduce computational complexity, a multiplierless filter design based on the canonical signed digit (CSD) representation is proposed. In this technique, at first, a prototype filter is designed. Then a pre‐optimization step is employed to adapt the prototype filter coefficients to system objectives and produce an enriched initial seed for genetic algorithm (GA) optimization. Finally, a customized GA is employed to jointly optimize and synthesize filter coefficients into the finite precision CSD representation so that the system objectives such as intersymbol‐interference, interchannel‐interference, and stop‐band attenuation remain unchanged against the full‐precision representation of coefficients. Copyright © 2014 John Wiley & Sons, Ltd. Mohammad Aliasgari, Yeganeh M. Marghi, Mohammadreza Baharani, Sied Mehdi Fakhraie |
Wirel. Commun. Mob. Comput. | 4 |
| 2015 | Power Efficient High-Level Synthesis by Centralized and Fine-Grained Clock GatingabstractNowadays, power is a primary concern in digital circuits and clock distribution networks are particularly a significant power consumer. Therefore, clock gating is an effective technique in saving dynamic power by reducing the switching activities. In this paper, we propose a centralized and fine-grained microarchitecture-level clock gating for low power hardware accelerators which are automatically designed by high-level synthesis (HLS) tool. The basic principium of our idea is not to use any extra computation for generating clock enabled signals and exploit exiting signals of finite state machine for controlling the datapath clock network. After determining the current state in finite state machine, clock sub-tree of current state is enabled and the other sub-trees are disabled with a slight increase in circuit area. Our approach is implemented within an HLS design flow for automatic low power hardware accelerator generation in application specific integrated circuit design. Experimental results are obtained on a set of representative benchmark programs. Depending on the circuit size and number of registers, it is shown that 47%–86% reduction in power dissipation is observed. Mohsen Riahi Alam, Mostafa Ersali Salehi Nasab, Sied Mehdi Fakhraie |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | A 256-kb 9T Near-Threshold SRAM With 1k Cells per Bitline and Enhanced Write and Read OperationsabstractIn this paper, we present a new 9T SRAM cell that has good write ability and improves read stability at the same time. Simulation results show that the proposed design increases read static noise margin and ION/IOFF of read path by 219% and 113%, respectively, at supply voltage of 300-mV over conventional 6T SRAM cell in a 90-nm CMOS technology. The proposed design lets us reduce the minimum operating voltage of SRAM (VDDmin) to 350 mV, whereas conventional 6T SRAM cannot operate successfully with an acceptable failure rate at supply voltages below 725 mV. We also compared our design with three other SRAM cells from recent literature. To verify the proposed design, a 256-kb SRAM is designed using new 9T and conventional 6T SRAM cells. Operating at their minimum possible VDDs, the proposed design decreases write and read power per operation by 92% and 93%, respectively, over the conventional rival. The area of the proposed SRAM cell is increased by 83% over a conventional 6T one. However, due to large ION/IOFF of read path for 9T cell, we are able to put 1k cells in each column of 256-kb SRAM block, resulting in the possibility for sharing write and read circuitries of each column between more cells compared with conventional 6T. Thus, the area overhead of 256kb SRAM based on new 9T cell is reduced to 37% compared with 6T SRAM. Ghasem Pasandi, Sied Mehdi Fakhraie |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Algorithm and FPGA implementation of interpolation-based soft output mmse mimo detector for 3GPP LTEabstractThe number of orthogonal frequency division multiplexing (OFDM) sub‐carriers in modern wireless communication systems such as Third Generation Partnership Project (3GPP) long term evolution (LTE) can be as high as 2048. Channel matrix of data sub‐carriers in conventional multiple‐input–multiple‐output (MIMO) receivers is derived by interpolating channel matrix of pilot sub‐carriers and MIMO detection is performed on each interpolated channel matrix. Symbol detection on sub‐carrier by sub‐carrier basis demands a very high computational power. This study presents the first very‐large scale integration implementation of a novel 4 × 4 interpolation‐based soft output minimum‐mean‐squared error (MMSE) MIMO detector which combines channel estimation and MIMO detection. The detector computes MMSE matrix on pilot sub‐carriers using proposed division‐free matrix inversion algorithm and obtains MMSE matrix of data sub‐carriers by interpolation, reducing computational complexity by more than 66%. The throughput of the implemented detector on Virtex‐7 field programmable gate array (FPGA) exceeds 1 Gbps with 2.1 µs latency. Amin Salari, Sied Mehdi Fakhraie, Aliazam Abbasfar |
IET Commun. | 2 |
| 2014 | An analytical method for reliability aware instruction set extension
Ali Azarpeyvand, Mostafa E. Salehi, Sied Mehdi Fakhraie |
J. Supercomput. | 3 |
| 2014 | Customized pipeline and instruction set architecture for embedded processing engines
Amir Yazdanbakhsh, Mostafa E. Salehi, Sied Mehdi Fakhraie |
J. Supercomput. | 3 |
| 2013 | Reliability-aware cross-layer custom instruction screeningabstractBias Temperature Instability (BTI) and process variation introduce remarkable unpredictability to Custom Instructions (CIs) manufactured at nano-scale technology. Moreover, shrinking the feature size to nanometer levels makes soft error another critical issue of CIs. To tackle these factors, we propose a reliability-aware cross-layer CI screening method. By adding an intermediate phase between the CI generation and CI selection phases, this method enables designers to prune the outputs of the generation phase in order to guarantee that synthesized CIs meet the required reliability constraints. For this purpose, a holistic framework is developed to analyze the combined effects of the BTI and process variation as well as the soft error on the CIs by making a link between circuit-level and system-level information. Based on this information collected from different layers of abstraction, the screening method prunes those CIs which cannot meet the reliability constraints. Experiments illustrate that BTI-unaware CI selection techniques may not meet the desired lifetime because of BTI-induced delay shift of CIs. Moreover, according to the results, a remarkable percentage of CIs is vulnerable to soft error and should not be fed into CI selection phase. Bahareh J. Farahani, Ali Azarpeyvand, Saeed Safari, Sied Mehdi Fakhraie |
DDECS | 4 |
| 2013 | Defuzzification block: New algorithms, and efficient hardware and software implementation issues
Hamid Reza Mahdiani, Abbas BanaiyanMofrad, Mohammad Haji Seyed Javadi, Sied Mehdi Fakhraie, Caro Lucas |
Eng. Appl. Artif. Intell. | 4 |
| 2012 | CIVA: Custom instruction vulnerability analysis frameworkabstractThis paper describes a methodology for analyzing the vulnerability of custom instructions against the electronic faults, considering different operations and the custom instruction graph topology. Our approach enables designers to optionally constrain the operand types and also the custom functional unit structure to reach an acceptable vulnerability. We have developed a framework to evaluate the desired goal. The presented framework explores the effects of different operations and their dependencies on overall vulnerability of the custom functional units. Our experiments show that, in most cases, custom functional units with similar speed-ups in performance present different vulnerability to soft errors. Ali Azarpeyvand, Mostafa E. Salehi, Sied Mehdi Fakhraie |
DDECS | 3 |
| 2012 | Vulnerability Analysis for Custom InstructionsabstractToday circuits are becoming more vulnerable to electronic noises and reliable system design has emerged as a key challenge to embedded system design. Logic fault in terms of soft errors or transient faults are now a serious problem for embedded processors. Recent developments in customized embedded processors significantly focus on improving the performance and area of the processor by augmenting it with application specific custom functional units that implement custom instructions. This paper analyzes the effect of type, order, and bit-width of the operations of different custom instruction sub-graphs on the vulnerability of extensible processors. We have developed a framework for studying the effects of different operations and their dependencies on overall vulnerability of the custom functional units and our experiments show that, in most cases, similar custom functional units could have different vulnerabilities to soft errors. Our approach enables designers to optionally constrain the operand types and also the custom functional unit structure to reach an acceptable vulnerability. Ali Azarpeyvand, Mostafa E. Salehi, Sied Mehdi Fakhraie |
DSD | 3 |
| 2012 | Instruction set architectural guidelines for embedded packet-processing engines
Mostafa E. Salehi, Sied Mehdi Fakhraie, Amir Yazdanbakhsh |
J. Syst. Archit. | 2 |
| 2012 | Relaxed Fault-Tolerant Hardware Implementation of Neural Networks in the Presence of Multiple Transient ErrorsabstractReliability should be identified as the most important challenge in future nano-scale very large scale integration (VLSI) implementation technologies for the development of complex integrated systems. Normally, fault tolerance (FT) in a conventional system is achieved by increasing its redundancy, which also implies higher implementation costs and lower performance that sometimes makes it even infeasible. In contrast to custom approaches, a new class of applications is categorized in this paper, which is inherently capable of absorbing some degrees of vulnerability and providing FT based on their natural properties. Neural networks are good indicators of imprecision-tolerant applications. We have also proposed a new class of FT techniques called relaxed fault-tolerant (RFT) techniques which are developed for VLSI implementation of imprecision-tolerant applications. The main advantage of RFT techniques with respect to traditional FT solutions is that they exploit inherent FT of different applications to reduce their implementation costs while improving their performance. To show the applicability as well as the efficiency of the RFT method, the experimental results for implementation of a face-recognition computationally intensive neural network and its corresponding RFT realization are presented in this paper. The results demonstrate promising higher performance of artificial neural network VLSI solutions for complex applications in faulty nano-scale implementation environments. Hamid Reza Mahdiani, Sied Mehdi Fakhraie, Caro Lucas |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2011 | Dynamic Soft Error Hardening via Joint Body Biasing and Dynamic Voltage ScalingabstractShrinking feature sizes, reduced voltages, and higher transistor count of nano-scale silicon chips challenge designers in terms of performance, power consumption, and reliability. This paper investigates the effect of simultaneous use of dynamic voltage and frequency scaling (DVFS) and body biasing (BB) on power consumption, reliability, and performance. An analytical model of reliability as a function of body bias voltage, supply voltage, and frequency is proposed. We derive a three dimensional optimization problem by exploiting proposed reliability model in conjunction with power consumption and performance model. The resulting problem is solved using widely-used geometric optimization to identify optimal supply voltage and body bias voltage and then is validated using accurate simulation. Afterwards, it is demonstrated how joint energy-performance-reliability space optimization method can be used in an adaptive reliability-aware power management systems. Finally, we show that combined soft error aware BB and DVFS is capable of improving power consumption about 30% in comparison to reliability-aware DVFS only for the same level of reliability and performance constraints. Farshad Firouzi, Amir Yazdanbakhsh, Hamed Dorosti, Sied Mehdi Fakhraie |
DSD | 4 |
| 2011 | Dynamic Voltage and Frequency Scheduling for Embedded Processors Considering Power/Performance TradeoffsabstractAn adaptive method to perform dynamic voltage and frequency scheduling (DVFS) for minimizing the energy consumption of microprocessor chips is presented. Instead of using a fixed update interval, the proposed DVFS system makes use of adaptive update intervals for optimal frequency and voltage scheduling. The optimization enables the system to rapidly track the workload changes so as to meet soft real-time deadlines. The technique, which can be realized with very simple hardware, is completely transparent to the application. The results of applying the method to some real application workloads demonstrate considerable power savings and fewer frequency updates compared to DVFS systems based on fixed update intervals. Mostafa E. Salehi, Mehrzad Samadi, Mehrdad Najibi, Ali Afzali-Kusha, Massoud Pedram, Sied Mehdi Fakhraie |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2010 | Instruction reliability analysis for embedded processorsabstractAdvances in silicon technology and shrinking the feature size to nanometer scale make unreliability of nano devices the most important concern of fault-tolerant designs. Soft error analysis has been greatly aided by the concept of architectural vulnerability factor (AVF) and architecturally correct execution (ACE). In this work, we exploit the techniques of AVF analysis to introduce the instruction-level vulnerability metric for software reliability analysis. The proposed metric can be used to make judgments about the reliability of different programs on different processors with regard to architectural and compiler guidelines for improving the processor reliability. Ali Azarpeyvand, Mostafa E. Salehi, Farshad Firouzi, Amir Yazdanbakhsh, Sied Mehdi Fakhraie |
DDECS | 5 |
| 2010 | Architecture-Level Design Space Exploration of Super Scalar Microarchitecture for Network ApplicationsabstractIncreasing diversity in packet-processing applications and rapid increases in channel bandwidth lead to greater complexity in communication protocols. These factors result in larger computational loads for packet-processing engines that introduce high performance microprocessor designs as an important solution. This paper presents an exhaustive simulation for exploring the performance of instruction-level parallel super scalar processors executing packet-processing applications. Based on the simulation results, a design space exploration has been used to derive performance-efficient application-specific super scalar processor architecture based on MIPS instruction set architecture. Simple Scalar architecture toolset has been used for design space exploration and network applications have been investigated to guide the architecture exploration. The optimizations achieve up to 80% improvement in performance for representative packet-processing applications. Mostafa E. Salehi, Hamed Dorosti, Sied Mehdi Fakhraie |
DSD | 3 |
| 2010 | Parallel scalable hardware implementation of asynchronous discrete particle swarm optimization
Amin Farmahini Farahani, Shervin Vakili, Sied Mehdi Fakhraie, Saeed Safari, Caro Lucas |
Eng. Appl. Artif. Intell. | 3 |
| 2009 | Computationally efficient active rule detection method: Algorithm and architecture
Mahdi Hamzeh, Hamid Reza Mahdiani, Ahmad Saghafi, Sied Mehdi Fakhraie, Caro Lucas |
Fuzzy Sets Syst. | 4 |
| 2009 | Quantitative analysis of packet-processing applications regarding architectural guidelines for network-processing-engine development
Mostafa E. Salehi, Sied Mehdi Fakhraie |
J. Syst. Archit. | 2 |
| 2008 | Scalable Architecture for on-Chip Neural Network Training using Swarm IntelligenceabstractThis paper presents a novel architecture for on-chip neural network training using particle swarm optimization (PSO). PSO is an evolutionary optimization algorithm with a growing field of applications which has been recently used to train neural networks. The architecture exploits PSO algorithm to evolve network weights as well as a method called layer partitioning to implement neural networks. In the proposed method, a neural network is partitioned into groups of neurons and the groups are sequentially mapped to available functional units. Thus, the architecture is reconfigurable for training and implementing different multilayer feedforward neural networks without the need for modifying the architecture. The implementation is intended for real-time applications regarding hardware cost and speed. The results show that the proposed system provides a trade-off between resource requirements and speed. Amin Farmahini Farahani, Sied Mehdi Fakhraie, Saeed Safari |
DATE | 2 |
| 2008 | A 65nm 10GHz pipelined MAC structureabstractIn this paper a pipelined 16times16+32 MAC structure is explained. This is a fused MAC which is using modified Booth encoding technique and a new low-voltage-swing 4:2 compressor. For the final adder, it is using low-voltage-swing carry-select structure. With this topology, we achieved a 5-stage pipelined MAC with 10 GHz clock frequency and 15 mW/GHz average power dissipation in 65 nm CMOS technology with 1.2 V power supply. Fatemeh Kashfi, Sied Mehdi Fakhraie, Saeed Safari |
ISCAS | 2 |
| 2007 | HW/SW partitioning using discrete particle swarmabstractHardware/Software partitioning is one of the most important issues of codesign of embedded systems, since the costs and delays of the final results of a design will strongly depend on partitioning. We present an algorithm based on Particle Swarm Optimization to perform the hardware/software partitioning of a given task graph for minimum cost subject to timing constraint. By novel evolving strategy, we enhance the efficiency and result's quality of our partitioning algorithm in an acceptable run-time. Also, we compare our results with those of Genetic Algorithm on different task graphs. Experimental results show the algorithm's effectiveness in achieving the optimal solution of the HW/SW partitioning problem even in large task graphs. Amin Farmahini Farahani, Mehdi Kamal, Sied Mehdi Fakhraie, Saeed Safari |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | Tertiary-Tree 12-GHz 32-bit Adder in 65nm TechnologyabstractThis paper presents a new 32-bit adder structure with 12 GHz low-power operation in 65nm technology. The fast conditional sparse-tree logic (FCSL) is based on modifying the initial sparse-tree architecture (Mathew et al., 2003) to enhance its speed using tertiary trees and applying a carry-select scheme in some of the more significant bits. This design has been compared with the sparse-tree adder and the low-voltage swing adder in terms of speed and power. It has been shown that speed can be improved using FCSL architecture while keeping the power at a comparable level. Amir Agah, Sied Mehdi Fakhraie, Azita Emami-Neyestanak |
ISCAS | 2 |
| 2007 | MDST: Multiprocessor DSP Simulation Toolkit for Voice Processing ApplicationsabstractIn this paper, we propose a multiprocessor DSP simulation toolkit suitable for performance evaluation of data-parallel applications like voice processing. The proposed toolkit uses the benefits of multi-level parallelism and clustering. Different DSP clusters are considered for the multiprocessor DSP simulation engine in which the DSP processors are grouped to cooperate. Satisfying the communication requirements, two global and local communication engines (GCE and LCE) implement the real behavior of intra-and inter-cluster communications. Using efficient abstraction levels for interconnections reduces the simulation time significantly. Abstract communication modeling, cycle-accurate behavior, and multi-level controlling are the most important features of the proposed simulation platform. Performance of the simulator is verified by standard single- and multi-channel voice processing applications such as ITU-T G.729a speech codec. Naser Sedaghati, Mahdi Nazm Bojnordi, Sied Mehdi Fakhraie |
MASCOTS | 3 |
| 2007 | A New Search Space Reduction Technique for Acquisition of UWB Signals in Multipath ChannelsabstractRapid acquisition of ultra-wideband (UWB) signals in multipath channels is a challenging task due to the following reasons. First, fine time resolution of UWB signals introduces a large search space. Second, due to the very low power of the received signal, it needs to be observed for a long time to make a reliable decision. These two problems make the process of acquisition unacceptably slow. In this paper, we introduce a new two-stage algorithm to reduce the search space and consequently the acquisition time. The performance of the proposed method is analytically evaluated in terms of mean acquisition time and the results are validated through simulations. The results show a significant improvement in the speed of acquisition with no additional complexity. Ahmad Saghafi, Sied Mehdi Fakhraie |
VTC Spring | 2 |
| 2006 | A Representation for Genetic-Algorithm-Based Multiprocessor Task SchedulingabstractA multiprocessor scheduling problem is defined as the assignment of a given set of tasks to a set of processors. These tasks should be assigned in a way such that the total execution time is minimized and certain criteria are met. A wide range of solutions and heuristics have been proposed to solve this important system optimization problem. In this paper, we propose a novel representation to solve the task scheduling problem using genetic algorithm (GA). This representation is novel not only in the way it presents task scheduling, but also in that the length of that representation is intelligently adaptable to the given problem. Task duplication is allowed in our method and it is capable of spanning a large proportion of the solution space without the need for penalty/rewards or adding repair mechanisms whilst always generating valid chromosomes. Due to this new representation, order of the search space has been reduced; consequently, the proposed approach outperforms some recently studied GA based scheduling methods over 120 times with respect to the number of fitness evaluations. Mehdi Salmani Jelodar, S. Najmeh Fakhraie, Faezeh Montazeri, Sied Mehdi Fakhraie, Majid Nili Ahmadabadi |
IEEE Congress on Evolutionary Computation | 4 |
| 2006 | SOPC-Based Parallel Genetic AlgorithmabstractThe ever-growing complexity of the modern chips is forcing fundamental changes in the way systems are designed. System-on-a-Programmable-Chip (SOPC) concept is bringing a major revolution in the design of integrated circuits, due to the fact that it makes unprecedented levels of in-field integration possible. Genetic Algorithm (GA) is a powerful function optimizer that is used successfully to solve problems in many different disciplines. A major drawback of GA is that it needs huge computation time for sequential execution on PCs. Therefore, the hardware implementation of GA has been the focus of some recent studies. Parallel GA (PGA) is particularly important for efficient hardware implementation and promise substantial gains in performance and results. In this paper, a SOPC-based PGA framework is proposed. Our proposed framework can be used in real-time applications. We have implemented our proposed system on an Altera ® Stratix Development Kit and we compare its performance with the corresponding software simulation. The results obtained indicate a speedup of up to 50 times in the elapsed computation time. Mehdi Salmani Jelodar, Mehdi Kamal, Sied Mehdi Fakhraie, Majid Nili Ahmadabadi |
IEEE Congress on Evolutionary Computation | 3 |
| 2006 | Software Implementation Issues of Existing and New Defuzzification MethodsabstractThis paper discusses software implementation issues of different defuzzification procedures in fuzzy systems. Three new defuzzification methods are introduced which are suitable for efficient software and also hardware implementations. A set of seven important existing defuzzification methods are reviewed and compared with these new methods for different software implementation approaches. The C models of all methods are prepared to perform a comprehensive analysis on the output accuracy of different methods. The results prove the superiority of our new proposed methods. In another study, three categories of low-level assembly models are developed for each method to evaluate its software execution time and instruction count when executed on each of three chosen popular processors. Namely, Texas Instruments C6xcopy DSP, Intel's Pentiumcopy IV, and IBM's PowerPC PPC405copy processors are used as the running engines for this comparison. Some accuracy-speed analysis diagrams are then introduced to guide the designers for choosing the defuzzification method which best suites their application requirements. Abbas BanaiyanMofrad, Hamid Reza Mahdiani, Sied Mehdi Fakhraie |
FUZZ-IEEE | 3 |
| 2006 | Dynamic voltage and frequency management based on variable update intervals for frequency settingabstractAn efficient adaptive method to perform dynamic voltage and frequency management (DVFM) for minimizing the energy consumption of microprocessor chips is presented. Instead of using a fixed update interval, the proposed DVFM system makes use of adaptive update intervals for optimal frequency and voltage scheduling. The optimization enables the system to rapidly track the workload changes so as to meet soft real-time deadlines. The method, which is based on introducing the concept of an effective deadline, utilizes the correlation between consecutive values of the workload. In practice because the frequency and voltage update rates are dynamically set based on variable update interval lengths, voltage fluctuations on the power network are also minimized. The technique, which may be implemented by simple hardware and is completely transparent from the application, leads to power savings of up to 60% for highly correlated workloads compared to DVFM systems based on fixed update intervals. Mehrdad Najibi, Mostafa E. Salehi, Ali Afzali-Kusha, Massoud Pedram, Sied Mehdi Fakhraie, Hossein Pedram |
ICCAD | 5 |
| 2006 | Neural network stream processing core (NnSP) for embedded systemsabstractNnSP is a stream-based programmable and code-level statically reconfigurable processor for realization of neural networks in embedded systems. NnSP is provided with a neural-network-to-stream compiler and a hardware core builder. The NnSP stream compiler makes it possible to realize various neural networks using NnSP. On the other hand, the NnSP builder makes the NnSP processor an IP core that can be restructured to satisfy different demands and constraints. This paper presents the architecture of the NnSP processor, the streaming mechanism, and the builder facilities. Also, synthesis results of a 64-PE NnSP on a 0.18 mum standard-cell library are presented. The obtained results show that a 64-PE NnSP can perform computations of 25.6 giga connections in a second, while its throughput is upto 51.2 giga 32-bit fixed point operations per second. Comparing with high performance parallel architectures locates 64-PE NnSP among the best state of the art parallel processors Hadi Esmaeilzadeh, Pooya Saeedi, Babak Nadjar Araabi, Caro Lucas, Sied Mehdi Fakhraie |
ISCAS | 5 |
| 2006 | Implementation of a high-speed low-power 32-bit adder in 70nm technologyabstractIn this article, the performance and power dissipation of two differential logic circuits in deep sub-micron technologies are obtained and compared together, and the superior topology is introduced. Low voltage swing (LVS) technique which improves circuit performance and lowers power consumption is described in detail. We conclude this article with the design, simulation and optimization of a high speed low-power 32-bit adder using the LVS technique in 70nm technology. This circuit can operate at 10GHz clock frequency with power dissipation as low as 2.58 mW/GHz. Fatemeh Kashfi, Sied Mehdi Fakhraie |
ISCAS | 2 |
| 2006 | Hardware implementation and comparison of new defuzzification techniques in fuzzy processorsabstractThis paper deals with hardware implementation aspects of the defuzzification block in fuzzy controllers and processors. Three new defuzzification methods are introduced which are suitable for low cost hardware implementation. A complete set of common existing defuzzification methods are reviewed to be compared with these new methods from different hardware implementation aspects. Two different hardware models with different structures are developed for each method. The first model realizes the full combinational or fastest possible hardware implementation, and the second model demonstrates the full sequential or the smallest possible hardware implementation of each method. All models are synthesized on 0.18 micron CMOS technology cells to analyze and compare the area, delay and power consumption of different realizations of defuzzification methods. Some area-power-delay-accuracy analysis diagrams are then introduced according to the synthesis results to guide the designers through choosing the defuzzification method and also the implementation structure which best suites their application from implementation cost, speed, power consumption and output accuracy points of view Hamid Reza Mahdiani, Abbas BanaiyanMofrad, Sied Mehdi Fakhraie |
ISCAS | 3 |
| 2005 | Genetic-algorithm Memory Minimisation for Designing Reconfigurable Ip Address Lookup EngineabstractIP address lookup engine is the beating heart of a router. For meeting the requirements of a desirable high-speed router, speed, memory consumption, scalability, and reconfigurability of its IP lookup engine are critical. This paper uses a genetic-algorithm approach to optimise the structure of fixed-stride multibit-trie IP lookup methods. In this work, the genetic algorithm is used first in the design phase as an offline optimisation mechanism. This nature-inspired simulation finds the most memory-efficient configuration of IP address segmentation for a fixed number of address segments. Then, for adapting to network variations, the proposed method dynamically changes the number of address segments to compromise between speed and memory consumption. Each time the number of segments changes, an online genetic program is run to optimise the segmentation. Therefore, the lookup engine reconfigures itself to cover more prefixes during the time. The reconfigurability in response to the network variations, and scalability to the number of prefixes improves the life time of a router that uses this method. Saeed Shamshiri, Sied Mehdi Fakhraie |
Int. J. Comput. Intell. Appl. | 2 |
| 2004 | A Low-Cost At-Speed BIST Architecture for Embedded Processor and SRAM Cores
Mohammad H. Tehranipour, Sied Mehdi Fakhraie, Zainalabedin Navabi, M. R. Movahedin |
J. Electron. Test. | 2 |
| 2004 | Scalable closed-boundary analog neural networksabstractIn many pattern-classification and recognition problems, separation of different swarms of class representatives is necessary. As well, in function-approximation problems, neurons with a local area of influence have demonstrated measurable success. In our previous work, we have shown how intrinsic quadratic characteristics of traditional metal-oxide-semiconductor (MOS) devices can be used to implement hyperspherical discriminating surfaces in hardware-implemented neurons. In this work, we further extend the concept from quadratic forms to more-arbitrary closed-boundary shapes. Accordingly, we demonstrate how intrinsic characteristics of submicron MOS devices can be utilized to implement efficient pattern discriminators for various applications and, through representative simulations, show their success in some typical function-approximation problems. Further, we offer two mathematical interpretations of possible roles for these networks: Geometrically, we show that our networks employ closed hypercone shapes as their discriminating surfaces; analytically, we show that a set of these synapses connected to a common integrating body calculates the distance between their inputs and weight vectors using a power norm. The feasibility of the idea is practically investigated by design, implementation, and test of a three-dimensional (3-D) closed-boundary pattern classifier, fabricated in 0.35-microm complimentary MOS, whose results are reflected in this work. Sied Mehdi Fakhraie, Hamed Farshbaf, Kenneth C. Smith |
IEEE Trans. Neural Networks | 1 |
| 2000 | Kalman-filtering timing recovery scheme for orthogonal frequency domain multiplexing (OFDM) systemsabstractThis paper describes a new technique for timing synchronization in orthogonal frequency domain multiplexing (OFDM) transceiver systems. The proposed technique is based on non-synchronized sampling rate and no pilot is required to be transmitted. Therefore the transmission capacity is increased. The proposed algorithm employs the angles of the received symbols in the OFDM subchannels and provides a computationally efficient technique for estimation of the timing error. An equivalent state space model is also derived for timing error and then Kalman filtering method is exploited for tracking purposes. The proposed technique is very robust, particularly under low signal to noise ratio conditions and has been verified by means of computer simulations. Mohamad Hajirostam, Mohammad Ali Maddah-Ali, A. Haft-Baradaran, M. T. Kilani, Sied Mehdi Fakhraie, M. Sharif-Khani, Omid Shoaei |
ICASSP | 5 |
| 2000 | A rail-to-rail, constant-Gm, 1-volt CMOS opampabstractA rail-to-rail, constant-g/sub m/, 1-volt only, full-CMOS opamp is presented. The opamp has a complementary PMOS/bulk-driven input stage with a feedback circuit to equalize g/sub m/, and a class AB output stage, which provide input and output rail-to-rail operation. HSPICE simulations are performed using BSIM 3.3 models of a 0.8 /spl mu/m CMOS process. This opamp has a DC gain of 45.1 dB, unity-gain bandwidth of 1.7 MHz, and phase margin of 63.4/spl deg/. Faramarz Bahmani, Sied Mehdi Fakhraie, Ali Khaki-Firooz |
ISCAS | 2 |
| 2000 | Low-power data-driven dynamic logic (D3L) [CMOS devices]abstractIn this paper a new family of low-power dynamic logic called data-driven dynamic logic (D/sup 3/L) is introduced. In this logic family, the synchronization clock has been eliminated, and correct sequencing is maintained by appropriate use of data instances. Then, it is shown that replacement of the clock with input data implies less power dissipation without speed degradation compared to conventional dynamic logic. Ramin Rafati, Sied Mehdi Fakhraie, Kenneth C. Smith |
ISCAS | 2 |