Morteza Saheb Zamani

dblp:69/5980 · DBLP profile ↗
← Back
51ranked-venue papers
5as first author
4since 2021 · last 2023
0000-0002-0826-1091ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 2 · 2 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2023 A parallel computing architecture based on cellular automata for hydraulic analysis of water distribution networks
Ali Suvizi, Azim Farghadan, Morteza Saheb Zamani
J. Parallel Distributed Comput.3
2023 An energy-efficient and accuracy-aware edge computing framework for heart arrhythmia detection: A joint model selection and task offloading approach
Vahid Amini, Mahmoud Momtazpour, Morteza Saheb Zamani
J. Supercomput.3
2022 FIFA: A Fully Invertible FPGA Architecture to Reduce BTI-Induced Aging Effects
abstract
FPGAs are increasingly becoming sensitive to aging effects mainly through BTI phenomena in the latest technology nodes. This phenomenon can be modeled as a shift in the threshold voltage of transistors which leads to performance degradation and reduction in SNM of SRAM cells. To reduce the aging effects on FPGA building blocks, we propose FIFA, a fully invertible FPGA architecture. In this architecture, two small modules are introduced to make the bitstream of logic and routing resources of the FPGA tiles fully invertible. Accordingly, to insert recovery cycles in the lifetime of the transistors, the bitstream stored in the configuration SRAM cells can be inverted occasionally without any changes in the functionality of the design. Using the proposed architecture, it is neither required to repeat the placement and routing procedures to generate multiple configuration bitstreams, nor extra memory is needed to save alternative bitstreams for changing the SRAM cells contents. Our experimental results over a set of industrial benchmarks show that by choosing an appropriate switch block arrangement and an optimized flipping frequency, the proposed architecture can improve the aging induced performance degradation by up to 62.5% with acceptable power and area overheads on logic and routing resources.
Mohammad Hadi Mottaghi, Mehdi Sedighi, Morteza Saheb Zamani
IEEE Trans. Computers3
2022 Runtime hardware Trojan detection by reconfigurable monitoring circuits
Reza Fani, Morteza Saheb Zamani
J. Supercomput.2
2020 Aging Mitigation in FPGAs Considering Delay, Power, and Temperature
abstract
Field programmable gate array (FPGA) devices are highly susceptible to transistor aging, mainly through the bias temperature instability (BTI) phenomenon. BTI can be modeled as threshold voltage increase in MOSFET transistors, which leads to the degradation in device performance. As technology scales, leakage power has turned into a major portion of FPGA total power consumption. In this paper, we study the mutual effects of BTI and leakage power by considering the temperature changes in the basic components of FPGAs. Our analysis shows that while the leakage-power reduction caused by BTI may be considered desirable, a bit-flipping scheme should still be employed to mitigate device degradation. We present an optimization problem to optimize the device performance over the device's lifetime. A postrouting aging-aware timing analysis method is also proposed to find the best flipping frequency. The simulation results show that bit flipping at a proper frequency may reduce the power consumption of device by about 5% while keeping the critical path delay below a given constraint.
Mohammad Hadi Mottaghi, Mehdi Sedighi, Morteza Saheb Zamani
IEEE Trans. Reliab.3
2017 Quantum Circuit Synthesis Targeting to Improve One-Way Quantum Computation Pattern Cost Metrics
abstract
One-way quantum computation (1WQC) is a model of universal quantum computations in which a specific highly entangled state called acluster stateallows for quantum computation by single-qubit measurements. The needed computations in this model are organized as measurement patterns. The traditional approach to obtain a measurement pattern is by translating a quantum circuit that solely consists of CZ andJ(α) gates into the corresponding measurement patterns and then performing some optimizations by using techniques proposed for the 1WQC model. However, in these cases, the input of the problem is a quantum circuit, not an arbitrary unitary matrix. Therefore, in this article, we focus on the first phase—that is, decomposing a unitary matrix into CZ andJ(α) gates. Two well-known quantum circuit synthesis methods, namely cosine-sine decomposition and quantum Shannon decomposition are considered and then adapted for a library of gates containing CZ andJ(α), equipped with optimizations. By exploring the solution space of the combinations of these two methods in a bottom-up approach of dynamic programming, a multiobjective quantum circuit synthesis method is proposed that generates a set of quantum circuits. This approach attempts to simultaneously improve the measurement pattern cost metrics after the translation from this set of quantum circuits.
Mahboobeh Houshmand, Mehdi Sedighi, Morteza Saheb Zamani, Kourosh Marjoei
ACM J. Emerg. Technol. Comput. Syst.3
2017 Latch-Based Structure: A High Resolution and Self-Reference Technique for Hardware Trojan Detection
abstract
Hardware Trojan detection has been the subject of many studies in the realm of hardware security in the recent years. The effectiveness of current techniques proposed for Trojan detection is limited by some factors, process variation noise being a major one. This paper introduces latch-based structures as a self-reference detection technique which uses in-circuit path delays as golden reference models. By addressing process variation, these structures can achieve high accuracy in detection resolution. The proposed method is a complementary approach to current side-channel techniques to cover their poor performance in detecting small Trojans. Simulation results show that this technique can detect small Trojans at the scale of only one logical gate with 90 percent probability on average.
Ghobad Zarrinchian, Morteza Saheb Zamani
IEEE Trans. Computers2
2016 FogLight: an efficient matrix-based approach to construct metabolic pathways by search space reduction
abstract
MOTIVATION: A fundamental computational problem in the area of metabolic engineering is finding metabolic pathways between a pair of source and target metabolites efficiently. We present an approach, namely FogLight, for searching metabolic networks utilizing Boolean (AND-OR) operations represented in matrix notation to efficiently reduce the search space. This enables the enumeration of all pathways between metabolites that are too distant for the application of brute-force methods. RESULTS: Benchmarking tests run with FogLight show that it can reduce the search space by up to 98%, after which the accelerated search for high accurate results is guaranteed. Using FogLight, several pathways between eight given pairs of metabolites are found of which the pathways from CO2 to ethanol are specifically discussed. Additionally, in comparison with three path-finding tools, namely PHT, FMM and RouteSearch, FogLight can find shorter and more pathways for attempted source-target metabolite pairs. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mehrshad Khosraviani, Morteza Saheb Zamani, Gholamreza Bidkhori
Bioinform.2
2016 Quantum-Logic Synthesis of Hermitian Gates
abstract
In this article, the problem of synthesizing a general Hermitian quantum gate into a set of primary quantum gates is addressed. To this end, an extended version of the Jacobi approach for calculating the eigenvalues of Hermitian matrices in linear algebra is considered as the basis of the proposed synthesis method. The quantum circuit synthesis method derived from the Jacobi approach and its optimization challenges are described. It is shown that the proposed method results in multiple-control rotation gates around the y axis, multiple-control phase shift gates, multiple-control NOT gates, and a middle diagonal Hermitian matrix, which can be synthesized to multiple-control Pauli Z gates. Using the proposed approach, it is shown how multiple-control U gates, where U is a single-qubit Hermitian quantum gate, can be implemented using a linear number of elementary gates in terms of circuit lines with the aid of one auxiliary qubit in an arbitrary state.
Mona Arabzadeh, Mahboobeh Houshmand, Mehdi Sedighi, Morteza Saheb Zamani
ACM J. Emerg. Technol. Comput. Syst.4
2014 Decomposition of Diagonal Hermitian Quantum Gates Using Multiple-Controlled Pauli Z Gates
abstract
Quantum logic decomposition refers to decomposing a given quantum gate to a set of physically implementable gates. An approach has been presented to decompose arbitrary diagonal quantum gates to a set of multiplexed-rotation gates around z axis. In this article, a special class of diagonal quantum gates, namely diagonal Hermitian quantum gates, is considered and a new perspective to the decomposition problem with respect to decomposing these gates is presented. It is first shown that these gates can be decomposed to a set that solely consists of multiple-controlled Z gates. Then a binary representation for the diagonal Hermitian gates is introduced. It is shown that the binary representations of multiple-controlled Z gates form a basis for the vector space that is produced by the binary representations of all diagonal Hermitian quantum gates. Moreover, the problem of decomposing a given diagonal Hermitian gate is mapped to the problem of writing its binary representation in the specific basis mentioned previously. Moreover, CZ gate is suggested to be the two-qubit gate in the decomposition library, instead of previously used CNOT gate. Experimental results show that the proposed approach can lead to circuits with lower costs in comparison with the previous ones.
Mahboobeh Houshmand, Morteza Saheb Zamani, Mehdi Sedighi, Mona Arabzadeh
ACM J. Emerg. Technol. Comput. Syst.2
2013 Improving bitstream compression by modifying FPGA architecture
abstract
The size of configuration bitstreams of field-programmable gate arrays (FPGA) is increasing rapidly. Compression techniques are used to decrease the size of bitstreams. In this paper, an appropriate bitstream format and variable symbol lengths are proposed to utilize the routing patterns for enhancing the compression efficiency. An order of inputs of multiplexers in switch modules is also proposed to improve the symbol statistics and hence, the compression efficiency. A framework to generate the bitstream and hardware description of FPGAs is developed as well. Experimental results over 20 MCNC benchmarks show that by applying the proposed approaches, the compression rate is improved by 46% on average compared to the methods with fixed symbol lengths without any area and performance degradation.
Seyyed Ahmad Razavi, Morteza Saheb Zamani
FPGA2
2012 Hardware Acceleration of STON Algorithm for Comparing 3-D Structure of Proteins
abstract
Comparing three-dimensional structure of proteins is one of the most fundamental problems in bioinformatics. In recent years, various algorithms have been proposed to solve this problem efficiently. The proposed algorithms are very time-consuming due to high complexity and large input data. In this paper, hardware acceleration is applied to minimize the execution time of a previous algorithm, called STON. An FPGA-based hardware is presented to perform this task in a fast manner. An average speed-up of 2.5x is achieved compared to the software execution.
Somayeh Kashi, Morteza Saheb Zamani
DSD2
2012 OWQS: One-Way Quantum Computation Simulator
abstract
In one-way quantum computation (1WQC) model, universal quantum computations are performed using measurements to designated qubits in a highly entangled state. The choices of basis for these measurements as well as the structure of the entanglements specify a quantum algorithm. Although a number of methods have been proposed to simulate quantum circuit model on classical computers, no efficient tool has been developed to simulate the 1WQC model directly. In this paper, some techniques such as qubit elimination, implicit and in-place matrix-vector multiplication and pattern reordering are utilized to considerably reduce the time and memory needed for the simulations. These techniques were implemented in a tool called One-Way Quantum computation Simulator (OWQS). Experimental results validate the efficiency of the proposed approach.
Eesa Nikahd, Mahboobeh Houshmand, Morteza Saheb Zamani, Mehdi Sedighi
DSD3
2012 Timing yield improvement of FPGAs utilizing enhanced architectures and multiple configurations under process variation (abstract only)
abstract
Designing with field-programmable gate arrays (FPGAs) can face with difficulties due to process variations. Some techniques use reconfigurability of FPGAs to reduce the effects of process variations in these chips. Furthermore, FPGA architecture enhancement is an effective way to degrade the impact of variation. In this paper, various FPGA architectures are examined to identify which architecture can achieve larger parametric yield improvement utilizing multiple configurations as opposed to single configuration. Experimental results show that by increasing cluster size from 4 to 10, yield improvement increases from 2.82X to 4.48X. However, changing look-up table (LUT) size from 4 to 7 results in yield improvement degradation from 2.82X to 1.45X, using 10 configurations compared to single configuration over 20 MCNC benchmark circuits. These results indicate that multi-configuration technique causes larger timing yield improvement in FPGAs with larger cluster size and smaller LUT size.
Fatemeh Sadat Pourhashemi, Morteza Saheb Zamani
FPGA2
2011 VMAP: A Variation Map-Aware Placement Algorithm for Leakage Power Reduction in FPGAs
abstract
In high frequency FPGAs with technology scale shrinking and threshold voltage value decreasing and based on existing large numbers of unused resources, leakage power has a considerable contribution in total power consumption. On the other hand, process variation, as an important challenge in nano-scale technologies, has a great impact on leakage power of FPGAs. Reconfigurability of FPGAs makes an unique opportunity to mitigate these challenges by their unique variation map extraction. In this paper, a per-chip process variation-aware placement (VMAP) algorithm is proposed to reduce the leakage power of FPGAs using the extracted variation map without neglecting dynamic power consumption. VMAP is adaptive to different process variation maps of various FPGA chips. Experimental results on attempted benchmarks show that power-delay-product (PDP) cost is reduced by 7.2% in the VMAP compared with conventional placement algorithms, with less than 16.8% standard deviation for different variation maps.
Behzad Salami 0001, Morteza Saheb Zamani, Ali Jahanian 0001
DSD2
2011 Evaluation of FPGA routing architectures under process variation
abstract
Uncertainty in performance of FPGAs is becoming an important issue due to increased process variations in nanometer regime. Therefore, it is vital to decrease the impact of variability in these devices. FPGA routing architecture enhancement can be an effective way, because as feature size scales down, routing delay dominates logic circuit delay. In this paper, unidirectional and bidirectional routing architectures are compared. We show that bidirectional architecture is better in terms of robustness against variation when short wire segments are considered. However as wire length increases, unidirectional routing architecture would be the preferred option. Experimental results show that in unidirectional routing architecture towards bidirectional for wire length of 8, it has obtained 36% and 20% improvement in standard deviation and 3¼+Ã of circuit delay, respectively.
Fatemeh Sadat Pourhashemi, Morteza Saheb Zamani
ACM Great Lakes Symposium on VLSI2
2011 Improved predictability, timing yield and power consumption using hierarchical highways-on-chip planning methodology
Ali Jahanian 0001, Morteza Saheb Zamani, Hamid Safizadeh
Integr.2
2010 Rule-based optimization of reversible circuits
abstract
Reversible logic has applications in various research areas including low-power design and quantum computation. In this paper, a rule-based optimization approach for reversible circuits is proposed which uses both negative and positive control Toffoli gates during the optimization. To this end, a set of rules for removing NOT gates and optimizing sub-circuits with common-target gates are proposed. To evaluate the proposed approach, the best-reported synthesized circuits and the results of a recent synthesis algorithm which uses both negative and positive controls are used. Our experiments reveal the potential of the proposed approach in optimizing synthesized circuits.
Mona Arabzadeh, Mehdi Saeedi, Morteza Saheb Zamani
ASP-DAC3
2010 A decoder-based switch box to mitigate soft errors in SRAM-based FPGAs
abstract
This paper proposes a new switch box architecture in SRAM-based FPGAs to mitigate soft error effects. In this switch box architecture, the number of SRAM bits required for programming switch box is reduced to 67% without any impact on routing capability of the switch box. This architecture does not require any modification of the existing placement and routing algorithms. The architecture was evaluated based on several MCNC benchmarks using VPR tool. The experimental results show that this architecture decreases the susceptibility of switch boxes to SEUs about 20% on average compared to the traditional ones.
Hassan Ebrahimi, Morteza Saheb Zamani, Hamid R. Zarandi
ASP-DAC2
2010 Reduction of process variation effect on FPGAs using multiple configurations
abstract
In recent years, parameter variations present critical challenges for manufacturability and yield on integrated circuits. In this paper, a new method for improving the timing yield of field programmable gate array (FPGA) devices affected by random and systematic within-die variation is proposed. By selection of an appropriate configuration from a set of functionally equivalent configurations average critical path delay is reduced under conditions of large random and systematic variation considering spatial correlation. Compared to the previous approach which is limited to a fixed placement, our method improves timing yield by attempting several placements and routings without lengthy placement and routing phases to handle systematic variations and spatial correlation. The average critical path delay is reduced by 7% compared to the previous work over 20 MCNC benchmarks. General Terms: FPGA, process variation, placement, routing
Delasa Aghamirzaie, Seyyed Ahmad Razavi, Morteza Saheb Zamani, Mahdi Nabiyouni
VLSI-SoC3
2010 Reversible circuit synthesis using a cycle-based approach
abstract
Reversible logic has applications in various research areas, including signal processing, cryptography and quantum computation. In this article, direct NCT-based synthesis of a given k -cycle in a cycle-based synthesis scenario is examined. To this end, a set of seven building blocks is proposed that reveals the potential of direct synthesis of a given permutation to reduce both quantum cost and average runtime. To synthesize a given large cycle, we propose a decomposition algorithm to extract the suggested building blocks from the input specification. Then, a synthesis method is introduced that uses the building blocks and the decomposition algorithm. Finally, a hybrid synthesis framework is suggested that uses the proposed cycle-based synthesis method in conjunction with one of the recent NCT-based synthesis approaches which is based on Reed-Muller (RM) spectra. The time complexity and the effectiveness of the proposed synthesis approach are analyzed in detail. Our analyses show that the proposed hybrid framework leads to a better quantum cost in the worst-case scenario compared to the previously presented methods. The proposed framework always converges and typically synthesizes a given specification very fast compared to the available synthesis algorithms. Besides, the quantum costs of benchmark functions are improved about 20% on average (55% in the best case).
Mehdi Saeedi, Morteza Saheb Zamani, Mehdi Sedighi, Zahra Sasanian
ACM J. Emerg. Technol. Comput. Syst.2
2009 A cycle-based synthesis algorithm for reversible logic
abstract
Several algorithms have been proposed for the synthesis of reversible circuits. In this paper, a cycle-based synthesis algorithm for reversible logic, based on the NCT library, has been proposed. In other words, direct implementation of a single 3-cycle, a pair of 3-cycles and a pair of 2-cycles have been explored and used to propose an efficient Toffoli-based synthesis algorithm for reversible circuits. The synthesis algorithm decomposes a given large cycle into a set of single 3-cycles, pairs of 3-cycles and pair of 2-cycles and synthesizes the resulted cycles directly. Our experimental results show that the proposed synthesis algorithm can outperform the available 2-cycle-based approach about 34% on average. In addition, several discussions for the generalization of the proposed method to the 2m-cycles are given.
Zahra Sasanian, Mehdi Saeedi, Mehdi Sedighi, Morteza Saheb Zamani
ASP-DAC4
2009 Multi-domain clock skew scheduling-aware register placement to optimize clock distribution network
abstract
Multi-domain clock skew scheduling is a cost effective technique for performance improvement. However, the required wire length and area overhead due to phase shifters for realizing such clock scheduler may be considerable if registers are placed without considering assigned skews. Focusing on this issue, in this paper, we propose a skew scheduling-aware register placement algorithm that enables clock tree optimization by considering domains assigned to registers in placement. Our experimental results show that the proposed approach remarkably decreases clock wire length and clock network power consumption at the cost of a slight increase in total wire length.
Naser MohammadZadeh, Minoo Mirsaeedi, Ali Jahanian 0001, Morteza Saheb Zamani
DATE4
2009 Improving Latency of Quantum Circuits by Gate Exchanging
abstract
Quantum circuit design flow consists of two main tasks: synthesis and physical design. In the current flows, two procedures are performed subsequently; synthesis converts the design description into a technology-dependent netlist and then physical design takes the fixed netlist, produces layout, and schedules the netlist on the layout. This style of design suffers from limiting the optimization process in the physical design stage, whereas using a flexible netlist and changing it locally during physical design using layout information often can provide more chance to optimize quantum circuit metrics. Focusing on this issue, in this paper, we propose an optimization flow using gate exchanging heuristic to improve the latency of quantum circuits. We have chosen ion trap technology as the underlying technology to study our flow. Our experimental results show that the proposed flow decreases the latency of quantum circuit by about 23% for the attempted benchmarks.
Naser MohammadZadeh, Morteza Saheb Zamani, Mehdi Sedighi
DSD2
2009 Improved performance and yield with chip master planning design methodology
abstract
Mis-prediction is a dominant problem in nano-scale design that may diminish the quality of physical design algorithms or may even result in failing the design cycle convergence. In this paper, a new planning methodology is presented in which a masterplan of the chip is constructed in early levels of physical design and the rest of succeeding physical design stags operate considering this masterplan. The proposed planning design flow is used to wire planning and buffer resource planning in order to compare with conventional contributions. Experimental results show the considerable improvements in terms of performance, timing yield and buffer usage.
Ali Jahanian 0001, Morteza Saheb Zamani
ACM Great Lakes Symposium on VLSI2
2008 Proposing an efficient method to estimate and reduce crosstalk after placement in VLSI circuits
abstract
Due to the increasing number of elements on a single chip area and the growing complexity of routing, existing methods to reduce crosstalk at the routing or post-routing stage do not seem efficient anymore. So crosstalk estimation should be considered in earlier design stages such as placement. To estimate crosstalk after placement information about the topology of a net and adjacency of its wire-segments should be available. Yet it is not because of the lack of routing information. In this paper, we propose a probabilistic method to estimate intra-grid wirelength of nets after placement and global routing. Results of incorporating this method with the previous crosstalk estimation schemes and our proposed crosstalk reduction method show its efficiency in detecting failing noisy nets before having detailed information of wire adjacency. Our general method improved the number of correctly detected failing noisy nets by 15% on average. In a second improvement, the directed version of this method increased the number of correct detections by 19%. It also decreased the number of false detections considerably.
Arash Mehdizadeh, Morteza Saheb Zamani
AICCSA2
2008 Design space exploration for a coarse grain accelerator
abstract
In the design process of a reconfigurable accelerator employing in an embedded system, multitude parameters may result in remarkable complexity and a large design space. Design space exploration as an alternative to the quantitative approach can be employed to find a right balance between the different design parameters. In this paper, a hybrid approach is introduced to analytically explore the design space for a coarse grain accelerator and determine a wise design point exploiting data extracted from applications, quantitatively. It also provides flexibility for taking into account new design constraints as well as new characteristics of applications. Furthermore, this approach is a methodological approach which reduces the design time and results in a point which satisfies the design goals.
Farhad Mehdipour, Hamid Noori, Morteza Saheb Zamani, Koji Inoue, Kazuaki J. Murakami
ASP-DAC3
2008 Moving forward: A non-search based synthesis method toward efficient CNOT-based quantum circuit synthesis algorithms
abstract
Quantum information processing is in the beginning stages. Among open research problems, quantum circuit synthesis has recently received significant attention. In this paper, we propose a new non-search based moving forward synthesis algorithm (MOSAIC) for CNOT-based quantum circuits. Compared with the widely used search-based methods, MOSAIC is guaranteed to produce a result and can lead to a solution with much fewer steps. To evaluate the proposed algorithms, different circuits taken from the literature are used. The experimental results show the efficiency of the proposed algorithm.
Mehdi Saeedi, Morteza Saheb Zamani, Mehdi Sedighi
ASP-DAC2
2008 A Fast Transformation-Based Synthesis Algorithm for Reversible Circuits
abstract
In this paper, a simple and fast algorithm for the synthesis of reversible circuits is presented. This algorithm considers the synthesis process as a kind of sorting problem, generating a reversible circuit composed of CNOT-based gates. We prove that the proposed algorithm converges for any given specification. The empirical results of realizing examples discussed in the literature are reported. The results show that the algorithm leads to a near optimum solution for all 3*3 specifications and very good results for other larger specifications in much fewer steps compared to the search based and other previous algorithms.
Ehsan K. Ardestani, Morteza Saheb Zamani, Mehdi Sedighi
DSD2
2008 Performance and Timing Yield Enhancement using Highway-on-Chip Planning
abstract
Interconnect mis-prediction is a dominant problem in nanoscale design that may weaken the quality of physical design algorithms or may even increase the design divergence possibility. In this paper, a new interconnect planning technique is presented based highway-on-chip approach. In this methodology, some highways are planned on chip and the location and amount of resource in highways are gradually determined during the placement process. Experimental results show that by using this technique, performance of the attempted benchmarks is improved by 13.66% on average and timing yield of the attempted circuits is improved by 10.02% on average. It is also shown that the results of this technique become better when the size of circuits grows.
Ali Jahanian 0001, Morteza Saheb Zamani
DSD2
2008 Multi-Objective Statistical Yield Enhancement using Evolutionary Algorithm
abstract
It has been shown that several process parameters encounter variation in the very deep submicron era. Due to the increased power and performance variability, a multi-objective variability-aware yield optimization method is crucial. However, most of current yield optimization methods use greedy single objective optimization approach. In this paper, a comprehensive and multi-objective yield optimization framework is proposed to consider the effects of both power and performance yield degradation. In other words, an evolutionary gate sizing optimization approach is introduced to be used in the proposed framework to enhance yield, significantly. Compared with recent yield optimization algorithms and, the proposed framework leads to better results for the attempted circuits. In addition, the independency of proposed framework on selected power and delay analysis techniques makes it suitable for future investigations as using dual threshold voltage assignment approach.
Minoo Mirsaeedi, Morteza Saheb Zamani, Mehdi Saeedi
DSD2
2008 Evaluation and Improvement of Quantum Synthesis Algorithms based on a Thorough Set of Metrics
abstract
Existing synthesis-related cost functions are explored and five fundamental properties of an efficient quantum circuit implementation are introduced. In addition, a thorough set of metrics for quantum circuit synthesis are proposed and applied on some well-known synthesis algorithms. Our analysis reveals the requirement of proposing new synthesis algorithms to produce realizable circuits. A new heuristic is also introduced which improves the results of a commonly used synthesis algorithm in terms of the proposed synthesis-related metrics.
Mehdi Saeedi, Naser MohammadZadeh, Mehdi Sedighi, Morteza Saheb Zamani
DSD4
2008 An Efficient Non-Tree Clock Routing Algorithm for Reducing Delay Uncertainty
abstract
The design of clock distribution networks in synchronous digital systems presents great challenges. In other words, controlling the clock signal delay in the presence of process parameter variations is a major problem in the design of high-speed synchronous circuits. In this paper, an efficient algorithm is presented to improve the tolerance of a clock distribution network against process variation. This algorithm generates a non-tree clock network for reducing uncertainty of the clock signal delay introduced by process variation. The non-tree clock network is generated by inserting cross links in appropriate points taking critical paths into account. To evaluate the effectiveness of the proposed algorithm several HSPICE-based Monte Carlo simulations were done. Experimental results show that the proposed approach can lead to a significant reduction in the deviation of signal propagation delay of critical paths and the skew variability without considerable increase in the total wirelength.
Morteza Saheb Zamani, Maryam Taajobian, Mehdi Saeedi
DSD1
2008 An architecture framework for an adaptive extensible processor
Hamid Noori, Farhad Mehdipour, Kazuaki J. Murakami, Koji Inoue, Morteza Saheb Zamani
J. Supercomput.5
2007 Algebraic Characterization of CNOT-Based Quantum Circuits with its Applications on Logic Synthesis
abstract
The exponential speed up of quantum algorithms and the fundamental limits of current CMOS process for future design technology have directed attentions toward quantum circuits. In this paper, the matrix specification of a broad category of quantum circuits, i.e. CNOT-based circuits, are investigated. We prove that the matrix elements of CNOT-based circuits can only be zeros or ones. In addition, the columns or rows of such a matrix have exactly one element with the value of 1. Furthermore, we show that these specifications can be used to synthesize CNOT-based quantum circuits. In other words, a new scheme is introduced to convert the matrix representation into its SOP equivalent using a novel quantum-based Karnaugh map extension. We then apply a search-based method to transform the obtained SOP into a CNOT-based circuit. Experimental results prove the correctness of the proposed concept.
Mehdi Saeedi, Morteza Saheb Zamani, Mehdi Sedighi
DSD2
2007 Improved timing closure by early buffer planning in floor-placement design flow
abstract
Buffer insertion plays an increasingly critical role on circuit performance and signal integrity especially in deep submicron technologies. Buffer insertion stage is very important for buffering efficiency. Early buffer insertion may cause misestimating due to unknown cell locations whereas buffer insertion after placement may not be very effective because the cell locations are fixed and buffer resources may be distributed inappropriately.In this paper, a buffer planning algorithm for floor-placement design flow is presented which creates a map of buffer requirements in various regions of the design at the floorplanning stage based on the statistical distribution of critical paths and enforces the placer to distribute white spaces with respect to the estimated buffer requirement map.Experimental results show that the proposed method improves the performance of experimented circuits with smaller number of buffers and better power consumption compare to conventional methods. Furthermore, power-delay product has been improved considerably, especially for large circuits with a small growth in CPU time.
Ali Jahanian 0001, Morteza Saheb Zamani
ACM Great Lakes Symposium on VLSI2
2007 An efficient net ordering algorithm for buffer insertion
abstract
There are efficient algorithms for net-based buffer insertion but they lead to sub-optimal path delays or unnecessarily large number of buffers due to their lack of global view. This can increase power consumption as well as die area. The ordering of nets for buffer insertion has a crucial impact on the quality of buffering in terms of path delay and the number of used buffers. A good net ordering can extend the local view of any net-based buffer insertion algorithm.In this paper, an efficient O(nlogn) algorithm for net ordering is presented. The net ordering problem is mapped to traditional knapsack problem to obtain an efficient ordering. Experimental results show that our algorithm can meet timing constraints with an 18.8% reduction in the number of buffers on average.
Hamid Reza Kheirabadi, Morteza Saheb Zamani
ACM Great Lakes Symposium on VLSI2
2007 A novel synthesis algorithm for reversible circuits
abstract
In this paper, a new non-search based synthesis algorithm for reversible circuits is proposed. Compared with the widely used search-based methods, our algorithm is guarantied to produce a result and can lead to a solution with much fewer steps. To evaluate the proposed method, several circuits taken from the literature are used. The experimental results corroborate the expected findings.
Mehdi Saeedi, Mehdi Sedighi, Morteza Saheb Zamani
ICCAD3
2007 An efficient heterogeneous reconfigurable functional unit for an adaptive dynamic extensible processor
abstract
Replacing functional units of an extensible processor with reconfigurable fnctional units enhances performance and flexibility ofprocessors to execute custom instructions. That is due to the ability ofreconfigurable fnctional units to perform computations in hardware to increase performance, while retaining much of the flexibility of a software solution. In this paper, we develop a heterogeneous architecture for the reconfigurable fnctional unit of an extensible processor. To verify the efficiency of our architecture, we applied it to 8 applications of Mibench. Our experiments show that compared to the similar architectures, ours supports a wide range of custom instructions. In addition, use of the new architecture improves execution time of custom instructions by 20% to 30% on average. Moreover, compared with the previous architecture, area is reduced by 15%.
Arash Mehdizadeh, Behnam Ghavami, Morteza Saheb Zamani, Hossein Pedram, Farhad Mehdipour
VLSI-SoC3
2006 Custom Instruction Generation Using Temporal Partitioning Techniques for a Reconfigurable Functional Unit
Farhad Mehdipour, Hamid Noori, Morteza Saheb Zamani, Kazuaki J. Murakami, Koji Inoue, Mehdi Sedighi
EUC3
2006 A Reconfigurable Functional Unit for an Adaptive Dynamic Extensible Processor
abstract
This paper presents a reconfigurable functional unit (RFU) for an adaptive dynamic extensible processor. The processor can tune its extended instructions to the target applications, after chip-fabrication. The custom instructions (CIs) are generated deploying the hot basic blocks during the training mode. In the normal mode, CIs are executed on the RFU. A quantitative approach was used for designing the RFU. The RFU is a matrix of functional units with 8 inputs and 6 outputs. Performance is enhanced up to 1.25 using the proposed RFU for 22 applications of Mibench. This processor needs no extra opcodes for CIs, new compiler, source code modification and recompilation.
Hamid Noori, Farhad Mehdipour, Kazuaki J. Murakami, Koji Inoue, Morteza Saheb Zamani
FPL5
2006 Reducing reconfiguration time of reconfigurable computing systems in integrated temporal partitioning and physical design framework
abstract
In reconfigurable systems, reconfiguration latency is a very important factor which impact the system performance. In this paper, a framework is proposed that integrates the temporal partitioning and physical design phases to perform a static compilation process for reconfigurable computing systems. A temporal partitioning algorithm is proposed which attempts to decrease the time of reconfiguration on a partially reconfigurable hardware. This algorithm attempts to find similar single or pair of operations between subsequent partitions. Considering similar pairs instead of single nodes brings about less complexity for routing process. By using this technique, smaller reconfiguration bit-stream is obtained, which directly decreases the reconfiguration overhead time at the run-time. A complementary algorithm attempts to increase the similarity of subsequent partitions by searching for similar pairs and using a technique called dummy node insertion. An incremental physical design process based on similar configurations produced in the partitioning stage improves the metrics over iterations.
Farhad Mehdipour, Morteza Saheb Zamani, Hamid Reza Ahmadifar, Mehdi Sedighi, Kazuaki J. Murakami
IPDPS2
2006 Prediction and reduction of routing congestion
abstract
Routing congestion is a critical issue in deep submicron design technology and it becomes one of the most challenging problems in today's design flow. This paper presents a true probabilistic congestion prediction method based on router's intelligence to be used in the placement stage of physical design flow. Experimental results show that for IBM-PLACE benchmarks, our prediction algorithm estimates the congestion more accurately than a recent method by about 19%. Furthermore, a new congestion reduction algorithm is presented which is based on contour plotting. Our experiments show that our algorithm reduces congestion by about 28% on average. In addition, comparing our results with a recent approach shows that our reduction technique reduces congestion more by about 13%.
Mehdi Saeedi, Morteza Saheb Zamani, Ali Jahanian 0001
ISPD2
2005 A novel reconfigurable hardware architecture for IP address lookup
abstract
IP address lookup is one of the most challenging problems of Internet routers. In this paper, an IP lookup rate of 263 Mlps (Million lookups per second) is achieved using a novel architecture on reconfigurable hardware platform. A partial reconfiguration may be needed for a small fraction of route updates. Prefixes can be added or removed at a rate of 2 million updates per second, including this hardware reconfiguration overhead. A route update may fail due to the physical resource limitations. In this case, which is rare if the architecture is properly configured initially, a full reconfiguration is needed to allocate more resources to the lookup unit.
Hamid Fadishei, Morteza Saheb Zamani, Masoud Sabaei
ANCS2
2005 Reducing Inter-Configuration Memory Usage and Performance Improvement in Reconfigurable Computing Systems
abstract
For running subsequent configurations in a reconfigurable computing system intermediate data must transfer between them. Reducing memory usage overhead can result in reduction in the array size and the number of input/output pins. In this paper, a new iterative design flow is proposed which integrates the synthesis and physical design aspects for performing a static compilation process. A new temporal partitioning algorithm for partitioning and scheduling is proposed, which tries to increase similarity of subsequent configurations in such a way that the reconfiguration time on a partially reconfigurable hardware decreases. In addition, we perform an iterative physical design process based on similar configurations produced in the previous stage. A modified algorithm improves our prior temporal partitioning algorithm, which usually had large overhead of memory usage and the number of input/output pins. This new approach performs partitioning in depth and tries to minimize the memory and 10 requirements.
Farhad Mehdipour, Morteza Saheb Zamani, Mehdi Sedighi
DSD2
2005 Parallel Hardware Implementation of Cellular Learning Automata Based Evolutionary Computing (CLA-EC) on FPGA
abstract
The CLA-EC is a model obtained by combining the concepts of cellular learning automata and evolutionary algorithms. The parallel structure of the CLA-EC makes it suitable for hardware-based applications including evolvable hardware. In this paper, based on the SIMD model, a parallel architecture is proposed and implemented on FPGA. Simulation results show that the proposed architecture can solve optimization problems thousands times faster than the sequential implementations.
Arash Hariri, Reza Rastegar, Morteza Saheb Zamani, Mohammad Reza Meybodi
FCCM3
2005 A Reconfigurable Architecture for Implementing Multiple Cipher Algorithms
Ali Valizadeh, Morteza Saheb Zamani, Babak Sadeghian, Farhad Mehdipour
FPT2
2003 Rectilinear floorplanning of FPGAs using Kohonen map
abstract
In this paper, we present an algorithm for the floorplanning of FPGs (field-programmable gate arrays). The algorithm uses Kohonen self-organizing map to floorplan regular structures with soft modules. An abstract specification of the design is converted to a set of appropriate input vectors, which are fed to the network. At the end of the process, the map shows a two-dimensional plane of the design in which the modules with high connectivity are placed adjacent to each other, hence minimizing total connection length in the design. Unlike conventional floorplanning algorithms, which are limited to rectangular shapes, our approach can produce rectilinear modules. Using Kohonen map enables the algorithm to do 3-dimensional floorplanning.
Morteza Saheb Zamani, Masoud Soleimani
IJCNN1
1999 An efficient method for placement of VLSI designs with Kohonen map
abstract
In this paper a Kohonen map-based algorithm for the placement of gate arrays and standard cells is presented. An abstract specification of the design is converted to a set of appropriate input vectors using a mathematical method, called "multidimensional scaling". These vectors which have, in general, higher dimensionality are fed to the self-organizing map at random in order to map them onto a 2D plane of the regular chip. The mapping is done in such a way that the cells with higher connectivity are placed close to each other, hence minimizing total connection length in the design. Two processes, called reassignment and rearrangement, are employed to make the algorithm applicable to the standard cell designs. In addition to the small examples introduced in other papers, two standard cell benchmarks were tried and better results were observed for these large designs compared to other neural net-barred approaches.
Morteza Saheb Zamani, Farhad Mehdipour
IJCNN1
1995 A neural network approach to the placement problem
abstract
No abstract available.
Morteza Saheb Zamani, Graham R. Hellestrand
ASP-DAC1
1995 A Stepwise Refinement Algorithm for Integrated Floorplanning, Placement and Routing of Hierarchical Designs
abstract
This paper presents a stepwise refinement approach to floorplanning, followed by the placement and routing of hierarchical VLSI circuits. The algorithm consists of several traversals of the design hierarchy, from each of which more accurate information about the design is obtained before the actual placement and routing is performed. Interdependencies between different levels of the hierarchy are considered in this approach. The algorithm is capable of being applied to large circuits with many levels of hierarchy and with many modules at each level.
Morteza Saheb Zamani, Graham R. Hellestrand
ISCAS1