Kiamal Z. Pekmestzi

dblp:03/1577 · DBLP profile ↗
← Back
39ranked-venue papers
4as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 6 · 1 first-authorSecurity and privacy · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 Systematic Embedded Development and Implementation Techniques on Intel Myriad VPUs
abstract
The worldwide demand for speed in applications challenges the deployment of compute-intensive algorithms at the power-constrained edge. Novel embedded devices such as the heterogeneous Vision Processing Units (VPUs) emerge as a promising solution for low-power embedded imaging/vision applications, as they accelerate computer vision algorithms and convolutional neural networks with only 1-2W. In this brief, we propose a development methodology for exploiting the full potential of the VPU heterogeneity and providing sufficient acceleration within their restricted power envelope. Based on this methodology, we demonstrate the development paradigm on the Myriad VPUs and report experimental results from the implementation of demanding image processing kernels.
Vasileios Leon, Kiamal Z. Pekmestzi, Dimitrios Soudris
VLSI-SoC2
2021 Exploiting the Potential of Approximate Arithmetic in DSP & AI Hardware Accelerators
abstract
Approximate computing is an emerging design paradigm, which exploits the inherent error resilience of numerous applications to improve their energy efficiency and/or performance. The current paper focuses on applications from the DSP and AI domains, and examines the impact of arithmetic approximations on accelerators for FPGA and ASIC technologies. Based on our design methodology, we implement and evaluate approximate architectures for image processing, signal filtering, telecommunication digital functions, and convolutional neural networks. The evaluation shows that sophisticated bit-level optimizations and disciplined approximations deliver significant gains in the hardware resources and performance of the accelerator in exchange for small errors and tunable accuracy loss.
Vasileios Leon, Kiamal Z. Pekmestzi, Dimitrios Soudris
FPL2
2021 Improving Power of DSP and CNN Hardware Accelerators Using Approximate Floating-point Multipliers
abstract
Approximate computing has emerged as a promising design alternative for delivering power-efficient systems and circuits by exploiting the inherent error resiliency of numerous applications. The current article aims to tackle the increased hardware cost of floating-point multiplication units, which prohibits their usage in embedded computing. We introduce AFMU (Approximate Floating-point MUltiplier), an area/power-efficient family of multipliers, which apply two approximation techniques in the resource-hungry mantissa multiplication and can be seamlessly extended to support dynamic configuration of the approximation levels via gating signals. AFMU offers large accuracy configuration margins, provides negligible logic overhead for dynamic configuration, and detects unexpected results that may arise due to the approximations. Our evaluation shows that AFMU delivers energy gains in the range 3.6%–53.5% for half-precision and 37.2%–82.4% for single-precision, in exchange for mean relative error around 0.05%–3.33% and 0.01%–2.20%, respectively. In comparison with state-of-the-art multipliers, AFMU exhibits up to 4–6× smaller error on average while delivering more energy-efficient computing. The evaluation in image processing shows that AFMU provides sufficient quality of service, i.e., more than 50 db PSNR and near 1 SSIM values, and up to 57.4% power reduction. When used in floating-point CNNs, the accuracy loss is small (or zero), i.e., up to 5.4% for MNIST and CIFAR-10, in exchange for up to 63.8% power gain.
Vasileios Leon, Theodora Paparouni, Evangelos Petrongonas, Dimitrios Soudris, Kiamal Z. Pekmestzi
ACM Trans. Embed. Comput. Syst.5
2020 Efficient design of magnitude and 2's complement comparators
Fotios Ntouskas, Costas Efstathiou, Kiamal Z. Pekmestzi
Integr.3
2019 Cooperative Arithmetic-Aware Approximation Techniques for Energy-Efficient Multipliers
abstract
Approximate computing appears as an emerging and promising solution for energy-efficient system designs, exploiting the inherent error-tolerant nature of various applications. In this paper, targeting multiplication circuits, i.e., the energy-hungry counterpart of hardware accelerators, an extensive exploration of the error--energy trade-off, when combining arithmetic-level approximation techniques, is performed for the first time. Arithmetic-aware approximations deliver significant energy reductions, while allowing to control the error values with discipline by setting accordingly a configuration parameter. Inspired from the promising results of prior works with one configuration parameter, we propose 5 hybrid design families for approximate and energy-friendly hardware multipliers, consisting of two independent parameters to tune the approximation levels. Interestingly, the resolution of the state-of-the-art Pareto diagram is improved, giving the flexibility to achieve better energy gains for a specific error constraint imposed by the system. Moreover, we outperform prior works in the field of approximate multipliers by up to 60% energy reduction, and thus, we define the new Pareto front.
Vasileios Leon, Konstantinos Asimakopoulos, Sotirios Xydis, Dimitrios Soudris, Kiamal Z. Pekmestzi
DAC5
2018 Approximate Hybrid High Radix Encoding for Energy-Efficient Inexact Multipliers
abstract
Approximate computing forms a design alternative that exploits the intrinsic error resilience of various applications and produces energy-efficient circuits with small accuracy loss. In this paper, we propose an approximate hybrid high radix encoding for generating the partial products in signed multiplications that encodes the most significant bits with the accurate radix-4 encoding and the least significant bits with an approximate higher radix encoding. The approximations are performed by rounding the high radix values to their nearest power of two. The proposed technique can be configured to achieve the desired energy-accuracy tradeoffs. Compared with the accurate radix-4 multiplier, the proposed multipliers deliver up to 56% energy and 55% area savings, when operating at the same frequency, while the imposed error is bounded by a Gaussian distribution with near-zero average. Moreover, the proposed multipliers are compared with state-of-the-art inexact multipliers, outperforming them by up to 40% in energy consumption, for similar error values. Finally, we demonstrate the scalability of our technique.
Vasileios Leon, Georgios Zervakis 0001, Dimitrios Soudris, Kiamal Z. Pekmestzi
IEEE Trans. Very Large Scale Integr. Syst.4
2018 VOSsim: A Framework for Enabling Fast Voltage Overscaling Simulation for Approximate Computing Circuits
Georgios Zervakis 0001, Fotios Ntouskas, Sotirios Xydis, Dimitrios Soudris, Kiamal Z. Pekmestzi
IEEE Trans. Very Large Scale Integr. Syst.5
2016 Design of Efficient 1's Complement Modified Booth Multiplier
abstract
The 2's complement representation is widely adopted, since compared to the other signed number systems has the advantage of simpler addition and single representation of zero. Sing-magnitude representation is used in digital signal processors for the representation of digital signals for low-power purposes. The 1's complement representation compared to the 2's complement one has the advantages of the simpler conversion to and from the sign-magnitude representation, simpler negation and that truncation of negative numbers is equivalent to that of the sign-magnitude representation. Therefore, the design of efficient arithmetic units for this system should be examined. In this work 1's complement modified Booth multipliers with complexity similar to that of the 2's complement ones are proposed.
Kiamal Z. Pekmestzi, Costas Efstathiou
DSD1
2016 Pre-Encoded Multipliers Based on Non-Redundant Radix-4 Signed-Digit Encoding
abstract
In this paper, we introduce an architecture of pre-encoded multipliers for digital signal processing applications based on off-line encoding of coefficients. To this extend, the Non-Redundant radix-4 Signed-Digit (NR4SD) encoding technique, which uses the digit values$\lbrace -1,0,+1,+2 \rbrace$or$\lbrace -2,-1,0,+1 \rbrace$, is proposed leading to a multiplier design with less complex partial products implementation. Extensive experimental analysis verifies that the proposed pre-encoded NR4SD multipliers, including the coefficients memory, are more area and power efficient than the conventional Modified Booth scheme.
Kostas Tsoumanis, Nicholas Axelos, Nikolaos Moschopoulos, Georgios Zervakis 0001, Kiamal Z. Pekmestzi
IEEE Trans. Computers5
2016 Flexible DSP Accelerator Architecture Exploiting Carry-Save Arithmetic
abstract
Hardware acceleration has been proved an extremely promising implementation strategy for the digital signal processing (DSP) domain. Rather than adopting a monolithic application-specific integrated circuit design approach, in this brief, we present a novel accelerator architecture comprising flexible computational units that support the execution of a large set of operation templates found in DSP kernels. We differentiate from previous works on flexible accelerators by enabling computations to be aggressively performed with carry-save (CS) formatted data. Advanced arithmetic design concepts, i.e., recoding techniques, are utilized enabling CS optimizations to be performed in a larger scope than in previous approaches. Extensive experimental evaluations show that the proposed accelerator architecture delivers average gains of up to 61.91% in area-delay product and 54.43% in energy consumption compared with the state-of-art flexible datapaths.
Kostas Tsoumanis, Sotirios Xydis, Georgios Zervakis 0001, Kiamal Z. Pekmestzi
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Design-Efficient Approximate Multiplication Circuits Through Partial Product Perforation
abstract
Approximate computing has received significant attention as a promising strategy to decrease power consumption of inherently error tolerant applications. In this paper, we focus on hardware-level approximation by introducing the partial product perforation technique for designing approximate multiplication circuits. We prove in a mathematically rigorous manner that in partial product perforation, the imposed errors are bounded and predictable, depending only on the input distribution. Through extensive experimental evaluation, we apply the partial product perforation method on different multiplier architectures and expose the optimal architecture-perforation configuration pairs for different error constraints. We show that, compared with the respective exact design, the partial product perforation delivers reductions of up to 50% in power consumption, 45% in area, and 35% in critical delay. In addition, the product perforation method is compared with the state-of-the-art approximation techniques, i.e., truncation, voltage overscaling, and logic approximation, showing that it outperforms them in terms of power dissipation and error.
Georgios Zervakis 0001, Kostas Tsoumanis, Sotirios Xydis, Dimitrios Soudris, Kiamal Z. Pekmestzi
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Approximate Multiplier Architectures Through Partial Product Perforation: Power-Area Tradeoffs Analysis
abstract
Approximate computing has received significant attention as a promising strategy to decrease power consumption of inherently error-tolerant applications. Hardware approximation mainly targets arithmetic units, e.g. adders and multipliers. In this paper, we design new approximate hardware multipliers and propose the Partial Product Perforation technique, which omits a number of consecutive partial products by perforating their generation. Through extensive experimental evaluation, we apply the partial product perforation method on different multiplier architectures and expose the optimal configurations for different error values. We show that the partial product perforation delivers reductions of up to 50% in power consumption, 45% in area and 35% in critical delay. Also, the product perforation method is compared with state-of-the-art works on approximate computing that consider the Voltage Over-Scaling (VOS) and logic approximation (i.e. design of approximate compressors) techniques, outperforming them in terms of power dissipation by up to 17% and 20% on average respectively. Finally, with respect to the aforementioned gains, the error value delivered by the proposed product perforation method is smaller by 70% and 99% than the VOS and logic approximation methods respectively.
Georgios Zervakis 0001, Kostas Tsoumanis, Sotirios Xydis, Nicholas Axelos, Kiamal Z. Pekmestzi
ACM Great Lakes Symposium on VLSI5
2015 Low leakage radiation tolerant CAM/TCAM cell
abstract
In this paper we propose a leakage-aware soft error tolerant storage element, implementable in standard CMOS technology and able to operate both as a CAM and as a TCAM cell. The proposed cell is immune to SNUs (Single Node Upsets) when operating as a CAM cell and demonstrates partial resilience (75%) when operating as a TCAM cell. Simulation results in SPICE at a 45nm PTM technology show a significant reduction in leakage dissipation compared to the standard but unprotected 6T-based TCAM cell as well as compared to conventional DICE-based CAM/TCAM solutions.
Nikolaos Eftaxiopoulos-Sarris, Nicholas Axelos, Kiamal Z. Pekmestzi
IOLTS3
2015 Hybrid approximate multiplier architectures for improved power-accuracy trade-offs
abstract
Approximate computing forms a promising design alternative for inherently error resilient applications, trading accuracy for power savings. In this paper, we exploit multi-level approximation, i.e. at the algorithmic, the logic and the circuit level, to design low power approximate arithmetic architectures for hardware multipliers. Motivated from the limited power savings that approximation techniques can achieve in isolation, we explore hybrid methods that apply simultaneously more than one techniques from different layers. We introduce the concept of perforation for approximate arithmetic circuit design and we explore the newly defined design space of hybrid designs showing that it leads to lower power consumption at every examined error range. To address the increased complexity of the target design space, we introduce an heuristic optimization technique and the corresponding design framework that automatically generates hybrid low-power approximate multipliers requiring a small number of design evaluations, i.e. synthesis, simulation, power and timing analysis. Through extensive experimentation, we show that the proposed techniques converge towards optimal solutions and deliver approximate designs that are always more efficient with respect to state-of-art approaches. Power savings of 11% are reported for small error bounds and more than 30% in case of more relaxed error constraints.
Georgios Zervakis 0001, Sotirios Xydis, Kostas Tsoumanis, Dimitrios Soudris, Kiamal Z. Pekmestzi
ISLPED5
2014 A segmentation-based BISR scheme
abstract
With memory estate increasing in System-On-Chips and highly integrated products, memory defects and wearout effects are the determining factor in the chip's yield loss and reliability. In this paper, a multiple cache-based Built-in Self-Repair scheme is proposed that is able to repair from the word level down to the bit level. Moreover, it is proved that the level of segmentation does not affect the repair efficiency. An exploration is then conducted to find the optimal scheme in terms of area overhead.
Georgios Zervakis 0001, Nikolaos Eftaxiopoulos-Sarris, Kostas Tsoumanis, Nicholas Axelos, Kiamal Z. Pekmestzi
ASP-DAC5
2014 FF-DICE: An 8T soft-error tolerant cell using Independent Dual Gate SOI FinFETs
abstract
In this paper we present FF-DICE (Footless-FinFET-DICE), an 8T footless storage element that exhibits soft error resilience characteristics to Single Event Upsets. The proposed cell utilises IDG (Independent Dual Gate) FinFETs to merge the functions of a typical cell's NMOS drivers and NMOS access transistors, thereby saving 33% of the transistors required for a typical DICE cell. Given the IDG FinFET dual gate mutual coupling and excellent control over the transistor conductive channel, simulations show that the proposed cell can operate as a regular memory cell, with Static Voltage Noise Margin of 341mV, Static Current Noise Margin of 15uA and provides soft-error immunity to particle strikes on single nodes, as well as considerable area savings compared to similar designs.
Nicholas Axelos, Nikolaos Eftaxiopoulos-Sarris, Georgios Zervakis 0001, Kostas Tsoumanis, Kiamal Z. Pekmestzi
IOLTS5
2014 Efficient modulo 2n+1 multiply and multiply-add units based on modified Booth encoding
Costas Efstathiou, Nikolaos Moschopoulos, Nicholas Axelos, Kiamal Z. Pekmestzi
Integr.4
2013 Efficient modulo 2n+1 multiplication for the idea block cipher
abstract
International Data Encryption Algorithm (IDEA) is a popular and secure cryptography algorithm, suitable for hardware implementation. IDEA comprises of modulo 216 additions, bitwise exclusive-OR operations and modulo 216+1 multiplications of 16-bit words. Among them, modulo 216+1 multiplication is the most time, space and power consuming operation. In this work, we propose an efficient modulo 2n+1 modified Booth multiplication algorithm which is adapted to operands used in the IDEA. The IDEA multiplier based on the proposed modulo 2n+1 multiplication algorithm yields area and power advantages of up to 12% and 14% respectively, compared to the already proposed modulo 2n+1 multiplier designs. The implementation of a single round of the IDEA block cipher based on the proposed multiplier verifies the area and power advantages over the implementations based on existing modulo 2n+1 multipliers.
Kiamal Z. Pekmestzi, Costas Efstathiou, Nikolaos Moschopoulos, Kostas Tsoumanis
ACM Great Lakes Symposium on VLSI1
2013 A radiation tolerant and self-repair memory cell
abstract
In this paper a new radiation tolerant memory cell is proposed. As CMOS technologies scale down, existing hardened cells, which are technology dependent, become more and more vulnerable to radiation effects. We tested some of the most known and effective hardened cells using the SPICE simulator LTspice but no one proved to ensure data integrity. The proposed cell consists of three standard 6T cells and proves to be 100% radiation tolerant in any technology, having however an expected area and power overhead comparing to the 6T and DICE cells. According to simulation results, these overheads are proportional to the number of transistors used, but the read time when no error has occurred and the write time are shorter.
Nikolaos Eftaxiopoulos-Sarris, Georgios Zervakis 0001, Kostas Tsoumanis, Kiamal Z. Pekmestzi
IOLTS4
2013 On the design of modulo 2n±1 residue generators
abstract
In this paper, we propose an efficient residue generator which concurrently computes the residues modulo 2n+1 and modulo 2n-1. The input operands are divided into n-bit vectors which are then grouped into two sets and added by two separate Carry Save Adder (CSA) trees. The output carry of each stage of the first CSA tree is used as input to the corresponding stage of the second CSA tree, while the output vectors of the trees are finally added modulo 2n+1 and 2n-1 to compute the residues. The proposed residue generator is well suited for Residue Number System (RNS) based applications which use both modulo 2n+1 and 2n-1 residues. An efficient configurable modulo 2n±1 residue generator is also proposed.
Kostas Tsoumanis, Costas Efstathiou, Nikolaos Moschopoulos, Kiamal Z. Pekmestzi
VLSI-SoC4
2013 A column parity based fault detection mechanism for FIFO buffers
Isidoros Sideris, Kiamal Z. Pekmestzi
Integr.2
2012 On the Design of Configurable Modulo 2n±1 Residue Generators
abstract
In this work new efficient modulo 2n+1 residue generators are proposed. The input operands are divided into n-bit vectors which are added by an inverted end around carry save adder tree and a final stage diminished-1 modulo 2n+1 adder. The conversion of the proposed residue generators to configurable modulo 2n±1 ones is also discussed. Modulo 2n±1 residue generators find applicability as forward converts from the binary to the residue number system, and in the design of self-checking digital systems.
Costas Efstathiou, Nikolaos Moschopoulos, Kostas Tsoumanis, Kiamal Z. Pekmestzi
DSD4
2012 Cost Effective Protection Techniques for TCAM Memory Arrays
abstract
This paper presents low cost techniques for error detection and correction in Ternary Content Addressable Memories (TCAMs). The techniques exploit the inherent redundancy of TCAM cells to allow for protection at lower cost. A fault detection technique with the cost of parity but with about the half probability of silent data corruption is proposed. This technique is then applied at both horizontal and vertical dimensions of the TCAM array, and a low cost error correction scheme is derived. Last, another error correction scheme is proposed, which employs a SECDED ECC of the half complexity, by making use of the TCAM redundancy, without compromising single bit error correction. The proposed schemes come with minimal area, power, and critical path overheads, in comparison with standard schemes, and they are good alternatives for TCAM arrays protection.
Isidoros Sideris, Kiamal Z. Pekmestzi
IEEE Trans. Computers2
2012 Compiler-in-the-loop exploration during datapath synthesis for higher quality delay-area trade-offs
abstract
Design space exploration during high-level synthesis targets the computation of those design solutions which form optimal trade-off points. This quest for optimal trade-offs has been focused on studying the impact of various architectural-level parameters during high-level synthesis algorithms, silently neglecting the trade-offs produced from the combined impact of behavioral-level together with architectural-level parameters. We propose a novel design space, exploration methodology that studies an extended instance of the solution space considering the effects of combining compiler- and architectural-level transformations. It is shown that exploring the design space in a global manner reveals new trade-off points, thus shifting towards higher quality design solutions. We use a combination of upper-bounding conditions together with gradient-based heuristic pruning to efficiently traverse the extended search space. Our exploration framework delivers significant quality improvements without compromising the optimality (Pareto accuracy) of the discovered solutions, together with significant runtime reductions compared to exploring exhaustively the solution space at every allocation scenario.
Sotirios Xydis, Kiamal Z. Pekmestzi, Dimitrios Soudris, George Economakos
ACM Trans. Design Autom. Electr. Syst.2
2012 Efficient Memory Repair Using Cache-Based Redundancy
abstract
In modern processes, conventional defect density and variability related yield losses are a major concern for the aggressive memory designs in integrated circuits. Synergistic action for memory repair at the circuit and architectural level is essential to maintain the yields and profitability of past technology nodes. In this paper, we propose a scalable memory repair architecture that utilizes a set of direct-mapped cache banks to replace faulty words. Statistical and mathematical probability analysis shows that the proposed scheme achieves high repairability levels with low area and static power dissipation overheads, the latter being a dominant issue in nanometer technologies. It is therefore a suitable solution along with other mature memory repair techniques, to enhance the overall repairability features and guarantee the correct and reliable operation of embedded memories in nanometer technologies.
Nicholas Axelos, Kiamal Z. Pekmestzi, Dimitris Gizopoulos
IEEE Trans. Very Large Scale Integr. Syst.2
2011 On the Design of Modulo 2^n+1 Multipliers
abstract
In this work a new efficient modulo 2n+1 modified Booth multiplication algorithm for operands in the weighted representation is proposed. According to our algorithm [n/2]+2 partial products are derived. The resulting partial products are reduced by an inverted end around carry save adder tree to two operands, which are finally added by a diminished-1 modulo 2n+1 adder. Our design compares favorably for both area and delay to the modulo 2n+1 modified Booth multipliers previously proposed.
Costas Efstathiou, Kiamal Z. Pekmestzi, Nicholas Axelos
DSD2
2011 High Performance and Area Efficient Flexible DSP Datapath Synthesis
abstract
This paper presents a new methodology for the synthesis of high performance flexible datapaths, targeting computationally intensive digital signal processing kernels of embedded applications. The proposed methodology is based on a novel coarse-grained reconfigurable/flexible architectural template, which enables the combined exploitation of the horizontal and vertical parallelism along with the operation chaining opportunities found in the application's behavioral description. Efficient synthesis techniques exploiting these architectural optimization concepts from a higher level of abstraction are presented and analyzed. Extensive experimentation showed average latency and area reductions up to 33.9% and 53.9%, respectively, and higher hardware area utilization, compared to previously published high performance coarse-grained reconfigurable datapaths.
Sotirios Xydis, George Economakos, Dimitrios Soudris, Kiamal Z. Pekmestzi
IEEE Trans. Very Large Scale Integr. Syst.4
2010 A bit level area aware cache-based architecture for memory repairs
abstract
In this paper an architecture for memory Built-In Self-Repair (BISR) is presented. The proposed scheme utilises a multiple bank cache-like memory for repairing defective bits and is a heavily modified version of a previously studied architecture that repaired at the word level, while the proposed repairs at the bit level. This scheme achieves high repair ratios at high defect densities with small overheads. The two architectures are compared on their footprint by means of an area approximation analysis. The proposed modified scheme shows significant area savings (greater than 4 times) while retaining the same high repairability qualities as its predecessor.
Nicholas Axelos, Kiamal Z. Pekmestzi
IOLTS2
2010 Designing efficient DSP datapaths through compiler-in-the-loop exploration methodology
abstract
This paper proposes a compiler-in-the-loop exploration framework during architectural DSP synthesis. We extend the conventional design space, considering code level transformations together with architectural level optimizations and their impact on the scheduled datapath. We show that the proposed methodology explores the design space more globally in comparison with existing methods. New trade-off points are revealed and Pareto curve shifting towards higher quality design solutions is performed.
Sotirios Xydis, Christodoulos Skouroumounis, Kiamal Z. Pekmestzi, Dimitrios Soudris, George Economakos
ISCAS3
2009 Designing coarse-grain reconfigurable architectures by inlining flexibility into custom arithmetic data-paths
Sotirios Xydis, George Economakos, Kiamal Z. Pekmestzi
Integr.3
2008 A high-speed radix-4 multiplexer-based array multiplier
abstract
This paper presents a new radix-4 multiplexer-based array multiplier, based on a multiplication scheme shown in a previous work, where 4-to-1 multiplexers are used for the computation of partial products. In the proposed design, the rows of the array are reduced to the half, compared to the initial multiplexer-based scheme, as two bits from both operands are processed at each step. The proposed scheme is compared to the Modified-Booth array multiplier and to the initial multiplexer-based array scheme. The compared designs are coded in VHDL and synthesized using the TSMC 0.13¼m technology library. The synthesis results of critical time and area show 11-22% improvement in critical time delay compared to the Modified-Booth array, in the expense of an area overhead of 3.8-16%. Compared to the initial multiplexer-based scheme, there is a significant improvement in terms of area and critical time.
Dimitris Bekiaris, Kiamal Z. Pekmestzi, Christos A. Papachristou
ACM Great Lakes Symposium on VLSI2
2008 A BISR Architecture for Embedded Memories
abstract
In this paper a BISR architecture for embedded memories is presented. The proposed scheme utilises a multiple bank cache-like memory for repairs. Statistical analysis is used for minimisation of the total resources required to achieve a very high fault coverage. Simulation results show that the proposed BISR scheme is characterised by high efficiency and low area overhead, even for high defect densities. On a 4 Mbit memory and an average number of 1024 memory defects per IC, a repair ratio of 100% and over 90% require less than 2% and 1% memory overhead respectively.
Kiamal Z. Pekmestzi, Nicholas Axelos, Isidoros Sideris, Nikolaos Moschopoulos
IOLTS1
2008 A predecoding technique for ILP exploitation in Java processors
Isidoros Sideris, Kiamal Z. Pekmestzi, George Economakos
J. Syst. Archit.2
2007 An Elliptic Curve Cryptosystem Design Based on FPGA Pipeline Folding
abstract
In this paper we present an efficient design technique for implementing the elliptic curve cryptographic (ECC) scheme in FPGAs. Our technique is based on a novel and efficient implementation of modular multiplication which is the core operation of ECC. To implement large bit-length multiplications we used a novel partitioning and pipeline folding scheme to fit at least 256-bit modular multiplications on a single Virtex-4 FPGA. Comparisons to several other schemes are presented.
Osama Daifallah Al-Khaleel, Christos A. Papachristou, Francis Wolff, Kiamal Z. Pekmestzi
IOLTS4
2006 FPGA-based Design of a Large Moduli Multiplier for Public Key Cryptographic Systems
abstract
High secure cryptographic systems require large bit-length encryption keys which presents a challenge to their efficient hardware implementation especially in embedded devices. Modular multiplication is the core operation in well known cryptosystems like RSA and elliptic curve (ECC). Therefore, it is important to employ efficient modular multiplications techniques to improve the overall performance of the cryptographic system. We present a modular multiplier based on the ordinary Montgomery's multiplication algorithm and a new array multiplication scheme to perform the multiplication. The new modular multiplier is scalable and can be used for large bit-lengths. We also implement the modular multiplier into the Virtex4 FPGA devices and we show that our technique has better performance when compared with other schemes. To implement large bit-length multiplications we used a novel partitioning and pipeline folding scheme to fit at least 512-bit modular multiplications on a single FPGA.
Osama Daifallah Al-Khaleel, Christos A. Papachristou, Francis Wolff, Kiamal Z. Pekmestzi
ICCD4
2006 Segmentation based design of serial parallel multipliers
abstract
In this paper, a novel architecture for the implementation of serial parallel multipliers (SPM) is proposed. The proposed multiplier is based on a segmentation technique of a simple SPM to blocks of equal bit length. This multiplier achieves higher throughput because it requires small number of zeros to start a new multiplication cycle at a moderate hardware expense and achieves significant hardware reduction compared to the double precision SPM. The proposed technique permits the optimization of the area time product
Paul Bougas, Andreas Tsirikos, Kostas Anagnostopoulos, Isidoros Sideris, Kiamal Z. Pekmestzi
ISCAS5
2005 Long Number Bit-Serial Squarers
abstract
New bit serial squarers for long numbers in LSB first form, are presented in this paper. The first presented scheme is a 50% operational efficient squarer than has the half number of cells compared to the traditional squarers. The second scheme is a 100% operational efficient squarer. In this scheme, the number of the cells remain unchanged compared to other proposed schemes but the number of the required registers is reduced significantly. Both schemes are presented in non-systolic and systolic form and are compared against other squarers presented in the bibliography from the aspect of hardware complexity.
E. Chaniotakis, Paraskevas Kalivas, Kiamal Z. Pekmestzi
IEEE Symposium on Computer Arithmetic3
2001 On the Hardware Implementation of the 3GPP Confidentiality and Integrity Algorithms
Kostas Marinis, Nikolaos Moschopoulos, Fotis Karoubalis, Kiamal Z. Pekmestzi
ISC4
2001 A bit-interleaved systolic architecture for a high-speed RSA system
Kiamal Z. Pekmestzi, Nikolaos Moschopoulos
Integr.1