VLDB 2026 Research / reviewers in the wild / expert
Ghassem Jaberipur
dblp:09/2313
· DBLP profile ↗
28ranked-venue papers
10as first author
7since 2021 · last 2025
0000-0001-8458-7627ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 6 first-author · 4 since 2021Theory of computation · 8 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Poster: Integration of Wearable and Affective Computing via Abstraction and Decision Fusion ArchitectureabstractThis paper introduces an efficient emotion detection method to integrate wearable and affective computing paradigms. Our research contributes to advancing emotion detection technologies, offering potential applications in diverse domains such as healthcare, human-computer interaction, and personalized computing experiences. Our approach addresses the increasing need for real-time emotion recognition while minimizing computational demands. By leveraging low-computation techniques, we propose a novel framework that achieves high accuracy in emotion detection. Besides, advanced data abstraction methods are developed to reduce data workload keeping detection performance. Experimental results demonstrate a notable accuracy rate of $89 .77$%, affirming the efficacy of our proposed method. Mohammadreza Najafi, Mohammad K. Fallah, Saeid Gorgin 0001, Ghassem Jaberipur, Jeong-A Lee |
WoWMoM | 4 |
| 2025 | Balanced Modular Addition for the Moduli Set $ \{2^{q},2^{q}\mp 1,2^{2q}+1\}${2q,2q∓1,22q+1} via Moduli-($ 2^{q}\mp \sqrt{-1}$2q∓-1) AddersabstractModuli-set$ \mathbf{\tau}=\{2^{\boldsymbol{q}},2^{\boldsymbol{q}}\pm 1\}$is often the base of choice for realization of digital computations via residue number systems. The optimum arithmetic performance in parallel residue channels, is generally achieved via equal bit-width residues (e.g.,$ \boldsymbol{q}~ \mathbf{i}\mathbf{n}~ \mathbf{\tau}$) that usually leads to equal computation speed within all the residue channels. However, the commonly difficult and costly task of reverse conversion (RC) is often eased in the existence of conjugate moduli. For example,$ 2^{\boldsymbol{q}}\mp 1\in \mathbf{\tau}$, lead to the efficient modulo-($ 2^{2\boldsymbol{q}}-1$) addition, as the bulk of$ \mathbf{\tau}$-RC, via the New-CRT reverse conversion method. Nevertheless, for additional dynamic range,$ \mathbf{\tau}$is augmented with other moduli. In particular,$ \mathbf{\phi}=\mathbf{\tau}\cup \{2^{2\boldsymbol{q}}+1\}$, leads to efficient RC, where the added modulo is conjugate with the product$ 2^{2\boldsymbol{q}}-1$of$ 2^{\boldsymbol{q}}\mp 1\in \mathbf{\tau}$. Therefore, the final step of$ \mathbf{\phi}$-RC would be fast and low cost/power modulo-($ 2^{4\boldsymbol{q}}-1$) addition. However, the$ 2\boldsymbol{q}$-bit channel-width jeopardizes the existing delay-balance in$ \mathbf{\tau}$. As a remedial solution, given that$ 2^{2\boldsymbol{q}}+1=\left(2^{\boldsymbol{q}}-\boldsymbol{j}\right)\left(2^{\boldsymbol{q}}+\boldsymbol{j}\right)$, with$ \boldsymbol{j}=\sqrt{-1}$, we design and implement modulo-($ 2^{2\boldsymbol{q}}+1$) adders via two parallel$ \boldsymbol{q}$-bit moduli-($ 2^{\boldsymbol{q}}\mp \boldsymbol{j}$) adders. The analytical and synthesis based evaluations of the proposed modulo-($ 2^{\boldsymbol{q}}\mp \boldsymbol{j}$) adders show that the delay-balance of$ \mathbf{\tau}$is preserved with no cost overhead vs.$ \mathbf{\phi}$. In particular, the binary-to-complex and complex-to-binary convertors are merely cost-free and immediate. Ghassem Jaberipur, Elham Rahman, Jeong-A Lee |
IEEE Trans. Computers | 1 |
| 2025 | Design of energy-efficient and high-speed hybrid decimal adder
Negin Mashayekhi, Ghassem Jaberipur, Mohammad Reza Reshadinezhad, Shekoofeh Moghimi |
J. Supercomput. | 2 |
| 2024 | Montgomery Modular Multiplication via Single-Base Residue Number Systems
Zabihollah Ahmadpour, Ghassem Jaberipur, Jeong-A Lee |
ARITH | 2 |
| 2023 | Modulo-(2q - 3) Multiplication with Fully Modular Partial Product Generation and ReductionabstractGiven the residue number systems that contain moduli of the form 2q± 1 and 2q± 3, it is desirable to employ delay-balanced adders and multipliers, in order to synchronize the operation of parallel residue channels. The required modulo-(2q± 3) adders, with compatible speed with modulo- (2q± 1) adders, already exist with parallel prefix architectures. However, the previously reported modulo-(2q± 3) multipliers, in one way or another, produce the non-modular products of the residues at the outset and work towards yielding the final modular product. This seems to be the main source of incompatible performance with the existing modulo- (2q± 1) fully modular multipliers. Therefore, as the first endeavor, we were motivated to design and implement efficient modulo-(2q− 3) multipliers with fully modular partial product generation and reduction that are more compatible with their modulo- (2q− 1) counterparts. However, unlike the case of modulo 2q− 1, it turns out that the straightforward modulo-(2q− 3) partial product reduction (e.g., via Wallace-tree reduction with greedy use of full adders and half adders) falls into an infinite loop of reduction stages. Therefore, we undertake a modified reduction algorithm that requires at most two reduction levels more than that of the modulo- (2q− 1) case to converge. To ensure the correct operation of the algorithm and ease the design process, an in-house software program produces the exact composition of reduction cells in each level of partial product reduction. Analytical and synthesis-based evaluations of the proposed design, and the previous ones, exhibit better figures of merit, as regards the delay (≥ 24%), area-delay (≥ 6%) and energy (≥ 10%) measures. Ghassem Jaberipur, Saeid Gorgin 0001, Navid Ahamadian, Jeong-A Lee |
ARITH | 1 |
| 2022 | Up to $8k$8k-bit Modular Montgomery Multiplication in Residue Number Systems With Fast 16-bit Residue ChannelsabstractHardware realization of public-key cryptosystems often entails Montgomery modular multiplication (MMM), which is more efficient in residue number systems (RNS). A large pool of co-prime moduli allows for higher number of dynamically changeable moduli-set pairs for the required base extension, leading to ultra-wide key-lengths to accommodate the indispensable resistance to differential power-analysis (DPA) attacks. The moduli are often of the form 2^r-, where r denotes the width of residue channels. In a previous relevant RNS MMM design, with r=64, probability of a successful DPA attack is less than 2^(-66), where efficient arithmetic is obtained only for a limited set of moduli that are insufficient for key-lengths over 1024 bits. Here we propose a free- RNS MMM scheme, for up-to 8192-bit key-lengths and fast 16-bit residue channels, based on the proposed -independent modulo-(2^r-) adders and multipliers. Moreover, we propose an especial method for moduli selection that is required for base extension, leading to the same aforementioned DPA-resistance measure and much lower measures for key-lengths over 1024. The implementation results show 82%,69%,44% less RSA delay, for key-lengths 512,1024,2048, respectively of the home designs versus the 512-bit main reference design, and more than 5%,100% for 4096,8192 key-lengths, respectively, all per 512-bit encrypted messages. Zabihollah Ahmadpour, Ghassem Jaberipur |
IEEE Trans. Computers | 2 |
| 2022 | Impact of Radix-10 Redundant Digit Set [-6, 9] on Basic Decimal Arithmetic OperationsabstractRedundant-digit decimal computer arithmetic, with the principal property of carry-free addition, has been the subject of several studies, as the relevant literature contains a wide variety of decimal digit sets and their binary representations. For example, symmetric decimal signed digit sets$[-\alpha, \alpha]$, for$\alpha \in \{5,6,7,8,9 \}$, and the asymmetric ones, such as [−8, 9], [−9, 7], and [0, 15], have been the basis of variant hardware architectures for decimal arithmetic operations. However, digit sets with the minimal 4-bit representations show better figures of merit. In this work, we present a new decimal digit set [−6, 9], called diminished-6 overloaded decimal digit set (DODDS). Its special 4-bit representation contains two posibits (weighted$2^{3}$and$2^{0}$) and two negabits (weighted$2^{2}$and$2^{1}$). Design and implementation of the corresponding carry-free decimal adders and subtractors are thoroughly discussed. Furthermore, to cover the requirements of all possible DODDS applications in decimal multipliers, dividers, and square rooters, we provide adder/subtractor designs for a variety of input combinations of DODDS and binary coded decimal (BCD) operands, all with DODDS output (e.g., DODDS + BCD = DODDS, which is useful in BCD partial product reduction). Analytical and synthesis-based evaluations and comparisons with similar previous works show the notable merits of DODDS over the other redundant decimal digit sets. Ghassem Jaberipur, Farzad Ghazanfari |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | Complex exponential functions: A high-precision hardware realization
Adel Hosseiny, Ghassem Jaberipur |
Integr. | 2 |
| 2019 | Modulo-(2^n+3) Parallel Prefix Addition via Diminished-3 Representation of ResiduesabstractDiminished-1 (D1) representation of modulo-(2n+ 1) residues in [1, 2n] uses the n-bit codes [0, 2n- 1] and maintains a zero-indicator bit. Such D1 encoding has led to efficient parallel prefix modulo-(2n+ 1) adders that perform as fast as the companion modulo-(2n- 1) and -2nadders with (3 + 2 log n)△ delay, where △ denotes the delay of a simple 2-input gate. Also similar, but slower (i.e., with one △ more delay) and slightly more complex, parallel prefix architectures have been offered for modulo-(2n- 3) adders. On the other hand, reverse conversion schemes for 4and 5-moduli sets that include conjugate moduli pairs 2n± 1 and 2n± 3 are already available, while we have not encountered any efficient modulo(2n+ 3) adder. Therefore, in this paper, we offer the diminished-3 (D3) representation of modulo-(2n+ 3) residues that maps the residue interval [3, 2n+ 2] to [0, 2n- 1] and maintains a 2-bit {0,1, 2}-indicator. The corresponding parallel prefix adder, which performs as fast as the fastest previous modulo-(2n- 3) adder is designed, where a 3-way compound architecture is devised as the bulk of modular addition that yields sum, sum+1, and sum+2. The proposed architecture is fully synthesized via Synopsis Design Compiler and tested for correctness, and its figures of merit compared with modulo(2n- 3) and -(2n+ 1) adders. Ghassem Jaberipur, Sahar Moradi Cherati |
ARITH | 1 |
| 2018 | Extended Redundant-Digit Instruction Set for Energy-Efficient ProcessorsabstractThe impact of extending the instruction set architecture (ISA) of a conventional binary processor by a set of redundant-digit arithmetic instructions is studied. Selected binary arithmetic instructions within a given code sequence are replaced with appropriate redundant-digit ones. The selection criteria is so enforced to lead to overall reduction of execution energy and energy-delay product (EDP). A special branch and bound algorithm is devised to modify the dataflow graph (DFG) to a new one that takes advantage of the extended redundant-digit instruction set. The DFG is obtained, via an in-house tool, from the intermediate code representation that is normally produced by the utilized compiler. The required redundant-digit arithmetic operations (including a multiplier, a multiply accumulator, and three- to four-operand redundant-digit adders specially designed for this work) have been synthesized on 45nm NanGate technology by a Synopsys Design Compiler. To evaluate the impact of the proposed ISA augmentation on actual code execution, the simulation and evaluation platform of our choice is an MIPS processor whose ISA is extended by the proposed redundant-digit instructions. Several digital signal processing benchmarks are utilized as the source of the baseline MIPS codes, which are converted (via the aforementioned algorithm) to the equivalent mixed binary/redundant-digit codes. Our experiments, as such, show up to 26% energy and 44% EDP savings. Saba Amanollahi, Ghassem Jaberipur |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Adapting Computer Arithmetic Structures to Sustainable Supercomputing in Low-Power, Majority-Logic NanotechnologiesabstractPetascale supercomputers are already pushing power boundaries that can be supplied or dissipated cost-effectively; greater challenges await us in the era of exascale machines. We are thus motivated to study methods of reducing the energy cost of arithmetic operations, which can be substantial in numerically intensive applications. Additionally, being both a widely-used operation in itself and an important building block for synthesizing other arithmetic operations, has received much attention in this regard. Circuit and energy costs of fast adders are dominated by their fast carry networks. The availability of simple and energy-efficient majority function in certain emerging nanotechnologies (such as quantum-dot cellular automata, single-electron tunneling, tunneling phase logic, magnetic tunnel junction, nanoscale bar magnets, and memristors) has motivated our work to reformulate the carry recurrence in terms of fully-utilized majority elements, with all three inputs usefully employed. We compare our novel designs and resulting circuits to prior proposals based on 3-input majority elements in quantum-dot cellular automata, demonstrating advantages in both speed and circuit complexity. We also show that the performance and cost advantages carry over to at least one other emerging, energy-efficient technology, single-electron tunneling, raising hopes for achieving similar benefits with other technologies, which we review very briefly. Ghassem Jaberipur, Behrooz Parhami, Dariush Abedi |
IEEE Trans. Sustain. Comput. | 1 |
| 2017 | Energy-Efficient VLSI Realization of Binary64 Division With Redundant Number SystemsabstractVLSI realizations of digit-recurrence binary division usually use redundant representation of partial remainders and quotient digits. The former allows for fast carry-free computation of the next partial remainder, and the latter leads to less number of the required divisor multiples. In studying the previous relevant works, we have noted that the binary carry-save (CS) number system is prevalent in the representation of partial remainders, and redundant high radix representation of quotient digits is popular in order to reduce the cycle count. In this paper, we explore a design space containing four division architectures. These are based on binary CS or radix-16 signed digit (SD) representations of partial remainders. On the other hand, they use full or partial precomputation of divisor multiples. The latter uses smaller multiplexer at the cost two extra adders, where one of the operands is constant within all cycles. The quotient digits are represented by radix-16 [-9, 9] SDs. Our synthesis-based evaluation of VLSI realizations of the best previous relevant work and the four proposed designs show reduced power and energy figures in the proposed designs at the cost of more silicon area and delay measures. However, our energy-delay product is 26%-35% less than that of the reference work. Saba Amanollahi, Ghassem Jaberipur |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Sign-Magnitude Encoding for Efficient VLSI Realization of Decimal MultiplicationabstractDecimal X × Y multiplication is a complex operation, where intermediate partial products (IPPs) are commonly selected from a set of precomputed radix-10 X multiples. Some works require only [0, 5] × X via recoding digits of Y to one-hot representation of signed digits in [-5,5]. This reduces the selection logic at the cost of one extra IPP. Two's complement signed-digit (TCSD) encoding is often used to represent IPPs, where dynamic negation (via one xor per bit of X multiples) is required for the recoded digits of Y in [-5, -1]. In this paper, despite generation of 17 IPPs, for 16-digit operands, we manage to start the partial product reduction (PPR) with 16 IPPs that enhance the VLSI regularity. Moreover, we save 75% of negating xors via representing precomputed multiples by sign-magnitude signed-digit (SMSD) encoding. For the first-level PPR, we devise an efficient adder, with two SMSD input numbers, whose sum is represented with TCSD encoding. Thereafter, multilevel TCSD 2:1 reduction leads to two TCSD accumulated partial products, which collectively undergo a special early initiated conversion scheme to get at the final binary-coded decimal product. As such, a VLSI implementation of 16 × 16-digit parallel decimal multiplier is synthesized, where evaluations show some performance improvement over previous relevant designs. Saeid Gorgin 0001, Ghassem Jaberipur |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | A Formulation of Fast Carry Chains Suitable for Efficient Implementation with Majority ElementsabstractCarry computation is a most important notion in computer arithmetic, because it dictates the speed of addition, which is in turn vital to high-speed computation, both as a directly used primitive and as a building block for synthesizing other operations. The theory of fast addition is well-established, but from time to time, changes in technology necessitate a reassessment of strategies for carry network implementation, even though the logical functions to be realized remain the same. We study the implications of the availability of simple, fast, and power-efficient majority gates (in technologies such as quantum-dot cellular automata, single-electron tunneling, tunneling phase logic, magnetic tunnel junction, and nanoscale bar magnets) to the design of carry networks, offering a reformulation of the carry recurrence that allows for building carry networks exclusively out of fully utilized majority elements. We compare our novel implementations based on 3-input majority elements to prior proposals based on these elements, demonstrating advantages in both speed and circuit complexity. Ghassem Jaberipur, Behrooz Parhami, Dariush Abedi |
ARITH | 1 |
| 2016 | Fast low energy RNS comparators for 4-moduli sets {2n±1, 2n, m} with m∈{2n+1±1, 2n-1-1}
Zeinab Torabi, Ghassem Jaberipur |
Integr. | 2 |
| 2016 | Low-Power/Cost RNS Comparison via Partitioning the Dynamic RangeabstractResidue number systems (RNSs) are the main choice in many comparisonand division-free applications (e.g., digital signal processing). However, the development of efficient RNS comparators can widen the spectrum of RNS applications. Such comparators can replace the straightforward, but slow and costly, practice of converting the comparison operands to binary, as inputs to a wide word binary comparator. This has motivated some researchers to design shortcut RNS comparison methods that obviate the need for full reverse conversions. However, the few actual realizations that we have encountered are based on moduli set τ = {2n- 1, 2n, 2n+ 1}. In this paper, after brief review and performance evaluation of the previous methods, we present a new τ-comparator with considerably reduced cost and power dissipation, with no delay penalty. The underlying comparison algorithm is based on ordering the dynamic range into consecutive partitions, and locating the partitions that own the corresponding comparison operands. The required circuitry includes two n-bit adders, which are replaced by one compound parallel prefix architecture, in order to save area and power. Postlayout performance evaluations, of the proposed work and the best previous one, show small latency improvement, 17%(46%) reduction in area consumption, 30%(41%) in power dissipation, and 31%(47%) in power-delay product, for n = 8(22). Zeinab Torabi, Ghassem Jaberipur |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Modulo-(2n - 2q - 1) Parallel Prefix Addition via Excess-Modulo Encoding of ResiduesabstractThe residue number system t = {2n -- 1, 2n, 2n + 1} has been extensively studied towards perfection in realization of efficient parallel prefix modular adders, with (3 + 2log n ?G) latency. Many applications, such as digital signal processing require fast modular operations. However, relying only on t limits the magnitude of n, and accordingly the dynamic range. Therefore, additional mutually prime moduli are required to accommodate for wider dynamic range. On the other hand, speed of modular arithmetic operations for the additional moduli should be as close as possible to those in t. This could be best met by the moduli of the form 2n -- (2q + 1), with 1 ≤ q ≤ n -- 2, such as 2n -- 3, 2n -- 5. However, the fastest parallel prefix realization of modulo-(2n -- 2q -- 1) adders that we have encountered in the relevant literature, claims (7 + 2 log n)?G latency. Motivated by the need to reduce the latter, we propose new designs of such adders with (5 + 2 log n)?G latency without any penalty in area consumption or power dissipation. The proposed modular addition algorithm entails supplementary representation of residues in [0,2q], as [2n -- (2q + 1), 2n -- 1]. This leads to additional performance efficiency similar to the effect of double zero representation in modulo-(2n -- 1) adders. The aforementioned analytically evaluated speed gain and improvements in other figures of merit are also supported via circuit simulation and synthesis. Hamed Fatemi Langroudi, Ghassem Jaberipur |
ARITH | 2 |
| 2015 | A New Residue Number System with 5-Moduli Set: {22q, 2q±3, 2q±1}abstractResidue number system (RNS) parameterized moduli sets almost always contain a power-of-two modulo (e.g. 2iq), where the corresponding computation channel and residue generator are the most efficient when compared with other non-power-of-two moduli (e.g. 2jq ± δ). Furthermore, inclusion of a power-of-two modulo leads to efficient use of the new Chinese remainder theorem for reverse conversion. However, few reverse conversion schemes and modulo-(2q±3) arithmetic operators have been recently reported for ℱ={2q ± 3, 2q ± 1}, where it appears that devising similar reverse conversion schemes for the more useful 5-moduli set ℱ⋃ {2iq} is too challenging that no such moduli set has been yet proposed. Therefore, we propose the arithmetically balanced moduli set 𝒫 = {22q, 2q ± 3, 2q ± 1} and study the corresponding problems of binary to RNS conversion and the reverse, where adder-only solutions (with neither costly read-only-memories nor multipliers) are presented. We state and prove some lemmas and theorems to obtain at the required infinite geometric series to express the multiplicative inverses as power-of-two polynomials. Different groupings of moduli are investigated and more feasible cases are set aside for realization of four reverse converters showing cost/speed trade-off that are evaluated analytically and by synthesis. Both forward and reverse converters are designed and implemented via multi-operand addition realized via fast parallel architectures. HamidReza Ahmadifar, Ghassem Jaberipur |
Comput. J. | 2 |
| 2015 | Comment on "High Speed Parallel Decimal Multiplication With Redundant Internal Encodings"abstractHanpropose a new method for parallel decimal multiplication with redundant partial products. They compare the performance of their multiplier with some previous relevant works, based on analytical and synthesis results. We have noted that the claimed critical delay path in (IEEE Trans. Computers, vol. 62, no. 5, pp. 956–968, May 2013) is faster than the actual critical delay path. Therefore, comparison results seem to be deceptive. For example, our accurate analytical evaluation devaluated the claimed speed advantage over the multiplier of (Microelectronics J., vol. 40, no. 10, pp. 1471–1481, Oct. 2009). Furthermore, we synthesized both multipliers, to show synthesis results confirm those of analytical evaluation. Saeid Gorgin 0001, Ghassem Jaberipur |
IEEE Trans. Computers | 2 |
| 2014 | Low area/power decimal addition with carry-select correction and carry-select sum-digits
Morteza Dorrigiv, Ghassem Jaberipur |
Integr. | 2 |
| 2011 | A Family of High Radix Signed Digit AddersabstractSigned digit (SD) number systems allow for high performance carry-free adders. Maximally redundant SD (MRSD) alternatives provide maximal encoding efficiency among Radix-2hSD number systems, whereby value of h tunes the area-time trade-off. Straightforward implementation of the conventional carry-free addition algorithm requires three O(log h) addition-like operations in sequence. However, there are several MRSD implementations with only one such operation. Some of them are delay optimized, but suffer from extensive hardware redundancy, while some other equally fast adders show less power/area consumption. A careful study of the latter cases hints on variety of improvement options, based on which and a new transfer computation technique, we develop a family of faster MRSD adders that consume less power/area than all the previous relevant works. They also fit efficiently within the redundant digit floating point addition scheme. However, similar to their relevant ancestor designs, suffer from an inherent property of MRSD adders, i.e., difficulty of handling hidden leading zero-digits. To remedy this problem, we use less redundant SD representations, where our transfer extraction method applies efficiently and leads to far less complex leading zero-digit detection. All the presented designs are supported by exhaustive correctness tests and performance evaluation via 0.13 micrometer CMOS technology synthesis. Saeid Gorgin 0001, Ghassem Jaberipur |
IEEE Symposium on Computer Arithmetic | 2 |
| 2011 | Decimal CORDIC Rotation based on Selection by Rounding: Algorithm and ArchitectureabstractHardware implementation of decimal floating-point arithmetic is a topic of great interest among the researchers in computer arithmetic and also the digital processor industry. Software packages for decimal arithmetic are actually being challenged by decimal hardware units. This spreading trend seems to include hardware implementation of elementary functions. The (Coordinate Rotation Digital Computer) CORDIC algorithm, due to its simplicity, is one of the most efficient methods for computing elementary functions. In this work, we develop a decimal CORDIC scheme with almost half number of equally long cycles with respect to the best previous design. This is achieved via retiming of the conventional CORDIC architecture and selection of the microrotation factors by rounding. However, the proposed design does not lead to a predetermined constant scaling factor. The solution that we use is to iteratively compute the logarithm of the scaling factor followed by a decimal exponentiation. The same CORDIC hardware is reused for performing the latter. The proposed CORDIC method requires 2n + 3 cycles for n-digit decimal operands vs. 4n cycles of the previous methods. Evaluations with 16-digit operands based on logical effort analysis conclude that the proposed architecture shows 82% speed advantage, at the cost of 60% more area and 2.5 KB more ROM. Amir Kaivani, Ghassem Jaberipur |
Comput. J. | 2 |
| 2010 | Fully redundant decimal addition and subtraction using stored-unibit encoding
Amir Kaivani, Ghassem Jaberipur |
Integr. | 2 |
| 2010 | Redundant-Digit Floating-Point Addition Scheme Based on a Stored Rounding ValueabstractDue to the widespread use and inherent complexity of floating-point addition, much effort has been devoted to its speedup via algorithmic and circuit techniques. We propose a new redundant-digit representation for floating-point numbers that leads to computation speedup in two ways: (1) Reducing the per-operation latency when multiple floating-point additions are performed before result conversion to nonredundant format and (2) Removing the addition associated with rounding. While the first of these advantages is offered by other redundant representations, the second one is unique to our approach, which replaces the power- and area-intensive rounding addition by low-latency insertion of a rounding two-valued digit, or twit, in a position normally assigned to a redundant twit within the redundant-digit format. Instead of conventional sign-magnitude representation, we use a sign-embedded encoding that leads to lower hardware redundancy, and thus, reduced power dissipation. While our intermediate redundant representations remain incompatible with the IEEE 754-2008 standard, many application-specific systems, such as those in DSP and graphics domains, can benefit from our designs. Description of our radix-16 redundant representation and its addition algorithm is followed by the architecture of a floating-point adder based on this representation. Detailed circuit designs are provided for many of the adder's critical subfunctions. Simulation and synthesis based on a 0.13 ¿m CMOS standard process show a latency reduction of 15 percent or better, and both area and power savings of around 58 percent, compared with the best designs reported in the literature. Ghassem Jaberipur, Behrooz Parhami, Saeid Gorgin 0001 |
IEEE Trans. Computers | 1 |
| 2009 | Fully Redundant Decimal ArithmeticabstractHardware implementation of all the basic radix-10 arithmetic operations is evolving as a new trend in the design and implementation of general purpose digital processors. Redundant representation of partial products and remainders is common in the multiplication and division hardware algorithms, respectively. Carry-free implementation of the more frequent add/subtract operations, with the byproduct of enhancing the speed of multiplication and division, is possible with redundant number representation. However, conversion of redundant results to conventional representations entails slow carry propagation that can be avoided if the results are kept in redundant format for later use as operands of other arithmetic operations. Given that redundant decimal representations, contrary to redundant binary, do not necessarily require extra storage, we are motivated to develop a framework for fully redundant decimal arithmetic, where all operands and results belong to the same redundant decimal number system and can be stored and later used as operands of further decimal operations. In this paper, we present a new faster decimal signed digit add/sub unit and show how it can be efficiently used in the design of decimal multipliers and dividers, where all operands and results are represented with the same redundant digit set [-7, 7]. Saeid Gorgin 0001, Ghassem Jaberipur |
IEEE Symposium on Computer Arithmetic | 2 |
| 2009 | Unified Approach to the Design of Modulo-(2n +/- 1) Adders Based on Signed-LSB Representation of ResiduesabstractModuli of the form 2nplusmn 1, which greatly simplify certain arithmetic operations in residue number systems (RNS), have been of longstanding interest. A steady stream of designs for modulo (2nplusmn 1) adders has rendered the latency of such adders quite competitive with ordinary adders. The next logical step is to approach the problem in a unified and systematic manner that does not require each design to be taken up from scratch and to undergo the error prone and labor intensive optimization for high speed and low power dissipation. Accordingly, we devise a new redundant representation of mod (2nplusmn 1) residues that allows ordinary fast adders and a small amount of peripheral logic to be used for mod (2nplusmn 1) addition. Advantages of the building block approach include shorter design time, easier exploration of the design space (area, speed, power tradeoffs), and greater confidence in the correctness of the resulting circuits. Advantages of the unified design include the possibility of fault tolerant and gracefully degrading RNS circuit realizations with fairly low hardware redundancy. Ghassem Jaberipur, Behrooz Parhami |
IEEE Symposium on Computer Arithmetic | 1 |
| 2009 | Improving the Speed of Parallel Decimal MultiplicationabstractHardware support for decimal computer arithmetic is regaining popularity. One reason is the recent growth of decimal computations in commercial, scientific, financial, and Internet-based computer applications. Newly commercialized decimal arithmetic hardware units use radix-10 sequential multipliers that are rather slow for multiplication-intensive applications. Therefore, the future relevant processors are likely to host fast parallel decimal multiplication circuits. The corresponding hardware algorithms are normally composed of three steps: partial product generation (PPG), partial product reduction (PPR), and final carry-propagating addition. The state of the art is represented by two recent full solutions with alternative designs for all the three aforementioned steps. In addition, PPR by itself has been the focus of other recent studies. In this paper, we examine both of the full solutions and the impact of a PPR-only design on the appropriate one. In order to improve the speed of parallel decimal multiplication, we present a new PPG method, fine-tune the PPR method of one of the full solutions and the final addition scheme of the other; thus, assembling a new full solution. Logical Effort analysis and 0.13 mum synthesis show at least 13 percent speed advantage, but at a cost of at most 36 percent additional area consumption. Ghassem Jaberipur, Amir Kaivani |
IEEE Trans. Computers | 1 |
| 2008 | Constant-time addition with hybrid-redundant numbers: Theory and implementations
Ghassem Jaberipur, Behrooz Parhami |
Integr. | 1 |