EDBT 2026 Demo / reviewers in the wild / expert
Tomás Lang
dblp:55/6764
· DBLP profile ↗
98ranked-venue papers
25as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 73 · 19 first-authorTheory of computation · 20 · 5 first-authorSoftware engineering, systems software and programming languages · 6Databases, data management, data science and information retrieval · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
41 papers |
Processor architecture and microarchitecture · 51% Integrated circuit design · 36% Memory systems · 7% | |
| Computer graphics and multimedia
1 paper |
Rendering · 100% |
Topics — the 30 heaviest of 73, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
computer arithmetic |
0.4 | 14 | 2012 | Radix-2 Multioperand and Multiformat Streaming Online Addition · IEEE Trans. Computers 2012 Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 Reciprocation, Square Root, Inverse Square Root, and Some Elementary Functions Using Small Multipliers · IEEE Trans. Computers 2000 |
Integrated circuit design
digital arithmetic circuits |
0.2 | 7 | 2007 | A Radix-10 Digit-Recurrence Division Unit: Algorithm and Architecture · IEEE Trans. Computers 2007 Floating-Point Multiply-Add-Fused with Reduced Latency · IEEE Trans. Computers 2004 Radix-4 Reciprocal Square-Root and Its Combination with Division and Square Root · IEEE Trans. Computers 2003 |
Processor architecture and microarchitecture › computer arithmetic
multioperand addition |
0.1 | 1 | 2012 | Radix-2 Multioperand and Multiformat Streaming Online Addition · IEEE Trans. Computers 2012 |
Integrated circuit design › digital arithmetic circuits
online arithmetic |
0.1 | 1 | 2012 | Radix-2 Multioperand and Multiformat Streaming Online Addition · IEEE Trans. Computers 2012 |
Processor architecture and microarchitecture › division algorithm
digit-recurrence divider |
0.1 | 2 | 2007 | A Radix-10 Digit-Recurrence Division Unit: Algorithm and Architecture · IEEE Trans. Computers 2007 Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture › computer arithmetic
elementary function evaluation |
0.1 | 2 | 2006 | Double-Residue Modular Range Reduction for Floating-Point Hardware Implementations · IEEE Trans. Computers 2006 Reciprocation, Square Root, Inverse Square Root, and Some Elementary Functions Using Small Multipliers · IEEE Trans. Computers 2000 |
Integrated circuit design › digital arithmetic circuits
CORDIC |
0.1 | 5 | 2000 | Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/Vectoring · IEEE Trans. Computers 2000 CORDIC Vectoring with Arbitrary Target Value · IEEE Trans. Computers 1998 Error Analysis and Reduction for Angle Calculation Using the CORDIC Algorithm · IEEE Trans. Computers 1997 |
Processor architecture and microarchitecture › computer arithmetic
floating-point arithmetic |
0.1 | 2 | 2006 | Double-Residue Modular Range Reduction for Floating-Point Hardware Implementations · IEEE Trans. Computers 2006 Boosting Very-High Radix Division with Prescaling and Selection by Rounding · IEEE Trans. Computers 2001 |
Integrated circuit design › digital circuit design › arithmetic circuit design
floating-point unit |
0.1 | 2 | 2004 | Floating-Point Multiply-Add-Fused with Reduced Latency · IEEE Trans. Computers 2004 Leading-One Prediction with Concurrent Position Correction · IEEE Trans. Computers 1999 |
Integrated circuit design › digital circuit design
arithmetic circuit design |
0.1 | 2 | 2001 | Boosting Very-High Radix Division with Prescaling and Selection by Rounding · IEEE Trans. Computers 2001 Low-Power Divider · IEEE Trans. Computers 1999 |
Processor architecture and microarchitecture
division algorithm |
0.1 | 1 | 2005 | Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 |
GPUs and heterogeneous computing
graphics accelerator |
0.1 | 1 | 2005 | High-Throughput CORDIC-Based Geometry Operations for 3D Computer Graphics · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture › division algorithm
quotient digit selection |
0.1 | 1 | 2005 | Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture › computer arithmetic
digit-recurrence algorithm |
0.1 | 3 | 2000 | Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/Vectoring · IEEE Trans. Computers 2000 Computation of sqrt(x/d) in a Very High Radix Combined Division/Square-Root Unit with Scaling · IEEE Trans. Computers 1998 On-the-Fly Rounding · IEEE Trans. Computers 1992 |
Integrated circuit design
digital circuit design |
0.0 | 4 | 2005 | Error Analysis and Reduction for Angle Calculation Using the CORDIC Algorithm · IEEE Trans. Computers 1997 Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/Vectoring · IEEE Trans. Computers 2000 |
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
fused multiply-add |
0.0 | 1 | 2004 | Floating-Point Multiply-Add-Fused with Reduced Latency · IEEE Trans. Computers 2004 |
Integrated circuit design › digital circuit design › arithmetic circuit design
division and square root unit |
0.0 | 2 | 1999 | Very High Radix Square Root with Prescaling and Rounding and a Combined Division/Square Root Unit · IEEE Trans. Computers 1999 Computation of sqrt(x/d) in a Very High Radix Combined Division/Square-Root Unit with Scaling · IEEE Trans. Computers 1998 |
Integrated circuit design › digital arithmetic circuits
division and square root |
0.0 | 1 | 2003 | Radix-4 Reciprocal Square-Root and Its Combination with Division and Square Root · IEEE Trans. Computers 2003 |
Processor architecture and microarchitecture › arithmetic unit
arithmetic unit design |
0.0 | 3 | 2005 | Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/Vectoring · IEEE Trans. Computers 2000 Computation of sqrt(x/d) in a Very High Radix Combined Division/Square-Root Unit with Scaling · IEEE Trans. Computers 1998 |
Memory systems
memory access |
0.0 | 2 | 1995 | Conflict-Free Access for Streams in Multimodule Memories · IEEE Trans. Computers 1995 Vector Multiprocessors with Arbitrated Memory Access · ISCA 1995 |
Integrated circuit design › digital arithmetic circuits › division and square root
divider |
0.0 | 1 | 1999 | Low-Power Divider · IEEE Trans. Computers 1999 |
Integrated circuit design › digital arithmetic circuits › floating-point unit design
floating-point adder |
0.0 | 1 | 1999 | Leading-One Prediction with Concurrent Position Correction · IEEE Trans. Computers 1999 |
Energy-efficient computing › low-power design
low-power arithmetic circuit |
0.0 | 1 | 1999 | Low-Power Divider · IEEE Trans. Computers 1999 |
Processor architecture and microarchitecture
arithmetic unit |
0.0 | 1 | 1997 | Error Analysis and Reduction for Angle Calculation Using the CORDIC Algorithm · IEEE Trans. Computers 1997 |
Memory systems
cache |
0.0 | 1 | 1996 | The Difference-bit Cache · ISCA 1996 |
Memory systems › memory access latency
cache access latency |
0.0 | 1 | 1996 | The Difference-bit Cache · ISCA 1996 |
Memory systems › cache
cache organization |
0.0 | 1 | 1996 | The Difference-bit Cache · ISCA 1996 |
Memory systems › cache › cache organization
set-associative cache |
0.0 | 1 | 1996 | The Difference-bit Cache · ISCA 1996 |
Memory systems › memory access patterns
conflict-free access |
0.0 | 1 | 1995 | Conflict-Free Access for Streams in Multimodule Memories · IEEE Trans. Computers 1995 |
Memory systems
memory architecture |
0.0 | 1 | 1995 | Conflict-Free Access for Streams in Multimodule Memories · IEEE Trans. Computers 1995 |
Methods — techniques the papers use, named apart from their topics
signed-digit representation · 0.2two's complement · 0.1carry-save representation · 0.1prescaling · 0.1radix-10 digit recurrence · 0.1speculation · 0.1modular range reduction · 0.1double-residue arithmetic · 0.1selection by rounding · 0.1logical effort timing model · 0.1CORDIC · 0.1constant-factor redundant CORDIC · 0.0hardware tracking of written registers · 0.0activation record placement · 0.0privilege states · 0.0hardware-software co-design · 0.0miss ratio analysis · 0.0analytical modeling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Radix-2 Multioperand and Multiformat Streaming Online AdditionabstractIn this paper, we present multioperand radix-2 online addition using different data representations (signed-digit, two's complement, and carry-save), in particular cases in which operands with different representations are added. We use the previously defined online full adder (olFA) as a component to build different multioperand online architectures. To merge data with different representations, an inner conversion of data is performed, eliminating any conversion stage and penalty time. We propose a technique to build multioperand trees efficiently and give six practical rules to deal with different kinds of data in the same adder. For addition of a stream of data, we determine the minimum number of separation cycles required to isolate two successive computations and propose a novel hardware technique that eliminates completely the separation cycles, resulting in the maximum throughput possible. Julio Villalba, Tomás Lang, Javier Hormigo |
IEEE Trans. Computers | 2 |
| 2009 | Division Unit for Binary Integer DecimalsabstractIn this work, we present a radix-10 division unit that is based on the digit-recurrence algorithm and implements binary encodings (binary integer decimal or BID) for significands. Recent decimal division designs are all based on the binary coded decimal (BCD) encoding. We adapt the radix-10 digit-recurrence algorithm to BID representation and implement the division unit in standard cell technology. The implementation of the proposed BID division unit is compared to that of a BCD based unit implementing the same algorithm. The comparison shows that for normalized operands the BID unit has the same latency as the BCD unit and reduced area, but the normalization is more expensive when implemented in BID. Tomás Lang, Alberto Nannarelli |
ASAP | 1 |
| 2007 | Improving the Throughput of On-line Addition for Data StreamsabstractIn this paper we deal with the throughput of on–line addition for a stream of data. This throughput is directly related to the initiation interval between two successive instances. The on–line delay for the addition of two signed–digit (or carry–save) numbers is two, and N+2 cycles are classically used to compute a new pair of N–digit data (initiation interval: N+2). In this paper we present some techniques to reduce the initiation interval to N (which is the theoretical minimum value) with a very small amount of hardware or N+1 with no hardware cost. For short operands, this might have a significant effect on the throughput. Julio Villalba, Javier Hormigo, Tomás Lang |
ASAP | 3 |
| 2007 | A Radix-10 Digit-Recurrence Division Unit: Algorithm and ArchitectureabstractIn this work, we present a radix-10 division unit that is based on the digit-recurrence algorithm. The previous decimal division designs do not include recent developments in the theory and practice of this type of algorithm, which were developed for radix-2kdividers. In addition to the adaptation of these features, the radix-10 quotient digit is decomposed into a radix-2 digit and a radix-5 digit in such a way that only five and two times the divisor are required in the recurrence. Moreover, the most significant slice of the recurrence, which includes the selection function, is implemented in radix-2, avoiding the additional delay introduced by the radix-10 carry-save additions and allowing the balancing of the paths to reduce the cycle delay. The results of the implementation of the proposed radix-10 division unit show that its latency is close to that of radix-16 division units (comparable dynamic range of significant) and it has a shorter latency than a radix-10 unit based on the Newton-Raphson approximation Tomás Lang, Alberto Nannarelli |
IEEE Trans. Computers | 1 |
| 2006 | Double-Residue Modular Range Reduction for Floating-Point Hardware ImplementationsabstractIn this paper, we present a novel algorithm and the corresponding architecture for performing range reduction, which is a preprocessing task required for the evaluation of some elementary functions such as trigonometric and exponential-based functions. The proposed algorithm introduces a modification to the modular range reduction algorithm which increases the speed of computation and allows us to design an architecture for the floating-point case. The implementation presented admits as an input argument any representable number of the standard single precision IEEE 754 floating-point representation and provides the maximum accuracy to the final result. This supposes a hardware solution to the problem of having an input argument close to a multiple of the constant. A final comparison with other implementations is presented. Julio Villalba, Tomás Lang, Mario A. González |
IEEE Trans. Computers | 2 |
| 2005 | Low Latency Digit-Recurrence Reciprocal and Square-Root Reciprocal Algorithm and ArchitectureabstractThe reciprocal and square-root reciprocal operations are important in several applications. For these operations, we present algorithms that combine a digit-by-digit module and one iteration of a quadratic-convergence approximation. The latter is implemented by a digit-recurrence, which uses the digits produced by the digit-by-digit part. In this way, both parts execute in an overlapped manner, so that the total number of cycles is about half of the number that would be required by the digit-by-digit part alone. Because of the approximation, correct rounding of the result cannot be obtained directly in all cases; we propose a variable-time implementation that produces the correctly rounded result with a small average overhead. Radix-4 implementations are described and have been synthesized. They achieve the same cycle time as the standard digit-by-digit implementation, resulting in a speed-up of about 2 and, because of the approximation part, the area factor is also about 2. We also show a combined implementation for both operations that has essentially the same complexity as that for square-root reciprocal alone. Elisardo Antelo, Tomás Lang, Paolo Montuschi, Alberto Nannarelli |
IEEE Symposium on Computer Arithmetic | 2 |
| 2005 | Floating-Point Fused Multiply-Add: Reduced Latency for Floating-Point AdditionabstractIn this paper we propose an architecture for the computation of the double-precision floating-point multiply-add fused (MAF) operation A+(B/spl times/C) that permits to compute the floating-point addition with lower latency than floating-point multiplication and MAF. While previous MAF architectures compute the three operations with the same latency, the proposed architecture permits to skip the first pipeline stages, those related with the multiplication B/spl times/C, in case of an addition. For instance, for a MAF unit pipelined into three or five stages, the latency of the floating-point addition is reduced to two or three cycles, respectively. To achieve the latency reduction for floating-point addition, the alignment shifter, which in previous organizations is in parallel with the multiplication, is moved so that the multiplication can be bypassed. To avoid that this modification increases the critical path, a double-datapath organization is used, in which the alignment and normalization are in separate paths. Moreover, we use the techniques developed previously of combining the addition and the rounding and of performing the normalization before the addition. Javier D. Bruguera, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 2005 | Digit-Recurrence Dividers with Reduced Logical DepthabstractIn this paper, we propose a class of division algorithms with the aim of reducing the delay of the selection of the quotient digit by introducing more concurrency and flexibility in its computation. From the proposed class of algorithms, we select one that moves part of the selection function out of the critical path, with a corresponding reduction in the critical path compared with existing alternatives: we present the algorithm and describe the architectures for radix 4 and for radix 16. For radix 16, we use the scheme of overlapping two radix-4 stages. In both cases, radix 4 and radix 16, we show that our algorithms allow the design of units with well-balanced critical paths with consequent decreases of the cycle times. Moreover, in the radix-16 case, we include some additional speculation techniques. To estimate the speedup, we used a rough timing model based on logical effort. For both radices, we estimate a speedup of about 25 percent with respect to previous implementations. In the radix-4 case, this is achieved by using roughly the same area, while, in the radix-16 case, the area is increased by about 30 percent. We verified our estimations by performing a synthesis of the radix-4 units. Elisardo Antelo, Tomás Lang, Paolo Montuschi, Alberto Nannarelli |
IEEE Trans. Computers | 2 |
| 2005 | High-Throughput CORDIC-Based Geometry Operations for 3D Computer GraphicsabstractGraphics processors require strong arithmetic support to perform computational kernels over data streams. Because of the current implementation using the basic arithmetic operations, the algorithms are given in algebraic terms. However, since the operations are really of a geometric nature, it seems to us that more flexibility in the implementation is obtained if the description is given in a high-level geometrical form. As a consequence of this line of thought, this paper is an attempt to reconsider some kernels in a graphics processor to obtain implementations that are potentially more scalable than just replicating the modules used in conventional implementations. We present the formulation of representative 3D computer graphics operations in terms of CORDIC-type primitives. Then, we briefly outline a stream processor based on CORDIC-type modules to efficiently implement these graphic operations. We perform a rough comparison with current implementations and conclude that the CORDIC-based alternative might be attractive. Tomás Lang, Elisardo Antelo |
IEEE Trans. Computers | 1 |
| 2004 | Floating-Point Multiply-Add-Fused with Reduced LatencyabstractWe propose architecture for the computation of the double-precision floating-point multiply-add-fused (MAP) operation A + (B /spl times/ C). This architecture is based on the combined addition and rounding (using a dual adder) and in the anticipation of the normalization step before the addition. Because the normalization is performed before the addition, it is not possible to overlap the leading-zero-anticipator with the adder. Consequently, to avoid the increase in delay, we modify the design of the LZA so that the leading bits of its output are produced first and can be used to begin the normalization. Moreover, parts of the addition are also anticipated. We have estimated the delay of the resulting architecture considering the load introduced by long connections, and we estimate a delay reduction of between 15 percent and 20 percent, with respect to previous implementations. Tomás Lang, Javier D. Bruguera |
IEEE Trans. Computers | 1 |
| 2003 | Radix-4 Reciprocal Square-Root and Its Combination with Division and Square RootabstractIn this work, we present a reciprocal square root algorithm by digit recurrence and selection by a staircase function and the radix-4 implementation. As in similar algorithms for division and square root, the results are obtained correctly rounded in a straightforward manner (in contrast to existing methods to compute the reciprocal square root). Although, apparently, a single selection function can only be used for j /spl ges/ 2 (the selection constants are different for j = 0, j = 1, and j /spl ges/ 2), we show that it is possible to use a single selection function for all iterations. We perform a rough comparison with existing methods and we conclude that our implementation is a low hardware complexity solution with moderate latency, especially for exactly rounded results. We also extend the unit to support division and square root with the same selection function and with slight modifications in the initialization of the reciprocal square root unit. Tomás Lang, Elisardo Antelo |
IEEE Trans. Computers | 1 |
| 2002 | Fast Radix-4 Retimed Division with Selection by ComparisonsabstractSince a large portion of the critical path in an implementation of radix-4 division corresponds to the delay of the quotient-digit selection module, it is of interest to reduce this delay. The proposal of this paper extends the approach presented recently of prestoring the selection constants corresponding to the actual value of the divisor and to perform the determination of the quotient digit by carry-free subtraction and sign detection. This extension consists in advancing the subtraction so that it is outside of the critical path. This advancement also provides the possibility of placing the registers so as to minimize the cycle time. We present the method and report results of synthesis using a family of standard cells. We conclude that the extension results in a speedup of 1.35 with respect to the basic implementation and of 1.3 with respect to the previously mentioned approach. We estimate that the areas of all three units are about the same. Elisardo Antelo, Tomás Lang, Paolo Montuschi, Alberto Nannarelli |
ASAP | 2 |
| 2002 | Floating-Point Fused Multiply-Add with Reduced LatencyabstractWe propose an architecture for the computation of the floating-point multiply-add-fused (MAF) operation A+ (B /spl times/ C). This architecture is based on the combined addition and rounding (using a dual adder) and on the anticipation of the normalization step before the addition. Because the normalization is performed before the addition, it is not possible to overlap the leading-zero-anticipator with the adder. Consequently, to avoid the increase in delay we modify the design of the LZA so that the leading bits of its output are produced first and can be used to begin the normalization. Moreover, parts of the addition are also anticipated. We have estimated the delay of the resulting architecture for double-precision format, considering the load introduced by long connections, and estimate a reduction of about 15% to 20% with respect to traditional implementations of the floating-point MAF unit. Tomás Lang, Javier D. Bruguera |
ICCD | 1 |
| 2001 | Using the Reverse-Carry Approach for Double Datapath Floating-Point AdditionabstractThe double-datapath organization of a floating-point adder results in reduced latency. One of the main characteristics of this organization is the combination of addition/subtraction with rounding into a single add/round module, which is implemented as one pipeline stage and might be responsible for the cycle time. We propose the utilization of the most-significant carry detector and the corresponding adder using the reverse-carry approach to reduce the latency of this add/round module. In addition, the particular organization of the reverse-carry adder is used to reduce the contribution on the delay of the row of half adders that is included in the FAR datapath to produce the sum plus two. Estimates for a 64 bit add/round module show a potential reduction of delay of about 15%. Javier D. Bruguera, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 2001 | Correctly Rounded Reciprocal Square-Root by Digit Recurrence and Radix-4 ImplementationabstractWe present a reciprocal square-root algorithm by digit recurrence and selection by a staircase function, and the radix-4 implementation. As similar algorithms for division and square-root, the results are obtained correctly rounded in a straightforward manner (in contrast to existing methods to compute the reciprocal square-root). Although apparently a single selection function can only be used for j/spl ges/2 (the selection constants are different for j=0, j=1 and j/spl ges/2), we show that it is possible to use a single selection function for all iterations. We perform a rough comparison with existing methods and we conclude that our implementation is a low hardware complexity solution with moderate latency, specially for exactly rounded results. Tomás Lang, Elisardo Antelo |
IEEE Symposium on Computer Arithmetic | 1 |
| 2001 | Bounds on Runs of Zeros and Ones for Algebraic FunctionsabstractThis paper presents upper bounds on the number of zeros and of ones after the rounding bit for algebraic functions. These functions include reciprocal, division, square root, and reciprocal square root, which have been considered in previous work. We propose simpler proofs for the previously given bounds and generalize to all algebraic functions. We also determine cases for which the bound is achieved for square root. As is mentioned in the previous work, these bounds are useful for determining the precision required in the computation of approximations in order to be able to perform correct rounding. We consider rounding to nearest, but the results can be easily extended to other rounding modes. Tomás Lang, Jean-Michel Muller |
IEEE Symposium on Computer Arithmetic | 1 |
| 2001 | Boosting Very-High Radix Division with Prescaling and Selection by RoundingabstractAn extension of the very-high radix division with prescaling and selection by rounding is presented. This extension consists of increasing the effective radix of the implementation by obtaining a few additional bits of the quotient per iteration, without increasing the complexity of the unit to obtain the prescaling factor or the delay of an iteration. As a consequence, for some values of the effective radix, it permits an implementation with a smaller area and the same execution time of the original scheme. Details of the algorithm and the implementation are presented. Estimations of the execution time and area are given for 54 bit and 114 bit quotients and compared with those of other division units. Paolo Montuschi, Tomás Lang |
IEEE Trans. Computers | 2 |
| 2001 | Multilevel reverse most-significant carry computationabstractA fast calculation of the most-significant carry in an addition is required in several applications. It has been proposed to calculate this carry by detecting the most-significant carry chain and collecting the carry after this chain. The detection can be implemented by a prefix tree of AND gates and the collecting by a multi-input OR. We propose a multilevel implementation, which allows the overlap of successive levels, thereby reducing the overall delay. For 64-bit operands we estimate a delay reduction of about 15% with respect to the traditional carry-lookahead-based method, with a similar hardware complexity. Javier D. Bruguera, Tomás Lang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Multilevel Reverse-Carry AdderabstractThe multilevel reverse-carry approach has been proposed previously for fast computation of the most-significant carry of an adder. We extend this approach to generate several carries and apply it to the implementation of the complete adder. Specifically, the operands are split into blocks and each block is added to produce the sum and the sum plus one. Concurrently with these additions the multilevel reverse-carry approach is used to generate the input carries of these blocks. Finally, these carries are used to select among the sum and the sum plus one. We have evaluated the resulting architecture for a 64-bit adder, considering the load introduced by long connections, and we estimate a reduction of about 15% in the critical path delay with respect to traditional implementations of prefix-tree based adders. Javier D. Bruguera, Tomás Lang |
ICCD | 2 |
| 2000 | Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/VectoringabstractA very-high radix algorithm and implementation for circular CORDIC is presented. We first present in depth the algorithm for the vectoring mode in which the selection of the digits is performed by rounding of the control variable. To assure convergence with this kind of selection, the operands are prescaled. However, in the CORDIC algorithm, the coordinate x varies during the execution so several scalings might be needed; we show that two scalings are sufficient. Moreover, the compensation of the variable scale factor (including the CORDIC scale factor and the prescaling factors) is done by computing the logarithm of the scale factor and performing the compensation by an exponential. Then, we combine, in a unified unit, the proposed vectoring algorithm and the very-high radix rotation algorithm, which was previously proposed by the authors. We compare with low-radix implementations in terms of latency and hardware complexity. Estimations of the delay for 32-bit precision show a speedup of about two with respect to the radix-4 case with redundant addition. This speedup is obtained at the cost of an increase in the hardware complexity, which is moderate for the pipelined implementation. We also compare at the algorithmic level with other very-high radix proposals, demonstrating the advantages of our algorithms. Elisardo Antelo, Tomás Lang, Javier D. Bruguera |
IEEE Trans. Computers | 2 |
| 2000 | Reciprocation, Square Root, Inverse Square Root, and Some Elementary Functions Using Small MultipliersabstractThis paper deals with the computation of reciprocals, square roots, inverse square roots, and some elementary functions using small tables, small multipliers, and, for some functions, a final "large" (almost full-length) multiplication. We propose a method, based on argument reduction and series expansion, that allows fast evaluation of these functions in high precision. The strength of this method is that the same scheme allows the computation of all these functions. We estimate the delay, the size/number of tables, and the size/number of multipliers and compare with other related methods. Milos D. Ercegovac, Tomás Lang, Jean-Michel Muller, Arnaud Tisserand |
IEEE Trans. Computers | 2 |
| 1999 | Very-High Radix CORDIC Vectoring with Scalings and Selection by RoundingabstractA very-high radix algorithm and implementation for circular CORDIC in vectoring mode is presented. As for division, to simplify the selection function, the operands are pre-scaled. However in the CORDIC algorithm the coordinate x varies during the execution so several scalings might be needed; we show that two scalings are sufficient. Moreover, the compensation of the variable scale factor is done by computing the logarithm of the scale factor and performing the compensation by an exponential. Estimations of the delay for 32 bit precision show a speed up of about two with respect to the radix-4 case with redundant addition. This speed up is obtained at the cost of an increase in the hardware complexity, which is moderate for the pipelined implementation. Elisardo Antelo, Tomás Lang, Javier D. Bruguera |
IEEE Symposium on Computer Arithmetic | 2 |
| 1999 | Boosting Very-High Radix Division with Prescaling and Selection by RoundingabstractAn extension of the very-high radix division with prescaling and selection by rounding is presented. This extension consists in increasing the effective radix of the implementation by obtaining a few additional bits of the quotient per iteration, without increasing the complexity of the unit to obtain the prescaling factor nor the delay of an iteration. As a consequence, for some values of the effective radix, it permits an implementation with a smaller area and the same execution time than the original scheme. Estimations are given for 54-bit and 114-bit quotients. Paolo Montuschi, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 1999 | Low-Power Division: Comparison among Implementations of Radix 4, 8 and 16abstractAlthough division is less frequent than addition and multiplication, because of its longer latency it dissipates a substantial part of the energy in floating-point units. In this paper we explore the relation between the radix and the energy dissipated. Previous work has been done an radix-4 and radix-8 division. Here we extend this study to a radix-4 scheme with two overlapped radix-4 stages and compare the latency, area, and energy of the three implementations. Results show that by applying the low-power techniques the energy dissipation is reduced from 30% to 40%, with respect to the standard implementation. An additional 20% reduction can be obtained using a dual voltage. Moreover the energy dissipated to complete the division is roughly the same for the three radices. However, the power dissipation, proportional to the average current, increases with the radix. If reducing the energy is the priority, for the same latency radix-16 with dual voltage produces the smallest energy dissipation. Alberto Nannarelli, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 1999 | Multilevel Reverse-Carry Computation for Comparison and for Sign and Overflow Detection in AdditionabstractA fast calculation of the most-significant carry in an addition is required in several applications, such as comparisons of two operands by performing their difference, sign detection, and overflow detection. It has been proposed to calculate this carry by detecting the most-significant carry chain and collecting the carry after this chain. The detection can be implemented by a prefix tree of AND gates and the collecting by a multi-input OR or by a connection with tristate buffers. We have performed an estimate of the delay of this implementation for a datapath width of 64 bits and conclude that it is not significantly faster than the traditional carry-lookahead based method. We propose a multilevel implementation, which allows the overlap of successive levels thereby reducing the overall delay. For 64-bit operands we estimate a delay reduction of about 15% with respect to the traditional carry-lookahead based method, with a similar number of gates and number and length of interconnections. Tomás Lang, Javier D. Bruguera |
ICCD | 1 |
| 1999 | Low-Power Radix-4 Combined Division and Square RootabstractBecause of the similarities in the algorithm it is quite common to implement division and square root in the same unit. The purpose of this work is to implement a low-power combined radix-4 division and square root floating-point double precision unit and to compare its performance and energy consumption with a radix-4 division only unit. Previous work has been done on reducing the energy dissipated in a divider. Here we apply the same techniques to the combined division and square root unit and consider modifications and tradeoffs. Results show that the energy dissipation for the combined division/square root unit can be reduced by about 35% without affecting the latency and an additional 20% reduction can be obtained using a dual voltage. Moreover the unit is 5% slower than a divider and its energy dissipation is 15% higher. Alberto Nannarelli, Tomás Lang |
ICCD | 2 |
| 1999 | Leading-One Prediction with Concurrent Position CorrectionabstractThis paper describes the design of a leading-one prediction (LOP) logic for floating-point addition with an exact determination of the shift amount for normalization of the adder result. Leading-one prediction is a technique to calculate the number of leading zeros of the result in parallel with the addition. However, the prediction might be in error by one bit and previous schemes to correct this error result in a delay increase. The design presented here incorporates a concurrent position correction logic, operating in parallel with the LOP, to detect the presence of that error and produce the correct shift amount. We describe the error detection as part of the overall LOP, perform estimates of its delay and complexity, and compare with previous schemes. Javier D. Bruguera, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1999 | Very High Radix Square Root with Prescaling and Rounding and a Combined Division/Square Root UnitabstractAn algorithm for square root with prescaling and selection by rounding is developed and combined with a similar scheme for division. Since division is usually more frequent than square root, the main concern of the combined implementation is to maintain the low execution time of division, while accepting a somewhat larger execution time for square root. The algorithm is presented in detail, including the mathematical development of bounds for the first square-root digit and for the scaling factor. The proposed implementation is described, evaluated and compared with other combined div/sqrt units. The comparisons show that the proposed scheme potentially produces a significant speed-up for division, whereas, for square root, the speed-up is small. Tomás Lang, Paolo Montuschi |
IEEE Trans. Computers | 1 |
| 1999 | Low-Power DividerabstractThe general objective of our work is to develop methods to reduce the energy consumption of arithmetic modules while maintaining the delay unchanged and keeping the increase in the area to a minimum. Here, we illustrate some techniques for dividers realized in CMOS technology. The energy dissipation reduction is carried out at different levels of abstraction: from the algorithm level down to the implementation, or gate, level. We describe the use of techniques such as switching-off not active blocks, retiming, dual voltage, and equalizing the paths to reduce glitches. Also, we describe modifications in the on-the-fly conversion and rounding algorithm and in the redundant representation of the residual in order to reduce the energy dissipation. The techniques and modifications mentioned above are applied to a radix-4, divider, realized with static CMOS standard cells, for which a reduction of 40 percent is obtained with respect to the standard implementation. This reduction is expected to be about 60 percent if low-voltage gates, for dual voltage implementation, are available. The techniques used here should be applicable to a variety of arithmetic modules which have similar characteristics. Alberto Nannarelli, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1998 | Leading-one prediction scheme for latency improvement in single datapath floating-point addersabstractThis paper describes the design of a Leading-one Predictor (LOP) for floating-point addition, with an exact determination of the shift amount required. Previous LOP proposals produce a shift amount which might be in error by one position, so that this error has to be corrected after the addition terminates, increasing the critical path. Our design incorporates a concurrent detection of this error so that the amount of shift is corrected before the actual shift, without increasing the latency. The scheme presented here is applicable to the common case of a single datapath floating-point addition in which the output of the adder is always positive. We estimate the reduction in the critical path and the increase in area. Javier D. Bruguera, Tomás Lang |
ICCD | 2 |
| 1998 | Extension of the working-zone-encoding method to reduce the energy on the microprocessor data busabstractThe energy at the I/O pins is a significant part of the overall consumption of a chip. To reduce this energy, this work extends to the data bus the working zone encoding method originally applied to encoding an external address bus. This method is based on the conjecture that programs favor a few working zones of their address space at each instant and that addresses to consecutive accesses for each zone frequently differ by a small amount. When the difference is small, instead of sending the whole address, the method sends an offset with respect to the previous address to that zone, together with an identifier for the zone. In this paper the same idea is extended to the data bus. The approach has been applied to several SPEC95 streams of references to memory along with the corresponding data values in a system with multiplexed address and multiplexed instruction/data buses. Moreover the effect of instruction and data caches is evaluated. Comparisons are given with previous methods for bus encoding, showing significant improvement in all cases except for the multiplexed address bus with instruction cache, where the best scheme depends on the overhead of the implementation. Tomás Lang, Enric Musoll, Jordi Cortadella |
ICCD | 1 |
| 1998 | Low-power radix-8 dividerabstractThis work describes the design of a double-precision radix-8 divider. Low-power techniques are applied in the design of the unit, and energy-delay tradeoffs considered. The energy dissipation in the divider can be reduced by up to 70% with respect to a standard implementation not optimized for energy, without penalizing the latency. The radix-8 divider is compared with the one obtained by overlapping three radix-2 stages and with a radix-4 divider. Results show that the latency of our divider is similar to that of the divider with overlapped stages, but the area is smaller. The speed-up of the radix-8 over the radix-4 is about 20% and the energy dissipated to complete a division is almost the same, although the area of the radix-8 is 50% larger. Alberto Nannarelli, Tomás Lang |
ICCD | 2 |
| 1998 | Power-delay tradeoffs for radix-4 and radix-8 dividersabstractThe use of higher radices in division reduces the n umber of iterations to complete the operation, but increases the complexity of the circuit. In this paper we explore the influence of the radix on the pow er dissipation of a floating-point divider and the pow er-dela y tradeoffs. We compare the performance and the energy consumption per operation for a radix-4 and a radix-8 divider, realized in CMOS technology. A reduction of about 40% in the energy consumption is obtained for both radices (about 70% if low-v oltage gates, for dual v oltage implementation, are a vailable). Also the results sho w that the radix-8 divider is about 20% faster and the energy dissipated to perform a division is about the same, with respect to the radix-4. Alberto Nannarelli, Tomás Lang |
ISLPED | 2 |
| 1998 | Computation of sqrt(x/d) in a Very High Radix Combined Division/Square-Root Unit with ScalingabstractA very-high radix digit-recurrence algorithm for the operation /spl radic/(x/d) is developed, with residual scaling and digit selection by rounding. This is an extension of the division and square-root algorithms presented previously, and for which a combined unit was shown to provide a fast execution of these operations. The architecture of a combined unit to execute division, square-root, and /spl radic/(x/d) is described, with inverse square-root as a special case. A comparison with the corresponding combined division and square-root unit shows a similar cycle time and an increase of one cycle for the extended operation with respect to square-root. To obtain an exactly rounded result for the extended operation a datapath of about 2n bits is needed. An alternative is proposed which requires approximately the same width as for square-root, but produces a result with an error of less than one ulp. The area increase with respect to the division and square root unit should be no greater than 15 percent. Consequently, whenever a very high radix unit for division and square-root seems suitable, it might be profitable to implement the extended unit instead. Elisardo Antelo, Tomás Lang, Javier D. Bruguera |
IEEE Trans. Computers | 2 |
| 1998 | CORDIC Vectoring with Arbitrary Target ValueabstractThe computation of additional functions in the CORDIC module increases its flexibility. We consider here the extension of the vectoring mode (angle calculation) so that the vector is rotated until one of the coordinates (for instance y) attains a target value t (in contrast to the value 0, as in standard vectoring). The main problem in the algorithm is that the modulus of the vector is scaled in each CORDIC iteration so that a direct comparison of y[j] with t does not assure convergence. We present a scheme that overcomes this and in which the implementation consists of a standard CORDIC module plus a module to determine the direction of rotation. This improves over a previous proposal in which more complex iterations are introduced as part of the CORDIC algorithm. Moreover, an error analysis is performed to determine the datapath width required for convergence. Since this width is large, we consider also the characteristics of the algorithm for a narrower datapath. Tomás Lang, Elisardo Antelo |
IEEE Trans. Computers | 1 |
| 1998 | Working-zone encoding for reducing the energy in microprocessor address busesabstractThe energy consumption due to input-output pins is a substantial part of the overall chip consumption. To reduce this energy, this work presents the working-zone encoding (WZE) method for encoding an external address bus, based on the conjecture that programs favor a few working zones of their address space at each instant. In such cases, the method identifies these zones and sends through the bus only the offset of this reference with respect to the previous reference to that zone, along with an identifier of the current working zone. This is combined with a one-hot encoding for the offset. Several improvements to this basic strategy are also described. The approach has been applied to several address streams, broken down into instruction-only, data-only, and instruction-data traces, to evaluate the effect on separate and shared address buses. Moreover, the effect of instruction and data caches is evaluated. For the case without caches, the proposed scheme is specially beneficial for data address and shared buses, which are the cases where other codings are less effective. On the other hand, for the case with caches the best scheme for the instruction-only and data-only traces is the WZE, whereas for the instruction-data traces it is either the WZE or the bus-invert with four groups (depending on the energy overhead of these techniques). Enric Musoll, Tomás Lang, Jordi Cortadella |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1997 | CORDIC Vectoring with Arbitrary Target ValueabstractThe computation of additional functions in the CORDIC module increases its flexibility. We consider the extension of the vectoring mode (angle calculation) so that the vector is rotated until one of the coordinates (for instance, y) attains a target value t (in contrast to the value 0, as in standard vectoring). The main problem in the algorithm is that the modulus of the vector is scaled in each CORDIC iteration so that a direct comparison of y[j] with t does not assure convergence. We present a scheme that overcomes this and in which the implementation consists of a standard CORDIC module plus a module to determine the direction of rotation. This improves over a previous proposal in which more complex iterations are introduced as part of the CORDIC algorithm. Tomás Lang, Elisardo Antelo |
IEEE Symposium on Computer Arithmetic | 1 |
| 1997 | CORDIC-based computation of arccos and arcsinabstractCORDIC-based algorithms to compute cos/sup -1/(t), sin/sup -1/(t) and /spl radic/(1-t/sup 2/) are proposed. The implementation requires a standard CORDIC module plus a module to compute the direction of rotation, this being the same hardware required for the extended CORDIC vectoring, recently proposed by the authors. Although these functions can be obtained as a special case of this extended vectoring, the specific algorithm we propose here presents two significant improvements: (1) it achieves an angle granularity of 2/sup -n/ using the same datapath width as the standard CORDIC algorithm (about n bits, instead of about 2n which would be required using the extended veetoring), and (2) no repetitions of iterations are needed. The proposed algorithm is compatible with the extended vectoring and, in contrast with previous implementations, the number of iterations and the delay of each iteration are the same as for the conventional CORDIC algorithm. Tomás Lang, Elisardo Antelo |
ASAP | 1 |
| 1997 | Low latency word serial CORDICabstractIn this paper we present a modification of the CORDIC algorithm which reduces the number of iterations almost to half by merging two successive iterations of the basic algorithm. The two coefficients per iteration are obtained with only a small increase in the cycle time by estimating one of the coefficients. A correcting iteration method is used to correct the possible errors produced by the estimate. Moreover, the modified iteration permits the reduction of the number of cycles required for the compensation of the scaling factor. The resulting architecture is word serial, working both in rotation and vectoring operation modes, presenting a low latency in comparison with the classical CORDIC approach. Julio Villalba, Tomás Lang |
ASAP | 2 |
| 1997 | Reducing TLB power requirementsabstractTranslation look-aside buffers (TLBs) are small caches to speed-up address translation in processors with virtual memory.This paper considers two issues: (1) a comparison of the power consumption of fully-associative, set-associative, and direct mapped TLBs for the same miss rate and (2) the proposal of modifications of the basic cells and of the structure of set-associative TLBs to reduce the power.The power evaluation is done using a model and the miss rates are obtained from simulations of the SPEC92 benchmark.With respect to ( 1) we conclude that for small TLBs (high miss rates) fully-associative TLBs consume less power but for larger TLBs (low miss rates) set-associative TLBs are better.Moreover, the proposed modifications produce significant reductions in power consumption.Our evaluations show a reduction of 40 to 60% compared to the best traditional TLB.The proposed TLB implementation produces an increase in delay and in area.However, these increases are tolerable because the cycle time is determined by the slower cache and because the TLB area corresponds to only a small portion of the chip area.196 __ __, ~-~-,-, 'II -F 'r-pT-yy-,~';ri '-I:.* : ,< 7, Toni Juan, Tomás Lang, Juan J. Navarro |
ISLPED | 2 |
| 1997 | Exploiting the locality of memory references to reduce the address bus energyabstractThe energy consumption at the I of the overall chip consumption.4 0 pins is a significant part his paper presents a method for encoding an external address bus which lowers its activity and, thus, decreases the energy.This method relies on the locality of memory references.Since applications favor a few working zones of their address space at each instant, for an address to one of these zones only the offset of this reference with respect to the previous reference to that zone needs to be sent over the bus, along with an identifier of the current working zone.This is combined with a modified one-hot encoding for the offset.An estimate of the area and energy overhead of the encoder/decoder are given; their effect is small.The approach has been applied to two memory-intensive examples, obtaining a bus-activity reduction of about.2 3 in both of them.Comparisons are given with previous metho L for bus encoding, showing significant improvement. I . Enric Musoll, Tomás Lang, Jordi Cortadella |
ISLPED | 2 |
| 1997 | Error Analysis and Reduction for Angle Calculation Using the CORDIC AlgorithmabstractIn this paper, we consider the errors appearing in angle computations with the CORDIC algorithm (circular and hyperbolic coordinate systems) using fixed-point arithmetic. We include errors arising not only from the finite number of iterations and the finite width of the data path, but also from the finite number of bits of the input. We show that this last contribution is significant when both operands are small and that the error is acceptable only if an input normalization stage is included, making unsatisfactory other previous proposals to reduce the error. We propose a method based on the prescaling of the input operands and a modified CORDIC recurrence and show that it is a suitable alternative to the input normalization with a smaller hardware cost. This solution can also be used in pipelined architectures with redundant carry-save arithmetic. Elisardo Antelo, Javier D. Bruguera, Tomás Lang, Emilio L. Zapata |
IEEE Trans. Computers | 3 |
| 1996 | The Difference-bit CacheabstractThe difference-bit cache is a two-way set-associative cache with an access time that is smaller than that of a conventional one and close or equal to that of a direct-mapped cache. This is achieved by noticing that the two tags for a set have to differ at least by one bit and by using this bit to select the way. In contrast with previous approaches that predict the way and have two types of hits (primary of one cycle and secondary of two to four cycles), all hits of the difference-bit cache are of one cycle. The evaluation of the access time of our cache organization has been performed using a recently proposed on-chip cache access model. Toni Juan, Tomás Lang, Juan J. Navarro |
ISCA | 2 |
| 1996 | Low-power radix-4 dividerabstractThe general objective of our work is to develop methods to reduce the power consumption of arithmetic modules, while maintaining the delay unchanged and keeping the increase in the area to a minimum. Here we illustrate some techniques for a radix-4 divider realized in 0.6 /spl mu/m CMOS technology. Using techniques such as switching-off not active blocks, retiming the recurrence, equalizing the paths to reduce glitches, using gates with lower drive capability, and changing the redundant representation, we obtained a power consumption reduction of 35% with respect to the standard implementation. The techniques used here should be applicable to a variety of arithmetic modules which have similar characteristics. Alberto Nannarelli, Tomás Lang |
ISLPED | 2 |
| 1995 | Sign detection and comparison networks with a small number of transitionsabstractWe present an approach to reducing the average number of signal transitions (T,,) in the design of sign-detection and comparison of magnitudes. Our approach reduces T/sub av/ from 21n/8 (n-operand precision in bits) to 4.5 in the case of iterative implementation, and from about n to roughly k+n/2/sup k-1/ in the tree network implemented with k-bit modules. We also discuss comparison of small numbers. The approach is applicable to other arithmetic problems.> Milos D. Ercegovac, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 1995 | Very-high radix combined division and square root with prescaling and selection by roundingabstractAn algorithm for square root with prescaling is developed and combined with a similar scheme for division. An implementation is described, evaluated and compared with other combined div/sqrt implementations.> Tomás Lang, Paolo Montuschi |
IEEE Symposium on Computer Arithmetic | 1 |
| 1995 | 2-D DCT using on-line arithmeticabstractPresents a VLSI architecture for the evaluation of the (8/spl times/8)-point 2-D DCT with on-line arithmetic. The utilization of on-line arithmetic, in combination with an algorithm based on FCT and matrix multiplication, reduces the total hardware maintaining a data rate and a latency similar to approaches based on distributed or parallel arithmetic. The architecture has been integrated in a chip using a 1 /spl mu/ CMOS technology, occupying an area of 56.7 mm/sup 2/. Javier D. Bruguera, Tomás Lang |
ICASSP | 2 |
| 1995 | Vector Multiprocessors with Arbitrated Memory AccessabstractThe high latency of memory accesses is one of the factors that most contribute to reduce the performance of current vector supercomputers. The conflicts that can occur in the memory modules plus the collisions in the interconnection network in the case of multiprocessors make that the execution time of applications increases significantly. In this work we propose a memory access method that for both cases of vector uniprocessors and multiprocessors allows to perform stream accesses with the smallest possible latency in the majority of the cases. The basic idea is to arbitrate the memory access by defining the order in which the memory modules are visited. The stream elements are requested out of order. In addition, the access method also reduces the cost of the interconnection network. KEYWORDS Vector Processor, Vector Multiprocessor, Multi-module Memories, Conflict-Free Access, Out-of-Order Access. 1 Introduction The high latency of the memory accesses is one of the main factors t... Montse Peiron, Mateo Valero, Eduard Ayguadé, Tomás Lang |
ISCA | 4 |
| 1995 | Conflict-Free Access for Streams in Multimodule MemoriesabstractAddress transformation schemes, such as skewing and linear transformations, have been proposed to achieve conflict-free access for streams with constant stride. However, this is achieved only for some strides. In this paper, we extend these schemes to achieve this conflict-free access for a larger number of strides. The basic idea is to perform an out-of-order access to a stream of fixed length. This stream is then stored in a local memory and used in subsequent instructions. This mode of operation is suitable for vector processors and for processors with decoupled access. The scheme and mode of operation proposed produce the largest possible number of conflict-free strides. Memory systems with any ratio between the number of memory modules and memory latency are considered. The hardware for address calculations and access control is described and shown to be of similar complexity as that required for access in order.> Mateo Valero, Tomás Lang, Montse Peiron, Eduard Ayguadé |
IEEE Trans. Computers | 2 |
| 1994 | MOB forms: a class of multilevel block algorithms for dense linear algebra operationsabstractMultilevel block algorithms exploit the data locality in linear algebra operations when executed in machines with several levels in the memory hierarchy. It is shown that the family we call Multilevel Orthogonal Block (MOB) algorithms is optimal and easy to design and that using the multilevel approach produces significant performance improvements. The effect of interference in the cache, of the TLB misses, and of page faults are also considered. The multilevel block algorithms are evaluated analytically for an ideal memory system with M cache levels without interferences. Moreover, experimental results of the MOB forms in some present high performance workstations are presented. Juan J. Navarro, Toni Juan, Tomás Lang |
International Conference on Supercomputing | 3 |
| 1994 | High-Radix Division and Square-Root with SpeculationabstractThe speed of high-radix digit-recurrence dividers and square-root units is mainly determined by the complexity of the result-digit selection. We present a scheme in which a simpler function speculates the result digit, and, when this speculation is incorrect, a rollback or a partial advance is performed. This results in operations with a shorter cycle time and a variable number of cycles. The scheme can be used in separate division and square-root units, or in a combined one. Several designs were realized and compared in terms of execution time and area. The fastest unit considered is a radix-512 divider with a partial advance of six bits.> Jordi Cortadella, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1994 | Very-High Radix Division with Prescaling and Selection by RoundingabstractA division algorithm in which the quotient-digit selection is performed by rounding the shifted residual in carry-save form is presented. To allow the use of this simple function, the divisor (and dividend) is prescaled to a range close to one. The implementation presented results in a fast iteration because of the use of carry-save forms and suitable recodings. The execution time is calculated and several convenient values of the radix are selected. Comparison with other dividers for radices 2/sup 9/ to 2/sup 18/ is performed using the same assumptions.> Milos D. Ercegovac, Tomás Lang, Paolo Montuschi |
IEEE Trans. Computers | 2 |
| 1993 | Division with speculation of quotient digitsabstractThe speed of SRT-type dividers is mainly determined by the complexity of the quotient-digit selection, so that implementations are limited to low-radix stages. A scheme is presented in which the quotient-digit is speculated and, when this speculation is incorrect, a rollback or a partial advance is performed. This results in a division operation with a shorter cycle time and a variable number of cycles. Several designs have been realized, and a radix-64 implementation that is 30% faster than the fastest conventional implementation (radix-8) at an increase of about 45% in area per quotient bit has been obtained. A radix-16 implementation that is about 10% faster than the radix-8 conventional one, with the additional advantage of requiring about 25% less area per quotient bit, is also shown.> Jordi Cortadella, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 1993 | Very high radix division with selection by rounding and prescalingabstractA division algorithm in which the quotient-digit selection is performed by rounding the shifted residual in carry-save form is presented. To allow the use of this simple function, the divisor (and dividend) is prescaled to a range close to one. The implementation presented results in a fast iteration because of the use of carry-save forms and suitable recodings. The execution time is calculated, and several convenient values of the radix are selected. Comparison with other high-radix dividers is performed using the same assumptions.> Milos D. Ercegovac, Tomás Lang, Paolo Montuschi |
IEEE Symposium on Computer Arithmetic | 2 |
| 1993 | Multiplication/ division/ square root module for massively parallel computers
Milos D. Ercegovac, Tomás Lang |
Integr. | 2 |
| 1993 | Conflict-free access to streams in multiprocessor systems
Montse Peiron, Mateo Valero, Eduard Ayguadé, Tomás Lang |
Microprocess. Microprogramming | 4 |
| 1992 | MAMACG: a tool for automatic mapping of matrix algorithms onto mesh array computational graphsabstractThe design of MAMACG, a software tool for automatically mapping an important class of matrix algorithms into mesh array computational graphs, is described. MAMACG is a concrete realization of the multimesh graph (MMG) method, implemented in Elk, a dialect of LISP with built-in X-graphics capabilities.> Dinh Lê, Milos D. Ercegovac, Tomás Lang, Jaime H. Moreno |
ASAP | 3 |
| 1992 | Conflict-free access of vectors with power-of-two stridesabstractAn address mapping and an access order is presented for conflict-free access to vectors with any initial address and power-of-two strides. We show that for this conflict-free access it is necessary that the memory be unmatched and present an implementation for M=2T, where M is the number of modules and T the module latency. Moreover, the implementation allows the masking of the latency of the address calculation, of the mapper, and of the bus arbiter. Mateo Valero, Tomás Lang, Eduard Ayguadé |
ICS | 2 |
| 1992 | Increasing the Number of Strides for Conflict-Free Vector AccessabstractAddress transformation schemes, such as skewing and linear transformations, have been proposed to achieve conflict-free vector access for some strides in vector processors with multi-module memories. In this paper, we extend these schemes to achieve this conflict-free access for a larger number of strides. The basic idea is to perform an out-of-order access to vectors of fixed length, equal to that of the vector registers of the processor. Both matched and unmatched memories are considered: we show that the number of strides is even larger for the latter case. The hardware for address calculations and access control is described and shown to be of similar complexity as that required for access in order. Mateo Valero, Tomás Lang, José María Llabería, Montse Peiron, Eduard Ayguadé, Juan J. Navarro |
ISCA | 2 |
| 1992 | On-the-Fly RoundingabstractIn implementations of operations based on digit-recurrence algorithms such as division, left-to-right multiplication and square root, the result is obtained in digit-serial form, from most significant digit to least significant. To reduce the complexity of the result-digit selection and allow the use of redundant addition, the result-digit has values from a signed-digit set. As a consequence, the result has to be converted to conventional representation, which can be done on-the-fly as the digits are produced, without the use of a carry-propagate adder. The authors describe three ways to modify this conversion process so that the result is rounded. The resulting operation is fast because no carry-propagate addition is needed. The schemes described apply also to online arithmetic operations.> Milos D. Ercegovac, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1992 | Higher Radix Square Root with PrescalingabstractA scheme for performing higher radix square root based on prescaling of the radicand is presented to reduce the complexity of the result-digit selection. The scheme requires several steps, namely multiplication for prescaling the radicand, square root, multiplication for prescaling for the division, and division. Online algorithms are used to reduce the overall time and pipelining to reuse the different modules. An estimate of the execution time for a radix-256 unit for double-precision square root and a comparison with other implementations indicate that the proposed approach is an alternative to consider when designing a square-root unit.> Tomás Lang, Paolo Montuschi |
IEEE Trans. Computers | 1 |
| 1992 | Constant-Factor Redundant CORDIC for Angle Calculation and RotationabstractA constant-factor redundant-CORDIC (CFR-CORDIC) scheme, where the scale factor is kept constant while an angle for plane rotations is computed, is developed. The direction of rotation is determined from an estimate of the sign, and convergence is assured by suitably placed correcting iterations. The number of iterations in the CORDIC rotation unit is reduced by about 25% by expressing the direction of the rotation in radix-2 and radix-4, and conversion to conventional representation is done on the fly. The performance of CFR-CORDIC is estimated and compared with that of previously proposed schemes. It is found to provide an execution time similar to that of redundant CORDIC with a variable scaling factor, with a significant saving in area.> Jeong-A Lee, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1991 | SVD by constant-factor-redundant-CORDICabstractA constant-factor-redundant-CORDIC (CFR-CORDIC) scheme is developed where the scale factor is forced to be constant while computing angles for SVD (singular value decomposition). Based on the scheme, a fixed-point implementation of SVD is presented with the following additional features: (1) the final scaling operation is done by shifting; (2) the number of iterations in the CORDIC rotation unit is reduced by about 25% by expressing the direction of the rotation in radix-2 and radix-4; and (3) the conventional number representation of rotated output is obtained on-the-fly, not from a carry-propagate adder. The authors compare this scheme with previously proposed ones and show that it provides an execution time similar to that of redundant CORDIC with variable scaling factor, with significant saving in area.> Jeong-A Lee, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 1991 | A Comparison of Redundant CORDIC Rotation EnginesabstractThe CMOS implementation of two high performance rotation processors using redundant CORDIC are reviewed and compared. One of the designs uses a variable scaling factor while the other is with constant scaling. The latter also incorporates some radix-4 CORDIC stages. Characteristics for 1.2 mu m CMOS implementations are given.> John A. Harding, Tomás Lang, Jeong-A Lee |
ICCD | 2 |
| 1991 | Module to Perform Multiplication, Division, and Square Root in Systolic Arrays for Matrix Computations
Milos D. Ercegovac, Tomás Lang |
J. Parallel Distributed Comput. | 2 |
| 1991 | Architectural Support for Reduced Register Saving / Restoring in Single-Window Register FilesabstractThe use of registers in a processor reduces the data and instruction memory traffic. Since this reduction is a significant factor in the improvement of the program execution time, recent VLSI processors have a large number of registers which can be used efficiently because of the advances in compiler technology. However, since registers have to be saved/restored across function calls, the corresponding register saving and restoring (RSR) memory traffic can almost eliminate the overall reduction. This traffic has been reduced by compiler optimizations and by providing multiple-window register files. Although these multiple-window architectures produce a large reduction in the RSR traffic, they have several drawbacks which make the single-window file preferable. We consider a combination of hardware support and compiler optimizations to reduce the RSR traffic for a single-window register file, beyond the reductions achieved by compiler optimizations alone. Basically, this hardware keeps track of the registers that are written during execution, so that the number of registers saved is minimized. Moreover, hardware is added so that a register is saved in the activation record of the function that uses it (instead of in the record of the current function); in this way a register is restored only when it is needed, rather than wholesale on procedure return. We present a register saving and restoring policy that makes use of this hardware, discuss its implementation, and evaluate the traffic reduction when the policy is combined with intraprocedural and interprocedural compiler optimizations. We show that, on the average for the four general-purpose programs measured, the RSR traffic is reduced by about 90 percent for a small register file (i.e., 32 registers), which results in an overall data memory traffic reduction of about 15 percent. Miquel Huguet, Tomás Lang |
ACM Trans. Comput. Syst. | 2 |
| 1990 | A graph-based approach to map matrix algorithms onto local-access processor arraysabstractThe authors describe the application of the multi-mesh graph (MMG) method to the mapping of large matrix algorithms onto class-specific local-access processor arrays. These arrays consist of cells with large local memory (i.e., memory size proportional to the size of the problems) and low cell bandwidth (much smaller than the cell computation rate). The results given indicate that the MMG method allows the analysis of such issues as allocation operations to cells, load balancing, scheduling, synchronization, and overhead in computations and data transfers. These aspects are illustrated by mapping the LU-decomposition algorithm onto a linear memory-linked array. Performance estimates indicate that mapping with the MMG method produces 94% utilization of cells in the target structure used. Therefore, the MMG is a suitable tool for mapping matrix algorithms onto pre-existing arrays.> Jaime H. Moreno, Tomás Lang |
ASAP | 2 |
| 1990 | The Performance of a Faulty Multistage Interconnection Network with Diverting Switches and Correction Links
Lance Kurisaki, Tomás Lang |
ICPP (1) | 2 |
| 1990 | An Analytical Characterization of Generalized Shuffle-Exchange NetworksabstractThe shuffle-exchange network can be generalized by the definition of three parameters (n, r', k). In a generalized shuffle-exchange (GSE) network, 2/sup n/ inputs are first permuted by a shuffle such that an n-bit source address label is rotated left r'-bit positions to yield the destination address label. The exchange performs arbitrary permutations on 2/sup k/*2/sup k/ exchange switches. The GSE networks can emulate a variety of other networks, including orthogonally connected multidimensional cubes of all sizes, and they provide the possibility of incorporating alternate paths into networks without the addition of extra processing nodes or interconnections. Generalized shuffle-exchange networks are characterized herein by their connectivity, their diameter, and the number of alternate paths they permit.> Isaac D. Scherson, Peter F. Corbett, Tomás Lang |
INFOCOM | 3 |
| 1990 | Architectural Support for the Management of Tightly-Coupled Fine-Grain Goals in Flat Concurrent Prolog
Leon Alkalaj, Tomás Lang, Milos D. Ercegovac |
ISCA | 2 |
| 1990 | Nonuniform Traffic Spots (NUTS) in Multistage Interconnection Networks
Tomás Lang, Lance Kurisaki |
J. Parallel Distributed Comput. | 1 |
| 1990 | Redundant and On-Line CORDIC: Application to Matrix Triangularization and SVDabstractSeveral modifications to the CORDIC method of computing angles and performing rotations are presented: (1) the use of redundant (carry-free) addition instead of a conventional (carry-propagate) one; (2) a representation of angles in a decomposed form to reduce area and communication bandwidth; (3) the use of on-line addition (left-to-right, digit-serial addition) to replace shifters by delays; and (4) the use of online multiplication, square root, and division to compute scaling factors and perform the scaling operations. The modifications improve the speed and the area of CORDIC implementations. The proposed scheme uses efficiently floating-point representations. The application of the modified CORDIC method to matrix triangularization by Givens' rotations and to the computation of the singular value decomposition (SVD) are discussed.> Milos D. Ercegovac, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1990 | Radix-4 Square Root Without Initial PLAabstractA systematic derivation of a radix-4 square-root algorithm using redundant residual and result is presented. Unlike other similar schemes it does not use a table lookup or PLA for the initial step, resulting in a simpler implementation without any time penalty. The scheme can be integrated with division and incorporates an on-the-fly conversion and rounding of the result, thus eliminating a carry-propagate step to obtain the final result. The result-digit selection uses 3 bits of the result and 7 bits of the estimate of the residual.> Milos D. Ercegovac, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1990 | Simple Radix-4 Division with Opterands ScalingabstractA radix-4 division algorithm with operands scaling is proposed. The algorithm uses a recurrence with redundant addition (carry-save or signed-digit) and combines simple scaling with a quotient-selection function that depends only on the estimate of the partial remainder and is independent of the divisor. The scheme results in a significant speedup with respect to both the radix-2 and radix-4 without scaling.> Milos D. Ercegovac, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1990 | Fast Multiplication Without Carry-Propagate AdditionabstractConventional schemes for fast multiplication accumulate the partial products in redundant form (carry-save or signed-digit) and convert the result to conventional representation in the last step. This step requires a carry-propagate adder which is comparatively slow and occupies a significant area of the chip in a VLSI implementation. A report is presented on a multiplication scheme (left-to-right, carry-free, LRCF) that does not require this carry-propagate step. The LRCF scheme performs the multiplication most-significant bit first and produces a conventional sign-and-magnitude product (most significant n bits) by means of an on-the-fly conversion. The resulting implementation is fast and regular and is very well suited for VLSI. The LRCF scheme for general radix r and a radix-4 signed-digit implementation are presented.> Milos D. Ercegovac, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1989 | Radix-4 square root without initial PLAabstractA systematic derivation of a radix-4 square root algorithm using redundance in the partial residuals and the result is presented. Unlike other similar schemes, the algorithm does not use a table-lookup or programmable logic array (PLA) for the initial step. The scheme can be integrated with division. It also performs on-the-fly conversion and rounding of the result, thus eliminating a carry-propagate step to obtain the final result. The selection function uses 4 b of the result and 8 b of the estimate of the partial residual.> Milos D. Ercegovac, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 1989 | On-the-fly rounding for division and square rootabstractIn division and square root implementation based on digit-recurrence algorithms, the result is obtained in digit-serial form, from most significant digit to least significant. To reduce the complexity of the result-digit selection and to allow the use of redundant addition, the result-digit has values from a signed-digit set. As a consequence, the result has to be converted to conventional representation. This conversion can be done on-the-fly as the digits are produced, without the use of a carry-propagate adder. The authors describe how to modify this conversion process so that the result is rounded. The resulting operation is faster than what is done conventionally because no carry-propagate addition is needed. Three rounding methods that differ in the rounding error and the hardware and time required are described.> Milos D. Ercegovac, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 1989 | Multistage Networks Including Traffic with Real-Time Constraints
Lance Kurisaki, Tomás Lang |
ICPP (1) | 2 |
| 1988 | Implementation of fast radix-4 division with operands scalingabstractA radix-4 divider can potentially achieve a speedup of two with respect to a radix-2 implementation by halving the number of steps. However, the complicated quotient-digit selection function increases the critical path and almost eliminates the speedup. The authors present an implementation of a scheme that scales the divisor close to unity, making the quotient-selection function independent of the divisor. They show a gate-array implementation that achieves a speedup of 1.5 with respect to the radix-2 case, doubling the number of gates. The speedup achieved is still considerably lower than the theoretical maximum of twice.> Milos D. Ercegovac, Tomás Lang, Ramin Modiri |
ICCD | 2 |
| 1988 | Nonuniform Traffic Spots in Multistage Interconnection Networks
Tomás Lang, Lance Kurisaki |
ICPP (1) | 1 |
| 1988 | Graph-based Partitioning of Matrix Algorithms for Systolic Arrays: Application to Transitive Closure
Jaime H. Moreno, Tomás Lang |
ICPP (1) | 2 |
| 1988 | On-Line Scheme for Computing Rotation Factors
Milos D. Ercegovac, Tomás Lang |
J. Parallel Distributed Comput. | 2 |
| 1987 | On-line scheme for computing rotation factorsabstractAn integrated radix-2 on-line algorithm for computing rotation factors for matrix transformations is presented. The inputs and outputs are in parallel form, conventional 2's complement, floating-point representation. The exponents are computed using conventional arithmetic while the significands are processed using on-line algorithms. The conventional result is obtained by using an on-the-fly conversion scheme. The rotation factors are computed in 9+n clock cycles for n -bit significands. The clock period is kept small by the use of carry-save adder schemes. The implementation and performance of the algorithm are discussed. Milos D. Ercegovac, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 1987 | A block-and-actions generator as an alternative to a simulator for collecting architecture measurementsabstractTo design a new processor or to modify an existing one, designers need to gather data to estimate the influence of specific architecture features on the performance of the proposed machine (PM). To obtain this data, it is necessary to measure on an existing machine (EM) the dynamic behavior of typical programs. Traditionally, simulators have been used to obtain measurements for PMs. Since several hundred EM instructions are required to decode, interpret, and measure each simulated (PM) instruction, the simulation time of typical programs is prohibitively large. Thus, designers tend to simulate only small programs and the results obtained might not be representative of a real system behavior. In this paper we present an alternative tool for collecting architecture measurements: the Block-and-Actions Generator (BKGEN). BKGEN produces a version of the program being measured which is directly executable by the EM. This executable version is obtained directly with the EM compiler or with the PM compiler and a assembly-to-assembly translator. The choice between these alternatives depends on the EM and PM compiler technology and the type of measurements to be obtained. BKGEN also collects the PM events to be measured (called actions). Each EM block of instructions is associated with a PM block of actions so that when the program is executed, it collects the measurements associated with the PM. The main advantage of BKGEN is that the execution time is substantially reduced compared to the execution time of a simulator while collecting similar data. Thus, large typical programs (compilers, assemblers, word processors, ...) can be used by the designer to obtain meaningful measurements. Miquel Huguet, Tomás Lang, Yuval Tamir |
PLDI | 2 |
| 1987 | On-the-Fly Conversion of Redundant into Conventional RepresentationsabstractAn algorithm to convert redundant number representations into conventional representations is presented. The algorithm is performed concurrently with the digit-by-digit generation of redundant forms by schemes such as SRT division. It has a step delay roughly equivalent to the delay of a carry-save adder and simple implementation. The conversion scheme is applicable in arithmetic algorithms such as nonrestoring division, square root, and on-line operations in which redundantly represented results are generated in a digit-by-digit manner, from most significant to least significant. Milos D. Ercegovac, Tomás Lang |
IEEE Trans. Computers | 2 |
| 1986 | Replication and Pipelining in Multiple-Instance Algorithms
Jaime H. Moreno, Tomás Lang |
ICPP | 2 |
| 1985 | A division algorithm with prediction of quotient digitsabstractA division algorithm with a simple selection of quotient digits including prediction is possible if the divisor is restricted to a suitable range. The conditions that the divisor must satisfy to have the quotient digit qi+1predicted while computing Ri+1are determined. Some implementation considerations are also given. Milos D. Ercegovac, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 2 |
| 1983 | A performance evaluation of the multiple bus network for multiprocessor systemsabstractIn this paper we present a mathematical model to compute the bandwidth of the multiple bus interconnection network. Due to the computational complexity associated with the exact solution, the processors are removed from the queues at the end of each memory cycle to facilitate the analysis. This leads to approximate solutions which are both easier to obtain and very accurate. Mateo Valero, José María Llabería, Jesús Labarta, Emilio Sanvicente, Tomás Lang |
SIGMETRICS | 5 |
| 1983 | Minimization of Demand Paging for the LRU Stack Model of Program Behavior
Christopher Wood, Eduardo B. Fernández, Tomás Lang |
Inf. Process. Lett. | 3 |
| 1983 | Reduction of Connections for Multibus OrganizationabstractThe multibus interconnection network is an attractive solution for connecting processors and memory modules in a multiprocessor with shared memory. It provides a throughput which is intermediate between the single bus and the crossbar, with a corresponding intermediate cost. Tomás Lang, Mateo Valero, Miguel Angel Fiol |
IEEE Trans. Computers | 1 |
| 1982 | M-users B-servers arbiter for multiple-busses multiprocessors
Tomás Lang, Mateo Valero |
Microprocessing and Microprogramming | 1 |
| 1982 | Bandwidth of Crossbar and Multiple-Bus Connections for MultiprocessorsabstractIn this paper we compare the effective bandwidth in a multiprocessor with shared memory using as interconnection networks the crossbar or the multiple-bus. We consider a system with N processors and N memory modules, in which the processor requests to the memory modules are independent and uniformly distributed random variables. We consider two cases: in the first the processor makes another request immediately after a memory service, and in the second there is some internal processing time. Tomás Lang, Mateo Valero, Ignacio Alegre |
IEEE Trans. Computers | 1 |
| 1978 | Architectural Support for System Protection and Database SecurityabstractA set of architectural extensions to a machine of the type of IBM System/370 is proposed. The proposal involves hardware/software interaction to constrain the execution-time behavior of application and higher authority programs. The extensions consist of new states of privilege, enforcement of disciplined transition between states, hardware distinction of information types, and a mechanism to control data transfers between main and external storage. Application of the extensions to a shared database system, where users interact through a high-level language, shows that protection of the operating system and the database can be enhanced significantly with respect to errors or deliberate attacks from users Eduardo B. Fernández, Rita C. Summers, Tomás Lang, Charles D. Coleman |
IEEE Trans. Computers | 3 |
| 1977 | Database Buffer Paging in Virtual Storage SystemsabstractThree models, corresponding to different sets of assumptions, are analyzed to study the behavior of a database buffer in a paging environment. The models correspond to practical situations and vary in their search strategies and replacement algorithms. The variation of I/O cost with respect to buffer size is determined for the three models. The analysis is valid for arbitrary database and buffer sizes, and the I/O cost is obtained in terms of the miss ratio, the buffer size, the number of main memory pages available for the buffer, and the relative buffer and database access costs. Tomás Lang, Christopher Wood, Eduardo B. Fernández |
ACM Trans. Database Syst. | 1 |
| 1976 | Scheduling of Unit-Length Independent Tasks with Execution Constraints
Tomás Lang, Eduardo B. Fernández |
Inf. Process. Lett. | 1 |
| 1976 | Interconnections Between Processors and Memory Modules Using the Shuffle-Exchange NetworkabstractThe shuffle-exchange network is considered as an interconnection network between processors and memory modules in an array computer. Lawrie showed that this network can be used to perform some important permutations in log2 N steps. This work is extended and a network is proposed that permits the realization of any permutation in 0([mi][/mi]N) shuffle-exchange steps. Additional modifications to the basic. procedure are presented that can be applied to perform efficiently some permutations that were not realizable with the original mechanism. Finally, an efficient procedure is described for the realization of a shuffle permutation of N elements on an array computer with M memory modules where M < N. Tomás Lang |
IEEE Trans. Computers | 1 |
| 1976 | A Shuffle-Exchange Network with Simplified ControlabstractIn this paper, a control mechanism for a shuffle-exchange interconnection network of N cells is proposed. With this network it is possible to realize some important permutations in log2N shuffle-exchange steps. In the control mechanism presented, the control variables at step k are determined by a Boolean operation of the control variables at step k - 1. The Boolean operation is very simple so that little additional hardware is required for this computation. This control scheme requires only one bit per cell instead of a destination tag of log2N bits required by a control mechanism presented previously. The network can be used for the interconnection of memory modules and processors in an array computer, and for the accessing of blocks of consecutive data in large dynamic memories. It is also shown that the shuffle-exchange interconnection network permits the efficient partitioning of an array computer into subarrays to allow for the simultaneous computation of several identical problems. Tomás Lang, Harold S. Stone |
IEEE Trans. Computers | 1 |
| 1975 | Definition and Evaluation of Access Rules in Data Management SystemsabstractA data structuring scheme for authorization purposes is presented, that: Eduardo B. Fernández, Rita C. Summers, Tomás Lang |
VLDB | 3 |