EDBT 2026 Demo / reviewers in the wild / expert
James E. Stine
dblp:23/1649
· DBLP profile ↗
33ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-8767-390XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 4 first-author · 5 since 2021Theory of computation · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Alchemy : A Methodology for Scalable RTL Design Space Exploration
Ryan Swann, James E. Stine |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Shared Recurrence Floating-Point Divide/Sqrt and Integer Divide/Remainder With Early TerminationabstractDivision, square root, and remainder are fundamental operations required by most computer systems. Floating-point and integer operations are commonly performed on separate datapaths. This paper presents the first detailed implementation of a shared recurrence unit that supports floating-point division/square root and integer division/remainder. It supports early termination and shares the normalization shifter needed for integer and subnormal inputs. Synthesis results show that shared double-precision dividers producing at least 4 bits per cycle are 9 - 18% smaller and 3 - 16% faster than separate integer and floating-point units. Kevin Kim, Katherine Parry, Cedar Turek, Alessandro Maiuolo, Rose Thompson, James E. Stine |
IEEE Trans. Computers | 7 |
| 2024 | Unified Digit Selection for Radix-4 Recurrence Division and Square RootabstractDivision and square root are fundamental operations required by most computer systems. They are commonly implemented in hardware using radix-4 recurrence, which produces a 2-bit result digit on each step. Unified digit selection logic chooses the next quotient or square root digit based on a residual and divisor or square root approximation. This paper presents the first derivation of digit selection constants for unified radix-4 recurrence division and square root. James E. Stine, Milos D. Ercegovac, Alberto Nannarelli, Katherine Parry, Cedar Turek |
IEEE Trans. Computers | 2 |
| 2022 | Point-Targeted Sparseness and Ling Transforms on Parallel Prefix Adder TreesabstractRephrasing binary addition as a parallel prefix tree problem allows for the generation of high-performance architectures with logarithmic delay. Modern literature and implementation seeks to explore this prefix tree design space in order to identify optimal circuits for each target application. This paper broadens the scope of the design space by treating both preprocessing and post-processing nodes as malleable parts of the tree structure. Structures obtained through this novel approach are shown to have superior performance. Implementation results are presented using the SkyWater Open Source 130nm PDK and the open-source tools developed by this paper are made available. Teodor-Dumitru Ene, James E. Stine |
ARITH | 2 |
| 2022 | Implementation of High Performance IEEE 754-Posit Conversion HardwareabstractThis paper demonstrates the implementation of conversion hardware between floating-point operands in the standardized IEEE 754 format to those in a Posit type-III unum format, as well as the reverse conversion from the Posit format to IEEE 754. High performance conversion between the two standards will encourage the use of dual-mode systems where IEEE 754 and Posit architectures can be used interchangeably when they are advantageous to each other, respectively. High performance, structural RTL architecture is shown in comparison to behavioral architecture in existing literature, yielding significant performance improvements. Synthesis comparisons to existing behavioral architectures are shown for standard IEEE 754 precisions and their closest Posit analogues. Results are given in cmos32soi32nm MTCMOS technology using ARM-based standard-cells and commercial EDA toolsets. Brett Mathis, James E. Stine |
ISCAS | 2 |
| 2021 | A Comprehensive Exploration of the Parallel Prefix Adder Tree SpaceabstractParallel prefix tree adders allow for high- performance computation due to their logarithmic delay. Modern literature focuses on a well-known group of adder tree networks, with adder taxonomies unable to adequately describe intermediary structures. Efforts to explore novel structures focus mainly on the hybridization of these widely-studied networks. This paper presents a method of generating any valid adder tree network by using a set of three, simple, point-targeted transforms. This method allows for possibilities such as the generation and classification of any hybrid or novel architecture, or the incremental refinement of pre-existing structures to better meet performance targets. Synthesis implementation results are presented on the SkyWater 90nm technology. Teodor-Dumitru Ene, James E. Stine |
ICCD | 2 |
| 2020 | A Novel Rounding Algorithm for a High Performance IEEE 754 Double-Precision Floating-Point MultiplierabstractThis paper proposes a new algorithm for IEEE 754 Floating point multiplication along with a complete implementation supporting normalized and denormalized numbers. The new rounder is based on injection rounding but instead adds two injections to the intermediate product. The first injection handles the case when the product does not overflow while the second handles the case when it does overflow. A special adder is developed to handle the two injection constants while minimizing duplicated hardware. Dual injection rounding eliminates the complex split between upper and lower bit paths in the single injection rounding algorithm [1] which in turn reduces all three key design targets; delay (1.2%), area (6.4%), and power (7.7%). Our novel design is compared against three designs, a standard injection rounder, Synopsys® DesignWare™, and Cadence® ChipWare™. S. Ross Thompson, James E. Stine |
ICCD | 2 |
| 2019 | A Well-Equipped Implementation: Normal/Denormalized Half/Single/Double Precision IEEE 754 Floating-Point Adder/SubtracterabstractThis paper shows the implementation and design of a completely IEEE 754-compliant floating-point adder and subtracter. This design focuses on maintaining low critical delay and power, while still containing hardware for full IEEE 754 compliance. A novel 64-bit prefix adder structure is used, where most of the performance benefits over a standard design come from parallelization. This adder and subtracter has full support for binary64, binary32, and binary16 operands. It also has the ability to convert integer values to the IEEE 754 standard. Integer conversion for binary64, binary32, and binary16 is supported, as well as conversion between any of the aforementioned IEEE 754 precisions. This design also has full support for denormalized operands. Synthesis results presented use a cmos32soi 32nm CMOS technology and ARM standard-cells. Brett Mathis, James E. Stine |
ASAP | 2 |
| 2019 | Fast and Area-Efficient SRAM Word-Line OptimizationabstractA word line driver controls the access of cells in a row in Static Random Access Memories (SRAMs) and has a significant impact on SRAM speed and power consumption. When gate delay is the dominant factor, simple models are a good guideline for fast word lines. However, routing wire delay is significant when the row size is large, which causes these designs to be suboptimal. This paper presents an analytical optimization technique using a delay model that includes gate delay, wire resistance, and wire capacitance to optimize high-performance word line driver topologies for SRAMs. The proposed methodology has a maximum 45% delay improvement and 42% buffer cost reduction. James E. Stine, Matthew R. Guthaus |
ISCAS | 2 |
| 2018 | Clarifications and Optimizations on Rounding for IEEE-compliant Floating-Point MultiplicationabstractWhen implementing rounding for floating-point multipliers, the final carry-propagate addition always will fall on the critical delay path. In the case of IEEE-754 floating-point multipliers it is, therefore, important that extra logic is needed to correctly round and post-normalize the result. Previous implementations of IEEE 754 floating-point rounding schemes for multiplication have created simple algorithms for implementing rounding logic that operate in parallel with the addition and contribute little to the critical delay. However, improvements can be made to make this process faster and more efficient. This paper proposes improvements upon previous algorithms to implement all four IEEE 754 rounding modes to efficiently support the implementation. Previous solutions are presented and improvements are introduced to speed up the process to support IEEE 754 floating-point rounding for multiplication. Results are shown using a 32nm standard cell library. Tuan D. Nguyen, Son Bui, James E. Stine |
ASAP | 3 |
| 2017 | A Reconfigurable Replica Bitline to Determine Optimum SRAM Sense Amplifier Set TimeabstractProcess variation in deep nanometer technologies causes read, write and retention failures in SRAM and results in yield loss. Although using larger transistors and higher supply voltages helps to mitigate these failures and improve the yield, they come at the cost of more power consumption, more area overhead and performance degradation. Reconfigurable control-signal circuits in SRAM not only help to reduce the failure rates, but also allow calibration after fabrication to improve the yield, reduce the power consumption and upgrade the performance. In this paper, a reconfigurable replica bitline scheme for sense amplifier set time is proposed to reduce the read failures and provide a wide operating-voltage range for SRAM. The proposed reconfigurable replica bitline also allows recalibration when SRAM performance deteriorates due to device aging degradation. Monte Carlo simulation results for a 64 kb SRAM array show that the proposed technique can reduce the access time by 20% compared to conventional replica bitline design in IBM/Global Foundries (GF) cmos32soi 32nm technology at 0.9 V with no area overhead. Samira Ataei, James E. Stine |
ACM Great Lakes Symposium on VLSI | 2 |
| 2016 | OpenRAM: an open-source memory compilerabstractComputer systems research is often inhibited by the availability of memory designs. Existing Process Design Kits (PDKs) frequently lack memory compilers, while expensive commercial solutions only provide memory models with immutable cells, limited configurations, and restrictive licenses. Manually creating memories can be time consuming and tedious and the designs are usually inflexible. This paper introduces OpenRAM, an open-source memory compiler, that provides a platform for the generation, characterization, and verification of fabricable memory designs across various technologies, sizes, and configurations. It enables research in computer architecture, system-on-chip design, memory circuit and device research, and computer-aided design. Matthew R. Guthaus, James E. Stine, Samira Ataei, Mehedi Sarwar |
ICCAD | 2 |
| 2016 | A 64 kb differential single-port 12T SRAM design with a bit-interleaving scheme for low-voltage operation in 32 nm SOI CMOSabstractIn this paper, a novel differential single-port 12T SRAM bitcell is presented. This bitcell uses a read buffer to eliminate read disturbance, improves the read stability and achieves read static noise margin equal to its hold static noise margin. Using a column-based select signal this bitcell provides a half-select free feature, facilitating a bit-interleaving structure to reduce multi-bit soft errors by conventional error correcting code techniques. By boosting the wordline and select signal voltage, this bitcell can read and write with no error at 300 mV while data can be held down to 250 mV in standby mode. Bitline leakage suppression in 12T bitcell allows more bitcells per bitline for high density SRAMs and provides faster read operation. This paper also introduces OpenRAM, an open-source memory compiler, that provides a platform for the generation, characterization, and verification of fabricable memory designs across various technologies, sizes, and configurations. Using OpenRAM, a 64 kb 12T SRAM macro is designed in IBM 32 nm SOI CMOS technology that operates down to 0.3 V with 50 MHz operating frequency while it functions at 0.9 V with 2.2 GHz operating frequency, as well. Samira Ataei, James E. Stine, Matthew R. Guthaus |
ICCD | 2 |
| 2015 | An IEEE 754 double-precision floating-point multiplier for denormalized and normalized floating-point numbersabstractThis paper discusses an optimized double-precision floating-point multiplier that can handle both denormalized and normalized IEEE 754 floating-point numbers. Discussions of the optimizations are given and compared versus similar implementations, however, the main objective is keeping compliant for denormalized IEEE 754 floating-point numbers while still maintaining high performance operations for normalized numbers. S. Ross Thompson, James E. Stine |
ASAP | 2 |
| 2015 | Multi Replica Bitline Delay Technique for Variation Tolerant Timing of SRAM Sense AmplifiersabstractTiming variation of sense amplifier enable (SAE) attributable to the random variation of transistor threshold Voltage is reduced by a novel Multi Replica Bitline Delay technique to provide the best tracking with process variations for SRAM applications. Multi replica bitline with a sufficient count of replica cells are utilized in parallel and delay of RBLs is added together to generate timing for sense amplifier (SA). Simulation results in IBM 65nm CMOS technology show that 50% timing variation is reduced at 1.0V supply Voltage. Samira Ataei, James E. Stine |
ACM Great Lakes Symposium on VLSI | 2 |
| 2014 | Additional optimizations for parallel squarer unitsabstractThis paper discusses modifications to algorithms to compute parallel squaring. The method described in this paper improves upon designs previously presented utilizing Boolean simplifications. The algorithms discussed in this paper significantly saves area and delay for squarers ranging from 8 bits to 32 bits. Results are shown for area, delay, and power using Virtex 5 Xilinx FPGAs. Son Bui, James E. Stine |
ISCAS | 2 |
| 2014 | Optimized cubic chebyshev interpolator for elementary function hardware implementationsabstractThis paper presents a cubic interpolator for computing elementary functions using truncated-matrix arithmetic units and an optimized number of coefficients bits. The proposed method optimizes the initial coefficient values found using a Chebyshev series approximation, minimizing the maximum absolute error of the interpolator output. The resulting designs can be utilized for approximating any function up to 53-bits of precision (IEEE double precision significant). Area, delay and power estimates are given for 16, 24 and 32-bit cubic interpolators that compute the reciprocal function, targeting a 65nm CMOS technology from IBM. Results indicate the proposed method uses smaller arithmetic units and has reduced lookup table sizes than previously proposed methods. Masoud Sadeghian, James E. Stine, E. George Walters III |
ISCAS | 2 |
| 2014 | Enhancing the Unified Logical Effort algorithm for branching and load distributionabstractIn this paper, the authors proposed a novel technique to calculate the capacitance distribution and branching effort of a multiple fan-out logic path for equal propagation delay in each path regardless of number of gates and lengths of the wire segments in those logic path. The authors utilize the prior methods for the Unified Logical Effort (ULE) methodology as the basis of delay estimation and transistor sizing for the work of this paper. The runtime for fan-out of 2 is logarithmic in n or O(log2(n)), where n is the precision index. Several examples are also analyzed and detailed using the proposed algorithm showing how the branching effort can easily be calculated in the presence of a complex circuit tree with arbitrary loading. Mehedi Sarwar, James E. Stine |
ISCAS | 2 |
| 2011 | A recursive-divide architecture for multiplication and divisionabstractMultipliers have been key and critical components for most application-specific and general-purpose computer architectures. However, these architectures have been transitioning towards multiple cores that can process large amounts of data through parallel approaches to computation. Unfortunately, traditional arithmetic functional units that worked well for single-core architectures have the side-effect of incurring large amounts of area and power. Consequently, multi-core computer architecture need new ways of thinking about increased through- put to handle large amounts of data. This paper presents a recursive high radix divide unit that is modified to handle both multiplication and division targeted at multi-core architectures. Results are obtained with a 65 nm technology and show a significant decrease in area and power while still maintaining a low total latency by utilizing high radix encoding within the functional unit. More importantly, because the datapath unit requires complex recoding, it does not increase its latency as the bit size increases. Therefore, these units can occupy low amounts of area while still maintaining high amounts of processing power. James E. Stine, Amey Phadke, Surpriya Tike |
ISCAS | 1 |
| 2009 | Parallel Prefix Ling Structures for Modulo 2^n-1 AdditionabstractParallel-prefix adders draw significant amounts of attention within general-purpose and application-specific architectures because of their logarithmic delay and efficient implementation in VLSI. This paper proposes a scheme to enhance parallel-prefix adders for modulo 2n- 1 addition by incorporating Ling equations into parallel-prefix structures. As opposed to previous research, this work clarifies the use of Ling equations for Modulo and provides enhancements to its implementation. Results are given in this work for a placed and routed design within a variation-aware 45 nm technology. The implementation results show a significant improvement in delay and even a reduction in power dissipation. James E. Stine |
ASAP | 2 |
| 2008 | Compressor trees for decimal partial product reductionabstractDecimal multiplication has grown in interest due to the recent announcement of new IEEE 754R standards and the availability of high-speed decimal computation hardware. Prior research enabled partial products to be coded more efficiently for their use in radix 10 architectures. This paper clarifies previous techniques for partial product reduction using carry-save adders and presents a new 4:2 compressor structure. This new structure improves performance at the expense of more gates, however, regularity is introduced into the circuit to promote implementations in Very Large Scale Integration (VLSI) Designs. Results are presented and compared for several designs using a TSMC SCN6M $0.18 mu m feature size. Ivan D. Castellanos, James E. Stine |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | A 64-bit Decimal Floating-Point ComparatorabstractDecimal arithmetic is growing in importance as scientific studies reveal that current financial and commercial applications spend a high percentage overhead in this type of calculations. Typically, software is utilized to emulate decimal floating point arithmetic in these applications. On the other hand, functional units that employ decimal floating point hardware can improve performance by two or three orders of magnitude. This paper presents the design and implementation of a novel decimal floating-point comparator compliant with the current draft revision of the IEEE-754 Standard for floating-point arithmetic. It utilizes a novel BCD magnitude comparator with logarithmic delay and it supports 64- bit decimal floating-point numbers. Area and delay results are examined for an implementation in TSMC SCN6M SCMOS technology. Ivan D. Castellanos, James E. Stine |
ASAP | 2 |
| 2006 | Low power binary addition using carry increment addersabstractSparse prefix tree adders like carry-select adders and carry-increment adders are commonly used in the implementation of high-speed parallel adders. This paper presents a novel Ling carry-increment adder, which further reduces the area and power consumption as compared to a conventional carry-increment adder. The proposed algorithm utilizes Ling pseudo-carries in both the prefix tree and the output sum blocks. The algorithm is verified with the implementation of two 64-bit adders using the conventional carry-increment and the proposed Ling carry-increment algorithms. An 8-bit sum block using the proposed algorithm uses 7% fewer devices and consumes 16% less energy, while the complete adder uses 8% fewer transistors and consumes 7% less energy Johannes Grad, James E. Stine |
ISCAS | 2 |
| 2006 | Compressed symmetric tables for accurate function approximation of reciprocalsabstractThis paper presents a high-speed method for accurate function approximations of reciprocals. This method employs parallel table lookups followed by multi-operand addition. Similar to previous methods, it takes advantage of leading zeros and symmetry in the table entries to reduce the table sizes. This method, although similar to previous methods, achieves lower memory sizes without having to increase the number of operands in the multi-operand addition. James E. Stine, Nitin Naresh |
ISCAS | 1 |
| 2005 | New algorithms for carry propagationabstractThis paper presents the analysis and implementation of different algorithms for carry propagation in binary adders. Besides traditional AND-OR-Invert based adders, NAND-type and NOR-type adders with single-gate delay per bit are studied as well. A unified classification based on work by Ling and Doran is presented. The new techniques are applied to ripple-carry adders, carry-skip adders and conditional-sum adders. Transistor level implementations in static and dynamic CMOS logic, as well as simulation results using a 90nm technology are presented. Johannes Grad, James E. Stine |
ACM Great Lakes Symposium on VLSI | 2 |
| 2004 | Modified booth truncated multipliersabstractTruncated multiplication provides an efficient method for reducing the power dissipation and area of rounded parallel multipliers in digital signal processing systems. With this technique, the products of parallel multipliers are rounded to a shorter word size and the least-significant columns of the multiplication matrix are not used. This technique provides significant savings in terms of power dissipation for unsigned multiplication. Although previous implementations involved unsigned and signed array and tree multipliers, this technique can be equally applied to multiplication using Booth-encoding. This paper presents the design and implementation of parallel and truncated multipliers that use Booth-encoding and compressors for signed multiplication. Initial estimates indicate that truncated parallel multipliers dissipate less power than standard parallel multipliers for operand sizes of 16 bits. Alok A. Katkar, James E. Stine |
ACM Great Lakes Symposium on VLSI | 2 |
| 2003 | Variations on Truncated MultiplicationabstractTruncated multiplication can be used to significantly reduce the power dissipation for applications that do not require correctly-rounded results. This paper presents an efficient method for truncated multiplication called hybrid-correction truncation that utilizes the advantages of two previous methods to obtain lower average and maximum absolute error. Comparisons are presented contrasting power, area, and delay for all three methods compared to standard parallel multipliers. Estimates indicate that hybrid truncated multipliers dissipate slightly less power and consume slightly less area than previous methods for truncated multiplication. In addition, utilization of the hybrid truncation method can provide a method for altering the implementation within certain limits to meet a given precision. James E. Stine, Oliver M. Duverne |
DSD | 1 |
| 2003 | A pipelined clock-delayed domino carry-lookahead adderabstractClock-delayed (CD) domino is a dynamic logic family developed to provide both inverting and non-inverting logic on single-rail gates. It is self-timed and can be easily pipelined for superior speed performance which makes it an attractive option in high-speed logic implementation. This paper presents the design of two different high-speed pipeline configurations of a 32-bit carry look-ahead adder using CD domino gates utilizing efficient clocking methodology to reduce the overall critical path delay. Bhushan A. Shinkre, James E. Stine |
ACM Great Lakes Symposium on VLSI | 2 |
| 2001 | Combined IEEE Compliant and Truncated Floating Point Multipliers for Reduced Power DissipationabstractTruncated multiplication can be used to significantly reduce power dissipation for applications that do not require correctly rounded results. This paper presents a power efficient method for designing floating point multipliers that can perform either correctly rounded IEEE compliant multiplication or truncated multiplication, based on an input control signal. Compared to conventional IEEE floating point multipliers, these multipliers require only a small amount of additional area and delay, yet provide a significant reduction in power dissipation for applications that do not require IEEE compliant results. Kent E. Wires, Michael J. Schulte, James E. Stine |
ICCD | 3 |
| 1999 | Approximating Elementary Functions with Symmetric Bipartite TablesabstractThis paper presents a high-speed method for function approximation that employs symmetric bipartite tables. This method performs two parallel table lookups to obtain a carry-save (borrow-save) function approximation, which is either converted to a two's complement number or is Booth encoded. Compared to previous methods for bipartite table approximations, this method uses less memory by taking advantage of symmetry and leading zeros in one of the two tables. It also has a closed-form solution for the table entries, provides tight bounds on the maximum absolute error, and can be applied to a wide range of functions. A variation of this method provides accurate initial approximations that are useful in multiplicative divide and square root algorithms. Michael J. Schulte, James E. Stine |
IEEE Trans. Computers | 2 |
| 1998 | A Combined Interval and Floating Point MultiplierabstractInterval arithmetic provides an efficient method for monitoring and controlling errors in numerical calculations. However, existing software packages for interval arithmetic are often too slow for numerically intensive computations. This paper presents the design of a multiplier that performs either interval or floating point multiplication. This multiplier requires only slightly more area and delay than a conventional floating point multiplier, and is one to two orders of magnitude faster than software implementations of interval multiplication. James E. Stine, Michael J. Schulte |
Great Lakes Symposium on VLSI | 1 |
| 1997 | Symmetric Bipartite Tables for Accurate Function ApproximationabstractThe paper presents a methodology for designing bipartite tables for accurate function approximation. Bipartite tables use two parallel table lookups to obtain a carry-save (borrow-save) function approximation. A carry propagate adder can then convert this approximation to a two's complement number or the approximation can be directly Booth encoded. Our method for designing bipartite tables, called the Symmetric Bipartite Table Method, utilizes symmetry in the table entries to reduce the overall memory requirements. It has several advantages over previous bipartite table methods in that it: (1) provides a closed form solution for the table entries; (2) has right bounds on the maximum absolute error; (3) requires smaller table lookups to achieve a given accuracy; and (4) can be applied to a wide range of functions. Compared to conventional table lookups, the symmetric bipartite tables presented are 15.0 to 41.7 times smaller when the operand size is 16 bits and 99.1 to 273.9 times smaller when the operand size is 24 bits. Michael J. Schulte, James E. Stine |
IEEE Symposium on Computer Arithmetic | 2 |
| 1997 | Accurate Function Approximations by Symmetric Table Lookup and AdditionabstractThis paper presents a high-speed method for accurate function approximations. This method employs parallel table lookups followed by multi-operand addition. It takes advantage of leading zeros and symmetry in the table entries to reduce the table sizes. By increasing the number of tables and the number of operands in the multi-operand addition, the amount of memory is significantly reduced. This method provides a closed form solution for the table entries and can be applied to a variety of elementary functions. Compared to conventional table lookups, it requires two to three orders of magnitude less memory. The design of elementary function generators that use this method are presented and compared to similar methods for elementary function generation. Michael J. Schulte, James E. Stine |
ASAP | 2 |