Graham A. Jullien

dblp:33/6775 · DBLP profile ↗
← Back
90ranked-venue papers
8as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-authorTheory of computation · 13 · 1 first-authorArtificial intelligence and machine learning · 2Computer networks · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
12 papers
Integrated circuit design · 65% Emerging computing paradigms · 19% Processor architecture and microarchitecture · 9%
Theoretical computer science
2 papers
Algorithms and data structures · 96% Coding theory · 4%
Network and information security
1 paper
Cryptographic primitives and cryptanalysis · 100%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design
digital circuit design
0.162005
Efficient Techniques for Binary-to-Multidigit Multidimensional Logarithmic Number System Conversion Using Range-Addressable Look-Up Tables · IEEE Trans. Computers 2005
A Number System with Continuous Valued Digits and Modulo Arithmetic · IEEE Trans. Computers 2002
A New Design Technique for Column Compression Multipliers · IEEE Trans. Computers 1995
Emerging computing paradigms › field-coupled nanocomputing
quantum-dot cellular automata
0.112006
Design Tools for an Emerging SoC Technology: Quantum-Dot Cellular Automata · Proc. IEEE 2006
Integrated circuit design
analog and mixed-signal circuits
0.012002
A Number System with Continuous Valued Digits and Modulo Arithmetic · IEEE Trans. Computers 2002
Processor architecture and microarchitecture › computer arithmetic
number representation
0.012002
A Number System with Continuous Valued Digits and Modulo Arithmetic · IEEE Trans. Computers 2002
Integrated circuit design › digital circuit design
arithmetic circuit design
0.051995
A New Design Technique for Column Compression Multipliers · IEEE Trans. Computers 1995
Large Dynamic Range Computations over Small Finite Rings · IEEE Trans. Computers 1994
On Modulus Replication for Residue Arithmetic Computations of Complex Inner Products · IEEE Trans. Computers 1990
Cryptographic primitives and cryptanalysis › public-key cryptography
modular exponentiation
0.012000
Complexity and Fast Algorithms for Multiexponentiations · IEEE Trans. Computers 2000
Integrated circuit design
residue number system arithmetic
0.051994
Large Dynamic Range Computations over Small Finite Rings · IEEE Trans. Computers 1994
On Modulus Replication for Residue Arithmetic Computations of Complex Inner Products · IEEE Trans. Computers 1990
High-speed signal processing using systolic arrays over finite rings · IEEE J. Sel. Areas Commun. 1988
Integrated circuit design
digital arithmetic circuits
0.011999
Theory and Applications of the Double-Base Number System · IEEE Trans. Computers 1999
Electronic design automation
logic synthesis
0.022006
Design Tools for an Emerging SoC Technology: Quantum-Dot Cellular Automata · Proc. IEEE 2006
A New Design Technique for Column Compression Multipliers · IEEE Trans. Computers 1995
Emerging computing paradigms
approximate and stochastic computing
0.021994
Large Dynamic Range Computations over Small Finite Rings · IEEE Trans. Computers 1994
On Modulus Replication for Residue Arithmetic Computations of Complex Inner Products · IEEE Trans. Computers 1990
Integrated circuit design › digital circuit design › arithmetic circuit design
multiplier design
0.011995
A New Design Technique for Column Compression Multipliers · IEEE Trans. Computers 1995
Integrated circuit design
low-power circuit design
0.012002
A Number System with Continuous Valued Digits and Modulo Arithmetic · IEEE Trans. Computers 2002
Integrated circuit design
VLSI design
0.021994
High-speed signal processing using systolic arrays over finite rings · IEEE J. Sel. Areas Commun. 1988
Large Dynamic Range Computations over Small Finite Rings · IEEE Trans. Computers 1994
Hardware accelerators and domain-specific architectures
systolic array
0.011988
High-speed signal processing using systolic arrays over finite rings · IEEE J. Sel. Areas Commun. 1988
Processor architecture and microarchitecture › pipelining
parallel-pipeline architecture
0.011994
Large Dynamic Range Computations over Small Finite Rings · IEEE Trans. Computers 1994
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator › convolution acceleration
convolution accelerator
0.011983
Processor Architectures for Two-Dimensional Convolvers Using a Single Multiplexed Computational Element with Finite Field Arithmetic · IEEE Trans. Computers 1983
Integrated circuit design › digital circuit design › arithmetic circuit design
modular multiplication
0.011980
Implementation of Multiplication, Modulo a Prime Number, with Applications to Number Theoretic Transforms · IEEE Trans. Computers 1980
Physical-layer communications
digital signal processing
0.011988
High-speed signal processing using systolic arrays over finite rings · IEEE J. Sel. Areas Commun. 1988
Integrated circuit design
digital signal processing circuits
0.011979
Implementation of FFT Structures Using the Residue Number System · IEEE Trans. Computers 1979
High-performance computing
fast fourier transform
0.011979
Implementation of FFT Structures Using the Residue Number System · IEEE Trans. Computers 1979
Integrated circuit design
finite field arithmetic
0.011983
Processor Architectures for Two-Dimensional Convolvers Using a Single Multiplexed Computational Element with Finite Field Arithmetic · IEEE Trans. Computers 1983
Coding theory › finite fields
finite field arithmetic
0.011980
Implementation of Multiplication, Modulo a Prime Number, with Applications to Number Theoretic Transforms · IEEE Trans. Computers 1980
Coding theory › finite fields › finite field arithmetic
number theoretic transform
0.011980
Implementation of Multiplication, Modulo a Prime Number, with Applications to Number Theoretic Transforms · IEEE Trans. Computers 1980

Methods — techniques the papers use, named apart from their topics

range-addressable look-up table · 0.1binary-like complex arithmetic · 0.1index calculus · 0.0quantization analysis · 0.0half adder · 0.0full adder · 0.0column compression · 0.0scaling · 0.0multivariate mapping · 0.0polynomial residue coding · 0.0lookup table · 0.0read-only memory · 0.0lookup tables · 0.0
YearPublicationVenuePosition
2010 Recursive architectures for 2DLNS multiplication
abstract
In the area of signal processing, digital circuits are advantageous in terms of lower sensitivity to noise and process variations, simplicity of design, programmability and test, while they attain higher speed, more functionality per chip, lower power dissipation or lower cost. Since some of DSP algorithms heavily rely on multiplication, there are constant demands for more efficient multiplication structures. In this paper, 2DLNS-based multiplication architectures with two different levels of recursion are presented. Our architectures combine some of the flexibility of software with the high performance of hardware through implementing the recursive multiplication schemes on a 2DLNS processing structure. The implementations demonstrate the efficiency of 2DLNS in DSP applications and show outstanding results in terms of operation delay and dynamic power consumption.
Mahzad Azarmehr, Majid Ahmadi, Graham A. Jullien
ISCAS3
2010 Video-Active RAM: A processor-in-memory architecture for video coding applications
abstract
This paper presents the Video-Active RAM (VA-RAM) architecture for video coding applications. VA-RAM is a processor-in-memory architecture customized for video coding applications. The VA-RAM architecture has been used to implement several video coding algorithms. A prototype of the VA-RAM for block-based integer-pixel ME has been fabricated using the TSMC 0.18 um CMOS technology. The architecture uses 89,687 gates and 18,976 bits of on-chip memory. At a maximum clock frequency of 125 MHz, the fabricated chip is able to process 15 4CIF fps. It consumes 84.68 mW at 125 MHz and has core area of 2.9 mm2.
Mohammed Sayed, Wael Badawy, Graham A. Jullien
ISCAS3
2009 Low-complexity algorithm for fractional-pixel motion estimation
abstract
This paper presents low-complexity algorithm for fractional-pixel motion estimation (FME). The proposed algorithm uses a mathematical model to approximate the matching error at fractional-pixel locations. Hence, no interpolation is required at fractional-pixel locations. The matching error values at integer-pixel locations are used to evaluate the model coefficients. The performance of the proposed algorithm has been compared with other FME algorithms including the full quarter-pixel search (FQPS) algorithm. The performance analysis shows that the proposed algorithm has about 93% less computational complexity than the FQPS algorithm for approximately 0.2 dB drop in the reconstruction PSNR values.
Mohammed Sayed, Wael Badawy, Graham A. Jullien
ICIP3
2008 On the refinement of the DCT/IDCT scaling factor sensitivity
abstract
This paper proposes to represent the floating-point multipliers required to perform IDCT implementations using a rational Diophantine (i.e. ratio of integers) approximation with a common denominator, which is not necessarily a power of two. A case study to support this proposal is presented by applying the proposed scheme to Chenpsilas IDCT algorithm. Results show better performance when applying the proposed scheme compared to the traditional shift process. Similar studies can be obtained for any other potential up-scaling factor, and by modifying any other potential IDCT fast algorithm.
Ihab Amer, Wael Badawy, Vassil S. Dimitrov, Graham A. Jullien
ICME4
2008 A Low noise CMOS image sensor with an emission filter for fluorescence applications
abstract
This paper presents a 128x128 low noise CMOS image sensor with emission filter for fluorescence detection. The imager, fabricated in 0.18 mum CMOS technology, provides low- noise operation by employing both the active reset (AR) technique and the active column sensor (ACS) readout method. The emission filter was fabricated using PDMS and Sudan II Blue dye mixed, spin-coated and deposited in the class 1000 clean room. The designed filter is suitable for excitation at wavelengths below 340 nm and emission at 450nm and above. Filter properties, such as thickness, transmission and efficiency of utilization with the fabricated imager are discussed. Preliminary measurements of the system using microbeads are also presented.
Marianna Beiderman, Terence Tam, Alexander Fish, Graham A. Jullien, Orly Yadid-Pecht
ISCAS4
2008 Robust analog neural network based on continuous valued number system
abstract
This paper explores properties of an analog artificial neural network architecture based on Continuous Valued Number System (CVNS) with distributed neurons. In conventional lumped neural networks, the effect of weight quantization errors effects the performance of the network as the network size increases. However, based on a stochastic model it is shown here that the CVNS capability in detecting and correcting errors along with the inherent self-scaling property of distributed neurons, can control the output quantization noise to signal ratio. This property contributes to a robust analog VLSI architecture based on analog distributed CVNS-Adaline neurons.
Mitra Mirhassani, Majid Ahmadi, Graham A. Jullien
ISCAS3
2008 Low-Power Mixed-Signal CVNS-Based 64-Bit Adder for Media Signal Processing
abstract
In this paper, design of a mixed-signal 64-bit adder based on the continuous valued number system (CVNS) is presented. The 64-bit adder is generated by cascading four 16-bit radix-2 CVNS adders. Truncated summation of the CVNS digits reduced the number of required interconnections in the system, which in turn reduced design complexity and hardware costs. This adder can perform one 64-bit, two 32-bit, four 16-bit, or eight 8-bit additions on demand for media signal processing applications. The compact and low-power and low-noise design of the adder is suitable for this type of application. The 64-bit adder designed in TSMC CMOS 0.18-mum technology, has a worst case delay of 1.5 ns, energy dissipation of about 14 pJ with the core area of 13 250mum2.
Mitra Mirhassani, Majid Ahmadi, Graham A. Jullien
IEEE Trans. Very Large Scale Integr. Syst.3
2007 Digital Multiplication using Continuous Valued Digits
abstract
Binary multiplication is one of the fundamental arithmetic operations, and it is used in digital filters and signal processing applications. The Continuous Valued Number System (CVNS) is a recently introduced number system that allows digital arithmetic, with arbitrary precision, to be implemented with analog circuitry. Due to the analog nature of the numbers system, CVNS reduces the total system and cross talk noise. An8×8digital multiplier is proposed, using CVNS compressors for reducing the digital partial products. A new definition for the CVNS compressor for the first time is introduced, along with a novel CMOS current mode circuit. The multiplier is realized in TSMC CMOS0.18μmtechnology, with a maximum delay of900ps, static power consumption of19mWand a core area of11200μm2. The example demonstrates that CVNS designs can yield fast, low power arithmetic circuits using low noise analog circuitry.
Mitra Mirhassani, Majid Ahmadi, Graham A. Jullien
ISCAS3
2007 A CMOS Contact Imager for Cell Detection in Bio-Sensing Applications
abstract
Many experimental procedures in cell biology rely on the use of biochemical light-emitting reporters to enhance structures or processes of interest in a cell sample. The ability to integrate these widely-used protocols on the surface of a CMOS biosensor array, combined with the capacity to perform cell actuation/stimulation, would provide researchers with a valuable tool to conduct high-throughput, high-density, accurate analyses. To this end, a low noise, high-sensitivity CMOS contact imager is presented for the detection of cell cultures coupled directly to the sensor surface. Design considerations are discussed for the application of neural activity recording of neuronal networks culturedin vitrofrom dissociated neuron cells. A 128 × 128 CMOS imager implemented in 0.18-μm CMOS technology is presented featuring the implementation of the active reset technique with the use of the active column sensor pixel architecture.
Terence Tam, Graham A. Jullien, Orly Yadid-Pecht
ISCAS2
2007 Simulation of random cell displacements in QCA
abstract
We analyze the behavior of quantum-dot cellular automata (QCA) building blocks in the presence of random cell displacements. The QCA cells are modeled using the coherence vector description and simulated using QCADesigner. We evaluate various fundamental circuits: the wire, the inverter, the majority gate, and the two-wire crossing approaches: the coplanar crossover and the multilayer crossover. Our results show that different building blocks have different displacement tolerances. The coplanar crossover and inverter perform the weakest. The wire is the most robust. We have found displacement tolerances to be a function of circuit layout and geometry rather than cell size.
Gabriel Schulhof, Konrad Walus, Graham A. Jullien
ACM J. Emerg. Technol. Comput. Syst.3
2007 Multi-mode operator for SHA-2 hash functions
Ryan Glabb, Laurent Imbert, Graham A. Jullien, Arnaud Tisserand, Nicolas Veyrat-Charvillon
J. Syst. Archit.3
2006 Array Processing Using Alternate Arithmetic - A 20 Year Legacy
abstract
20 years ago, the first Systolic Array Workshop was held at the University of Oxford. This became an annual event with the name being changed to the Application Specific Array Processor (ASAP) Conference at the Princeton Workshop in 1990. Under either name, the conference highlights the implementation of special purpose computational processors, a basic feature of which is performing large numbers of arithmetic computations per second. In this paper we discuss representations of numbers and, in particular, the properties and advantages of arithmetic processors using these representations. In a retrospective, this paper looks at our own attempts, over the past 2 decades, to find new ways of representing, and computing with, numbers in order to achieve some advantages at the implementation level.
Graham A. Jullien
ASAP1
2006 Passive reduced-order macromodeling algorithm for structure dynamics in MEMS systems
abstract
In this paper, a novel passive reduced-order macro-modeling algorithm is proposed for structure dynamics in MEMS systems that are described by discrete models through the use of finite element methods (FEM). In the proposed scheme, the system of equations given by the FEM formulation are converted to state-space form such that the state-space equations are compatible with passive Krylov subspace methods based on congruent transformations. As an example, modeling, simulation, and experimental studies are presented for a Butterfly gyro developed at the Imego Institute. Our simulation results for this example are in very good agreement with previous publications using traditional modeling techniques
Rumi Zhang, Graham A. Jullien, Wei Wang 0003, Anestis Dounavis
ISCAS2
2006 A proposed hardware reference model for spatial transformation and quantization in H.264
Ihab Amer, Wael Badawy, Graham A. Jullien
J. Vis. Commun. Image Represent.3
2006 On the Use of Hash Functions as Preprocessing Algorithms to Detect Defects on Repeating Definite Textures
Ibrahim Cem Baykal, Graham A. Jullien
Mach. Vis. Appl.2
2006 Design Tools for an Emerging SoC Technology: Quantum-Dot Cellular Automata
abstract
The future of system-on-chip (SoC) technologies, based on the scaling of current FET-based integrated circuitry, is being predicted to reach fabrication limits by the year 2015. Economic limits may be reached before that time. Continued scaling of electronic devices to molecular scales will undoubtedly require a paradigm shift from the FET-based switch to an alternative mechanism of information representation and processing. This paradigm shift will also have to encompass the tools and design culture that have made the current SoC technology possible-the ability to design monolithic integrated circuits with many hundreds of millions of transistors. In this paper, we examine the initial development of a tool to automate the design of one of the promising emerging nanoelectronic technologies, quantum-dot cellular automata, which has been proposed as a computing paradigm based on single electron effects within quantum dots and molecules.
Konrad Walus, Graham A. Jullien
Proc. IEEE2
2005 Parallel Montgomery Multiplication in GF(2k) Using Trinomial Residue Arithmetic
abstract
We propose the first general multiplication algorithm in GF(2/sup k/) with a subquadratic area complexity of O(k/sup 8/5/) = O(k/sup 1.6/). Using the Chinese remainder theorem, we represent the elements of GF(2/sup k/); i.e. the polynomials in GF(2) [X] of degree at most k-1, by their remainder modulo a set of n pairwise prime trinomials, T/sub 1/,...,T/sub n/, of degree d and such that nd /spl ges/ k. Our algorithm is based on Montgomery's multiplication applied to the ring formed by the direct product of the trinomials.
Jean-Claude Bajard, Laurent Imbert, Graham A. Jullien
IEEE Symposium on Computer Arithmetic3
2005 Error-Free Computation of 8x8 2-D DCT and IDCT Using Two-Dimensional Algebraic Integer Quantization
abstract
This paper presents a novel error-free (infinite-precision) architecture for the fast implementation of both 8/spl times/8 2D discrete cosine transform and inverse DCT. The architecture uses a new algebraic integer quantization of a 1D radix-8 DCT that allows the separable computation of a 2D 8/spl times/8 DCT without any intermediate number representation conversions. This is a considerable improvement on previously introduced algebraic integer encoding techniques to compute both DCT and IDCT which eliminates the requirements to approximate the transformation matrix elements by obtaining their exact representations and hence mapping the transcendental functions without any errors. Using this encoding scheme, an entire 8/spl times/8 1D DCT-SQ (scalar quantization) algorithm can be implemented with only 24 adders. Apart from the multiplication-free nature, this new mapping scheme fits to this algorithm, eliminating any computational or quantization errors and resulting short-word-length and high-speed-design.
Khan Wahid, Vassil S. Dimitrov, Graham A. Jullien
IEEE Symposium on Computer Arithmetic3
2005 A Fault-Tolerant Modulus Replication Complex FIR Filter
abstract
In this paper we propose an architecture for the implementation of fault-tolerant computation for a high throughput multirate equalizer used in a 1 Gbps asymmetrical wireless LAN. Exploiting the algebraic structure of the modulus replication residue number system (MRRNS) minimizes the area overhead, and the area cost to correct a fault in a single computational channel is 82.7%. Generalized results for single error correction showing significant area savings are also presented.
Ian Steiner, Laurent Imbert, Graham A. Jullien, Vassil S. Dimitrov, Grant McGibney
ASAP4
2005 Simple 4-Bit Processor Based On Quantum-Dot Cellular Automata (QCA)
abstract
We describe the design and layout of a simple 4-bit processor based on quantum dot cellular automata (QCA) using the QCADesigner design tool. The processor design is based on an accumulator architecture which reduces the required hardware complexity and allows for reasonable simulation times. Our aim is to provide evidence that QCA has potential applications in future computers provided that the underlying technology is made feasible.
Konrad Walus, Mike Mazur, Gabriel Schulhof, Graham A. Jullien
ASAP4
2005 A high-performance hardware implementation of the H.264 simplified 8×8 transformation and quantization [video coding]
abstract
The recently approved digital video standard known as H.264 promises to be an excellent video format for use with a large range of applications. Real-time encoding/decoding is a main requirement for adoption of the standard to take place in the consumer marketplace. Transformation and quantization in H.264 are relatively less complex than their correspondences in other video standards. Nevertheless, for real-time operation, a speedup is required for such processes. Especially after the recent proposal to use an 8/spl times/8 integer approximation of discrete cosine transform (DCT) to give significant compression performance at standard definition (SD) and high definition (HD) resolutions. This paper discusses a high-performance hardware implementation of the H.264 simplified 8/spl times/8 transformation and quantization. The results show that the architecture satisfies the real-time constraints required by different digital video applications.
Ihab Amer, Wael Badawy, Graham A. Jullien
ICASSP (2)3
2005 Efficient Techniques for Binary-to-Multidigit Multidimensional Logarithmic Number System Conversion Using Range-Addressable Look-Up Tables
abstract
The multidimensional logarithmic number system (MDLNS), which has similar properties to the classical logarithmic number system (LNS), provides more degrees of freedom than the LNS by virtue of having two, or more, orthogonal bases and has the ability to use multiple MDLNS components, or digits. Unlike the LNS, there is no monotonic relationship between standard binary representations and MDLNS representations. Using look-up tables (LUTs) to perform the mapping function can be unrealistic for hardware implementations when large binary ranges or multiple digits are used. This work proposes a novel range-addressable technique for using look-up tables that allows efficient conversion from binary-to-single or multidigit MDLNS with varying accuracies, depending on the selected implementation.
Roberto Muscedere, Vassil S. Dimitrov, Graham A. Jullien, William C. Miller
IEEE Trans. Computers3
2005 A New Time Distributed DCT Architecture for MPEG-4 Hardware Reference Model
abstract
This paper presents the design of a new time distributed architecture (TDA) which outlines the architecture (ISO/IEC JTC1/SC29/WG11 MPEG2002/M8565) submitted to MPEG4 Part9 committee and included in the ISO/IEC JTC1/SC29/WG11 MPEG2002/9115N document. The proposed TDA optimizes the two-dimensional discrete cosine transform (2-D-DCT) architecture performance. It uses a time distribution mechanism to exploit the computational redundancy within the inner product computation module. The application specific requirements of input, output and coefficients word length are met by scheduling the input data. The coefficient matrix uses linear mappings to assign necessary computation to processor elements in both space and time domains. The performance analysis shows performance savings in excess of 96% as compared to the direct implementation and more than 71% as compared to other optimized application specific architectures for DCT.
Mehboob Alam, Wael Badawy, Graham A. Jullien
IEEE Trans. Circuits Syst. Video Technol.3
2004 Hardware prototyping for the H.264 4×4 transformation [video coding]
abstract
This paper presents a hardware prototype of the H.264 transformation. The proposed architecture uses only add and shift operations to reduce the computational requirements for the 4/spl times/4 transform. The architecture is developed to be used in high-resolution applications such as high definition television (HDTV) and digital cinema. The developed architecture is prototyped and simulated using ModelSim 5.4/spl reg/. It is synthesized using Leonardo Spectrum/spl reg/. The results show that the architecture satisfies the real-time constraints required by different digital video applications.
Ihab Amer, Wael Badawy, Graham A. Jullien
ICASSP (5)3
2004 A VLSI prototype for Hadamard transform with application to MPEG-4 part 10
abstract
This paper presents a VLSI prototype for the 2times2 Hadamard transform that is applied to the DC coefficients of the four 4times4 blocks of each chroma component as described in the MPEG-4 part 10 advanced video coding (AVC) standard. A VLSI prototype fir the quantization process that is accompanied with the transform operation is given as well. The implemented transform represents a level in the hierarchical transform adopted in the new AVC standard. The transform is computed using add operations only. This reduces the computational requirements of the design. The architecture is prototyped and simulated using ModelSim 5.4reg. It is synthesized using Leonardo Spectrumreg. The results show that the architecture satisfies the real-time constraints required by high definition television (HDTV)
Ihab Amer, Wael Badawy, Graham A. Jullien
ICME3
2003 Error-Free Arithmetic for Discrete Wavelet Transforms Using Algebraic Integers
abstract
A novel encoding scheme is introduced with applications to error-free computation of discrete wavelet transforms (DWT) based on Daubechies wavelets. The encoding scheme is based on an algebraic integer decomposition of the wavelet coefficients. This work is a continuation of our research into error-free computation of DCTs and IDCTs, and this extension is timely since the DWT is part of the new standard for JPEG2000. This encoding technique eliminates the requirements to approximate the transformation matrix elements by obtaining their exact representations. As a result, we achieve error-free calculations up to the final reconstruction step where we are free to choose an approximate substitution precision based on a hardware/accuracy trade-off.
Khan A. Wahid, Vassil S. Dimitrov, Graham A. Jullien
IEEE Symposium on Computer Arithmetic3
2002 A Novel Pipelined Threads Architecture for AES Encryption Algorithm
abstract
This paper presents a single-chip parallel architecture for advanced encryption standard (AES). The proposed architecture uses the thread approach, which integrates fully pipelined parallel units, that process 128 bits/cycle and quadruples the data throughput. The threads architecture allows a reduction of the clock rate by a factor of four, while maintaining the data throughput, and consumes less power. The prototype runs at a data rate of 7.68 Gbps on a Xilinx xc2V1500 Virtex-II FPGA. The data rate shows that the proposed thread approach produces one of the fastest single-chip FPGA implementations currently available. In addition, the proposed architecture is scalable to 192, 256 and higher bits.
Mehboob Alam, Wael Badawy, Graham A. Jullien
ASAP3
2002 Efficient Conversion From Binary to Multi-Digit Multi-Dimensional Logarithmic Number Systems Using Arrays of Range Addressable Look-Up Tables
abstract
The multi-dimensional logarithmic number system (MDLNS), with similar properties to the logarithmic number system (LNS), provides more degrees of freedom than the LNS by virtue of having two orthogonal bases and the ability to use multiple digits. Unlike the LNS, there is no direct functional relationship between binary/floating point representation and the MDLNS representation. Traditionally look-up tables (LUTs) were used to move from the binary domain to the MDLNS domain. This method can be unrealistic for hardware implementation when large binary ranges or multiple digits are used. This paper introduces a range addressable technique for table look-up arrays that allows efficient conversion from binary to single or multi-digit MDLNS.
Roberto Muscedere, Vassil S. Dimitrov, Graham A. Jullien, William C. Miller
ASAP3
2002 A Number System with Continuous Valued Digits and Modulo Arithmetic
abstract
This paper presents a novel number system based on signed continuous valued digits. Arithmetic operations in this number system are performed using simple analog circuitry, in contrast to the conventional implementation of arithmetic units by Boolean or multiple-valued logic circuits. Unlike the limited precision offered by classical analog arithmetic circuits, the ensemble of continuous valued digits that comprises a number in this system provides arbitrary implementation precision with standard analog circuitry. The number system also provides almost-carry-free arithmetic structures with digit level redundancy. In this paper, we introduce the mathematical foundations for positive and negative numbers, addition, multiplication, redundancy, radix conversions, and also the digit value integrity for circuit implementations. Potential applications are in the area of low noise and low cross-talk circuitry for arithmetic circuits used in mixed-signal systems.
Aryan Saed, Majid Ahmadi, Graham A. Jullien
IEEE Trans. Computers3
2001 The Use of the Multi-Dimensional Logarithmic Number System in DSP Applications
abstract
A recently introduced double-base number representation has proved to be successful in improving the performance of several algorithms in cryptography and digital signal processing. The index-calculus version of this number system can be regarded as a two-dimensional extension of the classical logarithmic number system. This paper builds on previous special results by generalizing the number system both in multiple dimensions (multiple bases) and by the use of multiple digits. Adopting both generalizations the paper shows that large reductions in hardware complexity are achievable compared to an equivalent precision logarithmic number system.
Vassil S. Dimitrov, Jonathan Eskritt, Laurent Imbert, Graham A. Jullien, William C. Miller
IEEE Symposium on Computer Arithmetic4
2000 A New Algorithm for the Elimination of Common Subexpressions in Hardware Implementation of Digital Filters by Using Genetic Programming
abstract
A new algorithm based on Genetic Programming (GP) for the problem of optimization of Multiple Constant Multiplication (MCM) by Common Subexpression Elimination (CSE) is developed. This method is used for hardware optimization of DSP systems. A solution based on GP is shown in this paper. The performance of the technique is demonstrated in one- and multi-dimensional digital filters with constant coefficients.
H. Safiri, Majid Ahmadi, Graham A. Jullien, William C. Miller
ASAP3
2000 A MEMS micromagnetic actuator for use in a bionic interface
abstract
The design of a microelectromechanical (MEMS) device that forms part of a micro acousto-magnetic transducer for use with a hearing instrument is described in this paper. A MEMS realization of a microelectromagnetic actuator is used to generate a magnetic field that exerts a force on a permanent micromagnet that has been implanted on the round window of the cochlea. The motion of the implanted magnet will develop traveling waves on the basilar membrane inside the cochlea to give a hearing capability. A modular realization of the micro electromagnet has been proposed that offers the advantage that the magnet's magnetomotive force performance characteristics can be easily changed by depositing additional layers of the modular realization for the magnet core segments and their associated section of excitation winding. The MEMS structures have been designed and simulated using the IntelliSuite MEMS software package from the IntelliSense Corporation.
Sazzadur Chowdhury, Graham A. Jullien, Majid Ahmadi, William C. Miller
ISCAS2
2000 Defect detection in web inspection using fuzzy fusion of texture features
abstract
This paper describes a novel approach for real-time defect detection targeted to web inspection systems using real-time processing of raster scanned images. Our new method utilizes a fuzzy fusion of texture features to detect image rows that contain defects, without requiring access to the two dimensional frame data. The method is suitable for in-camera processing where an FPGA is directly connected to the digitized video stream. This method is successfully implemented in the limited resources of one FPGA without the use of a frame-grabber.
S. Hossain Hajimowlana, Roberto Muscedere, Graham A. Jullien, James W. Roberts
ISCAS3
2000 Complexity and Fast Algorithms for Multiexponentiations
abstract
In this paper, we propose new algorithms for multiple modular exponentiation operations. The major aim of these algorithms is to speed up the performance of some cryptographic protocols based on multiexponentiation. Our new algorithms are based on binary-like complex arithmetic, introduced by K. Pekmestzi (1989) and generalized in this paper.
Vassil S. Dimitrov, Graham A. Jullien, William C. Miller
IEEE Trans. Computers2
1999 Arithmetic with Signed Analog Digits
abstract
This paper presents mathematical foundations of the Overlap Resolution Number System (ORNS) which is based on signed continuous valued digits (CVDs). ORNS is a redundant number system employing residue arithmetic. In contrast to the implementation of arithmetic by binary or multiple-valued logic circuits, arithmetic operations in this novel number system are performed by analog digit manipulation circuitry. The redundancy in an ensemble of continuous valued digits that comprises a number provides tolerance to implementation imprecisions. Processing with these analog digits is performed by carry-free arithmetic structures with systematic circuit level redundancy.
Aryan Saed, Majid Ahmadi, Graham A. Jullien
IEEE Symposium on Computer Arithmetic3
1999 Theory and Applications of the Double-Base Number System
abstract
In this paper, we analyze some of the main properties of a double base number system, using bases 2 and 3; in particular, we emphasize the sparseness of the representation. A simple geometric interpretation allows an efficient implementation of the basic arithmetic operations and we introduce an index calculus for logarithmic-like arithmetic with considerable hardware reductions in lookup table size. We discuss the application of this number system in the area of digital signal processing; we illustrate the discussion with examples of finite impulse response filtering.
Vassil S. Dimitrov, Graham A. Jullien, William C. Miller
IEEE Trans. Computers2
1998 Digital Arithmetic Using Analog Arrays
abstract
This paper describes techniques for using locally connected analog cellular neural networks (CNNs) to implement digital arithmetic arrays; the arithmetic is implemented using a recently disclosed Double-Base Number System (DBNS). The CNN arrays are targeted for low power low-noise DSP applications where lower slew rate during transitions is a potential advantage. Specifically, we demonstrate that a CNN array, using a simple nonlinear feedback template, with hysteresis, can perform arbitrary length arithmetic with good performance in terms of stability and robustness. The principles presented in this paper can also be used to implement arithmetic in other number systems such as the binary number system.
Saeid Sadeghi-Emamchaie, Graham A. Jullien, Vassil S. Dimitrov, William C. Miller
Great Lakes Symposium on VLSI2
1998 A new DCT algorithm based on encoding algebraic integers
abstract
In this paper we introduce an algebraic integer encoding scheme for the basis matrix elements of 8/spl times/8 DCTs and IDCTs. In particular, we encode the function cos(/spl pi//16) and generate the other matrix elements using standard trigonometric identities. This encoding technique eliminates the requirement to approximate the matrix elements; rather we use algebraic 'placeholders' for them. Using this encoding scheme we are able to produce a multiplication free implementation of the Feig-Winograd algorithm.
Vassil S. Dimitrov, Graham A. Jullien, William C. Miller
ICASSP2
1998 An Algorithm for Modular Exponentiation
Vassil S. Dimitrov, Graham A. Jullien, William C. Miller
Inf. Process. Lett.2
1997 Theory and applications for a double-base number system
abstract
Presents a rigorous theoretical analysis of the main properties of a double-base number system, using bases 2 and 3. In particular, we emphasize the sparseness of the representation. A simple geometric interpretation allows an efficient implementation of the basic arithmetic operations, and we introduce an index calculus for logarithmic-like arithmetic with considerable hardware reductions in look-up table size. Two potential areas of applications are discussed: applications in digital signal processing for computation of inner products and in cryptography for computation of modular exponentiations.
Vassil S. Dimitrov, Graham A. Jullien, William C. Miller
IEEE Symposium on Computer Arithmetic2
1997 Algorithms for Multi-Exponentiation Based on Complex Arithmetic
abstract
In this paper, we propose new algorithms for multiple modular exponentiation operations. The major aim of these algorithms is to speed up the performance of some cryptographic protocols based on multi-exponentiation. The algorithms proposed are based on binary-like complex arithmetic, introduced by K. Pekmestzi (1989) and generalized in this paper.
Vassil S. Dimitrov, Graham A. Jullien, William C. Miller
IEEE Symposium on Computer Arithmetic2
1996 Design and VLSI Implementation of a Unified Synapse-Neuron Architecture
abstract
We describe the design and VLSI implementation of a unified synapse-neuron architecture for multi-layer neural networks. A new hybrid building block proposed for this purpose is formed by integrating a partial S-shape neural nonlinearity within a Multiplying DAC synapse. MDAC synapse contains modifications to simplify sign-bit circuit. Small analog circuits generate a distributed S-shape neural function by combining quadratic characteristics of four MOS transistors. The proposed modular neural network architecture features design simplicity and scalability, area efficiency, reduced interconnection problem, improved robustness and digital programmability. Based on the proposed scheme, we have considerably increased the synaptic density in the improved version of a programmable optically-coupled neural network.
Hormoz Djahanshahi, Majid Ahmadi, Graham A. Jullien, William C. Miller
Great Lakes Symposium on VLSI3
1996 VLSI Neural System Architecture for Finite Ring Recursive Reduction
David Zhang 0001, Graham A. Jullien
Int. J. Neural Syst.2
1996 On computing Chebyshev optimal nonuniform interpolation
Zhongde Wang, Graham A. Jullien, William C. Miller
Signal Process.2
1995 An array processor for inner product computations using a Fermat number ALU
abstract
This paper explores an architecture for parallel independent computations of inner products over the direct product ring /spl Rfr//sub 257/spl times/17/. The structure is based on the polynomial mapping of the Modulus Replication RNS for calculations over dynamic ranges much larger than the product of the computational moduli. We show that the computational ring is optimal for our purposes, and introduce basic cells for the efficient calculation of all elements of the polynomial ring computations.
Wenzhe Luo, Graham A. Jullien, Neil M. Wigley, William C. Miller, Zhongde Wang
ASAP2
1995 VLSI implementation of high throughput DSP using finite ring arithmetic
abstract
Reviews strategies in implementing DSP systems using residue computations; in particular the authors highlight work underway in the VLSI Research Group, at the University of Windsor, in the area of high throughput DSP systems on silicon. The paper reviews a series of issues including mapping techniques, fault tolerant architectures, and area and power efficient VLSI implementation procedures. CAD tools, developed for automating the design and layout of residue DSP systems on silicon, are also discussed. An extensive bibliography is provided for further reading.
Graham A. Jullien
ICASSP1
1995 A New Design Technique for Column Compression Multipliers
abstract
In this paper, a new design technique for column-compression (CC) multipliers is presented. Constraints for column compression with full and half adders are analyzed and, under these constraints, considerable flexibility for implementation of the CC multiplier, including the allocation of adders, and choosing the length of the final fast adder, is exploited. Using the example of an 8/spl times/8 bit CC multiplier, we show that architectures obtained from this new design technique are more area efficient, and have shorter interconnections than the classical Dadda CC multiplier. We finally show that our new technique is also suitable for the design of twos complement multipliers.>
Zhongde Wang, Graham A. Jullien, William C. Miller
IEEE Trans. Computers2
1994 A regular recursive algorithm for the discrete sine transform
abstract
We derive a new recursive algorithm for the discrete sine transform (DST) which possesses a very regular structure. The multiplication coefficients in our algorithm can be generated by a simple recursion without the requirement for trigonometric functions; also, no shifts of data are required. In comparison, the existing recursive algorithm for the DST proposed by Wang (1990) has an irregular structure and requires considerable data shifts.>
Zhongde Wang, Graham A. Jullien, William C. Miller
ICASSP (3)2
1994 Area-Time Analysis of Carry Lookahead Adders Using Enhanced Multiple Output Domino Logic
abstract
In order to improve the area and speed of the design of carry lookahead adders (CLAs) using enhanced multiple output domino logic (EMODL), we investigate the trade-off between the number of cascaded gate stages and the gate fan-in of each stage by varying these factors in four different architectural structures for a 32-bit CLA implemented in 1.2 micron CMOS technology. HSPICE simulation results show that the number of cascaded stages is a more critical factor than the gate fan-in.>
June Wang, Zhongde Wang, Graham A. Jullien, William C. Miller
ISCAS3
1994 Current Input TSPC Latch for High Speed, Complex Switching Trees
abstract
This paper discusses new techniques for obtaining high clock rates with complex n-blocks in True-Single-Phase dynamic latch structures. In this paper we present new dynamic current steering latch structures, and apply them to both CMOS and BiCMOS technologies. In the latter case, we exploit the superior properties of the available bipolar devices to achieve substantial speed increases. The latching technique allows complex n-FET blocks (fan-in between 10 and 20) to be used with the TSPC latch at high data rates (over 150 MHz for a 1.2 /spl mu/ CMOS process). The n-FET block is built as a minimized binary tree, which we have termed a switching tree, and interpreted as a general look-up table for use in a variety of bit-level systolic array processors.>
J. C. Czilli, Graham A. Jullien, William C. Miller
ISCAS3
1994 The generalized discrete W transform and its application to interpolation
Zhongde Wang, Graham A. Jullien, William C. Miller
Signal Process.2
1994 Recursive algorithms for the forward and inverse discrete cosine transform with arbitrary length
abstract
The authors first demonstrate that the forward and inverse discrete cosine transform (DCT, IDCT) can be represented by Chebyshev polynomials of the third and second kind, respectively. Then, they derive recursive algorithms for the DCT and IDCT with arbitrary length from the recursive formulae for the Chebyshev polynomials. The proposed algorithms are particularly suitable for VLSI implementation using array processing architectures.>
Zhongde Wang, Graham A. Jullien, William C. Miller
IEEE Signal Process. Lett.2
1994 Large Dynamic Range Computations over Small Finite Rings
abstract
Presents a new multivariate mapping strategy for the recently introduced Modulus Replication Residue Number System (MRRNS). This mapping allows computation over a large dynamic range using replications of extremely small rings. The technique maintains the useful features of the MRRNS, namely: ease of input coding; absence of a Chinese Remainder Theorem inverse mapping across the full dynamic range; replication of identical rings; and natural integration of complex data processing. The concepts are illustrated by a specific example of complex inner product processing associated with a radix-4 decimation in time fast Fourier transform algorithm. A complete quantization analysis is performed and an efficient scaling strategy chosen based on the analysis. The example processor uses replications of three rings: modulo-3, -5, and -7; the effective dynamic range is in excess of 32 b. The paper also includes very-large-scale-integration implementation strategies for the processor architecture that consists of arrays of massively parallel linear bit-level pipelines.>
Neil M. Wigley, Graham A. Jullien, Daniel Reaume
IEEE Trans. Computers2
1993 Integer mapping architectures for the polynomial ring engine
abstract
A finite polynomial ring structure for mapping inner product computations to parallel independent ring computations over 3-b moduli has been introduced by N.M. Wigley et al. (1992). The main algorithmic computation architecture can be implemented using well-established systolic array mapping principles, and a project to construct a Polynomial Ring Engine (PRE) is underway to exploit the VLSI implementation properties of such computations. A semi-systolic architecture for the input and output conversion mappings that are required in the engine is introduced here. It is shown that the entire mappings procedure can be carried out with pipelined six-input logic blocks and small, fast, binary adders. CMOS implementation techniques for the pipelined blocks are discussed, and the design procedure is illustrated with results from a recently completed module generator.>
Sami S. Bizzan, Graham A. Jullien, Neil M. Wigley, William C. Miller
IEEE Symposium on Computer Arithmetic2
1993 New Concepts for the Design of Carry Lookahaead Adders
Zhongde Wang, Graham A. Jullien, William C. Miller, June Wang
ISCAS2
1993 Pipelined analog multi-layer feedforward neural networks
Navid Yazdi, Majid Ahmadi, Graham A. Jullien, Malayappan Shridhar
ISCAS3
1993 VLSI implementations of number theoretic techniques in signal processing
Graham A. Jullien, Neil M. Wigley, William C. Miller
Integr.1
1991 Small moduli replications in the MRRNS
abstract
The authors describe mapping, scaling, and conversion processes using a new mapping strategy for the modulus replication residue number system (MRRNS). The strategy allows direct mapping of bits of either a purely real or multiplexed bit coded complex number to a set of independent rings, defined by moduli 3, 5, and 7. The MRRNS technique is superior to a large QRNS system operating with a computational dynamic range of over 27 b. A classical radix-4 implementation of a 1024 FFT is used for the comparison. The scaling and conversion procedure is shown to be a set of finite ring calculations followed by an array of ordinary binary adders. The VLSI implementation of the most complex finite ring circuit required (a Mod 7 multiplier) is shown to be easily implemented using the switching tree approach, and mask extracted simulations at 50 MHz demonstrate the embedding of the switching trees in a dynamic pipeline/evaluate circuit with restoring latch.>
Neil M. Wigley, Graham A. Jullien, Daniel Reaume, William C. Miller
IEEE Symposium on Computer Arithmetic2
1991 Arithmetic for digital neural networks
abstract
The implementation of large input digital neurons using designs based on parallel counters is described. The implementation of the design uses a two-cell library, in which each cell is implemented using switching trees which are pipelined binary trees of n-channel transistors. Results obtained from initial switching trees realized with a 3- mu m CMOS process indicate that the design is capable of being pipelined at 40 MHz sample rates, with better performance expected for more advanced technologies. It appears feasible to develop a wafer-scale implementation with 2000 neurons (each with 1000 inputs) that would perform 3*10/sup 12/ additions/s.>
David Zhang 0001, Graham A. Jullien, William C. Miller, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
1990 Array processing on finite polynomial rings
abstract
The disadvantage of computations using finite rings is the need to compute over many different rings in order to produce useful dynamic ranges of computation. By mapping integers into polynomial rings, one can replace the different rings by the replication of the same ring with considerable computational advantages. The authors present the methodology of such a mapping strategy, and discuss the application of the features of the technique to dense fabrication strategies such as WSI and ULSI, where redundancy is an integral part of the architecture. It is shown that all of the operations can be performed by the recently introduced bit-steered ROM technique, with attendant advantages of easy testability and fault detection. The polynomial mapping strategy allows the extra advantage of low-overhead redundancy when general computational blocks are employed.>
Neil M. Wigley, Graham A. Jullien
ASAP2
1990 On Modulus Replication for Residue Arithmetic Computations of Complex Inner Products
abstract
A technique is presented for coding weighted magnitude components (e.g. bits) of numbers directly into polynomial residue rings, such that repeated use may be made of the same set of moduli to effectively increase the dynamic range of the computation. This effectively limits the requirement for large sets of relatively prime moduli, For practical computations over quadratic residue rings, at least 6-bit moduli have to be considered. It is shown that 5-bit moduli can be effectively used for large dynamic range computations.>
Neil M. Wigley, Graham A. Jullien
IEEE Trans. Computers2
1989 VLSI implementation of a digital image threshold selection architecture
P. K. Lim, Maher A. Sid-Ahmed, Graham A. Jullien
Integr.3
1988 High-speed signal processing using systolic arrays over finite rings
abstract
A modular architecture for very fast digital signal processing (DSP) elements are presented. The computation is performed over finite rings (or fields) and is able to emulate processing over the integer ring using residue number systems. The computations are restricted to closed operations (ring or field binary operators) with the ability to perform limited scaling operations. Computations naturally defined over finite mathematical systems are also easily implemented using this approach. The technique evolves from the decomposition of each closed calculation using the ring/field associativity property. Linear systolic arrays, formed with multiple elements, each of a single generic form, are used for all calculations. The pipeline cycle is determined from the generic cell and is predicted to be very fast by a critical path analysis. The cells are matched to the VLSI medium, and the resulting array structures are very dense. Examples of DSP applications are given to illustrate the technique, and example cell and array VLSI layouts are presented for a 3- mu m CMOS process.>
Majid Taheri, Graham A. Jullien, William C. Miller
IEEE J. Sel. Areas Commun.2
1987 A signal processing cell architecture
abstract
A data flow general purpose digital signal processor has been previously developed [1] for real time applications of digital signal processing. The Data Flow Signal Processor (DFSP) is attached to a host computer, and is based on a binary tree structure. It employs two types of cells: processing and arithmetic cells, and utilizes residue number system [2] for arithmetic operations. The objective of this work is to develop architecture of the processing cell [3]. This processing cell is simulated on a VAX 11/785 computer system utilizing A Hardware Programming Language (AHPL) [4]. Simulation results shows that processing cell architecture is valid and DFSP is capable of high throughput rates. This paper describes structure and operation of the processing cell.
Mohsin M. Jamali, M. M. Hussain, Graham A. Jullien
ICASSP3
1987 Implementation of the generalized FIR filter structure using the residue arithmetic
abstract
Very recently, the Quadratic Residue System (QRNS) has been introduced [3,4,5]. Using the QRNS complex multiplication can be performed with two base field multiplication and zero additions. The primary restriction is the limited form of the moduli set for RNS operations. The QRNS has since been geralized for any type of moduli set with an increase in multiplication from 2 to 3 and the resulting number system has been termed Modified Quadratic Residue Number System (MQRNS) [1,2]. In [9] a recursive FIR filter has been developed using the Complex Number Theoretic z-transform (CNT z-transform). Recently, in [6], the implementation of this recursive FIR filter structure has been presented using the QRNS and the MQRNS. Extension of this implementation to generalized FIR filter (Lagrange) has also been briefly presented in [6]. In this paper, we consolidate the implementation aspects of the generalized FIR filter using the MQRNS and also prove that the QRNS is not a suitable medium for the implementation.
Ramasamy Krishnan, Graham A. Jullien, William C. Miller
ICASSP2
1987 VLSI Modular architectures for complex digital signal processing applications
abstract
Recently, the Quadratic Residue Number System (QRNS)[3,4] and Modified Quadratic Residue Number System (MQRNS)[1,2] have been introduced to perform complex multiplications efficiently. The growing number of complex digital signal processing applications will be implemented more efficiently and economically by using Very Large Scale Integration (VLSI) technology. In this paper we discuss VLSI implementation of complex multiplication using the QRNS and MQRNS. We also concentrate on the aspects of VLSI implementation of Finite Impulse Response (FIR) filter architectures.
Ramasamy Krishnan, Graham A. Jullien, William C. Miller
ICASSP2
1987 Systolic ROM arrays for implementing RNS FIR filters
abstract
The Residue Number System (RNS), its concept, computational power, and applications have been investigated in the past [1,6,7]. Most of the studies have resulted in realizations suitable for discrete implementation [2,8]. This paper introduces a linear systolic array architecture for an RNS based FIR filter suitable for VLSI fabrication. The array, which is completely pipelined, consists of modular cells which only communicate to their nearest neighbor. The connected cells constitute a linear systolic array, and the construction of the cell is such that it can be programmed to function in many different DSP tasks. The final result is the construction of a linear systolic ROM that effectively replaces the previously used discrete ROM arrays.
Majid Taheri, Graham A. Jullien, William C. Miller
ICASSP2
1986 A VLSI array for computing the DFT based on RNS
abstract
The Discrete Fourier Transform (DFT) has been adopted in a wide spectrum of Digital Signal Processing (DSP) applications due to the advances in VLSI technology, One dimensional systolic arrays are employed to implement the DFT algorithms where N DFT points can be computed in O(N) time using O(N) area. Residue Number System (RNS) is used to achieve parallelism on the mathematical level, as the arithmetic operations are performed independently for each modulus. Modularity has been realized on both functional and layout levels. Two types of arrays are described. The first array offers higher speed performance, while the second requires less area and is more general. The proposed structures are based on bit parallel processing and lend themselves to pipelining.
Magdy A. Bayoumi, Graham A. Jullien, William C. Miller
ICASSP2
1986 Computation of complex number theoretic transforms using quadratic residue number systems
abstract
Very recently, the Quadratic Residue Number System (QRNS) has been introduced [4,5]. The QRNS is obtained from a mapping of Gaussian integers over a finite ring to a ring of conjugate elements. The conjugate ring has the remarkable property that both addition and multiplication are performed component-wise, therefore complex multiplication only requires two base field multiplications and zero additions. The operations are performed over sub-rings, isomorphic to the conjugate ring via the Chinese Remainder Theorem isomorphism. The primary restriction is the limited form of the moduli set for RNS computations. The QRNS has since been generalized for any type of moduli set with an increase in multiplications from 2 to 3 and the resulting number system has been termed the Modified Quadratic Residue Number System (MQRNS) [1,2]. The direct FIR filter architecture and bit-slice architecture for FIR and recursive digital filters have, been presented using the QRNS and MQRNS [4]. In this paper, the computation of the Complex Number Theoretic Transform(CNTT) and the hardware implementation of a radix-2 butterfly structure, using high-density ROM arrays, are presented. This paper shows that both theQRNS and MQRNS require almost the same amount of hardware for the implementation of the butterfly structure. The computation of Cyclic Convolution in both the QRNS and MQRNS is also discussed.
Ramasamy Krishnan, Graham A. Jullien, William C. Miller
ICASSP2
1986 The implementation of the generalized Lagrange FIR filter structure defined over finite fields or rings
abstract
This paper discusses the use of the Complex Number Theoretic z-transform in implementing a recursive FIR filter structure for frequency samples spaced around the unit circle in the complex number theoretic z-domain. The complex arithmetic operations have been implemented using the Quadratic Residue Number System (QRNS) and Modified Quadratic Residue Number System (MQRNS) for uniformly spaced frequency samples around the unit circle. We discuss the extension of this technique to non-uniformly spaced samples around the unit circle and the resulting filter structure has been termed the generalized number theoretic FIR filter structure. We demonstrate that the MQRNS is the only suitable tool in implementing this recursive FIR filter structure for non-uniformly spaced frequency samples.
Ramasamy Krishnan, Graham A. Jullien, William C. Miller
ICASSP2
1985 A VLSI implementation of an FFT/NTT computational unit
abstract
The coupling of Residue Number System (RNS) with the recent advances in VLSI technology leads to an efficient implementation of many digital signal processing algorithms. This paper discusses modularity in implementing RNS systems, as modularity is considered an important criterion for VLSI design. An NTT/FFT computational unit is implemented using two multi-look-up table modules as building block units. The layout can be optimized using a look-up table layout procedure which supports the custom design approach. The modularity has been achieved on both functional and layout levels where the interconnection area is minimum.
Magdy A. Bayoumi, Graham A. Jullien, William C. Miller
ICASSP2
1985 An efficient VLSI adder for DSP architectures based on RNS
abstract
The implementation of Residue Number System (RNS) architectures using the VLSI technology is discussed. An example of implementing an RNS adder is presented in this paper. Two approaches; the look-up table and the binary adder, have been analyzed in the scope of VLSI criteria where the performance measures are area and time. Two models have been developed, they are flexible, support any modulus, and they provide custom design capabilities. Within the context of this paper, it has been found that the look-up table approach is superior in both area and time up to 5 bits, while the binary adder approach offers better performance for larger moduli.
Magdy A. Bayoumi, Graham A. Jullien, William C. Miller
ICASSP2
1985 Software techniques for programming a general purpose data flow signal processor
abstract
A real time general purpose signal processor architecture has been developed previously [1,2]. This architecture utilizes parallel, pipeline and distributed processing approaches to achieve high speed computation. Software techniques for programming the data flow signal processor are presented since conventional programming languages are not suitable for programming fast parallel machines. Data flow graphs (DFG) are used to develop an interactive programming environment which will shield the programmer from the internal structure of the data flow signal processor (DFSP). The programming of the DFSP is demonstrated with an image processing application. This example illustrates that the direct convolution can be used to perform computations at video rates.
Mohsin M. Jamali, Graham A. Jullien, William C. Miller, S. I. Ahmad
ICASSP2
1985 Complex digital signal processing using quadratic residue number systems
abstract
Recently, the Quadratic Residue Number System (QRNS) has been introduced [4,5,6], which allows the multiplication of complex integers with two real multiplications. Restrictions on the form of the moduli can be removed if an increase in real multiplications from two to three can be tolerated; the resulting number system has been termed the Modified Quadratic Residue Number System (MQRNS). In this paper the MQRNS is defined, and residue to binary conversion techniques in both the QRNS and MQRNS are presented. Hardware implementations of non-recursive and recursive digital filters are also presented where the QRNS and MQRNS structures are realized using a bit-slice architectures.
Ramasamy Krishnan, Graham A. Jullien, William C. Miller
ICASSP2
1984 A real time general purpose signal processor
abstract
A design of real time general purpose signal processor architecture is proposed in this paper. The processor is based upon a binary tree structure utilizing multiprocessing, pipeline and distributed processing techniques. A host computer distributes the individual tasks to each processor to perform parallel operations. The residue number system is used for carry free arithmetic operations stored in RAM's and to achieve smaller packet size, eliminating serial transmission of packets as proposed in other data flow machines. The processor is programmable and capable or performing real time signal processing operations.
Mohsin M. Jamali, Graham A. Jullien, William C. Miller, S. I. Ahmad
ICASSP2
1984 A VLSI model for residue number system architectures
Magdy A. Bayoumi, Graham A. Jullien, William C. Miller
Integr.2
1983 Models for VLSI implementation of residue number system arithmetic modules
abstract
This paper discusses the implementation of RNS arithmetic modules using VLSI technology. The modules are based on the interconnection of read-only memory look-up tables. The paper first outlines a memory model for a single look-up table which allows the selection of the most efficient layout for memories which do not have power of 2 dimensions. The paper then discusses various examples of interconnected memory modules with associated optimizing layout algorithms. Finally, an example is given of the application of one of the modules to a large prime modulus multiplier.
Magdy A. Bayoumi, Graham A. Jullien, William C. Miller
IEEE Symposium on Computer Arithmetic2
1983 An area-time efficient NMOS adder
Magdy A. Bayoumi, Graham A. Jullien, William C. Miller
Integr.2
1983 Processor Architectures for Two-Dimensional Convolvers Using a Single Multiplexed Computational Element with Finite Field Arithmetic
abstract
This paper describes the theory, simulation, and construction of a two-dimensional number theoretic transform (NTT) convolver. The convolver performs indirect convolution by using the cyclic convolution property of a class of generalized discrete Fourier transforms (DFT's) defined over rings isomorphic to direct sums of Galois fields. The paper first presents the theoretical development of the computational element required for computing the generalized discrete Fourier transform (GDFIT) and its inverse. The theory extends the use of base fields to second degree extension fields and provides efficient choices for transform parameters to minimize hardware. The paper next presents results of recent work in multidimensional transform memory structures, and extends this work to the complete convolution process. The two theories are then "married" to produce efficient, very high speed convolution architectures. Simulation results are presented for a second degree extension field image convolver and constructional details are presented for a fast image convolver using 2 base fields and designed to operate as a peripheral to a fast 32 bit minicomputer.
Hari K. Nagpal, Graham A. Jullien, William C. Miller
IEEE Trans. Computers2
1982 Quantization error and limit cycles analysis in residue number system coded recursive filters
abstract
The paper discusses the Residue Number System (RNS) implementation of second order recursive digital filter sections. The RNS offers the advantage of using integer based arithmetic operations and a simple hardware realization involving arrays of look up tables stored in high density ROMs, In residue number system scaling is necessary to keep the data within the limited dynamic range. The effects of residue scaling on the performance of second order sections are described here. An analysis of the quantization error accumulation associated with scaling in the proposed realizations shows that the residue architecture combines good quantization error performance with high speed implementation. The conditions for existance of zero input limit cycles in residue coded recursive sections is derived for certain cases. Regions on the coefficient space which allow limit cycles of period 1 (dc) and period 2 are determined and the bounds on their magnitude are derived.
Anna Z. Baraniecki, Graham A. Jullien
ICASSP2
1982 Memory architecture of a video-rate image convolver
abstract
This paper describes the design of an image convolver suitable for the convolution of a (256 × 256) image with large filter impulse responses in under 1/30 of a second. The memory architecture of the convolver is based on the structures associated with parallel and pipelined FFT/NTT processors that use ROM oriented implementation of the Residue Number System and a 2 D radix-r Ordered-input, Ordered-output FFT algorithm. The 2 D convolution operation is performed by the use of overlap-save technique of sectioned convolutions.
Hari K. Nagpal, Graham A. Jullien, William C. Miller
ICASSP2
1981 A two-dimensional finite field processor for image filtering
abstract
The paper describes the design of a high speed two-dimensional digital convolver for use as a preprocessor element in a quality control inspection system. The preprocessor architecture is based on the concepts and structures associates with parallel finite field transforms that have been implemented with read-only-memory arrays. The transform is computed over two second-degree extension Galois Fields to achieve a dynamic range for the multidimensional convolver that is in excess of sixteen bits. A dedicated memory structure that can support a multiplexed 2D NTT butterfly is described. Image processing results obtained by simulating the finite field processor are presented.
Graham A. Jullien, William C. Miller
ICASSP1
1980 A hardware realization of an NTT convolver using ROM arrays
abstract
This paper describes the construction of a digital signal convolver. The convolution is performed using a Number Theoretic Transform computed over a ring which is isomorphic to three extension fields of second degree:R \simeq GF(191^{2}) \oplus GF(193^{2}) \oplus GF(449^{2}). The transform is implemented using arrays of latched read-only-memories to provide a high throughput computational element. Special procedures are used to reduce ROM size by making use of indices and sub-modular addition techniques. Memory structures are described that allow two records to be convolved at the same time.
Graham A. Jullien, William C. Miller
ICASSP1
1980 Implementation of Multiplication, Modulo a Prime Number, with Applications to Number Theoretic Transforms
abstract
This paper discusses a technique for multiplying numbers, modulo a prime number, using look-up tables stored in read-only memories. The application is in the computation of number theoretic transforms implemented in a ring which is isomorphic to a direct sum of several Galois fields, parallel computations being performed in each field.
Graham A. Jullien
IEEE Trans. Computers1
1979 Hardware implementation of convolution using number theoretic transforms
abstract
Using the Residue Number System (RNS) as a basis for the hardware construction, two different hard-ware structures are discussed for implementing NTTs over a direct sum of Galois Fields GF(mi2). The first structure uses arrays of read only memories and the second uses arrays of microprocessors; in particular, single chip microprocessors are proposed. The two techniques offer trade offs between speed of operation, and cost. Using the RNS, rather than conventional binary arithmetic, allows more flexibility in the choice of the generator, α, and consequently more flexibility in allowable transform parameters. A selection of parameters, for the two realizations, are discussed when the Galois Field of mi2elements is a finite field of Gaussian or quadratic integers.
A. Baraniecka, Graham A. Jullien
ICASSP2
1979 Implementation of FFT Structures Using the Residue Number System
abstract
This paper considers the implementation of a fast Fourier transform (FFT) structure using arrays of read-only memories. The arithmetic operations are based entirely on the residue number system. The most important aspect of the structure relates to the scaling arrays, which are required to prevent overflow. Because of the limitations of the number system, scaling factors have to be chosen on an a priori basis. This paper develops optimum procedures for choosing both scaling factors and the position of scaling arrays in the structure. Some examples are presented relating to the filtering of speech via a convolutional filter structure.
Ben-Dau Tseng, Graham A. Jullien, William C. Miller
IEEE Trans. Computers2
1978 Application of the residue number system to computer processing of digital signals
abstract
The residue number system offers parallel processing, digital hardware, implementations for the binary operations of addition, subtraction and multiplication. This paper discusses the use of the residue number system in implementing digital signal processing functions, in which these binary operations abound. The paper covers implementations using arrays of read only memories, and briefly discusses the use of parallel microprocessor structures. ROM array implementations of scaling operations are also presented.
Graham A. Jullien, William C. Miller
IEEE Symposium on Computer Arithmetic1
1978 Recursive digital filters in image processing
abstract
Recently there has been a great deal of interest in the design and stability analysis of 2 dimensional recursive digital filters. Workers in the field have concentrated most of their efforts on these problems, but little has been reported on eigher the applications of their designs, or indeed if the designs were intended for any specific applications. In this paper, we examine the use of recursive filters in the image processing area, specifically that of (a) inverse filtering in image restoration and (b) image enhancement. Recent methods are used in the design of these filters and discussions as to their usefulness with supporting examples are presented.
A. Chottera, Graham A. Jullien
ICASSP2
1978 An error anaylsis of a FFT implementation using the residue number system
abstract
This paper considers an implementation of the FFT based upon the residue number system. This system offers the advantages of using integer based arithmetic operations and a simple hardware realization involving table look-up arrays. The proposed architecture is such that rapid evolutionary changes in read-only-memory technology can be easily incorporated into the hardware realization. In this paper, a generalized expression for predicting RMS relative error that includes A/D quantization, integer normalization and scaling rounding considerations, has been derived. This analysis leads to the optimal choice of several parameters, which are related to error minimization and simplification of the hardware realization. The generalized expression derived with respect to the FFT can also be used to predict errors in high-speed convolution filters.
Ben-Dau Tseng, William C. Miller, Graham A. Jullien, J. J. Soltis, A. Baraniecka
ICASSP3
1978 Residue Number Scaling and Other Operations Using ROM Arrays
abstract
Over the last two decades there has been considerable interest in the implementation of digital computer elements using hardware based on the residue number system. This paper considers implementing such systems with arrays of look-up tables placed in high density read-only memories. The type of system discussed is restricted to one in which the only operations are addition, subtraction, multiplication, and scaling by a predetermined constant. Special attention is given to the scaling algorithm, and two different scaling algorithms are developed.
Graham A. Jullien
IEEE Trans. Computers1