EDBT 2026 Demo / reviewers in the wild / expert
Bruce F. Cockburn
dblp:68/6781
· DBLP profile ↗
49ranked-venue papers
6as first author
3since 2021 · last 2024
0000-0002-4340-8394ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 6 first-author · 3 since 2021Computer networks · 8Software engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hardware-Efficient Logarithmic Floating-Point Multipliers for Error-Tolerant ApplicationsabstractThe increasing computational intensity of important new applications poses a challenge for their use in resource-restricted devices. Approximate computing using power-efficient arithmetic circuits is one of the emerging strategies to reach this objective. In this article, five hardware-efficient logarithmic floating-point (FP) multipliers are proposed, which all use simple operators, such as adders and multiplexers, to replace complex and more costly conventional FP multipliers. Radix-4 logarithms are used to further reduce the hardware complexity. These designs produce double-sided error distributions to mitigate error accumulation in complex computations. The proposed multipliers provide superior trade-offs between accuracy and hardware, with up to 30.8% higher accuracy than a recent logarithmic FP design or up to$68\times $less energy than the conventional FP multiplier. Using the proposed FP logarithmic multipliers in JPEG image compression achieves higher image quality than a recent logarithmic multiplier design with up to 4.7 dB larger peak signal-to-noise ratio. For training in benchmark NN applications, the proposed FP multipliers can slightly improve the classification accuracy while achieving$4.2\times $less energy and$2.2\times $smaller area than the state-of-the-art design. Zijing Niu, Honglan Jiang, Bruce F. Cockburn, Leibo Liu, Jie Han 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | A Logarithmic Floating-Point Multiplier for the Efficient Training of Neural NetworksabstractThe development of important applications of increasingly large neural networks (NNs) is spurring research that aims to increase the power efficiency of the arithmetic circuits that perform the huge amount of computation in NNs. The floating-point (FP) representation with a large dynamic range is usually used for training. In this paper, it is shown that the FP representation is naturally suited for the binary logarithm of numbers. Thus, it favors a design based on logarithmic arithmetic. Specifically, we propose an efficient hardware implementation of logarithmic FP multiplication that uses simpler operations to replace complex multipliers for the training of NNs. This design produces a double-sided error distribution that mitigates the accumulative effect of errors in iterative operations, so it is up to 45% more accurate than a recent logarithmic FP design. The proposed multiplier also consumes up to 23.5x less energy and 10.7x smaller area compared to exact FP multipliers. Benchmark NN applications, including a 922-neuron model for the MNIST dataset, show that the classification accuracy can be slightly improved using the proposed multiplier, while achieving up to 2.4x less energy and 2.8x smaller area with a better performance. Zijing Niu, Honglan Jiang, Mohammad Saeed Ansari, Bruce F. Cockburn, Leibo Liu, Jie Han 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2021 | An Improved Logarithmic Multiplier for Energy-Efficient Neural ComputingabstractMultiplication is the most resource-hungry operation in neural networks (NNs). Logarithmic multipliers (LMs) simplify multiplication to shift and addition operations and thus reduce the energy consumption. Since implementing the logarithm in a compact circuit often introduces approximation, some accuracy loss is inevitable in LMs. However, this inaccuracy accords with the inherent error tolerance of NNs and their associated applications. This article proposes an improved logarithmic multiplier (ILM) that, unlike existing designs, rounds both inputs to their nearest powers of two by using a proposed nearest-one detector (NOD) circuit. Considering that the output of the NOD uses a one-hot representation, some entries in the truth table of a conventional adder cannot occur. Hence, a compact adder is designed for the reduced truth table. The 8x8 ILM achieves up to 17.48 percent saving in power consumption compared to a recent LM in the literature while being almost 8 percent more accurate. Moreover, the evaluation of the ILM for two benchmark NN workloads shows up to 21.85 percent reduction in energy consumption compared to the NNs implemented with other LMs. Interestingly, using the ILM increases the classification accuracy of the considered NNs by up to 1.4 percent compared to a NN implementation that uses exact multipliers. Mohammad Saeed Ansari, Bruce F. Cockburn, Jie Han 0001 |
IEEE Trans. Computers | 2 |
| 2020 | Improving the Accuracy and Hardware Efficiency of Neural Networks Using Approximate MultipliersabstractImproving the accuracy of a neural network (NN) usually requires using larger hardware that consumes more energy. However, the error tolerance of NNs and their applications allow approximate computing techniques to be applied to reduce implementation costs. Given that multiplication is the most resource-intensive and power-hungry operation in NNs, more economical approximate multipliers (AMs) can significantly reduce hardware costs. In this article, we show that using AMs can also improve the NN accuracy by introducing noise. We consider two categories of AMs: 1) deliberately designed and 2) Cartesian genetic programing (CGP)-based AMs. The exact multipliers in two representative NNs, a multilayer perceptron (MLP) and a convolutional NN (CNN), are replaced with approximate designs to evaluate their effect on the classification accuracy of the Mixed National Institute of Standards and Technology (MNIST) and Street View House Numbers (SVHN) data sets, respectively. Interestingly, up to 0.63% improvement in the classification accuracy is achieved with reductions of 71.45% and 61.55% in the energy consumption and area, respectively. Finally, the features in an AM are identified that tend to make one design outperform others with respect to NN accuracy. Those features are then used to train a predictor that indicates how well an AM is likely to work in an NN. Mohammad Saeed Ansari, Vojtech Mrazek, Bruce F. Cockburn, Lukás Sekanina, Zdenek Vasícek, Jie Han 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | A Hardware-Efficient Logarithmic Multiplier with Improved AccuracyabstractLogarithmic multipliers take the base-2 logarithm of the operands and perform multiplication by only using shift and addition operations. Since computing the logarithm is often an approximate process, some accuracy loss is inevitable in such designs. However, the area, latency, and power consumption can be significantly improved at the cost of accuracy loss. This paper presents a novel method to approximate log2N that, unlike the existing approaches, rounds N to its nearest power of two instead of the highest power of two smaller than or equal to N. This approximation technique is then used to design two improved 16×16 logarithmic multipliers that use exact and approximate adders (ILM-EA and ILM-AA, respectively). These multipliers achieve up to 24.42% and 9.82% savings in area and power-delay product, respectively, compared to the state-of-the-art design in the literature with similar accuracy. The proposed designs are evaluated in the Joint Photographic Experts Group (JPEG) image compression algorithm and their advantages over other approximate logarithmic multipliers are shown. Mohammad Saeed Ansari, Bruce F. Cockburn, Jie Han 0001 |
DATE | 2 |
| 2019 | Characterizing Approximate Adders and Multipliers Optimized under Different Design ConstraintsabstractTaking advantage of the error resilience in many applications as well as the perceptual limitations of humans, numerous approximate arithmetic circuits have been proposed that trade off accuracy for higher speed or lower power in emerging applications that exploit approximate computing. However, characterizing the various approximate designs for a specific application under certain performance constraints becomes a new challenge. In this paper, approximate adders and multipliers are evaluated and compared for a better understanding of their characteristics when the implementations are optimized for performance or power. Although simple truncation can effectively reduce the hardware of an arithmetic circuit, it is shown that some other designs perform better in speed, power and power-delay product. For instance, many approximate adders have a higher performance than a truncated adder. A truncated multiplier is faster but consumes a higher power than most approximate designs for achieving a similar mean error magnitude. The logarithmic multipliers are very fast and power-efficient at a lower accuracy. Approximate multipliers can also be generated by an automated process to be very efficient while ensuring a sufficiently high accuracy. Honglan Jiang, Francisco J. H. Santiago, Mohammad Saeed Ansari, Leibo Liu, Bruce F. Cockburn, Fabrizio Lombardi, Jie Han 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2018 | Automatic Selection of Process Corner Simulations for Faster Design VerificationabstractIntegrated circuit designs are verified in simulation over a set of process corners, which are combinations of expected transistor properties, power supply voltages, and die temperatures. The simulation time per corner can be long and semiconductor processes can have more than 1000 corners. Simulation is thus a serious bottleneck in design verification. We propose an algorithm that selects the smallest number of process corner simulations that are required to estimate minimum and/or maximum values of the output functions that model circuit behavior. Using our best corner selection algorithm, the required number of process corner simulations is reduced by an average of 79% (a speed-up of 4.71) with respect to a set of 46 output functions from nine industrial benchmark circuits. Michael Shoniker, Oleg Oleynikov, Bruce F. Cockburn, Jie Han 0001, Manish Rana, Witold Pedrycz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Feedback-Based Low-Power Soft-Error-Tolerant Design for Dual-Modular Redundancy
Yufeng Li 0003, Jie Han 0001, Jianhao Hu, Fan Yang 0001, Xuan Zeng 0001, Bruce F. Cockburn, Jie Chen 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2017 | A true random number generator based on parallel STT-MTJsabstractRandom number generators are an essential part of cryptographic systems. For the highest level of security, true random number generators (TRNG) are needed instead of pseudorandom number generators. In this paper, the stochastic behavior of the spin transfer torque magnetic tunnel junction (STT-MTJ) is utilized to produce a TRNG design. A parallel structure with multiple MTJs is proposed that minimizes device variation effects. The design is validated in a 28-nm CMOS process with Monte Carlo simulation using a compact model of the MTJ. The National Institute of Standards and Technology (NIST) statistical test suite is used to verify the randomness quality when generating encryption keys for the Transport Layer Security or Secure Sockets Layer (TLS/SSL) cryptographic protocol. This design has a generation speed of 177.8 Mbit/s, and an energy of 0.64 pJ is consumed to set up the state in one MTJ. Yuanzhuo Qu, Jie Han 0001, Bruce F. Cockburn, Witold Pedrycz, Yue Zhang 0010, Weisheng Zhao 0001 |
DATE | 3 |
| 2017 | Optimization of Low-Density Parity Check decoder performance for OpenCL designs synthesized to FPGAs
Andrew J. Maier, Bruce F. Cockburn |
J. Parallel Distributed Comput. | 2 |
| 2016 | SRAM memory margin probability failure estimation using Gaussian Process regressionabstractEstimating the failure probabilities of SRAM memory cells using Monte Carlo or Importance Sampling techniques is expensive in the number of SPICE simulations needed. This paper presents a methodology for estimating the dynamic margin failure probabilities by building a surrogate model of the dynamic margin using Gaussian Process regression. Additive kernel functions that can extrapolate the margin values from the simulated samples are presented. These proposed kernel functions decrease the out-of-sample error of the surrogate model for a 6T cell by 32% compared with a six-dimensional universal kernel such as a Radial-Basis-Function kernel (RBF). Finally, the failure probability values predicted by a surrogate model built using 1250 SPICE simulations are reported and compared with Monte Carlo analysis with 106samples. The results show a relative error of 30% at 0.4V (predicted value of 4×10-6for the Monte Carlo estimate of 3×10-6) and a relative error of 172% at 0.3V (predicted value of 3×10-5for the Monte Carlo estimate of 1.1×10-5) for the dynamic read margin. These accuracy numbers are similar to those reported in previous proposals while the reduction in SPICE simulations is between 4× and 23× relative to these proposals and 800× compared to Monte Carlo method. Manish Rana, Ramon Canal, Jie Han 0001, Bruce F. Cockburn |
ICCD | 4 |
| 2016 | Stochastic Circuit Design and Performance Evaluation of Vector Quantization for Different Error MeasuresabstractVector quantization (VQ) is a general data compression technique that has a scalable implementation complexity and potentially a high compression ratio. In this paper, a novel implementation of VQ using stochastic circuits is proposed and its performance is evaluated against conventional binary designs. The stochastic and binary designs are compared for the same compression quality, and the circuits are synthesized for an industrial 28-nm cell library. The effects of varying the sequence length of the stochastic representation are studied with respect to throughput per area (TPA) and energy per operation (EPO). The stochastic implementations are shown to have higher EPOs than the conventional binary implementations due to longer latencies. When a shorter encoding sequence with 512 bits is used to obtain a lower quality compression measured by the L1-norm, squared L2-norm, and third-law errors, the TPA ranges from 1.16 to 2.56 times than that of the binary implementation with the same compression quality. Thus, although the stochastic implementation underperforms for a high compression quality, it outperforms the conventional binary design in terms of TPA for a reduced compression quality. By exploiting the progressive precision feature of a stochastic circuit, a readily scalable processing quality can be attained by halting the computation after different numbers of clock cycles. Jie Han 0001, Bruce F. Cockburn, Duncan G. Elliott |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Stochastic circuit design and performance evaluation of vector quantizationabstractVector quantization (VQ) is a general data compression technique that has a scalable implementation complexity and potentially a high compression ratio. In this paper, a novel implementation of VQ using stochastic circuits is proposed and its performance is evaluated. The stochastic and binary designs are compared for the same compression quality and the circuits are synthesized for an industrial 28-nm cell library. The effects of varying the sequence length of the stochastic design are studied with respect to the performance metric of throughput per area (TPA). When a shortened 512-bit encoding sequence is used to obtain a lower quality compression, the TPA is about 2.60 times that of the binary implementation with the same quality as that of the stochastic implementation measured by the L1norm error (i.e., the first-order error). Thus, the stochastic implementation outperforms the conventional binary design in terms of TPA for a relatively low compression quality. By exploiting the progressive precision feature of a stochastic circuit, a readily scalable processing quality can be attained by simply halting the computation after different numbers of clock cycles. Jie Han 0001, Bruce F. Cockburn, Duncan G. Elliott |
ASAP | 3 |
| 2015 | Minimizing the number of process corner simulations during design verification
Michael Shoniker, Bruce F. Cockburn, Jie Han 0001, Witold Pedrycz |
DATE | 2 |
| 2014 | Adaptive dual-threshold neural signal compression suitable for implantable recordingabstractThis paper presents a digital architecture for neural signal compression using adaptive two-threshold spike detection and a nonlinear discrete wavelet coefficient selection scheme. The circuits and algorithms are described and compared with the state-of-the-art. The proposed 16-channel digital architecture is capable of neural data compression to 0.5% of the original raw data rate while consuming 21μW, with 30-kHz 8-bit sampling, in a 0.8-V 130-nm low-power IBM process. Russell Dodd, Bruce F. Cockburn, Vincent C. Gaudet |
ICASSP | 2 |
| 2012 | Accurate simulation of non-isotropic fading channels with arbitrary temporal correlationabstractThe accurate simulation of wireless channels is important since it permits the realistic and repeatable performance measurement of wireless systems. A new technique is proposed for simulating Rayleigh fading channels with isotropic or non-isotropic scattering and with arbitrary temporal correlation. Fading samples are generated by passing Gaussian samples through a spectrum shaping filter. A new iterative algorithm is then presented for designing stable complex infinite impulse response (IIR) filters with quantised coefficients. The algorithm utilises a least-squares cost function augmented with a barrier function to ensure filter stability and to reduce quantisation noise. The performance of the proposed filter design algorithm is verified with 18-bit fixed-point simulations of different fading channel scenarios including isotropic and non-isotropic scattering and the IEEE 802.11n model F fading spectrum. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn |
IET Commun. | 3 |
| 2012 | Layered space-time multiple-input multiple-output detector with parameterisable performanceabstractThis brief presents an improved layered space-time symbol detection algorithm for multiple-input multiple-output (MIMO) wireless systems. The proposed detection scheme utilises a different layer-ordering strategy than the so-called optimal ordering used in the conventional Bell Laboratories Layered Space-Time (BLAST) algorithm. To calculate the nulling vectors and their associated ordering, instead of using direct numerical techniques, the authors utilise a numerically stable iterative solution. It is shown that the performance of the proposed detector approaches closely that of an optimal maximum-likelihood detector at the expense of greater computational complexity compared with BLAST. Utilising the proposed parameterisable detector, one can trade off between the computational requirement and the error-rate performance. Amirhossein Alimohammad 0001, Santosh V. Nagaraj, Saeed Fouladi Fard, Bruce F. Cockburn |
IET Commun. | 4 |
| 2012 | Hardware Implementation of Nakagami and Weibull Variate GeneratorsabstractAn efficient implementation of Nakagami-mand Weibull variate generators on a single field-programmable gate array (FPGA) is presented. The hardware model first generates a correlated Rayleigh fading variate sequence and then transforms it into a sequence of Nakagami-mor Weibull fading variates. A biquad processor facilitates the compact implementation of a Rayleigh variate generator with arbitrary autocorrelation properties. A combination of logarithmic and linear domain segmentations along with piece-wise linear approximations is used to accurately implement the nonlinear numerical functions required to transform the correlated Rayleigh fading process into Nakagami-m or Weibull fading processes. When implemented on a Xilinx Virtex-5 5VSX240TFF1738-2 FPGA, the fading simulator uses only 1.6% of the configurable slices, 1.2% of the DSP48E modules and 3 block memories, while operating at 120 MHz, generating 120 million complex variates per second. The throughput can be increased up to 373 MHz with this FPGA if two separate clock sources are utilized. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Accurate multiple-input multiple-output fading channel simulator using a compact and highthroughput reconfigurable architectureabstractThis article presents an ultra-compact and high-throughput reconfigurable fading channel simulator that supports a relatively large number of propagation paths. To closely reflect actual radio channels, the authors used a recently improved Rayleigh and Ricean fading channel model based on the sum-of-sinusoids technique. The improved model is optimised for hardware compactness. To achieve a fast fading variate generation rate with much less hardware and no significant loss in accuracy, the new scheme first generates fading samples at a lower rate using a time-multiplexed datapath that can be fit into a small fraction of a field-programmable gate array (FPGA). In the second step, the simulator uses a compact multiplication-free linear interpolator to produce the fading samples at the full symbol rate. Implementing a 64-path fading channel simulator on a Xilinx Virtex-4 XC4VLX200-11 FPGA requires only 13 044 (14%) of the configurable slices, 10 (2%) of the block memories and one (1%) of the dedicated DSP blocks, while generating 64 × 191 million complex-valued fading samples per second. The simulated paths can be readily combined to form high path count models for multiple-input multiple-output systems as well as frequency-selective channels. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn |
IET Commun. | 3 |
| 2011 | Single-field programmable gate array simulator for geometric multiple-input multiple-output fading channel modelsabstractThe authors propose a compact and fast fading channel simulator for the baseband verification and prototyping of multiple-input multiple-output (MIMO) wireless communication systems. The simulator accurately and efficiently implements models of both single-bounce and multiple-bounce geometric propagation conditions. Fading samples are generated at a low rate, comparable to the Doppler frequency, and then interpolated to match the desired sample rate. Bit-true simulations verify the accuracy of the hardware simulator. As an example, when implemented on a Xilinx XC5VLX110 field-programmable gate array (FPGA), the 4×4 MIMO geometric fading channel simulator occupies only 6.6% of the configurable slices while generating more than 16×324 million samples per second. The geometric MIMO fading channel simulator is well suited for use in an FPGA-based error rate performance verification system. Saeed Fouladi Fard, Amirhossein Alimohammad 0001, Bruce F. Cockburn |
IET Commun. | 3 |
| 2011 | Hardware Implementation of Rayleigh and Ricean Variate GeneratorsabstractCompact and fast implementations of digital Rayleigh and Ricean variate generators are presented. Polynomial curve fitting is utilized along with a combination of logarithmic and uniform domain segmentation to provide accuracy, compactness and fast variate generation. A typical instantiation of the proposed Rayleigh generator occupies 124 (<;1%) of the configurable slices, two dedicated multipliers (<;1%), and one on-chip block memory (<; 1%) of a Xilinx Virtex-5 field-programmable gate array (FPGA) and operates at 317 MHz, generating 317 million Rayleigh variates per second. The Ricean variate generator implementation on the same device utilizes 366 (<; 1%) of the logical slices, three on-chip block memories ( <;1%), and 11 (2.8%) of the dedicated multipliers. The application of the Rayleigh and Ricean variate generators is demonstrated in a FPGA-based bit error rate simulator that measures at hardware speeds the symbol error rate performance of a typical wireless communication system over Rayleigh and Ricean fading channels. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Design and Characterization of a Multilevel DRAMabstractMultilevel DRAM (MLDRAM) increases the storage density of DRAMs by using more than two signal levels in the data storage cells. Our MLDRAM uses reference and data cell signals that are generated in the cell array using charge sharing. The single-step sensing method uses multiple reference signals in parallel. We describe an operational 19200-cell MLDRAM test chip in 1.8-V, 180-nm mixed-signal CMOS that allows 1, 1.5, 2, 2.25, and 2.5 bits-per-cell operation using 2, 3, 4, 5, and 6 data signal levels, respectively. Characterization features include a partitioned memory array with four different data cell sizes, two sense amplifier sizes, and selective bitline shielding. New tests were developed using an MLDRAM fault model covering basic functionality, retention time, multilevel march, inter-bitline coupling and cell plate voltage bump tests. We show that, with short bitlines, MLDRAM's noise margins can be similar to DRAM's to more reliably store two bits in a 1T1C cell. John C. Koob, Sue Ann Ung, Bruce F. Cockburn, Duncan G. Elliott |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | A Unified Architecture for the Accurate and High-Throughput Implementation of Six Key Elementary FunctionsabstractThis paper presents a unified architecture for the compact implementation of several key elementary functions, including reciprocal, square root, and logarithm, in single-precision floating-point arithmetic. The proposed high-throughput design is based on uniform domain segmentation and curve fitting techniques. Numerically accurate least-squares regression is utilized to calculate the polynomial coefficients. The architecture is optimized by analyzing the trade-off between the size of the required memory and the precision of intermediate variables to achieve the minimum 23-bit accuracy required for single-precision floating-point representation. The efficiency of the proposed unified data path is demonstrated on a common field-programmable gate array. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn |
IEEE Trans. Computers | 3 |
| 2009 | FPGA-based accelerator for the verification of leading-edge wireless systemsabstractThe design of communication systems becomes increasingly challenging as product complexity and cost pressures increase and as the time-to-market is shortened more than ever before. This paper presents a bit error rate tester (BERT) for the hardware-based verification of the physical layer (PHY) layer of emerging wireless systems. We integrate fundamental modules of a typical PHY layer along with the channel simulator onto a single field-programmable gate array (FPGA). For a proof-of-concept, we present the results of a FPGA-based performance verification exercise for a multiple antenna system. The proposed BERT system significantly decreases the test time compared to conventional software-based verification, hence increasing designer productivity. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn |
DAC | 3 |
| 2009 | A flexible layered architecture for accurate digital baseband algorithm development and verificationabstractMany emerging communication technologies significantly increase the complexity of the physical layer and have dramatically increased the number of operating configurations. To ensure maximum performance, designers have to optimize their algorithm implementations, which requires for comprehensive performance testing in all possible operating modes various channel conditions. This paper presents a flexible and affordable framework for baseband algorithm development and performance verification for digital communication systems with an arbitrary number of modules, each operating at a possibly different sampling rate with various latencies. The proposed architecture is scalable to support complex scenarios, such as multiple antenna systems, and is compact enough to be implemented within a single field-programmable gate array. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn |
DATE | 3 |
| 2009 | A Single FPGA Filter-Based Multipath Fading EmulatorabstractEmulation of fading channels is a key step in the design and verification of wireless communication systems. Testing wireless transceivers with actual fading channels is inconvenient due to unrepeatable and uncontrollable channel conditions. In this paper we present a compact field-programmable gate array (FPGA) implementation for a circuit that generates temporally-correlated fading variates for emulating multipath fading radio channels. The implemented fading emulator is flexible enough to model different propagation scenarios accurately and is compact enough that it can be implemented on the same FPGA with the design under test (DUT) for greater emulation efficiency and speed-up. Several streams of Rayleigh or Rician fading variates are generated by passing independent samples of Gaussian noise through spectrum shaping filters. The new baseband emulator is fully parameterizable and can emulate a wide variety of single and multiple antenna scenarios. Saeed Fouladi Fard, Amirhossein Alimohammad 0001, Bruce F. Cockburn, Christian Schlegel |
GLOBECOM | 3 |
| 2009 | Compact Rayleigh and Rician fading simulator based on random walk processesabstractThis article describes a significantly improved sum-of-sinusoids-based model for the accurate simulation of time-correlated Rayleigh and Rician fading channels. The proposed model utilises random walk processes instead of random variables for some of the sinusoid parameters to more accurately reproduce the behaviour of wireless radio propagation. Every fading block generated using our model has accurate statistical properties on its own and hence, unlike previously proposed models, there is no need for time-consuming ensemble-averaging over multiple blocks. Using numerical simulation it is shown that the important statistical properties of the generated fading samples have excellent agreement with the theoretical reference functions. A fixed-point hardware implementation of the corresponding Rayleigh and Rician fading channel simulator on a field-programmable gate array (FPGA) is presented. By efficiently scheduling the operations, the reconfigurable fading channel simulator is compact enough that it can be efficiently used to simulate multipath scenarios and multiple-antenna systems (e.g. a 4×4 MIMO channel) using a single FPGA. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
IET Commun. | 3 |
| 2008 | A Novel Technique for Efficient Hardware Simulation of Spatiotemporally Correlated MIMO Fading ChannelsabstractWe present a fading model with a compact and fast hardware implementation suitable for correlated Rayleigh fading channel simulators. The proposed scheme is based on the sum-of-sinusoids model because of its flexibility and efficient mapping onto hardware. Using numerical simulation, it is shown that the statistical properties of the generated fading variates match the theoretical reference model. Since the cross-correlations between sequences of generated fading variates are small, this model can also be used to implement a time-correlated multiple-input multiple-output (MIMO) fading channel simulator on a single field-programmable gate array (FPGA). The MIMO channel simulator can also be extended to support spatial correlation between generated fading samples. An implementation of a spatiotemporally correlated (4, 4) MIMO channel simulator on a Xilinx Virtex-II Pro XC2VP100-6 FPGA uses 46% of the configurable slices, 30% of the dedicated multipliers, and 32% of the on-chip block memories while generating 4 times 201 million 2 times 16-bit complex-valued fading samples per second. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
ICC | 3 |
| 2008 | On the efficiency and accuracy of hybrid pseudo-random number generators for FPGA-based simulationsabstractMost commonly-used pseudo-random number generators (PNGs) in computer systems are based on linear recurrence. These deterministic PNGs have fast and compact implementations, andean ensure very long periods. However, the points generated by linear PNGs in fact have a regular lattice structure and are thus not suit able for applications that rely on the assumption of uniformly distributed pseudo-random numbers (PNs). In this paper we propose and evaluate several fast and compact linear, non-linear, and hybrid PNGs for a field- programmable gate array (FPGA). The PNGs have excellent equidistribution properties and very small autocorrelations, and have very long repetition periods. The distribution and long-range correlation properties of the new generators are efficiently, and much more rapidly, estimated at hardware speeds using designed modules within the FPGA. The results of these statistical tests confirm that the combination of several linear PNGs or the combination of even one small non-linear PNG with a linear PNG significantly improves the statistical properties of the generated PNs. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
IPDPS | 3 |
| 2008 | A single-FPGA multipath MIMO fading channel simulatorabstractWe present an accurate model for compact implementations of Rayleigh and Rician fading channels. Verification of the proposed fading simulator is performed by comparing the simulated statistics with those of the ideal reference models. A parameterizable field-programmable gate array (FPGA) implementation of the channel simulator is presented. The design is readily scalable to support multipath fading channels and multiple-input multiple-output (MIMO) systems. A 16-path fading channel, providing either Rician or Rayleigh fading, uses 41% of the configurable slices, 33% of the dedicated multipliers, and 32% of the on-chip block memories of a Xilinx Virtex-II Pro XC2VP100-6 FPGA while generating over 200 million complex- valued fading coefficients per second. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
ISCAS | 3 |
| 2008 | A 600-Mb/s encoder and decoder for low-density parity-check convolutional codesabstractA 600-Mb/s rate-1/2 (128,3,6) LDPC convolutional code encoder and decoder was implemented in a 90-nm CMOS process. The encoder operates at 1.1 GHz and includes built-in all-phase termination. The decoder design maximizes throughput while minimizing the number of memory banks and delivering an information throughput of 1 bit per clock cycle. The size of the decoder controller is minimized by sharing it among an arbitrary number of decoder processors. The decoder dissipates 0.61 nJ of energy per decoded information bit at an SNR of 2.0 and a throughput of 600 Mb/s. An integrated test system enables accurate power measurements for various SNR settings. Tyler L. Brandon, John C. Koob, Leendert van den Berg, Zhengang Chen, Amirhossein Alimohammad 0001, Ramkrishna Swamy, Jason Klaus, Stephen Bates, Vincent C. Gaudet, Bruce F. Cockburn, Duncan G. Elliott |
ISCAS | 10 |
| 2008 | Hardware-based Error Rate Testing of Digital Baseband Communication SystemsabstractWe present a flexible architecture for evaluating the bit-error-rate (BER) performance of prototype digital baseband communication systems. Using an efficient elastic buffer interface, an arbitrary baseband module can be added to the cascaded architecture of a digital baseband communication system, independent of the module's operating rate, its position in the cascade structure, and its latency. The proposed BER tester uses an accurate fading channel model and a Gaussian noise generator to provide a realistic and repeatable test environment in the laboratory. This evaluation environment should reduce the need for time-consuming field tests, hence reducing the time-to-market and increasing productivity. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn |
ITC | 3 |
| 2008 | An Accurate and Compact Rayleigh and Rician Fading Channel SimulatorabstractA stochastic sum-of-sinusoids based simulation model is proposed for Rayleigh and Rician fading channels. The time-averaged statistical properties of the new model have been significantly improved compared to existing models. Verification of the proposed fading simulator is carried out by comparing its measured statistical properties with the properties of the ideal reference models. The simulator utilizes a time-overlapped implementation strategy to provide a compact design suitable for multiple antenna simulators. An implementation of the resulting Rician fading simulator on a Xilinx Virtex-II Pro XC2VP100- 6 FPGA uses only 2% of the configurable slices, 1% of the dedicated multipliers, and 2% of the on-chip block memories while generating 201 million 2 times 16-bit complex-valued fading samples per second. The scalable design of the fading channel simulator enables a straightforward implementation of multiple antenna channels and different diversity schemes. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
VTC Spring | 3 |
| 2008 | A scalable LDPC decoder ASIC architecture with bit-serial message exchange
Tyler L. Brandon, Robert Hang, Gary Block, Vincent C. Gaudet, Bruce F. Cockburn, Sheryl L. Howard, Christian Giasson, Keith Boyle, Paul Goud, Siavash Sheikh Zeinoddin, Anthony Rapley, Stephen Bates, Duncan G. Elliott, Christian Schlegel |
Integr. | 5 |
| 2008 | A Compact and Accurate Gaussian Variate GeneratorabstractA compact, fast, and accurate realization of a digital Gaussian variate generator (GVG) based on the Box-Muller algorithm is presented. The proposed GVG has a faster Gaussian sample generation rate and higher tail accuracy with a lower hardware cost than published designs. The GVG design can be readily configured to achieve arbitrary tail accuracy (i.e., with a proposed 16-bit datapath up to plusmn15 times the standard deviation sigma) with only small variations in hardware utilization, and without degrading the output sample rate. Polynomial curve fitting is utilized along with a hybrid (i.e., combination of logarithmic and uniform) segmentation and a scaling scheme to maintain accuracy. A typical instantiation of the proposed GVG occupies only 534 configurable slices, two on-chip block memories, and three dedicated multipliers of the Xilinx Virtex-II XC2V4000-6 field-programmable gate array (FPGA) and operates at 248 MHz, generating 496 million Gaussian variates (GVs) per second within a range of plusmn6.66sigma. To accurately achieve a range of plusmn9.4sigma, the GVG uses 852 configurable slices, three block memories, and three on-chip dedicated multipliers of the same FPGA while still operating at 248 MHz, generating 496 million GVs per second. The core area and performance of a GVG implemented in a 90-nm CMOS technology are also given. The statistical characteristics of the GVG are evaluated and confirmed using multiple standard statistical goodness-of-fit tests. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | A Compact Fading Channel Simulator Using Timing-Driven Resource SharingabstractThis paper presents a computationally-efficient design and implementation technique for fading channel simulators. Our fixed-point implementation of a Rayleigh fading channel simulator on a field-programmable gate array (FPGA) utilizes only 4% of the configurable slices, 19% of the dedicated multipliers, and 2% of the on-chip memory blocks, while generating 12.5 million statistically accurate fading variates per second. The designed channel emulator can be parameterized to simulate a wide variety of different channel characteristics over bandwidths of up to 12.5 MHz. The compact simulator can also be instantiated multiple times to assess the performance of communication systems over multiple-input multiple-output (MIMO) channels. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
ASAP | 3 |
| 2007 | A Flexible Filter Processor for Fading Channel SimulationabstractA flexible and compact general-purpose filter processor is presented. This processor is intended for the hardware-based simulation of wireless channels on field-programmable gate arrays (FPGAs). When implemented on a Xilinx Virtex2P XC2VP100-6 FPGA, it utilizes 2% of the configurable slices, 9% of the dedicated 18times18-bitmultipliers, and 14 BlockRAMs. When paired with multiplicationfree interpolators, it can generate up to 300 million fading samples per second. The statistical properties of generated fading samples are shown to closely match the theoretical reference statistical properties. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
FCCM | 3 |
| 2007 | Compound Uniform Random Number Generators with On-Chhip Correlation and Distribution MeasurementsabstractA fast and compact implementation of nonlinear pseudo-random number generators (PNGs) on a field-programmable gate array (FPGA) is presented. The distribution and correlation properties of PNGs are efficiently, and very rapidly, estimated within the FPGA. It is shown that the pseudo-random output of nonlinear PNGs have significantly improved randomness properties compared to linear PNGs. To the best of our knowledge, this paper presents the first long-range correlation and distribution measurements in hardware for PNGs on FPGAs. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
FPT | 3 |
| 2007 | An Improved SOS-Based Fading Channel EmulatorabstractWe describe an improved scheme for simulating Rayleigh fading channels which accurately reproduces the required channel statistics. The new scheme is based on the sum- of-sinusoids Rayleigh fading model because of its flexibility and efficient mapping onto hardware. Using numerical simulation it is shown that the statistical properties of the generated fading variates, such as the probability density function, the autocorrelation, and the level crossing rate, follow the desired theoretical properties. A fixed-point implementation of the fading channel simulator on a field-programmable gate array utilizes only 5% of the configurable slices and generates over 200 million 16-bit fading variates per second. Amirhossein Alimohammad 0001, Saeed Fouladi Fard, Bruce F. Cockburn, Christian Schlegel |
VTC Fall | 3 |
| 2005 | Test and Characterization of a Variable-Capacity Multilevel DRAMabstractMultilevel DRAM (MLDRAM) increases the storage density of DRAMs by using more than two data signal levels in the storage cells. An operational 19200-cell MLDRAM in 1.8-V 0.18-/spl mu/m mixed-signal CMOS is described that allows 1, 1.5, 2, 2.25 and 2.5 bits-per-cell operation using 2, 3, 4, 5 and 6 data signal levels, respectively. The MLDRAM uses reference and data cell signals that are generated in the cell array using charge sharing. The single-step sensing method uses multiple reference signals in parallel. Test chip characterization features include four cell sizes, two sense amplifier sizes, and bitline shields for half of the cells. New tests were developed based on an MLDRAM fault model. These include basic functionality, retention time, multilevel march, inter-bitline coupling, and cell-plate voltage bump tests. Our results show that the data and reference signals are generated correctly and that MLDRAM is possible for up to six signal levels. John C. Koob, Sue Ann Ung, Ashwin S. Rao, Daniel A. Leder, Craig S. Joly, Kristopher C. Breen, Tyler L. Brandon, Michael Hume, Bruce F. Cockburn, Duncan G. Elliott |
VTS | 9 |
| 2005 | Design of a 3-D fully depleted SOI computational RAMabstractWe introduce a three-dimensional (3-D) processor-in-memory integrated circuit design that provides progressively increasing processing power as the number of stacked dies increases, while incurring no extra design effort or mask sets. Innovative techniques for processor/memory redundancy and fast global bus evaluation are described. The architecture can be augmented with a nearest-neighbor physical 3-D communications network that can substantially reduce interconnect lengths and relieve routing congestion. The test chip, with 128 Kb of memory and 512 processing elements (PEs) on two fully depleted silicon-on-insulator (SOI) dies, can achieve a peak of 170 billion-bit-operations per second at 400 MHz. John C. Koob, Daniel A. Leder, Raymond J. Sung, Tyler L. Brandon, Duncan G. Elliott, Bruce F. Cockburn, Lisa G. McIlrath |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 1999 | Voiceband signal classification using statistically optimal combinations of low-complexity discriminant variablesabstractThis paper considers the problem of classifying voiceband modem, facsimile, native binary data, and speech signals on a 64-kb/s digital channel using low complexity discriminant functions similar to those investigated by Benvenuto (1989). Benvenuto's method is simplified by replacing an approximate demodulation step with rectification of the passband signal and improved by using statistical analysis to select and combine the most effective variables into statistically optimal discriminant functions. Jeremy S. Sewall, Bruce F. Cockburn |
IEEE Trans. Commun. | 2 |
| 1998 | Transition Maximization Techniques for Enhancing the Two-Pattern Fault Coverage of Pseudorandom Test Pattern GeneratorsabstractThis paper presents simulation evidence supporting the use of bit transition maximization techniques in the design of hardware test pattern generators (TPGs). Bit transition maximization is a heuristic technique that involves increasing the probability that a bit will change values going from one test pattern to the next. For most of the ISCAS-85 benchmarks and many of the ISCAS-89 benchmarks bit transition maximization enhances the fault coverage of two-pattern faults such as gate delay faults and CMOS transistor stuck-open faults. It achieves these benefits without reducing the fault coverage with respect to classical stuck-at faults. Bruce F. Cockburn, Albert L.-C. Kwong |
VTS | 1 |
| 1995 | Synthesized Transparent BIST for Detecting Scrambled Pattern-Sensitive Faults in RAMsabstractThis paper describes a synthesizable, transparent, built-in self-test (BIST) scheme for random-access memories (RAMs). By altering only two parameters in a VHDL specification, BIST circuits can be automatically generated to detect 2-, 3- or 4-cell write-triggered coupling faults as well as two different classes of 5-cell faults. The 5-cell faults represent either unlinked scrambled active physical neighborhood pattern-sensitive faults (PNPSFs), or arbitrary combinations of unlinked scrambled active, static, and passive PNPSFs. The BIST scheme uses a modified version of Nicolaidis' method to make the applied tests transparent; thus the data that were held in the RAM at the start of the test will be restored by the end of the test, if no faults are present. All single faults of the above fault types, as well as most other standard fault types, are guaranteed to be detected because of the use of an aliasing-free signature analyzer. By comparing numerous intermediate signatures, the new design has a very low probability of aliasing when multiple faults are present. Bruce F. Cockburn, Y.-F. Nicole Sat |
ITC | 1 |
| 1994 | Deterministic tests for detecting singleV-coupling faults in RAMs
Bruce F. Cockburn |
J. Electron. Test. | 1 |
| 1994 | Tutorial on semiconductor memory testing
Bruce F. Cockburn |
J. Electron. Test. | 1 |
| 1992 | Near-optimal tests for classes of write-triggered coupling faults in RAMs
Bruce F. Cockburn, Janusz A. Brzozowski |
J. Electron. Test. | 1 |
| 1990 | Detection of coupling faults in RAMs
Janusz A. Brzozowski, Bruce F. Cockburn |
J. Electron. Test. | 2 |
| 1990 | Switch-level testability of the dynamic CMOS PLA
Bruce F. Cockburn, Janusz A. Brzozowski |
Integr. | 1 |