VLDB 2026 Research / reviewers in the wild / expert
Joseph R. Cavallaro
dblp:41/5458 · also Joe Cavallaro
· DBLP profile ↗
89ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-9841-1806ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 3 first-author · 3 since 2021Computer networks · 21 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-authorArtificial intelligence and machine learning · 3Security and privacy · 3 · 3 since 2021Theory of computation · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Clinical Features and Physiological Signals Fusion Network for Mechanical Circulatory Support Need Prediction in Pediatric Cardiac Intensive Care UnitabstractWe link the hemodynamic response to inotropic agents with outcomes related to Mechanical Circulatory Support (MCS) by analyzing physiological time series and clinical features using a Machine Learning/Deep Learning ensemble approach for multi-modal waveforms in the pediatric cardiac intensive care setting of a quaternary-care hospital. Unlike existing studies that typically process a single feature type or focus on short-term diagnoses from physiological signals, our novel system processes minute-by-minute multi-sensor data to identify the need for MCS in patients admitted with acute decompensated heart failure. The data used includes tabular clinical features, time series from hemodynamic monitors, and raw waveforms from electrocardiogram and arterial blood pressure signals. Our predictions support an early identification of high-risk patients after just two days of Intensive Care Unit (ICU) admission, with classification and feature importance results confirming the predictive ability of the early hemodynamic response to inotropic agent administration, achieving an AUC of 0.88 in the prediction classification task. This is particularly significant in cases where clinical decisions are not straightforward, such as those in the cohort for this study. Antonio Mendoza, Sebastian Tume, Kriti Puri, Sebastian Acosta, Joseph R. Cavallaro |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Towards Receiver-Agnostic and Collaborative Radio Frequency Fingerprint IdentificationabstractRadio frequency fingerprint identification (RFFI) is an emerging device authentication technique, which exploits the hardware characteristics of the RF front-end as device identifiers. The receiver hardware impairments interfere with the feature extraction of transmitter impairments, but their effect and mitigation have not been comprehensively studied. In this paper, we propose a receiver-agnostic RFFI system by employing adversarial training to learn the receiver-independent features. Moreover, when there are multiple receivers, collaborative inference are designed to enhance classification accuracy. Finally, we show how it is possible to leverage fine-tuning for further improvement with fewer collected signals. To validate the approach, we have conducted extensive experimental evaluation by applying the approach to a LoRaWAN case study involving ten LoRa devices and 20 software-defined radio (SDR) receivers. The results show that receiver-agnostic training enables the trained neural network to become robust to changes in receiver characteristics. The collaborative inference improves classification accuracy by up to 20% beyond a single-receiver RFFI system and fine-tuning can bring a 40% improvement for underperforming receivers. The system is further evaluated on a more practical testbed. By making additional use of online augmentation and multi-packet inference, the identification accuracy is improved from 50% to 90% at 10 dB. Guanxiong Shen, Junqing Zhang, Alan Marshall 0001, Roger F. Woods, Joseph R. Cavallaro, Liquan Chen |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | A Unified Parallel CORDIC-Based Hardware Architecture for LSTM Network AccelerationabstractDeep Neural Networks (DNNs) have recently become the standard tool for solving various practical problems in a wide range of applications with state-of-the-art performance. Recurrent Neural Networks (RNNs) such as Long Short-Term Memory (LSTM) are a subset of DNNs with fully connected single or multi-layer networks. The complex neurons and internal states of LSTM networks enable them to build a memory of events, making them ideal for time series applications. Despite the great potential of LSTM networks, their heterogeneous operations and computational resource requirements create a vast gap when it comes to the fast processing time required in real-time applications using low-power, low-cost edge devices. This work proposes a novel hardware architecture that combines serial-parallel computation with matrix algebra concepts and efficient low-power computer arithmetics for LSTM network acceleration. The hardware is based on a systolic ring of outer-product-based processing elements (PEs) and a reusable single activation function block (AFB). PEs and AFB are implemented using the coordinate rotation digital computer algorithm (CORDIC) in the linear and hyperbolic modes. Unlike most approaches, the proposed hardware can be configured to perform recurrent and non-recurrent fully connected layers (FC) computations, making it suitable for various low-power edge applications. The architecture is validated on the Xilinx PYNQ-Z1 development board using an open-source time series dataset. The implemented design achieves$\text{114} \mu \text{s}$average latency and$\text{1.8 GOPS}$throughput. The proposed design's low latency and$\text{0.438 W}$power consumption makes it suitable for resource-constrained edge platforms. Nadya A. Mohamed, Joseph R. Cavallaro |
IEEE Trans. Computers | 2 |
| 2023 | Toward Length-Versatile and Noise-Robust Radio Frequency Fingerprint IdentificationabstractRadio frequency fingerprint identification (RFFI) can classify wireless devices by analyzing the signal distortions caused by intrinsic hardware impairments. Recently, state-of-the-art neural networks have been adopted for RFFI. However, many neural networks, e.g., multilayer perceptron (MLP) and convolutional neural network (CNN), require fixed-size input data. In addition, many IoT devices work in low signal-to-noise ratio (SNR) scenarios but the RFFI performance in such scenarios is often unsatisfactory. In this paper, we analyze the reason why MLP- and CNN-based RFFI systems are constrained by the input size. To overcome this, we propose four neural networks that can process signals of variable lengths, namely flatten-free CNN, long short-term memory (LSTM) network, gated recurrent unit (GRU) network, and transformer. We adopt data augmentation during training which can significantly improve the model’s robustness to noise. We compare two augmentation schemes, namely offline and online augmentation. The results show the online one performs better. During the inference, a multi-packet inference approach is further leveraged to improve the classification accuracy in low SNR scenarios. We take LoRa as a case study and evaluate the system by classifying 10 commercial-off-the-shelf LoRa devices in various SNR conditions. The online augmentation can boost the low-SNR classification accuracy by up to 50% and the multi-packet inference approach can further increase the accuracy by over 20%. Guanxiong Shen, Junqing Zhang, Alan Marshall 0001, Mikko Valkama, Joseph R. Cavallaro |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | RT-RCG: Neural Network and Accelerator Search Towards Effective and Real-time ECG Reconstruction from Intracardiac ElectrogramsabstractThere exists a gap in terms of the signals provided by pacemakers (i.e., intracardiac electrogram (EGM)) and the signals doctors use (i.e., 12-lead electrocardiogram (ECG)) to diagnose abnormal rhythms. Therefore, the former, even if remotely transmitted, are not sufficient for doctors to provide a precise diagnosis, let alone make a timely intervention. To close this gap and make a heuristic step towards real-time critical intervention in instant response to irregular and infrequent ventricular rhythms, we propose a new framework dubbed RT-RCG to automatically search for (1) efficient Deep Neural Network (DNN) structures and then (2) corresponding accelerators, to enable R eal- T ime and high-quality R econstruction of E C G signals from E G M signals. Specifically, RT-RCG proposes a new DNN search space tailored for ECG reconstruction from EGM signals and incorporates a differentiable acceleration search (DAS) engine to efficiently navigate over the large and discrete accelerator design space to generate optimized accelerators. Extensive experiments and ablation studies under various settings consistently validate the effectiveness of our RT-RCG. To the best of our knowledge, RT-RCG is the first to leverage neural architecture search (NAS) to simultaneously tackle both reconstruction efficacy and efficiency. Yongan Zhang, Anton Banta, Yonggan Fu, Mathews John, Allison Post, Mehdi Razavi, Joseph R. Cavallaro, Behnaam Aazhang, Yingyan (Celine) Lin |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2022 | Towards Scalable and Channel-Robust Radio Frequency Fingerprint Identification for LoRaabstractRadio frequency fingerprint identification (RFFI) is a promising device authentication technique based on transmitter hardware impairments. The device-specific hardware features can be extracted at the receiver by analyzing the received signal and used for authentication. In this paper, we propose a scalable and channel-robust RFFI framework achieved by deep learning powered radio frequency fingerprint (RFF) extractor and channel independent features. Specifically, we leverage deep metric learning to train an RFF extractor, which has excellent generalization ability and can extract RFFs from previously unseen devices. Any devices can be enrolled via the pre-trained RFF extractor and the RFF database can be maintained efficiently for allowing devices to join and leave. Wireless channel impacts the RFF extraction and is tackled by exploiting channel independent features and data augmentation. We carried out extensive experimental evaluation involving 60 commercial off-the-shelf LoRa devices and a USRP N210 software defined radio platform. The results have successfully demonstrated that our framework can achieve excellent generalization abilities for rogue device detection and device classification as well as effective channel mitigation. Guanxiong Shen, Junqing Zhang, Alan Marshall 0001, Joseph R. Cavallaro |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | OFDM-Based Beam-Oriented Digital Predistortion for Massive MIMOabstractLinearization of massive MIMO arrays is a significant computational challenge that typically scales with the number of antennas. In this work, we introduce a beam-oriented digital predistortion (DPD) scheme for OFDM-based massive MIMO systems that applies predistortion before the precoder in the OFDM guard-band subcarriers. Using simulation results, we show that, for a 64 antenna massive MIMO array, our proposed method can achieve the same DPD performance as a conventional DPD method while requiring an order of magnitude fewer multiplications. Chance Tarver, Alexios Balatsoukas-Stimming, Christoph Studer, Joseph R. Cavallaro |
ISCAS | 4 |
| 2021 | Radio Frequency Fingerprint Identification for Narrowband Systems, Modelling and ClassificationabstractDevice authentication is essential for securing Internet of things. Radio frequency fingerprint identification (RFFI) is an emerging technique that exploits intrinsic and unique hardware impairments as the device identifier. The existing RFFI literature focuses on experimental exploration but comprehensive modelling is missing. This paper systematically models impairments of transmitter and receiver in narrowband systems and carries out extensive experiments and simulations to evaluate their effects on RFFI. The modelled impairments include oscillator imperfections, imbalance of inphase (I) and quadrature (Q) branches of mixers and power amplifier (PA) nonlinearity. We then propose a convolutional neural network-based RFFI protocol. We carry out experimental measurements over three months and demonstrate that oscillator imperfections are not suitable for RFFI due to their unpredictable time variation caused by temperature change. Our simulation results show that our protocol can classify 50 and 200 devices with uniformly and randomly distributed IQ imbalances and PA nonlinearities with high accuracy, namely 99% and 89%, respectively. We also show that the RFFI has some tolerance on different receiver imbalances during training and classification. Specifically, the accuracy is shown to degrade less than 20% when the residual receiver's gain and phase imbalances are small. Based on the experimental and simulation results, we made recommendations for designing a robust RFFI protocol, namely compensate carrier frequency offset and calibrate IQ imbalances of receivers. Junqing Zhang, Roger F. Woods, Magnus Sandell, Mikko Valkama, Alan Marshall 0001, Joseph R. Cavallaro |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2020 | Gbit/s Non-Binary LDPC Decoders: High-Throughput using High-Level SpecificationsabstractIt is commonly perceived that an HLS specification targeted for FPGAs cannot provide throughput performance in par with equivalent RTL descriptions. In this work we developed a complex design of a non-binary LDPC decoder, that although hard to generalise, shows that HLS provides sufficient architectural refinement options. They allow attaining performance above CPU- and GPU-based ones and excel at providing a faster design cycle when compared to RTL development. Oscar Ferraz, Srinivasan Subramaniyan, Joseph R. Cavallaro, Gabriel Falcão Paiva Fernandes, Madhura Purnaprajna |
FCCM | 4 |
| 2020 | GPU-Based LDPC Decoding for vRAN Systems in 5G and BeyondabstractNext-generation virtual radio access networks (vRAN) will benefit from the flexibility provided by virtualization in proposed Cloud-RAN configurations. These systems for 5G and beyond may consist of commodity hardware such as GPUs in data centers with multiple connected base stations (gNBs) flexibly receiving allocated resources depending on time-varying, real-time demands. In this paper, parallel reconfigurable algorithms and architectures for channel decoding are proposed. In particular, flexible rate and block length LDPC decoders for the new radio (NR) physical layer on GPU are characterized. We implement these GPU decoders using reduced word lengths of 8-bits to represent the log-likelihood ratios during decoding, and we utilize multiple GPU streams to process multiple blocks of codewords in parallel. These techniques allow our implementation to reduce the device transfer overhead and achieve the low-latency or high-throughput targets for 5G and beyond. Moreover, we integrate our decoder into the Open Air Interface (OAI) NR software stack to investigate virtualization capabilities when containerizing vRAN functionality such as the LDPC decoder. Chance Tarver, Matthew Jordan Tonnemacher, Hao Chen 0010, Jianzhong Zhang 0002, Joseph R. Cavallaro |
ISCAS | 5 |
| 2019 | Decentralized Coordinate-Descent Data Detection and Precoding for Massive MU-MIMOabstractMassive multiuser (MU) multiple-input multiple-output (MIMO) promises significant improvements in spectral efficiency compared to small-scale MIMO. Typical massive MU-MIMO base-station (BS) designs rely on centralized linear data detectors and precoders which entail excessively high complexity, interconnect data rates, and chip input/output (I/O) bandwidth when executed on a single computing fabric. To resolve these complexity and bandwidth bottlenecks, we propose new decentralized algorithms for data detection and precoding that use coordinate descent. Our methods parallelize computations across multiple computing fabrics, while minimizing interconnect and I/O bandwidth. The proposed decentralized algorithms achieve near-optimal error-rate performance and multi-Gbps throughput at sub-1 ms latency when implemented on a multi-GPU cluster with half-precision floating-point arithmetic. Kaipeng Li 0003, Oscar Castañeda, Charles Jeon, Joseph R. Cavallaro, Christoph Studer |
ISCAS | 4 |
| 2018 | Energy-efficient Convolutional Neural Networks via Statistical Error Compensated Near Threshold ComputingabstractThere has been a growing need for deploying machine learning algorithms such as convolutional neural networks (CNNs) on resource-constrained edge platforms to enable on-device local inference. Despite CNNs' excellent performance that approaches and sometimes exceeds humans in a large variety of tasks, their often prohibitive complexity remains a major inhibitor. To address the energy challenge, near threshold computing (NTC) has been proposed to aggressively reduce energy consumption, at the cost of increased performance variation due to circuit level statistical behavior. In this paper, we propose a variation-tolerant architecture for CNNs capable of robust operations in the NTC regime for energy efficiency. Specifically, we construct robust CNNs from two low-cost unreliable designs that have different error statistics: a NTC design with full precision, and a K-means approximated design where weight vectors in the CNN are clustered to reduce complexity. When evaluated in CNNs using the MNIST dataset, simulation results in 45 nm CMOS show that the proposed architecture enables robust CNNs operating in the NTC regime. Specifically, the proposed CNN can enhance variation tolerance by 10× and achieve up to 134× reduction in the standard deviation of inference accuracy Pdet while incurring marginal degradation in the median inference accuracy. Yingyan (Celine) Lin, Joseph R. Cavallaro |
ISCAS | 2 |
| 2017 | Multi component carrier, sub-band DPD and GNURadio implementationabstractDigital predistortion (DPD) is an effective way of mitigating spurious emission violations without the need of a significant backoff in the transmitter, thus providing better power efficiency and network coverage. In this paper, the IM3 subband DPD, proposed earlier by the authors, is extended to more than two component carriers (CCs) through a sequential learning solution. The DPD learning is iterated over each spurious emission generated by each pair and trio of CCs. We train and apply the DPD coefficients for the intermodulation distortion (IMD) products until a satisfactory performance is achieved. The algorithm is tested in simulations using MATLAB and in a novel, real-time implementation on a CPU via a software version of the algorithm using GNURadio. Chance Tarver, Mahmoud Abdelaziz, Lauri Anttila, Joseph R. Cavallaro |
ISCAS | 4 |
| 2017 | On the achievable rates of decentralized equalization in massive MU-MIMO systemsabstractMassive multi-user (MU) multiple-input multiple-output (MIMO) promises significant gains in spectral efficiency compared to traditional, small-scale MIMO technology. Linear equalization algorithms, such as zero forcing (ZF) or minimum mean-square error (MMSE)-based methods, typically rely on centralized processing at the base station (BS), which results in (i) excessively high interconnect and chip input/output data rates, and (ii) high computational complexity. In this paper, we investigate the achievable rates of decentralized equalization that mitigates both of these issues. We consider two distinct BS architectures that partition the antenna array into clusters, each associated with independent radio-frequency chains and signal processing hardware, and the results of each cluster are fused in a feed forward network. For both architectures, we consider ZF, MMSE, and a novel, non-linear equalization algorithm that builds upon approximate message passing (AMP), and we theoretically analyze the achievable rates of these methods. Our results demonstrate that decentralized equalization with our AMP-based methods incurs no or only a negligible loss in terms of achievable rates compared to that of centralized solutions. Charles Jeon, Kaipeng Li 0003, Joseph R. Cavallaro, Christoph Studer |
ISIT | 3 |
| 2016 | FPGA design of a coordinate descent data detector for large-scale MU-MIMOabstractWe propose a new, low-complexity data-detection algorithm and a corresponding high-throughput FPGA design for 3GPP LTE-based large-scale (or massive) multi-user (MU) multiple-input multiple-output (MIMO) wireless communication systems. Our algorithm performs approximate minimum mean-square error (MMSE) data detection using coordinate descent (CD), which enables near-MMSE performance at low computational complexity, even for systems with hundreds of antennas at the base station (BS). We design a high-throughput VLSI architecture for 3GPP LTE wideband systems with a deep and interleaved pipeline, which can be parametrized at design time to support various antenna configurations. Our CD-based data detector achieves 379Mb/s throughout, while using 24 k LUTs and 771 DSP units on a Xilinx Virtex-7 FPGA for a 128 BS antenna, 8 user large-scale MU-MIMO system. Michael Wu 0001, Chris Dick, Joseph R. Cavallaro, Christoph Studer |
ISCAS | 3 |
| 2015 | Robust Consensus-Based Cooperative Spectrum Sensing under Insistent Spectrum Sensing Data Falsification AttacksabstractIn this paper, we introduce Insistent Spectrum Sensing Data Falsification (ISSDF) as a new practical and destructive attack model aimed at distributed cooperative spectrum sensing schemes that are based on iterative average consensus. We compare various linear iteration-based and iterative gossip-based schemes in terms of primary user detection performance and convergence speed under this attack. Moreover, we devise a trust management scheme to mitigate the attack and we propose a practical trust-aware consensus-based scheme for distributed cooperative spectrum sensing which is resilient to ISSDF. Finally, we quantify the performance improvement due to trust management through extensive simulations. Aida Vosoughi, Joseph R. Cavallaro, Alan Marshall 0001 |
GLOBECOM | 2 |
| 2015 | Scale- and orientation-invariant keypoints in higher-dimensional dataabstractDescription of keypoints, or local image features, is widely employed in computer vision. However, the most successful techniques do not extend immediately to more than two spatial dimensions. In this paper, we describe robust methods for extracting local orientations and gradient histograms from higher-dimensional data, using these techniques to develop a three-dimensional analogue of the popular Scale-Invariant Feature Transform (SIFT). We apply our algorithm to intra-patient registration of magnetic resonance (MR) images, with promising results. Our implementation will be released as open-source software. Blaine Rister, Daniel Reiter, Daniel Volz, Mark Horowitz, Refaat E. Gabr, Joseph R. Cavallaro |
ICIP | 7 |
| 2015 | VLSI design of large-scale soft-output MIMO detection using conjugate gradientsabstractWe propose an FPGA design for soft-output data detection in orthogonal frequency-division multiplexing (OFDM)-based large-scale (multi-user) MIMO systems. To reduce the high computational complexity of data detection, our design uses a modified version of the conjugate gradient least square (CGLS) algorithm. In contrast to existing linear detection algorithms for massive MIMO systems, our method avoids two of the most complex tasks, namely Gram-matrix computation and matrix inversion, while still being able to compute soft-outputs. Our architecture uses an array of reconfigurable processing elements to compute the CGLS algorithm in a hardware-efficient manner. Implementation results on Xilinx Virtex-7 FPGA for a 128 antenna, 8 user large-scale MIMO system show that our design only uses 70% of the area-delay product of the competitive method, while exhibiting superior error-rate performance. Bei Yin, Michael Wu 0001, Joseph R. Cavallaro, Christoph Studer |
ISCAS | 3 |
| 2014 | Conjugate gradient-based soft-output detection and precoding in massive MIMO systemsabstractMassive multiple-input multiple-output (MIMO) promises improved spectral efficiency, coverage, and range, compared to conventional (small-scale) MIMO wireless systems. Unfortunately, these benefits come at the cost of significantly increased computational complexity, especially for systems with realistic antenna configurations. To reduce the complexity of data detection (in the uplink) and precoding (in the downlink) in massive MIMO systems, we propose to use conjugate gradient (CG) methods. While precoding using CG is rather straightforward, soft-output minimum mean-square error (MMSE) detection requires the computation of the post-equalization signal-to-interference-and-noise-ratio (SINR). To enable CG for soft-output detection, we propose a novel way of computing the SINR directly within the CG algorithm at low complexity. We investigate the performance/complexity trade-offs associated with CG-based soft-output detection and precoding, and we compare it to existing exact and approximate methods. Our results reveal that the proposed algorithm is able to outperform existing methods for massive MIMO systems with realistic antenna configurations. Bei Yin, Michael Wu 0001, Joseph R. Cavallaro, Christoph Studer |
GLOBECOM | 3 |
| 2014 | Low power implementation of digital predistortion filter on a heterogeneous application specific multiprocessorabstractPower-constrained mobile radio communication transmitters drive their transmit power amplifiers close to their saturation regions, which results in nonlinear intermodulation distortion that is especially harmful in multi-cluster and carrier aggregation transmission scenarios. Digital predistortion is a method for linearizing the transmitter and suppressing the most harmful spurious emissions at the transmitter power amplifier output. This paper describes a programmable implementation of a digital predistortion filter on a heterogeneous Transport Trigger Architecture (TTA) multiprocessor. The predistortion algorithm is based on a parallel Hammerstein polynomial model and the experimental results show that the proposed programmable architecture is capable of linearizing a 20 MHz LTE carrier in realtime with a power consumption that is suitable for mobile devices. Amanullah Ghazi, Jani Boutellier, Mahmoud Abdelaziz, Xiaojia Lu, Lauri Anttila, Joseph R. Cavallaro, Shuvra S. Bhattacharyya, Mikko Valkama, Markku Juntti |
ICASSP | 6 |
| 2014 | Parallel programming of a symmetric transport-triggered architecture with applications in flexible LDPC encodingabstractExposed-datapath architectures yield small, low-power processors that trade instruction word length for aggressive compile-time scheduling and a high degree of instruction-level parallelism. In this paper, we present a general-purpose parallel accelerator consisting of a main processor and eight symmetric clusters, all in a single core. Use of a lightweight and memory-efficient application programming interface allows for the first high-performance program executing both sequential and data-parallel code on the same TTA processor. We use the processor for LDPC encoding, a popular method of forward error correction. Demonstrating the flexibility of software-defined radio, we benchmark the processor with two programs, one which can handle almost any sort of LDPC code, and another which is optimized for a specific standard. We achieve a throughput of 5 Mb/s with the flexible program and 92 Mb/s with the standard-specific one, while consuming only 95 mW at a clock frequency of 1175 MHz. Blaine Rister, Pekka Jääskeläinen, Olli Silvén, Jari Hannuksela, Joseph R. Cavallaro |
ICASSP | 5 |
| 2014 | Efficient architecture mapping of FFT/IFFT for cognitive radio networksabstractCognitive radio networks require flexibility to support a variety of wireless communication system standards. Many modern systems utilize some form of orthogonal frequency division multiplexing (OFDM) and single-carrier frequency-division multiple access (SC-FDMA) often augmented with multiple input multiple output (MIMO) antenna schemes. A common module in these standards is the fast Fourier transform (FFT) and its inverse. Although many architectures exist for traditional power-of-two FFT lengths, the recent 3GPP LTE standards define non-power-of-two transform lengths. The various FFT and IFFT lengths for both the uplink and downlink processing require support for radix-2, radix-3, and radix-5 modules. In this paper, we propose a highly flexible FFT/IFFT architecture that can support a broad variety of transform sizes and efficient mapping to programmable testbed platforms for cognitive radio networks. This novel architecture will provide a range of transform sizes of the general form (2n3k5l), and for use in emerging algorithms for massive MIMO detectors. Bei Yin, Inkeun Cho, Joseph R. Cavallaro, Shuvra S. Bhattacharyya, Jarmo Takala |
ICASSP | 4 |
| 2014 | A 3.8Gb/s large-scale MIMO detector for 3GPP LTE-AdvancedabstractThis paper proposes - to the best of our knowledge - the first ASIC design for high-throughput data detection in single carrier frequency division multiple access (SC-FDMA)-based large-scale MIMO systems, such as systems building on future 3GPP LTE-Advanced standards. In order to substantially reduce the complexity of linear soft-output data detection in systems having hundreds of antennas at the base station (BS), the proposed detector builds upon a truncated Neumann series expansion to compute the necessary matrix inverse at low complexity. To achieve high throughput in the 3GPP LTE-A uplink, we develop a systolic VLSI architecture including all necessary processing blocks. We present a corresponding ASIC design that achieves 3.8 Gb/s for a 128 antenna, 8 user 3GPP LTE-A based large-scale MIMO system, while occupying 11.1 mm2in a TSMC 45nm CMOS technology. Bei Yin, Michael Wu 0001, Chris Dick, Joseph R. Cavallaro, Christoph Studer |
ICASSP | 5 |
| 2013 | Highly scalable on-the-fly interleaved address generation for UMTS/HSPA+ parallel turbo decoderabstractHigh throughput parallel interleaver design is a major challenge in designing parallel turbo decoders that conform to high data rate requirements of advanced standards such as HSPA+. The hardware complexity of the HSPA+ interleaver makes it difficult to scale to high degrees of parallelism. We propose a novel algorithm and architecture for on-the-fly parallel interleaved address generation in UMTS/HSPA+ standard that is highly scalable. Our proposed algorithm generates an interleaved memory address from an original input address without building the complete interleaving pattern or storing it; the generated interleaved address can be used directly for interleaved writing to memory blocks. We use an extended Euclidean algorithm for modular multiplicative inversion as a step towards reversed intra-row permutations in UMTS/HSPA+ standard. As a result, we can determine interleaved addresses from original addresses. We also propose an efficient and scalable hardware architecture for our method. Our design generates 32 interleaved addresses in one cycle and satisfies the data rate requirement of 672 Mbps in HSPA+ while the silicon area and frequency is improved compared to recent related works. Aida Vosoughi, Hao Shen 0013, Joseph R. Cavallaro, Yuanbin Guo |
ASAP | 4 |
| 2013 | High-throughput beamforming receiver for millimeter wave mobile communicationabstractIn this paper, we present a novel FPGA-based high-throughput beamforming MIMO receiver for millimeter wave mobile communication. With vast spectrum and small antenna element size, millimeter wave communication becomes very attractive and promising to support next generation mobile communication (5G). However, the high data rate requirement challenges both algorithm and architecture. In order to support the high data rate and to reduce the overhead of selecting the best beam pair, we propose a novel beamforming synchronization scheme more suitable for mobile communication. By further optimizing the algorithm and the architecture, we present a complete mobile receiver based on FPGA, which includes RF frontend, ADC, beamforming control, synchronization, channel estimator, soft MAP detector, and channel decoder. The design operates at 28 GHz carrier frequency with 500 MHz bandwidth. The throughput can reach 1.52 Gbps. We also performed the indoor and outdoor over-the-air transmission field tests. This work provides a platform for future millimeter wave mobile communication research. Bei Yin, Shadi Abu-Surra, Gary Xu, Thomas Henige, Eran Pisek, Zhouyue Pi, Joseph R. Cavallaro |
GLOBECOM | 7 |
| 2013 | A fast and efficient sift detector using the mobile GPUabstractEmerging mobile applications, such as augmented reality, demand robust feature detection at high frame rates. We present an implementation of the popular Scale-Invariant Feature Transform (SIFT) feature detection algorithm that incorporates the powerful graphics processing unit (GPU) in mobile devices. Where the usual GPU methods are inefficient on mobile hardware, we propose a heterogeneous dataflow scheme. By methodically partitioning the computation, compressing the data for memory transfers, and taking into account the unique challenges that arise out of the mobile GPU, we are able to achieve a speedup of 4-7x over an optimized CPU version, and a 6.4x speedup over a published GPU implementation. Additionally, we reduce energy consumption by 87 percent per image. We achieve near-realtime detection without compromising the original algorithm. Blaine Rister, Michael Wu 0001, Joseph R. Cavallaro |
ICASSP | 4 |
| 2013 | Accelerating computer vision algorithms using OpenCL framework on the mobile GPU - A case studyabstractRecently, general-purpose computing on graphics processing units (GPGPU) has been enabled on mobile devices thanks to the emerging heterogeneous programming models such as OpenCL. The capability of GPGPU on mobile devices opens a new era for mobile computing and can enable many computationally demanding computer vision algorithms on mobile devices. As a case study, this paper proposes to accelerate an exemplar-based inpainting algorithm for object removal on a mobile GPU using OpenCL. We discuss the methodology of exploring the parallelism in the algorithm as well as several optimization techniques. Experimental results demonstrate that our optimization strategies for mobile GPUs have significantly reduced the processing time and make computationally intensive computer vision algorithms feasible for a mobile device. To the best of the authors' knowledge, this work is the first published implementation of general-purpose computing using OpenCL on mobile GPUs. Yingen Xiong, Jay Yun, Joseph R. Cavallaro |
ICASSP | 4 |
| 2013 | Implementation trade-offs for linear detection in large-scale MIMO systemsabstractIn this paper, we analyze the VLSI implementation tradeoffs for linear data detection in the uplink of large-scale multiple-input multiple-output (MIMO) wireless systems. Specifically, we analyze the error incurred by using the sub-optimal, low-complexity matrix inverse proposed in Wu et al., 2013, ISCAS, and compare its performance and complexity to an exact matrix inversion algorithm. We propose a Cholesky-based reference architecture for exact matrix inversion and show corresponding implementation results on an Virtex-7 FPGA. Using this reference design, we perform a performance/complexity trade-off comparison with an FPGA implementation for the proposed approximate matrix inversion, which reveals that the inversion circuit of choice is determined by the antenna configuration (base-station antennas vs. number of users) of large-scale MIMO systems. Bei Yin, Michael Wu 0001, Christoph Studer, Joseph R. Cavallaro, Chris Dick |
ICASSP | 4 |
| 2013 | Parallel interleaver architecture with new scheduling scheme for high throughput configurable turbo decoderabstractParallel architecture is required for high throughput turbo decoder to meet the data rate requirements of the emerging wireless communication systems. However, due to the severe memory conflict problem caused by parallel architectures, the interleaver design has become a major challenge that limits the achievable throughput. Moreover, the high complexity of the interleaver algorithm makes the parallel interleaving address generation hardware very difficult to implement. In this paper, we propose a parallel interleaver architecture that can generate multiple interleaving addresses on-the-fly. We devised a novel scheduling scheme with which we can use more efficient buffer structures to eliminate memory contention. The synthesis results show that the proposed architecture with the new scheduling scheme can significantly reduce memory usage and hardware complexity. The proposed architecture also shows great flexibility and scalability compared to prior work. Aida Vosoughi, Hao Shen 0013, Joseph R. Cavallaro, Yuanbin Guo |
ISCAS | 4 |
| 2013 | Approximate matrix inversion for high-throughput data detection in the large-scale MIMO uplinkabstractThe high processing complexity of data detection in the large-scale multiple-input multiple-output (MIMO) uplink necessitates high-throughput VLSI implementations. In this paper, we propose - to the best of our knowledge - first matrix inversion implementation suitable for data detection in systems having hundreds of antennas at the base station (BS). The underlying idea is to carry out an approximate matrix inversion using a small number of Neumann-series terms, which allows one to achieve near-optimal performance at low complexity. We propose a novel VLSI architecture to efficiently compute the approximate inverse using a systolic array and show reference FPGA implementation results for various system configurations. For a system where 128 BS antennas receive data from 8 single-antenna users, a single instance of our design processes 1.9M matrices/s on a Xilinx Virtex-7 FPGA, while using only 3.9% of the available slices and 3.6% of the available DSP48 units. Michael Wu 0001, Bei Yin, Aida Vosoughi, Christoph Studer, Joseph R. Cavallaro, Chris Dick |
ISCAS | 5 |
| 2013 | VLSI Architecture for Layered Decoding of QC-LDPC Codes With High Circulant WeightabstractIn this brief, we propose a high-throughput layered decoder architecture to support a broader family of quasicyclic low-density parity-check (QC-LDPC) codes, whose parity-check matrices are constructed from arrays of circulant submatrices. Each nonzero circulant submatrix is a superposition of K cyclic-shifted identity matrices, where the circulant weight K ≥ 1. We propose a novel layered decoder architecture to support QC-LDPC codes with any circulant weight. We present a block-serial decoding architecture which processes a layer of a parity check matrix block by block, where each block is a Z×Z circulant matrix with a circulant weight of K. In the case study, we demonstrate an LDPC decoder design for the China Mobile Multimedia Broadcasting (CMMB) standard, which was synthesized for a TSMC 65-nm CMOS technology. With a core area of 3.9 mm2, the CMMB LDPC decoder achieves a maximum throughput of 1.1 Gb/s with 15 iterations. Yang Sun 0001, Joseph R. Cavallaro |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Baseband signal compression in wireless base stationsabstractTo comply with the evolving wireless standards, base stations must provide greater data rates over the serial data link between base station processor and RF unit. This link is especially important in distributed antenna systems and cooperating base stations settings. This paper explores the compression of baseband signal samples prior to transfer over the above-mentioned link. We study lossy and lossless compression of baseband signals and analyze the cost and gain of each approach. Sample quantizing is proposed as a lossy compression scheme and it is shown to be effective by experiments. With QPSK modulation, sample quantizing achieves a compression ratio of 4:1 and 3.5:1 in downlink and uplink, respectively. The corresponding compression ratios are 2.3:1 and 2:1 for 16-QAM. In addition, lossless compression algorithms including arithmetic coding, Elias-gamma coding, and unused significant bit removal, and also a recently proposed baseband signal compression scheme are evaluated. The best compression ratio achieved for lossless compression is 1.5:1 in downlink. Our simulations and over-the-air experiments suggest that compression of baseband signal samples is a feasible and promising solution for increasing the effective bit rates of the link to/from remote RF units without requiring much complexity and cost to the base station. Aida Vosoughi, Michael Wu 0001, Joseph R. Cavallaro |
GLOBECOM | 3 |
| 2012 | Reconfigurable multi-standard uplink MIMO receiver with partial interference cancellationabstractAs HSPA/HSPA+ and LTE/LTE-A evolve in parallel, the reconfigurability of a receiver to support multiple standards has become more and more important, especially for small cells. In this paper, we first suggest a reconfigurable multistandard uplink MIMO receiver based on a frequency domain equalizer. Then, to improve the performance, we propose two low-complexity partial iterative interference cancellation (IC) schemes to deal with the residual inter-chip and inter-antenna interference in HSPA/HSPA+ and the residual inter-symbol and inter-antenna interference in LTE/LTE-A. Compared with a receiver consisting of separate HSPA/HSPA+ and LTE/LTE-A uplink receivers, this reconfigurable receiver can save up to 66.9% complexity. Moreover, the two partial IC schemes have negligible performance loss compared with full IC scheme. They can achieve 2 dB gains in both standards with only 15.2% additional complexity to no IC scheme. Bei Yin, Kiarash Amiri, Joseph R. Cavallaro, Yuanbin Guo |
ICC | 3 |
| 2012 | High-Throughput Soft-Output MIMO Detector Based on Path-Preserving Trellis-Search AlgorithmabstractIn this paper, we propose a novel path-preserving trellis-search (PPTS) algorithm and its high-speed VLSI architecture for soft-output multiple-input-multiple-output (MIMO) detection. We represent the search space of the MIMO signal with an unconstrained trellis, where each node in stagekof the trellis maps to a possible complex-valued symbol transmitted by antennak. Based on the trellis model, we convert the soft-output MIMO detection problem into a multiple shortest paths problem subject to the constraint that every trellis node must be covered in this set of paths. The PPTS detector is guaranteed to have soft information for every possible symbol transmitted on every antenna so that the log-likelihood ratio (LLR) for each transmitted data bit can be more accurately formed. Simulation results show that the PPTS algorithm can achieve near-optimal error performance with a low search complexity. The PPTS algorithm is a hardware-friendly data-parallel algorithm because the search operations are evenly distributed among multiple trellis nodes for parallel processing. As a case study, we have designed and synthesized a fully-parallel systolic-array detector and two folded detectors for a 4 × 4 16-QAM system using a 1.08 V TSMC 65-nm CMOS technology. With a 1.18 mm2core area, the folded detector can achieve a throughput of 2.1 Gbps. With a 3.19 mm2core area, the fully-parallel systolic-array detector can achieve a throughput of 6.4 Gbps. Yang Sun 0001, Joseph R. Cavallaro |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | High-throughput Contention-Free concurrent interleaver architecture for multi-standard turbo decoderabstractTo meet the higher data rate requirement of emerging wireless communication technology, numerous parallel turbo decoder architectures have been developed. However, the interleaver has become a major bottleneck that limits the achievable throughput in the parallel decoders due to the massive memory conflicts. In this paper, we propose a flexible Double-Buffer based Contention-Free (DBCF) interleaver architecture that can efficiently solve the memory conflict problem for parallel turbo decoders with very high parallelism. The proposed DBCF architecture enables high throughput concurrent interleaving for multi-standard turbo decoders that support UMTS/HSPA+, LTE and WiMAX, with small datapath delays and low hardware cost. We implemented the DBCF interleaver with a 65nm CMOS technology. The implementation of this highly efficient DBCF interleaver architecture shows significant improvement in terms of the maximum throughput and occupied chip area compared to the previous work. Yang Sun 0001, Joseph R. Cavallaro, Yuanbin Guo |
ASAP | 3 |
| 2011 | Multi-layer parallel decoding algorithm and vlsi architecture for quasi-cyclic LDPC codesabstractWe propose a multi-layer parallel decoding algorithm and VLSI architecture for decoding of structured quasi-cyclic low-density parity-check codes. In the conventional layered decoding algorithm, the block-rows of the parity check matrix are processed sequentially, or layer after layer. The maximum number of rows that can be simultaneously processed by the conventional layered decoder is limited to the sub-matrix size. To remove this limitation and support layer-level parallelism, we extend the conventional layered decoding algorithm and architecture to enable simultaneously processing of multiple (K) layers of a parity check matrix, which will lead to a roughly K-fold throughput increase. As a case study, we have designed a double-layer parallel LDPC decoder for the IEEE 802.11n standard. The decoder was synthesized for a TSMC 45-nm CMOS technology. With a synthesis area of 0.81 mm2and a maximum clock frequency of 815 MHz, the decoder achieves a maximum throughput of 3.0 Gbps at 15 iterations. Yang Sun 0001, Joseph R. Cavallaro |
ISCAS | 3 |
| 2011 | Efficient hardware implementation of a highly-parallel 3GPP LTE/LTE-advance turbo decoder
Yang Sun 0001, Joseph R. Cavallaro |
Integr. | 2 |
| 2011 | Architecture Design and Implementation of the Metric First List Sphere Detector AlgorithmabstractSoft-output detection of a multiple-input-multiple-output (MIMO) signal pose a significant challenge in future wireless systems. In this paper, we introduce a soft-output modified metric first (MMF)-LSD algorithm for MIMO detection. We design a scalable architecture and address a method to decrease memory requirements. We provide implementation results for a spatial multiplexing (SM) system with four transmitted streams and with 16- and 64-quadrature amplitude modulation (QAM) on a 0.18-μ m CMOS application specific integrated circuit (ASIC) technology. The MFF-LSD implementation is more efficient than the depth first (DF)-LSD in the crucial low signal-to-noise rate (SNR) region and the detection rate of the 64-QAM implementation is 39.2 Mbps@26 db with 48.2 kGEs complexity. Markus Myllylä, Joseph R. Cavallaro, Markku Juntti |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Programming high performance signal processing systems in high level languagesabstractNo abstract available. Kees A. Vissers, Devada Varma, Vinod Kathail, Jeff Bier, Don MacMillen, Joseph R. Cavallaro |
FPGA | 6 |
| 2010 | Physical layer algorithm and hardware verification of MIMO relays using cooperative partial detectionabstractCooperative communication with multi-antenna relays can significantly increase the reliability and speed. However, cooperative MIMO detection would impose considerable complexity overhead onto the relay if a full detect-and-forward (FDF) strategy is employed. In order to address this challenge, we propose a novel cooperative partial detection (CPD) strategy to partition the detection task between the relay and the destination. CPD utilizes the inherent structure of the tree-based sphere detectors, and modifies the tree traversal so that instead of visiting all the levels of the tree, only a subset of the levels, thus a subset of the transmitted streams, are visited. Based on this methodology, the destination combines the source signal and the partial relay signal to perform the detection step. We show, in both simulation and hardware verification on the WARP platform, that using the CPD approach, the relay can avoid the considerable overhead of MIMO detection while helping the source-destination link to improve its performance. Kiarash Amiri, Michael Wu 0001, Melissa Duarte, Joseph R. Cavallaro |
ICASSP | 4 |
| 2010 | Low-complexity and high-performance soft MIMO detection based on distributed M-algorithm through trellis-diagramabstractThis paper presents a novel low-complexity multiple-input multiple-output (MIMO) detection scheme using a distributed M-algorithm (DM) to achieve high performance soft MIMO detection. To reduce the searching complexity, we build a MIMO trellis graph and split the searching operations among different nodes, where each node will apply the M-algorithm. Instead of keeping a global candidate list as the traditional detector does, this algorithm keeps multiple small candidate lists to generate soft information. Since the DM algorithm can achieve good BER performance with a small M, the sorting cost of the DM algorithm is lower than that of the conventional K-best MIMO algorithm. The proposed algorithm is very suitable for high speed parallel processing. Yang Sun 0001, Joseph R. Cavallaro |
ICASSP | 2 |
| 2010 | Implementation aspects of list sphere decoder algorithms for MIMO-OFDM systems
Markus Myllylä, Markku Juntti, Joseph R. Cavallaro |
Signal Process. | 3 |
| 2009 | High throughput VLSI architecture for soft-output mimo detection based on a greedy graph algorithmabstractMaximum-likelihood (ML) decoding is a very computational-intensive task for multiple-input multiple-output (MIMO) wireless channel detection. This paper presents a new graph based algorithm to achieve near ML performance for soft MIMO detection. Instead of using the traditional tree search based structure, we represent the search space of the MIMO signals with a directed graph and a greedy algorithm is applied to compute the a posteriori probability (APP) for each transmitted bit. The proposed detector has two advantages: 1) it keeps a fixed throughput and has a regular and parallel datapath structure which makes it amenable to high speed VLSI implementation, and 2) it attempts to maximize the a posteriori probability by making the locally optimum choice at each stage with the hope of finding the global minimum Euclidean distance for every transmitted bit. Compared to the soft K-best detector, the proposed solution significantly reduces the complexity because sorting is not required, while still maintaining good bit error rate (BER) performance. The proposed greedy detection algorithm has been designed and synthesized for a 4 x 4 16-QAM MIMO system in a TSMC 65 nm CMOS technology. The detector achieves a maximum throughput of 600 Mbps with a 0.79 mm2 core area. Yang Sun 0001, Joseph R. Cavallaro |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | Architecture design and implementation of the increasing radius - List sphere detector algorithmabstractA list sphere detector (LSD) is an enhancement of a sphere detector (SD) that can be used to approximate the optimal MAP detector. In this paper, we introduce a novel architecture for the increasing radius (IR)-LSD algorithm, which is based on the Dijkstra's algorithm. The parallelism possibilities are introduced in the presented architecture, which is also scalable for different multiple-input multiple-output (MIMO) systems. The novel architecture is implemented on a Virtex-IV field programmable gate array (FPGA) chip using high-level ANSI C++ language based Catapult C Synthesis tool from Mentor Graphics. The used word lengths, the latency of the design, and the required resources are presented and analyzed for 4 times 4 MIMO system with 16- quadrature amplitude modulation (QAM). The detector implementation achieves a maximum throughput of 12.1Mbps at high signal-to-noise ratio (SNR). Markus Myllylä, Markku Juntti, Joseph R. Cavallaro |
ICASSP | 3 |
| 2009 | Probabilistically bounded soft sphere detection for MIMO-OFDM receivers: algorithm and system architectureabstractIterative soft detection and channel decoding for MIMO OFDM downlink receivers is studied in this work. Proposed inner soft sphere detection employs a variable upper bound for number of candidates per transmit antenna and utilizes the breath-first candidate-search algorithm. Upper bounds are based on probability distribution of the number of candidates found inside the spherical region formed around the received symbolvector. Detection accuracy of unbounded breadth-first candidatesearch is preserved while significant reduction of the search latency and area cost is achieved. This probabilistically bounded candidate-search algorithm improves error-rate performance of non-probabilistically bounded soft sphere detection algorithms, while providing smaller detection latency with same hardware resources. Prototype architecture of soft sphere detector is synthesized on Xilinx FPGA and for an ASIC design. Using area-cost of a single soft sphere detector, a level of processing parallelism required to achieve targeted high data rates for future wireless systems (for example, 1 Gbps data rate) is determined. Predrag Radosavljevic, Yuanbin Guo, Joseph R. Cavallaro |
IEEE J. Sel. Areas Commun. | 3 |
| 2008 | Configurable and scalable high throughput turbo decoder architecture for multiple 4G wireless standardsabstractIn this paper, we propose a novel multi-code turbo decoder architecture for 4G wireless systems. To support various 4G standards, a configurable multi-mode MAP (maximum a posteriori) decoder is designed for both binary and duo-binary turbo codes with small resource overhead (less than 10%) compared to the single-mode architecture. To achieve high data rates in 4G, we present a parallel turbo decoder architecture with scalable parallelism tailored to the given throughput requirements. High-level parallelism is achieved by employing contention-free interleavers. Multi-banked memory structure and routing network among memories and MAP decoders are designed to operate at full speed with parallel interleavers. We designed a very low-complexity recursive on-line address generator supporting multiple interleaving patterns, which avoids the interleaver address memory. Design trade-offs in terms of area and power efficiency are explored to find the optimal architectures. A 711 Mbps data rate is feasible with 32 Radix-4 MAP decoders running at 200 MHz clock rate. Yang Sun 0001, Yuming Zhu, Manish Goel, Joseph R. Cavallaro |
ASAP | 4 |
| 2008 | Novel Sort-Free Detector with Modified Real-Valued Decomposition (M-RVD) Ordering in MIMO SystemsabstractK-best MIMO detection technique is the prominent method of simplifying the detection complexity in MIMO systems while maintaining BER performance comparable with the optimum maximum-likelihood (ML) detection technique. However, sorting the candidate nodes in the tree search of the conventional K-best detection can take a significant number of cycles which would reduce the achievable data rate of the detector. In order to reduce this delay, and keep high performance at the same time, we propose using a novel sort-free based MIMO detector which avoids the demanding sorting step. Moreover, this detector utilizes a novel modified real-valued decomposition (M-RVD) ordering that, when compared to the conventional real valued decomposition scheme, can improve the BER performance at no extra computational cost. We show that our proposed detector can outperform the conventional K-best detector with a smaller combination of computation and latency requirements. Kiarash Amiri, Chris Dick, Raghu Mysore Rao, Joseph R. Cavallaro |
GLOBECOM | 4 |
| 2008 | QRD-QLD Searching Based Sphere Detection for Emerging MIMO Downlink OFDM ReceiversabstractIn this paper, a detection algorithm with parallel partial candidate-search algorithm is presented. Two fully independent partial search processes are simultaneously employed for two groups of transmit antennas based on QR and QL decompositions of the channel matrix. Proposed QRD- QLD detection algorithm is compared with well-known QRD-M scheme adopted for several emerging wireless standards. Latency of the QRD-QLD candidate search is about twice as small for similar error-rate performance and for identical hardware resources. Total detection latency of QRD-QLD algorithm that also includes computation of soft information for outer decoder is also substantially smaller. Predrag Radosavljevic, Kyeong Jin Kim, Joseph R. Cavallaro |
GLOBECOM | 3 |
| 2008 | Design of block-structured LDPC codes for iterative receivers with soft sphere detectionabstractIn this paper we design block-structured LDPC codes for iterative MIMO receivers with soft sphere detection in particular channel environments. The receiver EXIT charts are used as a design tool. The main design constraint is to preserve the block-structure of parity-check matrices supporting modular semi-parallel decoder architectures. We show that newly designed block-structured code profiles provide improved error-rate performance of between 0.5 dB and 1 dB in different channel environments for moderate codeword size. Predrag Radosavljevic, Joseph R. Cavallaro |
ICASSP | 2 |
| 2008 | Cooperative Communications Using Scalable, Medium Block-length LDPC CodesabstractCooperative communications has received increasing attention in recent years because of the ubiquity of wireless devices. Each year more mobile computing devices enter the market, most of which have limitations in terms of size, number of antennas and battery power. Cooperation enables these devices to form a virtual multiple antenna system and benefit from diversity. Most of the literature on cooperation is information theoretic with often unrealistic assumptions. In this paper we examine decode and forward scheme from a practical point of view and utilize scalable, architecture-aware low density parity check (LDPC) codes. We present simulation results for two variations of decode and forward strategy and show that even with realistic assumptions about the system, relaying outperforms direct link communications with over 2.5 dB. Marjan Karkooti, Joseph R. Cavallaro |
WCNC | 2 |
| 2007 | Implementation Aspects of List Sphere Detector AlgorithmsabstractA list sphere detector (LSD) can be used to approximate the optimal maximum a posteriori (MAP) detection. The total complexity of the LSD algorithms is relative to the number of visited nodes in the search tree. We compare the differences between real and complex signal model in the LSD algorithm implementation and study its impact on the complexity and performance with different search strategies. In hardware implementation, the number of visited nodes needs to be bounded in order to determine the complexity and the latency of the implementation. Thus, we study the performance of LSD algorithms with a limited number of nodes in the search. We show that the algorithms with real signal model are less complex compared to the complex signal model, and that the performance may suffer significantly with limited search depending on the search strategy. Markus Myllylä, Markku Juntti, Joseph R. Cavallaro |
GLOBECOM | 3 |
| 2007 | VLSI Decoder Architecture for High Throughput, Variable Block-size and Multi-rate LDPC CodesabstractA low-density parity-check (LDPC) decoder architecture that supports variable block sizes and multiple code rates is presented. The proposed architecture is based on the structured quasi-cyclic (QC-LDPC) codes whose performance compares favorably with that of randomly constructed LDPC codes for short to moderate block sizes. The main contribution of this work is to address the variable block-size and multi-rate decoder hardware complexity that stems from the irregular LDPC codes. The overall decoder, which was synthesized, placed and routed on TSMC 0.13-micron CMOS technology with a core area of 4.5 square millimeters, supports variable code lengths from 360 to 4200 bits and multiple code rates between frac14 and 9/10. The average throughput can achieve 1 Gbps at 2.2 dB SNR. Yang Sun 0001, Marjan Karkooti, Joseph R. Cavallaro |
ISCAS | 3 |
| 2007 | A List Sphere Detector based on Dijkstra's Algorithm for MIMO-OFDM SystemsabstractA list sphere detector (LSD) based on Dijkstra's algorithm, namely increasing radius (IR) - LSD, is introduced for multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems. The complexity and the performance of the IR-LSD is compared to a breadth first and depth first search based LSDs. The required list sizes for considered LSDs are determined for a 4times4 system with quadrature amplitude modulations (QAMs). The complexity of the algorithms is relative to the number of visited nodes in the search tree structure. Thus, the algorithm complexities are compared by studying and comparing the number of visited nodes per received symbol vector by the algorithm via computer simulations. The IR-LSD is found to visit the least amount of nodes of the LSDs in the search tree in each studied case, and found the least complex especially with higher order constellations. Markus Myllylä, Markku Juntti, Joseph R. Cavallaro |
PIMRC | 3 |
| 2006 | Configurable, High Throughput, Irregular LDPC Decoder Architecture: Tradeoff Analysis and ImplementationabstractLow Density Parity Check (LDPC) codes are one of the best error correcting codes that enable the future generations of wireless devices to achieve higher data rates. This paper presents a novel flexible decoder architecture for irregular LDPC codes that supports twelve combinations of code lengths -648, 1296, 1944 bits- and code rates- 1/2, 2/3, 3/4, 5/6- based on the IEEE 802.11n standard. All the codes correspond to a block-structured parity check matrix, in which the sub-blocks are either a shifted identity matrix or a zero matrix. A prototype of the LDPC decoder has been implemented and tested on a Xilinx FPGA and has been synthesized for ASIC. Marjan Karkooti, Predrag Radosavljevic, Joseph R. Cavallaro |
ASAP | 3 |
| 2006 | Comparison of Two Novel List Sphere Detector Algorithms for MIMO-OFDM SystemsabstractIn this paper, the complexity and performance of two novel list sphere detector (LSD) algorithms are studied and evaluated in multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) system. The LSDs are based on the K-best and the Schnorr-Euchner enumeration (SEE) algorithms. The required list sizes for LSD algorithms are determined for a 2times2 system with 4-quadrature amplitude modulation (QAM), 16-QAM, and 64-QAM. The complexity of the algorithms is compared by studying the number of visited nodes per received symbol vector by the algorithm in computer simulations. The SEE based LSD algorithm is found to be a less complex and a feasible choice for implementation compared to the K-best based LSD algorithm Markus Myllylä, Pirkka Silvola, Markku Juntti, Joseph R. Cavallaro |
PIMRC | 4 |
| 2006 | Multi-Rate High-Throughput LDPC Decoder: Tradeoff Analysis Between Decoding Throughput and AreaabstractIn order to achieve high decoding throughput (hundreds of MBits/sec and above) for multiple code rates and moderate codeword lengths, several LDPC decoder solutions with different levels of processing parallelism are possible. Selection between these solutions is based on a threefold criterion: hardware complexity, decoding throughput, and error-correcting performance. In this work, we determine the multi-rate LDPC decoder architecture with the best tradeoff in terms of area cost, error-correcting performance, and decoding throughput. The prototype architecture of this decoder is implemented on an FPGA Predrag Radosavljevic, Alexandre de Baynast, Marjan Karkooti, Joseph R. Cavallaro |
PIMRC | 4 |
| 2006 | Special Issue on Reconfigurable Radio Technologies in Support of Ubiquitous Seamless Computing
Panagiotis Demestichas, Guillaume Vivier, Joseph R. Cavallaro |
Mob. Networks Appl. | 3 |
| 2006 | Truncated Online Arithmetic with Applications to Communication SystemsabstractTruncation in digit-precision is a very important and common operation in embedded system design for bounding the required finite precision and for area-time-power savings. In this paper, we present the use of online arithmetic to provide truncated computations with communication systems as one of the applications. In contrast to truncation in conventional arithmetic, online arithmetic can truncate dynamically and produce both area and time benefits due to the digit-serial nature of computations. This is of great advantage in communication systems where the precision requirements can change dynamically with the environment. While truncation in conventional arithmetic can have significant truncation errors, especially when the output precision is less than the input precision, the redundancy and most significant digit first nature of online arithmetic restricts the truncation error to only the least significant digit of the truncated result. As an application that uses significant truncation in precision, a code matched filter detector for wireless systems is designed using truncated online arithmetic. The detector can provide both hard decisions and soft(er) decisions dynamically as well as interface with other conventional arithmetic circuits or act as a DSP coprocessor. Thus, optimized communication receivers with coexisting conventional arithmetic for saturation and online arithmetic for truncation can now be built. The truncated online arithmetic detector was also verified with a VLSI implementation in an AMI 0.5 mu MOSIS tiny chip process Sridhar Rajagopal, Joseph R. Cavallaro |
IEEE Trans. Computers | 2 |
| 2005 | Displacement MIMO Kalman equalizer for CDMA downlink in fast fading channelsabstractIn this paper, a streamlined MIMO Kalman equalizer architecture is proposed to extract the commonality in the data path by jointly considering the displacement structure of the transition matrix and the block-Toeplitz structure of the channel matrix. Finally, an iterative conjugate-gradient based algorithm is proposed to avoid the inverse of the Hermitian symmetric innovation correlation matrix in Kalman gain processor. The proposed architecture not only reduces the numerical complexity to O(F log F) per chip, but also facilitates the parallel and pipelined VLSI implementation in real-time processing. Yuanbin Guo, Jianzhong Zhang 0002, Dennis McCain, Joseph R. Cavallaro |
GLOBECOM | 4 |
| 2005 | FFT-accelerated iterative MIMO chip equalizer architecture for CDMA downlinkabstractIn this paper, we present a novel FFT-accelerated iterative linear MMSE chip equalizer in the MIMO CDMA downlink receiver. The reversed form time-domain matrix multiplication in the conjugate gradient (CG) iteration is accelerated by an equivalent frequency-domain circular convolution with FFT-based "overlap-save" architecture. The iteration rapidly refines a crude initial approximation to the actual final equalizer taps. This avoids the direct-matrix-inverse with O((NL)/sup 3/) complexity, and reduces the standard CG complexity from O((NL)/sup 2/) to O(NLlog/sub 2/(NL)). Simulation demonstrates strong numerical stability and promising performance/complexity tradeoff, especially for very long channels. Yuanbin Guo, Dennis McCain, Joseph R. Cavallaro |
ICASSP (3) | 3 |
| 2005 | Low-complexity iterative multiuser detection and decoding for real-time applicationsabstractThis paper presents a low-complexity multiuser decoding technique that can be implemented in real time for a convolutionally coded direct sequence code division multiple access (DS-CDMA) system. The main contribution, denoted here as the iterative prior update (IPU), consists of iterative interference cancellation and prior updates on sequences of coded bits combined with M-algorithm and list decoding. We illustrate performance gains over other low-complexity sequence detection and decoding strategies and argue that the algorithm converges within a few iterations and requires only a small size buffer for keeping track of the priors along iterations. The fact that the we can use existing available architectures for Viterbi decoding with slight modifications and can meet the real-time processing constraints makes the IPU algorithm an attractive alternative for cellular systems. Elza Erkip, Joseph R. Cavallaro, Behnaam Aazhang |
IEEE Trans. Wirel. Commun. | 3 |
| 2004 | Chip-level LMMSE equalization for downlink MIMO CDMA in fast fading environmentsabstractIn this paper, we consider linear MMSE equalization for wireless downlink transmission with multiple transmit and receive antennas in fast fading environments. We propose a new algorithm based on the conjugate-gradient algorithm with enhanced channel estimation. In order to be robust to the channel variations, the channel coefficients are estimated by using a weighted sliding window. Two methods to determine optimal weights with respect to the Doppler frequency are proposed. The algorithm has been simulated in a fast fading environment (vehicular A with a velocity for the mobile station of 120 km/h). We show by simulations that good performance is obtained in a correlated fast fading environment with reasonable complexity. Moreover, this approach outperforms methods based on basic sliding window or forgetting factor and the LMS algorithm. Alexandre de Baynast, Predrag Radosavljevic, Joseph R. Cavallaro |
GLOBECOM | 3 |
| 2004 | Efficient MIMO equalization for downlink multi-code CDMA: complexity optimization and comparative studyabstractWe present an efficient LMMSE chip equalizer to suppress the interference caused by the multipath fading channel in the MIMO multi-code CDMA downlink. The block-Toeplitz structure in the correlation matrix is approximated with a block circulant matrix. An FFT-based algorithm is applied to avoid the direct-matrix-inverse (DMI) in the system equation. Hermitian optimization is proposed to further reduce the complexity. A comparative study in both performance and complexity with the conjugate-gradient (CG) algorithm is then presented. The simulation shows very promising results for the FFT-based equalizer compared with both the DMI and CG algorithms. Yuanbin Guo, Jianzhong Zhang 0002, Dennis McCain, Joseph R. Cavallaro |
GLOBECOM | 4 |
| 2003 | Reducing dynamic power consumption in next generation DS-CDMA mobile communication receiversabstractReduction of the power consumption in portable wireless receivers is an important consideration for next-generation cellular systems specified by standards such as the UMTS, IMT2000. We explore the architectural design-space and methodologies for reducing the dynamic power dissipation in the direct sequence code division multiple access (DS-CDMA) downlink RAKE receiver. Starting with a reference implementation of the DS-CDMA RAKE receiver, we demonstrate design methodologies for achieving significant power reduction, while highlighting the corresponding performance trade-offs. At the algorithm level, we investigate the tradeoffs of reduced precision and arithmetic complexity on the receiver performance. We then present two architectures for implementing the reference and reduced complexity receivers, and analyze these architectures with respect to their dynamic power dissipation. Our findings report that reduction in precision from a 16 bit to a 10 bit data-path is found to yield significant power savings of 25.6% in the reference RAKE receiver architecture, with a performance loss of less than 1 dB. Further, a power reduction of up to 24.65% is achieved in a 16 bit data-path for the reduced complexity RAKE receiver compared to the reference architecture, with a performance loss of less than 2 dB. Although there is a tradeoff in performance, adaptive power saving is very important for mobile wireless terminals. The combined effect of reduced precision and complexity reduction leads to a 37.44% savings in baseband processing power. Vikram Chandrasekhar, Frank Livingston, Joseph R. Cavallaro |
ASAP | 3 |
| 2003 | Viturbo: a reconfigurable architecture for Viterbi and turbo decodingabstractA runtime reconfigurable architecture for high speed Viterbi and turbo decoding is designed and implemented on an FPGA. The architecture can be reconfigured to decode a range of convolutionally coded data with constraint lengths varying from 3 to 9, rates 1/2 and 1/3, and various generator polynomials. It can also be reconfigured to decode turbo coded data with constraint length 4 and rate 1/3. Reconfiguration of the architecture requires a single clock cycle and does not require FPGA reprogramming. The proposed architecture can deliver data rates up to 60.5 Mbit/s for Viterbi decoding and 3.54 Mbit/s for turbo decoding, making it suitable for a range of wireless communication standards like IEEE 802.11a, 3GPP, GSM, GPRS, and many others. Joseph R. Cavallaro, Mani Vaya |
ICASSP (2) | 1 |
| 2002 | Robotic Fault Detection using Nonlinear Analytical RedundancyabstractIn this paper we discuss the application of our recently developed nonlinear analytical redundancy (NLAR) fault detection technique to a two-degree of freedom robot manipulator. NLAR extends the traditional linear AR technique to derive the maximum possible number of fault detection tests into the continuous nonlinear domain. The ability to handle nonlinear systems vastly expands the accuracy and viable applications of the AR technique. The effectiveness of the approach is demonstrated through an example. Martin L. Leuschen, Joseph R. Cavallaro, Ian D. Walker |
ICRA | 2 |
| 2002 | Real-time algorithms and architectures for multiuser channel estimation and detection in wireless base-station receiversabstractThis paper presents algorithms and architecture designs that can meet real-time requirements of multiuser channel estimation and detection in future code-division multiple-access-based wireless base-station receivers. Sophisticated algorithms proposed to implement multiuser channel estimation and detection make their real-time implementation difficult on current digital signal processor-based receivers. A maximum-likelihood based multiuser channel estimation scheme requiring matrix inversions is redesigned from an implementation perspective for a reduced complexity, iterative scheme with a simple fixed-point very large scale integration (VLSI) architecture. A reduced-complexity, bit-streaming multiuser detection algorithm that avoids the need for multishot detection is also developed for a simple, pipelined VLSI architecture. Thus, we develop real-time solutions for multiuser channel estimation and detection for third-generation wireless systems by: (1) designing the algorithms from a fixed-point implementation perspective, without significant loss in error rate performance; (2) task partitioning; and (3) designing bit-streaming fixed-point VLSI architectures that explore pipelining, parallelism, and bit-level computations to achieve real-time with minimum area overhead. Sridhar Rajagopal, Srikrishna Bhashyam, Joseph R. Cavallaro, Behnaam Aazhang |
IEEE Trans. Wirel. Commun. | 3 |
| 2001 | On-line Arithmetic for Detection in Digital Communication ReceiversabstractThis paper demonstrates the advantages of using online arithmetic for traditional and advanced detection algorithms for communication systems. Detection is one of the core computationally-intensive physical layer operations in a communication receiver and determines the communication data rates. Detection algorithms typically involve hard decisions (sign based testing) to find the sign of the transmitted information bit. This results in extraneous computations in a conventional number system as the sign is obtained only at the end due to the least significant digit first (LSDF) nature of computations. Online arithmetic, based on a signed digit number representation, provides most significant digit first (MSDF) computation. Hence, the computations can stop after the first non-zero MSD (sign) is computed and additional computations for the successive digits can be avoided. Back-conversion to a conventional number system is not required as the sign of the digit represents the detected bit. A comparison of a radix-4 serial digit on-line multiuser detector with an 8-bit parallel conventional arithmetic multiuser detector shows a decrease in latency by 1.79X, a 3X increase in throughput, and possible savings in area. Sridhar Rajagopal, Joseph R. Cavallaro |
IEEE Symposium on Computer Arithmetic | 2 |
| 2001 | VLSI implementation of Mallat's fast discrete wavelet transform algorithm with reduced complexityabstractThis paper proposes a novel VLSI architecture to compute the DWT (discrete wavelet transform) coefficients using Mallat's algorithm with reduced complexity. We studied the commonality embedded in the mirror low-pass and high-pass filters of the algorithm and use a PLA as an address generator (PAG) to load the data for cascaded FIR computation. By using an embedded downsampling process in the control signal design, we reduced the complexity by saving storage and computation. The prototyping design is implemented and fabricated using the AMI 1.5 micron CMOS process through the MOSIS service. Yuanbin Guo, Hongzhong Zhang, Joseph R. Cavallaro |
GLOBECOM | 4 |
| 2001 | On multipath channel estimation for CDMA systems using multiple sensorsabstractThis paper focuses on the design of a multiuser receiver structure for the reverse link of a code-division multiple-access communication system, in the presence of multipath effects and using an antenna array at the base station receiver. The algorithm presented solves the complex multidimensional problem of channel estimation in this complex scenario using a maximum-likelihood approach. This channel estimation technique requires the transmission of a training sequence or feedback of detected data. Once a composite channel-impulse response of each user is estimated, it is directly used in the detection process instead of first extracting the individual channel parameters, such as path delays and attenuation factors. The paper presents a framework that facilitates a computationally efficient solution to the combined problem of channel estimation and detection in a scenario involving multiple users, multiple paths, and multiple sensors at the receiver. Chaitali Sengupta, Joseph R. Cavallaro, Behnaam Aazhang |
IEEE Trans. Commun. | 2 |
| 2000 | Efficient VLSI Architectures for Baseband Signal Processing in Wireless Base-Station ReceiversabstractA real-time VLSI architecture is designed for multiuser channel estimation, one of the core baseband processing operations in wireless base-station receivers. Future wireless base-station receivers will need to use sophisticated algorithms to support extremely high data rates and multimedia. Current DSP architectures are unable to fully exploit the parallelism and bit level arithmetic present in these algorithms. These features can be revealed and efficiently implemented by task partitioning the algorithms for a VLSI solution. We modify the channel estimation algorithm for a reduced complexity fixed-point hardware implementation. We show the complexity and hardware required for three different area-time tradeoffs: an area-constrained, a time-constrained and an area-time efficient architecture. The area-constrained architecture achieves low data rates with minimum hardware, which may be used in pico-cell base-stations. The time-constrained solution exploits the entire available parallelism and determines the maximum theoretical data rates. The area-time efficient architecture meets real-time requirements with minimum area overhead. The orders-of-magnitude difference between area and time constrained solutions reveals significant inherent parallelism in the algorithm. All proposed VLSI solutions exhibit better time performance than a previous DSP implementation. Sridhar Rajagopal, Srikrishna Bhashyam, Joseph R. Cavallaro, Behnaam Aazhang |
ASAP | 3 |
| 2000 | Maximum weight basis decoding of convolutional codesabstractWe describe a new suboptimal decoding technique for linear codes based on the calculation of maximum weight basis of the code. The idea is based on estimating the maximum number locations in a codeword which have the least probability of estimation error without violating the codeword structure. In this paper we discuss the details of the algorithm for a convolutional code. The error correcting capability of the convolutional code increases with the constraint length of the code. Unfortunately the decoding complexity of Viterbi (1967) algorithm grows exponentially with the constraint length. We also augment the maximal weight basis algorithm by incorporating the ideas of list decoding technique. The complexity of the algorithm grows only quadratically with the constraint length and the performance of the algorithm is comparable to the optimal Viterbi decoding method. Elza Erkip, Joseph R. Cavallaro, Behnaam Aazhang |
GLOBECOM | 3 |
| 1999 | Keeping the Analog Genie in the Bottle: A Case for Digital RobotsabstractWe consider the case for adopting a truly 'digital' type of robot, which would evolve between a discrete and finite set of states. One distinct advantage of this philosophy is that a formal logical analysis can be applied to digital robots, since discrete-time models can now correctly and completely model the robot behavior. We argue that there are significant benefits to this strategy in numerous cases, especially with respect to fault detection and fault tolerance. However, there are also disadvantages-in order to guarantee digital behavior, constraints on the robot's operations are imposed. Essentially, we gain formality of digital analysis at the expense of precision of continuous movement. Using an analogy to digital electronics, we discuss ways in which the development of digital robots could revolutionize certain aspects of robotics. Ian D. Walker, Joseph R. Cavallaro, Martin L. Leuschen |
ICRA | 2 |
| 1999 | Efficient multiuser receivers for CDMA systemsabstractWe focus on the design of multiuser receiver structures for code division multiple access (CDMA) communication systems, in the presence of multipath effects and multiple sensors at the base station receiver. We present a flexible and extensible framework that allows the use of an estimated effective spreading code from the channel estimation phase, in the multiuser detection process. The effective spreading code captures all the channel parameters such as path delays, attenuation factors, and directions of arrival. Hence estimation of this one composite vector removes the necessity of estimating each individual parameter, thus reducing computational complexity. The results also show that this approach leads to better performance for multiuser detection, especially when the channel consists of a number of low energy paths in addition to a few discrete strong paths. Chaitali Sengupta, Joseph R. Cavallaro, Behnaam Aazhang |
WCNC | 3 |
| 1998 | Fixed point error analysis of multiuser detection and synchronization algorithms for CDMA communication systemsabstractConventional correlation based single-user techniques for direct sequence code division multiple access (DS-CDMA) wireless communication systems are susceptible to performance degradation due to interference from other users. Previous research has focused on development of several multiuser techniques where information about multiple users is used to improve performance for each individual user. Due to performance benefits of these methods, they are attractive candidates for implementation in future cellular systems. In this paper we present an error analysis of fixed point implementation of some of these techniques. Chaitali Sengupta, Joseph R. Cavallaro, Behnaam Aazhang |
ICASSP | 3 |
| 1998 | Maximum likelihood multipath channel parameter estimation in CDMA systems using antenna arraysabstractThe problem addressed in this paper is the estimation of the channel parameters in a code division multiple access (CDMA) communication system, in the presence of multipath effects and multiple sensors at the base station receiver. The algorithm presented solves the problem by estimating a composite channel impulse response of each user, which can be directly used in the detection process to appropriately modify the spreading code of the user. In addition, the algorithm combines the benefit of spatial processing in the form of an antenna array at the receiver to gain an increase in performance of the system. Chaitali Sengupta, Joseph R. Cavallaro, Behnaam Aazhang |
PIMRC | 2 |
| 1997 | Efficient Implementation of Rotation Operations for High Performance QRD-RLS FilteringabstractIn this paper we present practical techniques for implementing Givens rotations based on the well-known CORDIC algorithm. Rotations are the basic operation in many high performance adaptive filtering schemes as well as numerous other advanced signal processing algorithms relying on matrix decompositions. To improve the efficiency of these methods, we propose to use "approximate rotations", whereby only a few (i.e. r/spl Lt/b, where b is the operand word length) elementary angles of the original CORDIC sequence are applied, so as to reduce the total number of required shift add operations. This seamingly rather ad hoc and heuristic procedure constitutes a representative example of a very useful design concept termed "approximate signal processing" recently introduced and formally exposed by S.H. Nawab et al. (1997), concerning the trade-off between system performance and implementation complexity, i.e. between accuracy and resources. This is a subject of increasing importance with respect to the efficient realization of demanding signal processing tasks. We present the application of the described rotation schemes to QRD-RLS filtering in wireless communications, specifically high speed channel equalization and beamforming, i.e. for intersymbol and co-channel/interuser interference suppression, respectively. It is shown via computer simulations that the convergence behavior of the scheme using approximate Givens rotations is insensitive to the value of r, and that the misadjustment error decreases as r is increased, opening zip possibilities for "incremental refinement" strategies. B. Haller, Jürgen Götze, Joseph R. Cavallaro |
ASAP | 3 |
| 1997 | Solving the SVD updating problem for subspace tracking on a fixed sized linear array of processorsabstractThis paper addresses the problem of tracking the covariance matrix eigenstructure, based on SVD (singular value decomposition) updating, of a time-varying data matrix formed from the received vectors. This problem occurs frequently in signal processing applications such as adaptive beamforming, direction finding, spectral estimation, etc. As this problem needs to be solved in real time, it is natural to look for a parallel algorithm so that computation time can be reduced by distributing the work among a number of processing units. This paper proposes a parallel scheme for SVD updating that can be implemented on a fixed sized array of off-the-shelf processors, to get speedups close to the number of processors used. Chaitali Sengupta, Joseph R. Cavallaro, Behnaam Aazhang |
ICASSP | 2 |
| 1997 | Computationally efficient multiuser detectorsabstractCDMA is becoming an increasingly popular multiplexing scheme in wireless communication and this has necessitated the development of efficient detection techniques. The exponential complexity of the optimal detector on one end and inferior performance of conventional single-user detector at the other have led to the development of suboptimal multiuser detectors with lower complexity. Most of these detection techniques involve solution of a linear system. In their naive implementation this requires O(n/sup 3/) operations in the size of the matrix. This cost can be reduced if we move towards modern iterative techniques for solution of the system. However, maximum benefit can be achieved if we fully exploit the structure of the system. We propose several methods of reducing the computational complexity utilizing the above ideas. We have also come up with algorithms which computationally can achieve the lower bound in complexity. Joseph R. Cavallaro, Behnaam Aazhang |
PIMRC | 2 |
| 1997 | Tracking fading multipath channel parameters, in CDMA systems, using a subspace-based method-an implementation perspectiveabstractIn this paper, we evaluate several implementation issues in the application of subspace based methods to tracking channel parameters in code division multiple access (CDMA) communication systems, in the presence of multipath fading. We focus on the behavior of singular value decomposition (SVD) based schemes while tracking the time variations in the signal subspace, due to fading. We also evaluate the application of several techniques to reduce the complexity of the computationally expensive SVD procedure, to the channel estimation problem. Chaitali Sengupta, Joseph R. Cavallaro, Behnaam Aazhang |
PIMRC | 2 |
| 1995 | A dynamic fault tolerance framework for remote robotsabstractThis paper presents a layered fault tolerance framework containing new fault detection and tolerance schemes. The framework is divided into servo, interface, and supervisor layers. The servo layer is the continuous robot system and its normal controller. The interface layer monitors the servo layer for sensor or motor failures using analytical redundancy based fault detection tests. A newly developed algorithm generates the dynamic thresholds necessary to adapt the detection tests to the modeling inaccuracies present in robotic control. Depending on the initial conditions, the interface layer can provide some sensor fault tolerance automatically without direction from the supervisor. If the interface runs out of alternatives, the discrete event supervisor searches for remaining tolerance options and initiates the appropriate action based on the current robot structure indicated by the fault tree database. The layers form a hierarchy of fault tolerance which provide different levels of detection and tolerance capabilities for structurally diverse robots.> Monica L. Visinsky, Joseph R. Cavallaro, Ian D. Walker |
IEEE Trans. Robotics Autom. | 2 |
| 1994 | New Dynamic Model-Based Fault Detection Thresholds for Robot ManipulatorsabstractAutonomous robotic fault detection is becoming increasingly important as robots are used in more inaccessible and hazardous environments. Detection algorithms, however, are adversely effected by the model simplification, parameter uncertainty, and computational inaccuracy inherent in robotic control, leading to an unacceptable number of false alarms and overzealous fault tolerance. The algorithms must use thresholds to mask out these errors. Typically, the thresholds are empirically determined from a specific robot trajectory. The effect of modeling inaccuracy, however, fluctuates dynamically as the robot moves and failures occur. The thresholds need to be dynamic and respond to the changes in the robot system so as to differentiate between real failures and misalignment due to modeling errors. This paper first summarizes the reachable measurement intervals (RMI) method of computing dynamic thresholds and then, learning from the robot-oriented analysis of RMI, presents a more efficient threshold generation method using the manipulator dynamics property of linearity in parameters.> Monica L. Visinsky, Ian D. Walker, Joseph R. Cavallaro |
ICRA | 3 |
| 1994 | Redundant and On-Line CORDIC for Unitary TransformationsabstractTwo-sided unitary transformations of arbitrary 2/spl times/2 matrices are needed in parallel algorithms based on Jacobi-like methods for eigenvalue and singular value decompositions of complex matrices. This paper presents a two-sided unitary transformation structured to facilitate the integrated evaluation of parameters and application of the typically required transformations using only the primitives afforded by CORDIC; thus enabling significant speedup in the computation of these transformations on special-purpose processor array architectures implementing Jacobi-like algorithms. We discuss implementation in (nonredundant) CORDIC to motivate and lead up to implementation in the redundant and on-line enhancements to CORDIC. Both variable and constant scale factor redundant (CFR) CORDIC approaches are detailed and it is shown that the transformations may be computed in 10n+/spl delta/ time, where n is the data precision in bits and /spl delta/ is a constant accounting for accumulated on-line delays. A more area-intensive approach using a novel on-line CORDIC encoded angle summation/difference scheme reduces computation time to 6n+/spl delta/. The area/time complexities involved in the various approaches are detailed.> Nariankadu D. Hemkumar, Joseph R. Cavallaro |
IEEE Trans. Computers | 2 |
| 1993 | Efficient complex matrix transformations with CORDICabstractA two-sided unitary transformation (Q transformation) structured to permit integrated evaluation and application using CORDIC primitives is introduced. The Q transformation is shown to be useful as an atomic operation in parallel arrays for computing the eigenvalue/singular value decomposition of Hermitian/arbitrary matrices, and three specific Q transformations that are needed in such arrays are identified. Issues related to the use of CORDIC for complex arithmetic are addressed, and implementations in both conventional (nonredundant) CORDIC and redundant and online modifications to CORDIC are described. If the time to compute a CORDIC operation in nonredundant CORDIC is T/sub c/, the Q transformations identified here can be evaluated and/or applied in 2T/sub c/ using four CORDIC modules for maximum concurrency. In either case, 0.5 T/sub c/ is required to account for scale factor correction. It is shown that a Q transformation can be evaluated and/or applied in /spl ap/10n, where n is the desired bit-precision.> Nariankadu D. Hemkumar, Joseph R. Cavallaro |
IEEE Symposium on Computer Arithmetic | 2 |
| 1993 | Numerical Accuracy and Hardware Tradeoffs for CORDIC Arithmetic for Special-Purpose ProcessorsabstractThe coordinate rotation digital computer (CORDIC) algorithm is used in numerous special-purpose systems for real-time signal processing applications. An analysis of fixed-point CORDIC in the Y-reduction mode, which allows computation of the inverse tangent function, shows that unnormalized input values can result in large numerical errors. The authors describe two approaches for tackling the numerical accuracy problem. The first approach builds on a fixed-point CORDIC unit and eliminates the problem by including additional hardware for normalization. A method for integrating the normalization operation with the CORDIC iterations for efficient implementation in O(n/sup 1.5/) hardware is provided. The second solution to the accuracy problem is to use a floating-point CORDIC unit but reduce the implementation complexity by using a hybrid architecture. Arguments to support the use of such an architecture in certain special-purpose arrays are presented.> Kishore Kota, Joseph R. Cavallaro |
IEEE Trans. Computers | 2 |
| 1989 | Fault-tolerant VLSI processor array for the SVDabstractDynamic reconfiguration techniques are presented for a two-dimensional systolic array for the singular value decomposition (SVD) of a matrix. Extra computation time is not required, since idle time inherent in the array is exploited. This scheme does not require additional spare processors and is easily implemented in VLSI. Only minor hardware and communication time increases within each processing element are required.> Joseph R. Cavallaro, Christopher D. Near, M. Ümit Uyar |
ICCD | 1 |
| 1988 | Floating point CORDIC for matrix computationsabstractThe Coordinate Rotation Digital Computer (CORDIC) algorithms provide a VLSI hardware technique for computing the inverse tangents and vector rotations needed by many matrix decomposition algorithms. A novel simplified CORDIC processor composed of floating-point data paths with a fixed-point angle calculation is proposed. This hybrid processor possesses sufficient accuracy for matrix computations such as the QRD, eigenvalue decomposition, and the singular value decomposition. The simplified structure allows for efficient VLSI implementation.> Joseph R. Cavallaro, Franklin T. Luk |
ICCD | 1 |
| 1988 | CORDIC Arithmetic for an SVD ProcessorabstractArithmetic issues in the calculation of the Singular Value Decomposition (SVD) are discussed. Traditional algorithms using hardware division and square root are replaced with the special-purpose CORDIC algorithms for computing vector rotations and inverse tangents. The CORDIC 2 × 2 SVD processor can be twice as fast as one assembled from traditional hardware units. A CORDIC SVD processor array is suitable for VLSI implementation and is important for use in real-time signal processing applications. Joseph R. Cavallaro, Franklin T. Luk |
J. Parallel Distributed Comput. | 1 |
| 1987 | CORDIC arithmetic for an SVD processorabstractArithmetic issues in the calculation of the Singular Value Decomposition (SVD) are discussed. Traditional algorithms using hardware division and square root are replaced with the special purpose CORDIC algorithms for computing vector rotations and inverse tangents. The CORDIC 2×2 SVD processor can be twice as fast as one assembled from traditional hardware units. A prototype VLSI implementation of a CORDIC SVD processor array is planned for use in real-time signal processing applications. Joseph R. Cavallaro, Franklin T. Luk |
IEEE Symposium on Computer Arithmetic | 1 |