Pei-Yun Tsai 0001

dblp:90/7423-1 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-6088-3875ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 3 since 2021Computer networks · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 A Memory-Based Continuous-Flow FFT Processor with a Conflict-Free In-Place Addressing Scheme Supporting Composite Power Points
abstract
A configurable fast Fourier transform (FFT) processor capable of computing composite power points becomes prevailing in recent communication and signal processing systems. In this paper, a continuous-flow memory-based FFT processor is designed to support 63 modes from 32 to 4096 points. A conflict-free and in-place addressing scheme is proposed to minimize the number of memory banks and total memory sizes for input buffering, intermediate storage, and output re-ordering. The processing element handles radix-3, radix-4, radix-5, radix-(4•2), radix- 32, radix-42butterflies using a multi-path delay commutator (MDC) architecture with up to five parallel paths. Besides, the 3-tuple multi-radix representation is adopted for the virtual address generation and thus we can extend the concept of conventional conflict-free and in-place addressing for physical address and bank assignments beyond power-of-2 cases. From the comparison results, our proposed FFT processor architecture demonstrates strong computational efficiency by considering both the arithmetic complexity of butterfly operations and twiddle-factor multipliers as well as throughput.
Chih-Chia Chang, Hung-Shu Yu, Pei-Yun Tsai 0001
ISCAS3
2025 An OTFS-Based Synthetic Aperture Radar Imaging System for Joint Communication and Remote Sensing
abstract
Orthogonal time-frequency space (OTFS) modulation has demonstrated its robustness against high Doppler shift and is suitable for wireless communications under high mobility. We develop a joint communication and synthetic aperture radar (SAR) remote sensing system based on OTFS modulation for unmanned aerial vehicle (UAV) applications. A SAR imaging flow is proposed, including intra-pulse Doppler compensation and pilot-based channel estimation as range processing for an OTFS receiver operating in stripmap mode. The normalized fractional delay and Doppler shift are considered in the SAR channel model according to the practical conditions. The peak sidelobe ratio (PSLR) and impulse response width (IRW) of the imaging results are evaluated and compared to the conventional linear frequency modulation (LFM)-based SAR system. We show that the proposed processing flow can achieve good performance of SAR imaging without positioning deviation and supports applications of joint communication and remote sensing.
Chih-Hsien Lin, Pei-Yun Tsai 0001
VTC2025-Spring2
2024 MAML-Based 24-Hour Personalized Blood Pressure Estimation from Wrist Photoplethysmography Signals in Free-Living Context
abstract
Systolic blood pressure (BP) variability at daytime and nighttime, also known as BP dip, has shown its clinical value in diagnosis and treatment of cardiovascular diseases (CVDs). A model agnostic meta learning (MAML)-based 24-hour personalized approach is proposed in this paper for BP estimation using wrist photoplethysmography (PPG) signals from smart watch in free-living context. Detection of cuff-inflation is adopted to avoid reactive hyperemia effect after deflation and to accomplish synchronization between reference BP values measured by the ambulatory BP monitoring (ABPM) device and recorded 24-hour PPG signals in the processing flow during the experiment. The assessment of signal quality and activity occurrence is incorporated to indicate the applicability of BP estimation. The fast adaptability of personal model is enhanced by pre-training with the MAML technique and thus only few data for fine-tuning in the testing task are required. From the experimental results, our approach achieves diurnal and nocturnal SBP estimation error of 0.89 ± 7.78 mmHg from 14 recruited subjects with ages from 20 to 90 years old and shows the feasibility of tracking BP variability with smart watch in daily life.
Jia-Yu Yang, Chih-I Ho, Pei-Yun Tsai 0001, Hung-Ju Lin, Tzung-Dau Wang
ICASSP3
2024 Implementation of Group-Approximate Expectation Propagation Algorithm for Uplink MIMO-SCMA Detection Using 16-Point Codebook
abstract
The complexity of conventional massage propagation algorithm (MPA) for detection of sparse code multiple access (SCMA) grows exponentially as the size of the codebook increases, posing a challenge for hardware implementation of large-size codebooks. Expectation propagation algorithm (EPA) has shown its superiority owing to its linear complexity with respect to the codebook size. In this paper, we propose log-domain group-approximate EPA (GA-EPA) for further complexity reduction. The mother constellation points are partitioned into several groups, which can simplify the calculation of posterior probability. Compared to conventional EPA, log-domain GA-EPA can reduce approximately 76.4% of multiplications and 53.8% of divisions for MIMO-SCMA signal detection. A GA-EPA detector is then designed in 40nm CMOS technology, we use customized floating-point to shorten word-lengths and to exploit the property of exponential function for accomplishing 17% total area reduction and more than 99% table reduction. From the synthesis results, our design for MIMO-SCMA detection with 16-point codebook from 4 receiving antennas can achieve a throughput of 364Mbps at an operating frequency of 167MHz. Compared to the prior MPA-related implementations, our work outperforms in normalized hardware efficiency and demonstrates a promising solution for large codebook cardinality.
Mei-Hsuan Chang, Pei-Yun Tsai 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Design and Implementation of a Real-Time Imaging Processor for Spaceborne Synthetic Aperture Radar With Configurability
abstract
A real-time imaging processor for spaceborne synthetic aperture radar (SAR) is designed and implemented to realize the range Doppler algorithm (RDA) with configurability. The azimuth fast Fourier transform (FFT) decomposition is adopted for full utilization of data after fetching them from high bandwidth memory (HBM) by burst access to achieve streaming input–output for 2-D FFT/inverse FFT (IFFT) processing in all modes. Hybrid datapaths including fixed-point (FP), customized floating-point (CFP), and double-precision (DP) representations are used to achieve the desired signal-to-quantization-noise ratio (SQNR). The 2-D decoupling and scheduling technique is used for complexity reduction of computing spatially varying phase compensation terms. Besides, a multisegment second-order Taylor series expansion is proposed to approximate the migration factor for configurability, which is an essential component in cross-coupling compensation and azimuth compression (AC), especially when squint angle becomes large. Variable range FFT sizes from 8K to 32K are supported to cover different swath widths. The processing times for image sizes of 8K$\times8\text{K}$, 8K$\times16\text{K}$, and 8K$\times32\text{K}$are 0.34, 0.68, and 1.35 s, respectively, which meet the real-time processing requirement. Our implementation demonstrates significant improvement in processing efficiency and hardware efficiency with configurability compared with prior works.
Jia-Zhao Lin, Po-Ta Chen, Hung-Yuan Chin, Pei-Yun Tsai 0001, Sz-Yuan Lee
IEEE Trans. Very Large Scale Integr. Syst.4
2023 Modified Frequency-Domain Block Adaptive Quantizer for Synthetic Aperture Radar with Down-Sampling Requirement
abstract
A modified frequency-domain block adaptive quantizer (BAQ) is proposed for simultaneous sample-rate reduction and compression. Owing to practical hardware accelerators of fast Fourier transform (FFT) having sizes usually in powers of some simple primes, the influence on frequency-domain data compression caused by the trailing portion that follows echo signals to the FFT accelerator is first discussed. Then, both sample-rate reduction and compression are considered in frequency domain. Instead of zero padding, cyclic extension is employed for reserving the spectrum property of echo signals. Hamming windowing is applied on the cyclic extension for suppressing the sidelobe leakage. Furthermore, the distorted and noisy trailing portion can be discarded when the decompressed signal is transformed back to the time domain. From simulation results, the proposed approach achieves better signal-to-quantization noise ratio (SQNR) and smaller compressed data quantity than the conventional time-domain BAQ with anti-aliasing filter and conventional frequency-domain BAQ when sample-rate reduction is required.
Pei-Yun Tsai 0001, Hung-Shu Yu, Sz-Yuan Lee
IGARSS1
2023 Low Routing Complexity Multiframe Pipelined LDPC Decoder Based on a Novel Pseudo Marginalized Min-Sum Algorithm for High Throughput Applications
abstract
This article presents a high throughput and low routing complexity multiframe pipelined low-density parity check (LDPC) decoder design based on a novel pseudo marginalized min-sum (PMMS) message passing approach. The proposed PMMS approach reduces the required number of interconnections in the routing network allowing the design to be implemented with reduced hardware complexity, low power consumption and high throughput capability while supporting multiple coding rates with short and long codewords as defined in many application standards. Implementation results for IEEE802.11ad/ay standards show that the proposed design satisfies a target bit error rate (BER) requirement of$3 \times 10^{-7}$with 64 quadrature amp mod (QAM) targeting high throughput applications. Furthermore, the proposed design is able to achieve a throughput of 62 and 101.8 Gb/s with two pipelining stages under 28-nm CMOS and 16-nm FinFET CMOS process, respectively. As compared to the existing NMS algorithm, the proposed design based on the PMMS approach reduces the number of wires in the routing network by 45.5%, and the wirelength of the overall decoder by 17%. The area and power consumption are also reduced by 9.4% and 12%, respectively, as compared to the conventional normalized min-sum (NMS) algorithm.
Henry Lopez Davila, Tsung-Han Wu, Shyh-Jye Jou, Sau-Gee Chen, Pei-Yun Tsai 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2022 Tensor-Based Orthogonal Matching Pursuit with Phase Rotation for Channel Estimation In Hybrid Beamforming Mimo-Ofdm Systems
abstract
Tensor decomposition is often employed for channel estimation in hybrid beamforming MIMO-OFDM systems because of multiple dimensions and channel sparsity. We propose to incorporate phase rotation in factor matrices of tensor-based orthogonal matching pursuit (T-OMP) algorithm to solve the energy leakage problem caused by the grid constraint. The phase rotation can be applied in all the dimensions of virtual channel tensor including angle of arrival (AoA), angle of departure (AoD), and delay for grid refinement. Consequently, fewer iterations are required to estimate the sparse coefficients in the core tensor. In addition, the tensor fusion technique is also proposed to further improve the performance. With the grid refinement, the number of required coefficients in the core tensor is reduced and close to the number of paths. Hence, compared to the conventional T-OMP algorithms, less computation complexity is needed while better performance can be achieved.
Cheng-Hung Lo, Pei-Yun Tsai 0001
ICASSP2
2022 Compressive Sensing Based Hardware Design for Channel Estimation of Wideband Millimeter Wave Hybrid MIMO System
abstract
Channel estimation is a crucial issue for hybrid multiple-input multiple-output architecture of wideband wireless millimeter wave system. In this paper, we present a hardware implementation of channel estimation based on compressive sensing method. We exploit sparsity in channel space and take advantage of both the time domain and the frequency domain to effectively reduce the computational complexity. We then disassemble the sensing matrix into smaller dimensional matrices to further simply the sensing formula by utilizing the orthogonality of the DFT codebook. The sensing issue is solved by the generalized orthogonal matching pursuit with Cholesky decomposition techniques, which achieves a significant reduction in the number of computations. Finally, we evaluate the performance of the proposed method with the perfect channel state information, and hardware performance by fixed-point analysis and RTL design and synthesis results.
Chung-Lun Tu, Tse-Yuan Lin, Kang-Lun Chiu, Shyh-Jye Jou, Pei-Yun Tsai 0001
ISCAS5
2022 Design and Implementation of a 6.5-Gb/s Multiradix Simplified Viterbi-Sphere Decoder for Trellis-Coded Generalized Spatial Modulation With Spatial Multiplexing
abstract
A simplified Viterbi-sphere decoder (VSD) for trellis-coded generalized spatial modulation (GSM) with spatial multiplexing (TCGSMX) is designed and implemented to exploit the soft information for achieving better performance. Multiple code rates are supported to combat different levels of spatial correlations in fading channels and to offer options for various transmission efficiencies. Hence, the proposed multiradix simplified VSD is configurable to process four radix-2 stages for 1/2 code rate, two radix-4 stages for 2/3 code rate, and one radix-8 stage for 3/4 code rate in one clock cycle to upgrade the throughput by fully pipelined architecture. To reduce the complexity, several techniques are employed, including the efficient settings of survival nodes of sphere decoding and sequence length for backtracing by Viterbi decoding. The chip has been fabricated in TSMC 40-nm CMOS technology. Compared to related works, the implementation provides higher throughput up to 6.5 Gb/s at 1.1-V supply voltage as well as 180.8-MHz operating frequency and consumes 374.3 mW for TCGSMX systems using 16-QAM and code rate of 1/2.
Zhe-Yu Wang, Pei-Yun Tsai 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2021 Reduced-Complexity Channel Estimation by Hierarchical Interpolation Exploiting Sparsity for Massive MIMO Systems with Uniform Rectangular Array
abstract
Angle reciprocity is an important property adopted for massive MIMO channel estimation. For systems equipped with uniform rectangular antenna array, two-dimensional fast Fourier transform (2D-FFT) is often employed to transform the information in the spatial domain to the angular domain. To save the complexity, we propose hierarchical channel interpolation algorithm by exploiting the channel sparsity in the millimeter wave frequency band. Local interpolation with fine resolution is only applied to the selected regions determined from the coarse interpolation for obtaining required angular information so as to eliminate the unnecessary computations. Furthermore, the interference caused by the energy leakage from the adjacent paths is cancelled to acquire better estimation results of path gains. From the performance simulation results, the proposed algorithm can achieve better channel estimation performance than the zero-padded and phase-rotated 2D-FFT algorithms with reduced complexity for acquiring angular information.
Chi-Shiang Wang, Pei-Yun Tsai 0001
ICASSP2
2021 Low-Complexity Wideband Hybrid Precoding for mmWave MIMO-OFDM
abstract
Hybrid RF and baseband precoding can not only deal with low-scattering problem caused by severe signal attenuation of mmWave but also significantly reduce the power and cost of radio frequency (RF) chains and data converters in the massive MIMO transceiver. However, its baseband complexity is still very high for frequency-selective multi-carrier system such as orthogonal frequency division multiplexing (OFDM). Therefore, this paper proposes a computationally-efficient hybrid precoding algorithm for mmWave MIMO-OFDM systems. By decomposing the optimization of spectrum efficiency into many sub-optimization problems, we developed an iterative matrix- inversion-bypass algorithm to reduce the complexity. The baseband precoder was also designed to further reduce total complexity. The simulation and complexity analysis results showed that we can significantly reduce computational complexity for mmWave MIMO with negligible performance degradation compared to a state-of-the-art work in the literature.
Hsin-Wen Ku, Wei-Hao Fang, Huan-Lun Tso, Pei-Yun Tsai 0001, Yuan-Hao Huang
ISCAS4
2020 Design of A Convergence-Aware Based Expectation Propagation Algorithm for Uplink Mimo Scma Systems
abstract
Sparse code multiple access (SCMA) uses multidimensional sparse codewords to transmit user data. The expectation propagation algorithm (EPA) exploiting the sparse property shows linear complexity growth and thus is preferred for multi-user detection. To further reduce the complexity, a convergence-aware based EPA for uplink MIMO SCMA systems is proposed. Techniques including user termination, antenna termination, and codebook reduction are adopted. The user termination must be combined with the iteration constraint to avoid misjudgement. The antenna termination can stop the computations related with certain antennas having strong channel gains. Only possible codewords are considered in the reduced codebook to eliminate unnecessary calculation for posterior probability. From simulation results, we show that these three techniques can strike a balance between performance and complexity and more than 50% complexity can be saved.
Jih-Yang Lin, Pei-Yun Tsai 0001
ICASSP2
2020 A 538Mbps 2×64 Spatial Permutation Modulation Detector for MIMO Systems
abstract
Spatial permutation modulation (SPM) is a new multiple-input-multiple-output (MIMO) technology for next-generation communication systems, which is an extension of spatial modulation (SM) that conveys data information at multiple time instants. By encoding the permutation of activated antennas at several time instants, the SPM system gains benefits of transmit and time diversities enabling reliable mobile communications. This study proposed a low-complexity SPM detector, called multiple-candidate-selection matching maximal ratio combining detector (MCSMMRC), for the SPM system. The MCSMMRC detector has very low complexity, fixed throughput, and scalable computing structure, which makes MCSMMD suitable for hardware implementation. This study designed and implemented the proposed MCSMMD detector by using a Xilinx Virtex-7 FPGA chip. The FPGA implementation results showed that it achieved a maximum throughput of 538 Mbps and exhibited better normalized throughput than those of other SM-based detector chips in the literature.
Jung-Chun Chi, Yu-Cheng Yeh, I-Wei Lai, Pei-Yun Tsai 0001, Yuan-Hao Huang
ISCAS4
2020 Weighted Pulse Decomposition Analysis of Fingertip Photoplethysmogram Signals for Blood Pressure Assessment
abstract
A weighted pulse decomposition analysis is proposed for cuffless blood pressure assessment. Five Gaussian waves are used for decomposition. The weighted least squares criterion is adopted for optimization. Instead of applying weights to certain crest and trough points of the pulse signal from fingertip photoplethysmogram, we apply the weights to the pulse segment that is recognized informative for vascular age, vessel stiffness, and blood pressure. In addition, the boundary constraints of the Gaussian parameters are carefully set so that the normalized root mean square error (NRMSE) between the original PPG and synthesized PPG can be kept acceptable. From the results, we can see that weighting makes the decomposition stable and the variances of Gaussian parameters reduced. Besides, the modified boundary constraints improve the NRMSE. Furthermore, the correlation between the blood pressures and the features from weighted pulse decomposition analysis is enhanced and thus the proposed approach is helpful to blood pressure assessment.
Chiu-Hua Huang, Jia-Wei Guo 0001, Yu-Chia Yang, Pei-Yun Tsai 0001, An-Yeu Wu, Hung-Ju Lin, Tzung-Dau Wang
ISCAS4
2020 A 75-Gb/s/mm2 and Energy-Efficient LDPC Decoder Based on a Reduced Complexity Second Minimum Approximation Min-Sum Algorithm
abstract
This article presents a high-throughput and low-routing complexity low-density parity check (LDPC) decoder design based on a novel second minimum approximation min-sum (SAMS) algorithm. The routing congestion is mitigated by reducing the required interconnections in the critical path of the routing network. The implementation and postlayout results with 28-nm 1P9M CMOS process show that the proposed design can achieve a throughput of 10.5 Gb/s for a millimeter-wave 60-GHz baseband system while satisfying the low bit error rate (BER) requirements (10-7). The proposed design reduces the wiring in the routing network by 21% and improves the area by 12% compared to the conventional min-sum (MS) and normalized MS (NMS) algorithm. Additional hardware optimizations are obtained by considering the internal message passing resolution based on the BER and signal-to-noise ratio (SNR) requirements for a practical baseband system. The power consumption is efficiently reduced by the employment of a shared address generator that exploits the degree of parallelism to reduce the switching activity on a group of memory elements. The LDPC decoder is implemented with a core area of 0.14 mm2, power consumption of 81 mW at 312.5 MHz, and the area and power efficiency of 75 Gb/s/mm2and 10.2 pJ/bit, respectively.
Henry Lopez Davila, Hsun-Wei Chan, Kang-Lun Chiu, Pei-Yun Tsai 0001, Shyh-Jye Jou
IEEE Trans. Very Large Scale Integr. Syst.4
2019 A Double K-best Viterbi-sphere Decoder for Trellis-coded Generalized Spatial Modulation with Multiple Code Rates
abstract
The trellis-coded generalized spatial modulation (TCGSM) system with multiple code rates is investigated so as to provide configurability and to enhance efficiency in different fading channels. Various high-radix stages are then adopted to support the optimal sequence detection for different configurations. A double K-best technique is then proposed to strike a good balance between performance and complexity of Viterbi-sphere decoding. From the analysis and simulation results, 29%-69% complexity can be saved with about 1dB performance degradation for the code rate of 4/5 in the spatial domain with QPSK and 16-QAM constellations. The first K-best selection for generating branch metrics from sphere decoding can reduce the detection complexity of the constellation while the second K- best selection can be considered if a high-radix trellis is used when the number of transmit antennas is large.
Zhe-Yu Wang, Pei-Yun Tsai 0001, Yuan-Hao Huang, I-Wei Lai
ICASSP2
2019 Design of a Tunable Block Floating-Point Quantizer with Fractional Exponent
abstract
For the conventional block floating-point quantizer (BFPQ), usually a large block size causes performance degradation and thus small block sizes are preferred, especially when non-uniformly distributed signals are processed. A tunable BFPQ with fractional exponent is proposed in this paper to deal with the problem. We first examine the root cause of degradation through analytic equations and then propose to tune the thresholds for deriving the exponent and fractional exponent of the block so as to strike a good balance between the quantization error and saturation error. An optimal tuning value depending on the block size and mantissa word-length can be obtained. Thus, the tunable BFPQ can achieve better output signal-to-quantization-noise ratio (SQNR) in a wide dynamic range. The analytic equation for the output SQNR of the proposed BFPQ is derived to verify the simulated results. Only one extra multiplication is required for each block to implement the tunable BFPQ. Finally, we show the obvious SQNR improvements compared to the conventional scheme for various settings of block sizes and mantissa word-lengths.
Pei-Yun Tsai 0001, Tien-I Yang, Ching-Horng Lee, Li-Mei Chen, Sz-Yuan Lee
ISCAS1
2018 Fast-Convergence Singular Value Decomposition for Tracking Time-Varying Channels in Massive Mimo Systems
abstract
A fast-convergence singular value decomposition (SVD) algorithm is developed for tracking time-varying channels in massive MIMO precoding/beamforming systems. Since only strong eigen-modes are selected for data transmission in these systems, our SVD algorithm exploits the properties of partial decomposition and temporal correlation. Besides, the proposed self-adjusting inverse power method can achieve fast convergence by modifying the shift according to the intermediate result during each iteration. Furthermore, the singular vectors and values of the desired eigenmodes can be computed simultaneously. Thus, parallel processing is possible to facilitate high-throughput implementation. Compared to the self-power method with super linear convergence, the self-adjusting inverse power method has better convergence and lower complexity. Good channel tracking capability is also demonstrated.
Pei-Yun Tsai 0001, Jian-Lin Li
ICASSP1
2018 Theoretical Performance Analysis Assisted by Machine Learning for Spatial Permutation Modulation (SPM) in Slow-Fading Channels
abstract
Based on spatial modulation (SM), spatial permu- tation modulation (SPM) has been recently proposed to enhance the performance of the multiple-input multiple-output (MIMO) system. SPM maps data bits to both the QAM symbol and permutation array. At successive time instants, different transmit antennas are activated according to the mapped permutation array to transmit the QAM symbol. In this work, the error rate of SPM in slow-fading channels is analyzed. The performance is first analyzed with the closed-form expression for the special case, and then is generalized to arbitrary cases by using the approximation of Gamma random variables. The machine learning algorithm is adopted to simplify the generalization and estimate the diversity. Through the analyses, we discover that by simply adding transmit antennas, the performance of SPM in slow-fading channels can be greatly enhanced due to the reduction of the time dependency. Numerical simulations demonstrate the accuracy of our analyses and show that by adding one transmit antenna, the time dependency can almost be removed, leading to around 3 dB SNR gain for the BER performance.
Jhih-Wei Shih, Jung-Chun Chi, Yuan-Hao Huang, Pei-Yun Tsai 0001, I-Wei Lai
ICC4
2017 A generalized matrix-decomposition processor for joint MIMO transceiver design
abstract
A generalized matrix-decomposition processor is designed and implemented, which supports QR decomposition (QRD), eigenvalue decomposition (EVD), and geometric-mean decomposition (GMD), to accelerate computations in MIMO precoding/beamforming systems. The processor adopts memory-based architecture with 16 processing elements (PEs) each consisting of one CORDIC module. An improved GMD algorithm is proposed, which reduces 13.2% complexity and can be implemented by homogeneous CORDIC operations. The EVD adopts the Rayleigh quotient shift and deflation technique to accelerate convergence. The basis computations can be accomplished by mirrored operations during channel matrix decomposition. From the implementation results, the generalized processor achieves decomposition throughput of 10M, 0.99M, 2.96M matrixes per second for 4 × 4 complex QRD, EVD and GMD.
Yu-Chi Wu, Pei-Yun Tsai 0001
ICASSP2
2017 Design of an SVD engine for 8×8 MIMO precoding systems
abstract
A singular-value-decomposition (SVD) engine for 8×8 MIMO precoding systems is designed and implemented. The memory-based architecture is adopted with eight processing elements, each having two CORDIC modules. Two-phase operations are performed including bidiagonalization and Golub-Reinsch SVD (GR-SVD) with Rayleigh quotient shift. The split, deflation, and shift techniques of GR-SVD can effectively decrease the processed matrix size and accelerate the diagonalization to enhance the throughput. To cover the wide distribution of singular vector elements and singular values derived from 8 × 8 MIMO channel matrix, hybrid datapath representations are used. The thresholds for split and deflation can be adjusted and thus the accuracy of the SVD engine is variable according to the requirements. From the synthesis results, the SVD engine in 45nm CMOS technology is able to provide the throughput rate of 636K matrix/s and outperforms the previous design.
Chun-Hun Wu, Chin-Yi Liu, Pei-Yun Tsai 0001
ISCAS3
2016 Population Based Ant Colony Optimization for Reconstructing ECG Signals
Yih-Chun Cheng, Tom Hartmann, Pei-Yun Tsai 0001, Martin Middendorf
EvoApplications (1)3
2016 WHDVI: A wireless high definition video interface technique for digital home
Tsung-Han Tsai 0001, Pei-Yun Tsai 0001, Meng-Yuan Huang, Li-Yang Huang
Integr.2
2015 Multi-mode sorted QR decomposition for 4×4 and 8×8 single-user/multi-user MIMO precoding
abstract
This paper presents a configurable multi-mode QR decomposition (QRD) processor. It supports 4×4 and 8×8 QRD with multi-layer sorting for single-user MIMO precoding. Besides, it can perform block-based sorting for multi-user MIMO precoding. Both forward mode for decomposition and backward mode for signal precoding are provided. This QRD processor is designed in pipelined systolic array. An in-place strategy with pointer-based control mechanism is proposed for sorting buffers, which reduces 42.3% D flip-flops. The proposed processor implemented in 90nm CMOS technology can generate 9.45MQRD/s for decomposing 8×8 channel matrix with sorting, and outperforms the related works in terms of throughput and hardware efficiency.
Chi-Mao Chen, Chih-Hsiang Lin, Pei-Yun Tsai 0001
ISCAS3
2015 Low-complexity compressed sensing with variable orthogonal multi-matching pursuit and partially known support for ECG signals
abstract
In this paper, we present low-complexity compressed sensing (CS) techniques for monitoring electrocardiogram (ECG) signals in wireless body sensor network (WBSN). We first exploit ECG properties in the wavelet domain to extend the partially known support set (PKS) so as to reduce the support augmentation and estimation efforts in the iterative recovery algorithm. Then, variable orthogonal multi-matching pursuit (vOMMP) algorithm is proposed, which uses orthogonal matching pursuit (OMP) algorithm in the first phase to effectively augment the support set with reliable supports and adopts the orthogonal multi-matching pursuit (OMMP) in the second phase to rescue the missing supports. The reconstruction performance is thus enhanced. Furthermore, the computation-intensive pseudo-inverse operation for signal reconstruction is simplified by the matrix-inversion-free technique based on QR decomposition. The performance and complexity comparisons manifest the advantages of our proposed techniques.
Yih-Chun Cheng, Pei-Yun Tsai 0001
ISCAS2
2014 The GMD-based precoding and antenna selection schemes for CoMP joint processing
abstract
In this paper, we discuss the precoding schemes for coordinated multi-point (CoMP) joint processing (JP) using geometric mean decomposition (GMD). Unlike JP singular value decomposition (JP-SVD) whose performance is dominated by the spatial pipe with weak channel gain, JP-GMD is proposed to derive the precoding matrix of all the cooperative cell sites and the decoding matrix of the user equipment (UE) so that equal spatial channel gains can be obtained. Sphere decoding (SD) is then employed to achieve the maximum likelihood (ML) solution. Besides, centralized Tomlinson Harashima precoding (THP) can be adopted at the base station to remove interference, called JP-GMD-THP. We show that the JP-GMD plus SD scheme has significant performance improvement over JP-SVD, JP-zero forcing (ZF) and JP-minimum mean square error (MMSE). JP-GMD-THP that allows simple one-tap equalization has performance loss only about 0.7 dB compared to the JP-GMD plus SD scheme. Furthermore, to reduce the joint processing efforts at all cell sites and to take advantage of the spatial domain, two antenna selection techniques, best antenna selection and grouping antenna selection, are also proposed. As opposed to the best antenna selection, the grouping antenna selection technique can reduce search efforts with about 0.5 dB SNR degradation, but is still better than the non-cooperative schemes.
Ching-Heng Yeh, Pei-Yun Tsai 0001
PIMRC2
2014 Improvement of Explicit Channel Feedback for MIMO-OFDM WLAN and Its Implementation
abstract
In this paper, explicit channel feedback for MIMO- OFDM systems is investigated. We first show that by exploiting the finite signal space of the channel impulse response, sub-sampling smoothing filter can be designed to achieve both feedback reduction and noise suppression. Secondly, to improve the distribution of the signal to be quantized, it is better to extract the common scaling term from the spectral block to utilize frequency correlation rather than to extract from the spatial block. Simulation results show that the proposed channel feedback scheme with the good sub-sampling smoothing filter brings substantial performance gain and helps the reduction of feedback overhead compared to feedback of the noisy CSI with the grouping strategy. We also discuss the implementation of the feedback encoding block to evaluate the hardware requirements given the feedback throughput constraint. The feedback reduction ratio is also shown to be essential to hardware saving and processing cycles. The proposed scheme has been implemented in 90nm CMOS technology with a maximum operating clock frequency of 160 MHz. We demonstrate the feasibility of the proposed feedback encoding scheme with good performance and efficient hardware.
Min-Ching Chen, Pei-Yun Tsai 0001
VTC Spring2
2014 BD-QRD, Block THP and Constrained Sphere Decoding for Multi-User MIMO Systems
abstract
This paper presents a multi-user MIMO transceiver design with a block-based decomposition and precoding scheme. To exploit spatial diversity, we propose to use block-diagonal QR decomposition (BD- QRD) to decompose the channel matrix. To eliminate multiple access interference (MAI), block- Tomlinson-Harashima precoding (B-THP) is further proposed to be combined with BD-QRD so that the equivalent channel matrix after precoding at the transmitter becomes a block diagonal matrix. On the other hand, the block-based sorting is adopted to balance the energy spread among all spatial pipes for BD-QRD and thus the performance can be further enhanced. With these decomposition and precoding techniques at the transmitter, the sphere decoding (SD) techniques can be employed at the receiver with a small revision to constrain the search space. We show that the proposed BD-QRD, B-THP, and constrained SD for multi-user MIMO systems retaining the spatial diversity, outperform the conventional QRD-THP and BD-SVD in the interested SNR region.
Chi-Mao Chen, Pei-Yun Tsai 0001, Chia-Wei Chen
VTC Spring2
2013 Equal-rate QR decomposition based on MMSE technique for multi-user MIMO precoding
abstract
In recent years, the research on multiple-input multiple-output (MIMO) wireless communications has attracted much interest. This study investigates precoding techniques for multi-user MIMO communications. By applying QR decomposition to augmented channel matrix based on minimum mean squared error (MMSE) approach plus Tomlinson-Harashima precoding (THP) and equal-rate power allocation, we show that the proposed equal-rate QR-MMSE-THP precoder has better performance than the one based on zero-forcing (ZF) channel matrix, called equal-rate QR-ZF-THP precoder, and the conventional MMSE precoder. In addition, sorting strategies can be adopted to enhance performance with more freedom for QR decomposition compared to block diagonalized decompositions such as block diagonalized-geometric mean decomposition (BD-GMD) and BD-GMD-MMSE schemes while complexity can still be kept practical. Also, the sorting strategies such as full sorting (FS), per-layer sorting (PLS) and per-user sorting (PUS) are discussed and compared in this paper. Simulation results show that for large MIMO systems, the proposed equal-rate QR-MMSE-THP with PLS strategy outperforms the equal-rate BD-GMD-MMSE-THP precoder with PUS.
Chia-Wei Chen, Hen-Wai Tsao, Pei-Yun Tsai 0001
PIMRC3
2013 A non-coherent neighbor cell search scheme for LTE/LTE-A systems
abstract
A new neighbor cell search algorithm for LTE/LTE-A systems is presented in this paper. To improve the interference problem in channel estimation for coherent SSS detection in the conventional neighbor cell search approaches, we propose a non-coherent scheme that takes advantage of the similarity of channel responses at adjacent subcarriers. The proposed neighbor cell search procedure not only includes both PSS and SSS detection, but also can combat different carrier frequency offsets that the home cell signal and the neighbor cell signal may suffer. The removal of the home cell synchronization signals in our algorithm converts the neighbor cell PSS and SSS into new sequences for recognition, respectively. By examining the cross-correlation properties of the new sequences, we show that partial correlation can well detect the neighbor cell sector ID and group ID through the new sequences. From simulation results, it is also clear that the proposed algorithm has good detection results and outperforms the conventional coherent approaches.
Shun-Fang Liu, Pei-Yun Tsai 0001
WCNC2
2013 Design and implementation of a distributed WLS localization and tracking algorithm in wireless sensor network
abstract
In this paper, a distributed algorithm for localization and tracking in wireless sensor network is designed and implemented. Combining the fingerprint approach and subgradient optimization, the proposed algorithm, called golf algorithm, uses different strategies to attain fast movement for tracking a moving object and a fine tune to approach a stationary object, like a drive and a putt on the putting green in the golf. Based on the WLS objective function, a recursive form for both the fingerprint approach and subgradient optimization is derived and can be accomplished by anchor nodes in collaboration. From simulation results, we can see the proposed algorithm has better estimation performance than the conventional decentralized algorithms to localize a stationary or moving target in large wireless sensor networks. In addition, its hardware architecture is designed. The CORDIC operation is employed to compute the vector norm. Parallel processing is used to handle the calculation of fingerprint data. The tracking trajectories of the floating-point program and fixed-point hardware implementation are shown to be quite similar with small finite precision effect, which verifies the feasibility of this hardware solution.
Ching-Hsien Wang, Pei-Yun Tsai 0001
WCNC2
2012 Precoder selection under K-best MIMO detection for 3GPP-LTE/LTE-A systems
abstract
Precoding by using a finite-set codebook at transmitters is an effective approach to improve the performance of multiple-input multiple-output (MIMO) systems with reduced feedback overhead. K-Best sphere decoders are popular for high-throughput spatially-multiplexed MIMO signal detection at receivers. In this paper, we propose a precoding-matrix-index (PMI) selection criterion suitable for K-best MIMO detectors by minimizing the trace of the upper triangular matrix after QR decomposition (QRD). To further improve the system performance, fully-sorted QRD, per-layer sorted QRD and partially-sorted QRD are also proposed to be incorporated in the minimum-trace PMI selection criterion. Simulation results based on the MIMO codebooks of 3GPP-LTE/LTE-A systems show the improvements of the proposed scheme compared to the non-precoding MIMO schemes and the precoding schemes using the conventional PMI selection criterions.
Pei-Yun Tsai 0001
APCC2
2011 A Generalized Conflict-Free Memory Addressing Scheme for Continuous-Flow Parallel-Processing FFT Processors With Rescheduling
abstract
This paper presents a generalized conflict-free memory addressing scheme for memory-based fast Fourier transform (FFT) processors with parallel arithmetic processing units made up of radix-$2^{q}$multi-path delay commutator (MDC). The proposed addressing scheme considers the continuous-flow operation with minimum shared memory requirements. To improve throughput, parallel high-radix processing units are employed. We prove that the solution to non-conflict memory access satisfying the constraints of the continuous-flow, variable-size, higher-radix, and parallel-processing operations indeed exists. In addition, a rescheduling technique for twiddle-factor multiplication is developed to reduce hardware complexity and to enhance hardware efficiency. From the results, we can see that the proposed processor has high utilization and efficiency to support flexible configurability for various FFT sizes with fewer computation cycles than the conventional radix-2/radix-4 memory-based FFT processors.
Pei-Yun Tsai 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2010 High-throughput QR decomposition for MIMO detection in OFDM systems
abstract
In this paper, we aim to design and implement a high-throughput QR decomposition architecture for 4 × 4 MIMO signal detection problems. A real-value decomposed MIMO system model is handled and thus the channel matrix to be processed is extended to the size 8×8. Instead of direct factorization, we propose a QR decomposition scheme by cascading one complex-value and one real-value Givens rotation blocks, which can save 44% hardware complexity. The systolic array is adopted for hardware implementation to facilitate pipeline design. Then, the requirement of skewed inputs to the conventional complex-value QR-decomposition systolic array is improved and 37% of delay elements are removed. The real-value Givens rotation stage is implemented by a stacked triangular systolic array to match with the throughput of the complex-value one. We have implemented the proposed design in 0.18 μm CMOS technology with 152K gates. From post-layout simulations, the maximum operating frequency can achieve 90.09MHz. The proposed scheme not only reduces the hardware complexity, but also supports high throughput for MIMO-OFDM signal detection up to 2.16 Gbps under stationary channels.
Zheng-Yu Huang, Pei-Yun Tsai 0001
ISCAS2
2010 A 4×4 64-QAM reduced-complexity K-best MIMO detector up to 1.5Gbps
abstract
In this paper, a VLSI architecture of a reduced-complexity K-best sphere decoder is designed, which aims to solve the 4 × 4 64-QAM multiple-input multiple-output (MIMO) signal detection problems in high-speed applications. We propose a fully-pipelined sorter, which can generate one result per clock cycle and thus greatly enhance the detection throughput. On the other hand, various K values are adopted at each layer to save the hardware complexity. The proposed design has been implemented in 0.18 Jim CMOS technology and has 366K gates. From post-layout simulation, this work achieves a detection rate of 1.5 Gbps at 62.5-MHz clock frequency.
Pei-Yun Tsai 0001, Wei-Tzuo Chen, Xing-Cheng Lin, Meng-Yuan Huang
ISCAS1
2009 A Low-Power Delay Buffer Using Gated Driver Tree
abstract
This paper presents circuit design of a low-power delay buffer. The proposed delay buffer uses several new techniques to reduce its power consumption. Since delay buffers are accessed sequentially, it adopts a ring-counter addressing scheme. In the ring counter, double-edge-triggered (DET) flip-flops are utilized to reduce the operating frequency by half and the C-element gated-clock strategy is proposed. A novel gated-clock-driver tree is then applied to further reduce the activity along the clock distribution network. Moreover, the gated-driver-tree idea is also employed in the input and output ports of the memory block to decrease their loading, thus saving even more power. Both simulation results and experimental results show great improvement in power consumption. A 256 times 8 delay buffer is fabricated and verified in 0.18 mum CMOS technology and it dissipates only 2.56 mW when operating at 135 MHz from 1.8-V supply voltage.
Po-Chun Hsieh, Jing-Siang Jhuang, Pei-Yun Tsai 0001, Tzi-Dar Chiueh
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Adaptive raised-cosine channel interpolation for pilot-aided OFDM systems
abstract
In this paper, we first show equivalence of OFDM channel estimation using time-domain windowing and using frequency-domain interpolation. Based on this equivalence, a new frequency-domain channel interpolator featuring the advantage of time-domain channel impulse response windowing is proposed. Furthermore, the proposed raised-cosine channel interpolator can adaptively adjust its coefficients to accommodate various channel power delay profiles. Both theoretical and simulation results verify that the proposed interpolator achieves more accurate channel estimation performance than other existing solutions.
Pei-Yun Tsai 0001, Tzi-Dar Chiueh
IEEE Trans. Wirel. Commun.1