VLDB 2026 Research / reviewers in the wild / expert
Yuan-Hao Huang
dblp:69/5896
· DBLP profile ↗
36ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-6781-7312ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Computer networks · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multilinear Generalized SVD Processor for Multicast Multiuser MIMO Systems
Jung-Chun Chi, Yi-Chieh Hsu, Yuan-Hao Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Algorithm and Architecture of a Channel-Correlation-Resistant Massive MIMO Detector Based on Expectation PropagationabstractExpectation propagation (EP) and its approximation algorithms achieve near-optimal signal detection in massive multiple-input multiple-output (MIMO) systems with substantially reduced complexity under uncorrelated channels. However, approximate EP (EPA) detectors often suffer from severe performance degradation in the presence of channel correlation. This article presents a novel EPA-correlation-resistant Chebyshev-accelerated Richardson iteration (EPA-CR-CARI) MIMO detection algorithm that achieves near-optimal performance without requiring channel-specific processing, which incurs considerable latency and area overhead even under correlated channels. Furthermore, a corresponding low-latency and low-complexity VLSI architecture is presented. Compared to the state-of-the-art design targeting correlated channels, the proposed MIMO detector achieves an 18.7% reduction in latency and a 17.3% decrease in hardware complexity. Yi-Ling Tsai, Wenn-Yi Lin, Chung-An Shen, Yuan-Hao Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Two-stage Adaptive Compressive Sensing and Reconstruction for Terahertz Single-Pixel ImagingabstractTerahertz (THz) single-pixel imaging system based on compressive sensing (CS) technique is recently studied widely because it can significantly reduce the cost of THz sensor and processing latency. This paper proposes a two-stage adaptive CS and reconstruction algorithm for single-pixel THz imaging systems. The proposed algorithm incorporates the object profile information, which is acquired from the first-stage sensing and reconstruction, in the mask patterns of the second-stage sensing and reconstruction, thereby improving the image quality with smaller measurement number and reduced computational complexity. Simulation results show that, when the target mean square error is set to be 0.08, the proposed two-stage CS and reconstruction algorithm can reduce 43.9% measurements and 36% complexity for 64 × 64 images compared with the traditional one-stage single-pixel CS imaging system. Yu-Kai Zhang, Che-Yu Chou, Shang-Hua Yang, Yuan-Hao Huang |
ISCAS | 4 |
| 2024 | The Algorithm and VLSI Architecture of High-Throughput and Highly Efficient Tensor Decomposition EngineabstractTensor decomposition is critical for compressing data and extracting key features in novel high-dimensional signal processing systems. However, due to the enormous amount of data and the highly complicated computations, designing an efficient tensor decomposition processor is very challenging. This paper presents the algorithm and VLSI architecture design of a low-latency and high-throughput tensor decomposition processor. A parallel higher-order orthogonal iteration (P-HOOI) algorithm is proposed where multiple updated matrices are computed concurrently. Thus, a low-latency tensor decomposition is achieved. Furthermore, a novel VLSI architecture is presented so that the efficiency of the component utilization is improved and the hardware complexity is greatly reduced. Therefore, the proposed tensor decomposition processor enhances the processing throughput with minimum employment of hardware components. Performance evaluations based on the post-layout estimations in the ASIC flow and based on the FPGA platform are reported in this paper. Compared with state-of-the-art designs in the literature the proposed tensor decomposition engine greatly enhances the throughput and hardware efficiency. Ting-Yu Tsai, Chung-An Shen, Tsung-Lin Wu, Yuan-Hao Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | -Complexity Low-Rank Approximation SVD for Massive Matrix in Tensor Train FormatabstractTensor train decomposition (TTD) has recently been proposed for high-dimensional signals because it can save significant storage in various signal processing applications. This paper presents a low-rank approximation algorithm of singular value decomposition (SVD) for large-scale matrices in tensor train format (TT-format). The proposed alternating least square block power SVD (ALS-BPSVD) algorithm can reduce the computational complexity by decomposing the large-scale SVD into a low-rank approximation scheme with a fixed-iteration block power method for searching singular values and vectors. Moreover, a low-complexity two-step truncation scheme was proposed to reduce more complexity and facilitate the parallel processing. The proposed ALS-BPSVD algorithm can support the low-rank approximation SVD for matrices with dimension higher than 211× 211. The simulation results show that the ALS-BPSVD achieved up to 21.3 times speed-up compared to the benchmark ALS-SVD algorithm for the random matrices with prescribed singular values. Jung-Chun Chi, Chiao-En Chen, Yuan-Hao Huang |
ICASSP | 3 |
| 2022 | Tensor-Based Hybrid Precoding Processor for 8 × 8 × 8 mmWave 3D-MIMO SystemsabstractHybrid baseband precoding and RF beamforming is a highly efficient technology for millimeter-wave (mmWave) massive multiple-input multiple-output (MIMO) systems. Threedimensional (3D) MIMO system with uniform planar array (UPA) of transit antennas and uniform linear array (ULA) of receive antennas can provide more flexible and efficient beamforming capability in sparse mmWave channels. Tensor is a compact multi-way algebraic model that can describe high-dimension systems such as the sparse mmWave 3D-MIMO system. This paper proposes a tensor-based hybrid precoding algorithm for continuous time-drifting 3D-MIMO systems which can achieves better performance in high bit-stream and low SNR systems. The FPGA implementation of the tensor-based hybrid precoding processor can support the 3D-MIMO system with 8 × 8 UPA transmitter and 8-antenna ULA receiver with a maximal normalized throughput of 17.0 M matrices/sec compared to the existing counterparts. Tsung-Lin Wu, Chung-An Shen, Yuan-Hao Huang |
ISCAS | 3 |
| 2022 | Chaos LiDAR Based RGB-D Face Classification System With Embedded CNN Accelerator on FPGAsabstractFace classification is important in many applications such as surveillance, border control, and security systems. However, wide variations in environments such as insufficient light, large distances or pose angles make the task challenging. Depth sensors are added with RGB cameras for improving classification accuracy but commercial RGB-D sensors are most targeted for indoors applications. In this paper, we present and design a Chaos LiDAR depth sesnor that provides high-precision depth images through intelligent correlation processing for both indoors and outdoors applications. Our Chaos LiDAR depth sensor detects range from 2 to 40 meters with precision around 8mm at 20-meter. With the Chaos LiDAR depth as input, we design a RGB-D based face classification embedded CNN (eCNN) model for wide range applications such as dim illumination, various distances and large poses. Our Chaos LiDAR increases around 14.27% classification accuracy compared to RealSense D435i for distance from 3 to 5 meter. The eCNN face classification subsystem is implemented in Xilinx ZCU 102 and achieves 11.11 ms inference time. The eCNN engine achieves a peak throughput at 614.4 GOPS. The overall system including Chaos LiDAR, correlation and eCNN FPGA achieves face classification inference rate of 10fps. Ching-Te Chiu, Yu-Chun Ding, Wei-Jyun Chen, Shu-Yun Wu, Chao-Tsung Huang, Chun-Yeh Lin, Chia-Yu Chang, Meng-Jui Lee, Shimazu Tatsunori, Tsung Chen, Fan-Yi Lin, Yuan-Hao Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 13 |
| 2021 | Low-Complexity Wideband Hybrid Precoding for mmWave MIMO-OFDMabstractHybrid RF and baseband precoding can not only deal with low-scattering problem caused by severe signal attenuation of mmWave but also significantly reduce the power and cost of radio frequency (RF) chains and data converters in the massive MIMO transceiver. However, its baseband complexity is still very high for frequency-selective multi-carrier system such as orthogonal frequency division multiplexing (OFDM). Therefore, this paper proposes a computationally-efficient hybrid precoding algorithm for mmWave MIMO-OFDM systems. By decomposing the optimization of spectrum efficiency into many sub-optimization problems, we developed an iterative matrix- inversion-bypass algorithm to reduce the complexity. The baseband precoder was also designed to further reduce total complexity. The simulation and complexity analysis results showed that we can significantly reduce computational complexity for mmWave MIMO with negligible performance degradation compared to a state-of-the-art work in the literature. Hsin-Wen Ku, Wei-Hao Fang, Huan-Lun Tso, Pei-Yun Tsai 0001, Yuan-Hao Huang |
ISCAS | 5 |
| 2020 | A 538Mbps 2×64 Spatial Permutation Modulation Detector for MIMO SystemsabstractSpatial permutation modulation (SPM) is a new multiple-input-multiple-output (MIMO) technology for next-generation communication systems, which is an extension of spatial modulation (SM) that conveys data information at multiple time instants. By encoding the permutation of activated antennas at several time instants, the SPM system gains benefits of transmit and time diversities enabling reliable mobile communications. This study proposed a low-complexity SPM detector, called multiple-candidate-selection matching maximal ratio combining detector (MCSMMRC), for the SPM system. The MCSMMRC detector has very low complexity, fixed throughput, and scalable computing structure, which makes MCSMMD suitable for hardware implementation. This study designed and implemented the proposed MCSMMD detector by using a Xilinx Virtex-7 FPGA chip. The FPGA implementation results showed that it achieved a maximum throughput of 538 Mbps and exhibited better normalized throughput than those of other SM-based detector chips in the literature. Jung-Chun Chi, Yu-Cheng Yeh, I-Wei Lai, Pei-Yun Tsai 0001, Yuan-Hao Huang |
ISCAS | 5 |
| 2020 | Precoder Design for Transmitter Preprocessing Aided Spatial Modulated QPSK Systems using One-bit DACs and Quantized Phase ShiftersabstractIn this paper we investigate the precoder design problem of a preprocessing-aided spatial modulation multiple-input-multiple-output (PSM-MIMO) system with quadrature-phase-shift-keying (QPSK) signaling that facilitates hardware implementation using one-bit digital-to-analog-converters (DACs) and quantized phase shifters. We formulate the new design problem and propose an efficient algorithm based on a combination of alternating minimization method and mathematical programming with equilibrium constraints (MPECs). The bit-error-rate (BER) performance of the proposed one-bit precoder in comparison to a number of benchmark PSM-MIMO precoders is quantified via numerical simulations. Chiao-En Chen, Hsin-Ching Yang, Kelvin Kuang-Chi Lee, Yuan-Hao Huang |
VTC Spring | 4 |
| 2020 | The Configurable Hybrid Precoding Processor for Bit-Stream-Based mmWave MIMO SystemsabstractBeamforming technology plays an essential role in the promising millimeter wave (mmWave) massive multiple-input and multiple-output (MIMO) communications for fifth generation (5G) new radio system. Specifically, hybrid analog beamforming and digital precoding scheme can be employed to reduce the excessive radio frequency (RF) chains and data converters in the massive MIMO transceiver while still maintaining optimal spectral efficiency. However, traditional hybrid precoding architecture cannot be configured to support different system specifications, such as the number of bit streams or transmit antennas, leading to the limitation to hardware flexibility and efficiency. This article presents a configurable and low-complexity hybrid precoder based on the parallel data-stream processing in respect of system, algorithm, and architecture. The proposed algorithm was designed to avoid signal dependence between data streams so as to realize configurable precoding architecture. The performance and complexity of the proposed algorithm were also simulated and analyzed in detail in this article. Moreover, the hybrid precoding processor chip was designed and implemented based on the proposed algorithm. A novel data-processing flow was designed in the precoder so as to increase the hardware efficiency and configurability. The designed hybrid precoding chip can be configured to support one to four data streams for 16 × 16 mmWave MIMO systems. The designed precoder was implemented by using TSMC 40-nm CMOS technology. The normalized throughput achieved 11.1 M channel-matrices per second at a maximum clock frequency of 300 MHz. The area complexity is 263.5 kGE, and power consumption is 119.1 mW. Hao-Yu Cheng, Chen-Wei Chen, Chung-An Shen, Yuan-Hao Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | A Double K-best Viterbi-sphere Decoder for Trellis-coded Generalized Spatial Modulation with Multiple Code RatesabstractThe trellis-coded generalized spatial modulation (TCGSM) system with multiple code rates is investigated so as to provide configurability and to enhance efficiency in different fading channels. Various high-radix stages are then adopted to support the optimal sequence detection for different configurations. A double K-best technique is then proposed to strike a good balance between performance and complexity of Viterbi-sphere decoding. From the analysis and simulation results, 29%-69% complexity can be saved with about 1dB performance degradation for the code rate of 4/5 in the spatial domain with QPSK and 16-QAM constellations. The first K-best selection for generating branch metrics from sphere decoding can reduce the detection complexity of the constellation while the second K- best selection can be considered if a high-radix trellis is used when the number of transmit antennas is large. Zhe-Yu Wang, Pei-Yun Tsai 0001, Yuan-Hao Huang, I-Wei Lai |
ICASSP | 3 |
| 2019 | Given-Rotation-Based Generalized Eigenvalue Decomposition Processor for MU-MIMO PrecodingabstractMulti-user multiple-input multiple-output (MU-MIMO) is an important transmission technique for wireless communication systems, which increases spectral efficiency by transmitting data to several users at the same time and frequency band. The main design issue for the MU-MIMO system is to suppress co-channel interference among all users by using appropriate MIMO precoding technique. Generalized eigenvalue decomposition (GEVD) is an inevitable and high-complexity processing unit in the leakage-based precoding system with multi-antenna users. This paper presents a Givens rotation-based algorithm for hardware implementation, which can be realized by only coordinate rotation digital computer (CORDIC) processors. The proposed GEVD processor was designed for the MU-MIMO mode in the IEEE 802.11ac system. It supports eight transmitting antennas at the base station and two receiving antennas at each of four users. The GEVD processor was designed and implemented by using a TSMC 40-nm process technology. The chip synthesis results show that the proposed GEVD processor achieves a user throughput of 0.98 M matrix/sec/user, equivalent to 0.25 M channel matrices/sec, conforming to the IEEE 802.11ac standard. Zao-Fu Yang, Jung-Chun Chi, Po-Wei Fu, Jenwei Liang, Chiao-En Chen, Yuan-Hao Huang |
ISCAS | 6 |
| 2018 | Theoretical Performance Analysis Assisted by Machine Learning for Spatial Permutation Modulation (SPM) in Slow-Fading ChannelsabstractBased on spatial modulation (SM), spatial permu- tation modulation (SPM) has been recently proposed to enhance the performance of the multiple-input multiple-output (MIMO) system. SPM maps data bits to both the QAM symbol and permutation array. At successive time instants, different transmit antennas are activated according to the mapped permutation array to transmit the QAM symbol. In this work, the error rate of SPM in slow-fading channels is analyzed. The performance is first analyzed with the closed-form expression for the special case, and then is generalized to arbitrary cases by using the approximation of Gamma random variables. The machine learning algorithm is adopted to simplify the generalization and estimate the diversity. Through the analyses, we discover that by simply adding transmit antennas, the performance of SPM in slow-fading channels can be greatly enhanced due to the reduction of the time dependency. Numerical simulations demonstrate the accuracy of our analyses and show that by adding one transmit antenna, the time dependency can almost be removed, leading to around 3 dB SNR gain for the BER performance. Jhih-Wei Shih, Jung-Chun Chi, Yuan-Hao Huang, Pei-Yun Tsai 0001, I-Wei Lai |
ICC | 3 |
| 2018 | A Portable Monitoring System with Automatic Event Detection for Sleep Apnea Level-IV EvaluationabstractTo meet the demands on a comfortable screening, or even diagnostic, equipment without interfering with the sleep, this study develops a level IV portable system, equipped with two tri-axial accelerometers (TAA) measuring the thoracic and abdominal respiratory efforts, and one oximeter measuring the oxygen saturation (SpO2), to identify obstructive sleep apnea (OSA), central sleep apnea (CSA), and hypopnea (HYP) events. The prototype integrates all the hardware and software for physiological information extraction. In addition, an automatic event detection algorithm is proposed to reduce the labor-intensive work on scoring the events. Based on 63 subjects, with 80% data for training and 20% for validation, the classification accuracy of the apnea hypopnea-index (AHI) is 84.13%. The results indicate that the proposed algorithm has great potential to classify the severity of patients in clinical examinations for both the screening and the homecare purposes. Jhao-Cheng Wu, Chia-Wei Wang, Yuan-Hao Huang, Hau-Tieng Wu, Po-Chiun Huang, Yu-Lun Lo |
ISCAS | 3 |
| 2017 | A classification-based elephant flow detection method using application round on SDN environmentsabstractWe propose in this paper a classification-based elephant flow detection method for SDN environments. The proposed elephant flow detection method consists of two classifiers running on the switch and the controller, respectively. For better performance, the concept of application round is used to build the classifier running on the controller. Experimental results show that our elephant flow detection method is able to obtain high recall and F-measure. Yuan-Hao Huang, Wen-Yueh Shih, Jiun-Long Huang |
APNOMS | 1 |
| 2017 | Sphere decoding for spatial permutation modulation MIMO systemsabstractMultiple-input multiple-output (MIMO) is an essential technology for modern wireless communication systems. Spatial modulation (SM) is an evolving MIMO transmission scheme for energy-efficient massive MIMO systems. SM conveys the information of transmit antenna indices and modulated symbols in MIMO systems. A variant spatial permutation modulation (SPM) was further proposed to include transmit and time diversities by transmitting a permutation array of antenna indices during several time instants. The SPM achieves better error rate performance than the SM especially in fast fading channel. This paper investigates sphere decoding algorithms for the SPM receiver. An ordering scheme was proposed to reduce the number of visited nodes in the spherical tree search. The improved ordered sphere decoding saves about 95.6% computational complexity corresponding to 22.7 times throughput in the software implementation. Jung-Chun Chi, Yu-Cheng Yeh, I-Wei Lai, Yuan-Hao Huang |
ICC | 4 |
| 2017 | Power-aware space-time-trellis-coded MIMO detector with SNR estimation and state-purgingabstractSpace-time trellis codes (STTCs) combine channel coding and multiple-input multiple-output (MIMO) techniques to provide coding and diversity gains for wireless communication systems. The decoding complexity is extremely high because of the high density of branch metric calculations. Thus, this study presents a state-purging mechanism based on the T-algorithm to reduce the computational complexity. In the proposed mechanism, an embedded code-aided (CA) signal-to-noise (SNR) estimator provides the SNR information to determine the optimal threshold based on the count of metric normalization. The state-purging mechanism greatly reduces the complexity of branch metric calculations and produces only negligible coding gain degradation. The fabricated 4×4 STTC MIMO detector using 90 nm 1P9M CMOS technology reduces 17.62%, 15.81%, and 13.45% power consumptions for QPSK, 8PSK, 16QAM modulations, respectively, when SNR is 20 dB. Kai-Ting Shr, Chieh-Yu Chen, Jin-Wei Jhang, Yuan-Hao Huang |
ISCAS | 4 |
| 2017 | Sleep Apnea Detection Based on Thoracic and Abdominal Movement Signals of Wearable Piezoelectric BandsabstractPhysiologically, the thoracic (THO) and abdominal (ABD) movement signals, captured using wearable piezo-electric bands, provide information about various types of apnea, including central sleep apnea (CSA) and obstructive sleep apnea (OSA). However, the use of piezo-electric wearables in detecting sleep apnea events has been seldom explored in the literature. This study explored the possibility of identifying sleep apnea events, including OSA and CSA, by solely analyzing one or both the THO and ABD signals. An adaptive non-harmonic model was introduced to model the THO and ABD signals, which allows us to design features for sleep apnea events. To confirm the suitability of the extracted features, a support vector machine was applied to classify three categories - normal and hypopnea, OSA, and CSA. According to a database of 34 subjects, the overall classification accuracies were on average 75.9%±11.7% and 73.8%±4.4%, respectively, based on the cross validation. When the features determined from the THO and ABD signals were combined, the overall classification accuracy became 81.8%±9.4%. These features were applied for designing a state machine for online apnea event detection. Two event-byevent accuracy indices, S and I, were proposed for evaluating the performance of the state machine. For the same database, the S index was 84.01%±9.06%, and the I index was 77.21%±19.01%. The results indicate the considerable potential of applying the proposed algorithm to clinical examinations for both screening and homecare purposes. Yin-Yan Lin, Hau-Tieng Wu, Chi-An Hsu, Po-Chiun Huang, Yuan-Hao Huang, Yu-Lun Lo |
IEEE J. Biomed. Health Informatics | 5 |
| 2016 | Low-complexity hybrid beam-tracking algorithms and architectures for mmWave MIMO systemsabstractIn the next-generation 5G system, millimeter wave (mmWave) multiple-input multiple-output (MIMO) system utilizes massive antennas at transmitter and receiver to increase the throughput and reliability. Due to the short wavelength of mmWave signals, more antennas can be adopted to alleviate severe channel path loss. However, the increase of RF chains raises the chip cost of the transceiver. Thus, hybrid RF beamforming and baseband precoding scheme was proposed to reduce the RF chain number and complexity. This study presents a modified mmWave channel model by adding continuous drifting feature for mmWave system. This paper also proposes two pre-parallel index selection algorithms and architectures that can reduce the complexity by reusing previous precoder results and enables a parallel hardware architecture for the hybrid beamtracking. The PPIS-MIB-SOMP and bit-stream-based PPISMIB-SOMP algorithms can reduce the complexity by 74% and 49%, respectively, in the 16x16 mmWave MIMO continuous tracking system. Kai-Neng Hsu, Cheng-Gang He, Yuan-Hao Huang |
ISCAS | 3 |
| 2016 | A List Orthogonal Matching Pursuit Detector for Generalized Space Shift Keying MIMO SystemsabstractIn this paper, a new generalized space shift keying (GSSK) multiple-input-multiple-output (MIMO) detector is proposed. The proposed detector improves the conventional orthogonal matching pursuit (OMP) detector by a list detection strategy, and admits a regular structure with fixed throughput characteristic. Simulation results show that our proposed list OMP detector can achieve a more favourable performance- complexity tradeoff compared to other GSSK-MIMO detectors. Kuan-Hua Chen, Chiao-En Chen, Yuan-Hao Huang |
VTC Fall | 3 |
| 2016 | A High-SNR Projection-Based Atom Selection OMP Processor for Compressive SensingabstractCompressive sensing (CS) has recently become a critical technique to reduce the high computational cost of signal processing systems. One of the main signal processing modules of CS systems is the signal reconstruction processor. Orthogonal matching pursuit (OMP) is one of the main signal reconstruction algorithms because of its low complexity and high regularity compared with the greedy $l1$ algorithm. However, the main problem of an OMP-based signal reconstruction processor is its low reconstruction performance. Therefore, this paper proposed an enhanced projection-based atom selection orthogonal matching pursuit (POMP) algorithm to solve this problem. The original POMP algorithm has more favorable signal-to-reconstruction-noise ratio (SRNR) performance than the OMP algorithm, but the computational complexity of the POMP is too high to be realized in hardware. The proposed algorithm greatly reduces the computational complexity of the original POMP algorithm without any SRNR performance loss by using multiple index candidates and matrix-inversion-bypass method. The proposed reconstruction processor architecture was designed and implemented using the TSMC 90-nm 1P9M CMOS technology. The postlayout results showed that the proposed processor exhibited a lower normalized processing latency and achieved a higher SRNR of 34.9 dB compared with the processors in previous studies. Jin-Wei Jhang, Yuan-Hao Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | An HMM-based eye movement detection system using EEG brain-computer interfaceabstractThis paper presents an eye movement detection system based on EEG-based brain-computer interface (BCI). The proposed system performs feature extraction of the EEG signals and uses hidden Markov model (HMM) to perform eye movement detection. The proposed system is implemented on an FPGA board to form a real time BCI control system with a laptop. The users wearing a EEG headset can control a game on the computer by moving their eyeballs toward directions of left/right and up/down. This proposed system achieves a nearly 90% detection rate without training. Chi-Hsuan Hsieh, Hao-Ping Chu, Yuan-Hao Huang |
ISCAS | 3 |
| 2014 | A 2-D Interpolation-Based QRD Processor With Partial Layer Mapping for MIMO-OFDM SystemsabstractThe rapid growth of wideband communication devices, such as smart phones and tablets, has dramatically increased the throughput requirements of communication systems. Multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) has been widely recognized as the potential technology to achieve this requirement. The main implementation bottleneck lies in the high-complexity MIMO detections for thousands of subcarrier symbols in the MIMO-OFDM receiver. The QR decomposition processor is an essential preprocessing module for many MIMO decoders because the QR decomposition helps improve the decoding efficiency significantly by providing the tree-search structure for the MIMO detection algorithms. This paper presents a low-complexity 2-D interpolation-based QR decomposition (2-D-IQRD) with a partial layer mapping scheme to reduce computational complexity. The proposed 2-D-IQRD processor also adopts a scaling technique to reduce the signal dynamic-range problem in the traditional IQRD algorithm. These proposed methods successfully address the limitation of the interpolation-based QRD processor caused by successive multiplications in layer mapping and inverse layer mapping. Therefore, the 2-D interpolation can greatly reduce complexity and improve the BER performance of fast-fading MIMO-OFDM systems. This paper designs and implements the proposed architecture using a 90-nm CMOS technology. The implementation results show that the proposed architecture achieves 45.6 MQRD/s. Li-Wei Chai, Po-Lin Chiu, Yuan-Hao Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | A 3.1 Gb/s 8 × 8 Sorting Reduced K-Best Detector With Lattice Reduction and QR DecompositionabstractThis paper presents the VLSI implementation of a lattice-reduction-aided (LRA) detection system. The proposed system includes a QR decomposition, lattice reduction (LR) processor, and sorting-reduced (SR) K-best detector for 8 × 8 multiple-input multiple-output (MIMO) systems. The bit error rate of the proposed MIMO detection system only incurs approximately 3 dB of implementation loss compared with optimal maximum likelihood detection with 64-quadratic-amplitude modulation. The proposed processor can also support different throughput requirements by adjusting the stage number of LR. The SR K-best detector can achieve 3.1 Gb/s throughput with 0.24-ns latency. The throughput of the system reaches 585 Mb/s if one channel preprocessing can support 72 symbol detections. The corresponding energy per bit is 63 pJ/bit, which is the smallest value achieved to date. This paper presents the first VLSI implementation of a complete LRA K-best detector with an 8 × 8 dimension. Chun-Fu Liao, Jhong-Yu Wang, Yuan-Hao Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | 4×4/8×8/12×12 reconfigurable MIMO detector on multi-core DSP based on Eigen decomposition of dataflow graphabstractThis paper presents a 4×4/8×8/12×12 and QPSK/16-QAM/64-QAM reconfigurable MIMO detector on a multi-core DSP platform for different MIMO detection algorithms, including multiple-candidate-selection QRSIC (MCS-QRSIC), distributed K-best, and sorting-reduced K-best detectors. This study uses an Eigen-value decomposition method to investigate the intrinsic degree of parallelism of MIMO detectors and allocate the processing elements to DSP cores. MCS-QRSIC outperforms other detectors in 4 × 4 detection, and distributed K-best has the best performance in 8×8 and 12×12 MIMO systems. The normalized throughput achieves 188/153/142 Mbps for 64-QAM 4×4/8×8/12×12 MIMO detections at a 1GHz clock rate with 448 cores. Sheng-Hung Wu, Chien-Yu Kao, Jen-Yuan Hsu, Pangan Ting, Gwo Giun Lee, Yuan-Hao Huang |
ICASSP | 6 |
| 2013 | Human respiratory feature extraction on an UWB radar signal processing platformabstractThis paper presents a human respiratory feature extraction algorithm and its implementation on an ultra-wideband (UWB) impulse-radio radar signal processing platform. The conventional human detection algorithms only extract the respiration rate by the radar system. However, there is more information that is never explored in the radar-detected respiratory signals. Thus, this study proposes a modified raised cosine waveform as the respiration model and an iterative feature extraction algorithm to acquire more respiratory features, such as inspiration and expiration speeds, respiration intensity, and respiration holding ratio. These extracted features can be regarded as the compressed signals for the long-term remote medical monitoring system. The proposed respiratory feature extraction algorithm is designed and implemented on a radar signal processing platform with an Radar front-end chip, an ARM processor, and an FPGA chip. The proposed circuit can detect human respiratory signals from 0.1 to 1 Hz rate and analyze the respiratory features for each period of the respiratory signal. Chi-Hsuan Hsieh, Yi-Hsiang Shen, Yu-Fang Chiu, Ta-Shun Chu, Yuan-Hao Huang |
ISCAS | 5 |
| 2012 | A constant-throughput LLL algorithm with deep insertion for LR-aided MIMO detectionabstractLattice reduction (LR) has emerged as a popular technique to improve the performance of many low complexity MIMO detectors. Among the existing LR algorithms, the LLL algorithm has been applied to MIMO communications almost exclusively to date. However, the conventional LLL algorithm has variable reduction time, and hence non-constant throughput for the resulting communications system. In order to address this issue, this paper proposed a new variant of the LLL-algorithm (CT-LLL-deep) having constant throughput by introducing the concept of deep insertion into a parallel variant of LLL. Simulation results show that LR-aided detectors using the proposed CT-LLL-deep are capable of achieving lower bit error rate with much less reduction time, comparing to those using other existing LR algorithms. Chiao-En Chen, Chun-Fu Liao, Yuan-Hao Huang |
ISCAS | 4 |
| 2011 | Latency-constrained low-complexity lattice reduction for MIMO-OFDM systemsabstractRecent studies have investigated lattice-reduction (LR) pre processing technique for multiple-input multiple-output (MIMO) detection. However, if LR is applied to the orthogonal frequency-division-multiplexing (OFDM) system, its complexity and latency increase greatly because of the large number of sub-carriers. This paper proposes a new processing architecture for LR-aided MIMO-OFDM system. This LR processing architecture reduces the number of iteration loops by using preprocessing matrix of adjacent sub-carrier. Beside, the grouping of sub-carriers can break the long critical computational path so as to comprise the computational complexity and latency. We simulate the proposed LR-aided MIMO-OFDM processing in the 3GPP-LTE system. The pro posed method not only reduces the computational complexity but also shortens the latency for the lattice reduction. Chun-Fu Liao, Fang-Chun Lan, Yuan-Hao Huang, Po-Lin Chiu |
ICASSP | 3 |
| 2011 | Reduced-complexity interpolation-based QR decomposition using partial layer mappingabstractThis paper presents a double one-dimension interpolation-based QR decomposition (DOD-IQRD) algorithm and a partial layer-mapping (PLM) scheme for multiple input multiple-output (MIMO) orthogonal-frequency-division multiplexing (OFDM) systems. The proposed algorithm integrates the calculations of frequency-domain channel estimation and QR decomposition for MIMO-OFDM systems, and thus requires less computational complexity than traditional scheme. Besides, the proposed partial layer-mapping scheme can further reduce computational complexity with negligible performance degradation. The analysis results show that the double one dimension interpolation-based QR decomposition algorithm reduces the complexity to 45.2% of the traditional algorithm in the 3GPP-LTE 4×4 MIMO-OFDM system. The proposed QR decomposition algorithm combined with the partial layer mapping scheme can further save up to 70.5% computational complexity for MIMO-OFDM systems. Li-Wei Chai, Po-Lin Chiu, Yuan-Hao Huang |
ISCAS | 3 |
| 2010 | Design of 4 × 4 MIMO-OFDMA receiver with precode codebook search for 3GPP-LTEabstractIn this paper, we propose a VLSI design of a 4 × 4 MIMO-OFDMA down-link baseband receiver for the mobile communications. The receiver consists of synchronization, channel estimation/compensation, precode codebook search, and MIMO detection. To realize the closed-loop MIMO system in the mobile environment, we utilize the MMSE-OSIC algorithm for MIMO detector and capacity selection criterion for precode codebook searcher, in which the Strassen's algorithm is employed to reduce their matrix inversion and multiplication complexity. The proposed MIMO-OFDMA receiver is designed and evaluated using a standard 90nm CMOS technology. The results show that the receiver achieves 150Mbps data rate and consumes HOmW power at 7.81MHz, which meets the requirement of Mode I~III OFDMA modes for the 3GPP-LTE system. Chia-Ching Lee, Chun-Fu Liao, Chao-Ming Chen, Yuan-Hao Huang |
ISCAS | 4 |
| 2010 | High-Efficiency Soft-Error-Tolerant Digital Signal Processing Using Fine-Grain Subword-Detection ProcessingabstractThe soft error problem in digital circuits is becoming increasingly important as the IC fabrication technology progresses from the deep submicrometer scale to the nanometer scale. This paper proposes a subword-detection processing (SDP) technique and a fine-grain soft-error-tolerance (FGSET) architecture to improve the performance of the digital signal processing circuit. In the SDP technique, the logic masking property of the soft error in the combinational circuit is utilized to mask the single-event upset (SEU) caused by disturbing particles in the inactive area. To further improve the performance, the masked portion of the datapath can be used as the estimation redundancy in the algorithmic softerror-tolerance (ASET) technique. This technique is called subword-detection and redundant processing (SDRP). In the FGSET architecture, the soft error in each processing element (fine grain) can be recovered by the arithmetic datapath-level ASET technique. Analysis of the fast Fourier transform processor example shows that the proposed FGSET architecture can improve the performance of the coarse-grain SET (CGSET) by 8.5 dB. The low-cost SDP technique (1.03x) yields a noise reduction of 5.3 dB over the CGSET approach (1.40x), while the efficient SDRP I (1.57x) and SDRP II (1.88x) techniques outperform the CGSET approach by 24.5 and 30.5 dB, respectively. Yuan-Hao Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | A Scalable MIMO Detection Architecture with Non-sorted Multiple-candidate SelectionabstractIn this paper, we propose a QR-based MIMO detection algorithm and its architecture based on a non-sorted multiple-candidate selection process. The proposed multiple-candidate selection process can mitigate the error propagation problem in the general QR-SIC detection, and therefore the detection probability is increased. This algorithm requires only 24% of the computational complexity of the V-BLAST, which is only slightly larger than that of the conventional QR-SIC algorithm. Besides, the proposed algorithm features high flexibility between the complexity and the performance, and it can even reach the performance of ML detection for the high performance system. Furthermore, the flexible selection approach requires no sorting operation like traditional K-best algorithm. Thus, a simple scalable VLSI architecture can be constructed for different MIMO configurations. Po-Lin Chiu, Yuan-Hao Huang |
ISCAS | 2 |
| 2009 | A Modified Sorted-QR Decomposition Algorithm for Parallel Processing in MIMO DetectionabstractThis paper proposes a modified sorted-QR decomposition algorithm for the high-dimensional multiple-input multiple-output (MIMO) detection. Due to the growing demands of high-dimension MIMO channels and large number of OFDM subcarriers, the sorted-QR decomposition becomes one of the computational bottlenecks in the QR-based MIMO detection. The proposed Givens-rotation-based algorithm aims to improve the throughput and the hardware utilization efficiency by relaxing the sorting condition and allowing the cross-column parallel rotation operations. The simulation results show that our proposed parallel algorithm can significantly reduce the minimal computation latency from 28% to 21% of the original algorithm, and the hardware utilization rate can be enhanced from 71% to 90% for 16times16 MIMO using four Given-rotation processing elements. The detection performance degradation is negligible and better hardware efficiency can be obtained for the larger number of MIMO antennas. Ren-Hao Lai, Cheng-Ming Chen, Pangan Ting, Yuan-Hao Huang |
ISCAS | 4 |
| 2002 | A new audio coding scheme using a forward masking model and perceptually weighted vector quantizationabstractThis paper presents a new audio coder that includes two techniques to improve the sound quality of the audio coding system. First, a forward masking model is proposed. This model exploits adaptation of the peripheral sensory and neural elements in the auditory system, which is often deemed as the cause of forward masking. In the proposed audio coder, the forward masking is first modeled by a nonlinear analog circuit and then difference equations for finding the solution of this circuit are formulated. The parameters of the circuit are derived from several factors, including time difference between masker and maskee, masker level, masker frequency, and masker duration. Inclusion of this model in the coding process will remove more redundancy inaudible to humans and thus improves the coding efficiency. Secondly, we propose a new vector quantization technique, whose codebooks are generated by a perceptually weighted binary-tree self-organizing feature maps (PW-BTSOFM) algorithm. This vector quantization technique adopts a perceptually weighted error criterion to train and select codewords so that the quantization error is kept below the just-noticed distortion (JND) while using the smallest possible codebook, again reducing the required coded bit rate. Experimental objective and subjective sound quality measurements show that the proposed audio coding scheme requires about 30% less bits than the MPEG layer III audio coding standard. Yuan-Hao Huang, Tzi-Dar Chiueh |
IEEE Trans. Speech Audio Process. | 1 |
| 1999 | A new forward masking model and its application to perceptual audio codingabstractThis paper presents a new forward masking model for perceptual audio coding. This model exploits adaptation of the peripheral sensory and neural elements in the auditory system, which is often deemed as the cause of forward masking. Nonlinearity of the ear is modeled by a nonlinear analog circuit with difference equations. We incorporate this model in the MPEG layer III audio coding scheme and construct a masking plane in the frequency-time space. With some extra computations, the new audio coding scheme can improve the sound quality of the decoded audio signals. In our experiments, subjective and objective sound quality measurements show that, to achieve the same reconstructed sound quality, the new scheme requires 12% to 23% less bits than the original MPEG layer III scheme. Yuan-Hao Huang, Tzi-Dar Chiueh |
ICASSP | 1 |