Sau-Gee Chen

dblp:31/793 · DBLP profile ↗
← Back
38ranked-venue papers
4as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 2 since 2021Computer networks · 5

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
3 papers
Physical-layer communications · 93% Wireless networking · 7%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Physical-layer communications › synchronization › frequency synchronization
carrier frequency offset estimation
0.212014
A Blind Fine Synchronization Scheme for SC-FDE Systems · IEEE Trans. Commun. 2014
Physical-layer communications
equalization
0.212014
A Blind Fine Synchronization Scheme for SC-FDE Systems · IEEE Trans. Commun. 2014
Physical-layer communications › equalization › frequency-domain equalization
single-carrier frequency-domain equalization
0.212014
A Blind Fine Synchronization Scheme for SC-FDE Systems · IEEE Trans. Commun. 2014
Physical-layer communications › synchronization › timing estimation
symbol timing estimation
0.212014
A Blind Fine Synchronization Scheme for SC-FDE Systems · IEEE Trans. Commun. 2014
Physical-layer communications
synchronization
0.212014
A Blind Fine Synchronization Scheme for SC-FDE Systems · IEEE Trans. Commun. 2014
Physical-layer communications › MIMO
antenna selection
0.112011
Energy-Bandwidth Efficiency Tradeoff in MIMO Multi-Hop Wireless Networks · IEEE J. Sel. Areas Commun. 2011
Physical-layer communications
MIMO
0.112011
Energy-Bandwidth Efficiency Tradeoff in MIMO Multi-Hop Wireless Networks · IEEE J. Sel. Areas Commun. 2011
Wireless networking › wireless mesh network
multihop wireless network
0.112011
Energy-Bandwidth Efficiency Tradeoff in MIMO Multi-Hop Wireless Networks · IEEE J. Sel. Areas Commun. 2011
Physical-layer communications › cooperative communication
relay networks
0.112011
Energy-Bandwidth Efficiency Tradeoff in MIMO Multi-Hop Wireless Networks · IEEE J. Sel. Areas Commun. 2011
Physical-layer communications › signal detection
noncoherent detection
0.112010
Non-Coherent Cell Identification Detection Methods and Statistical Analysis for OFDM Systems · IEEE Trans. Commun. 2010
Physical-layer communications › channel estimation
blind estimation
0.112014
A Blind Fine Synchronization Scheme for SC-FDE Systems · IEEE Trans. Commun. 2014
Physical-layer communications › modulation › multicarrier modulation
OFDM
0.012010
Non-Coherent Cell Identification Detection Methods and Statistical Analysis for OFDM Systems · IEEE Trans. Commun. 2010

Methods — techniques the papers use, named apart from their topics

optimization · 0.2weighted least squares · 0.2decision feedback · 0.2cramer-rao bound · 0.2performance analysis · 0.1statistical analysis · 0.1
YearPublicationVenuePosition
2025 HDR-CNF: single-image high dynamic range imaging based on conditional normalizing flows
Kai-Wei Peng, Jui-Chiu Chiang, Sau-Gee Chen, Yu-Shan Lin
Multim. Tools Appl.3
2023 Low Routing Complexity Multiframe Pipelined LDPC Decoder Based on a Novel Pseudo Marginalized Min-Sum Algorithm for High Throughput Applications
abstract
This article presents a high throughput and low routing complexity multiframe pipelined low-density parity check (LDPC) decoder design based on a novel pseudo marginalized min-sum (PMMS) message passing approach. The proposed PMMS approach reduces the required number of interconnections in the routing network allowing the design to be implemented with reduced hardware complexity, low power consumption and high throughput capability while supporting multiple coding rates with short and long codewords as defined in many application standards. Implementation results for IEEE802.11ad/ay standards show that the proposed design satisfies a target bit error rate (BER) requirement of$3 \times 10^{-7}$with 64 quadrature amp mod (QAM) targeting high throughput applications. Furthermore, the proposed design is able to achieve a throughput of 62 and 101.8 Gb/s with two pipelining stages under 28-nm CMOS and 16-nm FinFET CMOS process, respectively. As compared to the existing NMS algorithm, the proposed design based on the PMMS approach reduces the number of wires in the routing network by 45.5%, and the wirelength of the overall decoder by 17%. The area and power consumption are also reduced by 9.4% and 12%, respectively, as compared to the conventional normalized min-sum (NMS) algorithm.
Henry Lopez Davila, Tsung-Han Wu, Shyh-Jye Jou, Sau-Gee Chen, Pei-Yun Tsai 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2022 HDR-AGAN: Ghost-Free High Dynamic Range Imaging with Attention Guided Adversarial Network
abstract
Creating a high dynamic range (HDR) image from a set of low dynamic range (LDR) images with multiple exposures is always a challenging task when there is motion among the LDR image stacks. In this paper, we propose a generative adversarial network-based method with an attention module, called HDR- AGAN, to deal with the issue caused by object motions. Given three dynamic LDR images with different exposures, the well-exposed one is treated as the reference image, while the other two images are aligned to the reference one so that two additional inputs are obtained in the proposed system. An attention module operated between the reference image and the non-reference image is employed for extracting useful features. The generated HDR image is judged by the discriminator which operates on specified image areas derived by the ghost detection module. Experimental results show that, compared to existing algorithms, our method generates HDR images with less ghost artifacts and also yields better image quality in terms of several objective metrics.
Jui-Chiu Chiang, Sau-Gee Chen
ICIP3
2020 A 128-Point Multi-Path SC FFT Architecture
abstract
This paper presents a new radix-2kmulti-path FFT architecture, named MSC FFT, which is based on a single-path radix-2 serial commutator (SC) FFT architecture. The proposed multi-path architecture has a very high hardware utilization that results in a small chip area, while providing high throughput. In addition, the adoption of radix-2kFFT algorithms allows for simplifying the rotators even further. It is achieved by optimizing the structure of the processing element (PE). The implemented architecture is a 128-point 4-parallel multi-path SC FFT using 90 nm process. Its area and power consumption at 250 MHz are only 0.167 mm2and 14.81 mW, respectively. Compared with existing works, the proposed design reduces significantly the chip area and the power consumption, while providing high throughput.
Shun-Che Hsu, Shen-Jui Huang, Sau-Gee Chen, Shin-Che Lin, Mario Garrido
ISCAS3
2019 Reconfigurable Radix-2k×3 Feedforward FFT Architectures
abstract
Due to the increasing demand for high-throughput and low-cost mobile devices, design of high-parallel reconfigurable FFT processors has become more and more important. However, FFT lengths varied, designing a multi-length FFT processor with the requirement meet has become unprecedentedly challenging, especially as the FFT lengths includes non-power-of-two. In this paper, reconfigurable mixed-radix 2k×3-point feedforward FFT architectures are proposed. It can be realized as any power-of-two parallelism to achieve the sweet spot, with performs high enough to meet the requirement and still promise a reasonable cost. A proposed feedforward radix-3 FFT is applied in the architecture, empowering the FFT processor to achieve high parallelisms. An 8-parallel 128-2048/1536-point FFT processor for the 4G LTE system is implemented with TSMC 90nm technology. Compared to the existing designs, this work offers a high-throughput and high area-efficiency solution for mixed-radix FFT operation.
Wei-Lun Tsai, Sau-Gee Chen, Shen-Jui Huang
ISCAS2
2016 A high-parallelism memory-based FFT processor with high SQNR and novel addressing scheme
abstract
This paper presents an area-efficient memory-based FFT processor for long FFT lengths. To achieve high throughput, radix-42 FFT algorithm is adopted to reduce number of FFT stages. For low-complexity realization of the main butterfly processing element, a folded-by-2 scheme along with an optimized scheduling is designed. Variable FFT lengths (i.e., 1024 ∼ 32768 points) can be supported through flexible switch configurations. Moreover, a conflict-free memory addressing scheme is devised to support 16-way parallel and normal-order data input/output without re-ordering buffers. An optimized block floating-point (BFP) scheme is employed for long-length FFT operations. The EDA synthesis results with TSMC-90nm process show that the area of proposed FFT processor is 2.98 mm2, and the power consumption is 29 mW @160MHz clock frequency. The SQNR performance is over 70dB for all supported FFT lengths with 16-bit wordlength.
Shen-Jui Huang, Sau-Gee Chen
ISCAS2
2014 An IEEE 802.15.3c/802.11ad compliant SC/OFDM dual-mode baseband receiver for 60 GHz Band
abstract
In this paper, a dual-standard, dual-mode baseband receiver for 60 GHz wireless communication is presented. The receiver is designed to support SC and OFDM modes for both IEEE 802.15.3c and IEEE 802.11ad standards. The receiver is integrated with all-digital synchronization, radix-16 FFT, phase noise cancellation and low-complexity time-domain equalizer for line-of-sight channel application. The hardware utilization achieves 70% by leveraging hardware sharing between two modes and two standards for area efficiency. The receiver is implemented with 65 nm 1P9M process in 7.95 mm2core area. With feed-through architecture, the throughput rate supports up to 7.04 Gb/s and 15.84 Gb/s for SC mode (220 MHz) and OFDM mode (330 MHz), respectively.
Wei-Chang Liu, Fu-Chun Yeh, Chia-Yi Wu, Ting-Chen Wei, Ya-Shiue Huang, Shen-Jui Huang, Ching-Da Chan, Shyh-Jye Jou, Sau-Gee Chen
ISCAS9
2014 A Blind Fine Synchronization Scheme for SC-FDE Systems
abstract
This work presents a blind fine synchronization scheme, which estimates and compensates residual carrier-frequency offset (RCFO) and symbol timing offset (STO) , for single-carrier frequency-domain equalization (SC-FDE) systems. Existing fine synchronization schemes for SC-FDE systems rely on time-domain unique words (UW) sequences as reference signals to assure the estimation accuracy, at the cost of decreased system throughput. The proposed technique, named simplified weighted least-square method for single-carrier systems (SWLS-SC), combines the decision feedback structure and SWLS estimator for OFDM systems. Together with specifically derived weighting factors, it has much better estimation accuracy than the well-known linear least-square (LLS) method for SC-FDE systems, and its BER performance can approach that of the ideal synchronization condition. The proposed technique is more effective than existing techniques, in terms of both performance and throughput. Theoretical estimation bounds are also derived to verify the effectiveness of the proposed method.
Ying-Tsung Lin, Sau-Gee Chen
IEEE Trans. Commun.2
2013 Low-complexity and high-performance non-coherent cell identification detection schemes for OFDM-based systems
abstract
This work proposes two low-complexity and high-performance cell ID detection schemes for cellular communication systems. The first one, called real-correlation multiple differential detection (RMDD), derived from our previous work on cell ID detection called CERCD method, has much less complex multiplication operations while maintains the same performance. Although CERCD algorithm is more robust than existing cell detection methods in AWGN and multipath channel conditions, its performance still can be further improved. As such, the second scheme, called multiple differential detection (MDD), is proposed to improve CERCD method. Simulation results show that MDD has much better performance in frequency-selective channels. Performances and computational complexities of proposed schemes are also evaluated and analyzed under different channel environments to demonstrate their effectiveness.
Ying-Tsung Lin, Yi-Hsiang Wang, Sau-Gee Chen, Chih-Liang Chen
ICASSP3
2013 A SC/HSI dual-mode baseband receiver with frequency-domain equalizer for IEEE 802.15.3c
abstract
In this paper, an 8X-parallelism digital baseband receiver is proposed for IEEE 802.15.3c application. The baseband receiver consists of all-digital synchronization, radix-16 FFT and LS-LMS equalizer modules. It supports SC and HSI dual-mode in IEEE 802.15.3c with single hardware for area efficiency. The chip is implemented with 65 nm 1P9M process. The fabricated area is 12.96 mm2with 3463 K gate counts. The post-layout verification shows the throughput rate under QPSK modulation achieves 3.52 Gb/s and 5.28 Gb/s for SC mode (220 MHz) and HSI mode (330 MHz), respectively.
Wei-Chang Liu, Fu-Chun Yeh, Ting-Chen Wei, Ya-Shiue Huang, Tai-Yang Liu, Shen-Jui Huang, Ching-Da Chan, Shyh-Jye Jou, Sau-Gee Chen
ISCAS9
2012 A memory-efficient continuous-flow FFT processor for Wimax application
abstract
This paper presents an area-efficient continuous-flow FFT/IFFT processor for Wimax application. Especially, the integration design of FFT processors with other functional blocks is considered according to Wiamx specification. A new memory scheduling scheme targeted for UL-PUSC transmission mode is proposed to reduce about 50% total memory space requirement compared with the conventional ping-pong based approach. In addition, a cascaded and pipelined PE structure is designed for high speed operation. The structure is configurable to support different FFT size efficiently, especially for those non-power-of-8 DFT operations. The EDA synthesis results show that the proposed FFT processor occupies only 1.62 mm2area based on TSMC 0.18-um process.
Shen-Jui Huang, Sau-Gee Chen
ISCAS2
2012 An efficient blind fine synchronization scheme for SCBT systems
abstract
This work presents a new technique which can blindly perform the fine synchronization for single-carrier block-transmission (SCBT) systems without using unique words (UW). The proposed technique, called simplified weighted LS method for single-carrier systems (SWLS-SC), combines the concept of the decision feedback structure and SWLS estimator for OFDM systems. Together with specifically derived weighting factors, the proposed scheme has much better estimation accuracy than the well-known linear least-square (LLS) method for SC systems, while its BER performance can approach that of the ideal synchronization condition. With the proposed estimator, the system throughput can be further increased, because there is no need to insert UW for fine synchronization. The theoretical estimation bounds are also derived to prove the effectiveness of the proposed scheme.
Ying-Tsung Lin, Sau-Gee Chen
ISCAS2
2012 A New Genetics-Aided Message Passing Decoding Algorithm for LDPC Codes
abstract
The popular LDPC decoding algorithms based on the message passing (MP) algorithm have high decoding performances. However, they are noticeably inferior to the maximum likelihood (ML) decoding algorithm. This work proposes a genetics-aided message passing (GA-MP) algorithm by applying a new genetic algorithm to MP algorithm. As a result, significantly performance improvement over MP algorithm can be achieved. Besides, compared with other genetic-aided decoding algorithms, the proposed algorithm has much better performances and much lower computational complexity. Simulations show that the decoding performance of GA-MP algorithm can achieve performances very close to the algorithm, while outperform MP algorithm. Besides, its performance will grow proportionally with the generation number without leveling off as observed in conventional MP algorithms, under high SNR condition.
Jui-Hui Hung, Yi-De Lu, Sau-Gee Chen
VTC Fall3
2011 Energy-Bandwidth Efficiency Tradeoff in MIMO Multi-Hop Wireless Networks
abstract
This paper considers a MIMO multi-hop network and analyzes the relationship between its energy consumption and bandwidth efficiency. Its minimum energy consumption is formulated as an optimization problem. By taking both transmit antennas (TAs) and receive antennas (RAs) into consideration, the energy-bandwidth efficiency tradeoff in the networks is investigated. Moreover, the minimum energy of an equally-spaced relaying strategy is investigated for various numbers of antennas. In addition, the minimum energy over all possible antenna pairs is derived. Finally, the effect of the number of hops on the energy-bandwidth efficiency tradeoff is considered. For a fixed antenna pair, the minimum energy over all possible rates and hop numbers are obtained. Generally, the routes with more hops minimize the energy consumption in the low effective rate region. On the other hand, in the high effective rate region, the routes with fewer hops minimize the energy consumption.
Chih-Liang Chen, Wayne E. Stark, Sau-Gee Chen
IEEE J. Sel. Areas Commun.3
2010 A Real-Time High-Throughput LDPC Decoder for IEEE 802.3an Standard
abstract
The existing LDPC decoders are mostly based on the belief-propagation (BP) algorithms, due to good BER performances. However, they demand large chip areas. This paper proposes a high-throughput LDPC decoder based on the bit-flipping algorithms, for the (2048, 1723) RS-LDPC code adopted in the IEEE 802.3an standard. High decoding performances and low iteration numbers are achieved by introducing a strategy of flipping low-correlation bits and an additional syndrome vote scheme. As a result, the decoding performance is very close to the most popular BP-based min-sum algorithm (MSA) but with much lower computational complexity. Besides, the decoder achieves high hardware utilization with real-time processing capability. Synthesized with UMC 90nm process, the decoder chip area, throughput and average power dissipation are 1.42M gates, 16Gbps and 368mW, respectively, at 500MHz clock rate. Compared with existing BP-based designs, it has much smaller chip area and lower power dissipation, with comparable performances.
Jui-Hui Hung, Li-Wei Kao, Sau-Gee Chen
VTC Spring3
2010 A Data Detection Scheme for Single-Carrier Block Transmission Using Sphere Decoding Algorithm
abstract
In this paper, an efficient sphere decoding algorithm (SDA) applied to solve the inter-symbol-interference (ISI) data detection problem is proposed. This proposed algorithm takes advantages of the SDA to obtain the solution close to ML solution with the simplified K-best tree search method for the ISI data detection under ISI effect. Compared with the frequency-domain MMSE or Zero-Forcing equalization techniques for SCBT (Single-Carrier Block Transmission) systems, this algorithm can perform at least 1.5dB better than frequency- domain equalizers under the environment of randomly generated channel impulse responses.
Ying-Tsung Lin, Chia-Hsun Kuo, Sau-Gee Chen, Wai-Chi Fang
VTC Spring3
2010 Non-Coherent Cell Identification Detection Methods and Statistical Analysis for OFDM Systems
abstract
To achieve high-performance cell identification detection with low complexity for orthogonal frequency division multiplexing (OFDM) systems, this work proposes an efficient non-coherent cell identification (ID) detection technique based on a new optimization metric. Furthermore, the metric is simplified to two lower-complexity metrics. As such, two more modified cell ID detection methods extended from the first one are also proposed. Experiments show that all the three new cell ID methods achieve better performances than the conventional methods in multipath Rayleigh fading channels. This work also conducts the statistical analysis to characterize the proposed techniques completely. It is shown that the results of the theoretical analysis are close to the simulation results. Among the new techniques, specifically, the third proposed method using the second simplified new metric has much lower complexity and higher performance than the conventional methods.
Chih-Liang Chen, Sau-Gee Chen
IEEE Trans. Commun.2
2009 Symbol Time Synchronization Based on SINR Maximization for OFDM
abstract
This work presents a symbol time (ST) synchronization algorithm for orthogonal frequency-division multiplexing systems based on a metric of signal-to-interference-and-noise ratio (SINR) which drops drastically with the increasing ST error. The ST is intuitively estimated by maximizing the SINR metric that exploits the correlation results of the separated-by-Nsamples. Unlike most existing techniques considering only AWGN or time-invariant multipath channels, the proposed maximum-SINR (MSINR) technique considers the time-variant multipath channels. Compared with conventional techniques, the proposed MSINR technique is less sensitive to the carrier frequency offset and doubly-selective channel effects.
Wen-Long Chin, Sau-Gee Chen
VTC Spring2
2009 On the Analysis of Combined Synchronization Error Effects in OFDM Systems
abstract
This work analyzes combined effects of major synchronization errors, including the symbol time offset (STO), carrier frequency offset (CFO) and sampling clock frequency offset (SCFO) of orthogonal frequency-division multiplexing (OFDM) systems. Such errors degrade the performance of an OFDM receiver by introducing inter-carrier interference (ICI) and inter-symbol interference (ISI) into the systems. Traditionally, designing an OFDM receiver needs plenty of Monte Carlo simulations because the synchronization errors are simultaneously inevitable in practical environments. Therefore, we formulate the theoretical signal-to-interference-and-noise ratio (SINR) to assist the design of OFDM receivers. By knowing the required SINR of specific application, all combinations of allowable errors can be derived. Then, cost-effective algorithms could be easily designed.
Wen-Long Chin, Sau-Gee Chen
VTC Spring2
2008 A Blind Maximum-SINR Synchronization Technique for OFDM Systems
abstract
This work presents a blind synchronizer for orthogonal frequency-division multiplexing (OFDM) systems based on signal-to-interference-and-noise-ratio (SINR) maximization. Due to the incurred losses from inter-symbol interference (ISI) and inter-carrier interference (ICI) introduced by synchronization errors, the SINR of the received data drops drastically. By taking advantage of this characteristic, both the symbol time and carrier frequency offsets are intuitively estimated by maximizing the SINR metric. For the SINR metric, we propose a blind SINR estimation, which does not need prior knowledge of the channel profiles and transmitted data. As such, the proposed maxi-mum-SINR (MSINR) synchronization algorithm is non-data-aided (NDA) so that the transmission efficiency can be improved. Moreover, to reduce the computational complexity, the early-late gate (ELG) technique is proposed for the implementation of the synchronizer. Simulation results exhibit better performance for the MSINR algorithm than conventional techniques in multipath fading channels.
Wen-Long Chin, Sau-Gee Chen
GLOBECOM2
2008 A low-complexity symbol time estimation for OFDM systems
abstract
Conventional symbol time (ST) synchronization algorithms for orthogonal frequency-division multiplexing (OFDM) systems mostly are based on maximum correlation result of the cyclic prefix. Due to the channel effect, one needs to further identify the channel impulse response (CIR) so as to obtain a better ST estimation. Overall, the required computational complexity is high because it involves correlation operation, as well as the fast Fourier transform (FFT) and inverse FFT (IFFT) operations. In this work, without the FFT/IFFT operations and the knowledge of CIR, a low-complexity time-domain ST estimation is proposed. We first characterize the frequency-domain interference effect as a function of the ST by deriving some analytical equations considering the channel effect. Based on the derivation, the new method locates the symbol boundary at the sampling point with the minimum interference in the frequency-domain. Moreover, for reducing the computational complexity, the proposed frequency-domain minimum-interference metric is converted into a low-complexity time-domain metric by utilizing the Parseval’s theorem and the sampling theory. Simulation results exhibit high performances for the proposed algorithm in the multipath fading channels.
Wen-Long Chin, Sau-Gee Chen
ICASSP2
2008 An efficient cell search algorithm for preamble-based OFDM systems
abstract
For cellular mobile telecommunication, cell identification is an indispensable initial process to establish a communication link. Cell identification requires high computational complexity and accuracy in a short time, especially during the handover process. To achieve high-performance cell identification with low-complexity for the orthogonal frequency division multiplexing (OFDM) systems, this work proposes an efficient cell ID search technique. Simulation results show that the proposed cell ID search method has better performance and lower computational complexity than conventional methods, in multipath Rayleigh fading channels.
Chih-Liang Chen, Sau-Gee Chen
PIMRC2
2008 Theoretical Analysis of Joint Synchronization Error Effects for OFDMA Systems
abstract
Few papers investigate synchronization error effects for the uplink of orthogonal frequency-division multiple access (OFDMA) systems. By contrast, this paper simultaneously analyzes the joint effects of major synchronization errors, including symbol time (ST) offset (STO), carrier frequency offset (CFO) and sampling clock frequency offset (SCFO) for the uplink of OFDMA systems in time-variant multipath fading channels. Such errors degrade the performance of an OFDMA receiver by introducing inter-carrier interference (ICI), inter-symbol interference (ISI) and multiple-access interference (MAI) into the systems. A theoretical signal-to-interference-and-noise ratio (SINR) is formulated to characterize the losses due to synchronization errors in time-variant multipath fading channels. The results provide designers a useful reference in designing suitable synchronization algorithms for the OFDMA applications.
Wen-Long Chin, Sau-Gee Chen
VTC Spring2
2007 A Joint Synchronization Algorithm for OFDM Systems
abstract
This work presents a joint maximum-likelihood (ML) synchronization algorithm for symbol time (ST) offset (STO), carrier frequency offset (CFO) and sampling clock frequency offset (SCFO) for OFDM systems. Unlike most existing techniques which are time-domain approaches considering only AWGN and/or static multipath conditions, the proposed algorithm is developed in frequency-domain under time-variant multipath channels. By analyzing the received frequency- domain data, a mathematical model for the joint effects of STO, CFO and SCFO is derived. The results are used to formulate a log-likelihood function of two consecutive symbols. Based on the function, a joint ML algorithm is proposed. The joint algorithm is both efficient and robust, because the three main synchronization issues are treated altogether. Simulation results exhibit high performances in time-variant multipath fading channels.
Wen-Long Chin, Sau-Gee Chen, Chih-Liang Chen
PIMRC2
2006 Efficient CORDIC Designs for Multi-Mode OFDM FFT
abstract
In this paper, we propose a new CORDIC algorithm and architectures which can generate close-to- optimum rotation sequences easily with small lookup table sizes. This new design is particularly suitable for the applications of adjust- able-length FFT. In all, the required number of shift-and- add operations for micro-rotations and scale-factor compensations is only n/2, where n is the output precision. For design sign verification, we synthesized both serial and pipelined architectures, by using Synopsys Design Complier based on UMC 0.18 µm, 1P6M CMOS technology. The synthesized 16-bit pipelined FFT PE runs at 222MHz, with a total gate count of 89263 and a low-power consumption of 26.75 mW. It meets the FFT speed requirements of most OFDM-based communication systems, including DAB, DVB, 802.16 and VDSL. Compared with a conventional multiplier-based FFT PE and the existing CORDIC-based FFT PE's, the proposed designs has better performances in terms of area, speed and power consumption.
Cheng-Ying Yu, Sau-Gee Chen, Jen-Chuan Chih
ICASSP (3)2
2006 An efficient LDPC code structure combined with the concept of difference family
abstract
In this paper, we construct a new irregular LDPC code structure combined with a concept of difference family [5]. The corresponding decoders for the proposed LDPC code structure achieve better performance than those decoders for random-structured codes and the quasi-cyclic codes, when simulated with the commonly used sum-product iterative decoding algorithm and the min-sum based algorithms. Meanwhile the resulting codes have both low encoding and decoding complexities comparable to those of the efficient quasi-cycle LDPC codes. Specifically, we propose a new code structure for parity-check matrices and devise new sets of the difference families for the new matrices. An LDPC decoder based on the proposed code structure has been realized with UMC 0.18 μm ASIC process technology. The decoder assumes a code rate of 3/4, a code length of 960 bits, with maximum 10 decoding iterations. The decoder can achieve a decoding throughput rate up to 370Mbps and take an area of 800k gates.
Yuan-Jih Chu, Sau-Gee Chen
IWCMC2
2004 DCT-based channel estimation for OFDM systems
abstract
In this paper, based on the property of channel frequency response and the concept of interpolation in transform domain, we propose two discrete cosine transform (DCT)-based pilot-symbol-aided channel estimators, which can mitigate the aliasing error and high-frequency distortion of the direct discrete Fourier transform (DFT)-based channel estimators when the multipath fading channels have non-sample-spaced path delays. Both proposed estimators outperform the conventional DFT-based channel estimators. Of these two DCT-based estimators, one has its performance close to MMSE estimator, while the other one has the advantage of easy implementation with a little performance degradation. Furthermore, in implementation, the DCT-based estimators have the advantages of utilizing mature fast DCT algorithms and architectures, which is favorable to matrix-based channel estimators.
Yen-Hui Yeh, Sau-Gee Chen
ICC2
2004 Reduction of Doppler-induced ICI by interference prediction
abstract
Although wireless OFDM systems are robust against frequency-selective fading and ISI effects, their relatively long symbol lengths are vulnerable to time-selective fading and ICI (intercarrier interference) effects in mobile environments due to Doppler spread. However, commonly, a time-selective fading channel is linearly time-varying from symbol to symbol for existing wireless communication systems and vehicle speed limits. The linear property greatly reduces channel parameters. Based on this condition, we propose an ICI-reduction method. The ICI-reduction method first solves the linear channel parameters of each path, and then combines with a decision-feedback method to predict the ICI values from the estimated channel information, followed by the subtraction of the ICI components. Further, the proposed algorithm adopts an iterative ICI refinement technique to enhance the ICI prediction accuracy. The simulation results show that the new method can effectively reduce the ICI and error floor of the symbol error rate. The proposed algorithm is further simplified. The simplified version still maintains almost the same performance as before, with a much reduced complexity.
Yen-Hui Yeh, Sau-Gee Chen
PIMRC2
2003 Efficient channel estimation based on discrete cosine transform
abstract
Channel impairment caused by multipath reflections can deeply degrade the transmission efficiency in wireless communication systems. Based on the property of the channel frequency response and the concept of interpolation, a DCT-based pilot-aided channel estimator for orthogonal frequency division multiplexing is proposed. This approach can mitigate the aliasing effect in the DFT-based channel estimator when there is non-sample-spaced path delay. Compared with a DFT-based estimator, the DCT-based estimator significantly improves the performance with a comparable complexity. In addition, a noise reduction scheme is introduced and combined with the estimator. In implementation, the DCT-based estimator has the advantages of utilizing mature fast DCT algorithms and compatible FFT algorithms, which is favorable to other matrix-based channel estimation methods.
Yen-Hui Yeh, Sau-Gee Chen
ICASSP (4)2
2000 A fast CORDIC algorithm based on a novel angle recoding scheme
abstract
This work proposes a new CORDIC algorithm, which considerably reduces rotation number. It is achieved by combining several design techniques. Particularly, a new angle recoding scheme for table lookup is developed to speed up the convergence rate of the rotation angle. The required lookup table is small. Other design techniques used include: the leading-one bit detection and the variable-scale-factor compensation algorithm. The number of the shift-and-add operations required in the compensation algorithm can be also further reduced by using the same residue recoding scheme. Simulations show that on average the new design needs only 3.5 iterations to generate results with 22-bit precision, which is much less than the existing designs (normally need 22 iterations). The new encoding scheme can be applied to other iterative convergence computation function such as the division operation.
Jen-Chuan Chih, Sau-Gee Chen
ISCAS2
2000 A novel iterative design technique for linear-phase FIR half-band filters
abstract
This paper proposes a novel iterative design scheme for generating linear-phase, half-band FIR filters with considerably reduced arithmetic operations and narrow transition bandwidth. Starting from a short half-band subfilter, preferably a multiplier-free one, the scheme can evolve this subfilter to a better composite half-band filter by using simple algebraic composition transformation iteratively. Compared to the existing design methods, the proposed scheme is free of filter optimization processes and is much simpler. In implementation, the designed half-band filters can be realized as a multi-stage cascaded structure, composed of several short half-band subfilters and delay elements. Since its building elements have trivial coefficients with short wordlengths, the proposed structure for half-band filters has low complexity and low coefficient dynamic range. Thus it is suitable for finite-precision implementation. Furthermore, the structure is highly modular, repetitive in all stages, and suitable for VLSI implementation.
Min-Chi Kao, Sau-Gee Chen
ISCAS2
2000 Design of finite-word-length FIR filters with least-squares error
Yung-An Kao, Sau-Gee Chen
Signal Process.2
1997 On the convergence and MSE of Chen's LMS adaptive algorithm
abstract
The previously proposed Chen's (see IEEE Trans. Circuits Syst.-II: Analog and Digital Signal Processing, vol.43, p.372-8, May 1996) LMS algorithm costs only half the multiplications of the conventional direct-form LMS algorithm (DLMS). Despite of this merit, the algorithm lacked a rigorous theoretical analysis. This article characterizes its properties and conditions for mean and mean-square convergences. The closed-form MSE is derived, which is slightly larger than that of the DLMS algorithm. It is shown, under the condition that the LMS step size /spl mu/ is very small and an extra compensation step size /spl alpha/ is properly chosen, that Chen's algorithm has a comparable performance to that of the DLMS algorithm. For the algorithm to converge, a tighter bound than before is also derived. The derived properties and conditions are verified by simulations.
Sau-Gee Chen, Yung-An Kao, Ching-Yeu Chen
ICASSP1
1997 A radix-4 redundant CORDIC algorithm with fast on-line variable scale factor compensation
abstract
In this work, a fast radix-4 redundant CORDIC algorithm with variable scale factor is proposed. The algorithm includes an on-line scale factor decomposition algorithm that transforms the complicated variable scale factor into a sequence of simple shift-and-add operations and does the variable scale factor compensation in the same fashion. On the other hand, the on-line decomposition algorithm itself can be realized with a simple and fast hardware. The new CORDIC algorithm has the smallest number of 0.8 n iterations among all the CORDIC algorithms, which requires only about two-third rotation number that of the existing best (hybrid radix-2 and radix-4) redundant algorithms. Therefore, the new algorithm achieves fast rotation iterations, high-speed and low-overhead scale factor compensations, which are hard to attain simultaneously for the existing algorithms. The on-line scale factor compensation can be also applied to the existing on-line CORDIC algorithms.
Chieh-Chih Li, Sau-Gee Chen
ICASSP2
1995 High quality and low complexity pitch modification of acoustic signals
abstract
A high-quality and low-complexity algorithm for pitch modification of acoustic signals is proposed. It is high quality, because time-domain waveform shape and phase synchrony of a synthesized signal closely resemble that of its original signal. Low complexity is attained by performing the algorithm entirely in time domain with fast algorithm, and without resorting to complicated frequency domain analysis and synthesis. The time domain synthesis mostly consists of fast algorithms of finding minimum absolute error (MAE) and a cross fading operation for high correlation gain and phase synchrony. The MAE-based algorithm is shown to yield smaller complexity, and better performance, than other well-known correlation cost functions. Pitch modification is performed by selecting and synthesizing apposite frames from original input signals in a way such that the original waveform shape and phase synchrony are preserved. Subjective tests showed a comparable performance to that of the best known algorithms, but at much reduced complexity.
Gang-Janp Lin, Sau-Gee Chen, Terry Y. Wu
ICASSP2
1994 New Systolic Arrays for Matrix Multiplication
abstract
In this paper, three new systolic arrays for matrix multiplication are proposed. The first systolic array has the minimum number of 3n-2 clock cycles in completing a matrix multiplication among the known structures, with n^2 processors elements (PE's). It is achieved by applying a new input data flow and deposition scheme. The second array is derived by combining the data flow technique with the simple Blahut's matrix multiplication algorithm. Not only the second array has the least amount of processing time of3n-2 clock cycles, it has the least area complexity of about n^2 /2 PE's. By further modifying its input data flow patterns, the third array is obtained. Its processing time is further reduced to 2.5n-2 clock cycles. The proposed architectures exhibit better performances than the known structures, according to several standard performance measures.
Sau-Gee Chen, Jiann-Cherng Lee, Chieh-Chih Li
ICPP (2)1
1994 A New Efficient 2-D LMS Adaptive Filtering Algorithm
abstract
A new efficient 2-D LMS adaptive filtering algorithm is proposed in this work. The new algorithm costs N/sup 2//2-1 less multiplications, at the expense of N/sup 2//2 more additions than the existing 2-D direct-form LMS algorithms per adaptive iteration. Performance of the new algorithm is shown to be comparable to that of the known algorithms by theoretical analysis and simulation results, in the aspects of SNR and convergence rate. The new algorithm has N/sup 2/+1 coefficients, in contrast to N/sup 2/ coefficients of the conventional algorithms.>
Sau-Gee Chen, Yung-An Kao
ISCAS1
1991 Efficient implementation of the normalized recursive least-square lattice filter
abstract
An efficient hardware implementation is proposed for the optimal, computationally intensive, normalized recursive least-squares lattice (NLSL) adaptive filter. NLSL operation steps are optimized for better hardware utilization. To best match the execution steps, a combined processing unit for radix-2 division and square-root operations, and a radix-4 MSB-first multiplication unit based on signed digit (SD) arithmetic are proposed. The proposed NLSL implementation achieves both good area and time performance. The approach is shown to be better than the CORDIC approach, and SD arithmetic is a good choice for implementing the complicated signal processing algorithms.>
Sau-Gee Chen, Jih-Feng Lin
ICASSP1