EDBT 2026 Demo / reviewers in the wild / expert
Chung-An Shen
dblp:05/8309
· DBLP profile ↗
27ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-0628-5129ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 4 first-author · 9 since 2021Computer networks · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design and Implementation of a Low Latency and Low Complexity Reed-Solomon Decoder for IEEE 802.15.4a/z IR-UWB
You-Jie Siao, Chung-An Shen |
ISCAS | 2 |
| 2026 | Efficient 3-D Convolutional Neural Network Model and VLSI Architecture for Low-Latency and Low-Complexity ProcessingabstractThree-dimensional convolutional neural networks (3D-CNNs) have shown remarkable performance in applications such as point cloud labeling, semantic segmentation, and human action recognition (HAR). However, conventional 3D-CNN models, such as C3D, suffer from high computational and storage demands, incurring significant challenges for hardware acceleration. This article presents an efficient 3D-CNN framework based on 3-D depthwise separable convolution (3DDSC), which significantly reduces the number of parameters and computational complexity. Furthermore, a novel pooling-ahead (PA) technique is proposed to minimize memory access overhead. Based on the proposed 3D-CNN framework, a low-latency and low-complexity VLSI architecture is designed. The proposed design is implemented using the TSMC 40-nm technology and validated on a field-programmable gate array (FPGA) platform. The experimental results show that the FPGA implementation achieves a$6\times $reduction in hardware resource utilization and a$1.7\times $improvement in latency compared to state-of-the-art designs. Moreover, the application-specific integrated circuit (ASIC) implementation attains$1.86\times $lower complexity and$1.1\times $faster latency, further demonstrating the efficiency and scalability of the proposed architecture. Yung-Yi Gu, Liang-Yu Li, Chung-An Shen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2026 | Algorithm and Architecture of a Channel-Correlation-Resistant Massive MIMO Detector Based on Expectation PropagationabstractExpectation propagation (EP) and its approximation algorithms achieve near-optimal signal detection in massive multiple-input multiple-output (MIMO) systems with substantially reduced complexity under uncorrelated channels. However, approximate EP (EPA) detectors often suffer from severe performance degradation in the presence of channel correlation. This article presents a novel EPA-correlation-resistant Chebyshev-accelerated Richardson iteration (EPA-CR-CARI) MIMO detection algorithm that achieves near-optimal performance without requiring channel-specific processing, which incurs considerable latency and area overhead even under correlated channels. Furthermore, a corresponding low-latency and low-complexity VLSI architecture is presented. Compared to the state-of-the-art design targeting correlated channels, the proposed MIMO detector achieves an 18.7% reduction in latency and a 17.3% decrease in hardware complexity. Yi-Ling Tsai, Wenn-Yi Lin, Chung-An Shen, Yuan-Hao Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | The Algorithmic and Architectural Optimizations for Highly Efficient Tensor Decomposition EnginesabstractThis paper presents efficient algorithm and architectures for tensor decomposition based on field programmable gate array (FPGA). Based on the parallel HOOI (PHOOI) algorithm, a circuit structure is presented to reduce the complexity and to improve the efficiency of tensor decomposition engine. Furthermore, a novel algorithm and architecture are proposed to reduce the number of iterations and enhance the throughput. The proposed circuit achieves parallel processing with a minimum increment of hardware complexity. Both proposed architectures are designed and implemented on the FPGA platform. Evaluation results show that the efficiency of proposed architectures is enhanced by 40% and 60% respectively. I-Ting Tsai, Chung-An Shen |
ISCAS | 2 |
| 2025 | Single-Cycle Independent Component Analysis Processor for In-Band Full Duplex SystemsabstractThis paper presents a single-cycle independent component analysis (ICA) algorithm and architecture for self-interference cancellation (SIC) in In-band Full-duplex (IBFD) communication systems. The proposed algorithm, AICA-EBM, incorporates the adaptive momentum (ADAM) approach with the entropy-bound estimation (EBM) to achieve rapid convergence. The simulation results show that the proposed AICA-EBM algorithm leads to a constant single-cycle processing time to achieve satisfactory SIC optimality in IBFD systems. The architecture of the ICA processor based on the proposed AICA-EBM algorithm is designed. Novel circuit structures are proposed to reduce the complexity of the ICA processor. The AICA-EBM processor is implemented by following an application-specific integrated circuit (ASIC) flow with the TSMC 90 nm process. The post-layout estimations show that Compared to previous ICA processors, the proposed design achieves a 2x improvement in processing throughput and leads to the best hardware efficiency. Hao-Lun Weng, Chung-An Shen, Mohamed E. Fouda, Ahmed M. Eltawil |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | The Design of a Low-latency Tensor Decomposition Algorithm and VLSI ArchitectureabstractTensor decomposition has emerged as an essential means for applications with multidimensional signals such as data compression and feature extraction. However, due to the enormous amount of signal processing and storage, it is challenging to design a low-latency tensor decomposition processor with high hardware efficiency. This paper presents a low-latency algorithm and VLSI architecture for a tensor decomposition processor. The designed circuit is implemented based on the FPGA platform. The estimation results demonstrate that the proposed architecture achieves a low-latency performance and high hardware efficiency. Yu-An Chen, Chung-An Shen |
ISCAS | 2 |
| 2024 | The Algorithm and VLSI Architecture of High-Throughput and Highly Efficient Tensor Decomposition EngineabstractTensor decomposition is critical for compressing data and extracting key features in novel high-dimensional signal processing systems. However, due to the enormous amount of data and the highly complicated computations, designing an efficient tensor decomposition processor is very challenging. This paper presents the algorithm and VLSI architecture design of a low-latency and high-throughput tensor decomposition processor. A parallel higher-order orthogonal iteration (P-HOOI) algorithm is proposed where multiple updated matrices are computed concurrently. Thus, a low-latency tensor decomposition is achieved. Furthermore, a novel VLSI architecture is presented so that the efficiency of the component utilization is improved and the hardware complexity is greatly reduced. Therefore, the proposed tensor decomposition processor enhances the processing throughput with minimum employment of hardware components. Performance evaluations based on the post-layout estimations in the ASIC flow and based on the FPGA platform are reported in this paper. Compared with state-of-the-art designs in the literature the proposed tensor decomposition engine greatly enhances the throughput and hardware efficiency. Ting-Yu Tsai, Chung-An Shen, Tsung-Lin Wu, Yuan-Hao Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | High-Throughput Independent Component Analysis Processor for Full Duplex SystemsabstractThis paper presents the algorithm and very-large-scale integration (VLSI) architecture of a high-throughput and highly efficient independent component analysis (ICA) processor for self-interference cancellation (SIC) in in-band full-duplex (IBFD) systems. This is the first VLSI architecture reported in the literature based on the state-of-the-art entropy bound minimization (EBM) approach. A novel ICA algorithm is presented in this paper with momentum gradient descent optimization. Simulation results show that the number of iterations for the proposed algorithm is significantly reduced compared to the conventional ICA algorithms. Furthermore, a novel early-distribution estimation scheme is proposed in the designed ICA processor to compute multiple distribution functions with low latency and low complexity. The processing flow and the efficiency for the hardware utilization are specifically designed so that the processing speed is maximized with minimum employment of hardware components. The proposed ICA processor is designed and implemented based on the application-specific-integrated circuit (ASIC) flow. The post-layout estimations show that compared with the conventional EBM-based scheme, the proposed design improves the throughput and efficiency by 30x. In addition, compared to prior designs shown in the literature, the proposed ICA processor also demonstrates a significant enhancement in terms of throughput and efficiency. Jen-Hao Cheng, Tien-Min Chang, Chung-An Shen, Mohamed E. Fouda, Ahmed M. Eltawil |
IEEE J. Sel. Areas Commun. | 3 |
| 2022 | Tensor-Based Hybrid Precoding Processor for 8 × 8 × 8 mmWave 3D-MIMO SystemsabstractHybrid baseband precoding and RF beamforming is a highly efficient technology for millimeter-wave (mmWave) massive multiple-input multiple-output (MIMO) systems. Threedimensional (3D) MIMO system with uniform planar array (UPA) of transit antennas and uniform linear array (ULA) of receive antennas can provide more flexible and efficient beamforming capability in sparse mmWave channels. Tensor is a compact multi-way algebraic model that can describe high-dimension systems such as the sparse mmWave 3D-MIMO system. This paper proposes a tensor-based hybrid precoding algorithm for continuous time-drifting 3D-MIMO systems which can achieves better performance in high bit-stream and low SNR systems. The FPGA implementation of the tensor-based hybrid precoding processor can support the 3D-MIMO system with 8 × 8 UPA transmitter and 8-antenna ULA receiver with a maximal normalized throughput of 17.0 M matrices/sec compared to the existing counterparts. Tsung-Lin Wu, Chung-An Shen, Yuan-Hao Huang |
ISCAS | 2 |
| 2022 | Configurable Independent Component Analysis Preprocessing AcceleratorabstractAn independent component analysis (ICA) has been used in many applications, including self-interference cancellation (SIC) for in-band full-duplex (IBFD) wireless systems and anomaly detection in industrial Internet of Things (IoT). This article presents a high-throughput and highly efficient configurable preprocessing accelerator for the ICA algorithm. The proposed ICA accelerator has three major blocks that perform data centering, covariance matrix for computation, and eigenvalue decomposition (EVD). Specifically, the proposed accelerator is based on a high-performance matrix multiplication array (MMA). The proposed MMA architecture uses time-multiplexed processing, so that the efficiency of hardware utilization is greatly enhanced. Furthermore, the processing flow utilizes parallel processing, such that the centering, the calculation of the covariance matrix, and the EVD are conducted simultaneously and are individually pipelined to maximize throughput. This article presents the architecture, circuit design, and performance estimates based on post-layout extraction of the proposed preprocessing ICA accelerator. The proposed design achieves a throughput of 40.7 kMatrices/s at a complexity of 73.3 kGE. Hsi-Hung Lu, Chung-An Shen, Mohamed E. Fouda, Ahmed M. Eltawil |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | The Configurable Hybrid Precoding Processor for Bit-Stream-Based mmWave MIMO SystemsabstractBeamforming technology plays an essential role in the promising millimeter wave (mmWave) massive multiple-input and multiple-output (MIMO) communications for fifth generation (5G) new radio system. Specifically, hybrid analog beamforming and digital precoding scheme can be employed to reduce the excessive radio frequency (RF) chains and data converters in the massive MIMO transceiver while still maintaining optimal spectral efficiency. However, traditional hybrid precoding architecture cannot be configured to support different system specifications, such as the number of bit streams or transmit antennas, leading to the limitation to hardware flexibility and efficiency. This article presents a configurable and low-complexity hybrid precoder based on the parallel data-stream processing in respect of system, algorithm, and architecture. The proposed algorithm was designed to avoid signal dependence between data streams so as to realize configurable precoding architecture. The performance and complexity of the proposed algorithm were also simulated and analyzed in detail in this article. Moreover, the hybrid precoding processor chip was designed and implemented based on the proposed algorithm. A novel data-processing flow was designed in the precoder so as to increase the hardware efficiency and configurability. The designed hybrid precoding chip can be configured to support one to four data streams for 16 × 16 mmWave MIMO systems. The designed precoder was implemented by using TSMC 40-nm CMOS technology. The normalized throughput achieved 11.1 M channel-matrices per second at a maximum clock frequency of 300 MHz. The area complexity is 263.5 kGE, and power consumption is 119.1 mW. Hao-Yu Cheng, Chen-Wei Chen, Chung-An Shen, Yuan-Hao Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | The Design and Implementation of a Highly Efficient Motion Estimation Engine for HEVC SystemsabstractThis paper presents the algorithm and VLSI architecture of a highly efficient Motion Estimation (ME) for High Efficiency Video Coding (HEVC) systems. To be specific, this paper proposes an Adaptive Fast Search Algorithm where the search space is adaptive to the characteristics of the video and the number of search candidates is greatly reduced. The experimental results show that, compared to the conventional approaches, this algorithm reduces the computational complexity by 54% with a marginal 2.01% performance degradation. Furthermore, the VLSI architecture and circuit implementation of the proposed ME engine is presented in this paper. The architectural and circuit-level optimizations for enhancing the throughput and reducing the complexity are illustrated. The proposed ME system is underwent the ASIC design flow and realized based on the Xilinx Zynq UltraScale+ FPGA platform. The results show that the, for ASIC and FPGA, the implemented ME achieves 60 frames per second (fps) and 30 fps with resolution of 3840×2160 respectively. The efficiency is also significantly enhanced. Yu-Hao Tseng, Chung-An Shen |
ISCAS | 2 |
| 2019 | Semantic Multi-Keyword Search over Encrypted Cloud Data with Privacy PreservationabstractCloud storage provides the great convenience for people to access their data at anytime from any place. Since cloud storage is usually run by the third-party service provider, keyword search over cloud data with privacy protection is of great importance. Many studies in the literature have proposed keyword search scheme for document search, but, in most schemes, the query keywords must exactly match those in the document indexes. However, it is impractical to restrict query keywords provided by the user when performing the search. This paper proposes the scheme for semantic multi-keyword search over encrypted cloud data. Users are able to select query keywords on their own choice. In addition, the query privacy of the user and the security of the documents are protected simultaneously through encrypted document search to prevent snooping from the cloud service provider. Experiments are conducted using a dataset of massive real world papers. The results show that the proposed scheme can effectively perform the semantic multi-keyword search over encrypted cloud data with great efficiency. Fei-Ju Hsieh, Tai-Lin Chin, Chin-Ya Huang, Shan-Hsiang Shen, Chung-An Shen |
VTC Fall | 5 |
| 2019 | An Efficient Joint Node and Link Mapping Approach Based on Genetic Algorithm for Network VirtualizationabstractNetwork virtualization is a promising technology for the emerging 5G and cloud computing networks where the virtual network is a logical topology consisting of virtual nodes and virtual links. In network virtualization, how to efficiently assign resources of the physical network to the virtual networks is of great significance and is known as the Virtual Network Embedding (VNE) problem. This paper presents an efficient algorithm tackling with the coordinated VNE problem. Specifically, a Mod-MaxMatch approach is presented which takes the global link resources into considerations when mapping the virtual nodes. Furthermore, a path splitting scheme based on the genetic algorithm is proposed while mapping the virtual links. The proposed algorithm minimizes the redundant reutilization of physical links and mitigates the demand for network bandwidths. A well-known link cost function is used to evaluate the network performance. The experimental results show that the link cost for the proposed approach is reduced by 77% compared to the traditional methodology and by 21% compared to the state-of-art design. Chia-Wei Huang, Chung-An Shen, Chin-Ya Huang, Tai-Lin Chin, Shan-Hsiang Shen |
VTC Fall | 2 |
| 2019 | User Centric Low Latency Data Transmission in Ultra Dense Vehicular NetworksabstractIn this paper, we propose a user centric bandwidth allocation scheme for low latency data transmission in ultra dense vehicular networks. Various mobile devices such as mobile phones, sensors of vehicles or autonomous driving systems,require low latency and bandwidth intensive packet delivery between the devices and the Internet aiming to support real-time applications. In the ultra dense vehicular network, small base stations (SBSs) are densely deployed in a fixed geographic area to provides higher date rate. In further, each SBS cooperates with others to form clusters to better support seamless wireless data transmissions, and each mobile device dynamically plans its wireless connectivity for data transmission when it moves in the network. Specifically, each mobile device pre-allocates the amount of bandwidth from a cluster, formed by several SBSs, based on its expected movement, the delay and band-width requirement of the packet transmission and the resource availability of each cluster. Moreover, to effectively utilize the available network resource, each cluster also redistributes its residual bandwidth to the mobile devices pre-allocate bandwidth from it. Consequently, the latency of the data transmission can be better sustained in the ultra dense vehicular network. Wei-Tsang Teng, Chin-Ya Huang, Shan-Hsiang Shen, Tai-Lin Chin, Chung-An Shen |
VTC Fall | 5 |
| 2018 | The hardware and software co-design of a configurable QoS for video streaming based on OpenFlow protocol and NetFPGA platform
Teng-Wei Chu, Chung-An Shen, Chun-Wei Wu |
Multim. Tools Appl. | 2 |
| 2018 | Advanced Multimedia Power-Saving Method Using a Dynamic Pixel Dimmer on AMOLED DisplaysabstractAs an emissive display, the active matrix organic light-emitting diodes (AMOLEDs) endure lower power efficiency at the higher level of pixel intensities. In the existing techniques, the power consumption is lowered with a significant loss of details and hue alterations. This paper proposes a hue-preserving pixel-dimming technique using subtractive coefficients based on the decomposed hue-saturation-value color map that reduces the power consumption on AMOLED displays. An entropy-based scene detection is adopted to maintain the computational efficiency of the proposed method on video input. Experimental results on a 5.5-in 1080p AMOLED displays show that the proposed method conserves up to 73% of the displaying power with high perceptual qualities as compared with the existing methods. Peter Chondro, Chia-Hua Chang, Shanq-Jang Ruan, Chung-An Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | The VLSI architecture of a highly efficient configurable pre-processor for MIMO detectionsabstractThis paper presents the VLSI architecture of a highly efficient configurable pre-processor supporting QR decomposition (QRD), sorted QRD (SQRD), and MMSE-SQRD for MIMO detections. The proposed design is architected based on the Givens Rotation algorithm and a high-throughput pipelined systolic array structure. Moreover, for achieving low-complexity, a novel norm-calculation scheme is utilized so that the overhead for the sorting operation is minimized. The circuit elements are also realized with the considerations of hardware sharing. The proposed pre-processor has been synthesized, placed, and routed using TSMC 90nm technology. The post-layout estimations show that this configurable pre-processor can support QRD, SQRD, and MMSE-SQRD, and can process 44M matrices per second for 4×4 MIMO systems with manageable complexity. Tzu-Ting Tseng, Chung-An Shen |
IPCCC | 2 |
| 2017 | The VLSI Architecture of a Highly Efficient Deblocking Filter for HEVC SystemsabstractThis paper presents the VLSI architecture and hardware implementation of a highly efficient deblocking filter (DBF) for High Efficiency Video Coding systems. In order to reduce the number of data accesses and thus to enhance the timing efficiency, novel data structures and memory access schemes for image pixels are proposed. Furthermore, a novel edge-fetching order is presented to strike a balance between the processing throughput and complexity. Based on the proposed structure and access pattern, a six-stage pipelined two-line DBF engine with low-latency data access sequence is designed, aiming to achieve high processing throughput while at the same time maintaining low complexity. The detailed storage structure and data access scheme are illustrated and VLSI architecture for the DBF engine is depicted in this paper. In addition, the proposed DBF is implemented using TSMC 90-nm standard cell library. The experimental results based on postlayout estimations show that the proposed design can achieve 60 frames/s for a frame resolution of 4096 × 2048 pixels (ultra high definition resolution) assuming an operating frequency of 100 MHz. Moreover, this design occupies an area complexity of 466.5 kGE with a power consumption of 26.26 mW. In comparison with prior designs targeting similar system specification and throughput, the proposed design results in a significantly reduced area complexity. Po-Kai Hsu, Chung-An Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Algorithm and Architecture of Configurable Joint Detection and Decoding for MIMO Wireless Communications With Convolutional CodesabstractThis paper presents an algorithm and a VLSI architecture of a configurable joint detection and decoding (CJDD) scheme for multi-input multioutput (MIMO) wireless communication systems with convolutional codes. A novel tree-enumeration strategy is proposed such that the MIMO detection and decoding of convolutional codes can be conducted in single stage using a tree-searching engine. Moreover, this design can be configured to support different combinations of quadrature amplitude modulation (QAM) schemes as well as encoder code rates, and thus can be more practically deployed to real-world MIMO wireless systems. A formal outline of the proposed algorithm will be given and simulation results for 16-QAM and 64-QAM with rate-1/2 and rate-1/3 codes will be presented showing that, compared with the conventional separate scheme, the CJDD algorithm can greatly improve bit error rate (BER) performance with different system settings. In addition, the VLSI architecture and implementation of the CJDD approach will be illustrated. The architectures and circuits are designed to support configurability and flexibility while maintaining high efficiency and low complexity. The postlayout experimental results for 16-QAM and 64-QAM with rate-1/2 and rate-1/3 codes show that, compared with the previous configurable design, this architecture can achieve reduced or comparable complexity with improved BER performance. Chung-An Shen, Chia-Po Yu, Chien-Hao Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | The joint detect and decoding approach for MIMO systems with turbo codesabstractThis paper presents the approach and algorithm of joint detection decoding with turbo codes for MIMO wireless communication systems. To the best of the authors' knowledge, this is the first proposed approach that can perform MIMO detection and decoding of turbo codes in a single stage. The detailed introduction of the proposed method and outline of the algorithm are given, while extensive simulations have been conducted to investigate the performance of the proposed algorithm. The simulation results show that, comparing to the conventional separate scheme, the proposed joint algorithm can achieve comparable BER performance with reduced hardware complexity and processing latency. Po-Hsiang Hsiung, Chung-An Shen, Huan-Chun Wang |
ISCAS | 2 |
| 2014 | Low power reduced-complexity error-resilient MIMO detectorabstractThis paper presents a reduced-complexity low power error-resilient K-Best MIMO Detector. A novel tree-enumeration method is proposed such that the error-resilient detection processes a reduced search space and is more suitable for VLSI design. Moreover, a circuit-level optimization is employed to further simplify the complexity. Experimental results are given showing that the circuit-level optimization decreases the detector area by 15% and power consumption by 41%. Moreover, we show that the proposed error-resilient MIMO detector with reduced-voltage memory can achieve a total of 19% reduction in power consumption compared with the conventional scheme, while still maintaining close-to optimal PER performance. Chung-An Shen, Muhammed S. Khairy, Ahmed M. Eltawil, Fadi J. Kurdahi |
ISCAS | 1 |
| 2012 | Error resilient MIMO detector for memory-dominated wireless communication systemsabstractIn current broadband MIMO-OFDM systems such as 3GPP LTE, embedded buffering memories occupy a large portion of chip area and a significant amount of power consumption. Due to the dense structure of memories, they are especially vulnerable to scaling effects such as process variation. These effects (hardware errors) become more pronounced when aggressive voltage scaling is used due to the reduced voltage overhead. To address this issue, we present an error resilient MIMO detector. First, we derive a combined distribution of the received data in a MIMO-OFDM receiver that includes both the noise incurred by the wireless channel and errors introduced at the receiver buffering memory due to aggressive voltage scaling. Using the derived distribution, a modified MIMO detection algorithm based on the tree-searching structure is presented. A case study is presented showing that the proposed approach can achieve near-optimal performance in the presence of both channel noise and memory error, while 40% to 50% of memory power savings are realized. Muhammed S. Khairy, Chung-An Shen, Ahmed M. Eltawil, Fadi J. Kurdahi |
GLOBECOM | 2 |
| 2012 | A Best-First Soft/Hard Decision Tree Searching MIMO Decoder for a 4 × 4 64-QAM SystemabstractThis paper presents the algorithm and VLSI architecture of a configurable tree-searching approach that combines the features of classical depth-first and breadth-first methods. Based on this approach, techniques to reduce complexity while providing both hard and soft outputs decoding are presented. Furthermore, a single programmable parameter allows the user to tradeoff throughput versus BER performance. The proposed multiple-input-multiple-output decoder supports a 4 × 4 64-QAM system and was synthesized with 65-nm CMOS technology at 333 MHz clock frequency. For the hard output scheme the design can achieve an average throughput of 257.8 Mbps at 24 dB signal-to-noise ratio (SNR) with area equivalent to 54.2 Kgates and a power consumption of 7.26 mW. For the soft output scheme it achieves an average throughput of 83.3 Mbps across the SNR range of interest with an area equivalent to 64 Kgates and a power consumption of 11.5 mW. Chung-An Shen, Ahmed M. Eltawil, Khaled N. Salama, Sudip Mondal |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | An Adaptive Reduced Complexity K-Best Decoding Algorithm with Early TerminationabstractThis paper presents a K-Best decoding algorithm that requires a much smaller K while preserving advantages of the sphere decoding algorithm such as branch pruning and an adaptively updated pruning threshold. The proposed approach results in examining a much smaller set of modulation points with a significantly reduced complexity. Simulations are presented that quantify the BER performance and complexity in terms of the number of visited nodes. The variability in the required operations (hence run-time) due to branch pruning is studied and compared with the sphere decoding algorithm. Chung-An Shen, Ahmed M. Eltawil |
CCNC | 1 |
| 2010 | A best-first tree-searching approach for ML decoding in MIMO systemabstractIn MIMO communication systems maximum-likelihood (ML) decoding can be formulated as a tree-searching problem. This paper presents a tree-searching approach that combines the features of classical depth-first and breadth-first approaches to achieve close to ML performance while minimizing the number of visited nodes. A detailed outline of the algorithm is given, including the required storage. The effects of storage size on BER performance and complexity in terms of search space are also studied. Our result demonstrates that with a proper choice of storage size the proposed method visits 40% fewer nodes than a sphere decoding algorithm at signal to noise ratio (SNR) = 20dB and by an order of magnitude at 0 dB SNR. Chung-An Shen, Ahmed M. Eltawil, Sudip Mondal, Khaled N. Salama |
ISCAS | 1 |
| 2010 | Design and Implementation of a Sort-Free K-Best Sphere DecoderabstractThis paper describes the design and very-large-scale integration (VLSI) architecture for a 4 × 4 breadth-first K-best multiple-input-multiple-output (MIMO) decoder using a 64 quadrature-amplitude modulation (QAM) scheme. A novel sort-free approach to path extension, as well as quantized metrics result in a high-throughput VLSI architecture with lower power and area consumption compared to state-of-the-art published systems. Functionality is confirmed via a field-programmable gate array (FPGA) implementation on a Xilinx Virtex II Pro FPGA. Comparison of simulation and measurements are given, and FPGA utilization figures are provided. Finally, VLSI architectural tradeoffs are explored for a synthesized application-specific IC (ASIC) implementation in a 65-nm CMOS technology. Sudip Mondal, Ahmed M. Eltawil, Chung-An Shen, Khaled N. Salama |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |