EDBT 2026 Demo / reviewers in the wild / expert
Chia-Hsiang Yang
dblp:98/6552
· DBLP profile ↗
18ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0003-1163-321XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 since 2021Computer networks · 6 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A High-Throughput Constructive Interference Precoder for 16 × MU-MIMO SystemsabstractIn a multiuser multiple-input multiple-output (MU-MIMO) downlink system, users are susceptible to interuser interference (IUI) because of data being simultaneously transmitted over the same time-frequency resources. Conventionally, precoding algorithms aim to eliminate the IUI. However, constructive interference (CI) precoding can achieve better error performance by exploiting the IUI. This article presents a high-throughput CI precoder. Design optimization across the algorithm and the architecture layers is conducted, reducing the complexity for multiplications by 81.6%. As the number of iterations for convergence varies, dynamic resource allocation is utilized to support each modulation mode with maximized utilization: time-multiplexing for the 4-QAM mode and parallel-processing for the 16-QAM mode. The proposed symbol updater also allows more efficient scheduling. As a proof of concept, a CI precoder chip that supports up to$16 \times $MU-MIMO systems is designed based in a 40-nm CMOS technology. The performance gains at a bit error rate (BER)$= 10^{-4}$are 10.7 and 12.5 dB for 4-QAM and 16-QAM, respectively, compared with conventional regularized zero-forcing (RZF) schemes. The precoder delivers a maximum throughput of 3.2 Gb/s at a clock frequency of 200 MHz for the$16 \times $MU-MIMO configuration. Ren-Hao Chiou, Chia-Hsiang Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Hybrid Precoding Baseband Processor for 64 × 64 Millimeter Wave MIMO SystemsabstractThis paper presents a hybrid precoding processor for millimeter wave (mmWave) multiple-input-multiple-output (MIMO) communication systems. The proposed architecture supports 64 antennas with 4-to-8 RF chains, and 4-bit phase resolution for each phase shifter in the analog beamformer. A polar decomposition (PD) engine is designed to achieve a 2.4-to-$8.6\times $lower latency with 1.5-to-$1.7\times $less normalized gate count compared to the singular value decomposition (SVD)-based implementation for the same functionality. Multiplicands are combined to reduce 55% area by utilizing the distributive property. Approximate phase extraction and low-precision multiplication for quantized phase extraction are adopted, which leads to 43% and 41% area reductions, respectively. The proposed hybrid precoding processor is designed in a 40-nm CMOS technology. It delivers a throughput of up to 10,288 kMat/s and consumes 170.9 mW at 278 MHz. It outperforms the state-of-the-art hybrid precoding processors with 5.5-$13.7\times $higher normalized area efficiency and 6.9-$38.4\times $lower normalized energy. Chen-Chien Kao, Chiao-En Chen, Chia-Hsiang Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | Achieving Accurate Automatic Sleep Apnea/Hypopnea Syndrome Assessment Using Nasal Pressure SignalabstractAutomatic assessment of sleep apnea/ hypopnea syndrome (SAHS) based on fewer physiological signals is critical for the success of healthcare at home. However, previous studies that use such settings only achieve a lower assessment accuracy, causing fewer syndromes to be separated for effective diagnosis. This paper presents a 3-stage support vector machines (SVM)-based algorithm for SAHS assessment using a single-channel nasal pressure (NP) signal. In this work, NP signal is utilized for feature extraction. Amplitude features, as well as those extracted using discrete Fourier transform and discrete wavelet transform, are used for machine learning. A total of 58 sets of polysomnography recordings, each with approximately 7 h in duration, were analyzed. This work achieves a sensitivity of 95.7% and a positive predictive value of 90.9%, outperforming previous works using NP signal. Compared with prior studies using only SpO2 signal, this work still achieves better performance and supports more classification levels. Thanks to the low-complexity settings based only on the NP signal, the proposed approach provides a promising solution to SAHS assessment for remote healthcare. Ying-Sheng Lin, Yi-Pao Wu, Yi-Chung Wu, Pei-Lin Lee, Chia-Hsiang Yang |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | A Color Doppler Processing Engine with an Adaptive Clutter Filter for Portable Ultrasound Imaging DevicesabstractThis paper presents an optimized color Doppler processing engine for portable ultrasound devices. The imaging performance is improved by the proposed adaptive clutter filter that separates the clutter from the blood signal in the frequency-eigenvalue space. The adaptive clutter filter provides high attenuation on clutter and also reserves the low-velocity blood flow. Compared to the conventional eigenvalue-based and frequency-based clutter filters, this work achieves 2.5-to-17.5dB and 1-to-12.5dB improvements, respectively, in blood-to-clutter ratio (BCR) based on an in vivo dataset with 16 images. The blood-tissue discriminator is re-arranged and integrated with the clutter filter to further enhance the imaging performance with less computational complexity. The computational complexity is reduced by 48-78% by leveraging the reduced processed area. Randomized spatial down-sampling is adopted to further reduce the processing time. An overall reduction of 60% in processing time is achieved. This work provides a promising solution to color Doppler imaging for portable ultrasound devices. Yi-Lin Lo, Chia-Hsiang Yang |
ICASSP | 2 |
| 2021 | A High-Throughput FPGA Accelerator for Short-Read Mapping of the Whole Human GenomeabstractThe mapping of DNA subsequences to a known reference genome, referred to as “short-read mapping”, is essential for next-generation sequencing. Hundreds of millions of short reads need to be aligned to a tremendously long reference sequence, making short-read mapping very time consuming. In this article, a high-throughput hardware accelerator is proposed so as to accelerate this task. A Bloom filter-based candidate mapping location (CML) generator and a folded processing element (PE) array are proposed to address CML selection and the Smith-Waterman (SW) alignment algorithm, respectively. It is shown that the proposed CML generator reduces the required memory access by 40 percent by employing a down-sampling scheme when compared to the Ferragina-Manzini index (FM-index) solution. The proposed hierarchical Bloom filter (HBF) that includes optimized parameters achieves a 1.5×104times acceleration over the conventional Bloom filter. The proposed memory re-allocation scheme further reduces the memory access time for the HBF by a factor of 256. The proposed folded PE array delivers a 1.2-to-3.2 times higher giga cell updates per second (GCUPS). The processing time can be further reduced by 53-to-72 percent by employing a fully pipelined PE array that allows for a tailored shift amount for seeding. The accelerator is realized on a Stratix V GX FPGA with 16GB external SDRAM. Operated at 200MHz, the proposed FPGA accelerator delivers a 2.1-to-11 times higher throughput with the highest 99 percent accuracy and 98 percent sensitivity compared to the state-of-the-art FPGA-based solutions. Yen-Lung Chen, Bo-Yi Chang, Chia-Hsiang Yang, Tzi-Dar Chiueh |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Iterative Receiver with a Lattice-Reduction-Aided MIMO Detector for IEEE 802.11axabstractThis paper presents the first 802.11ax compliant iterative detection and decoding (IDD) receiver that supports up to 4×4 1024-QAM MIMO detection in the open literature. Soft-input-soft-output (SISO) MIMO detection is implemented with a lattice reduction aided (LRA) K-best searcher and a max-log list demapper. A hardware-efficient IDD receiver is proposed to achieve the required packet-rate (PER) with a feasible latency. The extrinsic information transfer (EXIT) chart is utilized to reduce the number of iterations for IDD. Given the 802.11ax latency constraint, the performance, power, area (PPA) design space is explored to identify the optimal IDD receiver architecture. 50% of IDD inner iterations are reduced with only a 0.05dB loss in PER. The proposed IDD receiver achieves a 1dB improvement in PER with 3.6× smaller area and 3.0× lower power consumption when compared to the best non-IDD receiver. Yao-Pin Wang, Chi-Chih Wen, Chen-Chien Kao, Chung-Jung Huang, Der-Zheng Liu, Chia-Hsiang Yang |
GLOBECOM | 6 |
| 2020 | A CycleGAN Accelerator for Unsupervised Learning on Mobile DevicesabstractCycle-consistent generative adversarial networks (CycleGANs) have been commonly used for unsupervised-learning applications, especially for image-to-image translation. A CycleGAN has more complex dataflow since it features two generator-discriminator pairs. Massive external memory access also results in a long latency for both training and inference. Data structure for transposed convolution also needs to be tailored. This paper presents the first dedicated CycleGAN accelerator for energy-constrained mobile applications. The numbers of external and internal memory accesses are reduced by 98.3% and 68.3% through spatial data reuse, input feature map reuse, and local data reuse. The computational complexity is reduced by 79.4% by skipping zeros in the transposed convolutional layers. An architecture with two processing cores is proposed to improve the utilization by 2×. Designed in a 40-nm CMOS technology, the proposed CycleGAN accelerator dissipates 445 mW at 227 MHz from a 0.9-V supply. It achieves a 38× higher throughput-to-area ratio and 127× higher energy efficiency than a GPU. Yi-Yen Hsieh, Yu-Chi Lee, Chia-Hsiang Yang |
ISCAS | 3 |
| 2018 | A 2×2-16×16 Reconfigurable GGMD Processor for MIMO Communication SystemsabstractThe generalized geometric mean decomposition (GGMD) is a recently proposed matrix decomposition which can be viewed as a computationally efficient counterpart of the conventional geometric mean decomposition (GMD). As the GMD is the core algorithm in many high performance precoders and equalizers such as the Tomlinson-Harashima precoder and decision feedback equalizer, GGMD facilitates a more computationally efficient implementation while exhibiting identical performance. This work presents the first GGMD processor in the open literature, supporting various matrix sizes by leveraging the reconfigurable processing element (PE) using coordinate rotation digital computers (CORDICs). The implemented GGMD processor supports matrix sizes of 2nwhich ranges from 2×2 to 16×16, and the throughput performance is maximized through a PE array architecture. The chip integrates 326.9K gates in an area of 1.65 mm2in a 90-nm CMOS technology with the maximum throughput achieving 450K matrices/sec for a 16×16 matrix at 125 MHz. It dissipates 20.7-28.5 mW at 125 MHz from a 1V supply. Compared to previous GMD designs, this work supports a larger MIMO system with lower hardware complexity and power consumption. Chih-Hsuan Chiang, Shuo-An Huang, Chiao-En Chen, Chia-Hsiang Yang |
ISCAS | 4 |
| 2017 | Integration of energy-recycling logic and wireless power transfer for ultra-low-power implantablesabstractThis paper presents an integration of energy-recycling logic circuits with a wireless power transfer receiving module for ultra-low-power applications, such as transcutaneous biomedical implantables. In the prototype design, one inductive coil implanted inside the body receives wireless power and supplies the following electronics. While part of the loading is composed of conventional CMOS logics, the rest is implemented with energy-recycling logic circuits. Energy-recycling logic and the associated adiabatic operation achieve excellent energy efficiency by transferring and recycling energy between digital logic blocks along with the signal propagation. The required AC supplies further lead to a natural integration with wireless power transfer and therefore obviate the need for a rectifier that contributes to substantial power loss. As a proof of concept, a finite-impulse-response filter is designed in 90-nm CMOS process. Simulation results show a 59.3% power reduction as compared to static CMOS counterpart. Hsin-Tzu Lin, Yi-Chung Wu, Ping-Hsuan Hsieh, Chia-Hsiang Yang |
ISCAS | 4 |
| 2017 | An Area-Efficient Multi-Mode LLR Computing Engine for MMSE-Based MIMO DetectorsabstractIt is known that by extracting log-likelihood ratio (LLR) values from the received MIMO signals, a MIMO detector is able to exchange a posteriori soft information with a channel decoder to improve the error performance. A minimum mean-square error with parallel interference cancellation (MMSE-PIC) detector is considered to be a practical MIMO detector, but the computational complexity for the LLR computation is still at a high level for higher-order modulations. This paper presents a low-complexity LLR computation algorithm together with the hardware implementation for MMSE-PIC detectors. The complexity is evaluated using logic synthesis based on a 90-nm CMOS technology. For 256-QAM modulation, a 50% reduction in area is achieved. In addition, a 44% reduction in area can be achieved for a multi-mode MMSE-PIC detector that supports multiple modulations for signals up to 256-QAM. Wei-Cheng Sun, Chia-Hsiang Yang, Yen-Ming Chen, Yeong-Luh Ueng |
VTC Spring | 2 |
| 2016 | Error-resilient sequential cells with successive time borrowing for stochastic computingabstractThis paper presents error-resilient sequential building blocks with time-borrowing capability without extra latches and generated clocks. The circuits are able to recover the timing errors caused by PVT variations and/or over-voltage scaling by up to half a cycle. Unlike prior works, the timing errors can be recovered dynamically through successive time borrowing without stalled cycles, retaining a constant throughput. The circuit structure can be applied to both ASICs and microprocessors. The proposed sequential cells are highly compatible with current cell-based IC design flow, for both feedforward and feedback datapaths. As a proof of concept, a design with key DSP building blocks has been verified. The results show that the performance of the DSP modules is improved by 13-15% in the worst-case operation condition, yielding a promising solution for stochastic computing under an unreliable operation condition. Wei-Chang Liu, Ching-Da Chan, Shuo-An Huang, Chi-Wei Lo, Chia-Hsiang Yang, Shyh-Jye Jou |
ICASSP | 5 |
| 2016 | sBWT: memory efficient implementation of the hardware-acceleration-friendly Schindler transform for the fast biological sequence mappingabstractMOTIVATION: The Full-text index in Minute space (FM-index) derived from the Burrows-Wheeler transform (BWT) is broadly used for fast string matching in large genomes or a huge set of sequencing reads. Several graphic processing unit (GPU) accelerated aligners based on the FM-index have been proposed recently; however, the construction of the index is still handled by central processing unit (CPU), only parallelized in data level (e.g. by performing blockwise suffix sorting in GPU), or not scalable for large genomes. RESULTS: To fulfill the need for a more practical, hardware-parallelizable indexing and matching approach, we herein propose sBWT based on a BWT variant (i.e. Schindler transform) that can be built with highly simplified hardware-acceleration-friendly algorithms and still suffices accurate and fast string matching in repetitive references. In our tests, the implementation achieves significant speedups in indexing and searching compared with other BWT-based tools and can be applied to a variety of domains. AVAILABILITY AND IMPLEMENTATION: sBWT is implemented in C ++ with CPU-only and GPU-accelerated versions. sBWT is open-source software and is available at http://jhhung.github.io/sBWT/Supplementary information: Supplementary data are available at Bioinformatics online. CONTACT: [email protected] or [email protected] (also [email protected]). Chia-Hua Chang, Min-Te Chou, Yi-Chung Wu, Ting-Wei Hong, Yun-Lung Li, Chia-Hsiang Yang, Jui-Hung Hung |
Bioinform. | 6 |
| 2015 | An Iterative Geometric Mean Decomposition Algorithm for MIMO Communications SystemsabstractThis paper presents an iterative geometric mean decomposition (IGMD) algorithm for multiple-input-multiple-output (MIMO) wireless communications. In contrast to the conventional geometric mean decomposition (GMD) algorithm, the proposed IGMD does not require the explicit Kth root computation in the preprocessing stage but depends on a carefully constructed iterative procedure that generates the GMD in its limit. We prove analytically that the proposed IGMD is guaranteed to converge to the exact GMD under certain sufficient conditions, and propose three different constructions achieving this condition. Both numerical simulations and complexity analysis of the proposed IGMD have been conducted and compared with the conventional GMD. Simulation results show that our new IGMD algorithm effectively reduces the complexity overhead and hence is more advantageous for low-complexity implementations. Chiao-En Chen, Yu-Cheng Tsai, Chia-Hsiang Yang |
IEEE Trans. Wirel. Commun. | 3 |
| 2013 | A 191μW BPSK demodulator for data and power telemetry in biomedical implantsabstractThis paper presents a fully digital binary-phase-shift keying (BPSK) demodulator for data and power telemetry. This demodulator recovers BPSK signals by detecting the symbol edge of the digitized received carrier. Parameters of the coupling coils, rectifier DC output, and data rate are taken into consideration in the early design stage. Given a limited coil size and quality factor, the demodulator achieves a data rate of 678kb/s with BER < 10-9 at a carrier frequency of 13.56MHz. Fabricated in a 0.18μm CMOS process, the chip area is 0.445mm square. The chip core dissipates 191μW at 13.56MHz. A system prototype was developed to transmit data and power simultaneously through a pair of coils. Li-Lan Wang, Chia-Hsiang Yang, Herming Chiueh |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | An Energy-Efficient VLSI Architecture for Cognitive Radio Wideband Spectrum SensingabstractSpectrum sensing over a wide bandwidth increases the probability of finding unutilized spectrum for cognitive radios. However, energy-efficient VLSI realization of wideband sensing algorithms is challenging due to complex signal processing and real-time requirement. In addition, strong primary users introduce spectral leakage in adjacent unused bands, resulting in sensing performance degradation. To address these challenges, we propose a cascaded filter-bank channelization scheme and its reconfigurable VLSI architecture that can be optimized for power, area, sensing time, and detection performance. In addition, the spectral leakage due to strong interference is compensated with an interference cancellation method. Compared to the conventional PSD-based energy detection, the proposed channelization scheme supports reliable wideband signal detection with 2.5× less power. Given a 0.5ms sensing time under 30-dB adjacent-band interference-to-noise ratio, a 30× sensing-time improvement is achieved while maintaining a false- alarm probability of 0.1 and a detection probability of 0.9. The energy efficiency is improved from lower power consumption and reduced sensing time. Tsung-Han Yu, Chia-Hsiang Yang, Dejan Markovic, Danijela Cabric |
GLOBECOM | 2 |
| 2011 | A Hardware-Efficient VLSI Architecture for Hybrid Sphere-MCMC DetectionabstractThis paper presents a hybrid soft-output MIMO detector that searches reliable soft-information in both deterministic and probabilistic ways. The fixed-complexity sphere detector (FSD) is first applied to provide near maximum-likelihood (ML) solutions. The solutions are next used to initialize the Markov Chain Monte Carlo (MCMC) detector that uses parallel Gibbs samplers (GSs) for remaining candidate enumeration. A low-complexity VLSI architecture is proposed to demonstrate the feasibility of hardware realization for high-throughput applications. Simulation results indicate that the hybrid detector has a 2.3x complexity reduction and a 2× throughput improvement compared to individual soft-output FSD and MCMC detectors. Fang-Li Yuan, Chia-Hsiang Yang, Dejan Markovic |
GLOBECOM | 2 |
| 2008 | A Multi-Core Sphere Decoder VLSI Architecture for MIMO CommunicationsabstractThe sphere decoding algorithm finds applications in multi-input multi-output (MIMO) decoding, because it achieves near maximum likelihood (ML) detection performance with significantly reduced computational complexity. Previous work has focused on implementations based on K-best or depth-first search, limiting the BER performance or the search speed. This paper presents a scalable multi-core sphere decoder architecture that can combine the advantages of the K-best and depth-first search methods. The proposed architecture demonstrated a 3-5 dB improvement in the BER performance for 16times16 systems using 16 processing elements (PEs) compared to the architecture with one PE. An improved search speed of the multi-core architecture also enables a 10times energy efficiency improvement over the single core architecture for the same data rate. Chia-Hsiang Yang, Dejan Markovic |
GLOBECOM | 1 |
| 2008 | A Flexible VLSI Architecture for Extracting Diversity and Spatial Multiplexing Gains in MIMO ChannelsabstractThe sphere decoding algorithm is able to approach maximum likelihood (ML) detection with significantly reduced computational complexity for multi-input multi-output (MIMO) communications. The computational reduction makes it attractive for hardware implementation. This paper presents a unified sphere decoder architecture that deploys diversity-multiplexing tradeoff in MIMO channels by taking advantage of the flexibility in the number of antennas and modulation schemes. Several signal processing and circuit techniques are constructively combined to reduce the hardware complexity: a 20 times area reduction is achieved even without interleaving of sub-carriers compared to the direct-mapped architecture. The proposed flexible architecture supports antenna arrays from 2x2 to 16x16, modulations from BPSK to 64-QAM, over 16 to 128 sub-carriers. The peak estimated data rate exceeds 1.5 Gbps over a 16 MHz bandwidth in just 0.55 mm2in a standard 90 nm CMOS process. Chia-Hsiang Yang, Dejan Markovic |
ICC | 1 |