Shyh-Jye Jou

dblp:15/534 · DBLP profile ↗
← Back
70ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-8821-3486ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 64 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Neural Network-Based Contiguous Carrier Aggregation Digital Predistortion Design for Sub-THz Power Amplifier in Baseband Transmitter
Cheng-Hsuan Lai, Chung-Lun Tu, Yi-Shan Huang, Shyh-Jye Jou
ISCAS4
2026 A 66.15 TFLOPS/W RRAM-based Near-memory-computing Macro using Log-multiplication and Error-aware Adaptive Adder Structure with 2-D Pipelined CNN Dataflow
Yue-Ci Lee, Yuan-Ping Huang, Shyh-Jye Jou
ISCAS3
2026 High Throughput LDPC Decoder with Ultra-low BER Using Hardware Sharing across Two Code Rates for IEEE Std. 802.15.3d
Ching Liang Yeh, Yi-Shan Huang, Chung-Lun Tu, Shyh-Jye Jou
ISCAS4
2026 A PVT-Resilient Subthreshold SRAM-Based In-Memory Computing Accelerator With In-Situ Regulation for Energy-Efficient Spiking Neural Networks
abstract
This paper presents a PVT-resilient, subthreshold SRAM-based computing-in-memory (CIM) macro tailored for energy-efficient spiking neural networks (SNNs). The macro integrates in-situ current sensors and distributed voltage regulators to enable robust large-scale (1024 wordlines, 1304 bitlines and 128 shared neuron cells) subthreshold current-mode CIM, mitigating energy overheads and process-voltage-temperature (PVT) sensitivity. The neuron cells adopt a programmable, memory cell-based firing threshold to enhance neuron robustness against PVT variations. The architecture uses a stride-tick batching schedule to significantly reduce buffer overhead with enhanced input data reuse. Exploiting the high sparsity of SNNs, the proposed system demonstrates significant improvements in energy efficiency and variation tolerance. Fabricated in 28-nm CMOS, the prototype attains 93.64\% accuracy on keyword spotting, delivers up to 1181.42 TOPS/W, and achieves 7.24 TOPS/mm^2, demonstrating a viable and efficient solution for high-performance edge SNN processing.
Shih-Hang Kao, Yang-Chan Hung, I-Wen Wang, Bing-Han Liu, Yu-Chia Chen, Tian-Sheuan Chang, Shyh-Jye Jou, Chien-Nan Jimmy Liu, Hung-Ming Chen, Wei-Zen Chen
IEEE Trans. Circuits Syst. I Regul. Pap.7
2025 A Deep Learning Accelerator for Modified YOLOv7-tiny with Shortcut-Aware Layer Fusion
abstract
Convolutional Neural Networks (CNNs) have become dominant in object detection tasks. However, off-chip data traffic, primarily from intermediate feature maps, remains a major performance bottleneck in CNN accelerators. This paper proposes a deep learning accelerator for YOLOv7-tiny with the following features: 1) A Shortcut-Aware Layer Fusion method that reduces off-chip feature map traffic by 32%. 2) A Multi-Bank Memory Control Scheme that efficiently manages all feature map data in one unified on-chip buffer, eliminating unnecessary data copying and improving the utilization of buffer resources. The design is implemented in TSMC 28 nm HPC+ process and occupies 2.19 mm2. The proposed accelerator achieves an energy efficiency of 3.0 TOPS/W and processes 640×640 input images at 34.5 FPS.
Wei-En Huang, Yi-Shan Huang, Shyh-Jye Jou
ISCAS3
2024 A Multi-Bit Near-RRAM based Computing Macro with Highly Computing Parallelism for CNN Application
abstract
Resistive random-access memory (RRAM) based compute-in-memory (CIM) is an emerging approach to address the demand for practical implementation of artificial intelligence (AI) on resource constrained edge devices by reducing the power-hungry data transfer between memory and processing unit. However, the state-of-the-art RRAM CIM designs fail to strike a balance between precision, energy efficiency, throughput, and latency. This work merges the techniques of CIM and compute-near-memory (CNM) to deliver high precision, high energy efficiency, high throughput, and low latency. In this paper, a 256Kb RRAM based CNM macro fabricated in TSMC 40 nm process is presented featuring: 1) opposite weight mapping with variation-robust SA to mitigate the impact of RRAM device variations on MAC (Multiply-Accumulate) results; 2) switched-capacitor-based analog multiplication circuit to achieve highly parallel computing of 128 4-bit by 4-bit MAC result with low power consumption and high operation speed; and 3) joint optimization of hardware and software to compensate for the accuracy loss after considering the non-idealities of circuits. The macro achieves a low latency of 17ns and high energy efficiency of 71 TOPS/W for MAC operations with 4-bit input, 4-bit weight and 4-bit output precision. It is used to accelerate the convolution process in the Light-CSPDenseN et AI model, resulting in a high accuracy of 86.33% on Visual Wake Words dataset.
Kuan-Chih Lin, Hao Zuo, Hsiang-Yu Wang, Yuan-Ping Huang, Ci-Hao Wu, Yan-Cheng Guo, Shyh-Jye Jou, Tuo-Hung Hou, Tian-Sheuan Chang
DATE7
2024 Online Self-Adaptive Estimation and Compensation Design for DC Voltage Offset, Frequency-Independent, and Frequency-Dependent IQ Mismatch in Sub-THz Digital Baseband Transceiver
abstract
This paper presents a baseband transceiver architecture working at Sub-THz based on IEEE 802.15.3d. We propose an online self-adaptive compensator (OSAC) to deal with DC voltage offset (DCVO), frequency-independent IQ (FIIQ) mismatch, and frequency-dependent IQ (FDIQ) mismatch. For DCVO compensation, the residual DCVO in I/Q branch is improved from 75 mV / 75 mV to 0.0027 mV / 0.0024 mV. For FIIQ mismatch, the image rejection ratio (IMRR) is improved from 12.80 dB to 55.09 dB. For FDIQ mismatch, the IMRR can be improved from 26.02 dB to 73.38 dB. The hardware implementation of OSAC is operated at 3.52 GHz with the TSMC 16-nm FinFET CMOS process to achieve the data rate of 14.02 Gbps. The core area and power consumption of OSAC are 12769 μm2and 84.7 mW, respectively.
Chung-Lun Tu, Shyh-Jye Jou
ISCAS3
2024 Channel Estimation and Equalization Design with SNR Decision Based Universal Threshold for Sub-THz Single Carrier Baseband Receiver
abstract
In this paper, a Golay-correlator SNR decision based universal threshold (GC-SNR-UT) method for channel estimation with frequency domain equalizer (FDE) in sub-THz band is proposed. The SNR decision based universal threshold denoising method combines the SNR decision mechanism with the threshold for mean square error optimization (TMSE) and the universal threshold formula. The simulation environment incorporates channel effect based on IEEE 802.15.3d standard and additive white Gaussian noise (AWGN). By implementing the proposed SNR decision based universal threshold (SNR-UT) denoising method, the computational complexity can be reduced by about 50% compared to the two-stage universal threshold, and non-linear operations are not required. Furthermore, it improves the required SNR by about 0.9 dB at the uncoded BER requirement of 1.9 ×10−4compared to results without any noise mitigation. Notably, it provides about a 2.2 dB margin to handle other non-ideal effects. In the hardware implementation, we use the TSMC 16 nm FinFET process to achieve a data transmission with a bandwidth of 3.52 GHz. The core area and power are 0.142 mm2and 654.2 mW, respectively for the proposed design.
Feng-Ju Liao, Chung-Lun Tu, Shyh-Jye Jou
ISCAS3
2024 A 128 Gb/s LDPC Decoder Using Partial Syndrome-based Dynamic Decoding Scheme for Terahertz Wireless Multi-Media Networks
abstract
This paper presents a low power and high throughput multi-framed pipelined LDPC decoder architecture based on a novel partial syndrome-based dynamic decoding (PSDD) approach. The proposed PSDD can reduce clock cycle to allow the LDPC decoder to be implemented with better energy efficiency. We propose a high throughput sorting method and implement the LDPC decoder with a pipelined multi-frame VLSI architecture. The implementation results for the IEEE 802.15.3d Thz standard shows that the proposed design has a coding gain of 10−8at the specified SNR of 18.1 dB with 16 QAM modulation. Furthermore, the proposed design can achieve a throughput rate of 128.5 Gbps with the 16nm FinFET CMOS process, respectively.
Tsung-Han Wu, Ching Liang Yeh, Yi-Shan Huang, Shyh-Jye Jou
ISCAS4
2023 On Automating Finger-Cap Array Synthesis with Optimal Parasitic Matching for Custom SAR ADC
abstract
Due to its excellent power efficiency, the successive-approximation-register (SAR) analog-to-digital converter (ADC) is an attractive design choice for low-power ADC implements. In analog layout design, the parasitics induced by interconnecting wires and elements affect the accuracy and performance of the device. Due to the requirement of low-power and high-speed, series of very small lateral metal-metal capacitor units are usually adopted as the architecture of capacitor array. Besides power consumption and area reduction, the parasitic capacitance would significantly affect the matching properties and settling time of capacitors. This work presents a framework to synthesize good-quality binary-weighted capacitors for custom SAR ADC. Also, this work proposes a parasitic-aware ILP-based weight-dynamic network routing algorithm to generate a layout considering parasitic capacitance and capacitance ratio mismatch simultaneously. The experimental result shows that the effective number of bits (ENOB) of the layout generated by our approach is comparable to or better than that of manual design and other automated works, closing the gap between pre-sim and post-sim results.
Cheng-Yu Chiang, Chia-Lin Hu, Mark Po-Hung Lin, Yu-Szu Chung, Shyh-Jye Jou, Jieh-Tsorng Wu, Shiuh-Hua Wood Chiang, Chien-Nan Jimmy Liu, Hung-Ming Chen
ASP-DAC5
2023 Offline and Time-variant EVD-based Closed-loop Digital Predistortion Design for Sub-THz Power Amplifier Array in Basedband Transmitter
abstract
In this paper, we propose an eigenvalue decomposition (EVD) based digital predistortion (DPD) design for nonlinear memory PAs in active antenna array structure with offline and time-variant scenario. The simulation environment incorporates various non-ideal effects based on IEEE Std 802.15.3d, the transmitter error vector magnitude (TX EVM) performance can be improved from -9.07dB to -20.64 dB in the presence of the proposed DPD. Furthermore, since the characteristics of the sub-THz power amplifiers are sensitive to temperature changes, we simulate the impact when the temperature rises from 27 °C to 125°C. This leads to inaccurate offline estimation, while with the proposed online tracking mechanism, the TX EVM can be improved from -13.34 dB to -20.05 dB. For the hardware implementation, we use 16-nm FinFET CMOS process with eight times parallelism architecture to achieve a data transmission with a bandwidth of 7.04GHz. The core area and power are 0.106 mm2 and 464.8 mW for the TX DPD, and 0.0576 mm2 and 252.0 mW for the receiver optimizer, respectively.
Chung-Lun Tu, Chin-Ming Chang, Shyh-Jye Jou
ISCAS3
2023 Low Routing Complexity Multiframe Pipelined LDPC Decoder Based on a Novel Pseudo Marginalized Min-Sum Algorithm for High Throughput Applications
abstract
This article presents a high throughput and low routing complexity multiframe pipelined low-density parity check (LDPC) decoder design based on a novel pseudo marginalized min-sum (PMMS) message passing approach. The proposed PMMS approach reduces the required number of interconnections in the routing network allowing the design to be implemented with reduced hardware complexity, low power consumption and high throughput capability while supporting multiple coding rates with short and long codewords as defined in many application standards. Implementation results for IEEE802.11ad/ay standards show that the proposed design satisfies a target bit error rate (BER) requirement of$3 \times 10^{-7}$with 64 quadrature amp mod (QAM) targeting high throughput applications. Furthermore, the proposed design is able to achieve a throughput of 62 and 101.8 Gb/s with two pipelining stages under 28-nm CMOS and 16-nm FinFET CMOS process, respectively. As compared to the existing NMS algorithm, the proposed design based on the PMMS approach reduces the number of wires in the routing network by 45.5%, and the wirelength of the overall decoder by 17%. The area and power consumption are also reduced by 9.4% and 12%, respectively, as compared to the conventional normalized min-sum (NMS) algorithm.
Henry Lopez Davila, Tsung-Han Wu, Shyh-Jye Jou, Sau-Gee Chen, Pei-Yun Tsai 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2022 DASC: A DRAM Data Mapping Methodology for Sparse Convolutional Neural Networks
abstract
The data transferring of sheer model size of CNN (Convolution Neural Network) has become one of the main performance challenges in modern intelligent systems. Although pruning can trim down substantial amount of non-effective neurons, the excessive DRAM accesses of the non-zero data in a sparse network still dominate the overall system performance. Proper data mapping can enable efficient DRAM accesses for a CNN. However, previous DRAM mapping methods focus on dense CNN and become less effective when handling the compressed format and irregular accesses of sparse CNN. The extensive design space search for mapping parameters also results in a time-consuming process. This paper proposes DASC: a DRAM data mapping methodology for sparse CNNs. DASC is designed to handle the data access patterns and block schedule of sparse CNN to attain good spatial locality and efficient DRAM accesses. The bank-group feature in modern DDR is further exploited to enhance processing parallelism. DASC also introduces an analytical model to facilitate fast exploration and quick convergence of parameter search in minutes instead of days from previous work. When compared with the state-of-the-art, DASC decreases the total DRAM latencies and attains an average of 17.1x, 14.3x, and 23.3x better DRAM performance for sparse AlexNet, VGG-16, and ResNet-50 respectively.
Bo-Cheng Lai, Tzu-Chieh Chiang, Po-Shen Kuo, Wan-Ching Wang, Yan-Lin Hung, Hung-Ming Chen, Chien-Nan Jimmy Liu, Shyh-Jye Jou
DATE8
2022 Design of a mmWave Digital Baseband Receiver Integrated with WOLA-CP-OFDM Technique
abstract
In this paper, we propose a triple-mode baseband (BB) receiver (RX) architecture for cyclic prefix orthogonal frequency division multiplexing (CP-OFDM) with weighted overlap and add (WOLA) technique. The system specifications follow the IEEE 802.11ad/ay standard for next-generation millimeter wave communications. This BBRX is a flexible and expandable architecture due to the modularized, parameterized, and feed-forward data flow design strategies. The non-ideal circuit effects, such as frequency offset, phase noise, power amplifier (PA) non-linearity, and adjacent channel interference are jointly considered for the overall performance evaluation. The simulation results show that CP-OFDM with WOLA (W-CP-OFDM) and filter bank multicarrier (FBMC) present similar out-of-band performance while considering PA non-linearity. The synthesis results indicate that the W-CP-OFDM/FBMC has 2.1%/97% gate count increased overhead as compared to CP-OFDM.
Kang-Lun Chiu, Hsun-Wei Chan, Hsuan-Ping Chiu, Chun-Yi Liu 0001, Chih-Wei Jen, Shyh-Jye Jou
ISCAS6
2022 Low-Complexity Pseudo Direct Learning Digital Pre-Distortion Architecture for Nonlinearity and Memory Effect of Power Amplifier in mmWave Baseband Transmitter
abstract
In this paper, we design a power amplifier (PA) digital pre-distortion (DPD) module at the baseband transmitter to pre-compensate the nonlinearity and memory effect of the PA. In terms of DPD, we propose a low-complexity pseudo direct learning (PDL) DPD architecture based on the system level point of view according to the IEEE 802.11ad/ay specifications. The compensated error vector magnitude (EVM) performance of nonlinearity and memory effects can be improved from -12.8 dB to -21.7 dB at 16-QAM mode. The gain flatness of the -3 dB bandwidth can be extended from 0.3$\pi$ to 0.78$\pi$ with improving of 2.6 times. For the hardware implementation, we use TSMC 28-nm HPC_PLUS CMOS technology with four times parallelism to achieve 2.5 GHz chip rate. The gate counts and power of the proposed DPD design are 273.1 K and 98.0 mW, respectively.
Shen-Zhe Lu, Nai-Cheng Xue, Hung-Chih Liu, Chih-Wei Jen, Shyh-Jye Jou
ISCAS5
2022 Compressive Sensing Based Hardware Design for Channel Estimation of Wideband Millimeter Wave Hybrid MIMO System
abstract
Channel estimation is a crucial issue for hybrid multiple-input multiple-output architecture of wideband wireless millimeter wave system. In this paper, we present a hardware implementation of channel estimation based on compressive sensing method. We exploit sparsity in channel space and take advantage of both the time domain and the frequency domain to effectively reduce the computational complexity. We then disassemble the sensing matrix into smaller dimensional matrices to further simply the sensing formula by utilizing the orthogonality of the DFT codebook. The sensing issue is solved by the generalized orthogonal matching pursuit with Cholesky decomposition techniques, which achieves a significant reduction in the number of computations. Finally, we evaluate the performance of the proposed method with the perfect channel state information, and hardware performance by fixed-point analysis and RTL design and synthesis results.
Chung-Lun Tu, Tse-Yuan Lin, Kang-Lun Chiu, Shyh-Jye Jou, Pei-Yun Tsai 0001
ISCAS4
2022 Joint Digital Online Compensation of TX and RX Time-Varying I/Q Mismatch and DC-Offset in mmWave Transceiver System
abstract
Millimeter-wave (mmWave) RF and analog front-end circuits are very susceptible to chip process and temperature variation, which cause I/Q mismatch and DC offset. As a consequence, the system performance can be seriously degraded, especially in wideband multi-Gb/s systems. This paper proposes an online estimation and compensation in the baseband receiver end of a single carrier (SC) system for TX and RX frequency independent (FI) initial and time-varying (TV) I/Q mismatch and DC-offset. Based on specification of IEEE 802.11ay, the compensated image rejection ratio (IRR) performance of the TX and RX FI I/Q mismatch effects can be improved from 16.43 dB to 54.70 dB at SNR of 25.52 dB. Moreover, we propose a novel algorithm based on the Golay sequence for estimation and compensation to cancel TX and RX initial/TV DC-offset. The DC-offset in the I/Q channel is improved from −56.46 dB/−58.56 dB to −97.30 dB/−98.70 dB at SNR of 25.52 dB. In the hardware implementation, a four-time parallelism architecture are proposed to work at a 625 MHz clock rate with a 28-nm HPC_PLUS CMOS process under 64-QAM mode for 15 Gbps transmission. Using a clock gating control scheme for FI I/Q mismatch and DC-offset estimator, the total power of the proposed the module can be reduced from 59.5 mW to 32.7 mW. The gate count and power of the proposed estimation/compensation design are only 2.94% and 1.69% of the overall digital RX baseband gate count and power consumption.
Hung-Chih Liu, Zheng-Chun Huang, Ngoc-Giang Doan, Chih-Wei Jen, Shyh-Jye Jou
IEEE Trans. Circuits Syst. I Regul. Pap.5
2022 A 14 μJ/Decision Keyword-Spotting Accelerator With In-SRAMComputing and On-Chip Learning for Customization
abstract
Keyword spotting (KWS) has gained popularity as a natural way to interact with consumer devices in recent years. However, because of its always on nature and the variety of speech, it necessitates a low-power design as well as user customization. This article describes a low-power, energy-efficient KWS accelerator with static random access memory (SRAM)-based in-memory computing (IMC) and on-chip learning for user customization. However, IMC is constrained by macro size, limited precision, and nonideal effects. To address the issues mentioned above, this article proposes bias compensation and fine-tuning using an IMC-aware model design. Furthermore, because learning with low-precision edge devices results in zero error and gradient values due to quantization, this article proposes error scaling and small gradient accumulation to achieve the same accuracy as ideal model training. The simulation results show that with user customization, we can recover the accuracy loss from 51.08% to 89.76% with compensation and fine-tuning and further improve to 96.71% with customization. The chip implementation can successfully run the model with only 14$\mu \text{J}$per decision. When compared to the state-of-the-art works, the presented design has higher energy efficiency with additional on-chip model customization capabilities for higher accuracy.
Yu-Hsiang Chiang, Tian-Sheuan Chang, Shyh-Jye Jou
IEEE Trans. Very Large Scale Integr. Syst.3
2021 A Low-Jitter ADPLL with Adaptive High-Order Loop Filter and Fine Grain Varactor Based DCO
abstract
An all-digital phase-locked loop (ADPLL) with adaptive higher-order filter is proposed in this paper. The proposed ADPLL can select the first to third order of the loop filter by turning the IIR filter to adjust the system performance and attenuate input noise. Moreover, the phenomenon that spurious tone is getting closer to the main tone at higher-order ADPLL will be analyzed in this paper. The chip has been designed and implemented in TSMC 40 nm GP 1P10M CMOS process technology. The total area of the ADPLL core is 0.0106 mm2. By turning on the IIR filter, the measured rms jitter is 0.6 ps (0.298 % UI) and the power consumption is 5.1 mW from a 0.9 V supply at 4.96 GHz output frequency with 40 MHz reference clock.
Chia-Chen Chang, Yu-Tung Chin, Hossameldin A. Ibrahim, Kang-Yu Chang, Shyh-Jye Jou
ISCAS5
2021 A Digital Two-Stage Phase Noise Compensation and rCFO/rSCO Tracking Module for mmW Single Carrier Systems
abstract
This article proposes a digital two-stage phase noise (PN) compensation and residual carrier frequency offset (rCFO) and residual sampling clock offset (rSCO) tracking module in the baseband receiver side for single carrier (SC) systems in the millimeter-wave (mmW) band. The proposed two-stage PN compensation and rCFO/rSCO tracking module (2STG-PNC&T) includes a correlation-based linear interpolation method for the common phase error (CPE) of the low-frequency PN, a subsymbol CPE compensation with a reformed low-complexity π/4robust PN slicer for the high-frequency PN, and a tracking mechanism for rCFO and rSCO estimation on 64-QAM under the harsh PN condition of the frequency synthesizers used in the mmW band. The proposed module can restrict the maximum rCFO/rSCO within ±0.4 ppm and track the phase error caused by the PN successfully. Besides, the compensated power spectral density of the residual PN is reduced from -90 to -125 dBc/Hz at 1-MHz offset. A four-time parallelism architecture of the new 2STG-PNC&T is proposed to work at a 625-MHz clock rate under 64-QAM for 15 Gb/s transmission. The gate count/power of the proposed design is only 4.0%/6.9% of the overall digital baseband, respectively.
Hsun-Wei Chan, Wei-Che Lee, Kang-Lun Chiu, Chih-Wei Jen, Shyh-Jye Jou
IEEE Trans. Very Large Scale Integr. Syst.5
2020 On EDA Solutions for Reconfigurable Memory-Centric AI Edge Applications
abstract
Memory-centric designs deploy computation to storage and enable efficient in-memory computation while avoiding massive amount of data movement. The in-memory-computing schemes have shown distinct advantages and concerns when applying to different types of memory technologies, from conventional SRAM, DRAM to emerging ReRAM. Moreover, the next-generation smart edge systems are expected to support various intelligent applications by employing multi-task machine learning models which would be dynamically activated. To attain an efficient design within short design cycle, it is imperative to have an integrated design framework with automated tools to support hybrid memory systems and perform effective optimization across design stages. This work will introduce a unified framework which integrates EDA solutions to address the design and optimization challenges at different aspects of next-generation memory-centric designs, including fast reconfiguring in-memory/near-memory computing designs to provide optimized solutions (behavioral models and APR cell layouts) for designers to choose the best suitable architectures for their applications.
Hung-Ming Chen, Chia-Lin Hu, Kang-Yu Chang, Alexandra Küster, Yu-Hsien Lin, Po-Shen Kuo, Wei-Tung Chao, Bo-Cheng Lai, Chien-Nan Jimmy Liu, Shyh-Jye Jou
ICCAD10
2020 Digital Self-Healing using Smart Sensing Technique for IQ Mismatch and LO Leakage against Non-Flat Path Response in mmWave Communication System
abstract
In this paper, we propose a self-healing technique with the smart sensing capability to exactly estimate and pre-compensate both the imperfect effects of IQ mismatch and LO leakage of the up-conversion mixer under the non-flat path impulse response (PIR) for improving TX EVM in 60 GHz broadband wireless communication systems. The proposed method may use the direct digital frequency synthesis (DDFS) module to send the test signals including a single-tone (ST) and a two-tone (TT) modes for sensing PIR and estimating the imperfect effects of the up-conversion mixer, respectively. The 8-bit fixed-point simulation result shows the proposed technique can achieve very excellent results.
Ngoc-Giang Doan, Hung-Chih Liu, Chih-Wei Jen, Shyh-Jye Jou
ISCAS4
2020 A 75-Gb/s/mm2 and Energy-Efficient LDPC Decoder Based on a Reduced Complexity Second Minimum Approximation Min-Sum Algorithm
abstract
This article presents a high-throughput and low-routing complexity low-density parity check (LDPC) decoder design based on a novel second minimum approximation min-sum (SAMS) algorithm. The routing congestion is mitigated by reducing the required interconnections in the critical path of the routing network. The implementation and postlayout results with 28-nm 1P9M CMOS process show that the proposed design can achieve a throughput of 10.5 Gb/s for a millimeter-wave 60-GHz baseband system while satisfying the low bit error rate (BER) requirements (10-7). The proposed design reduces the wiring in the routing network by 21% and improves the area by 12% compared to the conventional min-sum (MS) and normalized MS (NMS) algorithm. Additional hardware optimizations are obtained by considering the internal message passing resolution based on the BER and signal-to-noise ratio (SNR) requirements for a practical baseband system. The power consumption is efficiently reduced by the employment of a shared address generator that exploits the degree of parallelism to reduce the switching activity on a group of memory elements. The LDPC decoder is implemented with a core area of 0.14 mm2, power consumption of 81 mW at 312.5 MHz, and the area and power efficiency of 75 Gb/s/mm2and 10.2 pJ/bit, respectively.
Henry Lopez Davila, Hsun-Wei Chan, Kang-Lun Chiu, Pei-Yun Tsai 0001, Shyh-Jye Jou
IEEE Trans. Very Large Scale Integr. Syst.5
2019 A 50 Gb/s Adaptive Dual Data-Paths NS-EICL ADFE with 50 Parallelisms for 2-PAM Systems
abstract
A 50Gb/s all-digital adaptive noise-suppression (NS) feed-forward equalizer (AFFE) and adaptive decision feedback equalizer (ADFE) for 2-level pulse amplitude modulation (2-PAM) serial link systems is presented. Based on a parallel extended incremental coefficients-lookahead scheme (EICL), we propose a Dual Data-paths Self-Lookahead Filter (DD-SLF) for ADFE. DD-SLF architecture has better energy efficiency and hardware area than an original SLF architecture due to the number of delay elements in the feedback loop is reduced. Furthermore, gated clock technique with the design idea of register file architecture is used to replace the pipelined delay elements to save power. The whole equalizer which operates at 1GHz system clock rate with 50 parallelisms is implemented in 40nm CMOS technology with a 0.38mm2core area. The equalizer with 50Gb/s throughput rate achieves 2.6pJ/bit energy efficiency under 0.81V supply measurement results.
Chee-Kit Ng, Kang-Lun Chiu, Yu-Chun Lin, Shyh-Jye Jou
ISCAS4
2018 Digital Self-Interference Cancellation for OFDM Full-Duplex Transmission in 60 GHz Band
abstract
In this paper, we proposed an all-digital self-interference cancellation (DSIC) with correction vector for convolution (CVC) technique to eliminate serious multipath self-interference for orthogonal frequency-division multiplexing (OFDM) full-duplex transmission in 60 GHz band. The CVC can solve the non-equivalence of circular and linear convolution. This work based on IEEE 802.15.3c/IEEE 802.11ad specification is synthesized with 40 nm 1P9M general purposes process. Under the Access Point (AP) and the User Equipment (UE) horn antennas being aligned face to face, the 9-bit fixed-point simulation result shows the proposed technique can achieve the target of the uncoded bit-error-rate (BER) of 10-2at signal-to-noise ratio (SNR) of 24.8 dB satisfying the requirement of 26 dB.
Chih-Wei Jen, Hung-Wei Yang, Hsun-Wei Chan, Hung-Chih Liu, Henry Lopez Davila, Chun-Yi Liu 0001, Shyh-Jye Jou
ISCAS7
2018 A 40Gb/s All-Digital Adaptive Noise-Suppression Feed-Forward Filter and Adaptive Decision Feedback Equalizer with 40 parallelisms for 2-PAM Systems
abstract
A 40Gb/s all-digital adaptive noise-suppression feed-forward filter/equalizer (AFFE) and adaptive decision feedback equalizer (ADFE) for 2-level pulse amplitude modulation (2-PAM) systems is presented. Batch mode coefficients update (BMCU) unit together with coefficients-lookahead scheme are proposed to achieve high parallelism architecture for ADFE. With these schemes, new extended incremental coefficients-lookahead filter architecture is proposed to provide high throughput rate and to reduce hardware complexity of parallel ADFE. Besides, feed-forward noise-suppression architecture is proposed for AFFE to provide better signal-to-noise ratio (SNR). The equalizer operates at 1 GHz system clock with 40 parallelisms is implemented in 40nm CMOS technology with a core area 0.23mm2. The measurement results verify the equalizer performance and the maximum throughput of 40Gb/s is achieved under 0.9V supply with 4.35pJ/bit energy efficiency.
Chee-Kit Ng, Yu-Chun Lin, Wei-Chang Liu, Chih-Feng Wu, Shyh-Jye Jou
ISCAS5
2017 Residual sampling clocking offset estimation and compensation for FBMC-OQAM baseband receiver in the 60 GHz band
abstract
In this paper, we propose a tracking mechanism for the residual sampling clocking offset (SCO) with corresponding pilot and auxiliary pilot arrangements for filter bank multi-carrier (FBMC) offset QAM (OQAM) baseband receiver in the 60 GHz band. This work can effectively compensate the non-ideal effects on the high-band subcarriers while using more data subcarrier for increasing bandwidth efficiency. This work is designed based on IEEE 802.15.3c and IEEE 802.11ad standards and synthesized with 40 nm 1P9M general purposes (GP) process. The proposed SCO tracking module with the proposed alternated pilot and auxiliary pilot arrangement can improve the BER floor and reach to 10−2 at SNR of 22 dB which is only 1.75 dB more than that of no SCO effect in 64-QAM modulation.
Chun-Yi Liu 0001, Yu-Cheng Yao, Meng-Siou Sie, Edmund Wen Jen Leong, Henry Lopez Davila, Chih-Wei Jen, Shyh-Jye Jou
ISCAS7
2016 Error-resilient sequential cells with successive time borrowing for stochastic computing
abstract
This paper presents error-resilient sequential building blocks with time-borrowing capability without extra latches and generated clocks. The circuits are able to recover the timing errors caused by PVT variations and/or over-voltage scaling by up to half a cycle. Unlike prior works, the timing errors can be recovered dynamically through successive time borrowing without stalled cycles, retaining a constant throughput. The circuit structure can be applied to both ASICs and microprocessors. The proposed sequential cells are highly compatible with current cell-based IC design flow, for both feedforward and feedback datapaths. As a proof of concept, a design with key DSP building blocks has been verified. The results show that the performance of the DSP modules is improved by 13-15% in the worst-case operation condition, yielding a promising solution for stochastic computing under an unreliable operation condition.
Wei-Chang Liu, Ching-Da Chan, Shuo-An Huang, Chi-Wei Lo, Chia-Hsiang Yang, Shyh-Jye Jou
ICASSP6
2016 A subthreshold SRAM with embedded data-aware write-assist and adaptive data-aware keeper
abstract
We propose a data-aware power cut-off write-assist 12T SRAM cell (DPC12T) which improves the write-ability to improve the write minimum operating voltage (VMIN). Moreover, we propose an adaptive data-aware keeper (DAK) to lower the design conflicts among the keeper current, read current and the bit-line leakage current to improve the read stability and read VMINfor single-ended read operation. Fabricated 40nm 8kb test chip macro with 64 cells per bit-line can achieve VMIN250 mV and 230 mV without and with enabling DAK at 6 MHz and 4 MHz, respectively. The SRAM test macro with 256, 512 and 1024 cells per bit-line demonstrates that DAK improves the read VMIN by 9% to 21% at low supply voltages.
Yi-Wei Chiu, Yu-Hao Hu, Jun-Kai Zhao, Shyh-Jye Jou, Ching-Te Chuang
ISCAS4
2016 A memory access reordering polyphase network for 60 GHz FBMC-OQAM baseband receiver
abstract
In this paper, a novel memory access reordering polyphase network (PPN) for 60 GHz filter bank multi-carrier (FBMC) offset QAM (OQAM) baseband receiver is presented. The 8X-parallelism architecture of PPN is integrated into our 8X-parallelism digital baseband receiver. The PPN is designed based on IEEE 802.15.3c and IEEE 802.11ad standards. The PPN contains three key modules including memory bank, filter coefficient selector and 4-tap PPN filter. The proposed PPN is synthesized with 40 nm 1P9M general purposes (GP) process. The implementation result shows it can operate at specified 330/500 MHz clock rate with power consumption of 17/26 mW at 0.81 V supply voltage. With 8X-parallelism architecture, the sampling rate supports up to 2.64/4 GHz.
Chun-Yi Liu 0001, Meng-Siou Sie, Edmund Wen Jen Leong, Yu-Cheng Yao, Chih-Wei Jen, Shyh-Jye Jou
ISCAS6
2016 A Systematic ANSI S1.11 Filter Bank Specification Relaxation and Its Efficient Multirate Architecture for Hearing-Aid Systems
abstract
Recently, emerging mobile computing requires the high integration of hearing aids into a single system-on-chip. Modern hearing aid systems include a frequency decomposer, noise reduction, feedback cancellation, auditory compensation, and intelligent adaptation. The majority of existing works concentrated on improving the performance and efficiency on a single-signal processing block or one-sided effect. These works lacked comprehensive discussions on system-wide aspects regarding the overall impacts. To design an optimal hearing aid system, frequency decomposers, or the filter banks that dominate in hearing aid systems, are the first priority. We propose a systematic relaxation of the ANSI S1.11 specification and its design procedure for filter banks. The proposed design procedure overcomes the drawbacks of previous works and changes the five performance indices of the filter bank: delay, complexity, sub-band rate reduction, ripples of synthesized output, and prescription matching errors. These performance indices help system or algorithm designers in selecting a beneficial system-relaxed filter bank to achieve optimal hearing aids. The proposed multirate filter bank using the resampling method provides an efficient, low complexity, and delay-constrained computing architecture. Finally, seven design cases are used to demonstrate the proposed method, and comprehensive discussions of the five performance indices are presented.
Cheng-Yen Yang, Chih-Wei Liu, Shyh-Jye Jou
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Golay-Correlator Window-Based Noise Cancellation Equalization Technique for 60-GHz Wireless OFDM/SC Receiver
abstract
In this paper, a Golay-correlator window-based noise cancellation (GC-WNC) technique with frequency-domain equalizer (FDE) is proposed. The GC-WNC is a cooperative scheme in the time and frequency domains to combat the multipath effect in nonline-of-sight (NLOS) and LOS channels for orthogonal frequency-division multiplexing (OFDM) and single-carrier mode baseband inner receiver over 60-GHz environment for IEEE 802.15.3c and 802.11ad. According to mean-square error criterion, WNC approach is to minimize the estimation error between the ideal and the estimated channel frequency response (CFR) on each subchannel. The CFR is precisely obtained as coefficients of FDE to compensate multipath effect even in NLOS channel. The GC-WNC FDE with 8X-parallelism is designed as a part of digital baseband inner receiver with 40-nm CMOS general-purpose process. Because of area restriction of tape-out chip, only the OFDM mode is fabricated in the chip. The GC-WNC FDE has an equivalent gate count of 230k occupying 11.3% of the baseband inner receiver. Based on the chip measurement results, the baseband inner receiver with GC-WNC FDE provides 24-Gb/s throughput with 500-MHz operating clock and 0.94 V supply voltage. The power consumption of GC-WNC FDE is 69.79 mW. The baseband inner receiver with GC-WNC FDE can deliver a multigigabit per second throughput with the power dissipation of 2.91/2.26 mW/Gb/s at 500-/330-MHz operating clock for the OFDM mode.
Chih-Feng Wu, Wei-Chang Liu, Chia-Chun Tsui, Chun-Yi Liu 0001, Meng-Siou Sie, Shyh-Jye Jou
IEEE Trans. Very Large Scale Integr. Syst.6
2015 A 28nm 36kb high speed 6T SRAM with source follower PMOS read and bit-line under-drive
abstract
In this paper, we present source follower PMOS Read and bit-line under-drive techniques to improve the operation speed as compared to present commercial SRAM compilers. A source follower PMOS is utilized to connect local bit-lines (LBL) to global bit-lines (GBL) instead of using a NAND gate. To further improve the discharging time from LBL to GBL, we propose a bit-line under-drive circuit to reduce the voltage level of LBL. The simulated access time of the proposed macro is 445 ps at slow N slow P (SS) corner, -40°C, 0.81 V. As compared to the SRAM macro which is generated by commercial SRAM compilers with the fastest combination, the access time of the proposed SRAM macro is 12% faster than that of commercial SRAM compilers. A 36kb high speed 6T SRAM macros with source follower PMOS Read and bit-line under-drive techniques is fabricated in 28nm HKMG CMOS process. The measurement results of the chip in SS corner show the proposed SRAM macro passes all MBIST patterns at 500 MHz at 0.81 V, room temperature.
Chi-Hao Hong, Yi-Wei Chiu, Jun-Kai Zhao, Shyh-Jye Jou, Wen-Tai Wang, Reed Lee
ISCAS4
2015 A 0.325 V, 600-kHz, 40-nm 72-kb 9T Subthreshold SRAM with Aligned Boosted Write Wordline and Negative Write Bitline Write-Assist
abstract
This brief presents a two-port disturb-free 9T subthreshold static random access memory (SRAM) cell with independent single-ended read bitline and write bitline (WBL) and cross-point data-aware write structure to facilitate robust subthreshold operation and bit-interleaving architecture for enhanced soft error immunity. The design employs a variation-tolerant line-up write-assist scheme where the timing of areaefficient boosted write wordline and negative WBL are aligned and triggered/initiated by the same low-going global WBL to maximize the write-ability enhancement. A 72-kb test chip is implemented in United Microelectronics Corp. 40-nm low-power (40LP) CMOS. Full functionality is achieved for VDD ranging from 1.5 to 0.32 V without redundancy. The measured maximum operation frequency is 260 MHz (450 kHz) at 1.1 V (0.32 V) and 25 °C. At 0.325 V and 25 °C, the chip operates at 600 kHz with 5.78 μW total power and 4.69 μW leakage power, offering 2× frequency improvement compared with 300 kHz of our previous 72-kb 9T subthreshold SRAM design in the same 40LP technology. The energy efficiency (power/frequency/IO) at 0.325 V and 25 °C is 0.267 pJ/bit, a 23.7% improvement over the 0.350 pJ/bit of our previous design.
Chien-Yu Lu, Ching-Te Chuang, Shyh-Jye Jou, Ming-Hsien Tu, Ya-Ping Wu, Chung-Ping Huang, Paul-Sen Kan, Huan-Shun Huang, Kuen-Di Lee, Yung-Shin Kao
IEEE Trans. Very Large Scale Integr. Syst.3
2015 A Low-Jitter Cell-Based Digitally Controlled Oscillator With Differential Multiphase Outputs
abstract
A low-jitter digitally controlled oscillator (DCO) with multiphase differential outputs and good linearity is presented. The DCO is composed of four differential delay cells and can achieve linear tuning over a wide frequency range. The proposed fully differential delay cell comprises logic cells in standard library and varactors. The measured rms jitter and pk-pk jitter from 2.5-GHz carrier are 2.827 and 29 ps, respectively. The power consumption is 6 mW from a 1.2 V supply. An experimental prototype is designed using 65-nm CMOS technology, and the chip area is 156 μm × 92 μm2.
Ming-Chiuan Su, Shyh-Jye Jou, Wei-Zen Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2014 An efficient 18-band quasi-ANSI 1/3-octave filter bank using re-sampling method for digital hearing aids
abstract
This paper presents the multirate and re-sampling techniques to realize a low-delay, 18-band quasi-ANSI filter bank for digital hearing aids, which not only achieves a rather low computation complexity without a significant increase in the latency, but reduces greatly the total computation complexity for sub-band signal processing followed by the filter bank, such as noise reduction as well as wide dynamic range compression (WDRC). Researches done in the literature all focused on how to reduce the computation complexity of the filter bank. In particular, with the efficient multirate and interpolated FIR (IFIR) approaches for a 10-ms, 18-band quasi-ANSI filter bank, approximately 93% of the multiplications are saved, compared that with a straightforward parallel FIR filters architecture. However, they did not consider the computation complexity of the sub-band signal processing. In this paper, we first investigate realizing the FIR filter bank efficiently by using the multirate re-sampling techniques. To reduce the complexity, the optimized re-sampling factor for each filter band is explored carefully. Then, with the resampling technique, an efficient multirate quasi-ANSI FIR filter bank architecture is proposed. Compare to the state-of-the-art quasi-ANSI filter bank, approximately 17.7% of multiplicative complexity is reduced further and, up to 25% of the total computation complexity for sub-band signal processing followed by the filter bank is saved, but with only a slight increase in latency, i.e. 13.6 ms.
Cheng-Yen Yang, Chih-Wei Liu, Shyh-Jye Jou
ICASSP3
2014 An ultra-low voltage hearing aid chip using variable-latency design technique
abstract
This paper presents a low-power hearing aid chip which operates under near-threshold voltage region to minimize energy consumption. The proposed variable-latency design technique compensates the performance degradation under ultra-low voltage, while the proposed FIR filter computing datapath improves the energy efficiency for filter bank computation in hearing aids. The hearing aid system composed of four heterogeneous processing elements to optimize the flexibility and power consumption. The overall system was fabricated in TSMC 65nm LP process. The measured results show that the power consumption achieves 500 μW at 0.5V and 6 MHz
Kuo-Chiang Chang, Shien-Chun Luo, Ching-Ji Huang, Chih-Wei Liu, Yuan-Hua Chu, Shyh-Jye Jou
ISCAS6
2014 An IEEE 802.15.3c/802.11ad compliant SC/OFDM dual-mode baseband receiver for 60 GHz Band
abstract
In this paper, a dual-standard, dual-mode baseband receiver for 60 GHz wireless communication is presented. The receiver is designed to support SC and OFDM modes for both IEEE 802.15.3c and IEEE 802.11ad standards. The receiver is integrated with all-digital synchronization, radix-16 FFT, phase noise cancellation and low-complexity time-domain equalizer for line-of-sight channel application. The hardware utilization achieves 70% by leveraging hardware sharing between two modes and two standards for area efficiency. The receiver is implemented with 65 nm 1P9M process in 7.95 mm2core area. With feed-through architecture, the throughput rate supports up to 7.04 Gb/s and 15.84 Gb/s for SC mode (220 MHz) and OFDM mode (330 MHz), respectively.
Wei-Chang Liu, Fu-Chun Yeh, Chia-Yi Wu, Ting-Chen Wei, Ya-Shiue Huang, Shen-Jui Huang, Ching-Da Chan, Shyh-Jye Jou, Sau-Gee Chen
ISCAS8
2014 Low Complexity Formant Estimation Adaptive Feedback Cancellation for Hearing Aids Using Pitch Based Processing
abstract
This paper proposes a novel algorithm and architecture for the adaptive feedback cancellation (AFC) based on the pitch and the formant information for hearing aid (HA) applications. The proposed method, named as Pitch based Formant Estimation (PFE-AFC), has significantly low complexity compared to Prediction Error Method AFC (PEM-AFC). The proposed PFE-AFC consists of a forward and a backward path processing. The forward path processing includes a low complexity pitch based formant estimator for decorrelation filter coefficients update and a pitch based voice activity detector for speech detection, which facilitates the feedback cancellation filter in the backward path to reduce feedback component and maintain speech quality. From system point of view, the PFE-AFC has low complexity overhead since it is easy to share computation resource with other components in the HA system, such as noise reduction and auditory compensation. In addition, the PFE-AFC is suitable for hardware implementation owing to its regular structure. Complexity evaluations show that the PFE-AFC has four orders lower complexity than the PEM-AFC. Simulation results show that the PFE-AFC and the PEM-AFC can achieve similar PESQ (perceptual evaluation speech quality) and ASG (added stable gain). Moreover, the proposed PFE-AFC can outperform the conventional AFC.
Yi FanChiang, Cheng-Wen Wei, Yi-Le Meng, Yu-Wen Lin, Shyh-Jye Jou, Tian-Sheuan Chang
IEEE ACM Trans. Audio Speech Lang. Process.5
2014 Correction to "Low complexity formant estimation adaptive feedback cancellation for hearing aids using pitch based processing"
abstract
In the above paper (ibid., vol. 22, no. 8, pp.1248-1259, Aug. 2014), Table IV is incorrect. The correct table is presented here.
Yi FanChiang, Cheng-Wen Wei, Yi-Le Meng, Yu-Wen Lin, Shyh-Jye Jou, Tian-Sheuan Chang
IEEE ACM Trans. Audio Speech Lang. Process.5
2013 A 40nm 1.0Mb pipeline 6T SRAM with variation-tolerant Step-Up Word-Line and Adaptive Data-Aware Write-Assist
abstract
We present a 1.0Mb pipeline 6T SRAM in 40nm Low-Power CMOS technology. The design employs a variation-tolerant Step-Up Word-Line (SUWL) to improve the Read Static Noise Margin (RSNM) without compromising the Read performance and Write-ability. The Write-ability is further enhanced by an Adaptive Data-Aware Write-Assist (ADAWA) scheme. The 1.0Mb test chip operates from 1.5V to 0.7V, with operating frequency of [email protected] and 25°C. The measured power consumption is 23.21mW (Active)/2.42mW (Leakage) at 1.2V, TT, 25°C; and 6.01mW (Active)/0.35mW (Leakage) at 0.7V, TT, 25°C.
Chi-Shin Chang, Hao-I Yang, Wei-Nan Liao, Yi-Wei Lin, Nan-Chun Lien, Chien-Hen Chen, Ching-Te Chuang, Wei Hwang, Shyh-Jye Jou, Ming-Hsien Tu, Huan-Shun Huang, Yong-Jyun Hu, Paul-Sen Kan, Cheng-Yo Cheng, Wei-Chang Wang, Jian-Hao Wang, Kuen-Di Lee, Chia-Cheng Chen, Wei-Chiang Shih
ISCAS9
2013 A SC/HSI dual-mode baseband receiver with frequency-domain equalizer for IEEE 802.15.3c
abstract
In this paper, an 8X-parallelism digital baseband receiver is proposed for IEEE 802.15.3c application. The baseband receiver consists of all-digital synchronization, radix-16 FFT and LS-LMS equalizer modules. It supports SC and HSI dual-mode in IEEE 802.15.3c with single hardware for area efficiency. The chip is implemented with 65 nm 1P9M process. The fabricated area is 12.96 mm2with 3463 K gate counts. The post-layout verification shows the throughput rate under QPSK modulation achieves 3.52 Gb/s and 5.28 Gb/s for SC mode (220 MHz) and HSI mode (330 MHz), respectively.
Wei-Chang Liu, Fu-Chun Yeh, Ting-Chen Wei, Ya-Shiue Huang, Tai-Yang Liu, Shen-Jui Huang, Ching-Da Chan, Shyh-Jye Jou, Sau-Gee Chen
ISCAS8
2013 A 40 nm 0.32 V 3.5 MHz 11T single-ended bit-interleaving subthreshold SRAM with data-aware write-assist
abstract
This paper presents a new bit-interleaving 11T subthreshold SRAM cell with Data-Aware Power-Cutoff (DAPC) Write-assist to mitigate the leakage and variation and improve the Write-ability in deep sub-100nm technology. Measurement results from a 4 Kb test chip implemented in 40 nm General Purpose (40GP) CMOS technology operates for VDDdown to 0.32 V (~0.69X of threshold voltage) with VDDMINlimited by Read operation. The measured maximum operation frequency is 3.5 MHz (16.5 MHz) at 0.32 V (0.38 V) with total power consumption of 15.2 μW (27.2 μW) at 25 °C.
Yi-Wei Chiu, Yu-Hao Hu, Ming-Hsien Tu, Jun-Kai Zhao, Shyh-Jye Jou, Ching-Te Chuang
ISLPED5
2013 STBC-OFDM Downlink Baseband Receiver for Mobile WMAN
abstract
This paper proposes a space time block code-orthogonal frequency division multiplexing downlink baseband receiver for mobile wireless metropolitan area network. The proposed baseband receiver applied in the system with two transmit antennas and one receive antenna aims to provide high performance in outdoor mobile environments. It provides a simple and robust synchronizer and an accurate but hardware affordable channel estimator to overcome the challenge of multipath fading channels. The coded bit error rate performance for 16 quadrature amplitude modulation can achieve less than 10-6under the vehicle speed of 120 km/hr. The proposed baseband receiver designed in 90-nm CMOS technology can support up to 27.32 Mb/s uncoded data transmission under 10 MHz channel bandwidth. It requires a core area of 2.41 × 2.41 mm2and dissipates 68.48 mW at 78.4 MHz with 1 V power supply.
Hsiao-Yun Chen, Jyun-Nan Lin, Hsiang-Sheng Hu, Shyh-Jye Jou
IEEE Trans. Very Large Scale Integr. Syst.4
2012 An all-digital bit transistor characterization scheme for CMOS 6T SRAM array
abstract
We present an all-digital bit transistor characterization scheme for CMOS 6T SRAM array. The scheme employs an on-chip operational amplifier feedback loop to measure the individual threshold voltage (VTH) of 6T SRAM bit cell transistors (holding PMOS, pull-down NMOS, and access NMOS) in SRAM cell array environment. The measured voltage is converted to frequency with dual VCO and counter based digital read-out to facilitate data extraction, processing, and statistical analysis. A 512Kb test chip is implemented in 55nm 1P10M Standard Performance (SP) CMOS technology. Monte Carlo simulations indicate that the accuracy of the VTHmeasurement scheme is about 2-7mV at TT corner across temperature range from 85°C to -45°C, and post-layout simulations show the resolution of the digital read-out scheme isTHdistributions agree well with Monte Carlo simulation results.
Geng-Cing Lin, Shao-Cheng Wang, Yi-Wei Lin, Ming-Chien Tsai, Ching-Te Chuang, Shyh-Jye Jou, Nan-Chun Lien, Wei-Chiang Shih, Kuen-Di Lee, Jyun-Kai Chu
ISCAS6
2012 High-performance 0.6V VMIN 55nm 1.0Mb 6T SRAM with adaptive BL bleeder
abstract
This paper presents a 1.0Mb high-performance 0.6V VMIN6T SRAM design implemented in UMC 55nm Standard Performance (SP) CMOS technology. This design utilizes an adaptive LBL bleeder technique to reduce Read disturb and Half-Select disturb of 6T cells while maintaining adequate sensing margin. A bleeder timing control circuit adaptively adjusts the LBL voltage level prior to Read/Write operation to facilitate wide operation voltage range. Hierarchical WL, hierarchical BL, and distributed replica timing control scheme are used to improve SRAM performance. Based on measurement results, the SRAM operates from 1.5V down to 0.6V. The maximum operating frequency is [email protected] and [email protected].
Hao-I Yang, Yi-Wei Lin, Mao-Chih Hsia, Geng-Cing Lin, Chi-Shin Chang, Yin-Nien Chen, Ching-Te Chuang, Wei Hwang, Shyh-Jye Jou, Nan-Chun Lien, Hung-Yu Li, Kuen-Di Lee, Wei-Chiang Shih, Ya-Ping Wu, Wen-Ta Lee, Chih-Chiang Hsu
ISCAS9
2012 Testing strategies for a 9T sub-threshold SRAM
abstract
Due to the increasing demands of lower-power devices, a lot of research effort has been devoted to develop new SRAM cell designs that can be effectively and economically operated at the subthreshold region. However, each new SRAM cell design has its own cell structure and design techniques, which may result in different faulty behaviors than the conventional 6T SRAMs and require specialized test methods to detect those uncovered fault models. In this paper, we focus on developing the test methods for testing a new 9T subthreshold SRAM design, which utilizes single bit-line read/write, two write word-lines for writing different values, and a separate read path. A mixed march algorithm with different background and address-traverse directions is proposed to detect various uncovered fault models and validated through real test chips. A new specialized technique of floating bit-line attacking is also presented to detect the stability faults, which cannot be effectively detected by applying the conventional test methods, for the new 9T SRAM design.
Hao-Yu Yang, Chen-Wei Lin, Hung-Hsin Chen, Mango Chia-Tso Chao, Ming-Hsien Tu, Shyh-Jye Jou, Ching-Te Chuang
ITC6
2012 Sub µW Noise Reduction for CIC Hearing Aids
abstract
This paper presents a sub noise reduction design to enhance speech for completely-in-the-canal (CIC) type hearing aids by optimizing its algorithm and associated architecture. In algorithm optimization, a low-complexity mixed perceptual-discrete wavelet packet transform (P-DWPT) and fast Hartley transform (FHT) are adopted for spectral decomposition and reconstruction. A simple yet efficient denoise method with 4-zone-voice activity detection (VAD) supports a consonant protection to improve speech quality and a skip scheme to reduce power consumption. In the designed architecture, mixed P-DWPT and FHT are folded into one 8-by-8 configurable butterfly computation unit with on-time scheduling for low power operation. The circuit is implemented with 0.18 μm CMOS process and consumes only 0.65 μW power at 1.0 V with a speech quality that is comparable to that achieved using other high-complexity algorithms.
Cheng-Wen Wei, Sheng-Jie Su, Tian-Sheuan Chang, Shyh-Jye Jou
IEEE Trans. Very Large Scale Integr. Syst.4
2011 Low power InfomaxICA with compensation strategy for binaural hearing-aid
abstract
Binaural hearing-aids are under intensive study nowadays. The information exchange across both ears provides an opportunity to perform the blind source separation to enhance the signal SNR. This study combines the conventional InfomaxICA with our proposed novel algorithms: binaural delay compensation, minimum-interference initial de-mixing guess and low power design concept, for a real-time binaural hearing-aid application. Simulations demonstrate our algorithms provide 20.5 to 30.6 dB gain at 0 to 35 points delay at SNR= 0 dB in our assumed scenario. Therefore, our algorithms are robust under various practical wearing conditions.
Fan-Chiang Yi, Ching-Wen Huang, Tai-Shih Chi, Shyh-Jye Jou
ISCAS4
2011 8T single-ended sub-threshold SRAM with cross-point data-aware write operation
Yi-Wei Chiu, Jihi-Yu Lin, Ming-Hsien Tu, Shyh-Jye Jou, Ching-Te Chuang
ISLPED4
2009 10Gbps Decision Feedback Equalizer with Dynamic Lookahead Decision Loop
abstract
Decision feedback equalizer (DFE) uses a feedback path to cancel post-cursor ISI, and this feedback path will also cause the limitation of its maximum throughput rate. This paper proposes a new lookahead method to break the feedback path for multi-gigabit DFE design. After lookahead computation, each paralleled sub-circuit has the same throughput rate as original one. Therefore, the total throughput rate is proportional to the parallelization factor. The computation complexity of the proposed architecture is lower than that of multiplexer-based lookahead DFE if the tap number of the feedback filter is large. It is shown that the new method saves 10% hardware complexity for an 8 taps feedback filter DFE and 98% hardware complexity for a 12 taps feedback filter DFE in comparison to a 10 Gbps multiplexer-based lookahead DFE.
Yu-Chun Lin, Muh-Tian Shiue, Shyh-Jye Jou
ISCAS3
2008 Symbol and carrier frequency offset synchronization for IEEE802.16e
abstract
IEEE 802.16e standard has been proposed as a specification for the next generation wireless communication system. Synchronization plays an important role for a wireless receiver. Because the repetition of preamble is unapparent in orthogonal frequency division multiple access (OFDMA) modulation mode, a correlation based scheme like match filter is used to estimate the symbol boundary and integer carrier frequency offset (ICFO). In this paper, an effective hardware architecture of correlation is proposed. By adopting the modified algorithm, mass of multipliers are removed from hardware implementation. Design results show that 56% area reduction and 64% power saving are achieved. Moreover, a ping-pong algorithm is proposed to increase the accuracy of ICFO synchronization about two orders at most.
Jyun-Nan Lin, Hsiao-Yun Chen, Ting-Chen Wei, Shyh-Jye Jou
ISCAS4
2008 A reconfigurable MAC architecture implemented with mixed-Vt standard cell library
abstract
In this paper, a 32-bit reconfigurable multiplication-accumulation architecture, which can execute flexibly one 32×32, two 16×16 or four 8x8 two’s complement multiply-accumulation, is proposed and demonstrated. It is based on the modified Booth encoding scheme and designed with techniques of reducing sign-extension bits, removing one extra partial product row and adjusting the positions of hot signals but with elegant modifications. It is implemented with a 130nm mixed-VtCMOS standard cell library and shows saving of area and power consumption by approximately 16% and 14% respectively as compared to the previous design.
Li-Rong Wang, Yi-Wei Chiu, Chia-Lin Hu, Ming-Hsien Tu, Shyh-Jye Jou
ISCAS5
2007 Novel Programmable FIR Filter Based on Higher Radix Recoding for Low-Power and High-Performance Applications
abstract
This paper proposes the novel programmable digital finite impulse response (FIR) filters for low-power and high-performance applications. The architecture is based on the higher radix recoding scheme which specifically targets the reduction of power consumption by decreasing partial products and precomputation sharing in each partial product. We extend higher radix recoding scheme with secondary radix recoding to further improve performance by reducing the propagation delay of precomputing odd multiples of the multiplicand. A 10-tap programmable FIR filter based on proposed schemes was design in TSMC 0.18-μm technology. The performance and power consumption of the proposed schemes can improve about 40.9-65.5% and 28.3-50.2% over existing designs.
Hsiao-Yun Chen, Shyh-Jye Jou
ICASSP (3)2
2007 Blind Mode/GI Detection and Coarse Symbol Synchronization for DVB-T/H
abstract
In a non-data-aided (NDA) broadcasting system such as DVB-T/H, blind transmission mode, guard interval (GI) length detection and coarse symbol synchronization (CSS) play important roles to estimate the transmitted OFDM symbol parameters and start the synchronization processes. In this paper, a single hardware and division-free architecture, modified from normalized-maximum-correlation (NMC) architecture, for DVB-T/H blind mode/GI detection and coarse symbol synchronization (CSS) is proposed. By adopting the proposed twister memory access scheme and sequential blind mode detection scheme, the architecture reduces 33% of memory costs and at most 58.18% mode detection latencies.
Wei-Chang Liu, Ting-Chen Wei, Shyh-Jye Jou
ISCAS3
2006 Memory reduction ICFO estimation architecture for DVB-T
abstract
OFDM system is sensitive to CFO (carriers frequency offset) which is divided into fractional part and integer part. In usually, estimation of integer CFO needs to store one OFDM symbol. However, the maximum length of FFT in DVB-T is 8192, so it requires a lot of storage units. In this paper, we propose a memory reduction algorithm for integer CFO estimation used in DVB-T. This memory reduction algorithm reduces the usage of the continual pilots and only sign bit is stored. Thus, the usage of memory is reduced by 90%, the access number is reduced by 94% and the estimation time is reduced by 37%.
Ting-Zhen Wei, Shyh-Jye Jou, Muh-Tian Shiue
ISCAS2
2003 5Gbps serial link transmitter with pre-emphasis
abstract
High-speed serial link that achieves Gbps has the advantage of low cost and thus become popular. In this paper, we will implement the high-speed data serial link transceiver and demonstrate the pre-emphasis circuit. The overall circuit is implemented in TSMC 0.18um 1P6M 1.8v CMOS process. The performance of the transceiver can reach 5Gbps over the 10-meter long cable.
Chih-Hsien Lin, Chung-Hong Wang, Shyh-Jye Jou
ASP-DAC3
2003 Low-power digital CDMA receiver
abstract
The advanced design task of the digital CDMA receiver are presented in this report. A biased number system and architecture is used to reduce the switching activity to reduce power consumption. Carry-save adder tree is used to speed up the summation of 127 data (3 bits) in the synchronization and date extraction process. Verilog HDL is used to describe this system and design compiler of Synopsys is used to synthesize our design. Design results show that it can work at 155MHz (Chip rate) with 9913 Gate counts by using Compass 0.35 um CMOS Cell library.
Ja-Sheng Liu, I-Hsin Chen, Yi-Chen Tsai, Shyh-Jye Jou
ASP-DAC4
2001 Intrinsic response for analog module testing using an analog testability bus
abstract
A parasitic effect removal methodology is proposed to handle the large parasitic effects in analog testability buses. The removal is done by an on-chip test generation technique and an intrinsic response extraction algorithm. On-chip test generation creates test signals on-chip to avoid the parasitic effects of the test application bus. The intrinsic response extraction cross-checks and cancels the parasitic effects of both test application and response observation paths. The tests using both SPICE simulation and MNABST-1 P1149.4 test chip reveal that the proposed algorthm can not only remove the parasitic effects of the test buses but also tolerate test signal variations. Furthermore, it is robust enough to handle loud environmental noise and the nonlinearity of the switching devices.
Chauchin Su, Yue-Tsang Chen, Shyh-Jye Jou
ACM Trans. Design Autom. Electr. Syst.3
2000 Fixed-Width Multiplier for DSP Application
abstract
A new compensation method that reduces the error of fixed-width multiplier for digital signal processing (DSP) application is proposed. The designs of using this input number based compensation method are carried out on array multiplier and Booth multiplier. The hardware complexity is reduced to about 50% of the original multiplier. Design results show that the new architectures have lower hardware overhead, lower error and fast operation speed as compared with other proposed architectures.
Shyh-Jye Jou, Hui-Hsuan Wang
ICCD1
1999 Decentralized BIST Methodology for System Level Interconnects
Chauchin Su, Shyh-Jye Jou
J. Electron. Test.2
1997 Structural approach for performance driven ECC circuit synthesis
abstract
ECCGen is a logic synthesizer for error control coding circuits. It takes H matrices as inputs and produces circuit schematics in two steps, literal minimization, and gate/pin assignment. Different from conventional logic synthesis tools, it takes a structural approach to avoid the combinatorial explosion problem in Boolean function and/or true table representations of ECC circuits. Moreover, the structural approach also reduce the complexity of timing and area optimization significantly when multiple-input exclusive-or gates are used. The test results show that ECCGen achieves a reduction of 57% in transistor count and 15% in delay time on thirteen industrial ECC circuits.
Chauchin Su, Kathy Y. Chen, Shyh-Jye Jou
ASP-DAC3
1997 Parasitic Effect Removal for Analog Measurement in P1149.4 Environment
abstract
An intrinsic response extraction algorithm is derived and implemented to remove the parasitic effects in P1149.4 analog measurement. The methodology is tested on and verified by SPICE simulation results and real measurement data.
Chauchin Su, Yue-Tsang Chen, Shyh-Jye Jou
ITC3
1996 Syndrome Simulation And Syndrome Test For Unscanned Interconnects
abstract
In this paper, we present a syndrome test methodology for the testing of unscanned interconnects in a boundary scan environment. Mathematical equations are derived for the relationship of test length, fault-free and faulty syndromes, and tolerable error rate. To calculate fault-free and faulty syndromes, we propose an event driven syndrome simulation algorithm. To shorten testing time and reduce test cost, we transform and solve the problem as a set covering problem.
Chauchin Su, Shyh-Shen Hwang, Shyh-Jye Jou, Yuan-Tzu Ting
Asian Test Symposium3
1996 Metrology for analog module testing using analog testability bus
abstract
In this paper, we propose a method to generate high quality test waveform on chip to avoid the parasitic effects in an analog testability bus test environment. For the test response analysis, we derive an extraction methodology to remove the parasitic effects and obtain the intrinsic response of the CUT. The test results show that the algorithm is robust such that the intrinsic responses remain the same regardless of the small variation in the test waveforms. With the concept of intrinsic responses, we are able to use a single library for the testing and diagnosis of multiple instantiation of an analog module.
Chauchin Su, Yue-Tsang Chen, Shyh-Jye Jou, Yuan-Tzu Ting
ICCAD3
1995 Impulse response fault model and fault extraction for functional level analog circuit diagnosis
abstract
In this paper, a functional fault model for analog circuit diagnosis is proposed. A faulty module is modeled as a fault-free module in serial or in parallel with a fault module. To extract such a fault module, we adopt an iterative deconvolution technique to deconvolute the impulse response of the fault module from the faulty response. The test results show that with such a fault model and fault extraction technique the diagnostic resolution is improved significantly due to the separation of the fault and the system function. Moreover, such a fault model allows single-module fault tables to be applied to the diagnosis of a multi-module system.
Chauchin Su, Shenshung Chiang, Shyh-Jye Jou
ICCAD3
1995 A Parallel Event-Driven MOS Timing Simulator on Distributed-Memory Multiprocessors
abstract
In this paper, PMOTA, a parallel event-driven relaxation-based algorithm for the timing simulation of MOS circuits, is proposed. Two new schemes, static data distribution and dynamic event assignment, are derived to reduce the communication overhead and achieve high speedup ratio, high processor utilization and balanced load among processor elements. Implementation results show that parallel simulations on distributed memory multiprocessor are feasible for large scale circuit simulation.
Wen-Hsing Hsieh, Shyh-Jye Jou, Chauchin Su
ISCAS2
1995 Circuits Design Optimization Using Symbolic Approach
abstract
A methodology and integrated system for the optimization of integrated circuits is presented. We propose the method that combines the symbolic analyzer and effective optimization algorithms to shorten the computation time. After selecting a circuit topology, the symbolic simulator SAGA2 is invoked to model the circuit. SAGA2 generates exact analytic expressions which describe the circuit's behavior. The model is then passed to the design optimization program. The optimization program sizes all elements to satisfy the performance constraints and a user-defined design objective. A forms-based user interface is provided to allow the designer to specify the problem easily. The system had been applied in the filter circuit design.
Shyh-Jye Jou, Kou-Fong Liu, Chauchin Su
ISCAS1
1994 Hierarchical Techniques for Symbolic Analysis of Large Electronic Circuits
abstract
A hierarchical symbolic analyzer (SAGA2) for the analysis of electronic circuits is presented. SAGA2 analyzes lumped, linear, or linearized (small-signal) circuits in the S- and Z-domain. For large circuits, a hierarchical two-port method is used that is two to three order faster than that without using the hierarchical method. Also, the memory used is dramatically reduced.>
Shyh-Jye Jou, Mei-Fang Perng, Chauchin Su, C. K. Wang
ISCAS1
1994 An IDDQ Based Built-in Concurrent Test Technique for Interconnects in a Boundary-Scan Environment
abstract
An I/sub DDQ/ based scheme has been presented for concurrent built-in self-test of MCM interconnects. The scheme detects interconnect faults while the system is on-line.
Chauchin Su, Kychin Hwang, Shyh-Jye Jou
ITC3