VLDB 2026 Research / reviewers in the wild / expert
Ming-Der Shieh
dblp:21/4938
· DBLP profile ↗
56ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-7361-1860ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Computer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Unified FFT/NTT Design for Efficient NTRU Equation Solving in FALCON CryptographyabstractPost-Quantum Cryptography has garnered significant attention with the rapid development of quantum computing. FALCON, a lattice-based quantum-resistant digital signature scheme, has been standardized. During the key pair generation process, solving the NTRU equation based on Galois theory is required. The simultaneous use of two distinct algebraic structures, $\mathbb{Q}[x]/\left( {{x^n} + 1} \right)$ and $\mathbb{Z}[x]/\left( {{x^n} + 1} \right)$, where n is a power of 2, presents challenges in hardware design. This paper proposes a hardware accelerator specifically customized for FFT/IFFT and NTT/INTT operations on rings under different point numbers. The design incorporates conflict-free scheduling for memory access and the selection of twiddle factors across various radices. The proposed design occupies an area of 0.54 mm2and operates at a frequency of 333 MHz based on the synthesis results using 40nm technology. Compared to existing FFT/IFFT design on rings, the proposed one reduces the area requirement of processing elements by 33% without increasing the overall execution time. Moreover, our design supports both 64-bit radix-4 and 128-bit radix-2 NTT/INTT. This solution effectively addresses the two transformations required for key pair generation in FALCON. Jun-Hao Liao, Kuan-Ting Cho, Ming-Der Shieh |
ISCAS | 3 |
| 2025 | A Low-complexity and Reconfigurable Design for Nonlinear Function Approximation in TransformersabstractNonlinear function approximation has been studied for decades. With the rise of transformer-based AI models, the need for efficient, low-complexity circuit implementations for functions like softmax, GELU, and layer normalization has intensified due to non-negligible hardware overhead. Existing methods reuse softmax hardware for GELU or reconfigure both, but the preprocessing required for their GELU approximation results in area inefficiency. To address this, we propose a novel successive approximation technique that reduces preprocessing complexity and compensates for errors in successive steps. Additionally, our reconfigurable design supports softmax, GELU, and square root functions, optimizing hardware area and flexibility. Experimental results show a 2.48x increase in throughput per area for softmax and a 4.96x increase for GELU, with only a 0.09% accuracy loss under BERT-base models in comparison to related work. Qi-Xian Wu, Shu-Sian Teng, Ming-Der Shieh, Chih-Tsun Huang, Juin-Ming Lu |
ISCAS | 3 |
| 2022 | Battery Pack Reliability and Endurance Enhancement for Electric Vehicles by Dynamic ReconfigurationabstractIn the fast-growing electric vehicle (EV) industry, key technology challenges include the improvement of battery efficiency, reliability, and endurance. In this paper, we propose a novel battery pack design methodology that supports dynamic reconfiguration of the battery pack architecture, which partitions the battery modules into the primary group and secondary group. The proposed reconfigurable design effectively improves the battery pack reliability and endurance, especially for battery packs that contain modules with uneven aging conditions. Simulation results show that, with our approach, the equivalent average aging speed among all battery modules is slowed down, and the battery pack's endurance is increased. Yu-You Chou, Cheng-Wen Wu, Ming-Der Shieh, Chao-Hsun Chen |
ATS | 3 |
| 2022 | Aging Impact of Power MOSFETs in Charger with Different Operation FrequencyabstractThe global warming and pollution issues have reached an alarming level that requires immediate actions by all governments in the world to reduce fossil fuel vehicles, and even ban them in the foreseeable future. As a result, electric vehicles (EVs) are gaining ground rapidly. Inside the EVs, especially their power train, semiconductor power devices are considered the key components. Increasing the performance, energy efficiency, and reliability of power MOSFET, therefore, is critical in the EV industry. In this work, we stress the reliability and lifetime of power MOSFET devices and propose an aging model for them. We will show the procedure to construct the efficient aging model for a power MOSFET, supported by circuit-level simulation results. The aging conditions, including high temperature and/or high voltage on the gate oxide, will result in threshold voltage shift, which in turn decreases the MOSFET's switching speed and driving capability. With the proposed aging assessment tool for power MOSFET, the device's performance under different aging conditions can be predicted. The experiment results show that aging will reduces its switching speed and driving power. Experimental results by our tool also show that the circuit efficiency, ripple voltage, switching loss, etc., are negatively affected by aging. Kuan-Hsun Duh, Cheng-Wen Wu, Ming-Der Shieh, Chao-Hsun Chen, Ming-Yan Fan |
ATS | 3 |
| 2022 | Efficient VLSI Architecture of Bluestein's FFT for Fully Homomorphic EncryptionabstractFully homomorphic encryption (FHE) is a powerful scheme that allows computations to be performed on encrypted data. To reduce the computational complexity, double-CRT representation has been adopted in the BGV-FHE cryptosystem, in which the 2nd-CRT, also known as the polynomial-CRT, can be viewed as performing Discrete Fourier Transform (DFT). Since the point size of the DFT is usually non power of two, the traditional Cooley-Tukey FFT algorithm cannot be directly applied to reduce the complexity. This paper explores efficient VLSI architecture of Bluestein’s FFT for BGV-FHE applications. A mixed-radix single-port merged-bank memory addressing algorithm is presented to increase the effective memory bandwidth and to reduce the required memory area concurrently. The evaluation was conducted by implementing a Bluestein’s FFT compiler that can be configured to generate different point sizes of DFT designs for BGV-FHE. Analytical results also show that the proposed Bluestein’s FFT design can lead to a more area-efficient solution as compared to the mixed-radix counterpart in general cases. Shi-Yong Wu, Ming-Der Shieh |
ISCAS | 3 |
| 2022 | A Decision Tree-Based Screening Method for Improving Test Quality of Memory ChipsabstractThere is a growing demand for high-reliability and high-quality integrated circuit (IC) products, while their test costs should be kept as low as possible. We investigate the test process of advanced memory chips, where the high temperature operating life (HTOL) test has been used to determine their intrinsic reliability. This high temperature sampling test can run from 168 to 1,000 hours, so it is time-consuming and expensive. Recently, machine learning (ML) algorithms have been used to solve classification problems, so far as good training data can be obtained. In our case, there is already a large amount of parametric test data generated from the existing test flow. Therefore, in this work, we propose a decision tree (DT)-based screening method to predict weak (unreliable) dies that would fail the HTOL test. We show that experienced test engineers can prioritize the parametric test data for better use of the DT model. Finally, we take advantage of the high interpretability of DT to develop the multi-feature heuristics, which can be used to improve the quality of final test (FT). Keeping the overkill rate at 0%, our heuristics can screen out 25% more bad dies, i.e., we can improve the FT quality without additional cost. Ya-Chi Cheng, Pai-Yu Tan, Cheng-Wen Wu, Ming-Der Shieh, Chien-Hui Chuang, Gordon Liao |
ITC-Asia | 4 |
| 2022 | Weak Die Screening by Feature Prioritized Random Forest for Improving Semiconductor Quality and Reliabilityabstractwith the increasing demand for safety-critical products, the quality and reliability of semiconductor components are among the top priorities. In recent years, test data analytics by machine learning (ML) algorithms are widely considered to have great potential for improving the quality and reliability of semiconductor chips. In this work, we inspect a typical test flow of advanced semiconductor products, and propose an ML-based weak die screening method for improving the quality and reliability of shipped products. We propose the feature prioritized random forest (FPRF) model, which can fit smoothly into the existing test flow. We perform experiments on an advanced SRAM product using the FPRF model. We perform feature analysis based on the test data obtained from the final test (FT). After the FPRF screening, we are able to screen out more bad dies from those that have passed the FT. For an overkill rate of 12.93 %, the bad die hit rate can be as high as 96.55%. One can explore the FPRF model for other products as well. Shian-Yu Lin, Pai-Yu Tan, Cheng-Wen Wu, Ming-Der Shieh, Chien-Hui Chuang, Gordon Liao |
ITC-Asia | 4 |
| 2022 | Improving Test Quality of Memory Chips by a Decision Tree-Based Screening MethodabstractThere is a growing demand for high-reliability and high-quality integrated circuit (IC) products, while their test costs should be kept as low as possible. We investigate the test process of advanced memory chips, where the high temperature operating life (HTOL) test has been used to determine their intrinsic reliability. This high temperature sampling test can run from 168 to 1,000 hours, so it is time-consuming and expensive. Recently, machine learning (ML) algorithms have been used to solve classification problems, so far as good training data can be obtained. In our case, there is already a large amount of parametric test data generated from the existing test flow. Therefore, in this work, we propose a decision tree (DT)-based screening method to predict weak (unreliable) dies that would fail the HTOL test. We show that experienced test engineers can prioritize the parametric test data for better use of the DT model. Finally, we take advantage of the high interpretability of DT to develop the multi-feature heuristics, which can be used to improve the quality of final test (FT). Keeping the overkill rate at 0%, we can screen out 25% more bad dies in the 5nm SRAM case with the heuristics, and in the 4nm case, we can screen out 14% more bad dies, i.e., we can improve the FT quality without additional cost. Ya-Chi Cheng, Pai-Yu Tan, Cheng-Wen Wu, Ming-Der Shieh, Chien-Hui Chuang, Gordon Liao |
ITC | 4 |
| 2022 | Locating Image Objects With Probability DistributionsabstractIn this letter, we predict the locations as probability distributions for the tasks of image object detection. We adopt the Kullback-Leibler divergences as the regression losses to train the deep neural networks. Since most existing evaluations label the objects with rectangular bounding boxes, we propose the Nearest Distribution Converter to find the closest uniform distributions from the predicted ones. Our proposed method can improve the detected accuracy measured in mAP by 0.57%, 0.75%, and 0.48% on the models YOLOv3, the YOLOv4-tiny, and the YOLOv4, respectively. Yu-Hsiang Lin, Chih-Jen Hsu, Chih-Hung Kuo, Ming-Der Shieh |
IEEE Signal Process. Lett. | 4 |
| 2021 | On Compare-and-Swap Optimization for Fully Homomorphic Encrypted DataabstractFully homomorphic encryption (FHE) is a powerful scheme that can be applied to perform computations directly on encrypted data for ensuring the security of cloud computing. Using the concept of aggregate plaintext, this paper explores how to optimize the compare-and-swap operation, commonly used for sorting and searching in cloud computing, for FHE data. The resulting performance is optimized by properly arranging the operand locations, corresponding to the desired plaintext slots, and then scheduling a sequence of fundamental homomorphic operations designed to accomplish the targeted operation. Moreover, the hypercube plaintext slot structure is considered to efficiently manipulate the required shift operation and merge multiple operands in an aggregate plaintext. Analytical results show that the number of homomorphic multiplication and shift operations needed in the proposed compare-and-swap operation are logn + 4 and 2logn + 1, respectively, for n-bit data, which is at least 16 times faster than related works for n = 64. Applying the proposed scheme can not only reduce the size of required FHE data, but also improve the total computation time of the chosen operation. Chien-Chih Huang, Jyun-Neng Ji, Ming-Der Shieh |
ISCAS | 3 |
| 2020 | VLSI Architecture of Polynomial Multiplication for BGV Fully Homomorphic EncryptionabstractFully homomorphic encryption (FHE) has attracted much attention because computations can be directly performed on ciphertexts. This work explores the hardware architecture of polynomial multiplication defined in BGV-FHE, targeting the applications of aggregate plaintext using cyclotomic polynomials. We show how to effectively combine the characteristics of cyclotomic polynomials and the prime-factor FFT algorithm to obtain a novel design derived by the concept of Chinese Remainder Theorem. Experimental results reveal that a significant speedup in terms of operation reduction can be achieved by adopting the proposed schemes as compared to existing works assuming a comparable security level. For example, about 2.44 and 6.34 times improvement in the total number of required modular addition and multiplication, respectively, can be obtained by using 32 one-bit aggregate slots as compared to Chen's work when the 21845-th cyclotomic polynomial is considered. The improvement could be huge if all of the available slots are involved in applications. Hsuan-Jui Hsu, Ming-Der Shieh |
ISCAS | 2 |
| 2019 | Architecture-aware Memory Access Scheduling for High-throughput Cascaded ClassifiersabstractCascaded classifier based object detectors are popular for many applications because of their high efficiency. Many researches have been devoted to developing the corresponding hardware accelerators. To reduce the circuit complexity while maintaining sufficient throughput, on-chip memories are commonly partitioned into several banks for parallel data access. However, since the coefficients of feature extraction are irregular, memory access conflict would frequently occur without proper scheduling. The proposed scheme explicitly schedules the access sequence as a post-processing for managing the coefficient memory. By formulating the desired sequence as a graph model, the classical graph coloring theory can then be adopted to solve the scheduling problem. In addition, the proposed graph model also considers the resource constraint on intermediate storage. Experimental results show that the throughput and area-efficiency of the target cascaded classifier can be greatly improved by adopting the proposed scheme as compared to the related work. Hsiang-Chih Hsiao, Chun-Wei Chen, Jonas Wang, Ming-Der Shieh, Pei-Yin Chen |
DDECS | 4 |
| 2019 | Efficient Comparison and Swap on Fully Homomorphic Encrypted DataabstractFully homomorphic encryption (FHE) allows arbitrary computations to be performed directly on encrypted data for ensuring the security of cloud computing. In contrast to encrypting the plaintext in bit level as done in existing works, this paper explores how to reduce the computation complexity of encrypted data by adopting the concept of aggregate plaintext and proposes an efficient scheme to handle the comparison and swap operation, which is commonly used for sorting and searching in cloud computing. Experimental results reveal that employing the proposed scheme can not only reduce the size of required FHE data, but also improve the total computation time of the chosen operation. For 32-bit data comparison, the proposed one can operate 2.3 times faster and achieve about 52 times reduction in the required FHE data size as well as the transmission bandwidth to the cloud in comparison to the related work. Jyun-Neng Ji, Ming-Der Shieh |
ISCAS | 2 |
| 2018 | Fast Keyframe Selection and Switching for ICP-based Camera Pose EstimationabstractSimultaneous localization and mapping algorithms are important for high-quality registration used in augmented reality applications. Keyframe based SLAM can effectively reduce local drift by aligning a frame to the corresponding keyframe; however, it still suffers from losing trace for frames far from the keyframe. This work presents a fast keyframe selection and switching algorithm to replace unsuitable keyframes with qualified backup frames. The overhead of using backup process is greatly reduced by only inspecting the inlier information produced at the first iteration of iterative closest point (ICP) algorithm. Moreover, several useful criteria considering spatial and/or temporal relationships are also presented to evaluate the quality of keyframes and backup frames. Experimental results show that about 11.37% of relative pose error and 16.79% of the I CP iterations can be reduced by applying the proposed schemes as compared to the traditional keyframe-based approach. The reduction in computational time is achieved by speeding up the convergence of ICP, which is an additional benefit from applying the proposed schemes. Chun-Wei Chen, Wen-Yuan Hsiao, Jonas Wang, Ming-Der Shieh |
ISCAS | 5 |
| 2018 | Minimizing ESOP Expressions for Fully Homomorphic EncryptionabstractWith the rapid growth of cloud computing services, fully homomorphic encryption (FHE) has attracted much attention because a homomorphic evaluation can be directly performed on ciphertexts to ensure data privacy. This work explores the constraints and an associated minimization algorithm for exclusive sum-of-products (ESOP) expressions based on a homomorphism map and features of FHE. An efficient ESOP minimization algorithm based on the Sierpinski gasket and the triangle rule is proposed to reduce the maximum degree of the ESOP expression, thus relaxing the need to perform the recryption operation on ciphertexts. Experimental results show that about 23% of all candidate 4-variable functions can be further improved by the proposed algorithm to minimize the maximum degree of ESOP as compared with the results obtained using the traditional minimization algorithm. Jheng-Hao Ye, Si-Quan Chen, Ming-Der Shieh |
ISCAS | 3 |
| 2018 | Low-Complexity VLSI Design of Large Integer Multipliers for Fully Homomorphic Encryption
Jheng-Hao Ye, Ming-Der Shieh |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Fast model searching and combining for example learning-based super-resolutionabstractSingle-image super-resolution is an important technique for high resolution display related applications. Example learning-based approaches can provide plenty of image details by using trained dataset. The regression based methods reduce the memory storage size by training mapping functions rather than using a huge dictionary. However, the speed of searching the nearest cluster for the desired mapping function is still the bottleneck of the system. This problem is getting critical when the number of mapping functions is increased. This work presents an operator denoted as local multi-gradient level pattern to fast yet effectively describe the patch local geometry for a cluster of patches. The corresponding cluster can then be quickly identified by a simple lookup table. Furthermore, the potential cluster misclassification problem, induced by adopting the simplified clustering feature, is relaxed by applying the proposed model combining scheme. Simulation results show that the proposed one can achieve about 8 times speedup with even higher SSIM as compared to the related k-mean based method. Chun-Wei Chen, Fang-Kai Hsu, Der-Wei Yang, Jonas Wang, Ming-Der Shieh |
ISCAS | 5 |
| 2015 | High-quality texture compression using adaptive color grouping and selection algorithmabstractTexture compression is an important technique to reduce memory storage and increase rendering speed in graphic processing units (GPUs). To ensure fast decoding and retain random memory access characteristics, the industry standard DXTC compression finds an approximate line segment in color space for each 4×4 block. The idea has been extended to multiline solutions by segmenting the 4×4 block into several partitions and adopted by the new format BPTC. BPTC can provide the highest quality than other texture compression standards. However, the limited partition combinations used in BPTC would smooth the texture detail in some cases. This work introduces an adaptive color grouping and selection algorithm to relax the texture smoothing issue. Simulation results show that the proposed algorithm can improve the average PSNR by 0.78 dB for those blocks. Chun-Wei Chen, Ching-Heng Su, Der-Wei Yang, Jonas Wang, Chia-Cheng Lo, Ming-Der Shieh |
ISCAS | 6 |
| 2015 | Blind Channel Estimation for CP/CP-Free OFDM Systems Using Subspace ApproachabstractThe existing subspace-based channel estimation approach for orthogonal frequency division multiplexing (OFDM) systems with an expanding factor, called the repetition index, suffers from the problem of low probability of full row rank of the signal matrix with few OFDM blocks, which may fail to estimate the channel impulse response (CIR). In this paper, a subspace-based channel estimation algorithm with a two-step signal matrix construction method is proposed. The probability of full row rank of the proposed signal matrix is very close to one even if few OFDM blocks are available. The proposed approach can also be applied to both cyclic prefix (CP)-based and non-CP-based OFDM systems whereas the conventional approach is only suitable for CP-based systems. Simulation results show that the proposed method outperforms related blind channel estimation approaches in terms of normalized mean square error (NMSE). Shih-Hao Fang, Ju-Ya Chen, Jing-Shiun Lin, Ming-Der Shieh, Jen-Yuan Hsu |
VTC Spring | 4 |
| 2015 | Depth-Reliability-Based Stereo-Matching Algorithm and Its VLSI Architecture DesignabstractA low-complexity depth-reliability-based stereomatching algorithm and an efficient scanline memory-merging implementation scheme are proposed in this paper. The developed algorithm analyzes the accuracy of disparity results by using simple local window-based methods and preserves reliable information only. A bidirectional depth propagation flow is then adopted to fill the unreliable segments by using reliable information. Moreover, a set of predefined function-specific reliability variables are extracted to further improve depth quality in the occluded and smooth regions, which can reduce 39% bad pixels obtained by applying the basic 7 × 7 window-based matching. The proposed scanline memory-merging scheme along with data prefetching can lead to 32.7% savings on the scanline memory area and relax the requirements of external frame buffer size and bandwidth. Experimental results show that the implemented stereo-matching hardware has a gate count of 223 k including the scanline memory, and can achieve up to 70 frames/s for 480 × 540 resolution (2 × 2 downsampling of FullHD side-by-side 3-D format) with 56 disparity levels. Der-Wei Yang, Li-Chia Chu, Chun-Wei Chen, Jonas Wang, Ming-Der Shieh |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Low-Complexity High-Throughput QR Decomposition Design for MIMO SystemsabstractQR decomposition is a fundamental operation widely used in various signal detection schemes for multiple-input multiple-output (MIMO) systems. In this brief, a high-throughput converted form QR factorization (QRF) is investigated. The proposed two-phase computing scheme starts with a direct form factorization followed by a postprocessing using simple row/column permutations. Using coordinate rotation digital computer (CORDIC) algorithm, a massively parallel array architecture consisting of pipelined and folded CORDIC modules is developed to enhance the throughput. In addition, chaining mode operation is supported, where five successive signal vector updates can be performed the following every factorization. The chip implementation result in TSMC 0.18-μm CMOS process indicates the design, with an equivalent gate count 192100, can operate at 200 MHz and accomplish 25-M complex-valued QRF per second. This suggests a highest 3-Gb/s data rate for signal detections in a 4 × 4 MIMO system. The proposed design also outperforms other designs in two compound performance indices, i.e., data rate normalized with respect to gate count and power consumption, respectively. Jing-Shiun Lin, Yin-Tsung Hwang, Shih-Hao Fang, Po-Han Chu, Ming-Der Shieh |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | Efficient Memory-Addressing Algorithms for FFT Processor DesignabstractThis paper explores efficient memory management schemes for memory-based architectures of the fast Fourier transform (FFT). A data relocation scheme that merges multiple banks to lower the area requirement and power dissipation of memory-based FFT architectures is proposed. The proposed memory-addressing method can effectively deal with single-port, merged-bank memory with high-radix processing elements. Compared with conventional memory-based FFT designs using dual-port memory, the derived architecture has better performance in terms of area and power consumption. The proposed scheme is extended to a cached-memory FFT architecture to further reduce power dissipation. An 8192-point cached-memory FFT processor is implemented for digital video broadcasting-terrestrial/handheld applications by using 0.18-μm 1P6M CMOS technology. Experimental results show that the proposed memory scheme consumes 10.1%-29.3% less area and 9.6%-67.9% less power compared with those of the multibank design. Hsin-Fu Luo, Yi-Jun Liu, Ming-Der Shieh |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | An efficient countermeasure against power attacks for ECC over GF(p)abstractPower attacks are serious threats to cryptographic devices, and most countermeasures against power attacks result in a large time overhead for hardware implementation. This work presents an efficient countermeasure against power attacks for elliptic curve cryptography over GF(p). The proposed algorithm adopts the Montgomery ladder scalar multiplication algorithm as a basic framework to protect SPA. Then, a new scheme is presented to effectively manipulate the key so as to reduce the resulting time overhead for preventing differential power attack (DPA) and zero power attack (ZPA). Particularly, the base point blinding technique and half key splitting scheme are used to protect the upper and the lower halves of the key, respectively. Experimental results show the proposed countermeasure exhibit a time advantage over related works. Compared to other countermeasures against SPA, DPA, and ZPA, the proposed one can achieve up to 15% time improvement for accomplishing one 160-bit GF(p) scalar multiplication. Jheng-Hao Ye, Szu-Han Huang, Ming-Der Shieh |
ISCAS | 3 |
| 2014 | Subspace-Based Blind Channel Estimation for MIMO-OFDM Systems with New Signal Permutation MethodabstractSubspace-based blind channel estimation for multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) systems with repetition index has been proposed in the literature. However, the full-row-rank probability of the signal matrix using the MIMO-based repetition index method is extremely low, especially when low-order modulation or few received OFDM symbols are used. In this paper, the MIMO-based repetition index method with a new signal permutation method is proposed. The proposed method solves the problem of previous MIMO-based repetition index method. Additionally, the proposed method performs well even if low-order modulation or few revived OFDM symbols are applied. Simulation results show that proposed method not only has high full-row-rank probability, but also outperforms conventional blind channel estimation approaches in terms of normalized mean-square error (NMSE). Shih-Hao Fang, Ju-Ya Chen, Jing-Shiun Lin, Ming-Der Shieh, Jen-Yuan Hsu |
VTC Spring | 4 |
| 2014 | Scalable Montgomery Modular Multiplication Architecture with Low-Latency and Low-Memory Bandwidth RequirementabstractMontgomery modular multiplication is widely used in public-key cryptosystems. This work shows how to relax the data dependency in conventional word-based algorithms to maximize the possibility of reusing the current words of variables. With the greatly relaxed data dependency, we then proposed a novel scheduling scheme to alleviate the number of memory access in the developed scalable architecture. Analytical results show that the memory bandwidth requirement of the proposed scalable architecture is almost $(1/(w-1))$ times that of conventional scalable architectures, where $(w)$ denotes word size. The proposed one also retains a latency of exactly one cycle between the operations of the same words in two consecutive iterations of the Montgomery modular multiplication algorithm when employing enough processing elements. Compared to the design in the related work, experimental results demonstrate that the proposed one achieves an almost 54 percent reduction in power consumption with no degradation in throughput. The reduced number of memory access not only leads to lower power consumption, but also facilitates the design of scalable architectures for any precision of operands. Wen-Ching Lin, Jheng-Hao Ye, Ming-Der Shieh |
IEEE Trans. Computers | 3 |
| 2013 | Subspace-Based Blind Channel Estimation for MIMO-OFDM Systems with Repetition IndexabstractA subspace-based blind channel estimation algorithm for multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) systems is proposed in this paper. The proposed algorithm for MIMO-OFDM systems generalizes the conventional repetition index approach, which is only suitable for single-input single-output (SISO)-OFDM systems. Additionally, the necessary conditions of the proposed algorithm to be full row rank are discussed. Simulation results show that the probability of full row rank of the proposed signal matrix is close to one for different constellation schemes. The proposed generalized approach also outperforms the conventional approach with virtual carriers (VCs) in terms of normalized mean-squared error (NMSE) under static channel environments even if the number of OFDM symbols is small. Shih-Hao Fang, Ju-Ya Chen, Jing-Shiun Lin, Ming-Der Shieh, Dung-Rung Hsieh, Jen-Yuan Hsu |
VTC Fall | 4 |
| 2013 | Low-complexity multi-standard variable length coding decoder using tree-based partition and classificationabstractMPEG‐2 and H.264/AVC use variable length coding (VLC) to remove statistical redundancy. Representing the codeword table efficiently is thus an important issue to reduce hardware complexity, especially for multi‐standard applications. In this study, the VLC tree is decomposed using sub‐tree classification. The proposed algorithm reduces the amount of storage required for codewords. The proposed MPEG‐2/H.264 VLC decoder has 11.9 K gates when synthesised to operate at 180 MHz. The gate count is 20% lower than the sum of gate count of individual H.264 and MPEG‐2 VLC decoders. Chia-Cheng Lo, Chia-Wei Hsu, Ming-Der Shieh |
IET Image Process. | 3 |
| 2012 | An efficient QR decomposition design for MIMO systemsabstractMultiple-input multiple-output (MIMO) techniques have been widely used in various wireless communication systems these days. QR factorization is a fundamental module yet computationally intensive used in many MIMO detection schemes. In this paper, a complex-valued QR factorization (CQRF) scheme realized via a sequence of real-value Givens rotations is first presented. An efficient CQRF design using coordinate rotation digital computer (CORDIC) modules is next developed. The design features a highly parallel architecture to support high throughput operations. One CQRF can be obtained in every 8 clock cycles. To reduce the circuit complexity, a pipelined CORDIC structure is also applied. The implementation results in TSMC 0.18-µm CMOS process indicate that the proposed design can achieve a throughput rate of 25MCQRFs per second while consuming only 103.7k gates in circuit complexity. Performance evaluation based on a composite index consisting of area and throughput rate also shows the advantages of the proposed design against other similar works. Jing-Shiun Lin, Yin-Tsung Hwang, Po-Han Chu, Ming-Der Shieh, Shih-Hao Fang |
ISCAS | 4 |
| 2012 | Efficient scissoring scheme for scanline-based rendering of 2D vector graphicsabstractThis work presents a look-up table-based (LUT-based) algorithm for scanline-based rendering of OpenVG. The proposed method can deal with arbitrary number of scissoring rectangles. The rasterization and scissoring in the proposed architecture can be performed concurrently to reduce rendering time. The scanline-size buffers used as scissoring LUTs result in low area overhead. Moreover, a linked list structure of scissoring rectangles is proposed in order that only the scissoring rectangles interacted with the processing scanline are accessed to increase bus efficiency and reduce power consumption. Implementation results based on TSMC 0.13-μm CMOS technology show that the proposed rasterization design with LUT-based scissoring can operate at 200 MHz with 77K gate counts. The proposed design can render 16.8 tiger images with 392×483 resolution per second assuming ideal bus latency. Compared to existing works, the proposed design achieves a smaller area and more functionality for higher display resolution with comparable throughput. Wen-Ching Lin, Jheng-Hao Ye, Der-Wei Yang, Si-Yu Huang, Ming-Der Shieh, Jonas Wang |
ISCAS | 5 |
| 2012 | Fast scalable radix-4 Montgomery modular multiplierabstractMontgomery modular multiplication is widely applied to public key cryptosystems like Rivest-Sharmir-Adleman (RSA) and elliptic curve cryptography (ECC). This work presents a word-based Booth encoded radix-4 Montgomery modular multiplication algorithm for low-latency scalable architecture. The data dependency resulting from the inherent right shifting of the intermediate results in the conventional radix-4 Montgomery modular multiplication algorithm is alleviated; thus the latency between the neighboring process elements (PEs) is exactly one cycle. The number of the equivalent operands in the accumulation is not increased with operand reduction scheme. Implementation results based on the same technology show that compared to other Booth encoded radix-4 Montgomery modular multipliers, the proposed design achieves at least 23% time reduction for accomplishing one 1024-bit Montgomery modular multiplication. Sheng-Hong Wang, Wen-Ching Lin, Jheng-Hao Ye, Ming-Der Shieh |
ISCAS | 4 |
| 2012 | Blind channel estimation for cyclic prefix-free orthogonal frequency-division multiplexing systems with particular input symbolsabstractSubspace algorithms are commonly used for blind channel estimation in orthogonal frequency-division multiplexing (OFDM) systems as they can establish the relation between the channel impulse response and singular vectors of the noise subspace. However, they are not workable under the conventional cyclic prefix-free OFDM signal model as the corresponding necessary condition of subspace algorithms is not satisfied. To solve this problem, this study proposes two subspace-based blind channel estimation approaches using the real and periodic symbols of time-domain transmitted OFDM blocks. From the simulation results, we can obtain the proposed approaches outperform conventional ones in terms of normalised mean-squared error and bit error rate under time-invariant channel environments. Shih-Hao Fang, Ju-Ya Chen, Ming-Der Shieh, Jing-Shiun Lin |
IET Commun. | 3 |
| 2010 | Design of high-speed bit-serial divider in GF(2m)abstractIn this paper, we reformulated the conventional iterative division algorithm by substituting the pre-defined variable and then updating its initial value accordingly. The reformulated division algorithm allows a restructuring of the divider architecture to further improve its operating speed without increasing latency and area cost. Using the proposed fast algorithm, we developed a high-speed bit-serial GF(2m) divider. Analytical results show that the cost of the initial value update and variable transformation in the reformulated algorithm is almost negligible in the hardware implementation. Our divider reduces the critical path delay. Compared with related divider designs, the proposed design has time and area advantages. Wen-Ching Lin, Ming-Der Shieh, Chien-Ming Wu |
ISCAS | 2 |
| 2010 | Low-complexity Reed-Solomon decoder for optical communicationsabstractThis paper presents a low-complexity Reed-Solomon (RS) decoder design based on the modified Euclidean (ME) algorithm proposed by Truong. The low-complexity feature is achieved by first reformulating Truong's ME algorithm using the proposed polynomial manipulation scheme so that a more compact polynomial representation can be derived. Together with the developed folding scheme and the simplified boundary cell, the resulting design can effectively reduce the hardware complexity and meet the throughput requirement of optical communication systems. Compared with the related works, our development not only provides the minimum area requirement but also has the smallest area-time complexity. Experimental results demonstrate that the developed RS(255, 239) decoder, implemented in TSMC 0.18 μm process, can operate up to 430 MHz and achieve a throughput rate of 3.44 Gbps with a total gate count of 11,763. Yung-Kuei Lu, Ming-Der Shieh, Chien-Ming Wu |
ISCAS | 2 |
| 2010 | Efficient memory management for FFT processorsabstractThis paper presents an efficient memory management scheme for memory-based architecture of Fast Fourier Transform (FFT). A new data relocating scheme is proposed to merge multiple banks for alleviating the area requirement as well as the power dissipation of memory-based FFT. The proposed memory addressing method can effectively deal with merged-banked single-ported memory with high-radix processing elements. Compared to conventional memory-based FFT designs using two-ported memories, the derived architecture reveals better performance in area requirement and power consumption. Hsin-Fu Luo, Ming-Der Shieh, Yi-Jun Liu, Chien-Ming Wu |
ISCAS | 2 |
| 2010 | Subspace-Based Blind Channel Estimation for OFDM Systems with Conjugate-Symmetric PropertyabstractA subspace-based blind channel estimation algorithm for OFDM systems with the conjugate-symmetric property is proposed in this paper. In subspace-based algorithms, construction of the signal or noise subspace is an important issue. Different to other subspace-based approaches exploiting the signal or noise subspace from cyclic prefix, virtual subcarriers, or oversampling techniques, we develop a blind channel estimation algorithm by establishing noise subspace from time-domain real symbols, which can be obtained by using the conjugate-symmetric property of OFDM symbols in frequency-domain. Simulation results demonstrate that the proposed algorithm outperforms conventional subspace-based channel estimation algorithms in mean-squared error. Shih-Hao Fang, Ju-Ya Chen, Ming-Der Shieh, Jing-Shiun Lin |
VTC Spring | 3 |
| 2010 | Word-Based Montgomery Modular Multiplication Algorithm for Low-Latency Scalable ArchitecturesabstractModular multiplication is a crucial operation in public key cryptosystems like RSA and elliptic curve cryptography (ECC). This paper presents a new word-based Montgomery modular multiplication algorithm which can be used to achieve a low-latency scalable architecture for efficient hardware implementations. We show how to relax the data dependency in conventional word-based algorithms so that a latency of exactly one cycle can be obtained regardless of the chosen word size w (w > 1). With the presented operand reduction scheme, the proposed scalable architecture can operate at high speeds and suitable data paths can be chosen for specific applications. Complexity analysis shows that the proposed architecture has the lowest latency and area complexity compared to related scalable architectures. Experimental results demonstrate that our design has area, speed, and flexibility advantages over related schemes. Ming-Der Shieh, Wen-Ching Lin |
IEEE Trans. Computers | 1 |
| 2010 | A High-Performance Unified-Field Reconfigurable Cryptographic ProcessorabstractWith rapid increases in communication and network applications, cryptography has become a crucial issue to ensure the security of transmitted data. In this paper, we propose a microcode-based architecture with a novel reconfigurable datapath which can perform either prime field GF(p) operations or binary extension field GF(2m) operations for arbitrary prime numbers, irreducible polynomials, and precision. Using these field arithmetic units, users are capable of programming cryptographic algorithms in microcode sequences for full compliance with a majority of public-key cryptographic algorithms such as Rivest-Shamir-Adleman (RSA) and elliptic curve cryptosystems. An algorithmic optimization or refinement can thus be made at a higher level based on the reconfigurable datapath. Experimental results show that the developed processor has full cryptography algorithm flexibility, high hardware utilization, and high performance. Jun-Hong Chen, Ming-Der Shieh, Wen-Ching Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Hardware/Software Codesign of Resource Constrained Real-Time SystemsabstractSystem-level design methods can provide a systematic and effective way of evaluating various design options, thus shortening the product development time. This paper relaxes the HC algorithm by considering the K best candidates in each clustering iteration to alleviate the possibility of being trapped in local minimum during hardware/software (HW/SW) partition. We also present an architecture mapping algorithm together with the defined sensitivity measure to further reduce the hardware requirement. Simulation results show that the complexity of exploration time can be greatly reduced with only little performance loss as compared to the exhaustive search. The proposed algorithm can thus provide a good compromise between exploration time and accuracy. Chia-Cheng Lo, Jung-Guan Luo, Ming-Der Shieh |
IAS | 3 |
| 2009 | Efficient Software-Based Self-Test Methods for Embedded Digital Signal ProcessorsabstractEmbedded processors are ubiquitous in today's system-on-chip design. In addition to designing digital signal processors (DSPs) for various applications, developing efficient test methods with little overhead and desired fault coverage for DSPs are also crucial and practical. Compared with the scan-based test methods, the software-based self-test (SBST) method does not suffer from area overhead and performance degradation, and can provide at-speed test for DSPs with the potential drawbacks of lower fault coverage and a larger amount of test vectors. This paper explores techniques to improve the fault coverage of SBST methods for the developed DSP core with instructions fully compatible with those of the TI TMS320C54x. Experimental results exhibit that applying the developed SBST test flow obtains more than 96% fault coverage for our DSP core, which is higher than the reported values in related works. Jun-Jie Zhu, Wen-Ching Lin, Jheng-Hao Ye, Ming-Der Shieh |
Asian Test Symposium | 4 |
| 2009 | A Generalized Blind Channel Estimation Algorithm for OFDM Systems with Cyclic PrefixabstractA subspace-based blind channel estimation algorithm with cyclic prefix is proposed in this paper. A systematic approach is used to construct a new signal matrix in the proposed algorithm. Compared with conventional blind channel estimation algorithm, the proposed algorithm has lower computational complexity and higher probability of full row rank for the corresponding signal matrix. In addition, fewer OFDM symbols can be used to satisfy the necessary condition for achieving a full-row-rank signal matrix. Simulation results show that the proposed algorithm outperforms conventional methods in mean-squared error and bit error rate under static channel. Even with a smaller number of received OFDM symbols, the proposed algorithm can perform well. Shih-Hao Fang, Ju-Ya Chen, Ming-Der Shieh, Jing-Shiun Lin |
ISCAS | 3 |
| 2009 | Flexible GF(2m) Divider Design for Cryptographic ApplicationsabstractIn cryptographic applications, private key algorithms usually aim at high-throughput data communication, while public key algorithms require much lower throughput for private key exchange and authentication. To increase hardware utilization and reduce area overhead, this paper presents a flexible divider design in GF(2m), which can be configured to operate in either SIMD or SISD mode. When applied to SIMD applications, the divider can perform multiple divisions in parallel and output results per cycle; thus, it is suitable for AES cryptosystems demanding high throughput. In SISD applications, the divider is scalable and can handle different sizes of operand such as those specified in ECC standards. A scalable design can also relax the potential problem of high fanout control signals. Complexity analysis shows the proposed divider, operated in SIMD mode, has lower area complexity and higher throughput in comparison with related work. Wen-Ching Lin, Ming-Der Shieh, Chien-Ming Wu |
ISCAS | 2 |
| 2009 | Modified Subspace Based Channel Estimation Algorithm for OFDM SystemsabstractA modified subspace-based blind channel estimation algorithm for OFDM systems is proposed in this paper. A systematic approach is used to construct the signal matrix in the proposed algorithm. Compared with conventional subspace- based blind channel estimation algorithm, the proposed algorithm has lower computational complexity and higher probability of full row rank for the corresponding signal matrix. In addition, fewer OFDM symbols can be used to satisfy the necessary condition for achieving a full-row-rank signal matrix. Simulation results show that the proposed algorithm outperforms conventional method in mean-squared error and bit error rate under static channel. Even with a smaller number of received OFDM symbols, the proposed algorithm can perform well. Shih-Hao Fang, Ju-Ya Chen, Ming-Der Shieh, Jing-Shiun Lin |
VTC Spring | 3 |
| 2008 | High-speed modular multiplication design for public-key cryptosystemsabstractModular exponentiation for public-key cryptosystems is usually accomplished by repeated modular multiplications on large integers. A high-speed design of modular multiplication is thus very crucial to speed up the decryption/encryption process. In this paper, we first explore how to relax the data dependency existing among the multiplication, quotient determination, and modular reduction in conventional Montgomery modular multiplication algorithm. Then we proposed a new modular reduction algorithm with a smaller critical path delay in hardware implementation. The speed improvement is achieved by reducing the critical path delay from the 4-to-2 to 3-to-2 carry-save addition, and the resulting time complexity of our development is decreased by simultaneously performing the multiplication and modular reduction processes. Experimental results show that our modular exponentiation can obtain both time and area-time (AT) advantages compared with existing work. Jun-Hong Chen, Wen-Ching Lin, Hao-Hsuan Wu, Ming-Der Shieh |
ISCAS | 4 |
| 2008 | A new look-up table-based multiplier/squarer design for cryptosystems over GF(2m)abstractThis paper presents a high-speed multiplier/squarer design over finite field GF(2m) for large m. We extended the look-up table (LUT) based multiplication algorithm introduced by Hasan to reduce the LUT generation time and then showed how to effectively add the squaring operation to the developed multiplier. The unified multiplication/squaring module is very suitable for applications like Elliptic Curve Cryptography (ECC) in which these two types of operations are operated alternately. Experimental results exhibit that using the proposed sub-group, multiple look-up tables (SG-MLUT) based scheme, up to 29% improvement in the total computation time of multiplication can be achieved in comparison with that using Hasan’s algorithm. When employing the unified multiplier/squarer module instead of Hasan’s design in ECC applications, we can gain further improvement in the scalar multiplication time because no LUT generation is needed using our design, and obtain about 24.5% reduction on the resulting area-time (AT) complexity. Wen-Ching Lin, Jun-Hong Chen, Ming-Der Shieh |
ISCAS | 3 |
| 2008 | A New Modular Exponentiation Architecture for Efficient Design of RSA CryptosystemabstractModular exponentiation with a large modulus, which is usually accomplished by repeated modular multiplications, has been widely used in public key cryptosystems for secured data communications. To speed up the computation, the Montgomery modular multiplication algorithm is used to relax the process of quotient determination, and the carry-save addition (CSA) is employed to reduce the critical path delay. In this paper, based on the inherent data dependency between the modular multiplication and square operations in the H-algorithm of modular exponentiation, we present a new modular exponentiation architecture with a unified modular multiplication/square module and show how to reduce the number of input operands for the CSA tree by mathematical manipulation. The developed architecture has the following advantages. 1) There is no need to convert the carry-save form of an operand into its binary representation at the end of each modular multiplication. In this way, except the final step to get the result of modular exponentiation, the time-consuming carry propagation can then be eliminated. 2) The number of input operands for the CSA tree is reduced in a very efficient way. 3) The hardware saving is achieved with very limited impact on the original critical path delay when designed with two distinct modular multiplication and square components. Experimental results show that our modular exponentiation design obtains the least hardware complexity compared with the existing work and outperforms them in terms of area-time (AT) complexity as well. Ming-Der Shieh, Jun-Hong Chen, Hao-Hsuan Wu, Wen-Ching Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | A New Montgomery Modular Multiplication Algorithm and its VLSI Design for RSA CryptosystemabstractModular exponentiation for RSA cryptosystem is usually accomplished by repeated modular multiplications on large integers, which is considerably time-consuming. To speed up the operation, the Montgomery modular multiplication algorithm is employed to eliminate the trial division, and the carry-save addition is used to alleviate the carry propagation delay. In this paper, we propose a unified Montgomery modular multiplication algorithm that can be applied to fulfil either the conventional modular multiplication or squaring operation in carry-save form so as to achieve area-efficient design of modular exponentiation. Meanwhile, we reduce the number of input operands for carry-save addition by mathematical manipulation to minimize the resulting critical path delay. Compared with the existing works, our modular exponentiation design obtains the least hardware complexity and outperforms them in terms of area-time (AT) complexity. Jun-Hong Chen, Haw-Shiuan Wu, Ming-Der Shieh, Wen-Ching Lin |
ISCAS | 3 |
| 2006 | High-speed CRC design for 10 Gbps applicationsabstractThe use of cyclic redundancy codes (CRCs) in many high-throughput applications has made the design of parallel CRC circuitry an important research topic. Parallel implementation of the linear feedback shift registers (LFSRs) requires multiple input bits being processed at the same time; therefore, is much faster than the serial implementation. The common way to process M input bits simultaneously is to multiply the companion metric M times and put the resulting circuit in the feedback loop. This, however, will increase the circuit complexity within the loop so as to limit the final speedup ratio. In this paper, based on the state-space transformation, we investigate how to design high-speed CRC circuitry for 10 Gbps applications. Our design can efficiently deal with the case that the length of the message bits is not a multiple of M and achieves low-cost solution by sharing the input block with the output block outside the feedback loop. Jing-Shiun Lin, Chung-Kung Lee, Ming-Der Shieh, Jun-Hong Chen |
ISCAS | 3 |
| 2006 | Design and implementation of efficient Reed-Solomon decoders for multi-mode applicationsabstractWe present a multi-mode Reed-Solomon decoder based on the reformulated inversionless Berlekamp-Massey algorithm, which can retain the throughput rate of the reformulated architecture in many practical applications. With the developed coefficient-selector-free multi-mode arrangement, the resulting design possesses not only area-efficient property but also very simple and regular interconnect topology that makes it very suitable for VLSI realization. Implementation results exhibit that the achievable throughput rate of the developed decoder for n/spl les/255 and 0/spl les/t/spl les/8, implemented in UMC 0.18/spl mu/m 1P6M process, is 3.2Gbps at the maximum clock rate of 400MHz and the total gate count is 22,931. Compared with the existing work based on extended Euclidean algorithm, our development provides both area and speed advantages and can be used for multi-standard applications. Ming-Der Shieh, Yung-Kuei Lu, Shen-Ming Chung, Jun-Hong Chen |
ISCAS | 1 |
| 2006 | Efficient path metric access for reducing interconnect overhead in Viterbi decodersabstractEfficient management of the path metric memory and minimization of interconnection networks between the memory and addcompareselect unit (ACSU) are always the key concerns on the design and implementation of Viterbi decoders. In this paper, we derive a set of simple equations to partition the memory into P banks such that the equivalent memory bandwidth can be increased with very simple interconnection networks. Compared with the previous work, our proposed approach reveals the following superiority: (1) Each memory bank can be treated as a local memory of a specific ACS; thus, the interconnection network is simplified. (2) The P memory banks can be merged into only two pseudo-banks regardless of the number of ACS operations. This not only further reduces the hardware requirements of address generation, but also makes smaller the required memory space. Ming-Der Shieh, Tai-Ping Wang, Chien-Ming Wu, Chun-Ming Huang |
ISCAS | 1 |
| 2005 | VLSI architectural design tradeoffs for sliding-window log-MAP decodersabstractTurbo codes have received tremendous attention and have commenced their practical applications due to their excellent error-correcting capability. Investigation of efficient iterative decoder realizations is of particular interest because the underlying soft-input soft-output decoding algorithms usually lead to highly complicated implementation. This paper describes the architectural design and analysis of sliding-window (SW) Log-MAP decoders in terms of a set of predetermined parameters. The derived mathematical representations can be applied to construct a variety of VLSI architectures for different applications. Based on our development, a SW-Log-MAP decoder complying with the specification of third-generation mobile radio systems is realized to demonstrate the performance tradeoffs among latency, average decoding rate, area/computation complexity, and memory power consumption. This paper thus provides useful and general information on practical implementation of SW-Log-MAP decoders. Chien-Ming Wu, Ming-Der Shieh, Chien-Hsing Wu 0002, Yin-Tsung Hwang, Jun-Hong Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | High-Speed, Low-Complexity Systolic Designs of Novel Iterative Division Algorithms in GF(2^m)abstractWe extend the binary algorithm invented by Stein and propose novel iterative division algorithms over GF(2/sup m/) for systolic VLSI realization. While algorithm EBg is a basic prototype with guaranteed convergence in at most 2m - 1 iterations, its variants, algorithms EBd and EBdf, are designed for reduced complexity and fixed critical path delay, respectively. We show that algorithms EBd and EBdf can be mapped to parallel-in parallel-out systolic circuits with low area-time complexities of O(m/sup 2/loglogm) and O(m/sup 2/), respectively. Compared to the systolic designs based on the extended Euclid's algorithm, our circuits exhibit significant speed and area advantages. Chien-Hsing Wu 0002, Chien-Ming Wu, Ming-Der Shieh, Yin-Tsung Hwang |
IEEE Trans. Computers | 3 |
| 2000 | High-speed generation of LFSR signaturesabstractWe investigate techniques for speeding up the compaction simulation of a single-input signature register based on its equivalent multiple-input implementation. Our approach is to systematically decompose the original input sequence into a set of subsequences based on the theory of finite field. High-speed signature computations can then be achieved by inputting those subsequences at the same time and employing the lookahead technique for those subsequences to speed up compaction. Both the internal-XOR and external-XOR LFSRs are implemented to demonstrate the flexibility of our development. Compared with the existing methods that were mainly developed for software programming, our results are suitable for both software and hardware implementation and have the potential of reducing the memory requirement of off-line determination of signatures. Ming-Der Shieh, Hsin-Fu Lo, Ming-Hwa Sheu |
Asian Test Symposium | 1 |
| 2000 | An efficient approach for in-place scheduling of path metric update in Viterbi decodersabstractThe in-place path metric updating is a well-known technique for efficiently dealing with the management of path metric memory in Viterbi decoders. In this paper, we present a simple but efficient technique to partition the path metric memory into 2/sup i/ banks and then distribute a set of path metrics into scheduled add compare select (ACS) units. Results show that applying the presented scheduling technique the equivalent memory bandwidth can be increased with limited hardware overhead. The resulting architecture has the following characteristics: (1) the interconnection overhead between ACS units and the memory bank structure can be significantly reduced, (2) the control circuit is regular and the implementation can be derived in a systematic way. Therefore, the architecture can be easily applied to handle the convolutional code with a long constraint length and it is suitable to be implemented in VLSI applications. Chien-Ming Wu, Ming-Der Shieh, Chien-Hsing Wu 0002, Ming-Hwa Sheu |
ISCAS | 2 |
| 1998 | Design of a High-Speed Square GeneratorabstractGiven a binary number N, the simplest way for evaluating its square N/sup 2/ is the use of ROM look-up tables. For example, the squares of 12-bit numbers can be stored in a ROM of (2/sup 12//spl times/24) bits, which takes an area of 3.5 mm/sup 2/ and an access time of 9.96 ns with 0.8 /spl mu/m CMOS process. However, the conventional ROM table approaches are limited only for small bit size applications due to the unmanageable increase of the ROM table size. A novel design of square generator circuit using a folding approach is presented for high speed performance applications. Results show that, with the same process, the proposed square generator circuit takes 12.27 ns to generate the squares of 40 bit numbers with an area of about 2.88 times that of the (2/sup 12//spl times/24) ROM, i.e., 10 mm/sup 2/ a design trade-off between speed and area. A nested structure is also presented to achieve a 103 bit square generator with a delay of 15.82 ns. The bit size can be further increased by adding more levels of the nested structure. The results are promising and thus the proposed approach is well suitable for large bit size and high speed applications. Chin-Long Wey, Ming-Der Shieh |
IEEE Trans. Computers | 2 |
| 1996 | A CAM-Based VLSI Architecture for Shared Buffer ATM Switch with Fuzzy Controlled Buffer ManagementabstractThis paper proposes a CAM-based shared buffer ATM switch-on-a-chip architecture that takes network-element internal congestion control into consideration. This internal congestion control includes selective cell discard, priority service scheduling, and fuzzy controlled buffer management. To provide "fair" access to the network resources for all users, the SMXQ buffer control scheme is adopted. The SMXQ scheme assumes a queue length threshold is chosen for the logical queue pertaining to each output port, and if the queue length exceeds the threshold the arriving cells are discarded. Selective cell discard is performed per port basis when the shared buffer is full or the queue length of a particular output port exceeds its threshold. For each output port, {CLP=1} cells will be discarded before any {CLP=0} cell is discarded. The set of chosen queue length thresholds are computed directly by an on-chip fuzzy congestion controller (FCC) in sub microsecond intervals. Chie Dou, Ming-Der Shieh |
ICCD | 2 |
| 1993 | ASLCScan: A Scan Design Technique for Asynchronous Sequential Logic CircuitsabstractAsynchronous sequential logic circuits (ASLCs) are synthesized with either the Huffman model, referred to as HMASLCs, or with the signal transition graph (STG), referred to as STGASLCs. Based on a single stuck-at fault model, this paper describes fault effects for both HMASLCs and STGASLCs and addresses the similarities and differences between them. The fault effects include redundant faults and state oscillations. Input/output redundancy is a special feature of STGASLCs which relaxes the fundamental mode in HMASLCs. Results of this study show that the faults due to the input/output concurrency cannot be tested without a scan structure. This paper presents a scan design technique. ASLCScan. With this structure, the test generation problem is reduced to one of just testing the combinational logic.> Chin-Long Wey, Ming-Der Shieh, P. David Fisher |
ICCD | 2 |