VLDB 2026 Research / reviewers in the wild / expert
Wonyong Sung
dblp:22/1975
· DBLP profile ↗
95ranked-venue papers
10as first author
9since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 4 first-author · 5 since 2021Systems, architecture and hardware · 24 · 5 first-authorArtificial intelligence and machine learning · 16 · 6 since 2021Computer networks · 3Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Teacher Intervention: Improving Convergence of Quantization Aware Training for Ultra-Low Precision TransformersabstractPre-trained Transformer models such as BERT have shown great success in a wide range of applications, but at the cost of substantial increases in model complexity.Quantizationaware training (QAT) is a promising method to lower the implementation cost and energy consumption.However, aggressive quantization below 2-bit causes considerable accuracy degradation due to unstable convergence, especially when the downstream dataset is not abundant.This work proposes a proactive knowledge distillation method called Teacher Intervention (TI) for fast converging QAT of ultralow precision pre-trained Transformers.TI intervenes layer-wise signal propagation with the intact signal from the teacher to remove the interference of propagated quantization errors, smoothing loss surface of QAT and expediting the convergence.Furthermore, we propose a gradual intervention mechanism to stabilize the recovery of subsections of Transformer layers from quantization.The proposed schemes enable fast convergence of QAT and improve the model accuracy regardless of the diverse characteristics of downstream fine-tuning tasks.We demonstrate that TI consistently achieves superior accuracy with significantly lower finetuning iterations on well-known Transformers of natural language processing as well as computer vision compared to the state-of-the-art QAT methods. Kyuhong Shim, Seongmin Park 0003, Wonyong Sung, Jungwook Choi |
EACL | 4 |
| 2023 | Enhancing Computation Efficiency in Large Language Models through Weight and Activation QuantizationabstractLarge Language Models (LLMs) are proficient in natural language processing tasks, but their deployment is often restricted by extensive parameter sizes and computational demands.This paper focuses on post-training quantization (PTQ) in LLMs, specifically 4-bit weight and 8-bit activation (W4A8) quantization, to enhance computational efficiency-a topic less explored compared to weight-only quantization.We present two innovative techniques: activation-quantization-aware scaling (AQAS) and sequence-length-aware calibration (SLAC) to enhance PTQ by considering the combined effects on weights and activations and aligning calibration sequence lengths to target tasks.Moreover, we introduce dINT, a hybrid data format combining integer and denormal representations, to address the underflow issue in W4A8 quantization, where small values are rounded to zero.Through rigorous evaluations of LLMs, including OPT and LLaMA, we demonstrate that our techniques significantly boost task accuracies to levels comparable with full-precision models.By developing arithmetic units compatible with dINT, we further confirm that our methods yield a 2× hardware efficiency improvement compared to 8-bit integer MAC unit. Janghwan Lee, Seungcheol Baek, Seok Joong Hwang, Wonyong Sung, Jungwook Choi |
EMNLP | 5 |
| 2023 | Token-Scaled Logit Distillation for Ternary Weight Generative Language ModelsabstractGenerative Language Models (GLMs) have shown impressive performance in tasks such as text generation, understanding, and reasoning. However, the large model size poses challenges for practical deployment. To solve this problem, Quantization-Aware Training (QAT) has become increasingly popular. However, current QAT methods for generative models have resulted in a noticeable loss of accuracy. To counteract this issue, we propose a novel knowledge distillation method specifically designed for GLMs. Our method, called token-scaled logit distillation, prevents overfitting and provides superior learning from the teacher model and ground truth. This research marks the first evaluation of ternary weight quantization-aware training of large-scale GLMs with less than 1.0 degradation in perplexity and achieves enhanced accuracy in tasks like common-sense QA and arithmetic reasoning as well as natural language understanding. Our code is available at https://github.com/aiha-lab/TSLD. Sihwa Lee, Janghwan Lee, Sukjin Hong, Du-Seong Chang, Wonyong Sung, Jungwook Choi |
NeurIPS | 6 |
| 2022 | Understanding the Role of Self Attention for Efficient Speech Recognition
Kyuhong Shim, Jungwook Choi, Wonyong Sung |
ICLR | 3 |
| 2022 | Similarity and Content-based Phonetic Self Attention for Speech Recognition
Kyuhong Shim, Wonyong Sung |
INTERSPEECH | 2 |
| 2022 | Macro-Block Dropout for Improved Regularization in Training End-to-End Speech Recognition ModelsabstractThis paper proposes a new regularization algorithm referred to as macro-block dropout. The overfitting issue has been a difficult problem in training large neural network models. The dropout technique has proven to be simple yet very effective for regularization by preventing complex co-adaptations during training. In our work, we define a macro-block that contains a large number of units from the input to a Recurrent Neural Network (RNN). Rather than applying dropout to each unit, we apply random dropout to each macro-block. This algorithm has the effect of applying different drop out rates for each layer even if we keep a constant average dropout rate, which has better regularization effects. In our experiments using Recurrent Neural Network-Transducer (RNN-T), this algorithm shows relatively 4.30 % and 6.13 % Word Error Rates (WERs) improvement over the conventional dropout on LibriSpeech test-clean and test-other. With an Attention-based Encoder-Decoder (AED) model, this algorithm shows relatively 4.36 % and 5.85 % WERs improvement over the conventional dropout on the same test sets. Chanwoo Kim 0001, Sathish Indurti, Jinhwan Park, Wonyong Sung |
SLT | 4 |
| 2021 | Stochastic Precision Ensemble: Self-Knowledge Distillation for Quantized Deep Neural NetworksabstractThe quantization of deep neural networks (QDNNs) has been actively studied for deployment in edge devices. Recent studies employ the knowledge distillation (KD) method to improve the performance of quantized networks. In this study, we propose stochastic precision ensemble training for QDNNs (SPEQ). SPEQ is a knowledge distillation training scheme; however, the teacher is formed by sharing the model parameters of the student network. We obtain the soft labels of the teacher by randomly changing the bit precision of the activation stochastically at each layer of the forward-pass computation. The student model is trained with these soft labels to reduce the activation quantization noise. The cosine similarity loss is employed, instead of the KL-divergence, for KD training. As the teacher model changes continuously by random bit-precision assignment, it exploits the effect of stochastic ensemble KD. SPEQ outperforms the existing quantization training methods in various tasks, such as image classification, question-answering, and transfer learning without the need for cumbersome teacher networks. Yoonho Boo, Sungho Shin, Jungwook Choi, Wonyong Sung |
AAAI | 4 |
| 2021 | SQWA: Stochastic Quantized Weight Averaging For Improving The Generalization Capability Of Low-Precision Deep Neural NetworksabstractLow-precision deep neural networks (DNNs) are very needed for efficient implementations, but severe quantization of weights often sacrifices the generalization capability and lowers the test accuracy. We present a new quantized neural network optimization approach, stochastic quantized weight averaging (SQWA), to design low-precision DNNs with good generalization capability using model averaging. The proposed approach includes (1) floating-point model training, (2) direct quantization of weights, (3) capturing multiple low precision models during retraining with cyclical learning rates, (4) averaging the captured models, and (5) re-quantizing the averaged model and fine-tuning it with low-learning rates. With SQWA training, we could develop the best performing QDNNs for image classification on ImageNet datasets and also for semantic segmentation on Pascal VOC 2012 dataset. Sungho Shin, Yoonho Boo, Wonyong Sung |
ICASSP | 3 |
| 2021 | Convolution-Based Attention Model With Positional Encoding For Streaming Speech Recognition On Embedded DevicesabstractOn-device automatic speech recognition (ASR) is much more preferred over server-based implementations owing to its low latency and privacy protection. Many server-based ASRs employ recurrent neural networks (RNNs) to exploit their ability to recognize long sequences with a limited number of states; however, they are inefficient for single-stream implementations in embedded devices. In this study, a highly efficient convolutional model-based ASR with monotonic chunkwise attention is developed. Although temporal convolution-based models allow more efficient implementations, they demand a long filter-length to avoid looping or skipping problems. To remedy this problem, we add positional encoding, while shortening the filter length, to a convolution-based ASR encoder. It is demonstrated that the accuracy of the short filter-length convolutional model is significantly improved. In addition, the effect of positional encoding is analyzed by visualizing the attention energy and encoder outputs. The proposed model achieves the word error rate of 11.20% on TED-LIUMv2 for an end-to-end speech recognition task. Jinhwan Park, Chanwoo Kim 0001, Wonyong Sung |
SLT | 3 |
| 2020 | HLHLp: Quantized Neural Networks Training for Reaching Flat Minima in Loss SurfaceabstractQuantization of deep neural networks is extremely essential for efficient implementations. Low-precision networks are typically designed to represent original floating-point counterparts with high fidelity, and several elaborate quantization algorithms have been developed. We propose a novel training scheme for quantized neural networks to reach flat minima in the loss surface with the aid of quantization noise. The proposed training scheme employs high-low-high-low precision in an alternating manner for network training. The learning rate is also abruptly changed at each stage for coarse- or fine-tuning. With the proposed training technique, we show quite good performance improvements for convolutional neural networks when compared to the previous fine-tuning based quantization scheme. We achieve the state-of-the-art results for recurrent neural network based language modeling with 2-bit weight and activation. Sungho Shin, Jinhwan Park, Yoonho Boo, Wonyong Sung |
AAAI | 4 |
| 2020 | Fixed-Point Optimization of Transformer Neural NetworkabstractThe Transformer model adopts a self-attention structure and shows very good performance in various natural language processing tasks. However, it is difficult to implement the Transformer in embedded systems because of its very large model size. In this study, we quantize the parameters and hidden signals of the Transformer for complexity reduction. Not only matrices for weights and embedding but the input and the softmax outputs are also quantized to utilize low-precision matrix multiplication. The fixed-point optimization steps consist of quantization sensitivity analysis, hardware conscious word-length assignment, quantization and retraining, and post-training for improved generalization. We achieved 27.51 BLEU score on the WMT English-to-German translation task with 4-bit weights and 6-bit hidden signals. Yoonho Boo, Wonyong Sung |
ICASSP | 2 |
| 2020 | Low-Latency Lightweight Streaming Speech Recognition with 8-Bit Quantized Simple Gated Convolutional Neural NetworksabstractAutomatic speech recognition (ASR) is very important for mobile devices. However, deep neural network-based ASR demands a large number of computations, while the memory bandwidth and battery capacity of mobile devices are limited. Server-based implementations are mostly employed, but this increases latency or privacy concerns. Efficient on-device ASR is the solution for these issues. In this paper, we propose a low-latency on-device speech recognition system with a simple gated convolutional network (SGCN). The SGCN shows a competitive recognition accuracy even with 1M parameters. In addition, SGCN is advantageous for parallelization which enables efficient cache utilization. 8-bit quantization is applied to reduce the memory size and computation time. The proposed system features online recognition fulfilling the 0.4s latency limit and operates with the real-time factor of 0.2 using only a single 900MHz CPU core. The system occupying 1.2MB memory footprint shows 19.75% word error rate (WER) with greedy decoding. Jinhwan Park, Xue Qian, Youngmin Jo, Wonyong Sung |
ICASSP | 4 |
| 2020 | Effect of Adding Positional Information on Convolutional Neural Networks for End-to-End Speech Recognition
Jinhwan Park, Wonyong Sung |
INTERSPEECH | 2 |
| 2019 | Simple Gated Convnet for Small Footprint Acoustic ModelingabstractAcoustic modeling with recurrent neural networks has shown very good performance, especially for end-to-end speech recognition. However, most recurrent neural networks require sequential computation of the output, which results in large memory access overhead when implemented in embedded devices. Convolution-based sequential modeling does not suffer from this problem; however, the model usually requires a large number of parameters. We propose simple gated convolutional neural networks (Simple Gated ConvNet) for acoustic modeling and show that the network performs very well even when the number of parameters is fairly small, less than 3 million. The Simple Gated ConvNet (SGCN) is constructed by combining the simplest form of Gated ConvNet and one-dimensional (1-D) depthwise convolution. The model has been evaluated using the Wall Street Journal (WSJ) Corpus and has shown a performance competitive to RNN-based ones. The performance of the SGCN has also been evaluated using the LibriSpeech Corpus. The developed model was implemented in ARM CPU based systems and showed the real time factor (RTF) of around 0.05. Lukas Lee, Jinhwan Park, Wonyong Sung |
ASRU | 3 |
| 2019 | Memorization Capacity of Deep Neural Networks under Parameter QuantizationabstractMost deep neural networks (DNNs) require complex models to achieve high performance. Parameter quantization is widely used for reducing the implementation complexities. Previous studies on quantization were mostly based on extensive simulation using training data on a specific model. We choose a different approach and attempt to measure the per-parameter capacity of DNN models and interpret the results to obtain insights on optimum quantization of parameters. This research uses artificially generated data and generic forms of fully connected DNNs, convolutional neural networks, and recurrent neural networks. We conduct memorization and classification tests to study the effects of the number and precision of the parameters on the performance. The model and the per-parameter capacities are assessed by measuring the mutual information between the input and the classified output. To get insight for parameter quantization when performing real tasks, the training and test performances are compared. Yoonho Boo, Sungho Shin, Wonyong Sung |
ICASSP | 3 |
| 2019 | Workload-aware Automatic Parallelization for Multi-GPU DNN TrainingabstractDeep neural networks (DNNs) have emerged as successful solutions for variety of artificial intelligence applications, but their very large and deep models impose high computational requirements during training. Multi-GPU parallelization is a popular option to accelerate demanding computations in DNN training, but most state-of-the-art multi-GPU deep learning frameworks not only require users to have an in-depth understanding of the implementation of the frameworks themselves, but also apply parallelization in a straight-forward way without optimizing GPU utilization. In this work, we propose a workload-aware auto-parallelization framework (WAP) for DNN training, where the work is automatically distributed to multiple GPUs based on the workload characteristics. We evaluate WAP using TensorFlow with popular DNN benchmarks (AlexNet and VGG-16), and show competitive training throughput compared with the state-of-the-art frameworks, and also demonstrate that WAP automatically optimizes GPU assignment based on the workload's compute requirements, thereby improving energy efficiency. Sungho Shin, Youngmin Jo, Jungwook Choi, Swagath Venkataramani, Vijayalakshmi Srinivasan, Wonyong Sung |
ICASSP | 6 |
| 2018 | Character-level Language Modeling with Gated Hierarchical Recurrent Neural Networks
Iksoo Choi, Jinhwan Park, Wonyong Sung |
INTERSPEECH | 3 |
| 2018 | Hierarchical Recurrent Neural Networks for Acoustic Modeling
Jinhwan Park, Iksoo Choi, Yoonho Boo, Wonyong Sung |
INTERSPEECH | 4 |
| 2018 | Fully Neural Network Based Speech Recognition on Mobile and Embedded DevicesabstractReal-time automatic speech recognition (ASR) on mobile and embedded devices has been of great interests for many years. We present real-time speech recognition on smartphones or embedded systems by employing recurrent neural network (RNN) based acoustic models, RNN based language models, and beam-search decoding. The acoustic model is end-to-end trained with connectionist temporal classification (CTC) loss. The RNN implementation on embedded devices can suffer from excessive DRAM accesses because the parameter size of a neural network usually exceeds that of the cache memory and the parameters are used only once for each time step. To remedy this problem, we employ a multi-time step parallelization approach that computes multiple output samples at a time with the parameters fetched from the DRAM. Since the number of DRAM accesses can be reduced in proportion to the number of parallelization steps, we can achieve a high processing speed. However, conventional RNNs, such as long short-term memory (LSTM) or gated recurrent unit (GRU), do not permit multi-time step parallelization. We construct an acoustic model by combining simple recurrent units (SRUs) and depth-wise 1-dimensional convolution layers for multi-time step parallelization. Both the character and word piece models are developed for acoustic modeling, and the corresponding RNN based language models are used for beam search decoding. We achieve a competitive WER for WSJ corpus using the entire model size of around 15MB and achieve real-time speed using only a single core ARM without GPU or special hardware. Jinhwan Park, Yoonho Boo, Iksoo Choi, Sungho Shin, Wonyong Sung |
NeurIPS | 5 |
| 2018 | On-Device End-to-end Speech Recognition with Multi-Step Parallel RnnsabstractMost of the current automatic speech recognition is performed on a remote server. However, the demand for speech recognition on personal devices is increasing, owing to the requirement of shorter recognition latency and increased privacy. End-to-end speech recognition that employs recurrent neural networks (RNNs) shows good accuracy, but the execution of conventional RNNs, such as the long short-term memory (LSTM) or gated recurrent unit (GRU), demands many memory accesses, thus hindering its real-time execution on smart-phones or embedded systems. To solve this problem, we built an end-to-end acoustic model (AM) using linear recurrent units instead of LSTM or GRU and employed a multi-step parallel approach for reducing the number of DRAM accesses. The AM is trained with the connectionist temporal classification (CTC) loss, and the decoding is conducted using weighted finite-state transducers (WFSTs). The proposed system achieves x4.8 real-time speed when executed on a single core of an ARM CPU-based system. Yoonho Boo, Jinhwan Park, Lukas Lee, Wonyong Sung |
SLT | 4 |
| 2017 | Character-level language modeling with hierarchical recurrent neural networksabstractRecurrent neural network (RNN) based character-level language models (CLMs) are extremely useful for modeling out-of-vocabulary words by nature. However, their performance is generally much worse than the word-level language models (WLMs), since CLMs need to consider longer history of tokens to properly predict the next one. We address this problem by proposing hierarchical RNN architectures, which consist of multiple modules with different timescales. Despite the multi-timescale structures, the input and output layers operate with the character-level clock, which allows the existing RNN CLM training approaches to be directly applicable without any modifications. Our CLM models show better perplexity than Kneser-Ney (KN) 5-gram WLMs on the One Billion Word Benchmark with only 2% of parameters. Also, we present real-time character-level end-to-end speech recognition examples on the Wall Street Journal (WSJ) corpus, where replacing traditional mono-clock RNN CLMs with the proposed models results in better recognition accuracies even though the number of parameters are reduced to 30%. Kyuyeon Hwang, Wonyong Sung |
ICASSP | 2 |
| 2017 | Fixed-point optimization of deep neural networks with adaptive step size retrainingabstractFixed-point optimization of deep neural networks plays an important role in hardware based design and low-power implementations. Many deep neural networks show fairly good performance even with 2- or 3-bit precision when quantized weights are fine-tuned by retraining. We propose an improved fixed-point optimization algorithm that estimates the quantization step size dynamically during the retraining. In addition, a gradual quantization scheme is also tested, which sequentially applies fixed-point optimizations from high- to low-precision. The experiments are conducted for feed-forward deep neural networks (FFDNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). Sungho Shin, Yoonho Boo, Wonyong Sung |
ICASSP | 3 |
| 2017 | SVD-Softmax: Fast Softmax Approximation on Large Vocabulary Neural NetworksabstractWe propose a fast approximation method of a softmax function with a very large vocabulary using singular value decomposition (SVD). SVD-softmax targets fast and accurate probability estimation of the topmost probable words during inference of neural network language models. The proposed method transforms the weight matrix used in the calculation of the output vector by using SVD. The approximate probability of each word can be estimated with only a small part of the weight matrix by using a few large singular values and the corresponding elements for most of the words. We applied the technique to language modeling and neural machine translation and present a guideline for good approximation. The algorithm requires only approximately 20\% of arithmetic operations for an 800K vocabulary case and shows more than a three-fold speedup on a GPU. Kyuhong Shim, Iksoo Choi, Yoonho Boo, Wonyong Sung |
NIPS | 5 |
| 2017 | Structured Pruning of Deep Convolutional Neural NetworksabstractReal-time application of deep learning algorithms is often hindered by high computational complexity and frequent memory accesses. Network pruning is a promising technique to solve this problem. However, pruning usually results in irregular network connections that not only demand extra representation efforts but also do not fit well on parallel computation. We introduce structured sparsity at various scales for convolutional neural networks: feature map-wise, kernel-wise, and intra-kernel strided sparsity. This structured sparsity is very advantageous for direct computational resource savings on embedded computers, in parallel computing environments, and in hardware-based systems. To decide the importance of network connections and paths, the proposed method uses a particle filtering approach. The importance weight of each particle is assigned by assessing the misclassification rate with a corresponding connectivity pattern. The pruned network is retrained to compensate for the losses due to pruning. While implementing convolutions as matrix products, we particularly show that intra-kernel strided sparsity with a simple constraint can significantly reduce the size of the kernel and feature map tensors. The proposed work shows that when pruning granularities are applied in combination, we can prune the CIFAR-10 network by more than 70% with less than a 1% loss in accuracy. Kyuyeon Hwang, Wonyong Sung |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2016 | Evaluation of block turbo codes for long-haul optical networksabstractOptical network traffic has been increasing significantly over the last decade, and strong forward error correction (FEC) is essential to deal with the traffic. While hard-decision codes have been used for the first and the second generation FECs, soft-decision codes are being actively studied as candidates for the third generation codes. The requirements for the third generation codes are over 10 dB net coding gain (NCG) with about 20 % overhead. According to Tzimpragos et al. [1], block turbo codes (BTCs) with 15 % and 20 % overhead obtained the NCGs of over 11 dB. In this paper, we study the BTC pool that targets the long-haul optical networks and evaluate the BTC candidates. In order to achieve the target bit error rate (BER) of 10-15, a large minimum distance is essential for the BTCs. Thus, we estimate the BER performances of the BTCs with the minimum distances of 24 and 36. Since the target BER requires extensive simulations, we estimate the NCGs by observing the BER down to 10-9and conjecturing about the rest of the BER. We also compare the candidate BTCs in terms of the decoding-complexity. Junhee Cho 0002, Wonyong Sung |
APCC | 2 |
| 2016 | Learning separable fixed-point kernels for deep convolutional neural networksabstractDeep convolutional neural networks have shown outstanding performance in several speech and image recognition tasks. However they demand high computational complexity which limits their deployment in resource limited machines. The proposed work lowers the hardware complexity by constraining the learned convolutional kernels to be separable and also reducing the word-length of these kernels and other weights in the fully connected layers. To compensate for the effect of direct quantization, a retraining scheme that includes filter separation and quantization inside of the adaptation procedure is developed in this work. The filter separation reduces the number of parameters and arithmetic operations by 60% for a 5 × 5 kernel, and the quantization further lowers the precision of storage and arithmetic by more than 80 to 90°% when compared to a floating-point algorithm. Experimental results on MNIST and CIFAR-10 datasets are presented. Kyuyeon Hwang, Wonyong Sung |
ICASSP | 3 |
| 2016 | Character-level incremental speech recognition with recurrent neural networksabstractIn real-time speech recognition applications, the latency is an important issue. We have developed a character-level incremental speech recognition (ISR) system that responds quickly even during the speech, where the hypotheses are gradually improved while the speaking proceeds. The algorithm employs a speech-to-character unidirectional recurrent neural network (RNN), which is end-to-end trained with connectionist temporal classification (CTC), and an RNN-based character-level language model (LM). The output values of the CTC-trained RNN are character-level probabilities, which are processed by beam search decoding. The RNN LM augments the decoding by providing long-term dependency information. We propose tree-based online beam search with additional depth-pruning, which enables the system to process infinitely long input speech with low latency. This system not only responds quickly on speech but also can dictate out-of-vocabulary (OOV) words according to pronunciation. The proposed model achieves the word error rate (WER) of 8.90% on the Wall Street Journal (WSJ) Nov'92 20K evaluation set when trained on the WSJ SI-284 training set. Kyuyeon Hwang, Wonyong Sung |
ICASSP | 2 |
| 2016 | FPGA based implementation of deep neural networks using on-chip memory onlyabstractDeep neural networks (DNNs) demand a very large amount of computation and weight storage, and thus efficient implementation using special purpose hardware is highly desired. In this work, we have developed an FPGA based fixed-point DNN system using only on-chip memory not to access external DRAM. The execution time and energy consumption of the developed system is compared with a GPU based implementation. Since the capacity of memory in FPGA is limited, only 3-bit weights are used for this implementation, and training based fixed-point weight optimization is employed. The implementation using Xilinx XC7Z045 is tested for the MNIST handwritten digit recognition benchmark and a phoneme recognition task on TIMIT corpus. The obtained speed is about one quarter of a GPU based implementation and much better than that of a PC based one. The power consumption is less than 5 Watt at the full speed operation resulting in much higher efficiency compared to GPU based systems. Jinhwan Park, Wonyong Sung |
ICASSP | 2 |
| 2016 | Fixed-point performance analysis of recurrent neural networksabstractRecurrent neural networks have shown excellent performance in many applications; however they require increased complexity in hardware or software based implementations. The hardware complexity can be much lowered by minimizing the word-length of weights and signals. This work analyzes the fixed-point performance of recurrent neural networks using a retrain based quantization method. The quantization sensitivity of each layer in RNNs is studied, and the overall fixed-point optimization results minimizing the capacity of weights while not sacrificing the performance are presented. A language model and a phoneme recognition examples are used. Sungho Shin, Kyuyeon Hwang, Wonyong Sung |
ICASSP | 3 |
| 2016 | Sequence to Sequence Training of CTC-RNNs with Partial WindowingabstractConnectionist temporal classification (CTC) based supervised sequence training of recurrent neural networks (RNNs) has shown great success in many machine learning areas including end-to-end speech and handwritten character recognition. For the CTC training, however, it is required to unroll (or unfold) the RNN by the length of an input sequence. This unrolling requires a lot of memory and hinders a small footprint implementation of online learning or adaptation. Furthermore, the length of training sequences is usually not uniform, which makes parallel training with multiple sequences inefficient on shared memory models such as graphics processing units (GPUs). In this work, we introduce an expectation-maximization (EM) based online CTC algorithm that enables unidirectional RNNs to learn sequences that are longer than the amount of unrolling. The RNNs can also be trained to process an infinitely long input sequence without pre-segmentation or external reset. Moreover, the proposed approach allows efficient parallel training on GPUs. Our approach achieves 20.7% phoneme error rate (PER) on the very long input sequence that is generated by concatenating all 192 utterances in the TIMIT core test set. In the end-to-end speech recognition task on the Wall Street Journal corpus, a network can be trained with only 64 times of unrolling with little performance loss. Kyuyeon Hwang, Wonyong Sung |
ICML | 2 |
| 2016 | Dynamic hand gesture recognition for wearable devices with low complexity recurrent neural networksabstractGesture recognition is a very essential technology for many wearable devices. While previous algorithms are mostly based on statistical methods including the hidden Markov model, we develop two dynamic hand gesture recognition techniques using low complexity recurrent neural network (RNN) algorithms. One is based on video signal and employs a combined structure of a convolutional neural network (CNN) and an RNN. The other uses accelerometer data and only requires an RNN. Fixed-point optimization that quantizes most of the weights into two bits is conducted to optimize the amount of memory size for weight storage and reduce the power consumption in hardware and software based implementations. Sungho Shin, Wonyong Sung |
ISCAS | 2 |
| 2015 | Fixed point optimization of deep convolutional neural networks for object recognitionabstractDeep convolutional neural networks have shown promising results in image and speech recognition applications. The learning capability of the network improves with increasing depth and size of each layer. However this capability comes at the cost of increased computational complexity. Thus reduction in hardware complexity and faster classification are highly desired. This work proposes an optimization method for fixed point deep convolutional neural networks. The parameters of a pre-trained high precision network are first directly quantized using L2 error minimization. We quantize each layer one by one, while other layers keep computation with high precision, to know the layer-wise sensitivity on word-length reduction. Then the network is retrained with quantized weights. Two examples on object recognition, MNIST and CIFAR-10, are presented. Our results indicate that quantization induces sparsity in the network which reduces the effective number of network parameters and improves generalization. This work reduces the required memory storage by a factor of 1/10 and achieves better classification results than the high precision networks. Kyuyeon Hwang, Wonyong Sung |
ICASSP | 3 |
| 2015 | Single stream parallelization of generalized LSTM-like RNNs on a GPUabstractRecurrent neural networks (RNNs) have shown outstanding performance on processing sequence data. However, they suffer from long training time, which demands parallel implementations of the training procedure. Parallelization of the training algorithms for RNNs are very challenging because internal recurrent paths form dependencies between two different time frames. In this paper, we first propose a generalized graph-based RNN structure that covers the most popular long short-term memory (LSTM) network. Then, we present a parallelization approach that automatically explores parallelisms of arbitrary RNNs by analyzing the graph structure. The experimental results show that the proposed approach shows great speed-up even with a single training stream, and further accelerates the training when combined with multiple parallel training streams. Kyuyeon Hwang, Wonyong Sung |
ICASSP | 2 |
| 2014 | X1000 real-time phoneme recognition VLSI using feed-forward deep neural networksabstractDeep neural networks show very good performance in phoneme and speech recognition applications when compared to previously used GMM (Gaussian Mixture Model)-based ones. However, efficient implementation of deep neural networks is difficult because the network size needs to be very large when high recognition accuracy is demanded. In this work, we develop a digital VLSI for phoneme recognition using deep neural networks and assess the design in terms of throughput, chip size, and power consumption. The developed VLSI employs a fixed-point optimization method that only uses +Δ, 0, and -Δ for representing each of the weight. The design employs 1,024 simple processing units in each layer, which however can be scaled easily according to the needed throughput, and the throughput of the architecture varies from 62.5 to 1,000 times of the real-time processing speed. Jonghong Kim, Kyuyeon Hwang, Wonyong Sung |
ICASSP | 3 |
| 2014 | Fault tolerance analysis of digital feed-forward deep neural networksabstractAs the homeostatis characteristics of nerve systems show, artificial neural networks are considered to be robust to variation of circuit components and interconnection faults. However, the tolerance of neural networks depends on many factors, such as the fault model, the network size, and the training method. In this study, we analyze the fault tolerance of fixed-point feed-forward deep neural networks for the implementation in CMOS digital VLSI. The circuit errors caused by the interconnection as well as the processing units are considered. In addition to the conventional and dropout training methods, we develop a new technique that randomly disconnects weights during the training to increase the error resiliency. Feed-forward deep neural networks for phoneme recognition are employed for the experiments. Kyuyeon Hwang, Wonyong Sung |
ICASSP | 3 |
| 2014 | Power Modeling for GPU Architectures Using McPATabstractGraphics Processing Units (GPUs) are very popular for both graphics and general-purpose applications. Since GPUs operate many processing units and manage multiple levels of memory hierarchy, they consume a significant amount of power. Although several power models for CPUs are available, the power consumption of GPUs has not been studied much yet. In this article we develop a new power model for GPUs by utilizing McPAT, a CPU power tool. We generate initial power model data from McPAT with a detailed GPU configuration, and then adjust the models by comparing them with empirical data. We use the NVIDIA's Fermi architecture for building the power model, and our model estimates the GPU power consumption with an average error of 7.7% and 12.8% for the microbenchmarks and Merge benchmarks, respectively. Jieun Lim 0001, Nagesh B. Lakshminarayana, Hyesoon Kim, William J. Song, Sudhakar Yalamanchili, Wonyong Sung |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2014 | Rate-0.96 LDPC Decoding VLSI for Soft-Decision Error Correction of NAND Flash MemoryabstractThe reliability of data stored in high-density Flash memory devices tends to decrease rapidly because of the reduced cell size and multilevel cell technology. Soft-decision error correction algorithms that use multiple-precision sensing for reading memory can solve this problem; however, they require very complex hardware for high-throughput decoding. In this paper, we present a rate-0.96 (68254, 65536) shortened Euclidean geometry low-density parity-check code and its VLSI implementation for high-throughput NAND Flash memory systems. The design employs the normalized a posteriori probability (APP)-based algorithm, serial schedule, and conditional update, which lead to simple functional units, halved decoding iterations, and low-power consumption, respectively. A pipelined-parallel architecture is adopted for high-throughput decoding, and memory-reduction techniques are employed to minimize the chip size. The proposed decoder is implemented in 0.13-μm CMOS technology, and the chip size and energy consumption of the decoder are compared with those of a BCH (Bose-Chaudhuri-Hocquenghem) decoding circuit showing comparable error-correcting performance and throughput. Jonghong Kim, Wonyong Sung |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Soft-decision decoding with cell to cell interference removed signal in nand flash memoryabstractAs the semiconductor feature size continuously decreases, the signal integrity of MLC (multi-level cell) NAND flash memory becomes degraded, which causes increased bit error rate and short retention time limit. Since it is well known that cell-to-cell interference (CCI) is one of the major sources of bit errors, several signal processing solutions to mitigate or remove the CCI have been developed. Even though these works show BER (bit error rate) performance improvement with hard-decision error correction, combining an interference canceller and soft-decision decoding still needs to be studied. In this research, we derive a mathematical formulation for computing the likelihood function of the CCI removed signal, and then propose a conditional LLR (log-likelihood ratio) to represent soft-reliability information for soft-decision error correction. The proposed soft-information computation scheme is applied to simulated NAND flash memory, and it is demonstrated that soft-decision error correction with the CCI removed signal significantly increases the retention time limit of MLC NAND flash memory. Dong-hwan Lee, Wonyong Sung |
ICASSP | 2 |
| 2013 | GPU based implementation of recursive digital filtering algorithmsabstractRecursive filtering is widely used for many signal processing applications. Speeding-up the computation of recursive filtering using many processing elements is difficult because of the dependency problem. In this paper, massively parallel computation of recursive filtering algorithms using GPGPUs (General Purpose Graphics Processing Units) is studied. The proposed method uses the multi-block parallel processing algorithm, where each thread executes one block of data as independently as possible. To resolve the dependency among threads, we develop a fast look-ahead method that shows high efficiency even when thousands of threads are used. The developed method has been implemented using Nvidia GTX 285 GPU and shows over 15 times of speed-up when compared to sequential CPU based implementations. Dong-hwan Lee, Wonyong Sung |
ICASSP | 2 |
| 2013 | DRAM access reduction in GPUs by thread-block scheduling for overlapped data reuseabstractGeneral Purpose Graphics Processing Units (GPG-PUs) show very high throughput when executing parallel programs. However, they usually demand very large DRAM bandwidth and consume much power for memory access. Although recent high performance GPGPUs equip L2 cache to absorb some of DRAM accesses, the cache hit ratio can hardly be very high because of the limited cache size. We propose a GPU thread-block scheduling method that can better utilize L2 cache and reduce the DRAM memory access. This scheduling method exploits the inter-block locality in the scheduling of GPU thread-blocks. This method can easily be implemented by modifying application programs. This technique is applied to the Hotspot benchmark programs, and reduces the DRAM access by up to 39%. Seungyeol Lee, Wonyong Sung |
ISCAS | 2 |
| 2012 | Multi-user real-time speech recognition with a GPUabstractWe have developed a multi-user large vocabulary speech recognition system employing a fully composed one-level weighted finite state transducer (WFST) based network on a Graphics Processing Unit (GPU). This system improves the overall throughput and latency of speech recognition engine which processes multiple users' utterances at the same time with efficient scheduling, parameter sharing, and communication overhead reduction techniques. We conduct both batch speech simulation and trace driven online simulation to access the performance of the developed system. Traces are generated based on a queueing model. Jungsuk Kim, Wonyong Sung |
ICASSP | 2 |
| 2012 | Least squares based cell-to-cell interference cancelation technique for multi-level cell nand flash memoryabstractCell-to-cell interference becomes a major source of bit errors in NAND flash memories as the semiconductor technology continuously shrinks down. Recently, signal processing approaches to mitigate the interference have been proposed, and the least mean square (LMS) adaptive filtering based method [1] offers a promising solution. In this research, we propose a least squares based cell-to-cell interference cancelation method, which is more suitable for NAND flash memory devices where one page of data is accessed at a time. With a simulation model, we show that this approach outperforms the LMS filtering based one whether the interference is severe or not. In order to simplify the algorithm, the input data to compute the channel characteristics is decimated, and as a result the arithmetic intensity of the proposed algorithm is comparable to the LMS based one. Dong-hwan Lee, Wonyong Sung |
ICASSP | 2 |
| 2012 | Performance of rate 0.96 (68254, 65536) EG-LDPC code for NAND Flash memory error correctionabstractAs the process technology scales down and the number of bits per cell increases, NAND Flash memory is more prone to bit errors. In this paper, we employ a rate-0.96 (68254, 65536) Euclidean geometry (EG) low-density parity-check (LDPC) code for NAND Flash memory error correction, and evaluate the performance under binary input (BI) additive white Gaussian noise (AWGN) and NAND Flash memory channels. The performance effect of output signal quantization is also studied. We show the strategies for determining the optimum quantization boundaries and computing the quantized log-likelihood ratio (LLR) for the NAND Flash channel model that is approximated as a mixture of Gaussian distributions. Simulation results show that the error performance with the NAND Flash memory channel is much different from that with the BI-AWGN channel. Since the distribution of NAND Flash memory output signal is not stationary, it is important to accurately assess the stochastic distribution of the signal for optimum sensing. Jonghong Kim, Dong-hwan Lee, Wonyong Sung |
ICC | 3 |
| 2012 | Performance analysis of multi-bank DRAM with increased clock frequencyabstractAs the performance of computer systems improves, the peak bandwidth of the DRAM system needs to be increased. In this study, we analyze the performance of multi-bank DRAMs when increasing the clock frequency by employing three metrics: data bus busy time, bank busy time and inter-bank interference time. We use a cycle-accurate DRAM model simulator to quantitatively measure each metric. Increasing the DRAM clock frequency obviously contributes to lowering the data bus busy time. From the analysis result, we find that raising the number of banks is needed when increasing the DRAM clock frequency. However, the inter-bank interference time becomes the performance bottleneck as the number of banks increases. We suggest that future multi-bank DRAM system should tackle this side-effect to efficiently exploit the faster clock frequency. Su-Jin Cho, Jae-Woo Ahn, Hyojin Choi, Wonyong Sung |
ISCAS | 4 |
| 2012 | A simulation-based study for DRAM power reduction strategies in GPGPUsabstractGeneral Purpose Graphics Processing Units (GPGPUs) operate many threads concurrently, however they demand large DRAM access because of small internal memory size assigned to each thread. As a result, the power consumption in DRAM components becomes increasingly significant. We have examined a few techniques that can reduce DRAM power consumption in GPGPUs. A GPGPU simulator supporting L2 cache is used for this study. The effects of changing the memory channel organization, DRAM clock frequency, row buffer management policy, open or closed, and the L2 cache memory system are studied. Not only the total DRAM energy consumption but also that due to each DRAM operation, such as active-precharge, burst, and background, are estimated. The examined DRAM power reduction techniques bring negligible execution time changes for solving compute-bound problems, but they result in 12–27% savings of DRAM power consumption. Hyojin Choi, Kyuyeon Hwang, Jae-Woo Ahn, Wonyong Sung |
ISCAS | 4 |
| 2011 | H- and C-level WFST-based large vocabulary continuous speech recognition on Graphics Processing UnitsabstractWe have implemented 20,000-word large vocabulary continuous speech recognition (LVCSR) systems employing Hand C-level weighted finite state transducer (WFST) based networks on Graphics Processing Units (GPUs). Both the emission probability computation and the Viterbi beam search are implemented on the GPU in a data-parallel manner to minimize the extra data transfer time between the host CPU and the GPU. This study utilizes word-length optimization techniques to reduce the synchronization overhead in the Viterbi beam search. We achieve 18.6% to 21.9% of speed up by using an efficient data packing method with less than 0.2% accuracy degradation. Furthermore, we explore different levels of abstraction in recognition network generation to reduce the number of synchronization operations as well as to minimize the memory usage. The experimental results show that the implemented systems on the GPU perform speech recognition 4.07 to 4.55 times faster than highly optimized sequential implementations on a CPU. Jungsuk Kim, Kisun You, Wonyong Sung |
ICASSP | 3 |
| 2011 | Parallel computation of adaptive lattice filtersabstractParallel computation of the adaptive lattice filtering algorithm is difficult due to the dependency problem caused by feedback operations. The conventional control-level parallel computation method that exploits the modular structure of the filtering algorithm can only utilize a limited degree of parallelism even when the algorithm is pipeline-transformed. In order to increase the degree of parallelism, we apply the data-level parallel processing method that computes multiple output samples at a time by parallelizing the computation of time-varying linear recursive equations. The control-level parallel processing approach is useful for SIMD (Single Instruction Multiple Data) processor based implementations. However, the data-level parallel processing method is indispensable for multicore based implementations not only to utilize the increased number of processing cores but also to overcome the communication delay between cores. Dong-hwan Lee, Wonyong Sung |
ICASSP | 2 |
| 2011 | Memory access pattern-aware DRAM performance model for multi-core systemsabstractThe DRAM latency modeling is complex because most chips contain row-buffers and multiple banks to exploit patterns of DRAM accesses. As a result, the latency of DRAM access not only depends on the circuit timing parameters but also memory access patterns. This study derives an analytical model that predicts the DRAM access performance using DRAM timing and memory access pattern parameters. As a performance metric, the bank busy time of DRAM is used. The pattern parameters employed represent memory access characteristics such as the number of row-buffer misses, the number of read or write requests that hit the row-buffers, etc. The proposed model not only relates the DRAM access performance with the memory access pattern but also provides information for timing optimization of next generation DRAMs. The model is evaluated with SPLASH-2 benchmark by using cycle-accurate timing simulations with DDR3 timings. The evaluation results show that, in memory-bounded cases, the execution time is limited by bank utilization, not by the data bus occupation ratio. Hyojin Choi, Jongbok Lee, Wonyong Sung |
ISPASS | 3 |
| 2010 | An FPGA implementation of speech recognition with weighted finite state transducersabstractIn this paper we present a hardware architecture for large vocabulary continuous speech recognition that conducts a search over a weighted finite state transducer (WFST) network. A pipelined architecture is proposed to fully utilize the memory bandwidth. A hash table is used to manage small sized working sets efficiently. We also applied a parallelization technique that increases the traversal speed by 17%. The recognition system is fully functional on an FPGA, which runs at 100MHz. The experimental result on the Wall Street Journal 5,000 vocabulary task shows that the recognition speed of the system is 5.3× faster than real-time. Jungwook Choi, Kisun You, Wonyong Sung |
ICASSP | 3 |
| 2010 | Multi-core and SIMD architecture based implementation of recursive digital filtering algorithmsabstractImplementation of recursive filtering equations using parallel computer architecture is difficult because of the dependency problem. In this paper, parallel computation of recursive filtering equations is studied for multi-core architecture with SIMD (Single Instruction Multiple Data) arithmetic support. In order to exploit both SIMD and multi-core features, a multi-block parallel processing algorithm is employed, which computes the particular solutions for multiple blocks simultaneously and then compensates for the homogeneous solutions. The implementation results are obtained using a Pentium quad-core CPU supporting four-way SIMD processing. Dong-hwan Lee, Wonyong Sung |
ICASSP | 2 |
| 2010 | Parallel implementation of an error diffusion halftoning algorithm with a general purpose graphics processing unitabstractGeneral purpose graphics processing units (GPGPUs) contain many execution units, thus they are very attractive for high speed image processing. However, the error diffusion halftoning algorithm can hardly exploit the benefit of massively parallel processing architecture because this algorithm uses feedback of the output error as well as the results of neighboring pixels. In this study, pixels that can be processed without dependency are found by examining the dependency graph. Also, a parallel processing method requiring less synchronization overhead is developed by considering the characteristics of GPGPUs. Becksang Seong, Jae-Woo Ahn, Wonyong Sung |
ICIP | 3 |
| 2010 | VLSI Implementation of BCH Error Correction for Multilevel Cell NAND Flash MemoryabstractBit-error correction is crucial for realizing cost-effective and reliable NAND Flash-memory-based storage systems. In this paper, low-power and high-throughput error-correction circuits have been developed for multilevel cell (MLC) nand Flash memories. The developed circuits employ the Bose-Chaudhuri-Hocquenghem code to correct multiple random bit errors. The error-correcting codes for them are designed based on the bit-error characteristics of MLC NAND Flash memories for solid-state drives. To trade the code rate, circuit complexity, and power consumption, three error-correcting architectures, named as whole-page, sector-pipelined, and multistrip ones, are proposed. The VLSI design applies both algorithmic and architectural-level optimizations that include parallel algorithm transformation, resource sharing, and time multiplexing. The chip area, power consumption, and throughput results for these three architectures are presented. Hyojin Choi, Wonyong Sung |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | VLSI for 5000-word continuous speech recognitionabstractWe have developed a VLSI chip for 5,000 word speaker-independent continuous speech recognition. This chip employs a context-dependent HMM (hidden Markov model) based speech recognition algorithm, and contains emission probability and Viterbi beam search pipelined hardware units. The feature vector for speech recognition is computed using a host processor in software in order to adopt various enhancement algorithms. The amount of internal SRAM size is minimized by moving data out to the external DRAM, and a custom DRAM controller module is designed to efficiently read and write consecutive data. The experimental result shows that the implemented system has a real-time factor of 0.77 and 0.55 using SDRAM and DDR SDRAM, respectively. Kisun You, Jungwook Choi, Wonyong Sung |
ICASSP | 4 |
| 2009 | OpenMP-based parallel implementation of a continuous speech recognizer on a multi-core systemabstractWe have implemented a 20,000-word continuous speech recognizer on a multi-core based system. A fine grain parallel processing approach is employed for good scalability, and the OpenMP library is used for enhanced portability. In the emission probability computation, a dynamic workload distribution method is employed for good load balancing. However, the search network involved in the Viterbi beam search is statically partitioned into independent subtrees to reduce memory synchronization overhead. In order to further improve the performance, a workload predictive thread assignment strategy as well as a false cache line sharing prevention method are employed. The test was conducted using WSJ1 20 k test and development set. We achieved the speed-up of 3.90 by utilizing four threads parallelization in a four-core system compared to four copies of the baseline single thread speech recognizer running simultaneously. The final recognition system runs about twice the speed of the real-time requirement. Kisun You, Youngjoon Lee, Wonyong Sung |
ICASSP | 3 |
| 2009 | Scalable HMM based inference engine in large vocabulary continuous speech recognitionabstractParallel scalability allows an application to efficiently utilize an increasing number of processing elements. In this paper we explore a design space for application scalability for an inference engine in large vocabulary continuous speech recognition (LVCSR). Our implementation of the inference engine involves a parallel graph traversal through an irregular graph-based knowledge network with millions of states and arcs. The challenge is not only to define a software architecture that exposes sufficient fine-grained application concurrency, but also to efficiently synchronize between an increasing number of concurrent tasks and to effectively utilize the parallelism opportunities in today's highly parallel processors. We propose four application-level implementation alternatives we call ldquoalgorithm stylesrdquo, and construct highly optimized implementations on two parallel platforms: an Intel Core i7 multicore processor and a NVIDIA GTX280 manycore processor. The highest performing algorithm style varies with the implementation platform. On 44 minutes of speech data set, we demonstrate substantial speedups of 3.4times on Core i7 and 10.5times on GTX280 compared to a highly optimized sequential implementation on Core i7 without sacrificing accuracy. The parallel implementations contain less than 2.5% sequential overhead, promising scalability and significant potential for further speedup on future platforms. Jike Chong, Kisun You, Youngmin Yi, Ekaterina Gonina, Christopher J. Hughes, Wonyong Sung, Kurt Keutzer |
ICME | 6 |
| 2009 | VLSI Implementation of a Soft Bit-flipping Decoder for PG-LDPC CodesabstractImplementation of high throughput VLSI chips for low-density parity-check codes has been considered very difficult especially when the row or column weight of the code is high. In this paper, a projective-geometry (PG) LDPC code is implemented in VLSI employing the proposed soft bit flipping (SBF) algorithm. The SBF algorithm requires only simple interconnections, but its error correcting performance is close to the sum-product algorithm (SPA). Parallel processing architecture is employed for increasing the throughput. With the (1057, 813) PG-LDPC code, the implemented 4-bit SBF decoder consumes only a small area of 2.5mm2while providing 6.5Gbps and good performance close to the floating-point SPA by 0.6dB at the frame error rate of 10−4. Jonghong Kim, Hyunwoo Ji, Wonyong Sung |
ISCAS | 4 |
| 2009 | Efficient Software-Based Encoding and Decoding of BCH CodesabstractError correction software for Bose-Chaudhuri-Hochquenghem (BCH) codes is optimized for general purpose processors that do not equip hardware for Galois field arithmetic. The developed software applies parallelization with a table lookup method to reduce the number of iterations, and maximum parallelization under a cache size limitation is sought for a high throughput implementation. Since this method minimizes the number of lookup tables for encoding and decoding processes, a large parallel factor can be chosen for a given cache size. The naive word length of a general purpose CPU is used as a whole by employing the developed mask elimination method. The tradeoff of the algorithm complexity and the regularity is examined for several syndrome generation methods, which leads to a simple error detection scheme that reuses the encoder and a simplified syndrome generation method requiring only a small number of Galois field multiplications. The parallel factor for Chien search is increased much by transforming the error locator polynomial so that it contains symmetric exponents of positive and negative signs. The experimental results demonstrate that the developed software cannot only provide sufficient throughput for real-time error correction of NAND flash memory in embedded systems but also enhance the reliability of file systems in general purpose computers. Wonyong Sung |
IEEE Trans. Computers | 2 |
| 2009 | Access-Pattern-Aware On-Chip Memory Allocation for SIMD ProcessorsabstractThe number of cycles for each external memory access in Single Instruction Multiple Data (SIMD) processors is heavily affected by the access pattern, such as aligned, unaligned, or stride. We developed a high-performance dynamic on-chip memory-allocation method for SIMD processors by considering the memory access pattern as well as the access frequency. The access pattern and the access count for an array of a loop are determined by both code analysis and profiling, which are performed on a developed compiler framework. This framework not only conducts dynamic on-chip memory allocation but also generates optimized codes for a target processor. The proposed allocation method has been tested with several multimedia benchmarks including motion estimation, 2-D discrete cosine transform, and MPEG2 encoder programs. Hoseok Chang, Wonyong Sung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Efficient vectorization of SIMD programs with non-aligned and irregular data access hardwareabstractAutomatic vectorization of programs for partitioned-ALU SIMD (Single Instruction Multiple Data) processors has been difficult because of not only data dependency issues but also non-aligned and irregular data access problems. A non-aligned or irregular data access operation incurs many overhead cycles for data alignment. Moreover, this causes difficulty in efficient code generation and hinders automatic vectorization. In this paper, we employ special memory access hardware for improving the performance of SIMD processors; one is the split line buffer and the other is the packing buffer. The former solves the non-aligned memory access problem, while the latter simplifies irregular and stride data access. The addition of these hardware units not only requires very small changes to the instruction set architecture but also contributes to the significant performance improvement by vectorizing more loops and reducing the overhead cycles. We have also developed an auto-vectorization compiler which utilizes these special hardware units. Experiments have been conducted to compare the proposed method with the conventional one, which show 50% increase in the number of vectorized loops and 77% increase in the total performance of an MPEG2 encoder program. Hoseok Chang, Wonyong Sung |
CASES | 2 |
| 2008 | Software implementation of Chien search process for strong BCH codesabstractThe Chien search process is the most time-consuming portion in the software decoding of Bose-Chaudhuri-Hochquenghem (BCH) codes. It can be parallelized using a table look-up method, but a large memory space is required for a high parallel factor. In this paper, we propose to perform the Chien search process not on the standard basis but on the shifted one. By simply shifting the basis, we can reduce the size of the look-up table significantly, thereby shortening the execution time with a much larger parallel factor. The experimental results show that the parallel factor is approximately doubled while using the same memory space, and up to 35% of speedup is achieved. Wonyong Sung |
ISCAS | 2 |
| 2007 | Mobile CPU Based Optimization of Fast Likelihood Computation for Continuous Speech RecognitionabstractWe studied the real-time implementation of continuous speech recognition algorithm on a mobile CPU, Intel PXA270, platform. Especially, the optimization of fast likelihood computation, which takes the largest part in most speech recognizers, is conducted by employing SIMD (single instruction multiple data) programming and software pipelining. The overhead of exhaustive memory accesses is also minimized by placing frequently used acoustic model data at the fast internal SRAM. The number of execution cycles for the fast likelihood computation of the 1000-word vocabulary resource management (RM) task has been reduced by 56.42%. The resulting performance shows approximately four times faster processing speed than the real-time implementation requirement on a 520 MHz Intel XScale-based system. Kisun You, Youngjoon Lee, Wonyong Sung |
ICASSP (4) | 3 |
| 2007 | Memory Access Reduced Software Implementation of H.264/AVC Sub-pixel Motion Estimation Using Differential Data EncodingabstractWe studied an efficient software implementation of H.264/AVC sub-pixel motion estimation (ME) algorithm on a VLIW-SIMD digital signal processor, TMS320C6416. The sub-pixel ME algorithm demands large memory accesses while the required arithmetic operations are fairly simple. Although the CPU clock cycles for arithmetic operations can be reduced much by employing sub-word operations and applying software pipelining techniques, the limited memory bandwidth of the architecture restricts the overall performance. Moreover, aggressive VLIW-SIMD optimization results in the degradation of the performance by causing excessive CPU stalls during memory accesses. In this paper, we relieved the memory bandwidth requirements for creating quarter-pixel images by reducing the precision of image data, from 8 bits to 4 bits. As a result, the amount of memory accesses is much reduced at the cost of some increase of the arithmetic operations, which contributes to the balance of arithmetic and memory access operations. The experimental result shows that the memory stall cycles are decreased by 80% and the speed-up of 260% is obtained. The bit rate of the encoded video stream is increased slightly, about 2% on the average, due to the effects of quantization. Hyojin Choi, Wonchul Lee, Wonyong Sung |
ISCAS | 3 |
| 2006 | Design and Implementation of Speech Recognition on a Softcore Based FpgaabstractA speech recognition algorithm is implemented on a Xilinx's MicroBlaze softcore based FPGA platform. When compared to programmable DSP or RISC processor based design, this approach not only can support large vocabulary speech recognition but also leads to a multi-channel system implementation utilizing an off-the-shelf FPGA device. The design space is explored in a few ways; firstly by configuring the datapath of the MicroBlaze softcore and secondly by adding a custom hardware block to speed up the emission probability computation. The hardware block designed for emission probability computation incorporates the BRAM where the reference data is accessed directly thereby reduces the overhead of transferring data from the processor to the hardware block. Hyunjin Lim, Kisun You, Wonyong Sung |
ICASSP (3) | 3 |
| 2006 | An FPGA based SIMD processor with a vector memory unitabstractA SIMD processor that contains a 16-way partitioned data-path is designed for efficient multimedia data processing. In order to automatically align data needed for SIMD processing, the architecture adopts a vector memory unit that consists of 17-bank memory blocks. The vector memory unit also has address generation and rearrangement units for eliminating bank conflicts. The MicroBlaze FPGA based RISC processor is used for program control and scalar data processing. The architecture has been implemented on a Xilinx FPGA, and the implementation performance for several multimedia kernels is obtained. Hoseok Chang, Wonyong Sung |
ISCAS | 3 |
| 2004 | Implementation of a digital color copier using a VLIW SIMD architectureabstractWe have developed real-time image processing programs for a digital color copier using TMS320C6416 digital signal processor. The processor is good for realtime image processing because of multiple and packed-data processing functional units. However, it needs careful programming to exploit deep pipelining, multiple functional units and packed-data instructions. To improve the unit utilization ratio, we developed a few techniques, such as pack/unpack reduction, resource maximization and operation reduction, and multi-pixel processing. All the critical functions for the implementation of a digital color copier, which include shading correction, X-zooming, 2D filtering, and vector half-toning, are implemented. It is shown that a 720 MHz C6416 CPU can perform all the real-time processing needed for a 20 PPM (page per minute) 600 DPI A4 size color copier. We used C programming with intrinsic functions and linear assembly programming followed by the assembly optimizer. Seokhoon Ju, Wonyong Sung |
ICASSP (5) | 2 |
| 2004 | Implementation of an intonational quality assessment system for a handheld deviceabstractIn this paper, we describe an implementation of an intonational quality assessment system for foreign language learning using a handheld portable device. The Viterbi algorithm is employed to conduct the forced alignments that indicate the boundary of each phonemes and a pitch detector is used to extract the intonational features. The tonal pitch type of the segmented syllables is classified and the tendency of the pitch movement is measured. Then, the score of the spoken sentence is generated based on this information. We have implemented this system on an ARM7 RISC processor based system. For real time operation, we applied fixed-point arithmetic to the signal processing kernels and rearranged the algorithm flow of the system. As a result, the system runs in real time on a 60MHz CPU clock frequency. Reference DB Kisun You, Hoyoun Kim, Wonyong Sung |
INTERSPEECH | 3 |
| 2003 | Implementation of a digital copier using TMS320C6414 VLIW DSP processorabstractIn this paper, we developed real-time image processing programs for a digital copier using a TMS320C6414 CPU. The CPU is good for real-time image processing because of multiple and packed-data processing functional units. However, it needs careful programming to exploit deep pipelining, multiple functional units and packed-data instructions. All the critical functions for the implementation of a digital copier, which include shading correction, X-zoom, 2D filtering, and halftoning, are implemented through assembly programming. Programs using linear assembly programming followed by the assembly optimization in software are compared with the manual assembly coded versions. The results show that explicit disambiguation of memory dependency is most critical for the assembly optimization. The cache miss effects are also evaluated. Taeksang Hwang, Wonyong Sung |
ICASSP (2) | 2 |
| 2002 | Implementation of an intonational quality assessment systemabstractIn this paper, we describe an intonational quality scoring system for foreign language learning. We employed the segmental K-means algorithm to segment the speech into syllables and a pitch detection algorithm to extract the intonational features. We classified the segmented syllables into 5 types according to the shapes of the pitch contours in them. To present visual aids to students, we displayed the classified tonal pitch type of each syllable and the overall pitch movement tendency of the test and reference sentences. We devised an algorithm to obtain a score from the spoken sentence and used this value as a measure for assessing the intonational quality. 1. Wonyong Sung |
INTERSPEECH | 2 |
| 2001 | A block priority based instruction caching scheme for multimedia processorsabstractAn instruction caching scheme that utilizes the block priority information is proposed mainly targeted for embedded multimedia processors. The block priority information is obtained by profiling application programs. The goal of this caching scheme is to keep more important code blocks longer using the block priority information, which programmers provide by analyzing the profiling results of multimedia applications. In addition to a new caching scheme, the methods for determining the priority of each code block are also developed and their performances are evaluated using real multimedia applications. The experimental results show that the cache miss ratio can be reduced up to nearly a half of that of the normal LRU replacement scheme although the improvement depends on the cache size. Jiyang Kang, Wonyong Sung |
ICASSP | 2 |
| 2001 | Combined word-length optimization and high-level synthesis ofdigital signal processing systemsabstractConventional approaches for fixed-point implementation of digital signal processing algorithms require the scaling and word-length (WL) optimization at the algorithm level and the high-level synthesis for functional unit sharing at the architecture level. However, the algorithm-level WL optimization has a few limitations because it can neither utilize the functional unit sharing information for signal grouping nor estimate the hardware cost for each operation accurately. In this study, we develop a combined WL optimization and high-level synthesis algorithm not only to minimize the hardware implementation cost, but also to reduce the optimization time significantly. This software initially finds the WL sensitivity or minimum WL of each signal throughout fixed-point simulations of a signal flow graph, performs the WL conscious high-level synthesis where signals having the similar WL sensitivity are assigned to the same functional unit, and then conducts the final WL optimization by iteratively modifying the WLs of the synthesized hardware model. A list-scheduling-based and an integer linear-programming-based algorithms are developed for the WL conscious high-level synthesis. The hardware cost function to minimize is generated by using a synthesized hardware model. Since fixed-point simulation is used to measure the performance, this method can be applied to general, including nonlinear and time-varying, digital signal processing systems. A fourth-order infinite-impulse response filter, a fifth-order elliptic filter, and a 12th-order adaptive least mean square filter are implemented using this software. Ki-Il Kum, Wonyong Sung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Multimedia processor-based implementation of an error-diffusion halftoning algorithm exploiting subword parallelismabstractMultimedia processor-based implementations of digital image processing algorithms have become important since several multimedia processors are now available and can replace special-purpose hardware-based systems because of their flexibility. Multimedia processors increase throughput by processing multiple pixels simultaneously using a subword-parallel arithmetic and logic unit architecture. The error-diffusion halftoning algorithm employs feedback of quantized output signals to faithfully convert a multi-level image to a binary image or to one with fewer levels of quantization. This makes it difficult to achieve speedup by utilizing the multimedia extension. In this study, the error-diffusion halftoning algorithm is implemented for a multimedia processor using three methods: single-pixel, single-line, and multiple-line processing. The single-pixel approach is the closest to conventional implementations, but the multimedia extension is used only in the filter kernel. The single-line approach computes multiple pixels in one scan-line simultaneously, but requires a complex algorithm transformation to remove dependencies between pixels. The multiple-line method exploits parallelism by employing a skewed data structure and processing multiple pixels in different scan-lines. The Pentium MMX instruction set is used for quantitative performance evaluation including run-time overheads and misaligned memory accesses. A speedup of more than ten times is achieved compared to the software (integer C) implementation on a conventional processor for the structurally sequential error-diffusion halftoning algorithm. Jae-Woo Ahn, Wonyong Sung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Variable dimensional algebraic CELP coding of prototype waveformsabstractWe propose a variable dimensional algebraic codebook structure in order to quantize the prototype waveforms efficiently. The proposed algorithm adjusts the interval of candidate pulse positions in a codevector according to the pitch period. The analysis-by-synthesis search procedure is computationally efficient due to the characteristics of the algebraic codebook structure. We also develop a method to perceptually improve the basis pulse of codevectors, which enhances the reconstructed speech quality with little increase in computational complexity. The improved prototype waveform interpolation coder adopting the proposed methods achieves a high quality of speech at 4 kbps. Jongseo Sohn, Wonyong Sung |
ICASSP | 2 |
| 2000 | Memory efficient software synthesis with mixed coding style from dataflow graphsabstractThis paper presents a set of techniques to reduce the code and data sizes for software synthesis from graphical digital signal-processing programs based on the synchronous dataflow model. By sharing the kernel code among multiple instances of a block with a shared function, we can further reduce the code size below the previous results based on inline coding style. A systematic approach also is devised to give up the single appearance schedule for reducing the data buffer requirement. The proposed techniques have been evaluated with two real-life examples to prove their significance. Wonyong Sung, Soonhoi Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1999 | A floating-point to integer C converter with shift reduction for fixed-point digital signal processorsabstractA floating-point to integer C program translator is developed for convenient programming and efficient use of fixed-point programmable digital signal processors (DSPs). It not only converts data types and supports automatic scaling, but also conducts shift optimization to enhance execution speed. Since the input and output of this translator are ANSI C compliant programs, it can be used for any fixed-point DSP that supports ANSI C compiler. A shift reduction method is developed for minimizing the scaling overhead of translated integer C programs. It considers the data-path of a target processor and profiling results. Using the shift reduction method, 4% to 37% speedup is obtained. The translated integer C codes are 20 to 400 times faster than the floating-point versions when applied to TMS320C50, TMS320C60 and Motorola 56000 DSPs. Ki-Il Kum, Jiyang Kang, Wonyong Sung |
ICASSP | 3 |
| 1999 | A low resolution pulse position coding method for improved excitation modeling of speech transitionabstractWe propose a new excitation model for transitional speech to reduce the distortion due to the traditional two-excitation source, voiced unvoiced, model. The proposed low resolution pulse position coding (LRPPC) algorithm detects the existence of pulses at frames of weak periodicity, which are determined as unvoiced, and transmits the approximate pulse positions. In the decoder, dispersed pulses that have a flat magnitude spectrum are synthesized at the decoded positions to form the excitation signal. A subjective quality test shows that the vocoder employing the LRPPC algorithm produces better quality of speech, and is very robust to mode decision errors. Jongseo Sohn, Wonyong Sung |
ICASSP | 2 |
| 1999 | An enhanced two-level adaptive multiple branch prediction for superscalar processors
Jongbok Lee, Soo-Mook Moon, Wonyong Sung |
J. Syst. Archit. | 3 |
| 1999 | A statistical model-based voice activity detectionabstractIn this letter, we develop a robust voice activity detector (VAD) for the application to variable-rate speech coding. The developed VAD employs the decision-directed parameter estimation method for the likelihood ratio test. In addition, we propose an effective hang-over scheme which considers the previous observations by a first-order Markov process modeling of speech occurrences. According to our simulation results, the proposed VAD shows significantly better performances than the G.729B VAD in low signal-to-noise ratio (SNR) and vehicular noise environments. Jongseo Sohn, Nam Soo Kim, Wonyong Sung |
IEEE Signal Process. Lett. | 3 |
| 1998 | A Hardware Software Cosimulation Backplane with Automatic Interface GenerationabstractA hardware software cosimulation environment is developed using the backplane approach. This paper defines the backplane protocol for communication and synchronization between client simulators to seamlessly integrate a new simulator without modification. Automatic interface generation facility is also devised for more effective cosimulation environment. The environment is implemented based on Ptolemy and validated with QAM example run on different configurations. Wonyong Sung, Soonhoi Ha |
ASP-DAC | 1 |
| 1998 | Optimized Timed Hardware Software Cosimulation without Roll-backabstractAn optimized hardware software cosimulation method based on the backplane approach is presented in this paper. To enhance the performance of cosimulation, efforts are focused on reducing control packets between simulators as well as concurrent execution of simulators without roll-back. Wonyong Sung, Soonhoi Ha |
DATE | 1 |
| 1998 | A voice activity detector employing soft decision based noise spectrum adaptationabstractIn this paper, a voice activity detector (VAD) for variable rate speech coding is decomposed into two parts, a decision rule and a background noise statistic estimator, which are analysed separately by applying a statistical model. A robust decision rule is derived from the generalized likelihood ratio test by assuming that the noise statistics are known a priori. To estimate the time-varying noise statistics, allowing for the occasional presence of the speech signal, a novel noise spectrum adaptation algorithm using the soft decision information of the proposed decision rule is developed. The algorithm is robust, especially for the time-varying noise such as babble noise. Jongseo Sohn, Wonyong Sung |
ICASSP | 2 |
| 1998 | A synchronization scheme for multi-carrier CDMA systemsabstractA new time-domain synchronization scheme for multi-carrier CDMA systems is developed, which does not require the DFT or FFT of the whole OFDM block. This scheme is particularly proposed for being used with the time-domain data detection techniques which do not need the whole DFT analysis of the input signal (Seunghyeon Nahm and Wonyong Sung, 1996, 1998). Timing detection is performed using the correlation property of the guard interval in the time domain. The carrier frequency synchronization is performed in two steps: tracking to the nearest subcarrier frequency followed by correction of the remaining frequency offset, which is a multiple of the subcarrier frequency interval. The guard interval-based frequency offset detection method is employed for the first step (Daffara and Adami, 1995). A synchronization block containing null symbols on a small number of subcarriers is proposed for the second step, which enables efficient implementation of frequency offset detection using the decimation-in-frequency technique. Seunghyeon Nahm, Wonyong Sung |
ICC | 2 |
| 1998 | Time- and frequency-domain hybrid detection scheme for OFDM-CDMA systemsabstractAn efficient receiver structure for OFDM-CDMA, or MC-CDMA, systems is developed by combining the frequency- and the time-domain detection methods. The frequency-domain detection method uses the transformed signal, thus is efficient for large parallel data transmission, such as in typical OFDM systems, while the time-domain detection method does not transform the signal and uses a correlator receiver, which is efficient for single data transmission with a large processing gain. The developed method generates the decimated-in-frequency (DIF) signal using the small size of the DFT or FFT operations, so that each DIF time-domain signal contains only one data symbol, and uses a correlator receiver for detecting each data symbol. This scheme requires less or an equal number of computations than both methods in any variations of the system parameters, such as the processing gain, the number of parallel data, and the OFDM block size. Seunghyeon Nahm, Wonyong Sung |
ICC | 2 |
| 1998 | Fixed-point error analysis and word length optimization of 8×8 IDCT architecturesabstractComplete fixed-point error models that include the coefficient quantization are derived for two popular 8/spl times/8 two-dimensional (2-D) IDCT architectures; one is based on distributed arithmetic, and the other is the multiplier-adder chain. The error models are evaluated in the integer domain to accurately measure the effects of rounding. The analysis results show that the overall mean-square error performance (OMSE) is the most critical condition for meeting the IEEE specification (IEEE Std. 1180-1990) when the rounding scheme is employed. On the other hand, the mean error effects (OME and PME) are dominant for truncation. Finally, the analysis results are compared with those of bit-accurate simulation. Seehyun Kim, Wonyong Sung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | An Enhanced Two-Level Adaptive Multiple Branch Prediction for Superscalar Processors
Jongbok Lee, Wonyong Sung, Soo-Mook Moon |
Euro-Par | 2 |
| 1997 | Fixed-point C compiler for TMS320C50 digital signal processorabstractA fixed-point C compiler is developed for convenient and efficient programming of TMS320C50 fixed-point digital signal processor. This compiler supports the 'fix' data type that can have an individual integer word-length according to the range of a variable. It can add or subtract two data having different integer word-lengths by automatically inserting shift operations. The accuracy of fixed-point multiply operation is significantly increased by storing the upper part of the multiplied double-precision result instead of keeping the lower part as conducted in the integer multiplication. Several target specific code optimization techniques are employed to improve the compiler efficiency. The empirical results show that the execution speed of a fixed-point C program is much, about an order of magnitude, faster than that of a floating-point C program in a fixed-point digital signal processor. Jiyang Kang, Wonyong Sung |
ICASSP | 2 |
| 1997 | A fast direction sequence generation method for CORDIC processorsabstractThis paper describes a new direction sequence generation method for the circular CORDIC algorithm. A conventional approach employs an angle computation algorithm to control the direction of rotation in the form of a sign sequence, where the sign generation is a bottle-neck for the fast implementations. The proposed method reduces the number of sequential computations by employing a new angle representation model and linearizing the arctangent function in small angles. The direction sequence can be generated by about a third of the iterative computations required in the conventional algorithm, which also reduces the hardware requirements as much. Especially, this algorithm is attractive when pipelining is not allowed for feedback control, such as found in phase tracking applications. A VLSI implementation example for a high-speed quadrature demodulator is also discussed. Seunghyeon Nahm, Wonyong Sung |
ICASSP | 2 |
| 1997 | Adaptive Threshold Error Diffusion Technique for Color Inkjet Printing abstractAn adaptive error diffusion algorithm which exploits the statistical characteristics of the modified input is developed to match the grey-levels of the input and the halftoned images. We adjust the threshold by using the quantization error to reflect the local characteristics and reduce the memory requirements. The proposed algorithm is combined with a color inkjet printer model to compensate for the distortion caused by the dot size differences in each color. Wonyong Sung, Byonghyo Shim |
ICIP (1) | 1 |
| 1995 | An integrated hardware-software cosimulation environment for heterogeneous systems prototypingabstractNo abstract available. Kyuseok Kim, Youngsoo Shin, Taekyoon Ahn, Wonyong Sung, Kiyoung Choi, Soonhoi Ha |
ASP-DAC | 5 |
| 1994 | Word-length determination and scaling software for a signal flow block diagramabstractSoftware for word-length determination and scaling of a general signal flow block diagram is developed for aiding the development of fixed-point digital signal processing algorithms. A fixed-point performance measure, such as signal to quantization noise ratio, and a hardware cost model are used as constraints to the optimization. Simulation using a real input signal is conducted for the evaluation of fixed-point performance. The number of simulations required for the optimization process is greatly reduced by employing netlist preprocessing and efficient search methods.> Wonyong Sung, Ki-Il Kum |
ICASSP (2) | 1 |
| 1992 | Mapping locally recursive SEGs upon a multiprocessor system in a ring networkabstractA multiprocessor code generation method for digital signal processing algorithms represented by SFGs (signal flow graphs) is developed. For reducing the number of communication operations as well as distributing the workload evenly among the processors, a multiprocessor scheduling method based on a parallel block processing scheme, which processes multiple blocks of input data concurrently, is employed. The developed method first divides an SFG into graph segments to reduce the dependency time. A segment merging process is followed, which results less number of temporary data storages and data transfers. A multiprocessor code is generated by applying a single processor code generation method to each of these segments. The implementation result for QR-RLS algorithm using the developed method is included.> Wonyong Sung, Sanjit K. Mitra, Ki-Il Kum |
ASAP | 1 |
| 1992 | Multiprocessor Implementation of Digital Filtering Algorithms Using a Parallel Block Processing MethodabstractAn efficient real-time implementation of digital filtering algorithms using a multiprocessor system in a ring network is investigated. This method is based on a parallel block processing approach, where a continuously supplied input data is divided into blocks, and the blocks are processed concurrently by being assigned to each processor in the system. This approach requires only a simple interconnection network and reduces significantly the number of communications among the processors, making the system easily expandable and highly efficient. In addition, various digital signal processing algorithms can be implemented on the same multiprocessor system. The data dependency of the blocks to be processed concurrently brings on dependency problems between the processors. A systematic scheduling method has been developed by using a precedence graph for the analysis of the dependency relation. Methods for solving the dependency problems between the processors are also investigated. Implementation procedures and results for FIR, recursive, and adaptive filtering algorithms are illustrated.> Wonyong Sung, Sanjit K. Mitra, Branko Jeren |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1987 | Implementation of digital filtering algorithms using pipelined vector processorsabstractThe implementation of digital filtering algorithms using pipelined vector processors is investigated. Modeling of vector processors and vectorization methods are explained, and then the performances of several implementation methods are evaluated based on the model. Vector processor implementation of FIR filtering algorithms using the outer product method and the indirect convolution method is evaluated. Recursive and adaptive filtering algorithms, which lead to dependency problems in direct vector processor implementations, are implemented very efficiently using a newly developed vectorization method. The proposed method computes multiple output samples at a time, making the vector length independent of the filter order. Illustrative examples comparing theoretical results with Cray X-MP simulation results are included. Wonyong Sung, Sanjit K. Mitra |
Proc. IEEE | 1 |
| 1986 | Efficient multi-processor implementation of recursive digital filtersabstractEfficient computation method for the recursive digital filtering is studied in the multi-processor environment. The method solves the dependency problem by separate computations of the particular and transient solutions. The throughput of the algorithm increases linearly with the number of processors, making it possible to increase the throughput effectively by using multiple number of processors. The implementations of the algorithm using a vector-processor and a multiprocessor in a ring network are also studied. Wonyong Sung, Sanjit K. Mitra |
ICASSP | 1 |
| 1985 | Efficient FIR filter design using differential coding of filter coefficientsabstractA general framework for the design of an FIR filter based on the cascade of an FIR pre-filter with quantized coefficients followed by an all pole IIR stage is proposed. A simple design method results by designing the pre-filter stage coefficients via a differential coding scheme. The IIR stage is then simply a linear predictor reconstruction stage. The differential coding method eliminates the redundancies of the filter coefficient sequence and makes it possible to quantize the coefficients into a few levels. A linear prediction algorithm is applied to optimize the differential coding method using the statistical characteristics of the given filter. Relatively wideband filters as well as narrowband filters have been sucessfully designed without interpolating the original filters. Overall design procedures and design examples for a variety of filters are included, and the theoretical aspects of this approach are studied. Wonyong Sung, Sanjit K. Mitra |
ICASSP | 1 |
| 1980 | A 4800 bps LPC vocoder with improved excitationabstractWe present an improved 4800 bps LPC vocoder system that virtually eliminates the buzzy effect from synthetic speech. Excitation signal in the new system is formed by adding high-pass filtered pitch pulses or random noise to a baseband residual signal (0 - 600 Hz) that has been coded by pitch predictive PCM. Since the baseband residual is used as a part of excitation, the system is also robust to V/UV and pitch errors. According to our informal listening tests, the synthetic speech of the new system does not have the buzzy effect. As a result the vocoder speech quality is more natural than that of a conventional LPC vocoder. Chong Kwan Un, Wonyong Sung |
ICASSP | 2 |