Kyung-Tae Kim

dblp:18/6012 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
12since 2021 · last 2024
0000-0003-1200-5282ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Computer networks · 4 · 4 since 2021
YearPublicationVenuePosition
2024 Fusion-Vital: Video-RF Fusion Transformer for Advanced Remote Physiological Measurement
abstract
Remote physiology, which involves monitoring vital signs without the need for physical contact, has great potential for various applications. Current remote physiology methods rely only on a single camera or radio frequency (RF) sensor to capture the microscopic signatures from vital movements. However, our study shows that fusing deep RGB and RF features from both sensor streams can further improve performance. Because these multimodal features are defined in distinct dimensions and have varying contextual importance, the main challenge in the fusion process lies in the effective alignment of them and adaptive integration of features under dynamic scenarios. To address this challenge, we propose a novel vital sensing model, named Fusion-Vital, that combines the RGB and RF modalities through the new introduction of pairwise input formats and transformer-based fusion strategies. We also perform comprehensive experiments based on a newly collected and released remote vital dataset comprising synchronized video-RF sensors, showing the superiority of the fusion approach over the previous single-sensor baselines in various aspects.
Jae-Ho Choi 0004, Ki-Bong Kang, Kyung-Tae Kim
AAAI3
2024 RF-Vital: Radio-Based Contactless Respiration Monitoring for a Moving Individual
abstract
The noncontact respiration rate measurement (nRRM) method allows a system to monitor the breathing patterns of an individual without physical contact, which is crucial for regular health monitoring. Current nRRM approaches primarily depend on detecting minor variations in RGB profiles reflected from a camera to remotely extract respiration signals. However, these methods require continuous pixel-level tracking, which restricts their use on individuals in quasi-stationary sitting positions. To address this limitation, we propose a radiofrequency (RF)-Vital model, which leverages RF signals to extend the applicability of nRRM methods to individuals who exhibit global motions (GMs) and even walk around. The core idea of the RF-Vital model lies in the unique characteristics of RF signals: the RF signals received from a moving individual capture both their respiratory motions (RMs) and GMs through linear superposition while simultaneously providing the reflections of GM alone. To fully utilize such unique properties, we introduce a new RF modality that allows stable inclusion of micro-level respiration signatures, even when GMs are present. Additionally, we optimize the RF-Vital model using a novel multitask adversarial learning framework combined with a new loss function, which facilitates the direct mapping of the desired RMs as well as the self-supervised removal of GMs, thereby effectively filtering out RMs from mixtures of GMs and RMs. The proposed RFvital model was evaluated using newly published data sets. It demonstrated state-of-the-art performance in static conditions and achieved the significant milestone of enabling nRRM under moving conditions.
Jae-Ho Choi 0004, Ki-Bong Kang, Kyung-Tae Kim
IEEE Internet Things J.3
2024 Multipath Signal Mitigation for Indoor Localization Based on MIMO FMCW Radar System
abstract
Indoor human localization (IHL) is an important application of the Internet of Things. Multipath effects are the most challenging problem for successful IHL when using radar, and thus, they should be mitigated properly. Therefore, effective solutions for indoor multipath mitigation are devised in this article based on a co-located multiple-input multiple-output (MIMO) frequency-modulated continuous-wave (FMCW) radar without a priori information of reflection geometry and training data sets. To this end, a velocity–azimuth domain suitable for false alarm reduction under severe multipath propagations rather than conventional domains is suggested based on an in-depth analysis of radar echoes from the indoor environment. Then, intra- and inter-frame integration are exploited to enhance the human signal-to-interference-plus-noise ratio, which helps improve the human detection probability. Finally, the detected signal is identified as a multipath based on whether the direction-of-arrival and direction-of-departure are different. The proposed framework is demonstrated through experiments conducted in a seminar room using a commercial MIMO FMCW radar. The results indicate that the multipath components are explicitly discriminated from the backscattering of an individual using the proposed scheme. Ghost targets induced by multipath can be suppresed owing to the proposed framework, and therefore robust IHL can be achieved.
JeongKi Park, Kyung-Tae Kim
IEEE Internet Things J.3
2024 Radar-Based Crowd Counting in Real-World Environments With Spatiotemporal Transformer
abstract
With the advent of deep learning (DL) for signal processing, the deployment of DL for radar-based crowd counting has yielded significant performance enhancement. Despite these advancements, current methodologies predominantly undergo validation in controlled conditions with limited subject movement variability, posing a challenge for practical usage. Addressing this gap, this letter first attempts the application of radar-based crowd counting in an unregulated and dense setting, capturing the radar reflections of up to 31 subjects in real-world scenarios, such as queues at restaurant kiosks. Furthermore, to address the complexities of such a challenging condition, we introduce a novel radar crowd counting model that utilizes a spatiotemporal transformer. The expremental results demonstrate the potentiality of the proposed model as a robust crowd counting system under the full realistic scenarios, as well as establish its superiority over the conventional radar-based crowd counting models.
Jae-Ho Choi 0004, Kyung-Tae Kim
IEEE Signal Process. Lett.2
2023 Range Resolution Improvement Using Cross-Track Interferometry for FMCW Radar
abstract
The conventional range resolution improvement (RRI) technique combines two spectra of two received pulses with linear frequency modulation in the range-frequency domain. However, its application to frequency-modulation continuous-wave (FMCW) radar is not straightforward. This study proposes an RRI method for two FMCW radars using cross-track interferometry, which are placed at different positions along the cross-track axis. Through the combination of the deramped signals of the two FMCW radars with the RRI, an improved high-resolution range profile was obtained. Furthermore, the simulation and experimental results validated the effectiveness of the proposed method.
Inhyeok Lee, Min-Gon Cho, Kyung-Tae Kim
IEEE Geosci. Remote. Sens. Lett.4
2022 Remote Respiration Monitoring of Moving Person Using Radio Signals
Jae-Ho Choi 0004, Ki-Bong Kang, Kyung-Tae Kim
ECCV (37)3
2022 Deep Learning Approach for Radar-Based People Counting
abstract
With the development of deep learning (DL) frameworks in the field of pattern recognition, DL-based algorithms have outperformed handcrafted feature (HF)-based ones in various applications. However, there still exist several challenges in applying the DL framework to a radar-based people counting (RPC) task: The powerful representation capacity of a deep neural network (DNN) learns not only the desired human-induced components but also unwanted nuisance factors, and available data for RPC is usually insufficient to train a huge-sized DNN, leading to an increased possibility of overfitting. To tackle this problem, we propose novel solutions for the successful application of the DL framework to the RPC task from various perspectives. First, we newly formulate the preprocessing pipelines to transform the raw received radar echoes into a better-matched form for a DNN. Second, we devise a novel backbone architecture that reflects the spatiotemporal characteristics of the radar signals, while relieving the burden on training through a parameter efficient design. Finally, an unsupervised pretraining process and a newly defined loss function are proposed for further stabilized network convergence. Several experimental results using real measured data show that the proposed scheme enables an effective utilization of DL for RPC, achieving a significant performance improvement compared to conventional RPC methods.
Jae-Ho Choi 0004, Kyung-Tae Kim
IEEE Internet Things J.3
2022 Direction Finding for Multiple Wideband Chirp Signal Sources using Blind Signal Separation and Matched Filtering
Min Kim 0002, Seong-Hyeon Lee, In-Oh Choi, Kyung-Tae Kim
Signal Process.4
2022 Fusion of Target and Shadow Regions for Improved SAR ATR
abstract
Synthetic aperture radar (SAR) systems, which operate under a slant-viewing geometry, inevitably entail shadow regions in the resulting radar image. Such shadow profiles contain backprojected signatures of an object’s configuration as with target profiles; however, they are rarely utilized in current SAR-based recognition techniques. A major challenge in leveraging shadow information together lies in the intrinsic limitation of current single-pathway approaches, in which the target and shadow cannot be addressed simultaneously because of their incompatible domain properties. Hence, we herein propose novel solutions that enable the successful fusion of target and shadow regions within SAR for the first time. First, we devise new image preprocessing techniques specifically customized for shadows to compensate for their unique domain characteristics, which are distinct from the target. Second, we introduce a parallelized SAR processing mechanism such that a network can independently extract features oriented toward each conflicting modality. Third, adaptive fusion strategies are proposed for the optimal integration of features from each region while considering their relative significance layer by layer. Extensive experiments on public benchmark datasets demonstrate that the proposed framework allows a network to effectively employ shadow signatures and targets, thereby outperforming previous methods significantly for all setups.
Jae-Ho Choi 0004, Myung-Jun Lee, Nam-Hoon Jeong, Kyung-Tae Kim
IEEE Trans. Geosci. Remote. Sens.5
2022 Length Prediction of Moving Vehicles Using a Commercial FMCW Radar
abstract
The classification of moving vehicles on roadways is an important application of intelligent transportation systems. For successful vehicle classification using radar, suitable feature extraction with high accuracy from the returned echoes of moving vehicles is necessary. In this paper, a novel framework for predicting the length of each moving vehicle is presented based on two different scenarios of the frequency-modulated continuous-waveform (FMCW) radar, i.e., the radar is stationary or moving. Different clutter reduction techniques were considered depending on the motion of the FMCW radar. Using clutter-free signals, moving vehicles are separated on a range-angle map, rather than on a range-Doppler (R-D) map, enabling the successful separation whether the vehicles are overlapping or not on the R-D map. Finally, the scattering center information across multiple frames, rather than a single frame, is fused to significantly increase the accuracy of the predicted length. The proposed framework was demonstrated through experiments in an actual road environment using a commercial FMCW radar. The results indicated a good match with the true lengths of the vehicles.
JeongKi Park, In-Oh Choi, Kyung-Tae Kim
IEEE Trans. Intell. Transp. Syst.3
2021 People Counting Using IR-UWB Radar Sensor in a Wide Area
abstract
Conventional radar-based people counting systems are designed mainly for dense spatial distributions in a small region of interest (ROI). Therefore, a system with only conventional energy-based features, which are effective for a small ROI with a limited spatial distribution of individuals, generally fails to cope with the diverse and complex spatial distributions that arise from the freer movements of individuals as the ROI widens. To address this problem, a novel approach that achieves robust people counting in both wide and small ROIs is presented in this study. The proposed technique incorporates modified CLEAN-based features in the range domain and energy-based features in the frequency domain to efficiently address both dense and dispersed distributions of individuals. Subsequently, principal component analysis and an appropriate normalization of the proposed features are performed for improving the people counting system further. Based on several experiments in practical environments with wide ROIs and severe multipath effects, we observed that the proposed approach yields significantly improved performance compared with traditional people counting systems.
Jae-Ho Choi 0004, Kyung-Tae Kim
IEEE Internet Things J.3
2021 Reduction of False Alarm Rate in SAR-MTI Based on Weighted Kurtosis
abstract
Moving target indication (MTI) is considered one of the most important applications of synthetic aperture radar (SAR) in military operations and traffic monitoring. Although many studies have focused on detectors based on the statistical clutter models, a mismatch between the measured data and the applied statistical model often leads to unreliable MTI results, particularly in terms of false alarm rates. To reduce the number of false alarms, we propose an efficient MTI scheme consisting of conventional MTI techniques (the displaced phase center antenna (DPCA) and the interferogram's magnitude and phase (IMP) methods), which operate in dual-receive antenna (DRA) mode in an SAR system. These techniques are coupled with a new discrimination stage based on a new detection metric-weighted kurtosis. Using simulated and real measured data from TerraSAR-X, the proposed MTI scheme demonstrates robust performance in the reduction of false alarm rates with only a slight increase in the computation time compared with conventional schemes.
Myung-Jun Lee, Seungjae Lee 0001, Bo-Hyun Ryu, Byoung-Gyun Lim, Kyung-Tae Kim
IEEE Trans. Geosci. Remote. Sens.5
2017 ISAR Imaging of High-Speed Maneuvering Target Using Gapped Stepped-Frequency Waveform and Compressive Sensing
abstract
In the case of a stepped-frequency waveform (SFW) inverse synthetic aperture radar (ISAR) system, the translational motion (TM) of a target can be usually divided into two parts: 1) target motion within a pulse repetition interval, called the inter-pulse translational motion (IPTM) and 2) target motion between bursts, called the inter-burst translational motion (IBTM). The former induces severe blurring in the ISAR images as well as range-compressed data (i.e., range profile), and the latter also causes dramatic degradation of the ISAR image quality. In this paper, a novel framework for high-resolution gapped SFW (GSFW) ISAR imaging of high-speed maneuvering target is proposed. The main novelty of the proposed method is twofold: 1) accurate TM parameter estimation in conjunction with a compressive sensing theory using a newly devised cost function and particle swarm optimization and 2) compensation for both the IPTM and IBTM phase errors simultaneously even with the GSFW data set. Simulation results using ideal point scatterers show that the proposed method is capable of precise reconstruction of ISAR image and accurate TM parameter estimation. Experimental results using real measured data verify the robustness and the effectiveness of the proposed method.
Seungjae Lee 0001, Seong-Hyeon Lee, Kyung-Tae Kim
IEEE Trans. Image Process.4
2015 Performance of Sparse Recovery Algorithms for the Reconstruction of Radar Images From Incomplete RCS Data
abstract
In this letter, we compare the performances of sparse recovery algorithms (SRAs) for the reconstruction of a 2-D inverse synthetic aperture radar (ISAR) image from incomplete radar-cross-section (RCS) data. The three methods considered for the SRA include the basis pursuit (BP), the BP denoising, and the orthogonal matching pursuit methods. The performances of the methods in terms of the reconstruction accuracy of the ISAR image are compared using the incomplete RCS data. In addition, traditional interpolation methods such as nearest-neighbor interpolation, linear interpolation, and spline interpolation are applied to the incomplete RCS data to reconstruct ISAR images, and their performances are compared to that of the SRAs.
Ji-Hoon Bae 0001, Byung-Soo Kang, Kyung-Tae Kim, Eunjung Yang
IEEE Geosci. Remote. Sens. Lett.3
2013 New Discrimination Features for SAR Automatic Target Recognition
abstract
We propose new features and a redundancy-free feature selection scheme for discriminating targets from clutter in high-resolution synthetic aperture radar imagery. These were found to be very useful for discriminating targets from clutter and well combined with various classical discriminative features. If a feature set with a larger dimension is used, its redundancy may increase. In order to reduce this redundancy, a rank-based feature selection scheme is devised. Our discriminating features and feature selection algorithm are evaluated using the moving and stationary target acquisition and recognition (MSTAR) data set.
Jong-Il Park, Sang-Hong Park, Kyung-Tae Kim
IEEE Geosci. Remote. Sens. Lett.3
2011 Cross-Range Scaling Algorithm for ISAR Images Using 2-D Fourier Transform and Polar Mapping
abstract
This paper proposes a new method that solves the problem of inverse synthetic aperture radar image cross-range scaling by estimating the rotational velocity (RV) using the expansion-rotation-scale relationship between two range-Doppler (RD) images. This method is composed of three steps. The first step is preprocessing to construct 2-D Fourier transform images and initial polar-mapped images. In this step, two RD images are 2-D Fourier transformed to avoid the necessity of finding the rotation center; then, the transformed images are polar mapped with identical polar grids to convert rotation into translation in the θ-direction only. The second step is a coarse search that finds the angular shift that provides the maximum correlation between two polar images. The angular shift found is used as the initial relative rotation angle (RA), and the initial relative scaling factor (RSF) is calculated using the RV which is equal to the angular shift divided by the time delay between the images. The third step is the optimization of the relative RA using the Nelder-Mead approach, with the RSF updated using the relative RA derived during each iteration. In simulations using a Mig-25 aircraft, composed of ideal point scatterers, and the measured data from a Boeing 747-400, the targets were properly rescaled in the range-cross-range domain due to the accurate estimation of the RV.
Sang-Hong Park, Hyo-Tae Kim, Kyung-Tae Kim
IEEE Trans. Geosci. Remote. Sens.3
2010 Robust automatic speech recognition with decoder oriented ideal binary mask estimation
abstract
In this paper, we propose a joint optimal method for automatic speech recognition (ASR) and ideal binary mask (IBM) estimation in transformed into the cepstral domain through a newly derived generalized expectation maximization algorithm. First, cepstral domain missing feature marginalization is established using a linear transformation, after tying the mean and variance of non-existing cepstral coefficients. Second, IBM estimation is formulated using a generalized expectation maximization algorithm directly to optimize the ASR performance. Experimental results show that even in highly non-stationary mismatch condition (dance music as background noise), the proposed method achieves much higher absolute ASR accuracy improvement ranging from 14.69 % at 0 dB SNR to 40.10 % at 15 dB SNR compared with the conventional noise suppression method. Index Terms: robust speech recognition, ideal binary mask classification, missing feature 1.
Lae-Hoon Kim, Kyung-Tae Kim, Mark Hasegawa-Johnson
INTERSPEECH2
2010 Segmentation of ISAR Images of Targets Moving in Formation
abstract
This paper proposes a new method to separate inverse synthetic aperture radar (ISAR) images using the range-Doppler algorithm when multiple targets fly closely spaced in a formation. This method is composed of three steps. The first step is an initial range-Doppler imaging of the whole set of targets to obtain a ¿bulk¿ image. In this step, range profiles are aligned using a new cost function, which is the sum of the amplitudes of the pixels lying along a polynomial that models the trajectory of the bulk image. To reduce the probability of incorrectly aligning high-amplitude pixels from different targets, pixel amplitudes are converted to binary form: A number ¿ of pixels with the greatest amplitudes are identified in the bulk range profile, and their amplitudes are converted to one; the amplitudes of the remaining pixels are converted to zero. This process gives each of the ¿ pixels the same weight on the cost function. Then, phase adjustment is used to coarsely separate targets in a 2-D image. The second step is the separation of the bulk image into component images using a window of the target size. The third step is a second range-Doppler imaging, in which each ISAR image is enhanced using range alignment and phase adjustment. Simulations using three targets composed of point scattering centers prove that the proposed method can effectively segment three targets flying in a formation.
Sang-Hong Park, Hyo-Tae Kim, Kyung-Tae Kim
IEEE Trans. Geosci. Remote. Sens.3
2008 Speech Bandwidth Extension Using Temporal Envelope Modeling
abstract
Speech bandwidth extension (SBE) assumes that high-frequency components of a speech signal, e.g., the frequency band of 47 kHz, can be estimated by parameters extracted from the narrowband signal (04 kHz). Therefore, it is very important to understand the characteristics of the highband signal as well as perceptual cues to represent the highband signal. This letter proposes a new SBE algorithm using a temporal envelope model. The temporal envelope model considers band-limited temporal envelopes as the perceptual cue of the 47 kHz band signal while it deemphasizes the importance of rapidly varying components. To implement the SBE with no additional bits, the proposed method adopts a Gaussian mixture model (GMM) to estimate the temporal envelope of the highband signal from that of the narrowband one. Simulation results confirm that the proposed SBE algorithm shows better perceptual quality than a conventional source-filter model-based approach.
Kyung-Tae Kim, Min-Ki Lee, Hong-Goo Kang
IEEE Signal Process. Lett.1
2007 Speech quality estimation using packet loss effects in CELP-type speech coders
Min-Ki Lee, Kyung-Tae Kim, Hong-Goo Kang, Dae Hee Youn
INTERSPEECH2
2005 A fast adaptive-codebook search algorithm for G.723.1 speech coder
abstract
This letter presents a new fast search algorithm for the multitap adaptive codebook used in the G.723.1 standard speech coder. In contrast with the standard method that a closed-loop pitch lag and gains for a fifth-order pitch predictor are searched simultaneously, the proposed algorithm adopts a sequential and restricted approach to determine the parameters. In other words, the proposed scheme first determines a couple of pitch lag candidates using a first-order pitch predictor and then computes the pitch gains of the fifth-order predictor within a restricted search area. Experimental results confirm that the proposed algorithm reduces the total complexity by 30.69% in the encoding process and provides speech quality equivalent to the standard method.
Sung-Kyo Jung, Kyung-Tae Kim, Young-Cheol Park, Hong-Goo Kang
IEEE Signal Process. Lett.2
2004 A bit-rate/bandwidth scalable speech coder based on ITU-T G.723.1 standard
abstract
The paper presents a new scalable coder based on the ITU-T G.723.1 standard which is one of the most famous speech coders for VoIP applications. In order to support both bit-rate scalability and bandwidth scalability, the proposed coder adopts a split-band approach, where the input signal, sampled at 16 kHz, is decomposed into two equal frequency bands. The lower-band speech is coded with a standard coder, such as the G.723.1 standard. In addition, the low-band enhancement layer for lower-band speech improves the perceptual quality of decoded speech by employing additional coding units based on a cascaded codebook approach. The higher-band signal is encoded using an MDCT-based transform coding scheme. The proposed coder, at a bit-rate of 19.4 kbit/s, provides speech quality comparable to the ITU-T 24 kbit/s G.722.1 coder, while it also has interoperability with G.723.1.
Sung-Kyo Jung, Kyung-Tae Kim, Hong-Goo Kang
ICASSP (1)2
2004 Temporal normalization techniques for transform-type speech coding and application to split-band wideband coders
abstract
In this paper we present an efficient coding method for the upper band(4-7kHz) of wideband(0.5-7kHz) speech coding based on a band-split approach. Due to the impulselike characteristics in upper band signal, it is very difficult to efficiently quantize the signal at low bit-rate when we use transform coding techniques. We propose two temporal normalization techniques, direct temporal energy normalization and frequency domain linear prediction, to reduce the extremely noticeable artifacts. Simulation results show that the proposed algorithm successfully encodes the upper band signal, and the new split-band type wideband coder adopting the proposed technology provides better quality than 56 kbit/s ITU-T G.722 at the bitrate of 20 kbit/s.
Kyung-Tae Kim, Sung-Kyo Jung, MiSuk Lee, Hong-Goo Kang, Dae Hee Youn
INTERSPEECH1
2002 A new bandwidth scalable wideband speech/audio coder
abstract
In this paper, we present a new bandwidth-scalable coder for wide band speech and audio signals. The proposed coder splits 8 kHz signal bandwidth into two narrow bands, and different coding schemes are applied to each band. The lower-band speech is coded with ITU-T G.729 Annex E, and the higher-band signal is compressed using a new algorithm based on the gammatone filter band with an invertible auditory model. Due to the split-band architecture and completely independent coding schemes for each band, the output speech of the decoder can be selected to be a narrowband or wideband coding according to the channel conditions. Subjective tests showed that, for wideband speech and audio signals, the proposed coder at 18 kbit/s produces superior quality to ITU-T 24 kbit/s G.722.1 with the shorter algorithmic delay.
Kyung-Tae Kim, Sung-Kyo Jung, Young-Cheol Park, Dae Hee Youn
ICASSP1
2001 Efficient implementation of ITU-t g.723.1 speech coder for multichannel voice transmission and storage
abstract
Dual-rate G.723.1 speech coder has been widely applied to real-time video and teleconferencing applications where reduced bandwidth and good voice quality is required. This paper presents an efficient implementation of G.723.1 speech coder. To simplify the excitation quantization procedure which is the most computationally demanding, we propose fast algorithms for adaptive codebook and fixed codebook search. In the fast adaptive codebook search, pitch delay and pitch gains are computed sequentially. In the fast fixed codebook search, the codebook structure is redesigned based on the interleaved single-pulse permutation (ISPP) design at high rate mode and the depth-first tree search is applied instead of nested-loop search at low rate mode. A real-time implementation is achieved using a 16-bit fixed-point TMS320C62x DSP. The implemented G.723.1 speech coder requires 8.70 and 10.29 MHz clock cycles at low and high rate, respectively, 57.8 kByte of program memory and 55 kByte of data memory. Thus, more than 16 channels of G.723.1 coder can be operated in real-time using a single TMS320C62x DSP.
Sung-Kyo Jung, Young-Cheol Park, Sung-Wan Yoon, Kyung-Tae Kim, Dae Hee Youn
INTERSPEECH4
2001 An efficient transcoding algorithm for G.723.1 and EVRC speech coders
abstract
Interoperability is one the most important factors for a successful integration of the speech network. To operate speech networks employing different speech coders, but integrated as one, bitstreams generated by one coder should be translated seamlessly to those of the other coders. Connecting two coders in tandem may be the simplest way to accomplish this. However, coders in tandem connection often produce problems such as poor speech quality, high computational load, and additional transmission delay. We propose an efficient transcoding algorithm that can provide interoperability to the networks employing ITU-T G.723.1 and TIA IS-127 EVRC speech coders. Subjective and objective quality evaluation have confirmed that the speech quality produced by the proposed transcoding algorithm was equivalent to, or better than the tandem coding, while it had shorter processing delay and less computational complexity.
Kyung-Tae Kim, Sung-Kyo Jung, Young-Cheol Park, Yong-Soo Choi, Dae Hee Youn
VTC Fall1
1994 Generation of multi-syllable nonsense words for the assessment of Korean text-to-speech system
Cheol-Woo Jo, Kyung-Tae Kim
ICSLP2
1990 Construction of a large Korean speech database and its management system in ETRI
Joon-Hyuk Choi, Kyung-Tae Kim
ICSLP2
1990 Speech synthesis using demisyllables for Korean: a preliminary system
Jung-Chul Lee, Hee-il Han, Eung-Bae Kim, Chang-Joo Kim, Kyung-Tae Kim
ICSLP6