EDBT 2026 Demo / reviewers in the wild / expert
Ki-Seung Lee
dblp:51/5655
· DBLP profile ↗
18ranked-venue papers
17as first author
0since 2021 · last 2020
0000-0001-6723-6913ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 12 first-authorArtificial intelligence and machine learning · 7 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Audio and music processing · 93% Image and video processing · 7% | |
| Artificial intelligence
2 papers |
Speech recognition and synthesis · 100% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › spatial audio › binaural reproduction
crosstalk cancellation |
0.2 | 1 | 2013 | Position-Dependent Crosstalk Cancellation Using Space Partitioning · IEEE Trans. Speech Audio Process. 2013 |
Audio and music processing
spatial audio |
0.2 | 1 | 2013 | Position-Dependent Crosstalk Cancellation Using Space Partitioning · IEEE Trans. Speech Audio Process. 2013 |
Audio and music processing › spatial audio
head-related transfer function |
0.1 | 1 | 2011 | A Relevant Distance Criterion for Interpolation of Head-Related Transfer Functions · IEEE Trans. Speech Audio Process. 2011 |
Audio and music processing › spatial audio
spatial audio generation |
0.1 | 1 | 2011 | A Relevant Distance Criterion for Interpolation of Head-Related Transfer Functions · IEEE Trans. Speech Audio Process. 2011 |
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis |
0.1 | 1 | 2006 | MLP-based phone boundary refining for a TTS database · IEEE Trans. Speech Audio Process. 2006 |
Natural language and speech › Speech recognition and synthesis › speech coding
low-bit-rate speech coding |
0.0 | 1 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 |
Natural language and speech › Speech recognition and synthesis
speech coding |
0.0 | 1 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 |
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
unit selection |
0.0 | 1 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 |
Image and video processing › image restoration
denoising |
0.0 | 1 | 1999 | Image enhancement based on signal subspace approach · IEEE Trans. Image Process. 1999 |
Image and video processing
image enhancement |
0.0 | 1 | 1999 | Image enhancement based on signal subspace approach · IEEE Trans. Image Process. 1999 |
Audio and music processing › speech enhancement
noise reduction |
0.0 | 1 | 1999 | Image enhancement based on signal subspace approach · IEEE Trans. Image Process. 1999 |
Audio and music processing › acoustic signal processing › audio signal reconstruction › audio restoration
subspace-based noise reduction |
0.0 | 1 | 1999 | Image enhancement based on signal subspace approach · IEEE Trans. Image Process. 1999 |
Natural language and speech › Speech recognition and synthesis › speech corpus
speech corpus construction |
0.0 | 1 | 2006 | MLP-based phone boundary refining for a TTS database · IEEE Trans. Speech Audio Process. 2006 |
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
concatenative speech synthesis |
0.0 | 1 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.0 | 1 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 |
Methods — techniques the papers use, named apart from their topics
space partitioning · 0.2artificial neural network · 0.2mel-cepstral distortion · 0.1interpolation · 0.1ROC curves · 0.1multi-layer perceptron · 0.1hidden markov model · 0.1rate-distortion piecewise linear approximation · 0.0harmonic+noise model · 0.0signal subspace method · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Joint Audio-Ultrasound Food Recognition for Noisy EnvironmentsabstractContinuous recognition of ingested foods without user intervention is very useful for the pre-screening of obesity and diet-related disease. An automatic food recognition method that combines the two modalities of audio and ultrasonic signals (US) is proposed in this study. Under a noise-free environment, classification accuracy of an audio-only recognizer is generally higher than that of US-only recognizers, but the performance of US recognizers is unaffected by acoustic noise levels. In the recognition system presented herein, the likelihood score of the audio-US feature was given by a linear combination of class-conditional observation log-likelihoods for two classifiers, using the appropriate weights. We developed a weighting process adaptive to signal-to-noise ratios (SNRs). The main objective here involves determining the optimal SNR classification boundaries and constructing a set of optimum stream weights for each SNR class. A feasibility test was conducted to verify the usefulness of the proposed method by conducting recognition experiments on seven types of food. The performance was compared with conventional methods that use in-ear and throat microphones. The proposed method yielded remarkable levels of recognition performance of 90.13% for artificially added noise and 89.67% under actual noisy environments, when the SNR ranged from 0 to 20 dB. Ki-Seung Lee |
IEEE J. Biomed. Health Informatics | 1 |
| 2019 | Speech enhancement using ultrasonic doppler sonar
Ki-Seung Lee |
Speech Commun. | 1 |
| 2017 | Restricted Boltzmann Machine-Based Voice Conversion for Nonparallel CorpusabstractA large amount of parallel training corpus is necessary for robust, high-quality voice conversion. However, such parallel data may not always be available. This letter presents a new voice conversion method that needs no parallel speech corpus, and adopts a restricted Boltzmann machine (RBM) to represent the distribution of the spectral features derived from a target speaker. A linear transformation was employed to convert the spectral and delta features. A conversion function was obtained by maximizing the conditional probability density function with respect to the target RBM. A feasibility test was carried out on the OGI VOICES corpus. Results from the subjective listening tests and the objective results both showed that the proposed method outperforms the conventional GMM-based method. Ki-Seung Lee |
IEEE Signal Process. Lett. | 1 |
| 2014 | A unit selection approach for voice transformation
Ki-Seung Lee |
Speech Commun. | 1 |
| 2013 | Position-Dependent Crosstalk Cancellation Using Space PartitioningabstractThe present study tested a new stereo playback system that effectively cancels cross-talk signals at an arbitrary listening position. Such a playback system was implemented by integrating listener position tracking techniques and crosstalk cancellation techniques. The entire listening space was partitioned into a number of non-overlapped cells and a crosstalk cancellation filter was assigned to each cell. The listening space partitions and the corresponding crosstalk cancellation filters were constructed by maximizing the average channel separation ratio (CSR). Since the proposed method employed cell-based crosstalk cancellation, estimation of the exact position of the listener was not necessary. Instead, it was only necessary to determine the cell in which the listener was located. This was achieved by simply employing an artificial neural network (ANN) where the time delay to each pair of microphones was used as the ANN input and the ANN output corresponded to the index of cells. The experimental results showed that more than 95% of the experimental listening space had a CSR ≥ 10 dB when the number of clusters exceeded 12. Under these conditions, the correlation between the true directions of the virtual sound sources and the directions recognized by the subjects was greater than 0.9. Ki-Seung Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | A Relevant Distance Criterion for Interpolation of Head-Related Transfer FunctionsabstractIn binaural synthesis, in order to realize more precise and accurate spatial sound, it would be desirable to measure a large number of the head-related transfer functions (HRTFs) in various directions. To reduce the size of the HRTFs, interpolation is often employed, where the HRTF for any direction can be obtained by a limited number of the representative HRTFs. In this paper, it is determined which distortion measure for interpolation of the HRTFs in the horizontal plane is most suitable for predicting audible differences in sound location. Four kinds of HRTF sets, measured using three human heads and one mannequin (KEMAR), were prepared for this study. Using various objective distortion criteria, the differences between interpolated and measured HRTFs were computed. These were then related to the results from the listening tests through receiver operator characteristic (ROC) curves. The results of the present study indicated that for the HRTF sets measured from three human heads, the best predictor of performance was obtained using the distortion measurement computed from the mel-cepstral coefficients, whereas the distortion measurement associated with interaural time delay predicted audible differences in sound location reasonably well for the KEMAR HRTF set. A feasibility test was conducted to verify the usefulness of the selected distortion measurement. Ki-Seung Lee, Seok-Pil Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | MLP-based phone boundary refining for a TTS databaseabstractThe automatic labeling of a large speech corpus plays an important role in the development of a high-quality Text-To-Speech (TTS) synthesis system. This paper describes a method for the automatic labeling of speech signals, which mainly involves the construction of a large database for a TTS synthesis system. The main objective of the work involves the refinement of an initial estimation of phone boundaries which are provided by an alignment, based on a Hidden Markov Model. A multilayer perceptron (MLP) was employed to refine the phone boundaries. To increase the accuracy of phoneme segmentation, several specialized MLPs were individually trained based on phonetic transition. The optimum partitioning of the entire phonetic transition space and the corresponding MLPs were constructed from the standpoint of minimizing the overall deviation from the hand-labeling position. The experimental results showed that more than 93% of all phone boundaries have a boundary deviation from a reference position smaller than 20 ms. We also confirmed that the database constructed using the proposed method produced results that were perceptually comparable to a hand-labeled database, based on subjective listening tests. Ki-Seung Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Temporal decomposition based on a rate-distortion criterionabstractThis letter addresses a temporal decomposition (TD) technique that is based on a rate-distortion criterion. In the proposed TD scheme, a set of interpolation functions is constructed from a given training corpus, and the optimum target points are found in the sense of minimizing, not only spectral distortion, but also bit rates. The results of the simulation show that an average spectral distortion of about 1.4 dB can be achieved at an average bit rate of about 8 bits/frame. Ki-Seung Lee |
IEEE Signal Process. Lett. | 1 |
| 2003 | Context-adaptive phone boundary refining for a TTS databaseabstractA method for the automatic segmentation of speech signals is described. The method is dedicated to the construction of a large database for a Text-To-Speech (TTS) synthesis system. The main issue of the work involves the refinement of an initial estimation of phone boundaries which are provided by an alignment, based on a Hidden Markov Model (HMM). Multi-layer perceptron (MLP) was used as a phone boundary detector. To increase the performance of segmentation, a technique which individually trains an MLP according to phonetic transition is proposed. The optimum partitioning of the entire phonetic transition space is constructed from the standpoint of minimizing the overall deviation from hand labelling positions. With single speaker stimuli, the experimental results showed that more than 95% of all phone boundaries have a boundary deviation from the reference position smaller than 20 ms, and the refinement of the boundaries reduces the root mean square error by about 25%. Ki-Seung Lee, Jeongsu Kim |
ICASSP (1) | 1 |
| 2002 | A segmental speech coder based on a concatenative TTS
Ki-Seung Lee, Richard V. Cox |
Speech Commun. | 1 |
| 2002 | Context-adaptive smoothing for concatenative speech synthesisabstractIn text-to-speech synthesis, spectral smoothing is often employed to reduce artifacts at unit-joining points. A context-adaptive smoothing method is proposed in this letter, where the amount of smoothing is determined according to context information. Discontinuities at unit boundaries are predicted by a regression tree, and smoothing factors are computed by using predicted discontinuities and real discontinuities at unit boundaries. Experimental results are presented to demonstrate the effectiveness of the proposed method. Ki-Seung Lee, Sang-Ryong Kim |
IEEE Signal Process. Lett. | 1 |
| 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigmabstractPrevious studies have shown that a concatenative speech synthesis system with a large database produces more natural sounding speech. We apply this paradigm to the design of improved very low bit rate speech coders (sub 1000 b/s). The proposed speech coder consists of unit selection, prosody coding, prosody modification and waveform concatenation. The encoder selects the best unit sequence from a large database and compresses the prosody information. The transmitted parameters include unit indices and the prosody information. To increase naturalness as well as intelligibility, two costs are considered in the unit selection process: an acoustic target cost and a concatenation cost. A rate-distortion-based piecewise linear approximation is proposed to compress the pitch contour. The decoder concatenates the set of units, and then synthesizes the resultant sequence of speech frames using the harmonic+noise model (HNM) scheme. Before concatenating units, prosody modification which includes pitch shifting and gain modification is applied to match those of the input speech. With single speaker stimuli, a comparison category rating (CCR) test shows that the performance of the proposed coder is close to that of the 2400-b/s MELP coder at an average bit rate of about 800-b/s during talk spurts. Ki-Seung Lee, Richard V. Cox |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Corpus-based techniques in the AT&t nextgen synthesis systemabstractThe AT&T text-to-speech (TTS) synthesis system has been used as a framework for experimenting with a perceptuallyguided data-driven approach t o s p e e c h s y n thesis, with primary focus on data-driven elements in the \back end".Statistical training techniques applied to a large corpus are used to make decisions about predicted speech e v ents and selected speech i n ventory units.Our recent a d v ances in automatic phonetic and prosodic labeling and a new faster harmonic plus noise model (HNM) and unit preselection implementations have signi cantly improved TTS quality and speeded up both development time and runtime. Ann K. Syrdal, Colin W. Wightman, Alistair Conkie, Yannis Stylianou, Marc C. Beutnagel, Juergen Schroeter, Volker Strom, Ki-Seung Lee, Matthew J. Makashay |
INTERSPEECH | 8 |
| 1999 | TTS based very low bit rate speech coderabstractThis paper addresses a speech coder which uses a text-to-speech (TTS) synthesis system to achieve very low bit rates (sub 1 kbps). The main issue of the work is the accurate coding of the pitch (f/sub 0/) and gain contours which are principle components of prosody. This is of paramount interest since the correct prosody will increase naturalness and an efficient coding scheme will provide high coding gain. Together with the phonetic transcription, the f/sub 0/ and gain contour constitute the parameters that are necessary for the TTS system to synthesize the speech signal. Piecewise linear approximation is used to code the f/sub 0/ parameter. A technique which minimizes the bit rate while maintaining f/sub 0/ error below a given threshold are described. To obtain both high compression and smoothly changing gain contours, the variance of the signal is averaged over each half phoneme length is transmitted as gain information. With single speaker stimuli, and a priori text transcription information, we obtained natural sounding speech at an average bit rate of about 300 bps. Ki-Seung Lee, Richard V. Cox |
ICASSP | 1 |
| 1999 | Image enhancement based on signal subspace approachabstractThis paper describes a block-by-block basis image enhancement algorithm which uses the signal subspace method to enhance images corrupted by uncorrelated additive noise. The enhancement is performed by eliminating the noise components in the noise subspace and estimating the clean image from the remaining components in the signal subspace. Ki-Seung Lee, Eun Suk Kim, Won Doh, Dae Hee Youn |
IEEE Trans. Image Process. | 1 |
| 1996 | Image enhancement based on signal subspace approachabstractA newly developed image enhancement algorithm is described in this contribution. The proposed algorithm makes use of the signal subspace method to enhance images corrupted by uncorrelated additive noise. This enhancement is performed by eliminating the noise subspace and estimating clean image from the remaining signal subspace. We propose the block-adaptive Wiener filtering which engages properties of the human visual system to estimate clean image. This criterion enables one to not only preserve the detailed structure of the given image, but to reduce the level of background noise as well. Subjective evaluation tests show the superiority of the method proposed here. In particular, edge blurring effects are noticeably reduced compared to the conventional methods. Ki-Seung Lee, Won Doh, Kun Jong Park, Dae Hee Youn |
ICIP (1) | 1 |
| 1996 | A new voice transformation method based on both linear and nonlinear prediction analysisabstractIn this paper, we describe a voice transformation meth-od which c hanges source speaker's acoustic features to those of a target speaker.The method developed here, acoustic features are divided into two parts, linear and nonlinear parts.Linear parts are characterized by LPC cepstrum coecients which are obtained from LP analysis.As for nonlinear part, which represent the excitation signal, is modelled by the long-delay nonlinear predictor using a neural net.Conversion rules for excitation signal are generated by the average pitch ratio and the mapping codebook, and those for LPC cepstrum coecients are based on the orthogonal vetctor space conversion.In addition, the spectral envelope compensation is proposed to correct spectral distortion in the transformed speech.A listening test shows that the proposed method makes it possible to convert speaker's individuality while maintaining high quality. Ki-Seung Lee, Dae Hee Youn, Il-Whan Cha |
ICSLP | 1 |
| 1995 | Voice personality transformation using an orthogonal vector space conversion
Ki-Seung Lee, Dae Hee Youn, Il-Whan Cha |
EUROSPEECH | 1 |