EDBT 2026 Demo / reviewers in the wild / expert
Taesu Kim
dblp:44/6997
· DBLP profile ↗
33ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 5 since 2021Systems, architecture and hardware · 9 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MEGA.mini-S: A 1.9-to-15.0 TOPS/W Generative AI Processor with Bit-serial Weight Scalability and Bank-Conflict-free Sparsity Compression
Taesu Kim, Donghyeon Han |
ISCAS | 1 |
| 2025 | Debunking the CUDA Myth Towards GPU-based AI Systems: Evaluation of the Performance and Programmability of Intel's Gaudi NPU for AI Model ServingabstractThis paper presents a comprehensive evaluation of Intel Gaudi NPUs as an alternative to NVIDIA GPUs, which is currently the de facto standard in AI system design.First, we create microbenchmarks to compare Intel Gaudi-2 with NVIDIA A100, showing that Gaudi-2 achieves competitive performance not only in primitive AI compute, memory, and communication operations but also in executing several important AI workloads end-to-end.We then assess Gaudi NPU's programmability by discussing several softwarelevel optimization strategies to employ for implementing critical FBGEMM operators and vLLM, evaluating their efficiency against GPU-optimized counterparts.Results indicate that Gaudi-2 achieves energy efficiency comparable to A100, though there are notable areas for improvement in terms of software maturity.Overall, we conclude that, with effective integration into high-level AI frameworks, Gaudi NPUs could challenge NVIDIA GPU's dominance in * Both authors contributed equally to this research. Yunjae Lee, Juntaek Lim, Jehyeon Bang, Eunyeong Cho, Huijong Jeong, Taesu Kim, Joonhyung Lee, Jinseop Im, Ranggi Hwang, Se Jung Kwon, Dongsoo Lee, Minsoo Rhu |
ISCA | 6 |
| 2025 | GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-TuningabstractLow-Rank Adaptation (LoRA) is a popular method for parameter-efficient fine-tuning (PEFT) of generative models, valued for its simplicity and effectiveness. Despite recent enhancements, LoRA still suffers from a fundamental limitation: overfitting when the bottleneck is widened. It performs best at ranks 32–64, yet its accuracy stagnates or declines at higher ranks, still falling short of full fine-tuning (FFT) performance. We identify the root cause as LoRA’s structural bottleneck, which introduces gradient entanglement to the unrelated input channels and distorts gradient propagation. To address this, we introduce a novel structure, Granular Low-Rank Adaptation (GraLoRA) that partitions weight matrices into sub-blocks, each with its own low-rank adapter. With negligible computational or storage cost, GraLoRA overcomes LoRA’s limitations, effectively increases the representational capacity, and more closely approximates FFT behavior. Experiments on code generation, commonsense reasoning, mathematical reasoning, general language understanding, and image generation benchmarks show that GraLoRA consistently outperforms LoRA and other baselines, achieving up to +8.5\% absolute gain in Pass@1 on HumanEval+. These improvements hold across model sizes and rank settings, making GraLoRA a scalable and robust solution for PEFT. Yeonjoon Jung, Daehyun Ahn, Taesu Kim, Eunhyeok Park |
NeurIPS | 4 |
| 2024 | OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language ModelsabstractLarge language models (LLMs) with hundreds of billions of parameters require powerful server-grade GPUs for inference, limiting their practical deployment. To address this challenge, we introduce the outlier-aware weight quantization (OWQ) method, which aims to minimize LLM's footprint through low-precision representation. OWQ prioritizes a small subset of structured weights sensitive to quantization, storing them in high-precision, while applying highly tuned quantization to the remaining dense weights. This sensitivity-aware mixed-precision scheme reduces the quantization error notably, and extensive experiments demonstrate that 3.1-bit models using OWQ perform comparably to 4-bit models optimized by OPTQ. Furthermore, OWQ incorporates a parameter-efficient fine-tuning for task-specific adaptation, called weak column tuning (WCT), enabling accurate task-specific LLM adaptation with minimal memory overhead in the optimized format. OWQ represents a notable advancement in the flexibility, efficiency, and practicality of LLM optimization literature. The source code is available at https://github.com/xvyaward/owq. Changhun Lee, Jungyu Jin, Taesu Kim, Eunhyeok Park |
AAAI | 3 |
| 2024 | SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer BlocksabstractLarge language models (LLMs) have proven to be highly effective across various natural language processing tasks. However, their large number of parameters poses significant challenges for practical deployment. Pruning, a technique aimed at reducing the size and complexity of LLMs, offers a potential solution by removing redundant components from the network. Despite the promise of pruning, existing methods often struggle to achieve substantial end-to-end LLM inference speedup. In this paper, we introduce SLEB, a novel approach designed to stream- line LLMs by eliminating redundant transformer blocks. We choose the transformer block as the fundamental unit for pruning, because LLMs exhibit block-level redundancy with high similarity between the outputs of neighboring blocks. This choice allows us to effectively enhance the processing speed of LLMs. Our experimental results demonstrate that SLEB outperforms previous LLM pruning methods in accelerating LLM inference while also maintaining superior perplexity and accuracy, making SLEB as a promising technique for enhancing the efficiency of LLMs. The code is available at: https://github.com/jiwonsong-dev/SLEB. Jiwon Song, Kyungseok Oh, Taesu Kim, Yulhwa Kim, Jae-Joon Kim |
ICML | 3 |
| 2024 | Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language ModelsabstractBinarization, which converts weight parameters to binary values, has emerged as an effective strategy to reduce the size of large language models (LLMs). However, typical binarization techniques significantly diminish linguistic effectiveness of LLMs.
To address this issue, we introduce a novel binarization technique called Mixture of Scales (BinaryMoS). Unlike conventional methods, BinaryMoS employs multiple scaling experts for binary weights, dynamically merging these experts for each token to adaptively generate scaling factors. This token-adaptive approach boosts the representational power of binarized LLMs by enabling contextual adjustments to the values of binary weights. Moreover, because this adaptive process only involves the scaling factors rather than the entire weight matrix, BinaryMoS maintains compression efficiency similar to traditional static binarization methods. Our experimental results reveal that BinaryMoS surpasses conventional binarization techniques in various natural language processing tasks and even outperforms 2-bit quantization methods, all while maintaining similar model size to static binarization techniques. Dongwon Jo, Taesu Kim, Yulhwa Kim, Jae-Joon Kim |
NeurIPS | 2 |
| 2023 | Cross-Speaker Emotion Transfer by Manipulating Speech Style LatentsabstractIn recent years, emotional text-to-speech has shown considerable progress. However, it requires a large amount of labeled data, which is not easily accessible. Even if it is possible to acquire an emotional speech dataset, there is still a limitation in controlling emotion intensity. In this work, we propose a novel method for cross-speaker emotion transfer and manipulation using vector arithmetic in latent style space. By leveraging only a few labeled samples, we generate emotional speech from reading-style speech without losing the speaker identity. Furthermore, emotion strength is readily controllable using a scalar value, providing an intuitive way for users to manipulate speech. Experimental results show the proposed method affords superior performance in terms of expressiveness, naturalness, and controllability, preserving speaker identity. Suhee Jo, Younggun Lee, Yookyung Shin, Yeongtae Hwang, Taesu Kim |
ICASSP | 5 |
| 2023 | Leveraging Early-Stage Robustness in Diffusion Models for Efficient and High-Quality Image SynthesisabstractWhile diffusion models have demonstrated exceptional image generation capabilities, the iterative noise estimation process required for these models is compute-intensive and their practical implementation is limited by slow sampling speeds. In this paper, we propose a novel approach to speed up the noise estimation network by leveraging the robustness of early-stage diffusion models. Our findings indicate that inaccurate computation during the early-stage of the reverse diffusion process has minimal impact on the quality of generated images, as this stage primarily outlines the image while later stages handle the finer details that require more sensitive information. To improve computational efficiency, we combine our findings with post-training quantization (PTQ) to introduce a method that utilizes low-bit activation for the early reverse diffusion process while maintaining high-bit activation for the later stages. Experimental results show that the proposed method can accelerate the early-stage computation without sacrificing the quality of the generated images. Yulhwa Kim, Dongwon Jo, Hyesung Jeon, Taesu Kim, Daehyun Ahn, Jae-Joon Kim |
NeurIPS | 4 |
| 2023 | Searching for Robust Binary Neural Networks via Bimodal Parameter PerturbationabstractBinary neural networks (BNNs) are advantageous in performance and memory footprint but suffer from low accuracy due to their limited expression capability. Recent works have tried to enhance the accuracy of BNNs via a gradient-based search algorithm and showed promising results. However, the mixture of architecture search and binarization induce the instability of the search process, resulting in convergence to the suboptimal point. To address this issue, we propose a BNN architecture search framework with bimodal parameter perturbation. The bimodal parameter perturbation can improve the stability of gradient-based architecture search by reducing the sharpness of the loss surface along both weight and architecture parameter axes. In addition, we refine the inverted bottleneck convolution block for having robustness with BNNs. The synergy of the refined space and the stabilized search process allows us to find out the accurate BNNs with high computation efficiency. Experimental results show that our framework finds the best architecture on CIFAR-100 and ImageNet datasets in the existing search space for BNNs. We also tested our framework on another search space based on the inverted bottleneck convolution block, and the selected BNN models using our approach achieved the highest accuracy on both datasets with a much smaller number of equivalent operations than previous works. Daehyun Ahn, Taesu Kim, Eunhyeok Park, Jae-Joon Kim |
WACV | 3 |
| 2023 | V-LSTM: An Efficient LSTM Accelerator Using Fixed Nonzero-Ratio Viterbi-Based PruningabstractLong short-term memory (LSTM) has been widely adopted in tasks with sequence data, such as speech recognition and language modeling. LSTM brought significant accuracy improvement by introducing additional parameters to recurrent neural network (RNN). However, increasing number of parameters and computations also led to inefficiency in computing LSTM on edge devices with limited on-chip memory size and DRAM bandwidth. In order to reduce the latency and energy of LSTM computations, there has been a pressing need for model compression schemes and suitable hardware accelerators. In this article, we first propose the Fixed Nonzero-ratio Viterbi-based Pruning, which can reduce the memory footprint of LSTM models by 96% with negligible accuracy loss. By applying additional constraints on the distribution of surviving weights in Viterbi-based Pruning, the proposed pruning scheme mitigates the load-imbalance problem and thereby increases the processing engine utilization rate. Then, we propose the V-LSTM, an efficient sparse LSTM accelerator based on the proposed pruning scheme. High compression ratio of the proposed pruning scheme allows the proposed accelerator to achieve 24.9% lower per-sample latency than that of state-of-the-art accelerators. The proposed accelerator is implemented on Xilinx VC-709 FPGA evaluation board running at 200 MHz for evaluation. Taesu Kim, Daehyun Ahn, Dongsoo Lee, Jae-Joon Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS
Yookyung Shin, Younggun Lee, Suhee Jo, Yeongtae Hwang, Taesu Kim |
INTERSPEECH | 5 |
| 2022 | EdiTTS: Score-based Editing for Controllable Text-to-Speech
Jaesung Tae, Hyeongju Kim, Taesu Kim |
INTERSPEECH | 3 |
| 2021 | SPRITE: Sparsity-Aware Neural Processing Unit with Constant Probability of Index-MatchingabstractSparse neural networks are widely used for memory savings. However, irregular indices of non-zero input activations and weights tend to degrade the overall system performance. This paper presents a scheme to maintain constant probability of index-matching for weight and input over a wide range of sparsity overcoming a critical limitation in previous works. A sparsity-aware neural processing unit based on the proposed scheme improves the system performance up to 6.1× compared to previous sparse convolutional neural network hardware accelerators. Sungju Ryu, Youngtaek Oh, Taesu Kim, Daehyun Ahn, Jae-Joon Kim |
DATE | 3 |
| 2020 | V-LSTM: An Efficient LSTM Accelerator Using Fixed Nonzero-Ratio Viterbi-Based PruningabstractLong Short-Term Memory (LSTM) has been widely adopted in tasks with sequence data, such as speech recognition and language modeling. LSTM brought significant accuracy improvement by introducing additional parameters to Recurrent Neural Network (RNN). However, increasing number of parameters and computations also led to inefficiency in computing LSTM on edge devices with limited on-chip memory size and DRAM bandwidth. In order to reduce the latency and energy of LSTM computations, there has been a pressing need for model compression schemes and suitable hardware accelerators. In this paper, we first propose the Fixed Nonzero-ratio Viterbi-based Pruning, which can reduce the memory footprint of LSTM models by 96% with negligible accuracy loss. By applying additional constraints on the distribution of surviving weights in Viterbi-based Pruning, the proposed pruning scheme mitigates the load-imbalance problem and thereby increases the processing engine utilization rate. Then, we propose the V-LSTM, an efficient sparse LSTM accelerator based on the proposed pruning scheme. High compression ratio of the proposed pruning scheme allows the proposed accelerator to achieve 24.9% lower per-sample latency than that of state-of-the-art accelerators. The proposed accelerator is implemented on Xilinx VC-709 FPGA evaluation board running at 200MHz for evaluation. Taesu Kim, Daehyun Ahn, Jae-Joon Kim |
FPGA | 1 |
| 2020 | Time-step interleaved weight reuse for LSTM neural network computingabstractIn Long Short-Term Memory (LSTM) neural network models, a weight matrix tends to be repeatedly loaded from DRAM if the size of on-chip storage of the processor is not large enough to store the entire matrix. To alleviate heavy overhead of DRAM access for weight loading in LSTM computations, we propose a weight reuse scheme which utilizes the weight sharing characteristics in two adjacent time-step computations. Experimental results show that the proposed weight reuse scheme reduces the energy consumption by 28.4-57.3% and increases the overall throughput by 110.8% compared to the conventional schemes. Naebeom Park, Yulhwa Kim, Daehyun Ahn, Taesu Kim, Jae-Joon Kim |
ISLPED | 4 |
| 2019 | Robust and Fine-grained Prosody Control of End-to-end Speech SynthesisabstractWe propose prosody embeddings for emotional and expressive speech synthesis networks. The proposed methods introduce temporal structures in the embedding networks, thus enabling fine-grained control of the speaking style of the synthesized speech. The temporal structures can be designed either on the speech side or the text side, leading to different control resolutions in time. The prosody embedding networks are plugged into end-to-end speech synthesis networks and trained without any other supervision except for the target speech for synthesizing. It is demonstrated that the prosody embedding networks learned to extract prosodic features. By adjusting the learned prosody features, we could change the pitch and amplitude of the synthesized speech both at the frame level and the phoneme level. We also introduce the temporal normalization of prosody embeddings, which shows better robustness against speaker perturbations during prosody transfer tasks. Younggun Lee, Taesu Kim |
ICASSP | 2 |
| 2019 | Double Viterbi: Weight Encoding for High Compression Ratio and Fast On-Chip Reconstruction for Deep Neural Network
Daehyun Ahn, Dongsoo Lee, Taesu Kim, Jae-Joon Kim |
ICLR (Poster) | 3 |
| 2019 | Large-Scale Speaker Retrieval on Random Speaker Variability SubspaceabstractThis paper describes a fast speaker search system to retrieve segments of the same voice identity in the large-scale data.A recent study shows that Locality Sensitive Hashing (LSH) enables quick retrieval of a relevant voice in the large-scale data in conjunction with i-vector while maintaining accuracy.In this paper, we proposed Random Speaker-variability Subspace (RSS) projection to map a data into LSH based hash tables.We hypothesized that rather than projecting on completely random subspace without considering data, projecting on randomly generated speaker variability space would give more chance to put the same speaker representation into the same hash bins, so we can use less number of hash tables.Multiple RSS can be generated by randomly selecting a subset of speakers from a large speaker cohort.From the experimental result, the proposed approach shows 100 times and 7 times faster than the linear search and LSH, respectively. Suwon Shon, Younggun Lee, Taesu Kim |
INTERSPEECH | 3 |
| 2018 | c.light: A Tool for Exploring Light Properties in Early Design StageabstractAlthough a light becomes an important design element, there are little techniques available to explore shapes and light effects in early design stages. We present c.light, a design tool that consists of a set of modules and a mobile application for visualizing the light in a physical world. It allows designers to easily fabricate both tangible and intangible properties of a light without a technical barrier. We analyzed how c.light contributes to the ideation process of light design through a workshop. The results showed that c.light largely expands designers' capability to manipulate intangible properties of light and, by doing so, it facilitates collaborative and inverted ideation process in early design stages. It is expected that the results of this study could enhance our understanding of how designers manipulate light in a physical world in early design stages and could be a good stepping stone for future tool development. Kyeong-Ah Jeong, EunJin Kim, Taesu Kim, Hyeon-Jeong Suk |
CHI | 3 |
| 2018 | Viterbi-based Pruning for Sparse Matrix with Fixed and High Index Compression Ratio
Dongsoo Lee, Daehyun Ahn, Taesu Kim, Pierce Chuang, Jae-Joon Kim |
ICLR (Poster) | 3 |
| 2018 | Deep Neural Network Optimized to Resistive Memory with Nonlinear Current-Voltage CharacteristicsabstractArtificial Neural Network computation relies on intensive vector-matrix multiplications. Recently, the emerging nonvolatile memory (NVM) crossbar array showed a feasibility of implementing such operations with high energy efficiency. Thus, there have been many works on efficiently utilizing emerging NVM crossbar arrays as analog vector-matrix multipliers. However, nonlinear I-V characteristics of NVM restrain critical design parameters, such as the read voltage and weight range, resulting in substantial accuracy loss. In this article, instead of optimizing hardware parameters to a given neural network, we propose a methodology of reconstructing the neural network itself to be optimized to resistive memory crossbar arrays. To verify the validity of the proposed method, we simulated various neural networks with MNIST and CIFAR-10 dataset using two different Resistive Random Access Memory models. Simulation results show that our proposed neural network produces inference accuracies significantly higher than conventional neural network when the network is mapped to synapse devices with nonlinear I-V characteristics. Taesu Kim, Jinseok Kim 0004, Jae-Joon Kim |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2017 | Speaker Clustering by Iteratively Finding Discriminative Feature Space and Cluster Labels
Sungrack Yun, Hye Jin Jang, Taesu Kim |
INTERSPEECH | 3 |
| 2012 | Revisiting level-0 caches in embedded processorsabstractLevel-0 (L0) caches have been proposed in the past as an inexpensive way to improve performance and reduce energy consumption in resource-constrained embedded processors. This paper proposes new L0 data cache organizations using the assumption that an L0 hit/miss determination can be completed prior to the L1 access. This is a realistic assumption for very small L0 caches that can nevertheless deliver significant miss rate and/or energy reduction. The key issue for such caches is how and when to move data between the L0 and L1 caches. The first new cache, a flow cache, targets a conflict miss reduction in a direct-mapped L1 cache. It offers a simpler hardware design and uses on average 10% less dynamic energy than the victim cache with nearly identical performance. The second new cache, a hit cache, reduces the dynamic energy consumption in a set-associative L1 cache by 30% without impacting performance. A variant of this policy reduces the dynamic energy consumption by up to 50%, with 5% performance degradation. Nam Duong, Taesu Kim, Dali Zhao, Alexander V. Veidenbaum |
CASES | 2 |
| 2012 | Improving Cache Management Policies Using Dynamic Reuse DistancesabstractCache management policies such as replacement, bypass, or shared cache partitioning have been relying on data reuse behavior to predict the future. This paper proposes a new way to use dynamic reuse distances to further improve such policies. A new replacement policy is proposed which prevents replacing a cache line until a certain number of accesses to its cache set, called a Protecting Distance (PD). The policy protects a cache line long enough for it to be reused, but not beyond that to avoid cache pollution. This can be combined with a bypass mechanism that also relies on dynamic reuse analysis to bypass lines with less expected reuse. A miss fetch is bypassed if there are no unprotected lines. A hit rate model based on dynamic reuse history is proposed and the PD that maximizes the hit rate is dynamically computed. The PD is recomputed periodically to track a program's memory access behavior and phases. Next, a new multi-core cache partitioning policy is proposed using the concept of protection. It manages lifetimes of lines from different cores (threads) in such a way that the overall hit rate is maximized. The average per-thread lifetime is reduced by decreasing the thread's PD. The single-core PD-based replacement policy with bypass achieves an average speedup of 4.2% over the DIP policy, while the average speedups over DIP are 1.5% for dynamic RRIP (DRRIP) and 1.6% for sampling dead-block prediction (SDP). The 16-core PD-based partitioning policy improves the average weighted IPC by 5.2%, throughput by 6.4% and fairness by 9.9% over thread-aware DRRIP (TA-DRRIP). The required hardware is evaluated and the overhead is shown to be manageable. Nam Duong, Dali Zhao, Taesu Kim, Rosario Cammarota, Mateo Valero, Alexander V. Veidenbaum |
MICRO | 3 |
| 2011 | Self-similarity Based Lightweight Intrusion Detection Method for Cloud Computing
Hyukmin Kwon, Taesu Kim, Song Jin Yu, Huy Kang Kim |
ACIIDS (2) | 2 |
| 2010 | Binaural loudness based speech reinforcement with a closed-form solutionabstractThis paper addresses a perceptual signal processing to far-end speech signal in communication systems under near-end environmental noise conditions. Based on the binaural perceptual loudness model, the proposed speech reinforcement system achieves better speech quality and clearness. To effectively reflect the noise influence to both ears, the proposed method utilizes a noise level difference between open and receiver side ear. Its computational complexity is also reduced by deriving an approximated closed-form solution while computing frequency-dependent gain factors. Test results confirm that the proposed system significantly enhances the clearness of target speech while maintaining speech quality compared to conventional monaural-based one. Ho Seon Shin, Min-Seok Choi, Taesu Kim, Hong-Goo Kang |
ICASSP | 3 |
| 2007 | Fast fixed-point independent vector analysis algorithms for convolutive blind source separation
Intae Lee, Taesu Kim, Te-Won Lee |
Signal Process. | 2 |
| 2007 | Blind Source Separation Exploiting Higher-Order Frequency DependenciesabstractBlind source separation (BSS) is a challenging problem in real-world environments where sources are time delayed and convolved. The problem becomes more difficult in very reverberant conditions, with an increasing number of sources, and geometric configurations of the sources such that finding directionality is not sufficient for source separation. In this paper, we propose a new algorithm that exploits higher order frequency dependencies of source signals in order to separate them when they are mixed. In the frequency domain, this formulation assumes that dependencies exist between frequency bins instead of defining independence for each frequency bin. In this manner, we can avoid the well-known frequency permutation problem. To derive the learning algorithm, we define a cost function, which is an extension of mutual information between multivariate random variables. By introducing a source prior that models the inherent frequency dependencies, we obtain a simple form of a multivariate score function. In experiments, we generate simulated data with various kinds of sources in various environments. We evaluate the performances and compare it with other well-known algorithms. The results show the proposed algorithm outperforms the others in most cases. The algorithm is also able to accurately recover six sources with six microphones. In this case, we can obtain about 16-dB signal-to-interference ratio (SIR) improvement. Similar performance is observed in real conference room recordings with three human speakers reading sentences and one loudspeaker playing music Taesu Kim, Hagai Attias, Soo-Young Lee, Te-Won Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Frequency Domain Blind Source Separation Exploiting Higher-Order DependenciesabstractWe propose a novel approach to the blind source separation (BSS) that exploits frequency dependencies within a source. In contrast to conventional algorithms that separate the sources independently in each frequency bin, we assume that dependencies exist between frequency bins in a source signal. In this manner, we can reduce or eliminate the well-known frequency permutation problem. We derive the learning algorithm by defining a cost function as an extension of mutual information between multivariate random variables and by introducing a source prior that models the inherent frequency dependencies. This results in a simple form of a multivariate score function. In simulations and real recording experiments, we evaluate the performance of the proposed method and compare it against other well-known algorithms under various conditions. Our results indicate that modeling dependencies yields improved performance and robust scaling to higher number of sources and mixtures. Taesu Kim, Hagai Attias, Soo-Young Lee, Te-Won Lee |
ICASSP (5) | 1 |
| 2006 | On the multivariate Laplace distributionabstractIn this letter, we discuss the multivariate Laplace probability model in the context of a normal variance mixture model. We briefly review the derivation of the probability density function (pdf) and discuss a few important properties. We then present two methods for estimating its parameters from data and include an example of usage, where we apply the model to represent the statistics of the discrete Fourier transform coefficients of a speech signal. Since the pdf is given in closed form, and the model parameters can be easily obtained, this distribution may be useful for representing multivariate, sparsely distributed data, with mutually dependent components. Torbjørn Eltoft, Taesu Kim, Te-Won Lee |
IEEE Signal Process. Lett. | 2 |
| 2005 | Robust time delay estimation in noisy reverberant environments with a probabilistic graphical modelabstractThe paper presents a fast new algorithm for estimating the relative time delay between a microphone pair in noisy reverberant conditions. The algorithm is derived from a novel approach to this problem, which we formulate in the framework of probabilistic graphical models. We construct a probabilistic model of the microphone signals, using models of speech and noise as building blocks. The reverberation filter coefficients and the relative time delay appear as model parameters, and are estimated by our algorithm from data. The resulting delay estimate is Bayes optimal and takes into account noise and reverberation in a principled manner. We demonstrate very good performance on data from real and simulated room environments. Taesu Kim, Hagai Attias, Soo-Young Lee, Te-Won Lee |
ICASSP (5) | 1 |
| 2005 | Learning self-organized topology-preserving complex speech features at primary auditory cortex
Taesu Kim, Soo-Young Lee |
Neurocomputing | 1 |
| 2003 | FPGA implementation of ICA algorithm for blind signal separation and adaptive noise cancelingabstractAn field programmable gate array (FPGA) implementation of independent component analysis (ICA) algorithm is reported for blind signal separation (BSS) and adaptive noise canceling (ANC) in real time. In order to provide enormous computing power for ICA-based algorithms with multipath reverberation, a special digital processor is designed and implemented in FPGA. The chip design fully utilizes modular concept and several chips may be put together for complex applications with a large number of noise sources. Experimental results with a fabricated test board are reported for ANC only, BSS only, and simultaneous ANC/BSS, which demonstrates successful speech enhancement in real environments in real time. Chang-Min Kim, Hyung-Min Park, Taesu Kim, Yoon-Kyung Choi, Soo-Young Lee |
IEEE Trans. Neural Networks | 3 |