VLDB 2026 Research / reviewers in the wild / expert
Tsun-An Hsieh
dblp:262/7175
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On the Importance of Neural Wiener Filter for Resource Efficient Multichannel Speech EnhancementabstractWe introduce a time-domain framework for efficient multichannel speech enhancement, emphasizing low latency and computational efficiency. This framework incorporates two compact deep neural networks (DNNs) surrounding a multichannel neural Wiener filter (NWF). The first DNN enhances the speech signal to estimate NWF coefficients, while the second DNN refines the output from the NWF. The NWF, while conceptually similar to the traditional frequency-domain Wiener filter, undergoes a training process optimized for low-latency speech enhancement, involving fine-tuning of both analysis and synthesis transforms. Our research results illustrate that the NWF output, having minimal nonlinear distortions, attains performance levels akin to those of the first DNN, deviating from conventional Wiener filter paradigms. Training all components jointly outperforms sequential training, despite its simplicity. Consequently, this framework achieves superior performance with fewer parameters and reduced computational demands, making it a compelling solution for resource-efficient multichannel speech enhancement. Tsun-An Hsieh, Jacob Donley, Buye Xu, Ashutosh Pandey 0004 |
ICASSP | 1 |
| 2024 | Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
Tsun-An Hsieh, Heeyoul Choi |
INTERSPEECH | 1 |
| 2022 | Speech Recovery For Real-World Self-Powered Intermittent DevicesabstractThe incompleteness of speech inputs severely degrades the performance of all the related speech signal processing applications. Although many researches have been proposed to address this issue, they controlled the data missing conditions by simulation with self-defined masking lengths or sizes. Besides, the masking definitions are different among all these experimental settings. This paper presents a novel intermittent speech recovery (ISR) system for real-world self-powered intermittent devices. Three contributive stages: interpolation, enhancement, and combination are applied to the ISR system for speech reconstruction. The experimental results show that our recovery system increases speech quality by up to 591.7%, while increasing speech intelligibility by up to 80.5%. Most importantly, the proposed ISR system improves the WER scores by up to 52.6%. The promising results not only confirm the effectiveness of the reconstruction but also encourage the utilization of these battery-free wearable/IoT devices. Yuchen Lin 0003, Tsun-An Hsieh, Kuo-Hsuan Hung, Harinath Garudadri, Yu Tsao 0001, Tei-Wei Kuo |
ICASSP | 2 |
| 2022 | OSSEM: one-shot speaker adaptive speech enhancement using meta learningabstractAlthough deep learning (DL) has achieved notable progress in speech enhancement (SE), further research is still required for a DL-based SE system to adapt effectively and efficiently to particular speakers.In this study, we propose a novel meta-learning-based speaker-adaptive SE approach (called OSSEM) that aims to achieve SE model adaptation in a one-shot manner.OSSEM consists of a modified transformer SE network and a speaker-specific masking (SSM) network.In practice, the SSM network takes an enrolled speaker embedding extracted using ECAPA-TDNN to adjust the input noisy feature through masking.To evaluate OSSEM, we designed a modified Voice Bank-DEMAND dataset, in which one utterance from the testing set was used for model adaptation, and the remaining utterances were used for testing the performance.Moreover, we set restrictions allowing the enhancement process to be conducted in real time, and thus designed OSSEM to be a causal SE system.Experimental results first show that OSSEM can effectively adapt a pretrained SE model to a particular speaker with only one utterance, thus yielding improved SE results.Meanwhile, OSSEM exhibits a competitive performance compared to state-of-the-art causal SE systems. Szu-Wei Fu, Tsun-An Hsieh, Yu Tsao 0001, Mirco Ravanelli |
INTERSPEECH | 3 |
| 2021 | MetricGAN+: An Improved Version of MetricGAN for Speech EnhancementabstractThe discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory.Objective evaluation metrics which consider human perception can hence serve as a bridge to reduce the gap.Our previously proposed MetricGAN was designed to optimize objective metrics by connecting the metric with a discriminator.Because only the scores of the target evaluation functions are needed during training, the metrics can even be non-differentiable.In this study, we propose a MetricGAN+ in which three training techniques incorporating domainknowledge of speech processing are proposed.With these techniques, experimental results on the VoiceBank-DEMAND dataset show that MetricGAN+ can increase PESQ score by 0.3 compared to the previous MetricGAN and achieve stateof-the-art results (PESQ score = 3.15). Szu-Wei Fu, Tsun-An Hsieh, Peter Plantinga, Mirco Ravanelli, Xugang Lu, Yu Tsao 0001 |
Interspeech | 3 |
| 2021 | Improving Perceptual Quality by Phone-Fortified Perceptual Loss Using Wasserstein Distance for Speech EnhancementabstractSpeech enhancement (SE) aims to improve speech quality and intelligibility, which are both related to a smooth transition in speech segments that may carry linguistic information, e.g.phones and syllables.In this study, we propose a novel phonefortified perceptual loss (PFPL) that takes phonetic information into account for training SE models.To effectively incorporate the phonetic information, the PFPL is computed based on latent representations of the wav2vec model, a powerful selfsupervised encoder that renders rich phonetic information.To more accurately measure the distribution distances of the latent representations, the PFPL adopts the Wasserstein distance as the distance measure.Our experimental results first reveal that the PFPL is more correlated with the perceptual evaluation metrics, as compared to signal-level losses.Moreover, the results showed that the PFPL can enable a deep complex U-Net SE model to achieve highly competitive performance in terms of standardized quality and intelligibility evaluations on the Voice Bank-DEMAND dataset. Tsun-An Hsieh, Szu-Wei Fu, Xugang Lu, Yu Tsao 0001 |
Interspeech | 1 |
| 2020 | WaveCRN: An Efficient Convolutional Recurrent Neural Network for End-to-End Speech EnhancementabstractDue to the simple design pipeline, end-to-end (E2E) neural models for speech enhancement (SE) have attracted great interest. In order to improve the performance of the E2E model, the local and sequential properties of speech should be efficiently taken into account when modelling. However, in most current E2E models for SE, these properties are either not fully considered or are too complex to be realized. In this letter, we propose an efficient E2E SE model, termed WaveCRN. Compared with models based on convolutional neural networks (CNN) or long short-term memory (LSTM), WaveCRN uses a CNN module to capture the speech locality features and a stacked simple recurrent units (SRU) module to model the sequential property of the locality features. Different from conventional recurrent neural networks and LSTM, SRU can be efficiently parallelized in calculation, with even fewer model parameters. In order to more effectively suppress noise components in the noisy speech, we derive a novel restricted feature masking approach, which performs enhancement on the feature maps in the hidden layers; this is different from the approaches that apply the estimated ratio mask to the noisy spectral features, which is commonly used in speech separation methods. Experimental results on speech denoising and compressed speech restoration tasks confirm that with the SRU and the restricted feature map, WaveCRN performs comparably to other state-of-the-art approaches with notably reduced model complexity and inference time. Tsun-An Hsieh, Hsin-Min Wang, Xugang Lu, Yu Tsao 0001 |
IEEE Signal Process. Lett. | 1 |