Tai-Shih Chi

dblp:42/8052 · DBLP profile ↗
← Back
34ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-0584-8399ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 6 since 2021Artificial intelligence and machine learning · 15 · 6 since 2021Systems, architecture and hardware · 3 · 1 since 2021
YearPublicationVenuePosition
2025 Tonality-Based Accompaniment-Guided Automatic Singing Evaluation
Pei-Chin Hsieh, Yih-Liang Shen, Ngoc Son Tran, Tai-Shih Chi
INTERSPEECH4
2025 Spectro-Temporal Modulations Incorporated Two-Stream Robust Speech Emotion Recognition
abstract
Deep learning based speech emotion recognition (SER) models have shown impressive results in controlled environments, but their performance significantly degrades in noisy conditions. This paper proposes a robust two-stream SER model by combining spectro-temporal modulation features with conventional acoustic features. Experiments were conducted on German (EMODB) and English (RAVDESS) datasets using the clean-train-noisy-test paradigm. The results demonstrate that spectro-temporal modulation features offer superior robustness in noisy conditions compared with conventional acoustic features such as MFCCs and time-frequency features from Mel-spectrograms. Additionally, we analyze weights of modulation features and demonstrate the model emphasizes contours of formants and harmonics, which are crucial features for speech perception in noise, for robust SER. Incorporating the stream of spectro-temporal modulations not only enhances the robustness of the model but also provides deeper insights into the task of SER in noise.
Yih-Liang Shen, Pei-Chin Hsieh, Tai-Shih Chi
IEEE Trans. Affect. Comput.3
2024 SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
Chun Yin, Tai-Shih Chi, Yu Tsao 0001, Hsin-Min Wang
INTERSPEECH2
2023 Mandarin Electrolaryngeal Speech Voice Conversion using Cross-domain Features
Hsin-Hao Chen 0006, Yung-Lun Chien, Ming-Chi Yen, Shu-Wei Tsai, Tai-Shih Chi, Hsin-Min Wang, Yu Tsao 0001
INTERSPEECH5
2023 Audio-Visual Mandarin Electrolaryngeal Speech Voice Conversion
Yung-Lun Chien, Hsin-Hao Chen 0006, Ming-Chi Yen, Shu-Wei Tsai, Hsin-Min Wang, Yu Tsao 0001, Tai-Shih Chi
INTERSPEECH7
2022 Perceptual Characteristics Based Multi-objective Model for Speech Enhancement
Chiang-Jen Peng, Yun-Ju Chan, Yih-Liang Shen, Yu Tsao 0001, Tai-Shih Chi
INTERSPEECH6
2021 Extending Music Based On Emotion And Tonality Via Generative Adversarial Network
abstract
We propose a generative model for music extension in this paper. The model is composed of two classifiers, one for music emotion and one for music tonality, and a generative adversarial network (GAN). Therefore, it can generate symbolic music not only based on low level spectral and temporal characteristics, but also on high level emotion and tonality attributes of previously observed music pieces. The generative model works in a universal latent space constructed by the variational autoencoder (VAE) for representing music pieces. We conduct subjective listening tests and derive objective measures for performance evaluation. Experimental results show that the proposed model produces much smoother and more authentic music pieces than the baseline model in terms of all subjective and objective measures.
Bo-Wei Tseng, Yih-Liang Shen, Tai-Shih Chi
ICASSP3
2021 Attention-Based Multi-Task Learning for Speech-Enhancement and Speaker-Identification in Multi-Speaker Dialogue Scenario
abstract
Multi-task learning (MTL) and attention mechanism have been proven to effectively extract robust acoustic features for various speech-related tasks in noisy environments. In this study, we propose an attention-based MTL (ATM) approach that integrates MTL and the attention-weighting mechanism to simultaneously realize a multi-model learning structure that performs speech enhancement (SE) and speaker identification (SI). The proposed ATM system consists of three parts: SE, SI, and attention-Net (AttNet). The SE part is composed of a long-short-term memory (LSTM) model, and a deep neural network (DNN) model is used to develop the SI and AttNet parts. The overall ATM system first extracts the representative features and then enhances the speech signals in LSTM-SE and specifies speaker identity in DNN-SI. The AttNet computes weights based on DNN-SI to prepare better representative features for LSTM-SE. We tested the proposed ATM system on Taiwan Mandarin hearing in noise test sentences. The evaluation results confirmed that the proposed system can effectively enhance speech quality and intelligibility of a given noisy input. Moreover, the accuracy of the SI can also be notably improved by using the proposed ATM system.
Chiang-Jen Peng, Yun-Ju Chan, Syu-Siang Wang, Yu Tsao 0001, Tai-Shih Chi
ISCAS6
2020 A Multi-Dilation and Multi-Resolution Fully Convolutional Network for Singing Melody Extraction
abstract
Each human cognitive function involves bottom-up and top-down processes. Several methods have been proposed for singing melody extraction by emphasizing either the bottom-up or top-down processes. For hearing, the bottom-up processes include spectral and spectro-temporal decomposition of the sound by the cochlea and the auditory cortex. In this paper, we propose a neural network, which includes spectro-temporal multi-resolution decomposition of the log-spectrogram of the sound and a semantic segmentation model to respectively address the bottom-up and top-down processing of hearing, for singing melody extraction. Simulation results show the proposed model outperforms all previously proposed methods, emphasizing either bottom-up or top-down processing, in almost all objective evaluation metrics.
Cheng-You You, Tai-Shih Chi
ICASSP3
2019 Autoencoding HRTFS for DNN Based HRTF Personalization Using Anthropometric Features
abstract
We proposed a deep neural network (DNN) based approach to synthesize the magnitude of personalized head-related transfer functions (HRTFs) using anthropometric features of the user. To mitigate the over-fitting problem when training dataset is not very large, we built an autoencoder for dimensional reduction and establishing a crucial feature set to represent the raw HRTFs. Then we combined the decoder part of the autoencoder with a smaller DNN to synthesize the magnitude HRTFs. In this way, the complexity of the neural networks was greatly reduced to prevent unstable results with large variance due to overfitting. The proposed approach was compared with a baseline DNN model with no autoencoder. The log-spectral distortion (LSD) metric was used to evaluate the performance. Experiment results show that the proposed approach can reduce LSD of estimated HRTFs with greater stability.
Tzu-Yu Chen, Tzu-Hsuan Kuo, Tai-Shih Chi
ICASSP3
2019 CNN Based Two-stage Multi-resolution End-to-end Model for Singing Melody Extraction
abstract
Inspired by human hearing perception, we propose a two-stage multi-resolution end-to-end model for singing melody extraction in this paper. The convolutional neural network (CNN) is the core of the proposed model to generate multi-resolution representations. The 1-D and 2-D multi-resolution analysis on waveform and spectrogram-like graph are successively carried out by using 1-D and 2-D CNN kernels of different lengths and sizes. The 1-D CNNs with kernels of different lengths produce multi-resolution spectrogram-like graphs without suffering from the trade-off between spectral and temporal resolutions. The 2-D CNNs with kernels of different sizes extract features from spectro-temporal envelopes of different scales. Experiment results show the proposed model outperforms three compared systems in three out of five public databases.
Ming-Tso Chen, Bo-Jun Li, Tai-Shih Chi
ICASSP3
2019 Reinforcement Learning Based Speech Enhancement for Robust Speech Recognition
abstract
Conventional deep neural network (DNN)-based speech enhancement (SE) approaches aim to minimize the mean square error (MSE) between enhanced speech and clean reference. The MSE-optimized model may not directly improve the performance of an automatic speech recognition (ASR) system. If the target is to minimize the recognition error, the recognition results should be used to design the objective function for optimizing the SE model. However, the structure of an ASR system, which consists of multiple units, such as acoustic and language models, is usually complex and not differentiable. In this study, we propose to adopt the reinforcement learning (RL) algorithm to optimize the SE model based on the recognition results. We evaluated the proposed RL-based SE system on the Mandarin Chinese broadcast news corpus (MATBN). Experimental results demonstrate that the proposed SE system can effectively improve the ASR results with a notable 12:40% and 19:23% error rate reductions for signal to noise ratio (SNR) at 0 dB and 5 dB conditions, respectively.
Yih-Liang Shen, Chao-Yuan Huang, Syu-Siang Wang, Yu Tsao 0001, Hsin-Min Wang, Tai-Shih Chi
ICASSP6
2018 A Hybrid Neural Network Based on the Duplex Model of Pitch Perception for Singing Melody Extraction
abstract
In this paper, we build up a hybrid neural network (NN) for singing melody extraction from polyphonic music by imitating human pitch perception. For human hearing, there are two pitch perception models, the spectral model and the temporal model, in accordance with whether harmonics are resolved or not. Here, we first use NNs to implement individual models and evaluate their performance in the task of singing melody extraction. Then, we combine the NNs to constitute the composite NN to simulate the duplex model, which complements the pitch perception from unresolved harmonics of the spectral model using the temporal model. Simulation results show the proposed composite NN outperforms other conventional methods in singing melody extraction.
Hsin Chou, Ming-Tso Chen, Tai-Shih Chi
ICASSP3
2018 A Generative Auditory Model Embedded Neural Network for Speech Processing
abstract
Before the era of the neural network (NN), features extracted from auditory models have been applied to various speech applications and been demonstrated more robust against noise than conventional speech-processing features. What's the role of auditory models in the current NN era? Are they obsolete? To answer this question, we construct a NN with a generative auditory model embedded to process speech signals. The generative auditory model consists of two stages, the stage of spectrum estimation in the logarithmic-frequency axis by the cochlea and the stage of spectral-temporal analysis in the modulation domain by the auditory cortex. The NN is evaluated in a simple speaker identification task. Experiment results show that the auditory model embedded NN is still more robust against noise, especially in low SNR conditions, than the randomly-initialized NN in speaker identification.
Yu-Wen Lo, Yih-Liang Shen, Yuan-Fu Liao, Tai-Shih Chi
ICASSP4
2018 Singing Voice Correction Using Canonical Time Warping
abstract
Expressive singing voice correction is an appealing but challenging problem. A robust time-warping algorithm which synchronizes two singing recordings can provide a promising solution. We thereby propose to address the problem by canonical time warping (CTW) which aligns amateur singing recordings to professional ones. A new pitch contour is generated given the alignment information, and a pitch-corrected singing is synthesized back through the vocoder. The objective evaluation shows that CTW is robust against pitch-shifting and time-stretching effects, and the subjective test demonstrates that CTW prevails the other methods including DTW and the commercial auto-tuning software. Finally, we demonstrate the applicability of the proposed method in a practical, real-world scenario.
Yin-Jyun Luo, Ming-Tso Chen, Tai-Shih Chi, Li Su 0004
ICASSP3
2017 Dereverberation based on bin-wise temporal variations of complex spectrogram
abstract
Humans analyze sounds not only based on their frequency contents, but also on the temporal variations of the frequency contents. Inspired by auditory perception, we propose a deep neural network (DNN) based dereverberation algorithm in the rate domain, which presents the temporal variations of frequency contents, in this paper. We show convolutional noise in the time domain can be approximated to multiplicative noise in the rate domain. To remove the multiplicative noise, we adopt the rate-domain complex-valued ideal ratio mask (RDcIRM) as the training target of the DNN. Simulation results show that the proposed rate-domain DNN algorithm is more capable of recovering high-intelligible and high-quality speech from reverberant speech than the compared state-of-the-art dereverberation algorithm. Hence, it is highly suitable for speech applications involving human listeners.
Tai-Shih Chi
ICASSP3
2017 Simulations of High-Frequency Vocoder on Mandarin Speech Recognition for Acoustic Hearing Preserved Cochlear Implant
Tsung-Chen Wu, Tai-Shih Chi, Chia-Fone Lee
INTERSPEECH2
2016 A Spectral Modulation Sensitivity Weighted Pre-Emphasis Filter for Active Noise Control System
Kah-Meng Cheong, Yuh-Yuan Wang, Tai-Shih Chi
INTERSPEECH3
2016 Discriminative Layered Nonnegative Matrix Factorization for Speech Separation
Chung-Chien Hsu, Tai-Shih Chi, Jen-Tzung Chien
INTERSPEECH2
2015 Modulation Wiener filter for improving speech intelligibility
abstract
This paper presents a single-channel high-dimensional Wiener filter in the spectro-temporal modulation domain. Unlike other conventional noise reduction techniques, the proposed algorithm not only reduces noise but also enhances the “textures” of the speech signal. A non-iterative decision-directed noise estimation method is adopted to estimate the modulation SNR for the modulation-domain Wiener filter. The efficacy of the proposed algorithm in enhancing speech intelligibility is assessed using the short-time objective intelligibility (STOI) measure. Statistical analysis results demonstrate that our proposed algorithm can improve STOI scores in speech-shape noise (SSN) and white noise conditions, but not in babble noise condition, while the conventional Wiener filter fails to improve STOI scores in all three noise conditions.
Chung-Chien Hsu, Kah-Meng Cheong, Jen-Tzung Chien, Tai-Shih Chi
ICASSP4
2015 A hearing model to estimate mandarin speech intelligibility for the hearing impaired patients
abstract
A hearing model, which is parameterized by hearing thresholds, degrees of loudness recruitment and reductions of frequency resolution of a hearing-impaired (HI) patient, is proposed in this paper. The model is developed in the filter-bank framework and is flexible for fitting hearing-loss conditions of HI patients. Psychoacoustic experiments were conducted under clean and noisy conditions to validate the model's capability in predicting Mandarin speech intelligibility for HI patients. Statistical analysis on the hearing-test results suggests that the proposed model can predict Mandarin speech intelligibility for HI patients to a certain degree.
Pei-Chun Tsai, Shih-Ting Lin, Wen-Chung Lee, Chung-Chien Hsu, Tai-Shih Chi, Chia-Fone Lee
ICASSP5
2015 Layered nonnegative matrix factorization for speech separation
Chung-Chien Hsu, Jen-Tzung Chien, Tai-Shih Chi
INTERSPEECH3
2015 A two-stage singing voice separation algorithm using spectro-temporal modulation features
Frederick Z. Yen, Mao-Chang Huang, Tai-Shih Chi
INTERSPEECH3
2014 Binary mask estimation based on frequency modulations
Chung-Chien Hsu, Jen-Tzung Chien, Tai-Shih Chi
INTERSPEECH3
2013 Voice activity detection based on frequency modulation of harmonics
abstract
In this paper, we propose a voice activity detection (VAD) algorithm based on spectro-temporal modulation structures of input sounds. A multi-resolution spectro-temporal analysis framework is used to inspect prominent speech structures. By comparing with an adaptive threshold, the proposed VAD distinguishes speech from non-speech based on the energy of the frequency modulation of harmonics. Compared with three standard VADs, ITU-T G.729B, ETSI AMR1 and AMR2, our proposed VAD significantly outperforms them in non-stationary noises in terms of the receiver operating characteristic (ROC) curves and the recognition rates from a practical distributed speech recognition (DSR) system.
Chung-Chien Hsu, Tse-En Lin, Jian-Hueng Chen, Tai-Shih Chi
ICASSP4
2013 Spectral modulation sensitivity based perceptual acoustic echo cancellation
Wei-Lun Chuang, Kah-Meng Cheong, Chung-Chien Hsu, Tai-Shih Chi
INTERSPEECH4
2013 Spectro-temporal modulation based singing detection combined with pitch-based grouping for singing voice separation
Tse-En Lin, Chung-Chien Hsu, Jian-Hueng Chen, Tai-Shih Chi
INTERSPEECH5
2013 A precedence effect based far-field DoA estimation algorithm
abstract
A robust far-field DoA estimation algorithm is proposed in this paper. The algorithm is inspired by the precedence effect which results in humans robust capability in localizing sound sources. The algorithm implements the concept of the precedence effect by applying a proper threshold and an onset detection mechanism in the cross-correlation domain. Experiment results show that by cascading the proposed algorithm to a conventional cross-correlation based DoA algorithm, the accuracy of the DoA estimation is significantly improved in far-field test conditions.
Wen-Sheng Chou, Tai-Shih Chi
ISCAS2
2012 Spectro-temporal subband Wiener filter for speech enhancement
abstract
In this paper, we propose a signal-channel speech enhancement algorithm by applying the conventional Wiener filter in the spectro-temporal modulation domain. The multi-resolution spectro-temporal analysis and synthesis framework for Fourier spectrograms [12] is extended to the analysis-modification-synthesis (AMS) framework for speech enhancement. Compared with conventional speech enhancement algorithms, a Wiener filter and an extended minimum mean-square error (MMSE) algorithm, our proposed method outperforms them by a large/small margin in white/babble noise conditions from both objective and subjective evaluations.
Chung-Chien Hsu, Tse-En Lin, Jian-Hueng Chen, Tai-Shih Chi
ICASSP4
2011 A binaural algorithm for space and pitch detection
abstract
A binaural algorithm to simultaneously detect the azimuth angle and the pitch of the sound source is proposed in this paper. This algorithm is extended from the stereausis model with two-dimensional coincidence detectors in the joint Space-Pitch domain. In our simulations, sounds from different locations are produced by passing through the Head-Related-Transfer-Function (HRTF). Simulation results show that estimated azimuth angles from our proposed algorithm are more accurate than those from the stereausis model in the single sound source testing condition. Satisfactory results in streaming sound sources from the two-sound mixture by using estimated Space-Pitch information are also demonstrated in our pilot experiments.
Wen-Sheng Chou, Kah-Meng Cheong, Tai-Shih Chi
ICASSP3
2011 FFT-based spectro-temporal analysis and synthesis of sounds
abstract
The concept of the two-dimensional spectro-temporal modulation filtering of the auditory model is implemented for the FFT spectrogram. It analyzes the spectrogram in terms of the temporal dynamics and the spectral structures of the sound. The overlap and add (OTA) method, which is more convenient and reliable than the iterative-projection method proposed in, is used to invert the FFT spectrogram back to sounds. The Non-Negative Sparse Coding (NNSC) method is adopted to demonstrate the benefit of our analysis-synthesis procedures in a noise suppression application. Even without fine-tuning parameters, our proposed analysis-synthesis procedures offer benefits in de-noising especially under low SNR conditions.
Chung-Chien Hsu, Ting-Han Lin, Tai-Shih Chi
ICASSP3
2011 Low power InfomaxICA with compensation strategy for binaural hearing-aid
abstract
Binaural hearing-aids are under intensive study nowadays. The information exchange across both ears provides an opportunity to perform the blind source separation to enhance the signal SNR. This study combines the conventional InfomaxICA with our proposed novel algorithms: binaural delay compensation, minimum-interference initial de-mixing guess and low power design concept, for a real-time binaural hearing-aid application. Simulations demonstrate our algorithms provide 20.5 to 30.6 dB gain at 0 to 35 points delay at SNR= 0 dB in our assumed scenario. Therefore, our algorithms are robust under various practical wearing conditions.
Fan-Chiang Yi, Ching-Wen Huang, Tai-Shih Chi, Shyh-Jye Jou
ISCAS3
2010 Spectro-temporal modulations for robust speech emotion recognition
Lan-Ying Yeh, Tai-Shih Chi
INTERSPEECH2
2009 Perception-based objective speech quality assessment
abstract
A joint spectro-temporal auditory model is utilized to assess speech quality objectively. The model mimics early and central auditory functions and serves as a spectro-temporal modulation filterbank. Three perceptual relevant parameters, intelligibility, clarity and naturalness, are addressed by the model and are combined to estimate the subjective mean opinion score (MOS) for speech quality measure. Through a simple multiple linear regression analysis, we demonstrate the performance of our proposed perception-based objective speech quality measure is better than that of the state-of-the-art P.563 standard in estimating MOS of the codec-distorted speech in ITU-T Supp. 23 database.
Ting-Yu Yen, Jian-Hueng Chen, Tai-Shih Chi
ICASSP3