VLDB 2026 Research / reviewers in the wild / expert
Tai-Shih Chi
dblp:42/8052
· DBLP profile ↗
34ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-0584-8399ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 6 since 2021Artificial intelligence and machine learning · 15 · 6 since 2021Systems, architecture and hardware · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tonality-Based Accompaniment-Guided Automatic Singing Evaluation
Pei-Chin Hsieh, Yih-Liang Shen, Ngoc Son Tran, Tai-Shih Chi |
INTERSPEECH | 4 |
| 2025 | Spectro-Temporal Modulations Incorporated Two-Stream Robust Speech Emotion RecognitionabstractDeep learning based speech emotion recognition (SER) models have shown impressive results in controlled environments, but their performance significantly degrades in noisy conditions. This paper proposes a robust two-stream SER model by combining spectro-temporal modulation features with conventional acoustic features. Experiments were conducted on German (EMODB) and English (RAVDESS) datasets using the clean-train-noisy-test paradigm. The results demonstrate that spectro-temporal modulation features offer superior robustness in noisy conditions compared with conventional acoustic features such as MFCCs and time-frequency features from Mel-spectrograms. Additionally, we analyze weights of modulation features and demonstrate the model emphasizes contours of formants and harmonics, which are crucial features for speech perception in noise, for robust SER. Incorporating the stream of spectro-temporal modulations not only enhances the robustness of the model but also provides deeper insights into the task of SER in noise. Yih-Liang Shen, Pei-Chin Hsieh, Tai-Shih Chi |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
Chun Yin, Tai-Shih Chi, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 2 |
| 2023 | Mandarin Electrolaryngeal Speech Voice Conversion using Cross-domain Features
Hsin-Hao Chen 0006, Yung-Lun Chien, Ming-Chi Yen, Shu-Wei Tsai, Tai-Shih Chi, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 5 |
| 2023 | Audio-Visual Mandarin Electrolaryngeal Speech Voice Conversion
Yung-Lun Chien, Hsin-Hao Chen 0006, Ming-Chi Yen, Shu-Wei Tsai, Hsin-Min Wang, Yu Tsao 0001, Tai-Shih Chi |
INTERSPEECH | 7 |
| 2022 | Perceptual Characteristics Based Multi-objective Model for Speech Enhancement
Chiang-Jen Peng, Yun-Ju Chan, Yih-Liang Shen, Yu Tsao 0001, Tai-Shih Chi |
INTERSPEECH | 6 |
| 2021 | Extending Music Based On Emotion And Tonality Via Generative Adversarial NetworkabstractWe propose a generative model for music extension in this paper. The model is composed of two classifiers, one for music emotion and one for music tonality, and a generative adversarial network (GAN). Therefore, it can generate symbolic music not only based on low level spectral and temporal characteristics, but also on high level emotion and tonality attributes of previously observed music pieces. The generative model works in a universal latent space constructed by the variational autoencoder (VAE) for representing music pieces. We conduct subjective listening tests and derive objective measures for performance evaluation. Experimental results show that the proposed model produces much smoother and more authentic music pieces than the baseline model in terms of all subjective and objective measures. Bo-Wei Tseng, Yih-Liang Shen, Tai-Shih Chi |
ICASSP | 3 |
| 2021 | Attention-Based Multi-Task Learning for Speech-Enhancement and Speaker-Identification in Multi-Speaker Dialogue ScenarioabstractMulti-task learning (MTL) and attention mechanism have been proven to effectively extract robust acoustic features for various speech-related tasks in noisy environments. In this study, we propose an attention-based MTL (ATM) approach that integrates MTL and the attention-weighting mechanism to simultaneously realize a multi-model learning structure that performs speech enhancement (SE) and speaker identification (SI). The proposed ATM system consists of three parts: SE, SI, and attention-Net (AttNet). The SE part is composed of a long-short-term memory (LSTM) model, and a deep neural network (DNN) model is used to develop the SI and AttNet parts. The overall ATM system first extracts the representative features and then enhances the speech signals in LSTM-SE and specifies speaker identity in DNN-SI. The AttNet computes weights based on DNN-SI to prepare better representative features for LSTM-SE. We tested the proposed ATM system on Taiwan Mandarin hearing in noise test sentences. The evaluation results confirmed that the proposed system can effectively enhance speech quality and intelligibility of a given noisy input. Moreover, the accuracy of the SI can also be notably improved by using the proposed ATM system. Chiang-Jen Peng, Yun-Ju Chan, Syu-Siang Wang, Yu Tsao 0001, Tai-Shih Chi |
ISCAS | 6 |
| 2020 | A Multi-Dilation and Multi-Resolution Fully Convolutional Network for Singing Melody ExtractionabstractEach human cognitive function involves bottom-up and top-down processes. Several methods have been proposed for singing melody extraction by emphasizing either the bottom-up or top-down processes. For hearing, the bottom-up processes include spectral and spectro-temporal decomposition of the sound by the cochlea and the auditory cortex. In this paper, we propose a neural network, which includes spectro-temporal multi-resolution decomposition of the log-spectrogram of the sound and a semantic segmentation model to respectively address the bottom-up and top-down processing of hearing, for singing melody extraction. Simulation results show the proposed model outperforms all previously proposed methods, emphasizing either bottom-up or top-down processing, in almost all objective evaluation metrics. Cheng-You You, Tai-Shih Chi |
ICASSP | 3 |
| 2019 | Autoencoding HRTFS for DNN Based HRTF Personalization Using Anthropometric FeaturesabstractWe proposed a deep neural network (DNN) based approach to synthesize the magnitude of personalized head-related transfer functions (HRTFs) using anthropometric features of the user. To mitigate the over-fitting problem when training dataset is not very large, we built an autoencoder for dimensional reduction and establishing a crucial feature set to represent the raw HRTFs. Then we combined the decoder part of the autoencoder with a smaller DNN to synthesize the magnitude HRTFs. In this way, the complexity of the neural networks was greatly reduced to prevent unstable results with large variance due to overfitting. The proposed approach was compared with a baseline DNN model with no autoencoder. The log-spectral distortion (LSD) metric was used to evaluate the performance. Experiment results show that the proposed approach can reduce LSD of estimated HRTFs with greater stability. Tzu-Yu Chen, Tzu-Hsuan Kuo, Tai-Shih Chi |
ICASSP | 3 |
| 2019 | CNN Based Two-stage Multi-resolution End-to-end Model for Singing Melody ExtractionabstractInspired by human hearing perception, we propose a two-stage multi-resolution end-to-end model for singing melody extraction in this paper. The convolutional neural network (CNN) is the core of the proposed model to generate multi-resolution representations. The 1-D and 2-D multi-resolution analysis on waveform and spectrogram-like graph are successively carried out by using 1-D and 2-D CNN kernels of different lengths and sizes. The 1-D CNNs with kernels of different lengths produce multi-resolution spectrogram-like graphs without suffering from the trade-off between spectral and temporal resolutions. The 2-D CNNs with kernels of different sizes extract features from spectro-temporal envelopes of different scales. Experiment results show the proposed model outperforms three compared systems in three out of five public databases. Ming-Tso Chen, Bo-Jun Li, Tai-Shih Chi |
ICASSP | 3 |
| 2019 | Reinforcement Learning Based Speech Enhancement for Robust Speech RecognitionabstractConventional deep neural network (DNN)-based speech enhancement (SE) approaches aim to minimize the mean square error (MSE) between enhanced speech and clean reference. The MSE-optimized model may not directly improve the performance of an automatic speech recognition (ASR) system. If the target is to minimize the recognition error, the recognition results should be used to design the objective function for optimizing the SE model. However, the structure of an ASR system, which consists of multiple units, such as acoustic and language models, is usually complex and not differentiable. In this study, we propose to adopt the reinforcement learning (RL) algorithm to optimize the SE model based on the recognition results. We evaluated the proposed RL-based SE system on the Mandarin Chinese broadcast news corpus (MATBN). Experimental results demonstrate that the proposed SE system can effectively improve the ASR results with a notable 12:40% and 19:23% error rate reductions for signal to noise ratio (SNR) at 0 dB and 5 dB conditions, respectively. Yih-Liang Shen, Chao-Yuan Huang, Syu-Siang Wang, Yu Tsao 0001, Hsin-Min Wang, Tai-Shih Chi |
ICASSP | 6 |
| 2018 | A Hybrid Neural Network Based on the Duplex Model of Pitch Perception for Singing Melody ExtractionabstractIn this paper, we build up a hybrid neural network (NN) for singing melody extraction from polyphonic music by imitating human pitch perception. For human hearing, there are two pitch perception models, the spectral model and the temporal model, in accordance with whether harmonics are resolved or not. Here, we first use NNs to implement individual models and evaluate their performance in the task of singing melody extraction. Then, we combine the NNs to constitute the composite NN to simulate the duplex model, which complements the pitch perception from unresolved harmonics of the spectral model using the temporal model. Simulation results show the proposed composite NN outperforms other conventional methods in singing melody extraction. Hsin Chou, Ming-Tso Chen, Tai-Shih Chi |
ICASSP | 3 |
| 2018 | A Generative Auditory Model Embedded Neural Network for Speech ProcessingabstractBefore the era of the neural network (NN), features extracted from auditory models have been applied to various speech applications and been demonstrated more robust against noise than conventional speech-processing features. What's the role of auditory models in the current NN era? Are they obsolete? To answer this question, we construct a NN with a generative auditory model embedded to process speech signals. The generative auditory model consists of two stages, the stage of spectrum estimation in the logarithmic-frequency axis by the cochlea and the stage of spectral-temporal analysis in the modulation domain by the auditory cortex. The NN is evaluated in a simple speaker identification task. Experiment results show that the auditory model embedded NN is still more robust against noise, especially in low SNR conditions, than the randomly-initialized NN in speaker identification. Yu-Wen Lo, Yih-Liang Shen, Yuan-Fu Liao, Tai-Shih Chi |
ICASSP | 4 |
| 2018 | Singing Voice Correction Using Canonical Time WarpingabstractExpressive singing voice correction is an appealing but challenging problem. A robust time-warping algorithm which synchronizes two singing recordings can provide a promising solution. We thereby propose to address the problem by canonical time warping (CTW) which aligns amateur singing recordings to professional ones. A new pitch contour is generated given the alignment information, and a pitch-corrected singing is synthesized back through the vocoder. The objective evaluation shows that CTW is robust against pitch-shifting and time-stretching effects, and the subjective test demonstrates that CTW prevails the other methods including DTW and the commercial auto-tuning software. Finally, we demonstrate the applicability of the proposed method in a practical, real-world scenario. Yin-Jyun Luo, Ming-Tso Chen, Tai-Shih Chi, Li Su 0004 |
ICASSP | 3 |
| 2017 | Dereverberation based on bin-wise temporal variations of complex spectrogramabstractHumans analyze sounds not only based on their frequency contents, but also on the temporal variations of the frequency contents. Inspired by auditory perception, we propose a deep neural network (DNN) based dereverberation algorithm in the rate domain, which presents the temporal variations of frequency contents, in this paper. We show convolutional noise in the time domain can be approximated to multiplicative noise in the rate domain. To remove the multiplicative noise, we adopt the rate-domain complex-valued ideal ratio mask (RDcIRM) as the training target of the DNN. Simulation results show that the proposed rate-domain DNN algorithm is more capable of recovering high-intelligible and high-quality speech from reverberant speech than the compared state-of-the-art dereverberation algorithm. Hence, it is highly suitable for speech applications involving human listeners. Tai-Shih Chi |
ICASSP | 3 |
| 2017 | Simulations of High-Frequency Vocoder on Mandarin Speech Recognition for Acoustic Hearing Preserved Cochlear Implant
Tsung-Chen Wu, Tai-Shih Chi, Chia-Fone Lee |
INTERSPEECH | 2 |
| 2016 | A Spectral Modulation Sensitivity Weighted Pre-Emphasis Filter for Active Noise Control System
Kah-Meng Cheong, Yuh-Yuan Wang, Tai-Shih Chi |
INTERSPEECH | 3 |
| 2016 | Discriminative Layered Nonnegative Matrix Factorization for Speech Separation
Chung-Chien Hsu, Tai-Shih Chi, Jen-Tzung Chien |
INTERSPEECH | 2 |
| 2015 | Modulation Wiener filter for improving speech intelligibilityabstractThis paper presents a single-channel high-dimensional Wiener filter in the spectro-temporal modulation domain. Unlike other conventional noise reduction techniques, the proposed algorithm not only reduces noise but also enhances the “textures” of the speech signal. A non-iterative decision-directed noise estimation method is adopted to estimate the modulation SNR for the modulation-domain Wiener filter. The efficacy of the proposed algorithm in enhancing speech intelligibility is assessed using the short-time objective intelligibility (STOI) measure. Statistical analysis results demonstrate that our proposed algorithm can improve STOI scores in speech-shape noise (SSN) and white noise conditions, but not in babble noise condition, while the conventional Wiener filter fails to improve STOI scores in all three noise conditions. Chung-Chien Hsu, Kah-Meng Cheong, Jen-Tzung Chien, Tai-Shih Chi |
ICASSP | 4 |
| 2015 | A hearing model to estimate mandarin speech intelligibility for the hearing impaired patientsabstractA hearing model, which is parameterized by hearing thresholds, degrees of loudness recruitment and reductions of frequency resolution of a hearing-impaired (HI) patient, is proposed in this paper. The model is developed in the filter-bank framework and is flexible for fitting hearing-loss conditions of HI patients. Psychoacoustic experiments were conducted under clean and noisy conditions to validate the model's capability in predicting Mandarin speech intelligibility for HI patients. Statistical analysis on the hearing-test results suggests that the proposed model can predict Mandarin speech intelligibility for HI patients to a certain degree. Pei-Chun Tsai, Shih-Ting Lin, Wen-Chung Lee, Chung-Chien Hsu, Tai-Shih Chi, Chia-Fone Lee |
ICASSP | 5 |
| 2015 | Layered nonnegative matrix factorization for speech separation
Chung-Chien Hsu, Jen-Tzung Chien, Tai-Shih Chi |
INTERSPEECH | 3 |
| 2015 | A two-stage singing voice separation algorithm using spectro-temporal modulation features
Frederick Z. Yen, Mao-Chang Huang, Tai-Shih Chi |
INTERSPEECH | 3 |
| 2014 | Binary mask estimation based on frequency modulations
Chung-Chien Hsu, Jen-Tzung Chien, Tai-Shih Chi |
INTERSPEECH | 3 |
| 2013 | Voice activity detection based on frequency modulation of harmonicsabstractIn this paper, we propose a voice activity detection (VAD) algorithm based on spectro-temporal modulation structures of input sounds. A multi-resolution spectro-temporal analysis framework is used to inspect prominent speech structures. By comparing with an adaptive threshold, the proposed VAD distinguishes speech from non-speech based on the energy of the frequency modulation of harmonics. Compared with three standard VADs, ITU-T G.729B, ETSI AMR1 and AMR2, our proposed VAD significantly outperforms them in non-stationary noises in terms of the receiver operating characteristic (ROC) curves and the recognition rates from a practical distributed speech recognition (DSR) system. Chung-Chien Hsu, Tse-En Lin, Jian-Hueng Chen, Tai-Shih Chi |
ICASSP | 4 |
| 2013 | Spectral modulation sensitivity based perceptual acoustic echo cancellation
Wei-Lun Chuang, Kah-Meng Cheong, Chung-Chien Hsu, Tai-Shih Chi |
INTERSPEECH | 4 |
| 2013 | Spectro-temporal modulation based singing detection combined with pitch-based grouping for singing voice separation
Tse-En Lin, Chung-Chien Hsu, Jian-Hueng Chen, Tai-Shih Chi |
INTERSPEECH | 5 |
| 2013 | A precedence effect based far-field DoA estimation algorithmabstractA robust far-field DoA estimation algorithm is proposed in this paper. The algorithm is inspired by the precedence effect which results in humans robust capability in localizing sound sources. The algorithm implements the concept of the precedence effect by applying a proper threshold and an onset detection mechanism in the cross-correlation domain. Experiment results show that by cascading the proposed algorithm to a conventional cross-correlation based DoA algorithm, the accuracy of the DoA estimation is significantly improved in far-field test conditions. Wen-Sheng Chou, Tai-Shih Chi |
ISCAS | 2 |
| 2012 | Spectro-temporal subband Wiener filter for speech enhancementabstractIn this paper, we propose a signal-channel speech enhancement algorithm by applying the conventional Wiener filter in the spectro-temporal modulation domain. The multi-resolution spectro-temporal analysis and synthesis framework for Fourier spectrograms [12] is extended to the analysis-modification-synthesis (AMS) framework for speech enhancement. Compared with conventional speech enhancement algorithms, a Wiener filter and an extended minimum mean-square error (MMSE) algorithm, our proposed method outperforms them by a large/small margin in white/babble noise conditions from both objective and subjective evaluations. Chung-Chien Hsu, Tse-En Lin, Jian-Hueng Chen, Tai-Shih Chi |
ICASSP | 4 |
| 2011 | A binaural algorithm for space and pitch detectionabstractA binaural algorithm to simultaneously detect the azimuth angle and the pitch of the sound source is proposed in this paper. This algorithm is extended from the stereausis model with two-dimensional coincidence detectors in the joint Space-Pitch domain. In our simulations, sounds from different locations are produced by passing through the Head-Related-Transfer-Function (HRTF). Simulation results show that estimated azimuth angles from our proposed algorithm are more accurate than those from the stereausis model in the single sound source testing condition. Satisfactory results in streaming sound sources from the two-sound mixture by using estimated Space-Pitch information are also demonstrated in our pilot experiments. Wen-Sheng Chou, Kah-Meng Cheong, Tai-Shih Chi |
ICASSP | 3 |
| 2011 | FFT-based spectro-temporal analysis and synthesis of soundsabstractThe concept of the two-dimensional spectro-temporal modulation filtering of the auditory model is implemented for the FFT spectrogram. It analyzes the spectrogram in terms of the temporal dynamics and the spectral structures of the sound. The overlap and add (OTA) method, which is more convenient and reliable than the iterative-projection method proposed in, is used to invert the FFT spectrogram back to sounds. The Non-Negative Sparse Coding (NNSC) method is adopted to demonstrate the benefit of our analysis-synthesis procedures in a noise suppression application. Even without fine-tuning parameters, our proposed analysis-synthesis procedures offer benefits in de-noising especially under low SNR conditions. Chung-Chien Hsu, Ting-Han Lin, Tai-Shih Chi |
ICASSP | 3 |
| 2011 | Low power InfomaxICA with compensation strategy for binaural hearing-aidabstractBinaural hearing-aids are under intensive study nowadays. The information exchange across both ears provides an opportunity to perform the blind source separation to enhance the signal SNR. This study combines the conventional InfomaxICA with our proposed novel algorithms: binaural delay compensation, minimum-interference initial de-mixing guess and low power design concept, for a real-time binaural hearing-aid application. Simulations demonstrate our algorithms provide 20.5 to 30.6 dB gain at 0 to 35 points delay at SNR= 0 dB in our assumed scenario. Therefore, our algorithms are robust under various practical wearing conditions. Fan-Chiang Yi, Ching-Wen Huang, Tai-Shih Chi, Shyh-Jye Jou |
ISCAS | 3 |
| 2010 | Spectro-temporal modulations for robust speech emotion recognition
Lan-Ying Yeh, Tai-Shih Chi |
INTERSPEECH | 2 |
| 2009 | Perception-based objective speech quality assessmentabstractA joint spectro-temporal auditory model is utilized to assess speech quality objectively. The model mimics early and central auditory functions and serves as a spectro-temporal modulation filterbank. Three perceptual relevant parameters, intelligibility, clarity and naturalness, are addressed by the model and are combined to estimate the subjective mean opinion score (MOS) for speech quality measure. Through a simple multiple linear regression analysis, we demonstrate the performance of our proposed perception-based objective speech quality measure is better than that of the state-of-the-art P.563 standard in estimating MOS of the codec-distorted speech in ITU-T Supp. 23 database. Ting-Yu Yen, Jian-Hueng Chen, Tai-Shih Chi |
ICASSP | 3 |