Jing-Tong Tzeng

dblp:382/2299 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-2053-0581ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Personalized Federated Learning with Fuzzy Clustering for Dysarthric Speech Recognition
abstract
Pathological speech recognition is challenging because clinical datasets are scarce, variable, and subject to strict privacy constraints preventing cross-institutional data sharing. These regulations necessitate federated learning (FL) for collaborative training without sharing raw data. However, FL degrades under non-IID data. Hard-clustering FL addresses this by partitioning clients into groups but imposes rigid boundaries, discards boundary samples, and suffers performance drops as cluster numbers increase. We propose Fuzzy Cluster-Based Personalized Federated Learning (FCPFL), using fuzzy C-means to softly group clients and pseudo-label-guided feature selection to identify discriminative features. FCPFL weights client updates by membership degree, allowing boundary samples to participate in multiple clusters and increasing training data by $25 \%$. Experiments show FCPFL reduces word error rate (WER) by $\mathbf{4 . 8 2 \%}$ and $\mathbf{1 . 7 4 \%}$ on ADReSS and TORGO, compared to hardclustered FL baselines.
Jie-Shiang Yang, Jing-Tong Tzeng, Chi-Chun Lee
ASRU2
2025 Valve Token Masked Autoencoder for Missing Recordings on Cardiac Abnormality Classification
abstract
Automated auscultation and cardiovascular screening systems for cardiac abnormalities have received growing interest in clinical applications. Still, they face challenges due to missing or invalid recordings caused by technical issues. To address this, we introduce a novel framework leveraging the masked autoencoder strategy, uniquely treating each heart valve recording as a distinct token. Our approach reconstructs missing valve data using existing representation and learnable mask tokens, achieving inter-valve integration through positional embeddings and TCNs. We demonstrate state-of-the-art (SOTA) performance on the CirCor DigiScope dataset, outperforming top participants and SOTA imputation methods in terms of mean cost of patient outcome, accuracy, F1-measure, and macro F1 score. Furthermore, our analysis highlights improved predictive accuracy on limited input data, while generative results indicate our capability to provide comprehensive reconstruction of auscultation recordings for further clinical evaluations.
An-Yan Chang, Jing-Tong Tzeng, Huan-Yu Chen, Chun-Hsiang Huang, Edward Pei-Chuan Huang, Chi-Chun Lee
ICASSP2
2025 Noise-Robust Speech Emotion Recognition Using Shared Self-Supervised Representations with Integrated Speech Enhancement
abstract
Recent studies have demonstrated the effectiveness of fine-tuning self-supervised speech representation models for speech emotion recognition (SER). However, applying SER in real-world environments remains challenging due to pervasive noise. Relying on low-accuracy predictions due to noisy speech can undermine the user’s trust. This paper proposes a unified self-supervised speech representation framework for enhanced speech emotion recognition designed to increase noise robustness in SER while generating enhanced speech. Our framework integrates speech enhancement (SE) and SER tasks, leveraging shared self-supervised learning (SSL)-derived features to improve emotion classification performance in noisy environments. This strategy encourages the SE module to enhance discriminative information for SER tasks. Additionally, we introduce a cascade unfrozen training strategy, where the SSL model is gradually unfrozen and fine-tuned alongside the SE and SER heads, ensuring training stability and preserving the generalizability of SSL representations. This approach demonstrates improvements in SER performance under unseen noisy conditions without compromising SE quality. When tested at a 0 dB signal-to-noise ratio (SNR) level, our proposed method outperforms the original baseline by 3.7% in F1-Macro and 2.7% in F1-Micro scores, where the differences are statistically significant.
Jing-Tong Tzeng, Seong-Gyun Leem, Ali N. Salman, Chi-Chun Lee, Carlos Busso
ICASSP1
2025 Mask Augmentation For Tumor Classification In Medical Images
abstract
Tumor detection and classification in medical images are critical for guiding patient management and treatment decisions. However, accurate segmentation and classification of tumors remain challenging due to their small size relative to the overall image. Existing approaches often face difficulties with limited data and potential segmentation errors, resulting in suboptimal performance in real-world applications. To address these challenges, we propose a novel approach leveraging the Segmentation Mask Augmentation (SMA) framework. Our framework enhances the robustness of tumor classification models by generating diverse and imprecise segmentation masks during training, thereby simulating real-world scenarios. Experimental results across four distinct datasets demonstrate the effectiveness of our approach. Our framework presents a promising solution for robust tumor classification, with potential implications for improving clinical diagnosis and patient management.
Chun-Chieh Weng, Huan-Yu Chen, Jing-Tong Tzeng, Ching-Heng Lin, Po-Chih Kuo, Chi-Chun Lee
ICASSP3
2025 Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
abstract
In this study, we revisit key training strategies in machine learning often overlooked in favor of deeper architectures. Specifically, we explore balancing strategies, activation functions, and fine-tuning techniques to enhance speech emotion recognition (SER) in naturalistic conditions. Our findings show that simple modifications improve generalization with minimal architectural changes. Our multi-modal fusion model, integrating these optimizations, achieves a valence CCC of 0.6953, the best valence score in Task 2: Emotional Attribute Regression. Notably, fine-tuning RoBERTa and WavLM separately in a single-modality setting, followed by feature fusion without training the backbone extractor, yields the highest valence performance. Additionally, focal loss and activation functions significantly enhance performance without increasing complexity. These results suggest that refining core components, rather than deepening models, leads to more robust SER in-the-wild.
Jing-Tong Tzeng, Bo-Hao Su, Ya-Tse Wu, Hsing-Hang Chou, Chi-Chun Lee
INTERSPEECH1
2024 GaP-Aug: Gamma Patch-Wise Correction Augmentation Method for Respiratory Sound Classification
abstract
Automated auscultation analysis using electronic stethoscope has received growing interest in clinical applications. Recently, researchers showed successes by using deep learning methods to distinguish between pathological respiratory sound classes. Nevertheless, the challenge persists due to the scarcity of abnormal samples, and the distinct characteristics between low-pitched and discontinuous crackles and high-pitched and continuous wheezes. In this study, we proposed a novel augmentation method, namely gamma patch-wise correction augmentation, which directly operates on spectrograms to handle with these two challenges. We achieved state-of-the-art performances on both 60-40 official split and 80-20 cross-validation of the public ICBHI dataset, outperforming previous top-performing studies by 11.82% in sensitivity and 5.27% in ICBHI score. Furthermore, Grad-CAM analysis shows that our approach better preserves the distinctive characteristics of crackles and wheezes than SpecAug.
An-Yan Chang, Jing-Tong Tzeng, Huan-Yu Chen, Chih-Wei Sung, Chun-Hsiang Huang, Edward Pei-Chuan Huang, Chi-Chun Lee
ICASSP2