Szu-Jui Chen

dblp:217/2921 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-6406-9280ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
Szu-Jui Chen, John H. L. Hansen
Speech Commun.1
2025 Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
abstract
One challenge of integrating speech input with large language models (LLMs) stems from the discrepancy between the continuous nature of audio data and the discrete tokenbased paradigm of LLMs. To mitigate this gap, we propose a method for integrating vector quantization (VQ) into LLM-based automatic speech recognition (ASR). Using the LLM embedding table as the VQ codebook, the VQ module aligns the continuous representations from the audio encoder with the discrete LLM inputs, enabling the LLM to operate on a discretized audio representation that better reflects the linguistic structure. We further create a “soft discretization” of the audio representation by updating the codebook and performing a weighted sum over the codebook embeddings. Empirical results demonstrate that our proposed method significantly improves upon the LLMbased ASR baseline, particularly in out-of-domain conditions. This work highlights the potential of soft discretization as a modality bridge in LLM-based ASR.
Mu Yang, Szu-Jui Chen, Jiamin Xie, John H. L. Hansen
ASRU2
2025 A Neural Codec Approach for Noise-Robust Bandwidth Expansion
Mu Yang, Szu-Jui Chen, John H. L. Hansen
INTERSPEECH3
2024 Fearless Steps Apollo: Team Communications Based Community Resource Development for Science, Technology, Education, and Historical Preservation
abstract
The Fearless Steps Apollo (FS-APOLLO) resource is a collection of 150,000 hours of audio, associated meta-data, and supplemental speech technology infrastructure intended to benefit the (i) speech processing technology, (ii) communication science, team-based psychology, and (iii) education/STEM, history/preservation/archival communities. The FS-APOLLO initiative which started in 2014 has since resulted in the preservation of over 75,000 hours of NASA Apollo Missions audio. Systems created for this audio collection have led to the emergence of several new Speech and Language Technologies (SLT). This paper seeks to provide an overview of the latest advancements in the FS-Apollo effort and explore upcoming strategies in big-data deployment, outreach, and novel avenues of K-12 and STEM education facilitated through this resource.
John H. L. Hansen, Aditya Joglekar, Meena Chandra Shekar, Szu-Jui Chen
ICASSP4
2024 Dual-Path Minimum-Phase and All-Pass Decomposition Network for Single Channel Speech Dereverberation
abstract
With the development of deep neural networks (DNN), many DNN-based speech dereverberation approaches have been proposed to achieve significant improvement over the traditional methods. However, most deep learning-based dereverberation methods solely focus on suppressing time-frequency domain reverberations without utilizing cepstral domain features which are potentially useful for dereverberation. In this paper, we propose a dual-path neural network structure to separately process minimum-phase and all-pass components of single channel speech. First, we decompose speech signal into minimum-phase and all-pass components in cepstral domain, then Conformer embedded U-Net is used to remove reverberations of both components. Finally, we combine these two processed components together to synthesize the enhanced output. The performance of proposed method is tested on REVERB-Challenge evaluation dataset in terms of commonly used objective metrics. Experimental results demonstrate that our method outperforms other compared methods.
Szu-Jui Chen, John H. L. Hansen
ICASSP2
2023 Language Agnostic Data-Driven Inverse Text Normalization
Szu-Jui Chen, Debjyoti Paul, Yutong Pang
INTERSPEECH1
2022 FeaRLESS: Feature Refinement Loss for Ensembling Self-Supervised Learning Features in Robust End-to-end Speech Recognition
abstract
Self-supervised learning representations (SSLR) have resulted in robust features for downstream tasks in many fields.Recently, several SSLRs have shown promising results on automatic speech recognition (ASR) benchmark corpora.However, previous studies have only shown performance for solitary SSLRs as an input feature for ASR models.In this study, we propose to investigate the effectiveness of diverse SSLR combinations using various fusion methods within end-to-end (E2E) ASR models.In addition, we will show there are correlations between these extracted SSLRs.As such, we further propose a feature refinement loss for decorrelation to efficiently combine the set of input features.For evaluation, we show that the proposed "FeaRLESS learning features" perform better than systems without the proposed feature refinement loss for both the WSJ and Fearless Steps Challenge (FSC) corpora.
Szu-Jui Chen, Jiamin Xie, John H. L. Hansen
INTERSPEECH1
2022 Improving Data Driven Inverse Text Normalization using Data Augmentation and Machine Translation
Debjyoti Paul, Yutong Pang, Szu-Jui Chen
INTERSPEECH3
2021 Scenario Aware Speech Recognition: Advancements for Apollo Fearless Steps & CHiME-4 Corpora
abstract
In this study, we propose to investigate triplet loss for the purpose of an alternative feature representation for ASR. We consider a general non-semantic speech representation, which is trained with a self-supervised criteria based on triplet loss called TRILL, for acoustic modeling to represent the acoustic characteristics of each audio. This strategy is then applied to the CHiME-4 corpus and CRSS-UTDallas Fearless Steps Corpus, with emphasis on the 100-hour challenge corpus which consists of 5 selected NASA Apollo-11 channels. An analysis of the extracted embeddings provides the foundation needed to characterize training utterances into distinct groups based on acoustic distinguishing properties. Moreover, we also demonstrate that triplet-loss based embedding performs better than i-Vector in acoustic modeling, confirming that the triplet loss is more effective than a speaker feature. With additional techniques such as pronunciation and silence probability modeling, plus multi-style training, we achieve a +5.42% and +3.18% relative WER improvement for the development and evaluation sets of the Fearless Steps Corpus. To explore generalization, we further test the same technique on the 1 channel track of CHiME-4 and observe a +11.90% relative WER improvement for real test data.
Szu-Jui Chen, John H. L. Hansen
ASRU1
2019 Acoustic Modeling for Overlapping Speech Recognition: Jhu Chime-5 Challenge System
abstract
This paper summarizes our acoustic modeling efforts in the Johns Hopkins University speech recognition system for the CHiME-5 challenge to recognize highly-overlapped dinner party speech recorded by multiple microphone arrays. We explore data augmentation approaches, neural network architectures, front-end speech dereverberation, beamforming and robust i-vector extraction with comparisons of our in-house implementations and publicly available tools. We finally achieved a word error rate of 69.4% on the development set, which is a 11.7% absolute improvement over the previous baseline of 81.1%, and release this improved baseline with refined techniques/tools as an advanced CHiME-5 recipe.
Vimal Manohar, Szu-Jui Chen, Yusuke Fujita, Shinji Watanabe 0001, Sanjeev Khudanpur
ICASSP2
2018 Building State-of-the-art Distant Speech Recognition Using the CHiME-4 Challenge with a Setup of Speech Enhancement Baseline
abstract
This paper describes a new baseline system for automatic speech recognition (ASR) in the CHiME-4 challenge to promote the development of noisy ASR in speech processing communities by providing 1) state-of-the-art system with a simplified single system comparable to the complicated top systems in the challenge, 2) publicly available and reproducible recipe through the main repository in the Kaldi speech recognition toolkit.The proposed system adopts generalized eigenvalue beamforming with bidirectional long short-term memory (LSTM) mask estimation.We also propose to use a time delay neural network (TDNN) based on the lattice-free version of the maximum mutual information (LF-MMI) trained with augmented all six microphones plus the enhanced data after beamforming.Finally, we use a LSTM language model for lattice and n-best re-scoring.The final system achieved 2.74% WER for the real test set in the 6-channel track, which corresponds to the 2nd place in the challenge.In addition, the proposed baseline recipe includes four different speech enhancement measures, short-time objective intelligibility measure (STOI), extended STOI (eSTOI), perceptual evaluation of speech quality (PESQ) and speech distortion ratio (SDR) for the simulation test set.Thus, the recipe also provides an experimental platform for speech enhancement studies with these performance measures.
Szu-Jui Chen, Aswin Shanmugam Subramanian, Hainan Xu, Shinji Watanabe 0001
INTERSPEECH1
2018 Student-Teacher Learning for BLSTM Mask-based Speech Enhancement
abstract
Spectral mask estimation using bidirectional long short-term memory (BLSTM) neural networks has been widely used in various speech enhancement applications, and it has achieved great success when it is applied to multichannel enhancement techniques with a mask-based beamformer.However, when these masks are used for single channel speech enhancement they severely distort the speech signal and make them unsuitable for speech recognition.This paper proposes a studentteacher learning paradigm for single channel speech enhancement.The beamformed signal from multichannel enhancement is given as input to the teacher network to obtain soft masks.An additional cross-entropy loss term with the soft mask target is combined with the original loss, so that the student network with single-channel input is trained to mimic the soft mask obtained with multichannel input through beamforming.Experiments with the CHiME-4 challenge single channel track data shows improvement in ASR performance.
Aswin Shanmugam Subramanian, Szu-Jui Chen, Shinji Watanabe 0001
INTERSPEECH2