Jiarui Hai

dblp:321/6772 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-9968-7372ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
abstract
In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both audio-oriented and language-oriented TSE by utilizing a CLAP model as the feature extractor for target sounds. Furthermore, SoloAudio leverages synthetic audio generated by state-of-the-art text-to-audio models for training, demonstrating strong generalization to out-of-domain data and unseen sound events. We evaluate this approach on the FSD Kaggle 2018 mixture dataset and real data from AudioSet, where SoloAudio achieves the state-of-the-art results on both in-domain and out-of-domain data, and exhibits impressive zero-shot and few-shot capabilities. Source code1and demos2are released.
Helin Wang, Jiarui Hai, Yen-Ju Lu, Karan Thakkar, Mounya Elhilali, Najim Dehak
ICASSP2
2025 SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
abstract
In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. The source code1and demos2are released.
Helin Wang, Meng Yu 0003, Jiarui Hai, Chen Chen 0075, Rilin Chen, Najim Dehak, Dong Yu 0001
ICASSP3
2025 EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Jiarui Hai, Yong Xu 0004, Hao Zhang 0112, Chenxing Li, Helin Wang, Mounya Elhilali, Dong Yu 0001
INTERSPEECH1
2024 DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction
abstract
Common target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in separating the target from the background. This study introduces DPM-TSE, a generative method based on diffusion probabilistic modeling (DPM) for Target Sound Extraction (TSE), to achieve both cleaner target renderings as well as improved separability from unwanted sounds. The technique also tackles the noise floor of DPM by introducing a correction method for noise schedules and sample steps. This approach is evaluated using both objective and subjective quality metrics on the FSD Kaggle 2018 dataset. The results show that DPM-TSE has a significant improvement in perceived quality in terms of target extraction and purity.
Jiarui Hai, Helin Wang, Dongchao Yang, Karan Thakkar, Najim Dehak, Mounya Elhilali
ICASSP1
2024 Investigating Self-Supervised Deep Representations for EEG-Based Auditory Attention Decoding
abstract
Auditory Attention Decoding (AAD) algorithms play a crucial role in isolating desired sound sources within challenging acoustic environments directly from brain activity. Although recent research has shown promise in AAD using shallow representations such as auditory envelope and spectrogram, there has been limited exploration of deep Self-Supervised (SS) representations on a larger scale. In this study, we undertake a comprehensive investigation into the performance of linear decoders across 12 deep and 2 shallow representations, applied to EEG data from multiple studies spanning 57 subjects and multiple languages. Our experimental results consistently reveal the superiority of deep features for AAD at decoding background speakers, regardless of the datasets and analysis windows. This result indicates possible nonlinear encoding of unattended signals in the brain that are revealed using deep nonlinear features. Additionally, we analyze the impact of different layers of SS representations and window sizes on AAD performance. These findings underscore the potential for enhancing EEG-based AAD systems through the integration of deep feature representations.
Karan Thakkar, Jiarui Hai, Mounya Elhilali
ICASSP2
2024 DreamVoice: Text-Guided Voice Conversion
Jiarui Hai, Karan Thakkar, Helin Wang, Zengyi Qin, Mounya Elhilali
INTERSPEECH1
2024 Noise-robust Speech Separation with Fast Generative Correction
Helin Wang, Jesús Villalba 0001, Laureano Moro-Velázquez, Jiarui Hai, Thomas Thebaud, Najim Dehak
INTERSPEECH4
2023 Boosting Modality Representation With Pre-Trained Models and Multi-Task Training for Multimodal Sentiment Analysis
abstract
Sentiment analysis has traditionally leveraged information from text data. More recently, it has become increasingly clear that multimodal data provides a rich space to drastically boost interpretation of human sentiments by harnessing information across multiple modalities. In this study, we incorporate pre-trained feature extractors and propose a multitask training strategy to improve modality representations for Multimodal Sentiment Analysis (MSA). The experimental results on the CH-SIMS v2 dataset demonstrate the superior performance of the proposed system compared to existing state-of-the-art methods, validating the effectiveness of our proposed approach. Furthermore, our framework reduces reliance on textual data, achieving competitive outcomes even when utilizing only auditory and visual modalities.
Jiarui Hai, Yu-Jeh Liu, Mounya Elhilali
ASRU1
2022 Leveraging Natural Language Processing and Time Series Models to Analyze COVID-19 Vaccination Sentiment Dynamics from Tweets
Jiancheng Ye, Jiarui Hai, Zidan Wang, Chumei Wei
AMIA2
2022 Progressive Teacher-Student Training Framework for Music Tagging
abstract
Music tagging is the task of predicting multiple tags of a music excerpt, and plays an important role in modern music recommendation systems. To obtain superior performance, recent approaches of music tagging focus on developing sophisticated models or exploiting additional multi-modal information. However, none of them deal with the problem of label noise during the training process despite of its ubiquitous presence. In this paper, we propose a progressive two-stage teacher-student training framework to prevent the music tagging model from overfitting label noise. Experimental results suggest that the proposed method surpasses conventional label-noise-robust methods and exhibits scalability across different tagging models. Moreover, detailed analyses demonstrate that the two teachers in the framework gradually improve student model’s generalization performance and effectively avoid the impairment from label noise.
Rui Lu 0003, Baigong Zheng, Jiarui Hai, Zhiyao Duan, Ji Liu 0002
ICASSP3