Ismail Rasim Ülgen

dblp:283/9046 · also Ismail Rasim Ulgen · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Can Emotion Fool Anti-spoofing?
Aurosweta Mahapatra, Ismail Rasim Ülgen, Abinay Reddy Naini, Carlos Busso, Berrak Sisman
INTERSPEECH2
2025 The Interspeech 2025 Challenge on Speech Emotion Recognition in Naturalistic Conditions
Abinay Reddy Naini, Lucas Goncalves, Ali N. Salman, Pravin Mote, Ismail Rasim Ülgen, Thomas Thebaud, Laureano Moro-Velázquez, L. Paola García-Perera, Najim Dehak, Berrak Sisman, Carlos Busso
INTERSPEECH5
2024 Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
abstract
Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assumed that emotion information is indirectly embedded within speaker embeddings, leading to their under-utilization. Our study reveals a direct and useful link between emotion and state-of-the-art speaker embeddings in the form of intra-speaker clusters. By conducting a thorough clustering analysis, we demonstrate that emotion information can be readily extracted from speaker embeddings. In order to leverage this information, we introduce a novel contrastive pretraining approach applied to emotion-unlabeled data for speech emotion recognition. The proposed approach involves the sampling of positive and the negative examples based on the intra-speaker clusters of speaker embeddings. The proposed strategy, which leverages extensive emotion-unlabeled data, leads to a significant improvement in SER performance, whether employed as a standalone pretraining task or integrated into a multi-task pretraining setting.
Ismail Rasim Ülgen, Zongyang Du, Carlos Busso, Berrak Sisman
ICASSP1
2024 Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
Ali N. Salman, Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim Ülgen, Carlos Busso, Berrak Sisman
INTERSPEECH4
2024 Discrete Unit Based Masking For Improving Disentanglement in Voice Conversion
abstract
Voice conversion (VC) aims to modify the speaker’s identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling the speaker’s identity from linguistic information is crucial. However, the disentanglement approaches used in these methods are limited as the speaker features depend on the phonetic content of the utterance, compromising disentanglement. This dependency is amplified with attention-based methods. To address this, we introduce a novel masking mechanism in the input before speaker encoding, masking certain discrete speech units that correspond highly with phoneme classes. Our work aims to reduce the phonetic dependency of speaker features by restricting access to some phonetic information. Furthermore, since our approach is at the input level, it is applicable to any encoder-decoder based VC framework. Our approach improves disentanglement and conversion performance across multiple VC methods, showing significant effectiveness, particularly in attention-based method, with 44% relative improvement in objective intelligibility.
Philip H. Lee, Ismail Rasim Ülgen, Berrak Sisman
SLT2
2022 Unsupervised Domain Adaptation of Neural PLDA Using Segment Pairs for Speaker Verification
abstract
Probabilistic linear discriminant analysis (PLDA) is a popular tool in speaker recognition. PLDA parameters are estimated from labeled data, and domain shift causes performance degradation. Since obtaining a labeled dataset for target domain is costly, we propose an unsupervised domain adaptation of PLDA which does not require clustering or alignment. Our proposed method relies on discriminative training of PLDA using pseudo target/non-target labels. Segment pairs extracted from the same utterance that is likely to contain a single speaker are labeled as target, and pairs from randomly sampled utterances that are likely to belong to different speakers are labeled as non-target. We applied this approach to Neural PLDA. Our approach significantly improved the performance of the out-of-domain PLDA on the target domain with a relatively small amount of unlabeled data, performing on par with a baseline adaptation approach on SRE18 full-full scenario while outperforming it on full-10s and 10s-10s cases.
Ismail Rasim Ülgen, Levent M. Arslan
SLT1