Filip Sedlak

dblp:84/7972 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorArtificial intelligence and machine learning · 2Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Kernel, tree and ensemble methods · 67% Speech recognition and synthesis · 33%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Kernel, tree and ensemble methods
classifier combination
0.212013
Sparse Classifier Fusion for Speaker Verification · IEEE Trans. Speech Audio Process. 2013
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble selection
0.212013
Sparse Classifier Fusion for Speaker Verification · IEEE Trans. Speech Audio Process. 2013
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker verification
0.212013
Sparse Classifier Fusion for Speaker Verification · IEEE Trans. Speech Audio Process. 2013
Audio and music processing › speaker recognition
speaker verification
0.112012
Low-Variance Multitaper MFCC Features: A Case Study in Robust Speaker Verification · IEEE Trans. Speech Audio Process. 2012
Audio and music processing › audio feature extraction
mel-frequency cepstral coefficients
0.012012
Low-Variance Multitaper MFCC Features: A Case Study in Robust Speaker Verification · IEEE Trans. Speech Audio Process. 2012
Audio and music processing › audio feature extraction
speech feature extraction
0.012012
Low-Variance Multitaper MFCC Features: A Case Study in Robust Speaker Verification · IEEE Trans. Speech Audio Process. 2012

Methods — techniques the papers use, named apart from their topics

sparse regularization · 0.2logistic regression · 0.2support vector machine · 0.1multitaper spectral estimation · 0.1joint factor analysis · 0.1gaussian mixture model · 0.1
YearPublicationVenuePosition
2013 Sparse Classifier Fusion for Speaker Verification
abstract
State-of-the-art speaker verification systems take advantage of a number of complementary base classifiers by fusing them to arrive at reliable verification decisions. In speaker verification, fusion is typically implemented as a weighted linear combination of the base classifier scores, where the combination weights are estimated using a logistic regression model. An alternative way for fusion is to use classifier ensemble selection, which can be seen as sparse regularization applied to logistic regression. Even though score fusion has been extensively studied in speaker verification, classifier ensemble selection is much less studied. In this study, we extensively study a sparse classifier fusion on a collection of twelve I4U spectral subsystems on the NIST 2008 and 2010 speaker recognition evaluation (SRE) corpora.
Ville Hautamäki, Tomi Kinnunen, Filip Sedlak, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001
IEEE Trans. Speech Audio Process.3
2012 Vulnerability of speaker verification systems against voice conversion spoofing attacks: The case of telephone speech
abstract
Voice conversion - the methodology of automatically converting one's utterances to sound as if spoken by another speaker - presents a threat for applications relying on speaker verification. We study vulnerability of text-independent speaker verification systems against voice conversion attacks using telephone speech. We implemented a voice conversion systems with two types of features and nonparallel frame alignment methods and five speaker verification systems ranging from simple Gaussian mixture models (GMMs) to state-of-the-art joint factor analysis (JFA) recognizer. Experiments on a subset of NIST 2006 SRE corpus indicate that the JFA method is most resilient against conversion attacks. But even it experiences more than 5-fold increase in the false acceptance rate from 3.24 % to 17.33 %.
Tomi Kinnunen, Zhizheng Wu 0001, Kong-Aik Lee, Filip Sedlak, Chng Eng Siong, Haizhou Li 0001
ICASSP4
2012 Low-Variance Multitaper MFCC Features: A Case Study in Robust Speaker Verification
abstract
In speech and audio applications, short-term signal spectrum is often represented using mel-frequency cepstral coefficients (MFCCs) computed from a windowed discrete Fourier transform (DFT). Windowing reduces spectral leakage but variance of the spectrum estimate remains high. An elegant extension to windowed DFT is the so-called multitaper method which uses multiple time-domain windows (tapers) with frequency-domain averaging. Multitapers have received little attention in speech processing even though they produce low-variance features. In this paper, we propose the multitaper method for MFCC extraction with a practical focus. We provide, first, detailed statistical analysis of MFCC bias and variance using autoregressive process simulations on the TIMIT corpus. For speaker verification experiments on the NIST 2002 and 2008 SRE corpora, we consider three Gaussian mixture model based classifiers with universal background model (GMM-UBM), support vector machine (GMM-SVM) and joint factor analysis (GMM-JFA). Multitapers improve MinDCF over the baseline windowed DFT by relative 20.4% (GMM-SVM) and 13.7% (GMM-JFA) on the interview-interview condition in NIST 2008. The GMM-JFA system further reduces MinDCF by 18.7% on the telephone data. With these improvements and generally noncritical parameter selection, multitaper MFCCs are a viable candidate for replacing the conventional MFCCs.
Tomi Kinnunen, Rahim Saeidi, Filip Sedlak, Kong-Aik Lee, Johan Sandberg, Maria Sandsten, Haizhou Li 0001
IEEE Trans. Speech Audio Process.3
2011 Classifier subset selection and fusion for speaker verification
abstract
State-of-the-art speaker verification systems consists of a number of complementary subsystems whose outputs are fused, to arrive at more accurate and reliable verification decision. In speaker verification, fusion is typically implemented as a linear combination of the subsystem scores. Parameters of the linear model are commonly estimated using the logistic regression method, as implemented in the popular FoCal toolkit. In this paper, we study simultaneous use of classifier selection and fusion. We study four alternative fusion strategies, three score warping techniques, and provide interesting experimental bounds on optimal classifier subset selection. Detailed experiments are carried out on the NIST 2008 and 2010 SRE corpora.
Filip Sedlak, Tomi Kinnunen, Ville Hautamäki, Kong-Aik Lee, Haizhou Li 0001
ICASSP1
2010 Towards task-independent person authentication using eye movement signals
abstract
We propose a person authentication system using eye movement signals. In security scenarios, eye-tracking has earlier been used for gaze-based password entry. A few authors have also used physical features of eye movement signals for authentication in a task-dependent scenario with matched training and test samples. We propose and implement a task-independent scenario whereby the training and test samples can be arbitrary. We use short-term eye gaze direction to construct feature vectors which are modeled using Gaussian mixtures. The results suggest that there are personspecific features in the eye movements that can be modeled in a task-independent manner. The range of possible applications extends beyond the security-type of authentication to proactive and user-convenience systems.
Tomi Kinnunen, Filip Sedlak, Roman Bednarik
ETRA2