Sergey Novoselov

dblp:148/9761 · DBLP profile ↗
← Back
24ranked-venue papers
12as first author
7since 2021 · last 2025
0009-0009-9354-8967ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 first-author · 7 since 2021Artificial intelligence and machine learning · 16 · 7 first-author · 4 since 2021
YearPublicationVenuePosition
2025 ITMO language diarization and identification systems for the DISPLACE 2024 challenge
abstract
This paper describes our language diarization and identification systems developed for far-field recorded group conversations. Our approach has a two-stage design and relies on classical methods, such as spectral clustering of language embeddings. The heuristic bypass (HBP) method was utilized to generate the similarity matrix required for spectral clustering used in the first stage. In the second stage the language identification block predicts language labels for a specific set of target languages. Users can manually determine the number of clusters for spectral clustering when using the identification block into the processing pipeline. Various language embedding extractors, including those based on ResNet34 and wav2vec 2.0 architectures, were utilized. We used these systems, as well as their fusion, into submission for Track 2 on language diarization of the DISPLACE 2024 challenge. Our system achieved 5 % relative improvements on eval set compared to the organizer-provided baseline system, securing the second place for Track 2 of the challenge.
Egor Ausev, Vladimir Volokhov, Sergey Novoselov, Vladislav Marchevskiy, Ekaterina Shangina, Alexey Logunov
ICASSP3
2025 In Search of Optimal Pretraining Strategy for Robust Speaker Recognition
abstract
While demonstrating state-of-the-art results in the microphone channel domain (VoxCeleb protocols), contemporary speaker verification systems are not often tested in challenging acoustic environments such as telephone channel or far-field microphone. This paper compares modern pretraining strategies, proven beneficial for the speaker verification task. It follows wav2vec 2.0, HuBERT, ASR procedures, and aims to identify the most effective, robust approach. We conduct a range of experiments with pretraining on the LibriSpeech corpus and finetuning on the VoxCeleb dataset. The systems are evaluated on multiple protocols with the microphone, telephone, and cross-channel tasks. Our empirical results show that ASR pretraining demonstrates superior in-domain performance but fails to match HuBERT/wav2vec 2.0 in out-of-domain NIST SRE assessment. Adoption of wav2vec 2.0 strategy achieves a 34% average improvement in out-of-domain evaluations compared to the baseline systems. We employ UMAP visualization of models’ embedding space to further understand the reasons for unstable performance in adversarial conditions. We also conclude that while a choice of a pretraining scheme is important, the impact of a speaker verification backend is negligible.
Nikita Khmelev, Stepan Malykh, Alexander Anikin, Anastasia Korenevskaya, Sergey Novoselov, Vladimir Volokhov, Anastasia Zorkina, Vladislav Marchevskiy, Galina Lavrentyeva
ICASSP5
2025 STCON NIST SRE24 System: Composite Speaker Recognition Solution for Challenging Scenarios
Stepan Malykh, Alexander Anikin, Nikita Khmelev, Anastasia Korenevskaya, Anastasia Zorkina, Sergey Novoselov, Vladislav Marchevskiy, Vladimir Volokhov, Andrey Shulipa, Alexander Kozlov, Alexander Melnikov, Vasiliy Galyuk, Timur Pekhovsky
INTERSPEECH6
2025 Cryfish: On deep audio analysis with Large Language Models
abstract
The recent revolutionary progress in text-based large language models (LLMs) has contributed to the growth of interest in extending capabilities of such models to multimodal perception and understanding tasks. Hearing is an essential capability that is highly desired to be integrated into LLMs. However, effective integrating listening capabilities into LLMs is a significant challenge lying in generalizing complex auditory tasks across speech and sounds. To address these issues, we introduce Cryfish, our version of auditory-capable LLM. The model integrates WavLM audio-encoder features into Qwen2 model using a transformer-based connector. Cryfish is adapted to various auditory tasks through a specialized training strategy. We evaluate the model on the new Dynamic SUPERB Phase-2 comprehensive multitask benchmark specifically designed for auditory-capable models. The paper presents an in-depth analysis and detailed comparison of Cryfish with the publicly available models.
Anton Mitrofanov, Sergey Novoselov, Tatiana Prisyach, Vladislav Marchevskiy, Arseniy Karelin, Nikita Khmelev, Dmitry Dutov, Stepan Malykh, Igor Agafonov, Aleksandr Nikitin, Oleg Petrov
INTERSPEECH2
2023 Universal Speaker Recognition Encoders for Different Speech Segments Duration
abstract
Creating universal speaker encoders which are robust for different acoustic and speech duration conditions is a big challenge today. According to our observations systems trained on short speech segments are optimal for short phrase speaker verification and systems trained on long segments are superior for long segments verification. A system trained simultaneously on pooled short and long speech segments does not give optimal verification results and usually degrades both for short and long segments. This paper addresses the problem of creating universal speaker encoders for different speech segments duration. We describe our simple recipe for training universal speaker encoder for any type of selected neural network architecture. According to our evaluation results of wav2vec-TDNN based systems obtained for NIST SRE and VoxCeleb1 benchmarks the proposed universal encoder provides speaker verification improvements in case of different enrollment and test speech segment duration. The key feature of the proposed encoder is that it has the same inference time as the selected neural network architecture.
Sergey Novoselov, Vladimir Volokhov, Galina Lavrentyeva
ICASSP1
2023 On the robustness of wav2vec 2.0 based speaker recognition systems
Sergey Novoselov, Galina Lavrentyeva, Anastasia Avdeeva, Vladimir Volokhov, Nikita Khmelev, Artem Akulov, Polina Leonteva
INTERSPEECH1
2021 SdSVC Challenge 2021: Tips and Tricks to Boost the Short-Duration Speaker Verification System Performance
Aleksei Gusev, Alisa Vinogradova, Sergey Novoselov, Sergei Astapov
Interspeech3
2020 STC-Innovation Speaker Recognition Systems for Far-Field Speaker Verification Challenge 2020
Aleksei Gusev, Vladimir Volokhov, Alisa Vinogradova, Andzhukaev Tseren, Andrey Shulipa, Sergey Novoselov, Timur Pekhovsky, Alexander Kozlov
INTERSPEECH6
2020 Blind Speech Signal Quality Estimation for Speaker Verification Systems
Galina Lavrentyeva, Marina Volkova, Anastasia Avdeeva, Sergey Novoselov, Artem Gorlanov, Andzhukaev Tseren, Artem Ivanov, Alexander Kozlov
INTERSPEECH4
2019 Phonespoof: A New Dataset for Spoofing Attack Detection in Telephone Channel
abstract
The results of spoofing detection systems proposed during ASVspoof Challenges 2015 and 2017 confirmed the perspective in detection of unforseen spoofing trials in microphone channel. However, telephone channel presents much more challenging conditions for spoofing detection, due to limited bandwidth, various coding standards and channel effects. Research on the topic has thus far only made use of program codecs and other telephone channel emulations. Such emulations does not quite match the real telephone spoofing attacks. In order to asses spoofing detection methods in real scenario we present the PHONESPOOF dataset - spoofing data collected through realistic telephone channels. The PHONE-SPOOF data collection represents most threatening types of spoofing attacks and is publicly available dataset1. This work2aimed to investigate robustness of the state-of-the-art deep learning based antispoofing systems under telephone spoofing attacks conditions based on the PHONESPOOF data. Moreover newly collected dataset makes it possible to analize language dependency issue for the Anti-Spoofing methods. In the work we also focused on the development of a unified LCNN-based approach for spoofing attack detection. The goal was to train a single system able to detect various types of spoofing attacks in telephone channel. The obtained results approve the effectiveness of such solution.
Galina Lavrentyeva, Sergey Novoselov, Marina Volkova, Yuri Matveev, Maria De Marsico
ICASSP2
2019 STC Antispoofing Systems for the ASVspoof2019 Challenge
abstract
This paper describes the Speech Technology Center (STC) antispoofing systems submitted to the ASVspoof 2019 challenge. The ASVspoof2019 is the extended version of the previous challenges and includes 2 evaluation conditions: logical access use-case scenario with speech synthesis and voice conversion attack types and physical access use-case scenario with replay attacks. During the challenge we developed anti-spoofing solutions for both scenarios. The proposed systems are implemented using deep learning approach and are based on different types of acoustic features. We enhanced Light CNN architecture previously considered by the authors for replay attacks detection and which performed high spoofing detection quality during the ASVspoof2017 challenge. In particular here we investigate the efficiency of angular margin based softmax activation for training robust deep Light CNN classifier to solve the mentioned-above tasks. Submitted systems achieved EER of 1.86% in logical access scenario and 0.54% in physical access scenario on the evaluation part of the Challenge corpora. High performance obtained for the unknown types of spoofing attacks demonstrates the stability of the offered approach in both evaluation conditions.
Galina Lavrentyeva, Sergey Novoselov, Andzhukaev Tseren, Marina Volkova, Artem Gorlanov, Alexander Kozlov
INTERSPEECH2
2019 Speaker Diarization with Deep Speaker Embeddings for DIHARD Challenge II
Sergey Novoselov, Aleksei Gusev, Artem Ivanov, Timur Pekhovsky, Andrey Shulipa, Anastasia Avdeeva, Artem Gorlanov, Alexander Kozlov
INTERSPEECH1
2019 STC Speaker Recognition Systems for the VOiCES from a Distance Challenge
abstract
This paper presents the Speech Technology Center (STC) speaker recognition (SR) systems submitted to the VOiCES From a Distance challenge 2019. The challenge's SR task is focused on the problem of speaker recognition in single channel distant/far-field audio under noisy conditions. In this work we investigate different deep neural networks architectures for speaker embedding extraction to solve the task. We show that deep networks with residual frame level connections outperform more shallow architectures. Simple energy based speech activity detector (SAD) and automatic speech recognition (ASR) based SAD are investigated in this work. We also address the problem of data preparation for robust embedding extractors training. The reverberation for the data augmentation was performed using automatic room impulse response generator. In our systems we used discriminatively trained cosine similarity metric learning model as embedding backend. Scores normalization procedure was applied for each individual subsystem we used. Our final submitted systems were based on the fusion of different subsystems. The results obtained on the VOiCES development and evaluation sets demonstrate effectiveness and robustness of the proposed systems when dealing with distant/far-field audio under noisy conditions.
Sergey Novoselov, Aleksei Gusev, Artem Ivanov, Timur Pekhovsky, Andrey Shulipa, Galina Lavrentyeva, Vladimir Volokhov, Alexander Kozlov
INTERSPEECH1
2019 STC Speaker Recognition Systems for the VOiCES from a Distance Challenge
Sergey Novoselov, Aleksei Gusev, Artem Ivanov, Timur Pekhovsky, Andrey Shulipa, Galina Lavrentyeva, Vladimir Volokhov, Alexander Kozlov
INTERSPEECH1
2018 Deep CNN Based Feature Extractor for Text-Prompted Speaker Recognition
abstract
Deep learning is still not a very common tool in speaker verification field. We study deep convolutional neural network performance in the text-prompted speaker verification task. The prompted passphrase is segmented into word states - i.e. digits - to test each digit utterance separately. We train a single high-level feature extractor for all states and use cosine similarity metric for scoring. The key feature of our network is the Max-Feature-Map activation function, which acts as an embedded feature selector. By using multitask learning scheme to train the high-level feature extractor we were able to surpass the classic baseline systems in terms of quality and achieved impressive results for such a novice approach, getting 2.85% EER on the RSR2015 evaluation set. Fusion of the proposed and the baseline systems improves this result.
Sergey Novoselov, Oleg Kudashev, Vadim Shchemelinin, Ivan Kremnev, Galina Lavrentyeva
ICASSP1
2018 Triplet Loss Based Cosine Similarity Metric Learning for Text-independent Speaker Recognition
Sergey Novoselov, Vadim Shchemelinin, Andrey Shulipa, Alexander Kozlov, Ivan Kremnev
INTERSPEECH1
2017 Audio Replay Attack Detection with Deep Learning Frameworks
Galina Lavrentyeva, Sergey Novoselov, Egor Malykh, Alexander Kozlov, Oleg Kudashev, Vadim Shchemelinin
INTERSPEECH2
2016 STC anti-spoofing systems for the ASVspoof 2015 challenge
abstract
This paper presents the Speech Technology Center (STC) systems submitted to Automatic Speaker Verification Spoofing and Countermeasures (ASVspoof) Challenge 2015. In this work we investigate different acoustic feature spaces to determine reliable and robust countermeasures against spoofing attacks. In addition to the commonly used front-end MFCC features we explored features derived from phase spectrum and features based on applying the multiresolution wavelet transform. Similar to state-of-the-art ASV systems, we used the standard TV approach for probability modelling in spoofing detection systems. Experiments performed on the development and evaluation datasets of the Challenge demonstrate that the use of phase-related and wavelet-based features provides a substantial input into the efficiency of the resulting STC systems. In our research we also focused on the comparison of the linear (SVM) and nonlinear (DBN) classifiers.
Sergey Novoselov, Alexander Kozlov, Galina Lavrentyeva, Konstantin Simonchik, Vadim Shchemelinin
ICASSP1
2016 A Speaker Recognition System for the SITW Challenge
Oleg Kudashev, Sergey Novoselov, Konstantin Simonchik, Alexander Kozlov
INTERSPEECH2
2016 Usage of DNN in Speaker Recognition: Advantages and Problems
Oleg Kudashev, Sergey Novoselov, Timur Pekhovsky, Konstantin Simonchik, Galina Lavrentyeva
ISNN2
2015 Plda-based system for text-prompted password speaker verification
abstract
Recently we have proposed a new State-GMM-supervector extractor for solving the problem of text-dependent speaker recognition. We demonstrated that segmenting the passphrase into word states for supervector extraction makes it possible to create more accurate statistical models of speech signals and to achieve reduction of EER compared to the best state-of-the-art systems of text-dependent verification for a text-prompted passphrase. In this paper we used a similar approach for creating a text-dependent verification system based on PLDA. The proposed system is easy to implement. Although the PLDA system is inferior to the baseline systems, fusing it with the baseline systems leads to improved quality of text-prompted password speaker verification in comparison to the fusion of baseline systems.
Sergey Novoselov, Timur Pekhovsky, Andrey Shulipa, Oleg Kudashev
AVSS1
2015 Non-linear PLDA for i-vector speaker verification
Sergey Novoselov, Timur Pekhovsky, Oleg Kudashev, Valentin Mendelev, Alexey Prudnikov
INTERSPEECH1
2014 Text-dependent GMM-JFA system for password based speaker verification
abstract
We propose a new State-GMM-supervector extractor for solving the problem of text-dependent speaker recognition. The proposed scheme for supervector extraction makes it easy to implement a text-dependent JFA system for passphrase verification. We examine the conditions of both a global and a text-prompted passphrase. The experiments conducted on the Wells Fargo Bank speech database show that the proposed method makes it possible to create more accurate statistical models of speech signals and to achieve a 44% relative reduction of EER compared to the best state-of-the-art systems of text-dependent verification for a text-prompted passphrase.
Sergey Novoselov, Timur Pekhovsky, Andrey Shulipa, Alexey Sholokhov
ICASSP1
2014 RBM-PLDA subsystem for the NIST i-vector challenge
abstract
This paper presents the Speech Technology Center (STC) system submitted to NIST i-vector challenge. The system includes different subsystems based on TV-PLDA, TV-SVM, and RBM-PLDA. In this paper we focus on examining the third RBM-PLDA subsystem. Within this subsystem, we present our RBM extractor of the pseudo i-vector. Experiments performed on the test dataset of NIST-2014 demonstrate that although the RBM-PLDA subsystem is inferior to the former two subsystems in terms of absolute minDCF, during the final fusion it provides a substantial input into the efficiency of the resulting STC system reaching 0.241 at the minDCF point.
Sergey Novoselov, Timur Pekhovsky, Konstantin Simonchik, Andrey Shulipa
INTERSPEECH1