VLDB 2026 Research / reviewers in the wild / expert
Vladimir Volokhov
dblp:181/4366
· DBLP profile ↗
8ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ITMO language diarization and identification systems for the DISPLACE 2024 challengeabstractThis paper describes our language diarization and identification systems developed for far-field recorded group conversations. Our approach has a two-stage design and relies on classical methods, such as spectral clustering of language embeddings. The heuristic bypass (HBP) method was utilized to generate the similarity matrix required for spectral clustering used in the first stage. In the second stage the language identification block predicts language labels for a specific set of target languages. Users can manually determine the number of clusters for spectral clustering when using the identification block into the processing pipeline. Various language embedding extractors, including those based on ResNet34 and wav2vec 2.0 architectures, were utilized. We used these systems, as well as their fusion, into submission for Track 2 on language diarization of the DISPLACE 2024 challenge. Our system achieved 5 % relative improvements on eval set compared to the organizer-provided baseline system, securing the second place for Track 2 of the challenge. Egor Ausev, Vladimir Volokhov, Sergey Novoselov, Vladislav Marchevskiy, Ekaterina Shangina, Alexey Logunov |
ICASSP | 2 |
| 2025 | In Search of Optimal Pretraining Strategy for Robust Speaker RecognitionabstractWhile demonstrating state-of-the-art results in the microphone channel domain (VoxCeleb protocols), contemporary speaker verification systems are not often tested in challenging acoustic environments such as telephone channel or far-field microphone. This paper compares modern pretraining strategies, proven beneficial for the speaker verification task. It follows wav2vec 2.0, HuBERT, ASR procedures, and aims to identify the most effective, robust approach. We conduct a range of experiments with pretraining on the LibriSpeech corpus and finetuning on the VoxCeleb dataset. The systems are evaluated on multiple protocols with the microphone, telephone, and cross-channel tasks. Our empirical results show that ASR pretraining demonstrates superior in-domain performance but fails to match HuBERT/wav2vec 2.0 in out-of-domain NIST SRE assessment. Adoption of wav2vec 2.0 strategy achieves a 34% average improvement in out-of-domain evaluations compared to the baseline systems. We employ UMAP visualization of models’ embedding space to further understand the reasons for unstable performance in adversarial conditions. We also conclude that while a choice of a pretraining scheme is important, the impact of a speaker verification backend is negligible. Nikita Khmelev, Stepan Malykh, Alexander Anikin, Anastasia Korenevskaya, Sergey Novoselov, Vladimir Volokhov, Anastasia Zorkina, Vladislav Marchevskiy, Galina Lavrentyeva |
ICASSP | 6 |
| 2025 | STCON NIST SRE24 System: Composite Speaker Recognition Solution for Challenging Scenarios
Stepan Malykh, Alexander Anikin, Nikita Khmelev, Anastasia Korenevskaya, Anastasia Zorkina, Sergey Novoselov, Vladislav Marchevskiy, Vladimir Volokhov, Andrey Shulipa, Alexander Kozlov, Alexander Melnikov, Vasiliy Galyuk, Timur Pekhovsky |
INTERSPEECH | 8 |
| 2023 | Universal Speaker Recognition Encoders for Different Speech Segments DurationabstractCreating universal speaker encoders which are robust for different acoustic and speech duration conditions is a big challenge today. According to our observations systems trained on short speech segments are optimal for short phrase speaker verification and systems trained on long segments are superior for long segments verification. A system trained simultaneously on pooled short and long speech segments does not give optimal verification results and usually degrades both for short and long segments. This paper addresses the problem of creating universal speaker encoders for different speech segments duration. We describe our simple recipe for training universal speaker encoder for any type of selected neural network architecture. According to our evaluation results of wav2vec-TDNN based systems obtained for NIST SRE and VoxCeleb1 benchmarks the proposed universal encoder provides speaker verification improvements in case of different enrollment and test speech segment duration. The key feature of the proposed encoder is that it has the same inference time as the selected neural network architecture. Sergey Novoselov, Vladimir Volokhov, Galina Lavrentyeva |
ICASSP | 2 |
| 2023 | On the robustness of wav2vec 2.0 based speaker recognition systems
Sergey Novoselov, Galina Lavrentyeva, Anastasia Avdeeva, Vladimir Volokhov, Nikita Khmelev, Artem Akulov, Polina Leonteva |
INTERSPEECH | 4 |
| 2020 | STC-Innovation Speaker Recognition Systems for Far-Field Speaker Verification Challenge 2020
Aleksei Gusev, Vladimir Volokhov, Alisa Vinogradova, Andzhukaev Tseren, Andrey Shulipa, Sergey Novoselov, Timur Pekhovsky, Alexander Kozlov |
INTERSPEECH | 2 |
| 2019 | STC Speaker Recognition Systems for the VOiCES from a Distance ChallengeabstractThis paper presents the Speech Technology Center (STC) speaker recognition (SR) systems submitted to the VOiCES From a Distance challenge 2019. The challenge's SR task is focused on the problem of speaker recognition in single channel distant/far-field audio under noisy conditions. In this work we investigate different deep neural networks architectures for speaker embedding extraction to solve the task. We show that deep networks with residual frame level connections outperform more shallow architectures. Simple energy based speech activity detector (SAD) and automatic speech recognition (ASR) based SAD are investigated in this work. We also address the problem of data preparation for robust embedding extractors training. The reverberation for the data augmentation was performed using automatic room impulse response generator. In our systems we used discriminatively trained cosine similarity metric learning model as embedding backend. Scores normalization procedure was applied for each individual subsystem we used. Our final submitted systems were based on the fusion of different subsystems. The results obtained on the VOiCES development and evaluation sets demonstrate effectiveness and robustness of the proposed systems when dealing with distant/far-field audio under noisy conditions. Sergey Novoselov, Aleksei Gusev, Artem Ivanov, Timur Pekhovsky, Andrey Shulipa, Galina Lavrentyeva, Vladimir Volokhov, Alexander Kozlov |
INTERSPEECH | 7 |
| 2019 | STC Speaker Recognition Systems for the VOiCES from a Distance Challenge
Sergey Novoselov, Aleksei Gusev, Artem Ivanov, Timur Pekhovsky, Andrey Shulipa, Galina Lavrentyeva, Vladimir Volokhov, Alexander Kozlov |
INTERSPEECH | 7 |