VLDB 2026 Research / reviewers in the wild / expert
Erik Visser
dblp:64/9262
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
Kyungguen Byun, Jason Filos, Erik Visser, Sunkuk Moon |
INTERSPEECH | 3 |
| 2024 | Parameter Efficient Audio Captioning with Faithful Guidance Using Audio-Text Shared Latent RepresentationabstractThere has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks. Albeit performance improvements, such models frequently suffer from hallucination and large memory footprint making them challenging to deploy on edge devices. In this paper, we address both these issues for the application of automated audio captioning. First, we propose a data augmentation technique for generating hallucinated audio captions and show that similarity based on an audio-text shared latent space is suitable for detecting hallucination. Then, we propose a parameter efficient inference time faithful decoding algorithm that enables smaller audio captioning models with performance equivalent to larger models trained with more data. During the beam decoding step, the smaller model utilizes an audio-text shared latent representation to semantically align the generated text with corresponding input audio. Faithful guidance is introduced into the beam probability by incorporating the cosine similarity between latent representation projections of greedy rolled out intermediate beams and audio clip. We show the efficacy of our algorithm on benchmark datasets and evaluate the proposed scheme against baselines using conventional audio captioning and semantic similarity metrics while illustrating tradeoffs between performance and complexity. Arvind Krishna Sridhar, Yinyi Guo, Erik Visser, Rehana Mahfuz |
ICASSP | 3 |
| 2023 | Improving Audio Captioning Using Semantic Similarity MetricsabstractAudio captioning quality metrics which are typically borrowed from the machine translation and image captioning areas measure the degree of overlap between predicted tokens and gold reference tokens. In this work, we consider a metric measuring semantic similarities between predicted and reference captions instead of measuring exact word overlap. We first evaluate its ability to capture similarities among captions corresponding to the same audio file and compare it to other established metrics. We then propose a fine-tuning method to directly optimize the metric by backpropagating through a sentence embedding extractor and audio captioning network. Such fine-tuning results in an improvement in predicted captions as measured by both traditional metrics and the proposed semantic similarity captioning metric. Rehana Mahfuz, Yinyi Guo, Erik Visser |
ICASSP | 3 |
| 2023 | Application of Knowledge Distillation to Multi-Task Speech Representation Learning
Mine Kerpicci, Erik Visser |
INTERSPEECH | 4 |
| 2022 | Multi-Task Voice Activated Framework Using Self-Supervised LearningabstractSelf-supervised learning methods such as wav2vec 2.0 have shown promising results in learning speech representations from unlabelled and untranscribed speech data that are useful for speech recognition. Since these representations are learned without any task-specific supervision, they can also be useful for other voice activated tasks like speaker verification, keyword spotting, emotion classification etc. In our work, we propose a general purpose framework for adapting a pre-trained wav2vec 2.0 model for different voice activated tasks. We develop downstream network architectures that operate on the contextualized speech representations of wav2vec 2.0 to adapt the representations for solving a given task. Finally, we extend our framework to perform multi-task learning by jointly optimizing the network parameters on multiple voice activated tasks using a shared transformer backbone. Both of our single and multi-task frameworks achieve state-of-the-art results in speaker verification and keyword spotting benchmarks. Our best performing models achieve 1.98% and 3.15% EER on VoxCeleb1 test set when trained on Vox-Celeb2 and VoxCeleb1 respectively, and 98.23% accuracy on Google Speech Commands v1.0 keyword spotting dataset. Shehzeen Hussain, Erik Visser |
ICASSP | 4 |
| 2020 | Incremental Learning Algorithm For Sound Event DetectionabstractThis paper presents a new learning strategy for the Sound Event Detection (SED) system to tackle the issues of i) knowledge migration from a pre-trained model to a new target model and ii) learning new sound events without forgetting the previously learned ones without re-training from scratch. In order to migrate the previously learned knowledge from the source model to the target one, a neural adapter is employed on the top of the source model. The source model and the target model are merged via this neural adapter layer. The neural adapter layer facilitates the target model to learn new sound events with minimal training data and maintaining the performance of the previously learned sound events similar to the source model. Our extensive analysis on the DCASE16 and US-SED dataset reveals the effectiveness of the proposed method in transferring knowledge between source and target models without introducing any performance degradation on the previously learned sound events while obtaining a competitive detection performance on the newly learned sound events. Eunjeong Koh, Fatemeh Saki, Yinyi Guo, Cheng-Yu Hung, Erik Visser |
ICME | 5 |
| 2007 | Frequency Domain Passive Broadband Speaker Localization using a Permutation-Free Blind Source Separation AlgorithmabstractTraditional passive broadband source localization techniques like maximum likelihood estimation and MUSIC have shown difficulties in situations where multiple correlating source signals are interfering with each other. Blind source separation (BSS) algorithms on the other hand have demonstrated good performance in separating correlated mixture signals into independent sources. In this paper it will be shown that the performance of traditional source localization algorithms can be improved by using a permutation-free frequency domain BSS algorithm as a front end. In addition a source localization method based solely on information gained from the separated BSS solution and sensor array architecture is presented. The methodologies are illustrated in an undercomplete acoustic scenario involving 3 speech sources and a 6 element microphone array. Erik Visser |
ICASSP (2) | 1 |
| 2006 | Geometrically constrained permutation-free source separation in an undercomplete speech unmixing scenario
Erik Visser |
INTERSPEECH | 1 |