Emil Solsbæk Ottosen

dblp:192/1725 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
0since 2021 · last 2017
0000-0001-5404-4541ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › spectral processing
phase vocoder
0.312017
A Phase Vocoder Based on Nonstationary Gabor Frames · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing
time-frequency analysis
0.312017
A Phase Vocoder Based on Nonstationary Gabor Frames · IEEE ACM Trans. Audio Speech Lang. Process. 2017

Methods — techniques the papers use, named apart from their topics

phase locking · 0.3adaptive time-frequency representation · 0.3
YearPublicationVenuePosition
2017 A Phase Vocoder Based on Nonstationary Gabor Frames
abstract
We propose a new algorithm for time stretching music signals based on the theory of nonstationary Gabor frames (NSGFs). The algorithm extends the techniques of the classical phase vocoder (PV) by incorporating adaptive time-frequency (TF) representations and adaptive phase locking. The adaptive TF representations imply good time resolution for the onsets of attack transients and good frequency resolution for the sinusoidal components. We estimate the phase values only at peak channels and the remaining phases are then locked to the values of the peaks in an adaptive manner. During attack transients we keep the stretch factor equal to one and we propose a new strategy for determining which channels are relevant for reinitializing the corresponding phase values. In contrast to previously published algorithms we use a non-uniform NSGF to obtain a low redundancy of the corresponding TF representation. We show that with just three times as many TF coefficients as signal samples, artifacts such as phasiness and transient smearing can be greatly reduced compared to the classical PV. The proposed algorithm is tested on both synthetic and real-world signals and compared with state-of-the-art algorithms in a reproducible manner.
Emil Solsbæk Ottosen, Monika Dörfler
IEEE ACM Trans. Audio Speech Lang. Process.1