Jakob Abeßer

dblp:64/8761 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
1since 2021 · last 2023
0000-0003-4689-7944ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-authorArtificial intelligence and machine learning · 3 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Audio and music processing · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › music analysis
expressive performance analysis
0.312017
Score-Informed Analysis of Tuning, Intonation, Pitch Modulation, and Dynamics in Jazz Solos · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing
music information retrieval
0.312017
Score-Informed Analysis of Tuning, Intonation, Pitch Modulation, and Dynamics in Jazz Solos · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing
music transcription
0.312017
Instrument-Centered Music Transcription of Solo Bass Guitar Recordings · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing
audio representation learning
0.212023
How Robust are Audio Embeddings for Polyphonic Sound Event Tagging? · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Audio and music processing › source separation
music source separation
0.112017
Score-Informed Analysis of Tuning, Intonation, Pitch Modulation, and Dynamics in Jazz Solos · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing › music analysis
playing technique classification
0.112017
Instrument-Centered Music Transcription of Solo Bass Guitar Recordings · IEEE ACM Trans. Audio Speech Lang. Process. 2017

Methods — techniques the papers use, named apart from their topics

sensitivity measures · 0.7embedding space analysis · 0.7support vector machine · 0.3score-informed analysis · 0.3onset detection · 0.3fundamental frequency tracking · 0.3fundamental frequency estimation · 0.3
YearPublicationVenuePosition
2023 How Robust are Audio Embeddings for Polyphonic Sound Event Tagging?
abstract
Sound classification algorithms are challenged by the natural variability of everyday sounds, particularly for large sound class taxonomies. In order to be applicable in real-life environments, such algorithms must also be able to handle polyphonic scenarios, where simultaneously occurring and overlapping sound events need to be classified. With the rapid progress of deep learning, several deep audio embeddings (DAEs) have been proposed as pre-trained feature representations for sound classification. In this article, we analyze the embedding spaces of two non-trainable audio representations (NTARs) and five DAEs for sound classification in polyphonic scenarios (sound event tagging) and make several contributions. First, we compare general properties like the inter-correlation between feature dimensions and the scattering of sound classes in the embedding spaces. Second, we test the robustness of the embeddings against several audio degradations and propose two sensitivity measures based on a class-agnostic and a class-centric view on the resulting drift in the embedding space. Finally, as a central contribution, we study how a blending between pairs of sounds maps to embedding space trajectories and how the path of these trajectories can cause classification errors due to their proximity to other sound classes. Throughout our analyses, the PANN embeddings have shown the best overall performance for low-polyphony sound event tagging.
Jakob Abeßer, Sascha Grollmisch, Meinard Müller
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Fundamental Frequency Contour Classification: A Comparison between Hand-crafted and CNN-based Features
abstract
In this paper, we evaluate hand-crafted features as well as features learned from data using a convolutional neural network (CNN) for different fundamental frequency classification tasks. We compare classification based on full (variable-length) contours and classification based on fixed-sized subcontours in combination with a fusion strategy. Our results show that hand-crafted and learned features lead to comparable results for both classification scenarios. Aggregating contour-level to file-level classification results generally improves the results. In comparison to the hand-crafted features, our examination indicates that the CNN-based features show a higher degree of redundancy across feature dimensions, where multiple filters (convolution kernels) specialize on similar contour shapes.
Jakob Abeßer, Meinard Müller
ICASSP1
2017 Data-driven solo voice enhancement for jazz music retrieval
abstract
Retrieving short monophonic queries in music recordings is a challenging research problem in Music Information Retrieval (MIR). In jazz music, given a solo transcription, one retrieval task is to find the corresponding (potentially polyphonic) recording in a music collection. Many conventional systems approach such retrieval tasks by first extracting the predominant F0-trajectory from the recording, then quantizing the extracted trajectory to musical pitches and finally comparing the resulting pitch sequence to the monophonic query. In this paper, we introduce a data-driven approach that avoids the hard decisions involved in conventional approaches: Given pairs of time-frequency (TF) representations of full music recordings and TF representations of solo transcriptions, we use a DNN-based approach to learn a mapping for transforming a “polyphonic” TF representation into a “monophonic” TF representation. This transform can be considered as a kind of solo voice enhancement. We evaluate our approach within a jazz solo retrieval scenario and compare it to a state-of-the-art method for predominant melody extraction.
Stefan Balke, Christian Dittmar, Jakob Abeßer, Meinard Müller
ICASSP3
2017 Score-Informed Analysis of Tuning, Intonation, Pitch Modulation, and Dynamics in Jazz Solos
abstract
Both the collection and analysis of large music repertoires constitute major challenges within musicological disciplines such as jazz research. Automatic methods of music analysis based on audio signal processing have the potential to assist researchers and to accelerate the transcription and analysis of music recordings significantly. In this paper, we propose a framework for analyzing improvised monophonic solos in multi-instrumental jazz recordings with special focus on reed and brass instruments. The analysis algorithms rely on prior score-information, which is taken from high quality manual solo transcriptions. Following an initial solo and accompaniment source separation, we propose algorithms for tone-wise extraction of fundamental frequency and intensity contours. Based on this fine-grained representation of recorded jazz solos, we perform several exploratory experiments motivated by questions relating to jazz research in order to analyze the use of expressive stylistic devices such as intonation, pitch modulation, and dynamics in jazz solos. The results show that a score-informed audio analysis of jazz recordings can provide valuable insights into the individual stylistic characteristics of jazz musicians.
Jakob Abeßer, Klaus Frieler, Estefanía Cano, Martin Pfleiderer, Wolf-Georg Zaddach
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Instrument-Centered Music Transcription of Solo Bass Guitar Recordings
abstract
This paper deals with the automatic transcription of solo bass guitar recordings with an additional estimation of playing techniques and fretboard positions used by the musician. Our goal is to first develop a system for a robust estimation of the note parameters pitch, onset, and duration (score-level parameters). As a second step, we aim to automatically detect the applied plucking and expression style as well as the fret and string positions for each note (instrument-level parameters). Our approach is to first apply a note onset detection followed by a tracking of the fundamental frequency contours based on a reassigned magnitude spectrogram. Then, we model the spectral envelope of each note and derive various timbre-related audio features. Using a support vector machine classifier, we automatically classify the instrument-level parameters for each detected note event. Our results show that the proposed system achieves accuracy values above 0.88 for the estimation of the plucking style, expression style, and string number for isolated note samples. As an additional contribution, we analyze the influence of the note duration characteristics in the classification performance. In a score-level evaluation on a novel public dataset of solo bass guitar tracks, our method outperforms three existing transcription algorithms for bass transcription in polyphonic music as well as a melody transcription algorithm for monophonic music.
Jakob Abeßer, Gerald Schuller
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Parameter extraction for bass guitar sound models including playing styles
abstract
We present a system to realistically model the sound of bass guitars, and how to estimate the corresponding parameters from the sound of a bass guitar alone, without other physical measurements. Our model includes plucking and expression styles of the musician, like vibrato or bending, and the string number for a realistic modeling and reproduction of the sound. We show that we can estimate the playing techniques and the string number with relatively high accuracy.
Gerald Schuller, Jakob Abeßer, Christian Kehling
ICASSP2
2014 Score-Informed Tracking and Contextual Analysis of Fundamental Frequency Contours in Trumpet and Saxophone Jazz Solos
Jakob Abeßer, Martin Pfleiderer, Klaus Frieler, Wolf-Georg Zaddach
DAFx1
2014 Automatic Tablature Transcription of Electric Guitar Recordings by Estimation of Score- and Instrument-Related Parameters
Christian Kehling, Jakob Abeßer, Christian Dittmar, Gerald Schuller
DAFx2
2012 A digitalwaveguide model of the electric bass guitar including different playing techniques
abstract
In this paper, we present a novel audio synthesis model that allows us to simulate bass guitar tones with 11 different playing techniques to choose from. In contrast, previous approaches focussing on bass guitar synthesis only implemented the two slap techniques. We apply a digital waveguide model extended by different modular parts to imitate the sound production process on this instrument. The results of aMUSHRA listening test reveal that an audio coding scheme based on the presented algorithm offers a high perceived sound quality in comparison to conventional low bit-rate coding schemes while requiring a much lower bit-rate.
Patrick Kramer, Jakob Abeßer, Christian Dittmar, Gerald Schuller
ICASSP2
2011 Modeling musical attributes to characterize ensemble recordings using rhythmic audio features
abstract
In this paper, we present the results of a pre-study on music performance analysis of ensemble music. Our aim is to implement a music classification system for the description of live recordings, for instance to help musicologist and musicians to analyze improvised ensemble performances. The main problem we deal with is the extraction of a suitable set of audio features from the recorded instrument tracks. Our approach is to extract rhythm-related audio features and to apply them for regression-based modeling of eight more general musical attributes. The model based on Partial Least-Squares Regression without preceding Principal Component Analysis performed best for all of the eight attributes.
Jakob Abeßer, Olivier Lartillot, Christian Dittmar, Tuomas Eerola, Gerald Schuller
ICASSP1
2010 Feature-based extraction of plucking and expression styles of the electric bass guitar
abstract
In this paper,we present a feature-based approach for the classification of different playing techniques in bass guitar recordings. The applied audio features are chosen to capture typical instrument sounds induced by 10 different playing techniques. A novel database that consists of approx. 4300 isolated bass notes was assembled for the purpose of evaluation. The usage of domain-specific features in a combination of feature selection and feature space transformation techniques improved the classification accuracy by over 27% points in comparison to a state-of-the-art baseline system. Classification accuracy reached 93.25% and 95.61% for the recognition of plucking and expression styles respectively.
Jakob Abeßer, Hanna M. Lukashevich, Gerald Schuller
ICASSP1