EDBT 2026 Demo / reviewers in the wild / expert
Meinard Müller
dblp:01/1071
· DBLP profile ↗
81ranked-venue papers
8as first author
18since 2021 · last 2025
0000-0001-6062-7524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 6 since 2021Theory of computation · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dense-Sparse Dynamic Time Warping for Customizing Piano Concerto AccompanimentsabstractIn this study, we explore how pianists can customize Music Minus One (MMO) concerto accompaniments to match their playing style. Bypassing the need for a symbolic score, often not available digitally, we use three types of audio data: solo piano recordings, MMO orchestra-only recordings, and mixed recordings of both piano and orchestra (e.g., from YouTube). The mixed recording serves as an intermediary reference to align the solo and orchestra parts, with only the orchestral part being adjusted through time-scale modification to synchronize with the user’s playing. The main challenge with estimating these alignments is the spectral mismatch between recordings containing different musical parts. Motivated by this application scenario, we introduce Dense-Sparse DTW, a variant of Dynamic Time Warping (DTW) that is designed to improve robustness of alignments to spectral mismatch by focusing on aligning a selected subset of audio frames containing prominent timing cues. We collect and annotate data from four piano concerto movements and establish a framework for generating and evaluating customized accompaniment recordings. On this benchmark, we show that Dense-Sparse DTW has better or comparable performance than more complex approaches based on source separation and spectral subtraction techniques. T. J. Tsai 0001, Kavi Dey, Yigitcan Özer, Meinard Müller |
ICASSP | 4 |
| 2024 | Performance Conditioning for Diffusion-Based Multi-Instrument Music SynthesisabstractGenerating multi-instrument music from symbolic music representations is an important task in Music Information Retrieval (MIR). A central but still largely unsolved problem in this context is musically and acoustically informed control in the generation process. As the main contribution of this work, we propose enhancing control of multi-instrument synthesis by conditioning a generative model on a specific performance and recording environment, thus allowing for better guidance of timbre and style. Building on state-of-the-art diffusion-based music generative models, we introduce performance conditioning - a simple tool indicating the generative model to synthesize music with style and timbre of specific instruments taken from specific performances. Our prototype is evaluated using uncurated performances with diverse instrumentation and achieves state-of-the-art FAD realism scores while allowing novel timbre and style control. Our project page, including samples and demonstrations, is available at benadar293.github.io/midipm. Ben Maman, Johannes Zeitler, Meinard Müller, Amit Bermano |
ICASSP | 3 |
| 2024 | Soft Dynamic Time Warping with Variable Step WeightsabstractIn computer vision and audio processing, soft dynamic time warping (SDTW) techniques have been used as a differentiable loss function to train deep neural networks (DNNs) on weakly aligned data. In existing SDTW algorithms, the horizontal, vertical, and diagonal alignment steps all have the same weight, i.e., they contribute equally to the alignment cost. This equal weighting scheme for all step sizes can lead to degenerated alignments by, e.g., aligning most predictions to a single target frame in the early stages of training. Problems with equal step weights are known from classical DTW and have been addressed by assigning different weights to different step sizes. In this paper, we extend SDTW to allow for variable step weights and provide efficient dynamic programming algorithms for the forward and backward passes. As an example, we demonstrate the potential of the method on the task of training a DNN for pitch class estimation from music recordings, using step weight parameters that reduce the influence of outliers in repetitions of the same target frame. Johannes Zeitler, Michael Krause 0002, Meinard Müller |
ICASSP | 3 |
| 2024 | Source Separation of Piano Concertos Using Musically Motivated Augmentation TechniquesabstractIn this work, we address the novel and rarely considered source separation task of decomposing piano concerto recordings into separate piano and orchestral tracks. Being a genre written for a pianist typically accompanied by an ensemble or orchestra, piano concertos often involve an intricate interplay of the piano and the entire orchestra, leading to high spectro–temporal correlations between the constituent instruments. Moreover, in the case of piano concertos, the lack of multi-track data for training constitutes another challenge in view of data-driven source separation approaches. As a basis for our work, we adapt existing deep learning (DL) techniques, mainly used for the separation of popular music recordings. In particular, we investigate spectrogram- and waveform-based approaches as well as hybrid models operating in both spectrogram and waveform domains. As a main contribution, we introduce a musically motivated data augmentation approach for training based on artificially generated samples. Furthermore, we systematically investigate the effects of various augmentation techniques for DL-based models. For our experiments, we use a recently published, open-source dataset of multi-track piano concerto recordings. Our main findings demonstrate that the best source separation performance is achieved by a hybrid model when combining all augmentation techniques. Yigitcan Özer, Meinard Müller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Soft Dynamic Time Warping for Multi-Pitch Estimation and BeyondabstractMany tasks in music information retrieval (MIR) involve weakly aligned data, where exact temporal correspondences are unknown. The connectionist temporal classification (CTC) loss is a standard technique to learn feature representations based on weakly aligned training data. However, CTC is limited to discrete-valued target sequences and can be difficult to extend to multi-label problems. In this article, we show how soft dynamic time warping (SoftDTW), a differentiable variant of classical DTW, can be used as an alternative to CTC. Using multi-pitch estimation as an example scenario, we show that SoftDTW yields results on par with a state-of-the-art multi-label extension of CTC. In addition to being more elegant in terms of its algorithmic formulation, SoftDTW naturally extends to real-valued target sequences. Michael Krause 0002, Christof Weiß, Meinard Müller |
ICASSP | 3 |
| 2023 | TAPE: An End-to-End Timbre-Aware Pitch EstimatorabstractPitch estimation of a target musical source within a multi-source polyphonic signal is of great interest for music performance analysis. One possible approach for extracting the pitch of a target source is to first perform source separation and then estimate the pitch of the separated track. However, as we will show, this typically leads to poor results. As an alternative to this approach, we introduce a timbre-aware pitch estimator (TAPE), which estimates the pitch of a target source in an end-to-end manner without the need for an explicit source separation step. Opposed to existing approaches that assume the predominance of a lead voice, our approach builds upon other cues that only rely on the timbral characteristics. Our results on real violin–piano duets show that, without any pre-processing step, TAPE trained on synthetic mixes outperforms the sequential procedure of source separation and pitch estimation under many settings, even if the target source is not predominant. Nazif Can Tamer, Yigitcan Özer, Meinard Müller, Xavier Serra |
ICASSP | 3 |
| 2023 | Evaluating Speech-Phoneme Alignment and its Impact on Neural Text-To-Speech SynthesisabstractIn recent years, the quality of text-to-speech (TTS) synthesis vastly improved due to deep-learning techniques, with parallel architectures, in particular, providing excellent synthesis quality at fast inference. Training these models usually requires speech recordings, corresponding phoneme-level transcripts, and the temporal alignment of each phoneme to the utterances. Since manually creating such fine-grained alignments requires expert knowledge and is time-consuming, it is common practice to estimate them using automatic speech–phoneme alignment methods. In the literature, either the estimation methods’ accuracy or their impact on the TTS system’s synthesis quality is evaluated. In this study, we perform experiments with five state-of-the-art speech–phoneme aligners and evaluate their output with objective and subjective measures. As our main result, we show that small alignment errors (below 75 ms error) do not decrease the synthesis quality, which implies that the alignment error may not be the crucial factor when choosing an aligner for TTS training. Frank Zalkow, Prachi Govalkar, Meinard Müller, Emanuël A. P. Habets, Christian Dittmar |
ICASSP | 3 |
| 2023 | Multi-Scale Spectral Loss RevisitedabstractThe Multi-Scale Spectral (MSS) loss is commonly used for comparing audio signals, as it provides a good trade-off between temporal and spectral resolution. However, some configuration choices, including window type and size, magnitude compression, as well as the distance between spectrograms, are often set implicitly, even though they can significantly impact the loss properties and the convergence of trained models. Particularly in the context of differentiable digital signal processing (DDSP), where learned parameters may explicitly control the frequency of synthesis components, the MSS loss often fails to provide informative gradients. The main goal of this paper is to gain a better understanding of how different configurations of the MSS loss affect this problem. As an illustrative example, we analyze the task of sinusoid frequency estimation via gradient descent to compare different configurations and their effect on the loss properties. Furthermore, we show that favorable configurations can also facilitate unsupervised training of a more complex DDSP additive synthesis autoencoder. Our results indicate that a careful configuration may benefit many applications where the MSS loss is utilized. Simon J. Schwär, Meinard Müller |
IEEE Signal Process. Lett. | 2 |
| 2023 | How Robust are Audio Embeddings for Polyphonic Sound Event Tagging?abstractSound classification algorithms are challenged by the natural variability of everyday sounds, particularly for large sound class taxonomies. In order to be applicable in real-life environments, such algorithms must also be able to handle polyphonic scenarios, where simultaneously occurring and overlapping sound events need to be classified. With the rapid progress of deep learning, several deep audio embeddings (DAEs) have been proposed as pre-trained feature representations for sound classification. In this article, we analyze the embedding spaces of two non-trainable audio representations (NTARs) and five DAEs for sound classification in polyphonic scenarios (sound event tagging) and make several contributions. First, we compare general properties like the inter-correlation between feature dimensions and the scattering of sound classes in the embedding spaces. Second, we test the robustness of the embeddings against several audio degradations and propose two sensitivity measures based on a class-agnostic and a class-centric view on the resulting drift in the embedding space. Finally, as a central contribution, we study how a blending between pairs of sounds maps to embedding space trajectories and how the path of these trajectories can cause classification errors due to their proximity to other sound classes. Throughout our analyses, the PANN embeddings have shown the best overall performance for low-polyphony sound event tagging. Jakob Abeßer, Sascha Grollmisch, Meinard Müller |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Local Periodicity-Based Beat Tracking for Expressive Classical Piano MusicabstractTo model the periodicity of beats, state-of-the-art beat tracking systems use “post-processing trackers” (PPTs) that rely on several empirically determined global assumptions for tempo transition, which work well for music with a steady tempo. For expressive classical music, however, these assumptions can be too rigid. With two large datasets of Western classical piano music, namely the Aligned Scores and Performances (ASAP) dataset and a dataset of Chopin's Mazurkas (Maz-5), we report on experiments showing the failure of existing PPTs to cope with local tempo changes, thus calling for new methods. In this paper, we propose a new local periodicity-based PPT, called predominant local pulse-based dynamic programming (PLPDP) tracking, that allows for more flexible tempo transitions. Specifically, the new PPT incorporates a method called “predominant local pulses” (PLP) in combination with a dynamic programming (DP) component to jointly consider the locally detected periodicity and beat activation strength at each time instant. Accordingly, PLPDP accounts for the local periodicity, rather than relying on a global tempo assumption. Compared to existing PPTs, PLPDP particularly enhances the recall values at the cost of a lower precision, resulting in an overall improvement of F1-score for beat tracking in ASAP (from 0.473 to 0.493) and Maz-5 (from 0.595 to 0.838). Ching-Yu Chiu, Meinard Müller, Matthew E. P. Davies, Alvin Wen-Yu Su, Yi-Hsuan Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Hierarchical Classification for Instrument Activity Detection in Orchestral Music RecordingsabstractInstrument activity detection is a fundamental task in music information retrieval, serving as a basis for many applications, such as music recommendation, music tagging, or remixing. Most published works on this task cover popular music and music for smaller ensembles. In this paper, we embrace orchestral and opera music recordings as a rarely considered scenario for automated instrument activity detection. Orchestral music is particularly challenging since it consists of intricate polyphonic and polytimbral sound mixtures where multiple instruments are playing simultaneously. Orchestral instruments can naturally be arranged in hierarchical taxonomies, according to instrument families. As the main contribution of this paper, we show that a hierarchical classification approach can be used to detect instrument activity in our scenario, even if only few fine-grained, instrument-level annotations are available. We further consider additional loss terms for improving the hierarchical consistency of predictions. For our experiments, we collect a dataset containing 14 hours of orchestral music recordings with aligned instrument activity annotations. Finally, we perform an analysis of the behavior of our proposed approach with regard to potential confounding errors. Michael Krause 0002, Meinard Müller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Theme Transformer: Symbolic Music Generation With Theme-Conditioned TransformerabstractAttention-based Transformer models have been increasingly employed for automatic music generation. To condition the generation process of such a model with a user-specified sequence, a popular approach is to take that conditioning sequence as a priming sequence and ask a Transformer decoder to generate a continuation. However, thisprompt-based conditioningcannot guarantee that the conditioning sequence would develop or even simply repeat itself in the generated continuation. In this paper, we propose an alternative conditioning approach, calledtheme-based conditioning, that explicitly trains the Transformer to treat the conditioning sequence as a thematic material that has to manifest itself multiple times in its generation result. This is achieved with two main technical contributions. First, we propose a deep learning-based approach that uses contrastive representation learning and clustering to automatically retrieve thematic materials from music pieces in the training data. Second, we propose a novel gated parallel attention module to be used in a sequence-to-sequence (seq2seq) encoder/decoder architecture to more effectively account for a given conditioning thematic material in the generation process of the Transformer decoder. We report on objective and subjective evaluations of variants of the proposed Theme Transformer and the conventional prompt-based baseline, showing that our best model can generate, to some extent, polyphonic pop piano music with repetition and plausible variations of a given condition. Yi-Jen Shih, Shih-Lun Wu, Frank Zalkow, Meinard Müller, Yi-Hsuan Yang |
IEEE Trans. Multim. | 4 |
| 2022 | Hierarchical Classification of Singing Activity, Gender, and Type in Complex Music RecordingsabstractTraditionally, work on singing voice detection has focused on identifying singing activity in music recordings. In this work, our aim is to extend this task towards simultaneously detecting the presence of singing voice as well as determining singer gender and voice type. We describe and compare four strategies for exploiting the hierarchical relationships between these levels. In particular, we introduce a novel loss term that promotes consistency across hierarchy levels. We evaluate the strategies on a dataset containing over 200 hours of complex opera recordings with various singers of different genders and voice types, with a particular focus on hierarchical consistency. Our experiments show that by adding our loss term, a joint classification strategy using a single neural network achieves slightly improved evaluation scores and significantly more consistent results. Michael Krause 0002, Meinard Müller |
ICASSP | 2 |
| 2022 | An Analysis Method for Metric-Level Switching in Beat TrackingabstractFor expressive music, the tempo may change over time, posing challenges to tracking the beats by an automatic model. The model may first tap to the correct tempo, but then may fail to adapt to a tempo change, or switch between several incorrect but perceptually plausible ones (e.g., half- or double-tempo). Existing evaluation metrics for beat tracking do not reflect such behaviors, as they typically assume a fixed relationship between the reference beats and estimated beats. In this letter, we propose a new performance analysis method, called annotation coverage ratio (ACR), that accounts for a variety of possible metric-level switching behaviors of beat trackers. The idea is to derive sequences of modified reference beats of all metrical levels for every two consecutive reference beats, and compare every sequence of modified reference beats to the subsequences of estimated beats. We show via experiments on three datasets of different genres the usefulness of ACR when being utilized alongside existing metrics, and discuss the new insights that can be gained. Ching-Yu Chiu, Meinard Müller, Matthew E. P. Davies, Alvin Wen-Yu Su, Yi-Hsuan Yang |
IEEE Signal Process. Lett. | 2 |
| 2021 | Adaptive Pitch-Shifting with Applications to Intonation Adjustment in a Cappella RecordingsabstractA central challenge for a cappella singers is to adjust their intonation and to stay in tune relative to their fellow singers. During editing of a cappella recordings, one may want to adjust local intonation of individual singers or account for global intonation drifts over time. This requires applying a time-varying pitch-shift to the audio recording, which we refer to as adaptive pitch-shifting. In this context, existing (semi-)automatic approaches are either labor-intensive or face technical and musical limitations. In this work, we present automatic methods and tools for adaptive pitch-shifting with applications to intonation adjustment in a cappella recordings. To this end, we show how to incorporate time-varying information into existing pitch-shifting algorithms that are based on resampling and time-scale modification (TSM). Furthermore, we release an open-source Python toolbox, which includes a variety of TSM algorithms and an implementation of our method. Finally, we show the potential of our tools by two case studies on global and local intonation adjustment in a cappella recordings using a publicly available multitrack dataset of amateur choral singing. Sebastian Rosenzweig, Simon J. Schwär, Jonathan Driedger, Meinard Müller |
DAFx | 4 |
| 2021 | An evolutionary multi-objective feature selection approach for detecting music segment boundaries of specific typesabstractThe goal of music segmentation is to identify boundaries between parts of music pieces which are perceived as entities. Segment boundaries often go along with a change in musical properties including instrumentation, key, and tempo (or a combination thereof). One can consider different types (or classes) of boundaries according to these musical properties. In contrast to existing datasets with missing specifications of changing properties for annotated boundaries, we have created a set of artificial music tracks with precise annotations for boundaries of different types. This allows for a profound analysis and interpretation of annotated and predicted boundaries and a more exhaustive comparison of different segmentation algorithms. For this scenario, we formulate a novel multi-objective optimisation task that identifies boundaries of only a specific type. The optimisation is conducted by means of evolutionary multi-objective feature selection and a novelty-based segmentation approach. Furthermore, we provide lists of audio features from non-dominated fronts which most significantly contribute to the estimation of given boundaries (the first objective) and most significantly reduce the performance of the prediction of other boundaries (the second objective). Igor Vatolkin, Fabian Ostermann, Meinard Müller |
GECCO | 3 |
| 2021 | Reliability Assessment of Singing Voice F0-Estimates Using Multiple AlgorithmsabstractOver the last decades, various conceptually different approaches for fundamental frequency (F0) estimation in monophonic audio recordings have been developed. The algorithms’ performances vary depending on the acoustical and musical properties of the input audio signal. A common strategy to assess the reliability (correctness) of an estimated F0-trajectory is to evaluate against an annotated reference. However, such annotations may not be available for a particular audio collection and are typically labor-intensive to generate. In this work, we consider an approach to automatically assess the reliability of F0-trajectories estimated from monophonic singing voice recordings. As main contribution, we propose three reliability indicators that are based on the outputs of multiple algorithms. Besides providing a mathematical description of the indicators, we analyze the indicators’ behavior using a set of annotated vocal F0-trajectories. Furthermore, we show the potential of the proposed indicators for exploring unlabeled audio collections. Sebastian Rosenzweig, Frank Scherbaum, Meinard Müller |
ICASSP | 3 |
| 2021 | CTC-Based Learning of Chroma Features for Score-Audio Music RetrievalabstractThis paper deals with a scoreaudio music retrieval task where the aim is to find relevant audio recordings of Western classical music, given a short monophonic musical theme in symbolic notation as a query. Strategies for comparing score and audio data are often based on a common mid-level representation, such as chroma features, which capture melodic and harmonic properties. Recent studies demonstrated the effectiveness of deep neural networks that learn task-specific mid-level representations. Usually, such supervised learning approaches require scoreaudio pairs where individual note events of the score are aligned to the corresponding time positions of the audio excerpt. However, in practice, it is tedious to generate such strongly aligned training pairs. As one contribution of this paper, we show how to apply the Connectionist Temporal Classification (CTC) loss in the training procedure, which only uses weakly aligned training pairs. In such a pair, only the time positions of the beginning and end of a theme occurrence are annotated in an audio recording, rather than requiring local alignment annotations. We evaluate the resulting features in our theme retrieval scenario and show that they improve the state of the art for this task. As a main result, we demonstrate that with the CTC-based training procedure using weakly annotated data, we can achieve results almost as good as with strongly annotated data. Furthermore, we assess our chroma features in depth by inspecting their temporal smoothness or granularity as an important property and by analyzing the impact of different degrees of musical complexity on the features. Frank Zalkow, Meinard Müller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Local Key Estimation In Classical Music Recordings: A Cross-Version Study on Schubert's WinterreiseabstractWhile global key and chord estimation for both popular and classical music recordings have received a lot of attention, little research has been devoted to estimating the local key for classical music. In this work, we approach local key estimation on a unique cross-version dataset comprising nine performances (versions) of Schubert's song cycle Winterreise-a challenging scenario of high musical ambiguity and subjectivity. We compare an HMM-based system with a CNN-based approach. For both models, we employ a similar training procedure including the optimization of hyperparameters on a validation split. We systematically evaluate the model predictions and provide musical explanations for key confusions. As our main contribution, we explore how different training-test splits affect the models' efficacy. Splitting along the song axis, we find that both methods perform similarly well. Splitting along the version axis leads to clearly higher results, especially for the CNN, which seems to effectively learn the harmonic progressions of the songs ("cover song effect") and successfully generalizes to unseen versions. Hendrik Schreiber 0001, Christof Weiß, Meinard Müller |
ICASSP | 3 |
| 2020 | Local Key Estimation in Music Recordings: A Case Study Across Songs, Versions, and AnnotatorsabstractWhile global key and chord estimation for both popular and classical music recordings have received a lot of attention, little research has been devoted to estimating the local key for classical music. Partly, this may be due to its inherent ambiguity and subjectivity, which makes annotating local keys a challenging task. In this article, we approach local key estimation with a cross-version dataset comprising nine performances (versions) of Schubert's song cycle Winterreise annotated by three different music theory experts. We consider two baseline methods that are representative for common types of signal processing algorithms: an HMM-based system and a CNN-based approach. For both models, we employ a similar training procedure including the optimization of hyperparameters on a validation split. We systematically evaluate the model predictions and provide musical explanations for key confusions. As a central contribution, we explore how different training-test splits affect the models' efficacy. Splitting along the song axis, we find that both methods perform similarly well. Splitting along the version axis, we obtain substantially higher accuracies, especially for the CNN, which seems to effectively learn the harmonic progressions of the songs (“cover song effect”) and successfully generalizes to unseen versions. We further discuss the results for several songs in detail and assess our results from the perspective of multiple annotators. This cross-annotator study reveals that a substantial part of the systems' errors coincides with annotator disagreement and an even larger part can be traced back to musically explainable relationships among different keys. Christof Weiß, Hendrik Schreiber 0001, Meinard Müller |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Fundamental Frequency Contour Classification: A Comparison between Hand-crafted and CNN-based FeaturesabstractIn this paper, we evaluate hand-crafted features as well as features learned from data using a convolutional neural network (CNN) for different fundamental frequency classification tasks. We compare classification based on full (variable-length) contours and classification based on fixed-sized subcontours in combination with a fusion strategy. Our results show that hand-crafted and learned features lead to comparable results for both classification scenarios. Aggregating contour-level to file-level classification results generally improves the results. In comparison to the hand-crafted features, our examination indicates that the CNN-based features show a higher degree of redundancy across feature dimensions, where multiple filters (convolution kernels) specialize on similar contour shapes. Jakob Abeßer, Meinard Müller |
ICASSP | 2 |
| 2019 | Mid-level Chord Transition Features for Musical Style AnalysisabstractChords and their progressions are an important aspect of Western tonal music. Specifically, transitions between subsequent chords within a piece carry style-relevant information. To extract such information from audio recordings, a naive approach first peforms automatic chord estimation for computing chord labels explicitly and then derives transition statistics. Often, this is done with Hidden Markov Models involving the Viterbi decoding algorithm. However, since chords are often ambiguous, deciding on one “optimal” chord sequence can be problematic, which heavily affects the subsequent derivation of transition features. In this paper, we propose novel mid-level features that capture chord transitions in a “soft” way. Our method exploits the Baum-Welch algorithm, which does not involve hard decisions on chord labels. Instead, we obtain probabilistic features that account for ambiguities among chords and chord transitions. In several experiments, we evaluate these features within a style classification scenario discriminating four historical periods of Western classical music. Our soft transition features consistently achieve higher accuracies than comparable hard-decision features, thus demonstrating the descriptive power of the novel features. Christof Weiß, Fabian Brand, Meinard Müller |
ICASSP | 3 |
| 2019 | Evaluating Salience Representations for Cross-modal Retrieval of Western Classical Music RecordingsabstractIn this paper, we consider a cross-modal retrieval scenario of Western classical music. Given a short monophonic musical theme in symbolic notation as query, the objective is to find relevant audio recordings in a database. A major challenge of this retrieval task is the possible difference in the degree of polyphony between the monophonic query and the music recordings. Previous studies for popular music addressed this issue by performing the cross-modal comparison based on predominant melodies extracted from the recordings. For Western classical music, however, this approach is problematic since the underlying assumption of a single pre-dominant melody is often violated. Instead of extracting the melody explicitly, another strategy is to perform the cross-modal comparison directly on the basis of melody-enhanced salience representations. As the main contribution of this paper, we evaluate several conceptually different salience representations for our cross-modal retrieval scenario. Our extensive experimental results, which have been made available on a website, comprise more than 2000 musical themes and 100 hours of audio recordings. Frank Zalkow, Stefan Balke, Meinard Müller |
ICASSP | 3 |
| 2018 | Unifying Local and Global Methods for Harmonic-Percussive Source SeparationabstractThis paper addresses the separation of drums from music recordings, a task closely related to harmonic-percussive source separation (HPSS). In previous works, two families of algorithms have been prominently applied to this problem. They are based either on local filtering and diffusion schemes, or on global low-rank models. In this paper, we propose to combine the advantages of both paradigms. To this end, we use a local approach based on Kernel Additive Modeling (KAM) to extract an initial guess for the percussive and harmonic parts. Subsequently, we use Non-Negative Matrix Factorization (NMF) with soft activation constraints as a global approach to jointly enhance both estimates. As an additional contribution, we introduce a novel constraint for enhancing percussive activations and a scheme for estimating the percussive weight of NMF components. Throughout the paper, we use a real-world music example to illustrate the ideas behind our proposed method. Finally, we report promising BSS Eval results achieved with the publicly available test corpora ENST-Drums and QUASI, which contain isolated drum and accompaniment tracks. Christian Dittmar, Patricio López-Serrano, Meinard Müller |
ICASSP | 3 |
| 2018 | A Review of Automatic Drum TranscriptionabstractIn Western popular music, drums and percussion are an important means to emphasize and shape the rhythm, often defining the musical style. If computers were able to analyze the drum part in recorded music, it would enable a variety of rhythm-related music processing tasks. Especially the detection and classification of drum sound events by computational methods is considered to be an important and challenging research problem in the broader field of music information retrieval. Over the last two decades, several authors have attempted to tackle this problem under the umbrella term automatic drum transcription (ADT). This paper presents a comprehensive review of ADT research, including a thorough discussion of the task-specific challenges, categorization of existing techniques, and evaluation of several state-of-the-art systems. To provide more insights on the practice of ADT systems, we focus on two families of ADT techniques, namely methods based on non-negative matrix factorization and recurrent neural networks. We explain the methods’ technical details and drum-specific variations and evaluate these approaches on publicly available data sets with a consistent experimental setup. Finally, the open issues and underexplored areas in ADT research are identified and discussed, providing future directions in this field. Chih-Wei Wu, Christian Dittmar, Carl Southall, Richard Vogl, Gerhard Widmer, Jason Hockman, Meinard Müller, Alexander Lerch 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2017 | Data-driven solo voice enhancement for jazz music retrievalabstractRetrieving short monophonic queries in music recordings is a challenging research problem in Music Information Retrieval (MIR). In jazz music, given a solo transcription, one retrieval task is to find the corresponding (potentially polyphonic) recording in a music collection. Many conventional systems approach such retrieval tasks by first extracting the predominant F0-trajectory from the recording, then quantizing the extracted trajectory to musical pitches and finally comparing the resulting pitch sequence to the monophonic query. In this paper, we introduce a data-driven approach that avoids the hard decisions involved in conventional approaches: Given pairs of time-frequency (TF) representations of full music recordings and TF representations of solo transcriptions, we use a DNN-based approach to learn a mapping for transforming a “polyphonic” TF representation into a “monophonic” TF representation. This transform can be considered as a kind of solo voice enhancement. We evaluate our approach within a jazz solo retrieval scenario and compare it to a state-of-the-art method for predominant melody extraction. Stefan Balke, Christian Dittmar, Jakob Abeßer, Meinard Müller |
ICASSP | 4 |
| 2017 | Known-Artist Live Song Identification Using Audio HashprintsabstractThe goal of live song identification is to allow concertgoers to identify a live performance by recording a few seconds of the performance on their cell phone. This paper proposes a multistep approach to address this problem for popular bands. In the first step, GPS data are used to associate the audio query with a concert in order to infer who the musical artist is. This reduces the search space to a dataset containing the artist's studio recordings. In the next step, the known-artist search is solved by representing the audio as a sequence of binary codes called hashprints, which can be efficiently matched against the database using a two-stage cross-correlation approach. The hashprint representation is derived from a set of spectrotemporal filters that are learned in an unsupervised artist-specific manner. On the Gracenote live song identification benchmark, the proposed system outperforms five other baseline systems and improves the mean reciprocal rank of the previous state of the art from 0.68 to 0.79, while simultaneously reducing the average runtime per query from 10 to 0.9 s. We conduct extensive analyses of major factors affecting system performance. T. J. Tsai 0001, Thomas Prätzlich, Meinard Müller |
IEEE Trans. Multim. | 3 |
| 2016 | Retrieving audio recordings using musical themesabstractIn 1948, Barlow and Morgenstern released a collection of about 10,000 themes of well-known instrumental pieces from the corpus of Western Classical music [1]. These monophonic themes (usually four bars long) are often the most memorable parts of a piece of music. In this paper, we report on a systematic study considering a cross-modal retrieval scenario. Using a musical theme as a query, the objective is to identify all related music recordings from a given audio collection. By adapting well-known retrieval techniques, our main goal is to get a better understanding of the various challenges including tempo deviations, musical tunings, key transpositions, and differences in the degree of polyphony between the symbolic query and the audio recordings to be retrieved. In particular, we present an oracle fusion approach that indicates upper performance limits achievable by a combination of current retrieval techniques. Stefan Balke, Vlora Arifi-Müller, Lukas Lamprecht, Meinard Müller |
ICASSP | 4 |
| 2016 | Harmonic-percussive-residual sound separation using the structure tensor on spectrogramsabstractHarmonic-percussive-residual (HPR) sound separation is a useful preprocessing tool for applications such as pitched instrument transcription or rhythm extraction. Recent methods rely on the observation that in a spectrogram representation, harmonic sounds lead to horizontal structures and percussive sounds lead to vertical structures. Furthermore, these methods associate structures that are neither horizontal nor vertical (i.e., non-harmonic, non-percussive sounds) with a residual category. However, this assumption does not hold for signals like frequency modulated tones that show fluctuating spectral structures, while nevertheless carrying tonal information. Therefore, a strict classification into horizontal and vertical is inappropriate for these signals and might lead to leakage of tonal information into the residual component. In this work, we propose a novel method that instead uses the structure tensor—a mathematical tool known from image processing—to calculate predominant orientation angles in the magnitude spectrogram. We show how this orientation information can be used to distinguish between harmonic, percussive, and residual signal components, even in the case of frequency modulated signals. Finally, we verify the effectiveness of our method by means of both objective evaluation measures as well as audio examples. Richard Fug, Andreas Niedermeier, Jonathan Driedger, Sascha Disch, Meinard Müller |
ICASSP | 5 |
| 2016 | Memory-restricted multiscale dynamic time warpingabstractDynamic Time Warping (DTW) is an established method for finding a global alignment between two feature sequences. However, having a computational complexity that is quadratic in the input length, memory consumption becomes a major issue when dealing with long feature sequences. Various strategies have been proposed to reduce the memory requirements of DTW. For example, online alignment approaches often have a constant memory consumption by applying forward path estimation strategies. However, this comes at the cost of robustness. Efficient offline DTW based on multiscale strategies constitutes another approach. While methods built on this principle are usually robust, their memory requirements are still dependent on the input length. By combining ideas from online alignment approaches and offline multiscale strategies, we introduce a novel alignment procedure that allows for specifying a constant upper bound on its memory requirements. This is an important aspect when working on devices with limited computational resources. Experiments show that when restricting the memory consumption of our proposed procedure to eight megabytes, it basically yields the same alignments as the standard DTW procedure. Thomas Prätzlich, Jonathan Driedger, Meinard Müller |
ICASSP | 3 |
| 2016 | Triple-based analysis of music alignments without the need of ground-truth annotationsabstractThe goal of music alignment methods is to temporally align different versions of the same piece of music. These methods are typically evaluated by comparing the computed alignments to given ground-truth annotations. Creating such annotations is usually very labor intensive. For many musical pieces, especially in classical music, there exists a multitude of different recordings. In this work, we investigate whether an evaluation of music alignment algorithms can be performed without ground-truth annotations when at least a triplet of recordings of the same piece of music is available. The main idea is to align the time points of a fixed reference version, in a circular way, back through a second and third version by using their pairwise alignments. A triple error is then computed by comparing these time points with their circularly aligned version. In this paper, we formalize the idea of the triple error and discuss its potential and limitations. We present typical examples for the triple error and compare it to the pairwise alignment error based on ground-truth. Furthermore, we present a case study to indicate the potential of the triple error to analyze alignments and to compare different alignment methods without the need of ground-truth annotations. Thomas Prätzlich, Meinard Müller |
ICASSP | 2 |
| 2016 | Reverse Engineering the Amen Break - Score-Informed Separation and Restoration Applied to Drum RecordingsabstractThis work addresses the extraction of high-quality component signals from drum solo recordings (breakbeats) for music production and remixing purposes. Specifically, we employ audio source separation techniques to recover sound events from the drum sound mixture corresponding to the individual drum strokes. Our separation approach is based on an informed variant of non-negative matrix factor deconvolution (NMFD) that has been proposed and applied to drum transcription and separation in earlier works. In this paper, we systematically study the suitability of NMFD and the impact of audio- and score-based side information in the context of drum separation. In the case of imperfect decompositions, we observe different cross-talk artifacts appearing during the attack and the decay segment of the extracted drum sounds. Based on these findings, we propose and evaluate two extensions to the core technique. The first extension is based on applying a cascaded NMFD decomposition while retaining selected side information. The second extension is a time–frequency selective restoration approach using a dictionary of single note drum sounds. For all our experiments, we use a publicly available data set consisting of multitrack drum recordings and corresponding annotations that allows us to evaluate the source separation quality. Using this test set, we show that our proposed methods improve the quality of the component signals. Christian Dittmar, Meinard Müller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Matching Musical Themes based on noisy OCR and OMR inputabstractIn the year 1948, Barlow and Morgenstern published the book “A Dictionary of Musical Themes”, which contains 9803 important musical themes from the Western classical music literature. In this paper, we deal with the problem of automatically matching these themes to other digitally available sources. To this end, we introduce a processing pipeline that automatically extracts from the scanned pages of the printed book textual metadata using Optical Character Recognition (OCR) as well as symbolic note information using Optical Music Recognition (OMR). Due to the poor printing quality of the book, the OCR and OMR results are quite noisy containing numerous extraction errors. As one main contribution, we adjust alignment techniques for matching musical themes based on the OCR and OMR input. In particular, we show how the matching quality can be substantially improved by fusing the OCR- and OMR-based matching results. Finally, we report on our experiments within the challenging Barlow and Morgenstern scenario, which also indicates the potential of our techniques when considering other sources of musical themes such as digital music archives and the world wide web. Stefan Balke, Sanu Pulimootil Achankunju, Meinard Müller |
ICASSP | 3 |
| 2015 | Extracting singing voice from music recordings by cascading audio decomposition techniquesabstractThe problem of extracting singing voice from music recordings has received increasing research interest in recent years. Many proposed decomposition techniques are based on one of the following two strategies. The first approach is to directly decompose a given music recording into one component for the singing voice and one for the accompaniment by exploiting knowledge about specific characteristics of singing voice. Procedures following the second approach disassemble the recording into a large set of fine-grained components, which are classified and reassembled afterwards to yield the desired source estimates. In this paper, we propose a novel approach that combines the strengths of both strategies. We first apply different audio decomposition techniques in a cascaded fashion to disassemble the music recording into a set of mid-level components. This decomposition is fine enough to model various characteristics of singing voice, but coarse enough to keep an explicit semantic meaning of the components. These properties allow us to directly reassemble the singing voice and the accompaniment from the components. Our objective and subjective evaluations show that this strategy can compete with state-of-the-art singing voice separation algorithms and yields perceptually appealing results. Jonathan Driedger, Meinard Müller |
ICASSP | 2 |
| 2015 | Estimating double thumbnails for music recordingsabstractAudio thumbnailing, which aims at finding the most representative audio segment of a music recording, is an important task in music information retrieval. In general, the notion of a thumbnail is not well-defined and several musical parts may be good thumbnail candidates. For example, for popular music, both a verse and a refrain section may serve as suitable thumbnail candidates. Instead of considering only one thumbnail, we consider in this paper the problem of finding the two most representative segments that correspond to different musical parts. We denote these two segments as double thumbnails. As our main technical contributions, we propose two approaches for computing double thumbnails, both extending a previously introduced repetition-based thumbnailing procedure. In the first approach, which is straightforward, we simply apply the original thumbnailing procedure two times in an iterative fashion. In the second approach, we introduce a novel method for jointly estimating the two thumbnails within one optimization procedure. Finally, we report on experimental results demonstrating the performances of the two double thumbnailing procedures and indicate directions towards full music structure analysis. Nanzhu Jiang, Meinard Müller |
ICASSP | 2 |
| 2015 | Kernel Additive Modeling for interference reduction in multi-channel music recordingsabstractWhen recording a live musical performance, the different voices, such as the instrument groups or soloists of an orchestra, are typically recorded in the same room simultaneously, with at least one microphone assigned to each voice. However, it is difficult to acoustically shield the microphones. In practice, each one contains interference from every other voice. In this paper, we aim to reduce these interferences in multi-channel recordings to recover only the isolated voices. Following the recently proposed Kernel Additive Modeling framework, we present a method that iteratively estimates both the power spectral density of each voice and the corresponding strength in each microphone signal. With this information, we build an optimal Wiener filter, strongly reducing interferences. The trade-off between distortion and separation can be controlled by the user through the number of iterations of the algorithm. Furthermore, we present a computationally effective approximation of the iterative procedure. Listening tests demonstrate the effectiveness of the method. Thomas Prätzlich, Rachel M. Bittner, Antoine Liutkus, Meinard Müller |
ICASSP | 4 |
| 2015 | Novel audio features for capturing tempo salience in music recordingsabstractIn music compositions, certain parts may be played in an improvisational style with a rather vague notion of tempo, while other parts are characterized by having a clearly perceivable tempo. Based on this observation, we introduce in this paper some novel audio features for capturing tempo-related information. Rather than measuring the specific tempo of a local section of a given recording, our objective is to capture the existence or absence of a notion of tempo, a kind of tempo salience. By a quantitative analysis within an Indian music scenario, we demonstrate that our audio features capture the aspect of tempo salience well, while being independent of continuous fluctuations and local changes in tempo. Balaji Thoshkahna, Meinard Müller, Venkatesh Kulkarni, Nanzhu Jiang |
ICASSP | 2 |
| 2015 | Tonal complexity features for style classification of classical musicabstractWe propose a set of novel audio features for classifying the style of classical music. The features rely on statistical measures based on a chroma feature representation of the audio data and describe the tonal complexity of the music, independently from the orchestration or timbre of the music. To analyze this property, we use a dataset containing piano and orchestral music from four general historical periods including Baroque, Classical, Romantic, and Modern. By applying dimensionality reduction techniques, we derive visualizations that demonstrate the discriminative power of the features with regard to the music styles. In classification experiments, we evaluate the features' performance using an SVM classifier. We investigate the influence of artist filtering with respect to the individual composers on the classification performance. In all experiments, we compare the results to the performance of standard features. We show that the introduced features capture meaningful properties of musical style and are robust to timbral variations. Christof Weiß, Meinard Müller |
ICASSP | 2 |
| 2014 | TSM Toolbox: MATLAB Implementations of Time-Scale Modification Algorithms
Jonathan Driedger, Meinard Müller |
DAFx | 2 |
| 2014 | Towards efficient audio thumbnailingabstractAudio thumbnailing, which aims at finding the most representative audio segment of a music recording, is an important task in music information retrieval. In this paper, we show how the computational efficiency of a recently proposed state-of-the-art thumbnailing approach can be improved significantly. The basic idea of the previous approach is to compute for each possible segment a fitness value that expresses repetitiveness and then to define the thumbnail as the fitness-maximizing segment. As a first acceleration strategy, we propose an efficient multi-level sampling strategy to reduce the number of segments the fitness has to be computed for. Second, we obtain further accelerations by suitably adjusting the resolution used in the fitness computation depending on the level of the segment. As a third contribution, we exploit an intrinsic property of the fitness computation that allows us to estimate the fitness for certain segments without any further computation. Our experimental results show that combining these three strategies leads to accelerations by a factor of 20 to 200 depending on the duration of the song while keeping the overall accuracy for the thumbnail estimation. Nanzhu Jiang, Meinard Müller |
ICASSP | 2 |
| 2014 | Exploiting global features for tempo octave correctionabstractTempo estimation is a fundamental problem in music information retrieval. Most approaches attempt to solve two problems: first finding a dominant pulse and second correcting the metrical level of this pulse. The latter has also been dubbed fixing the octave error. We propose an algorithm for tempo estimation that addresses both problems mostly independently. While using a standard pulse detection technique, for octave error correction, we exploit a simple relationship between a single global feature, average spectral novelty, and listener perception of musical tempo. The proposed method is extremely simple. Nevertheless, it outperforms most existing tempo estimation methods and is on par with the best-performing ones. It thus exemplifies that a global feature-based approach can significantly improve tempo estimation. Hendrik Schreiber 0001, Meinard Müller |
ICASSP | 2 |
| 2014 | Improving Time-Scale Modification of Music Signals Using Harmonic-Percussive SeparationabstractA major problem in time-scale modification (TSM) of music signals is that percussive transients are often perceptually degraded. To prevent this degradation, some TSM approaches try to explicitly identify transients in the input signal and to handle them in a special way. However, such approaches are problematic for two reasons. First, errors in the transient detection have an immediate influence on the final TSM result and, second, a perceptual transparent preservation of transients is by far not a trivial task. In this paper we present a TSM approach that handles transients implicitly by first separating the signal into a harmonic component as well as a percussive component which typically contains the transients. While the harmonic component is modified with a phase vocoder approach using a large frame size, the noise-like percussive component is modified with a simple time-domain overlap-add technique using a short frame size, which preserves the transients to a high degree without any explicit transient detection. Jonathan Driedger, Meinard Müller, Sebastian Ewert |
IEEE Signal Process. Lett. | 2 |
| 2014 | Accelerating Index-Based Audio IdentificationabstractIn view of rapidly growing digital music collections and ubiquitous music consumption, the development of technologies for identifying, browsing, and managing audio content has become a major strand of research. In this context, audio identification (ID) systems for identifying audio recordings by means of short query audio clips have become of commercial relevance. In this paper, we take a closer look at a widely used audio ID system originally developed by Haitsma and Kalker and propose several modifications that yield significant improvements with regard to retrieval speed and storage requirements. As the main contribution, we introduce a measure that establishes a connection between the temporal correlation of hash values (used for indexing) and their ability to survive in the presence of noise and signal distortions. Based on this measure, we improve the overall performance of the audio ID system by means of four strategies. First, we change the way fingerprints (audio features) are generated to increase their reliability. Second, by prioritizing more reliable hash values when searching for reference entries, we achieve substantial gains in retrieval speed by a factor of almost seven. Third, by enlarging the query fingerprint, we increase our chances of identifying reliable hash values. Fourth, by indexing only the most reliable hashes, thus applying a sub-sampling strategy, we significantly lower the server side storage requirements by a factor of ten. Hendrik Schreiber 0001, Meinard Müller |
IEEE Trans. Multim. | 2 |
| 2014 | Unsupervised Music Structure Annotation by Time Series Structure Features and Segment SimilarityabstractAutomatically inferring the structural properties of raw multimedia documents is essential in today's digitized society. Given its hierarchical and multi-faceted organization, musical pieces represent a challenge for current computational systems. In this article, we present a novel approach to music structure annotation based on the combination of structure features with time series similarity. Structure features encapsulate both local and global properties of a time series, and allow us to detect boundaries between homogeneous, novel, or repeated segments. Time series similarity is used to identify equivalent segments, corresponding to musically meaningful parts. Extensive tests with a total of five benchmark music collections and seven different human annotations show that the proposed approach is robust to different ground truth choices and parameter settings. Moreover, we see that it outperforms previous approaches evaluated under the same framework. Joan Serrà, Meinard Müller, Peter Grosche, Josep Lluís Arcos |
IEEE Trans. Multim. | 2 |
| 2013 | Personalization and Evaluation of a Real-Time Depth-Based Full Body TrackerabstractReconstructing a three-dimensional representation of human motion in real-time constitutes an important research topic with applications in sports sciences, human-computer-interaction, and the movie industry. In this paper, we contribute with a robust algorithm for estimating a personalized human body model from just two sequentially captured depth images that is more accurate and runs an order of magnitude faster than the current state-of-the-art procedure. Then, we employ the estimated body model to track the pose in real-time from a stream of depth images using a tracking algorithm that combines local pose optimization and a stabilizing dataBase look-up. Together, this enables accurate pose tracking that is more accurate than previous approaches. As a further contribution, we evaluate and compare our algorithm to previous work on a comprehensive benchmark dataset containing more than 15 minutes of challenging motions. This dataset comprises calibrated marker-Based motion capture data, depth data, as well as ground truth tracking results and is publicly available for research purposes. Thomas Helten, Andreas Baak, Gaurav Bharaj, Meinard Müller, Hans-Peter Seidel, Christian Theobalt |
3DV | 4 |
| 2013 | Efficient data adaption for musical source separation methods based on parametric modelsabstractThe decomposition of a monaural audio recording into musically meaningful sound sources constitutes one of the central research topics in music signal processing. In this context, many recent approaches employ parametric models that describe a recording in a highly structured and musically informed way. However, a major drawback of such approaches is that the parameter learning process typically relies on computationally expensive data adaption methods. In this paper, the main idea is to distinguish parameters in which the model is linear explicitly from the remaining parameters. Exploiting the linearity we translate the data adaption problem into a sparse linear least squares problem with box constraints (SLLS-BC), a class of problems for which highly efficient numerical solvers exist. First experiments show that our approach based on modified SLLS-BC methods accelerates the data adaption by a factor of four or more compared to recently proposed methods. Sebastian Ewert, Meinard Müller, Mark B. Sandler |
ICASSP | 2 |
| 2013 | Real-Time Body Tracking with One Depth Camera and Inertial SensorsabstractIn recent years, the availability of inexpensive depth cameras, such as the Microsoft Kinect, has boosted the research in monocular full body skeletal pose tracking. Unfortunately, existing trackers often fail to capture poses where a single camera provides insufficient data, such as non-frontal poses, and all other poses with body part occlusions. In this paper, we present a novel sensor fusion approach for real-time full body tracking that succeeds in such difficult situations. It takes inspiration from previous tracking solutions, and combines a generative tracker and a discriminative tracker retrieving closest poses in a database. In contrast to previous work, both trackers employ data from a low number of inexpensive body-worn inertial sensors. These sensors provide reliable and complementary information when the monocular depth information alone is not sufficient. We also contribute by new algorithmic solutions to best fuse depth and inertial data in both trackers. One is a new visibility model to determine global body pose, occlusions and usable depth correspondences and to decide what data modality to use for discriminative tracking. We also contribute with a new inertial-based pose retrieval, and an adapted late fusion step to calculate the final body pose. Thomas Helten, Meinard Müller, Hans-Peter Seidel, Christian Theobalt |
ICCV | 2 |
| 2013 | Score-informed audio decomposition and applicationsabstractThe separation of different sound sources from polyphonic music recordings constitutes a complex task since one has to account for different musical and acoustical aspects. In the last years, various score-informed procedures have been suggested where musical cues such as pitch, timing, and track information are used to support the source separation process. In this paper, we discuss a framework for decomposing a given music recording into notewise audio events which serve as elementary building blocks. In particular, we introduce an interface that employs the additional score information to provide a natural way for a user to interact with these audio events. By simply selecting arbitrary note groups within the score a user can access, modify, or analyze corresponding events in a given audio recording. In this way, our framework not only opens up new ways for audio editing applications, but also serves as a valuable tool for evaluating and better understanding the results of source separation algorithms. Jonathan Driedger, Harald Grohganz, Thomas Prätzlich, Sebastian Ewert, Meinard Müller |
ACM Multimedia | 5 |
| 2013 | Towards cover group thumbnailingabstractIn this paper we investigate whether we can extract the commonalities shared by a group of cover songs or versions of the same musical piece. As a main contribution, we introduce the concept of cover group thumbnail, which is the most representative, essential subsequence for an entire group of versions. Opposed to previous approaches, we jointly consider all versions of a given song to compute a single cover group template, which then shows a high degree of robustness against version-specific aspects. To compute such a template, we introduce a modification of a recent audio thumbnailing technique. To evaluate the reliability of our conceptual contribution, we consider the task of template-based version identification, where we show comparable accuracies to existing systems. Peter Grosche, Meinard Müller, Joan Serrà |
ACM Multimedia | 2 |
| 2013 | A Robust Fitness Measure for Capturing Repetitions in Music Recordings With Applications to Audio ThumbnailingabstractThe automatic extraction of structural information from music recordings constitutes a central research topic. In this paper, we deal with a subproblem of audio structure analysis called audio thumbnailing with the goal to determine the audio segment that best represents a given music recording. Typically, such a segment has many (approximate) repetitions covering large parts of the recording. As the main technical contribution, we introduce a novel fitness measure that assigns a fitness value to each segment that expresses how much and how well the segment “explains” the repetitive structure of the entire recording. The thumbnail is then defined to be the fitness-maximizing segment. To compute the fitness measure, we describe an optimization scheme that jointly performs two error-prone steps, path extraction and grouping, which are usually performed successively. As a result, our approach is even able to cope with strong musical and acoustic variations that may occur within and across related segments. As a further contribution, we introduce the concept of fitness scape plots that reveal global structural properties of an entire recording. Finally, to show the robustness and practicability of our thumbnailing approach, we present various experiments based on different audio collections that comprise popular music, classical music, and folk song field recordings. Meinard Müller, Nanzhu Jiang, Peter Grosche |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Unsupervised Detection of Music Boundaries by Time Series Structure FeaturesabstractLocating boundaries between coherent and/or repetitive segments of a time series is a challenging problem pervading many scientific domains. In this paper we propose an unsupervised method for boundary detection, combining three basic principles: novelty, homogeneity, and repetition. In particular, the method uses what we call structure features, a representation encapsulating both local and global properties of a time series. We demonstrate the usefulness of our approach in detecting music structure boundaries, a task that has received much attention in recent years and for which exist several benchmark datasets and publicly available annotations. We find our method to significantly outperform the best accuracies published so far. Importantly, our boundary approach is generic, thus being applicable to a wide range of time series beyond the music and audio domains. Joan Serrà, Meinard Müller, Peter Grosche, Josep Lluís Arcos |
AAAI | 2 |
| 2012 | Using score-informed constraints for NMF-based source separationabstractTechniques based on non-negative matrix factorization (NMF) can be used to efficiently decompose a magnitude spectrogram into a set of template (column) vectors and activation (row) vectors. To better control this decomposition, NMF has been extended using prior knowledge and parametric models. In this paper, we present such an extended approach that uses additional score information to guide the decomposition process. Here, opposed to previous methods, our main idea is to impose constraints on both the template as well as the activation side. We show that using such double constraints results in musically meaningful decompositions similar to parametric approaches, while being computationally less demanding and easier to implement. Furthermore, additional onset constraints can be incorporated in a straightforward manner without sacrificing robustness. We evaluate our approach in the context of separating note groups (e. g. the left or right hand) from monaural piano recordings. Sebastian Ewert, Meinard Müller |
ICASSP | 2 |
| 2012 | Toward musically-motivated audio fingerprintsabstractIn this paper, we investigate to which extent well-known audio fingerprinting techniques, which aim at identifying a specific audio recording, can be modified to also deal with more musical variations. To this end, we replace the standard peak fingerprints based on a spectrogram by peak fingerprints based on other more “musical” feature representations. Our systematic experiments show that such modified peak fingerprints allow for a robust identification of different versions and performances of the same piece of music if the query length is at least 15 seconds. This indicates that highly efficient audio fingerprinting techniques can also be applied to accelerate tasks such as audio matching or cover song identification. Peter Grosche, Meinard Müller |
ICASSP | 2 |
| 2012 | Toward characteristic audio shingles for efficient cross-version music retrievalabstractThe general goal of cross-version music retrieval is to identify all versions of a given piece of music by means of a short query audio fragment. To speed up the retrieval process, hashing techniques have been proposed, where the audio material is split up into small overlapping shingles (used as hashes) that consist of short feature subsequences. In this paper, we extend this work with the goal to minimize the number of hash lookups. To this end, one requires larger shingles that characterize the underlying piece of music to a high degree, while being robust to variations that occur across different versions. As our main contribution, we report on extensive experiments to highlight the delicate trade-off between the query length, feature parameters, shingle dimension, and index settings. These insights are of fundamental importance for building efficient cross-version retrieval systems that scale to millions of songs. Peter Grosche, Meinard Müller |
ICASSP | 2 |
| 2012 | 2nd international ACM workshop on music information retrieval with user-centered and multimodal strategies (MIRUM)abstractThe International ACM Workshop on Music Information Retrieval with User-Centered and Multimodal Strategies (MIRUM) at ACM Multimedia was proposed in order to gather experts from the Music and Multimedia Information Retrieval communities, as well as other neighboring fields, and to provide a high-profile platform for presenting current work on Music Information Retrieval with a strong focus on user-centered and multimodal approaches. Following a successful first edition at ACM Multimedia 2011, a second edition of MIRUM was held at ACM Multimedia 2012, which is the focus of this overview. After a description of the rationale and focus areas of the workshop, the accepted submissions and other program elements are summarized. Cynthia C. S. Liem, Meinard Müller, Steven K. Tjoa, George Tzanetakis |
ACM Multimedia | 2 |
| 2012 | Towards Cross-Version Harmonic Analysis of MusicabstractFor a given piece of music, there often exist multiple versions belonging to the symbolic (e.g., MIDI representations), acoustic (audio recordings), or visual (sheet music) domain. Each type of information allows for applying specialized, domain-specific approaches to music analysis tasks. In this paper, we formulate the idea of a cross-version analysis for comparing and/or combining analysis results from different representations. As an example, we realize this idea in the context of harmonic analysis to automatically evaluate MIDI-based chord labeling procedures using annotations given for corresponding audio recordings. To this end, one needs reliable synchronization procedures that automatically establish the musical relationship between the multiple versions of a given piece. This becomes a hard problem when there are significant local deviations in these versions. We introduce a novel late-fusion approach that combines different alignment procedures in order to identify reliable parts in synchronization results. Then, the cross-version comparison of the various chord labeling results is performed only on the basis of the reliable parts. Finally, we show how inconsistencies in these results across the different versions allow for a quantitative and qualitative evaluation, which not only indicates limitations of the employed chord labeling strategies but also deepens the understanding of the underlying music material. Sebastian Ewert, Meinard Müller, Verena Konz, Daniel Müllensiefen, Geraint A. Wiggins |
IEEE Trans. Multim. | 2 |
| 2011 | Estimating note intensities in music recordingsabstractIn this paper, we present automated methods for estimating note intensities in music recordings. Given a MIDI file (representing the score) and an audio recording (representing an interpretation) of a piece of music, our idea is to parametrize the spectrogram of the audio recording by exploiting the MIDI information and then to estimate the note intensities from the resulting model. The model is based on the idea of note-event spectrograms describing the part of a spectrogram that can be attributed to a given note event. After initializing our model with note events provided by the MIDI, we adapt all model parameters such that our model spectrogram approximates the audio spectrogram as accurately as possible. While note-wise intensity estimation is a very challenging task for general music, our experiments indicate promising results on polyphonic piano music. Sebastian Ewert, Meinard Müller |
ICASSP | 2 |
| 2011 | A data-driven approach for real-time full body pose reconstruction from a depth cameraabstractIn recent years, depth cameras have become a widely available sensor type that captures depth images at real-time frame rates. Even though recent approaches have shown that 3D pose estimation from monocular 2.5D depth images has become feasible, there are still challenging problems due to strong noise in the depth data and self-occlusions in the motions being captured. In this paper, we present an efficient and robust pose estimation framework for tracking full-body motions from a single depth image stream. Following a data-driven hybrid strategy that combines local optimization with global retrieval techniques, we contribute several technical improvements that lead to speed-ups of an order of magnitude compared to previous approaches. In particular, we introduce a variant of Dijkstra's algorithm to efficiently extract pose features from the depth data and describe a novel late-fusion scheme based on an efficiently computable sparse Hausdorff distance to combine local and global pose estimates. Our experiments show that the combination of these techniques facilitates real-time tracking with stable results even for fast and complex motions, making it applicable to a wide range of inter-active scenarios. Andreas Baak, Meinard Müller, Gaurav Bharaj, Hans-Peter Seidel, Christian Theobalt |
ICCV | 2 |
| 2011 | Outdoor human motion capture using inverse kinematics and von mises-fisher samplingabstractHuman motion capturing (HMC) from multiview image sequences is an extremely difficult problem due to depth and orientation ambiguities and the high dimensionality of the state space. In this paper, we introduce a novel hybrid HMC system that combines video input with sparse inertial sensor input. Employing an annealing particle-based optimization scheme, our idea is to use orientation cues derived from the inertial input to sample particles from the manifold of valid poses. Then, visual cues derived from the video input are used to weight these particles and to iteratively derive the final pose. As our main contribution, we propose an efficient sampling procedure where the particles are derived analytically using inverse kinematics on the orientation cues. Additionally, we introduce a novel sensor noise model to account for uncertainties based on the von Mises-Fisher distribution. Doing so, orientation constraints are naturally fulfilled and the number of needed particles can be kept very small. More generally, our method can be used to sample poses that fulfill arbitrary orientation or positional kinematic constraints. In the experiments, we show that our system can track even highly dynamic motions in an outdoor environment with changing illumination, background clutter, and shadows. Gerard Pons-Moll, Andreas Baak, Juergen Gall, Laura Leal-Taixé, Meinard Müller, Hans-Peter Seidel, Bodo Rosenhahn |
ICCV | 5 |
| 2011 | 1st international ACM workshop on music information retrieval with user-centered and multimodal strategies (MIRUM)abstractThe 1st International ACM Workshop on Music Information Retrieval with User-Centered and Multimodal Strategies (MIRUM) at ACM Multimedia was proposed in order to gather experts from the Music and Multimedia Information Retrieval communities, as well as other neighboring fields. The workshop aims to provide a high-profile platform for presenting current work on Music Information Retrieval, with strong focus on user-centered and multimodal approaches. These focus areas are not only relevant to the Music Information Retrieval field, but equally recognized as emerging and relevant in the Multimedia domain. This way, a cross-disciplinary dialogue on open challenges can be initiated, facilitating new bridging opportunities and increased exchanges of expertise between communities. In this summary, we provide an overview of the 1st MIRUM workshop. After a description of the rationale and focus areas of the workshop, the accepted submissions and other program elements are summarized. Cynthia C. S. Liem, Meinard Müller, Douglas Eck, George Tzanetakis |
ACM Multimedia | 2 |
| 2011 | Extracting Predominant Local Pulse Information From Music RecordingsabstractThe extraction of tempo and beat information from music recordings constitutes a challenging task in particular for non-percussive music with soft note onsets and time-varying tempo. In this paper, we introduce a novel mid-level representation that captures musically meaningful local pulse information even for the case of complex music. Our main idea is to derive for each time position a sinusoidal kernel that best explains the local periodic nature of a previously extracted note onset representation. Then we employ an overlap-add technique accumulating all these kernels over time to obtain a single function that reveals the predominant local pulse (PLP). Our concept introduces a high degree of robustness to noise and distortions resulting from weak and blurry onsets. Furthermore, the resulting PLP curve reveals the local pulse information even in the presence of continuous tempo changes and indicates a kind of confidence in the periodicity estimation. As further contribution, we show how our PLP concept can be used as a flexible tool for enhancing tempo estimation and beat tracking. The practical relevance of our approach is demonstrated by extensive experiments based on music recordings of various genres. Peter Grosche, Meinard Müller |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Motion reconstruction using sparse accelerometer dataabstractThe development of methods and tools for the generation of visually appealing motion sequences using prerecorded motion capture data has become an important research area in computer animation. In particular, data-driven approaches have been used for reconstructing high-dimensional motion sequences from low-dimensional control signals. In this article, we contribute to this strand of research by introducing a novel framework for generating full-body animations controlled by only four 3D accelerometers that are attached to the extremities of a human actor. Our approach relies on a knowledge base that consists of a large number of motion clips obtained from marker-based motion capturing. Based on the sparse accelerometer input a cross-domain retrieval procedure is applied to build up a lazy neighborhood graph in an online fashion. This graph structure points to suitable motion fragments in the knowledge base, which are then used in the reconstruction step. Supported by a kd-tree index structure, our procedure scales to even large datasets consisting of millions of frames. Our combined approach allows for reconstructing visually plausible continuous motion streams, even in the presence of moderate tempo variations which may not be directly reflected by the given knowledge base. Jochen Tautges, Arno Zinke, Björn Krüger, Jan Baumann, Andreas Weber 0004, Thomas Helten, Meinard Müller, Hans-Peter Seidel, Bernd Eberhardt |
ACM Trans. Graph. | 7 |
| 2010 | Introducing the Interpretation Switcher Interface to Music Education
Verena Konz, Meinard Müller |
CSEDU (1) | 2 |
| 2010 | Multisensor-fusion for 3D full-body human motion captureabstractIn this work, we present an approach to fuse video with orientation data obtained from extended inertial sensors to improve and stabilize full-body human motion capture. Even though video data is a strong cue for motion analysis, tracking artifacts occur frequently due to ambiguities in the images, rapid motions, occlusions or noise. As a complementary data source, inertial sensors allow for drift-free estimation of limb orientations even under fast motions. However, accurate position information cannot be obtained in continuous operation. Therefore, we propose a hybrid tracker that combines video with a small number of inertial units to compensate for the drawbacks of each sensor type: on the one hand, we obtain drift-free and accurate position information from video data and, on the other hand, we obtain accurate limb orientations and good performance under fast motions from inertial sensors. In several experiments we demonstrate the increased performance and stability of our human motion tracker. Gerard Pons-Moll, Andreas Baak, Thomas Helten, Meinard Müller, Hans-Peter Seidel, Bodo Rosenhahn |
CVPR | 4 |
| 2010 | Cyclic tempogram - A mid-level tempo representation for musicsignalsabstractThe extraction of local tempo and beat information from audio recordings constitutes a challenging task, particularly for music that reveals significant tempo variations. Furthermore, the existence of various pulse levels such as measure, tactus, and tatum often makes the determination of absolute tempo problematic. In this paper, we present a robust mid-level representation that encodes local tempo information. Similar to the well-known concept of cyclic chroma features, where pitches differing by octaves are identified, we introduce the concept of cyclic tempograms, where tempi differing by a power of two are identified. Furthermore, we describe how to derive cyclic tempograms from music signals using two different methods for periodicity analysis and finally sketch some applications to tempo-based audio segmentation. Peter Grosche, Meinard Müller, Frank Kurth |
ICASSP | 2 |
| 2010 | Perceptual audio features for unsupervised key-phrase detectionabstractWe propose a new type of audio feature (HFCC-ENS) as well as an unsupervised method for detecting short sequences of spoken words (key-phrases) within long speech recordings. Our technical contributions are threefold: Firstly, we propose to use bandwidth-adapted filterbanks instead of classical MFCC-style filters in the feature extraction step. Secondly, the time resolution of the resulting features is adapted to account for the temporal characteristics of the spoken phrases. Thirdly, the key-phrase detection step is performed by matching sequences of the resulting HFCC-ENS features with features extracted from a target speech recording. We evaluate the proposed method using the German Kiel Corpus and furthermore investigate speech-related properties of the proposed feature. Dirk von Zeddelmann, Frank Kurth, Meinard Müller |
ICASSP | 3 |
| 2010 | Towards Timbre-Invariant Audio Features for Harmony-Based MusicabstractChroma-based audio features are a well-established tool for analyzing and comparing harmony-based Western music that is based on the equal-tempered scale. By identifying spectral components that differ by a musical octave, chroma features possess a considerable amount of robustness to changes in timbre and instrumentation. In this paper, we describe a novel procedure that further enhances chroma features by significantly boosting the degree of timbre invariance without degrading the features' discriminative power. Our idea is based on the generally accepted observation that the lower mel-frequency cepstral coefficients (MFCCs) are closely related to timbre. Now, instead of keeping the lower coefficients, we discard them and only keep the upper coefficients. Furthermore, using a pitch scale instead of a mel scale allows us to project the remaining coefficients onto the 12 chroma bins. We present a series of experiments to demonstrate that the resulting chroma features outperform various state-of-the art features in the context of music matching and retrieval applications. As a final contribution, we give a detailed analysis of our enhancement procedure revealing the musical meaning of certain pitch-frequency cepstral coefficients. Meinard Müller, Sebastian Ewert |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | High resolution audio synchronization using chroma onset featuresabstractThe general goal of music synchronization is to automatically align the multiple information sources such as audio recordings, MIDI files, or digitized sheet music related to a given musical work. In computing such alignments, one typically has to face a delicate tradeoff between robustness and accuracy. In this paper, we introduce novel audio features that combine the high temporal accuracy of onset features with the robustness of chroma features. We show how previous synchronization methods can be extended to make use of these new features. We report on experiments based on polyphonic Western music demonstrating the improvements of our proposed synchronization framework. Sebastian Ewert, Meinard Müller, Peter Grosche |
ICASSP | 2 |
| 2009 | Making chroma features more robust to timbre changesabstractChroma-based audio features are a well-established tool for analyzing and comparing music data. By identifying spectral components that differ by a musical octave, chroma features show a high degree of invariance to variations in timbre. In this paper, we describe a novel procedure for making chroma features even more robust to changes in timbre and instrumentation while keeping their discriminative power. Our idea is based on the generally accepted observation that the lower mel-frequency cepstral coefficients (MFCCs) are closely related to timbre. Now, instead of keeping the lower coefficients, we discard them and only keep the upper coefficients. Furthermore, using a pitch scale instead of a mel scale allows us to project the remaining coefficients onto the twelve chroma bins. Our systematic experiments show that the resulting chroma features have indeed gained a significant boost towards timbre invariance. Meinard Müller, Sebastian Ewert, Sebastian Kreuzer |
ICASSP | 1 |
| 2009 | Stabilizing motion tracking using retrieved motion priorsabstractIn this paper, we introduce a novel iterative motion tracking framework that combines 3D tracking techniques with motion retrieval for stabilizing markerless human motion capturing. The basic idea is to start human tracking without prior knowledge about the performed actions. The resulting 3D motion sequences, which may be corrupted due to tracking errors, are locally classified according to available motion categories. Depending on the classification result, a retrieval system supplies suitable motion priors, which are then used to regularize and stabilize the tracking in the next iteration step. Experiments with the HumanEVA-II benchmark show that tracking and classification are remarkably improved after few iterations. Andreas Baak, Bodo Rosenhahn, Meinard Müller, Hans-Peter Seidel |
ICCV | 3 |
| 2008 | Path-constrained partial music synchronizationabstractDigital music collections often contain different versions and interpretations of a single musical work. In view of music retrieval and browsing applications, one important task, also referred to as audio synchronization, is to automatically time-align two given audio recordings of the same underlying piece. In this paper, we present a novel synchronization procedure, which can compute meaningful audio alignments even in the presence of structural variations. Such variations include the omission of repetitions, the insertion of additional parts (soli, cadenzas), or differences in the number of stanzas in popular, folk, or art songs. As one main contribution, we introduce the concept of path-constrained similarity matrices. This enables us to employ a flexible and efficiently computable partial matching procedure in the optimization step of our synchronization algorithm. Our overall strategy aims at aligning preferably long consecutive runs while avoiding an over-fragmentation of the audio material. Meinard Müller, Daniel Appelt |
ICASSP | 1 |
| 2008 | Multimodal presentation and browsing of musicabstractRecent digitization efforts have led to large music collections, which contain music documents of various modes comprising textual, visual and acoustic data. In this paper, we present a multimodal music player for presenting and browsing digitized music collections consisting of heterogeneous document types. In particular, we concentrate on music documents of two widely used types for representing a musical work, namely visual music representation (scanned images of sheet music) and associated interpretations (audio recordings). We introduce novel user interfaces for multimodal (audio-visual) music presentation as well as intuitive navigation and browsing. Our system offers high quality audio playback with time-synchronous display of the digitized sheet music associated to a musical work. Furthermore, our system enables a user to seamlessly crossfade between various interpretations belonging to the currently selected musical work. David Damm, Christian Fremerey, Frank Kurth, Meinard Müller, Michael Clausen |
ICMI | 4 |
| 2008 | Efficient Index-Based Audio MatchingabstractGiven a large audio database of music recordings, the goal of classical audio identification is to identify a particular audio recording by means of a short audio fragment. Even though recent identification algorithms show a significant degree of robustness towards noise, MP3 compression artifacts, and uniform temporal distortions, the notion of similarity is rather close to the identity. In this paper, we address a higher level retrieval problem, which we refer to as audio matching: given a short query audio clip, the goal is to automatically retrieve all excerpts from all recordings within the database that musically correspond to the query. In our matching scenario, opposed to classical audio identification, we allow semantically motivated variations as they typically occur in different interpretations of a piece of music. To this end, this paper presents an efficient and robust audio matching procedure that works even in the presence of significant variations, such as nonlinear temporal, dynamical, and spectral deviations, where existing algorithms for audio identification would fail. Furthermore, the combination of various deformation- and fault-tolerance mechanisms allows us to employ standard indexing techniques to obtain an efficient, index-based matching procedure, thus providing an important step towards semantically searching large-scale real-world music collections. Frank Kurth, Meinard Müller |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Bounds and Constructions for Optimal Constant Weight Conflict-Avoiding CodesabstractA conflict-avoiding code (CAC) C of length n with weight k is a family of binary sequences of length n and weight k satisfying Sigma0lestlesn-1xitxj,t+sles lambda for any distinct codewords xj= (xi0,xi1,hellip,xi,n-1) and xj= (xj0, xj1,hellip, xj,n-1) in C and for any integer s, where the subscripts are taken modulo n. A CAC with maximal code size for given n and k is said to be optimal. A CAC has been studied for sending messages correctly through a multiple-access channel. The use of an optimal CAC enables the largest possible number of asynchronous users to transmit information efficiently and reliably. In this paper, the case lambda = 1 is treated, and various direct and recursive constructions of optimal CACs for weight k = 4 and 5 are obtained by providing constructions of CACs for general weight k. In particular, the maximum code size of CACs satisfying certain sufficient conditions is determined through number theoretical and combinatorial approaches. Koji Momihara, Meinard Müller, Junya Satoh, Masakazu Jimbo |
ISIT | 2 |
| 2007 | Constant Weight Conflict-Avoiding CodesabstractA conflict-avoiding code (CAC) C of length n with weight k is a family of binary sequences of length n and weight k satisfying $\sum_{0\le t\le n-1}x_{it}x_{j,t+s}\le \lambda$ for any distinct codewords $x_i=(x_{i0},x_{i1},\ldots,x_{i,n-1})$ and $x_j=(x_{j0},x_{j1},\ldots,x_{j,n-1})$ in C and for any integer s, where the subscripts are taken modulo n. A CAC with maximum code size for given n and k is said to be optimal. A CAC has been studied for sending messages correctly through a multiple-access channel. The use of an optimal CAC enables the largest possible number of potential users to transmit information efficiently and reliably. In this paper, the case $\lambda=1$ is treated, and various direct and recursive constructions of optimal CACs for weight $k=4$ and 5 are obtained by providing constructions of CACs for general weight k. In particular, the maximum code size of CACs satisfying certain sufficient conditions is determined through number theoretical and combinatorial approaches. Koji Momihara, Meinard Müller, Junya Satoh, Masakazu Jimbo |
SIAM J. Discret. Math. | 2 |
| 2006 | An Information Retrieval System for Motion Capture Data
Bastian Demuth, Tido Röder, Meinard Müller, Bernd Eberhardt |
ECIR | 3 |
| 2006 | Enhancing Similarity Matrices for Music Audio AnalysisabstractSimilarity matrices have become an important tool in music audio analysis. However, the quadratic time and space complexity as well as the intricacy of extracting the desired structural information from these matrices are often prohibitive with regard to real-world applications. In this paper, we describe an approach for enhancing the structural properties of similarity matrices based on two concepts: first, we introduce a new class of robust and scalable audio features which absorb local temporal variations. As a second contribution, we then incorporate contextual information into the local similarity measure. The resulting enhancement leads to significant reduction in matrix size and also eases the structure extraction step. As an example, we sketch the application of our techniques to the problems of audio summarization and audio synchronization, obtaining effective and computationally feasible algorithms Meinard Müller, Frank Kurth |
ICASSP (5) | 1 |
| 2005 | Cluttered orderings for the complete bipartite graph
Meinard Müller, Tomoko Adachi, Masakazu Jimbo |
Discret. Appl. Math. | 1 |
| 2005 | Efficient content-based retrieval of motion capture dataabstractThe reuse of human motion capture data to create new, realistic motions by applying morphing and blending techniques has become an important issue in computer animation. This requires the identification and extraction of logically related motions scattered within some data set. Such content-based retrieval of motion capture data, which is the topic of this paper, constitutes a difficult and time-consuming problem due to significant spatio-temporal variations between logically related motions. In our approach, we introduce various kinds of qualitative features describing geometric relations between specified body points of a pose and show how these features induce a time segmentation of motion capture data streams. By incorporating spatio-temporal invariance into the geometric features and adaptive segments, we are able to adopt efficient indexing methods allowing for flexible and efficient content-based retrieval and browsing in huge motion capture databases. Furthermore, we obtain an efficient preprocessing method substantially accelerating the cost-intensive classical dynamic time warping techniques for the time alignment of logically similar motion data streams. We present experimental results on a test data set of more than one million frames, corresponding to 180 minutes of motion. The linearity of our indexing algorithms guarantees the scalability of our results to much larger data sets. Meinard Müller, Tido Röder, Michael Clausen |
ACM Trans. Graph. | 1 |
| 2004 | Erasure-resilient codes from affine spaces
Meinard Müller, Masakazu Jimbo |
Discret. Appl. Math. | 1 |
| 2004 | Generating fast Fourier transforms of solvable groups
Michael Clausen, Meinard Müller |
J. Symb. Comput. | 2 |