EDBT 2026 Demo / reviewers in the wild / expert
Michael Wohlmayr
dblp:05/7252
· DBLP profile ↗
20ranked-venue papers
10as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-authorDatabases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Audio and music processing · 100% | |
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 50% Kernel, tree and ensemble methods · 50% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › music transcription
multipitch tracking |
0.3 | 2 | 2013 | Model-Based Multiple Pitch Tracking Using Factorial HMMs: Model Adaptation and Inference · IEEE Trans. Speech Audio Process. 2013 A Probabilistic Interaction Model for Multipitch Tracking With Factorial Hidden Markov Models · IEEE Trans. Speech Audio Process. 2011 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › bayesian network
bayesian network classifiers |
0.1 | 1 | 2012 | Maximum Margin Bayesian Network Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Machine learning › Kernel, tree and ensemble methods › large margin methods
max-margin parameter learning |
0.1 | 1 | 2012 | Maximum Margin Bayesian Network Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Audio and music processing › source separation › speech separation
single-channel speech separation |
0.1 | 1 | 2011 | Source-Filter-Based Single-Channel Speech Separation Using Pitch Information · IEEE Trans. Speech Audio Process. 2011 |
Audio and music processing › source separation
speech separation |
0.1 | 1 | 2011 | Source-Filter-Based Single-Channel Speech Separation Using Pitch Information · IEEE Trans. Speech Audio Process. 2011 |
Audio and music processing › source separation
single-channel source separation |
0.1 | 2 | 2013 | Model-Based Multiple Pitch Tracking Using Factorial HMMs: Model Adaptation and Inference · IEEE Trans. Speech Audio Process. 2013 A Probabilistic Interaction Model for Multipitch Tracking With Factorial Hidden Markov Models · IEEE Trans. Speech Audio Process. 2011 |
Methods — techniques the papers use, named apart from their topics
observation likelihood pruning · 0.2expectation-maximization · 0.2convex relaxation · 0.1conjugate gradient · 0.1vector quantization · 0.1nonnegative matrix factorization · 0.1multi-pitch estimation · 0.1minimum description length · 0.1loopy max-sum algorithm · 0.1gaussian mixture model · 0.1factorial HMM · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Self-adaption in single-channel source separationabstractSingle-channel source separation (SCSS) usually uses pre-trained source-specific models to separate the sources. These models capture the characteristics of each source and they perform well when matching the test conditions. In this paper, we extend the applicability of SCSS. We develop an EM-like iterative adaption algorithm which is capable to adapt the pre-trained models to the changed characteristics of the specific situation, such as a different acoustic channel introduced by variation in the room acoustics or changed speaker position. The adaption framework requires signal mixtures only, i.e. specific single source signals are not necessary. We consider speech/noise mixtures and we restrict the adaption to the speech model only. Model adaption is empirically evaluated using mixture utterances from the CHiME 2 challenge. We perform experiments using speaker dependent (SD) and speaker independent (SI) models trained on clean or reverberated single speaker utterances. We successfully adapt SI source models trained on clean utterances and achieve almost the same performance level as SD models trained on reverberated utterances. Michael Wohlmayr, Ludwig Mohr, Franz Pernkopf |
INTERSPEECH | 1 |
| 2013 | Model adaptation of factorial HMMS for multipitch trackingabstractFactorial hidden Markov models (FHMMs) are used for tracking the pitch of two interacting speakers [1]. In this statistical approach, the characteristics of each speaker are captured by pre-trained models. Speaker models that match the test conditions well allow for high tracking performance, however the availability of such models is unrealistic. To extend the applicabiliy of the FHMM framework, we develop an EM-like iterative adaptation algorithm which is capable to adapt the model parameters to the specific situation, e.g. acoustic channel, using only speech mixture data. Model adaptation is empirically evaluated using real room recordings of mixture utterances from the GRID corpus. Michael Wohlmayr, Franz Pernkopf |
ICASSP | 1 |
| 2013 | Stochastic margin-based structure learning of Bayesian network classifiersabstractThe margin criterion for parameter learning in graphical models gained significant impact over the last years. We use the maximum margin score for discriminatively optimizing the structure of Bayesian network classifiers. Furthermore, greedy hill-climbing and simulated annealing search heuristics are applied to determine the classifier structures. In the experiments, we demonstrate the advantages of maximum margin optimized Bayesian network structures in terms of classification performance compared to traditionally used discriminative structure learning methods. Stochastic simulated annealing requires less score evaluations than greedy heuristics. Additionally, we compare generative and discriminative parameter learning on both generatively and discriminatively structured Bayesian network classifiers. Margin-optimized Bayesian network classifiers achieve similar classification performance as support vector machines. Moreover, missing feature values during classification can be handled by discriminatively optimized Bayesian network classifiers, a case where purely discriminative classifiers usually require mechanisms to complete unknown feature values in the data first. Franz Pernkopf, Michael Wohlmayr |
Pattern Recognit. | 2 |
| 2013 | Model-Based Multiple Pitch Tracking Using Factorial HMMs: Model Adaptation and InferenceabstractRobustness against noise and interfering audio signals is one of the challenges in speech recognition and audio analysis technology. One avenue to approach this challenge is single-channel multiple-source modeling. Factorial hidden Markov models (FHMMs) are capable of modeling acoustic scenes with multiple sources interacting over time. While these models reach good performance on specific tasks, there are still serious limitations restricting the applicability in many domains. In this paper, we generalize these models and enhance their applicability. In particular, we develop an EM-like iterative adaptation framework which is capable to adapt the model parameters to the specific situation (e.g. actual speakers, gain, acoustic channel, etc.) using only speech mixture data. Currently, source-specific data is required to learn the model. Inference in FHMMs is an essential ingredient for adaptation. We develop efficient approaches based on observation likelihood pruning. Both adaptation and efficient inference are empirically evaluated for the task of multipitch tracking using the GRID corpus. Michael Wohlmayr, Franz Pernkopf |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Maximum Margin Bayesian Network ClassifiersabstractWe present a maximum margin parameter learning algorithm for Bayesian network classifiers using a conjugate gradient (CG) method for optimization. In contrast to previous approaches, we maintain the normalization constraints on the parameters of the Bayesian network during optimization, i.e., the probabilistic interpretation of the model is not lost. This enables us to handle missing features in discriminatively optimized Bayesian networks. In experiments, we compare the classification performance of maximum margin parameter learning to conditional likelihood and maximum likelihood learning approaches. Discriminative parameter learning significantly outperforms generative maximum likelihood estimation for naive Bayes and tree augmented naive Bayes structures on all considered data sets. Furthermore, maximizing the margin dominates the conditional likelihood approach in terms of classification performance in most cases. We provide results for a recently proposed maximum margin optimization approach based on convex relaxation. While the classification results are highly similar, our CG-based optimization is computationally up to orders of magnitude faster. Margin-optimized Bayesian network classifiers achieve classification performance comparable to support vector machines (SVMs) using fewer parameters. Moreover, we show that unanticipated missing feature values during classification can be easily processed by discriminatively optimized Bayesian network classifiers, a case where discriminative classifiers usually require mechanisms to complete unknown feature values in the data first. Franz Pernkopf, Michael Wohlmayr, Sebastian Tschiatschek |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Gain-robust multi-pitch tracking using sparse nonnegative matrix factorizationabstractWhile nonnegative matrix factorization (NMF) has successfully been applied for gain-robust multi-pitch detection, a method to track pitch values over time was not provided. We embed NMF-based pitch detection into a recently proposed pitch-tracking system, based on a factorial hidden Markov model (FHMM). The original system models speech spectra with Gaussian mixture models, which is sensitive to a gain mismatch between training and test data. We therefore combine the advantages of these two approaches and derive a gain-adaptive observation model for the FHMM. As training algorithm we use a modification of ℓ0-sparse NMF, which represents the short-time spectrum with scalable basis vectors. In experiments we show that the new approach significantly increases the gain-robustness of the original tracking system. Robert Peharz, Michael Wohlmayr, Franz Pernkopf |
ICASSP | 2 |
| 2011 | Maximum margin structure learning of Bayesian network classifiersabstractRecently, the margin criterion has been successfully used for parameter optimization in graphical models. We introduce maximum margin based structure learning for Bayesian network classifiers and demonstrate its advantages in terms of classification performance compared to traditionally used discriminative structure learning methods. In particular, we provide empirical results for generative structure learning and two discriminative structure learning approaches on handwritten digit recognition tasks. We show that maximum margin structure learning outperforms other structure learning methods. Furthermore, we present classification results achieved with different bitwidth for representing the parameters of the classifiers. Franz Pernkopf, Michael Wohlmayr, Manfred Mücke |
ICASSP | 2 |
| 2011 | Efficient implementation of probabilistic multi-pitch trackingabstractWe significantly improve the computational efficiency of a probabilistic approach for multiple pitch tracking. This method is based on a factorial hidden Markov model and two alternative interaction models for magnitude and log-magnitude spectra, respectively. The main computational bottleneck comprises the determination of observation likelihoods. However, we show that up to 99.5% of the smallest likelihood values can be discarded at each time frame with out affecting the overall tracking accuracy. For both interaction models, we present a heuristic to efficiently find the largest likelihood values. Experiments on the GRID database show that the proposed methods result in a major speedup without significantly changing tracking accuracy. Michael Wohlmayr, Robert Peharz, Franz Pernkopf |
ICASSP | 1 |
| 2011 | A Pitch Tracking Corpus with Evaluation on Multipitch Tracking ScenarioabstractIn this paper, we introduce a novel pitch tracking database (PTDB) including ground truth signals obtained from a laryngograph. The database, referenced as PTDB-TUG, consists of 2342 phonetically rich sentences taken from the TIMIT corpus. Each sentence was at least recorded once by a male and a female native speaker. In total, the database contains 4720 recordings from 10 male and 10 female speakers. Furthermore, we evaluated two multipitch tracking systems on a subset of speakers to provide a benchmark for further research activities. The database can be downloaded at http://www.spsc.tugraz.at/tools. Gregor Pirker, Michael Wohlmayr, Stefan Petrik, Franz Pernkopf |
INTERSPEECH | 2 |
| 2011 | EM-Based Gain Adaptation for Probabilistic Multipitch TrackingabstractWe introduce an EM algorithm for automatic speaker gain adaptation, and use this approach for probabilistic multipitch tracking. We derive a lower bound on the log-likelihood of the gain parameters and use a fast pruning method to make lower bound optimization efficient. We evaluate the performance of gain adapted multipitch tracking on the GRID database, where 3000 speech mixtures were generated for each mixing level. For gain differences in the range of zero up to 18dB, the proposed method achieves almost the same performance as for the case where the gain is assumed to be known. Michael Wohlmayr, Franz Pernkopf |
INTERSPEECH | 1 |
| 2011 | Source-Filter-Based Single-Channel Speech Separation Using Pitch InformationabstractIn this paper, we investigate the source-filter-based approach for single-channel speech separation. We incorporate source-driven aspects by multi-pitch estimation in the model-driven method. For multi-pitch estimation, the factorial HMM is utilized. For modeling the vocal tract filters either vector quantization (VQ) or non-negative matrix factorization are considered. For both methods, the final combination of the source and filter model results in an utterance dependent model that finally enables speaker independent source separation. The contributions of the paper are the multi-pitch tracker, the gain estimation for the VQ based method which accounts for different mixing levels, and a fast approximation for the likelihood computation. Additionally, a linear relationship between pitch tracking performance and speech separation performance is shown. Michael Stark 0004, Michael Wohlmayr, Franz Pernkopf |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | A Probabilistic Interaction Model for Multipitch Tracking With Factorial Hidden Markov ModelsabstractWe present a simple and efficient feature modeling approach for tracking the pitch of two simultaneously active speakers. We model the spectrogram features of single speakers using Gaussian mixture models in combination with the minimum description length model selection criterion. To obtain a probabilistic representation for the speech mixture spectrogram features of both speakers, we employ the mixture maximization model (MIXMAX) and, as an alternative, a linear interaction model. A factorial hidden Markov model is applied for tracking pitch over time. This statistical model can be used for applications beyond speech, whenever the interaction between individual sources can be represented as MIXMAX or linear model. For tracking, we use the loopy max-sum algorithm, and provide empirical comparisons to exact methods. Furthermore, we discuss a scheduling mechanism of loopy belief propagation for online tracking. We demonstrate experimental results using Mocha-TIMIT as well as data from the speech separation challenge provided by Cooke We show the excellent performance of the proposed method in comparison to a well known multipitch tracking algorithm based on correlogram features. Using speaker-dependent models, the proposed method improves the accuracy of correct speaker assignment, which is important for single-channel speech separation. In particular, we are able to reduce the overall tracking error by 51% relative for the speaker-dependent case. Moreover, we use the estimated pitch trajectories to perform single-channel source separation, and demonstrate the beneficial effect of correct speaker assignment on speech separation performance. Michael Wohlmayr, Michael Stark 0004, Franz Pernkopf |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | A mixture maximization approach to multipitch tracking with factorial hidden Markov modelsabstractWe present a simple and efficient feature modeling approach for tracking the pitch of two speakers speaking simultaneously. We model the spectrogram features of single speakers using Gaussian mixture models in combination with the minimum description length model selection criterion. Furthermore, the mixture maximization (MIXMAX) interaction model is employed to yield a probabilistic representation for the mixture of both speakers. Finally, a factorial hidden Markov model is applied for tracking. We demonstrate experimental results on two databases, and show the excellent performance of the proposed method in comparison to a well known multipitch tracking algorithm based on correlogram features. Michael Wohlmayr, Michael Stark 0004, Franz Pernkopf |
ICASSP | 1 |
| 2010 | Single Channel Speech Separation Using Source-Filter RepresentationabstractWe propose a fully probabilistic model for source-filter based single channel source separation. In particular, we perform separation in a sequential manner, where we estimate the source-driven aspects by a factorial HMM used for multi-pitch estimation. Afterwards, these pitch tracks are combined with the vocal tract filter model to form an utterance dependent model. Additionally, we introduce a gain estimation approach to enable adaptation to arbitrary mixing levels in the speech mixtures. We thoroughly evaluate this system and finally end up in a speaker independent model. Michael Stark 0004, Michael Wohlmayr, Franz Pernkopf |
ICPR | 2 |
| 2010 | Large Margin Learning of Bayesian Classifiers Based on Gaussian Mixture Models
Franz Pernkopf, Michael Wohlmayr |
ECML/PKDD (3) | 2 |
| 2009 | Finite mixture spectrogram modeling for multipitch tracking using a factorial hidden Markov model
Michael Wohlmayr, Franz Pernkopf |
INTERSPEECH | 1 |
| 2009 | On Discriminative Parameter Learning of Bayesian Network Classifiers
Franz Pernkopf, Michael Wohlmayr |
ECML/PKDD (2) | 2 |
| 2008 | Multipitch tracking using a factorial hidden Markov modelabstractIn this paper, we present an approach to track the pitch of two simultaneous speakers. Using a well-known feature extraction method based on the correlogram, we track the resulting data using a factorial hidden Markov model (FHMM). In contrast to the recently developed multipitch determination algorithm [1], which is based on a HMM, we can accurately associate estimated pitch points with their corresponding source speakers. We evalute our approach on the “Mocha-TIMIT” database [2] of speech utterances mixed at 0dB, and compare the results to the multipitch determination algorithm [1] used as a baseline. Experiments show that our FHMM tracker yields good performance for both pitch estimation and correct speaker assignment. Michael Wohlmayr, Franz Pernkopf |
INTERSPEECH | 1 |
| 2007 | Speech-nonspeech discrimination using the information bottleneck method and spectro-temporal modulation indexabstractIn this work, we adopt an information theoretic approach- the Information Bottleneck method- to extract the relevant spectro-temporal modulations for the task of speech / non-speech dis-crimination- non-speech events include music, noise and an-imal vocalizations. A compact representation (a “cluster pro-totype”) is built for each class consisting of the maximally in-formative features with respect to the classification task. We assess the similarity of a sound to each representative cluster using the spectro-temporal modulation index (STMI) adapted to handle the contribution of different frequency bands. A sim-ple threshold check is then used for discriminating speech from non-speech events. Conducted experiments have shown that the proposed method has low complexity and high accuracy of dis-crimination in low SNR conditions compared to recently pro-posed methods for the same task. Index Terms: audio classification, speech discrimination, au-ditory model Maria E. Markaki, Michael Wohlmayr, Yannis Stylianou |
INTERSPEECH | 2 |
| 2007 | Joint position-pitch extraction from multichannel audioabstractRecently, a method for joint extraction of pitch and location information from two-channel recordings has been introduced. This framework offers a new, natural representation of all acoustic sources in the auditory scene, and has potential to be used as front-end in applications such as advanced tracking of multiple speakers in conference rooms. In this paper, we explore basic properties of this method and propose improvements in performance by using circular arrangements of multiple microphones. Index Terms: Acoustic arrays, pitch estimation, source localization Michael Wohlmayr, Marián Képesi |
INTERSPEECH | 1 |