EDBT 2026 Demo / reviewers in the wild / expert
Franz Pernkopf
dblp:97/887
· DBLP profile ↗
130ranked-venue papers
18as first author
25since 2021 · last 2026
0000-0002-6356-3367ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 87 · 13 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 66 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GaPaTTA: Gaussian Entropy-Guided Prompt Placement for Test-Time Adaptation in Semantic SegmentationabstractAbstract Test-time adaptation (TTA) aims to improve model robustness under domain shifts without access to source data–an essential capability for real-world applications such as autonomous driving and robotics. Existing TTA methods for semantic segmentation often rely on stochastic techniques like Monte Carlo dropout or augmentation-averaged predictions to estimate uncertainty or stabilize outputs. However, these approaches typically require multiple forward passes, which are computationally expensive and limit real-time applicability. We propose GaPaTTA, a lightweight and deterministic TTA framework built on SegFormer. Unlike previous methods, GaPaTTA adopts a single forward pass with a traditional augmentation strategy, avoiding repeated inference required by ensemble-based TTA approaches. Key innovations include: (1) Grad-CAM-based global prompt placement identifies the most relevant encoder layers for adaptation; (2) Gaussian entropy-guided local prompt injection selects the top-K most uncertain pixels; (3) Shannon entropy-based filtering suppresses unreliable pseudo-labels; and (4) cross-stage consistency aligns mid- and high-level features for structural coherence. Experiments on ACDC (A-Fog, A-Night, A-Rain, A-Snow), Cityscapes-Foggy (CS-Fog) and Cityscapes-Rainy (CS-Rain) demonstrate that GaPaTTA consistently outperforms previous TTA methods in mean intersection over union (mIoU) while reducing inference time by over 50%. The source code is available at https://github.com/ml4papers/GaPaTTA . Jixiang Lei, Franz Pernkopf |
Mach. Learn. | 2 |
| 2025 | Effective Bayesian Causal Inference via Structural Marginalisation and Autoregressive OrdersabstractThe traditional two-stage approach to causal inference first identifies a \emph{single} causal model (or equivalence class of models), which is then used to answer causal queries. However, this neglects any epistemic model uncertainty. In contrast, \emph{Bayesian} causal inference does incorporate epistemic uncertainty into query estimates via Bayesian marginalisation (posterior averaging) over \emph{all} causal models. While principled, this marginalisation over entire causal models, i.e., both causal structures (graphs) and mechanisms, poses a tremendous computational challenge. In this work, we address this challenge by decomposing structure marginalisation into the marginalisation over (i) causal orders and (ii) directed acyclic graphs (DAGs) given an order. We can marginalise the latter in closed form by limiting the number of parents per variable and utilising Gaussian Processes to model mechanisms. To marginalise over orders, we use a sampling-based approximation, for which we devise a novel auto-regressive distribution over causal orders (ARCO). Our method outperforms state-of-the-art in structure learning on simulated non-linear additive noise benchmarks, and yields competitive results on real-world data. Furthermore, we can accurately infer interventional distributions and average causal effects. Christian Toth, Christian Knoll 0002, Franz Pernkopf, Robert Peharz |
AISTATS | 3 |
| 2025 | Input Uncertainty Attribution by Uncertainty PropagationabstractAttributing uncertainties to the input space elevates the trustworthiness and explainability of machine learning applications. This paper proposes a novel method called Smoothness Constrained Attribution (SCA), which uses the uncertainty propagation mechanism to propagate the output uncertainty back to the input space. This input uncertainty attribution relies solely on test-time data, the trained uncertainty-aware Machine Learning (ML) model, and assumes a smooth input space, resulting in an efficient and simple system. SCA is compared to existing input Uncertainty Attribution Mechanisms (iUCAMs) based on eXplainable Artificial Intelligence (XAI) and an oracle reference using heteroscedastic noise in different synthetic datasets. These evaluations demonstrate the robustness and improvements of SCA compared to existing methods. Benedikt Kantz, Sophie Steger, Clemens Staudinger, Christoph Feilmayr, Johannes Wachlmayr, Alexander Haberl, Stefan Schuster, Franz Pernkopf |
ICASSP | 8 |
| 2025 | Avoiding Domain Drift and Constant Predictions with Diffusion Enhanced Vector-Quantized Autoencoders for Temperature PredictionsabstractAccurate temperature prediction in rotary cement kilns is crucial for process stability and equipment longevity. However, the repeated application of the predicted values creates an error accumulation over the length of the forecast, causing a domain drift of the predictions. This issue is exacerbated for image prediction, as more degrees of freedom lead to a higher sensitivity to small errors as local structures are lost. Using a vector-quantized autoencoder can mitigate the problem as it can map predictions back to the source domain, but it leads to almost constant predictions. Thus, we propose the usage of an additional diffusion model to avoid the local minimum of constant predictions. Our method maintains reliable in-domain predictions, preventing localized temperature peaks and ensuring stable kiln operation. Our experiments show, that the proposed vector-quantized diffusion model (VQ-Diff) can forecast much longer time sequences than reference methods with high accuracy, by being limited to the generation of in-domain images. Nina Lampl, João Machado de Freitas, Alexander Fuchs 0009, Benedikt Brezina, Michael Klitzsch, Franz Pernkopf |
ICASSP | 6 |
| 2025 | Uncertainty prediction for prominence classification with chroma featuresabstractThis paper presents methods for prominence classification in conversational speech. Most existing tools rely on prosodic features extracted at syllable- or phone-level, performing well on read speech. This is not the case for conversational speech, where the quality of automatic segmentation is significantly worse. We introduce entropy-based chroma features, requiring only word-level segmentations. They perform equally well as a random forest classifier with prosodic features (requiring phone-level segmentation), with accuracies in the range of the human inter-rater agreement. We further use Bayesian deep learning to quantify the epistemic and aleatoric uncertainty of the prediction for prosodic and chroma features. Whereas the aleatoric uncertainty is, as expected, consistent with inter-rater agreement and similarly high for both feature sets, the epistemic uncertainty is lower for the classifier based on chroma features, indicating higher classification consistency across the corpus. Julian Linke, Sophie Steger, Philipp Steinwender, Gernot Kubin, Franz Pernkopf, Barbara Schuppler |
ICASSP | 5 |
| 2025 | Two-Level Test-Time Adaptation in Multimodal LearningabstractTest-time adaptation (TTA) aims to adjust the parameters of a pre-trained source model using samples from the target domain, without requiring access to the source data. While recent studies have shown the potential of TTA across various computer vision tasks, most TTA methods are limited to uni-modal adaptation, and the domain shift caused by unimodal data corruption in multimodal tasks is not adequately addressed. Although some recent approaches have reduced cross-modal information discrepancy through modality-sharing modules, the domain adaptation for modality-specific modules has been overlooked. In this paper, we introduce a two-level test-time adaptation method (2LTTA) that accounts for both intra-modal distribution shifts and cross-modal reliability bias in multimodal learning (MML). Unlike conventional TTA methods, which focus primarily on fine-tuning normalization layers, 2LTTA modulates all normalization layers, self-Attention modules of the encoder related to the corrupted modality, and the modality-sharing block. Additionally, we design a two-level objective function that addresses both intra-modal distribution shift and cross-modal reliability bias in the modality fusion block. First, Shannon entropy with sample reweighting is used to mitigate intra-modal distribution shifts caused by data corruption. Second, a diversity-promoting loss is incorporated to reduce cross-modal information discrepancy. Our experiments show that 2LTTA outperforms baseline methods across various datasets. Jixiang Lei, Franz Pernkopf |
IJCNN | 2 |
| 2025 | HR-OTTA: Robust Online Test-Time Adaptation with Hybrid Fine-TuningabstractTest-time adaptation (TTA) aims to adapt a model trained on a source domain to an unlabeled target domain without requiring source data. Traditional TTA assumes that the entire target domain data is available for observation during adaptation. In contrast, online test-time adaptation (OTTA) processes target domain data sequentially in mini-batches, enabling real-time model updates as new data arrives. While recent studies have demonstrated the potential of OTTA for various computer vision tasks, several challenges remain. This paper addresses two key issues: (1) While existing methods adapt different model components (e.g., batch normalization (BN) layers, fully connected (FC) layers, and the whole feature extractor), the most effective strategy is unclear. We propose a hybrid fine-tuning approach that modulates only the shallow convolutional (Conv) layers and all BN layers, enabling a more precise and effective adaptation compared to conventional methods. (2) The objective function plays a crucial role in enhancing prediction accuracy. Unlike previous TTA frameworks that focus solely on entropy minimization, we introduce a robust objective function that combines a diversity-promoting regularization term to balance certainty and diversity, along with a confidence-enhanced loss term that prioritizes high-confidence predictions. Our experiments on various corruption datasets, using both CNN and vision transformer (ViT)-based models, demonstrate that our approach outperforms state-of-the-art methods across a wide range of applications. Our implementation is publicly available.1 Jixiang Lei, Franz Pernkopf |
IJCNN | 2 |
| 2025 | PromptCAL: Entropy-Calibrated and Prompt-Tuned Test-Time Adaptation for Semantic SegmentationabstractTest-time adaptation (TTA) aims to improve the robustness of segmentation models to an unlabeled target domain without requiring access to the source data. While existing TTA methods have achieved promising results on image classification, they often fail to translate effectively to semantic segmentation due to the spatial complexity and fine-grained nature of dense predictions. We propose PromptCAL, a lightweight and effective TTA framework tailored for semantic segmentation, built upon the SegFormer architecture. Our method addresses two central challenges: (1) Which model component to adapt remains underexplored. Using Grad-CAM visualization and sensitivity analysis, we identify Stage 2 of the transformer backbone as the most domain-sensitive and restrict adaptation to this stage. (2) How to identify reliable supervision during adaptation is critical. We introduce a confidence-aware self-training mechanism based on per-pixel entropy filtering to guide pixel selection for model adaptation, ensuring label quality and model transferability. In addition, we incorporate lightweight prompt injection to enhance the adaptability of mid-level features. Our method achieves competitive improvements over the state-of-the-art while maintaining high adaptation efficiency and significantly reducing runtime overhead. Extensive experiments on corrupted semantic segmentation benchmarks, including ACDC (A-fog, A-night, A-rain, and A-snow), Cityscapes-foggy (CS-fog) and Cityscapes-rainy (CS-rain) demonstrate that PromptCAL achieves comparable or superior accuracy to state-of-the-art TTA baselines, while reducing adaptation time by over 50% per domain. This makes it a practical solution for efficient TTA in smart cities and edge-deployed vision systems. The source code is available at https://github.com/ml4papers/PromptCAL. Jixiang Lei, Franz Pernkopf |
SMC | 2 |
| 2025 | Acoustic COVID-19 Detection Using Multiple Instance LearningabstractIn the COVID-19 pandemic, a rigorous testing scheme was crucial. However, tests can be time-consuming and expensive. A machine learning-based diagnostic tool for audio recordings could enable widespread testing at low costs. In order to achieve comparability between such algorithms, the DiCOVA challenge was created. It is based on the Coswara dataset offering the recording categories cough, speech, breath and vowel phonation. Recording durations vary greatly, ranging from one second to over a minute. A base model is pre-trained on random, short time intervals. Subsequently, a Multiple Instance Learning (MIL) model based on self-attention is incorporated to make collective predictions for multiple time segments within each audio recording, taking advantage of longer durations. In order to compete in the fusion category of the DiCOVA challenge, we utilize a linear regression approach among other fusion methods to combine predictions from the most successful models associated with each sound modality. The application of the MIL approach significantly improves generalizability, leading to an AUC ROC score of 86.6% in the fusion category. By incorporating previously unused data, including the sound modality 'sustained vowel phonation' and patient metadata, we were able to significantly improve our previous results reaching a score of 92.2%. Michael Reiter, Franz Pernkopf |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Multiantenna Radar Signal Interference Mitigation Using Complex-Valued Convolutional Neural NetworksabstractModern vehicles increasingly rely on sensors to monitor their environment and to support driver assistance and safety systems. Most vehicles use a variety of different sensors to improve robustness. A vital part of these is the radar sensor. It provides the vehicle not only with location but also with valuable velocity information from surrounding objects. The increasing usage of radar systems in road traffic also causes problems in terms of mutual interference between different radar sensors. This interference leads to broadband disturbances in the signal which must be mitigated to ensure reliable object detection and object angle estimation. In this article, we compare different variants of convolutional neural networks (CNNs) in their ability to mitigate mutual interference for multiantenna radar data. We analyze the potential of using multiantenna data for real-valued CNN (RVCNN) and complex-valued (CVCNN) models, comparing detection, phase reconstruction, and angle estimation performances. Furthermore, we propose a complex-valued CVCNN (CVCNN) architecture using a modified batch normalization method that omits activation scaling. Our experiments show, that using multiantenna data in combination with CVCNNs can greatly improve detection, phase, as well as angle estimation performance and that activation scaling is detrimental to our CVCNN architecture. Alexander Fuchs 0009, Johanna Rock, Máté Tóth, Paul Meissner, Franz Pernkopf |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2024 | Data-Scarce Condition Modeling Requires Model-Based Prior RegularizationabstractIn the metallurgical industry, taking measurements during production can be infeasible or undesired, and only the terminated process can be measured. This poses problems for regression models, as the intermediate target values for a time series are hidden in the accumulated end-of-process measurement. The lack of data quality and quantity also often limits the modeling to linear estimators, as neural networks struggle to converge and/or overfit on scarce noisy data. In this paper, we present a model-based prior for regularized training of neural networks for refractory wear modeling to handle scarce datasets with partially hidden targets. We use an iterative least-squares approach for mutual estimation of intermediate target values, which are then further used as a regularization prior for neural network training. We provide experimental results for refractory wear modeling of two distinct steel processing vessel types. Our results show substantial improvements in wear prediction performance. Nikolaus Mutsam, Alexander Fuchs 0009, Fabio Ziegler, Franz Pernkopf |
ICASSP | 4 |
| 2024 | Resource-Efficient Neural Networks for Embedded SystemsabstractWhile machine learning is traditionally a resource intensive task, embedded systems, autonomous navigation, and the vision of the Internet of Things fuel the interest in resource-efficient approaches. These approaches aim for a carefully chosen trade-off between performance and resource consumption in terms of computation and energy. The development of such approaches is among the major challenges in current machine learning research and key to ensure a smooth transition of machine learning technology from a scientific environment with virtually unlimited computing resources into everyday's applications. In this article, we provide an overview of the current state of the art of machine learning techniques facilitating these real-world requirements. In particular, we focus on resource-efficient inference based on deep neural networks (DNNs), the predominant machine learning models of the past decade. We give a comprehensive overview of the vast literature that can be mainly split into three non-mutually exclusive categories: (i) quantized neural networks, (ii) network pruning, and (iii) structural efficiency. These techniques can be applied during training or as post-processing, and they are widely used to reduce the computational demands in terms of memory footprint, inference speed, and energy efficiency. We also briefly discuss different concepts of embedded hardware for DNNs and their compatibility with machine learning techniques as well as potential for energy and latency reduction. We substantiate our discussion with experiments on well-known benchmark data sets using compression techniques (quantization, pruning) for a set of resource-constrained embedded systems, such as CPUs, GPUs and FPGAs. The obtained results highlight the difficulty of finding good trade-offs between resource efficiency and prediction quality. Wolfgang Roth, Günther Schindler, Bernhard Klein, Robert Peharz, Sebastian Tschiatschek, Holger Fröning, Franz Pernkopf, Zoubin Ghahramani |
J. Mach. Learn. Res. | 7 |
| 2023 | Self-Attention for Enhanced OAMP Detection in MIMO SystemsabstractMultiple-Input Multiple-Output (MIMO) systems are essential for wireless communications. Since classical algorithms for symbol detection in MIMO setups require large computational resources or provide poor results, data-driven algorithms are becoming more popular. Most of the proposed algorithms, however, introduce approximations leading to degraded performance for realistic MIMO systems. In this paper, we introduce a neural-enhanced hybrid model, augmenting the analytic backbone algorithm with state-of-the-art neural network components. In particular, we introduce a self-attention model for the enhancement of the iterative Orthogonal Approximate Message Passing (OAMP)-based decoding algorithm. In our experiments, we show that the proposed model can outperform existing data-driven approaches for OAMP while having improved generalization to other SNR values at limited computational overhead. Alexander Fuchs 0009, Christian Knoll 0002, Nima N. Moghadam, Alexey Pak, Jinliang Huang, Erik Leitinger, Franz Pernkopf |
ICASSP | 7 |
| 2023 | Variational Message Passing-Based Respiratory Motion Estimation and Detection Using Radar SignalsabstractWe present a variational message passing (VMP)-based approach to detect the presence of a person based on their respiratory chest motion using multistatic ultra-wideband (UWB) radar. In the process, the respiratory motion is estimated for contact-free vital sign monitoring. The received signal is modeled as a backscatter channel and the respiratory motion and propagation channels are estimated using VMP. We use the evidence lower bound (ELBO) to approximate the model evidence for the detection. Numerical analyses and measurements demonstrate that the proposed method leads to a significant improvement in the detection performance compared to a fast Fourier transform (FFT)-based detector or an estimator-correlator in low-signal-to-noise ratio (SNR) conditions, since the multipath components (MPCs) are better incorporated into the detection procedure. Specifically, the proposed method has a detection probability of 0.95 at −20dB SNR, while the estimator-correlator and FFT-based detector have 0.32 and 0.05, respectively. Jakob Möderl, Erik Leitinger, Franz Pernkopf, Klaus Witrisal |
ICASSP | 3 |
| 2023 | Self-Guided Belief Propagation - A Homotopy Continuation MethodabstractBelief propagation (BP) is a popular method for performing probabilistic inference on graphical models. In this work, we enhance BP and propose self-guided belief propagation (SBP) that incorporates the pairwise potentials only gradually. This homotopy continuation method converges to a unique solution and increases the accuracy without increasing the computational burden. We provide a formal analysis to demonstrate that SBP finds the global optimum of the Bethe approximation for attractive models where all variables favor the same state. Moreover, we apply SBP to various graphs with random potentials and empirically show that: (i) SBP is superior in terms of accuracy whenever BP converges, and (ii) SBP obtains a unique, stable, and accurate solution whenever BP does not converge. Christian Knoll 0002, Adrian Weller, Franz Pernkopf |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Example or Prototype? Learning Concept-Based Explanations in Time-Series
Christoph Obermair, Alexander Fuchs 0009, Franz Pernkopf, Lukas Felsberger, Andrea Apollonio, Daniel Wollmann |
ACML | 3 |
| 2022 | End-to-End Keyword Spotting Using Neural Architecture Search and QuantizationabstractThis paper introduces neural architecture search (NAS) for the automatic discovery of end-to-end keyword spotting (KWS) models for limited resource environments. We employ a differentiable NAS approach to optimize the structure of convolutional neural networks (CNNs) operating on raw audio waveforms. After a suitable KWS model is found with NAS, we conduct quantization of weights and activations to reduce the memory footprint. We conduct extensive experiments on the Google speech commands dataset. In particular, we compare our end-to-end models to mel-frequency cepstral coefficient (MFCC) based CNNs. For quantization, we compare fixed bit-width quantization and trained bit-width quantization. Using NAS only, we were able to obtain a highly efficient model with an accuracy of 95.55% using 75.7k parameters and 13.6M operations. Using trained bit-width quantization, the same model achieves a test accuracy of 93.76% while using on average only 2.91 bits per activation and 2.51 bits per weight. David Peter, Wolfgang Roth, Franz Pernkopf |
ICASSP | 3 |
| 2022 | Homophone Disambiguation Profits from Durational Information
Barbara Schuppler, Emil Berger, Xenia Kogler, Franz Pernkopf |
INTERSPEECH | 4 |
| 2022 | Active Bayesian Causal InferenceabstractCausal discovery and causal reasoning are classically treated as separate and consecutive tasks: one first infers the causal graph, and then uses it to estimate causal effects of interventions. However, such a two-stage approach is uneconomical, especially in terms of actively collected interventional data, since the causal query of interest may not require a fully-specified causal model. From a Bayesian perspective, it is also unnatural, since a causal query (e.g., the causal graph or some causal effect) can be viewed as a latent quantity subject to posterior inference—quantities that are not of direct interest ought to be marginalized out in this process, thus contributing to our overall uncertainty. In this work, we propose Active Bayesian Causal Inference (ABCI), a fully-Bayesian active learning framework for integrated causal discovery and reasoning, i.e., for jointly inferring a posterior over causal models and queries of interest. In our approach to ABCI, we focus on the class of causally-sufficient nonlinear additive Gaussian noise models, which we model using Gaussian processes. To capture the space of causal graphs, we use a continuous latent graph representation, allowing our approach to scale to practically relevant problem sizes. We sequentially design experiments that are maximally informative about our target causal query, collect the corresponding interventional data, update our beliefs, and repeat. Through simulations, we demonstrate that our approach is more data-efficient than existing methods that only focus on learning the full causal graph. This allows us to accurately learn downstream causal queries from fewer samples, while providing well-calibrated uncertainty estimates of the quantities of interest. Christian Toth, Lars Lorch, Christian Knoll 0002, Andreas Krause 0001, Franz Pernkopf, Robert Peharz, Julius von Kügelgen |
NeurIPS | 5 |
| 2022 | Fixing the Bethe approximation: How structural modifications in a graph improve belief propagationabstractBelief propagation is an iterative method for inference in probabilistic graphical models. Its well-known relationship to a classical concept from statistical physics, the Bethe free energy, puts it on a solid theoretical foundation. If belief propagation fails to approximate the marginals, then this is often due to a failure of the Bethe approximation. In this work, we show how modifications in a graphical model can be a great remedy for fixing the Bethe approximation. Specifically, we analyze how the removal of edges influences and improves belief propagation, and demonstrate that this positive effect is particularly distinct for dense graphs. Harald Leisenberger, Franz Pernkopf, Christian Knoll 0002 |
UAI | 2 |
| 2022 | Blind Speech Separation and Dereverberation using neural beamforming
Lukas Pfeifenberger, Franz Pernkopf |
Speech Commun. | 2 |
| 2021 | Autonomous Robot for Measuring Room Impulse Responses
Stefan Fragner, Tobias Topar, Maximilian Giller, Lukas Pfeifenberger, Franz Pernkopf |
Interspeech | 5 |
| 2021 | Acoustic Echo Cancellation with Cross-Domain Learning
Lukas Pfeifenberger, Matthias Zöhrer, Franz Pernkopf |
Interspeech | 3 |
| 2021 | Convergence behavior of belief propagation: estimating regions of attraction via Lyapunov functionsabstractIn this work, we estimate the regions of attraction for belief propagation. This extends existing stability analysis and provides initial message values for which belief propagation is guaranteed to converge. Our approach utilizes the theory of Lyapunov functions that, however, does not readily yield useful regions of attraction. Therefore, we utilize polynomial sum-of-squares relaxations and provide an algorithm that computes valid Lyapunov functions. This admits a novel way of studying the solution space of belief propagation. Finally, we apply our approach to small-scale models and discuss the effect of the potentials on the regions of attraction. Harald Leisenberger, Christian Knoll 0002, Richard Seeber, Franz Pernkopf |
UAI | 4 |
| 2021 | Synthesis and Analysis-By-Synthesis of Modulated Diplophonic Glottal Area WaveformsabstractDiplophonia is a type of disordered voice in which two simultaneous pitches are perceived. Most commonly in diplophonic voices, the vocal folds are divided into two parts that vibrate at different frequencies. The glottal area is the projected area of the space between the vocal folds. The glottal area in time is referred to as the glottal area waveform (GAW). The GAW is modeled for diplophonic voice by superimposing two partial GAWs (pGAWs) that are trains of single-peak pulses with different pulse frequencies, i.e., fundamental frequencies ($f_o$s). In current kinematic models of diplophonic vocal fold vibration, the pGAWs are assumed to be quasiperiodic. This assumption is mitigated here by modulating pulse-to-pulse cycle length and amplitude. Both random and deterministic modulations are considered. Deterministic modulations depend on the difference of the pGAWs' instantaneous phases. Model GAWs are fitted to input GAWs using an analysis-by-synthesis approach which we refer to as `modulated pulse trains decomposition' (MPD). MPD is shown to be applicable to diplophonic as well as to nondiplophonic types of dysphonia, which include multi-pulse patterns, random timing behaviours, and chaos. It is mostly robust against modulations but degraded by large random modulations. MPD is compared to a deep autoencoder neural network, and the WaveGlow neural network. In terms of time-domain fitting errors, MPD outperforms the other two approaches unless random modulations are large. MPD outperforms the best of the other two approaches by up to approximately 5 dB. For large random modulations, the deep autoencoder network achieves the smallest fitting errors. In terms of magnitude spectrum fitting errors, WaveGlow is superior except for natural input GAWs containing only nondiplophonic types of dysphonia. Also pulse timing errors are shown to be advantageous for MPD. Philipp Aichinger, Franz Pernkopf |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Deep Structured Mixtures of Gaussian ProcessesabstractGaussian Processes (GPs) are powerful non-parametric Bayesian regression models that allow exact posterior inference, but exhibit high computational and memory costs. In order to improve scalability of GPs, approximate posterior inference is frequently employed, where a prominent class of approximation techniques is based on local GP experts. However, local-expert techniques proposed so far are either not well-principled, come with limited approximation guarantees, or lead to intractable models. In this paper, we introduce deep structured mixtures of GP experts, a stochastic process model which i) allows exact posterior inference, ii) has attractive computational and memory costs, and iii) when used as GP approximation, captures predictive uncertainties consistently better than previous expert-based approximations. In a variety of experiments, we show that deep structured mixtures have a low approximation error and often perform competitive or outperform prior work. Martin Trapp 0001, Robert Peharz, Franz Pernkopf, Carl E. Rasmussen |
AISTATS | 3 |
| 2020 | Towards Real-Time Single-Channel Singing-Voice Separation with Pruned Multi-Scaled DensenetsabstractModern musical source separation systems based on deep neural networks reach unprecedented levels of separation quality. However, harnessing the power of these large-scale models in typical audio production environments, which frequently offer only limited computing resources while demanding real-time processing, remains challenging. We extend the multi-scaled DenseNet in several aspects to facilitate real-time source separation scenarios. Specifically, we reduce the computational requirements by inferring Mel-scaled masks and decrease the model size via effective use of bottleneck layers, while improving performance using a deep clustering objective. In addition, we are able to further increase the model efficiency by applying parameterized structured pruning of convolutional weights without any significant impact on the separation performance. We significantly reduce the model size and increase the computational efficiency by a factor of 1.6 and 4.3, respectively, while maintaining the separation performance. Günther Schindler, Christian Schörkhuber, Wolfgang Roth, Franz Pernkopf, Holger Fröning |
ICASSP | 5 |
| 2020 | Acoustic Scene Classification for Mismatched Recording Devices Using Heated-Up Softmax and Spectrum CorrectionabstractDeep neural networks (DNNs) are successful in applications with matching inference and training distributions. In realworld scenarios, DNNs have to cope with truly new data samples during inference, potentially coming from a shifted data distribution. This usually causes a drop in performance. Acoustic scene classification (ASC) with different recording devices is one of this situation. Furthermore, an imbalance in quality and amount of data recorded by different devices causes severe challenges. In this paper, we introduce two calibration methods to tackle these challenges. In particular, we applied scaling of the features to deal with varying frequency response of the recording devices. Furthermore, to account for the shifted data distribution, a heated-up softmax is embedded to calibrate the predictions of the model. We use robust and resource-efficient models, and show the efficiency of heated-up softmax. Our ASC system reaches state-of-the-art performance on the development set of DCASE challenge 2019 task 1B with only ~70K parameters. It achieves 70.1% average classification accuracy for device B and device C. It performs on par with the best single model system of the DCASE 2019 challenge and outperforms the baseline system by 28.7% (absolute). Truc Nguyen, Franz Pernkopf, Michal Kosmider |
ICASSP | 2 |
| 2020 | Resource-Efficient DNNs for Keyword Spotting using Neural Architecture Search and QuantizationabstractThis paper introduces neural architecture search (NAS) for the automatic discovery of small models for keyword spotting (KWS) in limited resource environments. We employ a differentiable NAS approach to optimize the structure of convolutional neural networks (CNNs) to maximize the classification accuracy while minimizing the number of operations per inference. Using NAS only, we were able to obtain a highly efficient model with 95.4% accuracy on the Google speech commands dataset with 494.8 kB of memory usage and 19.6 million operations. Additionally, weight quantization is used to reduce the memory consumption even further. We show that weight quantization to low bit-widths (e.g. 1 bit) can be used without substantial loss in accuracy. By increasing the number of input features from 10 MFCC to 20 MFCC we were able to increase the accuracy to 96.3% at 340.1 kB of memory usage and 27.1 million operations. David Peter, Wolfgang Roth, Franz Pernkopf |
ICPR | 3 |
| 2020 | On Resource-Efficient Bayesian Network Classifiers and Deep Neural NetworksabstractWe present two methods to reduce the complexity of Bayesian network (BN) classifiers. First, we introduce quantization-aware training using the straight-through gradient estimator to quantize the parameters of BNs to few bits. Second, we extend a recently proposed differentiable tree-augmented naive Bayes (TAN) structure learning approach by also considering the model size. Both methods are motivated by recent developments in the deep learning community, and they provide effective means to trade off between model size and prediction accuracy, which is demonstrated in extensive experiments. Furthermore, we contrast quantized BN classifiers with quantized deep neural networks (DNNs) for small-scale scenarios which have hardly been investigated in the literature. We show Pareto optimal models with respect to model size, number of operations, and test error and find that both model classes are viable options. Wolfgang Roth, Franz Pernkopf, Günther Schindler, Holger Fröning |
ICPR | 2 |
| 2020 | Nonlinear Residual Echo Suppression Using a Recurrent Neural Network
Lukas Pfeifenberger, Franz Pernkopf |
INTERSPEECH | 2 |
| 2020 | Bayesian Neural Networks with Weight Sharing Using Dirichlet ProcessesabstractWe extend feed-forward neural networks with a Dirichlet process prior over the weight distribution. This enforces a sharing on the network weights, which can reduce the overall number of parameters drastically. We alternately sample from the posterior of the weights and the posterior of assignments of network connections to the weights. This results in a weight sharing that is adopted to the given data. In order to make the procedure feasible, we present several techniques to reduce the computational burden. Experiments show that our approach mostly outperforms models with random weight sharing. Our model is capable of reducing the memory footprint substantially while maintaining a good performance compared to neural networks without weight sharing. Wolfgang Roth, Franz Pernkopf |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Complex Signal Denoising and Interference Mitigation for Automotive Radar Using Convolutional Neural Networks
Johanna Rock, Máté Tóth, Elmar Messner, Paul Meissner, Franz Pernkopf |
FUSION | 5 |
| 2019 | Deep Complex-valued Neural BeamformersabstractWe propose a complex-valued deep neural network (cDNN) for speech enhancement and source separation. While existing end-to-end systems use complex-valued gradients to pass the training error to a real-valued DNN used for gain mask estimation, we use the full potential of complex-valued LSTMs, MLPs and activation functions to estimate complex-valued beamforming weights directly from complex-valued microphone array data. By doing so, our cDNN is able to locate and track different moving sources by exploiting the phase information in the data. In our experiments, we use a typical living room environment, mixtures of the WallStreet Journal corpus, and YouTube noise. We compare our cDNN against the BeamformIt toolkit as a baseline, and a mask-based beamformer as a state-of-the-art reference system. We observed a significant improvement in terms of PESQ, STOI and WER. Lukas Pfeifenberger, Matthias Zöhrer, Franz Pernkopf |
ICASSP | 3 |
| 2019 | Acoustic Scene Classification with Mismatched Recording Devices Using Mixture of Experts LayerabstractRecently, a mismatch in acoustic conditions such as a temporal recording gap as well as different recording devices for the development and the evaluation data has been considered in Acoustic Scene Classification (ASC). This brings ASC closer to real-world conditions. In this paper, we address ASC with mismatching recording devices. This has been introduced as task 1B of the DCASE 2018 challenge. We proposed a flexible and robust model that uses a mixture of experts (MoE) layer replacing the fully connected dense layer such that each expert can adapt to the specific domains of the data. Furthermore, we observe different Convolutional Neural Network (CNN) models as well as the number of the experts of the MoE dense layer using log mel features. In addition, we perform mixup data augmentation to enhance the robustness of our models. In experiments, the classification performance is 66.1% using 15 experts in the MoE dense layer with approximately 2M parameters. This outperforms the best model of task 1B of the DCASE 2018 challenge by 2.5% (absolute). This model uses an ensemble selection of 12 individual models with ~12M parameters. Truc Nguyen, Franz Pernkopf |
ICME | 2 |
| 2019 | Recurrent Dilated DenseNets for a Time-Series Segmentation TaskabstractEfficient real-time segmentation and classification of time-series data is key in many applications, including sound and measurement analysis. We propose an efficient convolutional recurrent neural network (CRNN) architecture that is able to deliver improved segmentation performance at lower computational cost than plain RNN methods. We develop a CNN architecture, using dilated DenseNet-like kernels and implement it within the proposed CRNN architecture. For the task of online wafer-edge measurement analysis, we compare our proposed methods to standard RNN methods, such as Long Short Term Memory (LSTM) and Gated Recurrent Units (GRUs). We focus on small models with a low computational complexity, in order to run our model on an embedded device. We show that frame-based methods generally perform better than RNNs in our segmentation task and that our proposed recurrent dilated DenseNet achieves a substantial improvement of over 1.1 % F1-score compared to other frame-based methods. Alexander Fuchs 0009, Robin Priewald, Franz Pernkopf |
ICMLA | 3 |
| 2019 | Acoustic Scene Classification Using Deep Mixtures of Pre-trained Convolutional Neural NetworksabstractWe propose a heterogeneous system of Deep Mixture of Experts (DMoEs) models using different Convolutional Neural Networks (CNNs) for acoustic scene classification (ASC). Each DMoEs module is a mixture of different parallel CNN structures weighted by a gating network. All CNNs use the same input data. The CNN architectures play the role of experts extracting a variety of features. The experts are pre-trained, and kept fixed (frozen) for the DMoEs model. The DMoEs is post-trained by optimizing weights of the gating network, which estimates the contribution of the experts in the mixture. In order to enhance the performance, we use an ensemble of three DMoEs modules each with different pairs of inputs and individual CNN models. The input pairs are spectrogram combinations of binaural audio and mono audio as well as their pre-processed variations using harmonic-percussive source separation (HPSS) and nearest neighbor filters (NNFs). The classification result of the proposed system is 72.1% improving the baseline by around 12% (absolute) on the development data of DCASE 2018 challenge task 1A. Truc Nguyen, Alexander Fuchs 0009, Franz Pernkopf |
ICMLA | 3 |
| 2019 | Acoustic Scene Classification with Mismatched Devices Using CliqueNets and Mixup Data Augmentation
Truc Nguyen, Franz Pernkopf |
INTERSPEECH | 2 |
| 2019 | Bayesian Learning of Sum-Product NetworksabstractSum-product networks (SPNs) are flexible density estimators and have received significant attention due to their attractive inference properties. While parameter learning in SPNs is well developed, structure learning leaves something to be desired: Even though there is a plethora of SPN structure learners, most of them are somewhat ad-hoc and based on intuition rather than a clear learning principle. In this paper, we introduce a well-principled Bayesian framework for SPN structure learning. First, we decompose the problem into i) laying out a computational graph, and ii) learning the so-called scope function over the graph. The first is rather unproblematic and akin to neural network architecture validation. The second represents the effective structure of the SPN and needs to respect the usual structural constraints in SPN, i.e. completeness and decomposability. While representing and learning the scope function is somewhat involved in general, in this paper, we propose a natural parametrisation for an important and widely used special case of SPNs. These structural parameters are incorporated into a Bayesian model, such that simultaneous structure and parameter learning is cast into monolithic Bayesian posterior inference. In various experiments, our Bayesian SPNs often improve test likelihoods over greedy SPN learners. Further, since the Bayesian framework protects against overfitting, we can evaluate hyper-parameters directly on the Bayesian model score, waiving the need for a separate validation set, which is especially beneficial in low data regimes. Bayesian SPNs can be applied to heterogeneous domains and can easily be extended to nonparametric formulations. Moreover, our Bayesian approach is the first, which consistently and robustly learns SPN structures under missing data. Martin Trapp 0001, Robert Peharz, Franz Pernkopf, Zoubin Ghahramani |
NeurIPS | 4 |
| 2019 | Training Discrete-Valued Neural Networks with Sign Activations Using Weight Distributions
Wolfgang Roth, Günther Schindler, Holger Fröning, Franz Pernkopf |
ECML/PKDD (2) | 4 |
| 2019 | Learning a Behavior Model of Hybrid Systems Through Combining Model-Based Testing and Machine Learning
Bernhard K. Aichernig, Roderick Bloem, Masoud Ebrahimi 0002, Martin Horn, Franz Pernkopf, Wolfgang Roth, Astrid Rupp, Martin Tappler, Markus Tranninger |
ICTSS | 5 |
| 2019 | Belief Propagation: Accurate Marginals or Accurate Partition Function - Where is the Difference?
Christian Knoll 0002, Franz Pernkopf |
UAI | 2 |
| 2019 | Eigenvector-Based Speech Mask Estimation for Multi-Channel Speech EnhancementabstractWe present the Eigennet architecture for estimating a gain mask from noisy, multi-channel microphone observations. While existing mask estimators use magnitude features, our system also exploits the spatial information embedded in the phase of the data. The mask is used to obtain the Minimum Variance Distortionless Response (MVDR) and Generalized Eigenvalue (GEV) beamformers. We also derive the Phase Aware Normalization (PAN) postfilter, which corrects both magnitude and phase distortions caused by the GEV. Further, we demonstrate the properties of our eigenvector features, and compare their performance with three state-of-the-art reference systems. We report their performance in terms of SNR improvement and Word Error Rate (WER) using Google and Kaldi Speech-to-Text API. Experiments are performed on the WSJ0 and CHiME4 corpora, where competitive performance in both WER and SNR is achieved. Lukas Pfeifenberger, Matthias Zöhrer, Franz Pernkopf |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Resource Efficient Deep Eigenvector BeamformingabstractWe propose binary neural networks (BNN s) for acoustic beamforming. This makes the speech enhancement approach resource efficient and applicable for embedded applications. Using CHiME4 data, we use BNN s to estimate the speech presence probability mask for GEV-PAN beamformers. By doing so, we achieve audio quality and ASR scores on par to single-precision deep neural networks (DNNs), while the computational requirements and the memory footprint are significantly reduced. Matthias Zöhrer, Lukas Pfeifenberger, Günther Schindler, Holger Fröning, Franz Pernkopf |
ICASSP | 5 |
| 2018 | Towards Efficient Forward Propagation on Resource-Constrained Systems
Günther Schindler, Matthias Zöhrer, Franz Pernkopf, Holger Fröning |
ECML/PKDD (1) | 3 |
| 2018 | Fixed Points of Belief Propagation - An Analysis via Polynomial Homotopy ContinuationabstractBelief propagation (BP) is an iterative method to perform approximate inference on arbitrary graphical models. Whether BP converges and if the solution is a unique fixed point depends on both the structure and the parametrization of the model. To understand this dependence it is interesting to find all fixed points. In this work, we formulate a set of polynomial equations, the solutions of which correspond to BP fixed points. To solve such a nonlinear system we present the numerical polynomial-homotopy-continuation (NPHC) method. Experiments on binary Ising models and on error-correcting codes show how our method is capable of obtaining all BP fixed points. On Ising models with fixed parameters we show how the structure influences both the number of fixed points and the convergence properties. We further asses the accuracy of the marginals and weighted combinations thereof. Weighting marginals with their respective partition function increases the accuracy in all experiments. Contrary to the conjecture that uniqueness of BP fixed points implies convergence, we find graphs for which BP fails to converge, even though a unique fixed point exists. Moreover, we show that this fixed point gives a good approximation, and the NPHC method is able to obtain this fixed point. Christian Knoll 0002, Dhagash Mehta, Tianran Chen, Franz Pernkopf |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Hybrid generative-discriminative training of Gaussian mixture models
Wolfgang Roth, Robert Peharz, Sebastian Tschiatschek, Franz Pernkopf |
Pattern Recognit. Lett. | 4 |
| 2018 | Tracking of Multiple Fundamental Frequencies in Diplophonic VoicesabstractDiplophonia is a type of pathological voice in which two fundamental frequencies (fo) are present simultaneously. Specialized audio analyzers that can handle up to two fos in diplophonic voices are in their infancy. We propose the tracking of up to two fos in diplophonic voices by audio waveform modeling (AWM), which involves obtaining candidates by repetitive execution of the Viterbi algorithm, followed by waveform Fourier synthesis, and heuristic candidate selection with majority voting. Our approach is evaluated with reference fo-tracks obtained from laryngeal highspeed videos of 29 sustained phonations and compared to state-of-the-art tracking algorithms for multiple fos. An accurate and a fast variant of our algorithm are tested. The median error rate of the accurate variant is 6.52%, whereas the most accurate benchmark achieves 11.11%. The fast variant is more than twice as fast as the fastest relevant benchmark, and the median error rate is 9.52%. Furthermore, illustrative results of connected speech analysis are reported. Our approach may help to improve detection and analysis of diplophonia in clinical research and practice, as well as to advance synthesis of disordered voices. Philipp Aichinger, Martin Hagmüller, Berit Schneider-Stickler, Jean Schoentgen, Franz Pernkopf |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Respiratory airflow estimation from lung sounds based on regressionabstractThe aim of this work is the estimation of respiratory flow from lung sound recordings, i.e. acoustic airflow estimation. With a 16-channel lung sound recording device, we simultaneously record the respiratory flow and the lung sounds on the posterior chest from six lung-healthy subjects in supine position. For the recordings of four selected sensor positions, we extract linear frequency cepstral coefficient (LFCC) features and map these on the airflow signal. We use multivariate polynomial regression to fit the features to the airflow signal. Compared to most of the previous approaches, the proposed method uses lung sounds instead of trachea sounds. Furthermore, our method masters the estimation of the airflow without prior knowledge of the respiratory phase, i.e. no additional algorithm for phase detection is required. Another benefit is the avoidance of time-consuming calibration. In experiments, we evaluate the proposed method for various selections of sensor positions in terms of mean squared error (MSE) between estimated and actual airflow. Moreover, we show the accuracy of the method regarding a frame-based breathing-phase detection. Elmar Messner, Martin Hagmüller, Paul Swatek, Freyja-Maria Smolle-Jüttner, Franz Pernkopf |
ICASSP | 5 |
| 2017 | DNN-based speech mask estimation for eigenvector beamformingabstractIn this paper, we present an optimal multi-channel Wiener filter, which consists of an eigenvector beamformer and a single-channel postfilter. We show that both components solely depend on a speech presence probability, which we learn using a deep neural network, consisting of a deep autoencoder and a softmax regression layer. To prevent the DNN from learning specific speaker and noise types, we do not use the signal energy as input feature, but rather the cosine distance between the dominant eigenvectors of consecutive frames of the power spectral density of the noisy speech signal. We compare our system against the BeamformIt toolkit, and state-of-the-art approaches such as the front-end of the best system of the CHiME3 challenge. We show that our system yields superior results, both in terms of perceptual speech quality and classification error. Lukas Pfeifenberger, Matthias Zöhrer, Franz Pernkopf |
ICASSP | 3 |
| 2017 | Eigenvector-Based Speech Mask Estimation Using Logistic Regression
Lukas Pfeifenberger, Matthias Zöhrer, Franz Pernkopf |
INTERSPEECH | 3 |
| 2017 | Frame and Segment Level Recurrent Neural Networks for Phone ClassificationabstractWe introduce a simple and efficient frame and segment levelRNN model (FS-RNN) for phone classification. It processesthe input atframe levelandsegment levelby bidirectional gatedRNNs. This type of processing is important to exploit the(temporal) information more effectively compared to(i)mod-els which solely process the input at frame level and(ii)mod-els which process the input on segment level using features ob-tained by heuristic aggregation of frame level features. Further-more, we incorporated the activations of the last hidden layerof the FS-RNN as an additional feature type in a neural higher-order CRF (NHO-CRF). In experiments, we demonstrated ex-cellent performance on the TIMIT phone classification task, re-porting a performance of13.8%phone error rate for the FS-RNN model and11.9%when combined with the NHO-CRF. Inboth cases we significantly exceeded the state-of-the-art perfor-mance. Martin Ratajczak, Sebastian Tschiatschek, Franz Pernkopf |
INTERSPEECH | 3 |
| 2017 | Virtual Adversarial Training and Data Augmentation for Acoustic Event Detection with Gated Recurrent Neural Networks
Matthias Zöhrer, Franz Pernkopf |
INTERSPEECH | 2 |
| 2017 | On Loopy Belief Propagation - Local Stability Analysis for Non-Vanishing Fields
Christian Knoll 0002, Franz Pernkopf |
UAI | 2 |
| 2017 | Safe Semi-Supervised Learning of Sum-Product Networks
Martin Trapp 0001, Tamas Madl, Robert Peharz, Franz Pernkopf, Robert Trappl |
UAI | 4 |
| 2017 | On the Latent Variable Interpretation in Sum-Product NetworksabstractOne of the central themes in Sum-Product networks (SPNs) is the interpretation of sum nodes as marginalized latent variables (LVs). This interpretation yields an increased syntactic or semantic structure, allows the application of the EM algorithm and to efficiently perform MPE inference. In literature, the LV interpretation was justified by explicitly introducing the indicator variables corresponding to the LVs' states. However, as pointed out in this paper, this approach is in conflict with the completeness condition in SPNs and does not fully specify the probabilistic model. We propose a remedy for this problem by modifying the original approach for introducing the LVs, which we call SPN augmentation. We discuss conditional independencies in augmented SPNs, formally establish the probabilistic interpretation of the sum-weights and give an interpretation of augmented SPNs as Bayesian networks. Based on these results, we find a sound derivation of the EM algorithm for SPNs. Furthermore, the Viterbi-style algorithm for MPE proposed in literature was never proven to be correct. We show that this is indeed a correct algorithm, when applied to selective SPNs, and in particular when applied to augmented SPNs. Our theoretical results are confirmed in experiments on synthetic data and 103 real-world datasets. Robert Peharz, Robert Gens, Franz Pernkopf, Pedro M. Domingos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Phase-Aware Signal Processing for Automatic Speech Recognition
Johannes Fahringer, Tobias Schrank, Johannes Stahl 0003, Pejman Mowlaee, Franz Pernkopf |
INTERSPEECH | 5 |
| 2016 | Manual versus Automated: The Challenging Routine of Infant Vocalisation Segmentation in Home Videos to Study Neuro(mal)developmentabstractIn recent years, voice activity detection has been a highly researched field, due to its importance as input stage in many real-world applications. Automated detection of vocalisations in the very first year of life is still a stepchild of this field. On our quest defining acoustic parameters in pre-linguistic vocalisations as markers for neuro(mal)development, we are confronted with the challenge of manually segmenting and annotating hours of variable quality home video material for sequences of infant voice/vocalisations. While in total our corpus comprises video footage of typically developing infants and infants with various neurodevelopmental disorders of more than a year running time, only a small proportion has been processed so far. This calls for automated assistance tools for detecting and/or segmenting infant utterances from real-live video recordings. In this paper, we investigated several approaches of infant voice detection and segmentation, including a rule-based voice activity detector, hidden Markov models with Gaussian mixture observation models, support vector machines, and random forests. Results indicate that the applied methods could be well applied in a semi-automated retrieval of infant utterances from highly non-standardised footage. At the same time, our results show that, a fully automated approach for this problem is yet to come. Florian B. Pokorny, Robert Peharz, Wolfgang Roth, Matthias Zöhrer, Franz Pernkopf, Peter B. Marschik, Björn W. Schuller |
INTERSPEECH | 5 |
| 2016 | Virtual Adversarial Training Applied to Neural Higher-Order Factors for Phone ClassificationabstractWe explore virtual adversarial training (VAT) applied to neu-ral higher-order conditional random fields for sequence label-ing. VAT is a recently introduced regularization method pro-moting local distributional smoothness: It counteracts the prob-lem that predictions of many state-of-the-art classifiers are un-stable to adversarial perturbations. Unlike random noise, ad-versarial perturbations are minimal and bounded perturbationsthat flip the predicted label. We utilize VAT to regularize neuralhigher-order factors in conditional random fields. These fac-tors are for example important for phone classification wherephone representations strongly depend on the context phones.However, without using VAT for regularization, the use of suchfactors was limited as they were prone to overfitting. In exten-sive experiments, we successfully apply VAT to improve per-formance on the TIMIT phone classification task. In particular,we achieve a phone error rate of13.0%, exceeding the state-of-the-art performance by a wide margin.Index Terms: Virtual adversarial training, local distributionalsmoothing, deep higher-order factors, neural higher-order con-ditional random field, phone classificatio. Martin Ratajczak, Sebastian Tschiatschek, Franz Pernkopf |
INTERSPEECH | 3 |
| 2016 | Maximum margin hidden Markov models for sequence classification
Nikolaus Mutsam, Franz Pernkopf |
Pattern Recognit. Lett. | 2 |
| 2015 | Detection of negative emotions in speech signals using bags-of-audio-wordsabstractBoosted by a wide potential application spectrum, emotional speech recognition, i.e., the automatic computer-aided identification of human emotional states based on speech signals, currently describes a popular field of research. However, a variety of studies especially concentrating on the recognition of negative emotions often neglected the specific requirements of real-world scenarios, for example, robustness, real-time capability, and realistic speech corpora. Motivated by these facts, a robust, low-complex classification system for the detection of negative emotions in speech signals was implemented on the basis of a spontaneous, strongly emotionally colored speech corpus. Therefore, an innovative approach in the field of emotion recognition was applied as the core of the system - the bag-of-words approach that is originally known from text and image document retrieval applications. Thorough performance evaluations were carried out and a promising recognition accuracy of 65.6 % for the 2-class paradigm negative versus non-negative emotional states attests to the potential of bags-of-words in speech emotion recognition in the wild. Florian B. Pokorny, Franz Graf 0002, Franz Pernkopf, Björn W. Schuller |
ACII | 3 |
| 2015 | On Theoretical Properties of Sum-Product NetworksabstractSum-product networks (SPNs) are a promising avenue for probabilistic modeling and have been successfully applied to various tasks. However, some theoretic properties about SPNs are not yet well understood. In this paper we fill some gaps in the theoretic foundation of SPNs. First, we show that the weights of any complete and consistent SPN can be transformed into locally normalized weights without changing the SPN distribution. Second, we show that consistent SPNs cannot model distributions significantly (exponentially) more compactly than decomposable SPNs. As a third contribution, we extend the inference mechanisms known for SPNs with finite states to generalized SPNs with arbitrary input distributions. Robert Peharz, Sebastian Tschiatschek, Franz Pernkopf, Pedro M. Domingos |
AISTATS | 3 |
| 2015 | Multi-channel speech processing architectures for noise robust speech recognition: 3rd CHiME challenge resultsabstractRecognizing speech under noisy condition is an ill-posed problem. The CHiME 3 challenge targets robust speech recognition in realistic environments such as street, bus, caffee and pedestrian areas. We study variants of beamformers used for pre-processing multi-channel speech recordings. In particular, we investigate three variants of generalized side-lobe canceller (GSC) beamformers, i.e. GSC with sparse blocking matrix (BM), GSC with adaptive BM (ABM), and GSC with minimum variance distortionless response (MVDR) and ABM. Furthermore, we apply several post-filters to further enhance the speech signal. We introduce MaxPower postfilters and deep neural postfilters (DPFs). DPFs outperformed our baseline systems significantly when measuring the overall perceptual score (OPS) and the perceptual evaluation of speech quality (PESQ). In particular DPFs achieved an average relative improvement of 17.54% OPS points and 18.28% in PESQ, when compared to the CHiME 3 baseline. DPFs also achieved the best WER when combined with an ASR engine on simulated development and evaluation data, i.e. 8.98% and 10.82% WER. The proposed MaxPower beamformer achieved the best overall WER on CHiME 3 real development and evaluation data, i.e. 14.23% and 22.12%, respectively. Lukas Pfeifenberger, Tobias Schrank, Matthias Zöhrer, Martin Hagmüller, Franz Pernkopf |
ASRU | 5 |
| 2015 | Representation models in single channel source separationabstractModel-based single-channel source separation (SCSS) is an ill-posed problem requiring source-specific prior knowledge. In this paper, we use representation learning and compare general stochastic networks (GSNs), Gauss Bernoulli restricted Boltzmann machines (GBRBMs), conditional Gauss Bernoulli restricted Boltzmann machines (CGBRBMs), and higher order contractive autoencoders (HCAEs) for modeling the source-specific knowledge. In particular, these models learn a mapping from speech mixture spectrogram representations to single-source spectrogram representations, i.e. we apply them as filter for the speech mixture. In the test case, the individual source spectrograms of both models are inferred and the softmask for re-synthesis of the time signals is determined thereof. We evaluate the deep architectures on data of the 2nd CHiME speech separation challenge and provide results for a speaker dependent, a speaker independent, a matched noise condition and an unmatched noise condition task. Our experiments show the best PESQ and overall perceptual score on average for GSNs in all four tasks. Matthias Zöhrer, Franz Pernkopf |
ICASSP | 2 |
| 2015 | Neural higher-order factors in conditional random fields for phoneme classificationabstractWe explore neural higher-order input-dependent factors inlinear-chain conditional random fields (LC-CRFs) for sequencelabeling. It is a fusion of two powerful models as higher-orderLC-CRFs with linear factors are well-established for sequencelabeling tasks, but they lack to model non-linear dependencies.Therefore, we present neural higher-order input-dependent fac-tors which map sub-sequences of inputs to sub-sequences ofoutputs using distinct multilayer perceptron sub-networks. Thisis important in many tasks, in particular, for phoneme classifi-cation where the phone representation strongly depends on thecontext phonemes. Experimental results for phoneme classifi-cation with LC-CRFs and neural higher-order factors confirmthis fact and we achieve the best ever reported phoneme clas-sification performance on TIMIT, i.e. a phoneme error rate of15:8%. Furthermore, we show that the success is not obviousas linear high-order factors degrade phoneme classification per-formance on TIMIT. Martin Ratajczak, Sebastian Tschiatschek, Franz Pernkopf |
INTERSPEECH | 3 |
| 2015 | On representation learning for artificial bandwidth extensionabstractRecently, sum-product networks (SPNs) showed convincing results on the ill-posed task of artificial bandwidth extension (ABE). However, SPNs are just one type of many architectures which can be summarized as representational models. In this paper, using ABE as benchmark task, we perform a comparative study of Gauss Bernoulli restricted Boltzmann machines, conditional restricted Boltzmann machines, higher order contractive autoencoders, SPNs and generative stochastic networks (GSNs). Especially the latter ones are promising architectures in terms of its reconstruction capabilities. Our experiments show impressive results of GSNs, achieving on average an improvement of 3.90dB and 4.08dB in segmental SNR on a speaker dependent (SD) and speaker independent (SI) scenario compared to SPNs, respectively. Matthias Zöhrer, Robert Peharz, Franz Pernkopf |
INTERSPEECH | 3 |
| 2015 | Message Scheduling Methods for Belief Propagation
Christian Knoll 0002, Michael Rath 0001, Sebastian Tschiatschek, Franz Pernkopf |
ECML/PKDD (2) | 4 |
| 2015 | Structured Regularizer for Neural Higher-Order Sequence Models
Martin Ratajczak, Sebastian Tschiatschek, Franz Pernkopf |
ECML/PKDD (1) | 3 |
| 2015 | Parameter Learning of Bayesian Network Classifiers Under Computational Constraints
Sebastian Tschiatschek, Franz Pernkopf |
ECML/PKDD (1) | 2 |
| 2015 | On Bayesian Network Classifiers with Reduced Precision ParametersabstractBayesian network classifier (BNCs) are typically implemented on nowadays desktop computers. However, many real world applications require classifier implementation on embedded or low power systems. Aspects for this purpose have not been studied rigorously. We partly close this gap by analyzing reduced precision implementations of BNCs. In detail, we investigate the quantization of the parameters of BNCs with discrete valued nodes including the implications on the classification rate (CR). We derive worst-case and probabilistic bounds on the CR for different bit-widths. These bounds are evaluated on several benchmark datasets. Furthermore, we compare the classification performance and the robustness of BNCs with generatively and discriminatively optimized parameters, i.e. parameters optimized for high data likelihood and parameters optimized for classification, with respect to parameter quantization. Generatively optimized parameters are more robust for very low bit-widths, i.e. less classifications change because of quantization. However, classification performance is better for discriminatively optimized parameters for all but very low bit-widths. Additionally, we perform analysis for margin-optimized tree augmented network (TAN) structures which outperform generatively optimized TAN structures in terms of CR and robustness. Sebastian Tschiatschek, Franz Pernkopf |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Representation Learning for Single-Channel Source Separation and Bandwidth ExtensionabstractIn this paper, we use deep representation learning for model-based single-channel source separation (SCSS) and artificial bandwidth extension (ABE). Both tasks are ill-posed and source-specific prior knowledge is required. In addition to well-known generative models such as restricted Boltzmann machines and higher order contractive autoencoders two recently introduced deep models, namely generative stochastic networks (GSNs) and sum-product networks (SPNs), are used for learning spectrogram representations. For SCSS we evaluate the deep architectures on data of the 2ndCHiME speech separation challenge and provide results for a speaker dependent, a speaker independent, a matched noise condition and an unmatched noise condition task. GSNs obtain the best PESQ and overall perceptual score on average in all four tasks. Similarly, frame-wise GSNs are able to reconstruct the missing frequency bands in ABE best, measured in frequency-domain segmental SNR. They outperform SPNs embedded in hidden Markov models and the other representation models significantly. Matthias Zöhrer, Robert Peharz, Franz Pernkopf |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Modeling speech with sum-product networks: Application to bandwidth extensionabstractSum-product networks (SPNs) are a recently proposed type of probabilistic graphical models allowing complex variable interactions while still granting efficient inference. In this paper we demonstrate the suitability of SPNs for modeling log-spectra of speech signals using the application of artificial bandwidth extension, i.e. artificially replacing the high-frequency content which is lost in telephone signals. We use SPNs as observation models in hidden Markov models (HMMs), which model the temporal evolution of log short-time spectra. Missing frequency bins are replaced by the SPNs using most-probable-explanation inference, where the state-dependent reconstructions are weighted with the HMM state posterior. According to subjective listening and objective evaluation, our system consistently and significantly improves the state of the art. Robert Peharz, Georg Kapeller, Pejman Mowlaee, Franz Pernkopf |
ICASSP | 4 |
| 2014 | Blind source extraction based on a direction-dependent a-priori SNRabstractIn many hands-free applications, we encounter a speaker located in the near-field embedded in diffuse far-field noise. In this paper, we contribute an algorithm to estimate the speech and noise power spectral density (PSD) based on a directiondependent SNR (DD-SNR). The only prior knowledge needed is a model of the diffuse noise sound field. The enhanced speech signal is obtained by a parametric multi-channel Wiener filter (PMWF), which is constructed without any speech presence or absence probabilities, or smoothing in frequency. We achieve high speech quality and sufficient noise reduction by iteratively improving the speech PSD estimate using the output of the PMWF. The performance of our algorithm is demonstrated by using the PESQ and PEASS measures. Lukas Pfeifenberger, Franz Pernkopf |
INTERSPEECH | 2 |
| 2014 | Self-adaption in single-channel source separationabstractSingle-channel source separation (SCSS) usually uses pre-trained source-specific models to separate the sources. These models capture the characteristics of each source and they perform well when matching the test conditions. In this paper, we extend the applicability of SCSS. We develop an EM-like iterative adaption algorithm which is capable to adapt the pre-trained models to the changed characteristics of the specific situation, such as a different acoustic channel introduced by variation in the room acoustics or changed speaker position. The adaption framework requires signal mixtures only, i.e. specific single source signals are not necessary. We consider speech/noise mixtures and we restrict the adaption to the speech model only. Model adaption is empirically evaluated using mixture utterances from the CHiME 2 challenge. We perform experiments using speaker dependent (SD) and speaker independent (SI) models trained on clean or reverberated single speaker utterances. We successfully adapt SI source models trained on clean utterances and achieve almost the same performance level as SD models trained on reverberated utterances. Michael Wohlmayr, Ludwig Mohr, Franz Pernkopf |
INTERSPEECH | 3 |
| 2014 | Single channel source separation with general stochastic networksabstractSingle channel source separation (SCSS) is ill-posed and thus challenging. In this paper, we apply general stochastic networks (GSNs) – a deep neural network architecture – to SCSS. We extend GSNs to be capable of predicting a time-frequency representation, i.e. softmask by introducing a hybrid generative-discriminative training objective to the network. We evaluate GSNs on data of the 2nd CHiME speech separation challenge. In particular, we provide results for a speaker dependent, a speaker independent, a matched noise condition and an unmatched noise condition task. Empirically, we compare to other deep architectures, namely a deep belief network (DBN) and a multi-layer perceptron (MLP). In general, deep architectures perform well on SCSS tasks. Matthias Zöhrer, Franz Pernkopf |
INTERSPEECH | 2 |
| 2014 | General Stochastic Networks for Classification
Matthias Zöhrer, Franz Pernkopf |
NIPS | 2 |
| 2014 | Integer Bayesian Network Classifiers
Sebastian Tschiatschek, Karin Paul, Franz Pernkopf |
ECML/PKDD (3) | 3 |
| 2013 | On the Asymptotic Optimality of Maximum Margin Bayesian NetworksabstractMaximum margin Bayesian networks (MMBNs) are Bayesian networks with discriminatively optimized parameters. They have shown good classification performance in various applications. However, there has not been any theoretic analysis of their asymptotic performance, e.g. their Bayes consistency. For specific classes of MMBNs, i.e. MMBNs with fully connected graphs and discrete-valued nodes, we show Bayes consistency for binary-class problems and a sufficient condition for Bayes consistency in the multi-class case. We provide simple examples showing that MMBNs in their current formulation are not Bayes consistent in general. These examples are especially interesting, as the model used for the MMBNs can represent the assumed true distributions. This indicates that the current formulations of MMBNs may be deficient. Furthermore, experimental results on the generalization performance are presented. Sebastian Tschiatschek, Franz Pernkopf |
AISTATS | 2 |
| 2013 | Generalization of pre-image iterations for speech enhancementabstractIn this paper, we extend the pre-image iteration method for speech de-noising by automatic determination of the kernel variance. The kernel variance needs to be adapted in different noise conditions. In previous work, the signal-to-noise ratio (SNR) was assumed to be known and the kernel variance was pre-defined using a development set. In the proposed method, a function is derived that maps a noise estimate to a potentially good value for the kernel variance. Hence, the SNR is not required to be known. Furthermore, the method is adapted for scenarios with colored noise, where - due to the properties of the noise - a different kernel variance for each frequency leads to better performance. We compare the proposed methods to the original pre-image iteration method and show an increase in performance in terms of the PEASS quality measures. Christina Leitner, Franz Pernkopf |
ICASSP | 2 |
| 2013 | Bounds for Bayesian network classifiers with reduced precision parametersabstractBayesian network classifiers are probabilistic classifiers achieving good classification rates in various applications. These classifiers consist of a directed acyclic graph and a set of conditional probability densities, which in case of discrete-valued nodes can be represented by conditional probability tables. In this paper, we investigate the effect of quantizing these conditional probabilities. We derive worst-case and best-case bounds on the classification rate using interval arithmetic. Furthermore, we determine performance bounds that hold with a user specified confidence using quantization theory. Our results emphasize that only small bit-widths are necessary to achieve good classification rates. Sebastian Tschiatschek, Carlos Eduardo Cancino-Chacón, Franz Pernkopf |
ICASSP | 3 |
| 2013 | Model adaptation of factorial HMMS for multipitch trackingabstractFactorial hidden Markov models (FHMMs) are used for tracking the pitch of two interacting speakers [1]. In this statistical approach, the characteristics of each speaker are captured by pre-trained models. Speaker models that match the test conditions well allow for high tracking performance, however the availability of such models is unrealistic. To extend the applicabiliy of the FHMM framework, we develop an EM-like iterative adaptation algorithm which is capable to adapt the model parameters to the specific situation, e.g. acoustic channel, using only speech mixture data. Model adaptation is empirically evaluated using real room recordings of mixture utterances from the GRID corpus. Michael Wohlmayr, Franz Pernkopf |
ICASSP | 2 |
| 2013 | The Most Generative Maximum Margin Bayesian NetworksabstractAlthough discriminative learning in graphical models generally improves classification results, the generative semantics of the model are compromised. In this paper, we introduce a novel approach of hybrid generative-discriminative learning for Bayesian networks. We use an SVM-type large margin formulation for discriminative training, introducing a likelihood-weighted \ell^1-norm for the SVM-norm-penalization. This simultaneously optimizes the data likelihood and therefore partly maintains the generative character of the model. For many network structures, our method can be formulated as a convex problem, guaranteeing a globally optimal solution. In terms of classification, the resulting models outperform state-of-the art generative and discriminative learning methods for Bayesian networks, and are comparable with linear and kernelized SVMs. Furthermore, the models achieve likelihoods close to the maximum likelihood solution and show robust behavior in classification experiments with missing features. Robert Peharz, Sebastian Tschiatschek, Franz Pernkopf |
ICML (3) | 3 |
| 2013 | Greedy Part-Wise Learning of Sum-Product Networks
Robert Peharz, Bernhard C. Geiger, Franz Pernkopf |
ECML/PKDD (2) | 3 |
| 2013 | Stochastic margin-based structure learning of Bayesian network classifiersabstractThe margin criterion for parameter learning in graphical models gained significant impact over the last years. We use the maximum margin score for discriminatively optimizing the structure of Bayesian network classifiers. Furthermore, greedy hill-climbing and simulated annealing search heuristics are applied to determine the classifier structures. In the experiments, we demonstrate the advantages of maximum margin optimized Bayesian network structures in terms of classification performance compared to traditionally used discriminative structure learning methods. Stochastic simulated annealing requires less score evaluations than greedy heuristics. Additionally, we compare generative and discriminative parameter learning on both generatively and discriminatively structured Bayesian network classifiers. Margin-optimized Bayesian network classifiers achieve similar classification performance as support vector machines. Moreover, missing feature values during classification can be handled by discriminatively optimized Bayesian network classifiers, a case where purely discriminative classifiers usually require mechanisms to complete unknown feature values in the data first. Franz Pernkopf, Michael Wohlmayr |
Pattern Recognit. | 1 |
| 2013 | Model-Based Multiple Pitch Tracking Using Factorial HMMs: Model Adaptation and InferenceabstractRobustness against noise and interfering audio signals is one of the challenges in speech recognition and audio analysis technology. One avenue to approach this challenge is single-channel multiple-source modeling. Factorial hidden Markov models (FHMMs) are capable of modeling acoustic scenes with multiple sources interacting over time. While these models reach good performance on specific tasks, there are still serious limitations restricting the applicability in many domains. In this paper, we generalize these models and enhance their applicability. In particular, we develop an EM-like iterative adaptation framework which is capable to adapt the model parameters to the specific situation (e.g. actual speakers, gain, acoustic channel, etc.) using only speech mixture data. Currently, source-specific data is required to learn the model. Inference in FHMMs is an essential ingredient for adaptation. We develop efficient approaches based on observation likelihood pruning. Both adaptation and efficient inference are empirically evaluated for the task of multipitch tracking using the GRID corpus. Michael Wohlmayr, Franz Pernkopf |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Speech enhancement using pre-image iterationsabstractIn this paper, we present a new method to de-noise speech in the complex spectral domain. The method is derived from kernel principal component analysis (kPCA). Instead of applying PCA in a high-dimensional feature space and then going back to the original input space by using a solution to the pre-image problem, only the pre-image step is applied for de-noising. We show that the de-noised audio sample is a convex combination of the noisy input data and that the resulting algorithm is closely related to the soft k-means algorithm. Compared to kPCA, this method reduces the computational costs while the audio quality is similar and speech quality measures do not degrade. Christina Leitner, Franz Pernkopf |
ICASSP | 2 |
| 2012 | On linear and mixmax interaction models for single channel source separationabstractFor model-based single channel source separation, one typically assumes a linear interaction model, i.e. that the mixture magnitude spectrogram is the sum of the individual source magnitude spectrograms. In the log-domain, the MIXMAX interaction model is the corresponding approximation for the linear model. Hence, one would expect similar performance for both approaches. However, in this paper we empirically show that this is not the case for vector-quantizer-based (VQ) single channel source separation. We propose factorial linear-VQ, the linear counterpart to factorial max-VQ, and compare the two methods in systematic source separation experiments. Linear-VQ performs significantly better than max-VQ for comparable code-book sizes and behaves more robustly in the presence of additive white noise. Furthermore, we compare resynthesis properties of binary and continuous time-frequency masks. While binary masks achieve a higher interference suppression, the use of continuous masks results in a consistently better signal quality. Robert Peharz, Franz Pernkopf |
ICASSP | 2 |
| 2012 | Exact Maximum Margin Structure Learning of Bayesian Networks
Robert Peharz, Franz Pernkopf |
ICML | 2 |
| 2012 | Convex Combinations of Maximum Margin Bayesian Network Classifiers
Sebastian Tschiatschek, Franz Pernkopf |
ICPRAM (1) | 2 |
| 2012 | Bayesian Network Classifiers with Reduced Precision Parameters
Sebastian Tschiatschek, Peter Reinprecht, Manfred Mücke, Franz Pernkopf |
ECML/PKDD (1) | 4 |
| 2012 | Sparse nonnegative matrix factorization with ℓ0-constraintsabstractAlthough nonnegative matrix factorization (NMF) favors a sparse and part-based representation of nonnegative data, there is no guarantee for this behavior. Several authors proposed NMF methods which enforce sparseness by constraining or penalizing the [Formula: see text] of the factor matrices. On the other hand, little work has been done using a more natural sparseness measure, the [Formula: see text]. In this paper, we propose a framework for approximate NMF which constrains the [Formula: see text] of the basis matrix, or the coefficient matrix, respectively. For this purpose, techniques for unconstrained NMF can be easily incorporated, such as multiplicative update rules, or the alternating nonnegative least-squares scheme. In experiments we demonstrate the benefits of our methods, which compare to, or outperform existing approaches. Robert Peharz, Franz Pernkopf |
Neurocomputing | 2 |
| 2012 | Maximum Margin Bayesian Network ClassifiersabstractWe present a maximum margin parameter learning algorithm for Bayesian network classifiers using a conjugate gradient (CG) method for optimization. In contrast to previous approaches, we maintain the normalization constraints on the parameters of the Bayesian network during optimization, i.e., the probabilistic interpretation of the model is not lost. This enables us to handle missing features in discriminatively optimized Bayesian networks. In experiments, we compare the classification performance of maximum margin parameter learning to conditional likelihood and maximum likelihood learning approaches. Discriminative parameter learning significantly outperforms generative maximum likelihood estimation for naive Bayes and tree augmented naive Bayes structures on all considered data sets. Furthermore, maximizing the margin dominates the conditional likelihood approach in terms of classification performance in most cases. We provide results for a recently proposed maximum margin optimization approach based on convex relaxation. While the classification results are highly similar, our CG-based optimization is computationally up to orders of magnitude faster. Margin-optimized Bayesian network classifiers achieve classification performance comparable to support vector machines (SVMs) using fewer parameters. Moreover, we show that unanticipated missing feature values during classification can be easily processed by discriminatively optimized Bayesian network classifiers, a case where discriminative classifiers usually require mechanisms to complete unknown feature values in the data first. Franz Pernkopf, Michael Wohlmayr, Sebastian Tschiatschek |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Gain-robust multi-pitch tracking using sparse nonnegative matrix factorizationabstractWhile nonnegative matrix factorization (NMF) has successfully been applied for gain-robust multi-pitch detection, a method to track pitch values over time was not provided. We embed NMF-based pitch detection into a recently proposed pitch-tracking system, based on a factorial hidden Markov model (FHMM). The original system models speech spectra with Gaussian mixture models, which is sensitive to a gain mismatch between training and test data. We therefore combine the advantages of these two approaches and derive a gain-adaptive observation model for the FHMM. As training algorithm we use a modification of ℓ0-sparse NMF, which represents the short-time spectrum with scalable basis vectors. In experiments we show that the new approach significantly increases the gain-robustness of the original tracking system. Robert Peharz, Michael Wohlmayr, Franz Pernkopf |
ICASSP | 3 |
| 2011 | Maximum margin structure learning of Bayesian network classifiersabstractRecently, the margin criterion has been successfully used for parameter optimization in graphical models. We introduce maximum margin based structure learning for Bayesian network classifiers and demonstrate its advantages in terms of classification performance compared to traditionally used discriminative structure learning methods. In particular, we provide empirical results for generative structure learning and two discriminative structure learning approaches on handwritten digit recognition tasks. We show that maximum margin structure learning outperforms other structure learning methods. Furthermore, we present classification results achieved with different bitwidth for representing the parameters of the classifiers. Franz Pernkopf, Michael Wohlmayr, Manfred Mücke |
ICASSP | 1 |
| 2011 | Efficient implementation of probabilistic multi-pitch trackingabstractWe significantly improve the computational efficiency of a probabilistic approach for multiple pitch tracking. This method is based on a factorial hidden Markov model and two alternative interaction models for magnitude and log-magnitude spectra, respectively. The main computational bottleneck comprises the determination of observation likelihoods. However, we show that up to 99.5% of the smallest likelihood values can be discarded at each time frame with out affecting the overall tracking accuracy. For both interaction models, we present a heuristic to efficiently find the largest likelihood values. Experiments on the GRID database show that the proposed methods result in a major speedup without significantly changing tracking accuracy. Michael Wohlmayr, Robert Peharz, Franz Pernkopf |
ICASSP | 3 |
| 2011 | Kernel PCA for Speech EnhancementabstractIn this paper, we apply kernel principal component analysis (kPCA), which has been successfully used for image denoising, to speech enhancement. In contrast to other enhancement methods which are based on the magnitude spectrum, we rather apply kPCA to complex spectral data. This is facilitated by Gaussian kernels. In the experiments, we show good noise reduction with few artifacts for noise corrupted speech at different SNR levels using additive white Gaussian noise. We compared kPCA with linear PCA and spectral subtraction and evaluated all algorithms with perceptually motivated quality measures. Christina Leitner, Franz Pernkopf, Gernot Kubin |
INTERSPEECH | 2 |
| 2011 | A Pitch Tracking Corpus with Evaluation on Multipitch Tracking ScenarioabstractIn this paper, we introduce a novel pitch tracking database (PTDB) including ground truth signals obtained from a laryngograph. The database, referenced as PTDB-TUG, consists of 2342 phonetically rich sentences taken from the TIMIT corpus. Each sentence was at least recorded once by a male and a female native speaker. In total, the database contains 4720 recordings from 10 male and 10 female speakers. Furthermore, we evaluated two multipitch tracking systems on a subset of speakers to provide a benchmark for further research activities. The database can be downloaded at http://www.spsc.tugraz.at/tools. Gregor Pirker, Michael Wohlmayr, Stefan Petrik, Franz Pernkopf |
INTERSPEECH | 4 |
| 2011 | EM-Based Gain Adaptation for Probabilistic Multipitch TrackingabstractWe introduce an EM algorithm for automatic speaker gain adaptation, and use this approach for probabilistic multipitch tracking. We derive a lower bound on the log-likelihood of the gain parameters and use a fast pruning method to make lower bound optimization efficient. We evaluate the performance of gain adapted multipitch tracking on the GRID database, where 3000 speech mixtures were generated for each mixing level. For gain differences in the range of zero up to 18dB, the proposed method achieves almost the same performance as for the case where the gain is assumed to be known. Michael Wohlmayr, Franz Pernkopf |
INTERSPEECH | 2 |
| 2011 | Semantic and phonetic automatic reconstruction of medical dictations
Stefan Petrik, Christina Drexel, Leo Fessler, Jeremy Jancsary, Alexandra Klein, Gernot Kubin, Johannes Matiasek, Franz Pernkopf, Harald Trost |
Comput. Speech Lang. | 8 |
| 2011 | Source-Filter-Based Single-Channel Speech Separation Using Pitch InformationabstractIn this paper, we investigate the source-filter-based approach for single-channel speech separation. We incorporate source-driven aspects by multi-pitch estimation in the model-driven method. For multi-pitch estimation, the factorial HMM is utilized. For modeling the vocal tract filters either vector quantization (VQ) or non-negative matrix factorization are considered. For both methods, the final combination of the source and filter model results in an utterance dependent model that finally enables speaker independent source separation. The contributions of the paper are the multi-pitch tracker, the gain estimation for the VQ based method which accounts for different mixing levels, and a fast approximation for the likelihood computation. Additionally, a linear relationship between pitch tracking performance and speech separation performance is shown. Michael Stark 0004, Michael Wohlmayr, Franz Pernkopf |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | A Probabilistic Interaction Model for Multipitch Tracking With Factorial Hidden Markov ModelsabstractWe present a simple and efficient feature modeling approach for tracking the pitch of two simultaneously active speakers. We model the spectrogram features of single speakers using Gaussian mixture models in combination with the minimum description length model selection criterion. To obtain a probabilistic representation for the speech mixture spectrogram features of both speakers, we employ the mixture maximization model (MIXMAX) and, as an alternative, a linear interaction model. A factorial hidden Markov model is applied for tracking pitch over time. This statistical model can be used for applications beyond speech, whenever the interaction between individual sources can be represented as MIXMAX or linear model. For tracking, we use the loopy max-sum algorithm, and provide empirical comparisons to exact methods. Furthermore, we discuss a scheduling mechanism of loopy belief propagation for online tracking. We demonstrate experimental results using Mocha-TIMIT as well as data from the speech separation challenge provided by Cooke We show the excellent performance of the proposed method in comparison to a well known multipitch tracking algorithm based on correlogram features. Using speaker-dependent models, the proposed method improves the accuracy of correct speaker assignment, which is important for single-channel speech separation. In particular, we are able to reduce the overall tracking error by 51% relative for the speaker-dependent case. Moreover, we use the estimated pitch trajectories to perform single-channel source separation, and demonstrate the beneficial effect of correct speaker assignment on speech separation performance. Michael Wohlmayr, Michael Stark 0004, Franz Pernkopf |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | On optimizing the computational complexity for VQ-based single channel source separationabstractWe introduce two suboptimal search heuristics for reducing the computational burden in single channel source separation. The heuristics approximating the observation likelihood are evaluated using the speaker dependent factorial-max VQ model. One approach extends beam search, whereas the second relies on the iterated conditional modes algorithm. We compare the methods to the hierarchically structured VQ model [1] and to the full search using the Grid Corpus [2]. The first two algorithms reduce the computational costs by almost two orders of magnitude compared to full search, whereas the separation performance shows a slight and insignificant decrease in terms of target-to-masker ratio. Additionally, the heuristics are compared in terms of execution time. Michael Stark 0004, Franz Pernkopf |
ICASSP | 2 |
| 2010 | A mixture maximization approach to multipitch tracking with factorial hidden Markov modelsabstractWe present a simple and efficient feature modeling approach for tracking the pitch of two speakers speaking simultaneously. We model the spectrogram features of single speakers using Gaussian mixture models in combination with the minimum description length model selection criterion. Furthermore, the mixture maximization (MIXMAX) interaction model is employed to yield a probabilistic representation for the mixture of both speakers. Finally, a factorial hidden Markov model is applied for tracking. We demonstrate experimental results on two databases, and show the excellent performance of the proposed method in comparison to a well known multipitch tracking algorithm based on correlogram features. Michael Wohlmayr, Michael Stark 0004, Franz Pernkopf |
ICASSP | 3 |
| 2010 | Single Channel Speech Separation Using Source-Filter RepresentationabstractWe propose a fully probabilistic model for source-filter based single channel source separation. In particular, we perform separation in a sequential manner, where we estimate the source-driven aspects by a factorial HMM used for multi-pitch estimation. Afterwards, these pitch tracks are combined with the vocal tract filter model to form an utterance dependent model. Additionally, we introduce a gain estimation approach to enable adaptation to arbitrary mixing levels in the speech mixtures. We thoroughly evaluate this system and finally end up in a speaker independent model. Michael Stark 0004, Michael Wohlmayr, Franz Pernkopf |
ICPR | 3 |
| 2010 | A factorial sparse coder model for single channel source separationabstractWe propose a probabilistic factorial sparse coder model for single channel source separation in the magnitude spectrogram domain. The mixture spectrogram is assumed to be the sum of the sources, which are assumed to be generated frame-wise as the output of sparse coders plus noise. For dictionary training we use an algorithm which can be described as non-negative matrix factorization with ℓ0 sparseness constraints. In order to infer likely source spectrogram candidates, we approximate the intractable exact inference by maximizing the posterior over a plausible subset of solutions. We compare our system to the factorial-max vector quantization model, where the proposed method shows a superior performance in terms of signal-to-interference ratio. Finally, the low computational requirements of the algorithm allows close to real time applications. © 2010 ISCA. Robert Peharz, Michael Stark 0004, Franz Pernkopf, Yannis Stylianou |
INTERSPEECH | 3 |
| 2010 | Large Margin Learning of Bayesian Classifiers Based on Gaussian Mixture Models
Franz Pernkopf, Michael Wohlmayr |
ECML/PKDD (3) | 1 |
| 2010 | Efficient Heuristics for Discriminative Structure Learning of Bayesian Network Classifiers
Franz Pernkopf, Jeff A. Bilmes |
J. Mach. Learn. Res. | 1 |
| 2010 | Joint Time-Frequency Segmentation Algorithm for Transient Speech Decomposition and Speech EnhancementabstractWe develop an algorithm, the joint time-frequency segmentation algorithm, where the wavelet packet coefficients of the analyzed speech signal are represented as tiles of a time-frequency representation adapted to the characteristics of the signal itself. Further, our algorithm enables the decomposition of the speech signal into transient and non-transient components, respectively. Any block of wavelet packet coefficients, whose tiling height is larger than or equal to the tiling width belongs to the transient component and vice versa for the non-transient component. The transient component is selectively amplified and recombined with the original speech to generate the modified speech with energy adjusted to be equal to the original speech. The intelligibility of the original and modified speech is evaluated by 16 human listeners. Word recognition rate results show that the modified speech significantly improves speech intelligibility in background noise, i.e., by 10% absolute at 0 dB to 27% absolute at -30 dB. Charturong Tantibundhit, Franz Pernkopf, Gernot Kubin |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Effective metric-based speaker segmentation in the frequency domainabstractIn this paper, we present an approach, called FREQDIST, for speaker segmentation based on a distance measurement applied in the frequency domain. To enhance the detection performance, the spectrum is reweighted using normalization techniques. Additionally, noise-like (i.e. flat) spectra are removed based on the entropy. Experiments using the TIMIT database [1] and Westdeutscher Rundfunk broadcast data show that our segmentation approach yields a good performance compared to the DISTBIC algorithm [2]. In particular, for the TIMIT data our algorithm reaches a false alarm rate (FAR) less than half of the value of the DISTBIC algorithm and a missed detection rate (MDR) of 7.0% instead of 13.1%. Christoph Böhm 0003, Franz Pernkopf |
ICASSP | 2 |
| 2009 | Towards source-filter based single sensor speech separationabstractWe present a new source-filter based method to separate two speakers talking simultaneously at equal level mixed into a single sensor. First, the relation between the spectral whitened mixture and the speakers excitation signals is analyzed. Therefore, a factorial HMM capturing also time dependencies is exploited. Then, the estimated excitation signals are combined with best fitting vocal tract information taken from a trained dictionary. We report results on the database of Cooke considering 108 speech mixtures. The average improvement of 2.9 dB in SIR for all data is lower but not significantly lower compared to the Gaussian mixture method which relies on known pitch-tracks. Although the performance is currently moderate we believe in this approach and its significance towards the development of speaker independent single sensor speech separation. Michael Stark 0004, Franz Pernkopf |
ICASSP | 2 |
| 2009 | Speech enhancement based on joint time-frequency segmentationabstractWe present an algorithm to decompose speech into transient and non-transient components. Our algorithm, the joint time-frequency segmentation algorithm, uses the wavelet packet coefficients of the speech signal and represents them as tiles of a time-frequency representation adapted to the characteristics of the signal itself. Any wavelet packet coefficient, whose tiling height is larger than or equal to the tiling width is characterized as a transient coefficient and vice versa for the non-transient coefficient. The transient component is selectively amplified and recombined with the original speech to generate the modified speech with energy adjusted to be equal to the energy of the original speech. The psychoacoustic tests performed with fourteen human listeners show that the speech modification significantly improves speech intelligibility in background noise, i.e., for 10% absolute at 0d B to 31% absolute at -30 dB. Charturong Tantibundhit, Franz Pernkopf, Gernot Kubin |
ICASSP | 2 |
| 2009 | Wavelet-based speaker change detection in single channel speech data
Michael Wiesenegger, Franz Pernkopf |
INTERSPEECH | 2 |
| 2009 | Finite mixture spectrogram modeling for multipitch tracking using a factorial hidden Markov model
Michael Wohlmayr, Franz Pernkopf |
INTERSPEECH | 2 |
| 2009 | On Discriminative Parameter Learning of Bayesian Network Classifiers
Franz Pernkopf, Michael Wohlmayr |
ECML/PKDD (2) | 1 |
| 2009 | Broad phonetic classification using discriminative Bayesian networks
Franz Pernkopf, Tuan Van Pham, Jeff A. Bilmes |
Speech Commun. | 1 |
| 2008 | Automatic phonetics-driven reconstruction of medical dictations on multiple levels of segmentationabstractAutomatic phonetic reconstruction of medical dictations from non- literal and automatically recognized speech transcripts leads to closer-to-literal transcripts for training. In this paper, we introduce an extended alignment method assessing multiple levels of text segmentation and show how open issues like wrong segmentation in the recognized transcript can be resolved. Furthermore, the effect of context-dependent reconstruction and the phonetic similarity threshold on the quality of the reconstructed transcription is measured. Experiments show an increase in precision between 0.7% and 4.7% absolute without loss in recall for the combined system incorporating all of these techniques in comparison to the system in the previous work. Stefan Petrik, Franz Pernkopf |
ICASSP | 2 |
| 2008 | Voice activity detection algorithms using subband power distance feature for noisy environments
Tuan Van Pham, Michael Stadtschnitzer, Franz Pernkopf, Gernot Kubin |
INTERSPEECH | 3 |
| 2008 | Multipitch tracking using a factorial hidden Markov modelabstractIn this paper, we present an approach to track the pitch of two simultaneous speakers. Using a well-known feature extraction method based on the correlogram, we track the resulting data using a factorial hidden Markov model (FHMM). In contrast to the recently developed multipitch determination algorithm [1], which is based on a HMM, we can accurately associate estimated pitch points with their corresponding source speakers. We evalute our approach on the “Mocha-TIMIT” database [2] of speech utterances mixed at 0dB, and compare the results to the multipitch determination algorithm [1] used as a baseline. Experiments show that our FHMM tracker yields good performance for both pitch estimation and correct speaker assignment. Michael Wohlmayr, Franz Pernkopf |
INTERSPEECH | 2 |
| 2008 | Tracking of Multiple Targets Using Online Learning for Reference Model AdaptationabstractRecently, much work has been done in multiple object tracking on the one hand and on reference model adaptation for a single-object tracker on the other side. In this paper, we do both tracking of multiple objects (faces of people) in a meeting scenario and online learning to incrementally update the models of the tracked objects to account for appearance changes during tracking. Additionally, we automatically initialize and terminate tracking of individual objects based on low-level features, i.e., face color, face size, and object movement. Many methods unlike our approach assume that the target region has been initialized by hand in the first frame. For tracking, a particle filter is incorporated to propagate sample distributions over time. We discuss the close relationship between our implemented tracker based on particle filters and genetic algorithms. Numerous experiments on meeting data demonstrate the capabilities of our tracking approach. Additionally, we provide an empirical verification of the reference model learning during tracking of indoor and outdoor scenes which supports a more robust tracking. Therefore, we report the average of the standard deviation of the trajectories over numerous tracking runs depending on the learning rate. Franz Pernkopf |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2006 | Bayesian networks for phonetic classification using time-scale features
Franz Pernkopf, Tuan Van Pham |
INTERSPEECH | 1 |
| 2005 | On Initialization of Gaussian Mixtures: A Hybrid Genetic EM AlgorithmabstractWe propose a genetic-based expectation-maximization (GA-EM) algorithm for learning Gaussian mixture models from multivariate data. This algorithm is capable of selecting the number of components of the model using the minimum description length (MDL) criterion. We combine EM and GA into a single procedure. The population-based stochastic search of the GA explores the search space more thoroughly than the EM method. Therefore, our algorithm enables us to escape from local optimal solutions since the algorithm becomes less sensitive to its initialization. The GA-EM algorithm is elitist which maintains the monotonic convergence property of the EM algorithm. The experiments show that the GA-EM outperforms the EM method since: (i) we have obtained a better MDL score while using exactly the same initialization and termination condition for both algorithms; (ii) our approach identifies the number of components which were used to generate the underlying data more often as the EM algorithm. Franz Pernkopf |
ICASSP (1) | 1 |
| 2005 | Discriminative versus generative parameter and structure learning of Bayesian network classifiersabstractIn this paper, we compare both discriminative and generative parameter learning on both discriminatively and generatively structured Bayesian network classifiers. We use either maximum likelihood (ML) or conditional maximum likelihood (CL) to optimize network parameters. For structure learning, we use either conditional mutual information (CMI), the explaining away residual (EAR), or the classification rate (CR) as objective functions. Experiments with the naive Bayes classifier (NB), the tree augmented naive Bayes classifier (TAN), and the Bayesian multinet have been performed on 25 data sets from the UCI repository (Merz et al., 1997) and from (Kohavi & John, 1997). Our empirical study suggests that discriminative structures learnt using CR produces the most accurate classifiers on almost half the data sets. This approach is feasible, however, only for rather small problems since it is computationally expensive. Discriminative parameter learning produces on average a better classifier than ML parameter learning. Franz Pernkopf, Jeff A. Bilmes |
ICML | 1 |
| 2005 | 3D surface analysis using coupled HMMs
Franz Pernkopf |
Mach. Vis. Appl. | 1 |
| 2005 | Genetic-Based EM Algorithm for Learning Gaussian Mixture ModelsabstractWe propose a genetic-based expectation-maximization (GA-EM) algorithm for learning Gaussian mixture models from multivariate data. This algorithm is capable of selecting the number of components of the model using the minimum description length (MDL) criterion. Our approach benefits from the properties of Genetic algorithms (GA) and the EM algorithm by combination of both into a single procedure. The population-based stochastic search of the GA explores the search space more thoroughly than the EM method. Therefore, our algorithm enables escaping from local optimal solutions since the algorithm becomes less sensitive to its initialization. The GA-EM algorithm is elitist which maintains the monotonic convergence property of the EM algorithm. The experiments on simulated and real data show that the GA-EM outperforms the EM method since: 1) We have obtained a better MDL score while using exactly the same termination condition for both algorithms. 2) Our approach identifies the number of components which were used to generate the underlying data more often than the EM algorithm. Franz Pernkopf, Djamel Bouchaffra |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Bayesian network classifiers versus selective k-NN classifier
Franz Pernkopf |
Pattern Recognit. | 1 |
| 2004 | Bayesian Network Classifiers Versus k-NN Classifier Using Sequential Feature Selection
Franz Pernkopf |
AAAI | 1 |
| 2004 | Prosody-based recognition of spoken German varietiesabstractAn approach to the recognition of regional language varieties is presented. The algorithm is tested on utterances of 3 to 6 seconds duration taken from large speech databases (SpeechDat) of Austrian and German German. The features are based only on the prosody of the speech and include parameters derived from the Fujisaki model and statistics of the fundamental frequency. Classification is performed using a multilayer perceptron and yields a rate of 64% correct. identification of the regional variety. Those results are then further evaluated for the use of a regional variety recognizer as a front-end of an automatic speech recognizer for different regional varieties. In case there is no a priori information of the distribution of the regional varieties spoken by the users, this approach yields a considerable improvement in the robustness of the speech recognition rates. Vedran Dizdarevic, Martin Hagmüller, Gernot Kubin, Franz Pernkopf, Micha Baum |
ICASSP (1) | 4 |
| 2004 | Detection of surface defects on raw steel blocks using Bayesian network classifiers
Franz Pernkopf |
Pattern Anal. Appl. | 1 |
| 2003 | Floating search algorithm for structure learning of Bayesian network classifiers
Franz Pernkopf, Paul O'Leary |
Pattern Recognit. Lett. | 1 |
| 2001 | Feature Selection for Classification Using Genetic Algorithms with a Novel Encoding
Franz Pernkopf, Paul O'Leary |
CAIP | 1 |