VLDB 2026 Research / reviewers in the wild / expert
Alfred Mertins
dblp:93/2424
· DBLP profile ↗
121ranked-venue papers
13as first author
9since 2021 · last 2023
0000-0001-5718-577XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 93 · 10 first-author · 5 since 2021Artificial intelligence and machine learning · 28 · 2 first-author · 2 since 2021Computer networks · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Improving Automatic Sleep Staging Via Temporal Smoothness RegularizationabstractWe propose a regularization method, so-called temporal smoothness regularization, for training deep neural networks for automatic sleep staging in small data settings. In intuition, we constrain the cross-entropy losses of any two adjacent epochs in the sequential input to be as close to each other as possible. The regularization closely reflects the slow transition nature of sleep process which implies small information changes between two consecutive sleep epochs. Via the regularization, we essentially discourage the network from overfitting to these small changes. Our experiments show that training the SeqSleepNet base network with the proposed regularization leads to performance improvement over the baseline without the regularization applied. Furthermore, our developed method achieves the performance on par with the state-of-the-art performance while outperforming other existing methods. Huy Phan, Elisabeth R. M. Heremans, Oliver Y. Chén, Philipp Koch, Alfred Mertins, Maarten De Vos |
ICASSP | 5 |
| 2023 | L-SeqSleepNet: Whole-cycle Long Sequence Modeling for Automatic Sleep StagingabstractHuman sleep is cyclical with a period of approximately 90 minutes, implying long temporal dependency in the sleep data. Yet, exploring this long-term dependency when developing sleep staging models has remained untouched. In this work, we show that while encoding the logic of a whole sleep cycle is crucial to improve sleep staging performance, the sequential modelling approach in existing state-of-the-art deep learning models are inefficient for that purpose. We thus introduce a method for efficient long sequence modelling and propose a new deep learning model, L-SeqSleepNet, which takes into account whole-cycle sleep information for sleep staging. Evaluating L-SeqSleepNet on four distinct databases of various sizes, we demonstrate state-of-the-art performance obtained by the model over three different EEG setups, including scalp EEG in conventional Polysomnography (PSG), in-ear EEG, and around-the-ear EEG (cEEGrid), even with a single EEG channel input. Our analyses also show that L-SeqSleepNet is able to alleviate the predominance of N2 sleep (the major class in terms of classification) to bring down errors in other sleep stages. Moreover the network becomes much more robust, meaning that for all subjects where the baseline method had exceptionally poor performance, their performance are improved significantly. Finally, the computation time only grows at a sub-linear rate when the sequence length increases. Huy Phan, Kristian P. Lorenzen, Elisabeth R. M. Heremans, Oliver Y. Chén, Minh C. Tran, Philipp Koch, Alfred Mertins, Mathias Baumert, Kaare B. Mikkelsen, Maarten De Vos |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Polyphonic Audio Event Detection: Multi-Label or Multi-Class Multi-Task Classification Problem?abstractPolyphonic events are the main error source of audio event detection (AED) systems. In deep-learning context, the most common approach to deal with event overlaps is to treat the AED task as a multi-label classification problem. By doing this, we inherently consider multiple one-vs.-rest classification problems, which are jointly solved by a single (i.e. shared) network. In this work, to better handle polyphonic mixtures, we propose to frame the task as a multi-class classification problem by considering each possible label combination as one class. To circumvent the large number of arising classes due to combinatorial explosion, we divide the event categories into multiple groups and construct a multi-task problem in a divide-and-conquer fashion, where each of the tasks is a multi-class classification problem. A network architecture is then devised for multi-class multi-task modelling. The network is composed of a backbone subnet and multiple task-specific subnets. The task-specific subnets are designed to learn time-frequency and channel attention masks to extract features for the task at hand from the common feature maps learned by the backbone. Experiments on the TUT-SED-Synthetic-2016 with high degree of event overlap show that the proposed approach results in more favorable performance than the common multi-label approach. Huy Phan, Thi Ngoc Tho Nguyen, Philipp Koch, Alfred Mertins |
ICASSP | 4 |
| 2022 | XSleepNet: Multi-View Sequential Model for Automatic Sleep StagingabstractAutomating sleep staging is vital to scale up sleep assessment and diagnosis to serve millions experiencing sleep deprivation and disorders and enable longitudinal sleep monitoring in home environments. Learning from raw polysomnography signals and their derived time-frequency image representations has been prevalent. However, learning from multi-view inputs (e.g., both the raw signals and the time-frequency images) for sleep staging is difficult and not well understood. This work proposes a sequence-to-sequence sleep staging model, XSleepNet,1that is capable of learning a joint representation from both raw signals and time-frequency images. Since different views may generalize or overfit at different rates, the proposed network is trained such that the learning pace on each view is adapted based on their generalization/overfitting behavior. In simple terms, the learning on a particular view is speeded up when it is generalizing well and slowed down when it is overfitting. View-specific generalization/overfitting measures are computed on-the-fly during the training course and used to derive weights to blend the gradients from different views. As a result, the network is able to retain the representation power of different views in the joint features which represent the underlying distribution better than those learned by each individual view alone. Furthermore, the XSleepNet architecture is principally designed to gain robustness to the amount of training data and to increase the complementarity between the input views. Experimental results on five databases of different sizes show that XSleepNet consistently outperforms the single-view baselines and the multi-view baseline with a simple fusion strategy. Finally, XSleepNet also outperforms prior sleep staging methods and improves previous state-of-the-art results on the experimental databases. Huy Phan, Oliver Y. Chén, Minh C. Tran, Philipp Koch, Alfred Mertins, Maarten De Vos |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Spherical Harmonic Representation for Dynamic Sound-Field MeasurementsabstractContinuously moving microphones produce a high number of spatially dense sound-field samples with low effort in hardware and acquisition time. By interpreting the dynamic procedure as the non-uniform sampling of spatial basis functions, a system of linear equations can be set up. Its solution encodes sound-field parameters that allow for the spatio-temporal reconstruction within the measurement area at bandwidths where static methods would require impractical setups. An existing framework considers such basis functions from a signal processing point of view. It uses sinc-function based interpolation filters which are highly localized around sampled trajectories and may lead to ill-posed problems unless sparsity constraints are made, especially for locations that are away from microphone trajectories. In this paper, we present a new physical interpretation of the dynamic sampling problem. Transferring the problem into frequency domain, we describe samples of a moving microphone in terms of sampled spherical harmonic functions. The use of these global basis functions leads to dynamic measurements that inherently encode expanded sound-field information and, thus, allow for robust reconstruction at off-trajectory positions. Fabrice Katzberg, Marco Maaß, Alfred Mertins |
ICASSP | 3 |
| 2021 | Self-Attention Generative Adversarial Network for Speech EnhancementabstractExisting generative adversarial networks (GANs) for speech enhancement solely rely on the convolution operation, which may obscure temporal dependencies across the sequence input. To remedy this issue, we propose a self-attention layer adapted from non-local attention, coupled with the convolutional and deconvolutional layers of a speech enhancement GAN (SEGAN) using raw signal input. Further, we empirically study the effect of placing the self-attention layer at the (de)convolutional layers with varying layer indices as well as at all of them when memory allows. Our experiments show that introducing self-attention to SEGAN leads to consistent improvement across the objective evaluation metrics of enhancement performance. Furthermore, applying at different (de)convolutional layers does not significantly alter performance, suggesting that it can be conveniently applied at the highest-level (de)convolutional layer with the smallest memory overhead1. Huy Phan, Oliver Y. Chén, Philipp Koch, Ngoc Q. K. Duong, Ian McLoughlin 0001, Alfred Mertins |
ICASSP | 7 |
| 2021 | Multi-View Audio And Music ClassificationabstractWe propose in this work a multi-view learning approach for audio and music classification. Considering four typical low-level representations (i.e. different views) commonly used for audio and music recognition tasks, the proposed multi-view network consists of four subnetworks, each handling one input types. The learned embedding in the subnetworks are then concatenated to form the multi-view embedding for classification similar to a simple concatenation network. However, apart from the joint classification branch, the network also maintains four classification branches on the single-view embedding of the subnetworks. A novel method is then proposed to keep track of the learning behavior on the classification branches and adapt their weights to proportionally blend their gradients for network training. The weights are adapted in such a way that learning on a branch that is generalizing well will be encouraged whereas learning on a branch that is overfitting will be slowed down. Experiments on three different audio and music classification tasks show that the proposed multi-view network not only outperforms the single-view baselines but also is superior to the multi-view baselines based on concatenation and late fusion. Huy Phan, Oliver Y. Chén, Lam Dang Pham, Philipp Koch, Ian McLoughlin 0001, Alfred Mertins |
ICASSP | 7 |
| 2021 | Room Impulse Response Reshaping and Crosstalk Cancellation Using Convex OptimizationabstractIn this article, a new convex formulation for the acoustic channel-equalization problem is proposed and efficient ways for solving it are presented. Both the alternating direction method of multipliers and a proximal algorithm are studied for optimization. Time-domain and frequency-domain constraints are included in an equal way, allowing for simultaneous reverberation reduction and frequency-response equalization. Rather than aiming at full channel equalization, we follow the channel-reshaping paradigm proposed in earlier works and try to push the reverberation tail under the temporal masking limit of the human auditory system, thus making reverberation inaudible for human listeners. While the algorithm is presented for the multiple-input-multiple-output (MIMO) scenario, the single-channel case is included as a special case. Comparisons with methods from the literature for the single- and multi-channel cases show that the new algorithm has faster convergence than the known ones. The proximal algorithm even allows for solving extremely large problems with very low memory demand. Alfred Mertins, Marco Maaß, Fabrice Katzberg |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | CNN-MoE Based Framework for Classification of Respiratory Anomalies and Lung Disease DetectionabstractThis paper presents and explores a robust deep learning framework for auscultation analysis. This aims to classify anomalies in respiratory cycles and detect diseases, from respiratory sound recordings. The framework begins with front-end feature extraction that transforms input sound into a spectrogram representation. Then, a back-end deep learning network is used to classify the spectrogram features into categories of respiratory anomaly cycles or diseases. Experiments, conducted over the ICBHI benchmark dataset of respiratory sounds, confirm three main contributions towards respiratory-sound analysis. Firstly, we carry out an extensive exploration of the effect of spectrogram types, spectral-time resolution, overlapping/non-overlapping windows, and data augmentation on final prediction accuracy. This leads us to propose a novel deep learning system, built on the proposed framework, which outperforms current state-of-the-art methods. Finally, we apply a Teacher-Student scheme to achieve a trade-off between model performance and model complexity which holds promise for building real-time applications. Lam Pham, Huy Phan, Ramaswamy Palaniappan, Alfred Mertins, Ian McLoughlin 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Deep Feature Embedding and Hierarchical Classification for Audio Scene ClassificationabstractIn this work, we propose an approach that features deep feature embedding learning and hierarchical classification with triplet loss function for Acoustic Scene Classification (ASC). In the one hand, a deep convolutional neural network is firstly trained to learn a feature embedding from scene audio signals. Via the trained convolutional neural network, the learned embedding embeds an input into the embedding feature space and transforms it into a high-level feature vector for representation. In the other hand, in order to exploit the structure of the scene categories, the original scene classification problem is structured into a hierarchy where similar categories are grouped into meta-categories. Then, hierarchical classification is accomplished using deep neural network classifiers associated with triplet loss function. Our experiments show that the proposed system achieves good performance on both the DCASE 2018 Task 1A and 1B datasets, resulting in accuracy gains of 15.6% and 16.6% absolute over the DCASE 2018 baseline on Task 1A and 1B, respectively. Lam Dang Pham, Ian McLoughlin 0001, Huy Phan, Ramaswamy Palaniappan, Alfred Mertins |
IJCNN | 5 |
| 2020 | Improving GANs for Speech EnhancementabstractGenerative adversarial networks (GAN) have recently been shown to be efficient for speech enhancement. However, most, if not all, existing speech enhancement GANs (SEGAN) make use of a single generator to perform one-stage enhancement mapping. In this work, we propose to use multiple generators that are chained to perform multi-stage enhancement mapping, which gradually refines the noisy input signals in a stage-wise fashion. Furthermore, we study two scenarios: (1) the generators share their parameters and (2) the generators' parameters are independent. The former constrains the generators to learn a common mapping that is iteratively applied at all enhancement stages and results in a small model footprint. On the contrary, the latter allows the generators to flexibly learn different enhancement mappings at different stages of the network at the cost of an increased model size. We demonstrate that the proposed multi-stage enhancement approach outperforms the one-stage SEGAN baseline, where the independent generators lead to more favorable results than the tied generators. The source code is available at http://github.com/pquochuy/idsegan. Huy Phan, Ian McLoughlin 0001, Lam Dang Pham, Oliver Y. Chén, Philipp Koch, Maarten De Vos, Alfred Mertins |
IEEE Signal Process. Lett. | 7 |
| 2019 | Forked Recurrent Neural Network for Hand Gesture Classification Using Inertial Measurement DataabstractFor many applications of hand gesture recognition, a delay-free, affordable, and mobile system relying on body signals is mandatory. Therefore, we propose an approach for hand gestures classification given signals of inertial measurement units (IMUs) that works with extremely short windows to avoid delays. With a simple recurrent neural network the suitability of the sensor modalities of an IMU (accelerometer, gyroscope, magnetometer) are evaluated by only providing data of one modality. For the multi-modal data a second network with mid-level fusion is proposed. Its forked architecture allows us to process data of each modality individually before carrying out a joint analysis for classification. Experiments on three databases reveal that even when relying on a single modality our proposed system outperforms state-of-the-art systems significantly. With the forked network classification accuracy can be further improved by over 10 % absolute compared to the best reported system while causing a fraction of the delay. Philipp Koch, Nele Sophie Brügge, Huy Phan, Marco Maaß, Alfred Mertins |
ICASSP | 5 |
| 2019 | Robust Room Equalization Using Sparse Sound-field ReconstructionabstractThe perceived quality of a sound played in a closed room is often degraded by the added reverberation. In order to combat these effects, the methods of room impulse response equalization may be used. Typically, properties of the human auditory system, such as temporal masking, are used for a better control of the late echoes. There, a prefilter is used to modify the played signal, which renders the echoes inaudible for a given position. When targeting a human listener who is always performing small movement, the method fails due to the spatial mismatch. In such a case, a whole volume needs to be equalized, which requires a huge amount of measurements. In this work, we propose to reduce this burden by measuring only a random subset of the required room impulse responses and reconstruct a grid employing sparse methods. By further interpolation an equalization of the whole target area is achieved. Radoslaw Mazur, Fabrice Katzberg, Alfred Mertins |
ICASSP | 3 |
| 2019 | Unifying Isolated and Overlapping Audio Event Detection with Multi-label Multi-task Convolutional Recurrent Neural NetworksabstractWe propose a multi-label multi-task framework based on a convolutional recurrent neural network to unify detection of isolated and overlapping audio events. The framework leverages the power of convolutional recurrent neural network architectures; convolutional layers learn effective features over which higher recurrent layers perform sequential modelling. Furthermore, the output layer is designed to handle arbitrary degrees of event overlap. At each time step in the recurrent output sequence, an output triple is dedicated to each event category of interest to jointly model event occurrence and temporal boundaries. That is, the network jointly determines whether an event of this category occurs, and when it occurs, by estimating onset and offset positions at each recurrent time step. We then introduce three sequential losses for network training: multi-label classification loss, distance estimation loss, and confidence loss. We demonstrate good generalization on two datasets: ITC-Irst for isolated audio event detection, and TUT-SED-Synthetic-2016 for overlapping audio event detection. Huy Phan, Oliver Y. Chén, Philipp Koch, Lam Dang Pham, Ian McLoughlin 0001, Alfred Mertins, Maarten De Vos |
ICASSP | 6 |
| 2019 | Spatio-Temporal Attention Pooling for Audio Scene ClassificationabstractAcoustic scenes are rich and redundant in their content. In this work, we present a spatio-temporal attention pooling layer coupled with a convolutional recurrent neural network to learn from patterns that are discriminative while suppressing those that are irrelevant for acoustic scene classification. The convolutional layers in this network learn invariant features from time-frequency input. The bidirectional recurrent layers are then able to encode the temporal dynamics of the resulting convolutional features. Afterwards, a two-dimensional attention mask is formed via the outer product of the spatial and temporal attention vectors learned from two designated attention layers to weigh and pool the recurrent output into a final feature vector for classification. The network is trained with between-class examples generated from between-class data augmentation. Experiments demonstrate that the proposed method not only outperforms a strong convolutional neural network baseline but also sets new state-of-the-art performance on the LITIS Rouen dataset. Huy Phan, Oliver Y. Chén, Lam Dang Pham, Philipp Koch, Maarten De Vos, Ian McLoughlin 0001, Alfred Mertins |
INTERSPEECH | 7 |
| 2018 | Compressive Sampling of Sound Fields Using Moving MicrophonesabstractFor conventional sampling of sound-fields, the measurement in space by use of stationary microphones is impractical for high audio frequencies. Satisfying the Nyquist-Shannon sampling theorem requires a huge number of sampling points and entails other difficulties, such as the need for exact calibration and spatial positioning of a large number of microphones. Dynamic sound-field measurements involving tracked microphones may weaken this spatial sampling problem. However, for aliasing-free reconstruction, there is still the need of sampling a huge number of unknown sound-field variables. Thus in real-world applications, the trajectories may be expected to lead to underdetermined sampling problems. In this paper, we present a compressed sensing framework that allows for stable and robust sub-Nyquist sampling of sound fields by use of moving microphones. Fabrice Katzberg, Radoslaw Mazur, Marco Maaß, Philipp Koch, Alfred Mertins |
ICASSP | 5 |
| 2018 | Weighted and Multi-Task Loss for Rare Audio Event DetectionabstractWe present in this paper two loss functions tailored for rare audio event detection in audio streams. The weighted loss is designed to tackle the common issue of imbalanced data in background/foreground classification while the multi-task loss enables the networks to simultaneously model the class distribution and the temporal structures of the target events for recognition. We study the proposed loss functions with deep neural networks (DNNs) and convolutional neural networks (CNNs) coupled with state-of-the-art phase-aware signal enhancement. Experiments on the DCASE 2017 challenge's data show that our system with the proposed losses significantly outperforms not only the DCASE 2017 baseline but also our baseline which has a similar network architecture and a standard loss function. Huy Phan, Martin Krawczyk-Becker, Timo Gerkmann, Alfred Mertins |
ICASSP | 4 |
| 2018 | Enabling Early Audio Event Detection with Neural NetworksabstractThis paper presents a methodology for early detection of audio events from audio streams. Early detection is the ability to infer an ongoing event during its initial stage. The proposed system consists of a novel inference step coupled with dual parallel tailored-loss deep neural networks (DNNs). The DNNs share a similar architecture except for their loss functions, i.e. weighted loss and multitask loss, which are designed to efficiently cope with issues common to audio event detection. The inference step is newly introduced to make use of the network outputs for recognizing ongoing events. The monotonicity of the detection function is required for reliable early detection, and will also be proved. Experiments on the ITC-Irst database show that the proposed system achieves state-of-the-art detection performance. Furthermore, even partial events are sufficient to achieve good performance similar to that obtained when an entire event is observed, enabling early event detection. Huy Phan, Philipp Koch, Ian McLoughlin 0001, Alfred Mertins |
ICASSP | 4 |
| 2018 | Boosting Black-Box Variational Inference by Incorporating the Natural GradientabstractIn this paper we present a modification of the popular Black-Box Variational Inference (BBVI) approach which significantly improves the computational efficiency of the inference. We achieve this performance boost by replacing the standard gradient in the stochastic gradient ascent framework of BBVI with the natural gradient. Our experimental results (e.g. training of neutral networks) show that the proposed method outperforms the original BBVI algorithm on both synthetic and real data. Felix Trusheim, Alexandru Condurache, Alfred Mertins |
ICPR | 3 |
| 2018 | A Compressed Sensing Framework for Dynamic Sound-Field MeasurementsabstractThe conventional sampling of sound fields by use of stationary microphones is impractical for large bandwidths. Satisfying the Nyquist-Shannon sampling theorem in three-dimensional space requires a huge number of sampling positions. Dynamic sound-field measurements with moving microphones together with a compressed-sensing recovery allow for weakening the spatial sampling problem. For bandlimited signals, the dynamic samples taken along the microphone trajectory may be related to the room impulse responses on a virtual grid in space via spatial interpolation. The tracking of the microphone positions and the knowledge of the excitation sequence allow for setting up a linear system of equations that can be solved for the room impulse responses on the modeled virtual grid. Nevertheless, there is still the necessity for recovering a huge number of sound-field variables, in order to ensure aliasing-free reconstruction. Thus, for practical applications, random or suboptimally chosen trajectories may be expected to lead to underdetermined sampling problems for a given volume of interest. In this paper, we present a compressed sensing framework that enables us to uniquely solve the dynamic sampling problem despite having underdetermined variables. The spatio-temporal sampling problem is integrated into compressed sensing models that allow for stable and robust sub-Nyquist sampling given incoherent measurements. For a modeled equidistant grid and sparse Fourier representations, the influence of the microphone trajectories on the compressed sensing problem is investigated and a simple expression is derived for evaluating trajectories with regard to compressed-sensing based recovery. Fabrice Katzberg, Radoslaw Mazur, Marco Maaß, Philipp Koch, Alfred Mertins |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Visual Landmark Based 3D Road Course Estimation with Black Box Variational Inference
Felix Trusheim, Alexandru Condurache, Alfred Mertins |
CAIP (1) | 3 |
| 2017 | Measurement of sound fields using moving microphonesabstractThe sampling of sound fields involves the measurement of spatially dependent room impulse responses, where the Nyquist-Shannon sampling theorem applies in both the temporal and spatial domains. Therefore, sampling inside a volume of interest requires a huge number of sampling points in space, which comes along with further difficulties such as exact microphone positioning and calibration of multiple microphones. In this paper, we present a method for measuring sound fields using moving microphones whose trajectories are known to the algorithm. At that, the number of microphones is customizable by trading measurement effort against sampling time. Through spatial interpolation of the dynamic measurements, a system of linear equations is set up which allows for the reconstruction of the entire sound field inside the volume of interest. Fabrice Katzberg, Radoslaw Mazur, Marco Maaß, Philipp Koch, Alfred Mertins |
ICASSP | 5 |
| 2017 | CNN-LTE: A class of 1-X pooling convolutional neural networks on label tree embeddings for audio scene classificationabstractWe present in this work an approach for audio scene classification. Firstly, given the label set of the scenes, a label tree is automatically constructed where the labels are grouped into meta-classes. This category taxonomy is then used in the feature extraction step in which an audio scene instance is transformed into a label tree embedding image. Elements of the image indicate the likelihoods that the scene instances belong to different meta-classes. A class of simple 1-X (i.e. 1-max, 1-mean, and 1-mix) pooling convolutional neural networks, which are tailored for the task at hand, are finally learned on top of the image features for scene recognition. Experimental results on the DCASE 2013 and DCASE 2016 datasets demonstrate the efficiency of the proposed method. Huy Phan, Philipp Koch, Lars Hertel, Marco Maaß, Radoslaw Mazur, Alfred Mertins |
ICASSP | 6 |
| 2017 | Audio Scene Classification with Deep Recurrent Neural NetworksabstractWe introduce in this work an efficient approach for audio scene classification using deep recurrent neural networks. An audio scene is firstly transformed into a sequence of high-level label tree embedding feature vectors. The vector sequence is then divided into multiple subsequences on which a deep GRU-based recurrent neural network is trained for sequence-to-label classification. The global predicted label for the entire sequence is finally obtained via aggregation of subsequence classification outputs. We will show that our approach obtains an F1-score of 97.7% on the LITIS Rouen dataset, which is the largest dataset publicly available for the task. Compared to the best previously reported result on the dataset, our approach is able to reduce the relative classification error by 35.3%. Huy Phan, Philipp Koch, Fabrice Katzberg, Marco Maaß, Radoslaw Mazur, Alfred Mertins |
INTERSPEECH | 6 |
| 2017 | Improved Audio Scene Classification Based on Label-Tree Embeddings and Convolutional Neural NetworksabstractIn this paper, we present an efficient approach for audio scene classification. We aim at learning representations for scene examples by exploring the structure of their class labels. A category taxonomy is automatically learned by collectively optimizing a tree-structured clustering of the given labels into multiple metaclasses. A scene recording is then transformed into a label-tree embedding image. Elements of the image represent the likelihoods that the scene instance belongs to the metaclasses. We investigate classification with label-tree embedding features learned from different low-level features as well as their fusion. We show that the combination of multiple features is essential to obtain good performance. While averaging label-tree embedding images over time yields good performance, we argue that average pooling possesses an intrinsic shortcoming. We alternatively propose an improved classification scheme to bypass this limitation. We aim at automatically learning common templates that are useful for the classification task from these images using simple but tailored convolutional neural networks. The trained networks are then employed as a feature extractor that matches the learned templates across a label-tree embedding image and produce the maximum matching scores as features for classification. Since audio scenes exhibit rich content, template learning and matching on low-level features would be inefficient. With label-tree embedding features, we have quantized and reduced the low-level features into the likelihoods of the metaclasses, on which the template learning and matching are efficient. We study both training convolutional neural networks on stacked label-tree embedding images and multistream networks. Experimental results on the DCASE2016 and LITIS Rouen datasets demonstrate the efficiency of the proposed methods. Huy Phan, Lars Hertel, Marco Maaß, Philipp Koch, Radoslaw Mazur, Alfred Mertins |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2016 | Learning compact structural representations for audio events using regressor banksabstractWe introduce a new learned descriptor for audio signals which is efficient for event representation. The entries of the descriptor are produced by evaluating a set of regressors on the input signal. The regressors are class-specific and trained using the random regression forests framework. Given an input signal, each regressor estimates the onset and offset positions of the target event. The estimation confidence scores output by a regressor are then used to quantify how the target event aligns with the temporal structure of the corresponding category. Our proposed descriptor has two advantages. First, it is compact, i.e. the dimensionality of the descriptor is equal to the number of event classes. Second, we show that even simple linear classification models, trained on our descriptor, yield better accuracies on audio event classification task than not only the nonlinear baselines but also the state-of-the-art results. Huy Phan, Marco Maaß, Lars Hertel, Radoslaw Mazur, Ian McLoughlin 0001, Alfred Mertins |
ICASSP | 6 |
| 2016 | Comparing time and frequency domain for audio event recognition using deep learningabstractRecognizing acoustic events is an intricate problem for a machine and an emerging field of research. Deep neural networks achieve convincing results and are currently the state-of-the-art approach for many tasks. One advantage is their implicit feature learning, opposite to an explicit feature extraction of the input signal. In this work, we analyzed whether more discriminative features can be learned from either the time-domain or the frequency-domain representation of the audio signal. For this purpose, we trained multiple deep networks with different architectures on the Freiburg-106 and ESC-10 datasets. Our results show that feature learning from the frequency domain is superior to the time domain. Moreover, additionally using convolution and pooling layers, to explore local structures of the audio signal, significantly improves the recognition performance and achieves state-of-the-art results. Lars Hertel, Huy Phan, Alfred Mertins |
IJCNN | 3 |
| 2016 | Robust Audio Event Recognition with 1-Max Pooling Convolutional Neural NetworksabstractWe present in this paper a simple, yet efficient convolutional neural network (CNN) architecture for robust audio event recognition. Opposing to deep CNN architectures with multiple convolutional and pooling layers topped up with multiple fully connected layers, the proposed network consists of only three layers: convolutional, pooling, and softmax layer. Two further features distinguish it from the deep architectures that have been proposed for the task: varying-size convolutional filters at the convolutional layer and 1-max pooling scheme at the pooling layer. In intuition, the network tends to select the most discriminative features from the whole audio signals for recognition. Our proposed CNN not only shows state-of-the-art performance on the standard task of robust audio event recognition but also outperforms other deep architectures up to 4.5% in terms of recognition accuracy, which is equivalent to 76.3% relative error reduction. Huy Phan, Lars Hertel, Marco Maaß, Alfred Mertins |
INTERSPEECH | 4 |
| 2016 | Label Tree Embeddings for Acoustic Scene ClassificationabstractWe present in this paper an efficient approach for acoustic scene classification by exploring the structure of class labels. Given a set of class labels, a category taxonomy is automatically learned by collectively optimizing a clustering of the labels into multiple meta-classes in a tree structure. An acoustic scene instance is then embedded into a low-dimensional feature representation which consists of the likelihoods that it belongs to the meta-classes. We demonstrate state-of-the-art results on two different datasets for the acoustic scene classification task, including the DCASE 2013 and LITIS Rouen datasets. Huy Phan, Lars Hertel, Marco Maaß, Philipp Koch, Alfred Mertins |
ACM Multimedia | 5 |
| 2016 | Error performance and energy efficiency analyses of fully cooperative OFDM communication in frequency selective fadingabstractThis study devises for the first time a comprehensive analysis of decode‐and‐forward, space–time coded, fully cooperative orthogonal frequency division multiplexing (OFDM) systems from both error performance and energy efficiency perspectives, in either identically or non‐identically distributed frequency selective Rayleigh fading channels. The authors’ analyses show the interesting result that the fully cooperative OFDM is better than direct OFDM transmission from both perspectives in many cases, especially at a low power and low signal‐to‐noise ratio regime, making it particularly useful for emerging low power applications, such as wireless body area networks. Le Chung Tran, Alfred Mertins |
IET Commun. | 2 |
| 2016 | Learning Representations for Nonspeech Audio Events Through Their Similarities to Speech PatternsabstractThe human auditory system is very well matched to both human speech and environmental sounds. Therefore, the question arises whether human speech material may provide useful information for training systems for analyzing nonspeech audio signals, e.g., in a classification task. In order to answer this question, we consider speech patterns as basic acoustic concepts, which embody and represent the target nonspeech signal. To find out how similar the nonspeech signal is to speech, we classify it with a classifier trained on the speech patterns and use the classification posteriors to represent the closeness to the speech bases. The speech similarities are finally employed as a descriptor to represent the target signal. We further show that a better descriptor can be obtained by learning to organize the speech categories hierarchically with a tree structure. Furthermore, these descriptors are generic. That is, once the speech classifier has been learned, it can be employed as a feature extractor for different datasets without retraining. Lastly, we propose an algorithm to select a sufficient subset, which provides an approximate representation capability of the entire set of available speech patterns. We conduct experiments for the application of audio event analysis. Phone triplets from the TIMIT dataset were used as speech patterns to learn the descriptors for audio events of three different datasets with different complexity, including UPC-TALP, Freiburg-106, and NAR. The experimental results on the event classification task show that a good performance can be easily obtained even if a simple linear classifier is used. Furthermore, fusion of the learned descriptors as an additional source leads to state-of-the-art performance on all the three target datasets. Huy Phan, Lars Hertel, Marco Maaß, Radoslaw Mazur, Alfred Mertins |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2015 | Cosine-Sine Modulated Filter Banks for Motion Estimation and Correction
Marco Maaß, Huy Phan, Anita Möller, Alfred Mertins |
ACIVS | 4 |
| 2015 | Joint time- and frequency-domain reshaping of room impulse responsesabstractIn listening room compensation, the aim is to compensate for the degradations that are rendered to an audio signal by transmission in a closed room. Due to multiple reflections of the soundwaves, the listener receives a superposition of delayed and attenuated versions of the source signal. A filter is designed so that the convolution of the room impulse response and the equalizer contains better acoustic properties than the original acoustic channel. For dereverberation, the filter is usually designed by optimizing a solely time-domain based cost function and, therefore, may introduce spectral distortions. Recent methods also consider the frequency-domain representation of the overall system and aim at yielding a flat overall frequency response. However, in some cases, it may be desirable to follow a predefined frequency response rather than obtaining a flat one. Hearing impaired persons, for example, may prefer or even need an amplification of certain frequency ranges. In this work, we propose a new method to jointly shape the time- and the frequency-domain representations of the overall acoustic channel according to prescribed curves. Furthermore, we integrate the concept of auditory scales into the filter design. Jan Ole Jungmann, Radoslaw Mazur, Alfred Mertins |
ICASSP | 3 |
| 2015 | Early event detection in audio streamsabstractAudio event detection has been an active field of research in recent years. However, most of the proposed methods, if not all, analyze and detect complete events and little attention has been paid for early detection. In this paper, we present a system which enables early audio event detection in continuous audio recordings in which an event can be reliably recognized when only a partial duration is observed. Our evaluation on the ITC-Irst database, one of the standard database of the CLEAR 2006 evaluation, shows that: on one hand, the proposed system outperforms the best baseline system by 16% and 8% in terms of detection error rate and detection accuracy respectively; on the other hand, even partial events are enough to achieve the performance that is obtainable when the whole events are observed. Huy Phan, Marco Maaß, Radoslaw Mazur, Alfred Mertins |
ICME | 4 |
| 2015 | Representing nonspeech audio signals through speech classification modelsabstractThe human auditory system is very well matched to both human speech and environmental sounds. Therefore, the question arises whether human speech material may provide useful information for training systems for analyzing nonspeech audio signals, such as in a recognition task. To find out how similar nonspeech signals are to speech, we measure the closeness between target nonspeech signals and different basis speech categories via a speech classification model. The speech similarities are finally employed as a descriptor to represent the target signal. We further show that a better descriptor can be obtained by learning to organize the speech categories hierarchically with a tree structure. We conduct experiments for the audio event analysis application by using speech words from the TIMIT dataset to learn the descriptors for the audio events of the Freiburg-106 dataset. Our results on the event recognition task outperform those achieved by the best system even though a simple linear classifier is used. Furthermore, integrating the learned descriptors as an additional source leads to improved performance. Huy Phan, Lars Hertel, Marco Maaß, Radoslaw Mazur, Alfred Mertins |
INTERSPEECH | 5 |
| 2015 | Random Regression Forests for Acoustic Event Detection and ClassificationabstractDespite the success of the automatic speech recognition framework in its own application field, its adaptation to the problem of acoustic event detection has resulted in limited success. In this paper, instead of treating the problem similar to the segmentation and classification tasks in speech recognition, we pose it as a regression task and propose an approach based on random forest regression. Furthermore, event localization in time can be efficiently handled as a joint problem. We first decompose the training audio signals into multiple interleaved superframes which are annotated with the corresponding event class labels and their displacements to the temporal onsets and offsets of the events. For a specific event category, a random-forest regression model is learned using the displacement information. Given an unseen superframe, the learned regressor will output the continuous estimates of the onset and offset locations of the events. To deal with multiple event categories, prior to the category-specific regression phase, a superframe-wise recognition phase is performed to reject the background superframes and to classify the event superframes into different event categories. While jointly posing event detection and localization as a regression problem is novel, the superior performance on two databases ITC-Irst and UPC-TALP demonstrates the efficiency and potential of the proposed approach. Huy Phan, Marco Maaß, Radoslaw Mazur, Alfred Mertins |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Joint time-domain reshaping and frequency-domain equalization of room impulse responsesabstractIn listening room compensation, the aim is to compensate for the degradations that are rendered to an audio signal by transmission in a closed room. Due to multiple reflections of the soundwaves, the listener receives a superposition of delayed and attenuated versions of the source signal. A filter is designed so that the convolution of the room impulse response and the equalizer contains better acoustic properties than the original acoustic channel. Common approaches for derever-beration optimize only the time-domain representation of the overall impulse response and may introduce distortions in the frequency domain. Equalization of the frequency response, on the other hand, often does not consider the time-domain behavior in an ideal way. In this paper, we propose a novel method to jointly consider both the time- and frequency-domain behavior. It outperforms the methods known from literature in terms of dereverberation and equalization performance. Results are presented for a room impulse response measured in a real living room. Jan Ole Jungmann, Radoslaw Mazur, Alfred Mertins |
ICASSP | 3 |
| 2014 | Acoustic event detection and localization with regression forestsabstractThis paper proposes an approach for the efficient automatic joint detection and localization of single-channel acoustic events us-ing random forest regression. The audio signals are decom-posed into multiple densely overlapping superframes annotated with event class labels and their displacements to the temporal starting and ending points of the events. Using the displacement information, a multivariate random forest regression model is learned for each event category to map each superframe to con-tinuous estimates of onset and offset locations of the events. In addition, two classifiers are trained using random forest clas-sification to classify superframes of background and different event categories. On testing, based on the detection of category-specific superframes using the classifiers, the learned regressor provides the estimates of onset and offset locations in time of the corresponding event. While posing event detection and lo-calization as a regression problem is novel, the quantitative eval-uation on ITC-Irst database of highly variable acoustic events shows the efficiency and potential of the proposed approach. Index Terms: acoustic event detection, regression forest, ran-dom forest, superframe Huy Phan, Marco Maaß, Radoslaw Mazur, Alfred Mertins |
INTERSPEECH | 4 |
| 2013 | Perturbation of room impulse responses and its application in robust listening room compensationabstractThe purpose of room impulse response reshaping is to reduce reverberation and thus to improve the perceived quality of the received signal by prefiltering the source signal before it is played with a loudspeaker. The filter design is usually carried out by solving an optimization problem. There are, in general, two possibilities to improve the robustness of the equalizers against small movements of the listener and/or receiver; namely multi-position approaches or the utilization of a regularization term. Multi-position approaches suffer from the extensive effort of measuring multiple room impulse responses. Stochastic models may describe the average system error due to spatial mismatch, but only quadratic penalty terms have been considered so far. In this contribution we propose a third method to improve robustness against spatial misalignment. We combine the two approaches by generating multiple realizations of distorted room impulse responses and feeding them into the multiposition algorithm. Based on our previous work, we propose a model to capture the perturbations with respect to the assumed displacement. Jan Ole Jungmann, Radoslaw Mazur, Alfred Mertins |
ICASSP | 3 |
| 2013 | Feature extraction with a multiscale modulation analysis for robust automatic speech recognitionabstractIn this work we present a new feature extraction method that is robust against the effects of varying vocal tract lengths. The principle of the method is based on invariant integration and makes use of a modulation filtering approach, similar to the recently proposed scattering transform. In particular, we show how the transform can be used to obtain features that are robust against variations of the vocal tract length. Phoneme recognition experiments show a clearly increased robustness in case of mismatching average vocal tract lengths. Florian Müller 0002, Alfred Mertins |
ICASSP | 2 |
| 2013 | Accelerated Nonlinear Gaussianization for Feature Extraction
Alexandru Condurache, Alfred Mertins |
ICPRAM | 2 |
| 2013 | Unitary Differential Space-Time-Frequency Codes for MB-OFDM UWB Wireless CommunicationsabstractIn a multiple-input multiple-output (MIMO), multiband orthogonal frequency division multiplexing (MB-OFDM) ultra-wideband (UWB) system, coherent detection requires the transmission of a large number of symbols for channel estimation, thus reducing the bandwidth efficiency. For the first time, this paper proposes unitary differential space-time-frequency codes (DSTFCs) for MB-OFDM UWB communications, which increase the system bandwidth efficiency because no channel state information (CSI) is required. The proposed system would be useful when CSI is unavailable at the receiver, such as when the transmission of multiple channel estimation symbols is impractical or uneconomical. The coding and decoding algorithms for the proposed DSTFCs are then derived for both constant envelope modulation scheme, such as PSK (phase shift keying) and 4QAM (quadrature amplitude modulation), and multi-dimensional modulation scheme, such as DCM (dual carrier modulation). The paper also quantifies for the first time the diversity order of a DSTFC MB-OFDM system. Simulation results show that the application of DSTFCs can significantly improve the bit error performance of conventional differential MB-OFDM system (without MIMO), and even provide much better bit error performance than the conventional coherent MB-OFDM system (without MIMO) at high signal-to-noise ratios. Le Chung Tran, Alfred Mertins, Tadeusz A. Wysocki |
IEEE Trans. Wirel. Commun. | 2 |
| 2012 | Optimized gradient calculation for room impulse response reshaping algorithm based on p-norm optimizationabstractBy using room impulse response shortening and reshaping it is possible to reduce the reverberation effects and therefore improve the perceived quality. This may be achieved by a prefilter that modifies the overall impulse response to have a faster decay. The traditional filter shortening approach using least-squares methods is fast and directly computable, but it suffers from late echoes. Newer approaches using the p-norm overcome this drawback but are computationally very demanding, as the optimization process uses a gradient-descent approach with slow convergence. In this work we propose a modification to this approach that results in a significantly faster convergence. With this modification, the algorithm is less likely to be trapped in a local minimum and therefore also leads to a better convergence point. The method will be demonstrated on simulated and real-world room impulse responses. Radoslaw Mazur, Jan Ole Jungmann, Alfred Mertins |
ICASSP | 3 |
| 2012 | On using the auditory image model and invariant-integration for noise robust automatic speech recognitionabstractCommonly used feature extraction methods for automatic speech recognition (ASR) incorporate only rudimentary psychoacoustic findings. Several works showed that a physiologically closer auditory processing during the feature extraction stage can enhance the robustness of an ASR system in noisy environments. The “auditory image model” (AIM) is such a more sophisticated computational model. In this work we show how invariant integration can be applied in the feature space given by the AIM, and we analyze the performance of the resulting features under noisy conditions on the Aurora-2 task. Furthermore, we show that previously presented features based on power-normalization and invariant integration benefit from the AIM-based integration features when the feature vectors are combined with each other. Florian Müller 0002, Alfred Mertins |
ICASSP | 2 |
| 2012 | Event Detection using Log-linear Models for Coronary Contrast Agent Injections
Dierck Matern, Alexandru Condurache, Alfred Mertins |
ICPRAM (2) | 3 |
| 2012 | Enhancing Vocal Tract Length Normalization with Elastic Registration for Automatic Speech RecognitionabstractVocal tract length normalization (VTLN) is commonly applied utterance-wise with a warping function that makes the assumption of a linear dependence between the vocal tract length and the location of the formants. In this work we propose a datadriven method for enhancing the performance of systems that already use standard VTLN. The method is based on elastic registration to estimate optimal non-parametric transformations to further reduce inter-speaker variabilities. Results show that the proposed method can increase the performance of monophone systems such that it reaches that of a triphone system. Index Terms: automatic speech recognition, vocal tract length normalization, elastic registration 1. Florian Müller 0002, Alfred Mertins |
INTERSPEECH | 2 |
| 2012 | Feature Analysis for Parkinson's Disease Detection Based on Transcranial Sonography Image
Lei Chen 0012, Johann Hagenah, Alfred Mertins |
MICCAI (3) | 3 |
| 2012 | Combined Acoustic MIMO Channel Crosstalk Cancellation and Room Impulse Response ReshapingabstractVirtual 3-D sound can be easily delivered to a listener by binaural audio signals that are reproduced via headphones, which guarantees that only the correct signals reach the corresponding ears. Reproducing the binaural audio signal by two or more loudspeakers introduces the problems of crosstalk on the one hand, and, of reverberation on the other hand. In crosstalk cancellation, the audio signals are fed through a network of prefilters prior to loudspeaker reproduction to ensure that only the designated signal reaches the corresponding ear of the listener. Since room impulse responses are very sensitive to spatial mismatch, and since listeners might slightly move while listening, robust designs are needed. In this paper, we present a method that jointly handles the three problems of crosstalk, reverberation reduction, and spatial robustness with respect to varying listening positions for one or more binaural source signals and multiple listeners. The proposed method is based on a multichannel room impulse response reshaping approach by optimizing a -norm based criterion. Replacing the well-known least-squares technique by a -norm based method employing a large value for allows us to explicitly control the amount of crosstalk and to shape the remaining reverberation effects according to a desired decay. Jan Ole Jungmann, Radoslaw Mazur, Markus Kallinger, Tiemin Mei, Alfred Mertins |
IEEE Trans. Speech Audio Process. | 5 |
| 2011 | Multi image super resolution using compressed sensingabstractIn this paper we present a new compressed sensing model and reconstruction method for multi-detector signal acquisition. We extend the concept of the famous single-pixel camera to a multi-detector device with the benefit of reducing measurement time, while still providing resolution enhancement and deblurring. We provide a scalable model which allows the trade off between system complexity (number of detectors) and time (number of measurements). Torsten Edeler, Kevin Ohliger, Stephan Hussmann, Alfred Mertins |
ICASSP | 4 |
| 2011 | A sparsity based criterion for solving the permutation ambiguity in convolutive blind source separationabstractIn this paper, we present a new algorithm for solving the permutation ambiguity in convolutive blind source separation. A common approach for separation of convolutive mixtures is the transformation to the time-frequency domain, where the convolution becomes a multiplication. This allows for the use of well-known instantaneous ICA algorithms independently in each frequency bin. However, this simplification leads to the problem of correctly aligning these single bins previously to the transformation to the time domain. Here, we propose a new criterion for solving this ambiguity. The new approach is based on the sparsity of the speech signals and yields a robust depermutation algorithm. The results will be shown on real world examples. Radoslaw Mazur, Alfred Mertins |
ICASSP | 2 |
| 2011 | Noise Robust Speaker-Independent Speech Recognition with Invariant-Integration Features Using Power-Bias SubtractionabstractThis paper presents new results about the robustness of invariantintegration features (IIF) in noisy conditions. Furthermore, it is shown that a feature-enhancement method known as “powerbias subtraction ” for noisy conditions can be combined with the IIF approach to improve its performance in noisy environments while keeping the robustness of the IIFs to mismatching vocaltract length training-testing conditions. Results of experiments with training on clean speech only as well as experiments with matched-condition training are presented. Index Terms: speech recognition, speaker independency, noise robustness, invariant integration, power normalization Florian Müller 0002, Alfred Mertins |
INTERSPEECH | 2 |
| 2011 | Contextual invariant-integration features for improved speaker-independent speech recognition
Florian Müller 0002, Alfred Mertins |
Speech Commun. | 2 |
| 2011 | Elastic-Transform Based Multiclass GaussianizationabstractThe concept of “Gaussianization” implies a transformation aimed at changing the distribution of the input random variable to Gaussian. It has been used until now as a means to achieve independence among components in multivariate distributions that in turn was used as a tool for various purposes ranging from density estimation to normalization. In this contribution we propose Gaussianization for pattern recognition applications in support of Gaussianity assumptions made by various classifiers. Previous approaches completely ignore separability considerations, the Gaussianization being conducted over the entire data, irrespective of class affiliation and are not useful for recognition purposes. We instead propose a transform such that the output random variable is distributed according to a Gaussian mixture, where each class accounts for one mixture component. We successfully test our method on both synthetic and real data. Alexandru Condurache, Alfred Mertins |
IEEE Signal Process. Lett. | 2 |
| 2011 | A Dynamic Fine-Grain Scalable Compression Scheme With Application to Progressive Audio CodingabstractThis paper studies the fine-grain scalable compression problem with emphasis on 1-D signals such as audio signals. Like in the successful 2-D still image compression techniques embedded zerotree wavelet coder (EZW) and set partitioning in hierarchical trees (SPIHT), the desired fine-granular scalability and high coding efficiency are benefited from a tree-based significance mapping technique. A significance tree serves to quickly locate and efficiently encode the important coefficients in the transform domain. The aim of this paper is to find such suitable significance trees for compressing dynamically variant 1-D signals. The proposed solution is a novel dynamic significance tree (DST) where, unlike in existing solutions with a single type of tree, a significance tree is chosen dynamically out of a set of trees by taking into account the actual coefficients distribution. We show how a set of possible DSTs can be derived that is optimized for a given (training) dataset. The method outperforms the existing scheme for lossy audio compression based on a single-type tree (SPIHT) and the scalable audio coding schemes MPEG-4 BSAC and MPEG-4 SLS. For bitrates less than 32 kbps, it results in an improved perceived audio quality compared to the fixed-bitrate MPEG-2/4 AAC audio coding scheme while providing progressive transmission and finer scalability. Stefan Strahl, Heiko Hansen, Alfred Mertins |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Quality assessment for listening-room compensation algorithmsabstractIn this contribution various objective measures that can be used to evaluate speech dereverberation algorithms by means of listening-room compensation (LRC) are compared to subjective listening tests. It is shown that technical measures describing the impulse responses are suitable for evaluation of such algorithms. Most signal-based objective measures fail to judge the specific distortions that may be introduced by LRC algorithms like late reverberation since these artifacts are small in amplitude but perceptually relevant due to the loss of masking of the room impulse response. Only one signal-based measure, the so-called perceptual similarity measure (PSM), showed high correlation with subjective rating for the given test setup. Stefan Goetze, Eugen Albertin, Markus Kallinger, Alfred Mertins, Karl-Dirk Kammeyer |
ICASSP | 4 |
| 2010 | Multiple feature extraction for early Parkinson risk assessment based on transcranial sonography imageabstractTranscranial sonography (TCS) is a new tool for the diagnosis of Parkinson's disease (PD) at a very early state. The TCS image of mesencephalon shows a distinct hyperechogenic pattern in about 90% PD patients. This pattern is usually manually segmented and the substantia nigra (SN) region can be used as an early PD indicator. However this method is based on manual evaluation of examined images. The extraction of multiple features from TCS images characterizing the half mesencephalon morphology and structure can be used to validate the observer-independent PD indicator. We propose hybrid feature extraction methods which includes statistical, geometrical and texture features for the early PD risk assessment. These features are tested with support vector machines (SVMs). Furthermore five features are selected with the sequential feature selection methods. The results show that the correct rate of the classification with these five features is reaching 96%. Lei Chen 0012, Günter Seidel, Alfred Mertins |
ICIP | 3 |
| 2010 | An LDA-based Relative Hysteresis Classifier with Application to Segmentation of Retinal VesselsabstractIn a pattern classification setup, image segmentation is achieved by assigning each pixel to one of two classes: object or background. The special case of vessel segmentation is characterized by a strong disproportion between the number of representatives of each class (i.e. class skew) and also by a strong overlap between classes. These difficulties can be solved using problem-specific knowledge. The proposed hysteresis classification makes use of such knowledge in an efficient way. We describe a novel, supervised, hysteresis-based classification method that we apply to the segmentation of retina photographies. This procedure is fast and achieves results that comparable or even superior to other hysteresis methods and, for the problem of retina vessel segmentation, to known dedicated methods on similar data sets. Alexandru Condurache, Florian Müller 0002, Alfred Mertins |
ICPR | 3 |
| 2010 | Invariant integration features combined with speaker-adaptation methodsabstractSpeaker-normalization and -adaptation methods are essential components of state-of-the-art speech recognition systems nowadays. Recently, so-called invariant integration features were presented which are motivated by the theory of invariants. While it was shown that the integration features outperform MFCCs when used with a basic monophone recognition system, it was left open, if their benefits still can be observed when a more sophisticated recognition system with speaker-normalization and/or speaker-adaptation components is used. This work investigates the combination of the integration features with standard speaker-normalization and -adaptation methods. We show that the integration features benefit from adaptation methods and significantly outperform MFCCs in matching, as well as in mismatching training-test conditions. Index Terms: Speaker independency, invariant integration, speaker normalization, speaker adaptation Florian Müller 0002, Alfred Mertins |
INTERSPEECH | 2 |
| 2010 | Room Impulse Response Shortening/Reshaping With Infinity- and p -Norm OptimizationabstractThe purpose of room impulse response (RIR) shortening and reshaping is usually to improve the intelligibility of the received signal by prefiltering the source signal before it is played with a loudspeaker in a closed room. In an alternative, but mathematically equivalent setting, one may aim to postfilter a recorded microphone signal to remove audible echoes. While least-squares methods have mainly been used for the design of shortening/reshaping filters for RIRs until now, we propose to use the infinity- orp-norm as optimization criteria. In our method, design errors will be uniformly distributed over the entire temporal range of the shortened/reshaped global impulse response. In addition, the psychoacoustic property of masking effects is considered during the filter design, which makes it possible to significantly reduce the filter length, compared to standard approaches, without affecting the perceived performance. Alfred Mertins, Tiemin Mei, Markus Kallinger |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Generalized cyclic transformations in speaker-independent speech recognitionabstractA feature extraction method is presented that is robust against vocal tract length changes. It uses the generalized cyclic transformations primarily used within the field of pattern recognition. In matching training and testing conditions the resulting accuracies are comparable to the ones of MFCCs. However, in mismatching training and testing conditions with respect to the mean vocal tract length the presented features significantly outperform the MFCCs. Florian Müller 0002, Eugene Belilovsky, Alfred Mertins |
ASRU | 3 |
| 2009 | Using the scaling ambiguity for filter shortening in convolutive blind source separationabstractIn this paper, we propose to use the scaling ambiguity of convolutive blind source separation for shortening the unmixing filters. An often used approach for separating convolutive mixtures is the transformation to the time-frequency domain where an instantaneous ICA algorithm can be applied for each frequency separately. This approach leads to the so called permutation and scaling ambiguity. While different methods for the permutation problem have been widely studied, the solution for the scaling problem is usually based on the minimal distortion principle. We propose an alternative approach that allows the unmixing filters to be as short as possible. Shorter unmixing filters will suffer less from circular-convolution effects that are inherent to unmixing approaches based on bin-wise ICA followed by permutation and scaling correction. The results for the new algorithm will be shown on a real-world example. Radoslaw Mazur, Alfred Mertins |
ICASSP | 2 |
| 2009 | Room impulse response shortening with infinity-norm optimizationabstractThe purpose of room impulse response (RIR) shortening is to improve the intelligibility of the received signal by prefiltering the source signal before it is played with a loudspeaker in a closed room. In this paper, we propose to use the infinity-norm as optimization criterion for the design of shortening filters of RIRs. Similar to the equiripple filter design method, design errors will be uniformly distributed over the unwanted temporal range of the shortened global impulse response. The D50 measure is exploited during the design of the shortening filter, which makes it possible to significantly reduce the length of the prefilter without affecting the perceived performance. Tiemin Mei, Alfred Mertins, Markus Kallinger |
ICASSP | 2 |
| 2009 | Differential Space-Time-Frequency Codes for MB-OFDM UWB with Dual Carrier ModulationabstractIn a multiple-input multiple-output (MIMO) multiband orthogonal frequency division multiplexing (MB-OFDM) ultra-wideband (UWB) system, coherent detection where the channel state information (CSI) is assumed to be exactly known at the receiver requires the transmission of a large number of symbols for channel estimation, thus reducing the bandwidth efficiency. This paper examines the use of unitary differential space-time frequency codes (DSTFCs) in MB-OFDM UWB, which increases the system bandwidth efficiency due to the fact that no CSI is required for differential detection. The proposed DSTFC MB-OFDM system would be useful when the transmission of multiple channel estimation symbols is impractical or uneconomical. Simulation results show that the application of DSTFCs associated with dual carrier modulation (DCM) can significantly improve the bit error performance of conventional differential MB-OFDM system (without MIMO), and even provide better bit error performance than the DSTFC MB-OFDM system associated with constant envelope modulation schemes. Le Chung Tran, Alfred Mertins |
ICC | 2 |
| 2009 | Super resolution of time-of-flight depth images under consideration of spatially varying noise varianceabstractIn this paper we propose a novel way of using time-of-flight camera depth and intensity images to produce a higher resolution depth image with prior knowledge of spatial noise distribution, which is correlated with the incident light falling on each pixel. The proposed method is compared to well-established methods, and results with real image data are presented. Torsten Edeler, Kevin Ohliger, Stephan Hussmann, Alfred Mertins |
ICIP | 4 |
| 2009 | Invariant-integration method for robust feature extraction in speaker-independent speech recognitionabstractThe vocal tract length (VTL) is one of the variabilities that speaker-independent automatic speech recognition (ASR) systems encounter. Standard methods to compensate for the effects of different VTLs within the processing stages of the ASR systems often have a high computational effort. By using an appropriate warping scheme for the frequency centers of the timefrequency analysis, a change in VTL can be approximately described by a translation in the subband-index space. We present a new type of features that is based on the principle of invariant integration, and an according feature selection method is described. ASR experiments show the increased robustness of the proposed features in comparison to standard MFCCs. Index Terms: speech recognition, speaker-independency, invariant integration, monomials Florian Müller 0002, Alfred Mertins |
INTERSPEECH | 2 |
| 2009 | Improved Modelling of Ultrasound Contrast Agent Diminution for Blood Perfusion Analysis
Christian Kier, Karsten Meyer-Wiethe, Günter Seidel, Alfred Mertins |
MICCAI (1) | 4 |
| 2009 | An Approach for Solving the Permutation Problem of Convolutive Blind Source Separation Based on Statistical Signal ModelsabstractIn this paper, we present a new algorithm for solving the permutation ambiguity in convolutive blind source separation. Transformed to the frequency domain, existing algorithms can efficiently solve the reduction of the source separation problem into independent instantaneous separation in each frequency bin. However, this independency leads to the problem of correctly aligning these single bins. The new algorithm models the frequency-domain separated signals by means of the generalized Gaussian distribution and employs the small deviation of the parameters between neighboring bins for the detection of correct permutations. The performance of the algorithm will be demonstrated on synthetic and real-world data. Radoslaw Mazur, Alfred Mertins |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Novel constructions of improved square complex orthogonal designs for eight transmit antennasabstractConstructions of square, maximum rate complex orthogonal space–time block codes (CO STBCs) are well known, however codes constructed via the known methods include numerous zeros, which impede their practical implementation. By modifying the Williamson and Wallis-Whiteman arrays to apply to complex matrices, we propose two methods of construction ofsquare, order-$4n$CO STBCs fromsquare, order-$n$codes which satisfy certain properties. Applying the proposed methods, we constructsquare,maximum rate, order-8 CO STBCs with no zeros, such that the transmitted symbols are equally dispersed through transmit antennas. Those codes, referred to as theimproved square CO STBCs, have the advantages that the power is equally transmitted via each transmit antenna during every symbol time slot and that a lower peak-to-average power ratio (PAPR) is required to achieve the same bit error rates as the conventional CO STBCs with zeros. Le Chung Tran, Tadeusz A. Wysocki, Jennifer Seberry, Alfred Mertins, Sarah Spence Adams |
IEEE Trans. Inf. Theory | 4 |
| 2009 | Space-Time-Frequency Code implementation in MB-OFDM UWB communications: Design criteria and performanceabstractThis paper proposes a general framework of space-time-frequency codes (STFCs) for multi-band orthogonal frequency division multiplexing (MB-OFDM) ultra-wide band (UWB) communications systems. A great similarity between the STFC MB-OFDM UWB systems and conventional wireless complex orthogonal space-time block code (CO STBC) multiple-input multiple-output (MIMO) systems is discovered. This allows us to quantify the pairwise error probability (PEP) of the proposed system and derive the general decoding method for the implemented STFCs. Based on the theoretical analysis results of PEP, we can further quantify the diversity order and coding gain of MB-OFDM UWB systems, and derive the design criteria for STFCs, namely diversity gain criterion and coding gain criterion. The maximum achievable diversity order is found to be the product of the number of transmit antennas, the number of receive antennas, and the FFT size. We also show that all STFCs constructed based on the conventional CO STBCs can satisfy the diversity gain criterion. Various baseband simulation results are shown for the Alamouti code and a code of order 8. Simulation results indicate the significant improvement achieved in the proposed STFC MB-OFDM UWB systems, compared to the conventional MB-OFDM UWB ones. Le Chung Tran, Alfred Mertins |
IEEE Trans. Wirel. Commun. | 2 |
| 2008 | Cooperative Communication in Space-Time-Frequency Coded MB-OFDM UWBabstractThough cooperative communication has been intensively examined for general wireless systems, such as mobile and ad-hoc networks, it has been almost unexplored in the case of space-time-frequency coded multi-band OFDM ultra-wideband (STFC MB-OFDM UWB). This paper proposes a framework of cooperative communication in such systems where nodes are equipped with only one antenna. Simulation results show that cooperative communication might, in some cases, provide better error performance than non-cooperative communication without any additional transmission power. Le Chung Tran, Alfred Mertins, Tadeusz A. Wysocki |
VTC Fall | 2 |
| 2008 | Blind source separation for convolutive mixtures based on the joint diagonalization of power spectral density matrices
Tiemin Mei, Alfred Mertins, Fuliang Yin, Jiangtao Xi, Joe F. Chicharo |
Signal Process. | 2 |
| 2008 | Convolutive Blind Source Separation Based on Disjointness Maximization of Subband SignalsabstractThe concept of disjoint component analysis (DCA) is based on the fact that different speech or audio signals are typically more disjoint than mixtures of them. This letter studies the problem of blind separation of convolutive mixtures through the subband-wise maximization of the disjointness of time-frequency representations of the signals. In our approach, we first define a frequency-dependent measure representing the closeness to disjointness of a group of subband signals. Then, this frequency-dependent measure is integrated to form an objective function that only depends on the time-domain parameters of the separation system. Lastly, an efficient natural-gradient-based learning rule is developed for the update of the separation-system coefficients. Tiemin Mei, Alfred Mertins |
IEEE Signal Process. Lett. | 2 |
| 2008 | Local Region Descriptors for Active Contours EvolutionabstractEdge-based and region-based active contours are frequently used in image segmentation. While edges characterize small neighborhoods of pixels, region descriptors characterize entire image regions that may have overlapping probability densities. In this paper, we propose to characterize image regions locally by defining Local Region Descriptors (LRDs). These are essentially feature statistics from pixels located within windows centered on the evolving contour, and they may reduce the overlap between distributions. LRDs are used to define general-form energies based on level sets. In general, a particular energy is associated with an active contour by means of the logarithm of the probability density of features conditioned on the region. In order to reduce the number of local minima of such energies, we introduce two novel functions for constructing the energy functional which are both based on the assumption that local densities are approximately Gaussian. The first uses a similarity measure between features of pixels that involves confidence intervals. The second employs a local Markov Random Field (MRF) model. By minimizing the associated energies, we obtain active contours that can segment objects that have largely overlapping global probability densities. Our experiments show that the proposed method can accurately segment natural large images in very short time when using a fast level-set implementation. Cristina Darolti, Alfred Mertins, Christoph Bodensteiner, Ulrich G. Hofmann |
IEEE Trans. Image Process. | 2 |
| 2007 | A Fast Level-Set Method for Accurate Tracking of Articulated Objects with an Edge-Based Binary Speed Term
Cristina Darolti, Alfred Mertins, Ulrich G. Hofmann |
ACIVS | 2 |
| 2007 | A Spatially Robust Least Squares Crosstalk CancellerabstractCrosstalk cancellation is a well-known technique to generate virtual 3D sound via loudspeakers. Usually, headphones are used to playback audio material, which has been filtered with HRTFs. Audio playback in a reverberant room using loudspeakers and crosstalk cancellation is our intended application. This raises the need for a robust design, since listeners might slightly move their heads during a listening session. The novel design is based on a known least squares crosstalk canceller design technique, but with added robustness. The robustness is achieved with the help of assumed stochastic perturbation systems, which lie in parallel to the actual propagation impulse responses, during the design process. Markus Kallinger, Alfred Mertins |
ICASSP (1) | 2 |
| 2007 | Optimization of Gabor Features for Text-Independent Speaker IdentificationabstractFor text-independent speaker identification a prominent combination is to use Gaussian mixture models (GMM) for classification while relying on Mel-frequency cepstral coefficients (MFCC) as features. To take temporal information into account the time difference of features of adjacent speech frames are appended to the initial features. In this paper we investigate the applicability of spectro-temporal features obtained from Gabor-filters and present an algorithm for optimizing the possible parameters. Simulation results on a database show that spectro-temporal features achieve higher recognition rates than purely temporal features for clean speech as well as for disturbed speech. Volker Mildner, Stefan Goetze, Karl-Dirk Kammeyer, Alfred Mertins |
ISCAS | 4 |
| 2007 | Automatic speech recognition and speech variability: A review
Mohamed Benzeghiba, Renato De Mori, Olivier Deroo, Stéphane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christophe Ris, Richard Rose, Vivek Tyagi, Christian Wellekens |
Speech Commun. | 9 |
| 2007 | Introduction to the Special Issue on Intrinsic Speech Variations
Renato De Mori, Olivier Deroo, Stéphane Dupont, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christian Wellekens |
Speech Commun. | 7 |
| 2006 | Multiresolution Lossy-to-Lossless Coding of MRI Objects
Habibollah Danyali, Alfred Mertins |
ACIVS | 2 |
| 2006 | Automatic Speech Recognition and Intrinsic Speech VariationabstractThis paper briefly reviews state of the art related to the topic of speech variability sources in automatic speech recognition systems. It focuses on some variations within the speech signal that make the ASR task difficult. The variations detailed in the paper are intrinsic to the speech and affect the different levels of the ASR processing chain. For different sources of speech variation, the paper summarizes the current knowledge and highlights specific feature extraction or modeling weaknesses and current trends Mohamed Benzeghiba, Renato De Mori, Olivier Deroo, Stéphane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christophe Ris, Richard Rose, Vivek Tyagi, Christian Wellekens |
ICASSP (5) | 9 |
| 2006 | Multi-Channel Room Impulse Response Shaping - A StudyabstractThis paper addresses the usability of channel shortening equalizers known for data transmission systems for the equalization of acoustic systems. In multicarrier systems, equalization filters are used to shorten the channel's effective length to the size of a cyclic prefix or the guard interval. In most applications the equalizer succeeds the channel. In acoustic systems, an equalizer is placed in front of a playback loudspeaker to generate a desired impulse response for the concatenation of the equalizer, a loudspeaker, a room impulse response, and a reference microphone. In this paper, we show that shaping the desired impulse response to a shorter reverberation time is more appropriate for acoustical systems than trying to exactly truncate it to a maximum length. Investigations are carried out using a multi-loudspeaker-multi-microphone system Markus Kallinger, Alfred Mertins |
ICASSP (5) | 2 |
| 2006 | Frequency-Warping Invariant Features for Automatic Speech RecognitionabstractBased on the well-known relationship between vocal tract length (VTL) variation and linear frequency warping, we present a method for generating vocal tract length invariant (VTLI) features. These features are computed as translation invariant, correlation-type features in a log-frequency domain. In phoneme classification and recognition experiments on the TIMIT database, their discrimination capabilities and robustness to mismatches between training and test conditions turned out to be considerably better than for Mel-frequency cepstral coefficients (MFCCs). The best results are obtained when VTLI features and MFCCs are combined Alfred Mertins, Jan Rademacher |
ICASSP (5) | 1 |
| 2006 | Improved warping-invariant features for automatic speech recognitionabstractIn this paper, we extend a previously introduced method for the generation of vocal tract length invariant (VTLI) features. The novelty is a reduction of the number of obtained invariances to the more desired ones, which results in a significant improvement of recognition rates. In experiments on the TIMIT database, the enhanced discrimination capabilities and robustness to mismatches between training and test conditions are shown. Jan Rademacher, Matthias Wächter, Alfred Mertins |
INTERSPEECH | 3 |
| 2006 | Scalable multiresolution color image segmentation
Fardin Akhlaghian Tab, Golshah Naghdy, Alfred Mertins |
Signal Process. | 3 |
| 2006 | Blind Source Separation Based on Time-Domain Optimization of a Frequency-Domain Independence CriterionabstractA new technique for the blind separation of convolutive mixtures is proposed in this paper. Inspired by the works of Amari, Sabala , and Rahbar, we firstly start from the application of Kullback-Leibler divergence in frequency domain, and then we integrate Kullback-Leibler divergence over the whole frequency range of interest to yield a new objective function which turns out to be time-domain variable dependent. In other words, the objective function is derived in frequency domain which can be optimized with respect to time domain variables. The proposed technique has the advantages of frequency domain approaches and is suitable for very long mixing channels, but does not suffer from the local permutation problem as the separation is achieved in time-domain Tiemin Mei, Jiangtao Xi, Fuliang Yin, Alfred Mertins, Joe F. Chicharo |
IEEE Trans. Speech Audio Process. | 4 |
| 2005 | Low latency joint source-channel coding using overcomplete expansions and residual source redundancyabstractIn this paper, we present a joint source-channel coding method which employs quantized overcomplete frame expansions that are binary transmitted through noisy channels. The frame expansions can be interpreted as real-valued block codes that are directly applied to waveform signals prior to quantization. At the decoder, first the index-based redundancy is used by a soft-input soft-output source decoder to determine the a posteriori probabilities for all possible symbols. Given these symbol probabilities, we then determine least-squares estimates for the reconstructed symbols. The performance of the proposed approach is evaluated for code constructions based on the DFT and is compared to other decoding approaches as well as to classical BCH block codes. The results show that the new technique is superior for a wide range of channel conditions, especially when strict delay constraints for the transmission system are given Jörg Kliewer, Alfred Mertins |
GLOBECOM | 2 |
| 2005 | New aspects of combining echo cancellers with beamformersabstractIn this contribution, we introduce a novel framework for combining approaches for acoustic echo cancellation and beamforming. Classical schemes of combination incorporate the simple concatenation of both subsystems. However, if the echo cancellers come first, they cannot exploit the noise reduction capabilities of the beamformer. In the other setup, a time-variant beamformer can heavily disturb the succeeding echo canceller's convergence. The new approach establishes a possibility to smoothly switch between two possible target functions: the mean squared errors at the beamformer's inputs and its output, respectively. Thus, one combined system benefits from advantages of both classical schemes. The paper contains a theoretical analysis of possible solutions for preceding echo cancellers, which are adapted using the beamformer output. The appropriate modification of the normalized least mean squares algorithm is derived too. Karl-Dirk Kammeyer, Markus Kallinger, Alfred Mertins |
ICASSP (3) | 3 |
| 2005 | Oldenburg logatome speech corpus (OLLO) for speech recognition experiments with humans and machinesabstractThis paper introduces the new OLdenburg LOgatome speech corpus (OLLO) and outlines design considerations during its creation. OLLO is distinct from previous ASR corpora as it specifically targets (1) the fair comparison between human and machine speech recognition performance, and (2) the realistic representation of intrinsic variabilities in speech that are significant for automatic speech recognition (ASR) systems. To enable an unbiased human-machine comparison, OLLO is designed for recognition of individual phonemes that are embedded in logatomes, specifically, three-phoneme sequences with no semantic information. A balanced set of target-phonemes important for human and automatic speech recognition has been chosen, drawing on pilot ASR studies and cross-fertilization from the field of human speech intelligibility testing. Several intrinsic variabilities in speech are represented in OLLO, by recording from 40 speakers from four German dialect regions, and by covering six articulation characteristics. Results from preliminary phonetic time-labeling and ASR experiments are promising and consistent with corpus variabilities. 1. Thorsten Wesker, Bernd T. Meyer, Kirsten Wagener, Jörn Anemüller, Alfred Mertins, Birger Kollmeier |
INTERSPEECH | 5 |
| 2005 | Generalized Williamson and Wallis-Whiteman constructions for improved square order-8 CO STBCsabstractConstructions of square, maximum rate complex orthogonal space-time block codes (CO STBCs) are well known, however codes constructed via the known methods include numerous zeros, which impede their practical implementation. By modifying the Williamson and Wallis-Whiteman arrays to apply to complex matrices, we propose two methods of construction or square, order-4n CO STBCs from square, order-n codes, which satisfy certain properties. Applying the proposed methods, we construct square, maximum rate, order-8 CO STBCs with no zeros, such that the transmitted symbols equally disperse through transmit antennas. Those codes, referred to as the improved square CO STBCs, have the advantages that the power is equally transmitted via each transmit antenna during every symbol time slot and that a lower peak-to-mean power ratio per each antenna is required to achieve the same bit error rates as for the conventional CO STBCs with zeros Le Chung Tran, Tadeusz A. Wysocki, Jennifer Seberry, Alfred Mertins, Sarah Spence Adams |
PIMRC | 4 |
| 2005 | A Generalized Algorithm for the Generation of Correlated Rayleigh Fading EnvelopesabstractAlthough the generation of correlated Rayleigh fading envelopes has been intensively considered in the literature, all conventional methods have their own particular shortcomings, which seriously impedes their applicability. A very general, straightforward algorithm is proposed for the generation of an arbitrary number of Rayleigh envelopes with any desired, equal or unequal power in wireless channels, either with or without Doppler frequency shifts. The proposed algorithm can be applied in the case of spatial correlation, such as with antenna arrays in multiple input multiple output (MIMO) systems, or spectral correlation between random processes, as in orthogonal frequency division multiplexing (OFDM) systems. It can also be used for generating correlated Rayleigh fading envelopes in either discrete-time instants or a real-time scenario. Besides being more generalized, our proposed algorithm is more precise, while overcoming all the shortcomings of conventional methods. Le Chung Tran, Tadeusz A. Wysocki, Jennifer Seberry, Alfred Mertins |
WOWMOM | 4 |
| 2005 | Sketch-based image matching Using Angular partitioningabstractThis work presents a novel method for image similarity measure, where a hand-drawn rough black and white sketch is compared with an existing data base of full color images (art works and photographs). The proposed system creates ambient intelligence in terms of the evaluation of nonprecise, easy to input sketched information. The system can then provide the user with options of either retrieving similar images in the database or ranking the quality of the sketch against a given standard, i.e., the original image model. Alternatively, the inherent pattern-matching capability of the system can be utilized to allow detection of distortion in any given real time-image sequences in vision-driven ambient intelligence applications. The proposed method can cope with images containing several complex objects in an inhomogeneous background. Two abstract images are obtained using strong edges of the model image and the morphologically thinned outline of the sketched image. The angular-spatial distribution of pixels in the abstract images is then employed to extract new compact and effective features using the Fourier transform. The extracted features are rotation and scale invariant and robust against translation. Experimental results from seven different approaches confirm the efficacy of the proposed method in both the retrieval performance and the time required for feature extraction and search. Abdolah Chalechale, Golshah Naghdy, Alfred Mertins |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2004 | On iterative source-channel image decoding with Markov random field source modelsabstractIn this paper, we propose a novel iterative source-channel decoding approach for robust transmission of compressed still images over noisy communication channels. Besides the explicit redundancy introduced by channel encoding, also implicit residual source redundancy is exploited for error protection. The source redundancy is modeled by a Markov random field (MRF) source model, which considers the residual spatial correlation after source encoding. The resulting MRF-based soft-input/soft-output source decoder is used as outer constituent decoder in the proposed iterative source-channel decoding scheme, where due to the link between MRFs and the Gibbs distribution, the source decoder can be implemented with very low complexity. We show that this iterative decoding scheme can be successfully employed for recovering the image data, especially when the channel is highly corrupted. Jörg Kliewer, Norbert Goertz, Alfred Mertins |
ICASSP (4) | 3 |
| 2004 | Blind source separation of nonstationary convolutively mixed signals in the subband domainabstractThe paper proposes a new technique for blind source separation (BSS) in the subband domain using an extended lapped transform (ELT) decomposition for nonstationary, convolutively mixed signals. As identified by S. Araki et al. (see Proc. 4th Int. Symp. on Independent Component Analysis and Blind Signal Separation - ICA2003, p.499-504, 2003), the motivation for subband-based BSS is the drawback of frequency domain BSS when dealing with separating mixed speech signals over a few seconds resulting in few samples in individual frequency bins leading to poor separation performance. In the proposed approach, mixed signals are decomposed into subband components by an ELT and within each subband a time domain Newton BSS algorithm is employed based on the nonstationarity property of the input signals and the joint diagonalization of output correlation matrices with time varying second order statistics (SOS). This subband version is compared to a fullband version using the same BSS algorithm. Iain Russell, Jiangtao Xi, Alfred Mertins, Joe F. Chicharo |
ICASSP (5) | 3 |
| 2004 | Document image analysis and verification using cursive signatureabstractA new approach for document image analysis and verification is presented. The approach utilizes connected component analysis and geometric properties of labelled regions for region of interest extraction. Document images containing Persian/Arabic text combined with English text, headlines, ruling lines, trade mark and cursive signature are used as a test data. Persian/Arabic signature extraction is investigated as a case study. The proposed method uses special characteristics of such signatures for extraction and verification procedures. A set of efficient, invariant and compact features is extracted utilizing spatial partitioning of the signature region. Comparative results exhibit high extraction and verification rates Abdolah Chalechale, Golshah Naghdy, Prashan Premaratne, Alfred Mertins |
ICME | 4 |
| 2004 | Soft-input reconstruction of binary transmitted quantized overcomplete expansionsabstractWe propose a soft-decoding method for quantized overcomplete frame expansions that are binary transmitted through noisy channels. The frame expansions can be viewed as real-valued block codes that are directly applied to waveform signals prior to quantization. The explicit redundancy introduced in the continuous amplitude domain is exploited by the decoder in two stages. First, the index-based redundancy is used by a soft-input soft-output source decoding approach that outputs decoded symbols together with their reliability information. In a second stage, the soft information on the symbols and the structure of the introduced redundancy are used to correct errors. The performance of the proposed approach is evaluated for different code constructions based on the discrete Fourier transform (DFT), the discrete cosine transform (DCT), and the discrete Hadamard transform (DHT), and is compared to standard approaches without soft decoding. Jörg Kliewer, Alfred Mertins |
IEEE Signal Process. Lett. | 2 |
| 2003 | Fully scalable texture coding of arbitrarily shaped video objectsabstractThis paper presents a fully scalable texture coding algorithm for arbitrarily shaped video object based on the 3D-SPIHT algorithm. The proposed algorithm, called Fully Scalable Object-based 3D-SPIHT (FSOB-3DSPIHT), modifies the 3D-SPIHT algorithm to code video objects with arbitrary shape and adds spatial and temporal scalability features to it. It keeps important features of the original 3D-SPIHT coder such as compression efficiency, full embeddedness, and rate scalability. The full scalability feature of the modified algorithm is achieved through the introduction of multiple resolution dependent lists for the sorting stage of the algorithm. The idea of bitstream transcoding without decoding to obtain different bitstreams for various spatial and temporal resolutions and bit rates is completely supported by the algorithm. Habibollah Danyali, Alfred Mertins |
ICASSP (3) | 2 |
| 2003 | Scalable to lossless audio compression based on perceptual set partitioning in hierarchical trees (PSPIHT)abstractThe paper proposes a technique for scalable to lossless audio compression. The scheme presented is perceptually scalable and also provides for lossless compression. It produces smooth objective scalability, in terms of SegSNR, from lossy to lossless compression. The proposal is built around the introduced perceptual SPIHT algorithm, which is a modification of the SPIHT algorithm. Both objective and subjective results are reported and demonstrate both perceptual and objective measure scalability. The subjective results indicate that the proposed method performs comparably with the MPEG-4 AAC coder at 16, 32 and 64 kbps, yet also achieves a scalable-to-lossless architecture. Mohammed Raad, Alfred Mertins, Ian S. Burnett |
ICASSP (5) | 2 |
| 2003 | Scalable to lossless audio compression based on perceptual set partitioning in hierarchical trees (PSPIHT)abstractThis paper proposes a technique for scalable and also provides for lossless compression. It reduces smooth objective scalability, in terms of SegSNR, from lossy to lossless compression. The proposal is built around the perceptual SPIHT algorithm, which is a modification of the SPIHT algorithm and is introduced in this paper. Both objective and subjective results are reported and demonstrate that the proposed method performs comparably with the MPEG-4 AAC coder at 16, 32 and 64 kbps, yet also achieves a scalable-to-lossless architecture. Mohammed Raad, Alfred Mertins, Ian S. Burnett |
ICME | 2 |
| 2003 | Multi-rate extension of the scalable to lossless PSPIHT audio coderabstractThis paper extends a scalable to lossless compression scheme to allow scalability in terms of sampling rate as well as quantization resolution. The scheme presented is an extension of a perceptually scalable scheme that scales to lossless compression, producing smooth objective scalability, in terms of SNR, until lossless compression is achieved. The scheme is built around the Perceptual SPIHT algorithm, which is a modification of the SPIHT algorithm. An analysis of the expected limitations of scaling across sampling rates is given as well as lossless compression results showing the competitive performance of the presented technique. 1. Mohammed Raad, Ian S. Burnett, Alfred Mertins |
INTERSPEECH | 3 |
| 2003 | Coefficient Selection Methods for Scalable Spread Spectrum Watermarking
Angela Piper, Reihaneh Safavi-Naini, Alfred Mertins |
IWDW | 3 |
| 2002 | Frame bounds for biorthogonal cosine-modulated filter banksabstractIn this paper, we derive explicit expressions for the eigenvalues of the frame operator for cosine-modulated filter. banks. The filter banks may be critically sampled or oversampled by an integer factor. The analysis of low-delay, biorthogonal filter banks shows that prototypes solely designed to minimize the stopband energy may lead to wide open frames and thus to an undesirable numerical behavior. Because the computational cost of determining the frame bounds with the proposed method is very low, we can directly use the bounds during prototype optimization and obtain prototypes with minimum stopband energy under the condition of fixed frame bounds. Alfred Mertins |
ICASSP | 1 |
| 2002 | Design of redundant FIR precoders for arbitrary channel lengths based on an MMSE criterionabstractThe joint design of transmitter and receiver for multichannel data transmission over dispersive channels is considered. In particular, the practically important case where the transmitter consists of FIR filters and the channel impulse response has arbitrary length is addressed. The design criterion is the minimization of the mean squared error at the receiver output under the constraint of a fixed transmit power. The proposed algorithm allows a straightforward transmitter design and yields (in general) a near-optimal solution for the transmit filters. Under certain conditions, the exact solution for the optimal transmitter is obtained. Alfred Mertins |
ICC | 1 |
| 2002 | Highly scalable image compression based on SPIHT for network applicationsabstractWe propose a highly scalable image compression scheme based on the set partitioning in hierarchical trees (SPIHT) algorithm. Our algorithm, called highly scalable SPIHT (HS-SPIHT), supports spatial and SNR scalability and provides a bitstream that can be easily adapted (reordered) to given bandwidth and resolution requirements by a simple transcoder (parser). The HS-SPIHT algorithm adds the spatial scalability feature without sacrificing the SNR embeddedness property as found in the original SPIHT bitstream. HS-SPIHT finds applications in progressive Web browsing, flexible image storage and retrieval, and image transmission over heterogeneous networks. Habibollah Danyali, Alfred Mertins |
ICIP (1) | 2 |
| 2002 | Decoding of images using soft-bits and Markov random field modelingabstractWe combine soft-bit decoding with Markov random field modeling of source signals and apply the proposed technique to image communication over noisy channels. No redundancy is added by the encoder in the form of channel codes. For error correction and concealment, the decoder relies only on the bit-reliability information extracted at the channel output and the random field modeling of natural images. The results obtained show that the proposed method yields excellent performance even under extremely noisy conditions. Alfred Mertins, Olivier Jamart |
ICIP (1) | 1 |
| 2001 | Perfect reconstruction integer-modulated filter banksabstractWe present design methods for perfect reconstruction (PR) integer-modulated filter banks, including biorthogonal (low-delay) filter banks. Both the prototype filter and the modulation sequences are composed of integers, thus allowing efficient hardware implementations. To derive such filter banks, we first extend the PR conditions known for cosine modulation to other, more general, modulation schemes. We present solutions where the PR conditions on the prototype and the modulation are entirely decoupled and where some simple coupling is introduced. The conditions are derived for both even and odd numbers of channels. Design examples are presented for both cases. Alfred Mertins, Tanja Karp |
ICASSP | 1 |
| 2001 | Efficient biorthogonal cosine-modulated filter banks
Tanja Karp, Alfred Mertins, Gerald Schuller |
Signal Process. | 2 |
| 2000 | Embedded wavelet coding of arbitrarily shaped objects
Alfred Mertins, Sudhir Singh |
VCIP | 1 |
| 2000 | Processing arbitrary-length signals with linear-phase cosine-modulated filter banks
Jörg Kliewer, Tanja Karp, Alfred Mertins |
Signal Process. | 3 |
| 1998 | Biorthogonal cosine-modulated filter banks without DC leakageabstractIn this paper, we present a structure for implementing the polyphase filters of biorthogonal modulated filter banks that automatically guarantees perfect reconstruction (PR) of the filter bank and furthermore allows to specify the values of the filters' frequency responses at certain frequencies. Thus, modulated filter banks without DC leakage can be designed. The new structure is based on lifting schemes for the polyphase filters and DC leakage can be avoided very easily when reducing the number of lifting coefficients that can be freely chosen and used for filter optimization. The great advantage of the new method is that we do not have to take constraints into consideration when optimizing the prototype filter, but PR and specified zeros are structure inherent. Tanja Karp, Alfred Mertins |
ICASSP | 2 |
| 1998 | Discrete-coefficient linear-phase prototypes for PR cosine-modulated filter banksabstractA method for the design of perfect reconstruction (PR) linear-phase prototypes for cosine-modulated filter banks with discrete coefficients is presented. Such prototypes are of great interest for efficient hardware implementations. The design procedure is based on a subspace approach that allows to linearly combine PR prototype filters in such a way that the resulting filter also is a PR prototype. Within a given subspace the weights of the optimal linear combination can be easily computed via an eigenanalysis. The filter design is carried out iteratively, while the PR property is guaranteed throughout the design process. No non-linear optimization routine is needed. Alfred Mertins |
ICASSP | 1 |
| 1998 | Optimized Biorthogonal Shape Adaptive Wavelets
Alfred Mertins |
ICIP (3) | 1 |
| 1997 | Efficiently VLSI-realizable prototype filters for modulated filter banksabstractThis paper presents methods for the efficient realization of prototype filters for modulated filter banks. The implementation is based on the lattice structure of the polyphase filters. The lattice coefficients, representing rotations, are approximated by a small number of simple /spl mu/-rotations each of which can be realized by some shift and add operations instead of a multiplication. Since the lattice structure is robust against coefficient quantization we do not loose the perfect reconstruction (PR) property of the filter bank when doing this approximation. The frequency responses of the original and approximated prototype filters are compared in terms of complexity and stopband attenuation. Tanja Karp, Alfred Mertins, Truong Q. Nguyen |
ICASSP | 2 |
| 1997 | Design of paraunitary oversampled cosine-modulated filter banksabstractIn this paper we derive perfect reconstruction (PR) conditions for oversampled cosine-modulated filter banks. The results can be regarded as a generalization of the known work for critical subsampling. We show that in the oversampled case we gain some additional degree of freedom, which can be exploited in the filter design process. This leads to PR prototypes with stopband attenuations being much higher than in the critically subsampled PR case. The filters designed as PR filters for the oversampled case can also serve as prototypes for critically subsampled cosine-modulated pseudo QMF banks. Jörg Kliewer, Alfred Mertins |
ICASSP | 2 |
| 1996 | Edge sensitive subband coding of imagesabstractIn this paper, we introduce a family of novel edge preserving image coding techniques based on the discrete wavelet transform (DWT). These techniques utilize the edge map (which not necessarily contains closed curves) of the processed image or its subbands in the way that no filtering over edges is performed. This results in a reduction of power in the higher frequency bands. The achievable savings are higher than the additional bit rate needed for transmission of the edge map. Kalman Cinkler, Alfred Mertins |
ICASSP | 2 |
| 1996 | Processing arbitrary-length signals with MDFT filter banksabstractIn this paper, methods for processing arbitrary-length input signals with MDFT filter banks are presented. The MDFT filter bank can be regarded as the most general type of a special class of filter banks. These filter banks have linear phase analysis filters, but different centers of symmetry due to subsampling with and without a phase shift. Already known extension methods cannot be applied to these filter banks in their original form. We first discuss the symmetric extension for two special input signal lengths. These cases are then incorporated in the general solution for arbitrary-length input signals. Tanja Karp, Jörg Kliewer, Alfred Mertins, Norbert J. Fliege |
ICASSP | 3 |
| 1996 | Time-delay estimation for compound point-processes using hidden Markov modelsabstractA new time-delay estimation algorithm for compound point-processes is presented. Compound point-processes, a generalization of temporal point-processes, describe processes with discrete events, where each occurrence time is associated with certain features. It is shown that, although the events are not observable, the time delays from events at one location to the same events at a second location can be estimated using hidden Markov models based on the associated features. We demonstrate the performance of this time-delay estimation algorithm with an application to the estimation of section-related traffic data in road traffic monitoring and control systems. Jens Wohlers, Alfred Mertins, Norbert J. Fliege |
ICASSP | 2 |
| 1996 | Coding of digital video with the edge-sensitive discrete wavelet transformabstractIn this paper we present an image sequence coding scheme for very low bit rate coding which is based on spatial redundancy reduction via the new edge sensitive subband coding method and temporal redundancy reduction via windowed overlapped block-matching motion compensation. The scheme has the main advantage, that only the significant regions of difference-images are coded. Thus the computational cost can be kept low. Due to the properties of both the temporal and the spatial coding used, there are no blocking effects in the coded images, and the overall visual performance of the coding scheme is very good. Kalman Cinkler, Alfred Mertins |
ICIP (1) | 2 |
| 1995 | Time-varying and support preservative filter banks: design of optimal transition and boundary filters via SVDabstractMethods for switching filter coefficients and filter bank structures and methods for processing finite length signals are studied. The problem of designing optimal boundary and transition filters is solved directly via singular value decomposition (SVD) while the optimality criterion is based on the subband statistics. The optimized filters provide a good match between the subband statistics in the transition regions (and at the boundaries) to the statistics in the steady state. The filter banks considered are maximally decimated M-channel linear and non-linear phase (biorthogonal and paraunitary) filter banks with real filter coefficients. Alfred Mertins |
ICASSP | 1 |
| 1994 | Statistical optimization of PR-QMF banks and waveletsabstractThe design of optimal power complementary filters for multiresolution signal decomposition is considered. As a new performance measure in statistical optimization of filter banks the norm of cross-correlation (all subbands including all delays) will be used for optimization. The obtained filters will be compared to filters having minimum reconstruction error when the high-pass branches are dropped. In the paraunitary case, which is considered here, this reconstruction error is closely related to the coding gain of PCM in subbands over direct PCM. While the KL-transform combines both minimum reconstruction error and perfect decorrelation, filter banks do not. It is clear that the norm of cross-correlation becomes very important when better techniques than PCM are considered for the coding of subband signals.> Alfred Mertins |
ICASSP (3) | 1 |
| 1994 | Non-linear regression based feature extraction for connected-word recognition in noiseabstractThis paper shows the application of non-linear regression to robust feature extraction for noisy speech recognition. In this approach, a non-linear estimator is used to compute noise invariant features from non-linear combinations of noise contaminated observations. The observations may be short-term subband-energies obtained from a filter bank analysis, cepstral coefficients of linear prediction coefficients. Instead of training the hidden Markov models (HMMs) under various noise conditions, they can be trained with clean data. The results show that this method leads to error rates comparable to those achieved by training in the presence of noise.> Frank Seide, Alfred Mertins |
ICASSP (2) | 2 |
| 1992 | Multiple parameter estimation of composite signals in uncharacterized impulsive noiseabstractThe problem of estimating amplitudes and arrival times of composite signals in non-Gaussian noise is addressed. The maximum likelihood estimator (MLE) in Gaussian mixture noise with statistical independent noise samples is presented, and it is shown that the MLE is related to a nonlinear form of the orthogonality principle. Some important properties of the multiple parameter estimation problem are extracted from Fisher's information matrix. The problem of the MLE is that the noise density is seldom known in practice. A nonlinear nonparametric iterative method which does not require the knowledge of the noisy density is presented and compared to the MLE. The basic difference between the nonparametric method and the MLE is that the nonlinear manipulations are based on the actual input samples, and therefore the nonparametric method is applicable in nonstationary noise processes.> Alfred Mertins, Norbert J. Fliege |
ICASSP | 1 |