Nikolaos Mitianoudis

dblp:49/537 · also Nikos Mitianoudis · DBLP profile ↗
← Back
31ranked-venue papers
12as first author
5since 2021 · last 2024
0000-0003-0898-6102ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
7 papers
Image and video processing · 68% Audio and music processing · 30% Geometric modeling and processing · 2%
Artificial intelligence
1 paper
3D vision · 77% Deep learning architectures and training · 23%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
image fusion
1.022023
A Convolutional Neural Network-Based Conditional Random Field Model for Structured Multi-Focus Image Fusion Robust to Noise · IEEE Trans. Image Process. 2023
Conditional Random Field Model for Robust Multi-Focus Image Fusion · IEEE Trans. Image Process. 2019
Image and video processing
image restoration
1.022023
A Convolutional Neural Network-Based Conditional Random Field Model for Structured Multi-Focus Image Fusion Robust to Noise · IEEE Trans. Image Process. 2023
Conditional Random Field Model for Robust Multi-Focus Image Fusion · IEEE Trans. Image Process. 2019
Image and video processing › image fusion
multi-focus image fusion
1.022023
A Convolutional Neural Network-Based Conditional Random Field Model for Structured Multi-Focus Image Fusion Robust to Noise · IEEE Trans. Image Process. 2023
Conditional Random Field Model for Robust Multi-Focus Image Fusion · IEEE Trans. Image Process. 2019
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.812024
Monocular Depth Estimation: A Thorough Review · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Audio and music processing
source separation
0.742020
A novel Directional Framework for Source Counting and Source Separation in Instantaneous Underdetermined Audio Mixtures · IEEE ACM Trans. Audio Speech Lang. Process. 2020
A Generalized Directional Laplacian Distribution : Estimation, Mixture Models and Audio Source Separation · IEEE Trans. Speech Audio Process. 2012
Batch and Online Underdetermined Source Separation Using Laplacian Mixture Models · IEEE Trans. Speech Audio Process. 2007
Image and video processing › image restoration
image denoising
0.712023
A Convolutional Neural Network-Based Conditional Random Field Model for Structured Multi-Focus Image Fusion Robust to Noise · IEEE Trans. Image Process. 2023
Audio and music processing › source separation › blind source separation
underdetermined source separation
0.732020
A novel Directional Framework for Source Counting and Source Separation in Instantaneous Underdetermined Audio Mixtures · IEEE ACM Trans. Audio Speech Lang. Process. 2020
A Generalized Directional Laplacian Distribution : Estimation, Mixture Models and Audio Source Separation · IEEE Trans. Speech Audio Process. 2012
Batch and Online Underdetermined Source Separation Using Laplacian Mixture Models · IEEE Trans. Speech Audio Process. 2007
Audio and music processing › computational auditory scene analysis
source counting
0.412020
A novel Directional Framework for Source Counting and Source Separation in Instantaneous Underdetermined Audio Mixtures · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Image and video processing › image fusion
multi-modal image fusion
0.412019
Conditional Random Field Model for Robust Multi-Focus Image Fusion · IEEE Trans. Image Process. 2019
Image and video processing › image representation
moment-based shape description
0.112009
A Unifying Approach to Moment-Based Shape Orientation and Symmetry Classification · IEEE Trans. Image Process. 2009
Geometric modeling and processing
shape analysis
0.112009
A Unifying Approach to Moment-Based Shape Orientation and Symmetry Classification · IEEE Trans. Image Process. 2009
Audio and music processing › source separation
blind source separation
0.012003
Audio source separation of convolutive mixtures · IEEE Trans. Speech Audio Process. 2003
Audio and music processing › source separation › blind source separation
convolutive mixture separation
0.012003
Audio source separation of convolutive mixtures · IEEE Trans. Speech Audio Process. 2003
Audio and music processing › source separation › blind source separation
independent component analysis
0.012003
Audio source separation of convolutive mixtures · IEEE Trans. Speech Audio Process. 2003

Methods — techniques the papers use, named apart from their topics

graph cuts · 1.0conditional random field · 1.0deep learning · 0.8convolutional neural network · 0.7directional fuzzy c-means · 0.4clustering · 0.4energy minimization · 0.4mixture model · 0.1maximum likelihood estimation · 0.1expectation-maximization · 0.1fourier series analysis · 0.1
YearPublicationVenuePosition
2024 Monocular Depth Estimation: A Thorough Review
abstract
Estimation of depth in two-dimensional images is among the challenging topics in Computer Vision. This is a well-studied but also an ill-posed problem, which has long been the focus of intense research. This paper is an in-depth review of the topic, presenting two aspects, one that considers the mechanisms of human depth perception, and another that includes the various Deep Learning approaches. The methods are presented in a compact and structured way that outlines the topic and categorizes the approaches according to the line of research followed in the recent decade. Although there has been significant advancement in the topic, it was without any connection with human depth perception and the potential benefits from this sector.
Vasileios Arampatzakis, George Pavlidis, Nikolaos Mitianoudis, Nikos Papamarkos
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 A Dilated MultiRes Visual Attention U-Net for historical document image binarization
Nikolaos Detsikas, Nikolaos Mitianoudis, Nikos Papamarkos
Signal Process. Image Commun.2
2023 A non-contact SpO2 estimation using video magnification and infrared data
abstract
Peripheral oxygen saturation (SpO2) is one important vital sign to be monitored in individuals, whose health is fragile, such as the elderly. Contactless SpO2monitoring using RGB cameras has been already developed with satisfactory results. This work explores the case of achieving an acceptable level of performance, when the lightning conditions are not optimal, particularly during night time, by processing solely infrared low-cost camera recordings. The Eulerian Video Magnification (EVM) technique was used to enhance the subtle differences in skin pixel intensity in the facial area. Two approaches were explored for performing regression: one using 12 novel features extracted from the amplified photoplethysmography (PPG) signal and Generalized Additive Models and a second using a 3D Convolution Neural Network (CNN) architecture on the raw amplified forehead video. The root mean square error in the estimated SpO2levels using both methods is minimal and in the accepted range for these applications.
Thomas Stogiannopoulos, Grigorios-Aris Cheimariotis, Nikolaos Mitianoudis
ICASSP3
2023 A Convolutional Neural Network-Based Conditional Random Field Model for Structured Multi-Focus Image Fusion Robust to Noise
abstract
The limited depth of field of optical lenses, makes multi-focus image fusion (MFIF) algorithms of vital importance. Lately, Convolutional Neural Networks (CNN) have been widely adopted in MFIF methods, however their predictions mostly lack structure and are limited by the size of the receptive field. Moreover, since images have noise due to various sources, the development of MFIF methods robust to image noise is required. A novel robust to noise Convolutional Neural Network-based Conditional Random Field (mf-CNNCRF) model is introduced. The model takes advantage of the powerful mapping between input and output of CNN networks and the long range interactions of the CRF models in order to reach structured inference. Rich priors for both unary and smoothness terms are learned by training CNN networks. The α -expansion graph-cut algorithm is used to reach structured inference for MFIF. A new dataset, which includes clean and noisy image pairs, is introduced and is used to train the networks of both CRF terms. A low-light MFIF dataset is also developed to demonstrate real-life noise introduced by the camera sensor. Qualitative and quantitative evaluation prove that mf-CNNCRF outperforms state-of-the-art MFIF methods for clean and noisy input images, while being more robust to different noise types without requiring prior knowledge of noise.
Odysseas Bouzos, Ioannis Andreadis, Nikolaos Mitianoudis
IEEE Trans. Image Process.3
2022 Low-Cost Online Convolution Checksum Checker
abstract
Managing random hardware faults requires the faults to be detected online, thus simplifying recovery. Algorithm-based fault tolerance has been proposed as a low-cost mechanism to check online the result of computations against random hardware failures. In this case, the checksum of the actual result is checked against a predicted checksum computed in parallel by a hardware checker. In this work, we target the design of such checkers for convolution engines that are currently the most critical building block in image processing and computer vision applications. The proposed convolution checksum checker, named ConvGuard, utilizes a newly introduced invariance condition of convolution to predictimplicitlythe output checksum using only the pixels at the border of the input image. In this way, ConvGuard reduces the power required for accumulating the input pixels without requiring large buffers to hold intermediate checksum results. The design of ConvGuard is generic and can be configured for different output sizes and strides. The experimental results show that ConvGuard utilizes only a small percentage of the area/power of an efficient convolution engine while being significantly smaller and more power efficient than a state-of-the-art checksum checker for various practical cases.
Dionysios Filippas, Nikolaos Margomenos, Nikolaos Mitianoudis, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
IEEE Trans. Very Large Scale Integr. Syst.3
2020 Deep Person Identification Using Spatiotemporal Facial Motion Amplification
abstract
We explore the capabilities of a new biometric trait, which is based on information extracted through facial motion amplification. Unlike traditional facial biometric traits, the new biometric does not require the visibility of facial features, such as the eyes or nose, that are critical in common facial biometric algorithms. In this paper we propose the formation of a spatiotemporal facial blood flow map, constructed using small motion amplification. Experiments show that the proposed approach provides significant discriminatory capacity over different training and testing days and can be potentially used in situations where traditional facial biometrics may not be applicable.
K. Gkentsidis, Theodora Pistola, Nikolaos Mitianoudis, Nikolaos V. Boulgouris
ICIP3
2020 A novel Directional Framework for Source Counting and Source Separation in Instantaneous Underdetermined Audio Mixtures
abstract
The audio source separation problem is a well-known problem that was addressed using a variety of techniques. A common setback in these techniques is that the total number of sound sources in the audio mixture must be known beforehand. However, this knowledge is not always available and thus needs to be estimated. Many approaches have attempted to estimate the number of sources in an audio mixture. There are several clustering techniques that can count the sources in an audio mixture, nonetheless, there are cases, where the directionality of the audio data in the mixture may lead these techniques to failure. In this article, we propose a generalised Directional Fuzzy C-Means (DFCM) framework that offers a complete multi-dimensional, directional solution to this problem. Our proposal shows remarkably high performance in estimating the correct number of sources in the majority of the cases and in addition, it can be used as an effective mechanism to separate the sources. The complete source counting-separation framework can act as a robust low-complexity simultaneous solution to both problems.
Thomas Sgouros, Nikolaos Mitianoudis
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Biometric Identification Using Facial Motion Amplification
abstract
We propose a new biometric trait based on facial motion amplification. The main advantage of the new biometric characteristic is that it does not rely on the visibility of critical facial features, such as nose, mouth, iris, or eyebrows. This makes it effective even when the respective areas are covered. Using the proposed system, facial image sequences are captured using an ordinary video camera and facial blood flow is calculated by means of small motion amplification. The calculated blood flow is captured from limited facial areas and is represented as a template that is suitable for identification purposes. Experiments on a new database show promising performance of the proposed approach, and provide evidence of the discriminatory capacity of the proposed biometric.
Theodora Pistola, Anastasios Papadopoulos, Nikolaos Mitianoudis, Nikolaos V. Boulgouris
ICIP3
2019 Conditional Random Field Model for Robust Multi-Focus Image Fusion
abstract
In this paper, a novel multi-focus image fusion algorithm based on conditional random field optimization (mf-CRF) is proposed. It is based on an unary term that includes the combined activity estimation of both high and low frequencies of the input images, while a spatially varying smoothness term is introduced, in order to align the graph-cut solution with boundaries of focused and defocused pixels. The proposed model retains the advantages of both spatial-domain methods and multi-spectral methods and by solving an energy minimization problem and finds an optimal solution for the multi-focus image fusion problem. Experimental results demonstrate the effectiveness of the proposed method that outperforms current state-of-the-art multi-focus image fusion algorithms in both qualitative and quantitative comparisons. In this paper, the successful application of the mf-CRF model in multi-modal image fusion (visible-infrared and medical) is also presented.
Odysseas Bouzos, Ioannis Andreadis, Nikolaos Mitianoudis
IEEE Trans. Image Process.3
2018 A Hermite neural network incorporating artificial bee colony optimization to model shoreline realignment at a reef-fronted beach
George E. Tsekouras, Vasilis Trygonis, Andreas Maniatopoulos, Anastasios Rigos, Antonios Chatzipavlis, John V. Tsimikas, Nikolaos Mitianoudis, Adonis Velegrakis
Neurocomputing7
2018 Multidimensional directional steerable filters - Theory and application to 3D flow estimation
Dimitrios S. Alexiadis, Nikolaos Mitianoudis, Tania Stathaki
Image Vis. Comput.2
2015 Document image binarization using local features and Gaussian mixture modeling
Nikolaos Mitianoudis, Nikos Papamarkos
Image Vis. Comput.1
2014 Local Co-occurrence and Contrast Mapping for Document Image Binarization
abstract
Document Image Binarization refers to the task of transforming a scanned image of a handwritten or printed document into a bi-level representation containing only characters and background. Here, we address the historic document image binarization problem using a three-stage methodology. Firstly, we remove possible stains and noise from the document image by estimating the document background image. The remaining background and character pixels are separated using a Local Co-occurrence Mapping, local contrast and a two-state Gaussian Mixture Model. In the last stage, possible isolated misclassified blobs are removed by a morphology operator. The proposed scheme offers robust and fast performance, especially for handwritten documents.
Nikolaos Mitianoudis, Nikos Papamarkos
ICFHR1
2014 Multidimensional steerable filters and 3D flow estimation
abstract
In this work, the 3D flow estimation problem is formulated in the 4D spatiotemporal frequency domain, and it is shown that 3D motion manifests itself as energy concentration along hyper-planes in that domain. Based on this, the construction and use of appropriate directional multidimensional “steerable” filters, which can extract directional energy in spacetime, is proposed. Steerable filters have been constructed for up to 3 dimensions. We extend the relevant mathematical definitions to multiple dimensions and formulate filter-based algorithms for 3D flow estimation. Experimental results on simulated and real data verify the efficiency of the algorithms.
Dimitrios S. Alexiadis, Nikolaos Mitianoudis, Tania Stathaki
ICIP2
2014 Multi-spectral document image binarization using image fusion and background subtraction techniques
abstract
In this paper, the authors exploit a multispectral image representation to perform more accurate document image binarisation compared to previous color representations. In the first stage, image fusion is employed to create a “document” and a “background” image. In the second stage, the FastICA algorithm is used to perform background subtraction. In the third stage, a spatial kernel K-harmonic means classifier binarizes the FastICA output. The proposed system outperforms previous efforts on document image binarization.
Nikolaos Mitianoudis, Nikos Papamarkos
ICIP1
2014 Real time hand detection in a complex background
Ekaterini Stergiopoulou, Kyriakos Sgouropoulos, Nikos A. Nikolaou, Nikos Papamarkos, Nikolaos Mitianoudis
Eng. Appl. Artif. Intell.5
2012 A Generalized Directional Laplacian Distribution : Estimation, Mixture Models and Audio Source Separation
abstract
Directional or Circular statistics are pertaining to the analysis and interpretation of directions or rotations. In this work, a novel probability distribution is proposed to model multidimensional sparse directional data. The Generalized Directional Laplacian Distribution (DLD) is a hybrid between the Laplacian distribution and the von Mises-Fisher distribution. The distribution's parameters are estimated using Maximum-Likelihood Estimation over a set of training data points. Mixtures of Directional Laplacian Distributions (MDLD) are also introduced in order to model multiple concentrations of sparse directional data. The author explores the application of the derived DLD mixture model to cluster sound sources that exist in an underdetermined instantaneous sound mixture. The proposed model can solve the general${K\times L~(K< L)}$underdetermined instantaneous source separation problem, offering a fast and stable solution.
Nikolaos Mitianoudis
IEEE Trans. Speech Audio Process.1
2011 Region-based image fusion using a combinatory Chebyshev-ICA method
abstract
The aim of this paper is to provide an algorithm for image fusion which combines the techniques of Chebyshev polynomial (CP) approximation and independent component analysis (ICA), based on the regional information of input images. We present a region-based method that combines the merits of both techniques. It utilises segmentation to identify edges, texture and other important features in the input image and subsequently apply the different fusion methods according to regions. The proposed method exhibits better perceptual performance than individual CP and ICA fusion approaches especially in noise corrupted images.
Zaid Omar, Nikolaos Mitianoudis, Tania Stathaki
ICASSP2
2010 A Directional Laplacian Density for Underdetermined Audio Source Separation
Nikolaos Mitianoudis
ICANN (1)1
2010 Two-dimensional Chebyshev polynomials for image fusion
abstract
This report documents in detail the research carried out by the author throughout his first year. The paper presents a novel method for fusing images in a domain concerning multiple sensors and modalities. Using Chebyshev polynomials as basis functions, the image is decomposed to perform fusion at feature level. Results show favourable performance compared to previous efforts on image fusion, namely ICA and DT-CWT, in noise affected images. The work presented here aims at providing a novel framework for future studies in image analysis and may introduce innovations in the fields of surveillance, medical imaging and remote sensing.
Zaid Omar, Nikolaos Mitianoudis, Tania Stathaki
PCS2
2009 A Unifying Approach to Moment-Based Shape Orientation and Symmetry Classification
abstract
In this paper, the problem of moment-based shape orientation and symmetry classification is jointly considered. A generalization and modification of current state-of-the-art geometric moment-based functions is introduced. The properties of these functions are investigated thoroughly using Fourier series analysis and several observations and closed-form solutions are derived. We demonstrate the connection between the results presented in this work and symmetry detection principles suggested from previous complex moment-based formulations. The proposed analysis offers a unifying framework for shape orientation/symmetry detection. In the context of symmetry classification and matching, the second part of this work presents a frequency domain method, aiming at computing a robust moment-based feature set based on a true polar Fourier representation of image complex gradients and a novel periodicity detection scheme using subspace analysis. The proposed approach removes the requirement for accurate shape centroid estimation, which is the main limitation of moment-based methods, operating in the image spatial domain. The proposed framework demonstrated improved performance, compared to state-of-the-art methods.
Georgios Tzimiropoulos, Nikolaos Mitianoudis, Tania Stathaki
IEEE Trans. Image Process.2
2008 Optimal contrast for color image fusion using ICA bases
Nikolaos Mitianoudis, Tania Stathaki
FUSION1
2007 An Affine Invariant Function using PCA Bases with an Application to Within-Class Object Recognition
abstract
The problem of shape-based recognition of objects under affine transformations is considered. We focus on the construction of a robust and highly discriminative affine invariant function that can be used for within-class object recognition applications. Using the boundaries of the objects of interest, a training scheme, based on principal component analysis (PCA), is proposed to derive a set of basis functions with desired properties. The derived bases are then used for the construction of a novel affine invariant function. The proposed invariant function is evaluated for the problem of aircraft silhouette identification and appears to achieve comparable performance to a popular wavelet-based affine invariant function. At the same time, the proposed framework is much simpler than that based on wavelet analysis.
Georgios Tzimiropoulos, Nikolaos Mitianoudis, Tania Stathaki
ICASSP (1)2
2007 Applied Multi-Dimensional Fusion
abstract
The purpose of the Applied Multi-dimensional Fusion Project is to investigate the benefits that data fusion and related techniques may bring to future military Intelligence Surveillance Target Acquisition and Reconnaissance systems. In the course of this work, it is intended to show the practical application of some of the best multi-dimensional fusion research in the UK. This paper highlights the work done in the area of multi-spectral synthetic data generation, super-resolution, joint fusion and blind image restoration, multi-resolution target detection and identification and assessment measures for fusion. The paper also delves into the future aspirations of the work to look further at the use of hyper-spectral data and hyper-spectral fusion. The paper presents a wide work base in multi-dimensional fusion that is brought together through the use of common synthetic data, posing real-life problems faced in the theatre of war. Work done to date has produced practical pertinent research products with direct applicability to the problems posed.
Asher Mahmood, Philip M. Tudor, William Oxford, Robert Hansford, James D. B. Nelson, Nick G. Kingsbury, Antonis Katartzis, Maria Petrou, Nikolaos Mitianoudis, Tania Stathaki, Alin Achim, David Bull 0001, Cedric Nishan Canagarajah, Stavri G. Nikolov, Artur Loza, Nedeljko Cvejic
Comput. J.9
2007 Joint Fusion and Blind Restoration For Multiple Image Scenarios With Missing Data
abstract
Image fusion systems aim at transferring ‘interesting’ information from the input sensor images to the fused image. The common assumption for most fusion approaches is the existence of a high-quality reference image signal for all image parts in all input sensor images. In the case that there are common degraded areas in at least one of the input images, the fusion algorithms cannot improve the information provided there, but simply convey a combination of this degraded information to the output. The authors propose a combined spatial-domain method of fusion and restoration in order to identify these common degraded areas in the fused image and use a regularized restoration approach to enhance the content in these areas. The proposed approach was tested on both multi-focus and multi-modal image sets and produced interesting results.
Nikolaos Mitianoudis, Tania Stathaki
Comput. J.1
2007 Smooth Signal Extraction From Instantaneous Mixtures
abstract
The problem of blind separation of statistically independent sources from instantaneous mixtures, using the efficient framework of independent component analysis (ICA), has been widely addressed in the literature. In this letter, the authors propose a sequential blind signal extraction algorithm that attempts to identify smooth sources in instantaneous mixtures. The approach incorporates smoothness constraints in the traditional negentropy cost function to extract smooth components, using an approximate second-order optimization method
Nikolaos Mitianoudis, Tania Stathaki, Anthony G. Constantinides
IEEE Signal Process. Lett.1
2007 Robust Recognition of Planar Shapes Under Affine Transforms Using Principal Component Analysis
abstract
A scheme, based on principal component analysis (PCA), is proposed that can be used for the recognition of 2-D planar shapes under affine transformations. A PCA step is first used to map the object boundary to its canonical form, reducing the problem of the nonuniform sampling of the object contour introduced by the affine transformation. Then, a PCA-based scheme is employed to train a set of basis functions on the signals extracted from the objects' boundaries. The derived bases are used to analyze the boundary locally. Based on the theory of invariants and local boundary analysis, a novel invariant function is constructed. The performance of the proposed framework is compared with a standard wavelet-based approach with promising results.
Georgios Tzimiropoulos, Nikolaos Mitianoudis, Tania Stathaki
IEEE Signal Process. Lett.2
2007 Batch and Online Underdetermined Source Separation Using Laplacian Mixture Models
abstract
In this paper, we explore the problem of sound source separation and identification from a two-sensor instantaneous mixture. The estimation of the mixing and the sources is performed using Laplacian mixture models (LMM). The proposed algorithm fits the model using batch processing of the observed data and performs separation using either a hard or a soft decision scheme. An extension of the algorithm to online source separation, where the samples are arriving in a real-time fashion, is also presented. The online version demonstrates several promising source separation possibilities in the case of nonstationary mixing.
Nikolaos Mitianoudis, Tania Stathaki
IEEE Trans. Speech Audio Process.1
2006 Adaptive Image Fusion Using Ica Bases
abstract
Image fusion can be viewed as a process that incorporates essential information from different modality sensors into a composite image. The use of bases trained using independent component analysis (ICA) for image fusion has been highlighted recently. Common fusion rules can be used in the ICA fusion framework with promising results. In this paper, the authors propose an adaptive fusion scheme, based on the ICA fusion framework, that maximises the sparsity of the fusion image in the transform domain
Nikolaos Mitianoudis, Tania Stathaki
ICASSP (2)1
2005 Overcomplete source separation using Laplacian mixture models
abstract
The authors explore the use of Laplacian mixture models (LMMs) to address the overcomplete blind source separation problem in the case that the source signals are very sparse. A two-sensor setup was used to separate an instantaneous mixture of sources. A hard and a soft decision scheme were introduced to perform separation. The algorithm exhibits good performance as far as separation quality and convergence speed are concerned.
Nikolaos Mitianoudis, Tania Stathaki
IEEE Signal Process. Lett.1
2003 Audio source separation of convolutive mixtures
abstract
The problem of separation of audio sources recorded in a real world situation is well established in modern literature. A method to solve this problem is blind source separation (BSS) using independent component analysis (ICA). The recording environment is usually modeled as convolutive. Previous research on ICA of instantaneous mixtures provided solid background for the separation of convolved mixtures. The authors revise current approaches on the subject and propose a fast frequency domain ICA framework, providing a solution for the apparent permutation problem encountered in these methods.
Nikolaos Mitianoudis, Mike E. Davies 0001
IEEE Trans. Speech Audio Process.1