Sayeh Mirzaei

dblp:139/7487 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 A novel framework for under-determined blind source separation based on adaptive source counting using mixed linear and circular data clustering algorithm for low latency applications
Mahdi Khademi, Sayeh Mirzaei, Yaser Norouzi
Multim. Tools Appl.2
2025 EEG Motor imagery classification based on a ConvLSTM Autoencoder framework augmented by attention BiLSTM
Sayeh Mirzaei, Parisa Ghasemi, Mohammadreza Bakhtyari
Multim. Tools Appl.1
2025 Enhancing aspect-based sentiment analysis through graph attention networks and supervised contrastive learning
Akram Karimi Zarandi, Sayeh Mirzaei
Multim. Tools Appl.2
2024 A survey of aspect-based sentiment analysis classification with a focus on graph neural network methods
Akram Karimi Zarandi, Sayeh Mirzaei
Multim. Tools Appl.2
2023 Hyperspectral image classification using K-plane clustering and kernel principal component analysis
Sayeh Mirzaei
Multim. Tools Appl.1
2023 Acoustic scene classification with multi-temporal complex modulation spectrogram features and a convolutional LSTM network
Sayeh Mirzaei, Iman Khani Jazani
Multim. Tools Appl.1
2019 Hyperspectral image classification using Non-negative Tensor Factorization and 3D Convolutional Neural Networks
Sayeh Mirzaei, Hugo Van hamme, Shima Khosravani
Signal Process. Image Commun.1
2016 Under-determined reverberant audio source separation using Bayesian Non-negative Matrix Factorization
Sayeh Mirzaei, Hugo Van hamme, Yaser Norouzi
Speech Commun.1
2015 Who's Speaking?: Audio-Supervised Classification of Active Speakers in Video
abstract
Active speakers have traditionally been identified in video by detecting their moving lips. This paper demonstrates the same using spatio-temporal features that aim to capture other cues: movement of the head, upper body and hands of active speakers. Speaker directional information, obtained using sound source localization from a microphone array is used to supervise the training of these video features.
Punarjay Chakravarty, Sayeh Mirzaei, Tinne Tuytelaars, Hugo Van hamme
ICMI2
2015 Two-stage blind audio source counting and separation of stereo instantaneous mixtures using Bayesian tensor factorisation
abstract
In this paper, the authors address the tasks of audio source counting and separation for two‐channel instantaneous mixtures. This goal is achieved in two steps. First, a novel scheme is proposed for estimating the number of sources and the corresponding channel intensity difference (CID) values. For this purpose, an angular spectrum is evaluated as a function of the ratio of the magnitude spectrogram of the two channels and the peak locations of that spectrum are obtained. In the second stage, a new approach is developed for extracting the individual source signals exploiting a Bayesian non‐parametric modelling. The mean field variational Bayesian approach is applied for inferring the unknown parameters. Classification is then performed on the inferred active CID values to obtain the individual source magnitude spectrograms. This way, the number of spectral components used for modelling each source is found automatically from the data. The Bayesian approach is compared with the standard Kullback–Leibler non‐negative tensor factorisation method to illustrate the effectiveness of Bayesian modelling. The performance of the source separation is measured by obtaining the existing metrics for multichannel blind source separation evaluation. The experiments are performed on instantaneous mixtures from the dev2 database.
Sayeh Mirzaei, Yaser Norouzi, Hugo Van hamme
IET Signal Process.1
2015 Blind audio source counting and separation of anechoic mixtures using the multichannel complex NMF framework
abstract
In this paper, we address the tasks of audio source counting and separation for a stereo anechoic mixture of audio signals. This will be achieved in two stages. In the first stage, a novel approach is introduced for estimating the number of sources as well as the channel mixing coefficients. For this purpose, a 2-D spectrum is evaluated against both the phase and amplitude differences of the two channels. Hence, obtaining the peak locations of the spectrum yields the number of the sources and the corresponding channel coefficients. In the second stage, an extension of a single channel complex matrix factorization method to multichannel is developed to extract the individual source signals. We find primary estimates of the sources via binary masking and then apply the complex factorization to the complex spectrogram of each source. The obtained factors are then utilized as initial values in the complex multichannel factorization model. We also suggest a method for estimating the number of required components for modeling each source. The separation performance improvement over the conventional methods is investigated by calculating BSS evaluation metrics. The comparison is also carried out in terms of source counting and localization with the recently proposed DeMIX-Anechoic method.
Sayeh Mirzaei, Hugo Van hamme, Yaser Norouzi
Signal Process.1
2014 Blind speech source localization, counting and separation for 2-channel convolutive mixtures in a reverberant environment
abstract
Copyright © 2014 ISCA. In this paper, the tasks of speech source localization, source counting and source separation are addressed for an unknown number of sources in a stereo recording scenario. In the first stage, the angles of arrival of individual source signals are estimated through a peak finding scheme applied to the angular spectrum which has been derived using non-linear GCC-PHAT. Then, based on the known channel mixture coefficients, we propose an approach for separating the sources based on Maximum Likelihood (ML) estimation. The predominant source in each time-frequency bin is identified through ML assuming a diffuse noise model. The separation performance is improved over a binary time-frequency masking method. The performance is measured by obtaining the existing metrics for blind source separation evaluation. The experiments are performed on synthetic speech mixtures in both anechoic and reverberant environments.
Sayeh Mirzaei, Hugo Van hamme, Yaser Norouzi
INTERSPEECH1
2013 Model order estimation using Bayesian NMF for discovering phone patterns in spoken utterances
abstract
In earlier work, we have shown that vocabulary discovery from spoken utterances and subsequent recognition of the acquired vocabulary can be achieved through Non-negative Matrix Factorization (NMF). An open issue for this task is to determine automatically how many different word representations should be included in the model. In this paper, Bayesian NMF is applied to estimate the model order. The per-utterance word activations are given a gamma prior while the word models are assumed deterministic. Two Bayesian approaches are applied for obtaining optimal parameter values. First, the penalized joint log-likelihood of the parameters is considered as the objective function. Then, maximal marginal likelihood estimator (MMLE) is implemented which obtains the word models maximizing the likelihood after integration over the activations. The variational Bayesian algorithm, which maximizes a lower bound of the marginal log-likelihood, is applied to this optimization problem. The number of required latent components or basis vectors (model order) is estimated by evaluating likelihood metrics. The inferred model order is validated by observing error criteria on a test set. Experiments on synthetic data as well as real speech show that MMLE is more effective for the purpose of model order selection. Copyright © 2013 ISCA.
Sayeh Mirzaei, Hugo Van hamme, Yaser Norouzi
INTERSPEECH1