EDBT 2026 Demo / reviewers in the wild / expert
Yassine Ben Ayed
dblp:44/3444 · also Yassine Benayed
· DBLP profile ↗
55ranked-venue papers
1as first author
25since 2021 · last 2026
0000-0002-3676-3670ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-level fusion framework for Alzheimer's recognition using speech and textual features
Dhouha Guesmi, Hasna Njah, Yassine Ben Ayed |
Neurocomputing | 3 |
| 2026 | Transformer encoder and data augmentation for real-time speech emotion recognition
Chawki Barhoumi, Yassine Ben Ayed |
Multim. Tools Appl. | 2 |
| 2025 | Optimized Hybrid Deep Learning Model for Accurate Classification of Alzheimer's Stages
Maysam Chaari, Yassine Ben Ayed |
AINA (4) | 2 |
| 2025 | MLP-Mixer for Automatic Detection of Parkinson's Disease from Speech
Rania Khaskhoussy, Yassine Ben Ayed |
AINA (3) | 2 |
| 2025 | Person Identification with Arrhythmic and Normal ECG Signals Using Hybrid Machine Learning and Deep Learning Models
Sihem Hamza, Yassine Ben Ayed |
ITS (2) | 2 |
| 2025 | A New Attention-based Architecture for Robust Face-Age RecognitionabstractFace age recognition task has become essential in several applications like minors security and marketing. However, despite the recent advances in deep learning algorithms, developing accurate and efficient architecture for face age classification remains difficult due to the unconstrained environment and variation on aging effect. Requiring both a distinctive and descriptive face age representation model. Requiring both a distinctive and descriptive face age representation models. We introduce, in this paper, a three-stage system to enhance face age recognition using the transformer architecture and pre-trained Convolutional Neural Networks (CNN). The system begins with extracting the age representation using Vision Transforms (ViT) and FaceNet model. Subsequently, the Pearson correlation coefficient is considered for selecting the face age pertinent features. Hence, a Multi-Head Attention (MHA) block is used to extract the final descriptive representation. Finally, the Support Vector Machine Classifier (SVM) is trained to distinguish between the different age classes. The challenging FG-NET dataset was considered in our work for evaluating the proposed architecture. The conducted experimental study shows evidence that our approach ranks top of the list of recent potential published work with an 63% recognition rate. Amal Abbes, Yassine Ben Ayed |
KES | 2 |
| 2025 | Cross-Attention-Enhanced Multimodal Fake News Detection using Autoencoder-based Fusion and Transformer-based modelsabstractWith the rapid spread of information online, distinguishing between real and fake content is crucial. Fake news often manipulates both textual and visual modalities to mislead readers, posing significant challenges for detection systems. A key limitation of many existing methods lies in their inability to effectively learn a unified representation of multimodal data. In this study, we propose a new neural network architecture for fake news detection that integrates a bimodal transformer-based model with a binary classification module. Our model comprises three main components: (1) feature extraction from textual and visual inputs(2) a hybrid multimodal fusion module that combines autoencoder-based early fusion with a cross-attention mechanism to capture deep inter-modal relationships, and (3) a classifier for final fake news prediction. We evaluate our model on benchmark datasets, including Fakedit and FakeNewsNet, and demonstrate its superiority over existing methods. The experimental results show significant improvements in accuracy, precision, recall, and F1-score, validating the effectiveness and robustness of our cross-attention-enhanced multimodal framework. Rahma Ghorbel, Hanen Ameur, Yassine Ben Ayed |
KES | 3 |
| 2025 | Efficient Deep Learning Model for Heart Disease ClassificationabstractHeart disease remains a leading global health challenge, emphasizing the urgent need for precise and efficient diagnostic techniques. This study proposes an advanced approach for the automated classification of heart diseases using ElectroCardioGram (ECG) signals and deep learning. Specifically, it examines the effectiveness of two modified models, namely EfficientNet B0 and EfficientNet B3, to categorize cardiac conditions into six distinct classes. To adapt these models to our specific task, we employ transfer learning by fine-tuning both architectures. The methodology consists of three main stages: signal pre-processing, feature extraction using Short-Time Fourier Transform (STFT) to generate spectrogram representations, and classification. The performance of the proposed approach was evaluated using the Massachusetts Institute of Technology Beth Israel Hospital Arrhythmia database (MIT-BIH database). Notably, EfficientNet B0 achieved the highest accuracy, reaching 99.14% while EfficientNet B3 achieve accuracy equal to 98.62%. Hend Karoui, Sihem Hamza, Yassine Ben Ayed |
KES | 3 |
| 2024 | Deep Learning for Cardiac Diseases Classification
Hend Karoui, Sihem Hamza, Yassine Ben Ayed |
ICCCI (1) | 3 |
| 2024 | Age-API: are landmarks-based features still distinctive for invariant facial age recognition?
Amal Abbes, Wael Ouarda, Yassine Ben Ayed |
Multim. Tools Appl. | 3 |
| 2024 | Shot boundary detection using multimodal Siamese network
Mohamed Bouyahi 0002, Yassine Ben Ayed |
Multim. Tools Appl. | 2 |
| 2024 | Automated detection of COVID-19 based on transfer learning
Amira Echtioui, Yassine Ben Ayed |
Multim. Tools Appl. | 2 |
| 2023 | Improving Speech Emotion Recognition Using Data Augmentation and Balancing TechniquesabstractSpeech Emotion Recognition (SER) is a challenging task due to the complexity and variability of human emotions. In this paper, we propose an innovative approach to improve SER performance on the EMODB dataset. Our approach employs data augmentation techniques, such as noise addition and spectrogram shift, as well as balancing techniques, including random oversampling. We also extract five different features from the dataset samples: MFCC, Chroma, Mel Spectrogram, ZCR, and RMS. We compare the performance of four different classifiers - MLP, SVM, KNN, and CNN - with and without the use of our proposed approach. Our results demonstrate that the proposed approach significantly enhances the accuracy of speech emotion recognition compared to the approach without data augmentation and balancing techniques. Our experiments reveal that the proposed approach achieves higher accuracy and F1-score compared to other approaches, with MLP and CNN models achieving 100% accuracy. These findings highlight the effectiveness of data augmentation and balancing techniques in improving the performance of speech emotion recognition. Moreover, our approach holds great potential for application in various real-life scenarios, including mental health monitoring, human-robot interaction, and speech-based virtual assistants. Chawki Barhoumi, Yassine Ben Ayed |
CW | 2 |
| 2023 | A New Deep Learning Method for Delineating Early Gastrointestinal Cancer
Intissar Dhrari Hajsalem, Amal Abbes, Yassine Ben Ayed |
HIS (5) | 3 |
| 2023 | Unlocking the Potential of Deep Learning and Filter Gabor for Facial Emotion Recognition
Chawki Barhoumi, Yassine Ben Ayed |
ICCCI | 2 |
| 2023 | Improving Parkinson's disease recognition through voice analysis using deep learning
Rania Khaskhoussy, Yassine Ben Ayed |
Pattern Recognit. Lett. | 2 |
| 2022 | A Deep Convolutional Autoencoder-Based Approach for Parkinson's Disease Diagnosis Through Speech Signals
Rania Khaskhoussy, Yassine Ben Ayed |
ADMA (1) | 2 |
| 2022 | Detection of Heart Diseases Using CNN-LSTM
Hend Karoui, Sihem Hamza, Yassine Ben Ayed |
HIS | 3 |
| 2022 | Toward improving person identification using the ElectroCardioGram (ECG) signal based on non-fiducial features
Sihem Hamza, Yassine Ben Ayed |
Multim. Tools Appl. | 2 |
| 2022 | Voice spoofing detection based on acoustic and glottal flow features using conventional machine learning techniques
Raoudha Rahmeni, Anis Ben Aicha, Yassine Ben Ayed |
Multim. Tools Appl. | 3 |
| 2021 | An I-vector-based approach for discriminating between patients with Parkinson's disease and healthy peopleabstractParkinson's disease (PD) is one of the neurological disorders that affect the central nervous system leads to cognitive, emotional and speech disorders. Many methods have been proposed over time for discriminating between people with PD and healthy people using signals processing. In this paper, a new approach is defined using i-vector subspace modelling to discriminate healthy people from people with PD. The i-vectors features is one of the crucial parameters that prove promising results in the domain of speech recognition. In this study two i-vectors dimensionality (100 and 200 dimensions) extracted from voice recordings using Gaussian Mixture Models based on Universal Background Model (GMM-UBM) size (64, 128 and 256 Gaussians). To the end, we assess the effect of the i-vectors features by using Support Vector Machine (SVM). The results reveal show that the proposed approach can be strongly recommended for classifying Parkinson's patient from healthy individuals. Rania Khaskhoussy, Yassine Ben Ayed |
ICMV | 2 |
| 2021 | Age Estimation and Gender Recognition Using Biometric Modality
Amal Abbes, Randa Boukhris Trabelsi, Yassine Ben Ayed |
ISDA | 3 |
| 2021 | Recognition of Person Using ECG Signals Based on Single Heartbeat
Sihem Hamza, Yassine Ben Ayed |
ISDA | 2 |
| 2021 | Detecting Parkinson's Disease According to Gender Using Speech Signals
Rania Khaskhoussy, Yassine Ben Ayed |
KSEM | 2 |
| 2021 | A Language-Based Approach for Predicting Alzheimer Disease Severity
Randa Ben Ammar, Yassine Ben Ayed |
RCIS | 2 |
| 2020 | Multimodal features for shots boundary detectionabstractShot Boundary Detection (SBD) also known as a temporal video segmentation is a preprocessing task for multiple videos applications, such as indexing and retrieval. The SBD output provides coherent temporal units which are easy to manipulate. The Most previous works implement theirs frameworks based on visual features to measure similarity for transition detection task. However, the video is very enriched by data which could be beneficial. In this paper, referring to recent multimodal works, we propose to introduce the audio components to increase the SBD task. Firstly, we worked on candidate segments obtained by measuring similarity between low features (SURF, HSF) from original video. Then we used deep features obtained from trained model (Resnet-50) for visual similarity and we introduced the audio segmentation based on Power Spectrum Density (PSD) to contribute for transition detection. The proposed method is evaluated on the clip shots dataset. Experiments on this data show that the proposed multimodal approach can achieve a better performance compared with the state-of-the-art of methods that used visual approach. Mohamed Bouyahi 0002, Yassine Ben Ayed |
ICMV | 2 |
| 2020 | Abstractive meeting summarization based on an attentional neural modelabstractThrough the ages, in all nations, at all times, people spend a lot of their time on discussing new and important issues either on meetings or in conferences. With the evolution and the abundance of Automatic Speech Recognition (ASR) frameworks, automatic transcripts and even automatic meeting summarization are getting more and more interesting. Recently, automatic summarization faces deeper progresses on speech summarization. Neural models had been introduced to tackle with many difficulties of abstractive summarization. Our contribution in this paper focuses on these weaknesses of neural abstractive meeting summarization and suggests an encoder-decoder model based on an attentional algorithm on the decoding sequence. We proposed a deep encoder-decoder model based on attention mechanism (DEDA) for ASR transcripts. Experiments on the AMI Dataset demonstrates that our proposed method ensured competitive results with the state of the art even on extractive or abstractive models. The experimental analyses also put the stress on the performance of the summarized utterances as well as the reduction of the occurrence repetition in summaries. Nouha Dammak, Yassine Ben Ayed |
ICMV | 2 |
| 2020 | Bimodal person recognition using dorsal-vein and finger-vein imagesabstractNowadays, human recognition with biometric characteristic have been investigated in many researchers. However, traditional biometric methods are not always robust against counterfeit and spoof attacks. The challenge of this work is to develop a bimodal biometric system, based on more reliable and robust characteristics against counterfeiting. Indeed, biometric system based on the venous network of the hand gives higher recognition rate compared to other systems. Furthermore, the fusion of two prints minimums the error rate and counterfeits. In our biometric system, the most challenging phase is the feature extraction step. For this, we propose a new feature extraction approach based on the concatenation of the pyramid of Difference of Gaussian (DoG) and Local Line Binary Patterns (LLBP) histograms (H-DoG_LLBP). To evaluate the proposed system, we opted for Support Vector Machine (SVM) and Artificial Neural Network (ANN) classifiers. The experimental study is based on the dorsal hand vein BOSPHORUS database and finger vein MMCBNU_6000 database. Our system presents an area under curve (AUC) equal to 0.99, and 0.001 mean square error (MSE) with ANN, and 0.0042 equal error rate (EER) with SVM. Amal Abbes, Randa Boukhris Trabelsi, Yassine Ben Ayed |
KES | 3 |
| 2020 | Language-related features for early detection of Alzheimer DiseaseabstractAlzheimer’s disease (AD) is the most leading symptom of neurodegenerative dementia; AD is defined now as one of the most costly chronic diseases. For that automatic diagnosis and control of Alzheimer’s disease may have a significant effect on society along with patient well-being. Language disorder is regarded to be among the most common symptoms of AD, as a direct and natural result of cognitive impairment. Hence, the diagnosis of Alzheimer’s disease using speech-based features is gaining growing attention. The aim of this study is to extract linguistic features following a proposed taxonomy of language impairment of AD patients. Obtained results indicate that the proposed taxonomy of the linguistic features extracted from the speech samples can be used to differentiate between Alzheimer’s disease patients and the healthy control group. Support Vector Machine (SVM) classifier obtained classification accuracy over 90 percent. Randa Ben Ammar, Yassine Ben Ayed |
KES | 2 |
| 2020 | Speech Emotion Recognition with deep learningabstractThis paper proposes an emotion recognition system based on speech signals in two-stage approach, namely feature extraction and classification engine. Firstly, two sets of feature are investigated which are: the first one, we extract an 42-dimensional vector of audio features including 39 coefficients of Mel Frequency Cepstral Coefficients (MFCC), Zero Crossing Rate(ZCR), Harmonic to Noise Rate (HNR) and Teager Energy Operator (TEO). And the second one, we propose the use of the method Auto-Encoder for the selection of pertinent parameters from the parameters previously extracted. Secondly, we use the Support Vector Machines (SVM) as a classifier method. Experiments are conducted on the Ryerson Multimedia Laboratory (RML). Hadhami Aouani, Yassine Ben Ayed |
KES | 2 |
| 2020 | Video Scenes Segmentation Based on Multimodal Genre PredictionabstractRecent technologies’ understanding videos content remain limited due to its complexity and length. However, videos segmentation into small coherent units facilitates indexing and searching task. The subjectivity remains the essential constraint of videos, but the genre (drama, action...) does not present any conflict. In this paper, we present a new approach to video segmentation into scenes based on genre prediction. Initially, the video is divided into shots of equal duration. We used architecture, based on audio-visuals deep features extracted from trained neural networks for genre prediction, and we introduced a transition detection method based on the similarity calculation between shots genre. The originality of this method consists in using the highly level semantic relationship between successive shots for transition detection. We reached good performances on videos of the multi varied genre. We used the RAI dataset and BBC dataset to evaluate our method through a comparison with other state-of-the-art approaches. Mohamed Bouyahi 0002, Yassine Ben Ayed |
KES | 2 |
| 2020 | Svm for human identification using the ECG signalabstractIn this paper, a person identification system has been simulated using electrocardiogram (ECG) signals as biometrics. In this work, we propose a two-phase method to conduct human identification using the ECG signal, which are the feature extraction and the classification. In the first phase, it makes a fusion of three new types of characteristics: cepstral coefficients, ZCR, and entropy. In the second phase, the support vector machines (SVM) has been applied for the classification system. The proposed methods are evaluated using two public databases namely MIT-BIH arrhythmia and ECG-ID database obtained from the Physionet database. Experimental results show that our features can achieve high subject identification accuracy of 100% on ECG signals that are from the MIT-BIH database, ECG-ID (Five recording), and ECG-ID (Two recording), indicating that our features makes it possible to improve the efficiency of our identification system. Sihem Hamza, Yassine Ben Ayed |
KES | 2 |
| 2020 | Acoustic features exploration and examination for voice spoofing counter measures with boosting machine learning techniquesabstractAutomatic Speaker verification systems are vulnerable to spoofing attacks. We propose our anti-spoofing system. It uses some acoustic features. The set of binary classifiers includes XGBoost tree boosting algorithm. ASV Spoof 2015 corpus is utilized in the experiments as the main database for anti-spoofing systems training. The pretreatment of acoustic features is essential for better performance of the system. Obtained results demonstrate that the proposed system can provide a good accuracy. High evaluation performance can be obtained using the combination between MFCC, LogFBE and SSC features. The attained metrics values such 97.80% for the accuracy, 97.12% for the precision, 98.50% for the Recall and 97.81% for the F1-mesure validate the performance of the proposed technique. Raoudha Rahmeni, Anis Ben Aicha, Yassine Ben Ayed |
KES | 3 |
| 2019 | ADAL System: Aspect Detection for Arabic Language
Sana Trigui, Ines Boujelben, Salma Jamoussi, Yassine Ben Ayed |
HIS | 4 |
| 2019 | Evaluation of Acoustic Features for Early Diagnosis of Alzheimer Disease
Randa Ben Ammar, Yassine Ben Ayed |
ISDA | 2 |
| 2019 | Deep Support Vector Machines for Speech Emotion Recognition
Hadhami Aouani, Yassine Ben Ayed |
ISDA | 2 |
| 2019 | Histogram Based Method for Unsupervised Meeting Speech Summarization
Nouha Dammak, Yassine Ben Ayed |
ISDA | 2 |
| 2019 | Biometric Individual Identification System Based on the ECG Signal
Sihem Hamza, Yassine Ben Ayed |
ISDA | 2 |
| 2019 | Automatic Detection of Parkinson's Disease from Speech Using Acoustic, Prosodic and Phonetic Features
Rania Khaskhoussy, Yassine Ben Ayed |
ISDA | 2 |
| 2019 | Speech spoofing countermeasures based on source voice analysis and machine learning techniquesabstractAutomatic speaker verification (ASV) [7] systems are susceptible to malicious attacks. It discredit the performance of a standard ASV system by increasing its false acceptance rates. This paper presents a new countermeasure for the protection of automatic speaker verification systems from spoofed signals. The new countermeasure is based on the analysis of a sequence of acoustic feature vectors using the glottal inverse filtering. In the proposed method, speech is decomposed into a glottal source signal and model the vocal tract filter through glottal inverse filtering. The IAIF desriptors are constructed and are used as features. Support Vector Machines (SVM) classifier and Extreme learning machine (ELM) are used to classify the obtained features as genuine or spoofed. It is hoped that the proposed method can help to detect the genuine speech from the spoofed one. Raoudha Rahmeni, Anis Ben Aicha, Yassine Ben Ayed |
KES | 3 |
| 2019 | Spoken keyword search system using improved ASR engine and novel template-based keyword scoring
Ilyes Rebai, Yassine Ben Ayed, Walid Mahdi |
Multim. Tools Appl. | 2 |
| 2018 | Speech Processing for Early Alzheimer Disease Diagnosis: Machine Learning Based ApproachabstractAlzheimer's disease (AD) is a neurodegenerative disease characterized by the insidious onset of cognitive, emotional and language disorders. These attacks are sufficiently intense to affect the daily social and professional lives of patients. Today, in the absence of a reliable diagnosis and effective curative treatments, fighting this disease is becoming a real public health issue, prompting research to consider non-drug techniques. Among these techniques, speech processing is proving to be a relevant and innovative field of investigation. Several Machine Learning algorithms achieved promising results in distinguishing AD from healthy control subjects. Alternatively, many other factors such as feature extraction, the number of attributes for feature selection, used classifiers, may affect the prediction accuracy evaluation. To surmount these weaknesses, a model is suggested which include a feature extraction step followed by imperative attribute selection and classification is achieved using a machine learning classifiers. The current findings show that the proposed model can be strongly recommended for classifying Alzheimer's patient from healthy individuals with an accuracy of 79%. Randa Ben Ammar, Yassine Ben Ayed |
AICCSA | 2 |
| 2018 | A novel keyword rescoring method for improved spoken keyword spottingabstractIn this paper, we present a spoken KeyWord Spotting (KWS) system which creates a search index from word lattices generated by a deep speech recognizer. Basic KWS systems estimate word posteriors from the lattices and use them to make “correct/false alarm” decisions. The main issue of lattice-based posterior probability is that a putative detection can have very low posterior probability so that the decider fails to detect it and considers it as a false alarm. Therefore, our goal is to enhance the keyword decision by detecting and boosting the score of missed detections. Accordingly, inspired by template matching approach, we propose a new keyword rescoring method. More precisely, detected hits are rescored based on the acoustic similarity and the new score are used then by the decider to make the final decision. Experiments demonstrate that the proposed method potentially leads to more accurate keyword results than the conventional KWS system. Ilyes Rebai, Yassine Ben Ayed, Walid Mahdi |
KES | 2 |
| 2017 | Improving of Open-Set Language Identification by Using Deep SVM and Thresholding FunctionsabstractState-of-the-art language identification (LID) systems are based on an iVector feature extractor front-end followed by a multi-class recognition back-end. Identification accuracy degrades considerably when LID systems face open-set languages. As compared to in-set identification task, the open-set task is adequate to mimic the real challenge of language identification. In this paper, we propose an approach to the problem of out-of-set (OOS) data detection in the context of open-set language identification with zero-knowledge for OOS languages. The main feature of this study is the emphasis on the in-set (target) language identification, on the one hand, and on OOS language detection, on the other hand. Accordingly, we propose a deep SVM based LID back-end system to improve the target languages identification. Along with that, we define three OOS thresholding formulations. These formulations are used to decide whether the speech segment is a target or OOS language. The experimental results demonstrate the effectiveness of the deep SVM back-end system as compared to state-of-the-art techniques. Besides that, the thresholding functions perfectly detect and reject the OOS data. A relative decrease of 6% in Equal Error Rate (EER) is reported over classical OOS detection methods, in discriminating target and OOS languages. Ilyes Rebai, Yassine Ben Ayed, Walid Mahdi |
AICCSA | 2 |
| 2017 | Hierarchical vs non-hierarchical audio indexation and classification for video genresabstractIn this paper, Support Vector Machines (SVMs) are used for segmenting and indexing video genres based on only audio features extracted at block level, which has a prominent asset by capturing local temporal information. The main contribution of our study is to show the wide effect on the classification accuracies while using an hierarchical categorization structure based on Mel Frequency Cepstral Coefficients (MFCC) audio descriptor. In fact, the classification consists in three common video genres: sports videos, music clips and news scenes. The sub-classification may divide each genre into several multi-speaker and multi-dialect sub-genres. The validation of this approach was carried out on over 360 minutes of video span yielding a classification accuracy of over 99%. Nouha Dammak, Yassine Ben Ayed |
ICMV | 2 |
| 2017 | Improving speech recognition using data augmentation and acoustic model fusionabstractDeep learning based systems have greatly improved the performance in speech recognition tasks, and various deep architectures and learning methods have been developed in the last few years. Along with that, Data Augmentation (DA), which is a common strategy adopted to increase the quantity of training data, has been shown to be effective for neural network training to make invariant predictions. On the other hand, Ensemble Method (EM) approaches have received considerable attention in the machine learning community to increase the effectiveness of classifiers. Therefore, we propose in this work a new Deep Neural Network (DNN) speech recognition architecture which takes advantage from both DA and EM approaches in order to improve the prediction accuracy of the system. In this paper, we first explore an existing approach based on vocal tract length perturbation and we propose a different DA technique based on feature perturbation to create a modified training data sets. Finally, EM techniques are used to integrate the posterior probabilities produced by different DNN acoustic models trained on different data sets. Experimental results demonstrate an increase in the recognition performance of the proposed system. Ilyes Rebai, Yassine Ben Ayed, Walid Mahdi, Jean-Pierre Lorré |
KES | 2 |
| 2016 | Deep kernel-SVM networkabstractDeep learning techniques have claimed state-of-the-art results in a wide range of tasks, including classification. Despite the promising results, there are limitations for these large networks. In fact, deep neural networks have a poor generalisation performance on small data sets, such as biologic data. This paper describes a new machine learning algorithm for classification tasks. We introduce a Multi-Layer Multiple Kernel Learning (ML-MKL) framework. The input data are first transformed through a set of weighted non-linear kernel functions in a multilayer structure. Then, an SVM classifier is used to make the final decision. The proposed network is trained to minimize the error function. Indeed, we propose to optimize the network over an adaptive backpropagation algorithm. The generalization performance of the proposed method is compared over various state-of-the-art multiple kernel algorithms on several benchmark and two real world applications, including object recognition and spoken language recognition. Experimental results show that the ML-MKL generally outperforms existing kernel methods. Ilyes Rebai, Yassine Ben Ayed, Walid Mahdi |
IJCNN | 2 |
| 2016 | Deep multilayer multiple kernel learning
Ilyes Rebai, Yassine Ben Ayed, Walid Mahdi |
Neural Comput. Appl. | 2 |
| 2015 | Indexing and classifiying video genres using Support Vector MachinesabstractIn this paper, classifying and indexing hierarchical video genres using Support Vector Machines (SVMs) are based on only audio features. In fact, segmentation parameters are extracted at block levels, which have a major benefit by capturing local temporal information. The main contribution of our study is to present a powerful combination between the two employed audio descriptors; Mel Frequency Cepstral Coefficients (MFCC) and signal energy in order to classify a big YouTube dataset that includes multi-Arabic dialects video genres and even sub-genres: several sports analysis and various matches categories (foot-ball, basket-ball, hand-ball and volley-ball), both studio and fields news scenes over and above various multi-singer and multi-instruments music clips. Validation of this approach was carried out on over 18 hours of video span yielding a classification accuracy of 98,5% for genres, 97% for sports sub-genres and 76% for music sub-genres. Finally we discuss SVM kernels performance on our proposed dataset. Nouha Dammak, Yassine Ben Ayed |
AICCSA | 2 |
| 2015 | Deep architecture using Multi-Kernel Learning and multi-classifier methodsabstractKernel Methods have been successfully applied in different tasks and used on a variety of data sample sizes. Multiple Kernel Learning (MKL) and Multilayer Multiple Kernel Learning (MLMKL), as new families of kernel methods, consist of learning the optimal kernel from a set of predefined kernels by using an optimization algorithm. However, learning this optimal combination is considered to be an arduous task. Furthermore, existing algorithms often do not converge to the optimal solution (i.e., weight distribution). They achieve worse results than the simplest method, which is based on the average combination of base kernels, for some real-world applications. In this paper, we present a hybrid model that integrates two methods: Support Vector Machine (SVM) and Multiple Classifier (MC) methods. More precisely, we propose a multiple classifier framework of deep SVMs for classification tasks. We adopt the MC approach to train multiple SVMs based on multiple kernel in a multi-layer structure in order to avoid solving the complicated optimization tasks. Since the average combination of kernels gives high performance, we train multiple models with a predefined combination of kernels. Indeed, we apply a specific distribution of weights for each model. To evaluate the performance of the proposed method, we conducted an extensive set of classification experiments on a number of benchmark data sets. Experimental results show the effectiveness and efficiency of the proposed method as compared to various state-of-the-art MKL and MLMKL algorithms. Ilyes Rebai, Yassine Ben Ayed, Walid Mahdi |
AICCSA | 2 |
| 2015 | Graphical Models for Multi-dialect Arabic Isolated Words RecognitionabstractThis paper presents the use of multiple hybrid systems for the recognition of isolated words from a large multi-dialect Arabic vocabulary. Such as the Hidden Markov models (HMM), Dynamic Bayesian networks (DBN) lack a discriminatory ability especially on speech recognition even if their progress is huge. Multi-Layer perceptrons (MLP) was applied in literature as an estimator of emission probabilities in HMM and proves it effectiveness. In order to ameliorate the results of recognition systems, we apply Support Vectors Machine (SVM) as an estimator of posterior probabilities since they are characterized by a high predictive power and discrimination. Moreover, they are based on a structural risk minimization (SRM) where the aim is to set up a classifier that minimizes a bound on the expected risk, rather than the empirical risk. In this work we have done a comparative study between three hybrid systems MLP/HMM, SVM/HMM and SVM/DBN and the standards models of HMM and DBN. In this paper, we describe the use of the hybrid model SVM/DBN for multi-dialect Arabic isolated words recognition. So, by using 67,132 speech files of Arabic isolated words, this work arises a comparative study of our acknowledgment system of it as the following: the use of especially the HMM standards leads to a recognition rate of 74.18%.as the average rate of 8 domains for everyone of the 4 dialects. Also, with the hybrid systems MLP/HMM and SVM/HMM we succeed in achieving the value of 77.74%.and 7806% respectively. Moreover, our proposed system SVM/DBN realizes the best performances, whereby, we achieve 87.67% as a recognition rate more than 83.01% obtained by GMM/DBN. Elyes Zarrouk, Yassine Ben Ayed, Faïez Gargouri |
KES | 2 |
| 2015 | Graphical models for the recognition of Arabic continuous speech based triphones modelingabstractRecent developments in inference and learning in Dynamic Bayesian networks (DBN) allow their use in real-world applications is the first successful application of DBNs to a large scale speech recognition problem. Even if their progress is huge, those models lack a discriminatory ability especially on speech recognition such as the Hidden Markov models (HMM). In this paper, we present the performance of the hybridization of Supports Vectors machine with Dynamic Bayesian networks for Arabic triphones-based continuous speech. In fact, SVM are based on a structural risk minimization (SRM) where the aim is to set up a classifier that minimizes a bound on the expected risk, rather than the empirical risk. The best results are obtained with the proposed system SVM/DBN when we achieve 78.87% as the best recognition rate of a tested speaker. The speech recognizer was evaluated with ARABIC_DB corpus and performs at 8.04% WER as compared to 10.08% with triphones mixture-Gaussian DBN system, 10.54% with hybrid model SVM/HMM and 12.03% with HMM standards. Elyes Zarrouk, Yassine Ben Ayed, Faïez Gargouri |
SNPD | 2 |
| 2015 | Text-to-speech synthesis system with Arabic diacritic recognition system
Ilyes Rebai, Yassine Ben Ayed |
Comput. Speech Lang. | 2 |
| 2014 | A Multi-objective Genetic Algorithm for Model Selection for Support Vector Machines
Amal Bouraoui, Yassine Ben Ayed, Salma Jamoussi |
PRICAI | 2 |
| 2003 | Confidence measures for keyword spotting using support vector machinesabstractSupport vector machines (SVM) is a new and very promising classification technique developed from the theory of structural risk minimisation. We propose an alternative out-of-vocabulary word detection method relying on confidence measures and support vector machines. Confidence measures are computed from phone level information provided by a hidden Markov model (HMM) based speech recognizer. We use three kinds of average techniques as arithmetic, geometric and harmonic averages to compute a confidence measure for each word. The acceptance/rejection decision of a word is based on the confidence feature vector which is processed by a SVM classifier. The performance of the proposed SVM classifier is compared with methods based on the averaging of confidence measures. Yassine Ben Ayed, Dominique Fohr, Jean Paul Haton, Gérard Chollet |
ICASSP (1) | 1 |