Alessandro L. Koerich

dblp:37/4335 · also Alessandro Lameiras Koerich · DBLP profile ↗
← Back
88ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0001-5879-7014ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 5 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 13 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021
YearPublicationVenuePosition
2026 MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
abstract
Personalized expression recognition (ER) involves adapting a machine learning model to subject-specific data for improved recognition of expressions with considerable inter-personal variability. Subject-specific ER can benefit significantly from multi-source domain adaptation (MSDA) methods – where each domain corresponds to a specific subject – to improve model accuracy and robustness. Despite promising results, state-of-the-art MSDA approaches often overlook multimodal information or blend sources into a single domain, limiting subject diversity and failing to explicitly capture unique subject-specific characteristics. To address these limitations, we introduce MuSACo, a multi-modal subject-specific selection and adaptation method for ER based on co-training. It leverages complementary information across multiple modalities and multiple source domains for subject-specific adaptation. This makes MuSACo particularly relevant for affective computing applications in digital health, such as patient-specific assessment for stress or pain, where subject-level nuances are crucial. MuSACo selects source subjects relevant to the target and generates pseudo-labels using the dominant modality for class-aware learning, in conjunction with a class-agnostic loss to learn from less confident target samples. Finally, source features from each modality are aligned, while only confident target features are combined. Experimental results on challenging multimodal ER datasets – BioVid, StressID, and BAH – show that MuSACo outperforms UDA (blending) and state-of-the-art MSDA methods. Our code is available: https://github.com/osamazeeshan/MuSACo
Muhammad Osama Zeeshan, Natacha Gillet, Alessandro L. Koerich, Marco Pedersoli, François Brémond, Eric Granger
WACV3
2026 Progressive Multi-Source Domain Adaptation for Personalized Facial Expression Recognition
abstract
Personalized facial expression recognition (FER) involves adapting a machine learning model using samples from labeled sources and unlabeled target domains. Given the challenges of recognizing subtle expressions with considerable interpersonal variability, state-of-the-art unsupervised domain adaptation (UDA) methods focus on the multi-source UDA (MSDA) setting, where each domain corresponds to a specific subject, and improve model accuracy and robustness. However, when adapting to a specific target, the diverse nature of multiple source domains translates to a large shift between source and target data. State-of-the-art MSDA methods for FER address this domain shift by considering all the sources to adapt to the target representations. Nevertheless, adapting to a target subject presents significant challenges due to large distributional differences between source and target domains, often resulting in negative transfer. In addition, integrating all sources simultaneously increases computational costs and causes misalignment with the target. To address these issues, we propose a progressive MSDA approach that gradually introduces information from subjects (source domains) based on their similarity to the target subject. This will ensure that only the most relevant sources from the target are selected, which helps avoid the negative transfer caused by dissimilar sources. During adaptation, the source domains are introduced in a curriculum manner. We first exploit the closest sources to reduce the distribution shift with the target and then move towards the furthest while only considering the most relevant sources based on the predetermined threshold. Furthermore, to mitigate catastrophic forgetting caused by the incremental introduction of source subjects, we implemented a density-based memory mechanism that preserves the most relevant historical source samples for adaptation. Our extensive experiments 1 show the effectiveness of our proposed method on challenging FER datasets: Biovid, UNBC-McMaster, Aff-Wild2, and BAH. Further, performance is evaluated on a cross-dataset setting (UNBC-McMaster → BioVid), showing the importance of gradually adapting to source subjects.
Muhammad Osama Zeeshan, Marco Pedersoli, Alessandro L. Koerich, Eric Granger
IEEE Trans. Affect. Comput.3
2025 Disentangled Source-Free Personalization for Facial Expression Recognition with Neutral Target Data
abstract
Facial Expression Recognition (FER) from videos is a crucial task in various application areas, such as human-computer interaction and health diagnosis and monitoring (e.g., assessing pain and depression). Beyond the challenges of recognizing subtle emotional or health states, the effectiveness of deep FER models is often hindered by the considerable inter-subject variability in expressions. Source-free (unsupervised) domain adaptation (SFDA) methods may be employed to adapt a pre-trained source model using only unlabeled target domain data, thereby avoiding data privacy, storage, and transmission issues. Typically, SFDA methods adapt to a target domain dataset corresponding to an entire population and assume it includes data from all recognition classes. However, collecting such comprehensive target data can be difficult or even impossible for FER in healthcare applications. In many real-world scenarios, it may be feasible to collect a short neutral control video (which displays only neutral expressions) from target subjects before deployment. These videos can be used to adapt a model to better handle the variability of expressions among subjects. This paper introduces the Disentangled SFDA (DSFDA) method to address the challenge posed by adapting models with missing target expression data. DSFDA leverages data from a neutral target control video for end-to-end generation and adaptation of target data with missing non-neutral data. Our method learns to disentangle features related to expressions and identity while generating the missing non-neutral expression data for the target subject, thereby enhancing model accuracy. Additionally, our self-supervision strategy improves model adaptation by reconstructing target images that maintain the same identity and source expression. Experimental results1on the challenging BioVid, UNBC-McMaster and StressID datasets indicate that our DSFDA approach can outperform state-of-the-art adaptation methods.1https://github.com/MasoumehSharafi/DSFDA/
Masoumeh Sharafi, Emma Ollivier, Muhammad Osama Zeeshan, Soufiane Belharbi, Alessandro L. Koerich, Marco Pedersoli, Simon Bacon, Eric Granger
FG5
2025 Learning from Stochastic Teacher Representations Using Student-Guided Knowledge Distillation
Muhammad Haseeb Aslam, Clara Martinez, Marco Pedersoli, Alessandro L. Koerich, Ali Etemad, Eric Granger
ECML/PKDD (6)4
2025 Representation ensemble learning applied to facial expression recognition
Bruna Rossetto Delazeri, Andre G. Hochuli, Jean Paul Barddal, Alessandro L. Koerich, Alceu S. Britto Jr.
Neural Comput. Appl.4
2025 Concept Drift Adaptation in Text Stream Mining Settings: A Systematic Review
abstract
The society produces textual data online in several ways, e.g., via reviews and social media posts. Therefore, numerous researchers have been working on discovering patterns in textual data that can indicate peoples’ opinions, interests, and so on. Most tasks regarding natural language processing are addressed using traditional machine learning methods and static datasets. This setting can lead to several problems, e.g., outdated datasets and models, which degrade in performance over time. This is particularly true regarding concept drift, in which the data distribution changes over time. Furthermore, text streaming scenarios also exhibit further challenges, such as the high speed at which data arrive over time. Models for stream scenarios must adhere to the aforementioned constraints while learning from the stream, thus storing texts for limited periods and consuming low memory. This study presents a systematic literature review regarding concept drift adaptation in text stream scenarios. Considering well-defined criteria, we selected 48 papers published between 2018 and August 2024 to unravel aspects such as text drift categories, detection types, model update mechanisms, stream mining tasks addressed, and text representation methods and their update mechanisms. Furthermore, we discussed drift visualization and simulation and listed real-world datasets used in the selected papers. Finally, we brought forward a discussion on existing works in the area, also highlighting open challenges and future research directions for the community.
Cristiano Mesquita Garcia, Ramon Abílio, Alessandro L. Koerich, Alceu S. Britto Jr., Jean Paul Barddal
ACM Trans. Intell. Syst. Technol.3
2024 Is it Fine to Tune? Evaluating SentenceBERT Fine-tuning for Brazilian Portuguese Text Stream Classification
abstract
Pre-trained language models (LMs) have been used in several scenarios and data mining tasks due to their good-quality representations and their use readiness. Although LMs constitute a significant gain in usability, they are frequently utilized statically over time, meaning that these models can suffer from concept drift and semantic shift, which correspond to changes in data distribution and word meanings. These phenomena are more noticeable when new texts become gradually available. This paper evaluates the impact of updating pre-trained SentenceBERT models overtime on a Brazilian news post classification task in text streaming fashion, a paradigm suitable for learning from data streams. While we update the SBERT model yearly with a reduced number of recent posts, we compare it with scenarios using static LMs. We used the adaptive random forest for classification and evaluated it regarding macro F1-score and elapsed time. The experimental results show that regularly leveraging sampled texts from the recent past for fine-tuning LMs can improve performance metrics over time, reaching better results than using static LMs in most years analyzed. We also evaluated the run times, which suggests that fine-tuning LMs over time provides a good trade-off between performance and run time.
Bruno Yuiti Leão Imai, Cristiano Mesquita Garcia, Marcio Vinicius Rocha, Alessandro L. Koerich, Alceu S. Britto Jr., Jean Paul Barddal
IEEE Big Data4
2024 Distilling Privileged Multimodal Information for Expression Recognition using Optimal Transport
abstract
Deep learning models for multimodal expression recognition have reached remarkable performance in controlled laboratory environments because of their ability to learn complementary and redundant semantic information. However, these models struggle in the wild, mainly because of the unavailability and quality of modalities used for training. In practice, only a subset of the training-time modalities may be available at test time. Learning with privileged information enables models to exploit data from additional modalities that are only available during training. State-of-the-art knowledge distillation (KD) methods have been proposed to distill information from multiple teacher models (each trained on a modality) to a common student model. These privileged KD methods typically utilize point-to-point matching, yet have no explicit mechanism to capture the structural information in the teacher representation space formed by introducing the privileged modality. We argue that encoding this same structure in the student space may lead to enhanced student performance. This paper introduces a new structural KD mechanism based on optimal transport (OT), where entropy-regularized OT distills the structural dark knowledge. Our privileged KD with OT (PKDOT) method captures the local structures in the multimodal teacher representation by calculating a cosine similarity matrix and selecting the top-k anchors to allow for sparse OT solutions, resulting in a more stable distillation process. Experiments1were performed on two challenging problems - pain estimation on the Biovid dataset (ordinal classification) and arousal-valance prediction on the Affwild2 dataset (regression). Results show that our proposed method can outperform state-of-the-art privileged KD methods on these problems. The diversity among modalities and fusion architectures indicates that PKDOT is modality-and model-agnostic.
Muhammad Haseeb Aslam, Muhammad Osama Zeeshan, Soufiane Belharbi, Marco Pedersoli, Alessandro L. Koerich, Simon Bacon, Eric Granger
FG5
2024 Guided Interpretable Facial Expression Recognition via Spatial Action Unit Cues
abstract
Although state-of-the-art classifiers for facial expression recognition (FER) can achieve a high level of accuracy, they lack interpretability, an important feature for end-users. Experts typically associate spatial action units (AUs) from a codebook to facial regions for the visual interpretation of expressions. In this paper, the same expert steps are followed. A new learning strategy is proposed to explicitly incorporate AU cues into classifier training, allowing to train deep interpretable models. During training, this AU codebook is used, along with the input image expression label, and facial landmarks, to construct a AU heatmap that indicates the most discriminative image regions of interest w.r.t the facial expression. This valuable spatial cue is leveraged to train a deep interpretable classifier for FER. This is achieved by constraining the spatial layer features of a classifier to be correlated with AU heatmaps. Using a composite loss, the classifier is trained to correctly classify an image while yielding interpretable visual layer-wise attention correlated with AU maps, simulating the expert decision process. Our strategy only relies on image class expression for supervision, without additional manual annotations. Our new strategy is generic, and can be applied to any deep CNN - or transformer-based classifier without requiring any architectural change or significant additional training time. Our extensive evaluation11Our code is available at:https://github.com/sbelharbi/interpretable-fer-aus. on two public benchmarks RAF-DB, and AffectNet datasets shows that our proposed strategy can improve layer-wise interpretability without degrading classification performance. In addition, we explore a common type of interpretable classifiers that rely on class activation mapping (CAM) methods, and show that our approach can also improve CAM interpretability.
Soufiane Belharbi, Marco Pedersoli, Alessandro L. Koerich, Simon Bacon, Eric Granger
FG3
2024 Subject-Based Domain Adaptation for Facial Expression Recognition
abstract
Adapting a deep learning model to a specific target individual is a challenging facial expression recognition (FER) task that may be achieved using unsupervised domain adaptation (UDA) methods. Although several UDA methods have been proposed to adapt deep FER models across source and target data sets, multiple subject-specific source domains are needed to accurately represent the intra-and inter-person variability in subject-based adaption. This paper considers the setting where domains correspond to individuals, not entire datasets. Unlike UDA, multi-source domain adaptation (MSDA) methods can leverage multiple source datasets to improve the accuracy and robustness of the target model. However, previous methods for MSDA adapt image classification models across datasets and do not scale well to a more significant number of source domains. This paper introduces a new MSDA method for subject-based domain adaptation in FER. It efficiently leverages information from multiple source subjects (labeled source domain data) to adapt a deep FER model to a single target individual (unlabeled target domain data). During adaptation, our subject-based MSDA first computes a between-source discrepancy loss to mitigate the domain shift among data from several source subjects. Then, a new strategy is employed to generate augmented confident pseudo-labels for the target subject, allowing a reduction in the domain shift between source and target subjects. Experiments1performed on the challenging BioVid heat and pain dataset with 87 subjects and the UNBC-McMaster shoulder pain dataset with 25 subjects show that our subject-based MSDA can outperform state-of-the-art methods yet scale well to multiple subject-based source domains.
Muhammad Osama Zeeshan, Muhammad Haseeb Aslam, Soufiane Belharbi, Alessandro L. Koerich, Marco Pedersoli, Simon Bacon, Eric Granger
FG4
2024 Improving Sampling Methods for Fine-Tuning SentenceBERT in Text Streams
Cristiano Mesquita Garcia, Alessandro L. Koerich, Alceu S. Britto Jr., Jean Paul Barddal
ICPR (19)2
2024 Alleviating Catastrophic Forgetting in Facial Expression Recognition with Emotion-Centered Models
Israel A. Laurensi R., Alceu S. Britto Jr., Jean Paul Barddal, Alessandro L. Koerich
ICPR (9)4
2023 Multiresolution texture analysis of histopathologic images using ecological diversity measures
Steve Ataky, Alessandro L. Koerich
Expert Syst. Appl.2
2023 Large-margin representation learning for texture classification
Jonathan de Matos, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Alessandro L. Koerich
Pattern Recognit. Lett.4
2022 Towards Robust Speech-to-Text Adversarial Attack
abstract
This paper introduces a novel adversarial algorithm for attacking the advanced speech-to-text transcription systems. Our proposed approach is based on developing an extension for the conventional distortion condition of the general adversarial optimization formulation using the Cramér integral probability metric. Minimizing over such a metric contributes to crafting signals very close to the subspace of legitimate speech recordings. That helps yield more robust adversarial signals against over-the-air playbacks without employing neither costly expectation over transformations nor static room impulse response simulations. Our approach considerably outperforms other targeted and non-targeted algorithms in terms of word error rate and sentence-level accuracy. Furthermore compared to seven other strong white and black-box adversarial attacks, our proposed approach is considerably more resilient against multiple consecutive over-the-air playbacks, corroborating its higher robustness in noisy environments.
Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich
ICASSP3
2022 Named Entity Recognition for Audio De-Identification
abstract
Data anonymization is often a task carried out by humans. Automating it would reduce the cost and time required to complete this task. This paper presents a pipeline to automate the anonymization of audio data in French. We propose a pipeline, which takes audio files with their transcriptions and removes the named entities (NEs) present in the audio. Our pipeline is made up of a forced aligner, which aligns words in an audio transcript with speech and a model that performs named entity recognition (NER). Then, the audio segments that correspond to NEs are substituted with silence to anonymize audio. We compared forced aligners and NER models to find the best ones for our scenario. We evaluated our pipeline on a small hand-annotated dataset, achieving an F1 score of 0.769. This result shows that automating this task is feasible.
Guillaume Baril, Patrick Cardinal, Alessandro L. Koerich
IJCNN3
2022 Evaluation of Self-taught Learning-based Representations for Facial Emotion Recognition
abstract
This work describes different strategies to generate unsupervised representations obtained through the concept of self-taught learning for facial emotion recognition (FER). The idea is to create complementary representations promoting diver-sity by varying the autoencoders' initialization, architecture, and training data. SVM, Bagging, Random Forest, and a dynamic ensemble selection method are evaluated as final classification methods. Experimental results on JAFFE and Cohn-Kanade datasets using a leave-one-subject-out protocol show that FER methods based on the proposed diverse representations compare favorably against state-of-the-art approaches that also explore unsupervised feature learning.
Bruna Rossetto Delazeri, Leonardo León Vera, Jean Paul Barddal, Alessandro L. Koerich, Alceu S. Britto Jr.
IJCNN4
2022 Pattern Spotting and Image Retrieval in Historical Documents using Deep Hashing
abstract
This paper presents a deep learning approach for image retrieval and pattern spotting in digital collections of historical documents. First, a region proposal algorithm detects object candidates in the document page images. Next, deep learning models are used for feature extraction, considering two distinct variants, which provide either real-valued or binary code representations. Finally, candidate images are ranked by computing the feature similarity with a given input query. A robust experimental protocol evaluates the proposed approach considering each representation scheme (real-valued and binary code) on the DocExplore image database. The experimental results show that the proposed deep models compare favorably to the state-of-the-art image retrieval approaches for images of historical documents, outperforming other deep models by 2.56 percentage points using the same techniques for pattern spotting. Besides, the proposed approach also reduces the search time up to 200$\times$, and the storage cost up to 6,000$\times$ when compared to related works based on real-valued representations.
Caio da S. Dias, Alceu S. Britto Jr., Jean Paul Barddal, Laurent Heutte, Alessandro L. Koerich
SMC5
2022 Two-view fine-grained classification of plant species
Voncarlos Araujo, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich
Neurocomputing4
2022 A novel bio-inspired texture descriptor based on biodiversity and taxonomic measures
Steve Ataky, Alessandro L. Koerich
Pattern Recognit.2
2022 A human-in-the-loop recommendation-based framework for reconstruction of mechanically shredded documents
Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Maria Cláudia Silva Boeres, Alessandro L. Koerich, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
Pattern Recognit. Lett.4
2022 Multidiscriminator Sobolev Defense-GAN Against Adversarial Attacks for End-to-End Speech Systems
abstract
This paper introduces a defense approach against end-to-end adversarial attacks developed for cutting-edge speech-to-text systems. The proposed defense algorithm has four steps. First, we use the short-time Fourier transform to represent speech signals with 2D spectrograms. Second, we iteratively find a safe vector using a spectrogram subspace projection operation. This operation minimizes the chordal distance adjustment between spectrograms with an additional regularization term. Third, we synthesize a spectrogram with such a safe vector using a novel GAN architecture trained with Sobolev integral probability metric. We impose an additional constraint on the generator network to improve the model’s performance in terms of stability and the total number of learned modes. Finally, we reconstruct the signal from the synthesized spectrogram and the Griffin-Lim phase approximation technique. We evaluate the proposed defense approach against six strong white and black-box adversarial attacks on DeepSpeech, Kaldi, and Lingvo models. The experimental results show that our algorithm outperforms other state-of-the-art defense algorithms in terms of accuracy and signal quality.
Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich
IEEE Trans. Inf. Forensics Secur.3
2021 Class-Conditional Defense GAN Against End-To-End Speech Attacks
abstract
In this paper we propose a novel defense approach against end-to-end adversarial attacks developed to fool advanced speech-to-text systems such as DeepSpeech and Lingvo. Unlike conventional defense approaches, the proposed approach does not directly employ low-level transformations such as autoencoding a given input signal aiming at removing potential adversarial perturbation. Instead of that, we find an optimal input vector for a class conditional generative adversarial network through minimizing the relative chordal distance adjustment between a given test input and the generator network. Then, we reconstruct the 1D signal from the synthesized spectrogram and the original phase information derived from the given input signal. Hence, this reconstruction does not add any extra noise to the signal and according to our experimental results, our defense-GAN considerably outperforms conventional defense algorithms both in terms of word error rate and sentence level recognition accuracy.
Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich
ICASSP3
2021 Tensor analysis with n-mode generalized difference subspace
Bernardo Bentes Gatto, Eulanda M. dos Santos, Alessandro L. Koerich, Kazuhiro Fukui, Waldir S. S. Júnior
Expert Syst. Appl.3
2021 Cyclic Defense GAN Against Speech Adversarial Attacks
abstract
This paper proposes a new defense approach for counteracting state-of-the-art white and black-box adversarial attack algorithms. Our approach fits into the implicit reactive defense algorithm category since it does not directly manipulate the potentially malicious input signals. Instead, it reconstructs a similar signal with a synthesized spectrogram using a cyclic generative adversarial network. This cyclic framework helps to yield a stable generative model. Finally, we feed the reconstructed signal into the speech-to-text model for transcription. The conducted experiments on targeted and non-targeted adversarial attacks developed for attacking DeepSpeech, Kaldi, and Lingvo models demonstrate the proposed defense's effectiveness in adverse scenarios.
Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich
IEEE Signal Process. Lett.3
2020 Fast(er) Reconstruction of Shredded Text Documents via Self-Supervised Deep Asymmetric Metric Learning
abstract
The reconstruction of shredded documents consists in arranging the pieces of paper (shreds) in order to reassemble the original aspect of such documents. This task is particularly relevant for supporting forensic investigation as documents may contain criminal evidence. As an alternative to the laborious and time-consuming manual process, several researchers have been investigating ways to perform automatic digital reconstruction. A central problem in automatic reconstruction of shredded documents is the pairwise compatibility evaluation of the shreds, notably for binary text documents. In this context, deep learning has enabled great progress for accurate reconstructions in the domain of mechanically-shredded documents. A sensitive issue, however, is that current deep model solutions require an inference whenever a pair of shreds has to be evaluated. This work proposes a scalable deep learning approach for measuring pairwise compatibility in which the number of inferences scales linearly (rather than quadratically) with the number of shreds. Instead of predicting compatibility directly, deep models are leveraged to asymmetrically project the raw shred content onto a common metric space in which distance is proportional to the compatibility. Experimental results show that our method has accuracy comparable to the state-of-the-art with a speed-up of about 22 times for a test instance with 505 shreds (20 mixed shredded-pages from different documents).
Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Maria Cláudia Silva Boeres, Alessandro L. Koerich, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
CVPR4
2020 Detection of Adversarial Attacks and Characterization of Adversarial Subspace
abstract
Adversarial attacks have always been a serious threat for any data-driven model. In this paper, we explore subspaces of adversarial examples in unitary vector domain, and we propose a novel detector for defending our models trained for environmental sound classification. We measure chordal distance between legitimate and malicious representation of sounds in unitary space of generalized Schur decomposition and show that their manifolds lie far from each other. Our front-end detector is a regularized logistic regression which discriminates eigenvalues of legitimate and adversarial spectrograms. The experimental results on three benchmarking datasets of environmental sounds represented by spectrograms reveal high detection rate of the proposed detector for eight types of adversarial attacks and it also outperforms other detection approaches.
Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich
ICASSP3
2020 Continuous Emotion Recognition via Deep Convolutional Autoencoder and Support Vector Regressor
abstract
Automatic facial expression recognition (FER) is an important research area in the emotion recognition and computer vision. Applications can be found in several domains such as medical treatment, driver fatigue surveillance, sociable robotics, and several other human-computer interaction systems. Therefore, it is crucial that the machine should be able to recognize the emotional state of the user with high accuracy. In recent years, deep neural networks have been used with great success in recognizing emotions. In this paper, we present a new model for continuous emotion recognition based on FER by using an unsupervised learning approach based on transfer learning and autoencoders. The proposed approach also includes preprocessing and post-processing techniques which contribute favorably to improving the performance of predicting the concordance correlation coefficient for arousal and valence dimensions. Experimental results for predicting spontaneous and natural emotions on the RECOLA 2016 dataset have shown that the proposed approach based on visual information can achieve concordance correlation coefficient of 0.516 and 0.264 for valence and arousal, respectively.
Sevegni Odilon Clement Allognon, Alceu S. Britto Jr., Alessandro L. Koerich
IJCNN3
2020 Data Augmentation for Histopathological Images Based on Gaussian-Laplacian Pyramid Blending
abstract
Data imbalance is a major problem that affects several machine learning (ML) algorithms. Such a problem is troublesome because most of the ML algorithms attempt to optimize a loss function that does not take into account the data imbalance. Accordingly, the ML algorithm simply generates a trivial model that is biased toward predicting the most frequent class in the training data. In the case of histopathologic images (HIs), both low-level and high-level data augmentation (DA) techniques still present performance issues when applied in the presence of inter-patient variability; whence the model tends to learn color representations, which is related to the staining process. In this paper, we propose a novel approach capable of not only augmenting HI dataset but also distributing the inter-patient variability by means of image blending using the Gaussian-Laplacian pyramid. The proposed approach consists of finding the Gaussian pyramids of two images of different patients and finding the Laplacian pyramids thereof. Afterwards, the left-half side and the right-half side of different HIs are joined in each level of the Laplacian pyramid, and from the joint pyramids, the original image is reconstructed. This composition combines the stain variation of two patients, avoiding that color differences mislead the learning process. Experimental results on the BreakHis dataset have shown promising gains vis-à-vis the majority of DA techniques presented in the literature.
Steve Ataky, Jonathan de Matos, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich
IJCNN5
2020 Cross-Representation Transferability of Adversarial Attacks: From Spectrograms to Audio Waveforms
abstract
This paper shows the susceptibility of spectrogram-based audio classifiers to adversarial attacks and the transferability of such attacks to audio waveforms. Some commonly used adversarial attacks to images have been applied to Mel-frequency and short-time Fourier transform spectrograms, and such perturbed spectrograms are able to fool a 2D convolutional neural network (CNN). Such attacks produce perturbed spectrograms that are visually imperceptible by humans. Furthermore, the audio waveforms reconstructed from the perturbed spectrograms are also able to fool a 1D CNN trained on the original audio. Experimental results on a dataset of western music have shown that the 2D CNN achieves up to 81.87% of mean accuracy on legitimate examples and such performance drops to 12.09% on adversarial examples. Likewise, the 1D CNN achieves up to 78.29% of mean accuracy on original audio samples and such performance drops to 27.91% on adversarial audio waveforms reconstructed from the perturbed spectrograms.
Karl M. Koerich, Mohammad Esmaeilpour, Sajjad Abdoli, Alceu S. Britto Jr., Alessandro L. Koerich
IJCNN5
2020 Incremental and decremental fuzzy bounded twin support vector machine
Alexandre Reeberg de Mello, Marcelo Ricardo Stemmer, Alessandro L. Koerich
Inf. Sci.3
2020 Self-supervised deep reconstruction of mixed strip-shredded text documents
Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Maria Cláudia Silva Boeres, Alessandro L. Koerich, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
Pattern Recognit.4
2020 A Robust Approach for Securing Audio Classification Against Adversarial Attacks
abstract
Adversarial audio attacks can be considered as a small perturbation unperceptive to human ears that is intentionally added to an audio signal and causes a machine learning model to make mistakes. This poses a security concern about the safety of machine learning models since the adversarial attacks can fool such models toward the wrong predictions. In this paper we first review some strong adversarial attacks that may affect both audio signals and their 2D representations and evaluate the resiliency of deep learning models and support vector machines (SVM) trained on 2D audio representations such as short time Fourier transform, discrete wavelet transform (DWT) and cross recurrent plot against several state-of-the-art adversarial attacks. Next, we propose a novel approach based on pre-processed DWT representation of audio signals and SVM to secure audio systems against adversarial attacks. The proposed architecture has several preprocessing modules for generating and enhancing spectrograms including dimension reduction and smoothing. We extract features from small patches of the spectrograms using the speeded up robust feature (SURF) algorithm which are further used to transform into cluster distance distribution using the K-Means++ algorithm. Finally, SURF-generated vectors are encoded by this codebook and the resulting codewords are used for training a SVM. All these steps yield to a novel approach for audio classification that provides a good tradeoff between accuracy and resilience. Experimental results on three environmental sound datasets show the competitive performance of the proposed approach compared to the deep neural networks both in terms of accuracy and robustness against strong adversarial attacks.
Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich
IEEE Trans. Inf. Forensics Secur.3
2019 Texture CNN for Histopathological Image Classification
abstract
Biopsies are the gold standard for breast cancer diagnosis. This task can be improved by the use of Computer Aided Diagnosis (CAD) systems, reducing the time of diagnosis and reducing the inter and intra-observer variability. The advances in computing have brought this type of system closer to reality. However, datasets of Histopathological Images (HI) from biopsies are quite small and unbalanced what makes difficult to use modern machine learning techniques such as deep learning. In this paper we propose a compact architecture based on texture filters that has fewer parameters than traditional deep models but is able to capture the difference between malignant and benign tissues with relative accuracy. The experimental results on the BreakHis dataset have show that the proposed texture CNN achieves almost 90% of accuracy for classifying benign and malignant tissues.
Jonathan de Matos, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich
CBMS4
2019 Texture CNN for Thermoelectric Metal Pipe Image Classification
abstract
In this paper, the concept of representation learning based on deep neural networks is applied as an alternative to the use of handcrafted features in a method for automatic visual inspection of corroded thermoelectric metallic pipes. A texture convolutional neural network (TCNN) replaces hand-crafted features based on Local Phase Quantization (LPQ) and Haralick descriptors (HD) with the advantage of learning an appropriate textural representation and the decision boundaries into a single optimization process. Experimental results have shown that it is possible to reach the accuracy of 99.20% in the task of identifying different levels of corrosion in the internal surface of thermoelectric pipe walls, while using a compact network that requires much less effort in tuning parameters when compared to the handcrafted approach since the TCNN architecture is compact regarding the number of layers and connections. The observed results open up the possibility of using deep neural networks in real-time applications such as the automatic inspection of thermoelectric metal pipes.
Daniel Vriesman, Alceu S. Britto Jr., Alessandro Zimmer, Alessandro L. Koerich
ICTAI4
2019 Double Transfer Learning for Breast Cancer Histopathologic Image Classification
abstract
This work proposes a classification approach for breast cancer histopathologic images (HI) that uses transfer learning to extract features from HI using an Inception-v3 CNN pre-trained with ImageNet dataset. We also use transfer learning on training a support vector machine (SVM) classifier on a tissue labeled colorectal cancer dataset aiming to filter the patches from a breast cancer HI and remove the irrelevant ones. We show that removing irrelevant patches before training a second SVM classifier, improves the accuracy for classifying malign and benign tumors on breast cancer images. We are able to improve the classification accuracy in 3.7% using the feature extraction transfer learning and an additional 0.7% using the irrelevant patch elimination. The proposed approach outperforms the state-of-the-art in three out of the four magnification factors of the breast cancer dataset.
Jonathan de Matos, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich
IJCNN4
2019 Image Retrieval and Pattern Spotting using Siamese Neural Network
abstract
This paper presents a novel approach for image retrieval and pattern spotting in document image collections. The manual feature engineering is avoided by learning a similarity-based representation using a Siamese Neural Network trained on a previously prepared subset of image pairs from the ImageNet dataset. The learned representation is used to provide the similarity-based feature maps used to find relevant image candidates in the data collection given an image query. A robust experimental protocol based on the public Tobacco800 document image collection shows that the proposed method compares favor-ably against state-of-the-art document image retrieval methods, reaching 0.94 and 0.83 of mean average precision (mAP) for retrieval and pattern spotting (IoU=0.7), respectively. Besides, we have evaluated the proposed method considering feature maps of different sizes, showing the impact of reducing the number of features in the retrieval performance and time-consuming.
Kelly Lais Wiggers, Alceu S. Britto Jr., Laurent Heutte, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira
IJCNN4
2019 Emotion Recognition Using Fusion of Audio and Video Features
abstract
In this paper we propose a fusion approach to continuous emotion recognition that combines visual and auditory modalities in their representation spaces to predict the arousal and valence levels. The proposed approach employs a pre-trained convolution neural network and transfer learning to extract features from video frames that capture the emotional content. For the auditory content, a minimalistic set of parameters such as prosodic, excitation, vocal tract, and spectral descriptors are used as features. The fusion of these two modalities is carried out at a feature level, before training a single support vector regressor (SVR) or at a prediction level, after training one SVR for each modality. The proposed approach also includes preprocessing and post-processing techniques which contribute favorably to improving the concordance correlation coefficient (CCC). Experimental results for predicting spontaneous and natural emotions on the RECOLA dataset have shown that the proposed approach takes advantage of the complementary information of visual and auditory modalities and provides CCC of 0.749 and 0.565 for arousal and valence, respectively. The proposed approach outperforms the baseline system and several traditional approaches based on auditory and visual handcrafted features.
Juan D. S. Ortega, Patrick Cardinal, Alessandro L. Koerich
SMC3
2019 Memory Integrity of CNNs for Cross-Dataset Facial Expression Recognition
abstract
Facial expression recognition is a major problem in the domain of artificial intelligence. One of the best ways to solve this problem is the use of convolutional neural networks (CNNs). However, a large amount of data is required to train properly these networks but most of the datasets available for facial expression recognition are relatively small. A common way to circumvent the lack of data is to use CNNs trained on large datasets of different domains and fine-tuning the layers of such networks to the target domain. However, the fine-tuning process does not preserve the memory integrity as CNNs have the tendency to forget patterns they have learned. In this paper, we evaluate different strategies of fine-tuning a CNN with the aim of assessing the memory integrity of such strategies in a cross-dataset scenario. A CNN pre-trained on a source dataset is used as the baseline and four adaptation strategies have been evaluated: fine-tuning its fully connected layers; fine-tuning its last convolutional layer and its fully connected layers; retraining the CNN on a target dataset; and the fusion of the source and target datasets and retraining the CNN. Experimental results on four datasets have shown that the fusion of the source and the target datasets provides the best trade-off between accuracy and memory integrity.
Dylan C. Tannugi, Alceu S. Britto Jr., Alessandro L. Koerich
SMC3
2019 End-to-end environmental sound classification using a 1D convolutional neural network
Sajjad Abdoli, Patrick Cardinal, Alessandro L. Koerich
Expert Syst. Appl.3
2018 Fine-Grained Hierarchical Classification of Plant Leaf Images Using Fusion of Deep Models
abstract
A fine-grained plant leaf classification method based on the fusion of deep models is described. Complementary global and patch-based leaf features are combined at each hierarchical level (genus and species) by pre-trained CNNs. The deep models are adapted for plant recognition by using data augmentation techniques to face the problem of plant classes with very few samples for training in the available imbalanced dataset. Experimental results have shown that the proposed coarse-to-fine classification strategy is a very promising alternative to deal with the low inter-class and high intra-class variability inherent to the problem of plant identification. The proposed method was able to surpass other state-of-the-art approaches on the ImageCLEF 2015 plant recognition dataset in terms of average classification scores.
Voncarlos Araujo, Alceu S. Britto Jr., Andre L. Brun, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira
ICTAI4
2018 Document Image Retrieval Using Deep Features
abstract
This paper proposes a novel approach for content based graphical object retrieval in document images. The challenge is to search for occurrences of a queried graphical objects in document images that can vary in terms of color, shape, texture and quality, increasing considerably the level of difficulty of the retrieval process. To that end, the manual feature engineering is avoided by learning the image representation for the retrieval task using a Convolutional Neural Network (CNN). However, such a representation should be as compact as possible to allow a fast document image retrieval and storage. Thus, a pretrained CNN model is used to cope with the lack of training data, which is fine tuned to achieve a compact yet discriminant representation of the graphical objects. From experiments conducted on the public Tobacco800 document image collection, we show that the proposed method compares favorably against state-of-the-art document image retrieval methods, reaching 0.72 of average precision (mAP). In addition, an increase of 4 percentage points in the average precision is observed using a compact deep representation in which the number of features is reduced by 16 times, thus allowing a reduction of 47% in terms of computation time by the image retrieval task.
Kelly Lais Wiggers, Alceu S. Britto Jr., Laurent Heutte, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira
IJCNN4
2017 Two-stage facial age prediction using group-specific features
abstract
A novel two-stage age prediction approach with group-specific features is proposed in this paper. Aging process is captured through a highly discriminating feature representation that models shape, appearance, skin spots, and wrinkles. The two-stage method consists of a multi-class Support Vector Machine (SVM) to predict the age bracket while the final age prediction is carried out using Support Vector Regression (SVR). The novelty of our work is that the feature extraction is group-specific and can therefore be tailored to each age bracket in the specific age prediction step. The FG-NET Aging dataset was used to evaluate the proposed method and an impressive mean absolute error (MAE) of 3.98 was achieved. Our approach outperforms the current state-of-the-art while increasing the robustness to blur, expression and lighting variation with local phase features.
Jhony K. Pontes, Clinton Fookes, Alceu S. Britto Jr., Alessandro L. Koerich
ICASSP4
2017 A two-step cascade classification method
abstract
This paper proposes a classification approach in which monolithic and multiple classifier systems are combined in a cascading fashion. The rationale behind that is to deal with the existing trade-off between the need for increasing the accuracy, while reducing the complexity of the classification method. In other words, the idea is to offer an interesting strategy to conciliate the different levels of efforts necessary to deal with easy and hard patterns usually observed in a classification problem. The experimental results have shown that for some problems more than 90% of the instances can be processed in the first step of the cascade, saving efforts by avoiding the use of the second step in which a more complex classification method is used. It means that for some problems the reduction of the classification cost achieved more than 70% when compared to the use of an MCS. In addition to this interesting classification cost reduction, the cascade approach has shown to be able of improving the classification accuracy up to 15.19 percentage points.
Eunelson Jose da Silva Junior, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Fabrício Enembreck, Robert Sabourin, Alessandro L. Koerich
IJCNN6
2017 Multiple classifier system for plant leaf recognition
abstract
This paper presents a multiple classifier system (MCS) to identify plants species based on the texture and shape features extracted from leaf images. A diverse pool of SVM and Neural Network classifiers is trained on four different feature sets, namely, Local Binary Pattern (LBP), Histogram of Gradients (HOG), Speed of Robust Features (SURF) and Zernike Moments (ZM). Then, a static classifier selection method is used to search for the ensembles that maximize the average classification score. Experimental results on ImageCLEF 2011 and 2012 datasets have shown that combining different kind of classifiers trained on shape and texture features is an effective strategy for the plant automatic identification. The MCS improves the identification performance in up to 28% relative to the monolithic approach. Furthermore, the proposed approach also compares favourably with the best results reported in the literature for those datasets.
Voncarlos Araujo, Alceu S. Britto Jr., Andre L. Brun, Alessandro L. Koerich, Rosane Palate
SMC4
2016 Facial expression recognition using a pairwise feature selection and classification approach
abstract
This paper proposes a novel approach that combines specialized pairwise classifiers trained with different feature subsets for facial expression classification. The proposed approach first detects and extracts automatically faces from images. Next, the face is split into several regular zones and textural features are extracted from each zone to capture local information. The features extracted from all zones are concatenated to model the whole face. A pairwise approach that considers all pairs of classes and a hybrid feature selection strategy is used to both reduce the dimensionality and to select relevant features to discriminate between specific pairs of classes. Several pairwise classifiers are then trained with such pairwise feature subsets. At the end, given a new face image, all features are extracted from such a face, but only the previously selected subset of features is inputted to each pairwise classifier. The output of all pairwise classifiers is combined using a majority voting rule to decide on the facial expression. Experiments have been carried out on three publicly available datasets (JAFFE, CK and TFEID) and the correct classification rates of 99.05%, 98.07% and 99.63% were achieved respectively. Therefore, the pairwise approach is effective to discriminate between different facial expressions and the results achieved by the proposed approach are slightly better than several current approaches.
Marcelo J. Cossetin, Júlio C. Nievola, Alessandro L. Koerich
IJCNN3
2016 Native Language Detection Using the I-Vector Framework
Mohammed Senoussaoui, Patrick Cardinal, Najim Dehak, Alessandro L. Koerich
INTERSPEECH4
2016 A flexible hierarchical approach for facial age estimation based on multiple features
Jhony K. Pontes, Alceu S. Britto Jr., Clinton Fookes, Alessandro L. Koerich
Pattern Recognit.4
2015 Visual and acoustic identification of bird species
abstract
This paper presents a novel approach for bird species identification that relies on both visual features extracted from unconstrained bird images and acoustic features extracted from bird vocalizations. The Scale Invariant Feature Transform (SIFT) detects local features in bird images, which are then used to train a support vector machine classifier. The instances that are not classified with a certain degree of certainty are then rejected and reclassified using Mel-frequency cepstral coefficients (MFCCs) extracted from the bird songs if available. Experiments conducted on a dataset of 50 bird species that comprise images from the CUB200-2011 and audio samples from Xeno-Canto have shown that improvements between 1.2 and 15.7 percentage points are achieved when using an acoustic classifier to re-process the instances rejected by the visual classifier, depending on the rejection level.
Andreia Marini, Alef J. Turatti, Alceu S. Britto Jr., Alessandro L. Koerich
ICASSP4
2015 Combining overall and local class accuracies in an oracle-based method for dynamic ensemble selection
abstract
This paper presents a k-nearest oracle-based dynamic ensemble selection method in which overall local accuracy (OLA) and local class accuracy (LCA) are combined into a twostep selection scheme. The OLA and LCA are computed on the neighborhood of the test pattern in a validation set to filter out the classifiers selected by the k-nearest oracles. The complementary information of OLA and LCA has shown to be an interesting alternative to approximate the classification performance to that estimated for the oracle of the initial pool of classifiers. The results were compared with the recognition rates of the majority voting of all classifiers in the initial pool, and also with the recognition rates of related classifier and ensemble selection methods which have inspired the proposed method and its variants. The proposed method achieved the best results on 5 out of 8 experiments using small and large datasets of different applications.
Leila Maria Vriesmann, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Robert Sabourin
IJCNN4
2015 PKLot - A robust dataset for parking lot classification
Paulo R. L. Almeida, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Eunelson Jose da Silva Junior, Alessandro L. Koerich
Expert Syst. Appl.5
2014 An HMM-Based Gesture Recognition Method Trained on Few Samples
abstract
This paper addresses the problem of recognizing gestures which are captured using the Kinect sensor in a educational game devoted to the deaf community. Different strategies are evaluated to deal with the problem of having few samples for training. We have experimented a Leave One Out Training and Testing (LOOT) strategy and an HMM-based ensemble of classifiers. A dataset containing 181 videos of gestures related to nine signs commonly used in educational games is introduced, which is available for research purposes. The experimental results have shown that the proposed ensemble-based method is a promising strategy to deal with problems where few training samples are available.
Vinicius Godoy, Alceu S. Britto Jr., Alessandro L. Koerich, Jacques Facon, Luiz Eduardo Soares de Oliveira
ICTAI3
2014 Automatic detection of musicians' ancillary gestures based on video analysis
Rodrigo A. Seger, Marcelo M. Wanderley, Alessandro L. Koerich
Expert Syst. Appl.3
2013 Music Genre Recognition Using Gabor Filters and LPQ Texture Descriptors
Yandre M. G. Costa, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Fabien Gouyon
CIARP (2)3
2013 Parking Space Detection Using Textural Descriptors
abstract
In this paper we assess the use of textural de-scriptors for the problem of parking space detection. We focus our experiments on two descriptors (Local Binary Patterns and Local Phase Quantization) that have attracted a great deal of attention because of their outstanding performance in a number of applications. We show through a series of comprehensive experiments that both descriptors are able to achieve very low error rates on a database composed of 105,837 images of parking spaces. We also show that the combination of the diverse classifiers developed in this work can bring further improvement achieving an error rate of 0.16%. The results reached in this work compare favorably to other published methods.
Paulo R. L. Almeida, Luiz Eduardo Soares de Oliveira, Eunelson Jose da Silva Junior, Alceu S. Britto Jr., Alessandro L. Koerich
SMC5
2013 Bird Species Classification Based on Color Features
abstract
This paper presents a novel approach for bird species classification based on color features extracted from unconstrained images. This means that the birds may appear in different scenarios as well may present different poses, sizes and angles of view. Besides, the images present strong variations in illuminations and parts of the birds may be occluded by other elements of the scenario. The proposed approach first applies a color segmentation algorithm in an attempt to eliminate background elements and to delimit candidate regions where the bird may be present within the image. Next, the image is split into component planes and from each plane, normalized color histograms are computed from these candidate regions. After aggregation processing is employed to reduce the number of the intervals of the histograms to a fixed number of bins. The histogram bins are used as feature vectors to by a learning algorithm to try to distinguish between the different numbers of bird species. Experimental results on the CUB-200 dataset show that the segmentation algorithm achieves 75% of correct segmentation rate. Furthermore, the bird species classification rate varies between 90% and 8%, depending on the number of classes taken into account.
Andreia Marini, Jacques Facon, Alessandro L. Koerich
SMC3
2013 Network infrastructure design with a multilevel algorithm
Hideson Alves da Silva, Alceu S. Britto Jr., Luis Eduardo Soares de Oliveira, Alessandro L. Koerich
Expert Syst. Appl.4
2013 Fusion of feature sets and classifiers for facial expression recognition
Thiago H. H. Zavaschi, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich
Expert Syst. Appl.4
2012 Comparing textural features for music genre classification
abstract
In this paper we compare two different textural feature sets for automatic music genre classification. The idea is to convert the audio signal into spectrograms and then extract features from this visual representation. Two textural descriptors are explored in this work: the Gray Level Co-Occurrence Matrix (GLCM) and Local Binary Patterns (LBP). Besides, two different strategies of extracting features are considered: a global approach where the features are extracted from the entire spectrogram image and then classified by a single classifier; a local approach where the spectrogram image is split into several zones which are classified independently and final decision is then obtained by combining all the partial results. The database used in our experiments was the Latin Music Database, which contains music pieces categorized into 10 musical genres, and has been used for MIREX (Music Information Retrieval Evaluation eXchange) competitions. After a comprehensive series of experiments we show that the SVM classifier trained with LBP is able to achieve a recognition rate of 80%. This rate not only outperforms the GLCM by a fair margin but also is slightly better than the results reported in the literature.
Yandre M. G. Costa, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Fabien Gouyon
IJCNN3
2012 Music genre classification using dynamic selection of ensemble of classifiers
abstract
This paper presents a dynamic ensemble selection method for music genre classification which employs two pools of diverse classifiers. The pools of classifiers are created by using different features types extracted from three distinct segments of each music piece. From these initial pools of weak classifiers, ensembles of classifiers are dynamically selected for each test pattern using the k-nearest oracles method. The experiments compare the performance of different selection strategies on the Latin Music Database to those related to the use of best single classifier, and to the combination of all classifiers in the pool. It was possible to observe that the most promising selection strategy evaluated allows improving the classification accuracy from 63.71% to 70.31%.
Paulo R. L. Almeida, Eunelson Jose da Silva Junior, Tatiana Montes Celinski, Alceu S. Britto Jr., Luis Eduardo Soares de Oliveira, Alessandro L. Koerich
SMC6
2012 Music genre classification using LBP textural features
Yandre M. G. Costa, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Fabien Gouyon, Jefferson G. Martins
Signal Process.3
2011 Facial expression recognition using ensemble of classifiers
abstract
This paper presents a novel method for facial expression classification that employs the combination of two different feature sets in an ensemble approach. A pool of base classifiers is created using two feature sets: Gabor filters and local binary patterns (LBP). Then a multi-objective genetic algorithm is used to search for the best ensemble using as objective functions the accuracy and the size of the ensemble. The experimental results on two databases have shown the efficiency of the proposed strategy by finding powerful ensembles, which improves the recognition rates between 5% and 10%.
Thiago H. H. Zavaschi, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira
ICASSP2
2011 Automatic Bird Species Identification for Large Number of Species
abstract
In this paper we focus on the automatic identification of bird species from their audio recorded song. Bird monitoring is important to perform several tasks, such as to evaluate the quality of their living environment or to monitor dangerous situations to planes caused by birds near airports. We deal with the bird species identification problem using signal processing and machine learning techniques. First, features are extracted from the bird recorded songs using specific audio treatment, next the problem is performed according to a classical machine learning scenario, where a labeled database of previously known bird songs are employed to create a decision procedure that is used to predict the species of a new bird song. Experiments are conducted in a dataset of recorded songs of bird species which appear in a specific region. The experimental results compare the performance obtained in different situations, encompassing the complete audio signals, as recorded in the field, and short audio segments (pulses) obtained from the signals by a split procedure. The influence of the number of classes (bird species) in the identification accuracy is also evaluated.
Marcelo Teider Lopes, Lucas L. Gioppo, Thiago T. Higushi, Celso A. A. Kaestner, Carlos Nascimento Silla Jr., Alessandro L. Koerich
ISM6
2011 Feature set comparison for automatic bird species identification
abstract
This paper deals with the automated bird species identification problem, in which it is necessary to identify the species of a bird from its audio recorded song. This is a clever way to monitor biodiversity in ecosystems, since it is an indirect non-invasive way of evaluation. Different features sets which summarize in different aspects the audio properties of the audio signal are evaluated in this paper together with machine learning algorithms, such as probabilistic, instance-based, decision trees, neural networks and support vector machines. Experiments are conducted in a dataset of recorded songs of three bird species. The experimental results compare the performance of the features sets and different classifiers showing that it is possible to obtain very promising results in the automated bird species identification problem.
Marcelo Teider Lopes, Carlos Nascimento Silla Jr., Alessandro L. Koerich, Celso A. A. Kaestner
SMC3
2010 Verification of Unconstrained Handwritten Words at Character Level
abstract
In this paper we present a verification module that has as input the output provided by a word recognizer which is based on the segmentation-recognition paradigm. The word recognizer models words as the concatenation of character hidden Markov models (HMMs) and it provides at the output a list with the Top N best word hypotheses, including their likelihoods and the segmentation points of the words into sub words, which ideally should be characters. The verification module uses the segmentation points provided by the word recognizer for each word hypothesis to extract different features from each sub word. A classifier based on a multilayer perceptron neural network assigns a character class (A-Z) and estimates the a posteriori probability to each sub word that make up a word. Further, both the character class and the a posteriori probabilities are combined with the original output of the word recognizer to re-rank the word hypothesis into the Top N list. Experimental results show that the verification module improves the Top 1 recognition rate in 3.9% for an 85,092-word recognition task.
Alessandro L. Koerich, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira
ICFHR1
2010 Selection of Training Instances for Music Genre Classification
abstract
In this paper we present a method for the selection of training instances based on the classification accuracy of a SVM classifier. The instances consist of feature vectors representing short-term, low-level characteristics of music audio signals. The objective is to build, from only a portion of the training data, a music genre classifier with at least similar performance as when the whole data is used. The particularity of our approach lies in a pre-classification of instances prior to the main classifier training: i.e. we select from the training data those instances that show better discrimination with respect to class memberships. On a very challenging dataset of 900 music pieces divided among 10 music genres, the instance selection method slightly improves the music genre classification in 2.4 percentage points. On the other hand, the resulting classification model is significantly reduced, permitting much faster classification over test data.
Miguel Lopes, Fabien Gouyon, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira
ICPR3
2010 On the suitability of state-of-the-art music information retrieval methods for analyzing, categorizing and accessing non-Western and ethnic music collections
Thomas Lidy, Carlos Nascimento Silla Jr., Olmo Cornelis, Fabien Gouyon, Andreas Rauber, Celso A. A. Kaestner, Alessandro L. Koerich
Signal Process.7
2009 Evaluation of Different Strategies to Optimize an HMM-Based Character Recognition System
abstract
Different strategies for combination of complementary features in an HMM-based method for handwritten character recognition are evaluated. In addition, a noise reduction method is proposed to deal with the negative impact of low probability symbols in the training database. New sequences of observations are generated based on the original ones, but considering a noise reduction process. The experimental results based on 52 classes of alphabetic characters and more than 23,000 samples have shown that the strategies proposed to optimize the HMM-based recognition method are very promising.
Murilo Santos, Albert Hung-Ren Ko, Luiz Eduardo Soares de Oliveira, Robert Sabourin, Alessandro L. Koerich, Alceu S. Britto Jr.
ICDAR5
2009 Combining different biometric traits with one-class classification
Cheila Bergamini, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Robert Sabourin
Signal Process.3
2008 Fusion of biometric systems using one-class classification
abstract
One of the main requirements of biometric systems is the ability of producing very low false acceptation rate, which very often can be achieved only by combining different biometric traits. The literature has shown that the pattern classification approach usually surpasses the classifier combination approach for this task. In this work we take into account the pattern classification approach, but considering the one-class classification approach. We show that one-class classification could be considered as an alternative for biometric fusion specially when the data is highly unbalanced or data from a single class is available. The results for one-class classification reported in this paper compares to the standard two-class SVM and surpasses all the conventional classifier combination rules tested.
Cheila Bergamini, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Robert Sabourin
IJCNN3
2008 Feature Selection in Automatic Music Genre Classification
abstract
This paper presents the results of the application of a feature selection procedure to an automatic music genre classification system. The classification system is based on the use of multiple feature vectors and an ensemble approach, according to time and space decomposition strategies. Feature vectors are extracted from music segments from the beginning, middle and end of the original music signal (time decomposition). Despite being music genre classification a multi-class problem, we accomplish the task using a combination of binary classifiers, whose results are merged in order to produce the final music genre label (space decomposition). As individual classifiers several machine learning algorithms were employed: naive-Bayes, decision trees, support vector machines and multi-layer perceptron neural nets. Experiments were carried out on a novel dataset called Latin music database, which contains 3,227 music pieces categorized in 10 musical genres. The experimental results show that the employed features have different importance according to the part of the music signal from where the feature vectors were extracted. Furthermore, the ensemble approach provides better results than the individual segments in most cases.
Carlos Nascimento Silla Jr., Alessandro L. Koerich, Celso A. A. Kaestner
ISM2
2008 Filtering segmentation cuts for digit string recognition
Eduardo Vellasques, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Alessandro L. Koerich, Robert Sabourin
Pattern Recognit.4
2007 WEB Image Classification Based on the Fusion of Image and Text Classifiers
abstract
This paper presents a novel method for the classification of images that combines information extracted from the images and contextual information. The main hypothesis is that contextual information related to an image can contribute in the image classification process. First, independent classifiers are designed to deal with images and text. From the images color, shape and texture features are extracted. These features are used with a neural network (NN) classifier to carry out image classification. On the other hand, contextual information is processed and used with a Naive Bayes (NB) classifier. At the end, the outputs of both classifiers are combined through heuristic rules. Experimental results on a database of more than 5,000 HTML documents have shown that the combination of classifiers provides a meaningful improvement (about 16%) in the correct image classification rate relative to the results provided by the NN classifier alone.
Pedro R. Kalva, Fabrício Enembreck, Alessandro L. Koerich
ICDAR3
2007 Detection and Classification of Human Movements in Video Scenes
Andre G. Hochuli, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Alessandro L. Koerich
PSIVT4
2007 People Counting in Low Density Video Sequences
Jaime Dalla Valle, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Alceu S. Britto Jr.
PSIVT3
2007 Detection of non-conventional events on video scenes
abstract
This article presents a novel approach for detection of non-conventional events in videos scenes. This novel approach consists in analyzing in real-time video from a security camera to detect, segment and tracking objects in movement to further classify its movement as conventional or non-conventional. From each tracked object in the scene features such as position, speed, changes in directions and in the bounding box sizes are extracted. These features make up a feature vector. At the classification step, feature vectors generated from objects in movement in the scene are matched almost in real-time against reference feature vectors previously labeled which are stored in a database and an algorithm based on the instance-based learning paradigm is used to classify the object movement as conventional or non-conventional. Experimental results on video clips from two databases (Parking Lot and CAVIAR) have shown that the proposed approach is able to detect non-conventional events with accuracies between 77% and 82%.
Andre G. Hochuli, Alceu S. Britto Jr., Alessandro L. Koerich
SMC3
2007 Automatic music genre classification using ensemble of classifiers
abstract
This paper presents a novel approach to the task of automatic music genre classification which is based on multiple feature vectors and ensemble of classifiers. Multiple feature vectors are extracted from a single music piece. First, three 30-second music segments, one from the beginning, one from the middle and one from end part of a music piece are selected and feature vectors are extracted from each segment. Individual classifiers are trained to account for each feature vector extracted from each music segment. At the classification, the outputs provided by each individual classifier are combined through simple combination rules such as majority vote, max, sum and product rules, with the aim of improving music genre classification accuracy. Experiments carried out on a large dataset containing more than 3,000 music samples from ten different Latin music genres have shown that for the task of automatic music genre classification, the features extracted from the middle part of the music provide better results than using the segments from the beginning or end part of the music. Furthermore, the proposed ensemble approach, which combines the multiple feature vectors, provides better accuracy than using single classifiers and any individual music segment.
Carlos Nascimento Silla Jr., Celso A. A. Kaestner, Alessandro L. Koerich
SMC3
2005 Unconstrained handwritten character recognition using metaclasses of characters
abstract
In this paper we tackle the problem of unconstrained handwritten character recognition using different classification strategies. For such an aim, four multilayer perceptron classifiers (MLP) were built and used into three different classification strategies: combination of two 26-class classifiers; 26-metaclass classifier; 52-class classifier. Experimental results on the NIST SD19 database have shown that the recognition rate achieved by the metaclass classifier (87.8%) outperforms the other approaches (82.9% and 86.3%).
Alessandro L. Koerich, Pedro R. Kalva
ICIP (2)1
2005 Combination of homogeneous classifiers for musical genre classification
abstract
Content-based music genre classification is a useful tool for multimedia indexing and retrieval. In this paper a novel content-based music genre classification approach that employs combination of homogeneous classifiers is proposed. First, musical surface features and beat-related features are extracted from different pans of music tracks and three 15-dimensional feature vectors are generated. The features are extracted from the beginning, middle and end parts of the music. These features vectors are used to train three multilayer perceptron neural network classifiers. At the classification step, the outputs provided by each neural network based classifier are combined using max, sum and product rules. Experimental results show that the proposed combination of homogeneous classifiers outperforms single feature vectors and single classifiers, achieving higher correct music genre classification rates.
Alessandro L. Koerich, Cleverson Poitevin
SMC1
2005 Recognition and Verification of Unconstrained Handwritten Words
abstract
This paper presents a novel approach for the verification of the word hypotheses generated by a large vocabulary, offline handwritten word recognition system. Given a word image, the recognition system produces a ranked list of the N-best recognition hypotheses consisting of text transcripts, segmentation boundaries of the word hypotheses into characters, and recognition scores. The verification consists of an estimation of the probability of each segment representing a known class of character. Then, character probabilities are combined to produce word confidence scores which are further integrated with the recognition scores produced by the recognition system. The N-best recognition hypothesis list is reranked based on such composite scores. In the end, rejection rules are invoked to either accept the best recognition hypothesis of such a list or to reject the input word image. The use of the verification approach has improved the word recognition rate as well as the reliability of the recognition system, while not causing significant delays in the recognition process. Our approach is described in detail and the experimental results on a large database of unconstrained handwritten words extracted from postal envelopes are presented.
Alessandro L. Koerich, Robert Sabourin, Ching Y. Suen
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Improving classification performance using metaclasses
abstract
In this paper we propose a new methodology to improve the performance of classifiers on relatively difficult classification problems with complex boundaries between classes, overlapping classes, and a lack of sufficient number of samples for some classes. We investigate the use of contextual information to overcome such problems, especially in the case of class overlapping, high number of classes and high dimensional feature spaces. The proposed methodology assumes that contextual information is available and that it can be used to disambiguate overlapping classes. In fact, the contextual information is used to reduce the number of classes as well as at the design of the classifier. This new methodology was applied to the problem of unconstrained handwritten character recognition where we have up to 52 different classes (A-Z, a-z). Experimental results on a 100,000-character database show that it is possible to reduce the number of classes and the complexity of the classifier and, at the same time, to improve the recognition accuracy in more than 17%.
Alessandro L. Koerich
SMC1
2003 Lexicon-driven HMM decoding for large vocabulary handwriting recognition with multiple character models
Alessandro L. Koerich, Robert Sabourin, Ching Y. Suen
Int. J. Document Anal. Recognit.1
2003 Large vocabulary off-line handwriting recognition: A survey
Alessandro L. Koerich, Robert Sabourin, Ching Y. Suen
Pattern Anal. Appl.1
2002 Fast two-level Viterbi search algorithm for unconstrained handwriting recognition
abstract
This paper describes a fast two-level Viterbi search algorithm for recognizing handwritten words as a sequence of characters concatenated according to a lexicon. The algorithm is based on hidden Markov model (HMM) representations of characters and it breaks up the computation of word likelihood scores into two levels: state level and character level. This enables the reuse of likelihood scores of characters to decode all words in the lexicon, avoiding repeated computation of state sequences. Experimental results with an 85,000-word vocabulary indicate that the computational cost of an off-line handwritten word recognition system may be reduced by more than a factor of 20 while not introducing search errors.
Alessandro L. Koerich, Robert Sabourin, Ching Y. Suen
ICASSP1
2001 A Distributed Scheme for Lexicon-Driven Handwritten Word Recognition and its Application to Large Vocabulary Problems
abstract
Many offline handwritten word recognition systems have been proposed since the early nineties. Most systems reported high recognition rates, however, they overlooked a very important factor in the process: speed factor. The authors explore the potential for speeding up an offline handwritten word recognition system via concurrency. The goal of the system is to achieve both full accuracy and high speed when taking into account large vocabularies. This was accomplished by integrating the recognition process with multiprocessing and distributed computing concepts. Experimental results showed that the multiprocessing environment is very promising in enhancing a sequential offline handwritten word recognition system performance.
Alessandro L. Koerich, Robert Sabourin, Ching Y. Suen
ICDAR1
1999 Automatic Storage, Retrieval, and Visualization of Bank Check Images
abstract
This paper presents an automated system for storage and retrieval of bank checks in contrast with the microfilming techniques that are currently used. The bank check images are introduced into an extraction module where the filled in information is segmented. This information is indexed via keywords derived from the MICR line and stored in a database under a hybrid structure where hash tables, trees and inverted files are employed. For the information retrieval and visualization, make-up bank check images are generated. The experimental results reveal a good performance of the proposed method in terms of compactness of stored information and high visual quality of the reconstructed images.
Alessandro L. Koerich, Lee Luan Ling
ICDAR1
1997 Compression of bank cheque images based on layout knowledge
abstract
In this paper a scheme for bank cheque images compression based on layout knowledge is proposed. The layout structure of the cheques is analyzed and the nonessential parts are located. These parts, viz., the background and the printed information, are eliminated from the original image. The resulting image contains some noise that are eliminated by a filtering operation. The image is enclosed to eliminate some uninformative parts. The final image has only the filled information. The digitized image can be easily reconstructed by restoring the filled information and summing it with background and printed information. The proposed compression scheme is tested by Brazilian bank cheques. Comparisons with other compression schemes, shows that the proposed scheme performs significantly better in terms of the compression efficiency, maintaining the visual quality.
Alessandro L. Koerich, Lee Luan Ling
ICASSP1
1997 A Prototype for Brazilian Bankcheck Recognition
abstract
This paper describes a prototype for Brazilian bankcheck recognition. The description is divided into three topics: bankcheck information extraction, digit amount recognition and signature verification. In bankcheck information extraction, our algorithms provide signature and digit amount images free of background patterns and bankcheck printed information. In digit amount recognition, we dealt with the digit amount segmentation and implementation of a complete numeral character recognition system involving image processing, feature extraction and neural classification. In signature verification, we designed and implemented a static signature verification system suitable for banking and commercial applications. Our signature verification algorithm is capable of detecting both simple, random and skilled forgeries. The proposed automatic bankcheck recognition prototype was intensively tested by real bankcheck data as well as simulated data providing the following performance results: for skilled forgeries, 4.7% equal error rate; for random forgeries, zero Type I error and 7.3% Type II error; for bankcheck numerals, 92.7% correct recognition rate.
Lee Luan Ling, Miguel Gustavo Lizárraga, Natanael Rodrigues Gomes, Alessandro L. Koerich
Int. J. Pattern Recognit. Artif. Intell.4