Ing Ren Tsang

dblp:98/5344 · also Tsang Ing Ren · DBLP profile ↗
← Back
105ranked-venue papers
6as first author
15since 2021 · last 2025
0000-0002-3677-0264ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 64 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20Human-computer interaction and ubiquitous computing · 16 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2025 Improving COVID-19 Detection in Chest X-Rays Using EfficientNet with Self-Supervised Contrastive Learning
abstract
We propose a self-supervised framework for COVID-19 detection from chest X-rays (CXRs) that combines EfficientNet with four contrastive learning techniques: DINOCXR, BYOL, SimCLR, and VICReg. By pretraining on over 13,000 unlabeled CXRs from the ChestX-ray14-v3 dataset and fine-tuning on the COVIDGR dataset, our approach addresses the challenge of scarce labeled data in medical imaging. Experimental results show that SimCLR achieves the highest recall (77.78%), making it well-suited for initial screening, while EfficientNet finetuned directly provides the best precision$(87.18 \%)$, suggesting its use for confirmatory diagnosis. Our findings reveal that this twostage combination outperforms state-of-the-art models such as COVIDNet-CXR and DINO-CXR in terms of F1-score and data efficiency, using only 6% of the labeled data. This hybrid strategy offers a compelling balance between sensitivity and specificity, making it particularly valuable for clinical deployment.
Tales T. Alves, Ing Ren Tsang, Cleber Zanchettin
ICTAI2
2025 LIGA: a LIghtweight CNN architecture designed to classify popular music Genres from the Amazonian region
abstract
Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) can deliver outstanding performance when appropriately optimized for specific tasks. Although these architectures typically require substantial computational resources, lightweight CNN models may achieve comparable or even superior efficiency in well-defined and constrained scenarios. Therefore, the effectiveness of the approach depends on the careful tuning of the hyperparameter and the deliberate selection of an architecture aligned with the characteristics of the target problem. This work presents LIGA, a LIghtweight CNN-based architecture designed to classify popular music Genres from the Amazonian region, including andean, brega, carimbó, cumbia, merengue, pasillo, salsa, and vaqueirada, originating from countries such as Bolivia, Brazil, Colombia, Ecuador, French Guiana, Peru, the Dominican Republic, and Venezuela. In addition to low computational resource usage and improved training speed, LIGA achieved higher precision and accuracy compared to the EfficientNet, MobileNet, ResNet, VGG, Xception, MobileViT, and MaxViT models.
Cláudio Gomes 0004, Ing Ren Tsang
SMC2
2023 A Neurally Guided Patch-Based Style Transfer for Mobile Devices
abstract
Style transfer is an application that has increased interest, primarily because of the impressive results obtained using neural networks. However, this application demands a lot of computational resources, thus preventing its use in low-end mobiles. The patch-based approach is an interesting alternative that consumes less memory. This work uses two methods to generate high-resolution stylized images. Gated convolution in the neural network and the Halide language in the patch-based implementation are used for optimization. As a result, we were able to apply the style transfer method on a mobile device with 4 GB of memory RAM running under 6 seconds for$1920\times 1080$image size preserving the high-frequency details.
Jose Ivson S. Silva, Kevin Ian Ruiz Vargas, Antônio A. Carlos, Lucas P. de Albuquerque, Mateus Baltazar de Almeida, Allan Soares Vasconcelos, Victor Ximenes C. Oliveira, José Gabriel P. Tavares, Danilo Vaz Marcolino Alves, Diêgo J. C. Santiago, Bernardo Augusto de Oliveira, Carlos Padilha, Ing Ren Tsang
IJCNN13
2023 Vector Representation and Machine Learning for Short-Term Photovoltaic Power Prediction
abstract
Short-term photovoltaic (PV) energy production forecasting is critical for managing grid-connected systems and energy trading. Machine learning models are widely used for accurate prediction, and this study proposes using Time2Vec as an embedding for a transformer-based neural network architecture. Experiments on two PV power plants in India showed significant improvements comparing our proposed architecture to MLP, LSTM, and the persis-tence model, which is a standard baseline prediction in this type of forecasting, with over 20 % improvements in some horizons. These findings demonstrate the effectiveness of the proposed approach for short-term PV forecasting using machine learning models.
Renan Costa, Olga Vilela, Ing Ren Tsang
SMC4
2022 Convolutional Decoupled cVAE-GANs for Pseudo-Replay Based Continual Learning
abstract
Continual Learning is the concept of having a model able to sequentially learn to solve new tasks without losing the ability to solve previous tasks. Achieving this is challenging because neural networks usually suffer from catastrophic forgetting of the preceding tasks when they are learning new ones. To handle this issue, pseudo-replay approaches leverages the performance of generative networks using them to generate samples related to past data to serve as input to the model when it is learning new tasks. In this work, we propose an improved architecture and training strategy based on the state-of-the-art pseudo-replay IRCL method. We use a cVAE-GAN as the generative model and train it decoupled from the other components of the architecture. Also, we make use of convolutional layers for the architecture components instead of linear ones. Our experimental results show that the proposed method outperforms the state-of-the-art IRCL method by up to 10% in Average Accuracy and up to 8.3% in Average Backward Transfer on both Split MNIST and Split FashionMNIST datasets.
Lucas A. M. De Alcantara, Jose Ivson S. Silva, Miguel L. P. C. Silva, Saulo C. S. Machado, Ing Ren Tsang
ICTAI5
2022 U-Net Based Discriminator for Real-World Super-Resolution
abstract
In recent years, single image super-resolution (SISR) deep learning techniques have achieved remarkable improvements in recovering a high-resolution (HR) image from an observed low-resolution (LR) input. Nevertheless, these proposed methods fail in many real-world scenarios since their models are usually trained using a pre-defined degradation process from HR ground truth images to LR ones. To address this issue, new architectures have been proposed focusing on adopting more complicated degradation models to emulate real-world degradation achieving prominent performance but still limited to certain kinds of inputs and dropping considerably in other cases. In this paper, we present a GAN structure for blind super-resolution tasks, applying a technique that has not been very commonly used in SISR proposals: a U-Net architecture as a discriminator of the GAN network. Adding this structural change will encourage the discriminator to focus more on semantic and structural changes between real and fake images and to attend less to domain-preserving perturbations. In addition, the loss function of the generator was modified by adding the LPIPS loss function for the perceptual loss and a per-pixel consistency regularization technique based on the CutMix data augmentation. Numerous novel solutions that have been proposed recently involve powerful deep learning techniques. The proposed model was trained using the DF2K dataset employing a degradation framework for real-world images by estimating blur kernels and real noise distributions to obtain more realistic LR samples. Finally, we present a benchmark comparing our results with other methods in the state-of-the-art. The commonly-used evaluation metrics for image restoration PSNR, SSIM, and LPIPS were used for this evaluation.
Kevin Ian Ruiz Vargas, Fidel A. Guerrero-Peña, Pedro D. Marrero-Fernández, Leonardo Lanfranchi, Ing Jyh Tsang, Ing Ren Tsang
ICTAI6
2022 Low-discrepancy Sampling for Full Reference Image Quality Assessment Speed Up
abstract
Image quality assessment (IQA) aims to predict the image quality perceived by the human visual system (HVS). Full Reference (FR) image quality assessment is an objective algorithm requiring information about the reference image for quality assessment. Consequently, the FR algorithms may need a high number of operations to complete the evaluation. Another relevant point is that IQA based on Convolutional Neural Networks (CNN) requires a long training. Considering the high computational cost of FR assessment, we propose to use sampling methods as an alternative to the conventional IQA. First, we apply Van der Corput-Halton, Sobol, and uniform sampling methods to obtain a small representation of the images. Afterwards, we evaluate the sampled image using the Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Deep Image Quality Measure for FR (DIQaM-FR) metrics. The experimental results reveal that 7.8% of image pixels of the Live database are sufficient to obtain approximate values of SSIM and low mean error of PSNR. The sampling blocks used in the training of DIQaM-FR demonstrate to be adequate for training the model showing a correlation of 0.968 for SROCC applied on the Live database and a considerably lower training time.
Jair Galvão, Ing Ren Tsang, Francisco Madeiro, Emerson A. O. Lima
IJCNN2
2022 Ensemble of Convolutional Neural Networks for Sparse-View Cone-Beam Computed Tomography
abstract
Risks related to excessive exposure of patients to ionizing radiation are a significant concern in the medical community. Several approaches based on Convolutional Neural Networks (CNNs) have been proposed to develop safer and more reliable sparse-view Computed Tomography (SVCT) systems. Most of those solutions process tomographic data within 2D slices individually. However, recent works have shown that 3D models - that exploit data correlation among adjacent slices - can outperform previous 2D models. Once the kernel size in most of those 3D models is not bigger than 5 x 5 x 5, such inter-slice exploration is restricted to a limited neighborhood, resulting in minor inter-slice analysis during the training/validation phase. To efficiently exploit data correlation among the coronal, axial, and sagittal views of the SVCT volume, we propose an ensemble of four 2D CNNs. Three of them are used to process the orthogonal SVCT volume views separately, and the fourth CNN combines the outputs from the previous three networks. Since our final architecture is highly deep, we also present a training method in stages to avoid the non-convergence of the deepest layers. We conducted experiments using head cone-beam Computed Tomography (CBCT) scans extensively used in imaged guided radiotherapy (IGTR) during brain tumor treatment. Our method presented superior results in reducing reconstruction artifacts of SVCT volumes compared to the state-of-the-art 2D and 3D models.
Carlos A. Alves Júnior, Luis Filipe Alves Pereira, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN4
2022 Wind Power Forecast Based on Transformers and Clustering of Wind Farms with Temporal and Spatial Interdependence
abstract
Wind power forecasting remains a challenging area of research due to the complex relationships present in the conversion of wind into electrical power and the uncertainty associated with the physical modeling of the wind, which is highly dependent on the initial conditions. The need for a precise forecast is essential to meet the goals of high integration of renewable sources in the energy matrix to optimize non-renewable resources and improve the reliability in the operation of the electric system. In this work, we applied a wind power forecasting model based on Transformers that uses a self-attention encoder and a seq-2-seq decoder with global attention to capturing the nonlinear relationships of the power time series of wind farms. This evaluation considers the behavior of groups of wind farms with similar behavior. In terms of the similarity metric, five distant and close wind farms are clustered by the k-means method, using the autocorrelation features for time lags from 3 to 24 hours, mean and standard deviation, selected among 447 wind farms located in the NE region of Brazil. Compared to traditional linear autoregressive and LSTM models, the results show a lower prediction error. In this way, the model can capture temporal correlations arising from autoregressive and spatial behaviors. This shows the interdependence between hundreds of wind farms spread over a vast region with high wind potential, different altitudes, orography, distance from the coast, and wind regimes.
Fábio Lima, Ing Ren Tsang
IJCNN2
2022 An Ensemble Learning Method for Segmentation Fusion
abstract
The segmentation of cells present in microscope images is an essential step in many tasks, including determining protein concentration and analysis of gene expression per cell. In single-cell genomics studies, cell segmentations are vital to assess the genetic makeup of individual cells and their relative spatial location. Several methods and tools have been developed to offer robust segmentation, with deep learning models currently being the most promising solutions. As an alternative to developing another cell segmentation targeted model, we propose a learning ensemble strategy that aggregates many independent candidate segmentations of the same image to produce a single consensus segmentation. We are particularly interested in learning how to ensemble crowdsource image segmentations created by experts and non-experts in laboratories and data houses. We compare our trained ensemble model with other fusion methods adopted by the biomedical community and assess the robustness of the results on three aspects: fusion with outliers, missing data, and synthetic deformations. Our approach outperforms these methods in efficiency and quality, especially when there is a high disagreement among candidate segmentations of the same image.
Carlos H. C. Pena, Ing Ren Tsang, Pedro D. Marrero-Fernández, Fidel A. Guerrero-Peña, Alexandre Cunha
IJCNN2
2022 A Library and Web Platform for RoboCup Soccer Matches Data Analysis
Felipe N. A. Pereira, Mateus F. B. Soares, Conceição Rocha, Tales T. Alves, Tiago H. R. P. Gonçalves, José R. da Silva, Ing Ren Tsang, Paulo S. G. de Mattos Neto, Edna Barros
RoboCup7
2022 Web Soccer Monitor: An Open-Source 2D Soccer Simulation Monitor for the Web and the Foundation for a New Ecosystem
Mateus F. B. Soares, Ing Ren Tsang, Paulo S. G. de Mattos Neto, Edna Barros
RoboCup2
2022 Entropic Out-of-Distribution Detection: Seamless Detection of Unknown Examples
abstract
In this article, we argue that the unsatisfactory out-of-distribution (OOD) detection performance of neural networks is mainly due to the SoftMax loss anisotropy and propensity to produce low entropy probability distributions in disagreement with the principle of maximum entropy. On the one hand, current OOD detection approaches usually do not directly fix the SoftMax loss drawbacks, but rather build techniques to circumvent it. Unfortunately, those methods usually produce undesired side effects (e.g., classification accuracy drop, additional hyperparameters, slower inferences, and collecting extra data). On the other hand, we propose replacing SoftMax loss with a novel loss function that does not suffer from the mentioned weaknesses. The proposed IsoMax loss is isotropic (exclusively distance-based) and provides high entropy posterior probability distributions. Replacing the SoftMax loss by IsoMax loss requires no model or training changes. Additionally, the models trained with IsoMax loss produce as fast and energy-efficient inferences as those trained using SoftMax loss. Moreover, no classification accuracy drop is observed. The proposed method does not rely on outlier/background data, hyperparameter tuning, temperature calibration, feature extraction, metric learning, adversarial training, ensemble procedures, or generative models. Our experiments showed that IsoMax loss works as a seamless SoftMax loss drop-in replacement that significantly improves neural networks' OOD detection performance. Hence, it may be used as a baseline OOD detection approach to be combined with current or future OOD detection techniques to achieve even higher results.
David Macedo, Ing Ren Tsang, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira, Teresa Bernarda Ludermir
IEEE Trans. Neural Networks Learn. Syst.2
2021 Entropic Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection approaches usually present special requirements (e.g., hyperparameter validation, collection of outlier data) and produce side effects (e.g., classification accuracy drop, slower energy-inefficient inferences). We argue that these issues are a consequence of the SoftMax loss anisotropy and disagreement with the maximum entropy principle. Thus, we propose the IsoMax loss and the entropic score. The seamless drop-in replacement of the SoftMax loss by IsoMax loss requires neither additional data collection nor hyperparameter validation. The trained models do not exhibit classification accuracy drop and produce fast energy-efficient inferences. Moreover, our experiments show that training neural networks with IsoMax loss significantly improves their OOD detection performance. The IsoMax loss exhibits state-of-the-art performance under the mentioned conditions (fast energy-efficient inference, no classification accuracy drop, no collection of outlier data, and no hyperparameter validation), which we call the seamless OOD detection task. In future work, current OOD detection methods may replace the SoftMax loss with the IsoMax loss to improve their performance on the commonly studied non-seamless OOD detection problem.
David Macedo, Ing Ren Tsang, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira, Teresa Bernarda Ludermir
IJCNN2
2021 Variational DNN embeddings for text-independent speaker verification
Hector N. B. Pinheiro, Ing Ren Tsang, André Adami, George D. C. Cavalcanti
Pattern Recognit. Lett.2
2020 Perturbation-based classifier
Edson L. Araújo, George D. C. Cavalcanti, Ing Ren Tsang
Soft Comput.3
2020 Burst Ranking for Blind Multi-Image Deblurring
abstract
We propose a new incremental aggregation algorithm for multi-image deblurring with automatic image selection. The primary motivation is that current burst deblurring methods do not handle well situations in which misalignment or out-of-context frames are present in the burst. These real-life situations result in poor reconstructions or manual selection of the images that are used to deblur. Automatically selecting the best frames within the burst to improve the base reconstruction is challenging because the number of possible images fusions is equal to the power set cardinal. Here, we approach the multi-image deblurring problem as a two steps process. First, we successfully learn a comparison function to rank a burst of images using a deep convolutional neural network. Then, an incremental Fourier burst accumulation with a reconstruction degradation mechanism is applied fusing only less blurred images that are sufficient to maximize the reconstruction quality. Experiments with the proposed algorithm have shown superior results when compared to other similar approaches, outperforming other methods described in the literature in previously described situations. We validate our findings on several synthetic and real datasets.
Fidel A. Guerrero-Peña, Pedro D. Marrero-Fernández, Ing Ren Tsang, Jorge J. G. Leandro, Ricardo Nishihara
IEEE Trans. Image Process.3
2020 Single image HDR reconstruction using a CNN with masked features and perceptual loss
abstract
Digital cameras can only capture a limited range of real-world scenes' luminance, producing images with saturated pixels. Existing single image high dynamic range (HDR) reconstruction methods attempt to expand the range of luminance, but are not able to hallucinate plausible textures, producing results with artifacts in the saturated areas. In this paper, we present a novel learning-based approach to reconstruct an HDR image by recovering the saturated pixels of an input LDR image in a visually pleasing way. Previous deep learning-based methods apply the same convolutional filters on wellexposed and saturated pixels, creating ambiguity during training and leading to checkerboard and halo artifacts. To overcome this problem, we propose a feature masking mechanism that reduces the contribution of the features from the saturated areas. Moreover, we adapt the VGG-based perceptual loss function to our application to be able to synthesize visually pleasing textures. Since the number of HDR images for training is limited, we propose to train our system in two stages. Specifically, we first train our system on a large number of images for image inpainting task and then fine-tune it on HDR reconstruction. Since most of the HDR examples contain smooth regions that are simple to reconstruct, we propose a sampling strategy to select challenging training patches during the HDR fine-tuning stage. We demonstrate through experimental results that our approach can reconstruct visually pleasing HDR results, better than the current state of the art on a wide range of scenes.
Marcel Santana Santos, Ing Ren Tsang, Nima Khademi Kalantari
ACM Trans. Graph.2
2019 Fast and robust multiple ColorChecker detection using deep convolutional neural networks
Pedro D. Marrero-Fernández, Fidel A. Guerrero-Peña, Ing Ren Tsang, Jorge J. G. Leandro
Image Vis. Comput.3
2019 Natural image segmentation with non-extensive mixture models
Dusan Stosic, Darko Stosic, Teresa Bernarda Ludermir, Ing Ren Tsang
J. Vis. Commun. Image Represent.4
2018 Multiclass Weighted Loss for Instance Segmentation of Cluttered Cells
abstract
We propose a new multiclass weighted loss function for instance segmentation of cluttered cells. We are primarily motivated by the need of developmental biologists to quantify and model the behavior of blood T -cells which might help us in understanding their regulation mechanisms and ultimately help researchers in their quest for developing an effective immunotherapy cancer treatment. Segmenting individual touching cells in cluttered regions is challenging as the feature distribution on shared borders and cell foreground are similar thus difficulting discriminating pixels into proper classes. We present two novel weight maps applied to the weighted cross entropy loss function which take into account both class imbalance and cell geometry. Binary ground truth training data is augmented so the learning model can handle not only foreground and background but also a third touching class. This framework allows training using U - N et. Experiments with our formulations have shown superior results when compared to other similar schemes, outperforming binary class models with significant improvement of boundary adequacy and instance detection. We validate our results on manually annotated microscope images of T-cells.
Fidel A. Guerrero-Peña, Pedro D. Marrero-Fernández, Ing Ren Tsang, Mary Yui, Ellen Rothenberg, Alexandre Cunha
ICIP3
2017 Speaker segmentation using i-vector in meetings domain
abstract
In this paper, we propose a speaker segmentation method for meeting audio based on i-vector. The motivation is to utilize the Total Variability (TV) framework as a feature extractor and to exploit the potential of modeling the speaker and channel variabilities for speaker segmentation in meetings. A distance-based segmentation method is designed with the cosine distance. A sliding window with variable length searches for speaker turns, through the distance between the i-vectors extracted from two segments with the same size. The experiments are conducted on the AMI Meeting Corpus, covering several conversation scenarios. For the training data of the UBM and TV matrix, 5 conversations from AMI Meeting Corpus are sampled. Other 10 conversations from AMI Meeting Corpus to compose the test data. The experiments show an improvement in the MDR and FAR curves compared with the FixSlid approach with different distance metrics, and for most of the operating points when compared with the classical BIC based WinGrow. The proposed method has on average a better computational performance, improving in 61.5% compared with the XBIC based FixSlid, and improving in 86.7% compared with the BIC based WinGrow.
Leonardo Valeriano Neri, Hector N. B. Pinheiro, Ing Ren Tsang, George D. C. Cavalcanti, André Adami
ICASSP3
2017 Optimizing speaker-specific filter banks for speaker verification
abstract
In this work, we investigate speaker-specific filter banks for text-independent speaker verification. The proposed method performs an heuristic search for the best filter-bank configuration using the Artificial Bee Colony (ABC) algorithm and a proper fitness function for the standard i-vectors/PLDA-based speaker verification system. Furthermore, filter-bank decorrelated amplitudes are used instead of the cepstral coefficients produced by Discrete Cosine Transform (DCT). In the experiments, the proposed method is compared to standard Mel and linear scales in both cases where the decorrelation is performed using DCT and high-pass filtering. The comparison is performed on the MIT Mobile Device Speaker Verification Corpus in a gender-dependent trial scheme. The proposed method outperformed the baseline systems in almost all the test sets for both genders. Performance gains of 4.6% and 26.0% are achieved for male and female speakers, respectively.
Hector N. B. Pinheiro, Fernando M. de Paula Neto, Adriano Lorena Inácio de Oliveira, Ing Ren Tsang, George D. C. Cavalcanti, André Adami
ICASSP4
2017 Combining dissimilarity spaces for text categorization
Roberto H. W. Pinheiro, George D. C. Cavalcanti, Ing Ren Tsang
Inf. Sci.3
2016 Type-2 fuzzy GMM for text-independent speaker verification under unseen noise conditions
abstract
This paper describes a novel GMM-UBM based system that deals with the session noise variability problem. The system uses the Type-2 Fuzzy GMM framework by considering the speaker GMM parameters to be uncertain in an interval. The parameters intervals are estimated using a multicondition model training on noisy speeches that are synthesized from the speaker's utterances. Experiments were conducted using the MIT Device Speaker Verification Corpus with utterances having the lowest noise level as training data. The result shows an improvement in the EER of 24.11% for the proposed method compared to the GMM-UBM when evaluated over the noisiest utterances. This shows that the method reduces the effects of the session variability.
Hector N. B. Pinheiro, Sergio R. F. Vieira, Ing Ren Tsang, George D. C. Cavalcanti, Paulo S. G. de Mattos Neto
ICASSP3
2016 Facial expression Recognition based on Motion Estimation
abstract
In this paper, we propose a novel facial expression recognition method based on features of the motion, Facial Expression Recognition based on Motion Estimation (FERME). The proposed approach encodes the directional information of the facial expression. The facial motion is encoded by using the motion estimation between different images from the same (or similar) face. The facial expression image is compared against the most similar image from each facial expression of training database. The best match is obtained using the Structural Similarity Index (SSIM). We propose a modified version of the Adaptive Reduction Search Area algorithm (MARSA) for motion vector calculation. FERME compares the motion vectors to the vectors of the highest occurrences obtained from each facial expression. From this comparison, the Euclidean distances vectors are generated. Support Vector Machine (SVM) is used to classify the facial expression. The experimental results show the effectiveness of the proposed approach.
Hemir da Cunha Santiago, Ing Ren Tsang, George D. C. Cavalcanti
IJCNN2
2016 Class-wise feature extraction technique for multimodal data
Elias Rodrigues da Silva Júnior, George D. C. Cavalcanti, Ing Ren Tsang
Neurocomputing3
2015 Supervised fractional eigenfaces
abstract
Supervised Fractional Eigenfaces (SFE) is an extension of Principal Component Analysis (PCA), which uses the fractional covariance matrix, class label information, and nonlinear data transformation to extract discriminant features. The proposed method combines techniques of two state-of-the-art feature extractors: Fractional Eigenfaces and Dual Supervised PCA. Supervised Fractional Eigenfaces was evaluated in three known face datasets and it achieved significant smaller recognition error.
Tiago Buarque Assunção de Carvalho, Maria A. A. Sibaldo, Ing Ren Tsang, George D. C. Cavalcanti
ICIP4
2015 Efficient 2×2 block-based connected components labeling algorithms
abstract
This paper presents three new efficient 2×2 block-based algorithms for connected components labeling: a two-scan which assigns provisional labels to blocks, a two-scan which assigns provisional labels to pixels and a one-and-a-half-scan which assigns provisional labels to blocks. A new stripe image representation is designed in order to perform the second pass only through the blocks containing some foreground pixel. We also improved the existing 2×2 block-based algorithms by utilizing information of a pixel during a transition in the mask, which allows checking of four neighbor pixels in the mask at most. Thus, the average number of checking operations needed to inspect the neighbor pixels in the first scan is reduced from 1.459 to 1.156, an improvement of 21%. We conducted experiments using synthetic and real images to evaluate the performance of the proposed methods compared to the existing methods. The proposed block-based one-and-a-half-scan algorithm presents the best performance in the real images dataset, which is composed of 1290 documents. Our block-based two-scan algorithm which assigns provisional labels to pixels showed to be the fastest in the synthetic dataset, especially in high density images.
Diêgo J. C. Santiago, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang
ICIP2
2015 Evolutionary Adaptive Self-Generating Prototypes for imbalanced datasets
abstract
The nearest neighbor (NN) is one of the most well known classifiers in pattern recognition. Despite the high classification accuracy, the NN has several drawbacks: high storage requirements, bad time of response, and high noise sensitivity. Prototype Generation (PG) is one of the most well-known solutions to tackle these shortcomings. In supervised classification, many real world datasets do not have an equitable distribution among the different classes, these are called imbalanced datasets. Many PG techniques that have a high classification accuracy in regular datasets, have a poor performance when dealing with imbalanced datasets. The Self-Generating Prototypes (SGP) is one of these techniques. The Adaptive Self-Generating Prototypes was proposed to tackle the SGP problem with imbalanced datasets, but, in doing so, the reduction rate is compromised. This paper proposes the Evolutionary Adaptive Self-Generating Prototypes (EASGP), a SGP based technique with iterative merging and evolutionary pruning to help find the optimal solution. An experimental analysis is performed with datasets of different levels of imbalance ratio and statistical tests are used to evaluate the proposed technique. The results obtained show that EASGP outperforms previous SGP based algorithms in classification accuracy and reduction.
Dayvid V. R. Oliveira, George D. C. Cavalcanti, Ing Ren Tsang, Ricardo Martins de Abreu Silva
IJCNN3
2015 A bootstrap-based iterative selection for ensemble generation
abstract
We propose a bootstrap-based iterative method for generating classifier ensembles called Iterative Classifier Selection Bagging (ICS-Bagging). Each iteration of ICS-Bagging has two phases: i) bootstrap sampling to generate a pool of classifiers; and, ii) selection of the best classifier of the pool using a fitness function based on the ensemble accuracy and diversity. The selected classifier is added to the final ensemble. The bootstrap sampling runs on each iteration and updates the probability of sampling per class based on the class accuracy. This process is repeated until the number of classifiers in the final ensemble is reached. For the specific case of imbalanced datasets, we also propose the SMOTE-ICS-Bagging, a variation of the ICS-Bagging that runs SMOTE at the beginning of each iteration in order to reduce the class imbalance before data sampling. We compared the proposed techniques with Bagging, Random Subspace and SMOTEBagging, using 15 imbalanced datasets from KEEL. The results show the proposed techniques outperform all other techniques in accuracy. Ranking diagrams revealed that the proposed algorithms achieved the highest rankings in accuracy, outperforming SMOTEBagging, a renowned ensemble generation method for imbalanced datasets.
Dayvid V. R. Oliveira, Thyago N. Porpino, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN4
2015 Perceptual video quality assessment for adaptive streaming encoding
abstract
Adaptive video streaming has become prominent due to the rising diversity of Web-enabled personal devices. Common limitations in bandwidth and decoding power challenge the efficiency of content encoders to preserve visual quality at reduced data rates over a wide range of display resolutions. Objective assessment of perceptual video quality has greatly improved in the past decade but remains an open problem. Among the most relevant metrics are the many variations of the Structural Similarity (SSIM) index. In this work, several SSIM-based metrics are compared, optimized and improved towards better correlation with human perception by testing the HD content of the LIVE Mobile Video Quality Database. A shifted gradient is proposed to preserve more image feature information for similarity comparison, thus increasing accuracy, along with a down-sampling box pooling filter that coherently emulates Gaussian pooling while reducing computation complexity by a factor of four and providing broader scalability.
Estevao C. Monteiro, Ricardo E. P. Scholz, Carlos A. G. Ferraz, Ing Ren Tsang, Roberto S. M. Barros
VCIP4
2015 Data-driven global-ranking local feature selection methods for text categorization
Roberto H. W. Pinheiro, George D. C. Cavalcanti, Ing Ren Tsang
Expert Syst. Appl.3
2015 META-DES: A dynamic ensemble selection framework using meta-learning
Rafael M. O. Cruz, Robert Sabourin, George D. C. Cavalcanti, Ing Ren Tsang
Pattern Recognit.4
2014 Fractional Eigenfaces
abstract
The proposed Fractional Eigenfaces method is a feature extraction technique for high dimensional data. It is related to Fractional PCA (FPCA), which is based on the theory of fractional covariance matrix, and it is an extension of the classical Eigenfaces. Like FPCA, it computes projections for a low dimensional space from the fractional covariance matrix and similar to the Eigenfaces, it is suited for high dimensional data. Moreover, the proposed technique extends the fractional transformation of the data for more stages of the feature extractions than FPCA. The Fractional Eigenfaces is evaluated in three different face databases. Results show that it achieves a higher accuracy rate than FPCA and Eigenfaces according to the Wilcoxon hypothesis test.
Tiago Buarque Assunção de Carvalho, Maria A. A. Sibaldo, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang, Jan Sijbers
ICIP3
2014 Type-2 Fuzzy GMMs for Robust Text-Independent Speaker Verification in Noisy Environments
abstract
This paper proposes the use of the type-2 fuzzy GMM (T2FGMM) framework in order to improve the verification rates of the standard GMM-UBM text-independent speaker verification system in noisy environments. Based on type-2 fuzzy sets, the T2FGMM framework describes GMMs with uncertain parameters and provides likelihood intervals for them. The proposed method (T2F-GMM-UBM) estimates the parameter intervals using the noisy speeches from the speakers and the Bayesian estimation used in the standard GMM-UBM system. The proposed method was evaluated using the MIT Device Speaker Verification Corpus (MITDSVC) which contains speeches from 48 speakers recorded in three different locations: a quiet office, a mildly noisy lobby, and a busy street intersection. The Equal Error Rate (EER) was computed for each speaker and the mean and standard deviation were analyzed. Although the proposed method did not achieve better performance in the office location, significant improvements were achieved in both lobby and street intersection locations. The improvement in the lobby was 14.21% while in the street intersection location was 10.47%. The left tailed paired Rank Sign Wilcox on Test was also performed in both locations and the p-values found were 0.0127 and 0.0230, respectively. The proposed method proved to have better performance in noisy environments compared to the standard GMM-UBM system.
Hector N. B. Pinheiro, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang, Jan Sijbers
ICPR2
2014 A modular neural network architecture that selects a different set of features per module
abstract
Modular Neural Network (MNN) divides a problem into smaller and easier sub-problems, and each sub-problem is solved by a neural network called expert. In previous MNN architectures, all experts used the same set of features. This work proposes a modular neural network architecture in which a specialized set of features is selected per expert. As each expert deals with a different sub-problem, it is expected an improvement in the accuracy rate when different and specialized features are selected per expert. The feature selection procedure is an optimization method based on the binary particle swarm optimization. Experimental results over public datasets show that the proposed modular neural network obtains better accuracy rates than literature MNNs.
Diogo da Silva Severo, Everson Verissimo, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN4
2014 Semi-supervised clustering for MR brain image segmentation
Nara M. Portela, George D. C. Cavalcanti, Ing Ren Tsang
Expert Syst. Appl.3
2013 Image deblurring using maps of highlights
abstract
In deblurring an image, we seek to recover the original sharp image. However, without knowledge of the blurring process, we cannot expect to recover the image perfectly. We propose a deblurring method of a single-image where the blur kernel is directly estimated from highlight spots or streaks with high intensity value. These highlighted points can be represented by specular reflection of light that may appear in the eye, on a shiny surface, or in small light sources present in the image. In this work, we detect automatically these high-lighted points in a blurred image. Therefore, creating a map of highlight, which is used as a guide to extract automatically a single highlight from the blurred image. Due to its unique nature, it is demosntrated that the highlighted points are a good estimation of the blur kernel and a sharp image is restored using these kernel. The experimental results show the performance of this method in comparison to several other deblurring methods.
Fabiane Queiroz, Ing Ren Tsang, Lior Shapira, Ron Banner
ICASSP2
2013 Fast block-based algorithms for connected components labeling
abstract
Block-based algorithms are considered the fastest approach to label connected components in binary images. However, the existing algorithms are two-scan which would need more comparisons if they were used as one-and-a-half-scan algorithms. Here, we proposed a new mask that enables the design of a block-based one-and-a-half-scan algorithm without any extra comparison. Furthermore, three new efficient algorithms for connected components labeling are presented: a block-based two-scan, a pixel-based one-and-a-half-scan and a block-based one-and-a-half-scan. We conducted experiments using synthetic and realistic images to evaluate the performance of the proposed methods compared to the existing methods. The proposed block-based one-and-a-half-scan algorithm presents the best performance in the realistic images dataset composed of 1290 documents. Our block-based two-scan algorithm proved to be the fastest in the synthetic dataset, especially in low density images.
Diêgo J. C. Santiago, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang
ICASSP2
2013 Contextual Image Segmentation Based on the Potts Model
abstract
Image segmentation is one of the basic steps in image analysis. Clustering methods are an unsupervised way to provide image segmentation. This paper proposes a clustering algorithm for contextual image segmentation, called spatially variant finite mixture model (SVFMM). For the case of spatially varying mixture of Gaussian density functions with unknown means and variances, an expectation-maximization (EM) algorithm is derived for maximum likelihood estimation of the parameters of the mixture model. In this paper, the Potts model is adopted as a priori density function for the spatially variant mixture proportions to imposes spatial smoothness constraints in the model. Experimental results on a set of different real images show the effectiveness of the proposed method.
Nara M. Portela, George D. C. Cavalcanti, Ing Ren Tsang
ICTAI3
2013 Hybrid Feature Selection and Weighting Method Based on Binary Particle Swarm Optimization
abstract
This work proposes an optimization technique based on binary particle swarm optimization that performs feature selection and feature weighting simultaneously. In the optimization process, each member of the population is described as a vector having three parts: i) one weight per feature (feature weighting), ii) one binary value per feature indicating the presence or the absence of the feature (feature selection), and, iii) the number of neighbors of the kNN classifier. After optimization, this vector is used as a mask to generate a new subset of features that is evaluated using the kNN classifier. The experimental study was performed on public datasets and showed that the proposed technique obtains better accuracy and reduction rates than state-of-the-art techniques.
Diogo da Silva Severo, Everson Verissimo, George D. C. Cavalcanti, Ing Ren Tsang
ICTAI4
2013 Diversity in task decomposition: A strategy for combining mixtures of experts
abstract
The “no free lunch” theorem has stated that learning algorithms cannot be universally good. An alternative to alleviate the weakness of using only one classifier is to combine several of them. Mixture of Experts is a learning algorithm that combines classifiers, in which each classifier or expert is dedicated to solve part of the problem. The partition of the problem is defined by a step called Task Decomposition where the problem is divided in subproblems. This paper proposes an approach to combine mixture of experts, in which different task decomposition methods are used to divide the problem. This strategy aims to increase the diversity of the ensemble, since different task decomposition methods generate different partitions of the database. The experimental study shows that the proposed method obtains better accuracy rates when compared with the traditional mixture of experts.
Everson Verissimo, Diogo da Silva Severo, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN4
2013 A Combined Features Approach for Speaker Segmentation Using BIC and Artificial Neural Networks
abstract
We present a combined features approach for speaker segmentation task. This approach utilizes different acoustic features extracted from audio stream. The Bayesian Information Criterion (BIC) is used for each acoustic feature as a distance measure to verify the merging of two audio segments. An Artificial Neural Network (ANN) combines the time index from each ?BIC with the highest value, and estimates the change point. In the experiments, a data set containing examples with several speakers is used to compare our approach with the Chen and Gopalakrishnan's window-growing-based approach, using different acoustic features sets. The results show an improvement in both the Miss Detection Rate (MDR) and the False Alarm Rate (FAR) compared to the window-growing-based approach.
Leonardo Valeriano Neri, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang, Jan Sijbers
SMC2
2013 Type-2 Fuzzy GMM-UBM for Text-Independent Speaker Verification
abstract
This paper proposes the use of a type-2 fuzzy framework in the standard GMM-UBM based text-independent speaker verification systems. Based on type-2 fuzzy sets, the framework provides pertinence intervals for the models. The decision process is obtained using a Support Vectors Machine (SVM) that processes the interval likelihoods. A Voice Activity Detection (VAD) algorithm was also used to discard the parts of the speech signal without voice. The proposed method was tested on the MIT Device Speaker Verification Corpus which contains several different mobile devices used in different environments. The result shows the robustness of the system and the improvements in the verification ratios of the T2F-GMM-UBM compared to the classical GMM-UBM based systems.
Hector N. B. Pinheiro, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang, Jan Sijbers
SMC2
2013 Motion Compensation Techniques in Permutation-Based Video Encryption
abstract
This paper presents a motion compensation technique applied in the permutation-based digital video encryption and compression method introduced by Socek et al. The encryption method is based in permutations that can improve the spatial correlation on each video frame, making them more compressible by a spatial encoder. However, the compression performance of this method depends on the temporal correlation between consecutive frames and the algorithm does not provide a way to explore non-trivial temporal correlation properly. Consequently, the compression ratio of an encrypted video is very sensitive to the motion in the scene. Here, we propose a motion compensation method to be applied in both encryption and decryption process. Experiments using the H.264 codec show a significant improvement of the compression performance in high motion video sequences.
Caio C. Sabino, Laís Andrade, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang, Jan Sijbers
SMC3
2013 Pedestrian Detection under Progressive Occlusion
abstract
Pedestrian detection is a very promising area in computer vision, since it enables interesting and a variety of applications such as car assistance, surveillance systems and robot vision. During the last years, a variety of new techniques were proposed which greatly improved the detection rates. However, the performance of such systems rapidly deteriorates when pedestrians are under occlusion. This paper analyze how the detection rates of HOG, HOG-LBP, and two new combinations, HOG-LTP and HOG-LMEBP, are affected when occlusion area are progressively added to pedestrian images. Using the INRIA dataset, occlusions were synthetically generated by merging different sizes of non-pedestrian images from different directions. We show that detection of pedestrian under occlusion can be improved by simply combining features.
Silvio G. O. Santos, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang, Jan Sijbers
SMC2
2013 Weighted Modular Image Principal Component Analysis for face recognition
George D. C. Cavalcanti, Ing Ren Tsang, José Francisco Pereira
Expert Syst. Appl.2
2013 ATISA: Adaptive Threshold-based Instance Selection Algorithm
George D. C. Cavalcanti, Ing Ren Tsang, Cesar Lima Pereira
Expert Syst. Appl.2
2013 Feature representation selection based on Classifier Projection Space and Oracle analysis
Rafael M. O. Cruz, George D. C. Cavalcanti, Ing Ren Tsang, Robert Sabourin
Expert Syst. Appl.3
2013 AutoAssociative Pyramidal Neural Network for one class pattern classification with implicit feature extraction
Bruno J. T. Fernandes, George D. C. Cavalcanti, Ing Ren Tsang
Expert Syst. Appl.3
2013 Lateral Inhibition Pyramidal Neural Network for Image Classification
abstract
The human visual system is one of the most fascinating and complex mechanisms of the central nervous system that enables our capacity to see. It is through the visual system that we are able to accomplish from the most simple task such as object recognition to the most complex visual interpretation, understanding and perception. Inspired by this sophisticated system, two models based on the properties of the human visual system are proposed. These models are designed based on the concepts of receptive and inhibitory fields. The first model is a pyramidal neural network with lateral inhibition, called lateral inhibition pyramidal neural network. The second proposed model is a supervised image segmentation system, called segmentation and classification based on receptive fields. This work shows that the combination of these two models is beneficial, and the results obtained are better than that of other state-of-the-art methods.
Bruno J. T. Fernandes, George D. C. Cavalcanti, Ing Ren Tsang
IEEE Trans. Cybern.3
2012 L2-Norm metric learning applied to unconstrained face pair-matching
abstract
This paper proposes a metric learning algorithm based on the L2-Norm (L2ML) in the context of the face pair-matching problem as an attempt to overcome the low discriminatory power of most current descriptors when the operating conditions are unconstrained. The L2ML differs from other similar techniques by giving an efficient closed-form solution to a relatively simple optimization objective. As the experiments show, despite the simplicity, the performance of the proposed method is comparable to that of more complex state of the art techniques. In fact, the combination of only two descriptors in the L2ML space reaches an average accuracy of 84.97% in the challenging Image Restricted benchmark of the aligned Labeled Faces in the Wild (LFW) dataset.
Rafael M. Barreto, Ing Ren Tsang, George D. C. Cavalcanti
ICIP2
2012 Retinal vessel segmentation using Average of Synthetic Exact Filters and Hessian matrix
abstract
The segmentation of blood vessels in retinal images is an important procedure for the prediction and diagnosis of cardiovascular diseases, such as hypertension and diabetes, which are known to affect the retinal blood vessels appearance. This work aims to develop an effective method of retinal vessels segmentation by combining correlation filters and measures extracted from the eigenvalues of the Hessian matrix. The approach uses a threshold to segment the image generated by this combination and is evaluated on two public image databases, Drive and Stare. The results are compared to other state-of-the-art methods described in the literature.
Wendeson S. Oliveira, Ing Ren Tsang, George D. C. Cavalcanti
ICIP2
2012 Image Fusion Combining Frequency Domain Techniques Based on Focus
abstract
Image focus is a property closely related to image quality. In some images it is not possible to obtain a clear focus in all regions simultaneously, so an alternative is to use image fusion to combine pictures with different focus into one with all the best-focused regions. This paper describes two image fusion algorithms in the frequency domain that are based on focus: Contrast in DCT domain and Spatial Frequency. The algorithms divide the images in fixed size blocks to decide which image should be selected to constitute the final result. Improvements are made to both techniques to decide when to choose an entire block or pixels individually. The proposed approach combines the different techniques (or different settings of a single technique), by comparing evaluation metrics (PSNR) values obtained for each block independently and selecting the technique that performs better for the analyzed block. The final image quality, evaluated using PSNR and RMSE, is superior compared to the results of the individual techniques.
Hugo R. Albuquerque, Ing Ren Tsang, George D. C. Cavalcanti
ICTAI2
2012 Data Complexity Measures and Nearest Neighbor Classifiers: A Practical Analysis for Meta-learning
abstract
The classifier accuracy is affected by the properties of the data sets used to train it. Nearest neighbor classifiers are known for being simple and accurate in several domains, but their behavior is strongly dependent on data complexity. On the other hand, there are data complexity measures which aim to describe properties of the data sets. This work aims to show how data complexity measures can be efficiently used to predict the behavior of the Nearest Neighbor classifier. Seven data complexity measures and seventeen real datasets are used in the experimental study. Each data complexity measure is analyzed individually in order to find a relationship between its value and the accuracy of the classifier on a given dataset. No single measure used is good enough to predict the behavior of the Nearest Neighbor classifier. However, the combination of these measures provides a powerful tool to predict the accuracy of the Nearest Neighbor classifier.
George D. C. Cavalcanti, Ing Ren Tsang, Breno A. Vale
ICTAI2
2012 Model Representation for Facial Expression Recognition Based on Shape and Texture
abstract
In this paper, we present an efficient method for facial expression recognition. Three features extraction methods are combined to form a model representation for facial expressions. Once the feature and the model representation are defined a Support Vector Machine (SVM) is used for the classification task. The proposed method is tested using the Yale and Cohn-Kanade databases, which contains 165 images and 1480 images, respectively. The method presented a recognition rate of 98.1% and 93% for the Yale and Cohn-Kanade respectively.
Adriana Cruz de Gois, Victor Oliveira Antonino, Ing Ren Tsang, George D. C. Cavalcanti
ICTAI3
2012 Improved Self-Generating Prototypes Algorithm for Imbalanced Datasets
abstract
Some real world datasets have different proportions of classes, too many instances of the majority classes and only a few of the minority classes, those are called imbalanced datasets. Many applications, like medical diagnosis and risk analysis, are interested in the under-represented class, but classifiers and prototype generation techniques usually have a bias towards the majority classes. Because of that, the problem of classification with imbalanced datasets has become an important topic in Pattern Recognition. The Self-Generating Prototypes (SGP) have a high reduction power and an excellent performance with balanced datasets, but, with imbalanced datasets, the generated prototypes do not have a good representation of the training dataset. This algorithm generates many prototypes of the majority classes and only a few, or even none, of the minority classes. The aim of this paper is to propose the Adaptive Self-Generating Prototypes (ASGP), an improvement of the SGP2, the second version of the SGP, designed to handle imbalanced datasets. This paper also exposes the reasons for the low performance of the SGP2 with such datasets. Empirical results show that the ASGP has a higher performance with imbalanced datasets than the SGP2, especially when it comes to classification accuracy of the minority classes.
Dayvid V. R. Oliveira, Guilherme R. Magalhaes, George D. C. Cavalcanti, Ing Ren Tsang
ICTAI4
2012 An Unsupervised Segmentation Method for Retinal Vessel Using Combined Filters
abstract
Image segmentation of retinal blood vessels is an important procedure for the prediction and diagnosis of cardiovascular related diseases, such as hypertension and diabetes, which are known to affect the retinal blood vessels appearance. This work develops an unsupervised segmentation procedure for the segmentation of retinal vessels images using a combined matched filter, Frangi filter and Gabor Wavelet Filter. After the vessel enhancement, two segmentation methods are tested. The first method uses an approach based on deformable models and the second uses fuzzy C-means for the image segmentation. The procedure is evaluated using two public image databases, Drive and Stare. The results are compared to other state-of-the-art methods described in the literature.
Wendeson S. Oliveira, Ing Ren Tsang, George D. C. Cavalcanti
ICTAI2
2012 Class-Dependent Locality Preserving Projections for Multimodal Scenarios
abstract
This paper proposes a method for linear feature extraction called Class-dependent Locality Preserving Projections. It is a supervised extension of the Locality Preserving Projection algorithm and it aims to work in scenarios with within-class multimodality, which are those scenarios where the scattering of the patterns follows more than one modal distribution. Differently from the classical feature extraction techniques that build their solutions based on the whole dataset, the Class-dependent Locality Preserving Projections looks at each class separately, building a specific projection for each class. The proposed technique analyses a query pattern based on the output of each class and chooses the class that better fit the pattern. The experimental study shows that the Class-dependent Locality Preserving Projections is a feature extraction technique for general purposes, however, it is particularly well succeed when applied to within-class multimodal scenarios.
Elias Rodrigues da Silva Júnior, George D. C. Cavalcanti, Ing Ren Tsang
ICTAI3
2012 A Dimensionality Reduction Approach for Modular Neural Networks
abstract
A modular neural network architecture is composed by independent neural networks that focus on different parts of the whole task. This work proposes the Intrinsic Modular Neural Networks that aims not only to reduce the number of classes and patterns in each independent neural network, but also to reduce the dimensionality of the data. The task decomposition is performed by the High-Dimensional Data Clustering algorithm. After the clustering, the training patterns are divided in groups and each group is used to train an independent neural network. Experiments on public databases show promising results.
Everson Verissimo, Diogo da Silva Severo, George D. C. Cavalcanti, Ing Ren Tsang
ICTAI4
2012 Iris Segmentation and Recognition Using 2D Log-Gabor Filters
Carlos A. C. M. Bastos, Ing Ren Tsang, George D. C. Cavalcanti
IDEAL2
2012 Face Detection under Illumination Variance Using Combined AdaBoost and Gradientfaces
João Paulo Magalhães, Ing Ren Tsang, George D. C. Cavalcanti
IDEAL2
2012 Real-Time Head Pose Estimation for Mobile Devices
Euclides N. Arcoverde Neto, Rafael M. Barreto, Rafael M. Duarte, João Paulo Magalhães, Carlos A. C. M. Bastos, Ing Ren Tsang, George D. C. Cavalcanti
IDEAL6
2012 A Hybrid GMM Speaker Verification System for Mobile Devices in Variable Environments
Ing Ren Tsang, George D. C. Cavalcanti, Dimas Gabriel, Hector N. B. Pinheiro
IDEAL1
2012 A Neural Network Based Approach for GPCR Protein Prediction Using Pattern Discovery
Ing Ren Tsang, George D. C. Cavalcanti, Francisco Nascimento Junior, Gabriela Espadas
IDEAL1
2012 A fingerprint spoof detection based on MLP and SVM
abstract
We introduce a fingerprint spoof detection technique based on MLP and SVM that combines several features. The proposed technique is evaluated on two scenarios: (i) when an impostor can perform consecutive attempts to be considered authentic; and, (ii) when the system deals with fingerprints from elderly people. In order to analyze these scenarios, a database was developed. The results show that the proposed combination of features increases the system performance in at least 33.56% and that the average error increases as more attempts for acceptance are allowed. The SVM classifier presents better performance in almost all the tested configurations. However, MLP is more accurate with biometrics from elderly people.
Luis Filipe A. Pereira, Hector N. B. Pinheiro, Jose Ivson S. Silva, Anderson G. Silva, Thais M. L. Pina, George D. C. Cavalcanti, Ing Ren Tsang, Joao Paulo Nogueira de Oliveir
IJCNN7
2012 Pupil segmentation using Pulling & Pushing and BSOM neural network
abstract
Segmentation is a preliminary step for many computer vision systems. Several segmentation algorithms have been developed for different tasks. Here, we are interested in the pupil segmentation, an important procedure in iris recognition systems. In most of the pupil segmentation algorithms it is assumed that the pupil has a circular shape. These methods inaccurate identify pupil borders that do not have a circular shape. In iris recognition, the error caused by an imprecise segmentation can lead to poor recognition rates. In this paper we propose a new method for pupil segmentation based on the Pulling & Pushing method and a batch-SOM neural network in order to improve the segmentation. We tested the proposed method in the MMU1 and Casia V3 iris databases, obtaining accurate results.
Carlos A. C. M. Bastos, Ing Ren Tsang, Gabriel S. Vasconcelos, George D. C. Cavalcanti
SMC2
2012 Neighborhood coding for image representation and neighborhood operations
abstract
Neighborhood coding is a binary image representation method that has been performing successfully for a variety of applications such as features extraction for image recognition, shape descriptor, and image compression. Despite the success of this coding method, the representation lacked a formal notation. Here, we proposed a formal mathematical notation to represent any binary image into a neighborhood coding scheme. Using this representation we also introduce the concept of neighborhood operations, a procedure akin to mathematical morphology, however having a lower computational cost.
Tiago Buarque Assunção de Carvalho, Maria A. A. Sibaldo, Denise Jaeger Tenório, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang
SMC4
2012 MLPBoost: A combined AdaBoost / multi-layer perceptron network approach for face detection
abstract
Face detection is a research area in computer vision of great interest. Even though several different methods have been developed, improvements can still be made in the false-positive detection and increase in the speed of the detector. In this work, we investigate the AdaBoost technique as an artificial neural network. We propose a new model called MLPBoost, which is an hybridization between AdaBoost and Multi-Layer Perceptron (MLP) networks. This algorithm has shown improvements in the performance of classifiers already trained with AdaBoost, either by the increase in the detection rate and the reduction of false positive rates, or by decreasing the processing time of these classifiers.
George D. C. Cavalcanti, João Paulo Magalhães, Rafael M. Barreto, Ing Ren Tsang
SMC4
2012 A modular architecture based on image quality for fingerprint spoof detection
abstract
This work proposes an improvement for fingerprint spoof detection in order to reduce the occurrence of live fingerprints from elderly people taken as spoof. The novel architecture combines classifiers working independently into two distinct image quality groups. The results show that high quality spoof images are easily detected, the misclassification rate for live fingerprints is reduced by 64.0% and the overall system performance increases 49.61%.
George D. C. Cavalcanti, Luis Filipe A. Pereira, Hector N. B. Pinheiro, Jose Ivson S. Silva, Anderson G. Silva, Thais M. L. Pina, Daniel B. O. Carvalho, Ing Ren Tsang
SMC8
2012 Recognition of partially occluded face using Gradientface and Local Binary Patterns
abstract
Currently one of the most important challenges of face recognition systems is the problem of occlusion, which is quite common in real applications. There are several studies in the literature treating this problem, but no defined or robust solution is agreed. The focus of this work is to develop face recognition method with sunglasses and scarf occlusion. We propose a robust approach which consists in detecting the face region that does not have occlusion and uses this region to obtain the recognition. To classify the occluded and non-occluded parts, a Multi-Layer Perceptron (MLP) is applied. While for the recognition a combined Gradientface and Local Binary Pattern (LBP) are used. Gradientface is applied to address the variation in the illumination of the image. Experiments are shown using the AR Face and ORL databases.
George D. C. Cavalcanti, Ing Ren Tsang, Josivan R. Reis
SMC2
2012 Speaker verification using type-2 Fuzzy Gaussian Mixture Models
abstract
This paper proposes the use of a Type-2 Fuzzy GMM (T2FGMM) speaker verification system for mobile devices in variable environments. This model is an extension of Gaussian mixture models based on the type-2 fuzzy set, which provide pertinence intervals for the trained samples. The decision process is obtained using the Generalized Linear Model (GLM) that processes the interval likelihoods. A Voice Activity Detection (VAD) algorithm was also used to improve the speaker verification ratio. The proposed method was tested on the MIT mobile device speaker verification database which contains several different mobile devices used in different environments. The result shows the robustness of the system and the improvements in the verification ratios of the T2GMM over the classical GMM.
Ing Ren Tsang, Dimas Gabriel, Hector N. B. Pinheiro, George D. C. Cavalcanti
SMC1
2012 Video colortoning
abstract
The growing amount of information transferred and stored in graphical and video format, and the forthcoming of new devices that have a limited hardware capabilities to either store data or display shades (such as electronic paper), creates a demand for higher computational resources or new techniques that consume less of these resources. In this paper is proposed a new method for video coding called Video Colortoning, which reduces the size of the image representation in a sequence of color image, with a small loss of quality and in real time. The proposed method treats the problem of quantization noise and the flicker effect, optimizes the data for further compression and considers the human visual system regarding the handling of colors.
Ing Ren Tsang, Diogo C. Lemos, Dario S. M. Pinheiro, George D. C. Cavalcanti, Ing Jyh Tsang
SMC1
2012 Combined AdaBoost and gradientfaces for face detection under illumination problems
abstract
Regardless of several different methods for face detection have been developed in the last years, there are still situations that requires more improvements especially in issues related to variations in illumination and face occlusion. Illumination problems are normally handled by using preprocessing, and model or training-based approaches. We propose here a face detection method combining the well-known AdaBoost with Gradientfaces following a model-based approach, which was not yet used for the face detection problem. We have applied Gradientfaces before training an AdaBoost Haar-based cascade classifier to overcome the problem of strong variations in illumination. Cited approaches were evaluated first in a data set containing artificial and then real illumination problems. Experiments show that proposed method is stable when facing different lighting conditions, and better than others when dealing with strong and uncontrolled illumination problems.
Ing Ren Tsang, João Paulo Magalhães, George D. C. Cavalcanti
SMC1
2012 A global-ranking local feature selection method for text categorization
Roberto H. W. Pinheiro, George D. C. Cavalcanti, Renato Fernandes Corrêa, Ing Ren Tsang
Expert Syst. Appl.4
2011 Fuzzy Active Contour Models
abstract
This paper presents a Fuzzy Active Contour Model for image segmentation using three variations. The proposed models are based on the Fuzzy Energy-Based Active Contour model introduced by Krinidis and Chatzis. First, an update criteria that changes only localized membership values at each iteration is introduced. Second, the model is extended to a type-2 fuzzy logic. And finally, a multiple object segmentation schema is applied to the original model. We present some experimental results, showing the performance for each modification and some of its advantages.
Cesar Lima Pereira, Carlos A. C. M. Bastos, Ing Ren Tsang, George D. C. Cavalcanti
FUZZ-IEEE3
2011 A robust feature extraction algorithm based on class-Modular Image Principal Component Analysis for face verification
abstract
Face verification systems reach good performance on ideal environmental conditions. Conversely, they are very sensitive to non-controlled environments. This work proposes the class-Modular Image Principal Component Analysis (cMIMPCA) algorithm for face verification. It extracts local and global information of the user faces aiming to reduce the effects caused by illumination, facial expression and head pose changes. Experimental results performed over three well-known face databases showed that cMIMPCA obtains promising results for the face verification task.
José Francisco Pereira, Rafael M. Barreto, George D. C. Cavalcanti, Ing Ren Tsang
ICASSP4
2011 A weighted image reconstruction based on PCA for pedestrian detection
abstract
Pedestrian detection is a task usually associated with security and surveillance systems. The development of a pedestrian detection system poses a hard challenge, because of its inherently complex nature. In this work, we present an analysis of an existing pedestrian detection model based on PCA reconstruction errors. We investigate how the method works and where changes can be made to improve its original performance. The proposed improvements enhance the system's accuracy by using weights, that are found in an automated way using a genetic algorithm. We also found that some reconstruction errors used by the original method are not strictly necessary and therefore they can be eliminated to reduce the classifying time by half.
Guilherme V. Carvalho, Lailson B. Moraes, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN4
2011 A method for dynamic ensemble selection based on a filter and an adaptive distance to improve the quality of the regions of competence
abstract
Dynamic classifier selection systems aim to select a group of classifiers that is most adequate for a specific query pattern. This is done by defining a region around the query pattern and analyzing the competence of the classifiers in this region. However, the regions are often surrounded by noise which can difficult the classifier selection. This fact makes the performance of most dynamic selection systems no better than static selections. In this paper we demonstrate that the performance of dynamic selection systems end up limited by the quality of the regions extracted. Thereafter, we propose a new dynamic classifier selection system that improves the regions of competence in order to achieve higher recognition rates. Results obtained from several classification databased show the proposed method not only significantly increase the recognition performance, but also decreases the computational cost.
Rafael M. O. Cruz, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN3
2011 Autoassociative Pyramidal Neural Network for face verification
abstract
In this paper, the face verification problem is addressed. A neural network with autoassociation memory and receptive fields based architecture is proposed. It is called AAPNet (AutoAssociative Pyramidal Neural Network). The proposed neural network integrates feature extraction and image reconstruction in the same structure. For a given recognition task, at least one instance of the AAPNet must be trained for each known class. Thus, the AAPNet outputs how similar is a given probe image to its class. The AAPNet is applied in a face verification task using thumbnail-sized faces and achieves better results when compared to state-of-the-art models.
Bruno J. T. Fernandes, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN3
2011 GA-PAT-KNN: Framework for time series forecasting
abstract
A novel framework for time series prediction that integrates Genetic Algorithm (GA), Partial Axis Search Tree (PAT) and K-Nearest Neighbors (KNN) is proposed. This methodology is based on the information obtained from Technical analysis of a stock. Experiments have shown that GAs can capture the most relevant variables and improve the accuracy of predicting the direction of daily change in a stock price index. A comparison with other models shows the advantage of the proposed framework.
Armando A. Gonçalves, Igor Alencar, Ing Ren Tsang, George D. C. Cavalcanti
IJCNN3
2011 Lag selection for time series forecasting using Particle Swarm Optimization
abstract
The time series forecasting is an useful application for many areas of knowledge such as biology, economics, climatology, biology, among others. A very important step for time series prediction is the correct selection of the past observations (lags). This paper uses a new algorithm based in swarm of particles to feature selection on time series, the algorithm used was Frankenstein's Particle Swarm Optimization (FPSO). Many forms of filters and wrappers were proposed to feature selection, but these approaches have their limitations in relation to properties of the data set, such as size and whether they are linear or not. Optimization algorithms, such as FPSO, make no assumption about the data and converge faster. Hence, the FPSO may to find a good set of lags for time series forecasting and produce most accurate forecastings. Two prediction models were used: Multilayer Perceptron neural network (MLP) and Support Vector Regression (SVR). The results show that the approach improved previous results and that the forecasting using SVR produced best results, moreover its showed that the feature selection with FPSO was better than the features selection with original Particle Swarm Optimization.
Gustavo H. T. Ribeiro, Paulo S. G. de Mattos Neto, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN4
2011 BSOM network for pupil segmentation
abstract
Segmentation is a preliminary step in many computer vision systems. In most of pupil segmentation algorithms it is assumed that the pupil has a predefined shape, usually circular. This parametrization might lead to errors when the eye image is distorted or deformed and when the pupil is partially occluded by eyelids or eyelashes. In this work, we propose a new method for pupil segmentation based on a batch-SOM (BSOM) neural network composed by three steps: (1) definition of the initial neurons position; (2) use BSOM to extract the contour; and (3) perform a contour adjustment. The method is capable of finding the pupil contour in a flexible manner, independently of a predefined shape. We modified the BSOM algorithm in three points: (1) in the update process, introducing the neighborhood constraint; (2) removal of the neurons, and (3) in the convergence criteria. Experiments were performed using Casia-IrisV3 Interval, Casia-IrisV4 Syn, and MMU1 iris image databases.
Gabriel S. Vasconcelos, Carlos A. C. M. Bastos, Ing Ren Tsang, George D. C. Cavalcanti
IJCNN3
2011 DifFocus: An approach for image segmentation
abstract
A new multi-focus analysis method for image segmentation, based on human visual system focal attention and focusing phenomena, is presented. The main goal is to develop a fast and accurate segmentation algorithm, suitable for real-time applications, such as robotics and autonomous vehicles. The difFocus shows to be fast enough and it achieved the 8th place in the Berkeley Segmentation Benchmark global ranking for both color and grayscale image segmentation. It also shows to be robust to parameter selection, achieving promising results.
Diogo C. Costa, Ing Ren Tsang, Carlos A. B. Mello
SMC2
2010 A graph-based friend recommendation system using Genetic Algorithm
abstract
A social network is composed by communities of individuals or organizations that are connected by a common interest. Online social networking sites like Twitter, Facebook and Orkut are among the most visited sites in the Internet. Presently, there is a great interest in trying to understand the complexities of this type of network from both theoretical and applied point of view. The understanding of these social network graphs is important to improve the current social network systems, and also to develop new applications. Here, we propose a friend recommendation system for social network based on the topology of the network graphs. The topology of network that connects a user to his friends is examined and a local social network called Oro-Aro is used in the experiments. We developed an algorithm that analyses the sub-graph composed by a user and all the others connected people separately by three degree of separation. However, only users separated by two degree of separation are candidates to be suggested as a friend. The algorithm uses the patterns defined by their connections to find those users who have similar behavior as the root user. The recommendation mechanism was developed based on the characterization and analyses of the network formed by the user's friends and friends-of-friends (FOF).
Nitai B. Silva, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang
IEEE Congress on Evolutionary Computation2
2010 A combined Pulling & pushing and Active Contour method for pupil segmentation
abstract
Pupil segmentation is usually the first step used for searching iris regions. Iris localization is an extremely important procedure in iris biometrics systems, since the correct segmentation of inner and outer boundaries is critical to achieve high recognition rates. An iris localization method based on a spring force-driven iterative scheme, called Pulling & Pushing have been proposed by He et al. 2006. Here, we propose a pupil segmentation procedure that combines Pulling & Pushing and Active Contour Models, overcoming and improving the results of the previous method. We also developed a new strategy to identify and fill reflection points that appear inside the pupil. We tested our method in MMU1 and Casia V1 and V3 iris databases, obtaining accurate results.
Carlos A. C. M. Bastos, Ing Ren Tsang, George D. C. Cavalcanti
ICASSP2
2010 Neighborhood coding for bilevel image compression and shape recognition
abstract
Neighborhood coding was proposed to encode binary images. Previously, this coding scheme presented good results in the problem of handwritten character recognition. In this article, we extended this coding scheme so that it can be applied as an image shape descriptor and in a bilevel image compression method. An algorithm to reduce the number of codes needed to reconstruct the image without loss of information is presented. Using the exactly same set of reduced codes, a lossless compression method and a shape recognition system are proposed. The reduced codes are used with Huffman coding and RLE (Run-Length Encoding) to obtain a compression rate comparable to well-known image compression algorithms such as LZW and CCITT Group 4. For the shape recognition task we applied a template matching algorithm to the set of strings generated by the coding reduction procedure. We tested this method in the MPEG-7 Core Experiment Shape 1 part A2 and the binary image compression challenge database.
Tiago Buarque Assunção de Carvalho, Denise Jaeger Tenório, Ing Ren Tsang, George D. C. Cavalcanti, Ing Jyh Tsang
ICASSP3
2010 Analysis of 2D log-Gabor Filters to Encode Iris Patterns
abstract
This paper presents an analysis of the parameters used to construct 2D log-Gabor filters to encode iris patterns. This filter is a band-pass complex filter composed by four parameters that are used to extract information direct in the 2D domain. An iris recognition system, composed by segmentation, normalization, encoding and matching is also described. The segmentation module combines the Pulling & Pushing and Active Contour Model and the Circular Hough Transform to find the inner and the outer boundaries of the iris. The experiments were performed using the CASIA v1 iris database and the results are analyzed using ROC curves. We conclude that 2D log-Gabor filters are also an effective alternative to encode the features present on iris patterns. The combination of the modified segmentation procedure and the use of 2D Log-Gabor filters showed good results for certain important regions of the ROC curves.
Carlos A. C. M. Bastos, Ing Ren Tsang, George D. C. Cavalcanti
ICTAI (2)2
2010 A New Heterogeneous Dissimilarity Measure for Data Classification
abstract
Instance-based learning algorithms typically suffer influences of dissimilarity functions. The problem is frequently related to the Nearest Neighbor rules of these algorithms. This paper will introduce a new dissimilarity measure, called Heterogeneous Centered Difference Measure, which is tested over many known databases. The results are compared with other distance functions.
Cesar Lima Pereira, George D. C. Cavalcanti, Ing Ren Tsang
ICTAI (2)3
2010 Off-line Signature Verification: An Approach Based on Combining Distances and One-class Classifiers
abstract
This paper presents an off-line signature verification system composed of a combination of several different classifiers. Identity authentication is a very important characteristics specially in systems that requires a high degree of security such as in bank transactions. In our experiments, one-class classifier was used to create a signature verification system, consequently only genuine signatures were necessary for the training phase. We proposed five distances measurement as features for the classification system. The distances extracted from the signature database were: furthest, nearest, template, central and ncentral. Also, a normalization procedure was applied to turn the distance scale invariant. These distances were combined using four operation: product, mean, maximum and minimum. The calculated distances were used as a feature vector to represent the signatures. Finally, the distances measurement and their combinations were used as input vector for different classifiers. The proposed signature verification method obtained very good rates.
Milena R. P. Souza, George D. C. Cavalcanti, Ing Ren Tsang
ICTAI (1)3
2010 An ensemble classifier for offline cursive character recognition using multiple feature extraction techniques
abstract
This paper presents a novel approach for cursive character recognition by using multiple feature extraction algorithms and a classifier ensemble. Several feature extraction techniques, using different approaches, are extracted and evaluated. Two techniques, Modified Edge Maps and Multi Zoning, are proposed. The former one presents the best overall result. Based on the results, a combination of the feature sets is proposed in order to achieve high recognition performance. This combination is motivated by the observation that the feature sets are both, independent and complementary. The ensemble is performed by combining the outputs generated by the classifier in each feature set separately. Both fixed and trained combination rules are evaluated using the C-Cube database. A trained combination scheme using a MLP network as combiner achieves the best results which is also the best results for the C-Cube database by a good margin.
Rafael M. O. Cruz, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN3
2010 Improving financial time series prediction using exogenous series and neural networks committees
abstract
Time series forecasting is useful in many researches areas. The use of models that provide a reliable prediction in financial time series may bring valuable profits for the investors. This paper proposes a methodology based on information obtained from exogenous series used in combination with neural networks to predict stock series. The best trained neural networks were used in combination to improve the prediction capacity of a single networks. To evaluate the proposed prediction models, some known metrics were applied. Moreover, we also proposed one new metric called Prediction in Direction and Accuracy (PDA), which benefits models with great performance in prediction accuracy and trend. Addictionally, there was used an evolutionary algorithm to choose the best trained models that maximize PDA. Experiments with two of the most important Brazilian companies stock quotes have shown the usefulness of the proposed prediction system to generate profits in investments.
Manoel C. Amorim Neto, Gustavo Tavares, Victor Medeiros Outtes Alves, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN5
2010 Does the affinity matrix influence the performance of the Locality Preserving Projection algorithm?
abstract
Classical feature extraction techniques, like PCA and LDA, do not deal properly with multimodal problems. Such techniques create projections that do not preserve the multimodal structure of the original data distribution. Locality Preserving Projection (LPP) is a feature extraction technique which looks for a transformation matrix that minimizes the changes into the structure of the data after the transformation. This local structure is captured by the affinity matrix. However, there many ways to calculate this affinity matrix. The main aim of this paper is to evaluate the influence of different affinity matrices over the LPP accuracy. The experiments showed that the correct choice of the affinity matrix can lead to a performance gain. Among the analyzed affinity matrices, Local Scaling and Nearest Neighbor reached the best results.
Elias Rodrigues da Silva Júnior, George D. C. Cavalcanti, Ing Ren Tsang
SMC3
2009 Text Line Segmentation Based on Morphology and Histogram Projection
abstract
Text extraction is an important phase in document recognition systems. In order to segment text from a page document it is necessary to detect all the possible manuscript text regions. In this article we propose an efficient algorithm to segment handwritten text lines. The text line algorithm uses a morphological operator to obtain the features of the images. Following, a sequence of histogram projection and recovery is proposed to obtain the line segmented region of the text. First, an Y histogram projection is performed which results in the text lines positions. To divide the lines in different regions a threshold is applied. After that, another threshold is used to eliminate false lines. These procedures, however, cause some loss on the text line area. So, a recovery method is proposed to minimize this effect. In order to detect the extreme positions of the text in the horizontal direction, an X histogram projection is applied. Then, as in the Y direction, another threshold is used to eliminate false words. Finally, in order to optimize the area of the manuscript text line, a text selection is carried out. Experimental results using the IAM-database showed that this new approach is robust, fast and produces very good score rates.
Rodolfo P. dos Santos, Gabriela S. Clemente, Ing Ren Tsang, George D. C. Cavalcanti
ICDAR3
2009 A receptive field based approach for face detection
abstract
This paper presents a new neural network to perform the visual pattern classification task. The neural network is calledI-PyraNetwhich is a hybrid implementation of the PyraNet and the concepts of the inhibitory fields. In order to improve the results obtained by this neural network, it is also presented the 2-D Gabor filter. Furthermore, both, the neural network and the filter, are applied over a face detection task and are compared to the results obtained by a SVM.
Bruno J. T. Fernandes, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN3
2009 Financial time series prediction using exogenous series and combined neural networks
abstract
Time series forecasting have been a subject of interest in several different areas of research such as: meteorology, demography, health, computer and finance. Since it can be applied to various practical problems in real world, techniques to predict time series have been a topic of increasing research activities, especially in the financial sector that has a great interest in the forecast of the stock market. In this article, we are interested in the forecast of the time series related to the Brazilian oil company, Petrobras (PETR4). A methodology based on information obtained from exogenous series was used in combination with a neural network to predict the PETR4 stock series. Exogenous series were selected by analyzing the correlation between the series with the Petrobras stocks series. In this way, the prediction was obtained by not just using the previous values of the series but also by using information external to the PETR4 series. The values of the selected series were used as features for a prediction stage based on combined neural networks. To evaluate the performance of the system classical measurements were used, however we also introduce a new performance index called Sum of the Losses and Gains (SLG).
Manoel C. Amorim Neto, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN3
2009 A data mining approach to solve the goal scoring problem
abstract
In soccer, scoring goals is a fundamental objective which depends on many conditions and constraints. Considering the RoboCup soccer 2D-simulator, this paper presents a data mining-based decision system to identify the best time and direction to kick the ball towards the goal to maximize the overall chances of scoring during a simulated soccer match. Following the CRISP-DM methodology, data for modeling were extracted from matches of major international tournaments (10691 kicks), knowledge about soccer was embedded via transformation of variables and a Multilayer Perceptron was used to estimate the scoring chance. Experimental performance assessment to compare this approach against previous LDA-based approach was conducted from 100 matches. Several statistical metrics were used to analyze the performance of the system and the results showed an increase of 7.7% in the number of kicks, producing an overall increase of 78% in the number of goals scored.
Renato Oliveira, Paulo J. L. Adeodato, Arthur Carvalho, Icamaan Viegas, Christian Diego, Ing Ren Tsang
IJCNN6
2009 Modular Image Principal Component Analysis for face recognition
abstract
One of the most successful process to accomplish human face recognition are the methods based on the principal component analysis (PCA), also known as eigenfaces. Recently, novel PCA approaches have been proposed: modular (MPCA) and two-dimensional (IMPCA). These approaches have achieved outstanding result in feature extraction and recognition. IMPCA is used for feature extraction based on 2D matrix representation and MPCA is based on image division to improve face recognition with variations like facial expressions, light and head pose. In this work we use some aspects of these methods to build a new technique called modular Image PCA (MIMPCA). The results achieved with the proposed method are superior in all experiments compared with the original techniques under different conditions of head pose angle, illumination and facial expression.
José Francisco Pereira, George D. C. Cavalcanti, Ing Ren Tsang
IJCNN3
2008 Classification and Segmentation of Visual Patterns Based on Receptive and Inhibitory Fields
abstract
This paper presents a new model to realize a supervised image segmentation task. It is based on the concept of receptive fields that intends to analyze pieces of an image considering not only the pixels or group of them, but also the relationship between them and their neighbors, called segmentation and classification with receptive fields (SCRF). Also, in order to work with the SCRF model, is proposed here a new artificial neural network, called IPyraNet, which is a hybrid implementation of the recently described PyraNet and the nonclassical receptive fields inhibition. Furthermore, the model and the network are applied together in order to realize a satellite image segmentation task.
Bruno J. T. Fernandes, George D. C. Cavalcanti, Ing Ren Tsang
HIS3
2008 A SVM for GPCR Protein Prediction Using Pattern Discovery
abstract
Machines learning techniques have been applied in several different problems in bioinformatics. Similarly, pattern discovery algorithms have also been used to uncover hidden motifs in protein sequences, contributing greatly to the understanding of the problem of protein classification. G-protein coupled receptors (GPCRs) represent one of the largest protein families in Human Genome. Most of these receptors are major target for drug discovery and development. Therefore, they are of interest to the pharmaceutical industry. The technique used in this paper combine machine learning and pattern discovery methods to develop a protein prediction procedure in relation to its functional class, more specifically to predict GPCR protein class. Vilo[2]proposed an algorithm in order to extract pattern of regular expressions from known protein GPCR sequences and used them to predict coupling specificity of G protein coupled receptors to their G proteins. We analyze these patterns and combine them as features for feeding a SVM to predict the GPCR super class. We demonstrate the results using ROC curves, which are well-indicated to evaluate the performance of this kind of classifiers. The experiments, based on the GPCRDB database, also showed that we were able to find some novel GPCR sequences that were not described in the PROSITE database.
Francisco Nascimento Junior, Ing Ren Tsang, George D. C. Cavalcanti
HIS2
2007 Analysis of mammogram using self-organizing neural networks based on spatial isomorphism
abstract
The correct segmentation and measurement of mammography images is of fundamental importance for the development of automatic or computer-aided cancer detection systems. In this paper we propose a method to segment mammogram image using a self-organizing neural network based on spatial isomorphism. The method used is a modified version of the algorithm proposed by Venkatesh and Rishikesh [1] to extract object boundaries in an image. This model explores the principle of spatial isomorphism and self-organization in order to create flexible contours that characterize shapes in images. We modified the original algorithm to overcame problems of local minimum, poor performance for image object with large concavity and imprecise results when simple or far from object border contour are chosen. A comparison of both algorithm and original segmentation used by the MIAS database [9] is presented.
Aida Araujo Ferreira, Francisco Nascimento Jr., Ing Ren Tsang, George D. C. Cavalcanti, Teresa Bernarda Ludermir, Ronaldo Ribeiro Barbosa de Aquino
IJCNN3
2006 Neighbourhood Vector as Shape Parameter for Pattern Recognition
abstract
We present a neighbourhood vector representation as shape parameter for binary images. This method is based on the pixel neighbourhood relation. Each pixel is transformed into a vector, V = (n, e, s, w), where each element of the vector represents the total number of neighbour pixels in the respective direction, north, east, south, west. A binary object is represented by a set of neighbourhood vectors (NV), in which the information of the shape structure is retained. The k-means and fuzzy c-means clustering methods are used to reduce the total amount of NV and the probability distribution of the reduced NV is used to characterize a class of image. We applied this method for handwritten character recognition, using neural network as classifiers. The results show that the shape parameter can be used as a general method of feature extraction for problems in image processing and pattern recognition. In addition, we present an application of this representation scheme for the neighbourhood image operator.
Ing Ren Tsang, Ing Jyh Tsang
IJCNN1
1999 Image coding using neighbourhood relations
Ing Jyh Tsang, Ing Ren Tsang, Dirk Van Dyck
Pattern Recognit. Lett.2
1998 Handwritten Character Recognition based on Moment Features Derived from Image Partition
abstract
In this work we present a novel approach to handwritten character recognition which is based on the intuitive way in which characters are written as one or a few continuous lines. Therefore we calculate the zeroth, first and second radial moment as a function of the angle. In practice this is done by dividing the character into 32 angular sections. The three obtained curves can be used for pattern recognition using statistical analysis. The method has been evaluated using the NIST handwritten character data set. At first, a simple chi-square test gave a result of 80.81% recognition rate at zero rejection rate for digits. Using a back-propagation algorithm the recognition rate obtained was 87.54% also at zero rejection rate, showing that the features are sufficient to discriminate the characters.
Ing Jyh Tsang, Ing Ren Tsang, Dirk Van Dyck
ICIP (2)2