VLDB 2026 Research / reviewers in the wild / expert
João Paulo Pordeus Gomes
dblp:163/4376 · also João Gomes 0002, João P. P. Gomes
· DBLP profile ↗
50ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0003-1686-595XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorSystems, architecture and hardware · 3 · 1 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Long-Term HDD RUL Prediction Using a Novel Feature Engineering ApproachabstractHard Disk Drive (HDD) failures are a persistent and costly problem in data centers, making the accurate prediction of Remaining Useful Life (RUL) essential for mitigating downtime and data loss. While machine learning models, particularly LSTMs, have been widely applied to this task using Self-Monitoring, Analysis, and Reporting Technology (SMART) data, their effectiveness is largely confined to short-term predictions (e.g., 30 days), lacking the ability to accurately forecast RUL over longer horizons, such as one year. To address these limitations, this paper proposes a novel methodology. First, we create a Binary Classification Model that discriminates input data as healthy or faulty. Then, we use its output to generate failure likelihood features to append to the original dataset. We evaluate our approach on a real-world dataset from Backblaze, comparing its performance against baseline models to demonstrate that incorporating the auxiliary features improves the MSE score of the RUL predictions up to 38%. Francisco L. F. Pereira, Victor A. E. de Farias, Felipe T. Brito, João Paulo Pordeus Gomes, Javam C. Machado |
ICMLA | 4 |
| 2025 | Minimal learning machine for multi-label learningabstractAbstract Distance-based supervised method, the minimal learning machine, constructs a predictive model from data by learning a mapping between input and output distance matrices. In this paper, we propose new methods and evaluate how their core component, the distance mapping, can be adapted to multi-label learning. The proposed approach is based on combining the distance mapping with an inverse distance weighting. Although the proposal is one of the simplest methods in the multi-label learning literature, it achieves state-of-the-art performance for small to moderate-sized multi-label learning problems. In addition to its simplicity, the proposed method is fully deterministic: Its hyper-parameter can be selected via ranking loss-based statistic which has a closed form, thus avoiding conventional cross-validation-based hyper-parameter tuning. In addition, due to its simple linear distance mapping-based construction, we demonstrate that the proposed method can assess the uncertainty of the predictions for multi-label classification, which is a valuable capability for data-centric machine learning pipelines. Joonas Hämäläinen, Antoine Hubermont, Amauri H. Souza, César Lincoln C. Mattos, João Paulo Pordeus Gomes, Tommi Kärkkäinen |
Mach. Learn. | 5 |
| 2024 | Amortized Variational Deep Kernel LearningabstractDeep kernel learning (DKL) marries the uncertainty quantification of Gaussian processes (GPs) and the representational power of deep neural networks. However, training DKL is challenging and often leads to overfitting. Most notably, DKL often learns “non-local” kernels — incurring spurious correlations. To remedy this issue, we propose using amortized inducing points and a parameter-sharing scheme, which ties together the amortization and DKL networks. This design imposes an explicit dependency between the ELBO’s model fit and capacity terms. In turn, this prevents the former from dominating the optimization procedure and incurring the aforementioned spurious correlations. Extensive experiments show that our resulting method, amortized varitional DKL (AVDKL), i) consistently outperforms DKL and standard GPs for tabular data; ii) achieves significantly higher accuracy than DKL in node classification tasks; and iii) leads to substantially better accuracy and negative log-likelihood than DKL on CIFAR100. Alan Lucas Silva Matias, César Lincoln C. Mattos, João Paulo Pordeus Gomes, Diego Mesquita |
ICML | 3 |
| 2024 | Towards automatic labeling of exception handling bugs: A case study of 10 years bug-fixing in Apache Hadoop
Antônio da Silva, Renan Gomes Vieira, Diego Mesquita, João Paulo Pordeus Gomes, Lincoln S. Rocha |
Empir. Softw. Eng. | 4 |
| 2024 | Spatio-temporal wind speed forecasting with approximate Bayesian uncertainty quantification
Airton F. Souza Neto, César Lincoln C. Mattos, João Paulo Pordeus Gomes |
Neural Comput. Appl. | 3 |
| 2023 | Thin and deep Gaussian processesabstractGaussian processes (GPs) can provide a principled approach to uncertainty quantification with easy-to-interpret kernel hyperparameters, such as the lengthscale, which controls the correlation distance of function values.However, selecting an appropriate kernel can be challenging.
Deep GPs avoid manual kernel engineering by successively parameterizing kernels with GP layers, allowing them to learn low-dimensional embeddings of the inputs that explain the output data.
Following the architecture of deep neural networks, the most common deep GPs warp the input space layer-by-layer but lose all the interpretability of shallow GPs. An alternative construction is to successively parameterize the lengthscale of a kernel, improving the interpretability but ultimately giving away the notion of learning lower-dimensional embeddings. Unfortunately, both methods are susceptible to particular pathologies which may hinder fitting and limit their interpretability.
This work proposes a novel synthesis of both previous approaches: {Thin and Deep GP} (TDGP). Each TDGP layer defines locally linear transformations of the original input data maintaining the concept of latent embeddings while also retaining the interpretation of lengthscales of a kernel. Moreover, unlike the prior solutions, TDGP induces non-pathological manifolds that admit learning lower-dimensional representations.
We show with theoretical and experimental results that i) TDGP is, unlike previous models, tailored to specifically discover lower-dimensional manifolds in the input data, ii) TDGP behaves well when increasing the number of layers, and iii) TDGP performs well in standard benchmark datasets. Daniel Augusto R. M. A. de Souza, Alexander Nikitin 0002, St John, Magnus Ross, Mauricio A. Álvarez, Marc Peter Deisenroth, João Paulo Pordeus Gomes, Diego Mesquita, César Lincoln C. Mattos |
NeurIPS | 7 |
| 2022 | Predicting Test Execution Times with Asymmetric Random ForestsabstractBeing able to estimate a test execution time is of fundamental importance when you need to prioritize tests.Furthermore, it is also important that an estimation algorithm do not underestimate the execution time, since time can be a hard constraint in many problems.If a test take longer than expected, some test that is planned to be executed in the future may have to be cancelled.Under such scenario, in this paper, we developed two simple variants of the Random Forest regression algorithm to predict test execution times in storage diagnostics tests.The proposed methods are compared to a baseline time estimation method (already available in a commercial product) and other machine learning based models.On the basis of our experiments we can state that the proposed variants achieved promising results when considering an asymmetric error metric. Francisco L. F. Pereira, Helio Silva, João Paulo Pordeus Gomes, Javam C. Machado |
ESANN | 3 |
| 2022 | Bayesian Analysis of Bug-Fixing Time using Report DataabstractBackground: Bug-fixing is the crux of software maintenance. It entails tending to heaps of bug reports using limited resources. Using historical data, we can ask questions that contribute to better-informed allocation heuristics. The caveat here is that often there is not enough data to provide a sound response. This issue is especially prominent for young projects. Also, answers may vary from project to project. Consequently, it is impossible to generalize results without assuming a notion of relatedness between projects. Renan Gomes Vieira, Diego Mesquita, César Lincoln C. Mattos, Ricardo Britto 0001, Lincoln S. Rocha, João Paulo Pordeus Gomes |
ESEM | 6 |
| 2022 | Detecting Customer Induced Damages in Motherboards with Deep Neural NetworksabstractIdentifying Customer Induced Damage (CID) is a key part in warranty programs of electronics manufacturers. CID is defined as any damage in the unit performed by an unauthorized person including the customer. In such cases, damaged units are not covered by warranty. An important aspect of CID inspection activities is that, in most industries, the task is performed manually. Since such task demands high attention to details, human mistakes occur very often. For such reasons, CID detection can be considered as a good candidate to be handled by automatic detection methods with computer vision tools. In this work, we evaluate modern computer vision methods based on deep neural networks (SwinSoft Teacher and Mask-RCNN as prediction networks with swin and resnet backbones) for automatic CID detection in motherboard images. All images were originated from private company repair centers around the globe. We highlight the following results for Average Precision (AP) and Average Recall (AR):$AP{@}0.5 =\mathbf{80.2}$and$AR^{m}=\mathbf{57.9}$, for Mask-RCNN with Swin-S backbone, where the variant$m$means medium, and$AP{@}0.5 =\mathbf{74.2}$and$AR^{m}=\mathbf{45.4}$, for Mask-RCNN Swin-T backbone. Our results demonstrate that Swin-based methods had the best overall results and exhibit interesting characteristics for our application. Danilo Alves, Victor A. E. de Farias, Iago C. Chaves, Richard Chao, João P. V. Madeiro, João Paulo Pordeus Gomes, Javam C. Machado |
IJCNN | 6 |
| 2022 | The role of bug report evolution in reliable fixing estimation
Renan Gomes Vieira, César Lincoln C. Mattos, Lincoln S. Rocha, João Paulo Pordeus Gomes, Matheus Paixão |
Empir. Softw. Eng. | 4 |
| 2022 | Self-tuning portfolio-based Bayesian optimization
Thiago de P. Vasconcelos, Daniel Augusto R. M. A. de Souza, Gustavo C. de M. Virgolino, César Lincoln C. Mattos, João Paulo Pordeus Gomes |
Expert Syst. Appl. | 5 |
| 2022 | On the design of a similarity function for sparse binary data with application on protein function annotation
Marcelo B. A. Veras, Bishnu Sarker, Sabeur Aridhi, João Paulo Pordeus Gomes, José A. F. de Macêdo, Engelbert Mephu Nguifo, Marie-Dominique Devignes, Malika Smaïl-Tabbone |
Knowl. Based Syst. | 4 |
| 2022 | Bayesian MultilaterationabstractMultilateration (MLAT) is thede factotechnique to localize points of interest (POIs) in navigation and surveillance systems. Despite sensors being inherently noisy, most existing techniques i) are oblivious to noise patterns in sensor measurements; and ii) only provide point estimates of the POI. This often results in unreliable estimates with high variance,i.e., that are highly sensitive to measurement noise. To overcome this caveat, we advocate the use of Bayesian modeling. Using Bayesian statistics, we provide a comprehensive guide to handle uncertainties in MLAT, including principled choices for the likelihood function and the prior distributions. Notably, the resulting model is easy to implement and can leverage off-the-shelf Markov Chain Monte Carlo (MCMC) software for inference. Besides coping with unreliable measurements, our framework can also deal with sensors whose location is not completely known, which is an asset in mobile systems. Our solution also naturally incorporates multiple measurements per reference point, a common practical situation that is usually not handled directly by other approaches. Comprehensive experiments with both synthetic and real-world data indicate that our Bayesian approach to the MLAT task provides better position estimation and uncertainty quantification when compared to the available alternatives. Alisson S. C. Alencar, César Lincoln C. Mattos, João Paulo Pordeus Gomes, Diego Mesquita |
IEEE Signal Process. Lett. | 3 |
| 2021 | Learning GPLVM with arbitrary kernels using the unscented transformationabstractGaussian Process Latent Variable Model (GPLVM) is a flexible framework to handle uncertain inputs in Gaussian Processes (GPs) and incorporate GPs as components of larger graphical models. Nonetheless, the standard GPLVM variational inference approach is tractable only for a narrow family of kernel functions. The most popular implementations of GPLVM circumvent this limitation using quadrature methods, which may become a computational bottleneck even for relatively low dimensions. For instance, the widely employed Gauss-Hermite quadrature has exponential complexity on the number of dimensions. In this work, we propose using the unscented transformation instead. Overall, this method presents comparable, if not better, performance than off-the-shelf solutions to GPLVM, and its computational complexity scales only linearly on dimension. In contrast to Monte Carlo methods, our approach is deterministic and works well with quasi-Newton methods, such as the Broyden-Fletcher-Goldfarb-Shanno (BFGS) algorithm. We illustrate the applicability of our method with experiments on dimensionality reduction and multistep-ahead prediction with uncertainty propagation. Daniel Augusto R. M. A. de Souza, Diego Mesquita, João Paulo Pordeus Gomes, César Lincoln C. Mattos |
AISTATS | 3 |
| 2021 | A novel fuzzy ARTMAP with area of influence
Alan Lucas Silva Matias, Ajalmar R. da Rocha Neto, César Lincoln C. Mattos, João Paulo Pordeus Gomes |
Neurocomputing | 4 |
| 2021 | Predicting the Health Degree of Hard Disk Drives With Asymmetric and Ordinal Deep Neural ModelsabstractPredicting failures in Hard Disk Drives (HDD) is a major challenge that has been faced by both industry and academy in recent years. Being able to predict failure events may incur in avoiding data losses and also improve service availability. Among all failure prediction strategies, the health degree prediction is one of the most popular. The task of health degree prediction consists of, given a finite set of health states that are related to the degradation of the equipment, estimate which state reflects the actual degradation of the equipment. This problem is usually modeled as a classification task. Although many health degree prediction methods have been proposed, some practical details regarding this prediction task have been neglected in previous works. In this work we tackle two of these aspects: the ordinal nature of the problem and the different costs associated with miss-classifications. The problem can be considered as ordinal since classifying a HDD in a health level that is far from is true health level shall be more penalized than classifying in a near health level, thus a classical classification framework is not recommended. The different costs associated with mis-classifications are related to the fact that early predictions are preferred than late prediction since the later can result in failures. Such aspects are considered in a framework based on Deep Recurrent Neural Networks (DRNN). The choice of DRNN is given its remarkable performances in many applications including HDDs failure prediction. The resulting methods outperformed state-of-the-art approaches in a metric that consider the new aspects that motivated our proposal. Fernando Dione S. Lima, Francisco L. F. Pereira, Iago C. Chaves, Javam C. Machado, João Paulo Pordeus Gomes |
IEEE Trans. Computers | 5 |
| 2020 | On the Use of Cultural Enhancement Strategies to Improve the NEAT AlgorithmabstractKnowledge transmitted between generations by non-genetic means can be understood as culture. The capacity of individuals from certain species to teach and learn plays a fundamental role in directing the evolutionary process. The Neuroevolution of Augmenting Topologies (NEAT) framework enables evolving neural structures to iteratively solve a given learning problem. However, the NEAT approach does not consider cultural aspects in its formulation. In such a context, the aim of this paper is to propose and evaluate ways of enhancing the NEAT framework with additional learning approaches. The parameters involved in the analysis comprise the Backpropagation and the Extreme Learning Machine (ELM) learning algorithms, the individuals to be taught, the moment when culture manifests in the system, and the nature of the lessons to be learned. Empirical results on sequential learning tasks indicate that cultural enhancements, as well as some of the proposed variations, accelerate the neuroevolution convergence. Arthur L. A. Paulino, Yuri Lenon Barbosa Nogueira, João Paulo Pordeus Gomes, César Lincoln C. Mattos, Leonardo Ramos Rodrigues |
CEC | 3 |
| 2020 | A Hybrid TLBO-Particle Filter Algorithm Applied to Remaining Useful Life Prediction in the Presence of Multiple Degradation FactorsabstractOne of the end goals of a Prognostic and Health Monitoring (PHM) algorithm is to provide accurate Remaining Useful Life (RUL) predictions for the monitored component or system. Most of the PHM algorithms found in the literature are based on the assumption that the degradation process is governed by only one degradation factor. However, some components and systems may be subject to multiple degradation factors. In this paper, we propose a hybrid algorithm that incorporates a Teaching-Learning Based optimization (TLBO) step into a Particle Filter (PF) framework. PF is an algorithm that can handle multiple degradation factors. However, it has some drawbacks such as sample degeneracy and sample impoverishment. The hybrid TLBO-PF algorithm proposed in this paper improves the performance of the standard PF algorithm by reducing the effects of sample degeneracy and sample impoverishment. A case study is presented to evaluate the performance of the proposed algorithm for estimating the degradation factors and predicting the RUL of a Lithium-ion battery, which is affected by two degradation factors. The results show that the proposed algorithm presented a better performance for both the tasks (degradation factor estimation and RUL prediction) when compared with the standard Particle Filter algorithm. Leonardo Ramos Rodrigues, Daniel B. P. Coelho, João Paulo Pordeus Gomes |
CEC | 3 |
| 2020 | A New Methodology for Classifying QRS Morphology in ECG SignalsabstractThe electrocardiogram (ECG) is a non-invasive method to detect cardiovascular diseases (CVD), the most common cause of death in the world. The recognition of heartbeat morphologies present in the ECG signal is an effective way to detect CVDs prematurely. Many approaches were developed for this purpose, such as the use of Wavelets, High Order Statistics (HOS), Local Binary Patterns (LBP), Random Projection, Fiducial points, and Hermite Polynomials. Unfortunately, most parts of these approaches suffer from the high variability of ECG signal features and conditions. Also, it is common to use more than one of them simultaneously, which makes it hard to infer the contributions of each one. This work presents a new robust methodology to extract features for heartbeat morphology classification. Moreover, we introduce new labels for a small set of morphologies present in MIT-BIH Arrhythmia database, taking into account only the QRS complex (the 3 more representative waves of a heartbeat) instead of the whole heartbeat. We evaluate each approach in isolation and the results show that our method outperforms other well-known strategies. Weslley L. Caldas, João P. V. Madeiro, César Lincoln C. Mattos, João Paulo Pordeus Gomes |
IJCNN | 4 |
| 2020 | Using Autoencoders for Anomaly Detection in Hard Disk DrivesabstractNowadays, predicting failures in Hard Disk Drives (HDD) is of key importance for storage service providers and end users. Being able to detect in advance that a disk is going to fail may enable maintenance actions that can avoid severe data losses. For that reason, many researchers had devoted attention to this research topic. Recently, several authors have reported promising results by using attributes collected by the SMART (Self-Monitoring, Analysis and Reporting Technology) system along with machine learning methods. Although the best results were obtained by supervised machine learning methods, it is important to notice that data from degraded HDDs is scarce. Hence, anomaly detection methods arise a promising solution. Among such methods, recent works reported that reconstruction based anomaly detection algorithms had the best performance on HDDs fault detection. In line with such results, in this paper we aim to further investigate the performance of such methods. We conducted tests with classical PCA based methods and neural autoencoder based methods. In addition to testing with the popular reconstruction based autoencoder method we also evaluated a method that analyzes the distribution of the latent space. Additionally we propose a simple formulation to combine both methods. On the basis of our experiments, we verified that autoencoder based methods had the best performances according to the two evaluation metrics. Among such methods, the combination approach had the best overall performance. Francisco L. F. Pereira, Iago C. Chaves, João Paulo Pordeus Gomes, Javam C. Machado |
IJCNN | 3 |
| 2020 | Parsimonious Minimal Learning Machine via Multiresponse Sparse RegressionabstractThe training procedure of the minimal learning machine (MLM) requires the selection of two sets of patterns from the training dataset. These sets are called input reference points (IRP) and output reference points (ORP), which are used to build a mapping between the input geometric configurations and their corresponding outputs. In the original MLM, the number of input reference points is the hyper-parameter and the patterns are chosen at random. Therefore, the conventional proposal does not consider which patterns will belong to each reference point group, since the model does not implement an appropriate way of selecting the most suitable patterns as reference points. Such an approach can impact on the decision function in terms of smoothness, resulting in high complexity models. This paper introduces a new approach to select IRP for MLM applied to classification tasks. The optimally selected minimal learning machine (OS-MLM) relies on the multiresponse sparse regression (MRSR) ranking method and the leave-one-out (LOO) criterion to sort the patterns in terms of relevance and select an appropriate number of input reference points, respectively. The experimental assessment conducted on UCI datasets reports the proposal was able to produce sparser models and achieve competitive performance when compared to the regular strategy of selecting MLM input RPs. Madson L. D. Dias, Átilla N. Maia, Ajalmar R. da Rocha Neto, João Paulo Pordeus Gomes |
Int. J. Neural Syst. | 4 |
| 2020 | A new perspective for Minimal Learning Machines: A lightweight approach
José A. V. Florêncio, Saulo A. F. Oliveira, João Paulo Pordeus Gomes, Ajalmar R. da Rocha Neto |
Neurocomputing | 3 |
| 2020 | A sparse linear regression model for incomplete datasets
Marcelo B. A. Veras, Diego Mesquita, César Lincoln C. Mattos, João Paulo Pordeus Gomes |
Pattern Anal. Appl. | 4 |
| 2020 | LS-SVR as a Bayesian RBF Network
Diego Mesquita, Luis A. Freitas, João Paulo Pordeus Gomes, César Lincoln C. Mattos |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Sparse minimal learning machine using a diversity measure minimization
Madson L. D. Dias, Lucas Silva de Sousa, Ajalmar R. da Rocha Neto, César Lincoln C. Mattos, João Paulo Pordeus Gomes, Tommi Kärkkäinen |
ESANN | 5 |
| 2019 | No-PASt-BO: Normalized Portfolio Allocation Strategy for Bayesian OptimizationabstractBayesian Optimization (BO) is a framework for black-box optimization that is especially suitable for expensive cost functions. Among the main parts of a BO algorithm, the acquisition function is of fundamental importance, since it guides the optimization algorithm by translating the uncertainty of the regression model in a utility measure for each point to be evaluated. Considering such aspect, selection and design of acquisition functions are one of the most popular research topics in BO. Since no single acquisition function was proved to have better performance in all tasks, a well-established approach consists of selecting different acquisition functions along the iterations of a BO execution. In such an approach, the GP-Hedge algorithm is a widely used option given its simplicity and good performance. Despite its success in various applications, GP-Hedge shows an undesirable characteristic of accounting on all past performance measures of each acquisition function to select the next function to be used. In this case, good or bad values obtained in an initial iteration may impact the choice of the acquisition function for the rest of the algorithm. This fact may induce a dominant behavior of an acquisition function and impact the final performance of the method. Aiming to overcome such limitation, in this work we propose a variant of GP-Hedge, named No-PASt-BO, that reduce the influence of far past evaluations. Moreover, our method presents a built-in normalization that avoids the functions in the portfolio to have similar probabilities, thus improving the exploration. The obtained results on both synthetic and real-world optimization tasks indicate that No-PASt-BO presents competitive performance and always outperforms GP-Hedge. Thiago de P. Vasconcelos, Daniel Augusto R. M. A. de Souza, César Lincoln C. Mattos, João Paulo Pordeus Gomes |
ICTAI | 4 |
| 2019 | Artificial Neural Networks with Random Weights for Incomplete Datasets
Diego Mesquita, João Paulo Pordeus Gomes, Leonardo Ramos Rodrigues |
Neural Process. Lett. | 2 |
| 2018 | A Modified Symbiotic Organisms Search Algorithm Applied to Flow Shop Scheduling ProblemsabstractThe Symbiotic Organism Search (SOS) algorithm is an optimization metaheuristic inspired by the symbiotic relationships that occur among organisms in nature. In the last few years, the SOS algorithm attracted increasing attention due to its good performance on various real-world problems, despite the fact that no specific parameter adjustment is required. In this paper, we propose an improved version of SOS by modifying the organisms selection strategy. In the proposed version of the algorithm, three organisms are selected from the population without having a predefined symbiotic relationship. Once the organisms are selected, an assignment step is conducted to assign each organism to a symbiotic relationship. We tested the performance of the proposed algorithm using twenty benchmark instances of the flow shop scheduling problem. We compared the results with the results obtained using the original SOS algorithm. The proposed modification improved the performance of the SOS algorithm in the search for the global optimum value in most of the instances. Leonardo Ramos Rodrigues, João Paulo Pordeus Gomes, Ajalmar R. da Rocha Neto, Amauri H. Souza |
CEC | 2 |
| 2018 | Optimally Selected Minimal Learning Machine
Átilla N. Maia, Madson L. D. Dias, João Paulo Pordeus Gomes, Ajalmar R. da Rocha Neto |
IDEAL (1) | 3 |
| 2018 | Hard Disk Drive Failure Prediction Method Based On A Bayesian NetworkabstractThe ability to predict failures in Hard Disk Drives (HDD) is a major objective of HDD manufacturers since avoiding unexpected failures may prevent data loss. As a consequence, failure prediction in HDDs became a topic that attracted much attention in recent years. Nowadays, most HDDs are equipped with a threshold-based monitoring system named Self-Monitoring, Analysis and Reporting Technology (SMART). The system collects several performance parameters and detects anomalies that may indicate incipient failures. Although the SMART system is very popular, it achieves failure detection rates of 3% to 10%. Moreover, SMART works as an incipient failure detection method and does not provide an estimate of the remaining life of the HDD. In this paper, we propose a failure prediction method using SMART attributes and a Bayesian Network. The proposed method, named Bayesian network based Method for Failure prediction in HDDs (BNFH) uses a subset of the SMART attributes and a set of SMART trend related attributes to provide remaining life estimates of HDDs. To demonstrate practical usefulness, this method was applied to a dataset consisting of 49,056 hard drives from Backblaze's data centers. Iago C. Chaves, Manoel Rui P. de Paula, Lucas G. M. Leite, João Paulo Pordeus Gomes, Javam C. Machado |
IJCNN | 4 |
| 2018 | Remaining Useful Life Estimation of Hard Disk Drives based on Deep Neural NetworksabstractNowadays Hard Disk Drives (HDDs) are essential storage devices in most large-scale storage systems. As a consequence, HDD failures have severe effects that may range from data loss to service unavailability. Considering such scenario, academy and industry have driven its attention to HDDs failure prognostics solutions. In this work, we evaluate two of the most common deep learning architecture in the task of HDD failure prediction. Our tests were conducted on real-world datasets and the proposals were compared to a recurrent neural network. The results of this study showed that deep learning models are a valid alternative to solve this problem since they achieved good results and a significant amount of data is available. Fernando Dione S. Lima, Francisco L. F. Pereira, Lucas G. M. Leite, João Paulo Pordeus Gomes, Javam C. Machado |
IJCNN | 4 |
| 2018 | Fuzzy ART-Based Classification via Sparse Bayesian LearningabstractIn this paper, a new Fuzzy ART-based architecture called SBFA is introduced. The SBFA is composed of the modules ART and Target (T), and a third novel module called SC2that links these two modules. The new module SC2formulates the multiresponse classification problem as a linear system computed from the input data, the ART categories, and the targets labels in the module T. The module SC2then solves the linear system through the multiresponse Sparse Bayesian Learning (M-SBL) algorithm. The solution of such a linear system is the prediction field W. Since we use the M-SBL algorithm, the prediction field is sparse and, thus, the number of categories also decreases. From the achievements, our proposal is in general equivalent to state-of-art models and outperforms other ARTMAP-based models. Also, the SBFA successfully incorporates the M-SBL in the learning process. Alan Lucas Silva Matias, Lucas Silva de Sousa, Ajalmar R. da Rocha Neto, João Paulo Pordeus Gomes |
IJCNN | 4 |
| 2018 | A bi-directional evaluation-based approach for image retargeting quality assessment
Saulo A. F. Oliveira, Shara Shami Araújo Alves, João Paulo Pordeus Gomes, Ajalmar R. da Rocha Neto |
Comput. Vis. Image Underst. | 3 |
| 2018 | Regression based performance modeling and provisioning for NoSQL cloud databases
Victor A. E. de Farias, Flávio R. C. Sousa, José G. R. Maia, João Paulo Pordeus Gomes, Javam C. Machado |
Future Gener. Comput. Syst. | 4 |
| 2018 | Sparse Least-Squares Support Vector Machines via Accelerated Segmented Test: A dual approach
Saulo A. F. Oliveira, João Paulo Pordeus Gomes, Ajalmar R. da Rocha Neto |
Neurocomputing | 2 |
| 2018 | Using Degradation Messages to Predict Hydraulic System Failures in a Commercial AircraftabstractThis paper presents a failure prognostics methodology based on degradation messages using a particle filter framework. In the proposed method, the degradation messages are interpreted as quantized measurements. The use of degradation messages reduces the need for data communication bandwidth, since they only need to be transmitted or stored when some predefined threshold is crossed. In contrast, the direct monitoring of degradation-related measurements requires frequent updates, usually at a fixed sampling rate. This characteristic is fundamental on failure prognostics applications in real aircraft due to the high infrastructure costs associated with data transmission. A case study using field data recorded from commercial aircraft is presented to illustrate the proposed methodology. The problem under consideration consists of estimating the time remaining until the fluid level reaches unacceptably low values in the hydraulic system. Two quantization steps are considered in the evaluation. Predictions employing direct measurements of fluid level are also performed to establish a comparative performance baseline. The results show that it is possible to choose a quantization step that allows a reduction in the transmission costs without significant loss of prognostics performance. João Paulo Pordeus Gomes, Leonardo Ramos Rodrigues, Bruno P. Leao, Roberto Kawakami Harrop Galvão, Takashi Yoneyama |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2018 | Spare Parts List Recommendations for Multiple-Component Redundant Systems Using a Modified Pareto Ant Colony Optimization ApproachabstractHigh availability requirements for systems have become common in many industry sectors. In order to comply with these requirements, companies must keep spare parts to reduce system downtime when a failure event occurs. A problem that comes up is the definition of the set of spare parts that should be acquired to ensure that the desired availability will be achieved. Efficient algorithms are available to solve this problem for the cases in which the failure of any component within the system leads to a system failure. However, highly integrated systems for safety-critical applications commonly have complex interactions among their components. Also, redundancies are commonly used in system architecture. For these systems, finding nondominated recommended spare parts lists (RSPLs) is a challenging task. The problem addressed in this paper consists in finding nondominated solutions for a bi-objective RSPL problem considering RSPL cost and system availability. We propose an algorithm to solve the RSPL problem using a modified Pareto ant colony optimization (MPACO) approach. In the proposed MPACO, we use marginal analysis to generate the initial population. Also, we propose a new approach to perform the pheromone update step based on the distances between consecutive points in the current solution. Numerical experiments are carried out to illustrate the application of the proposed MPACO in three different multiple-component redundant systems with different complexity levels. The results show that the proposed MPACO presents a better performance when compared with the original PACO and the nondominated sorting genetic algorithm II. Leonardo Ramos Rodrigues, João Paulo Pordeus Gomes |
IEEE Trans. Ind. Informatics | 2 |
| 2017 | A Robust Minimal Learning Machine based on the M-Estimator
João Paulo Pordeus Gomes, Diego Mesquita, Ananda Freire, Amauri H. Souza, Tommi Kärkkäinen |
ESANN | 1 |
| 2017 | Euclidean distance estimation in incomplete datasets
Diego Mesquita, João Paulo Pordeus Gomes, Amauri H. Souza, Juvêncio S. Nobre |
Neurocomputing | 2 |
| 2017 | Ensemble of Efficient Minimal Learning Machines for Classification and Regression
Diego Mesquita, João Paulo Pordeus Gomes, Amauri H. Souza |
Neural Process. Lett. | 2 |
| 2017 | A stochastic framework for K-SVD with applications on face recognition
Gustavo Malkomes, Carlos Brito 0001, João Paulo Pordeus Gomes |
Pattern Anal. Appl. | 3 |
| 2017 | A Fault Detection Method for Hard Disk Drives Based on Mixture of Gaussians and Nonparametric StatisticsabstractHard Disk Drives (HDD) failure prediction is a challenging topic that has attracted much attention in recent years. Predicting failures in HDD may avoid losing data thus improving data reliability. Previous works on failure prediction are based on parametric approaches that model healthy drives with a Gaussian distribution. Although they achieved good results, the Gaussianity assumption may not hold true. The following work proposes a method for fault detection in HDD based on a Gaussian Mixture Model. A self-monitoring, analysis, and reporting technology dataset is used to evaluate the proposed method. Results show that the method outperforms previous works in both fault detection and time before failure. Lucas P. Queiroz, Francisco Caio M. Rodrigues, João Paulo Pordeus Gomes, Felipe T. Brito, Iago C. Chaves, Manoel Rui P. de Paula, Marcos Rogério Salvador, Javam C. Machado |
IEEE Trans. Ind. Informatics | 3 |
| 2016 | Machine Learning Approach for Cloud NoSQL Databases Performance ModelingabstractCloud computing is a successful, emerging paradigm that supports on-demand services with pay-as-you-go model. With the exponential growth of data, NoSQL databases have been used to manage data in the cloud. In these newly emerging settings, mechanisms to guarantee Quality of Service heavily relies on performance predictability, i.e., the ability to estimate the impact of concurrent query execution on the performance of individual queries in a continuously evolving workload. This paper presents a performance modeling approach for NoSQL databases in terms of performance metrics which is capable of capturing the non-linear effects caused by concurrency and distribution aspects. Experimental results confirm that our performance modeling can accurately predict mean response time measurements under a wide range of workload configurations. Victor A. E. de Farias, Flávio R. C. Sousa, José G. R. Maia, João Paulo Pordeus Gomes, Javam C. Machado |
CCGrid | 4 |
| 2016 | K-means for Datasets with Missing Attributes: Building Soft Constraints with Observed and Imputed Values
Diego Mesquita, João Paulo Pordeus Gomes, Leonardo Ramos Rodrigues |
ESANN | 2 |
| 2016 | Using Robust Extreme Learning Machines to Predict Cotton Yarn Strength and Hairiness
Diego Mesquita, Antônio C. Araújo Neto, Jose Queiroz Neto, João Paulo Pordeus Gomes, Leonardo Ramos Rodrigues |
ESANN | 4 |
| 2016 | Radial Basis Function Neural Networks for Datasets with Missing Values
Diego Mesquita, João Paulo Pordeus Gomes |
ISDA | 2 |
| 2016 | A New Degradation Indicator Based on a Statistical Anomaly ApproachabstractThis paper presents a new method for combining measured parameters into a single indicator for monitoring the condition of systems subjected to degradation effects. The proposed approach integrates the use of nonparametric density estimation techniques into Runger's U2method, which allows for the separation of variables that are directly related to degradation effects from those which are not. Two simulated case studies are presented for illustration, namely, the monitoring of a flap extension and retraction system, and a gas turbine employed as an auxiliary power unit. For comparison, degradation indicators are also calculated by using Hotteling's T2and Runger's U2methods, as well as a nonparametric method without separation of variables. In both case studies, the proposed method provided the best results in terms of fault detection performance and suitability for remaining useful life prediction. João Paulo Pordeus Gomes, Roberto Kawakami Harrop Galvão, Takashi Yoneyama, Bruno P. Leao |
IEEE Trans. Reliab. | 1 |
| 2016 | Asymmetric Unscented Transform for Failure PrognosisabstractThis paper presents a novel method for sigma point (SP) selection in the context of the unscented transform (UT). Such a method yields an asymmetric SP set, which can be especially useful in failure prognosis applications. The proposed method can be employed to overcome performance limitations associated with the application of conventional UT for the estimation of remaining useful life (RUL) of equipment, while maintaining the benefits yielded by the use of the UT in this task. Sample applications based on simulated and real data of aircraft systems degradation are presented to illustrate and validate the concept. Results indicate that the proposed method yields improved RUL estimation results at lower computational cost when compared to the original UT. Bruno P. Leao, Takashi Yoneyama, João Paulo Pordeus Gomes |
IEEE Trans. Reliab. | 3 |
| 2015 | A Cost Sensitive Minimal Learning Machine for Pattern Classification
João Paulo Pordeus Gomes, Amauri H. Souza, Francesco Corona, Ajalmar R. da Rocha Neto |
ICONIP (1) | 1 |
| 2015 | A Minimal Learning Machine for Datasets with Missing Values
Diego Mesquita, João Paulo Pordeus Gomes, Amauri H. Souza |
ICONIP (1) | 2 |