VLDB 2026 Research / reviewers in the wild / expert
Júlio C. Nievola
dblp:62/6863 · also Júlio César Nievola
· DBLP profile ↗
29ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-2212-4499ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Hybrid Software Testing Recommendation System Based on Bugs and Tests DescriptionsabstractDuring the software development cycle, there is often a need to run tests after correcting a bug. The tests to be executed are only a few times automated, demanding even more time from the tester. In addition to the testing time, there is also the difficulty in selecting the tests that should be executed at that moment, as depending on the size of the software, the number of tests reaches thousands. In this research, we propose a hybrid model for recommending manual tests, which, based on the bug description and test steps, selects the best tests to be carried out to check whether a given bug has been resolved. This work was developed using real-world data from software under development, which features manual user interface tests. During development, our application presented an accuracy rate of 60.82% through the Recall@20 metric using the training data, which was confirmed through the use of human users, where the system presented an accuracy rate of 59.70%. Bruno Kostiuk, Gregory Moro Puppi Wanderley, Cesar Augusto Tacla, Antônio David Viniski, Júlio C. Nievola, Emerson Cabrera Paraiso |
CSCWD | 5 |
| 2024 | Biclustering-based multi-label classification
Luiz Rafael Schmitke, Emerson Cabrera Paraiso, Júlio C. Nievola |
Knowl. Inf. Syst. | 3 |
| 2022 | Classifying Hierarchical Data Streams using Global Classifiers and Summarization TechniquesabstractThe hierarchical classification of data streams requires models capable of handling a class hierarchy and updating themselves whenever a new example arrives, within restrained processing time and memory consumption. Current state-of-the-art models store raw instances and handle the hierarchy locally, performing a high number of computations at every hierarchy level and with all, eventually redundant, data. This paper introduces Global k-Nearest Centroids (kNC) and Global Dribble, two novel methods for the hierarchical classification of data streams. Both methods use summarization techniques to represent data with constant computational resources usage and a global classification approach to process instances in less time when compared to local strategies. We compare both methods with a state-of-the-art local classifier, and the proposed methods achieved a higher number of correct predictions and process instances nearly twice as fast. Eduardo Tieppo, Jean Paul Barddal, Júlio C. Nievola |
IJCNN | 3 |
| 2022 | Improving Data Stream Classification using Incremental Yeo-Johnson Power TransformationabstractData transformation plays an essential role as a preprocessing step in learning models. Several classification techniques have premises about the underlying data distribution, such as normal distribution assumed in Bayesians classifiers. However, applying data transformation in a streaming setting requires processing an infinite and continuous flow of data. In this paper, we propose the Incremental Yeo-Johnson Power Transformation, a variant of the well-known batch Yeo-Johnson transformation that is tailored for streaming settings, i.e., it supports streaming data via statistical sampling and hypothesis testing. Experimental results show that our proposal achieves the same data normality as its batch counterpart. In addition, it improves the prediction performance of a data stream classifier based on Bayesian statistical models. Overall, learning models obtained 3 percentage points improvement. Eduardo Tieppo, Jean Paul Barddal, Júlio C. Nievola |
SMC | 3 |
| 2021 | Adaptive Global k-Nearest Neighbors for Hierarchical Classification of Data StreamsabstractData stream classification differs from batch learning classification methods as data is made available sequentially and may drift over time. Therefore, data stream classification can be simultaneous to all other kinds of classification problems, and it has been revisiting many aspects related to classification in the last years. So far, hierarchical classification was weakly addressed in streaming scenarios despite being a well-established research topic. To fill in this gap between such areas, in this paper, we propose the adaptive global k-Nearest Neighbors for the hierarchical classification of data streams (Global kNN-hDS). Our proposal classifies hierarchical data streams using a constrained memory buffer and a global classification approach. We compare our method against a stateof-the-art local kNN also tailored for streaming scenarios, and results show that our method obtains competitive prediction rates while being statistically faster. Eduardo Tieppo, Jean Paul Barddal, Júlio C. Nievola |
SMC | 3 |
| 2021 | BNPA: An R package to learn path analysis input models from a data set semi-automatically using Bayesian networks
Elias Cesar Araujo De Carvalho, Joao Ricardo Nickenig Vissoci, Luciano de Andrade, Wagner de Lara Machado, Emerson Cabrera Paraiso, Júlio C. Nievola |
Knowl. Based Syst. | 6 |
| 2018 | A generalized financial time series forecasting model based on automatic feature engineering using genetic algorithms and support vector machineabstractWe propose the genetic algorithm for time window optimization, which is an embedded genetic algorithm (GA), to optimize the time window (TW) of the attributes using feature selection and support vector machine. This GA is evolved using the results of a trading simulation, and it determines the best TW for each technical indicator. An appropriate evaluation was conducted using a walk-forward trading simulation, and the trained model was verified to be generalizable for forecasting other stock data. The results show that using the GA to determine the TW can improve the rate of return, leading to better prediction models than those resulting from using the default TW. Norberto Ritzmann Junior, Júlio C. Nievola |
IJCNN | 2 |
| 2017 | Combining Process Mining with Trace Clustering: Manufacturing Shop Floor Process - An Applied CaseabstractProcess mining allows observing process execution based on real event data and proposes methods and tools to provide diagnostics, reducing the gaps between practice and conceptual models. When process mining discovery techniques are applied in flexible processes with many decisions at runtime, the results are often semi-structured or unstructured process models that are difficult to understand. In this context, trace clustering is an approach for reducing the complexity of process models and improving the accuracy and comprehensibility. This paper presents an applied case in industrial manufacturing production with an unstructured process and issues in production performance indicators. A set of techniques were used to understand how the process occurs in practice, how many trace clusters should be identified as homogeneous process variants, and what causes production inefficiency. Finally, the results identify bottlenecks caused by erroneous decisions at runtime and serve to support process improvement. Alex Meincheim, Cleiton dos Santos Garcia, Júlio C. Nievola, Edson Emílio Scalabrin |
ICTAI | 3 |
| 2016 | Facial expression recognition using a pairwise feature selection and classification approachabstractThis paper proposes a novel approach that combines specialized pairwise classifiers trained with different feature subsets for facial expression classification. The proposed approach first detects and extracts automatically faces from images. Next, the face is split into several regular zones and textural features are extracted from each zone to capture local information. The features extracted from all zones are concatenated to model the whole face. A pairwise approach that considers all pairs of classes and a hybrid feature selection strategy is used to both reduce the dimensionality and to select relevant features to discriminate between specific pairs of classes. Several pairwise classifiers are then trained with such pairwise feature subsets. At the end, given a new face image, all features are extracted from such a face, but only the previously selected subset of features is inputted to each pairwise classifier. The output of all pairwise classifiers is combined using a majority voting rule to decide on the facial expression. Experiments have been carried out on three publicly available datasets (JAFFE, CK and TFEID) and the correct classification rates of 99.05%, 98.07% and 99.63% were achieved respectively. Therefore, the pairwise approach is effective to discriminate between different facial expressions and the results achieved by the proposed approach are slightly better than several current approaches. Marcelo J. Cossetin, Júlio C. Nievola, Alessandro L. Koerich |
IJCNN | 2 |
| 2016 | A constructive algorithm for neural networks inspired on decision trees and evolutionary algorithmsabstractInspired on decision trees and evolutionary algorithms, this paper proposes a learning algorithm of constructive neural networks that relies on three principles: to layout the neurons in a tree-like structure; to train each neuron individually; and, to optimize all the weights using an evolutionary approach. This way, it is expected to advance in two main questions concerning multilayer perceptrons (MLPs): how to determine the network architecture and how to build models that are more comprehensible. Based on the normalized information gain of each attribute, the algorithm builds the network architecture. In the process, it automatically creates a set of training examples for each individual neuron and executes single-cell learning. Once the network is created and trained, particle swarm optimization is utilized to evolve the connections of the network. Five metrics were utilized to validate the method when compared to decision trees and MLPs: accuracy, sensitivity, specificity, precision and comprehensibility. The experiments were executed in thirteen different databases and the results suggest that the proposed algorithm can generate neural networks with good classification performance and more comprehensible. Marcus Vinícius Mazega Figueredo, Emerson Cabrera Paraiso, Júlio C. Nievola |
IJCNN | 3 |
| 2015 | A hierarchical neural network for predicting protein functionsabstractThis paper introduces the use of a modified feedforward neural network to cope with the problem of predicting protein functions. Since this kind of classification task is inherently hierarchical, this work proposes the use of two different architectures for the modified feedforward neural network, both mimicking the hierarchical nature of the classes (protein functions) to be predicted. The first approach consists of four feed-forward neural networks in cascade, each one taking as input the classification obtained by the previous network, which means, the input to a network is the classes that could be assigned to the protein at the immediately higher (parent) level in the class hierarchy. The second approach is an extension of the first one, which also adds as input to each sub-network the attributes of the protein being classified. In both situations, it was used two kinds of feed-forward architectures: an Adaline network, which is composed of a single layer of adjustable weights, and a MLP ("Multi-Layer Perceptron"), composed by two layers of adjustable weights. Both approaches were compared with a baseline consisting of a single MLP that maps the input attributes to the classes of the lowest level in the hierarchy. The MLP was built with the input layer, plus one hidden layer and one output layer. The three approaches were compared on eight datasets, the first four involving the prediction of GPCR (G-Protein Coupled Receptor) functions and the second four datasets involving the prediction of enzymes functions. The results show that a big-bang hierarchical neural network, based on the MLP paradigm, using a top-down evaluation for new instances has better behavior in hierarchical problems, when compared to its flat version. Júlio C. Nievola, Emerson Cabrera Paraiso, Alex Alves Freitas |
BIBE | 1 |
| 2013 | Dynamically modeling users with MODUS-SD and Kohonen's mapabstractThe lack of tools for small software development teams leads us to propose an architecture to better support it. The multi-agent system (MAS) based architecture is called CSCW-SD. CSCW-SD has a module (MODUS-SD) that models users for better system customization and usability. In this paper, we present an enhancement to this module to help dynamically modeling the users as they interact with the system. In order to do that, we introduce an algorithm called GSOM that clusters user's behavior (the extracted features) and automatically detects cluster boundaries in the resulting trained SOM network. Lucas Galete, Milton Pires Ramos, Júlio C. Nievola, Emerson Cabrera Paraiso |
CSCWD | 3 |
| 2012 | Predicting GPCR and enzymes function with a global approach based on LCSabstractThe families of G-Protein Coupled Receptor (GPCR) and enzymes are among the main protein family. They represent to the scientific and medical communities, a significant target for bioactive and drug discovery programs. The model of classification of enzymes and GPCR is characterized by its hierarchical structure in format of tree and this makes more difficult its prediction. In this work we propose an adapted version of Learning Classifier Systems (LCS) data mining algorithm which tends to be more efficient than statistical methods based on homology used in tool such as PSI-BLAST. Hence, a new global model approach, called HLCS (Hierarchical Learning Classifier System) is used to predict the function of enzymes and GPCR, respecting its organizational structure of classes throughout the model development. The HLCS is expressed as a set of IF-THEN classification rules, which have the advantage of representing comprehensible knowledge to biologist users. The HLCS is evaluated with eight datasets from enzymes and GPCR, and compared with a Global Naive Bayes algorithm, named GMNB. In the tests realized the HLCS outperformed the GMNB in the databases of the GPCR proteins group type. Luiz Melo Romão, Júlio C. Nievola |
BIBE | 2 |
| 2012 | Multi-Label Hierarchical Classification using a Competitive Neural Network for protein function predictionabstractHierarchical classification is a problem with applications in many areas as protein function prediction where the dates are hierarchically structured. Therefore, it is necessary the development of algorithms able to induce hierarchical classification models. This paper presents an algorithm for hierarchical classification using the global approach, called Multilabel Hierarchical Classification using a Competitive Neural Network (MHC-CNN). It was tested on some datasets from the bioinformatics field and its results are promising. Helyane Bronoski Borges, Júlio C. Nievola |
IJCNN | 2 |
| 2012 | New models for long-term Internet traffic forecasting using artificial neural networks and flow based informationabstractThis paper investigates the use of ensembles of artificial neural networks in predicting long-term Internet traffic. It discusses a method for collecting traffic information based on flows, obtained with the NetFlow protocol, to build the time series. It also proposes four traffic forecasting models based on ensembles of TLFNs (Time-Lagged FeedFoward Networks), each one differing from the others by the way it reads the training data and by the number of artificial neural networks used in the forecasts. The proposed prediction models are confronted with the classic method of Holt-Winters, by comparing the mean absolute percentage error (MAPE) of the forecasts. It is concluded that the proposed models perform well, and can be considered a good option for planning network links that transport Internet traffic. Marcio L. F. Miguel, Manoel Camillo Penna, Júlio C. Nievola, Marcelo Eduardo Pellenz |
NOMS | 3 |
| 2012 | Comparing the dimensionality reduction methods in gene expression databases
Helyane Bronoski Borges, Júlio C. Nievola |
Expert Syst. Appl. | 2 |
| 2012 | Predicting published news effect in the Brazilian stock market
P. S. M. Nizer, Júlio C. Nievola |
Expert Syst. Appl. | 2 |
| 2011 | Noctua: A tool for Knowledge Acquisition and Collaborative Knowledge Construction with a virtual catalystabstractThis paper presents Noctua, a tool to assist in Knowledge Acquisition and Collaborative Knowledge Construction processes. Noctua contains an innovation: a virtual catalyst designed to facilitate the task of eliciting and validating knowledge. The virtual catalyst queries collaborators, proposing new knowledge, seeking confirmation to the knowledge already elicited, and showing conflicting opinions. Noctua takes into account collaborators' profiles in order to automatically ask them questions related to each one's field of knowledge or interest. G. Boz, Milton Pires Ramos, Gilson Yukio Sato, Júlio C. Nievola, Emerson Cabrera Paraiso |
CSCWD | 4 |
| 2011 | Multiobjective Optimization of Indexes Obtained by Clustering for Feature Selection Methods Evaluation in Genes Expression Microarrays
Rodolfo Garcia, Emerson Cabrera Paraiso, Júlio C. Nievola |
IDEAL | 3 |
| 2011 | A Virtual Catalyst in the Knowledge Acquisition Process
Geraldo Boz Jr., Milton Pires Ramos, Gilson Yukio Sato, Cesar Augusto Tacla, Júlio C. Nievola, Emerson Cabrera Paraiso |
SEKE | 5 |
| 2009 | Dimensionality Reduction in Gene Expression Database through the Random Projection MethodabstractDimensionality reduction applied to gene expression is challenging for machine learning algorithms due to a small number of samples and a high number of attributes. This paper proposes a preprocessing phase by means of random projection method in microarray data. Experimental results are promising and it shows that the use of this method improves the performance of classification algorithms. Helyane Bronoski Borges, Júlio C. Nievola |
ICMLA | 2 |
| 2008 | Pattern recognition for brain-computer interface on disabled subjects using a wavelet transformationabstractThe main objective of this work is to present an exploratory approach on electroencephalographic (EEG) signal, analyzing the patterns on the time-frequency plane. This work also aims to optimize the EEG signal analysis through the improvement of classifiers and, eventually, of the BCI performance. In this paper a novel exploratory approach for data mining on EEG signal based on continuous wavelet transformation (CWT) and wavelet coherence (WC) statistical analysis is introduced and applied. The CWT allows the signal underlying information content illustration by representing time-frequency patterns on Wavelet Coherence qualitative analysis. Results suggest that the proposed methodology is capable of identifying regions on time-frequency spectrum during the specified task on BCI. Furthermore, an example of a region is identified, and the patterns were classified using a radial basis function neural network (RBF-NN). This innovative characteristic of the process justify the feasibility of the proposed approach on another data mining applications. It can open new physiologic researches on this field and researches on different non-stationary time series analysis. Thiago Bassani, Júlio C. Nievola |
CIBCB | 2 |
| 2008 | Comparative of data base evolution in rule association algorithms in incremental and conventional wayabstractMany results in the literature indicate that the incremental approach to association mining leads to gain regarding the time needed to obtain the rules, but there is no evaluation about their quality, compared to non-incremental algorithms. This paper presents the comparison of usage of two typical algorithms representing each approach: APriori andZigZag. Execution time clearly shows the advantage of incremental approaches, but when someone needs accurate results concerning the association rules obtained, the matter should be taken with more caution, because the rules obtained are not necessarily in a relation one-to-one, according to the results obtained. Euclides Peres Farias, Júlio C. Nievola |
IJCNN | 2 |
| 2007 | Statistical and Biological Validation Methods in Cluster Analysis of Gene ExpressionabstractData clustering methods have become standard techniques in the analysis of gene expression data. They are used in a variety of tasks ranging from simple data pre- treatment for posterior analysis to the identification of important information, such as gene function and/or the participation of a group of genes in a given biological process. Data clustering methods also offer advantages to the biologist from the economic point of view and given the time that would be necessary to obtain this type of information without the aid of intelligent computational methods. This work aims at guiding the choices in order to get the best possible solution from data clustering. To do so, algorithms from different approaches were used, i.e. k-means and SOM algorithms belong to the unidimentional approach and SAMBA algorithm, a bidimentional approach. Methods of statistical and biological validation were employed in order to choose the best data clustering solution. Results presented here demonstrated that the statistic validation methods were hardly in agreement with the biology validation method. Furthermore, some advantages of the SOM algorithm over the k-means algorithm were observed. Use of the bidimentional algorithm SAMBA revealed dataset structure not identified by the unidimentional algorithms. It was possible to aggregate meaningfull biological information to genes of unknown function. All the content of this work, including all the data clustering and detailed analysis are available at the URL http://www.ppgia.pucpr.br/~nievola/clusteranalysis. Daniele Yumi Sunaga, Júlio C. Nievola, Milton Pires Ramos |
ICMLA | 2 |
| 2007 | Feature Selection as a Preprocessing Step for Classification in Gene Expression DataabstractMany times, when studying gene expression data, unknown attributes, which can be redundant and even, in certain cases, irrelevant, are manipulated. The application of selection attributes algorithms as a preprocessing can help in the knowledge discovery database process. This paper is about applying selection attributes algorithms in two gene expression databases. The result shows that the use of these algorithms can improve the classification algorithms performance. Helyane Bronoski Borges, Júlio C. Nievola |
ISDA | 2 |
| 2005 | Attribute selection methods comparison for classification of diffuse large B-cell lymphomaabstractThe use of data mining techniques has helped to solve many problems in the rapidly growing field of bioinformatics. Despite that, the presence of thousands of attributes makes the results unclear and also contributes to the decrease of the accuracy of the classifier used. This paper presents a comparison of the use of various attribute selection methods aiming to reduce the number of genes to be searched. The results show that most of the combinations from search algorithms and evaluation algorithms within the attribute selection algorithm work well, reducing the number of attributes and leading to improved classification rates. Júlio C. Nievola, Helyane Bronoski Borges |
ICMLA | 1 |
| 2003 | Genetic Programming for Attribute Construction in Data Mining
Fernando E. B. Otero, Monique M. S. Silva, Alex Alves Freitas, Júlio C. Nievola |
EuroGP | 4 |
| 2002 | Constructing X-of-n Attributes With A Genetic Algorithm
Otavio Larsen, Alex Alves Freitas, Júlio C. Nievola |
GECCO | 3 |
| 2001 | Discovering Fuzzy Classification Rules with Genetic Programming and Co-evolution
Roberto R. F. Mendes, Fabricio de B. Voznika, Alex Alves Freitas, Júlio C. Nievola |
PKDD | 4 |