VLDB 2026 Research / reviewers in the wild / expert
Francisco de A. T. de Carvalho
dblp:65/6740 · also Francisco de Assis Tenório de Carvalho
· DBLP profile ↗
113ranked-venue papers
31as first author
18since 2021 · last 2025
0000-0003-1128-745XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 25 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 first-authorHuman-computer interaction and ubiquitous computing · 11 · 4 first-authorDatabases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Computer networks · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep contrastive variational subspace clustering
Marcos de Souza Oliveira, Sérgio Ricardo de Melo Queiroz, Cleber Zanchettin, Francisco de A. T. de Carvalho |
Neurocomputing | 4 |
| 2025 | Kernel clustering with automatic variable weighting for interval data
José Nataniel Andrade de Sá, Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho |
Neurocomputing | 3 |
| 2025 | A cluster-wise regression method for distribution-valued data
Antonio Balzanella, Rosanna Verde, Francisco de A. T. de Carvalho |
Knowl. Based Syst. | 3 |
| 2025 | Fuzzy and Crisp Gaussian Kernel-Based Co-Clustering With Automatic Width ComputationabstractCo-clustering algorithms separate a data matrix in blocks, by grouping, simultaneously, objects according to variables and variables according to objects, and has gained widespread attention in the last few years. At the same time, kernel-based clustering is a well-developed topic of research. These methods can efficiently group nonlinear clusters through transformations in the data space. The research involving co-clustering and kernel function is still in the initial stage. In this article, we proposed the first kernel-based algorithms that can learn the width hyperparameter of the Gaussian kernel automatically for hard and fuzzy co-clustering. The main advantages of the proposed methods are that there is no need for a previous additional step to tune the width hyperparameter, and we consider width hyperparameters that are the same for all clusters, varying only with respect to objects or variables (global methods), or they can also vary across clusters (local methods). As a consequence, our methods can rescale the objects and variables separately, according to their distribution, and in the local case, also according to the distribution in each variable cluster and object cluster, respectively. Experiments conducted over 14 real datasets, and compared with traditional clustering methods and previous state-of-the-art co-clustering algorithms, showed the efficiency of the proposed algorithms. José Nataniel Andrade de Sá, Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho |
IEEE Trans. Fuzzy Syst. | 3 |
| 2024 | Novel L1-Based Neural Gas Clustering AlgorithmsabstractClustering algorithms of the Neural Gas (NG) type take into consideration the dissimilarities between prototypes in the original input space. It has been successfully applied in vector quantization, topology creation as well as clustering. NG algorithms conventionally are based on the squared Euclidean distance or L2 distance, which has several known setbacks (not robust to noise and outliers). Our goal is to introduce new NG clustering algorithms (online and batch) based on the L1 distance (more robust to noise and outliers). We propose three Neural Gas algorithms based on the L1 distance using two different algorithms to find the optimal prototypes and compare them with another well-known clustering algorithm. Given the experiments performed, the proposed methods showed a competitive performance. Preliminary results indicate that research on Neural Gas algorithms based on L1 distance is promising. Nicomedes L. Cavalcanti, Francisco de A. T. de Carvalho |
ICMLA | 2 |
| 2024 | Self-organizing maps with adaptive distances for multiple dissimilarity matrices
Laura M. P. Mariño, Francisco de A. T. de Carvalho |
Mach. Learn. | 2 |
| 2023 | Medoid based semi-supervised fuzzy clustering algorithms for multi-view relational data
Diogo P. P. Branco, Francisco de A. T. de Carvalho |
Fuzzy Sets Syst. | 2 |
| 2023 | Gaussian kernel fuzzy c-means with width parameter computation and regularizationabstract• The paper provides fuzzy c-means algorithms based on Gaussian kernel functions. • The first algorithm computes the width parameters though suitable constraints. • The second algorithm computes the width parameters though entropy regularization . • Experiments with benchmark data sets shows the usefulness of the algorithms. The conventional Gaussian kernel fuzzy c-means clustering algorithms require selecting the width hyper-parameter, which is data-dependent and fixed for the entire execution. Not only that, but these parameters are the same for every dataset variable. Therefore, the variables have the same importance in the clustering task , including irrelevant variables. This paper proposes a Gaussian kernel fuzzy c-means with kernelization of the metric and automated computation of width parameters. These width parameters change at each iteration of the algorithm and vary from each variable and from each cluster. Thus, this algorithm can re-scale the variables differently, thus highlighting those that are relevant to the clustering task. Fuzzy clustering algorithms with regularization have become popular due to their high performance in large-scale data clustering , robustness for initialization, and low computational complexity . Because the width parameters of the variables can also be controlled by entropy, this paper also proposes Gaussian kernel fuzzy c-means algorithms with kernelization of the metric and automated computation of width parameters through entropy regularization. To demonstrate their usefulness, the proposed algorithms are compared with the conventional KFCM-K algorithm and previous algorithms that automatically compute the width parameter of the Gaussian kernel . Eduardo C. Simões, Francisco de A. T. de Carvalho |
Pattern Recognit. | 2 |
| 2022 | Kernel-based Fuzzy Co-clustering in Feature Space with Automated Variable WeightingabstractKernel functions have been used successfully in clustering algorithms to deal with the separability of clusters efficiently. Bringing this idea to co-clustering, we propose two kernel-based fuzzy co-clustering algorithms based on the fuzzy double Kmeans (FDK). The first proposed algorithm, the Gaussian kernel fuzzy double Kmeans (GKFDK), is based on FDK and computes the cluster prototypes in the original feature space. The second algorithm, the Weighted gaussian kernel fuzzy double Kmeans (WGKFDK), is an extension of the GKFDK with automated variable weighting, that distinguishes the relevance of the variables in each cluster. Experiments performed with both synthetic and real data, in comparison with previous state-of-the-art co-clustering algorithms, showed the effectiveness of the proposed algorithms. José Nataniel Andrade de Sá, Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 3 |
| 2022 | Unsupervised feature selection method based on iterative similarity graph factorization and clustering by modularity
Marcos de Souza Oliveira, Sérgio Ricardo de Melo Queiroz, Francisco de A. T. de Carvalho |
Expert Syst. Appl. | 3 |
| 2022 | Clustering interval-valued data with adaptive Euclidean and City-Block distances
Sara Inés Rizo Rodríguez, Francisco de A. T. de Carvalho |
Expert Syst. Appl. | 2 |
| 2022 | Two weighted c-medoids batch SOM algorithms for dissimilarity data
Laura M. P. Mariño, Francisco de A. T. de Carvalho |
Inf. Sci. | 2 |
| 2022 | Vector batch SOM algorithms for multi-view dissimilarity data
Laura M. P. Mariño, Francisco de A. T. de Carvalho |
Knowl. Based Syst. | 2 |
| 2021 | Weighted Clusterwise Linear Regression based on adaptive quadratic form distance
Ricardo A. M. da Silva, Francisco de A. T. de Carvalho |
Expert Syst. Appl. | 2 |
| 2021 | Interval joint robust regression method
Francisco de A. T. de Carvalho, Eufrásio de Andrade Lima Neto, Ullysses da N. Rosendo |
Neurocomputing | 1 |
| 2021 | Co-clustering algorithms for distributional data with automated variable weighting
Francisco de A. T. de Carvalho, Antonio Balzanella, Antonio Irpino, Rosanna Verde |
Inf. Sci. | 1 |
| 2021 | A clusterwise nonlinear regression algorithm for interval-valued data
Francisco de A. T. de Carvalho, Eufrásio de Andrade Lima Neto, Kassio C. F. da Silva |
Inf. Sci. | 1 |
| 2021 | Soft subspace clustering of interval-valued data with regularizations
Sara Inés Rizo Rodríguez, Francisco de A. T. de Carvalho |
Knowl. Based Syst. | 2 |
| 2020 | A new batch SOM algorithm for relational data with weighted medoidsabstractThe great majority of previous works on SOM concern quantitative vectorial data. Nowadays, relatively few SOM algorithms are able to manage relational data despite their usefulness. This paper proposes a new batch SOM algorithm for relational data with weighted medoids. The particularity of the proposed approach is to consider the cluster representatives as vectors of weights whose components measure how objects are weighted as a medoid in a given cluster. From an initial solution and for a fixed epoch and radius, the proposed training batch SOM algorithm provides a partition and cluster representatives by optimizing a suitable objective function aiming to preserve the topological properties of the data on the map. Experiments with datasets of UCI machine learning repository, in comparison with relevant medoid-based batch SOM for relational data algorithms, showed the usefulness of the proposed method. Laura M. P. Mariño, Francisco de A. T. de Carvalho |
IJCNN | 2 |
| 2020 | PSO for Fuzzy Clustering of Multi-view Relational DataabstractParticle Swarm Optimization (PSO) is a population-based meta-heuristic known for its simplicity, being successfully used in clustering task with interesting performance. Clustering of multi-view data sets has received increasing attention since it explores multiple sources or views of data sets aiming at improving clustering accuracy. Previous studies mainly focused on PSO-based clustering of single-view vector data, neither single- nor multi-view PSO-based clustering of relational received proper attention. This paper introduces a PSO-based approach to the fuzzy clustering of multi-view relational data, which can cluster data sets described by several dissimilarity matrices, each of them representing a particular view. In this work, ten fitness functions were considered, in which eight of them were adapted to deal with multi-view relational data and to consider the relevance weights of views. These fitness functions were compared to evaluate which best fit to cluster multi-view relational data. The performance and usefulness of the proposed approach, in comparison with previous single- and multi-view relational fuzzy clustering algorithms, are illustrated with several multi-view data sets. The Adjusted Rand Index (ARI) and F-measure were used to assess the quality of fuzzy partitions provided by clustering algorithms. The results have shown that the proposed methods significantly outperformed the compared algorithms in the majority of cases. Renê Pereira de Gusmão, Francisco de A. T. de Carvalho |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2019 | A new fuzzy clustering algorithm for interval-valued data based on City-Block distanceabstractInterval-valued data are needed, for example, when an object represents a group of individuals and the variables used to describe it need to assume a value which expresses the variability inherent to the description of a group. Interval-valued data arise in practical situations such as recording monthly interval temperatures at meteorological stations, daily interval stock prices, etc. In this paper is proposed a robust partitioning fuzzy clustering algorithm for interval-valued data based on adaptive City-Block distance that takes into account the relevance of the variables according to the boundaries. This distance changes at each iteration of the algorithm and is different from one cluster to another. The method optimizes an objective function by alternating three steps to compute the representatives of each group, the fuzzy partition, as well as relevance weights for the interval-valued variables for each boundary. Experiments on synthetic and real interval-valued datasets corroborate the usefulness and robustness of the proposed algorithm. Sara Inés Rizo Rodríguez, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 2 |
| 2019 | Adaptive- L_2 L 2 Batch Neural Gas
Nicomedes L. Cavalcanti, Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho |
ICANN (2) | 3 |
| 2019 | Clustering interval-valued data with automatic variables weightingabstractOver the past few years, Symbolic Data Analysis has gained popularity providing suitable methods for managing aggregated data represented by lists, intervals, histograms or even distributions. This paper proposes a partitioning clustering algorithm for interval-valued data based on the suitable adaptive Euclidean distance that takes into account the relevance of the variables according to the boundaries. The proposed distance changes at each algorithm iteration and is different from one cluster to another. The method provides a partition and a prototype for each cluster by optimizing an adequacy criterion that measures the fitting between groups and their representatives. Experiments on synthetic and real interval-valued datasets corroborate the usefulness of the proposed algorithm. Sara Inés Rizo Rodríguez, Francisco de A. T. de Carvalho |
IJCNN | 2 |
| 2019 | A Fuzzy Clustering Algorithm with Multi-medoids for Multi-view Relational Data
Eduardo C. Simões, Francisco de A. T. de Carvalho |
ISNN (1) | 2 |
| 2019 | Clustering of multi-view relational data based on particle swarm optimization
Renê Pereira de Gusmão, Francisco de A. T. de Carvalho |
Expert Syst. Appl. | 2 |
| 2018 | Fuzzy clustering Algorithm based on Adaptive City-block distance and Entropy RegularizationabstractThe Euclidean distance is traditionally used to compare the objects and the prototypes in the Fuzzy C-Means algorithms, but theoretical studies indicate that methods based on City-Block distances are more robust concerning the presence of outliers in the dataset than those based on Euclidean distances. Moreover, most often conventional Fuzzy C-Means clustering algorithms consider that all variables are equally important for the clustering task. However, in real situations, some variables may be more or less relevant or even irrelevant for clustering. This paper proposes a partitioning fuzzy clustering algorithm based on Adaptive City-block distances and entropy regularization. The proposed method optimizes an objective function by alternating three steps aiming to compute the fuzzy cluster representatives, the fuzzy partition, as well as relevance weights for the variables. Several experiments on synthetic and real-world datasets including its application to noisy image texture segmentation are presented to corroborate both clustering and robustness capabilities of the proposed algorithm over conventional approaches. Sara Inés Rizo Rodríguez, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 2 |
| 2018 | On Combining Fuzzy C-Regression Models and Fuzzy C-Means with Automated Weighting of the Explanatory VariablesabstractThis paper presents a fuzzy clusterwise regression method aiming to provide linear regression models that are based on homogeneous clusters of observations with respect to the explanatory variables and that are well fitted with respect to the response variable. To achieve this aim, this method combines Fuzzy C-Regression Models and Fuzzy C-Means with automatic computation of relevance weights to the explanatory variables. Because it learns simultaneously a prototype and a linear regression model for each cluster it is able to provide an appropriated regression model for unknown observations based on their description by the explanatory variables. We also discussed both, a heuristic procedure to automatically tune one of the hyperparameters of the proposed method in order to obtain more useful (explanation) models in a prediction task, and a way of making a fuzzy combination of each intra-cluster fitted model as a more natural and appropriate response to the problem of choosing the best regression model for a prediction task. Experiments with synthetic and real datasets corroborate the usefulness of the proposed method. Ricardo A. M. da Silva, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 2 |
| 2018 | Gaussian Kernel-Based Fuzzy Clustering with Automatic Bandwidth Computation
Francisco de A. T. de Carvalho, Lucas V. C. Santana, Marcelo Rodrigo Portela Ferreira |
ICANN (1) | 1 |
| 2018 | Fuzzy Clustering Algorithm Based on Adaptive Euclidean Distance and Entropy Regularization for Interval-Valued Data
Sara Inés Rizo Rodríguez, Francisco de A. T. de Carvalho |
ICANN (1) | 2 |
| 2018 | An exponential-type kernel robust regression model for interval-valued variables
Eufrásio de Andrade Lima Neto, Francisco de A. T. de Carvalho |
Inf. Sci. | 2 |
| 2018 | Gaussian kernel c-means hard clustering algorithms with automated computation of the width hyper-parameters
Francisco de A. T. de Carvalho, Eduardo C. Simões, Lucas V. C. Santana, Marcelo Rodrigo Portela Ferreira |
Pattern Recognit. | 1 |
| 2017 | Fuzzy clustering of multi-view relational data with pairwise constraintsabstractThvs paper presents SS-MVFCVSMdd, a semi-supervised multiview fuzzy clustering algorithm for relational data described by multiple dissimilarity matrices. SS-MVFCVSMdd provides a fuzzy partition in a predetermined number of fuzzy clusters, a representative for each fuzzy cluster, learns a relevance weight for each dissimilarity matrix, and takes into account pairwise constraints must-link and cannot-link, by optimizing a suitable objective function. Experiments with multiview real-valued data sets described by multiple dissimilarity matrices show the usefulness of the proposed algorithm. Diogo P. P. Branco, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 2 |
| 2017 | Fuzzy clustering algorithm with automatic variable selection and entropy regularizationabstractThis paper proposes a partitioning fuzzy clustering algorithm with automatic variable selection and entropy regularization. The proposed method is an iterative three steps algorithm which provides a fuzzy partition, a representative for each fuzzy cluster, and learns a relevance weight for each variable in each cluster by minimizing a suitable objective function that includes a multi-dimensional distance function as the dissimilarity measure and entropy as the regularization term. Experiments on real-world datasets corroborate the usefulness of the proposed algorithm. Sara Inés Rizo Rodríguez, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 2 |
| 2017 | On Combining Clusterwise Linear Regression and K-Means with Automatic Weighting of the Explanatory Variables
Ricardo A. M. da Silva, Francisco de A. T. de Carvalho |
ICANN (2) | 2 |
| 2017 | Multi-view hard c-means with automated weighting of views and variablesabstractMulti-View Clustering models can be viewed as a way to extract information from different data representations to improve the clustering accuracy. In multi-view clustering, some views are irrelevant and among the relevant ones, some may be more or less relevant than others. This is why the most part of existing algorithms assign a weight to each view aiming to compute its relevance in the clustering process. However very few algorithms computes also the relevance weight of variables inside each view aiming to achieve automated feature selection. This paper proposes a muti-view hard c-means clustering algorithm with automated computation of weights for both views and variables in such a way that the relevant views as well as the relevant variables in each view are selected for clustering. Compared to previous similar works, an advantage of the proposed method is that, apart the need to know previously the number of clusters, there are no additional parameters to tune. Experiments with benchmark data sets corroborate the usefulness of the proposed method. Rodrigo C. de Araujo, Francisco de A. T. de Carvalho, Yves Lechevallier |
IJCNN | 2 |
| 2017 | A robust regression method based on exponential-type kernel functions
Francisco de A. T. de Carvalho, Eufrásio de Andrade Lima Neto, Marcelo Rodrigo Portela Ferreira |
Neurocomputing | 1 |
| 2017 | Fuzzy clustering of interval-valued data with City-Block and Hausdorff distances
Francisco de A. T. de Carvalho, Eduardo C. Simões |
Neurocomputing | 1 |
| 2017 | Fuzzy clustering of distributional data with automatic weighting of variable components
Antonio Irpino, Rosanna Verde, Francisco de A. T. de Carvalho |
Inf. Sci. | 3 |
| 2017 | Nonlinear regression applied to interval-valued data
Eufrásio de Andrade Lima Neto, Francisco de A. T. de Carvalho |
Pattern Anal. Appl. | 2 |
| 2016 | A Gaussian Kernel-based Clustering Algorithm with Automatic Hyper-parameters Computation
Francisco de A. T. de Carvalho, Marcelo Rodrigo Portela Ferreira, Eduardo C. Simões |
ISNN | 1 |
| 2016 | Particle Swarm Optimization applied to relational data clusteringabstractThis work introduces a hard clustering algorithm based on Particle Swarm Optimization metaheuristic that is able to partition objects considering their relational descriptions given by a single dissimilarity matrix. The PSO is a metaheuristic based on population which is well known for its simplicity, good performance and it was already designed as clustering algorithm for vector data. The proposed PSO algorithm uses a modified version of the HCMdd algorithm as local search. The HCMdd algorithm is a variant of the well known hard K-medoids clustering algorithm for relational data, that is designed to provide a partition and a representative for each cluster. The performance and the usefullness of the proposed algorithm, in comparison with HCMdd, RHCM and Spectral clustering algorithms, these last two are also able to work with relational data, are illustrated with suitable normalized data sets from the UCI Machine Learning Repository. Renê Pereira de Gusmão, Francisco de A. T. de Carvalho |
SMC | 2 |
| 2016 | Batch SOM algorithms for interval-valued data with automatic weighting of the variables
Francisco de A. T. de Carvalho, Patrice Bertrand, Eduardo C. Simões |
Neurocomputing | 1 |
| 2016 | Kernel-based hard clustering methods with kernelization of the metric and automatic weighting of the variables
Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho, Eduardo C. Simões |
Pattern Recognit. | 2 |
| 2016 | Guest Editorial Special Issue on Granular/Symbolic Data ProcessingabstractGranular/symbolic data processing is an emerging conceptual and computing paradigm of information processing. In the era of big data, the emergence of granular/symbolic processing has been motivated by the urgent need for intelligent transformation of empirical data that are now commonly available in vast quantities, into a human-manageable knowledge. In such an aggregation process, we hope to retain as much information as possible while making the findings easily understood and well-supported by the existing experimental evidence. Those aggregated entities are often referred to as symbolic or granular data. Research areas referred to as symbolic data analysis in statistics and multivariate data analysis address some of the fundamental or applied facets of granular computing. The theoretical fundamentals of granular/symbolic data processing are well-established. They involve set theory (interval mathematics), fuzzy sets, rough sets, and random sets linked together in a highly comprehensive treatment of this emerging paradigm. In addition to interval-based formalism of information granules, we also encounter histograms, distributions, lists of values, etc. Hence, granular/symbolic data processing hinges on a general computation theory that effectively uses granules such as classes, clusters, subsets, groups, and intervals to build an efficient computational model for complex applications realized in the presence of huge amounts of data, information, and knowledge. This research arises as a substantial shift from the current machine-centric to human-centric approach to information and knowledge. Shun-Feng Su, Witold Pedrycz, Tzung-Pei Hong, Francisco de A. T. de Carvalho |
IEEE Trans. Cybern. | 4 |
| 2015 | Fuzzy clustering of distribution-valued data using an adaptive L2 Wasserstein distanceabstractIn this paper, a fuzzy c-means algorithm based on an adaptive L2-Wasserstein distance for histogram-valued data is proposed. The adaptive distance induces a set of weights associated with the components of histogram-valued data and thus of the variables. The minimization of the criterion in the fuzzy c-means algorithm is performed according three steps such that the representation, the allocation and the weights associated to the components of the variables are alternately computed until a the convergence of the solution to a local optimum. The effectiveness of the proposed algorithm is demonstrated through experiments with synthetic and real-world datasets. Francisco de A. T. de Carvalho, Antonio Irpino, Rosanna Verde |
FUZZ-IEEE | 1 |
| 2015 | Fuzzy co-clustering with automated variable weightingabstractWe propose two fuzzy co-clustering algorithms based on the double Kmeans algorithm. Fuzzy approaches are known to require more computation time than hard ones but the fuzziness principle allows a description of uncertainties that often appears in real world applications. The first algorithm proposed, fuzzy double Kmeans (FDK) is a fuzzy version of double Kmeans (DK). The second algorithm, weighted fuzzy double Kmeans (W-FDK), is an extension of FDK with automated variable weighting allowing co-clustering and feature selection simultaneously. We illustrate our contribution using Monte Carlo simulations on datasets with different parameters and real datasets commonly used in the co-clustering context. Charlotte Laclau, Francisco de A. T. de Carvalho, Mohamed Nadif |
FUZZ-IEEE | 2 |
| 2015 | A multi-view relational fuzzy c-medoid vectors clustering algorithm
Francisco de A. T. de Carvalho, Filipe M. de Melo, Yves Lechevallier |
Neurocomputing | 1 |
| 2014 | An adjustable p-exponential clustering algorithm
Valmir Macario, Francisco de A. T. de Carvalho |
ESANN | 2 |
| 2014 | A kernel k-means clustering algorithm based on an adaptive Mahalanobis kernelabstractIn this paper, a kernel k-means algorithm based on an adaptive Mahalanobis kernel is proposed. This kernel is built based on an adaptive quadratic distance defined by a symmetric positive definite matrix that changes at each algorithm iteration and takes into account the correlations between variables, allowing the discovery of clusters with non-hyperspherical shapes. The effectiveness of the proposed algorithm is demonstrated through experiments with synthetic and benchmark datasets. Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho |
IJCNN | 2 |
| 2014 | Dynamic clustering of histogram data based on adaptive squared Wasserstein distances
Antonio Irpino, Rosanna Verde, Francisco de A. T. de Carvalho |
Expert Syst. Appl. | 3 |
| 2014 | Kernel fuzzy c-means with automatic variable weighting
Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho |
Fuzzy Sets Syst. | 2 |
| 2014 | Kernel-based hard clustering methods in the feature space with automatic variable weighting
Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho |
Pattern Recognit. | 2 |
| 2013 | Semi-supervised fuzzy c-medoids clustering algorithm with multiple prototype representationabstractSemi-supervised clustering is a special form of classification that uses a large amount of unlabeled data together with labeled data to achieve better classification results. This paper introduces a semi-supervised fuzzy clustering algorithm of relational data with multiple prototype representation (SS-CLAMP) that aims to furnish a partition and a set of prototypes for each fuzzy cluster as well as to learn a relevance weight for each dissimilarity matrix by optimizing an adequacy criterion that measures the fit between the fuzzy clusters and their representatives in a competitive way and that takes into account pairwise constraints must-link and cannot-link. Experiments with real-valued data sets show the usefulness of the proposed algorithm. Filipe M. de Melo, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 2 |
| 2013 | Batch self-organizing maps for mixed feature-type symbolic dataabstractThe Kohonen Self-Organizing Map (SOM) is an unsupervised neural network method with a competitive learning strategy which has both clustering and visualization properties. In this paper, we present batch SOM algorithms based on adaptive and non-adaptive distances for mixed feature-type symbolic data that, for a fixed epoch, optimize a cost function. The performance, and usefulness of these SOM algorithms are illustrated with real mixed feature-type symbolic data sets. Francisco de A. T. de Carvalho, Gibson B. N. Barbosa |
IJCNN | 1 |
| 2013 | Relational partitioning fuzzy clustering algorithms based on multiple dissimilarity matrices
Francisco de A. T. de Carvalho, Yves Lechevallier, Filipe M. de Melo |
Fuzzy Sets Syst. | 1 |
| 2013 | Nonlinear multicriteria clustering based on multiple dissimilarity matrices
Sérgio Ricardo de Melo Queiroz, Francisco de A. T. de Carvalho, Yves Lechevallier |
Pattern Recognit. | 2 |
| 2012 | A fuzzy clustering algorithm based on adaptive city-block distancesabstractThis paper gives an adaptive version of the fuzzy clustering algorithm based on city-block distances. The proposed method gives a fuzzy partition and a prototype for each cluster by optimizing an adequacy criterion based on an adaptive city-block distance that changes at each algorithm's iteration and is different from one cluster to another. Experiments with real data sets show the usefulness of this algorithm. Francisco de A. T. de Carvalho, Julio T. Pimentel |
FUZZ-IEEE | 1 |
| 2012 | Kernel fuzzy clustering methods based on local adaptive distancesabstractThis paper presents kernel fuzzy clustering methods in which dissimilarity measures are obtained as sums of squared Euclidean distances between patterns and centroids computed individually for each variable by means of kernel functions. The advantage of the proposed approach over the conventional kernel clustering methods is that it allows us to use adaptive distances which changes at each algorithm iteration and can be different from one cluster to another. This kind of dissimilarity measure is suitable to learn the weights of the variables during the clustering process, improving the performance of the algorithms. Another advantage of this approach is that it allows the introduction of various fuzzy partition and cluster interpretations tools. Experiments with benchmark data sets illustrate the usefulness of our algorithms and the merit of the fuzzy partition and cluster interpretation tools. Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 2 |
| 2012 | An adaptive semi-supervised fuzzy clustering algorithm based on objective function optimizationabstractSemi-supervised learning uses large amount of unlabeled data, combined with the labeled data, to guide the learning process. This paper introduces a new semi-supervised clustering algorithm based on an adaptive distance. The proposed method furnishes a fuzzy partition and a prototype for each cluster by optimizing a criterion based on an adaptive distance allowing the construction of partitions in ellipsoids format, in addition to spherical shape generated by the Euclidean distance. Experiments with real and synthetic data sets show the usefulness of the proposed method by comparing with others adaptive and non-adaptive semi-supervised clustering algorithms in a clustering task. Valmir Macario, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 2 |
| 2012 | Multicriteria clustering with weighted Tchebycheff distances for relational dataabstractWe present a new algorithm capable of partitioning sets of objects by taking simultaneously into account their relational descriptions given by multiple dissimilarity matrices. The algorithm uses a nonlinear aggregation criterion, weighted Tchebycheff distances, more appropriate than linear combinations (such as weighted averages) for the construction of compromise solutions. We obtain a partition of the set of objects, the prototype of each cluster and a weight vector that indicates the relevance of each criterion in each cluster. Since this is a clustering algorithm for relational data, it is compatible with any distance function used to measure the dissimilarity between objects. Some practical applications are shown, the good results obtained indicate the interest of the presented algorithm. Sérgio Ricardo de Melo Queiroz, Francisco de A. T. de Carvalho, Yves Lechevallier |
IJCNN | 2 |
| 2012 | A pattern classifier for interval-valued data based on multinomial logistic regression modelabstractInterval-valued data arise in practical situations such as recording monthly interval temperatures at meteorological stations, daily interval stock prices, etc. This paper introduces a multinomial logistic regression method for interval-valued data in order to classify items described by interval-valued variables into a pre-defined number of a priori classes. Applications of the proposed approach on real as well as synthetic interval-valued data sets showed the usefulness of this approach. Alberto Pereira de Barros, Francisco de A. T. de Carvalho, Eufrásio de Andrade Lima Neto |
SMC | 2 |
| 2012 | Partitioning fuzzy clustering algorithms for mixed feature-type symbolic dataabstractThis paper presents partitioning fuzzy clustering algorithms for mixed feature-type symbolic data. The proposed algorithms need a previous pre-processing step in order to obtain a suitable homogenization of the mixed feature-type symbolic data into histogram-valued symbolic data. These fuzzy clustering algorithms give a fuzzy partition and a prototype for each fuzzy cluster by optimizing an adequacy criterion based on suitable adaptive and non-adaptive Euclidean distances between vectors of histogram-valued data. The adaptive Euclidean distances change at each algorithm iteration and are different from one fuzzy cluster to another. Experiments with real mixed feature-type symbolic data sets show the usefulness of these fuzzy clustering algorithms. Francisco de A. T. de Carvalho, Lucas F. S. Cambuim |
SMC | 1 |
| 2012 | Partitioning fuzzy clustering algorithms for interval-valued data based on Hausdorff distancesabstractThis paper presents partitioning fuzzy clustering algorithms for interval-valued data. These fuzzy clustering algorithms give a fuzzy partition and a prototype for each fuzzy cluster by optimizing an adequacy criterion based on suitable adaptive and non-adaptive Hausdorff distances between vectors of intervals. The adaptive Hausdorff distances change at each algorithm iteration and are different from one fuzzy cluster to another. Experiments with real interval-valued data sets show the usefulness of these fuzzy clustering algorithms. Francisco de A. T. de Carvalho, Julio T. Pimentel |
SMC | 1 |
| 2012 | Partitioning hard kernel clustering methods based on local adaptive distancesabstractThis paper presents partitioning hard kernel clustering methods in which dissimilarity measures are obtained as sums of squared Euclidean distances between patterns and centroids computed individually for each variable by means of kernel functions. The advantage of the proposed approach over the conventional kernel clustering methods is that it allows to learn the weights of the variables during the clustering process, improving the performance of the algorithms. Another advantage of this approach is that it allows the introduction of various partition and cluster interpretations tools. Experiments with benchmark data sets illustrate the usefulness of our algorithms and the merit of the partition and cluster interpretation tools. Marcelo Rodrigo Portela Ferreira, Francisco de A. T. de Carvalho |
SMC | 2 |
| 2012 | An adaptive isodata fuzzy clustering algorithm with partial supervisionabstractSemi-supervised learning uses large amount of unlabeled data, combined with labeled data, to guide the learning process. This paper introduces a new clustering algorithm with partial supervision based on an adaptive distance. The proposed method furnishes a fuzzy partition and a prototype for each cluster by optimizing a criterion based on an adaptive distance allowing the construction of partitions in ellipsoids format, in addition to spherical shape generated by the Euclidean distance. Experiments with real data sets show the usefulness of the proposed method by comparing with others adaptive and non-adaptive semi-supervised clustering algorithms in a clustering task. Valmir Macario, Francisco de A. T. de Carvalho |
SMC | 2 |
| 2012 | Exponential smoothing methods for forecasting bar diagram-valued time seriesabstractWhen a set of categories with related frequencies of the observed variable is available for each time point we have a bar diagram-valued time series. This paper introduces exponential smoothing methods to forecast bar diagram-valued time series data. The proposed method is inspired in the approach introduced by Maia and De Carvalho (2011) to deal with inteval-valued time series. The smoothing parameters are estimated by using techniques for non-linear optimization problems with bound constraints. The results are discussed based on two wellknown classical performance measurements, which have been adapted here for this particular type of data: the U of Theil statistics and average relative variance (ARV) in the framework of a Monte Carlo experiment. The synthetic data sets take into account differents aspects, e.g., sample size and forecast horizons among others. Applications using real bar diagram-valued time series also were considered to demonstrate the practicality of the methods. The results demonstrate that the proposed approaches are useful in forecasting bar diagram-valued times series. C. A. G. de Araujo Junior, Francisco de A. T. de Carvalho, André Luis Santiago Maia |
SMC | 2 |
| 2012 | Inferring epigenetic and transcriptional regulation during blood cell development with a mixture of sparse linear modelsabstractMOTIVATION: Blood cell development is thought to be controlled by a circuit of transcription factors (TFs) and chromatin modifications that determine the cell fate through activating cell type-specific expression programs. To shed light on the interplay between histone marks and TFs during blood cell development, we model gene expression from regulatory signals by means of combinations of sparse linear regression models. RESULTS: The mixture of sparse linear regression models was able to improve the gene expression prediction in relation to the use of a single linear model. Moreover, it performed an efficient selection of regulatory signals even when analyzing all TFs with known motifs (>600). The method identified interesting roles for histone modifications and a selection of TFs related to blood development and chromatin remodelling. AVAILABILITY: The method and datasets are available from http://www.cin.ufpe.br/~igcf/SparseMix. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thaís Gaudencio do Rêgo, Helge G. Roider, Francisco de A. T. de Carvalho, Ivan G. Costa |
Bioinform. | 3 |
| 2012 | Partitioning hard clustering algorithms based on multiple dissimilarity matrices
Francisco de A. T. de Carvalho, Yves Lechevallier, Filipe M. de Melo |
Pattern Recognit. | 1 |
| 2011 | Adaptive Batch SOM for Multiple Dissimilarity Data TablesabstractThis paper introduces a clustering algorithm based on batch Self-Organizing Maps to partition objects taking into account their relational descriptions given by multiple dissimilarity matrices. The presented approach provides a partition of the objects and a prototype for each cluster, moreover the method is able to learn relevance weights for each dissimilarity matrix by optimizing an adequacy criterion that measures the fit between clusters and the respective prototypes. These relevance weights change at each iteration and are different from one cluster to another. Anderson B. dos S. Dantas, Francisco de A. T. de Carvalho |
ICTAI | 2 |
| 2011 | A batch self-organizing maps algorithm based on adaptive distancesabstractClustering methods aims to organize a set of items into clusters such that items within a given cluster have a high degree of similarity, while items belonging to different clusters have a high degree of dissimilarity. The self-organizing map (SOM) introduced by Kohonen is an unsupervised competitive learning neural network method which has both clustering and visualization properties, using a neighborhood lateral interaction function to discover the topological structure hidden in the data set. In this paper, we introduce a batch self-organizing map algorithm based on adaptive distances. Experimental results obtained in real benchmark datasets show the effectiveness of our approach in comparison with traditional batch self-organizing map algorithms. Luciano D. S. Pacífico, Francisco de A. T. de Carvalho |
IJCNN | 2 |
| 2011 | Predicting gene expression in T cell differentiation from histone modifications and transcription factor binding affinities by linear mixture modelsabstractBACKGROUND: The differentiation process from stem cells to fully differentiated cell types is controlled by the interplay of chromatin modifications and transcription factor activity. Histone modifications or transcription factors frequently act in a multi-functional manner, with a given DNA motif or histone modification conveying both transcriptional repression and activation depending on its location in the promoter and other regulatory signals surrounding it. RESULTS: To account for the possible multi functionality of regulatory signals, we model the observed gene expression patterns by a mixture of linear regression models. We apply the approach to identify the underlying histone modifications and transcription factors guiding gene expression of differentiated CD4+ T cells. The method improves the gene expression prediction in relation to the use of a single linear model, as often used by previous approaches. Moreover, it recovered the known role of the modifications H3K4me3 and H3K27me3 in activating cell specific genes and of some transcription factors related to CD4+ T differentiation. Ivan G. Costa, Helge G. Roider, Thaís Gaudencio do Rêgo, Francisco de A. T. de Carvalho |
BMC Bioinform. | 4 |
| 2011 | Symbolic data analysis tools for recommendation systems
Byron L. D. Bezerra, Francisco de A. T. de Carvalho |
Knowl. Inf. Syst. | 2 |
| 2010 | A new approach for semi-supervised clustering based on Fuzzy C-MeansabstractIn traditional machine learning applications, only labeled data is used to train the classifier. Labeled data are difficult, expensive, time-consuming and require human experts to be obtained in several real applications. Semi-supervised learning address this issue. Semi-supervised learning uses large amount of unlabeled data, combined with the labeled data, to build better classifiers. The semi-supervised algorithm could be an extension of an unsupervised algorithm. Such algorithm would be based on unsupervised clustering algorithms, adding a term in its objective function that makes use of labeled information to guide the learning process. This study presents a new algorithm for semi-supervised clustering based on Fuzzy C-Means algorithm. The classifier was evaluated and compared against two semi-supervised clustering algorithms in the context of learning from partially labeled data. The behavior of the proposed algorithm is discussed and the results are validated using cross-validation and the confidence interval. Thus, it was possible to certify the better accuracy performance of the new algorithm when a few labeled data are available. Valmir Macario, Francisco de A. T. de Carvalho |
FUZZ-IEEE | 2 |
| 2010 | A relational fuzzy c-means clustering algorithm based on multiple dissimilarity matricesabstractThis paper introduces a relational fuzzy c-means clustering algorithm that is able to partition objects taking into account simultaneously several dissimilarity matrices. The aim is to obtain a collaborative role of the different dissimilarity matrices in order to obtain a final consensus partition. These matrices could have been obtained using different sets of variables and dissimilarity functions. This algorithm is designed to give a fuzzy partition and a prototype for each cluster as well as to learn a relevance weight for each dissimilarity matrix by optimizing an objective function. These relevance weights change at each algorithm's iteration and are different from one cluster to another. Experiments with datasets from UCI machine learning repository show the usefulness of the proposed algorithm. Francisco de A. T. de Carvalho, Filipe M. de Melo, Yves Lechevallier |
ISDA | 1 |
| 2010 | Fuzzy K-means clustering algorithms for interval-valued data based on adaptive quadratic distances
Francisco de A. T. de Carvalho, Camilo P. Tenorio |
Fuzzy Sets Syst. | 1 |
| 2010 | Unsupervised pattern recognition models for mixed feature-type symbolic data
Francisco de A. T. de Carvalho, Renata M. C. R. de Souza |
Pattern Recognit. Lett. | 1 |
| 2009 | An Analysis of Meta-learning Techniques for Ranking Clustering Algorithms Applied to Artificial Data
Rodrigo G. F. Soares, Teresa Bernarda Ludermir, Francisco de A. T. de Carvalho |
ICANN (1) | 3 |
| 2009 | Bivariate Generalized Linear Model for Interval-Valued VariablesabstractCurrent symbolic regression methods visualize problems from an optimization point of view and do not consider the probabilistic aspects related to regression models. In this paper, we present the bivariate generalized linear model (BGLM) proposed by Iwasaki and Tsubaki [5] in the context of interval-valued data sets. Important aspects related to the BGLM that remain open or can be improved will be considered. The performance of this new approach in relation to symbolic regression methods proposed by Billard and Diday [1] and Lima Neto and De Carvalho [7] will be considered through real interval data sets. Eufrásio de Andrade Lima Neto, Gauss M. Cordeiro, Francisco de A. T. de Carvalho, Ulisses Umbelino dos Anjos, Abner Gomes da Costa |
IJCNN | 3 |
| 2009 | Clustering of symbolic data using the assignment-prototype algorithmabstractThis paper shows a fuzzy relational clustering method in order to perform the clustering of symbolic data. The presented method yields a fuzzy partition and prototype for each cluster by optimizing an adequacy criterion based on suitable dissimilarity measures. This work considers two volume-based measures that may be applied to data described by set-valued, list-valued or interval-valued symbolic variables. Experiments with real and synthetic symbolic data sets show the usefulness of the proposed approach. The accuracy of the results were assessed by the corrected Rand index and the overall error rate of classification. Kelly P. Silva, Francisco de A. T. de Carvalho, Marc Csernel |
IJCNN | 2 |
| 2009 | Partitional clustering algorithms for symbolic interval data based on single adaptive distances
Francisco de A. T. de Carvalho, Yves Lechevallier |
Pattern Recognit. | 1 |
| 2009 | Clustering constrained symbolic data
Francisco de A. T. de Carvalho, Marc Csernel, Yves Lechevallier |
Pattern Recognit. Lett. | 1 |
| 2009 | Dynamic Clustering of Interval-Valued Data Based on Adaptive Quadratic DistancesabstractThis paper presents partitioning dynamic clustering methods for interval-valued data based on suitable adaptive quadratic distances. These methods furnish a partition and a prototype for each cluster by optimizing an adequacy criterion that measures the fitting between the clusters and their representatives. These adaptive quadratic distances change at each algorithm iteration and can either be the same for all clusters or different from one cluster to another. Moreover, various tools for the partition and cluster interpretation of interval-valued data are also presented. Experiments with real and synthetic interval-valued data sets show the usefulness of these adaptive clustering methods and the merit of the partition and cluster interpretation tools. Francisco de A. T. de Carvalho, Yves Lechevallier |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2008 | A Weighted Partitioning Dynamic Clustering Algorithm for Quantitative Feature Data Based on Adaptive Euclidean DistancesabstractThis paper introduces a weighted partitioning dynamic clustering algorithm for quantitative feature data based on adaptive euclidean distances. The proposed method is an iterative four-steps relocation algorithm involving the determination of the clusters representatives (prototypes), the weight of each individual, the distance associated to each cluster and the construction of the clusters, at each iteration. Moreover, the algorithm furnishes automatically the best weight of each individual in such a way that as close it is an individual from the prototype of the cluster it belongs as high it is its weight. Experiments with real and synthetic datasets show the usefulness of the proposed method. Francisco de A. T. de Carvalho, Luciano D. S. Pacífico |
HIS | 1 |
| 2008 | Neural Networks and Exponential Smoothing Models for Symbolic Interval Time Series Processing - Applications in Stock MarketabstractThe need to consider data that contain information that cannot be represented by classical models has led to the development of symbolic data analysis (SDA). As a particular case of symbolic data, symbolic interval time series are interval-valued data which are collected in a chronological sequence through time. This paper presents two approaches to symbolic interval time series analysis. The first approach is based on artificial neural networks. The second, is a new model based on exponential smoothing methods, where the smoothing parameters are estimated by using techniques for nonlinear optimization problems with bound constraints. The practicality of the methods is demonstrated by applications on real interval time series. André Luis Santiago Maia, Francisco de A. T. de Carvalho |
HIS | 2 |
| 2008 | Clustering of symbolic data through a dissimilarity volume based measureabstractThe recording of symbolic data has become a common practice with the advances in database technologies. This paper shows hard and fuzzy relational clustering in order to partition symbolic data. These methods optimize objective functions based on a dissimilarity function. The distance used is a volume based measure and may be applied to data described by set-valued, list-valued or interval-valued symbolic variables. Experiments with real and synthetic symbolic data sets show the usefulness of the proposed approach. Kelly P. Silva, Francisco de A. T. de Carvalho, Marc Csernel |
IJCNN | 2 |
| 2008 | Evolving both size and accuracy of RBF networks using Memetic AlgorithmabstractOne of the main obstacles to obtain an artificial neural network with reasonable performance is the parameter setting. This work proposes a methodology to the automatic definition of RBF (radial basis function) networks with an appropriate configuration for the selected classification problems. We propose the use of a memetic algorithm in order to perform the search for networks with minimum architecture and error rate. A set of experiments was made with four datasets and we were able to show the effectiveness of the method. Kelly P. Silva, Rodrigo G. F. Soares, Francisco de A. T. de Carvalho, Teresa Bernarda Ludermir |
IJCNN | 3 |
| 2008 | An evolutionary approach for the clustering data problemabstractThe clustering problem consists in the discovery of interesting groups in a dataset. Such task is very important and widely tackled in the literature. In this paper, we propose an evolutionary method in order to obtain well formed and spatially separated clusters. The proposed algorithm uses a complete solution representation, each partition is represented by a length-variablechromosome. The variation operators were chosen to facilitate the exchange of clustering information between individuals. We have put two complementary clustering criteria together in the fitness function, so that the method can find clusters with arbitrary shapes. The k-means algorithm was the basis of the local search operator, such operator might refine the clustering solutions. The population diversity was an important issue for the algorithm, so a diversity maintenance scheme was employed. Differently from other existing clustering algorithms, our algorithm does not need the setting of the number of clusters in advance. We evaluated the method in different contexts, using both real and simulated data. Rodrigo G. F. Soares, Kelly P. Silva, Teresa Bernarda Ludermir, Francisco de A. T. de Carvalho |
IJCNN | 4 |
| 2008 | Nonlinear regression model to symbolic interval-valued variablesabstractThis paper introduces a nonlinear regression method to fit a regression model to symbolic interval-valued data set. The nonlinear method will be inspired in the method proposed by and will consider two independent nonlinear regression models fitted over the midpoint and range of the intervals. The assessment of the proposed prediction methods is based on the average behavior of the root mean square error and of the square of the correlation coefficient in the framework of a Monte Carlo experiment. The synthetic data sets taking into account the different degree of nonlinearity between the dependent and the independent interval variables, among others aspects. Eufrásio de Andrade Lima Neto, Francisco de A. T. de Carvalho |
SMC | 2 |
| 2008 | Forecasting models for interval-valued time series
André Luis Santiago Maia, Francisco de A. T. de Carvalho, Teresa Bernarda Ludermir |
Neurocomputing | 2 |
| 2007 | Application of a Hybrid Classifier to the Recognition of Petrochemical OdorsabstractNowadays there are several data mining algorithms applied to the resolution of many different problems, such as the classification of patterns. However, when these algorithms are used separately to classify they usually present an inferior performance compared to the performance obtained by combined models. The Bagging and Boosting techniques combine models of the same kind in a competitive form, in other words, the output is generally provided by the winning classifier. Alternatively, Stacking usually combines different algorithms, constituting a hybrid model. Nevertheless, stacking has a high cost, due to the search for the best models that will be combined to solve a certain problem. Thus, we present a Hybrid Classifier (HC) to be applied to the recognition of gases derived from petrol at a lower cost and in a cooperative way. Eleonora Ma. Jesus Oliveira, Paulemir G. Campos, Teresa Bernarda Ludermir, Francisco de A. T. de Carvalho, Wilson Rosa de Oliveira |
HIS | 4 |
| 2007 | A Clustering Method for Mixed Feature-Type Symbolic Data using Adaptive Squared Euclidean DistancesabstractThis work presents a clustering method for mixed feature-type symbolic data. The presented method needs a previous pre-processing step to transform mixed symbolic data into modal symbolic data. The dynamic clustering algorithm with adaptive distances has then as input a set of vectors of modal symbolic data (weight distributions) and furnishes a partition and a prototype to each class by optimizing an adequacy criterion that measures the fitting between the clusters and their representatives based on adaptive squared Euclidean distances. Examples with synthetic symbolic data sets and an application with a real symbolic data sets show the usefulness of this method. Renata M. C. R. de Souza, Francisco de A. T. de Carvalho |
HIS | 2 |
| 2007 | A Partitioning Fuzzy Clustering Algorithm for Symbolic Interval Data based on Adaptive Mahalanobis DistancesabstractThe recording of symbolic interval data has become a common practice with the recent advances in database technologies. This paper introduces a fuzzy clustering algorithm to partitioning symbolic interval data. The proposed method furnish a fuzzy partition and a prototype (a vector of intervals) for each cluster by optimizing an adequacy criterion that measures the fitting between the clusters and their representatives. To compare symbolic interval data, the method use a suitable adaptive Mahalanobis disance defined on vectors of intervals. Experiments with real and synthetic symbolic interval data sets showed the usefulness of the proposed method. Camilo P. Tenorio, Francisco de A. T. de Carvalho, Julio T. Pimentel |
HIS | 2 |
| 2007 | Clustering of symbolic interval data based on a single adaptive L1 distanceabstractThe recording of symbolic interval data has become a common practice with the recent advances in database technologies. This paper introduces a dynamic clustering method to partitioning symbolic interval data. This method furnishes a partition and a prototype for each cluster by optimizing an adequacy criterion that measures the fitting between the clusters and their representatives. To compare symbolic interval data, the method uses a single adaptive L1distance that at each iteration changes but is the same for all the clusters. Experiments with real and synthetic symbolic interval data sets showed the usefulness of the proposed method. Francisco de A. T. de Carvalho, Julio T. Pimentel, Lucas X. T. Bezerra |
IJCNN | 1 |
| 2007 | Inequality Constraints in Regression Models to Symbolic Interval VariablesabstractThis paper introduces some approaches to fitting a constrained linear regression model to interval-valued data. The new methods show the importance of the range's information in their prediction performance and the use of inequality constraints to guarantee mathematical coherence between the predicted values of the lower bound (gammaUi) and the upper bound (ŷLi)-The authors also propose expressions to the goodness of fit measure calleddetermination coefficient. The assessment of the proposed prediction methods is based on the estimation of the average behaviour of theroot mean square errorand of thesquare of the correlation coefficientin the framework of a Monte Carlo experiment with differents data sets configurations. Finally, the approaches proposed in this paper are applied in a real data-set. Eufrásio de Andrade Lima Neto, Francisco de A. T. de Carvalho, Jose F. Coelho Neto |
IJCNN | 2 |
| 2007 | Analyzing Distance Measures for Symbolic Data Based on Fuzzy ClusteringabstractVarious propositions to solve the problem of symbolic data clustering are available in the literature. This paper introduces a comparative study among some well known dissimilarity functions treating symbolic data. An extension of the fuzzy c-means clustering algorithm is used to create groups of individuals characterized by symbolic variables of mixed types. The proposed method furnishes a fuzzy partition and a prototype for each cluster by optimizing a criterion dependent on the dissimilarity function. Experiments involving benchmark data sets are carried out in order to compare the accuracy of each function. Alzennyr Da Silva, Yves Lechevallier, Francisco de A. T. de Carvalho |
ISDA | 3 |
| 2007 | Construction and Analysis of Evolving Data Summaries: An Application on Web Usage DataabstractTaking the temporal dimension into account during the analysis of Web usage data has become a necessity since the way a site is visited may well evolve due to modifications in the structure and content of the site, or even due to changes in the behavior of certain user groups. Consequently, the models associated with these behaviors must be continuously updated. One solution to this problem is to update these models using summaries obtained by means of an evolutionary approach based on clustering methods. To do this, we carry out various clustering strategies that are applied on time sub-periods. We compare the results obtained using this method with those reached by traditional global analysis. Alzennyr Da Silva, Yves Lechevallier, Fabrice Rossi, Francisco de A. T. de Carvalho |
ISDA | 4 |
| 2007 | Clustering symbolic interval data based on a single adaptive hausdorff distanceabstractThe recording of symbolic interval data has become popular with the recent advances in database technologies. This paper introduces a dynamic clustering method to partitioning symbolic interval data. This method furnishes a partition and a prototype for each cluster by optimizing an adequacy criterion that measures the fitting between the clusters and their representatives. To compare symbolic interval data, the method uses a single adaptive Hausdorff distance that at each iteration changes but is the same for all the clusters. Experiments with real and synthetic symbolic interval data sets showed the usefulness of the proposed method. Francisco de A. T. de Carvalho, Julio T. Pimentel, Lucas X. T. Bezerra, Renata M. C. R. de Souza |
SMC | 1 |
| 2007 | Constrained linear regression models for interval-valued data with dependenceabstractThis paper introduces some approaches to fitting a constrained linear regression model to interval-valued data. The use of inequality constraints guarantee mathematical coherence between the predicted values of the lower bound (y circLi) and the upper bound (y circUi). The authors also propose expressions to the goodness of fit measure called determination coefficient. The assessment of the proposed prediction methods is based on the average behaviour of the root mean square error and of the square of the correlation coefficient in the framework of a Monte Carlo experiment. The synthetic data sets takes into account the dependence or not between the midpoint and range of the intervals, among others aspects. Finally, the approaches are applied in a real data-set. Eufrásio de Andrade Lima Neto, Francisco de A. T. de Carvalho, Jose F. Coelho Neto |
SMC | 2 |
| 2007 | Fuzzy c-means clustering methods for symbolic interval data
Francisco de A. T. de Carvalho |
Pattern Recognit. Lett. | 1 |
| 2006 | Fuzzy Clustering Algorithms for Symbolic Interval Data based on L2 NormabstractThe recording of symbolic interval data has become a common practice with the recent advances in database technologies. This paper introduces fuzzy clustering algorithms to partitioning symbolic interval data. The proposed methods furnish a fuzzy partition and a prototype (a vector of intervals) for each cluster by optimizing an adequacy criterion that measures the fitting between the clusters and their representatives. To compare symbolic interval data, the methods use a suitable (adaptive and non-adaptive) L2norm defined on vectors of intervals. Experiments with real and synthetic symbolic interval data sets showed the usefulness of the proposed method. Francisco de A. T. de Carvalho, Nicomedes L. Cavalcanti |
FUZZ-IEEE | 1 |
| 2006 | A Fuzzy Clustering Algorithm for Symbolic Interval Data Based on a Single Adaptive Euclidean Distance
Francisco de A. T. de Carvalho |
ICONIP (3) | 1 |
| 2006 | A Hybrid Model for Symbolic Interval Time Series Forecasting
André Luis Santiago Maia, Francisco de A. T. de Carvalho, Teresa Bernarda Ludermir |
ICONIP (2) | 2 |
| 2006 | A Modal Symbolic Classifier for Interval Data
Fabio C. D. Silva, Francisco de A. T. de Carvalho, Renata M. C. R. de Souza, Joyce Q. Silva |
ICONIP (2) | 2 |
| 2006 | Hybrid model with dynamic architecture for forecasting time seriesabstractNonlinear artificial neural network models are very attractive for modeling and forecasting time series. The use of such models in these types of applications is motivated by experimental results that show a high capacity of approximation for functions with high accuracy. However, many researchers have used feedforward and/or backpropagation models for time series predictions. In this paper, a model is applied for neural networks with the dynamic architecture proposed by Ghiassi and Saidane (2005), known as the DAN2 model. The results of DAN2 are compared with auto-regressive integrated mobile average (ARIMA) models. As the main result of the paper, we propose a hybrid model with dynamic architecture (HAD) based on combinations of individual forecasts from the DAN2 and ARIMA models with the aim of obtaining more precise forecasts for poorly behaved time series. The results suggest that for this kind of series, the HAD hybrid model outperforms the individual DAN2 and ARIMA models. Gecynalda Soares da Silva Gomes, André Luis Santiago Maia, Teresa Bernarda Ludermir, Francisco de A. T. de Carvalho, Aluízio F. R. Araújo |
IJCNN | 4 |
| 2006 | Mining Web Usage Data for Discovering Navigation ClustersabstractThe analysis of a web site based on usage data is an important task as it provides insight into the organization of the site and its adequacy regarding user needs. This allows the relationship between prior categories and user browsing patterns to be explored. In this paper we propose an approach for discovering the profiles of visitor groups. To this end, we begin by mapping user interests into symbolic objects, which is the basis of the Symbolic Data Analysis and represents here a successful interaction of the user with the web site. We then identify groups of users with similar behavior by means of a dynamic clustering approach applying a context dependent dissimilarity measure. The method was applied to identify visitor groups of a web site in the educational domain and also to analyze the traces of different user behavior. Alzennyr Da Silva, Yves Lechevallier, Francisco de A. T. de Carvalho, Brigitte Trousse |
ISCC | 3 |
| 2006 | C^2: : A Collaborative Recommendation System Based on Modal Symbolic User ProfileabstractRecommendation Systems have become an important tool to cope with the information overload problem by acquiring information about the user behavior. However, the process of getting user personal data may vary in many different ways, and can be done implicitly (through actions) or explicitly (through rates). After tracing actions or getting rates of the user, Computational Recommendation Technologies use information filtering techniques to recommend items. In this paper we describe an approach to improve the recommendation quality in the first moments the user interacts with the system. The main idea is: (1) first of all, we describe the items with the general users opinion about them; and (2) after this, we use modal symbolic structures to save this content in the user profile. The proposed methodology outperforms, concerning the Find Good Items task measured by half-life utility metric, other approaches based on the following techniques: Cognitive Filtering, Social Filtering and hybrid methods. Byron L. D. Bezerra, Francisco de A. T. de Carvalho, Valmir Macario |
Web Intelligence | 2 |
| 2006 | Partitional fuzzy clustering methods based on adaptive quadratic distances
Francisco de A. T. de Carvalho, Camilo P. Tenorio, Nicomedes L. Cavalcanti Junior |
Fuzzy Sets Syst. | 1 |
| 2006 | Adaptive Hausdorff distances and dynamic clustering of symbolic interval data
Francisco de A. T. de Carvalho, Renata M. C. R. de Souza, Marie Chavent, Yves Lechevallier |
Pattern Recognit. Lett. | 1 |
| 2004 | Classification of SAR Images Through a Convex Hull Region Oriented Approach
Simith T. D'Oliveira Junior, Francisco de A. T. de Carvalho, Renata M. C. R. de Souza |
ICONIP | 2 |
| 2004 | Clustering of Interval-Valued Data Using Adaptive Squared Euclidean Distances
Renata M. C. R. de Souza, Francisco de A. T. de Carvalho, Fabio C. D. Silva |
ICONIP | 2 |
| 2004 | A symbolic approach for content-based information filtering
Byron L. D. Bezerra, Francisco de A. T. de Carvalho |
Inf. Process. Lett. | 2 |
| 2004 | A Modal Symbolic Classifier for selecting time series models
Ricardo B. C. Prudêncio, Teresa Bernarda Ludermir, Francisco de A. T. de Carvalho |
Pattern Recognit. Lett. | 3 |
| 2004 | Clustering of interval data based on city-block distances
Renata M. C. R. de Souza, Francisco de A. T. de Carvalho |
Pattern Recognit. Lett. | 2 |