Alneu de Andrade Lopes

dblp:60/5946 · DBLP profile ↗
← Back
42ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0003-3112-4746ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 6 since 2021Databases, data management, data science and information retrieval · 12 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Multi-view Graph Condensation via Tensor Decomposition
abstract
Graph Neural Networks (GNNs) have demonstrated remarkable results in various real-world applications, including drug discovery, object detection, social media analysis, recommender systems, and text classification. In contrast to their vast potential, training them on large-scale graphs presents significant computational challenges due to the resources required for their storage and processing. Graph Condensation has emerged as a promising solution to reduce these demands by learning a synthetic compact graph that preserves the essential information of the original one while maintaining the GNN's predictive performance. Despite their efficacy, current graph condensation approaches frequently rely on a computationally intensive bi-level optimization. Moreover, they fail to maintain a mapping between synthetic and original nodes, limiting the interpretability of the model's decisions. In this sense, a wide range of decomposition techniques have been applied to learn linear or multi-linear functions from graph data, offering a more transparent and less resource-intensive alternative. However, their applicability to graph condensation remains unexplored. This paper addresses this gap and proposes a novel method called Multi-view Graph Condensation via Tensor Decomposition (GCTD) to investigate to what extent such techniques can synthesize an informative smaller graph and achieve comparable downstream task performance. Extensive experiments on six real-world datasets demonstrate that GCTD effectively reduces graph size while preserving GNN performance, achieving up to a 4.0% improvement in accuracy on three out of six datasets and competitive performance on large graphs compared to existing approaches. Our code is available at https://github.com/nicolasrsantos/gctd.
Nícolas Roque dos Santos, Dawon Ahn, Diego Minatel, Alneu de Andrade Lopes, Evangelos E. Papalexakis
WSDM4
2025 Predicting Land Sensing Indicators With Geolocalized Complex Network
abstract
Sensing a land area and collecting the indicators data is a task that can require investment resources. However, the possibility of making a prediction based on a few indicators means a considerable saving of resources. Here, we present a technique to predict indicators using a semi-supervised graph-based regression. We present the Geolocalized Network Based on Pearson Correlation (GNB-PC) for this task. To create the network, we combine two different topologies of graphs, and each vertex stores information such as the Pearson coefficient and the discretized value of the indicator. We propose a combination of Random Walks with linear Regression and Semi-supervised Pearson correlation for the inference. We adapted the Gibbs sampling for the regression task to evaluate the method. The experiments show that our solution outperforms the baselines when employing information from the dataset with geolocalization data in the graph construction step and comparing the proposed framework with state-of-the-art baselines.
Renan Guilherme Nespolo, Alan Valejo, Alneu de Andrade Lopes
IEEE Geosci. Remote. Sens. Lett.3
2024 Semi-Supervised Coarsening of Bipartite Graphs for Text Classification via Graph Neural Network
abstract
Graph Neural Networks (GNNs) have recently received extensive attention due to their applicability in a wide range of tasks, including drug discovery, text classification, traffic forecasting, hardware design, and recommendation. However, GNNs face significant challenges regarding scalability and the ability to handle large-scale graphs. Several strategies have been proposed to address these challenges, with multilevel optimization being a prominent approach. This technique involves hierarchically generating compact graphs through a coarsening step, applying a target algorithm (e.g., community detection) to the coarsest graph, and then projecting the initial solution back to the original input to derive the final solution. In this work, we introduce a method for graph-based text classification using GNNs. Our approach involves generating ten smaller graphs from an input bipartite graph using the coarsening step within the multilevel optimization and applying a GNN to learn node representations at various levels of granularity. Moreover, we propose a novel semi-supervised coarsening algorithm called Greedy Sorted Matching using Class and Split Information for Bipartite Graphs (GMCb). GMCb leverages class and train-test split information to select document nodes to merge during the graph coarsening step. We perform three types of reductions by either coarsening only one of the partitions of the graph or both simultaneously. Our method is evaluated on eight diverse datasets using three different GNN architectures. We assess each model's performance, memory usage, and training time to understand the impacts of graph reduction. Our experiments demonstrate that contracting the document nodes can improve performance while reducing memory consumption and training time.
Nícolas Roque dos Santos, Diego Minatel, Alan Valejo, Alneu de Andrade Lopes
DSAA4
2023 DIF-SR: A Differential Item Functioning-Based Sample Reweighting Method
Diego Minatel, Antonio Rafael Sabino Parmezan, Mariana Curi, Alneu de Andrade Lopes
CIARP4
2023 Bipartite Graph Coarsening for Text Classification Using Graph Neural Networks
Nícolas Roque dos Santos, Diego Minatel, Alan Valejo, Alneu de Andrade Lopes
CIARP4
2023 Fairness-Aware Model Selection Using Differential Item Functioning
abstract
Differential Item Functioning (DIF) is a powerful tool for developing fairer tests and mitigating bias in applicant selection tests. DIF aims to detect items in a test that favor or harm groups of people based on aspects such as gender, age, and race, which should be irrelevant to the assessment. Likewise, in machine learning, selecting a model from a pool of candidates is essential to identify the one that minimizes or eliminates discriminatory effects in its decision-making process. As far as we know, research into knowledge discovery through supervised machine learning has predominantly focused on including fair-ness notions at the lowest level of the pre-processing, pattern extraction, and post-processing phases. This fact evidences a need for studies on the impact of model selection on the development of impartial models. Herein, we present a novel approach to fairness-aware model selection to fill the mentioned gap. Our proposal introduces ABC, the first group fairness metric based on DIF concepts. We experimentally evaluated our approach against two model selection strategies by employing ten datasets, six classification algorithms, one performance measure, four group fairness measures, and one statistical significance test. According to the results, our proposal stands out for achieving a trade-off between improving the sense of justice and good classifier performance. Consequently, ABC is a promising metric for selecting fairer models with high predictive power.
Diego Minatel, Antonio Rafael Sabino Parmezan, Mariana Curi, Alneu de Andrade Lopes
ICMLA4
2022 A survey of the extraction and applications of causal relations
abstract
Abstract Causationin written natural language can express a strong relationship between events and facts. Causation in the written form can be referred to as a causal relation where a cause event entails the occurrence of an effect event. A cause and effect relationship is stronger than a correlation between events, and therefore aggregated causal relations extracted from large corpora can be used in numerous applications such as question-answering and summarisation to produce superior results than traditional approaches. Techniques like logical consequence allow causal relations to be used in niche practical applications such as event prediction which is useful for diverse domains such as security and finance. Until recently, the use of causal relations was a relatively unpopular technique because the causal relation extraction techniques were problematic, and the relations returned were incomplete, error prone or simplistic. The recent adoption of language models and improved relation extractors for natural language such as Transformer-XL (Daiet al. (2019).Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860 ) has seen a surge of research interest in the possibilities of using causal relations in practical applications. Until now, there has not been an extensive survey of the practical applications of causal relations; therefore, this survey is intended precisely to demonstrate the potential of causal relations. It is a comprehensive survey of the work on the extraction of causal relations and their applications, while also discussing the nature of causation and its representation in text.
Brett Drury, Hugo Gonçalo Oliveira, Alneu de Andrade Lopes
Nat. Lang. Eng.3
2021 A graph-based approach for positive and unlabeled learning
abstract
Positive and Unlabeled Learning (PUL) uses unlabeled documents and a few positive documents for retrieving a set of "interest" documents from a text collection. Usually, PUL approaches are based on the vector space model. However, when dealing with semi-supervised learning for text classification or information retrieval, graph-based approaches have been proved to outperform vector space model-based approaches. So, in this article, a graph-based approach for PUL is proposed: Label Propagation for Positive and Unlabeled Learning (LP-PUL). The proposed framework consists of three steps: (i) building a similarity graph, (ii) identifying reliable negative documents, and (iii) performing label propagation to classify the remaining unlabeled documents as positive or negative. We carried out experiments to measure the impact of the different choices in each step of the proposed framework. We also demonstrated that the proposal surpasses the classification performance of other PUL (RC-SVM, PU-LP, and PE-PUC) or one-class learning (k-NN-based, k-Means-based, and Dense Autoencoder) algorithms in terms of F1. Considering the best results of any algorithm used in the experimental evaluation, PU-PUL can improve the classification performance from 2%, when using only 1 labeled document, to 28%, when 30 labeled documents are employed.
Julio César Carnevali, Rafael Geraldeli Rossi, Evangelos E. Milios, Alneu de Andrade Lopes
Inf. Sci.4
2021 Local-entity resolution for building location-based social networks by using stay points
abstract
The quality of a location-based social network (LBSN) is mainly related to the granularity of information on the users' location. When LBSN is built using stay points, it presents much more information since GPS logs convey more users' mobility information. However, the main challenge in building LBSN using stay points is to define local-vertices. This problem is known as local-entity resolution. This local-vertices could represent venues with semantic information like parks, restaurants, among others. The most common way to resolve local-entity is by applying clustering algorithms to group nearby stay points into local-vertices. However, in this case, only geographic information is used, which makes it very difficult to separate geographically close venues into distinct local-vertices. This paper addresses this gap and presents a novel approach that uses the coarsening stage of a multilevel optimization scheme to build LBSNs by using stay points. The experimental evaluation carried out indicates that our approach has advantages compared to usual clustering methods to represent real-world features.
Diego Minatel, Vinícius Ferreira 0001, Alneu de Andrade Lopes
Theor. Comput. Sci.3
2020 Unsupervised learning of textual pattern based on Propagation in Bipartite Graph
abstract
Graph-based algorithms have aroused considerable interests in recent years by facilitating pattern recognition and learning via information propagation process through the graph. Here, we propose an unsupervised learning algorithm based on propagatio
Thiago de Paulo Faleiros, Alan Valejo, Alneu de Andrade Lopes
Intell. Data Anal.3
2020 A benchmarking tool for the generation of bipartite network models with overlapping communities
Alan Valejo, Fabiana Góes, Luzia Romanetto, Maria Cristina Ferreira de Oliveira, Alneu de Andrade Lopes
Knowl. Inf. Syst.5
2020 A coarsening method for bipartite networks via weight-constrained label propagation
abstract
A multilevel method is a scalable strategy to solve optimization problems in large bipartite networks, which operates in three stages. Initially the input network is iteratively coarsened into a hierarchy of gradually smaller networks. Coarsening implies in collapsing vertices into so-called super-vertices which inherit properties of their originating vertices. An initial solution is obtained executing the target algorithm in the coarsest network. Finally, this solution is successively projected back over the inverse sequence of coarsened networks, up to the initial one, yielding an approximate final solution. Despite its potential applicability, the strategy faces several theoretical and practical limitations. Coarsening is usually attained following a user-defined policy to match vertices pairwise. However, the network reduction process is extremely slow and may yield degraded solutions due to propagation of poor matches. Additionally, proper parameterization of coarsening algorithms is difficult, as well as ensuring the super-vertices preserve the relevant properties. We address these issues with a near-linear complexity coarsening strategy based on weight-constrained label propagation. Our strategy collapses groups of vertices, rather than pairs, yielding faster and more extensive network reduction. Moreover, users may specify the desired size of the coarsest network and control super-vertex weights. The applicability of our solution is illustrated in multiple scenarios, namely: multilevel implementation of an existing high-cost community detection algorithm; as a direct community detection algorithm; finally, network visualization, in connection with force-directed graph drawing algorithms. Results provide empirical evidence on the potential of our proposal to foster novel applications of the multilevel method in bipartite networks.
Alan Valejo, Thiago de Paulo Faleiros, Maria Cristina Ferreira de Oliveira, Alneu de Andrade Lopes
Knowl. Based Syst.4
2018 A Comparison of Graph Construction Methods for Semi-Supervised Learning
abstract
Graph-based methods are among the most active approaches to semi-supervised learning. This occurs mainly due to their ability to deal with local and global characteristics of available data, identify classes or groups regardless the data shape, and to represent submanifold in Euclidean space. Graph-based methods are sensitive to graph construction and a challenge in the area is the construction of a graph to represent data patterns. Several unsupervised graph construction methods have been proposed for dealing with different issues. However, it lacks a detailed study that evaluates their properties and effectiveness. Here, we analyze the robustness of such methods for SSL classification. The graph construction methods analyzed include k-nearest neighbor (kNN), mutual kNN combined with minimum spanning tree (M-kNN), b-matching by belief propagation (BP), b-matching by greedy approximation and sequential kNN (S-kNN). Statistical analyses are carried out with respect to classification accuracy. We observe that robustness of the methods varies according to external factors, such as labeled data representativeness and parameters k or b, and internal factors, related to the topology of the network. Regular graph construction methods, b-matching and S-kNN, achieve the best results on classification and generate more homogeneous and sparse networks.
Lilian Berton, Alneu de Andrade Lopes, Didier Augusto Vega-Oliveros
IJCNN2
2018 The role of location and social strength for friendship prediction in location-based social networks
abstract
International audience
Jorge Carlos Valverde-Rebaza, Mathieu Roche, Pascal Poncelet, Alneu de Andrade Lopes
Inf. Process. Manag.4
2018 Word sense disambiguation: A complex network approach
abstract
In recent years, concepts and methods of complex networks have been employed to tackle the word sense disambiguation (WSD) task by representing words as nodes, which are connected if they are semantically similar. Despite the increasingly number of studies carried out with such models, most of them use networks just to represent the data, while the pattern recognition performed on the attribute space is performed using traditional learning techniques. In other words, the structural relationship between words have not been explicitly used in the pattern recognition process. In addition, only a few investigations have probed the suitability of representations based on bipartite networks and graphs (bigraphs) for the problem, as many approaches consider all possible links between words. In this context, we assess the relevance of a bipartite network model representing both feature words (i.e. the words characterizing the context) and target (ambiguous) words to solve ambiguities in written texts. Here, we focus on the semantical relationships between these two type of words, disregarding the relationships between feature words. In special, the proposed method not only serves to represent texts as graphs, but also constructs a structure on which the discrimination of senses is accomplished. Our results revealed that the proposed learning algorithm in such bipartite networks provides excellent results mostly when topical features are employed to characterize the context. Surprisingly, our method even outperformed the support vector machine algorithm in particular cases, with the advantage of being robust even if a small training dataset is available. Taken together, the results obtained here show that the proposed representation/classification method might be useful to improve the semantical characterization of written texts.
Edilson Anselmo Corrêa Júnior, Alneu de Andrade Lopes, Diego R. Amancio
Inf. Sci.2
2018 Multilevel approach for combinatorial optimization in bipartite network
abstract
Multilevel approaches aim at reducing the cost of a target algorithm over a given network by applying it to a coarsened (or reduced) version of the original network. They have been successfully employed in a variety of problems, most notably community detection. However, current solutions are not directly applicable to bipartite networks and the literature lacks studies that illustrate their application for solving multilevel optimization problems in such networks. This article addresses this gap and introduces a multilevel optimization approach for bipartite networks and the implementation of a general multilevel framework including novel algorithms for coarsening and uncorsening, applicable to a variety of problems. We analyze how the proposed multilevel strategy affects the topological features of bipartite networks and show that a controlled coarsening strategy can preserve properties such as degree and clustering coefficient centralities. The applicability of the general framework is illustrated in two optimization problems, one for solving the Barber's modularity for community detection and the second for dimensionality reduction in text classification. We show that the solutions thus obtained are statistically equivalent, regarding accuracy, to those of conventional approaches, whilst requiring considerably lower execution times.
Alan Valejo, Maria Cristina Ferreira de Oliveira, Geraldo P. R. Filho, Alneu de Andrade Lopes
Knowl. Based Syst.4
2017 A survey of the applications of Bayesian networks in agriculture
abstract
The application of machine learning to agriculture is currently experiencing a "surge of interest" from the academic community as well as practitioners from industry. This increased attention has produced a number of differing approaches that use varying machine learning frameworks. It is arguable that Bayesian Networks are particularly suited to agricultural research due to their ability to reason with incomplete information and incorporate new information. Bayesian Networks are currently underrepresented in the machine learning applied to agriculture research literature, and to date there are no survey papers that currently centralize the state of the art. The aim of this paper is rectify the lack of a survey paper in this area by providing a self-contained resource that will: centralize the current state of the art, document the historical progression of Bayesian Networks in agriculture and indicate possible future lines of research as well as providing an introduction to Bayesian Networks for researchers who are new to the area.
Brett Drury, Jorge Carlos Valverde-Rebaza, Maria-Fernanda Moura, Alneu de Andrade Lopes
Eng. Appl. Artif. Intell.4
2017 RGCLI: Robust Graph that Considers Labeled Instances for Semi-Supervised Learning
abstract
Graph-based semi-supervised learning (SSL) provides a powerful framework for the modeling of manifold structures in high-dimensional spaces. Additionally, graph representation is effective for the propagation of the few initial labels existing in training data. Graph-based SSL requires robust graphs as input for an accurate data mining task, such as classification. In contrast to most graph construction methods, which ignore the labeled instances available in SSL scenarios, a previous study proposed a graph-construction method, named GBILI, to exploit the informativeness conveyed by such instances available in a semi-supervised classification domain. Here, we have improved the method proposing an optimized algorithm referred to as Robust Graph that Considers Labeled Instances (RGCLI) for the generation of more robust graphs. The contributions of this paper are threefold: i) reduction of GBILI time complexity from quadratic to O(nklogn). This enhancement allows addressing large datasets; ii) demonstration of RGCLI mathematical properties, proving the constructed graph is an optimal graph to model the smoothness assumption of SSL; and iii) evaluation of the efficacy of the proposed approach in a comprehensive semi-supervised classification scenario with several datasets, including an image segmentation task, which needs a large graph to represent the image. Such experiments show the use of labeled vertices in the graph construction process improves the graph topology, hence, the learning task in which it will be employed.
Lilian Berton, Thiago de Paulo Faleiros, Alan Valejo, Jorge Carlos Valverde-Rebaza, Alneu de Andrade Lopes
Neurocomputing5
2017 Using bipartite heterogeneous networks to speed up inductive semi-supervised learning and improve automatic text categorization
abstract
Due to the volume of texts available in digital form, the organization, management and knowledge extraction are laborious and frequently impossible to be handled. To automatically cope with these tasks, usually classification models are generated through supervised learning techniques. Unfortunately, this type of learning usually demands a huge human effort to label large volume of texts to build accurate classification models. Since collecting unlabeled texts is easy and inexpensive in several domains, the generation of classification models through inductive semi-supervised learning has been highlighted in recent years. Inductive semi-supervised learning allows to build a classification model using labeled and unlabeled texts. In this scenario, the goal is to augment the set of labeled documents with unlabeled documents to better discriminate class patterns. Hence, fewer texts must be previously labeled. However, semi-supervised learning algorithms that consider texts represented in a vector space model usually obtain unsatisfactory classification performances and are surpassed by semi-supervised learning algorithms that consider texts represented in a network. Nevertheless, despite the classification performances, effective approaches based on networks are generated through the similarities among documents and the classification of a new document are also based on the computation of similarities. This implies to set parameters and compute similarities to both generation the networks and classification of new documents. This approach is not feasible to generate fast responses and consequently to classify a huge volume of texts. In this article, we propose an approach to induce a classification model through semi-supervised learning considering text collections represented by bipartite heterogeneous networks. Bipartite networks are easily and quickly generated, leading to classification performance equivalent or better than other approaches based on network or vector space model and allows a fast classification of new documents. The results presented in this article demonstrate that the proposed approach is able to (i) speed up semi-supervised learning, (ii) speed up the classification of new documents and (iii) surpass classification performance of other existing inductive semi-supervised learning techniques.
Rafael Geraldeli Rossi, Alneu de Andrade Lopes, Solange Oliveira Rezende
Knowl. Based Syst.2
2017 Optimizing the class information divergence for transductive classification of texts using propagation in bipartite graphs
abstract
Scalable algorithm based on bipartite graphs to perform transduction learning.Label propagation procedure that uses class information associated with vertices and edges.Better performance than state-of-the-art algorithms based on vector space or graphs.Comprehensive evaluation showing the proposal performance with few labeled instances.Optimization process using KL-Divergence. Transductive classification is an useful way to classify a collection of unlabeled textual documents when only a small fraction of this collection can be manually labeled. Graph-based algorithms have aroused considerable interests in recent years to perform transductive classification since the graph-based representation facilitates label propagation through the graph edges. In a bipartite graph representation, nodes represent objects of two types, here documents and terms, and the edges between documents and terms represent the occurrences of the terms in the documents. In this context, the label propagation is performed from documents to terms and then from terms to documents iteratively. In this paper we propose a new graph-based transductive algorithm that use the bipartite graph structure to associate the available class information of labeled documents and then propagate these class information to assign labels for unlabeled documents. By associating the class information to edges linking documents to terms we guarantee that a single term can propagate different class information to its distinct neighbors. We also demonstrated that the proposed method surpasses the algorithms for transductive classification based on vector space model or graphs when only a small number of labeled documents is available.
Thiago de Paulo Faleiros, Rafael Geraldeli Rossi, Alneu de Andrade Lopes
Pattern Recognit. Lett.3
2016 On the equivalence between algorithms for Non-negative Matrix Factorization and Latent Dirichlet Allocation
Thiago de Paulo Faleiros, Alneu de Andrade Lopes
ESANN2
2016 Exploiting social and mobility patterns for friendship prediction in location-based social networks
abstract
Link prediction is a “hot topic” in network analysis and has been largely used for friendship recommendation in social networks. With the increased use of location-based services, it is possible to improve the accuracy of link prediction methods by using the mobility of users. The majority of the link prediction methods focus on the importance of location for their visitors, disregarding the strength of relationships existing between these visitors. We, therefore, propose three new methods for friendship prediction by combining, efficiently, social and mobility patterns of users in location-based social networks (LBSNs). Experiments conducted on real-world datasets demonstrate that our proposals achieve a competitive performance with methods from the literature and, in most of the cases, outperform them. Moreover, our proposals use less computational resources by reducing considerably the number of irrelevant predictions, making the link prediction task more efficient and applicable for real world applications.
Jorge Carlos Valverde-Rebaza, Mathieu Roche, Pascal Poncelet, Alneu de Andrade Lopes
ICPR4
2016 The Extraction from News Stories a Causal Topic Centred Bayesian Graph for Sugarcane
abstract
Sugarcane is an important product to the Brazilian economy because it is the primary ingredient of ethanol which is used as a gasoline substitute. Sugarcane is affected by many factors which can be modelled in a Bayesian Graph. This paper describes a technique to build a Causal Bayesian Network from information in news stories. The technique: extracts causal relations from news stories, converts them into an event graph, removes irrelevant information, solves structure problems, and clusters the event graph by topic distribution. Finally, the paper describes a method for generating inferences from the graph based upon evidence in agricultural news stories. The graph is evaluated through a manual inspection and with a comparison with the EMBRAPA sugarcane taxonomy.
Brett Drury, Conceição Rocha, Maria-Fernanda Moura, Alneu de Andrade Lopes
IDEAS4
2016 Optimization and label propagation in bipartite heterogeneous networks to improve transductive classification of texts
abstract
Transductive classification is a useful way to classify texts when labeled training examples are insufficient. Several algorithms to perform transductive classification considering text collections represented in a vector space model have been proposed. However, the use of these algorithms is unfeasible in practical applications due to the independence assumption among instances or terms and the drawbacks of these algorithms. Network-based algorithms come up to avoid the drawbacks of the algorithms based on vector space model and to improve transductive classification. Networks are mostly used for label propagation, in which some labeled objects propagate their labels to other objects through the network connections. Bipartite networks are useful to represent text collections as networks and perform label propagation. The generation of this type of network avoids requirements such as collections with hyperlinks or citations, computation of similarities among all texts in the collection, as well as the setup of a number of parameters. In a bipartite heterogeneous network, objects correspond to documents and terms, and the connections are given by the occurrences of terms in documents. The label propagation is performed from documents to terms and then from terms to documents iteratively. Nevertheless, instead of using terms just as means of label propagation, in this article we propose the use of the bipartite network structure to define the relevance scores of terms for classes through an optimization process and then propagate these relevance scores to define labels for unlabeled documents. The new document labels are used to redefine the relevance scores of terms which consequently redefine the labels of unlabeled documents in an iterative process. We demonstrated that the proposed approach surpasses the algorithms for transductive classification based on vector space model or networks. Moreover, we demonstrated that the proposed algorithm effectively makes use of unlabeled documents to improve classification and it is faster than other transductive algorithms.
Rafael Geraldeli Rossi, Alneu de Andrade Lopes, Solange Oliveira Rezende
Inf. Process. Manag.2
2015 Term Network Approach for Transductive Classification
Rafael Geraldeli Rossi, Solange Oliveira Rezende, Alneu de Andrade Lopes
CICLing (2)3
2015 Graph Construction for Semi-Supervised Learning
Lilian Berton, Alneu de Andrade Lopes
IJCAI2
2015 Bipartite Graph for Topic Extraction
Thiago de Paulo Faleiros, Alneu de Andrade Lopes
IJCAI2
2015 Link prediction in graph construction for supervised and semi-supervised learning
abstract
Many real-world domains are relational in nature since they consist of a set of objects related to each other in complex ways. However, there are also flat data sets and if we want to apply graph-based algorithms, it is necessary to construct a graph from this data. This paper aims to: i) increase the exploration of graph-based algorithms and ii) proposes new techniques for graph construction from flat data. Our proposal focuses on constructing graphs using link prediction measures for predicting the existence of links between entities from an initial graph. Starting from a basic graph structure such as a minimum spanning tree, we apply a link prediction measure to add new edges in the graph. The link prediction measures considered here are based on structural similarity of the graph that improves the graph connectivity. We evaluate our proposal for graph construction in supervised and semi-supervised classification and we confirm the graphs achieve better accuracy.
Lilian Berton, Jorge Carlos Valverde-Rebaza, Alneu de Andrade Lopes
IJCNN3
2015 Graph-based measures to assist user assessment of multidimensional projections
Robson Motta, Rosane Minghim, Alneu de Andrade Lopes, Maria Cristina Ferreira de Oliveira
Neurocomputing3
2014 Link Prediction in Online Social Networks Using Group Information
Jorge Carlos Valverde-Rebaza, Alneu de Andrade Lopes
ICCSA (6)2
2014 Graph Construction Based on Labeled Instances for Semi-supervised Learning
abstract
Semi-Supervised Learning (SSL) techniques have become very relevant since they require a small set of labeled data. In this context, graph-based algorithms have gained prominence in the area due to their capacity to exploiting, besides information about data points, the relationships among them. Moreover, data represented in graphs allow the use of collective inference (vertices can affect each other), propagation of labels (autocorrelation among neighbors) and use of neighborhood characteristics of a vertex. An important step in graph-based SSL methods is the conversion of tabular data into a weighted graph. The graph construction has a key role in the quality of the classification in graph-based methods. This paper explores a method for graph construction that uses available labeled data. We provide extensive experiments showing the proposed method has many advantages: good classification accuracy, quadratic time complexity, no sensitivity to the parameter k > 10, sparse graph formation with average degree around 2 and hub formation from the labeled points, which facilitates the propagation of labels.
Lilian Berton, Alneu de Andrade Lopes
ICPR2
2014 Multilevel refinement based on neighborhood similarity
abstract
The multilevel graph partitioning strategy aims to reduce the computational cost of the partitioning algorithm by applying it on a coarsened version of the original graph. This strategy is very useful when large-scale networks are analyzed. To improve the multilevel solution, refinement algorithms have been used in the uncorsening phase. Typical refinement algorithms exploit network properties, for example minimum cut or modularity, but they do not exploit features from domain specific networks. For instance, in social networks partitions with high clustering coefficient or similarity between vertices indicate a better solution. In this paper, we propose a refinement algorithm (RSim) which is based on neighborhood similarity. We compare RSim with: 1. two algorithms from the literature and 2. one baseline strategy, on twelve real networks. Results indicate that RSim is competitive with methods evaluated for general domains, but for social networks it surpasses the competing refinement algorithms.
Alan Valejo, Jorge Carlos Valverde-Rebaza, Brett Drury, Alneu de Andrade Lopes
IDEAS4
2014 Inductive Model Generation for Text Classification Using a Bipartite Heterogeneous Network
Rafael Geraldeli Rossi, Alneu de Andrade Lopes, Thiago de Paulo Faleiros, Solange Oliveira Rezende
J. Comput. Sci. Technol.2
2013 Learning Bayesian Network Using Parse Trees for Extraction of Protein-Protein Interaction
Pedro Nelson Shiguihara-Juárez, Alneu de Andrade Lopes
CICLing (2)2
2013 An incremental learning algorithm based on the K-associated graph for non-stationary data classification
João Roberto Bertini Jr., Liang Zhao 0001, Alneu de Andrade Lopes
Inf. Sci.3
2012 Inductive Model Generation for Text Categorization Using a Bipartite Heterogeneous Network
abstract
Usually, algorithms for categorization of numeric data have been applied for text categorization after a preprocessing phase which assigns weights for textual terms deemed as attributes. However, due to characteristics of textual data, some algorithms for data categorization are not efficient for text categorization. Characteristics of textual data such as sparsity and high dimensionality sometimes impair the quality of general purpose classifiers. Here, we propose a text classifier based on a bipartite heterogeneous network used to represent textual document collections. Such algorithm induces a classification model assigning weights to objects that represents terms of the textual document collection. The induced weights correspond to the influence of the terms in the classification of documents they appear. The least-mean-square algorithm is used in the inductive process. Empirical evaluation using a large amount of textual document collections shows that the proposed IMBHN algorithm produces significantly better results than the k-NN, C4.5, SVM and Naïve Bayes algorithms.
Rafael Geraldeli Rossi, Thiago de Paulo Faleiros, Alneu de Andrade Lopes, Solange Oliveira Rezende
ICDM3
2012 Multidimensional Projections for Visual Analysis of Social Networks
Rafael Messias Martins, Gabriel de Faria Andery, Henry Heberle, Fernando Vieira Paulovich, Alneu de Andrade Lopes, Hélio Pedrini, Rosane Minghim
J. Comput. Sci. Technol.5
2011 A nonparametric classification method based on K-associated graphs
João Roberto Bertini Jr., Liang Zhao 0001, Robson Motta, Alneu de Andrade Lopes
Inf. Sci.4
2010 Combining Local and Global KNN With Cotraining
abstract
Semi-supervised learning is a machine learning paradigm in which the induced hypothesis is improved by taking advantage of unlabeled data. It is particularly useful when labeled data is scarce. Cotraining is a widely adopted semi-supervised approach that assumes availability of two views of the training data a restrictive assumption for most real world tasks. In this paper, we propose a one-view Cotraining approach that combines two different k-Nearest Neighbors (KNN) strategies referred to as global and local k-NN. In global KNN, the nearest neighbors selected to classify a new instance are given by the training examples which include this instance as one of their own k-nearest neighbors. In local KNN, on the other hand, the neighborhood considered when classifying a new instance is computed with the traditional KNN approach. We carried out experiments showing that a combination of these strategies significantly improves the classification accuracy in Cotraining, particularly when one single view of training data is available. We also introduce an optimized algorithm to cope with time complexity of computing the global KNN, which enables tackling real classification problems.
Víctor Laguna, Alneu de Andrade Lopes
ECAI2
2010 An incremental space to visualize dynamic data sets
Roberto Pinho, Maria Cristina Ferreira de Oliveira, Alneu de Andrade Lopes
Multim. Tools Appl.3
2009 Centrality Measures from Complex Networks in Active Learning
Robson Motta, Alneu de Andrade Lopes, Maria Cristina Ferreira de Oliveira
Discovery Science2
2007 Visual text mining using association rules
Alneu de Andrade Lopes, Roberto Pinho, Fernando Vieira Paulovich, Rosane Minghim
Comput. Graph.1