EDBT 2026 Demo / reviewers in the wild / expert
Adriano Veloso
dblp:12/919 · also Adriano Alonso Veloso
· DBLP profile ↗
75ranked-venue papers
13as first author
17since 2021 · last 2025
0000-0002-9177-4954ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 40 · 11 first-author · 3 since 2021Artificial intelligence and machine learning · 36 · 5 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Genetic Algorithms for Feature Selection with Language Models
Francisco Galuppo, Gianlucca L. Zuin, Guilherme Drummond, Denis Oliveira, Adriano Veloso |
IEEE Big Data | 5 |
| 2025 | Enhancing Authorship Attribution with Synthetic PaintingsabstractAttributing authorship to paintings is a historically complex task, and one of its main challenges is the limited availability of real artworks for training computational models. This study investigates whether synthetic images, generated through DreamBooth fine-tuning of Stable Diffusion, can improve the performance of classification models in this context. We propose a hybrid approach that combines real and synthetic data to enhance model accuracy and generalization across similar artistic styles. Experimental results show that adding synthetic images leads to higher ROC-AUC and accuracy compared to using only real paintings. By integrating generative and discriminative methods, this work contributes to the development of computer vision techniques for artwork authentication in data-scarce scenarios. Clarissa Loures, Caio Hosken, Luan Oliveira, Gianlucca L. Zuin, Adriano Veloso |
ICMLA | 5 |
| 2025 | High Performance Machine Learning for Nation Wide Crop ClassificationabstractIn this work, we present an automated system for analyzing agricultural productivity using machine learning and remote sensing. Satellite images are employed to classify crops and planting periods, addressing challenges such as infrequent image captures, cloud interference, and spatial resolution through geometric expansion and data-filling techniques. Focusing on Brazil’s primary crops, soybeans and corn, we construct a dataset of productive zones to delineate regions, identify planting areas, and determine planting and harvesting timelines. The system achieves an AUROC of 0.94 with manual labeling and 0.88 using a semi-supervised heuristic approach, demonstrating its effectiveness in optimizing agricultural productivity assessment despite data limitations. Lucas Borges Aquino, Gianlucca L. Zuin, Adriano Veloso, Nivio Ziviani |
IJCNN | 3 |
| 2025 | Leveraging Large Language Models for Tacit Knowledge Discovery in Organizational ContextsabstractDocumenting tacit knowledge in organizations can be a challenging task due to incomplete initial information, difficulty in identifying knowledgeable individuals, the interplay of formal hierarchies and informal networks, and the need to ask the right questions. To address this, we propose an agent-based framework leveraging large language models (LLMs) to iteratively reconstruct dataset descriptions through interactions with employees. Modeling knowledge dissemination as a Susceptible-Infectious (SI) process with waning infectivity, we conduct 864 simulations across various synthetic company structures and different dissemination parameters. Our results show that the agent achieves 94.9% full-knowledge recall, with self-critical feedback scores strongly correlating with external literature critic scores. We analyze how each simulation parameter affects the knowledge retrieval process for the agent. In particular, we find that our approach is able to recover information without needing to access directly the only domain specialist. These findings highlight the agent’s ability to navigate organizational complexity and capture fragmented knowledge that would otherwise remain inaccessible. Gianlucca L. Zuin, Saulo Martiello Mastelini, Túlio C. Loures, Adriano Veloso |
IJCNN | 4 |
| 2025 | "A 6 or a 9?": Ensemble Learning through the Multiplicity of Performant Models and ExplanationsabstractCreating models from past observations and ensuring their effectiveness on new data is the essence of machine learning. However, selecting models that generalize well remains a challenging task. Related to this topic, the Rashomon Effect refers to cases where multiple models perform similarly well for a given learning problem. This often occurs in real-world scenarios, like the manufacturing process or medical diagnosis, where diverse patterns in data lead to multiple high-performing solutions. We propose the Rashomon Ensemble, a method that strategically selects models from these diverse high-performing solutions to improve generalization. By grouping models based on both their performance and explanations, we construct ensembles that maximize diversity while maintaining predictive accuracy. This selection ensures that each model covers a distinct region of the solution space, making the ensemble more robust to distribution shifts and variations in unseen data. We validate our approach on both open and proprietary collaborative real-world datasets, demonstrating up to 0.20+ AUROC improvements in scenarios where the Rashomon ratio is large. Additionally, we demonstrate tangible benefits for businesses in various real-world applications, highlighting the robustness, practicality, and effectiveness of our approach. Gianlucca L. Zuin, Adriano Veloso |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | Navigating Time's Possibilities: Plausible Counterfactual Explanations for Multivariate Time-Series Forecast through Genetic AlgorithmsabstractCounterfactual learning has become promising for understanding and modeling causality in complex and dynamic systems. This paper presents a novel method for counterfactual learning in the context of multivariate time series analysis and forecast. The primary objective is to uncover hidden causal relationships and identify potential interventions to achieve desired outcomes. The proposed methodology integrates genetic algorithms and rigorous causality tests to infer and validate counterfactual dependencies within temporal sequences. More specifically, we employ Granger causality to enhance the reliability of identified causal relationships, rigorously assessing their statistical significance. Then, genetic algorithms, in conjunction with quantile regression, are used to exploit these intricate causal relationships to project future scenarios. The synergy between genetic algorithms and causality tests ensures a thorough exploration of the temporal dynamics present in the data, revealing hidden dependencies and enabling the projection of outcomes under hypothetical interventions. We evaluate the performance of our algorithm on real-world data, showcasing its ability to handle complex causal relationships, revealing meaningful counterfactual insights, and allowing for the prediction of outcomes under hypothetical interventions. Gianlucca L. Zuin, Adriano Veloso |
TrustCom | 2 |
| 2023 | Mitigating bias in facial analysis systems by incorporating label diversity
Camila Kolling, Victor Flavio de Andrade Araujo, Adriano Veloso, Soraia Raupp Musse |
Comput. Graph. | 3 |
| 2022 | The Subtle Art of Digging for Defects: Analyzing Features for Defect Prediction in Java Projects
Geanderson E. dos Santos, Adriano Veloso, Eduardo Figueiredo 0001 |
ENASE | 2 |
| 2022 | Drug Repurposing Opportunities in Shapley SpaceabstractRepurposing is the process of finding new indications for already available drugs. Mathematical and computational modeling facilitates the complex and time-consuming process of identifying new uses for known compounds. Motivated by the study of drug repurposing, we present an unsupervised node embedding algorithm that learns latent representations of drugs and diseases based on adaptive sampling of nodes' nearest neighbors. The learned latent representations were then fed to train a model in order to predict the existence of possible links between pairs of nodes. We studied the proximities in the model's decision space to identify hidden similarities among drugs and diseases and managed to find dozens of drug repurposing candidates. Amir Jalilifard, Adriano Veloso |
IJCNN | 2 |
| 2022 | Automatic Model Evaluation using Feature Importance Patterns on Unlabeled DataabstractRecent studies have shown that the estimated efficacy of a classification model in a specific training set can be very different from the same model efficacy after deployment, or when the model is evaluated in a dataset with a different distribution from the training set. This situation is known as distribution shift or dataset shift, and an emerging strategy for this problem is to estimate the efficacy of the classification model in unlabeled data with unknown distribution (i.e., aka AutoEval approaches). Most of the recent works study how distribution shift affects Deep Learning Models applied to Computer Vision tasks (i.e., unstructured data). However, distribution shift also can affect the efficacy of a model on tabular/structured data. In this work, we proposed and analyzed AutoEval approaches on tabular data. We proposed an AutoEval method based on the use of feature importance, which are typically used as model explanations, to detect patterns of correct and incorrect classifications that can be used to estimate model efficacy. We conducted experiments using six real-world datasets related to three different subjects. Our results indicated that the proposed method outperforms all baselines, with a reduction up to 89% in the gap between estimated and real efficacy, in comparison with standard 10-fold cross-validation (CV) error estimation. Besides, we evaluated our AutoEval approaches as indicators to model selection in the feature selection task. In this task, compared to CV, our proposed method achieved gains up to 35%. Ismael S. Silva, Adriano Veloso |
IJCNN | 2 |
| 2022 | A data-centric approach for predicting individual outcomes in a multi-party legislative systemabstractCongressional roll-call voting is highly predictable in strong two-party systems, and simple spatial voting models are sufficient to explain most of individual voting choices. Such spatial predictability is mainly due to the high degree of ideological stability exhibited by members of the same party. In multi-party legislative systems even if we also find a high degree of party discipline, individuals can be more susceptible to defection. This is a more complex and harder to monitor environment, and as such, models that predict “yea” or “nay” votes are the need of the hour for advocacy groups. In order to learn effective prediction models, the training data must entail diverse aspects of the modus-operandi of the congress. In this paper we discuss a data-centric approach for the problem, with emphasis on the features we built. Each congressperson is represented by a diverse set of features which are computed using linguistic models, causal models, and structured information. Finally, we show that meta-features derived from feature importance information lead to highly effective prediction models. Roberta Viola, Guilherme Drummond, Adriano Veloso, Mauricio Zuardi |
IJCNN | 3 |
| 2022 | Early identification of ICU patients at risk of complications: Regularization based on robustness and stability of explanations
Tiago Amador, Saulo Saturnino, Adriano Veloso, Nivio Ziviani |
Artif. Intell. Medicine | 3 |
| 2022 | Deep Learning Techniques for Explainable Resource Scales in Collectible Card GamesabstractIn collectible card games, developers face the challenge of creating new, and interesting cards that are not too strong or game-breaking, retaining the game’s overall balance. Over time, this becomes challenging due to the sheer volume of the published content. In this article, we propose a framework for generating models capable of recommending resource scales, a pivotal point in balancing. We evaluate the usage of several state-of-the-art neural architectures to learn representations for text followed by gradient boosting decision trees to incorporate remaining features. Throughout our analysis, we present various explanation tools that should empower game developers, and aid them with new insights. In particular, we present the sets of words that drive the model in diverse situations, such as when it was inaccurate by a small margin. We also exhibit instances where textual features cannot give an accurate prediction, requiring additional information. Our method achieves a mean reciprocal rank of$+0.8$when evaluated on popular card games, even though superficially identical cards might have distinct costs. Gianlucca L. Zuin, Luiz Chaimowicz, Adriano Veloso |
IEEE Trans. Games | 3 |
| 2021 | Assessing Media Bias in Cross-Linguistic and Cross-National Populations
Allan Sales da Costa Melo, Albin Zehe, Leandro Balby Marinho, Adriano Veloso, Andreas Hotho, Janna Omeliyanenko |
ICWSM | 4 |
| 2021 | Predicting Heating Sliver in Duplex Stainless Steels Manufacturing through Rashomon SetsabstractA particular challenge while designing duplex steel is the minimization of surface defects such as heating slivers, since these defects may significantly increase production costs as they remain mostly undetected, being observed only during the final product inspection. Heating slivers may be originated in different stages of the steelmaking process, and to identify their formation in duplex stainless steel, we propose a feature decomposition approach that resulted in hundreds of thousands of predictive models learned from defective and non-defective duplex plates, thus offering diverse models regarding their incidence. We grouped these models based on their competing explanations, as a means to isolate the different possible root causes for heating sliver. Finally, we search for optimal models in the graphs derived from intra-cluster feature relationships. Gianlucca L. Zuin, Felipe Marcelino, Lucas Borges, João Couto, Victor Jorge, Mychell Laurindo, Glaucio Barcelos, Márcio Cunha, Valdeci Alvarenga, Henrique Rodrigues, Paulo Balsamo, Adriano Veloso |
IJCNN | 12 |
| 2021 | Finding reduced Raman spectroscopy fingerprint of skin samples for melanoma diagnosis through machine learning
Daniella Castro Araújo, Adriano Veloso, Renato Santos de Oliveira Filho, Marie-Noelle Giraud, Leandro José Raniero, Lydia Masako Ferreira, Renata Andrade Bitar |
Artif. Intell. Medicine | 2 |
| 2021 | Predicting the Evolution of Pain Relief: Ensemble Learning by Diversifying Model ExplanationsabstractModeling from data usually has two distinct facets: building sound explanatory models or creating powerful predictive models for a system or phenomenon. Most of recent literature does not exploit the relationship between explanation and prediction while learning models from data. Recent algorithms are not taking advantage of the fact that many phenomena are actually defined by diverse sub-populations and local structures, and thus there are many possible predictive models providing contrasting interpretations or competing explanations for the same phenomenon. In this article, we propose to explore a complementary link between explanation and prediction. Our main intuition is that models having their decisions explained by the same factors are likely to perform better predictions for data points within the same local structures. We evaluate our methodology to model the evolution of pain relief in patients suffering from chronic pain under usual guideline-based treatment. The ensembles generated using our framework are compared with all-in-one approaches of robust algorithms to high-dimensional data, such as Random Forests and XGBoost. Chronic pain can be primary or secondary to diseases. Its symptomatology can be classified as nociceptive, nociplastic, or neuropathic, and is generally associated with many different causal structures, challenging typical modeling methodologies. Our data includes 631 patients receiving pain treatment. We considered 338 features providing information about pain sensation, socioeconomic status, and prescribed treatments. Our goal is to predict, using data from the first consultation only, if the patient will be successful in treatment for chronic pain relief. As a result of this work, we were able to build ensembles that are able to consistently improve performance by up to 33% when compared to models trained using all the available features. We also obtained relevant gains in interpretability, with resulting ensembles using only 15% of the total number of features. We show we can effectively generate ensembles from competing explanations, promoting diversity in ensemble learning and leading to significant gains in accuracy by enforcing a stable scenario in which models that are dissimilar in terms of their predictions are also dissimilar in terms of their explanation factors. Anderson Bessa Da Costa, Larissa Moreira, Daniel Ciampi De Andrade, Adriano Veloso, Nivio Ziviani |
ACM Trans. Comput. Heal. | 4 |
| 2020 | Modeling Pharmacological Effects with Multi-Relation Unsupervised Graph EmbeddingabstractA pharmacological effect of a drug on cells, organs and systems refers to the specific biochemical interaction produced by a drug substance, which is called its mechanism of action. Drug repositioning (or drug repurposing) is a fundamental problem for the identification of new opportunities for the use of already approved or failed drugs. In this paper, we present a method based on a multi-relation unsupervised graph embedding model that learns latent representations for drugs and diseases so that the distance between these representations reveals repositioning opportunities. Once representations for drugs and diseases are obtained we learn the likelihood of new links (that is, new indications) between drugs and diseases. Known drug indications are used for learning a model that predicts potential indications. Compared with existing unsupervised graph embedding methods our method shows superior prediction performance in terms of area under the ROC curve, and we present examples of repositioning opportunities found on recent biomedical literature that were also predicted by our method. Dehua Chen, Amir Jalilifard, Adriano Veloso, Nivio Ziviani |
IJCNN | 3 |
| 2020 | Explainable Deep CNNs for MRI-Based Diagnosis of Alzheimer's DiseaseabstractDeep Convolutional Neural Networks (CNNs) are becoming prominent models for semi-automated diagnosis of Alzheimer's Disease (AD) using brain Magnetic Resonance Imaging (MRI). Although being highly accurate, deep CNN models lack transparency and interpretability, precluding adequate clinical reasoning and not complying with most current regulatory demands. One popular choice for explaining deep image models is occluding regions of the image to isolate their influence on the prediction. However, existing methods for occluding patches of brain scans generate images outside the distribution to which the model was trained for, thus leading to unreliable explanations. In this paper, we propose an alternative explanation method that is specifically designed for the brain scan task. Our method, which we refer to as Swap Test, produces heatmaps that depict the areas of the brain that are most indicative of AD, providing interpretability for the model's decisions in a format understandable to clinicians. Experimental results using an axiomatic evaluation show that the proposed method is more suitable for explaining the diagnosis of AD using MRI while the opposite trend was observed when using a typical occlusion test. Therefore, we believe our method may address the inherent black-box nature of deep neural networks that are capable of diagnosing AD. Eduardo Nigri, Nivio Ziviani, Fabio A. M. Cappabianco, Augusto Antunes, Adriano Veloso |
IJCNN | 5 |
| 2020 | Deep Active Learning for Anomaly DetectionabstractAnomalies are intuitively easy for human experts to understand, but they are hard to define mathematically. Therefore, in order to have performance guarantees in unsupervised anomaly detection, priors need to be assumed on what the anomalies are. By contrast, active learning provides the necessary priors through appropriate expert feedback. Thus, in this work we present an active learning method that can be built upon existing deep learning solutions for unsupervised anomaly detection, so that outliers can be separated from normal data effectively. We introduce a new layer that can be easily attached to any deep learning model designed for unsupervised anomaly detection to transform it into an active method. We report results on both synthetic and real anomaly detection datasets, using multi-layer perceptrons and autoencoder architectures empowered with the proposed active layer, and we discuss their performance on finding clustered and low density anomalies. Tiago Pimentel, Marianne Monteiro, Adriano Veloso, Nivio Ziviani |
IJCNN | 3 |
| 2020 | Assessing the Reliability of Visual Explanations of Deep Models with Adversarial PerturbationsabstractThe interest in complex deep neural networks for computer vision applications is increasing. This leads to the need for improving the interpretable capabilities of these models. Recent explanation methods present visualizations of the relevance of pixels from input images, thus enabling the direct interpretation of properties of the input that lead to a specific output. These methods produce maps of pixel importance, which are commonly evaluated by visual inspection. This means that the effectiveness of an explanation method is assessed based on human expectation instead of actual feature importance. Thus, in this work we propose an objective measure to evaluate the reliability of explanations of deep models. Specifically, our approach is based on changes in the network's outcome resulting from the perturbation of input images in an adversarial way. We present a comparison between widely-known explanation methods using our proposed approach. Finally, we also propose a straightforward application of our approach to clean relevance maps, creating more interpretable maps without any loss in essential explanation (as per our proposed measure). Dan Valle, Tiago Pimentel, Adriano Veloso |
IJCNN | 3 |
| 2020 | Automatic Tag Recommendation for Painting Artworks Using Diachronic DescriptionsabstractIn this paper, we deal with the problem of automatic tag recommendation for painting artworks. Diachronic descriptions containing deviations on the vocabulary used to describe each painting usually occur when the work is done by many experts over time. The objective of this work is to provide a framework that produces a more accurate and homogeneous set of tags for each painting in a large collection. To validate our method we build a model based on a weakly-supervised neural network for over 5,300 paintings with hand-labeled descriptions made by experts for the paintings of the Brazilian painter Candido Portinari. This work takes place with the Portinari Project which started in 1979 intending to recover and catalog the paintings of the Brazilian painter. The Portinari paintings at that time were in private collections and museums spread around the world and thus inaccessible to the public. The descriptions of each painting were made by a large number of collaborators over 40 years as the paintings were recovered and these diachronic descriptions caused deviations on the vocabulary used to describe each painting. Our proposed framework consists of (i) a neural network that receives as input the image of each painting and uses frequent itemsets as possible tags, and (ii) a clustering step in which we group related tags based on the output of the pre-trained classifiers. Gianlucca L. Zuin, Adriano Veloso, João Cândido Portinari, Nivio Ziviani |
IJCNN | 2 |
| 2020 | Computing with Subjectivity LexiconsabstractIn this paper, we introduce a new set of lexicons for expressing subjectivity in text documents written in Brazilian Portuguese. Besides the non-English idiom, in contrast to other subjectivity lexicons available, these lexicons represent different subjectivity dimensions (other than sentiment) and are more compact in number of terms. This last feature was designed intentionally to leverage the power of word embedding techniques, i.e., with the words mapped to an embedding space and the appropriate distance measures, we can easily capture semantically related words to the ones in the lexicons. Thus, we do not need to build comprehensive vocabularies and can focus on the most representative words for each lexicon dimension. We showcase the use of these lexicons in three highly non-trivial tasks: (1) Automated Essay Scoring in the Presence of Biased Ratings, (2) Subjectivity Bias in Brazilian Presidential Elections and (3) Fake News Classification Based on Text Subjectivity. All these tasks involve text documents written in Portuguese. Caio Libânio Melo Jerônimo, Cláudio Elízio Calazans Campelo, Leandro Balby Marinho, Allan Sales da Costa Melo, Adriano Veloso, Roberta Viola |
LREC | 5 |
| 2020 | Understanding machine learning software defect predictions
Geanderson E. dos Santos, Eduardo Figueiredo 0001, Adriano Veloso, Markos Viggiato, Nivio Ziviani |
Autom. Softw. Eng. | 3 |
| 2019 | Learning a Resource Scale for Collectible Card GamesabstractIn Collectible Card Games like "Magic: the Gathering", one of the developers' main challenges is creating new and interesting cards that are not too strong or game-braking, pertaining the game's overall balance. One way to address this issue is through the analysis of the cards resource costs. Powerful cards need more resource to be played while weaker ones need less resource. This work proposes a recommender system to a card's resource scale. In summary, we model the problem as a classification task and present and in-depth analysis of our results. We propose using LSTMs to learn a vector representation for text followed by XGBoost models to incorporate remaining features. Our approach is capable of reaching a Mean Reciprocal Rank of 0.8064 despite superficially identical cards having different mana costs. The analysis provided indicate that the model was able to learn useful rules for predicting a card's resource cost and highlight key insights for future research. Gianlucca L. Zuin, Adriano Veloso |
CoG | 2 |
| 2019 | A Sequential Approach for Pain Recognition Based on Facial Representations
Antoni Mauricio, Fabio A. M. Cappabianco, Adriano Veloso, Guillermo Cámara Chávez |
ICVS | 3 |
| 2019 | Fake News Classification Based on Subjective LanguageabstractWhile many works investigate spread patterns of fake news in social networks, we focus on the textual content. Instead of relying on syntactic representations of documents (aka Bag of Words) as many works do, we seek more robust representations that may better differentiate fake from legitimate news. We propose to consider the subjectivity of news under the assumption that the subjectivity levels of legitimate and fake news are significantly different. For computing the subjectivity level of news, we rely on a set subjectivity lexicons built by Brazilian linguists. We then build subjectivity feature vectors for each news article by calculating the Word Mover's Distance (WMD) between the news and these lexicons considering the embedding the news words lie in, in order to classify the documents. The results demonstrate that our method is more robust than classical text classification approaches, especially in scenarios where training and test domains are different. Caio Libânio Melo Jerônimo, Leandro Balby Marinho, Cláudio Elízio Calazans Campelo, Adriano Veloso, Allan Sales da Costa Melo |
iiWAS | 4 |
| 2019 | Efficient Estimation of Node Representations in Large Graphs using Linear ContextsabstractLearning distributed representations in graphs has a rising interest in the neural network community. Recent works have proposed new methods for learning low dimensional embeddings of nodes and edges in graphs and networks. Several of these methods rely on the SkipGram algorithm to learn distributed representations, and they usually process a large number of multi-hop neighbors in order to produce the context from which node representations are learned. This is a limiting factor for these methods as graphs and networks keep growing in size. In this paper, we propose a simple alternate method which is as effective as previous methods, but being much faster at learning node representations. Our proposed method employs a restricted number of permutations over the immediate neighborhood of a node as context to generate its representation, thus avoiding long walks and large contexts while learning the representations. We present a thorough evaluation showing that our method outperforms state-of-the-art methods in six different datasets related to the problems of link prediction and node classification, being one to three orders of magnitude faster than baselines when generating node embeddings for very large graphs. Tiago Pimentel, Rafael Castro, Adriano Veloso, Nivio Ziviani |
IJCNN | 3 |
| 2018 | Dynamic Prediction of ICU Mortality Risk Using Domain AdaptationabstractEarly recognition of risky trajectories during an Intensive Care Unit (ICU) stay is one of the key steps towards improving patient survival. Learning trajectories from physiological signals continuously measured during an ICU stay requires learning time-series features that are robust and discriminative across diverse patient populations. Patients within different ICU populations (referred here as domains) vary by age, conditions and interventions. Thus, mortality prediction models using patient data from a particular ICU population may perform suboptimally in other populations because the features used to train such models have different distributions across the groups. In this paper, we explore domain adaptation strategies in order to learn mortality prediction models that extract and transfer complex temporal features from multivariate time-series ICU data. Features are extracted in a way that the state of the patient in a certain time depends on the previous state. This enables dynamic predictions and creates a mortality risk space that describes the risk of a patient at a particular time. Experiments based on cross-ICU populations reveals that our model outperforms all considered baselines. Gains in terms of AUC range from 4% to 8% for early predictions when compared with a recent state-of-the-art representative for ICU mortality prediction. In particular, models for the Cardiac ICU population achieve AUC numbers as high as 0.88, showing excellent clinical utility for early mortality prediction. Finally, we present an explanation of factors contributing to the possible ICU outcomes, so that our models can be used to complement clinical reasoning. Tiago Alves 0002, Alberto H. F. Laender, Adriano Veloso, Nivio Ziviani |
IEEE BigData | 3 |
| 2018 | Learning to Rank with Deep Autoencoder FeaturesabstractLearning to rank in Information Retrieval is the problem of learning the full order of a set of documents from their partially observed order. Datasets used by learning to rank algorithms are growing enormously in terms of number of features, but it remains costly and laborious to reliably label large datasets. This paper is about learning feature transformations using inexpensive unlabeled data and available labeled data, that is, building alternate features so that it becomes easier for existing learning to rank algorithms to find better ranking models from labeled datasets that are limited in size and quality. Deep autoencoders have proven powerful as nonlinear feature extractors, and thus we exploit deep autoencoder features for semi-supervised learning to rank. Typical approaches for learning autoencoder features are based on updating model parameters using either unlabeled data only, or unlabeled data first and then labeled data. We propose a novel approach which updates model parameters using unlabeled and labeled data simultaneously, enabling label propagation from labeled to unlabeled data. We present a comprehensive study on how deep autoencoder features improve the ranking performance of representative learning to rank algorithms, revealing the importance of building an effective feature set to describe the input data. Alberto Albuquerque, Tiago Amador, Renato Ferreira 0001, Adriano Veloso, Nivio Ziviani |
IJCNN | 4 |
| 2018 | Effective Fashion Retrieval Based on Semantic Compositional NetworksabstractTypical approaches for fashion retrieval rank clothing images according to the similarity to a user-provided query image. Similarity is usually assessed by encoding images in terms of visual elements such as color, shape and texture. In this work, we proceed differently and consider that the semantics of an outfit is mainly comprised of environmental and cultural concepts such as occasion, style and season. Thus, instead of retrieving outfits using strict visual elements, we find semantically similar outfits that fall into similar clothing styles and are adequate for the same occasions and seasons. We propose a compositional approach for fashion retrieval by arguing that the semantics of an outfit can be recognised by their constituents (i.e., clothing items and accessories). Specifically, we present a semantic compositional network (Comp-Net) in which clothing items are detected from the image and the probability of each item is used to compose a vector representation for the outfit. Comp-Net employs a normalization layer so that weights are updated by taking into consideration the previously known co-occurrence patterns between clothing items. Further, Comp-Net minimizes a cost-sensitive loss function as errors have different costs depending on the clothing item that is misclassified. This results in a space in which semantically related outfits are placed next to each other, enabling to find semantically similar outfits that may not be visually similar. We designed an evaluation setup that takes into account the association between different styles, occasions and seasons, and show that our compositional approach significantly outperforms a variety of recently proposed baselines. Dan Valle, Nivio Ziviani, Adriano Veloso |
IJCNN | 3 |
| 2018 | Learning Transferable Features For Open-Domain Question AnsweringabstractCorpora used to learn open-domain Question-Answering (QA) models are typically collected from a wide variety of topics or domains. Since QA requires understanding natural language, open-domain QA models generally need very large training corpora. A simple way to alleviate data demand is to restrict the domain covered by the QA model, leading thus to domain-specific QA models. While learning improved QA models for a specific domain is still challenging due to the lack of sufficient training data in the topic of interest, additional training data can be obtained from related topic domains. Thus, instead of learning a single open-domain QA model, we investigate domain adaptation approaches in order to create multiple improved domain-specific QA models. We demonstrate that this can be achieved by stratifying the source dataset, without the need of searching for complementary data unlike many other domain adaptation approaches. We propose a deep architecture that jointly exploits convolutional and recurrent networks for learning domain-specific features while transferring domain-shared features. That is, we use transferable features to enable model adaptation from multiple source domains. We consider different transference approaches designed to learn span-level and sentence-level QA models. We found that domain-adaptation greatly improves sentence-level QA performance, and span-level QA benefits from sentence information. Finally, we also show that a simple clustering algorithm may be employed when the topic domains are unknown and the resulting loss in accuracy is negligible. Gianlucca L. Zuin, Luiz Chaimowicz, Adriano Veloso |
IJCNN | 3 |
| 2018 | Automated Essay Scoring in the Presence of Biased RatingsabstractStudies in Social Sciences have revealed that when people evaluate someone else, their evaluations often reflect their biases.As a result, rater bias may introduce highly subjective factors that make their evaluations inaccurate.This may affect automated essay scoring models in many ways, as these models are typically designed to model (potentially biased) essay raters.While there is sizeable literature on rater effects in general settings, it remains unknown how rater bias affects automated essay scoring.To this end, we present a new annotated corpus containing essays and their respective scores.Different from existing corpora, our corpus also contains comments provided by the raters in order to ground their scores.We present features to quantify rater bias based on their comments, and we found that rater bias plays an important role in automated essay scoring.We investigated the extent to which rater bias affects models based on hand-crafted features.Finally, we propose to rectify the training set by removing essays associated with potentially biased scores while learning the scoring model. Evelin Amorim, Márcia Cançado, Adriano Veloso |
NAACL-HLT | 3 |
| 2018 | Fast and Effective Neural Networks for Translating Natural Language into Denotations
Tiago Pimentel, Juliano Viana, Adriano Veloso, Nivio Ziviani |
SPIRE | 3 |
| 2018 | Website replica detection with distant supervision
Cristiano R. de Carvalho, Edleno Silva de Moura, Adriano Veloso, Nivio Ziviani |
Inf. Retr. J. | 3 |
| 2017 | Exploiting item co-utility to improve collaborative filtering recommendationsabstractIn this article we study the extent to which the interplay between recommended items affect recommendation effectiveness. We introduce and formalize the concept of co‐utility as the property that any pair of recommended items has of being useful to a user, and exploit it to improve collaborative filtering recommendations. We present different techniques to estimate co‐utility probabilities, all of them independent of content information, and compare them with each other. We use these probabilities, as well as normalized predicted ratings, in an instance of an ‐hard problem termed the Max‐Sum Dispersion Problem (MSDP). A solution to MSDP hence corresponds to a set of items for recommendation. We study one heuristic and one exact solution to MSDP and perform comparisons among them. We also contrast our solutions (the best heuristic to MSDP) to different baselines by comparing the ratings users give to different recommendations. We obtain expressive gains in the utility of recommendations and our solutions also recommend higher‐rated items to the majority of users. Finally, we show that our co‐utility solutions are scalable in practice and do not harm recommendations' diversity. Aline Bessa, Rodrygo L. T. Santos, Adriano Veloso, Nivio Ziviani |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2016 | Pointwise and pairwise clothing annotation: combining features from social media
Keiller Nogueira, Adriano Veloso, Jefersson A. dos Santos |
Multim. Tools Appl. | 2 |
| 2015 | Reverse Engineering Socialbot Infiltration Strategies in TwitterabstractOnline Social Networks (OSNs) such as Twitter and Facebook have become a significant testing ground for Artificial Intelligence developers who build programs, known as socialbots, that imitate actual users by automating their social-network activities such as forming social links and posting content. Particularly, Twitter users have shown difficulties in distinguishing these socialbots from the human users in their social graphs. Frequently, legitimate users engage in conversations with socialbots. More impressively, socialbots are effective in acquiring human users as followers and exercising influence within them. While the success of socialbots is certainly a remarkable achievement for AI practitioners, their proliferation in the Twitter-sphere opens many possibilities for cybercrime. The proliferation of socialbots in the Twitter-sphere motivates us to assess the characteristics or strategies that make socialbots most likely to succeed. In this direction, we created 120 socialbot accounts in Twitter, which have a profile, follow other users, and generate tweets either by reposting messages that others have posted or by creating their own synthetic tweets. Then, we employ a 2k factorial design experiment in order to quantify the infiltration effectiveness of different socialbot strategies. Our analysis is the first of a kind, and reveals what strategies make socialbots successful in the Twitter-sphere. Carlos Alessandro Sena de Freitas, Fabrício Benevenuto, Saptarshi Ghosh 0001, Adriano Veloso |
ASONAM | 4 |
| 2015 | Learning sequential classifiers from long and noisy discrete-event sequences efficiently
Gessé Dafé, Adriano Veloso, Mohammed J. Zaki, Wagner Meira Jr. |
Data Min. Knowl. Discov. | 2 |
| 2015 | Improving daily deals recommendation using explore-then-exploit strategies
Anísio Lacerda, Rodrygo L. T. Santos, Adriano Veloso, Nivio Ziviani |
Inf. Retr. J. | 3 |
| 2015 | Mining citizen emotions to estimate the urgency of urban issues
Christian Masdeval, Adriano Veloso |
Inf. Syst. | 2 |
| 2014 | Learning to Rank Similar Apparel Styles with Economically-Efficient Rule-Based Active LearningabstractIncreasingly, people define and express themselves in online social networks, such as Facebook and Instagram, by uploading photos showing the clothes they wear. As a result, such online social networks are becoming major sources of inspiration, with users looking for others with similar clothing style. In this paper, we propose a novel learning to rank (L2R) algorithm for finding similar apparel style given a query image. L2R algorithms use a labeled training set to generate a ranking model that can later be used to rank new query results. These training sets, however, are costly and laborious to produce, requiring human annotators to assess the relevance of candidate images in relation to a query. Active learning algorithms are able to reduce the labeling effort by selectively sampling an unlabeled set of images and choosing the subset that maximizes a learning function's effectiveness. Specifically, our proposed L2R algorithm employs an association rule active sampling algorithm to select very small but effective training sets. Further, our algorithm operates on visual (e.g., image descriptors) and textual (e.g., comments associated with the image) elements, in a way that makes it able (i) to expand the query image (for which only visual elements are available) with textual elements, and (ii) to combine multiple elements, being visual or textual, using basic economic efficiency concepts. We conducted a systematic evaluation of the proposed algorithm using every-day photos collected from Instagram, and we show that our L2R algorithm reduces by two orders of magnitude the need for labeled images, and still improves upon the state-of-the-art models by 4-8% in terms of mean average precision. Mariane Moreira, Jefersson A. dos Santos, Adriano Veloso |
ICMR | 3 |
| 2014 | Economically-efficient sentiment stream analysisabstractText-based social media channels, such as Twitter, produce torrents of opinionated data about the most diverse topics and entities. The analysis of such data (aka. sentiment analysis) is quickly becoming a key feature in recommender systems and search engines. A prominent approach to sentiment analysis is based on the application of classification techniques, that is, content is classified according to the attitude of the writer. A major challenge, however, is that Twitter follows the data stream model, and thus classifiers must operate with limited resources, including labeled data and time for building classification models. Also challenging is the fact that sentiment distribution may change as the stream evolves. In this paper we address these challenges by proposing algorithms that select relevant training instances at each time step, so that training sets are kept small while providing to the classifier the capabilities to suit itself to, and to recover itself from, different types of sentiment drifts. Simultaneously providing capabilities to the classifier, however, is a conflicting-objective problem, and our proposed algorithms employ basic notions of Economics in order to balance both capabilities. We performed the analysis of events that reverberated on Twitter, and the comparison against the state-of-the-art reveals improvements both in terms of error reduction (up to 14%) and reduction of training resources (by orders of magnitude). Roberto L. de Oliveira Jr., Adriano Veloso, Adriano C. M. Pereira, Wagner Meira Jr., Renato Ferreira 0001, Srinivasan Parthasarathy 0001 |
SIGIR | 2 |
| 2014 | Context-Aware Deal Size Prediction
Anísio Lacerda, Adriano Veloso, Rodrygo L. T. Santos, Nivio Ziviani |
SPIRE | 2 |
| 2014 | Self-training author name disambiguation for information scarce scenariosabstractWe present a novel 3‐step self‐training method for author name disambiguation—SAND (self‐training associative name disambiguator)—which requires no manual labeling, no parameterization (in real‐world scenarios) and is particularly suitable for the common situation in which only the most basic information about a citation record is available (i.e., author names, and work and venue titles). During the first step, real‐world heuristics on coauthors are able to produce highly pure (although fragmented) clusters. The most representative of these clusters are then selected to serve as training data for the third supervised author assignment step. The third step exploits a state‐of‐the‐art transductive disambiguation method capable of detecting unseen authors not included in any training example and incorporating reliable predictions to the training data. Experiments conducted with standard public collections, using the minimum set of attributes present in a citation, demonstrate that our proposed method outperforms all representative unsupervised author grouping disambiguation methods and is very competitive with fully supervised author assignment methods. Thus, different from other bootstrapping methods that explore privileged, hard to obtain information such as self‐citations and personal information, our proposed method produces topnotch performance with no (manual) training data or parameterization and in the presence of scarce information. Anderson A. Ferreira, Adriano Veloso, Marcos André Gonçalves, Alberto H. F. Laender |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2014 | A Two-stage active learning method for learning to rankabstractLearning to rank (L2R) algorithms use a labeled training set to generate a ranking model that can later be used to rank new query results. These training sets are costly and laborious to produce, requiring human annotators to assess the relevance or order of the documents in relation to a query. Active learning algorithms are able to reduce the labeling effort by selectively sampling an unlabeled set and choosing data instances that maximize a learning function's effectiveness. In this article, we propose a novel two‐stage active learning method for L2R that combines and exploits interesting properties of its constituent parts, thus being effective and practical. In the first stage, an association rule active sampling algorithm is used to select a very small but effective initial training set. In the second stage, a query‐by‐committee strategy trained with the first‐stage set is used to iteratively select more examples until a preset labeling budget is met or a target effectiveness is achieved. We test our method with various LETOR benchmarking data sets and compare it with several baselines to show that it achieves good results using only a small portion of the original training sets. Rodrigo M. Silva, Marcos André Gonçalves, Adriano Veloso |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2014 | Multiobjective Pareto-Efficient Approaches for Recommender SystemsabstractRecommender systems are quickly becoming ubiquitous in applications such as e-commerce, social media channels, and content providers, among others, acting as an enabling mechanism designed to overcome the information overload problem by improving browsing and consumption experience. A typical task in many recommender systems is to output a ranked list of items, so that items placed higher in the rank are more likely to be interesting to the users. Interestingness measures include how accurate, novel, and diverse are the suggested items, and the objective is usually to produce ranked lists optimizing one of these measures. Suggesting items that are simultaneously accurate, novel, and diverse is much more challenging, since this may lead to a conflicting-objective problem, in which the attempt to improve a measure further may result in worsening other measures. In this article, we propose new approaches for multiobjective recommender systems based on the concept of Pareto efficiency—a state achieved when the system is devised in the most efficient manner in the sense that there is no way to improve one of the objectives without making any other objective worse off. Given that existing multiobjective recommendation algorithms differ in their level of accuracy, diversity, and novelty, we exploit the Pareto-efficiency concept in two distinct manners: (i) the aggregation of ranked lists produced by existing algorithms into a single one, which we call Pareto-efficient ranking, and (ii) the weighted combination of existing algorithms resulting in a hybrid one, which we call Pareto-efficient hybridization. Our evaluation involves two real application scenarios: music recommendation with implicit feedback (i.e., Last.fm) and movie recommendation with explicit feedback (i.e., MovieLens). We show that the proposed Pareto-efficient approaches are effective in suggesting items that are likely to be simultaneously accurate, diverse, and novel. We discuss scenarios where the system achieves high levels of diversity and novelty without compromising its accuracy. Further, comparison against multiobjective baselines reveals improvements in terms of accuracy (from 10.4% to 10.9%), novelty (from 5.7% to 7.5%), and diversity (from 1.6% to 4.2%). Marco Túlio Ribeiro, Nivio Ziviani, Edleno Silva de Moura, Itamar Hata, Anísio Lacerda, Adriano Veloso |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2013 | Exploratory and interactive daily deals recommendationabstractDaily deals sites (DDSs), such as Groupon and LivingSocial, attract millions of customers in the hunt for products and services at significantly reduced prices. A typical approach to increase revenue is to send email messages featuring the deals of the day. Such daily messages, however, are usually not centered on the customers, instead, all registered users typically receive similar messages with almost the same deals. Traditional recommendation algorithms are innocuous in DDSs because: (i) most of the users are sporadic bargain hunters, and thus past preference data is extremely sparse, (ii) deals have a short living period, and thus data is extremely volatile, and (iii) user taste and interest may undergo temporal drifts. In order to address such particularly challenging scenario, we propose new algorithms for daily deals recommendation based on the explore-then-exploit strategy.Users are split into exploration and exploitation sets -- in the exploration set the users receive non-personalized messages and a co-purchase network is updated with user feedback for purchases of the day, while in the exploitation set the updated network is used for recommending personalized messages for the remaining users.A thorough evaluation of our algorithms using real data obtained from a large daily deals website in Brazil in contrast to state-of-the-art recommendation algorithms show gains in precision ranging from 18% to 34%. Anísio Lacerda, Adriano Veloso, Nivio Ziviani |
RecSys | 2 |
| 2013 | Using Mutual Influence to Improve Recommendations
Aline Bessa, Adriano Veloso, Nivio Ziviani |
SPIRE | 2 |
| 2013 | A KDD-Based Methodology to Rank Trust in e-Commerce SystemsabstractDue to the growing popularity of the Web, there is an increasing number of people who perform e-business transactions. On the other hand, this popularity has also attracted the attention of criminals, raising the number of frauds on the Web and associated financial losses, which reach billions of dollars per year. This paper proposes a KDD-based methodology to detect fraud in e-payment systems. In order to evaluate this methodology we defined the concept of economic efficiency and applied it to an actual dataset of one of the largest Latin American electronic payment systems. The results show a very good performance, providing gains of up to 46.5% in comparison with the strategy currently employed by the company. José Felipe Júnior, Adriano C. M. Pereira, Wagner Meira Jr., Adriano Veloso |
Web Intelligence | 4 |
| 2012 | Named Entity Disambiguation in Streaming Data
Alexandre Davis, Adriano Veloso, Altigran S. da Silva, Alberto H. F. Laender, Wagner Meira Jr. |
ACL (1) | 2 |
| 2012 | Automatic Vandalism Detection in Wikipedia with Active Associative Classification
Maria I. M. Sumbana, Marcos André Gonçalves, Rodrigo Silva Oliveira, Jussara M. Almeida, Adriano Veloso |
TPDL | 5 |
| 2012 | Pareto-efficient hybridization for multi-objective recommender systemsabstractPerforming accurate suggestions is an objective of paramount importance for effective recommender systems. Other important and increasingly evident objectives are novelty and diversity, which are achieved by recommender systems that are able to suggest diversified items not easily discovered by the users. Different recommendation algorithms have particular strengths and weaknesses when it comes to each of these objectives, motivating the construction of hybrid approaches. However, most of these approaches only focus on optimizing accuracy, with no regard for novelty and diversity. The problem of combining recommendation algorithms grows significantly harder when multiple objectives are considered simultaneously. For instance, devising multi-objective recommender systems that suggest items that are simultaneously accurate, novel and diversified may lead to a conflicting-objective problem, where the attempt to improve an objective further may result in worsening other competing objectives. In this paper we propose a hybrid recommendation approach that combines existing algorithms which differ in their level of accuracy, novelty and diversity. We employ an evolutionary search for hybrids following the Strength Pareto approach, which isolates hybrids that are not dominated by others (i.e., the so called Pareto frontier). Experimental results on two recommendation scenarios show that: (i) we can combine recommendation algorithms in order to improve an objective without significantly hurting other objectives, and (ii) we allow for adjusting the compromise between accuracy, diversity and novelty, so that the recommendation emphasis can be adjusted dynamically according to the needs of different users. Marco Túlio Ribeiro, Anísio Lacerda, Adriano Veloso, Nivio Ziviani |
RecSys | 3 |
| 2012 | Cost-effective on-demand associative author name disambiguation
Adriano Veloso, Anderson A. Ferreira, Marcos André Gonçalves, Alberto H. F. Laender, Wagner Meira Jr. |
Inf. Process. Manag. | 1 |
| 2012 | A tool for generating synthetic authorship records for evaluating author name disambiguation methods
Anderson A. Ferreira, Marcos André Gonçalves, Jussara M. Almeida, Alberto H. F. Laender, Adriano Veloso |
Inf. Sci. | 5 |
| 2012 | Practical Detection of Spammers and Content Promoters in Online Video Sharing SystemsabstractA number of online video sharing systems, out of which YouTube is the most popular, provide features that allow users to post a video as a response to a discussion topic. These features open opportunities for users to introduce polluted content, or simply pollution, into the system. For instance, spammers may post an unrelated video as response to a popular one, aiming at increasing the likelihood of the response being viewed by a larger number of users. Moreover, content promoters may try to gain visibility to a specific video by posting a large number of (potentially unrelated) responses to boost the rank of the responded video, making it appear in the top lists maintained by the system. Content pollution may jeopardize the trust of users on the system, thus compromising its success in promoting social interactions. In spite of that, the available literature is very limited in providing a deep understanding of this problem. In this paper, we address the issue of detecting video spammers and promoters. Towards that end, we first manually build a test collection of real YouTube users, classifying them as spammers, promoters, and legitimate users. Using our test collection, we provide a characterization of content, individual, and social attributes that help distinguish each user class. We then investigate the feasibility of using supervised classification algorithms to automatically detect spammers and promoters, and assess their effectiveness in our test collection. While our classification approach succeeds at separating spammers and promoters from legitimate users, the high cost of manually labeling vast amounts of examples compromises its full potential in realistic scenarios. For this reason, we further propose an active learning approach that automatically chooses a set of examples to label, which is likely to provide the highest amount of information, drastically reducing the amount of required training data while maintaining comparable classification effectiveness. Fabrício Benevenuto, Adriano Veloso, Jussara M. Almeida, Marcos André Gonçalves, Virgílio A. F. Almeida |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | From bias to opinion: a transfer-learning approach to real-time sentiment analysisabstractReal-time interaction, which enables live discussions, has become a key feature of most Web applications. In such an environment, the ability to automatically analyze user opinions and sentiments as discussions develop is a powerful resource known as real time sentiment analysis. However, this task comes with several challenges, including the need to deal with highly dynamic textual content that is characterized by changes in vocabulary and its subjective meaning and the lack of labeled data needed to support supervised classifiers. In this paper, we propose a transfer learning strategy to perform real time sentiment analysis. We identify a task - opinion holder bias prediction - which is strongly related to the sentiment analysis task; however, in constrast to sentiment analysis, it builds accurate models since the underlying relational data follows a stationary distribution. Pedro Henrique Calais Guerra, Adriano Veloso, Wagner Meira Jr., Virgílio A. F. Almeida |
KDD | 2 |
| 2011 | Rule-Based Active Sampling for Learning to Rank
Rodrigo M. Silva, Marcos André Gonçalves, Adriano Veloso |
ECML/PKDD (3) | 3 |
| 2011 | Effective sentiment stream analysis with self-augmenting training and demand-driven projectionabstractHow do we analyze sentiments over a set of opinionated Twitter messages? This issue has been widely studied in recent years, with a prominent approach being based on the application of classification techniques. Basically, messages are classified according to the implicit attitude of the writer with respect to a query term. A major concern, however, is that Twitter (and other media channels) follows the data stream model, and thus the classifier must operate with limited resources, including labeled data for training classification models. This imposes serious challenges for current classification techniques, since they need to be constantly fed with fresh training messages, in order to track sentiment drift and to provide up-to-date sentiment analysis. Ismael S. Silva, Janaína Gomide, Adriano Veloso, Wagner Meira Jr., Renato Ferreira 0001 |
SIGIR | 3 |
| 2011 | Calibrated lazy associative classification
Adriano Veloso, Wagner Meira Jr., Marcos André Gonçalves, Humberto Mossri de Almeida, Mohammed J. Zaki |
Inf. Sci. | 1 |
| 2010 | Active Learning Genetic programming for record deduplicationabstractThe great majority of genetic programming (GP) algorithms that deal with the classification problem follow a supervised approach, i.e., they consider that all fitness cases available to evaluate their models are labeled. However, in certain application domains, a lot of human effort is required to label training data, and methods following a semi-supervised approach might be more appropriate. This is because they significantly reduce the time required for data labeling while maintaining acceptable accuracy rates. This paper presents the Active Learning GP (AGP), a semi-supervised GP, and instantiates it for the data deduplication problem. AGP uses an active learning approach in which a committee of multi-attribute functions votes for classifying record pairs as duplicates or not. When the committee majority voting is not enough to predict the class of the data pairs, a user is called to solve the conflict. The method was applied to three datasets and compared to two other deduplication methods. Results show that AGP guarantees the quality of the deduplication while reducing the number of labeled examples needed. Junio de Freitas, Gisele L. Pappa, Altigran S. da Silva, Marcos André Gonçalves, Edleno Silva de Moura, Adriano Veloso, Alberto H. F. Laender, Moisés G. de Carvalho |
IEEE Congress on Evolutionary Computation | 6 |
| 2010 | Demand-Driven Tag Recommendation
Guilherme Vale Menezes, Jussara M. Almeida, Fabiano Muniz Belém, Marcos André Gonçalves, Anísio Lacerda, Edleno Silva de Moura, Gisele L. Pappa, Adriano Veloso, Nivio Ziviani |
ECML/PKDD (2) | 8 |
| 2009 | The Metric Dilemma: Competence-Conscious Associative ClassificationabstractThe classification performance of an associative classifier is strongly dependent on the statistic measure or metric that is used to quantify the strength of the association between features and classes (i.e., confidence, correlation etc.). Previous studies have shown that classifiers produced by different metrics may provide conflicting predictions, and that the best metric to use is data-dependent and rarely known while designing the classifier. This uncertainty concerning the optimal match between metrics and problems is a dilemma, and prevents associative classifiers to achieve their maximal performance. This dilemma is the focus of this paper. A possible solution to this dilemma is to learn the competence, expertise, or assertiveness of metrics. The basic idea is that each metric has a specific sub-domain for which it is most competent (i.e., it consistently produces more accurate classifiers than the ones produced by other metrics). Particularly, we investigate stacking-based meta-learning methods, which use the training data to find the domain of competence of each metric. The meta-classifier describes the domains of competence (or areas of expertise) of each metric, enabling a more sensible use of these metrics so that competence-conscious classifiers can be produced (i.e., a metric is only used to produce classifiers for test instances that belong to its domain of competence). We conducted a systematic evaluation, using different datasets and evaluation measures, of classifiers produced by different metrics. The result is that, while no metric is always superior than all others, the selection of appropriate metrics according to their competence/expertise (i.e., competence-conscious associative classifiers) seems very effective, showing gains that range from 7% to 26% when compared to the baselines (SVMs and an existing ensemble method). Adriano Veloso, Mohammed J. Zaki, Wagner Meira Jr., Marcos André Gonçalves |
SDM | 1 |
| 2008 | Learning to rank at query-time using association rulesabstractSome applications have to present their results in the form of ranked lists. This is the case of many information retrieval applications, in which documents must be sorted according to their relevance to a given query. This has led the interest of the information retrieval community in methods that automatically learn effective ranking functions. In this paper we propose a novel method which uncovers patterns (or rules) in the training data associating features of the document with its relevance to the query, and then uses the discovered rules to rank documents. To address typical problems that are inherent to the utilization of association rules (such as missing rules and rule explosion), the proposed method generates rules on a demand-driven basis, at query-time. The result is an extremely fast and effective ranking method. We conducted a systematic evaluation of the proposed method using the LETOR benchmark collections. We show that generating rules on a demand-driven basis can boost ranking performance, providing gains ranging from 12 % to 123%, outperforming the state-of-the-art methods that learn to rank, with no need of time-consuming and laborious pre-processing. As a highlight, we also show that additional information, such as query terms, can make the generated rules more discriminative, further improving ranking performance. Adriano Veloso, Humberto Mossri de Almeida, Marcos André Gonçalves, Wagner Meira Jr. |
SIGIR | 1 |
| 2007 | Automatic Moderation of Comments in a Large On-line Journalistic Environment
Adriano Veloso, Wagner Meira Jr., Tiago Alves Macambira, Dorgival O. Guedes, Hélio Marcos Paz de Almeida |
ICWSM | 1 |
| 2007 | Multi-label Lazy Associative Classification
Adriano Veloso, Wagner Meira Jr., Marcos André Gonçalves, Mohammed J. Zaki |
PKDD | 1 |
| 2006 | Multi-evidence, multi-criteria, lazy associative document classificationabstractWe present a novel approach for classifying documents that combines different pieces of evidence (e.g., textual features of documents, links, and citations) transparently, through a data mining technique which generates rules associating these pieces of evidence to predefined classes. These rules can contain any number and mixture of the available evidence and are associated with several quality criteria which can be used in conjunction to choose the "best" rule to be applied at classification time. Our method is able to perform evidence enhancement by link forwarding/backwarding (i.e., navigating among documents related through citation), so that new pieces of link-based evidence are derived when necessary. Furthermore, instead of inducing a single model (or rule set) that is good on average for all predictions, the proposed approach employs a lazy method which delays the inductive process until a document is given for classification, therefore taking advantage of better qualitative evidence coming from the document. We conducted a systematic evaluation of the proposed approach using documents from the ACM Digital Library and from a Brazilian Web directory. Our approach was able to outperform in both collections all classifiers based on the best available evidence in isolation as well as state-of-the-art multi-evidence classifiers. We also evaluated our approach using the standard WebKB collection, where our approach showed gains of 1% in accuracy, being 25 times faster. Further, our approach is extremely efficient in terms of computational performance, showing gains of more than one order of magnitude when compared against other multi-evidence classifiers. Adriano Veloso, Wagner Meira Jr., Marco Cristo, Marcos André Gonçalves, Mohammed J. Zaki |
CIKM | 1 |
| 2006 | Lazy Associative ClassificationabstractDecision tree classifiers perform a greedy search for rules by heuristically selecting the most promising features. Such greedy (local) search may discard important rules. Associative classifiers, on the other hand, perform a global search for rules satisfying some quality constraints (i.e., minimum support). This global search, however, may generate a large number of rules. Further, many of these rules may be useless during classification, and worst, important rules may never be mined. Lazy (non-eager) associative classification overcomes this problem by focusing on the features of the given test instance, increasing the chance of generating more rules that are useful for classifying the test instance. In this paper we assess the performance of lazy associative classification. First we demonstrate that an associative classifier performs no worse than the corresponding decision tree classifier. Also we demonstrate that lazy classifiers outperform the corresponding eager ones. Our claims are empirically confirmed by an extensive set of experimental results. We show that our proposed lazy associative classifier is responsible for an error rate reduction of approximately 10 % when compared against its eager counterpart, and for a reduction of 20 % when compared against a decision tree classifier. A simple caching mechanism makes lazy associative classification fast, and thus improvements in the execution time are also observed. 1 Adriano Veloso, Wagner Meira Jr., Mohammed J. Zaki |
ICDM | 1 |
| 2004 | Asynchronous and Anticipatory Filter-Stream Based Parallel Algorithm for Frequent Itemset Mining
Adriano Veloso, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Srinivasan Parthasarathy 0001 |
PKDD | 1 |
| 2004 | Parallel and distributed methods for incremental frequent itemset miningabstractTraditional methods for data mining typically make the assumption that the data is centralized, memory-resident, and static. This assumption is no longer tenable. Such methods waste computational and input/output (I/O) resources when data is dynamic, and they impose excessive communication overhead when data is distributed. Efficient implementation of incremental data mining methods is, thus, becoming crucial for ensuring system scalability and facilitating knowledge discovery when data is dynamic and distributed. In this paper, we address this issue in the context of the important task of frequent itemset mining. We first present an efficient algorithm which dynamically maintains the required information even in the presence of data updates without examining the entire dataset. We then show how to parallelize this incremental algorithm. We also propose a distributed asynchronous algorithm, which imposes minimal communication overhead for mining distributed dynamic datasets. Our distributed approach is capable of generating local models (in which each site has a summary of its own database) as well as the global model of frequent itemsets (in which all sites have a summary of the entire database). This ability permits our approach not only to generate frequent itemsets, but also to generate high-contrast frequent itemsets, which allows one to examine how the data is skewed over different sites. Matthew Eric Otey, Srinivasan Parthasarathy 0001, Chao Wang 0050, Adriano Veloso, Wagner Meira Jr. |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2003 | Parallel and Distributed Frequent Itemset Mining on Dynamic Datasets
Adriano Veloso, Matthew Eric Otey, Srinivasan Parthasarathy 0001, Wagner Meira Jr. |
HiPC | 1 |
| 2003 | Mining Frequent Itemsets in Distributed and Dynamic DatabasesabstractTraditional methods for frequent itemset mining typically assume that data is centralized and static. Such methods impose excessive communication overhead when data is distributed, and they waste computational resources when data is dynamic. We present what we believe to be the first unified approach that overcomes these assumptions. Our approach makes use of parallel and incremental techniques to generate frequent itemsets in the presence of data updates without examining the entire database, and imposes minimal communication overhead when mining distributed databases. Further, our approach is able to generate both local and global frequent itemsets. This ability permits our approach to identify high-contrast frequent itemsets, which allows one to examine how the data is skewed over different sites. Matthew Eric Otey, Chao Wang 0050, Srinivasan Parthasarathy 0001, Adriano Veloso, Wagner Meira Jr. |
ICDM | 4 |
| 2003 | New Parallel Algorithms for Frequent Itemset Mining in Very Large DatabasesabstractFrequent itemset mining is a classic problem in data mining. It is a nonsupervised process which concerns in finding frequent patterns (or itemsets) hidden in large volumes of data in order to produce compact summaries or models of the database. These models are typically used to generate association rules, but recently they have also been used in far reaching domains like e-commerce and bio-informatics. Because databases are increasing in terms of both dimension (number of attributes) and size (number of records), one of the main issues in a frequent itemset mining algorithm is the ability to analyze very large databases. Sequential algorithms do not have this ability, especially in terms of run-time performance, for such very large databases. Therefore, we must rely on high performance parallel and distributed computing. We present new parallel algorithms for frequent itemset mining. Their efficiency is proven through a series of experiments on different parallel environments, that range from shared-memory multiprocessors machines to a set of SMP clusters connected together through a high speed network. We also briefly discuss an application of our algorithms to the analysis of large databases collected by a Brazilian Web portal. Adriano Veloso, Wagner Meira Jr., Srinivasan Parthasarathy 0001 |
SBAC-PAD | 1 |
| 2002 | Efficiently Mining Approximate Models of Associations in Evolving Databases
Adriano Veloso, Bruno Gusmão Rocha, Wagner Meira Jr., Márcio de Carvalho, Srinivasan Parthasarathy 0001, Mohammed J. Zaki |
PKDD | 1 |
| 2002 | Mining Frequent Itemsets in Evolving Databasesabstract1 Introduction The field of knowledge discovery and data mining (KDD), spurred by advances in data collection technology, is concerned with the process of deriving interesting and useful patterns from large datasets. The KDD process is computational and data-intensive and is inherently interactive and iterative in nature. In fact, interactivity is often the key to facilitating effective data understanding and knowledge discovery. In such an environment, response time is crucial because lengthy time delay between responses of consecutive user requests can disturb the flow of human perception and formation of insight. The task of guaranteeing quick response times is more complicated in dynamic datasets, where there is a constant influx of data. Changes to the data can invalidate existing patterns or introduce new. Simply re-executing algorithms from scratch when a database is updated can result in an explosion in the computational and I/O resources required. What is needed is a way to process the data incrementally and update the information that is gleaned while being cognizant of the interactive requirements of the process. In this paper we present such an approach for a key data mining task: association rule mining. Adriano Veloso, Wagner Meira Jr., Márcio de Carvalho, Bruno Pôssas, Srinivasan Parthasarathy 0001, Mohammed J. Zaki |
SDM | 1 |