EDBT 2026 Demo / reviewers in the wild / expert
Wagner Meira Jr.
dblp:m/WagnerMeiraJr
· DBLP profile ↗
62ranked-venue papers in the field
0as first author
5since 2021 · last 2025
0000-0002-2614-2723ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 25Data Mining & Knowledge Discovery · 21Database Systems & Data Management · 13Knowledge Engineering, Semantic Web & Information Systems · 2Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A CNN-Based Local-Global Self-attention via Averaged Window Embeddings for Hierarchical ECG Analysis
Arthur Buzelin, Pedro Robles Dutenhefner, Turi Rezende, Luisa G. Porfírio, Pedro Bento, Yan Aquino, Jose Fernandes, Caio Santana, Gabriela Miana, Gisele L. Pappa, Antônio L. P. Ribeiro, Wagner Meira Jr. |
ECML/PKDD (3) | 12 |
| 2024 | Topic Shifts as a Proxy for Assessing Politicization in Social MediaabstractPoliticization is a social phenomenon studied by political science characterized by the extent to which ideas and facts are given a political tone. A range of topics, such as climate change, religion and vaccines has been subject to increasing politicization in the media and social media platforms. In this work, we propose a computational method for assessing politicization in online conversations based on topic shifts, i.e., the degree to which people switch topics in online conversations. The intuition is that topic shifts from a non-political topic to politics are a direct measure of politicization – making something political, and that the more people switch conversations to politics, the more they perceive politics as playing a vital role in their daily lives. A fundamental challenge that must be addressed when one studies politicization in social media is that, a priori, any topic may be politicized. Hence, any keyword-based method or even machine learning approaches that rely on topic labels to classify topics are expensive to run and potentially ineffective. Instead, we learn from a seed of political keywords and use Positive-Unlabeled (PU) Learning to detect political comments in reaction to non-political news articles posted on Twitter, YouTube, and TikTok during the 2022 Brazilian presidential elections. Our findings indicate that all platforms show evidence of politicization as discussion around topics adjacent to politics such as economy, crime and drugs tend to shift to politics. Even the least politicized topics had the rate in which their topics shift to politics increased in the lead up to the elections and after other political events in Brazil – an evidence of politicization. The code is available at https://github.com/marceloslo/Topic-Shifts-as-a-Proxy-for-Assessing-Politicization-in-Social-Media. Marcelo Sartori Locatelli, Pedro H. Calais, Matheus Prado Miranda, João Pedro Junho, Tomas Lacerda Muniz, Wagner Meira Jr., Virgílio A. F. Almeida |
ICWSM | 6 |
| 2022 | Counterfactual inference with latent variable and its application in mental health care
Guilherme F. Marchezini, Anísio Lacerda, Gisele L. Pappa, Wagner Meira Jr., Débora M. Miranda, Marco Aurélio Romano-Silva, Danielle S. Costa, Leandro Malloy Diniz |
Data Min. Knowl. Discov. | 4 |
| 2022 | Sequential stratified regeneration: MCMC for large state spaces with an application to subgraph count estimation
Carlos H. C. Teixeira, Mayank Kakodkar, Vinícius Vitor dos Santos Dias, Wagner Meira Jr., Bruno Ribeiro 0001 |
Data Min. Knowl. Discov. | 4 |
| 2021 | Analyzing topic attention in online small groupsabstractAttention is a scarce resource disputed by algorithms and people on the Internet. This competition for attention is part of online spaces especially online small groups where there is a limited number of individuals interacting with each other using text and media content that is not controlled by algorithms or human curators. In these groups, as certain participants and piece of content can catch the collective attention, a question that naturally arises is: how to analyze topic attention in online small groups? In this paper, we propose a methodology aimed at answering this question. Our proposal consists of sets of analyses over topical (obtained from topic analysis) transition graphs for characterizing attention allocation, permanence and shifting as well as participant role characterization during discussions in online small groups. We experimented with our methodology using WhatsApp groups as a case study. Among other results, we identified and characterized abrupt and smooth topic transitions as well as patterns of participant activity related to certain topics. Josemar Alves Caetano, Jussara M. Almeida, Marcos André Gonçalves, Wagner Meira Jr., Humberto Torres Marques-Neto, Virgílio A. F. Almeida |
ASONAM | 4 |
| 2019 | Detecting Spatial Clusters of Disease Infection Risk Using Sparsely Sampled Social Media Mobility PatternsabstractStandard spatial cluster detection methods used in public health surveillance assign each disease case to a single location (typically, the patient's home address), aggregate locations to small areas, and monitor the number of cases in each area over time. However, such methods cannot detect clusters of disease resulting from visits to non-residential locations, such as a park or a university campus. Thus we develop two new spatial scan methods, the unconditional and conditional spatial logistic models, to search for spatial clusters of increased infection risk. We use mobility data from two sets of individuals, disease cases and healthy individuals, where each individual is represented by a sparse sample of geographical locations (e.g., from geo-tagged social media data). The methods account for the multiple, varying number of spatial locations observed per individual, either by non-parametric estimation of the odds of being a case, or by matching case and control individuals with similar numbers of observed locations. Applying our methods to synthetic and real-world scenarios, we demonstrate robust performance on detecting spatial clusters of infection risk from mobility data, outperforming competing baselines. Roberto C. S. N. P. Souza, Renato Assunção, Daniel B. Neill, Wagner Meira Jr. |
SIGSPATIAL/GIS | 4 |
| 2019 | Fractal: A General-Purpose Graph Pattern Mining SystemabstractIn this paper we propose Fractal, a high performance and high productivity system for supporting distributed graph pattern mining (GPM) applications. Fractal employs a dynamic (auto-tuned) load-balancing based on a hierarchical and locality-aware work stealing mechanism, allowing the system to adapt to different workload characteristics. Additionally, Fractal enumerates subgraphs by combining a depth-first strategy with a from scratch processing paradigm to avoid storing large amounts of intermediate state and, thus, improves memory efficiency. Regarding programmer productivity, Fractal presents an intuitive, expressive and modular API, allowing for rapid compositional expression of many GPM algorithms. Fractal-based implementations outperform both existing systemic solutions and specialized distributed solutions on many problems - from frequent graph mining to subgraph querying, over a range of datasets. Vinícius Vitor dos Santos Dias, Carlos H. C. Teixeira, Dorgival O. Guedes, Wagner Meira Jr., Srinivasan Parthasarathy 0001 |
SIGMOD Conference | 4 |
| 2018 | Graph Pattern Mining and Learning through User-Defined RelationsabstractIn this work we propose R-GPM, a parallel computing framework for graph pattern mining (GPM) through a user-defined subgraph relation. More specifically, we enable the computation of statistics of patterns through their subgraph classes, generalizing traditional GPM methods. R-GPM provides efficient estimators for these statistics by employing a MCMC sampling algorithm combined with several optimizations. We provide both theoretical guarantees and empirical evaluations of our estimators in application scenarios such as stochastic optimization of deep high-order graph neural network models and pattern (motif) counting. We also propose and evaluate optimizations that enable improvements of our estimators accuracy, while reducing their computational costs in up to 3-orders-of-magnitude. Finally, we show that R-GPM is scalable, providing near-linear speedups. Carlos H. C. Teixeira, Leornado Cotta, Bruno Ribeiro 0001, Wagner Meira Jr. |
ICDM | 4 |
| 2018 | Characterizing and Detecting Hateful Users on Twitter
Manoel Horta Ribeiro, Pedro H. Calais, Yuri A. Santos, Virgílio A. F. Almeida, Wagner Meira Jr. |
ICWSM | 5 |
| 2018 | An Unsupervised Boosting Strategy for Outlier Detection Ensembles
Guilherme Oliveira Campos, Arthur Zimek, Wagner Meira Jr. |
PAKDD (1) | 3 |
| 2018 | NetClass: A network-based relational model for document classification
Fernando Mourão, Leonardo Rocha 0001, Felipe Viegas, Thiago Salles, Marcos André Gonçalves, Srinivasan Parthasarathy 0001, Wagner Meira Jr. |
Inf. Sci. | 7 |
| 2018 | Scalable and Efficient Data Analytics and Mining with LemonadeabstractProfessionals outside of the area of Computer Science have an increasing need to analyze large bodies of data. This analysis often demands high level of security and has to be done in the cloud. However, current data analysis tools that demand little proficiency in systems programming struggle to deliver solutions which are scalable and safe. In this context we present Lemonade, a platform which focuses on creating data analysis and mining flows in the cloud, with authentication, authorization and accounting (AAA) guarantees. Lemonade provides an interface for the visual construction of flows, and encapsulates storage and data processing environment details, providing higher-level abstractions for data source access and algorithms. We illustrate its usage through a demo, where a data processing flow builds a classification model for detecting fake-news, also extracting some insights along the way. Walter Santos, Gustavo de P. Avelar, Manoel Horta Ribeiro, Dorgival O. Guedes, Wagner Meira Jr. |
Proc. VLDB Endow. | 5 |
| 2017 | Antagonism Also Flows Through Retweets: The Impact of Out-of-Context Quotes in Opinion Polarization Analysis
Pedro Henrique Calais Guerra, Roberto Nalon, Renato Assunção, Wagner Meira Jr. |
ICWSM | 4 |
| 2017 | What surprises does your past have for you?
Fernando Mourão, Leonardo Rocha 0001, Camila Souza Araujo, Wagner Meira Jr., Joseph A. Konstan |
Inf. Syst. | 4 |
| 2017 | A Two-Stage Machine learning approach for temporally-robust text classification
Thiago Salles, Leonardo Rocha 0001, Fernando Mourão, Marcos André Gonçalves, Felipe Viegas, Wagner Meira Jr. |
Inf. Syst. | 6 |
| 2016 | Infection Hot Spot Mining from Social Media Trajectories
Roberto C. S. N. P. Souza, Renato Assunção, Derick M. de Oliveira, Denise E. F. de Brito, Wagner Meira Jr. |
ECML/PKDD (2) | 5 |
| 2016 | A quantitative analysis of the temporal effects on automatic text classificationabstractAutomatic text classification (TC) continues to be a relevant research topic and several TC algorithms have been proposed. However, the majority of TC algorithms assume that the underlying data distribution does not change over time. In this work, we are concerned with the challenges imposed by the temporal dynamics observed in textual data sets. We provide evidence of the existence of temporal effects in three textual data sets, reflected by variations observed over time in the class distribution, in the pairwise class similarities, and in the relationships between terms and classes. We then quantify, using a series of full factorial design experiments, the impact of these effects on four well‐known TC algorithms. We show that these temporal effects affect each analyzed data set differently and that they restrict the performance of each considered TC algorithm to different extents. The reported quantitative analyses, which are the original contributions of this article, provide valuable new insights to better understand the behavior of TC algorithms when faced with nonstatic (temporal) data distributions and highlight important requirements for the proposal of more accurate classification models. Thiago Salles, Leonardo Rocha 0001, Marcos André Gonçalves, Jussara M. Almeida, Fernando Mourão, Wagner Meira Jr., Felipe Viegas |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2015 | Managing Infrastructure-Based Vehicular NetworksabstractIn this thesis work the authors exploit the management of infrastructure-based vehicular networks. The authors begin to work by investigating the most basic problem faced by the network designers when planning an infrastructure-based vehicular network: given a road network, a flow, and α available RSUs, where the RSUs must be located in order to maximize the network performance? The authors propose a novel approach for locating the RSUs: the authors develop a deployment algorithm based on partial mobility information. By partial mobility information, we mean: i) density of vehicles along the road network; and, ii) migration ratios from distinct locations of the road network. Cristiano M. Silva, Wagner Meira Jr. |
MDM (2) | 2 |
| 2015 | Learning sequential classifiers from long and noisy discrete-event sequences efficiently
Gessé Dafé, Adriano Veloso, Mohammed J. Zaki, Wagner Meira Jr. |
Data Min. Knowl. Discov. | 4 |
| 2014 | Reachability Queries in Very Large Graphs: A Fast Refined Online Search ApproachabstractA key problem in many graph-based applications is the need to know, given a directed graph G and two vertices u,v ∈ G, whether there is a path between u and v, i.e., if u reaches v. This problem is particularly challenging in the case of very large real-world graphs. A common approach is the preprocessing of the graphs, in order to produce an efficient index structure, which allows fast access to the reachability information of the vertices. However, the majority of existing methods can not handle very large graphs. We propose, in this paper, a novel indexing method called FELINE (Fast rEfined onLINE search), which is inspired by Dominance Graph Drawing. FELINE creates an index from the graph representation in a two-dimensional plane, which provides reachability information in constant time for a significant portion of queries. Experiments demonstrate the efficiency of FELINE compared to state-of-the-art approaches. Renê Rodrigues Veloso, Loïc Cerf, Wagner Meira Jr., Mohammed J. Zaki |
EDBT | 3 |
| 2014 | Complete discovery of high-quality patterns in large numerical tensorsabstractMany datasets are numerical tensors, i. e., associate n-tuples with numerical values. Until recently, the discovery of relevant local patterns in such numerical and multidimensional data has received little attention despite the broad applicative perspectives offered by this general framework. Even in the simpler 2-dimensional case, almost every proposal so far is either incomplete (i. e., it does not list every pattern) or relies on binning and mines Boolean tensors. In both cases, some information is lost during the process. In uncertain tensors, n-tuples satisfy the studied predicate to a certain extent and no information is lost w.r.t. the original data. Given an uncertain tensor, the closed patterns are its maximal “sub-tensors” covering n-tuples that “mostly” satisfy the predicate. Defining “mostly” is the key problem: the patterns should be both relevant given the data and efficiently extractable. The proposed complete extractor reuses the enumeration principles of the state-of-the-art miner for closed n-sets but incrementally enforces the newly designed definition. In this way, the proposed algorithm runs orders of magnitude faster than its only competitor and large datasets are tractable. The experimental section reports the discovery of dynamic patterns of influence in Twitter as well as usage patterns in a transportation network. Additional experiments on synthetic data quantitatively assess the quality of the chosen definition for the patterns. Loïc Cerf, Wagner Meira Jr. |
ICDE | 2 |
| 2014 | Of Pins and Tweets: Investigating How Users Behave Across Image- and Text-Based Social Networks
Raphael Ottoni, Diego B. Las Casas, João Paulo Pesce, Wagner Meira Jr., Christo Wilson, Alan Mislove, Virgílio A. F. Almeida |
ICWSM | 4 |
| 2014 | Economically-efficient sentiment stream analysisabstractText-based social media channels, such as Twitter, produce torrents of opinionated data about the most diverse topics and entities. The analysis of such data (aka. sentiment analysis) is quickly becoming a key feature in recommender systems and search engines. A prominent approach to sentiment analysis is based on the application of classification techniques, that is, content is classified according to the attitude of the writer. A major challenge, however, is that Twitter follows the data stream model, and thus classifiers must operate with limited resources, including labeled data and time for building classification models. Also challenging is the fact that sentiment distribution may change as the stream evolves. In this paper we address these challenges by proposing algorithms that select relevant training instances at each time step, so that training sets are kept small while providing to the classifier the capabilities to suit itself to, and to recover itself from, different types of sentiment drifts. Simultaneously providing capabilities to the classifier, however, is a conflicting-objective problem, and our proposed algorithms employ basic notions of Economics in order to balance both capabilities. We performed the analysis of events that reverberated on Twitter, and the comparison against the state-of-the-art reveals improvements both in terms of error reduction (up to 14%) and reduction of training resources (by orders of magnitude). Roberto L. de Oliveira Jr., Adriano Veloso, Adriano C. M. Pereira, Wagner Meira Jr., Renato Ferreira 0001, Srinivasan Parthasarathy 0001 |
SIGIR | 4 |
| 2014 | Sentiment analysis on evolving social streams: how self-report imbalances can helpabstractReal-time sentiment analysis is a challenging machine learning task, due to scarcity of labeled data and sudden changes in sentiment caused by real-world events that need to be instantly interpreted. In this paper we propose solutions to acquire labels and cope with concept drift in this setting, by using findings from social psychology on how humans prefer to disclose some types of emotions. In particular, we use findings that humans are more motivated to report positive feelings rather than negative feelings and also prefer to report extreme feelings rather than average feelings. Pedro Henrique Calais Guerra, Wagner Meira Jr., Claire Cardie |
WSDM | 2 |
| 2014 | Approximate similarity search for online multimedia services on distributed CPU-GPU platforms
George Teodoro, Eduardo Valle, Nathan Mariano, Ricardo da Silva Torres, Wagner Meira Jr., Joel H. Saltz |
VLDB J. | 5 |
| 2013 | A Measure of Polarization on Social Media Networks Based on Community Boundaries
Pedro Henrique Calais Guerra, Wagner Meira Jr., Claire Cardie, Robert D. Kleinberg |
ICWSM | 2 |
| 2013 | Ladies First: Analyzing Gender Roles and Behaviors in Pinterest
Raphael Ottoni, João Paulo Pesce, Diego B. Las Casas, Geraldo Franciscani Jr., Wagner Meira Jr., Ponnurangam Kumaraguru, Virgílio A. F. Almeida |
ICWSM | 5 |
| 2013 | Exploiting non-content preference attributes through hybrid recommendation methodabstractThis paper explores a method for incorporating into a recommender system explicit representations of user's preferences over non-content attributes such as popularity, recency, and similarity of recommended items. We show how such attributes can be modeled as a preference vector that can be used in a vector-space content-based recommender, and how that content-based recommender can be integrated with various collaborative filtering techniques through re-weighting of Top-M recommendations. We evaluate this approach on several recommender systems datasets and collaborative filtering methods, and find that incorporating the three preference attributes can lead to a substantial increase in Top-50 precision while also enhancing diversity and novelty. Fernando Mourão, Leonardo Rocha 0001, Joseph A. Konstan, Wagner Meira Jr. |
RecSys | 4 |
| 2013 | A KDD-Based Methodology to Rank Trust in e-Commerce SystemsabstractDue to the growing popularity of the Web, there is an increasing number of people who perform e-business transactions. On the other hand, this popularity has also attracted the attention of criminals, raising the number of frauds on the Web and associated financial losses, which reach billions of dollars per year. This paper proposes a KDD-based methodology to detect fraud in e-payment systems. In order to evaluate this methodology we defined the concept of economic efficiency and applied it to an actual dataset of one of the largest Latin American electronic payment systems. The results show a very good performance, providing gains of up to 46.5% in comparison with the strategy currently employed by the company. José Felipe Júnior, Adriano C. M. Pereira, Wagner Meira Jr., Adriano Veloso |
Web Intelligence | 3 |
| 2013 | Temporal contexts: Effective text classification in evolving document collections
Leonardo Rocha 0001, Fernando Mourão, Hilton de Oliveira Mota, Thiago Salles, Marcos André Gonçalves, Wagner Meira Jr. |
Inf. Syst. | 6 |
| 2012 | Studying User Footprints in Different Online Social NetworksabstractWith the growing popularity and usage of online social media services, people now have accounts (some times several) on multiple and diverse services like Facebook, Linked In, Twitter and You Tube. Publicly available information can be used to create a digital footprint of any user using these social media services. Generating such digital footprints can be very useful for personalization, profile management, detecting malicious behavior of users. A very important application of analyzing users' online digital footprints is to protect users from potential privacy and security risks arising from the huge publicly available user information. We extracted information about user identities on different social networks through Social Graph API, Friend Feed, and Profilactic, we collated our own dataset to create the digital footprints of the users. We used username, display name, description, location, profile image, and number of connections to generate the digital footprints of the user. We applied context specific techniques (e.g. Jaro Winkler similarity, Word net based ontologies) to measure the similarity of the user profiles on different social networks. We specifically focused on Twitter and Linked In. In this paper, we present the analysis and results from applying automated classifiers for disambiguating profiles belonging to the same user from different social networks. User ID and Name were found to be the most discriminative features for disambiguating user profiles. Using the most promising set of features and similarity metrics, we achieved accuracy, precision and recall of 98%, 99%, and 96%, respectively. Anshu Malhotra, Luam C. Totti, Wagner Meira Jr., Ponnurangam Kumaraguru, Virgílio A. F. Almeida |
ASONAM | 3 |
| 2012 | Cost-effective on-demand associative author name disambiguation
Adriano Veloso, Anderson A. Ferreira, Marcos André Gonçalves, Alberto H. F. Laender, Wagner Meira Jr. |
Inf. Process. Manag. | 5 |
| 2012 | Mining Attribute-structure Correlated Patterns in Large Attributed GraphsabstractIn this work, we study the correlation between attribute sets and the occurrence of dense subgraphs in large attributed graphs, a task we call structural correlation pattern mining. A structural correlation pattern is a dense subgraph induced by a particular attribute set. Existing methods are not able to extract relevant knowledge regarding how vertex attributes interact with dense subgraphs. Structural correlation pattern mining combines aspects of frequent itemset and quasi-clique mining problems. We propose statistical significance measures that compare the structural correlation of attribute sets against their expected values using null models. Moreover, we evaluate the interestingness of structural correlation patterns in terms of size and density. An efficient algorithm that combines search and pruning strategies in the identification of the most relevant structural correlation patterns is presented. We apply our method for the analysis of three real-world attributed graphs: a collaboration, a music, and a citation network, verifying that it provides valuable knowledge in a feasible time. Arlei Silva, Wagner Meira Jr., Mohammed J. Zaki |
Proc. VLDB Endow. | 2 |
| 2011 | Adaptive parallel approximate similarity search for responsive multimedia retrievalabstractThis paper introduces Hypercurves, a flexible framework for pro- viding similarity search indexing to high throughput multimedia services. Hypercurves efficiently and effectively answers k-nearest neighbor searches on multigigabyte high-dimensional databases. It supports massively parallel processing and adapts at runtime its parallelization regimens to keep answer times optimal for either low and high demands. In order to achieve its goals, Hypercurves introduces new techniques for selecting parallelism configurations and allocating threads to computation cores, including hyperthreaded cores. Its efficiency gains are throughly validated on a large database of multimedia descriptors, where it presented near linear speedups and superlinear scaleups. The adaptation reduces query response times in 43% and 74% for both platforms tested, when compared to the best static parallelism regimens. George Teodoro, Eduardo Valle, Nathan Mariano, Ricardo da Silva Torres, Wagner Meira Jr. |
CIKM | 5 |
| 2011 | From bias to opinion: a transfer-learning approach to real-time sentiment analysisabstractReal-time interaction, which enables live discussions, has become a key feature of most Web applications. In such an environment, the ability to automatically analyze user opinions and sentiments as discussions develop is a powerful resource known as real time sentiment analysis. However, this task comes with several challenges, including the need to deal with highly dynamic textual content that is characterized by changes in vocabulary and its subjective meaning and the lack of labeled data needed to support supervised classifiers. In this paper, we propose a transfer learning strategy to perform real time sentiment analysis. We identify a task - opinion holder bias prediction - which is strongly related to the sentiment analysis task; however, in constrast to sentiment analysis, it builds accurate models since the underlying relational data follows a stationary distribution. Pedro Henrique Calais Guerra, Adriano Veloso, Wagner Meira Jr., Virgílio A. F. Almeida |
KDD | 3 |
| 2011 | Is There a Best Quality Metric for Graph Clusters?
Hélio Marcos Paz de Almeida, Dorgival O. Guedes, Wagner Meira Jr., Mohammed J. Zaki |
ECML/PKDD (1) | 3 |
| 2011 | Data Integration via Constrained Clustering: An Application to Enzyme ClusteringabstractWhen multiple data sources are available for clustering, an a priori data integration process is usually required. This process may be costly and may not lead to good clusterings, since important information is likely to be discarded. In this paper we propose constrained clustering as a strategy for integrating data sources without losing any information. It basically consists of adding the complementary data sources as constraints that the algorithm must satisfy. As a concrete application of our approach, we focus on the problem of enzyme function prediction, which is a hard task usually performed by intensive experimental work. We use constrained clustering as a means of integrating information from diverse sources as constraints, and analyze how this additional information impacts clustering quality in an enzyme clustering application scenario. Our results show that constraints generally improve the clustering quality when compared to an unconstrained clustering algorithm. Elisa Boari de Lima, Raquel Cardoso de Melo Minardi, Wagner Meira Jr., Mohammed J. Zaki |
SDM | 3 |
| 2011 | Effective sentiment stream analysis with self-augmenting training and demand-driven projectionabstractHow do we analyze sentiments over a set of opinionated Twitter messages? This issue has been widely studied in recent years, with a prominent approach being based on the application of classification techniques. Basically, messages are classified according to the implicit attitude of the writer with respect to a query term. A major concern, however, is that Twitter (and other media channels) follows the data stream model, and thus the classifier must operate with limited resources, including labeled data for training classification models. This imposes serious challenges for current classification techniques, since they need to be constantly fed with fresh training messages, in order to track sentiment drift and to provide up-to-date sentiment analysis. Ismael S. Silva, Janaína Gomide, Adriano Veloso, Wagner Meira Jr., Renato Ferreira 0001 |
SIGIR | 4 |
| 2011 | Word co-occurrence features for text classification
Fábio Figueiredo, Leonardo Rocha 0001, Thierson Couto, Thiago Salles, Marcos André Gonçalves, Wagner Meira Jr. |
Inf. Syst. | 6 |
| 2011 | Calibrated lazy associative classification
Adriano Veloso, Wagner Meira Jr., Marcos André Gonçalves, Humberto Mossri de Almeida, Mohammed J. Zaki |
Inf. Sci. | 2 |
| 2010 | Temporally-aware algorithms for document classificationabstractAutomatic Document Classification (ADC) is still one of the major information retrieval problems. It usually employs a supervised learning strategy, where we first build a classification model using pre-classified documents and then use this model to classify unseen documents. The majority of supervised algorithms consider that all documents provide equally important information. However, in practice, a document may be considered more or less important to build the classification model according to several factors, such as its timeliness, the venue where it was published in, its authors, among others. In this paper, we are particularly concerned with the impact that temporal effects may have on ADC and how to minimize such impact. In order to deal with these effects, we introduce a temporal weighting function (TWF) and propose a methodology to determine it for document collections. We applied the proposed methodology to ACM-DL and Medline and found that the TWF of both follows a lognormal. We then extend three ADC algorithms (namely kNN, Rocchio and Naïve Bayes) to incorporate the TWF. Experiments showed that the temporally-aware classifiers achieved significant gains, outperforming (or at least matching) state-of-the-art algorithms. Thiago Salles, Leonardo Rocha 0001, Gisele L. Pappa, Fernando Mourão, Wagner Meira Jr., Marcos André Gonçalves |
SIGIR | 5 |
| 2010 | Distance-Based Outlier Detection: Consolidation and Renewed BearingabstractDetecting outliers in data is an important problem with interesting applications in a myriad of domains ranging from data cleaning to financial fraud detection and from network intrusion detection to clinical diagnosis of diseases. Over the last decade of research, distance-based outlier detection algorithms have emerged as a viable, scalable, parameter-free alternative to the more traditional statistical approaches. In this paper we assess several distance-based outlier detection approaches and evaluate them. We begin by surveying and examining the design landscape of extant approaches, while identifying key design decisions of such approaches. We then implement an outlier detection framework and conduct a factorial design experiment to understand the pros and cons of various optimizations proposed by us as well as those proposed in the literature, both independently and in conjunction with one another, on a diverse set of real-life datasets. To the best of our knowledge this is the first such study in the literature. The outcome of this study is a family of state of the art distance-based outlier detection algorithms. Our detailed empirical study supports the following observations. The combination of optimization strategies enables significant efficiency gains. Our factorial design study highlights the important fact that no single optimization or combination of optimizations (factors) always dominates on all types of data. Our study also allows us to characterize when a certain combination of optimizations is likely to prevail and helps provide interesting and useful insights for moving forward in this domain. Gustavo Henrique Orair, Carlos H. C. Teixeira, Wagner Meira Jr., Srinivasan Parthasarathy 0001 |
Proc. VLDB Endow. | 4 |
| 2009 | The Metric Dilemma: Competence-Conscious Associative ClassificationabstractThe classification performance of an associative classifier is strongly dependent on the statistic measure or metric that is used to quantify the strength of the association between features and classes (i.e., confidence, correlation etc.). Previous studies have shown that classifiers produced by different metrics may provide conflicting predictions, and that the best metric to use is data-dependent and rarely known while designing the classifier. This uncertainty concerning the optimal match between metrics and problems is a dilemma, and prevents associative classifiers to achieve their maximal performance. This dilemma is the focus of this paper. A possible solution to this dilemma is to learn the competence, expertise, or assertiveness of metrics. The basic idea is that each metric has a specific sub-domain for which it is most competent (i.e., it consistently produces more accurate classifiers than the ones produced by other metrics). Particularly, we investigate stacking-based meta-learning methods, which use the training data to find the domain of competence of each metric. The meta-classifier describes the domains of competence (or areas of expertise) of each metric, enabling a more sensible use of these metrics so that competence-conscious classifiers can be produced (i.e., a metric is only used to produce classifiers for test instances that belong to its domain of competence). We conducted a systematic evaluation, using different datasets and evaluation measures, of classifiers produced by different metrics. The result is that, while no metric is always superior than all others, the selection of appropriate metrics according to their competence/expertise (i.e., competence-conscious associative classifiers) seems very effective, showing gains that range from 7% to 26% when compared to the baselines (SVMs and an existing ensemble method). Adriano Veloso, Mohammed J. Zaki, Wagner Meira Jr., Marcos André Gonçalves |
SDM | 3 |
| 2009 | Analyzing seller practices in a Brazilian marketplaceabstractE-commerce is growing at an exponential rate. In the last decade, there has been an explosion of online commercial activity enabled by World Wide Web (WWW). These days, many consumers are less attracted to online auctions, preferring to buy merchandise quickly using fixed-price negotiations. Sales at Amazon.com, the leader in online sales of fixed-price goods, rose 37% in the first quarter of 2008. At eBay, where auctions make up 58% of the site's sales, revenue rose 14%. In Brazil, probably by cultural influence, online auctions are not been popular. This work presents a characterization and analysis of fixed-price online negotiations. Using actual data from a Brazilian marketplace, we analyze seller practices, considering seller profiles and strategies. We show that different sellers adopt strategies according to their interests, abilities and experience. Moreover, we confirm that choosing a selling strategy is not simple, since it is important to consider the seller's characteristics to evaluate the applicability of a strategy. The work also provides a comparative analysis of some selling practices in Brazil with popular worldwide marketplaces. Adriano C. M. Pereira, Diego Duarte, Wagner Meira Jr., Virgílio A. F. Almeida, Paulo B. Góes |
WWW | 3 |
| 2008 | Exploiting temporal contexts in text classificationabstractDue to the increasing amount of information being stored and accessible through the Web, Automatic Document Classification (ADC) has become an important research topic. ADC usually employs a supervised learning strategy, where we first build a classification model using pre-classified documents and then use it to classify unseen documents. One major challenge in building classifiers is dealing with the temporal evolution of the characteristics of the documents and the classes to which they belong. However, most of the current techniques for ADC do not consider this evolution while building and using the models. Previous results show that the performance of classifiers may be affected by three different temporal effects (class distribution, term distribution and class similarity). Further, it is shown that using just portions of the pre-classified documents, which we call contexts, for building the classifiers, result in better performance, as a consequence of the minimization of the aforementioned effects. Leonardo Rocha 0001, Fernando Mourão, Adriano C. M. Pereira, Marcos André Gonçalves, Wagner Meira Jr. |
CIKM | 5 |
| 2008 | Learning to rank at query-time using association rulesabstractSome applications have to present their results in the form of ranked lists. This is the case of many information retrieval applications, in which documents must be sorted according to their relevance to a given query. This has led the interest of the information retrieval community in methods that automatically learn effective ranking functions. In this paper we propose a novel method which uncovers patterns (or rules) in the training data associating features of the document with its relevance to the query, and then uses the discovered rules to rank documents. To address typical problems that are inherent to the utilization of association rules (such as missing rules and rule explosion), the proposed method generates rules on a demand-driven basis, at query-time. The result is an extremely fast and effective ranking method. We conducted a systematic evaluation of the proposed method using the LETOR benchmark collections. We show that generating rules on a demand-driven basis can boost ranking performance, providing gains ranging from 12 % to 123%, outperforming the state-of-the-art methods that learn to rank, with no need of time-consuming and laborious pre-processing. As a highlight, we also show that additional information, such as query terms, can make the generated rules more discriminative, further improving ranking performance. Adriano Veloso, Humberto Mossri de Almeida, Marcos André Gonçalves, Wagner Meira Jr. |
SIGIR | 4 |
| 2008 | Understanding temporal aspects in document classificationabstractDue to the increasing amount of information present on the Web, Automatic Document Classification (ADC) has become an important research topic. ADC usually follows a standard supervised learning strategy, where we first build a model using preclassified documents and then use it to classify new unseen documents. One major challenge for ADC in many scenarios is that the characteristics of the documents and the classes to which they belong may change over time. However, most of the current techniques for ADC are applied without taking into account the temporal evolution of the collection of documents Fernando Mourão, Leonardo Rocha 0001, Renata Braga Araújo, Thierson Couto, Marcos André Gonçalves, Wagner Meira Jr. |
WSDM | 6 |
| 2007 | Automatic Moderation of Comments in a Large On-line Journalistic Environment
Adriano Veloso, Wagner Meira Jr., Tiago Alves Macambira, Dorgival O. Guedes, Hélio Marcos Paz de Almeida |
ICWSM | 2 |
| 2007 | Multi-label Lazy Associative Classification
Adriano Veloso, Wagner Meira Jr., Marcos André Gonçalves, Mohammed J. Zaki |
PKDD | 2 |
| 2006 | Multi-evidence, multi-criteria, lazy associative document classificationabstractWe present a novel approach for classifying documents that combines different pieces of evidence (e.g., textual features of documents, links, and citations) transparently, through a data mining technique which generates rules associating these pieces of evidence to predefined classes. These rules can contain any number and mixture of the available evidence and are associated with several quality criteria which can be used in conjunction to choose the "best" rule to be applied at classification time. Our method is able to perform evidence enhancement by link forwarding/backwarding (i.e., navigating among documents related through citation), so that new pieces of link-based evidence are derived when necessary. Furthermore, instead of inducing a single model (or rule set) that is good on average for all predictions, the proposed approach employs a lazy method which delays the inductive process until a document is given for classification, therefore taking advantage of better qualitative evidence coming from the document. We conducted a systematic evaluation of the proposed approach using documents from the ACM Digital Library and from a Brazilian Web directory. Our approach was able to outperform in both collections all classifiers based on the best available evidence in isolation as well as state-of-the-art multi-evidence classifiers. We also evaluated our approach using the standard WebKB collection, where our approach showed gains of 1% in accuracy, being 25 times faster. Further, our approach is extremely efficient in terms of computational performance, showing gains of more than one order of magnitude when compared against other multi-evidence classifiers. Adriano Veloso, Wagner Meira Jr., Marco Cristo, Marcos André Gonçalves, Mohammed J. Zaki |
CIKM | 2 |
| 2006 | Lazy Associative ClassificationabstractDecision tree classifiers perform a greedy search for rules by heuristically selecting the most promising features. Such greedy (local) search may discard important rules. Associative classifiers, on the other hand, perform a global search for rules satisfying some quality constraints (i.e., minimum support). This global search, however, may generate a large number of rules. Further, many of these rules may be useless during classification, and worst, important rules may never be mined. Lazy (non-eager) associative classification overcomes this problem by focusing on the features of the given test instance, increasing the chance of generating more rules that are useful for classifying the test instance. In this paper we assess the performance of lazy associative classification. First we demonstrate that an associative classifier performs no worse than the corresponding decision tree classifier. Also we demonstrate that lazy classifiers outperform the corresponding eager ones. Our claims are empirically confirmed by an extensive set of experimental results. We show that our proposed lazy associative classifier is responsible for an error rate reduction of approximately 10 % when compared against its eager counterpart, and for a reduction of 20 % when compared against a decision tree classifier. A simple caching mechanism makes lazy associative classification fast, and thus improvements in the execution time are also observed. 1 Adriano Veloso, Wagner Meira Jr., Mohammed J. Zaki |
ICDM | 2 |
| 2005 | Maximal termsets as a query structuring mechanismabstractSearch engines process queries conjunctively to restrict the size of the answer set. Further, it is not rare to observe a mismatch between the vocabulary used in the text of Web pages and the terms used to compose the Web queries. The combination of these two features might lead to irrelevant query results, particularly in the case of more specific queries composed of three or more terms. To deal with this problem we propose a new technique for automatically structuring Web queries as a set of smaller subqueries. To select representative subqueries we use information on their distributions in the document collection. This can be adequately modeled using the concept of maximal termsets derived from the formalism of association rules theory. Experimentation shows that our technique leads to improved results. For the TREC-8 test collection, for instance, our technique led to gains in average precision of roughly 28% with regard to a BM25 ranking formula. Bruno Pôssas, Nivio Ziviani, Berthier A. Ribeiro-Neto, Wagner Meira Jr. |
CIKM | 4 |
| 2005 | Set-based vector model: An efficient approach for correlation-based rankingabstractThis work presents a new approach for ranking documents in the vector space model. The novelty lies in two fronts. First, patterns of term co-occurrence are taken into account and are processed efficiently. Second, term weights are generated using a data mining technique called association rules. This leads to a new ranking mechanism called the set-based vector model . The components of our model are no longer index terms but index termsets, where a termset is a set of index terms. Termsets capture the intuition that semantically related terms appear close to each other in a document. They can be efficiently obtained by limiting the computation to small passages of text. Once termsets have been computed, the ranking is calculated as a function of the termset frequency in the document and its scarcity in the document collection. Experimental results show that the set-based vector model improves average precision for all collections and query types evaluated, while keeping computational costs small. For the 2-gigabyte TREC-8 collection, the set-based vector model leads to a gain in average precision figures of 14.7% and 16.4% for disjunctive and conjunctive queries, respectively, with respect to the standard vector space model. These gains increase to 24.9% and 30.0%, respectively, when proximity information is taken into account. Query processing times are larger but, on average, still comparable to those obtained with the standard vector model (increases in processing time varied from 30% to 300%). Our results suggest that the set-based vector model provides a correlation-based ranking formula that is effective with general collections and computationally practical. Bruno Pôssas, Nivio Ziviani, Wagner Meira Jr., Berthier A. Ribeiro-Neto |
ACM Trans. Inf. Syst. | 3 |
| 2004 | Asynchronous and Anticipatory Filter-Stream Based Parallel Algorithm for Frequent Itemset Mining
Adriano Veloso, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Srinivasan Parthasarathy 0001 |
PKDD | 2 |
| 2004 | Processing Conjunctive and Phrase Queries with the Set-Based Model
Bruno Pôssas, Nivio Ziviani, Berthier A. Ribeiro-Neto, Wagner Meira Jr. |
SPIRE | 4 |
| 2003 | Mining Frequent Itemsets in Distributed and Dynamic DatabasesabstractTraditional methods for frequent itemset mining typically assume that data is centralized and static. Such methods impose excessive communication overhead when data is distributed, and they waste computational resources when data is dynamic. We present what we believe to be the first unified approach that overcomes these assumptions. Our approach makes use of parallel and incremental techniques to generate frequent itemsets in the presence of data updates without examining the entire database, and imposes minimal communication overhead when mining distributed databases. Further, our approach is able to generate both local and global frequent itemsets. This ability permits our approach to identify high-contrast frequent itemsets, which allows one to examine how the data is skewed over different sites. Matthew Eric Otey, Chao Wang 0050, Srinivasan Parthasarathy 0001, Adriano Veloso, Wagner Meira Jr. |
ICDM | 5 |
| 2002 | Efficiently Mining Approximate Models of Associations in Evolving Databases
Adriano Veloso, Bruno Gusmão Rocha, Wagner Meira Jr., Márcio de Carvalho, Srinivasan Parthasarathy 0001, Mohammed J. Zaki |
PKDD | 3 |
| 2002 | Mining Frequent Itemsets in Evolving Databasesabstract1 Introduction The field of knowledge discovery and data mining (KDD), spurred by advances in data collection technology, is concerned with the process of deriving interesting and useful patterns from large datasets. The KDD process is computational and data-intensive and is inherently interactive and iterative in nature. In fact, interactivity is often the key to facilitating effective data understanding and knowledge discovery. In such an environment, response time is crucial because lengthy time delay between responses of consecutive user requests can disturb the flow of human perception and formation of insight. The task of guaranteeing quick response times is more complicated in dynamic datasets, where there is a constant influx of data. Changes to the data can invalidate existing patterns or introduce new. Simply re-executing algorithms from scratch when a database is updated can result in an explosion in the computational and I/O resources required. What is needed is a way to process the data incrementally and update the information that is gleaned while being cognizant of the interactive requirements of the process. In this paper we present such an approach for a key data mining task: association rule mining. Adriano Veloso, Wagner Meira Jr., Márcio de Carvalho, Bruno Pôssas, Srinivasan Parthasarathy 0001, Mohammed J. Zaki |
SDM | 2 |
| 2002 | Set-based model: a new approach for information retrievalabstractThe objective of this paper is to present a new technique for computing term weights for index terms, which leads to a new ranking mechanism, referred to as set-based model. The components in our model are no longer terms, but termsets. The novelty is that we compute term weights using a data mining technique called association rules, which is time efficient and yet yields nice improvements in retrieval effectiveness. The set-based model function for computing the similarity between a document and a query considers the termset frequency in the document and its scarcity in the document collection. Experimental results show that our model improves the average precision of the answer set for all three collections evaluated. For the TReC-3 collection, our set-based model led to a gain, relative to the standard vector space model, of 37% in average precision curves and of 57% in average precision for the top 10 documents. Like the vector space model, the set-based model has time complexity that is linear in the number of documents in the collection. Bruno Pôssas, Nivio Ziviani, Wagner Meira Jr., Berthier A. Ribeiro-Neto |
SIGIR | 3 |
| 2002 | Enhancing the Set-Based Model Using Proximity Information
Bruno Pôssas, Nivio Ziviani, Wagner Meira Jr. |
SPIRE | 3 |
| 2001 | Rank-Preserving Two-Level Caching for Scalable Search EnginesabstractArticle Rank-preserving two-level caching for scalable search engines Share on Authors: Patricia Correia Saraiva Federal Univ. of Minas Gerais, Belo Horizonte, Brazil and Federal Univ. of Amazonas, Manaus, Brazil Federal Univ. of Minas Gerais, Belo Horizonte, Brazil and Federal Univ. of Amazonas, Manaus, BrazilView Profile , Edleno Silva de Moura Akwan Information Technologies, Belo Horizonte, Brazil Akwan Information Technologies, Belo Horizonte, BrazilView Profile , Nivio Ziviani Federal Univ. of Minas Gerais, Belo Horizonte, Brazil Federal Univ. of Minas Gerais, Belo Horizonte, BrazilView Profile , Wagner Meira Federal Univ. of Minas Gerais, Belo Horizonte, Brazil Federal Univ. of Minas Gerais, Belo Horizonte, BrazilView Profile , Rodrigo Fonseca Univ. of Minas, Belo Horizonte, Brazil Univ. of Minas, Belo Horizonte, BrazilView Profile , Berthier Ribeiro-Neto Federal Univ. of Minas Gerias, Belo Horizonte, Brazil Federal Univ. of Minas Gerias, Belo Horizonte, BrazilView Profile Authors Info & Claims SIGIR '01: Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrievalSeptember 2001 Pages 51–58https://doi.org/10.1145/383952.383959Published:01 September 2001 89citation981DownloadsMetricsTotal Citations89Total Downloads981Last 12 Months8Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Patricia Correia Saraiva, Edleno Silva de Moura, Rodrigo Fonseca, Wagner Meira Jr., Berthier A. Ribeiro-Neto, Nivio Ziviani |
SIGIR | 4 |
| 2000 | Dynamic Aspects of Documents of the Brazilian WebabstractThe huge number and the textual nature of Web documents have created the need for search engines. In this kind of system the user's queries are answered based on a static view of the whole or part of the Web. However, the Web is a dynamic system and its documents are inserted, changed and removed frequently, thus creating an inconsistency between the state of the Web and the static view of the documents of the search engine. We describe the dynamic aspects of the HTML documents of the Brazilian Web, namely the rate of insertion, change and removal of its documents, from the point of view of search engines. We also describe the tools we used to support this study. Whenever possible, we compare measures of the Brazilian Web with related measures of the World Wide Web. We show how the study of these aspects can be used in search engines to improve the quality of its services, in particular, aspects related to robots or spiders. Nahur Fonseca, Rodolfo F. Resende, Wagner Meira Jr. |
WISE | 3 |