Wagner Meira Jr.

dblp:m/WagnerMeiraJr · DBLP profile ↗
← Back
165ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0002-2614-2723ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 62 · 5 since 2021Artificial intelligence and machine learning · 38 · 9 since 2021Systems, architecture and hardware · 25 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 6 since 2021Computer networks · 15 · 1 since 2021Human-computer interaction and ubiquitous computing · 12 · 3 since 2021Software engineering, systems software and programming languages · 6Security and privacy · 3Theory of computation · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Monotonic Scaffolding as a Diagnostic Lens for Legal Reasoning in LLMs
abstract
Modern evaluation of Legal QA systems is shifting from terminal accuracy toward processaware analyses of model reasoning.We propose a diagnostic framework grounded in monotonic pedagogical scaffolding, where language models receive gold-standard, caserelevant information across stages aligned with the canonical legal framework FIRAC -Facts, Issue, Rules, Application, Conclusion.By strictly adding solution-relevant content at each step, we introduce a controlled monotonic intervention that allows for the evaluation of reasoning trajectories rather than isolated outcomes.This longitudinal design enables the introduction of two transition-based diagnostics: Errorsto-Success (E2S) quantifies the guidance required to reach correctness, while Success-to-Errors (S2E) measures the fragility of that correctness under additional structure.These local patterns define a global robustness criterion termed Stable Accuracy, which credits a response only if the model maintains correctness throughout all scaffolding stages and enforces a higher bar for correctness by distinguishing sustained reasoning from transient patterns.We instantiate the framework on 3,123 Brazilian Bar Exam questions paired with expertannotated explanations.Our findings reveal model instability patterns hidden from accuracy-only metrics and demonstrate that terminal accuracy systematically overestimates legal reasoning competence.To test the robustness of our diagnostics, we also evaluate a majority-vote aggregation across multiple reasoning samples, finding that the observed instability patterns persist under this stronger inference setting.Furthermore, principal component analysis indicates that legal domains cluster into distinct regions, suggesting systematic differences in reasoning demands across domains.While focused on the legal domain, our evaluation protocol is generalizable to any task with a staged reasoning structure.
Pedro Henrique Calais Guerra, Janderson Santos, Anísio Lacerda, Wagner Meira Jr.
ACL (1)4
2026 Evo-Reasoner: Evolutionary Optimization of Structured Reasoning in LLMs
Pedro Bento, Arthur Buzelin, Arthur Chagas, Yan Aquino, Wagner Meira Jr., Gisele L. Pappa
EvoApplications5
2025 Evolutionary Bias Identification with Embeddings
Arthur Buzelin, Yan Aquino, Victoria Estanislau, Pedro Bento, Lucas Dayrell, Samira Malaquias, Caio Santana, Guilherme H. G. Evangelista, Caio Souza Grossi, Pedro B. Rigueira, Luisa G. Porfírio, Marcelo Sartori Locatelli, Wagner Meira Jr., Gisele L. Pappa
EvoApplications (2)13
2025 Automatic Aesthetic Evaluation in Generative Image Models
Larissa D. Gomide, Lucas Nascimento Ferreira, Wagner Meira Jr.
ICCC3
2025 Evaluation of Medical Large Language Models: Taxonomy, Review, and Directions
abstract
The integration of Large Language Models (LLMs) into medicine presents both great opportunities and significant challenges, particularly in ensuring these models are accurate, reliable, and safe. While LLMs have shown impressive capabilities in understanding and generating human language, their application in the medical domain requires careful evaluation due to the critical nature of medical applications which are inherently linked to patient life and health. Current evaluations of LLMs in medicine are often fragmented and insufficient, with a lack of standardized performance metrics, limited use of real patient data, and insufficient attention to important applications, such as documentation, education, and research. Furthermore, traditional NLP-based evaluations are often inadequate for assessing the text generated by LLMs. Therefore, a robust evaluation is essential to ensure the responsible and effective use of LLMs in medical settings, and to address the inherent challenges associated with their implementation. This paper explores the various dimensions of LLM evaluation in the medical domain, proposes a new taxonomy for categorizing medical applications, and discusses directions for future research in this critical area.
Anísio Lacerda, Gisele L. Pappa, Adriano C. M. Pereira, Wagner Meira Jr., Alexandre Guimarães de Almeida Barros
IJCAI4
2025 A CNN-Based Local-Global Self-attention via Averaged Window Embeddings for Hierarchical ECG Analysis
Arthur Buzelin, Pedro Robles Dutenhefner, Turi Rezende, Luisa G. Porfírio, Pedro Bento, Yan Aquino, Jose Fernandes, Caio Santana, Gabriela Miana, Gisele L. Pappa, Antônio L. P. Ribeiro, Wagner Meira Jr.
ECML/PKDD (3)12
2025 STEval: A framework for evaluating spatio-temporal crime prediction models
Gabriel Amarante, Matheus Pimenta, Yan Andrade, Matheus Senna, Rainer Menezes, Antônio Hot Faria, Marcelo Vilas-Boas, Frederico Martins de Paula Neto, João Paulo da Silva, Everton Renato de Sousa, Jamicel da Silva, Wagner Meira Jr., George Teodoro, Leonardo Rocha 0001, Renato Ferreira 0001
Eng. Appl. Artif. Intell.12
2024 Examining the Behavior of LLM Architectures Within the Framework of Standardized National Exams in Brazil
abstract
The Exame Nacional do Ensino Médio (ENEM) is a pivotal test for Brazilian students, required for admission to a significant number of universities in Brazil. The test consists of four objective high-school level tests on Math, Humanities, Natural Sciences and Languages, and one writing essay. Students' answers to the test and to the accompanying socioeconomic status questionnaire are made public every year (albeit anonymized) due to transparency policies from the Brazilian Government. In the context of large language models (LLMs), these data lend themselves nicely to comparing different groups of humans with AI, as we can have access to human and machine answer distributions. We leverage these characteristics of the ENEM dataset and compare GPT-3.5 and 4, and MariTalk, a model trained using Portuguese data, to humans, aiming to ascertain how their answers relate to real societal groups and what that may reveal about the model biases. We divide the human groups by using socioeconomic status (SES), and compare their answer distribution with LLMs for each question and for the essay. We find no significant biases when comparing LLM performance to humans on the multiple-choice Brazilian Portuguese tests, as the distance between model and human answers is mostly determined by the human accuracy. A similar conclusion is found by looking at the generated text as, when analyzing the essays, we observe that human and LLM essays differ in a few key factors, one being the choice of words where model essays were easily separable from human ones. The texts also differ syntactically, with LLM generated essays exhibiting, on average, smaller sentences and less thought units, among other differences. These results suggest that, for Brazilian Portuguese in the ENEM context, LLM outputs represent no group of humans, being significantly different from the answers from Brazilian students across all tests. The appendices may be found at https://arxiv.org/abs/2408.05035.
Marcelo Sartori Locatelli, Matheus Prado Miranda, Igor Joaquim da Silva Costa, Matheus Torres Prates, Victor Thomé, Mateus Zaparoli Monteiro, Tomas Lacerda, Adriana S. Pagano, Eduardo Rios Neto, Wagner Meira Jr., Virgílio A. F. Almeida
AIES (1)10
2024 A Descriptive and Predictive Analysis Tool for Criminal Data: A Case Study from Brazil
Yan Andrade, Matheus Pimenta, Gabriel Amarante, Antônio Hot Faria, Marcelo Vilas-Boas, João Paulo da Silva, Felipe Rocha, Jamicel da Silva, Wagner Meira Jr., George Teodoro, Leonardo Rocha 0001, Renato Ferreira 0001
ICCSA (2)9
2024 Topic Shifts as a Proxy for Assessing Politicization in Social Media
abstract
Politicization is a social phenomenon studied by political science characterized by the extent to which ideas and facts are given a political tone. A range of topics, such as climate change, religion and vaccines has been subject to increasing politicization in the media and social media platforms. In this work, we propose a computational method for assessing politicization in online conversations based on topic shifts, i.e., the degree to which people switch topics in online conversations. The intuition is that topic shifts from a non-political topic to politics are a direct measure of politicization – making something political, and that the more people switch conversations to politics, the more they perceive politics as playing a vital role in their daily lives. A fundamental challenge that must be addressed when one studies politicization in social media is that, a priori, any topic may be politicized. Hence, any keyword-based method or even machine learning approaches that rely on topic labels to classify topics are expensive to run and potentially ineffective. Instead, we learn from a seed of political keywords and use Positive-Unlabeled (PU) Learning to detect political comments in reaction to non-political news articles posted on Twitter, YouTube, and TikTok during the 2022 Brazilian presidential elections. Our findings indicate that all platforms show evidence of politicization as discussion around topics adjacent to politics such as economy, crime and drugs tend to shift to politics. Even the least politicized topics had the rate in which their topics shift to politics increased in the lead up to the elections and after other political events in Brazil – an evidence of politicization. The code is available at https://github.com/marceloslo/Topic-Shifts-as-a-Proxy-for-Assessing-Politicization-in-Social-Media.
Marcelo Sartori Locatelli, Pedro H. Calais, Matheus Prado Miranda, João Pedro Junho, Tomas Lacerda Muniz, Wagner Meira Jr., Virgílio A. F. Almeida
ICWSM6
2024 DuMato: An efficient warp-centric subgraph enumeration system for GPU
Samuel Ferraz, Vinícius Vitor dos Santos Dias, Carlos H. C. Teixeira, Srinivasan Parthasarathy 0001, George Teodoro, Wagner Meira Jr.
J. Parallel Distributed Comput.6
2023 Graph Pattern Mining Paradigms: Consolidation and Renewed Bearing
abstract
Graph Pattern Mining (GPM) refers to a class of problems involving the processing of sub graphs extracted from larger graphs. Applications to GPM algorithms include querying subgraphs, identifying motif structures in biological networks, characterizing social media, among others. G PM algorithms are challenging to develop due to subroutines that include non-trivial graph theory concepts and methods such as isomorphism. General-purpose GPM systems have emerged as a solution to improve the user experience with such algorithms. However, existing general-purpose GPM systems are heterogeneous in terms of implementation details, hardware environment and algorithmic paradigms for sub graph exploration and thus, observations taken from the experimental results alone may not clearly identify when a particular paradigm prevails over another. In this work we present an experimentation analysis of popular paradigms used in existing GPM systems. In order to provide a fair and comprehensive evaluation of various algorithmic paradigms we implement all of them within a single GPM framework. Our results show that no single paradigm is best for every application scenario, and we believe that our findings may guide practitioner towards more optimized GPM systems in the future.
Vinícius Vitor dos Santos Dias, Samuel Ferraz, Aditya Vadlamani, Mahdi Erfanian, Carlos H. C. Teixeira, Dorgival O. Guedes, Wagner Meira Jr., Srinivasan Parthasarathy 0001
HiPC7
2023 Algorithmic Recourse in Mental Healthcare
abstract
This paper explores using algorithmic recourse as a tool in mental healthcare. Algorithmic recourse provides explanations and recommendations to individuals who want to reverse a machine learning prediction and has been widely used in various domains such as finance and marketing. However, its application in mental healthcare has been restricted. This paper addresses this by examining the potential benefits and challenges of using algorithmic recourse in mental healthcare, specifically in how changing one's behavior may affect their quality of life and well-being. The paper proposes a new classification-based framework for algorithmic recourse in mental healthcare. The proposed framework considers both observed and latent variables to account for the individuality of individuals and provides a more comprehensive understanding of mental health outcomes. The research results can provide valuable insights for future work and help bridge the gap between machine learning and mental healthcare.
Anísio Lacerda, Claudio Almeida, Leonardo Augusto Ferreira, Adriano C. M. Pereira, Gisele L. Pappa, Wagner Meira Jr., Débora M. Miranda, Marco Aurélio Romano-Silva, Leandro Malloy Diniz
IJCNN6
2022 Efficient Strategies for Graph Pattern Mining Algorithms on GPUs
abstract
Graph Pattern Mining (GPM) is an important, rapidly evolving, and computation demanding area. GPM computation relies on subgraph enumeration, which consists in extracting subgraphs that match a given property from an input graph. Graphics Processing Units (GPUs) have been an effective platform to accelerate applications in many areas. However, the irregularity of subgraph enumeration makes it challenging for efficient execution on GPU due to typical uncoalesced memory access, divergence, and load imbalance. Unfortunately, these aspects have not been fully addressed in previous work. Thus, this work proposes novel strategies to design and implement subgraph enumeration efficiently on GPU. We support a depth-first search style search (DFS-wide) that maximizes memory performance while providing enough parallelism to be exploited by the GPU, along with a warp-centric design that minimizes execution divergence and improves utilization of the computing capabilities. We also propose a low-cost load balancing layer to avoid idleness and redistribute work among thread warps in a GPU. Our strategies have been deployed in a system named DuMato, which provides a simple programming interface to allow efficient implementation of GPM algorithms. Our evaluation has shown that DuMato is often an order of magnitude faster than state-of-the-art GPM systems and can mine larger subgraphs (up to 12 vertices).
Samuel Ferraz, Vinícius Vitor dos Santos Dias, Carlos H. C. Teixeira, George Teodoro, Wagner Meira Jr.
SBAC-PAD5
2022 Counterfactual inference with latent variable and its application in mental health care
Guilherme F. Marchezini, Anísio Lacerda, Gisele L. Pappa, Wagner Meira Jr., Débora M. Miranda, Marco Aurélio Romano-Silva, Danielle S. Costa, Leandro Malloy Diniz
Data Min. Knowl. Discov.4
2022 Sequential stratified regeneration: MCMC for large state spaces with an application to subgraph count estimation
Carlos H. C. Teixeira, Mayank Kakodkar, Vinícius Vitor dos Santos Dias, Wagner Meira Jr., Bruno Ribeiro 0001
Data Min. Knowl. Discov.4
2021 Analyzing topic attention in online small groups
abstract
Attention is a scarce resource disputed by algorithms and people on the Internet. This competition for attention is part of online spaces especially online small groups where there is a limited number of individuals interacting with each other using text and media content that is not controlled by algorithms or human curators. In these groups, as certain participants and piece of content can catch the collective attention, a question that naturally arises is: how to analyze topic attention in online small groups? In this paper, we propose a methodology aimed at answering this question. Our proposal consists of sets of analyses over topical (obtained from topic analysis) transition graphs for characterizing attention allocation, permanence and shifting as well as participant role characterization during discussions in online small groups. We experimented with our methodology using WhatsApp groups as a case study. Among other results, we identified and characterized abrupt and smooth topic transitions as well as patterns of participant activity related to certain topics.
Josemar Alves Caetano, Jussara M. Almeida, Marcos André Gonçalves, Wagner Meira Jr., Humberto Torres Marques-Neto, Virgílio A. F. Almeida
ASONAM4
2021 Towards automatic diagnosis of rheumatic heart disease on echocardiographic exams through video-based deep learning
abstract
OBJECTIVE: Rheumatic heart disease (RHD) affects an estimated 39 million people worldwide and is the most common acquired heart disease in children and young adults. Echocardiograms are the gold standard for diagnosis of RHD, but there is a shortage of skilled experts to allow widespread screenings for early detection and prevention of the disease progress. We propose an automated RHD diagnosis system that can help bridge this gap. MATERIALS AND METHODS: Experiments were conducted on a dataset with 11 646 echocardiography videos from 912 exams, obtained during screenings in underdeveloped areas of Brazil and Uganda. We address the challenges of RHD identification with a 3D convolutional neural network (C3D), comparing its performance with a 2D convolutional neural network (VGG16) that is commonly used in the echocardiogram literature. We also propose a supervised aggregation technique to combine video predictions into a single exam diagnosis. RESULTS: The proposed approach obtained an accuracy of 72.77% for exam diagnosis. The results for the C3D were significantly better than the ones obtained by the VGG16 network for videos, showing the importance of considering the temporal information during the diagnostic. The proposed aggregation model showed significantly better accuracy than the majority voting strategy and also appears to be capable of capturing underlying biases in the neural network output distribution, balancing them for a more correct diagnosis. CONCLUSION: Automatic diagnosis of echo-detected RHD is feasible and, with further research, has the potential to reduce the workload of experts, enabling the implementation of more widespread screening programs worldwide.
Joao Francisco B. S. Martins, Erickson R. Nascimento, Bruno Ramos Nascimento, Craig A. Sable, Andrea Z. Beaton, Antônio L. P. Ribeiro, Wagner Meira Jr., Gisele L. Pappa
J. Am. Medical Informatics Assoc.7
2021 Identifying Networks Vulnerable to IP Spoofing
abstract
The lack of authentication in the Internet's data plane allows hosts to falsify (spoof) the source IP address in packet headers. IP source spoofing is the basis for amplification denial-of-service (DoS) attacks. Current approaches to locate sources of spoofed traffic lack coverage or are not deployable today. We propose a mechanism that a network with multiple peering links can use to coarsely locate the sources of spoofed traffic in the Internet. The idea behind our approach is that a network can monitor and map spoofed traffic arriving on a peering link to the set of sources routed toward that link. We propose mechanisms the network can use to systematically vary BGP announcement configurations to induce changes to Internet routes and to the set of sources routed to each peering link. A network using our technique can correlate observations over multiple configurations to more precisely delineate regions sending spoofed traffic. Evaluation of our techniques on the Internet shows that they can partition the Internet into small regions, allowing targeted intervention.
Osvaldo L. H. M. Fonseca, Ítalo S. Cunha, Elverton C. Fazzion, Wagner Meira Jr., Brivaldo Alves da Silva, Ronaldo A. Ferreira, Ethan Katz-Bassett
IEEE Trans. Netw. Serv. Manag.4
2020 Building The First English-Brazilian Portuguese Corpus for Automatic Post-Editing
abstract
This paper introduces the first corpus for Automatic Post-Editing of English and a low-resource language, Brazilian Portuguese. The source English texts were extracted from the WebNLG corpus and automatically translated into Portuguese using a state-of-the-art industrial neural machine translator. Post-edits were then obtained in an experiment with native speakers of Brazilian Portuguese. To assess the quality of the corpus, we performed error analysis and computed complexity indicators measuring how difficult the APE task would be. We report preliminary results of Phrase-Based and Neural Machine Translation Models on this new corpus. Data and code publicly available in our repository.
Felipe Almeida Costa, Thiago Castro Ferreira, Adriana S. Pagano, Wagner Meira Jr.
COLING4
2020 Tracking Down Sources of Spoofed IP Packets
Osvaldo L. H. M. Fonseca, Ítalo S. Cunha, Elverton C. Fazzion, Wagner Meira Jr., Brivaldo Junior, Ronaldo A. Ferreira, Ethan Katz-Bassett
Networking4
2020 Federated and secure cloud services for building medical image classifiers on an intercontinental infrastructure
Ignacio Blanquer, Francisco Vilar Brasileiro, Andrey Brito, Amanda Calatrava, Christof Fetzer, Flavio Figueiredo, Ronny Petterson Guimarães, Leandro Bezerra Marinho, Wagner Meira Jr., Altigran S. da Silva, Angel Alberich-Bayarri, Eduardo Camacho-Ramos, Ana Jimenez-Pastor, Antonio Luiz L. Ribeiro, Bruno Ramos Nascimento
Future Gener. Comput. Syst.10
2019 A Framework for Benchmarking Discrimination-Aware Models in Machine Learning
abstract
Discrimination-aware models in machine learning are a recent topic of study that aim to minimize the adverse impact of machine learning decisions for certain groups of people due to ethical and legal implications. We propose a benchmark framework for assessing discrimination-aware models. Our framework consists of systematically generated biased datasets that are similar to real world data, created by a Bayesian network approach. Experimental results show that we can assess the quality of techniques through known metrics of discrimination, and our flexible framework can be extended to most real datasets and fairness measures to support a diversity of assessments.
Rodrigo L. Cardoso, Wagner Meira Jr., Virgílio A. F. Almeida, Mohammed J. Zaki
AIES2
2019 Medical Imaging Processing Architecture on ATMOSPHERE Federated Platform
abstract
[EN] This paper describes the development of applications in the frame of the ATMOSPHERE platform. ATMOSPHERE provides means for developing container-based applications over a federated cloud offering measurin he trustworthiness of the applications. In this paper we show the design of a transcontinental application in the frame of medical imaging that keeps the data at one end and uses the processing capabilities of the resources available at the other end. The applications are described using TOSCA blueprints and the federation of IaaS resources is performed by the Fogbow middleware. Privacy guarantees are provided by means of SCONE and intensive computing resources are integrated through the use of GPUs directly mounted on the containers.
Ignacio Blanquer, Angel Alberich-Bayarri, Fabio García-Castro, George Teodoro, André Meirelles, Bruno Nascimento, Wagner Meira Jr., Antônio L. P. Ribeiro
CLOSER7
2019 Detecting Spatial Clusters of Disease Infection Risk Using Sparsely Sampled Social Media Mobility Patterns
abstract
Standard spatial cluster detection methods used in public health surveillance assign each disease case to a single location (typically, the patient's home address), aggregate locations to small areas, and monitor the number of cases in each area over time. However, such methods cannot detect clusters of disease resulting from visits to non-residential locations, such as a park or a university campus. Thus we develop two new spatial scan methods, the unconditional and conditional spatial logistic models, to search for spatial clusters of increased infection risk. We use mobility data from two sets of individuals, disease cases and healthy individuals, where each individual is represented by a sparse sample of geographical locations (e.g., from geo-tagged social media data). The methods account for the multiple, varying number of spatial locations observed per individual, either by non-parametric estimation of the odds of being a case, or by matching case and control individuals with similar numbers of observed locations. Applying our methods to synthetic and real-world scenarios, we demonstrate robust performance on detecting spatial clusters of infection risk from mobility data, outperforming competing baselines.
Roberto C. S. N. P. Souza, Renato Assunção, Daniel B. Neill, Wagner Meira Jr.
SIGSPATIAL/GIS4
2019 Identifying and Characterizing Bashlite and Mirai C&C Servers
abstract
IoT devices are often a vector for assembling massive botnets, as a consequence of being broadly available, having limited security protections, and significant challenges in deploying software upgrades. Such botnets are usually controlled by centralized Command-and-Control (C&C) servers, which need to be identified and taken down to mitigate threats. In this paper we propose a framework to infer C&C server IP addresses using four heuristics. Our heuristics employ static and dynamic analysis to automatically extract information from malware binaries. We use active measurements to validate inferences, and demonstrate the efficacy of our framework by identifying and characterizing C&C servers for 62% of 1050 malware binaries collected using 47 honeypots.
Gabriel Bastos, Wagner Meira Jr., Artur Marzano, Osvaldo L. H. M. Fonseca, Elverton C. Fazzion, Cristine Hoepers, Klaus Steding-Jessen, Marcelo H. P. Chaves, Ítalo S. Cunha, Dorgival O. Guedes
ISCC2
2019 Fractal: A General-Purpose Graph Pattern Mining System
abstract
In this paper we propose Fractal, a high performance and high productivity system for supporting distributed graph pattern mining (GPM) applications. Fractal employs a dynamic (auto-tuned) load-balancing based on a hierarchical and locality-aware work stealing mechanism, allowing the system to adapt to different workload characteristics. Additionally, Fractal enumerates subgraphs by combining a depth-first strategy with a from scratch processing paradigm to avoid storing large amounts of intermediate state and, thus, improves memory efficiency. Regarding programmer productivity, Fractal presents an intuitive, expressive and modular API, allowing for rapid compositional expression of many GPM algorithms. Fractal-based implementations outperform both existing systemic solutions and specialized distributed solutions on many problems - from frequent graph mining to subgraph querying, over a range of datasets.
Vinícius Vitor dos Santos Dias, Carlos H. C. Teixeira, Dorgival O. Guedes, Wagner Meira Jr., Srinivasan Parthasarathy 0001
SIGMOD Conference4
2019 BIGSEA: A Big Data analytics platform for public transportation information
Andy S. Alic, Jussara M. Almeida, Giovanni Aloisio, Nazareno Andrade, Nuno Antunes, Danilo Ardagna, Rosa M. Badia, Tânia Basso, Ignacio Blanquer, Tarciso Braz, Andrey Brito, Donatello Elia, Sandro Fiore, Dorgival O. Guedes, Marco Lattuada 0001, Daniele Lezzi, Matheus Maciel, Wagner Meira Jr., Demetrio Gomes Mestre, Regina Lúcia de Oliveira Moraes, Fábio Morais 0001, Carlos Eduardo S. Pires, Nádia P. Kozievitch, Walter Santos, Paulo Silva 0002, Marco Vieira
Future Gener. Comput. Syst.18
2018 Graph Pattern Mining and Learning through User-Defined Relations
abstract
In this work we propose R-GPM, a parallel computing framework for graph pattern mining (GPM) through a user-defined subgraph relation. More specifically, we enable the computation of statistics of patterns through their subgraph classes, generalizing traditional GPM methods. R-GPM provides efficient estimators for these statistics by employing a MCMC sampling algorithm combined with several optimizations. We provide both theoretical guarantees and empirical evaluations of our estimators in application scenarios such as stochastic optimization of deep high-order graph neural network models and pattern (motif) counting. We also propose and evaluate optimizations that enable improvements of our estimators accuracy, while reducing their computational costs in up to 3-orders-of-magnitude. Finally, we show that R-GPM is scalable, providing near-linear speedups.
Carlos H. C. Teixeira, Leornado Cotta, Bruno Ribeiro 0001, Wagner Meira Jr.
ICDM4
2018 Online Orthogonal Regression Based on a Regularized Squared Loss
abstract
Orthogonal regression extends the classical regression framework by assuming that the data may contain errors in both the dependent and independent variables. Often, this approach tends to outperform classical regression in real-world scenarios. However, the algorithms used to determine a solution to the orthogonal regression problem require the computation of singular value decompositions (SVD), which may be computationally expensive and impractical for real-world problems. In this work, we propose a new approach to the orthogonal regression problem based on a regularized squared loss. The method follows an online learning strategy which makes it more flexible for different types of applications. The algorithm is derived in primal and dual variables and the later formulation allows the introduction of kernels for nonlinear modeling. We compare our proposed orthogonal regression algorithm to a corresponding classical regression algorithm using both synthetic and real-world datasets from different applications. Our algorithm achieved better results for most of the datasets.
Roberto Souza 0001, Saul de Castro Leite, Wagner Meira Jr., Eduardo R. Hruschka
ICMLA3
2018 Characterizing and Detecting Hateful Users on Twitter
Manoel Horta Ribeiro, Pedro H. Calais, Yuri A. Santos, Virgílio A. F. Almeida, Wagner Meira Jr.
ICWSM5
2018 The Evolution of Bashlite and Mirai IoT Botnets
abstract
Vulnerable IoT devices are powerful platforms for building botnets that cause billion-dollar losses every year. In this work, we study Bashlite botnets and their successors, Mirai botnets. In particular, we focus on the evolution of the malware as well as changes in botnet operator behavior. We use monitoring logs from 47 honeypots collected over 11 months. Our results shed new light on those botnets, and complement previous findings by providing evidence that malware, botnet operators, and malicious activity are becoming more sophisticated. Compared to its predecessor, we find Mirai uses more resilient hosting and control infrastructures, and supports more effective attacks.
Artur Marzano, David Alexander, Osvaldo L. H. M. Fonseca, Elverton C. Fazzion, Cristine Hoepers, Klaus Steding-Jessen, Marcelo H. P. Chaves, Ítalo S. Cunha, Dorgival O. Guedes, Wagner Meira Jr.
ISCC10
2018 An Unsupervised Boosting Strategy for Outlier Detection Ensembles
Guilherme Oliveira Campos, Arthur Zimek, Wagner Meira Jr.
PAKDD (1)3
2018 NetClass: A network-based relational model for document classification
Fernando Mourão, Leonardo Rocha 0001, Felipe Viegas, Thiago Salles, Marcos André Gonçalves, Srinivasan Parthasarathy 0001, Wagner Meira Jr.
Inf. Sci.7
2018 Janus: Diagnostics and reconfiguration of data parallel programs
Vinícius Vitor dos Santos Dias, Wagner Meira Jr., Dorgival O. Guedes
J. Parallel Distributed Comput.2
2018 Scalable and Efficient Data Analytics and Mining with Lemonade
abstract
Professionals outside of the area of Computer Science have an increasing need to analyze large bodies of data. This analysis often demands high level of security and has to be done in the cloud. However, current data analysis tools that demand little proficiency in systems programming struggle to deliver solutions which are scalable and safe. In this context we present Lemonade, a platform which focuses on creating data analysis and mining flows in the cloud, with authentication, authorization and accounting (AAA) guarantees. Lemonade provides an interface for the visual construction of flows, and encapsulates storage and data processing environment details, providing higher-level abstractions for data source access and algorithms. We illustrate its usage through a demo, where a data processing flow builds a classification model for detecting fake-news, also extracting some insights along the way.
Walter Santos, Gustavo de P. Avelar, Manoel Horta Ribeiro, Dorgival O. Guedes, Wagner Meira Jr.
Proc. VLDB Endow.5
2017 PRIVAaaS: privacy approach for a distributed cloud-based data analytics platforms
abstract
Data privacy is a key challenge that is exacerbated by Big Data storage and analytics processing requirements. Big Data and Cloud Computing are related and allow the users to access data from any device, making data privacy essential as the data sets are exposed through the web. Organizations care about data privacy as it directly affects the confidence that clients have that their personal data are safe. This paper presents a data privacy approach - PRIVAaaS - and its inte-gration to the LEMONADE Web-based platform, developed to compose ETL (Extract, Transform, Load) process and Machine Learning workflows. The 3-level approach of PRIVAaaS, based on data anonymization policies, is implemented in a software toolkit that provides a set of libraries and tools which allows controlling and reducing data leakage in the context of Big Data processing.
Tânia Basso, Regina Lúcia de Oliveira Moraes, Nuno Antunes, Marco Vieira, Walter Santos, Wagner Meira Jr.
CCGrid6
2017 Lemonade: A scalable and efficient Spark-based platform for data analytics
abstract
Data Analytics is a concept related to pattern and relevant knowledge discovery from large amounts of data. In general, the task is complex and demands knowledge in very specific areas, such as massive data processing and parallel programming languages. However, analysts are usually not versed in Computer Science, but in the original data domain. In order to support them in such analysis, we present Lemonade — Live Exploration and Mining Of a Non-trivial Amount of Data from Everywhere — a platform for visual creation and execution of data analysis workflows. Lemonade encapsulates storage and data processing environment details, providing higher-level abstractions for data source access and algorithms coding. The goal is to enable batch and interactive execution of data analysis tasks, from basic ETL to complex data mining algorithms, in parallel, in a distributed environment. The current version supports HDFS (the Hadoop filesystem), local filesystems and distributed environments such as Apache Spark, the state-of-art framework for Big Data analysis.
Walter Santos, Luiz F. M. Carvalho, Gustavo de P. Avelar, Átila Silva Jr., Lucas M. Ponce, Dorgival O. Guedes, Wagner Meira Jr.
CCGrid7
2017 Antagonism Also Flows Through Retweets: The Impact of Out-of-Context Quotes in Opinion Polarization Analysis
Pedro Henrique Calais Guerra, Roberto Nalon, Renato Assunção, Wagner Meira Jr.
ICWSM4
2017 A Characterization of Load Balancing on the IPv6 Internet
Rafael Almeida, Osvaldo L. H. M. Fonseca, Elverton C. Fazzion, Dorgival O. Guedes, Wagner Meira Jr., Ítalo S. Cunha
PAM5
2017 What surprises does your past have for you?
Fernando Mourão, Leonardo Rocha 0001, Camila Souza Araujo, Wagner Meira Jr., Joseph A. Konstan
Inf. Syst.4
2017 A Two-Stage Machine learning approach for temporally-robust text classification
Thiago Salles, Leonardo Rocha 0001, Fernando Mourão, Marcos André Gonçalves, Felipe Viegas, Wagner Meira Jr.
Inf. Syst.6
2016 Diagnosing Performance Bottlenecks in Massive Data Parallel Programs
abstract
The increasing amount of data being stored and the variety of applications being proposed recently to make use of those data enabled a whole new generation of parallel programming environments and paradigms. Although most of these novel environments provide abstract programming interfaces and embed several run-time strategies that simplify several typical tasks in parallel and distributed systems, achieving good performance is still a challenge. In this paper we identify some common sources of performance degradation in the Spark programming environment and discuss some diagnosis dimensions that can be used to better understand such degradation. We then describe our experience in the use of those dimensions to drive the identification performance problems, and suggest how their impact may be minimized considering real applications.
Vinícius Vitor dos Santos Dias, Ruens Moreira, Wagner Meira Jr., Dorgival O. Guedes
CCGrid3
2016 Faster: A Low Overhead Framework for Massive Data Analysis
abstract
With the recent accelerated increase in the amount of social data available in the Internet, several big data distributed processing frameworks have been proposed and implemented. Hadoop has been used widely to process all kinds of data, not only from social media. Spark is gaining popularity for offering a more flexible, object-functional, programming interface, and also by improving performance in many cases. However, not all data analysis algorithms perform well on Hadoop or Spark. For instance, graph algorithms tend to generate large amounts of messages between processing elements, which may result in poor performance even in Spark. We introduce Faster, a low latency distributed processing framework, designed to explore data locality to reduce processing costs in such algorithms. It offers an API similar to Spark, but with a slightly different execution model and new operators. Our results show that it can significantly outperform Spark on large graphs, being up to one orders of magnitude faster when running PageRank in a partial Google+ friendship graph with more than one billion edges.
Matheus Santos, Wagner Meira Jr., Dorgival O. Guedes, Virgílio A. F. Almeida
CCGrid2
2016 An architecture integrating stationary and mobile roadside units for providing communication on intelligent transportation systems
abstract
In this work we investigate the benefits of an hybrid architecture integrating both mobile roadside units, and stationary roadside units supporting the operation of vehicular networks. Since traffic fluctuates, an architecture employing just stationary roadside units might not be able to properly support the network operation all the time. Similarly, an architecture composed just of mobile roadside units may lack part of the robustness provided by stationary roadside units. Furthermore, the traffic fluctuations are limited by the underlying road network, and the road networks do not change so often as traffic does. Thus, it seems straight full to assume that a set of roadside units will always be left stationary, while other roadside units will need to roam along the road network. As major roads counts on a higher transportation capacity, they tend to be very popular routes, and they are natural candidates for receiving the stationary roadside units. On the other hand, we may rely on mobile roadside units for handling roads presenting a high traffic variation. In this work we use the realistic vehicular mobility trace of Cologne, Germany, and we model the allocation of the roadside units as a Maximum Coverage Problem. Our results demonstrate the hybrid deployment increases the number of covered vehicles up to 45%.
Cristiano M. Silva, Wagner Meira Jr.
NOMS2
2016 Infection Hot Spot Mining from Social Media Trajectories
Roberto C. S. N. P. Souza, Renato Assunção, Derick M. de Oliveira, Denise E. F. de Brito, Wagner Meira Jr.
ECML/PKDD (2)5
2016 Dynamic Reconfiguration of Data Parallel Programs
abstract
Given the large amount of data from different sources that have become available to researchers in multiple fields, Data Science has emerged as a new paradigm for exploring and getting value from that data. In that context, new parallel processing environments with abstract programming interfaces, like Spark, were proposed to try to simplify the development of distributed programs. Although such solutions have become widely used, achieving the best performance with them is still not always straight-forward, despite the multiple run-time strategies they use. In this work we analyze some of the causes of performance degradation in such systems and, based on that analysis, we propose a tool to improve performance by dynamically adjusting data partitioning and parallelism degree in recurrent applications based on previous executions. Our results applying that methodology show consistent reductions in execution time for the applications considered, with gains of up to 50%.
Vinícius Vitor dos Santos Dias, Wagner Meira Jr., Dorgival O. Guedes
SBAC-PAD2
2016 Efficient Remapping of Internet Routing Events
abstract
Routing events impact multiple paths in the Internet, but current active topology mapping techniques monitor paths independently. Detecting a routing event on one Internet path does not trigger any measurements on other possibly-impacted paths. This approach leads to outdated and inconsistent routing information. We characterize routing events in the Internet and investigate probing strategies to efficiently identify paths impacted by a routing event. Our results indicate that targeted probing can help us quickly remap routing events and maintain more up-to-date and consistent topology maps.
Elverton C. Fazzion, Ítalo S. Cunha, Dorgival O. Guedes, Wagner Meira Jr., Renata Teixeira, Darryl Veitch, Christophe Diot
SIGCOMM4
2016 Watershed-ng: an extensible distributed stream processing framework
abstract
Summary Most high‐performance data processing (a.k.a. big data) systems allow users to express their computation using abstractions (like MapReduce), which simplify the extraction of parallelism from applications. Most frameworks, however, do not allow users to specify how communication must take place: That element is deeply embedded into the run‐time system abstractions, making changes hard to implement. In this work, we describe Wathershed‐ng, our re‐engineering of the Watershed system, a framework based on the filter–stream paradigm and originally focused on continuous stream processing. Like other big‐data environments, Watershed provided object‐oriented abstractions to express computation (filters), but the implementation of streams was a run‐time system element. By isolating stream functionality into appropriate classes, combination of communication patterns and reuse of common message handling functions (like compression and blocking) become possible. The new architecture even allows the design of new communication patterns, for example, allowing users to choose MPI, TCP, or shared memory implementations of communication channels as their problem demands. Applications designed for the new interface showed reductions in code size on the order of 50%and above in some cases. The performance results also showed significant improvements, because some implementation bottlenecks were removed in the re‐engineering process. Copyright © 2016 John Wiley & Sons, Ltd.
Rodrigo Caetano Rocha, Bruno Hott, Vinícius Vitor dos Santos Dias, Renato Ferreira 0001, Wagner Meira Jr., Dorgival O. Guedes
Concurr. Comput. Pract. Exp.5
2016 Exploring multiple evidence to infer users' location in Twitter
Erica C. Rodrigues, Renato Assunção, Gisele L. Pappa, Diogo Rennó, Wagner Meira Jr.
Neurocomputing5
2016 A quantitative analysis of the temporal effects on automatic text classification
abstract
Automatic text classification (TC) continues to be a relevant research topic and several TC algorithms have been proposed. However, the majority of TC algorithms assume that the underlying data distribution does not change over time. In this work, we are concerned with the challenges imposed by the temporal dynamics observed in textual data sets. We provide evidence of the existence of temporal effects in three textual data sets, reflected by variations observed over time in the class distribution, in the pairwise class similarities, and in the relationships between terms and classes. We then quantify, using a series of full factorial design experiments, the impact of these effects on four well‐known TC algorithms. We show that these temporal effects affect each analyzed data set differently and that they restrict the performance of each considered TC algorithm to different extents. The reported quantitative analyses, which are the original contributions of this article, provide valuable new insights to better understand the behavior of TC algorithms when faced with nonstatic (temporal) data distributions and highlight important requirements for the proposal of more accurate classification models.
Thiago Salles, Leonardo Rocha 0001, Marcos André Gonçalves, Jussara M. Almeida, Fernando Mourão, Wagner Meira Jr., Felipe Viegas
J. Assoc. Inf. Sci. Technol.6
2016 Isofunctional Protein Subfamily Detection Using Data Integration and Spectral Clustering
abstract
As increasingly more genomes are sequenced, the vast majority of proteins may only be annotated computationally, given experimental investigation is extremely costly. This highlights the need for computational methods to determine protein functions quickly and reliably. We believe dividing a protein family into subtypes which share specific functions uncommon to the whole family reduces the function annotation problem's complexity. Hence, this work's purpose is to detect isofunctional subfamilies inside a family of unknown function, while identifying differentiating residues. Similarity between protein pairs according to various properties is interpreted as functional similarity evidence. Data are integrated using genetic programming and provided to a spectral clustering algorithm, which creates clusters of similar proteins. The proposed framework was applied to well-known protein families and to a family of unknown function, then compared to ASMC. Results showed our fully automated technique obtained better clusters than ASMC for two families, besides equivalent results for other two, including one whose clusters were manually defined. Clusters produced by our framework showed great correspondence with the known subfamilies, besides being more contrasting than those produced by ASMC. Additionally, for the families whose specificity determining positions are known, such residues were among those our technique considered most important to differentiate a given group. When run with the crotonase and enolase SFLD superfamilies, the results showed great agreement with this gold-standard. Best results consistently involved multiple data types, thus confirming our hypothesis that similarities according to different knowledge domains may be used as functional similarity evidence. Our main contributions are the proposed strategy for selecting and integrating data types, along with the ability to work with noisy and incomplete data; domain knowledge usage for detecting subfamilies in a family with different specificities, thus reducing the complexity of the experimental function characterization problem; and the identification of residues responsible for specificity.
Elisa Boari de Lima, Wagner Meira Jr., Raquel Cardoso de Melo Minardi
PLoS Comput. Biol.2
2016 Non-Intrusive Planning the Roadside Infrastructure for Vehicular Networks
abstract
In this article, we describe a strategy for planning the roadside infrastructure for vehicular networks based on the global behavior of drivers. Instead of relying on the trajectories of all vehicles, our proposal relies on the migration ratios of vehicles between urban regions in order to infer the better locations for deploying the roadside units. By relying on the global behavior of drivers, our strategy does not incur in privacy concerns. Given a set of α available roadside units, our goal is to select those α-better locations for placing the roadside units in order to maximize the number of distinct vehicles experiencing at least one V2I contact opportunity. Our results demonstrate that full knowledge of the vehicle trajectories are not mandatory for achieving a close-to-optimal deployment performance when we intend to maximize the number of distinct vehicles experiencing (at least one) V2I contact opportunities.
Cristiano M. Silva, Wagner Meira Jr., João F. M. Sarubbi
IEEE Trans. Intell. Transp. Syst.2
2015 Design of roadside communication infrastructure with QoS guarantees
abstract
There are several kinds of envisioned vehicular applications: video delivery, accidents detection, dissemination of traffic announcements, and so forth. Such applications demand minimal (and possibly distinct) QoS guarantees that must couple the vehicular network. Given that vehicular networks will soon become reality, we demand strategies for planning and managing such networks. In this work we propose Delta (A), a QoS-based strategy for planning the roadside infrastructure supporting a vehicular network. Thus, the network provider may employ our strategy to design a new network, compare the performance of distinct vehicular networks, and even evaluate the adherence between vehicular applications and the network. Delta is based on two metrics: i) connectivity duration, and ii) percentage of vehicles presenting such connectivity duration. For instance, if a given vehicular application requires that 20% of the vehicles are connected during 30% of the trip, we say that such application requires a deployment Delta (0.3, 0.2). Complementary, we also present Delta-g, a greedy heuristic for solving Delta. A deployment Delta (0.1, 0.1) requires the coverage of 0.09% of the road network, while a deployment Delta (0.9, 0.9) requires 21.67% of coverage.
Cristiano M. Silva, Wagner Meira Jr.
ISCC2
2015 Managing Infrastructure-Based Vehicular Networks
abstract
In this thesis work the authors exploit the management of infrastructure-based vehicular networks. The authors begin to work by investigating the most basic problem faced by the network designers when planning an infrastructure-based vehicular network: given a road network, a flow, and α available RSUs, where the RSUs must be located in order to maximize the network performance? The authors propose a novel approach for locating the RSUs: the authors develop a deployment algorithm based on partial mobility information. By partial mobility information, we mean: i) density of vehicles along the road network; and, ii) migration ratios from distinct locations of the road network.
Cristiano M. Silva, Wagner Meira Jr.
MDM (2)2
2015 Evaluating the Performance of Heterogeneous Vehicular Networks
abstract
There are several kinds of envisioned vehicular applications: video delivery, accidents detection, dissemination of traffic announcements, and so forth. Such applications demand minimal (and possibly distinct) Quality of Service guarantees that must couple the vehicular network. Since the vehicular networks will become reality soon, we demand strategies for planning and managing such networks. In this work we propose the concept of a Delta Network: the Delta Network is a metric for evaluating the performance of heterogeneous vehicular networks. By using the concept of a Network Delta, we expect network providers being able to measure and compare the performance of distinct heterogeneous vehicular networks (Delta is technology-independent). Finally, the concept of a Delta Network may also be applied to couple vehicular applications and vehicular networks in order to support the network provider in the decision of deploying (or not) a novel vehicular application.
Cristiano M. Silva, Wagner Meira Jr.
VTC Fall2
2015 Planning the communication infrastructure for vehicular networks without tracking vehicles
abstract
This work presents a novel algorithm for the deployment of roadside units based on partial mobility information. Instead of relying on the individual vehicles trajectories, our proposal relies on the migration ratios between urban regions in order to infer the better locations for the deployment of the roadside units. Our goal is to identify those α locations maximizing the number of distinct vehicles experiencing at least one vehicle-to-infrastructure contact opportunity. We compare our strategy to two deployment algorithms: MCP-g relies on full mobility information (i.e., full knowledge of the individual vehicles trajectories), while MCP-kp does not assume any mobility information at all (simply places the roadside units at the densest locations of the road network). The results demonstrate that full knowledge of the vehicles trajectories is not mandatory for achieving a close-to-optimal deployment performance: by covering 1.0% of the Cologne's road network, our approach yields 89.8% of all vehicles experiencing at least one vehicle-to-infrastructure contact opportunity. If we assume previous knowledge of individual vehicles trajectories, we improve the coverage performance in just 2.3%. Complementary, our strategy provides a deployment layout very similar to the one obtained when considering the individual vehicles trajectories.
Cristiano M. Silva, João F. M. Sarubbi, Wagner Meira Jr.
WiMob3
2015 PDBest: a user-friendly platform for manipulating and enhancing protein structures
abstract
UNLABELLED: PDBest (PDB Enhanced Structures Toolkit) is a user-friendly, freely available platform for acquiring, manipulating and normalizing protein structures in a high-throughput and seamless fashion. With an intuitive graphical interface it allows users with no programming background to download and manipulate their files. The platform also exports protocols, enabling users to easily share PDB searching and filtering criteria, enhancing analysis reproducibility. AVAILABILITY AND IMPLEMENTATION: PDBest installation packages are freely available for several platforms at http://www.pdbest.dcc.ufmg.br CONTACT: [email protected], [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wellisson R. S. Gonçalves, Valdete M. Gonçalves-Almeida, Aleksander L. Arruda, Wagner Meira Jr., Carlos H. Silveira, Douglas E. V. Pires, Raquel Cardoso de Melo Minardi
Bioinform.4
2015 Deployment of roadside units based on partial mobility information
Cristiano M. Silva, André L. L. de Aquino, Wagner Meira Jr.
Comput. Commun.3
2015 Learning sequential classifiers from long and noisy discrete-event sequences efficiently
Gessé Dafé, Adriano Veloso, Mohammed J. Zaki, Wagner Meira Jr.
Data Min. Knowl. Discov.4
2015 Smart Traffic Light for Low Traffic Conditions - A Solution for Improving the Drivers Safety
Cristiano M. Silva, André L. L. de Aquino, Wagner Meira Jr.
Mob. Networks Appl.3
2014 Reachability Queries in Very Large Graphs: A Fast Refined Online Search Approach
abstract
A key problem in many graph-based applications is the need to know, given a directed graph G and two vertices u,v ∈ G, whether there is a path between u and v, i.e., if u reaches v. This problem is particularly challenging in the case of very large real-world graphs. A common approach is the preprocessing of the graphs, in order to produce an efficient index structure, which allows fast access to the reachability information of the vertices. However, the majority of existing methods can not handle very large graphs. We propose, in this paper, a novel indexing method called FELINE (Fast rEfined onLINE search), which is inspired by Dominance Graph Drawing. FELINE creates an index from the graph representation in a two-dimensional plane, which provides reachability information in constant time for a significant portion of queries. Experiments demonstrate the efficiency of FELINE compared to state-of-the-art approaches.
Renê Rodrigues Veloso, Loïc Cerf, Wagner Meira Jr., Mohammed J. Zaki
EDBT3
2014 Complete discovery of high-quality patterns in large numerical tensors
abstract
Many datasets are numerical tensors, i. e., associate n-tuples with numerical values. Until recently, the discovery of relevant local patterns in such numerical and multidimensional data has received little attention despite the broad applicative perspectives offered by this general framework. Even in the simpler 2-dimensional case, almost every proposal so far is either incomplete (i. e., it does not list every pattern) or relies on binning and mines Boolean tensors. In both cases, some information is lost during the process. In uncertain tensors, n-tuples satisfy the studied predicate to a certain extent and no information is lost w.r.t. the original data. Given an uncertain tensor, the closed patterns are its maximal “sub-tensors” covering n-tuples that “mostly” satisfy the predicate. Defining “mostly” is the key problem: the patterns should be both relevant given the data and efficiently extractable. The proposed complete extractor reuses the enumeration principles of the state-of-the-art miner for closed n-sets but incrementally enforces the newly designed definition. In this way, the proposed algorithm runs orders of magnitude faster than its only competitor and large datasets are tractable. The experimental section reports the discovery of dynamic patterns of influence in Twitter as well as usage patterns in a transportation network. Additional experiments on synthetic data quantitatively assess the quality of the chosen definition for the patterns.
Loïc Cerf, Wagner Meira Jr.
ICDE2
2014 Of Pins and Tweets: Investigating How Users Behave Across Image- and Text-Based Social Networks
Raphael Ottoni, Diego B. Las Casas, João Paulo Pesce, Wagner Meira Jr., Christo Wilson, Alan Mislove, Virgílio A. F. Almeida
ICWSM4
2014 Design of roadside infrastructure for information dissemination in vehicular networks
abstract
This work presents a probabilistic constructive heuristic to design the roadside infrastructure for information dissemination in vehicular networks. We formulate the problem as a Probabilistic Maximum Coverage Problem (PMCP) and we use them to maximize the number of vehicles in contact with the infrastructure. We compare our approach to a non-probabilistic MCP in simulated urban areas considering Manhattan-style topology with variable traffic conditions. The results reveal that our approach (Probabilistic MCP) increases the number of contacts between vehicles and dissemination points, optimizes the allocation of dissemination points, distributes the dissemination points in a layout that better fits the traffic flow and provides more regularity in the number of contacts experienced by vehicles.
Cristiano M. Silva, André L. L. de Aquino, Wagner Meira Jr.
NOMS3
2014 Economically-efficient sentiment stream analysis
abstract
Text-based social media channels, such as Twitter, produce torrents of opinionated data about the most diverse topics and entities. The analysis of such data (aka. sentiment analysis) is quickly becoming a key feature in recommender systems and search engines. A prominent approach to sentiment analysis is based on the application of classification techniques, that is, content is classified according to the attitude of the writer. A major challenge, however, is that Twitter follows the data stream model, and thus classifiers must operate with limited resources, including labeled data and time for building classification models. Also challenging is the fact that sentiment distribution may change as the stream evolves. In this paper we address these challenges by proposing algorithms that select relevant training instances at each time step, so that training sets are kept small while providing to the classifier the capabilities to suit itself to, and to recover itself from, different types of sentiment drifts. Simultaneously providing capabilities to the classifier, however, is a conflicting-objective problem, and our proposed algorithms employ basic notions of Economics in order to balance both capabilities. We performed the analysis of events that reverberated on Twitter, and the comparison against the state-of-the-art reveals improvements both in terms of error reduction (up to 14%) and reduction of training resources (by orders of magnitude).
Roberto L. de Oliveira Jr., Adriano Veloso, Adriano C. M. Pereira, Wagner Meira Jr., Renato Ferreira 0001, Srinivasan Parthasarathy 0001
SIGIR4
2014 Sentiment analysis on evolving social streams: how self-report imbalances can help
abstract
Real-time sentiment analysis is a challenging machine learning task, due to scarcity of labeled data and sudden changes in sentiment caused by real-world events that need to be instantly interpreted. In this paper we propose solutions to acquire labels and cope with concept drift in this setting, by using findings from social psychology on how humans prefer to disclose some types of emotions. In particular, we use findings that humans are more motivated to report positive feelings rather than negative feelings and also prefer to report extreme feelings rather than average feelings.
Pedro Henrique Calais Guerra, Wagner Meira Jr., Claire Cardie
WSDM2
2014 Thread scheduling and memory coalescing for dynamic vectorization of SPMD workloads
Teo Milanez, Caroline Collange, Fernando Magno Quintão Pereira, Wagner Meira Jr., Renato Ferreira 0001
Parallel Comput.4
2014 Approximate similarity search for online multimedia services on distributed CPU-GPU platforms
George Teodoro, Eduardo Valle, Nathan Mariano, Ricardo da Silva Torres, Wagner Meira Jr., Joel H. Saltz
VLDB J.5
2013 A Measure of Polarization on Social Media Networks Based on Community Boundaries
Pedro Henrique Calais Guerra, Wagner Meira Jr., Claire Cardie, Robert D. Kleinberg
ICWSM2
2013 Ladies First: Analyzing Gender Roles and Behaviors in Pinterest
Raphael Ottoni, João Paulo Pesce, Diego B. Las Casas, Geraldo Franciscani Jr., Wagner Meira Jr., Ponnurangam Kumaraguru, Virgílio A. F. Almeida
ICWSM5
2013 Exploiting non-content preference attributes through hybrid recommendation method
abstract
This paper explores a method for incorporating into a recommender system explicit representations of user's preferences over non-content attributes such as popularity, recency, and similarity of recommended items. We show how such attributes can be modeled as a preference vector that can be used in a vector-space content-based recommender, and how that content-based recommender can be integrated with various collaborative filtering techniques through re-weighting of Top-M recommendations. We evaluate this approach on several recommender systems datasets and collaborative filtering methods, and find that incorporating the three preference attributes can lead to a substantial increase in Top-50 precision while also enhancing diversity and novelty.
Fernando Mourão, Leonardo Rocha 0001, Joseph A. Konstan, Wagner Meira Jr.
RecSys4
2013 A KDD-Based Methodology to Rank Trust in e-Commerce Systems
abstract
Due to the growing popularity of the Web, there is an increasing number of people who perform e-business transactions. On the other hand, this popularity has also attracted the attention of criminals, raising the number of frauds on the Web and associated financial losses, which reach billions of dollars per year. This paper proposes a KDD-based methodology to detect fraud in e-payment systems. In order to evaluate this methodology we defined the concept of economic efficiency and applied it to an actual dataset of one of the largest Latin American electronic payment systems. The results show a very good performance, providing gains of up to 46.5% in comparison with the strategy currently employed by the company.
José Felipe Júnior, Adriano C. M. Pereira, Wagner Meira Jr., Adriano Veloso
Web Intelligence3
2013 aCSM: noise-free graph-based signatures to large-scale receptor-based ligand prediction
abstract
MOTIVATION: Receptor-ligand interactions are a central phenomenon in most biological systems. They are characterized by molecular recognition, a complex process mainly driven by physicochemical and structural properties of both receptor and ligand. Understanding and predicting these interactions are major steps towards protein ligand prediction, target identification, lead discovery and drug design. RESULTS: We propose a novel graph-based-binding pocket signature called aCSM, which proved to be efficient and effective in handling large-scale protein ligand prediction tasks. We compare our results with those described in the literature and demonstrate that our algorithm overcomes the competitor's techniques. Finally, we predict novel ligands for proteins from Trypanosoma cruzi, the parasite responsible for Chagas disease, and validate them in silico via a docking protocol, showing the applicability of the method in suggesting ligands for pockets in a real-world scenario. AVAILABILITY AND IMPLEMENTATION: Datasets and the source code are available at http://www.dcc.ufmg.br/∼dpires/acsm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Douglas E. V. Pires, Raquel Cardoso de Melo Minardi, Carlos H. Silveira, Frederico F. Campos, Wagner Meira Jr.
Bioinform.5
2013 Profiling divergences in GPU applications
abstract
SUMMARY The increasing programmability and the high computational power of graphics processing units make them attractive to general purpose programming. However, taking full benefit of this execution environment is a challenging task. One of these challenges stems from divergences, a phenomenon that occurs when threads that execute in lock‐step are forced to take different program paths because of branches in the code. In face of divergences, some threads will have to wait, idly, while their diverging siblings execute. Optimizing the code to avoid divergences is difficult because this task demands a deep understanding of programs that might be large and convoluted. To facilitate the detection of divergences, this paper introduces the divergence map, a data structure that indicates the location and the volume of divergences in a program. We build this map via dynamic profiling techniques, which we have implemented on top of an open source Parallel Thread Execution compiler. To illustrate the importance of the divergence map, we have used it to pinpoint the core regions that must be optimized in well‐known public applications. By hand optimizing some applications, we have added 9–11% speedups onto kernels that have already gone through the sieve of many programmers. Copyright © 2012 John Wiley & Sons, Ltd.
Bruno Coutinho, Diogo Sampaio, Fernando Magno Quintão Pereira, Wagner Meira Jr.
Concurr. Comput. Pract. Exp.4
2013 Temporal contexts: Effective text classification in evolving document collections
Leonardo Rocha 0001, Fernando Mourão, Hilton de Oliveira Mota, Thiago Salles, Marcos André Gonçalves, Wagner Meira Jr.
Inf. Syst.6
2012 Named Entity Disambiguation in Streaming Data
Alexandre Davis, Adriano Veloso, Altigran S. da Silva, Alberto H. F. Laender, Wagner Meira Jr.
ACL (1)5
2012 Studying User Footprints in Different Online Social Networks
abstract
With the growing popularity and usage of online social media services, people now have accounts (some times several) on multiple and diverse services like Facebook, Linked In, Twitter and You Tube. Publicly available information can be used to create a digital footprint of any user using these social media services. Generating such digital footprints can be very useful for personalization, profile management, detecting malicious behavior of users. A very important application of analyzing users' online digital footprints is to protect users from potential privacy and security risks arising from the huge publicly available user information. We extracted information about user identities on different social networks through Social Graph API, Friend Feed, and Profilactic, we collated our own dataset to create the digital footprints of the users. We used username, display name, description, location, profile image, and number of connections to generate the digital footprints of the user. We applied context specific techniques (e.g. Jaro Winkler similarity, Word net based ontologies) to measure the similarity of the user profiles on different social networks. We specifically focused on Twitter and Linked In. In this paper, we present the analysis and results from applying automated classifiers for disambiguating profiles belonging to the same user from different social networks. User ID and Name were found to be the most discriminative features for disambiguating user profiles. Using the most promising set of features and similarity metrics, we achieved accuracy, precision and recall of 98%, 99%, and 96%, respectively.
Anshu Malhotra, Luam C. Totti, Wagner Meira Jr., Ponnurangam Kumaraguru, Virgílio A. F. Almeida
ASONAM3
2012 An Adaptive Mesh Algorithm for the Numerical Solution of Electrical Models of the Heart
Rafael Sachetto Oliveira, Bernardo M. Rocha, Denise Burgarelli, Wagner Meira Jr., Rodrigo Weber dos Santos
ICCSA (1)4
2012 Data and Instruction Uniformity in Minimal Multi-threading
abstract
Simultaneous Multi-Threading (SMT) is a hardware model in which different threads share the same instruction fetching unit. This model is a compromise between high parallelism and low hardware cost. Minimal Multi-Threading (MMT) is a technique recently proposed to share instructions and execution between threads in a SMT machine. In this paper we propose new ways to explore redundancies in the MMT execution model. First, we propose and evaluate a new thread reconvergence heuristics that handles function calls better than previous approaches. Second, we demonstrate the existence of substantial regularity in inter-thread memory access patterns. We validate our results on the four data-parallel applications present in the PARSEC benchmark suite. The new thread reconvergence heuristics is, on the average, 82% more efficient than MMT's original reconvergence method. Furthermore, about 69% to 87% of all the memory addresses are either the same for all the threads, or are affine expressions of the thread identifier. This observation motivates the design of newly proposed hardware that benefits from regularity in inter-thread memory accesses.
Teo Milanez, Caroline Collange, Fernando Magno Quintão Pereira, Wagner Meira Jr., Renato Ferreira 0001
SBAC-PAD4
2012 HydroPaCe: understanding and predicting cross-inhibition in serine proteases through hydrophobic patch centroids
abstract
MOTIVATION: Protein-protein interfaces contain important information about molecular recognition. The discovery of conserved patterns is essential for understanding how substrates and inhibitors are bound and for predicting molecular binding. When an inhibitor binds to different enzymes (e.g. dissimilar sequences, structures or mechanisms what we call cross-inhibition), identification of invariants is a difficult task for which traditional methods may fail. RESULTS: To clarify how cross-inhibition happens, we model the problem, propose and evaluate a methodology called HydroPaCe to detect conserved patterns. Interfaces are modeled as graphs of atomic apolar interactions and hydrophobic patches are computed and summarized by centroids (HP-centroids), and their conservation is detected. Despite sequence and structure dissimilarity, our method achieves an appropriate level of abstraction to obtain invariant properties in cross-inhibition. We show examples in which HP-centroids successfully predicted enzymes that could be inhibited by the studied inhibitors according to BRENDA database. AVAILABILITY: www.dcc.ufmg.br/~raquelcm/hydropace CONTACT: [email protected]; [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Valdete M. Gonçalves-Almeida, Douglas E. V. Pires, Raquel Cardoso de Melo Minardi, Carlos H. Silveira, Wagner Meira Jr., Marcelo M. Santoro
Bioinform.5
2012 Cost-effective on-demand associative author name disambiguation
Adriano Veloso, Anderson A. Ferreira, Marcos André Gonçalves, Alberto H. F. Laender, Wagner Meira Jr.
Inf. Process. Manag.5
2012 Mining Attribute-structure Correlated Patterns in Large Attributed Graphs
abstract
In this work, we study the correlation between attribute sets and the occurrence of dense subgraphs in large attributed graphs, a task we call structural correlation pattern mining. A structural correlation pattern is a dense subgraph induced by a particular attribute set. Existing methods are not able to extract relevant knowledge regarding how vertex attributes interact with dense subgraphs. Structural correlation pattern mining combines aspects of frequent itemset and quasi-clique mining problems. We propose statistical significance measures that compare the structural correlation of attribute sets against their expected values using null models. Moreover, we evaluate the interestingness of structural correlation patterns in terms of size and density. An efficient algorithm that combines search and pruning strategies in the identification of the most relevant structural correlation patterns is presented. We apply our method for the analysis of three real-world attributed graphs: a collaboration, a music, and a citation network, verifying that it provides valuable knowledge in a feasible time.
Arlei Silva, Wagner Meira Jr., Mohammed J. Zaki
Proc. VLDB Endow.2
2011 Divergence Analysis and Optimizations
abstract
The growing interest in GPU programming has brought renewed attention to the Single Instruction Multiple Data (SIMD) execution model. SIMD machines give application developers a tremendous computational power, however, the model also brings restrictions. In particular, processing elements (PEs) execute in lock-step, and may lose performance due to divergences caused by conditional branches. In face of divergences, some PEs execute, while others wait, this alternation ending when they reach a synchronization point. In this paper we introduce divergence analysis, a static analysis that determines which program variables will have the same values for every PE. This analysis is useful in three different ways: it improves the translation of SIMD code to non-SIMD CPUs, it helps developers to manually improve their SIMD applications, and it also guides the compiler in the optimization of SIMD programs. We demonstrate this last point by introducing branch fusion, a new compiler optimization that identifies, via a gene sequencing algorithm, chains of similarities between divergent program paths, and weaves these paths together as much as possible. Our implementation has been accepted in the Ocelot open-source CUDA compiler, and is publicly available. We have tested it on many industrial-strength GPU benchmarks, including Rodinia and the Nvidia's SDK. Our divergence analysis has a 34% false-positive rate, compared to the results of a dynamic profiler. Our automatic optimization adds a 3% speed-up onto parallel quick sort, a heavily optimized benchmark. Our manual optimizations extend this number to over 10%.
Bruno Coutinho, Diogo Sampaio, Fernando Magno Quintão Pereira, Wagner Meira Jr.
PACT4
2011 Assessing documents' credibility with genetic programming
abstract
The concept of example credibility evaluates how much a classifier can trust an example when building a classification model. It is given by a credibility function, which is application dependent and estimated according to a series of factors that influence the credibility of the examples. Here we deal with automatic document classification and study the credibility of a document according to three factors: content, authorship and citations. We propose a genetic programming algorithm to estimate the credibility of training examples, and then add this estimation to a credibility-aware classifier. For that, we model the authorship and citation data as a complex network, and select a set of structural metrics that can be used to estimate credibility. These metrics are then merged with other content-related ones, and used as terminals for the GP. The GP was tested in a subset of the ACM-DL, and results showed that the credibility-aware classifier obtained results of micro and macroF1from 5% to 8% better than the traditional classifiers.
João R. M. Palotti, Thiago Salles, Gisele L. Pappa, Marcos André Gonçalves, Wagner Meira Jr.
IEEE Congress on Evolutionary Computation5
2011 Adaptive parallel approximate similarity search for responsive multimedia retrieval
abstract
This paper introduces Hypercurves, a flexible framework for pro- viding similarity search indexing to high throughput multimedia services. Hypercurves efficiently and effectively answers k-nearest neighbor searches on multigigabyte high-dimensional databases. It supports massively parallel processing and adapts at runtime its parallelization regimens to keep answer times optimal for either low and high demands. In order to achieve its goals, Hypercurves introduces new techniques for selecting parallelism configurations and allocating threads to computation cores, including hyperthreaded cores. Its efficiency gains are throughly validated on a large database of multimedia descriptors, where it presented near linear speedups and superlinear scaleups. The adaptation reduces query response times in 43% and 74% for both platforms tested, when compared to the best static parallelism regimens.
George Teodoro, Eduardo Valle, Nathan Mariano, Ricardo da Silva Torres, Wagner Meira Jr.
CIKM5
2011 Semi-supervised genetic programming for classification
abstract
Learning from unlabeled data provides innumerable advantages to a wide range of applications where there is a huge amount of unlabeled data freely available. Semi-supervised learning, which builds models from a small set of labeled examples and a potential large set of unlabeled examples, is a paradigm that may effectively use those unlabeled data. Here we propose KGP, a semi-supervised transductive genetic programming algorithm for classification. Apart from being one of the first semi-supervised algorithms, it is transductive (instead of inductive), i.e., it requires only a training dataset with labeled and unlabeled examples, which should represent the complete data domain. The algorithm relies on the three main assumptions on which semi-supervised algorithms are built, and performs both global search on labeled instances and local search on unlabeled instances. Periodically, unlabeled examples are moved to the labeled set after a weighted voting process performed by a committee. Results on eight UCI datasets were compared with Self-Training and KNN, and showed KGP as a promising method for semi-supervised learning.
Filipe de Lima Arcanjo, Gisele L. Pappa, Paulo Viana Bicalho, Wagner Meira Jr., Altigran S. da Silva
GECCO4
2011 From bias to opinion: a transfer-learning approach to real-time sentiment analysis
abstract
Real-time interaction, which enables live discussions, has become a key feature of most Web applications. In such an environment, the ability to automatically analyze user opinions and sentiments as discussions develop is a powerful resource known as real time sentiment analysis. However, this task comes with several challenges, including the need to deal with highly dynamic textual content that is characterized by changes in vocabulary and its subjective meaning and the lack of labeled data needed to support supervised classifiers. In this paper, we propose a transfer learning strategy to perform real time sentiment analysis. We identify a task - opinion holder bias prediction - which is strongly related to the sentiment analysis task; however, in constrast to sentiment analysis, it builds accurate models since the underlying relational data follows a stationary distribution.
Pedro Henrique Calais Guerra, Adriano Veloso, Wagner Meira Jr., Virgílio A. F. Almeida
KDD3
2011 Credibility of web applications
abstract
The popularization of Web has given rise to new services every day, demanding mechanisms to ensure the credibility of these services. Since now, little has been done to measure and understand the credibility of this complex Web environment, which itself is a major research challenge. From the challenges related to the task of assigning a credibility value to an online service in Web 2.0 applications, we propose a framework for the design, implementation and evaluation of credibility models. We call a credibility model a function capable of assigning a credibility value to a transaction of a Web application, considering different criteria of this service and its supplier. To validate this framework and models, we perform experiments using an actual dataset, from which we evaluated different credibility models using distinct types of information sources, and it allows to compare and evaluate these credibility models. The obtained results are very good, showing representative gains, when compared to a baseline and also with a known state-of-the-art approach. The results confirm that the credibility framework can be used to enforce trust to users of services on the Web.
Sara Guimarães, Adriano C. M. Pereira, Arlei Silva, Wagner Meira Jr.
MEDES4
2011 Is There a Best Quality Metric for Graph Clusters?
Hélio Marcos Paz de Almeida, Dorgival O. Guedes, Wagner Meira Jr., Mohammed J. Zaki
ECML/PKDD (1)3
2011 Watershed: A High Performance Distributed Stream Processing System
abstract
The task of extracting information from datasets that become larger at a daily basis, such as those collected from the web, is an increasing challenge, but also provides more interesting insights and analysis. Current analyses went beyond content and now focus on tracking and understanding users' relationships and interactions. Such computation is intensive both in terms of the processing demand imposed by the algorithms and also the sheer amount of data that has to handled. In this paper we introduce Watershed, a distributed computing framework designed to support the analysis of very large data streams online and in real-time. Data are obtained from streams by the system's processing components, transformed, and directed to other streams, creating large flows of information. The processing components are decoupled from each other and their connections are strictly data-driven. They can be dynamically inserted and removed, providing an environment in which it is feasible that different applications share intermediate results or cooperate to a global purpose. Our experiments demonstrate the flexibility in creating a set of data analysis algorithms and their composition into a powerful stream analysis environment.
Thatyene Louise Alves de Souza Ramos, Rodrigo Silva Oliveira, Ana Paula de Carvalho, Renato Ferreira 0001, Wagner Meira Jr.
SBAC-PAD5
2011 Distributed Skycube Computation with Anthill
abstract
Recently skyline queries have gained considerable attention and are among the most important tools for multi-criteria analysis. In order to process all possible combinations of criteria along with their inherent analysis, researchers introduced and studied the notion of \emph{skycube}. Simply put, a skycube is a pre-materialization of all possible subspaces with their associated skylines. An efficient skycube computation relies on the detection of redundancies in the different processing steps and enhanced result sharing between subspaces. Lately, the Orion algorithm was proposed to compute the skycube in a very efficient way. The approach relies on the derivation of skyline points over different subspaces. Nevertheless, because there are 2^{|D|} - 1 subspaces (where D is the set of dimensions) in a skycube, the running time still grows exponentially with the number of dimensions and easily becomes intractable on real-world datasets. In this study, we detail the distribution of Orion within a \emph{filter-stream} framework and we conduct an extensive set of experiments on large datasets collected from Twitter to demonstrate the efficiency of our method.
Renê Rodrigues Veloso, Loïc Cerf, Chedy Raïssi, Wagner Meira Jr.
SBAC-PAD4
2011 Data Integration via Constrained Clustering: An Application to Enzyme Clustering
abstract
When multiple data sources are available for clustering, an a priori data integration process is usually required. This process may be costly and may not lead to good clusterings, since important information is likely to be discarded. In this paper we propose constrained clustering as a strategy for integrating data sources without losing any information. It basically consists of adding the complementary data sources as constraints that the algorithm must satisfy. As a concrete application of our approach, we focus on the problem of enzyme function prediction, which is a hard task usually performed by intensive experimental work. We use constrained clustering as a means of integrating information from diverse sources as constraints, and analyze how this additional information impacts clustering quality in an enzyme clustering application scenario. Our results show that constraints generally improve the clustering quality when compared to an unconstrained clustering algorithm.
Elisa Boari de Lima, Raquel Cardoso de Melo Minardi, Wagner Meira Jr., Mohammed J. Zaki
SDM3
2011 Effective sentiment stream analysis with self-augmenting training and demand-driven projection
abstract
How do we analyze sentiments over a set of opinionated Twitter messages? This issue has been widely studied in recent years, with a prominent approach being based on the application of classification techniques. Basically, messages are classified according to the implicit attitude of the writer with respect to a query term. A major concern, however, is that Twitter (and other media channels) follows the data stream model, and thus the classifier must operate with limited resources, including labeled data for training classification models. This imposes serious challenges for current classification techniques, since they need to be constantly fed with fresh training messages, in order to track sentiment drift and to provide up-to-date sentiment analysis.
Ismael S. Silva, Janaína Gomide, Adriano Veloso, Wagner Meira Jr., Renato Ferreira 0001
SIGIR4
2011 A Characterization Methodology of Evolutionary Behavior in Recommender Systems
Alan Cardoso, Daniel Rocha, Rafael Sachetto Oliveira, Leonardo Rocha 0001, Fernando Mourão, Wagner Meira Jr.
WEBIST6
2011 Word co-occurrence features for text classification
Fábio Figueiredo, Leonardo Rocha 0001, Thierson Couto, Thiago Salles, Marcos André Gonçalves, Wagner Meira Jr.
Inf. Syst.6
2011 Calibrated lazy associative classification
Adriano Veloso, Wagner Meira Jr., Marcos André Gonçalves, Humberto Mossri de Almeida, Mohammed J. Zaki
Inf. Sci.2
2010 Tuning Genetic Programming parameters with factorial designs
abstract
Parameter setting of Evolutionary Algorithms is a time consuming task with two main approaches: parameter tuning and parameter control. In this work we describe a new methodology for tuning parameters of Genetic Programming algorithms using factorial designs, one-factor designs and multiple linear regression. Our experiments show that factorial designs can be used to determine which parameters have the largest effect on the algorithm's performance. This way, parameter setting efforts can focus on them, largely reducing the parameter search space. Two classical GP problems were studied, with six parameters for the first problem and seven for the second. The results show the maximum tree depth as the parameter with the largest effect on both problems. A one-factor design was performed to fine-tune tree depth on the first problem and a multiple linear regression to fine-tune tree depth and number of generations on the second.
Elisa Boari de Lima, Gisele L. Pappa, Jussara M. Almeida, Marcos André Gonçalves, Wagner Meira Jr.
IEEE Congress on Evolutionary Computation5
2010 CredibilityRank: A Framework for the Design and Evaluation of Rank-based Credibility Models for Web Applications
abstract
The popularization of Web has given rise to new services every day, demanding mechanisms to ensure the credibility of these services. Since now, little has been done to measure and understand the credibility of this complex Web environment, which itself is a major research challenge. From the challenges related to the task of assigning a credibility value to an online service in Web 2.0 applications, we propose a framework for the design, implementation and evaluation of credibility models. To validate the framework, we perform experiments using an actual dataset, from which we evaluated different credibility models using distinct types of information sources, and it allows to compare and evaluate these credibility models. The results show that the credibility framework is applicable and is capable of supporting decision making by users of Web services.
Sara Guimarães, Arlei Silva, Wagner Meira Jr., Adriano C. M. Pereira
EUC3
2010 Performance Debugging of GPGPU Applications with the Divergence Map
abstract
The increasing programability and the high computational power of Graphical Processing Units (GPU) make them attractive to general purpose programming. However, taking full benefit of this execution environment is a challenging task. One of these challenges stem from divergences, a phenomenon that occurs when threads that execute in lock-step are forced to take different program paths due to branches in the code. In face of divergences, some threads will have to wait, idly, while their diverging siblings execute. Optimizing the code to avoid divergences is difficult, because this task demands a deep understanding of programs that might be large and convoluted. In order to facilitate the detection of divergences, this paper introduces the divergence map, a data structure that indicates the location and the volume of divergences in a program. We build this map via dynamic profiling techniques, which we have implemented on top of an open source CUDA compiler. To illustrate the importance of the divergence map, we have used it to pin-point the core regions that must be optimized in well known public applications. By hand optimizing some applications, we have added 9-11% speedups onto kernels that have already gone through the sieve of many programmers.
Bruno Coutinho, Diogo Sampaio, Fernando Magno Quintão Pereira, Wagner Meira Jr.
SBAC-PAD4
2010 Tree Projection-Based Frequent Itemset Mining on Multicore CPUs and GPUs
abstract
Frequent itemset mining (FIM) is a core operation for several data mining applications as association rules computation, correlations, document classification, and many others, which has been extensively studied over the last decades. Moreover, databases are becoming increasingly larger, thus requiring a higher computing power to mine them in reasonable time. At the same time, the advances in high performance computing platforms are transforming them into hierarchical parallel environments equipped with multi-core processors and many-core accelerators, such as GPUs. Thus, fully exploiting these systems to perform FIM tasks poses as a challenging and critical problem that we address in this paper. We present efficient multi-core and GPU accelerated parallelizations of the Tree Projection, one of the most competitive FIM algorithms. The experimental results show that our Tree Projection implementation scales almost linearly in a CPU shared-memory environment after careful optimizations, while the GPU versions are up to 173 times faster than standard the CPU version.
George Teodoro, Nathan Mariano, Wagner Meira Jr., Renato Ferreira 0001
SBAC-PAD3
2010 Temporally-aware algorithms for document classification
abstract
Automatic Document Classification (ADC) is still one of the major information retrieval problems. It usually employs a supervised learning strategy, where we first build a classification model using pre-classified documents and then use this model to classify unseen documents. The majority of supervised algorithms consider that all documents provide equally important information. However, in practice, a document may be considered more or less important to build the classification model according to several factors, such as its timeliness, the venue where it was published in, its authors, among others. In this paper, we are particularly concerned with the impact that temporal effects may have on ADC and how to minimize such impact. In order to deal with these effects, we introduce a temporal weighting function (TWF) and propose a methodology to determine it for document collections. We applied the proposed methodology to ACM-DL and Medline and found that the TWF of both follows a lognormal. We then extend three ADC algorithms (namely kNN, Rocchio and Naïve Bayes) to incorporate the TWF. Experiments showed that the temporally-aware classifiers achieved significant gains, outperforming (or at least matching) state-of-the-art algorithms.
Thiago Salles, Leonardo Rocha 0001, Gisele L. Pappa, Fernando Mourão, Wagner Meira Jr., Marcos André Gonçalves
SIGIR5
2010 Distance-Based Outlier Detection: Consolidation and Renewed Bearing
abstract
Detecting outliers in data is an important problem with interesting applications in a myriad of domains ranging from data cleaning to financial fraud detection and from network intrusion detection to clinical diagnosis of diseases. Over the last decade of research, distance-based outlier detection algorithms have emerged as a viable, scalable, parameter-free alternative to the more traditional statistical approaches. In this paper we assess several distance-based outlier detection approaches and evaluate them. We begin by surveying and examining the design landscape of extant approaches, while identifying key design decisions of such approaches. We then implement an outlier detection framework and conduct a factorial design experiment to understand the pros and cons of various optimizations proposed by us as well as those proposed in the literature, both independently and in conjunction with one another, on a diverse set of real-life datasets. To the best of our knowledge this is the first such study in the literature. The outcome of this study is a family of state of the art distance-based outlier detection algorithms. Our detailed empirical study supports the following observations. The combination of optimization strategies enables significant efficiency gains. Our factorial design study highlights the important fact that no single optimization or combination of optimizations (factors) always dominates on all types of data. Our study also allows us to characterize when a certain combination of optimizations is likely to prevail and helps provide interesting and useful insights for moving forward in this domain.
Gustavo Henrique Orair, Carlos H. C. Teixeira, Wagner Meira Jr., Srinivasan Parthasarathy 0001
Proc. VLDB Endow.4
2009 Coordinating the use of GPU and CPU for improving performance of compute intensive applications
abstract
GPUs have recently evolved into very fast parallel co-processors capable of executing general purpose computations extremely efficiently. At the same time, multi-core CPUs evolution continued and today's CPUs have 4-8 cores. These two trends, however, have followed independent paths in the sense that we are aware of very few works that consider both devices cooperating to solve general computations. In this paper we investigate the coordinated use of CPU and GPU to improve efficiency of applications even further than using either device independently. We use Anthill runtime environment, a data-flow oriented framework in which applications are decomposed into a set of event-driven filters, where for each event, the runtime system can use either GPU or CPU for its processing. For evaluation, we use a histopathology application that uses image analysis techniques to classify tumor images for neuroblas-toma prognosis. Our experimental environment includes dual and octa-core machines, augmented with GPUs and we evaluate our approach's performance for standalone and distributed executions. Our experiments show that a pure GPU optimization of the application achieved a factor of 15 to 49 times improvement over the single core CPU version, depending on the versions of the CPUs and GPUs. We also show that the execution can be further reduced by a factor of about 2 by using our runtime system that effectively choreographs the execution to run cooperatively both on GPU and on a single core of CPU. We improve on that by adding more cores, all of which were previously neglected or used ineffectively. In addition, the evaluation on a distributed environment has shown near linear scalability to multiple hosts.
George Teodoro, Rafael Sachetto Oliveira, Olcay Sertel, Metin Nafi Gürcan, Wagner Meira Jr., Ümit V. Çatalyürek, Renato Ferreira 0001
CLUSTER5
2009 From an artificial neural network to a stock market day-trading system: A case study on the BM&F BOVESPA
abstract
Predicting trends in the stock market is a subject of major interest for both scholars and financial analysts. The main difficulties of this problem are related to the dynamic, complex, evolutive and chaotic nature of the markets. In order to tackle these problems, this work proposes a day-trading system that “translates” the outputs of an artificial neural network into business decisions, pointing out to the investors the best times to trade and make profits. The ANN forecasts the lowest and highest stock prices of the current trading day. The system was tested with the two main stocks of the BM&FBOVESPA, an important and understudied market. A series of experiments were performed using different data input configurations, and compared with four benchmarks. The results were evaluated using both classical evaluation metrics, such as the ANN generalization error, and more general metrics, such as the annualized return. The ANN showed to be more accurate and give more return to the investor than the four benchmarks. The best results obtained by the ANN had an mean absolute percentage error around 50% smaller than the best benchmark, and doubled the capital of the investor.
Leonardo C. Martinez, Diego N. da Hora, João R. M. Palotti, Wagner Meira Jr., Gisele L. Pappa
IJCNN4
2009 Assessing success factors of selling practices in electronic marketplaces
abstract
Electronic markets have early emerged as an important topic inside e-commerce research. An e-market is a digital ecosystem intended to provide their users with online services that will facilitate information exchange and transactions. This work presents a characterization and analysis of fixed-price online negotiations. Using actual data from a Brazilian marketplace, we analyze selling practices, considering seller profiles and selling strategies. There are important factors that can be considered when analyzing selling practices, such as the seller's reputation and experience, offer's price, duration, among others. We evaluate which factors impact on the success of selling practices in e-markets, which can be used to support seller's decision and recommend selling practices. Moreover, we investigate some important hypotheses about selling practices in online marketplaces, which allow us to state interesting conclusions, such as: a seller profile can achieve success or not in a trade, depending on the adopted strategy; the offer's price and how it is being advertised are two important success factors.
Adriano C. M. Pereira, Diego Duarte, Wagner Meira Jr., Paulo B. Góes
MEDES3
2009 The Metric Dilemma: Competence-Conscious Associative Classification
abstract
The classification performance of an associative classifier is strongly dependent on the statistic measure or metric that is used to quantify the strength of the association between features and classes (i.e., confidence, correlation etc.). Previous studies have shown that classifiers produced by different metrics may provide conflicting predictions, and that the best metric to use is data-dependent and rarely known while designing the classifier. This uncertainty concerning the optimal match between metrics and problems is a dilemma, and prevents associative classifiers to achieve their maximal performance. This dilemma is the focus of this paper. A possible solution to this dilemma is to learn the competence, expertise, or assertiveness of metrics. The basic idea is that each metric has a specific sub-domain for which it is most competent (i.e., it consistently produces more accurate classifiers than the ones produced by other metrics). Particularly, we investigate stacking-based meta-learning methods, which use the training data to find the domain of competence of each metric. The meta-classifier describes the domains of competence (or areas of expertise) of each metric, enabling a more sensible use of these metrics so that competence-conscious classifiers can be produced (i.e., a metric is only used to produce classifiers for test instances that belong to its domain of competence). We conducted a systematic evaluation, using different datasets and evaluation measures, of classifiers produced by different metrics. The result is that, while no metric is always superior than all others, the selection of appropriate metrics according to their competence/expertise (i.e., competence-conscious associative classifiers) seems very effective, showing gains that range from 7% to 26% when compared to the baselines (SVMs and an existing ensemble method).
Adriano Veloso, Mohammed J. Zaki, Wagner Meira Jr., Marcos André Gonçalves
SDM3
2009 Quantifying the Impact of Information Aggregation on Complex Networks: A Temporal Perspective
Fernando Mourão, Leonardo Rocha 0001, Lucas C. O. Miranda, Virgílio A. F. Almeida, Wagner Meira Jr.
WAW5
2009 Analyzing seller practices in a Brazilian marketplace
abstract
E-commerce is growing at an exponential rate. In the last decade, there has been an explosion of online commercial activity enabled by World Wide Web (WWW). These days, many consumers are less attracted to online auctions, preferring to buy merchandise quickly using fixed-price negotiations. Sales at Amazon.com, the leader in online sales of fixed-price goods, rose 37% in the first quarter of 2008. At eBay, where auctions make up 58% of the site's sales, revenue rose 14%. In Brazil, probably by cultural influence, online auctions are not been popular. This work presents a characterization and analysis of fixed-price online negotiations. Using actual data from a Brazilian marketplace, we analyze seller practices, considering seller profiles and strategies. We show that different sellers adopt strategies according to their interests, abilities and experience. Moreover, we confirm that choosing a selling strategy is not simple, since it is important to consider the seller's characteristics to evaluate the applicability of a strategy. The work also provides a comparative analysis of some selling practices in Brazil with popular worldwide marketplaces.
Adriano C. M. Pereira, Diego Duarte, Wagner Meira Jr., Virgílio A. F. Almeida, Paulo B. Góes
WWW3
2008 A seller's perspective characterization methodology for online auctions
abstract
Online auction services have reached great popularity and revenue over the last years. A key component for this success is the seller. Few studies proposed analyzing how the seller and the auction configuration affect the negotiation results. In this work we propose a methodology to characterize online auctions by the seller's perspective. This methodology is based on: (1) recognizing the characteristics of the variables related to the auction results and (2) capturing the correlation among these variables to identify seller profiles and selling strategies. We applied our methodology to a real case study, using an eBay dataset, to validate two hypotheses about sellers and their practices. These results are useful to understand the complex mechanisms that guide ending prices, success (or failure), and the attraction of bids in online auctions, which can support decision strategies for buyers and sellers.
Arlei Silva, Pedro H. Calais, Adriano C. M. Pereira, Fernando Mourão, Jussara M. Almeida, Wagner Meira Jr., Paulo B. Góes
ICEC6
2008 Exploiting temporal contexts in text classification
abstract
Due to the increasing amount of information being stored and accessible through the Web, Automatic Document Classification (ADC) has become an important research topic. ADC usually employs a supervised learning strategy, where we first build a classification model using pre-classified documents and then use it to classify unseen documents. One major challenge in building classifiers is dealing with the temporal evolution of the characteristics of the documents and the classes to which they belong. However, most of the current techniques for ADC do not consider this evolution while building and using the models. Previous results show that the performance of classifiers may be affected by three different temporal effects (class distribution, term distribution and class similarity). Further, it is shown that using just portions of the pre-classified documents, which we call contexts, for building the classifiers, result in better performance, as a consequence of the minimization of the aforementioned effects.
Leonardo Rocha 0001, Fernando Mourão, Adriano C. M. Pereira, Marcos André Gonçalves, Wagner Meira Jr.
CIKM5
2008 Achieving Multi-Level Parallelism in the Filter-Labeled Stream Programming Model
abstract
New architectural trends in chip design resulted in machines with multiple processing units as well as efficient communication networks, leading to the wide availability of systems that provide multiple levels of parallelism, both inter- and intra-machine. Developing applications that efficiently make use of such systems is a challenge, specially for application-domain programmers. In this paper we present a new version of the Anthill programming environment that efficiently exploits multi-level parallelism and experimental results that demonstrate such efficiency. Anthill is based on the filter-stream model; in this model, applications are decomposed into a set of filters communicating through streams, which has already been shown to be efficient for expressing inter-machine parallelism. We replaced the filter run-time environment, originally process-oriented, with an event-oriented version. This new version allow programmers to efficiently express opportunities for parallelism within each compute node through a higher-level programming abstraction. We evaluated our solution on dual- and quad-core machines with two data mining applications: Eclat and KNN. Both had drops in execution time nearly proportional to the number of cores on a single machine. When using a cluster of dual-core machines, speed-ups were close to linear on the number of available cores for both applications, confirming event-oriented Anthill performs well both on the inter- and intra-machine parallelism levels.
George Teodoro, Daniel Fireman, Dorgival O. Guedes, Wagner Meira Jr., Renato Ferreira 0001
ICPP4
2008 Learning to rank at query-time using association rules
abstract
Some applications have to present their results in the form of ranked lists. This is the case of many information retrieval applications, in which documents must be sorted according to their relevance to a given query. This has led the interest of the information retrieval community in methods that automatically learn effective ranking functions. In this paper we propose a novel method which uncovers patterns (or rules) in the training data associating features of the document with its relevance to the query, and then uses the discovered rules to rank documents. To address typical problems that are inherent to the utilization of association rules (such as missing rules and rule explosion), the proposed method generates rules on a demand-driven basis, at query-time. The result is an extremely fast and effective ranking method. We conducted a systematic evaluation of the proposed method using the LETOR benchmark collections. We show that generating rules on a demand-driven basis can boost ranking performance, providing gains ranging from 12 % to 123%, outperforming the state-of-the-art methods that learn to rank, with no need of time-consuming and laborious pre-processing. As a highlight, we also show that additional information, such as query terms, can make the generated rules more discriminative, further improving ranking performance.
Adriano Veloso, Humberto Mossri de Almeida, Marcos André Gonçalves, Wagner Meira Jr.
SIGIR4
2008 Evaluating Longitudinal Aspects of Online Bidding Behavior
Leonardo Rocha 0001, Adriano C. M. Pereira, Fernando Mourão, Arlei Silva, Wagner Meira Jr., Paulo B. Góes
WEBIST (2)5
2008 Understanding temporal aspects in document classification
abstract
Due to the increasing amount of information present on the Web, Automatic Document Classification (ADC) has become an important research topic. ADC usually follows a standard supervised learning strategy, where we first build a model using preclassified documents and then use it to classify new unseen documents. One major challenge for ADC in many scenarios is that the characteristics of the documents and the classes to which they belong may change over time. However, most of the current techniques for ADC are applied without taking into account the temporal evolution of the collection of documents
Fernando Mourão, Leonardo Rocha 0001, Renata Braga Araújo, Thierson Couto, Marcos André Gonçalves, Wagner Meira Jr.
WSDM6
2008 Reactivity-based Approaches To Improve Web System's QoS
Adriano C. M. Pereira, Wagner Meira Jr., Walter D. S. Filho
J. Web Eng.3
2007 An Efficient and Reliable Scientific Workflow System
abstract
This paper presents a fault tolerance framework for applications that process data using a distributed network of user-defined operations in a pipelined fashion. The framework saves intermediate results and messages exchanged among application components in a distributed data management system to facilitate quick recovery from failures. The experimental results show that the framework scales well and our approach introduces very little overhead to application execution.
Tulio Tavares, George Teodoro, Tahsin M. Kurç, Renato Ferreira 0001, Dorgival O. Guedes, Wagner Meira Jr., Ümit V. Çatalyürek, Shannon Hastings, Scott Oster, Stephen Langella, Joel H. Saltz
CCGRID6
2007 Automatic Moderation of Comments in a Large On-line Journalistic Environment
Adriano Veloso, Wagner Meira Jr., Tiago Alves Macambira, Dorgival O. Guedes, Hélio Marcos Paz de Almeida
ICWSM2
2007 Limiting the power consumption of main memory
abstract
The peak power consumption of hardware components affects their powersupply, packaging, and cooling requirements. When the peak power consumption is high, the hardware components or the systems that use them can become expensive and bulky. Given that components and systems rarely (if ever) actually require peak power, it is highly desirable to limit power consumption to a less-than-peak power budget, based on which power supply, packaging, and cooling infrastructure scan be more intelligently provisioned.
Bruno Diniz, Dorgival O. Guedes, Wagner Meira Jr., Ricardo Bianchini
ISCA3
2007 Multi-label Lazy Associative Classification
Adriano Veloso, Wagner Meira Jr., Marcos André Gonçalves, Mohammed J. Zaki
PKDD2
2007 Fault-tolerance in filter-labeled-stream applications
abstract
Fault tolerance is a desirable feature in distributed high-performance systems, since applications tend to run for long periods of time and faults become more likely as the number of nodes in the system increase. However, most distributed environments lack any fault tolerant features, since they tend to be hard to implement and use, and often hurt performance dramatically. In this paper we discuss how we successfully added fault-tolerance to the Anthill distributed programming environment by using an application-level checkpoint/rollback solution. The programming model offers an abstraction where the programmer can easily identify points during the execution where the communication pattern is well defined, forming a consistent cut where checkpoints may be saved consistently without requiring extra communication, avoiding any domino effect during recovery from faults. We present the new abstractions for fault tolerance, describe how the solution was implemented and present performance results that show the efficiency of the solution with both regular and irregular applications.
Bruno Coutinho, Dorgival O. Guedes, Wagner Meira Jr., Renato Ferreira 0001
SBAC-PAD3
2007 A Scalable Parallel Deduplication Algorithm
abstract
The identification of replicas in a database is fundamental to improve the quality of the information. Deduplication is the task of identifying replicas in a database that refer to the same real world entity. This process is not always trivial, because data may be corrupted during their gathering, storing or even manipulation. Problems such as misspelled names, data truncation, data input in a wrong format, lack of conventions (like how to abbreviate a name), missing data or even fraud may lead to the insertion of replicas in a database. The deduplication process may be very hard, if not impossible, to be performed manually, since actual databases may have hundreds of millions of records. In this paper, we present our parallel deduplication algorithm, called FER- APARDA. By using probabilistic record linkage, we were able to successfully detect replicas in synthetic datasets with more than 1 million records in about 7 minutes using a 20- computer cluster, achieving an almost linear speedup. We believe that our results do not have similar in the literature when it comes to the size of the data set and the processing time.
Walter Santos, Thiago Teixeira, Carla Machado, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Altigran S. da Silva
SBAC-PAD4
2007 Multi-level Parallelism in the Computational Modeling of the Heart
abstract
Computational modeling of the heart has demonstrated to be a useful tool for the investigation and comprehension of the complex biophysical processes that underlie cardiac function. Unfortunately, large scale simulations, such as those resulting from the discretization of an entire heart, remain a computational challenge. In order to reduce simulation execution times, parallel implementations have traditionally exploited data parallelism via numerical schemes based on domain-decomposition. However, it has been verified that the parallel efficiency of these implementations severely degrades as the number of processors increases. In this work, we propose and implement a new parallel algorithm for the solution of cardiac models. By relaxing the coherence of the execution, a new level of parallelism could be identified and exploited: pipelining. A synchronous parallel algorithm that uses both pipelining and data decomposition techniques was implemented and used the MPI library for communication. Numerical tests were performed in a 8-node linux-cluster. Our preliminary results indicate that the proposed algorithm is able to increase the parallel efficiency up to 20% when compared to the traditional approach that uses pure data-level parallelism. In addition, the numerical precision was kept under control (relative errors under 4%) when the relaxed coherence execution was adopted.
Carolina Ribeiro Xavier, Rafael Sachetto Oliveira, Vinícius F. Vieira, Rodrigo Weber dos Santos, Wagner Meira Jr.
SBAC-PAD5
2007 Analyzing ebay Negotiation Patterns
Adriano C. M. Pereira, Leonardo Rocha 0001, Fernando Mourão, T. Torres, Wagner Meira Jr., Paulo B. Góes
WEBIST (3)5
2007 Workload models of spam and legitimate e-mails
Luíz Henrique Gomes, Cristiano Cazita, Jussara M. Almeida, Virgílio A. F. Almeida, Wagner Meira Jr.
Perform. Evaluation5
2006 Assessing Data Virtualization for Irregularly Replicated Large Datasets
abstract
Large volumes of data are generated every day by experiments, simulations and all sorts of applications. It is common to observe situations where portions of data are irregularly replicated and distributed in different data sources. It would be desirable to be able to handle these several pieces of irregular data (replicated or not) as a unique large dataset. This is called data virtualization and is the focus of this paper. In this paper, we present a system which is capable of dealing with irregularly replicated data and is able to create a virtual view of the union of the individual irregular portions of data hosted by each data source. Our system indexes the data intervals from each data source and allows clients to submit queries against the virtual dataset created. In order to select what server will be responsible for each data interval of a query, we use and compare three algorithms, namely Random, Round-Robin and Weighted Round-Robin. The comparison is driven by simulation and the parameters for the simulation are all taken from a real data-centered application (the Virtual Microscope).
Bruno Diniz, Diego L. Nogueira, André Cardoso, Renato Ferreira 0001, Dorgival O. Guedes, Wagner Meira Jr.
CCGRID6
2006 Multi-evidence, multi-criteria, lazy associative document classification
abstract
We present a novel approach for classifying documents that combines different pieces of evidence (e.g., textual features of documents, links, and citations) transparently, through a data mining technique which generates rules associating these pieces of evidence to predefined classes. These rules can contain any number and mixture of the available evidence and are associated with several quality criteria which can be used in conjunction to choose the "best" rule to be applied at classification time. Our method is able to perform evidence enhancement by link forwarding/backwarding (i.e., navigating among documents related through citation), so that new pieces of link-based evidence are derived when necessary. Furthermore, instead of inducing a single model (or rule set) that is good on average for all predictions, the proposed approach employs a lazy method which delays the inductive process until a document is given for classification, therefore taking advantage of better qualitative evidence coming from the document. We conducted a systematic evaluation of the proposed approach using documents from the ACM Digital Library and from a Brazilian Web directory. Our approach was able to outperform in both collections all classifiers based on the best available evidence in isolation as well as state-of-the-art multi-evidence classifiers. We also evaluated our approach using the standard WebKB collection, where our approach showed gains of 1% in accuracy, being 25 times faster. Further, our approach is extremely efficient in terms of computational performance, showing gains of more than one order of magnitude when compared against other multi-evidence classifiers.
Adriano Veloso, Wagner Meira Jr., Marco Cristo, Marcos André Gonçalves, Mohammed J. Zaki
CIKM2
2006 Lazy Associative Classification
abstract
Decision tree classifiers perform a greedy search for rules by heuristically selecting the most promising features. Such greedy (local) search may discard important rules. Associative classifiers, on the other hand, perform a global search for rules satisfying some quality constraints (i.e., minimum support). This global search, however, may generate a large number of rules. Further, many of these rules may be useless during classification, and worst, important rules may never be mined. Lazy (non-eager) associative classification overcomes this problem by focusing on the features of the given test instance, increasing the chance of generating more rules that are useful for classifying the test instance. In this paper we assess the performance of lazy associative classification. First we demonstrate that an associative classifier performs no worse than the corresponding decision tree classifier. Also we demonstrate that lazy classifiers outperform the corresponding eager ones. Our claims are empirically confirmed by an extensive set of experimental results. We show that our proposed lazy associative classifier is responsible for an error rate reduction of approximately 10 % when compared against its eager counterpart, and for a reduction of 20 % when compared against a decision tree classifier. A simple caching mechanism makes lazy associative classification fast, and thus improvements in the execution time are also observed. 1
Adriano Veloso, Wagner Meira Jr., Mohammed J. Zaki
ICDM2
2006 Evaluating the impact of reactivity on the performance of Web applications
abstract
The great success of the Internet has raised new challenges in terms of applications and the satisfaction of their users. In fact, there is strong evidence that a significant part of the user behavior depends on its satisfaction. Users reactions may affect the load of a server, establishing successive interactions where the user behavior affects the system behavior and vice-versa. It is important to understand this interactive process to design systems more suited to user requirements. In this work we study and explain how this reactive interaction is performed by users and how it affects the system's performance. We perform experiments using a real server under a TPC-W-based workload generated using a reactive version of httperf. We also simulate different workload configurations in order to evaluate the effects on the system's load. The results show that accounting for reactivity causes a significant impact on the server's performance in terms of throughput and response time, raising the possibility of performance improvement of Web systems by considering reactivity
Adriano C. M. Pereira, Wagner Meira Jr.
IPCCC3
2006 Assessing the impact of reactive workloads on the performance of Web applications
abstract
Designing systems with better performance and scalability is a real need to fulfill the user demands and generate profitable Web services. Being able to mimic user behavior and the workload they generate on the servers is fundamental to evaluate the performance of systems and their improvements. One aspect that is usually neglected by workload generators is the user reactivity, that is, how the users react to variable server response time. Further, it is not clear how the reactivity-related changes in the user generated workload affect the server and how these dependences converge. This paper addresses this problem by proposing, implementing, and validating a workload generator that accounts for reactivity while interacting with servers. Our workload generator is used, for instance, to generate workloads based on a TPC-W benchmark. These workloads are used to assess the impacts of reactivity on the performance of a Web application. The results show significant changes in terms of throughput and response time for the experiments, raising the possibility of improving the performance of Web systems considering user reactivity.
Adriano C. M. Pereira, Wagner Meira Jr., Walter Santos
ISPASS3
2006 ParTriCluster: A Scalable Parallel Algorithm for Gene Expression Analysis
abstract
Analyzing gene expression patterns is becoming a highly relevant task in the bio informatics area. This analysis makes it possible to determine the behavior patterns of genes under various conditions, a fundamental information for treating diseases, among other applications. An advance in this area is the tricluster algorithm, which is the first algorithm capable of determining 3D clusters, that is, it determines clusters of sets of genes that behave similarly in a set of samples and set of time stamps. However, while biological experiments collect an increasing amount of data to be analyzed and correlated, the triclustering problem is NP-complete, and its parallelization seems to be an essential step towards obtaining feasible solutions. In this paper we propose and evaluate the implementation of a parallel version of the tricluster algorithm using the filter-labeled-stream paradigm supported by the Anthill parallel programming environment. The results show that our parallelization scales linearly with the data size. Further, the parallelization strategy is applicable to any depth-first searches
Renata Braga Araújo, Guilherme Henrique Trielli Ferreira, Gustavo Henrique Orair, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes
SBAC-PAD4
2006 A Run-time System for Efficient Execution of Scientific Workflows on Distributed Environments
abstract
Scientific workflow systems have been introduced in response to the demand of researchers from several domains of science who need to process and analyze increasingly larger datasets. The design of these systems is largely based on the observation that data analysis applications can be composed as pipelines or networks of computations on data. In this paper we present a run-time support system that is designed to facilitate this type of computation in distributed computing environments. Our system is optimized for data-intensive workflows, in which efficient management and retrieval of data, coordination of data processing and data movement, and check-pointing of intermediate results are critical and challenging issues. Experimental evaluation of our system shows that linear speedups can be achieved for sophisticated applications, which are implemented as a network of multiple data processing components
George Teodoro, Tulio Tavares, Renato Ferreira 0001, Tahsin M. Kurç, Wagner Meira Jr., Dorgival O. Guedes, Tony Pan, Joel H. Saltz
SBAC-PAD5
2006 A hierarchical characterization of a live streaming media workload
Eveline Veloso, Virgílio A. F. Almeida, Wagner Meira Jr., Azer Bestavros, Shudong Jin
IEEE/ACM Trans. Netw.3
2005 Maximal termsets as a query structuring mechanism
abstract
Search engines process queries conjunctively to restrict the size of the answer set. Further, it is not rare to observe a mismatch between the vocabulary used in the text of Web pages and the terms used to compose the Web queries. The combination of these two features might lead to irrelevant query results, particularly in the case of more specific queries composed of three or more terms. To deal with this problem we propose a new technique for automatically structuring Web queries as a set of smaller subqueries. To select representative subqueries we use information on their distributions in the document collection. This can be adequately modeled using the concept of maximal termsets derived from the formalism of association rules theory. Experimentation shows that our technique leads to improved results. For the TREC-8 test collection, for instance, our technique led to gains in average precision of roughly 28% with regard to a BM25 ranking formula.
Bruno Pôssas, Nivio Ziviani, Berthier A. Ribeiro-Neto, Wagner Meira Jr.
CIKM4
2005 Scheduling Data Flow Applications Using Linear Programming
abstract
Grid environments are becoming cost-effective substitutes to supercomputers. Datacutter is one of several initiatives in creating mechanisms for applications to efficiently exploit the vast computation power of such environments. In Datacutter, applications are modeled as a set of communicating filters that may run on several nodes of a computational grid. To achieve high performance, a number of transparent copies of each of the filters that comprise the application need to be appropriately placed on different nodes of the grid. Such task is carried out by a scheduler which is the focus of this work. We present LPSched, a scheduler for Datacutter applications which uses linear programming to make decisions about the number of copies of each filter as well as the placement of each of the copies across the nodes. LPSched bases its decisions upon the performance behavior of each filter as well as the resources currently available on the grid.
Luiz Thomaz do Nascimento, Renato Ferreira 0001, Wagner Meira Jr., Dorgival O. Guedes
ICPP3
2005 AnthillSched: A Scheduling Strategy for Irregular and Iterative I/O-Intensive Parallel Jobs
Fabrício Góes, Pedro Henrique Calais Guerra, Bruno Coutinho, Leonardo Rocha 0001, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Walfredo Cirne
JSSPP5
2005 Energy conservation in heterogeneous server clusters
abstract
The previous research on cluster-based servers has focused on homogeneous systems. However, real-life clusters are almost invariably heterogeneous in terms of the performance, capacity, and power consumption of their hardware components. In this paper, we argue that designing efficient servers for heterogeneous clusters requires defining an efficiency metric, modeling the different types of nodes with respect to the metric, and searching for request distributions that optimize the metric. To concretely illustrate this process, we design a cooperative Web server for a heterogeneous cluster that uses modeling and optimization to minimize the energy consumed per request. Our experimental results for a cluster comprised of traditional and blade nodes show that our server can consume 42 % less energy than an energy-oblivious server, with only a negligible loss in throughput. The results also show that our server conserves 45 % more energy than an energy-conscious server that was previously proposed for homogeneous clusters. 1
Taliver Heath, Bruno Diniz, Enrique V. Carrera, Wagner Meira Jr., Ricardo Bianchini
PPoPP4
2005 Anthill: A Scalable Run-Time Environment for Data Mining Applications
abstract
Data mining techniques are becoming increasingly more popular as a reasonable means to collect summaries from the rapidly growing datasets in many areas. However, as the size of the raw data increases, parallel data mining algorithms are becoming a necessity. In this paper, we present a run-time support system that was designed to allow the efficient implementation of data-mining algorithms on heterogeneous distributed environments. We believe that the runtime framework is suitable for a broader class of applications, beyond data mining. We also present a parallelization strategy that is supported by the run-time system. We show scalability results of three different data-mining algorithms that were parallelized using our approach and our run-time support. All applications scale almost linearly up to a large number of nodes.
Renato Ferreira 0001, Wagner Meira Jr., Dorgival O. Guedes, Lúcia M. A. Drummond, Bruno Coutinho, George Teodoro, Tulio Tavares, Renata Braga Araújo, Guilherme T. Ferreira
SBAC-PAD2
2005 Set-based vector model: An efficient approach for correlation-based ranking
abstract
This work presents a new approach for ranking documents in the vector space model. The novelty lies in two fronts. First, patterns of term co-occurrence are taken into account and are processed efficiently. Second, term weights are generated using a data mining technique called association rules. This leads to a new ranking mechanism called the set-based vector model . The components of our model are no longer index terms but index termsets, where a termset is a set of index terms. Termsets capture the intuition that semantically related terms appear close to each other in a document. They can be efficiently obtained by limiting the computation to small passages of text. Once termsets have been computed, the ranking is calculated as a function of the termset frequency in the document and its scarcity in the document collection. Experimental results show that the set-based vector model improves average precision for all collections and query types evaluated, while keeping computational costs small. For the 2-gigabyte TREC-8 collection, the set-based vector model leads to a gain in average precision figures of 14.7% and 16.4% for disjunctive and conjunctive queries, respectively, with respect to the standard vector space model. These gains increase to 24.9% and 30.0%, respectively, when proximity information is taken into account. Query processing times are larger but, on average, still comparable to those obtained with the standard vector model (increases in processing time varied from 30% to 300%). Our results suggest that the set-based vector model provides a correlation-based ranking formula that is effective with general collections and computationally practical.
Bruno Pôssas, Nivio Ziviani, Wagner Meira Jr., Berthier A. Ribeiro-Neto
ACM Trans. Inf. Syst.3
2005 Quantifying the Performability of Cluster-Based Services
abstract
In this paper, we propose a two-phase methodology for systematically evaluating the performability (performance and availability) of cluster-based Internet services. In the first phase, evaluators use a fault-injection infrastructure to characterize the service's behavior in the presence of faults. In the second phase, evaluators use an analytical model to combine an expected fault load with measurements from the first phase to assess the service's performability. Using this model, evaluators can study the service's sensitivity to different design decisions, fault rates, and other environmental factors. To demonstrate our methodology, we study the performability of a multitier Internet service. In particular, we evaluate the performance and availability of three soft state maintenance strategies for an online bookstore service in the presence of seven classes of faults. Among other interesting results, we clearly isolate the effect of different faults, showing that the tier of Web servers is responsible for an often dominant fraction of the service unavailability. Our results also demonstrate that storing the soft state in a database achieves better performability than storing it in main memory (even when the state is efficiently replicated) when we weight performance and availability equally. Based on our results, we conclude that service designers may want an unbalanced system in which they heavily load highly available components and leave more spare capacity for components that are likely to fail more often.
Kiran Nagaraja, Gustavo Machado Campagnani Gama, Ricardo Bianchini, Richard P. Martin, Wagner Meira Jr., Thu D. Nguyen
IEEE Trans. Parallel Distributed Syst.5
2004 Characterizing a spam traffic
abstract
The rapid increase in the volume of unsolicited commercial e-mails, also known as spam, is beginning to take its toll in system administrators, business corporations and end-users. Widely varying estimates of the cost associated with spam are available in the literature. However, a quantitative analysis of the determinant characteristics of spam traffic is still an open problem. This work fills this gap and presents what we believe to be the first extensive characterization of a spam traffic.
Luíz Henrique Gomes, Cristiano Cazita, Jussara M. Almeida, Virgílio A. F. Almeida, Wagner Meira Jr.
Internet Measurement Conference5
2004 Asynchronous and Anticipatory Filter-Stream Based Parallel Algorithm for Frequent Itemset Mining
Adriano Veloso, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Srinivasan Parthasarathy 0001
PKDD2
2004 MASKS: Managing Anonymity while Sharing Knowledge to Servers
abstract
This work presents an architecture that allows users to enhance their privacy control over the computational environment. Web privacy is a topic that is raising, nowadays, many discussions. Usually, people do not know how their privacy can be violated or what can be done to protect it. Among the generated conflicts, we would like to show up the one that happens between privacy and personalization: by one side, users appreciate the idea of receiving personalized services and do not approve the collection, tracing and analysis of their actions; by the other side, personalization services need this type of information in order to profile their users. The architecture presented in this article helps users to understand better how their privacy can be invaded and, at the same time, gives them a better control of their privacy, through anonymity, without preventing them from receiving personalized services.
Robert Pinto, Lucila Ishitani, Virgílio A. F. Almeida, Wagner Meira Jr., Fabiano A. Fonseca, Fernando D. O. Castro
SEC4
2004 Processing Conjunctive and Phrase Queries with the Set-Based Model
Bruno Pôssas, Nivio Ziviani, Berthier A. Ribeiro-Neto, Wagner Meira Jr.
SPIRE4
2004 State Maintenance and its Impact on the Performability of Multi-tiered Internet Services
abstract
In this paper, we evaluate the performance, availability, and combined performability of four soft state maintenance strategies in two multitier Internet services, an online book store and an auction service. To take soft state and service latency into account, we propose an extension of our previous quantification methodology, and novel availability and performability metrics. Our results demonstrate that storing the soft state in a database can achieve better performability than storing it in main memory, even when the state is efficiently replicated. Strategies that offload the handling of soft state from the database increase the load on other tiers and, consequently, increase the impact of faults in these tiers on service availability. Based on these results, we conclude that service designers need to provision the cluster and balance the load with availability and cost, as well as performance, in mind.
Gustavo Machado Campagnani Gama, Kiran Nagaraja, Ricardo Bianchini, Richard P. Martin, Wagner Meira Jr., Thu D. Nguyen
SRDS5
2004 Parallel and distributed methods for incremental frequent itemset mining
abstract
Traditional methods for data mining typically make the assumption that the data is centralized, memory-resident, and static. This assumption is no longer tenable. Such methods waste computational and input/output (I/O) resources when data is dynamic, and they impose excessive communication overhead when data is distributed. Efficient implementation of incremental data mining methods is, thus, becoming crucial for ensuring system scalability and facilitating knowledge discovery when data is dynamic and distributed. In this paper, we address this issue in the context of the important task of frequent itemset mining. We first present an efficient algorithm which dynamically maintains the required information even in the presence of data updates without examining the entire dataset. We then show how to parallelize this incremental algorithm. We also propose a distributed asynchronous algorithm, which imposes minimal communication overhead for mining distributed dynamic datasets. Our distributed approach is capable of generating local models (in which each site has a summary of its own database) as well as the global model of frequent itemsets (in which all sites have a summary of the entire database). This ability permits our approach not only to generate frequent itemsets, but also to generate high-contrast frequent itemsets, which allows one to examine how the data is skewed over different sites.
Matthew Eric Otey, Srinivasan Parthasarathy 0001, Chao Wang 0050, Adriano Veloso, Wagner Meira Jr.
IEEE Trans. Syst. Man Cybern. Part B5
2003 Parallel and Distributed Frequent Itemset Mining on Dynamic Datasets
Adriano Veloso, Matthew Eric Otey, Srinivasan Parthasarathy 0001, Wagner Meira Jr.
HiPC4
2003 Mining Frequent Itemsets in Distributed and Dynamic Databases
abstract
Traditional methods for frequent itemset mining typically assume that data is centralized and static. Such methods impose excessive communication overhead when data is distributed, and they waste computational resources when data is dynamic. We present what we believe to be the first unified approach that overcomes these assumptions. Our approach makes use of parallel and incremental techniques to generate frequent itemsets in the presence of data updates without examining the entire database, and imposes minimal communication overhead when mining distributed databases. Further, our approach is able to generate both local and global frequent itemsets. This ability permits our approach to identify high-contrast frequent itemsets, which allows one to examine how the data is skewed over different sites.
Matthew Eric Otey, Chao Wang 0050, Srinivasan Parthasarathy 0001, Adriano Veloso, Wagner Meira Jr.
ICDM5
2003 Load Balancing on Stateful Clustered Web Servers
abstract
One of the main challenges to the wide use of the Internet is the scalability of the servers, that is, their ability to handle the increasing demand. Scalability in stateful servers, which comprise e-commerce and other transaction-oriented servers, is even more difficult, since it is necessary to keep transaction data across requests from the same user. One common strategy for achieving scalability is to employ clustered servers, where the load is distributed among the various servers. However, as a consequence of the workload characteristics and the need of maintaining data coherent among the servers that compose the cluster, load imbalance arise among servers, reducing the efficiency of the server as a whole. We propose and evaluate a strategy for load balancing in stateful clustered servers. Our strategy is based on control theory and allowed significant gains over configurations that do not employ the load balancing strategy, reducing the response time in up to 50% and increasing the throughput in up to 16%.
George Teodoro, Tulio Tavares, Bruno Coutinho, Wagner Meira Jr., Dorgival O. Guedes
SBAC-PAD4
2003 New Parallel Algorithms for Frequent Itemset Mining in Very Large Databases
abstract
Frequent itemset mining is a classic problem in data mining. It is a nonsupervised process which concerns in finding frequent patterns (or itemsets) hidden in large volumes of data in order to produce compact summaries or models of the database. These models are typically used to generate association rules, but recently they have also been used in far reaching domains like e-commerce and bio-informatics. Because databases are increasing in terms of both dimension (number of attributes) and size (number of records), one of the main issues in a frequent itemset mining algorithm is the ability to analyze very large databases. Sequential algorithms do not have this ability, especially in terms of run-time performance, for such very large databases. Therefore, we must rely on high performance parallel and distributed computing. We present new parallel algorithms for frequent itemset mining. Their efficiency is proven through a series of experiments on different parallel environments, that range from shared-memory multiprocessors machines to a set of SMP clusters connected together through a high speed network. We also briefly discuss an application of our algorithms to the analysis of large databases collected by a Brazilian Web portal.
Adriano Veloso, Wagner Meira Jr., Srinivasan Parthasarathy 0001
SBAC-PAD2
2003 Extending UML to Specify and Verify E-commerce Systems
Mark A. J. Song, Adriano C. M. Pereira, Gustavo Gorgulho, Sérgio Vale Aguiar Campos, Wagner Meira Jr.
SEKE6
2003 A hierarchical and multiscale approach to analyze E-business workloads
Daniel A. Menascé, Virgílio A. F. Almeida, Rudolf H. Riedi, Flávia Ribeiro, Rodrigo Fonseca, Wagner Meira Jr.
Perform. Evaluation6
2002 A Formal Methodology to Specify E-commerce Systems
Adriano C. M. Pereira, Mark A. J. Song, Gustavo Gorgulho, Wagner Meira Jr., Sérgio Vale Aguiar Campos
ICFEM4
2002 A hierarchical characterization of a live streaming media workload
abstract
Abstract—We present a thorough characterization of what we believe to be the first significant live Internet streaming media workload in the scientific literature. Our characterization of over 3.5 million requests spanning a 28-day period is done at three increasingly granular levels, corresponding to clients, sessions, and transfers. Our findings support two important conclusions. First, we show that the nature of interactions between users and objects is fundamentally different for live versus stored objects. Access to stored objects is user driven, whereas access to live objects is object driven. This reversal of active/passive roles of users and objects leads to interesting dualities. For instance, our analysis underscores a Zipf-like profile for user interest in a given object, which is in contrast to the classic Zipf-like popularity of objects for a given user. Also, our analysis reveals that transfer lengths are highly variable and that this variability is due to client stickiness to a particular live object, as opposed to structural (size) properties of objects. Second, by contrasting two live streaming workloads from two radically different applications, we conjecture that some characteristics of live media access workloads are likely to be highly dependent on the nature of the live content being accessed. This dependence is clear from the strong temporal correlation observed in the traces, which we attribute to the impact of synchronous access to live content. Based on our analysis, we present a model for live media workload generation that incorporates many of our findings, and which we implement in GISMO. Index Terms—Internet, live streaming, measurement, multimedia, workload characterization. I.
Eveline Veloso, Virgílio A. F. Almeida, Wagner Meira Jr., Azer Bestavros, Shudong Jin
Internet Measurement Workshop3
2002 Efficiently Mining Approximate Models of Associations in Evolving Databases
Adriano Veloso, Bruno Gusmão Rocha, Wagner Meira Jr., Márcio de Carvalho, Srinivasan Parthasarathy 0001, Mohammed J. Zaki
PKDD3
2002 Mining Frequent Itemsets in Evolving Databases
abstract
1 Introduction The field of knowledge discovery and data mining (KDD), spurred by advances in data collection technology, is concerned with the process of deriving interesting and useful patterns from large datasets. The KDD process is computational and data-intensive and is inherently interactive and iterative in nature. In fact, interactivity is often the key to facilitating effective data understanding and knowledge discovery. In such an environment, response time is crucial because lengthy time delay between responses of consecutive user requests can disturb the flow of human perception and formation of insight. The task of guaranteeing quick response times is more complicated in dynamic datasets, where there is a constant influx of data. Changes to the data can invalidate existing patterns or introduce new. Simply re-executing algorithms from scratch when a database is updated can result in an explosion in the computational and I/O resources required. What is needed is a way to process the data incrementally and update the information that is gleaned while being cognizant of the interactive requirements of the process. In this paper we present such an approach for a key data mining task: association rule mining.
Adriano Veloso, Wagner Meira Jr., Márcio de Carvalho, Bruno Pôssas, Srinivasan Parthasarathy 0001, Mohammed J. Zaki
SDM2
2002 Set-based model: a new approach for information retrieval
abstract
The objective of this paper is to present a new technique for computing term weights for index terms, which leads to a new ranking mechanism, referred to as set-based model. The components in our model are no longer terms, but termsets. The novelty is that we compute term weights using a data mining technique called association rules, which is time efficient and yet yields nice improvements in retrieval effectiveness. The set-based model function for computing the similarity between a document and a query considers the termset frequency in the document and its scarcity in the document collection. Experimental results show that our model improves the average precision of the answer set for all three collections evaluated. For the TReC-3 collection, our set-based model led to a gain, relative to the standard vector space model, of 37% in average precision curves and of 57% in average precision for the top 10 documents. Like the vector space model, the set-based model has time complexity that is linear in the number of documents in the collection.
Bruno Pôssas, Nivio Ziviani, Wagner Meira Jr., Berthier A. Ribeiro-Neto
SIGIR3
2002 Enhancing the Set-Based Model Using Proximity Information
Bruno Pôssas, Nivio Ziviani, Wagner Meira Jr.
SPIRE3
2001 Resource placement in distributed E-commerce servers
abstract
E-commerce services have become a promising and profitable application of the Internet. In order to keep them growing, solutions must be found to deal with unreliable connections and high latencies, among other problems. The best solutions to such problems tend to depend on the distribution of the service over the network, placing servers in multiple locations, closer to customers. If placement of servers is effective it tends to reduce delays and traffic-related costs. In this paper we discuss the distribution of e-commerce services by introducing a traffic-aware cost model and evaluating it using an actual log from an e-tailer. The results show that the model yields good placement solutions, which perform better than simpler ad-hoc solutions.
Gustavo Machado Campagnani Gama, Wagner Meira Jr., Márcio L. B. Carvalho, Dorgival O. Guedes, Virgílio A. F. Almeida
GLOBECOM2
2001 Rank-Preserving Two-Level Caching for Scalable Search Engines
abstract
Article Rank-preserving two-level caching for scalable search engines Share on Authors: Patricia Correia Saraiva Federal Univ. of Minas Gerais, Belo Horizonte, Brazil and Federal Univ. of Amazonas, Manaus, Brazil Federal Univ. of Minas Gerais, Belo Horizonte, Brazil and Federal Univ. of Amazonas, Manaus, BrazilView Profile , Edleno Silva de Moura Akwan Information Technologies, Belo Horizonte, Brazil Akwan Information Technologies, Belo Horizonte, BrazilView Profile , Nivio Ziviani Federal Univ. of Minas Gerais, Belo Horizonte, Brazil Federal Univ. of Minas Gerais, Belo Horizonte, BrazilView Profile , Wagner Meira Federal Univ. of Minas Gerais, Belo Horizonte, Brazil Federal Univ. of Minas Gerais, Belo Horizonte, BrazilView Profile , Rodrigo Fonseca Univ. of Minas, Belo Horizonte, Brazil Univ. of Minas, Belo Horizonte, BrazilView Profile , Berthier Ribeiro-Neto Federal Univ. of Minas Gerias, Belo Horizonte, Brazil Federal Univ. of Minas Gerias, Belo Horizonte, BrazilView Profile Authors Info & Claims SIGIR '01: Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrievalSeptember 2001 Pages 51–58https://doi.org/10.1145/383952.383959Published:01 September 2001 89citation981DownloadsMetricsTotal Citations89Total Downloads981Last 12 Months8Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Patricia Correia Saraiva, Edleno Silva de Moura, Rodrigo Fonseca, Wagner Meira Jr., Berthier A. Ribeiro-Neto, Nivio Ziviani
SIGIR4
2000 In search of invariants for e-business workloads
abstract
Understanding the nature and characteristics of e-business workloads is a crucial step to improve the quality of service offered to customers in electronic business environments.However, the variety and complexity of the interactions between customers and sites make the characterization of ebusiness workloads a challenging problem.Using a multilayer hierarchical model, this paper presents a detailed characterization of the workload of two actual e-business sites: an online bookstore and an electronic auction site.Through the characterization process, we found the presence of autonomous agents, or robots, in the workload and used the hierarchical structure to determine their characteristics.We also found that search terms follow a Zipf distribution.
Daniel A. Menascé, Virgílio A. F. Almeida, Rudolf H. Riedi, Flávia Ribeiro, Rodrigo Fonseca, Wagner Meira Jr.
EC6
2000 Dynamic Aspects of Documents of the Brazilian Web
abstract
The huge number and the textual nature of Web documents have created the need for search engines. In this kind of system the user's queries are answered based on a static view of the whole or part of the Web. However, the Web is a dynamic system and its documents are inserted, changed and removed frequently, thus creating an inconsistency between the state of the Web and the static view of the documents of the search engine. We describe the dynamic aspects of the HTML documents of the Brazilian Web, namely the rate of insertion, change and removal of its documents, from the point of view of search engines. We also describe the tools we used to support this study. Whenever possible, we compare measures of the Brazilian Web with related measures of the World Wide Web. We show how the study of these aspects can be used in search engines to improve the quality of its services, in particular, aspects related to robots or spiders.
Nahur Fonseca, Rodolfo F. Resende, Wagner Meira Jr.
WISE3
1999 Efficiency Analysis of Brokers in the Electronic Marketplace
Virgílio A. F. Almeida, Wagner Meira Jr., Victor F. Ribeiro, Nivio Ziviani
Comput. Networks2
1998 The Influence of Geographical and Cultural Issues on the Cache Proxy Server Workload
Virgílio A. F. Almeida, Márcio G. Cesário, Rodrigo Fonseca, Wagner Meira Jr., Cristina D. Murta
Comput. Networks4
1997 VM-Based Shared Memory on Low-Latency, Remote-Memory-Access Networks
abstract
Recent technological advances have produced network interfaces that provide users with very low-latency access to the memory of remote machines. We examine the impact of such networks on the implementation and performance of software DSM. Specifically, we compare two DSM systems---Cashmere and TreadMarks---on a 32-processor DEC Alpha cluster connected by a Memory Channel network.Both Cashmere and TreadMarks use virtual memory to maintain coherence on pages, and both use lazy, multi-writer release consistency. The systems differ dramatically, however, in the mechanisms used to track sharing information and to collect and merge concurrent updates to a page, with the result that Cashmere communicates much more frequently, and at a much finer grain.Our principal conclusion is that low-latency networks make DSM based on fine-grain communication competitive with more coarse-grain approaches, but that further hardware improvements will be needed before such systems can provide consistently superior performance. In our experiments, Cashmere scales slightly better than TreadMarks for applications with false sharing. At the same time, it is severely constrained by limitations of the current Memory Channel hardware. In general, performance is better for TreadMarks.
Leonidas I. Kontothanassis, Galen C. Hunt, Robert Stets, Nikos Hardavellas, Michal Cierniak, Srinivasan Parthasarathy 0001, Wagner Meira Jr., Sandhya Dwarkadas, Michael L. Scott
ISCA7