VLDB 2026 Research / reviewers in the wild / expert
Giovanni Comarela
dblp:79/9882 · also Giovanni V. Comarela
· DBLP profile ↗
22ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-7612-9650ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building realistic cybersecurity datasets through a edge cloud testbed with cost-aware monitoringabstractThis paper presents a distributed and autonomous experimental infrastructure designed to support the execution of cybersecurity experiments and the generation of flow and cloud monitoring datasets. The proposed system enables the execution of diverse and reproducible security experiments across geographically separated institutions with distinct physical and logical infrastructures. The infrastructure integrates real applications and networks to emulate both benign and malicious traffic, supporting the generation of flow-based and cloud-level datasets under varied monitoring configurations. Through collaborative deployment at two universities in Brazil, the proposed testbed shows its adaptability and scalability across multiple environments. The experimental results demonstrate that monitoring intervals ranging from 5 to 10 s achieve an effective balance between the detection performance of machine learning models for malicious activities in cloud services and the operational costs associated with network and cloud monitoring, maintaining high classification accuracy across diverse attack types. The generated datasets provide a consistent basis for evaluating monitoring strategies and developing data-driven detection models in cloud-native environments. Willen Borges Coelho, Giovanni Comarela, Rodolfo V. Valentim, Idilio Drago, Rodolfo da Silva Villaça |
Comput. Networks | 2 |
| 2025 | The Value of Complaints: Churn Prediction in a Major Residential Internet Service Provider Using Textual Data
Wadham Bottacin, Vitor F. Zanotelli, Matheus S. De Martin, Pedro de Morais, Rodolfo da Silva Villaça, Vinícius F. S. Mota, Magnos Martinello, Antônio Augusto de Aragão Rocha, Giovanni Comarela |
AINA (3) | 9 |
| 2025 | A Dynamic-Adaptive Architecture for Immersive and Interactive 3D Collaborative Virtual Environments
Diego R. C. Dias, Marcelo de Paiva Guimarães, Giovanni Comarela, Vinícius F. S. Mota, Vitor F. Zanotelli, Luís Carlos Trevelin |
AINA (1) | 3 |
| 2025 | Racist Aesthetics in Digital Extremist Communities
Victoria Ferro, Athus Cavalini, Thamya Donadia, Fabio Goveia, Fabio Malini, Giovanni Comarela |
AINA (2) | 6 |
| 2025 | Strategies for Customer Churn Mitigation with Attendant Recommendation Systems
Adolpho O. dos Santos Filho, Vitor F. Zanotelli, Giovanni Comarela, Antônio Augusto de Aragão Rocha |
AINA (3) | 3 |
| 2025 | Characterization and Prediction of Customer (Dis)satisfaction of a Mobile Internet Provider
Luiza B. Laquini, Vitor F. Zanotelli, Paulo H. L. Rettore, Antônio Augusto de Aragão Rocha, Giovanni Comarela, Vinícius F. S. Mota |
AINA (1) | 5 |
| 2024 | Analysis of Bias in GPT Language Models through Fine-tuning Containing Divergent DataabstractIn this study, we examined the effects of integrating data that contains divergent information, especially concerning anti-vaccination narratives, into the training of a GPT-2 language model. The model was fine-tuned using content sourced from anti-vaccination groups and channels on Telegram, aiming to analyze its ability to generate coherent and rationalized texts in comparison to a model pre-trained on OpenAI’s WebText dataset. The results demonstrate that fine-tuning a GPT-2 model with biased data leads the model to perpetuate these biases in its responses, albeit with a certain degree of rationalization. This finding underscores the importance of using high-quality and reliable data in training natural language processing models, highlighting the implications for information dissemination through these models. It also provides social scientists with a tool to explore and understand the complexities and challenges associated with public health misinformation via the use of language models, particularly in the context of vaccines. Leandro Furlam Turi, Athus Cavalini, Giovanni Comarela, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza |
IJCNN | 3 |
| 2023 | Machine Learning-Based Early Attack Detection Using Open RAN Intelligent ControllerabstractWe design and demonstrate a method for early detection of Denial-of-Service attacks. The proposed approach takes advantage of the OpenRAN framework to collect measurements from the air interface (for attack detection) and to dynamically control the operation of the Radio Access Network (RAN). For that purpose, we developed our near-Real Time (RT) RAN Intelligent Controller (RIC) interface. We apply and analyze a wide range of Machine Learning algorithms to data traffic analysis that satisfy the accuracy and latency requirements set by the near-RT RIC. Our results show that the proposed framework is able to correctly classify genuine vs. malicious traffic with high accuracy (i.e., 95%) in a realistic testbed environment, allowing us to detect attacks already at the Distributed Unit (DU), before malicious traffic even enters the Centralized Unit (CU). Bruno Missi Xavier, Merim Dzaferagic, Diarmuid Collins, Giovanni Comarela, Magnos Martinello, Marco Ruffini |
ICC | 4 |
| 2023 | Dependence and Model Selection in LLP: The Problem of VariantsabstractThe problem of Learning from Label Proportions (LLP) has received considerable research attention and has numerous practical applications. In LLP, a hypothesis assigning labels to items is learned using knowledge of only the proportion of labels found in predefined groups, called bags. While a number of algorithmic approaches to learning in this context have been proposed, very little work has addressed the model selection problem for LLP. Nonetheless, it is not obvious how to extend straightforward model selection approaches to LLP, in part because of the lack of item labels. More fundamentally, we argue that a careful approach to model selection for LLP requires consideration of the dependence structure that exists between bags, items, and labels. In this paper we formalize this structure and show how it affects model selection. We show how this leads to improved methods of model selection that we demonstrate outperform the state of the art over a wide range of datasets and LLP algorithms. Gabriel Franco, Mark Crovella, Giovanni Comarela |
KDD | 3 |
| 2022 | MAP4: A Pragmatic Framework for In-Network Machine Learning Traffic ClassificationabstractSelf-driving networks guided by machine-learning (ML) algorithms are the driving force for building networks of the future. ML is effective at making inferences about data that is too complex or too unpredictable for humans. The network softwarization enabled by a deep programmability approach opens up new opportunities to deploy ML at the programmable data plane. In this paper, we introduce the MAP4 as a framework that explores the feasibility of mapping ML models in programmable network devices. To achieve this, we rely on the P4 language to deploy a pre-trained model into a programmable switch, utilizing the ML model to accurately classify flows at line rate. Our approach demonstrates that ML models working as classifiers can better fit the data by using the new levels of network programmability from the P4 language. The results showed that with few packets, most of the flows are properly classified. In some use cases, with two packets in the flow, 97% of traffic can be correctly classified, and all classes are properly labeled with a maximum of four packets. Bruno Missi Xavier, Rafael S. Guimarães, Giovanni Comarela, Magnos Martinello |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2021 | Tracking Knowledge Propagation Across Wikipedia Languages
Rodolfo V. Valentim, Giovanni Comarela, Souneil Park, Diego Sáez-Trumper |
ICWSM | 2 |
| 2021 | Programmable Switches for in-Networking ClassificationabstractDeploying accurate machine learning algorithms into a high-throughput networking environment is a challenging task. On the one hand, machine learning has proved itself useful for traffic classification in many contexts (e.g., intrusion detection, application classification, and early heavy hitter identification). On the other hand, most of the work in the area is related to post-processing (i.e., training and testing are performed offline on previously collected samples) or to scenarios where the traffic has to leave the data plane to be classified (i.e., high latency). In this work, we tackle the problem of creating simple and reasonably accurate machine learning models that can be deployed into the data plane in a way that performance degradation is acceptable. To that purpose, we introduce a framework and discuss issues related to the translation of simple models, for handling individual packets or flows, into the P4 language. We validate our framework with an intrusion detection use case and by deploying a single decision tree into a Netronome SmartNIC (Agilio CX 2x10GbE). Our results show that high-accuracy is achievable (above 95%) with minor performance degradation, even for a large number of flows. Bruno Missi Xavier, Rafael S. Guimarães, Giovanni Comarela, Magnos Martinello |
INFOCOM | 3 |
| 2020 | A Focused Crawler for Web Feature Service and Web Map Service Discovering
Víctor Macêdo Alexandrino, Giovanni Comarela, Altigran S. da Silva, Jugurta Lisboa Filho |
W2GIS | 2 |
| 2020 | ppiGReMLIN: a graph mining based detection of conserved structural arrangements in protein-protein interfacesabstractBACKGROUND: Protein-protein interactions (PPIs) are fundamental in many biological processes and understanding these interactions is key for a myriad of applications including drug development, peptide design and identification of drug targets. The biological data deluge demands efficient and scalable methods to characterize and understand protein-protein interfaces. In this paper, we present ppiGReMLIN, a graph based strategy to infer interaction patterns in a set of protein-protein complexes. Our method combines an unsupervised learning strategy with frequent subgraph mining in order to detect conserved structural arrangements (patterns) based on the physicochemical properties of atoms on protein interfaces. To assess the ability of ppiGReMLIN to point out relevant conserved substructures on protein-protein interfaces, we compared our results to experimentally determined patterns that are key for protein-protein interactions in 2 datasets of complexes, Serine-protease and BCL-2. RESULTS: ppiGReMLIN was able to detect, in an automatic fashion, conserved structural arrangements that represent highly conserved interactions at the specificity binding pocket of trypsin and trypsin-like proteins from Serine-protease dataset. Also, for the BCL-2 dataset, our method pointed out conserved arrangements that include critical residue interactions within the conserved motif LXXXXD, pivotal to the binding specificity of BH3 domains of pro-apoptotic BCL-2 proteins towards apoptotic suppressors. Quantitatively, ppiGReMLIN was able to find all of the most relevant residues described in literature for our datasets, showing precision of at least 69% up to 100% and recall of 100%. CONCLUSIONS: ppiGReMLIN was able to find highly conserved structures on the interfaces of protein-protein complexes, with minimum support value of 60%, in datasets of similar proteins. We showed that the patterns automatically detected on protein interfaces by our method are in agreement with interaction patterns described in the literature. Felippe C. Queiroz, Adriana M. P. Vargas, Maria G. A. Oliveira, Giovanni Comarela, Sabrina de Azevedo Silveira |
BMC Bioinform. | 4 |
| 2018 | Using supervised learning successful descriptors to perform protein structural classification through unsupervised learning
Cleiton Monteiro, Vinício F. Mendes, Giovanni Comarela, Sabrina de Azevedo Silveira |
BIBM | 3 |
| 2018 | Assessing Candidate Preference through Web Browsing HistoryabstractPredicting election outcomes is of considerable interest to candidates, political scientists, and the public at large. We propose the use of Web browsing history as a new indicator of candidate preference among the electorate, one that has potential to overcome a number of the drawbacks of election polls. However, there are a number of challenges that must be overcome to effectively use Web browsing for assessing candidate preference - including the lack of suitable ground truth data and the heterogeneity of user populations in time and space. We address these challenges, and show that the resulting methods can shed considerable light on the dynamics of voters' candidate preferences in ways that are difficult to achieve using polls. Giovanni Comarela, Ramakrishnan Durairajan, Paul Barford, Dino P. Christenson, Mark Crovella |
KDD | 1 |
| 2016 | Detecting Unusually-Routed ASes: Methods and Applications
Giovanni Comarela, Evimaria Terzi, Mark Crovella |
Internet Measurement Conference | 1 |
| 2014 | Identifying and Analyzing High Impact Routing Events with PathMinerabstractUnderstanding the dynamics of the interdomain routing system is challenging. One reason is that a single routing or policy change can have far reaching and complex effects. Connecting observed behavior with its underlying causes is made even more difficult by the amount of noise in the BGP system. In this paper we address these challenges by presenting PathMiner, a system to extract large scale routing events from background noise and identify the AS or link responsible for the event. Giovanni Comarela, Mark Crovella |
Internet Measurement Conference | 1 |
| 2013 | Studying interdomain routing over long timescalesabstractThe dynamics of interdomain routing have traditionally been studied through the analysis of BGP update traffic. However, such studies tend to focus on the volume of BGP updates rather than their effects, and tend to be local rather than global in scope. Studying the global state of the Internet routing system over time requires the development of new methods, which we do in this paper. We define a new metric, MRSD, that allows us to measure the similarity between two prefixes with respect to the state of the global routing system. Applying this metric over time yields a measure of how the set of total paths to each prefix varies at a given timescale. We implement this analysis method in a MapReduce framework and apply it to a dataset of more than 1TB, collected daily over 3 distinct years and monthly over 8 years. We show that this analysis method can uncover interesting aspects of how Internet routing has changed over time. We show that on any given day, approximately 1% of the next-hop decisions made in the Internet change, and this property has been remarkably constant over time; the corresponding amount of change in one month is 10% and in two years is 50%. Digging deeper, we can decompose next-hop decision changes into two classes: churn, and structural (persistent) change. We show that structural change shows a strong 7-day periodicity and that it represents approximately 2/3 of the total amount of changes. Giovanni Comarela, Gonca Gürsun, Mark Crovella |
Internet Measurement Conference | 1 |
| 2012 | New kid on the block: exploring the google+ social graphabstractThis paper presents a detailed analysis of the Google+ social network. We identify the key differences and similarities with other popular networks like Facebook and Twitter, in order to determine whether Google+ is a new paradigm or yet another social network. This work is based on large-scale crawls of over 27 million user profiles that represented nearly 50% of the entire network in 2011. We observe that the average path length between users is slightly higher than other networks, possibly because Google+ is a new system where relationships are still rapidly growing. Google+ shows a higher level of reciprocity than Twitter, which also has directed social links. The newly available "places lived" field could be used to study how users are distributed around the world and how aggressively the service has been adopted in different countries. We find that Google+ is popular in countries with relatively low Internet penetration rate. Based on the amount and types of information publicly shared in user profiles, we also find that the notion of privacy varies significantly across different cultures. Gabriel Magno, Giovanni Comarela, Diego Sáez-Trumper, Meeyoung Cha, Virgílio A. F. Almeida |
Internet Measurement Conference | 2 |
| 2012 | Finding trendsetters in information networksabstractInfluential people have an important role in the process of information diffusion. However, there are several ways to be influential, for example, to be the most popular or the first that adopts a new idea. In this paper we present a methodology to find trendsetters in information networks according to a specific topic of interest. Trendsetters are people that adopt and spread new ideas influencing other people before these ideas become popular. At the same time, not all early adopters are trendsetters because only few of them have the ability of propagating their ideas by their social contacts through word-of-mouth. Differently from other influence measures, a trendsetter is not necessarily popular or famous, but the one whose ideas spread over the graph successfully. Other metrics such as node in-degree or even standard Pagerank focus only in the static topology of the network. We propose a ranking strategy that focuses on the ability of some users to push new ideas that will be successful in the future. To that end, we combine temporal attributes of nodes and edges of the network with a Pagerank based algorithm to find the trendsetters for a given topic. To test our algorithm we conduct innovative experiments over a large Twitter dataset. We show that nodes with high in-degree tend to arrive late for new trends, while users in the top of our ranking tend to be early adopters that also influence their social contacts to adopt the new trend. Diego Sáez-Trumper, Giovanni Comarela, Virgílio A. F. Almeida, Ricardo Baeza-Yates, Fabrício Benevenuto |
KDD | 2 |
| 2011 | Modeling the Performance of the Hadoop Online PrototypeabstractMapReduce is an important paradigm to support modern data-intensive applications. In this paper we address the challenge of modeling performance of one implementation of MapReduce called Hadoop Online Prototype (HOP), with a specific target on the intra-job pipeline parallelism. We use a hierarchical model that combines a precedence model and a queuing network model to capture the intra-job synchronization constraints. We first show how to build a precedence graph that represents the dependencies among multiple tasks of the same job. We then apply it jointly with an approximate Mean Value Analysis (aMVA) solution to predict mean job response time and resource utilization. We validate our solution against a queuing network simulator in various scenarios, finding that our performance model presents a close agreement, with maximum relative difference under 15%. Emanuel Vianna, Giovanni Comarela, Tatiana Pontes, Jussara M. Almeida, Virgílio A. F. Almeida, Kevin Wilkinson, Harumi A. Kuno, Umeshwar Dayal |
SBAC-PAD | 2 |