VLDB 2026 Research / reviewers in the wild / expert
Rodolfo V. Valentim
dblp:237/7754 · also Rodolfo Vieira Valentim
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-7702-2991ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building realistic cybersecurity datasets through a edge cloud testbed with cost-aware monitoringabstractThis paper presents a distributed and autonomous experimental infrastructure designed to support the execution of cybersecurity experiments and the generation of flow and cloud monitoring datasets. The proposed system enables the execution of diverse and reproducible security experiments across geographically separated institutions with distinct physical and logical infrastructures. The infrastructure integrates real applications and networks to emulate both benign and malicious traffic, supporting the generation of flow-based and cloud-level datasets under varied monitoring configurations. Through collaborative deployment at two universities in Brazil, the proposed testbed shows its adaptability and scalability across multiple environments. The experimental results demonstrate that monitoring intervals ranging from 5 to 10 s achieve an effective balance between the detection performance of machine learning models for malicious activities in cloud services and the operational costs associated with network and cloud monitoring, maintaining high classification accuracy across diverse attack types. The generated datasets provide a consistent basis for evaluating monitoring strategies and developing data-driven detection models in cloud-native environments. Willen Borges Coelho, Giovanni Comarela, Rodolfo V. Valentim, Idilio Drago, Rodolfo da Silva Villaça |
Comput. Networks | 3 |
| 2026 | FedScope - Federated Host Embeddings From Telescope Traffic: Design and Implementationabstractnetwork telescope is a range of IP addresses that host no services. Millions of bots and scanners contact it to look for vulnerable systems, and the traffic it exposes is fundamental to understanding malicious activities. The visibility a telescope offers depends on its size and geolocation, and merging the information from multiple telescopes could help increase visibility and uncover more malicious activities. However, sharing raw telescope data is complicated, calling for solutions that allow one to directly share the knowledge rather than the data obtained from multiple deployments. In this paper, we explore the application of Federated Learning (FL) to create and share such global knowledge from the malicious activities seen in distributed telescopes. For that, we introduce FedScope, an FL-based solution for generatinghost embeddingsin a distributed way. We compare FedScope to local and distributed alternatives in downstream tasks, such as sender classification or coordinated activities detection. We show that FedScope (i) produces embeddings of equal or higher quality than those of a single telescope; (ii) increases coverage, allowing the global model to monitor more malicious actors; (iii) avoids the sharing of the raw data, limiting exchanged data. Andrea Sordello, Rodolfo V. Valentim, Luca Vassio, Idilio Drago, Marco Mellia |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | LogPrécis: Unleashing language models for automated malicious log analysisabstractSecurity logs are the key to understanding attacks and diagnosing vulnerabilities. Often coming in the form of text logs, their analysis remains a daunting challenge. Language Models (LMs) have demonstrated unmatched potential in understanding natural and programming languages. The question arises as to whether and how LMs could be also used to automatise the analysis of security logs. We here systematically study how to benefit from the state-of-the-art LM to support the analysis of text-like Unix shell attack logs automatically. For this, we thoroughly designed LogPrécis. LogPrécis receives as input malicious shell sessions. It then automatically identifies and assigns the attacker tactic to each portion of the session, i.e., unveiling the sequence of the attacker's goals. This creates a unique attack fingerprint. We demonstrate LogPrécis capability to support the analysis of two large datasets containing about 400,000 unique Unix shell attacks recorded in a 2-year-long honeypot deployment. LogPrécis reduces the analysis to about 3,000 unique fingerprints. Such abstraction lets us better understand attacks, extract attack prototypes, detect novelties, and track families and mutations. Overall, LogPrécis, released as open source, demonstrates the potential of adopting LMs for security analysis and paves the way for better and more responsive defence against cyberattacks. Matteo Boffa, Idilio Drago, Marco Mellia, Luca Vassio, Danilo Giordano, Rodolfo V. Valentim, Zied Ben-Houidi |
Comput. Secur. | 6 |
| 2024 | X-squatter: AI Multilingual Generation of Cross-Language Sound-squattingabstractSound-squatting is a squatting technique that exploits similarities in word pronunciation to trick users into accessing malicious resources. It is an understudied threat that has gained traction with the popularity of smart speakers and audio-only content, such as podcasts. The picture gets even more complex when multiple languages are involved. We here introduce X-squatter, a multi- and cross-language AI-based system that relies on a Transformer Neural Network for generating high-quality sound-squatting candidates. We illustrate the use of X-squatter by searching for domain name squatting abuse across hundreds of millions of issued TLS certificates, alongside other squatting types. Key findings unveil that approximately 15% of generated sound-squatting candidates have associated TLS certificates, well above the prevalence of other squatting types (7%). Furthermore, we employ X-squatter to assess the potential for abuse in PyPI packages, revealing the existence of hundreds of candidates within a 3-year package history. Notably, our results suggest that the current platform checks cannot handle sound-squatting attacks, calling for better countermeasures. We believe X-squatter uncovers the usage of multilingual sound-squatting phenomena on the Internet and it is a crucial asset for proactive protection against the threat. Rodolfo V. Valentim, Idilio Drago, Marco Mellia, Federico Cerutti 0001 |
ACM Trans. Priv. Secur. | 1 |
| 2023 | URLGEN - Toward Automatic URL Generation Using GANsabstractURLs play an essential role on the Internet, allowing access to Web resources. Automatically generating URLs is helpful in various tasks, such as application debugging, API testing, and blocklist creation for security applications. Current testing suites deeply embed experts’ domain knowledge to generate suitable URLs, resulting in an ad-hoc solution for each given application. These tools thus require heavy manual intervention, with the expensive coding of rules that are hard to maintain. We here introduce URLGEN, a system that uses Generative Adversarial Networks (GANs) to tackle the automatic URL generation problem. URLGEN is designed for web API testing and generates URL samples for an application without any system expertise, complementing the existing tools. It leverages Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN) architectures, augmented by an embedding layer that simplifies the URL learning and generation process. We show that URLGEN learns to generate new valid URLs from samples of real URLs without requiring any domain knowledge and following a purely data-driven approach. We compare the GAN architecture of URLGEN against other design options and show that the LSTM architecture can better capture the correlation among URL characters, outperforming previously proposed solutions. Finally, we show that the URLGEN approach can be extended to other scenarios, which we illustrate with two use cases, i.e., cybersquatting domain prediction and URL classification. Rodolfo V. Valentim, Idilio Drago, Martino Trevisan, Marco Mellia |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Tracking Knowledge Propagation Across Wikipedia Languages
Rodolfo V. Valentim, Giovanni Comarela, Souneil Park, Diego Sáez-Trumper |
ICWSM | 1 |
| 2020 | KeySFC: Traffic steering using strict source routing for dynamic and efficient network orchestration
Cristina K. Dominicini, Gilmar L. Vassoler, Rodolfo V. Valentim, Rodolfo da Silva Villaça, Moisés R. N. Ribeiro, Magnos Martinello, Eduardo Zambon |
Comput. Networks | 3 |