VLDB 2026 Research / reviewers in the wild / expert
Daniel Perdices
dblp:214/6167
· DBLP profile ↗
11ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-3421-7633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Internet usage and performance in GEO satellite networks: A large-scale study across Europe and AfricaabstractSatellite Communication (SatCom) offers internet connectivity where traditional infrastructures are too expensive to deploy. When using satellites in a geostationary orbit, the distance from Earth forces a round-trip time of at least 550 ms. Coupled with the constrained capacity of the physical link, this challenges the traditional internet access quality we are used to. In this paper, we present a complete passive characterization of the traffic carried by an operational SatCom provider. With this unique vantage point, we observe the performance of the SatCom technology, as well as the usage habits of subscribers in different countries in Europe and Africa. We highlight the implications of such technology on Internet usage and functioning, and we pinpoint technical challenges due to the CDN and DNS resolution issues, while discussing possible optimizations that the ISP could implement to improve the service offered to SatCom subscribers. We complete the characterization of the adoption and performance of newer protocols with a focus on IPv6 and QUIC. Gabriele Merlach, Daniel Perdices, Gianluca Perna, Martino Trevisan, Danilo Giordano, Marco Mellia |
Comput. Networks | 2 |
| 2025 | GPT on the wire: Towards realistic network traffic conversations generated with large language modelsabstractRealistic network traffic generation is essential for evaluating the performance, security, and scalability of modern communication systems. Traditional methods, such as traffic replay systems and statistical models, while useful, often fall short in capturing the complexity and variability of real-world network scenarios. Recent advancements in Artificial Intelligence (AI), especially Large Language Models (LLMs) like ChatGPT, have introduced new approaches to synthetic traffic generation. This paper presents a novel architecture using OpenAI’s GPT-3.5 Turbo to generate synthetic network traffic, with a focus on creating multi-protocol conversations that are indistinguishable from real-world interactions. Through fine-tuning and prompt engineering, the proposed system successfully generates packet- and conversation-level network traffic for ICMP, ARP, DNS, TCP and HTTP protocols. Additionally, by integrating a Mixture of Experts (MoE) architecture, this model simulates real-world network conversations with high accuracy, being able to generate a conversation combining ARP, DNS, TCP and HTTP without packet or protocol errors. The results show how the application of LLMs in network traffic generation improves realism and adaptability, establishing this approach as a valuable tool for future security testing and network performance evaluation. In addition, the proposed methodology is easily adaptable to other LLMs available both through APIs and to be downloaded and executed on your own computer. Javier Aday Delgado-Soto, Jorge E. López de Vergara, Iván González 0004, Daniel Perdices, Luis de Pedro |
Comput. Networks | 4 |
| 2025 | An expert-aware Markovian system for end-user proactive troubleshooting in the Network and Security Operations CenterabstractCompanies’ Network Operations Centers continuously monitor network health to keep activity fully operational in the current scenario of decentralization of digital workplaces. In this task, network managers have a diverse set of tools to proactively troubleshoot network changes that could potentially lead to network outages. Unfortunately, these tools primarily focus on servers and high-volume services, while changes in end-users’ traffic can also be a symptom of relevant issues such as the misuse and misconfiguration of resources, service slowdowns, and cybersecurity breaches. End-user behavior, in particular, tends to be more erratic and chaotic than server traffic, especially when different users may share the same IP address over time (e.g., DHCP and Wi-Fi environments). To address these challenges, we propose modeling end-user behavior as an unordered collection of activities (i.e., pieces of regular behavior) rather than as time series. These activities, such as downloading a file or browsing the Internet, yield an identifiable sequence of measurements in terms of network metrics over time (e.g., throughput or number of connections). Deviations from typical sequences of measurements are flagged as irregular behaviors, with higher priority assigned to increasingly uncommon patterns. This methodology leverages Markov chains to represent activities as sequences of expert-aware discrete states of network metrics. When applied to a multi-year dataset from a global enterprise, our proposal has been able to cope with the chaotic and challenging environment of end-user behavioral modeling, identifying critical issues such as DNS misconfigurations, compromised printers, and cybersecurity threats. • Users’ behavior is chaotic compared to more predictable servers’ profiles. • Expert system is proposed to identify end-user anomalies. • The system is based on an expert-tailored Markovian approach. • System has been evaluated on a years-long dataset from international company. • This identified DNS misuse, compromised printers, and cybersecurity threats. Alejandro León 0003, Daniel Perdices, José Luis García-Dorado, Javier Ramos 0002, Javier Aracil 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Monitoring Web QoE in Satellite Networks from Passive MeasurementsabstractSatellite Communication (SatCom) is the only choice to access the Internet in remote regions and is characterized by extreme latency and constrained capacity. For SatCom operators, it is thus fundamental to monitor the Quality of Experience (QoE) of subscribers, to measure their satisfaction, spot anomalies and optimize the peculiar network setup. The Web has become the primary source of Internet content, and Web browsing is the main activity of internauts. This paper addresses the challenge of monitoring Web QoE in SatCom environments, proposing a tailored system that employs a supervised approach to predict Web QoE using passive measurements. The system collects training data through Test Agents that mimic real subscribers' traffic patterns and uses them to build Machine Learning (ML) models that predict performance metrics. The findings demonstrate the feasibility of monitoring Web QoE in SatCom environments, with limitations on website applicability and temporal stability. The need for periodic data generation and the development of a general machine learning model for unseen websites remain open challenges. This research contributes to enhancing web browsing experiences in SatCom and expanding understanding of Web QoE monitoring in diverse network settings. Gianluca Perna, Martino Trevisan, Danilo Giordano, Daniel Perdices, Marco Mellia |
CCNC | 4 |
| 2023 | Web browsing privacy in the deep learning era: Beyond VPNs and encryptionabstractWeb browsing privacy is a matter of paramount importance for the Internet users. While they try to protect themselves from being monitored by getting advantage of encryption or VPNs, users’ privacy is still unaccomplished, even taking into account the tangled web, with several domains visited at the same time in a single web page, or IP addresses of a cloud provider shared by several sites. In this work, we provide a novel approach to identify user web browsing that only takes into account the IP addresses that the user has connected to and without performing any DNS reverse resolutions. We use this sequence of addresses as an input of different state-of-the-art deep learning models, such as multi-layer perceptron and transformers, which are able to accurately identify which was the website actually visited among Alexa’s World Top 500 most visited domains. Moreover, we have also studied other factors, such as the dependence on the DNS server used to resolve the visited IP addresses, the accuracy for the top domains (e.g., Google, YouTube, Facebook, etc.), data augmentation by packet sampling simulation to improve our results, the impact on packet sampling and the fine-tuning and possible impact of model parameters or the scalability of our approach. We conclude that, using only a 10% of the packets, we can identify the visited website with an accuracy and F1 score between 94% and 95%. Daniel Perdices, Jorge E. López de Vergara, Iván González 0004, Luis de Pedro |
Comput. Networks | 1 |
| 2022 | When satellite is all you have: watching the internet from 550 msabstractSatellite Communication (SatCom) offers internet connectivity where traditional infrastructures are too expensive to deploy. When using satellites in a geostationary orbit, the distance from Earth forces a round trip time higher than 550 ms. Coupled with the limited and shared capacity of the physical link, this poses a challenge to the traditional internet access quality we are used to. Daniel Perdices, Gianluca Perna, Martino Trevisan, Danilo Giordano, Marco Mellia |
IMC | 1 |
| 2021 | Assessing the Limits of Privacy and Data Usage for Web Browsing AnalyticsabstractWeb browsing analytics provides insights on the websites that users access, which affects their privacy. Although this analysis might be seen as an easy task, different problems, such as encryption, the tangled web, with several domains visited at the same time in a single web page, or IP addresses of a cloud provider shared by several sites, make it a though job. However, despite these issues, users' privacy is still unaccomplished, as we show in this work. We provide a novel approach that only takes into account the IP addresses that the user has connected to without performing any reverse DNS lookup. We use this sequence of addresses as an input of a neural network, which is able to identify accurately which was the website actually visited among Alexa's World Top 500 most visited domains. Moreover, we have also studied other factors, such as the dependence on the DNS server used to resolve the visited IP addresses, the accuracy for the top domains (e.g., Google, YouTube, Facebook, etc.), data augmentation to improve our results, or the impact on packet sampling. In this last case, we conclude that, using only a 10% of the packets, we can identify the visited website with an accuracy of 93%, whereas it can be over 97% if there is no packet sampling and we use data augmentation. Daniel Perdices, Jorge E. López de Vergara, Iván González 0004 |
CNSM | 1 |
| 2021 | Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunitiesabstractIn an Internet arena where the search engines and other digital marketing firms’ revenues peak, other actors still have open opportunities to monetize their users’ data. After the convenient anonymization, aggregation, and agreement, the set of websites users visit may result in exploitable data for ISPs. Uses cover from assessing the scope of advertising campaigns to reinforcing user fidelity among other marketing approaches, as well as security issues. However, sniffers based on HTTP, DNS, TLS or flow features do not suffice for this task. Modern websites are designed for preloading and prefetching some contents in addition to embedding banners, social networks’ links, images, and scripts from other websites. This self-triggered traffic makes it confusing to assess which websites users visited on purpose. Moreover, DNS caches prevent some queries of actively visited websites to be even sent. On this limited input, we propose to handle such domains as words and the sequences of domains as documents. This way, it is possible to identify the visited websites by translating this problem to a text classification context and applying the most promising techniques of the natural language processing and neural networks fields. After applying different representation methods such as TF–IDF, Word2vec, Doc2vec, and custom neural networks in diverse scenarios and with several datasets, we can state websites visited on purpose with accuracy figures over 90%, with peaks close to 100%, being processes that are fully automated and free of any human parametrization. Daniel Perdices, Javier Ramos 0002, José Luis García-Dorado, Iván González 0004, Jorge E. López de Vergara |
Comput. Networks | 1 |
| 2021 | Deep-FDA: Using Functional Data Analysis and Neural Networks to Characterize Network Services Time SeriesabstractIn network management, it is important to model baselines, trends, and regular behaviors to adequately deliver network services. However, their characterization is complex, so network operation and system alarming become a challenge. Several problems exist: Gaussian assumptions cannot be made, time series have different trends, and it is difficult to reduce their dimensionality. To overcome this situation, we propose Deep-FDA, a novel approach for network service modeling that combines functional data analysis (FDA) and neural networks. Specifically, we explore the use of functional clustering and functional depth measurements to characterize network services with time series generated from enriched flow records, showing how this method can detect different separated trends. Moreover, we augment this statistical approach with the use of autoencoder neural networks, improving the classification results. To evaluate and check the applicability of our proposal, we performed experiments with synthetic and real-world data, where we show graphically and numerically the performance of our method compared to other state-of-the-art alternatives. We also exemplify its application in different network management use cases. The results show that FDA and neural networks are complementary, as they can help each other to improve the drawbacks that both analysis methods have when are applied separately. Daniel Perdices, Jorge E. López de Vergara, Javier Ramos 0002 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2019 | On the Modeling of Multi-Point RTT Passive Measurements for Network Delay MonitoringabstractMany network management actions need a simultaneous consideration of several elements' state. This is becoming an even more complex matter with the advent of reconfigurable deployments, where scaling functions up can prevent performance bottlenecks. Therefore, fine-grained detection of significant burdens arises as a cornerstone to optimize their monitoring and operation. We present advanced distributed passive retrieval of information, and statistical multi-point analysis (AdPRISMA), a passive monitoring system intended to fit models for network delay measurements with clustering elements to improve representation of central and extreme behaviors. As distinguishing features, it relies on cost-effective multi-point round-trip time (RTT) passive network measurements, and is able to select a suitable parametric model optimizing the trade-off between fitting and complexity. AdPRISMA can correlate records collected from several vantage points and detect where performance issues are most likely to appear; adjust alarms in terms of the probability of events; and adapt its behavior to dynamic network conditions while presenting a fair identification of anomalous situations. We evaluate AdPRISMA with experiments both in virtual environments and with real-world data to provide evidences of its applicability and capabilities to represent network elements' delay. Daniel Perdices, David Muelas, Iria Prieto, Luis de Pedro, Jorge E. López de Vergara |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2018 | Network Performance Monitoring with Flexible Models of Multi-Point Passive Measurements
Daniel Perdices, David Muelas, Luis de Pedro, Jorge E. López de Vergara |
CNSM | 1 |