VLDB 2026 Research / reviewers in the wild / expert
Luca Vassio
dblp:136/5733
· DBLP profile ↗
32ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-2920-1856ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021Computer networks · 9 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedScope - Federated Host Embeddings From Telescope Traffic: Design and Implementationabstractnetwork telescope is a range of IP addresses that host no services. Millions of bots and scanners contact it to look for vulnerable systems, and the traffic it exposes is fundamental to understanding malicious activities. The visibility a telescope offers depends on its size and geolocation, and merging the information from multiple telescopes could help increase visibility and uncover more malicious activities. However, sharing raw telescope data is complicated, calling for solutions that allow one to directly share the knowledge rather than the data obtained from multiple deployments. In this paper, we explore the application of Federated Learning (FL) to create and share such global knowledge from the malicious activities seen in distributed telescopes. For that, we introduce FedScope, an FL-based solution for generatinghost embeddingsin a distributed way. We compare FedScope to local and distributed alternatives in downstream tasks, such as sender classification or coordinated activities detection. We show that FedScope (i) produces embeddings of equal or higher quality than those of a single telescope; (ii) increases coverage, allowing the global model to monitor more malicious actors; (iii) avoids the sharing of the raw data, limiting exchanged data. Andrea Sordello, Rodolfo V. Valentim, Luca Vassio, Idilio Drago, Marco Mellia |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2025 | Dominance or Fair Play in Social Networks? A Model of Influencer Popularity Dynamics
Franco Galante, Chiara Ravazzi, Luca Vassio, Michele Garetto, Emilio Leonardi |
ASONAM (2) | 3 |
| 2025 | Contagious Rhythms: A Wave-Based Epidemic Approach for Music Virality on Social Platforms
Gabriel P. Oliveira, Luca Vassio, Ana Paula Couto da Silva, Mirella M. Moro |
ASONAM (1) | 2 |
| 2025 | Join the Chat: How Curiosity Sparks Participation in Telegram GroupsabstractThis study delves into the mechanisms that spark user curiosity driving active engagement within public Telegram groups. By analyzing approximately 6 million messages from 29,196 users across 409 groups, we identify and quantify the key factors that stimulate users to actively participate (i.e., send messages) in group discussions. These factors include social influence, novelty, complexity, uncertainty, and conflict, all measured through metrics derived from message sequences and user participation over time. After clustering the messages, we apply explainability techniques to assign meaningful labels to the clusters. This approach uncovers macro categories representing distinct curiosity stimulation profiles, each characterized by a unique combination of various stimuli. Social influence from peers and influencers drives engagement for some users, while for others, rare media types or a diverse range of senders and media sparks curiosity. Analyzing patterns, we found that user curiosity stimuli are mostly stable, but, as the time between the initial message increases, curiosity occasionally shifts. A graph-based analysis of influence networks reveals that users motivated by direct social influence tend to occupy more peripheral positions, while those who are not stimulated by any specific factors are often more central, potentially acting as initiators and conversation catalysts. These findings contribute to understanding information dissemination and spread processes on social media networks, potentially contributing to more effective communication strategies. Giordano Paoletti, Jussara M. Almeida, Luca Vassio, Marcos André Gonçalves, Marco Mellia |
ICWSM | 3 |
| 2025 | The Sweet Danger of Sugar: Debunking Representation Learning for Encrypted Traffic ClassificationabstractRecently we have witnessed the explosion of proposals that, inspired by Language Models like BERT, exploit Representation Learning models to create traffic representations. All of them promise astonishing performance in encrypted traffic classification (up to 98% accuracy). In this paper, with a networking expert mindset, we critically reassess their performance. Through extensive analysis, we demonstrate that the reported successes are heavily influenced by data preparation problems, which allow these models to find easy shortcuts - spurious correlation between features and labels - during fine-tuning that unrealistically boost their performance. When such shortcuts are not present - as in real scenarios - these models perform poorly. We also introduce Pcap-Encoder, an LM-based representation learning model that we specifically design to extract features from protocol headers. Pcap-Encoder appears to be the only model that provides an instrumental representation for traffic classification. Yet, its complexity questions its applicability in practical settings. Our findings reveal flaws in dataset preparation and model training, calling for a better and more conscious test design. We propose a correct evaluation methodology and stress the need for rigorous benchmarking. Giovanni Dettori, Matteo Boffa, Luca Vassio, Marco Mellia |
SIGCOMM | 4 |
| 2025 | CoDÆN: Benchmarks and Comparison of Evolutionary Community Detection Algorithms for Dynamic NetworksabstractWeb data are often modelled as complex networks in which entities interact and form communities. Nevertheless, web data evolves over time, and network communities change alongside it. This makes Community Detection (CD) in dynamic graphs a relevant problem, calling for evolutionary CD algorithms. The choice and evaluation of such algorithm performance is challenging because of the lack of a comprehensive set of benchmarks and specific metrics. To address these challenges, we propose CoDÆN—Community Detection Algorithms in Evolving Networks—a benchmarking framework for evolutionary CD algorithms in dynamic networks, that we offer as open source to the community. CoDÆN allows us to generate synthetic community-structured graphs with known ground truth and design evolving scenarios combining nine basic graph transformations that modify edges, nodes, and communities. We propose three complementary metrics (i.e., Correctness, Delay, and Stability) to compare evolutionary CD algorithms. Armed with CoDÆN, we consider three evolutionary modularity-based CD approaches, dissecting their performance to gauge the trade-off between the stability of the communities and their correctness. Next, we compare the algorithms in real Web-oriented datasets, confirming such a trade-off. Our findings reveal that algorithms that introduce memory in the graph maximise stability but add delay when abrupt changes occur. Conversely, algorithms that introduce memory by initialising the CD algorithms with the previous solution fail to identify the split and birth of new communities. These observations underscore the value of CoDÆN in facilitating the study and comparison of alternative evolutionary community detection algorithms. Giordano Paoletti, Luca Gioacchini, Marco Mellia, Luca Vassio, Jussara M. Almeida |
ACM Trans. Web | 4 |
| 2024 | LogPrécis: Unleashing language models for automated malicious log analysisabstractSecurity logs are the key to understanding attacks and diagnosing vulnerabilities. Often coming in the form of text logs, their analysis remains a daunting challenge. Language Models (LMs) have demonstrated unmatched potential in understanding natural and programming languages. The question arises as to whether and how LMs could be also used to automatise the analysis of security logs. We here systematically study how to benefit from the state-of-the-art LM to support the analysis of text-like Unix shell attack logs automatically. For this, we thoroughly designed LogPrécis. LogPrécis receives as input malicious shell sessions. It then automatically identifies and assigns the attacker tactic to each portion of the session, i.e., unveiling the sequence of the attacker's goals. This creates a unique attack fingerprint. We demonstrate LogPrécis capability to support the analysis of two large datasets containing about 400,000 unique Unix shell attacks recorded in a 2-year-long honeypot deployment. LogPrécis reduces the analysis to about 3,000 unique fingerprints. Such abstraction lets us better understand attacks, extract attack prototypes, detect novelties, and track families and mutations. Overall, LogPrécis, released as open source, demonstrates the potential of adopting LMs for security analysis and paves the way for better and more responsive defence against cyberattacks. Matteo Boffa, Idilio Drago, Marco Mellia, Luca Vassio, Danilo Giordano, Rodolfo V. Valentim, Zied Ben-Houidi |
Comput. Secur. | 4 |
| 2024 | TASP: Topic-based abstractive summarization of Facebook text posts
Irene Benedetto, Moreno La Quatra, Luca Cagliero, Luca Vassio, Martino Trevisan |
Expert Syst. Appl. | 4 |
| 2024 | Cross-Network Embeddings Transfer for Traffic AnalysisabstractArtificial Intelligence (AI) approaches have emerged as powerful tools to improve traffic analysis for network monitoring and management. However, the lack of large labeled datasets and the ever-changing networking scenarios make a fundamental difference compared to other domains where AI is thriving. We believe the ability to transfer the specific knowledge acquired in one network (or dataset) to a different network (or dataset) would be fundamental to speed up the adoption of AI-based solutions for traffic analysis and other networking applications (e.g., cybersecurity). We here propose and evaluate different options to transfer the knowledge built from a provider network, owning data and labels, to a customer network that desires to label its traffic but lacks labels. We formulate this problem as a domain adaptation problem that we solve with embedding alignment techniques and canonical transfer learning approaches. We present a thorough experimental analysis to assess the performance considering both supervised (e.g., classification) and unsupervised (e.g., novelty detection) downstream tasks related to darknet and honeypot traffic. Our experiments show the proper transfer techniques to use the models obtained from a network in a different network. We believe our contribution opens new opportunities and business models where network providers can successfully share their knowledge and AI models with customers. Luca Gioacchini, Marco Mellia, Luca Vassio, Idilio Drago, Giulia Milan, Zied Ben-Houidi, Dario Rossi 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2023 | GCNH: A Simple Method For Representation Learning On Heterophilous GraphsabstractGraph Neural Networks (GNNs) are well-suited for learning on homophilous graphs, i.e., graphs in which edges tend to connect nodes of the same type. Yet, achievement of consistent GNN performance on heterophilous graphs remains an open research problem. Recent works have proposed extensions to standard GNN architectures to improve performance on heterophilous graphs, trading off model simplicity for prediction accuracy. However, these models fail to capture basic graph properties, such as neighborhood label distribution, which are fundamental for learning. In this work, we propose GCN for Heterophily (GCNH), a simple yet effective GNN architecture applicable to both heterophilous and homophilous scenarios. GCNH learns and combines separate representations for a node and its neighbors, using one learned importance coefficient per layer to balance the contributions of center nodes and neighborhoods. We conduct extensive experiments on eight real-world graphs and a set of synthetic graphs with varying degrees of heterophily to demonstrate how the design choices for GCNH lead to a sizable improvement over a vanilla GCN. Moreover, GCNH outperforms state-of-the-art models of much higher complexity on four out of eight benchmarks, while producing comparable results on the remaining datasets. Finally, we discuss and analyze the lower complexity of GCNH, which results in fewer trainable parameters and faster training times than other methods, and show how GCNH mitigates the oversmoothing problem. Andrea Cavallo, Claas Grohnfeldt, Michele Russo, Giulio Lovisotto, Luca Vassio |
IJCNN | 5 |
| 2023 | Data driven scalability and profitability analysis in free floating electric car sharing systems
Alessandro Ciociola, Danilo Giordano, Luca Vassio, Marco Mellia |
Inf. Sci. | 3 |
| 2023 | Practical anonymization for data streams: z-anonymity and relation with k-anonymity
Nikhil Jha, Luca Vassio, Martino Trevisan, Emilio Leonardi, Marco Mellia |
Perform. Evaluation | 2 |
| 2023 | i-DarkVec: Incremental Embeddings for Darknet Traffic AnalysisabstractDarknets are probes listening to traffic reaching IP addresses that host no services. Traffic reaching a darknet results from the actions of internet scanners, botnets, and possibly misconfigured hosts. Such peculiar nature of the darknet traffic makes darknets a valuable instrument to discover malicious online activities, e.g., identifying coordinated actions performed by bots or scanners. However, the massive amount of packets and sources that darknets observe makes it hard to extract meaningful insights, calling for scalable tools to automatically identify and group sources that share similar behaviour. We here present i-DarkVec, a methodology to learn meaningful representations of Darknet traffic. i-DarkVec leverages Natural Language Processing techniques (e.g., Word2Vec) to capture the co-occurrence patterns that emerge when scanners or bots launch coordinated actions. As in NLP problems, the embeddings learned with i-DarkVec enable several new machine learning tasks on the darknet traffic, such as identifying clusters of senders engaged in similar activities. We extensively test i-DarkVec and explore its design space in a case study using real darknets. We show that with a proper definition of services , the learned embeddings can be used to (i) solve the classification problem to associate unknown sources’ IP addresses to the correct classes of coordinated actors and (ii) automatically identify clusters of previously unknown sources performing similar attacks and scans, easing the security analyst’s job. i-DarkVec leverages a novel incremental embedding learning approach that is scalable and robust to traffic changes, making it applicable to dynamic and large-scale scenarios. Luca Gioacchini, Luca Vassio, Marco Mellia, Idilio Drago, Zied Ben-Houidi, Dario Rossi 0001 |
ACM Trans. Internet Techn. | 2 |
| 2022 | The Internet with Privacy Policies: Measuring The Web Upon ConsentabstractTo protect user privacy, legislators have regulated the use of tracking technologies, mandating the acquisition of users’ consent before collecting data. As a result, websites started showing more and more consent management modules–i.e., Consent Banners–the visitors have to interact with to access the website content. Since these banners change the content the browser loads, they challenge web measurement collection, primarily to monitor the extent of tracking technologies, but also to measure web performance. If not correctly handled, Consent Banners prevent crawlers from observing the actual content of the websites. In this paper, we present a comprehensive measurement campaign focusing on popular websites in Europe and the US, visiting both landing and internal pages from different countries around the world. We engineer Priv-Accept , a Web crawler able to accept the Consent Banners, as most users would do in practice. It lets us compare how webpages change before and after accepting such policies, if present. Our results show that all measurements performed ignoring the Consent Banners offer a biased and partial view of the Web. After accepting the privacy policies, web tracking is far more pervasive, and webpages are larger and slower to load. Nikhil Jha, Martino Trevisan, Luca Vassio, Marco Mellia |
ACM Trans. Web | 3 |
| 2021 | The stock exchange of influencers: a financial approach for studying fanbase variation trendsabstractIn many online social networks (OSNs), a limited portion of profiles emerges and reaches a large base of followers, i.e., the so-called social influencers. One of their main goals is to increase their fanbase to increase their visibility, engaging users through their content. In this work, we propose a novel parallel between the ecosystem of OSNs and the stock exchange market. Followers act as private investors, and they follow influencers, i.e., buy stocks, based on their individual preferences and on the information they gather through external sources. In this preliminary study, we show how the approaches proposed in the context of the stock exchange market can be successfully applied to social networks. Our case study focuses on 60 Italian Instagram influencers and shows how their followers short-term trends obtained through Bollinger bands become close to those found in external sources, Google Trends in our case, similarly to phenomena already observed in the financial market. Besides providing a strong correlation between these different trends, our results pose the basis for studying social networks with a new lens, linking them with a different domain. Fabio Bertone, Luca Vassio, Martino Trevisan |
ASONAM | 2 |
| 2021 | Temporal dynamics of posts and user engagement of influencers on Facebook and InstagramabstractA relevant fraction of human interactions occurs on online social networks. Freshness of content seems to play an important role, with content popularity rapidly vanishing over time. In this paper, we investigate how influencers' generated content (i.e., posts) attracts interactions, measured by number of likes or reactions. We analyse the activity of Italian influencers and followers over more than 5 years, focusing on two popular social networks: Facebook and Instagram, including more than 13 billion interactions and about 4 million posts. We characterise the influencers' and followers' behaviour over time, show that influencers' posts are short-lived with an exponential temporal decay, and characterise the time evolution of the interactions from their initial peak till the end of a post lifetime. Finally, leveraging our findings, we discuss how they can be exploited to develop an analytical model of the interactions temporal dynamics. Luca Vassio, Michele Garetto, Carla Fabiana Chiasserini, Emilio Leonardi |
ASONAM | 1 |
| 2021 | DarkVec: automatic analysis of darknet traffic with word embeddingsabstractDarknets are passive probes listening to traffic reaching IP addresses that host no services. Traffic reaching them is unsolicited by nature and often induced by scanners, malicious senders and misconfigured hosts. Its peculiar nature makes it a valuable source of information to learn about malicious activities. However, the massive amount of packets and sources that reach darknets makes it hard to extract meaningful insights. In particular, multiple senders contact the darknet while performing similar and coordinated tasks, which are often commanded by common controllers (botnets, crawlers, etc.). How to automatically identify and group those senders that share similar behaviors remains an open problem. Luca Gioacchini, Luca Vassio, Marco Mellia, Idilio Drago, Zied Ben-Houidi, Dario Rossi 0001 |
CoNEXT | 2 |
| 2021 | Understanding web pornography usage from traffic analysis
Andrea Morichetta 0002, Martino Trevisan, Luca Vassio, Julia Krickl |
Comput. Networks | 3 |
| 2021 | Towards website domain name classification using graph based semi-supervised learning
Azadeh Faroughi, Andrea Morichetta 0002, Luca Vassio, Flavio Figueiredo, Marco Mellia, Reza Javidan |
Comput. Networks | 3 |
| 2021 | Characterizing client usage patterns and service demand for car-sharing systems
Victor Aquiles Alencar, Felipe Rooke, Michele Cocca, Luca Vassio, Jussara M. Almeida, Alex Borges Vieira |
Inf. Syst. | 4 |
| 2020 | z-anonymity: Zero-Delay Anonymization for Data StreamsabstractWith the advent of big data and the birth of the data markets that sell personal information, individuals' privacy is of utmost importance. The classical response is anonymization, i.e., sanitizing the information that can directly or indirectly allow users' re-identification. The most popular solution in the literature is the k-anonymity. However, it is hard to achieve k-anonymity on a continuous stream of data, as well as when the number of dimensions becomes high.In this paper, we propose a novel anonymization property called z-anonymity. Differently from k-anonymity, it can be achieved with zero-delay on data streams and it is well suited for high dimensional data. The idea at the base of z-anonymity is to release an attribute (an atomic information) about a user only if at least z - 1 other users have presented the same attribute in a past time window. z-anonymity is weaker than k-anonymity since it does not work on the combinations of attributes, but treats them individually. In this paper, we present a probabilistic framework to map the z-anonymity into the k-anonymity property. Our results show that a proper choice of the z-anonymity parameters allows the data curator to likely obtain a k-anonymized dataset, with a precisely measurable probability. We also evaluate a real use case, in which we consider the website visits of a population of users and show that z-anonymity can work in practice for obtaining the k-anonymity too. Nikhil Jha, Thomas Favale, Luca Vassio, Martino Trevisan, Marco Mellia |
IEEE BigData | 3 |
| 2020 | E-Scooter Sharing: Leveraging Open Data for System DesignabstractWith the shift toward a Mobility-as-a-Service paradigm, electric scooter sharing systems are becoming a popular transportation mean in cities. Given their novelty, we lack of consolidated approaches to study and compare different system design options. In this work, we propose a simulation approach that leverages open data to create a demand model that captures and generalises the usage of this transportation mean in a city. This calls for ingenuity to deal with coarse open data granularity. In particular, we create a flexible, data-driven demand model by using modulated Poisson processes for temporal estimation, and Kernel Density Estimation (KDE) for spatial estimation. We next use this demand model alongside a configurable e-scooter sharing simulator to compare performance of different electric scooter sharing design options, such as the impact of the number of scooters and the cost of managing their charging. We focus on the municipalities of Minneapolis and Louisville which provide large scale open data about e-scooter sharing rides. Our approach let researchers, municipalities and scooter sharing providers to follow a data driven approach to compare and improve the design of e-scooter sharing system in smart cities. Alessandro Ciociola, Michele Cocca, Danilo Giordano, Luca Vassio, Marco Mellia |
DS-RT | 4 |
| 2019 | Data Analysis and Modelling of Users' Behavior on the Web
Luca Vassio, Marco Mellia |
IM | 1 |
| 2019 | Mining Patterns in Mobile Network Logs
Golnazsadat Zargarian, Luca Vassio, Maurizio M. Munafò, Marco Mellia |
IM | 2 |
| 2019 | Characterizing Web Pornography Consumption from Passive Measurements
Andrea Morichetta 0002, Martino Trevisan, Luca Vassio |
PAM | 3 |
| 2019 | Free floating electric car sharing design: Data driven optimisation
Michele Cocca, Danilo Giordano, Marco Mellia, Luca Vassio |
Pervasive Mob. Comput. | 4 |
| 2019 | Free Floating Electric Car Sharing: A Data Driven Approach for System DesignabstractIn this paper, we study the design of a free floating car sharing system based on electric vehicles. We rely on data about millions of rentals of a free floating car sharing operator based on internal combustion engine cars that we recorded in four cities. We characterize the nature of rentals, highlighting the non-stationary, and highly dynamic nature of usage patterns. Building on this data, we develop a discrete-event trace-driven simulator to study the usage of a hypothetical electric car sharing system. We use it to study the charging station placement problem, modeling different return policies, car battery charge and discharge due to trips, and the stochastic behavior of customers for plugging a car to a pole. Our data-driven approach helps car sharing providers to gauge the impact of different design solutions. Our simulations show that it is preferred to place charging stations within popular parking areas where cars are parked for short time (e.g., downtown). By smartly placing charging stations in just 8% of city zones, no trip ends with a discharged battery, i.e., all trips are feasible. Customers shall collaborate by bringing the car to a charging station when the battery level goes below a minimum threshold. This may reroute the customer to a different destination zone than the desired one; however, this happens in less than 10% of all trips. Michele Cocca, Danilo Giordano, Marco Mellia, Luca Vassio |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2018 | Free Floating Electric Car Sharing in Smart Cities: Data Driven System DimensioningabstractCar sharing is a popular means of transport in smart cities. The free floating paradigm lets the customers autonomously pick and drop available cars freely, within city limits. In this work we study the different policies when designing an electric Free Floating Car Sharing (FFCS) system. This system has the need to guarantee battery charge, a time-consuming operation, for which charging stations availability becomes a key factor for the sustainability of the whole system. We harvest the data of an already operative FFCS provider, and extract information about actual users' driving patterns. We implement a trace driven simulator to replay collected users' trips and simulate car batteries consumption for different design parameters. In this work, we limit the study to a single city, Turin (Italy), where we leverage actual trips registered over 2 months. We analyse and discuss several system design alternatives: the number of charging stations, their placement, and when to force users to return cars for charge. We identify regimes where cars never discharge and users can freely drop cars anywhere, albeit they are rarely rerouted to a charging station, possibly located in a nearby area to their original destination. Surprisingly, our data shows that even few charging stations (15 or more, i.e., 6% of city areas) guarantees the system to work almost autonomously, making thus possible free floating car sharing a feasible solution with electric cars. Michele Cocca, Danilo Giordano, Marco Mellia, Luca Vassio |
SMARTCOMP | 4 |
| 2018 | You, the Web, and Your Device: Longitudinal Characterization of Browsing HabitsabstractUnderstanding how people interact with the web is key for a variety of applications, e.g., from the design of effective web pages to the definition of successful online marketing campaigns. Browsing behavior has been traditionally represented and studied by means of clickstreams , i.e., graphs whose vertices are web pages, and edges are the paths followed by users. Obtaining large and representative data to extract clickstreams is, however, challenging. The evolution of the web questions whether browsing behavior is changing and, by consequence, whether properties of clickstreams are changing. This article presents a longitudinal study of clickstreams from 2013 to 2016. We evaluate an anonymized dataset of HTTP traces captured in a large ISP, where thousands of households are connected. We first propose a methodology to identify actual URLs requested by users from the massive set of requests automatically fired by browsers when rendering web pages. Then, we characterize web usage patterns and clickstreams, taking into account both the temporal evolution and the impact of the device used to explore the web. Our analyses precisely quantify various aspects of clickstreams and uncover interesting patterns, such as the typical short paths followed by people while navigating the web, the fast increasing trend in browsing from mobile devices, and the different roles of search engines and social networks in promoting content. Finally, we contribute a dataset of anonymized clickstreams to the community to foster new studies.1 Luca Vassio, Idilio Drago, Marco Mellia, Zied Ben-Houidi, Mohamed Lamine Lamali |
ACM Trans. Web | 1 |
| 2017 | Mining and modeling web trajectories from passive tracesabstractIn modern web, users contact lots of services, identified by the domain name of the server. The temporal sequence and transitions of visited domains form a trajectory of the user on the web. In this work, we analyze 4 weeks of such trajectories, extracted from logs collected in our university network, and mine them via big data and machine learning methodologies to extract the interests of users. Our goal is to create a model of such trajectories and find similarities so to observe peculiarity of users' browsing. Thanks to the model, we propose a methodology to automatically group together the trajectories of single users and/or communities into highly descriptive environments which in turn allow the analyst to identify the topic of interest. We propose an automatic way to highlight differences in terms of popularity and content of environments. Lastly, we analyze the transition among environments, showing how people in smaller communities, e.g., in the same department, have a much more homogeneous behaviour than people at large, e.g., in the university. Luca Vassio, Marco Mellia, Flavio Figueiredo, Ana Paula Couto da Silva, Jussara M. Almeida |
IEEE BigData | 1 |
| 2016 | A hybrid ABC for expensive optimizations: CEC 2016 competition benchmarkabstractAn evolution of the Artificial Bee Colony (ABC) optimization algorithm, called the Artificial super-Bee enhanced Colony (AsBeC), is presented for leading to the best improvement with a low number of analyses. AsBeC is designed to provide fast convergence speed, high solution accuracy and robust performance over a wide range of problems. It implements enhancements of ABC structure and original hybridizations with interpolation strategies. The aforementioned techniques are tested on the expensive benchmark of the Special Session on RealParameter Single Objective Optimization at CEC 2016. In this specific case, the hybridization with a quadratic trust region approach assumes a major importance. Moreover, the AsBeC results are compared to the algorithms tested on the same benchmark at CEC 2015, showing remarkable competitiveness and robustness. Enrico Ampellio, Luca Vassio |
CEC | 2 |
| 2016 | Detecting user actions from HTTP traces: Toward an automatic approachabstractDetecting explicit user actions, i.e., requests for web pages such as hyper-link clicks, from passive traces is fundamental for many applications, such as network forensics or content popularity estimation. Every URL explicitly visited by a user usually triggers further automatic URL requests to obtain all objects that compose the web page. HTTP traces provide a summary of all URLs requested by users, but no information that could be used to separate explicit from automatic requests. Previous works have targeted this problem and ad-hoc heuristics have been proposed. Validation has been typically done using synthetic traces. This paper investigates whether an approach based solely on machine learning can successfully detect user actions from HTTP traces. A machine learning approach would come with many advantages - e.g., it minimizes manual tuning of parameters and can easily adapt to page structure changes. We build both real and synthetic traces to assess the performance and gain insights on the features that bring most advantages in classification. Our results show that machine learning reaches similar or better performance as previous heuristics. Furthermore, we show that models built with machine learning algorithms are robust, presenting consistent performance in different scenarios. Luca Vassio, Idilio Drago, Marco Mellia |
IWCMC | 1 |