VLDB 2026 Research / reviewers in the wild / expert
Alessandro Vespignani
dblp:65/1550
· DBLP profile ↗
28ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0003-3419-4205ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 5 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 5 since 2021Computer networks · 2Theory of computation · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Human-AI Coevolution (Abstract Reprint)abstractHuman-AI coevolution, defined as a process in which humans and AI algorithms continuously influence each other, increasingly characterises our society, but is understudied in artificial intelligence and complexity science literature. Recommender systems and assistants play a prominent role in human-AI coevolution, as they permeate many facets of daily life and influence human choices through online platforms. The interaction between users and AI results in a potentially endless feedback loop, wherein users' choices generate data to train AI models, which, in turn, shape subsequent user preferences. This human-AI feedback loop has peculiar characteristics compared to traditional human-machine interaction and gives rise to complex and often “unintended” systemic outcomes. This paper introduces human-AI coevolution as the cornerstone for a new field of study at the intersection between AI and complexity science focused on the theoretical, empirical, and mathematical investigation of the human-AI feedback loop. In doing so, we: (i) outline the pros and cons of existing methodologies and highlight shortcomings and potential ways for capturing feedback loop mechanisms; (ii) propose a reflection at the intersection between complexity science, AI and society; (iii) provide real-world examples for different human-AI ecosystems; and (iv) illustrate challenges to the creation of such a field of study, conceptualising them at increasing levels of abstraction, i.e., scientific, legal and socio-political. Dino Pedreschi, Luca Pappalardo, Emanuele Ferragina, Ricardo Baeza-Yates, Albert-László Barabási, Frank Dignum, Virginia Dignum, Tina Eliassi-Rad, Fosca Giannotti, János Kertész, Alistair Knott, Yannis E. Ioannidis, Paul Lukowicz, Andrea Passarella, Alex Pentland, John Shawe-Taylor, Alessandro Vespignani |
IJCAI | 17 |
| 2025 | Human-AI coevolutionabstractHuman-AI coevolution, defined as a process in which humans and AI algorithms continuously influence each other, increasingly characterises our society, but is understudied in artificial intelligence and complexity science literature. Recommender systems and assistants play a prominent role in human-AI coevolution, as they permeate many facets of daily life and influence human choices through online platforms. The interaction between users and AI results in a potentially endless feedback loop, wherein users' choices generate data to train AI models, which, in turn, shape subsequent user preferences. This human-AI feedback loop has peculiar characteristics compared to traditional human-machine interaction and gives rise to complex and often “unintended” systemic outcomes. This paper introduces human-AI coevolution as the cornerstone for a new field of study at the intersection between AI and complexity science focused on the theoretical, empirical, and mathematical investigation of the human-AI feedback loop. In doing so, we: (i) outline the pros and cons of existing methodologies and highlight shortcomings and potential ways for capturing feedback loop mechanisms; (ii) propose a reflection at the intersection between complexity science, AI and society; (iii) provide real-world examples for different human-AI ecosystems; and (iv) illustrate challenges to the creation of such a field of study, conceptualising them at increasing levels of abstraction, i.e., scientific, legal and socio-political. Dino Pedreschi, Luca Pappalardo, Emanuele Ferragina, Ricardo Baeza-Yates, Albert-László Barabási, Frank Dignum, Virginia Dignum, Tina Eliassi-Rad, Fosca Giannotti, János Kertész, Alistair Knott, Yannis E. Ioannidis, Paul Lukowicz, Andrea Passarella, Alex Pentland, John Shawe-Taylor, Alessandro Vespignani |
Artif. Intell. | 17 |
| 2025 | Epydemix: An open-source Python package for epidemic modeling with integrated approximate Bayesian calibrationabstractWe present Epydemix, an open-source Python package for the development and calibration of stochastic compartmental epidemic models. The framework supports flexible model structures that incorporate demographic information, age-stratified contact matrices, and dynamic public health interventions. A key feature of Epydemix is its integration of Approximate Bayesian Computation (ABC) techniques to perform parameter inference and model calibration through comparison between observed and simulated data. The package offers a range of ABC methods such as simple rejection sampling, simulation-budget-constrained rejection, and Sequential Monte Carlo (ABC-SMC). Epydemix is modular, and supports ABC-based calibration both for models defined within the package and for those developed externally. To demonstrate the computational framework capabilities, we discuss usage examples that include (i) simulating an intervention-driven model with time-varying parameters, and (ii) benchmarking calibration performance using synthetic epidemic data. We further illustrate the use of the package in a retrospective case study that includes scenario projections under alternative intervention assumptions. By lowering the barrier for the implementation of computational and inference approaches, Epydemix makes epidemic modeling more accessible to a wider range of users, from academic researchers to public health professionals. Nicolò Gozzi, Matteo Chinazzi, Jessica T. Davis, Corrado Gioannini, Luca Rossi 0010, Marco Ajelli, Nicola Perra, Alessandro Vespignani |
PLoS Comput. Biol. | 8 |
| 2024 | Challenges of COVID-19 Case Forecasting in the US, 2020-2021abstractDuring the COVID-19 pandemic, forecasting COVID-19 trends to support planning and response was a priority for scientists and decision makers alike. In the United States, COVID-19 forecasting was coordinated by a large group of universities, companies, and government entities led by the Centers for Disease Control and Prevention and the US COVID-19 Forecast Hub (https://covid19forecasthub.org). We evaluated approximately 9.7 million forecasts of weekly state-level COVID-19 cases for predictions 1-4 weeks into the future submitted by 24 teams from August 2020 to December 2021. We assessed coverage of central prediction intervals and weighted interval scores (WIS), adjusting for missing forecasts relative to a baseline forecast, and used a Gaussian generalized estimating equation (GEE) model to evaluate differences in skill across epidemic phases that were defined by the effective reproduction number. Overall, we found high variation in skill across individual models, with ensemble-based forecasts outperforming other approaches. Forecast skill relative to the baseline was generally higher for larger jurisdictions (e.g., states compared to counties). Over time, forecasts generally performed worst in periods of rapid changes in reported cases (either in increasing or decreasing epidemic phases) with 95% prediction interval coverage dropping below 50% during the growth phases of the winter 2020, Delta, and Omicron waves. Ideally, case forecasts could serve as a leading indicator of changes in transmission dynamics. However, while most COVID-19 case forecasts outperformed a naïve baseline model, even the most accurate case forecasts were unreliable in key phases. Further research could improve forecasts of leading indicators, like COVID-19 cases, by leveraging additional real-time data, addressing performance across phases, improving the characterization of forecast confidence, and ensuring that forecasts were coherent across spatial scales. In the meantime, it is critical for forecast users to appreciate current limitations and use a broad set of indicators to inform pandemic-related decision making. Velma K. Lopez, Estee Y. Cramer, Robert Pagano, John M. Drake, Eamon B. O'Dea, Madeline Adee, Turgay Ayer, Jagpreet Chhatwal, Ozden O. Dalgic, Mary A. Ladd, Benjamin P. Linas, Peter P. Mueller, Jade Xiao, Johannes Bracher, Alvaro J. Castro Rivadeneira, Aaron Gerding, Tilmann Gneiting, Yuxin Huang 0009, Dasuni Jayawardena, Abdul H. Kanji, Khoa Le, Anja Mühlemann, Jarad Niemi, Evan L. Ray, Ariane Stark, Nutcha Wattanachit, Martha W. Zorn, Sen Pei, Jeffrey Shaman, Teresa K. Yamana, Samuel R. Tarasewicz, Daniel J. Wilson 0002, Sid Baccam, Heidi Gurung, Steve Stage, Brad Suchoski, Lei Gao 0011, Zhiling Gu, Myungjin Kim, Guannan Wang, Li Wang 0035, Yueying Wang, Lauren Gardner, Sonia Jindal, Maximilian Marshall, Kristen Nixon, Juan Dent, Alison L. Hill, Joshua Kaminsky, Elizabeth C. Lee, Joseph Chadi Lemaitre, Justin Lessler, Claire P. Smith, Shaun Truelove, Matt Kinsey, Luke C. Mullany, Kaitlin Rainwater-Lovett, Lauren Shin, Katharine Tallaksen, Shelby Wilson, Dean Karlen, Lauren A. Castro, Geoffrey Fairchild, Isaac Michaud, Dave Osthus, Jiang Bian 0002, Wei Cao 0007, Zhifeng Gao, Juan M. Lavista Ferres, Chaozhuo Li, Tie-Yan Liu, Xing Xie 0001, Shun Zheng 0001, Matteo Chinazzi, Jessica T. Davis, Kunpeng Mu, Ana L. Pastore y Piontti, Alessandro Vespignani, Xinyue Xiong, Robert Walraven, Quanquan Gu, Lingxiao Wang 0001, Pan Xu 0002, Difan Zou, Graham Casey Gibson, Daniel Sheldon, Ajitesh Srivastava, Aniruddha Adiga, Benjamin Hurt, Gursharn Kaur, Bryan L. Lewis, Madhav V. Marathe, Akhil Sai Peddireddy, Przemyslaw J. Porebski, Srinivasan Venkatramanan, Lijing Wang 0001, Pragati V. Prasad, Jo W. Walker, Alexander E. Webber, Rachel B. Slayton, Matthew Biggerstaff, Nicholas G. Reich, Michael A. Johansson |
PLoS Comput. Biol. | 82 |
| 2023 | Deep Bayesian Active Learning for Accelerating Stochastic SimulationabstractStochastic simulations such as large-scale, spatiotemporal, age-structured epidemic models are computationally expensive at fine-grained resolution. While deep surrogate models can speed up the simulations, doing so for stochastic simulations and with active learning approaches is an underexplored area. We propose Interactive Neural Process (INP), a deep Bayesian active learning framework for learning deep surrogate models to accelerate stochastic simulations. INP consists of two components, a spatiotemporal surrogate model built upon Neural Process (NP) family and an acquisition function for active learning. For surrogate modeling, we develop Spatiotemporal Neural Process (STNP) to mimic the simulator dynamics. For active learning, we propose a novel acquisition function, Latent Information Gain (LIG), calculated in the latent space of NP based models. We perform a theoretical analysis and demonstrate that LIG reduces sample complexity compared with random sampling in high dimensions. We also conduct empirical studies on three complex spatiotemporal simulators for reaction diffusion, heat flow, and infectious disease. The results demonstrate that STNP outperforms the baselines in the offline learning setting and LIG achieves the state-of-the-art for Bayesian active learning. Dongxia Wu, Ruijia Niu, Matteo Chinazzi, Alessandro Vespignani, Yi-An Ma, Rose Yu |
KDD | 4 |
| 2022 | Multi-fidelity Hierarchical Neural ProcessesabstractScience and engineering fields use computer simulation extensively. These simulations are often run at multiple levels of sophistication to balance accuracy and efficiency. Multi-fidelity surrogate modeling reduces the computational cost by fusing different simulation outputs. Cheap data generated from low-fidelity simulators can be combined with limited high-quality data generated by an expensive high-fidelity simulator. Existing methods based on Gaussian processes rely on strong assumptions of the kernel functions and can hardly scale to high-dimensional settings. We propose Multi-fidelity Hierarchical Neural Processes (MF-HNP), a unified neural latent variable model for multi-fidelity surrogate modeling. MF-HNP inherits the flexibility and scalability of Neural Processes. The latent variables transform the correlations among different fidelity levels from observations to latent space. The predictions across fidelities are conditionally independent given the latent states. It helps alleviate the error propagation issue in existing methods. MF-HNP is flexible enough to handle non-nested high dimensional data at different fidelity levels with varying input and output dimensions. We evaluate MF-HNP on epidemiology and climate modeling tasks, achieving competitive performance in terms of accuracy and uncertainty estimation. In contrast to deep Gaussian Processes with only low-dimensional (< 10) tasks, our method shows great promise for speeding up high-dimensional complex simulations (over 7000 for epidemiology modeling and 45000 for climate modeling). Dongxia Wu, Matteo Chinazzi, Alessandro Vespignani, Yi-An Ma, Rose Yu |
KDD | 3 |
| 2022 | Anatomy of the first six months of COVID-19 vaccination campaign in ItalyabstractWe analyze the effectiveness of the first six months of vaccination campaign against SARS-CoV-2 in Italy by using a computational epidemic model which takes into account demographic, mobility, vaccines data, as well as estimates of the introduction and spreading of the more transmissible Alpha variant. We consider six sub-national regions and study the effect of vaccines in terms of number of averted deaths, infections, and reduction in the Infection Fatality Rate (IFR) with respect to counterfactual scenarios with the actual non-pharmaceuticals interventions but no vaccine administration. Furthermore, we compare the effectiveness in counterfactual scenarios with different vaccines allocation strategies and vaccination rates. Our results show that, as of 2021/07/05, vaccines averted 29, 350 (IQR: [16, 454-42, 826]) deaths and 4, 256, 332 (IQR: [1, 675, 564-6, 980, 070]) infections and a new pandemic wave in the country. During the same period, they achieved a -22.2% (IQR: [-31.4%; -13.9%]) IFR reduction. We show that a campaign that would have strictly prioritized age groups at higher risk of dying from COVID-19, besides frontline workers and the fragile population, would have implied additional benefits both in terms of avoided fatalities and reduction in the IFR. Strategies targeting the most active age groups would have prevented a higher number of infections but would have been associated with more deaths. Finally, we study the effects of different vaccination intake scenarios by rescaling the number of available doses in the time period under study to those administered in other countries of reference. The modeling framework can be applied to other countries to provide a mechanistic characterization of vaccination campaigns worldwide. Nicolò Gozzi, Matteo Chinazzi, Jessica T. Davis, Kunpeng Mu, Ana L. Pastore y Piontti, Marco Ajelli, Nicola Perra, Alessandro Vespignani |
PLoS Comput. Biol. | 8 |
| 2021 | Quantifying Uncertainty in Deep Spatiotemporal ForecastingabstractDeep learning is gaining increasing popularity for spatiotemporal forecasting. However, prior works have mostly focused on point estimates without quantifying the uncertainty of the predictions. In high stakes domains, being able to generate probabilistic forecasts with confidence intervals is critical to risk assessment and decision making. Hence, a systematic study of uncertainty quantification (UQ) methods for spatiotemporal forecasting is missing in the community. In this paper, we describe two types of spatiotemporal forecasting problems: regular grid-based and graph-based. Then we analyze UQ methods from both the Bayesian and the frequentist point of view, casting in a unified framework via statistical decision theory. Through extensive experiments on real-world road network traffic, epidemics, and air quality forecasting tasks, we reveal the statistical and computational trade-offs for different UQ methods: Bayesian methods are typically more robust in mean prediction, while confidence levels obtained from frequentist methods provide more extensive coverage over data variations. Computationally, quantile regression type methods are cheaper for a single confidence interval but require re-training for different intervals. Sampling based methods generate samples that can form multiple confidence intervals, albeit at a higher computational cost. Dongxia Wu, Liyao Gao, Matteo Chinazzi, Xinyue Xiong, Alessandro Vespignani, Yi-An Ma, Rose Yu |
KDD | 5 |
| 2021 | Estimating the cumulative incidence of COVID-19 in the United States using influenza surveillance, virologic testing, and mortality data: Four complementary approachesabstractEffectively designing and evaluating public health responses to the ongoing COVID-19 pandemic requires accurate estimation of the prevalence of COVID-19 across the United States (US). Equipment shortages and varying testing capabilities have however hindered the usefulness of the official reported positive COVID-19 case counts. We introduce four complementary approaches to estimate the cumulative incidence of symptomatic COVID-19 in each state in the US as well as Puerto Rico and the District of Columbia, using a combination of excess influenza-like illness reports, COVID-19 test statistics, COVID-19 mortality reports, and a spatially structured epidemic model. Instead of relying on the estimate from a single data source or method that may be biased, we provide multiple estimates, each relying on different assumptions and data sources. Across our four approaches emerges the consistent conclusion that on April 4, 2020, the estimated case count was 5 to 50 times higher than the official positive test counts across the different states. Nationally, our estimates of COVID-19 symptomatic cases as of April 4 have a likely range of 2.3 to 4.8 million, with possibly as many as 7.6 million cases, up to 25 times greater than the cumulative confirmed cases of about 311,000. Extending our methods to May 16, 2020, we estimate that cumulative symptomatic incidence ranges from 4.9 to 10.1 million, as opposed to 1.5 million positive test counts. The proposed combination of approaches may prove useful in assessing the burden of COVID-19 during resurgences in the US and other countries with comparable surveillance systems. Fred S. Lu, André T. Nguyen, Nicholas B. Link, Mathieu Molina, Jessica T. Davis, Matteo Chinazzi, Xinyue Xiong, Alessandro Vespignani, Marc Lipsitch, Mauricio Santillana |
PLoS Comput. Biol. | 8 |
| 2021 | Predicting seasonal influenza using supermarket retail recordsabstractIncreased availability of epidemiological data, novel digital data streams, and the rise of powerful machine learning approaches have generated a surge of research activity on real-time epidemic forecast systems. In this paper, we propose the use of a novel data source, namely retail market data to improve seasonal influenza forecasting. Specifically, we consider supermarket retail data as a proxy signal for influenza, through the identification of sentinel baskets, i.e., products bought together by a population of selected customers. We develop a nowcasting and forecasting framework that provides estimates for influenza incidence in Italy up to 4 weeks ahead. We make use of the Support Vector Regression (SVR) model to produce the predictions of seasonal flu incidence. Our predictions outperform both a baseline autoregressive model and a second baseline based on product purchases. The results show quantitatively the value of incorporating retail market data in forecasting models, acting as a proxy that can be used for the real-time analysis of epidemics. Ioanna Miliou, Xinyue Xiong, Salvatore Rinzivillo, Qian Zhang 0016, Giulio Rossetti, Fosca Giannotti, Dino Pedreschi, Alessandro Vespignani |
PLoS Comput. Biol. | 8 |
| 2020 | Keynote Speaker: Alessandro VespignaniabstractAlessandro Vespignani research activity is focused on the study of "techno-social" systems, where infrastructures composed of different technological layers are interoperating within the social component that drives their use and development. In this context we aim at understanding how the very same elements assembled in large number can give rise - according to the various forces and elements at play - to different macroscopic and dynamical behaviors, opening the path to quantitative computational approaches and forecasting power. The main research lines pursued at the moment are: Develop analytical and computational models for the co-evolution and interdependence of large-scale social, technological and biological networks. Modeling contagion processes in structured populations. Developing predictive computational tools for the analysis of the spatial spread of emerging diseases. Analyze the dynamics and evolution of information and social networks. Model the adaptive behavior of social systems. Prof. Vespignani is a joint appointment between the College of Science, the College of Computer and Information Science, and the Bouvé College of Health Sciences. Alessandro Vespignani |
KDD | 1 |
| 2020 | Detecting critical slowing down in high-dimensional epidemiological systemsabstractDespite medical advances, the emergence and re-emergence of infectious diseases continue to pose a public health threat. Low-dimensional epidemiological models predict that epidemic transitions are preceded by the phenomenon of critical slowing down (CSD). This has raised the possibility of anticipating disease (re-)emergence using CSD-based early-warning signals (EWS), which are statistical moments estimated from time series data. For EWS to be useful at detecting future (re-)emergence, CSD needs to be a generic (model-independent) feature of epidemiological dynamics irrespective of system complexity. Currently, it is unclear whether the predictions of CSD-derived from simple, low-dimensional systems-pertain to real systems, which are high-dimensional. To assess the generality of CSD, we carried out a simulation study of a hierarchy of models, with increasing structural complexity and dimensionality, for a measles-like infectious disease. Our five models included: i) a nonseasonal homogeneous Susceptible-Exposed-Infectious-Recovered (SEIR) model, ii) a homogeneous SEIR model with seasonality in transmission, iii) an age-structured SEIR model, iv) a multiplex network-based model (Mplex) and v) an agent-based simulator (FRED). All models were parameterised to have a herd-immunity immunization threshold of around 90% coverage, and underwent a linear decrease in vaccine uptake, from 92% to 70% over 15 years. We found evidence of CSD prior to disease re-emergence in all models. We also evaluated the performance of seven EWS: the autocorrelation, coefficient of variation, index of dispersion, kurtosis, mean, skewness, variance. Performance was scored using the Area Under the ROC Curve (AUC) statistic. The best performing EWS were the mean and variance, with AUC > 0.75 one year before the estimated transition time. These two, along with the autocorrelation and index of dispersion, are promising candidate EWS for detecting disease emergence. Tobias S. Brett, Marco Ajelli, Quanhui Liu, Mary G. Krauland, John J. Grefenstette, Wilbert Van Panhuis, Alessandro Vespignani, John M. Drake, Pejman Rohani |
PLoS Comput. Biol. | 7 |
| 2020 | The COVID-19 outbreak in Sichuan, China: Epidemiology and impact of interventionsabstractIn January 2020, a COVID-19 outbreak was detected in Sichuan Province of China. Six weeks later, the outbreak was successfully contained. The aim of this work is to characterize the epidemiology of the Sichuan outbreak and estimate the impact of interventions in limiting SARS-CoV-2 transmission. We analyzed patient records for all laboratory-confirmed cases reported in the province for the period of January 21 to March 16, 2020. To estimate the basic and daily reproduction numbers, we used a Bayesian framework. In addition, we estimated the number of cases averted by the implemented control strategies. The outbreak resulted in 539 confirmed cases, lasted less than two months, and no further local transmission was detected after February 27. The median age of local cases was 8 years older than that of imported cases. We estimated R0 at 2.4 (95% CI: 1.6-3.7). The epidemic was self-sustained for about 3 weeks before going below the epidemic threshold 3 days after the declaration of a public health emergency by Sichuan authorities. Our findings indicate that, were the control measures be adopted four weeks later, the epidemic could have lasted 49 days longer (95% CI: 31-68 days), causing 9,216 more cases (95% CI: 1,317-25,545). Quanhui Liu, Ana I. Bento, Kexin Yang 0002, Hang Zhang 0029, Stefano Merler, Alessandro Vespignani, Jiancheng Lv 0001, Tao Zhou 0001, Marco Ajelli |
PLoS Comput. Biol. | 7 |
| 2017 | Forecasting Seasonal Influenza Fusing Digital Indicators and a Mechanistic Disease ModelabstractThe availability of novel digital data streams that can be used as proxy for monitoring infectious disease incidence is ushering in a new era for real-time forecast approaches to disease spreading. Here, we propose the first seasonal influenza forecast framework based on a stochastic, spatially structured mechanistic model (individual level microsimulation) initialized with geo-localized microblogging data. The framework provides for more than 600 census areas in the United States, Italy and Spain, the initial conditions for a stochastic epidemic computational model that generates an ensemble of forecasts for the main indicators of the epidemic season: peak time and intensity. We evaluate the forecasts accuracy and reliability by comparing the results with the data from the official influenza surveillance systems in the US, Italy and Spain in the seasons 2014/15 and 2015/16. In all countries studied, the proposed framework provides reliable results with leads of up to 6 weeks that became more stable and accurate with progression of the season. The results for the United States have been generated in real-time in the context of the Centers for Disease Control and Prevention ``Forecasting the Influenza Season Challenge''. A characteristic feature of the mechanistic modeling approach is in the explicit estimate of key epidemiological parameters relevant for public health decision-making that cannot be achieved with statistical models that do not consider the disease dynamic. Furthermore, the presented framework allows the fusion of multiple data streams in the initialization stage and can be enriched with census, weather and socioeconomic data. Qian Zhang 0016, Nicola Perra, Daniela Perrotta, Michele Tizzoni, Daniela Paolotti, Alessandro Vespignani |
WWW | 6 |
| 2015 | Social Data Mining and Seasonal Influenza Forecasts: The FluOutlook Platform
Qian Zhang 0016, Corrado Gioannini, Daniela Paolotti, Nicola Perra, Daniela Perrotta, Marco Quaggiotto, Michele Tizzoni, Alessandro Vespignani |
ECML/PKDD (3) | 8 |
| 2013 | Host Mobility Drives Pathogen Competition in Spatially Structured PopulationsabstractInteractions among multiple infectious agents are increasingly recognized as a fundamental issue in the understanding of key questions in public health regarding pathogen emergence, maintenance, and evolution. The full description of host-multipathogen systems is, however, challenged by the multiplicity of factors affecting the interaction dynamics and the resulting competition that may occur at different scales, from the within-host scale to the spatial structure and mobility of the host population. Here we study the dynamics of two competing pathogens in a structured host population and assess the impact of the mobility pattern of hosts on the pathogen competition. We model the spatial structure of the host population in terms of a metapopulation network and focus on two strains imported locally in the system and having the same transmission potential but different infectious periods. We find different scenarios leading to competitive success of either one of the strain or to the codominance of both strains in the system. The dominance of the strain characterized by the shorter or longer infectious period depends exclusively on the structure of the population and on the the mobility of hosts across patches. The proposed modeling framework allows the integration of other relevant epidemiological, environmental and demographic factors, opening the path to further mathematical and computational studies of the dynamics of multipathogen systems. Chiara Poletto, Sandro Meloni, Vittoria Colizza, Yamir Moreno, Alessandro Vespignani |
PLoS Comput. Biol. | 5 |
| 2012 | Complex dynamic networks: Tools and methods
J. Ignacio Alvarez-Hamelin, Eric Fleury, Alessandro Vespignani, Artur Ziviani |
Comput. Networks | 3 |
| 2012 | Inferring the Structure of Social Contacts from Demographic Data in the Analysis of Infectious Diseases SpreadabstractSocial contact patterns among individuals encode the transmission route of infectious diseases and are a key ingredient in the realistic characterization and modeling of epidemics. Unfortunately, the gathering of high quality experimental data on contact patterns in human populations is a very difficult task even at the coarse level of mixing patterns among age groups. Here we propose an alternative route to the estimation of mixing patterns that relies on the construction of virtual populations parametrized with highly detailed census and demographic data. We present the modeling of the population of 26 European countries and the generation of the corresponding synthetic contact matrices among the population age groups. The method is validated by a detailed comparison with the matrices obtained in six European countries by the most extensive survey study on mixing patterns. The methodology presented here allows a large scale comparison of mixing patterns in Europe, highlighting general common features as well as country-specific differences. We find clear relations between epidemiologically relevant quantities (reproduction number and attack rate) and socio-demographic characteristics of the populations, such as the average age of the population and the duration of primary school cycle. This study provides a numerical approach for the generation of human mixing patterns that can be used to improve the accuracy of mathematical models in the absence of specific experimental data. Laura Fumanelli, Marco Ajelli, Piero Manfredi, Alessandro Vespignani, Stefano Merler |
PLoS Comput. Biol. | 4 |
| 2012 | Digital EpidemiologyabstractMobile, social, real-time: the ongoing revolution in the way people communicate has given rise to a new kind of epidemiology. Digital data sources, when harnessed appropriately, can provide local and timely information about disease and health dynamics in populations around the world. The rapid, unprecedented increase in the availability of relevant data from various digital sources creates considerable technical and computational challenges. Marcel Salathé, Linus Bengtsson, Todd J. Bodnar, Devon D. Brewer, John S. Brownstein, Caroline O. Buckee, Ellsworth M. Campbell, Ciro Cattuto, Shashank Khandelwal, Patricia L. Mabry, Alessandro Vespignani |
PLoS Comput. Biol. | 11 |
| 2011 | Properties and Evolution of Internet Traffic Networks from Anonymized Flow DataabstractMany projects have tried to analyze the structure and dynamics of application overlay networks on the Internet using packet analysis and network flow data. While such analysis is essential for a variety of network management and security tasks, it is infeasible on many networks: either the volume of data is so large as to make packet inspection intractable, or privacy concerns forbid packet capture and require the dissociation of network flows from users’ actual IP addresses. Our analytical framework permits useful analysis of network usage patterns even under circumstances where the only available source of data is anonymized flow records. Using this data, we are able to uncover distributions and scaling relations in host-to-host networks that bear implications for capacity planning and network application design. We also show how to classify network applications based entirely on topological properties of their overlay networks, yielding a taxonomy that allows us to accurately identify the functions of unknown applications. We repeat this analysis on a more recent dataset, allowing us to demonstrate that the aggregate behavior of users is remarkably stable even as the population changes. Mark R. Meiss, Filippo Menczer, Alessandro Vespignani |
ACM Trans. Internet Techn. | 3 |
| 2008 | Ranking web sites with real user trafficabstractWe analyze the traffic-weighted Web host graph obtained from a large sample of real Web users over about seven months. A number of interesting structural properties are revealed by this complex dynamic network, some in line with the well-studied boolean link host graph and others pointing to important differences. We find that while search is directly involved in a surprisingly small fraction of user clicks, it leads to a much larger fraction of all sites visited. The temporal traffic patterns display strong regularities, with a large portion of future requests being statistically predictable by past ones. Given the importance of topological measures such as PageRank in modeling user navigation, as well as their role in ranking sites for Web search, we use the traffic data to validate the PageRank random surfing model. The ranking obtained by the actual frequency with which a site is visited by users differs significantly from that approximated by the uniform surfing/teleportation behavior modeled by PageRank, especially for the most important sites. To interpret this finding, we consider each of the fundamental assumptions underlying PageRank and show how each is violated by actual user behavior Mark R. Meiss, Filippo Menczer, Santo Fortunato, Alessandro Flammini, Alessandro Vespignani |
WSDM | 5 |
| 2007 | Decoding the structure of the WWW: A comparative analysis of Web crawlsabstractThe understanding of the immense and intricate topological structure of the World Wide Web (WWW) is a major scientific and technological challenge. This has been recently tackled by characterizing the properties of its representative graphs, in which vertices and directed edges are identified with Web pages and hyperlinks, respectively. Data gathered in large-scale crawls have been analyzed by several groups resulting in a general picture of the WWW that encompasses many of the complex properties typical of rapidly evolving networks. In this article, we report a detailed statistical analysis of the topological properties of four different WWW graphs obtained with different crawlers. We find that, despite the very large size of the samples, the statistical measures characterizing these graphs differ quantitatively, and in some cases qualitatively, depending on the domain analyzed and the crawl used for gathering the data. This spurs the issue of the presence of sampling biases and structural differences of Web crawls that might induce properties not representative of the actual global underlying graph. In short, the stability of the widely accepted statistical description of the Web is called into question. In order to provide a more accurate characterization of the Web graph, we study statistical measures beyond the degree distribution, such as degree-degree correlation functions or the statistics of reciprocal connections. The latter appears to enclose the relevant correlations of the WWW graph and carry most of the topological information of the Web. The analysis of this quantity is also of major interest in relation to the navigability and searchability of the Web. M. Ángeles Serrano, Ana Gabriela Maguitman, Marián Boguñá, Santo Fortunato, Alessandro Vespignani |
ACM Trans. Web | 5 |
| 2006 | Exploring networks with traceroute-like probes: Theory and simulations
Luca Dall'Asta, J. Ignacio Alvarez-Hamelin, Alain Barrat, Alexei Vázquez, Alessandro Vespignani |
Theor. Comput. Sci. | 5 |
| 2006 | Algorithmic Computation and Approximation of Semantic Similarity
Ana Gabriela Maguitman, Filippo Menczer, Fulya Erdinc, Heather Roinestad, Alessandro Vespignani |
World Wide Web | 5 |
| 2005 | Large scale networks fingerprinting and visualization using the k-core decompositionabstractWe use the k-core decomposition to develop algorithms for the analysis of large scale complex networks. This decomposition, based on a re- cursive pruning of the least connected vertices, allows to disentangle the hierarchical structure of networks by progressively focusing on their cen- tral cores. By using this strategy we develop a general visualization algo- rithm that can be used to compare the structural properties of various net- works and highlight their hierarchical structure. The low computational complexity of the algorithm, O(n + e), where n is the size of the net- work, and e is the number of edges, makes it suitable for the visualization of very large sparse networks. We show how the proposed visualization tool allows to find specific structural fingerprints of networks. J. Ignacio Alvarez-Hamelin, Luca Dall'Asta, Alain Barrat, Alessandro Vespignani |
NIPS | 4 |
| 2005 | Algorithmic detection of semantic similarityabstractAutomatic extraction of semantic information from text and links in Web pages is key to improving the quality of search results. However, the assessment of automatic semantic measures is limited by the coverage of user studies, which do not scale with the size, heterogeneity, and growth of the Web. Here we propose to leverage human-generated metadata --- namely topical directories --- to measure semantic relationships among massive numbers of pairs of Web pages or topics. The Open Directory Project classifies millions of URLs in a topical ontology, providing a rich source from which semantic relationships between Web pages can be derived. While semantic similarity measures based on taxonomies (trees) are well studied, the design of well-founded similarity measures for objects stored in the nodes of arbitrary ontologies (graphs) is an open problem. This paper defines an information-theoretic measure of semantic similarity that exploits both the hierarchical and non-hierarchical structure of an ontology. An experimental study shows that this measure improves significantly on the traditional taxonomy-based approach. This novel measure allows us to address the general question of how text and link analyses can be combined to derive measures of relevance that are in good agreement with semantic similarity. Surprisingly, the traditional use of text similarity turns out to be ineffective for relevance ranking. Ana Gabriela Maguitman, Filippo Menczer, Heather Roinestad, Alessandro Vespignani |
WWW | 4 |
| 2005 | On the lack of typical behavior in the global Web traffic networkabstractWe offer the first large-scale analysis of Web traffic based on network flow data. Using data collected on the Internet2 network, we constructed a weighted bipartite clientserver host graph containing more than 18 × 10^6 vertices and 68 × 10^6 edges valued by relative traffic flows. When considered as a traffic map of the World-Wide Web, the generated graph provides valuable information on the statistical patterns that characterize the global information flow on the Web. Statistical analysis shows that client-server connections and traffic flows exhibit heavy-tailed probability distributions lacking any typical scale. In particular, the absence of an intrinsic average in some of the distributions implies the absence of a prototypical scale appropriate for server design, Web-centric network design, or traffic modeling. The inspection of the amount of traffic handled by clients and servers and their number of connections highlights non-trivial correlations between information flow and patterns of connectivity as well as the presence of anomalous statistical patterns related to the behavior of users on the Web. The results presented here may impact considerably the modeling, scalability analysis, and behavioral study of Web applications. Mark R. Meiss, Filippo Menczer, Alessandro Vespignani |
WWW | 3 |
| 2004 | Traffic-Driven Model of the World Wide Web Graph
Alain Barrat, Marc Barthelemy, Alessandro Vespignani |
WAW | 3 |