VLDB 2026 Research / reviewers in the wild / expert
Shweta Bansal
dblp:94/2947 · also Shweta A. Bansal
· DBLP profile ↗
10ranked-venue papers
1as first author
6since 2021 · last 2023
0000-0002-1740-5421ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Charting the spatial dynamics of early SARS-CoV-2 transmission in Washington stateabstractThe spread of SARS-CoV-2 has been geographically uneven. To understand the drivers of this spatial variation in SARS-CoV-2 transmission, in particular the role of stochasticity, we used the early stages of the SARS-CoV-2 invasion in Washington state as a case study. We analysed spatially-resolved COVID-19 epidemiological data using two distinct statistical analyses. The first analysis involved using hierarchical clustering on the matrix of correlations between county-level case report time series to identify geographical patterns in the spread of SARS-CoV-2 across the state. In the second analysis, we used a stochastic transmission model to perform likelihood-based inference on hospitalised cases from five counties in the Puget Sound region. Our clustering analysis identifies five distinct clusters and clear spatial patterning. Four of the clusters correspond to different geographical regions, with the final cluster spanning the state. Our inferential analysis suggests that a high degree of connectivity across the region is necessary for the model to explain the rapid inter-county spread observed early in the pandemic. In addition, our approach allows us to quantify the impact of stochastic events in determining the subsequent epidemic. We find that atypically rapid transmission during January and February 2020 is necessary to explain the observed epidemic trajectories in King and Snohomish counties, demonstrating a persisting impact of stochastic events. Our results highlight the limited utility of epidemiological measures calculated over broad spatial scales. Furthermore, our results make clear the challenges with predicting epidemic spread within spatially extensive metropolitan areas, and indicate the need for high-resolution mobility and epidemiological data. Tobias S. Brett, Shweta Bansal, Pejman Rohani |
PLoS Comput. Biol. | 2 |
| 2022 | Secure e-Voting System - A Review
Urmila Devi, Shweta Bansal |
HIS | 2 |
| 2022 | Dissecting recurrent waves of pertussis across the boroughs of LondonabstractPertussis has resurfaced in the UK, with incidence levels not seen since the 1980s. While the fundamental causes of this resurgence remain the subject of much conjecture, the study of historical patterns of pathogen diffusion can be illuminating. Here, we examined time series of pertussis incidence in the boroughs of Greater London from 1982 to 2013 to document the spatial epidemiology of this bacterial infection and to identify the potential drivers of its percolation. The incidence of pertussis over this period is characterized by 3 distinct stages: a period exhibiting declining trends with 4-year inter-epidemic cycles from 1982 to 1994, followed by a deep trough until 2006 and the subsequent resurgence. We observed systematic temporal trends in the age distribution of cases and the fade-out profile of pertussis coincident with increasing national vaccine coverage from 1982 to 1990. To quantify the hierarchy of epidemic phases across the boroughs of London, we used the Hilbert transform. We report a consistent pattern of spatial organization from 1982 to the early 1990s, with some boroughs consistently leading epidemic waves and others routinely lagging. To determine the potential drivers of these geographic patterns, a comprehensive parallel database of borough-specific features was compiled, comprising of demographic, movement and socio-economic factors that were used in statistical analyses to predict epidemic phase relationships among boroughs. Specifically, we used a combination of a feed-forward neural network (FFNN), and SHapley Additive exPlanations (SHAP) values to quantify the contribution of each covariate to model predictions. Our analyses identified a number of predictors of a borough's historical epidemic phase, specifically the age composition of households, the number of agricultural and skilled manual workers, latitude, the population of public transport commuters and high-occupancy households. Univariate regression analysis of the 2012 epidemic identified the ratio of cumulative unvaccinated children to the total population and population of Pakistan-born population to have moderate positive and negative association, respectively, with the timing of epidemic. In addition to providing a comprehensive overview of contemporary pertussis transmission in a large metropolitan population, this study has identified the characteristics that determine the spatial spread of this bacterium across the boroughs of London. Arash Saeidpour, Shweta Bansal, Pejman Rohani |
PLoS Comput. Biol. | 2 |
| 2022 | Spatial clustering in vaccination hesitancy: The role of social influence and social selectionabstractThe phenomenon of vaccine hesitancy behavior has gained ground over the last three decades, jeopardizing the maintenance of herd immunity. This behavior tends to cluster spatially, creating pockets of unprotected sub-populations that can be hotspots for outbreak emergence. What remains less understood are the social mechanisms that can give rise to spatial clustering in vaccination behavior, particularly at the landscape scale. We focus on the presence of spatial clustering, and aim to mechanistically understand how different social processes can give rise to this phenomenon. In particular, we propose two hypotheses to explain the presence of spatial clustering: (i) social selection, in which vaccine-hesitant individuals share socio-demographic traits, and clustering of these traits generates spatial clustering in vaccine hesitancy; and (ii) social influence, in which hesitant behavior is contagious and spreads through neighboring societies, leading to hesitant clusters. Adopting a theoretical spatial network approach, we explore the role of these two processes in generating patterns of spatial clustering in vaccination behaviors under a range of spatial structures. We find that both processes are independently capable of generating spatial clustering, and the more spatially structured the social dynamics in a society are, the higher spatial clustering in vaccine-hesitant behavior it realizes. Together, we demonstrate that these processes result in unique spatial configurations of hesitant clusters, and we validate our models of both processes with fine-grain empirical data on vaccine hesitancy, social determinants, and social connectivity in the US. Finally, we propose, and evaluate the effectiveness of two novel intervention strategies to diminish hesitant behavior. Our generative modeling approach informed by unique empirical data provides insights on the role of complex social processes in driving spatial heterogeneity in vaccine hesitancy. Lucila G. Alvarez Zuzek, Casey M. Zipfel, Shweta Bansal |
PLoS Comput. Biol. | 3 |
| 2021 | Revealing mechanisms of infectious disease spread through empirical contact networksabstractThe spread of pathogens fundamentally depends on the underlying contacts between individuals. Modeling the dynamics of infectious disease spread through contact networks, however, can be challenging due to limited knowledge of how an infectious disease spreads and its transmission rate. We developed a novel statistical tool, INoDS (Identifying contact Networks of infectious Disease Spread) that estimates the transmission rate of an infectious disease outbreak, establishes epidemiological relevance of a contact network in explaining the observed pattern of infectious disease spread and enables model comparison between different contact network hypotheses. We show that our tool is robust to incomplete data and can be easily applied to datasets where infection timings of individuals are unknown. We tested the reliability of INoDS using simulation experiments of disease spread on a synthetic contact network and find that it is robust to incomplete data and is reliable under different settings of network dynamics and disease contagiousness compared with previous approaches. We demonstrate the applicability of our method in two host-pathogen systems: Crithidia bombi in bumblebee colonies and Salmonella in wild Australian sleepy lizard populations. INoDS thus provides a novel and reliable statistical tool for identifying transmission pathways of infectious disease spread. In addition, application of INoDS extends to understanding the spread of novel or emerging infectious disease, an alternative approach to laboratory transmission experiments, and overcoming common data-collection constraints. Pratha Sah, Michael C. Otterstatter, Stephan T. Leu, Sivan Leviyang, Shweta Bansal |
PLoS Comput. Biol. | 5 |
| 2021 | Health inequities in influenza transmission and surveillanceabstractThe lower an individual's socioeconomic position, the higher their risk of poor health in low-, middle-, and high-income settings alike. As health inequities grow, it is imperative that we develop an empirically-driven mechanistic understanding of the determinants of health disparities, and capture disease burden in at-risk populations to prevent exacerbation of disparities. Past work has been limited in data or scope and has thus fallen short of generalizable insights. Here, we integrate empirical data from observational studies and large-scale healthcare data with models to characterize the dynamics and spatial heterogeneity of health disparities in an infectious disease case study: influenza. We find that variation in social and healthcare-based determinants exacerbates influenza epidemics, and that low socioeconomic status (SES) individuals disproportionately bear the burden of infection. We also identify geographical hotspots of influenza burden in low SES populations, much of which is overlooked in traditional influenza surveillance, and find that these differences are most predicted by variation in susceptibility and access to sickness absenteeism. Our results highlight that the effect of overlapping factors is synergistic and that reducing this intersectionality can significantly reduce inequities. Additionally, health disparities are expressed geographically, and targeting public health efforts spatially may be an efficient use of resources to abate inequities. The association between health and socioeconomic prosperity has a long history in the epidemiological literature; addressing health inequities in respiratory-transmitted infectious disease burden is an important step towards social justice in public health, and ignoring them promises to pose a serious threat. Casey M. Zipfel, Vittoria Colizza, Shweta Bansal |
PLoS Comput. Biol. | 3 |
| 2018 | Deploying digital health data to optimize influenza surveillance at national and local scalesabstractThe surveillance of influenza activity is critical to early detection of epidemics and pandemics and the design of disease control strategies. Case reporting through a voluntary network of sentinel physicians is a commonly used method of passive surveillance for monitoring rates of influenza-like illness (ILI) worldwide. Despite its ubiquity, little attention has been given to the processes underlying the observation, collection, and spatial aggregation of sentinel surveillance data, and its subsequent effects on epidemiological understanding. We harnessed the high specificity of diagnosis codes in medical claims from a database that represented 2.5 billion visits from upwards of 120,000 United States healthcare providers each year. Among influenza seasons from 2002-2009 and the 2009 pandemic, we simulated limitations of sentinel surveillance systems such as low coverage and coarse spatial resolution, and performed Bayesian inference to probe the robustness of ecological inference and spatial prediction of disease burden. Our models suggest that a number of socio-environmental factors, in addition to local population interactions, state-specific health policies, as well as sampling effort may be responsible for the spatial patterns in U.S. sentinel ILI surveillance. In addition, we find that biases related to spatial aggregation were accentuated among areas with more heterogeneous disease risk, and sentinel systems designed with fixed reporting locations across seasons provided robust inference and prediction. With the growing availability of health-associated big data worldwide, our results suggest mechanisms for optimizing digital data streams to complement traditional surveillance in developed settings and enhance surveillance opportunities in developing countries. Elizabeth C. Lee, Ali Arab 0002, Sandra M. Goldlust, Cécile Viboud, Bryan T. Grenfell, Shweta Bansal |
PLoS Comput. Biol. | 6 |
| 2014 | Statistical Analysis of Multilingual Text Corpus and Development of Language Models
Shyam S. Agrawal, Abhimanue, Shweta Bansal, Minakshi Mahajan |
LREC | 3 |
| 2014 | Exploring community structure in biological networks with random graphsabstractBACKGROUND: Community structure is ubiquitous in biological networks. There has been an increased interest in unraveling the community structure of biological systems as it may provide important insights into a system's functional components and the impact of local structures on dynamics at a global scale. Choosing an appropriate community detection algorithm to identify the community structure in an empirical network can be difficult, however, as the many algorithms available are based on a variety of cost functions and are difficult to validate. Even when community structure is identified in an empirical system, disentangling the effect of community structure from other network properties such as clustering coefficient and assortativity can be a challenge. RESULTS: Here, we develop a generative model to produce undirected, simple, connected graphs with a specified degrees and pattern of communities, while maintaining a graph structure that is as random as possible. Additionally, we demonstrate two important applications of our model: (a) to generate networks that can be used to benchmark existing and new algorithms for detecting communities in biological networks; and (b) to generate null models to serve as random controls when investigating the impact of complex network features beyond the byproduct of degree and modularity in empirical biological networks. CONCLUSION: Our model allows for the systematic study of the presence of community structure and its impact on network function and dynamics. This process is a crucial step in unraveling the functional consequences of the structural properties of biological systems and uncovering the mechanisms that drive these systems. Pratha Sah, Lisa Singh, Aaron Clauset, Shweta Bansal |
BMC Bioinform. | 4 |
| 2009 | Exploring biological network structure with clustered random networksabstractBACKGROUND: Complex biological systems are often modeled as networks of interacting units. Networks of biochemical interactions among proteins, epidemiological contacts among hosts, and trophic interactions in ecosystems, to name a few, have provided useful insights into the dynamical processes that shape and traverse these systems. The degrees of nodes (numbers of interactions) and the extent of clustering (the tendency for a set of three nodes to be interconnected) are two of many well-studied network properties that can fundamentally shape a system. Disentangling the interdependent effects of the various network properties, however, can be difficult. Simple network models can help us quantify the structure of empirical networked systems and understand the impact of various topological properties on dynamics. RESULTS: Here we develop and implement a new Markov chain simulation algorithm to generate simple, connected random graphs that have a specified degree sequence and level of clustering, but are random in all other respects. The implementation of the algorithm (ClustRNet: Clustered Random Networks) provides the generation of random graphs optimized according to a local or global, and relative or absolute measure of clustering. We compare our algorithm to other similar methods and show that ours more successfully produces desired network characteristics.Finding appropriate null models is crucial in bioinformatics research, and is often difficult, particularly for biological networks. As we demonstrate, the networks generated by ClustRNet can serve as random controls when investigating the impacts of complex network features beyond the byproduct of degree and clustering in empirical networks. CONCLUSION: ClustRNet generates ensembles of graphs of specified edge structure and clustering. These graphs allow for systematic study of the impacts of connectivity and redundancies on network function and dynamics. This process is a key step in unraveling the functional consequences of the structural properties of empirical biological systems and uncovering the mechanisms that drive these systems. Shweta Bansal, Shashank Khandelwal, Lauren Ancel Meyers |
BMC Bioinform. | 1 |