VLDB 2026 Research / reviewers in the wild / expert
Sen Pei
dblp:129/1503
· DBLP profile ↗
17ranked-venue papers
6as first author
14since 2021 · last 2024
0000-0002-7072-2995ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Exploring Domain Incremental Video Highlights Detection with the LiveFood BenchmarkabstractVideo highlights detection (VHD) is an active research field in computer vision, aiming to locate the most user-appealing clips given raw video inputs. However, most VHD methods are based on the closed world assumption, i.e., a fixed number of highlight categories is defined in advance and all training data are available beforehand. Consequently, existing methods have poor scalability with respect to increasing highlight domains and training data. To address above issues, we propose a novel video highlights detection method named Global Prototype Encoding (GPE) to learn incrementally for adapting to new domains via parameterized prototypes. To facilitate this new research direction, we collect a finely annotated dataset termed LiveFood, including over 5,100 live gourmet videos that consist of four domains: ingredients, cooking, presentation, and eating. To the best of our knowledge, this is the first work to explore video highlights detection in the incremental learning setting, opening up new land to apply VHD for practical scenarios where both the concerned highlight domains and training data increase over time. We demonstrate the effectiveness of GPE through extensive experiments. Notably, GPE surpasses popular domain incremental learning methods on LiveFood, achieving significant mAP improvements on all domains. Concerning the classic datasets, GPE also yields comparable performance as previous arts. The code is available at: https://github.com/ForeverPs/IncrementalVHD_GPE. Sen Pei, Shixiong Xu, Xiaojie Jin 0004 |
AAAI | 1 |
| 2024 | Image Background Serves as Good Proxy for Out-of-distribution DataabstractOut-of-distribution (OOD) detection empowers the model trained on the closed image set to identify unknown data in the open world. Though many prior techniques have yielded considerable improvements in this research direction, two crucial obstacles still remain. Firstly, a unified perspective has yet to be presented to view the developed arts with individual designs, which is vital for providing insights into future work. Secondly, we expect sufficient natural OOD supervision to promote the generation of compact boundaries between the in-distribution (ID) and OOD data without collecting explicit OOD samples. To tackle these issues, we propose a general probabilistic framework to interpret many existing methods and an OOD-data-free model, namely $\textbf{S}$elf-supervised $\textbf{S}$ampling for $\textbf{O}$OD $\textbf{D}$etection (SSOD). SSOD efficiently exploits natural OOD signals from the ID data based on the local property of convolution. With these supervisions, it jointly optimizes the OOD detection and conventional ID classification in an end-to-end manner. Extensive experiments reveal that SSOD establishes competitive state-of-the-art performance on many large-scale benchmarks, outperforming the best previous method by a large margin, e.g., reporting $\textbf{-6.28}$% FPR95 and $\textbf{+0.77}$% AUROC on ImageNet, $\textbf{-19.01}$% FPR95 and $\textbf{+3.04}$% AUROC on CIFAR-10, and top-ranked performance on hard OOD datasets, i.e., ImageNet-O and OpenImage-O. Sen Pei |
ICLR | 1 |
| 2024 | epiDAMIK 2024: The 7th International Workshop on Epidemiology meets Data Mining and Knowledge DiscoveryabstractWhile the worst of COVID-19 pandemic has most likely passed us, an occurrence of equally devastating global pandemic or regional epidemic cannot be ruled out in future. H1N1, Zika, SARS, MERS, and Ebola outbreaks over the past few decades have sharply illustrated our enormous vulnerability to emerging infectious diseases. While the data mining research community has demonstrated increased interest in epidemiological applications, much is still left to be desired. For example, there is an urgent need to develop sound theoretical principles and transformative computational approaches that will allow us to address the escalating threat of current and future pandemics. Data mining and knowledge discovery have an important role to play in this regard. Different aspects of infectious disease modeling, analysis, and control have traditionally been studied within the confines of individual disciplines, such as mathematical epidemiology and public health, and data mining and machine learning. Coupled with increasing data generation across multiple domains/sources (e.g., wastewater surveillance, electronic medical records, and social media), there is a clear need for analyzing them to inform public health policies and outcomes timely. Recent advances in disease surveillance and forecasting, and initiatives such as the CDC Flu Challenge, CDC COVID-19 Forecasting Hub etc., have brought these disciplines closer together. On the one hand, public health practitioners seek to use novel datasets, such as Safegraph, Unacast, and Google mobility data, and techniques like Graph Neural Networks. On the other hand, researchers from data mining and machine learning develop novel tools for solving many fundamental problems in the public health policy planning and decision-making process, leveraging novel datasets (e.g., COVID-19 behavioral health surveys, contact tracing trees, and satellite images of urban streets) and combining them with more traditional time series information (e.g., surveillance, hospitalization, and death records). We believe the next stage of advances will result from closer collaborations between these two groups, which is the main objective of epiDAMIK. Alexander Rodríguez, Bijaya Adhikari, Ajitesh Srivastava, Sen Pei, Marie-Laure Charpignon, Kai Wang 0040, Serina Chang, Anil Vullikanti, B. Aditya Prakash |
KDD | 4 |
| 2024 | Challenges of COVID-19 Case Forecasting in the US, 2020-2021abstractDuring the COVID-19 pandemic, forecasting COVID-19 trends to support planning and response was a priority for scientists and decision makers alike. In the United States, COVID-19 forecasting was coordinated by a large group of universities, companies, and government entities led by the Centers for Disease Control and Prevention and the US COVID-19 Forecast Hub (https://covid19forecasthub.org). We evaluated approximately 9.7 million forecasts of weekly state-level COVID-19 cases for predictions 1-4 weeks into the future submitted by 24 teams from August 2020 to December 2021. We assessed coverage of central prediction intervals and weighted interval scores (WIS), adjusting for missing forecasts relative to a baseline forecast, and used a Gaussian generalized estimating equation (GEE) model to evaluate differences in skill across epidemic phases that were defined by the effective reproduction number. Overall, we found high variation in skill across individual models, with ensemble-based forecasts outperforming other approaches. Forecast skill relative to the baseline was generally higher for larger jurisdictions (e.g., states compared to counties). Over time, forecasts generally performed worst in periods of rapid changes in reported cases (either in increasing or decreasing epidemic phases) with 95% prediction interval coverage dropping below 50% during the growth phases of the winter 2020, Delta, and Omicron waves. Ideally, case forecasts could serve as a leading indicator of changes in transmission dynamics. However, while most COVID-19 case forecasts outperformed a naïve baseline model, even the most accurate case forecasts were unreliable in key phases. Further research could improve forecasts of leading indicators, like COVID-19 cases, by leveraging additional real-time data, addressing performance across phases, improving the characterization of forecast confidence, and ensuring that forecasts were coherent across spatial scales. In the meantime, it is critical for forecast users to appreciate current limitations and use a broad set of indicators to inform pandemic-related decision making. Velma K. Lopez, Estee Y. Cramer, Robert Pagano, John M. Drake, Eamon B. O'Dea, Madeline Adee, Turgay Ayer, Jagpreet Chhatwal, Ozden O. Dalgic, Mary A. Ladd, Benjamin P. Linas, Peter P. Mueller, Jade Xiao, Johannes Bracher, Alvaro J. Castro Rivadeneira, Aaron Gerding, Tilmann Gneiting, Yuxin Huang 0009, Dasuni Jayawardena, Abdul H. Kanji, Khoa Le, Anja Mühlemann, Jarad Niemi, Evan L. Ray, Ariane Stark, Nutcha Wattanachit, Martha W. Zorn, Sen Pei, Jeffrey Shaman, Teresa K. Yamana, Samuel R. Tarasewicz, Daniel J. Wilson 0002, Sid Baccam, Heidi Gurung, Steve Stage, Brad Suchoski, Lei Gao 0011, Zhiling Gu, Myungjin Kim, Guannan Wang, Li Wang 0035, Yueying Wang, Lauren Gardner, Sonia Jindal, Maximilian Marshall, Kristen Nixon, Juan Dent, Alison L. Hill, Joshua Kaminsky, Elizabeth C. Lee, Joseph Chadi Lemaitre, Justin Lessler, Claire P. Smith, Shaun Truelove, Matt Kinsey, Luke C. Mullany, Kaitlin Rainwater-Lovett, Lauren Shin, Katharine Tallaksen, Shelby Wilson, Dean Karlen, Lauren A. Castro, Geoffrey Fairchild, Isaac Michaud, Dave Osthus, Jiang Bian 0002, Wei Cao 0007, Zhifeng Gao, Juan M. Lavista Ferres, Chaozhuo Li, Tie-Yan Liu, Xing Xie 0001, Shun Zheng 0001, Matteo Chinazzi, Jessica T. Davis, Kunpeng Mu, Ana L. Pastore y Piontti, Alessandro Vespignani, Xinyue Xiong, Robert Walraven, Quanquan Gu, Lingxiao Wang 0001, Pan Xu 0002, Difan Zou, Graham Casey Gibson, Daniel Sheldon, Ajitesh Srivastava, Aniruddha Adiga, Benjamin Hurt, Gursharn Kaur, Bryan L. Lewis, Madhav V. Marathe, Akhil Sai Peddireddy, Przemyslaw J. Porebski, Srinivasan Venkatramanan, Lijing Wang 0001, Pragati V. Prasad, Jo W. Walker, Alexander E. Webber, Rachel B. Slayton, Matthew Biggerstaff, Nicholas G. Reich, Michael A. Johansson |
PLoS Comput. Biol. | 29 |
| 2023 | Domain Decorrelation with Potential Energy RankingabstractMachine learning systems, especially the methods based on deep learning, enjoy great success in modern computer vision tasks under ideal experimental settings. Generally, these classic deep learning methods are built on the i.i.d. assumption, supposing the training and test data are drawn from the same distribution independently and identically. However, the aforementioned i.i.d. assumption is, in general, unavailable in the real-world scenarios, and as a result, leads to sharp performance decay of deep learning algorithms. Behind this, domain shift is one of the primary factors to be blamed. In order to tackle this problem, we propose using Potential Energy Ranking (PoER) to decouple the object feature and the domain feature in given images, promoting the learning of label-discriminative representations while filtering out the irrelevant correlations between the objects and the background. PoER employs the ranking loss in shallow layers to make features with identical category and domain labels close to each other and vice versa. This makes the neural networks aware of both objects and background characteristics, which is vital for generating domain-invariant features. Subsequently, with the stacked convolutional blocks, PoER further uses the contrastive loss to make features within the same categories distribute densely no matter domains, filtering out the domain information progressively for feature alignment. PoER reports superior performance on domain generalization benchmarks, improving the average top-1 accuracy by at least 1.20% compared to the existing methods. Moreover, we use PoER in the ECCV 2022 NICO Challenge, achieving top place with only a vanilla ResNet-18 and winning the jury award. The code has been made publicly available at: https://github.com/ForeverPs/PoER. Sen Pei, Jiaxi Sun, Shiming Xiang, Gaofeng Meng |
AAAI | 1 |
| 2023 | epiDAMIK 6.0: The 6th International Workshop on Epidemiology meets Data Mining and Knowledge DiscoveryabstractThe epiDAMIK workshop serves as a platform for advancing the utilization of data-driven methods in the fields of epidemiology and public health research. These fields have seen relatively limited exploration of data-driven approaches compared to other disciplines. Therefore, our primary objective is to foster the growth and recognition of the emerging discipline of data-driven and computational epidemiology, providing a valuable avenue for sharing state-of-the-art research and ongoing projects. The workshop also seeks to showcase results that are not typically presented at major computing conferences, including valuable insights gained from practical experiences. Our target audience encompasses researchers in AI, machine learning, and data science from both academia and industry, who have a keen interest in applying their work to epidemiological and public health contexts. Additionally, we welcome practitioners from mathematical epidemiology and public health, as their expertise and contributions greatly enrich the discussions. Homepage: https://epidamik.github.io/ Bijaya Adhikari, Alexander Rodríguez, Amulya Yadav, Sen Pei, Ajitesh Srivastava, Marie-Laure Charpignon, Anil Vullikanti, B. Aditya Prakash |
KDD | 4 |
| 2023 | Inference of transmission dynamics and retrospective forecast of invasive meningococcal diseaseabstractThe pathogenic bacteria Neisseria meningitidis, which causes invasive meningococcal disease (IMD), predominantly colonizes humans asymptomatically; however, invasive disease occurs in a small proportion of the population. Here, we explore the seasonality of IMD and develop and validate a suite of models for simulating and forecasting disease outcomes in the United States. We combine the models into multi-model ensembles (MME) based on the past performance of the individual models, as well as a naive equally weighted aggregation, and compare the retrospective forecast performance over a six-month forecast horizon. Deployment of the complete vaccination regimen, introduced in 2011, coincided with a change in the periodicity of IMD, suggesting altered transmission dynamics. We found that a model forced with the period obtained by local power wavelet decomposition best fit and forecast observations. In addition, the MME performed the best across the entire study period. Finally, our study included US-level data until 2022, allowing study of a possible IMD rebound after relaxation of non-pharmaceutical interventions imposed in response to the COVID-19 pandemic; however, no evidence of a rebound was found. Our findings demonstrate the ability of process-based models to retrospectively forecast IMD and provide a first analysis of the seasonality of IMD before and after the complete vaccination regimen. Jaime Cascante-Vega, Marta Galanti, Katharina Schley, Sen Pei, Jeffrey Shaman |
PLoS Comput. Biol. | 4 |
| 2023 | Ensemble inference of unobserved infections in networks using partial observationsabstractUndetected infections fuel the dissemination of many infectious agents. However, identification of unobserved infectious individuals remains challenging due to limited observations of infections and imperfect knowledge of key transmission parameters. Here, we use an ensemble Bayesian inference method to infer unobserved infections using partial observations. The ensemble inference method can represent uncertainty in model parameters and update model states using all ensemble members collectively. We perform extensive experiments in both model-generated and real-world networks in which individuals have differential but unknown transmission rates. The ensemble method outperforms several alternative approaches for a variety of network structures and observation rates, despite that the model is mis-specified. Additionally, the computational complexity of this algorithm scales almost linearly with the number of nodes in the network and the number of observations, respectively, exhibiting the potential to apply to large-scale networks. The inference method may support decision-making under uncertainty and be adapted for use for other dynamical models in networks. Renquan Zhang, Jilei Tai, Sen Pei |
PLoS Comput. Biol. | 3 |
| 2022 | Out-of-distribution Detection with Boundary Aware Learning
Sen Pei, Xin Zhang 0093, Bin Fan 0001, Gaofeng Meng |
ECCV (24) | 1 |
| 2022 | epiDAMIK 5.0: The 5th International Workshop on Epidemiology meets Data Mining and Knowledge DiscoveryabstractSimilar to previous iterations, the epiDAMIK @ KDD workshop is a forum to promote data driven approaches in epidemiology and public health research. Even after the devastating impact of COVID-19 pandemic, data driven approaches are not as widely studied in epidemiology, as they are in other spaces. We aim to promote and raise the profile of the emerging research area of data-driven and computational epidemiology, and create a venue for presenting state-of-the-art and in-progress results-in particular, results that would otherwise be difficult to present at a major data mining conference, including lessons learnt in the 'trenches'. The current COVID-19 pandemic has only showcased the urgency and importance of this area. Our target audience consists of data mining and machine learning researchers from both academia and industry who are interested in epidemiological and public-health applications of their work, and practitioners from the areas of mathematical epidemiology and public health. Homepage: https://epidamik.github.io/. Bijaya Adhikari, Amulya Yadav, Sen Pei, Ajitesh Srivastava, Sarah Kefayati, Alexander Rodríguez, Marie-Laure Charpignon, Anil Vullikanti, B. Aditya Prakash |
KDD | 3 |
| 2022 | An ensemble forecast system for tracking dynamics of dengue outbreaks and its validation in ChinaabstractAs a common vector-borne disease, dengue fever remains challenging to predict due to large variations in epidemic size across seasons driven by a number of factors including population susceptibility, mosquito density, meteorological conditions, geographical factors, and human mobility. An ensemble forecast system for dengue fever is first proposed that addresses the difficulty of predicting outbreaks with drastically different scales. The ensemble forecast system based on a susceptible-infected-recovered (SIR) type of compartmental model coupled with a data assimilation method called the ensemble adjusted Kalman filter (EAKF) is constructed to generate real-time forecasts of dengue fever spread dynamics. The model was informed by meteorological and mosquito density information to depict the transmission of dengue virus among human and mosquito populations, and generate predictions. To account for the dramatic variations of outbreak size in different seasons, the effective population size parameter that is sequentially updated to adjust the predicted outbreak scale is introduced into the model. Before optimizing the transmission model, we update the effective population size using the most recent observations and historical records so that the predicted outbreak size is dynamically adjusted. In the retrospective forecast of dengue outbreaks in Guangzhou, China during the 2011-2017 seasons, the proposed forecast model generates accurate projections of peak timing, peak intensity, and total incidence, outperforming a generalized additive model approach. The ensemble forecast system can be operated in real-time and inform control planning to reduce the burden of dengue fever. Yuliang Chen, Tao Liu 0078, Xiaolin Yu, Qinghui Zeng, Zixi Cai, Haisheng Wu, Qingying Zhang, Jianpeng Xiao, Wenjun Ma, Sen Pei, Pi Guo |
PLoS Comput. Biol. | 10 |
| 2022 | Epidemic management and control through risk-dependent individual contact interventionsabstractTesting, contact tracing, and isolation (TTI) is an epidemic management and control approach that is difficult to implement at scale because it relies on manual tracing of contacts. Exposure notification apps have been developed to digitally scale up TTI by harnessing contact data obtained from mobile devices; however, exposure notification apps provide users only with limited binary information when they have been directly exposed to a known infection source. Here we demonstrate a scalable improvement to TTI and exposure notification apps that uses data assimilation (DA) on a contact network. Network DA exploits diverse sources of health data together with the proximity data from mobile devices that exposure notification apps rely upon. It provides users with continuously assessed individual risks of exposure and infection, which can form the basis for targeting individual contact interventions. Simulations of the early COVID-19 epidemic in New York City are used to establish proof-of-concept. In the simulations, network DA identifies up to a factor 2 more infections than contact tracing when both harness the same contact data and diagnostic test data. This remains true even when only a relatively small fraction of the population uses network DA. When a sufficiently large fraction of the population (≳ 75%) uses network DA and complies with individual contact interventions, targeting contact interventions with network DA reduces deaths by up to a factor 4 relative to TTI. Network DA can be implemented by expanding the computational backend of existing exposure notification apps, thus greatly enhancing their capabilities. Implemented at scale, it has the potential to precisely and effectively control future epidemics while minimizing economic disruption. Tapio Schneider, Oliver R. A. Dunbar, Lucas Böttcher, Dmitry Burov, Alfredo Garbuno-Inigo, Gregory LeClaire Wagner, Sen Pei, Chiara Daraio, Raffaele Ferrari, Jeffrey Shaman |
PLoS Comput. Biol. | 8 |
| 2021 | The 4th International Workshop on Epidemiology meets Data Mining and Knowledge Discovery (epiDAMIK 4.0 @ KDD2021)abstractThe 4th [email protected] workshop is a forum to discuss new insights into how data mining can play a bigger role in epidemiology and public health research. While the integration of data science methods into epidemiology has significant potential, it remains under studied. We aim to raise the profile of this emerging research area of data-driven and computational epidemiology, and create a venue for presenting state-of-the-art and in-progress results-in particular, results that would otherwise be difficult to present at a major data mining conference, including lessons learnt in the 'trenches'. The current COVID-19 pandemic has only showcased the urgency and importance of this area. Our target audience consists of data mining and machine learning researchers from both academia and industry who are interested in epidemiological and public-health applications of their work, and practitioners from the areas of mathematical epidemiology and public health. Bijaya Adhikari, Ajitesh Srivastava, Sen Pei, Sarah Kefayati, Rose Yu, Amulya Yadav, Alexander Rodríguez, Arvind Ramanathan, Anil Vullikanti, B. Aditya Prakash |
KDD | 3 |
| 2021 | StoCast: Stochastic Disease Forecasting With Progression UncertaintyabstractForecasting patients' disease progressions with rich longitudinal clinical data has drawn much attention in recent years due to its impactful application in healthcare and the medical field. Researchers have tackled this problem by leveraging traditional machine learning, statistical techniques and deep learning based models. However, existing methods suffer from either deterministic internal structures or over-simplified stochastic components, failing to deal with complex uncertain scenarios such as progression uncertainty (i.e., multiple possible trajectories) and data uncertainty (i.e., imprecise observations and misdiagnosis). To overcome these major uncertainty issues, we propose a novel deep generative model, Stochastic Disease Forecasting Model (StoCast), along with an associated neural network architecture StoCastNet, that can be trained efficiently via stochastic optimization techniques. Our StoCast model uses internal stochastic components to deal with departures of observed data from patients' true health states, and more importantly, is able to produce a comprehensive estimate of future disease progression trajectories. Based on two public datasets related to Alzheimer's disease and Parkinson's disease, we demonstrate how our StoCast model achieves robust and superior performance than deterministic baseline approaches, and conveys richer information that can potentially assist doctors to make decisions with greater confidence in a complex uncertain scenario. Xian Teng, Sen Pei, Yu-Ru Lin |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Forecasting influenza in Europe using a metapopulation model incorporating cross-border commuting and air travelabstractPast work has shown that models incorporating human travel can improve the quality of influenza forecasts. Here, we develop and validate a metapopulation model of twelve European countries, in which international translocation of virus is driven by observed commuting and air travel flows, and use this model to generate influenza forecasts in conjunction with incidence data from the World Health Organization. We find that, although the metapopulation model fits the data well, it offers no improvement over isolated models in forecast quality. We discuss several potential reasons for these results. In particular, we note the need for data that are more comparable from country to country, and offer suggestions as to how surveillance systems might be improved to achieve this goal. Sarah C. Kramer, Sen Pei, Jeffrey Shaman |
PLoS Comput. Biol. | 2 |
| 2020 | Aggregating forecasts of multiple respiratory pathogens supports more accurate forecasting of influenza-like illnessabstractInfluenza-like illness (ILI) is a commonly measured syndromic signal representative of a range of acute respiratory infections. Reliable forecasts of ILI can support better preparation for patient surges in healthcare systems. Although ILI is an amalgamation of multiple pathogens with variable seasonal phasing and attack rates, most existing process-based forecasting systems treat ILI as a single infectious agent. Here, using ILI records and virologic surveillance data, we show that ILI signal can be disaggregated into distinct viral components. We generate separate predictions for six contributing pathogens (influenza A/H1, A/H3, B, respiratory syncytial virus, and human parainfluenza virus types 1-2 and 3), and develop a method to forecast ILI by aggregating these predictions. The relative contribution of each pathogen to the total ILI signal is estimated using a Markov Chain Monte Carlo (MCMC) method upon forecast aggregation. We find highly variable overall contributions from influenza type A viruses across seasons, but relatively stable contributions for the other pathogens. Using historical data from 1997 to 2014 at US national and regional levels, the proposed forecasting system generates improved predictions of both seasonal and near-term targets relative to a baseline method that simulates ILI as a single pathogen. The hierarchical forecasting system can generate predictions for each viral component, as well as infer and predict their contributions to ILI, which may additionally help physicians determine the etiological causes of ILI in clinical settings. Sen Pei, Jeffrey Shaman |
PLoS Comput. Biol. | 1 |
| 2019 | Predictability in process-based ensemble forecast of influenzaabstractProcess-based models have been used to simulate and forecast a number of nonlinear dynamical systems, including influenza and other infectious diseases. In this work, we evaluate the effects of model initial condition error and stochastic fluctuation on forecast accuracy in a compartmental model of influenza transmission. These two types of errors are found to have qualitatively similar growth patterns during model integration, indicating that dynamic error growth, regardless of source, is a dominant component of forecast inaccuracy. We therefore examine the nonlinear growth of model initial error and compute the fastest growing directions using singular vector analysis. Using this information, we generate perturbations in an ensemble forecast system of influenza to obtain more optimal ensemble spread. In retrospective forecasts of historical outbreaks for 95 US cities from 2003 to 2014, this approach improves short-term forecast of incidence over the next one to four weeks. Sen Pei, Mark A. Cane, Jeffrey Shaman |
PLoS Comput. Biol. | 1 |