VLDB 2026 Research / reviewers in the wild / expert
Samuel P. Fraiberger
dblp:182/2204 · also Samuel Fraiberger
· DBLP profile ↗
8ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0003-3582-8978ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterabstractManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale, Samuel Fraiberger, Victor Orozco-Olvera, Paul Röttger. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Manuel Tonneau, Niyati Malhotra, Scott A. Hale, Samuel P. Fraiberger, Víctor Orozco-Olvera, Paul Röttger |
ACL (1) | 5 |
| 2024 | NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataabstractManuel Tonneau, Pedro Quinta De Castro, Karim Lasri, Ibrahim Farouq, Lakshmi Subramanian, Victor Orozco-Olvera, Samuel Fraiberger. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Manuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq, Lakshminarayanan Subramanian, Víctor Orozco-Olvera, Samuel P. Fraiberger |
ACL (1) | 7 |
| 2024 | 280 Characters to Employment: Using Twitter to Quantify Job VacanciesabstractAccurate assessment of workforce needs is critical for designing well-informed economic policy and improving market efficiency. While surveys are the gold standard for estimating when and where workers are needed, they also have important limitations, most notably their substantial costs, dependence on existing and extensive surveying infrastructure, and limited temporal, geographical, and sectorial resolution. Here, we investigate the potential of social media to provide a complementary signal for estimating labor market demand. We introduce a novel statistical approach for extracting information about the location and occupation advertised in job vacancies posted on Twitter. We then construct an aggregate index of labor market demand by occupational class in every major U.S. city from 2015 to 2022, which we evaluate against two sources of official statistics and an index from a large aggregator of online job postings. We find that the newly constructed index is strongly correlated with official statistics and, in some cases, advantageous compared to statistics from job aggregators. Moreover, we demonstrate that our index can robustly improve the prediction of official statistics across occupations and states. Boris Sobol, Manuel Tonneau, Samuel P. Fraiberger, Do Lee, Nir Grinberg |
ICWSM | 3 |
| 2023 | Large-Scale Demographic Inference of Social Media Users in a Low-Resource ScenarioabstractCharacterizing the demographics of social media users enables a diversity of applications, from better targeting of policy interventions to the derivation of representative population estimates of social phenomena. Achieving high performance with supervised learning, however, can be challenging as labeled data is often scarce. Alternatively, rule-based matching strategies provide well-grounded information but only offer partial coverage over users. It is unclear, therefore, what features and models are best suited to maximize coverage over a large set of users while maintaining high performance. In this paper, we develop a cost-effective strategy for large-scale demographic inference by relying on minimal labeling efforts. We combine a name-matching strategy with graph-based methods to map the demographics of 1.8 million Nigerian Twitter users. Specifically, we compare a purely graph-based propagation model, namely Label Propagation (LP), with Graph Convolutional Networks (GCN), a graph model that also incorporates node features based on user content. We find that both models largely outperform supervised learning approaches based purely on user content that lack graph information. Notably, we find that LP achieves comparable performance to the state-of-the-art GCN while providing greater interpretability at a lower computing cost. Moreover, performance does not significantly improve with the addition of user-specific features, such as textual representations of user tweets and user geolocation. Leveraging our data collection effort, we describe the demographic composition of Nigerian Twitter finding that it is a highly non-uniform sample of the general Nigerian population. Karim Lasri, Manuel Tonneau, Haaya Naushan, Niyati Malhotra, Ibrahim Farouq, Víctor Orozco-Olvera, Samuel P. Fraiberger |
ICWSM | 7 |
| 2022 | Multilingual Detection of Personal Employment Status on TwitterabstractDetecting disclosures of individuals' employment status on social media can provide valuable information to match job seekers with suitable vacancies, offer social protection, or measure labor market flows.However, identifying such personal disclosures is a challenging task due to their rarity in a sea of social media content and the variety of linguistic forms used to describe them.Here, we examine three Active Learning (AL) strategies in real-world settings of extreme class imbalance, and identify five types of disclosures about individuals' employment status (e.g.job loss) in three languages using BERT-based classification models.Our findings show that, even under extreme imbalance settings, a small number of AL iterations is sufficient to obtain large and significant gains in precision, recall, and diversity of results compared to a supervised baseline with the same number of labels.We also find that no AL strategy consistently outperforms the rest.Qualitative analysis suggests that AL helps focus the attention mechanism of BERT on core terms and adjust the boundaries of semantic expansion, highlighting the importance of interpretable models to provide greater control and visibility into this dynamic learning process. Manuel Tonneau, Dhaval Adjodah, João Palotti, Nir Grinberg, Samuel P. Fraiberger |
ACL (1) | 5 |
| 2022 | Targeted Policy Recommendations using Outcome-aware ClusteringabstractPolicy recommendations using observational data typically rely on estimating an econometric model on a sample of observations drawn from an entire population. However, different policy actions could potentially be optimal for different subgroups of a population. In this paper, we propose outcome-aware clustering, a new methodology to segment a population into different clusters and derive cluster-level policy recommendations. Outcome-aware clustering differs from conventional clustering algorithms across two basic dimensions. First, given a specific outcome of interest, outcome-aware clustering segments the population based on selecting a small set of features that closely relate with the outcome variable. Second, the clustering algorithm aims to generate near-homogeneous clusters based on a combination of cluster size-balancing constraints, inter and intra-cluster distances in the reduced feature space. We generate targeted policy recommendations for each outcome-aware cluster based on a standard multivariate regression of a condensed set of actionable policy features (which may partially overlap or differ from the features used for segmentation) from the observational data. We implement our outcome-aware clustering method on the Living Standards Measurement Study - Integrated Surveys on Agriculture (LSMS-ISA) dataset to generate targeted policy recommendations for improving farmers outcomes in sub-Saharan Africa. Based on a detailed analysis of the LSMS-ISA, we derive outcome-aware clusters of farmer populations across three sub-Saharan African countries and show that the targeted policy recommendations at the cluster level significantly differ from policies that are generated at the population level. Ananth Balashankar, Samuel P. Fraiberger, Eric Deregt, Marelize Gorgens, Lakshminarayanan Subramanian |
COMPASS | 2 |
| 2019 | Reconstructing the MERS disease outbreak from newsabstractDisease surveillance is critical for mobilizing health care resources and deciding on isolation measures to contain the spread of infectious diseases. Because ground truth signals of rare and deadly diseases are sparse, it can be useful to enrich surveillance systems using measures of social and environmental factors which are known to influence the spread of a disease. One approach to measure such factors is by using real time news streams. In this study, we model the epidemiological transmission of the Middle Eastern Respiratory Syndrome (MERS) disease during the outbreak that occurred from 2013 to 2018 in the Arabian peninsula. Using the GDELT news event database, we show that conflict related signals allow us to reconstruct the time series of newly infected cases per week. This reduces the residual sum of squared errors by a factor of 3.36 as compared to a standard epidemiological model. We also capture interpretable time-sensitive factors which illustrate the importance of using real time news stream to model the evolution of a disease such as MERS and facilitate early and effective policy interventions. Ananth Balashankar, Aashish Dugar, Lakshminarayanan Subramanian, Samuel P. Fraiberger |
COMPASS | 4 |
| 2019 | Identifying Predictive Causal Factors from News StreamsabstractAnanth Balashankar, Sunandan Chakraborty, Samuel Fraiberger, Lakshminarayanan Subramanian. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ananth Balashankar, Sunandan Chakraborty, Samuel P. Fraiberger, Lakshminarayanan Subramanian |
EMNLP/IJCNLP (1) | 3 |