VLDB 2026 Research / reviewers in the wild / expert
Aron Culotta
dblp:17/551
· DBLP profile ↗
21ranked-venue papers in the field
2as first author
8since 2021 · last 2024
0000-0003-2660-7575ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 13 (1 first)Data Mining & Knowledge Discovery · 7 (1 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | How Does Empowering Users with Greater System Control Affect News Filter Bubbles?abstractWhile recommendation systems enable users to find articles of interest, they can also create "filter bubbles" by presenting content that reinforces users' pre-existing beliefs. Users are often unaware that the system placed them in a filter bubble and, even when aware, they often lack direct control over it. To address these issues, we first design a political news recommendation system augmented with an enhanced interface that exposes the political and topical interests the system inferred from user behavior. This allows the user to adjust the recommendation system to receive more articles on a particular topic or presenting a particular political stance. We then conduct a user study to compare our system to a traditional interface and found that the transparent approach helped users realize that they were in a filter bubble. Additionally, the enhanced system led to less extreme news for most users but also allowed others to move the system to more extremes. Similarly, while many users moved the system from extreme liberal/conservative to the center, this came at the expense of reducing political diversity of the articles shown. These findings suggest that, while the proposed system increased awareness of the filter bubbles, it had heterogeneous effects on news consumption depending on user preferences. Ping Liu 0002, Karthik Shivaram, Aron Culotta, Matthew A. Shapiro, Mustafa Bilgic 0001 |
ICWSM | 3 |
| 2024 | Characterizing Online Criticism of Partisan News Media Using Weakly Supervised LearningabstractWe propose novel methods to identify tweets that criticize partisan news sources. Prior work suggests that criticism, ridicule, and distrust of news media all play important roles in hyperpartisanship, misinformation, and filter bubble formation. Thus, understanding the prevalence and temporal dynamics of media-targeted criticism can provide us with updated tools to assess the health of the information ecosystem. There is a scarcity of labeled data for this task, and we develop a weakly supervised learning approach that leverages multiple noisy labeling functions based on both the content of the tweet as well as the historical news sharing behavior of the user. Using this classifier, we explore how tweets expressing criticism vary by user, news source, and time, finding substantial spikes in media criticism during politically polarizing events, such as the investigation into Russian interference in the 2016 U.S. elections and the 2017 "unite the right" rally in Charlottesville. This type of media-targeting criticism is also more likely to occur after users have been exposed to unreliable and hyperpartisan media. Karthik Shivaram, Mustafa Bilgic 0001, Matthew A. Shapiro, Aron Culotta |
ICWSM | 4 |
| 2024 | Forecasting Political News Engagement on Social MediaabstractUnderstanding how political news consumption changes over time can provide insights into issues such as hyperpartisanship, filter bubbles, and misinformation. To investigate long-term trends of news consumption, we curate a collection of over 60M tweets from politically engaged users over seven years, annotating ~10% with mentions of news outlets and their political leaning. We then train a neural network to forecast the political lean of news articles Twitter users will engage with, considering both past news engagements as well as tweet content. Using the learned representation of this model, we cluster users to discover salient patterns of long-term news engagement. Our findings include the following: (1) hyperpartisan users are more engaged with news; (2) right-leaning users engage with contra-partisan sources more than left-leaning users; (3) topics such as immigration, COVID-19, Islamaphobia, and gun control are salient indicators of engagement with low quality news sources. Karthik Shivaram, Mustafa Bilgic 0001, Matthew A. Shapiro, Aron Culotta |
ICWSM | 4 |
| 2023 | Online Reviews Are Leading Indicators of Changes in K-12 School AttributesabstractSchool rating websites are increasingly used by parents to assess the quality and fit of U.S. K-12 schools for their children. These online reviews often contain detailed descriptions of a school’s strengths and weaknesses, which both reflect and inform perceptions of a school. Existing work on these text reviews has focused on finding words or themes that underlie these perceptions, but has stopped short of using the textual reviews as leading indicators of school performance. In this paper, we investigate to what extent the language used in online reviews of a school is predictive of changes in the attributes of that school, such as its socio-economic makeup and student test scores. Using over 300K reviews of 70K U.S. schools from a popular ratings website, we apply language processing models to predict whether schools will significantly increase or decrease in an attribute of interest over a future time horizon. We find that using the text improves predictive performance significantly over a baseline model that does not include text but only the historical time-series of the indicators themselves, suggesting that the review text carries predictive power. A qualitative analysis of the most predictive terms and phrases used in the text reviews indicates a number of topics that serve as leading indicators, such as diversity, changes in school leadership, a focus on testing, and school safety. Linsen Li 0003, Aron Culotta, Douglas N. Harris, Nicholas Mattei |
WWW | 2 |
| 2022 | Leaders or Followers? A Temporal Analysis of Tweets from IRA Trolls
Siva K. Balasubramanian, Mustafa Bilgic 0001, Aron Culotta, Libby Hemphill, Anita Nikolich, Matthew A. Shapiro |
ICWSM | 3 |
| 2022 | Identifying Hurricane Evacuation Intent on Twitter
Xintian Li, Samiul Hasan, Aron Culotta |
ICWSM | 3 |
| 2022 | Reducing Cross-Topic Political Homogenization in Content-Based News RecommendationabstractContent-based news recommenders learn words that correlate with user engagement and recommend articles accordingly. This can be problematic for users with diverse political preferences by topic — e.g., users that prefer conservative articles on one topic but liberal articles on another. In such instances, recommenders can have a homogenizing effect by recommending articles with the same political lean on both topics, particularly if both topics share salient, politically polarized terms like “far right” or “radical left.” In this paper, we propose attention-based neural network models to reduce this homogenization effect by increasing attention on words that are topic specific while decreasing attention on polarized, topic-general terms. We find that the proposed approach results in more accurate recommendations for simulated users with such diverse preferences. Karthik Shivaram, Ping Liu 0002, Matthew A. Shapiro, Mustafa Bilgic 0001, Aron Culotta |
RecSys | 5 |
| 2021 | The Interaction between Political Typology and Filter Bubbles in News Recommendation AlgorithmsabstractAlgorithmic personalization of news and social media content aims to improve user experience; however, there is evidence that this filtering can have the unintended side effect of creating homogeneous “filter bubbles,” in which users are over-exposed to ideas that conform with their preexisting perceptions and beliefs. In this paper, we investigate this phenomenon in the context of political news recommendation algorithms, which have important implications for civil discourse. Ping Liu 0002, Karthik Shivaram, Aron Culotta, Matthew A. Shapiro, Mustafa Bilgic 0001 |
WWW | 3 |
| 2020 | Characterizing Variation in Toxic Language by Social Context
Bahar Radfar, Karthik Shivaram, Aron Culotta |
ICWSM | 3 |
| 2019 | Collecting representative social media samples from a search engine by adaptive query generationabstractStudies in computational social science often require collecting data about users via a search engine interface: a list of keywords is provided as a query to the interface and documents matching this query are returned. The validity of a study will hence critically depend on the representativeness of the data returned by the search engine. In this paper, we develop a multi-objective approach to build queries yielding documents that are both relevant to the study and representative of the larger population of documents. We then specify measures to evaluate the relevance and the representativeness of documents retrieved by a query system. Using these measures, we experiment on three real-world datasets and show that our method outperforms baselines commonly used to solve this data collection problem. Virgile Landeiro, Aron Culotta |
ASONAM | 2 |
| 2019 | Estimating tie strength in follower networks to measure brand perceptionsabstractAs public entities like brands and politicians increasingly rely on social media to engage their constituents, analyzing who follows them can reveal information about how they are perceived. Whereas most prior work considers following networks as unweighted directed graphs, in this paper we use a tie strength model to place weights on follow links to estimate the strength of relationship between users. We use conversational signals (retweets, mentions) as a proxy class label for a binary classification problem, using social and linguistic features to estimate tie strength. We then apply this approach to a case study estimating how brands are perceived with respect to certain issues (e.g., how environmentally friendly is Patagonia perceived to be?). We compute weighted follower overlap scores to measure the similarity between brands and exemplar accounts (e.g., environmental non-profits), finding that the tie strength scores can provide more nuanced estimates of consumer perception. Aron Culotta |
ASONAM | 3 |
| 2019 | Discovering and Controlling for Latent Confounds in Text Classification Using Adversarial Domain AdaptationabstractIn text classification, the testing data often systematically differ from the training data, a problem called dataset shift. In this paper, we investigate a type of dataset shift we call confounding shift. Such a setting exists when two conditions are met: (a) there is a confound variable Z that influences both text features X and class label Y; (b) the relationship between Z and Y changes from training to testing. While recent work in this area has required confounds to be known ahead of time, this is unrealistic for many settings. To address this shortcoming, we propose a method both to discover and to control for potential confounds. The approach first uses neural network-based topic modeling to discover potential confounds that differ between training and testing data, then uses adversarial training to fit a classification model that is invariant to these discovered confounds. We find the resulting method to improve over state-of-the-art domain adaptation method, while also producing results that are competitive with those obtained when confounds are known ahead of time. Virgile Landeiro, Aron Culotta |
SDM | 3 |
| 2018 | Forecasting the Presence and Intensity of Hostility on Instagram Using Linguistic and Social Features
Ping Liu 0002, Joshua Guberman, Libby Hemphill, Aron Culotta |
ICWSM | 4 |
| 2017 | Mining the Demographics of Political Sentiment from Twitter Using Learning from Label ProportionsabstractOpinion mining and demographic attribute inference have many applications in social science. In this paper, we propose models to infer daily joint probabilities of multiple latent attributes from Twitter data, such as political sentiment and demographic attributes. Since it is costly and time-consuming to annotate data for traditional supervised classification, we instead propose scalable Learning from Label Proportions (LLP) models for demographic and opinion inference using U.S. Census, national and state political polls, and Cook partisan voting index as population level data. In LLP classification settings, the training data is divided into a set of unlabeled bags, where only the label distribution of each bag is known, removing the requirement of instance-level annotations. Our proposed LLP model, Weighted Label Regularization (WLR), provides a scalable generalization of prior work on label regularization to support weights for samples inside bags, which is applicable in this setting where bags are arranged hierarchically (e.g., county-level bags are nested inside of state-level bags). We apply our model to Twitter data collected in the year leading up to the 2016 U.S. presidential election, producing estimates of the relationships among political sentiment and demographics over time and place. We find that our approach closely tracks traditional polling data stratified by demographic category, resulting in error reductions of 28-44% over baseline approaches. We also provide descriptive evaluations showing how the model may be used to estimate interactions among many variables and to identify linguistic temporal variation, capabilities which are typically not feasible using traditional polling methods. Ehsan Mohammady Ardehaly, Aron Culotta |
ICDM | 2 |
| 2017 | Identifying Leading Indicators of Product Recalls from Online Reviews Using Positive Unlabeled Learning and Domain Adaptation
Shreesh Kumara Bhat, Aron Culotta |
ICWSM | 2 |
| 2017 | Controlling for Unobserved Confounds in Classification Using Correlational Constraints
Virgile Landeiro, Aron Culotta |
ICWSM | 2 |
| 2017 | WSDM 2017 Workshop on Mining Online Health Reports: MOHRS 2017abstractThe workshop on Mining Online Health Reports (MOHRS) draws upon the rapidly developing field of Computational Health, focusing on textual content that has been generated through the various facets of Web activity. Online user-generated information mining, especially from social media platforms and search engines, has been in the forefront of many research efforts, especially in the fields of Information Retrieval and Natural Language Processing. The incorporation of such data and techniques in a number of health-oriented applications has provided strong evidence about the potential benefits, which include better population coverage, timeliness and the operational ability in places with less established health infrastructure. The workshop aims to create a platform where relevant state-of-the-art research is presented, but at the same time discussions among researchers with cross-disciplinary backgrounds can take place. It will focus on the characterisation of data sources, the essential methods for mining this textual information, as well as potential real-world applications and the arising ethical issues. MOHRS '17 will feature 3 keynote talks and 4 accepted paper presentations, together with a panel discussion session. Nigel Collier, Nut Limsopatham, Aron Culotta, Mike Conway, Ingemar J. Cox, Vasileios Lampos |
WSDM | 3 |
| 2015 | Data deidentification in medical transcriptions using regular expressions and machine learningabstractA system is developed to redact personally identifiable information (PII) through a combination of entity recognition, regular expressions, and machine learning with very high precision from millions of medical transcriptions. This system is trained and tested with manually redacted medical transcriptions using an internally developed coding system, providing double blind classification capabilities. Joshua Seeger, Aron Culotta, Jason Keller, Patrick van Kessel, Michael Jugovich |
IEEE BigData | 2 |
| 2009 | An Entity Based Model for Coreference ResolutionabstractRecently, many advanced machine learning approaches have been proposed for coreference resolution; however, all of the discriminatively-trained models reason over mentions rather than entities. That is, they do not explicitly contain variables indicating the “canonical” values for each attribute of an entity (e.g., name, venue, title, etc.). This canonicalization step is typically implemented as a post-processing routine to coreference resolution prior to adding the extracted entity to a database. In this paper, we propose a discriminatively-trained model that jointly performs coreference resolution and canonicalization, enabling features over hypothesized entities. We validate our approach on two different coreference problems: newswire anaphora resolution and research paper citation matching, demonstrating improvements in both tasks and achieving an error reduction of up to 62% when compared to a method that reasons about mentions only. Michael L. Wick, Aron Culotta, Khashayar Rohanimanesh, Andrew McCallum |
SDM | 2 |
| 2007 | Canonicalization of database records using adaptive similarity measuresabstractIt is becoming increasingly common to construct databases from information automatically culled from many heterogeneous sources. For example, a research publication database can be constructed by automatically extracting titles, authors, and conference information from online papers. A common difficulty in consolidating data from multiple sources is that records are referenced in a variety of ways (e.g. abbreviations, aliases, and misspellings). Therefore, it can be difficult to construct a single, standard representation to present to the user. We refer to the task of constructing this representation as canonicalization. Despite its importance, there is little existing work on canonicalization. Aron Culotta, Michael L. Wick, Rob Hall 0001, Matthew Marzilli, Andrew McCallum |
KDD | 1 |
| 2005 | Joint deduplication of multiple record types in relational dataabstractRecord deduplication is the task of merging database records that refer to the same underlying entity. In relational data-bases, accurate deduplication for records of one type is often dependent on the decisions made for records of other types. Whereas nearly all previous approaches have merged records of different types independently, this work models these inter-dependencies explicitly to collectively deduplicate records of multiple types. We construct a conditional random field model of deduplication that captures these relational dependencies, and then employ a novel relational partitioning algorithm to jointly deduplicate records. For two citation matching datasets, we show that collectively deduplicating paper and venue records results in up to a 30% error reduction in venue deduplication, and up to a 20% error reduction in paper deduplication. Aron Culotta, Andrew McCallum |
CIKM | 1 |