David García 0001

dblp:37/3550 · also David García Becerra · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-2820-9151ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 R.U.Psycho? A Framework for Robust Unified Psychometric Testing of Language Models
Julian Schelb, Orr Borin, David García 0001, Andreas Spitz
LREC3
2025 Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models
abstract
Prompt-based language models like GPT4 and LLaMa have been used for a wide variety of use cases such as simulating agents, searching for information, or for content analysis.For all of these applications and others, political biases in these models can affect their performance.Several researchers have attempted to study political bias in language models using evaluation suites based on surveys, such as the Political Compass Test (PCT), often finding a particular leaning favored by these models.However, there is some variation in the exact prompting techniques, leading to diverging findings, and most research relies on constrained-answer settings to extract model responses.Moreover, the Political Compass Test is not a scientifically valid survey instrument.In this work, we contribute a political bias measured informed by political science theory, building on survey design principles to test a wide variety of input prompts, while taking into account prompt sensitivity.We then prompt 11 different open and commercial models, differentiating between instruction-tuned and non-instructiontuned models, and automatically classify their political stances from 88,110 responses.Leveraging this dataset, we compute political bias profiles across different prompt variations and find that while PCT exaggerates bias in certain models like GPT3.5, measures of political bias are often unstable, but generally more leftleaning for instruction-tuned models.Code and data are available on GitHub 1 .
Mats Faulborn, Indira Sen, Max Pellert, Andreas Spitz, David García 0001
ACL (1)5
2025 Extracting Affect Aggregates from Longitudinal Social Media Data with Temporal Adapters for Large Language Models
abstract
This paper proposes temporally aligned Large Language Models (LLMs) as a tool for longitudinal analysis of social media data. We fine-tune Temporal Adapters for Llama 3 8B on full timelines from a panel of British Twitter users and extract longitudinal aggregates of emotions and attitudes with established questionnaires. We focus our analysis on the beginning of the COVID-19 pandemic that had a strong impact on public opinion and collective emotions. We validate our estimates against representative British survey data and find strong positive and significant correlations for several collective emotions. The estimates obtained are robust across multiple training seeds and prompt formulations, and in line with collective emotions extracted using a traditional classification model trained on labeled data. We demonstrate the flexibility of our method on questions of public opinion for which no pre-trained classifier is available. Our work extends the analysis of affect in LLMs to a longitudinal setting through Temporal Adapters. It enables flexible and new approaches to the longitudinal analysis of social media data.
Georg Ahnert, Max Pellert, David García 0001, Markus Strohmaier
ICWSM3
2023 Just Another Day on Twitter: A Complete 24 Hours of Twitter Data
abstract
At the end of October 2022, Elon Musk concluded his acquisition of Twitter. In the weeks and months before that, several questions were publicly discussed that were not only of interest to the platform's future buyers, but also of high relevance to the Computational Social Science research community. For example, how many active users does the platform have? What percentage of accounts on the site are bots? And, what are the dominating topics and sub-topical spheres on the platform? In a globally coordinated effort of 80 scholars to shed light on these questions, and to offer a dataset that will equip other researchers to do the same, we have collected all 375 million tweets published within a 24-hour time period starting on September 21, 2022. To the best of our knowledge, this is the first complete 24-hour Twitter dataset that is available for the research community. With it, the present work aims to accomplish two goals. First, we seek to answer the aforementioned questions and provide descriptive metrics about Twitter that can serve as references for other researchers. Second, we create a baseline dataset for future research that can be used to study the potential impact of the platform's ownership change.
Jürgen Pfeffer, Daniel Matter, Kokil Jaidka, Onur Varol, Afra J. Mashhadi, Jana Lasser, Dennis Assenmacher, Diyi Yang, Cornelia Brantner, Daniel M. Romero, Jahna Otterbacher, Carsten Schwemmer, Kenneth Joseph, David García 0001, Fred Morstatter
ICWSM15
2023 This Sample Seems to Be Good Enough! Assessing Coverage and Temporal Reliability of Twitter's Academic API
abstract
Because of its willingness to share data with academia and industry, Twitter has been the primary social media platform for scientific research as well as for consulting businesses and governments in the last decade. In recent years, a series of publications have studied and criticized Twitter's APIs and Twitter has partially adapted its existing data streams. The newest Twitter API for Academic Research allows to "access Twitter's real-time and historical public data with additional features and functionality that support collecting more precise, complete, and unbiased datasets. The main new feature of this API is the possibility of accessing the full archive of all historic Tweets. In this article, we will take a closer look at the Academic API and will try to answer two questions. First, are the datasets collected with the Academic API complete? Secondly, since Twitter's Academic API delivers historic Tweets as represented on Twitter at the time of data collection, we need to understand how much data is lost over time due to Tweet and account removal from the platform. Our work shows evidence that Twitter's Academic API can indeed create (almost) complete samples of Twitter data based on a wide variety of search terms. We also provide evidence that Twitter's data endpoint v2 delivers better samples than the previously used endpoint v1.1. Furthermore, collecting Tweets with the Academic API at the time of studying a phenomenon rather than creating local archives of stored Tweets, allows for a straightforward way of following Twitter's developer agreement. Finally, we will also discuss technical artifacts and implications of the Academic API. We hope that our work can add another layer of understanding of Twitter data collections leading to more reliable studies of human behavior via social media data.
Jürgen Pfeffer, Angelina Voggenreiter, Jana Lasser, Luca Hammer, Oliver Stritzel, David García 0001
ICWSM6
2021 Unique on Facebook: formulation and evidence of (nano)targeting individual users with non-PII data
abstract
The privacy of an individual is bounded by the ability of a third party to reveal their identity. Certain data items such as a passport ID or a mobile phone number may be used to uniquely identify a person. These are referred to as Personal Identifiable Information (PII) items. Previous literature has also reported that, in datasets including millions of users, a combination of several non-PII items (which alone are not enough to identify an individual) can uniquely identify an individual within the dataset. In this paper, we define a data-driven model to quantify the number of interests from a user that make them unique on Facebook. To the best of our knowledge, this represents the first study of individuals' uniqueness at the world population scale. Besides, users' interests are actionable non-PII items that can be used to define ad campaigns and deliver tailored ads to Facebook users. We run an experiment through 21 Facebook ad campaigns that target three of the authors of this paper to prove that, if an advertiser knows enough interests from a user, the Facebook Advertising Platform can be systematically exploited to deliver ads exclusively to a specific user. We refer to this practice as nanotargeting. Finally, we discuss the harmful risks associated with nanotargeting such as psychological persuasion, user manipulation, or blackmailing, and provide easily implementable countermeasures to preclude attacks based on nanotargeting campaigns on Facebook.
José González Cabañas, Ángel Cuevas, Rubén Cuevas Rumín, Juan López-Fernández, David García 0001
Internet Measurement Conference5
2017 Bias in Online Freelance Marketplaces: Evidence from TaskRabbit and Fiverr
abstract
Online freelancing marketplaces have grown quickly in recent years. In theory, these sites offer workers the ability to earn money without the obligations and potential social biases associated with traditional employment frameworks. In this paper, we study whether two prominent online freelance marketplaces - TaskRabbit and Fiverr - are impacted by racial and gender bias. From these two platforms, we collect 13,500 worker profiles and gather information about workers' gender, race, customer reviews, ratings, and positions in search rankings. In both marketplaces, we find evidence of bias: we find that gender and race are significantly correlated with worker evaluations, which could harm the employment opportunities afforded to the workers. We hope that our study fuels more research on the presence and implications of discrimination in online environments.
Aniko Hannak, Claudia Wagner 0001, David García 0001, Alan Mislove, Markus Strohmaier, Christo Wilson
CSCW3
2017 A survey of multimodal sentiment analysis
Mohammad Soleymani 0001, David García 0001, Brendan Jou, Björn W. Schuller, Shih-Fu Chang, Maja Pantic
Image Vis. Comput.2
2016 Emotions, Demographics and Sociability in Twitter Interactions
Kristina Lerman, Megha Arora, Luciano Gallegos, Ponnurangam Kumaraguru, David García 0001
ICWSM5
2016 The QWERTY Effect on the Web: How Typing Shapes the Meaning of Words in Online Human-Computer Interaction
abstract
The QWERTY effect postulates that the keyboard layout influences word meanings by linking positivity to the use of the right hand and negativity to the use of the left hand. For example, previous research has established that words with more right hand letters are rated more positively than words with more left hand letters by human subjects in small scale experiments. In this paper, we perform large scale investigations of the QWERTY effect on the web. Using data from eleven web platforms related to products, movies, books, and videos, we conduct observational tests whether a hand-meaning relationship can be found in text interpretations by web users. Furthermore, we investigate whether writing text on the web exhibits the QWERTY effect as well, by analyzing the relationship between the text of online reviews and their star ratings in four additional datasets. Overall, we find robust evidence for the QWERTY effect both at the point of text interpretation (decoding) and at the point of text creation (encoding). We also find under which conditions the effect might not hold. Our findings have implications for any algorithmic method aiming to evaluate the meaning of words on the web, including for example semantic or sentiment analysis, and show the existence of "dactilar onomatopoeias" that shape the dynamics of word-meaning associations. To the best of our knowledge, this is the first work to reveal the extent to which the QWERTY effect exists in large scale human-computer interaction on the web.
David García 0001, Markus Strohmaier
WWW1
2015 It's a Man's Wikipedia? Assessing Gender Inequality in an Online Encyclopedia
Claudia Wagner 0001, David García 0001, Mohsen Jadidi, Markus Strohmaier
ICWSM2
2014 Gender Asymmetries in Reality and Fiction: The Bechdel Test of Social Media
David García 0001, Ingmar Weber, Venkata Rama Kiran Garimella
ICWSM1
2014 Who watches (and shares) what on youtube? and when?: using twitter to understand youtube viewership
abstract
By combining multiple social media datasets, it is possible to gain insight into each dataset that goes beyond what could be obtained with either individually. In this paper we combine user-centric data from Twitter with video-centric data from YouTube to build a rich picture of who watches and shares what on YouTube. We study 87K Twitter users, 5.6 million YouTube videos and 15 million video sharing events from user-, video- and sharing-event-centric perspectives. We show that features of Twitter users correlate with YouTube features and sharing-related features. For example, urban users are quicker to share than rural users. We find a superlinear relationship between initial Twitter shares and the final amounts of views. We discover that Twitter activity metrics play more role in video popularity than mere amount of followers. We also reveal the existence of correlated behavior concerning the time between video creation and sharing within certain timescales, showing the time onset for a coherent response, and the time limit after which collective responses are extremely unlikely. Response times depend on the category of the video, suggesting Twitter video sharing is highly dependent on the video content. To the best of our knowledge, this is the first large-scale study combining YouTube and Twitter data, and it reveals novel, detailed insights into who watches (and shares) what on YouTube, and when.
Adiya Abisheva, Venkata Rama Kiran Garimella, David García 0001, Ingmar Weber
WSDM3
2013 Damping Sentiment Analysis in Online Communication: Discussions, Monologs and Dialogs
Mike Thelwall, Kevan Buckley, Georgios Paltoglou, Marcin Skowron, David García 0001, Stéphane Gobron, Junghyun Ahn, Arvid Kappas, Dennis Küster, Janusz A. Holyst
CICLing (2)5