VLDB 2026 Research / reviewers in the wild / expert
Eni Mustafaraj
dblp:97/592 · also Eniana Mustafaraj
· DBLP profile ↗
15ranked-venue papers
4as first author
4since 2021 · last 2024
0000-0003-2243-5892ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Tracing the Evolution of Information Transparency for OpenAI's GPT Models through a Biographical ApproachabstractInformation transparency, the open disclosure of information about models, is crucial for proactively evaluating the potential societal harm of large language models (LLMs) and developing effective risk mitigation measures. Adapting the biographies of artifacts and practices (BOAP) method from science and technology studies, this study analyzes the evolution of information transparency within OpenAI’s Generative Pre-trained Transformers (GPT) model reports and usage policies from its inception in 2018 to GPT-4, one of today’s most capable LLMs. To assess the breadth and depth of transparency practices, we develop a 9-dimensional, 3-level analytical framework to evaluate the comprehensiveness and accessibility of information disclosed to various stakeholders. Findings suggest that while model limitations and downstream usages are increasingly clarified, model development processes have become more opaque. Transparency remains minimal in certain aspects, such as model explainability and real-world evidence of LLM impacts, and the discussions on safety measures such as technical interventions and regulation pipelines lack in-depth details. The findings emphasize the need for enhanced transparency to foster accountability and ensure responsible technological innovations. Zhihan Xu, Eni Mustafaraj |
AIES (1) | 2 |
| 2024 | Online News Coverage of Critical Race Theory Controversies: A Dataset of Annotated HeadlinesabstractIn this paper, we introduce an annotated dataset of 11,704 unique U.S. news headlines related to critical race theory and its controversies from August 2020 through December 2022. Annotations generated by GPT-4 specify the headline stance and the primary actor in the headline. GPT-4 annotations performed well on the validation dataset, with weighted average F-scores of 0.8339 for headline stance annotations and 0.7625 for primary actor annotations. Along with the annotated headlines and URLs to the full article, we augment the dataset with metrics that are relevant to future research on political polarization, news frame analysis, and regional news coverage. The dataset includes partisan audience bias scores by news source domain, tags for mentions of U.S. states in the article body, and exposure and engagement metrics for articles shared on Reddit. Among other preliminary descriptive analyses, we find that the most prevalent headline stance in our headlines dataset is anti-CRT (43.06%), and the most prevalent primary actor in our headlines dataset is political influencers (56.56%). This paper describes the data collection methodology, preliminary descriptive analysis, and possible uses of the dataset for future research in political science, computational social sciences, and natural language processing. Our dataset and replication code is available to access on Zenodo at zenodo.org/doi/10.5281/zenodo.10516190 Anna Lieb, Maneesh Arora, Eni Mustafaraj |
ICWSM | 3 |
| 2023 | Assessing Google Search's New Features in Supporting Credibility judgments of Unknown WebsitesabstractThis study assesses the awareness and perceived utility of two features Google Search introduced in February 2021: “About this result” and “More about this page”. Google stated that the goal of these features is to help users vet unfamiliar web domains (or sources). We investigated whether the features were sufficiently prominent to be detected by frequent users of Google Search, and their perceived utility for making credibility judgments of sources, in one-on-one user studies with 25 undergraduate college students, who identify as frequent users of Google Search. Our results indicate a lack of adoption or awareness of these features by our participants and neutral-positive perceptions of their utility in evaluating web sources. We also examined the perceived usefulness of nine other domain credibility signals collected from the W3C. Ace Wang, Liz Maylin De Jesus Sanchez, Anya Wintner, Yuanxin Zhu, Eni Mustafaraj |
CHIIR | 5 |
| 2023 | Capturing the Aftermath of the Dobbs v. Jackson Women's Health Organization Decision in Google Search Results across the U.SabstractHow do Google Search results change following an impactful real-world event, such as the U.S. Supreme Court decision on June 24, 2022 to overturn Roe v. Wade? And what do they tell us about the nature of event-driven content, generated by various participants in the online information environment? In this paper, we present a dataset of more than 1.74 million Google Search results pages collected between June 24 and July 17, 2022, intended to capture what Google Search surfaced in response to queries about this event of national importance. These search pages were collected for 65 locations in 13 U.S. states, a mix of red, blue, and purple states, with respect to their voting patterns. We describe the process of building a set of circa 1,700 phrases used for searching Google, how we gathered the search results for each location, and how these results were parsed to extract information about the most frequently encountered web domains. We believe that this dataset, which comprises raw data (search results as HTML files) and processed data (extracted links organized as CSV files) can be used to answer research questions that are of interest to computational social scientists as well as communication and media studies scholars. Brooke Perreault, Lan Dau, Anya Wintner, Eni Mustafaraj |
ICWSM | 4 |
| 2020 | The Media Coverage of the 2020 US Presidential Election Candidates through the Lens of Google's Top Stories
Anna Kawakami, Khonzoda Umarova, Eni Mustafaraj |
ICWSM | 3 |
| 2016 | HABITAT EXPLORER: Designing Educational Games for Collaborative Learning on Interactive SurfacesabstractData science is an interdisciplinary field at the intersection of computer science, mathematics, and subject matter expertise with the purpose to gain insights from relatively large sets of data. With elementary school age children as the target audience, we developed an educational collaborative game, Habitat Explorer for the MultiTaction display to introduce users to the core data science cycle components of data collection, exploration, and visualization. Users capture "sea creatures" in a collection jar, sort their collection into a data table, and then are able to create a bar chart visualization using their data. These components are interspersed with questions assessing user comprehension of the simple biological and data science concepts introduced during the game. The main goal of this research is to understand the potential of multi-touch displays to facilitate data science education with future aims to develop more complicated data exploration applications. Anne Schwartz, Clara Sorensen, Eni Mustafaraj |
ISS | 3 |
| 2015 | What Do Retweets Indicate? Results from User Survey and Meta-Review of Research
Panagiotis Takis Metaxas, Eni Mustafaraj, Kily Wong, Laura Zeng, Megan O'Keefe, Samantha Finn |
ICWSM | 2 |
| 2015 | The Visible and Invisible in a MOOC Discussion ForumabstractOnline discussion forums in a MOOC setting allow students to become aware of other students enrolled in the course. However, what is (usually) visible in the forums is the output of ``active'' students who engage in asking and answering questions. In addition to such active participants, there is (as always in online communities) a large group of ``passive'' users (so-called lurkers), who might find the forum useful to their learning, and read it regularly, despite remaining ``invisible''. Our analysis of a large MOOC online forum shows that for every active participant in the forum there are two passive ones. 30% of active participants complete the course, compared to only 6.6% of the passive participants. Vice-versa, 67% of students who complete the course are also active in the forum. However, ``invisible activity'' (e.g. reading or searching the forum) is something that both groups practice equally and more frequently, while only 3.3% of forum actions are visible. Eni Mustafaraj, Jessica Bu |
L@S | 1 |
| 2014 | What does enrollment in a MOOC mean?abstractIn 2012, when MOOCs became largely known, media reports were fascinated with the big number of enrollments. The number 150,000 students was mentioned for both Stanford's Artificial Intelligence course and MIT's Circuits and Electronics, to be later followed by the underwhelming completion rates, that often are in the single digit percentages. But what kind of enrollment do these large numbers really show? We try to answer this question by breaking this number into its components, while comparing two successive iterations of the same MOOC offered on the edX platform. Eni Mustafaraj |
L@S | 1 |
| 2014 | The Co-retweeted Network and Its Applications for Measuring the Perceived Political PolarizationabstractThis paper introduces a novel network, the co-retweeted network, that is constructed as the undirected weighted graph that connects highly visible accounts who have been retweeted by members of the audience during some real-time event. Like bibliographics co-citation used to indicate that two papers treat a related subject matter, co-retweeting is used to indicate that two accounts present similar opinions in an online discussion. Thus, the co-retweeted network can be seen as a form of consulting the opinion of the crowd that is following the discussion about the similarity (or difference) of positions expressed by the highly visible accounts. When applied on political conversations related to some event, the co-retweeted network enables the measurement of the polarity of political orientation of major players (including news organizations) based on the views of the audience. It can also measure the degree of polarization of the event itself. Samantha Finn, Eni Mustafaraj, Panagiotis Takis Metaxas |
WEBIST (1) | 2 |
| 2013 | Panel: mobile application development in computing curriculaabstractMany institutions are considering offering a course on mobile application development to harness its popularity to attract new majors, retain those we have, and to motivate learning. The panelists present four experiences in teaching a mobile application development course. They share their experiences in an effort to start a discussion about mobile application development in computing curricula. In the first half of the session, each panelist presents their experience including: an overview of the course; its audience, position in the curriculum, and pre-requisites; the platform, language, and development environment used; positives about the course; and roadblocks and negatives about the course. This provides a foundation for an audience directed discussion in the second half. Stoney Jackson, Stanislav Kurkovsky, Eni Mustafaraj, Lori Postner |
SIGCSE | 3 |
| 2012 | Hiding in Plain Sight: A Tale of Trust and Mistrust inside a Community of Citizen Reporters
Eni Mustafaraj, Panagiotis Takis Metaxas, Samantha Finn, Andrés Monroy-Hernández |
ICWSM | 1 |
| 2011 | Can Collective Sentiment Expressed on Twitter Predict Political Elections?abstractResearch examining the predictive power of social media (especially Twitter) displays conflicting results, particularly in the domain of political elections. This paper applies methods used in studies that have shown a direct correlation between volume/sentiment of Twitter chatter and future electoral results in a new dataset about political elections. We show that these methods display a series of shortcomings that make them inadequate for determining whether social media messages can predict the outcome of elections. Jessica Elan Chung, Eni Mustafaraj |
AAAI | 2 |
| 2011 | Limits of Electoral Predictions Using Twitter
Daniel Gayo-Avello, Panagiotis Takis Metaxas, Eni Mustafaraj |
ICWSM | 3 |
| 2007 | Knowledge Extraction and Summarization for an Application of Textual Case-Based Interpretation
Eni Mustafaraj, Martin Hoof, Bernd Freisleben |
ICCBR | 1 |