Alex Berke

dblp:241/7229 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-5996-0557ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 4 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Uncovering Relationships Between Android Developers, User Privacy, and Developer Willingness to Reduce Fingerprinting Risks
abstract
The major mobile platforms, Android and iOS, have introduced changes that restrict user tracking to improve user privacy, yet apps continue to covertly track users via device fingerprinting. We study the opportunity to improve this dynamic with a case study on mobile fingerprinting that evaluates developers’ perceptions of how well platforms protect user privacy and how developers perceive platform privacy interventions. Specifically, we study developers’ willingness to make changes to protect users from fingerprinting and how developers consider trade-offs between user privacy and developer effort. We do this via a survey of 246 Android developers, presented with a hypothetical Android change that protects users from fingerprinting at the cost of additional developer effort.
Alex Berke, Güliz Seray Tuncay, Michael A. Specter, Mihai Christodorescu
CHI1
2026 An Improved Entropy Measure for Web Browser Fingerprinting Risk
abstract
Browser fingerprinting is the practice of tracking users across the Web by collecting attributes from their devices and combining them to create unique identifiers. This practice poses major privacy risks to users, and more than a decade of research has quantified fingerprinting risks due to various attributes, leading browser developers to implement many privacy-enhancing changes. Early work used Shannon entropy to quantify risks. However, Shannon entropy can grow with dataset size, limiting the ability to compare datasets and results. Researchers then introduced normalized entropy as a measure for comparing browser fingerprinting datasets of different sizes and numerous works followed using normalized entropy for this purpose. We identify and address a resulting problem in the fingerprinting literature. We show normalized entropy is ill-suited to compare datasets of different sizes --- it decreases as dataset size increases. We show this both analytically and empirically, leveraging a recently published dataset of browser attributes commonly used for fingerprinting. Given the unmet need for a better fingerprinting risk measure, we define a minimal set of desired properties for such a measure: scale-invariance, monotonicity and estimability. We then propose to use Tsallis entropy as a more interpretable fingerprinting risk measure. We evaluate Shannon, normalized, and Tsallis entropy with respect to the properties, and prove that only Tsallis entropy satisfies all of them.
Alex Berke, Enrico Bacis, Umar Syed
Proc. Priv. Enhancing Technol.1
2025 How Unique is Whose Web Browser? The role of demographics in browser fingerprinting among US users
abstract
Browser fingerprinting can be used to identify and track users across the Web, even without cookies, by collecting attributes from users' devices to create unique 'fingerprints'. This technique and resulting privacy risks have been studied for over a decade. Yet further research is limited because prior studies used data not publicly available. Additionally, data in prior studies lacked user demographics. Here we provide a first-of-its-kind dataset to enable further research. It includes browser attributes with users' demographics and survey responses, collected with informed consent from 8,400 US study participants. We use this dataset to demonstrate how fingerprinting risks differ across demographic groups. For example, we find lower income users are more at risk, and find that as users' age increases, they are both more likely to be concerned about fingerprinting and at real risk of fingerprinting. Furthermore, we demonstrate an overlooked risk: user demographics, such as gender, age, income level and race, can be inferred from browser attributes commonly used for fingerprinting, and we identify which browser attributes most contribute to this risk. Our data collection process also conducted an experiment to study what impacts users' likelihood to share browser data for open research, in order to inform future data collection efforts, with responses from 12,461 total participants. Female participants were significantly less likely to share their browser data, as were participants who were shown the browser data we asked to collect. Overall, we show the important role of user demographics in the ongoing work that intends to assess fingerprinting risks and improve user privacy, with findings to inform future privacy enhancing browser developments. The dataset and data collection tool we provide can be used to further study research questions not addressed in this work.
Alex Berke, Badih Ghazi, Enrico Bacis, Pritish Kamath, Ravi Kumar 0001, Robin Lassonde, Pasin Manurangsi, Umar Syed
Proc. Priv. Enhancing Technol.1
2024 Poster: zkTax: A Pragmatic Way to Support Zero-Knowledge Tax Disclosures
abstract
Tax returns contain financial information of interest to third parties: public officials are asked to share financial data for transparency, companies seek to assess the financial status of business partners, and individuals need to prove their income to third-parties.Tax returns also contain sensitive data such that sharing them in their entirety undermines privacy.We outline how zero-knowledge cryptography may be applied to address this tension by allowing individuals and organizations to make provable claims about select information in their tax returns without revealing additional information, in a way that can be independently verified by third parties.We highlight key system goals and design specifications for this zero-knowledge tax disclosure system (zkTax) and present a prototype implementation.The prototype consists of three distinct services that can be distributed: a tax authority that provides signed tax documents; a Redact & Prove Service that enables users to redact tax documents and produce a zero-knowledge proof attesting the provenance of the redacted data; and a Verify Service to check the validity of claims.We demonstrate how zkTax could be implemented with minimal changes to existing tax infrastructure, allowing the system to be extensible to other contexts and jurisdictions.This work provides a practical example of how distributed tools leveraging cryptography can enhance existing government or financial infrastructures, providing immediate transparency alongside privacy without system overhauls.
Alex Berke, Tobin South, Robert Mahari, Kent Larson, Alex Pentland
CCS1
2024 Insights from an Experiment Crowdsourcing Data from Thousands of US Amazon Users: The importance of transparency, money, and data use
abstract
Data generated by users on digital platforms are a crucial resource for advocates and researchers interested in uncovering digital inequities, auditing algorithms, and understanding human behavior. Yet data access is often restricted. How can researchers both effectively and ethically collect user data? This paper shares an innovative approach to crowdsourcing user data to collect otherwise inaccessible Amazon purchase histories, spanning 5 years, from more than 5,000 U.S. users. We developed a data collection tool that prioritizes participant consent and includes an experimental study design. The design allows us to study multiple important aspects of privacy perception and user data sharing behavior, including how socio-demographics, monetary incentives and transparency can impact share rates. Experiment results (N=6,325) reveal both monetary incentives and transparency can significantly increase data sharing. Age, race, education, and gender also played a role, where female and less-educated participants were more likely to share. Our study design enables a unique empirical evaluation of the 'privacy paradox', where users claim to value their privacy more than they do in practice. We set up both real and hypothetical data sharing scenarios and find measurable similarities and differences in share rates across these contexts. For example, increasing monetary incentives had a 6 times higher impact on share rates in real scenarios. In addition, we study participants' opinions on how data should be used by various third parties, again finding that gender, age, education, and race have a significant impact. Notably, the majority of participants disapproved of government agencies using purchase data yet the majority approved of use by researchers. Overall, our findings highlight the critical role that transparency, incentive design, and user demographics play in ethical data collection practices, and provide guidance for future researchers seeking to crowdsource user generated data.
Alex Berke, Robert Mahari, Alex Pentland, Kent Larson, Dana Calacci
Proc. ACM Hum. Comput. Interact.1
2022 Privacy Limitations of Interest-based Advertising on The Web: A Post-mortem Empirical Analysis of Google's FLoC
abstract
In 2020, Google announced it would disable third-party cookies in the Chrome browser to improve user privacy. In order to continue to enable interest-based advertising while mitigating risks of individualized user tracking, Google proposed FLoC. The FLoC algorithm assigns users to "cohorts" that represent groups of users with similar browsing behaviors so that ads can be served to users based on their cohort. In 2022, after testing FLoC in a real world trial, Google canceled the proposal with little explanation in favor of another way to enable interest-based advertising. This work provides a post-mortem analysis of two critical privacy risks for FloC by applying an implementation of FLoC to a real-world browsing history dataset collected from over 90,000 U.S. devices over a one year period.
Alex Berke, Dana Calacci
CCS1
2022 Mobility and COVID-19 in Andorra: Country-Scale Analysis of High-Resolution Mobility Patterns and Infection Spread
abstract
Throughout the COVID-19 pandemic, nonpharmaceutical interventions, such as mobility restrictions, have been globally adopted as critically important strategies to curb the spread of infection. However, such interventions come with immense social and economic costs and the relative effectiveness of different mobility restrictions are not well understood. Some recent works have used telecoms data sources that cover fractions of a population to understand behavioral changes and how these changes have impacted case growth. This study analyzed uniquely comprehensive datasets in order to examine the relationship between mobility and transmission of COVID-19 in the country of Andorra. The data consisted of spatio-temporal telecoms data for all mobile subscribers in the country, serology screening results for 91% of the population, and COVID-19 case reports. A comprehensive set of mobility metrics was developed using the telecoms data to indicate entrances to the country, contact with tourists, stay-at-home rates, trip-making and levels of crowding. Mobility metrics were compared to infection rates across communities and transmission rate over time. All metrics dropped sharply at the start of the country's lockdown and gradually rose again as the restrictions were gradually lifted. Several of these metrics were highly correlated with lagged transmission rate. There was a stronger correlation for measures of indoor crowding and inter-community trip-making, and a weaker correlation for total trips (including intra-community trips) and stay-at-homes rates. These findings provide support for policies which aim to discourage gathering indoors while lifting the most restrictive mobility limitations.
Ronan Doorley, Alex Berke, Ariel Noyman, Josep Ribó, Vanesa Arroyo, Marc Pons 0002, Kent Larson
IEEE J. Biomed. Health Informatics2