VLDB 2026 Research / reviewers in the wild / expert
Afra J. Mashhadi
dblp:71/7815 · also Afra Jahanbakhsh Mashhadi, Afra Mashhadi
· DBLP profile ↗
25ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0003-4631-4438ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 13 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Computer networks · 5 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How Does Memorization Impact LLMs' Social Reasoning? An Assessment using Seen and Unseen QueriesabstractAs Large Language Models (LLMs) have rapidly advanced in social reasoning tasks, their applications have expanded to domains such as healthcare and psychology. Given the direct interaction of users with these applications, it is essential to evaluate the performance of LLMs, particularly in human-like social reasoning capabilities. While previous studies have explored human-aligned social reasoning in LLMs, they have not adequately assessed whether the generated reasoning answers stem from the LLMs' memorization of training data or their natural language understanding. In this study, we aim to address this gap by assessing the impact of training data memorization on the human-aligned social reasoning capabilities of LLMs. We introduce IR+CoT (Information Retrieval (IR) + Chain of Thought (CoT)), a framework that leverages retrieved information from input questions to fine-tune prompt templates and employs CoT methods. IR+CoT mitigates the effects of memorization and enhances the LLMs' social reasoning performance. Experiments on three LLMs, using seen (present during the training of the LLMs) and unseen (introduced post-training) questions from Reddit and Lemmy, show that IR+CoT enhances social reasoning and reduces memorization effects. This research's novelty lies in using old and new questions to assess memorization's impact on social reasoning. Maryam Amirizaniani, Maryna Sivachenko, Adrian Lavergne, Chirag Shah 0001, Afra J. Mashhadi |
WSDM | 5 |
| 2024 | A Model- and Data-Agnostic Debiasing System for Achieving Equalized OddsabstractAs reliance on Machine Learning (ML) systems in real-world decision-making processes grows, ensuring these systems are free of bias against sensitive demographic groups is of increasing importance. Existing techniques for automatically debiasing ML models generally require access to either the models’ internal architectures, the models’ training datasets, or both. In this paper we outline the reasons why such requirements are disadvantageous, and present an alternative novel debiasing system that is both data- and model-agnostic. We implement this system as a Reinforcement Learning Agent and through extensive experiments show that we can debias a variety of target ML model architectures over three benchmark datasets. Our results show performance comparable to data- and/or model-gnostic state-of-the-art debiasers. Thomas Pinkava, Jack W. McFarland, Afra J. Mashhadi |
AIES (1) | 3 |
| 2024 | Can LLMs Reason Like Humans? Assessing Theory of Mind Reasoning in LLMs for Open-Ended QuestionsabstractTheory of mind (ToM) reasoning involves understanding that others have intentions, emotions, and thoughts, which is crucial for regulating one's reasoning. Although large language models (LLMs) excel in tasks such as summarization, question answering, and translation, they still face challenges with ToM reasoning, especially in open-ended questions. Despite advancements, the extent to which LLMs truly understand ToM reasoning and how closely it aligns with human ToM reasoning remains inadequately explored in open-ended scenarios. Motivated by this gap, we assess the abilities of LLMs to perceive and integrate human intentions and emotions into their ToM reasoning processes within open-ended questions. Our study utilizes posts from Reddit's ChangeMyView platform, which demands nuanced social reasoning to craft persuasive responses. Our analysis, comparing semantic similarity and lexical overlap metrics between responses generated by humans and LLMs, reveals clear disparities in ToM reasoning capabilities in open-ended questions, with even the most advanced models showing notable limitations. To enhance LLM capabilities, we implement a prompt tuning method that incorporates human intentions and emotions, resulting in improvements in ToM reasoning performance. However, despite these improvements, the enhancement still falls short of fully achieving human-like reasoning. This research highlights the deficiencies in LLMs' social reasoning and demonstrates how integrating human intentions and emotions can boost their effectiveness. Maryam Amirizaniani, Elias Martin, Maryna Sivachenko, Afra J. Mashhadi, Chirag Shah 0001 |
CIKM | 4 |
| 2023 | Just Another Day on Twitter: A Complete 24 Hours of Twitter DataabstractAt the end of October 2022, Elon Musk concluded his acquisition of Twitter. In the weeks and months before that, several questions were publicly discussed that were not only of interest to the platform's future buyers, but also of high relevance to the Computational Social Science research community. For example, how many active users does the platform have? What percentage of accounts on the site are bots? And, what are the dominating topics and sub-topical spheres on the platform? In a globally coordinated effort of 80 scholars to shed light on these questions, and to offer a dataset that will equip other researchers to do the same, we have collected all 375 million tweets published within a 24-hour time period starting on September 21, 2022. To the best of our knowledge, this is the first complete 24-hour Twitter dataset that is available for the research community. With it, the present work aims to accomplish two goals. First, we seek to answer the aforementioned questions and provide descriptive metrics about Twitter that can serve as references for other researchers. Second, we create a baseline dataset for future research that can be used to study the potential impact of the platform's ownership change. Jürgen Pfeffer, Daniel Matter, Kokil Jaidka, Onur Varol, Afra J. Mashhadi, Jana Lasser, Dennis Assenmacher, Diyi Yang, Cornelia Brantner, Daniel M. Romero, Jahna Otterbacher, Carsten Schwemmer, Kenneth Joseph, David García 0001, Fred Morstatter |
ICWSM | 5 |
| 2023 | Privacy-Aware Adversarial Network in Human Mobility PredictionabstractAs mobile devices and location-based services are increasingly developed in different smart city scenarios and applications, many unexpected privacy leakages have arisen due to geolocated data collection and sharing. User re-identification and other sensitive inferences are major privacy threats when geolocated data are shared with cloud-assisted applications. Significantly, four spatio-temporal points are enough to uniquely identify 95% of the individuals, which exacerbates personal information leakages. To tackle malicious purposes such as user re-identification, we propose an LSTM-based adversarial mechanism with representation learning to attain a privacy-preserving feature representation of the original geolocated data (i.e., mobility data) for a sharing purpose. These representations aim to maximally reduce the chance of user re-identification and full data reconstruction with a minimal utility budget (i.e., loss). We train the mechanism by quantifying privacy-utility trade-off of mobility datasets in terms of trajectory reconstruction risk, user re-identification risk, and mobility predictability. We report an exploratory analysis that enables the user to assess this trade-off with a specific loss function and its weight parameters. The extensive comparison results on four representative mobility datasets demonstrate the superiority of our proposed architecture in mobility privacy protection and the efficiency of the proposed privacy-preserving features extractor. We show that the privacy of mobility traces attains decent protection at the cost of marginal mobility utility. Our results also show that by exploring the Pareto optimal setting, we can simultaneously increase both privacy (45%) and utility (32%). Yuting Zhan, Hamed Haddadi 0001, Afra J. Mashhadi |
Proc. Priv. Enhancing Technol. | 3 |
| 2022 | CloudFL: A Zero-Touch Federated Learning Framework for Privacy-aware Sensor CloudabstractIntelligent sensing solutions bridge the gap between the physical world and the cyber-physical systems by digitizing the sensor data collected from sensor devices. Sensor cloud networks provide physical and virtual sensing device resources and enable uninterrupted intelligent solutions to end-users. Thanks to advancements in machine learning algorithms and big data, the automation of mundane tasks with artificial intelligence is becoming a reliable smart option. However, existing approaches based on centralized Machine Learning (ML) on sensor cloud networks fail to ensure data privacy. Moreover, centralized ML works with the pre-requisite to transfer the entire training dataset from end devices to a central server. To address this, we propose a Quantized Federated Learning (FL) based approach, called CloudFL, to ensure data privacy on end devices in a sensor cloud network. Our framework enables a personalized version of FL implementation and enhances privacy and security with cryptosystem tools to obfuscate the information of the FL process from unauthorized access. Furthermore, microservices of our approach provide software as a service implementation of FL with instances of cloud servers that require zero-touch on local data for training. Viraaji Mothukuri, Reza M. Parizi, Seyed Amin Pouriyeh, Afra J. Mashhadi |
ARES | 4 |
| 2022 | Causal Impact Model to Evaluate the Diffusion Effect of Social Media Campaigns
Xinchen Yu, Afra J. Mashhadi, Jeremy Boy, René Clausen Nielsen, Lingzi Hong |
ECSCW | 2 |
| 2022 | Quantifying fairness of federated learning LPPM modelsabstractDespite the great potential offered by Artificial Intelligence in the context of smart mobility, it comes with the greater challenge of preserving the privacy of users. Federated Learning (FL) has gained popularity as a privacy-friendly approach, however, an equally important aspect rarely addressed in the literature, is its fairness. In this work we audit a FL-based privacy-preserving model. We use Entropy to determine similarity within the system's input data and compare its value against that of the output to detect unfair treatment. Amina Ben Salem, Besma Khalfoun, Sonia Ben Mokhtar, Afra J. Mashhadi |
MobiSys | 4 |
| 2021 | No Walk in the Park: The Viability and Fairness of Social Media Analysis for Parks and Recreational Policy Making
Afra J. Mashhadi, Samantha G. Winder, Emilia H. Lia, Spencer A. Wood |
ICWSM | 1 |
| 2021 | Deep Embedded Clustering of Urban Communities Using Federated LearningabstractDeep clustering utilizes representation learning to learn features in an unsupervised setting. Although successful, the current models rely on the assumption of the centralized dataset, which due to the privacy concerns is becoming less realistic. To address this challenge, we propose a federated deep convolutional embedded clustering framework. Our framework relies on a federated server to orchestrate the training between workers where each participant individually trains the model with the objective of decreasing clustering loss using Kullback–Leibler divergence. To avoid feature space being distorted by the clustering loss, each worker maintains their own local decoder which for privacy reasons is not shared with the federated server. Empirical results with both IID and non-IID client data on benchmark datasets demonstrates the feasibility of our federated training when compared to the centralized counterpart. We also evaluate our model on a real world application of community detection using GPS traces and measure the computational complexity and energy consumption on a smartphone. Afra J. Mashhadi, Joshua Sterner, Jeffrey Murray |
IJCNN | 1 |
| 2021 | Caring Without Sharing: A Federated Learning Crowdsensing Framework for Diversifying Representation of Cities
Michael Cho, Afra J. Mashhadi |
MobiQuitous | 2 |
| 2016 | Exploring space syntax on entrepreneurial opportunities with Wi-Fi analyticsabstractIndustrial events and exhibitions play a powerful role in creating social relations amongst individuals and firms, enabling them to expand their social network so to acquire resources. However, often these events impose a spatial structure which impacts encounter opportunities. In this paper, we study the impact that the spatial configuration has on the formation of network relations. We designed, developed and deployed a Wi-Fi analytics solution comprising of wearable Wi-Fi badges and gateways in a large scale industrial exhibition event to study the spatio-temporal trajectories of the 2.5K+ attendees including two special groups: 34 investors and 27 entrepreneurs. Our results suggest that certain zones with designated functionalities play a key role in forming social ties across attendees and the different behavioural properties of investors and entrepreneurs can be explained through a spatial lens. Based on our findings we offer three concrete recommendations for future organisers of networking events. Afra J. Mashhadi, Utku Günay Acer, Aidan Boran, Philipp M. Scholl, Claudio Forlivesi, Geert Vanderhulst, Fahim Kawsar |
UbiComp | 1 |
| 2016 | Understanding the impact of personal feedback on face-to-face interactions in the workplaceabstractFace-to-face interactions have proven to accelerate team and larger organisation success. Many past research has explored the benefits of quantifying face-to-face interactions for informed workplace management, however to date, little attention has been paid to understand how the feedback on interaction behaviour is perceived at a personal scale. In this paper, we offer a reflection on the automated feedback of personal interactions in a workplace through a longitudinal study. We designed and developed a mobile system that captured, modelled, quantified and visualised face-to-face interactions of 47 employees for 4 months in an industrial research lab in Europe. Then we conducted semi-structured interviews with 20 employees to understand their perception and experience with the system. Our findings suggest that the short-term feedback on personal face-to-face interactions was not perceived as an effective external cue to promote self-reflection and that employees desire long-term feedback annotated with actionable attributes. Our findings provide a set of implications for the designers of future workplace technology and also opens up avenues for future HCI research on promoting self-reflection among employees. Afra J. Mashhadi, Akhil Mathur, Marc Van den Broeck, Geert Vanderhulst, Fahim Kawsar |
ICMI | 1 |
| 2016 | Phantom cascades: The effect of hidden nodes on information diffusion
Václav Belák, Afra J. Mashhadi, Alessandra Sala, Donn Morrison |
Comput. Commun. | 2 |
| 2015 | Tiny habits in the giant enterprise: understanding the dynamics of a quantified workplaceabstractWe offer a reflection on the technology usage for workplace quantification through an in the wild study. Using a prototype Quantified Workplace system equipped with passive and participatory sensing modalities, we collected and visualized different workplace metrics (noise, color, air quality, self reported mood, and self reported activity) in two European offices of a research organization for a period of 4 months. Next we surveyed 70 employees to understand their engagement experience with the system. We then conducted semi-structured interviews with 20 employees in which they explained which workplace metrics are useful and why, how they engage with the system and what privacy concerns they have. Our findings suggest that sense of inclusion acts as the initial incentive for engagement which gradually translates into a habitual routine. We found that incorporation of an anonymous participatory sensing aspect into the system could lead to sustained user engagement. Compared to past studies we observed a shift in the privacy concerns, due to the trust and transparency of our prototype system. We conclude by providing a set of design principles for building future Quantified Workplace systems. Akhil Mathur, Marc Van den Broeck, Geert Vanderhulst, Afra J. Mashhadi, Fahim Kawsar |
UbiComp | 4 |
| 2015 | Detecting human encounters from WiFi radio signalsabstractWe present the design, implementation and evaluation of a novel human encounter detection framework for measuring and analysing human behaviour in social settings. We propose the use of WiFi probes, management frames of WiFi, that periodically radiate from mobile devices (as proxies for humans), and existing WiFi access points to automatically capture radio signals and detect human copresence. Based on the spatio-temporal properties of this copresence and their interplay we defined a model, borrowing theories from sociology, to detect human encounters -- short-lived, spontaneous human interactions. We evaluated our framework using controlled and in-the-wild experiments yielding a detection performance of 96% and 86% respectively. As such, our framework opens up interesting opportunities for designing proxemic and group applications, as well as conducting large-scale studies in the areas of computational social sciences. Geert Vanderhulst, Afra J. Mashhadi, Marzieh Dashti, Fahim Kawsar |
MUM | 2 |
| 2014 | Poverty on the cheap: estimating poverty maps using aggregated mobile communication networksabstractGovernments and other organisations often rely on data collected by household surveys and censuses to identify areas in most need of regeneration and development projects. However, due to the high cost associated with the data collection process, many developing countries conduct such surveys very infrequently and include only a rather small sample of the population, thus failing to accurately capture the current socio-economic status of the country's population. In this paper, we address this problem by means of a methodology that relies on an alternative source of data from which to derive up to date poverty indicators, at a very fine level of spatio-temporal granularity. Taking two developing countries as examples, we show how to analyse the aggregated call detail records of mobile phone subscribers and extract features that are strongly correlated with poverty indexes currently derived from census data. Chris Smith-Clarke, Afra J. Mashhadi, Licia Capra |
CHI | 2 |
| 2014 | Mind the map: the impact of culture and economic affluence on crowd-mapping behavioursabstractCrowd-mapping is a form of collaborative work that empowers citizens to collect and share geographic knowledge. OpenStreetMap (OSM) is a successful example of such paradigm, where the goal of building and maintaining an accurate global map of the changing world is being accomplished by means of local contributions made by over 1.2M citizens. While OSM has been subject to many country-specific studies, the relationship between national culture and economic affluence and users' participation has been so far unexplored. In this work, we systematically study the link between them: we characterise OSM users in terms of who they are, how they contribute, during what period of time, and across what geographic areas. We find strong correlations between these characteristics and national culture factors (e.g., power distance, individualism, pace of life, self expression), and well as Gross Domestic Product per capita. Based on these findings, we discuss design issues that developers of crowd-mapping services should consider to account for cross-cultural differences. Giovanni Quattrone, Afra J. Mashhadi, Licia Capra |
CSCW | 2 |
| 2014 | Modelling growth of urban crowd-sourced informationabstractUrban crowd-sourcing has become a popular paradigm to harvest spatial information about our evolving cities directly from citizens. OpenStreetMap is a successful example of such paradigm, with an accuracy of its user-generated content comparable to that of curated databases (e.g., Ordnance Survey). Coverage is however low and most importantly non-uniformly distributed across the city. Being able to model the spontaneous growth of digital information in these domains is required, so to be able to plan interventions aimed at gathering content about areas that would otherwise be neglected. Inspired by models of physical urban growth developed by urban planners, we build a model of digital growth of crowd-sourced spatial information that is both easy to interpret and dynamic, so to be able to determine what factors impact growth and how these change over time. We build and test the model against five years of OpenStreetMap data for the city of London, UK. We then run the model against two other cities, chosen for their different physical and digital growth's characteristics, so to stress-test the model. We conclude with a discussion of the implications of this work on both developers and users of urban crowd-sourcing applications. Giovanni Quattrone, Afra J. Mashhadi, Daniele Quercia, Chris Smith-Clarke, Licia Capra |
WSDM | 2 |
| 2013 | Putting ubiquitous crowd-sourcing into contextabstractUbiquitous crowd-sourcing has become a popular mechanism to harvest knowledge from the masses. OpenStreetMap (OSM) is a successful example of ubiquitous crowd-sourcing, where citizens volunteer geographic information in order to build and maintain an accurate map of the changing world. Research has shown that OSM information is accurate, by comparing it with centrally maintained spatial information such as Ordnance Survey. However, we find that coverage is low and non uniformly distributed, thus challenging the suitability of ubiquitous crowd-sourcing as a mechanism to map the whole world. In this paper, we investigate what contextual factors correlate with coverage of OSM information in urban settings. We find that, although there is a direct correlation between population density and information coverage, other socio-economic factors also play an important role. We discuss the implications of these findings with respect to the design of urban crowd-sourcing applications. Afra J. Mashhadi, Giovanni Quattrone, Licia Capra |
CSCW | 1 |
| 2013 | The Life of the Party: Impact of Social Mapping in OpenStreetMap
Desislava Hristova, Giovanni Quattrone, Afra J. Mashhadi, Licia Capra |
ICWSM | 3 |
| 2013 | Temporal analysis of activity patterns of editors in collaborative mapping project of OpenStreetMapabstractIn the recent years Wikis have become an attractive platform for social studies of the human behaviour. Containing millions records of edits across the globe, collaborative systems such as Wikipedia have allowed researchers to gain a better understanding of editors participation and their activity patterns. However, contributions made to Geo-wikis_wiki-based collaborative mapping projects_ differ from systems such as Wikipedia in a fundamental way due to spatial dimension of the content that limits the contributors to a set of those who posses local knowledge about a specific area and therefore cross-platform studies and comparisons are required to build a comprehensive image of online open collaboration phenomena. In this work, we study the temporal behavioural pattern of OpenStreetMap editors, a successful example of geo-wiki, for two European capital cities. We categorise different type of temporal patterns and report on the historical trend within a period of 7 years of the project age. We also draw a comparison with the previously observed editing activity patterns of Wikipedia. Taha Yasseri, Giovanni Quattrone, Afra J. Mashhadi |
OpenSym | 3 |
| 2012 | Fair content dissemination in participatory DTNs
Afra J. Mashhadi, Sonia Ben Mokhtar, Licia Capra |
Ad Hoc Networks | 1 |
| 2011 | Priority scheduling for participatory Delay Tolerant NetworksabstractDelay Tolerant Networking (DTN) protocols have been investigated as effective ways to distribute content in scenarios where the producers and consumers of content belong to the same geographical area. Common focus of all these approaches has been on maximising network delivery while treating messages as if they were all worth the same to recipients (i.e., messages are forwarded in a first-encountered/first-forwarded fashion). However, because of battery limitations on mobile devices, it is often the case that not all messages can be delivered. In this paper, we propose a new approach for priority-scheduling in participatory DTNs, whereby both message's value to the end-user as well as the likelihood of future encounters are combined to make the forwarding decision. Afra J. Mashhadi, Licia Capra |
WOWMOM | 1 |
| 2009 | Habit: Leveraging human mobility and social network for efficient content dissemination in Delay Tolerant NetworksabstractThis paper proposes Habit, an efficient multi-layered approach to content dissemination in Delay Tolerant Networks (DTN) that leverages information about nodes' colocation (physical layer) and their social network (application layer). More precisely, the regularity of users' colocation is learned based on historical colocation observations; also, the users' social network (or `network of interest') is dynamically propagated during periods of colocation; finally, these distinct pieces of information are locally combined and used to compute the paths that content should follow, in a way that maximises both precision (i.e., nodes receive only content they are interested in) and recall (i.e., all relevant content is received by interested nodes). Afra J. Mashhadi, Sonia Ben Mokhtar, Licia Capra |
WOWMOM | 1 |