VLDB 2026 Research / reviewers in the wild / expert
Sanja Scepanovic
dblp:151/9322
· DBLP profile ↗
18ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0002-1534-8128ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How Your Location Relates to Health: Variable Importance and Interpretable Machine Learning for Environmental and Sociodemographic DataabstractHealth outcomes depend on complex environmental and sociodemographic factors whose effects change over location and time. Only recently has fine-grained spatial and temporal data become available to study these effects, namely the MEDSAT dataset of English health, environmental, and sociodemographic information. Leveraging this new resource, we use a variety of variable importance techniques to robustly identify the most informative predictors across multiple health outcomes. We then develop an interpretable machine learning framework based on Generalized Additive Models (GAMs) and Multiscale Geographically Weighted Regression (MGWR) to analyze both local and global spatial dependencies of each variable on various health outcomes. Our findings identify NO2 as a global predictor for asthma, hypertension, and anxiety, alongside other outcome-specific predictors related to occupation, marriage, and vegetation. Regional analyses reveal local variations with air pollution and solar radiation, with notable shifts during COVID. This comprehensive approach provides actionable insights for addressing health disparities, and advocates for the integration of interpretable machine learning in public health. Ishaan Maitra, Raymond Lin, Jon Donnelly, Sanja Scepanovic, Cynthia Rudin |
AAAI | 5 |
| 2025 | RiskRAG: A Data-Driven Solution for Improved AI Model Risk ReportingabstractRisk reporting is essential for documenting AI models, yet only 14% of model cards mention risks, out of which 96% copying content from a small set of cards, leading to a lack of actionable insights. Existing proposals for improving model cards do not resolve these issues. To address this, we introduce RiskRAG, a Retrieval Augmented Generation based risk reporting solution guided by five design requirements we identified from literature, and co-design with 16 developers: identifying diverse model-specific risks, clearly presenting and prioritizing them, contextualizing for real-world uses, and offering actionable mitigation strategies. Drawing from 450K model cards and 600 real-world incidents, RiskRAG pre-populates contextualized risk reports. A preliminary study with 50 developers showed that they preferred RiskRAG over standard model cards, as it better met all the design requirements. A final study with 38 developers, 40 designers, and 37 media professionals showed that RiskRAG improved their way of selecting the AI model for a specific application, encouraging a more careful and deliberative decision-making. The RiskRAG project page is accessible at: https://social-dynamics.net/ai-risks/card. Pooja S. B. Rao, Sanja Scepanovic, Ke Zhou 0003, Edyta Paulina Bogucka, Daniele Quercia |
CHI | 2 |
| 2025 | Impact Assessment Card: Communicating Risks and Benefits of AI UsesabstractCommunicating the risks and benefits of AI is important for regulation and public understanding. Yet current methods such as technical reports often exclude people without technical expertise. Drawing on HCI research, we developed an Impact Assessment Card to present this information more clearly. We held three focus groups with a total of 12 participants who helped identify design requirements and create early versions of the card. We then tested a refined version in an online study with 235 participants, including AI developers, compliance experts, and members of the public selected to reflect the U.S. population by age, sex, and race. Participants used either the card or a full impact assessment report to write an email supporting or opposing a proposed AI system. The card led to faster task completion and higher-quality emails across all groups. We discuss how design choices can improve accessibility and support AI governance. Examples of cards are available at: https://social-dynamics.net/ai-risks/impact-card/ Edyta Paulina Bogucka, Marios Constantinides, Sanja Scepanovic, Daniele Quercia |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | Co-designing an AI Impact Assessment Report Template with AI Practitioners and AI Compliance ExpertsabstractIn the evolving landscape of AI regulation, it is crucial for companies to conduct impact assessments and document their compliance through comprehensive reports. However, current reports lack grounding in regulations and often focus on specific aspects like privacy in relation to AI systems, without addressing the real-world uses of these systems. Moreover, there is no systematic effort to design and evaluate these reports with both AI practitioners and AI compliance experts. To address this gap, we conducted an iterative co-design process with 14 AI practitioners and 6 AI compliance experts and proposed a template for impact assessment reports grounded in the EU AI Act, NIST's AI Risk Management Framework, and ISO 42001 AI Management System. We evaluated the template by producing an impact assessment report for an AI-based meeting companion at a major tech company. A user study with 8 AI practitioners from the same company and 5 AI compliance experts from industry and academia revealed that our template effectively provides necessary information for impact assessments and documents the broad impacts of AI systems. Participants envisioned using the template not only at the pre-deployment stage for compliance but also as a tool to guide the design stage of AI uses. Edyta Paulina Bogucka, Marios Constantinides, Sanja Scepanovic, Daniele Quercia |
AIES (1) | 3 |
| 2024 | ExploreGen: Large Language Models for Envisioning the Uses and Risks of AI TechnologiesabstractResponsible AI design is increasingly seen as an imperative by both AI developers and AI compliance experts. One of the key tasks is envisioning AI technology uses and risks. Recent studies on the model and data cards reveal that AI practitioners struggle with this task due to its inherently challenging nature. Here, we demonstrate that leveraging a Large Language Model (LLM) can support AI practitioners in this task by enabling reflexivity, brainstorming, and deliberation, especially in the early design stages of the AI development process. We developed an LLM framework, ExploreGen, which generates realistic and varied uses of AI technology, including those overlooked by research, and classifies their risk level based on the EU AI Act regulation. We evaluated our framework using the case of Facial Recognition and Analysis technology in nine user studies with 25 AI practitioners. Our findings show that ExploreGen is helpful to both developers and compliance experts. They rated the uses as realistic and their risk classification as accurate (94.5%). Moreover, while unfamiliar with many of the uses, they rated them as having high adoption potential and transformational impact. Viviane Herdel, Sanja Scepanovic, Edyta Paulina Bogucka, Daniele Quercia |
AIES (1) | 2 |
| 2024 | Characterizing Fake News Targeting CorporationsabstractMisinformation proliferates in the online sphere, with evident impacts on the political and social realms, influencing democratic discourse and posing risks to public health and safety. The corporate world is also a prime target for fake news dissemination. While recent studies have attempted to characterize corporate misinformation and its effects on companies, their findings often suffer from limitations due to qualitative or narrative approaches and a narrow focus on specific industries. To address this gap, we conducted an analysis utilizing social media quantitative methods and crowd-sourcing studies to investigate corporate misinformation across a diverse array of industries within the S&P 500 companies. Our study reveals that corporate misinformation encompasses topics such as products, politics, and societal issues. We discovered companies affected by fake news also get reputable news coverage but less social media attention, leading to heightened negativity in social media comments, diminished stock growth, and increased stress mentions among employee reviews. Additionally, we observe that a company is not targeted by fake news all the time, but there are particular times when a critical mass of fake news emerges. These findings hold significant implications for regulators, business leaders, and investors, emphasizing the necessity to vigilantly monitor the escalating phenomenon of corporate misinformation. Ke Zhou 0003, Sanja Scepanovic, Daniele Quercia |
ICWSM | 2 |
| 2024 | Exploratory Analysis of Recommending Urban Parks for Health-Promoting ActivitiesabstractParks are essential spaces for promoting urban health, and recommender systems could assist individuals in discovering parks for leisure and health-promoting activities. This is particularly important in large cities like London, which has over 1,500 named parks, making it challenging to understand what each park offers. Due to the lack of datasets and the diverse health-promoting activities parks can support (e.g., physical, social, nature-appreciation), it is unclear which recommendation algorithms are best suited for this task. To explore the dynamics of recommending parks for specific activities, we created two datasets: one from a survey of over 250 London residents, and another by inferring visits from over 1 million geotagged Flickr images taken in London parks. Analyzing the geographic patterns of these visits revealed that recommending nearby parks is ineffective, suggesting that this recommendation task is distinct from Point of Interest recommendation. We then tested various recommendation models, identifying a significant popularity bias in the results. Additionally, we found that personalized models have advantages in recommending parks beyond the most popular ones. The data and findings from this study provide a foundation for future research on park recommendations. Linus W. Dietz, Sanja Scepanovic, Ke Zhou 0003, Daniele Quercia |
RecSys | 2 |
| 2024 | Good Intentions, Risky Inventions: A Method for Assessing the Risks and Benefits of AI in Mobile and Wearable UsesabstractIntegrating Artificial Intelligence (AI) into mobile and wearables offers numerous benefits at individual, societal, and environmental levels. Yet, it also spotlights concerns over emerging risks. Traditional assessments of risks and benefits have been sporadic, and often require costly expert analysis. We developed a semi-automatic method that leverages Large Language Models (LLMs) to identify AI uses in mobile and wearables, classify their risks based on the EU AI Act, and determine their benefits that align with globally recognized long-term sustainable development goals; a manual validation of our method by two experts in mobile and wearable technologies, a legal and compliance expert, and a cohort of nine individuals with legal backgrounds who were recruited from Prolific, confirmed its accuracy to be over 85%. We uncovered that specific applications of mobile computing hold significant potential in improving well-being, safety, and social equality. However, these promising uses are linked to risks involving sensitive data, vulnerable groups, and automated decision-making. To avoid rejecting these risky yet impactful mobile and wearable uses, we propose a risk assessment checklist for the Mobile HCI community. Marios Constantinides, Edyta Paulina Bogucka, Sanja Scepanovic, Daniele Quercia |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2023 | Heart Rate Extraction from Abdominal Audio SignalsabstractAbdominal sounds (ABS) have been traditionally used for assessing gastrointestinal (GI) disorders. However, the assessment requires a trained medical professional to perform multiple abdominal auscultation sessions, which is resource-intense and may fail to provide an accurate picture of patients’ continuous GI wellbeing. This has generated a technological interest in developing wearables for continuous capture of ABS, which enables a fuller picture of patient’s GI status to be obtained at reduced cost. This paper seeks to evaluate the feasibility of extracting heart rate (HR) from such ABS monitoring devices. The collection of HR directly from these devices would enable gathering vital signs alongside GI data without the need for additional wearable devices, providing further cost benefits and improving general usability. We utilised a dataset containing 104 hours of ABS audio, collected from the abdomen using an e-stethoscope, and electrocardiogram as ground truth. Our evaluation shows for the first time that we can successfully extract HR from audio collected from a wearable on the abdomen. As heart sounds collected from the abdomen suffer from significant noise from GI and respiratory tracts, we leverage wavelet denoising for improved heart beat detection. The mean absolute error of the algorithm for average HR is 3.4BPM with mean directional error of -1.2BPM over the whole dataset. A comparison to photoplethysmography-based wearable HR sensors shows that our approach exhibits comparable accuracy to consumer wrist-worn wearables for average and instantaneous heart rate. Jake Stuchbury-Wass, Erika Bondareva, Kayla-Jade Butkow, Sanja Scepanovic, Zoran Radivojevic, Cecilia Mascolo |
ICASSP | 4 |
| 2023 | How Circadian Rhythms Extracted from Social Media Relate to Physical Activity and SleepabstractCircadian rhythm has been linked to both physical and mental health at an individual level in prior research. Such a link at population level has been long hypothesized but has never been tested, largely because of lack of data. To partly fix this literature gap, we need: a dataset on population-level circadian rhythms, a dataset on population-level health conditions, and strong associations between these two partly independent sets. Recent work has shown that affect on social media data relates to population-level circadian rhythms. Building upon that work, we extracted five circadian rhythm metrics from 6M Reddit posts across 18 major cities (for which the number of residents is highly correlated with the number of users), and paired them with three ground-truth health metrics (daily number of steps, sleep quantity, and sleep quality) extracted from 233K wearable users in these cities. We found that rhythms of online activity approximated sleeping patterns rather than, what the literature previously hypothesized, alertness levels. Despite that, we found that these rhythms, when computed in two specific times of the day (i.e., late at night and early morning), were still predictive of the three ground-truth health metrics: in general, healthier cities had morning spikes on social media, night dips, and expressions of positive affect. These results suggest that circadian rhythms on social media, if taken at two specific times of the day and operationalized with literature-driven metrics, can approximate the temporal evolution of people's shared underlying biological rhythm as it relates to physical activity (R2=0.492), sleep quantity (R2=0.765), and sleep quality (R2=0.624). Ke Zhou 0003, Marios Constantinides, Daniele Quercia, Sanja Scepanovic |
ICWSM | 4 |
| 2023 | Responsible AI for Earth Observation: Attitides Among ExpertsabstractAs AI permeates industries and reaches the general public, the significance of responsible AI (RAI) principles becomes increasingly vital. This study offers valuable insights from 27 Earth Observation (EO) experts in 11 countries, unveiling diverse attitudes towards RAI principles and nuanced perspectives across genders and age groups. It highlights the variation in definitions and interpretations of core principles such as fairness, reliability, privacy, transparency, accountability, and sustainability. Moreover, the study identifies RAI concerns specific to the EO domain, emphasizing the need to integrate domain knowledge and effectively communicate issues of inequality and failure use cases. These findings contribute to a preliminary understanding of RAI attitudes in the AI for EO context. Future research should involve a larger pool of experts and investigate the attitudes of EO users and the general public to complement these initial findings. Sanja Scepanovic, Edyta Paulina Bogucka, Daniele Quercia, Cristiano Nattero |
IGARSS | 1 |
| 2023 | MedSat: A Public Health Dataset for England Featuring Medical Prescriptions and Satellite ImageryabstractAs extreme weather events become more frequent, understanding their impact on human health becomes increasingly crucial. However, the utilization of Earth Observation to effectively analyze the environmental context in relation to health remains limited. This limitation is primarily due to the lack of fine-grained spatial and temporal data in public and population health studies, hindering a comprehensive understanding of health outcomes. Additionally, obtaining appropriate environmental indices across different geographical levels and timeframes poses a challenge. For the years 2019 (pre-COVID) and 2020 (COVID), we collected spatio-temporal indicators for all Lower Layer Super Output Areas in England. These indicators included: i) 111 sociodemographic features linked to health in existing literature, ii) 43 environmental point features (e.g., greenery and air pollution levels), iii) 4 seasonal composite satellite images each with 11 bands, and iv) prescription prevalence associated with five medical conditions (depression, anxiety, diabetes, hypertension, and asthma), opioids and total prescriptions. We combined these indicators into a single MedSat dataset, the availability of which presents an opportunity for the machine learning community to develop new techniques specific to public health. These techniques would address challenges such as handling large and complex data volumes, performing effective feature engineering on environmental and sociodemographic factors, capturing spatial and temporal dependencies in the models, addressing imbalanced data distributions, developing novel computer vision methods for health modeling based on satellite imagery, ensuring model explainability, and achieving generalization beyond the specific geographical region. Sanja Scepanovic, Ivica Obadic, Sagar Joglekar 0001, Laura Giustarini, Cristiano Nattero, Daniele Quercia, Xiao Xiang Zhu 0001 |
NeurIPS | 1 |
| 2022 | Depression at Work: Exploring Depression in Major US Companies from Online ReviewsabstractStudies on depression in the workplace have mostly investigated its impact on individual employees. Little is known about its association with the company as a whole, or the state where the company is based. This is due to the lack of scalable methodologies operationalizing depression in the specific context of the workplace, and of data documenting potential distress. In this work, we adapted a work-related depression scale called Occupational Depression Inventory (ODI), gathered more than 350K employee reviews of 104 major companies across the whole US for the (2008-2020) years, and developed a deep-learning framework (called AutoODI) scoring these reviews on a composite ODI score. Presence of ODI mentions manifested itself not only at micro-level (companies scoring high in ODI suffered from low stock growth) but also at macro-level (states hosting these companies were associated with high depression rates, talent shortage, and economic deprivation). This new way of applying AutoODI onto company reviews offers both theoretical implications for the literature in computational social science, occupational health and economic geography, and practical implications for companies and policy makers. Indira Sen, Daniele Quercia, Marios Constantinides, Matteo Montecchi, Licia Capra, Sanja Scepanovic, Renzo Bianchi |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2021 | The Healthy States of America: Creating a Health Taxonomy with Social Media
Sanja Scepanovic, Luca Maria Aiello, Ke Zhou 0003, Sagar Joglekar 0001, Daniele Quercia |
ICWSM | 1 |
| 2021 | Jane Jacobs in the Sky: Predicting Urban Vitality with Open Satellite DataabstractThe presence of people in an urban area throughout the day -- often called 'urban vitality' -- is one of the qualities world-class cities aspire to the most, yet it is one of the hardest to achieve. Back in the 1970s, Jane Jacobs theorized urban vitality and found that there are four conditions required for the promotion of life in cities: diversity of land use, small block sizes, the mix of economic activities, and concentration of people. To build proxies for those four conditions and ultimately test Jane Jacobs's theory at scale, researchers have had to collect both private and public data from a variety of sources, and that took decades. Here we propose the use of one single source of data, which happens to be publicly available: Sentinel-2 satellite imagery. In particular, since the first two conditions (diversity of land use and small block sizes) are visible to the naked eye from satellite imagery, we tested whether we could automatically extract them with a state-of-the-art deep-learning framework and whether, in the end, the extracted features could predict vitality. In six Italian cities for which we had call data records, we found that our framework is able to explain on average 55% of the variance in urban vitality extracted from those records. Sanja Scepanovic, Sagar Joglekar 0001, Stephen Law, Daniele Quercia |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2021 | The Language of Situational EmpathyabstractEmpathy is the tendency to understand and share others' thoughts and feelings. Literature in psychology has shown through surveys potential beneficial implications of empathy. Prior psychology literature showed that a particular type of empathy called "situational empathy" --- an immediate empathic response to a triggering situation (e.g., a distressing situation) --- is reflected in the language people use in response to the situation. However, this has not so far been properly measured at scale. In this work, we collected 4k textual reactions (and corresponding situational empathy labels) to different stories. Driven by theoretical concepts, we developed computational models to predict situational empathy from text and, in so doing, we built and made available a list of empathy-related words. When applied to Reddit posts and movie transcripts, our models produced results that matched prior theoretical findings, offering evidence of external validity and suggesting its applicability to unstructured data. The capability of measuring proxies for empathy at scale might benefit a variety of areas such as social media, digital healthcare, and workplace well-being. Ke Zhou 0003, Luca Maria Aiello, Sanja Scepanovic, Daniele Quercia, Sara Konrath |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2021 | Humane Visual AI: Telling the Stories Behind a Medical ConditionabstractA biological understanding is key for managing medical conditions, yet psychological and social aspects matter too. The main problem is that these two aspects are hard to quantify and inherently difficult to communicate. To quantify psychological aspects, this work mined around half a million Reddit posts in the sub-communities specialised in 14 medical conditions, and it did so with a new deep-learning framework. In so doing, it was able to associate mentions of medical conditions with those of emotions. To then quantify social aspects, this work designed a probabilistic approach that mines open prescription data from the National Health Service in England to compute the prevalence of drug prescriptions, and to relate such a prevalence to census data. To finally visually communicate each medical condition's biological, psychological, and social aspects through storytelling, we designed a narrative-style layered Martini Glass visualization. In a user study involving 52 participants, after interacting with our visualization, a considerable number of them changed their mind on previously held opinions: 10% gave more importance to the psychological aspects of medical conditions, and 27% were more favourable to the use of social media data in healthcare, suggesting the importance of persuasive elements in interactive visualizations. Wonyoung So, Edyta Paulina Bogucka, Sanja Scepanovic, Sagar Joglekar 0001, Ke Zhou 0003, Daniele Quercia |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Classification of Wide-Area SAR Mosaics: Deep Learning Approach for Corine Based Mapping of Finland Using Multitemporal Sentinel-1 DataabstractHere, we examine a deep learning approach to perform land cover classification using country-wide SAR mosaics compiled using multitemporal Sentinel-1 imagery. We capitalize on our earlier study [1], demonstrating the suitability of deep learning models for land cover mapping using satellite C-band SAR images. A set of SAR mosaics compiled from consecutive Sentinel-1 IW mode acquisitions covering the whole territory of Finland was used in production of the whole-country land cover map. The imagery were used as an input to the state-of-the-art deep-learning model for semantic segmentation called FC-DenseNet. This model was pre-trained on the ImageNet dataset and further fine-tuned in this study. CORINE land cover map was used as a reference, and the model was trained to distinguish between 5 Level-1 CORINE classes. Upon the evaluation and benchmarking, we found that the FC-DenseNet model is able to achieve nearly 90% overall classification accuracy. These results indicate the suitability of deep learning approaches to support efficient operational wide-area mapping using satellite SAR imagery. Oleg Antropov, Yrjö Rauste, Sanja Scepanovic, Vladimir Ignatenko, Anne Lönnqvist, Jaan Praks |
IGARSS | 3 |