VLDB 2026 Research / reviewers in the wild / expert
Rumi Chunara
dblp:130/0470
· DBLP profile ↗
20ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-5346-7259ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 10 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identity-Robust Language Model Generation via Content Integrity PreservationabstractLarge Language Model (LLM) outputs often vary across user sociodemographic attributes, leading to disparities in factual accuracy, utility, and safety, even for objective questions where demographic information is irrelevant.Unlike prior work on stereotypical or representational bias, this paper studies identitydependent degradation of core response quality.We show empirically that such degradation arises from biased generation behavior, despite factual knowledge being robustly encoded across identities.Motivated by this mismatch, we propose a lightweight, training-free framework for identity-robust generation that selectively neutralizes non-critical identity information while preserving semantically essential attributes, thus maintaining output content integrity.Experiments across four benchmarks and 18 sociodemographic identities demonstrate an average 66.3% reduction in identitydependent bias compared to vanilla prompting and outperforms existing prompt-based defenses.Our work addresses a critical gap in mitigating the impact of user identity cues in prompts on core generation quality. Miao Zhang 0030, Kelly Chen, Md Mehrab Tanjim, Rumi Chunara |
ACL (1) | 4 |
| 2025 | Machine learning based prediction of medication adherence in heart failure using large electronic health record cohort with linkages to pharmacy-fill and neighborhood-level dataabstractOBJECTIVE: While timely interventions can improve medication adherence, it is challenging to identify which patients are at risk of nonadherence at point-of-care. We aim to develop and validate flexible machine learning (ML) models to predict a continuous measure of adherence to guideline-directed medication therapies (GDMTs) for heart failure (HF). MATERIALS AND METHODS: We utilized a large electronic health record (EHR) cohort of 34,697 HF patients seen at NYU Langone Health with an active prescription for ≥1 GDMT between April 01, 2021 and October 31, 2022. The outcome was adherence to GDMT measured as proportion of days covered (PDC) at 6 months following a clinical encounter. Over 120 predictors included patient-, therapy-, healthcare-, and neighborhood-level factors guided by the World Health Organization's model of barriers to adherence. We compared performance of several ML models and their ensemble (superlearner) for predicting PDC with traditional regression model (OLS) using mean absolute error (MAE) averaged across 10-fold cross-validation, % increase in MAE relative to superlearner, and predictive-difference across deciles of predicted PDC. RESULTS: Superlearner, a flexible nonparametric prediction approach, demonstrated superior prediction performance. Superlearner and quantile random forest had the lowest MAE (mean [95% CI] = 18.9% [18.7%-19.1%] for both), followed by MAEs for quantile neural network (19.5% [19.3%-19.7%]) and kernel support vector regression (19.8% [19.6%-20.0%]). Gradient boosted trees and OLS were the 2 worst performing models with 17% and 14% higher MAEs, respectively, relative to superlearner. Superlearner demonstrated improved predictive difference. CONCLUSION: This development phase study suggests potential of linked EHR-pharmacy data and ML to identify HF patients who will benefit from medication adherence interventions. DISCUSSION: Fairness evaluation and external validation are needed prior to clinical integration. Samrachana Adhikari, Tyrel Stokes, Xiyue Li, Yunan Zhao, Cassidy Fitchett, Nathalia Ladino, Steven Lawrence, Young S. Cho, Carine Hamo, John A Dodson, Rumi Chunara, Ian M. Kronish, Amrita Mukhopadhyay, Saul Blecker |
J. Am. Medical Informatics Assoc. | 12 |
| 2024 | Mitigating Urban-Rural Disparities in Contrastive Representation Learning with Satellite ImageryabstractSatellite imagery is being leveraged for many societally critical tasks across climate, economics, and public health. Yet, because of heterogeneity in landscapes (e.g. how a road looks in different places), models can show disparate performance across geographic areas. Given the important potential of disparities in algorithmic systems used in societal contexts, here we consider the risk of urban-rural disparities in identification of land-cover features. This is via semantic segmentation (a common computer vision task in which image regions are labelled according to what is being shown) which uses pre-trained image representations generated via contrastive self-supervised learning. We propose fair dense representation with contrastive learning (FairDCL) as a method for de-biasing the multi-level latent space of a convolution neural network. The method improves feature identification by removing spurious latent representations which are disparately distributed across urban and rural areas, and is achieved in an unsupervised way by contrastive pre-training. The pre-trained image representation mitigates downstream urban-rural prediction disparities and outperforms state-of-the-art baselines on real-world satellite images. Embedding space evaluation and ablation studies further demonstrate FairDCL’s robustness. As generalizability and robustness in geographic imagery is a nascent topic, our work motivates researchers to consider metrics beyond average accuracy in such applications. Miao Zhang 0030, Rumi Chunara |
AIES (1) | 2 |
| 2023 | Measures of Disparity and their Efficient EstimationabstractQuantifying disparities, that is differences in outcomes among population groups, is an important task in public health, economics, and increasingly in machine learning. In this work, we study the question of how to collect data to measure disparities. The field of survey statistics provides extensive guidance on sample sizes necessary to accurately estimate quantities such as averages. However, there is limited guidance for estimating disparities. We consider a broad class of disparity metrics including those used in machine learning for measuring fairness of model outputs. For each metric, we derive the number of samples to be collected per group that increases the precision of disparity estimates given a fixed data collection budget. We also provide sample size calculations for hypothesis tests that check for significant disparities. Our methods can be used to determine sample sizes for fairness evaluations. We validate the methods on two nationwide surveys, used for understanding population-level attributes like employment and health, and a prediction model. Absent a priori information on the groups, we find that equally sampling the groups typically performs well. Harvineet Singh, Rumi Chunara |
AIES | 2 |
| 2023 | When do Minimax-fair Learning and Empirical Risk Minimization Coincide?abstractMinimax-fair machine learning minimizes the error for the worst-off group. However, empirical evidence suggests that when sophisticated models are trained with standard empirical risk minimization (ERM), they often have the same performance on the worst-off group as a minimax-trained model. Our work makes this counter-intuitive observation concrete. We prove that if the hypothesis class is sufficiently expressive and the group information is recoverable from the features, ERM and minimax-fairness learning formulations indeed have the same performance on the worst-off group. We provide additional empirical evidence of how this observation holds on a wide range of datasets and hypothesis classes. Since ERM is fundamentally easier than minimax optimization, our findings have implications on the practice of fair machine learning. Harvineet Singh, Matthäus Kleindessner, Volkan Cevher, Rumi Chunara, Chris Russell 0001 |
ICML | 4 |
| 2021 | Uncertainty as a Form of Transparency: Measuring, Communicating, and Using UncertaintyabstractAlgorithmic transparency entails exposing system properties to various stakeholders for purposes that include understanding, improving, and contesting predictions. Until now, most research into algorithmic transparency has predominantly focused on explainability. Explainability attempts to provide reasons for a machine learning model's behavior to stakeholders. However, understanding a model's specific behavior alone might not be enough for stakeholders to gauge whether the model is wrong or lacks sufficient knowledge to solve the task at hand. In this paper, we argue for considering a complementary form of transparency by estimating and communicating the uncertainty associated with model predictions. First, we discuss methods for assessing uncertainty. Then, we characterize how uncertainty can be used to mitigate model unfairness, augment decision-making, and build trustworthy systems. Finally, we outline methods for displaying uncertainty to stakeholders and recommend how to collect information required for incorporating uncertainty into existing ML pipelines. This work constitutes an interdisciplinary review drawn from literature spanning machine learning, visualization/HCI, design, decision-making, and fairness. We aim to encourage researchers and practitioners to measure, communicate, and use uncertainty as a form of transparency. Umang Bhatt, Javier Antorán, Qingzi Vera Liao, Prasanna Sattigeri, Riccardo Fogliato, Gabrielle Gauthier Melançon, Ranganath Krishnan, Jason Stanley, Omesh Tickoo, Lama Nachman, Rumi Chunara, Madhulika Srikumar, Adrian Weller, Alice Xiang |
AIES | 12 |
| 2021 | Causal Multi-level FairnessabstractAlgorithmic systems are known to impact marginalized groups severely, and more so, if all sources of bias are not considered. While work in algorithmic fairness to-date has primarily focused on addressing discrimination due to individually linked attributes, social science research elucidates how some properties we link to individuals can be conceptualized as having causes at macro (e.g. structural) levels, and it may be important to be fair to attributes at multiple levels. For example, instead of simply considering race as a causal, protected attribute of an individual, the cause may be distilled as perceived racial discrimination an individual experiences, which in turn can be affected by neighborhood-level factors. This multi-level conceptualization is relevant to questions of fairness, as it may not only be important to take into account if the individual belonged to another demographic group, but also if the individual received advantaged treatment at the macro-level. In this paper, we formalize the problem of multi-level fairness using tools from causal inference in a manner that allows one to assess and account for effects of sensitive attributes at multiple levels. We show importance of the problem by illustrating residual unfairness if macro-level sensitive attributes are not accounted for, or included without accounting for their multi-level nature. Further, in the context of a real-world task of predicting income based on macro and individual-level attributes, we demonstrate an approach for mitigating unfairness, a result of multi-level sensitive attributes. Vishwali Mhasawade, Rumi Chunara |
AIES | 2 |
| 2021 | Telemedicine and healthcare disparities: a cohort study in a large healthcare system in New York City during COVID-19abstractOBJECTIVE: Through the coronavirus disease 2019 (COVID-19) pandemic, telemedicine became a necessary entry point into the process of diagnosis, triage, and treatment. Racial and ethnic disparities in healthcare have been well documented in COVID-19 with respect to risk of infection and in-hospital outcomes once admitted, and here we assess disparities in those who access healthcare via telemedicine for COVID-19. MATERIALS AND METHODS: Electronic health record data of patients at New York University Langone Health between March 19th and April 30, 2020 were used to conduct descriptive and multilevel regression analyses with respect to visit type (telemedicine or in-person), suspected COVID diagnosis, and COVID test results. RESULTS: Controlling for individual and community-level attributes, Black patients had 0.6 times the adjusted odds (95% CI: 0.58-0.63) of accessing care through telemedicine compared to white patients, though they are increasingly accessing telemedicine for urgent care, driven by a younger and female population. COVID diagnoses were significantly more likely for Black versus white telemedicine patients. DISCUSSION: There are disparities for Black patients accessing telemedicine, however increased uptake by young, female Black patients. Mean income and decreased mean household size of a zip code were also significantly related to telemedicine use. CONCLUSION: Telemedicine access disparities reflect those in in-person healthcare access. Roots of disparate use are complex and reflect individual, community, and structural factors, including their intersection-many of which are due to systemic racism. Evidence regarding disparities that manifest through telemedicine can be used to inform tool design and systemic efforts to promote digital health equity. Rumi Chunara, Katharine Lawrence, Paul A. Testa, Oded Nov, Devin M. Mann |
J. Am. Medical Informatics Assoc. | 1 |
| 2020 | Quasi-Experimental Designs for Assessing Response on Social Media to Policy Changes
Yijun Tian 0001, Rumi Chunara |
ICWSM | 2 |
| 2020 | No Computation without Representation: Avoiding Data and Algorithm Biases through DiversityabstractThe emergence and growth of research on issues of ethics in Artificial Intelligence, and in particular algorithmic fairness, has roots in an essential observation that structural inequalities in our society are reflected in the data used to train predictive models and in the design of objective functions. While research aiming to mitigate these issues is inherently interdisciplinary, the design of unbiased algorithms and fair socio-technical systems are key desired outcomes which depend on practitioners from the fields of data science and computing. However, these computing fields broadly also suffer from the same under-representation issues that are found in the datasets we analyze. This disconnect affects the design of both the desired outcomes and metrics by which we measure success. If the ethical AI research community accepts this, we tacitly endorse the status quo and contradict the goals of non-discrimination and equity which work on algorithmic fairness, accountability, and transparency seeks to address. Caitlin Kuhlman, Latifa Jackson, Rumi Chunara |
KDD | 3 |
| 2020 | COVID-19 transforms health care through telemedicine: Evidence from the fieldabstractThis study provides data on the feasibility and impact of video-enabled telemedicine use among patients and providers and its impact on urgent and nonurgent healthcare delivery from one large health system (NYU Langone Health) at the epicenter of the coronavirus disease 2019 (COVID-19) outbreak in the United States. Between March 2nd and April 14th 2020, telemedicine visits increased from 102.4 daily to 801.6 daily. (683% increase) in urgent care after the system-wide expansion of virtual urgent care staff in response to COVID-19. Of all virtual visits post expansion, 56.2% and 17.6% urgent and nonurgent visits, respectively, were COVID-19-related. Telemedicine usage was highest by patients 20 to 44 years of age, particularly for urgent care. The COVID-19 pandemic has driven rapid expansion of telemedicine use for urgent care and nonurgent care visits beyond baseline periods. This reflects an important change in telemedicine that other institutions facing the COVID-19 pandemic should anticipate. Devin M. Mann, Rumi Chunara, Paul A. Testa, Oded Nov |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Race, Ethnicity and National Origin-Based Discrimination in Social Media and Hate Crimes across 100 U.S. Cities
Kunal Relia, Stephanie H. Cook, Rumi Chunara |
ICWSM | 4 |
| 2018 | Creating full individual-level location timelines from sparse social media dataabstractIn many domain applications, a continuous timeline of human locations is critical; for example for understanding possible locations wherea disease may spread, or the flow of traffic. While data sources such as GPS trackers or Call Data Records are temporally-rich, they are expensive, often not publicly available or garnered only in select locations, restricting their wide use. Conversely, geo-located social media data are publicly and freely available, but present challenges especially for full timeline inference due to their sparse nature. We propose a stochastic framework, Intermediate Location Computing (ILC) which uses prior knowledge about human mobility patterns to predict every missing location from an individual's social media timeline. We compare ILC with a state-of-the-art RNN baseline as well as methods that are optimized for next-location prediction only. For three major cities, ILC predicts the top 1 location for all missing locations in a timeline, at 1 and 2-hour resolution, with up to 77.2% accuracy (up to 6% better accuracy than all compared methods). Specifically, ILC also outperforms the RNN in settings of low data; both cases of very small number of users (under 50), as well as settings with more users, but with sparser timelines. In general, the RNN model needs a higher number of users to achieve the same performance as ILC. Overall, this work illustrates the tradeoff between prior knowledge of heuristics and more data, for an important societal problem of filling in entire timelines using freely available, but sparse social media data. Nabeel Abdur Rehman, Kunal Relia, Rumi Chunara |
SIGSPATIAL/GIS | 3 |
| 2018 | From the User to the Medium: Neural Profiling Across Web Communities
Mohammad Akbari 0001, Kunal Relia, Anas Elghafari, Rumi Chunara |
ICWSM | 4 |
| 2018 | Socio-spatial Self-organizing Maps: Using Social Media to Assess Relevant Geographies for Exposure to Social ProcessesabstractSocial media offers a unique window into attitudes like racism and homophobia, exposure to which are important, hard to measure and understudied social determinants of health. However, individual geo-located observations from social media are noisy and geographically inconsistent. Existing areas by which exposures are measured, like Zip codes, average over irrelevant administratively-defined boundaries. Hence, in order to enable studies of online social environmental measures like attitudes on social media and their possible relationship to health outcomes, first there is a need for a method to define the collective, underlying degree of social media attitudes by region. To address this, we create the Socio-spatial-Self organizing map, "SS-SOM" pipeline to best identify regions by their latent social attitude from Twitter posts. SS-SOMs use neural embedding for text-classification, and augment traditional SOMs to generate a controlled number of non-overlapping, topologically-constrained and topically-similar clusters. We find that not only are SS-SOMs robust to missing data, the exposure of a cohort of men who are susceptible to multiple racism and homophobia-linked health outcomes, changes by up to 42% using SS-SOM measures as compared to using Zip code-based measures. Kunal Relia, Mohammad Akbari 0001, Dustin Duncan, Rumi Chunara |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2017 | New data paradigms: From the crowd and backabstractKnowledge generation from citizens is becoming both more feasible as well as important. Data directly from individuals can be critical as it can add information beyond what is available otherwise. Crowdsourced data also is very amenable in open data efforts given the nature of its generation. In this talk I will describe several efforts in which we are generating crowdsourced knowledge from open data and using it to more readily improve knowledge in public health. Rumi Chunara |
IEEE BigData | 1 |
| 2017 | Assessing Behavior Stage Progression From Social Media DataabstractImportant work rooted in psychological theory posits that health behavior change occurs through a series of discrete stages. Our work builds on the field of social computing by identifying how social media data can be used to resolve behavior stages at high resolution (e.g. hourly/daily) for key population subgroups and times. In essence this approach opens new opportunities to advance psychological theories and better understand how our health is shaped based on the real, dynamic, and rapid actions we make every day. To do so, we bring together domain knowledge and machine learning methods to form a hierarchical classification of Twitter data that resolves different stages of behavior. We identify and examine temporal patterns of the identified stages, with alcohol as a use case (planning or looking to drink, currently drinking, and reflecting on drinking). Known seasonal trends are compared with findings from our methods. We discuss the potential health policy implications of detecting high frequency behavior stages. Elissa R. Weitzman, Rumi Chunara |
CSCW | 3 |
| 2017 | High-resolution Temporal Representations of Alcohol and Tobacco Behaviors from Social Media DataabstractUnderstanding tobacco- and alcohol-related behavioral patterns is critical for uncovering risk factors and potentially designing targeted social computing intervention systems. Given that we make choices multiple times per day, hourly and daily patterns are critical for better understanding behaviors. Here, we combine natural language processing, machine learning and time series analyses to assess Twitter activity specifically related to alcohol and tobacco consumption and their sub-daily, daily and weekly cycles. Twitter self-reports of alcohol and tobacco use are compared to other data streams available at similar temporal resolution. We assess if discussion of drinking by inferred underage versus legal age people or discussion of use of different types of tobacco products can be differentiated using these temporal patterns. We find that time and frequency domain representations of behaviors on social media can provide meaningful and unique insights, and we discuss the types of behaviors for which the approach may be most useful. Tom Huang, Anas Elghafari, Kunal Relia, Rumi Chunara |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2016 | Network inference from multimodal data: A review of approaches from infectious disease transmission
Bisakha Ray, Elodie Ghedin, Rumi Chunara |
J. Biomed. Informatics | 3 |
| 2014 | First Feasibility of a Surveillance Platform Combining Community-Submitted Symptoms and Specimens for Molecular Diagnostic Testing
Jennifer Goff, Aaron A. Rowe, Rumi Chunara |
AMIA | 3 |