VLDB 2026 Research / reviewers in the wild / expert
Rasika Bhalerao
dblp:201/7550 · also Rasika Vinayak Bhalerao
· DBLP profile ↗
11ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-4815-313XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 since 2021Security and privacy · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Model AI Assignments 2025abstractThe Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of thirteen AI assignments from the 2025 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment specifications and supporting resources may be found at http://modelai.gettysburg.edu Todd W. Neller, Rasika Bhalerao, Eun Kyung Ko, Vishodana Thamotharan, Lisa Zhang 0003, Sonya Allin, Mahdi Haghifam, Michael Pawliuk, Rutwa Engineer, Florian Shkurti, Cunyan Ma, Daniella DiPaola, Cynthia Breazeal, Loreto Alonzi, Brian Wright, Ali Rivera, Kristin Fasiang, Duri Long, Shruthi Chockkalingam, Giulia Toti, Evan Shieh, Princewill Okoroafor, Thema Monroe-White, Mustafa Haiderbhai, Carolyn Quinlan, Ashwin R. Bharadwaj, Anio Zhang, Rajagopal Venkatesaramani, Sarah Wharton, John Masla, Lydia Guterman, Mary Cate Gustafson-Quiett, Christina A. Bosch, Samar Abu Hegley, Calvin Macatantan, Eric Klopfer, Harold Abelson, Shira Wein, Mercy Wairimu Gachoka, Li-Hsin Chang, Maryam Mirzaei, Mohammad Mahdi Ajallooeian |
AAAI | 2 |
| 2024 | My Learnings from Allowing Large Language Models in Introductory Computer Science ClassesabstractMany instructors want to allow their students to use large language models (LLMs) in their introductory computer science courses, but they first want to see other instructors' results from doing so before taking on the risk in their own courses. Presented here are the results from allowing students to use LLMs in the second course in a sequence of intensive introductory courses designed to prepare students with a non-computational background for entry into a masters' degree program. We allowed students to use the internet and LLMs (such as ChatGPT or Github Copilot) to help with assignments, with guidelines to avoid plagiarism and encourage learning. We then surveyed students to ask about how they used LLMs, whether they saw others cheating, how they generally used internet-based resources on assignments and exams, and their feedback on the policies. We found that students are overwhelmingly using LLMs (and the internet generally) to learn and code "better" rather than cheat. These results are intended to be a starting point to spark discussion on the adoption of new technologies in introductory computer science courses. The authors themselves will continue teaching courses with the policy that students should interact with an LLM the way they interact with a person: students are encouraged to discuss and collaborate with it, but copying code from it is considered plagiarism. Rasika Bhalerao |
SIGCSE (2) | 1 |
| 2024 | What to Expect When You're Accessing: An Exploration of User Privacy Rights in People Search WebsitesabstractPeople Search Websites, a category of data brokers, collect, catalog, monetize and often publicly display individuals' personally identifiable information (PII). We present a study of user privacy rights in 20 such websites assessing the usability of data access and data removal mechanisms. We combine insights from these two processes to determine connections between sites, such as shared access mechanisms or removal effects. We find that data access requests are mostly unsuccessful. Instead, sites cite a variety of legal exceptions or misinterpret the nature of the requests. By purchasing reports, we find that only one set of connected sites provided access to the same report they sell to customers. We leverage a multiple step removal process to investigate removal effects between suspected connected sites. In general, data removal is more streamlined than data access, but not very transparent; questions about the scope of removal and reappearance of information remain. Confirming and expanding the connections observed in prior phases, we find that four main groups are behind 14 of the sites studied, indicating the need to further catalog these connections to simplify removal. Kejsi Take, Jordyn Young, Rasika Bhalerao, Kevin Gallagher 0001, Andrea Forte, Damon McCoy, Rachel Greenstadt |
Proc. Priv. Enhancing Technol. | 3 |
| 2023 | Pretraining Language Models with Human PreferencesabstractLanguage models (LMs) are pretrained to imitate text from large and diverse datasets that contain content that would violate human preferences if generated by an LM: falsehoods, offensive comments, personally identifiable information, low-quality or buggy code, among others. Here, we explore alternative objectives for pretraining LMs in a way that also guides them to generate text aligned with human preferences. We benchmark five objectives for pretraining with human feedback across three tasks and study how they affect the alignment and capabilities of pretrained LMs. We find a Pareto-optimal and simple approach among those we explored: conditional training, or learning distribution over tokens conditional on their human preference scores. Conditional training reduces the rate of undesirable content by up to an order of magnitude, both when generating without a prompt and with an adversarially-chosen prompt. Moreover, conditional training maintains the downstream task performance of standard LM pretraining, both before and after task-specific finetuning. Pretraining with human feedback results in much better preference satisfaction than standard LM pretraining followed by finetuning with feedback, i.e., learning and then unlearning undesirable behavior. Our results suggest that we should move beyond imitation learning when pretraining LMs and incorporate human preferences from the start of training. Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, Ethan Perez |
ICML | 4 |
| 2023 | The Digital-Safety Risks of Financial Technologies for Survivors of Intimate Partner Violence
Rosanna Bellini, Kevin Lee 0001, Megan A. Brown, Jeremy Shaffer, Rasika Bhalerao, Thomas Ristenpart |
USENIX Security Symposium | 5 |
| 2022 | An Analysis of Terms of Service and Official Policies with Respect to Sex WorkabstractPolicymakers who design the rules that govern the internet and the technologists who implement them can often be disconnected from some of the populations affected by their products. In this study, we analyze the terms of service, community guidelines, privacy policies, and other documents officially issued by online platforms in the United States to discuss their implications with regards to a marginalized population of interest: workers in the sex industry, ranging in autonomy from sex workers with a high degree of autonomy to survivors of sex trafficking. While criminalized and stigmatized populations such as sex industry workers are underrepresented among technologists, we show how technological decision makers without subject matter knowledge or understanding of the motivations and effects on the population can unintentionally lead to harming sex industry workers. Our analysis is in line with sex industry worker-led movements to stop arresting sex industry workers, de-stigmatize sex work, and let sex industry workers remain and flourish in online life. We study over 100 online platforms from 13 platform types and discuss the laws, perceptions, and motivations behind their policies regarding the sex industry, and how these policies affect sex industry workers. We find that platforms generally view sex industry workers as either criminals, victims, spam, or entrepreneurs; we show how using the first three paradigms to characterize the entire industry can lead to stigmatization, overly general and restrictive rules, and decreased accessibility to online life. We use this study as an example to illustrate the need for a cultural shift in the technology community towards empathy and social education and provide concrete research directions towards a solution. Rasika Bhalerao, Damon McCoy |
ISTAS | 1 |
| 2022 | Learning Outcomes and Assessments for Ethical ComputingabstractThe inclusion of computing ethics and social impact curricula in computer science programs has received increasing attention in research and teaching. While students explicitly interested in the topic may choose to take courses designed only to teach ethical thinking in computer science, we have a societal need to "teach ethics" in a wider variety of computing courses. With the powerful tools that we give students comes responsibility, and students should know how to consider ethical implications of the things they build. The difficulty of developing strategies to "teach ethics" in computing courses is that 1) teaching computing ethics is different than teaching other computational courses, requiring more than the simple transmission of information or knowledge 2) we lack a way to assess the efficacy of the strategies we use to accomplish the "more." We aim to foster a discussion on the current research and instructional approaches, led by researchers who are actively teaching responsible computer science and data science in a variety of institutional settings (e.g., private, public, small, large, etc.). In addition to forming new research and teaching collaborations, we hope this discussion will inspire the larger community of computer science educators to embed computing ethics and social impact in their technical courses. Rasika Bhalerao, Emanuelle Burton, Stacy A. Doore, Judy Goldsmith |
SIGCSE (2) | 1 |
| 2022 | Ethics and Efficacy of Unsolicited Anti-Trafficking SMS OutreachabstractThe sex industry exists on a continuum based on the degree of work autonomy present in one's labor conditions: a high degree of autonomy exists on one side of the continuum where certain independent sex workers have a great deal of agency, while much less autonomy exists on the other side, where sex is traded under conditions of human trafficking. Various organizations across North America perform outreach to sex workers to offer assistance in the form of services (e.g., healthcare, financial assistance, housing) as well as prayer and intervention. Increasingly, technology is used to look for trafficking victims and/or facilitate the provision of assistance or services, for example through scraping and parsing sex industry workers' advertisements into a database of contact information that can be used by outreach organizations. However, little is known about the efficacy of anti-trafficking outreach technology, nor the potential risks of using such technology to identify and contact the highly stigmatized and marginalized population of those working in the sex industry. In this work, we investigate the use, context, benefits, and harms of an anti-trafficking technology platform via qualitative interviews with multiple stakeholders: the technology developers (n=6), organizations that use the technology (n=17), and sex industry workers who have been contacted or wish to be contacted (n=24). Our findings illustrate misalignment between developers, users of the platform, and sex industry workers they are attempting to assist. In their current state, anti-trafficking outreach tools such as the one we investigate are ineffective and, at best, serve as a mechanism for spam and, at worst, scale and exacerbate harm against the population they aim to serve. We conclude with a discussion of best practices -- and the feasibility of their implementation -- for technology-facilitated outreach efforts to minimize risk or harm to sex industry workers while efficiently providing needed services. Rasika Bhalerao, Nora McDonald, Hanna Barakat, Vaughn Hamilton, Damon McCoy, Elissa M. Redmiles |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2021 | Swiped: Analyzing Ground-truth Data of a Marketplace for Stolen Debit and Credit Cards
Max Aliapoulios, Cameron Ballard, Rasika Bhalerao, Tobias Lauinger, Damon McCoy |
USENIX Security Symposium | 3 |
| 2020 | CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsabstractWarning: This paper contains explicit statements of offensive stereotypes and may be upsetting.Pretrained language models, especially masked language models (MLMs) have seen success across many NLP tasks.However, there is ample evidence that they use the cultural biases that are undoubtedly present in the corpora they are trained on, implicitly creating harm with biased representations.To measure some forms of social bias in language models against protected demographic groups in the US, we introduce the Crowdsourced Stereotype Pairs benchmark (CrowS-Pairs).CrowS-Pairs has 1508 examples that cover stereotypes dealing with nine types of bias, like race, religion, and age.In CrowS-Pairs a model is presented with two sentences: one that is more stereotyping and another that is less stereotyping.The data focuses on stereotypes about historically disadvantaged groups and contrasts them with advantaged groups.We find that all three of the widelyused MLMs we evaluate substantially favor sentences that express stereotypes in every category in CrowS-Pairs.As work on building less biased models advances, this dataset can be used as a benchmark to evaluate progress. Nikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. Bowman |
EMNLP (1) | 3 |
| 2017 | Pre-processing and classification of hyperspectral imagery via selective inpaintingabstractWe propose a semi-supervised algorithm for processing and classification of hyperspectral imagery. For initialization, we keep 20% of the data intact, and use Principal Component Analysis to discard voxels from noisier bands and pixels. Then, we use either an Accelerated Proximal Gradient algorithm (APGL), or a modified APGL algorithm with a penalty term for distance between inpainted pixels and endmembers (APGL Hyp), on the initialized datacube to inpaint the missing data. APGL and APGL Hyp are distinguished by performance on datasets with full pixels removed or extreme noise. This inpainting technique results in band-by-band datacube sharpening and removal of noise from individual spectral signatures. We can also classify the inpainted cube by assigning each pixel to its nearest endmember via Euclidean distance. We demonstrate improved accuracy in classification over data-mining techniques like k-means, unmixing techniques like Hierarchical Non-Negative Matrix Factorization, and graph-based methods like Non-Local Total Variation. Victoria Chayes, Rasika Bhalerao, Wei Zhu 0007, Andrea L. Bertozzi, Wenzi Liao, Stanley J. Osher |
ICASSP | 3 |