Saloni Dash

dblp:250/1496 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0005-7292-8275ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Anecdoctoring: Automated Red-Teaming Across Language and Place
abstract
Disinformation is among the top risks of generative artificial intelligence (AI) misuse.Global adoption of generative AI necessitates redteaming evaluations (i.e., systematic adversarial probing) that are robust across diverse languages and cultures, but red-teaming datasets are commonly US-and English-centric.To address this gap, we propose "anecdoctoring", a novel red-teaming approach that automatically generates adversarial prompts across languages and cultures.We collect misinformation claims from fact-checking websites in three languages (English, Spanish, and Hindi) and two geographies (US and India).We then cluster individual claims into broader narratives and characterize the resulting clusters with knowledge graphs, with which we augment an attacker LLM.Our method produces higher attack success rates and offers interpretability benefits relative to few-shot prompting.Results underscore the need for disinformation mitigations that scale globally and are grounded in realworld adversarial misuse.
Alejandro Cuevas, Saloni Dash, Bharat Nayak, Dan Vann, Madeleine I. G. Daepp
EMNLP2
2022 Poster: Leveraging Question Answering to Understand Context Specific Patterns in Fact Checked Articles in the Global South
abstract
Propagation of misinformation on various social media platforms is a common occurrence, especially around political events, religious beliefs, and public health. Fact checked articles, which investigate the credibility of dubious claims online, provide a reliable source of debunked misinformation. However, existing (older) fact checked articles remain an underutilized resource for understanding patterns in fake stories. We propose the use of Question Answering (QA) for analysing fact checked articles for systematically extracting metadata, potentially useful for downstream tasks such as misinformation detection, using a range of simple to nuanced questions. We find that the method gives us a context-specific understanding of common patterns and themes in misinformation, which is especially important in the Global South, where misinformation is layered with propagandist underpinning. Our findings suggest that this method can be extended by fine tuning on any event specific data set of fact checked articles to yield more robust and accurate results.
Arshia Arya, Saloni Dash, Syeda Zainab Akbar, Joyojeet Pal, Anirban Sen
COMPASS2
2022 Insights Into Incitement: A Computational Perspective on Dangerous Speech on Twitter in India
abstract
Dangerous speech on social media platforms can be framed as blatantly inflammatory, or be couched in innuendo. It is also centrally tied to who engages it – it can be driven by openly sectarian social media accounts, or through subtle nudges by influential accounts, allowing for complex means of reinforcing vilification of marginalized groups, an increasingly significant problem in the media environment in the Global South. We identify dangerous speech by influential accounts on Twitter in India around three key events, examining both the language and networks of messaging that condones or actively promotes violence against vulnerable groups. We characterize dangerous speech users by assigning Danger Amplification Belief scores and show that dangerous users are more active on Twitter as compared to other users as well as most influential in the network, in terms of a larger following as well as volume of verified accounts. We find that dangerous users have a more polarized viewership, suggesting that their audience is more susceptible to incitement. Using a mix of network centrality measures and qualitative analysis, we find that most dangerous accounts tend to either be in mass media related occupations or allied with low-ranking, right-leaning politicians, and act as “broadcasters” in the network, where they are best positioned to spearhead the rapid dissemination of dangerous speech across the platform.
Saloni Dash, Rynaa Grover, Gazal Shekhawat, Sukhnidh Kaur, Dibyendu Mishra, Joyojeet Pal
COMPASS1
2022 DISMISS: Database of Indian Social Media Influencers on Twitter
Arshia Arya, Soham De, Dibyendu Mishra, Gazal Shekhawat, Anmol Panda, Faisal M. Lalani, Parantak Singh, Ramaravind Kommiya Mothilal, Rynaa Grover, Sachita Nishal, Saloni Dash, Shehla Rashid Shora, Syeda Zainab Akbar, Joyojeet Pal
ICWSM12
2022 Divided We Rule: Influencer Polarization on Twitter during Political Crises in India
Saloni Dash, Dibyendu Mishra, Gazal Shekhawat, Joyojeet Pal
ICWSM1
2022 Evaluating and Mitigating Bias in Image Classifiers: A Causal Perspective Using Counterfactuals
abstract
Counterfactual examples for an input—perturbations that change specific features but not others—have been shown to be useful for evaluating bias of machine learning models, e.g., against specific demographic groups. However, generating counterfactual examples for images is nontrivial due to the underlying causal structure on the various features of an image. To be meaningful, generated perturbations need to satisfy constraints implied by the causal model. We present a method for generating counterfactuals by incorporating a structural causal model (SCM) in an improved variant of Adversarially Learned Inference (ALI), that generates counterfactuals in accordance with the causal relationships between attributes of an image. Based on the generated counterfactuals, we show how to explain a pre-trained machine learning classifier, evaluate its bias, and mitigate the bias using a counterfactual regularizer. On the Morpho-MNIST dataset, our method generates counterfactuals comparable in quality to prior work on SCM-based counterfactuals (DeepSCM), while on the more complex CelebA dataset our method outperforms DeepSCM in generating high-quality valid counterfactuals. Moreover, generated counterfactuals are indistinguishable from reconstructed images in a human evaluation experiment and we subsequently use them to evaluate the fairness of a standard classifier trained on CelebA data. We show that the classifier is biased w.r.t. skin and hair color, and how counterfactual regularization can remove those biases.
Saloni Dash, Vineeth N. Balasubramanian, Amit Sharma 0007
WACV1
2021 Quantifying Resemblance of Synthetic Medical Time-Series
abstract
Access to medical data is often restricted due to privacy laws e.g.HIPAA and GDPR.We address the viability of substituting real data with synthetic data to protect privacy while maintaining utility.Medical data records are fundamentally longitudinal, with one patient having multiple health events influenced by covariates like gender, age etc. Synthesis of medical data, hence, falls under time-series generative modeling.We demonstrate methods to measure synthetic medical time-series quality on datasets from previously published synthetic data research.We deploy four time-series metrics to quantify resemblance in synthetic and real covariate plots while comparing baseline data generation methods.
Karan Bhanot, Saloni Dash, Joseph Pedersen, Isabelle Guyon, Kristin P. Bennett
ESANN2
2020 Medical Time-Series Data Generation Using Generative Adversarial Networks
Saloni Dash, Andrew Yale, Isabelle Guyon, Kristin P. Bennett
AIME1
2020 Generation and evaluation of privacy preserving synthetic health data
Andrew Yale, Saloni Dash, Ritik Dutta, Isabelle Guyon, Adrien Pavão, Kristin P. Bennett
Neurocomputing2
2019 Privacy Preserving Synthetic Health Data
Andrew Yale, Saloni Dash, Ritik Dutta, Isabelle Guyon, Adrien Pavão, Kristin P. Bennett
ESANN2