Ritwik Sinha

dblp:127/3163 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Image Difference Captioning via Adversarial Preference Optimization
abstract
Image Difference Captioning (IDC) aims to generate natural language descriptions that highlight subtle differences between two visually similar images.While recent advances leverage pre-trained vision-language models to align fine-grained visual differences with textual semantics, existing supervised approaches often overly focus on dataset-specific language patterns and fail to capture fine-grained and context-aware preferences on IDC, due to limited annotation diversity and a lack of semantically informative negative examples during training, To address these limitations, we propose an adversarial direct preference optimization (ADPO) framework for IDC, which formulates IDC as a preference optimization problem under the Bradley-Terry-Luce model, directly aligning the captioning policy with pairwise difference preferences via Direct Preference Optimization (DPO).To model more accurate and diverse IDC preferences, we introduce an adversarially trained hard negative retriever that selects counterfactual captions, This results in a minimax optimization problem, which we solve via policy-gradient reinforcement learning, enabling the policy and retriever to improve jointly.By dynamically generating semantically challenging negatives, our method reduces reliance on dataset-specific patterns.Experiments on benchmark IDC datasets show that our approach outperforms existing baselines, especially in generating fine-grained and accurate difference descriptions.
Zihan Huang, Junda Wu, Rohan Surana, Tong Yu 0001, David T. Arbour, Ritwik Sinha, Julian J. McAuley
EMNLP6
2025 Diversify-verify-adapt: Efficient and Robust Retrieval-Augmented Ambiguous Question Answering
abstract
Yeonjun In, Sungchul Kim, Ryan A. Rossi, Mehrab Tanjim, Tong Yu, Ritwik Sinha, Chanyoung Park. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yeonjun In, Sungchul Kim, Ryan Rossi, Md. Mehrab Tanjim, Tong Yu 0001, Ritwik Sinha, Chanyoung Park 0001
NAACL (Long Papers)6
2025 Leveraging semantic similarity for experimentation with AI-generated treatments
abstract
Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main methodological challenge in this setting is representing these high-dimensional treatments without losing their semantic meaning or rendering analysis intractable. Here we address this problem by focusing on learning low-dimensional representations that capture the underlying structure of such treatments. These representations enable downstream applications such as guiding generative models to produce meaningful treatment variants and facilitating adaptive assignment in online experiments. We propose double kernel representation learning, which models the causal effect through the inner product of kernel-based representations of treatments and user covariates. We develop an alternating-minimization algorithm that learns these representations efficiently from data and provide convergence guarantees under a low-rank factor model. As an application of this framework, we introduce an adaptive design strategy for online experimentation and demonstrate the method's effectiveness through numerical experiments.
David T. Arbour, Raghavendra Addanki, Ritwik Sinha, Avi Feller
NeurIPS4
2025 Experimentation under Treatment Dependent Network Interference
abstract
Randomized Controlled Trials (RCTs) are a fundamental aspect of data-driven decision-making. RCTs often assume that the units are not influenced by each other. Traditional approaches addressing such effects assume a fixed network structure between the interfering units. However, real-world networks are rarely static, and treatment assignments can actively reshape the interference structure itself, as seen in financial access interventions that alter informal lending networks or healthcare programs that modify peer influence dynamics. This creates a novel and unexplored problem: estimating treatment effects when the interference network is determined by treatment allocation. In this work, we address this gap by proposing two single-experiment estimators for scenarios where network edges depend on nodal treatments constructed from instrumental variables derived from neighbourhood treatments. We prove their unbiasedness and experimentally validate the proposed estimators both on synthetic and real data.
Shiv Shankar, Ritwik Sinha, Madalina Fiterau
UAI2
2024 A/B testing under Interference with Partial Network Information
abstract
A/B tests are often required to be conducted on subjects that might have social connections. For e.g., experiments on social media, or medical and social interventions to control the spread of an epidemic. In such settings, the SUTVA assumption for randomized-controlled trials is violated due to network interference, or spill-over effects, as treatments to group A can potentially also affect the control group B. When the underlying social network is known exactly, prior works have demonstrated how to conduct A/B tests adequately to estimate the global average treatment effect (GATE). However, in practice, it is often impossible to obtain knowledge about the exact underlying network. In this paper, we present UNITE: a novel estimator that relax this assumption and can identify GATE while only relying on knowledge of the superset of neighbors for any subject in the graph. Through theoretical analysis and extensive experiments, we show that the proposed approach performs better in comparison to standard estimators.
Shiv Shankar, Ritwik Sinha, Yash Chandak, Saayan Mitra, Madalina Fiterau
AISTATS2
2024 On Online Experimentation without Device Identifiers
abstract
Measuring human feedback via randomized experimentation is a cornerstone of data-driven decision-making. The methodology used to estimate user preferences from their online behaviours is critically dependent on user identifiers. However, in today's digital landscape, consumers frequently interact with content across multiple devices, which are often recorded with different identifiers for the same consumer. The inability to match different device identities across consumers poses significant challenges for accurately estimating human preferences and other causal effects. Moreover, without strong assumptions about the device-user graph, the causal effects might not be identifiable. In this paper, we propose HIFIVE, a variational method to solve the problem of estimating global average treatment effects (GATE) from a fragmented view of exposures and outcomes. Experiments show that our estimator is superior to standard estimators, with a lower bias and greater robustness to network uncertainty.
Shiv Shankar, Ritwik Sinha, Madalina Fiterau
ICML2
2024 Continuous Treatment Effects with Surrogate Outcomes
abstract
In many real-world causal inference applications, the primary outcomes (labels) are often partially missing, especially if they are expensive or difficult to collect. If the missingness depends on covariates (i.e., missingness is not completely at random), analyses based on fully observed samples alone may be biased. Incorporating surrogates, which are fully observed post-treatment variables related to the primary outcome, can improve estimation in this case. In this paper, we study the role of surrogates in estimating continuous treatment effects and propose a doubly robust method to efficiently incorporate surrogates in the analysis, which uses both labeled and unlabeled data and does not suffer from the above selection bias problem. Importantly, we establish the asymptotic normality of the proposed estimator and show possible improvements on the variance compared with methods that solely use labeled data. Extensive simulations show our methods enjoy appealing empirical performance.
Zhenghao Zeng, David T. Arbour, Avi Feller, Raghavendra Addanki, Ryan Rossi, Ritwik Sinha, Edward H. Kennedy
ICML6
2024 Discovering and Mitigating Biases in CLIP-based Image Editing
abstract
In recent years, the use of CLIP (Contrastive Language-Image Pre-Training) has become increasingly popular in a wide range of downstream applications, including zero-shot image classification and text-to-image synthesis. Despite being trained on a vast dataset, the CLIP model has been found to exhibit biases against certain protected attributes, such as gender and race. While previous research has focused on the impact of such biases on image classification, there has been little investigation into their effects on CLIP-based generative tasks. In this paper, we aim to address this gap in the literature by uncovering the queries for which the CLIP model introduces biases in the text-based image editing task. Through a series of experiments, we demonstrate that these biases can have a significant impact on the quality and content of the generated images. To mitigate these biases, we propose a debiasing technique that does not require retraining either the CLIP model or the underlying generative model. Our results show that our proposed framework can effectively reduce the impact of biases in CLIP-based image editing models. Overall, this paper highlights the importance of addressing biases in CLIP-based generative tasks and provides practical solutions that can be readily adopted by researchers and practitioners working in this area.1
Md. Mehrab Tanjim, Krishna Kumar Singh, Kushal Kafle, Ritwik Sinha, Garrison W. Cottrell
WACV4
2023 Direct Inference of Effect of Treatment (DIET) for a Cookieless World
abstract
Brands use cookies and device identifiers to link different web visits to the same consumer. However, with increasing demands for privacy, these identifiers are about to be phased out, making identity fragmentation a permanent feature of the online world. Assessing treatment effects via randomized experiments (A/B testing) in such a scenario is challenging because identity fragmentation causes a) users to receive hybrid/mixed treatments, and b) hides the causal link between the historical treatments and the outcome. In this work, we address the problem of estimating treatment effects when a lack of identification leads to incomplete knowledge of historical treatments. This is a challenging problem which has not been addressed in literature yet. We develop a new method called DIET, which can adjust for users being exposed to mixed treatments without the entire history of treatments being available. Our method takes inspiration from the Cox model, and uses a proportional outcome approach under which we prove that one can obtain consistent estimates of treatment effects even under identity fragmentation. Our experiments, on one simulated and two real datasets, show that our method leads to up to 20% reduction in error and 25% reduction in bias over the naive estimate.
Shiv Shankar, Ritwik Sinha, Saayan Mitra, Moumita Sinha, Madalina Fiterau
AISTATS2
2023 VADER: Video Alignment Differencing and Retrieval
abstract
We propose VADER, a spatio- temporal matching, alignment, and change summarization method to help fight misinformation spread via manipulated videos. VADER matches and coarsely aligns partial video fragments to candidate videos using a robust visual descriptor and scalable search over adaptively chunked video content. A transformer- based alignment module then refines the temporal localization of the query fragment within the matched video. A space- time comparator module identifies regions of manipulation between aligned content, invariant to any changes due to any residual temporal misalignments or artifacts arising from non- editorial changes of the content. Robustly matching video to a trusted source enables conclusions to be drawn on video provenance, enabling informed trust decisions on content encountered. Code and data are available at https://github.com/AlexBlck/vader
Alexander Black 0001, Simon Jenni, Tu Bui, Md. Mehrab Tanjim, Stefano Petrangeli, Ritwik Sinha, Viswanathan (Vishy) Swaminathan, John P. Collomosse
ICCV6
2023 SODA: Protecting Proprietary Information in On-Device Machine Learning Models
abstract
The growth of low-end hardware has led to a proliferation of machine learning-based services in edge applications. These applications gather contextual information about users and provide some services, such as personalized offers, through a machine learning (ML) model. A growing practice has been to deploy such ML models on the user's device to reduce latency, maintain user privacy, and minimize continuous reliance on a centralized source. However, deploying ML models on the user's edge device can leak proprietary information about the service provider. In this work, we investigate on-device ML models that are used to provide mobile services and demonstrate how simple attacks can leak proprietary information of the service provider. We show that different adversaries can easily exploit such models to maximize their profit and accomplish content theft. Motivated by the need to thwart such attacks, we present an end-to-end framework, SODA, for deploying and serving on edge devices while defending against adversarial usage. Our results demonstrate that SODA can detect adversarial usage with 89% accuracy in less than 50 queries with minimal impact on service performance, latency, and storage.
Akanksha Atrey, Ritwik Sinha, Saayan Mitra, Prashant J. Shenoy
SEC2
2023 Privacy Aware Experiments without Cookies
abstract
Consider two brands that want to jointly test alternate web experiences for their customers with an A/B test. Such collaborative tests are today enabled usingthird-party cookies, where each brand has information on the identity of visitors to another website, ensuring a consistent treatment experience. With the imminent elimination of third-party cookies, such A/B tests will become untenable. We propose a two-stage experimental design, where the two brands only need to agree on high-level aggregate parameters of the experiment to test the alternate experiences. Our design respects the privacy of customers. We propose an unbiased estimator of the Average Treatment Effect (ATE), and provide a way to use regression adjustment to improve this estimate. On real and simulated data, we show that the approach provides valid estimate of the ATE and is robust to the proportion of visitors overlapping across the brands. Our demonstration describes how a marketer can design such an experiment and analyze the results.
Shiv Shankar, Ritwik Sinha, Saayan Mitra, Viswanathan (Vishy) Swaminathan, Sridhar Mahadevan, Moumita Sinha
WSDM2
2022 Debiasing Image-to-Image Translation Models
Md. Mehrab Tanjim, Krishna Kumar Singh, Kushal Kafle, Ritwik Sinha, Garrison W. Cottrell
BMVC4
2022 Contextualized Styling of Images for Web Interfaces using Reinforcement Learning
abstract
Content personalization is one of the foundations of today’s digital marketing. Often the same image needs to be adapted for different design schemes for content that is created for different occasions, geographic locations or other aspects of the target population. We present a novel reinforcement learning (RL) based method for automatically stylizing images to complement the design scheme of media, e.g., interactive websites, apps, or posters. Our approach considers attributes related to the design of the media and adapts the style of the input image to match the context. We do so using a preferential reward system in the RL framework that learns a reward function using human feedback. We conducted several user studies to evaluate our approach and demonstrate that we are able to effectively adapt image styles to different design schemes. In user studies, images stylized through our approach were the most preferred variation across a majority of our experiments. Additionally, we also release a dataset consisting of perceptual associations of web context with the associated image style.
Pooja Guhan, Saayan Mitra, Somdeb Sarkhel, Stefano Petrangeli, Ritwik Sinha, Viswanathan (Vishy) Swaminathan, Aniket Bera, Dinesh Manocha
ISM5
2022 Machine Learning-Based Hard/Soft Logic Trade-offs in VTR
abstract
Circuit optimization, in any application, is of high importance since it not only improves the efficiency of the intended purpose but also enhances the quality of the final product. It enables the circuit designer to cater to the specific needs of the customer. For circuit optimization to occur, we need to elaborate these circuits on a primary level and perform synthesis operations. Previous research shows that the investigation of improvements to different Hardware Description Language (HDL) elaboration phases, was completely closed source. Verilog To Routing (VTR) is an open-source Electronic Design Automation (EDA) tool. ODIN II is the VTR synthesizer that parses the input Verilog, elaborates its Abstract Syntax Tree (AST), performs the partial mapping according to the architecture file, and performs optimizations such as unused logic removal. To that end, the hard versus soft logic trade-off aims to optimize the performance of the circuit. This project focuses on using machine learning approaches to make synthesis tools intelligent enough to decide this ratio on their own, without the need for human intervention, and based on some predefined criteria. This paper discusses the criteria for having less latency or less critical path delay in the circuit. Also, it aims at providing this level of intelligence at an earlier stage in the VTR pipeline to make better use of this information.
Ritwik Sinha, Seyed Alireza Damghani, Kenneth B. Kent
RSP1
2022 Generating and Controlling Diversity in Image Search
abstract
In our society, generations of systemic biases have led to some professions being more common among certain genders and races. This bias is also reflected in image search on stock image repositories and search engines, e.g., a query like “male Asian administrative assistant” may produce limited results. The pursuit of a utopian world demands providing content users with an opportunity to present any profession with diverse racial and gender characteristics. The limited choice of existing content for certain combinations of profession, race, and gender presents a challenge to content providers. Current research dealing with bias in search mostly focuses on re-ranking algorithms. However, these methods cannot create new content or change the overall distribution of protected attributes in photos. To remedy these problems, we propose a new task of high-fidelity image generation conditioning on multiple attributes from imbalanced datasets. Our proposed task poses new sets of challenges for the state-of-the-art Generative Adversarial Networks (GANs). In this paper, we also propose a new training framework to better address the challenges. We evaluate our framework rigorously on a real-world dataset and perform user studies that show our model is preferable to the alternatives.
Md. Mehrab Tanjim, Ritwik Sinha, Krishna Kumar Singh, Sridhar Mahadevan, David T. Arbour, Moumita Sinha, Garrison W. Cottrell
WACV2
2021 BOhance: Bayesian Optimization for Content Enhancement
abstract
We present BOhance, an efficient solution for optimizing digital content like images. Our approach enhances the standard and widely-used method for optimizing content, A/B testing, by using Bayesian Optimization. Our work effectively extends A/B testing in the continuous domain where A/B testing cannot efficiently test infinitely many variants. We test our approach on an image enhancement task where we use iterative human feedback on different variants of an image to arrive at the optimal variant. BOhance auto-generates candidate content variants to be tested based on the human feedback on prior variants. We demonstrate with user-studies conducted on Amazon Mechanical Turk that BOhance can be both time and cost-efficient; and a superior alternative to existing solutions. Furthermore, we conduct a Visual Turing Test to obtain human impressions on the optimum variants generated by BOhance. Our experiments show that given a human-enhanced image and an image generated by BOhance, 53% users think that the BOhance image was generated by a human expert.
Trisha Mittal, Viswanathan (Vishy) Swaminathan, Somdeb Sarkhel, Ritwik Sinha, David T. Arbour, Saayan Mitra, Dinesh Manocha
ISM4
2020 Attribution IQ: Scalable Game Theoretic Attribution in Web Analytics
abstract
Attribution in digital marketing is the task of assigning the credit due to each marketing interaction toward a marketing outcome. Such information helps the brand decide on marketing strategies for the future. In web analytics, attribution goes beyond marketing channels, and can be performed across a broad range of dimensions (e.g. images displayed on a website). This requires attribution algorithms to operate on dimensions with a large number of dimensional elements (hundreds of thousands in the extreme). Additionally, given the many possible metrics a marketer may be interested in, it is infeasible to perform ahead of time computation for all combinations. In this work, we propose a game theoretic attribution model that can be computed at query-time. Our demo of Attribution IQ runs in the Adobe Analytics production system. Our system is highly scalable (it operates on millions of customer journeys), interactive (the user can change the parameters and get updated results instantaneously) and requires no pre-computation (all computation being performed at query time). We demonstrate the effectiveness of Attribution IQ on real-world e-commerce web analytics datasets.
Ritwik Sinha, Ivan Andrus, Trevor Paulsen
CIKM1
2020 Bayesian Estimation of the Effect of Television Advertising on Web Metrics
abstract
Aggregate advertising-presenting a single ad to large groups of individuals through traditional media such as television and print-presents a unique challenge to measuring efficacy because treatment and outcome are observed from two disparate sources (interaction and revenue realization). In this work, we propose a Bayesian model to estimate the impact of an ad on observable web metrics that are readily available in many modern analytics suites. The proposed model controls for three sources of possible confounding: the time, geography, and content of the advertisement. The proposed model is easily applicable to a wide variety of problems and readily generates error bounds for the estimates. We evaluate our approach on a real dataset for a set of TV ads for an advertiser.
Ritwik Sinha, Shiv Kumar Saini, Moumita Sinha, David T. Arbour
DSAA1
2019 RAPID: Rapid and Precise Interpretable Decision Sets
abstract
Interpretable Decision Sets (IDS) is an approach to building transparent and interpretable supervised machine learning models. Unfortunately, IDS does not scale to most commonly encountered big data sets. In this paper, we propose Rapid And Precise Interpretable Decision Sets (RAPID), a faster alternative to IDS. We use the existing formulation of decision set learning and propose a time-efficient learning framework. RAPID has two major improvements over IDS. First, it uses a linear-time randomized Unconstrained Submodular Maximization algorithm to optimize the objective function. Second, we design special data structures, based on Frequent-Pattern (FP) trees to achieve better computational efficiency. In this work, we first perform a time complexity analysis of IDS and RAPID, and show the significant advantages of the proposed method. Next we run our algorithm, along with baselines, on three public datasets. We show comparable accuracy for RAPID, with 10, 000x improvement in running time over IDS. Additionally, due to the significant improvements in running time of RAPID, we can run more extensive hyperparameter search algorithms, leading to comparable accuracy with competitive baseline models.
Sunny Dhamnani, Dhruv Singal, Ritwik Sinha, Tharun Mohandoss, Manish Dash
IEEE BigData3
2019 On Densification for Minwise Hashing
Tung Mai, Anup B. Rao, Matt Kapilevich, Ryan Rossi, Yasin Abbasi-Yadkori, Ritwik Sinha
UAI6
2018 Saliency Prediction for Mobile User Interfaces
abstract
We introduce models for saliency prediction for mobile user interfaces. A mobile interface may include elements like buttons and text in addition to natural images which enable performing a variety of tasks. Saliency in natural images is a well studied topic. However, given the difference in what constitutes a mobile interface, and the usage context of these devices, we postulate that saliency prediction for mobile interface images requires a fresh approach. Mobile interface design involves operating on elements, the building blocks of the interface. We first collected eye-gaze data from mobile devices for a free viewing task. Using this data, we develop a novel autoencoder based multi-scale deep learning model that provides saliency prediction at the mobile interface element level. Compared to saliency prediction approaches developed for natural images, we show that our approach performs significantly better on a range of established metrics.
Prakhar Gupta, Shubh Gupta, Ajaykrishnan Jayagopal, Sourav Pal, Ritwik Sinha
WACV5
2015 Improving Marketing Interactions by Mining Sequences
Ritwik Sinha, Sanket Mehta, Tapan Bohra, Adit Krishnan
WISE (1)1
2015 A Non-parametric Approach to the Multi-channel Attribution Problem
Meghanath Macha Yadagiri, Shiv Kumar Saini, Ritwik Sinha
WISE (1)3