EDBT 2026 Demo / reviewers in the wild / expert
Atanu R. Sinha
dblp:120/2076
· DBLP profile ↗
14ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-7949-7982ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Subjective Behaviors and Preferences in LLM: Language of BrowsingabstractSai Sundaresan, Harshita Chopra, Atanu R. Sinha, Koustava Goswami, Nagasai Saketh Naidu, Raghav Karan, N Anushka. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Sai Sundaresan, Harshita Chopra, Atanu R. Sinha, Koustava Goswami, Nagasai Saketh Naidu, Raghav Karan, N. Anushka |
EMNLP | 3 |
| 2025 | Guidance Source Matters: How Guidance from AI, Expert, or a Group of Analysts Impacts Visual Data Preparation and AnalysisabstractThe progress in generative AI has fueled AI-powered tools like co-pilots and assistants to provision better guidance, particularly during data analysis. However, research on guidance has not yet examined the perceived efficacy of the source from which guidance is offered and the impact of this source on the user's perception and usage of guidance. We ask whether users perceive all guidance sources as equal, with particular interest in three sources: (i) AI, (ii) human expert, and (iii) a group of human analysts. As a benchmark, we consider a fourth source, (iv) unattributed guidance, where guidance is provided without attribution to any source, enabling isolation of and comparison with the effects of source-specific guidance. We design a five-condition between-subjects study, with one condition for each of the four guidance sources and an additional (v) no-guidance condition, which serves as a baseline to evaluate the influence of any kind of guidance. We situate our study in a custom data preparation and analysis tool wherein we task users to select relevant attributes from an unfamiliar dataset to inform a business report. Depending on the assigned condition, users can request guidance, which the system then provides in the form of attribute suggestions. To ensure internal validity, we control for the quality of guidance across source-conditions. Through several metrics of usage and perception, we statistically test five preregistered hypotheses and report on additional analysis. We find that the source of guidance matters to users, but not in a manner that matches received wisdom. For instance, users utilize guidance differently at various stages of analysis, including expressing varying levels of regret, despite receiving guidance of similar quality. Notably, users in the AI condition reported both higher post-task benefit and regret. Arpit Narechania, Alex Endert, Atanu R. Sinha |
IUI | 3 |
| 2025 | Handling Missing Responses under Cluster Dependence with Applications to Language Model EvaluationabstractHuman annotations play a crucial role in evaluating the performance of GenAI models. Two common challenges in practice, however, are missing annotations (the response variable of interest) and cluster dependence among human-AI interactions (e.g., questions asked by the same user may be highly correlated). Reliable inference must address both issues to achieve unbiased estimation and appropriately quantify uncertainty when estimating average scores from human annotations. In this paper, we analyze the doubly robust estimator, a widely used method in missing data analysis and causal inference, applied to this setting and establish novel theoretical properties under cluster dependence. We further illustrate our findings through simulations and a real-world conversation quality dataset. Our theoretical and empirical results underscore the importance of incorporating cluster dependence in missing response problems to perform valid statistical inference. Zhenghao Zeng, David T. Arbour, Avi Feller, Ishita Dasgupta 0002, Atanu R. Sinha, Edward H. Kennedy |
NeurIPS | 5 |
| 2024 | Fundamental Limits of Throughput and Availability: Applications to prophet inequalities and transaction fee mechanism designabstractThis paper studies the fundamental limits of availability and throughput for independent and heterogeneous demands of a limited resource. Availability is the probability that the demands are below the capacity of the resource. Throughput is the expected fraction of the resource that is utilized by the demands. We offer a concentration inequality generator that gives lower bounds on feasible availability and throughput pairs with a given capacity and independent but not necessarily identical distributions of up-to-unit demands. We show that availability and throughput cannot both be poor. These bounds are analogous to tail inequalities on sums of independent random variables, but hold throughout the support of the demand distribution. This analysis gives analytically tractable bounds supporting the unit-demand characterization of Chawla et al. [2023] and generalizes to up-to-unit demands. Our bounds also provide an approach towards improved multi-unit prophet inequalities [Hajiaghayi et al., 2007]. They have applications to transaction fee mechanism design (for blockchains) where high availability limits the probability of profitable user-miner coalitions [Chung and Shi, 2023]. Aadityan Ganesh, Jason D. Hartline, Atanu R. Sinha, Matthew vonAllmen |
EC | 3 |
| 2023 | DataCockpit: A Toolkit for Data Lake Navigation and Monitoring Utilizing Quality and Usage InformationabstractModern organizations amass their datasets into centralized repositories called data lakes, affording analytics as needed. The resultant scale and complexity of these data lakes, however, can make data navigation and monitoring challenging for users. We present DataCockpit, a Python toolkit that leverages datasets, usage logs, and associated meta-data to provision data usage and quality characteristics. DataCockpit computes these characteristics for each attribute (e.g., number of times it was queried for subsequent use in downstream applications) and record (e.g., number of non-missing, valid values) and aggregates them at the level of datasets. We develop a visual monitoring tool, powered by DataCockpit, and demonstrate how it can assist data / system administrators as well as end-users to effectively navigate and monitor a data lake. DataCockpit and the monitoring tool are available as open source software for developers to build custom monitoring applications on top of data lakes. Arpit Narechania, Surya Chakraborty, Shivam Agarwal, Atanu R. Sinha, Ryan Rossi, Fan Du, Jane Hoffswell, Shunan Guo, Eunyee Koh, Alex Endert, Shamkant B. Navathe |
IEEE Big Data | 4 |
| 2023 | DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data PreparationabstractSelecting relevant data subsets from large, unfamiliar datasets can be difficult. We address this challenge by modeling and visualizing two kinds of auxiliary information: (1) quality – the validity and appropriateness of data required to perform certain analytical tasks; and (2) usage – the historical utilization characteristics of data across multiple users. Through a design study with 14 data workers, we integrate this information into a visual data preparation and analysis tool, DataPilot. DataPilot presents visual cues about “the good, the bad, and the ugly” aspects of data and provides graphical user interface controls as interaction affordances, guiding users to perform subset selection. Through a study with 36 participants, we investigate how DataPilot helps users navigate a large, unfamiliar tabular dataset, prepare a relevant subset, and build a visualization dashboard. We find that users selected smaller, effective subsets with higher quality and usage, and with greater success and confidence. Arpit Narechania, Fan Du, Atanu R. Sinha, Ryan Rossi, Jane Hoffswell, Shunan Guo, Eunyee Koh, Shamkant B. Navathe, Alex Endert |
CHI | 3 |
| 2023 | Delivery Optimized Discovery in Behavioral User Segmentation under Budget ConstraintabstractUsers' behavioral footprints online enable firms to discover behavior-based user segments (or, segments) and deliver segment specific messages to users. Following the discovery of segments, delivery of messages to users through preferred media channels like Facebook and Google can be challenging, as only a portion of users in a behavior segment find match in a medium, and only a fraction of those matched actually see the message (exposure). Even high quality discovery becomes futile when delivery fails. Many sophisticated algorithms exist for discovering behavioral segments; however, these ignore the delivery component. The problem is compounded because (i) the discovery is performed on the behavior data space in firms' data (e.g., user clicks), while the delivery is predicated on the static data space (e.g., geo, age) as defined by media; and (ii) firms work under budget constraint. We introduce a stochastic optimization based algorithm for delivery optimized discovery of behavioral user segments and offer new metrics to address the joint optimization. We leverage optimization under a budget constraint for delivery combined with a learning-based component for discovery. Extensive experiments on a public dataset from Google and a proprietary dataset show the effectiveness of our approach by simultaneously improving delivery metrics, reducing budget spend and achieving strong predictive performance in discovery. Harshita Chopra, Atanu R. Sinha, Sunav Choudhary, Ryan Rossi, Paavan Kumar Indela, Veda Pranav Parwatala, Srinjayee Paul, Aurghya Maiti |
CIKM | 2 |
| 2023 | The Role of Unattributed Behavior Logs in Predictive User SegmentationabstractOnline browsing on firms' sites generates user behavior logs (or, logs). These logs are mainstays that drive several user modeling tasks. The logs that inform user modeling are the ones that are attributed to each user, termed Attributed Behaviors (AB). But, a lot more logs are anonymous, upwards of 85%. For example, many users do not sign in while browsing. These logs are not attributed to users, termed Unattributed Behaviors (UB), and are not recognized in user modeling. We examine whether and how UB can benefit user modeling. We focus on a common task, that of user segmentation, for which the prior art uses only AB. We demonstrate that information from UBs, although unattributed to any individual, when used along with ABs, enriches performance of machine learning model for user segmentation. We perform predictive segmentation, whereby predicted outcomes for each segment are evaluated against actual outcomes. Multiple evaluations on two datasets, one of which is public, relative to state of the art baseline, show strong performance of our model in predicting outcomes and in reducing user segmentation error. Atanu R. Sinha, Harshita Chopra, Aurghya Maiti, Atishay Ganesh, Sarthak Kapoor, Saili Myana, Saurabh Mahapatra |
CIKM | 1 |
| 2021 | Data-Sharing Economy: Value-Addition from Data meets PrivacyabstractThe need for improved segmentation, targeting, personalization fuel the practice of data sharing among companies. Concurrently, data sharing faces the headwind of new laws emphasizing users' privacy in data. Under the premise that sharing of data occurs from a provider to a recipient, we propose a practicable approach of generating representational data for sharing that achieves value-addition for the recipient's tasks while preserving privacy of users. Prior art shows that the mechanism to improve value-addition inevitably weakens privacy in the generated data. In a first of a kind contribution, our system offers tunable controls to adjust the extent of privacy desired by the provider and the extent of value-addition expected by the recipient. Our experiments on a public data show that under common organizational practice of data-sharing, data generation for value-addition is achievable while preserving privacy. Our demonstration starkly shows the trade-off between privacy-protection and value addition, through user-controlled knobs and offers a prototype of a platform for data sharing which is mindful of this trade-off. Piyush Bagad, Subrata Mitra, Sunny Dhamnani, Atanu R. Sinha, Raunak Gautam, Haresh Khanna |
WSDM | 4 |
| 2019 | Surveys without Questions: A Reinforcement Learning Approach
Atanu R. Sinha, Deepali Jain, Nikhil Sheoran, Sopan Khosla, Reshmi Sasidharan |
AAAI | 1 |
| 2019 | Mentor Pattern Identification from Product Usage Logs
Ankur Garg, Aman Kharb, Yash H. Malviya, J. P. Sagar, Atanu R. Sinha, Iftikhar Ahamath Burhanuddin, Sunav Choudhary |
PAKDD (3) | 5 |
| 2018 | Measurement of Users' Experience on Online Platforms from Their Behavior Logs
Deepali Jain, Atanu R. Sinha, Deepali Gupta, Nikhil Sheoran, Sopan Khosla |
PAKDD (1) | 2 |
| 2017 | Automatic Assignment of Topical Icons to Documents for Faster File NavigationabstractSeveral computer users neither assign names to their documents systematically nor organize them into suitable folders, making it difficult to search for relevant files when needed. While this problem can be addressed in several ways, we explore the novel approach of automated assignment of topical icons to documents in order to cue memory for faster navigation. Specifically, we overlay the currently available generic software-oriented file association icons like Acrobat, Word or Powerpoint on documents with algorithmically assigned icons that are specific to the topical content of the documents. Our pipeline method uses document clustering, significant-phrase extraction, phrase generalization and phrase vector matching for assigning icons to documents. Experimental results show that topical iconification significantly speeds up document navigation time vis-á-vis content-based file naming, in both a controlled laboratory setup as well as in a crowdsourced study. Icons assigned by our algorithm are observed to have satisfactory inter-annotator agreement with respect to their meanings. Rishiraj Saha Roy, Abhijeet Singh, Prashant Chawla, Shubham Saxena, Atanu R. Sinha |
ICDAR | 5 |
| 2016 | Summarizing Multimedia Content
Natwar Modani, Pranav Maneriker, Gaurush Hiranandani, Atanu R. Sinha, Utpal, Vaishnavi Subramanian, Shivani Gupta |
WISE (2) | 4 |