EDBT 2026 Demo / reviewers in the wild / expert
Sanmay Das
dblp:30/2277
· DBLP profile ↗
50ranked-venue papers
13as first author
14since 2021 · last 2026
0000-0002-6814-871XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 9 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 1 since 2021Theory of computation · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimally Auditing Adversarial AgentsabstractFraud can pose a challenge in many resource allocation domains, including social service delivery and credit provision. For example, agents may misreport private information in order to gain benefits or access to credit. To mitigate this, a principal can design strategic audits to verify claims and penalize misreporting. In this paper, we introduce a general model of audit policy design as a principal-agent game with multiple agents, where the principal commits to an audit policy, and agents collectively choose an equilibrium that minimizes the principal’s utility. We examine both adaptive and non-adaptive settings, depending on whether the principal's policy can be responsive to the distribution of agent reports. Our work provides efficient algorithms for computing optimal audit policies in both settings and extends these results to a setting with limited audit budgets. Sanmay Das, Fang-Yi Yu |
AAAI | 1 |
| 2025 | Active Geospatial Search for Efficient Tenant Eviction OutreachabstractTenant evictions threaten housing stability and are a major concern for many cities. An open question concerns whether data-driven methods enhance outreach programs that target at-risk tenants to mitigate their risk of eviction. We propose a novel active geospatial search (AGS) modeling framework for this problem. AGS integrates property-level information in a search policy that identifies a sequence of rental units to canvas to both determine their eviction risk and provide support if needed. We propose a hierarchical reinforcement learning approach to learn a search policy for AGS that scales to large urban areas containing thousands of parcels, balancing exploration and exploitation and accounting for travel costs and a budget constraint. Crucially, the search policy adapts online to newly discovered information about evictions. Evaluation using eviction data for a large urban area demonstrates that the proposed framework and algorithmic approach are considerably more effective at sequentially identifying eviction cases than baseline methods. Anindya Sarkar, Alex DiChristofano, Sanmay Das, Patrick J. Fowler, Nathan Jacobs, Yevgeniy Vorobeychik |
AAAI | 3 |
| 2024 | Discretionary Trees: Understanding Street-Level Bureaucracy via Machine LearningabstractStreet-level bureaucrats interact directly with people on behalf of government agencies to perform a wide range of functions, including, for example, administering social services and policing. A key feature of street-level bureaucracy is that the civil servants, while tasked with implementing agency policy, are also granted significant discretion in how they choose to apply that policy in individual cases. Using that discretion could be beneficial, as it allows for exceptions to policies based on human interactions and evaluations, but it could also allow biases and inequities to seep into important domains of societal resource allocation. In this paper, we use machine learning techniques to understand street-level bureaucrats' behavior. We leverage a rich dataset that combines demographic and other information on households with information on which homelessness interventions they were assigned during a period when assignments were not formulaic. We find that caseworker decisions in this time are highly predictable overall, and some, but not all of this predictivity can be captured by simple decision rules. We theorize that the decisions not captured by the simple decision rules can be considered applications of caseworker discretion. These discretionary decisions are far from random in both the characteristics of such households and in terms of the outcomes of the decisions. Caseworkers typically only apply discretion to households that would be considered less vulnerable. When they do apply discretion to assign households to more intensive interventions, the marginal benefits to those households are significantly higher than would be expected if the households were chosen at random; there is no similar reduction in marginal benefit to households that are discretionarily allocated less intensive interventions, suggesting that caseworkers are using their knowledge and experience to improve outcomes for households experiencing homelessness. Gaurab Pokharel, Sanmay Das, Patrick J. Fowler |
AAAI | 2 |
| 2024 | The Impact of Features Used by Algorithms on Perceptions of Fairness
Andrew Estornell, Tina Zhang, Sanmay Das, Chien-Ju Ho, Brendan Juba, Yevgeniy Vorobeychik |
IJCAI | 3 |
| 2024 | Clinical risk prediction using language models: benefits and considerationsabstractOBJECTIVE: The use of electronic health records (EHRs) for clinical risk prediction is on the rise. However, in many practical settings, the limited availability of task-specific EHR data can restrict the application of standard machine learning pipelines. In this study, we investigate the potential of leveraging language models (LMs) as a means to incorporate supplementary domain knowledge for improving the performance of various EHR-based risk prediction tasks. METHODS: We propose two novel LM-based methods, namely "LLaMA2-EHR" and "Sent-e-Med." Our focus is on utilizing the textual descriptions within structured EHRs to make risk predictions about future diagnoses. We conduct a comprehensive comparison with previous approaches across various data types and sizes. RESULTS: Experiments across 6 different methods and 3 separate risk prediction tasks reveal that employing LMs to represent structured EHRs, such as diagnostic histories, results in significant performance improvements when evaluated using standard metrics such as area under the receiver operating characteristic (ROC) curve and precision-recall (PR) curve. Additionally, they offer benefits such as few-shot learning, the ability to handle previously unseen medical concepts, and adaptability to various medical vocabularies. However, it is noteworthy that outcomes may exhibit sensitivity to a specific prompt. CONCLUSION: LMs encompass extensive embedded knowledge, making them valuable for the analysis of EHRs in the context of risk prediction. Nevertheless, it is important to exercise caution in their application, as ongoing safety concerns related to LMs persist and require continuous consideration. Angeela Acharya, Sulabh Shrestha, Anyi Chen, Joseph Conte, Sanja Avramovic, Siddhartha Sikdar, Antonios Anastasopoulos, Sanmay Das |
J. Am. Medical Informatics Assoc. | 8 |
| 2023 | Popularizing Fairness: Group Fairness and Individual WelfareabstractGroup-fair learning methods typically seek to ensure that some measure of prediction efficacy for (often historically) disadvantaged minority groups is comparable to that for the majority of the population. When a principal seeks to adopt a group-fair approach to replace another, the principal may face opposition from those who feel they may be harmed by the switch, and this, in turn, may deter adoption. We propose that a potential mitigation to this concern is to ensure that a group-fair model is also popular, in the sense that, for a majority of the target population, it yields a preferred distribution over outcomes compared with the conventional model. In this paper, we show that state of the art fair learning approaches are often unpopular in this sense. We propose several efficient algorithms for postprocessing an existing group-fair learning scheme to improve its popularity while retaining fairness. Through extensive experiments, we demonstrate that the proposed postprocessing approaches are highly effective in practice. Andrew Estornell, Sanmay Das, Brendan Juba, Yevgeniy Vorobeychik |
AAAI | 2 |
| 2023 | Incentivizing Recourse through Auditing in Strategic ClassificationabstractThe increasing automation of high-stakes decisions with direct impact on the lives and well-being of individuals raises a number of important considerations. Prominent among these is strategic behavior by individuals hoping to achieve a more desirable outcome. Two forms of such behavior are commonly studied: 1) misreporting of individual attributes, and 2) recourse, or actions that truly change such attributes. The former involves deception, and is inherently undesirable, whereas the latter may well be a desirable goal insofar as it changes true individual qualification. We study misreporting and recourse as strategic choices by individuals within a unified framework. In particular, we propose auditing as a means to incentivize recourse actions over attribute manipulation, and characterize optimal audit policies for two types of principals, utility-maximizing and recourse-maximizing. Additionally, we consider subsidies as an incentive for recourse over manipulation, and show that even a utility-maximizing principal would be willing to devote a considerable amount of audit budget to providing such subsidies. Finally, we consider the problem of optimizing fines for failed audits, and bound the total cost incurred by the population as a result of audits. Andrew Estornell, Sanmay Das, Yang Liu 0018, Yevgeniy Vorobeychik |
IJCAI | 3 |
| 2023 | Fair and Efficient Allocation of Scarce Resources Based on Predicted Outcomes: Implications for Homeless Service DeliveryabstractArtificial intelligence, machine learning, and algorithmic techniques in general, provide two crucial abilities with the potential to improve decision-making in the context of allocation of scarce societal resources. They have the ability to flexibly and accurately model treatment response at the individual level, potentially allowing us to better match available resources to individuals. In addition, they have the ability to reason simultaneously about the effects of matching sets of scarce resources to populations of individuals. In this work, we leverage these abilities to study algorithmic allocation of scarce societal resources in the context of homelessness. In communities throughout the United States, there is constant demand for an array of homeless services intended to address different levels of need. Allocations of housing services must match households to appropriate services that continuously fluctuate in availability, while inefficiencies in allocation could “waste” scarce resources as households will remain in-need and re-enter the homeless system, increasing the overall demand for homeless services. This complex allocation problem introduces novel technical and ethical challenges. Using administrative data from a regional homeless system, we formulate the problem of “optimal” allocation of resources given data on households with need for homeless services. The optimization problem aims to allocate available resources such that predicted probabilities of household re-entry are minimized. The key element of this work is its use of a counterfactual prediction approach that predicts household probabilities of re-entry into homeless services if assigned to each service. Through these counterfactual predictions, we find that this approach has the potential to improve the efficiency of the homeless system by reducing re-entry, and, therefore, system-wide demand. However, efficiency comes with trade-offs - a significant fraction of households are assigned to services that increase probability of re-entry. To address this issue as well as the inherent fairness considerations present in any context where there are insufficient resources to meet demand, we discuss the efficiency, equity, and fairness issues that arise in our work and consider potential implications for homeless policies. Amanda R. Kube, Sanmay Das, Patrick J. Fowler |
J. Artif. Intell. Res. | 2 |
| 2023 | Community- and data-driven homelessness prevention and service delivery: optimizing for equityabstractOBJECTIVE: The study tests a community- and data-driven approach to homelessness prevention. Federal policies call for efficient and equitable local responses to homelessness. However, the overwhelming demand for limited homeless assistance is challenging without empirically supported decision-making tools and raises questions of whom to serve with scarce resources. MATERIALS AND METHODS: System-wide administrative records capture the delivery of an array of homeless services (prevention, shelter, short-term housing, supportive housing) and whether households reenter the system within 2 years. Counterfactual machine learning identifies which service most likely prevents reentry for each household. Based on community input, predictions are aggregated for subpopulations of interest (race/ethnicity, gender, families, youth, and health conditions) to generate transparent prioritization rules for whom to serve first. Simulations of households entering the system during the study period evaluate whether reallocating services based on prioritization rules compared with services-as-usual. RESULTS: Homelessness prevention benefited households who could access it, while differential effects exist for homeless households that partially align with community interests. Households with comorbid health conditions avoid homelessness most when provided longer-term supportive housing, and families with children fare best in short-term rentals. No additional differential effects existed for intersectional subgroups. Prioritization rules reduce community-wide homelessness in simulations. Moreover, prioritization mitigated observed reentry disparities for female and unaccompanied youth without excluding Black and families with children. DISCUSSION: Leveraging administrative records with machine learning supplements local decision-making and enables ongoing evaluation of data- and equity-driven homeless services. CONCLUSIONS: Community- and data-driven prioritization rules more equitably target scarce homeless resources. Amanda R. Kube, Sanmay Das, Patrick J. Fowler |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | Local Justice and the Algorithmic Allocation of Scarce Societal ResourcesabstractAI is increasingly used to aid decision-making about the allocation of scarce societal resources, for example housing for homeless people, organs for transplantation, and food donations. Recently, there have been several proposals for how to design objectives for these systems that attempt to achieve some combination of fairness, efficiency, incentive compatibility, and satisfactory aggregation of stakeholder preferences. This paper lays out possible roles and opportunities for AI in this domain, arguing for a closer engagement with the political philosophy literature on local justice, which provides a framework for thinking about how societies have over time framed objectives for such allocation problems. It also discusses how we may be able to integrate into this framework the opportunities and risks opened up by the ubiquity of data and the availability of algorithms that can use them to make accurate predictions about the future. Sanmay Das |
AAAI | 1 |
| 2022 | GenSyn: A Multi-stage Framework for Generating Synthetic Microdata using Macro Data SourcesabstractIndividual-level data (microdata) that characterizes a population, is essential for studying many real-world problems. However, acquiring such data is not straightforward due to cost and privacy constraints, and access is often limited to aggregated data (macro data) sources. In this study, we examine synthetic data generation as a tool to extrapolate difficult-to-obtain high-resolution data by combining information from multiple easier-to-obtain lower-resolution data sources. In particular, we introduce a framework that uses a combination of univariate and multivariate frequency tables from a given target geographical location in combination with frequency tables from other auxiliary locations to generate synthetic microdata for individuals in the target location. Our method combines the estimation of a dependency graph and conditional probabilities from the target location with the use of a Gaussian copula to leverage the available information from the auxiliary locations. We perform extensive testing on two real-world datasets and demonstrate that our approach outperforms prior approaches in preserving the overall dependency structure of the data while also satisfying the constraints defined on the different variables. Angeela Acharya, Siddhartha Sikdar, Sanmay Das, Huzefa Rangwala |
IEEE Big Data | 3 |
| 2022 | Just Resource Allocation? How Algorithmic Predictions and Human Notions of Justice InteractabstractWe examine justice in data-aided decisions in the context of a scarce societal resource allocation problem. Non-experts (recruited on Amazon Mechanical Turk) have to determine which homeless households to serve with limited housing assistance. We empirically elicit decision-maker preferences for whether to prioritize more vulnerable households or households who would best take advantage of more intensive interventions. We present three main findings. (1) When vulnerability or outcomes are quantitatively conceptualized and presented, humans (at a single point in time) are remarkably consistent in making either vulnerability- or outcome-oriented decisions. (2) Prior exposure to quantitative outcome predictions has a significant effect and changes the preferences of human decision-makers from vulnerability-oriented to outcome-oriented about one-third of the time. (3) Presenting algorithmically-derived risk predictions in addition to household descriptions reinforces decision-maker preferences. Among the vulnerability-oriented, presenting the risk predictions leads to a significant increase in allocations to the more vulnerable household, whereas among the outcome-oriented it leads to a significant decrease in allocations to the more vulnerable household. These findings emphasize the importance of explicitly aligning data-driven decision aids with system-wide allocation goals. Amanda R. Kube, Sanmay Das, Patrick J. Fowler, Yevgeniy Vorobeychik |
EC | 2 |
| 2021 | Incentivizing Truthfulness Through Audits in Strategic ClassificationabstractIn many societal resource allocation domains, machine learning methods are increasingly used to either score or rank agents in order to decide which ones should receive either resources (e.g., homeless services) or scrutiny (e.g., child welfare investigations) from social services agencies. An agency's scoring function typically operates on a feature vector that contains a combination of self-reported features and information available to the agency about individuals or households. This can create incentives for agents to misrepresent their self-reported features in order to receive resources or avoid scrutiny, but agencies may be able to selectively audit agents to verify the veracity of their reports. We study the problem of optimal auditing of agents in such settings. When decisions are made using a threshold on an agent's score, the optimal audit policy has a surprisingly simple structure, uniformly auditing all agents who could benefit from lying. While this policy can, in general be hard to compute because of the difficulty of identifying the set of agents who could benefit from lying given a complete set of reported types, we also present sufficient conditions under which it is tractable. We show that the scarce resource setting is more difficult, and exhibit an approximately optimal audit policy in this case. In addition, we show that in either setting verifying whether it is possible to incentivize exact truthfulness is hard even to approximate. However, we also exhibit sufficient conditions for solving this problem optimally, and for obtaining good approximations. Andrew Estornell, Sanmay Das, Yevgeniy Vorobeychik |
AAAI | 2 |
| 2021 | Scarce Societal Resource Allocation and the Price of (Local) JusticeabstractWe consider the allocation of scarce societal resources, where a central authority decides which individuals receive which resources under capacity or budget constraints. Several algorithmic fairness criteria have been proposed to guide these procedures, each quantifying a notion of local justice to ensure the allocation is aligned with the principles of the local institution making the allocation. For example, the efficient allocation maximizes overall social welfare, whereas the leximin assignment seeks to help the “neediest first.” Although the “price of fairness” (PoF) of leximin has been studied in prior work, we expand on these results by exploiting the structure inherent in real-world scenarios to provide tighter bounds. We further propose a novel criterion – which we term LoINC (leximin over individually normalized costs) – that maximizes a different but commonly used notion of local justice: prioritizing those benefiting the most from receiving the resources. We derive analogous PoF bounds for LoINC, showing that the price of LoINC is typically much lower than that of leximin. We provide extensive experimental results using both synthetic data and in a real-world setting considering the efficacy of different homelessness interventions. These results show that the empirical PoF tends to be substantially lower than worst-case bounds would imply and allow us to characterize situations where the price of LoINC fairness can be high. Sanmay Das, Roman Garnett |
AAAI | 2 |
| 2020 | Deception through Half-TruthsabstractDeception is a fundamental issue across a diverse array of settings, from cybersecurity, where decoys (e.g., honeypots) are an important tool, to politics that can feature politically motivated “leaks” and fake news about candidates. Typical considerations of deception view it as providing false information. However, just as important but less frequently studied is a more tacit form where information is strategically hidden or leaked. We consider the problem of how much an adversary can affect a principal's decision by “half-truths”, that is, by masking or hiding bits of information, when the principal is oblivious to the presence of the adversary. The principal's problem can be modeled as one of predicting future states of variables in a dynamic Bayes network, and we show that, while theoretically the principal's decisions can be made arbitrarily bad, the optimal attack is NP-hard to approximate, even under strong assumptions favoring the attacker. However, we also describe an important special case where the dependency of future states on past states is additive, in which we can efficiently compute an approximately optimal attack. Moreover, in networks with a linear transition function we can solve the problem optimally in polynomial time. Andrew Estornell, Sanmay Das, Yevgeniy Vorobeychik |
AAAI | 2 |
| 2020 | Election Control by Manipulating Issue SignificanceabstractIntegrity of elections is vital to democratic systems, but it is frequently threatened by malicious actors.The study of algorithmic complexity of the problem of manipulating election outcomes by changing its structural features is known as election control Rothe [2016].One means of election control that has been proposed, pertinent to the spatial voting model, is to select a subset of issues that determine voter preferences over candidates.We study a variation of this model in which voters have judgments about relative importance of issues, and a malicious actor can manipulate these judgments.We show that computing effective manipulations in this model is NP-hard even with two candidates or binary issues.However, we demonstrate that the problem becomes tractable with a constant number of voters or issues.Additionally, while it remains intractable when voters can vote stochastically, we exhibit an important special case in which stochastic voting behavior enables tractable manipulation. Andrew Estornell, Sanmay Das, Edith Elkind, Yevgeniy Vorobeychik |
UAI | 2 |
| 2019 | Allocating Interventions Based on Predicted Outcomes: A Case Study on Homelessness ServicesabstractModern statistical and machine learning methods are increasingly capable of modeling individual or personalized treatment effects. These predictions could be used to allocate different interventions across populations based on individual characteristics. In many domains, like social services, the availability of different possible interventions can be severely resource limited. This paper considers possible improvements to the allocation of such services in the context of homelessness service provision in a major metropolitan area. Using data from the homeless system, we use a counterfactual approach to show potential for substantial benefits in terms of reducing the number of families who experience repeat episodes of homelessness by choosing optimal allocations (based on predicted outcomes) to a fixed number of beds in different types of homelessness service facilities. Such changes in the allocation mechanism would not be without tradeoffs, however; a significant fraction of households are predicted to have a higher probability of re-entry in the optimal allocation than in the original one. We discuss the efficiency, equity and fairness issues that arise and consider potential implications for policy. Amanda R. Kube, Sanmay Das, Patrick J. Fowler |
AAAI | 2 |
| 2019 | Revenue Enhancement via Asymmetric Signaling in Interdependent-Value AuctionsabstractWe consider the problem of designing the information environment for revenue maximization in a sealed-bid second price auction with two bidders. Much of the prior literature has focused on signal design in settings where bidders are symmetrically informed, or on the design of optimal mechanisms under fixed information structures. We study commonand interdependent-value settings where the mechanism is fixed (a second-price auction), but the auctioneer controls the signal structure for bidders. We show that in a standard common-value auction setting, there is no benefit to the auctioneer in terms of expected revenue from sharing information with the bidders, although there are effects on the distribution of revenues. In an interdependent-value model with mixed private- and common-value components, however, we show that asymmetric, information-revealing signals can increase revenue. Zhuoshu Li, Sanmay Das |
AAAI | 2 |
| 2018 | The Promise and Perils of Myopia in Dynamic Pricing With Censored InformationabstractA seller with unlimited inventory of a digital good interacts with potential buyers with i.i.d. valuations. The seller can adaptively quote prices to each buyer to maximize long-term profits, but does not know the valuation distribution exactly. Under a linear demand model, we consider two information settings: partially censored, where agents who buy reveal their true valuations after the purchase is completed, and completely censored, where agents never reveal their valuations. In the partially censored case, we prove that myopic pricing with a Pareto prior is Bayes optimal and has finite regret. In both settings, we evaluate the myopic strategy against more sophisticated look-aheads using three valuation distributions generated from real data on auctions of physical goods, keyword auctions, and user ratings, where the linear demand assumption is clearly violated. For some datasets, complete censoring actually helps, because the restricted data acts as a "regularizer" on the posterior, preventing it from being affected too much by outliers. Meenal Chhabra, Sanmay Das, Ilya O. Ryzhov |
IJCAI | 2 |
| 2018 | Equilibrium Behavior in Competing Dynamic Matching MarketsabstractRival markets like rideshare services, universities, and organ exchanges compete to attract participants, seeking to maximize their own utility at potential cost to overall social welfare. Similarly, individual participants in such multi-market systems also seek to maximize their individual utility. If entry is costly, they should strategically enter only a subset of the available markets. All of this decision making---markets competitively adapting their matching strategies and participants arriving, choosing which market(s) to enter, and departing from the system---occurs dynamically over time. This paper provides the first analysis of equilibrium behavior in dynamic competing matching market systems---first from the points of view of individual participants when market policies are fixed, and then from the points of view of markets when agents are stochastic. When compared to single markets running social-welfare-maximizing matching policies, losses in overall social welfare in competitive systems manifest due to both market fragmentation and the use of non-optimal matching policies. We quantify such losses and provide policy recommendations to help alleviate them in fielded systems. Zhuoshu Li, Neal Gupta, Sanmay Das, John Dickerson 0001 |
IJCAI | 3 |
| 2017 | Coordinated Versus Decentralized Exploration In Multi-Agent Multi-Armed BanditsabstractIn this paper, we introduce a multi-agent multi-armed bandit-based model for ad hoc teamwork with expensive communication. The goal of the team is to maximize the total reward gained from pulling arms of a bandit over a number of epochs. In each epoch, each agent decides whether to pull an arm, or to broadcast the reward it obtained in the previous epoch to the team and forgo pulling an arm. These decisions must be made only on the basis of the agent’s private information and the public information broadcast prior to that epoch. We first benchmark the achievable utility by analyzing an idealized version of this problem where a central authority has complete knowledge of rewards acquired from all arms in all epochs and uses a multiplicative weights update algorithm for allocating arms to agents. We then introduce an algorithm for the decentralized setting that uses a value-of-information based communication strategy and an exploration-exploitation strategy based on the centralized algorithm, and show experimentally that it converges rapidly to the performance of the centralized method. Mithun Chakraborty, Kai Yee Phoebe Chua, Sanmay Das, Brendan Juba |
IJCAI | 3 |
| 2016 | Trading on a Rigged Game: Outcome Manipulation in Prediction Markets
Mithun Chakraborty, Sanmay Das |
IJCAI | 2 |
| 2016 | A Symbolic Closed-Form Solution to Sequential Market Making with Inventory
Shamin Kinathil, Scott Sanner, Sanmay Das, Nicolás Della Penna |
IJCAI | 3 |
| 2016 | Manipulation among the Arbiters of Collective Intelligence: How Wikipedia Administrators Mold Public OpinionabstractOur reliance on networked, collectively built information is a vulnerability when the quality or reliability of this information is poor. Wikipedia, one such collectively built information source, is often our first stop for information on all kinds of topics; its quality has stood up to many tests, and it prides itself on having a “neutral point of view.” Enforcement of neutrality is in the hands of comparatively few, powerful administrators. In this article, we document that a surprisingly large number of editors change their behavior and begin focusing more on a particular controversial topic once they are promoted to administrator status. The conscious and unconscious biases of these few, but powerful, administrators may be shaping the information on many of the most sensitive topics on Wikipedia; some may even be explicitly infiltrating the ranks of administrators in order to promote their own points of view. In addition, we ask whether administrators who change their behavior in this suspicious manner can be identified in advance. Neither prior history nor vote counts during an administrator’s election are useful in doing so, but we find that an alternative measure, which gives more weight to influential voters, can successfully reject these suspicious candidates. This second result has important implications for how we harness collective intelligence: even if wisdom exists in a collective opinion (like a vote), that signal can be lost unless we carefully distinguish the true expert voter from the noisy or manipulative voter. Sanmay Das, Allen Lavoie, Malik Magdon-Ismail |
ACM Trans. Web | 1 |
| 2015 | Price Evolution in a Continuous Double Auction Prediction Market With a Scoring-Rule Based Market MakerabstractThe logarithmic market scoring rule (LMSR), the most common automated market making rule for prediction markets, is typically studied in the framework of dealer markets, where the market maker takes one side of every transaction. The continuous double auction (CDA) is a much more widely used microstructure for general financial markets in practice. In this paper, we study the properties of CDA prediction markets with zero-intelligence traders in which an LMSR-style market maker participates actively. We extend an existing idea of Robin Hanson for integrating LMSR with limit order books in order to provide a new, self-contained market making algorithm that does not need “special” access to the order book and can participate as another trader. We find that, as expected, the presence of the market maker leads to generally lower bid-ask spreads and higher trader surplus (or price improvement), but, surprisingly, does not necessarily improve price discovery and market efficiency; this latter effect is more pronounced when there is higher variability in trader beliefs. Mithun Chakraborty, Sanmay Das, Justin Peabody |
AAAI | 2 |
| 2015 | Actions Are Louder than Words in Social MediaabstractWe study the relationship between the level of chatter on a social medium (like Twitter) and the level of the observed actions related to the chatter. For example, in a disaster, how does relief-donation chatter on Twitter correlate with the dollar amount received? One hypothesis is that a fraction of those who act will also tweet about it, which implies linear scaling, action ∝ chatter. On the other hand, if there is a contagion effect (those who tweet about donation incite others to donate) and these incited donors tend to be "quiet" and not broadcast their actions, then we expect superlinear scaling, Rostyslav Korolov, Justin Peabody, Allen Lavoie, Sanmay Das, Malik Magdon-Ismail, William A. Wallace |
ASONAM | 4 |
| 2015 | Prediction of Systemic-to-Pulmonary Artery shunt surgery outcomes using administrative dataabstractSystemic-to-Pulmonary Artery (SPA) shunt surgery, one of the most common cardiac surgical procedures in the newborn period, provides a means to palliate children with limited pulmonary blood flow, such as in Tetralogy of Fallot. Despite the simplicity of the procedure, it is associated with significant morbidity (such as need for extracorporeal membrane oxygenation (ECMO), and long post-operative length of stay (PLOS) in the hospital following surgery) and mortality. These outcomes are known to be impacted by a number of complex factors (including patient specific and procedure specific factors, perioperative related factors, etc.), whose relative importance in clinical decision making remains the domain of clinical judgment. The increasing availability of multi-modal data on patient care and outcomes opens up the opportunity to assess clinical practices from a more data-driven perspective. In this paper, we report results from a study of 1036 patients (from 44 children's hospitals across the US) during 2009-2014 that applies a machine learning approach to predicting post-operative outcomes for patients in the Pediatric Health Information System (PHIS) database. We demonstrate that it is feasible to achieve significant prediction benefits using a standard machine learning approach (random forests) on a carefully constructed dataset, showing the value of applying machine learning even with noisy administrative databases. The methods we describe can be used to identify potential important variables that lead to good clinical judgment as defined by desirable clinical outcomes. Sara Moein, Sanmay Das, Pirooz A. Eghtesady |
BIBM | 3 |
| 2015 | Market Scoring Rules Act As Opinion Pools For Risk-Averse AgentsabstractA market scoring rule (MSR) – a popular tool for designing algorithmic prediction markets – is an incentive-compatible mechanism for the aggregation of probabilistic beliefs from myopic risk-neutral agents. In this paper, we add to a growing body of research aimed at understanding the precise manner in which the price process induced by a MSR incorporates private information from agents who deviate from the assumption of risk-neutrality. We first establish that, for a myopic trading agent with a risk-averse utility function, a MSR satisfying mild regularity conditions elicits the agent’s risk-neutral probability conditional on the latest market state rather than her true subjective probability. Hence, we show that a MSR under these conditions effectively behaves like a more traditional method of belief aggregation, namely an opinion pool, for agents’ true probabilities. In particular, the logarithmic market scoring rule acts as a logarithmic pool for constant absolute risk aversion utility agents, and as a linear pool for an atypical budget-constrained agent utility with decreasing absolute risk aversion. We also point out the interpretation of a market maker under these conditions as a Bayesian learner even when agent beliefs are static. Mithun Chakraborty, Sanmay Das |
NIPS | 2 |
| 2015 | Two-sided search with experts
Yinon Nahum, David Sarne, Sanmay Das, Onn Shehory |
Auton. Agents Multi Agent Syst. | 3 |
| 2014 | Automated inference of point of view from user interactions in collective intelligence venuesabstractEmpirical evaluation of trust and manipulation in large-scale collective intelligence processes is challenging. The datasets involved are too large for thorough manual study, and current automated options are limited. We introduce a statistical framework which classifies point of view based on user interactions. The framework works on Web-scale datasets and is applicable to a wide variety of collective intelligence processes. It enables principled study of such issues as manipulation, trustworthiness of information, and potential bias. We demonstrate the model’s effectiveness in determining point of view on both synthetic data and a dataset of Wikipedia user interactions. We build a combined model of topics and points-of-view on the entire history of English Wikipedia, and show how it can be used to find potentially biased articles and visualize user interactions at a high level. Sanmay Das, Allen Lavoie |
ICML | 1 |
| 2014 | The Role of Common and Private Signals in Two-Sided Matching with Interviews
Sanmay Das, Zhuoshu Li |
WINE | 1 |
| 2013 | On the Social Welfare of Mechanisms for Repeated Batch MatchingabstractWe study hybrid online-batch matching problems, where agents arrive continuously, but are only matched in periodic rounds, when many of them can be considered simultaneously. Agents not getting matched in a given round remain in the market for the next round. This setting models several scenarios of interest, including many job markets as well as kidney exchange mechanisms. We consider the social utility of two commonly used mechanisms for such markets: one that aims for stability in each round (greedy), and one that attempts to maximize social utility in each round (max-weight). Surprisingly, we find that in the long term, the social utility of the greedy mechanism can be higher than that of the max-weight mechanism. We hypothesize that this is because the greedy mechanism behaves similarly to a soft threshold mechanism, where all connections below a certain threshold are rejected by the participants in favor of waiting until the next round. Motivated by this observation, we propose a method to approximately calculate the optimal threshold for an individual agent to use based on characteristics of the other agents participating, and demonstrate experimentally that social utility is high when all agents use this strategy. Thresholding can also be applied by the mechanism itself to improve social welfare; we demonstrate this with an example on graphs that model pairwise kidney exchange. Elliot Anshelevich, Meenal Chhabra, Sanmay Das, Matthew Gerrior |
AAAI | 3 |
| 2013 | Instructor Rating MarketsabstractWe describe the design of Instructor Rating Markets (IRMs) where human participants interact through intelligent automated market-makers in order to provide dynamic collective feedback to instructors on the progress of their classes. The markets are among the first to enable the empirical study of prediction markets where traders can affect the very outcomes they are trading on. More than 200 students across the Rensselaer campus participated in markets for ten classes in the Fall 2010 semester. In this paper, we describe how we designed these markets in order to elicit useful information, and analyze data from the deployment. We show that market prices convey useful information on future instructor ratings and contain significantly more information than do past ratings. The bulk of useful information contained in the price of a particular class is provided by students who are in that class, showing that the markets are serving to disseminate insider information. At the same time, we find little evidence of attempted manipulation by raters. The markets are also a laboratory for comparing different market designs and the resulting price dynamics, and we show how they can be used to compare market making algorithms. Mithun Chakraborty, Sanmay Das, Allen Lavoie, Malik Magdon-Ismail, Yonatan Naamad |
AAAI | 2 |
| 2013 | Manipulation among the arbiters of collective intelligence: how wikipedia administrators mold public opinionabstractOur reliance on networked, collectively built information is a vulnerability when the quality or reliability of this information is poor. Wikipedia, one such collectively built information source, is often our first stop for information on all kinds of topics; its quality has stood up to many tests, and it prides itself on having a "Neutral Point of View". Enforcement of neutrality is in the hands of comparatively few, powerful administrators. We find a surprisingly large number of editors who change their behavior and begin focusing more on a particular controversial topic once they are promoted to administrator status. The conscious and unconscious biases of these few, but powerful, administrators may be shaping the information on many of the most sensitive topics on Wikipedia; some may even be explicitly infiltrating the ranks of administrators in order to promote their own points of view. Neither prior history nor vote counts during an administrator's election can identify those editors most likely to change their behavior in this suspicious manner. We find that an alternative measure, which gives more weight to influential voters, can successfully reject these suspicious candidates. This has important implications for how we harness collective intelligence: even if wisdom exists in a collective opinion (like a vote), that signal can be lost unless we carefully distinguish the true expert voter from the noisy or manipulative voter. Sanmay Das, Allen Lavoie, Malik Magdon-Ismail |
CIKM | 1 |
| 2013 | Anarchy, stability, and utopia: creating better matchings
Elliot Anshelevich, Sanmay Das, Yonatan Naamad |
Auton. Agents Multi Agent Syst. | 2 |
| 2012 | A bayesian market makerabstractEnsuring sufficient liquidity is one of the key challenges for designers of prediction markets. Variants of the logarithmic market scoring rule (LMSR) have emerged as the standard. LMSR market makers are loss-making in general and need to be subsidized. Proposed variants, including liquidity sensitive market makers, suffer from an inability to react rapidly to jumps in population beliefs. In this paper we propose a Bayesian Market Maker for binary outcome (or continuous 0-1) markets that learns from the informational content of trades. By sacrificing the guarantee of bounded loss, the Bayesian Market Maker can simultaneously offer: (1) significantly lower expected loss at the same level of liquidity, and, (2) rapid convergence when there is a jump in the underlying true value of the security. We present extensive evaluations of the algorithm in experiments with intelligent trading agents and in human subject experiments. Our investigation also elucidates some general properties of market makers in prediction markets. In particular, there is an inherent tradeoff between adaptability to market shocks and convergence during market equilibrium. Aseem Brahma, Mithun Chakraborty, Sanmay Das, Allen Lavoie, Malik Magdon-Ismail |
EC | 3 |
| 2012 | Two-sided search with expertsabstractIn this paper we study distributed agent matching in environments characterized by uncertain signals, costly exploration, and the presence of an information broker. Each agent receives information about the potential value of matching with others. This information signal may, however be noisy, and the agent incurs some cost in receiving it. If all candidate agents agree to the matching the team is formed and each agent receives the true unknown utility of the matching, and leaves the market. We consider the effect of the presence of information brokers, or experts, on the outcomes of such matching processes. Experts can, upon payment of a fee, perform the service of disambiguating noisy signals and revealing the true value of a match to any agent. We analyze equilibrium behavior given the fee set by a monopolist expert and use this analysis to derive the revenue maximizing strategy for the expert as the first mover in a Stackelberg game. Surprisingly, we find that better information can hurt: the presence of the expert, even if the use of its services is optional, can degrade both individual agents' utilities and overall social welfare. While in one-sided search the presence of the expert can only help, in two-sided (and general k-sided) search the externality imposed by the fact that others are consulting the expert can lead to a situation where the equilibrium outcome is that everyone consults the expert, even though all agents would be better off if the expert were not present. As an antidote, we show how market designers can enhance welfare by taxing use of expert services. Yinon Nahum, David Sarne, Sanmay Das, Onn Shehory |
EC | 3 |
| 2012 | Market mechanisms for resource allocation in pervasive sensor applications
Sahin Cem Geyik, Syed Yousaf Shah, Boleslaw K. Szymanski, Sanmay Das, Petros Zerfos |
Pervasive Mob. Comput. | 4 |
| 2012 | A Model for Information Growth in Collective Wisdom ProcessesabstractCollaborative media such as wikis have become enormously successful venues for information creation. Articles accrue information through the asynchronous editing of users who arrive both seeking information and possibly able to contribute information. Most articles stabilize to high-quality, trusted sources of information representing the collective wisdom of all the users who edited the article. We propose a model for information growth which relies on two main observations: (i) as an article’s quality improves, it attracts visitors at a faster rate (a rich-get-richer phenomenon); and, simultaneously, (ii) the chances that a new visitor will improve the article drops (there is only so much that can be said about a particular topic). Our model is able to reproduce many features of the edit dynamics observed on Wikipedia; in particular, it captures the observed rise in the edit rate, followed by 1/ t decay. Despite differences in the media, we also document similar features in the comment rates for a segment of the LiveJournal blogosphere. Sanmay Das, Malik Magdon-Ismail |
ACM Trans. Knowl. Discov. Data | 1 |
| 2011 | Near-Optimal Target Learning With Stochastic Binary Signals
Mithun Chakraborty, Sanmay Das, Malik Magdon-Ismail |
UAI | 2 |
| 2011 | Identifying Relevant Data for a Biological Database: Handcrafted Rules versus Machine LearningabstractWith well over 1,000 specialized biological databases in use today, the task of automatically identifying novel, relevant data for such databases is increasingly important. In this paper, we describe practical machine learning approaches for identifying MEDLINE documents and Swiss-Prot/TrEMBL protein records, for incorporation into a specialized biological database of transport proteins named TCDB. We show that both learning approaches outperform rules created by hand by a human expert. As one of the first case studies involving two different approaches to updating a deployed database, both the methods compared and the results will be of interest to curators of many specialized databases. Aditya Kumar Sehgal, Sanmay Das, Keith Noto, Milton H. Saier Jr., Charles Elkan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | Predictive State Representations for grounding human-robot communicationabstractAllowing robots to communicate naturally with humans is an important goal for social robotics. Most approaches have focused on building high-level probabilistic cognitive models. However, research in cognitive science shows that people often build common ground for communication with each other by seeking and providing evidence of understanding through behaviors like mimicry. Predictive State Representations (PSRs) allow one to build explicit, low-level models of the expected outcomes of actions, and are therefore well-suited for tasks that require providing such evidence of understanding. Using human-robot shadow puppetry as a prototype interaction study, we show that PSRs can be used successfully to both model human interactions, and to allow a robot to learn on-line how to engage a human in an interesting interaction. Eric M. Meisner, Sanmay Das, Volkan Isler, Jeffrey C. Trinkle, Selma Sabanovic, Linnda R. Caporael |
ICRA | 2 |
| 2010 | Collective wisdom: information growth in wikis and blogsabstractWikis and blogs have become enormously successful media for collaborative information creation. Articles and posts accrue information through the asynchronous editing of users who arrive both seeking information and possibly able to contribute information. Most articles stabilize to high quality, trusted sources of information representing the collective wisdom of all the users who edited the article. We propose a model for information growth which relies on two main observations: (i) as an article's quality improves, it attracts visitors at a faster rate (a rich get richer phenomenon); and, simultaneously, (ii) the chances that a new visitor will improve the article drops (there is only so much that can be said about a particular topic). Our model is able to reproduce many features of the edit dynamics observed on Wikipedia and on blogs collected from LiveJournal; in particular, it captures the observed rise in the edit rate, followed by 1/t decay. Sanmay Das, Malik Magdon-Ismail |
EC | 1 |
| 2009 | Anarchy, Stability, and Utopia: Creating Better Matchings
Elliot Anshelevich, Sanmay Das, Yonatan Naamad |
SAGT | 2 |
| 2008 | Adapting to a Market Shock: Optimal Sequential Market-MakingabstractWe study the profit-maximization problem of a monopolistic market-maker who sets two-sided prices in an asset market. The sequential decision problem is hard to solve because the state space is a function. We demonstrate that the belief state is well approximated by a Gaussian distribution. We prove a key monotonicity property of the Gaussian state update which makes the problem tractable, yielding the first optimal sequential market-making algorithm in an established model. The algorithm leads to a surprising insight: an optimal monopolist can provide more liquidity than perfectly competitive market-makers in periods of extreme uncertainty, because a monopolist is willing to absorb initial losses in order to learn a new valuation rapidly so she can extract higher profits later. Sanmay Das, Malik Magdon-Ismail |
NIPS | 1 |
| 2007 | Learning to trade with insider informationabstractThis paper introduces algorithms for learning how to trade using insider (superior) information in Kyle's model of financial markets. Prior results in finance theory relied on the insider having perfect knowledge of the structure and parameters of the market. I show here that it is possible to learn the equilibrium trading strategy when its form is known even without knowledge of the parameters governing trading in the model. However, the rate of convergence to equilibrium is slow, and an approximate algorithm that does not converge to the equilibrium strategy achieves better utility when the horizon is limited. I analyze this approximate algorithm from the perspective of reinforcement learning and discuss the importance of domain knowledge in designing a successful learning algorithm. Sanmay Das |
ICEC | 1 |
| 2007 | Finding Transport Proteins in a General Protein Database
Sanmay Das, Milton H. Saier Jr., Charles Elkan |
PKDD | 1 |
| 2005 | Two-Sided Bandits and the Dating Market
Sanmay Das, Emir Kamenica |
IJCAI | 1 |
| 2002 | The influence of social norms and social consciousness on intention reconciliation
Barbara J. Grosz, Sarit Kraus, David G. Sullivan, Sanmay Das |
Artif. Intell. | 4 |
| 2001 | Filters, Wrappers and a Boosting-Based Hybrid for Feature Selection
Sanmay Das |
ICML | 1 |