Aritra Ghosh 0001

dblp:143/7228-1 · DBLP profile ↗
← Back
16ranked-venue papers
13as first author
8since 2021 · last 2023
0000-0003-2024-2173ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 9 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2023 DiFA: Differentiable Feature Acquisition
abstract
Feature acquisition in predictive modeling is an important task in many practical applications. For example, in patient health prediction, we do not fully observe their personal features and need to dynamically select features to acquire. Our goal is to acquire a small subset of features that maximize prediction performance. Recently, some works reformulated feature acquisition as a Markov decision process and applied reinforcement learning (RL) algorithms, where the reward reflects both prediction performance and feature acquisition cost. However, RL algorithms only use zeroth-order information on the reward, which leads to slow empirical convergence, especially when there are many actions (number of features) to consider. For predictive modeling, it is possible to use first-order information on the reward, i.e., gradients, since we are often given an already collected dataset. Therefore, we propose differentiable feature acquisition (DiFA), which uses a differentiable representation of the feature selection policy to enable gradients to flow from the prediction loss to the policy parameters. We conduct extensive experiments on various real-world datasets and show that DiFA significantly outperforms existing feature acquisition methods when the number of features is large.
Aritra Ghosh 0001, Andrew S. Lan
AAAI1
2023 Balancing Test Accuracy and Security in Computerized Adaptive Testing
Wanyong Feng, Aritra Ghosh 0001, Stephen Sireci, Andrew S. Lan
AIED2
2023 A Conceptual Model for End-to-End Causal Discovery in Knowledge Tracing
Nischal Ashok Kumar, Wanyong Feng, Jaewook Lee 0006, Hunter McNichols, Aritra Ghosh 0001, Andrew S. Lan
EDM5
2022 DiPS: Differentiable Policy for Sketching in Recommender Systems
abstract
In sequential recommender system applications, it is important to develop models that can capture users' evolving interest over time to successfully recommend future items that they are likely to interact with. For users with long histories, typical models based on recurrent neural networks tend to forget important items in the distant past. Recent works have shown that storing a small sketch of past items can improve sequential recommendation tasks. However, these works all rely on static sketching policies, i.e., heuristics to select items to keep in the sketch, which are not necessarily optimal and cannot improve over time with more training data. In this paper, we propose a differentiable policy for sketching (DiPS), a framework that learns a data-driven sketching policy in an end-to-end manner together with the recommender system model to explicitly maximize recommendation quality in the future. We also propose an approximate estimator of the gradient for optimizing the sketching algorithm parameters that is computationally efficient. We verify the effectiveness of DiPS on real-world datasets under various practical settings and show that it requires up to 50% fewer sketch items to reach the same predictive quality than existing sketching policies.
Aritra Ghosh 0001, Saayan Mitra, Andrew S. Lan
AAAI1
2022 Automated Scoring for Reading Comprehension via In-context BERT Tuning
Nigel Fernandez, Aritra Ghosh 0001, Naiming Liu, Zichao Wang 0001, Benoît Choffin, Richard G. Baraniuk, Andrew S. Lan
AIED (1)2
2021 Option Tracing: Beyond Correctness Analysis in Knowledge Tracing
Aritra Ghosh 0001, Jay Raspat, Andrew S. Lan
AIED (1)1
2021 BOBCAT: Bilevel Optimization-Based Computerized Adaptive Testing
abstract
Computerized adaptive testing (CAT) refers to a form of tests that are personalized to every student/test taker. CAT methods adaptively select the next most informative question/item for each student given their responses to previous questions, effectively reducing test length. Existing CAT methods use item response theory (IRT) models to relate student ability to their responses to questions and static question selection algorithms designed to reduce the ability estimation error as quickly as possible; therefore, these algorithms cannot improve by learning from large-scale student response data. In this paper, we propose BOBCAT, a Bilevel Optimization-Based framework for CAT to directly learn a data-driven question selection algorithm from training data. BOBCAT is agnostic to the underlying student response model and is computationally efficient during the adaptive testing process. Through extensive experiments on five real-world student response datasets, we show that BOBCAT outperforms existing CAT methods (sometimes significantly) at reducing test length.
Aritra Ghosh 0001, Andrew S. Lan
IJCAI1
2021 Do We Really Need Gold Samples for Sample Weighting under Label Noise?
abstract
Learning with labels noise has gained significant traction recently due to the sensitivity of deep neural networks under label noise under common loss functions. Losses that are theoretically robust to label noise, however, often makes training difficult. Consequently, several recently proposed methods, such as Meta-Weight-Net (MW-Net), use a small number of unbiased, clean samples to learn a weighting function that downweights samples that are likely to have corrupted labels under the meta-learning framework. However, obtaining such a set of clean samples is not always feasible in practice. In this paper, we analytically show that one can easily train MW-Net without access to clean samples simply by using a loss function that is robust to label noise, such as mean absolute error, as the meta objective to train the weighting network. We experimentally show that our method beats all existing methods that do not use clean samples and performs on-par with methods that use gold samples on benchmark datasets across various noise types and noise rates.
Aritra Ghosh 0001, Andrew S. Lan
WACV1
2020 Skill-based Career Path Modeling and Recommendation
abstract
The development of new technologies at an unprecedented rate is rapidly changing the landscape of the labor market. Therefore, for workers who want to build a successful career, acquiring new skills required by new jobs through lifelong learning is crucial. In this paper, we propose a novel and interpretable monotonic nonlinear state-space model to analyze online user professional profiles and provide actionable feedback and recommendations to users on how they can reach their career goals. Specifically, we use a series of binary-valued and non-decreasing latent states to represent the expanding skill set of each user throughout their career and propose an efficient inference method under our model. Using a series of experiments on two large real-world datasets, we show that our model (sometimes significantly) outperforms existing methods on the tasks of company, job title, and skill prediction. More importantly, our model is interpretable and can be used for other important tasks including skill gap identification and career path planning. Using a series of case studies, we show that our model can provide i) actionable feedback to users and guide them through their upskilling and reskilling processes and ii) recommendations of feasible paths for users to reach their career goals.
Aritra Ghosh 0001, Beverly P. Woolf, Shlomo Zilberstein, Andrew S. Lan
IEEE BigData1
2020 Context-Aware Attentive Knowledge Tracing
abstract
Knowledge tracing (KT) refers to the problem of predicting future learner performance given their past performance in educational applications. Recent developments in KT using flexible deep neural network-based models excel at this task. However, these models often offer limited interpretability, thus making them insufficient for personalized learning, which requires using interpretable feedback and actionable recommendations to help learners achieve better learning outcomes. In this paper, we propose attentive knowledge tracing (AKT), which couples flexible attention-based neural network models with a series of novel, interpretable model components inspired by cognitive and psychometric models. AKT uses a novel monotonic attention mechanism that relates a learner's future responses to assessment questions to their past responses; attention weights are computed using exponential decay and a context-aware relative distance measure, in addition to the similarity between questions. Moreover, we use the Rasch model to regularize the concept and question embeddings; these embeddings are able to capture individual differences among questions on the same concept without using an excessive number of parameters. We conduct experiments on several real-world benchmark datasets and show that AKT outperforms existing KT methods (by up to $6%$ in AUC in some cases) on predicting future learner responses. We also conduct several case studies and show that AKT exhibits excellent interpretability and thus has potential for automated feedback and personalization in real-world educational settings.
Aritra Ghosh 0001, Neil T. Heffernan, Andrew S. Lan
KDD1
2020 Optimal Bidding Strategy without Exploration in Real-time Bidding
abstract
Maximizing utility with a budget constraint is the primary goal for advertisers in real-time bidding (RTB) systems. The policy maximizing the utility is referred to as the optimal bidding strategy. Earlier works on optimal bidding strategy apply model-based batch reinforcement learning methods which can not generalize to unknown budget and time constraint. Further, the advertiser observes a censored market price which makes direct evaluation infeasible on batch test datasets. Previous works ignore the losing auctions to alleviate the difficulty with censored states; thus significantly modifying the test distribution. We address the challenge of lacking a clear evaluation procedure as well as the error propagated through batch reinforcement learning methods in RTB systems. We exploit two conditional independence structures in the sequential bidding process that allow us to propose a novel practical framework using the maximum entropy principle to imitate the behavior of the true distribution observed in real-time traffic. Moreover, the framework allows us to train a model that can generalize to the unseen budget conditions than limit only to those observed in history. We compare our methods on two real-world RTB datasets with several baselines and demonstrate significantly improved performance under various budget settings.
Aritra Ghosh 0001, Saayan Mitra, Somdeb Sarkhel, Viswanathan (Vishy) Swaminathan
SDM1
2019 Scalable Bid Landscape Forecasting in Real-Time Bidding
abstract
In programmatic advertising, ad slots are usually sold using second-price (SP) auctions in real-time. The highest bidding advertiser wins but pays only the second-highest bid (known as the winning price). In SP, for a single item, the dominant strategy of each bidder is to bid the true value from the bidder's perspective. However, in a practical setting, with budget constraints, bidding the true value is a sub-optimal strategy. Hence, to devise an optimal bidding strategy, it is of utmost importance to learn the winning price distribution accurately. Moreover, a demand-side platform (DSP), which bids on behalf of advertisers, observes the winning price if it wins the auction. For losing auctions, DSPs can only treat its bidding price as the lower bound for the unknown winning price. In literature, typically censored regression is used to model such partially observed data. A common assumption in censored regression is that the winning price is drawn from a fixed variance (homoscedastic) uni-modal distribution (most often Gaussian). However, in reality, these assumptions are often violated. We relax these assumptions and propose a heteroscedastic fully parametric censored regression approach, as well as a mixture density censored network. Our approach not only generalizes censored regression but also provides flexibility to model arbitrarily distributed real-world data. Experimental evaluation on the publicly available dataset for winning price estimation demonstrates the effectiveness of our method. Furthermore, we evaluate our algorithm on one of the largest demand-side platforms and significant improvement has been achieved in comparison with the baseline solutions.
Aritra Ghosh 0001, Saayan Mitra, Somdeb Sarkhel, Jason Xie, Gang Wu 0013, Viswanathan (Vishy) Swaminathan
ECML/PKDD (3)1
2017 Robust Loss Functions under Label Noise for Deep Neural Networks
abstract
In many applications of classifier learning, training data suffers from label noise. Deep networks are learned using huge training data where the problem of noisy labels is particularly relevant. The current techniques proposed for learning deep networks under label noise focus on modifying the network architecture and on algorithms for estimating true labels from noisy labels. An alternate approach would be to look for loss functions that are inherently noise-tolerant. For binary classification there exist theoretical results on loss functions that are robust to label noise. In this paper, we provide some sufficient conditions on a loss function so that risk minimization under that loss function would be inherently tolerant to label noise for multiclass classification problems. These results generalize the existing results on noise-tolerant loss functions for binary classification. We study some of the widely used loss functions in deep networks and show that the loss function based on mean absolute value of error is inherently robust to label noise. Thus standard back propagation is enough to learn the true classifier even under label noise. Through experiments, we illustrate the robustness of risk minimization with such loss functions for learning neural networks.
Aritra Ghosh 0001, P. S. Sastry 0001
AAAI1
2017 On the Robustness of Decision Tree Learning Under Label Noise
Aritra Ghosh 0001, Naresh Manwani, P. S. Sastry 0001
PAKDD (1)1
2016 A Preference Approach to Reputation in Sponsored Search
abstract
Determining reputation of an advertiser in sponsored search is a recent important problem with direct impact on revenue for web publishers and relevance of ads. Individual performance of advertisers is usually expressed through observed click through rate, which depends on advertiser reputation, ad relevance and position. However, advertiser reputation has not been explicitly modeled in click prediction literature. Using traditional approaches in web page popularity for organic search in this context is not reasonable as the notion of link-structure in web is not directly applicable to sponsored search. In this study, we motivate and propose a pairwise preference relation model to study the advertiser reputation problem. Pairwise comparisons of advertisers give information over and above the information available in their individual historical performances. We relate the notion of preference among the advertisers to the spectral properties of the preference graph. We provide empirical evidence of the existence of reputation bias in click behavior. Consequently, we experiment with this signal to improve click prediction.
Aritra Ghosh 0001, Dinesh Gaurav, Rahul Agrawal
CIKM1
2015 Making risk minimization tolerant to label noise
Aritra Ghosh 0001, Naresh Manwani, P. S. Sastry 0001
Neurocomputing1