Aman Agarwal

dblp:25/8468 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 4 first-authorArtificial intelligence and machine learning · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Information retrieval · 74% Recommender systems · 26%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › ranking
learning to rank
1.542019
Addressing Trust Bias for Unbiased Learning-to-Rank · WWW 2019
Estimating Position Bias without Intrusive Interventions · WSDM 2019
Intervention Harvesting for Context-Dependent Examination-Bias Estimation · SIGIR 2019
Information retrieval › ranking › learning to rank
unbiased learning to rank
1.132019
Addressing Trust Bias for Unbiased Learning-to-Rank · WWW 2019
Intervention Harvesting for Context-Dependent Examination-Bias Estimation · SIGIR 2019
A General Framework for Counterfactual Learning-to-Rank · SIGIR 2019
Information retrieval › user behavior › search behavior
click model
0.822019
Addressing Trust Bias for Unbiased Learning-to-Rank · WWW 2019
Intervention Harvesting for Context-Dependent Examination-Bias Estimation · SIGIR 2019
Information retrieval › ranking › learning to rank › unbiased learning to rank
counterfactual learning to rank
0.822019
Estimating Position Bias without Intrusive Interventions · WSDM 2019
A General Framework for Counterfactual Learning-to-Rank · SIGIR 2019
Recommender systems › debiased recommendation
propensity estimation
0.722019
Estimating Position Bias without Intrusive Interventions · WSDM 2019
Effective Evaluation Using Logged Bandit Feedback from Multiple Loggers · KDD 2017
Bioinformatics and computational biology
functional genomics
0.712023
TIVAN-indel: a computational framework for annotating and predicting non-coding regulatory small insertions and deletions · Bioinform. 2023
Bioinformatics and computational biology › gene regulation
gene regulation analysis
0.712023
DeepPHiC: predicting promoter-centered chromatin interactions using a novel deep learning approach · Bioinform. 2023
Bioinformatics and computational biology
genomics
0.712023
TIVAN-indel: a computational framework for annotating and predicting non-coding regulatory small insertions and deletions · Bioinform. 2023
Bioinformatics and computational biology › gene regulation
regulatory genomics
0.712023
DeepPHiC: predicting promoter-centered chromatin interactions using a novel deep learning approach · Bioinform. 2023
Bioinformatics and computational biology › gene regulation › regulatory variant interpretation
regulatory variant effect prediction
0.712023
TIVAN-indel: a computational framework for annotating and predicting non-coding regulatory small insertions and deletions · Bioinform. 2023
Bioinformatics and computational biology › genomics
variant annotation
0.712023
TIVAN-indel: a computational framework for annotating and predicting non-coding regulatory small insertions and deletions · Bioinform. 2023
Recommender systems › recommender system evaluation › off-policy evaluation
inverse propensity scoring
0.522019
Intervention Harvesting for Context-Dependent Examination-Bias Estimation · SIGIR 2019
A General Framework for Counterfactual Learning-to-Rank · SIGIR 2019
Recommender systems › collaborative filtering
implicit feedback
0.412019
Intervention Harvesting for Context-Dependent Examination-Bias Estimation · SIGIR 2019
Information retrieval › user behavior › search behavior › click model
position-based model
0.412019
Addressing Trust Bias for Unbiased Learning-to-Rank · WWW 2019
Information retrieval › user behavior › search behavior › click model
position bias
0.412019
Estimating Position Bias without Intrusive Interventions · WSDM 2019
Information retrieval
search engines
0.412019
Estimating Position Bias without Intrusive Interventions · WSDM 2019
Recommender systems › recommender system evaluation
off-policy evaluation
0.312017
Effective Evaluation Using Logged Bandit Feedback from Multiple Loggers · KDD 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.112019
Addressing Trust Bias for Unbiased Learning-to-Rank · WWW 2019

Methods — techniques the papers use, named apart from their topics

inverse propensity scoring · 0.8expectation-maximization · 0.8bayes rule · 0.8transfer learning · 0.7supervised machine learning · 0.7multimodal deep learning · 0.7multi-task learning · 0.7epigenomic profiling · 0.7enrichment analysis · 0.7inverse propensity score weighting · 0.4intervention harvesting · 0.4intervention data harvesting · 0.4gradient-based optimization · 0.4extremum estimator · 0.4convex-concave procedure · 0.4SVM · 0.4
YearPublicationVenuePosition
2024 CBAR-UNet: A novel methodology for segmentation of cardiac magnetic resonance images using block attention-based deep residual neural network
Rakesh Kumar 0009, Meenu Gupta, Aman Agarwal, Anand Nayyar
Multim. Tools Appl.3
2023 DeepPHiC: predicting promoter-centered chromatin interactions using a novel deep learning approach
abstract
MOTIVATION: Promoter-centered chromatin interactions, which include promoter-enhancer (PE) and promoter-promoter (PP) interactions, are important to decipher gene regulation and disease mechanisms. The development of next-generation sequencing technologies such as promoter capture Hi-C (pcHi-C) leads to the discovery of promoter-centered chromatin interactions. However, pcHi-C experiments are expensive and thus may be unavailable for tissues/cell types of interest. In addition, these experiments may be underpowered due to insufficient sequencing depth or various artifacts, which results in a limited finding of interactions. Most existing computational methods for predicting chromatin interactions are based on in situ Hi-C and can detect chromatin interactions across the entire genome. However, they may not be optimal for predicting promoter-centered chromatin interactions. RESULTS: We develop a supervised multi-modal deep learning model, which utilizes a comprehensive set of features such as genomic sequence, epigenetic signal, anchor distance, evolutionary features and DNA structural features to predict tissue/cell type-specific PE and PP interactions. We further extend the deep learning model in a multi-task learning and a transfer learning framework and demonstrate that the proposed approach outperforms state-of-the-art deep learning methods. Moreover, the proposed approach can achieve comparable prediction performance using predefined biologically relevant tissues/cell types compared to using all tissues/cell types in the pretraining especially for predicting PE interactions. The prediction performance can be further improved by using computationally inferred biologically relevant tissues/cell types in the pretraining, which are defined based on the common genes in the proximity of two anchors in the chromatin interactions. AVAILABILITY AND IMPLEMENTATION: https://github.com/lichen-lab/DeepPHiC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Aman Agarwal, Li Chen 0029
Bioinform.1
2023 TIVAN-indel: a computational framework for annotating and predicting non-coding regulatory small insertions and deletions
abstract
MOTIVATION: Small insertion and deletion (sindel) of human genome has an important implication for human disease. One important mechanism for non-coding sindel (nc-sindel) to have an impact on human diseases and phenotypes is through the regulation of gene expression. Nevertheless, current sequencing experiments may lack statistical power and resolution to pinpoint the functional sindel due to lower minor allele frequency or small effect size. As an alternative strategy, a supervised machine learning method can identify the otherwise masked functional sindels by predicting their regulatory potential directly. However, computational methods for annotating and predicting the regulatory sindels, especially in the non-coding regions, are underdeveloped. RESULTS: By leveraging labeled nc-sindels identified by cis-expression quantitative trait loci analyses across 44 tissues in Genotype-Tissue Expression (GTEx), and a compilation of both generic functional annotations and large-scale epigenomic profiles, we develop TIssue-specific Variant Annotation for Non-coding indel (TIVAN-indel), which is a supervised computational framework for predicting non-coding regulatory sindels. As a result, we demonstrate that TIVAN-indel achieves the best prediction performance in both with-tissue prediction and cross-tissue prediction. As an independent evaluation, we train TIVAN-indel from the 'Whole Blood' tissue in GTEx and test the model using 15 immune cell types from an independent study named Database of Immune Cell Expression. Lastly, we perform an enrichment analysis for both true and predicted sindels in key regulatory regions such as chromatin interactions, open chromatin regions and histone modification sites, and find biologically meaningful enrichment patterns. AVAILABILITY AND IMPLEMENTATION: https://github.com/lichen-lab/TIVAN-indel. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Aman Agarwal, Fengdi Zhao, Li Chen 0029
Bioinform.1
2019 A General Framework for Counterfactual Learning-to-Rank
abstract
Implicit feedback (e.g., click, dwell time) is an attractive source of training data for Learning-to-Rank, but its naive use leads to learning results that are distorted by presentation bias. For the special case of optimizing average rank for linear ranking functions, however, the recently developed SVM-PropRank method has shown that counterfactual inference techniques can be used to provably overcome the distorting effect of presentation bias. Going beyond this special case, this paper provides a general and theoretically rigorous framework for counterfactual learning-to-rank that enables unbiased training for a broad class of additive ranking metrics (e.g., Discounted Cumulative Gain (DCG)) as well as a broad class of models (e.g., deep networks). Specifically, we derive a relaxation for propensity-weighted rank-based metrics which is subdifferentiable and thus suitable for gradient-based optimization. We demonstrate the effectiveness of this general approach by instantiating two new learning methods. One is a new type of unbiased SVM that optimizes DCG - called SVM PropDCG - and we show how the resulting optimization problem can be solved via the Convex Concave Procedure (CCP). The other is Deep PropDCG, where the ranking function can be an arbitrary deep network. In addition to the theoretical support, we empirically find that SVM PropDCG significantly outperforms existing linear rankers in terms of DCG. Moreover, the ability to train non-linear ranking functions via Deep PropDCG further improves performance.
Aman Agarwal, Kenta Takatsu, Ivan Zaitsev, Thorsten Joachims
SIGIR1
2019 Intervention Harvesting for Context-Dependent Examination-Bias Estimation
abstract
Accurate estimates of examination bias are crucial for unbiased learning-to-rank from implicit feedback in search engines and recommender systems, since they enable the use of Inverse Propensity Score (IPS) weighting techniques to address selection biases and missing data. Unfortunately, existing examination-bias estimators are limited to the Position-Based Model (PBM), where the examination bias may only depend on the rank of the document. To overcome this limitation, we propose a Contextual Position-Based Model (CPBM) where the examination bias may also depend on a context vector describing the query and the user. Furthermore, we propose an effective estimator for the CPBM based on intervention harvesting. A key feature of the estimator is that it does not require disruptive interventions but merely exploits natural variation resulting from the use of multiple historic ranking functions. Real-world experiments on the ArXiv search engine and semi-synthetic experiments on the Yahoo Learning-To-Rank dataset demonstrate the superior effectiveness and robustness of the new approach.
Zhichong Fang, Aman Agarwal, Thorsten Joachims
SIGIR2
2019 Estimating Position Bias without Intrusive Interventions
abstract
Presentation bias is one of the key challenges when learning from implicit feedback in search engines, as it confounds the relevance signal. While it was recently shown how counterfactual learning-to-rank (LTR) approaches \citeJoachims/etal/17a can provably overcome presentation bias when observation propensities are known, it remains to show how to effectively estimate these propensities. In this paper, we propose the first method for producing consistent propensity estimates without manual relevance judgments, disruptive interventions, or restrictive relevance modeling assumptions. First, we show how to harvest a specific type of intervention data from historic feedback logs of multiple different ranking functions, and show that this data is sufficient for consistent propensity estimation in the position-based model. Second, we propose a new extremum estimator that makes effective use of this data. In an empirical evaluation, we find that the new estimator provides superior propensity estimates in two real-world systems -- Arxiv Full-text Search and Google Drive Search. Beyond these two points, we find that the method is robust to a wide range of settings in simulation studies.
Aman Agarwal, Ivan Zaitsev, Xuanhui Wang, Cheng Li 0012, Marc Najork, Thorsten Joachims
WSDM1
2019 Addressing Trust Bias for Unbiased Learning-to-Rank
abstract
Existing unbiased learning-to-rank models use counterfactual inference, notably Inverse Propensity Scoring (IPS), to learn a ranking function from biased click data. They handle the click incompleteness bias, but usually assume that the clicks are noise-free, i.e., a clicked document is always assumed to be relevant. In this paper, we relax this unrealistic assumption and study click noise explicitly in the unbiased learning-to-rank setting. Specifically, we model the noise as the position-dependent trust bias and propose a noise-aware Position-Based Model, named TrustPBM, to better capture user click behavior. We propose an Expectation-Maximization algorithm to estimate both examination and trust bias from click data in TrustPBM. Furthermore, we show that it is difficult to use a pure IPS method to incorporate click noise and thus propose a novel method that combines a Bayes rule application with IPS for unbiased learning-to-rank. We evaluate our proposed methods on three personal search data sets and demonstrate that our proposed model can significantly outperform the existing unbiased learning-to-rank methods.
Aman Agarwal, Xuanhui Wang, Cheng Li 0012, Michael Bendersky, Marc Najork
WWW1
2018 Analyzing Behavioral Trends in Community Driven Discussion Platforms Like Reddit
abstract
The aim of this paper is to present methods to systematically analyze individual and group behavioral patterns observed in community driven discussion platforms like Reddit where users exchange information and views on various topics of current interest. We conduct this study by analyzing the statistical behavior of posts and modeling user interactions around them. We have chosen Reddit as an example, since it has grown exponentially from a small community to one of the biggest social network platforms in the recent times. Due to its large user base and popularity, a variety of behavior is present among users in terms of their activity. Our study provides interesting insights about a large number of inactive posts which fail to gather attention despite their authors exhibiting Cyborg-like behavior to draw attention. We also present interesting insights about short-lived but extremely active posts emulating a phenomenon like Mayfly Buzz. Further, we present methods to find the nature of activity around highly active posts to determine the presence of Limelight hogging activity, if any. We analyzed over 2 million posts and more than 7 million user responses to them during entire 2008 and over 63 million posts and over 608 million user responses to them from August 2014 to July 2015 amounting to two one-year periods, in order to understand how social media space has evolved over the years.
Sachin Thukral, Hardik Meisheri, Tushar Kataria, Aman Agarwal, Ishan Verma, Lipika Dey
ASONAM4
2017 Effective Evaluation Using Logged Bandit Feedback from Multiple Loggers
abstract
Accurately evaluating new policies (e.g. ad-placement models, ranking functions, recommendation functions) is one of the key prerequisites for improving interactive systems. While the conventional approach to evaluation relies on online A/B tests, recent work has shown that counterfactual estimators can provide an inexpensive and fast alternative, since they can be applied offline using log data that was collected from a different policy fielded in the past. In this paper, we address the question of how to estimate the performance of a new target policy when we have log data from multiple historic policies. This question is of great relevance in practice, since policies get updated frequently in most online systems. We show that naively combining data from multiple logging policies can be highly suboptimal. In particular, we find that the standard Inverse Propensity Score (IPS) estimator suffers especially when logging and target policies diverge -- to a point where throwing away data improves the variance of the estimator. We therefore propose two alternative estimators which we characterize theoretically and compare experimentally. We find that the new estimators can provide substantially improved estimation accuracy.
Aman Agarwal, Soumya Basu 0003, Tobias Schnabel, Thorsten Joachims
KDD1