VLDB 2026 Research / reviewers in the wild / expert
Asish Ghoshal
dblp:35/10890
· DBLP profile ↗
16ranked-venue papers
10as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 8 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Information extraction and text analysis · 24% Probabilistic and Bayesian machine learning · 20% Learning theory · 13% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 24 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › retrieval models › neural retrieval
multi-vector retrieval |
0.7 | 1 | 2023 | CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval · ACL (1) 2023 |
Information retrieval
retrieval models |
0.7 | 1 | 2023 | CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval · ACL (1) 2023 |
Information retrieval › interactive information retrieval
explainable search |
0.6 | 1 | 2022 | QUASER: Question Answering with Scalable Extractive Rationalization · SIGIR 2022 |
Information retrieval
question answering |
0.6 | 1 | 2022 | QUASER: Question Answering with Scalable Extractive Rationalization · SIGIR 2022 |
Machine learning › Trustworthy machine learning › interpretability
explainable AI |
0.5 | 1 | 2021 | FiD-Ex: Improving Sequence-to-Sequence Models for Extractive Rationale Generation · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
fact-checking |
0.5 | 1 | 2021 | FiD-Ex: Improving Sequence-to-Sequence Models for Extractive Rationale Generation · EMNLP (1) 2021 |
Machine learning › Deep learning architectures and training › regularization
label smoothing |
0.5 | 1 | 2021 | Learning Better Structured Representations Using Low-rank Adaptive Label Smoothing · ICLR 2021 |
Machine learning › Representation and self-supervised learning
structured representation |
0.5 | 1 | 2021 | Learning Better Structured Representations Using Low-rank Adaptive Label Smoothing · ICLR 2021 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.4 | 1 | 2020 | Low-Resource Domain Adaptation for Compositional Task-Oriented Semantic Parsing · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › semantic parsing
task-oriented semantic parsing |
0.4 | 1 | 2020 | Low-Resource Domain Adaptation for Compositional Task-Oriented Semantic Parsing · EMNLP (1) 2020 |
Machine learning › Learning theory
generalization bounds |
0.3 | 1 | 2018 | Learning Maximum-A-Posteriori Perturbation Models for Structured Prediction in Polynomial Time · ICML 2018 |
Machine learning › Learning theory › generalization bounds
rademacher complexity |
0.3 | 1 | 2018 | Learning Maximum-A-Posteriori Perturbation Models for Structured Prediction in Polynomial Time · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.3 | 1 | 2018 | Learning Maximum-A-Posteriori Perturbation Models for Structured Prediction in Polynomial Time · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
bayesian network structure learning |
0.3 | 1 | 2017 | Learning Identifiable Gaussian Bayesian Networks in Polynomial Time and Sample Complexity · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
0.3 | 1 | 2017 | Learning Identifiable Gaussian Bayesian Networks in Polynomial Time and Sample Complexity · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.3 | 1 | 2017 | Learning Identifiable Gaussian Bayesian Networks in Polynomial Time and Sample Complexity · NIPS 2017 |
Knowledge, reasoning and agents › Multi-agent systems
multi-robot systems |
0.2 | 1 | 2013 | Covering space with simple robots: From chains to random trees · ICRA 2013 |
Robotics › Motion planning and robot control
path planning |
0.2 | 1 | 2013 | Covering space with simple robots: From chains to random trees · ICRA 2013 |
Robotics › Motion planning and robot control › motion planning › sampling-based motion planning
RRT |
0.2 | 1 | 2013 | Covering space with simple robots: From chains to random trees · ICRA 2013 |
Robotics › Motion planning and robot control › motion planning › sampling-based motion planning
sampling-based path planning |
0.2 | 1 | 2013 | Covering space with simple robots: From chains to random trees · ICRA 2013 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model |
0.1 | 1 | 2021 | FiD-Ex: Improving Sequence-to-Sequence Models for Extractive Rationale Generation · EMNLP (1) 2021 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
low-resource domain adaptation |
0.1 | 1 | 2020 | Low-Resource Domain Adaptation for Compositional Task-Oriented Semantic Parsing · EMNLP (1) 2020 |
Machine learning › Learning theory
sample complexity |
0.1 | 1 | 2017 | Learning Identifiable Gaussian Bayesian Networks in Polynomial Time and Sample Complexity · NIPS 2017 |
Robotics › Robot navigation and mapping
mobile robot navigation |
0.0 | 1 | 2013 | Covering space with simple robots: From chains to random trees · ICRA 2013 |
Methods — techniques the papers use, named apart from their topics
conditional token interaction · 0.7unsupervised generative models · 0.6pre-trained language model · 0.6multi-task learning · 0.6sentence markers · 0.5label smoothing · 0.5intermediate fine-tuning · 0.5fusion-in-decoder · 0.5meta-learning · 0.4BART · 0.4randomized algorithms · 0.3polynomial-time algorithm · 0.3conditional independence testing · 0.3infrared sensing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Identifying Causal Changes Between Linear Structural Equation ModelsabstractLearning the structures of structural equation models (SEMs) as directed acyclic graphs (DAGs) from data is crucial for representing causal relationships in various scientific domains. Instead of estimating individual DAG structures, it is often preferable to directly estimate changes in causal relations between conditions, such as changes in genetic expression between healthy and diseased subjects. This work studies the problem of directly estimating the difference between two linear SEMs, i.e. *without estimating the individual DAG structures*, given two sets of samples drawn from the individual SEMs. We consider general classes of linear SEMs where the noise distributions are allowed to be Gaussian or non-Gaussian and have different noise variances across the variables in the individual SEMs. We rigorously characterize novel conditions related to the topological layering of the structural difference that lead to the *identifiability* of the difference DAG (DDAG). Moreover, we propose an *efficient* algorithm to identify the DDAG via sequential re-estimation of the difference of precision matrices. A surprising implication of our results is that causal changes can be identifiable even between *non-identifiable* models such as Gaussian SEMs with unequal noise variances. Synthetic experiments are presented to validate our theoretical results and to show the scalability of our method. Vineet Malik, Kevin Bello, Asish Ghoshal, Jean Honorio |
UAI | 3 |
| 2023 | CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector RetrievalabstractMinghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, Xilun Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Minghan Li 0002, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Scott Yih, Xilun Chen 0002 |
ACL (1) | 4 |
| 2022 | QUASER: Question Answering with Scalable Extractive RationalizationabstractDesigning natural language processing (NLP) models that produce predictions by first extracting a set of relevant input sentences, i.e., rationales, is gaining importance for improving model interpretability and producing supporting evidence for users. Current unsupervised approaches are designed to extract rationales that maximize prediction accuracy, which is invariably obtained by exploiting spurious correlations in datasets, and leads to unconvincing rationales. In this paper, we introduce unsupervised generative models to extract dual-purpose rationales, which must not only be able to support a subsequent answer prediction, but also support a reproduction of the input query. We show that such models can produce more meaningful rationales, that are less influenced by dataset artifacts, and as a result, also achieve the state-of-the-art on rationale extraction metrics on four datasets from the ERASER benchmark, significantly improving upon previous unsupervised methods. Our multi-task model is scalable and enables using state-of-the-art pretrained language models to design explainable question answering systems. Asish Ghoshal, Srinivasan Iyer 0001, Bhargavi Paranjape, Kushal Lakhotia, Scott Yih, Yashar Mehdad |
SIGIR | 1 |
| 2021 | Towards Understanding the Behaviors of Optimal Deep Active Learning AlgorithmsabstractActive learning (AL) algorithms may achieve better performance with fewer data because the model guides the data selection process. While many algorithms have been proposed, there is little study on what the optimal AL algorithm looks like, which would help researchers understand where their models fall short and iterate on the design. In this paper, we present a simulated annealing algorithm to search for this optimal oracle and analyze it for several tasks. We present qualitative and quantitative insights into the behaviors of this oracle, comparing and contrasting them with those of various heuristics. Moreover, we are able to consistently improve the heuristics using one particular insight. We hope that our findings can better inform future active learning research. The code is available at https://github.com/YilunZhou/optimal-active-learning. Yilun Zhou, Adithya Renduchintala, Xian Li 0003, Sida I. Wang, Yashar Mehdad, Asish Ghoshal |
AISTATS | 6 |
| 2021 | FiD-Ex: Improving Sequence-to-Sequence Models for Extractive Rationale GenerationabstractNatural language (NL) explanations of model predictions are gaining popularity as a means to understand and verify decisions made by large black-box pre-trained models, for tasks such as Question Answering (QA) and Fact Verification.Recently, pre-trained sequence to sequence (seq2seq) models have proven to be very effective in jointly making predictions, as well as generating NL explanations.However, these models have many shortcomings; they can fabricate explanations even for incorrect predictions, they are difficult to adapt to long input documents, and their training requires a large amount of labeled data.In this paper, we develop FiD-Ex 1 , which addresses these shortcomings for seq2seq models by: 1) introducing sentence markers to eliminate explanation fabrication by encouraging extractive generation, 2) using the fusion-in-decoder architecture to handle long input contexts, and 3) intermediate fine-tuning on re-structured open domain QA datasets to improve few-shot performance.FiD-Ex significantly improves over prior work in terms of explanation metrics and task accuracy on five tasks from the ERASER explainability benchmark in both fully supervised and few-shot settings. Kushal Lakhotia, Bhargavi Paranjape, Asish Ghoshal, Scott Yih, Yashar Mehdad, Srinivasan Iyer 0001 |
EMNLP (1) | 3 |
| 2021 | Learning Better Structured Representations Using Low-rank Adaptive Label Smoothing
Asish Ghoshal, Xilun Chen 0002, Sonal Gupta, Luke Zettlemoyer, Yashar Mehdad |
ICLR | 1 |
| 2020 | Minimax Bounds for Structured Prediction Based on Factor GraphsabstractStructured prediction can be considered as a generalization of many standard supervised learning tasks, and is usually thought as a simultaneous prediction of multiple labels. One standard approach is to maximize a score function on the space of labels, which usually decomposes as a sum of unary and pairwise potentials, each depending on one or two specific labels, respectively.For this approach, several learning and inference algorithms have been proposed over the years, ranging from exact to approximate methods while balancing the computational complexity.However, in contrast to binary and multiclass classification, results on the necessary number of samples for achieving learning are still limited, even for a specific family of predictors such as factor graphs.In this work, we provide minimax lower bounds for a class of general factor-graph inference models in the context of structured prediction.That is, we characterize the necessary sample complexity for any conceivable algorithm to achieve learning of general factor-graph predictors. Kevin Bello, Asish Ghoshal, Jean Honorio |
AISTATS | 2 |
| 2020 | Low-Resource Domain Adaptation for Compositional Task-Oriented Semantic ParsingabstractTask-oriented semantic parsing is a critical component of virtual assistants, which is responsible for understanding the user's intents (set reminder, play music, etc.).Recent advances in deep learning have enabled several approaches to successfully parse more complex queries (Gupta et al., 2018;Rongali et al., 2020), but these models require a large amount of annotated training data to parse queries on new domains (e.g.reminder, music).In this paper, we focus on adapting taskoriented semantic parsers to low-resource domains, and propose a novel method that outperforms a supervised neural model at a 10-fold data reduction.In particular, we identify two fundamental factors for low-resource domain adaptation: better representation learning and better training techniques.Our representation learning uses BART (Lewis et al., 2020) to initialize our model which outperforms encoder-only pre-trained representations used in previous work.Furthermore, we train with optimization-based meta-learning (Finn et al., 2017) to improve generalization to lowresource domains.This approach significantly outperforms all baseline methods in the experiments on a newly collected multi-domain taskoriented semantic parsing dataset (TOPv2 1 ). Xilun Chen 0002, Asish Ghoshal, Yashar Mehdad, Luke Zettlemoyer, Sonal Gupta |
EMNLP (1) | 2 |
| 2018 | Learning linear structural equation models in polynomial time and sample complexityabstractThe problem of learning structural equation models (SEMs) from data is a fundamental problem in causal inference. We develop a new algorithm — which is computationally and statistically efficient and works in the high-dimensional regime — for learning linear SEMs from purely observational data with arbitrary noise distribution. We consider three aspects of the problem: identifiability, computational efficiency, and statistical efficiency. We show that when data is generated from a linear SEM over p nodes and maximum Markov blanket size d, our algorithm recovers the directed acyclic graph (DAG) structure of the SEM under an identifiability condition that is more general than those considered in the literature, and without faithfulness assumptions. In the population setting, our algorithm recovers the DAG structure in $O(p(d + \log p))$ operations. In the finite sample setting, if the estimated precision matrix is sparse, our algorithm has a smoothed complexity of $\tilde{O}(p^3 + pd^{4})$, while if the estimated precision matrix is dense, our algorithm has a smoothed complexity of $\tilde{O}(p^5)$. For sub-Gaussian and bounded ($4m$-th, $m$ being positive integer) moment noise, our algorithm has a sample complexity of $\mathcal{O}(\frac{d^4}{\varepsilon^2} \log (\frac{p}{\sqrt{δ}}))$ and $\mathcal{O}(\frac{d^4}{\varepsilon^2} (\frac{p^2}{δ})^{\nicefrac{1}{m}})$ resp., to achieve $\varepsilon$ element-wise additive error with respect to the true autoregression matrix with probability at least $1 - δ$. Asish Ghoshal, Jean Honorio |
AISTATS | 1 |
| 2018 | Learning Sparse Polymatrix Games in Polynomial Time and Sample ComplexityabstractWe consider the problem of learning sparse polymatrix games from observations of strategic interactions. We show that a polynomial time method based on $\ell_{1,2}$-group regularized logistic regression recovers a game, whose Nash equilibria are the $ε$-Nash equilibria of the game from which the data was generated (true game), in $O(m^4 d^4 \log (pd))$ samples of strategy profiles — where $m$ is the maximum number of pure strategies of a player, $p$ is the number of players, and $d$ is the maximum degree of the game graph. Under slightly more stringent separability conditions on the payoff matrices of the true game, we show that our method learns a game with the exact same Nash equilibria as the true game. We also show that $Ω(d \log (pm))$ samples are necessary for any method to consistently recover a game, with the same Nash-equilibria as the true game, from observations of strategic interactions. Asish Ghoshal, Jean Honorio |
AISTATS | 1 |
| 2018 | Learning Maximum-A-Posteriori Perturbation Models for Structured Prediction in Polynomial TimeabstractMAP perturbation models have emerged as a powerful framework for inference in structured prediction. Such models provide a way to efficiently sample from the Gibbs distribution and facilitate predictions that are robust to random noise. In this paper, we propose a provably polynomial time randomized algorithm for learning the parameters of perturbed MAP predictors. Our approach is based on minimizing a novel Rademacher-based generalization bound on the expected loss of a perturbed MAP predictor, which can be computed in polynomial time. We obtain conditions under which our randomized learning algorithm can guarantee generalization to unseen examples. Asish Ghoshal, Jean Honorio |
ICML | 1 |
| 2018 | A Distributed Classifier for MicroRNA Target Prediction with Validation Through TCGA Expression DataabstractBACKGROUND: MicroRNAs (miRNAs) are approximately 22-nucleotide long regulatory RNA that mediate RNA interference by binding to cognate mRNA target regions. Here, we present a distributed kernel SVM-based binary classification scheme to predict miRNA targets. It captures the spatial profile of miRNA-mRNA interactions via smooth B-spline curves. This is accomplished separately for various input features, such as thermodynamic and sequence-based features. Further, we use a principled approach to uniformly model both canonical and non-canonical seed matches, using a novel seed enrichment metric. Finally, we verify our miRNA-mRNA pairings using an Elastic Net-based regression model on TCGA expression data for four cancer types to estimate the miRNAs that together regulate any given mRNA. RESULTS: We present a suite of algorithms for miRNA target prediction, under the banner Avishkar, with superior prediction performance over the competition. Specifically, our final kernel SVM model, with an Apache Spark backend, achieves an average true positive rate (TPR) of more than 75 percent, when keeping the false positive rate of 20 percent, for non-canonical human miRNA target sites. This is an improvement of over 150 percent in the TPR for non-canonical sites, over the best-in-class algorithm. We are able to achieve such superior performance by representing the thermodynamic and sequence profiles of miRNA-mRNA interaction as curves, devising a novel seed enrichment metric, and learning an ensemble of miRNA family-specific kernel SVM classifiers. We provide an easy-to-use system for large-scale interactive analysis and prediction of miRNA targets. All operations in our system, namely candidate set generation, feature generation and transformation, training, prediction, and computing performance metrics are fully distributed and are scalable. CONCLUSIONS: We have developed an efficient SVM-based model for miRNA target prediction using recent CLIP-seq data, demonstrating superior performance, evaluated using ROC curves for different species (human or mouse), or different target types (canonical or non-canonical). We analyzed the agreement between the target pairings using CLIP-seq data and using expression data from four cancer types. To the best of our knowledge, we provide the first distributed framework for miRNA target prediction based on Apache Hadoop and Spark. AVAILABILITY: All source code and sample data are publicly available at https://bitbucket.org/cellsandmachines/avishkar. Our scalable implementation of kernel SVM using Apache Spark, which can be used to solve large-scale non-linear binary classification problems, is available at https://bitbucket.org/cellsandmachines/kernelsvmspark. Asish Ghoshal, Michael A. Roth, Kevin Xia 0001, Ananth Grama, Somali Chaterji |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Information-theoretic limits of Bayesian network structure learningabstractIn this paper, we study the information-theoretic limits of learning the structure of Bayesian networks (BNs), on discrete as well as continuous random variables, from a finite number of samples. We show that the minimum number of samples required by any procedure to recover the correct structure grows as $Ω(m)$ and $Ω(k \log m + (k^2)/m)$ for non-sparse and sparse BNs respectively, where m is the number of variables and k is the maximum number of parents per node. We provide a simple recipe, based on an extension of the Fano’s inequality, to obtain information-theoretic limits of structure recovery for any exponential family BN. We instantiate our result for specific conditional distributions in the exponential family to characterize the fundamental limits of learning various commonly used BNs, such as conditional probability table based networks, Gaussian BNs, noisy-OR networks, and logistic regression networks. En route to obtaining our main results, we obtain tight bounds on the number of sparse and non-sparse essential-DAGs. Finally, as a byproduct, we recover the information-theoretic limits of sparse variable selection for logistic regression. Asish Ghoshal, Jean Honorio |
AISTATS | 1 |
| 2017 | Learning Graphical Games from Behavioral Data: Sufficient and Necessary ConditionsabstractIn this paper we obtain sufficient and necessary conditions on the number of samples required for exact recovery of the pure-strategy Nash equilibria (PSNE) set of a graphical game from noisy observations of joint actions. We consider sparse linear influence games — a parametric class of graphical games with linear payoffs, and represented by directed graphs of n nodes (players) and in-degree of at most k. We show that one can efficiently recover the PSNE set of a linear influence game with $O(k^2 \log n)$ samples, under very general observation models. On the other hand, we show that $Ω(k \log n)$ samples are necessary for any procedure to recover the PSNE set from observations of joint actions. Asish Ghoshal, Jean Honorio |
AISTATS | 1 |
| 2017 | Learning Identifiable Gaussian Bayesian Networks in Polynomial Time and Sample ComplexityabstractLearning the directed acyclic graph (DAG) structure of a Bayesian network from observational data is a notoriously difficult problem for which many non-identifiability and hardness results are known. In this paper we propose a provably polynomial-time algorithm for learning sparse Gaussian Bayesian networks with equal noise variance --- a class of Bayesian networks for which the DAG structure can be uniquely identified from observational data --- under high-dimensional settings. We show that $O(k^4 \log p)$ number of samples suffices for our method to recover the true DAG structure with high probability, where $p$ is the number of variables and $k$ is the maximum Markov blanket size. We obtain our theoretical guarantees under a condition called \emph{restricted strong adjacency faithfulness} (RSAF), which is strictly weaker than strong faithfulness --- a condition that other methods based on conditional independence testing need for their success. The sample complexity of our method matches the information-theoretic limits in terms of the dependence on $p$. We validate our theoretical findings through synthetic experiments. Asish Ghoshal, Jean Honorio |
NIPS | 1 |
| 2013 | Covering space with simple robots: From chains to random treesabstractInspired by the Rapidly-exploring random tree data-structure and algorithm for path planning, we introduce an approach for spanning physical space with a group of simple mobile robots. Emphasizing minimalism and using only InfraRed and contact sensors for communication, our position unaware robots physically embody elements of the tree. Although robots are fundamentally constrained in the spatial operations they may perform, we show that the approach-implemented on physical robots- remains consistent with the original data-structure idea. In particular, we show that a generalized form of Voronoi bias is present in the construction of the tree, and that such trees have an approximate space-filling property. We present an analysis of the physical system via sets of coupled stochastic equations: the first being the rate-equation for the transitions made by the robot controllers, and the second to capture the spatial process describing tree formation. We are able to provide an understanding of the control parameters in terms of a process mixing-time and show the dependence of the Voronoi bias on an interference parameter which grows as O(√N). Asish Ghoshal, Dylan A. Shell |
ICRA | 1 |