Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mrigank Raman

dblp:271/4429 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 38% Learning theory · 14% Deep learning architectures and training · 14%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
model selection
0.812024
Post-Hoc Reversal: Are We Selecting Models Prematurely? · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.712023
Model-tuning Via Prompts Makes NLP Models Adversarially Robust · EMNLP 2023
Computer vision › Vision and language › vision-language model
prompt learning
0.712023
Model-tuning Via Prompts Makes NLP Models Adversarially Robust · EMNLP 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
Model-tuning Via Prompts Makes NLP Models Adversarially Robust · EMNLP 2023
Machine learning › Trustworthy machine learning › robustness
adversarial attack
0.512021
Learning to Deceive Knowledge Graph Augmented Models via Targeted Perturbation · ICLR 2021
Machine learning › Transfer learning and domain adaptation
domain generalization
0.512021
Generalization on Unseen Domains via Inference-Time Label-Preserving Target Projections · CVPR 2021
Natural language and speech › Language models and text generation › large language model inference
inference-time adaptation
0.512021
Generalization on Unseen Domains via Inference-Time Label-Preserving Target Projections · CVPR 2021
Machine learning › Trustworthy machine learning
uncertainty estimation
0.212024
Post-Hoc Reversal: Are We Selecting Models Prematurely? · NeurIPS 2024
Natural language and speech › Language models and text generation
pre-trained language model
0.212023
Model-tuning Via Prompts Makes NLP Models Adversarially Robust · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

temperature scaling · 0.8stochastic weight averaging · 0.8ensembling · 0.8early stopping · 0.8prompt tuning · 0.7fine-tuning · 0.7adversarial training · 0.7targeted perturbation · 0.5optimization · 0.5metric learning · 0.5generative model · 0.5
YearPublicationVenuePosition
2024 Post-Hoc Reversal: Are We Selecting Models Prematurely?
abstract
Trained models are often composed with post-hoc transforms such as temperature scaling (TS), ensembling and stochastic weight averaging (SWA) to improve performance, robustness, uncertainty estimation, etc. However, such transforms are typically applied only after the base models have already been finalized by standard means. In this paper, we challenge this practice with an extensive empirical study. In particular, we demonstrate a phenomenon that we call post-hoc reversal, where performance trends are reversed after applying post-hoc transforms. This phenomenon is especially prominent in high-noise settings. For example, while base models overfit badly early in training, both ensembling and SWA favor base models trained for more epochs. Post-hoc reversal can also prevent the appearance of double descent and mitigate mismatches between test loss and test error seen in base models. Preliminary analyses suggest that these transforms induce reversal by suppressing the influence of mislabeled examples, exploiting differences in their learning dynamics from those of clean examples. Based on our findings, we propose post-hoc selection, a simple technique whereby post-hoc metrics inform model development decisions such as early stopping, checkpointing, and broader hyperparameter choices. Our experiments span real-world vision, language, tabular and graph datasets. On an LLM instruction tuning dataset, post-hoc selection results in >1.5x MMLU improvement compared to naive selection.
Rishabh Ranjan, Mrigank Raman, Carlos Guestrin, Zachary C. Lipton
NeurIPS3
2023 Model-tuning Via Prompts Makes NLP Models Adversarially Robust
abstract
In recent years, NLP practitioners have converged on the following practice: (i) import an off-the-shelf pretrained (masked) language model; (ii) append a multilayer perceptron atop the CLS token's hidden representation (with randomly initialized weights); and (iii) finetune the entire model on a downstream task (MLP-FT).This procedure has produced massive gains on standard NLP benchmarks, but these models remain brittle, even to mild adversarial perturbations.In this work, we demonstrate surprising gains in adversarial robustness enjoyed by Model-tuning Via Prompts (MVP), an alternative method of adapting to downstream tasks.Rather than appending an MLP head to make output prediction, MVP appends a prompt template to the input, and makes prediction via text infilling/completion. Across 5 NLP datasets, 4 adversarial attacks, and 3 different models, MVP improves performance against adversarial substitutions by an average of 8% over standard methods and even outperforms adversarial training-based state-of-art defenses by 3.5%.By combining MVP with adversarial training, we achieve further improvements in adversarial robustness while maintaining performance on unperturbed examples.Finally, we conduct ablations to investigate the mechanism underlying these gains.Notably, we find that the main causes of vulnerability of MLP-FT can be attributed to the misalignment between pre-training and fine-tuning tasks, and the randomly initialized MLP parameters. 1
Mrigank Raman, Pratyush Maini, J. Zico Kolter, Zachary C. Lipton, Danish Pruthi
EMNLP1
2021 Generalization on Unseen Domains via Inference-Time Label-Preserving Target Projections
abstract
Generalization of machine learning models trained on a set of source domains on unseen target domains with different statistics, is a challenging problem. While many approaches have been proposed to solve this problem, they only utilize source data during training but do not take advantage of the fact that a single target example is available at the time of inference. Motivated by this, we propose a method that effectively uses the target sample during inference beyond mere classification. Our method has three components - (i) A label-preserving feature or metric transformation on source data such that the source samples are clustered in accordance with their class irrespective of their domain (ii) A generative model trained on the these features (iii) A label-preserving projection of the target point on the source-feature manifold during inference via solving an optimization problem on the input space of the generative model using the learned metric. Finally, the projected target is used in the classifier. Since the projected target feature comes from the source manifold and has the same label as the real target by design, the classifier is expected to perform better on it than the true target. We demonstrate that our method outperforms the state-of-the-art Domain Generalization methods on multiple datasets and tasks.
Prashant Pandey 0002, Mrigank Raman, Sumanth Varambally, Prathosh A. P.
CVPR2
2021 Learning to Deceive Knowledge Graph Augmented Models via Targeted Perturbation
Mrigank Raman, Aaron Chan, Siddhant Agarwal, Peifeng Wang, Hansen Wang, Sungchul Kim, Ryan Rossi, Handong Zhao, Nedim Lipka, Xiang Ren 0001
ICLR1
2020 Centralized active tracking of a Markov chain with unknown dynamics
abstract
In this paper, selection of an active sensor subset for tracking a discrete time, finite state Markov chain having an unknown transition probability matrix (TPM) is considered. A total of N sensors are available for making observations of the Markov chain, out of which a subset of sensors are activated each time in order to perform reliable estimation of the process. The trade-off is between activating more sensors to gather more observations for the remote estimation, and restricting sensor usage in order to save energy and bandwidth consumption. The problem is formulated as a constrained minimization problem, where the objective is the long-run averaged mean-squared error (MSE) in estimation, and the constraint is on sensor activation rate. A Lagrangian relaxation of the problem is solved by an artful blending of two tools: Gibbs sampling for MSE minimization and an on-line version of expectation maximization (EM) to estimate the unknown TPM. Finally, the Lagrange multiplier is updated using slower timescale stochastic approximation in order to satisfy the sensor activation rate constraint. The on-line EM algorithm, though adapted from literature, can estimate vector-valued parameters even under time-varying dimension of the sensor observations. Numerical results demonstrate approximately 1 dB better error performance than uniform sensor sampling and comparable error performance (within 2 dB bound) against complete sensor observation. This makes the proposed algorithm amenable to practical implementation.
Mrigank Raman, Ojal Kumar, Arpan Chattopadhyay
MASS1