VLDB 2026 Research / reviewers in the wild / expert
Marjan Ghazvininejad
dblp:50/10813
· DBLP profile ↗
27ranked-venue papers
7as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed InputsabstractReward models have become a staple in modern NLP, serving as not only a scalable text evaluator, but also an indispensable component in many alignment recipes and inference-time algorithms.However, while recent reward models increase performance on standard benchmarks, this may partly be due to overfitting effects, which would confound an understanding of their true capability.In this work, we scrutinize the robustness of reward models and the extent of such overfitting.We build re-WordBench, which systematically transforms reward model inputs in meaning-or rankingpreserving ways.We show that state-of-theart reward models suffer from substantial performance degradation even with minor input transformations, sometimes dropping to significantly below-random accuracy, suggesting brittleness.To improve reward model robustness, we propose to explicitly train them to assign similar scores to paraphrases, and find that this approach also improves robustness to other distinct kinds of transformations.For example, our robust reward model reduces such degradation by roughly half for the Chat Hard subset in RewardBench.Furthermore, when used in alignment, our robust reward models demonstrate better utility and lead to higher-quality outputs, winning in up to 59% of instances against a standardly trained RM. Zhaofeng Wu, Michihiro Yasunaga, Andrew Cohen, Asli Celikyilmaz, Marjan Ghazvininejad |
EMNLP | 6 |
| 2025 | Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-JudgeabstractLLM-as-a-Judge models generate chain-of-thought (CoT) sequences intended to capture the step-by-step reasoning process that underlies the final evaluation of a response. However, due to the lack of human-annotated CoTs for evaluation, the required components and structure of effective reasoning traces remain understudied. Consequently, previous approaches often (1) constrain reasoning traces to hand-designed components, such as a list of criteria, reference answers, or verification questions and (2) structure them such that planning is intertwined with the reasoning for evaluation. In this work, we propose EvalPlanner, a preference optimization algorithm for Thinking-LLM-as-a-Judge that first generates an unconstrained evaluation plan, followed by its execution, and then the final judgment. In a self-training loop, EvalPlanner iteratively optimizes over synthetically constructed evaluation plans and executions, leading to better final verdicts. Our method achieves a new state-of-the-art performance for generative reward models on RewardBench and PPE, despite being trained on fewer amount of, and synthetically generated, preference pairs. Additional experiments on other benchmarks like RM-Bench, JudgeBench, and FollowBenchEval further highlight the utility of both planning and reasoning for building robust LLM-as-a-Judge reasoning models. Swarnadeep Saha, Xian Li 0003, Marjan Ghazvininejad, Jason Weston |
ICML | 3 |
| 2024 | Representation Deficiency in Masked Language ModelingabstractMasked Language Modeling (MLM) has been one of the most prominent approaches for pretraining bidirectional text encoders due to its simplicity and effectiveness. One notable concern about MLM is that the special $\texttt{[MASK]}$ symbol causes a discrepancy between pretraining data and downstream data as it is present only in pretraining but not in fine-tuning. In this work, we offer a new perspective on the consequence of such a discrepancy: We demonstrate empirically and theoretically that MLM pretraining allocates some model dimensions exclusively for representing $\texttt{[MASK]}$ tokens, resulting in a representation deficiency for real tokens and limiting the pretrained model's expressiveness when it is adapted to downstream data without $\texttt{[MASK]}$ tokens. Motivated by the identified issue, we propose MAE-LM, which pretrains the Masked Autoencoder architecture with MLM where $\texttt{[MASK]}$ tokens are excluded from the encoder. Empirically, we show that MAE-LM improves the utilization of model dimensions for real token representations, and MAE-LM consistently outperforms MLM-pretrained models on the GLUE and SQuAD benchmarks. Yu Meng 0001, Jitin Krishnan, Sinong Wang, Qifan Wang 0001, Yuning Mao, Marjan Ghazvininejad, Jiawei Han 0001, Luke Zettlemoyer |
ICLR | 7 |
| 2024 | David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMsabstractXiaochuang Han, Sachin Kumar, Yulia Tsvetkov, Marjan Ghazvininejad. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xiaochuang Han, Sachin Kumar 0009, Yulia Tsvetkov, Marjan Ghazvininejad |
NAACL-HLT | 4 |
| 2023 | XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language ModelsabstractLarge multilingual language models typically rely on a single vocabulary shared across 100+ languages.As these models have increased in parameter count and depth, vocabulary size has remained largely unchanged.This vocabulary bottleneck limits the representational capabilities of multilingual models like XLM-R.In this paper, we introduce a new approach for scaling to very large multilingual vocabularies by de-emphasizing token sharing between languages with little lexical overlap and assigning vocabulary capacity to achieve sufficient coverage for each individual language.Tokenizations using our vocabulary are typically more semantically meaningful and shorter compared to XLM-R.Leveraging this improved vocabulary, we train XLM-V, a multilingual language model with a one million token vocabulary.XLM-V outperforms XLM-R on every task we tested on ranging from natural language inference (XNLI), question answering (MLQA, XQuAD, TyDiQA), to named entity recognition (WikiAnn).XLM-V is particularly effective on low-resource language tasks and outperforms XLM-R by 11.2% and 5.8% absolute on MasakhaNER and Americas NLI, respectively. Davis Liang, Hila Gonen, Yuning Mao, Naman Goyal 0001, Marjan Ghazvininejad, Luke Zettlemoyer, Madian Khabsa |
EMNLP | 6 |
| 2022 | Discourse-Aware Soft Prompting for Text GenerationabstractCurrent efficient fine-tuning methods (e.g., adapters (Houlsby et al., 2019), prefix-tuning (Li and Liang, 2021), etc.) have optimized conditional text generation via training a small set of extra parameters of the neural language model, while freezing the rest for efficiency.While showing strong performance on some generation tasks, they don't generalize across all generation tasks.We show that soft-prompt based conditional text generation can be improved with simple and efficient methods that simulate modeling the discourse structure of human written text.We investigate two design choices: First, we apply hierarchical blocking on the prefix parameters to simulate a higherlevel discourse structure of human written text.Second, we apply attention sparsity on the prefix parameters at different layers of the network and learn sparse transformations on the softmax-function.We show that structured design of prefix parameters yields more coherent, faithful and relevant generations than the baseline prefix-tuning on all generation tasks. Marjan Ghazvininejad, Vladimir Karpukhin, Vera Gor, Asli Celikyilmaz |
EMNLP | 1 |
| 2022 | Natural Language to Code Translation with ExecutionabstractGenerative models of code, pretrained on large corpora of programs, have shown great success in translating natural language to code (Chen et al., 2021;Austin et al., 2021; Li et al., 2022, inter alia).While these models do not explicitly incorporate program semantics (i.e., execution results) during training, they are able to generate correct solutions for many problems.However, choosing a single correct program from a generated set for each problem remains challenging.In this work, we introduce execution resultbased minimum Bayes risk decoding (MBR-EXEC) for program selection and show that it improves the few-shot performance of pretrained code models on natural-language-tocode tasks.We select output programs from a generated candidate set by marginalizing over program implementations that share the same semantics.Because exact equivalence is intractable, we execute each program on a small number of test inputs to approximate semantic equivalence.Across datasets, execution or simulated execution significantly outperforms the methods that do not involve program semantics.We find that MBR-EXEC consistently improves over all execution-unaware selection methods, suggesting it as an effective approach for natural language to code translation.1 Freda Shi, Daniel Fried, Marjan Ghazvininejad, Luke Zettlemoyer, Sida I. Wang |
EMNLP | 3 |
| 2021 | Recipes for Adapting Pre-trained Monolingual and Multilingual Models to Machine TranslationabstractThere has been recent success in pre-training on monolingual data and fine-tuning on Machine Translation (MT), but it remains unclear how to best leverage a pre-trained model for a given MT task.This paper investigates the benefits and drawbacks of freezing parameters, and adding new ones, when fine-tuning a pre-trained model on MT.We focus on 1) Fine-tuning a model trained only on English monolingual data, BART.2) Fine-tuning a model trained on monolingual data from 25 languages, mBART.For BART we get the best performance by freezing most of the model parameters, and adding extra positional embeddings.For mBART we match or outperform the performance of naive fine-tuning for most language pairs with the encoder, and most of the decoder, frozen.The encoder-decoder attention parameters are most important to finetune.When constraining ourselves to an outof-domain training set for Vietnamese to English we see the largest improvements over the fine-tuning baseline. Asa Cooper Stickland, Xian Li 0003, Marjan Ghazvininejad |
EACL | 3 |
| 2021 | Distributionally Robust Multilingual Machine TranslationabstractMultilingual neural machine translation (MNMT) learns to translate multiple language pairs with a single model, potentially improving both the accuracy and the memoryefficiency of deployed models.However, the heavy data imbalance between languages hinders the model from performing uniformly across language pairs.In this paper, we propose a new learning objective for MNMT based on distributionally robust optimization, which minimizes the worst-case expected loss over the set of language pairs.We further show how to practically optimize this objective for large translation corpora using an iterated best response scheme, which is both effective and incurs negligible additional computational cost compared to standard empirical risk minimization.We perform extensive experiments on three sets of languages from two datasets and show that our method consistently outperforms strong baseline methods in terms of average and per-language performance under both many-to-one and one-to-many translation settings.1 Chunting Zhou, Daniel Levy 0002, Xian Li 0003, Marjan Ghazvininejad, Graham Neubig |
EMNLP (1) | 4 |
| 2021 | DeLighT: Deep and Light-weight Transformer
Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer 0001, Luke Zettlemoyer, Hannaneh Hajishirzi |
ICLR | 2 |
| 2021 | Non-Autoregressive Semantic Parsing for Compositional Task-Oriented DialogabstractArun Babu, Akshat Shrivastava, Armen Aghajanyan, Ahmed Aly, Angela Fan, Marjan Ghazvininejad. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Arun Babu, Akshat Shrivastava, Armen Aghajanyan, Ahmed Aly, Angela Fan, Marjan Ghazvininejad |
NAACL-HLT | 6 |
| 2021 | Improving Zero and Few-Shot Abstractive Summarization with Intermediate Fine-tuning and Data AugmentationabstractAlexander Fabbri, Simeng Han, Haoyuan Li, Haoran Li, Marjan Ghazvininejad, Shafiq Joty, Dragomir Radev, Yashar Mehdad. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Alexander R. Fabbri, Simeng Han, Haoran Li 0007, Marjan Ghazvininejad, Shafiq R. Joty, Dragomir R. Radev, Yashar Mehdad |
NAACL-HLT | 5 |
| 2020 | Simple and Effective Retrieve-Edit-Rerank Text GenerationabstractRetrieve-and-edit seq2seq methods typically retrieve an output from the training set and learn a model to edit it to produce the final output.We propose to extend this framework with a simple and effective post-generation ranking approach.Our framework (i) retrieves several potentially relevant outputs for each input, (ii) edits each candidate independently, and (iii) re-ranks the edited candidates to select the final output.We use a standard editing model with simple task-specific reranking approaches, and we show empirically that this approach outperforms existing, significantly more complex methodologies.Experiments on two machine translation (MT) datasets show new state-of-art results.We also achieve near state-of-art performance on the Gigaword summarization dataset, where our analyses show that there is significant room for performance improvement with better candidate output selection in future work. Nabil Hossain, Marjan Ghazvininejad, Luke Zettlemoyer |
ACL | 2 |
| 2020 | BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionabstractMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Mike Lewis, Yinhan Liu, Naman Goyal 0001, Marjan Ghazvininejad, Abdel-rahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer |
ACL | 4 |
| 2020 | Aligned Cross Entropy for Non-Autoregressive Machine TranslationabstractNon-autoregressive machine translation models significantly speed up decoding by allowing for parallel prediction of the entire target sequence. However, modeling word order is more challenging due to the lack of autoregressive factors in the model. This difficultly is compounded during training with cross entropy loss, which can highly penalize small shifts in word order. In this paper, we propose aligned cross entropy (AXE) as an alternative loss function for training of non-autoregressive models. AXE uses a differentiable dynamic program to assign loss based on the best possible monotonic alignment between target tokens and model predictions. AXE-based training of conditional masked language models (CMLMs) substantially improves performance on major WMT benchmarks, while setting a new state of the art for non-autoregressive models. Marjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer Levy |
ICML | 1 |
| 2020 | Non-autoregressive Machine Translation with Disentangled Context TransformerabstractState-of-the-art neural machine translation models generate a translation from left to right and every step is conditioned on the previously generated tokens. The sequential nature of this generation process causes fundamental latency in inference since we cannot generate multiple tokens in each sentence in parallel. We propose an attention-masking based model, called Disentangled Context (DisCo) transformer, that simultaneously generates all tokens given different contexts. The DisCo transformer is trained to predict every output token given an arbitrary subset of the other reference tokens. We also develop the parallel easy-first inference algorithm, which iteratively refines every token in parallel and reduces the number of required iterations. Our extensive experiments on 7 translation directions with varying data sizes demonstrate that our model achieves competitive, if not better, performance compared to the state of the art in non-autoregressive machine translation while significantly reducing decoding time on average. Jungo Kasai, James Cross 0003, Marjan Ghazvininejad, Jiatao Gu |
ICML | 3 |
| 2020 | Pre-training via ParaphrasingabstractWe introduce MARGE, a pre-trained sequence-to-sequence model learned with an unsupervised multi-lingual multi-document paraphrasing objective. MARGE provides an alternative to the dominant masked language modeling paradigm, where we self-supervise the \emph{reconstruction} of target text by \emph{retrieving} a set of related texts (in many languages) and conditioning on them to maximize the likelihood of generating the original. We show it is possible to jointly learn to do retrieval and reconstruction, given only a random initialization. The objective noisily captures aspects of paraphrase, translation, multi-document summarization, and information retrieval, allowing for strong zero-shot performance on several tasks. For example, with no additional task-specific training we achieve BLEU scores of up to 35.8 for document translation. We further show that fine-tuning gives strong performance on a range of discriminative and generative tasks in many languages, making MARGE the most generally applicable pre-training method to date. Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida I. Wang, Luke Zettlemoyer |
NeurIPS | 2 |
| 2020 | Multilingual Denoising Pre-training for Neural Machine TranslationabstractThis paper demonstrates that multilingual denoising pre-training produces significant performance gains across a wide variety of machine translation (MT) tasks. We present mBART—a sequence-to-sequence denoising auto-encoder pre-trained on large-scale monolingual corpora in many languages using the BART objective (Lewis et al., 2019 ). mBART is the first method for pre-training a complete sequence-to-sequence model by denoising full texts in multiple languages, whereas previous approaches have focused only on the encoder, decoder, or reconstructing parts of the text. Pre-training a complete model allows it to be directly fine-tuned for supervised (both sentence-level and document-level) and unsupervised machine translation, with no task- specific modifications. We demonstrate that adding mBART initialization produces performance gains in all but the highest-resource settings, including up to 12 BLEU points for low resource MT and over 5 BLEU points for many document-level and unsupervised models. We also show that it enables transfer to language pairs with no bi-text or that were not in the pre-training corpus, and present extensive analysis of which factors contribute the most to effective pre-training. 1 Yinhan Liu, Jiatao Gu, Naman Goyal 0001, Xian Li 0003, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, Luke Zettlemoyer |
Trans. Assoc. Comput. Linguistics | 6 |
| 2019 | Translating Translationese: A Two-Step Approach to Unsupervised Machine TranslationabstractGiven a rough, word-by-word gloss of a source language sentence, target language natives can uncover the latent, fully-fluent rendering of the translation.In this work we explore this intuition by breaking translation into a two step process: generating a rough gloss by means of a dictionary and then 'translating' the resulting pseudo-translation, or 'Translationese' into a fully fluent translation.We build our Translationese decoder once from a mish-mash of parallel data that has the target language in common and then can build dictionaries on demand using unsupervised techniques, resulting in rapidly generated unsupervised neural MT systems for many source languages.We apply this process to 14 test languages, obtaining better or comparable translation results on high-resource languages than previously published unsupervised MT studies, and obtaining good quality results for low-resource languages that have never been used in an unsupervised MT scenario. Nima Pourdamghani, Nada Aldarrab, Marjan Ghazvininejad, Kevin Knight, Jonathan May |
ACL (1) | 3 |
| 2019 | Mask-Predict: Parallel Decoding of Conditional Masked Language ModelsabstractMarjan Ghazvininejad, Omer Levy, Yinhan Liu, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Marjan Ghazvininejad, Omer Levy, Yinhan Liu, Luke Zettlemoyer |
EMNLP/IJCNLP (1) | 1 |
| 2018 | A Knowledge-Grounded Neural Conversation ModelabstractNeural network models are capable of generating extremely natural sounding conversational interactions. However, these models have been mostly applied to casual scenarios (e.g., as “chatbots”) and have yet to demonstrate they can serve in more useful conversational applications. This paper presents a novel, fully data-driven, and knowledge-grounded neural conversation model aimed at producing more contentful responses. We generalize the widely-used Sequence-to-Sequence (Seq2Seq) approach by conditioning responses on both conversation history and external “facts”, allowing the model to be versatile and applicable in an open-domain setting. Our approach yields significant improvements over a competitive Seq2Seq baseline. Human judges found that our outputs are significantly more informative. Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, William B. Dolan, Jianfeng Gao 0001, Scott Yih, Michel Galley |
AAAI | 1 |
| 2016 | Generating Topical PoetryabstractWe describe Hafez, a program that generates any number of distinct poems on a usersupplied topic.Poems obey rhythmic and rhyme constraints.We describe the poetrygeneration algorithm, give experimental data concerning its parameters, and show its generality with respect to language and poetic form. Marjan Ghazvininejad, Yejin Choi 0001, Kevin Knight |
EMNLP | 1 |
| 2015 | How Much Information Does a Human Translator Add to the Original?abstractWe ask how much information a human translator adds to an original text, and we provide a bound.We address this question in the context of bilingual text compression: given a source text, how many bits of additional information are required to specify the target text produced by a human translator?We develop new compression algorithms and establish a benchmark task. Barret Zoph, Marjan Ghazvininejad, Kevin Knight |
EMNLP | 2 |
| 2015 | How to Memorize a Random 60-Bit StringabstractUser-generated passwords tend to be memorable, but not secure.A random, computergenerated 60-bit string is much more secure.However, users cannot memorize random 60bit strings.In this paper, we investigate methods for converting arbitrary bit strings into English word sequences (both prose and poetry), and we study their memorability and other properties. Marjan Ghazvininejad, Kevin Knight |
HLT-NAACL | 1 |
| 2015 | Hierarchical Active Transfer LearningabstractWe describe a unified active transfer learning framework called Hierarchical Active Transfer Learning (HATL). HATL exploits cluster structure shared between different data domains to perform transfer learning by imputing labels for unlabeled target data and to generate effective label queries during active learning. The resulting framework is flexible enough to perform not only adaptive transfer learning and accelerated active learning but also unsupervised and semi-supervised transfer learning. We derive an intuitive and useful upper bound on HATL's error when used to infer labels for unlabeled target points. We also present results on synthetic data that confirm both intuition and our analysis. Finally, we demonstrate HATL's empirical effectiveness on a benchmark data set for sentiment classification. David C. Kale, Marjan Ghazvininejad, Anil Ramakrishna, Jingrui He, Yan Liu 0002 |
SDM | 2 |
| 2013 | From Local Similarity to Global Coding: An Application to Image ClassificationabstractBag of words models for feature extraction have demonstrated top-notch performance in image classification. These representations are usually accompanied by a coding method. Recently, methods that code a descriptor giving regard to its nearby bases have proved efficacious. These methods take into account the nonlinear structure of descriptors, since local similarities are a good approximation of global similarities. However, they confine their usage of the global similarities to nearby bases. In this paper, we propose a coding scheme that brings into focus the manifold structure of descriptors, and devise a method to compute the global similarities of descriptors to the bases. Given a local similarity measure between bases, a global measure is computed. Exploiting the local similarity of a descriptor and its nearby bases, a global measure of association of a descriptor to all the bases is computed. Unlike the locality-based and sparse coding methods, the proposed coding varies smoothly with respect to the underlying manifold. Experiments on benchmark image classification datasets substantiate the superiority of the proposed method over its locality and sparsity based rivals. Amirreza Shaban, Hamid R. Rabiee 0001, Mehrdad Farajtabar, Marjan Ghazvininejad |
CVPR | 4 |
| 2011 | Isograph: Neighbourhood Graph Construction Based on Geodesic Distance for Semi-supervised LearningabstractSemi-supervised learning based on manifolds has been the focus of extensive research in recent years. Convenient neighbourhood graph construction is a key component of a successful semi-supervised classification method. Previous graph construction methods fail when there are pairs of data points that have small Euclidean distance, but are far apart over the manifold. To overcome this problem, we start with an arbitrary neighbourhood graph and iteratively update the edge weights by using the estimates of the geodesic distances between points. Moreover, we provide theoretical bounds on the values of estimated geodesic distances. Experimental results on real-world data show significant improvement compared to the previous graph construction methods. Marjan Ghazvininejad, Mostafa Mahdieh, Hamid R. Rabiee 0001, Parisa Khanipour Roshan, Mohammad H. Rohban |
ICDM | 1 |