VLDB 2026 Research / reviewers in the wild / expert
Tahira Naseem
dblp:44/642
· DBLP profile ↗
23ranked-venue papers
6as first author
13since 2021 · last 2025
0009-0009-0603-2760ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Latent Principle Discovery for Language Model Self-ImprovementabstractWhen language model (LM) users aim to improve the quality of its generations, it is crucial to specify concrete behavioral attributes that the model should strive to reflect. However, curating such principles across many domains, even non-exhaustively, requires a labor-intensive annotation process. To automate this process, we propose eliciting these latent attributes that guide model reasoning toward human-preferred responses by explicitly modeling them in a self-correction setting. Our approach mines new principles from the LM itself and compresses the discovered elements to an interpretable set via clustering. Specifically, we employ a form of posterior-regularized Monte Carlo Expectation-Maximization to both identify a condensed set of the most effective latent principles and teach the LM to strategically invoke them in order to intrinsically refine its responses. We demonstrate that bootstrapping our algorithm over multiple iterations enables smaller language models (7-8B parameters) to self-improve, achieving +8-10\% in AlpacaEval win-rate, an average of +0.3 on MT-Bench, and +19-23\% in principle-following win-rate on IFEval. We also show that clustering the principles yields interpretable and diverse model-generated constitutions while retaining model performance. The gains that our method achieves highlight the potential of automated, principle-driven post-training recipes toward continual self-improvement. Keshav Ramji, Tahira Naseem, Ramón Fernandez Astudillo |
NeurIPS | 2 |
| 2024 | BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedbackabstractDistribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks. Gaurav Pandey 0001, Yatin Nandwani, Tahira Naseem, Guangxuan Xu, Dinesh Raghu, Sachindra Joshi, Asim Munawar, Ramón Fernandez Astudillo |
ICML | 3 |
| 2023 | Laziness Is a Virtue When It Comes to Compositionality in Neural Semantic ParsingabstractMaxwell Crouse, Pavan Kapanipathi, Subhajit Chaudhury, Tahira Naseem, Ramon Fernandez Astudillo, Achille Fokoue, Tim Klinger. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Maxwell Crouse, Pavan Kapanipathi, Subhajit Chaudhury, Tahira Naseem, Ramón Fernandez Astudillo, Achille Fokoue, Tim Klinger |
ACL (1) | 4 |
| 2023 | Alignment via Mutual InformationabstractMany language learning tasks require learners to infer correspondences between data in two modalities.Often, these alignments are manyto-many and context-sensitive.For example, translating into morphologically rich languages requires learning not just how words, but morphemes, should be translated; words and morphemes may have different meanings (or groundings) depending on the context in which they are used.We describe an informationtheoretic approach to context-sensitive, manyto-many alignment.Our approach first trains a masked sequence model to place distributions over missing spans in (source, target) sequences.Next, it uses this model to compute pointwise mutual information between source and target spans conditional on context.Finally, it aligns spans with high mutual information.We apply this approach to two learning problems: character-based word translation (using alignments for joint morphological segmentation and lexicon learning) and visually grounded reference resolution (using alignments to jointly localize referents and learn word meanings).In both cases, our proposed approach outperforms both structured and neural baselines, showing that conditional mutual information offers an effective framework for formalizing alignment problems in general domains. Shinjini Ghosh, Ramón Fernandez Astudillo, Tahira Naseem, Jacob Andreas |
CoNLL | 4 |
| 2022 | X-FACTOR: A Cross-metric Evaluation of Factual Correctness in Abstractive SummarizationabstractSubhajit Chaudhury, Sarathkrishna Swaminathan, Chulaka Gunasekara, Maxwell Crouse, Srinivas Ravishankar, Daiki Kimura, Keerthiram Murugesan, Ramón Fernandez Astudillo, Tahira Naseem, Pavan Kapanipathi, Alexander Gray. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Subhajit Chaudhury, Sarathkrishna Swaminathan, R. Chulaka Gunasekara, Maxwell Crouse, Srinivas Ravishankar, Daiki Kimura, Keerthiram Murugesan, Ramón Fernandez Astudillo, Tahira Naseem, Pavan Kapanipathi, Alexander G. Gray |
EMNLP | 9 |
| 2022 | Inducing and Using Alignments for Transition-based AMR ParsingabstractAndrew Drozdov, Jiawei Zhou, Radu Florian, Andrew McCallum, Tahira Naseem, Yoon Kim, Ramón Astudillo. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Andrew Drozdov, Jiawei Zhou 0001, Radu Florian, Andrew McCallum, Tahira Naseem, Ramón Fernandez Astudillo |
NAACL-HLT | 5 |
| 2022 | Maximum Bayes Smatch Ensemble Distillation for AMR ParsingabstractYoung-Suk Lee, Ramón Astudillo, Hoang Thanh Lam, Tahira Naseem, Radu Florian, Salim Roukos. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Young-Suk Lee 0001, Ramón Fernandez Astudillo, Hoang Thanh Lam, Tahira Naseem, Radu Florian, Salim Roukos |
NAACL-HLT | 4 |
| 2022 | DocAMR: Multi-Sentence AMR Representation and EvaluationabstractTahira Naseem, Austin Blodgett, Sadhana Kumaravel, Tim O’Gorman, Young-Suk Lee, Jeffrey Flanigan, Ramón Astudillo, Radu Florian, Salim Roukos, Nathan Schneider. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tahira Naseem, Austin Blodgett, Sadhana Kumaravel, Tim O'Gorman, Young-Suk Lee 0001, Jeffrey Flanigan, Ramón Fernandez Astudillo, Radu Florian, Salim Roukos, Nathan Schneider 0001 |
NAACL-HLT | 1 |
| 2021 | Structural Guidance for Transformer Language ModelsabstractPeng Qian, Tahira Naseem, Roger Levy, Ramón Fernandez Astudillo. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tahira Naseem, Roger Levy, Ramón Fernandez Astudillo |
ACL/IJCNLP (1) | 2 |
| 2021 | Bootstrapping Multilingual AMR with Contextual Word AlignmentsabstractJanaki Sheth, Young-Suk Lee, Ramón Fernandez Astudillo, Tahira Naseem, Radu Florian, Salim Roukos, Todd Ward. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Janaki Sheth, Young-Suk Lee 0001, Ramón Fernandez Astudillo, Tahira Naseem, Radu Florian, Salim Roukos, Todd Ward |
EACL | 4 |
| 2021 | Structure-aware Fine-tuning of Sequence-to-sequence Transformers for Transition-based AMR ParsingabstractPredicting linearized Abstract Meaning Representation (AMR) graphs using pre-trained sequence-to-sequence Transformer models has recently led to large improvements on AMR parsing benchmarks.These parsers are simple and avoid explicit modeling of structure but lack desirable properties such as graph well-formedness guarantees or built-in graph-sentence alignments.In this work we explore the integration of general pre-trained sequence-to-sequence language models and a structure-aware transition-based approach.We depart from a pointer-based transition system and propose a simplified transition set, designed to better exploit pre-trained language models for structured fine-tuning.We also explore modeling the parser state within the pre-trained encoder-decoder architecture and different vocabulary strategies for the same purpose.We provide a detailed comparison with recent progress in AMR parsing and show that the proposed parser retains the desirable properties of previous transition-based approaches, while being simpler and reaching the new parsing state of the art for AMR 2.0, without the need for graph re-categorization. Jiawei Zhou 0001, Tahira Naseem, Ramón Fernandez Astudillo, Young-Suk Lee 0001, Radu Florian, Salim Roukos |
EMNLP (1) | 2 |
| 2021 | AMR Parsing with Action-Pointer TransformerabstractJiawei Zhou, Tahira Naseem, Ramón Fernandez Astudillo, Radu Florian. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jiawei Zhou 0001, Tahira Naseem, Ramón Fernandez Astudillo, Radu Florian |
NAACL-HLT | 2 |
| 2021 | Generative Relation Linking for Question Answering over Knowledge Bases
Gaetano Rossiello, Nandana Mihindukulasooriya, Ibrahim Abdelaziz, Mihaela A. Bornea, Alfio Massimiliano Gliozzo, Tahira Naseem, Pavan Kapanipathi |
ISWC | 6 |
| 2020 | GPT-too: A Language-Model-First Approach for AMR-to-Text GenerationabstractManuel Mager, Ramón Fernandez Astudillo, Tahira Naseem, Md Arafat Sultan, Young-Suk Lee, Radu Florian, Salim Roukos. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Manuel Mager, Ramón Fernandez Astudillo, Tahira Naseem, Md. Arafat Sultan, Young-Suk Lee 0001, Radu Florian, Salim Roukos |
ACL | 3 |
| 2019 | Rewarding Smatch: Transition-Based AMR Parsing with Reinforcement LearningabstractOur work involves enriching the Stack-LSTM transition-based AMR parser (Ballesteros and Al-Onaizan, 2017) by augmenting training with Policy Learning and rewarding the Smatch score of sampled graphs.In addition, we also combined several AMR-to-text alignments with an attention mechanism and we supplemented the parser with pre-processed concept identification, named entities and contextualized embeddings.We achieve a highly competitive performance that is comparable to the best published results.We show an indepth study ablating each of the new components of the parser. Tahira Naseem, Abhishek Shah, Hui Wan 0001, Radu Florian, Salim Roukos, Miguel Ballesteros |
ACL (1) | 1 |
| 2012 | Selective Sharing for Multilingual Dependency Parsing
Tahira Naseem, Regina Barzilay, Amir Globerson |
ACL (1) | 1 |
| 2011 | Using Semantic Cues to Learn SyntaxabstractWe present a method for dependency grammar induction that utilizes sparse annotations of semantic relations. This induction set-up is attractive because such annotations provide useful clues about the underlying syntactic structure, and they are readily available in many domains (e.g., info-boxes and HTML markup). Our method is based on the intuition that syntactic realizations of the same semantic predicate exhibit some degree of consistency. We incorporate this intuition in a directed graphical model that tightly links the syntactic and semantic structures. This design enables us to exploit syntactic regularities while still allowing for variations. Another strength of the model lies in its ability to capture non-local dependency relations. Our results demonstrate that even a small amount of semantic annotations greatly improves the accuracy of learned dependencies when tested on both in-domain and out-of-domain texts. Tahira Naseem, Regina Barzilay |
AAAI | 1 |
| 2011 | In-domain Relation Discovery with Meta-constraints via Posterior Regularization
Harr Chen, Edward Benson, Tahira Naseem, Regina Barzilay |
ACL | 3 |
| 2010 | Using Universal Linguistic Knowledge to Guide Grammar Induction
Tahira Naseem, Harr Chen, Regina Barzilay |
EMNLP | 1 |
| 2009 | Unsupervised Multilingual Grammar Induction
Benjamin Snyder, Tahira Naseem, Regina Barzilay |
ACL/IJCNLP | 2 |
| 2009 | Adding More Languages Improves Unsupervised Multilingual Part-of-Speech Tagging: a Bayesian Non-Parametric Approach
Benjamin Snyder, Tahira Naseem, Jacob Eisenstein, Regina Barzilay |
HLT-NAACL | 2 |
| 2009 | Multilingual Part-of-Speech Tagging: Two Unsupervised ApproachesabstractWe demonstrate the effectiveness of multilingual learning for unsupervised part-of-speech tagging. The central assumption of our work is that by combining cues from multiple languages, the structure of each becomes more apparent. We consider two ways of applying this intuition to the problem of unsupervised part-of-speech tagging: a model that directly merges tag structures for a pair of languages into a single sequence and a second model which instead incorporates multilingual context using latent variables. Both approaches are formulated as hierarchical Bayesian models, using Markov Chain Monte Carlo sampling techniques for inference. Our results demonstrate that by incorporating multilingual evidence we can achieve impressive performance gains across a range of scenarios. We also found that performance improves steadily as the number of available languages increases. Tahira Naseem, Benjamin Snyder, Jacob Eisenstein, Regina Barzilay |
J. Artif. Intell. Res. | 1 |
| 2008 | Unsupervised Multilingual Learning for POS Tagging
Benjamin Snyder, Tahira Naseem, Jacob Eisenstein, Regina Barzilay |
EMNLP | 2 |