VLDB 2026 Research / reviewers in the wild / expert
Mahdi Namazifar
dblp:97/4945
· DBLP profile ↗
18ranked-venue papers
5as first author
14since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Split-Merge: Scalable and Memory-Efficient Merging of Expert LLMsabstractWe introduce a zero-shot merging framework for large language models (LLMs) that consolidates specialized domain experts into a single model without any further training.Our core contribution lies in leveraging relative task vectors-difference representations encoding each expert's unique traits with respect to a shared base model-to guide a principled and efficient merging process.By dissecting parameters into common dimensions (averaged across experts) and complementary dimensions (unique to each expert), we strike an optimal balance between generalization and specialization.We further devise a compression mechanism for the complementary parameters, retaining only principal components and scalar multipliers per expert, thereby minimizing overhead.A dynamic router then selects the most relevant domain at inference, ensuring that domain-specific precision is preserved.Experiments on code generation, mathematical reasoning, medical question answering, and instruction-following benchmarks confirm the versatility and effectiveness of our approach.Altogether, this framework enables truly adaptive and scalable LLMs that seamlessly integrate specialized knowledge for improved zero-shot performance. Sruthi Gorantla, Aditya Rawal, Devamanyu Hazarika, Kaixiang Lin, Mingyi Hong 0001, Mahdi Namazifar |
EMNLP | 6 |
| 2025 | Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMsabstractModern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with traditional textbook-style problem solving or factual question answering. Although existing benchmarks have driven progress in areas such as language understanding, code synthesis, and scientific problem solving, real-world engineering design demands the synthesis of domain knowledge, navigation of complex trade-offs, and management of the tedious processes that consume much of practicing engineers' time. Despite these shared challenges across engineering disciplines, no benchmark currently captures the unique demands of engineering design work. In this work, we introduce EngDesign, an Engineering Design benchmark that evaluates LLMs' abilities to perform practical design tasks across nine engineering domains. Unlike existing benchmarks that focus on factual recall or question answering, EngDesign uniquely emphasizes LLMs' ability to synthesize domain knowledge, reason under constraints, and generate functional, objective-oriented engineering designs. Each task in EngDesign represents a real-world engineering design problem, accompanied by a detailed task description specifying design goals, constraints, and performance requirements. EngDesign pioneers a simulation-based evaluation paradigm that moves beyond textbook knowledge to assess genuine engineering design capabilities and shifts evaluation from static answer checking to dynamic, simulation-driven functional verification, marking a crucial step toward realizing the vision of engineering Artificial General Intelligence (AGI). Xingang Guo, Xiangyi Kong, Yilan Jiang, Xiayu Zhao, Zhihua Gong, Daixuan Li, Tianle Sang, Beixiao Zhu, Gregory Jun, Yingbing Huang, Yuqi Xue, Rahul Dev Kundu, Qi Jian Lim, Luke Alexander Granger, Mohamed Badr Younis, Darioush Keivan, Nippun Sabharwal, Shreyanka Sinha, Prakhar Agarwal, Kojo Vandyck, Hanlin Mai, Aditya Venkatesh, Ayush Barik, Jiankun Yang, Chongying Yue, Jingjie He, Licheng Xu, Liujun Xu, Rushabh Shetty, Ziheng Guo, Dahui Song, Manvi Jha, Weijie Liang, Weiman Yan, Bryan Zhang, Sahil Bhandary Karnoor, Rutva Pandya, Xinyi Gong, Mithesh Ballae Ganesh, Feize Shi, Ruiling Xu, Yanfeng Ouyang, Lianhui Qin, Elyse Rosenbaum, Corey Snyder, Peter J. Seiler, Geir E. Dullerud, Xiaojia Shelly Zhang, Zuofu Cheng, Pavan Kumar Hanumolu, Mayank Kulkarni, Mahdi Namazifar, Bin Hu 0002 |
NeurIPS | 63 |
| 2023 | KILM: Knowledge Injection into Encoder-Decoder Language ModelsabstractYan Xu, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yan Xu 0012, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar, Yang Liu 0004, Dilek Hakkani-Tür |
ACL (1) | 2 |
| 2023 | Selective In-Context Data Augmentation for Intent Detection using Pointwise V-InformationabstractYen-Ting Lin, Alexandros Papangelis, Seokhwan Kim, Sungjin Lee, Devamanyu Hazarika, Mahdi Namazifar, Di Jin, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Alexandros Papangelis, Seokhwan Kim, Devamanyu Hazarika, Mahdi Namazifar, Di Jin 0005, Yang Liu 0004, Dilek Hakkani-Tür |
EACL | 6 |
| 2023 | CESAR: Automatic Induction of Compositional Instructions for Multi-turn DialogsabstractInstruction-based multitasking has played a critical role in the success of large language models (LLMs) in multi-turn dialog applications.While publicly-available LLMs have shown promising performance, when exposed to complex instructions with multiple constraints, they lag against state-of-the-art models like Chat-GPT.In this work, we hypothesize that the availability of large-scale complex demonstrations is crucial in bridging this gap.Focusing on dialog applications, we propose a novel framework, CESAR, that unifies a large number of dialog tasks in the same format and allows programmatic induction of complex instructions without any manual effort.We apply CESAR on InstructDial, a benchmark for instruction-based dialog tasks.We further enhance InstructDial with new datasets and tasks and utilize CESAR to induce complex tasks with compositional instructions.This results in a new benchmark called InstructDial++, which includes 63 datasets with 86 basic tasks and 68 composite tasks.Through rigorous experiments, we demonstrate the scalability of CESAR in providing rich instructions.Models trained on InstructDial++ can follow compositional prompts, such as prompts that ask for multiple stylistic constraints. Taha Aksu, Devamanyu Hazarika, Shikib Mehri, Seokhwan Kim, Dilek Hakkani-Tür, Yang Liu 0004, Mahdi Namazifar |
EMNLP | 7 |
| 2023 | Role of Bias Terms in Dot-Product AttentionabstractDot-product attention is a core module in the present generation of neural network models, particularly transformers, and is being leveraged across numerous areas such as natural language processing and computer vision. This attention module is comprised of three linear transformations, namely query, key, and value linear transformations, each of which has a bias term. In this work, we study the role of these bias terms, and mathematically show that the bias term of the key linear transformation is redundant and could be omitted without any impact on the attention module. Moreover, we argue that the bias term of the value linear transformation has a more prominent role than that of the bias term of the query linear transformation. We empirically verify these findings through multiple experiments on language modeling, natural language understanding, and natural language generation tasks. Mahdi Namazifar, Devamanyu Hazarika, Dilek Hakkani-Tür |
ICASSP | 1 |
| 2023 | "What do others think?": Task-Oriented Conversational Modeling with Subjective KnowledgeabstractChao Zhao, Spandana Gella, Seokhwan Kim, Di Jin, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Spandana Gella, Seokhwan Kim, Di Jin 0005, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 8 |
| 2022 | Attention Biasing and Context Augmentation for Zero-Shot Control of Encoder-Decoder Transformers for Natural Language GenerationabstractControlling neural network-based models for natural language generation (NLG) to realize desirable attributes in the generated outputs has broad applications in numerous areas such as machine translation, document summarization, and dialog systems. Approaches that enable such control in a zero-shot manner would be of great importance as, among other reasons, they remove the need for additional annotated data and training. In this work, we propose novel approaches for controlling encoder-decoder transformer-based NLG models in zero shot. While zero-shot control has previously been observed in massive models (e.g., GPT3), our method enables such control for smaller models. This is done by applying two control knobs, attention biasing and context augmentation, to these models directly during decoding and without additional training or auxiliary models. These knobs control the generation process by directly manipulating trained NLG models (e.g., biasing cross-attention layers). We show that not only are these NLG models robust to such manipulations but also their behavior could be controlled without an impact on their generation performance. Devamanyu Hazarika, Mahdi Namazifar, Dilek Hakkani-Tür |
AAAI | 2 |
| 2022 | ALFRED-L: Investigating the Role of Language for Action Learning in Interactive Visual EnvironmentsabstractArjun Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tur. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Arjun R. Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tür |
EMNLP | 4 |
| 2022 | Inducer-tuning: Connecting Prefix-tuning and Adapter-tuningabstractPrefix-tuning, or more generally continuous prompt tuning, has become an essential paradigm of parameter-efficient transfer learning.Using a large pre-trained language model (PLM), prefix-tuning can obtain strong performance by training only a small portion of parameters.In this paper, we propose to understand and further develop prefix-tuning through the kernel lens.Specifically, we make an analogy between prefixes and inducing variables in kernel methods and hypothesize that prefixes serving as inducing variables would improve their overall mechanism.From the kernel estimator perspective, we suggest a new variant of prefix-tuning-inducer-tuning, which shares the exact mechanism as prefix-tuning while leveraging the residual form found in adaptertuning.This mitigates the initialization issue in prefix-tuning.Through comprehensive empirical experiments on natural language understanding and generation tasks, we demonstrate that inducer-tuning can close the performance gap between prefix-tuning and fine-tuning. Yifan Chen 0004, Devamanyu Hazarika, Mahdi Namazifar, Yang Liu 0004, Di Jin 0005, Dilek Hakkani-Tür |
EMNLP | 3 |
| 2022 | Enhancing Knowledge Selection for Grounded Dialogues via Document Semantic GraphsabstractSha Li, Mahdi Namazifar, Di Jin, Mohit Bansal, Heng Ji, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Mahdi Namazifar, Di Jin 0005, Mohit Bansal, Heng Ji 0001, Yang Liu 0004, Dilek Hakkani-Tür |
NAACL-HLT | 2 |
| 2021 | Language Model is all You Need: Natural Language Understanding as Question AnsweringabstractDifferent flavors of transfer learning have shown tremendous impact in advancing research and applications of machine learning. In this work we study the use of a certain family of transfer learning, where the target domain is mapped to the source domain. Specifically we map Natural Language Understanding (NLU) problems to Question Answering (QA) problems and we show that in low data regimes this approach offers significant improvements compared to other approaches to NLU. Moreover, we show that these gains could be increased through sequential transfer learning across NLU problems from different domains. We show that our approach could reduce the amount of required data for the same performance by up to a factor of 10. Mahdi Namazifar, Alexandros Papangelis, Gökhan Tür, Dilek Hakkani-Tür |
ICASSP | 1 |
| 2021 | Correcting Automated and Manual Speech Transcription Errors Using Warped Language ModelsabstractMasked language models have revolutionized natural language processing systems in the past few years. A recently introduced generalization of masked language models called warped language models are trained to be more robust to the types of errors that appear in automatic or manual transcriptions of spoken language by exposing the language model to the same types of errors during training. In this work we propose a novel approach that takes advantage of the robustness of warped language models to transcription noise for correcting transcriptions of spoken language. We show that our proposed approach is able to achieve up to 10% reduction in word error rates of both automatic and manual transcriptions of spoken language. Mahdi Namazifar, John Malik, Li Erran Li, Gökhan Tür, Dilek Hakkani-Tür |
Interspeech | 1 |
| 2021 | Warped Language Models for Noise Robust Language UnderstandingabstractMasked Language Models (MLM) are self-supervised neural networks trained to fill in the blanks in a given sentence with masked tokens. Despite the tremendous success of MLMs for various text based tasks, they are not robust for spoken language understanding, especially for spontaneous conversational speech recognition noise. In this work we introduce Warped Language Models (WLM) in which input sentences at training time go through the same modifications as in MLM, plus two additional modifications, namely inserting and dropping random tokens. These two modifications extend and contract the sentence in addition to the modifications in MLMs, hence the word "warped" in the name. The insertion and drop modification of the input text during training of WLM resemble the types of noise due to Automatic Speech Recognition (ASR) errors, and as a result WLMs are likely to be more robust to ASR noise. Through computational results we show that natural language understanding systems built on top of WLMs perform better compared to those built based on MLMs, especially in the presence of ASR errors. Mahdi Namazifar, Gökhan Tür, Dilek Hakkani-Tür |
SLT | 1 |
| 2020 | Joint Contextual Modeling for ASR Correction and Language UnderstandingabstractThe quality of automatic speech recognition (ASR) is critical to Dialogue Systems as ASR errors propagate to and directly impact downstream tasks such as language understanding (LU). In this paper, we propose multi-task neural approaches to perform contextual language correction on ASR outputs jointly with LU to improve the performance of both tasks simultaneously. To measure the effectiveness of this approach we used a public benchmark, the 2nd Dialogue State Tracking (DSTC2) corpus. As a baseline approach, we trained task specific Statistical Language Models (SLM) and fine-tuned state-of-the-art Generative Pre-training (GPT) Language Model to re-rank the n-best ASR hypotheses, followed by a model to identify the dialog act and slots. i) We further trained ranker models using GPT and Hierarchical CNN-RNN models with discriminatory losses to detect the best output given n-best hypotheses. We extended these ranker models to first select the best ASR output and then identify the dialogue act and slots in an end to end fashion. ii) We also proposed a novel joint ASR error correction and LU model, a word confusion pointer network (WCN-Ptr) with multihead self attention on top, which consumes the word confusions populated from the n-best. We show that the error rates of off the shelf ASR and following LU systems can be reduced significantly by 14% relative with joint models trained using small amounts of in-domain data. Yue Weng, Sai Sumanth Miryala, Chandra Khatri, Huaixiu Zheng, Piero Molino, Mahdi Namazifar, Alexandros Papangelis, Hugh Williams, Franziska Bell, Gökhan Tür |
ICASSP | 7 |
| 2020 | Exploration Based Language Learning for Text-Based GamesabstractThis work presents an exploration and imitation-learning-based agent capable of state-of-the-art performance in playing text-based computer games. These games are of interest as they can be seen as a testbed for language understanding, problem-solving, and language generation by artificial agents. Moreover, they provide a learning setting in which these skills can be acquired through interactions with an environment rather than using fixed corpora. One aspect that makes these games particularly challenging for learning agents is the combinatorially large action space. Existing methods for solving text-based games are limited to games that are either very simple or have an action space restricted to a predetermined set of admissible actions. In this work, we propose to use the exploration approach of Go-Explore for solving text-based games. More specifically, in an initial exploration phase, we first extract trajectories with high rewards, after which we train a policy to solve the game by imitating these trajectories. Our experiments show that this approach outperforms existing solutions in solving text-based games, and it is more sample efficient in terms of the number of interactions with the environment. Moreover, we show that the learned policy can generalize better than existing solutions to unseen games without using any restriction on the action space. Andrea Madotto, Mahdi Namazifar, Joost Huizinga, Piero Molino, Adrien Ecoffet, Huaixiu Zheng, Alexandros Papangelis, Dian Yu 0002, Chandra Khatri, Gökhan Tür |
IJCAI | 2 |
| 2019 | Flexibly-Structured Model for Task-Oriented DialoguesabstractThis paper proposes a novel end-to-end architecture for task-oriented dialogue systems.It is based on a simple and practical yet very effective sequence-to-sequence approach, where language understanding and state tracking tasks are modeled jointly with a structured copy-augmented sequential decoder and a multi-label decoder for each slot.The policy engine and language generation tasks are modeled jointly following that.The copyaugmented sequential decoder deals with new or unknown values in the conversation, while the multi-label decoder combined with the sequential decoder ensures the explicit assignment of values to slots.On the generation part, slot binary classifiers are used to improve performance.This architecture is scalable to real-world scenarios and is shown through an empirical evaluation to achieve state-of-the-art performance on both the Cambridge Restaurant dataset and the Stanford in-car assistant dataset 1 . Lei Shu 0004, Piero Molino, Mahdi Namazifar, Hu Xu 0001, Bing Liu 0001, Huaixiu Zheng, Gökhan Tür |
SIGdial | 3 |
| 2008 | A Parallel Macro Partitioning Framework for Solving Mixed Integer Programs
Mahdi Namazifar, Andrew J. Miller |
CPAIOR | 1 |