VLDB 2026 Research / reviewers in the wild / expert
Nadir Durrani
dblp:54/9012
· DBLP profile ↗
47ranked-venue papers
15as first author
20since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 15 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMsabstractArabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern Standard Arabic (MSA), created using Machine Translation (MT) combined with human post-editing. We present AraDiCE, a benchmark for Arabic Dialect and Cultural Evaluation. We evaluate LLMs on dialect comprehension and generation, focusing specifically on low-resource Arabic dialects. Additionally, we introduce the first-ever fine-grained benchmark designed to evaluate cultural awareness across the Gulf, Egypt, and Levant regions, providing a novel dimension to LLM evaluation. Our findings demonstrate that while Arabic-specific models like Jais and AceGPT outperform multilingual models on dialectal tasks, significant challenges persist in dialect identification, generation, and translation. This work contributes ≈45K post-edited samples, a cultural benchmark, and highlights the importance of tailored training to improve LLM performance in capturing the nuances of diverse Arabic dialects and cultural contexts. We have released the dialectal translation models and benchmarks developed in this study (https://huggingface.co/datasets/QCRI/AraDiCE) Basel Mousi, Nadir Durrani, Fatema Ahmad, Md. Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, Firoj Alam |
COLING | 2 |
| 2025 | Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model DiffingabstractAs fine-tuning becomes the dominant paradigm for improving large language models (LLMs), understanding what changes during this process is increasingly important.Traditional benchmarking often fails to explain why one model outperforms another.In this work, we use model diffing, a mechanistic interpretability approach, to analyze the specific capability differences between Gemma-2-9b-it and a SimPO-enhanced variant.Using crosscoders, we identify and categorize latent representations that differentiate the two models.We find that SimPO acquired latent concepts predominantly enhance safety mechanisms (+32.8%),multilingual capabilities (+43.8%), and instruction-following (+151.7%),while its additional training also reduces emphasis on model self-reference (-44.1%) and hallucination management (-68.5%).Our analysis shows that model diffing can yield fine-grained insights beyond leaderboard metrics, attributing performance gaps to concrete mechanistic capabilities.This approach offers a transparent and targeted framework for comparing LLMs. Sabri Boughorbel, Fahim Dalvi, Nadir Durrani, Majd Hawasly |
EMNLP | 3 |
| 2025 | Editing Across Languages: A Survey of Multilingual Knowledge EditingabstractWhile Knowledge Editing has been extensively studied in monolingual settings, it remains underexplored in multilingual contexts.This survey systematizes recent research on Multilingual Knowledge Editing (MKE), a growing subdomain of model editing focused on ensuring factual edits generalize reliably across languages.We present a comprehensive taxonomy of MKE methods, covering parameter-based, memory-based, fine-tuning, and hypernetwork approaches.We survey available benchmarks, summarize key findings on method effectiveness and transfer patterns, and identify persistent challenges such as cross-lingual propagation, language anisotropy, and limited evaluation for low-resource and culturally specific languages.We also discuss broader concerns such as stability and scalability of multilingual edits.Our analysis consolidates a rapidly evolving area and lays the groundwork for future progress in editable language-aware LLMs. Nadir Durrani, Basel Mousi, Fahim Dalvi |
EMNLP | 1 |
| 2025 | From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
Asim Ersoy, Basel Mousi, Shammur Absar Chowdhury, Firoj Alam, Fahim Dalvi, Nadir Durrani |
INTERSPEECH | 6 |
| 2024 | Exploring Alignment in Shared Cross-lingual SpacesabstractDespite their remarkable ability to capture linguistic nuances across diverse languages, questions persist regarding the degree of alignment between languages in multilingual embeddings. Drawing inspiration from research on high-dimensional representations in neural language models, we employ clustering to uncover latent concepts within multilingual models. Our analysis focuses on quantifying the alignment and overlap of these concepts across various languages within the latent space. To this end, we introduce two metrics CALIGN and COLAP aimed at quantifying these aspects, enabling a deeper exploration of multilingual embeddings. Our study encompasses three multilingual models (mT5, mBERT, and XLM-R) and three downstream tasks (Machine Translation, Named Entity Recognition, and Sentiment Analysis). Key findings from our analysis include: i) deeper layers in the network demonstrate increased cross-lingual alignment due to the presence of language-agnostic concepts, ii) fine-tuning of the models enhances alignment within the latent space, and iii) such task-specific calibration helps in explaining the emergence of zero-shot capabilities in the models. Basel Mousi, Nadir Durrani, Fahim Dalvi, Majd Hawasly, Ahmed Abdelali |
ACL (1) | 2 |
| 2024 | LAraBench: Benchmarking Arabic AI with Large Language ModelsabstractAhmed Abdelali, Hamdy Mubarak, Shammur Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Youssef Elshahawy, Ahmed Ali, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ahmed Abdelali, Hamdy Mubarak, Shammur Absar Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Yousseif Elshahawy, Ahmed Ali 0002, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam |
EACL (1) | 15 |
| 2024 | Scaling up Discovery of Latent Concepts in Deep NLP ModelsabstractDespite the revolution caused by deep NLP models, they remain black boxes, necessitating research to understand their decision-making processes.A recent work by Dalvi et al. (2022) carried out representation analysis through the lens of clustering latent spaces within pretrained models (PLMs), but that approach is limited to small scale due to the high cost of running Agglomerative hierarchical clustering.This paper studies clustering algorithms in order to scale the discovery of encoded concepts in PLM representations to larger datasets and models.We propose metrics for assessing the quality of discovered latent concepts and use them to compare the studied clustering algorithms.We found that K-Means-based concept discovery significantly enhances efficiency while maintaining the quality of the obtained concepts.Furthermore, we demonstrate the practicality of this newfound efficiency by scaling latent concept discovery to LLMs and phrasal concepts.1 Majd Hawasly, Fahim Dalvi, Nadir Durrani |
EACL (1) | 3 |
| 2024 | Latent Concept-based Explanation of NLP ModelsabstractInterpreting and understanding the predictions made by deep learning models poses a formidable challenge due to their inherently opaque nature.Many previous efforts to explain these predictions rely on input features, specifically, the words within NLP models.However, such explanations are often less informative due to the discrete nature of the words and their lack of contextual verbosity.To address this limitation, we introduce Latent Concept Attribution (LACOAT), which generates explanations for predictions based on latent concepts.Our intuition is that a word can exhibit multiple facets depending on the context in which it is used.Therefore, given a word in context, the latent space derived from our training process reflects a specific facet of that word.LACOAT functions by mapping the representations of salient input words into the training latent space, enabling it to provide latent contextbased explanations of the prediction. 1 Xuemin Yu, Fahim Dalvi, Nadir Durrani, Marzia Nouri, Hassan Sajjad 0001 |
EMNLP | 3 |
| 2024 | What do end-to-end speech models learn about speaker, language and channel information? A layer-wise and neuron-level analysisabstractDeep neural networks are inherently opaque and challenging to interpret. Unlike hand-crafted feature-based models, we struggle to comprehend the concepts learned and how they interact within these models. This understanding is crucial not only for debugging purposes but also for ensuring fairness in ethical decision-making. In our study, we conduct a post-hoc functional interpretability analysis of pretrained speech models using the probing framework (Hupkes et al., 2018). Specifically, we analyze utterance-level representations of speech models trained for various tasks such as speaker recognition and dialect identification. We conduct layer and neuron-wise analyses, probing for speaker, language, and channel properties. Our study aims to answer the following questions: (i) what information is captured within the representations? (ii) how is it represented and distributed? and (iii) can we identify a minimal subset of the network that possesses this information? Our results reveal several novel findings, including: (i) channel and gender information are distributed across the network, (ii) the information is redundantly available in neurons with respect to a task, (iii) complex properties such as dialectal information are encoded only in the task-oriented pretrained network, (iv) and is localised in the upper layers, (v) we can extract a minimal subset of neurons encoding the pre-defined property, (vi) salient neurons are sometimes shared between properties, (vii) our analysis highlights the presence of biases (for example gender) in the network. Our cross-architectural comparison indicates that: (i) the pretrained models capture speaker-invariant information, and (ii) CNN models are competitive with Transformer models in encoding various understudied properties. Shammur Absar Chowdhury, Nadir Durrani, Ahmed Ali 0002 |
Comput. Speech Lang. | 2 |
| 2023 | ConceptX: A Framework for Latent Concept AnalysisabstractThe opacity of deep neural networks remains a challenge in deploying solutions where explanation is as important as precision. We present ConceptX, a human-in-the-loop framework for interpreting and annotating latent representational space in pre-trained Language Models (pLMs). We use an unsupervised method to discover concepts learned in these models and enable a graphical interface for humans to generate explanations for the concepts. To facilitate the process, we provide auto-annotations of the concepts (based on traditional linguistic ontologies). Such annotations enable development of a linguistic resource that directly represents latent concepts learned within deep NLP models. These include not just traditional linguistic concepts, but also task-specific or sensitive concepts (words grouped based on gender or religious connotation) that helps the annotators to mark bias in the model. The framework consists of two parts (i) concept discovery and (ii) annotation platform. Firoj Alam, Fahim Dalvi, Nadir Durrani, Hassan Sajjad 0001, Abdul Rafae Khan, Jia Xu 0004 |
AAAI | 3 |
| 2023 | Can LLMs Facilitate Interpretation of Pre-trained Language Models?abstractWork done to uncover the knowledge encoded within pre-trained language models rely on annotated corpora or human-in-the-loop methods.However, these approaches are limited in terms of scalability and the scope of interpretation.We propose using a large language model, Chat-GPT, as an annotator to enable fine-grained interpretation analysis of pre-trained language models.We discover latent concepts within pre-trained language models by applying agglomerative hierarchical clustering over contextualized representations and then annotate these concepts using ChatGPT.Our findings demonstrate that ChatGPT produces accurate and semantically richer annotations compared to human-annotated concepts.Additionally, we showcase how GPT-based annotations empower interpretation analysis methodologies of which we demonstrate two: probing frameworks and neuron interpretation.To facilitate further exploration and experimentation in the field, we make available a substantial Concept-Net dataset (TCN) comprising 39,000 annotated concepts.1 Basel Mousi, Nadir Durrani, Fahim Dalvi |
EMNLP | 2 |
| 2023 | Evaluating Neuron Interpretation Methods of NLP ModelsabstractNeuron interpretation offers valuable insights into how knowledge is structured within a deep neural network model. While a number of neuron interpretation methods have been proposed in the literature, the field lacks a comprehensive comparison among these methods. This gap hampers progress due to the absence of standardized metrics and benchmarks. The commonly used evaluation metric has limitations, and creating ground truth annotations for neurons is impractical. Addressing these challenges, we propose an evaluation framework based on voting theory. Our hypothesis posits that neurons consistently identified by different methods carry more significant information. We rigorously assess our framework across a diverse array of neuron interpretation methods. Notable findings include: i) despite the theoretical differences among the methods, neuron ranking methods share over 60% of their rankings when identifying salient neurons, ii) the neuron interpretation methods are most sensitive to the last layer representations, iii) Probeless neuron ranking emerges as the most consistent method. Yimin Fan, Fahim Dalvi, Nadir Durrani, Hassan Sajjad 0001 |
NeurIPS | 3 |
| 2023 | On the effect of dropping layers of pre-trained transformer models
Hassan Sajjad 0001, Fahim Dalvi, Nadir Durrani, Preslav Nakov |
Comput. Speech Lang. | 3 |
| 2023 | Discovering Salient Neurons in deep NLP modelsabstractWhile a lot of work has been done in understanding representations learned within deep NLP models and what knowledge they capture, work done towards analyzing individual neurons is relatively sparse. We present a technique called Linguistic Correlation Analysis to extract salient neurons in the model, with respect to any extrinsic property, with the goal of understanding how such knowledge is preserved within neurons. We carry out a fine-grained analysis to answer the following questions: (i) can we identify subsets of neurons in the network that learn a specific linguistic property? (ii) is a certain linguistic phenomenon in a given model localized (encoded in few individual neurons) or distributed across many neurons? (iii) how redundantly is the information preserved? (iv) how does fine-tuning pre-trained models towards downstream NLP tasks impact the learned linguistic knowledge? (v) how do models vary in learning different linguistic properties? Our data-driven, quantitative analysis illuminates interesting findings: (i) we found small subsets of neurons that can predict different linguistic tasks; (ii) neurons capturing basic lexical information, such as suffixation, are localized in the lowermost layers; (iii) neurons learning complex concepts, such as syntactic role, are predominantly found in middle and higher layers; (iv) salient linguistic neurons are relocated from higher to lower layers during transfer learning, as the network preserves the higher layers for task-specific information; (v) we found interesting differences across pre-trained models regarding how linguistic information is preserved within them; and (vi) we found that concepts exhibit similar neuron distribution across different languages in the multilingual transformer models. Our code is publicly available as part of the NeuroX toolkit (Dalvi et al., 2023). Nadir Durrani, Fahim Dalvi, Hassan Sajjad 0001 |
J. Mach. Learn. Res. | 1 |
| 2022 | Effect of Post-processing on Contextualized Word RepresentationsabstractPost-processing of static embedding has been shown to improve their performance on both lexical and sequence-level tasks. However, post-processing for contextualized embeddings is an under-studied problem. In this work, we question the usefulness of post-processing for contextualized embeddings obtained from different layers of pre-trained language models. More specifically, we standardize individual neuron activations using z-score, min-max normalization, and by removing top principal components using the all-but-the-top method. Additionally, we apply unit length normalization to word representations. On a diverse set of pre-trained models, we show that post-processing unwraps vital information present in the representations for both lexical tasks (such as word similarity and analogy) and sequence classification tasks. Our findings raise interesting points in relation to the research studies that use contextualized representations, and suggest z-score normalization as an essential step to consider when using them in an application. Hassan Sajjad 0001, Firoj Alam, Fahim Dalvi, Nadir Durrani |
COLING | 4 |
| 2022 | On the Transformation of Latent Space in Fine-Tuned NLP ModelsabstractWe study the evolution of latent space in finetuned NLP models.Different from the commonly used probing-framework, we opt for an unsupervised method to analyze representations.More specifically, we discover latent concepts in the representational space using hierarchical clustering.We then use an alignment function to gauge the similarity between the latent space of a pre-trained model and its finetuned version.We use traditional linguistic concepts to facilitate our understanding and also study how the model space transforms towards task-specific information.We perform a thorough analysis, comparing pre-trained and finetuned models across three models and three downstream tasks.The notable findings of our work are: i) the latent space of the higher layers evolve towards task-specific concepts, ii) whereas the lower layers retain generic concepts acquired in the pre-trained model, iii) we discovered that some concepts in the higher layers acquire polarity towards the output class, and iv) that these concepts can be used for generating adversarial triggers. Nadir Durrani, Hassan Sajjad 0001, Fahim Dalvi, Firoj Alam |
EMNLP | 1 |
| 2022 | Discovering Latent Concepts Learned in BERT
Fahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani, Jia Xu 0004, Hassan Sajjad 0001 |
ICLR | 4 |
| 2022 | Analyzing Encoded Concepts in Transformer Language ModelsabstractHassan Sajjad, Nadir Durrani, Fahim Dalvi, Firoj Alam, Abdul Khan, Jia Xu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Hassan Sajjad 0001, Nadir Durrani, Fahim Dalvi, Firoj Alam, Abdul Rafae Khan, Jia Xu 0004 |
NAACL-HLT | 2 |
| 2022 | Neuron-level Interpretation of Deep NLP Models: A SurveyabstractAbstract The proliferation of Deep Neural Networks in various domains has seen an increased need for interpretability of these models. Preliminary work done along this line, and papers that surveyed such, are focused on high-level representation analysis. However, a recent branch of work has concentrated on interpretability at a more granular level of analyzing neurons within these models. In this paper, we survey the work done on neuron analysis including: i) methods to discover and understand neurons in a network; ii) evaluation methods; iii) major findings including cross architectural comparisons that neuron analysis has unraveled; iv) applications of neuron probing such as: controlling the model, domain adaptation, and so forth; and v) a discussion on open issues and future research directions. Hassan Sajjad 0001, Nadir Durrani, Fahim Dalvi |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | Fighting the COVID-19 Infodemic in Social Media: A Holistic Perspective and a Call to Arms
Firoj Alam, Fahim Dalvi, Shaden Shaar, Nadir Durrani, Hamdy Mubarak, Alex Nikolov, Giovanni Da San Martino, Ahmed Abdelali, Hassan Sajjad 0001, Kareem Darwish, Preslav Nakov |
ICWSM | 4 |
| 2020 | Similarity Analysis of Contextual Word Representation ModelsabstractThis paper investigates contextual word representation models from the lens of similarity analysis.Given a collection of trained models, we measure the similarity of their internal representations and attention.Critically, these models come from vastly different architectures.We use existing and novel similarity measures that aim to gauge the level of localization of information in the deep models, and facilitate the investigation of which design factors affect model similarity, without requiring any external linguistic annotation.The analysis reveals that models within the same family are more similar to one another, as may be expected.Surprisingly, different architectures have rather similar representations, but different individual neurons.We also observed differences in information localization in lower and higher layers and found that higher layers are more affected by fine-tuning on downstream tasks. 1 John M. Wu, Yonatan Belinkov, Hassan Sajjad 0001, Nadir Durrani, Fahim Dalvi, James R. Glass |
ACL | 4 |
| 2020 | AraBench: Benchmarking Dialectal Arabic-English Machine TranslationabstractLow-resource machine translation suffers from the scarcity of training data and the unavailability of standard evaluation sets. While a number of research efforts target the former, the unavailability of evaluation benchmarks remain a major hindrance in tracking the progress in low-resource machine translation. In this paper, we introduce AraBench, an evaluation suite for dialectal Arabic to English machine translation. Compared to Modern Standard Arabic, Arabic dialects are challenging due to their spoken nature, non-standard orthography, and a large variation in dialectness. To this end, we pool together already available Dialectal Arabic-English resources and additionally build novel test sets. AraBench offers 4 coarse, 15 fine-grained and 25 city-level dialect categories, belonging to diverse genres, such as media, chat, religion and travel with varying level of dialectness. We report strong baselines using several training settings: fine-tuning, back-translation and data augmentation. The evaluation suite opens a wide range of research frontiers to push efforts in low-resource machine translation, particularly Arabic dialect translation. The evaluation suite and the dialectal system are publicly available for research purposes. Hassan Sajjad 0001, Ahmed Abdelali, Nadir Durrani, Fahim Dalvi |
COLING | 3 |
| 2020 | Analyzing Redundancy in Pretrained Transformer ModelsabstractTransformer-based deep NLP models are trained using hundreds of millions of parameters, limiting their applicability in computationally constrained environments.In this paper, we study the cause of these limitations by defining a notion of Redundancy, which we categorize into two classes: General Redundancy and Task-specific Redundancy.We dissect two popular pretrained models, BERT and XLNet, studying how much redundancy they exhibit at a representation-level and at a more fine-grained neuron-level.Our analysis reveals interesting insights, such as: i) 85% of the neurons across the network are redundant and ii) at least 92% of them can be removed when optimizing towards a downstream task.Based on our analysis, we present an efficient feature-based transfer learning procedure, which maintains 97% performance while using at-most 10% of the original neurons. 1 Fahim Dalvi, Hassan Sajjad 0001, Nadir Durrani, Yonatan Belinkov |
EMNLP (1) | 3 |
| 2020 | Analyzing Individual Neurons in Pre-trained Language ModelsabstractWhile a lot of analysis has been carried to demonstrate linguistic knowledge captured by the representations learned within deep NLP models, very little attention has been paid towards individual neurons.We carry out a neuron-level analysis using core linguistic tasks of predicting morphology, syntax and semantics, on pre-trained language models, with questions like: i) do individual neurons in pretrained models capture linguistic information?ii) which parts of the network learn more about certain linguistic phenomena?iii) how distributed or focused is the information?and iv) how do various architectures differ in learning these properties?We found small subsets of neurons to predict linguistic tasks, with lower level tasks (such as morphology) localized in fewer neurons, compared to higher level task of predicting syntax.Our study reveals interesting cross architectural comparisons.For example, we found neurons in XLNet to be more localized and disjoint when predicting properties compared to BERT and others, where they are more distributed and coupled. Nadir Durrani, Hassan Sajjad 0001, Fahim Dalvi, Yonatan Belinkov |
EMNLP (1) | 1 |
| 2020 | On the Linguistic Representational Power of Neural Machine Translation ModelsabstractDespite the recent success of deep neural networks in natural language processing and other spheres of artificial intelligence, their interpretability remains a challenge. We analyze the representations learned by neural machine translation (NMT) models at various levels of granularity and evaluate their quality through relevant extrinsic properties. In particular, we seek answers to the following questions: (i) How accurately is word structure captured within the learned representations, which is an important aspect in translating morphologically rich languages? (ii) Do the representations capture long-range dependencies, and effectively handle syntactically divergent languages? (iii) Do the representations capture lexical semantics? We conduct a thorough investigation along several parameters: (i) Which layers in the architecture capture each of these linguistic phenomena; (ii) How does the choice of translation unit (word, character, or subword unit) impact the linguistic properties captured by the underlying representations? (iii) Do the encoder and decoder learn differently and independently? (iv) Do the representations learned by multilingual NMT models capture the same amount of linguistic information as their bilingual counterparts? Our data-driven, quantitative evaluation illuminates important aspects in NMT models and their ability to capture various linguistic phenomena. We show that deep NMT models trained in an end-to-end fashion, without being provided any direct supervision during the training process, learn a non-trivial amount of linguistic information. Notable findings include the following observations: (i) Word morphology and part-of-speech information are captured at the lower layers of the model; (ii) In contrast, lexical semantics or non-local syntactic and semantic dependencies are better represented at the higher layers of the model; (iii) Representations learned using characters are more informed about word-morphology compared to those learned using subword units; and (iv) Representations learned by multilingual models are richer compared to bilingual models. Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad 0001, James R. Glass |
Comput. Linguistics | 2 |
| 2019 | What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP ModelsabstractDespite the remarkable evolution of deep neural networks in natural language processing (NLP), their interpretability remains a challenge. Previous work largely focused on what these models learn at the representation level. We break this analysis down further and study individual dimensions (neurons) in the vector representation learned by end-to-end neural models in NLP tasks. We propose two methods: Linguistic Correlation Analysis, based on a supervised method to extract the most relevant neurons with respect to an extrinsic task, and Cross-model Correlation Analysis, an unsupervised method to extract salient neurons w.r.t. the model itself. We evaluate the effectiveness of our techniques by ablating the identified neurons and reevaluating the network’s performance for two tasks: neural machine translation (NMT) and neural language modeling (NLM). We further present a comprehensive analysis of neurons with the aim to address the following questions: i) how localized or distributed are different linguistic properties in the models? ii) are certain neurons exclusive to some properties and not others? iii) is the information more or less distributed in NMT vs. NLM? and iv) how important are the neurons identified through the linguistic correlation method to the overall task? Our code is publicly available as part of the NeuroX toolkit (Dalvi et al. 2019a). This paper is a non-archived version of the paper published at AAAI (Dalvi et al. 2019b). Fahim Dalvi, Nadir Durrani, Hassan Sajjad 0001, Yonatan Belinkov, Anthony Bau, James R. Glass |
AAAI | 2 |
| 2019 | NeuroX: A Toolkit for Analyzing Individual Neurons in Neural NetworksabstractWe present a toolkit to facilitate the interpretation and understanding of neural network models. The toolkit provides several methods to identify salient neurons with respect to the model itself or an external task. A user can visualize selected neurons, ablate them to measure their effect on the model accuracy, and manipulate them to control the behavior of the model at the test time. Such an analysis has a potential to serve as a springboard in various research directions, such as understanding the model, better architectural choices, model distillation and controlling data biases. The toolkit is available for download.1 Fahim Dalvi, Avery Nortonsmith, Anthony Bau, Yonatan Belinkov, Hassan Sajjad 0001, Nadir Durrani, James R. Glass |
AAAI | 6 |
| 2019 | Identifying and Controlling Important Neurons in Neural Machine Translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad 0001, Nadir Durrani, Fahim Dalvi, James R. Glass |
ICLR (Poster) | 4 |
| 2017 | What do Neural Machine Translation Models Learn about Morphology?abstractNeural machine translation (MT) models obtain state-of-the-art performance while maintaining a simple, end-to-end architecture.However, little is known about what these models learn about source and target languages during the training process.In this work, we analyze the representations learned by neural MT models at various levels of granularity and empirically evaluate the quality of the representations for learning morphology through extrinsic part-of-speech and morphological tagging tasks.We conduct a thorough investigation along several parameters: word-based vs. character-based representations, depth of the encoding layer, the identity of the target language, and encoder vs. decoder representations.Our data-driven, quantitative evaluation sheds light on important aspects in the neural MT system and its ability to capture word structure.1 Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad 0001, James R. Glass |
ACL (1) | 2 |
| 2017 | Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging TasksabstractWhile neural machine translation (NMT) models provide improved translation quality in an elegant framework, it is less clear what they learn about language. Recent work has started evaluating the quality of vector representations learned by NMT models on morphological and syntactic tasks. In this paper, we investigate the representations learned at different layers of NMT encoders. We train NMT systems on parallel data and use the models to extract features for training a classifier on two tasks: part-of-speech and semantic tagging. We then measure the performance of the classifier as a proxy to the quality of the original NMT model for the given task. Our quantitative analysis yields interesting insights regarding representation learning in NMT models. For instance, we find that higher layers are better at learning semantics while lower layers tend to be better for part-of-speech tagging. We also observe little effect of the target language on source-side representations, especially in higher quality models. Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad 0001, Nadir Durrani, Fahim Dalvi, James R. Glass |
IJCNLP(1) | 4 |
| 2017 | Understanding and Improving Morphological Learning in the Neural Machine Translation DecoderabstractEnd-to-end training makes the neural machine translation (NMT) architecture simpler, yet elegant compared to traditional statistical machine translation (SMT). However, little is known about linguistic patterns of morphology, syntax and semantics learned during the training of NMT systems, and more importantly, which parts of the architecture are responsible for learning each of these phenomenon. In this paper we i) analyze how much morphology an NMT decoder learns, and ii) investigate whether injecting target morphology in the decoder helps it to produce better translations. To this end we present three methods: i) simultaneous translation, ii) joint-data learning, and iii) multi-task learning. Our results show that explicit morphological information helps the decoder learn target language morphology and improves the translation quality by 0.2–0.6 BLEU points. Fahim Dalvi, Nadir Durrani, Hassan Sajjad 0001, Yonatan Belinkov, Stephan Vogel |
IJCNLP(1) | 2 |
| 2017 | Domain adaptation using neural network joint model
Shafiq R. Joty, Nadir Durrani, Hassan Sajjad 0001, Ahmed Abdelali |
Comput. Speech Lang. | 2 |
| 2016 | Enabling Medical Translation for Low-Resource Languages
Ahmad Musleh, Nadir Durrani, Irina P. Temnikova, Preslav Nakov, Stephan Vogel, Osama Alsaad |
CICLing (2) | 2 |
| 2016 | A Deep Fusion Model for Domain Adaptation in Phrase-based MTabstractWe present a novel fusion model for domain adaptation in Statistical Machine Translation. Our model is based on the joint source-target neural network Devlin et al., 2014, and is learned by fusing in- and out-domain models. The adaptation is performed by backpropagating errors from the output layer to the word embedding layer of each model, subsequently adjusting parameters of the composite model towards the in-domain data. On the standard tasks of translating English-to-German and Arabic-to-English TED talks, we observed average improvements of +0.9 and +0.7 BLEU points, respectively over a competition grade phrase-based system. We also demonstrate improvements over existing adaptation methods. Nadir Durrani, Hassan Sajjad 0001, Shafiq R. Joty, Ahmed Abdelali |
COLING | 1 |
| 2016 | Eyes Don't Lie: Predicting Machine Translation Quality Using Eye MovementabstractHassan Sajjad, Francisco Guzmán, Nadir Durrani, Ahmed Abdelali, Houda Bouamor, Irina Temnikova, Stephan Vogel. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Hassan Sajjad 0001, Francisco Guzmán, Nadir Durrani, Ahmed Abdelali, Houda Bouamor, Irina P. Temnikova, Stephan Vogel |
HLT-NAACL | 3 |
| 2015 | How to Avoid Unwanted Pregnancies: Domain Adaptation using Neural Network ModelsabstractWe present novel models for domain adaptation based on the neural network joint model (NNJM).Our models maximize the cross entropy by regularizing the loss function with respect to in-domain model.Domain adaptation is carried out by assigning higher weight to out-domain sequences that are similar to the in-domain data.In our alternative model we take a more restrictive approach by additionally penalizing sequences similar to the outdomain data.Our models achieve better perplexities than the baseline NNJM models and give improvements of up to 0.5 and 0.6 BLEU points in Arabic-to-English and English-to-German language pairs, on a standard task of translating TED talks. Shafiq R. Joty, Hassan Sajjad 0001, Nadir Durrani, Kamla Al-Mannai, Ahmed Abdelali, Stephan Vogel |
EMNLP | 3 |
| 2015 | Using joint models or domain adaptation in statistical machine translation
Nadir Durrani, Hassan Sajjad 0001, Shafiq R. Joty, Ahmed Abdelali, Stephan Vogel |
MTSummit | 1 |
| 2015 | The Operation Sequence Model - Combining N-Gram-Based and Phrase-Based Statistical Machine TranslationabstractIn this article, we present a novel machine translation model, the Operation Sequence Model (OSM), which combines the benefits of phrase-based and N-gram-based statistical machine translation (SMT) and remedies their drawbacks. The model represents the translation process as a linear sequence of operations. The sequence includes not only translation operations but also reordering operations. As in N-gram-based SMT, the model is: (i) based on minimal translation units, (ii) takes both source and target information into account, (iii) does not make a phrasal independence assumption, and (iv) avoids the spurious phrasal segmentation problem. As in phrase-based SMT, the model (i) has the ability to memorize lexical reordering triggers, (ii) builds the search graph dynamically, and (iii) decodes with large translation units during search. The unique properties of the model are (i) its strong coupling of reordering and translation where translation and reordering decisions are conditioned on n previous translation and reordering decisions, and (ii) the ability to model local and long-range reorderings consistently. Using BLEU as a metric of translation accuracy, we found that our system performs significantly better than state-of-the-art phrase-based systems (Moses and Phrasal) and N-gram-based systems (Ncode) on standard translation tasks. We compare the reordering component of the OSM to the Moses lexical reordering model by integrating it into Moses. Our results show that OSM outperforms lexicalized reordering on all translation tasks. The translation quality is shown to be improved further by learning generalized representations with a POS-based OSM. Nadir Durrani, Helmut Schmid, Alexander Fraser 0001, Philipp Koehn, Hinrich Schütze |
Comput. Linguistics | 1 |
| 2014 | Improving Egyptian-to-English SMT by Mapping Egyptian into MSA
Nadir Durrani, Yaser Al-Onaizan, Abraham Ittycheriah |
CICLing (2) | 1 |
| 2014 | Investigating the Usefulness of Generalized Word Representations in SMT
Nadir Durrani, Philipp Koehn, Helmut Schmid, Alexander Fraser 0001 |
COLING | 1 |
| 2014 | Integrating an Unsupervised Transliteration Model into Statistical Machine TranslationabstractWe investigate three methods for integrating an unsupervised transliteration model into an end-to-end SMT system.We induce a transliteration model from parallel data and use it to translate OOV words.Our approach is fully unsupervised and language independent.In the methods to integrate transliterations, we observed improvements from 0.23-0.75(∆ 0.41) BLEU points across 7 language pairs.We also show that our mined transliteration corpora provide better rule coverage and translation quality compared to the gold standard transliteration corpora. Nadir Durrani, Hassan Sajjad 0001, Hieu Hoang, Philipp Koehn |
EACL | 1 |
| 2014 | Improving machine translation via triangulation and transliteration
Nadir Durrani, Philipp Koehn |
EAMT | 1 |
| 2013 | Model With Minimal Translation Units, But Decode With Phrases
Nadir Durrani, Alexander Fraser 0001, Helmut Schmid |
HLT-NAACL | 1 |
| 2011 | A Joint Sequence Translation Model with Integrated Reordering
Nadir Durrani, Helmut Schmid, Alexander Fraser 0001 |
ACL | 1 |
| 2011 | Comparing Two Techniques for Learning Transliteration Models Using a Parallel Corpus
Hassan Sajjad 0001, Nadir Durrani, Helmut Schmid, Alexander Fraser 0001 |
IJCNLP | 2 |
| 2010 | Hindi-to-Urdu Machine Translation through Transliteration
Nadir Durrani, Hassan Sajjad 0001, Alexander Fraser 0001, Helmut Schmid |
ACL | 1 |
| 2010 | Urdu Word Segmentation
Nadir Durrani, Sarmad Hussain |
HLT-NAACL | 1 |