EDBT 2026 Demo / reviewers in the wild / expert
Mehdi Rezagholizadeh
dblp:134/0625
· DBLP profile ↗
33ranked-venue papers
2as first author
29since 2021 · last 2025
0000-0003-4014-6007ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 1 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Balcony: A Lightweight Approach to Dynamic Inference of Generative Language ModelsabstractBenyamin Jamialahmadi, Parsa Kavehzadeh, Mehdi Rezagholizadeh, Parsa Farinneya, Hossein Rajabzadeh, Aref Jafari, Boxing Chen, Marzieh S. Tahaei. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Benyamin Jamialahmadi 0001, Parsa Kavehzadeh, Mehdi Rezagholizadeh, Parsa Farinneya, Hossein Rajabzadeh, Aref Jafari, Boxing Chen, Marzieh S. Tahaei |
EMNLP | 3 |
| 2025 | ReGLA: Refining Gated Linear AttentionabstractPeng Lu, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Peng Lu 0006, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais |
NAACL (Long Papers) | 3 |
| 2025 | Zebra-Llama: Towards Extremely Efficient Hybrid ModelsabstractWith the growing demand for deploying large language models (LLMs) across diverse applications, improving their inference efficiency is crucial for sustainable and democratized access. However, retraining LLMs to meet new user-specific requirements is prohibitively expensive and environmentally unsustainable. In this work, we propose a practical and scalable alternative: composing efficient hybrid language models from existing pre-trained models.
Our approach, X-EcoMLA, introduces a family of 1B, 3B, and 8B hybrid models by combining State Space Models (SSMs) and Multi-head Latent Attention (MLA) layers, using a refined initialization and post-training pipeline to efficiently transfer knowledge from pre-trained Transformers.
X-EcoMLA achieves Transformer-level accuracy with near-SSM efficiency using only 7–11 billion training tokens (compared to the trillions required for pre-training) and an 8B teacher. Moreover, it dramatically reduces KV cache size—down to 3.9%, 2%, and 2.73% of the original for the 1B, 3B, and 8B variants, respectively—while preserving 100%, 100%, and over 97% of average zero-shot performance on LM Harness tasks.
Compared to models like MambaInLLaMA, X-EcoMLA, Minitron, and Llamba, our approach consistently delivers competitive or superior accuracy while using significantly fewer tokens, smaller teachers, and vastly reduced KV cache memory. Notably, X-EcoMLA-8B surpasses Minitron-8B in few-shot accuracy by 7%, while using 8× fewer training tokens, over 12× smaller KV cache, and a smaller teacher (8B vs. 15B).
It also achieves 1.4x–3.3x higher throughput (tokens/s) than MambaInLlama. The source code is
released at https://github.com/AMD-AGI/AMD-Hybrid-Models. Mehdi Rezagholizadeh, Guihong Li, Vikram Appia, Emad Barsoum |
NeurIPS | 2 |
| 2024 | EWEK-QA : Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering SystemsabstractMohammad Dehghan, Mohammad Alomrani, Sunyam Bagga, David Alfonso-Hermelo, Khalil Bibi, Abbas Ghaddar, Yingxue Zhang, Xiaoguang Li, Jianye Hao, Qun Liu, Jimmy Lin, Boxing Chen, Prasanna Parthasarathi, Mahdi Biparva, Mehdi Rezagholizadeh. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mohammad Dehghan, Mohammad Ali Alomrani, Sunyam Bagga, David Alfonso-Hermelo, Khalil Bibi, Abbas Ghaddar, Yingxue Zhang 0001, Jianye Hao, Qun Liu 0001, Jimmy Lin, Boxing Chen, Prasanna Parthasarathi, Mahdi Biparva, Mehdi Rezagholizadeh |
ACL (1) | 15 |
| 2024 | Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language ModelsabstractDespite their widespread adoption, large language models (LLMs) remain prohibitive to use under resource constraints, with their ever growing sizes only increasing the barrier for use.One noted issue is the high latency associated with auto-regressive generation, rendering large LLMs use dependent on advanced computing infrastructure.Assisted decoding, where a smaller draft model guides a larger target model's generation, has helped alleviate this, but remains dependent on alignment between the two models.Thus if the draft model is insufficiently capable on some domain relative to the target model, performance can degrade.Alternatively, one can leverage multiple draft models to better cover the expertise of the target, but when multiple black-box draft models are available, selecting an assistant without details about its construction can be difficult.To better understand this decision making problem, we observe it as a contextual bandit, where a policy must choose a draft model based on a context.We show that even without prior knowledge of the draft models, creating an offline dataset from only outputs of independent draft/target models and training a policy over the alignment of these outputs can accelerate performance on multiple domains provided the candidates are effective.Further results show this to hold on various settings with multiple assisted decoding candidates, highlighting its flexibility and the advantageous role that such decision making can play. Jerry Huang, Prasanna Parthasarathi, Mehdi Rezagholizadeh, Sarath Chandar |
EMNLP | 3 |
| 2024 | CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational SearchabstractIn this paper, we study how open-source large language models (LLMs) can be effectively deployed for improving query rewriting in conversational search, especially for ambiguous queries.We introduce CHIQ, a two-step method that leverages the capabilities of LLMs to resolve ambiguities in the conversation history before query rewriting.This approach contrasts with prior studies that predominantly use closed-source LLMs to directly generate search queries from conversation history.We demonstrate on five well-established benchmarks that CHIQ leads to state-of-the-art results across most settings, showing highly competitive performances with systems leveraging closed-source LLMs.Our study provides a first step towards leveraging open-source LLMs in conversational search, as a competitive alternative to the prevailing reliance on commercial LLMs for query rewriting. Fengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholizadeh, Boxing Chen, Qun Liu 0001, Jian-Yun Nie |
EMNLP | 4 |
| 2024 | Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
Mahsa Salmani, Parsa Omidi, Mehdi Rezagholizadeh, Armaghan Eshaghi |
IJCAI | 5 |
| 2024 | CIRAL: A Test Collection for CLIR Evaluations in African LanguagesabstractCross-lingual information retrieval (CLIR) continues to be an actively studied topic in information retrieval (IR), and there have been consistent efforts in curating test collections to support its research. However, there is a lack of high-quality human-annotated CLIR resources for African languages: the few existing collections are mostly curated synthetically or from sources with limited corpora for these languages. We present CIRAL, a test collection for cross-lingual retrieval with English queries and passages in four African languages: Hausa, Somali, Swahili, and Yoruba. CIRAL's corpora are obtained from Indigenous African websites and consist of a total of over 2.5 million passages. We gathered over 1,600 queries and 30k high-quality binary relevance judgments annotated by native speakers of the languages. Additional pools were also obtained at CIRAL's shared task, which was hosted at the Forum for Information Retrieval Evaluation 2023 to encourage community participation in CLIR for African languages. We describe the design and curation process of our test collection and provide reproducible baselines that demonstrate CIRAL's utility in evaluating the effectiveness of systems. CIRAL is available at https://github.com/ciralproject/ciral. Mofe Adeyemi, Akintunde Oladipo, Xinyu Zhang 0018, David Alfonso-Hermelo, Mehdi Rezagholizadeh, Boxing Chen, Abdul-Hakeem Omotayo, Idris Abdulmumin, Naome A. Etori, Toyib Babatunde Musa, Samuel Fanijo, Oluwabusayo Olufunke Awoyomi, Saheed Abdullahi Salahudeen, Labaran Adamu Mohammed, Daud Abolade, Falalu Ibrahim Lawan, Maryam Sabo Abubakar, Ruqayya Nasir Iro, Amina Abubakar Imam, Shafie Abdi Mohamed, Hanad Mohamud Mohamed, Tunde Ajayi, Jimmy Lin |
SIGIR | 5 |
| 2023 | On the utility of enhancing BERT syntactic bias with Token Reordering PretrainingabstractYassir El Mesbahi, Atif Mahmud, Abbas Ghaddar, Mehdi Rezagholizadeh, Phillippe Langlais, Prasanna Parthasarathi. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). 2023. Yassir El Mesbahi, Atif Mahmud, Abbas Ghaddar, Mehdi Rezagholizadeh, Philippe Langlais, Prasanna Parthasarathi |
CoNLL | 4 |
| 2023 | Do we need Label Regularization to Fine-tune Pre-trained Language Models?abstractIvan Kobyzev, Aref Jafari, Mehdi Rezagholizadeh, Tianda Li, Alan Do-Omri, Peng Lu, Pascal Poupart, Ali Ghodsi. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Ivan Kobyzev, Aref Jafari, Mehdi Rezagholizadeh, Tianda Li, Alan Do-Omri, Peng Lu 0006, Pascal Poupart, Ali Ghodsi 0001 |
EACL | 3 |
| 2023 | DyLoRA: Parameter-Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank AdaptationabstractWith the ever-growing size of pretrained models (PMs), fine-tuning them has become more expensive and resource-hungry.As a remedy, low-rank adapters (LoRA) keep the main pretrained weights of the model frozen and just introduce some learnable truncated SVD modules (so-called LoRA blocks) to the model.While LoRA blocks are parameter-efficient, they suffer from two major problems: first, the size of these blocks is fixed and cannot be modified after training (for example, if we need to change the rank of LoRA blocks, then we need to retrain them from scratch); second, optimizing their rank requires an exhaustive search and effort.In this work, we introduce a dynamic low-rank adaptation (DyLoRA) technique to address these two problems together.Our Dy-LoRA method trains LoRA blocks for a range of ranks instead of a single rank by sorting the representation learned by the adapter module at different ranks during training.We evaluate our solution on different natural language understanding (GLUE benchmark) and language generation tasks (E2E, DART and WebNLG) using different pretrained models such as RoBERTa and GPT with different sizes.Our results show that we can train dynamic search-free models with DyLoRA at least 4 to 7 times faster than LoRA without significantly compromising performance.Moreover, our models can perform consistently well on a much larger range of ranks compared to LoRA. 1 Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, Ali Ghodsi 0001 |
EACL | 2 |
| 2023 | Efficient Classification of Long Documents via State-Space ModelsabstractTransformer-based models have achieved stateof-the-art performance on numerous NLP applications.However, long documents which are prevalent in real-world scenarios cannot be efficiently processed by transformers with the vanilla self-attention module due to their quadratic computation complexity and limited length extrapolation ability.Instead of tackling the computation difficulty for self-attention with sparse or hierarchical structures, in this paper, we investigate the use of State-Space Models (SSMs) for long document classification tasks.We conducted extensive experiments on six long document classification datasets, including binary, multi-class, and multi-label classification, comparing SSMs (with and without pre-training) to self-attention-based models.We also introduce the SSM-pooler model and demonstrate that it achieves comparable performance while being on average 36% more efficient.Additionally our method exhibits higher robustness to the input noise even in the extreme scenario of 40%. * Research done during internship in Huawei Noah's Ark Lab (Montreal). Peng Lu 0006, Suyuchen Wang, Mehdi Rezagholizadeh, Bang Liu 0003, Ivan Kobyzev |
EMNLP | 3 |
| 2023 | Robustdistiller: Compressing Universal Speech Representations for Enhanced Environment RobustnessabstractSelf-supervised speech pre-training enables deep neural network models to capture meaningful and disentangled factors from raw waveform signals. The learned universal speech representations can then be used across numerous down-stream tasks. These representations, however, are sensitive to distribution shifts caused by environmental factors, such as noise and/or room reverberation. Their large sizes, in turn, make them unfeasible for edge applications. In this work, we propose a knowledge distillation methodology termed RobustDistiller which compresses universal representations while making them more robust against environmental artifacts via a multi-task learning objective. The proposed layer-wise distillation recipe is evaluated on top of three well-established universal representations, as well as with three downstream tasks. Experimental results show the proposed methodology applied on top of the WavLM Base+ teacher model outperforming all other benchmarks across noise types and levels, as well as reverberation times. Oftentimes, the obtained results with the student model (24M parameters) achieved results inline with those of the teacher model (95M). Heitor R. Guimarães, Arthur Pimentel, Anderson R. Avila, Mehdi Rezagholizadeh, Boxing Chen, Tiago H. Falk |
ICASSP | 4 |
| 2023 | MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse LanguagesabstractAbstract MIRACL is a multilingual dataset for ad hoc retrieval across 18 languages that collectively encompass over three billion native speakers around the world. This resource is designed to support monolingual retrieval tasks, where the queries and the corpora are in the same language. In total, we have gathered over 726k high-quality relevance judgments for 78k queries over Wikipedia in these languages, where all annotations have been performed by native speakers hired by our team. MIRACL covers languages that are both typologically close as well as distant from 10 language families and 13 sub-families, associated with varying amounts of publicly available resources. Extensive automatic heuristic verification and manual assessments were performed during the annotation process to control data quality. In total, MIRACL represents an investment of around five person-years of human annotator effort. Our goal is to spur research on improving retrieval across a continuum of languages, thus enhancing information access capabilities for diverse populations around the world, particularly those that have traditionally been underserved. MIRACL is available at http://miracl.ai/. Xinyu Zhang 0018, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Qun Liu 0001, Mehdi Rezagholizadeh, Jimmy Lin |
Trans. Assoc. Comput. Linguistics | 8 |
| 2022 | From Fully Trained to Fully Random Embeddings: Improving Neural Machine Translation with Compact Word Embedding TablesabstractEmbedding matrices are key components in neural natural language processing (NLP) models that are responsible to provide numerical representations of input tokens (i.e. words or subwords). In this paper, we analyze the impact and utility of such matrices in the context of neural machine translation (NMT). We show that detracting syntactic and semantic information from word embeddings and running NMT systems with random embeddings is not as damaging as it initially sounds. We also show how incorporating only a limited amount of task-specific knowledge from fully-trained embeddings can boost the performance NMT systems. Our findings demonstrate that in exchange for negligible deterioration in performance, any NMT model can be run with partially random embeddings. Working with such structures means a minimal memory requirement as there is no longer need to store large embedding tables, which is a significant gain in industrial and on-device settings. We evaluated our embeddings in translating English into German and French and achieved a 5.3x compression rate. Despite having a considerably smaller architecture, our models in some cases are even able to outperform state-of-the-art baselines. Krtin Kumar, Peyman Passban, Mehdi Rezagholizadeh, Yiu Sing Lau, Qun Liu 0001 |
AAAI | 3 |
| 2022 | CILDA: Contrastive Data Augmentation Using Intermediate Layer Knowledge DistillationabstractKnowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models. Recent years have seen a surge of research aiming to improve KD by leveraging Contrastive Learning, Intermediate Layer Distillation, Data Augmentation, and Adversarial Training. In this work, we propose a learning-based data augmentation technique tailored for knowledge distillation, called CILDA. To the best of our knowledge, this is the first time that intermediate layer representations of the main task are used in improving the quality of augmented samples. More precisely, we introduce an augmentation technique for KD based on intermediate layer matching using contrastive loss to improve masked adversarial data augmentation. CILDA outperforms existing state-of-the-art KD approaches on the GLUE benchmark, as well as in an out-of-domain evaluation. Md. Akmal Haidar, Mehdi Rezagholizadeh, Abbas Ghaddar, Khalil Bibi, Philippe Langlais, Pascal Poupart |
COLING | 2 |
| 2022 | Pro-KD: Progressive Distillation by Following the Footsteps of the TeacherabstractWith the ever growing scale of neural models, knowledge distillation (KD) attracts more attention as a prominent tool for neural model compression. However, there are counter intuitive observations in the literature showing some challenging limitations of KD. A case in point is that the best performing checkpoint of the teacher might not necessarily be the best teacher for training the student in KD. Therefore, one important question would be how to find the best checkpoint of the teacher for distillation? Searching through the checkpoints of the teacher would be a very tedious and computationally expensive process, which we refer to as the checkpoint-search problem. Moreover, another observation is that larger teachers might not necessarily be better teachers in KD, which is referred to as the capacity-gap problem. To address these challenging problems, in this work, we introduce our progressive knowledge distillation (Pro-KD) technique which defines a smoother training path for the student by following the training footprints of the teacher instead of solely relying on distilling from a single mature fully-trained teacher. We demonstrate that our technique is quite effective in mitigating the capacity-gap problem and the checkpoint search problem. We evaluate our technique using a comprehensive set of experiments on different tasks such as image classification (CIFAR-10 and CIFAR-100), natural language understanding tasks of the GLUE benchmark, and question answering (SQuAD 1.1 and 2.0) using BERT-based models and consistently got superior results over state-of-the-art techniques. Mehdi Rezagholizadeh, Aref Jafari, Puneeth S. M. Saladi, Pranav Sharma, Ali Saheb Pasand, Ali Ghodsi 0001 |
COLING | 1 |
| 2022 | Dynamic Position Encoding for TransformersabstractRecurrent models have been dominating the field of neural machine translation (NMT) for the past few years. Transformers have radically changed it by proposing a novel architecture that relies on a feed-forward backbone and self-attention mechanism. Although Transformers are powerful, they could fail to properly encode sequential/positional information due to their non-recurrent nature. To solve this problem, position embeddings are defined exclusively for each time step to enrich word information. However, such embeddings are fixed after training regardless of the task and word ordering system of the source and target languages. In this paper, we address this shortcoming by proposing a novel architecture with new position embeddings that take the order of the target words into consideration. Instead of using predefined position embeddings, our solution generates new embeddings to refine each word’s position information. Since we do not dictate the position of the source tokens and we learn them in an end-to-end fashion, we refer to our method as dynamic position encoding (DPE). We evaluated the impact of our model on multiple datasets to translate from English to German, French, and Italian and observed meaningful improvements in comparison to the original Transformer. Joyce Zheng, Mehdi Rezagholizadeh, Peyman Passban |
COLING | 2 |
| 2022 | Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingabstractAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Yasheng Wang, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Xin Jiang 0002, Qun Liu 0001, Philippe Langlais |
EMNLP | 6 |
| 2022 | KroneckerBERT: Significant Compression of Pre-trained Language Models Through Kronecker Decomposition and Knowledge DistillationabstractMarzieh Tahaei, Ella Charlaix, Vahid Nia, Ali Ghodsi, Mehdi Rezagholizadeh. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Marzieh S. Tahaei, Ella Charlaix, Vahid Partovi Nia, Ali Ghodsi 0001, Mehdi Rezagholizadeh |
NAACL-HLT | 5 |
| 2022 | Learning functions on multiple sets using multi-set transformersabstractWe propose a general deep architecture for learning functions on multiple permutation-invariant sets. We also show how to generalize this architecture to sets of elements of any dimension by dimension equivariance. We demonstrate that our architecture is a universal approximator of these functions, and show superior results to existing methods on a variety of tasks including counting tasks, alignment tasks, distinguishability tasks and statistical distance measurements. This last task is quite important in Machine Learning. Although our approach is quite general, we demonstrate that it can generate approximate estimates of KL divergence and mutual information that are more accurate than previous techniques that are specifically designed to approximate those statistical distances. Kira A. Selby, Ahmad Rashid, Ivan Kobyzev, Mehdi Rezagholizadeh, Pascal Poupart |
UAI | 4 |
| 2021 | ALP-KD: Attention-Based Layer Projection for Knowledge DistillationabstractKnowledge distillation is considered as a training and compression strategy in which two neural networks, namely a teacher and a student, are coupled together during training. The teacher network is supposed to be a trustworthy predictor and the student tries to mimic its predictions. Usually, a student with a lighter architecture is selected so we can achieve compression and yet deliver high-quality results. In such a setting, distillation only happens for final predictions whereas the student could also benefit from teacher’s supervision for internal components. Motivated by this, we studied the problem of distillation for intermediate layers. Since there might not be a one-to-one alignment between student and teacher layers, existing techniques skip some teacher layers and only distill from a subset of them. This shortcoming directly impacts quality, so we instead propose a combinatorial technique which relies on attention. Our model fuses teacher-side information and takes each layer’s significance into consideration, then it performs distillation between combined teacher layers and those of the student. Using our technique, we distilled a 12-layer BERT (Devlin et al. 2019) into 6-, 4-, and 2-layer counterparts and evaluated them on GLUE tasks (Wang et al. 2018). Experimental results show that our combinatorial approach is able to outperform other existing techniques. Peyman Passban, Yimeng Wu, Mehdi Rezagholizadeh, Qun Liu 0001 |
AAAI | 3 |
| 2021 | MATE-KD: Masked Adversarial TExt, a Companion to Knowledge DistillationabstractAhmad Rashid, Vasileios Lioutas, Mehdi Rezagholizadeh. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ahmad Rashid, Vasileios Lioutas, Mehdi Rezagholizadeh |
ACL/IJCNLP (1) | 3 |
| 2021 | Annealing Knowledge DistillationabstractSignificant memory and computational requirements of large deep neural networks restrict their application on edge devices.Knowledge distillation (KD) is a prominent model compression technique for deep neural networks in which the knowledge of a trained large teacher model is transferred to a smaller student model.The success of knowledge distillation is mainly attributed to its training objective function, which exploits the softtarget information (also known as "dark knowledge") besides the given regular hard labels in a training set.However, it is shown in the literature that the larger the gap between the teacher and the student networks, the more difficult is their training using knowledge distillation.To address this shortcoming, we propose an improved knowledge distillation method (called Annealing-KD) by feeding the rich information provided by the teacher's softtargets incrementally and more efficiently.Our Annealing-KD technique is based on a gradual transition over annealed soft-targets generated by the teacher at different temperatures in an iterative process, and therefore, the student is trained to follow the annealed teacher output in a step-by-step manner.This paper includes theoretical and empirical evidence as well as practical experiments to support the effectiveness of our Annealing-KD method.We did a comprehensive set of experiments on different tasks such as image classification (CIFAR-10 and 100) and NLP language inference with BERT-based models on the GLUE benchmark and consistently got superior results. Aref Jafari, Mehdi Rezagholizadeh, Pranav Sharma, Ali Ghodsi 0001 |
EACL | 2 |
| 2021 | Towards Zero-Shot Knowledge Distillation for Natural Language ProcessingabstractKnowledge distillation (KD) is a common knowledge transfer algorithm used for model compression across a variety of deep learning based natural language processing (NLP) solutions.In its regular manifestations, KD requires access to the teacher's training data for knowledge transfer to the student network.However, privacy concerns, data regulations and proprietary reasons may prevent access to such data.We present, to the best of our knowledge, the first work on Zero-shot Knowledge Distillation for NLP, where the student learns from the much larger teacher without any task specific data.Our solution combines out-ofdomain data and adversarial training to learn the teacher's output distribution.We investigate six tasks from the GLUE benchmark and demonstrate that we can achieve between 75% and 92% of the teacher's classification score (accuracy or F1) while compressing the model 30 times.* Work done during an internship at Huawei Noah's Ark Lab. Ahmad Rashid, Vasileios Lioutas, Abbas Ghaddar, Mehdi Rezagholizadeh |
EMNLP (1) | 4 |
| 2021 | Universal-KD: Attention-based Output-Grounded Intermediate Layer Knowledge DistillationabstractIntermediate layer matching is shown as an effective approach for improving knowledge distillation (KD).However, this technique applies matching in the hidden spaces of two different networks (i.e.student and teacher), which lacks clear interpretability.Moreover, intermediate layer KD cannot easily deal with other problems such as layer mapping search and architecture mismatch (i.e. it requires the teacher and student to be of the same model type).To tackle the aforementioned problems all together, we propose Universal-KD to match intermediate layers of the teacher and the student in the output space (by adding pseudo classifiers on intermediate layers) via the attention-based layer projection.By doing this, our unified approach has three merits: (i) it can be flexibly combined with current intermediate layer distillation techniques to improve their results (ii) the pseudo classifiers of the teacher can be deployed instead of extra expensive teacher assistant networks to address the capacity gap problem in KD which is a common issue when the gap between the size of the teacher and student networks becomes too large; (iii) it can be used in cross-architecture intermediate layer KD.We did comprehensive experiments in distilling BERT-base into BERT-4, RoBERTa-large into DistilRoBERTa and BERT-base into CNN and LSTM-based models.Results on the GLUE tasks show that our approach is able to outperform other KD techniques. Yimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md. Akmal Haidar, Ali Ghodsi 0001 |
EMNLP (1) | 2 |
| 2021 | Fine-Tuning of Pre-Trained End-to-End Speech Recognition with Generative Adversarial NetworksabstractAdversarial training of end-to-end (E2E) ASR systems using generative adversarial networks (GAN) has recently been explored for low-resource ASR corpora. GANs help to learn the true data representation through a two-player min-max game. However, training an E2E ASR model using a large ASR corpus with a GAN framework has never been explored, because it might take excessively long time due to high-variance gradient updates and face convergence issues. In this paper, we introduce a novel framework for fine-tuning a pre-trained ASR model using the GAN objective where the ASR model acts as a generator and a discriminator tries to distinguish the ASR output from the real data. Since the ASR model is pre-trained, we hypothesize that the ASR model output (soft distribution vectors) helps to get higher scores from the discriminator and makes the task of the discriminator harder within our GAN framework, which in turn improves the performance of the ASR model in the fine-tuning stage. Here, the pre-trained ASR model is fine-tuned adversarially against the discriminator using an additional adversarial loss. Experiments on full LibriSpeech dataset show that our proposed approach outperforms baselines and conventional GAN-based adversarial models. Md. Akmal Haidar, Mehdi Rezagholizadeh |
ICASSP | 2 |
| 2021 | Transformer-Based ASR Incorporating Time-Reduction Layer and Fine-Tuning with Self-Knowledge DistillationabstractEnd-to-end automatic speech recognition (ASR), unlike conventional ASR, does not have modules to learn the semantic representation from speech encoder. Moreover, the higher frame-rate of speech representation prevents the model to learn the semantic representation properly. Therefore, the models that are constructed by the lower frame-rate of speech encoder lead to better performance. For Transformer-based ASR, the lower frame-rate is not only important for learning better semantic representation but also for reducing the computational complexity due to the self-attention mechanism which has O(n^2) order of complexity in both training and inference. In this paper, we propose a Transformer-based ASR model with the time reduction layer, in which we incorporate time reduction layer inside transformer encoder layers in addition to traditional sub-sampling methods to input features that further reduce the frame-rate. This can help in reducing the computational cost of the self-attention process for training and inference with performance improvement. Moreover, we introduce a fine-tuning approach for pre-trained ASR models using self-knowledge distillation (S-KD) which further improves the performance of our ASR model. Experiments on LibriSpeech datasets show that our proposed methods outperform all other Transformer-based ASR systems. Furthermore, with language model (LM) fusion, we achieve new state-of-the-art word error rate (WER) results for Transformer-based ASR models with just 30 million parameters trained without any external data. Md. Akmal Haidar, Mehdi Rezagholizadeh |
Interspeech | 3 |
| 2021 | Context-aware Adversarial Training for Name Regularity Bias in Named Entity RecognitionabstractAbstract In this work, we examine the ability of NER models to use contextual information when predicting the type of an ambiguous entity. We introduce NRB, a new testbed carefully designed to diagnose Name Regularity Bias of NER models. Our results indicate that all state-of-the-art models we tested show such a bias; BERT fine-tuned models significantly outperforming feature-based (LSTM-CRF) ones on NRB, despite having comparable (sometimes lower) performance on standard benchmarks. To mitigate this bias, we propose a novel model-agnostic training method that adds learnable adversarial noise to some entity mentions, thus enforcing models to focus more strongly on the contextual signal, leading to significant gains on NRB. Combining it with two other training strategies, data augmentation and parameter freezing, leads to further gains. Abbas Ghaddar, Philippe Langlais, Ahmad Rashid, Mehdi Rezagholizadeh |
Trans. Assoc. Comput. Linguistics | 4 |
| 2020 | Why Skip If You Can Combine: A Simple Knowledge Distillation Technique for Intermediate LayersabstractWith the growth of computing power neural machine translation (NMT) models also grow accordingly and become better.However, they also become harder to deploy on edge devices due to memory constraints.To cope with this problem, a common practice is to distill knowledge from a large and accurately-trained teacher network (T ) into a compact student network (S).Although knowledge distillation (KD) is useful in most cases, our study shows that existing KD techniques might not be suitable enough for deep NMT engines, so we propose a novel alternative.In our model, besides matching T and S predictions we have a combinatorial mechanism to inject layer-level supervision from T to S. In this paper, we target low-resource settings and evaluate our translation engines for Portuguese→English, Turkish→English, and English→German directions.Students trained using our technique have 50% fewer parameters and can still deliver comparable results to those of 12-layer teachers. Yimeng Wu, Peyman Passban, Mehdi Rezagholizadeh, Qun Liu 0001 |
EMNLP (1) | 3 |
| 2020 | From Unsupervised Machine Translation to Adversarial Text GenerationabstractWe present a self-attention based bilingual adversarial text generator (B-GAN) which can learn to generate text from the encoder representation of an unsupervised neural machine translation system. B-GAN is able to generate a distributed latent space representation which can be paired with an attention based decoder to generate fluent sentences. When trained on an encoder shared between two languages and paired with the appropriate decoder, it can generate sentences in either language. B-GAN is trained using a combination of reconstruction loss for auto-encoder, a cross domain loss for translation and a GAN based adversarial loss for text generation. We demonstrate that B-GAN, trained on monolingual corpora only using multiple losses, generates more fluent sentences compared to monolingual baselines while effectively using half the number of parameters. Ahmad Rashid, Alan Do-Omri, Md. Akmal Haidar, Qun Liu 0001, Mehdi Rezagholizadeh |
ICASSP | 5 |
| 2019 | EditNTS: An Neural Programmer-Interpreter Model for Sentence Simplification through Explicit EditingabstractWe present the first sentence simplification model that learns explicit edit operations (ADD, DELETE, and KEEP) via a neural programmer-interpreter approach.Most current neural sentence simplification systems are variants of sequence-to-sequence models adopted from machine translation.These methods learn to simplify sentences as a byproduct of the fact that they are trained on complex-simple sentence pairs.By contrast, our neural programmer-interpreter is directly trained to predict explicit edit operations on targeted parts of the input sentence, resembling the way that humans might perform simplification and revision.Our model outperforms previous state-of-the-art neural sentence simplification models (without external knowledge) by large margins on three benchmark text simplification corpora in terms of SARI (+0.95 WikiLarge, +1.89 WikiSmall, +1.41 Newsela), and is judged by humans to produce overall better and simpler output sentences 1 . Yue Dong 0002, Zichao Li 0003, Mehdi Rezagholizadeh, Jackie Chi Kit Cheung |
ACL (1) | 3 |
| 2018 | Reg-Gan: Semi-Supervised Learning Based on Generative Adversarial Networks for RegressionabstractThis research concerns introducing a method to solve the semi-supervised learning problem with generative adversarial networks (GANs) for regression. In contrast to classification, where only a limited number of distinct classes is given, the regression task is defined as predicting continuous labels for a given dataset. This method will be of particular interest for the applications in which a small number of labeled samples is available, and the labels are continuous such as predicting steering angles from the front camera image in the end-to-end task of autonomous driving. Semi-supervised learning is of vital importance for the applications where a small number of labeled samples is available, or labeling samples is difficult or expensive to collect. A case in point is autonomous driving in which obtaining sufficient labeled samples covering all driving conditions is costly. In this context, we can take advantage of semi-supervised learning techniques with groundbreaking generative models, such as generative adversarial networks. However, currently almost all proposed GAN-based semi-supervised techniques in the literature are focused on solving the classification problem. Hence, developing a GAN-based semi-supervised method for the regression task is still an open problem. In this work, two different architectures will be proposed to address this problem. In summary, our introduced method is able to predict continuous labels for a training dataset which has only a limited number of labeled samples. Moreover, the application of this technique for solving the end-to-end task in autonomous driving will be presented. We performed several experiments to evaluate our proposed method, and the results are very promising compared with the state-of-the-art Improved-GAN technique [1]. Mehdi Rezagholizadeh, Md. Akmal Haidar |
ICASSP | 1 |