Saiful Bari

dblp:241/6155 · also M. Saiful Bari · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-8795-8040ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Models
abstract
Alhanoof Althnian, Norah A. Alzahrani, Shaykhah Z. Alsubaie, Eman Albilali, Ahmed Abdelali, Nouf M. Alotaibi, M Saiful Bari, Yazeed Alnumay, Abdulhamed Alothaimen, Maryam Saif, Shahad D. Alzaidi, Faisal Abdulrahman Mirza, Yousef Almushayqih, Mohammed Al Saleem, Ghadah Alabduljabbar, Abdulmohsen Al-Thubaity, Areeb Alowisheq, Nora Al-Twairesh. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Alhanoof Althnian, Norah A. Alzahrani, Shaykhah Alsubaie, Eman Albilali, Ahmed Abdelali, Nouf M. Alotaibi, Saiful Bari, Yazeed Alnumay, Abdulhamed Alothaimen, Maryam Saif, Shahad D. Alzaidi, Faisal Mirza, Yousef Almushayqih, Mohammed Al Saleem, Ghadah Alabduljabbar, AbdulMohsen Al-Thubaity, Areeb Alowisheq, Nora Al-Twairesh
EMNLP7
2025 ALLaM: Large Language Models for Arabic and English
abstract
In this work, we present ALLaM: Arabic Large Language Model, a series of large language models to support the ecosystem of Arabic Language Technologies (ALT). ALLaM is carefully trained, considering the values of language alignment and transferability of knowledge at scale. The models are based on an autoregressive decoder-only architecture and are pretrained on a mixture of Arabic and English texts. We illustrate how the second-language acquisition via vocabulary expansion can help steer a language model towards a new language without any major catastrophic forgetting in English. Furthermore, we highlight the effectiveness of using translation data and the process of knowledge encoding within the language model's latent space. Finally, we show that effective alignment with human preferences can significantly enhance the performance of a large language model (LLM) compared to less aligned models of a larger scale. Our methodology enables us to achieve state-of-the-art performance in various Arabic benchmarks, including MMLU Arabic, ACVA, and Arabic Exams. Our aligned models improve both in Arabic and English from its base aligned models.
Saiful Bari, Yazeed Alnumay, Norah A. Alzahrani, Nouf M. Alotaibi, Hisham Abdullah Alyahya, Sultan Alrashed, Faisal Mirza, Shaykhah Alsubaie, Hassan A. Alahmed, Ghadah Alabduljabbar, Raghad Alkhathran, Yousef Almushayqih, Raneem Alnajim, Salman Alsubaihi, Maryam Al Mansour, Saad Amin Hassan, Majed Alrubaian, Ali Alammari, Zaki Alawami, AbdulMohsen Al-Thubaity
ICLR1
2024 When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
abstract
Norah Alzahrani, Hisham Alyahya, Yazeed Alnumay, Sultan AlRashed, Shaykhah Alsubaie, Yousef Almushayqih, Faisal Mirza, Nouf Alotaibi, Nora Al-Twairesh, Areeb Alowisheq, M Saiful Bari, Haidar Khan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Norah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yousef Almushayqih, Faisal Mirza, Nouf Alotaibi, Nora Al-Twairesh, Areeb Alowisheq, Saiful Bari, Haidar Khan
ACL (1)11
2024 XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and Retrieval
abstract
Mohammad Abdullah Matin Khan, M Saiful Bari, Xuan Long Do, Weishi Wang, Md Rizwan Parvez, Shafiq Joty. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Mohammad Abdullah Matin Khan, Saiful Bari, Xuan Do Long, Weishi Wang, Md. Rizwan Parvez, Shafiq R. Joty
ACL (1)2
2024 BenLLM-Eval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP
abstract
Large Language Models (LLMs) have emerged as one of the most important breakthroughs in natural language processing (NLP) for their impressive skills in language generation and other language-specific tasks. Though LLMs have been evaluated in various tasks, mostly in English, they have not yet undergone thorough evaluation in under-resourced languages such as Bengali (Bangla). To this end, this paper introduces BenLLM-Eval, which consists of a comprehensive evaluation of LLMs to benchmark their performance in the low-resourced Bangla language. In this regard, we select various important and diverse Bangla NLP tasks, such as text summarization, question answering, paraphrasing, natural language inference, text classification, and sentiment analysis for zero-shot evaluation of popular LLMs, namely, ChatGPT, LLaMA-2, and Claude-2. Our experimental results demonstrate that while in some Bangla NLP tasks, zero-shot LLMs could achieve performance on par, or even better than current SOTA fine-tuned models; in most tasks, their performance is quite poor (with the performance of open-source LLMs like LLaMA-2 being significantly bad) in comparison to the current SOTA results. Therefore, it calls for further efforts to develop a better understanding of LLMs in low-resource languages like Bangla.
Mohsinul Kabir, Mohammed Saidul Islam, Md. Tahmid Rahman Laskar, Mir Tafseer Nayeem, Saiful Bari, Enamul Hoque Prince
LREC/COLING5
2024 A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
abstract
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty, Jimmy Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Md. Tahmid Rahman Laskar, Sawsan Alqahtani, Saiful Bari, Mohammad Abdullah Matin Khan, Haidar Khan, Amran Bhuiyan, Chee-Wei Tan 0001, Md. Rizwan Parvez, Enamul Hoque Prince, Shafiq R. Joty, Jimmy Huang 0001
EMNLP3
2023 Crosslingual Generalization through Multitask Finetuning
abstract
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, Colin Raffel. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, Saiful Bari, Sheng Shen 0001, Hailey Schoelkopf, Xiangru Tang, Dragomir R. Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, Colin Raffel
ACL (1)7
2023 BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting
abstract
Zheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Winata, Stella Biderman, Edward Raff, Dragomir Radev, Vassilina Nikoulina. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Indra Winata, Stella Biderman, Edward Raff, Dragomir R. Radev, Vassilina Nikoulina
ACL (1)7
2021 UXLA: A Robust Unsupervised Data Augmentation Framework for Zero-Resource Cross-Lingual NLP
abstract
M Saiful Bari, Tasnim Mohiuddin, Shafiq Joty. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Saiful Bari, Tasnim Mohiuddin, Shafiq R. Joty
ACL/IJCNLP (1)1
2021 Nearest Neighbour Few-Shot Learning for Cross-lingual Classification
abstract
Even though large pre-trained multilingual models (e.g.mBERT, XLM-R) have led to significant performance gains on a wide range of cross-lingual NLP tasks, success on many downstream tasks still relies on the availability of sufficient annotated data.Traditional fine-tuning of pre-trained models using only a few target samples can cause over-fitting.This can be quite limiting as most languages in the world are under-resourced.In this work, we investigate cross-lingual adaptation using a simple nearest neighbor few-shot (< 15 samples) inference technique for classification tasks.We experiment using a total of 16 distinct languages across two NLP tasks-XNLI and PAWS-X.Our approach consistently improves traditional fine-tuning using only a handful of labeled samples in target locales.We also demonstrate its generalization capability across tasks.* Work done while Saiful was interning at Amazon AI 1 We loosely use the term LM to describe unsupervised pretrained models including Masked-LMs and Causal-LMs
Saiful Bari, Batool Haider, Saab Mansour
EMNLP (1)1
2020 Zero-Resource Cross-Lingual Named Entity Recognition
abstract
Recently, neural methods have achieved state-of-the-art (SOTA) results in Named Entity Recognition (NER) tasks for many languages without the need for manually crafted features. However, these models still require manually annotated training data, which is not available for many languages. In this paper, we propose an unsupervised cross-lingual NER model that can transfer NER knowledge from one language to another in a completely unsupervised way without relying on any bilingual dictionary or parallel data. Our model achieves this through word-level adversarial learning and augmented fine-tuning with parameter sharing and feature augmentation. Experiments on five different languages demonstrate the effectiveness of our approach, outperforming existing models by a good margin and setting a new SOTA for each language pair.
Saiful Bari, Shafiq R. Joty, Prathyusha Jwalapuram
AAAI1
2020 LNMap: Departures from Isomorphic Assumption in Bilingual Lexicon Induction Through Non-Linear Mapping in Latent Space
abstract
Most of the successful and predominant methods for Bilingual Lexicon Induction (BLI) are mapping-based, where a linear mapping function is learned with the assumption that the word embedding spaces of different languages exhibit similar geometric structures (i.e., approximately isomorphic).However, several recent studies have criticized this simplified assumption showing that it does not hold in general even for closely related languages.In this work, we propose a novel semi-supervised method to learn cross-lingual word embeddings for BLI.Our model is independent of the isomorphic assumption and uses non-linear mapping in the latent space of two independently pre-trained autoencoders.Through extensive experiments on fifteen (15) different language pairs (in both directions) comprising resource-rich and low-resource languages from two different datasets, we demonstrate that our method outperforms existing models by a good margin.Ablation studies show the importance of different model components and the necessity of non-linear mapping.
Tasnim Mohiuddin, Saiful Bari, Shafiq R. Joty
EMNLP (1)2
2019 A Unified Linear-Time Framework for Sentence-Level Discourse Parsing
abstract
We propose an efficient neural framework for sentence-level discourse analysis in accordance with Rhetorical Structure Theory (RST). Our framework comprises a discourse segmenter to identify the elementary discourse units (EDU) in a text, and a discourse parser that constructs a discourse tree in a top-down fashion. Both the segmenter and the parser are based on Pointer Networks and operate in linear time. Our segmenter yields an F1 score of 95.4%, and our parser achieves an F1 score of 81.7% on the aggregated labeled (relation) metric, surpassing previous approaches by a good margin and approaching human agreement on both tasks (98.3 and 83.0 F1).
Shafiq R. Joty, Prathyusha Jwalapuram, Saiful Bari
ACL (1)4