Khalid Almubarak

dblp:313/2169 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0006-4869-6440ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 68% Transfer learning and domain adaptation · 15% Knowledge representation and reasoning · 10%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
multilingual language models
2.442025
Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion · ACL (1) 2025
BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting · ACL (1) 2023
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset · NeurIPS 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.912025
Commonsense Reasoning in Arab Culture · ACL (1) 2025
Natural language and speech › Language models and text generation › multilingual language models
cross-lingual generalization
0.712023
Crosslingual Generalization through Multitask Finetuning · ACL (1) 2023
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.712023
BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting · ACL (1) 2023
Machine learning › Transfer learning and domain adaptation
language adaptation
0.712023
BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting · ACL (1) 2023
Natural language and speech › Language models and text generation › large language model fine-tuning
multi-task fine-tuning
0.712023
Crosslingual Generalization through Multitask Finetuning · ACL (1) 2023
Natural language and speech › Language models and text generation › prompting
zero-shot prompting
0.712023
BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting · ACL (1) 2023
Natural language and speech › Language models and text generation
multilingual dataset
0.612022
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset · NeurIPS 2022
Machine learning › Efficient and distributed learning › data curation
training data curation
0.612022
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

continued pretraining · 1.5progressive vocabulary expansion · 0.9cultural knowledge elicitation · 0.9benchmark construction · 0.9vocabulary extension · 0.7multi-task fine-tuning · 0.7large language model · 0.7
YearPublicationVenuePosition
2025 Commonsense Reasoning in Arab Culture
abstract
Abdelrahman Sadallah, Junior Cedric Tonga, Khalid Almubarak, Saeed Almheiri, Farah Atif, Chatrine Qwaider, Karima Kadaoui, Sara Shatnawi, Yaser Alesh, Fajri Koto. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Abdelrahman Boda Sadallah, Junior Cedric Tonga, Khalid Almubarak, Saeed Almheiri, Farah Atif, Chatrine Qwaider, Karima Kadaoui, Sara Shatnawi, Yaser Alesh, Fajri Koto
ACL (1)3
2025 Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion
abstract
Jianqing Zhu, Huang Huang, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An, Juncai He, Xiangbo Wu, Fei Yu, Junying Chen, Ma Zhuoheng, Yuhao Du, He Zhang, Saied Alshahrani, Emad A. Alghamdi, Lian Zhang, Ruoyu Sun, Haizhou Li, Benyou Wang, Jinchao Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jianqing Zhu, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An 0004, Juncai He 0001, Xiangbo Wu, Fei Yu 0017, Zhuoheng Ma, Saied Alshahrani, Emad A. Alghamdi, Ruoyu Sun 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu
ACL (1)6
2025 Unlocking language boundaries: AraCLIP - transforming Arabic language and image understanding through cross-lingual models
Muhammad Al-Barham, Imad Afyouni, Khalid Almubarak, Ayad Mashaan Turky, Ibrahim Abaker Targio Hashem, Ali Bou Nassif, Ismail Shahin, Ashraf Elnagar
Eng. Appl. Artif. Intell.3
2023 Crosslingual Generalization through Multitask Finetuning
abstract
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, Colin Raffel. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, Saiful Bari, Sheng Shen 0001, Hailey Schoelkopf, Xiangru Tang, Dragomir R. Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, Colin Raffel
ACL (1)14
2023 BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting
abstract
Zheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Winata, Stella Biderman, Edward Raff, Dragomir Radev, Vassilina Nikoulina. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Indra Winata, Stella Biderman, Edward Raff, Dragomir R. Radev, Vassilina Nikoulina
ACL (1)6
2022 The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
abstract
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large language models as a values-driven undertaking, putting issues of ethics, harm, and governance in the foreground. This paper documents the data creation and curation efforts undertaken by BigScience to assemble the Responsible Open-science Open-collaboration Text Sources (ROOTS) corpus, a 1.6TB dataset spanning 59 languages that was used to train the 176-billion-parameter BigScience Large Open-science Open-access Multilingual (BLOOM) language model. We further release a large initial subset of the corpus and analyses thereof, and hope to empower large-scale monolingual and multilingual modeling projects with both the data and the processing tools, as well as stimulate research around this large multilingual corpus.
Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro von Werra, Chenghao Mou, Eduardo G. Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Sasko, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben Allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa 0001, Paulo Villegas, Tristan Thrush, Shayne Longpre, Sebastian Nagel 0005, Leon Weber-Genzel, Manuel Muñoz, Daniel van Strien, Zaid Alyafeai, Khalid Almubarak, Minh Chien Vu, Itziar Gonzalez-Dios, Aitor Soroa, Kyle Lo, Manan Dey, Pedro Ortiz Suarez, Aaron Gokaslan, Shamik Bose, David Ifeoluwa Adelani, Long Phan, Hieu Tran, Ian Yu, Suhas Pai, Jenny Chim, Violette Lepercq, Suzana Ilic, Margaret Mitchell, Sasha Luccioni, Yacine Jernite
NeurIPS35