Tarek Naous

dblp:268/1448 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-0049-9318ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs
abstract
Zohaib Khan, Mustafa Dogan, Ifeoma Okoh, Pouya Sadeghi, Siddhartha Shrestha, Sergius Justus Chesami Nyah, Mahmoud O. Mokhiamar, Michael J Ryan, Tarek Naous. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mustafa Dogan, Ifeoma Okoh, Pouya Sadeghi, Siddhartha Shrestha, Sergius Justus Nyah, Mahmoud O. Mokhiamar, Michael J. Ryan, Tarek Naous
ACL (1)9
2025 CARE: Multilingual Human Preference Learning for Cultural Awareness
abstract
Language Models (LMs) are typically tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied.In this paper, we systematically analyze how native human cultural preferences can be incorporated into the preference learning process to train more culturally aware LMs.We introduce CARE, a multilingual resource containing 3,490 culturally specific questions and 31.7kresponses with human judgments.We demonstrate how a modest amount of high-quality native preferences improves cultural awareness across various LMs, outperforming larger generic preference data.Our analyses reveal that models with stronger initial cultural performance benefit more from alignment, leading to gaps among models developed in different regions with varying access to culturally relevant data.CARE is publicly available at https://github.com/Guochry/ CARE.
Geyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura, Yuki Mitsufuji, Alan Ritter, Wei Xu 0004
EMNLP2
2025 What are Foundation Models Cooking in the Post-Soviet World?
abstract
The culture of the Post-Soviet states is complex, shaped by a turbulent history that continues to influence current events.In this study, we investigate the Post-Soviet cultural food knowledge of foundation models by constructing BORSCH, a multimodal dataset encompassing 1147 and 823 dishes in the Russian and Ukrainian languages, centered around the Post-Soviet region.We demonstrate that leading models struggle to correctly identify the origins of dishes from Post-Soviet nations in both text-only and multimodal Question Answering (QA), instead over-predicting countries linked to the language the question is asked in.Through analysis of pretraining data, we show that these results can be explained by misleading dish-origin co-occurrences, along with linguistic phenomena such as Russian-Ukrainian code mixing.Finally, to move beyond QA-based assessments, we test models' abilities to produce accurate visual descriptions of dishes.The weak correlation between this task and QA suggests that QA alone may be insufficient as an evaluation of cultural understanding.To foster further research, we will make BORSCH publicly available at github.com/alavrouk/BORSch.
Anton Lavrouk, Tarek Naous, Alan Ritter, Wei Xu 0004
EMNLP2
2025 On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
abstract
Tarek Naous, Wei Xu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Tarek Naous, Wei Xu 0004
NAACL (Long Papers)1
2025 Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AI
abstract
In pseudonymous online fora like Reddit, the benefits of self-disclosure are often apparent to users (e.g., I can vent about my in-laws to understanding strangers), but the privacy risks are more abstract (e.g., will my partner be able to tell that this is me?). Prior work has sought to develop natural language processing (NLP) tools that help users identify potentially risky self-disclosures in their text, but none have been designed for or evaluated with the users they hope to protect. Absent this assessment, these tools will be limited by the social-technical gap: users need assistive tools that help them make informed decisions, not paternalistic tools that tell them to avoid self-disclosure altogether. To bridge this gap, we conducted a study with N =21 Reddit users; we had them use a state-of-the-art NLP disclosure detection model on two of their authored posts and asked them questions to understand if and how the model helped, where it fell short, and how it could be improved to help them make more informed decisions. Despite its imperfections, users responded positively to the model and highlighted its use as a tool that can help them catch mistakes, inform them of risks they were unaware of, and encourage self-reflection. However, our work also shows how, to be useful and usable, AI for supporting privacy decision-making must account for posting context, disclosure norms, and users' lived threat models, and provide explanations that help contextualize detected risks.
Isadora Krsek, Anubha Kabra, Yao Dou, Tarek Naous, Laura A. Dabbish, Alan Ritter, Wei Xu 0004, Sauvik Das
Proc. ACM Hum. Comput. Interact.4
2024 Reducing Privacy Risks in Online Self-Disclosures with Language Models
abstract
Yao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, Wei Xu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, Wei Xu 0004
ACL (1)3
2024 Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
abstract
As the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial.Despite advancements in multilingual capabilities, models are not designed with appropriate cultural nuances.In this paper, we show that multilingual and Arabic monolingual LMs exhibit bias towards entities associated with Western culture.We introduce CAMeL, a novel resource of 628 naturally-occurring prompts and 20,368 entities spanning eight types that contrast Arab and Western cultures.CAMeL provides a foundation for measuring cultural biases in LMs through both extrinsic and intrinsic evaluations.Using CAMeL, we examine the cross-cultural performance in Arabic of 16 different LMs on tasks such as story generation, NER, and sentiment analysis, where we find concerning cases of stereotyping and cultural unfairness.We further test their text-infilling performance, revealing the incapability of appropriate adaptation to Arab cultural contexts.Finally, we analyze 6 Arabic pre-training corpora and find that commonly used sources such as Wikipedia may not be best suited to build culturally aware LMs, if used as they are without adjustment.We will make CAMeL publicly available at: https://github.com/tareknaous/camel
Tarek Naous, Michael J. Ryan, Alan Ritter, Wei Xu 0004
ACL (1)1
2024 ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment
abstract
We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This paper introduces ReadMe++, a multilingual multi-domain dataset with human annotations of 9757 sentences in Arabic, English, French, Hindi, and Russian, collected from 112 different data sources. This benchmark will encourage research on developing robust multilingual readability assessment methods. Using ReadMe++, we benchmark multilingual and monolingual language models in the supervised, unsupervised, and few-shot prompting settings. The domain and language diversity in ReadMe++ enable us to test more effective few-shot prompting, and identify shortcomings in state-of-the-art unsupervised methods. Our experiments also reveal exciting results of superior domain generalization and enhanced cross-lingual transfer capabilities by models trained on ReadMe++. We will make our data publicly available and release a python package tool for multilingual sentence readability prediction using our trained models at: https://github.com/tareknaous/readme.
Tarek Naous, Michael J. Ryan, Anton Lavrouk, Mohit Chandra, Wei Xu 0004
EMNLP1
2023 Revisiting non-English Text Simplification: A Unified Multilingual Benchmark
abstract
Recent advancements in high-quality, largescale English resources have pushed the frontier of English Automatic Text Simplification (ATS) research.However, less work has been done on multilingual text simplification due to the lack of a diverse evaluation benchmark that covers complex-simple sentence pairs in many languages.This paper introduces the MULTI-SIM benchmark, a collection of 27 resources in 12 distinct languages containing over 1.7 million complex-simple sentence pairs.This benchmark will encourage research in developing more effective multilingual text simplification models and evaluation metrics.Our experiments using MULTISIM with pre-trained multilingual language models reveal exciting performance improvements from multilingual training in non-English settings.We observe strong performance from Russian in zero-shot crosslingual transfer to low-resource languages.We further show that few-shot prompting with BLOOM-176b achieves comparable quality to reference simplifications outperforming finetuned models in most languages.We validate these findings through human evaluation.
Michael J. Ryan, Tarek Naous, Wei Xu 0004
ACL (1)2
2023 Open-Domain Response Generation in Low-Resource Settings using Self-Supervised Pre-Training of Warm-Started Transformers
abstract
Learning response generation models constitute the main component of building open-domain dialogue systems. However, training open-domain response generation models requires large amounts of labeled data and pre-trained language generation models that are often nonexistent for low-resource languages. In this article, we propose a framework for training open-domain response generation models in low-resource settings. We consider Dialectal Arabic (DA) as a working example. The framework starts by warm-starting a transformer-based encoder-decoder with pre-trained language model parameters. Next, the resultant encoder-decoder model is adapted to DA by employing self-supervised pre-training on large-scale unlabeled data in the desired dialect. Finally, the model is fine-tuned on a very small labeled dataset for open-domain response generation. The results show significant performance improvements on three spoken Arabic dialects after adopting the framework’s three stages, highlighted by higher BLEU and lower Perplexity scores compared with multiple baseline models. Specifically, our models are capable of generating fluent responses in multiple dialects with an average human-evaluated fluency score above 4. Our data is made publicly available.
Tarek Naous, Zahraa Bassyouni, Basel Mousi, Hazem M. Hajj, Wassim El-Hajj, Khaled B. Shaban
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2022 Clustering Plotted Data by Image Segmentation
abstract
Clustering is a popular approach to detecting patterns in unlabeled data. Existing clustering methods typically treat samples in a dataset as points in a metric space and compute distances to group together similar points. In this paper, we present a different way of clustering points in 2-dimensional space, inspired by how humans cluster data: by training neural networks to perform instance segmentation on plotted data. Our approach, Visual Clustering, has several advantages over traditional clustering algorithms: it is much faster than most existing clustering algorithms (making it suitable for very large datasets), it agrees strongly with human intuition for clusters, and it is by default hyperparameter free (although additional steps with hyperparameters can be introduced for more control of the algorithm). We describe the method and compare it to ten other clustering methods on synthetic data to illustrate its advantages and disadvantages. We then demonstrate how our approach can be extended to higher-dimensional data and illustrate its performance on real-world data. Our implementation of Visual Clustering is publicly available as a python package that can be installed and used on any dataset in a few lines of code11https://hithub.com/tareknaous/visual-clustering. A demo on synthetic datasets is provided22https://huggingface.co/spaces/CVPR/visual-clustering.
Tarek Naous, Srinjay Sarkar, Abubakar Abid, James Zou 0001
CVPR1
2022 Stanceosaurus: Classifying Stance Towards Multicultural Misinformation
abstract
We present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.As far as we are aware, it is the largest corpus annotated with stance towards misinformation claims.The claims in Stanceosaurus originate from 15 fact-checking sources that cover diverse geographical regions and cultures.Unlike existing stance datasets, we introduce a more fine-grained 5class labeling strategy with additional subcategories to distinguish implicit stance.Pretrained transformer-based stance classifiers that are fine-tuned on our corpus show good generalization on unseen claims and regional claims from countries outside the training data.Cross-lingual experiments demonstrate Stanceosaurus' capability of training multilingual models, achieving 53.1 F1 on Hindi and 50.4 F1 on Arabic without any targetlanguage fine-tuning.Finally, we show how a domain adaptation method can be used to improve performance on Stanceosaurus using additional RumourEval-2019 data.We make Stanceosaurus publicly available to the research community and hope it will encourage further work on misinformation identification across languages and cultures.1 Source Country & Regions Lang #Claims #Tweets Irr.Sup.Ref. Dis.Que.Snopes USA (80%), INT'L (16.7%),Other (3.3%) en 30 3197 1051 428 229 1447 42 Poynter Europe (5%), INT'L (90%), Other (5%) en 20 2197 949 274 97 844 33 FullFact UK (30%), INT'L (55%), Other (15%) en 20 2379 806 300 179 1057 37 AFP Fact Check CAN Canada (55%), INT'L (30%), Other (15%) en 20 2078 746 252 130 910 40 AAP Fact Check Australia (10%), INT'L (65%), Other (25%) en 20 2302 739 374 136 1019 34 AFP Fact Check NZ New Zealand (15%), INT'L (75%), Other (10%) en 20 2227 879 194 81 1044 29 Blackdotresearch Singapore (30%), INT'L (55%), Other (15%) en 20 2307 842 248 113 1076 28 Factly India (45%), INT'L (55%) en 20 1979 889 190 117 734 49 Politifact USA (20%), INT'L (35%), Other (45%) en 20 2041 984 289 8 753 7 Alt News India (90.4%),INT'L (4.8%),Other (4.8%) hi 21 1730 550 489 172 500 19 Aajtak India (67%), Other (33%) hi 9 806 456 110 40 193 7 Hindi Newschecker India (56%), Other (44%) hi 9 781 195 313 46 219 8 MISBAR Arab World (58.3%),INT'L (8.3%),Other (33.4%) ar 12 2283 454 514 203 1031 81 Fatabyyano Arab World (28.5%),INT'L (57.1%),Other (14.4%) ar 7 986 234 163 49 522 18 Maharat Fact-o-meter INT'L (100%) ar 3 740
Jonathan Zheng, Ashutosh Baheti, Tarek Naous, Wei Xu 0004, Alan Ritter
EMNLP3
2022 A survey on blockchain-based Recommender Systems: Integration architecture and taxonomy
Loubna Mekouar, Youssef Iraqi, Issam W. Damaj, Tarek Naous
Comput. Commun.4