Felermino D. M. A. Ali

dblp:340/7069 · also Felermino Ali, Felermino Dário Mário António Ali · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Information extraction and text analysis · 51% Machine translation · 37% Language models and text generation · 12%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
low-resource machine translation
1.932025
Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual Context · EMNLP 2025
Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks · EMNLP 2024
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? · EMNLP 2025
Natural language and speech › Information extraction and text analysis
emotion recognition
0.912025
BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages · ACL (1) 2025
Natural language and speech › Machine translation
machine translation evaluation
0.912025
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? · EMNLP 2025
Natural language and speech › Information extraction and text analysis
multilingual NLP
0.912025
BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages · ACL (1) 2025
Natural language and speech › Information extraction and text analysis
text classification
0.812024
Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks · EMNLP 2024
Natural language and speech › Information extraction and text analysis › sentiment analysis › sentiment classification
cross-lingual sentiment classification
0.712023
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages · EMNLP 2023
Natural language and speech › Language models and text generation
multilingual language models
0.712023
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages · EMNLP 2023
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.712023
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages · EMNLP 2023
Natural language and speech › Language models and text generation › large language model evaluation › automatic evaluation
learned evaluation metric
0.312025
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

large language model · 1.7supervised fine-tuning · 0.9loanword constraints · 0.9learned metrics · 0.9multilingual encoder-decoder · 0.8fine-tuning · 0.8back-translation · 0.8OCR post-correction · 0.8
YearPublicationVenuePosition
2025 BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages
abstract
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava 0001, Aura Cristina Udrea, Lilian Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou 0019, Saif M. Mohammad
ACL (1)13
2025 Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual Context
abstract
This research investigates how to improve machine translation systems for low-resource languages by integrating loanword constraints as external linguistic knowledge.Focusing on the Portuguese-Emakhuwa language pair, which exhibits significant lexical borrowing, we address the challenge of effectively adapting loanwords during the translation process.To tackle this, we propose a novel approach that augments source sentences with loanword constraints, explicitly linking source-language loanwords to their target-language equivalents.Then, we perform supervised fine-tuning on multilingual neural machine translation models and multiple Large Language Models of different sizes.Our results demonstrate that incorporating loanword constraints leads to significant improvements in translation quality as well as in handling loanword adaptation correctly in target languages, as measured by different machine translation metrics.This approach offers a promising direction for improving machine translation performance in low-resource settings characterized by frequent lexical borrowing.
Felermino D. M. A. Ali, Henrique Lopes Cardoso, Rui Sousa-Silva
EMNLP1
2025 SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
abstract
Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Senyu Li, Jiayi Wang 0010, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani
EMNLP3
2024 Detecting Loanwords in Emakhuwa: An Extremely Low-Resource Bantu Language Exhibiting Significant Borrowing from Portuguese
abstract
The accurate identification of loanwords within a given text holds significant potential as a valuable tool for addressing data augmentation and mitigating data sparsity issues. Such identification can improve the performance of various natural language processing tasks, particularly in the context of low-resource languages that lack standardized spelling conventions.This research proposes a supervised method to identify loanwords in Emakhuwa, borrowed from Portuguese. Our methodology encompasses a two-fold approach. Firstly, we employ traditional machine learning algorithms incorporating handcrafted features, including language-specific and similarity-based features. We build upon prior studies to extract similarity features and propose utilizing two external resources: a Sequence-to-Sequence model and a dictionary. This innovative approach allows us to identify loanwords solely by analyzing the target word without prior knowledge about its donor counterpart. Furthermore, we fine-tune the pre-trained CANINE model for the downstream task of loanword detection, which culminates in the impressive achievement of the F1-score of 93%. To the best of our knowledge, this study is the first of its kind focusing on Emakhuwa, and the preliminary results are promising as they pave the way to further advancements.
Felermino D. M. A. Ali, Henrique Lopes Cardoso, Rui Sousa-Silva
LREC/COLING1
2024 Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks
abstract
This paper introduces a comprehensive collection of NLP resources for Emakhuwa, Mozambique's most widely spoken language.The resources include the first manually translated news bitext corpus between Portuguese and Emakhuwa, news topic classification datasets, and monolingual data.We detail the process and challenges of acquiring this data and present benchmark results for machine translation and news topic classification tasks.Our evaluation examines the impact of different data types-originally clean text, postcorrected OCR, and back-translated data-and the effects of fine-tuning from pre-trained models, including those focused on African languages.Our benchmarks demonstrate good performance in news topic classification and promising results in machine translation.We fine-tuned multilingual encoder-decoder models using real and synthetic data and evaluated them on our test set and the FLORES evaluation sets.The results highlight the importance of incorporating more data and potential for future improvements.All models, code, and datasets are available in the https://huggingface.
Felermino D. M. A. Ali, Henrique Lopes Cardoso, Rui Sousa-Silva
EMNLP1
2023 AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages
abstract
Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, Stephen Arthur. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum, David Ifeoluwa Adelani, Seid Muhie Yimam, Ibrahim Said Ahmad, Meriem Beloucif, Saif M. Mohammad, Sebastian Ruder, Oumaima Hourrane, Alípio Mário Jorge, Pavel Brazdil, Felermino D. M. A. Ali, Davis David, Salomey Osei, Bello Shehu Bello, Falalu Ibrahim Lawan, Tajuddeen Rabiu Gwadabe, Samuel Rutunda, Tadesse Destaw Belay, Wendimu Baye Messelle, Hailu Beshada Balcha, Sisay Adugna Chala, Hagos Tesfahun Gebremichael, Bernard Opoku, Stephen Arthur
EMNLP14