VLDB 2026 Research / reviewers in the wild / expert
Felermino D. M. A. Ali
dblp:340/7069 · also Felermino Ali, Felermino Dário Mário António Ali
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Information extraction and text analysis · 51% Machine translation · 37% Language models and text generation · 12% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
low-resource machine translation |
1.9 | 3 | 2025 | Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual Context · EMNLP 2025 Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks · EMNLP 2024 SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
emotion recognition |
0.9 | 1 | 2025 | BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages · ACL (1) 2025 |
Natural language and speech › Machine translation
machine translation evaluation |
0.9 | 1 | 2025 | SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
multilingual NLP |
0.9 | 1 | 2025 | BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis
text classification |
0.8 | 1 | 2024 | Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis › sentiment analysis › sentiment classification
cross-lingual sentiment classification |
0.7 | 1 | 2023 | AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages · EMNLP 2023 |
Natural language and speech › Language models and text generation
multilingual language models |
0.7 | 1 | 2023 | AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.7 | 1 | 2023 | AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages · EMNLP 2023 |
Natural language and speech › Language models and text generation › large language model evaluation › automatic evaluation
learned evaluation metric |
0.3 | 1 | 2025 | SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.7supervised fine-tuning · 0.9loanword constraints · 0.9learned metrics · 0.9multilingual encoder-decoder · 0.8fine-tuning · 0.8back-translation · 0.8OCR post-correction · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 LanguagesabstractShamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava 0001, Aura Cristina Udrea, Lilian Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou 0019, Saif M. Mohammad |
ACL (1) | 13 |
| 2025 | Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual ContextabstractThis research investigates how to improve machine translation systems for low-resource languages by integrating loanword constraints as external linguistic knowledge.Focusing on the Portuguese-Emakhuwa language pair, which exhibits significant lexical borrowing, we address the challenge of effectively adapting loanwords during the translation process.To tackle this, we propose a novel approach that augments source sentences with loanword constraints, explicitly linking source-language loanwords to their target-language equivalents.Then, we perform supervised fine-tuning on multilingual neural machine translation models and multiple Large Language Models of different sizes.Our results demonstrate that incorporating loanword constraints leads to significant improvements in translation quality as well as in handling loanword adaptation correctly in target languages, as measured by different machine translation metrics.This approach offers a promising direction for improving machine translation performance in low-resource settings characterized by frequent lexical borrowing. Felermino D. M. A. Ali, Henrique Lopes Cardoso, Rui Sousa-Silva |
EMNLP | 1 |
| 2025 | SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?abstractSenyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Senyu Li, Jiayi Wang 0010, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani |
EMNLP | 3 |
| 2024 | Detecting Loanwords in Emakhuwa: An Extremely Low-Resource Bantu Language Exhibiting Significant Borrowing from PortugueseabstractThe accurate identification of loanwords within a given text holds significant potential as a valuable tool for addressing data augmentation and mitigating data sparsity issues. Such identification can improve the performance of various natural language processing tasks, particularly in the context of low-resource languages that lack standardized spelling conventions.This research proposes a supervised method to identify loanwords in Emakhuwa, borrowed from Portuguese. Our methodology encompasses a two-fold approach. Firstly, we employ traditional machine learning algorithms incorporating handcrafted features, including language-specific and similarity-based features. We build upon prior studies to extract similarity features and propose utilizing two external resources: a Sequence-to-Sequence model and a dictionary. This innovative approach allows us to identify loanwords solely by analyzing the target word without prior knowledge about its donor counterpart. Furthermore, we fine-tune the pre-trained CANINE model for the downstream task of loanword detection, which culminates in the impressive achievement of the F1-score of 93%. To the best of our knowledge, this study is the first of its kind focusing on Emakhuwa, and the preliminary results are promising as they pave the way to further advancements. Felermino D. M. A. Ali, Henrique Lopes Cardoso, Rui Sousa-Silva |
LREC/COLING | 1 |
| 2024 | Building Resources for Emakhuwa: Machine Translation and News Classification BenchmarksabstractThis paper introduces a comprehensive collection of NLP resources for Emakhuwa, Mozambique's most widely spoken language.The resources include the first manually translated news bitext corpus between Portuguese and Emakhuwa, news topic classification datasets, and monolingual data.We detail the process and challenges of acquiring this data and present benchmark results for machine translation and news topic classification tasks.Our evaluation examines the impact of different data types-originally clean text, postcorrected OCR, and back-translated data-and the effects of fine-tuning from pre-trained models, including those focused on African languages.Our benchmarks demonstrate good performance in news topic classification and promising results in machine translation.We fine-tuned multilingual encoder-decoder models using real and synthetic data and evaluated them on our test set and the FLORES evaluation sets.The results highlight the importance of incorporating more data and potential for future improvements.All models, code, and datasets are available in the https://huggingface. Felermino D. M. A. Ali, Henrique Lopes Cardoso, Rui Sousa-Silva |
EMNLP | 1 |
| 2023 | AfriSenti: A Twitter Sentiment Analysis Benchmark for African LanguagesabstractShamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, Stephen Arthur. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum, David Ifeoluwa Adelani, Seid Muhie Yimam, Ibrahim Said Ahmad, Meriem Beloucif, Saif M. Mohammad, Sebastian Ruder, Oumaima Hourrane, Alípio Mário Jorge, Pavel Brazdil, Felermino D. M. A. Ali, Davis David, Salomey Osei, Bello Shehu Bello, Falalu Ibrahim Lawan, Tajuddeen Rabiu Gwadabe, Samuel Rutunda, Tadesse Destaw Belay, Wendimu Baye Messelle, Hailu Beshada Balcha, Sisay Adugna Chala, Hagos Tesfahun Gebremichael, Bernard Opoku, Stephen Arthur |
EMNLP | 14 |