Yassine Benajiba

dblp:17/6428 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
5since 2021 · last 2025
0009-0008-8158-4867ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 8 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 32% Information extraction and text analysis · 23% Trustworthy machine learning · 16%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 87% Recommender systems · 13%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › active learning
active sampling
0.912025
Active Evaluation Acquisition for Efficient LLM Benchmarking · ICML 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Active Evaluation Acquisition for Efficient LLM Benchmarking · ICML 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025
Natural language and speech › Language models and text generation
memory augmentation
0.912025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025
Machine learning › Reinforcement learning
policy learning
0.912025
Active Evaluation Acquisition for Efficient LLM Benchmarking · ICML 2025
Information retrieval › retrieval-augmented generation
memory retrieval
0.912025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025
Information retrieval
retrieval-augmented generation
0.912025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025
Natural language and speech › Information extraction and text analysis
named entity recognition
0.832023
Taxonomy Expansion for Named Entity Recognition · EMNLP 2023
Arabic Named Entity Recognition: A Feature-Driven Study · IEEE Trans. Speech Audio Process. 2009
Arabic Named Entity Recognition using Optimized Feature Sets · EMNLP 2008
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.712023
Characterizing and Measuring Linguistic Dataset Drift · ACL (1) 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
Characterizing and Measuring Linguistic Dataset Drift · ACL (1) 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › ontology learning
taxonomy expansion
0.712023
Taxonomy Expansion for Named Entity Recognition · EMNLP 2023
Natural language and speech › Information extraction and text analysis › named entity recognition
mention detection
0.322014
Aligned-Parallel-Corpora Based Semi-Supervised Learning for Arabic Mention Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Enhancing Mention Detection Using Projection via Aligned Corpora · EMNLP 2010
Recommender systems › interactive recommendation
conversational recommendation
0.312025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual information extraction
0.112014
Aligned-Parallel-Corpora Based Semi-Supervised Learning for Arabic Mention Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual projection
0.012010
Enhancing Mention Detection Using Projection via Aligned Corpora · EMNLP 2010
Natural language and speech › Information extraction and text analysis
Arabic NLP
0.012009
Arabic Named Entity Recognition: A Feature-Driven Study · IEEE Trans. Speech Audio Process. 2009

Methods — techniques the papers use, named apart from their topics

autonomous memory augmentation · 1.7reinforcement learning · 0.9dependency modeling · 0.9active evaluation acquisition · 0.9language model prompting · 0.7dataset characterization · 0.7semi-supervised learning · 0.2feature propagation · 0.2annotation projection · 0.1conditional random field · 0.1
YearPublicationVenuePosition
2025 MemInsight: Autonomous Memory Augmentation for LLM Agents
abstract
Large language model (LLM) agents have evolved to intelligently process information, make decisions, and interact with users or tools. A key capability is the integration of long-term memory capabilities, enabling these agents to draw upon historical interactions and knowledge. However, the growing memory size and need for semantic structuring pose significant challenges. In this work, we propose an autonomous memory augmentation approach, MemInsight, to enhance semantic data representation and retrieval mechanisms. By leveraging autonomous augmentation to historical interactions, LLM agents are shown to deliver more accurate and contextualized responses. We empirically validate the efficacy of our proposed approach in three task scenarios; conversational recommendation, question answering and event summarization. On the LLM-REDIAL dataset, MemInsight boosts persuasiveness of recommendations by up to 14%. Moreover, it outperforms a RAG baseline by 34% in recall for LoCoMo retrieval. Our empirical results show the potential of MemInsight to enhance the contextual performance of LLM agents across multiple tasks.
Rana Salama, Jason Cai, Michelle Yuan, Anna Currey, Monica Sunkara, Yassine Benajiba
EMNLP7
2025 Active Evaluation Acquisition for Efficient LLM Benchmarking
abstract
As large language models (LLMs) become increasingly versatile, numerous large scale benchmarks have been developed to thoroughly assess their capabilities. These benchmarks typically consist of diverse datasets and prompts to evaluate different aspects of LLM performance. However, comprehensive evaluations on hundreds or thousands of prompts incur tremendous costs in terms of computation, money, and time. In this work, we investigate strategies to improve evaluation efficiency by selecting a subset of examples from each benchmark using a learned policy. Our approach models the dependencies across test examples, allowing accurate prediction of the evaluation outcomes for the remaining examples based on the outcomes of the selected ones. Consequently, we only need to acquire the actual evaluation outcomes for the selected subset. We rigorously explore various subset selection policies and introduce a novel RL-based policy that leverages the captured dependencies. Empirical results demonstrate that our approach significantly reduces the number of evaluation prompts required while maintaining accurate performance estimates compared to previous methods.
Jie Ma 0005, Miguel Ballesteros, Yassine Benajiba, Graham Horwood
ICML4
2023 Characterizing and Measuring Linguistic Dataset Drift
abstract
Tyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas, Yassine Benajiba, Miguel Ballesteros, Dan Roth. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Tyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas, Yassine Benajiba, Miguel Ballesteros, Dan Roth 0001
ACL (1)5
2023 Dynamic Benchmarking of Masked Language Models on Temporal Concept Drift with Multiple Views
abstract
Katerina Margatina, Shuai Wang, Yogarshi Vyas, Neha Anna John, Yassine Benajiba, Miguel Ballesteros. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Aikaterini Margatina, Yogarshi Vyas, Neha Anna John, Yassine Benajiba, Miguel Ballesteros
EACL5
2023 Taxonomy Expansion for Named Entity Recognition
abstract
Karthikeyan K, Yogarshi Vyas, Jie Ma, Giovanni Paolini, Neha John, Shuai Wang, Yassine Benajiba, Vittorio Castelli, Dan Roth, Miguel Ballesteros. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Karthikeyan K, Yogarshi Vyas, Jie Ma 0005, Giovanni Paolini, Neha Anna John, Yassine Benajiba, Vittorio Castelli, Dan Roth 0001, Miguel Ballesteros
EMNLP7
2020 Aspect On: an Interactive Solution for Post-Editing the Aspect Extraction based on Online Learning
abstract
The task of aspect extraction is an important component of aspect-based sentiment analysis. However, it usually requires an expensive human post-processing to ensure quality. In this work we introduce Aspect On, an interactive solution based on online learning that allows users to post-edit the aspect extraction with little effort. The Aspect On interface shows the aspects extracted by a neural model and, given a dataset, annotates its words with the corresponding aspects. Thanks to the online learning, Aspect On updates the model automatically and continuously improves the quality of the aspects displayed to the user. Experimental results show that Aspect On dramatically reduces the number of user clicks and effort required to post-edit the aspects extracted by the model.
Mara Chinea-Rios, Marc Franco-Salvador, Yassine Benajiba
LREC3
2014 Aligned-Parallel-Corpora Based Semi-Supervised Learning for Arabic Mention Detection
abstract
In the last two decades, significant effort has been put into annotating linguistic resources in several languages. Despite this valiant effort, there are still many languages left that have only small amounts of such resources. The goal of this article is to present and investigate a method of propagating information (specifically mentions) from a resource-rich language such as English into a relatively less-resource language such as Arabic. We compare also this approach to its equivalent counterpart using monolingual resources. Part of the investigation is to quantify the contribution of propagating information in different conditions - based on the availability of resources in the target language. Experiments on the language pair Arabic-English show that one can achieve relatively decent performance by propagating information from a language with richer resources such as English into Arabic alone (no resources or models in the source language Arabic). Furthermore, results show that propagated features from English do help improve the Arabic system performance even when used in conjunction with all feature types built from the source language. Experiments also show that using propagated features in conjunction with lexically-derived features only (as can be obtained directly from a mention annotated corpus) brings the system performance at the one obtained in the target language by using feature derived from many linguistic resources, therefore improving the system when such resources are not available.
Imed Zitouni, Yassine Benajiba
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 The development of a fine grained class set for Amazigh POS tagging
abstract
Like most of the languages which have only recently started being investigated for the Natural Language Processing (NLP) tasks, Amazigh lacks annotated corpora and tools and still suffers from the scarcity of linguistic tools and resources. The main aim of this paper is to present a tokenizer tool and a new part-of-speech (POS) tagger based on a new Amazigh tag set (AMTS) composed of 28 tag. In line with our goal we have trained two sequence classification models using Support Vector Machines (SVMs) and Conditional Random Fields (CRFs) to build a toknizer and a POS tagger for the Amazigh language. We have used the 10-fold technique to evaluate and validate our approach. We report that POS tagging results using SVMs and CRFs are very comparable. Across the board, CRFs outperformed SVMs on the fold level (91.18% vs. 90.75%) and CRFs outperformed SVMs on the 10 folds average level (87.95% vs. 87.11%). Regarding tokenization task, SVMs outperformed CRFs on the fold level (99.97% vs. 99.85%) and on the 10 folds average level (99.95% vs. 99.89%).
Mohamed Outahajala, Lahbib Zenkouar, Yassine Benajiba, Paolo Rosso
AICCSA3
2011 POS Tagging in Amazighe Using Support Vector Machines and Conditional Random Fields
Mohamed Outahajala, Yassine Benajiba, Paolo Rosso, Lahbib Zenkouar
NLDB2
2010 Enhancing Mention Detection Using Projection via Aligned Corpora
Yassine Benajiba, Imed Zitouni
EMNLP1
2010 Arabic Word Segmentation for Better Unit of Analysis
Yassine Benajiba, Imed Zitouni
LREC1
2010 Arabic Mention Detection: Toward Better Unit of Analysis
Yassine Benajiba, Imed Zitouni
HLT-NAACL1
2009 Morphology-Based Segmentation Combination for Arabic Mention Detection
abstract
The Arabic language has a very rich/complex morphology. Each Arabic word is composed of zero or more prefixes , one stem and zero or more suffixes . Consequently, the Arabic data is sparse compared to other languages such as English, and it is necessary to conduct word segmentation before any natural language processing task. Therefore, the word-segmentation step is worth a deeper study since it is a preprocessing step which shall have a significant impact on all the steps coming afterward. In this article, we present an Arabic mention detection system that has very competitive results in the recent Automatic Content Extraction (ACE) evaluation campaign. We investigate the impact of different segmentation schemes on Arabic mention detection systems and we show how these systems may benefit from more than one segmentation scheme. We report the performance of several mention detection models using different kinds of possible and known segmentation schemes for Arabic text: punctuation separation, Arabic Treebank, and morphological and character-level segmentations. We show that the combination of competitive segmentation styles leads to a better performance. Results indicate a statistically significant improvement when Arabic Treebank and morphological segmentations are combined.
Yassine Benajiba, Imed Zitouni
ACM Trans. Asian Lang. Inf. Process.1
2009 Arabic Named Entity Recognition: A Feature-Driven Study
abstract
The named entity recognition task aims at identifying and classifying named entities within an open-domain text. This task has been garnering significant attention recently as it has been shown to help improve the performance of many natural language processing applications. In this paper, we investigate the impact of using different sets of features in three discriminative machine learning frameworks, namely, support vector machines, maximum entropy and conditional random fields for the task of named entity recognition. Our language of interest is Arabic. We explore lexical, contextual and morphological features and nine data-sets of different genres and annotations. We measure the impact of the different features in isolation and incrementally combine them in order to evaluate the robustness to noise of each approach. We achieve the highest performance using a combination of 15 features in conditional random fields using broadcast news data (Fbeta=1=83.34).
Yassine Benajiba, Mona T. Diab, Paolo Rosso
IEEE Trans. Speech Audio Process.1
2008 Arabic Named Entity Recognition using Optimized Feature Sets
Yassine Benajiba, Mona T. Diab, Paolo Rosso
EMNLP1
2007 ANERsys: An Arabic Named Entity Recognition System Based on Maximum Entropy
Yassine Benajiba, Paolo Rosso, José-Miguel Benedí
CICLing1
2007 Adapting the JIRS Passage Retrieval System to the Arabic Language
Yassine Benajiba, Paolo Rosso, José Manuel Gómez Soriano
CICLing1