Adrian-Gabriel Chifu

dblp:124/0331 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0003-4680-5528ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 8 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Uncovering the Limitations of Query Performance Prediction: Failures, Insights, and Implications for Selective Query Processing
abstract
Query Performance Prediction (QPP) estimates the effectiveness of retrieval systems for a given query, offering valuable insights for search effectiveness and query processing. Despite extensive research, a critical gap remains in understanding how well QPPs generalize across diverse retrieval paradigms and collections, a question of robustness that has significant implications for their practical utility. This article provides the first comprehensive cross-paradigm evaluation of QPP robustness and generalization capabilities, examining state-of-the-art QPPs including NQC, WIG, LETOR-based features, and newly explored dense-based predictors MQPPF and BERT-QPP. We systematically assess their performance across diverse sparse (BM25, DFree with and without query expansion), hybrid (SPLADE), and dense (ColBERT, TCT-ColBERT) rankers on four benchmark collections: TREC Robust, GOV2, WT10G, and MS-MARCO. The results reveal fundamental robustness challenges: predictors exhibit significant variability in accuracy, with collection being the dominant factor, followed by ranker type. Some sparse predictors perform adequately on specific collections such as TREC Robust and GOV2, but critically fail to generalize to other collections like WT10G and MS-MARCO. Dense-based predictors, while showing promise in specific scenarios with dense rankers, similarly lack generalization to sparse contexts. We demonstrate that these generalization failures severely limit practical applications: QPP-driven selective query processing achieves only marginal gains ( \(\approx\) 4% NDCG improvement), with reliability varying dramatically across settings. Our findings underscore that current QPP methods lack the robustness necessary for real-world deployment and highlight the urgent need for predictors that generalize reliably across diverse collections, align with modern dense retrieval architectures, and provide consistent utility for downstream applications. We publicly release our data and code to facilitate future research on robust QPP methods ( https://github.com/adrianchifu/UncoveringTheLimitationsofQPP/ ).
Adrian-Gabriel Chifu, Sébastien Déjean, Moncef Garouani, Josiane Mothe, Diégo Ortiz, Md. Zia Ullah
ACM Trans. Inf. Syst.1
2025 PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction
abstract
Text-to-image generation has recently emerged as a viable alternative to text-to-image retrieval, driven by the visually impressive results of generative diffusion models. Although query performance prediction is an active research topic in information retrieval, to the best of our knowledge, there is no prior study that analyzes the difficulty of queries (referred to as prompts) in text-to-image generation, based on human judgments. To this end, we introduce the first dataset of prompts which are manually annotated in terms of image generation performance. Additionally, we extend these evaluations to text-to-image retrieval by collecting manual annotations that represent retrieval performance. We thus establish the first joint benchmark for prompt and query performance prediction (PQPP) across both tasks, comprising over 10K queries. Our benchmark enables (i) the comparative assessment of prompt/query difficulty in both image generation and image retrieval, and (ii) the evaluation of prompt/query performance predictors addressing both generation and retrieval. We evaluate several pre- and post-generation/retrieval performance predictors, thus providing competitive baselines for future research. Our benchmark and code are publicly available at https://github.com/Eduard6421/PQPP.
Eduard Gabriel Poesina, Adriana Valentina Costache, Adrian-Gabriel Chifu, Josiane Mothe, Radu Tudor Ionescu
CVPR3
2025 BioReadNet: A Transformer-Driven Hybrid Model for Target Audience-Aware Biomedical Text Readability Assessment
abstract
The perception of the readability of biomedical texts varies depending on the reader's profile, a disparity further amplified by the intrinsic complexity of these documents and the unequal distribution of health literacy within the population. Although 72% of Internet users consult medical information online, a significant proportion have difficulty understanding it. To ensure that texts are accessible to a diverse audience, it is essential to assess readability. However, conventional readability formulas, designed for general texts, do not take this diversity into account, underlining the need to adapt evaluation tools to the specific needs of biomedical texts and the heterogeneity of readers. To address this gap, we propose a novel readability assessment method tailored to three distinct audiences: expert adults, non-expert adults, and children. Our approach is built upon a structured, bilingual biomedical corpus of 20,008 documents (8,854 in French, 11,154 in English), compiled from multiple sources to ensure diversity in both content and audience. Specifically, the French corpus combines texts from Cochrane and Wikipedia/Vikidia, both of which are subsets of the CLEAR corpus, while the English corpus merges documents from the Cochrane Library, Plaba, and Science Journal for Kids. For each original expert-level text, domain specialists produced simplified variants calibrated specifically to the comprehension abilities of non-expert adults or children. Every document is therefore explicitly labeled by its target audience. Leveraging this resource, we trained a diverse suite of classifiers, from classical approaches (e.g., XGBoost, SVM) to classifiers built upon language models (e.g., BERT, CamemBERT, BioBERT, DrBERT). We then designed a hybrid architecture "BioReadNet" that integrates transformer embeddings with expert-driven linguistic features, achieving a macro-averaged F1 score of 0.987.
Anya Amel Nait Djoudi, Patrice Bellot, Adrian-Gabriel Chifu
DocEng3
2024 Can We Predict QPP? An Approach Based on Multivariate Outliers
Adrian-Gabriel Chifu, Sébastien Déjean, Moncef Garouani, Josiane Mothe, Diégo Ortiz, Md. Zia Ullah
ECIR (3)1
2023 FreCDo: A Large Corpus for French Cross-Domain Dialect Identification
abstract
We present a novel corpus for French dialect identification comprising 413,522 French text samples collected from public news websites in Belgium, Canada, France and Switzerland. To ensure an accurate estimation of the dialect identification performance of models, we designed the corpus to eliminate potential biases related to topic, writing style, and publication source. More precisely, the training, validation and test splits are collected from different news websites, while searching for different keywords (topics). This leads to a French cross-domain (FreCDo) dialect identification task. We conduct experiments with four competitive baselines, a fine-tuned CamemBERT model, an XGBoost based on fine-tuned CamemBERT features, a Support Vector Machines (SVM) classifier based on fine-tuned CamemBERT features, and an SVM based on word n-grams. Aside from presenting quantitative results, we also make an analysis of the most discriminative features learned by CamemBERT. Our corpus is available in open-source format at https://github.com/MihaelaGaman/FreCDo.
Mihaela Gaman, Adrian-Gabriel Chifu, William Domingues, Radu Tudor Ionescu
KES2
2022 DeepREF: A Framework for Optimized Deep Learning-based Relation Classification
abstract
The Relation Extraction (RE) is an important basic Natural Language Processing (NLP) for many applications, such as search engines, recommender systems, question-answering systems and others. There are many studies in this subarea of NLP that continue to be explored, such as SemEval campaigns (2010 to 2018), or DDI Extraction (2013).For more than ten years, different RE systems using mainly statistical models have been proposed as well as the frameworks to develop them. This paper focuses on frameworks allowing to develop such RE systems using deep learning models. Such frameworks should make it possible to reproduce experiments of various deep learning models and pre-processing techniques proposed in various publications. Currently, there are very few frameworks of this type, and we propose a new open and optimizable framework, called DeepREF, which is inspired by the OpenNRE and REflex existing frameworks. DeepREF allows the employment of various deep learning models, to optimize their use, to identify the best inputs and to get better results with each data set for RE and compare with other experiments, making ablation studies possible. The DeepREF Framework is evaluated on several reference corpora from various application domains.
Igor Nascimento, Rinaldo Lima, Adrian-Gabriel Chifu, Bernard Espinasse, Sébastien Fournier
LREC3
2021 FreSaDa: A French Satire Data Set for Cross-Domain Satire Detection
abstract
In this paper, we introduce FreSaDa, a French Satire Data Set11https://github.com/adrianchifu/FreSaDa, which is composed of 11,570 articles from the news domain. In order to avoid reporting unreasonably high accuracy rates due to the learning of characteristics specific to publication sources, we divided our samples into training, validation and test, such that the training publication sources are distinct from the validation and test publication sources. This gives rise to a cross-domain (cross-source) satire detection task. We employ two classification methods as baselines for our new data set, one based on low-level features (character n-grams) and one based on high-level features (average of CamemBERT word embeddings). As an additional contribution, we present an unsupervised domain adaptation method based on regarding the pairwise similarities (given by the dot product) between the training samples and the validation samples as features. By including these domain-specific features, we attain significant improvements for both character n-grams and CamemBERT embeddings.
Radu Tudor Ionescu, Adrian-Gabriel Chifu
IJCNN2
2020 DeepNLPF: A Framework for Integrating Third Party NLP Tools
abstract
Natural Language Processing (NLP) of textual data is usually broken down into a sequence of several subtasks, where the output of one the subtasks becomes the input to the following one, which constitutes an NLP pipeline. Many third-party NLP tools are currently available, each performing distinct NLP subtasks. However, it is difficult to integrate several NLP toolkits into a pipeline due to many problems, including different input/output representations or formats, distinct programming languages, and tokenization issues. This paper presents DeepNLPF, a framework that enables easy integration of third-party NLP tools, allowing the user to preprocess natural language texts at lexical, syntactic, and semantic levels. The proposed framework also provides an API for complete pipeline customization including the definition of input/output formats, integration plugin management, transparent ultiprocessing execution strategies, corpus-level statistics, and database persistence. Furthermore, the DeepNLPF user-friendly GUI allows its use even by a non-expert NLP user. We conducted runtime performance analysis showing that DeepNLPF not only easily integrates existent NLP toolkits but also reduces significant runtime processing compared to executing the same NLP pipeline in a sequential manner.
Francisco Rodrigues, Rinaldo Lima, William Domingues, Robson do Nascimento Fidalgo, Adrian-Gabriel Chifu, Bernard Espinasse, Sébastien Fournier
LREC5
2019 On the Use of Dependencies in Relation Classification of Text with Deep Learning
Bernard Espinasse, Sébastien Fournier, Adrian-Gabriel Chifu, Gaël Guibon, René Azcurra, Valentin Macé
CICLing (2)3
2018 Predicting Contradiction Intensity: Low, Strong or Very Strong?
abstract
Reviews on web resources (e.g. courses, movies) become increasingly exploited in text analysis tasks (e.g. opinion detection, controversy detection). This paper investigates contradiction intensity in reviews exploiting different features such as variation of ratings and variation of polarities around specific entities (e.g. aspects, topics). Firstly, aspects are identified according to the distributions of the emotional terms in the vicinity of the most frequent nouns in the reviews collection. Secondly, the polarity of each review segment containing an aspect is estimated. Only resources containing these aspects with opposite polarities are considered. Finally, some features are evaluated, using feature selection algorithms, to determine their impact on the effectiveness of contradiction intensity detection. The selected features are used to learn some state-of-the-art learning approaches. The experiments are conducted on the Massive Open Online Courses data set containing 2244 courses and their 73,873 reviews, collected from coursera.org. Results showed that variation of ratings, variation of polarities, and reviews quantity are the best predictors of contradiction intensity. Also, J48 was the most effective learning approach for this type of classification.
Ismail Badache, Sébastien Fournier, Adrian-Gabriel Chifu
SIGIR3
2018 Query Performance Prediction Focused on Summarized Letor Features
abstract
Query performance prediction (QPP) aims at automatically estimating the information retrieval system effectiveness for any user's query. Previous work has investigated several types of pre- and post-retrieval query performance predictors; the latter has been shown to be more effective. In this paper we investigate the use of features that were initially defined for learning to rank in the task of QPP. While these features have been shown to be useful for learning to rank documents, they have never been studied as query performance predictors. We developed more than 350 variants of them based on summary functions. Conducting experiments on four TREC standard collections, we found that Letor-based features appear to be better QPP than predictors from the literature. Moreover, we show that combining the best Letor features outperforms the state of the art query performance predictors. This is the first study that considers such an amount and variety of Letor features for QPP and that demonstrates they are appropriate for this task.
Adrian-Gabriel Chifu, Léa Laporte, Josiane Mothe, Md. Zia Ullah
SIGIR1
2017 Human-Based Query Difficulty Prediction
Adrian-Gabriel Chifu, Sébastien Déjean, Stefano Mizzaro, Josiane Mothe
ECIR1
2017 Harnessing Ratings and Aspect-Sentiment to Estimate Contradiction Intensity in Temporal-Related Reviews
abstract
Analysis of opinions (reviews) generated by users becomes increasingly exploited by a variety of applications. It allows to follow the evolution of the opinions or to carry out investigations on products. The detection of contradictory opinions about a web resource (e.g., courses, movies, products, etc.) is an important task to evaluate the latter. This paper focuses on the problem of detecting contradictions in reviews based on the sentiment analysis around specific aspects of a resource (document). In general, for web resources such as online courses (e.g. on Coursera or edX ), reviews are often generated during course sessions. Between each session users stop reviewing on the course, and this course may have updates. So, in order to avoid the confusion of contradictory reviews coming from two or more different sessions, the reviews related to a given resource should be firstly grouped according to their session. Secondly, certain aspects are extracted according to the distributions of the emotional terms in the vicinity of the most frequent names in the reviews collection. Thirdly, the polarity of each review segment containing an aspect is identified. Then taking only the resources containing these aspects with opposite polarities (positive, negative). Finally, we propose a measure of contradiction intensity based on the joint dispersion of the polarity and the rating of the reviews containing the aspects within each resource. The evaluation of our approach is conducted on the Massive Open Online Courses (MOOC) collection containing 2244 courses and their 73,873 reviews, collected from Coursera . The results of experiments revealed the effectiveness of the proposed approach to capture and quantify contradiction intensity.
Ismail Badache, Sébastien Fournier, Adrian-Gabriel Chifu
KES3
2016 SegChainW2V: Towards a Generic Automatic Video Segmentation Framework, Based on Lexical Chains of Audio Transcriptions and Word Embeddings
abstract
With the advances in multimedia broadcasting through a rich variety of channels and with the vulgarization of video production, it becomes essential to be able to provide reliable means of retrieving information within videos, not only the videos themselves. Research in this area has been widely focused on the context of TV news broadcasts, for which the structure itself provides clues for story segmentation. The systematic employment of these clues would lead to thematically driven systems that would not be easily adaptable in the case of videos of other types. The systems are therefore dependent on the type of videos for which they have been designed. In this paper we aim at introducing SegChainW2V, a generic unsupervised framework for story segmentation, based on lexical chains from transcriptions and their vectorization. SegChainW2V takes into account the topic changes by perceiving the fiuctuations of the most frequent terms throughout the video, as well as their semantics through the word embedding vectorization.
Adrian-Gabriel Chifu, Sébastien Fournier
KES1
2015 DeShaTo: Describing the Shape of Cumulative Topic Distributions to Rank Retrieval Systems Without Relevance Judgments
Radu Tudor Ionescu, Adrian-Gabriel Chifu, Josiane Mothe
SPIRE2
2015 Word sense discrimination in information retrieval: A spectral clustering-based approach
Adrian-Gabriel Chifu, Florentina Hristea, Josiane Mothe, Marius Popescu
Inf. Process. Manag.1