VLDB 2026 Research / reviewers in the wild / expert
Karima Boutalbi
dblp:353/5584
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0004-9571-3325ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conversational Transcription and Knowledge Graph Representation
Karima Boutalbi, Antoine Lévêque, Olivier Le Van, Rafika Boutalbi |
COMPSAC | 1 |
| 2026 | From Literature Overload to Knowledge Graph: An Automated Pipeline for Literature Reviews
Olivier Le Van, Pierre Dardouillet, Karima Boutalbi |
COMPSAC | 3 |
| 2025 | A Novel Assistant for Question-Answering from Training Video Sessions Using RAGabstractIn many organizations, internal training programs play a crucial role in helping employees learn new skills and use specific software tools. Training sessions are often delivered through various formats such as PDFs, slide decks, and videos. However, employees don’t always remember everything from training. Finding specific information later can take time and means digging through a lot of material. The problem gets worse when training content includes mostly diagrams and images with little text, making it hard to quickly find information. To address this problem, we propose an innovative pipeline that uses the oral content from training videos. Spoken explanations often provide more detail than text or images. By turning spoken content into organized text, we make training easier to access and use. This paper explores how to retrieve information from training videos using several modeling choices, including speech-to-text, LLMs, and embeddings. Our solution consists of a chatbot designed for company use that serves as a reliable assistant for employees, overcoming the limitations of traditional tutoring by providing personalized assistance. By integrating Natural Language Processing (NLP) techniques, our chatbot will assume an important role in supporting employees’ training, reducing the time spent searching for information. Quentin Kembellec, Karima Boutalbi, Olivier Le Van |
COMPSAC | 2 |
| 2024 | IEcons: A New Consensus Approach Using Multi-Text Representations for Clustering TaskabstractToday we are able to generate a large set of text representations from the simple Bag-of-word (BOW) to the recent transformers capturing the semantic and the contextual text meaning. It was proven that there is no best text representation for text clustering task. Consequently, some works combined text representations using a consensus clustering approach. Two consensus approach types exist, namely explicit and implicit consensus. In the explicit consensus, also known asensemble clustering, the consensus function is applied a posterior after obtaining cluster labels from each text representation clustering allowing to capture global mutual information between the partitions of all text representations. On the other hand, implicit consensus uses tensor clustering to optimize the clustering consensus partition that deals with similarity matrices of text representations. Karima Boutalbi, Rafika Boutalbi, Hervé Verjus, Kavé Salamatian, David Telisson, Olivier Le Van |
CIKM | 1 |
| 2023 | Machine Learning for Text Anomaly Detection: A Systematic ReviewabstractAnomaly detection is a common task in various domains, which has attracted significant research efforts in recent years. Existing reviews mainly focus on structured data, such as numerical or categorical data. Several studies treated review of anomaly detection in general on heterogeneous data or concerning a specific domain. However, anomaly detection on unstructured textual data is less treated. In this work, we target textual anomaly detection. Thus, we propose a systematic review of anomaly detection solutions in the text. To do so, we analyze the included papers in our survey in terms of anomaly detection types, feature extraction methods, and machine learning methods. We also introduce a web scrapping to collect papers from digital libraries and propose a clustering method to classify selected papers automatically. Finally, we compare the proposed automatic clustering approach with manual classification, and we show the interest of our contribution. Karima Boutalbi, Faiza Loukil, Hervé Verjus, David Telisson, Kavé Salamatian |
COMPSAC | 1 |