VLDB 2026 Research / reviewers in the wild / expert
Olivier Le Van
dblp:354/7821
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0007-3081-3810ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conversational Transcription and Knowledge Graph Representation
Karima Boutalbi, Antoine Lévêque, Olivier Le Van, Rafika Boutalbi |
COMPSAC | 3 |
| 2026 | From Literature Overload to Knowledge Graph: An Automated Pipeline for Literature Reviews
Olivier Le Van, Pierre Dardouillet, Karima Boutalbi |
COMPSAC | 1 |
| 2025 | A Novel Assistant for Question-Answering from Training Video Sessions Using RAGabstractIn many organizations, internal training programs play a crucial role in helping employees learn new skills and use specific software tools. Training sessions are often delivered through various formats such as PDFs, slide decks, and videos. However, employees don’t always remember everything from training. Finding specific information later can take time and means digging through a lot of material. The problem gets worse when training content includes mostly diagrams and images with little text, making it hard to quickly find information. To address this problem, we propose an innovative pipeline that uses the oral content from training videos. Spoken explanations often provide more detail than text or images. By turning spoken content into organized text, we make training easier to access and use. This paper explores how to retrieve information from training videos using several modeling choices, including speech-to-text, LLMs, and embeddings. Our solution consists of a chatbot designed for company use that serves as a reliable assistant for employees, overcoming the limitations of traditional tutoring by providing personalized assistance. By integrating Natural Language Processing (NLP) techniques, our chatbot will assume an important role in supporting employees’ training, reducing the time spent searching for information. Quentin Kembellec, Karima Boutalbi, Olivier Le Van |
COMPSAC | 3 |
| 2025 | Towards a More Efficient Sinkhorn Distance Computation in Neural Topic ModelsabstractIn natural language processing, topic modeling aims to extract a corpora latent structure. In recent years, optimal transport distances have improved the topic extraction capabilities of Neural Topic Models (NTMs). More precisely, the Sinkhorn-Knopp algorithm is used to compute the blurred Wasserstein distance with relatively low complexity and is fully differentiable. This algorithm ease of implementation and advantages are thus particularly interesting for enforcing desired properties in NTMs. However, the algorithm can be unstable and inefficient under low blur setups, hence hindering overall topic model performances. In this article, we first assess the stability and efficiency of the Sinkhorn-Knopp algorithm in NTM scenarios. We compare five of the most relevant variations of this algorithm, and three distinct usages in NTMs. We evaluate each specific Sinkhorn-Knopp algorithm variation and topic model architecture independently, under various quantitative and qualitative metrics. Furthermore, we propose a novel method that focuses on the Sinkhorn-Knopp algorithm initialization, by reusing its dual variables from previous model updates as warm-start values. Our experiments reveal that our method can drastically improve the computation efficiency of the algorithm by reducing its number of iterations by up to 70%, and is easily applicable to any topic model using the Sinkhorn distance. Pierre Dardouillet, Kavé Salamatian, Hervé Verjus, Faiza Loukil, David Telisson, Olivier Le Van |
IJCNN | 6 |
| 2024 | IEcons: A New Consensus Approach Using Multi-Text Representations for Clustering TaskabstractToday we are able to generate a large set of text representations from the simple Bag-of-word (BOW) to the recent transformers capturing the semantic and the contextual text meaning. It was proven that there is no best text representation for text clustering task. Consequently, some works combined text representations using a consensus clustering approach. Two consensus approach types exist, namely explicit and implicit consensus. In the explicit consensus, also known asensemble clustering, the consensus function is applied a posterior after obtaining cluster labels from each text representation clustering allowing to capture global mutual information between the partitions of all text representations. On the other hand, implicit consensus uses tensor clustering to optimize the clustering consensus partition that deals with similarity matrices of text representations. Karima Boutalbi, Rafika Boutalbi, Hervé Verjus, Kavé Salamatian, David Telisson, Olivier Le Van |
CIKM | 6 |
| 2024 | Strategic Integration of Context for Fine-Tuning Topic Model PerformanceabstractIssue Tracking Systems software serves as an interface between a company and its customers. Customers can report bugs and seek assistance, among other demands. Reported issues include textual description, along with company defined metadata, aim at simplifying issue treatment by experts. In the context of the rapid growth of customer-reported issues, the manual treatment process becomes tedious and time-consuming. As a result, more and more studies focus on automating parts of this process, using semantic extraction and topic modeling approaches to automatically classify issues. To this end, most approaches consider the issue of textual description along with metadata, which can be a source of uncertainty and misleading in many real-world scenarios. Besides, knowledge from the company experts is often neglected. In this paper, we propose a general taxonomy of information incorporation into topic models. This aims to assemble all existing techniques, to further detect literature gaps. In addition, we propose a technique to incorporate expert knowledge into neural topic models. We evaluate our techniques and others in the literature on a real-world dataset coming from the JIRA software of a French HR management company. Results show a significant increase of more than 22% in classification performances when using expert knowledge, in addition to the issue textual description. The results validate our approach's effectiveness in improving the automatic classification of issues. Pierre Dardouillet, Kavé Salamatian, Hervé Verjus, Faiza Loukil, David Telisson, Olivier Le Van |
COMPSAC | 6 |