David Telisson

dblp:353/5698 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0003-1417-954XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Towards a More Efficient Sinkhorn Distance Computation in Neural Topic Models
abstract
In natural language processing, topic modeling aims to extract a corpora latent structure. In recent years, optimal transport distances have improved the topic extraction capabilities of Neural Topic Models (NTMs). More precisely, the Sinkhorn-Knopp algorithm is used to compute the blurred Wasserstein distance with relatively low complexity and is fully differentiable. This algorithm ease of implementation and advantages are thus particularly interesting for enforcing desired properties in NTMs. However, the algorithm can be unstable and inefficient under low blur setups, hence hindering overall topic model performances. In this article, we first assess the stability and efficiency of the Sinkhorn-Knopp algorithm in NTM scenarios. We compare five of the most relevant variations of this algorithm, and three distinct usages in NTMs. We evaluate each specific Sinkhorn-Knopp algorithm variation and topic model architecture independently, under various quantitative and qualitative metrics. Furthermore, we propose a novel method that focuses on the Sinkhorn-Knopp algorithm initialization, by reusing its dual variables from previous model updates as warm-start values. Our experiments reveal that our method can drastically improve the computation efficiency of the algorithm by reducing its number of iterations by up to 70%, and is easily applicable to any topic model using the Sinkhorn distance.
Pierre Dardouillet, Kavé Salamatian, Hervé Verjus, Faiza Loukil, David Telisson, Olivier Le Van
IJCNN5
2024 IEcons: A New Consensus Approach Using Multi-Text Representations for Clustering Task
abstract
Today we are able to generate a large set of text representations from the simple Bag-of-word (BOW) to the recent transformers capturing the semantic and the contextual text meaning. It was proven that there is no best text representation for text clustering task. Consequently, some works combined text representations using a consensus clustering approach. Two consensus approach types exist, namely explicit and implicit consensus. In the explicit consensus, also known asensemble clustering, the consensus function is applied a posterior after obtaining cluster labels from each text representation clustering allowing to capture global mutual information between the partitions of all text representations. On the other hand, implicit consensus uses tensor clustering to optimize the clustering consensus partition that deals with similarity matrices of text representations.
Karima Boutalbi, Rafika Boutalbi, Hervé Verjus, Kavé Salamatian, David Telisson, Olivier Le Van
CIKM5
2024 Strategic Integration of Context for Fine-Tuning Topic Model Performance
abstract
Issue Tracking Systems software serves as an interface between a company and its customers. Customers can report bugs and seek assistance, among other demands. Reported issues include textual description, along with company defined metadata, aim at simplifying issue treatment by experts. In the context of the rapid growth of customer-reported issues, the manual treatment process becomes tedious and time-consuming. As a result, more and more studies focus on automating parts of this process, using semantic extraction and topic modeling approaches to automatically classify issues. To this end, most approaches consider the issue of textual description along with metadata, which can be a source of uncertainty and misleading in many real-world scenarios. Besides, knowledge from the company experts is often neglected. In this paper, we propose a general taxonomy of information incorporation into topic models. This aims to assemble all existing techniques, to further detect literature gaps. In addition, we propose a technique to incorporate expert knowledge into neural topic models. We evaluate our techniques and others in the literature on a real-world dataset coming from the JIRA software of a French HR management company. Results show a significant increase of more than 22% in classification performances when using expert knowledge, in addition to the issue textual description. The results validate our approach's effectiveness in improving the automatic classification of issues.
Pierre Dardouillet, Kavé Salamatian, Hervé Verjus, Faiza Loukil, David Telisson, Olivier Le Van
COMPSAC5
2023 Machine Learning for Text Anomaly Detection: A Systematic Review
abstract
Anomaly detection is a common task in various domains, which has attracted significant research efforts in recent years. Existing reviews mainly focus on structured data, such as numerical or categorical data. Several studies treated review of anomaly detection in general on heterogeneous data or concerning a specific domain. However, anomaly detection on unstructured textual data is less treated. In this work, we target textual anomaly detection. Thus, we propose a systematic review of anomaly detection solutions in the text. To do so, we analyze the included papers in our survey in terms of anomaly detection types, feature extraction methods, and machine learning methods. We also introduce a web scrapping to collect papers from digital libraries and propose a clustering method to classify selected papers automatically. Finally, we compare the proposed automatic clustering approach with manual classification, and we show the interest of our contribution.
Karima Boutalbi, Faiza Loukil, Hervé Verjus, David Telisson, Kavé Salamatian
COMPSAC4