Mira Ait Saada

dblp:289/8195 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2023
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2023 Unsupervised Anomaly Detection in Multi-Topic Short-Text Corpora
abstract
Unsupervised anomaly detection seeks to identify deviant data samples in a dataset without using labels and constitutes a challenging task, particularly when the majority class is heterogeneous.This paper addresses this topic for textual data and aims to determine whether a text sample is an outlier within a potentially multi-topic corpus.To this end, it is crucial to grasp the semantic aspects of words, particularly when dealing with short texts, since it is difficult to syntactically discriminate data samples based only on a few words.Thereby we make use of word embeddings to represent each sample by a dense vector, efficiently capturing the underlying semantics.Then, we rely on the Mixture Model approach to detect which samples deviate the most from the underlying distributions of the corpus.Experiments carried out on real datasets show the effectiveness of the proposed approach in comparison to state-ofthe-art techniques both in terms of performance and time efficiency, especially when more than one topic is present in the corpus.
Mira Ait Saada, Mohamed Nadif
EACL1
2023 Contextual Word Embeddings Clustering Through Multiway Analysis: A Comparative Study
Mira Ait Saada, Mohamed Nadif
IDA1
2022 Tensor-based Graph Modularity for Text Data Clustering
abstract
Graphs are used in several applications to represent similarities between instances. For text data, we can represent texts by different features such as bag-of-words, static embeddings (Word2vec, GloVe, etc.), and contextual embeddings (BERT, RoBERTa, etc.), leading to multiple similarities (or graphs) based on each representation. The proposal posits that incorporating the local invariance within every graph and the consistency across different graphs leads to a consensus clustering that improves the document clustering. This problem is complex and challenged with the sparsity and the noisy data included in each graph. To this end, we rely on the modularity metric, which effectively evaluates graph clustering in such circumstances. Therefore, we present a novel approach for text clustering based on both a sparse tensor representation and graph modularity. This leads to cluster texts (nodes) while capturing information arising from the different graphs. We iteratively maximize a Tensor-based Graph Modularity criterion. Extensive experiments on benchmark text clustering datasets are performed, showing that the proposed algorithm referred to as Tensor Graph Modularity -TGM- outperforms other baseline methods in terms of clustering task. The source code is available at https://github.com/TGMclustering/TGMclustering.
Rafika Boutalbi, Mira Ait Saada, Anastasiia Iurshina, Steffen Staab, Mohamed Nadif
SIGIR2
2021 How to Leverage a Multi-layered Transformer Language Model for Text Clustering: an Ensemble Approach
abstract
Pre-trained Transformer-based word embeddings are now widely used in text mining where they are known to significantly improve supervised tasks such as text classification, named entity recognition and question answering. Since the Transformer models create several different embeddings for the same input, one at each layer of their architecture, various studies have already tried to identify those of these embeddings that most contribute to the success of the above-mentioned tasks. In contrast the same performance analysis has not yet been carried out in the unsupervised setting. In this paper we evaluate the effectiveness of Transformer models on the important task of text clustering. In particular, we present a clustering ensemble approach that harnesses all the network's layers. Numerical experiments carried out on real datasets with different Transformer models show the effectiveness of the proposed method compared to several baselines.
Mira Ait Saada, François Role, Mohamed Nadif
CIKM1
2021 Unsupervised Methods for the Study of Transformer Embeddings
Mira Ait Saada, François Role, Mohamed Nadif
IDA1