Séverine Affeldt

dblp:147/0407 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
4since 2021 · last 2025
0000-0002-4107-0887ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Cluster Insight: A Weighted Clustering Tool for Large Textual Data Exploration
abstract
International audience
Amine Ferdjaoui, Séverine Affeldt, Mohamed Nadif
WSDM2
2024 WordGraph: A Python Package for Reconstructing Interactive Causal Graphical Models from Text Data
abstract
We present WordGraph, a Python package for exploring the topics of documents corpora. WordGraph provides causal graphical models from text data vocabulary and proposes interactive visualizations of terms networks. Our ease-to-use package is provided with a pre-built pipeline to access the main modules through jupyter widgets. It results in the encapsulation of a whole vocabulary exploration process within a single jupyter notebook cell, with straightforward parameters settings and interactive plots. WordGraph pipeline is fully customizable by adding/removing widgets or changing default parameters. To assist users with no background in Python nor jupyter notebook, but willing to explore large corpora topics, we also propose an automatic dashboard generation from the customizable jupyter notebook pipeline in a web application style. WordGraph is available through a GitHub repository at https://github.com/MLDS-software/WordGraph.
Amine Ferdjaoui, Séverine Affeldt, Mohamed Nadif
WSDM2
2022 An effective strategy for churn prediction and customer profiling
Louis Geiler, Séverine Affeldt, Mohamed Nadif
Data Knowl. Eng.2
2021 Regularized Dual-PPMI Co-clustering for Text Data
abstract
Co-clustering of document-term matrices has proved to be more effective than one-sided clustering. By their nature, text data are also generally unbalanced and directional. Recently, the von Mises-Fisher (vMF) mixture model was proposed to handle unbalanced data while harnessing the directional nature of text. In this paper we propose a novel co-clustering approach based on a matrix formulation of vMF model-based co-clustering. This formulation leads to a flexible method for text co-clustering that can easily incorporate both word-word semantic relationships and document-document similarities. By contrast with existing methods, which generally use an additive incorporation of similarities, we propose a dual multiplicative regularization that better encapsulates the underlying text data structure. Extensive evaluations on various real-world text datasets demonstrate the superior performance of our proposed approach over baseline and competitive methods, both in terms of clustering results and co-cluster topic coherence.
Séverine Affeldt, Lazhar Labiod, Mohamed Nadif
SIGIR1
2020 Ensemble Block Co-clustering: A Unified Framework for Text Data
abstract
In this paper, we propose a unified framework for Ensemble Block Co-clustering (EBCO), which aims to fuse multiple basic co-clusterings into a consensus structured affinity matrix. Each co-clustering to be fused is obtained by applying a co-clustering method on the same document-term dataset. This fusion process reinforces the individual quality of the multiple basic data co-clusterings within a single consensus matrix. Besides, the proposed framework enables a completely unsupervised co-clustering where the number of co-clusters is automatically inferred based on the non trivial generalized modularity. We first define an explicit objective function which allows the joint learning of the basic co-clusterings aggregation and the consensus block co-clustering. Then, we show that EBCO generalizes the one side ensemble clustering to an ensemble block co-clustering context. We also establish theoretical equivalence to spectral co-clustering and weighted double spherical k-means clustering for textual data. Experimental results on various real-world document-term datasets demonstrate that EBCO is an efficient competitor to some state-of-the-art ensemble and co-clustering methods.
Séverine Affeldt, Lazhar Labiod, Mohamed Nadif
CIKM1
2020 Spectral clustering via ensemble deep autoencoder learning (SC-EDAE)
Séverine Affeldt, Lazhar Labiod, Mohamed Nadif
Pattern Recognit.1
2018 MIIC online: a web server to reconstruct causal or non-causal networks from non-perturbative data
abstract
Summary: We present a web server running the MIIC algorithm, a network learning method combining constraint-based and information-theoretic frameworks to reconstruct causal, non-causal or mixed networks from non-perturbative data, without the need for an a priori choice on the class of reconstructed network. Starting from a fully connected network, the algorithm first removes dispensable edges by iteratively subtracting the most significant information contributions from indirect paths between each pair of variables. The remaining edges are then filtered based on their confidence assessment or oriented based on the signature of causality in observational data. MIIC online server can be used for a broad range of biological data, including possible unobserved (latent) variables, from single-cell gene expression data to protein sequence evolution and outperforms or matches state-of-the-art methods for either causal or non-causal network reconstruction. Availability and implementation: MIIC online can be freely accessed at https://miic.curie.fr. Supplementary information: Supplementary data are available at Bioinformatics online.
Nadir Sella, Louis Verny, Guido Uguzzoni, Séverine Affeldt, Hervé Isambert
Bioinform.4
2017 Efficient global network learning from local reconstructions
abstract
Discovering complex interactions is an important issue in numerous fields ranging from social sciences to systems biology. Over the past few decades, many network learning methods have exhibited competitive results on various types of data. A commonly reached conclusion is that some learning approaches are more advisable than others depending on the dataset type or the complexity of the underlying network. Another frequently encountered issue relates to the ever increasing number of variables that need to be simultaneously dealt with, especially when only a small number of observations is available. The ScaleNet, a novel reconstruction method which can embed different types of network discovery approaches within a spectral framework for large graphical model was introduced recently. The approach identifies sets of connected variables based on the magnitude and sign of the eigenvector elements of a normalized graph Laplacian matrix, and it learns in parallel multiple relevant sub-graphs of a large network. However, the number of eigenvectors to be used and the size of the sub-graphs are to be fixed by an expensive procedure of cross-validation. In this contribution, we propose heuristics to find both optimal number of eigenvectors and the number of nodes in the sub-networks. We illustrate by the results on standard large-scale data sets and on a real human gut graph reconstruction that the proposed approaches save computational time, i.e. are efficient, and reach the state-of-the-art performance.
Séverine Affeldt, Nataliya Sokolovska, Edi Prifti, Jean-Daniel Zucker
IJCNN1
2017 Learning causal networks with latent variables from multivariate information in genomic data
abstract
Learning causal networks from large-scale genomic data remains challenging in absence of time series or controlled perturbation experiments. We report an information- theoretic method which learns a large class of causal or non-causal graphical models from purely observational data, while including the effects of unobserved latent variables, commonly found in many genomic datasets. Starting from a complete graph, the method iteratively removes dispensable edges, by uncovering significant information contributions from indirect paths, and assesses edge-specific confidences from randomization of available data. The remaining edges are then oriented based on the signature of causality in observational data. The approach and associated algorithm, miic, outperform earlier methods on a broad range of benchmark networks. Causal network reconstructions are presented at different biological size and time scales, from gene regulation in single cells to whole genome duplication in tumor development as well as long term evolution of vertebrates. Miic is publicly available at https://github.com/miicTeam/MIIC.
Louis Verny, Nadir Sella, Séverine Affeldt, Param Priya Singh, Hervé Isambert
PLoS Comput. Biol.3
2016 Spectral consensus strategy for accurate reconstruction of large biological networks
abstract
BACKGROUND: The last decades witnessed an explosion of large-scale biological datasets whose analyses require the continuous development of innovative algorithms. Many of these high-dimensional datasets are related to large biological networks with few or no experimentally proven interactions. A striking example lies in the recent gut bacterial studies that provided researchers with a plethora of information sources. Despite a deeper knowledge of microbiome composition, inferring bacterial interactions remains a critical step that encounters significant issues, due in particular to high-dimensional settings, unknown gut bacterial taxa and unavoidable noise in sparse datasets. Such data type make any a priori choice of a learning method particularly difficult and urge the need for the development of new scalable approaches. RESULTS: We propose a consensus method based on spectral decomposition, named Spectral Consensus Strategy, to reconstruct large networks from high-dimensional datasets. This novel unsupervised approach can be applied to a broad range of biological networks and the associated spectral framework provides scalability to diverse reconstruction methods. The results obtained on benchmark datasets demonstrate the interest of our approach for high-dimensional cases. As a suitable example, we considered the human gut microbiome co-presence network. For this application, our method successfully retrieves biologically relevant relationships and gives new insights into the topology of this complex ecosystem. CONCLUSIONS: The Spectral Consensus Strategy improves prediction precision and allows scalability of various reconstruction methods to large networks. The integration of multiple reconstruction algorithms turns our approach into a robust learning method. All together, this strategy increases the confidence of predicted interactions from high-dimensional datasets without demanding computations.
Séverine Affeldt, Nataliya Sokolovska, Edi Prifti, Jean-Daniel Zucker
BMC Bioinform.1
2016 3off2: A network reconstruction algorithm based on 2-point and 3-point information statistics
abstract
BACKGROUND: The reconstruction of reliable graphical models from observational data is important in bioinformatics and other computational fields applying network reconstruction methods to large, yet finite datasets. The main network reconstruction approaches are either based on Bayesian scores, which enable the ranking of alternative Bayesian networks, or rely on the identification of structural independencies, which correspond to missing edges in the underlying network. Bayesian inference methods typically require heuristic search strategies, such as hill-climbing algorithms, to sample the super-exponential space of possible networks. By contrast, constraint-based methods, such as the PC and IC algorithms, are expected to run in polynomial time on sparse underlying graphs, provided that a correct list of conditional independencies is available. Yet, in practice, conditional independencies need to be ascertained from the available observational data, based on adjustable statistical significance levels, and are not robust to sampling noise from finite datasets. RESULTS: We propose a more robust approach to reconstruct graphical models from finite datasets. It combines constraint-based and Bayesian approaches to infer structural independencies based on the ranking of their most likely contributing nodes. In a nutshell, this local optimization scheme and corresponding 3off2 algorithm iteratively "take off" the most likely conditional 3-point information from the 2-point (mutual) information between each pair of nodes. Conditional independencies are thus derived by progressively collecting the most significant indirect contributions to all pairwise mutual information. The resulting network skeleton is then partially directed by orienting and propagating edge directions, based on the sign and magnitude of the conditional 3-point information of unshielded triples. The approach is shown to outperform both constraint-based and Bayesian inference methods on a range of benchmark networks. The 3off2 approach is then applied to the reconstruction of the hematopoiesis regulation network based on recent single cell expression data and is found to retrieve more experimentally ascertained regulations between transcription factors than with other available methods. CONCLUSIONS: The novel information-theoretic approach and corresponding 3off2 algorithm combine constraint-based and Bayesian inference methods to reliably reconstruct graphical models, despite inherent sampling noise in finite datasets. In particular, experimentally verified interactions as well as novel predicted regulations are established on the hematopoiesis regulatory networks based on single cell expression data.
Séverine Affeldt, Louis Verny, Hervé Isambert
BMC Bioinform.1
2015 Robust reconstruction of causal graphical models based on conditional 2-point and 3-point information
Séverine Affeldt, Hervé Isambert
UAI1
2014 On the expansion of "dangerous" gene families in vertebrates
abstract
“Dangerous” gene families, defined as prone to dominant (gain-of-function) mutations, have been greatly expanded in the course of vertebrate evolution by contrast to gene families more prone to recessive (loss-of-function) mutations. While the maintenance of “essential” genes is ensured by their lethal double null mutations, the expansion of “dangerous” gene families, implicated in cancer and other severe genetic diseases in human, remains puzzling. Could gene susceptibility to dominant deleterious mutations be somehow responsible for this striking evolutionary expansion of “dangerous” gene families? We proposed such an evolutionary model suggesting that this counterintuitive expansion of “dangerous” gene families is in fact a consequence of their susceptibility to deleterious mutations and purifying selection in polyploid species that arose from two rounds of whole genome duplication (WGD) events dating back from the onset of jawed vertebrates, some 500MY ago [ 1 , 2 ]. All WGD duplicates, so-called "ohnologs", were thus initially acquired by speciation without the need to provide evolutionary benefit to be fixed in post-WGD species. Our data mining analyses, based on the 20,506 human protein coding genes, first revealed a strong correlation between the retention of ohnologs and their susceptibility to dominant deleterious mutations in humans [ 3 ]. It appears that the human genes associated with the occurrence of cancer and other genetic diseases (8,095) have retained significantly more ohnologs than expected by chance (48% versus 35%; 48% : 3,844/8,095; P=1.3×10 , χ ). We also found that the retention of ohnologs is more strongly related to their “dangerousness” than their “essentiality” [ 3 ]. To go beyond mere correlations, we also performed mediation analyses, following the approach of Pearl [ 4 ], and quantified the direct and indirect effects of many genomic properties, such as essentiality, expression levels or divergence rates, on the retention of ohnologs. This enabled us to investigate an alternative hypothesis frequently invoked to account for the biased retention of ohnologs, namely the “dosage-balance” hypothesis [ 5 ]. While this hypothesis posits that the ohnologs are retained because their interactions with protein partners require to maintain balanced expression levels throughout evolution, we found that most of the ohnologs have in fact been eliminated from permanent complexes in human (7.5% versus 35%; 7.5% : 18/239; P=1.2×10 , χ ). These mediation analyses also showed (Fig. 1 ) that the gene susceptibility to deleterious mutations is more relevant than dosage-balance for the retention of ohnologs in more transient complexes. Quantitative Mediation analysis of direct versus indirect effects of deleterious mutations and dosage balance on the retention of human ohnologs using (A) all human protein coding genes (20,506) or (B) human protein coding genes without SSD nor CNV (8,215). The thickness of the arrows outlines the relative importance of the corresponding direct or indirect effects. Dir.< 0 or Ind.<0 corresponds to an anticorrelated direct or indirect effect, respectively. Gene prone to deleterious mutations and/or to dosage balance (including haploinsufficient genes and genes involved in multiprotein complexes) are taken from [ 3 ]. These results suggest that the retention of human ohnologs is primarily caused by their susceptibility to deleterious mutations. They further establish that the retention of many ohnologs suspected to be dosage balanced is in fact indirectly mediated by their susceptibility to dominant deleterious mutations. All in all, this supports a new evolutionary model relying on a non-adaptive mechanism that hinges on ( i ) the speciation event concomitant to WGD, and ( ii ) the dominance of deleterious mutations leading to purifying selection in post-WGD species.
Séverine Affeldt, Param Priya Singh, Giulia Malaguti, Hervé Isambert
BMC Bioinform.1
2014 Human Dominant Disease Genes Are Enriched in Paralogs Originating from Whole Genome Duplication
abstract
PLOS Computational Biology recently published an article by Chen, Zhao, van Noort, and Bork [1] reporting that, in contrast to duplicated nondisease genes, human monogenic disease (MD) genes are (1) enriched in duplicates (in agreement with earlier reports [2]–[5]) and (2) more functionally similar to their closest paralogs based on sequence conservation and expression profile similarity. Chen et al. then proposed that human MD genes frequently have functionally redundant paralogs that can mask the phenotypic effects of deleterious mutations.
Param Priya Singh, Séverine Affeldt, Giulia Malaguti, Hervé Isambert
PLoS Comput. Biol.2