Christine Sinoquet

dblp:02/7019 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-6358-9420ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorTheory of computation · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Assessing Graph Neural Networks for latency and power consumption prediction in application mappings on multicore architectures
abstract
Accurately estimating the latency and power consumption of software applications deployed on multicore systems remains a major challenge for early-stage optimization, as existing methods typically rely on slow and resource-intensive simulations.This paper explores modeling application-to-architecture mappings as heterogeneous graphs and investigates Graph Neural Networks (GNNs) for predicting their performance.Four GNN models are evaluated across eleven datasets, considering five neural network-based software applications.The two best models achieve mean absolute percentage errors of about 2% for power prediction and 15% for latency, with prediction times of only a few tens of milliseconds.These results indicate the potential of GNN-based prediction as an efficient alternative to simulation-driven estimation, paving the way for early-stage AI-assisted mapping optimization.
Oscar Roussel, Zainab Ghrayeb, Sébastien Le Nours, Christine Sinoquet
ESANN4
2026 Extending Temporal Case-Based Reasoning for Action-Conditioned Time Series Prediction from Mixed Asynchronous Data: Application to Data-Driven Medical Simulation
Hugo Boisaubert, Lucas Vincent, Corinne Lejus-Bourdeau, Christine Sinoquet
ICAART (2)4
2026 Graph Neural Networks for Graph-Level Regression on Heterogeneous Network Data: Use Case in Early-Stage Optimization of Software Mapping on Multicore Platforms
Oscar Roussel, Zainab Ghrayeb, Sébastien Le Nours, Christine Sinoquet
IDA4
2025 Investigating four deep learning approaches as candidates for unified models in time series forecasting and event prediction: application in anesthesia training
abstract
This paper explores deep learning architectures for the purposes of unsupervised representation learning of hybrid asynchronous data, and joint prediction tasks.We aim to forecast short-term multivariate time series contextualized by events and to predict events contextualized by time series.Our proof-of-concept examines a real-world case of digitally assisted training in anesthesia.We evaluate four different architectures, using two strategies to integrate both time series and event sequences in the models.We assess the prediction quality of the models, and demonstrate that only one of the four architectures achieves performance outcomes compatible with our application objective.
Quentin Victor, Ianis Clavier, Hugo Boisaubert, Fabien Picarougne, Corinne Lejus-Bourdeau, Christine Sinoquet
ESANN6
2025 Two-in-One Models for Event Prediction and Time Series Forecasting. Comparison of Four Deep Learning Approaches to Simulate a Digital Patient Under Anesthesia
Quentin Victor, Ianis Clavier, Hugo Boisaubert, Fabien Picarougne, Corinne Lejus-Bourdeau, Christine Sinoquet
IDA6
2024 LSTM encoder-decoder model for contextualized time series forecasting applied to the simulation of a digital patient's physiological variables
abstract
This paper explores utilizing an encoder-decoder neural architecture for unsupervised representation learning of mixed asynchronous data, presenting the jmetts (Joint Modelling of Event Traces and Time Series) model.Our goal is to forecast short-term multivariate time series within event contexts.As a proof of concept, we examine a real-world case in digitally assisted training for anaesthesiology.jmetts demonstrates high predictive performance, with a maximum prediction error percentage of approximately 5.5%, comparable to that of its only competitor published to date.The source code can be found at https://github.com/ jp3142/jmetts_models_and_pipeline.
Julien Paris, Christine Sinoquet, Fadoua Taia-Alaoui, Corinne Lejus-Bourdeau
ESANN2
2024 Deep joint modelling of mixed asynchronous streams - Proof of concept for data-driven simulation of a digital patient under anaesthesia
abstract
In this paper, we investigate the use of an encoder-decoder neural architecture for unsupervised representation learning of mixed asynchronous data, and we introduce the JMETTS (Joint Modelling of Event Traces and Time Series) model. Our aim is to perform short-term forecasts for multivariate time series contextualized by events. As a proof of concept, we focus on a real-world case related to digitally assisted training in anaesthesiology. With a maximal prediction error percentage around 5.5%, the high predictive performance of JMETTS is comparable to its only competitor published to date. The source code is publicly available.
Julien Paris, Christine Sinoquet, Fadoua Taia-Alaoui, Corinne Lejus-Bourdeau
KES2
2023 A Framework for Context-Sensitive Prediction in Time Series - Feasibility Study for Data-Driven Simulation in Medicine
abstract
The need to comprehend and predict the dynamics of complex systems has spurred developments of time series forecasting methods across several disciplines. Nowadays, time series are collected with event logs for an ever increasing number of systems. Event logs can contain prominent information about the system dynamics. This paper addresses joint modelling of time series and event traces, to enhance time series forecasting. We introduce the Non-Homogeneous Markov Chain AutoRegressive model, to best apprehend the combined influence of past events belonging to several event categories, on a system’s dynamics. The originality of our proposal stems from the synchronization of a Hawkes temporal point process with the classical first-order hidden Markov model, through contextual variables. We also instantiate a basic version whose contextual variables only take into account events’ latest occurrences. Our proof-of-concept experiments address a real-world case related to digitally assisted training in anaesthesiology. We demonstrate that the advanced instantiation outperforms the basic instantiation (maximal prediction error percentage: 5.6% versus 13.5%). Finally, we validate the suitability of the advanced instantiation for the desired simulation, by applying real sequences of medical actions to digital patients. We show that we obtain highly realistic simulations.
Fatoumata Dama, Christine Sinoquet, Corinne Lejus-Bourdeau
DSAA2
2023 A hidden Markov model with Hawkes process-derived contextual variables to improve time series prediction. Case study in medical simulation
abstract
So far, models that take advantage of sequences of events to refine time series prediction have only been designed for specific applications.In this paper, we introduce the Non-Homogeneous Markov Chain AutoRegressive (NHMC-AR) model.In our model, the innovation arises from the synchronization of a multivariate Hawkes temporal point process with an autoregressive first-order hidden Markov model, through contextual variables.Experiments on anaesthesia data demonstrate that NHMC-AR has substantially better predictive performance compared to two competing methods.
Fatoumata Dama, Christine Sinoquet, Corinne Lejus-Bourdeau
ESANN2
2023 Partially Hidden Markov Chain Multivariate Linear Autoregressive model: inference and forecasting - application to machine health prognostics
abstract
Abstract Time series subject to regime shifts have attracted much interest in domains such as econometry, finance or meteorology. For discrete-valued regimes, models such as the popular Hidden Markov Chain (HMC) describe time series whose state process isunknownat all time-steps. Sometimes, time series are annotated. Thus, another category of models handles the case with regimesobservedat all time-steps. We present a novel model which addresses the intermediate case: (i) state processes associated to such time series are modelled by Partially Hidden Markov Chains (PHMCs); (ii) a multivariate linear autoregressive (MLAR) model drives the dynamics of the time series, within each regime. We describe a variant of the expectation maximization (EM) algorithm devoted to PHMC-MLAR model learning. We propose a hidden state inference procedure and a forecasting function adapted to the semi-supervised framework. We first assess inference and prediction performances, and analyze EM convergence times for PHMC-MLAR, using simulated data. We show the benefits of using partially observed states as well as a fully labelled scheme with unreliable labels, to decrease EM convergence times. We highlight the robustness of PHMC-MLAR to labelling errors in inference and prediction tasks. Finally, using turbofan engine data from a NASA repository, we show that PHMC-MLAR outperforms or largely outperforms other models: PHMC and MSAR (Markov Switching AutoRegressive model) for the feature prediction task, PHMC and five out of six recent state-of-the-art methods for the prediction of machine useful remaining life.
Fatoumata Dama, Christine Sinoquet
Mach. Learn.2
2021 Prediction and Inference in a Partially Hidden Markov-switching Framework with Autoregression. Application to Machinery Health Diagnosis
abstract
Time series subject to changes in regime are encountered in multiple applications. Models such as the renowned Hidden Markov Model (HMM) describe time series whose states are unknown at all time-steps. In some situations, partial knowledge on states is available. In this paper, we describe the Partially Hidden Markov Chain Linear AutoRegressive (PHMC-LAR) model. This model combines a HMM framework with local state-specific linear autoregressive dynamics. Namely, this hybrid model extends two published models, the Markov-Switching AutoRegressive (MSAR) model and the Partially Hidden Markov Chain (PHMC). Our contributions in this paper address: (i) Expectation-Maximization-based semi-supervised parameter learning, (ii) time series prediction, (iii) latent state inference via a variant of the Viterbi algorithm. We first validate our model on synthetic data. We show that integrating relatively limited knowledge on states considerably accelerates model training, while still preserving good prediction and state inference performances. Further, we compare the state inference performances of PHMC-LAR, PHMC and MSAR on realistic data with ground truth, in the context of a machine health diagnosis application. PHMC and PHMC-LAR show comparable performances, while PHMC-LAR is mainly subject to state degradation anticipation errors. This is a desirable property to ensure system safety through early maintenance operations. Thus PHMC-LAR illustrates a contribution to machine learning with potential beneficial impacts on safety in industrial, transportation and health applications.
Fatoumata Dama, Christine Sinoquet
ICTAI2
2018 Random Forest Framework Customized to Handle Highly Correlated Variables: An Extensive Experimental Study Applied to Feature Selection in Genetic Data
abstract
The random forest model is a popular framework used in classification and regression. In cases where high correlations exist within the data, it may be beneficial to capture these dependencies through latent variables, for an enhanced use of the random forest framework. In this paper, we present Sylva, the second proposal of a random forest with latent variables after T-Trees, derived from the seminal works of Botta and co-workers (Botta et al., 2008). Sylva is an innovative hybrid approach in which the dynamic generation of latent variables used to learn the random forest is driven by an additional forest model, this time a forest of latent tree models. The latter forest model, a class of Bayesian networks devised in (Mourad et al., 2011), allows a flexible modeling of the dependencies existing within the data. In the comprehensive study reported here, three variants of Sylva, instantiated by different clustering methods (CAST, DBSCAN, Louvain method), are compared to T-Trees using high-dimensional real-world datasets (161 datasets each describing around 5,000 observations and between 5,700 and 39,000 variables) in the context of genetic association studies. We show that T-Trees and Sylva have comparable high predictive powers (aeras under the ROC curves), that lie in range [0.887, 0.961] (T-Trees), and in interval [0.885, 0.979] (over the three Sylva instantiations). Interestingly, T-Trees and Sylva are shown to differ significantly in their importance measure distributions: in Sylva, the importance measure distribution corresponding to top ranked variables is significantly skewed towards higher values than in T-Trees, which meets the feature selection enhancement objective. This property holds true for the three instantiations of Sylva. In addition, the thorough analysis of the number of top-ranked variables jointly identified by T-Trees and Sylva highlights the possibility to cross-validate the findings, in order to constitute a priorized list of features (e.g., to be further analyzed by biologists, in the context of genetic association studies). Finally, we conclude that it is recommended to use CAST or DBSCAN, and not the Louvain method, on the 161 datasets analyzed, to increase the probability of Sylva to detect top variables missedby T-Trees among its top ranked variables.
Christine Sinoquet, Kamel Mekhnacha
DSAA1
2018 Combining latent tree modeling with a random forest-based approach, for genetic association studies
Christine Sinoquet, Kamel Mekhnacha
ESANN1
2018 Enhancement of a stochastic Markov-blanket framework with ant colony optimization, to uncover epistasis in genetic association studies
Christine Sinoquet, Clément Niel
ESANN1
2018 Random Forests with Latent Variables to Foster Feature Selection in the Context of Highly Correlated Variables. Illustration with a Bioinformatics Application
Christine Sinoquet, Kamel Mekhnacha
IDA1
2018 SMMB: a stochastic Markov blanket framework strategy for epistasis detection in GWAS
abstract
Motivation: Large scale genome-wide association studies (GWAS) are tools of choice for discovering associations between genotypes and phenotypes. To date, many studies rely on univariate statistical tests for association between the phenotype and each assayed single nucleotide polymorphism (SNP). However, interaction between SNPs, namely epistasis, must be considered when tackling the complexity of underlying biological mechanisms. Epistasis analysis at large scale entails a prohibitive computational burden when addressing the detection of more than two interacting SNPs. In this paper, we introduce a stochastic causal graph-based method, SMMB, to analyze epistatic patterns in GWAS data. Results: We present Stochastic Multiple Markov Blanket algorithm (SMMB), which combines both ensemble stochastic strategy inspired from random forests and Bayesian Markov blanket-based methods. We compared SMMB with three other recent algorithms using both simulated and real datasets. Our method outperforms the other compared methods for a majority of simulated cases of 2-way and 3-way epistasis patterns (especially in scenarii where minor allele frequencies of causal SNPs are low). Our approach performs similarly as two other compared methods for large real datasets, in terms of power, and runs faster. Availability and implementation: Parallel version available on https://ls2n.fr/listelogicielsequipe/DUKe/128/. Supplementary information: Supplementary data are available at Bioinformatics online.
Clément Niel, Christine Sinoquet, Christian Dina, Ghislain Rocheleau
Bioinform.2
2018 A method combining a random forest-based technique with the modeling of linkage disequilibrium through latent variables, to run multilocus genome-wide association studies
abstract
BACKGROUND: Genome-wide association studies (GWASs) have been widely used to discover the genetic basis of complex phenotypes. However, standard single-SNP GWASs suffer from lack of power. In particular, they do not directly account for linkage disequilibrium, that is the dependences between SNPs (Single Nucleotide Polymorphisms). RESULTS: We present the comparative study of two multilocus GWAS strategies, in the random forest-based framework. The first method, T-Trees, was designed by Botta and collaborators (Botta et al., PLoS ONE 9(4):e93379, 2014). We designed the other method, which is an innovative hybrid method combining T-Trees with the modeling of linkage disequilibrium. Linkage disequilibrium is modeled through a collection of tree-shaped Bayesian networks with latent variables, following our former works (Mourad et al., BMC Bioinformatics 12(1):16, 2011). We compared the two methods, both on simulated and real data. For dominant and additive genetic models, in either of the conditions simulated, the hybrid approach always slightly performs better than T-Trees. We assessed predictive powers through the standard ROC technique on 14 real datasets. For 10 of the 14 datasets analyzed, the already high predicted power observed for T-Trees (0.910-0.946) can still be increased by up to 0.030. We also assessed whether the distributions of SNPs' scores obtained from T-Trees and the hybrid approach differed. Finally, we thoroughly analyzed the intersections of top 100 SNPs output by any two or the three methods amongst T-Trees, the hybrid approach, and the single-SNP method. CONCLUSIONS: The sophistication of T-Trees through finer linkage disequilibrium modeling is shown beneficial. The distributions of SNPs' scores generated by T-Trees and the hybrid approach are shown statistically different, which suggests complementary of the methods. In particular, for 12 of the 14 real datasets, the distribution tail of highest SNPs' scores shows larger values for the hybrid approach. Thus are pinpointed more interesting SNPs than by T-Trees, to be provided as a short list of prioritized SNPs, for a further analysis by biologists. Finally, among the 211 top 100 SNPs jointly detected by the single-SNP method, T-Trees and the hybrid approach over the 14 datasets, we identified 72 and 38 SNPs respectively present in the top25s and top10s for each method.
Christine Sinoquet
BMC Bioinform.1
2013 A Survey on Latent Tree Models and Applications
abstract
In data analysis, latent variables play a central role because they help provide powerful insights into a wide variety of phenomena, ranging from biological to human sciences. The latent tree model, a particular type of probabilistic graphical models, deserves attention. Its simple structure - a tree - allows simple and efficient inference, while its latent variables capture complex relationships. In the past decade, the latent tree model has been subject to significant theoretical and methodological developments. In this review, we propose a comprehensive study of this model. First we summarize key ideas underlying the model. Second we explain how it can be efficiently learned from data. Third we illustrate its use within three types of applications: latent structure discovery, multidimensional clustering, and probabilistic inference. Finally, we conclude and give promising directions for future researches in this field.
Raphaël Mourad, Christine Sinoquet, Nevin Lianwen Zhang, Philippe Leray 0001
J. Artif. Intell. Res.2
2012 Probabilistic graphical models for genetic association studies
abstract
Probabilistic graphical models have been widely recognized as a powerful formalism in the bioinformatics field, especially in gene expression studies and linkage analysis. Although less well known in association genetics, many successful methods have recently emerged to dissect the genetic architecture of complex diseases. In this review article, we cover the applications of these models to the population association studies' context, such as linkage disequilibrium modeling, fine mapping and candidate gene studies, and genome-scale association studies. Significant breakthroughs of the corresponding methods are highlighted, but emphasis is also given to their current limitations, in particular, to the issue of scalability. Finally, we give promising directions for future research in this field.
Raphaël Mourad, Christine Sinoquet, Philippe Leray 0001
Briefings Bioinform.2
2011 A hierarchical Bayesian network approach for linkage disequilibrium modeling and data-dimensionality reduction prior to genome-wide association studies
abstract
BACKGROUND: Discovering the genetic basis of common genetic diseases in the human genome represents a public health issue. However, the dimensionality of the genetic data (up to 1 million genetic markers) and its complexity make the statistical analysis a challenging task. RESULTS: We present an accurate modeling of dependences between genetic markers, based on a forest of hierarchical latent class models which is a particular class of probabilistic graphical models. This model offers an adapted framework to deal with the fuzzy nature of linkage disequilibrium blocks. In addition, the data dimensionality can be reduced through the latent variables of the model which synthesize the information borne by genetic markers. In order to tackle the learning of both forest structure and probability distributions, a generic algorithm has been proposed. A first implementation of our algorithm has been shown to be tractable on benchmarks describing 105 variables for 2000 individuals. CONCLUSIONS: The forest of hierarchical latent class models offers several advantages for genome-wide association studies: accurate modeling of linkage disequilibrium, flexible data dimensionality reduction and biological meaning borne by latent variables.
Raphaël Mourad, Christine Sinoquet, Philippe Leray 0001
BMC Bioinform.2
2006 A Novel Approach for Structured Consensus Motif Inference Under Specificity and Quorum Constraints
Christine Sinoquet
APBC1
2006 Maximal sub-triangulation in pre-processing phylogenetic data
Anne Berry, Alain Sigayret, Christine Sinoquet
Soft Comput.3