VLDB 2026 Research / reviewers in the wild / expert
Pierre Dupont
dblp:63/2403
· DBLP profile ↗
42ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-4835-6519ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 10 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Software engineering, systems software and programming languages · 6Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 96% Computational science and engineering · 4% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Artificial intelligence
3 papers |
Probabilistic and Bayesian machine learning · 78% Representation and self-supervised learning · 16% Trustworthy machine learning · 6% | |
| Software engineering, system software, and programming languages
1 paper |
Requirements engineering and software design · 91% Empirical software engineering · 9% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › gene expression analysis
biclustering |
1.3 | 2 | 2024 | scCross: efficient search for rare subpopulations across multiple single-cell samples · Bioinform. 2024 MicroCellClust: mining rare and highly specific subpopulations from single-cell expression data · Bioinform. 2021 |
Bioinformatics and computational biology
single-cell analysis |
1.3 | 2 | 2024 | scCross: efficient search for rare subpopulations across multiple single-cell samples · Bioinform. 2024 MicroCellClust: mining rare and highly specific subpopulations from single-cell expression data · Bioinform. 2021 |
Data mining › interpretable machine learning
feature importance estimation |
0.5 | 1 | 2021 | An Importance Weighted Feature Selection Stability Measure · J. Mach. Learn. Res. 2021 |
Data mining › dimensionality reduction
feature selection |
0.5 | 1 | 2021 | An Importance Weighted Feature Selection Stability Measure · J. Mach. Learn. Res. 2021 |
Data mining › dimensionality reduction › feature selection
stable feature selection |
0.5 | 1 | 2021 | An Importance Weighted Feature Selection Stability Measure · J. Mach. Learn. Res. 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian model selection
bayesian feature selection |
0.2 | 1 | 2013 | Generalized spike-and-slab priors for Bayesian group feature selection using expectation propagation · J. Mach. Learn. Res. 2013 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › sparse bayesian learning
spike-and-slab prior |
0.2 | 1 | 2013 | Generalized spike-and-slab priors for Bayesian group feature selection using expectation propagation · J. Mach. Learn. Res. 2013 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process classification |
0.1 | 1 | 2011 | Robust Multi-Class Gaussian Process Classification · NIPS 2011 |
Bioinformatics and computational biology
biomarker discovery |
0.1 | 1 | 2010 | Robust biomarker identification for cancer diagnosis with ensemble feature selection methods · Bioinform. 2010 |
Computational science and engineering › ensemble learning
ensemble feature selection |
0.1 | 1 | 2010 | Robust biomarker identification for cancer diagnosis with ensemble feature selection methods · Bioinform. 2010 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
metabolic pathway analysis |
0.1 | 1 | 2010 | Pathway discovery in metabolic networks by subgraph extraction · Bioinform. 2010 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
pathway discovery |
0.1 | 1 | 2010 | Pathway discovery in metabolic networks by subgraph extraction · Bioinform. 2010 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
0.1 | 1 | 2009 | Partially supervised feature selection with regularized linear models · ICML 2009 |
Requirements engineering and software design › model-driven engineering › model synthesis
behavior model synthesis |
0.1 | 1 | 2005 | Generating Annotated Behavior Models from End-User Scenarios · IEEE Trans. Software Eng. 2005 |
Requirements engineering and software design
requirements elicitation |
0.1 | 1 | 2005 | Generating Annotated Behavior Models from End-User Scenarios · IEEE Trans. Software Eng. 2005 |
Requirements engineering and software design › requirements elicitation
scenario-based elicitation |
0.1 | 1 | 2005 | Generating Annotated Behavior Models from End-User Scenarios · IEEE Trans. Software Eng. 2005 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
robustness to label noise |
0.0 | 1 | 2011 | Robust Multi-Class Gaussian Process Classification · NIPS 2011 |
Bioinformatics and computational biology › gene expression analysis
gene selection |
0.0 | 1 | 2009 | Partially supervised feature selection with regularized linear models · ICML 2009 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.0 | 1 | 2009 | Partially supervised feature selection with regularized linear models · ICML 2009 |
Automata and formal languages
grammatical inference |
0.0 | 1 | 2000 | Probabilistic DFA Inference using Kullback-Leibler Divergence and Minimality · ICML 2000 |
Information theory › information measures › divergence measures
kullback-leibler divergence |
0.0 | 1 | 2000 | Probabilistic DFA Inference using Kullback-Leibler Divergence and Minimality · ICML 2000 |
Computer animation and physical simulation › contact simulation
frictional contact |
0.0 | 1 | 1998 | Analysis of Frictional Contact Models for Dynamic Simulation · ICRA 1998 |
Computer animation and physical simulation
rigid body simulation |
0.0 | 1 | 1998 | Analysis of Frictional Contact Models for Dynamic Simulation · ICRA 1998 |
Empirical software engineering › software evaluation
model validation |
0.0 | 1 | 2005 | Generating Annotated Behavior Models from End-User Scenarios · IEEE Trans. Software Eng. 2005 |
Methods — techniques the papers use, named apart from their topics
global sum criterion · 0.8biclustering · 0.8stability measure · 0.5max-sum submatrix · 0.5feature weighting · 0.5constrained optimization · 0.5support vector machine · 0.3expectation propagation · 0.3sparsity regularization · 0.2regularized linear model · 0.2group feature selection · 0.2gaussian process · 0.1shortest path algorithm · 0.1random walk · 0.1ensemble methods · 0.1labeled transition system synthesis · 0.1grammar induction · 0.1minimality · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Detection of Large Language Model Contamination with Tabular Data
Benoît Ronval, Pierre Dupont, Siegfried Nijssen |
IDA | 2 |
| 2024 | scCross: efficient search for rare subpopulations across multiple single-cell samplesabstractMOTIVATION: Identifying rare cell types is an important task to capture the heterogeneity of single-cell data, such as scRNA-seq. The widespread availability of such data enables to aggregate multiple samples, corresponding for example to different donors, into the same study. Yet, such aggregated data is often subject to batch effects between samples. Clustering it therefore generally requires the use of data integration methods, which can lead to overcorrection, making the identification of rare cells difficult. We present scCross, a biclustering method identifying rare subpopulations of cells present across multiple single-cell samples. It jointly identifies a group of cells with specific marker genes by relying on a global sum criterion, computed over entire subpopulation of cells, rather than pairwise comparisons between individual cells. This proves robust with respect to the high variability of scRNA-seq data, in particular batch effects. RESULTS: We show through several case studies that scCross is able to identify rare subpopulations across multiple samples without performing prior data integration. Namely, it identifies a cilium subpopulation with potential new ciliary genes from lung cancer cells, which is not detected by typical alternatives. It also highlights rare subpopulations in human pancreas samples sequenced with different protocols, despite visible shifts in expression levels between batches. We further show that scCross outperforms typical alternatives at identifying a target rare cell type in a controlled experiment with artificially created batch effects. This shows the ability of scCross to efficiently identify rare cell subpopulations characterized by specific genes despite the presence of batch effects. AVAILABILITY AND IMPLEMENTATION: The R and Scala implementation of scCross is freely available on GitHub, at https://github.com/agerniers/scCross/. A snapshot of the code and the data underlying this article are available on Zenodo, at https://zenodo.org/doi/10.5281/zenodo.10471063. Alexander Gerniers, Siegfried Nijssen, Pierre Dupont |
Bioinform. | 3 |
| 2022 | MicroCellClust 2: a hybrid approach for multivariate rare cell mining in large-scale single-cell dataabstractIdentifying rare subpopulations in single-cell data is a key aspect when analyzing its heterogeneity. With large datasets now commonly generated, the focus went to scalability when designing rare cell mining methods, often relying on univariate approaches. Yet, MicroCellClust, an approach based on a multivariate optimization problem, has proven effective to jointly identify rare cells and specific genes in small-scale data. The proposed solver had a quadratic complexity, posing a practical limit to analyzing small or middle-scale data. Here, we present a new approach that scales MicroCellClust to larger datasets. It first performs a beam search among cells that are identified as rare to find an initial approximation. Then it uses simulated annealing, a classical derivative-free optimization algorithm which efficiently approaches the optimal solution. MicroCellClust 2 has a linear complexity in terms of the number of cells, which makes it scalable to large data (typically containing over 100000 cells). Our experiments report the identification of rare megakaryocytes within 68000 PBMCs, and rare ependymal cells within 160000 mouse brain cells. These results show that MicroCellClust 2 is more effective at identifying a subpopulation as a whole than typical alternatives, demonstrating the usefulness of jointly selecting cells and genes as opposed to other approaches. Alexander Gerniers, Pierre Dupont |
BIBM | 2 |
| 2021 | Robust Selection Stability Estimation in Correlated Spaces
Victor Hamer, Pierre Dupont |
ECML/PKDD (3) | 2 |
| 2021 | MicroCellClust: mining rare and highly specific subpopulations from single-cell expression dataabstractMOTIVATION: Identifying rare subpopulations of cells is a critical step in order to extract knowledge from single-cell expression data, especially when the available data is limited and rare subpopulations only contain a few cells. In this paper, we present a data mining method to identify small subpopulations of cells that present highly specific expression profiles. This objective is formalized as a constrained optimization problem that jointly identifies a small group of cells and a corresponding subset of specific genes. The proposed method extends the max-sum submatrix problem to yield genes that are, for instance, highly expressed inside a small number of cells, but have a low expression in the remaining ones. RESULTS: We show through controlled experiments on scRNA-seq data that the MicroCellClust method achieves a high F1 score to identify rare subpopulations of artificially planted human T cells. The effectiveness of MicroCellClust is confirmed as it reveals a subpopulation of CD4 T cells with a specific phenotype from breast cancer samples, and a subpopulation linked to a specific stage in the cell cycle from breast cancer samples as well. Finally, three rare subpopulations in mouse embryonic stem cells are also identified with MicroCellClust. These results illustrate the proposed method outperforms typical alternatives at identifying small subsets of cells with highly specific expression profiles. AVAILABILITYAND IMPLEMENTATION: The R and Scala implementation of MicroCellClust is freely available on GitHub, at https://github.com/agerniers/MicroCellClust/ The data underlying this article are available on Zenodo, at https://dx.doi.org/10.5281/zenodo.4580332. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexander Gerniers, Orian Bricard, Pierre Dupont |
Bioinform. | 3 |
| 2021 | An Importance Weighted Feature Selection Stability MeasureabstractCurrent feature selection methods, especially applied to high dimensional data, tend to suffer from instability since marginal modifications in the data may result in largely distinct selected feature sets. Such instability strongly limits a sound interpretation of the selected variables by domain experts. Defining an adequate stability measure is also a research question. In this work, we propose to incorporate into the stability measure the importances of the selected features in predictive models. Such feature importances are directly proportional to feature weights in a linear model. We also consider the generalization to a non-linear setting. We illustrate, theoretically and experimentally, that current stability measures are subject to undesirable behaviors, for example, when they are jointly optimized with predictive accuracy. Results on micro-array and mass-spectrometric data show that our novel stability measure corrects for overly optimistic stability estimates in such a bi-objective context, which leads to improved decision-making. It is also shown to be less prone to the under- or over-estimation of the stability value in feature spaces with groups of highly correlated variables. Victor Hamer, Pierre Dupont |
J. Mach. Learn. Res. | 2 |
| 2020 | Joint optimization of predictive performance and selection stability
Victor Hamer, Pierre Dupont |
ESANN | 2 |
| 2019 | The Maximum Weighted Submatrix Coverage Problem: A CP Approach
Guillaume Derval, Vincent Branders, Pierre Dupont, Pierre Schaus |
CPAIOR | 3 |
| 2019 | Mining a Maximum Weighted Set of Disjoint Submatrices
Vincent Branders, Guillaume Derval, Pierre Schaus, Pierre Dupont |
DS | 4 |
| 2019 | Identifying gene-specific subgroups: an alternative to biclusteringabstractBACKGROUND: Transcriptome analysis aims at gaining insight into cellular processes through discovering gene expression patterns across various experimental conditions. Biclustering is a standard approach to discover genes subsets with similar expression across subgroups of samples to be identified. The result is a set of biclusters, each forming a specific submatrix of rows (e.g. genes) and columns (e.g. samples). Relevant biclusters can, however, be missed when, due to the presence of a few outliers, they lack the assumed homogeneity of expression values among a few gene/sample combinations. The Max-Sum SubMatrix problem addresses this issue by looking at highly expressed subsets of genes and of samples, without enforcing such homogeneity. RESULTS: We present here the K-CPGC algorithm to identify K relevant submatrices. Our main contribution is to show that this approach outperforms biclustering algorithms to identify several gene subsets representative of specific subgroups of samples. Experiments are conducted on 35 gene expression datasets from human tissues and yeast samples. We report comparative results with those obtained by several biclustering algorithms, including CCA, xMOTIFs, ISA, QUBIC, Plaid and Spectral. Gene enrichment analysis demonstrates the benefits of the proposed approach to identify more statistically significant gene subsets. The most significant Gene Ontology terms identified with K-CPGC are shown consistent with the controlled conditions of each dataset. This analysis supports the biological relevance of the identified gene subsets. An additional contribution is the statistical validation protocol proposed here to assess the relative performances of biclustering algorithms and of the proposed method. It relies on a Friedman test and the Hochberg's sequential procedure to report critical differences of ranks among all algorithms. CONCLUSIONS: We propose here the K-CPGC method, a computationally efficient algorithm to identify K max-sum submatrices in a large gene expression matrix. Comparisons show that it identifies more significantly enriched subsets of genes and specific subgroups of samples which are easily interpretable by biologists. Experiments also show its ability to identify more reliable GO terms. These results illustrate the benefits of the proposed approach in terms of interpretability and of biological enrichment quality. Open implementation of this algorithm is available as an R package. Vincent Branders, Pierre Schaus, Pierre Dupont |
BMC Bioinform. | 3 |
| 2015 | Survival Analysis with Cox Regression and Random Non-linear Projections
Samuel Branders, Benoît Frénay, Pierre Dupont |
ESANN | 3 |
| 2015 | Inferring statistically significant features from random forests
Jérôme Paul, Pierre Dupont |
Neurocomputing | 2 |
| 2015 | Kernel methods for heterogeneous feature selection
Jérôme Paul, Roberto D'Ambrosio, Pierre Dupont |
Neurocomputing | 3 |
| 2014 | Kernel methods for mixed feature selection
Jérôme Paul, Pierre Dupont |
ESANN | 2 |
| 2013 | STAMINA: a competition to encourage the development and assessment of software model inference techniquesabstractModels play a crucial role in the development and maintenance of software systems, but are often neglected during the development process due to the considerable manual effort required to produce them. In response to this problem, numerous techniques have been developed that seek to automate the model generation task with the aid of increasingly accurate algorithms from the domain of Machine Learning. From an empirical perspective, these are extremely challenging to compare; there are many factors that are difficult to control (e.g. the richness of the input and the complexity of subject systems), and numerous practical issues that are just as troublesome (e.g. tool availability). This paper describes the StaMinA ( Sta te M achine In ference A pproaches) competiton, that was designed to address these problems. The competition attracted numerous submissions, many of which were improved or adapted versions of techniques that had not been subjected to extensive empirical evaluations, and had not been evaluated with respect to their ability to infer models of software systems. This paper shows how many of these techniques substantially improve on the state of the art, providing insights into some of the factors that could underpin the success of the best techniques. In a more general sense it demonstrates the potential for competitions to act as a useful basis for empirical software engineering by (a) spurring the development of new techniques and (b) facilitating their comparative evaluation to an extent that would usually be prohibitively challenging without the active participation of the developers. Neil Walkinshaw, Bernard Lambeau, Christophe Damas, Kirill Bogdanov 0002, Pierre Dupont |
Empir. Softw. Eng. | 5 |
| 2013 | Type 1 and 2 mixtures of Kullback-Leibler divergences as cost functions in dimensionality reduction based on similarity preservation
John A. Lee 0001, Emilie Renard, Guillaume Bernard 0001, Pierre Dupont, Michel Verleysen |
Neurocomputing | 4 |
| 2013 | Generalized spike-and-slab priors for Bayesian group feature selection using expectation propagation
Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Pierre Dupont |
J. Mach. Learn. Res. | 3 |
| 2012 | The stability of feature selection and class prediction from ensemble tree classifiers
Jérôme Paul, Michel Verleysen, Pierre Dupont |
ESANN | 3 |
| 2011 | Robust Multi-Class Gaussian Process ClassificationabstractMulti-class Gaussian Process Classifiers (MGPCs) are often affected by over-fitting problems when labeling errors occur far from the decision boundaries. To prevent this, we investigate a robust MGPC (RMGPC) which considers labeling errors independently of their distance to the decision boundaries. Expectation propagation is used for approximate inference. Experiments with several datasets in which noise is injected in the class labels illustrate the benefits of RMGPC. This method performs better than other Gaussian process alternatives based on considering latent Gaussian noise or heavy-tailed processes. When no noise is injected in the labels, RMGPC still performs equal or better than the other methods. Finally, we show how RMGPC can be used for successfully identifying data instances which are difficult to classify accurately in practice. Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Pierre Dupont |
NIPS | 3 |
| 2010 | Expectation Propagation for Bayesian Multi-task Feature Selection
Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Thibault Helleputte, Pierre Dupont |
ECML/PKDD (1) | 4 |
| 2010 | Robust biomarker identification for cancer diagnosis with ensemble feature selection methodsabstractMOTIVATION: Biomarker discovery is an important topic in biomedical applications of computational biology, including applications such as gene and SNP selection from high-dimensional data. Surprisingly, the stability with respect to sampling variation or robustness of such selection processes has received attention only recently. However, robustness of biomarkers is an important issue, as it may greatly influence subsequent biological validations. In addition, a more robust set of markers may strengthen the confidence of an expert in the results of a selection method. RESULTS: Our first contribution is a general framework for the analysis of the robustness of a biomarker selection algorithm. Secondly, we conducted a large-scale analysis of the recently introduced concept of ensemble feature selection, where multiple feature selections are combined in order to increase the robustness of the final set of selected features. We focus on selection methods that are embedded in the estimation of support vector machines (SVMs). SVMs are powerful classification models that have shown state-of-the-art performance on several diagnosis and prognosis tasks on biological data. Their feature selection extensions also offered good results for gene selection tasks. We show that the robustness of SVMs for biomarker discovery can be substantially increased by using ensemble feature selection techniques, while at the same time improving upon classification performances. The proposed methodology is evaluated on four microarray datasets showing increases of up to almost 30% in robustness of the selected biomarkers, along with an improvement of approximately 15% in classification performance. The stability improvement with ensemble methods is particularly noticeable for small signature sizes (a few tens of genes), which is most relevant for the design of a diagnosis or prognosis model from a gene signature. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thomas Abeel, Thibault Helleputte, Yves Van de Peer, Pierre Dupont, Yvan Saeys |
Bioinform. | 4 |
| 2010 | Pathway discovery in metabolic networks by subgraph extractionabstractMOTIVATION: Subgraph extraction is a powerful technique to predict pathways from biological networks and a set of query items (e.g. genes, proteins, compounds, etc.). It can be applied to a variety of different data types, such as gene expression, protein levels, operons or phylogenetic profiles. In this article, we investigate different approaches to extract relevant pathways from metabolic networks. Although these approaches have been adapted to metabolic networks, they are generic enough to be adjusted to other biological networks as well. RESULTS: We comparatively evaluated seven sub-network extraction approaches on 71 known metabolic pathways from Saccharomyces cerevisiae and a metabolic network obtained from MetaCyc. The best performing approach is a novel hybrid strategy, which combines a random walk-based reduction of the graph with a shortest paths-based algorithm, and which recovers the reference pathways with an accuracy of approximately 77%. AVAILABILITY: Most of the presented algorithms are available as part of the network analysis tool set (NeAT). The kWalks method is released under the GPL3 license. Karoline Faust, Pierre Dupont, Jérôme Callut, Jacques van Helden |
Bioinform. | 2 |
| 2009 | Partially supervised feature selection with regularized linear modelsabstractThis paper addresses feature selection techniques for classification of high dimensional data, such as those produced by microarray experiments. Some prior knowledge may be available in this context to bias the selection towards some dimensions (genes) a priori assumed to be more relevant. We propose a feature selection method making use of this partial supervision. It extends previous works on embedded feature selection with linear models including regularization to enforce sparsity. A practical approximation of this technique reduces to standard SVM learning with iterative rescaling of the inputs. The scaling factors depend here on the prior knowledge but the final selection may depart from it. Practical results on several microarray data sets show the benefits of the proposed approach in terms of the stability of the selected gene lists with improved classification performances. Thibault Helleputte, Pierre Dupont |
ICML | 2 |
| 2009 | Feature Selection by Transfer Learning with Linear Regularized Models
Thibault Helleputte, Pierre Dupont |
ECML/PKDD (1) | 2 |
| 2008 | Semi-supervised Classification from Discriminative Random Walks
Jérôme Callut, Kevin Françoisse, Marco Saerens, Pierre Dupont |
ECML/PKDD (1) | 4 |
| 2007 | Bound-Consistent Deviation Constraint
Pierre Schaus, Yves Deville, Pierre Dupont |
CP | 3 |
| 2007 | Filtering for Subgraph Isomorphism
Stéphane Zampelli, Yves Deville, Christine Solnon, Sébastien Sorlin, Pierre Dupont |
CP | 5 |
| 2007 | A Position-Based Propagator for the Open-Shop Problem
Jean-Noël Monette, Yves Deville, Pierre Dupont |
CPAIOR | 3 |
| 2007 | The Deviation Constraint
Pierre Schaus, Yves Deville, Pierre Dupont, Jean-Charles Régin |
CPAIOR | 3 |
| 2007 | Learning Partially Observable Markov Models from First Passage Times
Jérôme Callut, Pierre Dupont |
ECML | 2 |
| 2006 | Sequence Discrimination Using Phase-Type Distributions
Jérôme Callut, Pierre Dupont |
ECML | 2 |
| 2005 | CP(Graph): Introducing a Graph Computation Domain in Constraint Programming
Grégoire Dooms, Yves Deville, Pierre Dupont |
CP | 3 |
| 2005 | Approximate Constrained Subgraph Matching
Stéphane Zampelli, Yves Deville, Pierre Dupont |
CP | 3 |
| 2005 | Inducing Hidden Markov Models to Model Long-Term Dependencies
Jérôme Callut, Pierre Dupont |
ECML | 2 |
| 2005 | Fβ support vector machinesabstractWe introduce in this paper F/sub /spl beta// SVMs, a new parametrization of support vector machines. It allows to optimize a SVM in terms of F/sub /spl beta//, a classical information retrieval criterion, instead of the usual classification rate. Experiments illustrate the advantages of this approach with respect to the traditional 2-norm soft-margin SVM when precision and recall are of unequal importance. An automatic model selection procedure based on the generalization F/sub /spl beta// score is introduced. It relies on the results of Chapelle, Vapnjk et al. (2002) about the use of gradient-based techniques in SVM model selection. The derivatives of a F/sub /spl beta// loss function with respect to the hyperparameters C and the width /spl sigma/ of a gaussian kernel are formally defined. The model is then selected by performing a gradient descent of the F/sub /spl beta// loss function over the set of hyperparameters. Experiments on artificial and real-life data show the benefits of this method when the F/sub /spl beta// score is considered. Jérôme Callut, Pierre Dupont |
IJCNN | 2 |
| 2005 | Links between probabilistic automata and hidden Markov models: probability distributions, learning models and induction algorithms
Pierre Dupont, François Denis, Yann Esposito |
Pattern Recognit. | 1 |
| 2005 | Generating Annotated Behavior Models from End-User ScenariosabstractRequirements-related scenarios capture typical examples of system behaviors through sequences of desired interactions between the software-to-be and its environment. Their concrete, narrative style of expression makes them very effective for eliciting software requirements and for validating behavior models. However, scenarios raise coverage problems as they only capture partial histories of interaction among-system component instances. Moreover, they often leave the actual requirements implicit. Numerous efforts have therefore been made recently to synthesize requirements or behavior models inductively from scenarios. Two problems arise from those efforts. On the one hand, the, scenarios must be complemented with additional input such as state assertions along episodes or flowcharts on such episodes. This makes such techniques difficult to use by the nonexpert end-users who provide the scenarios. On the other hand, the generated state machines may be hard to understand as their nodes generally convey no domain- specific properties. Their validation by analysts, complementary to model checking and animation by may therefore be quite difficult. This paper describes tool-supported techniques that overcome those two problems. Our tool generates a labeled transition system (LTS) for each system component from simple forms of message sequence charts (MSC) taken as examples or counterexamples of desired behavior. No additional input is required. A global LTS for the entire system is synthesized first. This LTS covers all scenario examples and excludes all counterexamples. It is inductively generated through an interactive procedure that extends known learning techniques for grammar induction. The procedure is incremental on training examples. It interactively produces additional scenarios that the end-user has to classify as examples or counterexamples of desired behavior. The LTS synthesis procedure may thus also be used independently for requirements elicitation through scenario questions generated by the tool. The synthesized system LTS is then projected on local LTS for each system component. For model validation by analysts, the tool generates state invariants that decorate the nodes of the local LTS. Christophe Damas, Bernard Lambeau, Pierre Dupont, Axel van Lamsweerde |
IEEE Trans. Software Eng. | 3 |
| 2004 | The Principal Components Analysis of a Graph, and Its Relationships to Spectral Clustering
Marco Saerens, François Fouss, Luh Yen, Pierre Dupont |
ECML | 4 |
| 2002 | Improved Smoothing for Probabilistic Suffix Trees Seen as Variable Order Markov Chains
Christopher Kermorvant, Pierre Dupont |
ECML | 2 |
| 2000 | Probabilistic DFA Inference using Kullback-Leibler Divergence and Minimality
Franck Thollard, Pierre Dupont, Colin de la Higuera |
ICML | 2 |
| 1998 | Analysis of Frictional Contact Models for Dynamic SimulationabstractSimulation of dynamic systems possessing unilateral frictional contacts is important to many industrial applications. While rigid body models are often employed, it is well established that friction can cause problems with the existence and uniqueness of the forward dynamics problem. In these situations, we argue that compliant contact models, while increasing the length of the state vector, successfully resolve these ambiguities. The simplicity and efficiency of rigid body models, however, provide strong motivation for their use during those portions of a simulation when the compliant contact model indicates a unique and stable solution. We use singular perturbation theory in combination with linear complementarity theory to establish conditions for the validity of the rigid body model with rolling and sliding unilateral contacts for planar systems. The results are illustrated with a simple example. Peter R. Kraus, Pierre Dupont |
ICRA | 3 |
| 1993 | Dynamic use of syntactical knowledge in continuous speech recognitionabstractThe control of continuous speech recognition by a context-free based language model requires a parsing process which may overload the acoustic decoding algorithm. We present a new approach to integrate such a language model in the search process. This approach extends the beam search Viterbi algorithm. In our case, the pruning technique not only selects the most likely acoustic hypotheses but also governs the dynamic expansion of a network structure. This algorithm is general enough to cope with the self-embedded recursivity of context-free languages and it favourably compares with other parsing techniques applied to spoken inputs. We present results which show that the syntactical knowledge may be efficiently included at the frame level of an acoustic decoding algorithm. Keywords: Continuous Speech Recognition, Context-Free Language Models, Beam Search Viterbi Algorithm. 1. INTRODUCTION Many continuous speech understanding systems use different language models for recognition, on the ... Pierre Dupont |
EUROSPEECH | 1 |