VLDB 2026 Research / reviewers in the wild / expert
Kevin Bleakley
dblp:79/3618
· DBLP profile ↗
7ranked-venue papers
2as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorArtificial intelligence and machine learning · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
1 paper |
Learning theory · 50% Multi-agent systems · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 100% | |
| Theoretical computer science
1 paper |
Information theory · 50% Algorithms and data structures · 50% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number alteration detection |
0.3 | 2 | 2012 | Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data · Bioinform. 2012 Control-free calling of copy number alterations in deep-sequencing data using GC-content normalization · Bioinform. 2011 |
Bioinformatics and computational biology › genomics › genomic variant analysis
genomic variation detection |
0.3 | 2 | 2012 | Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data · Bioinform. 2012 Control-free calling of copy number alterations in deep-sequencing data using GC-content normalization · Bioinform. 2011 |
Knowledge, reasoning and agents › Multi-agent systems
distributed estimation |
0.2 | 1 | 2016 | The Statistical Performance of Collaborative Inference · J. Mach. Learn. Res. 2016 |
Machine learning › Learning theory
statistical estimation |
0.2 | 1 | 2016 | The Statistical Performance of Collaborative Inference · J. Mach. Learn. Res. 2016 |
Distributed systems › distributed machine learning
collaborative inference |
0.2 | 1 | 2016 | The Statistical Performance of Collaborative Inference · J. Mach. Learn. Res. 2016 |
Bioinformatics and computational biology › cancer genomics › chromosomal aberration detection
loss of heterozygosity detection |
0.1 | 1 | 2012 | Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data · Bioinform. 2012 |
Information theory › hypothesis testing
change-point detection |
0.1 | 1 | 2010 | Fast detection of multiple change-points shared by many signals using group LARS · NIPS 2010 |
Algorithms and data structures
signal processing algorithms |
0.1 | 1 | 2010 | Fast detection of multiple change-points shared by many signals using group LARS · NIPS 2010 |
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction |
0.1 | 1 | 2009 | Supervised prediction of drug-target interactions using bipartite local models · Bioinform. 2009 |
Methods — techniques the papers use, named apart from their topics
stochastic matrix · 0.5spectral analysis · 0.5segmentation · 0.3expander graphs · 0.2expander graph · 0.2b-allele frequency profiling · 0.1GC-content normalization · 0.1lasso · 0.1group LARS · 0.1kernel methods · 0.1bipartite graph learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | The Statistical Performance of Collaborative InferenceabstractThe statistical analysis of massive and complex data sets will require the development of algorithms that depend on distributed computing and collaborative inference. Inspired by this, we propose a collaborative framework that aims to estimate the unknown mean $\theta$ of a random variable $X$. In the model we present, a certain number of calculation units, distributed across a communication network represented by a graph, participate in the estimation of $\theta$ by sequentially receiving independent data from $X$ while exchanging messages via a stochastic matrix $A$ defined over the graph. We give precise conditions on the matrix $A$ under which the statistical precision of the individual units is comparable to that of a (gold standard) virtual centralized estimate, even though each unit does not have access to all of the data. We show in particular the fundamental role played by both the non-trivial eigenvalues of $A$ and the Ramanujan class of expander graphs, which provide remarkable performance for moderate algorithmic cost. Gérard Biau, Kevin Bleakley, Benoît Cadre |
J. Mach. Learn. Res. | 2 |
| 2012 | Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing dataabstractSUMMARY: More and more cancer studies use next-generation sequencing (NGS) data to detect various types of genomic variation. However, even when researchers have such data at hand, single-nucleotide polymorphism arrays have been considered necessary to assess copy number alterations and especially loss of heterozygosity (LOH). Here, we present the tool Control-FREEC that enables automatic calculation of copy number and allelic content profiles from NGS data, and consequently predicts regions of genomic alteration such as gains, losses and LOH. Taking as input aligned reads, Control-FREEC constructs copy number and B-allele frequency profiles. The profiles are then normalized, segmented and analyzed in order to assign genotype status (copy number and allelic content) to each genomic region. When a matched normal sample is provided, Control-FREEC discriminates somatic from germline events. Control-FREEC is able to analyze overdiploid tumor samples and samples contaminated by normal cells. Low mappability regions can be excluded from the analysis using provided mappability tracks. AVAILABILITY: C++ source code is available at: http://bioinfo.curie.fr/projects/freec/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Valentina Boeva, Tatiana G. Popova, Kevin Bleakley, Pierre Chiche, Julie Cappo, Gudrun Schleiermacher, Isabelle Janoueix-Lerosey, Olivier Delattre, Emmanuel Barillot |
Bioinform. | 3 |
| 2011 | Control-free calling of copy number alterations in deep-sequencing data using GC-content normalizationabstractSUMMARY: We present a tool for control-free copy number alteration (CNA) detection using deep-sequencing data, particularly useful for cancer studies. The tool deals with two frequent problems in the analysis of cancer deep-sequencing data: absence of control sample and possible polyploidy of cancer cells. FREEC (control-FREE Copy number caller) automatically normalizes and segments copy number profiles (CNPs) and calls CNAs. If ploidy is known, FREEC assigns absolute copy number to each predicted CNA. To normalize raw CNPs, the user can provide a control dataset if available; otherwise GC content is used. We demonstrate that for Illumina single-end, mate-pair or paired-end sequencing, GC-contentr normalization provides smooth profiles that can be further segmented and analyzed in order to predict CNAs. AVAILABILITY: Source code and sample data are available at http://bioinfo-out.curie.fr/projects/freec/. Valentina Boeva, Andrei Yu. Zinovyev, Kevin Bleakley, Jean-Philippe Vert, Isabelle Janoueix-Lerosey, Olivier Delattre, Emmanuel Barillot |
Bioinform. | 3 |
| 2010 | Fast detection of multiple change-points shared by many signals using group LARSabstractWe present a fast algorithm for the detection of multiple change-points when each is frequently shared by members of a set of co-occurring one-dimensional signals. We give conditions on consistency of the method when the number of signals increases, and provide empirical evidence to support the consistency results. Jean-Philippe Vert, Kevin Bleakley |
NIPS | 2 |
| 2009 | Supervised prediction of drug-target interactions using bipartite local modelsabstractMOTIVATION: In silico prediction of drug-target interactions from heterogeneous biological data is critical in the search for drugs for known diseases. This problem is currently being attacked from many different points of view, a strong indication of its current importance. Precisely, being able to predict new drug-target interactions with both high precision and accuracy is the holy grail, a fundamental requirement for in silico methods to be useful in a biological setting. This, however, remains extremely challenging due to, amongst other things, the rarity of known drug-target interactions. RESULTS: We propose a novel supervised inference method to predict unknown drug-target interactions, represented as a bipartite graph. We use this method, known as bipartite local models to first predict target proteins of a given drug, then to predict drugs targeting a given protein. This gives two independent predictions for each putative drug-target interaction, which we show can be combined to give a definitive prediction for each interaction. We demonstrate the excellent performance of the proposed method in the prediction of four classes of drug-target interaction networks involving enzymes, ion channels, G protein-coupled receptors (GPCRs) and nuclear receptors in human. This enables us to suggest a number of new potential drug-target interactions. AVAILABILITY: An implementation of the proposed algorithm is available upon request from the authors. Datasets and all prediction results are available at http://cbio.ensmp.fr/~yyamanishi/bipartitelocal/. Kevin Bleakley, Yoshihiro Yamanishi |
Bioinform. | 1 |
| 2009 | DNA barcode analysis: a comparison of phylogenetic and statistical classification methodsabstractBACKGROUND: DNA barcoding aims to assign individuals to given species according to their sequence at a small locus, generally part of the CO1 mitochondrial gene. Amongst other issues, this raises the question of how to deal with within-species genetic variability and potential transpecific polymorphism. In this context, we examine several assignation methods belonging to two main categories: (i) phylogenetic methods (neighbour-joining and PhyML) that attempt to account for the genealogical framework of DNA evolution and (ii) supervised classification methods (k-nearest neighbour, CART, random forest and kernel methods). These methods range from basic to elaborate. We investigated the ability of each method to correctly classify query sequences drawn from samples of related species using both simulated and real data. Simulated data sets were generated using coalescent simulations in which we varied the genealogical history, mutation parameter, sample size and number of species. RESULTS: No method was found to be the best in all cases. The simplest method of all, "one nearest neighbour", was found to be the most reliable with respect to changes in the parameters of the data sets. The parameter most influencing the performance of the various methods was molecular diversity of the data. Addition of genetically independent loci--nuclear genes--improved the predictive performance of most methods. CONCLUSION: The study implies that taxonomists can influence the quality of their analyses either by choosing a method best-adapted to the configuration of their sample, or, given a certain method, increasing the sample size or altering the amount of molecular diversity. This can be achieved either by sequencing more mtDNA or by sequencing additional nuclear genes. In the latter case, they may also have to modify their data analysis method. Frederic Austerlitz, Brigitte Schaeffer, Kevin Bleakley, Madalina Olteanu, Raphaël Leblois, Michel Veuille, Catherine Larédo |
BMC Bioinform. | 4 |
| 2008 | Recovering probabilities for nucleotide trimming processes for T cell receptor TRA and TRG V-J junctions analyzed with IMGT toolsabstractBACKGROUND: Nucleotides are trimmed from the ends of variable (V), diversity (D) and joining (J) genes during immunoglobulin (IG) and T cell receptor (TR) rearrangements in B cells and T cells of the immune system. This trimming is followed by addition of nucleotides at random, forming the N regions (N for nucleotides) of the V-J and V-D-J junctions. These processes are crucial for creating diversity in the immune response since the number of trimmed nucleotides and the number of added nucleotides vary in each B or T cell. IMGT sequence analysis tools, IMGT/V-QUEST and IMGT/JunctionAnalysis, are able to provide detailed and accurate analysis of the final observed junction nucleotide sequences (tool "output"). However, as trimmed nucleotides can potentially be replaced by identical N region nucleotides during the process, the observed "output" represents a biased estimate of the "true trimming process." RESULTS: A probabilistic approach based on an analysis of the standardized tool "output" is proposed to infer the probability distribution of the "true trimmming process" and to provide plausible biological hypotheses explaining this process. We collated a benchmark dataset of TR alpha (TRA) and TR gamma (TRG) V-J rearranged sequences and junctions analysed with IMGT/V-QUEST and IMGT/JunctionAnalysis, the nucleotide sequence analysis tools from IMGT, the international ImMunoGeneTics information system, http://imgt.cines.fr. The standardized description of the tool output is based on the IMGT-ONTOLOGY axioms and concepts. We propose a simple first-order model that attempts to transform the observed "output" probability distribution into an estimate closer to the "true trimming process" probability distribution. We use this estimate to test the hypothesis that Poisson processes are involved in trimming. This hypothesis was not rejected at standard confidence levels for three of the four trimming processes: TRAV, TRAJ and TRGV. CONCLUSION: By using trimming of rearranged TR genes as a benchmark, we show that a probabilistic approach, applied to IMGT standardized tool "outputs" opens the way to plausible hypotheses on the events involved in the "true trimming process" and eventually to an exact quantification of trimming itself. With increasing high-throughput of standardized immunogenetics data, similar probabilistic approaches will improve understanding of processes so far only characterized by the "output" of standardized tools. Kevin Bleakley, Marie-Paule Lefranc, Gérard Biau |
BMC Bioinform. | 1 |