Benjamin Audit

dblp:53/902 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
2since 2021 · last 2024
0000-0003-2683-9990ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
comparative genomics
0.112005
CoGenT++: an extensive and extensible data environment for computational genomics · Bioinform. 2005
Bioinformatics and computational biology
functional genomics
0.112005
CoGenT++: an extensive and extensible data environment for computational genomics · Bioinform. 2005
Bioinformatics and computational biology › genomics › genomic data management
genome database
0.012003
COmplete GENome Tracking (COGENT): A Flexible Data Environment for Computational Genomics · Bioinform. 2003
Bioinformatics and computational biology
genomics
0.012003
COmplete GENome Tracking (COGENT): A Flexible Data Environment for Computational Genomics · Bioinform. 2003
Bioinformatics and computational biology
protein function prediction
0.012002
Modeling the percolation of annotation errors in a database of protein sequences · Bioinform. 2002
Bioinformatics and computational biology › biological database
sequence database quality
0.012002
Modeling the percolation of annotation errors in a database of protein sequences · Bioinform. 2002
Information theory › probability theory › stochastic processes › time series analysis
long-range dependence
0.012002
Wavelet-based estimators of scaling behavior · IEEE Trans. Inf. Theory 2002
Information theory › estimation theory › signal estimation
wavelet-based estimation
0.012002
Wavelet-based estimators of scaling behavior · IEEE Trans. Inf. Theory 2002

Methods — techniques the papers use, named apart from their topics

phylogenetic profiling · 0.1all-against-all similarity · 0.1database design · 0.0wavelet transform modulus maxima · 0.0probabilistic model · 0.0dynamical model · 0.0detrended fluctuation analysis · 0.0FARIMA · 0.0
YearPublicationVenuePosition
2024 Space-Scale Hybrid Continuous-Discrete Sliding Frank-Wolfe Method
abstract
In this work, we focus on the challenging problem of designing an off-the-grid method for dictionaries involving both positional and scale shifts.To tackle this challenge, we introduce a novel algorithm inspired by the Sliding Frank-Wolfe approach.In our proposed algorithm, positions are treated as continuous variables, whereas scales are discretized.Such a strategy eliminates numerical instabilities inherent to the direct application of Sliding Frank-Wolfe.We successfully apply this algorithm to the study of DNA replication data.
Clara Lage, Nelly Pustelnik, Jean-Michel Arbona, Benjamin Audit
IEEE Signal Process. Lett.4
2023 Neural network and kinetic modelling of human genome replication reveal replication origin locations and strengths
abstract
In human and other metazoans, the determinants of replication origin location and strength are still elusive. Origins are licensed in G1 phase and fired in S phase of the cell cycle, respectively. It is debated which of these two temporally separate steps determines origin efficiency. Experiments can independently profile mean replication timing (MRT) and replication fork directionality (RFD) genome-wide. Such profiles contain information on multiple origins' properties and on fork speed. Due to possible origin inactivation by passive replication, however, observed and intrinsic origin efficiencies can markedly differ. Thus, there is a need for methods to infer intrinsic from observed origin efficiency, which is context-dependent. Here, we show that MRT and RFD data are highly consistent with each other but contain information at different spatial scales. Using neural networks, we infer an origin licensing landscape that, when inserted in an appropriate simulation framework, jointly predicts MRT and RFD data with unprecedented precision and underlies the importance of dispersive origin firing. We furthermore uncover an analytical formula that predicts intrinsic from observed origin efficiency combined with MRT data. Comparison of inferred intrinsic origin efficiencies with experimental profiles of licensed origins (ORC, MCM) and actual initiation events (Bubble-seq, SNS-seq, OK-seq, ORM) show that intrinsic origin efficiency is not solely determined by licensing efficiency. Thus, human replication origin efficiency is set at both the origin licensing and firing steps.
Jean-Michel Arbona, Hadi Kabalane, Jeremy Barbier, Arach Goldar, Olivier Hyrien, Benjamin Audit
PLoS Comput. Biol.6
2017 Multi-scale structural community organisation of the human genome
abstract
BACKGROUND: Structural interaction frequency matrices between all genome loci are now experimentally achievable thanks to high-throughput chromosome conformation capture technologies. This ensues a new methodological challenge for computational biology which consists in objectively extracting from these data the structural motifs characteristic of genome organisation. RESULTS: We deployed the fast multi-scale community mining algorithm based on spectral graph wavelets to characterise the networks of intra-chromosomal interactions in human cell lines. We observed that there exist structural domains of all sizes up to chromosome length and demonstrated that the set of structural communities forms a hierarchy of chromosome segments. Hence, at all scales, chromosome folding predominantly involves interactions between neighbouring sites rather than the formation of links between distant loci. CONCLUSIONS: Multi-scale structural decomposition of human chromosomes provides an original framework to question structural organisation and its relationship to functional regulation across the scales. By construction the proposed methodology is independent of the precise assembly of the reference genome and is thus directly applicable to genomes whose assembly is not fully determined.
Rasha E. Boulos, Nicolas Tremblay, Alain Arneodo, Pierre Borgnat, Benjamin Audit
BMC Bioinform.5
2015 Embryonic Stem Cell Specific "Master" Replication Origins at the Heart of the Loss of Pluripotency
abstract
Epigenetic regulation of the replication program during mammalian cell differentiation remains poorly understood. We performed an integrative analysis of eleven genome-wide epigenetic profiles at 100 kb resolution of Mean Replication Timing (MRT) data in six human cell lines. Compared to the organization in four chromatin states shared by the five somatic cell lines, embryonic stem cell (ESC) line H1 displays (i) a gene-poor but highly dynamic chromatin state (EC4) associated to histone variant H2AZ rather than a HP1-associated heterochromatin state (C4) and (ii) a mid-S accessible chromatin state with bivalent gene marks instead of a polycomb-repressed heterochromatin state. Plastic MRT regions (≲ 20% of the genome) are predominantly localized at the borders of U-shaped timing domains. Whereas somatic-specific U-domain borders are gene-dense GC-rich regions, 31.6% of H1-specific U-domain borders are early EC4 regions enriched in pluripotency transcription factors NANOG and OCT4 despite being GC poor and gene deserts. Silencing of these ESC-specific "master" replication initiation zones during differentiation corresponds to a loss of H2AZ and an enrichment in H3K9me3 mark characteristic of late replicating C4 heterochromatin. These results shed a new light on the epigenetically regulated global chromatin reorganization that underlies the loss of pluripotency and lineage commitment.
Hanna Julienne, Benjamin Audit, Alain Arneodo
PLoS Comput. Biol.2
2013 Human Genome Replication Proceeds through Four Chromatin States
abstract
Advances in genomic studies have led to significant progress in understanding the epigenetically controlled interplay between chromatin structure and nuclear functions. Epigenetic modifications were shown to play a key role in transcription regulation and genome activity during development and differentiation or in response to the environment. Paradoxically, the molecular mechanisms that regulate the initiation and the maintenance of the spatio-temporal replication program in higher eukaryotes, and in particular their links to epigenetic modifications, still remain elusive. By integrative analysis of the genome-wide distributions of thirteen epigenetic marks in the human cell line K562, at the 100 kb resolution of corresponding mean replication timing (MRT) data, we identify four major groups of chromatin marks with shared features. These states have different MRT, namely from early to late replicating, replication proceeds though a transcriptionally active euchromatin state (C1), a repressive type of chromatin (C2) associated with polycomb complexes, a silent state (C3) not enriched in any available marks, and a gene poor HP1-associated heterochromatin state (C4). When mapping these chromatin states inside the megabase-sized U-domains (U-shaped MRT profile) covering about 50% of the human genome, we reveal that the associated replication fork polarity gradient corresponds to a directional path across the four chromatin states, from C1 at U-domains borders followed by C2, C3 and C4 at centers. Analysis of the other genome half is consistent with early and late replication loci occurring in separate compartments, the former correspond to gene-rich, high-GC domains of intermingled chromatin states C1 and C2, whereas the latter correspond to gene-poor, low-GC domains of alternating chromatin states C3 and C4 or long C4 domains. This new segmentation sheds a new light on the epigenetic regulation of the spatio-temporal replication program in human and provides a framework for further studies in different cell types, in both health and disease.
Hanna Julienne, Azedine Zoufir, Benjamin Audit, Alain Arneodo
PLoS Comput. Biol.3
2012 Replication Fork Polarity Gradients Revealed by Megabase-Sized U-Shaped Replication Timing Domains in Human Cell Lines
abstract
In higher eukaryotes, replication program specification in different cell types remains to be fully understood. We show for seven human cell lines that about half of the genome is divided in domains that display a characteristic U-shaped replication timing profile with early initiation zones at borders and late replication at centers. Significant overlap is observed between U-domains of different cell lines and also with germline replication domains exhibiting a N-shaped nucleotide compositional skew. From the demonstration that the average fork polarity is directly reflected by both the compositional skew and the derivative of the replication timing profile, we argue that the fact that this derivative displays a N-shape in U-domains sustains the existence of large-scale gradients of replication fork polarity in somatic and germline cells. Analysis of chromatin interaction (Hi-C) and chromatin marker data reveals that U-domains correspond to high-order chromatin structural units. We discuss possible models for replication origin activation within U/N-domains. The compartmentalization of the genome into replication U/N-domains provides new insights on the organization of the replication program in the human genome.
Antoine Baker, Benjamin Audit, Chun-Long Chen, Benoit Moindrot, Antoine Leleu, Guillaume Guilbaud, Aurélien Rappailles, Cédric Vaillant, Arach Goldar, Fabien Mongelard, Yves D'Aubenton-Carafa, Olivier Hyrien, Claude Thermes, Alain Arneodo
PLoS Comput. Biol.2
2011 Evidence for Sequential and Increasing Activation of Replication Origins along Replication Timing Gradients in the Human Genome
abstract
Genome-wide replication timing studies have suggested that mammalian chromosomes consist of megabase-scale domains of coordinated origin firing separated by large originless transition regions. Here, we report a quantitative genome-wide analysis of DNA replication kinetics in several human cell types that contradicts this view. DNA combing in HeLa cells sorted into four temporal compartments of S phase shows that replication origins are spaced at 40 kb intervals and fire as small clusters whose synchrony increases during S phase and that replication fork velocity (mean 0.7 kb/min, maximum 2.0 kb/min) remains constant and narrowly distributed through S phase. However, multi-scale analysis of a genome-wide replication timing profile shows a broad distribution of replication timing gradients with practically no regions larger than 100 kb replicating at less than 2 kb/min. Therefore, HeLa cells lack large regions of unidirectional fork progression. Temporal transition regions are replicated by sequential activation of origins at a rate that increases during S phase and replication timing gradients are set by the delay and the spacing between successive origin firings rather than by the velocity of single forks. Activation of internal origins in a specific temporal transition region is directly demonstrated by DNA combing of the IGH locus in HeLa cells. Analysis of published origin maps in HeLa cells and published replication timing and DNA combing data in several other cell types corroborate these findings, with the interesting exception of embryonic stem cells where regions of unidirectional fork progression seem more abundant. These results can be explained if origins fire independently of each other but under the control of long-range chromatin structure, or if replication forks progressing from early origins stimulate initiation in nearby unreplicated DNA. These findings shed a new light on the replication timing program of mammalian genomes and provide a general model for their replication kinetics.
Guillaume Guilbaud, Aurélien Rappailles, Antoine Baker, Chun-Long Chen, Alain Arneodo, Arach Goldar, Yves D'Aubenton-Carafa, Claude Thermes, Benjamin Audit, Olivier Hyrien
PLoS Comput. Biol.9
2007 CORRIE: enzyme sequence annotation with confidence estimates
abstract
Using a previously developed automated method for enzyme annotation, we report the re-annotation of the ENZYME database and the analysis of local error rates per class. In control experiments, we demonstrate that the method is able to correctly re-annotate 91% of all Enzyme Classification (EC) classes with high coverage (755 out of 827). Only 44 enzyme classes are found to contain false positives, while the remaining 28 enzyme classes are not represented. We also show cases where the re-annotation procedure results in partial overlaps for those few enzyme classes where a certain inconsistency might appear between homologous proteins, mostly due to function specificity. Our results allow the interactive exploration of the EC hierarchy for known enzyme families as well as putative enzyme sequences that may need to be classified within the EC hierarchy. These aspects of our framework have been incorporated into a web-server, called CORRIE, which stands for Correspondence Indicator Estimation and allows the interactive prediction of a functional class for putative enzymes from sequence alone, supported by probabilistic measures in the context of the pre-calculated Correspondence Indicators of known enzymes with the functional classes of the EC hierarchy. The CORRIE server is available at: http://www.genomes.org/services/corrie/.
Benjamin Audit, Emmanuel D. Levy, Walter R. Gilks, Leon Goldovsky, Christos A. Ouzounis
BMC Bioinform.1
2005 CoGenT++: an extensive and extensible data environment for computational genomics
abstract
MOTIVATION: CoGenT++ is a data environment for computational research in comparative and functional genomics, designed to address issues of consistency, reproducibility, scalability and accessibility. DESCRIPTION: CoGenT++ facilitates the re-distribution of all fully sequenced and published genomes, storing information about species, gene names and protein sequences. We describe our scalable implementation of ProXSim, a continually updated all-against-all similarity database, which stores pairwise relationships between all genome sequences. Based on these similarities, derived databases are generated for gene fusions--AllFuse, putative orthologs--OFAM, protein families--TRIBES, phylogenetic profiles--ProfUse and phylogenetic trees. Extensions based on the CoGenT++ environment include disease gene prediction, pattern discovery, automated domain detection, genome annotation and ancestral reconstruction. CONCLUSION: CoGenT++ provides a comprehensive environment for computational genomics, accessible primarily for large-scale analyses as well as manual browsing.
Leon Goldovsky, Paul J. Janssen, Dag G. Ahrén, Benjamin Audit, Ildefonso Cases, Nikos Darzentas, Anton J. Enright, Núria López-Bigas, José M. Peregrín-Alvarez, Mike Smith 0001, Sophia Tsoka, Victor Kunin, Christos A. Ouzounis
Bioinform.4
2005 Probabilistic annotation of protein sequences based on functional classifications
abstract
BACKGROUND: One of the most evident achievements of bioinformatics is the development of methods that transfer biological knowledge from characterised proteins to uncharacterised sequences. This mode of protein function assignment is mostly based on the detection of sequence similarity and the premise that functional properties are conserved during evolution. Most automatic approaches developed to date rely on the identification of clusters of homologous proteins and the mapping of new proteins onto these clusters, which are expected to share functional characteristics. RESULTS: Here, we inverse the logic of this process, by considering the mapping of sequences directly to a functional classification instead of mapping functions to a sequence clustering. In this mode, the starting point is a database of labelled proteins according to a functional classification scheme, and the subsequent use of sequence similarity allows defining the membership of new proteins to these functional classes. In this framework, we define the Correspondence Indicators as measures of relationship between sequence and function and further formulate two Bayesian approaches to estimate the probability for a sequence of unknown function to belong to a functional class. This approach allows the parametrisation of different sequence search strategies and provides a direct measure of annotation error rates. We validate this approach with a database of enzymes labelled by their corresponding four-digit EC numbers and analyse specific cases. CONCLUSION: The performance of this method is significantly higher than the simple strategy consisting in transferring the annotation from the highest scoring BLAST match and is expected to find applications in automated functional annotation pipelines.
Emmanuel D. Levy, Christos A. Ouzounis, Walter R. Gilks, Benjamin Audit
BMC Bioinform.4
2003 COmplete GENome Tracking (COGENT): A Flexible Data Environment for Computational Genomics
abstract
Abstract Summary: We present a database of fully sequenced and published genomes to facilitate the re-distribution of data and ensure reproducibility of results in the field of computational genomics. For its design we have implemented an extremely simple yet powerful schema to allow linking of genome sequence data to other resources. Availability: http://maine.ebi.ac.uk:8000/services/cogent/ Contact: [email protected] * To whom correspondence should be addressed. † The authors wish it to be known that, in their opinion, the first two authors should be regarded as joint First Authors. ‡ Present Address: External Services Group, EMBL-EBI.
Paul J. Janssen, Anton J. Enright, Benjamin Audit, Ildefonso Cases, Leon Goldovsky, Nicola Harte, Victor Kunin, Christos A. Ouzounis
Bioinform.3
2002 Modeling the percolation of annotation errors in a database of protein sequences
abstract
Public sequence databases contain information on the sequence, structure and function of proteins. Genome sequencing projects have led to a rapid increase in protein sequence information, but reliable, experimentally verified, information on protein function lags a long way behind. To address this deficit, functional annotation in protein databases is often inferred by sequence similarity to homologous, annotated proteins, with the attendant possibility of error. Now, the functional annotation in these homologous proteins may itself have been acquired through sequence similarity to yet other proteins, and it is generally not possible to determine how the functional annotation of any given protein has been acquired. Thus the possibility of chains of misannotation arises, a process we term 'error percolation'. With some simple assumptions, we develop a dynamical probabilistic model for these misannotation chains. By exploring the consequences of the model for annotation quality it is evident that this iterative approach leads to a systematic deterioration of database quality.
Walter R. Gilks, Benjamin Audit, Daniela De Angelis, Sophia Tsoka, Christos A. Ouzounis
Bioinform.2
2002 Wavelet-based estimators of scaling behavior
abstract
Various wavelet-based estimators of self-similarity or long-range dependence scaling exponent are studied extensively. These estimators mainly include the (bi)orthogonal wavelet estimators and the wavelet transform modulus maxima (WTMM) estimator. This study focuses both on short and long time-series. In the framework of fractional autoregressive integrated moving average (FARIMA) processes, we advocate the use of approximately adapted wavelet estimators. For these "ideal" processes, the scaling behavior actually extends down to the smallest scale, i.e., the sampling period of the time series, if an adapted decomposition is used. But in practical situations, there generally exists a cutoff scale below which the scaling behavior no longer holds. We test the robustness of the set of wavelet-based estimators with respect to that cutoff scale as well as to the specific density of the underlying law of the process. In all situations, the WTMM estimator is shown to be the best or among the best estimators in terms of the mean-squared error (MSE). We also compare the wavelet estimators with the detrended fluctuation analysis (DFA) estimator which was previously proved to be among the best estimators which are not wavelet-based estimators. The WTMM estimator turns out to be a very competitive estimator which can be further generalized to characterize multiscaling behavior.
Benjamin Audit, Emmanuel Bacry, Jean-François Muzy, Alain Arneodo
IEEE Trans. Inf. Theory1