Emmanuel Bacry

dblp:71/5652 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorTheory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 78% Learning theory · 12% Graph learning · 10%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 77% Computational social science and digital humanities · 23%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
1.032020
Sparse and low-rank multivariate Hawkes processes · J. Mach. Learn. Res. 2020
tick: a Python Library for Statistical Learning, with an emphasis on Hawkes Processes and Time-Dependent Models · J. Mach. Learn. Res. 2017
Uncovering Causality from Multivariate Hawkes Integrated Cumulants · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process › temporal point process › hawkes process
multivariate hawkes process
0.722020
Sparse and low-rank multivariate Hawkes processes · J. Mach. Learn. Res. 2020
Uncovering Causality from Multivariate Hawkes Integrated Cumulants · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process › temporal point process
hawkes process
0.622017
tick: a Python Library for Statistical Learning, with an emphasis on Hawkes Processes and Time-Dependent Models · J. Mach. Learn. Res. 2017
Uncovering Causality from Multivariate Hawkes Integrated Cumulants · J. Mach. Learn. Res. 2017
Machine learning › Learning theory › high-dimensional statistics
matrix regularization
0.412020
Sparse and low-rank multivariate Hawkes processes · J. Mach. Learn. Res. 2020
Machine learning › Graph learning › graph inference
network structure inference
0.412020
Sparse and low-rank multivariate Hawkes processes · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.312017
Uncovering Causality from Multivariate Hawkes Integrated Cumulants · J. Mach. Learn. Res. 2017
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.312017
Uncovering Causality from Multivariate Hawkes Integrated Cumulants · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
granger causality
0.312017
Uncovering Causality from Multivariate Hawkes Integrated Cumulants · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.312017
Uncovering Causality from Multivariate Hawkes Integrated Cumulants · J. Mach. Learn. Res. 2017
Computational science and engineering › statistical computing
statistical software
0.312017
tick: a Python Library for Statistical Learning, with an emphasis on Hawkes Processes and Time-Dependent Models · J. Mach. Learn. Res. 2017
Computational social science and digital humanities
social network analysis
0.112017
Uncovering Causality from Multivariate Hawkes Integrated Cumulants · ICML 2017
Information theory › probability theory › stochastic processes › time series analysis
long-range dependence
0.012002
Wavelet-based estimators of scaling behavior · IEEE Trans. Inf. Theory 2002
Information theory › estimation theory › signal estimation
wavelet-based estimation
0.012002
Wavelet-based estimators of scaling behavior · IEEE Trans. Inf. Theory 2002

Methods — techniques the papers use, named apart from their topics

integrated cumulants · 0.9time-dependent models · 0.6statistical learning · 0.6moment matching · 0.6trace norm penalization · 0.4matrix martingale concentration inequality · 0.4least squares loss · 0.4l1 penalization · 0.4wavelet transform modulus maxima · 0.0detrended fluctuation analysis · 0.0FARIMA · 0.0
YearPublicationVenuePosition
2024 Sharing sensitive data in life sciences: an overview of centralized and federated approaches
abstract
Biomedical data are generated and collected from various sources, including medical imaging, laboratory tests and genome sequencing. Sharing these data for research can help address unmet health needs, contribute to scientific breakthroughs, accelerate the development of more effective treatments and inform public health policy. Due to the potential sensitivity of such data, however, privacy concerns have led to policies that restrict data sharing. In addition, sharing sensitive data requires a secure and robust infrastructure with appropriate storage solutions. Here, we examine and compare the centralized and federated data sharing models through the prism of five large-scale and real-world use cases of strategic significance within the European data sharing landscape: the French Health Data Hub, the BBMRI-ERIC Colorectal Cancer Cohort, the federated European Genome-phenome Archive, the Observational Medical Outcomes Partnership/OHDSI network and the EBRAINS Medical Informatics Platform. Our analysis indicates that centralized models facilitate data linkage, harmonization and interoperability, while federated models facilitate scaling up and legal compliance, as the data typically reside on the data generator's premises, allowing for better control of how data are shared. This comparative study thus offers guidance on the selection of the most appropriate sharing strategy for sensitive datasets and provides key insights for informed decision-making in data sharing efforts.
Maria Alexandra Rujano, Jan-Willem Boiten, Christian Ohmann, Steve Canham, Sergio Contrino, Romain David 0001, Jonathan Ewbank, Claudia Filippone, Claire Connellan, Ilse Custers, Rick van Nuland, Michaela Theresia Mayrhofer, Petr Holub, Eva García Álvarez, Emmanuel Bacry, Nigel Hughes, Mallory Ann Freeberg, Birgit Schaffhauser, Harald Wagener, Alex Sánchez-Pla, Guido Bertolini, Maria Panagiotopoulou
Briefings Bioinform.15
2020 ZiMM: A deep learning model for long term and blurry relapses with non-clinical claims data
abstract
International audience
Anastasiia Kabeshova, Yiyang Yu, Bertrand Lukacs, Emmanuel Bacry, Stéphane Gaïffas
J. Biomed. Informatics4
2020 Sparse and low-rank multivariate Hawkes processes
abstract
We consider the problem of unveiling the implicit network structure of node interactions (such as user interactions in a social network), based only on high-frequency timestamps. Our inference is based on the minimization of the least-squares loss associated with a multivariate Hawkes model, penalized by $\ell_1$ and trace norm of the interaction tensor. We provide a first theoretical analysis for this problem, that includes sparsity and low-rank inducing penalizations. This result involves a new data-driven concentration inequality for matrix martingales in continuous time with observable variance, which is a result of independent interest and a broad range of possible applications since it extends to matrix martingales former results restricted to the scalar case. A consequence of our analysis is the construction of sharply tuned $\ell_1$ and trace-norm penalizations, that leads to a data-driven scaling of the variability of information available for each users. Numerical experiments illustrate the significant improvements achieved by the use of such data-driven penalizations.
Emmanuel Bacry, Martin Bompaire, Stéphane Gaïffas, Jean-François Muzy
J. Mach. Learn. Res.1
2017 Uncovering Causality from Multivariate Hawkes Integrated Cumulants
abstract
We design a new nonparametric method that allows one to estimate the matrix of integrated kernels of a multivariate Hawkes process. This matrix not only encodes the mutual influences of each node of the process, but also disentangles the causality relationships between them. Our approach is the first that leads to an estimation of this matrix without any parametric modeling and estimation of the kernels themselves. A consequence is that it can give an estimation of causality relationships between nodes (or users), based on their activity timestamps (on a social network for instance), without knowing or estimating the shape of the activities lifetime. For that purpose, we introduce a moment matching method that fits the second-order and the third-order integrated cumulants of the process. A theoretical analysis allows to prove that this new estimation technique is consistent. Moreover, we show on numerical experiments that our approach is indeed very robust to the shape of the kernels, and gives appealing results on the MemeTracker database and on financial order book data.
Massil Achab, Emmanuel Bacry, Stéphane Gaïffas, Iacopo Mastromatteo, Jean-François Muzy
ICML2
2017 Uncovering Causality from Multivariate Hawkes Integrated Cumulants
Massil Achab, Emmanuel Bacry, Stéphane Gaïffas, Iacopo Mastromatteo, Jean-François Muzy
J. Mach. Learn. Res.2
2017 tick: a Python Library for Statistical Learning, with an emphasis on Hawkes Processes and Time-Dependent Models
Emmanuel Bacry, Martin Bompaire, Philip Deegan, Stéphane Gaïffas, Søren Poulsen
J. Mach. Learn. Res.1
2016 First- and Second-Order Statistics Characterization of Hawkes Processes and Non-Parametric Estimation
abstract
We show that the jumps correlation matrix of a multivariate Hawkes process is related to the Hawkes kernel matrix through a system of Wiener-Hopf integral equations. A Wiener-Hopf argument allows one to prove that this system (in which the kernel matrix is the unknown) possesses a unique causal solution and consequently that the first- and second-order properties fully characterize a Hawkes process. The numerical inversion of this system of integral equations allows us to propose a fast and efficient method, which main principles were initially sketched by Bacry and Muzy, to perform a non-parametric estimation of the Hawkes kernel matrix. In this paper, we perform a systematic study of this non-parametric estimation procedure in the general framework of marked Hawkes processes. We precisely describe this procedure step by step. We discuss the estimation error and explain how the values for the main parameters should be chosen. Various numerical examples are given in order to illustrate the broad possibilities of this estimation procedure ranging from monovariate (power-law or non-positive kernels) up to three-variate (circular dependence) processes. A comparison with other non-parametric estimation procedures is made. Applications to high-frequency trading events in financial markets and to earthquakes occurrence dynamics are finally considered.
Emmanuel Bacry, Jean-François Muzy
IEEE Trans. Inf. Theory1
2011 Modeling microstructure noise using Hawkes processes
abstract
Hawkes processes are used for modeling tick-by-tick variations of a single or of a pair of asset prices. For each asset, two counting processes (with stochastic intensities) are associated respectively with the positive and negative jumps of the price. We show that, by coupling these two intensities, one can re produce high-frequency mean reversion structure that is characteristic of the microstructure noise. Moreover, in the case of two assets, by coupling the stochastic intensities corresponding to the positive (resp. negative) jumps of each asset, we are able to reproduce the Epps effect, i.e., the decorrelation of the increments at microscopic scales. At large scale our model becomes diffusive and converge towards a standard Brownian motion. Analytical closed-form formulae for the mean signature plot, the diffusive correlation matrix and the cross-asset correlation function at any time-scale are given. Empirical results are shown on futures Euro-Bund and Euro-Bobl high frequency data.
Emmanuel Bacry, Sylvain Delattre, Marc Hoffmann, Jean-François Muzy
ICASSP1
2007 Audio Signal Denoising with Complex Wavelets and Adaptive Block Attenuation
abstract
We investigate a new audio denoising algorithm. Complex wavelets protect phase of signals and are thus preferred in audio signal processing to real wavelets. The block attenuation eliminates the residual noise artifacts in reconstructed signals and provides a good approximation of the attenuation with oracle. A connection between the block attenuation and the decision-directed a priori SNR estimator of Ephraim and Malah is studied. Finally we introduce an adaptive block technique based on the dyadic CART algorithm. The experiments show that not only the proposed method does eliminate the residual noise artifacts, but it also preserves transients of signals better than short-time Fourier based methods do.
Guoshen Yu, Emmanuel Bacry, Stéphane Mallat
ICASSP (3)2
2002 Wavelet-based estimators of scaling behavior
abstract
Various wavelet-based estimators of self-similarity or long-range dependence scaling exponent are studied extensively. These estimators mainly include the (bi)orthogonal wavelet estimators and the wavelet transform modulus maxima (WTMM) estimator. This study focuses both on short and long time-series. In the framework of fractional autoregressive integrated moving average (FARIMA) processes, we advocate the use of approximately adapted wavelet estimators. For these "ideal" processes, the scaling behavior actually extends down to the smallest scale, i.e., the sampling period of the time series, if an adapted decomposition is used. But in practical situations, there generally exists a cutoff scale below which the scaling behavior no longer holds. We test the robustness of the set of wavelet-based estimators with respect to that cutoff scale as well as to the specific density of the underlying law of the process. In all situations, the WTMM estimator is shown to be the best or among the best estimators in terms of the mean-squared error (MSE). We also compare the wavelet estimators with the detrended fluctuation analysis (DFA) estimator which was previously proved to be among the best estimators which are not wavelet-based estimators. The WTMM estimator turns out to be a very competitive estimator which can be further generalized to characterize multiscaling behavior.
Benjamin Audit, Emmanuel Bacry, Jean-François Muzy, Alain Arneodo
IEEE Trans. Inf. Theory2