Cédric Gouy-Pailler

dblp:28/6423 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-1298-7845ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3
YearPublicationVenuePosition
2026 MOSAIC-FL, a Micro-Service Based Privacy-Preserving Framework with Application to Genomics
Paul Largillier, Karl Paygambar, Cédric Gouy-Pailler, Vincent Meyer, Mallek Mziou, Oana Stan
SECRYPT (1)3
2024 Federated Dataset Dictionary Learning for Multi-Source Domain Adaptation
abstract
In this article, we propose an approach for federated domain adaptation, a setting where distributional shift exists among clients and some have unlabeled data. The proposed framework, FedDaDiL, tackles the resulting challenge through dictionary learning of empirical distributions. In our setting, clients’ distributions represent particular domains, and Fed-DaDiL collectively trains a federated dictionary of empirical distributions. In particular, we build upon the Dataset Dictionary Learning framework by designing collaborative communication protocols and aggregation operations. The chosen protocols keep clients’ data private, thus enhancing overall privacy compared to its centralized counterpart. We empirically demonstrate that our approach successfully generates labeled data on the target domain with extensive experiments on (i) Caltech-Office, (ii) TEP, and (iii) CWRU benchmarks. Furthermore, we compare our method to its centralized counterpart and other benchmarks in federated domain adaptation.
Fabiola Espinoza Castellon, Eduardo Fernandes Montesuma, Fred Maurice Ngolè Mboula, Aurélien Mayoue, Antoine Souloumiac, Cédric Gouy-Pailler
ICASSP6
2023 Combining homomorphic encryption and differential privacy in federated learning
abstract
Recent works have investigated the relevance and practicality of using techniques such as Differential Privacy (DP) or Homomorphic Encryption (HE) to strengthen training data privacy in the context of Federated Learning protocols. As these two techniques cover different sources of confidentiality threats (other clients/end-users for the former, aggregation server for the latter), there is a need to consistently combine them in order to bridge the gap towards more realistic deployment scenarios. In this paper, we achieve that goal by means of a novel stochastic quantization operator which allows us to establish DP guarantees when the noise is both quantized and bounded due to the use of HE. The paper is concluded by experiments on the FEMNIST dataset which show that the precision required to get state-of-the art privacy/utility trade-off (which directly impacts HE parameters and, hence, HE operations performances) results in a computation time overhead between 0.2% and 1.1% imputable to HE (depending on the key setup, either single key or threshold), for the whole training of a 500k parameters model and state-of-the-art privacy/utility trade-off.
Arnaud Grivet Sébert, Marina Checri, Oana Stan, Renaud Sirdey, Cédric Gouy-Pailler
PST5
2022 SecTL: Secure and Verifiable Transfer Learning-based inference
abstract
International audience
Abbass Madi, Oana Stan, Renaud Sirdey, Cédric Gouy-Pailler
ICISSP4
2022 Federated learning with incremental clustering for heterogeneous data
abstract
Federated learning enables different parties to collaboratively build a global model under the orchestration of a server while keeping the training data on clients' devices. However, performance is affected when clients have heterogeneous data. To cope with this problem, we assume that despite data heterogeneity, there are groups of clients who have similar data distributions that can be clustered. In previous approaches, in order to cluster clients the server requires clients to send their parameters simultaneously. However, this can be problematic in a context where there is a significant number of participants that may have limited availability. To prevent such a bottleneck, we propose FLIC (Federated Learning with Incremental Clustering), in which the server exploits the updates sent by clients during federated training instead of asking them to send their parameters simultaneously. Hence no additional communications between the server and the clients are necessary other than what classical federated learning requires. We empirically demonstrate for various non-IID cases that our approach successfully splits clients into groups following the same data distributions. We also identify the limitations of FLIC by studying its capability to partition clients at the early stages of the federated learning process efficiently. We further address attacks on models as a form of data heterogeneity and empirically show that FLIC is a robust defense against poisoning attacks even when the proportion of malicious clients is higher than 50%.
Fabiola Espinoza Castellon, Aurélien Mayoue, Jacques-Henri Sublemontier, Cédric Gouy-Pailler
IJCNN4
2022 On the robustness of randomized classifiers to adversarial examples
abstract
Abstract This paper investigates the theory of robustness against adversarial attacks. We focus on randomized classifiers (i.e. classifiers that output random variables) and provide a thorough analysis of their behavior through the lens of statistical learning theory and information theory. To this aim, we introduce a new notion of robustness for randomized classifiers, enforcing local Lipschitzness using probability metrics. Equipped with this definition, we make two new contributions. The first one consists in devising a new upper bound on the adversarial generalization gap of randomized classifiers. More precisely, we devise bounds on the generalization gap and the adversarial gap i.e. the gap between the risk and the worst-case risk under attack) of randomized classifiers. The second contribution presents a yet simple but efficient noise injection method to design robust randomized classifiers. We show that our results are applicable to a wide range of machine learning models under mild hypotheses. We further corroborate our findings with experimental results using deep neural networks on standard image datasets, namely CIFAR-10 and CIFAR-100. On these tasks, we manage to design robust models that simultaneously achieve state-of-the-art accuracy (over 0.82 clean accuracy on CIFAR-10) and enjoy guaranteed robust accuracy bounds (0.45 against $$\ell _{2}$$ ℓ 2 adversaries with magnitude 0.5 on CIFAR-10).
Rafael Pinot, Laurent Meunier, Florian Yger, Cédric Gouy-Pailler, Yann Chevaleyre, Jamal Atif
Mach. Learn.4
2021 SPEED: secure, PrivatE, and efficient deep learning
Arnaud Grivet Sébert, Rafael Pinot, Martin Zuber, Cédric Gouy-Pailler, Renaud Sirdey
Mach. Learn.4
2020 STREAMER: A Powerful Framework for Continuous Learning in Data Streams
abstract
With the proliferation of continuous data generation, data stream processing has become a key topic in research. As a consequence, the need for dedicated tools to apply continuous learning in streams emerges. This paper presents STREAMER, a flexible, scalable, and cross-platform machine learning experimenter with a realistic operational stream environment and visualization capabilities. Oriented to data scientists, this framework provides a set of machine learning algorithms and an API to easily integrate new ones. In order to illustrate how STREAMER works, we show a demonstration of an unsupervised anomaly detection of electrocardiograms (ECG) tested in a streaming context.
Sandra García-Rodríguez, Mohammad Alshaer, Cédric Gouy-Pailler
CIKM3
2020 Detecting Anomalies from Streaming Time Series using Matrix Profile and Shapelets Learning
abstract
Detecting anomalies in streaming time series data with no prior labels is considered a challenging issue, especially, when anomalies may vary with time. There is a need to deal with time series streams by identifying the anomalous patterns. These patterns can be described by representative features extracted from the data, which expresses abnormal behavior. This work addresses the challenge of performing online and continuous learning over time series data. In this paper, a solution based on the Matrix Profile algorithm and representation learning approach is developed. In light of that, we will show how the integration of these widely used approaches in the streaming context is quite important for learning and detecting anomalies in realtime.
Mohammad Alshaer, Sandra García-Rodríguez, Cédric Gouy-Pailler
ICTAI3
2019 Theoretical evidence for adversarial robustness through randomization
abstract
This paper investigates the theory of robustness against adversarial attacks. It focuses on the family of randomization techniques that consist in injecting noise in the network at inference time. These techniques have proven effective in many contexts, but lack theoretical arguments. We close this gap by presenting a theo- retical analysis of these approaches, hence explaining why they perform well in practice. More precisely, we make two new contributions. The first one relates the randomization rate to robustness to adversarial attacks. This result applies for the general family of exponential distributions, and thus extends and unifies the previous approaches. The second contribution consists in devising a new upper bound on the adversarial risk gap of randomized neural networks. We support our theoretical claims with a set of experiments.
Rafael Pinot, Laurent Meunier, Alexandre Araujo, Hisashi Kashima, Florian Yger, Cédric Gouy-Pailler, Jamal Atif
NeurIPS6
2018 Streaming Binary Sketching Based on Subspace Tracking and Diagonal Uniformization
abstract
In this paper, we address the problem of learning compact similarity-preserving embeddings for massive high-dimensional streams of data in order to perform efficient similarity search. We present a new online method for computing binary compressed representations -sketches- of high-dimensional real feature vectors. Given an expected code length c and high-dimensional input data points, our algorithm provides a c-bits binary code for preserving the distance between the points from the original high-dimensional space. Our algorithm does not require neither the storage of the whole dataset nor a chunk, thus it is fully adaptable to the streaming setting. It also provides low time complexity and convergence guarantees. We demonstrate the quality of our binary sketches through experiments on real data for the nearest neighbors search task in the online setting.
Anne Morvan, Antoine Souloumiac, Cédric Gouy-Pailler, Jamal Atif
ICASSP3
2018 Graph sketching-based Space-efficient Data Clustering
abstract
In this paper, we address the problem of recovering arbitrary-shaped data clusters from datasets while facing high space constraints, as this is for instance the case in many real-world applications when analysis algorithms are directly deployed on resources-limited mobile devices collecting the data. We present DBMSTClu a new space-efficient density-based non-parametric method working on a Minimum Spanning Tree (MST) recovered from a limited number of linear measurements i.e. a sketched version of the dissimilarity graph between the N objects to cluster. Unlike k-means, k-medians or k-medoids algorithms, it does not fail at distinguishing clusters with particular forms thanks to the property of the MST for expressing the underlying structure of a graph. No input parameter is needed contrarily to DBSCAN or the Spectral Clustering method. An approximate MST is retrieved by following the dynamic semi-streaming model in handling the dissimilarity graph as a stream of edge weight updates which is sketched in one pass over the data into a compact structure requiring O(N polylog(N)) space, far better than the theoretical memory cost O(N2) of . The recovered approximate MST as input, DBMSTClu then successfully detects the right number of nonconvex clusters by performing relevant cuts on in a time linear in N. We provide theoretical guarantees on the quality of the clustering partition and also demonstrate its advantage over the existing state-of-the-art on several datasets.
Anne Morvan, Krzysztof Choromanski, Cédric Gouy-Pailler, Jamal Atif
SDM3
2018 Graph-based Clustering under Differential Privacy
Rafael Pinot, Anne Morvan, Florian Yger, Cédric Gouy-Pailler, Jamal Atif
UAI4
2017 Structured adaptive and random spinners for fast machine learning computations
abstract
We consider an efficient computational framework for speeding up several machine learning algorithms with almost no loss of accuracy. The proposed framework relies on projections via structured matrices that we call Structured Spinners, which are formed as products of three structured matrix-blocks that incorporate rotations. The approach is highly generic, i.e. i) structured matrices under consideration can either be fully-randomized or learned, ii) our structured family contains as special cases all previously considered structured schemes, iii) the setting extends to the non-linear case where the projections are followed by non-linear functions, and iv) the method finds numerous applications including kernel approximations via random feature maps, dimensionality reduction algorithms,new fast cross-polytope LSH techniques, deep learning, convex optimization algorithms via Newton sketches, quantization with random projection trees, and more. The proposed framework comes with theoretical guarantees characterizing the capacity of the structured model in reference to its unstructured counterpart and is based on a general theoretical principle that we describe in the paper. As a consequence of our theoretical analysis, we provide the first theoretical guarantees for one of the most efficient existing LSH algorithms based on the HD 3 HD 2 HD 1 structured matrix [Andoni et al., 2015]. The exhaustive experimental evaluation confirms the accuracy and efficiency of structured spinners for a variety of different applications.
Mariusz Bojarski, Anna Choromanska, Krzysztof Choromanski, Francois Fagan, Cédric Gouy-Pailler, Anne Morvan, Nourhan Sakr, Tamás Sarlós, Jamal Atif
AISTATS5
2017 WHODID: Web-Based Interface for Human-Assisted Factory Operations in Fault Detection, Identification and Diagnosis
Pierre Blanchart, Cédric Gouy-Pailler
ECML/PKDD (3)2
2017 Multi-dimensional signal approximation with sparse structured priors using split Bregman iterations
Yoann Isaac, Quentin Barthélemy, Cédric Gouy-Pailler, Michèle Sebag, Jamal Atif
Signal Process.3
2013 Backward hidden Markov chain for outlier-robust filtering and fixed-interval smoothing
abstract
International audience
Boujemaa Ait-El-Fquih, Cédric Gouy-Pailler
ICASSP2
2013 Multi-dimensional sparse structured signal approximation using split bregman iterations
abstract
The paper focuses on the sparse approximation of signals using overcomplete representations, such that it preserves the (prior) structure of multi-dimensional signals. The underlying optimization problem is tackled using a multi-dimensional split Bregman optimization approach. An extensive empirical evaluation shows how the proposed approach compares to the state of the art depending on the signal features.
Yoann Isaac, Quentin Barthélemy, Jamal Atif, Cédric Gouy-Pailler, Michèle Sebag
ICASSP4
2013 Multi-scale test procedure for non-stationarity in short and long memory time series
abstract
In this paper, we develop a test procedure for non-stationarity for possibly long-memory processes. Contrary to most of the proposed methods, the test procedure has the same distribution for short-range and long-range dependence stationary processes. Such tests have been already proposed in [1], but these authors do not have taken into account the dependence of the wavelet coefficients within scales and between scales. We also propose an application to electric power consumption monitoring.
Olaf Kouamo, Cédric Gouy-Pailler
ICASSP2
2012 From neuronal cost-based metrics towards sparse coded signals classification
Anthony Mouraud, Quentin Barthélemy, Aurélien Mayoue, Cédric Gouy-Pailler, Anthony Larue, Hélène Paugam-Moisy
ESANN4
2009 Uncued brain-computer interfaces: a variational hidden markov model of mental state dynamics
Cédric Gouy-Pailler, Jérémie Mattout, Marco Congedo, Christian Jutten
ESANN1