VLDB 2026 Research / reviewers in the wild / expert
Thierry Rakotoarivelo
dblp:49/2168
· DBLP profile ↗
29ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-7698-6214ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 4 first-authorArtificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Security and privacy · 2 · 2 since 2021Theory of computation · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing DPSGD via Per-Sample Momentum and Low-Pass FilteringabstractDifferentially Private Stochastic Gradient Descent (DPSGD) is widely used to train deep neural networks with formal privacy guarantees. However, the addition of differential privacy (DP) often degrades model accuracy by introducing both noise and bias. Existing techniques typically address only one of these issues, as reducing DP noise can exacerbate clipping bias and vice-versa. In this paper, we propose a novel method, DP-PMLF, which integrates per-sample momentum with a low-pass filtering strategy to simultaneously mitigate DP noise and clipping bias. Our approach uses per-sample momentum to smooth gradient estimates prior to clipping, thereby reducing sampling variance. It further employs a post-processing low-pass filter to attenuate high-frequency DP noise without consuming additional privacy budget. We provide a theoretical analysis demonstrating an improved convergence rate under rigorous DP guarantees, and our empirical evaluations reveal that DP-PMLF significantly enhances the privacy-utility trade-off compared to several state-of-the-art DPSGD variants. Xincheng Xu, Thilina Ranbaduge, Thierry Rakotoarivelo, David B. Smith 0001 |
AAAI | 4 |
| 2026 | C2P-M: Critical Connection Protection in Multiplex GraphsabstractMultiplex graphs represent diverse real-world interactions among entities, where multiple relationship types coexist within the same set of entities. These graphs introduce privacy risks, as data collectors can exploit cross-layer dependencies to infer hidden and sensitive connections. In this work, we propose aC2P-Mframework that identifies and protects critical connections while preserving the structural information in multiplex graphs. Unlike conventional methods for single-layer graphs that perturb all edges uniformly,C2P-Mselectively protects critical connections, maintaining the analytical usability of the graph. To achieve this, we introduce the multiplex$p$-cohesion model, which incorporates new score functions that account for both intra-layer and inter-layer dependencies, enabling precise identification of critical connections for each vertex. For privacy protection, our method protects the identified critical connections, leveraging an adaptive Randomized Response (RR) mechanism to ensure$\varepsilon$-Local Differential Privacy (LDP). We formally prove thatC2P-Msatisfies$\varepsilon$-LDP. Extensive experiments on eight real-world multiplex graph datasets demonstrate thatC2P-Msignificantly outperforms baseline privacy-preserving methods, achieving a better privacy-utility trade-off. Conggai Li, Wei Ni 0001, Ming Ding 0001, Youyang Qu, Wenjie Zhang 0001, Thierry Rakotoarivelo |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2025 | SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and MitigationabstractLarge language models (LLMs) are sophisticated artificial intelligence systems that enable machines to generate human-like text with remarkable precision. While LLMs offer significant technological progress, their development using vast amounts of user data scraped from the web and collected from extensive user interactions poses risks of sensitive information leakage. Most existing surveys focus on the privacy implications of the training data but tend to overlook privacy risks from user interactions and advanced LLM capabilities. This paper aims to fill that gap by providing a comprehensive analysis of privacy in LLMs, categorizing the challenges into four main areas: (i) privacy issues in LLM training data, (ii) privacy challenges associated with user prompts, (iii) privacy vulnerabilities in LLM-generated outputs, and (iv) privacy challenges involving LLM agents. We evaluate the effectiveness and limitations of existing mitigation mechanisms targeting these proposed privacy challenges and identify areas for further research. Yashothara Shanmugarasa, Ming Ding 0001, Mahawaga Arachchige Pathum Chamikara, Thierry Rakotoarivelo |
AsiaCCS | 4 |
| 2024 | A NEW HOPE: Contextual Privacy Policies for Mobile Applications and An Approach Toward Automated Generation
Shidong Pan, Zhen Tao 0001, Thong Hoang, Dawen Zhang, Tianshi Li 0001, Zhenchang Xing, Xiwei Xu 0001, Mark Staples, Thierry Rakotoarivelo, David Lo 0001 |
USENIX Security Symposium | 9 |
| 2024 | Decentralized Privacy Preservation for Critical Connections in GraphsabstractMany real-world interconnections among entities can be characterized as graphs. Collecting local graph information with balanced privacy and data utility has garnered notable interest recently. This paper delves into the problem of identifying and protecting critical information of entity connections for individual participants in a graph based on cohesive subgraph searches. This problem has not been addressed in the literature. To address the problem, we propose to extract the critical connections of a queried vertex using a fortress-like cohesive subgraph model known as$p$-cohesion. A user's connections within a fortress are obfuscated when being released, to protect critical information about the user. Novel merit and penalty score functions are designed to measure each participant's critical connections in the minimal$p$-cohesion., facilitating effective identification of the connections. We further propose to preserve the privacy of a vertex enquired by only protecting its critical connections when responding to queries raised by data collectors. We prove that, under the decentralized differential privacy (DDP) mechanism, one's response satisfies$(\varepsilon , \delta )$-DDP when its critical connections are protected while the rest remains unperturbed. The effectiveness of our proposed method is demonstrated through extensive experiments on real-life graph datasets. Conggai Li, Wei Ni 0001, Ming Ding 0001, Youyang Qu, David B. Smith 0001, Wenjie Zhang 0001, Thierry Rakotoarivelo |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2023 | Improving the Utility of Differentially Private SGD by Employing Wavelet TransformsabstractDeep learning (DL) has become a powerful tool in many areas of research and industry, ranging from computer vision to natural language processing. Nonetheless, as DL models are trained on large amounts of sensitive data, concerns about data privacy have emerged. In light of this, differential privacy (DP) has emerged as a promising technique that provides strong privacy guarantees while allowing useful information to be extracted from the data. DP involves adding random noise to the training data or model parameters, which makes it difficult for an attacker to identify the contribution of any single data point to the final model. Despite the promising results, DP can significantly degrade the performance of DL models, especially when dealing with large datasets or complex models. To improve the balance between privacy and utility, this paper proposes a novel modification to the vanilla DP algorithm that uses a Haar wavelet transform. The proposed method achieves better utility while maintaining the same ($\varepsilon, \delta$) privacy guarantees as vanilla DP algorithms. The paper provides an analytical demonstration of the improved noise variance bounds compared to previous methods. The paper also provides a detailed analysis of the convergence performance of the proposed algorithm and shows that the Haar wavelet transform improves the accuracy and efficiency of the training process. The experimental evaluation demonstrates that the proposed method outperforms state-of-the-art algorithms on four widely used scientific benchmark datasets making this a significant contribution to DP techniques’ practical applications in DL. Kanishka Ranaweera, David B. Smith 0001, Dinh C. Nguyen, Pubudu N. Pathirana, Ming Ding 0001, Thierry Rakotoarivelo, Aruna Seneviratne |
IEEE Big Data | 6 |
| 2023 | PADDLES: Phase-Amplitude Spectrum Disentangled Early Stopping for Learning with Noisy LabelsabstractConvolutional Neural Networks (CNNs) are powerful in learning patterns of different vision tasks, but they are sensitive to label noise and may overfit to noisy labels during training. The early stopping strategy averts updating CNNs during the early training phase and is widely employed in the presence of noisy labels. Motivated by biological findings that the amplitude spectrum (AS) and phase spectrum (PS) in the frequency domain play different roles in the animal’s vision system, we observe that PS, which captures more semantic information, can increase the robustness of CNNs to label noise, more so than AS can. We thus propose early stops at different times for AS and PS by disentangling the features of some layer(s) into AS and PS using Discrete Fourier Transform (DFT) during training. Our proposed Phase-AmplituDe DisentangLed Early Stopping (PADDLES) method is shown to be effective on both synthetic and real-world label-noise datasets. PADDLES out-performs other early stopping methods and obtains state-of-the-art performance. Huaxi Huang, Olivier Salvado, Thierry Rakotoarivelo, Dadong Wang, Tongliang Liu |
ICCV | 5 |
| 2022 | Enhancing Utility In The Watchdog Privacy MechanismabstractThis paper is concerned with enhancing data utility in the privacy watchdog method for attaining information-theoretic privacy. For a specific privacy constraint, the watchdog method filters out the high-risk data symbols through applying a uniform data regulation scheme, e.g., merging all high-risk symbols together. While this method entirely trades the symbols resolution off for privacy, we show that the data utility can be greatly improved by partitioning the high-risk symbols set and individually privatizing each subset. We further propose an agglomerative merging algorithm that finds a suitable partition of high-risk symbols: it starts with a singleton high-risk symbol, which is iteratively fused with others until the resulting subsets are private. Numerical simulations demonstrate the efficacy of this algorithm in privately achieving higher utilities in the watchdog scheme. Mohammad A. Zarrabian, Ni Ding, Parastoo Sadeghi, Thierry Rakotoarivelo |
ICASSP | 4 |
| 2021 | Improving Computational Efficiency of Communication for Omniscience and Successive OmniscienceabstractCommunication for omniscience (CO) refers to the problem where the users in a finite set V observe a discrete multiple random source and want to exchange data over broadcast channels to reach omniscience, the state where everyone recovers the entire source. This paper studies how to improve the computational complexity for the problem of minimizing the sum-rate for attaining omniscience in V. While the existing algorithms rely on the submodular function minimization (SFM) techniques and complete in O(|V|2· SFM (|V|) time, we prove the strict strong map property of the nesting SFM problem. We propose a parametric (PAR) algorithm that utilizes the parametric SFM techniques and reduces the complexity to O(|V| · SFM (|V|). We propose efficient solutions to the successive omniscience (SO): attaining omniscience successively in user subsets. We first focus on how to determine a complimentary subset X*\subsetneq V in the existing two-stage SO such that if the local omniscience in X*is reached first, the global omniscience whereafter can still be attained with the minimum sum-rate. It is shown that such a subset can be extracted at one of the iterations of the PAR algorithm. We then propose a novel multi-stage SO strategy: a nesting sequence of complimentary user subsets X*(1)\subsetneq ...\subsetneq X*(K)= V, the omniscience in which is attained progressively by the monotonic rate vectorsrV(1)≤ ...≤rV(K). We propose algorithms to obtain this K-stage SO from the returned results by the PAR algorithm. The run time of these algorithms is the same as the PAR algorithm. Ni Ding, Parastoo Sadeghi, Thierry Rakotoarivelo |
IEEE Trans. Inf. Theory | 3 |
| 2020 | Privacy-Utility Tradeoff in a Guessing Framework Inspired by Index CodingabstractThis paper studies the tradeoff in privacy and utility in a single-trial multi-terminal guessing (estimation) framework using a system model that is inspired by index coding. There are n independent discrete sources at a data curator. There are m legitimate users and one adversary, each with some side information about the sources. The data curator broadcasts a distorted function of sources to legitimate users, which is also overheard by the adversary. In terms of utility, each legitimate user wishes to perfectly reconstruct some of the unknown sources and attain a certain gain in the estimation correctness for the remaining unknown sources. In terms of privacy, the data curator wishes to minimize the maximal leakage: the worst-case guessing gain of the adversary in estimating any target function of its unknown sources after receiving the broadcast data. Given the system settings, we derive fundamental performance lower bounds on the maximal leakage to the adversary, which are inspired by the notion of confusion graph and performance bounds for the index coding problem. We also detail a greedy privacy enhancing mechanism, which is inspired by the agglomerative clustering algorithms in the information bottleneck and privacy funnel problems. Yucheng Liu 0005, Ni Ding, Parastoo Sadeghi, Thierry Rakotoarivelo |
ISIT | 4 |
| 2020 | On Properties and Optimization of Information-theoretic Privacy WatchdogabstractWe study the problem of privacy preservation in data sharing, where S is a sensitive variable to be protected and X is a non-sensitive useful variable correlated with S. Variable X is randomized into variable Y, which will be shared or released according to pY |X(y|x). We measure privacy leakage by information privacy (also known as log-lift in the literature), which guarantees mutual information privacy and differential privacy (DP). Let ${\mathcal{X}}_\varepsilon ^c \subseteq {\mathcal{X}}$ contain elements in the alphabet of for which the absolute value of log-lift (abs-log-lift for short) is greater than a desired threshold ϵ. When elements $x \in {\mathcal{X}}_\varepsilon ^c$ are randomized into $y \in {\mathcal{Y}},$ we derive the best upper bound on the abs-log-lift across the resultant pairs (s, y). We then prove that this bound is achievable via an X-invariant randomization p(y|x) = R(y) for $x,y \in {\mathcal{X}}_\varepsilon ^c$. However, the utility measured by the mutual information I(X; Y) is severely damaged in imposing a strict upper bound ϵ on the abs-log-lift. To remedy this and inspired by the probabilistic (ϵ, δ)-DP, we propose a relaxed (ϵ, δ)-log-lift framework. To achieve this relaxation, we introduce a greedy algorithm which exempts some elements in ${\mathcal{X}}_\varepsilon ^c$ from randomization, as long as their abs-log-lift is bounded by ϵ with probability 1 − δ. Numerical results demonstrate efficacy of this algorithm in achieving a better privacy-utility tradeoff. Parastoo Sadeghi, Ni Ding, Thierry Rakotoarivelo |
ITW | 3 |
| 2018 | Fairness in Multiterminal Data Compression: A Splitting Method for the Egalitarian SolutionabstractThis paper proposes a novel splitting (SPLIT) algorithm to achieve fairness in the multiterminal lossless data compression problem. It finds the egalitarian solution in the Slepian-Wolf region and completes in strongly polynomial time. We show that the SPLIT algorithm adaptively updates the source coding rates to the optimal solution, while recursively splitting the terminal set, enabling parallel and distributed computation. The result of an experiment demonstrates a significant reduction in computation time by the parallel implementation when the number of terminals becomes large. The achieved egalitarian solution is also shown to be superior to the Shapley value in distributed networks, e.g., wireless sensor networks, in that it best balances the nodes' energy consumption and is far less computationally complex to obtain. Ni Ding, David B. Smith 0001, Parastoo Sadeghi, Thierry Rakotoarivelo |
ICASSP | 4 |
| 2018 | Fairness in Multiterminal Data Compression: Decomposition of Shapley ValueabstractWe consider the problem of how to attain fairness in the multiterminal data compression problem by a game-theoretic approach and present a decomposition method for obtaining the Shapley value, a fair source coding rate vector in the Slepian-Wolf achievable region. We model a discrete memoryless multiple random source (DMMS) by a coalitional game where the entropy function quantifies the cost incurred by the source coding rates in each coalition. In the typical case for which the game is decomposable, we show that the Shapley value can be obtained separately for each subgame. The complexity of this decomposition method is determined by the maximum size of subgames, which is strictly smaller than the total number of terminals in the DMMS and contributes to a considerable reduction in computational complexity. An experimental result demonstrates large complexity reduction when the number of terminals in the DMMS becomes large. Ni Ding, David B. Smith 0001, Thierry Rakotoarivelo, Parastoo Sadeghi |
ISIT | 3 |
| 2018 | Distributed Data Compression in Sensor Clusters: A Maximum Independent Flow Approach
Ni Ding, Parastoo Sadeghi, David B. Smith 0001, Thierry Rakotoarivelo |
ISIT | 4 |
| 2018 | Adaptive Online One-Class Support Vector Machines with Applications in Structural Health MonitoringabstractOne-class support vector machine (OCSVM) has been widely used in the area of structural health monitoring, where only data from one class (i.e., healthy) are available. Incremental learning of OCSVM is critical for online applications in which huge data streams continuously arrive and the healthy data distribution may vary over time. This article proposes a novel adaptive self-advised online OCSVM that incrementally tunes the kernel parameter and decides whether a model update is required or not. As opposed to existing methods, this novel online algorithm does not rely on any fixed threshold, but it uses the slack variables in the OCSVM to determine which new data points should be included in the training set and trigger a model update. The algorithm also incrementally tunes the kernel parameter of OCSVM automatically based on the spatial locations of the edge and interior samples in the training data with respect to the constructed hyperplane of OCSVM. This new online OCSVM algorithm was extensively evaluated using synthetic data and real data from case studies in structural health monitoring. The results showed that the proposed method significantly improved the classification error rates, was able to assimilate the changes in the positive data distribution over time, and maintained a high damage detection accuracy in all case studies. Ali Anaissi, Khoa L. D. Nguyen, Thierry Rakotoarivelo, Mehrisadat Makki Alamdari, Yang Wang 0002 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2017 | Self-advised Incremental One-Class Support Vector Machines: An Application in Structural Health Monitoring
Ali Anaissi, Khoa L. D. Nguyen, Thierry Rakotoarivelo, Mehrisadat Makki Alamdari, Yang Wang 0002 |
ICONIP (1) | 3 |
| 2015 | Design, architecture and implementation of a resource discovery, reservation and provisioning framework for testbedsabstractExperimental platforms (testbeds) play a significant role in the evaluation of new and existing technologies. Their popularity has been raised lately as more and more researchers prefer experimentation over simulation as a way for acquiring more accurate results. This imposes significant challenges in testbed operators since an efficient mechanism is needed to manage the testbed's resources and provision them according to the users' needs. In this paper we describe such a framework which was implemented for the management of networking testbeds. We present the design requirements and the implementation details, along with the challenges we encountered during its operation in the NITOS testbed. Significant results were extracted through the experiences of the every day operation of the testbed's management. Donatos Stavropoulos, Aris Dadoukis, Thierry Rakotoarivelo, Maximilian Ott, Thanasis Korakis, Leandros Tassiulas |
WiOpt | 3 |
| 2014 | HPC Applications Deployment on Distributed Heterogeneous Computing Platforms via OMF, OML and P2PDCabstractA new tool and web portal are presented for deployment of High Performance Computing applications on distributed heterogeneous computing platforms. This tool relies on the decentralized environment P2PDC and the OMF and OML multithreaded control, instrumentation and measurement libraries. Deployment on PlanetLab of a numerical simulation application is studied. A first series of computational results is displayed and analyzed. Didier El Baz, The Tung Nguyen, Guillaume Jourjon, Thierry Rakotoarivelo |
PDP | 4 |
| 2014 | Temporal random walk as a lightweight communication infrastructure for opportunistic networksabstractInternational audience Victor Ramiro, Emmanuel Lochin, Patrick Sénac, Thierry Rakotoarivelo |
WoWMoM | 4 |
| 2014 | Experimentation on end-to-end performance aware algorithms in the federated environment of the heterogeneous PlanetLab and NITOS testbeds
Stratos Keranidis, Dimitris Giatsios, Thanasis Korakis, Iordanis Koutsopoulos, Leandros Tassiulas, Thierry Rakotoarivelo, Maximilian Ott, Thierry Parmentelat |
Comput. Networks | 6 |
| 2014 | An instrumentation framework for the critical task of measurement collection in the future Internet
Olivier Mehani, Guillaume Jourjon, Thierry Rakotoarivelo, Maximilian Ott |
Comput. Networks | 3 |
| 2014 | Designing and orchestrating reproducible experiments on federated networking testbeds
Thierry Rakotoarivelo, Guillaume Jourjon, Maximilian Ott |
Comput. Networks | 1 |
| 2013 | Into the Moana1 - Hypergraph-based network layer indirectionabstractIn this paper, we introduce the Moana network infrastructure. It draws on well-adopted practices from the database and software engineering communities to provide a robust and expressive information-sharing service using hypergraph-based network indirection. Our proposal is twofold. First, we argue for the need for additional layers of indirection used in modern information systems to bring the network layer abstraction closer to the developer's world, allowing for expressiveness and flexibility in the creation of future services. Second, we present a modular and extensible design of the network fabric to support incremental architectural evolution and innovation, as well as its initial evaluation. Yan Shvartzshnaider, Maximilian Ott, Olivier Mehani, Guillaume Jourjon, Thierry Rakotoarivelo, David Levy 0001 |
INFOCOM | 5 |
| 2013 | On the limits of DTN monitoringabstractCompared to wired networks, Delay/Disruption Tolerant Networks (DTN) are challenging to monitor due to their lack of infrastructure and the absence of end-to-end paths. This work studies the feasibility, limits and convergence of monitoring such DTNs. More specifically, we focus on the efficient monitoring of intercontact time distribution (ICT) between DTN participants. Our contribution is two-fold. First we propose two schemes to sample data using monitors deployed within the DTN. In particular, we sample and estimate the ICT distribution. Second, we evaluate this scheme over both simulated DTN networks and real DTN traces. Our initial results show that (i) there is a high correlation between the quality of sampling and the sampled mobility type, and (ii) the number and placement of monitors impact the estimation of the ICT distribution of the whole DTN. Victor Ramiro, Emmanuel Lochin, Patrick Sénac, Thierry Rakotoarivelo |
WOWMOM | 4 |
| 2010 | Models for an Energy-Efficient P2P Delivery ServiceabstractData and service delivery have been historically based on a ''network centric'' model, with datacentres being the focal sources. The amount of energy consumed by these datacentres has become an emerging issue for the companies operating them. Thus, many contributions have proposed solutions to improve the energy efficiency of current datacentre architecture and deployments. A recently proposed approach argues for removing the datacentres from the delivery architecture. Their functionalities will instead be distributed at the edge of the network, directly within operator-managed home devices, such as Home Gateways, or Set-Top-Box (STB). This paper presents a study of the overall energy consumption required by such a community of STBs in order to provide the same services as datacentres. This paper also investigates a possible distributed algorithm to further reduce this overall energy consumption. This algorithm will be deployed over a managed peer-to-peer network of STBs. It will make optimized decisions and instruct unused STBs to switch Off to save energy without altering the general Service Level Agreement. We demonstrate the potential benefit of such an algorithm through an off-line scheduling. Finally, we propose a service-delivery model that allows us to integrate the service availability in the energy optimization problem. The combination of these two models is the first step in the development of our energy optimisation distributed algorithm. Guillaume Jourjon, Thierry Rakotoarivelo, Maximilian Ott |
PDP | 2 |
| 2007 | SPAD: A distributed middleware architecture for QoS enhanced alternate path discovery
Thierry Rakotoarivelo, Patrick Sénac, Aruna Seneviratne, Michel Diaz |
Comput. Networks | 1 |
| 2006 | A Proactive Scheme for QoS Enhanced Alternate Path Discovery in a Super-Peer ArchitectureabstractIn the next generation Internet, the network should evolve from a plain communication medium into an endless source of services available to the end-systems. We name these services "overlay applications". They would be composed of multiple cooperative distributed application elements that would build a dynamic communication mesh, namely "overlay association". In a former contribution, we proposed an unstructured super-peer architecture (SPAD) that provides enhanced quality of service (QoS) between end-points within an overlay association. This architecture aims at discovering and utilizing composite alternate end-to-end paths that experience better QoS than the path given by the default IP routing mechanisms. This paper presents a proactive information dissemination scheme that complements SPAD's mechanisms and significantly improves its performances. Thierry Rakotoarivelo, Patrick Sénac, Aruna Seneviratne, Michel Diaz |
GLOBECOM | 1 |
| 2005 | Discovering alternate internet paths to enhance end-to-end quality of serviceabstractIn the next generation Internet, the network should evolve from a plain communication medium into an endless source of services available to the end-systems. These services (i.e. Overlay Applications) would be composed of multiple cooperative distributed software elements that would build dynamic communication mesh (i.e. an Overlay Association). We propose an unstructured Super-Peer architecture (SPAD) to provide enhanced Quality of Service (QoS) between end-points within an Overlay Association. Thierry Rakotoarivelo, Patrick Sénac |
CoNEXT | 1 |
| 2003 | Integrated Personal Mobility Architecture: A Complete Personal Mobility Solution
Binh Thai, Rachel Wan, Aruna Seneviratne, Thierry Rakotoarivelo |
Mob. Networks Appl. | 4 |