VLDB 2026 Research / reviewers in the wild / expert
Sinem Sav
dblp:190/2655
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-9096-8768ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SoK: A Taxonomy of Attacks and Defenses in Split Learning
Aqsa Shabbir, Halil Ibrahim Kanpak, Alptekin Küpçü, Sinem Sav |
ACNS (3) | 4 |
| 2026 | CURE: Privacy-Preserving Split Learning Done RightabstractTraining deep neural networks often needs large datasets stored and processed in the cloud, and in sensitive fields like healthcare, these workflows must follow strict privacy rules. Split Learning (SL), a framework that divides model layers between client(s) and server(s), is widely adopted for distributed model training. While SL reduces privacy risks by limiting server access to the full parameter set, previous research has identified that intermediate outputs exchanged between server and client can compromise the client's data privacy. Homomorphic encryption (HE)-based solutions exist, but they often impose prohibitive computational burdens. To address these challenges, we propose CURE, a novel system based on HE for the single-client setting that encrypts only the server side of the model and optionally the data. CURE enables secure SL while substantially improving communication and parallelization. We propose packing schemes for efficient execution of deep learning algorithms and generalize them to MLPs and convolutional models, enabling the evaluation of large architectures using our implementations, such as ResNet blocks. We demonstrate that CURE can achieve similar accuracy to plaintext SL, while being up to 210x more efficient in terms of the runtime compared to the state-of-the-art privacy-preserving alternatives. Finally, we propose a novel estimator that enables efficient use of HE in SL settings by recommending an optimal server-client split. Halil Ibrahim Kanpak, Aqsa Shabbir, Esra Genç, Alptekin Küpçü, Sinem Sav |
Proc. Priv. Enhancing Technol. | 5 |
| 2025 | Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data
Atilla Akkus, Masoud Poorghaffar Aghdam, Mingjie Li 0007, Junjie Chu 0002, Michael Backes 0001, Yang Zhang 0016, Sinem Sav |
USENIX Security Symposium | 7 |
| 2025 | Beacon Reconstruction Attack: Reconstruction of genomes in genomic data-sharing beacons using summary statisticsabstractMOTIVATION: Genomic data-sharing beacon protocol, developed by the Global Alliance for Genomics and Health, offers a privacy-preserving mechanism for querying genomic datasets while restricting direct data access. Despite their design, beacons remain vulnerable to privacy attacks. This study introduces a novel privacy vulnerability of the protocol: one can reconstruct large portions of the genomes of all beacon participants by only using the summary statistics reported by the protocol. RESULTS: We introduce a novel optimization-based algorithm that leverages beacon responses and SNP correlations for reconstruction. By optimizing for the SNP correlations and allele frequencies, the proposed approach achieves genome reconstruction with a substantially higher F1-score (70%) compared to baseline methods (45%) on beacons generated using individuals from the HapMap and OpenSNP datasets. We show that reconstructed genomes can be used by downstream applications such as in membership inference attacks against other beacons. Our findings reveal that beacons releasing allele frequencies substantially increase the reconstruction risk, underscoring the need for enhanced privacy-preserving mechanisms to protect genomic data. AVAILABILITY AND IMPLEMENTATION: Our implementation is available at https://github.com/ASAP-Bilkent/Beacon-Reconstruction-Attack. Kousar Saleem, A. Ercüment Çiçek, Sinem Sav |
Bioinform. | 3 |
| 2023 | Privacy-Preserving Federated Recurrent Neural NetworksabstractWe present RHODE, a novel system that enables privacy-preserving training of and prediction on Recurrent Neural Networks (RNNs) in a cross-silo federated learning setting by relying on multiparty homomorphic encryption. RHODE preserves the confidentiality of the training data, the model, and the prediction data; and it mitigates federated learning attacks that target the gradients under a passive-adversary threat model. We propose a packing scheme, multi-dimensional packing, for a better utilization of Single Instruction, Multiple Data (SIMD) operations under encryption. With multi-dimensional packing, RHODE enables the efficient processing, in parallel, of a batch of samples. To avoid the exploding gradients problem, RHODE provides several clipping approximations for performing gradient clipping under encryption. We experimentally show that the model performance with RHODE remains similar to non-secure solutions both for homogeneous and heterogeneous data distributions among the data holders. Our experimental evaluation shows that RHODE scales linearly with the number of data holders and the number of timesteps, sub-linearly and sub-quadratically with the number of features and the number of hidden units of RNNs, respectively. To the best of our knowledge, RHODE is the first system that provides the building blocks for the training of RNNs and its variants, under encryption in a federated learning setting. Sinem Sav, Abdulrahman Diaa, Apostolos Pyrgelis, Jean-Philippe Bossuat, Jean-Pierre Hubaux |
Proc. Priv. Enhancing Technol. | 1 |
| 2021 | POSEIDON: Privacy-Preserving Federated Neural Network Learning
Sinem Sav, Apostolos Pyrgelis, Juan Ramón Troncoso-Pastoriza, David Froelicher, Jean-Philippe Bossuat, João Sá Sousa, Jean-Pierre Hubaux |
NDSS | 1 |
| 2021 | Scalable Privacy-Preserving Distributed LearningabstractAbstract In this paper, we address the problem of privacy-preserving distributed learning and the evaluation of machine-learning models by analyzing it in the widespread MapReduce abstraction that we extend with privacy constraints. We designspindle(Scalable Privacy-preservINg Distributed LEarning), the first distributed and privacy-preserving system that covers the complete ML workflow by enabling the execution of a cooperative gradient-descent and the evaluation of the obtained model and by preserving data and model confidentiality in a passive-adversary model with up to N −1 colluding parties.spindleuses multiparty homomorphic encryption to execute parallel high-depth computations on encrypted data without significant overhead. We instantiatespindlefor the training and evaluation of generalized linear models on distributed datasets and show that it is able to accurately (on par with non-secure centrally-trained models) and efficiently (due to a multi-level parallelization of the computations) train models that require a high number of iterations on large input data with thousands of features, distributed among hundreds of data providers. For instance, it trains a logistic-regression model on a dataset of one million samples with 32 features distributed among 160 data providers in less than three minutes. David Froelicher, Juan Ramón Troncoso-Pastoriza, Apostolos Pyrgelis, Sinem Sav, João Sá Sousa, Jean-Philippe Bossuat, Jean-Pierre Hubaux |
Proc. Priv. Enhancing Technol. | 4 |
| 2016 | Examining the annealing schedules for RNA design algorithmabstractRNA structures are important for many biological processes in the cell. One important function of RNA are as catalytic elements. Ribozymes are RNA sequences that fold to form active structures that catalyze important chemical reactions. The folded structure for these RNA are very important; only specific conformations maintain these active structures, so it is very important for RNA to fold in a specific way. The RNA design problem describes the prediction of an RNA sequence that will fold into a given RNA structure. Solving this problem allows researchers to design RNA; they can decide on what folded secondary structure is required to accomplish a task, and the algorithm will give them a primary sequence to assemble. However, there are far too many possible primary sequence combinations to test sequentially to see if they would fold into the structure. Therefore we must employ heuristics algorithms to attempt to solve this problem. This paper introduces SIMARD, an evolutionary algorithm that uses an optimization technique called simulated annealing to solve the RNA design problem. We analyzes three different cooling schedules for the annealing process: 1) An adaptive cooling schedule, 2) a geometric cooling schedule, and 3) a geometric cooling schedule with warm up. Our results show that an adaptive annealing schedule may not be more effective at minimizing the Hamming distance between the target structure and our folded sequence's structure when compared with geometric schedules. The results also show that warming up in a geometric cooling schedule may be useful for optimizing SIMARD. Halid Emre Erhan, Sinem Sav, Stas Kalashnikov, Herbert H. Tsang |
CEC | 2 |