VLDB 2026 Research / reviewers in the wild / expert
Olivera Kotevska
dblp:191/9030
· DBLP profile ↗
14ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-1677-2243ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XMark: Reliable Multi-Bit Watermarking for LLM-Generated TextsabstractMulti-bit watermarking has emerged as a promising solution for embedding imperceptible binary messages into Large Language Model (LLM)-generated text, enabling reliable attribution and tracing of malicious usage of LLMs.Despite recent progress, existing methods still face key limitations: some become computationally infeasible for large messages, while others suffer from a poor trade-off between text quality and decoding accuracy.Moreover, the decoding accuracy of existing methods drops significantly when the number of tokens in the generated text is limited, a condition that frequently arises in practical usage.To address these challenges, we propose XMARK, a novel method for encoding and decoding binary messages in LLM-generated texts.The unique design of XMARK's encoder produces a less distorted logit distribution for watermarked token generation, preserving text quality, and also enables its tailored decoder to reliably recover the encoded message with limited tokens.Extensive experiments across diverse downstream tasks show that XMARK significantly improves decoding accuracy while preserving the quality of watermarked text, outperforming prior methods. Rui Hu 0005, Olivera Kotevska, Zikai Zhang 0003 |
ACL (1) | 3 |
| 2026 | IntraShuffler: A Privacy Preserving Framework for Heterogeneous DP Federated Learning
Farhin Farhad Riya, Olivera Kotevska, Jinyuan Sun |
DBSec | 2 |
| 2026 | Optimal Client Sampling in Federated Learning With Client-Level Heterogeneous Differential Privacy
Rui Hu 0005, Olivera Kotevska |
IEEE Internet Things J. | 3 |
| 2025 | Engineering Privacy at the Edge: A Practical Guide to Differential Privacy in System ArchitecturesabstractThe rapid expansion of distributed and edge computing platforms—spanning autonomous vehicles, IoT sensors, and healthcare monitors—has heightened concerns about data privacy. Differential Privacy (DP) offers a rigorous mathematical framework to protect sensitive information while retaining analytical utility. This tutorial introduces the foundations of DP for both numerical and categorical datasets and extends the discussion to correlation-aware techniques tailored for structured and high-dimensional data. Hands-on demonstrations will begin with the PETINA (Privacy prEservaTIoN Algorithms) package for numerical data and continue with MIC-DP (Maximum Information Correlated Differential Privacy) for tabular data. Designed for researchers and practitioners in secure systems, embedded architectures, and AI accelerators, the tutorial emphasizes practical and scalable methods for integrating DP into real-world system designs. Olivera Kotevska, Eyhab Al-Masri |
ICCD | 1 |
| 2025 | Balancing Trade-offs: Adaptive Differential Privacy in Interpretable Machine Learning ModelsabstractIn the advancing field of machine learning, balancing accuracy, interpretability, and privacy represents a significant challenge. The problem is exacerbated by the widespread deployment of pre-trained models locally in diverse applications, which could lead to various amounts of privacy leakage. Conventional Differential Privacy strategies, in which uniform noises are applied to model gradients, guarantee data privacy at the expense of accuracy and interpretability. This paper introduces a Feature-Sensitive Adaptive Differential Privacy (FADP) framework with a unique noise-adding strategy. Noises are adaptively added based on feature importance clustering, where important features are considered for interpretability. By employing a unique masking technique, FADP selectively preserves crucial features with minimal noise interference, maintaining accuracy while enhancing interpretability. The FADP framework addresses the limitations of traditional DP methods by preserving critical channels and improving interpretability — a vital requirement in machine learning applications that demand transparency in model decisions. Through comprehensive testing, FADP is shown to balance the trade-offs among accuracy, privacy, and interpretability, marking a substantial advancement in the field of privacy-preserving machine learning. Farhin Farhad Riya, Shahinul Hoque, Yingyuan Yang, Jinyuan Sun, Olivera Kotevska |
PST | 5 |
| 2024 | Frequency Oracle for Sensitive Data Monitoring (Student Abstract)abstractAs data privacy issues grow, finding the best privacy preservation algorithm for each situation is increasingly essential. This research has focused on understanding the frequency oracles (FO) privacy preservation algorithms. FO conduct the frequency estimation of any value in the domain. The aim is to explore how each can be best used and recommend which one to use with which data type. We experimented with different data scenarios and federated learning settings. Results showed clear guidance on when to use a specific algorithm. Richard Sances, Olivera Kotevska, Paul Laiu |
AAAI | 2 |
| 2024 | Privacy-Preserving Federated Learning for Science: Challenges and Research DirectionsabstractThis paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific artificial intelligence models, in particular, foundation models (FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy—an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains. Kibaek Kim, Raghavan Krishnan, Olivera Kotevska, Matthieu Dorier, Ravi K. Madduri, Minseok Ryu, Todd S. Munson, Robert B. Ross, Thomas Flynn 0001, Ai Kagawa, Byung-Jun Yoon, Christian Engelmann, Farzad Yousefian |
IEEE Big Data | 3 |
| 2024 | Assessing Membership Inference Attacks under Distribution ShiftsabstractMembership inference attacks (MIAs) exploit machine learning models to infer whether a data point was in the training set, posing significant privacy risks even with limited black-box access. These attacks rely on the attacker approximating the target model’s training distribution, yet the impact of distribution shifts between target and shadow models on MIA success remains underexplored. We systematically evaluate five types of distribution shifts —-cutout, jitter, Gaussian noise, label shift, and attribute shift —- at varying intensities. Our results reveal that these shifts affect MIA effectiveness in nuanced ways, with some reducing attack success while others exacerbate vulnerabilities, and the same shift can have opposite effects depending on the type of MIA. This highlights the complex interplay between distributional differences and attack performance, offering critical insights for improving model defenses against MIAs. Yichuan Shi, Olivera Kotevska, Viktor Reshniak, Amir Sadovnik |
IEEE Big Data | 2 |
| 2024 | A Survey on Privacy in Graph Neural Networks: Attacks, Preservation, and ApplicationsabstractGraph Neural Networks (GNNs) have gained significant attention owing to their ability to handle graph-structured data and the improvement in practical applications. However, many of these models prioritize high utility performance, such as accuracy, with a lack of privacy consideration, which is a major concern in modern society where privacy attacks are rampant. To address this issue, researchers have started to develop privacy-preserving GNNs. Despite this progress, there is a lack of a comprehensive overview of the attacks and the techniques for preserving privacy in the graph domain. In this survey, we aim to address this gap by summarizing the attacks on graph data according to the targeted information, categorizing the privacy preservation techniques in GNNs, and reviewing the datasets and applications that could be used for analyzing/solving privacy issues in GNNs. We also outline potential directions for future research in order to build better privacy-preserving GNNs. Yuying Zhao, Zhaoqing Li, Xueqi Cheng 0002, Yu Wang 0160, Olivera Kotevska, Philip S. Yu, Tyler Derr |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | PAS: Privacy Algorithms in SystemsabstractToday we face an explosion of data generation, ranging from health monitoring to national security infrastructure systems. More and more systems are connected to the Internet that collects data at regular time intervals. These systems share data and use machine learning methods for intelligent decisions, which resulted in numerous real-world applications (e.g., autonomous vehicles, recommendation systems, and heart-rate monitoring) that have benefited from it. However, these approaches are prone to identity thief and other privacy related cyber-security attacks. So, how can data privacy be protected efficiently in these scenarios? More dedicated efforts are needed to propose the integration of privacy techniques into existing systems and develop more advanced privacy techniques to address the complex challenges of multi-system connectivity and data fusion. Therefore, we have introduced Privacy Algorithms in Systems (PAS) at CIKM which provides a venue to gather academic researchers and industry researchers/practitioners to present their research in an effort to advance the frontier of this critical direction of privacy algorithms in systems. Philip S. Yu, Olivera Kotevska, Tyler Derr |
CIKM | 2 |
| 2022 | Analyzing Data Privacy for Edge SystemsabstractInternet-of-Things (IoT)-based streaming applications are all around us. Currently, we are transitioning from IoT processing being performed on the cloud to the edge. While moving to the edge provides significant networking efficiency benefits, IoT edge computing creates significant data privacy concerns. We propose a methodology that can successfully privacy protect the continual data streams generated by sensors on the edge device. We implement local differential privacy on streaming data and incorporate Bayesian inference and Gaussian process to evaluate the privacy policy. We demonstrate our methodology on a real-world smart meter testbed and identify the optimal privacy protection settings. Olivera Kotevska, Jordan Johnson, Aaron Gilad Kusne |
SMARTCOMP | 1 |
| 2021 | Measurement of Local Differential Privacy Techniques for IoT-based Streaming DataabstractVarious Internet of Things (IoT) devices generate complex, dynamically changed, and infinite data streams. Adversaries can cause harm if they can access the user’s sensitive raw streaming data. For this reason, protecting the privacy of the data streams is crucial. In this paper, we explore local differential privacy techniques for streaming data. We compare the techniques and report the advantages and limitations. We also present the effect on component (e.g., smoother, perturber) variations of distribution-based local differential privacy. We find that combining distribution-based noise during perturbation provides more flexibility to the interested entity. Sharmin Afrose, Danfeng Yao, Olivera Kotevska |
PST | 3 |
| 2020 | Methodology for Interpretable Reinforcement Learning Model for HVAC Energy ControlabstractDeep reinforcement learning (DRL) approaches have been used in various application areas to improve efficiency, optimization, or automation. However, very little is known about how the DRL algorithms make decisions and what features affect their performance. Using a case study of a DRL based Heating, Ventilation and Air Conditioning (HVAC) optimization methodology, we demonstrate how we can address these challenges by applying interpretability tools and systematically exploring the model inputs for better understanding the DRL behaviour and decision making process. We developed a methodology for interpretable reinforcement learning and evaluated our approach in real-world house located in Knoxville, TN. Our findings explain the reasoning behind DRL-based optimization decisions under different circumstances which has been discussed and confirmed by the experts in the field. Olivera Kotevska, Jeffrey Munk, Kuldeep R. Kurte, Kadir Amasyali, Robert W. Smith 0004, Helia Zandi |
IEEE BigData | 1 |
| 2019 | Kensor: Coordinated Intelligence from Co-Located SensorsabstractInternet of Things (IoT) is becoming more pervasive in many installations, including homes, manufacturing plants, and industrial facilities of all kinds. The data that IoT produces is a reflection of usual behavior such as daily routines and scheduled tasks, but also from unexpected behavior due to unintentional or undesirable abnormalities. Here, we focus on achieving coordinated intelligence about normal and abnormal phenomena from multiple sensors that are geographically co-located in close proximity, monitoring and controlling a set of co-located devices. Given a set of co-located sensors, we seek an intelligent approach that would automatically determine the “normal” patterns of behaviors among the correlated sensors. After normal behavior is extracted, later monitoring should detect any deviant variations over time. An example application is an entry monitoring and alert system for facilities such as nuclear reactors, where badge readers, door locks, lights, weight trackers and other co-located sensors at the entry point are collectively tracked. To address this problem, we identify the possible solution approach that can be used to solve its different variants. The implemented model is developed as a combination of rules and Markov Chain methods. Olivera Kotevska, Kalyan S. Perumalla, Juan Lopez Jr. |
IEEE BigData | 1 |