Klim Kireev

dblp:250/9566 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image Models
abstract
We evaluate the effectiveness of filtering child images from training datasets of text-to-image models to prevent model misuse to create child sexual abuse material (CSAM). First, we capture the complexity of preventing CSAM generation using a game-based security definition. Second, we show that current detection methods cannot remove all children from a dataset. Third, using an ethical proxy for CSAM (a child wearing glasses), we show that even when only a small percentage of child images are left in the training dataset after filtering, there exist prompting strategies that generate a child wearing glasses using only a few more queries than when the model is trained on the unfiltered data. Fine-tuning the filtered model on child images further reduces the additional query overhead. We also show that re-introducing a concept is possible via fine-tuning even if filtering is perfect. Our results show that current child filtering methods offer limited protection to closed-weight models and no protection to open-weight models, while reducing the generality of the model by hindering the generation of child-related concepts or changing their representation. We conclude by outlining challenges in conducting evaluations that establish robust evidence on the impact of concept filtering defenses for CSAM.
Ana-Maria Cretu 0002, Klim Kireev, Amro Abdalla, Wisdom Obinna, Raphael Meier, Sarah Adel Bargal, Elissa M. Redmiles, Carmela Troncoso
SP2
2025 A Telegram Dataset of Propaganda and its Moderation
abstract
Messaging applications like Telegram have evolved into de facto social networking platforms as they add features like broadcast channels and large groups. Yet, research on these aspects of Telegram is sparse compared to more traditional social media platforms. In this paper, we present a dataset of Telegram messages collected using the export API that returns channel histories, complemented by messages collected in real-time. This dual collection methodology allows us to label deleted messages, i.e., messages that are present in the real-time dataset but not the historical dataset. Additionally, we provide labels indicating whether messages have been sent by accounts belonging to one of two distinct propaganda networks. We provide experiments that show how this rich dataset of Telegram messages can be used to study moderation in Telegram, stances and trends on different topics, and to shed light on malicious behaviours present on Telegram. Finally, we outline other use cases where our dataset could help the research community better understand Telegram as a social network.
Klim Kireev, Yevhen Mykhno, Carmela Troncoso, Rebekah Overdorf
ICWSM1
2025 Characterizing and Detecting Propaganda-Spreading Accounts on Telegram
Klim Kireev, Yevhen Mykhno, Carmela Troncoso, Rebekah Overdorf
USENIX Security Symposium1
2023 Adversarial Robustness for Tabular Data through Cost and Utility Awareness
Klim Kireev, Bogdan Kulynych, Carmela Troncoso
NDSS1
2023 Transferable Adversarial Robustness for Categorical Data via Universal Robust Embeddings
abstract
Research on adversarial robustness is primarily focused on image and text data. Yet, many scenarios in which lack of robustness can result in serious risks, such as fraud detection, medical diagnosis, or recommender systems often do not rely on images or text but instead on tabular data. Adversarial robustness in tabular data poses two serious challenges. First, tabular datasets often contain categorical features, and therefore cannot be tackled directly with existing optimization procedures. Second, in the tabular domain, algorithms that are not based on deep networks are widely used and offer great performance, but algorithms to enhance robustness are tailored to neural networks (e.g. adversarial training). In this paper, we tackle both challenges. We present a method that allows us to train adversarially robust deep networks for tabular data and to transfer this robustness to other classifiers via universal robust embeddings tailored to categorical data. These embeddings, created using a bilevel alternating minimization framework, can be transferred to boosted trees or random forests making them robust without the need for adversarial training while preserving their high accuracy on tabular data. We show that our methods outperform existing techniques within a practical threat model suitable for tabular data.
Klim Kireev, Maksym Andriushchenko, Carmela Troncoso, Nicolas Flammarion
NeurIPS1
2022 On the effectiveness of adversarial training against common corruptions
abstract
The literature on robustness towards common corruptions shows no consensus on whether adversarial training can improve the performance in this setting. First, we show that, when used with an appropriately selected perturbation radius, Lp adversarial training can serve as a strong baseline against common corruptions improving both accuracy and calibration. Then we explain why adversarial training performs better than data augmentation with simple Gaussian noise which has been observed to be a meaningful baseline on common corruptions. Related to this, we identify the sigma-overfitting phenomenon when Gaussian augmentation overfits to a particular standard deviation used for training which has a significant detrimental effect on common corruption accuracy. We discuss how to alleviate this problem and then how to further enhance Lp adversarial training by introducing an efficient relaxation of adversarial training with learned perceptual image patch similarity as the distance metric. Through experiments on CIFAR-10 and ImageNet-100, we show that our approach does not only improve the Lp adversarial training baseline but also has cumulative gains with data augmentation methods such as AugMix, DeepAugment, ANT, and SIN, leading to state-of-the-art performance on common corruptions.
Klim Kireev, Maksym Andriushchenko, Nicolas Flammarion
UAI1