Piotr Kawa

dblp:267/1858 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-2025-0547ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Chatbot to Validator: A Dual-Agent Strategy towards Trustworthy On-Premise Conversational LLMs
Michal Podpora, Marek Baranowski, Aleksandra Kawala-Sterniuk, Mariusz Pelc, Piotr Rogala, Piotr Kawa, Pawel Piróg, Anna Romaniewska, Wojciech Rogala
ICAART (1)6
2026 WSSSP-Net: Weakly Supervised Semantic Segmentation Plugin Network for Face Anti-Spoofing
abstract
Face anti-spoofing (FAS) is essential for protecting facial-biometric systems from presentation attacks. We propose WSSSP-Net, a Weakly Supervised Semantic Segmentation Plugin Network that integrates a lightweight, attention-based segmentation decoder at multiple depths of any CNN or transformer encoder. Serving only as an auxiliary training-time module, the decoder guides feature learning without increasing inference runtime. Pixel-wise spoof masks are automatically generated via a face-parsing pipeline, removing the need for manual annotations and enabling multiscale spoof-aware feature refinement. In leave-one-out evaluations on leading FAS benchmarks, WSSSP-Net reduces HTER by up to 24.9% and increases AUC by up to 3.2% over state-of-the-art methods. In out-of-distribution tests on a separate dataset, it lowers HTER by up to 18.4%. Across attack classes, it reduces average APCER by up to 12.9% and BPCER by up to 12.4%, achieving all improvements without added inference cost.
Krzysztof Galus, Piotr Syga, Piotr Kawa
WACV3
2025 As Good as It KAN Get: High-Fidelity Audio Representation
abstract
Implicit neural representations (INR) have gained prominence for efficiently encoding multimedia data, yet their applications in audio signals remain limited. This study introduces the Kolmogorov-Arnold Network (KAN), a novel architecture using learnable activation functions, as an effective INR model for audio representation. KAN demonstrates superior perceptual performance over previous INRs, achieving the lowest Log-Spectral Distance of 1.29 and the highest Perceptual Evaluation of Speech Quality of 3.57 for 1.5~s audio. To extend KAN's utility, we propose FewSound, a hypernetwork-based architecture that enhances INR parameter updates. FewSound outperforms the state-of-the-art HyperSound, with a 33.3% improvement in MSE and 60.87% in SI-SNR. These results show KAN as a robust and adaptable audio representation with the potential for scalability and integration into various hypernetwork frameworks.
Patryk Marszalek, Maciej Rut, Piotr Kawa, Przemyslaw Spurek, Piotr Syga
CIKM3
2025 Replay Attacks Against Audio Deepfake Detection
abstract
2245
Nicolas M. Müller, Piotr Kawa, Wei Herng Choong, Adriana Cornelia Stan, Aditya Tirumala Bukkapatnam, Karla Pizzi, Alexander Wagner, Philip Sperl
INTERSPEECH2
2024 MLAAD: The Multi-Language Audio Anti-Spoofing Dataset
abstract
Text-to-Speech (TTS) technology brings significant advantages, such as giving a voice to those with speech impairments, but also enables audio deepfakes and spoofs. The former mislead individuals and may propagate misinformation, while the latter undermine voice biometric security systems. AI-based detection can help to address these challenges by automatically differentiating between genuine and fabricated voice recordings. However, these models are only as good as their training data, which currently is severely limited due to an overwhelming concentration on English and Chinese audio in anti-spoofing databases, thus restricting its worldwide effectiveness.In response, this paper presents the Multi-Language Audio Anti-Spoof Dataset (MLAAD), created using 52 TTS models, comprising 22 different architectures, to generate 160.2 hours of synthetic voice in 23 different languages. We train and evaluate three state-of-the-art deepfake detection models with MLAAD, and observe that MLAAD demonstrates superior performance over comparable datasets like InTheWild or FakeOrReal when used as a training resource. Furthermore, in comparison with the renowned ASVspoof 2019 dataset, MLAAD proves to be a complementary resource. In tests across eight datasets, MLAAD and ASVspoof 2019 alternately outperformed each other, both excelling on four datasets.By publishing1MLAAD and making trained models accessible via an interactive webserver2, we aim to democratize antispoofing technology, making it accessible beyond the realm of specialists, thus contributing to global efforts against audio spoofing and deepfakes.
Nicolas M. Müller, Piotr Kawa, Wei Herng Choong, Edresson Casanova, Eren Gölge, Piotr Syga, Philip Sperl, Konstantin Böttinger
IJCNN2
2024 A New Approach to Voice Authenticity
abstract
2245
Nicolas M. Müller, Piotr Kawa, Shen Hu, Matthias Neu, Jennifer Williams 0001, Philip Sperl, Konstantin Böttinger
INTERSPEECH2
2023 Improved DeepFake Detection Using Whisper Features
Piotr Kawa, Marcin Plata, Michal Czuba, Piotr Szymanski, Piotr Syga
INTERSPEECH1
2023 Defense Against Adversarial Attacks on Audio DeepFake Detection
Piotr Kawa, Marcin Plata, Piotr Syga
INTERSPEECH1
2022 Attack Agnostic Dataset: Towards Generalization and Stabilization of Audio DeepFake Detection
abstract
Audio DeepFakes allow the creation of high-quality, convincing utterances and therefore pose a threat due to its potential applications such as impersonation or fake news. Methods for detecting these manipulations should be characterized by good generalization and stability leading to robustness against attacks conducted with techniques that are not explicitly included in the training. In this work, we introduce Attack Agnostic Dataset - a combination of two audio DeepFakes and one anti-spoofing datasets that, thanks to the disjoint use of attacks, can lead to better generalization of detection methods. We present a thorough analysis of current DeepFake detection methods and consider different audio features (front-ends). In addition, we propose a model based on LCNN with LFCC and mel-spectrogram front-end, which not only is characterized by a good generalization and stability results but also shows improvement over LFCC-based mode - we decrease standard deviation on all folds and EER in two folds by up to 5%.
Piotr Kawa, Marcin Plata, Piotr Syga
INTERSPEECH1
2022 SpecRNet: Towards Faster and More Accessible Audio DeepFake Detection
abstract
Audio DeepFakes are utterances generated with the use of deep neural networks. They are highly misleading and pose a threat due to use in fake news, impersonation, or extortion. In this work, we focus on increasing accessibility to the audio DeepFake detection methods by providing SpecRNet, a neural network architecture characterized by a quick inference time and low computational requirements. Our benchmark shows that SpecRNet, requiring up to about 40% less time to process an audio sample, provides performance comparable to LCNN architecture — one of the best audio DeepFake detection models. Such a method can not only be used by online multimedia services to verify a large bulk of content uploaded daily but also, thanks to its low requirements, by average citizens to evaluate materials on their devices. In addition, we provide benchmarks in three unique settings that confirm the correctness of our model. They reflect scenarios of low–resource datasets, detection on short utterances and limited attacks benchmark in which we take a closer look at the influence of particular attacks on given architectures.
Piotr Kawa, Marcin Plata, Piotr Syga
TrustCom1
2021 Verify It Yourself: A Note on Activation Functions' Influence on Fast DeepFake Detection
Piotr Kawa, Piotr Syga
SECRYPT1