Artur Janicki

dblp:61/3320 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0002-9937-4402ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 2 since 2021Security and privacy · 7 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Adversarial Confusion Attack: Disrupting Multimodal Large Language Models
abstract
We introduce the adversarial confusion attack as a new class of threats against multimodal large language models (MLLMs).Unlike jailbreaks or targeted misclassification, the goal is to induce systematic disruption that makes the model generate incoherent or confidently incorrect outputs.Practical applications include embedding such adversarial images into websites to prevent MLLM-powered AI Agents from operating reliably.The proposed attack maximizes next-token entropy using a small ensemble of open-source MLLMs.In the white-box setting, we show that a single adversarial image can disrupt all models in the ensemble, both in the full-image and CAPTCHA-style adversarial patch settings.Despite relying on a basic adversarial technique, such as projected gradient descent (PGD), the attack generates perturbations that transfer to both unseen open-source (e.g., Qwen3-VL) and proprietary (e.g., GPT-5.1)models.• We introduce the adversarial confusion attack, which maximizes output entropy to destabilize decoding, and we characterize five distinct modes of resulting model failure.• In the white-box setting, we show that a single perturbation disrupts all models in the ensemble, in both the full-image and Adversarial CAPTCHA settings.
Jakub Hoscilowicz, Artur Janicki
ESANN2
2025 Large Language Models as Carriers of Hidden Messages
Jakub Hoscilowicz, Pawel Popiolek, Jan Rudkowski, Jedrzej Bieniasz, Artur Janicki
SECRYPT5
2025 ClickAgent: Enhancing UI Location Capabilities of Autonomous Agents
abstract
With the growing reliance on digital devices with graphical user interfaces (GUIs) like computers and smartphones, the demand for smart voice assistants has grown significantly. While multimodal large language models (MLLM) like GPT-4V excel in many areas, they struggle with GUI interactions, limiting their effectiveness in automating everyday tasks. In this work, we introduce ClickAgent, a novel framework for building autonomous agents. ClickAgent combines MLLM-driven reasoning and action planning with a separate UI location model that identifies relevant UI elements on the screen. This approach addresses a key limitation of current MLLMs: their inability to accurately locate UI elements. Evaluations conducted using both an Android emulator and a real smartphone show that ClickAgent outperforms other autonomous agents (DigiRL, CogAgent, AppAgent) on the AITW benchmark.
Jakub Hoscilowicz, Artur Janicki
SIGDIAL2
2024 Detection of AI-Generated Emails - A Case Study
abstract
This work-in-progress paper investigates the problem of assessing and detecting if a text was written by a human or if it was generated by a language model. In our case study, we focused on email messages. For the purpose of experiments, we used a combination of publicly available email datasets with our in-house data, containing in total over 10k emails. Then, we generated their “copies” using large language models (LLMs) with specific prompts. We experimented with various classifiers and feature spaces. We achieved encouraging results, with the F1-scores of almost 0.99 for email messages in English and over 0.92 for the ones in Polish, using Random Forest as a classifier. We found that the detection model relied strongly on typographic and orthographic (spelling) imperfections of the analyzed emails and on statistics of sentence lengths. We also observed the inferior results obtained for Polish, highlighting a need for research in the direction of languages underrepresented in training models.
Pawel Gryka, Kacper T. Gradon, Marek Kozlowski, Milosz Kutyla, Artur Janicki
ARES5
2024 Impact of Spelling and Editing Correctness on Detection of LLM-Generated Emails
abstract
In this paper, we investigated the impact of spelling and editing correctness on the accuracy of detection if an email was written by a human or if it was generated by a language model.As a dataset, we used a combination of publicly available email datasets with our in-house data, with over 10k emails in total.Then, we generated their "copies" using large language models (LLMs) with specific prompts.As a classifier, we used random forest, which yielded the best results in previous experiments.For English emails, we found a slight decrease in evaluation metrics if error-related features were excluded.However, for the Polish emails, the differences were more significant, indicating a decline in prediction quality by around 2% relative.The results suggest that the proposed detection method can be equally effective for English even if spelling-and grammar-checking tools are used.As for Polish, to compensate for error-related features, additional measures have to be undertaken.
Pawel Gryka, Kacper T. Gradon, Marek Kozlowski, Milosz Kutyla, Artur Janicki
FedCSIS5
2024 Unconditional Token Forcing: Extracting Text Hidden Within LLM
abstract
With the help of simple fine-tuning, one can artificially embed hidden text into large language models (LLMs).This text is revealed only when triggered by a specific query to the LLM.Two primary applications are LLM fingerprinting and steganography.In the context of LLM fingerprinting, a unique text identifier (fingerprint) is embedded within the model to verify licensing compliance.In the context of steganography, the LLM serves as a carrier for hidden messages that can be disclosed through a designated trigger.Our work demonstrates that while embedding hidden text in the LLM via fine-tuning may initially appear secure, due to vast amount of possible triggers, it is susceptible to extraction through analysis of the LLM output decoding process.We propose a novel approach to extraction called Unconditional Token Forcing.It is premised on the hypothesis that iteratively feeding each token from the LLM's vocabulary into the model should reveal sequences with abnormally high token probabilities, indicating potential embedded text candidates.Additionally, our experiments show that when the first token of a hidden fingerprint is used as an input, the LLM not only produces an output sequence with high token probabilities, but also repetitively generates the fingerprint itself.Code is available at github.com/jhoscilowic/zurek-stegano.
Jakub Hoscilowicz, Pawel Popiolek, Jan Rudkowski, Jedrzej Bieniasz, Artur Janicki
FedCSIS5
2024 Non-Linear Inference Time Intervention: Improving LLM Truthfulness
Jakub Hoscilowicz, Adam Wiacek, Jan Chojnacki, Adam Cieslak, Leszek Michon, Artur Janicki
INTERSPEECH6
2024 Gaze-dependent response activation in dialogue agent for cognitive-behavioral therapy
abstract
To provide additional support to psychiatric patients who have limited access to medical staff, our team has developed a therapeutic dialogue agent called Terabot. It operates in Polish and is enhanced with text-based emotion and intention recognition. It considers the patient’s emotions so that it can respond in an ‘empathetic’ way. Experiments have already been conducted at the Institute of Psychiatry and Neurology in Warsaw, Poland, during which we observed several problems. In particular, when the patient’s utterances were too short or too long, the dialogue system reacted inappropriately, interrupting the patient’s answer or causing very long pauses. The problem was that without input other than speech, the dialogue system could not provide answers in good timing. This decreased comfort and natural dialogue flow during patient therapeutic sessions. This decreased comfort and natural dialogue flow during patient therapeutic sessions. As it is known from human communication research, when a speaker is finishing the utterance, at the end of the speech, and directly after, this person is gazing at the at the interlocutor. To solve these unwanted problems in our dialogue system, we analyzed the fixation parameters of the patients in selected areas to find out where they gazed at most while they ended their utterances. After analyses, we found that the most fixations and the longest duration occurred on Terabot’s face. We have found that the input to our dialogue system will be the stream of real-time fixation data in these areas. This information can help to index the end of the utterance explicitly for the dialogue system. As a result, a better timing for Terabot’s speech activation can be determined. This approach allows Terabot’s response time to be individually personalized to suit the patients’ different response styles. A more natural and human-like flow of conversation with our dialogue agent can be achieved.
Karolina Gabor-Siatkowska, Izabela Stefaniak, Artur Janicki
KES3
2023 Open Vocabulary Keyword Spotting with Small-Footprint ASR-based Architecture and Language Models
abstract
We present the results of experiments on minimizing the model size for the text-based Open Vocabulary Keyword Spotting task.The main goal is to perform inference on devices with limited computing power, such as mobile phones.Our solution is based on the acoustic model architecture adopted from the automatic speech recognition task.We extend the acoustic model with a simple yet powerful language model, which improves recognition results without impacting latency and memory footprint.We also present a method to improve the recognition rate of rare keywords based on the recordings generated by a text-to-speech system.Evaluations using a public testset prove that our solution can achieve a true positive rate in the range of 73%-86%, with a false positive rate below 24%.The model size is only 3.2 MB, and the real-time factor measured on contemporary mobile phones is 0.05.
Mikolaj Pudo, Mateusz Wosik, Artur Janicki
FedCSIS3
2023 Optimizing Machine Translation for Virtual Assistants: Multi-Variant Generation with VerbNet and Conditional Beam Search
abstract
In this paper, we introduce a domain-adapted machine translation (MT) model for intelligent virtual assistants (IVA) designed to translate natural language understanding (NLU) training data sets.This work uses a constrained beam search to generate multiple valid translations for each input sentence.The search for the best translations in the presented translation algorithm is guided by a verb-frame ontology we derived from VerbNet.To assess the quality of the presented MT models, we train NLU models on these multiverb-translated resources and compare their performance to models trained on resources translated with a traditional single-best approach.Our experiments show that multi-verb translation improves intent classification accuracy by 3.8% relative compared to singlebest translation.We release five MT models that translate from English to Spanish, Polish, Swedish, Portuguese, and French, as well as an IVA verb ontology that can be used to evaluate the quality of IVA-adapted MT.
Marcin Sowanski, Artur Janicki
FedCSIS2
2023 Can We Use Probing to Better Understand Fine-Tuning and Knowledge Distillation of the BERT NLU?
Jakub Hoscilowicz, Marcin Sowanski, Piotr Czubowski, Artur Janicki
ICAART (3)4
2023 MOCKS 1.0: Multilingual Open Custom Keyword Spotting Testset
Mikolaj Pudo, Mateusz Wosik, Adam Cieslak, Justyna Krzywdziak, Bozena Lukasiak, Artur Janicki
INTERSPEECH6
2020 Improved weighted loss function for training end-of-speech detection models
abstract
In this paper we propose an improved system for the detection of end of speech (EOS) events in noisy environments, needed, for example, in voice interfaces of mobile devices. Our solution is based on a deep neural network composed of convolutional, feed-forward and LSTM layers. For the input data we use mel-frequency cepstral coefficients (MFCC). The main novelty of our solution is the metric used during the training process of the model: our loss function returns higher values the later the model recognizes the EOS event. We confront this approach with the loss functions previously used, where such a delay was not considered. The experiments run on the TIMIT corpus, as well as additional evaluations on the other types of audio data, showed that our solution is significantly more robust to noisy and far-field environments compared to the baseline solution.
Mikolaj Pudo, Adrian Wisniewski, Artur Janicki
MoMM3
2017 Increasing anti-spoofing protection in speaker verification using linear prediction
abstract
This article addresses the problem of anti-spoofing protection in an automatic speaker verification (ASV) system. An improved version of a previously proposed spoofing countermeasure is presented. The presented method is based on the analysis of linear prediction error that results from both short- and long-term prediction of the input speech signal. It was observed that non-natural speech signals, i.e., synthetic or converted speech, were predicted in a different way than genuine speech. Therefore, in contrast to the classical linear prediction analysis, where usually only the prediction coefficients are analyzed, in the proposed approach the residual (error) signals were examined. During this analysis, 23 various prediction parameters were extracted, such as the energy of the prediction error, prediction gains and temporal parameters related to the prediction error signals. Various binary classifiers were researched to separate human and spoof classes, however the support vector machines with radial basis function (SVM-RBF) yielded the best results. When tested on the corpora provided for the ASVspoof 2015 Challenge, the proposed countermeasure returned better results than the previous version of the algorithm and, in most of the cases, the baseline spoofing detector based on the local binary patterns (LBP). It is hoped that the proposed method can be part of a generalized spoofing countermeasure helping to increase security of ASV systems.
Artur Janicki
Multim. Tools Appl.1
2016 YouSkyde: information hiding for Skype video traffic
abstract
In this paper a new information hiding method for Skype videoconference calls – YouSkyde – is introduced. A Skype traffic analysis revealed that introducing intentional losses into the Skype video traffic stream to provide the means for clandestine communication is the most favourable solution. A YouSkyde proof-of-concept implementation was carried out and its experimental evaluation is presented. The results obtained prove that the proposed method is feasible and offer a steganographic bandwidth as high as 0.93 kbps, while introducing negligible distortions into transmission quality and providing high undetectability.
Wojciech Mazurczyk, Maciej Karas, Krzysztof Szczypiorski, Artur Janicki
Multim. Tools Appl.4
2016 Pitch-based steganography for Speex voice codec
abstract
Abstract This paper presents an improved version of a steganographic algorithm for IP telephony called HideF0. It is based on approximating the F0 parameter, which is responsible for conveying information about the pitch of the speech signal. The bits saved due to simplification of the pitch contour are used for the hidden transmission. In our experiments, the proposed method was applied to the narrowband Speex codec working in five different modes, with bitrates between 5,950 bps and 24,600 bps. We showed that HideF0 was able to create hidden channels with steganographic bandwidths of around 200 bps at the expense of a steganographic cost of between 0.5 and 0.7 MOS, depending on the Speex mode. Because of placing the approximation flag in the voice packet header, the improved version of the proposed algorithm yielded a significantly lower decrease in speech quality, when compared with the original version of HideF0. In addition, for low bitrates of the hidden channel (i.e., below ca. 50 bps) it was able to operate without introducing any steganographic cost. Copyright © 2016 John Wiley & Sons, Ltd.
Artur Janicki
Secur. Commun. Networks1
2016 An assessment of automatic speaker verification vulnerabilities to replay spoofing attacks
abstract
Abstract This paper analyses the threat of replay spoofing or presentation attacks in the context of automatic speaker verification. As relatively high‐technology attacks, speech synthesis and voice conversion, which have thus far received far greater attention in the literature, are probably beyond the means of the average fraudster. The implementation of replay attacks, in contrast, requires no specific expertise nor sophisticated equipment. Replay attacks are thus likely to be the most prolific in practice, while their impact is relatively under‐researched. The work presented here aims to compare at a high level the threat of replay attacks with those of speech synthesis and voice conversion. The comparison is performed using strictly controlled protocols and with six different automatic speaker verification systems including a state‐of‐the‐art iVector/probabilistic linear discriminant analysis system. Experiments show that low‐effort replay attacks present at least a comparable threat to speech synthesis and voice conversion. The paper also describes and assesses two replay attack countermeasures. A relatively new approach based on the local binary pattern analysis of speech spectrograms is shown to outperform a competing approach based on the detection of far‐field recordings. Copyright © 2016 John Wiley & Sons, Ltd.
Artur Janicki, Federico Alegre, Nicholas W. D. Evans
Secur. Commun. Networks1
2016 Trends in modern information hiding: techniques, applications, and detection
abstract
As the production, storage, and exchange of information become more extensive and important in the functioning of societies, the problem of protecting the information from unintended and undesired usage becomes more complex. In modern societies, protection of information involves many interdependent technological and policy issues related to information confidentiality, integrity, anonymity, authenticity, utility, etc. Information hiding techniques are receiving much attention today. Digital audio, video, and images are increasingly furnished with distinguishing but imperceptible marks, which may contain a hidden copyright notice or serial number or even help to prevent unauthorized copying directly. Digital watermarking and steganography may protect information, conceal secrets, or are used as core primitives in digital rights' management schemes. Alongside the previously mentioned types of digital media steganography, currently, the target of increased interest is network steganography—a part of information hiding focused on modern networks. It is a method of hiding secret data in users' normal data transmissions. Steganographic techniques arise and evolve with the development of network protocols and mechanisms and are expected to be used in secret communication or information sharing. Presently, it becomes a hot topic because of the proliferation of information networks and multimedia services in networks and social networks. The purpose of establishing applications of Information Hiding may be varied—possible uses can fall into the category of legal actions or illicit activity. Frequently, the illegal aspect is accentuated—starting from the criminal communication, through information leakage from protected systems, cyber weapon exchange, up to industrial espionage. Recently discovered malware like Hammertoss or Stegoloader utilize various information hiding techniques for botnets purposes to enable covert communication for the C and C (Command and Control) channel. This makes detection of such malware even more difficult, and it poses a serious challenge also to investigators. On the other side of the spectrum lies legitimate uses, which include circumvention of web censorship and surveillance, computer forensics (tracing and identification), and copyright protection (e.g., watermarking images). In this special issue, we are delighted to present a selection of nine papers, which, in our opinion, will contribute to the enhancement of knowledge in information hiding. The collection of high-quality research papers provides a view on the latest research advances on covert communication, steganography, and steganalysis. In the first paper, Multi-bit watermarking of high dynamic range images based on perceptual models, Emanuele Maiorana and Patrizio Campisi describe a multi-bit watermarking method dedicated to high dynamic range images. The proposed method takes advantage of various perceptual features of the human eye. The authors present the results of imperceptibility and robustness tests, using a database of 15 high dynamic range images. Also two next articles concern image processing. In MDE-based image steganography with large embedding capacity, Zhaoxia Yin and Bin Luo propose a steganographic method based on modification of direction exploitation and pixel pair matching. The algorithm is explained in detail, and a numerical example is given. The authors analyze also the quality of the conveyed covert image and evaluate security of the proposed algorithm, using two steganalysis methods. Fengyong Li, Xinpeng Zhang, Hang Cheng, and Jiang Yu in Digital image steganalysis based on local textural features and double dimensionality reduction also deal with steganalysis of image steganography. They propose a spatial steganalysis scheme based on local textural features and double dimensionality reduction. The authors demonstrate effectiveness of their method using 5000 greyscale images and three different steganographic techniques. Next three articles concern using audio signals for steganographic transmission. In the first of them, Real-time audio steganography attack based on automatic objective quality feedback, Qilin Qi, Aaron Sharp, Dongming Peng, and Hamid Sharif propose an active warden steganographic attack based on discrete spring transform with the use of an objective quality assessment. The authors show effectiveness of such an attack against two different steganographic techniques—spread spectrum-based and a time-scale modification-robust steganography. Shanyu Tang, Qing Chen, Wei Zhang, and Yongfeng Huang in Universal steganography model for low bit-rate speech codec describe a universal steganography model for low bit-rate speech codec. The proposed method is based on using perceptual evaluation of speech quality algorithm to choose a proper data hiding algorithm. The authors employ proposed approach for the Internet Speech Audio Codec and present results for the steganographic bandwidth and cost. Rennie Archibald and Dipak Ghosal in Design and performance evaluation of a covert timing channel discuss covert timing channels, in which the steganographic transmission is realized by modulating the inter-packet delay times. The authors propose a method, which minimizes overruns and underruns of an Internet Protocol phone buffer, and then evaluate its performance using Skype traffic. Pawel Laka and Lukasz Maksymiuk in Steganographic transmission in optical networks with the use of direct spread spectrum technique describe their method of steganographic transmission to be used in the physical layer of optical networks. The proposed method is based on the spread spectrum technique. The authors show results of their experiments with adjusting the spreading code length and optical power level dependencies. In the next article, On importance of steganographic cost for network steganography, an analysis of steganographic cost for network steganography is presented by Wojciech Mazurczyk, Steffen Wendzel, Ignacio Azagra Villares, and Krzysztof Szczypiorski. In this paper, the metric of steganographic cost is defined as degradation or a distortion of the carrier caused by the application of the steganographic method. The authors analyze various approaches to steganographic cost in selected single-method and multi-method steganographic techniques. In the last of the presented articles, Matrix embedding in multicast steganography: analysis in privacy, security and immediacy, Weiwei Liu, Guangjie Liu, and Yuewei Dai discuss multicast steganography, in which a single sender delivers simultaneously different secret messages to several receivers within the same cover object. The authors propose both synchronous and asynchronous multicast matrix embedding frameworks, based on Slepian–Wolf coding and overlapped multi-embedding, respectively. Privacy, security, and immediacy of the proposed solutions are also discussed. To summarize, we believe that this Special Issue will contribute to enhancing knowledge in Information and Communication Technology (ICT) security and in information hiding in particular. In addition, we also hope that the presented results will stimulate further research in the important areas of information and network security, including steganography and covert communication. We also want to thank the editor-in-chief of the Security and Communication Networks journal, the leading researchers contributing to the special issue and excellent reviewers for their great help and support that made this special issue possible.
Wojciech Mazurczyk, Krzysztof Szczypiorski, Artur Janicki, Hui Tian 0002
Secur. Commun. Networks3
2015 Novel Method of Hiding Information in IP Telephony Using Pitch Approximation
abstract
In this paper a novel steganographic method, called Hide F0, dedicated to IP telephony is proposed. It is based on the approximation of the parameter that describes the F0 frequency (the pitch) of the speaker's voice. We show that thanks to approximating some fragments of the "fine pitch" parameter in the Speex codec we can create efficient hidden transmission channels. We determined that for Speex working in mode 5 the Hide F0 method can provide a hidden channel with a capacity of ca. 220 bps at the optimal operating point. We also demonstrated that the proposed method offers a significantly more advantageous trade-off between the steganographic bandwidth and steganographic cost than the classic least significant bit (LSB) approach.
Artur Janicki
ARES1
2015 Spoofing countermeasure based on analysis of linear prediction error
abstract
In this paper a novel speaker verification spoofing countermeasure based on analysis of linear prediction error is presented. The method analyses the energy of the prediction error, prediction gain and temporal parameters related to the prediction error signal. The idea of the proposed algorithm and its implementation is described in detail. Various binary classifiers were researched to separate human and spoof classes. When tested on the corpora provided for the ASVspoof 2015 Challenge, the proposed countermeasure yielded much better results than the baseline spoofing detector based on local binary patterns (LBP). It is hoped that the proposed method can help in developing a generalised countermeasure able to detect spoofing attacks based on different variants of speech synthesis, voice conversion, and, potentially, also other spoofing algorithms. Index Terms: speaker verification, spoofing, linear prediction, local binary patterns, binary classification
Artur Janicki
INTERSPEECH1
2015 On the undetectability of transcoding steganography
abstract
Abstract Transcoding Steganography (TranSteg) is a fairly new IP telephony steganographic method that is characterized by a high steganographic bandwidth, low introduced distortions, and high undetectability. TranSteg utilizes compression of the overt data to free space for the secret data bits. In this paper, we focus on evaluating different possibilities for TranSteg detection. Building on the previous works, we perform a wide analysis of different steganalysis methods to assess the possibility of TranSteg detection and identify the most ‘undetectable’ pairs of voice codecs. Copyright © 2015 John Wiley & Sons, Ltd.
Artur Janicki, Wojciech Mazurczyk, Krzysztof Szczypiorski
Secur. Commun. Networks1
2013 Non-linguistic vocalisation recognition based on hybrid GMM-SVM approach
abstract
This paper describes an algorithm for detection of nonlinguistic vocalisations, such as laughter or fillers, based on acoustic features. The algorithm proposed combines the benefits of Gaussian mixture models (GMM) and the advantages of support vector machines (SVMs). Three GMMs were trained for garbage, laughter, and fillers, and then an SVM model was trained in the GMM score space. Various experiments were run to tune the parameters of the proposed algorithm, using the data sets originating from the SSPNet Vocalisation Corpus (SVC) provided for the Social Signals Sub-Challenge of the INTERSPEECH 2013 Computational Paralinguistics Challenge. The results showed a remarkable growth of the unweighted average of the area under the receiver operating curve (UAAUC) compared to the baseline results (from 87.6% to over 94% for the development set), which confirmed the efficiency of the proposed method. Index Terms: paralingustics, social signals, laughter detection, filler, support vector machines, Gaussian mixture models, cepstrum
Artur Janicki
INTERSPEECH1
2011 Automatic Speech Recognition for Polish in a Computer Game Interface
Artur Janicki, Dariusz Wawer
FedCSIS1
2005 Reconstruction of Polish diacritics in a text-to-speech system
Artur Janicki, Piotr Herman
INTERSPEECH1
2000 Automatic construction of acoustic inventory for the concatenative speech synthesis for polish
Artur Janicki
INTERSPEECH1