Filip Kruzel

dblp:04/8297 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-3462-9144ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Performance Analysis Of Parallel XOR And AES Encryption On Architectures Apple Silicon M4 Pro vs NVIDIA RTX 3070
abstract
This paper presents a comparative analysis of performance between two fundamentally different computing architectures: Apple M4 Pro with unified memory and Intel i5-8600K paired with an NVIDIA RTX 3070 discrete GPU. We evaluate both memory-bound (XOR) and compute-bound (AES-256-CTR) cryptography algorithms across sequential, OpenMP, Metal, and CUDA implementations. Our results reveal that the M4 Pro outperforms the Intel platform in most tasks, though the discrete GPU achieves faster XOR. This advantage stems from dedicated ARMv8 cryptographic instructions and a unified memory architecture eliminating PCIe transfer bottlenecks. Energy measurements show the M4 Pro consumes 15–30 watts during peak operation compared to 315 watts for the Intel and RTX 3070 system, resulting in superior energy efficiency for the Apple platform across all tested workloads.
Maciej Biegan, Mateusz Nytko, Filip Kruzel
ECMS3
2026 Benchmarking Transformer-Based Time Series Models: PyTorch vs MLX On Apple Silicon
abstract
We present a comparative evaluation of three transformer-based time series models—iTransformer, PatchTST, and FEDformer—across two deep learning frameworks: the original PyTorch implementations and newly developed ports to Apple’s MLX framework. Experiments are conducted on the ETTm1 electricity transformer dataset, covering long-term forecasting and anomaly detection tasks. Using identical hyperparameters and a consistent 80/20 chronological train–validation split, we measure training throughput, per-epoch wall-clock time, memory consumption, and predictive accuracy (RMSE, MAE for forecasting; precision, recall, F1, AUC for anomaly detection). Results indicate that MLX achieves substantially higher training throughput than PyTorch on the M4 Pro GPU, with speedups ranging from 1.02× to 3.12× across model architectures and batch sizes. Forecasting accuracy remains comparable between frameworks, with differences in RMSE below 6%. For anomaly detection, MLX attains higher F1 and AUC scores under identical configurations. These findings suggest that MLX is a viable alternative for time series model development on Apple Silicon, offering competitive accuracy with improved training efficiency. We attribute the observed speedups to architectural differences in memory handling and execution models, highlighting how framework design interacts with Apple Silicon’s unified memory architecture.
Kamil Dziedzic, Mateusz Nytko, Filip Kruzel
ECMS3
2026 Automated Blood Cell ClassificationUsing Computer Vision And Convolutional Neural Networks
abstract
Automated medical image analysis increasingly relies on models that are not only accurate but also robust, interpretable, and efficient enough for practical inference pipelines. This paper presents an end-to-end modelling and inference workflow for automated blood cell classification using convolutional neural networks trained on the BloodMNIST benchmark. We systematically compare classical machine learning baselines (logistic regression, SVM, random forests, MLP) with custom CNN architectures and a transfer-learning model (fine-tuned ResNet-18). On the BloodMNIST test set, the best model (fine-tuned ResNet-18) achieves 97.2\% accuracy and a macro-averaged F1-score of 0.969, outperforming the strongest classical baseline (0.87 accuracy, 0.86 macro-F1). To support explainable modelling, Grad-CAM visualisations are used to highlight image regions driving predictions. Robustness is assessed by modelling typical microscopy acquisition variability (rotations, colour jitter, noise, blur), showing only moderate degradation under realistic perturbations. Finally, we demonstrate deployability as an end-to-end system aspect via a RESTful API and a lightweight GUI enabling real-time single-image inference on commodity hardware.
Maciej Labuz, Filip Kruzel
ECMS2
2026 Music Recommendation System for Narrative Texts Based On Semantic And Emotional Analysis
abstract
The exponential growth of digital music resources requires advanced retrieval methods beyond traditional metadata. These limitations are particularly evident when correlating disparate domains, such as acoustic features and literary texts, where the high dimensionality of data poses significant computational challenges. This paper presents an innovative computational system that maps musical pieces to narrative passages by simulating human-like cognitive correlations. The architecture utilises three transformer models for feature extraction: all-MiniLM-L6-v2 for vector embeddings, RoBERTa for emotion classification, and DistilBERT for sentiment analysis. To address the substantial computational cost of processing a large-scale corpus of 649,078 tracks, a novel Cascaded Filtering Architecture was implemented. This multi-stage approach serves as a necessary optimisation to ensure system interactivity and second-level responsiveness. All strategies employ vector metrics, including cosine similarity and Euclidean distance, to quantify proximity in a constructed joint space. Experimental results demonstrate that the cascade approach achieves up to a 16-fold increase in computational efficiency compared to standard hybrid models, while maintaining high recommendation relevance across diverse narrative contexts. This work highlights the potential of scalable multimodal matching, offering new perspectives for interactive, high-performance recommendation systems in the digital literature and media.
Jakub Pedry, Filip Kruzel
ECMS2
2026 Automatic Analysis And Summarisation Of Court Auction Notices Using Natural Language Processing
abstract
Court auction notices constitute a highly formalised class of legal documents characterised by structural heterogeneity, procedural redundancy, and dispersed key information. This paper presents a computational pipeline for the automatic analysis and extractive summarisation of such notices using classical machine learning methods. The proposed system integrates domain-aware preprocessing, token-level semantic classification, and supervised sentence relevance modelling based on TF–IDF features enriched with structural indicators. A one-vs-rest logistic regression classifier is employed to identify informative, highly relevant auction content. The approach is evaluated on real-world Polish court auction data. An ablation study confirms the contribution of manually engineered features, and a manual evaluation involving 250 summary assessments demonstrates improved readability and information coverage compared to baseline methods. The system can be interpreted as an information-reduction and decision-support component that enables structured data extraction for downstream analytical and simulation-based applications.
Grzegorz Piasny, Filip Kruzel
ECMS2
2026 Implementing Lightweight ARX Ciphers on a 6502-Based 8-Bit Platform: A Position Study on Performance and Code-Size Trade-Offs
Dominik Madej, Filip Kruzel
SECRYPT (1)2
2025 Intel Xe Architecture Automatic Parameter Tuning With FEM Numerical Integration
abstract
This article analyses the usability of Intel ARC 770 based on Intel Xe Architecture for scientific computational tasks based on the Numerical Integration in Finite Element Method (FEM). To achieve a thorough comparison, we employed our proprietary auto-tuning algorithm, allowing us to assess architectural differences between Intel’s solution and the widely adopted Nvidia architecture. As a reference point, we selected the GeForce RTX 3060 Laptop due to its comparable performance and similar price at the time of release. Intel’s latest GPU architecture represents a significant step in the company’s ongoing efforts to establish a foothold in the GPGPU market, challenging Nvidia’s longstanding dominance. While our benchmarking results indicate that the ARC 770 delivers performance on par with reference Nvidia architecture, we identified distinctive behavioural characteristics that set it apart from previously analyzed architectures. These unique attributes could affect computational efficiency and workload distribution, making Intel’s approach an intriguing alternative for scientific computing. Through a series of experiments and simulations, we aim to provide deeper insight into how architectural variations influence the efficiency of FEM numerical integration. Given that the tested Intel GPU shares similarities with the professional Xe-HPC lineup, our research also explores its potential applicability beyond consumer-grade hardware, positioning it as a viable option for high-performance computing tasks.
Filip Kruzel, Mateusz Nytko
ECMS1
2025 Analysis Of Virtual Threads In Spring Applications
abstract
Java virtual threads are supposed to reduce the effort of creating high-throughput concurrent applications. This article presents an in-depth analysis of the behaviour of virtual threads in REST API applications created using the Spring framework. Through a series of performance evaluations and controlled simulations, we investigate how virtual threads manage the execution of concurrent tasks with differing computational demands. The tests focus on performance regarding the request processing time, providing a practical benchmark for understanding the scalability and responsiveness offered by virtual threads in real-world spring applications. These findings offer valuable insights for developers who want to optimize their applications for maximum concurrency and minimal resource consumption with a single configuration change.
Patryk Likus, Filip Kruzel
ECMS2
2024 Analysis Of Performance Differences Of FEM Numerical Integration Algorithm On Two Generations Of Intel Xe-LP GPUs
abstract
This article analyses the performance differences between two generations of integrated GPUs on the Finite Element Method (FEM) Numerical Integration Algorithm. The algorithm employs linear approximation in the convection-diffusion problem, where performance is highly memory-dependent. We used Intel 11th and 12th-generation processors with integrated GPUs based on Intel Xe-LP architecture to conduct our research. The GPUs have the same parameters except for the slightly lower bandwidth and a narrower bus in the newer architecture. This makes it an ideal choice to test the significance of the performance loss in the highly-parallel memory-bound algorithm. We also compare the two generations of GPUs in terms of their computational power, memory access patterns, and other relevant factors to identify all the sources of performance loss. Our research shows how a slight change in one parameter can affect the overall performance of an algorithm. Our experiments aim to provide insight into how different GPU architectures and generations affect FEM numerical integration performance.
Filip Kruzel, Mateusz Nytko
ECMS1
2023 Exploring The Benefits Of OpenCL Shared Virtual Memory: A Comparative Analysis On Integrated And External GPUs
abstract
In this article, we test the feature of using a Shared Virtual Memory in OpenCL on Intel Iris Xe and Nvidia GeForce RTX 3060 GPUs. The first of these architectures is integrated into the CPU, so by definition, it uses the same RAM as the CPU. The second one uses a PCI-Express connection to transfer data between computer memory and separate GPU memory. OpenCL Shared Virtual Memory (SVM) feature allows zero-copy mechanisms to use the same address space for the CPU and Accelerator. In this work, the authors test the differences between classical implementations of the FEM numerical integration algorithm on GPU with explicit data sending between CPU and Accelerator and the SVM implementation with the transfer hidden from the user. Research should answer whether the advantages of using a weaker integrated card with faster data transfer will overcome the external graphics card's connection bottleneck.
Filip Kruzel, Mateusz Nytko
ECMS1
2019 Vectorized Implementation Of The FEM Numerical Integration Algorithm On A Modern CPU
Filip Kruzel
ECMS1
2014 Finite Element Numerical Integration on Xeon Phi coprocessor
abstract
In the present article we describe the implementation of the finite element numerical integration algorithm for the Xeon Phi coprocessor.The coprocessor is an extension of the idea of the many-core specialized unit for calculations and, by assumption, its performance has to be competitive with the current families of GPUs.Its main advantage is the built-in set of 512-bit vector registers and the ease of transferring existing codes from normal x86 architectures.However, the differences between standard x86 architectures and Xeon Phi do not guarantee performance portability.We choose an alternative approach and, instead of porting standard multithreaded code, we adapt to Xeon Phi previously developed OpenCL algorithms for finite element numerical integration.The algorithm is tested for standard FEM approximations of selected problems.The obtained timing results allow to compare the performance of the OpenCL kernels executed on the Xeon Phi and the contemporary GPUs.
Filip Kruzel, Krzysztof Banas
FedCSIS1