Wlodzimierz Kasprzak

dblp:86/3294 · DBLP profile ↗
← Back
20ranked-venue papers
13as first author
4since 2021 · last 2023
0000-0002-4840-8860ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 8 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2023 Gender-aware speaker's emotion recognition based on 1-D and 2-D features
abstract
An approach to speaker's emotion recognition based on several acoustic feature types and 1D convolutional neural networks is described.The focus is on selecting the best speech features, improving the baseline model configuration and integrating in the solution a gender classification network.Features include a Mel-scale spectrogram and MFCC-, Chroma-, prosodic-and pitch-related features.Especially, the question whether to use 2-D maps of features or reduce them to 1-D vectors by averaging, is experimentally resolved.Well-known speech datasets RAVDESS, Tess, Crema-D and Savee are used in experiments.It appeared, that the best performing model consists of two convolutional networks for gender-aware classification and one gender classifier.The Chroma features have been found to be obsolete, and even disturbing, given other speech features.The f1 accuracy of proposed solution reached 73.2% on the RAVDESS dataset and 66.5% on all four datasets combined, improving the baseline model by 7.8% and 3%, respectively.This approach is a serious alternative to other proposed models, which reported accuracy scores of 60% -71% on the RAVDESS dataset.
Wlodzimierz Kasprzak, Mateusz Hryciów
FedCSIS1
2023 Automatic speaker's age classification in the Common Voice database
abstract
An approach to speaker's age classification using deep neural networks is described.Preliminary signal features are extracted, based on mel-frequency cepstral coefficients (MFCC).For gender classification an MLP network appears to be a satisfactory lightweight solution.For the age modelling and classification problem, two network types, ResNet34 and xvectors, were tested and compared.The impact of signal processing parameters and gender information (both theoretic perfect realistic imperfect) onto the age classification performance was experimentally studied.The neural networks were trained and verified on the large "Common Voice" dataset of English speech recordings.
Adam Nowakowski, Wlodzimierz Kasprzak
FedCSIS2
2022 A lightweight approach to two-person interaction classification in sparse image sequences
abstract
A lightweight neural network-based approach to two-person interaction classification in image sequences, based on human skeletons detected in sparse video frames, is proposed.The idea is to use an ensemble of pose classifiers ("experts"), where every expert is trained on different time-indexed snapshots of an interaction.Thus, the expertise of "weak" classifiers is distributed over the time duration of an interaction.The overall classification result is a weighted combination of all the pose experts.Important element of proposed solution is the refinement of skeleton data, based on a merging-of-joints procedure.This allows the generation of reliable features being passed to the artificial neural network.This is the key to our lightweight solution, as ANN resources, needed for feature space transformation, can be significantly limited.Our network model was trained and tested on the interaction subset of the well-known NTU RGB+D dataset, although only 2D skeleton information is used, typical in video analysis.The test results show comparable performance of our method with some of the best so far reported STMand CNN-based classifiers for this dataset, when they process sparse frame sequences, like we did.The recently proposed multistream Graph CNNs have shown superior results but only when processing dense frame sequences.Considering the dominating processing time and resources needed for skeleton estimation in every frame of the sequence, the key to real-time interaction recognition is to limit the number of processed frames.
Wlodzimierz Kasprzak, Van Khanh Do, Pawel Piwowarski
FedCSIS1
2022 Feature engineering techniques for skeleton-based two-person interaction classification in video
abstract
An LSTM-based approach to two-person interaction classification in image sequences, based on human skeleton detection in single frames, is proposed. The main contribution is the design of different preliminary feature extraction techniques from many skeleton data sets. The approach consists of four main stages. First, we employ the OpenPose engine in order to estimate silhouettes of persons in single video frames. Next, the skeleton data sets are tracked over time and particular joints are corrected whenever needed. Then, three different versions of features are defined, based on the stream of refined skeleton joints. In the final stage, a classifier network, based on LSTM, is developed and experimentally studied. The models are trained and evaluated on the interaction subset of the NTU RGB+D data set. The most efficient model achieved a classification accuracy of 94.5% (in the CV mode). This is an excellent performance, comparable with the best reported Graph CNN-based classifiers for this subset.
Sebastian Puchala, Wlodzimierz Kasprzak, Pawel Piwowarski
ICARCV2
2019 Object detection in the police surveillance scenario
abstract
Police and various security services use video analysis when investigating criminal activity.One typical scenario is the selection of object in image sequence and search for similar objects in other images.Algorithms supporting this scenario must reconcile several seemingly contradicting factors: training and detection speed, detection reliability and learning from sparse data.In the system that we propose a combined SVM/Cascade detector is used for both speed and detection reliability.In addition, object tracking and background-foreground separation algorithm together with sample synthesis is used to collect rich training data.Experiments show that the system is effective, useful and suitable for selected tasks of police surveillance.
Artur Wilkowski, Wlodzimierz Kasprzak, Maciej Stefanczyk
FedCSIS2
2014 A hierarchical CSP search for path planning of cooperating self-reconfigurable mobile fixtures
Wlodzimierz Kasprzak, Wojciech Szynkiewicz, Dimiter Zlatanov, Teresa Zielinska
Eng. Appl. Artif. Intell.1
2013 The generation of letter-to-sound rules for grapheme-to-phoneme conversion
abstract
This paper presents an approach to letter-to-sound translation for the Polish language that is a part of a speech recognition system. It describes the process of automatic generation of Polish letter-to-sound (LTS) rules. The LTS rules were trained with a Polish phonetic lexicon, that was extracted from the “wictionary” - a Polish on-line dictionary. This lexicon contains 35.826 entries. We examined a novel method for creating the letter-to-phone allowable pairing, that applies the “IBM Model 1 algorithm. Such automatically generated allowed letter-to-sound pairs were compared with a second pairing map, created by an expert. Both allowable pairing maps were used separately to train the Polish LTS rules. The test results verify that our generated pairing map leads to a more compact LTS model than the expert-made one.
Pawel Przybysz, Wlodzimierz Kasprzak
HSI2
2012 Auditory scene analysis by time-delay analysis with three microphones
abstract
We propose two methods for the disambiguation of results in time-delay based detection and localization of sound sources, when a triangle of microphones is applied for signal acquisition. A standard approach is to create histograms of time differences of arrival (TDOA) for each microphone pair in a triangular array and to create an averaged histogram. But each individual histogram is designed to detect unique orientation of source only within the local range of [-π/2, π/2]. Hence, taking the average for different pairs is not appropriate and such method suffers from ambiguity of results in the full range of orientations: [0, 2π]. Our first proposition is a delay vector transformation method, that combines corresponding delay measurements into vectors and transforms them into a 2-D space in which a full-range orientation histogram can finally be established and analyzed. In our second method, individual orientation histograms obtained for pairs of microphones are analyzed first and for each detected source two competitive hypotheses are created. Due to a final clustering of the hypothesis set a unique orientation of each source can be estimated.
Nozomu Hamada, Wlodzimierz Kasprzak, Pawel Przybysz
ETFA2
2005 Global Color Image Features for Discrete Self-localization of an Indoor Vehicle
Wlodzimierz Kasprzak, Wojciech Szynkiewicz, Mikolaj Karolczak
CAIP1
2001 An Iconic Classification Scheme for Video-Based Traffic Sensor Tasks
Wlodzimierz Kasprzak
CAIP1
1999 Neural networks for blind separation with unknown number of sources
Andrzej Cichocki, Juha Karhunen, Wlodzimierz Kasprzak, Ricardo Vigário
Neurocomputing3
1998 Adaptive Road Recognition and Ego-state Tracking in the Presence of Obstacles
Wlodzimierz Kasprzak, Heinrich Niemann
Int. J. Comput. Vis.1
1997 On Neural Blind Separation with Noise Suppression and Redundancy Reduction
abstract
Noise is an unavoidable factor in real sensor signals. We study how additive and convolutive noise can be reduced or even eliminated in the blind source separation (BSS) problem. Particular attention is paid to cases in which the number of sensors is larger than the number of sources. We propose various methods and associated adaptive learning algorithms for such an extended BSS problem. Performance and validity of the proposed approaches are demonstrated by extensive computer simulations.
Juha Karhunen, Andrzej Cichocki, Wlodzimierz Kasprzak, Petteri Pajunen
Int. J. Neural Syst.3
1997 Blind Source Separation with Convolutive Noise Cancellation
Wlodzimierz Kasprzak, Andrzej Cichocki, Shun-ichi Amari
Neural Comput. Appl.1
1996 Recurrent least square learning for quasi-parallel principal component analysis
Wlodzimierz Kasprzak, Andrzej Cichocki
ESANN1
1996 Hidden image separation from incomplete image mixtures by independent component analysis
abstract
It is known that the independent component analysis (ICA) (also called blind source separation) can be applied only if the number of received signals (sensors) is at least equal to the number of mixed sources, contained in the sensor signals. In this paper an application of the ICA is proposed for hidden (secured) image transmission by communication channels. We assume that only a single image mixture is transmitted. A friendly receiver contains the remaining original sources and therefore it can separate the hidden image of lowest energy. The influence of two nonlossless signal reduction stages, compression by principal component analysis and signal quantization, onto the separation ability is tested. Constraints of the mixing process are discussed that make impossible the hidden image separation without the key images.
Wlodzimierz Kasprzak, Andrzej Cichocki
ICPR1
1994 Adaptive Road Parameter Estimation in Monocular Image Sequence
Wlodzimierz Kasprzak, Heinrich Niemann, Dirk Wetzel
BMVC1
1994 Adaptive Estimation Procedures for Dynamic Road Scene Analysis
abstract
The common task of several processing steps in a vision system for autonomous road vehicle guidance is to stabilize the single image measurements of following image or scene features: contour motion, vanishing point, road class, moving object state and egomotion. An adaptive estimation procedure with linear or extended Kalman filter is applied for adaptive feature estimation. Redundant measurements and weighted averaging of short sequence measurements are proposed for robust detection and measurement error estimation.>
Wlodzimierz Kasprzak, Heinrich Niemann, Dirk Wetzel
ICIP (1)1
1993 Visual Motion Estimation from Image Contour Tracking
Wlodzimierz Kasprzak, Heinrich Niemann
CAIP1
1987 A linguistic approach to 3-D object recognition
Wlodzimierz Kasprzak
Comput. Graph.1