Pavel Zemcík

dblp:69/3392 · DBLP profile ↗
← Back
59ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-7969-5877ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 3 since 2021Systems, architecture and hardware · 7 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 How color profile affects image and video compression efficiency
Tomás Chlubna, Pavel Zemcík
Multim. Tools Appl.2
2026 Focus-aware compression and image quality metric for 3D displays
Tomás Chlubna, Michal Vlnas, David Barina, Tomás Milet, Pavel Zemcík
Signal Process.5
2026 Survey of FOSS 3D/2D graphics software blender usage in science, academia, and industry
abstract
Abstract Free and open-source software (FOSS) is a preferred tool for individuals and companies. The advantages of FOSS are minimal expenses, multi-platform and community support, transparent privacy policies, no vendor-related limitations, etc. Blender is a FOSS computer graphics 3D and 2D editor. It offers functions such as modeling, animating, video editing, simulations, image processing, scripting, rendering, etc. It is widely used as a free alternative to existing commercial products. This comprehensive survey examines Blender and its features and explores its usage in research, academic, and industrial projects, based on hundreds of collected and referenced sources. A comparison of Blender with alternative proprietary tools was conducted in terms of rendering performance, popularity, support, and feature set. According to this survey, Blender can be used as an efficient tool in film industry for visual effects composition, for dataset production for scientific experiments or deep learning methods, for educational purposes as a 3D geometrical problems demonstration tool, for the design of industrial models prepared, for example, for 3D printing, usage in augmented or virtual reality applications, etc. More specialized features are available as community-developed add-ons. The main reason why Blender is not used more often is that many professionals are used to other software.
Tomás Chlubna, Michal Vlnas, Tomás Milet, Pavel Zemcík
Vis. Comput.4
2025 Unsupervised Mineral Segmentation with Graph Neural Networks and Multi-modal SEM Data
Samuel Repka, Tuomas Eerola, David Motl, Jakub Výravský, Pavel Zemcík
CAIP (2)5
2025 How color profile affects the visual quality in light field rendering and novel view synthesis
Tomás Chlubna, Tomás Milet, Pavel Zemcík
Multim. Tools Appl.3
2025 Mineral segmentation using electron microscope images and spectral sampling through multimodal graph neural networks
abstract
We propose a novel Graph Neural Network-based method for segmentation based on data fusion of multimodal Scanning Electron Microscope (SEM) images. In most cases, Backscattered Electron (BSE) images obtained using SEM do not contain sufficient information for mineral segmentation. Therefore, imaging is often complemented with point-wise Energy-Dispersive X-ray Spectroscopy (EDS) spectral measurements that provide highly accurate information about the chemical composition but that are time-consuming to acquire. This motivates the use of sparse spectral data in conjunction with BSE images for mineral segmentation. The unstructured nature of the spectral data makes most traditional image fusion techniques unsuitable for BSE-EDS fusion. We propose using graph neural networks to fuse the two modalities and segment the mineral phases simultaneously. Our results demonstrate that providing EDS data for as few as 1% of BSE pixels produces accurate segmentation, enabling rapid analysis of mineral samples. The proposed data fusion pipeline is versatile and can be adapted to other domains that involve image data and point-wise measurements.
Samuel Repka, Borek Reich, Fedor Zolotarev, Tuomas Eerola, Pavel Zemcík
Pattern Recognit. Lett.5
2025 Light field video streaming on GPU
Tomás Chlubna, Tomás Milet, Pavel Zemcík
Signal Process. Image Commun.3
2025 Low-Error Reconstruction of Directional Functions With Spherical Harmonics
abstract
This paper proposes a novel approach for the low-error reconstruction of directional functions with spherical harmonics. We introduce a modified version of Spherical Gaussians with adaptive narrowness and amplitude to represent the input data in an intermediate form. This representation is then projected into spherical harmonics using a closed-form analytical solution. Because of the spectral properties of the proposed representation, the amount of ringing artifacts is reduced, and the overall precision of the reconstructed function is improved. The proposed method is more precise comparing to existing methods. The presented solution can be used in several graphical applications, as discussed in this paper. For example, the method is suitable for sparse models such as indirect illumination or reflectance functions.
Michal Vlnas, Tomás Milet, Pavel Zemcík
IEEE Trans. Vis. Comput. Graph.3
2025 Out-of-focus artifacts mitigation and autofocus methods for 3D displays
abstract
This paper proposes a novel content-aware method for automatic focusing of the scene on a 3D display. The method addresses a common problem that visualized content is often out of focus, which adversely affects perceived 3D content. The method outperforms existing focusing method, having the error lower by almost 30%. The existing and novel focusing is extended with depth-of-field enhancement of the scene to mitigate out-of-focus artifacts. The relation between the total depth range of the scene and the visual quality of the result is discussed and evaluated according to human perception experiments. A space-warping method for synthetic scenes is proposed to reduce out-of-focus artifacts while maintaining the scene appearance. A user study was conducted to evaluate the proposed methods and identify the crucial parameters in the scene-focusing process on the 3D stereoscopic display by Looking Glass Factory. The study confirmed the efficiency of the proposals and discovered that the depth-of-field artifact mitigation might not be suitable for all scenes despite theoretical hypotheses. The overall proposal of this paper is a set of methods that can be used to produce the best user experience with an arbitrary scene displayed on a 3D display.
Tomás Chlubna, Tomás Milet, Pavel Zemcík
Vis. Informatics3
2024 Lightweight all-focused light field rendering
Tomás Chlubna, Tomás Milet, Pavel Zemcík
Comput. Vis. Image Underst.3
2024 Efficient random-access GPU video decoding for light-field rendering
Tomás Chlubna, Tomás Milet, Pavel Zemcík
J. Vis. Commun. Image Represent.3
2024 How Capturing Camera Trajectory Distortion Affects User Experience on Looking Glass 3D Display
Tomás Chlubna, Tomás Milet, Pavel Zemcík
Multim. Tools Appl.3
2024 Automatic 3D-display-friendly scene extraction from video sequences and optimal focusing distance identification
abstract
Abstract This paper proposes a method for an automatic detection of 3D-display-friendly scenes from video sequences. Manual selection of such scenes by a human user would be extremely time consuming and would require additional evaluation of the result on 3D display. The input videos can be intentionally captured or taken from other sources, such as films. First, the input video is analyzed and the camera trajectory is estimated. The optimal frame sequence that follows defined rules, based on optical attributes of the display, is then extracted. This ensures the best visual quality and viewing comfort. The following identification of a correct focusing distance is an important step to produce a sharp and artifact-free result on a 3D display. Two novel and equally efficient focus metrics for 3D displays are proposed and evaluated. Further scene enhancements are proposed to correct the unsuitably captured video. Multiple image analysis approaches used in the proposal are compared in terms of both quality and time performance. The proposal is experimentally evaluated on a state-of-the-art 3D display by Looking Glass Factory and is suitable even for other multi-view devices. The problem of optimal scene detection, which includes the input frames extraction, resampling, and focusing, was not addressed in any previous research. Separate stages of the proposal were compared with existing methods, but the results show that the proposed scheme is optimal and cannot be replaced by other state-of-the-art approaches.
Tomás Chlubna, Tomás Milet, Pavel Zemcík
Multim. Tools Appl.3
2022 Comparison of light field compression methods
David Barina, Marek Solony, Tomás Chlubna, Drahomir Dlabaja, Ondrej Klíma 0002, Pavel Zemcík
Multim. Tools Appl.6
2022 Vehicle Speed Measurement Using Stereo Camera Pair
abstract
We have proposed a novel method for vehicle speed estimation using a calibrated and synchronized pair of stereo cameras. In a newly proposed method, we first localize the vehicle by detecting and tracking its license plate in a series of stereo images; then, we triangulate the vehicle position along its trajectory; and finally, we compute its speed based on the trajectory and time. The experiments show that the proposed method overcomes state-of-the-art results with a mean error of approximately 0.05 km/h, a standard deviation of less than 0.20 km/h, and a maximum absolute error of less than 0.75 km/h. For the purpose of evaluation, we have recorded a dataset that contains over 600 vehicles whose trajectories were recorded and for which their ground truth speed was obtained from a pair of single beam LIDARs in optical gate configuration. Using the presented method, the speed was measured for over 99 % of the recorded vehicles. Others were rejected by the method mainly due to their short trajectories, obstructed license plates or frame errors that would adversely affect the precision of the measurement.
Pavel Najman, Pavel Zemcík
IEEE Trans. Intell. Transp. Syst.2
2021 Unconstrained License Plate Detection in Hardware
Petr Musil, Roman Juránek, Pavel Zemcík
VEHITS3
2021 Real-time per-pixel focusing method for light field rendering
abstract
Light field rendering is an image-based rendering method that does not use 3D models but only images of the scene as input to render new views. Light field approximation, represented as a set of images, suffers from so-called refocusing artifacts due to different depth values of the pixels in the scene. Without information about depths in the scene, proper focusing of the light field scene is limited to a single focusing distance. The correct focusing method is addressed in this work and a real-time solution is proposed for focusing of light field scenes, based on statistical analysis of the pixel values contributing to the final image. Unlike existing techniques, this method does not need precomputed or acquired depth information. Memory requirements and streaming bandwidth are reduced and real-time rendering is possible even for high resolution light field data, yielding visually satisfactory results. Experimental evaluation of the proposed method, implemented on a GPU, is presented in this paper.
Tomás Chlubna, Tomás Milet, Pavel Zemcík
Comput. Vis. Media3
2020 Fire Segmentation in Still Images
Jozef Mlích, Karel Koplík, Michal Hradis, Pavel Zemcík
ACIVS4
2020 Cascaded Stripe Memory Engines for Multi-Scale Object Detection in FPGA
abstract
Object detection in embedded systems is important for many contemporary applications that involve vision and scene analysis. In this paper, we propose a novel architecture for object detection implemented in FPGA, based on the Stripe Memory Engine (SME), and point out shortcomings of existing architectures. SME processes a stream of image data so that it stores a narrow stripe of the input image and its scaled versions and uses a detector unit which is efficiently pipelined across multiple image positions within the SME. We show how to process images with up to 4K resolution at high frame rates using cascades of SMEs. As a detector algorithm, the SMEs use boosted soft cascade with simple image features that require only pixel comparisons and look-up tables; therefore, they are well suitable for hardware implemenation. We describe the components of our architecture and compare it to several published works in several configurations. As an example, we implemented face detection and license plate detection applications that work with HD images (1280 ×720 pixels) running at over 60 frames/s on Xilinx Zynq platform. We analyzed their power consumption, evaluated the accuracy of our detectors, and compared them to Haar Cascades from OpenCV that are often used by other authors. We show that our detectors offer better accuracy as well as performance at lower power consumption.
Petr Musil, Roman Juránek, Martin Musil, Pavel Zemcík
IEEE Trans. Circuits Syst. Video Technol.4
2019 Comprehensive Data Set for Automatic Single Camera Visual Speed Measurement
abstract
In this paper, we focus on traffic camera calibration and a visual speed measurement from a single monocular camera, which is an important task of visual traffic surveillance. Existing methods addressing this problem are difficult to compare due to a lack of a common data set with reliable ground truth. Therefore, it is not clear how the methods compare in various aspects and what factors are affecting their performance. We captured a new data set of 18 full-HD videos, each around 1 hr long, captured at six different locations. Vehicles in the videos (20 865 instances in total) are annotated with the precise speed measurements from optical gates using LiDAR and verified with several reference GPS tracks. We made the data set available for download and it contains the videos and metadata (calibration, lengths of features in image, annotations, and so on) for future comparison and evaluation. Camera calibration is the most crucial part of the speed measurement; therefore, we provide a brief overview of the methods and analyze a recently published method for fully automatic camera calibration and vehicle speed measurement and report the results on this data set in detail.
Jakub Sochor, Roman Juránek, Jakub Spanhel, Lukas Marsik, Adam Siroký, Adam Herout, Pavel Zemcík
IEEE Trans. Intell. Transp. Syst.7
2018 Interactive Spatial Augmented Reality in Collaborative Robot Programming: User Experience Evaluation
abstract
This paper presents a novel approach to interaction between human workers and industrial collaborative robots. The proposed approach addresses problems introduced by existing solutions for robot programming. It aims to reduce the mental demands and attention switches by centering all interaction in a shared workspace, combining various modalities and enabling interaction with the system without any external devices. The concept allows simple programming in the form of setting program parameters using spatial augmented reality for visualization and a touch-enabled table and robotic arms as input devices. We evaluated the concept utilizing a user experience study with six participants (shop-floor workers). All participants were able to program the robot and to collaborate with it using the program they parametrized. The final goal is to create a distraction-free, usable and low-effort interface for effective human-robot collaboration, enabling any ordinary skilled worker to customize the robot's program to changes in production or to personal (e.g. ergonomic) needs.
Zdenek Materna, Michal Kapinus, Vítezslav Beran, Pavel Smrz, Pavel Zemcík
RO-MAN5
2018 Comparison of bubble detectors and size distribution estimators
abstract
Detection, counting and characterization of bubbles, that is, transparent objects in a liquid, is important in many industrial applications. These applications include monitoring of pulp delignification and multiphase dispersion processes common in the chemical, pharmaceutical, and food industries. Typically the aim is to measure the bubble size distribution. In this paper, we present a comprehensive comparison of bubble detection methods for challenging industrial image data. Moreover, we compare the detection-based methods to a direct bubble size distribution estimation method that does not require the detection of individual bubbles. The experiments showed that the approach based on a convolutional neural network (CNN) outperforms the other methods in detection accuracy. However, the boosting-based approaches were remarkably faster to compute. The power spectrum approach for direct bubble size distribution estimation produced accurate distributions and it is fast to compute, but it does not provide the spatial locations of the bubbles. Selecting the most suitable method depends on the specific application.
Jarmo Ilonen, Roman Juránek, Tuomas Eerola, Lasse Lensu, Markéta Dubská, Pavel Zemcík, Heikki Kälviäinen
Pattern Recognit. Lett.6
2017 Holistic recognition of low quality license plates by CNN using track annotated data
abstract
This work is focused on recognition of license plates in low resolution and low quality images. We present a methodology for collection of real world (non-synthetic) dataset of low quality license plate images with ground truth transcriptions. Our approach to the license plate recognition is based on a Convolutional Neural Network which holistically processes the whole image, avoiding segmentation of the license plate characters. Evaluation results on multiple datasets show that our method significantly outperforms other free and commercial solutions to license plate recognition on the low quality data. To enable further research of low quality license plate recognition, we make the datasets publicly available.
Jakub Spanhel, Jakub Sochor, Roman Juránek, Adam Herout, Lukas Marsik, Pavel Zemcík
AVSS6
2017 Accelerating discrete wavelet transforms on GPUS
abstract
The two-dimensional discrete wavelet transform has a huge number of applications in image-processing techniques. Until now, several papers compared the performance of such transform on graphics processing units (GPUs). However, all of them only dealt with lifting and convolution computation schemes. In this paper, we show that corresponding horizontal and vertical lifting parts of the lifting scheme can be merged into non-separable lifting units, which halves the number of steps. We also discuss an optimization strategy leading to a reduction in the number of arithmetic operations. The schemes were assessed using the OpenCL and pixel shaders. The proposed non-separable lifting scheme outperforms the existing schemes in many cases, irrespective of its higher complexity.
David Barina, Michal Kula, Michal Matysek, Pavel Zemcík
ICIP4
2017 Absolute pose estimation from line correspondences using direct linear transformation
Bronislav Pribyl, Pavel Zemcík, Martin Cadík
Comput. Vis. Image Underst.2
2016 Single-Loop Software Architecture for JPEG 2000
abstract
JPEG 2000 is an image coding system based on the discrete wavelet transform. Unfortunately, there exist several major issues with the effective implementation of the codec. For high resolution data decomposed by a separable transform, immensely many CPU cache misses occur. Following the procedure as defined in the standard, the coefficients of a single resolution appears all at once. Consequently, the entropy coder (EBCOT) needs to once again return to the data already touched. We present a software architecture designed for JPEG 2000 coders. The proposed method employs a strip-based data processing technique while it performs a single-pass multi-scale wavelet transform. The overall compression chain is driven by incoming data while the fragments of the resulting bitstream can be produced immediately after loading the corresponding data and additionally in parallel. The method is friendly to the CPU cache and can nicely exploit the SIMD extensions.
David Barina, Ondrej Klíma 0002, Pavel Zemcík
DCC3
2016 CNN for license plate motion deblurring
abstract
In this work we explore the previously proposed approach of direct blind deconvolution and denoising with convolutional neural networks (CNN) in a situation where the blur kernels are partially constrained. We focus on blurred images from a real-life traffic surveillance system, on which we, for the first time, demonstrate that neural networks trained on artificial data provide superior reconstruction quality on real images compared to traditional blind deconvolution methods. The training data is easy to obtain by blurring sharp photos from a target system with a very rough approximation of the expected blur kernels, thereby allowing custom CNNs to be trained for a specific application (image content and blur range). Additionally, we evaluate the behavior and limits of the CNNs with respect to blur direction range and length.
Pavel Svoboda, Michal Hradis, Lukas Marsik, Pavel Zemcík
ICIP4
2016 Single-Loop Architecture for JPEG 2000
David Barina, Ondrej Klíma 0002, Pavel Zemcík
ICISP3
2016 Evaluation of feature point detection in high dynamic range imagery
Bronislav Pribyl, Alan Chalmers, Pavel Zemcík, Lucy Hooberman, Martin Cadík
J. Vis. Commun. Image Represent.3
2016 Big Data Analysis for Media Production
abstract
A typical high-end film production generates several terabytes of data per day, either as footage from multiple cameras or as background information regarding the set (laser scans, spherical captures, etc). This paper presents solutions to improve the integration of the multiple data sources, and understand their quality and content, which are useful both to support creative decisions on-set (or near it) and enhance the postproduction process. The main cinema specific contributions, tested on a multisource production dataset made publicly available for research purposes, are the monitoring and quality assurance of multicamera set-ups, multisource registration and acceleration of 3-D reconstruction, anthropocentric visual analysis techniques for semantic content annotation, and integrated 2-D–3-D web visualization tools. We discuss as well improvements carried out in basic techniques for acceleration, clustering and visualization, which were necessary to deal with the very large multisource data, and can be applied to other big data problems in diverse application fields.
Josep Blat, Alun Evans, Hansung Kim 0001, Evren Imre, Lukás Polok, Viorela Ila, Nikos Nikolaidis 0001, Pavel Zemcík, Anastasios Tefas, Pavel Smrz, Adrian Hilton 0001, Ioannis Pitas
Proc. IEEE8
2015 Depth-Based Filtration for Tracking Boost
David Chrapek, Vítezslav Beran, Pavel Zemcík
ACIVS3
2015 Convolutional Neural Networks for Direct Text Deblurring
abstract
In this work we address the problem of blind deconvolution and denoising. We focus on restoration of text documents and we show that this type of highly structured data can be successfully restored by a convolutional neural network. The networks are trained to reconstruct high-quality images directly from blurry inputs without assuming any specific blur and noise models. We demonstrate the performance of the convolutional networks on a large set of text documents and on a combination of realistic de-focus and camera shake blur kernels. On this artificial data, the convolutional networks significantly outperform existing blind deconvolution methods, including those optimized for text, in terms of image quality and OCR accuracy. In fact, the networks outperform even state-of-the-art non-blind methods for anything but the lowest noise levels. The approach is validated on real photos taken by various devices.
Michal Hradis, Jan Kotera, Pavel Zemcík, Filip Sroubek
BMVC3
2015 Camera Pose Estimation from Lines using Plücker Coordinates
abstract
Correspondences between 3D lines and their 2D images captured by a camera are often used to determine position and orientation of the camera in space. In this work, we propose a novel algebraic algorithm to estimate the camera pose. We parameterize 3D lines using Pl\"ucker coordinates that allow linear projection of the lines into the image. A line projection matrix is estimated using Linear Least Squares and the camera pose is then extracted from the matrix. An algebraic approach to handle mismatched line correspondences is also included. The proposed algorithm is an order of magnitude faster yet comparably accurate and robust to the state-of-the-art, it does not require initialization, and it yields only one solution. The described method requires at least 9 lines and is particularly suitable for scenarios with 25 and more lines, as also shown in the results.
Bronislav Pribyl, Pavel Zemcík, Martin Cadík
BMVC2
2015 Real-Time Pose Estimation Piggybacked on Object Detection
abstract
We present an object detector coupled with pose estimation directly in a single compact and simple model, where the detector shares extracted image features with the pose estimator. The output of the classification of each candidate window consists of both object score and likelihood map of poses. This extension introduces negligible overhead during detection so that the detector is still capable of real time operation. We evaluated the proposed approach on the problem of vehicle detection. We used existing datasets with viewpoint/pose annotation (WCVP, 3D objects, KITTI). Besides that, we collected a new traffic surveillance dataset COD20k which fills certain gaps of the existing datasets and we make it public. The experimental results show that the proposed approach is comparable with state-of-the-art approaches in terms of accuracy, but it is considerably faster - easily operating in real time (Matlab with C++ code). The source codes and the collected COD20k dataset are made public along with the paper.
Roman Juránek, Adam Herout, Markéta Dubská, Pavel Zemcík
ICCV4
2015 Quality assurance in large collections of video sequences
abstract
In the modern digital cinema production, extremely large volumes (in order of 10s of TB) of footage data are captured every day. The process of cataloging and reviewing such footage is nowadays largely manual and time consuming process. In our work, we aim at technical quality aspects, such as correct exposure, color compatibility of adjacent shots, and focusing. The main goal is to assist the reviewing process by providing shot quality meta-data and possibly ordering or even culling significant portion of the data from the review, as better quality shot of the same scene exists. However, in order to meaningfully compare technical quality, temporal shot synchronization needs to be performed first. We propose a fast and robust method for time synchronization of video sequences, capturing similar scenes, which arise naturally in digital cinema production. The method is tested on an extensive library of sequences and its performance is evaluated. We further present a preview of application of the proposed method to detecting focusing errors.
Lukás Polok, Lukas Klicnar, Vítezslav Beran, Pavel Smrz, Pavel Zemcík
ICIP5
2015 Fast covariance recovery in incremental nonlinear least square solvers
abstract
Many estimation problems in robotics rely on efficiently solving nonlinear least squares (NLS). For example, it is well known that the simultaneous localisation and mapping (SLAM) problem can be formulated as a maximum likelihood estimation (MLE) and solved using NLS, yielding a mean state vector. However, for many applications recovering only the mean vector is not enough. Data association, active decisions, next best view, are only few of the applications that require fast state covariance recovery. The problem is not simple since, in general, the covariance is obtained by inverting the system matrix and the result is dense. The main contribution of this paper is a novel algorithm for fast incremental covariance update, complemented by a highly efficient implementation of the covariance recovery. This combination yields to two orders of magnitude reduction in computation time, compared to the other state of the art solutions. The proposed algorithm is applicable to any NLS solver implementation, and does not depend on incremental strategies described in our previous papers, which are not a subject of this paper.
Viorela Ila, Lukás Polok, Marek Solony, Pavel Smrz, Pavel Zemcík
ICRA5
2014 Diagonal vectorisation of 2-D wavelet lifting
abstract
With the start of the widespread use of discrete wavelet transform in image processing, the need for its efficient implementation is becoming increasingly more important. This work presents a novel SIMD vectorisation of 2-D discrete wavelet transform through a lifting scheme. For all of the tested platforms, this vectorisation is significantly faster than other known methods, as shown in the results of the experiments.
David Barina, Pavel Zemcík
ICIP2
2014 Human Action Recognition for Real-time Applications
abstract
Action recognition in video is an important part of many applications. While the performance of action recognition has been intensively investigated, not much research so far has been done in understanding of how long sequence of video frames is needed to correctly recognize certain actions. This paper presents a new method of measurement of the necessary length of the video sequence needed to recognize the actions based on space-time feature points - the critical information necessary to successfully recognize the actions in real-time. The proposed methods is experimentally evaluated on human action recognition dataset.
Ivo Reznícek, Pavel Zemcík
ICPRAM2
2013 Minimum Memory Vectorisation of Wavelet Lifting
David Barina, Pavel Zemcík
ACIVS2
2013 High performance architecture for object detection in streamed video (abstract only)
abstract
Object detection is one of the key tasks in computer vision. It is computationally intensive and it is reasonable to accelerate it in hardware. The possible benefits of the acceleration are reduction of the computational load of the host computer system, increase of the overall performance of the applications, and reduction of the power consumption. We present novel architecture for multi-scale object detection in video streams. The architecture uses scanning window classifiers produced by WaldBoost learning algorithm, and simple image features. It employs small image buffer for data under processing, and on-the-fly scaling units to enable detection of object in multiple scales. The whole processing chain is pipelined and thus more image windows are processed in parallel. We implemented the engine in Spartan 6 FPGA and we show that it can process 640x480 pixel video streams at over 160 frames per second without the need of external memory. The design takes only a fraction of resources, compared to similar state of the art approaches.
Pavel Zemcík, Roman Juránek, Petr Musil, Martin Musil, Michal Hradis
FPGA1
2013 High performance architecture for object detection in streamed videos
abstract
In this paper, we introduce a novel architecture of an engine for high performance multi-scale detection of objects in videos based on WaldBoost training algorithm. The key properties of the architecture include processing of streamed data and low resource consumption. We implemented the engine in FPGA and we show that it can process 640×480 pixel video streams at over 160 fps without the need of external memory. We evaluate the design on the face detection task, compare it to state of the art designs, and discuss its features and limitations.
Pavel Zemcík, Roman Juránek, Petr Musil, Martin Musil, Michal Hradis
FPL1
2013 High performance FPGA object detector: Hardware prototype
abstract
Summary form only given. In this demo, we introduce a novel architecture of an engine for high performance multi-scale detection of objects in videos based on WaldBoost training algorithm. The key properties of the architecture include processing of streamed data and low resource consumption. We implemented the engine in FPGA and we show that it can process 640 × 480 pixel video streams at over 160 FPS without the need of external memory.
Pavel Zemcík, Roman Juránek, Petr Musil, Martin Musil, Michal Hradis
FPL1
2013 Efficient implementation for block matrix operations for nonlinear least squares problems in robotic applications
abstract
A large number of robotic, computer vision and computer graphics applications rely on efficiently solving the associated sparse linear systems. Simultaneous localization and mapping (SLAM), structure from motion (SfM), non-rigid shape recovery, and elastodynamic simulations are only few examples in this direction. In general, these problems are nonlinear and the solution can be approximated by incrementally solving a series of linearized problems. In some applications, the size of the system considerably affects the performance, especially when the sparsity is low. This paper exploits the block structure of such problems and offers very efficient solutions to manipulate block matrices within iterative nonlinear solvers. The resulting method considerably speeds-up the execution of the implementation of the nonlinear optimization problem. In this work, in particular, we focus our effort on testing the method on SLAM applications, but the applicability of the technique remains general. Our implementation outperforms the state of the art SLAM implementations on all tested datasets. In incremental mode, where a larger portion of time is spent in updating the system, our implementation is on average two times faster than the others.
Lukás Polok, Marek Solony, Viorela Ila, Pavel Smrz, Pavel Zemcík
ICRA5
2012 Annotating Images with Suggestions - User Study of a Tagging System
Michal Hradis, Martin Kolár, Ales Láník, Jirí Král, Pavel Zemcík, Pavel Smrz
ACIVS5
2012 GPU Optimization of Convolution for Large 3-D Real Images
Pavel Karas, David Svoboda, Pavel Zemcík
ACIVS3
2012 Fast bilateral filter for HDR imaging
Michal Seeman, Pavel Zemcík, Roman Juránek, Adam Herout
J. Vis. Commun. Image Represent.2
2012 EnMS: early non-maxima suppression - Speeding up pattern localization and other tasks
Adam Herout, Michal Hradis, Pavel Zemcík
Pattern Anal. Appl.3
2011 Analysis of Wear Debris through Classification
Roman Juránek, Stanislav Machalík, Pavel Zemcík
ACIVS3
2011 Simple Single View Scene Calibration
Bronislav Pribyl, Pavel Zemcík
ACIVS2
2011 Nonnegative Tensor Factorization Accelerated Using GPGPU
abstract
This article presents an optimized algorithm for Nonnegative Tensor Factorization (NTF), implemented in the CUDA (Compute Uniform Device Architecture) framework, that runs on contemporary graphics processors and exploits their massive parallelism. The NTF implementation is primarily targeted for analysis of high-dimensional spectral images, including dimensionality reduction, feature extraction, and other tasks related to spectral imaging; however, the algorithm and its implementation are not limited to spectral imaging. The speedups measured on real spectral images are around 60-100{\times} compared to a traditional C implementation compiled with an optimizing compiler. Since common problems in the field of spectral imaging may take hours on a state-of-the-art CPU, the speedup achieved using a graphics card is attractive. The implementation is publicly available in the form of a dynamically linked library, including an interface to MATLAB, and thus may be of help to researchers and engineers using NTF on large problems.
Jukka Antikainen, Jirí Havel, Radovan Josth, Adam Herout, Pavel Zemcík, Markku Hauta-Kasari
IEEE Trans. Parallel Distributed Syst.5
2010 Exploiting Neighbors for Faster Scanning Window Detection in Images
Pavel Zemcík, Michal Hradis, Adam Herout
ACIVS (2)1
2010 Multi-resolution Next Location Prediction for Distributed Virtual Environments
abstract
Computing performance of today's graphics hardware grows fast as well as amount of rendered data. Modern graphics engines enable a possibility to use an arbitrary number of textures with arbitrary resolutions. On the other hand, high quality distributed 3D virtual environments can't exploit the computation power due to the limited network bandwidth. The problem mainly appears just in case the designers of such environments use high resolution textures. To overcome this streaming bottleneck an efficient prefetching scheme should be proposed. Instead of blind greedy scheduling policy we propose a scheme which exploits movement history of users to realize a look-ahead policy which enables the clients to retrieve potentially rendered data in advance. The prediction itself is established by Markov chains due to their ability to fast learning in conjunction with 2-state predictor which increases ability of the scheduling system to adapt to new habits of particular users.
Jaroslav Pribyl, Pavel Zemcík
EUC2
2010 Rendering fur directly into images
Tania Pouli, Martin Prazák, Pavel Zemcík, Diego Gutierrez, Erik Reinhard
Comput. Graph.3
2008 "Local Rank Differences" Image Feature Implemented on GPU
Lukás Polok, Adam Herout, Pavel Zemcík, Michal Hradis, Roman Juránek, Radovan Josth
ACIVS3
2007 AdaBoost Engine
abstract
This paper presents an application specific engine dedicated for acceleration of AdaBoost image classifier. AdaBoost and its modifications belong to the most successful algorithms of image classification. This class of algorithms can also be used for object detection through scanning of the image with a sliding window whose content is classified using AdaBoost. Such process, however, is very computationally demanding. The engine presented in this paper implements a novel feature extraction method, suitable specifically for hardware acceleration, whose classification performance is at the same time equal or better than performance of the more traditionally used features. The novel feature extraction is based on simultaneous processing of a small grid of picture elements that can be accessed from a memory in a single read operation. The architecture of engine utilizes principles of finemultithreading combinedwith pipelining. Preliminary tests of the engine on Xilinx Virtex II - 250 show better results comparing to existing hardware implementations of image classification in accuracy, speed, and chip utilization.
Pavel Zemcík, Martin Zádník
FPL1
2000 Multispectral Image Color Encoding
abstract
Advances in colour image sensors and computer technology have opened up new opportunities for colour-based image analysis by allowing it to use multispectral images - images that have tens or even hundreds of spectral colour channels. As multispectral images usually occupy large amounts of memory, some suitable way is needed to compress such images and to represent them efficiently. Image compression has been one of the mainstream research topics for a long time; however, the research usually focuses on compressing images that are intended to be finally seen by humans. While many traditional methods can also be reused for multispectral images, some features of the multispectral images can be addressed differently than in traditional images. This paper describes a way of how to represent the multispectral colour data as a linear combination of colours contained in the clusters of colours that most frequently occur in an image.
Pavel Zemcík, Jan Vorácek, Michael Frydrych, Heikki Kälviäinen, Pekka J. Toivanen
ICPR1
2000 Compression of multispectral remote sensing images using clustering and spectral reduction
abstract
Image compression has been one of the main research topics in the field of image processing for a long time. The research usually focuses on compressing images that are visible to humans. The images being compressed are usually gray-level images or RGB color images. Recent advances in technology, however, enable the authors to make the detailed processing of spectral features in the images. Therefore, the compression of images with many spectral channels, called multispectral images, is required. Many methods used in traditional lossy image compression can be reused also in the compression of multispectral images. In this paper, a new combination of clustering spectra, manipulating spectral vectors, and encoding and decoding for multispectral images is presented. In the manipulation of the spectral vectors PCA, ICA, and wavelets are used. The approach is based on extracting relevant spectral information. Furthermore, some quantitative quality measures for multispectral images are presented.
Arto Kaarna, Pavel Zemcík, Heikki Kälviäinen, Jussi Parkkinen
IEEE Trans. Geosci. Remote. Sens.2
1998 Multispectral image compression
abstract
Image compression has been one of the mainstream research topics in image processing. The research usually focuses on compressing images that are visible to humans. Images are usually gray-level images or RGB color images. Advances in technology enable one to make the detailed processing of spectral color features in the images. Therefore, compression of images with many spectral color channels, called multispectral images, is required. Many methods used in traditional lossy image compression can be reused also in the compression of multispectral images. In this paper a new combination of clustering of colors, manipulating spectral color encoding and decoding for multispectral images is presented. The approach is based on extracting relevant color information. Furthermore, some quantitative quality measures for multispectral images are presented.
Arto Kaarna, Pavel Zemcík, Heikki Kälviäinen, Jussi Parkkinen
ICPR2
1995 Optimised CSG Tree Evaluation for Space Subdivision
abstract
Abstract Ray tracing is a well known technique for producing realistic computer images. The computational requirements of this method are such that optimisation techniques, for example space subdivision, must be used if complex scenes are to be rendered in reasonable times. Constructive solid geometry (CSG) is a method for describing the geometry of complex scenes by applying set operations to primitive objects. The status tree approach has been used successfully within ray tracing to evaluate CSG structures. This paper proposes a combination of the status tree and space subdivision techniques as a means to improve further the efficiency of ray tracing.
Pavel Zemcík, Alan Chalmers
Comput. Graph. Forum1