Michal Hradis

dblp:31/4629 · DBLP profile ↗
← Back
29ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-6364-129XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 11 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-authorSystems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents
Martin Kostelník, Michal Hradis, Martin Docekal
ICDAR (2)2
2025 BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
Jan Kohút, Martin Docekal, Michal Hradis, Marek Vasko
ICDAR (3)3
2025 Practical Fine-Tuning of Autoregressive Models on Limited Handwritten Texts
Jan Kohút, Michal Hradis
ICDAR (5)2
2025 BenCzechMark : A Czech-Centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism
abstract
Abstract We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple evaluation metrics. Its duel scoring system is grounded in statistical significance theory and uses aggregation across tasks inspired by social preference theory. Our benchmark encompasses 50 challenging tasks, with corresponding test datasets, primarily in native Czech, with 14 newly collected ones. These tasks span 8 categories and cover diverse domains, including historical Czech news, essays from pupils or language learners, and spoken word. Furthermore, we collect and clean BUT-Large Czech Collection, the largest publicly available clean Czech language corpus, and use it for (i) contamination analysis and (ii) continuous pretraining of the first Czech-centric 7B language model with Czech-specific tokenization. We use our model as a baseline for comparison with publicly available multilingual models. Lastly, we release and maintain a leaderboard with existing 50 model submissions, where new model submissions can be made at https://huggingface.co/spaces/CZLC/BenCzechMark.
Martin Fajcik, Martin Docekal, Jan Dolezal, Karel Ondrej, Karel Benes, Jan Kapsa, Pavel Smrz, Alexander Polok, Michal Hradis, Zuzana Neverilová, Ales Horák, Radoslav Sabol, Michal Stefánik, Adam Jirkovsky, David Adamczyk, Petr Hyner, Jan Hula, Hynek Kydlícek
Trans. Assoc. Comput. Linguistics9
2024 Self-supervised Pre-training of Text Recognizers
Martin Kiss, Michal Hradis
ICDAR (4)2
2024 SoftCTC - semi-supervised learning for text recognition using soft pseudo-labels
Martin Kiss, Michal Hradis, Karel Benes, Petr Buchal, Michal Kula
Int. J. Document Anal. Recognit.2
2023 Fine-Tuning is a Surprisingly Effective Domain Adaptation Baseline in Handwriting Recognition
Jan Kohút, Michal Hradis
ICDAR (4)2
2023 Towards Writing Style Adaptation in Handwriting Recognition
Jan Kohút, Michal Hradis, Martin Kiss
ICDAR (4)2
2022 Importance of Textlines in Historical Document Classification
Martin Kiss, Jan Kohút, Karel Benes, Michal Hradis
DAS4
2022 $\hbox {TG}^2$: text-guided transformer GAN for restoring document readability and perceived quality
Oldrich Kodym, Michal Hradis
Int. J. Document Anal. Recognit.2
2021 AT-ST: Self-training Adaptation Strategy for OCR in Domains with Limited Transcriptions
abstract
This paper addresses text recognition for domains with limited manual annotations by a simple self-training strategy. Our approach should reduce human annotation effort when target domain data is plentiful, such as when transcribing a collection of single person's correspondence or a large manuscript. We propose to train a seed system on large scale data from related domains mixed with available annotated data from the target domain. The seed system transcribes the unannotated data from the target domain which is then used to train a better system. We study several confidence measures and eventually decide to use the posterior probability of a transcription for data selection. Additionally, we propose to augment the data using an aggressive masking scheme. By self-training, we achieve up to 55 % reduction in character error rate for handwritten data and up to 38 % on printed data. The masking augmentation itself reduces the error rate by about 10 % and its effect is better pronounced in case of difficult handwritten data.
Martin Kiss, Karel Benes, Michal Hradis
ICDAR (4)3
2021 Page Layout Analysis System for Unconstrained Historic Documents
Oldrich Kodym, Michal Hradis
ICDAR (2)2
2021 TS-Net: OCR Trained to Switch Between Text Transcription Styles
abstract
Users of OCR systems, from different institutions and scientific disciplines, prefer and produce different transcription styles. This presents a problem for training of consistent text recognition neural networks on real-world data. We propose to extend existing text recognition networks with a Transcription Style Block (TSB) which can learn from data to switch between multiple transcription styles without any explicit knowledge of transcription rules. TSB is an adaptive instance normalization conditioned by identifiers representing consistently transcribed documents (e.g. single document, documents by a single transcriber, or an institution). We show that TSB is able to learn completely different transcription styles in controlled experiments on artificial data, it improves text recognition accuracy on large-scale real-world data, and it learns semantically meaningful transcription style embedding. We also show how TSB can efficiently adapt to transcription styles of new documents from transcriptions of only a few text lines.
Jan Kohút, Michal Hradis
ICDAR (4)2
2020 Fire Segmentation in Still Images
Jozef Mlích, Karel Koplík, Michal Hradis, Pavel Zemcík
ACIVS3
2020 OCR, Classification& Machine Translation (OCCAM)
abstract
The OCCAM project (Optical Character recognition, ClassificAtion & Machine Translation) aims at integrating the CEF (Connecting Europe Facility) Automated Translation service with image classification, Translation Memories (TMs), Optical Character Recognition (OCR), and Machine Translation (MT). It will support the automated translation of scanned business documents (a document format that, currently, cannot be processed by the CEF eTranslation service) and will also lead to a tool useful for the Digital Humanities domain.
Joachim Van den Bogaert, Arne Defauw, Frederic Everaert, Koen Van Winckel, Alina Kramchaninova, Anna Bardadym, Tom Vanallemeersch, Pavel Smrz, Michal Hradis
EAMT9
2019 Brno Mobile OCR Dataset
abstract
We introduce the Brno Mobile OCR Dataset (B-MOD) for document Optical Character Recognition from low-quality images captured by handheld mobile devices. While OCR of high-quality scanned documents is a mature field where many commercial tools are available, and large datasets of text in the wild exist, no existing datasets can be used to develop and test document OCR methods robust to non-uniform lighting, image blur, strong noise, built-in denoising, sharpening, compression and other artifacts present in many photographs from mobile devices. This dataset contains 2 113 unique pages from random scientific papers, which were photographed by multiple people using 23 different mobile devices. The resulting 19 728 photographs of various visual quality are accompanied by precise positions and text annotations of 500k text lines. We further provide an evaluation methodology, including an evaluation server and a testset with non-public annotations. We provide a state-of-the-art text recognition baseline build on convolutional and recurrent neural networks trained with Connectionist Temporal Classification loss. This baseline achieves 2 %, 22 % and 73 % word error rates on easy, medium and hard parts of the dataset, respectively, confirming that the dataset is challenging. The presented dataset will enable future development and evaluation of document analysis for low-quality images. It is primarily intended for line-level text recognition, and can be further used for line localization, layout analysis, image restoration and text binarization.
Martin Kiss, Michal Hradis, Oldrich Kodym
ICDAR2
2018 MixedEmotions: An Open-Source Toolbox for Multimodal Emotion Analysis
abstract
Recently, there is an increasing tendency to embed functionalities for recognizing emotions from user-generated media content in automated systems such as call-centre operations, recommendations, and assistive technologies, providing richer and more informative user and content profiles. However, to date, adding these functionalities was a tedious, costly, and time-consuming effort, requiring identification and integration of diverse tools with diverse interfaces as required by the use case at hand. The MixedEmotions Toolbox leverages the need for such functionalities by providing tools for text, audio, video, and linked data processing within an easily integrable plug-and-play platform. These functionalities include: 1) for text processing: emotion and sentiment recognition; 2) for audio processing: emotion, age, and gender recognition; 3) for video processing: face detection and tracking, emotion recognition, facial landmark localization, head pose estimation, face alignment, and body pose estimation; and 4) for linked data: knowledge graph integration. Moreover, the MixedEmotions Toolbox is open-source and free. In this paper, we present this toolbox in the context of the existing landscape, and provide a range of detailed benchmarks on standard test-beds showing its state-of-the-art performance. Furthermore, three real-world use cases show its effectiveness, namely, emotion-driven smart TV, call center monitoring, and brand reputation analysis.
Paul Buitelaar, Ian D. Wood, Sapna Negi, Mihael Arcan, John P. McCrae, Andrejs Abele, Cécile Robin, Vladimir Andryushechkin, Housam Ziad, Hesam Sagha, Maximilian Schmitt, Björn W. Schuller, J. Fernando Sánchez-Rada, Carlos Angel Iglesias, Carlos Navarro, Andreas Giefer, Nicolaus Heise, Vincenzo Masucci, Francesco A. Danza, Ciro Caterino, Pavel Smrz, Michal Hradis, Filip Povolný, Marek Klimes, Pavel Matejka, Giovanni Tummarello
IEEE Trans. Multim.22
2016 CNN for license plate motion deblurring
abstract
In this work we explore the previously proposed approach of direct blind deconvolution and denoising with convolutional neural networks (CNN) in a situation where the blur kernels are partially constrained. We focus on blurred images from a real-life traffic surveillance system, on which we, for the first time, demonstrate that neural networks trained on artificial data provide superior reconstruction quality on real images compared to traditional blind deconvolution methods. The training data is easy to obtain by blurring sharp photos from a target system with a very rough approximation of the expected blur kernels, thereby allowing custom CNNs to be trained for a specific application (image content and blur range). Additionally, we evaluate the behavior and limits of the CNNs with respect to blur direction range and length.
Pavel Svoboda, Michal Hradis, Lukas Marsik, Pavel Zemcík
ICIP2
2015 Camera Elevation Estimation from a Single Mountain Landscape Photograph
abstract
This work addresses the problem of camera elevation estimation from a single photograph in an outdoor environment. We introduce a new benchmark dataset of one-hundred thousand images with annotated camera elevation called Alps100K. We propose and experimentally evaluate two automatic data-driven approaches to camera elevation estimation: one based on convolutional neural networks, the other on local features. To compare the proposed methods to human performance, an experiment with 100 subjects is conducted. The experimental results show that both proposed approaches outperform humans and that the best result is achieved by their combination.
Martin Cadík, Jan Vasícek, Michal Hradis, Filip Radenovic, Ondrej Chum
BMVC3
2015 Convolutional Neural Networks for Direct Text Deblurring
abstract
In this work we address the problem of blind deconvolution and denoising. We focus on restoration of text documents and we show that this type of highly structured data can be successfully restored by a convolutional neural network. The networks are trained to reconstruct high-quality images directly from blurry inputs without assuming any specific blur and noise models. We demonstrate the performance of the convolutional networks on a large set of text documents and on a combination of realistic de-focus and camera shake blur kernels. On this artificial data, the convolutional networks significantly outperform existing blind deconvolution methods, including those optimized for text, in terms of image quality and OCR accuracy. In fact, the networks outperform even state-of-the-art non-blind methods for anything but the lowest noise levels. The approach is validated on real photos taken by various devices.
Michal Hradis, Jan Kotera, Pavel Zemcík, Filip Sroubek
BMVC1
2013 High performance architecture for object detection in streamed video (abstract only)
abstract
Object detection is one of the key tasks in computer vision. It is computationally intensive and it is reasonable to accelerate it in hardware. The possible benefits of the acceleration are reduction of the computational load of the host computer system, increase of the overall performance of the applications, and reduction of the power consumption. We present novel architecture for multi-scale object detection in video streams. The architecture uses scanning window classifiers produced by WaldBoost learning algorithm, and simple image features. It employs small image buffer for data under processing, and on-the-fly scaling units to enable detection of object in multiple scales. The whole processing chain is pipelined and thus more image windows are processed in parallel. We implemented the engine in Spartan 6 FPGA and we show that it can process 640x480 pixel video streams at over 160 frames per second without the need of external memory. The design takes only a fraction of resources, compared to similar state of the art approaches.
Pavel Zemcík, Roman Juránek, Petr Musil, Martin Musil, Michal Hradis
FPGA5
2013 High performance architecture for object detection in streamed videos
abstract
In this paper, we introduce a novel architecture of an engine for high performance multi-scale detection of objects in videos based on WaldBoost training algorithm. The key properties of the architecture include processing of streamed data and low resource consumption. We implemented the engine in FPGA and we show that it can process 640×480 pixel video streams at over 160 fps without the need of external memory. We evaluate the design on the face detection task, compare it to state of the art designs, and discuss its features and limitations.
Pavel Zemcík, Roman Juránek, Petr Musil, Martin Musil, Michal Hradis
FPL5
2013 High performance FPGA object detector: Hardware prototype
abstract
Summary form only given. In this demo, we introduce a novel architecture of an engine for high performance multi-scale detection of objects in videos based on WaldBoost training algorithm. The key properties of the architecture include processing of streamed data and low resource consumption. We implemented the engine in FPGA and we show that it can process 640 × 480 pixel video streams at over 160 FPS without the need of external memory.
Pavel Zemcík, Roman Juránek, Petr Musil, Martin Musil, Michal Hradis
FPL5
2012 Annotating Images with Suggestions - User Study of a Tagging System
Michal Hradis, Martin Kolár, Ales Láník, Jirí Král, Pavel Zemcík, Pavel Smrz
ACIVS1
2012 What do you want to do next: a novel approach for intent prediction in gaze-based interaction
abstract
Interaction intent prediction and the Midas touch have been a longstanding challenge for eye-tracking researchers and users of gaze-based interaction. Inspired by machine learning approaches in biometric person authentication, we developed and tested an offline framework for task-independent prediction of interaction intents. We describe the principles of the method, the features extracted, normalization methods, and evaluation metrics. We systematically evaluated the proposed approach on an example dataset of gaze-augmented problem-solving sessions. We present results of three normalization methods, different feature sets and fusion of multiple feature types. Our results show that accuracy of up to 76% can be achieved with Area Under Curve around 80%. We discuss the possibility of applying the results for an online system capable of interaction intent prediction.
Roman Bednarik, Hana Vrzakova, Michal Hradis
ETRA3
2012 Voice activity detection from gaze in video mediated communication
abstract
This paper discusses estimation of active speaker in multi-party video-mediated communication from gaze data of one of the participants. In the explored settings, we predict voice activity of participants in one room based on gaze recordings of a single participant in another room. The two rooms were connected by high definition, low delay audio and video links and the participants engaged in different activities ranging from casual discussion to simple problem-solving games. We treat the task as a classification problem. We evaluate several types of features and parameter settings in the context of Support Vector Machine classification framework. The results show that using the proposed approach vocal activity of a speaker can be correctly predicted in 89 % of the time for which the gaze data are available.
Michal Hradis, Shahram Eivazi, Roman Bednarik
ETRA1
2012 EnMS: early non-maxima suppression - Speeding up pattern localization and other tasks
Adam Herout, Michal Hradis, Pavel Zemcík
Pattern Anal. Appl.2
2010 Exploiting Neighbors for Faster Scanning Window Detection in Images
Pavel Zemcík, Michal Hradis, Adam Herout
ACIVS (2)2
2008 "Local Rank Differences" Image Feature Implemented on GPU
Lukás Polok, Adam Herout, Pavel Zemcík, Michal Hradis, Roman Juránek, Radovan Josth
ACIVS4