VLDB 2026 Research / reviewers in the wild / expert
Ladislav Lenc
dblp:19/9854
· DBLP profile ↗
36ranked-venue papers
11as first author
14since 2021 · last 2025
0000-0002-1066-7269ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 9 first-author · 12 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Language Models for Summarizing Czech Historical Documents and BeyondabstractText summarization is the task of shortening a larger body of text into a concise version while retaining its essential meaning and key information. While summarization has been significantly explored in English and other high-resource languages, Czech text summarization, particularly for historical documents, remains underexplored due to linguistic complexities and a scarcity of annotated datasets. Large language models such as Mistral and mT5 have demonstrated excellent results on many natural language processing tasks and languages. Therefore, we employ these models for Czech summarization, resulting in two key contributions: (1) achieving new state-of-the-art results on the modern Czech summarization dataset SumeCzech using these advanced models, and (2) introducing a novel dataset called Posel od Čerchova for summarization of historical Czech documents with baseline results. Together, these contributions provide a great potential for advancing Czech text summarization and open new avenues for research in Czech historical text processing. Václav Tran, Jakub Smíd, Jirí Martínek, Ladislav Lenc, Pavel Král |
ICAART (2) | 4 |
| 2025 | On self-supervision in historical handwritten document segmentationabstractAbstract Historical document analysis plays a crucial role in understanding and preserving our past. However, this task is often hindered by challenges such as limited annotated training data and the diverse nature of historical handwritten documents. In this paper, we explore the potential of self-supervised learning (SSL) in historical document analysis, with a particular focus on historical handwritten document segmentation, to overcome the need for extensive annotated data while enhancing efficiency and robustness. We present an overview of SSL methods suitable for historical document analysis and discuss their potential applications and benefits. Furthermore, we present an approach for SSL in the document domain, considering various setups, augmentations, and resolutions. We also provide experimental results that demonstrate its feasibility and effectiveness. Our findings indicate that most document segmentation tasks can be effectively addressed using SSL features, highlighting the potential of SSL to advance historical document analysis and pave the way for more efficient and robust document processing workflows. Josef Baloun, Martin Prantl, Ladislav Lenc, Jirí Martínek, Pavel Král |
Int. J. Document Anal. Recognit. | 3 |
| 2024 | COMICORDA: Dialogue Act Recognition in Comic BooksabstractDialogue act (DA) recognition is usually realized from a speech signal that is transcribed and segmented into text. However, only a little work in DA recognition from images exists. Therefore, this paper concentrates on this modality and presents a novel DA recognition approach for image documents, namely comic books. To the best of our knowledge, this is the first study investigating dialogue acts from comic books and represents the first steps to building a model for comic book understanding. The proposed method is composed of the following steps: speech balloon segmentation, optical character recognition (OCR), and DA recognition itself. We use YOLOv8 for balloon segmentation, Google Vision for OCR, and Transformer-based models for DA classification. The experiments are performed on a newly created dataset comprising 1,438 annotated comic panels. It contains bounding boxes, transcriptions, and dialogue act annotation. We have achieved nearly 98% average precision for speech balloon segmentation and exceeded the accuracy of 70% for the DA recognition task. We also present an analysis of dialogue structure in the comics domain and compare it with the standard DA datasets, representing another contribution of this paper. Jirí Martínek, Pavel Král, Ladislav Lenc, Josef Baloun |
LREC/COLING | 3 |
| 2024 | Heimatkunde: Dataset for Multi-Modal Historical Document Analysis
Josef Baloun, Václav Honzík, Ladislav Lenc, Jirí Martínek, Pavel Král |
ICAART (3) | 3 |
| 2023 | Towards Automatic Medical Report Classification in Czech
Pavel Pribán, Josef Baloun, Jirí Martínek, Ladislav Lenc, Martin Prantl, Pavel Král |
ICAART (3) | 4 |
| 2023 | FCN-Boosted Historical Map Segmentation with Little Training Data
Josef Baloun, Ladislav Lenc, Pavel Král |
ICDAR (1) | 2 |
| 2022 | Historical Map Toponym Extraction for Efficient Information Retrieval
Ladislav Lenc, Jirí Martínek, Josef Baloun, Martin Prantl, Pavel Král |
DAS | 1 |
| 2022 | Robust Grid Detection in Historical Map ImagesabstractThis paper presents a novel method for grid detection in historical maps. The approach is based on Hough transform accompanied with a sophisticated post-processing. They are applied to detect the grid that consists of graticule lines. It works without any training and does not require any annotated data. The proposed approach is very efficient in detecting the rectangular grid and the intersection points as shown in the international "MapSeg" segmentation competition, where it won the Task 3 with a significant margin. The robustness of the proposed method has been demonstrated by evaluating on another dataset composed of significantly different cadastral map images with excellent results. Josef Baloun, Ladislav Lenc, Pavel Král |
ICIP | 2 |
| 2022 | Weak supervision for Question Type Detection with large language modelsabstractInternational audience Jirí Martínek, Christophe Cerisara, Pavel Král, Ladislav Lenc, Josef Baloun |
INTERSPEECH | 4 |
| 2022 | Correction to: Building an efficient OCR system for historical documents with little training dataabstractWith the author(s)’ decision to order Open Choice, the copyright of the article changed on 3rd December 2020 to [The Authors] [2020] and the article is forthwith distributed under the terms of copyright. Jirí Martínek, Ladislav Lenc, Pavel Král |
Neural Comput. Appl. | 2 |
| 2022 | Well-calibrated confidence measures for multi-label text classification with a large number of labels
Lysimachos Maltoudoglou, Andreas Paisios, Ladislav Lenc, Jirí Martínek, Pavel Král, Harris Papadopoulos |
Pattern Recognit. | 3 |
| 2021 | ChronSeg: Novel Dataset for Segmentation of Handwritten Historical Chronicles
Josef Baloun, Pavel Král, Ladislav Lenc |
ICAART (2) | 3 |
| 2021 | ICDAR 2021 Competition on Historical Map Segmentation
Joseph Chazalon, Edwin Carlinet, Yizi Chen, Julien Perret, Bertrand Dumenieu, Clément Mallet, Thierry Géraud, Vincent Nguyen 0001, Josef Baloun, Ladislav Lenc, Pavel Král |
ICDAR (4) | 11 |
| 2021 | Dialogue Act Recognition Using Visual Information
Jirí Martínek, Pavel Král, Ladislav Lenc |
ICDAR (2) | 3 |
| 2020 | Re-Ranking for Writer Identification and Writer Retrieval
Simon Jordan, Mathias Seuret, Pavel Král, Ladislav Lenc, Jirí Martínek, Barbara Wiermann, Tobias Schwinger, Andreas K. Maier, Vincent Christlein |
DAS | 4 |
| 2020 | Improving Face Recognition Methods based on POEM FeaturesabstractObvyklý způsob použití POEM deskriptorů je vytvoření příznaků v pravidelných obdélníkových regionech, které pokrývají celý snímek. Příznaky jsou spojeny do jednoho vektoru, který reprezentuje snímek obličeje. V článku je navržena vylepšená metoda, která využívá automaticky detekované body pro vytvoření příznaků. Zároveň je použita komplexnější metoda pro porovnávání příznakových vektorů. Navržená metoda nalezne uplatnění zejména v případech, kdy je k dispozici omezené množství dat a použití např. neuronových sítí by proto bylo obtížné. Metoda je testována na třech standardních obličejových korpusech. Dosažené výsledky ukazují, že použití POEM deskriptorů a příznaků, vytvořených v automaticky detekovaných bodech, dosahuje výrazně lepších výsledků, než základní metody. Ladislav Lenc, Pavel Král |
ICAART (2) | 1 |
| 2020 | Building an efficient OCR system for historical documents with little training dataabstractAbstract As the number of digitized historical documents has increased rapidly during the last a few decades, it is necessary to provide efficient methods of information retrieval and knowledge extraction to make the data accessible. Such methods are dependent on optical character recognition (OCR) which converts the document images into textual representations. Nowadays, OCR methods are often not adapted to the historical domain; moreover, they usually need a significant amount of annotated documents. Therefore, this paper introduces a set of methods that allows performing an OCR on historical document images using only a small amount of real, manually annotated training data. The presented complete OCR system includes two main tasks: page layout analysis including text block and line segmentation and OCR. Our segmentation methods are based on fully convolutional networks, and the OCR approach utilizes recurrent neural networks. Both approaches are state of the art in the relevant fields. We have created a novel real dataset for OCR from Porta fontium portal. This corpus is freely available for research, and all proposed methods are evaluated on these data. We show that both the segmentation and OCR tasks are feasible with only a few annotated real data samples. The experiments aim at determining the best way how to achieve good performance with the given small set of data. We also demonstrate that obtained scores are comparable or even better than the scores of several state-of-the-art systems. To sum up, this paper shows a way how to create an efficient OCR system for historical documents with a need for only a little annotated training data. Jirí Martínek, Ladislav Lenc, Pavel Král |
Neural Comput. Appl. | 2 |
| 2019 | Hybrid Training Data for Historical Text OCRabstractCurrent optical character recognition (OCR) systems commonly make use of recurrent neural networks (RNN) that process whole text lines. Such systems avoid the task of character segmentation necessary for character-based approaches. A disadvantage of this approach is a need of a large amount of annotated data. This can be solved by sing generated synthetic data instead of costly manually annotated ones. Unfortunately, such data is often not suitable for historical documents particularly for quality reasons. This work presents a hybrid approach for generating annotated data for OCR at a low cost. We first collect a small dataset of isolated characters from historical document images. Then, we generate historical looking text lines from the generated characters. Another contribution lies in the design and implementation of an OCR system based on a convolutional-LSTM network. We first pre-train this system on hybrid data. Afterwards, the network is fine-tuned with real printed text lines. We demonstrate that this training strategy is efficient for obtaining state-of-the-art results. We also show that the score of the proposed system is comparable or even better in comparison to several state-of-the-art systems. Jirí Martínek, Ladislav Lenc, Pavel Král, Anguelos Nicolaou, Vincent Christlein |
ICDAR | 2 |
| 2019 | Multi-Lingual Dialogue Act Recognition with Deep Learning MethodsabstractThis paper deals with multi-lingual dialogue act (DA) recognition. The proposed approaches are based on deep neural networks and use word2vec embeddings for word representation. Two multi-lingual models are proposed for this task. The first approach uses one general model trained on the embeddings from all available languages. The second method trains the model on a single pivot language and a linear transformation method is used to project other languages onto the pivot language. The popular convolutional neural network and LSTM architectures with different set-ups are used as classifiers. To the best of our knowledge this is the first attempt at multi-lingual DA recognition using neural networks. The multi-lingual models are validated experimentally on two languages from the Verbmobil corpus. Jirí Martínek, Pavel Král, Ladislav Lenc, Christophe Cerisara |
INTERSPEECH | 3 |
| 2019 | Automatic face recognition with well-calibrated confidence measures
Charalambos Eliades, Ladislav Lenc, Pavel Král, Harris Papadopoulos |
Mach. Learn. | 2 |
| 2018 | Neural Networks for Multi-lingual Multi-label Document Classification
Jirí Martínek, Ladislav Lenc, Pavel Král |
ICANN (1) | 2 |
| 2018 | Semantic Space Transformations for Cross-Lingual Document Classification
Jirí Martínek, Ladislav Lenc, Pavel Král |
ICANN (1) | 2 |
| 2018 | Czech Text Document Corpus v 2.0
Pavel Král, Ladislav Lenc |
LREC | 2 |
| 2018 | On the effects of using word2vec representations in neural networks for dialogue act recognition
Christophe Cerisara, Pavel Král, Ladislav Lenc |
Comput. Speech Lang. | 3 |
| 2017 | Evaluation of Local Descriptors for Automatic Image Annotation
Ladislav Lenc |
ICAART (2) | 1 |
| 2017 | Two-Level Neural Network for Multi-label Document Classification
Ladislav Lenc, Pavel Král |
ICANN (2) | 1 |
| 2017 | Combination of Neural Networks for Multi-label Document Classification
Ladislav Lenc, Pavel Král |
NLDB | 1 |
| 2016 | Deep Neural Networks for Czech Multi-label Document Classification
Ladislav Lenc, Pavel Král |
CICLing (2) | 1 |
| 2016 | Genetic Algorithm for Weight Optimization in Descriptor based Face Recognition MethodsabstractThis paper presents a novel algorithm for weight optimization in descriptor based face recognition methods. We aim at the local texture features that are currently very popular in the face recognition (FR) field. Common concept in such methods is creating histograms of the operator values in rectangular image regions and concatenating them into one large vector called histogram sequence (HS). Usually the facial regions are given equal weight which does not correspond with the reality. We deal with this issue in this work and propose a novel method that optimizes the weights of the regions. The optimization method is based on a genetic algorithm (GA). We test the method together with the local binary patterns (LBP) and patterns of oriented edge magnitudes (POEM) descriptors. We evaluate our algorithms on two real-world corpora: Unconstrained facial images (UFI) database and FaceScrub database. The evaluation results show that the weighted methods outperform the non-weighted ones. The best achieved scores are 68.93% on the UFI database and 57.81% on the FaceScrub database. Ladislav Lenc |
ICAART (2) | 1 |
| 2016 | LBP features for breast cancer detectionabstractCancer is nowadays considered as one of the most dangerous diseases in the world. Especially, breast cancer represents for women the second most common type of cancer and is a main cause of cancer dead. This paper presents a novel method for breast cancer detection from mammographic images based on Local Binary Patterns (LBP). This approach successfully uses LBP based features with a classifier and thresholding. The proposed method is evaluated on a set composed of images extracted from MIAS and DDSM databases. We have experimentally shown that the proposed method is efficient and effective because the achieved accuracy is about 84%. Pavel Král, Ladislav Lenc |
ICIP | 2 |
| 2015 | Confidence Measure for Czech Document Classification
Pavel Král, Ladislav Lenc |
CICLing (2) | 2 |
| 2014 | A Composed Confidence Measure for Automatic Face Recognition in Uncontrolled EnvironmentabstractThis paper is focused on automatic face recognition in order to annotate people in photographs taken in completely uncontrolled environment. Recognition accuracy of the current approaches is not sufficient in this case and it is thus beneficial to improve the results. We would like to solve this issue by proposing a novel confidence measure method to identify the incorrectly classified examples at the output of our classifier. The proposed approach combines two measures based on the posterior probability and two ones based on the predictor features in a supervised way. The experiments show that the proposed approach is very efficient, because it detects almost all erroneous examples. Pavel Král, Ladislav Lenc |
ICAART (1) | 2 |
| 2013 | Face Recognition under Real-world Conditions
Ladislav Lenc, Pavel Král |
ICAART (2) | 1 |
| 2013 | Automatic Face Corpus Creation
Ladislav Lenc, Pavel Král |
ICAART (2) | 1 |
| 2013 | A combined SIFT/SURF descriptor for automatic face recognitionabstractThis paper deals with Automatic Face Recognition (AFR). A novel approach which combines the SIFT and SURF features for the face representation is proposed. The obtained combined SIFT/SURF descriptor is then used for face comparison by the adapted Kepenekci matching method. The proposed method is evaluated on the FERET and CTK corpora. The obtained recognition rates are 98.4% and 64.6% respectively. These recognition scores show that our approach outperforms significantly all other methods on these corpora. The differences between recognition error rates of the proposed approach and the second best one are 41% and 7% in relative value respectively. Ladislav Lenc, Pavel Král |
ICMV | 1 |
| 2011 | Automatic Face Recognition - Methods Improvement and Evaluation
Ladislav Lenc, Pavel Král |
ICAART (1) | 1 |