Berat Kurar-Barakat

dblp:221/8961 · also Berat Kurar, Berat Kurar Barakat · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
4since 2021 · last 2026
0000-0002-7240-7286ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author
YearPublicationVenuePosition
2026 Segmentation of Ink and Parchment in Dead Sea Scroll Fragments
abstract
Abstract The discovery of the Dead Sea Scrolls over sixty years ago is widely regarded as one of the greatest archaeological breakthroughs in modern history. Recent study of the scrolls presents ongoing computational challenges, including determining the provenance of fragments, clustering fragments based on their degree of similarity, and pairing fragments that originate from the same manuscript—all tasks that require focusing on individual letter and fragment shapes. This paper presents a computational method for segmenting ink and parchment regions in multispectral images of Dead Sea Scroll fragments. Using the newly developed Qumran Segmentation Dataset (QSD) consisting of 20 fragments, we apply multispectral thresholding to isolate ink and parchment regions based on their unique spectral signatures. To refine segmentation accuracy, we introduce an energy minimization technique that leverages ink contours, which are more distinguishable from the background and less noisy than inner ink regions. Experimental results demonstrate that this Multispectral Thresholding and Energy Minimization (MTEM) method achieves significant improvements over traditional binarization approaches like Otsu and Sauvola in parchment segmentation and is successful at delineating ink borders, in distinction from holes and background regions.
Berat Kurar-Barakat, Nachum Dershowitz
Int. J. Document Anal. Recognit.1
2022 Hard and Soft Labeling for Hebrew Paleography: A Case Study
Ahmad Droby, Daria Vasyutinsky Shapira, Irina Rabaev, Berat Kurar-Barakat, Jihad El-Sana
DAS4
2021 Unsupervised Learning of Text Line Segmentation by Differentiating Coarse Patterns
Berat Kurar-Barakat, Ahmad Droby, Raid Saabni, Jihad El-Sana
ICDAR (2)1
2021 VML-HP: Hebrew Paleography Dataset
Ahmad Droby, Berat Kurar-Barakat, Daria Vasyutinsky Shapira, Irina Rabaev, Jihad El-Sana
ICDAR (4)2
2020 Unsupervised Deep Learning for Handwritten Page Segmentation
abstract
Segmenting handwritten document images into regions with homogeneous patterns is an important pre-processing step for many document images analysis tasks. Hand-labeling data to train a deep learning model for layout analysis requires significant human effort. In this paper, we present an unsupervised deep learning method for page segmentation, which revokes the need for annotated images. A siamese neural network is trained to differentiate between patches using their measurable properties such as number of foreground pixels, and average component height and width. The network is trained that spatially nearby patches are similar. The network's learned features are used for page segmentation, where patches are classified as main and side text based on the extracted features. We tested the method on a dataset of handwritten document images with quite complex layouts. Our experiments show that the proposed unsupervised method is as effective as typical supervised methods.
Ahmad Droby, Berat Kurar-Barakat, Boraq Madi, Reem Alaasam, Jihad El-Sana
ICFHR2
2020 The HHD Dataset
abstract
Benchmark datasets are important in document image processing field, as they allow to analyze different approaches and compare their performances in a fair manner. There exist benchmark datasets for several alphabets such as Latin, Arabic and Chinese, but not the Hebrew alphabet. In this paper, a handwritten Hebrew dataset, HHD, is introduced. The HHD dataset is collected from hand-filled forms, and accompanied by their ground truth at character, word and text line levels. Presently, the dataset contains around 1000 document images, and we continue to further enlarge it. To the best of our knowledge, this is the first comprehensive corpus of Hebrew handwritten documents, and we believe it will help leveraging Hebrew documents processing and document processing in general. The dataset can be useful for various research applications, such as word spotting, word recognition, text line alignment, and writer identification. The initial small subset of the HDD for character classification can be downloaded from https://www.cs.bgu.ac.illr-vberatldatalhhd_dataset.zip together with the training and test sets subdivisions. We also provide baseline results for character classification on this initial subset. In the near future, the full HHD dataset will be made freely available to the research community.
Irina Rabaev, Berat Kurar-Barakat, Alexander Churkin, Jihad El-Sana
ICFHR2
2020 Unsupervised deep learning for text line segmentation
abstract
We present an unsupervised deep learning method for text line segmentation that is inspired by the relative variance between text lines and spaces among text lines. Handwritten text line segmentation is important for the efficiency of further processing. A common method is to train a deep learning network for embedding the document image into an image of blob lines that are tracing the text lines. Previous methods learned such embedding in a supervised manner, requiring the annotation of many document images. This paper presents an unsupervised embedding of document image patches without a need for annotations. The number of foreground pixels over the text lines is relatively different from the number of foreground pixels over the spaces among text lines. Generating similar and different pairs relying on this principle definitely leads to outliers. However, as the results show, the outliers do not harm the convergence and the network learns to discriminate the text lines from the spaces between text lines. Remarkably, with a challenging Arabic handwritten text line segmentation dataset, VML-AHTE, we achieved superior performance over the supervised methods. Additionally, the proposed method was evaluated on the ICDAR 2017 and ICFHR 2010 handwritten text line segmentation datasets.
Berat Kurar-Barakat, Ahmad Droby, Reem Alaasam, Boraq Madi, Irina Rabaev, Raed Shammes, Jihad El-Sana
ICPR1
2019 Layout Analysis on Challenging Historical Arabic Manuscripts using Siamese Network
abstract
This paper presents layout analysis for historical Arabic documents using siamese network. Given pages from different documents, we divide them into patches of similar sizes. We train a siamese network model that takes as an input a pair of patches and gives as an output a distance that corresponds to the similarity between the two patches. We used the trained model to calculate a distance matrix which in turn is used to cluster the patches of a page as either main text, side text or a background patch. We evaluate our method on challenging historical Arabic manuscripts dataset and report the F-measure. We show the effectiveness of our method by comparing with other works that use deep learning approaches, and show that we have state of art results.
Reem Alaasam, Berat Kurar-Barakat, Jihad El-Sana
ICDAR2
2019 The Pinkas Dataset
abstract
In historical document image processing, datasets account for a significant part of any research, and are crucial for the diversity and abundance of experimental results, which contribute to the development of new algorithms to meet the new challenge. Moreover, they are very important for benchmarking processing algorithms. Numerous publicly available document image datasets of different languages have been emerged. However, current segmentation and recognition performances are nearly saturated with respect to the present publicly available datasets. As such, collecting and labelling historical document images is a burden on historical document image processing researchers. This paper introduces a public historical document image dataset, Pinkas dataset, with new challenges to open room for improvement and identify strengths and weaknesses of available processing algorithms. It is the first dataset in medieval handwritten Hebrew and fully labeled at word, line and page level by an expert of historical Hebrew manuscripts. Pinkas dataset contributes to the diversity of benchmarking standards. In this paper we present meta features of Pinkas dataset and apply recent word spotting algorithms to analyze the room for improvement in terms of performance.
Berat Kurar-Barakat, Jihad El-Sana, Irina Rabaev
ICDAR1
2018 Word Spotting Using Convolutional Siamese Network
abstract
We present a method for word spotting using convolutional siamese network. A convolutional siamese network employs two identical convolutional network to rank similarity between two input word images. Once the network is trained, it can then be used to spot not just words with varying writing styles and backgrounds but also to spot out of vocabulary words that are not in the training set. Experiments on the historical Arabic manuscript dataset VML, and on the George Washington dataset shows comparable results with the state of the art.
Berat Kurar-Barakat, Reem Alaasam, Jihad El-Sana
DAS1
2018 Text Line Segmentation for Challenging Handwritten Document Images using Fully Convolutional Network
abstract
This paper presents a method for text line segmentation of challenging historical manuscript images. These manuscript images contain narrow interline spaces with touching components, interpenetrating vowel signs and inconsistent font types and sizes. In addition, they contain curved, multi-skewed and multi-directed side note lines within a complex page layout. Therefore, bounding polygon labeling would be very difficult and time consuming. Instead we rely on line masks that connect the components on the same text line. Then these line masks are predicted using a Fully Convolutional Network (FCN). In the literature, FCN has been successfully used for text line segmentation of regular handwritten document images. The present paper shows that FCN is useful with challenging manuscript images as well. Using a new evaluation metric that is sensitive to over segmentation as well as under segmentation, testing results on a publicly available challenging handwritten dataset are comparable with the results of a previous work on the same dataset.
Berat Kurar-Barakat, Ahmad Droby, Majeed Kassis, Jihad El-Sana
ICFHR1