Irina Rabaev

dblp:25/10400 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-8542-8342ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Recent advances in text line segmentation and baseline detection in historical document images: a systematic review
abstract
Abstract The purpose of this survey is to provide a comprehensive overview of recent advancements in text line segmentation and baseline detection techniques within the analysis of historical document images. Text line extraction is an essential step in the historical documents image analysis pipeline, as its results significantly impact the accuracy of subsequent processes, such as handwritten text recognition (HTR). Through a multi-stage procedure, we carefully selected 49 peer-reviewed studies published since 2019. Based on careful analysis of these studies, we summarize the information of the existing datasets, describe and categorize different methods, and summarize evaluation protocols. In addition, we compare the results of various methods on benchmark datasets. Finally, we highlight the gaps and suggest directions for future research. We believe that this comprehensive survey will be of great assistance to researchers working in the field of historical document image analysis, as it offers critical insights into the latest advancements and developments, providing a foundation for future research.
Irina Rabaev, Marina Litvak
Int. J. Document Anal. Recognit.1
2025 Multi-task Learning for Hebrew Paleography: Script Classification and Date Estimation
Nour Atamni, Boraq Madi, Shoshana Bordman, Daria Vasyutinsky Shapira, Irina Rabaev, Jihad El-Sana
ICDAR (5)5
2025 ICDAR 2025 Competition on Automatic Classification of Literary Epochs
Irina Rabaev, Marina Litvak, Roza Bass, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt
ICDAR (5)1
2024 Detecting Spiral Text Lines in Aramaic Incantation Bowls
Said Naamneh, Boraq Madi, Nour Atamni, Shoshana Boardman, Daria Vasyutinsky Shapira, Irina Rabaev, Raid Saabni, Jihad El-Sana
ICPR (19)6
2023 The 1st International Workshop on Implicit Author Characterization from Texts for Search and Retrieval (IACT'23)
abstract
The first edition of the Implicit Author Characterization from Texts for Search and Retrieval (IACT'23) aims at bringing to the forefront the challenges involved in identifying and extracting from texts implicit information about authors (e.g., human or AI) and using it in IR tasks. The IACT workshop provides a common forum to consolidate multi-disciplinary efforts and foster discussions to identify the wide-ranging issues related to the task of extracting implicit author-related information from the textual content, including novel tasks and datasets. We will also discuss the ethical implications of implicit information extraction. In addition, we announce a shared task focused on automatically determining the literary epochs of written books.
Marina Litvak, Irina Rabaev, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt
SIGIR2
2023 Automated gender classification from handwriting: a systematic survey
Irina Rabaev, Marina Litvak
Appl. Intell.1
2022 Hard and Soft Labeling for Hebrew Paleography: A Case Study
Ahmad Droby, Daria Vasyutinsky Shapira, Irina Rabaev, Berat Kurar-Barakat, Jihad El-Sana
DAS3
2021 Automatic Gender Classification from Handwritten Images: A Case Study
Irina Rabaev, Marina Litvak, Sean Asulin, Or Haim Tabibi
CAIP (2)1
2021 VML-HP: Hebrew Paleography Dataset
Ahmad Droby, Berat Kurar-Barakat, Daria Vasyutinsky Shapira, Irina Rabaev, Jihad El-Sana
ICDAR (4)4
2021 Urban Planter: A Web App for Automatic Classification of Urban Plants
abstract
Plant classification requires an expert because subtle differences in leaves or petal forms might differentiate between different species. On the contrary, some species are characterized by high variability in appearance. This paper introduces a web app for assisting people in identifying plants for discovering the best growing methods. The uploaded picture is submitted to the back-end server, and a pre-trained neural network classifies it to one of the predefined classes. The classification label and confidence are displayed to the end user on the front-end page. The application focuses on the house and garden plant species that can be grown mainly in a desert climate and are not covered by existing datasets. For training a model, we collected the Urban Planter dataset. The installation code of the alpha version and the demo video of the app can be found on https://github.com/UrbanPlanter/urbanplanterapp.
Sarit Divekar, Irina Rabaev, Marina Litvak
VCIP2
2021 Telemoji: A video chat with automated recognition of facial expressions
abstract
Autism spectrum disorder (ASD) is frequently ac-companied by impairment in emotional expression recognition, and therefore individuals with ASD may find it hard to interpret emotions and interact. Inspired by this fact, we developed a web-based video chat to assist people with ASD, both for real-time recognition of facial emotions and for practicing. This real-time application detects the speaker's face in a video stream and classifies the expressed emotion into one of the seven categories: neutral, surprise, happy, angry, disgust, fear, and sad. The classification is then displayed as the text label below the speaker's face. We developed this application as a part of the undergraduate project for the B.Sc. degree in Software Engineering. Its development and testing were made with the cooperation of the local society for children and adults with autism. The application has been released for unrestricted use on https://telemojii.herokuapp.com/. The demo is available at http://www.filedropper.com/telemojishortdemoblur.
Alex Kreinis, Tom Damri, Tomer Leon, Marina Litvak, Irina Rabaev
VCIP5
2020 The HHD Dataset
abstract
Benchmark datasets are important in document image processing field, as they allow to analyze different approaches and compare their performances in a fair manner. There exist benchmark datasets for several alphabets such as Latin, Arabic and Chinese, but not the Hebrew alphabet. In this paper, a handwritten Hebrew dataset, HHD, is introduced. The HHD dataset is collected from hand-filled forms, and accompanied by their ground truth at character, word and text line levels. Presently, the dataset contains around 1000 document images, and we continue to further enlarge it. To the best of our knowledge, this is the first comprehensive corpus of Hebrew handwritten documents, and we believe it will help leveraging Hebrew documents processing and document processing in general. The dataset can be useful for various research applications, such as word spotting, word recognition, text line alignment, and writer identification. The initial small subset of the HDD for character classification can be downloaded from https://www.cs.bgu.ac.illr-vberatldatalhhd_dataset.zip together with the training and test sets subdivisions. We also provide baseline results for character classification on this initial subset. In the near future, the full HHD dataset will be made freely available to the research community.
Irina Rabaev, Berat Kurar-Barakat, Alexander Churkin, Jihad El-Sana
ICFHR1
2020 Unsupervised deep learning for text line segmentation
abstract
We present an unsupervised deep learning method for text line segmentation that is inspired by the relative variance between text lines and spaces among text lines. Handwritten text line segmentation is important for the efficiency of further processing. A common method is to train a deep learning network for embedding the document image into an image of blob lines that are tracing the text lines. Previous methods learned such embedding in a supervised manner, requiring the annotation of many document images. This paper presents an unsupervised embedding of document image patches without a need for annotations. The number of foreground pixels over the text lines is relatively different from the number of foreground pixels over the spaces among text lines. Generating similar and different pairs relying on this principle definitely leads to outliers. However, as the results show, the outliers do not harm the convergence and the network learns to discriminate the text lines from the spaces between text lines. Remarkably, with a challenging Arabic handwritten text line segmentation dataset, VML-AHTE, we achieved superior performance over the supervised methods. Additionally, the proposed method was evaluated on the ICDAR 2017 and ICFHR 2010 handwritten text line segmentation datasets.
Berat Kurar-Barakat, Ahmad Droby, Reem Alaasam, Boraq Madi, Irina Rabaev, Raed Shammes, Jihad El-Sana
ICPR5
2019 The Pinkas Dataset
abstract
In historical document image processing, datasets account for a significant part of any research, and are crucial for the diversity and abundance of experimental results, which contribute to the development of new algorithms to meet the new challenge. Moreover, they are very important for benchmarking processing algorithms. Numerous publicly available document image datasets of different languages have been emerged. However, current segmentation and recognition performances are nearly saturated with respect to the present publicly available datasets. As such, collecting and labelling historical document images is a burden on historical document image processing researchers. This paper introduces a public historical document image dataset, Pinkas dataset, with new challenges to open room for improvement and identify strengths and weaknesses of available processing algorithms. It is the first dataset in medieval handwritten Hebrew and fully labeled at word, line and page level by an expert of historical Hebrew manuscripts. Pinkas dataset contributes to the diversity of benchmarking standards. In this paper we present meta features of Pinkas dataset and apply recent word spotting algorithms to analyze the room for improvement in terms of performance.
Berat Kurar-Barakat, Jihad El-Sana, Irina Rabaev
ICDAR3
2016 Keyword Retrieval Using Scale-Space Pyramid
abstract
We propose a pyramid-based method for keyword spotting in historical document images. The documents are represented by a scale-space pyramid of their features. The search for a query keyword begins at the highest level of the pyramid, where the initial candidates for matching are located. The candidates are further refined at each level of the pyramid. The number of levels is adaptive and depends on the length of the query word. The results from all the document images are combined and ranked. We compare two feature representations, grid-based and continuous, and show that continuous feature representation outperforms the grid-based representation. In order to reduce the memory used to store the scale-space pyramid of features, we discuss and compare two compressing approaches. The proposed method was evaluated on four different collections of historical documents achieving state-of-the-art results.
Irina Rabaev, Klara Kedem, Jihad El-Sana
DAS1
2015 Aligning transcript of historical documents using energy minimization
abstract
An ongoing considerable effort for digitizing historical manuscripts has produced images of original manuscripts, some accompanied by transcripts. Aligning the text in the input image with the text in the transcript will allow learning, training and evaluating recognition algorithms. Here we propose a system that computes the alignment by formulating the problem as an energy minimization task, where the alignment is performed between the input line image to a synthetic one. The energy function works at a connected component level and it combines a visual similarity measure and a learned distance metric that separates between inter-word and intra-word connected components.
Rafi Cohen, Irina Rabaev, Jihad El-Sana, Klara Kedem, Its'hak Dinstein
ICDAR2
2013 Text Line Detection in Corrupted and Damaged Historical Manuscripts
abstract
Most of the algorithms proposed for text line detection are designed to process binary images as input. For severely degraded documents, binarization often introduces significant noise and other artifacts. In this work we present a novel method designed to detect text lines directly in gray scale images. The method consists of two stages. Potential characters are detected in the first stage. This is done by analyzing the evolution maps of connected components obtained by a sliding threshold. The detected potential characters are grouped into text lines in the second stage using sweep-line approach. The suggested method is especially powerful when applied to torn and damaged documents that other algorithms are not able to deal with.
Irina Rabaev, Ofer Biller, Jihad El-Sana, Klara Kedem, Its'hak Dinstein
ICDAR1
2011 Case Study in Hebrew Character Searching
abstract
Searching for a letter or a word in historical documents is a practical challenge due to the various degradations present in such documents and the wide variance of handwriting. Searching in historical Hebrew documents is somewhat harder because of high similarities among Hebrew characters. In order to determine the features and their combinations appropriate for recognizing Hebrew script, we study a range of known features using a Dynamic Time Warping algorithm. In addition we describe a novel meth od for feature-based searching, which uses a number of models for the same character. This method is based on our original DTW algorithm that can match fragments of several models of the same character to match a query character. Consequently, we are not limited to any particular model of the character set. Application of this method leads to a significant improvement, even when using a small set of models.
Irina Rabaev, Ofer Biller, Jihad El-Sana, Klara Kedem, Its'hak Dinstein
ICDAR1