Irina Rabaev

dblp:25/10400 · DBLP profile ↗
← Back
10ranked-venue papers in the field
4as first author
5since 2021 · last 2025
0000-0002-8542-8342ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 9 (4 first)Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 Multi-task Learning for Hebrew Paleography: Script Classification and Date Estimation
Nour Atamni, Boraq Madi, Shoshana Bordman, Daria Vasyutinsky Shapira, Irina Rabaev, Jihad El-Sana
ICDAR (5)5
2025 ICDAR 2025 Competition on Automatic Classification of Literary Epochs
Irina Rabaev, Marina Litvak, Roza Bass, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt
ICDAR (5)1
2023 The 1st International Workshop on Implicit Author Characterization from Texts for Search and Retrieval (IACT'23)
abstract
The first edition of the Implicit Author Characterization from Texts for Search and Retrieval (IACT'23) aims at bringing to the forefront the challenges involved in identifying and extracting from texts implicit information about authors (e.g., human or AI) and using it in IR tasks. The IACT workshop provides a common forum to consolidate multi-disciplinary efforts and foster discussions to identify the wide-ranging issues related to the task of extracting implicit author-related information from the textual content, including novel tasks and datasets. We will also discuss the ethical implications of implicit information extraction. In addition, we announce a shared task focused on automatically determining the literary epochs of written books.
Marina Litvak, Irina Rabaev, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt
SIGIR2
2022 Hard and Soft Labeling for Hebrew Paleography: A Case Study
Ahmad Droby, Daria Vasyutinsky Shapira, Irina Rabaev, Berat Kurar-Barakat, Jihad El-Sana
DAS3
2021 VML-HP: Hebrew Paleography Dataset
Ahmad Droby, Berat Kurar-Barakat, Daria Vasyutinsky Shapira, Irina Rabaev, Jihad El-Sana
ICDAR (4)4
2019 The Pinkas Dataset
abstract
In historical document image processing, datasets account for a significant part of any research, and are crucial for the diversity and abundance of experimental results, which contribute to the development of new algorithms to meet the new challenge. Moreover, they are very important for benchmarking processing algorithms. Numerous publicly available document image datasets of different languages have been emerged. However, current segmentation and recognition performances are nearly saturated with respect to the present publicly available datasets. As such, collecting and labelling historical document images is a burden on historical document image processing researchers. This paper introduces a public historical document image dataset, Pinkas dataset, with new challenges to open room for improvement and identify strengths and weaknesses of available processing algorithms. It is the first dataset in medieval handwritten Hebrew and fully labeled at word, line and page level by an expert of historical Hebrew manuscripts. Pinkas dataset contributes to the diversity of benchmarking standards. In this paper we present meta features of Pinkas dataset and apply recent word spotting algorithms to analyze the room for improvement in terms of performance.
Berat Kurar-Barakat, Jihad El-Sana, Irina Rabaev
ICDAR3
2016 Keyword Retrieval Using Scale-Space Pyramid
abstract
We propose a pyramid-based method for keyword spotting in historical document images. The documents are represented by a scale-space pyramid of their features. The search for a query keyword begins at the highest level of the pyramid, where the initial candidates for matching are located. The candidates are further refined at each level of the pyramid. The number of levels is adaptive and depends on the length of the query word. The results from all the document images are combined and ranked. We compare two feature representations, grid-based and continuous, and show that continuous feature representation outperforms the grid-based representation. In order to reduce the memory used to store the scale-space pyramid of features, we discuss and compare two compressing approaches. The proposed method was evaluated on four different collections of historical documents achieving state-of-the-art results.
Irina Rabaev, Klara Kedem, Jihad El-Sana
DAS1
2015 Aligning transcript of historical documents using energy minimization
abstract
An ongoing considerable effort for digitizing historical manuscripts has produced images of original manuscripts, some accompanied by transcripts. Aligning the text in the input image with the text in the transcript will allow learning, training and evaluating recognition algorithms. Here we propose a system that computes the alignment by formulating the problem as an energy minimization task, where the alignment is performed between the input line image to a synthetic one. The energy function works at a connected component level and it combines a visual similarity measure and a learned distance metric that separates between inter-word and intra-word connected components.
Rafi Cohen, Irina Rabaev, Jihad El-Sana, Klara Kedem, Its'hak Dinstein
ICDAR2
2013 Text Line Detection in Corrupted and Damaged Historical Manuscripts
abstract
Most of the algorithms proposed for text line detection are designed to process binary images as input. For severely degraded documents, binarization often introduces significant noise and other artifacts. In this work we present a novel method designed to detect text lines directly in gray scale images. The method consists of two stages. Potential characters are detected in the first stage. This is done by analyzing the evolution maps of connected components obtained by a sliding threshold. The detected potential characters are grouped into text lines in the second stage using sweep-line approach. The suggested method is especially powerful when applied to torn and damaged documents that other algorithms are not able to deal with.
Irina Rabaev, Ofer Biller, Jihad El-Sana, Klara Kedem, Its'hak Dinstein
ICDAR1
2011 Case Study in Hebrew Character Searching
abstract
Searching for a letter or a word in historical documents is a practical challenge due to the various degradations present in such documents and the wide variance of handwriting. Searching in historical Hebrew documents is somewhat harder because of high similarities among Hebrew characters. In order to determine the features and their combinations appropriate for recognizing Hebrew script, we study a range of known features using a Dynamic Time Warping algorithm. In addition we describe a novel meth od for feature-based searching, which uses a number of models for the same character. This method is based on our original DTW algorithm that can match fragments of several models of the same character to match a query character. Consequently, we are not limited to any particular model of the character set. Application of this method leads to a significant improvement, even when using a small set of models.
Irina Rabaev, Ofer Biller, Jihad El-Sana, Klara Kedem, Its'hak Dinstein
ICDAR1