EDBT 2026 Demo / reviewers in the wild / expert
Raid Saabni
dblp:08/7443
· DBLP profile ↗
18ranked-venue papers
9as first author
5since 2021 · last 2025
0000-0001-8844-8675ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 2 since 2021Databases, data management, data science and information retrieval · 9 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Text Image Super-Resolution for Improved OCR in Real-Life Scenarios using Swin TransformersabstractText recognition in real-life images poses a challenging task due to blur, distortion, and low resolution. This work presents an innovative method integrating image super-resolution, image restoration, and optical character recognition techniques to enhance text recognition in real-life photographs. We specifically reviewed the processing of the TextZoom dataset and utilized transfer learning on an improved version of the image super-resolution model, SwinIR. The findings of our experiment show that our text recognition scores are better than the current best scores, and there is a significant rise in the peak signal-to-noise ratio while dealing with deformed low-resolution images from the TextZoom dataset. This approach outperforms earlier research in the domain of scene text image super-resolution and offers a promising resolution for text recognition in real-life images. The code can be accessed at this location: https://github.com/Phimanu/TextSR Philipp Hildebrandt, Maximilian Schulze, Sarel Cohen, Vanja Doskoc, Raid Saabni, Tobias Friedrich 0001 |
DocEng | 5 |
| 2024 | Detecting Spiral Text Lines in Aramaic Incantation Bowls
Said Naamneh, Boraq Madi, Nour Atamni, Shoshana Boardman, Daria Vasyutinsky Shapira, Irina Rabaev, Raid Saabni, Jihad El-Sana |
ICPR (19) | 7 |
| 2022 | Optical character recognition guided image super resolutionabstractRecognizing disturbed text in real-life images is a difficult problem, as information that is missing due to low resolution or out-of-focus text has to be recreated. Combining text super-resolution and optical character recognition deep learning models can be a valuable tool to enlarge and enhance text images for better readability, as well as recognize text automatically afterwards. We achieve improved peak signal-to-noise ratio and text recognition accuracy scores over a state-of-the-art text super-resolution model TBSRN on the real-world low-resolution dataset TextZoom while having a smaller theoretical model size due to the usage of quantization techniques. In addition, we show how different training strategies influence the performance of the resulting model. Philipp Hildebrandt, Maximilian Schulze, Sarel Cohen, Vanja Doskoc, Raid Saabni, Tobias Friedrich 0001 |
DocEng | 5 |
| 2021 | Text line extraction using deep learning and minimal sub seamsabstractAccurate text line extraction is a vital prerequisite for efficient and successful text recognition systems ranging from keywords/phrases searching to complete conversion to text. In many cases, the proposed algorithms target binary pre-processed versions of the image, which may cause insufficient results due to poor quality document images. Recently, more papers present solutions that work directly on gray-level images [1,2,7,12,15]. In this paper, we present a novel robust, and efficient algorithm to extract text-lines directly from gray-level document images. The proposed approach uses a combination of two variants of Convolutional Neural Network (CNNs), followed by minimal energy seam extraction. The first ConvNet is a modified version of the autoencoder used for biomedical image segmentation [8]. The second is a deep convolutional Neural Network, working on overlapping vertical slices of the original image. The two variants are combined to one neural net after re-attaching the resulting slices of the second net. The merged results of the two nets are used as a preprocessed image to obtain an energy map for a second phase. In the second step, we use the algorithm presented in [2], to track minimal energy sub-seams accumulated to perform a full local minimal/maximal separating and medial seam defining the text baselines and the text line regions. We have tested our approach on multi-lingual various datasets written at a range of image quality based on the ICDAR datasets. Adi Azran, Alon Schclar, Raid Saabni |
DocEng | 3 |
| 2021 | Unsupervised Learning of Text Line Segmentation by Differentiating Coarse Patterns
Berat Kurar-Barakat, Ahmad Droby, Raid Saabni, Jihad El-Sana |
ICDAR (2) | 3 |
| 2020 | A Manifold Learning Framework for the Detection of Cardiac Disorders in Acoustic Signals
Keren Hochman, Amir Averbuch, Alon Schclar, Raid Saabni |
ICPRAM | 4 |
| 2020 | A Diffusion Dimensionality Reduction Approach to Background Subtraction in Video Sequences
Dina Dushnik, Alon Schclar, Amir Averbuch, Raid Saabni |
IJCCI | 4 |
| 2014 | Real-Time Segmentation of On-Line Handwritten Arabic ScriptabstractReal-time performance is necessary in applications involving on-line handwriting recognition. However, conventional approaches usually wait until the entire curve is traced out before starting the analysis, inevitably causing delays in the recognition process. In regards to the Arabic script, the postponed analysis may be attributed to the cursive and unconstrained nature of the Arabic writing system, in both printed and handwritten forms. Nevertheless, this paper proposes a real-time recognition-based segmentation technique of on-line Arabic script. It demonstrate the feasibility of carrying out the most time consuming tasks, required for the segmentation process, during the course of writing. The system has been designed and tested using the ADAB Database, and promising results were obtained. George Kour, Raid Saabni |
ICFHR | 2 |
| 2014 | Text line extraction for historical document images
Raid Saabni, Abedelkadir Asi, Jihad El-Sana |
Pattern Recognit. Lett. | 1 |
| 2013 | Efficient Word Image Retrieval Using Earth Movers Distance Embedded to Wavelets Coefficients DomainabstractIn this paper we use the Earth Movers Distance (EMD) algorithm to measure similarity between shapes for recognizing and searching Arabic words. We have used the Shape Context and the Angular Radial Partitioning descriptors to evaluate matching and recognizing with EMD. Based on the encouraging results of high accuracy and recall, we follow the low-distortion embedding of the Earth Mover's Distance to map the shapes in the database under the EMD distance, into a normed space of wavelet coefficients as differences of coefficients histograms. The approximate k-nearest neighbors in the database of the embedded shapes are retrieved in sub linear time using a Locality-Sensitive Hashing (LSH) and generate a short list of candidates. This short list of candidates is used in a filter and refine strategy and the exact results are achieved using the original EMD on this short list. We demonstrate our method on the MNIST dataset and the freely available Arabic Printed Text Image (APTI) database. Our method achieves a speedup of 4 orders of magnitude over the exact method, at the cost of only a 2.4% reduction in accuracy. Raid Saabni |
ICDAR | 1 |
| 2013 | Comprehensive synthetic Arabic database for on/off-line script recognition research
Raid Saabni, Jihad El-Sana |
Int. J. Document Anal. Recognit. | 1 |
| 2012 | Fast Keyword Searching Using 'BoostMap' Based EmbeddingabstractDynamic Time Warping (DTW), is a simple but efficient technique for matching sequences with rigid deformation. Therefore, it is frequently used for matching shapes in general, and shapes of handwritten words in Document Image Analysis tasks. As DTW is computationally expensive, efficient algorithms for fast computation are crucial. Retrieving images from large scale datasets using DTW, suffers from the constraint of linear searching of all sample in the datasets. Fast approximation algorithms for image retrieval are mostly based on normed spaces where the triangle inequality holds, which is unfortunately not the case with the DTW metric. In this paper we present a novel approach for fast search of handwritten words within large datasets of shapes. The presented approach is based on the Boost-Map [1] algorithm, for embedding the feature space with the DTW measurement to an euclidean space and use the Local Sensitivity Hashing algorithm (LSH) to rank the k-nearest neighbors of a query image. The algorithm, first, processes and embeds objects of the large data sets to a normed space. Fast approximation of k-nearest neighbors using LSH on the embedding space, generates the top kranked samples which are examined using the real DTW distance to give final accurate results. We demonstrate our method on a database of 45; 800 images of word-parts extracted from the IFN/ENIT database [11] and images collected from 51 different writers. Our method achieves a speedup of 4 orders of magnitude over the exact method, at the cost of only a 2:2% reduction in accuracy. Raid Saabni, Alexander M. Bronstein |
ICFHR | 1 |
| 2012 | Text Detection and Recognition in Real World ImagesabstractDetecting and recognizing texts in real world images such as sign boards and advertisements is an important part of computer vision applications. The complexity of the problem comes out of many factors such as nonuniform background, different languages and fonts, and non consistent text alignment and orientation. In this paper, we present a novel approach to detect characters and words in real-world images. The presented approach decompose the gray level image into sequence of images, each one includes pixels with gray level values from different disjoint ranges. This decomposition enables extracting connected components representing characters or other non textual objects separated from their neighborhood background. An interpolation of two classes of features translated to histograms is used by a support vector machine to classify and collect the textual objects generating the textual zones. The Shape Context Descriptor [1], is used by the Earth Movers Distance(EMD) method to recognize the characters within the image. The recognized characters are fed to heuristic rule based system to determine words and give final results. To optimize the speed of the system, we follow the embedding of the EMD metric presented in [22] to a normed space to enable fast approximation of the k-Nearest Neighbors using Local Sensitivity Hashing functions(LSH). Experiments show that our algorithm can detect and recognize text regions from the ICDAR 2005 datasets [17] with high rates. Raid Saabni, Moti Zwilling |
ICFHR | 1 |
| 2011 | Fast Key-Word Searching via Embedding and Active-DTWabstractIn this paper we present a novel approach for fast search of handwritten Arabic word-parts within large lexicons. The algorithm runs through three steps to achieve the required results. First it warps multiple appearances of each word-part in the lexicon for embedding into the same euclidean space. The embedding is done based on the warping path produced by the Dynamic Time Warping (DTW) process while calculating the similarity distance. In the next step, all samples of different word-parts are resampled uniformly to the same size. The kd-tree structure is used to store all shapes representing word parts in the lexicon. Fast approximation of k-nearest neighbors generates a short list of candidates to be presented to the next step. In the third step, the Active-DTW [15] algorithm is used to examine each sample in the short list and give final accurate results. We demonstrate our method on a database of 23,500 images of word-parts extracted from the IFN/ENIT database [6] and 22,000 images collected from 93 writers. Our method achieves a speedup of 5 orders of magnitude over the exact method, at the cost of only a 3.8% reduction in accuracy. Raid Saabni, Alexander M. Bronstein |
ICDAR | 1 |
| 2011 | Language-Independent Text Lines Extraction Using Seam CarvingabstractIn this paper, we present a novel language-independent algorithm for extracting text-lines from handwritten document images. Our algorithm is based on the seam carving approach for content aware image resizing. We adopted the signed distance transform to generate the energy map, where extreme points indicate the layout of text-lines. Dynamic programming is then used to compute the minimum energy left-to-right paths (seams), which pass along the ``middle`` of the text-lines. Each path intersects a set of components, which determine the extracted text-line and estimate its hight. The estimated hight determines the text-line's region, which guides splitting touching components among consecutive lines. Unassigned components that fall within the region of a text-line are added to the components list of the line. The components between two consecutive lines are processed when the two lines are extracted and assigned to the closest text-line, based on the attributes of extracted lines, the sizes and positions of components. Our experimental results on Arabic, Chinese, and English historical documents show that our approach manage to separate multi-skew text blocks into lines at high success rates. Raid Saabni, Jihad El-Sana |
ICDAR | 1 |
| 2011 | Segmentation-Free Online Arabic Handwriting RecognitionabstractArabic script is naturally cursive and unconstrained and, as a result, an automatic recognition of its handwriting is a challenging problem. The analysis of Arabic script is further complicated in comparison to Latin script due to obligatory dots/stokes that are placed above or below most letters. In this paper, we introduce a new approach that performs online Arabic word recognition on a continuous word-part level, while performing training on the letter level. In addition, we appropriately handle delayed strokes by first detecting them and then integrating them into the word-part body. Our current implementation is based on Hidden Markov Models (HMM) and correctly handles most of the Arabic script recognition difficulties. We have tested our implementation using various dictionaries and multiple writers and have achieved encouraging results for both writer-dependent and writer-independent recognition. Fadi Biadsy, Raid Saabni, Jihad El-Sana |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2009 | Hierarchical On-line Arabic Handwriting RecognitionabstractIn this paper, we present a multi-level recognizer for online Arabic handwriting. In Arabic script (handwritten and printed), cursive writing - is not a style - it is an inherent part of the script. In addition, the connection between letters is done with almost no ligatures, which complicates segmenting a word into individual letters. In this work, we have adopted the holistic approach and avoided segmenting words into individual letters. To reduce the search space, we apply a series of filters in a hierarchical manner. The earlier filters perform light processing on a large number of candidates, and the later filters perform heavy processing on a small number of candidates. In the first filter, global features and delayed strokes patterns are used to reduce candidate word-part models. In the second filter, local features are used to guide a dynamic time warping (DTW) classification. The resulting k top ranked candidates are sent for shape context based classifier, which determines the recognized word-part. In this work, we have modified the classic DTW to enable different costs for the different operations and control their behavior. We have performed several experimental tests and have received encouraging results. Raid Saabni, Jihad El-Sana |
ICDAR | 1 |
| 2009 | Efficient Generation of Comprehensive Database for Online Arabic Script RecognitionabstractThe difficulties in segmenting cursive words into individual characters have shifted the focus of handwriting recognition research from segmentation-based approaches to segmentation-free (holistic) methods. However, maintaining and training large number of prototypes (models) that represent the words in the dictionary make the training process extremely expensive and difficult in computing resources. In this paper we present an efficient system that automatically generates prototypes for each word in a given dictionary using multiple appearance of each letter shape. Multiple appearance allows for many permutation of shapes for each word and thus complicates searching for the right prototype. To simplify the training, reduce the maintained prototypes, and avoid over fitting, we used dimensionality reduction followed by clustering techniques to reduce the size of these sets without affecting their ability to represent the wide variations of the handwriting styles. A set of generated fonts are created by professional writers imitating all handwriting styles for each character in each position. These fonts are used to generate all shapes for writing each word-part in a comprehensive dictionary. Principal component analysis and k-means clustering techniques are performed to select the minimal number of shapes representing the wide variations of handwriting styles for a word-part. Experimental results using an online recognition system proves the credibility of this process compared to manually generated databases. Raid Saabni, Jihad El-Sana |
ICDAR | 1 |