Jihad El-Sana

dblp:88/2193 · DBLP profile ↗
← Back
74ranked-venue papers
12as first author
10since 2021 · last 2025
0000-0002-1164-7040ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 26 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 22 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 13 · 6 first-authorTheory of computation · 1
YearPublicationVenuePosition
2025 Multi-task Learning for Hebrew Paleography: Script Classification and Date Estimation
Nour Atamni, Boraq Madi, Shoshana Bordman, Daria Vasyutinsky Shapira, Irina Rabaev, Jihad El-Sana
ICDAR (5)6
2024 Text Enhancement for Historical Handwritten Documents
Reem Alaasam, Boraq Madi, Jihad El-Sana
ICDAR (2)3
2024 Detecting Spiral Text Lines in Aramaic Incantation Bowls
Said Naamneh, Boraq Madi, Nour Atamni, Shoshana Boardman, Daria Vasyutinsky Shapira, Irina Rabaev, Raid Saabni, Jihad El-Sana
ICPR (19)8
2023 Scheme for palimpsests reconstruction using synthesized dataset
Boraq Madi, Reem Alaasam, Raed Shammas, Jihad El-Sana
Int. J. Document Anal. Recognit.4
2022 Hard and Soft Labeling for Hebrew Paleography: A Case Study
Ahmad Droby, Daria Vasyutinsky Shapira, Irina Rabaev, Berat Kurar-Barakat, Jihad El-Sana
DAS5
2022 HST-GAN: Historical Style Transfer GAN for Generating Historical Text Images
Boraq Madi, Reem Alaasam, Ahmad Droby, Jihad El-Sana
DAS4
2022 Text Edges Guided Network for Historical Document Super Resolution
Boraq Madi, Reem Alaasam, Jihad El-Sana
ICFHR3
2022 Textline alignment on the image domain
Boraq Madi, Ahmad Droby, Jihad El-Sana
Int. J. Document Anal. Recognit.3
2021 Unsupervised Learning of Text Line Segmentation by Differentiating Coarse Patterns
Berat Kurar-Barakat, Ahmad Droby, Raid Saabni, Jihad El-Sana
ICDAR (2)4
2021 VML-HP: Hebrew Paleography Dataset
Ahmad Droby, Berat Kurar-Barakat, Daria Vasyutinsky Shapira, Irina Rabaev, Jihad El-Sana
ICDAR (4)5
2020 Unsupervised Deep Learning for Handwritten Page Segmentation
abstract
Segmenting handwritten document images into regions with homogeneous patterns is an important pre-processing step for many document images analysis tasks. Hand-labeling data to train a deep learning model for layout analysis requires significant human effort. In this paper, we present an unsupervised deep learning method for page segmentation, which revokes the need for annotated images. A siamese neural network is trained to differentiate between patches using their measurable properties such as number of foreground pixels, and average component height and width. The network is trained that spatially nearby patches are similar. The network's learned features are used for page segmentation, where patches are classified as main and side text based on the extracted features. We tested the method on a dataset of handwritten document images with quite complex layouts. Our experiments show that the proposed unsupervised method is as effective as typical supervised methods.
Ahmad Droby, Berat Kurar-Barakat, Boraq Madi, Reem Alaasam, Jihad El-Sana
ICFHR5
2020 The HHD Dataset
abstract
Benchmark datasets are important in document image processing field, as they allow to analyze different approaches and compare their performances in a fair manner. There exist benchmark datasets for several alphabets such as Latin, Arabic and Chinese, but not the Hebrew alphabet. In this paper, a handwritten Hebrew dataset, HHD, is introduced. The HHD dataset is collected from hand-filled forms, and accompanied by their ground truth at character, word and text line levels. Presently, the dataset contains around 1000 document images, and we continue to further enlarge it. To the best of our knowledge, this is the first comprehensive corpus of Hebrew handwritten documents, and we believe it will help leveraging Hebrew documents processing and document processing in general. The dataset can be useful for various research applications, such as word spotting, word recognition, text line alignment, and writer identification. The initial small subset of the HDD for character classification can be downloaded from https://www.cs.bgu.ac.illr-vberatldatalhhd_dataset.zip together with the training and test sets subdivisions. We also provide baseline results for character classification on this initial subset. In the near future, the full HHD dataset will be made freely available to the research community.
Irina Rabaev, Berat Kurar-Barakat, Alexander Churkin, Jihad El-Sana
ICFHR4
2020 Unsupervised deep learning for text line segmentation
abstract
We present an unsupervised deep learning method for text line segmentation that is inspired by the relative variance between text lines and spaces among text lines. Handwritten text line segmentation is important for the efficiency of further processing. A common method is to train a deep learning network for embedding the document image into an image of blob lines that are tracing the text lines. Previous methods learned such embedding in a supervised manner, requiring the annotation of many document images. This paper presents an unsupervised embedding of document image patches without a need for annotations. The number of foreground pixels over the text lines is relatively different from the number of foreground pixels over the spaces among text lines. Generating similar and different pairs relying on this principle definitely leads to outliers. However, as the results show, the outliers do not harm the convergence and the network learns to discriminate the text lines from the spaces between text lines. Remarkably, with a challenging Arabic handwritten text line segmentation dataset, VML-AHTE, we achieved superior performance over the supervised methods. Additionally, the proposed method was evaluated on the ICDAR 2017 and ICFHR 2010 handwritten text line segmentation datasets.
Berat Kurar-Barakat, Ahmad Droby, Reem Alaasam, Boraq Madi, Irina Rabaev, Raed Shammes, Jihad El-Sana
ICPR7
2019 Layout Analysis on Challenging Historical Arabic Manuscripts using Siamese Network
abstract
This paper presents layout analysis for historical Arabic documents using siamese network. Given pages from different documents, we divide them into patches of similar sizes. We train a siamese network model that takes as an input a pair of patches and gives as an output a distance that corresponds to the similarity between the two patches. We used the trained model to calculate a distance matrix which in turn is used to cluster the patches of a page as either main text, side text or a background patch. We evaluate our method on challenging historical Arabic manuscripts dataset and report the F-measure. We show the effectiveness of our method by comparing with other works that use deep learning approaches, and show that we have state of art results.
Reem Alaasam, Berat Kurar-Barakat, Jihad El-Sana
ICDAR3
2019 The Pinkas Dataset
abstract
In historical document image processing, datasets account for a significant part of any research, and are crucial for the diversity and abundance of experimental results, which contribute to the development of new algorithms to meet the new challenge. Moreover, they are very important for benchmarking processing algorithms. Numerous publicly available document image datasets of different languages have been emerged. However, current segmentation and recognition performances are nearly saturated with respect to the present publicly available datasets. As such, collecting and labelling historical document images is a burden on historical document image processing researchers. This paper introduces a public historical document image dataset, Pinkas dataset, with new challenges to open room for improvement and identify strengths and weaknesses of available processing algorithms. It is the first dataset in medieval handwritten Hebrew and fully labeled at word, line and page level by an expert of historical Hebrew manuscripts. Pinkas dataset contributes to the diversity of benchmarking standards. In this paper we present meta features of Pinkas dataset and apply recent word spotting algorithms to analyze the room for improvement in terms of performance.
Berat Kurar-Barakat, Jihad El-Sana, Irina Rabaev
ICDAR2
2019 Learning Free Line Detection in Manuscripts using Distance Transform Graph
abstract
We present a fully automated learning free method, for line detection in manuscripts. We begin by separating components that span over multiple lines, then we remove noise, and small connected components such as diacritics. We apply a distance transform on the image to create the image skeleton. The skeleton is pruned, its vertexes and edges are detected, in order to generate the initial document graph. We calculate the vertex v-score using its t-score and l-score quantifying its distance from being an absolute link in a line. In a greedy manner we classify each edge in the graph either a link, a bridge or a conflict edge. We merge every two edges classified as link together, then we merge the conflict edges next. Finally we remove the bridge edges from the graph generating the final form of the graph. Each edge in the graph equals to one extracted line. We applied the method on the DIVA-hisDB dataset on both public and private sections. The public section participated in the recently conducted Layout Analysis for Challenging Medieval Manuscripts Competition, and we have achieved results surpassing the vast majority of these systems.
Majeed Kassis, Jihad El-Sana
ICDAR2
2018 Word Spotting Using Convolutional Siamese Network
abstract
We present a method for word spotting using convolutional siamese network. A convolutional siamese network employs two identical convolutional network to rank similarity between two input word images. Once the network is trained, it can then be used to spot not just words with varying writing styles and backgrounds but also to spot out of vocabulary words that are not in the training set. Experiments on the historical Arabic manuscript dataset VML, and on the George Washington dataset shows comparable results with the state of the art.
Berat Kurar-Barakat, Reem Alaasam, Jihad El-Sana
DAS3
2018 Text Line Segmentation for Challenging Handwritten Document Images using Fully Convolutional Network
abstract
This paper presents a method for text line segmentation of challenging historical manuscript images. These manuscript images contain narrow interline spaces with touching components, interpenetrating vowel signs and inconsistent font types and sizes. In addition, they contain curved, multi-skewed and multi-directed side note lines within a complex page layout. Therefore, bounding polygon labeling would be very difficult and time consuming. Instead we rely on line masks that connect the components on the same text line. Then these line masks are predicted using a Fully Convolutional Network (FCN). In the literature, FCN has been successfully used for text line segmentation of regular handwritten document images. The present paper shows that FCN is useful with challenging manuscript images as well. Using a new evaluation metric that is sensitive to over segmentation as well as under segmentation, testing results on a publicly available challenging handwritten dataset are comparable with the results of a previous work on the same dataset.
Berat Kurar-Barakat, Ahmad Droby, Majeed Kassis, Jihad El-Sana
ICFHR4
2017 Alignment of Historical Handwritten Manuscripts Using Siamese Neural Network
abstract
Historical manuscript alignment is a widely known problem in historical document analysis, and the attempt of finding the differences between manuscript editions is mainly done by hand. Today, most of the computational tools coming to assist the historians are based on word recognition or spotting. These solutions are partial at best. In this paper, we present a Siamese neural network based system, which automatically identifies whether a pair of images contain the same text without the need of recognizing the text. The user is required to annotate several pages of two manuscripts, and with the assistance of synthetically generated data and affine distortions we can align two manuscripts written by different writers, achieving strong results.
Majeed Kassis, Jumana Nassour, Jihad El-Sana
ICDAR3
2017 On writer identification for Arabic historical manuscripts
Abedelkadir Asi, Alaa Abdalhaleem, Daniel Fecker, Volker Märgner, Jihad El-Sana
Int. J. Document Anal. Recognit.5
2016 Automatic Synthesis of Historical Arabic Text for Word-Spotting
abstract
We present a novel framework for automatic and efficient synthesis of historical handwritten Arabic text. The main purpose of this framework is to assist word spotting and keyword searching in handwritten historical documents. The proposed framework consists of two main procedures: building a letter connectivity map and synthesizing words. A letter connectivity map includes multiple instances of the various shape of each letter, since a letter in Arabic usually has multiple shapes depends in its position in the word. Each map represents one writer and encodes the specific handwriting style. The letter connectivity map is used to guide the synthesis of any Arabic continuous subword, word, or sentence. The proposed framework automatically generates the letter connectivity map annotation from a several pages historical pages previously annotated. Once the letter connectivity map is available our framework can synthesis the pictorial representation of any Arabic word or sentence from their text representation. The writing style of the synthesized text resembles the writing style of the input pages. The synthesized words can be used in word-spotting and many other historical document processing applications. The proposed approach provides an intuitive and easy-to-use framework to search for a keyword in the rest of the manuscript. Our experimental study shows that our approach enables accurate results in word spotting algorithms.
Majeed Kassis, Jihad El-Sana
DAS2
2016 Keyword Retrieval Using Scale-Space Pyramid
abstract
We propose a pyramid-based method for keyword spotting in historical document images. The documents are represented by a scale-space pyramid of their features. The search for a query keyword begins at the highest level of the pyramid, where the initial candidates for matching are located. The candidates are further refined at each level of the pyramid. The number of levels is adaptive and depends on the length of the query word. The results from all the document images are combined and ranked. We compare two feature representations, grid-based and continuous, and show that continuous feature representation outperforms the grid-based representation. In order to reduce the memory used to store the scale-space pyramid of features, we discuss and compare two compressing approaches. The proposed method was evaluated on four different collections of historical documents achieving state-of-the-art results.
Irina Rabaev, Klara Kedem, Jihad El-Sana
DAS3
2016 Scribble Based Interactive Page Layout Segmentation Using Gabor Filter
abstract
This paper presents an interactive approach for fast and accurate page layout segmentation. It is a scribble-based interactive segmentation approach, where the user draws scribbles on the various regions and the system performs page layout segmentation. The user can correct and refine the resulting segmentation by drawing new scribbles. To classify the various regions of the page, we apply a bank of Gabor filters, in several orientations and multiple frequencies, to capture the orientation, the stroke width, and size of the text. These properties also implicitly encode the writing style of the document. After combining the responses of the Gabor filter into a feature matrix, we classify various regions of the document by applying graph cuts, while taking into account the user made scribbles. The presented approach is very fast, easy to use, robust to user interaction, and provides accurate results.
Majeed Kassis, Jihad El-Sana
ICFHR2
2016 Word Spotting Using Radial Descriptor Graph
abstract
In this paper we present, the Radial Descriptor Graph, a novel approach to compare pictorial representation of handwritten text, which is based on the radial descriptor. To build a radial descriptor graph, we compute the radial descriptor and generate feature points. These points are the nodes of the graph, and each adjacent points are connected to its adjacent node to form a planar graph. Then we iteratively reduce the edges of the graph, by merging adjacent nodes, to form a multilevel hierarchical representation of the graph. To compare two pictorial representations, we measure the distance between their correspondence planar graphs, after calculating the dominant signal for each node. The graph matching is based on optimizing the function that takes into account the distance between the feature points and the structure of the graphs. The distance between two radial descriptors is computed by measuring the difference between their corresponding dominant signals. We have tested our approach on three different datasets and obtained encouraging results.
Majeed Kassis, Jihad El-Sana
ICFHR2
2015 Simplifying the reading of historical manuscripts
abstract
Complex document layouts pose prominent challenges for document image understanding algorithms. These layouts impose irregularities on the location of text paragraphs which consequently induces difficulties in reading the text. In this paper we present a robust framework for analyzing historical manuscripts with complex layouts. This framework aims to provide a convenient reading experience for historians through topnotch algorithms for text localization, classification and dewarping. We segment text into spatially coherent regions and text-lines using texture-based filters and refine this segmentation by exploiting Markov Random Fields (MRFs). A principled technique is presented for dewarping curvy text regions using a non-linear geometric transformation. The framework has been validated using a subset of a publicly available dataset of historical documents and it provided promising results.
Abedelkadir Asi, Rafi Cohen, Klara Kedem, Jihad El-Sana
ICDAR4
2015 Aligning transcript of historical documents using energy minimization
abstract
An ongoing considerable effort for digitizing historical manuscripts has produced images of original manuscripts, some accompanied by transcripts. Aligning the text in the input image with the text in the transcript will allow learning, training and evaluating recognition algorithms. Here we propose a system that computes the alignment by formulating the problem as an energy minimization task, where the alignment is performed between the input line image to a synthetic one. The energy function works at a connected component level and it combines a visual similarity measure and a learned distance metric that separates between inter-word and intra-word connected components.
Rafi Cohen, Irina Rabaev, Jihad El-Sana, Klara Kedem, Its'hak Dinstein
ICDAR3
2015 Word of blobs
abstract
In this paper, we present a novel scheme for subdividing a pictorial representation of a word or word-part into a sequence of blobs, that resemble the stroke representing the word. These blobs are generated by applying a bank of Gabor filters that capture the width of the strokes in multiple directions and segment the strong response regions. From the resulting blobs we extract representative features that are combined using bag-of-features. The proposed scheme is robust; i.e., insensitive to noise, and works directly on gray scale images. It represents the handwritten curves as a sequence of elliptic blobs, whose width is similar to that of the original handwriting. We incorporated the proposed approach in word spotting procedure and evaluated its performance on Arabic handwritten datasets.
Jihad El-Sana, Klara Kedem
ICDAR1
2015 Sphere intersection 3D shape descriptor (SID)
Kirill Pevzner, Andrei Sharf, Jihad El-Sana
Comput. Aided Geom. Des.3
2014 The Influence of Language Orthographic Characteristics on Digital Word Recognition
abstract
We study the effect of language orthographic characteristics on the performance of digital word recognition in degraded documents such as historical documents. We provide a rigorous scheme for quantifying the influence of the orthographic characteristics on the quality of word recognition in such documents. We study and compare several orthographic characteristics for four natural languages and measure the effect of each individual characteristic on the digital word recognition process. To this end we create synthetic languages, for which all characteristics, except the one we examine, are identical, and measure the performance of two word recognition algorithms on synthetic documents of these languages. We examine and summarize the influence of the values of each characteristic on the performance of these word recognition methods.
Ofer Biller, Jihad El-Sana, Klara Kedem
Document Analysis Systems2
2014 A Coarse-to-Fine Approach for Layout Analysis of Ancient Manuscripts
abstract
Many applications along the manuscript analysis pipeline rely on the accuracy of pre-processing steps. Perfectly detecting the main text area in ancient historical documents is of great importance for these applications. We propose a learning-free approach to detect the main text area in ancient manuscripts. First, we coarsely segment the main text area by using a texture-based filter. Then, we refine the segmentation by formulating the problem as an energy minimization task and achieving the minimum using graph cuts. The energy function is derived from properties of the text components. Spatial coherence of the segmented text regions is explicitly encouraged by the energy function. We evaluate the suggested method on a publicly available dataset of 38 historical document images. Experiments show that the suggested approach outperforms another state-of-the-art page segmentation method in terms of segmentation quality and time performance.
Abedelkadir Asi, Rafi Cohen, Klara Kedem, Jihad El-Sana, Its'hak Dinstein
ICFHR4
2014 Document Writer Analysis with Rejection for Historical Arabic Manuscripts
abstract
Determining the individuality of handwriting in ancient manuscripts is an important aspect of the manuscript analysis process. Automatic identification of writers in historical manuscripts can support historians to gain insights into manuscripts with missing metadata such as writer name, period, and origin. In this paper writer classification and retrieval approaches for multi-page documents in the context of historical manuscripts are presented. The main contribution is a learning-based rejection strategy which utilizes writer retrieval and support vector machines for rejecting a decision if no corresponding writer can be found for a query manuscript. Experiments using different feature extraction methods demonstrate the abilities of our proposed methods. A dedicated data set based on a publicly available database of historical Arabic manuscripts was used and the experiments show promising results.
Daniel Fecker, Abedelkadir Asi, Werner Pantke, Volker Märgner, Jihad El-Sana, Tim Fingscheidt
ICFHR5
2014 Word Spotting Using Radial Descriptor
abstract
Word spotting provides an efficient mechanism for word searching and indexing of historical documents. In this paper we present a novel feature descriptor, radial descriptor, and study its application for spotting word parts on Arabic historical documents. The radial descriptor aims to capture the intensity variance of the neighborhood of a point at various scale space levels. Features with high variance along multiple levels are used to describe the shape of a word according to the bag-of-features model. The distance between two word-parts is computed as the distance between their occurrence probability histograms. We have tested our approach on a large dataset of Arabic word-parts and received encouraging results.
Majeed Kassis, Jihad El-Sana
ICFHR2
2014 Writer Identification for Historical Arabic Documents
abstract
Identification of writers of handwritten historical documents is an important and challenging task. In this paper we present several feature extraction and classification approaches for the identification of writers in historical Arabic manuscripts. The approaches are able to successfully identify writers of multipage documents. The feature extraction methods rely on different principles, such as contour-, textural- and key point-based and the classification schemes are based on averaging and voting. For all experiments a dedicated data set based on a publicly available database is used. The experiments show promising results and the best performance was achieved using a novel feature extraction based on key point descriptors.
Daniel Fecker, Abedelkadir Asi, Volker Märgner, Jihad El-Sana, Tim Fingscheidt
ICPR4
2014 Text line extraction for historical document images
Raid Saabni, Abedelkadir Asi, Jihad El-Sana
Pattern Recognit. Lett.3
2013 Text Line Detection in Corrupted and Damaged Historical Manuscripts
abstract
Most of the algorithms proposed for text line detection are designed to process binary images as input. For severely degraded documents, binarization often introduces significant noise and other artifacts. In this work we present a novel method designed to detect text lines directly in gray scale images. The method consists of two stages. Potential characters are detected in the first stage. This is done by analyzing the evolution maps of connected components obtained by a sliding threshold. The detected potential characters are grouped into text lines in the second stage using sweep-line approach. The suggested method is especially powerful when applied to torn and damaged documents that other algorithms are not able to deal with.
Irina Rabaev, Ofer Biller, Jihad El-Sana, Klara Kedem, Its'hak Dinstein
ICDAR3
2013 Markerless 3D gesture-based interaction for handheld Augmented Reality interfaces
abstract
Conventional 2D touch-based interaction methods for handheld Augmented Reality (AR) cannot provide intuitive 3D interaction due to a lack of natural gesture input with real-time depth information. The goal of this research is to develop a natural interaction technique for manipulating virtual objects in 3D space on handheld AR devices. We present a novel method that is based on identifying the positions and movements of the user's fingertips, and mapping these gestures onto corresponding manipulations of the virtual objects in the AR scene. We conducted a user study to evaluate this method by comparing it with a common touch-based interface under different AR scenarios. The results indicate that although our method takes longer time, it is more natural and enjoyable to use.
Huidong Bai, Jihad El-Sana, Mark Billinghurst
ISMAR3
2013 Synthesizing reality for realistic physical behavior of virtual objects in augmented reality applications for smart-phones
abstract
This paper presents a framework for augmented reality applications that runs on smart mobile phones and enables realistic physical behavior of the virtual objects in the real-world. The used mobile phone is equipped with two cameras and provides stereo images and live video. The two images are used to reconstruct a 3D representation of the real world, which is good enough to enable physical interaction between the virtual and the real-world objects, but not fine enough to synthesize the real-world view for back projection. The visual synthesis of the real-world view is done through the video stream. The constructed 3D representation is registered in the real-world view and used to place the virtual objects, determine their physical behavior, and detect collision with objects in the realworld view. The 3D reconstruction is not performed at each frame, but applied when necessary based on the position of the dynamic objects. Pose estimation is determined based on the movement of the mobile phone (gyroscope and accelerometer) and the viewed images. A physics engine, which utilizes the gravity vector obtained from the accelerometer sensor of the mobile device, is integrated into our framework. The physics engine equips virtual objects with realistic physical behavior.
Nir Amar, Meir Raviv, Barak David, Oleg Chernoguz, Jihad El-Sana
VR5
2013 Comprehensive synthetic Arabic database for on/off-line script recognition research
Raid Saabni, Jihad El-Sana
Int. J. Document Anal. Recognit.2
2012 Evolution Maps for Connected Components in Text Documents
abstract
For highly degraded text documents, common tasks such as binarization and line extraction, remain difficult tasks. Equipped with a reliable information regarding the distribution of character dimensions in the document, one can improve results of these algorithms significantly. We introduce a novel perspective of the image data which maps the evolution of connected components along the change in gray scale threshold. We use these maps to provide a robust algorithm for extracting information about character dimensions in degraded documents, and demonstrate improvement in binarization results using this information. We analyze statistically the characteristics of the evolution maps for text documents, and compare our results with ground truth data.
Ofer Biller, Klara Kedem, Its'hak Dinstein, Jihad El-Sana
ICFHR4
2012 Layout Analysis for Arabic Historical Document Images Using Machine Learning
abstract
Page layout analysis is a fundamental step of any document image understanding system. We introduce an approach that segments text appearing in page margins (a.k.a side-notes text) from manuscripts with complex layout format. Simple and discriminative features are extracted in a connected-component level and subsequently robust feature vectors are generated. Multilayer perception classifier is exploited to classify connected components to the relevant class of text. A voting scheme is then applied to refine the resulting segmentation and produce the final classification. In contrast to state-of-the-art segmentation approaches, this method is independent of block segmentation, as well as pixel level analysis. The proposed method has been trained and tested on a dataset that contains a variety of complex side-notes layout formats, achieving a segmentation accuracy of about 95%.
Syed Saqib Bukhari, Thomas M. Breuel, Abedelkadir Asi, Jihad El-Sana
ICFHR4
2012 Occluded Character Restoration Using Active Contour with Shape Priors
abstract
Broken or partially visible characters is common phenomenon in historical documents. It stems from various factors, such as overlaid text or degradation. Restoring such characters is necessary for document analysis applications. This paper presents a new approach for restoring underlaying Hebrew broken characters that were partially occluded by Arabic text in a palimpsest. We apply text recognition on the fragments of the Hebrew letters and select the k candidate letters that match best the fragments of the Hebrew letters. We then complete the broken Hebrew characters using active contour with the k candidates as shape priors. We use a modified geodesic active contour, which we tailored to occluded text restoration. It is initialized on the fragments of the Hebrew text, then it undergoes an expansion phase and a contraction phase via the occluding Arabic text to form the restored Hebrew character. We measure the distance between the completed character and its corresponding priors and choose the shape with the minimal distance as the reconstructed character. Experimental results are presented. On average 68% of the characters were correctly reconstructed.
Rafi Cohen, Klara Kedem, Its'hak Dinstein, Jihad El-Sana
ICFHR4
2012 CudaHull: Fast parallel 3D convex hull on the GPU
Ayal Stein, Eran Geva, Jihad El-Sana
Comput. Graph.3
2012 Displacement patches for view-dependent rendering
Yotam Livny, Gilad Bauman, Jihad El-Sana
Vis. Comput.3
2011 Case Study in Hebrew Character Searching
abstract
Searching for a letter or a word in historical documents is a practical challenge due to the various degradations present in such documents and the wide variance of handwriting. Searching in historical Hebrew documents is somewhat harder because of high similarities among Hebrew characters. In order to determine the features and their combinations appropriate for recognizing Hebrew script, we study a range of known features using a Dynamic Time Warping algorithm. In addition we describe a novel meth od for feature-based searching, which uses a number of models for the same character. This method is based on our original DTW algorithm that can match fragments of several models of the same character to match a query character. Consequently, we are not limited to any particular model of the character set. Application of this method leads to a significant improvement, even when using a small set of models.
Irina Rabaev, Ofer Biller, Jihad El-Sana, Klara Kedem, Its'hak Dinstein
ICDAR3
2011 Language-Independent Text Lines Extraction Using Seam Carving
abstract
In this paper, we present a novel language-independent algorithm for extracting text-lines from handwritten document images. Our algorithm is based on the seam carving approach for content aware image resizing. We adopted the signed distance transform to generate the energy map, where extreme points indicate the layout of text-lines. Dynamic programming is then used to compute the minimum energy left-to-right paths (seams), which pass along the ``middle`` of the text-lines. Each path intersects a set of components, which determine the extracted text-line and estimate its hight. The estimated hight determines the text-line's region, which guides splitting touching components among consecutive lines. Unassigned components that fall within the region of a text-line are added to the components list of the line. The components between two consecutive lines are processed when the two lines are extracted and assigned to the closest text-line, based on the attributes of extracted lines, the sizes and positions of components. Our experimental results on Arabic, Chinese, and English historical documents show that our approach manage to separate multi-skew text blocks into lines at high success rates.
Raid Saabni, Jihad El-Sana
ICDAR2
2011 Segmentation-Free Online Arabic Handwriting Recognition
abstract
Arabic script is naturally cursive and unconstrained and, as a result, an automatic recognition of its handwriting is a challenging problem. The analysis of Arabic script is further complicated in comparison to Latin script due to obligatory dots/stokes that are placed above or below most letters. In this paper, we introduce a new approach that performs online Arabic word recognition on a continuous word-part level, while performing training on the letter level. In addition, we appropriately handle delayed strokes by first detecting them and then integrating them into the word-part body. Our current implementation is based on Hidden Markov Models (HMM) and correctly handles most of the Arabic script recognition difficulties. We have tested our implementation using various dictionaries and multiple writers and have achieved encouraging results for both writer-dependent and writer-independent recognition.
Fadi Biadsy, Raid Saabni, Jihad El-Sana
Int. J. Pattern Recognit. Artif. Intell.3
2011 Shape Recognition and Pose Estimation for Mobile Augmented Reality
abstract
Nestor is a real-time recognition and camera pose estimation system for planar shapes. The system allows shapes that carry contextual meanings for humans to be used as Augmented Reality (AR) tracking targets. The user can teach the system new shapes in real time. New shapes can be shown to the system frontally, or they can be automatically rectified according to previously learned shapes. Shapes can be automatically assigned virtual content by classification according to a shape class library. Nestor performs shape recognition by analyzing contour structures and generating projective-invariant signatures from their concavities. The concavities are further used to extract features for pose estimation and tracking. Pose refinement is carried out by minimizing the reprojection error between sample points on each image contour and its library counterpart. Sample points are matched by evolving an active contour in real time. Our experiments show that the system provides stable and accurate registration, and runs at interactive frame rates on a Nokia N95 mobile phone.
Nate Hagbi, Oriel Bergig, Jihad El-Sana, Mark Billinghurst
IEEE Trans. Vis. Comput. Graph.3
2010 In-Place Sketching for content authoring in Augmented Reality games
abstract
Sketching leverages human skills for various purposes. In-Place Augmented Reality Sketching experiences build on the intuitiveness and flexibility of hand sketching for tasks like content creation. In this paper we explore the design space of In-Place Augmented Reality Sketching, with particular attention to content authoring in games. We propose a contextual model that offers a framework for the exploration of this design space by the research community. We describe a sketch-based AR racing game we developed to demonstrate the proposed model. The game is developed on top of our shape recognition and 3D registration library for mobile AR.
Nate Hagbi, Raphaël Grasset, Oriel Bergig, Mark Billinghurst, Jihad El-Sana
VR5
2010 Carving for topology simplification of polygonal meshes
Nate Hagbi, Jihad El-Sana
Comput. Aided Des.2
2010 Automatic reconstruction of tree skeletal structures from point clouds
abstract
Trees, bushes, and other plants are ubiquitous in urban environments, and realistic models of trees can add a great deal of realism to a digital urban scene. There has been much research on modeling tree structures, but limited work on reconstructing the geometry of real-world trees -- even then, most works have focused on reconstruction from photographs aided by significant user interaction. In this paper, we perform active laser scanning of real-world vegetation and present an automatic approach that robustly reconstructs skeletal structures of trees, from which full geometry can be generated. The core of our method is a series of global optimizations that fit skeletal structures to the often sparse, incomplete, and noisy point data. A significant benefit of our approach is its ability to reconstruct multiple overlapping trees simultaneously without segmentation. We demonstrate the effectiveness and robustness of our approach on many raw scans of different tree varieties.
Yotam Livny, Feilong Yan, Matt Olson, Baoquan Chen, Hao (Richard) Zhang, Jihad El-Sana
ACM Trans. Graph.6
2009 Hierarchical On-line Arabic Handwriting Recognition
abstract
In this paper, we present a multi-level recognizer for online Arabic handwriting. In Arabic script (handwritten and printed), cursive writing - is not a style - it is an inherent part of the script. In addition, the connection between letters is done with almost no ligatures, which complicates segmenting a word into individual letters. In this work, we have adopted the holistic approach and avoided segmenting words into individual letters. To reduce the search space, we apply a series of filters in a hierarchical manner. The earlier filters perform light processing on a large number of candidates, and the later filters perform heavy processing on a small number of candidates. In the first filter, global features and delayed strokes patterns are used to reduce candidate word-part models. In the second filter, local features are used to guide a dynamic time warping (DTW) classification. The resulting k top ranked candidates are sent for shape context based classifier, which determines the recognized word-part. In this work, we have modified the classic DTW to enable different costs for the different operations and control their behavior. We have performed several experimental tests and have received encouraging results.
Raid Saabni, Jihad El-Sana
ICDAR2
2009 Efficient Generation of Comprehensive Database for Online Arabic Script Recognition
abstract
The difficulties in segmenting cursive words into individual characters have shifted the focus of handwriting recognition research from segmentation-based approaches to segmentation-free (holistic) methods. However, maintaining and training large number of prototypes (models) that represent the words in the dictionary make the training process extremely expensive and difficult in computing resources. In this paper we present an efficient system that automatically generates prototypes for each word in a given dictionary using multiple appearance of each letter shape. Multiple appearance allows for many permutation of shapes for each word and thus complicates searching for the right prototype. To simplify the training, reduce the maintained prototypes, and avoid over fitting, we used dimensionality reduction followed by clustering techniques to reduce the size of these sets without affecting their ability to represent the wide variations of the handwriting styles. A set of generated fonts are created by professional writers imitating all handwriting styles for each character in each position. These fonts are used to generate all shapes for writing each word-part in a comprehensive dictionary. Principal component analysis and k-means clustering techniques are performed to select the minimal number of shapes representing the wide variations of handwriting styles for a word-part. Experimental results using an online recognition system proves the credibility of this process compared to manually generated databases.
Raid Saabni, Jihad El-Sana
ICDAR2
2009 In-place 3D sketching for authoring and augmenting mechanical systems
abstract
We present a framework for authoring three-dimensional virtual scenes for Augmented Reality (AR) which is based on hand sketching. Sketches consisting of multiple components are used to construct a 3D virtual scene augmented on top of the real drawing. Model structure and properties can be modified by editing the sketch itself and printed content can be combined with hand sketches to form a single scene. Authoring by sketching opens up new forms of interaction that have not been previously explored in Augmented Reality. To demonstrate the technology, we implemented an application that constructs 3D AR scenes of mechanical systems from freehand sketches, and animates the scenes using a physics engine. We provide examples of scenes composed from trihedral solid models, forces, and springs. Finally, we describe how sketch interaction can be used to author complicated physics experiments in a natural way.
Oriel Bergig, Nate Hagbi, Jihad El-Sana, Mark Billinghurst
ISMAR3
2009 Shape recognition and pose estimation for mobile augmented reality
abstract
In this paper we present Nestor, a system for real-time recognition and camera pose estimation from planar shapes. The system allows shapes that carry contextual meanings for humans to be used as augmented reality (AR) tracking fiducials. The user can teach the system new shapes at runtime by showing them to the camera. The learned shapes are then maintained by the system in a shape library. Nestor performs shape recognition by analyzing contour structures and generating projective invariant signatures from their concavities. The concavities are further used to extract features for pose estimation and tracking. Pose refinement is carried out by minimizing the reprojection error between sample points on each image contour and its library counterpart. Sample points are matched by evolving an active contour in real time. Our experiments show that the system provides stable and accurate registration, and runs at interactive frame rates on a Nokia N95 mobile phone.
Nate Hagbi, Oriel Bergig, Jihad El-Sana, Mark Billinghurst
ISMAR3
2009 Seamless patches for GPU-based terrain rendering
Yotam Livny, Zvi Kogan, Jihad El-Sana
Vis. Comput.3
2008 A Carving Framework for Topology Simplification of Polygonal Meshes
Nate Hagbi, Jihad El-Sana
GMP2
2008 In-place Augmented Reality
abstract
In this paper we present a new vision-based approach for transmitting virtual models for augmented reality (AR). A two dimensional representation of the virtual models is embedded in a printed image. We apply image-processing techniques to interpret the printed image and extract the virtual models, which are then overlaid back on the printed image. The main advantages of our approach are: (1) the image of the embedded virtual models and their behaviors are understandable to a human without using an AR system, and (2) no database or network communication is required to retrieve the models. The latter is useful in scenarios with large numbers of users. We implemented an AR system that demonstrates the feasibility of our approach. Applications in education, advertisement, gaming, and other domains can benefit from our approach, since content providers need only to publish the printed content and all virtual information arrives with it.
Nate Hagbi, Oriel Bergig, Jihad El-Sana, Klara Kedem, Mark Billinghurst
ISMAR3
2008 Interactive GPU-based adaptive cartoon-style rendering
Yotam Livny, Michael Press, Jihad El-Sana
Vis. Comput.3
2008 A GPU persistent grid mapping for terrain rendering
Yotam Livny, Neta Sokolovsky, Tal Grinshpoun, Jihad El-Sana
Vis. Comput.4
2002 Optimized View-Dependent Rendering for Large Polygonal Datasets
abstract
In this paper we are presenting a novel approach for rendering large datasets in a view-dependent manner. In a typical view-dependent rendering framework, an appropriate level of detail is selected and sent to the graphics hardware for rendering at each frame. In our approach, we have successfully managed to speed up the selection of the level of detail as well as the rendering of the selected levels. We have accelerated the selection of the appropriate level of detail by not scanning active nodes that do not contribute to the incremental update of the selected level of detail. Our idea is based on imposing a spatial subdivision over the view-dependence trees data-structure, which allows spatial tree cells to refine and merge in real-time rendering to comply with the changes in the active nodes list. The rendering of the selected level of detail is accelerated by using vertex arrays. To overcome the dynamic changes in the selected levels of detail we use multiple small vertex arrays whose sizes depend on the memory on the graphics hardware. These multiple vertex arrays are attached to the active cells of the spatial tree and represent the active nodes of these cells. These vertex arrays, which are sent to the graphics hardware at each frame, merge and split with respect to the changes in the cells of the spatial tree.
Jihad El-Sana, Eitan Bachmant
IEEE Visualization1
2002 Integrating motion perception with view-dependent rendering for dynamic environments
Jihad El-Sana, Nir Asis, Ofer Hadar
Comput. Graph.1
2001 Integrating Occlusion Culling with View-Dependent Rendering
abstract
We present an approach that integrates occlusion culling within the view-dependent rendering framework. View-dependent rendering provides the ability to change level of detail over the surface seamlessly and smoothly in real-time. The exclusive use of view-parameters to perform level-of-detail selection causes even occluded regions to be rendered in high level of detail. To overcome this serious drawback we have integrated occlusion culling into the level selection mechanism. Because computing exact visibility is expensive and it is currently not possible to perform this computation in real time, we use a visibility estimation technique instead. Our approach reduces dramatically the resolution at occluded regions.
Jihad El-Sana, Neta Sokolovsky, Cláudio T. Silva
IEEE Visualization1
2000 Multi-user view-dependent rendering
abstract
We present a novel architecture which allows rendering of a large-shared dataset at interactive rates on an inexpensive workstation. The idea is based on view-dependent rendering on a client-server network. The server stores the large dataset and manages the selection of the various levels of detail while the inexpensive clients receive a stream of update operations that generate the appropriate level of detail in an incremental fashion. These update operations are based on changes in the clients' view-parameters. Our approach dramatically reduces the amount of memory needed by each client and the entire computing system since the dataset is stored only once on the server's local memory. In addition, it decreases the load on the network as results of the incremental update contributed by view-dependent rendering.
Jihad El-Sana
IEEE Visualization1
2000 Efficiently computing and updating triangle strips for real-time rendering
Jihad El-Sana, Francine Evans, Aravind Kalaiah, Amitabh Varshney, Steven Skiena, Elvir Azanli
Comput. Aided Des.1
2000 Directional Discretized Occluders for Accelerated Occlusion Culling
abstract
We present a technique for accelerating the rendering of high depth‐complexity scenes. In a preprocessing stage, we approximate the input model with a hierarchical data structure and compute simple view‐dependent polygonal occluders to replace the complex input geometry in subsequent visibility queries. When the user is inspecting and visualizing the input model, the computed occluders are used to avoid rendering geometry which cannot be seen. Our method has several advantages which allow it to perform conservative visibility queries efficiently and it does not require any special graphics hardware. The preprocessing step of our approach can also be used within the framework of other visibility culling methods which need to pre‐select or pre‐render occluders. In this paper, we describe our technique and its implementation in detail, and provide experimental evidence of its performance. In addition, we briefly discuss possible extensions of our algorithm.
Fausto Bernardini, James T. Klosowski, Jihad El-Sana
Comput. Graph. Forum3
2000 External Memory View-Dependent Simplification
abstract
In this paper, we propose a novel external‐memory algorithm to support view‐dependent simplification for datasets much larger than main memory. In the preprocessing phase, we use a new spanned sub‐meshes simplification technique to build view‐dependence trees I/O‐efficiently, which preserves the correct edge collapsing order and thus assures the run‐time image quality. We further process the resulting view‐dependence trees to build the meta‐node trees, which can facilitate the run‐time level‐of‐detail rendering and is kept in disk. During run‐time navigation, we keep in main memory only the portions of the meta‐node trees that are necessary to render the current level of details, plus some prefetched portions that are likely to be needed in the near future. The prefetching prediction takes advantage of the nature of the run‐time traversal of the meta‐node trees, and is both simple and accurate. We also employ the implicit dependencies for preventing incorrect foldovers, as well as main‐memory buffer management and parallel processes scheme to separate the disk accesses from the navigation operations, all in an integrated manner. The experiments show that our approach scales well with respect to the main memory size available, with encouraging preprocessing and run‐time rendering speeds and without sacrificing the image quality.
Jihad El-Sana, Yi-Jen Chiang
Comput. Graph. Forum1
1999 View-Dependent Topology Simplification
Jihad El-Sana, Amitabh Varshney
EGVE1
1999 Haptic sculpting of dynamic surfaces
abstract
Conventional free-form surface design usually require tedious control-point manipulation and/or painstaking constraint specification via unnatural mouse-based interfaces. This paper presents a novel haptic approach for the direct manipulation of physics-based B-spline surfaces. Our method permits users to interactively sculpt virtual yet real material with a standard haptic device, and feel the physically realistic presence of virtual B-spline objects with force feedback throughout the design process. We aim to develop various haptic sculpting tools to expedite the direct manipulation of B-spline surfaces with haptic feedback and constraints. One significant contribution of this paper is that point, normal, and curvature constraints can be specified interactively and modified naturally using forces. We propose and formulate a dual representation for Bspline surfaces in both physical and mathematical space. This massspring model is mathematically constrained by the B-spline surface throughout the sculpting session. The equations of motion controlling the physical behavior of the B-spline surface are solved using a tractable numerical solver in real-time. The integration of haptics with traditional geometric modeling will increase the bandwidth of human-computer interaction, and thus shorten the time-consuming design cycle. We envision that this integrated approach promises a much greater potential in computer-integrated design and manufacturing, haptic interface, interactive graphics, medical applications, and virtual environments.
Frank Dachille, Hong Qin 0001, Arie E. Kaufman, Jihad El-Sana
SI3D4
1999 Skip Strips: Maintaining Triangle Strips for View-Dependent Rendering
abstract
View-dependent simplification has emerged as a powerful tool for graphics acceleration in visualization of complex environments. However, view-dependent simplification techniques have not been able to take full advantage of the underlying graphics hardware. Specifically, triangle strips are a widely used hardware-supported mechanism to compactly represent and efficiently render static triangle meshes. However, in a view-dependent framework, the triangle mesh connectivity changes at every frame, making it difficult to use triangle strips. We present a novel data structure, Skip Strip, that efficiently maintains triangle strips during such view-dependent changes. A Skip Strip stores the vertex hierarchy nodes in a skip-list-like manner with path compression. We anticipate that Skip Strips will provide a road map to combine rendering acceleration techniques for static datasets, typical of retained-mode graphics applications, with those for dynamic datasets found in immediate-mode applications.
Jihad El-Sana, Elvir Azanli, Amitabh Varshney
IEEE Visualization1
1999 Generalized View-Dependent Simplification
abstract
We propose a technique for performing view‐dependent geometry and topology simplifications for level‐of‐detail‐based renderings of large models. The algorithm proceeds by preprocessing the input dataset into a binary tree, the view‐dependence tree of general vertex‐pair collapses. A subset of the Delaunay edges is used to limit the number of vertex pairs considered for topology simplification. Dependencies to avoid mesh foldovers in manifold regions of the input object are stored in the view‐dependence tree in an implicit fashion. We have observed that this not only reduces the space requirements by a factor of two, it also highly localizes the memory accesses at run time. The view‐dependence tree is used at run time to generate the triangles for display. We also propose a cubic‐spline‐based distance metric that can be used to unify the geometry and topology simplifications by considering the vertex positions and normals in an integrated manner.
Jihad El-Sana, Amitabh Varshney
Comput. Graph. Forum1
1998 Topology Simplification for Polygonal Virtual Environments
abstract
We present a topology simplifying approach that can be used for genus reductions, removal of protuberances, and repair of cracks in polygonal models in a unified framework. Our work is complementary to the existing work on geometry simplification of polygonal datasets and we demonstrate that using topology and geometry simplifications together yields superior multiresolution hierarchies than is possible by using either of them alone. Our approach can also address the important issue of repair of cracks in polygonal models, as well as for rapid identification and removal of protuberances based on internal accessibility in polygonal models. Our approach is based on identifying holes and cracks by extending the concept of /spl alpha/-shapes to polygonal meshes under the L/sub /spl infin// distance metric. We then generate valid triangulations to fill them using the intuitive notion of sweeping an L/sub /spl infin// cube over the identified regions.
Jihad El-Sana, Amitabh Varshney
IEEE Trans. Vis. Comput. Graph.1
1997 Controlled simplification of genus for polygonal models
abstract
Genus-reducing simplifications are important in constructing multiresolution hierarchies for level-of-detail-based rendering, especially for datasets that have several relatively small holes, tunnels, and cavities. We present a genus-reducing simplification approach that is complementary to the existing work on genus-preserving simplifications. We propose a simplification framework in which genus-reducing and genus-preserving simplifications alternate to yield much better multiresolution hierarchies than would have been possible by using either one of them. In our approach we first identify the holes and the concavities by extending the concept of /spl alpha/-hulls to polygonal meshes under the L/sub /spl infin// distance metric and then generate valid triangulations to fill them.
Jihad El-Sana, Amitabh Varshney
IEEE Visualization1
1997 Decomposing and Solving Timetabling Constraint Networks
abstract
The binary version of the school timetabling (STT) problem is a real‐world example of a constraint network that includes only constraints of inequality. A new and useful representation for this real‐world problem, the STT_Grid, leads to a generic decomposition technique. The paper presents proofs of necessary and sufficient conditions for the existence of a solution to decomposed STT_Grids. The decomposition procedure is of low enough complexity to be practical for large problems, such as a real‐world high school. To test the decomposition approach, a typical high school was analyzed and used as a model for generating STT_Grids of various sizes. Experiments were conducted to test the difficulty of large STT networks and their solution by decomposition. The experimental results show that the decomposition procedure enables the solution of large STT_Grids (620 variables for a real school) in reasonable time. The constraint network of a typical STT_Grid is sparse and belongs to the class of easy problems. Still, due to the sizes of STTs, good constraint satisfaction problem search techniques (i.e., BackJumping and ForwardChecking) do not terminate in reasonable times for STT_Grids that are larger than 300 variables.
Amnon Meisels, Jihad El-Sana, Ehud Gudes
Comput. Intell.2
1997 Adaptive Real-Time Level-of-Detail-Based Rendering for Polygonal Models
abstract
We present an algorithm for performing adaptive real-time level-of-detail-based rendering for triangulated polygonal models. The simplifications are dependent on viewing direction, lighting, and visibility and are performed by taking advantage of image-space, object-space, and frame-to-frame coherences. In contrast to the traditional approaches of precomputing a fixed number of level-of-detail representations for a given object, our approach involves statically generating a continuous level-of-detail representation for the object. This representation is then used at run time to guide the selection of appropriate triangles for display. The list of displayed triangles is updated incrementally from one frame to the next. Our approach is more effective than the current level-of-detail-based rendering approaches for most scientific visualization applications, where there are a limited number of highly complex objects that stay relatively close to the viewer. Our approach is applicable for scalar (such as distance from the viewer) as well as vector (such as normal direction) attributes.
Julie C. Xia, Jihad El-Sana, Amitabh Varshney
IEEE Trans. Vis. Comput. Graph.2