Simone Marinai

dblp:66/5842 · DBLP profile ↗
← Back
51ranked-venue papers
18as first author
12since 2021 · last 2027
0000-0002-6702-2277ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 13 first-author · 9 since 2021Databases, data management, data science and information retrieval · 30 · 12 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2027 The DocLap integrated system for understanding multi-page scholarly documents
abstract
Computational understanding of documents is focused on visual and text analysis, building upon computer vision and natural language processing. With the advent of transformers document understanding is even more shifted from the actual comprehension of documents, which relies on concurrent perception of text, layout elements and document structure, to convoluted feature representations. Recent trends for providing unified access to such representations go towards Large Language Models (LLMs). However, these models have limitations: they lack explainability, demand significant resources for training and inference, and are not well suited for processing extensive inputs nor for direct application in specialized domains. This paper aims at creating a comprehensive interface for document analysis, enabling multi-layered exploration and integrating diverse features and contextual information. By bridging diverse information, our work pursues the identification, characterization, and linking of visual elements to semantic and contextual data, leveraging LLMs for interoperability. This enables a unified access to textual, visual, and structural layers, embedding levels of structured knowledge directly in the LLM context. Recent advances in Retrieval Augmented Generation (RAG) are also exploited to address some LLM limitations related to context length, allowing access to latent information from document representations such as graph and vector embeddings. The association of structural information to visual data allows formal analysis of documents and is exploited in our model to enhance visual recognition, improved through multi-modal LLM correction supported by ontology-based constraint violation detection. The framework enables semantic retrieval over extracted information, providing direct access to the document structure which can be exploited in many applications such as Question Answering (QA) and document understanding. As a result of this work, the DocLap (Document Layout Parser) system for document analysis and retrieval is proposed, which enables the extraction of visual and semantic features from documents and makes them accessible through natural language in an integrated framework providing conversational reasoning. The system’s segmentation, error detection, and information retrieval capabilities are extensively evaluated through experiments.
Lorenzo Massai, Simone Marinai
Expert Syst. Appl.2
2026 Towards Non-Latin Text and Layout Personalization for Enhanced Readability
Rina Buoy, Dylan Berkamp Fouepe Dongmo, Vesal Khean, Simone Marinai, Koichi Kise
ICDAR (3)4
2026 BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
abstract
Abstract We present a unified dataset for document Question-Answering (QA), which is obtained combining several public datasets related to Document AI and visually rich document understanding (VRDU). Our main contribution is twofold: on the one hand we reformulate existing Document AI tasks, such as Information Extraction (IE), into a Question-Answering task, making it a suitable resource for training and evaluating Large Language Models; on the other hand, we release the OCR of all the documents and include the exact position of the answer to be found in the document image as a bounding box. Using this dataset, we explore the impact of different prompting techniques (that might include bounding box information) on the performance of open-weight models, identifying the most effective approaches for document comprehension.
Simone Giovannini, Fabio Coppini, Andrea Gemelli, Simone Marinai
Int. J. Document Anal. Recognit.4
2025 Visual Large Language Models for Graphics Understanding: A Case Study on Floorplan Images
abstract
This study explores the use of Vision Large Language Models (VLLMs) for identifying items in complex graphical documents. In particular, we focus on looking for furniture objects (e.g. beds, tables, and chairs) and structural items (doors and windows) in floorplan images. We evaluate one object detection model (YOLO) and state-of-the-art VLLMs on two datasets featuring diverse floorplan layouts and symbols. The experiments with VLLMs are performed with a zero-shot setting, meaning the models are tested without any training or fine-tuning, as well as with a few-shot approach, where examples of items to be found in the image are given to the models in the prompt. The results highlight the strengths and limitations of VLLMs in recognizing architectural elements, providing guidance for future research in the use multimodal vision-language models for graphics recognition.
Valeria Nardoni, Kimiya Noor Ali, Zahra Ziran, Simone Marinai
DocEng4
2025 Graph Convolutional Teacher-Student Framework for Writer Inspection from Intra-variable Handwritten Words
Kumari Priya, Aritra Dey, Chandranath Adak, Soumi Chattopadhyay, Sukalpa Chanda, Simone Marinai
ICDAR (3)7
2025 International journal on document analysis and recognition editorial leadership change
Daniel P. Lopresti, Koichi Kise, Simone Marinai
Int. J. Document Anal. Recognit.3
2024 Structure Matters: Analyzing Videos Via Graph Neural Networks for Social Media Platform Attribution
abstract
Detecting the origin of a digital video within a social network is a critical task that aids law enforcement and intelligence agencies in identifying the creators of misleading visual content. In this research, we introduce an innovative method for identifying the original social network of a video, even when the video has been altered through actions like group of frames removal and file container reconstruction. The proposed method takes advantage of the video encoding’s temporal uniformity, leveraging motion vectors to characterize the specific features associated to various social media platforms. Each video is represented by a graph where nodes correspond to macroblocks. These macroblocks are interconnected by following the inter-prediction rules outlined in the H.264/AVC codec standard. Such a structure can be then classified using a graph neural network to predict the platform on which the video has been shared. Experimental results demonstrate that this approach outperforms both codec- and content-based approaches, underscoring the effectiveness of a structural approach in attributing the social media platform from which videos originated.
Andrea Gemelli, Dasara Shullani, Daniele Baracchi, Simone Marinai, Alessandro Piva
ICASSP4
2024 Datasets and annotations for layout analysis of scientific articles
abstract
Abstract For a long time now, datasets containing scientific articles have been crucial to the analysis and recognition of document images. These document collections have frequently served as a testing ground for cutting-edge methods for optical character recognition, layout analysis, and document understanding in general. We thoroughly analyze and compare many datasets proposed for layout analysis of scientific documents, ranging from small collections of scanned papers to modern large-scale datasets containing digital-born papers, which have been proposed to train deep learning-based methods. Furthermore, we outline a detailed taxonomy of the annotation procedures used considering manual, automatic, and generative approaches, and we analyze their benefits and drawbacks. This survey is meant to provide the reader with a review of the most used benchmarks together with detailed information on data, annotations, and complexity, helping scholars to identify the most suitable dataset for their tasks of interest. We also discuss possible open problems to further enhance datasets to support research in the layout analysis of scientific articles.
Andrea Gemelli, Simone Marinai, Lorenzo Pisaneschi, Francesco Santoni
Int. J. Document Anal. Recognit.2
2024 Editorial for special issue on "advanced topics in document analysis and recognition"
Elisa H. Barney Smith, Marcus Liwicki, Liangrui Peng, Simone Marinai
Int. J. Document Anal. Recognit.4
2023 Deep-learning for dysgraphia detection in children handwritings
abstract
Early identification of dysgraphia in children is crucial for timely intervention and support. Traditional methods, such as the Brave Handwriting Kinder (BHK) test, which relies on manual scoring of handwritten sentences, are both time-consuming and subjective posing challenges in accurate and efficient diagnosis. In this paper, an approach for dysgraphia detection by leveraging smart pens and deep learning techniques is proposed, automatically extracting visual features from children's handwriting samples. To validate the solution, samples of children handwritings have been gathered and several interviews with domain experts have been conducted. The approach has been compared with an algorithmic version of the BHK test and with several elementary school teachers' interviews.
Andrea Gemelli, Simone Marinai, Emanuele Vivoli, Tamara Zappaterra
DocEng2
2023 Automatic generation of scientific papers for data augmentation in document layout analysis
abstract
Document layout analysis is an important task to extract information from scientific literature. Deep-learning solutions for document layout analysis require large collections of training data that are not always available. We generate a large number of synthetic pages to subsequently train a neural network to perform document object detection. The proposed pipeline allows users to deal with less common layouts for which it is not easy to find large annotated datasets. High-quality annotations for a small collection of papers are obtained through a semi-automatic approach. Then, a generative model, based on LayoutTransformer, is used to generate plausible layouts that are subsequently populated with random information to perform data augmentation. We evaluate the proposed method considering scientific articles with two different types of layouts: double and single columns. For double-column papers, we improve detection by 1% starting from 385 manually annotated scientific articles. For single-column papers, we improve detection by 49% starting from 218 articles.
Lorenzo Pisaneschi, Andrea Gemelli, Simone Marinai
Pattern Recognit. Lett.3
2022 Graph Neural Networks and Representation Embedding for Table Extraction in PDF Documents
abstract
Tables are widely used in several types of documents since they can bring important information in a structured way. In scientific papers, tables can sum up novel discoveries and summarize experimental results, making the research comparable and easily understandable by scholars. Several methods perform table analysis working on document images, losing useful information during the conversion from the PDF files since OCR tools can be prone to recognition errors, in particular for text inside tables. The main contribution of this work is to tackle the problem of table extraction, exploiting Graph Neural Networks. Node features are enriched with suitably designed representation embeddings. These representations help to better distinguish not only tables from the other parts of the paper, but also table cells from table headers. We experimentally evaluated the proposed approach on a new dataset obtained by merging the information provided in the PubLayNet and PubTables-1M datasets.
Andrea Gemelli, Emanuele Vivoli, Simone Marinai
ICPR3
2020 Text alignment in early printed books combining deep learning and dynamic programming
Zahra Ziran, Xavier Pic, Simone Undri Innocenti, Daniele Mugnai, Simone Marinai
Pattern Recognit. Lett.5
2019 Deep neural networks for record counting in historical handwritten documents
Samuele Capobianco, Simone Marinai
Pattern Recognit. Lett.2
2018 Offline Bengali Writer Verification by PDF-CNN and Siamese Net
abstract
Automated handwriting analysis is a popular area of research owing to the variation of writing patterns. In this research area, writer verification is one of the most challenging branches, having direct impact on biometrics and forensics. In this paper, we deal with offline writer verification on complex handwriting patterns. Therefore, we choose a relatively complex script, i.e., Indic Abugida script Bengali (or, Bangla) containing more than 250 compound characters. From a handwritten sample, the probability distribution functions (PDFs) of some handcrafted features are obtained and input to a convolutional neural network (CNN). For such a CNN architecture, we coin the term "PDFCNN", where handcrafted feature PDFs are hybridized with auto-derived CNN features. Such hybrid features are then fed into a Siamese neural network for writer verification. The experiments are performed on a Bengali offline handwritten dataset of 100 writers. Our system achieves encouraging results, which sometimes exceed the results of state-of-the-art techniques on writer verification.
Chandranath Adak, Simone Marinai, Bidyut B. Chaudhuri, Michael Blumenstein
DAS2
2017 DocEmul: A Toolkit to Generate Structured Historical Documents
abstract
We propose a toolkit to generate structured synthetic documents emulating the actual document production process. Synthetic documents can be used to train systems to perform document analysis tasks. In our case we address the record counting task on handwritten structured collections containing a limited number of examples. Using the DocEmul toolkit we can generate a larger dataset to train a deep architecture to predict the number of records for each page. The toolkit is able to generate synthetic collections and also perform data augmentation to create a larger trainable dataset. It includes one method to extract the page background from real pages which can be used as a substrate where records can be written on the basis of variable structures and using cursive fonts. Moreover, it is possible to extend the synthetic collection by adding random noise, page rotations, and other visual variations. We performed some experiments on two different handwritten collections using the toolkit to generate synthetic data to train a Convolutional Neural Network able to count the number of records in the real collections.
Samuele Capobianco, Simone Marinai
ICDAR2
2017 Partitioning Open Plan Areas in Floor Plans
abstract
We are developing an application to automatically generate an accessible graphic from a floor plan image. Floor plans generally contain large regions with functionally different sub-areas. A problem faced by visually impaired users in exploring such accessible floor plans is understanding the boundaries of these sub-areas. We present an effective method to partition such open plan areas. Initially, we conducted a formative user study to understand how people partition open plan areas. Based on the findings of the study, we identified a general set of guidelines for partitioning open plans. These guidelines were used to generate a set of candidate lines for sub-areas. An obstacle avoiding shortest-path Voronoi diagram was used to determine boundaries for each sub-area. Candidate lines such as wall extensions were automatically generated to replace the identified boundaries. We selected the best replacement for each boundary by scoring candidate lines using a set of criteria such as line length. Finally the proposed method was tested on a standard floor plan corpus using three novel measures.
Anuradha Madugalla, Kim Marriott, Simone Marinai
ICDAR3
2015 Deepdocclassifier: Document classification with deep Convolutional Neural Network
abstract
This paper presents a deep Convolutional Neural Network (CNN) based approach for document image classification. One of the main requirement of deep CNN architecture is that they need huge number of samples for training. To overcome this problem we adopt a deep CNN which is trained using big image dataset containing millions of samples i.e., ImageNet. The proposed work outperforms both the traditional structure similarity methods and the CNN based approaches proposed earlier. The accuracy of the proposed approach with merely 20 images per class outperforms the state-of-the-art by achieving classification accuracy of 68.25%. The best results on Tobbacoo-3428 dataset show that our proposed method outperforms the state-of-the-art method by a significant margin and achieved a median accuracy of 77.6% with 100 samples per class used for training and validation.
Muhammad Zeshan Afzal, Samuele Capobianco, Muhammad Imran Malik, Simone Marinai, Thomas M. Breuel, Andreas Dengel 0001, Marcus Liwicki
ICDAR4
2015 Accessible On-Line Floor Plans
abstract
Better access to on-line information graphics is a pressing need for people who are blind or have severe vision impairment. We present a new model for accessible presentation of on-line information graphics and demonstrate its use for presenting floor plans. While floor plans are increasingly provided on-line, people who are blind are at best provided with only a high-level textual description. This makes it difficult for them to understand the spatial arrangement of the objects on the floor plan. Our new approach provides users with significantly better access to such plans. The users can automatically generate an accessible version of a floor plan from an on-line floor plan image quickly and independently by using a web service. This generates a simplified graphic showing the rooms, walls, doors and windows in the original floor plan as well as a textual overview. The accessible floor plan is presented on an iPad using audio feedback. As the users touch graphic elements on the screen, the element they are touching is described by speech and non-speech audio in order to help them navigate the graphic.
Cagatay Goncu, Anuradha Madugalla, Simone Marinai, Kim Marriott
WWW3
2013 Reflowing and annotating scientific papers on eBook readers
abstract
Working with scientific and technical papers on small screen devices, such as tablets and eBook readers, is difficult since these works are often typeset in multiple columns with a relatively small font size.
Simone Marinai
ACM Symposium on Document Engineering1
2012 Displaying chemical structural formulae in ePub format
abstract
We describe one tool designed to enhance the visualization of chemical structural formulae in E-book readers. When dealing with small formulae, to avoid the pixelation effect with zoomed images, the formula is converted to a vectoral representation and then enlarged. On the opposite, large formulae are split in sub-images by cutting the image in suitable locations attempting to reduce the parts of the formula that are broken. In both cases the formulae are embedded in one ePub document that allows users to browse the chemical structure on most reading devices.
Simone Marinai, Stefano Quiriconi
ACM Symposium on Document Engineering1
2012 Object recognition in floor plans by graphs of white connected components
Alessio Barducci, Simone Marinai
ICPR2
2011 Conversion of PDF Books in ePub Format
abstract
In the last years the interest in e-book readers is significantly growing. Two main document formats are supported by most devices: PDF and ePub. The PDF format is widely used to share documents allowing a cross-platform readability. However, it is not ideal for a comfortable reading on small screens. On the opposite, the ePub format is re-flowable and it is well suited for e-book readers. In this paper we describe a system for the conversion of PDF books to the ePub format aiming at inverting the text formatting made during the pagination. To this purpose, layout analysis techniques are performed to identify the book's table of contents and the main functional regions such as chapters, paragraphs, and notes.
Simone Marinai, Emanuele Marino, Giovanni Soda
ICDAR1
2011 Using Earth Mover's Distance in the Bag-of-Visual-Words Model for Mathematical Symbol Retrieval
abstract
In this paper, the Earth Mover's Distance (EMD) is used as a similarity measure in the mathematical symbol retrieval task. The approach is based on the Bag-of-Visual-Words model. In our case the features extracted from each symbol are clustered by means of Self-Organizing Maps (SOM) and then occurrences of features in the clusters are accumulated in a vector of visual words. The comparison between the latter vectors is performed with the EMD which naturally allows to incorporate the topological organization of SOM clusters in the distance computation. The proposed approach is experimentally tested in a mathematical symbol retrieval task and compared with the cosine similarity and with some variants that have been recently proposed.
Simone Marinai, Beatrice Miotti, Giovanni Soda
ICDAR1
2011 Text retrieval from early printed books
Simone Marinai
Int. J. Document Anal. Recognit.1
2011 Report from the AND 2009 working group on noisy text datasets
Simone Marinai, Dimosthenis Karatzas
Int. J. Document Anal. Recognit.1
2010 Table of contents recognition for converting PDF documents in e-book formats
abstract
We describe one tool for Table of Content (ToC) identification and recognition from PDF books. This task is part of ongoing research on the development of tools for the semi-automatic conversion of PDF documents in the Epub format that can be read on several E-book devices. Among various sub-tasks, the ToC extraction and recognition is particularly useful for an easy navigation of book contents.
Simone Marinai, Emanuele Marino, Giovanni Soda
ACM Symposium on Document Engineering1
2010 Bag of Characters and SOM Clustering for Script Recognition and Writer Identification
abstract
In this paper, we describe a general approach for script (and language) recognition from printed documents and for writer identification in handwritten documents. The method is based on a bag of visual word strategy where the visual words correspond to characters and the clustering is obtained by means of Self Organizing Maps (SOM). Unknown pages (words in the case of script recognition) are classified comparing their vectorial representations with those of one training set using a cosine similarity. The comparison is improved using a similarity score that is obtained taking into account the SOM organization of cluster centroids. Promising results are presented for both printed documents and handwritten musical scores.
Simone Marinai, Beatrice Miotti, Giovanni Soda
ICPR1
2009 Metadata Extraction from PDF Papers for Digital Library Ingest
abstract
In this paper we analyze our recent research on the use of document analysis techniques for metadata extraction from PDF papers. We describe a package that is designed to extract basic metadata from these documents. The package is used in combination with a digital library software suite to easily build personal digital libraries. The proposed software is based on a suitable combination of several techniques that include PDF parsing, low level document image processing, and layout analysis. In addition, we use the information gathered from a widely known citation database (DBLP) to assist the tool in the difficult task of author identification. The system is tested on some paper collections selected from recent conference proceedings.
Simone Marinai
ICDAR1
2009 Mathematical Symbol Indexing Using Topologically Ordered Clusters of Shape Contexts
abstract
This paper addresses the indexing and retrieval of mathematical symbols from digitized documents. The proposed approach exploits Shape Contexts (SC) to describe the shape of mathematical symbols. Starting from the vector space method, that is based on SC clustering, we explore the use of topological ordered clusters to improve the retrieval performance. The clustering is computed by means of Self-Organizing Maps that organize the clusters in two dimensional topologically ordered feature maps. The retrieval performance are compared with those obtained using the K-means clustering on a large collection of mathematical symbols gathered from the widely used INFTY database.
Simone Marinai, Beatrice Miotti, Giovanni Soda
ICDAR1
2008 A Comparison of Clustering Methods for Word Image Indexing
abstract
In this paper we compare three clustering methods used to perform word image indexing. The three methods are: the Self-Organizing Map (SOM), the Growing Hierarchical Self-Organizing Map (GHSOM), and the Spectral Clustering. We test these methods on a real data set composed of word images extracted from an encyclopedia of the XIX-th Century. The word images are grouped on the basis of the clustering methods and subsequently retrieved identifying the closest clusters to a query word. The accuracy of the methods is compared evaluating the performance of the word retrieval algorithm. From the experimental results we conclude that methods designed to automatically determine the number and the structure of clusters, such as GHSOM, are particularly suitable in the context represented by our data set.
Simone Marinai, Emanuele Marino, Giovanni Soda
Document Analysis Systems1
2006 Efficient Word Retrieval by Means of SOM Clustering and PCA
Simone Marinai, Stefano Faini, Emanuele Marino, Giovanni Soda
Document Analysis Systems1
2006 Font Adaptive Word Indexing of Modern Printed Documents
abstract
We propose an approach for the word-level indexing of modern printed documents which are difficult to recognize using current OCR engines. By means of word-level indexing, it is possible to retrieve the position of words in a document, enabling queries involving proximity of terms. Web search engines implement this kind of indexing, allowing users to retrieve Web pages on the basis of their textual content. Nowadays, digital libraries hold collections of digitized documents that can be retrieved either by browsing the document images or relying on appropriate metadata assembled by domain experts. Word indexing tools would therefore increase the access to these collections. The proposed system is designed to index homogeneous document collections by automatically adapting to different languages and font styles without relying on OCR engines for character recognition. The approach is based on three main ideas: the use of Self Organizing Maps (SOM) to perform unsupervised character clustering, the definition of one suitable vector-based word representation whose size depends on the word aspect-ratio, and the run-time alignment of the query word with indexed words to deal with broken and touching characters. The most appropriate applications are for processing modern printed documents (17th to 19th centuries) where current OCR engines are less accurate. Our experimental analysis addresses six data sets containing documents ranging from books of the 17th century to contemporary journals.
Simone Marinai, Emanuele Marino, Giovanni Soda
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Layout based document image retrieval by means of XY tree reduction
abstract
We analyze a system for the retrieval of document images on the basis of layout similarity. Layout objects are extracted and represented with the XY tree. Page similarity is computed with a tree-edit distance algorithm. The peculiarity of the approach is the use of tree grammars to model the variations in the tree, which are due to segmentation algorithms or to structural differences between documents with similar layout. A few class-independent grammatical rules are used to modify each tree and obtain a reduced tree that is supposed to preserve the most relevant features of the page.
Simone Marinai, Emanuele Marino, Giovanni Soda
ICDAR1
2005 Artificial Neural Networks for Document Analysis and Recognition
abstract
Artificial neural networks have been extensively applied to document analysis and recognition. Most efforts have been devoted to the recognition of isolated handwritten and printed characters with widely recognized successful results. However, many other document processing tasks, like preprocessing, layout analysis, character segmentation, word recognition, and signature verification, have been effectively faced with very promising results. This paper surveys the most significant problems in the area of offline document image processing, where connectionist-based approaches have been applied. Similarities and differences between approaches belonging to different categories are discussed. A particular emphasis is given on the crucial role of prior knowledge for the conception of both appropriate architectures and learning algorithms. Finally, the paper provides a critical analysis on the reviewed approaches and depicts the most promising research guidelines in the field. In particular, a second generation of connectionist-based models are foreseen which are based on appropriate graphical representations of the learning environment.
Simone Marinai, Marco Gori, Giovanni Soda
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Preface
Simone Marinai, Marco Gori
Pattern Recognit. Lett.1
2003 Using tree-grammars for training set expansion in page classification
abstract
In this paper we describe a method for the expansionof training sets made by XY trees representing page layout.This approach is appropriate when dealing with page classificationbased on MXY tree page representations. The basicidea is the use of tree grammars to model the variationsin the tree which are caused by segmentation algorithms.A set of general grammatical rules are defined and used toexpand the training set. Pages are classified with a k - nnapproach where the distance between pages is computed bymeans of tree-edit distance.
Stefano Baldi, Simone Marinai, Giovanni Soda
ICDAR2
2003 Indexing and retrieval of words in old documents
abstract
This paper describes a system for efficient indexing and retrieval of words in collections of document images. The proposed method is based on two main principles: unsupervised prototype clustering, and string encoding for efficient string matching. During indexing, a self organizing map (SOM) is trained so as to cluster together similar symbols (character-like objects) in a sub-set of the documents to be stored. By using the trained SOM the words in the whole collection can be stored and represented with a fixed-length description that can be easily compared in order to score most similar words in response to a user query. The system can be automatically adapted to different languages and font styles. The most appropriate applications are for the processing of old documents (18th and 19th Centuries) where current OCRs have more difficulties. Experimental results describe three application scenarios having various levels of difficulty for current OCR systems.
Simone Marinai, Emanuele Marino, Giovanni Soda
ICDAR1
2003 Edge-backpropagation for noisy logo recognition
Marco Gori, Marco Maggini, Simone Marinai, Jianqing Sheng, Giovanni Soda
Pattern Recognit.3
2002 Retrieval by Layout Similarity of Documents Represented with MXY Trees
Francesca Cesarini, Simone Marinai, Giovanni Soda
Document Analysis Systems2
2001 Page Classification for Meta-data Extraction from Digital Collections
Francesca Cesarini, Marco Lastri, Simone Marinai, Giovanni Soda
DEXA3
2001 Encoding of Modified X-Y Trees for Document Classification
abstract
Describes a method for classifying document images on the basis of their physical layout. The layout is described by means of a hierarchical description, the modified X-Y tree, that is derived from the classical X-Y tree segmentation algorithm taking into account cuts along lines in addition to cuts along white spaces between blocks. In order to reduce problems due to noise and the skew of the input image, the modified X-Y tree is built on top of regions extracted by a commercial OCR package. The tree is afterwards coded into a fixed-size representation that takes into account occurrences of tree patterns in the tree representing the page. Lastly, this feature vector is fed to an artificial neural network that is trained to classify document images. The system is applied to the classification of documents belonging to digital libraries. Examples of classes taken into account are "title page", "index" and "regular page". Many tests have been carried out on a data set of more than 600 pages from an online digital library. These tests allowed us to conclude that the use of modified X-Y trees is advantageous with respect to the classical X-Y decomposition for this classification task.
Francesca Cesarini, Marco Lastri, Simone Marinai, Giovanni Soda
ICDAR3
2001 Automatic document classification and indexing in high-volume applications
Enrico Appiani, Francesca Cesarini, Anna Maria Colla, Michelangelo Diligenti, Marco Gori, Simone Marinai, Giovanni Soda
Int. J. Document Anal. Recognit.6
2001 A serial combination of connectionist-based classifiers for OCR
Enrico Francesconi, Marco Gori, Simone Marinai, Giovanni Soda
Int. J. Document Anal. Recognit.3
1999 Structured Document Segmentation and Representation by the Modified X-Y tree
abstract
We describe a top-down approach to the segmentation and representation of documents containing tabular structures. Examples of these documents are invoices and technical papers with tables. The segmentation is based on an extension of X-Y trees, where the regions are split by means of cuts along separators (e.g. lines), in addition to cuts along white spaces. The leaves describe regions containing homogeneous information and cutting separators. Adjacency links among leaves of the tree describe local relationships between corresponding regions.
Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda
ICDAR3
1999 Projection based Segmentation of Musical Sheets
abstract
The automatic recognition of music scores is a key process for the electronic treatment of music information. In this paper we present the segmentation module of an OMR system. The proposed approach is based on the use of projection profiles for the location of elementary symbols that constitute the music notation. An extensive experimentation was made which the help of a tool developed to this purpose. Reported results shown a high efficiency in the correct location of elementary symbols.
Simone Marinai, Paolo Nesi
ICDAR1
1998 INFORMys: A Flexible Invoice-Like Form-Reader System
abstract
We describe a flexible form-reader system capable of extracting textual information from accounting documents, like invoices and bills of service companies. In this kind of document, the extraction of some information fields cannot take place without having detected the corresponding instruction fields, which are only constrained to range in given domains. We propose modeling the document's layout by means of attributed relational graphs, which turn out to be very effective for form registration, as well as for performing a focused search for instruction fields. This search is carried out by means of a hybrid model, where proper algorithms, based on morphological operations and connected components, are integrated with connectionist models. Experimental results are given in order to assess the actual performance of the system.
Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda
IEEE Trans. Pattern Anal. Mach. Intell.3
1997 A Neural-Based Architecture for Spot-Noisy Logo Recognition
abstract
Much attention has recently been paid to the recognition of graphical objects, such as company logos and trademarks. Recognizing these objects facilitates the recognition of document classes. Some promising results have been achieved by using autoassociator-based artificial neural networks (AANN) in the presence of homogeneously distributed noise. However, the performance drops significantly when dealing with spot-noisy logos, where strips or blobs produce a partial obstruction of the pictures. We propose a new approach for training AANNs especially conceived for dealing with spot noise. The basic idea is to introduce new metrics for assessing the reproduction error in AANNs. The proposed algorithm, referred to as spot-backpropagation (S-BP), is significantly more robust with respect to spot-noise than classical Euclidean norm-based backpropagation (BP). Our experimental results are based on a database of 88 real logos that are artificially corrupted by spot-noise.
Francesca Cesarini, Enrico Francesconi, Marco Gori, Simone Marinai, Jianqing Sheng, Giovanni Soda
ICDAR4
1997 Rectangle Labelling for an Invoice Understanding System
abstract
We present a method for the logical labelling of physical rectangles, extracted from invoices, based on a conceptual model which describes, as generally as possible, the invoice universe. This general knowledge is used in the semi automatic construction of a model for each class of invoices. Once the model is constructed, it can be applied to understand an invoice instance, whose class is univocally identified by its logo. This approach is used to design a flexible system which is able to learn, from a nucleus of general knowledge, a monotonic set of specific knowledge for each class of invoices (document models), in terms of physical coordinates for each rectangle and related semantic label.
Francesca Cesarini, Enrico Francesconi, Marco Gori, Simone Marinai, Jianqing Sheng, Giovanni Soda
ICDAR4
1995 Data Extraction from Form Images
Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda
DEXA3
1995 A system for data extraction from forms of known class
abstract
In this paper, we describe a flexible and efficient system for processing forms of a known class. The model is based on attributed relational graphs and the system performs form registration and location of information fields using algorithms based on the hypothesize-and-verify paradigm. A special emphasis has been placed at the low level, where an autoassociator-based connectionist model has exhibited successful results in finding the instruction fields in very noisy forms.
Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda
ICDAR3