Simone Marinai

dblp:66/5842 · DBLP profile ↗
← Back
30ranked-venue papers in the field
12as first author
4since 2021 · last 2026
0000-0002-6702-2277ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 22 (9 first)Information Retrieval & Web Search · 6 (3 first)Database Systems & Data Management · 2
YearPublicationVenuePosition
2026 Towards Non-Latin Text and Layout Personalization for Enhanced Readability
Rina Buoy, Dylan Berkamp Fouepe Dongmo, Vesal Khean, Simone Marinai, Koichi Kise
ICDAR (3)4
2025 Visual Large Language Models for Graphics Understanding: A Case Study on Floorplan Images
abstract
This study explores the use of Vision Large Language Models (VLLMs) for identifying items in complex graphical documents. In particular, we focus on looking for furniture objects (e.g. beds, tables, and chairs) and structural items (doors and windows) in floorplan images. We evaluate one object detection model (YOLO) and state-of-the-art VLLMs on two datasets featuring diverse floorplan layouts and symbols. The experiments with VLLMs are performed with a zero-shot setting, meaning the models are tested without any training or fine-tuning, as well as with a few-shot approach, where examples of items to be found in the image are given to the models in the prompt. The results highlight the strengths and limitations of VLLMs in recognizing architectural elements, providing guidance for future research in the use multimodal vision-language models for graphics recognition.
Valeria Nardoni, Kimiya Noor Ali, Zahra Ziran, Simone Marinai
DocEng4
2025 Graph Convolutional Teacher-Student Framework for Writer Inspection from Intra-variable Handwritten Words
Kumari Priya, Aritra Dey, Chandranath Adak, Soumi Chattopadhyay, Sukalpa Chanda, Simone Marinai
ICDAR (3)7
2023 Deep-learning for dysgraphia detection in children handwritings
abstract
Early identification of dysgraphia in children is crucial for timely intervention and support. Traditional methods, such as the Brave Handwriting Kinder (BHK) test, which relies on manual scoring of handwritten sentences, are both time-consuming and subjective posing challenges in accurate and efficient diagnosis. In this paper, an approach for dysgraphia detection by leveraging smart pens and deep learning techniques is proposed, automatically extracting visual features from children's handwriting samples. To validate the solution, samples of children handwritings have been gathered and several interviews with domain experts have been conducted. The approach has been compared with an algorithmic version of the BHK test and with several elementary school teachers' interviews.
Andrea Gemelli, Simone Marinai, Emanuele Vivoli, Tamara Zappaterra
DocEng2
2018 Offline Bengali Writer Verification by PDF-CNN and Siamese Net
abstract
Automated handwriting analysis is a popular area of research owing to the variation of writing patterns. In this research area, writer verification is one of the most challenging branches, having direct impact on biometrics and forensics. In this paper, we deal with offline writer verification on complex handwriting patterns. Therefore, we choose a relatively complex script, i.e., Indic Abugida script Bengali (or, Bangla) containing more than 250 compound characters. From a handwritten sample, the probability distribution functions (PDFs) of some handcrafted features are obtained and input to a convolutional neural network (CNN). For such a CNN architecture, we coin the term "PDFCNN", where handcrafted feature PDFs are hybridized with auto-derived CNN features. Such hybrid features are then fed into a Siamese neural network for writer verification. The experiments are performed on a Bengali offline handwritten dataset of 100 writers. Our system achieves encouraging results, which sometimes exceed the results of state-of-the-art techniques on writer verification.
Chandranath Adak, Simone Marinai, Bidyut B. Chaudhuri, Michael Blumenstein
DAS2
2017 DocEmul: A Toolkit to Generate Structured Historical Documents
abstract
We propose a toolkit to generate structured synthetic documents emulating the actual document production process. Synthetic documents can be used to train systems to perform document analysis tasks. In our case we address the record counting task on handwritten structured collections containing a limited number of examples. Using the DocEmul toolkit we can generate a larger dataset to train a deep architecture to predict the number of records for each page. The toolkit is able to generate synthetic collections and also perform data augmentation to create a larger trainable dataset. It includes one method to extract the page background from real pages which can be used as a substrate where records can be written on the basis of variable structures and using cursive fonts. Moreover, it is possible to extend the synthetic collection by adding random noise, page rotations, and other visual variations. We performed some experiments on two different handwritten collections using the toolkit to generate synthetic data to train a Convolutional Neural Network able to count the number of records in the real collections.
Samuele Capobianco, Simone Marinai
ICDAR2
2017 Partitioning Open Plan Areas in Floor Plans
abstract
We are developing an application to automatically generate an accessible graphic from a floor plan image. Floor plans generally contain large regions with functionally different sub-areas. A problem faced by visually impaired users in exploring such accessible floor plans is understanding the boundaries of these sub-areas. We present an effective method to partition such open plan areas. Initially, we conducted a formative user study to understand how people partition open plan areas. Based on the findings of the study, we identified a general set of guidelines for partitioning open plans. These guidelines were used to generate a set of candidate lines for sub-areas. An obstacle avoiding shortest-path Voronoi diagram was used to determine boundaries for each sub-area. Candidate lines such as wall extensions were automatically generated to replace the identified boundaries. We selected the best replacement for each boundary by scoring candidate lines using a set of criteria such as line length. Finally the proposed method was tested on a standard floor plan corpus using three novel measures.
Anuradha Madugalla, Kim Marriott, Simone Marinai
ICDAR3
2015 Deepdocclassifier: Document classification with deep Convolutional Neural Network
abstract
This paper presents a deep Convolutional Neural Network (CNN) based approach for document image classification. One of the main requirement of deep CNN architecture is that they need huge number of samples for training. To overcome this problem we adopt a deep CNN which is trained using big image dataset containing millions of samples i.e., ImageNet. The proposed work outperforms both the traditional structure similarity methods and the CNN based approaches proposed earlier. The accuracy of the proposed approach with merely 20 images per class outperforms the state-of-the-art by achieving classification accuracy of 68.25%. The best results on Tobbacoo-3428 dataset show that our proposed method outperforms the state-of-the-art method by a significant margin and achieved a median accuracy of 77.6% with 100 samples per class used for training and validation.
Muhammad Zeshan Afzal, Samuele Capobianco, Muhammad Imran Malik, Simone Marinai, Thomas M. Breuel, Andreas Dengel 0001, Marcus Liwicki
ICDAR4
2015 Accessible On-Line Floor Plans
abstract
Better access to on-line information graphics is a pressing need for people who are blind or have severe vision impairment. We present a new model for accessible presentation of on-line information graphics and demonstrate its use for presenting floor plans. While floor plans are increasingly provided on-line, people who are blind are at best provided with only a high-level textual description. This makes it difficult for them to understand the spatial arrangement of the objects on the floor plan. Our new approach provides users with significantly better access to such plans. The users can automatically generate an accessible version of a floor plan from an on-line floor plan image quickly and independently by using a web service. This generates a simplified graphic showing the rooms, walls, doors and windows in the original floor plan as well as a textual overview. The accessible floor plan is presented on an iPad using audio feedback. As the users touch graphic elements on the screen, the element they are touching is described by speech and non-speech audio in order to help them navigate the graphic.
Cagatay Goncu, Anuradha Madugalla, Simone Marinai, Kim Marriott
WWW3
2013 Reflowing and annotating scientific papers on eBook readers
abstract
Working with scientific and technical papers on small screen devices, such as tablets and eBook readers, is difficult since these works are often typeset in multiple columns with a relatively small font size.
Simone Marinai
ACM Symposium on Document Engineering1
2012 Displaying chemical structural formulae in ePub format
abstract
We describe one tool designed to enhance the visualization of chemical structural formulae in E-book readers. When dealing with small formulae, to avoid the pixelation effect with zoomed images, the formula is converted to a vectoral representation and then enlarged. On the opposite, large formulae are split in sub-images by cutting the image in suitable locations attempting to reduce the parts of the formula that are broken. In both cases the formulae are embedded in one ePub document that allows users to browse the chemical structure on most reading devices.
Simone Marinai, Stefano Quiriconi
ACM Symposium on Document Engineering1
2011 Conversion of PDF Books in ePub Format
abstract
In the last years the interest in e-book readers is significantly growing. Two main document formats are supported by most devices: PDF and ePub. The PDF format is widely used to share documents allowing a cross-platform readability. However, it is not ideal for a comfortable reading on small screens. On the opposite, the ePub format is re-flowable and it is well suited for e-book readers. In this paper we describe a system for the conversion of PDF books to the ePub format aiming at inverting the text formatting made during the pagination. To this purpose, layout analysis techniques are performed to identify the book's table of contents and the main functional regions such as chapters, paragraphs, and notes.
Simone Marinai, Emanuele Marino, Giovanni Soda
ICDAR1
2011 Using Earth Mover's Distance in the Bag-of-Visual-Words Model for Mathematical Symbol Retrieval
abstract
In this paper, the Earth Mover's Distance (EMD) is used as a similarity measure in the mathematical symbol retrieval task. The approach is based on the Bag-of-Visual-Words model. In our case the features extracted from each symbol are clustered by means of Self-Organizing Maps (SOM) and then occurrences of features in the clusters are accumulated in a vector of visual words. The comparison between the latter vectors is performed with the EMD which naturally allows to incorporate the topological organization of SOM clusters in the distance computation. The proposed approach is experimentally tested in a mathematical symbol retrieval task and compared with the cosine similarity and with some variants that have been recently proposed.
Simone Marinai, Beatrice Miotti, Giovanni Soda
ICDAR1
2010 Table of contents recognition for converting PDF documents in e-book formats
abstract
We describe one tool for Table of Content (ToC) identification and recognition from PDF books. This task is part of ongoing research on the development of tools for the semi-automatic conversion of PDF documents in the Epub format that can be read on several E-book devices. Among various sub-tasks, the ToC extraction and recognition is particularly useful for an easy navigation of book contents.
Simone Marinai, Emanuele Marino, Giovanni Soda
ACM Symposium on Document Engineering1
2009 Metadata Extraction from PDF Papers for Digital Library Ingest
abstract
In this paper we analyze our recent research on the use of document analysis techniques for metadata extraction from PDF papers. We describe a package that is designed to extract basic metadata from these documents. The package is used in combination with a digital library software suite to easily build personal digital libraries. The proposed software is based on a suitable combination of several techniques that include PDF parsing, low level document image processing, and layout analysis. In addition, we use the information gathered from a widely known citation database (DBLP) to assist the tool in the difficult task of author identification. The system is tested on some paper collections selected from recent conference proceedings.
Simone Marinai
ICDAR1
2009 Mathematical Symbol Indexing Using Topologically Ordered Clusters of Shape Contexts
abstract
This paper addresses the indexing and retrieval of mathematical symbols from digitized documents. The proposed approach exploits Shape Contexts (SC) to describe the shape of mathematical symbols. Starting from the vector space method, that is based on SC clustering, we explore the use of topological ordered clusters to improve the retrieval performance. The clustering is computed by means of Self-Organizing Maps that organize the clusters in two dimensional topologically ordered feature maps. The retrieval performance are compared with those obtained using the K-means clustering on a large collection of mathematical symbols gathered from the widely used INFTY database.
Simone Marinai, Beatrice Miotti, Giovanni Soda
ICDAR1
2008 A Comparison of Clustering Methods for Word Image Indexing
abstract
In this paper we compare three clustering methods used to perform word image indexing. The three methods are: the Self-Organizing Map (SOM), the Growing Hierarchical Self-Organizing Map (GHSOM), and the Spectral Clustering. We test these methods on a real data set composed of word images extracted from an encyclopedia of the XIX-th Century. The word images are grouped on the basis of the clustering methods and subsequently retrieved identifying the closest clusters to a query word. The accuracy of the methods is compared evaluating the performance of the word retrieval algorithm. From the experimental results we conclude that methods designed to automatically determine the number and the structure of clusters, such as GHSOM, are particularly suitable in the context represented by our data set.
Simone Marinai, Emanuele Marino, Giovanni Soda
Document Analysis Systems1
2006 Efficient Word Retrieval by Means of SOM Clustering and PCA
Simone Marinai, Stefano Faini, Emanuele Marino, Giovanni Soda
Document Analysis Systems1
2005 Layout based document image retrieval by means of XY tree reduction
abstract
We analyze a system for the retrieval of document images on the basis of layout similarity. Layout objects are extracted and represented with the XY tree. Page similarity is computed with a tree-edit distance algorithm. The peculiarity of the approach is the use of tree grammars to model the variations in the tree, which are due to segmentation algorithms or to structural differences between documents with similar layout. A few class-independent grammatical rules are used to modify each tree and obtain a reduced tree that is supposed to preserve the most relevant features of the page.
Simone Marinai, Emanuele Marino, Giovanni Soda
ICDAR1
2003 Using tree-grammars for training set expansion in page classification
abstract
In this paper we describe a method for the expansionof training sets made by XY trees representing page layout.This approach is appropriate when dealing with page classificationbased on MXY tree page representations. The basicidea is the use of tree grammars to model the variationsin the tree which are caused by segmentation algorithms.A set of general grammatical rules are defined and used toexpand the training set. Pages are classified with a k - nnapproach where the distance between pages is computed bymeans of tree-edit distance.
Stefano Baldi, Simone Marinai, Giovanni Soda
ICDAR2
2003 Indexing and retrieval of words in old documents
abstract
This paper describes a system for efficient indexing and retrieval of words in collections of document images. The proposed method is based on two main principles: unsupervised prototype clustering, and string encoding for efficient string matching. During indexing, a self organizing map (SOM) is trained so as to cluster together similar symbols (character-like objects) in a sub-set of the documents to be stored. By using the trained SOM the words in the whole collection can be stored and represented with a fixed-length description that can be easily compared in order to score most similar words in response to a user query. The system can be automatically adapted to different languages and font styles. The most appropriate applications are for the processing of old documents (18th and 19th Centuries) where current OCRs have more difficulties. Experimental results describe three application scenarios having various levels of difficulty for current OCR systems.
Simone Marinai, Emanuele Marino, Giovanni Soda
ICDAR1
2002 Retrieval by Layout Similarity of Documents Represented with MXY Trees
Francesca Cesarini, Simone Marinai, Giovanni Soda
Document Analysis Systems2
2001 Page Classification for Meta-data Extraction from Digital Collections
Francesca Cesarini, Marco Lastri, Simone Marinai, Giovanni Soda
DEXA3
2001 Encoding of Modified X-Y Trees for Document Classification
abstract
Describes a method for classifying document images on the basis of their physical layout. The layout is described by means of a hierarchical description, the modified X-Y tree, that is derived from the classical X-Y tree segmentation algorithm taking into account cuts along lines in addition to cuts along white spaces between blocks. In order to reduce problems due to noise and the skew of the input image, the modified X-Y tree is built on top of regions extracted by a commercial OCR package. The tree is afterwards coded into a fixed-size representation that takes into account occurrences of tree patterns in the tree representing the page. Lastly, this feature vector is fed to an artificial neural network that is trained to classify document images. The system is applied to the classification of documents belonging to digital libraries. Examples of classes taken into account are "title page", "index" and "regular page". Many tests have been carried out on a data set of more than 600 pages from an online digital library. These tests allowed us to conclude that the use of modified X-Y trees is advantageous with respect to the classical X-Y decomposition for this classification task.
Francesca Cesarini, Marco Lastri, Simone Marinai, Giovanni Soda
ICDAR3
1999 Structured Document Segmentation and Representation by the Modified X-Y tree
abstract
We describe a top-down approach to the segmentation and representation of documents containing tabular structures. Examples of these documents are invoices and technical papers with tables. The segmentation is based on an extension of X-Y trees, where the regions are split by means of cuts along separators (e.g. lines), in addition to cuts along white spaces. The leaves describe regions containing homogeneous information and cutting separators. Adjacency links among leaves of the tree describe local relationships between corresponding regions.
Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda
ICDAR3
1999 Projection based Segmentation of Musical Sheets
abstract
The automatic recognition of music scores is a key process for the electronic treatment of music information. In this paper we present the segmentation module of an OMR system. The proposed approach is based on the use of projection profiles for the location of elementary symbols that constitute the music notation. An extensive experimentation was made which the help of a tool developed to this purpose. Reported results shown a high efficiency in the correct location of elementary symbols.
Simone Marinai, Paolo Nesi
ICDAR1
1997 A Neural-Based Architecture for Spot-Noisy Logo Recognition
abstract
Much attention has recently been paid to the recognition of graphical objects, such as company logos and trademarks. Recognizing these objects facilitates the recognition of document classes. Some promising results have been achieved by using autoassociator-based artificial neural networks (AANN) in the presence of homogeneously distributed noise. However, the performance drops significantly when dealing with spot-noisy logos, where strips or blobs produce a partial obstruction of the pictures. We propose a new approach for training AANNs especially conceived for dealing with spot noise. The basic idea is to introduce new metrics for assessing the reproduction error in AANNs. The proposed algorithm, referred to as spot-backpropagation (S-BP), is significantly more robust with respect to spot-noise than classical Euclidean norm-based backpropagation (BP). Our experimental results are based on a database of 88 real logos that are artificially corrupted by spot-noise.
Francesca Cesarini, Enrico Francesconi, Marco Gori, Simone Marinai, Jianqing Sheng, Giovanni Soda
ICDAR4
1997 Rectangle Labelling for an Invoice Understanding System
abstract
We present a method for the logical labelling of physical rectangles, extracted from invoices, based on a conceptual model which describes, as generally as possible, the invoice universe. This general knowledge is used in the semi automatic construction of a model for each class of invoices. Once the model is constructed, it can be applied to understand an invoice instance, whose class is univocally identified by its logo. This approach is used to design a flexible system which is able to learn, from a nucleus of general knowledge, a monotonic set of specific knowledge for each class of invoices (document models), in terms of physical coordinates for each rectangle and related semantic label.
Francesca Cesarini, Enrico Francesconi, Marco Gori, Simone Marinai, Jianqing Sheng, Giovanni Soda
ICDAR4
1995 Data Extraction from Form Images
Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda
DEXA3
1995 A system for data extraction from forms of known class
abstract
In this paper, we describe a flexible and efficient system for processing forms of a known class. The model is based on attributed relational graphs and the system performs form registration and location of information fields using algorithms based on the hypothesize-and-verify paradigm. A special emphasis has been placed at the low level, where an autoassociator-based connectionist model has exhibited successful results in finding the instruction fields in very noisy forms.
Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda
ICDAR3