VLDB 2026 Research / reviewers in the wild / expert
George Retsinas
dblp:171/5669 · also Georgios Retsinas
· DBLP profile ↗
41ranked-venue papers
20as first author
26since 2021 · last 2025
0000-0001-6734-3575ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 10 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 9 first-author · 10 since 2021Databases, data management, data science and information retrieval · 16 · 10 first-author · 9 since 2021Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Open-Ended Robotic Exploration Using Vision-Inspired Similarity and Foundation ModelsabstractIn the domain of robotics, achieving Lifelong Open-ended Learning Autonomy (LOLA) represents a significant milestone, especially in contexts where autonomous agents must adapt to unforeseen environmental variations and evolving objectives. This paper introduces VISOR (VisionSimilarity for Open-ended Robotic exploration), a vision-based framework designed to assist robotic agents in autonomously exploring and learning from new environments and objects, whether through guided or random exploration, without reliance on predefined design considerations. In that direction, VISOR acts as a perception mediator, classifying everything a robot encounters in a scene as either known or unknown. It further identifies potential distractors (e.g., background elements), known categories, or objects specified through text seeds. By leveraging recent advancements in vision foundation models, VISOR operates in a training-free manner. It begins by segmenting a scene into its constituent entities, regardless of familiarity, and then extracts robust visual representations for each one. These representations are compared against an adaptive memory system that evolves over time; unknown objects are assigned unique IDs and added to this memory as new classes, enriching the robot's understanding of its environment. We argue that this evolving memory can facilitate guided exploration through prior knowledge, enhancing the efficiency of robotic exploration, and validate this by designing two exploration scenarios and running both simulated and real-world experiments. Panagiotis Paraskevas Filntisis, Efthymios Tsaprazlis, Paraskevas Oikonomou, Francesco Mattioli 0003, Vieri G. Santucci, George Retsinas, Petros Maragos |
ICRA | 6 |
| 2025 | Proactive Tactile Exploration for Object-Agnostic Shape Reconstruction from Minimal Visual PriorsabstractThe perception of an object's surface is important for robotic applications enabling robust object manipulation. The level of accuracy in such a representation affects the outcome of the action planning, especially during tasks that require physical contact, e.g. grasping. In this paper, we propose a novel iterative method for 3D shape reconstruction consisting of two steps. At first, a mesh is fitted on data points acquired from the object's surface, based on a single primitive template. Subsequently, the mesh is properly adjusted to adequately represent local deformities. Moreover, a novel proactive tactile exploration strategy aims at minimizing the total uncertainty with the least number of contacts, while reducing the risk of contact failure in case the estimated surface differs significantly from the real one. The performance of the methodology is evaluated both in 3D simulation and on a real setup. Paris Oikonomou, George Retsinas, Petros Maragos, Costas S. Tzafestas |
ICRA | 2 |
| 2025 | Category-Level 6D Object Pose Estimation in Agricultural Settings Using a Lattice-Deformation Framework and Diffusion-Augmented Synthetic DataabstractAccurate 6D object pose estimation is essential for robotic grasping and manipulation, particularly in agriculture, where fruits and vegetables exhibit high intra-class variability in shape, size, and texture. The vast majority of existing methods rely on instance-specific CAD models or require depth sensors to resolve geometric ambiguities, making them impractical for real-world agricultural applications. In this work, we introduce PLANTPose, a novel framework for category-level 6D pose estimation that operates purely on RGB input. PLANT-Pose predicts both the 6D pose and deformation parameters relative to a base mesh, allowing a single category-level CAD model to adapt to unseen instances. This enables accurate pose estimation across varying shapes without relying on instance-specific data. To enhance realism and improve generalization, we also leverage Stable Diffusion to refine synthetic training images with realistic texturing, mimicking variations due to ripeness and environmental factors and bridging the domain gap between synthetic data and the real world. Our evaluations on a challenging benchmark that includes bananas of various shapes, sizes, and ripeness status demonstrate the effectiveness of our framework in handling large intraclass variations while maintaining accurate 6D pose predictions, significantly outperforming the state-of-the-art RGB-based approach MegaPose. Our code, data, and models are publicly available at https://github.com/mariosgly/PLANTPose. Marios Glytsos, Panagiotis Paraskevas Filntisis, George Retsinas, Petros Maragos |
IROS | 3 |
| 2025 | Instance-Level Composed Image RetrievalabstractThe progress of composed image retrieval (CIR), a popular research direction in image retrieval, where a combined visual and textual query is used, is held back by the absence of high-quality training and evaluation data. We introduce a new evaluation dataset, i-CIR, which, unlike existing datasets, focuses on an instance-level class definition. The goal is to retrieve images that contain the same particular object as the visual query, presented under a variety of modifications defined by textual queries. Its design and curation process keep the dataset compact to facilitate future research, while maintaining its challenge—comparable to retrieval among more than 40M random distractors—through a semi-automated selection of hard negatives. To overcome the challenge of obtaining clean, diverse, and suitable training data, we leverage pre-trained vision-and-language models (VLMs) in a training-free approach called BASIC. The method separately estimates query-image-to-image and query-text-to-image similarities, performing late fusion to upweight images that satisfy both queries, while down-weighting those that exhibit high similarity with only one of the two. Each individual similarity is further improved by a set of components that are simple and intuitive. BASIC sets a new state of the art on i-CIR but also on existing CIR datasets that follow a semantic-level class definition. Project page: https://vrg.fel.cvut.cz/icir/. Bill Psomas, George Retsinas, Nikos Efthymiadis, Panagiotis Paraskevas Filntisis, Yannis Avrithis, Petros Maragos, Ondrej Chum, Giorgos Tolias |
NeurIPS | 2 |
| 2025 | Pre-training for Action Recognition with Automatically Generated Fractal Datasets
Davyd Svyezhentsev, George Retsinas, Petros Maragos |
Int. J. Comput. Vis. | 2 |
| 2024 | 3D Facial Expressions through Analysis-by-Neural-SynthesisabstractWhile existing methods for 3D face reconstruction from in-the-wild images excel at recovering the overall face shape, they commonly miss subtle, extreme, asymmetric, or rarely observed expressions. We improve upon these meth-ods with SMIRK (Spatial Modeling for Image-based Reconstruction of Kinesics), which faithfully reconstructs expres-sive 3D faces from images. We identify two key limitations in existing methods: shortcomings in their self-supervised training formulation, and a lack of expression diversity in the training images. For training, most methods employ differentiable rendering to compare a predicted face mesh with the input image, along with a plethora of additional loss functions. This differentiable rendering loss not only has to provide supervision to optimize for 3D face geom-etry, camera, albedo, and lighting, which is an ill-posed optimization problem, but the domain gap between ren-dering and input image further hinders the learning pro-cess. Instead, SMIRK replaces the differentiable rendering with a neural rendering module that, given the ren-dered predicted mesh geometry, and sparsely sampled pix-els of the input image, generates a face image. As the neural rendering gets color information from sampled im-age pixels, supervising with neural rendering-based reconstruction loss can focus solely on the geometry. Further it enables us to generate images of the input identity with varying expressions while training. These are then utilized as input to the reconstruction model and used as supervision with ground truth geometry. This effectively augments the training data and enhances the generalization for di-verse expressions. Our qualitative, quantitative and partic-ularly our perceptual evaluations demonstrate that SMIRK achieves the new state-of-the art performance on accurate expression reconstruction. For our method's source code, demo video and more, please visit our project webpage: https://georgeretsi.github.io/smirk/. George Retsinas, Panagiotis Paraskevas Filntisis, Radek Danecek, Victoria Fernández Abrevaya, Anastasios Roussos, Timo Bolkart, Petros Maragos |
CVPR | 1 |
| 2024 | Bessarion: Medieval Greek Inscriptions on a Challenging Dataset for Vision and NLP Tasks
Giorgos Sfikas, Panagiotis Dimitrakopoulos, George Retsinas, Christophoros Nikou, Pinelopi Kitsiou |
DAS | 3 |
| 2024 | DiffusionPen: Towards Controlling the Style of Handwritten Text Generation
Konstantina Nikolaidou, George Retsinas, Giorgos Sfikas, Marcus Liwicki |
ECCV (85) | 2 |
| 2024 | Augmenting Transformer Autoencoders with Phenotype Classification for Robust Detection of Psychotic RelapsesabstractRecently, deep autoencoder architectures have received attention for the problem of unsupervised anomaly detection. Detecting psychotic relapses in mental health patients is a crucial challenge, often framed as anomaly detection, given the limited availability of data during relapsing states. In this paper, motivated by the fact that during relapses patients tend to undergo behavioral changes, we augment the classical autoencoder architecture with extra patient identification components. We show that formulating the problem as one of both signal reconstruction and patient identification largely improves the overall precision and robustness of relapse detection and significantly outperforms previous methods with a relative improvement of 15%. In addition, we also explore multiple ways to fuse the identification and reconstruction errors into a unified anomaly score that outperforms the results achieved by each error in isolation. Niki Efthymiou, George Retsinas, Panagiotis Paraskevas Filntisis, Petros Maragos |
ICASSP | 2 |
| 2024 | Matrix Factorization in Tropical and Mixed Tropical-Linear AlgebrasabstractMatrix Factorization (MF) has found numerous applications in Machine Learning and Data Mining, including collaborative filtering recommendation systems, dimensionality reduction, data visualization, and community detection. Motivated by the recent successes of tropical algebra and geometry in machine learning, we investigate two problems involving matrix factorization over the tropical algebra. For the first problem, Tropical Matrix Factorization (TMF), which has been studied already in the literature, we propose an improved algorithm that avoids many of the local optima. The second formulation considers the approximate decomposition of a given matrix into the product of three matrices where a usual matrix product is followed by a tropical product. This formulation has a very interesting interpretation in terms of the learning of the utility functions of multiple users. We also present numerical results illustrating the effectiveness of the proposed algorithms, as well as an application to recommendation systems with promising results. Ioannis Kordonis, Emmanouil Theodosis, George Retsinas, Petros Maragos |
ICASSP | 3 |
| 2024 | Enhancing CRNN HTR Architectures with Transformer Blocks
George Retsinas, Konstantina Nikolaidou, Giorgos Sfikas |
ICDAR (4) | 1 |
| 2024 | Sparkle: Deep Learning Driven Autotuning for Taming High-Dimensionality of Spark DeploymentsabstractThe exponential growth of data in the Cloud has highlighted the need for more efficient data processing. In-Memory Computing frameworks (e.g., Spark) offer improved efficiency for large-scale data analytics, however, they also provide a plethora of configuration parameters that affect the resource consumption and performance of applications. Manually optimizing these parameters is a time-consuming process, due toi)the high-dimensional configuration space,ii)the complex inter-relationship between different parameters,iii)the diverse nature of workloads andiv)the inherent data heterogeneity. We introduceSparkle, an end-to-end deep learning-based framework for automating the performance modeling and tuning of Spark applications. We introduce a modular DNN architecture that expands to the entire Spark parameter configuration space and provides a universal performance modeling approach, completely eliminating the need for human or statistical reasoning. By employing a genetic optimization process,Sparklequickly traverses the design space and identifies highly optimized Spark configurations. Our experiments on the HiBench benchmark suite show thatSparkledelivers an average prediction accuracy of 93%, with high generalization capabilities, i.e.,$\approx 80\%$accuracy for unseen workloads, dataset sizes and configurations, outperforming state-of-art. Regarding end-to-end optimization,Sparkleefficiently explores Spark's high-dimensional parameter space, delivering new dominant Spark configurations, which correspond to 65% Pareto coverage w.r.t its Spark native optimization counterpart. Dimosthenis Masouros, George Retsinas, Sotirios Xydis, Dimitrios Soudris |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | Feather: An Elegant Solution to Effective DNN Sparsification
Athanasios Glentis Georgoulakis, George Retsinas, Petros Maragos |
BMVC | 2 |
| 2023 | Newton-Based Trainable Learning RateabstractSelecting an appropriate learning rate for efficiently training deep neural networks is a difficult process that can be affected by numerous parameters, such as the dataset, the model architecture or even the batch size. In this work, we propose an algorithm for automatically adjusting the learning rate during the training process, assuming a gradient descent formulation. The rationale behind our approach is to train the learning rate along with the model weights. Specifically, we formulate first and second-order gradients w.r.t. the learning rate as functions of consecutive weight gradients, leading to a cost-effective implementation. Our extensive experimental evaluation validates the effectiveness of the proposed method for a plethora of different settings. The proposed method has proven to be robust to both the initial learning rate and the batch size, making it ideal for an off-the-shelf optimizing scheme. George Retsinas, Giorgos Sfikas, Panagiotis Paraskevas Filntisis, Petros Maragos |
ICASSP | 1 |
| 2023 | E-Prevention: The ICASSP-2023 Challenge on Person Identification and Relapse Detection from Continuous Recordings of BiosignalsabstractThe e-Prevention challenge concerns the analysis and processing of long-term continuous recordings of biosignals recorded from wearable sensors, i.e., accelerometers, gyroscopes and heart rate monitors embedded in smartwatches, as well as sleep information and daily step count, in order to extract high-level representations of the wearer’s activity and behavior, termed as digital phenotypes. The ability of these digital phenotypes to quantify behavioral patterns and traits will be evaluated in two different tasks: 1) Person Identification, and 2) Relapse Detection in patients in the psychotic spectrum. The long-term data that will be used in this challenge have been acquired during the course of the e-Prevention project, an innovative integrated system for medical support that facilitates effective monitoring and relapse prevention in patients with mental disorders (i.e, schizophrenia and bipolar disorder). Specifically, the data were continuously collected from patients for a monitoring period of up to 2.5 years, while from the control subgroup for a period of 3 months, constituting one of the largest of its kind ever recorded. Athanasia Zlatintsi, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Christos Garoufis, George Retsinas, Thomas Sounapoglou, Ilias Maglogiannis, Panayiotis Tsanakas, Nikolaos Smyrnis, Petros Maragos |
ICASSP | 5 |
| 2023 | WordStylist: Styled Verbatim Handwritten Text Generation with Latent Diffusion Models
Konstantina Nikolaidou, George Retsinas, Vincent Christlein, Mathias Seuret, Giorgos Sfikas, Elisa H. Barney Smith, Hamam Mokayed, Marcus Liwicki |
ICDAR (2) | 2 |
| 2023 | Keyword Spotting Simplified: A Segmentation-Free Approach Using Character Counting and CTC Re-scoring
George Retsinas, Giorgos Sfikas, Christophoros Nikou |
ICDAR (1) | 1 |
| 2023 | Shared-Operation Hypercomplex Networks for Handwritten Text Recognition
Giorgos Sfikas, George Retsinas, Panagiotis Dimitrakopoulos, Basilios Gatos, Christophoros Nikou |
ICDAR (4) | 2 |
| 2022 | Best Practices for a Handwritten Text Recognition System
George Retsinas, Giorgos Sfikas, Basilios Gatos, Christophoros Nikou |
DAS | 1 |
| 2022 | On-the-Fly Deformations for Keyword Spotting
George Retsinas, Giorgos Sfikas, Basilios Gatos, Christophoros Nikou |
DAS | 1 |
| 2022 | Keyword Spotting with Quaternionic ResNet: Application to Spotting in Greek Manuscripts
Giorgos Sfikas, George Retsinas, Angelos P. Giotis, Basilios Gatos, Christophoros Nikou |
DAS | 2 |
| 2022 | Neural Network Approximation based on Hausdorff distance of Tropical Zonotopes
Panagiotis Misiakos, Georgios Smyrnis, George Retsinas, Petros Maragos |
ICLR | 3 |
| 2021 | From Seq2Seq Recognition to Handwritten Word Embeddings
George Retsinas, Giorgos Sfikas, Christophoros Nikou, Petros Maragos |
BMVC | 1 |
| 2021 | Iterative Weighted Transductive Learning for Handwriting Recognition
George Retsinas, Giorgos Sfikas, Christophoros Nikou |
ICDAR (4) | 1 |
| 2021 | Online Weight Pruning Via Adaptive Sparsity LossabstractPruning neural networks has regained interest in recent years as a means to compress state-of-the-art deep neural networks and enable their deployment on resource-constrained devices. In this paper, we propose a robust sparsity controlling framework that efficiently prunes network parameters during training with minimal computational overhead. We incorporate fast mechanisms to prune individual layers and build upon these to automatically prune the entire network under a user-defined budget constraint. Key to our end-to-end network pruning approach is the formulation of an intuitive and easy-to-implement adaptive sparsity loss used to explicitly control sparsity during training, enabling efficient budget-aware optimization. George Retsinas, Athena Elafrou, Georgios I. Goumas, Petros Maragos |
ICIP | 1 |
| 2021 | Deformation-Invariant Networks For Handwritten Text RecognitionabstractImage deformations under simple geometric restrictions are crucial for Handwriting Text Recognition (HTR), since different writing styles can be viewed as simple geometrical deformations of the same textual elements. In this respect, the usefulness of including deformation invariance to an HTR system is indisputable. We explore different existing strategies for ensuring deformation invariance, including spatial transformers and deformable convolutions, under the context of text recognition, as well as introduce a new deformation-based algorithm, inspired by adversarial learning, which aims to reduce character output uncertainty during evaluation time. The resulting HTR system is shown to achieve state-of-the-art performance on the IAM and RIMES datasets. George Retsinas, Giorgos Sfikas, Christophoros Nikou, Petros Maragos |
ICIP | 1 |
| 2020 | Person Identification Using Deep Convolutional Neural Networks on Short-Term Signals from Wearable SensorsabstractIn this work, we explore the discriminating ability of short-term signal patterns (e.g. few minutes long) with respect to the person identification task. We focus on signals recorded by simple wearable devices, such as smart watches, which can measure movements (accelerometer and gyroscope sensors) and biosignals (heart rate monitor). To address the person identification problem, we develop a deep neural network, based on one-dimensional convolutions, which receives raw signals from three different smartwatch sensors and predicts the person wearing the smartwatch. Experimental results indicate that even with signals from wearable sensors collected at intervals of only 10 minutes, different users can be identified with notably high accuracy, revealing the existence of distinct short-term patterns of movement and heart rate between different persons. George Retsinas, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Emmanouil Theodosis, Athanasia Zlatintsi, Petros Maragos |
ICASSP | 1 |
| 2020 | Maxpolynomial Division with Application To Neural Network SimplificationabstractIn this work, we further the link between neural networks with piecewise linear activations and tropical algebra. To that end, we introduce the process of Maxpolynomial Division, a geometric method which simulates division of polynomials in the max-plus semiring, while highlighting its key properties and noting its connection to neural networks. Afterwards, we generalize this method and apply it in the context of neural network minimization, for two-layer networks used for binary classification problems, attempting to reduce the size of the hidden layer before the output. A tractable method to find an appropriate divisor and perform the division is introduced and evaluated in the IMDB Movie Review and MNIST datasets, with preliminary experiments demonstrating a capacity of this method to reduce the size of the network, without major loss of performance. Georgios Smyrnis, Petros Maragos, George Retsinas |
ICASSP | 3 |
| 2020 | Enhancing Handwritten Text Recognition with N-gram sequence decomposition and Multitask LearningabstractCurrent state-of-the-art approaches in the field of Handwritten Text Recognition are predominately single task with unigram, character level target units. In our work, we utilize a Multi-task Learning scheme, training the model to perform decompositions of the target sequence with target units of different granularity, from fine to coarse. We consider this method as a way to utilize n-gram information, implicitly, in the training process, while the final recognition is performed using only the unigram output. Unigram decoding of such a multi-task approach highlights the capability of the learned internal representations, imposed by the different n-grams at the training step. We select n-grams as our target units and we experiment from unigrams till fourgrams, namely subword level granularities. These multiple decompositions are learned from the network with task-specific CTC losses. Concerning network architectures, we propose two alternatives, namely the Hierarchical and the Block Multi-task. Overall, our proposed model, even though evaluated only on the unigram task, outperforms its counterpart single-task by absolute 2.52% WER and 1.02% CER, in the greedy decoding, without any computational overhead during inference, hinting towards successfully imposing an implicit language model. Vasiliki Tassopoulou, George Retsinas, Petros Maragos |
ICPR | 2 |
| 2019 | RecNets: Channel-wise Recurrent Convolutional Neural Networks
George Retsinas, Athena Elafrou, Georgios I. Goumas, Petros Maragos |
BMVC | 1 |
| 2019 | An Alternative Deep Feature Approach to Line Level Keyword SpottingabstractKeyword spotting (KWS) is defined as the problem of detecting all instances of a given word, provided by the user either as a query word image (Query-by-Example, QbE) or a query word string (Query-by-String, QbS) in a body of digitized documents. Keyword detection is typically preceded by a preprocessing step where the text is segmented into text lines (line-level KWS). Methods following this paradigm are monopolized by test-time computationally expensive handwritten text recognition (HTR)-based approaches; furthermore, they typically cannot handle image queries (QbE). In this work, we propose a time and storage-efficient, deep feature-based approach that enables both the image and textual search options. Three distinct components, all modeled as neural networks, are combined: normalization, feature extraction and representation of image and textual input into a common space. These components, even if designed on word level image representations, collaborate in order to achieve an efficient line level keyword spotting system. The experimental results indicate that the proposed system is on par with state-of-the-art KWS methods. George Retsinas, Georgios Louloudis, Nikolaos Stamatopoulos, Giorgos Sfikas, Basilios Gatos |
CVPR | 1 |
| 2019 | Efficient Learning-Free Keyword SpottingabstractIn this article, a method for segmentation-based learning-free Query by Example (QbE) keyword spotting on handwritten documents is proposed. The method consists of three steps, namely preprocessing, feature extraction and matching, which address critical variations of text images (e.g., skew, translation, different writing styles). During the feature extraction step, a sequence of descriptors is generated using a combination of a zoning scheme and a novel appearance descriptor, referred as modified Projections of Oriented Gradients. The preprocessing step, which includes contrast normalization and main-zone detection, aims to overcome the shortcomings of the appearance descriptor. Moreover, an uneven zoning scheme is introduced by applying a denser zoning only on query images for a more detailed representation. This leads to a significant reduction in storage requirements of a document collection. The distance between the query and word sequences is efficiently computed by the proposed Selective Matching algorithm. This algorithm is further extended to handle an augmented set of images originating from a single query image. The efficiency of the proposed method is demonstrated by experimentation conducted on seven publicly available datasets. In these experiments, the proposed method significantly outperforms all state-of-the-art learning-free techniques. George Retsinas, Georgios Louloudis, Nikolaos Stamatopoulos, Basilios Gatos |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Exploring Critical Aspects of CNN-based Keyword Spotting. A PHOCNet StudyabstractDeep convolutional neural networks are today the new baseline for a wide range of machine vision tasks. The problem of keyword spotting is no exception to this rule. Many successful network architectures and learning strategies have been adapted from other vision tasks to create successful keyword spotting systems. In this paper, we argue that various details concerning this adaptation could be re-examined, to the end of building stronger spotting models. In particular, we examine the usefulness of a pyramidal spatial pooling layer versus a simpler approach, and show that a zoning strategy combined with fixed-size inputs can be just as effective while less computationally expensive. We also examine the usefulness of augmentation, class balancing and ensemble learning strategies and propose an improved network. Our hypotheses are tested with numerical experiments on the IAM document collection, where the proposed network outperforms all other existing models. George Retsinas, Giorgos Sfikas, Nikolaos Stamatopoulos, Georgios Louloudis, Basilios Gatos |
DAS | 1 |
| 2018 | Compact Deep Descriptors for Keyword SpottingabstractIn this work, we present a novel approach for the extraction of deep features from a Convolutional Neural Network (CNN), designed for the task of Keyword Spotting (KWS). The main novelty of our work concerns the generation of a compact descriptor able to simulate the existence/absence of unigrams or bigrams. This is accomplished using a binary, attribute-based representation of a word string together with an appropriate training procedure. Deep features are extracted from the output of the last convolutional layer and are organized into zones in order to incorporate spatial information of the detected attributes. In addition, a novel optimization scheme is proposed which relies on a very effective initialization of the network generating the compact descriptors. Experiments conducted on the IAM dataset prove the efficiency of the novel compact descriptor since the proposed system's performance in on par with the state-of-the-art. George Retsinas, Giorgos Sfikas, Georgios Louloudis, Nikolaos Stamatopoulos, Basilios Gatos |
ICFHR | 1 |
| 2017 | Nonlinear Manifold Embedding on Keyword Spotting Using t-SNEabstractNonlinear manifold embedding has attracted considerable attention due to its highly-desired property of efficiently encoding local structure, i.e. intrinsic space properties, into a low-dimensional space. The benefit of such an approach is twofold: it leads to compact representations while addressing the often-encountered curse of dimensionality. The latter plays an important role in retrieval applications, such as keyword spotting, where a sorted list of retrieved objects with respect to a distance metric is required. In this work, we explore the efficiency of the popular manifold embedding method t-distributed Stochastic Neighbor Embedding (t-SNE) on the Query-by-Example keyword spotting task. The main contribution of this work is the extension of t-SNE in order to support out-of-sample (OOS) embedding which is essential for mapping query images to the embedding space. The experimental results demonstrate a significant increase in keyword spotting performance when the word similarity is calculated on the embedding space. George Retsinas, Nikolaos Stamatopoulos, Georgios Louloudis, Giorgos Sfikas, Basilios Gatos |
ICDAR | 1 |
| 2017 | A PHOC Decoder for Lexicon-Free Handwritten Word RecognitionabstractIn this paper, we propose a novel probabilistic model for lexicon-free handwriting recognition. Model inputs are word images encoded as Pyramidal Histogram Of Character (PHOC) vectors. PHOC vectors have been used as efficient attribute-based, multi-resolution representations of either text strings or word image contents. The proposed model formulates PHOC decoding as the problem of finding the most probable sequence of characters corresponding to the given PHOC. We model PHOC layers as Beta-distributed observations, linked to hidden states that correspond to character estimates. Characters are in turn linked to one another along a Markov chain, encoding language model information. The sequence of characters is estimated using the max-sum algorithm in a process that is akin to Viterbi decoding. Numerical experiments on the well-known George Washington database show competitive recognition results. Giorgos Sfikas, George Retsinas, Basilios Gatos |
ICDAR | 2 |
| 2016 | Efficient Document Image Segmentation Representation by Approximating Minimum-Link PolygonsabstractThe result of a document image segmentation task, e.g. text line or word segmentation, is usually a labeled image with each label corresponding to a different segmented region. For many applications, the segmented regions need to be stored and represented in an efficient way, using simple geometric shapes. A challenging task is to restrict all pixels corresponding to a specific label inside a polygon with a minimum number of vertices. Such a polygon promotes the description simplicity and the storage efficiency, while providing a much more user-friendly representation that can be edited easily. The proposed method is a cost-effective approximation of the minimum-edges polygon problem, computing a contour enclosing only pixels of a certain label and using a greedy algorithm in order to reduce the contour into a minimum-link polygon that retains the separability property between the labeled set of pixels. George Retsinas, Georgios Louloudis, Nikolaos Stamatopoulos, Basilios Gatos |
DAS | 1 |
| 2016 | Keyword Spotting in Handwritten Documents Using Projections of Oriented GradientsabstractIn this paper, we present a novel approach for segmentation-based handwritten keyword spotting. The proposed approach relies upon the extraction of a simple yet efficient descriptor which is based on projections of oriented gradients. To this end, a global and a local word image descriptors, together with their combination, are proposed. Retrieval is performed using to the euclidean distance between the descriptors of a query image and the segmented word images. The proposed methods have been evaluated on the dataset of the ICFHR 2014 Competition on handwritten keyword spotting. Experimental results prove the efficiency of the proposed methods compared to several state-of-the-art techniques. George Retsinas, Georgios Louloudis, Nikolaos Stamatopoulos, Basilios Gatos |
DAS | 1 |
| 2016 | Zoning Aggregated Hypercolumns for Keyword SpottingabstractIn this paper we present a novel descriptor and method for segmentation-based keyword spotting. We introduce Zoning-Aggregated Hypercolumn features as pixel-level cues for document images. Motivated by recent research in machine vision, we use an appropriately pretrained convolutional network as a feature extraction tool. The resulting local cues are subsequently aggregated to form word-level fixed-length descriptors. Encoding is computationally inexpensive and does not require learning a separate feature generative model, in contrast to other widely used encoding methods (such as Fisher Vectors). Keyword spotting trials on machine-printed and handwritten documents show that the proposed model gives very competitive results. Giorgos Sfikas, George Retsinas, Basilios Gatos |
ICFHR | 2 |
| 2015 | GRPOLY-DB: An old Greek polytonic document image databaseabstractRecognition of old Greek document images containing polytonic (multi accent) characters is a challenging task due to the large number of existing character classes (more than 270) which cannot be handled sufficiently by current OCR technologies. Taking into account that the Greek polytonic system was used from the late antiquity until recently, a large amount of scanned Greek documents still remains without full test search capabilities. In order to assist the progress of relevant research, this paper introduces the first publicly available old Greek polytonic database GRPOLY-DB for the evaluation of several document image processing tasks. It contains both machine-printed and handwritten documents as well as annotation with ground-truth information that can be used for training and evaluation of the most commou document image processing tasks, i.e.. text line and word segmentation, test recognition, isolated character recognition and word spotting. Results using several representative baseline technologies are also presented in order to help researchers evaluate their methods and advance the frontiers of old Greek document image recognition and word spotting. Basilios Gatos, Nikolaos Stamatopoulos, Georgios Louloudis, Giorgos Sfikas, George Retsinas, Vassilis Papavassiliou, Fotini Sunistira, Vassilis Katsouros |
ICDAR | 5 |
| 2015 | Isolated character recognition using projections of oriented gradientsabstractIn this paper, we present a new approach for off-line isolated character recognition. The proposed method relies upon the application of a projection-based feature extraction stage, which resembles the Radon transform, on both the original image and a set of generated images corresponding to different gradient orientations of the original image. For the classification stage, Support Vectors Machines (SVM) are used. The proposed method is evaluated using one typewritten (GRPOLY-DB - Historical Greek) and two handwritten (CIL - Greek, CEDAR - English) publicly available databases. Experimental results prove the efficiency of the proposed method compared to several state-of-the-art techniques. George Retsinas, Basilios Gatos, Nikolaos Stamatopoulos, Georgios Louloudis |
ICDAR | 1 |