Seiichi Uchida

dblp:07/2381 · DBLP profile ↗
← Back
101ranked-venue papers in the field
9as first author
23since 2021 · last 2026
0000-0001-8592-7566ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 99 (9 first)Database Systems & Data Management · 1Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2026 Hierarchical Co-embedding of Font Shapes and Impression Tags
Yugo Kubota, Kaito Shiku, Seiichi Uchida
ICDAR (2)3
2025 Total Disentanglement of Font Images Into Style and Character Class Features
Daichi Haraguchi, Wataru Shimoda, Kota Yamaguchi, Seiichi Uchida
ICDAR (2)4
2025 Computer-Aided Multi-stroke Character Simplification by Stroke Removal
Ryo Ishiyama, Shinnosuke Matsuo, Seiichi Uchida
ICDAR (2)3
2025 Inverse Scene Text Removal
Takumi Yoshimatsu, Shumpei Takezaki, Seiichi Uchida
ICDAR (5)3
2024 What Text Design Characterizes Book Genres?
Daichi Haraguchi, Brian Kenji Iwana, Seiichi Uchida
DAS3
2024 Font Impression Estimation in the Wild
Kazuki Kitajima, Daichi Haraguchi, Seiichi Uchida
ICDAR (2)3
2024 Font Style Interpolation with Diffusion Models
Tetta Kondo, Shumpei Takezaki, Daichi Haraguchi, Seiichi Uchida
ICDAR (2)4
2024 Impression-CLIP: Contrastive Shape-Impression Embedding for Fonts
Yugo Kubota, Daichi Haraguchi, Seiichi Uchida
ICDAR (2)3
2024 Learning to Kern: Set-Wise Estimation of Optimal Letter Space
Kei Nakatsuru, Seiichi Uchida
ICDAR (2)2
2024 Typographic Text Generation with Off-the-Shelf Diffusion Model
KhayTze Peong, Seiichi Uchida, Daichi Haraguchi
ICDAR (2)2
2024 Cross-Domain Image Conversion by CycleDM
Sho Shimotsumagari, Shumpei Takezaki, Daichi Haraguchi, Seiichi Uchida
ICDAR (4)4
2023 Contour Completion by Transformers and Its Application to Vector Font Data
Yusuke Nagata, Brian Kenji Iwana, Seiichi Uchida
ICDAR (5)3
2023 Ambigram Generation by a Diffusion Model
Takahiro Shirakawa, Seiichi Uchida
ICDAR (3)2
2023 Analyzing Font Style Usage and Contextual Factors in Real Images
Naoya Yasukochi, Hideaki Hayashi, Daichi Haraguchi, Seiichi Uchida
ICDAR (3)4
2022 Revealing Reliable Signatures by Learning Top-Rank Pairs
Xiaotong Ji, Daiki Suehiro, Seiichi Uchida
DAS4
2022 TrueType Transformer: Character and Font Style Recognition in Outline Format
Yusuke Nagata, Jinki Otao, Daichi Haraguchi, Seiichi Uchida
DAS4
2022 Font Shape-to-Impression Translation
Masaya Ueda, Akisato Kimura, Seiichi Uchida
DAS3
2021 Impressions2Font: Generating Fonts by Specifying Impressions
Seiya Matsuda, Akisato Kimura, Seiichi Uchida
ICDAR (3)3
2021 Attention to Warp: Deep Metric Learning for Multivariate Time Series
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
ICDAR (3)7
2021 Font Style that Fits an Image - Font Generation Based on Image Context
Taiga Miyazono, Brian Kenji Iwana, Daichi Haraguchi, Seiichi Uchida
ICDAR (3)4
2021 Meta-learning of Pooling Layers for Character Recognition
Takato Otsuzuki, Heon Song, Seiichi Uchida, Hideaki Hayashi
ICDAR (3)3
2021 Which Parts Determine the Impression of the Font?
Masaya Ueda, Akisato Kimura, Seiichi Uchida
ICDAR (3)3
2021 Towards Book Cover Design via Layout Graphs
Taiga Miyazono, Seiichi Uchida, Brian Kenji Iwana
ICDAR (3)4
2020 Neural Style Difference Transfer and Its Application to Font Generation
Gantugs Atarsaikhan, Brian Kenji Iwana, Seiichi Uchida
DAS3
2020 Character-Independent Font Identification
Daichi Haraguchi, Shota Harada, Brian Kenji Iwana, Yuto Shinahara, Seiichi Uchida
DAS5
2020 Effect of Text Color on Word Embeddings
Masaya Ikoma, Brian Kenji Iwana, Seiichi Uchida
DAS3
2020 ACMU-Nets: Attention Cascading Modular U-Nets Incorporating Squeeze and Excitation Blocks
Seokjun Kang, Brian Kenji Iwana, Seiichi Uchida
DAS3
2020 Lyric Video Analysis Using Text Detection and Tracking
Shota Sakaguchi, Jun Kato 0001, Masataka Goto, Seiichi Uchida
DAS4
2019 On the Ability of a CNN to Realize Image-to-Image Language Conversion
abstract
The purpose of this paper is to reveal the ability that Convolutional Neural Networks (CNN) have on the novel task of image-to-image language conversion. We propose a new network to tackle this task by converting images of Korean Hangul characters directly into images of the phonetic Latin character equivalent. The conversion rules between Hangul and the phonetic symbols are not explicitly provided. The results of the proposed network show that it is possible to perform image-to-image language conversion. Moreover, it shows that it can grasp the structural features of Hangul even from limited learning data. In addition, it introduces a new network to use when the input and output have significantly different features.
Kohei Baba, Seiichi Uchida, Brian Kenji Iwana
ICDAR2
2019 Cascading Modular U-Nets for Document Image Binarization
abstract
In recent years, U-Net has achieved good results in various image processing tasks. However, conventional U-Nets need to be re-trained for individual tasks with enough amount of images with ground-truth. This requirement makes U-Net not applicable to tasks with small amounts of data. In this paper, we propose to use "modular" U-Nets, each of which is pre-trained to perform an existing image processing task, such as dilation, erosion, and histogram equalization. Then, to accomplish a specific image processing task, such as binarization of historical document images, the modular U-Nets are cascaded with inter-module skip connections and fine-tuned to the target task. We verified the proposed model using the Document Image Binarization Competition (DIBCO) 2017 dataset.
Seokjun Kang, Brian Kenji Iwana, Seiichi Uchida
ICDAR3
2019 Logo Design Analysis by Ranking
abstract
In this paper, we analyze logo designs by using machine learning, as a promising trial of graphic design analysis. Specifically, we will focus on favicon images, which are tiny logos used as company icons on web browsers, and analyze them to understand their trends in individual industry classes. For example, if we can catch the subtle trends in favicons of financial companies, they will suggest to us how professional designers express the atmosphere of financial companies graphically. For the purpose, we will use top-rank learning, which is one of the recent machine learning methods for ranking and very suitable for revealing the subtle trends in graphic designs.
Takuro Karamatsu, Daiki Suehiro, Seiichi Uchida
ICDAR3
2019 Page Segmentation using a Convolutional Neural Network with Trainable Co-Occurrence Features
abstract
In document analysis, page segmentation is a fundamental task that divides a document image into semantic regions. In addition to local features, such as pixel-wise information, co-occurrence features are also useful for extracting texture-like periodic information for accurate segmentation. However, existing convolutional neural network (CNN)-based methods do not have any mechanisms that explicitly extract co-occurrence features. In this paper, we propose a method for page segmentation using a CNN with trainable multiplication layers (TMLs). The TML is specialized for extracting co-occurrences from feature maps, thereby supporting the detection of objects with similar textures and periodicities. This property is also considered to be effective for document image analysis because of regularity in text line structures, tables, etc. In the experiment, we achieved promising performance on a pixel-wise page segmentation task by combining TMLs with U-Net. The results demonstrate that TMLs can improve performance compared to the original U-Net. The results also demonstrate that TMLs are helpful for detecting regions with periodically repeating features, such as tables and main text.
Hideaki Hayashi, Wataru Ohyama, Seiichi Uchida
ICDAR4
2019 Scene Text Magnifier
abstract
Scene text magnifier aims to magnify text in natural scene images without recognition. It could help the special groups, who have myopia or dyslexia to better understand the scene. In this paper, we design the scene text magnifier through interacted four CNN-based networks: character erasing, character extraction, character magnify, and image synthesis. The architecture of the networks are extended based on the hourglass encoder-decoders. It inputs the original scene text image and outputs the text magnified image while keeps the background unchange. Intermediately, we can get the side-output results of text erasing and text extraction. The four sub-networks are first trained independently and fine-tuned in end-to-end mode. The training samples for each stage are processed through a flow with original image and text annotation in ICDAR2013 and Flickr dataset as input, and corresponding text erased image, magnified text annotation, and text magnified scene image as output. To evaluate the performance of text magnifier, the Structural Similarity is used to measure the regional changes in each character region. The experimental results demonstrate our method can magnify scene text effectively without effecting the background.
Toshiki Nakamura, Anna Zhu, Seiichi Uchida
ICDAR3
2019 Selective Super-Resolution for Scene Text Images
abstract
In this paper, we realize the enhancement of super-resolution using images with scene text. Specifically, this paper proposes the use of Super-Resolution Convolutional Neural Networks (SRCNN) which are constructed to tackle issues associated with characters and text. We demonstrate that standard SRCNNs trained for general object super-resolution is not sufficient and that the proposed method is a viable method in creating a robust model for text. To do so, we analyze the characteristics of SRCNNs through quantitative and qualitative evaluations with scene text data. In addition, analysis using the correlation between layers by Singular Vector Canonical Correlation Analysis (SVCCA) and comparison of filters of each SRCNN using t-SNE is performed. Furthermore, in order to create a unified super-resolution model specialized for both text and objects, a model using SRCNNs trained with the different data types and Content-wise Network Fusion (CNF) is used. We integrate the SRCNN trained for character images and then SRCNN trained for general object images, and verify the accuracy improvement of scene images which include text. We also examine how each SRCNN affects super-resolution images after fusion.
Ryo Nakao, Brian Kenji Iwana, Seiichi Uchida
ICDAR3
2019 Training Convolutional Autoencoders with Metric Learning
abstract
We propose a new Training method that enables an autoencoder to extract more useful features for retrieval or classification tasks with limited-size datasets. Some targets in document analysis and recognition (DAR) including signature verification, historical document analysis, and scene text recognition, involve a common problem in which the size of the dataset available for training is small against the intra-class variety of the target appearance. Recently, several approaches, such as variational autoencoders and deep metric learning, have been proposed to obtain a feature representation that is suitable for the tasks. However, these methods sometimes cause an overfitting problem in which the accuracy of the test data is relatively low, while the performance for the training dataset is quite high. Our proposed method obtains feature representations for such tasks in DAR using convolutional autoencoders with metric learning. The accuracy is evaluated on an image-based retrieval of ancient Japanese signatures.
Yosuke Onitsuka, Wataru Ohyama, Seiichi Uchida
ICDAR3
2019 Serif or Sans: Visual Font Analytics on Book Covers and Online Advertisements
abstract
In this paper, we conduct a large-scale study of font statistics in book covers and online advertisements. Through the statistical study, we try to understand how graphic designers relate fonts and content genres and identify the relationship between font styles, colors, and genres. We propose an automatic approach to extract font information from graphic designs by applying a sequence of character detection, style classification, and clustering techniques to the graphic designs. The extracted font information is accumulated together with genre information, such as romance or business, for further trend analysis. Through our unique empirical study, we show that the collected font statistics reveal interesting trends in terms of how typographic design represents the impression and the atmosphere of the content genres.
Yuto Shinahara, Takuro Karamatsu, Daisuke Harada, Kota Yamaguchi, Seiichi Uchida
ICDAR5
2019 Modality Conversion of Handwritten Patterns by Cross Variational Autoencoders
abstract
This research attempts to construct a network that can convert online and offline handwritten characters to each other. The proposed network consists of two Variational Auto-Encoders (VAEs) with a shared latent space. The VAEs are trained to generate online and offline handwritten Latin characters simultaneously. In this way, we create a cross-modal VAE (Cross-VAE). During training, the proposed Cross-VAE is trained to minimize the reconstruction loss of the two modalities, the distribution loss of the two VAEs, and a novel third loss called the space sharing loss. This third, space sharing loss is used to encourage the modalities to share the same latent space by calculating the distance between the latent variables. Through the proposed method mutual conversion of online and offline handwritten characters is possible. In this paper, we demonstrate the performance of the Cross-VAE through qualitative and quantitative analysis.
Taichi Sumi, Brian Kenji Iwana, Hideaki Hayashi, Seiichi Uchida
ICDAR4
2019 Deep Dynamic Time Warping: End-to-End Local Representation Learning for Online Signature Verification
abstract
Siamese networks have been shown to be successful in learning deep representations for multivariate time series verification. However, most related studies optimize a global distance objective and suffer from a low discriminative power due to the loss of temporal information. To address this issue, we propose an end-to-end, neural network-based framework for learning local representations of time series, and demonstrate its effectiveness for online signature verification. This framework optimizes a Siamese network with a local embedding loss, and learns a feature space that preserves the temporal location-wise distances between time series. To achieve invariance to non-linear temporal distortion, we propose building a dynamic time warping block on top of the Siamese network, which will greatly improve the accuracy for local correspondences across intra-personal variability. Validation with respect to online signature verification demonstrates the advantage of our framework over existing techniques that use either handcrafted or learned feature representations.
Xiaomeng Wu, Akisato Kimura, Brian Kenji Iwana, Seiichi Uchida, Kunio Kashino
ICDAR4
2019 Capturing Micro Deformations from Pooling Layers for Offline Signature Verification
abstract
In this paper, we propose a novel Convolutional Neural Network (CNN) based method that extracts the location information (displacement features) of the maximums in the max-pooling operation and fuses it with the pooling features to capture the micro deformations between the genuine signatures and skilled forgeries as a feature extraction procedure. After the feature extraction procedure, we apply support vector machines (SVMs) as writer-dependent classifiers for each user to build the signature verification system. The extensive experimental results on GPDS-150, GPDS-300, GPDS-1000, GPDS-2000, and GPDS-5000 datasets demonstrate that the proposed method can discriminate the genuine signatures and their corresponding skilled forgeries well and achieve state-of-the-art results on these datasets.
Yuchen Zheng 0001, Wataru Ohyama, Brian Kenji Iwana, Seiichi Uchida
ICDAR4
2019 RankSVM for Offline Signature Verification
abstract
Signature verification systems suffer from imbalanced learning, which imposes strict requirements on classifiers. The standard classification approaches, such as SVM, often degrade the performance for imbalanced data or require additional parameters for data balancing. In this study, as a new approach for signature verification, we use RankSVM as the writer-dependent classifiers, which theoretically guarantees the generalization performance for imbalanced data. To investigate the ability of RankSVM for solving imbalanced learning problems in signature verification tasks, the extensive experiments are conducted on bitmaps of GPDS-150, GPDS-300, GPDS-600, and GPDS-1000 datasets and deep features of GPDS-960 dataset. The experimental results demonstrate that the RankSVM-based approach obtains a nearly equivalent performance with the state-of-the-art method on deep features of the GPDS-960 dataset, and achieves significantly better performance than standard-SVM-based approach on bitmaps of GPDS-150, GPDS-300, GPDS-600, and GPDS-1000 datasets.
Yuchen Zheng 0001, Wataru Ohyama, Daiki Suehiro, Seiichi Uchida
ICDAR5
2018 Contained Neural Style Transfer for Decorated Logo Generation
abstract
Making decorated logos requires image editing skills, without sufficient skills, it could be a time-consuming task. While there are many on-line web services to make new logos, they have limited designs and duplicates can be made. We propose using neural style transfer with clip art and text for the creation of new and genuine logos. We introduce a new loss function based on distance transform of the input image, which allows the preservation of the silhouettes of text and objects. The proposed method contains style transfer to only a designated area. We demonstrate the characteristics of proposed method. Finally, we show the results of logo generation with various input images.
Gantugs Atarsaikhan, Brian Kenji Iwana, Seiichi Uchida
DAS3
2018 CNN Training with Graph-Based Sample Preselection: Application to Handwritten Character Recognition
abstract
In this paper, we present a study on sample preselection in large training data set for CNN-based classification. To do so, we structure the input data set in a network representation, namely the Relative Neighbourhood Graph, and then extract some vectors of interest. The proposed preselection method is evaluated in the context of handwritten character recognition, by using two data sets, up to several hundred thousands of images. It is shown that the graph-based preselection can reduce the training data set without degrading the recognition accuracy of a non pretrained CNN shallow model.
Frédéric Rayar, Masanori Goto, Seiichi Uchida
DAS3
2018 Text Line Extraction Based on Integrated K-Shortest Paths Optimization
abstract
Text in images can be utilized in many image understanding applications due to the exact semantic information. In this paper, we propose a novel integrated k-shortest paths optimization based text line extraction method. Firstly, the candidate text components are extracted by the Maximal Stable Extremal Region (MSER) algorithm on gray, red, green and blue channels. Secondly, one integrated directed graph on red, green, and blue channels are constructed upon the candidate text components, which can effectively incorporate different channels into one framework. Then, the integrated directed graph is transformed guided by the extracted text lines in gray channel to reduced the computational complexity. Finally, we use the k-shortest paths optimization algorithm to extract the text lines by taking advantage of the particular structure of the integrated directed graph. Experimental results demonstrate the effectiveness of the proposed method in comparison with state-of-the-art methods.
Liuan Wang, Jun Sun 0004, Seiichi Uchida
DAS3
2017 How Does a CNN Manage Different Printing Types?
abstract
In past OCR research, different OCR engines are used for different printing types, i.e., machine-printed characters, handwritten characters, and decorated fonts. A recent research, however, reveals that convolutional neural networks (CNN) can realize a universal OCR, which can deal with any printing types without pre-classification into individual types. In this paper, we analyze how CNN for universal OCR manage the different printing types. More specifically, we try to find where a handwritten character of a class and a machine-printed character of the same class are "fused" in CNN. For analysis, we use two different approaches. The first approach is statistical analysis for detecting the CNN units which are sensitive (or insensitive) to type difference. The second approach is network-based visualization of pattern distribution in each layer. Both analyses suggest the same trend that types are not fully fused in convolutional layers but the distributions of the same class from different types become closer in upper layers.
Shouta Ide, Seiichi Uchida
ICDAR2
2017 Component Awareness in Convolutional Neural Networks
abstract
In this work, we investigate the ability of Convolutional Neural Networks (CNN) to infer the presence of components that comprise an image. In recent years, CNNs have achieved powerful results in classification, detection, and segmentation. However, these models learn from instance-level supervision of the detected object. In this paper, we determine if CNNs can detect objects using image-level weakly supervised labels without localization. To demonstrate that a CNN can infer awareness of objects, we evaluate a CNN's classification ability with a database constructed of Chinese characters with only character-level labeled components. We show that the CNN is able to achieve a high accuracy in identifying the presence of these components without specific knowledge of the component. Furthermore, we verify that the CNN is deducing the knowledge of the target component by comparing the results to an experiment with the component removed. This research is important for applications with large amounts of data without robust annotation such as Chinese character recognition.
Brian Kenji Iwana, Letao Zhou, Kumiko Tanaka-Ishii, Seiichi Uchida
ICDAR4
2017 Scene Text Eraser
abstract
The character information in natural scene images contains various personal information, such as telephone numbers, home addresses, etc. It is a high risk of leakage the information if they are published. In this paper, we proposed a scene text erasing method to properly hide the information via an inpainting convolutional neural network (CNN) model. The input is a scene text image, and the output is expected to be text erased image with all the character regions filled up the colors of the surrounding background pixels. This work is accomplished by a CNN model through convolution to deconvolution with interconnection process. The training samples and the corresponding inpainting images are considered as teaching signals for training. To evaluate the text erasing performance, the output images are detected by a novel scene text detection method. Subsequently, the same measurement on text detection is utilized for testing the images in benchmark dataset ICDAR2013. Compared with direct text detection way, the scene text erasing process demonstrates a drastically decrease on the precision, recall and f-score. That proves the effectiveness of proposed method for erasing the text in natural scene images.
Toshiki Nakamura, Anna Zhu, Keiji Yanai, Seiichi Uchida
ICDAR4
2017 Scene Text Relocation with Guidance
abstract
Applying object proposal technique for scene text detection becomes popular for its significant improvement in speed and accuracy for object detection. However, some of the text regions after the proposal classification are overlapped and hard to remove or merge. In this paper, we present a scene text relocation system that refines the detection from text proposals to text. An object proposal-based deep neural network is employed to get the text proposals. To tackle the detection overlapping problem, a refinement deep neural network relocates the overlapped regions by estimating the text probability inside, and locating the accurate text regions by thresholding. Since the space between words indifferent text lines are various, a guidance mechanism is proposed in text relocation to guide where to extract the text regions in word level. This refinement procedure helps boost the precision after removing multiple overlapped text regions or joint cracked text regions. The experimental results on standard benchmark ICDAR 2013 demonstrate the effectiveness of the proposed approach.
Anna Zhu, Seiichi Uchida
ICDAR2
2016 Globally Optimal Text Line Extraction Based on K-Shortest Paths Algorithm
abstract
The task of text line extraction in images is a crucial prerequisite for content-based image understanding applications. In this paper, we propose a novel text line extraction method based on k-shortest paths global optimization in images. Firstly, the candidate connected components are extracted by reformulating it as Maximal Stable Extremal Region (MSER) results in images. Then, the directed graph is built upon the connected component nodes with edges comprising of unary and pairwise cost function. Finally, the text line extraction problem is solved using the k-shortest paths optimization algorithm by taking advantage of the particular structure of the directed graph. Experimental results on public dataset demonstrate the effectiveness of proposed method in comparison with state-of-the-art methods.
Liuan Wang, Seiichi Uchida, Wei Fan 0005, Jun Sun 0004
DAS2
2015 Similarity-based regularization for semi-supervised learning for handwritten digit recognition
abstract
This paper presents an experimental analysis on the use of semi-supervised learning in the handwritten digit recognition field. More specifically, two new feedback-based techniques for retraining individual classifiers in a multi-expert scenario are discussed. These new methods analyze the final decision provided by the multi-expert system so that sample classified with a confidence greater than a specific threshold is used to update the system itself. Experimental results carried out on the CEDAR (handwritten digits) database are presented. In particular, error rate, similarity index and a new correlation score among them are considered in order to evaluate the best retraining rule. For the experimental evaluation, an SVM classifier and five different combination techniques at abstract and measurement level have been used. Finally, the results show that iterating the feedback process, on different multi-expert systems built with the five combination techniques, one retraining rule is winning over the other respect to the best correlation score.
Donato Barbuzzi, Giuseppe Pirlo, Seiichi Uchida, Volkmar Frinken, Donato Impedovo
ICDAR3
2015 Deep BLSTM neural networks for unconstrained continuous handwritten text recognition
abstract
Recently, two different trends in neural network-based machine learning could be observed. The first one are the introduction of Bidirectional Long Short-Term Memory (BLSTM) neural networks (NN) which made sequences with long-distant dependencies amenable for neural network-based processing. The second one are deep learning techniques, which greatly increased the performance of neural networks, by making use of many hidden layers. In this paper, we propose to combine these two ideas for the task of unconstrained handwriting recognition. Extensive experimental evaluation on the IAM database demonstrate an increase of the recognition performance when using deep learning approaches over commonly used BLSTM neural networks, as well as insight into how different types of hidden layers affect the recognition accuracy.
Volkmar Frinken, Seiichi Uchida
ICDAR2
2015 True color distributions of scene text and background
abstract
Color feature, as one of the low level features, plays important role in image processing, object recognition and other fields. For example, in the task of scene text detection and recognition, lots of methodologies employ features that utilize color contrast of text and the corresponding background for connected component extraction. However, the true distributions of text and its background, in terms of color, is still not examined because it requires an enough number of scene text database with pixel-level labelled text/non-text ground truth. To clarify the relationship between text and its background, in this paper, we aim at investigating the color non-parametric distribution of text and its background using a large database that contains 3018 scene images and 98,600 characters. The results of our experiments show that text and its background can be discriminated by means of color, therefore color feature can be used for scene text detection.
Renwu Gao, Shoma Eguchi, Seiichi Uchida
ICDAR3
2015 Preselection of support vector candidates by relative neighborhood graph for large-scale character recognition
abstract
We propose a pre-selection method for training support vector machines (SVM) with a large-scale dataset. Specifically, the proposed method selects patterns around the class boundary and the selected data is fed to train an SVM. For the selection, that is, searching for boundary patterns, we utilize a relative neighborhood graph (RNG). An RNG has an edge for each pair of neighboring patterns and thus, we can find boundary patterns by looking for edges connecting patterns from different classes. Through large-scale handwritten digit pattern recognition experiments, we show that the proposed pre-selection method accelerates SVM training process 5-15 times faster without degrading recognition accuracy.
Masanori Goto, Ryosuke Ishida, Seiichi Uchida
ICDAR3
2015 Tackling temporal pattern recognition by vector space embedding
abstract
This paper introduces a novel method of reducing the number of prototype patterns necessary for accurate recognition of temporal patterns. The nearest neighbor (NN) method is an effective tool in pattern recognition, but the downside is it can be computationally costly when using large quantities of data. To solve this problem, we propose a method of representing the temporal patterns by embedding dynamic time warping (DTW) distance based dissimilarities in vector space. Adaptive boosting (AdaBoost) is then applied for classifier training and feature selection to reduce the number of prototype patterns required for accurate recognition. With a data set of handwritten digits provided by the International Unipen Foundation (iUF), we successfully show that a large quantity of temporal data can be efficiently classified produce similar results to the established NN method while performing at a much smaller cost.
Brian Kenji Iwana, Seiichi Uchida, Kaspar Riesen, Volkmar Frinken
ICDAR2
2015 Learning non-Markovian constraints for handwriting recognition
abstract
Recently, the horizon of dynamic time warping (DTW) for matching two sequential patterns has been extended to deal with non-Markovian constraints. The non-Markovian constraints regulate the matching in a wider scale, whereas Markovian constraints regulate the matching only locally. The global optimization of the non-Markovian DTW is proved to be solvable in polynomial time by a graph cut algorithm. The main contribution of this paper is to reveal what is the best constraint for handwriting recognition by using the non-Markovian DTW. The result showed that the best constraint is not a Markovian but a totally non-Markovian constraint that regulates the matching between very distant points; that is, it was proved that the conventional Markovian DTW has a clear limitation and the non- Markovian DTW should be more focused in future research.
Ryosuke Kakisako, Seiichi Uchida, Volkmar Frinken
ICDAR2
2015 ICDAR 2015 competition on Robust Reading
abstract
Results of the ICDAR 2015 Robust Reading Competition are presented. A new Challenge 4 on Incidental Scene Text has been added to the Challenges on Born-Digital Images, Focused Scene Images and Video Text. Challenge 4 is run on a newly acquired dataset of 1,670 images evaluating Text Localisation, Word Recognition and End-to-End pipelines. In addition, the dataset for Challenge 3 on Video Text has been substantially updated with more video sequences and more accurate ground truth data. Finally, tasks assessing End-to-End system performance have been introduced to all Challenges. The competition took place in the first quarter of 2015, and received a total of 44 submissions. Only the tasks newly introduced in 2015 are reported on. The datasets, the ground truth specification and the evaluation protocols are presented together with the results and a brief summary of the participating methods.
Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukás Neumann, Vijay Chandrasekhar 0001, Shijian Lu, Faisal Shafait, Seiichi Uchida, Ernest Valveny
ICDAR12
2015 Exploring the world of fonts for discovering the most standard fonts and the missing fonts
abstract
This paper has two contributions toward understanding the principles in font design. The first contribution of this paper is to discover the most standard font shape of each letter class by analyzing thousands of different fonts. For this analysis, two different methods are used. The first method is congealing for aligning multiple images based on a nonlinear geometric transformation model. The average of the aligned image is considered as a standard font shape. The second method is network analysis for representing font variations as a large-scale relative neighborhood graph (RNG) and then finding its center. The font corresponding to the center is considered as the standard font shape. Both of the standard font shapes given by the two methods are plain without decoration, serif, or slant, and thus give an objective reason why we consider the plain font as the typical font shape. The second contribution is to utilize the RNG and the pairwise congealing technique for discovering unexplored font designs and then generating totally new fonts automatically.
Seiichi Uchida, Yuji Egashira, Kota Sato
ICDAR1
2013 Analyzing the Distribution of a Large-Scale Character Pattern Set Using Relative Neighborhood Graph
abstract
The goal of this research is to understand the true distribution of character patterns. Advances in computer technology for mass storage and digital processing have paved way to process a massive dataset for various pattern recognition problems. If we can represent and analyze the distribution of a large-scale character pattern set directly and understand its relationships deeply, it should be helpful for improving character recognizer. For this purpose, we propose a network analysis method to represent the distribution of patterns using a relative neighborhood graph and its clustered version. In this paper, the properties and validity of the proposed method are confirmed on 410,564 machine-printed digit patterns and 622,660 handwritten digit patterns which were manually ground-truthed and resized to 16 times 16 pixels. Our network analysis method represents the distribution of the patterns without any assumption, approximation or loss.
Masanori Goto, Ryosuke Ishida, Yaokai Feng, Seiichi Uchida
ICDAR4
2013 Scene Character Detection by an Edge-Ray Filter
abstract
Edge is a type of valuable clues for scene character detection task. Generally, the existing edge-based methods rely on the assumption of straight text line to prune away the non-character candidates. This paper proposes a new edge-based method, called edge-ray filter, to detect the scene character. The main contribution of the proposed method lies in filtering out complex backgrounds by fully utilizing the essential spatial layout of edges instead of the assumption of straight text line. Edges are extracted by a combination of Canny and Edge Preserving Smoothing Filter (EPSF). To effectively boost the filtering strength of the designed edge-ray filter, we employ a new Edge Quasi-Connectivity Analysis (EQCA) to unify complex edges as well as contour of broken character. Label Histogram Analysis (LHA) then filters out non-character edges and redundant rays through setting proper thresholds. Finally, two frequently-used heuristic rules, namely aspect ratio and occupation, are exploited to wipe off distinct false alarms. In addition to have the ability to handle special scenarios, the proposed method can accommodate dark-on-bright and bright-on-dark characters simultaneously, and provides accurate character segmentation masks. We perform experiments on the benchmark ICDAR 2011 Robust Reading Competition dataset as well as scene images with special scenarios. The experimental results demonstrate the validity of our proposal.
Rong Huang 0003, Palaiahnakote Shivakumara, Seiichi Uchida
ICDAR3
2013 ICDAR 2013 Robust Reading Competition
abstract
This report presents the final results of the ICDAR 2013 Robust Reading Competition. The competition is structured in three Challenges addressing text extraction in different application domains, namely born-digital images, real scene images and real-scene videos. The Challenges are organised around specific tasks covering text localisation, text segmentation and word recognition. The competition took place in the first quarter of 2013, and received a total of 42 submissions over the different tasks offered. This report describes the datasets and ground truth specification, details the performance evaluation protocols used and presents the final results along with a brief summary of the participating methods.
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluís Gómez i Bigorda, Sergi Robles, Joan Mas Romeu, David Fernández Mota, Jon Almazán, Lluís-Pere de las Heras
ICDAR3
2013 The Reading-Life Log - Technologies to Recognize Texts That We Read
abstract
Reading life log is a type of techniques to automatically and unconsciously record people's reading intentions, interests and habits. Besides, it can also serve as various assistants in our daily life. In this paper, a reading-life log system is implemented by a head-mounted and unobtrusive video camera with a high resolution and a high shutter speed. We utilize DP matching, and propose a text-based frame mosaicing method to integrate multiple frames in a clip. The developed system is tested in the various environments indoor and outdoor. The experimental results show that our system can provide reliable outputs with respect to the most correct responses. The infrequent misregistration between lines also indicates the feasibility and validity of the text-based frame mosaicing.
Takashi Kimura, Rong Huang 0003, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
ICDAR3
2013 On the Possibility of Structure Learning-Based Scene Character Detector
abstract
In this paper, we propose a structure learning-based scene character detector which is inspired by the observation that characters have their own inherent structures compared with the background. Graphs are extracted from the thinned binary image to represent the topological line structures of scene contents. Then, a graph classifier, namely gBoost classifier, is trained with the intent to seek out the inherent structures of character and the counterparts of non-character. The experimental results show that the proposed detector achieves the remarkable classification performance with the accuracy of about 70%, which demonstrates the existence and separability of the inherent structures.
Yugo Terada, Rong Huang 0003, Yaokai Feng, Seiichi Uchida
ICDAR4
2013 Part-Based Recognition of Arbitrary Fonts
abstract
In this paper, the part-based recognition method is introduced and applied to the arbitrary font recognition. The principle of the part-based method is to represent the character image as a set of parts and then recognize the image by finding the most possible parts set from the reference database. Since the part-based method does not rely on the global structure of a character, it is supposed to be robust against the variant appearances of the character. The experiment results indicate that it is possible to apply the part-based method to the font recognition, which is always considered as a difficult task by most of the researchers.
Seiichi Uchida, Marcus Liwicki
ICDAR2
2012 How Important is Global Structure for Characters?
abstract
This paper studies the importance of the features that represent the global structure of character strokes to character recognition. Most existing character recognition methods based on character stroke features utilize a set or a sequence of local features such as xy-coordinates and local direction of strokes. This is natural from the viewpoint that each stroke is a trajectory and thus can be represented as a sequence of local features. This viewpoint, however, has a clear limitation in that local features cannot deal with global structure directly. For example, the sequence of local features cannot deal with the fact that the two end points of character "0" should be close to each other. In this paper we propose a simple and novel global feature that describes the global structure of the character shape of each class. We prove the importance of the global feature through a feature selection experiment. Specifically, we show that the global features are more often selected than local features to enhance classification accuracy under the AdaBoost-based machine learning framework. Recognition experiments using online numeral data show also that the use of global features improves recognition accuracy.
Minoru Mori, Seiichi Uchida, Hitoshi Sakano
Document Analysis Systems2
2012 How Salient is Scene Text?
abstract
Computational models of visual attention use image features to identify salient locations in an image that are likely to attract human attention. Attention models have been quite effectively used for various object detection tasks. However, their use for scene text detection is under-investigated. As a general observation, scene text often conveys important information and is usually prominent or salient in the scene itself. In this paper, we evaluate four state-of-the-art attention models for their response to scene text. Initial results indicate that saliency maps produced by these attention models can be used for aiding scene text detection algorithms by suppressing non-text regions.
Asif Shahab, Faisal Shafait, Andreas Dengel 0001, Seiichi Uchida
Document Analysis Systems4
2012 A Part-Based Skew Estimation Method
abstract
In this paper we propose a part-based skew estimation method which is more robust to larger varieties of text images, such as camera-captured scene images. Specifically, the skew angle at each local part of the input image is estimated independently by referring the local part of upright character images stored as a database. Then the global skew angle is estimated by aggregating the estimated local skews. The proposed method does not assume that characters are laid-out in straight lines and thus have more robustness to the varieties of text images than conventional methods. The experimental results show the advantage of the proposed method over the conventional methods under several conditions.
Soma Shiraishi, Yaokai Feng, Seiichi Uchida
Document Analysis Systems3
2012 Toward Part-Based Document Image Decoding
abstract
Document image decoding (DID) is a trial to understand the contents of a whole document without any reference information about font, language, etc. Typically, DID approaches assume the correct segmentation of the document and some a priori knowledge about the language or the script. Unfortunately, this assumption will not hold if we deal with various documents, such as documents with various sized fonts, camera-captured documents, free-layout documents, or historical documents. In this paper, we propose a part-based character identification method where no segmentation into characters is necessary and no a priori information about the document is needed. The approach clusters similar key points and groups frequent neighboring key point clusters. Then a second iteration is performed, i.e., the groups are again clustered and optionally pairs frequent group clusters are detected. Our first experimental results on multi font-size documents look already very promising. We could find nearly perfect correspondences between characters and detected group clusters.
Wang Song, Seiichi Uchida, Marcus Liwicki
Document Analysis Systems2
2011 Scenery Character Detection with Environmental Context
abstract
For scenery character detection, we introduce environmental context, which is modeled by scene components, such as sky and building. Environmental context is expected to regulate the probability of character existence at a specific region in a scenery image. For example, if a region looks like a part of a building, the region has a higher probability than another region like a part of the sky. In this paper, environmental context is represented by state-of-the-art texture and color features and utilized in two different ways. Through experimental results, it was clearly shown that the environmental context has an effect of improving detection accuracy.
Yasuhiro Kunishige, Yaokai Feng, Seiichi Uchida
ICDAR3
2011 Reliable Online Stroke Recovery from Offline Data with the Data-Embedding Pen
abstract
In this paper we propose a complete system for online stroke recovery from offline data. The key idea of our approach is to use a novel pen device which is able to embed meta information into the ink during writing the strokes. This pen-device overcomes the need to get access to any memory on the pen when trying to recover the information, which is especially useful in multi-writer or multi-pen scenarios. The actual data-embedding is achieved by an additional ink dot sequence along a handwritten pattern during writing. We design the ink-dot sequence in such a way that it is possible to retrieve the writing direction from a scanned image. Furthermore, we propose novel processing steps in order to retrieve the original writing direction and finally the embedded data. In our experiments we show that we can reliably recover the writing direction of various patterns. Our system is able to determine the writing direction of straight lines, simple patterns with crossings (e.g., "x" and "II"), and even more complex patterns like handwritten words and symbols.
Marcus Liwicki, Akira Yoshida, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
ICDAR3
2011 Look Inside the World of Parts of Handwritten Characters
abstract
Part-based recognition is expected to be robust in difficult handwritten character recognition tasks. This is because part-based recognition is based on aggregation of independent recognition results at individual local parts without considering their global relations and thus is robust against various deformations, such as partial occlusion, overlap, broken stroke, etc. Since part-based recognition is a new approach, there are still several open problems toward its practical use. For example, compared with entire images, local parts are more ambiguous, i.e., less discriminative. For better recognition accuracy and less computations, we need to know the characteristics of local parts and then, for example, discard less discriminative parts. The purpose of this paper is to conduct some experiments in order to observe and analyze how the local parts of multiple classes are distributed in feature spaces. By handling parts appropriately based on the analysis, we will be able to enhance the usefulness of the part-based method.
Wang Song, Seiichi Uchida, Marcus Liwicki
ICDAR2
2011 Comparative Study of Part-Based Handwritten Character Recognition Methods
abstract
The purpose of this paper is to introduce three part-based methods for handwritten character recognition and then compare their performances experimentally. All of those methods decompose handwritten characters into "parts". Then some recognition processes are done in a part-wise manner and, finally, the recognition results at all the parts are combined via voting to have the recognition result of the entire character. Since part-based methods do not rely on the global structure of the character, we can expect their robustness against various deformations. Three voting methods have been investigated for the combination: single voting, multiple voting, and class distance. All of them use different strategies for voting. Experimental results on the MNIST database showed the relative superiority of the class distance method and the robustness of the multiple voting method against the reduction of training set.
Wang Song, Seiichi Uchida, Marcus Liwicki
ICDAR2
2011 A Generative Model for Handwritings Based on Enhanced Feature Desynchronization
abstract
A new generative model of handwriting patterns is proposed for interpreting their deformations. The model is based on feature desynchronization, which is a coupling process of x and y coordinate features of different timings. By changing the timings to be coupled, the model can generate various deformed patterns from a single pattern. The model is further enhanced by incorporating an adaptive rotation at each timing for increasing the variety of deformed patterns. An important fact is that this enhanced desynchronization model can be interpreted intuitively as a deformation process in actual handwriting. Experimental results showed that the model can generate various handwriting patterns close to actual deformed patterns.
Seiichi Uchida, Toru Sasaki, Yaokai Feng
ICDAR1
2011 A Keypoint-Based Approach toward Scenery Character Detection
abstract
This paper proposes a new approach toward scenery character detection. This is a key point-based approach where local features and a saliency map are fully utilized. Local features, such as SIFT and SURF, have been commonly used for computer vision and object pattern recognition problems, however, they have been rarely employed in character recognition and detection problems. Local feature, however, is similar to directional features, which have been employed in character recognition applications. In addition, local feature can detect corners and thus it is suitable for detecting characters, which are generally comprised of many corners. For evaluating the performance of the local feature, an experimental result was done and its results showed that SURF, i.e., a simple gradient feature, can detect about 70% of characters in scenery images. Then the saliency map was employed as an additional feature to the local feature. This trial is based on the expectation that scenery characters are generally printed to be salient and thus higher salient area will have a higher probability to be a character area. An experimental result showed that this expectation was reasonable and we can have better discrimination accuracy with the saliency map.
Seiichi Uchida, Yuki Shigeyoshi, Yasuhiro Kunishige, Yaokai Feng
ICDAR1
2010 Expansion of queries and databases for improving the retrieval accuracy of document portions: an application to a camera-pen system
abstract
This paper presents a method of improving the accuracy of document image retrieval focusing on the application to a camera-pen system. In a camera-pen system, document image retrieval is employed for locating the pen-tip position on a page. A serious problem is that since the camera is mounted close to the pen-tip, the camera captures only a tiny portion of the page and the resultant image is under severe perspective distortion, resulting in lowering the retrieval accuracy. To solve this problem, we propose new geometrically invariant features as well as expansion techniques which increase the number of index features of either the database or the query images. From the experimental results, it has been found that the query expansion technique with features by combining affine and perspective invariants allows us the best performance that improves the accuracy of a baseline method more than 27%.
Koichi Kise, Megumi Chikano, Kazumasa Iwata, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi
Document Analysis Systems5
2010 Data-embedding pen: augmenting ink strokes with meta-information
abstract
In this paper we present the first operational version of the data-embedding pen. During writing a pattern, this pen produces an additional ink-dot sequence along the ink stroke of the pattern. The ink-dot sequence represents, for example, meta-information (such as the writer's name and the date of writing) and thus drastically increases the value of the handwriting on a physical paper. Since the information is placed on the paper, it can be extracted just by scanning or photographing the paper. There is no need to get access to any memory on the pen to recover the information. This is useful especially in multi-writer or multi-pen scenarios. The experiments using an encoding scheme and a decoding algorithm showed very promising results. For example, it was proved that we can embed 28 or more bits of information on simple handwritten patterns and decode them with a high reliability.
Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
Document Analysis Systems2
2010 How to Design Kansei Retrieval Systems?
Yaokai Feng, Seiichi Uchida
WAIM2
2009 Statistical Classification of Spatial Relationships among Mathematical Symbols
abstract
In this paper, a statistical decision method for automatic classification of spatial relationships between each adjacent pair is proposed. Each pair is composed of mathematical symbols and/or alphabetical characters. Special treatment of mathematical symbols with variable size is important. This classification is important to recognize an accurate structure analysis module of math OCR. Experimental results on a very large database showed that the proposed method worked well with an accuracy of 99.57% by two important geometric feature relative size and relative position.
Walaa Aly, Seiichi Uchida, Akio Fujiyoshi, Masakazu Suzuki
ICDAR2
2009 Syntactic Detection and Correction of Misrecognitions in Mathematical OCR
abstract
This paper proposes a syntactic method for detection and correction of misrecognized mathematical formulae for a practical mathematical OCR system. Linear monadic context-free tree grammar (LM-CFTG) is employed as a formal framework to define syntactically acceptable mathematical formulae.For the purpose of practical evaluation, a verification system is developed, and the effectiveness of the method is demonstrated by using the ground-truthed mathematical document database InftyCDB-1 and a misrecognition database newly constructed for this study.A satisfactory number of misrecognitions are detected and delivered to the correction process.
Akio Fujiyoshi, Masakazu Suzuki, Seiichi Uchida
ICDAR3
2009 Capturing Digital Ink as Retrieving Fragments of Document Images
abstract
This paper presents a new method of capturing digital ink for pen-based computing. Current technologies such as tablets, ultrasonic and the Anoto pens rely on special mechanisms for locating the pen tip,which result in limiting the applicability.Our proposal is to ease this problem --- a camera pen that allows us to write on ordinary paper for capturing digital ink. A document image retrieval method called LLAH is tuned to locate the pen tip efficiently and accurately on the coordinates of a document only by capturing its tiny fragment.In this paper, we report some results on captured digital ink as well as to evaluate their quality.
Kazumasa Iwata, Koichi Kise, Tomohiro Nakai, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi
ICDAR5
2009 Stochastic Model of Stroke Order Variation
abstract
A stochastic model of stroke order variation is proposed and applied to the stroke order free online Kanji character recognition.The proposed model is a hidden Markov model (HMM) with a special topology to represent all stroke order variations. A sequence of state transitions from the initial state to the final state of the model represents one stroke order and provides a probability of the stroke order.The distribution of the stroke order probability can be trained automatically by using an EM algorithm from a training set of on-line character patterns. Experimental results on large-scale test patterns showed that the proposed model could represent actual stroke order variations appropriately and improve recognition accuracy by penalizing incorrect stroke orders.
Yoshinori Katayama, Seiichi Uchida, Hiroaki Sakoe
ICDAR2
2009 Conspicuous Character Patterns
abstract
Detection of characters in scenery images is often a very difficult problem. Although many researchers have tackled this difficult problem and achieved a good performance, it is still difficult to suppress many false alarms and although missings. This paper investigates a conspicuous character pattern, which is a special pattern designed for easier detection. In order to have an example of the conspicuous character pattern, we select a character font with a larger distance from a non-character pattern distribution and, simultaneously, with a smaller distance from a character pattern distribution. Experimental results showed that the character font selected by this method is actually more conspicuous (i.e., detected more easily) than other fonts.
Seiichi Uchida, Ryoji Hattori, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
ICDAR1
2009 Hierarchical Decomposition of Handwriting Deformation Vector Field Using 2D Warping and Global/Local Affine Transformation
abstract
This paper addresses the basic problem of how to extract, describe, and evaluate handwriting deformation from not the statistical but the deterministic viewpoint. The key ideas are threefold. The first idea is to apply 2D warping to extraction of handwriting deformation vector field (DVF) between a pair of input and target images. The second idea is to hierarchically decompose the DVF by a parametric deformation model of global/local affine transformation. As a result, the DVF is expressed by a series of deformation components each of which is characterized by a window size of local affine transformation. The third idea is interrupting of the series of deformation components to obtain natural, reasonable handwriting deformation. Experiments using the handwritten numeral database IPTP CDROM1B show that 31.1% of the handwriting DVF is expressed by global affine transformation, and the subsequent few local affine transformations successfully discriminate natural handwriting deformation from unnatural one.
Toru Wakahara, Seiichi Uchida
ICDAR2
2008 A Large-Scale Analysis of Mathematical Expressions for an Accurate Understanding of Their Structure
abstract
A wide variety of mathematical expressions printed in scientific and technical reports can be recognized by analyzing the two-dimensional layout structure. In this paper, the position relation between adjacent characters is analyzed for the purpose of automatic discrimination between baseline, subscript, and superscript characters. This analyzing is one of the most important parts of structure analysis. The proposed method is very promising, as the results reached up to (99.76%) over a very large database by using distribution map. This distribution map is defined by two important features, i.e., relative size and relative position.
Walaa Aly, Seiichi Uchida, Masakazu Suzuki
Document Analysis Systems2
2008 Affine Invariant Recognition of Characters by Progressive Pruning
abstract
There are many problems to realize camera-based character recognition. One of the problems is that characters in scenes are often distorted by geometric transformations such as affine distortions. Although some methods that remove the affine distortions have been proposed, they cannot remove a rotation transformation of a character. Thus a skew angle of a character has to be determined by examining all the possible angles. However, this consumes quite a bit of time. In this paper, in order to reduce the processing time for an affine invariant recognition, we propose a set of affine invariant features and a new recognition scheme called "progressive pruning."' The progressive pruning gradually prunes less feasible categories and skew angles using multiple classifiers. We confirmed the progressive pruning with the affine invariant features reduced the processing time at least less than half without decreasing the recognition rate.
Akira Horimatsu, Ryo Niwa, Masakazu Iwamura, Koichi Kise, Seiichi Uchida, Shinichiro Omachi
Document Analysis Systems5
2008 Skew Estimation by Instances
abstract
This paper proposes a novel skew estimation method by instances. The instances to be learned (i.e., stored) are rotation invariants and a rotation variant for each character category. Using the instances, it is possible to estimate a skew angle of each individual character on a document. This fact implies that the proposed method can estimate the skew angle of a document where characters do not form long straight text lines. Thus, the proposed method will be applicable to various documents such as signboard images captured by a camera. Experimental evaluation using synthetic and real images revealed the expected robustness against various character layouts.
Seiichi Uchida, Megumi Sakai, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
Document Analysis Systems1
2007 Predictive DP Matching for On-Line Character Recognition
abstract
For on-line character recognition, predictive DP match- ing is proposed where two physically different features, co- ordinate features and directional features, are handled in a unified manner. For this unification, the distance of the directional features is converted into a distance of the co- ordinate features by a feature prediction technique. An ex- perimental result showed that the predictive DP matching could attain a recognition rate comparable to the rate by the conventional DP matching which requires the costly op- timization of the weight to balance the two features.
D. Baba, Seiichi Uchida, Hiroaki Sakoe
ICDAR2
2007 Non-Uniform Slant Correction for Handwritten Text Line Recognition
abstract
In this paper we apply a novel non-uniform slant correction preprocessing technique to improve the recognition of offline handwritten text lines. The local slant correction is expressed as a global optimisation problem of the sequence of local slant angles. This is different to conventional slant removal techniques that rely on the average slant angle. Experiments based on a state-of-the-art handwritten text line recogniser show a significant gain in word level accuracy for the investigated preprocessing methods.
Roman Bertolami, Seiichi Uchida, Matthias Zimmermann, Horst Bunke
ICDAR2
2007 Image Pixel Force Fields and their Application for Color Map Vectorisation
abstract
Pixel force field is a novel image representation where at each pixel a two-dimensional vector is defined for representing the circumstance of the pixel. The vector is oriented to the center of the region composed of vectors having the same qualitative property, such as color and gray-scale level. Using the pixel force field, that is, the orientation and the magnitude of the vector, many fundamental and specific image processing tasks can be solved. As examples of the tasks, the force field is applied to color image thinning, color image segmentation, and color map vectorisation.
V. Bucha, Sergey Ablameyko 0001, Seiichi Uchida
ICDAR3
2007 Extraction of Embedded Class Information from Universal Character Pattern
abstract
This paper is concerned with a universal pattern, which is defined as a character pattern designed to have high machine-readability. This universal pattern is a charac- ter pattern printed with stripes. The cross ratio calculated from the widths of the stripes represents the character class. Thus, if the boundaries of the stripes can be detected for measuring the widths, the class can be determined without ordinary recognition process. Furthermore, since the cross ratio is invariant to projective distortions, the correct class will be still determined under those distortions. This pa- per describes a practical scheme to recognize this universal pattern. The proposed scheme includes a novel algorithm to detect the stripe boundaries stably even from the universal pattern image contaminated by non-uniform lighting and noise. The algorithm is realized by a combination of a dy- namic programming-based optimal boundary detection and a finite state automaton which represents the property of the universal pattern. Experimental results showed the pro- posed scheme could recognize 99.6% of the universal pat- tern images which underwent heavy projective distortions and non-uniform lighting.
Seiichi Uchida, Megumi Sakai, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
ICDAR1
2006 Structural Analysis of Mathematical Formulae with Verification Based on Formula Description Grammar
Seiichi Toyota, Seiichi Uchida, Masakazu Suzuki
Document Analysis Systems2
2005 Dewarping of document image by global optimization
abstract
This paper proposes a novel dewarping technique for document images of bound volumes. This technique is a kind of model fitting techniques for estimating the warp of each text line by fitting some elastic curve model to the text line. Differing from conventional techniques, the proposed technique is applicable to document images including local irregularities such as formulae, short text lines, and figures, since the proposed technique dewarps whole document images by fitting splines while considering the global optimality that specifies the desirable relationship among the splines. The experimental results on several document images including the local irregularities indicated the effectiveness of the proposed technique. The experimental result also indicated the effectiveness of the vertical division of a document image into some partial document images for more accurate dewarping.
Hironori Ezaki, Seiichi Uchida, Akira Asano, Hiroaki Sakoe
ICDAR2
2005 Online Character Recognition Based on Elastic Matching and Quadratic Discrimination
abstract
We try to link elastic matching with a statistical discrimination framework to overcome the overfitting problem which often degrades the performance of elastic matching-based online character recognizers. In the proposed technique, elastic matching is used just as an extractor of a feature vector representing the difference between input and reference patterns. Then quadratic discrimination is performed under the assumption that the feature vector is governed by a Gaussian distribution. The result of a recognition experiment on UNIPEN database (Train-R01/V07, 1a) showed that the proposed technique can attain a high recognition rate (97.95%) and outperforms a recent elastic matching-based recognizer.
Hiroto Mitoma, Seiichi Uchida, Hiroaki Sakoe
ICDAR2
2005 Mosaicing-by-recognition: a technique for video-based text recognition
abstract
In this paper, a mosaicing-by-recognition technique is proposed where video mosaicing and text recognition are simultaneously and collaboratively optimized in a one-step manner. Specifically, multiple frames capturing a long text line are optimally concatenated with a guide of the text recognition framework. In this optimization process, rotation, scaling, vertical shift, and speed fluctuation, which often appear in video frames captured by hand-held cameras, are compensated. The optimization is performed by a DP-based algorithm. The results of experiments to evaluate not only the accuracy of text recognition but also that of video mosaicing indicate that the proposed technique is practical and can provide reasonable results in most cases.
Hiromitsu Miyazaki, Seiichi Uchida, Hiroaki Sakoe
ICDAR2
2005 An HMM Implementation for On-line Handwriting Recognition - Based on Pen-Coordinate Feature and Pen-Direction Feature
abstract
An on-line handwritten character recognition technique based on a new HMM is proposed. In the proposed HMM, not only pen-direction feature but also pen-coordinate feature are separately utilized for describing the shape variation of on-line characters accurately. Specifically speaking, the proposed HMM outputs a pen-coordinate feature at each inter-state transition and outputs a pen-direction feature at each intra-state transition, i.e., self-transition. Thus, each state of the proposed HMM can specify the starting position and the direction of a line segment by its incoming inter-state transition and intra-state transition, respectively. The results of recognition experiments on 10-stroke Chinese characters show that the proposed HMM outperforms the conventional HMM which does not use the pen-coordinate feature because of its non-stationarity.
Daiki Okumur, Seiichi Uchida, Hiroaki Sakoe
ICDAR2
2005 A Ground-Truthed Mathematical Character and Symbol Image Database
abstract
This paper describes the specifications for our ground-truthed mathematical character and symbol image database, called InftyCDB-1. The ground-truth of each character is composed of type, font, quality (touched/broken) and link (relative position), etc. The database includes all the characters and symbols of 467 pages of 30 articles on mathematics, and is organized so that it can be used as word image database or as mathematical formula image database. InftyCDB-1 is a public database that is freely usable for research and development purposes.
Masakazu Suzuki, Seiichi Uchida, Akihiro Nomura 0001
ICDAR2
2003 INFTY: an integrated OCR system for mathematical documents
abstract
An integrated OCR system for mathematical documents, called INFTY, is presented. INFTY consists of four procedures, i.e., layout analysis, character recognition, structure analysis of mathematical expressions, and manual error correction. In those procedures, several novel techniques are utilized for better recognition performance. Experimental results on about 500 pages of mathematical documents showed high character recognition rates on both mathematical expressions and ordinary texts, and sufficient performance on the structure analysis of the mathematical expressions.
Masakazu Suzuki, Fumikazu Tamari, Ryoji Fukuda, Seiichi Uchida, Toshihiro Kanahori
ACM Symposium on Document Engineering4
2003 Detection and Segmentation of Touching Characters in Mathematical Expression
abstract
A technique for the detection and the segmentation of touching characters in mathematical expressions is presented. In the detection stage, a connected component initially recognized into some category is judged as a candidate of touched characters if its feature values deviate from the standard feature values of the category. In the segmentation stage, two component characters of the candidate are decided by the comparison with touching character images synthesized from two single character images. Experimental results showed the effectiveness on the accuracy improvement of the recognition of mathematical expressions.
Akihiro Nomura 0001, Kazuyuki Michishita, Seiichi Uchida, Masakazu Suzuki
ICDAR3
2003 Handwritten character recognition using elastic matching based on a class-dependent deformation model
abstract
For handwritten character recognition, a new elastic image matching (EM) technique based on a class-dependent deformation model is proposed. In the deformation model, any deformation of a class is described by a linear combination of eigen-deformations, which are intrinsic deformation directions of the class. The eigen-deformations can be estimated statistically from the actual deformations of handwritten characters. Experimental results show that the proposed technique can attain higher recognition rates than conventional EM techniques based on class-independent deformation models. The results also show the superiority of the proposed technique over those conventional EM techniques in computational efficiency.
Seiichi Uchida, Hiroaki Sakoe
ICDAR1
2001 An Efficient Correlation Computation Method for Binary Images Based on Matrix Factorisation
abstract
A novel algorithm for complexity reduction in binary image processing, namely for computation of correlation between image and object template is proposed. This algorithm is based on direct computation of vector-matrix multiplication with utilisation of binary matrix factorisation approach. Comparison with other algorithms is given and it is shown that our approach allows to reduce time and complexity of this task.
Rykhard Bohush, S. Maltsev, Sergey Ablameyko 0001, Seiichi Uchida
ICDAR4
2001 Handwritten Character Recognition Using Piecewise Linear Two-Dimensional Warping
abstract
The effectiveness of piecewise linear 2D warping, a dynamic programming-based elastic image matching technique, in handwritten character recognition is investigated. The technique presented is capable of providing compensation for most variations in character patterns with tractable computation. The superiority of the present technique over several conventional 2D warping techniques in variation compensation is experimentally justified. Another comparison with monotonic and continuous 2D warping, a more flexible matching technique, reveals that the method presented takes far less computation than the latter, yet provides almost the same recognition accuracy for most categories.
Mohammad Asad Ronee, Seiichi Uchida, Hiroaki Sakoe
ICDAR2
2001 Nonuniform Slant Correction Using Dynamic Programming
abstract
Slant correction is an indispensable technique for handwritten word recognition systems. Conventional slant correction techniques estimate the average slant angle of component characters and then correct the slant uniformly. Thus these conventional techniques will perform successfully under the assumption that each word is written with a constant slant. However, it is more widely acceptable assumption that the slant angle fluctuates during writing a word. In this paper, a nonuniform slant correction technique is presented where the slant correction problem is formulated as an optimal estimation problem of local slant angles at all horizontal positions. The optimal estimation is governed by a criterion function and several constraints for the global and local validity of the local angles. The optimal local slant angles which maximize the criterion satisfying the constraints are searched for efficiently by a dynamic programming based algorithm. Experimental results show the advantageous characteristics of the present technique over the uniform slant correction techniques.
Seiichi Uchida, Eiji Taira, Hiroaki Sakoe
ICDAR1
1999 Handwritten Character Recognition using Monotonic and Continuous Two-dimensional Warping
abstract
In this paper, a handwritten character recognition experiment using a monotonic and continuous two-dimensional warping algorithm is reported. This warping algorithm is based on dynamic programming and searches for the optimal pixel-to-pixel mapping between given two images subject to two-dimensional monotonicity and continuity constraints. Experimental comparisons with rigid matching and local perturbation show the performance superiority of the monotonic and continuous warping in character recognition.
Seiichi Uchida, Hiroaki Sakoe
ICDAR1