Shinichiro Omachi

dblp:o/ShinichiroOmachi · also Shin'ichiro Omachi · DBLP profile ↗
← Back
47ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-7706-9995ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 11Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 A real-driven image synthesis framework for adapting pre-trained scene text detectors
Yaohou Fan, Zhengmi Tang, Tomo Miyazaki, Yongsong Huang, Shinichiro Omachi
Image Vis. Comput.5
2026 Generative image compression by prediction of optimal realism levels
abstract
• A novel LIC model uses a predictor that automatically determines realism levels • We propose a two-stage training framework for stable training of the LIC model • Optimal realism levels determine the quality of distortion and perception of the image • We achieved significant results in perceptual metrics Learned image compression (LIC) methods have demonstrated high-quality reconstruction, and even surpassed the traditional codec algorithms at rate-distortion performance. However, reconstructing natural and realistic textures remains a challenge. The main problem is the distortion-realism trade-off. Realism deteriorates quickly if we minimize distortion, and vice versa. The existing methods manually determine and apply a realism level, which is a degree of the trade-off, to an entire image, resulting in time-consuming training and suboptimal quality. In this paper, we propose a novel LIC model equipped with a predictor that automatically determines realism levels for regions of the image. Thus, we reconstruct images using the optimal realism levels. Moreover, we developed a two-stage training framework for stable training of the proposed model. The experimental results show that the proposed model achieved significant perceptual quality while the PSNR is comparable to the existing methods.
Hayato Fukuoka, Shoma Iwai, Tomo Miyazaki, Shinichiro Omachi
Pattern Recognit.4
2026 Multi-masking strategies for self-supervised Low- and High-level text representation learning
Zhengmi Tang, Yuto Mitsui, Tomo Miyazaki, Shinichiro Omachi
Pattern Recognit.4
2025 Scene Text Reconstructor: A Contextual-Aware Masking Framework for Pre-training Text Detectors
Yaohou Fan, Tomo Miyazaki, Zhengmi Tang, Yongsong Huang, Shinichiro Omachi
ICONIP (2)6
2025 VQ-STE: Scene text erasing with mask refinement and vector-quantized texture dictionary
Zhengmi Tang, Tomo Miyazaki, Yongsong Huang, Jonathan Pradana Mailoa, Shinichiro Omachi
Knowl. Based Syst.6
2025 Texture and noise dual adaptation for infrared image super-resolution
abstract
Recent efforts have explored leveraging visible light images to enrich texture details in infrared (IR) super-resolution. However, this direct adaptation approach often becomes a double-edged sword, as it improves texture at the cost of introducing noise and blurring artifacts. Such imperfections are inherent in the spatial domain of visible images and are accentuated during the imaging process. Enhancing IR image quality by integrating rich texture details from visible images, while minimizing noise transfer, presents a challenging research avenue. To address these challenges, we propose the Texture and Noise Dual Adaptation SRGAN (DASRGAN), an innovative framework specifically engineered for robust IR super-resolution model adaptation. DASRGAN operates on the synergy of two key components: (1) Texture-Oriented Adaptation (TOA) to refine texture details meticulously, and (2) Noise-Oriented Adaptation (NOA), dedicated to minimizing noise transfer. Specifically, TOA uniquely integrates a specialized discriminator, incorporating a prior extraction branch, and employs a Sobel-guided adversarial loss to align texture distributions effectively. Concurrently, NOA utilizes a noise adversarial loss to distinctly separate the generative and Gaussian noise pattern distributions during adversarial training. Our extensive experiments confirm DASRGAN’s superiority. Comparative analyses against leading methods across multiple benchmarks and upsampling factors reveal that DASRGAN sets new state-of-the-art performance standards. Code are available at https://github.com/yongsongH/DASRGAN . • DASRGAN improves infrared image super-resolution using visible textures and noise control through dual adaptations. • Texture-Oriented Adaptation uses Sobel-based discriminator and texture alignment loss for detail enhancement. • Noise-Oriented Adaptation employs domain-specific loss to suppress noise propagation between modalities. • DASRGAN achieves state-of-the-art performance in infrared super-resolution benchmarks across metrics.
Yongsong Huang, Tomo Miyazaki, Xiaofeng Liu 0001, Yafei Dong, Shinichiro Omachi
Pattern Recognit.5
2025 IRSRMamba: Infrared Image Super-Resolution via Mamba-Based Wavelet Transform Feature Modulation Model
abstract
Infrared image super-resolution (IRSR) is challenging due to weak structures and textures. While Mamba-based state-space models (SSMs) efficiently model long-range dependencies, their inherent block-wise processing disrupts spatial consistency, limiting direct IRSR applicability. We propose IRSRMamba, a novel Mamba-based framework overcoming this limitation via tailored structural and textural preservation. Integrated into a Mamba backbone, our key innovations are: 1) Wavelet Transform Feature Modulation (WTFM), enhancing multi-scale frequency-aware feature extraction to mitigate block-induced coherence loss; and 2) an SSMs-based Semantic Consistency Loss, enforcing cross-block alignment to restore fragmented context. IRSRMamba achieves superior global-local fusion, structural coherence, and fine-detail preservation. Experiments show state-of-the-art PSNR, SSIM, and perceptual quality on IR benchmarks, as well as robust generalization to remote sensing. This work establishes Mamba-based architectures as highly promising for high-fidelity IR image enhancement. Code is available at https://github.com/yongsongH/IRSRMamba.
Yongsong Huang, Tomo Miyazaki, Xiaofeng Liu 0001, Shinichiro Omachi
IEEE Trans. Geosci. Remote. Sens.4
2024 Layout-Corrector: Alleviating Layout Sticking Phenomenon in Discrete Diffusion Model
Shoma Iwai, Atsuki Osanai, Shunsuke Kitada, Shinichiro Omachi
ECCV (34)4
2024 DiCTI: Diffusion-based Clothing Designer via Text-guided Input
abstract
Recent developments in deep generative models have opened up a wide range of opportunities for image synthesis, leading to significant changes in various creative fields, including the fashion industry. While numerous methods have been proposed to benefit buyers, particularly in virtual try-on applications, there has been relatively less focus on facilitating fast prototyping for designers and customers seeking to order new designs. To address this gap, we introduce DiCTI (Diffusion-based Clothing Designer via Text-guided Input), a straightforward yet highly effective approach that allows designers to quickly visualize fashion-related ideas using text inputs only. Given an image of a person and a description of the desired garments as input, DiCTI automatically generates multiple high-resolution, photorealistic images that capture the expressed semantics. By leveraging a powerful diffusion-based inpainting model conditioned on text inputs, DiCTI is able to synthesize convincing, high-quality images with varied clothing designs that viably follow the provided text descriptions, while being able to process very diverse and challenging inputs, captured in completely unconstrained settings. We evaluate DiCTI in comprehensive experiments on two different datasets (VITON-HD and Fashionpedia) and in comparison to the state-of-the-art (SoTa). The results of our experiments show that DiCTI convincingly outperforms the SoTA competitor in generating higher quality images with more elaborate garments and superior text prompt adherence, both according to standard quantitative evaluation measures and human ratings, generated as part of a user study. The source code of DiCTI will be made publicly available.
Ajda Lampe, Julija Stopar, Deepak Kumar Jain 0001, Shinichiro Omachi, Peter Peer, Vitomir Struc
FG4
2024 Controlling Rate, Distortion, and Realism: Towards a Single Comprehensive Neural Image Compression Model
abstract
In recent years, neural network-driven image compression (NIC) has gained significant attention. Some works adopt deep generative models such as GANs and diffusion models to enhance perceptual quality (realism). A critical obstacle of these generative NIC methods is that each model is optimized for a single bit rate. Consequently, multiple models are required to compress images to different bit rates, which is impractical for real-world applications. To tackle this issue, we propose a variable-rate generative NIC model. Specifically, we explore several discriminator designs tailored for the variable-rate approach and introduce a novel adversarial loss. Moreover, by incorporating the newly proposed multi-realism technique, our method allows the users to adjust the bit rate, distortion, and realism with a single model, achieving ultra-controllability. Unlike existing variable-rate generative NIC models, our method matches or surpasses the performance of state-of-the-art single-rate generative NIC models while covering a wide range of bit rates using just one model.
Shoma Iwai, Tomo Miyazaki, Shinichiro Omachi
WACV3
2024 Japanese historical character recognition by focusing on character parts
abstract
Japanese historical documents provide valuable information. Character recognition is a critical technology for the digitalization of historical documents. Sample imbalance is a significant obstacle in recognizing Japanese historical characters, kuzushiji. Thousands of kuzushiji only have less than a few samples. Thus, recognition performance deteriorates greatly in kuzushiji with a few samples. In this study, we propose a framework for transferring knowledge of character parts from font to kuzushiji. The pretraining learns character parts from synthesized font images. However, fine-tuning to kuzushiji is more complex. We propose calculating a mean squared error loss between feature vectors of kuzushiji and font images, resulting in consistent feature vectors in kuzushiji and font. Consequently, we can perform zero-shot recognition for kuzushiji using the font images of zero-sampled kuzushiji. The experimental results show that the proposed method recognized zero-sampled kuzushiji at approximately 48% accuracy. Consequently, we significantly expand the number of recognizable kuzushiji.
Takuru Ishikawa, Tomo Miyazaki, Shinichiro Omachi
Pattern Recognit.3
2023 Deep image compression using scene text quality assessment
Shohei Uchigasaki, Tomo Miyazaki, Shinichiro Omachi
Pattern Recognit.3
2023 A Scene-Text Synthesis Engine Achieved Through Learning From Decomposed Real-World Data
abstract
Scene-text image synthesis techniques that aim to naturally compose text instances on background scene images are very appealing for training deep neural networks due to their ability to provide accurate and comprehensive annotation information. Prior studies have explored generating synthetic text images on two-dimensional and three-dimensional surfaces using rules derived from real-world observations. Some of these studies have proposed generating scene-text images through learning; however, owing to the absence of a suitable training dataset, unsupervised frameworks have been explored to learn from existing real-world data, which might not yield reliable performance. To ease this dilemma and facilitate research on learning-based scene text synthesis, we introduce DecompST, a real-world dataset prepared from some public benchmarks, containing three types of annotations: quadrilateral-level BBoxes, stroke-level text masks, and text-erased images. Leveraging the DecompST dataset, we propose a Learning-Based Text Synthesis engine (LBTS) that includes a text location proposal network (TLPNet) and a text appearance adaptation network (TAANet). TLPNet first predicts the suitable regions for text embedding, after which TAANet adaptively adjusts the geometry and color of the text instance to match the background context. After training, those networks can be integrated and utilized to generate the synthetic dataset for scene text analysis tasks. Comprehensive experiments were conducted to validate the effectiveness of the proposed LBTS along with existing methods, and the experimental results indicate the proposed LBTS can generate better pretraining data for scene text detectors. Our dataset and code are made available at: https://github.com/iiclab/DecompST.
Zhengmi Tang, Tomo Miyazaki, Shinichiro Omachi
IEEE Trans. Image Process.3
2021 Stroke-Based Scene Text Erasing Using Synthetic Data for Training
abstract
Scene text erasing, which replaces text regions with reasonable content in natural images, has drawn significant attention in the computer vision community in recent years. There are two potential subtasks in scene text erasing: text detection and image inpainting. Both subtasks require considerable data to achieve better performance; however, the lack of a large-scale real-world scene-text removal dataset does not allow existing methods to realize their potential. To compensate for the lack of pairwise real-world data, we made considerable use of synthetic text after additional enhancement and subsequently trained our model only on the dataset generated by the improved synthetic text engine. Our proposed network contains a stroke mask prediction module and background inpainting module that can extract the text stroke as a relatively small hole from the cropped text image to maintain more background content for better inpainting results. This model can partially erase text instances in a scene image with a bounding box or work with an existing scene-text detector for automatic scene text erasing. The experimental results from the qualitative and quantitative evaluation on the SCUT-Syn, ICDAR2013, and SCUT-EnsText datasets demonstrate that our method significantly outperforms existing state-of-the-art methods even when they are trained on real-world data.
Zhengmi Tang, Tomo Miyazaki, Yoshihiro Sugaya, Shinichiro Omachi
IEEE Trans. Image Process.4
2020 Fidelity-Controllable Extreme Image Compression with Generative Adversarial Networks
abstract
We propose a GAN-based image compression method working at extremely low bitrates below 0.1bpp. Most existing learned image compression methods suffer from blur at extremely low bitrates. Although GAN can help to reconstruct sharp images, there are two drawbacks. First, GAN makes training unstable. Second, the reconstructions often contain unpleasing noise or artifacts. To address both of the drawbacks, our method adopts two-stage training and network interpolation. The two-stage training is effective to stabilize the training. Moreover, the network interpolation utilizes the models in both stages and reduces undesirable noise and artifacts, while maintaining important edges. Hence, we can control the trade-off between perceptual quality and fidelity without re-training models. The experimental results show that our model can reconstruct high quality images. Furthermore, our user study confirms that our reconstructions are preferable to state-of-the-art GAN-based image compression model. Our source code is available at https://github.com/iwa-shi/fidelity_controllable_compression.
Shoma Iwai, Tomo Miyazaki, Yoshihiro Sugaya, Shinichiro Omachi
ICPR4
2018 Mackerel Classification Using Global and Local Features
abstract
Blue and chub mackerels are often caught at the same time, and their market prices are different. Thus, humans need to classify them by their hands. Therefore, the demand for automatic sorting machines of these mackerels using image processing technique is increasing. Classification for blue and chub mackerels is a challenging task due to their quite similar appearance. In this paper, we propose a method for classifying blue and chub mackerels by image processing technique. The difference of the two mackerels appears on the whole and the local specific parts. Therefore, we use global and local features that are extracted from the whole and the specific part in fish images, respectively. Experimental results showed that the proposed method is superior to the methods using only global or local feature.
Yoshito Nagaoka, Tomo Miyazaki, Yoshihiro Sugaya, Shinichiro Omachi
ETFA4
2017 Glyph-Based Data Augmentation for Accurate Kanji Character Recognition
abstract
In this paper, we address a problem of data augmentation for character recognition. Particularly, we focus on incorporating variation in glyph into data augmentation of character images, which is a simple approach for data augmentation. Generally, existing methods increase data size by distorting images, whereas the proposed method applies noise injection into glyphs, resulting in data with radical variation in glyph. The proposed method exploits public database of glyphs for kanji and augments glyphs by injecting noise into glyphs. Then, we generate images of kanji automatically by deploying stroke images on the augmented glyphs. We carried out experiments for kanji character recognition using augmented data. The results show the effectiveness of the proposed method.
Kenichiro Ofusa, Tomo Miyazaki, Yoshihiro Sugaya, Shinichiro Omachi
ICDAR4
2017 Analysis of floor map image in information board for indoor navigation
abstract
Various indoor navigation methods have been developed recently, but digitalized data of indoor map is not always available. Therefore, an indoor navigation framework using an image of information board has been proposed. In this method, the process to extract map regions from the image of an information board is necessary to be done by hands beforehand, and the process to estimate passageway regions is important because its information is used in map matching. However, the method of passageway discrimination is very heuristic, which is intended for a specific type of floor maps. Therefore, in this paper, we propose a semi-automatic method to extract map regions from the image of information board with simple user's operation. We use GrabCut method and Snakes method for the extraction method. In GrabCut method, we detect closed regions to prevent the degradation of accuracy when conducting GrabCut to the downsizing image. The proposed method can extract a map region with few deficits in short calculation time. In addition, we propose a machine learning based method to classify passageway regions and other regions from a segment image. We confirmed that the proposed methods are effective and promising by experiments.
Tomoya Honto, Yoshihiro Sugaya, Tomo Miyazaki, Shinichiro Omachi
IPIN4
2016 Graph model boosting for structural data recognition
abstract
This paper presents a novel method for structural data recognition using a large number of graph models. Broadly, existing methods for structural data recognition have two crucial problems: 1) only a single model is used to capture structural variation, 2) naive classification rules are used, such as nearest neighbor method. In this paper, we propose to strengthen both capturing structural variation and the classification ability. The proposed method constructs a large number of graph models and trains decision tree classifiers with the models. There are two contributions of this paper. The first contribution is a novel graph model which can be constructed by straightforward calculation. This calculation enables us to construct many models in feasible time. The second contribution is a novel approach to capture structural variation. We construct a large number of our models in a boosting framework so that we can capture structural variation comprehensively. Consequently, we are able to perform structural data recognition with the powerful classification ability and comprehensive structural variation. In experiments, we show that the proposed method achieves significant results and outperforms the existing methods.
Tomo Miyazaki, Shinichiro Omachi
ICPR2
2014 Recovery and localization of handwritings by a camera-pen based on tracking and document image retrieval
Megumi Chikano, Koichi Kise, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi
Pattern Recognit. Lett.5
2014 More than ink - Realization of a data-embedding pen
Marcus Liwicki, Seiichi Uchida, Akira Yoshida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
Pattern Recognit. Lett.5
2013 The Reading-Life Log - Technologies to Recognize Texts That We Read
abstract
Reading life log is a type of techniques to automatically and unconsciously record people's reading intentions, interests and habits. Besides, it can also serve as various assistants in our daily life. In this paper, a reading-life log system is implemented by a head-mounted and unobtrusive video camera with a high resolution and a high shutter speed. We utilize DP matching, and propose a text-based frame mosaicing method to integrate multiple frames in a clip. The developed system is tested in the various environments indoor and outdoor. The experimental results show that our system can provide reliable outputs with respect to the most correct responses. The infrequent misregistration between lines also indicates the feasibility and validity of the text-based frame mosaicing.
Takashi Kimura, Rong Huang 0003, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
ICDAR5
2012 A Study on Caption Recognition for Multi-color Characters on Complex Background
abstract
We propose a caption recognition method for multicolor characters on complex background. Caption characters are used for an efficient search on a large amount of recorded TV programs. In the caption character recognition, the caption appearance section and the area is extracted, the character strokes are extracted from the area, and recognized. This paper focuses on caption character strokes extraction and recognition for multi-color characters on complex background which is a very difficult task for the conventional methods. The proposed method extracts decomposed binary images from input color caption image by color clustering. Then character candidates that are composed of combination of connect components are extracted by using recognition certainty. Finally, characters are selected by beyond-color Dynamic Programming method in which weight on recognition certainty and character alignment are used. In the character recognition evaluation of one-line multi-color character string on a complex background, a great improvement was achieved from a conventional technique that can recognize only one-color characters on complex background image.
Yutaka Katsuyama, Akihiro Minagawa, Yoshinobu Hotta, Jun Sun 0004, Shinichiro Omachi
ISM5
2012 A Statistical Analysis on Operation Scheduling for an Energy Network Project
abstract
Distributed power generation, using renewable energy, has been attracting attention to cope with global environment issues; a microgrid is a promising configuration for distributed power generation. To augment the stability and efficiency of the microgrid, an intelligent control, which considers the restrictions and characteristics of each unit, is indispensable. It can be achieved by constructing an efficient operation schedule for each power plant in the microgrid, depending on energy demand, and predicting passive power generation. The operation scheduling is regarded as a constrained optimization problem, which must have nonlinear characteristics in case of actual systems. Although several methods using metaheuristic optimization have been proposed, it would be trapped into a local minimum in some cases. In this paper, we statistically analyze operation schedules, computed for an actual power network of the demonstrative project. In addition, we conduct an investigation of the relationship between the input parameter space and the solution space, which can be exploited to obtain more appropriate initial solutions leading to better and faster converging solutions.
Yoshihiro Sugaya, Shinichiro Omachi, Akira Takeuchi, Yousuke Nozaki
IEEE Trans. Parallel Distributed Syst.2
2011 Reliable Online Stroke Recovery from Offline Data with the Data-Embedding Pen
abstract
In this paper we propose a complete system for online stroke recovery from offline data. The key idea of our approach is to use a novel pen device which is able to embed meta information into the ink during writing the strokes. This pen-device overcomes the need to get access to any memory on the pen when trying to recover the information, which is especially useful in multi-writer or multi-pen scenarios. The actual data-embedding is achieved by an additional ink dot sequence along a handwritten pattern during writing. We design the ink-dot sequence in such a way that it is possible to retrieve the writing direction from a scanned image. Furthermore, we propose novel processing steps in order to retrieve the original writing direction and finally the embedded data. In our experiments we show that we can reliably recover the writing direction of various patterns. Our system is able to determine the writing direction of straight lines, simple patterns with crossings (e.g., "x" and "II"), and even more complex patterns like handwritten words and symbols.
Marcus Liwicki, Akira Yoshida, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
ICDAR5
2011 Handwriting on Paper as a Cybermedium
Akira Yoshida, Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
KES (4)5
2010 Expansion of queries and databases for improving the retrieval accuracy of document portions: an application to a camera-pen system
abstract
This paper presents a method of improving the accuracy of document image retrieval focusing on the application to a camera-pen system. In a camera-pen system, document image retrieval is employed for locating the pen-tip position on a page. A serious problem is that since the camera is mounted close to the pen-tip, the camera captures only a tiny portion of the page and the resultant image is under severe perspective distortion, resulting in lowering the retrieval accuracy. To solve this problem, we propose new geometrically invariant features as well as expansion techniques which increase the number of index features of either the database or the query images. From the experimental results, it has been found that the query expansion technique with features by combining affine and perspective invariants allows us the best performance that improves the accuracy of a baseline method more than 27%.
Koichi Kise, Megumi Chikano, Kazumasa Iwata, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi
Document Analysis Systems6
2010 Data-embedding pen: augmenting ink strokes with meta-information
abstract
In this paper we present the first operational version of the data-embedding pen. During writing a pattern, this pen produces an additional ink-dot sequence along the ink stroke of the pattern. The ink-dot sequence represents, for example, meta-information (such as the writer's name and the date of writing) and thus drastically increases the value of the handwriting on a physical paper. Since the information is placed on the paper, it can be extracted just by scanning or photographing the paper. There is no need to get access to any memory on the pen to recover the information. This is useful especially in multi-writer or multi-pen scenarios. The experiments using an encoding scheme and a decoding algorithm showed very promising results. For example, it was proved that we can embed 28 or more bits of information on simple handwritten patterns and decode them with a high reliability.
Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
Document Analysis Systems4
2010 Tracking and Retrieval of Pen Tip Positions for an Intelligent Camera Pen
abstract
This paper presents a method of recovering digital ink for an intelligent camera pen, which is characterized by the functions that (1) it works on ordinary paper and (2) if an electronic document is printed on the paper the recovered digital ink is associated with the document. Two technologies called paper fingerprint and document image retrieval are integrated for realizing the above functions. The key of the integration is the introduction of image mosaicing and fast retrieval of previously seen fingerprints based on hashing of SURF local features. From the experimental results of 50 handwritings, we have confirmed that the proposed method is effective to recover and locate the digital ink from the handwriting on a physical paper.
Kazumasa Iwata, Koichi Kise, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi
ICFHR5
2010 Embedding Meta-Information in Handwriting -- Reed-Solomon for Reliable Error Correction
abstract
In this paper a more compact and more reliable coding scheme for the data-embedding pen is proposed. The data-embedding pen produces an additional ink-dot sequence along a handwritten pattern during writing. The ink-dot sequence represents, for example, meta-information (such as the writer's name and the date of writing) and thus drastically increases the value of the handwriting on a physical paper. There is no need to get access to any memory on the pen to recover the information, which is especially useful in multi-writer or multi-pen scenarios. In this paper we focus on the compactness of the encoded information. The aim of this paper is to encode as much information as possible in short stroke sequences. In our experiments we show that we can embed more information in shorter strokes than in previous work. In straight lines as short as 5 cm, 32 bits can successfully be embedded. Furthermore, the new encoding scheme also works reliably on more complex patterns.
Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
ICFHR4
2009 Capturing Digital Ink as Retrieving Fragments of Document Images
abstract
This paper presents a new method of capturing digital ink for pen-based computing. Current technologies such as tablets, ultrasonic and the Anoto pens rely on special mechanisms for locating the pen tip,which result in limiting the applicability.Our proposal is to ease this problem --- a camera pen that allows us to write on ordinary paper for capturing digital ink. A document image retrieval method called LLAH is tuned to locate the pen tip efficiently and accurately on the coordinates of a document only by capturing its tiny fragment.In this paper, we report some results on captured digital ink as well as to evaluate their quality.
Kazumasa Iwata, Koichi Kise, Tomohiro Nakai, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi
ICDAR6
2009 Conspicuous Character Patterns
abstract
Detection of characters in scenery images is often a very difficult problem. Although many researchers have tackled this difficult problem and achieved a good performance, it is still difficult to suppress many false alarms and although missings. This paper investigates a conspicuous character pattern, which is a special pattern designed for easier detection. In order to have an example of the conspicuous character pattern, we select a character font with a larger distance from a non-character pattern distribution and, simultaneously, with a smaller distance from a character pattern distribution. Experimental results showed that the character font selected by this method is actually more conspicuous (i.e., detected more easily) than other fonts.
Seiichi Uchida, Ryoji Hattori, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
ICDAR4
2009 Variational Bayesian Mixture Model on a Subspace of Exponential Family Distributions
abstract
Exponential principal component analysis (e-PCA) has been proposed to reduce the dimension of the parameters of probability distributions using Kullback information as a distance between two distributions. It also provides a framework for dealing with various data types such as binary and integer for which the Gaussian assumption on the data distribution is inappropriate. In this paper, we introduce a latent variable model for the e-PCA. Assuming the discrete distribution on the latent variable leads to mixture models with constraint on their parameters. This provides a framework for clustering on the lower dimensional subspace of exponential family distributions. We derive a learning algorithm for those mixture models based on the variational Bayes (VB) method. Although intractable integration is required to implement the algorithm for a subspace, an approximation technique using Laplace's method allows us to carry out clustering on an arbitrary subspace. Combined with the estimation of the subspace, the resulting algorithm performs simultaneous dimensionality reduction and clustering. Numerical experiments on synthetic and real data demonstrate its effectiveness for extracting the structures of data as a visualization technique and its high generalization ability as a density estimation model.
Kazuho Watanabe, Shotaro Akaho, Shinichiro Omachi, Masato Okada
IEEE Trans. Neural Networks3
2008 Affine Invariant Recognition of Characters by Progressive Pruning
abstract
There are many problems to realize camera-based character recognition. One of the problems is that characters in scenes are often distorted by geometric transformations such as affine distortions. Although some methods that remove the affine distortions have been proposed, they cannot remove a rotation transformation of a character. Thus a skew angle of a character has to be determined by examining all the possible angles. However, this consumes quite a bit of time. In this paper, in order to reduce the processing time for an affine invariant recognition, we propose a set of affine invariant features and a new recognition scheme called "progressive pruning."' The progressive pruning gradually prunes less feasible categories and skew angles using multiple classifiers. We confirmed the progressive pruning with the affine invariant features reduced the processing time at least less than half without decreasing the recognition rate.
Akira Horimatsu, Ryo Niwa, Masakazu Iwamura, Koichi Kise, Seiichi Uchida, Shinichiro Omachi
Document Analysis Systems6
2008 Skew Estimation by Instances
abstract
This paper proposes a novel skew estimation method by instances. The instances to be learned (i.e., stored) are rotation invariants and a rotation variant for each character category. Using the instances, it is possible to estimate a skew angle of each individual character on a document. This fact implies that the proposed method can estimate the skew angle of a document where characters do not form long straight text lines. Thus, the proposed method will be applicable to various documents such as signboard images captured by a camera. Experimental evaluation using synthetic and real images revealed the expected robustness against various character layouts.
Seiichi Uchida, Megumi Sakai, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
Document Analysis Systems4
2007 Extraction of Embedded Class Information from Universal Character Pattern
abstract
This paper is concerned with a universal pattern, which is defined as a character pattern designed to have high machine-readability. This universal pattern is a charac- ter pattern printed with stripes. The cross ratio calculated from the widths of the stripes represents the character class. Thus, if the boundaries of the stripes can be detected for measuring the widths, the class can be determined without ordinary recognition process. Furthermore, since the cross ratio is invariant to projective distortions, the correct class will be still determined under those distortions. This pa- per describes a practical scheme to recognize this universal pattern. The proposed scheme includes a novel algorithm to detect the stripe boundaries stably even from the universal pattern image contaminated by non-uniform lighting and noise. The algorithm is realized by a combination of a dy- namic programming-based optimal boundary detection and a finite state automaton which represents the property of the universal pattern. Experimental results showed the pro- posed scheme could recognize 99.6% of the universal pat- tern images which underwent heavy projective distortions and non-uniform lighting.
Seiichi Uchida, Megumi Sakai, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise
ICDAR4
2007 Fast Template Matching With Polynomials
abstract
Template matching is widely used for many applications in image and signal processing. This paper proposes a novel template matching algorithm, called algebraic template matching. Given a template and an input image, algebraic template matching efficiently calculates similarities between the template and the partial images of the input image, for various widths and heights. The partial image most similar to the template image is detected from the input image for any location, width, and height. In the proposed algorithm, a polynomial that approximates the template image is used to match the input image instead of the template image. The proposed algorithm is effective especially when the width and height of the template image differ from the partial image to be matched. An algorithm using the Legendre polynomial is proposed for efficient approximation of the template image. This algorithm not only reduces computational costs, but also improves the quality of the approximated image. It is shown theoretically and experimentally that the computational cost of the proposed algorithm is much smaller than the existing methods.
Shinichiro Omachi, Masako Omachi
IEEE Trans. Image Process.1
2005 Isolated Character Recognition by Searching Feature Points
abstract
Conventional segmentation technique cannot extract difficult characters such as an isolated character and a touching character. In this paper, we propose a novel character recognition method which executes segmentation and recognition simultaneously. This method enables us to extract and recognize such difficult characters. The effectiveness of the proposed method is confirmed by experiments.
Masakazu Iwamura, Kazuya Negishi, Shinichiro Omachi, Hirotomo Aso
ICDAR3
2001 Structure Extraction from Decorated Characters Using Multiscale Images
abstract
Decorated characters are widely used in various documents. Practical optical character reader is required to deal with not only common fonts but also complex designed fonts. However, since the appearances of decorated characters are complicated, most general character recognition systems cannot give good performances on decorated characters. In this paper, an algorithm that can extract character's essential structure from a decorated character is proposed. This algorithm is applied in preprocessing of character recognition. The proposed algorithm consists of three procedures: global structure extraction, interpolation of structure and smoothing. By using multiscale images, topographical features, such as ridges and ravines are detected for structure extraction. Ridges are used for extracting global structure and ravines are used for interpolation. Experimental results show character structures can be clearly extracted from very complex decorated characters.
Shinichiro Omachi, Masaki Inoue, Hirotomo Aso
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 A Parallel Architecture for Quadtree-based Fractal Image Coding
abstract
This paper proposes a parallel architecture for quadtree-based fractal image coding. This architecture is capable of performing the fractal image coding based on quadtree partitioning without the external memory for the fixed domain pool. Since a large domain block consists of small domain blocks, the calculations of distortion for all kinds of domain blocks are performed by the summation of the distortions for the maximum-depth domain pool which is extracted from the smallest range blocks of the neighbor processors. Fast comparison module is proposed for this architecture. This module can compute the distortions between range blocks and their eight isometric transformations by one full rotation around the center.
Shinhaeng Lee, Shinichiro Omachi, Hirotomo Aso
ICPP2
2000 A Modification of Eigenvalues to Compensate Estimation Errors of Eigenvectors
abstract
In statistical pattern recognition, parameters of distributions are usually estimated from training samples. It is well known that shortage of training samples causes estimation errors which reduce recognition accuracy. By studying estimation errors of eigenvalues, various methods of avoiding recognition accuracy reduction have been proposed. However, estimation errors of eigenvectors have not been considered enough. In this paper, we investigate estimation errors of eigenvectors to show these errors are another factor of recognition performance reduction. We propose a new method for modifying eigenvalues in order to reduce bad influence caused by estimation errors of eigenvectors. Effectiveness of the method is shown by experimental results.
Masakazu Iwamura, Shinichiro Omachi, Hirotomo Aso
ICPR2
2000 Precise Hand-printed Character Recognition Using Elastic Models via Nonlinear Transformation
abstract
Distorted character recognition is a difficult but in-evitable problem in hand-printed character recognition. In this paper, we propose a character recognition method us-ing elastic models for recognizing cursive characters with intricate structure. The models are fitted to unknown in-put patterns by applying the EM algorithm to minimize a measure of fittness. To avoid falling into local minima, mul-tiresolutional approach is introduced. Moreover, nonlinear transformation is adopted to realize more flexible matching. Experiments performed on Japanese characters show effec-tiveness of the proposed method. 1.
Tsuyoshi Kato, Shinichiro Omachi, Hirotomo Aso
ICPR2
2000 Structure Extraction from Various Kinds of Decorated Characters Using Multi-Scale Images
abstract
Decorated characters are widely used in various documents. Practical optical character reader is required to deal with not only common fonts but also complex designed fonts. However, since appearances of decorated characters are complicated, most general character recognition systems cannot give good performances on decorated characters. In this paper, an algorithm that can extract character's essential structure from a decorated character is proposed. This algorithm is applied in preprocessing of character recognition. The proposed algorithm consists of three parts: global structure extraction, interpolation of structure, and smoothing. By using multi-scale images, topographical features such as ridges and ravines are detected for structure extraction. Ridges are used for extracting global structure, and ravines are used for interpolation. Experimental results show clear character structures are extracted from very complex decorated characters.
Shinichiro Omachi, Masaki Inoue, Hirotomo Aso
ICPR1
2000 Two-Stage Computational Cost Reduction Algorithm Based on Mahalanobis Distance Approximations
abstract
For many pattern recognition methods, high recognition accuracy is obtained at very high expense of computational cost. In this paper, a new algorithm that reduces the computational cost for calculating discriminant function is proposed. This algorithm consists of two stages which are feature vector. Division and dimensional reduction. The processing of feature division is based on characteristic of covariance matrix. The dimensional reduction in the second stage is done by an approximation of the Mahalanobis distance. Compared with the well-known dimensional reduction method of K-L expansion, experimental results show the proposed algorithm not only reduces the computational cost but also improves the recognition accuracy.
Shinichiro Omachi, Nei Kato, Hirotomo Aso, Shunichi Kono, Tasuku Takagi
ICPR2
2000 A Noise-Adaptive Discriminant Function and Its Application to Blurred Machine-Printed Kanji Recognition
abstract
Accurate recognition of blurred images is a practical but previously mostly overlooked problem. In the paper, we quantify the level of noise in blurred images and propose a modification of discriminant functions that adapts to the level of noise. Experimental results indicate that the proposed method actually enhances the existing statistical methods and has impressive ability to recognize blurred image patterns.
Shinichiro Omachi, Hirotomo Aso
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 A Handwritten Character Recognition System Using Directional Element Feature and Asymmetric Mahalanobis Distance
abstract
This paper presents a precise system for handwritten Chinese and Japanese character recognition. Before extracting directional element feature (DEF) from each character image, transformation based on partial inclination detection (TPID) is used to reduce undesired effects of degraded images. In the recognition process, city block distance with deviation (CBDD) and asymmetric Mahalanobis distance (AMD) are proposed for rough classification and fine classification. With this recognition system, the experimental result of the database ETL9B reaches to 99.42%.
Nei Kato, Masato Suzuki, Shinichiro Omachi, Hirotomo Aso, Yoshiaki Nemoto
IEEE Trans. Pattern Anal. Mach. Intell.3
1998 An algorithm for constructing a multi-template dictionary for character recognition considering distribution of feature vectors
abstract
An algorithm to construct a multi-template dictionary for character recognition is presented. It is a repetitive process of partitioning and estimating the distribution by observing both within-class distribution and between-class information. The effectiveness of the algorithm is shown by experimental results.
Shinichiro Omachi, Hirotomo Aso
ICPR2