Keiichiro Shirai

dblp:63/7629 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
1since 2021 · last 2022
0000-0003-2072-5087ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-authorArtificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Geometric modeling and processing · 42% Computational photography and imaging · 30% Image and video processing · 18%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Geometric modeling and processing
3d reconstruction
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Computational photography and imaging › image-based modeling
3d reconstruction from images
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Geometric modeling and processing
shape matching
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Mathematical optimization › combinatorial optimization
assignment problem
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Mathematical optimization
combinatorial optimization
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Information retrieval › hashing
binary code learning
0.312018
Hadamard Coding for Supervised Discrete Hashing · IEEE Trans. Image Process. 2018
Information retrieval
hashing
0.312018
Hadamard Coding for Supervised Discrete Hashing · IEEE Trans. Image Process. 2018
Information retrieval
image retrieval
0.312018
Hadamard Coding for Supervised Discrete Hashing · IEEE Trans. Image Process. 2018
Information retrieval › hashing › supervised hashing
supervised discrete hashing
0.312018
Hadamard Coding for Supervised Discrete Hashing · IEEE Trans. Image Process. 2018
Visual content generation and editing › style transfer
color transfer
0.212016
Misaligned Image Integration With Local Linear Model · IEEE Trans. Image Process. 2016
Image and video processing › image restoration › image denoising › spectral image denoising
hyperspectral image denoising
0.212016
Local Spectral Component Decomposition for Multi-Channel Image Denoising · IEEE Trans. Image Process. 2016
Image and video processing › image restoration
image denoising
0.212016
Local Spectral Component Decomposition for Multi-Channel Image Denoising · IEEE Trans. Image Process. 2016

Methods — techniques the papers use, named apart from their topics

iterative closest point · 1.1assignment algorithm · 1.1hadamard matrix · 0.3discrete cyclic coordinate descent · 0.3spectral component decomposition · 0.2local linear model · 0.2linear correlation analysis · 0.2convex optimization · 0.2
YearPublicationVenuePosition
2022 ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization
abstract
Abstract In April 2016, a massive earthquake with a magnitude of 7.3 struck Kumamoto region, Japan, causing major devastation. One of the structures that were damaged in Kumamoto was Kumamoto Castle, a cultural asset of great significance in Japan. The stone retaining wall “ishigaki” that formed the foundation of the castle collapsed, and the superstructure was destroyed. The number of stones is estimated to be more than 70,000, and restoration work is anticipated to take more than 20 years. Since each of the stones is an important cultural asset, the broken stone structure needed to be restored to its original state in order not to lose its cultural value forever. In addition, the fallen stones need to be returned to their original positions in the ishigakis. In similar cases, non-automatic visual verification was used. However, for Kumamoto Castle, this would have been impossible because a large number of stones were displaced as a result of the collapse. The purpose of this project is to provide support for the restoration work by matching the stones that fell down after the collapse with those before the collapse using information technology, such as computer vision and optimization technologies. Specifically, we captured photographic images of the stones before and after the collapse to match them. The technical contributions of this study are as follows: (a) To estimate the scale and surface orientation of the stones, we exploit 3D model construction from the images. (b) To solve the jigsaw-puzzle-like problem of reassembling the stone fragments, we exploit the combination of a customized iterative closest point (ICP) algorithm for shape position matching and an assignment algorithm to find the best pairs of stones before and after the collapse by using the matching degree obtained from ICP. Here, only the 2D shape of the stones before the collapse can be used due to the small number of photos available. In contrast, a detailed 3D shape can be obtained from the stones after their collapse. We matched these asymmetric data in 2D and 3D to enable a comprehensive reconstruction. (c) We developed a user-friendly graphical user interface system that was used by actual masons without special knowledge. The developed system was used to match the ishigaki of a turret, Iidamaru. As a result, we succeeded in identifying 337 stones, or approximately 90% of the 370 images. These results are expected to be useful for and were used as a blueprint during actual restoration work.
Gou Koutaki, Sakino Ando, Keiichiro Shirai, Tsuyoshi Kishigami
Int. J. Comput. Vis.3
2019 Reference-based local color distribution transformation method and its application to image integration
abstract
We propose a reference-based filtering method that transforms the local color distribution of individual image patches. With our method, the color of a reference image is transformed by patch-wise color transformation so that it comes close to a noisy input image. Our method can be regarded as an extended version of reference-based filtering methods, such as guided image filtering and our previously proposed linear local color distribution transformation (LCDT). The main contribution is to enable more flexible transformation, which is required in some applications, e.g. , flash/no-flash image integration and high dynamic range (HDR) image generation. We also propose an efficient framework for HDR image generation which is particularly suitable for generation of dark scenes without loss of image contrast. Experimental results show that the proposed quadratic LCDT achieves more robust color transformation, even in flash/no-flash image integration, and it was confirmed that the proposed framework based on quadratic LCDT is effective in HDR image generation.
Ryo Matsuoka, Keiichiro Shirai, Masahiro Okuda
Signal Process. Image Commun.2
2018 Color Affine Subspace Pursuit for Color Artifact Removal
abstract
This paper proposes color affine subspace pursuit (CASSP) for color artifact removal. Local patches in natural color images tend to exhibit a line distribution, so-called a color line. According to this characteristic, a convex-optimization-based image recovery with a local color nuclear norm (LCNN) has conventionally been introduced to promote the color line property of local patches and succeeded in removing color artifacts. It is, however, often the case that a local patch does not form a line distribution, but a union of affine subspaces (UoAS), e.g., a patch consisting of two different colors. In such regions, the LCNN often results in color fading or color smearing. This paper promotes the UoAS property, i.e., the color line or plane distribution for each affine subspace in local patches by using CASSP. Our cost function for the CASSP consists of the LCNN for each centered color distribution cluster. Experimental results show that the CASSP improves both numerical reconstruction error and subjective visual quality, compared with the LCNN.
Kazuki Yamanaka, Seisuke Kyochi, Shunsuke Ono, Keiichiro Shirai
ICASSP4
2018 Hadamard Coded Discrete Cross Modal Hashing
abstract
Cross-modal retrieval is a hot topic in the fields of machine learning and media retrieval, making it possible to relate different types of media, such as image, text, and audio. A powerful method for the cross-modal retrieval, discrete cross-modal hashing (DCH), has recently been proposed. The DCH can encode different types of media feature vectors to binary codes. When stored in a database, the binary code makes searches efficient because the Hamming distance between the corresponding sections of two binary codes can be computed via a specialized CPU operation. Moreover, it has recently been shown that when optimizing hash functions for supervised discrete hashing (SDH), Hadamard matrices can be used, and this technique is named “HC-SDH”. In this study, we apply the HC-SDH to cross-modal hashing. Experimental results demonstrate that the proposed cross-modal hashing can achieve the same performance as the conventional DCH with 1/24 of the training time.
Koichi Eto, Gou Koutaki, Keiichiro Shirai
ICIP3
2018 Hadamard Coding for Supervised Discrete Hashing
abstract
In this paper, we propose a learning-based supervised discrete hashing method. Binary hashing is widely used for large-scale image retrieval as well as video and document searches because the compact binary code representation is essential for data storage and reasonable for query searches using bit-operations. The recently proposed supervised discrete hashing (SDH) method efficiently solves mixed-integer programming problems by alternating optimization and the discrete cyclic coordinate descent (DCC) method. Based on some preliminary experiments, we show that the SDH method can be simplified without performance degradation. We analyze the simplified model and provide a mathematically exact solution thereof; we reveal that the exact binary code is provided by a "Hadamard matrix." Therefore, we named our method Hadamard codedsupervised discrete hashing (HC-SDH). In contrast to SDH, our model does not require an alternating optimization algorithm and does not depend on initial values. HC-SDH is also easier to implement than iterative quantization (ITQ). Experimental results involving a large-scale database show that Hadamard coding outperforms conventional SDH in terms of precision, recall, and computational time. On the large datasets SUN-397 and ImageNet, HC-SDH provides a superior mean average of precision (mAP) and top-accuracy compared to the conventional SDH methods with the same code length and FastHash. The training time of HC-SDH is 170 times faster than conventional SDH and the testing time including the encoding time is seven times faster than FastHash which encodes using a binary-tree.
Gou Koutaki, Keiichiro Shirai, Mitsuru Ambai
IEEE Trans. Image Process.2
2017 Principal noiseless color component extraction by linear color composition with optimal coefficients
abstract
In this paper, we propose a principal color component extraction method that is simply performed by linear color composition (transformation) of R, G, B colors, but its composite coefficients are calculated so as to obtain a noisy-texture-less principal component of RGB color images. Our method is related to principal component analysis (PCA) and edge preserving smoothing by total variation (TV) minimization. The resultant image becomes a principal color component image with the minimum total variation. We show this problem can be formulated as TV minimization on a spherical manifold for a whitened data matrix. Although this spherical constraint is non-convex, it can be solved by using alternating direction method of multipliers (ADMM). As its application, we show the results of text character extraction from ancient wooden tablets, and how our method extracts faint ink characters while reducing wood grain textures. Our method is unsupervised but has performance equivalent to a linear discriminant analysis (LDA) method with user-assisted information.
Takuya Sugimoto, Kazuhiro Fujimori, Keiichiro Shirai, Hidetoshi Miyao, Minoru Maruyama
ICIP3
2016 Vectorial total variation based on arranged structure tensor for multichannel image restoration
abstract
We propose a new regularization function, named as Arranged Structure tensor Total Variation (ASTV), for multichannel image restoration. Since the standard structure tensor is a matrix whose eigenvalues well encodes local neighborhood information of an image, there has been proposed vectorial total variation based on the structure tensor for image regularization. However, the correlation among the channels cannot be measured by the structure tensor because the discrete differences of all the channels are just summed up in the entries of the structure tensor. On the other hand, ASTV is based on a newly-defined arranged structure tensor that becomes an approximately low-rank matrix when multichannel images have strong correlation among their channels. This suggests that penalizing the nuclear norm of the arranged structure tensor is a reasonable regularization for multichannel images, leading to the definition of ASTV. Experimental results illustrate the advantage of ASTV over a state-of-the-art vectorial total variation based on the structure tensor.
Shunsuke Ono, Keiichiro Shirai, Masahiro Okuda
ICASSP2
2016 Image colorization based on ADMM with fast singular value thresholding by Chebyshev polynomial approximation
abstract
We propose an image colorization method using fast soft-thresholding of singular values (singular value thresholding). An image colorization method with nuclear norm minimization (NNM) has been proposed and brings good results. NNM usually requires iterative application of singular value decomposition (SVD) for singular value thresholding. However, the computational cost of SVD in the colorization method becomes too expensive to handle high-resolution images. In this paper, we reduce its computational cost by using Chebyshev polynomial approximation (CPA). Singular value thresholding is expressed by a multiplication of certain matrices derived from the characteristic of CPA. As a result, our CPA-based technique makes the image colorization method much more efficient. In addition, we replace the optimization method used in the image colorization method by alternating direction method of multipliers, which further accelerates the computation. Experimental results verify the effectiveness of our method with respect to the computation time and the approximation precision.
Masaki Onuki, Shunsuke Ono, Keiichiro Shirai, Yuichi Tanaka 0001
ICASSP3
2016 Performance evaluation of multi-target tracking for PhyC-SN
abstract
Physical conversion sensor networks (PhyC-SN) achieve simultaneous data collection in wireless sensor networks. Although the PhyC-SN can recognize the median and the outliers of all the sensing data simultaneously, their separation into each sensor data is a difficult task. To address this task, we have proposed a data separation method that utilizes multi-target tracking with use of the sensing data and the received spectrum power for the separation. Even if some of the sensor data are close to each other, the instantaneous power of the sensor data bearing signal is a useful feature for the separation. In this paper, we clarify the separation accuracy of our separation method under various wireless environments by computer simulation.
Minato Oriuchi, Osamu Takyu, Keiichiro Shirai, Fumihito Sasamori, Shiro Handa, Takeo Fujii, Mai Ohta
WCNC3
2016 Misaligned Image Integration With Local Linear Model
abstract
We present a new image integration technique for a flash and long-exposure image pair to capture a dark scene without incurring blurring or noisy artifacts. Most existing methods require well-aligned images for the integration, which is often a burdensome restriction in practical use. We address this issue by locally transferring the colors of the flash images using a small fraction of the corresponding pixels in the long-exposure images. We formulate the image integration as a convex optimization problem with the local linear model. The proposed method makes it possible to integrate the color of the long-exposure image with the detail of the flash image without causing any harmful effects to its contrast, where we do not need perfect alignment between the images by virtue of our new integration principle. We show that our method successfully outperforms the state of the art in the image integration and reference-based color transfer for challenging misaligned data sets.
Tatsuya Baba, Ryo Matsuoka, Keiichiro Shirai, Masahiro Okuda
IEEE Trans. Image Process.3
2016 Local Spectral Component Decomposition for Multi-Channel Image Denoising
abstract
We propose a method for local spectral component decomposition based on the line feature of local distribution. Our aim is to reduce noise on multi-channel images by exploiting the linear correlation in the spectral domain of a local region. We first calculate a linear feature over the spectral components of an M -channel image, which we call the spectral line, and then, using the line, we decompose the image into three components: a single M -channel image and two gray-scale images. By virtue of the decomposition, the noise is concentrated on the two images, and thus our algorithm needs to denoise only the two gray-scale images, regardless of the number of the channels. As a result, image deterioration due to the imbalance of the spectral component correlation can be avoided. The experiment shows that our method improves image quality with less deterioration while preserving vivid contrast. Our method is especially effective for hyperspectral images. The experimental results demonstrate that our proposed method can compete with the other state-of-the-art denoising methods.
Mia Rizkinia, Tatsuya Baba, Keiichiro Shirai, Masahiro Okuda
IEEE Trans. Image Process.3
2015 Performance analysis of retargeting pyramid and its applications
abstract
Improved retargeting pyramid (iRP) is a multiscale image pyramid using content-aware image resizing (also known as retargeting). It has been reported that the iRP outperforms conventional pyramids in linear approximation and denoising. However, the reason of its performance gain was not well discussed so far. In this paper, we reveal that retargeting in the iRP acts as locally underdecimated filters for significant region(s) while keeping its global downsampling ratio. Furthermore, our iRP is applied to various image processing applications to validate its performance.
Ryosuke Morita, Keiichiro Shirai, Yuichi Tanaka 0001
ICIP2
2015 Non-local/local image filters using fast eigenvalue filtering
abstract
In this paper, we propose a fast and an approximate solution of non-local/local filters using Chebyshev polynomial approximation (CPA). A non-local/local filter is generally expressible in a matrix form. From the matrix notation, image denoising performance is improved by filtering the eigenvalues of the filter matrix. However, it requires much execution time due to computational complexity of eigendecomposition. To reduce the computational cost, we apply the CPA to eigenvalue filtering, leading to an eigendecomposition-free procedure. Moreover, a fast SURE-based parameter optimization is possible by using the CPA. It enables us to determine a suitable filtering parameter efficiently. Numerical examples illustrate that the proposed method is significantly faster than conventional methods while it maintains high approximate precision.
Masaki Onuki, Shunsuke Ono, Keiichiro Shirai, Yuichi Tanaka 0001
ICIP3
2015 A regularization approach for bayer reconstruction in lossy image coding by inverse demosaicing
abstract
Color image coding by inverse demosaicing, which exploits the implicit Bayer structure in color images, has a potential to achieve superior performance compared to conventional color image coding methods. The previous framework of inverse demosaicing was limited to lossless and near-lossless data compression, while this paper explores its adaptation to lossy compression. To cope with distortions due to lossy compression, we propose a regularization approach using side color information for the Bayer recovery problem in the decoder. Thanks to careful design of the regularization, the resulting Bayer recovery problem becomes an unconstrained quadratic programming problem, and thus several efficient solvers can be used. A numerical example demonstrates the efficacy of our approach. It can significantly reduce distortions in the recovery of the Bayer data and keep the total bit rate.
Masao Yamagishi, Seisuke Kyochi, Keiichiro Shirai, Masahiro Okuda
ICIP3
2014 Flash/no-flash image integration using convex optimization
abstract
When high ISO sensitivity is used to acquire images of dark scenes, their detail textures are often deteriorated by sensor noise. On the other hand, using flash photography with artificial light, one can shorten the exposure time, and obtain a sharp image under the low ISO sensitivity. However, the use of flash light changes the color tone and often generates unnatural images due to a specific color temperature of the additional light. This paper presents a new efficient method for flash/no-flash image integration. In contrast to conventional integration methods assuming that the flash image has a sharp texture without any noise, our method can successfully remove noise. Specifically, our method separately handles regions within the reach of the flash light and other regions out of range of the flash light, because the two regions have much different characteristics. As for the former well-exposed regions, we transferred the detail of the flash image to no-flash image by optimization and component separation. As for the latter under-exposed regions, an optimization based joint bilateral filtering that uses information of a flash image is performed to remove noise. Experimental results show the effectiveness of our method compared to the conventional methods.
Tatsuya Baba, Ryo Matsuoka, Shunsuke Ono, Keiichiro Shirai, Masahiro Okuda
ICASSP4
2014 Lossless/near-lossless color image coding by inverse demosaicing
abstract
In this paper, we introduce a novel framework for lossless/near-lossless (LS/NLS) color image coding assisted by an inverse demosaicing. Conventional frameworks are typically based on prediction (and quantization for NLS coding) followed by entropy coding, such as the JPEG-LS for bit rate saving. The approach of this work is totally different from the conventional ones. Basically, color images are created by demosaicing Bayer-pattern color filter array (CFA) whose operator can be expressed as square matrices. By using the (pseudo) inverse matrix of a joint demosaicing and color-to-gray conversion, the proposed decoder can recover the color image from its corresponding gray image data which is losslessly transmitted by the proposed encoder. Thus, LS/NLS color image reconstruction can be achieved while saving a bit rate significantly. In addition, using the same framework of color image coding, LS/NLS CFA coding can be realized by a comparable bit rate with JPEG-LS.
Ryo Kuroiwa 0003, Ryo Matsuoka, Seisuke Kyochi, Keiichiro Shirai, Masahiro Okuda
ICASSP4
2014 Retargeting pyramid using direct decimation
abstract
The retargeting pyramid (RP) is a multiscale image pyramid using content-aware image resizing. In the previous implementation of the RP, a two-step interpolation is adopted to obtain the desired resolution. However, this interpolation leads to performance loss for image processing. In this paper, we improve the performance of the RP by replacing the two-step interpolation with a single interpolation using a matrix representation of the bilateral filter and Tikhonov regularization.
Ryosuke Morita, Keiichiro Shirai, Yuichi Tanaka 0001
ICASSP2
2014 FFT based solution for multivariable L2 equations using KKT system via FFT and efficient pixel-wise inverse calculation
abstract
When solving l2optimization problems based on linear filtering with some regularization in signal/image processing such as Wiener filtering, the fast Fourier transform (FFT) is often available to reduce its computational complexity. Most of the problems, in which the FFT is used to obtain their solutions, are based on single variable equations. On the other hand, the Karush-Kuhn-Tucker (KKT) system, which is often used for solving constrained optimization problems, generally results in multivariable equations. In this paper, we propose a FFT based computational method for multivariable l2equations. Our method applies a FFT to each block of the KKT system, and represents the equation as an image-wise simultaneous equation consisting of Fourier transformed filters and images. In our method, an inverse matrix calculation that consists of complex pixel values gathered from each transformed image is required for each pixel. We exploit the homogeneity of neighboring values and solve them efficiently.
Keiichiro Shirai, Masahiro Okuda
ICASSP1
2014 Color transform between image pair using covariance correspondences of local color distributions
abstract
In the guided filter that performs operations such as denoising and contrast correction with the help of a guide image, the positions of corresponding subjects need to be completely aligned, otherwise the misaligned regions in the output image are deteriorated by blur. In this paper, we propose a guided filter for images which include moving dynamic regions. Our filter uses correspondences of local covariance matrices instead of using the conventional pixel-to-pixel correspondences. In addition, we also propose a classification method to detect the dynamic regions by using the support vector machine. Combining two kinds of guided filters for static/dynamic regions, more natural resulting images are obtained.
Yusuke Tatesumi, Keisuke Iwata, Keiichiro Shirai, Masahiro Okuda
ICASSP3
2013 High dynamic range image acquisition using flash image
abstract
We propose a denoising technique using multiple exposure image integration. When acquiring a dark scene, the detail of the dark area is often deteriorated by sensor noise. For a high dynamic range image acquisition, denoising in dark areas is a critical issue, since the dark area is, in general, enhanced by a tone-mapping and the noise is made more visible when displaying it on an output devise. In our method, a flash image is utilized as well as no-flash multiple exposure images to further reduce the noise. Multiple exposure integration is performed in a wavelet domain, where noise removal is achieved by the wavelet-shrinkage for multiple exposures. Our method works well especially for noise in shadows. We show the validity of the proposed algorithm by simulating the method with some actual noisy images.
Ryo Matsuoka, Tatsuya Baba, Masahiro Okuda, Keiichiro Shirai
ICASSP4
2013 Scalable image representation using improved retargeting pyramid
abstract
The retargeting pyramid (RP) method is a good alternative to the well-known Laplacian pyramid (LP) approach for multiscale image decomposition. RP can be obtained by replacing the low-pass filtering and downsampling processes in LP with content-aware image resizing (a.k.a. retargeting), which is a technique being developed in computer vision research. In this paper, we improve RP so that it obtains good scalable image representation. The improved RP is then integrated with a well-known multiscale-multidirection (MSMD) transform, contourlet transform, to construct a saliency-oriented MSMD image representation. In the experiment, our decomposition outperforms the conventional pyramid structures.
Yuichi Tanaka 0001, Keiichiro Shirai
ICASSP2
2013 An OCR System with OCRopus for Scientific Documents Containing Mathematical Formulas
abstract
This paper describes the installation of a mathematical formula recognition module into an open source OCR system: OCRopus. In particular we consider the identification of inline formulas utilizing existing modules. Text lines including math formulas are first processed using a N-gram language model to reduce the number of formula candidates by thresholding the conditional probability of words. Then the formula candidates are classified into formulas and texts by SVM using geometric features associated with the bounding boxes of symbols.
Fumihiro Furukori, Shinpei Yamazaki, T. Miyagishi, Keiichiro Shirai, Masayuki Okamoto
ICDAR4
2012 Local Covariance Filtering for Color Images
Keiichiro Shirai, Masahiro Okuda, Takao Jinno, Masayuki Okamoto, Masaaki Ikehara
ACCV (4)1
2012 Removal of Background Patterns and Signatures for Magnetic Ink Character Recognition of Checks
abstract
This paper describes a method to extract the magnetic ink characters (MICR E-13B font) printed on bank-checks for the purpose of using OCR as supporting MICR. In the case of OCR, the colorful background patterns and the overlapped signatures on MICR characters make it difficult to extract characters respectively by using simple binarization and labeling. Our method estimates the color and pitch of MICR characters in order to separate the characters in contact with sign strokes, then the remaining sign strokes are removed by tracing them. In the experiment, we use circulated bank-checks and samples provided by SEIKO EPSON and show the performance of our method.
Keiichiro Shirai, Masashi Akita, Masayuki Okamoto, Kazuya Tanikawa, Takaaki Akiyama, Tetsuji Sakaguchi
Document Analysis Systems1
2012 Color-line vector field and local color component decomposition for smoothing and denoising of color images
Keiichiro Shirai, Masahiro Okuda, Masaaki Ikehara
ICPR1
2011 Embedding a Mathematical OCR Module into OCRopus
abstract
This paper describes embedding a mathematical formula recognition module into the OCR system OCRopus aiming at developing a OCR system for scientific and technical documents which include mathematical formulas. OCRopus is a open source OCR system emphasizing modularity, easy extensibility, and reuse. This system has several basic components such as preprocessing, layout analysis, and text line recognition, so it is a challenging project to embed the mathematical formula recognition module into the OCRopus system. We have developed the math OCR module, then report how to embed our module into the OCRopus system in order to realize a math OCR which can deal with wide variety of documents including mathematical formulas.
Shinpei Yamazaki, Fumihiro Furukori, Qinzheng Zhao, Keiichiro Shirai, Masayuki Okamoto
ICDAR4
2011 Noiseless no-flash photo creation by color transform of flash image
abstract
In the dark place photographing, the increase of noise is one of the major problems to be addressed. Although using a flash is effective to reduce the noise, natural colors are faded away due to increase of specular lights. In this paper, we present a method that generates a no-flash like flash image by approximating colors of a flash image by those of a no-flash image. Our method is based on the ”color-line” feature of images, a linear distribution of colors in a local region is transformed by a set of transform matrices automatically. This method is able to deal correctly with the occluded regions where the colors are saturated by reflections or shadows and the color distribution is squashed.
Keiichiro Shirai, Masayuki Okamoto, Masaaki Ikehara
ICIP1
2008 A Study for High Performance Character Extraction from Color Scene Images
abstract
This paper describes a method for extracting character strings from scene images. Most characters on scene images appear with the same color and font size at every word or text line. In our algorithm, a scene image is divided into several blocks based on edges in the color space at first. Then the blobs, which consist of similar color pixels, are extracted by a clustering in a color space for each block. Although these blobs are correspond to characters or background patterns, after connecting them using these aspect ratios and pitches, SVM (Support Vector Machine) on several textural features of these blobs will classify each connected blob into character or background patterns. Testing with 251 images from ICDAR 2003 Text Locating Competition shows effectiveness of our algorithm.
Keiichiro Shirai, Masanori Wakabayashi, Masayuki Okamoto, Hiroaki Yamamoto
Document Analysis Systems1