Brian Kenji Iwana

dblp:169/8958 · DBLP profile ↗
← Back
20ranked-venue papers in the field
3as first author
7since 2021 · last 2024
0000-0002-5146-6818ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 20 (3 first)
YearPublicationVenuePosition
2024 What Text Design Characterizes Book Genres?
Daichi Haraguchi, Brian Kenji Iwana, Seiichi Uchida
DAS2
2024 Test Time Augmentation as a Defense Against Adversarial Attacks on Online Handwriting
Yoh Yamashita, Brian Kenji Iwana
ICDAR (2)2
2023 Vision Conformer: Incorporating Convolutions into Vision Transformer Layers
Brian Kenji Iwana, Akihiro Kusuda
ICDAR (4)1
2023 Contour Completion by Transformers and Its Application to Vector Font Data
Yusuke Nagata, Brian Kenji Iwana, Seiichi Uchida
ICDAR (5)2
2021 Attention to Warp: Deep Metric Learning for Multivariate Time Series
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
ICDAR (3)6
2021 Font Style that Fits an Image - Font Generation Based on Image Context
Taiga Miyazono, Brian Kenji Iwana, Daichi Haraguchi, Seiichi Uchida
ICDAR (3)2
2021 Towards Book Cover Design via Layout Graphs
Taiga Miyazono, Seiichi Uchida, Brian Kenji Iwana
ICDAR (3)5
2020 Neural Style Difference Transfer and Its Application to Font Generation
Gantugs Atarsaikhan, Brian Kenji Iwana, Seiichi Uchida
DAS2
2020 Character-Independent Font Identification
Daichi Haraguchi, Shota Harada, Brian Kenji Iwana, Yuto Shinahara, Seiichi Uchida
DAS3
2020 Effect of Text Color on Word Embeddings
Masaya Ikoma, Brian Kenji Iwana, Seiichi Uchida
DAS2
2020 ACMU-Nets: Attention Cascading Modular U-Nets Incorporating Squeeze and Excitation Blocks
Seokjun Kang, Brian Kenji Iwana, Seiichi Uchida
DAS2
2019 On the Ability of a CNN to Realize Image-to-Image Language Conversion
abstract
The purpose of this paper is to reveal the ability that Convolutional Neural Networks (CNN) have on the novel task of image-to-image language conversion. We propose a new network to tackle this task by converting images of Korean Hangul characters directly into images of the phonetic Latin character equivalent. The conversion rules between Hangul and the phonetic symbols are not explicitly provided. The results of the proposed network show that it is possible to perform image-to-image language conversion. Moreover, it shows that it can grasp the structural features of Hangul even from limited learning data. In addition, it introduces a new network to use when the input and output have significantly different features.
Kohei Baba, Seiichi Uchida, Brian Kenji Iwana
ICDAR3
2019 Cascading Modular U-Nets for Document Image Binarization
abstract
In recent years, U-Net has achieved good results in various image processing tasks. However, conventional U-Nets need to be re-trained for individual tasks with enough amount of images with ground-truth. This requirement makes U-Net not applicable to tasks with small amounts of data. In this paper, we propose to use "modular" U-Nets, each of which is pre-trained to perform an existing image processing task, such as dilation, erosion, and histogram equalization. Then, to accomplish a specific image processing task, such as binarization of historical document images, the modular U-Nets are cascaded with inter-module skip connections and fine-tuned to the target task. We verified the proposed model using the Document Image Binarization Competition (DIBCO) 2017 dataset.
Seokjun Kang, Brian Kenji Iwana, Seiichi Uchida
ICDAR2
2019 Selective Super-Resolution for Scene Text Images
abstract
In this paper, we realize the enhancement of super-resolution using images with scene text. Specifically, this paper proposes the use of Super-Resolution Convolutional Neural Networks (SRCNN) which are constructed to tackle issues associated with characters and text. We demonstrate that standard SRCNNs trained for general object super-resolution is not sufficient and that the proposed method is a viable method in creating a robust model for text. To do so, we analyze the characteristics of SRCNNs through quantitative and qualitative evaluations with scene text data. In addition, analysis using the correlation between layers by Singular Vector Canonical Correlation Analysis (SVCCA) and comparison of filters of each SRCNN using t-SNE is performed. Furthermore, in order to create a unified super-resolution model specialized for both text and objects, a model using SRCNNs trained with the different data types and Content-wise Network Fusion (CNF) is used. We integrate the SRCNN trained for character images and then SRCNN trained for general object images, and verify the accuracy improvement of scene images which include text. We also examine how each SRCNN affects super-resolution images after fusion.
Ryo Nakao, Brian Kenji Iwana, Seiichi Uchida
ICDAR2
2019 Modality Conversion of Handwritten Patterns by Cross Variational Autoencoders
abstract
This research attempts to construct a network that can convert online and offline handwritten characters to each other. The proposed network consists of two Variational Auto-Encoders (VAEs) with a shared latent space. The VAEs are trained to generate online and offline handwritten Latin characters simultaneously. In this way, we create a cross-modal VAE (Cross-VAE). During training, the proposed Cross-VAE is trained to minimize the reconstruction loss of the two modalities, the distribution loss of the two VAEs, and a novel third loss called the space sharing loss. This third, space sharing loss is used to encourage the modalities to share the same latent space by calculating the distance between the latent variables. Through the proposed method mutual conversion of online and offline handwritten characters is possible. In this paper, we demonstrate the performance of the Cross-VAE through qualitative and quantitative analysis.
Taichi Sumi, Brian Kenji Iwana, Hideaki Hayashi, Seiichi Uchida
ICDAR2
2019 Deep Dynamic Time Warping: End-to-End Local Representation Learning for Online Signature Verification
abstract
Siamese networks have been shown to be successful in learning deep representations for multivariate time series verification. However, most related studies optimize a global distance objective and suffer from a low discriminative power due to the loss of temporal information. To address this issue, we propose an end-to-end, neural network-based framework for learning local representations of time series, and demonstrate its effectiveness for online signature verification. This framework optimizes a Siamese network with a local embedding loss, and learns a feature space that preserves the temporal location-wise distances between time series. To achieve invariance to non-linear temporal distortion, we propose building a dynamic time warping block on top of the Siamese network, which will greatly improve the accuracy for local correspondences across intra-personal variability. Validation with respect to online signature verification demonstrates the advantage of our framework over existing techniques that use either handcrafted or learned feature representations.
Xiaomeng Wu, Akisato Kimura, Brian Kenji Iwana, Seiichi Uchida, Kunio Kashino
ICDAR3
2019 Capturing Micro Deformations from Pooling Layers for Offline Signature Verification
abstract
In this paper, we propose a novel Convolutional Neural Network (CNN) based method that extracts the location information (displacement features) of the maximums in the max-pooling operation and fuses it with the pooling features to capture the micro deformations between the genuine signatures and skilled forgeries as a feature extraction procedure. After the feature extraction procedure, we apply support vector machines (SVMs) as writer-dependent classifiers for each user to build the signature verification system. The extensive experimental results on GPDS-150, GPDS-300, GPDS-1000, GPDS-2000, and GPDS-5000 datasets demonstrate that the proposed method can discriminate the genuine signatures and their corresponding skilled forgeries well and achieve state-of-the-art results on these datasets.
Yuchen Zheng 0001, Wataru Ohyama, Brian Kenji Iwana, Seiichi Uchida
ICDAR3
2018 Contained Neural Style Transfer for Decorated Logo Generation
abstract
Making decorated logos requires image editing skills, without sufficient skills, it could be a time-consuming task. While there are many on-line web services to make new logos, they have limited designs and duplicates can be made. We propose using neural style transfer with clip art and text for the creation of new and genuine logos. We introduce a new loss function based on distance transform of the input image, which allows the preservation of the silhouettes of text and objects. The proposed method contains style transfer to only a designated area. We demonstrate the characteristics of proposed method. Finally, we show the results of logo generation with various input images.
Gantugs Atarsaikhan, Brian Kenji Iwana, Seiichi Uchida
DAS2
2017 Component Awareness in Convolutional Neural Networks
abstract
In this work, we investigate the ability of Convolutional Neural Networks (CNN) to infer the presence of components that comprise an image. In recent years, CNNs have achieved powerful results in classification, detection, and segmentation. However, these models learn from instance-level supervision of the detected object. In this paper, we determine if CNNs can detect objects using image-level weakly supervised labels without localization. To demonstrate that a CNN can infer awareness of objects, we evaluate a CNN's classification ability with a database constructed of Chinese characters with only character-level labeled components. We show that the CNN is able to achieve a high accuracy in identifying the presence of these components without specific knowledge of the component. Furthermore, we verify that the CNN is deducing the knowledge of the target component by comparing the results to an experiment with the component removed. This research is important for applications with large amounts of data without robust annotation such as Chinese character recognition.
Brian Kenji Iwana, Letao Zhou, Kumiko Tanaka-Ishii, Seiichi Uchida
ICDAR1
2015 Tackling temporal pattern recognition by vector space embedding
abstract
This paper introduces a novel method of reducing the number of prototype patterns necessary for accurate recognition of temporal patterns. The nearest neighbor (NN) method is an effective tool in pattern recognition, but the downside is it can be computationally costly when using large quantities of data. To solve this problem, we propose a method of representing the temporal patterns by embedding dynamic time warping (DTW) distance based dissimilarities in vector space. Adaptive boosting (AdaBoost) is then applied for classifier training and feature selection to reduce the number of prototype patterns required for accurate recognition. With a data set of handwritten digits provided by the International Unipen Foundation (iUF), we successfully show that a large quantity of temporal data can be efficiently classified produce similar results to the established NN method while performing at a much smaller cost.
Brian Kenji Iwana, Seiichi Uchida, Kaspar Riesen, Volkmar Frinken
ICDAR1