Brian Kenji Iwana

dblp:169/8958 · DBLP profile ↗
← Back
51ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0002-5146-6818ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 7 first-author · 17 since 2021Databases, data management, data science and information retrieval · 20 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Time Series Generation Conditioned on Unstructured Natural Language
Jaeyun Woo, Brian Kenji Iwana
ICPR (1)3
2024 What Text Design Characterizes Book Genres?
Daichi Haraguchi, Brian Kenji Iwana, Seiichi Uchida
DAS2
2024 Test Time Augmentation as a Defense Against Adversarial Attacks on Online Handwriting
Yoh Yamashita, Brian Kenji Iwana
ICDAR (2)2
2024 Improving Online Handwriting Recognition with Transfer Learning Using Out-of-Domain and Different-Dimensional Sources
Jiseok Lee 0001, Masaki Akiba, Brian Kenji Iwana
ICPR (31)3
2024 Model Selection with a Shapelet-Based Distance Measure for Multi-source Transfer Learning in Time Series Classification
Jiseok Lee 0001, Brian Kenji Iwana
ICPR (27)2
2024 Improving the Robustness of Time Series Neural Networks from Adversarial Attacks Using Time Warping
Yoh Yamashita, Brian Kenji Iwana
ICPR (14)2
2024 Facial Gesture Classification with Few-shot Learning Using Limited Calibration Data from Photo-reflective Sensors on Smart Eyewear
abstract
This study investigates smart eyewear for facial gesture classification with low user calibration costs.The smart eyewear is equipped with low-cost, comfortable, energy-efficient photo-reflective sensors, which can detect changes in facial muscle movements.Although the sensor output is useful for facial gesture classification, individual user calibration is considered necessary.Moreover, re-calibration is required whenever the user's wearing position changes.Therefore, reducing the calibration cost is crucial for the wider applicability of the eyewear.To address this issue, we propose a few-shot domain adaptation approach using Convolutional Neural Networks (CNN).We evaluate the accuracy of classifying eight gestures with data augmentation and a supervised contrastive loss.Data augmentation is employed to make the model more robust to noise, while the supervised contrastive loss is introduced to learn user-invariant features.Our approach with data augmentation achieves robust gesture classification, with an average accuracy of 93.46% (SD = 8.34%) for three emotion-related gestures without user-specific data.Furthermore, for user-independent training, we demonstrated that using few-shot learning with pre-trained models and only four repetitions of calibration data per gesture achieved a practical accuracy of 91.38% (SD = 6.03%), showing that a small amount of user-specific data is sufficient for the accurate classification.Also, it works under different wearing conditions, achieving an accuracy of 90.19% (SD = 3.56%).These results illustrate the potential of our method to improve the practicality of smart eyewear for facial expression recognition in cases of limited user data, making it more accessible and user-friendly.
Katsutoshi Masai, Maki Sugimoto, Brian Kenji Iwana
MUM3
2024 Scene text recognition via dual character counting-aware visual and semantic modeling network
Anna Zhu, Brian Kenji Iwana
Sci. China Inf. Sci.3
2023 Few shot font generation via transferring similarity guided global style and quantization local style
abstract
Automatic few-shot font generation (AFFG), aiming at generating new fonts with only a few glyph references, reduces the labor cost of manually designing fonts. However, the traditional AFFG paradigm of style-content disentanglement cannot capture the diverse local details of different fonts. So, many component-based approaches are proposed to tackle this problem. The issue with component-based approaches is that they usually require special pre-defined glyph components, e.g., strokes and radicals, which is infeasible for AFFG of different languages. In this paper, we present a novel font generation approach by aggregating styles from character similarity-guided global features and stylized component-level representations. We calculate the similarity scores of the target character and the referenced samples by measuring the distance along the corresponding channels from the content features, and assigning them as the weights for aggregating the global style features. To better capture the local styles, a cross-attention-based style transfer module is adopted to transfer the styles of reference glyphs to the components, where the components are self-learned discrete latent codes through vector quantization without manual definition. With these designs, our AFFG method could obtain a complete set of component-level style representations, and also control the global glyph characteristics. The experimental results reflect the effectiveness and generalization of the proposed method on different linguistic scripts, and also show its superiority when compared with other state-of-the-art methods. The source code can be found at https://github.com/awei669/VQ-Font.
Anna Zhu, Brian Kenji Iwana, Shilin Li
ICCV4
2023 Vision Conformer: Incorporating Convolutions into Vision Transformer Layers
Brian Kenji Iwana, Akihiro Kusuda
ICDAR (4)1
2023 Contour Completion by Transformers and Its Application to Vector Font Data
Yusuke Nagata, Brian Kenji Iwana, Seiichi Uchida
ICDAR (5)2
2023 FETNet: Feature erasing and transferring network for scene text removal
Guangtao Lyu, Kun Liu 0027, Anna Zhu, Seiichi Uchida, Brian Kenji Iwana
Pattern Recognit.5
2023 Corrigendum to "FETNet: Feature Erasing and Transferring Network for Scene Text Removal": Pattern Recognition Volume 140 (2023) 109531
Guangtao Lyu, Kun Liu 0027, Anna Zhu, Seiichi Uchida, Brian Kenji Iwana
Pattern Recognit.5
2023 Deep attentive time warping
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
Pattern Recognit.6
2022 On Mini-Batch Training with Varying Length Time Series
abstract
In real-world time series recognition applications, it is possible to have data with varying length patterns. However, when using artificial neural networks (ANN), it is standard practice to use fixed-sized mini-batches. To do this, time series data with varying lengths are typically normalized so that all the patterns are the same length. Normally, this is done using zero padding or truncation without much consideration. We propose a novel method of normalizing the lengths of the time series in a dataset by exploiting the dynamic matching ability of Dynamic Time Warping (DTW). In this way, the time series lengths in a dataset can be set to a fixed size while maintaining features typical to the dataset. In the experiments, all 11 datasets with varying length time series from the 2018 UCR Time Series Archive are used. We evaluate the proposed method by comparing it with 18 other length normalization methods on a Convolutional Neural Network (CNN), a Long-Short Term Memory network (LSTM), and a Bidirectional LSTM (BLSTM). The code is publicly available at https://github.com/uchidalab/varyJength_time_series.
Brian Kenji Iwana
ICASSP1
2022 Dynamic Data Augmentation with Gating Networks for Time Series Recognition
abstract
Data augmentation is a technique to improve the generalization ability of machine learning methods by increasing the size of the dataset. However, since every augmentation method is not equally effective for every dataset, you need to select an appropriate method carefully. We propose a neural network that dynamically selects the best combination of data augmentation methods using a Gating Network and a mutually beneficial feature consistency loss. The Gating Network is able to control how much of each data augmentation is used for the representation within the network. The feature consistency loss gives a constraint that augmented features from the same input pattern should be in similar. In the experiments, we demonstrate the effectiveness of the proposed method on the 12 largest time-series datasets from 2018 UCR Time Series Archive and reveal the relationships between the data augmentation methods through analysis of the proposed method.
Daisuke Oba, Shinnosuke Matsuo, Brian Kenji Iwana
ICPR3
2022 Text Style Transfer based on Multi-factor Disentanglement and Mixture
abstract
Text style transfer aims to transfer the reference style of one text image to another text image. Previous works have only been able to transfer the style to a binary text image. In this paper, we propose a framework to disentangle the text images into three factors: text content, font, and style features, and then remix the factors of different images to transfer a new style. Both the reference and input text images have no style restrictions. Adversarial training through multi-factor cross recognition is adopted in the network for better feature disentanglement and representation. To decompose the input text images into a disentangled representation with swappable factors, the network is trained using similarity mining within pairs of exemplars. To train our model, we synthesized a new dataset with various text styles in both English and Chinese. Several ablation studies and extensive experiments on our designed and public datasets demonstrate the effectiveness of our approach for text style transfer.
Anna Zhu, Zhanhui Yin, Brian Kenji Iwana, Shengwu Xiong 0001
ACM Multimedia3
2021 Self-Augmented Multi-Modal Feature Embedding
abstract
Oftentimes, patterns can be represented through different modalities. For example, leaf data can be in the form of images or contours. Handwritten characters can also be either online or offline. To exploit this fact, we propose the use of self-augmentation and combine it with multi-modal feature embedding. In order to take advantage of the complementary information from the different modalities, the self-augmented multi-modal feature embedding employs a shared feature space. Through experimental results on classification with online handwriting and leaf images, we demonstrate that the proposed method can create effective embeddings.
Shinnosuke Matsuo, Seiichi Uchida, Brian Kenji Iwana
ICASSP3
2021 Attention to Warp: Deep Metric Learning for Multivariate Time Series
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
ICDAR (3)6
2021 Font Style that Fits an Image - Font Generation Based on Image Context
Taiga Miyazono, Brian Kenji Iwana, Daichi Haraguchi, Seiichi Uchida
ICDAR (3)2
2021 Towards Book Cover Design via Layout Graphs
Taiga Miyazono, Seiichi Uchida, Brian Kenji Iwana
ICDAR (3)5
2021 Complex image processing with less data - Document image binarization by integrating multiple pre-trained U-Net modules
Seokjun Kang, Brian Kenji Iwana, Seiichi Uchida
Pattern Recognit.2
2021 Learning the micro deformations by max-pooling for offline signature verification
Yuchen Zheng 0001, Brian Kenji Iwana, Muhammad Imran Malik, Sheraz Ahmed, Wataru Ohyama, Seiichi Uchida
Pattern Recognit.2
2020 Neural Style Difference Transfer and Its Application to Font Generation
Gantugs Atarsaikhan, Brian Kenji Iwana, Seiichi Uchida
DAS2
2020 Character-Independent Font Identification
Daichi Haraguchi, Shota Harada, Brian Kenji Iwana, Yuto Shinahara, Seiichi Uchida
DAS3
2020 Effect of Text Color on Word Embeddings
Masaya Ikoma, Brian Kenji Iwana, Seiichi Uchida
DAS2
2020 ACMU-Nets: Attention Cascading Modular U-Nets Incorporating Squeeze and Excitation Blocks
Seokjun Kang, Brian Kenji Iwana, Seiichi Uchida
DAS2
2020 Negative Pseudo Labeling Using Class Proportion for Semantic Segmentation in Pathology
Hiroki Tokunaga, Brian Kenji Iwana, Yuki Teramoto, Akihiko Yoshizawa, Ryoma Bise
ECCV (15)2
2020 What is the Reward for Handwriting? - A Handwriting Generation Model Based on Imitation Learning
abstract
Analyzing the handwriting generation process is an important issue and has been tackled by various generation models, such as kinematics based models and stochastic models. In this study, we use a reinforcement learning (RL) framework to realize handwriting generation with the careful future planning ability. In fact, the handwriting process of human beings is also supported by their future planning ability; for example, the ability is necessary to generate a closed trajectory like `0' because any shortsighted model, such as a Markovian model, cannot generate it. For the algorithm, we employ generative adversarial imitation learning (GAIL). Typical RL algorithms require the manual definition of the reward function, which is very crucial to control the generation process. In contrast, GAIL trains the reward function along with the other modules of the framework. In other words, through GAIL, we can understand the reward of the handwriting generation process from handwriting examples. Our experimental results qualitatively and quantitatively show that the learned reward catches the trends in handwriting generation and thus GAIL is well suited for the acquisition of handwriting behavior.
Keisuke Kanda, Brian Kenji Iwana, Seiichi Uchida
ICFHR2
2020 Time Series Data Augmentation for Neural Networks by Time Warping with a Discriminative Teacher
abstract
Neural networks have become a powerful tool in pattern recognition and part of their success is due to generalization from using large datasets. However, unlike other domains, time series classification datasets are often small. In order to address this problem, we propose a novel time series data augmentation called guided warping. While many data augmentation methods are based on random transformations, guided warping exploits the element alignment properties of Dynamic Time Warping (DTW) and shapeDTW, a high-level DTW method based on shape descriptors, to deterministically warp sample patterns. In this way, the time series are mixed by warping the features of a sample pattern to match the time steps of a reference pattern. Furthermore, we introduce a discriminative teacher in order to serve as a directed reference for the guided warping. We evaluate the method on all 85 datasets in the 2015 UCR Time Series Archive with a deep convolutional neural network (CNN) and a recurrent neural network (RNN). The code with an easy to use implementation can be found at https://github.com/uchidalab/time_series_augmentation.
Brian Kenji Iwana, Seiichi Uchida
ICPR1
2020 DTW-NN: A novel neural network for time series recognition using dynamic alignment between inputs and weights
Brian Kenji Iwana, Volkmar Frinken, Seiichi Uchida
Knowl. Based Syst.1
2020 Time series classification using local distance-based features in multi-modal fusion networks
Brian Kenji Iwana, Seiichi Uchida
Pattern Recognit.1
2020 Few-Shot Text Style Transfer via Deep Feature Similarity
abstract
Generating text to have a consistent style with only a few observed highly-stylized text samples is a difficult task for image processing. The text style involving the typography, i.e., font, stroke, color, decoration, effects, etc., should be considered for transfer. In this paper, we propose a novel approach to stylize target text by decoding weighted deep features from only a few referenced samples. The deep features, including content and style features of each referenced text, are extracted from a Convolutional Neural Network (CNN) that is optimized for character recognition. Then, we calculate the similarity scores of the target text and the referenced samples by measuring the distance along the corresponding channels from the content features of the CNN when considering only the content, and assign them as the weights for aggregating the deep features. To enforce the stylized text to be realistic, a discriminative network with adversarial loss is employed. We demonstrate the effectiveness of our network by conducting experiments on three different datasets which have various styles, fonts, languages, etc. Additionally, the coefficients for character style transfer, including the character content, the effect of similarity matrix, the number of referenced characters, the similarity between characters, and performance evaluation by a new protocol are analyzed for better understanding our proposed framework.
Anna Zhu, Xiongbo Lu, Xiang Bai, Seiichi Uchida, Brian Kenji Iwana, Shengwu Xiong 0001
IEEE Trans. Image Process.5
2019 Dynamic Weight Alignment for Temporal Convolutional Neural Networks
abstract
In this paper, we propose a method of improving temporal Convolutional Neural Networks (CNN) by determining the optimal alignment of weights and inputs using dynamic programming. Conventional CNN convolutions linearly match the shared weights to a window of the input. However, it is possible that there exists a more optimal alignment of weights. Thus, we propose the use of Dynamic Time Warping (DTW) to dynamically align the weights to the input of the convolutional layer. Specifically, the dynamic alignment overcomes issues such as temporal distortion by finding the minimal distance matching of the weights and the inputs under constraints. We demonstrate the effectiveness of the proposed architecture on the Unipen online handwritten digit and character datasets, the UCI Spoken Arabic Digit dataset, and the UCI Activities of Daily Life dataset.
Brian Kenji Iwana, Seiichi Uchida
ICASSP1
2019 On the Ability of a CNN to Realize Image-to-Image Language Conversion
abstract
The purpose of this paper is to reveal the ability that Convolutional Neural Networks (CNN) have on the novel task of image-to-image language conversion. We propose a new network to tackle this task by converting images of Korean Hangul characters directly into images of the phonetic Latin character equivalent. The conversion rules between Hangul and the phonetic symbols are not explicitly provided. The results of the proposed network show that it is possible to perform image-to-image language conversion. Moreover, it shows that it can grasp the structural features of Hangul even from limited learning data. In addition, it introduces a new network to use when the input and output have significantly different features.
Kohei Baba, Seiichi Uchida, Brian Kenji Iwana
ICDAR3
2019 Cascading Modular U-Nets for Document Image Binarization
abstract
In recent years, U-Net has achieved good results in various image processing tasks. However, conventional U-Nets need to be re-trained for individual tasks with enough amount of images with ground-truth. This requirement makes U-Net not applicable to tasks with small amounts of data. In this paper, we propose to use "modular" U-Nets, each of which is pre-trained to perform an existing image processing task, such as dilation, erosion, and histogram equalization. Then, to accomplish a specific image processing task, such as binarization of historical document images, the modular U-Nets are cascaded with inter-module skip connections and fine-tuned to the target task. We verified the proposed model using the Document Image Binarization Competition (DIBCO) 2017 dataset.
Seokjun Kang, Brian Kenji Iwana, Seiichi Uchida
ICDAR2
2019 Selective Super-Resolution for Scene Text Images
abstract
In this paper, we realize the enhancement of super-resolution using images with scene text. Specifically, this paper proposes the use of Super-Resolution Convolutional Neural Networks (SRCNN) which are constructed to tackle issues associated with characters and text. We demonstrate that standard SRCNNs trained for general object super-resolution is not sufficient and that the proposed method is a viable method in creating a robust model for text. To do so, we analyze the characteristics of SRCNNs through quantitative and qualitative evaluations with scene text data. In addition, analysis using the correlation between layers by Singular Vector Canonical Correlation Analysis (SVCCA) and comparison of filters of each SRCNN using t-SNE is performed. Furthermore, in order to create a unified super-resolution model specialized for both text and objects, a model using SRCNNs trained with the different data types and Content-wise Network Fusion (CNF) is used. We integrate the SRCNN trained for character images and then SRCNN trained for general object images, and verify the accuracy improvement of scene images which include text. We also examine how each SRCNN affects super-resolution images after fusion.
Ryo Nakao, Brian Kenji Iwana, Seiichi Uchida
ICDAR2
2019 Modality Conversion of Handwritten Patterns by Cross Variational Autoencoders
abstract
This research attempts to construct a network that can convert online and offline handwritten characters to each other. The proposed network consists of two Variational Auto-Encoders (VAEs) with a shared latent space. The VAEs are trained to generate online and offline handwritten Latin characters simultaneously. In this way, we create a cross-modal VAE (Cross-VAE). During training, the proposed Cross-VAE is trained to minimize the reconstruction loss of the two modalities, the distribution loss of the two VAEs, and a novel third loss called the space sharing loss. This third, space sharing loss is used to encourage the modalities to share the same latent space by calculating the distance between the latent variables. Through the proposed method mutual conversion of online and offline handwritten characters is possible. In this paper, we demonstrate the performance of the Cross-VAE through qualitative and quantitative analysis.
Taichi Sumi, Brian Kenji Iwana, Hideaki Hayashi, Seiichi Uchida
ICDAR2
2019 Deep Dynamic Time Warping: End-to-End Local Representation Learning for Online Signature Verification
abstract
Siamese networks have been shown to be successful in learning deep representations for multivariate time series verification. However, most related studies optimize a global distance objective and suffer from a low discriminative power due to the loss of temporal information. To address this issue, we propose an end-to-end, neural network-based framework for learning local representations of time series, and demonstrate its effectiveness for online signature verification. This framework optimizes a Siamese network with a local embedding loss, and learns a feature space that preserves the temporal location-wise distances between time series. To achieve invariance to non-linear temporal distortion, we propose building a dynamic time warping block on top of the Siamese network, which will greatly improve the accuracy for local correspondences across intra-personal variability. Validation with respect to online signature verification demonstrates the advantage of our framework over existing techniques that use either handcrafted or learned feature representations.
Xiaomeng Wu, Akisato Kimura, Brian Kenji Iwana, Seiichi Uchida, Kunio Kashino
ICDAR3
2019 Capturing Micro Deformations from Pooling Layers for Offline Signature Verification
abstract
In this paper, we propose a novel Convolutional Neural Network (CNN) based method that extracts the location information (displacement features) of the maximums in the max-pooling operation and fuses it with the pooling features to capture the micro deformations between the genuine signatures and skilled forgeries as a feature extraction procedure. After the feature extraction procedure, we apply support vector machines (SVMs) as writer-dependent classifiers for each user to build the signature verification system. The extensive experimental results on GPDS-150, GPDS-300, GPDS-1000, GPDS-2000, and GPDS-5000 datasets demonstrate that the proposed method can discriminate the genuine signatures and their corresponding skilled forgeries well and achieve state-of-the-art results on these datasets.
Yuchen Zheng 0001, Wataru Ohyama, Brian Kenji Iwana, Seiichi Uchida
ICDAR3
2019 Mining the displacement of max-pooling for text recognition
Yuchen Zheng 0001, Brian Kenji Iwana, Seiichi Uchida
Pattern Recognit.2
2018 Contained Neural Style Transfer for Decorated Logo Generation
abstract
Making decorated logos requires image editing skills, without sufficient skills, it could be a time-consuming task. While there are many on-line web services to make new logos, they have limited designs and duplicates can be made. We propose using neural style transfer with clip art and text for the creation of new and genuine logos. We introduce a new loss function based on distance transform of the input image, which allows the preservation of the silhouettes of text and objects. The proposed method contains style transfer to only a designated area. We demonstrate the characteristics of proposed method. Finally, we show the results of logo generation with various input images.
Gantugs Atarsaikhan, Brian Kenji Iwana, Seiichi Uchida
DAS2
2018 Introducing Local Distance-Based Features to Temporal Convolutional Neural Networks
abstract
In this paper, we propose the use of local distance-based features determined by Dynamic Time Warping (DTW) for temporal Convolutional Neural Networks (CNN). Traditionally, DTW is used as a robust distance metric for time series patterns. However, this traditional use of DTW only utilizes the scalar distance metric and discards the local distances between the dynamically matched sequence elements. This paper proposes recovering these local distances, or DTW features, and utilizing them for the input of a CNN. We demonstrate that these features can provide additional information for the classification of isolated handwritten digits and characters. Furthermore, we demonstrate that the DTW features can be combined with the spatial coordinate features in multi-modal fusion networks to achieve state-of-the-art accuracy on the Unipen online handwritten character datasets.
Brian Kenji Iwana, Minoru Mori, Akisato Kimura, Seiichi Uchida
ICFHR1
2018 Discovering Class-Wise Trends of Max-Pooling in Subspace
abstract
The traditional max-pooling operation in Convolutional Neural Networks (CNNs) only obtains the maximal value from a pooling window. However, it discards the information about the precise position of the maximal value. In this paper, we extract the location of the maximal value in a pooling window and transform it into "displacement feature". We analyze and discover the class-wise trend of the displacement features in many ways. The experimental results and discussion demonstrate that the displacement features have beneficial behaviors for solving the problems in max-pooling.
Yuchen Zheng 0001, Brian Kenji Iwana, Seiichi Uchida
ICFHR2
2018 How do Convolutional Neural Networks Learn Design?
abstract
In this paper, we aim to understand the design principles in book cover images which are carefully crafted by experts. Book covers are designed in a unique way, specific to genres which convey important information to their readers. By using Convolutional Neural Networks (CNN) to predict book genres from cover images, visual cues which distinguish genres can be highlighted and analyzed. In order to understand these visual clues contributing towards the decision of a genre, we present the application of Layer-wise Relevance Propagation (LRP) on the book cover image classification results. We use LRP to explain the pixel-wise contributions of book cover design and highlight the design elements contributing towards particular genres. In addition, with the use of state-of-the-art object and text detection methods, insights about genre-specific book cover designs are discovered.
Shailza Jolly, Brian Kenji Iwana, Ryohei Kuroki, Seiichi Uchida
ICPR2
2017 Component Awareness in Convolutional Neural Networks
abstract
In this work, we investigate the ability of Convolutional Neural Networks (CNN) to infer the presence of components that comprise an image. In recent years, CNNs have achieved powerful results in classification, detection, and segmentation. However, these models learn from instance-level supervision of the detected object. In this paper, we determine if CNNs can detect objects using image-level weakly supervised labels without localization. To demonstrate that a CNN can infer awareness of objects, we evaluate a CNN's classification ability with a database constructed of Chinese characters with only character-level labeled components. We show that the CNN is able to achieve a high accuracy in identifying the presence of these components without specific knowledge of the component. Furthermore, we verify that the CNN is deducing the knowledge of the target component by comparing the results to an experiment with the component removed. This research is important for applications with large amounts of data without robust annotation such as Chinese character recognition.
Brian Kenji Iwana, Letao Zhou, Kumiko Tanaka-Ishii, Seiichi Uchida
ICDAR1
2017 Globally Optimal Object Tracking with Complementary Use of Single Shot Multibox Detector and Fully Convolutional Network
Brian Kenji Iwana, Shouta Ide, Hideaki Hayashi, Seiichi Uchida
PSIVT2
2017 Efficient temporal pattern recognition by means of dissimilarity space embedding with discriminative prototypes
Brian Kenji Iwana, Volkmar Frinken, Kaspar Riesen, Seiichi Uchida
Pattern Recognit.1
2016 A Robust Dissimilarity-Based Neural Network for Temporal Pattern Recognition
abstract
Temporal pattern recognition is challenging because temporal patterns require extra considerations over other data types, such as order, structure, and temporal distortions. Recently, there has been a trend in using large data and deep learning, however, many of the tools cannot be directly used with temporal patterns. Convolutional Neural Networks (CNN) for instance are traditionally used for visual and image pattern recognition. This paper proposes a method using a neural network to classify isolated temporal patterns directly. The proposed method uses dynamic time warping (DTW) as a kernel-like function to learn dissimilarity-based feature maps as the basis of the network. We show that using the proposed DTW-NN, efficient classification of on-line handwritten digits is possible with accuracies comparable to state-of-the-art methods.
Brian Kenji Iwana, Volkmar Frinken, Seiichi Uchida
ICFHR1
2016 A Further Step to Perfect Accuracy by Training CNN with Larger Data
abstract
Convolutional Neural Networks (CNN) are on the forefront of accurate character recognition. This paper explores CNNs at their maximum capacity by implementing the use of large datasets. We show a near-perfect performance by using a dataset of about 820,000 real samples of isolated handwritten digits, much larger than the conventional MNIST database. In addition, we report a near-perfect performance on the recognition of machine-printed digits and multi-font digital born digits. Also, in order to progress toward a universal OCR, we propose methods of combining the datasets into one classifier. This paper reveals the effects of combining the datasets prior to training and the effects of transfer learning during training. The results of the proposed methods also show an almost perfect accuracy suggesting the ability of the network to generalize all forms of text.
Seiichi Uchida, Shouta Ide, Brian Kenji Iwana, Anna Zhu
ICFHR3
2015 Tackling temporal pattern recognition by vector space embedding
abstract
This paper introduces a novel method of reducing the number of prototype patterns necessary for accurate recognition of temporal patterns. The nearest neighbor (NN) method is an effective tool in pattern recognition, but the downside is it can be computationally costly when using large quantities of data. To solve this problem, we propose a method of representing the temporal patterns by embedding dynamic time warping (DTW) distance based dissimilarities in vector space. Adaptive boosting (AdaBoost) is then applied for classifier training and feature selection to reduce the number of prototype patterns required for accurate recognition. With a data set of handwritten digits provided by the International Unipen Foundation (iUF), we successfully show that a large quantity of temporal data can be efficiently classified produce similar results to the established NN method while performing at a much smaller cost.
Brian Kenji Iwana, Seiichi Uchida, Kaspar Riesen, Volkmar Frinken
ICDAR1