VLDB 2026 Research / reviewers in the wild / expert
Masakazu Iwamura
dblp:17/3272
· DBLP profile ↗
60ranked-venue papers
11as first author
12since 2021 · last 2026
0000-0003-2508-2869ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 8 first-author · 7 since 2021Databases, data management, data science and information retrieval · 29 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Khmer Camera-Captured Document Layout Detection
Marry Kong, Rina Buoy, Sovisal Chenda, Nguonly Taing, Masakazu Iwamura, Koichi Kise |
ICDAR (3) | 5 |
| 2026 | Addressing the attention drift problem for khmer long textline recognition
Rina Buoy, Sovisal Chenda, Nguonly Taing, Marry Kong, Masakazu Iwamura, Koichi Kise |
Int. J. Document Anal. Recognit. | 5 |
| 2025 | A Depth- and Luminance-Based 2.5D Relief to Expand the Expressive Potential of Tactile Photography for People with Visual ImpairmentsabstractPeople with visual impairments (PVI) face significant challenges in accessing photographs.Although recent generative AI tools provide their verbal descriptions, they are often insufficient to convey rich visual detail.Tactile representations offer an alternative, but conventional 2D tactile photographs have limited expressive capacity.To address this, we propose a novel method to generate 2.5D photographic reliefs integrating depth and luminance information.A user study with 12 participants with visual impairments revealed the advantages of the proposed method and how users' motivations for engaging with photographs influence their tactile explorations. Kosei Takaishi, Kazunori Minatani, Masakazu Iwamura |
ASSETS | 3 |
| 2025 | Disassembling Depth-Estimation-Based 2.5D Photographic Reliefs to Improve Accessibility for People with Visual Impairments
Kosei Takaishi, Kazunori Minatani, Masakazu Iwamura |
ASSETS | 3 |
| 2024 | Grouping Effect for Bar Graph Summarization for People with Visual ImpairmentsabstractWhen communicating numerical data to people with visual impairments (PVI), summaries provided by current data visualization solutions tend to lose important information during summarization. To address this issue, our work focuses on summarization through bar grouping in bar graphs. Neither the effect of grouping nor the appropriate granularity of grouping has been discussed so far. Therefore, we investigate the cognitive effects of grouping and its relationship to the number of groups. A user study involving nine PVI (five blind and four with low vision) revealed that summarization through bar grouping conveys information significantly more accurately compared to simply reading individual data points, despite the inherent error produced by grouping. Additionally, we propose a cognitive error model to explain the characteristics of the observed errors. Banri Kakehi, Masakazu Iwamura, Kazunori Minatani, Koichi Kise |
ASSETS | 2 |
| 2024 | Making 3D Printer Accessible for People with Visual Impairments by Reading Scrolling Text and Menusabstract3D printing has immense potential to enhance the lives of people with visual impairments (PVI) by enabling them to understand shapes and other details through touch that words alone cannot convey. Several initiatives have made 3D tactile models accessible to PVI, yet these models are typically created by sighted individuals. Our goal is to empower PVI to create 3D tactile models independently, making 3D printers accessible to them. The biggest bottleneck for PVI in using 3D printers is the inability to read text on their display. Our work specifically focuses on making scrolling text and menus readable. Through a user study with 13 PVI (five blind and eight with low vision), we confirmed the effectiveness of the implemented functions over conventional smartphone apps and wearable devices. Naoya Tagawa, Masakazu Iwamura, Kazunori Minatani, Koichi Kise |
ASSETS | 2 |
| 2024 | DCDM: Diffusion-Conditioned-Diffusion Model for Scene Text Image Super-Resolution
Shrey Singh, Prateek Keserwani, Masakazu Iwamura, Partha Pratim Roy 0001 |
ECCV (15) | 3 |
| 2024 | Enhanced Cross-Task EEG Classification: Domain Adaptation with EEGNet
Vishal Pandey, Nikhil Panwar, Atharva Kumbhar, Partha Pratim Roy 0001, Masakazu Iwamura |
ICPR (11) | 5 |
| 2024 | Awake at the Wheel: Enhancing Automotive Safety Through EEG-Based Fatigue Detection
Gourav Siddhad, Sayantan Dey, Partha Pratim Roy 0001, Masakazu Iwamura |
ICPR (11) | 4 |
| 2024 | Towards reduced-complexity scene text recognition (RCSTR) through a novel salient feature selection
Rina Buoy, Masakazu Iwamura, Sovila Srun, Koichi Kise |
Int. J. Document Anal. Recognit. | 2 |
| 2024 | Parstr: partially autoregressive scene text recognition
Rina Buoy, Masakazu Iwamura, Sovila Srun, Koichi Kise |
Int. J. Document Anal. Recognit. | 2 |
| 2023 | VisPhoto: Photography for People with Visual Impairments via Post-Production of Omnidirectional Camera ImagingabstractMany people with visual impairments would like to take photographs. However, they often have difficulty pointing the camera at the target. In this paper, we address this problem by proposing a novel photo-taking system called VisPhoto. Unlike conventional methods, VisPhoto generates a photograph in post-production. When the shutter button is pressed, VisPhoto captures an omnidirectional camera image that contains the surrounding scene of the camera. In post-production, the system outputs a cropped region as a “photograph” that satisfies the user’s preference. We conducted an experiment consisting of two parts. First, 24 people with visual impairments took photographs with a genuine iPhone camera app, a conventional method, and VisPhoto. Second, 20 sighted people evaluated the quality of the photographs. The experimental results showed that the participants with visual impairments preferred to use VisPhoto to take photographs of difficult targets, whereas they preferred the conventional method for easy targets. Moreover, we revealed that their preferences for photo-taking methods were influenced by the participants’ needs and values about photography and their confidence in their photographic abilities. Naoki Hirabayashi, Masakazu Iwamura, Kazunori Minatani, Koichi Kise |
ASSETS | 2 |
| 2020 | Quantum Speedup for the Minimum Steiner Tree Problem
Masayuki Miyamoto, Masakazu Iwamura, Koichi Kise, François Le Gall |
COCOON | 2 |
| 2020 | Suitable Camera and Rotation Navigation for People with Visual Impairment on Looking for Something Using Object Detection TechniqueabstractAbstract For people with visual impairment, smartphone apps that use computer vision techniques to provide visual information have played important roles in supporting their daily lives. However, they can be used under a specific condition only. That is, only when the user knows where the object of interest is . In this paper, we first point out the fact mentioned above by categorizing the tasks that obtain visual information using computer vision techniques. Then, in looking for something as a representative task in a category, we argue suitable camera systems and rotation navigation methods. In the latter, we propose novel voice navigation methods. As a result of a user study comprised of seven people with visual impairment, we found that (1) a camera with a wide field of view such as an omnidirectional camera was preferred, and (2) users have different preferences in navigation methods. Masakazu Iwamura, Yoshihiko Inoue, Kazunori Minatani, Koichi Kise |
ICCHP (1) | 1 |
| 2020 | End-to-end Triplet Loss based Emotion Embedding System for Speech Emotion RecognitionabstractIn this paper, an end-to-end neural embedding system based on triplet loss and residual learning has been proposed for speech emotion recognition. The proposed system learns the embeddings from the emotional information of the speech utterances. The learned embeddings are used to recognize the emotions portrayed by given speech samples of various lengths. The proposed system implements Residual Neural Network architecture. It is trained using softmax pretraining and triplet loss function. The weights between the fully connected and embedding layers of the trained network are used to calculate the embedding values. The embedding representations of various emotions are mapped onto a hyperplane, and the angles among them are computed using the cosine similarity. These angles are utilized to classify a new speech sample into its appropriate emotion class. The proposed system has demonstrated 91.67% and 64.44% accuracy while recognizing emotions for RAVDESS and IEMOCAP dataset, respectively. Puneet Kumar 0003, Sidharth Jain, Balasubramanian Raman, Partha Pratim Roy 0001, Masakazu Iwamura |
ICPR | 5 |
| 2020 | Distortion-Adaptive Grape Bunch Counting for Omnidirectional ImagesabstractThis paper proposes the first object counting method for omnidirectional images. Because conventional object counting methods cannot handle the distortion of omnidirectional images, we propose to process them using stereographic projection, which enables conventional methods to obtain a good approximation of the density function. However, the images obtained by stereographic projection are still distorted. Hence, to manage this distortion, we propose two methods. One is a new data augmentation method designed for the stereographic projection of omnidirectional images. The other is a distortion-adaptive Gaussian kernel that generates a density map ground truth while taking into account the distortion of stereographic projection. Using the counting of grape bunches as a case study, we constructed an original grape-bunch image dataset consisting of omnidirectional images and conducted experiments to evaluate the proposed method. The results show that the proposed method performs better than a direct application of the conventional method, improving mean absolute error by 14.7% and mean squared error by 10.5%. Ryota Akai, Yuzuko Utsumi, Yuka Miwa, Masakazu Iwamura, Koichi Kise |
ICPR | 4 |
| 2020 | EEG-Based Cognitive State Assessment Using Deep Ensemble Model and Filter Bank Common Spatial PatternabstractElectroencephalography (EEG) is the most used physiological measure to evaluate the cognitive state of a user efficiently. As EEG inherently suffers from a poor spatial resolution, features extracted from each EEG channel may not be efficiently used for the cognitive state assessment. In this paper, the EEG-based cognitive state assessment has been performed during the mental arithmetic experiment, which includes two cognitive states (task and rest) of a user. To obtain the temporal as well as the spatial resolution of the EEG signal, we combined the Filter Bank Common Spatial Pattern (FBCSP) method and Long Short-Term Memory (LSTM)-based deep ensemble model for classifying the cognitive state of a user. Subject-wise data distribution has been performed due to the execution of a large volume of data in a low computing environment. In the FBCSP method, the input EEG is decomposed into multiple equal-sized frequency bands, and spatial features of each frequency bands are extracted using the Common Spatial Pattern (CSP) algorithm. Next, a feature selection algorithm has been applied to identify the most informative features for classification. The proposed deep ensemble model consists of multiple similar structured LSTM networks that work in parallel. The output of the ensemble model (i.e., the cognitive state of a user) is computed using the average weighted combination of the individual model prediction. This proposed model achieves 87% classification accuracy, and it can also effectively estimate the cognitive state of a user in a low computing environment. Debashis Das Chakladar, Shubhashis Dey, Partha Pratim Roy 0001, Masakazu Iwamura |
ICPR | 4 |
| 2017 | ICDAR2017 Robust Reading Challenge on Omnidirectional VideoabstractResults of ICDAR 2017 Robust Reading Challenge on Omnidirectional Video are presented. This competition uses Downtown Osaka Scene Text (DOST) Dataset that was captured in Osaka, Japan with an omnidirectional camera. Hence, it consists of sequential images (videos) of different view angles. Regarding the sequential images as videos (video mode), two tasks of localisation and end-to-end recognition are prepared. Regarding them as a set of still images (still image mode), three tasks of localisation, cropped word recognition and end-to-end recognition are prepared. As the dataset has been captured in Japan, the dataset contains Japanese text but also include text consisting of alphanumeric characters (Latin text). Hence, a submitted result for each task is evaluated in three ways: using Japanese only ground truth (GT), using Latin only GT and using combined GTs of both. Finally, by the submission deadline, we have received two submissions in the text localisation task of the still image mode. We intend to continue the competition in the open mode. Expecting further submissions, in this report we provide baseline results in all the tasks in addition to the submissions from the community. Masakazu Iwamura, Naoyuki Morimoto, Keishi Tainaka, Dena Bazazian, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICDAR | 1 |
| 2016 | Automatic character labeling for camera captured document imagesabstractCharacter groundtruth for camera captured documents is crucial for training and evaluating advanced OCR algorithms. Manually generating character level groundtruth is a time consuming and costly process. This paper proposes a robust groundtruth generation method based on document retrieval and image registration for camera captured documents. We use an elastic non-rigid alignment method to fit the captured document image which relaxes the flat paper assumption made by conventional solutions. The proposed method allows building very large scale labeled camera captured documents dataset, without any human intervention. We construct a large labeled dataset consisting of 1 million camera captured Chinese character images. Evaluation of samples generated by our approach showed that 99.99% of the images were correctly labeled, even with different distortions specific to cameras such as blur, specularity and perspective distortion. Koichi Kise, Masakazu Iwamura |
ICIP | 3 |
| 2015 | ICDAR 2015 competition on Robust ReadingabstractResults of the ICDAR 2015 Robust Reading Competition are presented. A new Challenge 4 on Incidental Scene Text has been added to the Challenges on Born-Digital Images, Focused Scene Images and Video Text. Challenge 4 is run on a newly acquired dataset of 1,670 images evaluating Text Localisation, Word Recognition and End-to-End pipelines. In addition, the dataset for Challenge 3 on Video Text has been substantially updated with more video sequences and more accurate ground truth data. Finally, tasks assessing End-to-End system performance have been introduced to all Challenges. The competition took place in the first quarter of 2015, and received a total of 44 submissions. Only the tasks newly introduced in 2015 are reported on. The datasets, the ground truth specification and the evaluation protocols are presented together with the results and a brief summary of the participating methods. Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukás Neumann, Vijay Chandrasekhar 0001, Shijian Lu, Faisal Shafait, Seiichi Uchida, Ernest Valveny |
ICDAR | 6 |
| 2015 | Preface
Faisal Shafait, Dimosthenis Karatzas, Seiichi Uchida, Masakazu Iwamura |
Int. J. Document Anal. Recognit. | 4 |
| 2014 | Performance Improvement in Local Feature Based Camera-Captured Character RecognitionabstractConcerning camera-captured Japanese character recognition, we have proposed a method to recognize characters, both simple and complex, that may not be linearly aligned and may be printed with a complex background. Recognition is performed based on local features and their arrangement. The arrangement is validated with an algorithm called local RANSAC. However, at least four corresponding local features are required. To relax that condition, we propose a new recognition method making it possible to recognize a character region with at least three corresponding local features. This method enables recall and precision to be improved with the simpler characters using more corresponding local features and computation times to be reduced by 7%. Takahiro Matsuda 0001, Masakazu Iwamura, Koichi Kise |
Document Analysis Systems | 2 |
| 2014 | A mixed reality head-mounted text translation system using eye gaze inputabstractEfficient text recognition has recently been a challenge for augmented reality systems. In this paper, we propose a system with the ability to provide translations to the user in real-time. We use eye gaze for more intuitive and efficient input for ubiquitous text reading and translation in head mounted displays (HMDs). The eyes can be used to indicate regions of interest in text documents and activate optical-character-recognition (OCR) and translation functions. Visual feedback and navigation help in the interaction process, and text snippets with translations from Japanese to English text snippets, are presented in a see-through HMD. We focus on travelers who go to Japan and need to read signs and propose two different gaze gestures for activating the OCR text reading and translation function. We evaluate which type of gesture suits our OCR scenario best. We also show that our gaze-based OCR method on the extracted gaze regions provide faster access times to information than traditional OCR approaches. Other benefits include that visual feedback of the extracted text region can be given in real-time, the Japanese to English translation can be presented in real-time, and the augmentation of the synchronized and calibrated HMD in this mixed reality application are presented at exact locations in the augmented user view to allow for dynamic text translation management in head-up display systems. Takumi Toyama, Daniel Sonntag, Andreas Dengel 0001, Takahiro Matsuda 0001, Masakazu Iwamura, Koichi Kise |
IUI | 5 |
| 2014 | Fast Instance Search Based on Approximate Bichromatic Reverse Nearest Neighbor SearchabstractIn the TRECVID Instance Search (INS) task, it is known that use of BM25, which is an improvement of the TFIDF,greatly improves retrieval performance. Its calculation, however, requires tremendous amount of computational cost and this fact makes its use intractable. In this paper, we present its efficient computational method. Since the BM25 is obtained by solving the bichromatic reverse nearest neighbor (BRNN)search problem,we propose an approximate method for the problem based on the state-of-the-art approximate nearest neighbor search method, bucket distance hashing (BDH). An experiment using the TRECVID INS 2012 dataset showed that the proposed method reduced computational cost to less than 1/3500 of the brute-force search with keeping the accuracy. Masakazu Iwamura, Nobuaki Matozaki, Koichi Kise |
ACM Multimedia | 1 |
| 2014 | Recovery and localization of handwritings by a camera-pen based on tracking and document image retrieval
Megumi Chikano, Koichi Kise, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi |
Pattern Recognit. Lett. | 3 |
| 2014 | More than ink - Realization of a data-embedding pen
Marcus Liwicki, Seiichi Uchida, Akira Yoshida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
Pattern Recognit. Lett. | 4 |
| 2013 | What is the Most EfficientWay to Select Nearest Neighbor Candidates for Fast Approximate Nearest Neighbor Search?abstractApproximate nearest neighbor search (ANNS) is a basic and important technique used in many tasks such as object recognition. It involves two processes: selecting nearest neighbor candidates and performing a brute-force search of these candidates. Only the former though has scope for improvement. In most existing methods, it approximates the space by quantization. It then calculates all the distances between the query and all the quantized values (e.g., clusters or bit sequences), and selects a fixed number of candidates close to the query. The performance of the method is evaluated based on accuracy as a function of the number of candidates. This evaluation seems rational but poses a serious problem; it ignores the computational cost of the process of selection. In this paper, we propose a new ANNS method that takes into account costs in the selection process. Whereas existing methods employ computationally expensive techniques such as comparative sort and heap, the proposed method does not. This realizes a significantly more efficient search. We have succeeded in reducing computation times by one-third compared with the state-of-theart on an experiment using 100 million SIFT features. Masakazu Iwamura, Tomokazu Sato, Koichi Kise |
ICCV | 1 |
| 2013 | Automatic Ground Truth Generation of Camera Captured Documents Using Document Image RetrievalabstractIn this paper a novel method for automatic ground truth generation of camera captured document images is proposed. Currently, no dataset is available for camera captured documents. It is very difficult to build these datasets manually, as it is very laborious and costly. The proposed method is fully automatic, allowing building the very large scale (i.e., millions of images) labeled camera captured documents dataset, without any human intervention. Evaluation of samples generated by the proposed approach shows that 99.98% of the images are correctly labeled. Novelty of the proposed approach lies in the use of document image retrieval for automatic labeling, especially for camera captured documents, which contain different distortions specific to camera, e.g., blur, occlusion, perspective distortion, etc. Sheraz Ahmed, Koichi Kise, Masakazu Iwamura, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 3 |
| 2013 | Key-Region Detection for Document Images - Application to Administrative Document RetrievalabstractIn this paper we argue that a key-region detector designed to take into account the special characteristics of document images can result in the detection of less and more meaningful key-regions. We propose a fast key-region detector able to capture aspects of the structural information of the document, and demonstrate its efficiency by comparing against standard detectors in an administrative document retrieval scenario. We show that using the proposed detector results to a smaller number of detected key-regions and higher performance without any drop in speed compared to standard state of the art detectors. Hongxing Gao, Marçal Rusiñol, Dimosthenis Karatzas, Josep Lladós 0001, Tomokazu Sato, Masakazu Iwamura, Koichi Kise |
ICDAR | 6 |
| 2013 | Automatic Labeling for Scene Text DatabaseabstractIt is thought that a large quantity of data improve quality of recognition. A large database, however, is not easy to obtain. The hardest task is labeling (also known as ground truthing), which usually requires human intervention. Since labeling by human is laborious and costly, labeling without human (automatic labeling) or minimization of human intervention (semi-automatic labeling) are ideal scenarios. As a step toward realization of the scenarios, knowing how much an automatic labeling system can perform without human intervention is important. In the current paper we propose a comprehensive automatic labeling technique for a scene text database, which performs segmentation and labeling for unsegmented and unlabeled character images. To our best knowledge, this is the first method to realize the comprehensive process for automatic labeling for scene text databases In experiments, we confirmed that the proposed method could add new unlabeled data in parallel with improving recognition performance of the classifier. Masakazu Iwamura, Masaki Tsukada, Koichi Kise |
ICDAR | 1 |
| 2013 | ICDAR 2013 Robust Reading CompetitionabstractThis report presents the final results of the ICDAR 2013 Robust Reading Competition. The competition is structured in three Challenges addressing text extraction in different application domains, namely born-digital images, real scene images and real-scene videos. The Challenges are organised around specific tasks covering text localisation, text segmentation and word recognition. The competition took place in the first quarter of 2013, and received a total of 42 submissions over the different tasks offered. This report describes the datasets and ground truth specification, details the performance evaluation protocols used and presents the final results along with a brief summary of the participating methods. Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluís Gómez i Bigorda, Sergi Robles, Joan Mas Romeu, David Fernández Mota, Jon Almazán, Lluís-Pere de las Heras |
ICDAR | 4 |
| 2013 | The Reading-Life Log - Technologies to Recognize Texts That We ReadabstractReading life log is a type of techniques to automatically and unconsciously record people's reading intentions, interests and habits. Besides, it can also serve as various assistants in our daily life. In this paper, a reading-life log system is implemented by a head-mounted and unobtrusive video camera with a high resolution and a high shutter speed. We utilize DP matching, and propose a text-based frame mosaicing method to integrate multiple frames in a clip. The developed system is tested in the various environments indoor and outdoor. The experimental results show that our system can provide reliable outputs with respect to the most correct responses. The infrequent misregistration between lines also indicates the feasibility and validity of the text-based frame mosaicing. Takashi Kimura, Rong Huang 0003, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICDAR | 4 |
| 2013 | An Anytime Algorithm for Camera-Based Character RecognitionabstractIn a scene image, some characters are difficult to recognize and some others are recognized easily. Such difficult characters usually make the processing time long while easy characters are recognized in a short time. In this paper, we propose a system which recognizes each character with a proper cost for the difficulty. Through the process, easy characters are recognized early and difficult ones are recognized late. This is a desired property of an anytime algorithm that the recognition accuracy does not decrease as the time increases. In order to realize it, we propose a method which splits the recognition process into several times and accumulates the recognition results and extracted features. We also discuss what is required to realize the anytime algorithm for the scene character recognition task. Experiments reveal that the proposed method obtains recognition results of easy characters earlier than the conventional method. Takuya Kobayashi, Masakazu Iwamura, Takahiro Matsuda 0001, Koichi Kise |
ICDAR | 2 |
| 2012 | Recognizing Words in Scenes with a Head-Mounted Eye-TrackerabstractRecognition of scene text using a hand-held camera is emerging as a hot topic of research. In this paper, we investigate the use of a head-mounted eye-tracker for scene text recognition. An eye-tracker detects the position of the user's gaze. Using gaze information of the user, we can provide the user with more information about his region/object of interest in a ubiquitous manner. Therefore, we can realize a service such as the user gazes at a certain word and soon obtain the related information of the word by combining a word recognition system with eye-tracking technology. Such a service is useful since the user has to do nothing but gazes at interested words. With a view to realize the service, we experimentally evaluate the effectiveness of using the eye-tracker for word recognition. The initial results show the recognition accuracy was around 70% in our word recognition experiment and the average computational time was less than one second per a query image. Takuya Kobayashi, Takumi Toyama, Faisal Shafait, Masakazu Iwamura, Koichi Kise, Andreas Dengel 0001 |
Document Analysis Systems | 4 |
| 2012 | Real-Time Document Image Retrieval on a SmartphoneabstractThis paper presents a novel interface running on smart phones which is capable of seamlessly linking physical and digital worlds through paper documents. This interface is based on a real-time document image retrieval method called Locally Likely Arrangement Hashing. By just only pointing a smart phone to a paper document, the user can obtain its corresponding electronic document. This can easily provide the user with the information associated with the retrieved document. This relevant information can be superimposed on the display of smart phones. Therefore, we consider that with the help of this interface, the user can utilize paper documents as a new medium to display various information. Kazutaka Takeda, Koichi Kise, Masakazu Iwamura |
Document Analysis Systems | 3 |
| 2012 | Expanding Recognizable Distorted Characters Using Self-Corrective RecognitionabstractLarge datasets are always demanded for better recognition performance. However, it is not easy to produce them because costly and slow human operators have been necessary for labeling. In the current paper, in order to resolve the problem on yielding large datasets, we propose a scenario for automatic labeling based on the self-corrective recognition algorithm. The strong point of the proposed method is the capability of expanding recognizable distorted characters unlike existing methods. In the experiments, we show a possibility to realize automatic labeling by the method. Masaki Tsukada, Masakazu Iwamura, Koichi Kise |
Document Analysis Systems | 2 |
| 2011 | Recognition of Multiple Characters in a Scene Image Using Arrangement of Local FeaturesabstractRecognizing characters in a scene helps us obtain useful information. For the purpose, character recognition methods are required to recognize characters of various sizes, various rotation angles and complex layout on complex background. In this paper, we propose a character recognition method using local features having several desirable properties. The novelty of the proposed method is to take into account arrangement of local features so as to recognize multiple characters in an image unlike past methods. The effectiveness and possible improvement of the method are discussed. Masakazu Iwamura, Takuya Kobayashi, Koichi Kise |
ICDAR | 1 |
| 2011 | Reliable Online Stroke Recovery from Offline Data with the Data-Embedding PenabstractIn this paper we propose a complete system for online stroke recovery from offline data. The key idea of our approach is to use a novel pen device which is able to embed meta information into the ink during writing the strokes. This pen-device overcomes the need to get access to any memory on the pen when trying to recover the information, which is especially useful in multi-writer or multi-pen scenarios. The actual data-embedding is achieved by an additional ink dot sequence along a handwritten pattern during writing. We design the ink-dot sequence in such a way that it is possible to retrieve the writing direction from a scanned image. Furthermore, we propose novel processing steps in order to retrieve the original writing direction and finally the embedded data. In our experiments we show that we can reliably recover the writing direction of various patterns. Our system is able to determine the writing direction of straight lines, simple patterns with crossings (e.g., "x" and "II"), and even more complex patterns like handwritten words and symbols. Marcus Liwicki, Akira Yoshida, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICDAR | 4 |
| 2011 | Real-Time Document Image Retrieval for a 10 Million Pages Database with a Memory Efficient and Stability Improved LLAHabstractThis paper presents a real-time document image retrieval method for a large-scale database with Locally Likely Arrangement Hashing (LLAH). In general, when a database is scaled up, a large amount of memory is required and retrieval accuracy drops due to insufficient discrimination power of features. To solve these problems, we propose three improvements: memory reduction by sampling feature points, improvement of discrimination power by increasing the number of feature dimensions and stabilizing features by reducing redundancy. From the experimental results, we have confirmed that the proposed method realizes 50% memory reduction, and achieves 99.4% accuracy and 38ms processing time for a database of 10 million pages. Kazutaka Takeda, Koichi Kise, Masakazu Iwamura |
ICDAR | 3 |
| 2011 | Handwriting on Paper as a Cybermedium
Akira Yoshida, Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
KES (4) | 4 |
| 2010 | Memory-based recognition of camera-captured charactersabstractThis paper addresses how to quickly recognize a character pattern using a lot of case examples without learning. Here without learning means just finding the most similar example from the case examples, and pretend as if the OCR understands the definition of the character. This strategy is expected to work well in most cases with a large dataset, however, also expected to take a lot of time for finding the most similar example. In this paper, we show that a lot of case examples can be processed in a short time. As a testbed, we handle recognition problem of camera-captured printed characters. Using a database storing 100 fonts, the proposed method achieved 97.0% of recognition rate for images captured from the right angle and 95.8% for those from 45 deg. with 4.56ms of processing time, that is about 220 characters per second including every process. Masakazu Iwamura, Tomohiko Tsuji, Koichi Kise |
Document Analysis Systems | 1 |
| 2010 | Expansion of queries and databases for improving the retrieval accuracy of document portions: an application to a camera-pen systemabstractThis paper presents a method of improving the accuracy of document image retrieval focusing on the application to a camera-pen system. In a camera-pen system, document image retrieval is employed for locating the pen-tip position on a page. A serious problem is that since the camera is mounted close to the pen-tip, the camera captures only a tiny portion of the page and the resultant image is under severe perspective distortion, resulting in lowering the retrieval accuracy. To solve this problem, we propose new geometrically invariant features as well as expansion techniques which increase the number of index features of either the database or the query images. From the experimental results, it has been found that the query expansion technique with features by combining affine and perspective invariants allows us the best performance that improves the accuracy of a baseline method more than 27%. Koichi Kise, Megumi Chikano, Kazumasa Iwata, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi |
Document Analysis Systems | 4 |
| 2010 | Data-embedding pen: augmenting ink strokes with meta-informationabstractIn this paper we present the first operational version of the data-embedding pen. During writing a pattern, this pen produces an additional ink-dot sequence along the ink stroke of the pattern. The ink-dot sequence represents, for example, meta-information (such as the writer's name and the date of writing) and thus drastically increases the value of the handwriting on a physical paper. Since the information is placed on the paper, it can be extracted just by scanning or photographing the paper. There is no need to get access to any memory on the pen to recover the information. This is useful especially in multi-writer or multi-pen scenarios. The experiments using an encoding scheme and a decoding algorithm showed very promising results. For example, it was proved that we can embed 28 or more bits of information on simple handwritten patterns and decode them with a high reliability. Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
Document Analysis Systems | 3 |
| 2010 | Tracking and Retrieval of Pen Tip Positions for an Intelligent Camera PenabstractThis paper presents a method of recovering digital ink for an intelligent camera pen, which is characterized by the functions that (1) it works on ordinary paper and (2) if an electronic document is printed on the paper the recovered digital ink is associated with the document. Two technologies called paper fingerprint and document image retrieval are integrated for realizing the above functions. The key of the integration is the introduction of image mosaicing and fast retrieval of previously seen fingerprints based on hashing of SURF local features. From the experimental results of 50 handwritings, we have confirmed that the proposed method is effective to recover and locate the digital ink from the handwriting on a physical paper. Kazumasa Iwata, Koichi Kise, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi |
ICFHR | 3 |
| 2010 | Embedding Meta-Information in Handwriting -- Reed-Solomon for Reliable Error CorrectionabstractIn this paper a more compact and more reliable coding scheme for the data-embedding pen is proposed. The data-embedding pen produces an additional ink-dot sequence along a handwritten pattern during writing. The ink-dot sequence represents, for example, meta-information (such as the writer's name and the date of writing) and thus drastically increases the value of the handwriting on a physical paper. There is no need to get access to any memory on the pen to recover the information, which is especially useful in multi-writer or multi-pen scenarios. In this paper we focus on the compactness of the encoded information. The aim of this paper is to encode as much information as possible in short stroke sequences. In our experiments we show that we can embed more information in shorter strokes than in previous work. In straight lines as short as 5 cm, 32 bits can successfully be embedded. Furthermore, the new encoding scheme also works reliably on more complex patterns. Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICFHR | 3 |
| 2010 | Handwriting Reconstruction for a Camera Pen Using Random Dot PatternsabstractThis paper proposes a new method of handwriting reconstruction using a camera pen. We print random dot patterns on the document background to enable retrieval of both the current document and the pen position on this document. Dot arrangements are stored in a hash table using Locally Likely Arrangement Hashing. For retrieval, they are extracted from the camera image and matched to the corresponding points in the hash table. We were able to achieve high retrieval accuracy (81.1~100.0%), given a sufficient amount of visible dots. Using a two-step homography approximation, an accurate image of handwriting can be reconstructed. By using knowledge about document context and a client-server architecture, our method allows real-time processing on ordinary hardware. Matthias Sperber, Martin Klinkigt, Koichi Kise, Masakazu Iwamura, Benjamin Adrian, Andreas Dengel 0001 |
ICFHR | 4 |
| 2009 | Real-Time Camera-Based Recognition of Characters and PictogramsabstractCamera-based character recognition systems should have the capability of quick operation and recognizing perspectively distorted texts in a complex layout. In this paper, in order to realize such a system, we propose a simple but efficient implementation of camera-based recognition of characters and pictograms. With help of new hashing and voting techniques, the proposed method runs well in real-time even on a laptop PC with a web camera. Masakazu Iwamura, Tomohiko Tsuji, Akira Horimatsu, Koichi Kise |
ICDAR | 1 |
| 2009 | Capturing Digital Ink as Retrieving Fragments of Document ImagesabstractThis paper presents a new method of capturing digital ink for pen-based computing. Current technologies such as tablets, ultrasonic and the Anoto pens rely on special mechanisms for locating the pen tip,which result in limiting the applicability.Our proposal is to ease this problem --- a camera pen that allows us to write on ordinary paper for capturing digital ink. A document image retrieval method called LLAH is tuned to locate the pen tip efficiently and accurately on the coordinates of a document only by capturing its tiny fragment.In this paper, we report some results on captured digital ink as well as to evaluate their quality. Kazumasa Iwata, Koichi Kise, Tomohiro Nakai, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi |
ICDAR | 4 |
| 2009 | Real-Time Retrieval for Images of Documents in Various Languages Using a Web CameraabstractWe propose a real-time retrieval method for document images in various languages. In this method, queries are images of documents captured by a web-camera. The document images corresponding to the queries are retrieved from the document image database in real time. Since we have already proposed a document image retrieval method for English documents, the proposed method is an extension for retrieval of documents in various languages. In the previous English document image retrieval method, only centroids of word regions are used as feature points. Therefore it cannot be applied to some languages including Japanese and Chinese due to no separation between words and periodic arrangements of characters. In the proposed method, additional features are introduced to realize real-time retrieval for document images in various languages. Tomohiro Nakai, Koichi Kise, Masakazu Iwamura |
ICDAR | 3 |
| 2009 | Conspicuous Character PatternsabstractDetection of characters in scenery images is often a very difficult problem. Although many researchers have tackled this difficult problem and achieved a good performance, it is still difficult to suppress many false alarms and although missings. This paper investigates a conspicuous character pattern, which is a special pattern designed for easier detection. In order to have an example of the conspicuous character pattern, we select a character font with a larger distance from a non-character pattern distribution and, simultaneously, with a smaller distance from a character pattern distribution. Experimental results showed that the character font selected by this method is actually more conspicuous (i.e., detected more easily) than other fonts. Seiichi Uchida, Ryoji Hattori, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICDAR | 3 |
| 2008 | Affine Invariant Recognition of Characters by Progressive PruningabstractThere are many problems to realize camera-based character recognition. One of the problems is that characters in scenes are often distorted by geometric transformations such as affine distortions. Although some methods that remove the affine distortions have been proposed, they cannot remove a rotation transformation of a character. Thus a skew angle of a character has to be determined by examining all the possible angles. However, this consumes quite a bit of time. In this paper, in order to reduce the processing time for an affine invariant recognition, we propose a set of affine invariant features and a new recognition scheme called "progressive pruning."' The progressive pruning gradually prunes less feasible categories and skew angles using multiple classifiers. We confirmed the progressive pruning with the affine invariant features reduced the processing time at least less than half without decreasing the recognition rate. Akira Horimatsu, Ryo Niwa, Masakazu Iwamura, Koichi Kise, Seiichi Uchida, Shinichiro Omachi |
Document Analysis Systems | 3 |
| 2008 | Skew Estimation by InstancesabstractThis paper proposes a novel skew estimation method by instances. The instances to be learned (i.e., stored) are rotation invariants and a rotation variant for each character category. Using the instances, it is possible to estimate a skew angle of each individual character on a document. This fact implies that the proposed method can estimate the skew angle of a document where characters do not form long straight text lines. Thus, the proposed method will be applicable to various documents such as signboard images captured by a camera. Experimental evaluation using synthetic and real images revealed the expected robustness against various character layouts. Seiichi Uchida, Megumi Sakai, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
Document Analysis Systems | 3 |
| 2008 | Memory efficient recognition of specific objects with local featuresabstractBalancing the recognition rate, processing time and memory requirement is an important issue for object recognition based on local features. For the task of recognizing not generic but specific objects (object instances), a larger number of local features enable us to improve the recognition rate but pose a problem of processing time and memory requirement. For the problem of processing time, approximate nearest neighbor search is known to be extremely effective. In this paper, we propose a method of memory reduction by applying scalar quantization to local features. From experimental results on 100,000 images, we have found that the scalar quantization is of great help; As compared to the original representation with 16 bit/dimension, the recognition rate of 98.1 % was kept unchanged using the representation with 2 bit/dim. Koichi Kise, Kazuto Noguchi, Masakazu Iwamura |
ICPR | 3 |
| 2007 | Improvement of Retrieval Speed and Required Amount of Memory for Geometric Hashing by Combining Local InvariantsabstractThe geometric hashing (GH) is a well-known model-based object recognition technique with good properties both in retrieval speed and required amount of memory. However, it has a significant weak point; as the number of objects increases, both retrieval speed and required amount of memory increase in the cubic, fourth or higher order. Recently, a new technique “locally likely arrangement hashing (LLAH) ” whose computational cost is a linear order has been proposed. The objective of the current paper is to reveal how LLAH improves the performance. By comparing GH and LLAH, we describe four primary factors of the performance improvement. 1 Masakazu Iwamura, Tomohiro Nakai, Koichi Kise |
BMVC | 1 |
| 2007 | Simple Representation and Approximate Search of Feature Vectors for Large-Scale Object RecognitionabstractThis paper presents two methods of large-scale recognition of planar objects with a simple representation and approximate search of local feature vectors. A central problem of the use of local feature vectors is the burden of computation and memory for finding nearest neighbors. To solve this problem, the proposed methods embody the following: (1) a simple bit representation of feature vectors and hashing enable us to fast access with less memory, (2) approximate search with query perturbation allows us to find approximate nearest neighbors efficiently. From large-scale experiments using 10,000 objects in the database and 2,000 query images, it was found that only 10–20% of correct nearest neighbors were enough for achieving recognition rate of 98.0%. The processing time for achieving this rate was 8.3 ms / query (excluding time for feature extraction). We have also tested the scalability of a proposed method using the database of 100,000 objects and obtained the result of 92.3 % accuracy in 4.5 ms /query. 1 Koichi Kise, Kazuto Noguchi, Masakazu Iwamura |
BMVC | 3 |
| 2007 | Extraction of Embedded Class Information from Universal Character PatternabstractThis paper is concerned with a universal pattern, which is defined as a character pattern designed to have high machine-readability. This universal pattern is a charac- ter pattern printed with stripes. The cross ratio calculated from the widths of the stripes represents the character class. Thus, if the boundaries of the stripes can be detected for measuring the widths, the class can be determined without ordinary recognition process. Furthermore, since the cross ratio is invariant to projective distortions, the correct class will be still determined under those distortions. This pa- per describes a practical scheme to recognize this universal pattern. The proposed scheme includes a novel algorithm to detect the stripe boundaries stably even from the universal pattern image contaminated by non-uniform lighting and noise. The algorithm is realized by a combination of a dy- namic programming-based optimal boundary detection and a finite state automaton which represents the property of the universal pattern. Experimental results showed the pro- posed scheme could recognize 99.6% of the universal pat- tern images which underwent heavy projective distortions and non-uniform lighting. Seiichi Uchida, Megumi Sakai, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICDAR | 3 |
| 2006 | Use of Affine Invariants in Locally Likely Arrangement Hashing for Camera-Based Document Image Retrieval
Tomohiro Nakai, Koichi Kise, Masakazu Iwamura |
Document Analysis Systems | 3 |
| 2005 | Isolated Character Recognition by Searching Feature PointsabstractConventional segmentation technique cannot extract difficult characters such as an isolated character and a touching character. In this paper, we propose a novel character recognition method which executes segmentation and recognition simultaneously. This method enables us to extract and recognize such difficult characters. The effectiveness of the proposed method is confirmed by experiments. Masakazu Iwamura, Kazuya Negishi, Shinichiro Omachi, Hirotomo Aso |
ICDAR | 1 |
| 2005 | Camera-Based Document Image Retrieval as Voting for Partial Signatures of Projective InvariantsabstractWe propose a method of document image retrieval using digital cameras. The proposed method takes as input a part or the whole of a document acquired as a query by a digital camera, and retrieves a document image that includes the query. For this purpose, it is required to solve the problem of "perspective distortion" of images, as well as to establish a way of matching parts of document images flexibly. These are achieved based on the following characteristics of the proposed method: (1) indexing of document images using the projective invariants called the "cross-ratios", (2) retrieval as voting for partial signatures of document images defined by the cross-ratios. From experimental results using digital cameras with high and low resolutions, we demonstrate the effectiveness of the proposed method. Tomohiro Nakai, Koichi Kise, Masakazu Iwamura |
ICDAR | 3 |
| 2000 | A Modification of Eigenvalues to Compensate Estimation Errors of EigenvectorsabstractIn statistical pattern recognition, parameters of distributions are usually estimated from training samples. It is well known that shortage of training samples causes estimation errors which reduce recognition accuracy. By studying estimation errors of eigenvalues, various methods of avoiding recognition accuracy reduction have been proposed. However, estimation errors of eigenvectors have not been considered enough. In this paper, we investigate estimation errors of eigenvectors to show these errors are another factor of recognition performance reduction. We propose a new method for modifying eigenvalues in order to reduce bad influence caused by estimation errors of eigenvectors. Effectiveness of the method is shown by experimental results. Masakazu Iwamura, Shinichiro Omachi, Hirotomo Aso |
ICPR | 1 |