VLDB 2026 Research / reviewers in the wild / expert
Koichi Kise
dblp:k/KoichiKise
· DBLP profile ↗
103ranked-venue papers
19as first author
16since 2021 · last 2026
0000-0001-5779-6968ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 15 first-author · 9 since 2021Databases, data management, data science and information retrieval · 55 · 10 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-authorHuman-computer interaction and ubiquitous computing · 11 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSecurity and privacy · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Non-Latin Text and Layout Personalization for Enhanced Readability
Rina Buoy, Dylan Berkamp Fouepe Dongmo, Vesal Khean, Simone Marinai, Koichi Kise |
ICDAR (3) | 5 |
| 2026 | Prediction of Grade, Gender, and Academic Performance of Children and Teenagers from Handwriting Using the Sigma-Lognormal Model
Adrian Iste, Kazuki Nishizawa, Chisa Tanaka, Andrew W. Vargo, Anna Scius-Bertrand, Koichi Kise |
ICDAR (3) | 7 |
| 2026 | Towards Khmer Camera-Captured Document Layout Detection
Marry Kong, Rina Buoy, Sovisal Chenda, Nguonly Taing, Masakazu Iwamura, Koichi Kise |
ICDAR (3) | 6 |
| 2026 | Addressing the attention drift problem for khmer long textline recognition
Rina Buoy, Sovisal Chenda, Nguonly Taing, Marry Kong, Masakazu Iwamura, Koichi Kise |
Int. J. Document Anal. Recognit. | 6 |
| 2026 | International Journal on Document Analysis and Recognition editorial leadership change
Daniel P. Lopresti, Koichi Kise, Josep Lladós 0001 |
Int. J. Document Anal. Recognit. | 2 |
| 2025 | International journal on document analysis and recognition editorial leadership change
Daniel P. Lopresti, Koichi Kise, Simone Marinai |
Int. J. Document Anal. Recognit. | 2 |
| 2024 | Grouping Effect for Bar Graph Summarization for People with Visual ImpairmentsabstractWhen communicating numerical data to people with visual impairments (PVI), summaries provided by current data visualization solutions tend to lose important information during summarization. To address this issue, our work focuses on summarization through bar grouping in bar graphs. Neither the effect of grouping nor the appropriate granularity of grouping has been discussed so far. Therefore, we investigate the cognitive effects of grouping and its relationship to the number of groups. A user study involving nine PVI (five blind and four with low vision) revealed that summarization through bar grouping conveys information significantly more accurately compared to simply reading individual data points, despite the inherent error produced by grouping. Additionally, we propose a cognitive error model to explain the characteristics of the observed errors. Banri Kakehi, Masakazu Iwamura, Kazunori Minatani, Koichi Kise |
ASSETS | 4 |
| 2024 | Making 3D Printer Accessible for People with Visual Impairments by Reading Scrolling Text and Menusabstract3D printing has immense potential to enhance the lives of people with visual impairments (PVI) by enabling them to understand shapes and other details through touch that words alone cannot convey. Several initiatives have made 3D tactile models accessible to PVI, yet these models are typically created by sighted individuals. Our goal is to empower PVI to create 3D tactile models independently, making 3D printers accessible to them. The biggest bottleneck for PVI in using 3D printers is the inability to read text on their display. Our work specifically focuses on making scrolling text and menus readable. Through a user study with 13 PVI (five blind and eight with low vision), we confirmed the effectiveness of the implemented functions over conventional smartphone apps and wearable devices. Naoya Tagawa, Masakazu Iwamura, Kazunori Minatani, Koichi Kise |
ASSETS | 4 |
| 2024 | Towards reduced-complexity scene text recognition (RCSTR) through a novel salient feature selection
Rina Buoy, Masakazu Iwamura, Sovila Srun, Koichi Kise |
Int. J. Document Anal. Recognit. | 4 |
| 2024 | Parstr: partially autoregressive scene text recognition
Rina Buoy, Masakazu Iwamura, Sovila Srun, Koichi Kise |
Int. J. Document Anal. Recognit. | 4 |
| 2023 | VisPhoto: Photography for People with Visual Impairments via Post-Production of Omnidirectional Camera ImagingabstractMany people with visual impairments would like to take photographs. However, they often have difficulty pointing the camera at the target. In this paper, we address this problem by proposing a novel photo-taking system called VisPhoto. Unlike conventional methods, VisPhoto generates a photograph in post-production. When the shutter button is pressed, VisPhoto captures an omnidirectional camera image that contains the surrounding scene of the camera. In post-production, the system outputs a cropped region as a “photograph” that satisfies the user’s preference. We conducted an experiment consisting of two parts. First, 24 people with visual impairments took photographs with a genuine iPhone camera app, a conventional method, and VisPhoto. Second, 20 sighted people evaluated the quality of the photographs. The experimental results showed that the participants with visual impairments preferred to use VisPhoto to take photographs of difficult targets, whereas they preferred the conventional method for easy targets. Moreover, we revealed that their preferences for photo-taking methods were influenced by the participants’ needs and values about photography and their confidence in their photographic abilities. Naoki Hirabayashi, Masakazu Iwamura, Kazunori Minatani, Koichi Kise |
ASSETS | 5 |
| 2023 | Exploring Users' Ability to Choose a Proper Fit in Smart-Rings: A Year-Long "In the Wild" Study
Peter Neigel, Andrew W. Vargo, Yusuke Komatsu, Chris Blakely, Koichi Kise |
INTERACT (4) | 5 |
| 2023 | Perception Versus Reality: How User Self-reflections Compare to Actual Data
Hannah R. Nolasco, Andrew W. Vargo, Yusuke Komatsu, Motoi Iwata, Koichi Kise |
INTERACT (3) | 5 |
| 2023 | Intelligence Augmentation: Future Directions and Ethical Implications in HCI
Andrew W. Vargo, Benjamin Tag, Mathilde Hutin, Victoria Abou Khalil, Shoya Ishimaru, Olivier Augereau, Tilman Dingler, Motoi Iwata, Koichi Kise, Laurence Devillers, Andreas Dengel 0001 |
INTERACT (4) | 9 |
| 2023 | Editorial for special issue on "advanced topics in document analysis and recognition"
Koichi Kise, Richard Zanibbi, Rajiv Jain, Gernot A. Fink |
Int. J. Document Anal. Recognit. | 1 |
| 2021 | Quality Assessment of Crowdwork via Eye Gaze: Towards Adaptive Personalized Crowdsourcing
Md. Rabiul Islam 0008, Shun Nawa, Andrew W. Vargo, Motoi Iwata, Masaki Matsubara, Atsuyuki Morishima, Koichi Kise |
INTERACT (2) | 7 |
| 2020 | Quantum Speedup for the Minimum Steiner Tree Problem
Masayuki Miyamoto, Masakazu Iwamura, Koichi Kise, François Le Gall |
COCOON | 3 |
| 2020 | Suitable Camera and Rotation Navigation for People with Visual Impairment on Looking for Something Using Object Detection TechniqueabstractAbstract For people with visual impairment, smartphone apps that use computer vision techniques to provide visual information have played important roles in supporting their daily lives. However, they can be used under a specific condition only. That is, only when the user knows where the object of interest is . In this paper, we first point out the fact mentioned above by categorizing the tasks that obtain visual information using computer vision techniques. Then, in looking for something as a representative task in a category, we argue suitable camera systems and rotation navigation methods. In the latter, we propose novel voice navigation methods. As a result of a user study comprised of seven people with visual impairment, we found that (1) a camera with a wide field of view such as an omnidirectional camera was preferred, and (2) users have different preferences in navigation methods. Masakazu Iwamura, Yoshihiko Inoue, Kazunori Minatani, Koichi Kise |
ICCHP (1) | 4 |
| 2020 | Distortion-Adaptive Grape Bunch Counting for Omnidirectional ImagesabstractThis paper proposes the first object counting method for omnidirectional images. Because conventional object counting methods cannot handle the distortion of omnidirectional images, we propose to process them using stereographic projection, which enables conventional methods to obtain a good approximation of the density function. However, the images obtained by stereographic projection are still distorted. Hence, to manage this distortion, we propose two methods. One is a new data augmentation method designed for the stereographic projection of omnidirectional images. The other is a distortion-adaptive Gaussian kernel that generates a density map ground truth while taking into account the distortion of stereographic projection. Using the counting of grape bunches as a case study, we constructed an original grape-bunch image dataset consisting of omnidirectional images and conducted experiments to evaluate the proposed method. The results show that the proposed method performs better than a direct application of the conventional method, improving mean absolute error by 14.7% and mean squared error by 10.5%. Ryota Akai, Yuzuko Utsumi, Yuka Miwa, Masakazu Iwamura, Koichi Kise |
ICPR | 5 |
| 2019 | Towards Quality Assessment of Crowdworker Output Based on Behavioral DataabstractIn this paper, we show preliminary results on the quality assessment of crowdworker output based on the movements of the mouse and the eyes while the task is performed. We assume that the mouse and the eyes stop longer if the quality is lower due to the lack of knowledge, or confidence, etc. Because the mouse- and eye-stopping duration follows lognormal distribution, we estimate its parameters (mean and standard deviation) to evaluate the quality. Results of preliminary experiments with 10 participants show that the parameters of correct outputs are different from those of incorrect ones. As compared to the task duration, which is often used as a feature for assessment, we have found that the mouse-and the eyestopping duration is advantageous and complementary for the assessment. Shigeaki Yuasa, Takumi Nakai, Takanori Maruichi, Manuel Landsmann, Koichi Kise, Masaki Matsubara, Atsuyuki Morishima |
IEEE BigData | 5 |
| 2019 | Private Reader: Using Eye Tracking to Improve Reading Privacy in Public SpacesabstractReading in public spaces can often be tricky if we wish to keep the contents away from the prying eye. We propose Private Reader, an eye-tracking approach towards maintaining privacy while reading by rendering only the portion of text that is gazed by the reader. We conducted a user study by evaluating for both the reader and observer in terms of privacy, reading comfort, and reading speed for three reading modes; normal, underscored, and scrambled text. "Scrambled" performs best in terms of perceived effort and frustration for the shoulder surfer. Our contribution is threefold; we developed a system to preserve privacy by rendering only the text at gaze-point of the reader, we conducted a user study to evaluate user preferences and subjective task load, and we suggested several scenarios where Private Reader is useful in public spaces. Kirill Ragozin, Yun Suen Pai, Olivier Augereau, Koichi Kise, Jochen Kerdels, Kai Kunze |
MobileHCI | 4 |
| 2018 | Vocabulometer: A Web Platform for Document and Reader Mutual AnalysisabstractWe present the Vocabulometer, a reading assistant system designed to record the reading activity of a user with an eye tracker and to extract mutual information about the users and the read documents. The Vocabulometer stands as a web platform and can be used for analyzing the comprehension of the user, the comprehensibility of the document, predicting the difficult words, recommending document according to the reader's in order to increase his skills, etc. Since the last years, with the development of low-cost eye trackers, the technology is now accessible for many people, which will allow using data mining and machine learning algorithms for the mutual analysis of documents and readers. Olivier Augereau, Clément Jacquet, Koichi Kise, Nicholas Journet |
DAS | 3 |
| 2018 | Comics Story Representation System Based on GenreabstractComics is usually classified into broad categories called "genres" according to its contents such as comedy, horror, science fiction, etc. Because a genre expresses a comics story briefly, people read comics which has contents based on their interest, by relying on comics genres. However, giving only one genre to one comic cannot express the detailed difference of the story. In this paper, we propose a system for generating comics story representation as a sub-sequence of genres. Our comics story representation can be applied to a new search engine based on stories or to a recommendation system which analyzes the tastes of the user's favorite comics by finding comics with similar story representation. We use a deep neural network to classify each page into the corresponding genre. Experimental results confirm the advantage of the proposed system. Yuki Daiku, Motoi Iwata, Olivier Augereau, Koichi Kise |
DAS | 4 |
| 2017 | Identification of Reader Specific Difficult Words by Analyzing Eye Gaze and Document ContentabstractThis paper presents an approach for identifying reader specific difficult words while someone is reading a textual document. The work is motivated by the need of developing human-document interaction systems, in general and creating person-specific online educational content, in particular. Eye gaze information gives person specific behavior whereas textual content is analyzed to get general linguistic aspect of the document content. These two pieces of information are fused together through machine learning algorithms to identify the set of difficult words for a particular reader reading a particular document. An annotated dataset has been created where each word in a document is marked with its bounding box information and each reader identifies a set of difficult words while reading the document. The dataset consists of sixteen documents and each document is read by five subjects. The method is evaluated through recall-precision analysis. The impressive precision at high recall attests the feasibility of building a practical application based on this research. The experiment further brings out several interesting facts about human reading behaviour. Utpal Garain, Onkar Pandit, Olivier Augereau, Ayano Okoso, Koichi Kise |
ICDAR | 5 |
| 2016 | Camera-Based System for User Friendly Annotation of DocumentsabstractWe propose a system for document annotation using a camera mounted on a smartphone. It works both on paper documents and electronic documents thanks to the functionality of document image retrieval. An important characteristic of this system is in the way of annotating documents. The system employs simple character stickers which represent user's opinions ("hard to understand", "interesting", "boring", "surprising", "doubtful") for friendly annotation on documents. We evaluated our system by changing the way of annotation and found that users most liked the proposed way of annotation with stickers though it sometimes caused a confusion about the interpretation of stickers. We discuss the possible way of solving this issue as a result of the analysis of experimental results. Yusuke Oguma, Koichi Kise |
DAS | 2 |
| 2016 | Semi-automatic Text and Graphics Extraction of Manga Using Eye Tracking InformationabstractThe popularity of storing, distributing and reading comic books electronically has made the task of comics analysis an interesting research problem. Different work have been carried out aiming at understanding their layout structure and the graphic content. However the results are still far from universally applicable, largely due to the huge variety in expression styles and page arrangement, especially in manga (Japanese comics). In this paper, we propose a comic image analysis approach using eye-tracking data recorded during manga reading sessions. As humans are extremely capable of interpreting the structured drawing content, and show different reading behaviors based on the nature of the content, their eye movements follow distinguishable patterns over text or graphic regions. Therefore, eye gaze data can add rich information to the understanding of the manga content. Experimental results show that the fixations and saccades indeed form consistent patterns among readers, and can be used for manga textual and graphical analysis. Christophe Rigaud, Nam Le Thanh 0001, Jean-Christophe Burie, Jean-Marc Ogier, Shoya Ishimaru, Motoi Iwata, Koichi Kise |
DAS | 7 |
| 2016 | Automatic character labeling for camera captured document imagesabstractCharacter groundtruth for camera captured documents is crucial for training and evaluating advanced OCR algorithms. Manually generating character level groundtruth is a time consuming and costly process. This paper proposes a robust groundtruth generation method based on document retrieval and image registration for camera captured documents. We use an elastic non-rigid alignment method to fit the captured document image which relaxes the flat paper assumption made by conventional solutions. The proposed method allows building very large scale labeled camera captured documents dataset, without any human intervention. We construct a large labeled dataset consisting of 1 million camera captured Chinese character images. Evaluation of samples generated by our approach showed that 99.99% of the images were correctly labeled, even with different distortions specific to cameras such as blur, specularity and perspective distortion. Koichi Kise, Masakazu Iwamura |
ICIP | 2 |
| 2016 | Towards an automated estimation of English skill via TOEIC score based on reading analysisabstractEstimating automatically the degree of language skill by analyzing the eye movements is a promising way to help people from all over the world to learn a new language. In this study, we focus on the English skills of non-native speakers. Our aim is to provide an algorithm that can assess accurately and automatically the TOEIC score after reading English texts for few minutes. As a first step towards this direction, we propose an algorithm that can predict accurately this score after reading and answering some questions about the comprehension of few English texts. We use an eye tracker in order to record the eye gaze, i.e. the positions where the reader is looking at. Then we extract several features to characterize the behavior, and consequently the skill of the reader. We also add a feature based on the number of correct answers to the questions. By using a machine learning based on multivariate regression, the score is estimated user independently. A backward stepwise feature selection is used to select the relevant features and to optimize the estimation. As a main result, the TOEIC score is estimated with 21.7 points of mean absolute error for 21 subjects after reading and answering the questions of only 3 documents. Olivier Augereau, Hiroki Fujiyoshi, Koichi Kise |
ICPR | 3 |
| 2015 | Quantifying reading habits: counting how many words you readabstractReading is a very common learning activity, a lot of people perform it everyday even while standing in the subway or waiting in the doctors office. However, we know little about our everyday reading habits, quantifying them enables us to get more insights about better language skills, more effective learning and ultimately critical thinking. This paper presents a first contribution towards establishing a reading log, tracking how much reading you are doing at what time. We present an approach capable of estimating the words read by a user, evaluate it in an user independent approach over 3 experiments with 24 users over 5 different devices (e-ink reader, smartphone, tablet, paper, computer screen). We achieve an error rate as low as 5% (using a medical electrooculography system) or 15% (based on eye movements captured by optical eye tracking) over a total of 30 hours of recording. Our method works for both an optical eye tracking and an Electrooculography system. We provide first indications that the method works also on soon commercially available smart glasses. Kai Kunze, Katsutoshi Masai, Masahiko Inami, Ömer Sacakli, Marcus Liwicki, Andreas Dengel 0001, Shoya Ishimaru, Koichi Kise |
UbiComp | 8 |
| 2015 | A proposal of a document image reading-life log based on document image retrieval and eyetrackingabstractInstead of analyzing directly the document images, analyzing the document reading can offer new perspectives for extracting information about both the reader and the document. Analyzing how people read texts can help to understand the cognitive process of the reading and might lead to new approaches and new solutions for pattern recognition and document image analysis. It can also lead to create smart documents that can measure reading information, provide feedback and adapt themselves depending on the behavior of the readers. As a step towards document reading analysis, the authors propose in this paper a solution for extracting the reading information and creating a “reading-life log”. This reading-life log contains basic features that can be used for many different kinds of applications. A tag cloud evolving according to the reading is presented as a first application of the reading-life log. Olivier Augereau, Koichi Kise, Kensuke Hoshika |
ICDAR | 2 |
| 2015 | Speech balloon and speaker association for comics and manga understandingabstractComics and manga are one of the most important forms of publication and play a major role in spreading culture all over the world. In this paper we focus on balloons and their association to comic characters or more generally text and graphic links retrieval. This information is not directly encoded in the image, whether scanned or digital-born, it has to be understood according to other information present in the image. Such high level information allows new browsing experience and story understanding (e.g. dialog analysis, situation retrieval). We propose a speech balloon and comic character association method able to retrieve which character is emitting which speech balloon. The proposed method is based on geometric graph analysis and anchor point selection. This work has been evaluated over various comic book styles from the eBDtheque dataset and also a volume of the Kingdom manga series. Christophe Rigaud, Nam Le Thanh 0001, Jean-Christophe Burie, Jean-Marc Ogier, Motoi Iwata, Eiki Imazu, Koichi Kise |
ICDAR | 7 |
| 2015 | The eye as the window of the language ability: Estimation of English skills by analyzing eye movement while reading documentsabstractReading-life log is a research field of analyzing our activities of reading documents to know more about readers and documents. In this paper we propose an implementation of reading-life log which is to estimate the English language skill by analyzing the activities of reading English documents. As input for the analysis, we employ eye movement information, because we consider the eye movement of skillful readers is far different from that of novices. From the experiments, we have found that the following two features are informative: (1) the sum of fixation duration, and (2) the sum of the velocity of saccades. By using these features the proposed method is to estimate the class of English skill from among low, middle and high, which are defined based on the scores of English standardized test called TOEIC. From the experimental results with 11 subjects and 10 documents, we have been successful to estimate the class with the accuracy of 90.9%. Kazuyo Yoshimura, Koichi Kise, Kai Kunze |
ICDAR | 2 |
| 2015 | DCT-OFDM Based Watermarking Scheme Robust Against Clipping, Rotation, and Scaling Attacks
Hiroaki Ogawa, Minoru Kuribayashi, Motoi Iwata, Koichi Kise |
IWDW | 4 |
| 2014 | Fast and Optimal Binary Template Matching Application to Manga Copyright ProtectionabstractTemplate matching is a technique used in classifying an object by comparing portions of images with another image. Finding a given template in an image is typically performed by scanning the image and evaluating the similarity with the template. When the scanning is concerned with the entire image template matching is optimal. This paper considers a special case of template matching where the templates are binary. Although binary template matching has been studied extensively since the early days of pattern recognition, this technique seems not longer in use in Document Image Analysis (DIA). The major reasons arête time complexity, the no-invariance to scale and rotation and the lack of adaptability of similarity measures. However, different contributions have been investigated during the last years to improve these aspects: robustness and discrimination capability of similarity measures, their characterization, time-processing optimization with hardware support, etc. In this paper, we will review first some of the recent issues about binary template matching. We will present then a system exploiting bitwise operators and parallel processing supporting fast and accurate binary template matching for Manga copyright protection. This system is compared to a FFT-based template matching, and it outperforms both in processing-time and detection accuracy. Mathieu Delalandre, Motoi Iwata, Koichi Kise |
Document Analysis Systems | 3 |
| 2014 | A Study to Achieve Manga Character Retrieval Method for Manga ImagesabstractManga (Japanese style comics) is one of the most popular publications. Nowadays manga is often handled as digital images not only in consumers' use but also in digital media. However, they hardly handle manga as content-based materials. Some digital media use tags or text data for retrieval, where the tags and text data are produced by handmade input. Therefore our goal is achieving content-based retrieval method for manga images. As the first step to the goal, we investigate the performance of Sun's method applying to manga character retrieval. Manga character retrieval means a image retrieval of which the input and output are a character image and page images where the input character appears respectively. It is useful for convenient use of manga images, for example, character retrieval or auto-tagging. We modify Sun's method so as to be applicable to manga character retrieval and then investigate the performance. Motoi Iwata, Atsushi Ito, Koichi Kise |
Document Analysis Systems | 3 |
| 2014 | Performance Improvement in Local Feature Based Camera-Captured Character RecognitionabstractConcerning camera-captured Japanese character recognition, we have proposed a method to recognize characters, both simple and complex, that may not be linearly aligned and may be printed with a complex background. Recognition is performed based on local features and their arrangement. The arrangement is validated with an algorithm called local RANSAC. However, at least four corresponding local features are required. To relax that condition, we propose a new recognition method making it possible to recognize a character region with at least three corresponding local features. This method enables recall and precision to be improved with the simpler characters using more corresponding local features and computation times to be reduced by 7%. Takahiro Matsuda 0001, Masakazu Iwamura, Koichi Kise |
Document Analysis Systems | 3 |
| 2014 | A mixed reality head-mounted text translation system using eye gaze inputabstractEfficient text recognition has recently been a challenge for augmented reality systems. In this paper, we propose a system with the ability to provide translations to the user in real-time. We use eye gaze for more intuitive and efficient input for ubiquitous text reading and translation in head mounted displays (HMDs). The eyes can be used to indicate regions of interest in text documents and activate optical-character-recognition (OCR) and translation functions. Visual feedback and navigation help in the interaction process, and text snippets with translations from Japanese to English text snippets, are presented in a see-through HMD. We focus on travelers who go to Japan and need to read signs and propose two different gaze gestures for activating the OCR text reading and translation function. We evaluate which type of gesture suits our OCR scenario best. We also show that our gaze-based OCR method on the extracted gaze regions provide faster access times to information than traditional OCR approaches. Other benefits include that visual feedback of the extracted text region can be given in real-time, the Japanese to English translation can be presented in real-time, and the augmentation of the synchronized and calibrated HMD in this mixed reality application are presented at exact locations in the augmented user view to allow for dynamic text translation management in head-up display systems. Takumi Toyama, Daniel Sonntag, Andreas Dengel 0001, Takahiro Matsuda 0001, Masakazu Iwamura, Koichi Kise |
IUI | 6 |
| 2014 | Fast Instance Search Based on Approximate Bichromatic Reverse Nearest Neighbor SearchabstractIn the TRECVID Instance Search (INS) task, it is known that use of BM25, which is an improvement of the TFIDF,greatly improves retrieval performance. Its calculation, however, requires tremendous amount of computational cost and this fact makes its use intractable. In this paper, we present its efficient computational method. Since the BM25 is obtained by solving the bichromatic reverse nearest neighbor (BRNN)search problem,we propose an approximate method for the problem based on the state-of-the-art approximate nearest neighbor search method, bucket distance hashing (BDH). An experiment using the TRECVID INS 2012 dataset showed that the proposed method reduced computational cost to less than 1/3500 of the brute-force search with keeping the accuracy. Masakazu Iwamura, Nobuaki Matozaki, Koichi Kise |
ACM Multimedia | 3 |
| 2014 | Recovery and localization of handwritings by a camera-pen based on tracking and document image retrieval
Megumi Chikano, Koichi Kise, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi |
Pattern Recognit. Lett. | 2 |
| 2014 | More than ink - Realization of a data-embedding pen
Marcus Liwicki, Seiichi Uchida, Akira Yoshida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
Pattern Recognit. Lett. | 6 |
| 2013 | What is the Most EfficientWay to Select Nearest Neighbor Candidates for Fast Approximate Nearest Neighbor Search?abstractApproximate nearest neighbor search (ANNS) is a basic and important technique used in many tasks such as object recognition. It involves two processes: selecting nearest neighbor candidates and performing a brute-force search of these candidates. Only the former though has scope for improvement. In most existing methods, it approximates the space by quantization. It then calculates all the distances between the query and all the quantized values (e.g., clusters or bit sequences), and selects a fixed number of candidates close to the query. The performance of the method is evaluated based on accuracy as a function of the number of candidates. This evaluation seems rational but poses a serious problem; it ignores the computational cost of the process of selection. In this paper, we propose a new ANNS method that takes into account costs in the selection process. Whereas existing methods employ computationally expensive techniques such as comparative sort and heap, the proposed method does not. This realizes a significantly more efficient search. We have succeeded in reducing computation times by one-third compared with the state-of-theart on an experiment using 100 million SIFT features. Masakazu Iwamura, Tomokazu Sato, Koichi Kise |
ICCV | 3 |
| 2013 | Automatic Ground Truth Generation of Camera Captured Documents Using Document Image RetrievalabstractIn this paper a novel method for automatic ground truth generation of camera captured document images is proposed. Currently, no dataset is available for camera captured documents. It is very difficult to build these datasets manually, as it is very laborious and costly. The proposed method is fully automatic, allowing building the very large scale (i.e., millions of images) labeled camera captured documents dataset, without any human intervention. Evaluation of samples generated by the proposed approach shows that 99.98% of the images are correctly labeled. Novelty of the proposed approach lies in the use of document image retrieval for automatic labeling, especially for camera captured documents, which contain different distortions specific to camera, e.g., blur, occlusion, perspective distortion, etc. Sheraz Ahmed, Koichi Kise, Masakazu Iwamura, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 2 |
| 2013 | Key-Region Detection for Document Images - Application to Administrative Document RetrievalabstractIn this paper we argue that a key-region detector designed to take into account the special characteristics of document images can result in the detection of less and more meaningful key-regions. We propose a fast key-region detector able to capture aspects of the structural information of the document, and demonstrate its efficiency by comparing against standard detectors in an administrative document retrieval scenario. We show that using the proposed detector results to a smaller number of detected key-regions and higher performance without any drop in speed compared to standard state of the art detectors. Hongxing Gao, Marçal Rusiñol, Dimosthenis Karatzas, Josep Lladós 0001, Tomokazu Sato, Masakazu Iwamura, Koichi Kise |
ICDAR | 7 |
| 2013 | Automatic Labeling for Scene Text DatabaseabstractIt is thought that a large quantity of data improve quality of recognition. A large database, however, is not easy to obtain. The hardest task is labeling (also known as ground truthing), which usually requires human intervention. Since labeling by human is laborious and costly, labeling without human (automatic labeling) or minimization of human intervention (semi-automatic labeling) are ideal scenarios. As a step toward realization of the scenarios, knowing how much an automatic labeling system can perform without human intervention is important. In the current paper we propose a comprehensive automatic labeling technique for a scene text database, which performs segmentation and labeling for unsegmented and unlabeled character images. To our best knowledge, this is the first method to realize the comprehensive process for automatic labeling for scene text databases In experiments, we confirmed that the proposed method could add new unlabeled data in parallel with improving recognition performance of the classifier. Masakazu Iwamura, Masaki Tsukada, Koichi Kise |
ICDAR | 3 |
| 2013 | The Reading-Life Log - Technologies to Recognize Texts That We ReadabstractReading life log is a type of techniques to automatically and unconsciously record people's reading intentions, interests and habits. Besides, it can also serve as various assistants in our daily life. In this paper, a reading-life log system is implemented by a head-mounted and unobtrusive video camera with a high resolution and a high shutter speed. We utilize DP matching, and propose a text-based frame mosaicing method to integrate multiple frames in a clip. The developed system is tested in the various environments indoor and outdoor. The experimental results show that our system can provide reliable outputs with respect to the most correct responses. The infrequent misregistration between lines also indicates the feasibility and validity of the text-based frame mosaicing. Takashi Kimura, Rong Huang 0003, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICDAR | 6 |
| 2013 | An Anytime Algorithm for Camera-Based Character RecognitionabstractIn a scene image, some characters are difficult to recognize and some others are recognized easily. Such difficult characters usually make the processing time long while easy characters are recognized in a short time. In this paper, we propose a system which recognizes each character with a proper cost for the difficulty. Through the process, easy characters are recognized early and difficult ones are recognized late. This is a desired property of an anytime algorithm that the recognition accuracy does not decrease as the time increases. In order to realize it, we propose a method which splits the recognition process into several times and accumulates the recognition results and extracted features. We also discuss what is required to realize the anytime algorithm for the scene character recognition task. Experiments reveal that the proposed method obtains recognition results of easy characters earlier than the conventional method. Takuya Kobayashi, Masakazu Iwamura, Takahiro Matsuda 0001, Koichi Kise |
ICDAR | 4 |
| 2013 | The Wordometer - Estimating the Number of Words Read Using Document Image Retrieval and Mobile Eye TrackingabstractWe introduce the Wordometer, a novel method to estimate the number of words a user reads using a mobile eye tracker and document image retrieval. We present a reading detection algorithm which works with over 91 % accuracy over 10 test subjects using 10-fold cross validation. We implement two algorithms to estimate the read words using a line break detector. A simple version gives an average error rate of 13,5 % for 9 users over 10 documents. A more sophisticated word count algorithm based on support vector regression with an RBF kernel reaches an average error rate from only 8.2 % (6.5 % if one test subject with abnormal behavior is excluded). The achieved error rates are comparable to pedometers that count our steps in our daily life. Thus, we believe the Wordometer can be used as a step counter for the information we read to make our knowledge life healthier. Kai Kunze, Hitoshi Kawaichi, Kazuyo Yoshimura, Koichi Kise |
ICDAR | 4 |
| 2013 | Reading Activity Recognition Using an Off-the-Shelf EEG - Detecting Reading Activities and Distinguishing Genres of DocumentsabstractThe document analysis community spends substantial resources towards computer recognition of any type of text (e.g. characters, handwriting, document structure etc.). In this paper, we introduce a new paradigm focusing on recognizing the activities and habits of users while they are reading. We describe the differences to the traditional approaches of document analysis. We present initial work towards recognizing reading activities. We report our initial findings using a commercial, dry electrode Electroencephalography (EEG) system. We show the feasibility to distinguish reading tasks for 3 different document genres with one user and near perfect accuracy. Distinguishing reading tasks for 3 different document types we achieve 97 % with user specific training. We present evidence that reading and non-reading related activities can be separated over 3 users using 6 classes, perfectly separating reading from non-reading. A simple EEG system seems promising for distinguishing the reading of different document genres. Kai Kunze, Yuki Shiga, Shoya Ishimaru, Koichi Kise |
ICDAR | 4 |
| 2013 | Specific Comic Character Detection Using Local Feature MatchingabstractComic books are a kind of storytelling graphic publications mainly expressed by abstract line drawings. As a clue of story lines, comic characters play an important role in the story, and their detection is an essential part of comic book analysis. For this purpose, the task includes (1) locating characters in comics pages and (2) identifying them, which is called specific character detection. Corresponding to different scenes of comic books, one specific character can be represented by various expressions coupled with rotations, occlusions, and other perspective drawing effects, which challenge the detection. In this paper, we focus on stable features regarding the possible transformations and proposed a framework to detect them. Specifically, some discriminative features are selected as detectors for characterizing characters, on the basis of a training dataset. Based on the detectors, the drawings of the same characters in different scenes can be detected. The methodology has been experimented and validated on 6 titles of comics. Despite the terrific changes for different scenes, the proposed method achieved detection of 70% comic characters. Weihan Sun, Jean-Christophe Burie, Jean-Marc Ogier, Koichi Kise |
ICDAR | 4 |
| 2013 | Wearable Reading Assist System: Augmented Reality Document Combining Document Retrieval and Eye TrackingabstractWe present a new system that assists people's reading activity by combining a wearable eye tracker, a see-through head mounted display, and an image based document retrieval engine. An image based document retrieval engine is used for identification of the reading document, whereas an eye tracker is used to detect which part of the document the reader is currently reading. The reader can refer to the glossary of the latest viewed key word by looking at the see-through head mounted display. This novel document reading assist application, which is the integration of a document retrieval system into an everyday reading scenario for the first time, enriches people's reading life. In this paper, we i) investigate the performance of the state-of-the-art image based document retrieval method using a wearable camera, ii) propose a method for identification of the word the reader is attendant, and iii) conduct pilot studies for evaluation of the system in this reading context. The results show the potential of a document retrieval system in combination with a gaze based user-oriented system. Takumi Toyama, Andreas Dengel 0001, Wakana Suzuki, Koichi Kise |
ICDAR | 4 |
| 2013 | User attention oriented augmented reality on documents with document dependent dynamic overlayabstractWhen we read a document (any kind of, scientific papers, novels, etc.), we often encounter a situation that the information from the reading document is too less to comprehend what the author(s) would like to convey. In this paper, we demonstrate how the combination of a wearable eye tracker, a see-through head-mounted display (HMD) and an image based document retrieval engine enhances people's reading experiences. By using our proposed system, the reader can get supportive information in the see-through HMD when he wants. A wearable eye tracker and a document retrieval engine are used to detect which line in the document the reader is reading. We propose a method to detect the reader's attention on a word in a reading document, in order to present information at a preferable moment. Furthermore, we also propose a method to project a point of the document to a point of the HMD screen, by calculating the pose of the reading document in the camera image. This projection enables the system to overlay the information dynamically in an augmented view on the reading line. The results from the user study and the experiments show the potential of the proposed system in a practical use case. Takumi Toyama, Wakana Suzuki, Andreas Dengel 0001, Koichi Kise |
ISMAR | 4 |
| 2013 | Detection of exact and similar partial copies for copyright protection of manga
Weihan Sun, Koichi Kise |
Int. J. Document Anal. Recognit. | 2 |
| 2012 | Recognizing Words in Scenes with a Head-Mounted Eye-TrackerabstractRecognition of scene text using a hand-held camera is emerging as a hot topic of research. In this paper, we investigate the use of a head-mounted eye-tracker for scene text recognition. An eye-tracker detects the position of the user's gaze. Using gaze information of the user, we can provide the user with more information about his region/object of interest in a ubiquitous manner. Therefore, we can realize a service such as the user gazes at a certain word and soon obtain the related information of the word by combining a word recognition system with eye-tracking technology. Such a service is useful since the user has to do nothing but gazes at interested words. With a view to realize the service, we experimentally evaluate the effectiveness of using the eye-tracker for word recognition. The initial results show the recognition accuracy was around 70% in our word recognition experiment and the average computational time was less than one second per a query image. Takuya Kobayashi, Takumi Toyama, Faisal Shafait, Masakazu Iwamura, Koichi Kise, Andreas Dengel 0001 |
Document Analysis Systems | 5 |
| 2012 | Similar Fragment Retrieval of Animations by a Bag-of-Features ApproachabstractSerial animated cartoons, also called animations, is a kind of popular video documents that describe narratives by videos of cartoons. By using digital techniques, illegal users can divide animations into fragments and distribute them without any copyright permissions. In this paper we focus on the copyright problem of animations and try to retrieve copyright infringement fragments based on key frames. Because of the huge volumes and rapid release of animations, it is impossible to store all the episodes into the database. We propose applying bag-of-features model to retrieve similar fragments based on the visual words extracted from a limited data set. In the experiments, 12 titles of animations are employed and the similar fragments outside database are applied as queries. Our method has achieved above 98% precision at 80% recall for fragments whose durations are over 140 seconds. From the results, we show that a latest release can be retrieved based on the features of the former ones from the same title. Weihan Sun, Koichi Kise, Yoann Champeil |
Document Analysis Systems | 2 |
| 2012 | Real-Time Document Image Retrieval on a SmartphoneabstractThis paper presents a novel interface running on smart phones which is capable of seamlessly linking physical and digital worlds through paper documents. This interface is based on a real-time document image retrieval method called Locally Likely Arrangement Hashing. By just only pointing a smart phone to a paper document, the user can obtain its corresponding electronic document. This can easily provide the user with the information associated with the retrieved document. This relevant information can be superimposed on the display of smart phones. Therefore, we consider that with the help of this interface, the user can utilize paper documents as a new medium to display various information. Kazutaka Takeda, Koichi Kise, Masakazu Iwamura |
Document Analysis Systems | 2 |
| 2012 | Expanding Recognizable Distorted Characters Using Self-Corrective RecognitionabstractLarge datasets are always demanded for better recognition performance. However, it is not easy to produce them because costly and slow human operators have been necessary for labeling. In the current paper, in order to resolve the problem on yielding large datasets, we propose a scenario for automatic labeling based on the self-corrective recognition algorithm. The strong point of the proposed method is the capability of expanding recognizable distorted characters unlike existing methods. In the experiments, we show a possibility to realize automatic labeling by the method. Masaki Tsukada, Masakazu Iwamura, Koichi Kise |
Document Analysis Systems | 3 |
| 2011 | Recognition of Multiple Characters in a Scene Image Using Arrangement of Local FeaturesabstractRecognizing characters in a scene helps us obtain useful information. For the purpose, character recognition methods are required to recognize characters of various sizes, various rotation angles and complex layout on complex background. In this paper, we propose a character recognition method using local features having several desirable properties. The novelty of the proposed method is to take into account arrangement of local features so as to recognize multiple characters in an image unlike past methods. The effectiveness and possible improvement of the method are discussed. Masakazu Iwamura, Takuya Kobayashi, Koichi Kise |
ICDAR | 3 |
| 2011 | Reliable Online Stroke Recovery from Offline Data with the Data-Embedding PenabstractIn this paper we propose a complete system for online stroke recovery from offline data. The key idea of our approach is to use a novel pen device which is able to embed meta information into the ink during writing the strokes. This pen-device overcomes the need to get access to any memory on the pen when trying to recover the information, which is especially useful in multi-writer or multi-pen scenarios. The actual data-embedding is achieved by an additional ink dot sequence along a handwritten pattern during writing. We design the ink-dot sequence in such a way that it is possible to retrieve the writing direction from a scanned image. Furthermore, we propose novel processing steps in order to retrieve the original writing direction and finally the embedded data. In our experiments we show that we can reliably recover the writing direction of various patterns. Our system is able to determine the writing direction of straight lines, simple patterns with crossings (e.g., "x" and "II"), and even more complex patterns like handwritten words and symbols. Marcus Liwicki, Akira Yoshida, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICDAR | 6 |
| 2011 | Similar Manga Retrieval Using Visual Vocabulary Based on Regions of InterestabstractManga, a kind of graphic novels expressed by sequential line drawings, is an important document image publication which is invoking more and more attentions for their copyright protection. Because of simple constructions and abstract expression styles, similar copies are always applied in plagiarisms of mangas. In addition, the enormous volume of copyrighted manga publications issues a challenge to the task of detecting suspicious images. Considering only some regions of interests (ROIs) require copyright protection, we propose a bag-of-features method using visual words based on ROIs for similar manga retrieval. In the experiments, we applied real manga publications as data and proved the effectiveness of the proposed method. Weihan Sun, Koichi Kise |
ICDAR | 2 |
| 2011 | Real-Time Document Image Retrieval for a 10 Million Pages Database with a Memory Efficient and Stability Improved LLAHabstractThis paper presents a real-time document image retrieval method for a large-scale database with Locally Likely Arrangement Hashing (LLAH). In general, when a database is scaled up, a large amount of memory is required and retrieval accuracy drops due to insufficient discrimination power of features. To solve these problems, we propose three improvements: memory reduction by sampling feature points, improvement of discrimination power by increasing the number of feature dimensions and stabilizing features by reducing redundancy. From the experimental results, we have confirmed that the proposed method realizes 50% memory reduction, and achieves 99.4% accuracy and 38ms processing time for a database of 10 million pages. Kazutaka Takeda, Koichi Kise, Masakazu Iwamura |
ICDAR | 2 |
| 2011 | Generic and Specific Object Recognition for Semantic Retrieval of Images
Martin Klinkigt, Koichi Kise, Andreas Dengel 0001 |
KES (1) | 2 |
| 2011 | Semantic Retrieval of Images by Learning from Wikipedia
Martin Klinkigt, Koichi Kise, Heiko Maus, Andreas Dengel 0001 |
KES (4) | 2 |
| 2011 | Handwriting on Paper as a Cybermedium
Akira Yoshida, Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
KES (4) | 6 |
| 2010 | From Local Features to Global Shape Constraints: Heterogeneous Matching Scheme for Recognizing Objects under Serious Background Clutter
Martin Klinkigt, Koichi Kise |
ACCV (4) | 2 |
| 2010 | Memory-based recognition of camera-captured charactersabstractThis paper addresses how to quickly recognize a character pattern using a lot of case examples without learning. Here without learning means just finding the most similar example from the case examples, and pretend as if the OCR understands the definition of the character. This strategy is expected to work well in most cases with a large dataset, however, also expected to take a lot of time for finding the most similar example. In this paper, we show that a lot of case examples can be processed in a short time. As a testbed, we handle recognition problem of camera-captured printed characters. Using a database storing 100 fonts, the proposed method achieved 97.0% of recognition rate for images captured from the right angle and 95.8% for those from 45 deg. with 4.56ms of processing time, that is about 220 characters per second including every process. Masakazu Iwamura, Tomohiko Tsuji, Koichi Kise |
Document Analysis Systems | 3 |
| 2010 | Expansion of queries and databases for improving the retrieval accuracy of document portions: an application to a camera-pen systemabstractThis paper presents a method of improving the accuracy of document image retrieval focusing on the application to a camera-pen system. In a camera-pen system, document image retrieval is employed for locating the pen-tip position on a page. A serious problem is that since the camera is mounted close to the pen-tip, the camera captures only a tiny portion of the page and the resultant image is under severe perspective distortion, resulting in lowering the retrieval accuracy. To solve this problem, we propose new geometrically invariant features as well as expansion techniques which increase the number of index features of either the database or the query images. From the experimental results, it has been found that the query expansion technique with features by combining affine and perspective invariants allows us the best performance that improves the accuracy of a baseline method more than 27%. Koichi Kise, Megumi Chikano, Kazumasa Iwata, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi |
Document Analysis Systems | 1 |
| 2010 | Data-embedding pen: augmenting ink strokes with meta-informationabstractIn this paper we present the first operational version of the data-embedding pen. During writing a pattern, this pen produces an additional ink-dot sequence along the ink stroke of the pattern. The ink-dot sequence represents, for example, meta-information (such as the writer's name and the date of writing) and thus drastically increases the value of the handwriting on a physical paper. Since the information is placed on the paper, it can be extracted just by scanning or photographing the paper. There is no need to get access to any memory on the pen to recover the information. This is useful especially in multi-writer or multi-pen scenarios. The experiments using an encoding scheme and a decoding algorithm showed very promising results. For example, it was proved that we can embed 28 or more bits of information on simple handwritten patterns and decode them with a high reliability. Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
Document Analysis Systems | 5 |
| 2010 | Tracking and Retrieval of Pen Tip Positions for an Intelligent Camera PenabstractThis paper presents a method of recovering digital ink for an intelligent camera pen, which is characterized by the functions that (1) it works on ordinary paper and (2) if an electronic document is printed on the paper the recovered digital ink is associated with the document. Two technologies called paper fingerprint and document image retrieval are integrated for realizing the above functions. The key of the integration is the introduction of image mosaicing and fast retrieval of previously seen fingerprints based on hashing of SURF local features. From the experimental results of 50 handwritings, we have confirmed that the proposed method is effective to recover and locate the digital ink from the handwriting on a physical paper. Kazumasa Iwata, Koichi Kise, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi |
ICFHR | 2 |
| 2010 | Embedding Meta-Information in Handwriting -- Reed-Solomon for Reliable Error CorrectionabstractIn this paper a more compact and more reliable coding scheme for the data-embedding pen is proposed. The data-embedding pen produces an additional ink-dot sequence along a handwritten pattern during writing. The ink-dot sequence represents, for example, meta-information (such as the writer's name and the date of writing) and thus drastically increases the value of the handwriting on a physical paper. There is no need to get access to any memory on the pen to recover the information, which is especially useful in multi-writer or multi-pen scenarios. In this paper we focus on the compactness of the encoded information. The aim of this paper is to encode as much information as possible in short stroke sequences. In our experiments we show that we can embed more information in shorter strokes than in previous work. In straight lines as short as 5 cm, 32 bits can successfully be embedded. Furthermore, the new encoding scheme also works reliably on more complex patterns. Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICFHR | 5 |
| 2010 | Handwriting Reconstruction for a Camera Pen Using Random Dot PatternsabstractThis paper proposes a new method of handwriting reconstruction using a camera pen. We print random dot patterns on the document background to enable retrieval of both the current document and the pen position on this document. Dot arrangements are stored in a hash table using Locally Likely Arrangement Hashing. For retrieval, they are extracted from the camera image and matched to the corresponding points in the hash table. We were able to achieve high retrieval accuracy (81.1~100.0%), given a sufficient amount of visible dots. Using a two-step homography approximation, an accurate image of handwriting can be reconstructed. By using knowledge about document context and a client-server architecture, our method allows real-time processing on ordinary hardware. Matthias Sperber, Martin Klinkigt, Koichi Kise, Masakazu Iwamura, Benjamin Adrian, Andreas Dengel 0001 |
ICFHR | 3 |
| 2010 | Object Recognition Based on n-gram Expression of Human ActionsabstractIn this paper, we propose a novel method for recognizing objects by observing human actions based on bag-of-features. The key contribution of our method is that human actions are represented as n-grams of symbols and used to identify specific object categories. First, features of human actions taken on a object are extracted from video images and encoded to symbols. Then, n-grams are generated from the sequence of symbols and registered for corresponding object category. For recognition phase, actions taken on the object are converted into set of n-grams in the same way and compared with ones representing object categories. We performed experiments to recognize objects in an office environment and confirmed the effectiveness of our method. Atsuhiro Kojima, Hiroshi Miki, Koichi Kise |
ICPR | 3 |
| 2009 | Real-Time Camera-Based Recognition of Characters and PictogramsabstractCamera-based character recognition systems should have the capability of quick operation and recognizing perspectively distorted texts in a complex layout. In this paper, in order to realize such a system, we propose a simple but efficient implementation of camera-based recognition of characters and pictograms. With help of new hashing and voting techniques, the proposed method runs well in real-time even on a laptop PC with a web camera. Masakazu Iwamura, Tomohiko Tsuji, Akira Horimatsu, Koichi Kise |
ICDAR | 4 |
| 2009 | Capturing Digital Ink as Retrieving Fragments of Document ImagesabstractThis paper presents a new method of capturing digital ink for pen-based computing. Current technologies such as tablets, ultrasonic and the Anoto pens rely on special mechanisms for locating the pen tip,which result in limiting the applicability.Our proposal is to ease this problem --- a camera pen that allows us to write on ordinary paper for capturing digital ink. A document image retrieval method called LLAH is tuned to locate the pen tip efficiently and accurately on the coordinates of a document only by capturing its tiny fragment.In this paper, we report some results on captured digital ink as well as to evaluate their quality. Kazumasa Iwata, Koichi Kise, Tomohiro Nakai, Masakazu Iwamura, Seiichi Uchida, Shinichiro Omachi |
ICDAR | 2 |
| 2009 | Real-Time Retrieval for Images of Documents in Various Languages Using a Web CameraabstractWe propose a real-time retrieval method for document images in various languages. In this method, queries are images of documents captured by a web-camera. The document images corresponding to the queries are retrieved from the document image database in real time. Since we have already proposed a document image retrieval method for English documents, the proposed method is an extension for retrieval of documents in various languages. In the previous English document image retrieval method, only centroids of word regions are used as feature points. Therefore it cannot be applied to some languages including Japanese and Chinese due to no separation between words and periodic arrangements of characters. In the proposed method, additional features are introduced to realize real-time retrieval for document images in various languages. Tomohiro Nakai, Koichi Kise, Masakazu Iwamura |
ICDAR | 2 |
| 2009 | Detecting Printed and Handwritten Partial Copies of Line Drawings Embedded in Complex BackgroundsabstractThe partial copy is a kind of copy produced by cropping parts of original materials. Illegal users often use this technique for plagiarizing copyrighted materials. In addition, original parts are not necessarily copied intact but may be modified by various techniques, and embedded into other materials, which make the detection quite difficult. In this paper, we propose a method of copyright protection applicable to partial copies aiming at the protection of line drawings such as comics. In order to cope with handwritten partial copies, we apply local feature matching with a database of copyrighted line drawings. Experimental results show that the proposed method not only performs good for detecting printed copies of line drawings, but also has effectiveness on the detection of handwritten ones, even if partial copies are embedded in complex backgrounds. Weihan Sun, Koichi Kise |
ICDAR | 2 |
| 2009 | Conspicuous Character PatternsabstractDetection of characters in scenery images is often a very difficult problem. Although many researchers have tackled this difficult problem and achieved a good performance, it is still difficult to suppress many false alarms and although missings. This paper investigates a conspicuous character pattern, which is a special pattern designed for easier detection. In order to have an example of the conspicuous character pattern, we select a character font with a larger distance from a non-character pattern distribution and, simultaneously, with a smaller distance from a character pattern distribution. Experimental results showed that the character font selected by this method is actually more conspicuous (i.e., detected more easily) than other fonts. Seiichi Uchida, Ryoji Hattori, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICDAR | 5 |
| 2008 | Affine Invariant Recognition of Characters by Progressive PruningabstractThere are many problems to realize camera-based character recognition. One of the problems is that characters in scenes are often distorted by geometric transformations such as affine distortions. Although some methods that remove the affine distortions have been proposed, they cannot remove a rotation transformation of a character. Thus a skew angle of a character has to be determined by examining all the possible angles. However, this consumes quite a bit of time. In this paper, in order to reduce the processing time for an affine invariant recognition, we propose a set of affine invariant features and a new recognition scheme called "progressive pruning."' The progressive pruning gradually prunes less feasible categories and skew angles using multiple classifiers. We confirmed the progressive pruning with the affine invariant features reduced the processing time at least less than half without decreasing the recognition rate. Akira Horimatsu, Ryo Niwa, Masakazu Iwamura, Koichi Kise, Seiichi Uchida, Shinichiro Omachi |
Document Analysis Systems | 4 |
| 2008 | Accuracy Improvement and Objective Evaluation of Annotation Extraction from Printed DocumentsabstractThere is an approach of annotation extraction from printed documents in which annotations are extracted by comparing the image of an annotated document and its original document image. In one of the previous methods, the image of an original document is actually printed and scanned in order to reproduce image degradations of the image of the annotated document. However such a method lacks convenience since users have to use the same printer and scanner to obtain images of an annotated document and its original document. In this paper, we propose an improved annotation extraction method in which the image degradations are compensated by image processing. In the proposed method, the difference between original and annotated document images due to image degradations is reduced by not only removal of the degradations in the annotated document images but also reproduction of the degradations in the original document images. The proposed method consists of three steps of processing which are for dithering, for color change, and for local displacement. We also propose an objective evaluation of extracted annotations to compare the experimental results accurately. Experimental results of the proposed method have shown that the recall of extracted annotations was 80.94% and the precision was 85.59%. Tomohiro Nakai, Kazumasa Iwata, Koichi Kise |
Document Analysis Systems | 3 |
| 2008 | Skew Estimation by InstancesabstractThis paper proposes a novel skew estimation method by instances. The instances to be learned (i.e., stored) are rotation invariants and a rotation variant for each character category. Using the instances, it is possible to estimate a skew angle of each individual character on a document. This fact implies that the proposed method can estimate the skew angle of a document where characters do not form long straight text lines. Thus, the proposed method will be applicable to various documents such as signboard images captured by a camera. Experimental evaluation using synthetic and real images revealed the expected robustness against various character layouts. Seiichi Uchida, Megumi Sakai, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
Document Analysis Systems | 5 |
| 2008 | Memory efficient recognition of specific objects with local featuresabstractBalancing the recognition rate, processing time and memory requirement is an important issue for object recognition based on local features. For the task of recognizing not generic but specific objects (object instances), a larger number of local features enable us to improve the recognition rate but pose a problem of processing time and memory requirement. For the problem of processing time, approximate nearest neighbor search is known to be extremely effective. In this paper, we propose a method of memory reduction by applying scalar quantization to local features. From experimental results on 100,000 images, we have found that the scalar quantization is of great help; As compared to the original representation with 16 bit/dimension, the recognition rate of 98.1 % was kept unchanged using the representation with 2 bit/dim. Koichi Kise, Kazuto Noguchi, Masakazu Iwamura |
ICPR | 1 |
| 2007 | Improvement of Retrieval Speed and Required Amount of Memory for Geometric Hashing by Combining Local InvariantsabstractThe geometric hashing (GH) is a well-known model-based object recognition technique with good properties both in retrieval speed and required amount of memory. However, it has a significant weak point; as the number of objects increases, both retrieval speed and required amount of memory increase in the cubic, fourth or higher order. Recently, a new technique “locally likely arrangement hashing (LLAH) ” whose computational cost is a linear order has been proposed. The objective of the current paper is to reveal how LLAH improves the performance. By comparing GH and LLAH, we describe four primary factors of the performance improvement. 1 Masakazu Iwamura, Tomohiro Nakai, Koichi Kise |
BMVC | 3 |
| 2007 | Simple Representation and Approximate Search of Feature Vectors for Large-Scale Object RecognitionabstractThis paper presents two methods of large-scale recognition of planar objects with a simple representation and approximate search of local feature vectors. A central problem of the use of local feature vectors is the burden of computation and memory for finding nearest neighbors. To solve this problem, the proposed methods embody the following: (1) a simple bit representation of feature vectors and hashing enable us to fast access with less memory, (2) approximate search with query perturbation allows us to find approximate nearest neighbors efficiently. From large-scale experiments using 10,000 objects in the database and 2,000 query images, it was found that only 10–20% of correct nearest neighbors were enough for achieving recognition rate of 98.0%. The processing time for achieving this rate was 8.3 ms / query (excluding time for feature extraction). We have also tested the scalability of a proposed method using the database of 100,000 objects and obtained the result of 92.3 % accuracy in 4.5 ms /query. 1 Koichi Kise, Kazuto Noguchi, Masakazu Iwamura |
BMVC | 1 |
| 2007 | Extraction of Embedded Class Information from Universal Character PatternabstractThis paper is concerned with a universal pattern, which is defined as a character pattern designed to have high machine-readability. This universal pattern is a charac- ter pattern printed with stripes. The cross ratio calculated from the widths of the stripes represents the character class. Thus, if the boundaries of the stripes can be detected for measuring the widths, the class can be determined without ordinary recognition process. Furthermore, since the cross ratio is invariant to projective distortions, the correct class will be still determined under those distortions. This pa- per describes a practical scheme to recognize this universal pattern. The proposed scheme includes a novel algorithm to detect the stripe boundaries stably even from the universal pattern image contaminated by non-uniform lighting and noise. The algorithm is realized by a combination of a dy- namic programming-based optimal boundary detection and a finite state automaton which represents the property of the universal pattern. Experimental results showed the pro- posed scheme could recognize 99.6% of the universal pat- tern images which underwent heavy projective distortions and non-uniform lighting. Seiichi Uchida, Megumi Sakai, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICDAR | 5 |
| 2006 | Use of Affine Invariants in Locally Likely Arrangement Hashing for Camera-Based Document Image Retrieval
Tomohiro Nakai, Koichi Kise, Masakazu Iwamura |
Document Analysis Systems | 2 |
| 2005 | Camera-Based Document Image Retrieval as Voting for Partial Signatures of Projective InvariantsabstractWe propose a method of document image retrieval using digital cameras. The proposed method takes as input a part or the whole of a document acquired as a query by a digital camera, and retrieves a document image that includes the query. For this purpose, it is required to solve the problem of "perspective distortion" of images, as well as to establish a way of matching parts of document images flexibly. These are achieved based on the following characteristics of the proposed method: (1) indexing of document images using the projective invariants called the "cross-ratios", (2) retrieval as voting for partial signatures of document images defined by the cross-ratios. From experimental results using digital cameras with high and low resolutions, we demonstrate the effectiveness of the proposed method. Tomohiro Nakai, Koichi Kise, Masakazu Iwamura |
ICDAR | 2 |
| 2004 | Document Image Retrieval in a Question Answering System for Document Images
Koichi Kise, Shota Fukushima, Keinosuke Matsumoto |
Document Analysis Systems | 1 |
| 2003 | Stippling Data on Backgrounds of Pages - Toward Seamless Integration of Paper and Electronic DocumentsabstractIn order to realize seamless integration of paper and electronic documents, it is at least necessary to assure error free conversion from one to the other. In general, the conversion from paper to electronic documents is the task of document image understanding. Although its research has made remarkable progress, it is still a hard task without limiting the type of documents. This paper presents a completely different approach to this task on condition that printed documents have their originals in electronic form. The proposed method employs fine dots to represent data of electronic documents and places the dots on white space (backgrounds) of pages. Since the data is encoded with an error correcting code, it is guaranteed to be correctly recovered from the scanned images of documents. Experimental results show that a page with normal foreground objects (characters and other things) can contain more than 4KB of data, even when errors up to 20% of the data are permitted. 1. Koichi Kise, Yasuo Miki, Keinosuke Matsumoto |
ICDAR | 1 |
| 2003 | Document Image Retrieval Based on 2D Density Distributions of Terms with Pseudo Relevance FeedbackabstractDocument image retrieval is a task to retrieve document images relevant to a user's query. Most existing methods based on word-level indexing rely on the representation called "bag of words" which originated in the field of information retrieval. This paper presents a new representation of documents that utilizes additional information about the location of words in pages so as to improve the retrieval performance. We consider that pages are relevant to a query if they contain its terms densely. This notion is embodied as density distributions of terms calculated in the proposed method. Its performance is improved with the help of "pseudo relevance feedback", i.e., a method of expanding a query by analyzing pages. Experimental results on English document images show that the proposed method is superior to conventional methods of electronic document retrieval at recall levels 0.0-0.6. Koichi Kise, Wuotang Yin, Keinosuke Matsumoto |
ICDAR | 1 |
| 2003 | Reports of the DAS02 working groups
Elisa H. Barney Smith, David Monn, Sriharsha Veeramachaneni, Koichi Kise, Alessio Malizia, Leon Todoran, Adnan El-Nasan, Rolf Ingold |
Int. J. Document Anal. Recognit. | 4 |
| 2002 | Spotting Where to Read on Pages - Retrieval of Relevant Parts from Page Images
Koichi Kise, Masaaki Tsujino, Keinosuke Matsumoto |
Document Analysis Systems | 1 |
| 2001 | Passage-Based Document Retrieval as a Tool for Text Mining with User's Information Needs
Koichi Kise, Markus Junker 0002, Andreas Dengel 0001, Keinosuke Matsumoto |
Discovery Science | 1 |
| 2001 | Experimental Evaluation of Passage-Based Document RetrievalabstractRetrieval of electronic documents is a fundamental component for intelligent access to the contents of documents. For the retrieval of long documents, a method called passage-based document retrieval has proven to be effective. In this paper we experimentally show that the passage-based retrieval is also advantageous for dealing with short queries on condition that documents are long. We employ a passage-based method based on density distributions of query terms in documents, and compare it with three conventional methods: the vector space model, pseudo-feedback and latent semantic indexing. Koichi Kise, Markus Junker 0002, Andreas Dengel 0001, Keinosuke Matsumoto |
ICDAR | 1 |
| 2000 | Backgrounds as Information Carriers for Printed DocumentsabstractThis paper presents a method of embedding and recovery of a large amount of data on printed documents. The proposed method is characterized by the following points: (1) As the medium for recording the data, backgrounds (white space) of pages are utilized. Since white space is dominant in ordinary pages, this enables us to record a large amount of data. (2) As the representation of data, an arbitrary figure (such as a logo mark of a company) is stippled to obtain a set of points, each of which is superimposed on a page image. This allows us to prevent the appearance of a page from deteriorating, since stippled figures play a role of background patterns on printed documents. The experimental results show that the method is capable of both recording 7.8KB of data and reading the embedded data with the accuracy of 99%. Koichi Kise, Yasuo Miki, Keinosuke Matsumoto |
ICPR | 1 |
| 1999 | On the Use of Density Distribution of Keywords for Automated Generation of Hypertext Links from Arbitrary Parts of DocumentsabstractThis paper presents a method of automated generation of hypertext links for electronic documents. The goal is to generate links from an arbitrary part of a document (a source of a link) to its relevant parts of target documents (destinations). To achieve this goal, we assume that words are often shared by parts of documents if these parts are relevant with each other. In order to extract parts densely including words of a source (keywords), we employ density distributions of keywords. This enables us to determine destinations simply by extracting parts whose density exceeds a threshold. Experiments on generating links from figures/tables to parts of documents, as well as from texts to parts of different documents show that our method with the optimal parameters yields recall of 60% and precision of 50%. Koichi Kise, Hiroyuki Mizuno, Masashi Yamaguchi, Keinosuke Matsumoto |
ICDAR | 1 |
| 1998 | Text-Line Extraction as Selection of Paths in the Neighbor Graph
Koichi Kise, Motoi Iwata, Andreas Dengel 0001, Keinosuke Matsumoto |
Document Analysis Systems | 1 |
| 1998 | Segmentation of Page Images Using the Area Voronoi Diagram
Koichi Kise, Akinori Sato, Motoi Iwata |
Comput. Vis. Image Underst. | 1 |
| 1996 | Page segmentation based on thinning of backgroundabstractThis paper presents a new method of page segmentation based on the analysis of background (white areas). The proposed method is capable of segmenting pages with non-rectangular layout as well as with various angles of skew. The characteristics of the method are as follows: (1) thinning of the background enables us to represent white areas of any shape as connected thin lines or chains and the robustness for tilted page images is also achieved by the representation; and (2) based on this representation, the task of page segmentation is defined as to find the loops enclosing printed areas. The task is achieved by eliminating unnecessary chains using not only a feature of white areas, but also a feature of black areas divided by a chain. Based on the experimental results and the comparison with previous methods, we discuss the advantages and limitations of the proposed method. Koichi Kise, Osamu Yanagida, Shinobu Takamatsu |
ICPR | 1 |
| 1995 | Interpretation of conceptual diagrams from line segments and stringsabstractA conceptual diagram is a kind of diagram which represents a logical structure among concepts using simple physical objects (loops, lines and character strings). This paper presents a method of interpretation of conceptual diagrams. In conceptual diagrams, a single physical object plays various logical roles depending on surrounding physical objects, because there are no specific ruled of writing conceptual diagrams. To cope with this problem, we introduce the strategy of hypothesis generation and verification; hypothesized interpretations are verified by relaxation which takes account of logical relations to other physical objects. From the experimental results using 50 conceptual diagrams, we discuss the effectiveness of our method. Koichi Kise, Noriyoshi Yoneda, Shinobu Takamatsu, Kunio Fukunaga |
ICDAR | 1 |
| 1993 | Incremental acquisition of knowledge about layout structures from examples of documentsabstractDocument image analysis systems often utilize the knowledge about layout structures to extract layout objects labeled logically. However, the lack of the facility for knowledge acquisition limits the applicability of the systems. The authors propose a method of acquiring knowledge for document image analysis. Given examples of document images and their layout objects labeled logically, the method generates and modifies the knowledge. The method is incremental so that the knowledge can be efficiently modified using additional examples. Counterexamples generated as errors obtained from the analysis of an example image can also be reflected into the knowledge so that the system may no longer generate the errors. Experimental results on both knowledge acquisition and the analysis using the acquired knowledge are also presented.> Koichi Kise, Naoko Yajima, Noboru Babaguchi, Kunio Fukunaga |
ICDAR | 1 |
| 1992 | Model based system for analyzing document imagesabstractDocument image analysis is the process of deriving logically structured representation of a document by analyzing the layout structure of its image. This paper proposes a knowledge based system for document image analysis which is applicable to various kinds of documents. The characteristics of the system are as follows: (1) The knowledge base called document model encodes only object-level knowledge hierarchically, declaratively and symbolically, aiming at high expressivity and maintainability of the knowledge description; (2) the document model is automatically constructed by referring samples of document images, and incrementally refined by feedback of error information of analysis.> Koichi Kise, Masaki Yamaoka, Noboru Babaguchi, Yoshikazu Tezuka |
ICPR (2) | 1 |
| 1991 | Connectionist Model BinarizationabstractImage binarization is a task to convert gray-level images into bi-level ones. Its underlying notion can be simply thought of as threshold selection. However, the result of binarization will cause significant influence on the process of image recognition or understanding. In this paper we discuss a new binarization method, named CMB (connectionist model binarization), which uses the connectionist model. In the method a gray-level histogram is input to a multilayer network trained with the back-propagation algorithm to obtain a threshold which gives a visually suitable binarized image. From the experimental results, it was verified that CMB is an effective binarization method in comparison with other methods. Noboru Babaguchi, Koji Yamada, Koichi Kise, Yoshikazu Tezuka |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1990 | Connectionist model binarizationabstractThe application of a connectionist model to an image binarization method called connectionist model binarization (CMB) is discussed. CMB employs a multilayer network of a connectionist model whose input and output are a histogram and a desirable threshold for binarization, respectively. This network is trained with a back-propagation algorithm to output a threshold which gives a visually suitable binarised image against any histogram. The details of CMB are described, and its learning strategy and binarization performance are discussed.> Noboru Babaguchi, Koji Yamada, Koichi Kise, Yoshikazu Tezuka |
ICPR (2) | 3 |
| 1988 | Visiting card understanding systemabstractThe authors present the visiting card understanding system, whose output is suitable for the input of a visiting-card database. The system consists of two modules. One is a document model which represents the hierarchical knowledge about the layout structure of visiting cards. The other is an understanding module which interprets the document model to general and test hierarchical hypotheses about the contents of a visiting card. Since the understanding module is fundamentally independent of document type, the system is applicable to many kinds of documents.> Koichi Kise, Koji Yamada, Naoki Tanaka, Noboru Babaguchi, Yoshikazu Tezuka |
ICPR | 1 |