VLDB 2026 Research / reviewers in the wild / expert
Yusuke Matsui 0001
dblp:56/10540-1
· DBLP profile ↗
38ranked-venue papers
13as first author
19since 2021 · last 2025
0000-0003-1529-0154ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 11 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff TableabstractApproximate nearest neighbor search (ANNS) is an essential building block for applications like RAG but can sometimes yield results that are overly similar to each other. In certain scenarios, search results should be similar to the query and yet diverse. We propose LotusFilter, a post-processing module to diversify ANNS results. We precompute a cutoff table summarizing vectors that are close to each other. During the filtering, LotusFilter greedily looks up the table to delete redundant vectors from the candidates. We demonstrated that the LotusFilter operates fast (0.02 [ms/query]) in settings resembling real-world RAG applications, utilizing features such as OpenAI embeddings. Our code is publicly available at https://github.com/matsui528/lotf. Yusuke Matsui 0001 |
CVPR | 1 |
| 2025 | Noisy Label Refinement with Semantically Reliable Synthetic ImagesabstractSemantic noise in image classification datasets, where visually similar categories are frequently mislabeled, poses a significant challenge to conventional supervised learning approaches. In this paper, we explore the potential of using synthetic images generated by advanced text-to-image models to address this issue. Although these high-quality synthetic images come with reliable labels, their direct application in training is limited by domain gaps and diversity constraints. Unlike conventional approaches, we propose a novel method that leverages synthetic images as reliable reference points to identify and correct mislabeled samples in noisy datasets. Extensive experiments across multiple benchmark datasets show that our approach significantly improves classification accuracy under various noise conditions, especially in challenging scenarios with semantic label noise. Additionally, since our method is orthogonal to existing noise-robust learning techniques, when combined with state-of-the-art noise-robust training methods, it achieves superior performance, improving accuracy by 30% on CIFAR-10 and by 11% on CIFAR-100 under 70% semantic noise, and by 24% on ImageNet-100 under real-world noise conditions. Yingxuan Li, Jiafeng Mao, Yusuke Matsui 0001 |
ICIP | 3 |
| 2024 | Cross-Lingual Learning in Multilingual Scene Text RecognitionabstractIn this paper, we investigate cross-lingual learning (CLL) for multilingual scene text recognition (STR). CLL transfers knowledge from one language to another. We aim to find the condition that exploits knowledge from high-resource languages for improving performance in low-resource languages. To do so, we first examine if two general insights about CLL discussed in previous works are applied to multilingual STR: (1) Joint learning with high- and low-resource languages may reduce performance on low-resource languages, and (2) CLL works best between typologically similar languages. Through extensive experiments, we show that two general insights may not be applied to multilingual STR. After that, we show that the crucial condition for CLL is the dataset size of high-resource languages regardless of the kind of high-resource languages. Our code, data, and models are available at https://github.com/ku21fan/CLL-STR. Jeonghun Baek, Yusuke Matsui 0001, Kiyoharu Aizawa |
ICASSP | 2 |
| 2024 | Manga109Dialog: A Large-Scale Dialogue Dataset for Comics Speaker DetectionabstractThe expanding market for e-comics has driven the development of automated methods for analyzing comics. To enhance the machine’s understanding of comics, an automated method is essential for linking text in comics to characters that speak those words. In this study, we developed Manga109Dialog1, which is the world’s largest speaker-to-text annotation dataset for comics, containing 132,692 pairs. We proposed a novel deep learning-based method using scene graph generation models. To tailor the unique features of comics, we enhanced the performance by considering the frame reading order. Our experiments with Manga109Dialog show that our scene-graph-based approach outperforms existing methods, achieving a prediction accuracy of over 75%, thus establishing a robust benchmark for speaker detection in comics. Yingxuan Li, Kiyoharu Aizawa, Yusuke Matsui 0001 |
ICME | 3 |
| 2024 | Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
Yingxuan Li, Ryota Hinami, Kiyoharu Aizawa, Yusuke Matsui 0001 |
ACM Multimedia | 4 |
| 2024 | A Flow-Based Centralized Route Guidance System for Traffic Congestion MitigationabstractWe propose a new route guidance system (RGS) to mitigate traffic congestion called Flow-based Capacity-aware Rerouting (FCR), which achieves better global optimality by considering the capacity of detour routes. FCR introduces a flow-based traffic management method that enables us to compute detour paths in units of flow rather than per vehicle processing, which reduces the load to calculate detour paths and makes the centralized system practically feasible. With the flow-based traffic management, FCR also enables us to consider traffic volume that should bypass the congested roads to mitigate the congestion and suggest detour paths to the corresponding traffic volume of vehicles. By prioritizing travel time loss when choosing detour routes to apply, FCR leads to better global optimality with the limited volume of rerouting traffic while considering each rerouted vehicle’s benefit, such as travel time loss. FCR offers a new centralized system architecture for actual real-time traffic control, which improves both global optimality and each vehicle’s benefit within a feasible computational load. Yusuke Matsui 0001, Takuya Yoshihiro |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Subset Retrieval Nearest Neighbor Machine TranslationabstractHiroyuki Deguchi, Taro Watanabe, Yusuke Matsui, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hiroyuki Deguchi 0002, Taro Watanabe, Yusuke Matsui 0001, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita |
ACL (1) | 3 |
| 2023 | Relative NN-Descent: A Fast Index Construction for Graph-Based Approximate Nearest Neighbor SearchabstractApproximate Nearest Neighbor Search (ANNS) is the task of finding the database vector that is closest to a given query vector. Graph-based ANNS is the family of methods with the best balance of accuracy and speed for million-scale datasets. However, graph-based methods have the disadvantage of long index construction time. Recently, many researchers have improved the tradeoff between accuracy and speed during a search. However, there is little research on accelerating index construction. We propose a fast graph construction algorithm, Relative NN-Descent (RNN-Descent). RNN-Descent combines NN-Descent, an algorithm for constructing approximate K-nearest neighbor graphs (K-NN graphs), and RNG Strategy, an algorithm for selecting edges effective for search. This algorithm allows the direct construction of graph-based indexes without ANNS. Experimental results demonstrated that the proposed method had the fastest index construction speed, while its search performance is comparable to existing state-of-the-art methods such as NSG. For example, in experiments on the GIST1M dataset, the construction of the proposed method is 2x faster than NSG. Additionally, it was even faster than the construction speed of NN-Descent. Naoki Ono, Yusuke Matsui 0001 |
ACM Multimedia | 2 |
| 2023 | Fast Partitioned Learned Bloom FilterabstractA Bloom filter is a memory-efficient data structure for approximate membership queries used in numerous fields of computer science.
Recently, learned Bloom filters that achieve better memory efficiency using machine learning models have attracted attention.
One such filter, the partitioned learned Bloom filter (PLBF), achieves excellent memory efficiency.
However, PLBF requires a $\mathcal{O}(N^3k)$ time complexity to construct the data structure, where $N$ and $k$ are the hyperparameters of PLBF.
One can improve memory efficiency by increasing $N$, but the construction time becomes extremely long.
Thus, we propose two methods that can reduce the construction time while maintaining the memory efficiency of PLBF.
First, we propose fast PLBF, which can construct the same data structure as PLBF with a smaller time complexity $\mathcal{O}(N^2k)$.
Second, we propose fast PLBF++, which can construct the data structure with even smaller time complexity $\mathcal{O}(Nk\log N + Nk^2)$.
Fast PLBF++ does not necessarily construct the same data structure as PLBF.
Still, it is almost as memory efficient as PLBF, and it is proved that fast PLBF++ has the same data structure as PLBF when the distribution satisfies a certain constraint.
Our experimental results from real-world datasets show that (i) fast PLBF and fast PLBF++ can construct the data structure up to 233 and 761 times faster than PLBF, (ii) fast PLBF can achieve the same memory efficiency as PLBF, and (iii) fast PLBF++ can achieve almost the same memory efficiency as PLBF.
The codes are available at [this https URL](https://github.com/atsukisato/FastPLBF). Atsuki Sato, Yusuke Matsui 0001 |
NeurIPS | 2 |
| 2023 | General and Practical Tuning Method for Off-the-Shelf Graph-Based Index: SISAP Indexing Challenge Report by Team UTokyo
Yutaro Oguri, Yusuke Matsui 0001 |
SISAP | 2 |
| 2022 | COO: Comic Onomatopoeia Dataset for Recognizing Arbitrary or Truncated Texts
Jeonghun Baek, Yusuke Matsui 0001, Kiyoharu Aizawa |
ECCV (28) | 2 |
| 2022 | ARM 4-BIT PQ: SIMD-Based Acceleration for Approximate Nearest Neighbor Search on ARMabstractWe accelerate the 4-bit product quantization (PQ) on the ARM architecture. Notably, the drastic performance of the conventional 4-bit PQ strongly relies on x64-specific SIMD register, such as AVX2; hence, we cannot yet achieve such good performance on ARM. To fill this gap, we first bundle two 128-bit registers as one 256-bit component. We then apply shuffle operations for each using the ARM-specific NEON instruction. By making this simple but critical modification, we achieve a dramatic speedup for the 4-bit PQ on an ARM architecture. Experiments show that the proposed method consistently achieves a 10x improvement over the naive PQ with the same accuracy. Yusuke Matsui 0001, Yoshiki Imaizumi, Naoya Miyamoto, Naoki Yoshifuji |
ICASSP | 1 |
| 2022 | Translation of Illustration Artist Style Using Sailormoonredraw DataabstractThe decision to draw requires answers to two questions: what to draw and how to draw. The latter refers to artist style and is an important factor in creating any illustration. In this paper, we propose a novel task, artist style translation, which translates one artist’s illustration into that of another artist style using deep learning. To solve this task in a supervised manner, we created a novel illustration dataset, SailormoonDataset, which consists of more than 2,000 artist’s stylistic illustrations of the same content. In addition, we propose a method based on the Swapping Autoencoder by introducing a new loss function for supervised learning and using multiple images to represent an artist style. We translate the face illustration of Sailor Moon into styles of different artists. By comparing the current results to those of the Swapping Autoencoder, we find that the proposed method successfully achieves the artist style translation. Keita Awane, Daichi Horita, Hikaru Ikuta, Yusuke Matsui 0001, Kiyoharu Aizawa, Naohiro Yanase |
ICIP | 4 |
| 2021 | Towards Fully Automated Manga TranslationabstractWe tackle the problem of machine translation of manga, Japanese comics. Manga translation involves two important problems in machine translation: context-aware and multimodal translation. Since text and images are mixed up in an unstructured fashion in Manga, obtaining context from the image is essential for manga translation. However, it is still an open problem how to extract context from image and integrate into MT models. In addition, corpus and benchmarks to train and evaluate such model is currently unavailable. In this paper, we make the following four contributions that establishes the foundation of manga translation research. First, we propose multimodal context-aware translation framework. We are the first to incorporate context information obtained from manga image. It enables us to translate texts in speech bubbles that cannot be translated without using context information (e.g., texts in other speech bubbles, gender of speakers, etc.). Second, for training the model, we propose the approach to automatic corpus construction from pairs of original manga and their translations, by which large parallel corpus can be constructed without any manual labeling. Third, we created a new benchmark to evaluate manga translation. Finally, on top of our proposed methods, we devised a first compleheisive system for fully automated manga translation. Ryota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke Matsui 0001 |
AAAI | 4 |
| 2021 | Cascading Feature Extraction for Fast Point Cloud Registration
Yoichiro Hisadome, Yusuke Matsui 0001 |
BMVC | 2 |
| 2021 | What if We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition With Fewer LabelsabstractScene text recognition (STR) task has a common practice: All state-of-the-art STR models are trained on large synthetic data. In contrast to this practice, training STR models only on fewer real labels (STR with fewer labels) is important when we have to train STR models without synthetic data: for handwritten or artistic texts that are difficult to generate synthetically and for languages other than English for which we do not always have synthetic data. However, there has been implicit common knowledge that training STR models on real data is nearly impossible because real data is insufficient. We consider that this common knowledge has obstructed the study of STR with fewer labels. In this work, we would like to reactivate STR with fewer labels by disproving the common knowledge. We consolidate recently accumulated public real data and show that we can train STR models satisfactorily only with real labeled data. Subsequently, we find simple data augmentation to fully exploit real data. Furthermore, we improve the models by collecting unlabeled data and introducing semi- and self-supervised methods. As a result, we obtain a competitive model to state-of-the-art methods. To the best of our knowledge, this is the first study that 1) shows sufficient performance by only using real labels and 2) introduces semi- and self-supervised methods into STR with fewer labels. Our code and data are available: https://github.com/ku21fan/STR-Fewer-Labels. Jeonghun Baek, Yusuke Matsui 0001, Kiyoharu Aizawa |
CVPR | 2 |
| 2021 | Improving The Quality Of Illustrations: Transforming Amateur Illustrations To A Professional StandardabstractWe propose an amateur- to professional-level illustration translator that can modify amateur illustrations slightly to produce professional-level quality images. The proposed translator is a GAN-based image translation module. We focus only on the neighboring region of a contour to improve the quality of the illustration by applying image completion to the neighboring region of the extracted line drawing. We artificially augment amateur-level illustrations from professional-level illustrations to solve the lack of a pair of amateur-level and professional-level datasets. This enables us to automatically prepare a pair of amateur-level and professional-level images, through which we can train a translator network. Through experiments and user study, we show that the proposed method improves the quality of amateur illustrations. Keita Awane, Koki Tsubota, Hikaru Ikuta, Yusuke Matsui 0001, Kiyoharu Aizawa, Naohiro Yanase |
ICIP | 4 |
| 2021 | Efficient Nearest Neighbor Search by Removing Anti-hubabstractThe central research question of the nearest neighbor search is how to reduce the memory cost while maintaining its accuracy. Instead of compressing each vector as is done in the existing methods, we propose a way to subsample unnecessary vectors to save memory. We empirically found that such unnecessary vectors have low hubness scores and thus can be easily identified beforehand. Such points are called anti-hubs in the data mining community. By removing anti-hubs, we achieved a memory-efficient search while preserving accuracy. In million-scale experiments, we showed that any vector compression method improves search accuracy by partial replacement with anti-hub removal under the same memory usage. A billion-scale benchmark showed that our data reduction combined with the best search method achieves higher accuracy under the assumption of fixed memory consumption. For example, our method had a much higher [email protected] (0.53) compared with the existing method (0.23) for the same memory consumption (6GB). Kimihiro Tanaka, Yusuke Matsui 0001, Shin'ichi Satoh 0001 |
ICMR | 2 |
| 2021 | A Route Guidance Method for Vehicles to Improve Driver's Experienced Delay Against Traffic Congestion
Yusuke Matsui 0001, Takuya Yoshihiro |
MobiQuitous | 1 |
| 2020 | Effective and Efficient: Toward Open-world Instance Re-identificationabstractInstance Re-identification (ReID) system facilitates various applications that require painful and boring video watching. Its efficiency and effectiveness accelerate the process of video analysis. In this tutorial, we summarize ReID technologies and provide an overview. We'll introduce fundamental technologies, existing challenges, trends, etc. This tutorial would be useful for multimedia content analysis and system-level multimedia retrieval, especially for an effective and efficient open-world ReID system for the practical, large-scale, and open-set domain. Zheng Wang 0007, Wu Liu 0005, Yusuke Matsui 0001, Shin'ichi Satoh 0001 |
ACM Multimedia | 3 |
| 2019 | Efficient Image Retrieval via Decoupling Diffusion into Online and Offline ProcessingabstractDiffusion is commonly used as a ranking or re-ranking method in retrieval tasks to achieve higher retrieval performance, and has attracted lots of attention in recent years. A downside to diffusion is that it performs slowly in comparison to the naive k-NN search, which causes a non-trivial online computational cost on large datasets. To overcome this weakness, we propose a novel diffusion technique in this paper. In our work, instead of applying diffusion to the query, we precompute the diffusion results of each element in the database, making the online search a simple linear combination on top of the k-NN search process. Our proposed method becomes 10∼ times faster in terms of online search speed. Moreover, we propose to use late truncation instead of early truncation in previous works to achieve better retrieval performance. Fan Yang 0038, Ryota Hinami, Yusuke Matsui 0001, Steven Ly, Shin'ichi Satoh 0001 |
AAAI | 3 |
| 2018 | Revisiting Column-Wise Vector Quantization for Memory-Efficient Matrix MultiplicationabstractMatrix multiplication is a fundamental operation for data analysis. Memory-efficient matrix multiplication is always essential, especially for problems involving large-scale datasets. We present PQMM, a column-wise product quantization for memory-efficinet matrix multiplication. A simple idea of column-wise vector quantization is revisited. Each column of a data matrix is quantized into an index and the matrix product is approximated by the sum of the outer products obtained from the codewords fetched by the indices. The empirical evaluation illustrates that, with the same amount of memory, PQMM outperforms the state-of-the-art sketch-based method with respect to the error. We further show that a combination of the sketch-based method and PQMM results in a slightly less accurate data representation, but which is significantly more memory-efficient. Yusuke Matsui 0001, Shin'ichi Satoh 0001 |
ICIP | 1 |
| 2018 | Reconfigurable Inverted IndexabstractExisting approximate nearest neighbor search systems suffer from two fundamental problems that are of practical importance but have not received sufficient attention from the research community. First, although existing systems perform well for the whole database, it is difficult to run a search over a subset of the database. Second, there has been no discussion concerning the performance decrement after many items have been newly added to a system. We develop a reconfigurable inverted index (Rii) to resolve these two issues. Based on the standard IVFADC system, we design a data layout such that items are stored linearly. This enables us to efficiently run a subset search by switching the search method to a linear PQ scan if the size of a subset is small. Owing to the linear layout, the data structure can be dynamically adjusted after new items are added, maintaining the fast speed of the system. Extensive comparisons show that Rii achieves a comparable performance with state-of-the art systems such as Faiss. Yusuke Matsui 0001, Ryota Hinami, Shin'ichi Satoh 0001 |
ACM Multimedia | 1 |
| 2018 | PQTable: Nonexhaustive Fast Search for Product-Quantized Codes Using Hash TablesabstractIn this paper, we propose a product quantization table (PQTable)-a fast search method for product-quantized codes via hash tables. An identifier of each database vector is associated with the slot of a hash table by using its PQ-code as a key. For querying, an input vector is PQ-encoded and hashed, and the items associated with that code are then retrieved. The proposed PQTable produces the same results as a linear PQ scan, and is 102-105times faster. Although the state-of-the-art performance can be achieved by previous inverted-indexing-based approaches, such methods require manually designed parameter setting and significant training; our PQTable is free of these limitations, and therefore offers a practical and effective solution for real-world problems. Specifically, when the vectors are highly compressed, our PQTable achieves one of the fastest search performances on a single CPU to date with significantly efficient memory usage (0.059-ms per query over 109data points with just 5.5-GB memory consumption). Finally, we show that our proposed PQTable can naturally handle the codes of an optimized product quantization (OPQTable). Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
IEEE Trans. Multim. | 1 |
| 2017 | Energy based fast event retrieval in video with temporal match kernelabstractWe propose a fast event retrieval method in video databases based on temporal match kernel. The basic idea of the baseline is to maximize the score function for all possible relative timestamps. However, considering the stability and computational cost, we simplify the similarity score by calculating the energy of the score function. In this way, maximizing the score function which is time-consuming and sensitive to noise is avoided. We derive the simplified energy formulation by using Parseval's theorem, which describes the unitarity of Fourier transform. Further, by our formulation, the problem becomes a simple nearest neighbor search, and we can use product quantization (PQ) to accelerate the computation. We evaluate our method on EVVE dataset for event retrieval. The experimental results show that our approach is much faster than the baseline, with a higher mAP as well. Comparison with state of the art demonstrates the efficacy of our approach. Junfu Pu, Yusuke Matsui 0001, Fan Yang 0038, Shin'ichi Satoh 0001 |
ICIP | 2 |
| 2017 | Deep Image Retrieval Applied on Kotenseki Ancient Japanese LiteratureabstractKotenseki is a collection of classical and ancient Japanese literature. It is comprised of image books that express Japanese stories by using comic drawings of different characters, such as humans, nature, and animals. To effectively store them for posterity, a search system is important. We propose an efficient CBIR system to assist the users in easily accessing the information and have an enjoyable experience browsing Kotenseki images. There are two main functions comprising keyword-based and image-based queries. We also provide automatic detection of objects within the original images to create a database of feature vectors. Our study utilizes the benefits of deep-feature and object-detection techniques. We also present our fine-tuned model, which adds L2 normalization, dropout, and data augmentation to obtain better accuracy, which can reach 84% of mAP in our Kotenseki dataset. Chairath Sirirattanapol, Yusuke Matsui 0001, Shin'ichi Satoh 0001, Kuninori Matsuda, Kazuaki Yamamoto |
ISM | 2 |
| 2017 | Region-Based Image Retrieval RevisitedabstractRegion-based image retrieval (RBIR) technique is revisited. In early attempts at RBIR in the late 90s, researchers found many ways to specify region-based queries and spatial relationships; however, the way to characterize the regions, such as by using color histograms, were very poor at that time. Here, we revisit RBIR by incorporating semantic specification of objects and intuitive specification of spatial relationships. Our contributions are the following. First, to support multiple aspects of semantic object specification (category, instance, and attribute), we propose a multitask CNN feature that allows us to use deep learning technique and to jointly handle multi-aspect object specification. Second, to help users specify spatial relationships among objects in an intuitive way, we propose recommendation techniques of spatial relationships. In particular, by mining the search results, a system can recommend feasible spatial relationships among the objects. The system also can recommend likely spatial relationships by assigned object category names based on language prior. Moreover, object-level inverted indexing supports very fast shortlist generation, and re-ranking based on spatial constraints provides users with instant RBIR experiences. Ryota Hinami, Yusuke Matsui 0001, Shin'ichi Satoh 0001 |
ACM Multimedia | 2 |
| 2017 | PQk-means: Billion-scale Clustering for Product-quantized CodesabstractData clustering is a fundamental operation in data analysis. For handling large-scale data, the standard k-means clustering method is not only slow, but also memory-inefficient. We propose an efficient clustering method for billion-scale feature vectors, called PQk-means. By first compressing input vectors into short product-quantized (PQ) codes, PQk-means achieves fast and memory-efficient clustering, even for high-dimensional vectors. Similar to k-means, PQk-means repeats the assignment and update steps, both of which can be performed in the PQ-code domain. Experimental results show that even short-length (32 bit) PQ-codes can produce competitive results compared with k-means. This result is of practical importance for clustering in memory-restricted environments. Using the proposed PQk-means scheme, the clustering of one billion 128D SIFT features with K = 105 is achieved within 14 hours, using just 32 GB of memory consumption on a single computer. Yusuke Matsui 0001, Keisuke Ogaki, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 1 |
| 2017 | Sketch-based manga retrieval using manga109 datasetabstractManga (Japanese comics) are popular worldwide. However, current e-manga archives offer very limited search support, i.e., keyword-based search by title or author. To make the manga search experience more intuitive, efficient, and enjoyable, we propose a manga-specific image retrieval system. The proposed system consists of efficient margin labeling, edge orientation histogram feature description with screen tone removal, and approximate nearest-neighbor search using product quantization. For querying, the system provides a sketch-based interface. Based on the interface, two interactive reranking schemes are presented: relevance feedback and query retouch. For evaluation, we built a novel dataset of manga images, Manga109, which consists of 109 comic books of 21,142 pages drawn by professional manga artists. To the best of our knowledge, Manga109 is currently the biggest dataset of manga images available for research. Experimental results showed that the proposed framework is efficient and scalable (70 ms from 21,142 pages using a single computer with 204 MB RAM). Yusuke Matsui 0001, Kota Ito, Yuji Aramaki, Azuma Fujimoto, Toru Ogawa, Toshihiko Yamasaki, Kiyoharu Aizawa |
Multim. Tools Appl. | 1 |
| 2017 | DrawFromDrawings: 2D Drawing Assistance via Stroke Interpolation with a Sketch DatabaseabstractWe present DrawFromDrawings, an interactive drawing system that provides users with visual feedback for assistance in 2D drawing using a database of sketch images. Following the traditional imitation and emulation training from art education, DrawFromDrawings enables users to retrieve and refer to a sketch image stored in a database and provides them with various novel strokes as suggestive or deformation feedback. Given regions of interest (ROIs) in the user and reference sketches, DrawFromDrawings detects as-long-as-possible (ALAP) stroke segments and the correspondences between user and reference sketches that are the key to computing seamless interpolations. The stroke-level interpolations are parametrized with the user strokes, the reference strokes, and new strokes created by warping the reference strokes based on the user and reference ROI shapes, and the user study indicated that the interpolation could produce various reasonable strokes varying in shapes and complexity. DrawFromDrawings allows users to either replace their strokes with interpolated strokes (deformation feedback) or overlays interpolated strokes onto their strokes (suggestive feedback). The other user studies on the feedback modes indicated that the suggestive feedback enabled drawers to develop and render their ideas using their own stroke style, whereas the deformation feedback enabled them to finish the sketch composition quickly. Yusuke Matsui 0001, Takaaki Shiratori, Kiyoharu Aizawa |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Text detection in manga by combining connected-component-based and region-based classificationsabstractAs manga (Japanese comics) have become common content in many countries, it is necessary to search manga by text query or translate them automatically. For these applications, we must first extract texts from manga. In this paper, we develop a method to detect text regions in manga. Taking motivation from methods used in scene text detection, we propose an approach using classifiers for both connected components and regions. We have also developed a text region dataset of manga, which enables learning and detailed evaluations of methods used to detect text regions. Experiments using the dataset showed that our text detection method performs more effectively than existing methods. Yuji Aramaki, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 2 |
| 2016 | Interactive region segmentation for mangaabstractManga (Japanese comics) are popular all over the world, and are created digitally. In this paper, we propose an interactive segmentation method tailored for manga. The proposed method enables annotators to select areas in manga efficiently. Our experimental results showed that the proposed framework works better than Adobe Photoshop CC, which is the most widely used commercial image editing software. Kota Ito, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICPR | 2 |
| 2016 | Sketch simplification by classifying strokesabstractIn this paper, we propose a novel approach to creating clean line drawing from a scribbled sketch automatically. The main problem is determining which strokes of a scribbled sketch should be merged. We use a machine learning approach to solve this problem. Our method can automatically generate training data by comparing scribbled sketches with manually drawn line drawings without using annotations. In order to verify the generated training data, we merged strokes and created clean line drawings in accordance with the generated training data. In addition, we trained a support vector machine to estimate the pairs of strokes to be merged. Further, we verified that our method can create line drawings using this estimator. Toru Ogawa, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICPR | 2 |
| 2015 | PQTable: Fast Exact Asymmetric Distance Neighbor Search for Product Quantization Using Hash TablesabstractWe propose the product quantization table (PQTable), a product quantization-based hash table that is fast and requires neither parameter tuning nor training steps. The PQTable produces exactly the same results as a linear PQ search, and is 102 to 105 times faster when tested on the SIFT1B data. In addition, although state-of-the-art performance can be achieved by previous inverted-indexing-based approaches, such methods do require manually designed parameter setting and much training, whereas our method is free from them. Therefore, PQTable offers a practical and useful solution for real-world problems. Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICCV | 1 |
| 2015 | Searching for nearest neighbors with a dense space partitioningabstractProduct quantization based approximate nearest neighbor search with the use of inverted index structures have recently received increasing attention. In this paper, we propose a new inverted index structure for searching nearest neighbors in very large datasets of high dimensional data. For data indexing, our proposed method creates a dense space partitioning using multiple centroids based assigning, which generates shorter candidate lists and improves the search speed. Our experiments with a dataset of one billion SIFT features show that while achieving higher accuracy, our method demonstrates better performances on search speed compared to IV-FADC, the conventional product quantization based inverted index structure. Tuan Anh Nguyen 0004, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 2 |
| 2015 | Challenge for Manga Processing: Sketch-based Manga RetrievalabstractWe propose a sketch-based system for manga image retrieval. In the system, users simply draw sketches, and similar images are retrieved in real time from a manga database. The results are updated every time the user draws a stroke and therefore users can intuitively interact with the system. The proposed method consists of a simple and efficient sliding window-based feature description framework and interactive re-ranking schemes, which are introduced because the characteristics of manga images are different from those of naturalistic images, and thus, traditional image retrieval methods are not effective. Additionally, the future directions of improving the current retrieval system, including by using a combination of text features and image features, and the construction of a large database are discussed. Yusuke Matsui 0001 |
ACM Multimedia | 1 |
| 2015 | Selective K-means Tree SearchabstractIn object recognition and image retrieval, an inverted indexing method is used to solve the approximate nearest neighbor search problem. In these tasks, inverted indexing provides a nonexhaustive solution to large-scale search. However, a problem of previous inverted indexing methods is that a large-scale inverted index is required to achieve a high search recall rate. In this study, we address the problem of reducing the time required to build an inverted index without degrading the search accuracy and speed. Thus, we propose a selective k-means tree search method that combines the power of both hierarchical k-means tree and selective nonexhaustive search. Experiments based on approximate nearest neighbor search using a large dataset comprising one billion SIFT features showed that the hierarchical inverted file based on the selective k-means tree method could be built six times faster, while obtaining almost the same recall and search speed as the state-of-the-art inverted indexing methods. Tuan Anh Nguyen 0004, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 2 |
| 2014 | Sketch2Manga: Sketch-based manga retrievalabstractWe propose a sketch-based method for manga image retrieval, in which users draw sketches via a Web browser that enables the automatic retrieval of similar images from a database of manga titles. The characteristics of manga images are different from those of naturalistic images. Despite the widespread attention given to content-based image retrieval systems, the question of how to retrieve manga images effectively has been little studied. We propose a fine multi-scale edge orientation histogram (FMEOH) whereby a number of differently sized squares on a page can be indexed efficiently. Our experimental results show that FMEOH can achieve greater accuracy than a state-of-the-art sketch-based retrieval method [1]. Yusuke Matsui 0001, Kiyoharu Aizawa, Yushi Jing |
ICIP | 1 |