Ryota Hinami

dblp:180/0282 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
3since 2021 · last 2024
0000-0003-1542-2612ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Machine translation · 30% Image recognition and object detection · 28% Video understanding and tracking · 18%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 96% Query processing and optimization · 4%
Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 90% Image and video processing · 10%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval › multimedia analysis
comic analysis
0.812024
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion · ACM Multimedia 2024
Computer vision › Image recognition and object detection
object detection
0.622018
Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation · EMNLP 2018
Large-Scale R-CNN with Classifier Adaptive Quantization · ECCV (3) 2016
Natural language and speech › Machine translation › neural machine translation
context-aware neural machine translation
0.512021
Towards Fully Automated Manga Translation · AAAI 2021
Natural language and speech › Machine translation
multimodal machine translation
0.512021
Towards Fully Automated Manga Translation · AAAI 2021
Information retrieval › reranking
diffusion-based re-ranking
0.412019
Efficient Image Retrieval via Decoupling Diffusion into Online and Offline Processing · AAAI 2019
Information retrieval
efficient retrieval
0.412019
Efficient Image Retrieval via Decoupling Diffusion into Online and Offline Processing · AAAI 2019
Information retrieval › retrieval models
ranked retrieval
0.412019
Efficient Image Retrieval via Decoupling Diffusion into Online and Offline Processing · AAAI 2019
Machine learning › Representation and self-supervised learning › contrastive learning › negative sampling
hard negative mining
0.312018
Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation · EMNLP 2018
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
0.312018
Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation · EMNLP 2018
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
0.312018
Reconfigurable Inverted Index · ACM Multimedia 2018
Information retrieval › indexing
inverted index
0.312018
Reconfigurable Inverted Index · ACM Multimedia 2018
Information retrieval › similarity search
nearest neighbor search
0.312018
Reconfigurable Inverted Index · ACM Multimedia 2018
Computer vision › Video understanding and tracking
video anomaly detection
0.312017
Joint Detection and Recounting of Abnormal Events by Learning Deep Generic Knowledge · ICCV 2017
Multimedia analysis and retrieval
image retrieval
0.312017
Region-Based Image Retrieval Revisited · ACM Multimedia 2017
Multimedia analysis and retrieval › image retrieval › content-based image retrieval
region-based image retrieval
0.312017
Region-Based Image Retrieval Revisited · ACM Multimedia 2017
Machine learning › Efficient and distributed learning
model compression
0.212016
Large-Scale R-CNN with Classifier Adaptive Quantization · ECCV (3) 2016
Computer vision › Vision and language
multimodal fusion
0.212024
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion · ACM Multimedia 2024
Image and video processing › document image analysis
comic image analysis
0.112021
Towards Fully Automated Manga Translation · AAAI 2021
Query processing and optimization
subset query
0.112018
Reconfigurable Inverted Index · ACM Multimedia 2018

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 1.5multimodal context-aware translation · 1.0automatic corpus construction · 1.0late truncation · 0.4kNN search · 0.4diffusion · 0.4text embedding · 0.3product quantization · 0.3negative phrase augmentation · 0.3IVFADC · 0.3Faster R-CNN · 0.3re-ranking · 0.3multitask CNN · 0.3multi-task learning · 0.3inverted indexing · 0.3deep learning · 0.3convolutional neural network · 0.3quantization · 0.2
YearPublicationVenuePosition
2024 Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
Yingxuan Li, Ryota Hinami, Kiyoharu Aizawa, Yusuke Matsui 0001
ACM Multimedia2
2021 Towards Fully Automated Manga Translation
abstract
We tackle the problem of machine translation of manga, Japanese comics. Manga translation involves two important problems in machine translation: context-aware and multimodal translation. Since text and images are mixed up in an unstructured fashion in Manga, obtaining context from the image is essential for manga translation. However, it is still an open problem how to extract context from image and integrate into MT models. In addition, corpus and benchmarks to train and evaluate such model is currently unavailable. In this paper, we make the following four contributions that establishes the foundation of manga translation research. First, we propose multimodal context-aware translation framework. We are the first to incorporate context information obtained from manga image. It enables us to translate texts in speech bubbles that cannot be translated without using context information (e.g., texts in other speech bubbles, gender of speakers, etc.). Second, for training the model, we propose the approach to automatic corpus construction from pairs of original manga and their translations, by which large parallel corpus can be constructed without any manual labeling. Third, we created a new benchmark to evaluate manga translation. Finally, on top of our proposed methods, we devised a first compleheisive system for fully automated manga translation.
Ryota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke Matsui 0001
AAAI1
2021 Painting Style-Aware Manga Colorization Based On Generative Adversarial Networks
abstract
Japanese comics (called manga) are traditionally created in monochrome format. In recent years, in addition to monochrome comics, full color comics, a more attractive medium, have appeared. Unfortunately, color comics require manual colorization, which incurs high labor costs. Although automatic colorization methods have been recently proposed, most of them are designed for illustrations, not for comics. Unlike illustrations, since comics are composed of many consecutive images, the painting style must be consistent. To realize consistent colorization, we propose here a semi-automatic colorization method based on generative adversarial networks (GAN); the method learns the painting style of a specific comic from small amount of training data. The proposed method takes a pair of a screen tone image and a flat colored image as input, and outputs a colorized image. Experiments show that the proposed method achieves better performance than the existing alternatives.
Yugo Shimizu, Ryosuke Furuta, Delong Ouyang, Yukinobu Taniguchi, Ryota Hinami, Shonosuke Ishiwatari
ICIP5
2019 Efficient Image Retrieval via Decoupling Diffusion into Online and Offline Processing
abstract
Diffusion is commonly used as a ranking or re-ranking method in retrieval tasks to achieve higher retrieval performance, and has attracted lots of attention in recent years. A downside to diffusion is that it performs slowly in comparison to the naive k-NN search, which causes a non-trivial online computational cost on large datasets. To overcome this weakness, we propose a novel diffusion technique in this paper. In our work, instead of applying diffusion to the query, we precompute the diffusion results of each element in the database, making the online search a simple linear combination on top of the k-NN search process. Our proposed method becomes 10∼ times faster in terms of online search speed. Moreover, we propose to use late truncation instead of early truncation in previous works to achieve better retrieval performance.
Fan Yang 0038, Ryota Hinami, Yusuke Matsui 0001, Steven Ly, Shin'ichi Satoh 0001
AAAI2
2018 Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation
abstract
Thanks to the success of object detection technology, we can retrieve objects of the specified classes even from huge image collections.However, the current state-of-the-art object detectors (such as Faster R-CNN) can only handle pre-specified classes.In addition, large amounts of positive and negative visual samples are required for training.In this paper, we address the problem of open-vocabulary object retrieval and localization, where the target object is specified by a textual query (e.g., a word or phrase).We first propose Query-Adaptive R-CNN, a simple extension of Faster R-CNN adapted to open-vocabulary queries, by transforming the text embedding vector into an object classifier and localization regressor.Then, for discriminative training, we then propose negative phrase augmentation (NPA) to mine hard negative samples which are visually similar to the query and at the same time semantically mutually exclusive of the query.The proposed method can retrieve and localize objects specified by a textual query from one million images in only 0.5 seconds with high precision.
Ryota Hinami, Shin'ichi Satoh 0001
EMNLP1
2018 Tracked Instance Search
abstract
In this work we propose tracking as a generic addition to the instance search task. From video data perspective, much information that can be used is not taken into account in the traditional instance search approach. This work aims to provide insights on exploiting such existing information by means of tracking and the proper combination of the results, independently of the instance search system. We also present a study on the improvement of the system when using multiple independent instances (up to 4) of the same person. Experimental results show that our system improves substantially its performance when using tracking. Best configuration improves from mAP = 0.447 to mAP = 0.511 for a single example, and from mAP = 0.647 to mAP = 0.704 for multiple (4) given examples.
Andreu Girbau-Xalabarder, Ryota Hinami, Shin'ichi Satoh 0001
ICASSP2
2018 Reconfigurable Inverted Index
abstract
Existing approximate nearest neighbor search systems suffer from two fundamental problems that are of practical importance but have not received sufficient attention from the research community. First, although existing systems perform well for the whole database, it is difficult to run a search over a subset of the database. Second, there has been no discussion concerning the performance decrement after many items have been newly added to a system. We develop a reconfigurable inverted index (Rii) to resolve these two issues. Based on the standard IVFADC system, we design a data layout such that items are stored linearly. This enables us to efficiently run a subset search by switching the search method to a linear PQ scan if the size of a subset is small. Owing to the linear layout, the data structure can be dynamically adjusted after new items are added, maintaining the fast speed of the system. Extensive comparisons show that Rii achieves a comparable performance with state-of-the art systems such as Faiss.
Yusuke Matsui 0001, Ryota Hinami, Shin'ichi Satoh 0001
ACM Multimedia2
2017 Joint Detection and Recounting of Abnormal Events by Learning Deep Generic Knowledge
abstract
This paper addresses the problem of joint detection and recounting of abnormal events in videos. Recounting of abnormal events, i.e., explaining why they are judged to be abnormal, is an unexplored but critical task in video surveillance, because it helps human observers quickly judge if they are false alarms or not. To describe the events in the human-understandable form for event recounting, learning generic knowledge about visual concepts (e.g., object and action) is crucial. Although convolutional neural networks (CNNs) have achieved promising results in learning such concepts, it remains an open question as to how to effectively use CNNs for abnormal event detection, mainly due to the environment-dependent nature of the anomaly detection. In this paper, we tackle this problem by integrating a generic CNN model and environment-dependent anomaly detectors. Our approach first learns CNN with multiple visual tasks to exploit semantic information that is useful for detecting and recounting abnormal events. By appropriately plugging the model into anomaly detectors, we can detect and recount abnormal events while taking advantage of the discriminative power of CNNs. Our approach outperforms the state-of-the-art on Avenue and UCSD Ped2 benchmarks for abnormal event detection and also produces promising results of abnormal event recounting.
Ryota Hinami, Tao Mei 0001, Shin'ichi Satoh 0001
ICCV1
2017 Region-Based Image Retrieval Revisited
abstract
Region-based image retrieval (RBIR) technique is revisited. In early attempts at RBIR in the late 90s, researchers found many ways to specify region-based queries and spatial relationships; however, the way to characterize the regions, such as by using color histograms, were very poor at that time. Here, we revisit RBIR by incorporating semantic specification of objects and intuitive specification of spatial relationships. Our contributions are the following. First, to support multiple aspects of semantic object specification (category, instance, and attribute), we propose a multitask CNN feature that allows us to use deep learning technique and to jointly handle multi-aspect object specification. Second, to help users specify spatial relationships among objects in an intuitive way, we propose recommendation techniques of spatial relationships. In particular, by mining the search results, a system can recommend feasible spatial relationships among the objects. The system also can recommend likely spatial relationships by assigned object category names based on language prior. Moreover, object-level inverted indexing supports very fast shortlist generation, and re-ranking based on spatial constraints provides users with instant RBIR experiences.
Ryota Hinami, Yusuke Matsui 0001, Shin'ichi Satoh 0001
ACM Multimedia1
2016 Large-Scale R-CNN with Classifier Adaptive Quantization
Ryota Hinami, Shin'ichi Satoh 0001
ECCV (3)1
2016 Audience Behavior Mining by Integrating TV Ratings with Multimedia Contents
abstract
TV ratings are a widely used indicator in the TV broadcasting field. While TV ratings are mainly used in advertising, they can also be used as a social sensor that reflects the interests of people. This paper presents a framework for discovering audience behavior through the mining of TV ratings. We have established a framework that enables discovery of numerous patterns of audience behavior from TV ratings. Used along with other multimedia contents such as video and text, it enables various types of knowledge to be semi-automatically found, such as what types of news programs are of most interest and what are the key visual features for acquiring high TV ratings. The discovery of audience behavior is achieved by focusing on the change points in the rating data, i.e., the points in time where many people switch the channel or turn the television on or off. Rich descriptions that characterize these points are extracted from multimedia contents, and then various filtering techniques are used to extract specific patterns of interest. Several applications of this framework for discovering knowledge demonstrated that it can effectively extract various types of audience behavior. To the best of our knowledge, this work is the first work to analyze the use of ratings data in combination with video and other multimedia data.
Ryota Hinami, Shin'ichi Satoh 0001
ISM1
2016 Bidirectional extraction and recognition of scene text with layout consistency
Ryota Hinami, Xinhao Liu 0002, Naoki Chiba, Shin'ichi Satoh 0001
Int. J. Document Anal. Recognit.1