EDBT 2026 Demo / reviewers in the wild / expert
Ryota Hinami
dblp:180/0282
· DBLP profile ↗
12ranked-venue papers
7as first author
3since 2021 · last 2024
0000-0003-1542-2612ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Machine translation · 30% Image recognition and object detection · 28% Video understanding and tracking · 18% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 96% Query processing and optimization · 4% | |
| Computer graphics and multimedia
3 papers |
Multimedia analysis and retrieval · 90% Image and video processing · 10% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Multimedia analysis and retrieval › multimedia analysis
comic analysis |
0.8 | 1 | 2024 | Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion · ACM Multimedia 2024 |
Computer vision › Image recognition and object detection
object detection |
0.6 | 2 | 2018 | Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation · EMNLP 2018 Large-Scale R-CNN with Classifier Adaptive Quantization · ECCV (3) 2016 |
Natural language and speech › Machine translation › neural machine translation
context-aware neural machine translation |
0.5 | 1 | 2021 | Towards Fully Automated Manga Translation · AAAI 2021 |
Natural language and speech › Machine translation
multimodal machine translation |
0.5 | 1 | 2021 | Towards Fully Automated Manga Translation · AAAI 2021 |
Information retrieval › reranking
diffusion-based re-ranking |
0.4 | 1 | 2019 | Efficient Image Retrieval via Decoupling Diffusion into Online and Offline Processing · AAAI 2019 |
Information retrieval
efficient retrieval |
0.4 | 1 | 2019 | Efficient Image Retrieval via Decoupling Diffusion into Online and Offline Processing · AAAI 2019 |
Information retrieval › retrieval models
ranked retrieval |
0.4 | 1 | 2019 | Efficient Image Retrieval via Decoupling Diffusion into Online and Offline Processing · AAAI 2019 |
Machine learning › Representation and self-supervised learning › contrastive learning › negative sampling
hard negative mining |
0.3 | 1 | 2018 | Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation · EMNLP 2018 |
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection |
0.3 | 1 | 2018 | Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation · EMNLP 2018 |
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search |
0.3 | 1 | 2018 | Reconfigurable Inverted Index · ACM Multimedia 2018 |
Information retrieval › indexing
inverted index |
0.3 | 1 | 2018 | Reconfigurable Inverted Index · ACM Multimedia 2018 |
Information retrieval › similarity search
nearest neighbor search |
0.3 | 1 | 2018 | Reconfigurable Inverted Index · ACM Multimedia 2018 |
Computer vision › Video understanding and tracking
video anomaly detection |
0.3 | 1 | 2017 | Joint Detection and Recounting of Abnormal Events by Learning Deep Generic Knowledge · ICCV 2017 |
Multimedia analysis and retrieval
image retrieval |
0.3 | 1 | 2017 | Region-Based Image Retrieval Revisited · ACM Multimedia 2017 |
Multimedia analysis and retrieval › image retrieval › content-based image retrieval
region-based image retrieval |
0.3 | 1 | 2017 | Region-Based Image Retrieval Revisited · ACM Multimedia 2017 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2016 | Large-Scale R-CNN with Classifier Adaptive Quantization · ECCV (3) 2016 |
Computer vision › Vision and language
multimodal fusion |
0.2 | 1 | 2024 | Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion · ACM Multimedia 2024 |
Image and video processing › document image analysis
comic image analysis |
0.1 | 1 | 2021 | Towards Fully Automated Manga Translation · AAAI 2021 |
Query processing and optimization
subset query |
0.1 | 1 | 2018 | Reconfigurable Inverted Index · ACM Multimedia 2018 |
Methods — techniques the papers use, named apart from their topics
multimodal fusion · 1.5multimodal context-aware translation · 1.0automatic corpus construction · 1.0late truncation · 0.4kNN search · 0.4diffusion · 0.4text embedding · 0.3product quantization · 0.3negative phrase augmentation · 0.3IVFADC · 0.3Faster R-CNN · 0.3re-ranking · 0.3multitask CNN · 0.3multi-task learning · 0.3inverted indexing · 0.3deep learning · 0.3convolutional neural network · 0.3quantization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
Yingxuan Li, Ryota Hinami, Kiyoharu Aizawa, Yusuke Matsui 0001 |
ACM Multimedia | 2 |
| 2021 | Towards Fully Automated Manga TranslationabstractWe tackle the problem of machine translation of manga, Japanese comics. Manga translation involves two important problems in machine translation: context-aware and multimodal translation. Since text and images are mixed up in an unstructured fashion in Manga, obtaining context from the image is essential for manga translation. However, it is still an open problem how to extract context from image and integrate into MT models. In addition, corpus and benchmarks to train and evaluate such model is currently unavailable. In this paper, we make the following four contributions that establishes the foundation of manga translation research. First, we propose multimodal context-aware translation framework. We are the first to incorporate context information obtained from manga image. It enables us to translate texts in speech bubbles that cannot be translated without using context information (e.g., texts in other speech bubbles, gender of speakers, etc.). Second, for training the model, we propose the approach to automatic corpus construction from pairs of original manga and their translations, by which large parallel corpus can be constructed without any manual labeling. Third, we created a new benchmark to evaluate manga translation. Finally, on top of our proposed methods, we devised a first compleheisive system for fully automated manga translation. Ryota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke Matsui 0001 |
AAAI | 1 |
| 2021 | Painting Style-Aware Manga Colorization Based On Generative Adversarial NetworksabstractJapanese comics (called manga) are traditionally created in monochrome format. In recent years, in addition to monochrome comics, full color comics, a more attractive medium, have appeared. Unfortunately, color comics require manual colorization, which incurs high labor costs. Although automatic colorization methods have been recently proposed, most of them are designed for illustrations, not for comics. Unlike illustrations, since comics are composed of many consecutive images, the painting style must be consistent. To realize consistent colorization, we propose here a semi-automatic colorization method based on generative adversarial networks (GAN); the method learns the painting style of a specific comic from small amount of training data. The proposed method takes a pair of a screen tone image and a flat colored image as input, and outputs a colorized image. Experiments show that the proposed method achieves better performance than the existing alternatives. Yugo Shimizu, Ryosuke Furuta, Delong Ouyang, Yukinobu Taniguchi, Ryota Hinami, Shonosuke Ishiwatari |
ICIP | 5 |
| 2019 | Efficient Image Retrieval via Decoupling Diffusion into Online and Offline ProcessingabstractDiffusion is commonly used as a ranking or re-ranking method in retrieval tasks to achieve higher retrieval performance, and has attracted lots of attention in recent years. A downside to diffusion is that it performs slowly in comparison to the naive k-NN search, which causes a non-trivial online computational cost on large datasets. To overcome this weakness, we propose a novel diffusion technique in this paper. In our work, instead of applying diffusion to the query, we precompute the diffusion results of each element in the database, making the online search a simple linear combination on top of the k-NN search process. Our proposed method becomes 10∼ times faster in terms of online search speed. Moreover, we propose to use late truncation instead of early truncation in previous works to achieve better retrieval performance. Fan Yang 0038, Ryota Hinami, Yusuke Matsui 0001, Steven Ly, Shin'ichi Satoh 0001 |
AAAI | 2 |
| 2018 | Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase AugmentationabstractThanks to the success of object detection technology, we can retrieve objects of the specified classes even from huge image collections.However, the current state-of-the-art object detectors (such as Faster R-CNN) can only handle pre-specified classes.In addition, large amounts of positive and negative visual samples are required for training.In this paper, we address the problem of open-vocabulary object retrieval and localization, where the target object is specified by a textual query (e.g., a word or phrase).We first propose Query-Adaptive R-CNN, a simple extension of Faster R-CNN adapted to open-vocabulary queries, by transforming the text embedding vector into an object classifier and localization regressor.Then, for discriminative training, we then propose negative phrase augmentation (NPA) to mine hard negative samples which are visually similar to the query and at the same time semantically mutually exclusive of the query.The proposed method can retrieve and localize objects specified by a textual query from one million images in only 0.5 seconds with high precision. Ryota Hinami, Shin'ichi Satoh 0001 |
EMNLP | 1 |
| 2018 | Tracked Instance SearchabstractIn this work we propose tracking as a generic addition to the instance search task. From video data perspective, much information that can be used is not taken into account in the traditional instance search approach. This work aims to provide insights on exploiting such existing information by means of tracking and the proper combination of the results, independently of the instance search system. We also present a study on the improvement of the system when using multiple independent instances (up to 4) of the same person. Experimental results show that our system improves substantially its performance when using tracking. Best configuration improves from mAP = 0.447 to mAP = 0.511 for a single example, and from mAP = 0.647 to mAP = 0.704 for multiple (4) given examples. Andreu Girbau-Xalabarder, Ryota Hinami, Shin'ichi Satoh 0001 |
ICASSP | 2 |
| 2018 | Reconfigurable Inverted IndexabstractExisting approximate nearest neighbor search systems suffer from two fundamental problems that are of practical importance but have not received sufficient attention from the research community. First, although existing systems perform well for the whole database, it is difficult to run a search over a subset of the database. Second, there has been no discussion concerning the performance decrement after many items have been newly added to a system. We develop a reconfigurable inverted index (Rii) to resolve these two issues. Based on the standard IVFADC system, we design a data layout such that items are stored linearly. This enables us to efficiently run a subset search by switching the search method to a linear PQ scan if the size of a subset is small. Owing to the linear layout, the data structure can be dynamically adjusted after new items are added, maintaining the fast speed of the system. Extensive comparisons show that Rii achieves a comparable performance with state-of-the art systems such as Faiss. Yusuke Matsui 0001, Ryota Hinami, Shin'ichi Satoh 0001 |
ACM Multimedia | 2 |
| 2017 | Joint Detection and Recounting of Abnormal Events by Learning Deep Generic KnowledgeabstractThis paper addresses the problem of joint detection and recounting of abnormal events in videos. Recounting of abnormal events, i.e., explaining why they are judged to be abnormal, is an unexplored but critical task in video surveillance, because it helps human observers quickly judge if they are false alarms or not. To describe the events in the human-understandable form for event recounting, learning generic knowledge about visual concepts (e.g., object and action) is crucial. Although convolutional neural networks (CNNs) have achieved promising results in learning such concepts, it remains an open question as to how to effectively use CNNs for abnormal event detection, mainly due to the environment-dependent nature of the anomaly detection. In this paper, we tackle this problem by integrating a generic CNN model and environment-dependent anomaly detectors. Our approach first learns CNN with multiple visual tasks to exploit semantic information that is useful for detecting and recounting abnormal events. By appropriately plugging the model into anomaly detectors, we can detect and recount abnormal events while taking advantage of the discriminative power of CNNs. Our approach outperforms the state-of-the-art on Avenue and UCSD Ped2 benchmarks for abnormal event detection and also produces promising results of abnormal event recounting. Ryota Hinami, Tao Mei 0001, Shin'ichi Satoh 0001 |
ICCV | 1 |
| 2017 | Region-Based Image Retrieval RevisitedabstractRegion-based image retrieval (RBIR) technique is revisited. In early attempts at RBIR in the late 90s, researchers found many ways to specify region-based queries and spatial relationships; however, the way to characterize the regions, such as by using color histograms, were very poor at that time. Here, we revisit RBIR by incorporating semantic specification of objects and intuitive specification of spatial relationships. Our contributions are the following. First, to support multiple aspects of semantic object specification (category, instance, and attribute), we propose a multitask CNN feature that allows us to use deep learning technique and to jointly handle multi-aspect object specification. Second, to help users specify spatial relationships among objects in an intuitive way, we propose recommendation techniques of spatial relationships. In particular, by mining the search results, a system can recommend feasible spatial relationships among the objects. The system also can recommend likely spatial relationships by assigned object category names based on language prior. Moreover, object-level inverted indexing supports very fast shortlist generation, and re-ranking based on spatial constraints provides users with instant RBIR experiences. Ryota Hinami, Yusuke Matsui 0001, Shin'ichi Satoh 0001 |
ACM Multimedia | 1 |
| 2016 | Large-Scale R-CNN with Classifier Adaptive Quantization
Ryota Hinami, Shin'ichi Satoh 0001 |
ECCV (3) | 1 |
| 2016 | Audience Behavior Mining by Integrating TV Ratings with Multimedia ContentsabstractTV ratings are a widely used indicator in the TV broadcasting field. While TV ratings are mainly used in advertising, they can also be used as a social sensor that reflects the interests of people. This paper presents a framework for discovering audience behavior through the mining of TV ratings. We have established a framework that enables discovery of numerous patterns of audience behavior from TV ratings. Used along with other multimedia contents such as video and text, it enables various types of knowledge to be semi-automatically found, such as what types of news programs are of most interest and what are the key visual features for acquiring high TV ratings. The discovery of audience behavior is achieved by focusing on the change points in the rating data, i.e., the points in time where many people switch the channel or turn the television on or off. Rich descriptions that characterize these points are extracted from multimedia contents, and then various filtering techniques are used to extract specific patterns of interest. Several applications of this framework for discovering knowledge demonstrated that it can effectively extract various types of audience behavior. To the best of our knowledge, this work is the first work to analyze the use of ratings data in combination with video and other multimedia data. Ryota Hinami, Shin'ichi Satoh 0001 |
ISM | 1 |
| 2016 | Bidirectional extraction and recognition of scene text with layout consistency
Ryota Hinami, Xinhao Liu 0002, Naoki Chiba, Shin'ichi Satoh 0001 |
Int. J. Document Anal. Recognit. | 1 |