VLDB 2026 Research / reviewers in the wild / expert
Keiji Yanai
dblp:60/2410
· DBLP profile ↗
97ranked-venue papers
21as first author
28since 2021 · last 2026
0000-0002-0431-183XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 83 · 14 first-author · 26 since 2021Artificial intelligence and machine learning · 26 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 19 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
Junwen Chen 0003, Peilin Xiong, Keiji Yanai |
ICPR (8) | 3 |
| 2026 | SceneTextStylizer: a training-free diffusion framework for scene text style transfer
Honghui Yuan, Keiji Yanai |
Multim. Syst. | 2 |
| 2025 | Japanese Kuzushiji Font Generation Employing Differentiable Renderer
Honghui Yuan, Junwen Chen 0003, Keiji Yanai |
ICDAR (5) | 3 |
| 2025 | KuzushijiGen: A Real-Time Few-Shot Japanese Kuzushiji Generator via Differentiable RenderingabstractKuzushiji is an ancient form of Japanese script that has deteriorated over time, making many characters unrecognizable. Recent font generation methods have mainly focused on modern font generation using large-scale raster images, making it difficult to apply them to ancient handwritten scripts like Kuzushiji. To address this challenge, we propose a web-based real-time application for generating Kuzushiji fonts using vector images without requiring any model training, which incorporates a differentiable rasterizer, a conditional diffusion model, a discriminator, and a Multi-Stroke Encoder for stroke-level text control. Multiple loss functions are employed to optimize font parameters. This application allows users to freely input modern Japanese characters to generate corresponding Kuzushiji images. Our results are resolution-independent and are more effective for preserving historical text. Supplementary slides: http://bit.ly/41YOyug Honghui Yuan, Keiji Yanai |
MMAsia | 2 |
| 2025 | WaveFontStyler: Font Style Transfer Based on Sound
Kota Izumi, Keiji Yanai |
MMM (5) | 2 |
| 2025 | CalorieVoL: Integrating Volumetric Context Into Multimodal Large Language Models for Image-Based Calorie Estimation
Hikaru Tanabe, Keiji Yanai |
MMM (4) | 2 |
| 2025 | KuzushijiDiffuser: Japanese Kuzushiji Font Generation with FontDiffuser
Honghui Yuan, Keiji Yanai |
MMM (2) | 2 |
| 2025 | KuzushijiFontDiff: Diffusion Model for Japanese Kuzushiji Font Generation
Honghui Yuan, Keiji Yanai |
MMM (5) | 2 |
| 2025 | SceneTextStyler: Editing Text with Style Transformation
Honghui Yuan, Keiji Yanai |
MMM (5) | 2 |
| 2025 | Focusing on what to Decode and what to Train: SOV Decoding with Specific Target Guided DeNoising and Vision Language AdvisorabstractRecent transformer-based methods achieve notable gains in the Human-object Interaction Detection (HOID) task by leveraging the detection of DETR and the prior knowledge of Vision-Language Model (VLM). However, these methods suffer from extended training times and complex optimization due to the entanglement of object detection and HOI recognition during the decoding process. Especially, the query embeddings used to predict both labels and boxes suffer from ambiguous representations, and the gap between the prediction of HOI labels and verb labels is not considered. To address these challenges, we introduce SOV-STG-VLA with three key components: Subject-Object-Verb (SOV) decoding, Specific Target Guided (STG) denoising, and a Vision-Language Advisor (VLA). Our SOV decoders disentangle object detection and verb recognition with a novel interaction region representation. The STG denoising strategy learns label embeddings with ground-truth information to guide the training and inference. Our SOV-STG achieves a fast convergence speed and high accuracy and builds a foundation for the VLA to incorporate the prior knowledge of the VLM. We introduce a vision advisor decoder to fuse both the interaction region information and the VLM's vision knowledge and a Verb-HOI prediction bridge to promote interaction representation learning. Our VLA notably improves our SOV-STG and achieves SOTA performance with one-sixth of training epochs compared to recent SOTA. Code and models are available at https://github.com/cjw2021/SOV-STG-VLA. Junwen Chen 0003, Yingcheng Wang, Keiji Yanai |
WACV | 3 |
| 2024 | Act-ChatGPT: Introducing Action Features into Multi-modal Large Language Models for Video Understanding
Yuto Nakamizo, Keiji Yanai |
ICPR (23) | 2 |
| 2024 | Font Style Translation in Scene Text Images with CLIPstyler
Honghui Yuan, Keiji Yanai |
ICPR (19) | 2 |
| 2024 | Training-Free Region Prediction with Stable Diffusion
Yuma Honbu, Keiji Yanai |
MMM (4) | 2 |
| 2023 | CalorieCam360: Simultaneous Eating Action Recognition of Multiple People Using an Omnidirectional CameraabstractIn recent years, as people become more health-conscious, dietary management has become increasingly important. Existing methods record only one person’s meals or eating movements, but cannot record the meals of multiple people at the same time. Therefore, we aim to record the meals of all people around a dining table using an omnidirectional camera simultaneously. Kento Terauchi, Keiji Yanai |
ICMR | 2 |
| 2023 | MADiMa '23: 8th International Workshop on Multimedia Assisted Dietary ManagementabstractThis abstract provides a summary and overview of the 8th International Workshop on Multimedia Assisted Dietary Management. Stavroula G. Mougiakakou, Keiji Yanai, Dario Allegra |
ACM Multimedia | 2 |
| 2023 | Mask-based Food Image Synthesis with Cross-Modal Recipe EmbeddingsabstractIn this paper, we propose a Mask-based Recipe Embedding GAN (MRE-GAN), which enables us to generate a realistic food image based on a given mask image containing single or multiple food regions with cross-modal recipe embeddings for each food region. Thus, we can change meal shapes by modifying mask images, while by editing recipe text, we can change meal appearance. Our experimental findings confirmed that the proposed method could generate higher quality food images than the baselines, and we could change meal shapes and appearances by editing mask images and recipe texts as we liked. Zhongtao Chen, Yuma Honbu, Keiji Yanai |
MMAsia | 3 |
| 2023 | VQ-VDM: Video Diffusion Models with 3D VQGANabstractIn recent years, deep generative models have achieved impressive performance such as realizing image generation that is indistinguishable from real images. Particularly, Latent Diffusion Models, one of the image generation models, have had a significant impact on society. Therefore, video generation is attracting attention as the next modality. However, video generation is more challenging than image generation due to the consideration of temporal consistency and the increase in computational complexity, since a video is a sequence of multiple frames. In this study, we propose a video generation model based on diffusion models employing 3D VQGAN, which is called VQ-VDM. The proposed model is about nine times faster than the Video Diffusion Models which directly generate videos, since our model generates a latent representation which is decoded into a video by a VQGAN decoder. Moreover, our model can generate higher quality video than prior video generation methods exclude state-of-the-art method. Ryota Kaji, Keiji Yanai |
MMAsia | 2 |
| 2023 | Contextual Associated Triplet Queries for Panoptic Scene Graph GenerationabstractThe Panoptic Scene Graph generation (PSG) task aims to extract the triplets composed of subject, object, and relation based on panoptic segmentation. For one-stage methods, PSGTR predicts the subject, object, and relation by one query. However, the integrated query is too implicit to simultaneously ascertain pairs of instances and relations. In PSGFormer, it learns instances and relation queries separately and establishes matches between subject-relation and object-relation pairs by employing the relation as an index. Nevertheless, this method could potentially impede the accurate determination of the optimal match. To address the aforementioned issues, we propose a new one-stage method, Contextual Associated Triplet Queries (CATQ), which employs three branches to decode subject, object, and relation features separately. Additionally, we leverage instance information to guide the relation decoding process. Furthermore, we introduce the triplet context fusion block to enable the extraction of more comprehensive instance pairs and triplet relations. Our proposed method achieves 34.8 Recall@20 and 20.9 mRecall@20 respectively and surpasses the state-of-the-art baseline method by 22.5% and 26.0% with half of the training session. Jingbin Xu, Junwen Chen 0003, Keiji Yanai |
MMAsia | 3 |
| 2023 | Virtual Try-On Considering Temporal Consistency for Videoconferencing
Daiki Shimizu, Keiji Yanai |
MMM (2) | 2 |
| 2023 | Transformer-Based Cross-Modal Recipe Embeddings with Large Batch Training
Jing Yang 0064, Junwen Chen 0003, Keiji Yanai |
MMM (2) | 3 |
| 2022 | StyleGAN-based CLIP-guided Image Shape ManipulationabstractIn this paper, we propose a text-guided image manipulation method which focuses on editing shape attribute using text description. We combine an image generation model, StyleGAN2, and image-text matching model, CLIP, and we have achieved the goal of image shape attribute manipulation by modifying the parameters of the pretrained StyleGAN2 generator. Qualitative and quantitative evaluations are conducted to demonstrate the effectiveness of the proposed method. Yuchen Qian, Keiji Yanai |
CBMI | 3 |
| 2022 | Continual Learning in Vision TransformerabstractContinual learning aims to continuously learn new tasks from new data while retaining the knowledge of tasks learned in the past. Recently, the Vision Transformer, which utilizes the Transformer initially proposed in natural language processing for computer vision, has shown higher accuracy than Convolutional Neural Networks (CNN) in image recognition tasks. However, there are few methods that have achieved continual learning with Vision Transformer. In this paper, we compare and improve continual learning methods that can be applied to both CNN and Vision Transformers. In our experiments, we compare several continual learning methods and their combinations to show the differences in accuracy and the number of parameters. Mana Takeda, Keiji Yanai |
ICIP | 2 |
| 2022 | Unseen Food SegmentationabstractFood image segmentation is important for detailed analysis on food images, especially for classification of multiple food items and calorie amount estimation. However, there is a costly problem in training a semantic segmentation model because it requires a large number of images with pixel-level annotations. In addition, the existence of a myriad of food categories causes the problem of insufficient data in each category. Although several food segmentation datasets such as the UEC-FoodPix Complete has been released so far, the number of food categories is still limited to a small number. Yuma Honbu, Keiji Yanai |
ICMR | 2 |
| 2022 | MADiMa'22: 7th International Workshop on Multimedia Assisted Dietary ManagementabstractThis abstract provides a summary and overview of the 7th International Workshop on Multimedia Assisted Dietary Management. Stavroula G. Mougiakakou, Giovanni Maria Farinella, Keiji Yanai, Dario Allegra |
ACM Multimedia | 3 |
| 2022 | Parallel Queries for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) Detection requires localizing a pair of humans and objects. Recent transformer-based methods leverage the query embeddings to represent the entire HOI instances. The target embeddings after decoding are used to represent the object and human characteristics at the same time. However, it is ambiguous to use the highly integrated embeddings to localize the human and object simultaneously. To address this problem, we split the detection decoding process into subject decoding and object decoding to detect the humans and objects in parallel. Our proposed method, Parallel Query Network (PQNet) uses two transformer decoders to decode the subject embeddings and object embeddings in parallel, and a novel verb decoder is used to fuse the representation from the detection decoding and predict the interaction. The attention mechanisms in the verb decoder consist of the attention between human and object embeddings and the attention between the fused embeddings and global semantic features. As the transformer architecture maintains the permutation of the input query embeddings, the paired boxes of humans and objects are directly predicted by feed-forward networks. With the full usage of the object detection part, our proposed architecture outperforms the state-of-the-art baseline method with half of the training epochs. Junwen Chen 0003, Keiji Yanai |
MMAsia | 2 |
| 2022 | Zero-Shot Font Style Transfer with a Differentiable RendererabstractRecently, a large-scale language-image multi-modal model, CLIP, has been used to realize language-based image translation in a zero-shot manner without training. In this study, we attempted to generate language-based decorative fonts for font images using CLIP. By the existing image style transfer methods using CLIP, stylized font images are usually only surrounded by decorations, and the characters themselves do not change significantly. On the other hand, in this study, we use CLIP and vector graphics image representation using a differentiable renderer to achieve a style transfer of text images that matches the input text. The experimental results show that the proposed method transfers the style of font images to match the given texts. In addition to text images, we confirmed that the proposed method was also able to transform the style of simple logo patterns based on the given texts. Kota Izumi, Keiji Yanai |
MMAsia | 2 |
| 2022 | FASSD-Net: Fast and Accurate Real-Time Semantic Segmentation for Embedded SystemsabstractRecent works of real-time semantic segmentation, remove or make use of light decoders from dense deep neural networks to achieve fast inference speed. This strategy helps to achieve real-time performance; however, the accuracy is significantly compromised in comparison to non-real-time methods. In this paper, we introduce two key modules aimed to design a high-performance decoder for real-time semantic segmentation, which also reduces the accuracy gap between real-time and non-real-time networks. The first module, Dilated Asymmetric Pyramidal Fusion (DAPF), is designed to increase the receptive field on the top of the last stage of the encoder, obtaining richer contextual features. The second module, Multi-resolution Dilated Asymmetric (MDA) module, fuses and refines detail and contextual information from multi-scale feature maps coming from early and deeper stages of the network. Both modules are designed to keep a low computational complexity by using asymmetric convolutions. With these modules, we propose a network entitled “FASSD-Net,” which is based on a light-weight CNN backbone. Running on a single Nvidia GTX 1080Ti, our model reaches 77.5% and 69.3% of mIoU, at 41 and 80 FPS on the Cityscapes and CamVid datasets, respectively. We present an extensive analysis of the accuracy-speed tradeoffs of three FASSD-Net variations on different embedded systems, demonstrating that a light version of our network can run on the low-power consumption Jetson Xavier NX, at 32 FPS reaching 74% of mIoU with full resolution ($1024\times 2048$). The source code and pre-trained models are available at github.com/GibranBenitez/FASSD-Net. Leonel Rosas-Arias, Gibran Benitez-Garcia, José Portillo-Portillo, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Keiji Yanai |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Cross-Modal Recipe Embeddings by Disentangling Recipe Contents and Dish StylesabstractNowadays, cooking recipe sharing sites on the Web are widely used, and play a major role in everyday home cooking. Since cooking recipes consist of dish photos and recipe texts, cross-modal recipe search is being actively explored. To enable cross-modal search, both food image features and cooking text recipe features are embedded into the same shared space in general. However, in most of the existing studies, a one-to-one correspondence between a recipe text and a dish image in the embedding space is assumed, although an unlimited number of photos with different serving styles and different plates can be associated with the same recipe. In this paper, we propose a RDE-GAN (Recipe Disentangled Embedding GAN) which separates food image information into a recipe image feature and a non-recipe shape feature. In addition, we generate a food image by integrating both the recipe embedding and a shape feature. Since the proposed embedding is free from serving and plate styles which are unrelated to cooking recipes, the experimental results showed that it outperformed the existing methods on cross-modal recipe search. We also confirmed that only either shape or recipe elements can be changed at the time of food image generation. Yu Sugiyama, Keiji Yanai |
ACM Multimedia | 2 |
| 2020 | Weakly-Supervised Plate And Food Region SegmentationabstractIn this paper, we propose a novel method to infer plate regions of food images without any pixel-wise annotation. We synthesize plate segmentation masks using difference of visualization in food image classifiers. To be concrete, we use two types of classifiers: a food category classifier and a food/non-food classifier. Using the Class Activation Mapping (CAM) which is one of the basic visualization techniques of CNNs, a food category classifier can highlight food regions containing no plate regions, while a food/non-food category classifier can highlight food regions including plate regions. By taking advantage of the difference between the food regions estimated by visualization of two kinds of the classifiers, in this paper, we demonstrate that we can estimate plate regions without any pixel-wise annotation, and we proposed the approach for boosting the accuracy of weakly-supervised food segmentation using the plate segmentation. In experiments, we show the effectiveness of the proposed approach by evaluating and comparing the accuracy of the weakly-supervised segmentation. The proposed approaches certainly improved an image-level weakly-supervised segmentation method in the food domain and outperformed a well-known bounding box-level weakly-supervised segmentation method. Wataru Shimoda, Keiji Yanai |
ICME | 2 |
| 2020 | IPN Hand: A Video Dataset and Benchmark for Real-Time Continuous Hand Gesture RecognitionabstractContinuous hand gesture recognition (HGR) is an essential part of human-computer interaction with a wide range of applications in the automotive sector, consumer electronics, home automation, and others. In recent years, accurate and efficient deep learning models have been proposed for HGR. However, in the research community, the current publicly available datasets lack real-world elements needed to build responsive and efficient HGR systems. In this paper, we introduce a new benchmark dataset named IPN Hand with sufficient size, variety, and real-world elements able to train and evaluate deep neural networks. This dataset contains more than 4,000 gesture samples and 800,000 RGB frames from 50 distinct subjects. We design 13 different static and dynamic gestures focused on interaction with touchless screens. We especially consider the scenario when continuous gestures are performed without transition states, and when subjects perform natural movements with their hands as non-gesture actions. Gestures were collected from about 30 diverse scenes, with real-world variation in background and illumination. With our dataset, the performance of three 3D-CNN models is evaluated on the tasks of isolated and continuous realtime HGR. Furthermore, we analyze the possibility of increasing the recognition accuracy by adding multiple modalities derived from RGB frames, i.e., optical flow and semantic segmentation, while keeping the real-time performance of the 3D-CNN model. Our empirical study also provides a comparison with the publicly available nvGesture (NVIDIA) dataset. The experimental results show that the state-of-the-art ResNext-101 model decreases about 30% accuracy when using our real-world dataset, demonstrating that the IPN Hand dataset can be used as a benchmark, and may help the community to step forward in the continuous HGR. Our dataset and pre-trained models used in the evaluation are publicly available at github.com/GibranBenitez/IPN-hand. Gibran Benitez-Garcia, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Keiji Yanai |
ICPR | 4 |
| 2020 | Mask-based Style-Controlled Image Synthesis Using a Mask Style EncoderabstractIn recent years, the advances in Generative Adversarial Networks (GANs) have shown impressive results for image generation and translation tasks. In particular, the image-to-image translation is a method of learning mapping from a source domain to a target domain and synthesizing an image. Image-to-image translation can be applied to a variety of tasks, making it possible to quickly and easily synthesize realistic images from semantic segmentation masks. However, in the existing image-to-image translation method, there is a limitation on controlling the style of the translated image, and it is not easy to synthesize an image by controlling the style of each mask element in detail. Therefore, we propose an image synthesis method that controls the style of each element by improving the existing image-to-image translation method. In the proposed method, we implement a mask style encoder that extracts style features for each mask element. The extracted style features are concatenated to the semantic mask in the normalization layer, and used the style-controlled image synthesis of each mask element. In the experiments, we performed style-controlled images synthesis using the datasets consisting of semantic segmentation masks and real images. The results show that the proposed method has excellent performance for style-controlled images synthesis for each element. Jaehyeong Cho, Wataru Shimoda, Keiji Yanai |
ICPR | 3 |
| 2020 | Fast and Accurate Real-Time Semantic Segmentation with Dilated Asymmetric ConvolutionsabstractRecent works have shown promising results applied to real-time semantic segmentation tasks. To maintain fast inference speed, most of the existing networks make use of light decoders, or they simply do not use them at all. This strategy helps to maintain a fast inference speed; however, their accuracy performance is significantly lower in comparison to non-real-time semantic segmentation networks. In this paper, we introduce two key modules aimed to design a high-performance decoder for real-time semantic segmentation for reducing the accuracy gap between real-time and non-real-time segmentation networks. Our first module, Dilated Asymmetric Pyramidal Fusion (DAPF), is designed to substantially increase the receptive field on the top of the last stage of the encoder, obtaining richer contextual features. Our second module, Multi-resolution Dilated Asymmetric (MDA) module, fuses and refines detail and contextual information from multi-scale feature maps coming from early and deeper stages of the network. Both modules exploit contextual information without excessively increasing the computational complexity by using asymmetric convolutions. Our proposed network entitled “FASSD-Net” reaches 78.8 % of mIoU accuracy on the Cityscapes validation dataset at 41.1 FPS on full resolution images (1024 x 2048). Besides, with a light version of our network, we reach 74.1 % of mIoU at 133.1 FPS (full resolution) on a single NVIDIA GTX 1080Ti card with no additional acceleration techniques. The source code and pre-trained models are available at github.com/GibranBenitez/FASSD- Net. Leonel Rosas-Arias, Gibran Benitez-Garcia, José Portillo-Portillo, Gabriel Sanchez-Perez, Keiji Yanai |
ICPR | 5 |
| 2020 | Hungry networks: 3D mesh reconstruction of a dish and a plate from a single dish image for estimating food volumeabstractDietary calorie management has been an important topic in recent years, and various methods and applications on image-based food calorie estimation have been published in the multimedia community. Most of the existing methods of estimating food calorie amounts use 2D-based image recognition. On the other hand, in this paper, we would like to make inferences based on 3D volume for more accurate estimation. We performed 3D reconstruction of a dish (food and plate) and a plate (without foods), from a single image. We succeeded in restoring the 3D shape with high accuracy while maintaining the consistency between a plate part of an estimated 3D dish and an estimated 3D plate. To achieve this, the following contributions were made in this paper. (1) Proposal of "Hungry Networks," a new network that generates two kinds of 3D volumes from a single image. (2) Introduction of plate consistency loss that matches the shapes of the plate parts of the two reconstructed models. (3) Creating a new dataset of 3D food models that are 3D scanned of actual foods and plates. We also conducted an experiment to infer the volume of only the food region from the difference of the two reconstructed volumes. As a result, it was shown that the introduced new loss function not only matches the 3D shape of the plate, but also contributes to obtaining the volume with higher accuracy. Although there are some existing studies that consider 3D shapes of foods, this is the first study to generate a 3D mesh volume from a single dish image. Shu Naritomi, Keiji Yanai |
MMAsia | 2 |
| 2020 | Weakly supervised semantic segmentation using distinct class specific saliency maps
Wataru Shimoda, Keiji Yanai |
Comput. Vis. Image Underst. | 2 |
| 2019 | Analyzing Regional Food Trends with Geo-tagged Twitter Food PhotosabstractTwitter is a world-wide popular SNS where many images are being posted along with tweet messages and location information. It is known that large regional differences exist regarding posted images. The differences are expected to be prominent especially for food images, since there are many kinds of regional foods over the world. However, the regional difference on food images over Twitter has not been explored so far. Therefore, in this paper, to make the difference clear, we analyzed geo-tagged food image gathered from the Twitter stream on the basis of the six regions and 17 kinds of rough food categories. For image analysis, we used only image features without any textual analysis, since Twitter messages are not always directly related to image contents. In addition, we visualized discriminative parts of local food images by applying the visualization method of CNNs. Kaimu Okamoto, Keiji Yanai |
CBMI | 2 |
| 2019 | Self-Supervised Difference Detection for Weakly-Supervised Semantic SegmentationabstractTo minimize the annotation costs associated with the training of semantic segmentation models, researchers have extensively investigated weakly-supervised segmentation approaches. In the current weakly-supervised segmentation methods, the most widely adopted approach is based on visualization. However, the visualization results are not generally equal to semantic segmentation. Therefore, to perform accurate semantic segmentation under the weakly supervised condition, it is necessary to consider the mapping functions that convert the visualization results into semantic segmentation. For such mapping functions, the conditional random field and iterative re-training using the outputs of a segmentation model are usually used. However, these methods do not always guarantee improvements in accuracy; therefore, if we apply these mapping functions iteratively multiple times, eventually the accuracy will not improve or will decrease. In this paper, to make the most of such mapping functions, we assume that the results of the mapping function include noise, and we improve the accuracy by removing noise. To achieve our aim, we propose the self-supervised difference detection module, which estimates noise from the results of the mapping functions by predicting the difference between the segmentation masks before and after the mapping. We verified the effectiveness of the proposed method by performing experiments on the PASCAL Visual Object Classes 2012 dataset, and we achieved 64.9% in the val set and 65.5% in the test set. Both of the results become new state-of-the-art under the same setting of weakly supervised semantic segmentation. Wataru Shimoda, Keiji Yanai |
ICCV | 2 |
| 2019 | DeepTaste: Augmented Reality Gustatory Manipulation with GAN-Based Real-Time Food-to-Food TranslationabstractWe have been studying augmented reality (AR)-based gustatory manipulation interfaces and previously proposed a gustatory manipulation interface using generative adversarial network (GAN)-based real time image-to-image translation. Unlike three-dimensional (3D) food model-based systems that only change the color or texture pattern of a particular type of food in an inflexible manner, our GAN-based system changes the appearance of food into multiple types of food in real time flexibly, dynamically, and interactively. In the present paper, we first describe in detail a user study on a vision-induced gustatory manipulation system using a 3D food model and report its successful experimental results. We then summarize identified problems of the 3D model-based system and describe implementation details of the GAN-based system. We finally report in detail the main user study in which we investigated the impact of the GAN-based system on gustatory sensations and food recognition when somen noodles were turned into ramen noodles or fried noodles, and steamed rice into curry and rice or fried rice. The experimental results revealed that our system successfully manipulates gustatory sensations to some extent and that the effectiveness seems to depend on the original and target types of food as well as the experience of each individual with the food. Kizashi Nakano, Daichi Horita, Nobuchika Sakata, Kiyoshi Kiyokawa, Keiji Yanai, Takuji Narumi |
ISMAR | 5 |
| 2019 | Ramen as You Like: Sketch-based Food Image Generation and EditingabstractIn recent years, a large number of images are being posted on SNS. The users often synthesize or modify their photos before uploading them. However, the task of synthesizing and modifying photos requires a lot of time and skill. In this demo, we demonstrate easy and fast image synthesis and modification through "sketch-based food image generation''. The proposed system uses pix2pix to generate realistic food images based on sketched images, and DeepLab V3+ to obtain sketch masks from real photos. A user can create a realistic food image easily and fast by sketching a mask image consisting of food elements. In addition, a user can also edit a mask image automatically generated from a real photo food photo, and generate a modified food image. For training, we have created a new ramen image dataset consisting of 555 images with 15 kinds of pixel-wise labels. Jaehyeong Cho, Wataru Shimoda, Keiji Yanai |
ACM Multimedia | 3 |
| 2019 | MADiMA'19: 5th International Workshop on Multimedia Assisted Dietary ManagementabstractThis abstract provides a summary and overview of the 5th International Workshop on Multimedia Assisted Dietary Management. Stavroula G. Mougiakakou, Giovanni Maria Farinella, Keiji Yanai, Dario Allegra |
ACM Multimedia | 3 |
| 2019 | Enchanting Your Noodles: GAN-based Real-time Food-to-Food Translation and Its Impact on Vision-induced Gustatory ManipulationabstractWe propose a novel gustatory manipulation interface which utilizes the cross-modal effect of vision on taste elicited with augmented reality (AR)-based real-time food appearance modulation using a generative adversarial network (GAN). Unlike existing systems which only change color or texture pattern of a particular type of food in an inflexible manner, our system changes the appearance of food into multiple types of food in real-time flexibly, dynamically and interactively in accordance with the deformation of the food that the user is actually eating by using GAN-based image-to-image translation. The experimental results reveal that our system successfully manipulates gustatory sensations to some extent and that the effectiveness depends on the original and target types of food as well as each user's food experience. Kizashi Nakano, Kiyoshi Kiyokawa, Daichi Horita, Keiji Yanai, Nobuchika Sakata, Takuji Narumi |
VR | 4 |
| 2019 | Enchanting Your Noodles: A Gustatory Manipulation Interface by Using GAN-based Real-time Food-to-Food TranslationabstractIn this demonstration, we present a novel gustatory manipulation interface which utilizes the cross-modal effect of vision on taste elicited with real-time food appearance modulation using a generative adversarial network (GAN). Unlike existing systems which only change color or texture pattern of a particular type of food in an inflexible manner, our system changes the appearance of food into multiple types of food in real-time flexibly, dynamically and interactively in accordance with the deformation of the food that the user is actually eating by using GAN-based image-to-image translation. Our system can turn somen noodles into ramen noodles or fried noodles, or steamed rice into curry and rice or fried rice. Users of our demonstration system will taste what is visually presented to some extent rather than what they are actually eating. Kizashi Nakano, Kiyoshi Kiyokawa, Daichi Horita, Keiji Yanai, Nobuchika Sakata, Takuji Narurni |
VR | 4 |
| 2018 | Magical Rice Bowl: A Real-time Food Category ChangerabstractIn this demo, we demonstrate "Real-time Food Category Change'' based on a Conditional Cycle GAN (cCycle GAN) with a large-scale food image data collected from the Twitter Stream. Conditional Cycle GAN is an extension of CycleGAN, which enables "Food Category Change'' among ten kinds of typical foods served in bowl-type dishes such as beef rice bowl and ramen noodles. The proposed system enables us to change the appearance of a given food photo according to the given category keeping the shape of the given food but exchanging its textures. For training, we used two hundred and thirty thousand food images which achieved very natural food category change among ten kinds of typical Japanese foods: ramen noodle, curry rice, fried rice, beef rice bowl, chilled noodle, spaghetti with meat source, white rice, eel bowl, and fried noodle. Ryosuke Tanno, Daichi Horita, Wataru Shimoda, Keiji Yanai |
ACM Multimedia | 4 |
| 2018 | AR DeepCalorieCam: An iOS App for Food Calorie Estimation with Augmented Reality
Ryosuke Tanno, Takumi Ege, Keiji Yanai |
MMM (2) | 3 |
| 2018 | Ramen spoon eraser: CNN-based photo transformation for improving attractiveness of ramen photosabstractIn recent years, a large number of food photos are being posted globally on SNS. To obtain many views or "likes", attractive photos should be posted. However, some casual foods are served with utensils on a plate or a bowl at restaurants, which spoils attractiveness of meal photos. Especially in Japan where ramen noodle is the most popular casual food, ramen is usually served with a ramen spoon in a ramen bowl in a ramen noodle shop. This is a big problem for SNS photographers, because a ramen spoon soaked in a ramen bowl extremely degrades the appearance of ramen photos. Then, in this paper, we propose anapplication called "ramen spoon eraser" that erases a spoon from ramen photos with spoons using a CNN-based Image-to-Image translation network. In this application, it is possible to automatically erase ramen spoons from ramen photos, which extremely improve the attractiveness of ramen photos. In the experiment, we train models in two ways as CNN-based Image-to-Image translation networks with the dataset consisting of ramen images with / without spoons collected from the Web. Daichi Horita, Jaehyeong Cho, Takumi Ege, Keiji Yanai |
VRST | 4 |
| 2018 | AR DeepCalorieCam V2: food calorie estimation with CNN and AR-based actual size estimationabstractIn most of the cases, the estimated calories are just associated with the estimated food categories, or the relative size compared to the standard size of each food category which are usually provided by a user manually. In addition, in the case of calorie estimation based on the amount of meal, a user conventionally needs to register a size-known reference object in advance and to take a food photo with the registered reference object. In this demo, we propose a new approach for food calorie estimation with CNN and Augmented Reality (AR)-based actual size estimation. By using Apple ARKit framework, we can measure the actual size of the meal area by acquiring the coordinates on the real world as a three-dimensional vector, we implemented this demo app. As a result, it is possible to calculate the size more accurately than in the previous method by measuring the meal area directly, the calorie estimation accuracy has improved. Ryosuke Tanno, Takumi Ege, Keiji Yanai |
VRST | 3 |
| 2017 | Scene Text EraserabstractThe character information in natural scene images contains various personal information, such as telephone numbers, home addresses, etc. It is a high risk of leakage the information if they are published. In this paper, we proposed a scene text erasing method to properly hide the information via an inpainting convolutional neural network (CNN) model. The input is a scene text image, and the output is expected to be text erased image with all the character regions filled up the colors of the surrounding background pixels. This work is accomplished by a CNN model through convolution to deconvolution with interconnection process. The training samples and the corresponding inpainting images are considered as teaching signals for training. To evaluate the text erasing performance, the output images are detected by a novel scene text detection method. Subsequently, the same measurement on text detection is utilized for testing the images in benchmark dataset ICDAR2013. Compared with direct text detection way, the scene text erasing process demonstrates a drastically decrease on the precision, recall and f-score. That proves the effectiveness of proposed method for erasing the text in natural scene images. Toshiki Nakamura, Anna Zhu, Keiji Yanai, Seiichi Uchida |
ICDAR | 3 |
| 2017 | Conditional Fast Style Transfer NetworkabstractIn this paper, we propose a conditional fast neural style transfer network. We extend the network proposed as a fast neural style transfer network by Johnson et al. [1] so that the network can learn multiple styles at the same time. To do that, we add a conditional input which selects a style to be transferred out of the trained styles. In addition, we show that the proposed network can mix multiple styles, although the network is trained with each of the training styles independently. The proposed network can also transfer different styles to the different parts of a given image at the same time, which we call "spatial style transfer". In the experiments, we confirmed that no quality degradation occurred in the multi-style network compared to the single network, and linear-weighted multi-style fusion enabled us to generate various kinds of new styles which are different from the trained single styles. Keiji Yanai, Ryosuke Tanno |
ICMR | 1 |
| 2017 | DeepStyleCam: A Real-Time Style Transfer App on iOS
Ryosuke Tanno, Shin Matsuo, Wataru Shimoda, Keiji Yanai |
MMM (2) | 4 |
| 2017 | Guest Editorial Nutrition Informatics: From Food Monitoring to Dietary ManagementabstractThe papers in this special section address the concept of nutrition informatics from food monitoring to dietary management. Non-communicable diseases (NCD) account for a massively increasing proportion of the global health burden. A number of behavioral and physiological factors are related to the rising onset of NCD worldwide, with unhealthy eating playing a key role among them. In parallel, food allergies and associated acute and sometimes life-threatening reactions are a public health problem. Thus, balanced nutrition with a proper diet is the key to the prevention of diet related diseases. The recent advances in smartphone technologies, wearable sensors, computer vision and machine learning will bring the applications of nutrition informatics closer to the individuals and enable them to make better decisions regarding their daily lives. Stavroula G. Mougiakakou, Giovanni Maria Farinella, Keiji Yanai, Edward Sazonov |
IEEE J. Biomed. Health Informatics | 3 |
| 2016 | Distinct Class-Specific Saliency Maps for Weakly Supervised Semantic Segmentation
Wataru Shimoda, Keiji Yanai |
ECCV (4) | 2 |
| 2016 | Weakly-supervised segmentation by combining CNN feature maps and object saliency mapsabstractIn general, CNN based semantic segmentation methods assume pixel-wise annotation is available, which is costly to obtain in general. On the other hand, image-level annotations is much easier to obtain than pixel-level annotation. Then, in this work, we focus on weakly-supervised semantic segmentation which is known as task of using training data with only image-level annotations. In this paper, we propose a new CNN-based semantic segmentation method which uses both activation features calculated by feed-forwarding and object saliency maps obtained by back-propagation. As a CNN, we use the VGG-16 pre-trained with 1000-class ILSVRC datasets and fine-tuned it with multi-label training using only image-level labeled dataset. By the experiments, we show that the proposed method achieved state-of-the-art results with the PASCAL VOC 2012 dataset. Wataru Shimoda, Keiji Yanai |
ICPR | 2 |
| 2016 | CNN-based Style Vector for Style Image RetrievalabstractIn this paper, we have examined the effectiveness of "style matrix" which is used in the works on style transfer and texture synthesis by Gatys et al. in the context of image retrieval as image features. A style matrix is presented by Gram matrix of the feature maps in a deep convolutional neural network. We proposed a style vector which are generated from a style matrix with PCA dimension reduction. In the experiments, we evaluate image retrieval performance using artistic images downloaded from Wikiarts.org regarding both artistic styles ans artists. We have obtained 40.64% and 70.40% average precision for style search and artist search, respectively, both of which outperformed the results by common CNN features. In addition, we found PCA-compression boosted the performance. Shin Matsuo, Keiji Yanai |
ICMR | 2 |
| 2016 | Overview of the ACM MultiMedia 2016 International Workshop on Multimedia Assisted Dietary ManagementabstractThis abstract provides a summary and overview of the 2nd international workshop on multimedia assisted dietary management. Stavroula G. Mougiakakou, Giovanni Maria Farinella, Keiji Yanai |
ACM Multimedia | 3 |
| 2016 | Efficient Mobile Implementation of A CNN-based Object Recognition SystemabstractBecause of the recent progress on deep learning studies, Convolutional Neural Network (CNN) based method have outperformed conventional object recognition methods with a large margin. However, it requires much more memory and computational costs compared to the conventional methods. Therefore, it is not easy to implement a CNN-based object recognition system on a mobile device where memory and computational power are limited. In this paper, we examine CNN architectures which are suitable for mobile implementation, and propose multi-scale network-in-networks (NIN) in which users can adjust the trade-off between recognition time and accuracy. We implemented multi-threaded mobile applications on both iOS and Android employing either NEON SIMD instructions or the BLAS library for fast computation of convolutional layers, and compared them in terms of recognition time on mobile devices. As results, it has been revealed that BLAS is better for iOS, while NEON is better for Android, and that reducing the size of an input image by resizing is very effective for speedup of CNN-based recognition. Keiji Yanai, Ryosuke Tanno, Koichi Okamoto |
ACM Multimedia | 1 |
| 2016 | GrillCam: A Real-Time Eating Action Recognition System
Koichi Okamoto, Keiji Yanai |
MMM (2) | 2 |
| 2016 | Event photo mining from Twitter using keyword bursts and image clustering
Takamu Kaneko, Keiji Yanai |
Neurocomputing | 2 |
| 2015 | A visual analysis on recognizability and discriminability of onomatopoeia words with DCNN featuresabstractIn this paper, we examine the relation between onomatopoeia and images using a large number of Web images. The objective of this paper is to examine if the images corresponding to Japanese onomatopoeia words which express the feeling of visual appearance can be recognized by the state-of-the-art visual recognition methods. In our work, first, we collect the images corresponding to onomatopoeia words using an Web image search engine, and then we filter out noise images to obtain clean dataset with automatic image re-ranking method. Next, we analyze the recognizability of various kinds of onomatopoeia images using improved Fisher vector (IFV) and deep convolutional neural network (DCNN) features. In addition, we collect images corresponding to the pairs of nouns and onomatopoeia words, and we examine if the images associated with the same nouns and the different onomatopoeia words are visually discriminable or not. By the experiments, it has been shown that the DCNN features extracted from the layer 7 of Overfeat's network pre-trained with the ILSVRC 2013 data have prominent ability to represent onomatopoeia images, and most of the onomatopoeia words have visual characteristics which can be recognized. Wataru Shimoda, Keiji Yanai |
ICME | 2 |
| 2015 | Automatic Construction of Action Datasets Using Web Videos with Density-Based Cluster Analysis and Outlier Detection
Do Hang Nga, Keiji Yanai |
PSIVT | 2 |
| 2015 | FoodCam: A real-time food recognition system on a smartphone
Yoshiyuki Kawano, Keiji Yanai |
Multim. Tools Appl. | 2 |
| 2014 | Real-Time Photo Mining from the Twitter Stream: Event Photo Discovery and Food Photo DetectionabstractSo many people are posting photos as well as short messages to Twitter every minutes from everywhere on the earth. By monitoring the Twitter stream, we can obtain various kinds of photos with texts. In this paper, as case studies of real-time Twitter photo mining, we introduce our current on-going projects on event photo discovery and food photo mining from the Twitter stream. Keiji Yanai, Takamu Kaneko, Yoshiyuki Kawano |
ISM | 1 |
| 2014 | FoodCam-256: A Large-scale Real-time Mobile Food RecognitionSystem employing High-Dimensional Features and Compression of Classifier WeightsabstractIn the demo, we demonstrate a large-scale food recognition system employing high-dimensional Fisher Vector and liner one-vs-rest classifiers. Since all the processes on image recognition perform on a smartphone, the system does not require an external image recognition server, and runs on an ordinary smartphone in a real-time way. Yoshiyuki Kawano, Keiji Yanai |
ACM Multimedia | 2 |
| 2014 | FoodCam: A Real-Time Mobile Food Recognition System Employing Fisher Vector
Yoshiyuki Kawano, Keiji Yanai |
MMM (2) | 2 |
| 2014 | A Dense SURF and Triangulation Based Spatio-temporal Feature for Action Recognition
Do Hang Nga, Keiji Yanai |
MMM (1) | 2 |
| 2014 | Automatic extraction of relevant video shots of specific actions exploiting Web data
Do Hang Nga, Keiji Yanai |
Comput. Vis. Image Underst. | 2 |
| 2013 | Large-scale web video shot ranking based on visual features and tag co-occurrenceabstractIn this paper, we propose a novel ranking method, VisualTextualRank, which extends [1] and [2]. Our method is based on random walk over bipartite graph to integrate visual information of video shots and tag information of Web videos effectively. Note that instead of treating the textual information as an additional feature for shot ranking, we explore the mutual reinforcement between shots and textual information of their corresponding videos to improve shot ranking. We apply our proposed method to the system of extracting automatically relevant video shots of specific actions from Web videos [3]. Based on our experimental results, we demonstrate that our ranking method can improve the performance of video shot retrieval. Do Hang Nga, Keiji Yanai |
ACM Multimedia | 2 |
| 2013 | Visual Analysis of Tag Co-occurrence on Nouns and Adjectives
Yuya Kohara, Keiji Yanai |
MMM (1) | 2 |
| 2013 | Summarization of Egocentric Moving Videos for Generating Walking Route Guidance
Masaya Okamoto, Keiji Yanai |
PSIVT | 2 |
| 2012 | Recognition of Multiple-Food Images by Detecting Candidate RegionsabstractIn this paper, we propose a two-step method to recognize multiple-food images by detecting candidate regions with several methods and classifying them with various kinds of features. In the first step, we detect several candidate regions by fusing outputs of several region detectors including Felzenszwalb's deformable part model (DPM) [1], a circle detector and the JSEG region segmentation. In the second step, we apply a feature-fusion-based food recognition method for bounding boxes of the candidate regions with various kinds of visual features including bag-of-features of SIFT and CSIFT with spatial pyramid (SP-BoF), histogram of oriented gradient (HoG), and Gabor texture features. In the experiments, we estimated ten food candidates for multiple-food images in the descending order of the confidence scores. As results, we have achieved the 55.8% classification rate, which improved the baseline result in case of using only DPM by 14.3 points, for a multiple-food image data set. This demonstrates that the proposed two-step method is effective for recognition of multiple-food images. Yuji Matsuda, Hajime Hoashi, Keiji Yanai |
ICME | 3 |
| 2012 | Multiple-food recognition considering co-occurrence employing manifold ranking
Yuji Matsuda, Keiji Yanai |
ICPR | 2 |
| 2012 | World seer: a realtime geo-tweet photo mapping systemabstractTwitter is a unique microblog which is different from conventional social media in terms of its quickness. Many Twitter's users send messages to Twitter on the spot with mobile phones or smart phones, and some of them send tweets with photos and geotags, which can be regarded as being geotagged photos. Geotagged tweet photos are very useful to understand what happens currently over the world. In the demo, we introduce "World Seer" which is a real-time geo-tweet photo mapping system. Users can see the latest geo-tweet photos related to given keywords and areas on the online maps. The system shows geo-tweet photos not only on the map, but also on the street-view. In addition, for some parts of the geo-tweet photos, the system can show representative photos for the given locations and the given times employing the GeoVisualRank method which takes into account both visual features of photos and proximity of geotags. Keiji Yanai |
ICMR | 1 |
| 2011 | Automatic construction of an action video shot database using web videosabstractThere are a huge number of videos with text tags on the Web nowadays. In this paper, we propose a method of automatically extracting from Web videos video shots corresponding to specific actions with just only providing action keywords such as “walking” and “eating”. The proposed method consists of three steps: (1) tag-based video selection, (2) segmenting videos into shots and extracting features from the shots, and (3) visual-feature-based video shot selection with tag-based scores taken into account. Firstly, we gather video IDs and tag lists for 1000 Web videos corresponding to given keywords via Web API, and we calculate tag relevance scores for each video using a tag-co-occurrence dictionary which is constructed in advance. Secondly, we fetch the top 200 videos from the Web in the descending order of the tag relevance scores, and segment each downloaded video into several shots. From each shot we extract spatio-temporal features, global motion features and appearance features, and convert them into the bag-of-features representation. Finally, we apply the VisualRank method to select the video shots which describe the actions corresponding to the given keywords best after calculating a similarity matrix between video shots. In the experiments, we achieved the 49.5% precision at 100 shots over six kinds of human actions by just providing keywords without any supervision. In addition, we made large-scale experiments on 100 kinds of action keywords. Do Hang Nga, Keiji Yanai |
ICCV | 2 |
| 2010 | Geotagged Image Recognition by Combining Three Different Kinds of Geolocation Features
Keita Yaegashi, Keiji Yanai |
ACCV (2) | 2 |
| 2010 | Geotagged Photo Recognition Using Corresponding Aerial Photos with Multiple Kernel LearningabstractIn this paper, we treat with generic object recognition for geotagged images. As a recognition method for geotagged photos, we have already proposed exploiting aerial photos around geotag places as additional image features for visual recognition of geotagged photos. In the previous work, to fuse two kinds of features, we just concatenate them. Instead, in this paper, we introduce Multiple Kernel Learning (MKL) to integrate both features of photos and aerial images. MKL can estimate the contribution weights to integrate both kinds of features. In the experiments, we confirmed effectiveness of usage of aerial photos for recognition of geotagged photos, and we evaluated the weights of both features estimated by MKL for eighteen concepts. Keita Yaegashi, Keiji Yanai |
ICPR | 2 |
| 2010 | Image Recognition of 85 Food Categories by Feature FusionabstractRecognition of food images is challenging due to their diversity and practical for health care on foods for people. In this paper, we propose an automatic food image recognition system for 85 food categories by fusing various kinds of image features including bag-of-features (BoF), color histogram, Gabor features and gradient histogram with Multiple Kernel Learning (MKL). In addition, we implemented a prototype system to recognize food images taken by cellular-phone cameras. In the experiment, we have achieved the 62.52% classification rate for 85 food categories. Hajime Hoashi, Taichi Joutou, Keiji Yanai |
ISM | 3 |
| 2010 | Automatic Construction of a Folksonomy-Based Visual OntologyabstractRecently, Folksonomy attracts attentions as a new method to index large-scale image databases. In the Folksonomy-style image databases, they allows users to attach keywords to images as “tags”. Since tag words are uncontrolled, they have various and many kinds of tags associated with images. This is much different from conventional image databases. In this paper, we propose a novel method to extract hierarchical structure on relations between tags from Folksonomy. The tag structure we extract can be used as an ontology for image database search which reflects both textual and visual relations between tags. In the proposed method, at first, we collect millions of tag-attached-images from Flickr which is the world-largest Folksonomy-style image database, and remove noise images from them. Next, we estimate concept vectors for highly-frequent tags based on only visual features, only tag word features and combined features of both visual and textual features, and compute JS divergence and entropy for three kinds of concept vectors. Finally we estimate hierarchical structures between tags regarding three kinds of concept vectors. In the experiments, we show the obtained hierarchical structure, and it includes interesting relations which sometimes are difficult to be discovered by human. In addition, as its application, we used and evaluated the obtained ontology for query expansion of text-tag-based image search over Flickr. These results indicate that the proposed method is promising and the structure is expected to help image search and some other applications. Hidetoshi Kawakubo, Yuuta Akima, Keiji Yanai |
ISM | 3 |
| 2009 | Extracting Spatio-temporal Local Features Considering Consecutiveness of Motions
Akitsugu Noguchi, Keiji Yanai |
ACCV (2) | 2 |
| 2009 | A food image recognition system with Multiple Kernel LearningabstractSince health care on foods is drawing people's attention recently, a system that can record everyday meals easily is being awaited. In this paper, we propose an automatic food image recognition system for recording people's eating habits. In the proposed system, we use the Multiple Kernel Learning (MKL) method to integrate several kinds of image features such as color, texture and SIFT adaptively. MKL enables to estimate optimal weights to combine image features for each category. In addition, we implemented a prototype system to recognize food images taken by cellular-phone cameras. In the experiment, we have achieved the 61.34% classification rate for 50 kinds of foods. To the best of our knowledge, this is the first report of a food image classification system which can be applied for practical use. Taichi Joutou, Keiji Yanai |
ICIP | 2 |
| 2009 | An analysis of the relation between visual concepts and geo-locations using geotagged images on the webabstractRecently, a large number of geotagged images are available on photo sharing Web sites such as Flickr. In this paper, we propose image region entropy and geo-location entropy for analyzing the relation between visual concepts and geographical locations using a large-scale geotagged image database. Image region entropy represents to what extent concepts have visual characteristics, while geo-location entropy represents to what extent concepts are distributed over the world. In the experiment, we analyzed relations between image region entropy and geo-location entropy in terms of 230 nouns and 100 adjectives, and we found that the concepts with low image entropy tend to have high geo-location entropy and vice versa. Hidetoshi Kawakubo, Keiji Yanai |
ICME | 2 |
| 2009 | Web image gathering with region-based bag-of-features and multiple instance learningabstractWe propose a new Web image gathering system which employs the region-based bag-of-features representation and multiple instance learning. The contribution of this work is introducing the region-based bag-of-features representation into an Web image gathering task where training data is incomplete and having proved its effectiveness by comparing the proposed method with the normal whole-image-based bag-of-features representation. In our method, first, we perform region segmentation for an image, and next we generate a bag-of-features vector for each region. One image is represented by a set of bag-of-features vectors in this paper, while one image is represented by just one bag-of-features vector in the normal bag-of-features representation which is very popular for visual object categorization tasks recently. Several works on Web image selection with bag-of- features have been proposed so far. However, in case that the training data includes much noise, sufficient results could not be obtained. In this paper, we divide images into regions and classify each region with multiple-instance support vector machine (mi-SVM) instead of classifying whole images. By this region-based classification, we can separate foreground regions from background regions and achieve more effective image training from incomplete training data. By the experiments, we show that the results by the proposed methods outperformed the results by the whole-image-based bag-of-visual-words and the normal support vector machine. Keiji Yanai |
ICME | 1 |
| 2009 | Can Geotags Help Image Recognition?
Keita Yaegashi, Keiji Yanai |
PSIVT | 2 |
| 2009 | Mining cultural differences from a large number of geotagged photosabstractWe propose a novel method to detect cultural differences over the world automatically by using a large amount of geotagged images on the photo sharingWeb sites such as Flickr. We employ the state-of-the-art object recognition technique developed in the research community of computer vision to mine representative photos of the given concept for representative local regions from a large-scale unorganized collection of consumer-generated geotagged photos. The results help us understand how objects, scenes or events corresponding to the same given concept are visually different depending on local regions over the world. Keiji Yanai, Bingyu Qiu |
WWW | 1 |
| 2008 | Web video retrieval based on the Earth Mover's Distance by integrating color, motion and soundabstractIn this paper, we propose a novel content-based video retrieval method for short video clips which are stored on consumer video sharing Web sites. It is based on the Earth Mover’s Distance which enables us to evaluate dissimilarities among videos where the number of shots and time length are different. As features extracted from videos, we use color, motion, sound and position of shots. By defining the ground distance of EMD as the weighted sum of Euclid distances of these four kinds of features, we integrate them when calculating EMD. In the experiments on video retrieval for YouTube videos, we obtained the 0.98 average precision at most, which shows effectiveness of the proposed method. In addition, the results of integration of four kinds of features outperformed the ones of single features, which shows that feature combination is effective. Keisuke Takada, Keiji Yanai |
ICIP | 2 |
| 2008 | Web image selection with PLSAabstractIn this paper, we propose a new method to select relevant images to the given keywords from the images gathered from the Web. Our novel method is based on the Probabilistic Latent Semantic Analysis (PLSA) model, which is a generative probabilistic topic model. Firstly, we gather images related to the given keywords from the Web with Web search engines. Secondly, we choose pseudo-training images from them by simple heuristic HTML analysis, and train our PLSA-based probabilistic model with them. Finally, we select relevant images from all the gathered images with the learned model. The experimental results shows that the results by the proposed method is almost equivalent to the results by existing methods, although our method does not need to prepare negative training samples in advance unlike existing methods. Keiji Yanai |
ICME | 1 |
| 2008 | Web Image Gathering with a Part-Based Object Recognition Method
Keiji Yanai |
MMM | 1 |
| 2008 | Automatic web image selection with a probabilistic latent topic modelabstractWe propose a new method to select relevant images to the given keywords from images gathered from theWeb based on the Probabilistic Latent Semantic Analysis (PLSA) model which is a probabilistic latent topic model originally proposed for text document analysis. The experimental results shows that the results by the proposed method is almost equivalent to or outperforms the results by existing methods. In addition, it is proved that our method can select more various images compared to the existing SVM-based methods. Keiji Yanai |
WWW | 1 |
| 2007 | Image collector III: a web image-gathering system with bag-of-keypointsabstractWe propose a new system to mine visual knowledge on the Web.There are huge image data as well as text data on the Web. However, mining image data from the Web is paid less attention than mining text data, since treating semantics of images are much more difficult. In this paper, we propose introducing a latest image recognition technique, which is the bag-of-keypoints representation,into Web image-gathering task. By the experiments we show theproposed system outperforms our previous systems and Google Imagesearch greatly. Keiji Yanai |
WWW | 1 |
| 2006 | Automatic "Go" record generation from a TV programabstractWe present a video recognition system of a "Go" TV program. It generates a Go play record automatically from a broadcast of Go played by human professionals. "Go" is the ancient Asian board game played between two player, which is similar to Chess and Shogi. For an MPEG2 video of a TV Go program, the system distinguishes play or commentary board shots from other types of shots such as player's shots, and detects Go stones placed on the board from board shots. The system removes several types of noise such as a player's head or hand. In addition, it also detects Go stone from commentary board shots which are often inserted between play board shots, and compensates for the order of Go stones placed on the play board during commentary board shots. In the experimental results for eight TV Go program the system have achieved the 95.7% precision and the 95.7% recall rate. Keiji Yanai, Takehisa Hayashiyama |
MMM | 1 |
| 2006 | Finding visual concepts by web image miningabstractWe propose measuring "visualness" of concepts with images on the Web, that is, what extent concepts have visual characteristics. This is a new application of "Web image mining". To know which concept has visually discriminative power is important for image recognition, since not all concepts are related to visual contents. Mining image data on the Web with our method enables it. Our method performs probabilistic region selection for images and computes an entropy measure which represents "visualness" of concepts. In the experiments, we collected about forty thousand images from the Web for 150 concepts. We examined which concepts are suitable for annotation of image contents. Keiji Yanai, Kobus Barnard |
WWW | 1 |
| 2005 | Image region entropy: a measure of "visualness" of web images associated with one conceptabstractWe propose a new method to measure visualness of concepts, that is, what extent concepts have visual characteristics. To know which concept has visually discriminative power is important for image annotation, especially automatic image annotation by image recognition system, since not all concepts are related to visual contents. Our method performs probabilistic region selection for images which are labeled as concept X or non-X, and computes an entropy measure which represents visualness of concepts. In the experiments, we collected about forty thousand images from the World-Wide Web using the Google Image Search for 150 concepts. We examined which concepts are suitable for annotation of image contents. Keiji Yanai, Kobus Barnard |
ACM Multimedia | 1 |
| 2004 | A fast image-gathering system from the World-Wide Web using a PC cluster
Keiji Yanai, Masaya Shindo, Kohei Noshita |
Image Vis. Comput. | 1 |
| 2003 | Image collector II: a system for gathering more than one thousand images from the Web for one keywordabstractWe propose a system that enables us to gather more than one thousand images from the World Wide Web. The system is called Image Collector II. The image collector, which we proposed previously, can gather only several hundreds images. We made the two following improvements to extend the ability of our previous system in terms of the number of gathered images and their precision: (1) We extracted some words appearing with high frequency from all HTML files embedding output images in an initial image gathering, and using them as keywords, we made a second image gathering again. Through this, we obtained more than one thousand images for one keyword. (2) The more images we gathered, the more he precision of gathered images decreased. To raise the precision, we introduced word vectors of HTML files embedding images into the image selecting process in addition to image feature vectors. Keiji Yanai |
ICME | 1 |
| 2003 | Generic image classification using visual knowledge on the webabstractIn this paper, we describe a generic image classification system with an automatic knowledge acquisition mechanism from the World-Wide Web. Due to the recent spread of digital imaging devices, the demand for image recognition of various kinds of real world scenes becomes greater. For realizing it, visual knowledge on various kinds of scenes is required. Then, we propose gathering visual knowledge on real world scenes for generic image classification from the World-Wide Web. Our system gathers a large number of images from the Web automatically and makes use of them as training images for generic image classification. It consists of three modules, which are an image-gathering module, an image-learning module and an image classification module. The image-gathering module gathers images related to given class keywords from the Web automatically. The learning module extracts image features from gathered images and associates them with each class. The image classification module classifies an unknown image into one of the classes corresponding to the class keywords by using the association between image features and classes. In the experiments, we achieved a classification rate 44.6% for generic images by using images gathered from the World-Wide Web automatically as training images. Keiji Yanai |
ACM Multimedia | 1 |
| 2002 | Image Classification by Web Images
Keiji Yanai |
PRICAI | 1 |
| 2001 | Image Collector: An Image-Gathering System From The World-Wide Web Employing Keyword-Based Search EnginesabstractDue to the recent explosive progress of WWW (World-Wide Web), we can easily access a large number of images from WWW. There are, however, no established methods to make use of WWW as a large image database. In this paper, we describe an automatic image-gathering system from WWW employing keywords and image features, which is called the Image Collector. By exploiting some existing keyword-based search engines and selecting images by their image features, our system obtains, with high accuracy, images that are strongly related to query keywords. We have implemented the system that gathers more than one hundred images from WWW in about five minutes. Keiji Yanai |
ICME | 1 |
| 2001 | A Fast Image-Gathering System on the World-Wide Web Using a PC Cluster
Keiji Yanai, Masaya Shindo, Kohei Noshita |
Web Intelligence | 1 |
| 2000 | Recognition of Indoor Images Employing Qualitative Model Fitting and Supporting Relation between ObjectsabstractWe describe a new design of a recognition system for a single image of indoor scene including complex occlusions. In our system the system first estimates the 3D structure of an object by fitting a 3D structure model to the image qualitatively. Next, by checking the supporting relation between objects, it eliminates object candidates that are impossible to exist and estimates actual objects from their parts in the image. Finally, the system then recognizes the objects that are consistent with each other. We implemented the system as a multi-agent-based image understanding system. This paper describes an outline of the system and results of recognition experiments. Keiji Yanai, Koichiro Deguchi |
ICPR | 1 |
| 1998 | An architecture of object recognition system for various images based on multi-agentsabstractAn image understanding system for real world images, which has the ability to recognize various kinds of images, is proposed. We propose a multi-agent architecture to integrate object recognition modules for individual target objects. In our method, the recognition results by different agents are fused not only on the evaluations by each modules themselves but also on relations of object locations, sizes, etc. This is carried out autonomously between the agents concerned, and the most reliable result is selected after the arbitration between them. We implemented an experimental system on a parallel computer, and achieved recognition for both indoor and outdoor images. Keiji Yanai, Koichiro Deguchi |
ICPR | 1 |