VLDB 2026 Research / reviewers in the wild / expert
Yoko Yamakata
dblp:07/3918
· DBLP profile ↗
36ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0003-2752-6179ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 16 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JFC-Recipe: A Dataset for Nutrient Estimation from Japanese User-Generated Cooking Recipes
Keisuke Shirai, Yoko Yamakata, Hirotaka Kameko, Akiko Sunto, Jun Harashima, Shinsuke Mori |
LREC | 2 |
| 2026 | LLM-Based Explainable Detection of LLM-Generated Code in Python Programming Courses
Jeonghun Baek, Tetsuro Yamazaki, Akimasa Morihata, Junichiro Mori, Yoko Yamakata, Kenjiro Taura, Shigeru Chiba |
SIGCSE (1) | 5 |
| 2026 | MaskingAgent: Preventing LLM Tutor from Providing Full Solutions in Python Programming Courses
Jeonghun Baek, Tetsuro Yamazaki, Akimasa Morihata, Junichiro Mori, Yoko Yamakata, Kenjiro Taura, Shigeru Chiba |
SIGCSE (2) | 5 |
| 2025 | Redefining Image-to-Recipe Retrieval with Nutritional and Ingredient SimilarityabstractCross-modal retrieval models have shown impressive performance on the image-to-recipe retrieval task, a common benchmark in the multimedia field. However, the task assumes that an exact recipe match for a query image exists in the target database—an assumption that rarely holds true in real-world scenarios. When excluding exact matches from the target domain, our analysis revealed that relying solely on visual and textual similarity between recipes is insufficient to achieve good retrieval results. Other similarities should also be considered. Since ingredient similarity aligns with human intuition and nutritional similarity is crucial for health-conscious applications, we propose a model that incorporates ingredient and nutritional relevance into the retrieval process. We measured the similarity of unpaired recipes using three new metrics: mean absolute scaled error (MASE) for assessing nutritional similarity and IOU and weighted IOU (WIOU) for measuring ingredient overlap. Our proposed method can also be applied with existing image-recipe retrieval models and improved top-1 MASE, IOU, and WIOU by up to 18.13%, 9.91%, and 6.88%. Satayu Parinayok, Shin'ichi Satoh 0001, Kiyoharu Aizawa, Yoko Yamakata |
ICME | 4 |
| 2025 | A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
Mashiro Toyooka, Kiyoharu Aizawa, Yoko Yamakata |
ACM Multimedia | 3 |
| 2025 | FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management ApplicationsabstractFood image classification models are crucial for dietary management applications because they reduce the burden of manual meal logging. However, most publicly available datasets for training such models rely on web-crawled images, which often differ from users' real-world meal photos. In this work, we present FoodLogAthl-218, a food image dataset constructed from real-world meal records collected through the dietary management application FoodLog Athl. The dataset contains 6,925 images across 218 food categories, with a total of 14,349 bounding boxes. Rich metadata, including meal date and time, anonymized user IDs, and meal-level context, accompany each image. Unlike conventional datasets-where a predefined class set guides web-based image collection-our data begins with user-submitted photos, and labels are applied afterward. This yields greater intra-class diversity, a natural frequency distribution of meal types, and casual, unfiltered images intended for personal use rather than public sharing. In addition to (1) a standard classification benchmark, we introduce two FoodLog-specific tasks: (2) an incremental fine-tuning protocol that follows the temporal stream of users' logs, and (3) a context-aware classification task where each image contains multiple dishes, and the model must classify each dish by leveraging the overall meal context. We evaluate these tasks using large multimodal models (LMMs). The dataset is publicly available at https://huggingface.co/datasets/FoodLog/FoodLogAthl-218. Mitsuki Watanabe, Sosuke Amano, Kiyoharu Aizawa, Yoko Yamakata |
ACM Multimedia | 4 |
| 2025 | FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
Yuki Imajuku, Yoko Yamakata, Kiyoharu Aizawa |
MMM (1) | 2 |
| 2025 | Leveraging LLM for Detecting and Explaining LLM-generated Code in Python Programming Courses
Jeonghun Baek, Tetsuro Yamazaki, Akimasa Morihata, Junichiro Mori, Yoko Yamakata, Kenjiro Taura, Shigeru Chiba |
SIGCSE (2) | 5 |
| 2024 | Next Topic Recommendation for Influencers on Social MediaabstractTo maintain popularity on social media over the long term, users need to shift to a new topic instead of sticking to one topic. When selecting a new topic, a user needs to consider both its popularity on the entire social media and its popularity among the current followers. The former affects the expected number of new followers, and the latter affects the expected ratio of the current followers the user can retain after the topic change. The timing is also important. The user should change to a new topic before the current topic becomes less popular and the user loses many of the current followers. If the user change the topic after losing the followers, it is more difficult to obtain new followers. In this paper, we introduce a new task based on these observations: recommending appropriate new topics for currently popular social media users at appropriate timing. As an example of opportunities in the research on this task, we also propose a simple method of predicting the popularity a given user would gain after shifting to a given new topic. Our method predicts it based on the similarity between the user’s current topic and the given new topic. In our experiment with data collected from X (formerly Twitter), our method improves the prediction accuracy compared with a baseline method. Masafumi Iwanaga, Keishi Tajima, Yoko Yamakata |
IEEE Big Data | 3 |
| 2024 | Measure and Improve Your Food: Ingredient Estimation Based Nutrition Calculator
Yoko Yamakata, Ryoma Maeda, Kiyoharu Aizawa |
ACM Multimedia | 2 |
| 2024 | Adaptive Feature Inheritance and Thresholding for Ingredient Recognition in Multimedia Cooking Instructions
Yixin Zhang 0001, Yoko Yamakata, Keishi Tajima |
MMAsia | 2 |
| 2023 | Open-Vocabulary Segmentation Approach for Transformer-Based Food Nutrient EstimationabstractNutrition plays a vital role in overall health and well-being. With a highly accurate nutrient estimation model, we develop a tool that displays nutritional values from food images, thereby reducing the labor-intensiveness of dietary assessment. We propose a method that uses depth data with RGB images and incorporates an open-vocabulary segmentation process that separates food from non-food instances, coupled with two-stage self-attention Transformer decoder. Our model outperforms the current state-of-the-art method, with an average percent MAE of 17.2% on Nutrition5k, an RGB-D food image dataset with calories, mass, and three macronutrients annotated. Our study also focuses on the significance of the food and background regions for calorie, mass, and nutrient estimation. We analyze the impact of non-food regions on each estimation task, with results suggesting that background information is crucial for calorie, mass, and carbohydrate estimation but not as essential for protein and fat estimation. The qualitative results also show that the model attends to regions with a high corresponding nutritional value. Implementation codes and pre-trained models are provided at https://github.com/Oatsty/nutrition5k. Satayu Parinayok, Yoko Yamakata, Kiyoharu Aizawa |
MMAsia | 2 |
| 2023 | Automatic Dataset Creation from User-generated Recipes for Ingredient-centric Food Image AnalysisabstractWe aim to develop an application that automatically creates a nutrition facts label from food images for precise dietary control. Firstly, we constructed a new dataset with food category labels and a list of ingredients in a nutritionally calculable format using an image classification model and BERT for 1.6 million recipes accompanied by images. The nutritional value of the recipe can be calculated using a conversion table consisting of the food item number and unit class. Next, using deep learning techniques, we built models that estimate the list of food item numbers from food images. While the multi-task model that identifies the food category label and the ingredient list simultaneously is only effective within a limited number of recipes, the single-task model that only identified the ingredient list achieved a Micro-F1 of 53.32% in total. Yoko Yamakata, Kiyoharu Aizawa |
MMAsia | 2 |
| 2022 | CEA++'22: 1st International Workshop on Multimedia for Cooking, Eating, and related APPlicationsabstractThe International Workshop on Multimedia for Cooking, Eating, and related APPlications is the successor of the former CEA workshop series. The former CEA series started in 2009. 13 years later, we are witnessing various food-related applications enabled by emerging deep learning technologies and related hardware. Based on such a background, the organizing committee of CEA decided to renew the workshop and extend the scope to accept broader topics, especially industrial applications. Yoko Yamakata, Atsushi Hashimoto 0001, Jingjing Chen 0001 |
ACM Multimedia | 1 |
| 2022 | Recipe-oriented Food Logging for Nutritional ManagementabstractWe propose a recipe-oriented food logging method that records food by recipe, unlike the ordinary food logging method that records food by name. We also develop an application RecipeLog for this purpose. RecipeLog can create a "skeleton recipe," which is a standardized recipe representation suitable for estimating the nutritional value of a dish. This is represented by a list of ingredients linked to a Nutrition Facts table and a flow graph consisting of cooking actions such as cutting, mixing, baking, simmering, and frying the ingredients. The recipe log allows recipes to be written with fewer operations by editing only the differences from the already registered base recipe. Experiments have confirmed that the recipe log can effectively identify differences in recipes from household to household. The future work is to construct a multimedia recipe dataset consisting of structured recipes and their images using RecipeLog. Yoko Yamakata, Akihisa Ishino, Akiko Sunto, Sosuke Amano, Kiyoharu Aizawa |
ACM Multimedia | 1 |
| 2022 | FoodLog Athl: Multimedia Food Recording Platform for Dietary Guidance and Food MonitoringabstractThis paper presents a new food recording tool, FoodLog Athl, for the healthcare or physical enhancement of its users. Unlike existing food recording tools, we designed the system for dietitians or third parties who monitor the users. The tool not only supports the users by functions such as food image recognition, but also it helps the dietitians watch and communicate with users. Furthermore, it calculates nutritional values from food records - the use of the tool reduces the workload of dietitians and focuses their work on nutrition guidance. Kei Nakamoto, Kohei Kumazawa, Hiroaki Karasawa, Sosuke Amano, Yoko Yamakata, Kiyoharu Aizawa |
MMAsia | 5 |
| 2022 | Wearable Camera Based Food Logging SystemabstractRecently, meal management apps have allowed people to record food items and calories from photos automatically. These technologies include extracting food regions from photos of served meals, identifying the name of the food in each region, and calculating nutritional data. However, what you eat is not the only indicator that should be kept in the food record. How fast you eat and the order in which you eat is also significant information for dietary management. Therefore, we aim to construct a system that automatically generates a meal log from first-person videos that users capture of their eating behavior with a wearable camera. To tackle the complex problems that the data this system assumes contains, we constructed an eating behavior record dataset: 9.9 hours of first-person video that assume the natural diets of a user. To investigate the feasibility of our proposed system, we evaluated whether the first step, the detection of the meal area in the video during the meal, could be achieved with sufficient accuracy using this dataset. Using the limited number of frames assumed to be annotated by the user as training data, 30 frames were annotated for user-specific model training and four frames for online adaptation, resulting in detection accuracy of 72% for food regions. Our next goal is to create a multi-user dataset and service the application. Kenshiro Sato, Yoko Yamakata, Sosuke Amano, Kiyoharu Aizawa |
MMAsia | 2 |
| 2021 | Noisy Annotation Refinement for Object Detection
Jiafeng Mao, Qing Yu 0013, Yoko Yamakata, Kiyoharu Aizawa |
BMVC | 3 |
| 2021 | CEA'21: The 13th Workshop on Multimedia for Cooking and Eating ActivitiesabstractThe 13th Workshop on Multimedia for Cooking and Eating Activities presents This overview introduces the aim of the CEA'21 workshop and the list of papers presented in the workshop. Yoko Yamakata, Atsushi Hashimoto 0001 |
ICMR | 1 |
| 2021 | RecipeLog: Recipe Authoring App for Accurate Food RecordingabstractDiet management is usually conducted by recording the name of foods eaten, but in fact, the nutritional value of food in the same name varies greatly from recipe to recipe. To know accurate nutritional values of the foods, recording personal recipes is effective but time-consuming. Therefore, we are developing a mobile application "RecipeLog", that assists users to write their own recipes by modifying prepared ones. In our experiments, we show that with RecipeLog users create personal recipes with 45% less edit distance compared to writing from scratch. Akihisa Ishino, Yoko Yamakata, Hiroaki Karasawa, Kiyoharu Aizawa |
ACM Multimedia | 2 |
| 2021 | MIRecipe: A Recipe Dataset for Stage-Aware Recognition of Changes in Appearance of IngredientsabstractIn this paper, we introduce a new recipe dataset MIRecipe (Multimedia-Instructional Recipe). It has both text and image data for every cooking step, while the conventional recipe datasets only contain final dish images, and/or images only for some of the steps. It consists of 26,725 recipes, which include 239,973 steps in total. The recognition of ingredients in images associated with cooking steps poses a new challenge: Since ingredients are processed during cooking, the appearance of the same ingredient is very different in the beginning and finishing stages of the cooking. The general object recognition methods, which assume the constant appearance of objects, do not perform well for such objects. To solve the problem, we propose two stage-aware techniques: stage-wise model learning, which trains a separate model for each stage, and stage-aware curriculum learning, which starts with the training data from the beginning stage and proceeds to the later stages. Our experiment with our dataset shows that our method achieves higher accuracy than the model trained using all the data without considering the stages. Our dataset is available at our GitHub repository. Yixin Zhang 0001, Yoko Yamakata, Keishi Tajima |
MMAsia | 2 |
| 2020 | Visual Grounding Annotation of Recipe Flow GraphabstractIn this paper, we provide a dataset that gives visual grounding annotations to recipe flow graphs. A recipe flow graph is a representation of the cooking workflow, which is designed with the aim of understanding the workflow from natural language processing. Such a workflow will increase its value when grounded to real-world activities, and visual grounding is a way to do so. Visual grounding is provided as bounding boxes to image sequences of recipes, and each bounding box is linked to an element of the workflow. Because the workflows are also linked to the text, this annotation gives visual grounding with workflow’s contextual information between procedural text and visual observation in an indirect manner. We subsidiarily annotated two types of event attributes with each bounding box: “doing-the-action,” or “done-the-action”. As a result of the annotation, we got 2,300 bounding boxes in 272 flow graph recipes. Various experiments showed that the proposed dataset enables us to estimate contextual information described in recipe flow graphs from an image sequence. Taichi Nishimura, Suzushi Tomori, Hayato Hashimoto, Atsushi Hashimoto 0001, Yoko Yamakata, Jun Harashima, Yoshitaka Ushiku, Shinsuke Mori |
LREC | 5 |
| 2020 | English Recipe Flow Graph CorpusabstractWe present an annotated corpus of English cooking recipe procedures, and describe and evaluate computational methods for learning these annotations. The corpus consists of 300 recipes written by members of the public, which we have annotated with domain-specific linguistic and semantic structure. Each recipe is annotated with (1) ‘recipe named entities’ (r-NEs) specific to the recipe domain, and (2) a flow graph representing in detail the sequencing of steps, and interactions between cooking tools, food ingredients and the products of intermediate steps. For these two kinds of annotations, inter-annotator agreement ranges from 82.3 to 90.5 F1, indicating that our annotation scheme is appropriate and consistent. We experiment with producing these annotations automatically. For r-NE tagging we train a deep neural network NER tool; to compute flow graphs we train a dependency-style parsing procedure which we apply to the entire sequence of r-NEs in a recipe. In evaluations, our systems achieve 71.1 to 87.5 F1, demonstrating that our annotation scheme is learnable. Yoko Yamakata, Shinsuke Mori, John Carroll 0001 |
LREC | 1 |
| 2020 | CEA'20: The 12th Workshop on Multimedia for Cooking and Eating ActivitiesabstractThe 12th Workshop on Multimedia for Cooking and Eating Activities presents This overview introduces the aim of the CEA'20 workshop and the list of papers presented in the workshop. Ichiro Ide, Yoko Yamakata, Atsushi Hashimoto 0001 |
ICMR | 2 |
| 2018 | A Case Study on Start-up of Dataset Construction: In Case of Recipe Named Entity CorpusabstractIn this paper, we report our experience in constructing a cooking recipe text corpus. We describe problems we found and explain how we managed them. One of the problems we faced in the construction of our recipe corpus is the difficulty of establishing a clear, stable, and complete guideline instructing annotators how to annotate. During the annotation, we found many unexpected cases for which the pre-defined guideline is not clear enough, and even cases for which the pre-defined guideline provides no guidance at all. As a result, we needed to update the guideline twice during the annotation, and also needed to revise annotations we have done before the updates. During that process, we have several trade-offs, and it is not easy to decide when and how often we should revise the annotations. It is even unclear whether we should revise them or should instead use the human resource for annotating more data. We show an experiment, whose result suggests that we should revise the old annotations. Another problem we had is the management of versions of the guideline, sets of annotations corresponding to them, and communication between participants. Yoko Yamakata, Keishi Tajima, Shinsuke Mori |
IEEE BigData | 1 |
| 2014 | FlowGraph2Text: Automatic Sentence Skeleton Compilation for Procedural Text GenerationabstractIn this paper we describe a method for generating a procedural text given its flow graph representation. Our main idea is to automatically collect sen-tence skeletons from real texts by re-placing the important word sequences with their type labels to form a skeleton pool. The experimental results showed that our method is feasible and has a potential to generate natural sentences. 1 Shinsuke Mori, Hirokuni Maeta, Tetsuro Sasada, Koichiro Yoshino, Atsushi Hashimoto 0001, Takuya Funatomi, Yoko Yamakata |
INLG | 7 |
| 2014 | Flow Graph Corpus from Recipe Texts
Shinsuke Mori, Hirokuni Maeta, Yoko Yamakata, Tetsuro Sasada |
LREC | 3 |
| 2013 | Workshop summary for the 5th international workshop on multimedia for cooking and eating activities (CEA'13)abstractThis summary introduces the aim of the CEA'13 workshop and the list of papers presented in the workshop. Kiyoharu Aizawa, Yoko Yamakata, Takuya Funatomi |
ACM Multimedia | 2 |
| 2012 | Overview of the ACM multimedia 2012 workshop on multimedia for cooking and eating activities (CEA'12)abstractThis overview introduces the aim of the CEA'12 workshop and the list of papers presented in the workshop. Mutsuo Sano, Ichiro Ide, Yoko Yamakata |
ACM Multimedia | 3 |
| 2011 | Cooking Ingredient Recognition Based on the Load on a Chopping Board during CuttingabstractThis paper presents a method for recognizing recipe ingredients based on the load on a chopping board when ingredients are cut. The load is measured by four sensors attached to the board. Each chop is detected by indentifying a sharp falling edge in the load data. The load features, including the maximum value, duration, impulse, peak position, and kurtosis, are extracted and used for ingredient recognition. Experimental results showed a precision of 98.1% in chop detection and 67.4% in ingredient recognition with a support vector machine (SVM) classifier for 16 common ingredients. Yoko Yamakata, Yoshiki Tsuchimoto, Atsushi Hashimoto 0001, Takuya Funatomi, Mayumi Ueda, Michihiko Minoh |
ISM | 1 |
| 2010 | Object Recognition Based on Object's Identity for Cooking Recognition TaskabstractIn this paper, we address a novel task "cooking recognition task". Cooking recognition task is to keep recognizing a cooking action of a cook and the target food product of the action at any time of the cooking. The target object of the recognition could change its visual feature completely by peeling, cutting, and mixed with others during the task. Nonetheless, such food product is called by the name that is given at the beginning of the cooking. Because the existing image recognition methods recognizes an object based on such idea that objects of the same name have similar visual feature, such methods cannot solve this cooking recognition task. Therefore, we propose new methods that recognize an object based on such idea that object keeps its identity even after manufactured and/or mixed with others. Yoko Yamakata, Koh Kakusho, Michihiko Minoh |
ISM | 1 |
| 2009 | Design of large planar diaphragm incorporating multiple vibrators for sound directivity control via FEM and BEMabstractWe have realized a sound directivity control of a large planar diaphragm by controlling the bending vibrations using multiple vibrators. This paper proposes a method to determine the parameters such as the shape and thickness of the diaphragm and the position of the vibrators. In this method, the finite element method (FEM) is used to simulate the diaphragm vibrations and the boundary element method (BEM) is used to simulate the radiated sound. We first validate the simulation results by measuring the actual bending vibrations and sound directivity and comparing them with the simulated results. We then show that metaheuristics are effective in finding appropriate parameters because the sound directivity largely varies in a continuous manner with variations in the parameters. Yoko Yamakata, Michiaki Katsumoto, Toshiyuki Kimura |
ICASSP | 1 |
| 2008 | Directional sound radiation system using a large planar diaphragm incorporating multiple vibratorsabstractThis paper aims to construct a system that produces directional sound radiation. Directional sound is intrinsically radiated from a vibrating resonant body of such instrument as a violin. The directivity is said to give the sound a realistic and spatial effect. To reproduce a sound with such directivity, we propose a method that uses multiple vibrators to artificially induce bending vibrations on a large planar diaphragm. As the first step of this study, we constructed a prototype system and demonstrated that (i) the bending vibration of the diaphragm is controllable by adjusting accelerated vibrations and (ii) the radiated sound obtains directivity as a specified condition by a user using the algorithm we proposed. This directivity of the radiated sound was obvious enough for humans to perceive. Yoko Yamakata, Michiaki Katsumoto, Toshiyuki Kimura |
ICASSP | 1 |
| 2007 | Inference by aggregation of evidence with applications to fuzzy probabilities
Anca L. Ralescu, Dan A. Ralescu, Yoko Yamakata |
Inf. Sci. | 3 |
| 2003 | Toward the Human Communication Efficiency Monitoring from Captured Audio and Video Media in Real Environments
Tomasz M. Rutkowski, Susumu Seki, Yoko Yamakata, Koh Kakusho, Michihiko Minoh |
KES | 3 |
| 2002 | Belief network based disambiguation of object reference in spoken dialogue system for robotabstractWe are studying joint activity in which a remote robot finds an object by communicating with the user over a voice-only channel. We focus on how the robot disambiguates the reference of the uttered word or phrase to the target object. For example, by “cup”, one may refer to a “teacup”, a “coffee cup”, or even a “glass” under some situations. This reference (hereafter, “object reference”) is user-dependent. We confirm that a user model of object references is significant by conducting a survey of 12 subjects. In addition to ambiguity of object reference, actual systems should cope with two other sources of uncertainty in speech and image recognition. We present a Belief Network based probabilistic reasoning system to determine the object reference. The resulting system demonstrates that the number of interactions needed to find a common reference is reduced as the user model is refined. Yoko Yamakata, Tatsuya Kawahara, Hiroshi G. Okuno |
INTERSPEECH | 1 |