EDBT 2026 Demo / reviewers in the wild / expert
Atsushi Hashimoto 0001
dblp:89/6733-1
· DBLP profile ↗
36ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0002-0799-4269ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 19 · 11 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Safer Mobile Agents: Scalable Generation and Evaluation of Diverse Scenarios for VLMs
Takara Taniguchi, Kuniaki Saito, Atsushi Hashimoto 0001 |
ICPR (1) | 3 |
| 2026 | Evaluating the Capability of Video Question Generation for Expert Knowledge ElicitationabstractSkilled human interviewers can extract valuable information from experts. This raises a fundamental question: what makes some questions more effective than others? To address this, a quantitative evaluation of question-generation models is essential. Video question generation (VQG) is a topic for video question answering (VideoQA), where questions are generated for given answers. Their evaluation typically focuses on the ability to answer questions, rather than the quality of generated questions. In contrast, we focus on the question quality in eliciting unseen knowledge from human experts. For a continuous improvement of VQG models, we propose a protocol that evaluates the ability by simulating question-answering communication with experts using a question-to-answer retrieval. We obtain the retriever by constructing a novel dataset, EgoExoAsk, which comprises 27,666 QA pairs generated from Ego-Exo4D’s expert commentary annotation. The EgoExoAsk training set is used to obtain the retriever, and the benchmark is constructed on the validation set with Ego-Exo4D video segments. Experimental results demonstrate our metric reasonably aligns with question generation settings: models accessing richer context are evaluated better, supporting that our protocol works as intended. The EgoExoAsk dataset is available in our project page. Huaying Zhang, Atsushi Hashimoto 0001, Tosho Hirasawa |
WACV | 2 |
| 2025 | CaptionSmiths: Flexibly Controlling Language Pattern in Image CaptioningabstractAn image captioning model flexibly switching its language pattern, e.g., descriptiveness and length, should be useful since it can be applied to diverse applications. However, despite the dramatic improvement in generative vision-language models, fine-grained control over the properties of generated captions is not easy due to two reasons: (i) existing models are not given the properties as a condition during training and (ii) existing models cannot smoothly transition its language pattern from one state to the other. Given this challenge, we propose a new approach, CaptionSmiths, to acquire a single captioning model that can handle diverse language patterns. First, our approach quantifies three properties of each caption, length, descriptiveness, and uniqueness of a word, as continuous scalar values, without human annotation. Given the values, we represent the conditioning via interpolation between two endpoint vectors corresponding to the extreme states, e.g., one for a very short caption and one for a very long caption. Empirical results demonstrate that the resulting model can smoothly change the properties of the output captions and show higher lexical alignment than baselines. For instance, CaptionSmiths reduces the error in controlling caption length by 506\% despite better lexical alignment. Code will be available on https://github.com/omron-sinicx/captionsmiths. Kuniaki Saito, Donghyun Kim 0006, Kwanyong Park, Atsushi Hashimoto 0001, Yoshitaka Ushiku |
ICCV | 4 |
| 2025 | Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
Takehiko Ohkawa, Takuma Yagi, Taichi Nishimura, Ryosuke Furuta, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Yoichi Sato 0001 |
WACV | 5 |
| 2025 | Poet-Weaver: Reflecting on Communication Failure in Personal Relationships With Stylized AI-Generated Conversation DigestsabstractInterpersonal communication often involves navigating social challenges like managing expectations and conveying intentions. In practical contexts like productivity and navigation of social boundaries, agent- and AI-mediated communication (AIMC) has served as an effective social intermediary. Communication within personal relationships, especially between intercultural friends from different backgrounds, can also face significant challenges, such as miscommunication and expressive suppression. However, AIMC remains underutilized in the context of established personal relationships due to ethical concerns about agency and authenticity. We propose designing openly-interpretable AIMC output as an augmented context cue to reduce its active social involvement, balancing AI support with preservation of relational agency. We developed and evaluated Poet-Weaver, a Discord text chat plugin that presents AI-generated insights on user conversations in a stylized, openly-interpretable way to encourage reflection on communication challenges. We conducted a mixed-methods study with 30 intercultural friend pairs to assess Poet-Weaver's impact. Findings showed Poet-Weaver effectively helped participants address communication failures, although AIMC still influenced users' behavior even with openly-interpretable output. We recommend future uses of AIMC that support transformative, positive relationship changes while preserving individual responsibility and identity within relationships. Seraphina Yong, Chi-Lan Yang, Atsushi Hashimoto 0001, Hideaki Kuzuoka, Shigeo Yoshida |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | Learning 3D Point Cloud Registration as a Single Optimization Problem
Rintaro Yanagi, Atsushi Hashimoto 0001, Naoya Chiba, Shusaku Sone, Yoshitaka Ushiku |
ACCV (9) | 2 |
| 2024 | COM Kitchens: An Unedited Overhead-View Video Dataset as a Vision-Language Benchmark
Koki Maeda, Tosho Hirasawa, Atsushi Hashimoto 0001, Jun Harashima, Leszek Rybicki, Yusuke Fukasawa, Yoshitaka Ushiku |
ECCV (65) | 3 |
| 2024 | PolarDB: Formula-Driven Dataset for Pre-Training Trajectory EncodersabstractFormula-driven supervised learning (FDSL) is a growing research topic for finding simple mathematical formulas that generate synthetic data and labels for pre-training neural networks. The main advantage of FDSL is that there is no risk of generating data with ethical implications such as gender bias and racial bias because it does not rely on real data as discussed in previous studies using fractals and polygons for pre-training image encoders. While FDSL has been proposed for pre-training image encoders, it has not been considered for temporal trajectory data. In this paper, we introduce PolarDB, the first formula-driven dataset for pre-training trajectory encoders with an application to fine-grained cutting-method recognition using hand trajectories. More specifically, we generate 270k trajectories for 432 categories on the basis of polar equations and use them to pre-train a Transformer-based trajectory encoder in an FDSL manner. In the experiments, we show that pre-training on PolarDB improves the accuracy of fine-grained cutting-method recognition on cooking videos of EPIC-KITCHEN and Ego4D datasets, where the pre-trained trajectory encoder is used as a plug-in module for a video recognition network. Sota Miyamoto, Takuma Yagi, Yuto Makimoto, Mahiro Ukai, Yoshitaka Ushiku, Atsushi Hashimoto 0001, Nakamasa Inoue |
ICASSP | 6 |
| 2024 | Vision-Language Interpreter for Robot Task PlanningabstractLarge language models (LLMs) are accelerating the development of language-guided robot planners. Meanwhile, symbolic planners offer the advantage of interpretability. This paper proposes a new task that bridges these two trends, namely, multimodal planning problem specification. The aim is to generate a problem description (PD), a machine-readable file used by the planners to find a plan. By generating PDs from language instruction and scene observation, we can drive symbolic planners in a language-guided framework. We propose a Vision-Language Interpreter (ViLaIn), a new framework that generates PDs using state-of-the-art LLM and vision-language models. ViLaIn can refine generated PDs via error message feedback from the symbolic planner. Our aim is to answer the question: How accurately can ViLaIn and the symbolic planner generate valid robot plans? To evaluate ViLaIn, we introduce a novel dataset called the problem description generation (ProDG) dataset. The framework is evaluated with four new evaluation metrics. Experimental results show that ViLaIn can generate syntactically correct problems with more than 99% accuracy and valid plans with more than 58% accuracy. Our code and dataset are available at https://github.com/omron-sinicx/ViLaIn. Keisuke Shirai, Cristian C. Beltran-Hernandez, Masashi Hamaya, Atsushi Hashimoto 0001, Shohei Tanaka, Kento Kawaharazuka, Kazutoshi Tanaka, Yoshitaka Ushiku, Shinsuke Mori |
ICRA | 4 |
| 2024 | Visuo-Tactile Zero-Shot Object Recognition with Vision-Language ModelabstractTactile perception is vital, especially when distinguishing visually similar objects. We propose an approach to incorporate tactile data into a Vision-Language Model (VLM) for visuo-tactile zero-shot object recognition. Our approach leverages the zero-shot capability of VLMs to infer tactile properties from the names of tactilely similar objects. The proposed method translates tactile data into a textual description solely by annotating object names for each tactile sequence during training, making it adaptable to various contexts with low training costs. The proposed method was evaluated on the FoodReplica and Cube datasets, demonstrating its effectiveness in recognizing objects that are difficult to distinguish by vision alone. Shiori Ueda, Atsushi Hashimoto 0001, Masashi Hamaya, Kazutoshi Tanaka, Hideo Saito 0001 |
IROS | 2 |
| 2024 | AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question AnsweringabstractVisual question answering aims to provide responses to questions given visual input. Recently, visual programmatic models (VPMs), which generate programs to answer questions through large language models (LLMs), have attracted attention. However, they often require long input prompts to provide the LLM with sufficient API usage details to generate relevant code. To address this limitation, we propose AdaCoder, an adaptive prompt compression framework for VPMs. AdaCoder operates in two phases: a compression phase and an inference phase. In the compression phase, given a preprompt that describes all API definitions with example code snippets, a set of compressed preprompts is generated, each depending on a specific question type. In the inference phase, AdaCoder predicts the question type and chooses the appropriate corresponding compressed preprompt to generate code to answer the question. In experiments, we apply AdaCoder to ViperGPT and demonstrate that it reduces token length by 71.1%, while maintaining or even improving the performance of visual question answering. Mahiro Ukai, Shuhei Kurita, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Nakamasa Inoue |
ACM Multimedia | 3 |
| 2023 | Invertible Conditional GAN Revisited: Photo-to-Manga Face Translation with Modern Architectures (Student Abstract)abstractRecent style translation methods have extended their transferability from texture to geometry. However, performing translation while preserving image content when there is a significant style difference is still an open problem. To overcome this problem, we propose Invertible Conditional Fast GAN (IcFGAN) based on GAN inversion and cFGAN. It allows for unpaired photo-to-manga face translation. Experimental results show that our method could translate styles under significant style gaps, while the state-of-the-art methods could hardly preserve image content. Taro Hatakeyama, Ryusuke Saito, Komei Hiruta, Atsushi Hashimoto 0001, Satoshi Kurihara |
AAAI | 4 |
| 2023 | Learning Food Picking without Food: Fracture Anticipation by Breaking Reusable Fragile ObjectsabstractFood picking is trivial for humans but not for robots, as foods are fragile. Presetting foods' physical properties does not help robots much due to the objects' inter- and intra-category diversity. A recent study proved that learning-based fracture anticipation with tactile sensors could overcome this problem; however, the method trains the model for each food to deal with intra-category differences, and tuning robots for each food leads to an undesirable amount of food consumption. This study proposes a novel framework for learning food-picking tasks without consuming foods. The key idea is to leverage the object-breaking experiences of several reusable fragile objects instead of consuming real foods while making the picking ability object-invariant with domain generalization (DG). In real-robot experiments, we trained a model with reusable objects (toy blocks, ping-pong balls, and jellies), selected based on the three common fracture types (crack, rupture, and crush). We then tested the model with four real food objects (tofu, bananas, potato chips, and tomatoes). The results showed that the proposed combination of reusable objects' breaking experiences and DG is effective for the food-picking task. Rinto Yagawa, Reina Ishikawa, Masashi Hamaya, Kazutoshi Tanaka, Atsushi Hashimoto 0001, Hideo Saito 0001 |
ICRA | 5 |
| 2023 | Reference-based Dense Pose Estimation via Partial 3D Point Cloud MatchingabstractInteracting with real-world objects is one of the fundamental tasks in multimedia. Despite its importance, existing object pose estimation targets only rigid objects. This demonstration proposes a novel application for non-rigid object pose estimation. Inspired by human dense pose estimation, we represent a pose of a non-rigid object as an indexed point cloud, where each index corresponds to that in a template. The correspondence is identified by a machine-learning-based 3D point cloud matching. Finding correspondence to the template point cloud enables a dense pose estimation with no object-specific learning processes. In the demonstration, we visualize the correspondence of points in observed depth images and the template. We also provide a demonstration of template point cloud reconstruction. Through these systems, onsite visitors can test our system with objects brought by themselves and have an experience with a state-of-the-art 3D point cloud matching method as well as this novel task. Rintaro Yanagi, Atsushi Hashimoto 0001, Naoya Chiba, Yoshitaka Ushiku |
ACM Multimedia | 2 |
| 2023 | State-aware video procedural captioning
Taichi Nishimura, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Hirotaka Kameko, Shinsuke Mori |
Multim. Tools Appl. | 2 |
| 2022 | Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe FlowsabstractWe present a new multimodal dataset called Visual Recipe Flow, which enables us to learn a cooking action result for each object in a recipe text. The dataset consists of object state changes and the workflow of the recipe text. The state change is represented as an image pair, while the workflow is represented as a recipe flow graph. We developed a web interface to reduce human annotation costs. The dataset allows us to try various applications, including multimodal information retrieval. Keisuke Shirai, Atsushi Hashimoto 0001, Taichi Nishimura, Hirotaka Kameko, Shuhei Kurita, Yoshitaka Ushiku, Shinsuke Mori |
COLING | 2 |
| 2022 | Conditional GAN for Small DatasetsabstractGenerating high-quality images with Generative Adversarial Networks (GANs) generally requires 100k+ training data. The required data amount is too large when we consider using GANs to support professional art creators; they need to follow the specific art style while interactively controlling the results along with their theme. This research proposes Conditional FastGAN, which adds a condition vector to FastGAN to produce high-quality different domain images even on small datasets. In our experiments, the MUCT Face Database of images consisting of face photos in various orientations and manga face images extracted from Osamu Tezuka’s works were used as a small-scale dataset. Fine-tuning with manga face images to a model pre-trained with photo-only face images enabled control of the generated images according to explicit conditions, such as photos and manga, for the same latent variables. In addition, the proposed method improved the FID score by 2.55 from the original FastGAN in the case of manga face generation. Komei Hiruta, Ryusuke Saito, Taro Hatakeyama, Atsushi Hashimoto 0001, Satoshi Kurihara |
ISM | 4 |
| 2022 | CEA++'22: 1st International Workshop on Multimedia for Cooking, Eating, and related APPlicationsabstractThe International Workshop on Multimedia for Cooking, Eating, and related APPlications is the successor of the former CEA workshop series. The former CEA series started in 2009. 13 years later, we are witnessing various food-related applications enabled by emerging deep learning technologies and related hardware. Based on such a background, the organizing committee of CEA decided to renew the workshop and extend the scope to accept broader topics, especially industrial applications. Yoko Yamakata, Atsushi Hashimoto 0001, Jingjing Chen 0001 |
ACM Multimedia | 2 |
| 2021 | Divergence Optimization for Noisy Universal Domain AdaptationabstractUniversal domain adaptation (UniDA) has been proposed to transfer knowledge learned from a label-rich source domain to a label-scarce target domain without any constraints on the label sets. In practice, however, it is difficult to obtain a large amount of perfectly clean labeled data in a source domain with limited resources. Existing UniDA methods rely on source samples with correct annotations, which greatly limits their application in the real world. Hence, we consider a new realistic setting called Noisy UniDA, in which classifiers are trained with noisy labeled data from the source domain and unlabeled data with an unknown class distribution from the target domain. This paper introduces a two-head convolutional neural network framework to solve all problems simultaneously. Our network consists of one common feature generator and two classifiers with different decision boundaries. By optimizing the divergence between the two classifiers’ outputs, we can detect noisy source samples, find "unknown" classes in the target domain, and align the distribution of the source and target domains. In an extensive evaluation of different domain adaptation settings, the proposed method outperformed existing methods by a large margin in most settings. Qing Yu 0013, Atsushi Hashimoto 0001, Yoshitaka Ushiku |
CVPR | 2 |
| 2021 | Removing Word-Level Spurious Alignment between Images and Pseudo-Captions in Unsupervised Image CaptioningabstractUkyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto, Taro Watanabe, Yuji Matsumoto. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Ukyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto 0001, Taro Watanabe, Yuji Matsumoto 0001 |
EACL | 3 |
| 2021 | CEA'21: The 13th Workshop on Multimedia for Cooking and Eating ActivitiesabstractThe 13th Workshop on Multimedia for Cooking and Eating Activities presents This overview introduces the aim of the CEA'21 workshop and the list of papers presented in the workshop. Yoko Yamakata, Atsushi Hashimoto 0001 |
ICMR | 2 |
| 2021 | State-aware Video Procedural CaptioningabstractVideo procedural captioning (VPC), which generates procedural text from instructional videos, is an essential task for scene understanding and real-world applications. The main challenge of VPC is to describe how to manipulate materials accurately. This paper focuses on this challenge by designing a new VPC task, generating a procedural text from the clip sequence of an instructional video and material list. In this task, the state of materials is sequentially changed by manipulations, yielding their state-aware visual representations (e.g., eggs are transformed into cracked, stirred, then fried forms). The essential difficulty is to convert such visual representations into textual representations; that is, a model should track the material states after manipulations to better associate the cross-modal relations. To achieve this, we propose a novel VPC method, which modifies an existing textual simulator for tracking material states as a visual simulator and incorporates it into a video captioning model. Our experimental results show the effectiveness of the proposed method, which outperforms state-of-the-art video captioning models. We further analyze the learned embedding of materials to demonstrate that the simulators capture their state transition. The code and dataset are available from https://github.com/misogil0116/svpc Taichi Nishimura, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Hirotaka Kameko, Shinsuke Mori |
ACM Multimedia | 2 |
| 2020 | Partially-Shared Variational Auto-encoders for Unsupervised Domain Adaptation with Target Shift
Ryuhei Takahashi, Atsushi Hashimoto 0001, Motoharu Sonogashira, Masaaki Iiyama |
ECCV (16) | 2 |
| 2020 | Visual Grounding Annotation of Recipe Flow GraphabstractIn this paper, we provide a dataset that gives visual grounding annotations to recipe flow graphs. A recipe flow graph is a representation of the cooking workflow, which is designed with the aim of understanding the workflow from natural language processing. Such a workflow will increase its value when grounded to real-world activities, and visual grounding is a way to do so. Visual grounding is provided as bounding boxes to image sequences of recipes, and each bounding box is linked to an element of the workflow. Because the workflows are also linked to the text, this annotation gives visual grounding with workflow’s contextual information between procedural text and visual observation in an indirect manner. We subsidiarily annotated two types of event attributes with each bounding box: “doing-the-action,” or “done-the-action”. As a result of the annotation, we got 2,300 bounding boxes in 272 flow graph recipes. Various experiments showed that the proposed dataset enables us to estimate contextual information described in recipe flow graphs from an image sequence. Taichi Nishimura, Suzushi Tomori, Hayato Hashimoto, Atsushi Hashimoto 0001, Yoko Yamakata, Jun Harashima, Yoshitaka Ushiku, Shinsuke Mori |
LREC | 4 |
| 2020 | CEA'20: The 12th Workshop on Multimedia for Cooking and Eating ActivitiesabstractThe 12th Workshop on Multimedia for Cooking and Eating Activities presents This overview introduces the aim of the CEA'20 workshop and the list of papers presented in the workshop. Ichiro Ide, Yoko Yamakata, Atsushi Hashimoto 0001 |
ICMR | 3 |
| 2020 | Photometric Stereo in Participating Media Using an Analytical Solution for Shape-Dependent Forward ScatterabstractImages captured in participating media such as murky water, fog, or smoke are degraded by scattered light. Thus, the use of traditional three-dimensional (3D) reconstruction techniques in such environments is difficult. In this paper, we propose a photometric stereo method for participating media. The proposed method differs from prvious studies with respect to modeling shape-dependent forward scatter. In the proposed model, forward scatter is described as an analytical form using lookup tables and is represented by spatially-variant kernels. We also propose an approximation of a large-scale dense matrix as a sparse matrix, which enables the removal of forward scatter. We discuss the approximation in the proposed method using synthesized data. Then, experiments with real data demonstrate that the proposed method improves 3D reconstruction in participating media. Yuki Fujimura, Masaaki Iiyama, Atsushi Hashimoto 0001, Michihiko Minoh |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Procedural Text Generation from a Photo SequenceabstractMultimedia procedural texts, such as instructions and manuals with pictures, support people to share how-to knowledge.In this paper, we propose a method for generating a procedural text given a photo sequence allowing users to obtain a multimedia procedural text.We propose a single embedding space both for image and text enabling to interconnect them and to select appropriate words to describe a photo.We implemented our method and tested it on cooking instructions, i.e., recipes.Various experimental results showed that our method outperforms standard baselines. Taichi Nishimura, Atsushi Hashimoto 0001, Shinsuke Mori |
INLG | 2 |
| 2018 | Photometric Stereo in Participating Media Considering Shape-Dependent Forward ScatterabstractImages captured in participating media such as murky water, fog, or smoke are degraded by scattered light. Thus, the use of traditional three-dimensional (3D) reconstruction techniques in such environments is difficult. In this paper, we propose a photometric stereo method for participating media. The proposed method differs from previous studies with respect to modeling shape-dependent forward scatter. In the proposed model, forward scatter is described as an analytical form using lookup tables and is represented by spatially-variant kernels. We also propose an approximation of a large-scale dense matrix as a sparse matrix, which enables the removal of forward scatter. Experiments with real and synthesized data demonstrate that the proposed method improves 3D reconstruction in participating media. Yuki Fujimura, Masaaki Iiyama, Atsushi Hashimoto 0001, Michihiko Minoh |
CVPR | 3 |
| 2018 | Restoration of Sea Surface Temperature Satellite Images Using a Partially Occluded Training SetabstractSea surface temperature(SST) satellite images are often partially occluded by clouds. Image inpainting is one approach to restore the occluded region. Considering the sparseness of SST images, they can be restored via learning-based inpainting. However, state-of-the-art learning-based inpainting methods using deep neural networks require large amount of non-occluded images as a training set. Since most SST images contain occluded regions, it is hard to collect sufficient non-occluded images. In this paper, we propose a novel method that uses occluded images as training images hence we can enlarge the amount of available training images from a certain SST image set. This is realized by comprising a novel reconstruction loss and adversarial loss. Experimental results confirm the effectiveness of our method. Satoki Shibata, Masaaki Iiyama, Atsushi Hashimoto 0001, Michihiko Minoh |
ICPR | 3 |
| 2017 | Restoration of sea surface temperature images by learning-based and optical-flow-based inpaintingabstractSea surface temperature (SST) images taken from satellites are partially occluded by clouds. In this paper, we propose an inpainting approach for restoration of the partially occluded images. Assuming the sparseness of the SST images, we employ a learning based inpainting for filling the occluded parts. Images taken in the past several days is another clue for filling the occluded parts. These images are regarded as time series data and a video inpainting method is also available. We employ PCA-based inpainting as a learning-based approach and optical-flow-based inpainting as video inpainting, and combine the two restored images according to the expected their restoration error. Experimental results with real satellite images show the effectiveness of our method. Satoki Shibata, Masaaki Iiyama, Atsushi Hashimoto 0001, Michihiko Minoh |
ICME | 3 |
| 2017 | Procedural Text Generation from an Execution VideoabstractIn recent years, there has been a surge of interest in automatically describing images or videos in a natural language. These descriptions are useful for image/video search, etc. In this paper, we focus on procedure execution videos, in which a human makes or repairs something and propose a method for generating procedural texts from them. Since video/text pairs available are limited in size, the direct application of end-to-end deep learning is not feasible. Thus we propose to train Faster R-CNN network for object recognition and LSTM for text generation and combine them at run time. We took pairs of recipe and cooking video, generated a recipe from a video, and compared it with the original recipe. The experimental results showed that our method can produce a recipe as accurate as the state-of-the-art scene descriptions. Atsushi Ushiku, Hayato Hashimoto, Atsushi Hashimoto 0001, Shinsuke Mori |
IJCNLP(1) | 3 |
| 2016 | Intention-Sensing Recipe Guidance via User Accessing to ObjectsabstractSensing the intention of a user’s forthcoming action is a necessary function for systems that assist human physical activity. In this article, a strategy for recipe guidance systems that can predict the forthcoming intended subtask in a cooking task is investigated. The focus is on user accessing objects, that is, touching and releasing objects. Touching can indicate the start of the forthcoming subtask and releasing can indicate the end of the task. The main difficulty lies in the fact that humans may move objects because they are in the way and use cooking tools that are unanticipated by an assistive system. In such cases, the accessed object should not indicate the forthcoming subtask. A method is proposed to track the progress of a task based on the object access history. This enables to eliminate object accesses that are out of context. Simultaneously, the method predicts the forthcoming subtask based on a combination of progress and materials rather than tools and materials. Then, a guidance system that runs as a web service is developed. In experiments, real cooking activities navigated by this system are observed. The Wizard of OZ method is utilized to simulate a system that detects object accesses. The experimental results show that 73.6% accuracy is achieved in the selection of the displayed information. This result supports the use of “access to objects” realize effective intention-sensing systems. Atsushi Hashimoto 0001, Jin Inoue, Takuya Funatomi, Michihiko Minoh |
Int. J. Hum. Comput. Interact. | 1 |
| 2014 | FlowGraph2Text: Automatic Sentence Skeleton Compilation for Procedural Text GenerationabstractIn this paper we describe a method for generating a procedural text given its flow graph representation. Our main idea is to automatically collect sen-tence skeletons from real texts by re-placing the important word sequences with their type labels to form a skeleton pool. The experimental results showed that our method is feasible and has a potential to generate natural sentences. 1 Shinsuke Mori, Hirokuni Maeta, Tetsuro Sasada, Koichiro Yoshino, Atsushi Hashimoto 0001, Takuya Funatomi, Yoko Yamakata |
INLG | 5 |
| 2011 | Developing a Real-Time System for Measuring the Consumption of SeasoningabstractIn this paper, we propose a real-time system for measuring the consumption of various types of seasonings. In our system, all seasonings are placed on a scale, and we continuously take images of these items using a camera. Our system estimates the consumption of each condiment by calculating the difference between the weight when the seasoning was picked up and the weight when it was placed back on the scale. Our system identifies the type of seasoning that was used by determining whether or not the seasoning was present on the scale. By using our system, users can automatically log their usage of seasoning. Then, they can adjust the seasoning according to their desired taste. Mayumi Ueda, Takuya Funatomi, Atsushi Hashimoto 0001, Takahiro Watanabe, Michihiko Minoh |
ISM | 3 |
| 2011 | Cooking Ingredient Recognition Based on the Load on a Chopping Board during CuttingabstractThis paper presents a method for recognizing recipe ingredients based on the load on a chopping board when ingredients are cut. The load is measured by four sensors attached to the board. Each chop is detected by indentifying a sharp falling edge in the load data. The load features, including the maximum value, duration, impulse, peak position, and kurtosis, are extracted and used for ingredient recognition. Experimental results showed a precision of 98.1% in chop detection and 67.4% in ingredient recognition with a support vector machine (SVM) classifier for 16 common ingredients. Yoko Yamakata, Yoshiki Tsuchimoto, Atsushi Hashimoto 0001, Takuya Funatomi, Mayumi Ueda, Michihiko Minoh |
ISM | 3 |
| 2010 | Tracking Food Materials with Changing Their Appearance in Food PreparingabstractThis paper describes our work in computer vision to track food materials in the food preparation process. Tracking such food materials is difficult, because they are often hidden when moved by hand. Furthermore, their appearance may change in hand when they are cut or peeled. For tracking these objects in such situations, we propose a novel method that matches an object on a cooking table to one grasped in the past. We use the following three criteria to match the objects even when they are cut or peeled: the similarity in their appearance, the validity of their change in appearance, and the grasped order. We experimentally evaluated our method by applying it to the scenes of cutting and peeling food materials. As a result, we achieved an accuracy of 83.6% in matching the objects. Atsushi Hashimoto 0001, Naoyuki Mori, Takuya Funatomi, Masayuki Mukunoki, Koh Kakusho, Michihiko Minoh |
ISM | 1 |