EDBT 2026 Demo / reviewers in the wild / expert
Kiyoharu Aizawa
dblp:71/5426
· DBLP profile ↗
13ranked-venue papers in the field
0as first author
7since 2021 · last 2026
0000-0003-2146-6275ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 6Other / Interdisciplinary · 6Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Realistic Virtual Flood Experience System Using 360° Videos and 3D City Models Constructed from Building FootprintsabstractVirtual flood experience systems, which enable users to vividly experience flooding, are attracting increasing attention as effective tools for communicating flood risks. However, existing systems typically rely on virtual cities that do not correspond to real locations and often lack sufficient photorealism, limiting users’ ability to relate scenarios to their own surroundings. Although 360° video-based virtual environments offer a simple and scalable way to visually replicate real-world scenes, effective 3D flood visualization in these environments typically requires 3D building geometry of the target area, which is not readily available in many regions. To address this limitation, we propose a new virtual flood experience framework that integrates 360° videos with 3D models automatically constructed from widely available 2D building footprints. By extruding footprints to plausible heights and spatially aligning the constructed models with 360° videos, our framework enables 3D flood visualization in photorealistic environments without relying on pre-existing city models such as CityGML. We demonstrate the framework in Memuro, Hokkaido, Japan, an area vulnerable to river flooding. A user study with local residents showed that the proposed system enhances users’ ability to envision location-specific flood evacuation situations, demonstrating its potential as an effective tool for disaster risk communication and education. Tatsuro Banno, Koki Kawada, Mizuki Takenawa, Masatoshi Denda, Kiyoharu Aizawa |
ICMR | 5 |
| 2023 | Text-to-Image Fashion Retrieval with Fabric TexturesabstractIn this study, we proposed text-to-image fashion image retrieval that captures the texture of clothing fabrics. A fabric’s texture is a major factor governing the comfort and appearance of clothes and significantly influences user preferences. However, unlike patterns and shapes that can readily be captured from a global image of the entire piece of clothing, extracting the fine and ambiguous characteristics of textures is considerably more challenging. The key concept is that by focusing on the "local" regions of clothing, detailed fabric textures can be more accurately captured. To this end, we propose a framework for learning cross-modal features from both global (the entire garment) and local (a close-up detail) image-text pairs. To verify the idea, we constructed a new dataset named Global and Local FACAD (G&L FACAD) by modifying the existing large-scale public FACAD dataset used for fashion retrieval. The experimental results confirm that the retrieval accuracy is significantly improved compared to the baselines. The code is available at https://github.com/SuzukiDaichi-git/texture_aware_fashion_retrieval.git. Daichi Suzuki, Go Irie, Kiyoharu Aizawa |
ICMR | 3 |
| 2023 | Open-Vocabulary Segmentation Approach for Transformer-Based Food Nutrient EstimationabstractNutrition plays a vital role in overall health and well-being. With a highly accurate nutrient estimation model, we develop a tool that displays nutritional values from food images, thereby reducing the labor-intensiveness of dietary assessment. We propose a method that uses depth data with RGB images and incorporates an open-vocabulary segmentation process that separates food from non-food instances, coupled with two-stage self-attention Transformer decoder. Our model outperforms the current state-of-the-art method, with an average percent MAE of 17.2% on Nutrition5k, an RGB-D food image dataset with calories, mass, and three macronutrients annotated. Our study also focuses on the significance of the food and background regions for calorie, mass, and nutrient estimation. We analyze the impact of non-food regions on each estimation task, with results suggesting that background information is crucial for calorie, mass, and carbohydrate estimation but not as essential for protein and fat estimation. The qualitative results also show that the model attends to regions with a high corresponding nutritional value. Implementation codes and pre-trained models are provided at https://github.com/Oatsty/nutrition5k. Satayu Parinayok, Yoko Yamakata, Kiyoharu Aizawa |
MMAsia | 3 |
| 2023 | Automatic Dataset Creation from User-generated Recipes for Ingredient-centric Food Image AnalysisabstractWe aim to develop an application that automatically creates a nutrition facts label from food images for precise dietary control. Firstly, we constructed a new dataset with food category labels and a list of ingredients in a nutritionally calculable format using an image classification model and BERT for 1.6 million recipes accompanied by images. The nutritional value of the recipe can be calculated using a conversion table consisting of the food item number and unit class. Next, using deep learning techniques, we built models that estimate the list of food item numbers from food images. While the multi-task model that identifies the food category label and the ingredient list simultaneously is only effective within a limited number of recipes, the single-task model that only identified the ingredient list achieved a Micro-F1 of 53.32% in total. Yoko Yamakata, Kiyoharu Aizawa |
MMAsia | 3 |
| 2022 | SLGAN: Style- and Latent-Guided Generative Adversarial Network for Desirable Makeup Transfer and RemovalabstractThere are five features to consider when using generative adversarial networks to apply makeup to photos of the human face. These features include (1) facial components, (2) interactive color adjustments, (3) makeup variations, (4) robustness to poses and expressions, and the (5) use of multiple reference images. To tackle the key features, we propose a novel style- and latent-guided makeup generative adversarial network for makeup transfer and removal. We provide a novel, perceptual makeup loss and a style-invariant decoder that can transfer makeup styles based on histogram matching to avoid the identity-shift problem. In our experiments, we show that our SLGAN is better than or comparable to state-of-the-art methods. Furthermore, we show that our proposal can interpolate facial makeup images to determine the unique features, compare existing methods, and help users find desirable makeup configurations. Daichi Horita, Kiyoharu Aizawa |
MMAsia | 2 |
| 2022 | FoodLog Athl: Multimedia Food Recording Platform for Dietary Guidance and Food MonitoringabstractThis paper presents a new food recording tool, FoodLog Athl, for the healthcare or physical enhancement of its users. Unlike existing food recording tools, we designed the system for dietitians or third parties who monitor the users. The tool not only supports the users by functions such as food image recognition, but also it helps the dietitians watch and communicate with users. Furthermore, it calculates nutritional values from food records - the use of the tool reduces the workload of dietitians and focuses their work on nutrition guidance. Kei Nakamoto, Kohei Kumazawa, Hiroaki Karasawa, Sosuke Amano, Yoko Yamakata, Kiyoharu Aizawa |
MMAsia | 6 |
| 2022 | Wearable Camera Based Food Logging SystemabstractRecently, meal management apps have allowed people to record food items and calories from photos automatically. These technologies include extracting food regions from photos of served meals, identifying the name of the food in each region, and calculating nutritional data. However, what you eat is not the only indicator that should be kept in the food record. How fast you eat and the order in which you eat is also significant information for dietary management. Therefore, we aim to construct a system that automatically generates a meal log from first-person videos that users capture of their eating behavior with a wearable camera. To tackle the complex problems that the data this system assumes contains, we constructed an eating behavior record dataset: 9.9 hours of first-person video that assume the natural diets of a user. To investigate the feasibility of our proposed system, we evaluated whether the first step, the detection of the meal area in the video during the meal, could be achieved with sufficient accuracy using this dataset. Using the limited number of frames assumed to be annotated by the user as training data, 30 frames were annotated for user-specific model training and four frames for online adaptation, resulting in detection accuracy of 72% for food regions. Our next goal is to create a multi-user dataset and service the application. Kenshiro Sato, Yoko Yamakata, Sosuke Amano, Kiyoharu Aizawa |
MMAsia | 4 |
| 2020 | Urban Movie Map for Walkers: Route View Synthesis using 360° VideosabstractWe propose a movie map for walkers based on synthesized street walking views along routes in a particular area. From the perspectives of walkers, we captured a number of omnidirectional videos along streets in the target area (1km2 around Kyoto Station). We captured a separate video for each street. We then performed simultaneous localization and mapping to obtain camera poses from key video frames in all of the videos and adjusted the coordinates based on a map of the area using reference points. To join one video to another smoothly at intersections, we identified frames of video intersection based on camera locations and visual feature matching. Finally, we generated moving route views by connecting the omnidirectional videos based on the alignment of the cameras. To improve smoothness at intersections, we generated rotational views by mixing video intersection frames from two videos. The results demonstrate that our method can precisely identify intersection frames and generate smooth connections between videos at intersections. Naoki Sugimoto, Toru Okubo, Kiyoharu Aizawa |
ICMR | 3 |
| 2019 | Assist Users' Interactions in Font Search with Unexpected but Useful Concepts Generated by Multimodal LearningabstractWhen searching for suitable fonts for a digital graphic, users usually start with an ambiguous thought. For example, they would look for fonts that are suitable for a personal web page or party invitations for children. Their design concept becomes clearer as they interact with external interventions such as exposure to suitable images for use in their web page or the children's preferences regarding the party. Hence, it is important to support users' interactions with unexpected but useful concepts during their search. In this paper, we present a novel framework that helps users to explore a font dataset using the multimodal method that provides unexpected but useful font images or concept words in response to the user's input. We collect a large font dataset and the associated tags and propose the use of unsupervised generative model that jointly learns the correlation between the visual features of a font and the associated tags for the creative process. By examining the results of the model that change with various inputs, we observed that the model produces highly promising results. In the experiment, we verified that the generated concepts by the model are not only new but also relevant to the user input that appears to be useful for inspiring users. Saemi Choi, Shun Matsumura, Kiyoharu Aizawa |
ICMR | 3 |
| 2019 | Social Font Search by Multimodal Feature EmbeddingabstractA typical tag/keyword-based search system retrieves documents where, given a query term q, the query term q occurs in the dataset. However, when applying these systems to a real-world font web community setting, practical challenges arise --- font tags are more subjective than other benchmark datasets, which magnify the tag mismatch problem. To address these challenges, we propose a tag dictionary space leveraged by word embedding, which relates undefined words that have a similar meaning. Even if a query is not defined in the tag dictionary, we can represent it as a vector on the tag dictionary space. The proposed system facilitates multi-modal inputs that can use both textual and image queries. By integrating a visual sentiment concept model that classifies affective concepts as adjective--noun pairs for a given image and uses it as a query, users can interact with the search system in a multi-modal way. We used crowd sourcing to collect user ratings for the retrieved fonts and observed that the retrieved font with the proposed methods obtained a higher score compared to other methods. Saemi Choi, Shun Matsumura, Kiyoharu Aizawa |
MMAsia | 3 |
| 2019 | Face hallucination through differential evolution parameter map learning with facial structure prior
Junjun Jiang, Jiayi Ma 0001, Suhua Tang, Yi Yu 0001, Kiyoharu Aizawa |
Inf. Sci. | 5 |
| 2017 | Simple, Efficient and Effective Encodings of Local Deep Features for Video Action RecognitionabstractFor an action recognition system a decisive component is represented by the feature encoding part which builds the final representation that serves as input to a classifier. One of the shortcomings of the existing encoding approaches is the fact that they are built around hand-crafted features and they are not also highly competitive on encoding the current deep features, necessary in many practical scenarios. In this work we propose two solutions specifically designed for encoding local deep features, taking advantage of the nature of deep networks, focusing on capturing the highest feature response of the convolutional maps. The proposed approaches for deep feature encoding provide a solution to encapsulate the features extracted with a convolutional neural network over the entire video. In terms of accuracy our encodings outperform by a large margin the current most widely used and powerful encoding approaches, while being extremely efficient for the computational cost. Evaluated in the context of action recognition tasks, our pipeline obtains state-of-the-art results on three challenging datasets: HMDB51, UCF50 and UCF101. I. C. Duta, Bogdan Ionescu, Kiyoharu Aizawa, Nicu Sebe |
ICMR | 3 |
| 2015 | A Discourse Search Engine Based on Rhetorical Structure Theory
Pascal Kuyten, Danushka Bollegala, Bernd Hollerit, Helmut Prendinger, Kiyoharu Aizawa |
ECIR | 5 |