Daichi Horita

dblp:223/3292 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation
abstract
Content-aware graphic layout generation aims to automatically arrange visual elements along with a given content, such as an e-commerce product image. In this paper, we argue that the current layout generation approaches suffer from the limited training data for the high-dimensional layout structure. We show that a simple retrieval augmentation can significantly improve the generation quality. Our model, which is named Retrieval-Augmented Layout Transformer (RALF),retrieves nearest neighbor layout examples based on an input image and feeds these results into an autoregressive generator. Our model can apply retrieval augmentation to various controllable generation tasks and yield high-quality layouts within a unified architecture. Our extensive experiments show that RALF successfully generates content-aware layouts in both constrained and unconstrained settings and significantly outperforms the baselines.1
Daichi Horita, Naoto Inoue, Kotaro Kikuchi, Kota Yamaguchi, Kiyoharu Aizawa
CVPR1
2023 A Structure-Guided Diffusion Model for Large-Hole Image Completion
Daichi Horita, Jiaolong Yang, Dong Chen 0003, Yuki Koyama 0001, Kiyoharu Aizawa, Nicu Sebe
BMVC1
2023 Restorable Visible and Infrared Image Fusion
abstract
Image fusion aims to synthesize multiple source images into a single image to integrate and enhance information. Specifically, we tackle the fusion of visible and infrared images. Previous works generally use structural similarity between the fusion and the paired source images to train a deep-learning-based fusion model. However, only using the structure often results in a texture-insufficient image. In this study, we aim to generate an image rich in texture. This study is inspired by the ability of an autoencoder to learn a compressed representation of the input image. Specifically, we learn a fusion image with the structure and texture of the source images. We propose a novel framework–Restorable visible and infrared Image Fusion, which consists of a fusion and decoupling network. The fusion network synthesizes source images, and the decoupling network restores the source images by decomposing a fusion image. Our framework can be trained by minimizing the difference between the source and restored images. The experimental results demonstrate that the fusion image generated by the proposed method maintains the texture of the source images.
Daichi Horita, Koki Tsubota, Kiyoharu Aizawa
ICIP2
2022 Translation of Illustration Artist Style Using Sailormoonredraw Data
abstract
The decision to draw requires answers to two questions: what to draw and how to draw. The latter refers to artist style and is an important factor in creating any illustration. In this paper, we propose a novel task, artist style translation, which translates one artist’s illustration into that of another artist style using deep learning. To solve this task in a supervised manner, we created a novel illustration dataset, SailormoonDataset, which consists of more than 2,000 artist’s stylistic illustrations of the same content. In addition, we propose a method based on the Swapping Autoencoder by introducing a new loss function for supervised learning and using multiple images to represent an artist style. We translate the face illustration of Sailor Moon into styles of different artists. By comparing the current results to those of the Swapping Autoencoder, we find that the proposed method successfully achieves the artist style translation.
Keita Awane, Daichi Horita, Hikaru Ikuta, Yusuke Matsui 0001, Kiyoharu Aizawa, Naohiro Yanase
ICIP2
2022 SLGAN: Style- and Latent-Guided Generative Adversarial Network for Desirable Makeup Transfer and Removal
abstract
There are five features to consider when using generative adversarial networks to apply makeup to photos of the human face. These features include (1) facial components, (2) interactive color adjustments, (3) makeup variations, (4) robustness to poses and expressions, and the (5) use of multiple reference images. To tackle the key features, we propose a novel style- and latent-guided makeup generative adversarial network for makeup transfer and removal. We provide a novel, perceptual makeup loss and a style-invariant decoder that can transfer makeup styles based on histogram matching to avoid the identity-shift problem. In our experiments, we show that our SLGAN is better than or comparable to state-of-the-art methods. Furthermore, we show that our proposal can interpolate facial makeup images to determine the unique features, compare existing methods, and help users find desirable makeup configurations.
Daichi Horita, Kiyoharu Aizawa
MMAsia1
2022 Fast Nonlinear Image Unblending
abstract
Nonlinear color blending, which is advanced blending indicated by blend modes such as "overlay" and "multiply," is extensively employed by digital creators to produce attractive visual effects. To enjoy such flexible editing modalities on existing bitmap images like photographs, however, creators need a fast nonlinear blending algorithm that decomposes an image into a set of semi-transparent layers. To address this issue, we propose a neural-network-based method for nonlinear decomposition of an input image into linear and nonlinear alpha layers that can be separately modified for editing purposes, based on the specified color palettes and blend modes. Experiments show that our proposed method achieves an inference speed 370 times faster than the state-of-the-art method of nonlinear image unblending, which uses computationally intensive iterative optimization. Furthermore, our reconstruction quality is higher or comparable than other methods, including linear blending models. In addition, we provide examples that apply our method to image editing with nonlinear blend modes.
Daichi Horita, Kiyoharu Aizawa, Ryohei Suzuki, Taizan Yonetsuji, Huachun Zhu
WACV1
2019 DeepTaste: Augmented Reality Gustatory Manipulation with GAN-Based Real-Time Food-to-Food Translation
abstract
We have been studying augmented reality (AR)-based gustatory manipulation interfaces and previously proposed a gustatory manipulation interface using generative adversarial network (GAN)-based real time image-to-image translation. Unlike three-dimensional (3D) food model-based systems that only change the color or texture pattern of a particular type of food in an inflexible manner, our GAN-based system changes the appearance of food into multiple types of food in real time flexibly, dynamically, and interactively. In the present paper, we first describe in detail a user study on a vision-induced gustatory manipulation system using a 3D food model and report its successful experimental results. We then summarize identified problems of the 3D model-based system and describe implementation details of the GAN-based system. We finally report in detail the main user study in which we investigated the impact of the GAN-based system on gustatory sensations and food recognition when somen noodles were turned into ramen noodles or fried noodles, and steamed rice into curry and rice or fried rice. The experimental results revealed that our system successfully manipulates gustatory sensations to some extent and that the effectiveness seems to depend on the original and target types of food as well as the experience of each individual with the food.
Kizashi Nakano, Daichi Horita, Nobuchika Sakata, Kiyoshi Kiyokawa, Keiji Yanai, Takuji Narumi
ISMAR2
2019 Enchanting Your Noodles: GAN-based Real-time Food-to-Food Translation and Its Impact on Vision-induced Gustatory Manipulation
abstract
We propose a novel gustatory manipulation interface which utilizes the cross-modal effect of vision on taste elicited with augmented reality (AR)-based real-time food appearance modulation using a generative adversarial network (GAN). Unlike existing systems which only change color or texture pattern of a particular type of food in an inflexible manner, our system changes the appearance of food into multiple types of food in real-time flexibly, dynamically and interactively in accordance with the deformation of the food that the user is actually eating by using GAN-based image-to-image translation. The experimental results reveal that our system successfully manipulates gustatory sensations to some extent and that the effectiveness depends on the original and target types of food as well as each user's food experience.
Kizashi Nakano, Kiyoshi Kiyokawa, Daichi Horita, Keiji Yanai, Nobuchika Sakata, Takuji Narumi
VR3
2019 Enchanting Your Noodles: A Gustatory Manipulation Interface by Using GAN-based Real-time Food-to-Food Translation
abstract
In this demonstration, we present a novel gustatory manipulation interface which utilizes the cross-modal effect of vision on taste elicited with real-time food appearance modulation using a generative adversarial network (GAN). Unlike existing systems which only change color or texture pattern of a particular type of food in an inflexible manner, our system changes the appearance of food into multiple types of food in real-time flexibly, dynamically and interactively in accordance with the deformation of the food that the user is actually eating by using GAN-based image-to-image translation. Our system can turn somen noodles into ramen noodles or fried noodles, or steamed rice into curry and rice or fried rice. Users of our demonstration system will taste what is visually presented to some extent rather than what they are actually eating.
Kizashi Nakano, Kiyoshi Kiyokawa, Daichi Horita, Keiji Yanai, Nobuchika Sakata, Takuji Narurni
VR3
2018 Magical Rice Bowl: A Real-time Food Category Changer
abstract
In this demo, we demonstrate "Real-time Food Category Change'' based on a Conditional Cycle GAN (cCycle GAN) with a large-scale food image data collected from the Twitter Stream. Conditional Cycle GAN is an extension of CycleGAN, which enables "Food Category Change'' among ten kinds of typical foods served in bowl-type dishes such as beef rice bowl and ramen noodles. The proposed system enables us to change the appearance of a given food photo according to the given category keeping the shape of the given food but exchanging its textures. For training, we used two hundred and thirty thousand food images which achieved very natural food category change among ten kinds of typical Japanese foods: ramen noodle, curry rice, fried rice, beef rice bowl, chilled noodle, spaghetti with meat source, white rice, eel bowl, and fried noodle.
Ryosuke Tanno, Daichi Horita, Wataru Shimoda, Keiji Yanai
ACM Multimedia2
2018 Ramen spoon eraser: CNN-based photo transformation for improving attractiveness of ramen photos
abstract
In recent years, a large number of food photos are being posted globally on SNS. To obtain many views or "likes", attractive photos should be posted. However, some casual foods are served with utensils on a plate or a bowl at restaurants, which spoils attractiveness of meal photos. Especially in Japan where ramen noodle is the most popular casual food, ramen is usually served with a ramen spoon in a ramen bowl in a ramen noodle shop. This is a big problem for SNS photographers, because a ramen spoon soaked in a ramen bowl extremely degrades the appearance of ramen photos. Then, in this paper, we propose anapplication called "ramen spoon eraser" that erases a spoon from ramen photos with spoons using a CNN-based Image-to-Image translation network. In this application, it is possible to automatically erase ramen spoons from ramen photos, which extremely improve the attractiveness of ramen photos. In the experiment, we train models in two ways as CNN-based Image-to-Image translation networks with the dataset consisting of ramen images with / without spoons collected from the Web.
Daichi Horita, Jaehyeong Cho, Takumi Ege, Keiji Yanai
VRST1