Yuki Endo 0001

dblp:22/216-1 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0001-5132-3350ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 8 first-author · 17 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 EG-HumanNeRF: Efficient Generalizable Human NeRF Utilizing Human Prior for Sparse View
Zhaorong Wang, Yoshihiro Kanamori, Yuki Endo 0001
Comput. Vis. Media3
2026 Person-In-Situ: Scene-Consistent Human Image Insertion With Occlusion-Aware Pose Control
abstract
ABSTRACT Compositing human figures into scene images has broad applications in areas such as entertainment and advertising. However, existing methods often cannot handle occlusion of the inserted person by foreground objects and unnaturally place the person in the frontmost layer. Moreover, they offer limited control over the inserted person's pose. To address these challenges, we propose two methods. Both allow explicit pose control via a 3D body model and leverage latent diffusion models to synthesize the person at a contextually appropriate depth, naturally handling occlusions without requiring occlusion masks. The first is a two‐stage approach: the model first learns a depth map of the scene with the person through supervised learning, and then synthesizes the person accordingly. The second method learns occlusion implicitly and synthesizes the person directly from input data without explicit depth supervision. Quantitative and qualitative evaluations show that both methods outperform existing approaches by better preserving scene consistency while accurately reflecting occlusions and user‐specified poses.
Shun Masuda, Yuki Endo 0001, Yoshihiro Kanamori
Comput. Animat. Virtual Worlds2
2025 FreeUV: Ground-Truth-Free Realistic Facial UV Texture Recovery via Cross-Assembly Inference Strategy
abstract
Recovering high-quality 3D facial textures from single-view 2D images is a challenging task, especially under the constraints of limited data and complex facial details such as wrinkles, makeup, and occlusions. In this paper, we introduce FreeUV, a novel ground-truth-free UV texture recovery framework that eliminates the need for annotated or synthetic UV data. FreeUV leverages a pre-trained stable diffusion model alongside a Cross-Assembly inference strategy to fulfill this objective. In FreeUV, separate networks are trained independently to focus on realistic appearance and structural consistency, and these networks are combined during inference to generate coherent textures. Our approach accurately captures intricate facial features and demonstrates robust performance across diverse poses and occlusions. Extensive experiments validate FreeUV’s effectiveness, with results surpassing state-of-the-art methods in both quantitative and qualitative metrics. Additionally, FreeUV enables new applications, including local editing, facial feature interpolation, and texture recovery from multi-view images. By reducing data requirements, FreeUV offers a scalable solution for generating high-fidelity 3D facial textures suitable for real-world scenarios.
Xingchao Yang, Takafumi Taketomi, Yuki Endo 0001, Yoshihiro Kanamori
CVPR3
2025 3D View Optimization for Improving Image Aesthetics
abstract
Achieving aesthetically pleasing photography necessitates attention to multiple factors, including composition and capture conditions, which pose challenges to novices. Prior research has explored the enhancement of photo aesthetics post-capture through 2D manipulation techniques; however, these approaches offer limited search space for aesthetics. We introduce a pioneering method that employs 3D operations to simulate the conditions at the moment of capture retrospectively. Our approach extrapolates the input image and then reconstructs the 3D scene from the extrapolated image, followed by optimization to find camera parameters and image size that yield the best 3D view with enhanced aesthetics. Comparative qualitative and quantitative assessments reveal that our method surpasses traditional 2D editing techniques with superior aesthetics.
Taichi Uchida, Yoshihiro Kanamori, Yuki Endo 0001
ICASSP3
2025 All-frequency Full-body Human Image Relighting
abstract
Abstract Relighting of human images enables post‐photography editing of lighting effects in portraits. The current mainstream approach uses neural networks to approximate lighting effects without explicitly accounting for the principle of physical shading. As a result, it often has difficulty representing high‐frequency shadows and shading. In this paper, we propose a two‐stage relighting method that can reproduce physically‐based shadows and shading from low to high frequencies. The key idea is to approximate an environment light source with a set of a fixed number of area light sources. The first stage employs supervised inverse rendering from a single image using neural networks and calculates physically‐based shading. The second stage then calculates shadow for each area light and sums up to render the final image. We propose to make soft shadow mapping differentiable for the area‐light approximation of environment lighting. We demonstrate that our method can plausibly reproduce all‐frequency shadows and shading caused by environment illumination, which have been difficult to reproduce using existing methods.
Daichi Tajima, Yoshihiro Kanamori, Yuki Endo 0001
Comput. Graph. Forum3
2025 Selfage: personalized facial age transformation using self-reference images
Taishi Ito, Yuki Endo 0001, Yoshihiro Kanamori
Vis. Comput.2
2025 TerraFusion: Joint generation of terrain geometry and texture using latent diffusion models
abstract
Three-dimensional terrain models are essential in domains such as video game development and film production. Because surface color is often correlated with terrain geometry, capturing this relationship is critical for generating realistic results. However, most existing methods synthesize either a heightmap or a texture without adequately modeling their inherent correlation. We propose a method that jointly generates terrain heightmaps and textures using a latent diffusion model. First, we train the model in an unsupervised manner to randomly generate paired heightmaps and textures. Then, we perform supervised learning on an external adapter to enable user control via hand-drawn sketches. Experiments demonstrate that our approach supports intuitive terrain generation while preserving the correlation between heightmaps and textures. Our method outperforms the two-stage and GAN-based baselines by ensuring structural coherence, in which textures naturally align with geometry, successfully accommodating both realistic landscapes and extreme user-defined shapes.
Kazuki Higo, Toshiki Kanai, Yuki Endo 0001, Yoshihiro Kanamori
Virtual Real. Intell. Hardw.3
2024 Makeup Prior Models for 3D Facial Makeup Estimation and Applications
abstract
In this work, we introduce two types of makeup prior models to extend existing 3D face prior models: PCA-based and StyleGAN2-based priors. The PCA-based prior model is a linear model that is easy to construct and is computationally efficient. However, it retains only low-frequency information. Conversely, the StyleGAN2-based model can represent high-frequency information with relatively higher computational cost than the PCA-based model. Although there is a trade-off between the two models, both are applicable to 3D facial makeup estimation and related applications. By leveraging makeup prior models and designing a makeup consistency module, we effectively address the challenges that previous methods faced in robustly estimating makeup, particularly in the context of handling self-occluded faces. In experiments, we demonstrate that our approach reduces computational costs by several orders of magnitude, achieving speeds up to 180 times faster. In addition, by improving the accuracy of the estimated makeup, we confirm that our methods are highly advantageous for various 3D facial makeup applications such as 3D makeup face reconstruction, user-friendly makeup editing, makeup transfer, and interpolation.
Xingchao Yang, Takafumi Taketomi, Yuki Endo 0001, Yoshihiro Kanamori
CVPR3
2024 DiffBody: Diffusion-based Pose and Shape Editing of Human Images
abstract
Pose and body shape editing in a human image has received increasing attention. However, current methods often struggle with dataset biases and deteriorate realism and the person’s identity when users make large edits. We propose a one-shot approach that enables large edits with identity preservation. To enable large edits, we fit a 3D body model, project the input image onto the 3D model, and change the body’s pose and shape. Because this initial textured body model has artifacts due to occlusion and the inaccurate body shape, the rendered image undergoes a diffusion-based refinement, in which strong noise destroys body structure and identity whereas insufficient noise does not help. We thus propose an iterative refinement with weak noise, applied first for the whole body and then for the face. We further enhance the realism by fine-tuning text embeddings via self-supervised learning. Our quantitative and qualitative evaluations demonstrate that our method outperforms other existing methods across various datasets.
Yuta Okuyama, Yuki Endo 0001, Yoshihiro Kanamori
WACV2
2024 Masked-attention diffusion guidance for spatially controlling text-to-image generation
Yuki Endo 0001
Vis. Comput.1
2024 Seasonal terrain texture synthesis via Köppen periodic conditioning
Toshiki Kanai, Yuki Endo 0001, Yoshihiro Kanamori
Vis. Comput.2
2023 Age-dependent face diversification via latent space analysis
Taishi Ito, Yuki Endo 0001, Yoshihiro Kanamori
Vis. Comput.2
2022 User-Controllable Latent Transformer for StyleGAN Image Layout Editing
abstract
Abstract Latent space exploration is a technique that discovers interpretable latent directions and manipulates latent codes to edit various attributes in images generated by generative adversarial networks (GANs). However, in previous work, spatial control is limited to simple transformations (e.g., translation and rotation), and it is laborious to identify appropriate latent directions and adjust their parameters. In this paper, we tackle the problem of editing the StyleGAN image layout by annotating the image directly. To do so, we propose an interactive framework for manipulating latent codes in accordance with the user inputs. In our framework, the user annotates a StyleGAN image with locations they want to move or not and specifies a movement direction by mouse dragging. From these user inputs and initial latent codes, our latent transformer based on a transformer encoder‐decoder architecture estimates the output latent codes, which are fed to the StyleGAN generator to obtain a result image. To train our latent transformer, we utilize synthetic data and pseudo‐user inputs generated by off‐the‐shelf StyleGAN and optical flow models, without manual supervision. Quantitative and qualitative evaluations demonstrate the effectiveness of our method over existing methods.
Yuki Endo 0001
Comput. Graph. Forum1
2022 Controlling StyleGANs using rough scribbles via one-shot learning
abstract
Abstract This paper tackles the challenging problem of one‐shot semantic image synthesis from rough sparse annotations, which we call “semantic scribbles.” Namely, from only a single training pair annotated with semantic scribbles, we generate realistic and diverse images with layout control over, for example, facial part layouts and body poses. We present a training strategy that performs pseudo labeling for semantic scribbles using the StyleGAN prior. Our key idea is to construct a simple mapping between StyleGAN features and each semantic class from a single example of semantic scribbles. With such mappings, we can generate an unlimited number of pseudo semantic scribbles from random noise to train an encoder for controlling a pretrained StyleGAN generator. Even with our rough pseudo semantic scribbles obtained via one‐shot supervision, our method can synthesize high‐quality images thanks to our GAN inversion framework. We further offer optimization‐based postprocessing to refine the pixel alignment of synthesized images. Qualitative and quantitative results on various datasets demonstrate improvement over previous approaches in one‐shot settings.
Yuki Endo 0001, Yoshihiro Kanamori
Comput. Animat. Virtual Worlds1
2022 3D terrain estimation from a single landscape image
abstract
Abstract This article presents the first technique to estimate a 3D terrain model from a single landscape image. Although monocular depth estimation also offers single‐image 3D reconstruction, it assigns depth only to pixels visible in the input image, resulting in an incomplete 3D terrain output. Our method generates a complete 3D terrain model as a textured height map via a three‐stage framework using deep neural networks. First, to exploit the performance of pixel‐aligned estimation, we estimate terrain's per‐pixel depth and color free from shadows or lights in the perspective view. Second, we triangulate the RGB‐D data generated in the first stage and rasterize the triangular mesh from the top view to obtain an incomplete textured height map. Finally, we inpaint the depth and color in the missing regions. Because there are many possible ways to complete the missing regions, we synthesize diverse shapes and textures during inpainting using a variational autoencoder. Qualitative and quantitative experiments reveal that our method outperforms existing methods applying a direct perspective‐to‐top view transform as image‐to‐image translation.
Haruka Takahashi, Yoshihiro Kanamori, Yuki Endo 0001
Comput. Animat. Virtual Worlds3
2022 Diversifying detail and appearance in sketch-based face image synthesis
Takato Yoshikawa, Yuki Endo 0001, Yoshihiro Kanamori
Vis. Comput.2
2021 Relighting Humans in the Wild: Monocular Full-Body Human Relighting with Domain Adaptation
abstract
Abstract The modern supervised approaches for human image relighting rely on training data generated from 3D human models. However, such datasets are often small (e.g., Light Stage data with a small number of individuals) or limited to diffuse materials (e.g., commercial 3D scanned human models). Thus, the human relighting techniques suffer from the poor generalization capability and synthetic‐to‐real domain gap. In this paper, we propose a two‐stage method for single‐image human relighting with domain adaptation. In the first stage, we train a neural network for diffuse‐only relighting. In the second stage, we train another network for enhancing non‐diffuse reflection by learning residuals between real photos and images reconstructed by the diffuse‐only network. Thanks to the second stage, we can achieve higher generalization capability against various cloth textures, while reducing the domain gap. Furthermore, to handle input videos, we integrate illumination‐aware deep video prior to greatly reduce flickering artifacts even with challenging settings under dynamic illuminations.
Daichi Tajima, Yoshihiro Kanamori, Yuki Endo 0001
Comput. Graph. Forum3
2020 Diversifying Semantic Image Synthesis and Editing via Class- and Layer-wise VAEs
abstract
Abstract Semantic image synthesis is a process for generating photorealistic images from a single semantic mask. To enrich the diversity of multimodal image synthesis, previous methods have controlled the global appearance of an output image by learning a single latent space. However, a single latent code is often insufficient for capturing various object styles because object appearance depends on multiple factors. To handle individual factors that determine object styles, we propose a class‐ and layer‐wise extension to the variational autoencoder (VAE) framework that allows flexible control over each object class at the local to global levels by learning multiple latent spaces. Furthermore, we demonstrate that our method generates images that are both plausible and more diverse compared to state‐of‐the‐art methods via extensive experiments with real and synthetic datasets in three different domains. We also show that our method enables a wide range of applications in image synthesis and editing tasks.
Yuki Endo 0001, Yoshihiro Kanamori
Comput. Graph. Forum1
2019 Animating landscape: self-supervised learning of decoupled motion and appearance for single-image video synthesis
abstract
Automatic generation of a high-quality video from a single image remains a challenging task despite the recent advances in deep generative models. This paper proposes a method that can create a high-resolution, long-term animation using convolutional neural networks (CNNs) from a single landscape image where we mainly focus on skies and waters. Our key observation is that the motion (e.g., moving clouds) and appearance (e.g., time-varying colors in the sky) in natural scenes have different time scales. We thus learn them separately and predict them with decoupled control while handling future uncertainty in both predictions by introducing latent codes. Unlike previous methods that infer output frames directly, our CNNs predict spatially-smooth intermediate data, i.e., for motion, flow fields for warping, and for appearance, color transfer maps, via self-supervised learning, i.e., without explicitly-provided ground truth. These intermediate data are applied not to each previous output frame, but to the input image only once for each output frame. This design is crucial to alleviate error accumulation in long-term predictions, which is the essential problem in previous recurrent approaches. The output frames can be looped like cinemagraph, and also be controlled directly by specifying latent codes or indirectly via visual annotations. We demonstrate the effectiveness of our method through comparisons with the state-of-the-arts on video prediction as well as appearance manipulation. Resultant videos, codes, and datasets will be available at http://www.cgg.cs.tsukuba.ac.jp/~endo/projects/AnimatingLandscape.
Yuki Endo 0001, Yoshihiro Kanamori, Shigeru Kuriyama
ACM Trans. Graph.1
2018 Transferring pose and augmenting background for deep human-image parsing and its applications
abstract
Parsing of human images is a fundamental task for determining semantic parts such as the face, arms, and legs, as well as a hat or a dress. Recent deep-learning-based methods have achieved significant improvements, but collecting training datasets with pixel-wise annotations is labor-intensive. In this paper, we propose two solutions to cope with limited datasets. Firstly, to handle various poses, we incorporate a pose estimation network into an end-to-end human-image parsing network, in order to transfer common features across the domains. The pose estimation network can be trained using rich datasets and can feed valuable features to the human-image parsing network. Secondly, to handle complicated backgrounds, we increase the variation in image backgrounds automatically by replacing the original backgrounds of human images with others obtained from large-scale scenery image datasets. Individually, each solution is versatile and beneficial to human-image parsing, while their combination yields further improvement. We demonstrate the effectiveness of our approach through comparisons and various applications such as garment recoloring, garment texture transfer, and visualization for fashion analysis.
Takazumi Kikuchi, Yuki Endo 0001, Yoshihiro Kanamori, Taisuke Hashimoto, Jun Mitani
Comput. Vis. Media2
2018 Relighting humans: occlusion-aware inverse rendering for full-body human images
abstract
Relighting of human images has various applications in image synthesis. For relighting, we must infer albedo, shape, and illumination from a human portrait. Previous techniques rely on human faces for this inference, based on spherical harmonics (SH) lighting. However, because they often ignore light occlusion, inferred shapes are biased and relit images are unnaturally bright particularly at hollowed regions such as armpits, crotches, or garment wrinkles. This paper introduces the first attempt to infer light occlusion in the SH formulation directly. Based on supervised learning using convolutional neural networks (CNNs), we infer not only an albedo map, illumination but also a light transport map that encodes occlusion as nine SH coefficients per pixel. The main difficulty in this inference is the lack of training datasets compared to unlimited variations of human portraits. Surprisingly, geometric information including occlusion can be inferred plausibly even with a small dataset of synthesized human figures, by carefully preparing the dataset so that the CNNs can exploit the data coherency. Our method accomplishes more realistic relighting than the occlusion-ignored formulation.
Yoshihiro Kanamori, Yuki Endo 0001
ACM Trans. Graph.2
2017 Semi-Automatic Conversion of 3D Shape into Flat-Foldable Polygonal Model
abstract
Abstract This paper presents a method that can convert a given 3D mesh into a flat‐foldable model consisting of rigid panels. A previous work proposed a method to assist manual design of a single component of such flat‐foldable model, consisting of vertically‐connected side panels as well as horizontal top and bottom panels. Our method semi‐automatically generates a more complicated model that approximates the input mesh with multiple convex components. The user specifies the folding direction of each convex component and the fidelity of shape approximation. Given the user inputs, our method optimizes shapes and positions of panels of each convex component in order to make the whole model flat‐foldable. The user can check a folding animation of the output model. We demonstrate the effectiveness of our method by fabricating physical paper prototypes of flat‐foldable models.
Emi Miyamoto, Yuki Endo 0001, Yoshihiro Kanamori, Jun Mitani
Comput. Graph. Forum2
2017 Deep reverse tone mapping
abstract
Inferring a high dynamic range (HDR) image from a single low dynamic range (LDR) input is an ill-posed problem where we must compensate lost data caused by under-/over-exposure and color quantization. To tackle this, we propose the first deep-learning-based approach for fully automatic inference using convolutional neural networks. Because a naive way of directly inferring a 32-bit HDR image from an 8-bit LDR image is intractable due to the difficulty of training, we take an indirect approach; the key idea of our method is to synthesize LDR images taken with different exposures (i.e., bracketed images ) based on supervised learning, and then reconstruct an HDR image by merging them. By learning the relative changes of pixel values due to increased/decreased exposures using 3D deconvolutional networks, our method can reproduce not only natural tones without introducing visible noise but also the colors of saturated pixels. We demonstrate the effectiveness of our method by comparing our results not only with those of conventional methods but also with ground-truth HDR images.
Yuki Endo 0001, Yoshihiro Kanamori, Jun Mitani
ACM Trans. Graph.1
2016 DeepProp: Extracting Deep Features from a Single Image for Edit Propagation
abstract
Abstract Edit propagation is a technique that can propagate various image edits (e.g., colorization and recoloring) performed via user strokes to the entire image based on similarity of image features. In most previous work, users must manually determine the importance of each image feature (e.g., color, coordinates, and textures) in accordance with their needs and target images. We focus on representation learning that automatically learns feature representations only from user strokes in a single image instead of tuning existing features manually. To this end, this paper proposes an edit propagation method using a deep neural network (DNN). Our DNN, which consists of several layers such as convolutional layers and a feature combiner, extracts stroke‐adapted visual features and spatial features, and then adjusts the importance of them. We also develop a learning algorithm for our DNN that does not suffer from the vanishing gradient problem, and hence avoids falling into undesirable locally optimal solutions. We demonstrate that edit propagation with deep features, without manual feature tuning, can achieve better results than previous work.
Yuki Endo 0001, Satoshi Iizuka, Yoshihiro Kanamori, Jun Mitani
Comput. Graph. Forum1
2016 Single Image Weathering via Exemplar Propagation
abstract
Abstract This paper presents an efficient approach for generating weathering effects with detailed appearance variations in a single image. Previous approaches merely change chroma or reflectance of weathered objects, which is not sufficient for materials with detailed shading and texture variations, such as growing moss and peeling plaster. Our method propagates such detailed features via seamless patch‐based synthesis driven by weathering degree distribution. Unlike previous methods, the weathering degrees are calculated efficiently using Radial Basis Functions even for materials with wide color variations. We use graph cut‐based optimization to identify the most weathered region as a “weathering exemplar”, from which we sample weathering patches. We demonstrate our method enables us to generate various types of detailed weathering effects interactively.
Satoshi Iizuka, Yuki Endo 0001, Yoshihiro Kanamori, Jun Mitani
Comput. Graph. Forum2
2014 Object Repositioning Based on the Perspective in a Single Image
abstract
Abstract We propose an image editing system for repositioning objects in a single image based on the perspective of the scene. In our system, an input image is transformed into a layer structure that is composed of object layers and a background layer, and then the scene depth is computed from the ground region that is specified by the user using a simple boundary line. The object size and order of overlapping are automatically determined during the reposition based on the scene depth. In addition, our system enables the user to move shadows along with objects naturally by extracting the shadow mattes using only a few user‐specified scribbles. Finally, we demonstrate the versatility of our system through applications to depth‐of‐field effects, fog synthesis and 3D walkthrough in an image.
Satoshi Iizuka, Yuki Endo 0001, Masaki Hirose, Yoshihiro Kanamori, Jun Mitani, Yukio Fukui
Comput. Graph. Forum2
2014 Efficient Depth Propagation for Constructing a Layered Depth Image from a Single Image
abstract
Abstract In this paper, we propose an interactive technique for constructing a 3D scene via sparse user inputs. We represent a 3D scene in the form of a Layered Depth Image (LDI) which is composed of a foreground layer and a background layer, and each layer has a corresponding texture and depth map. Given user‐specified sparse depth inputs, depth maps are computed based on superpixels using interpolation with geodesic‐distance weighting and an optimization framework. This computation is done immediately, which allows the user to edit the LDI interactively. Additionally, our technique automatically estimates depth and texture in occluded regions using the depth discontinuity. In our interface, the user paints strokes on the 3D model directly. The drawn strokes serve as 3D handles with which the user can pull out or push the 3D surface easily and intuitively with real‐time feedback. We show our technique enables efficient modeling of LDI that produce sufficient 3D effects.
Satoshi Iizuka, Yuki Endo 0001, Yoshihiro Kanamori, Jun Mitani, Yukio Fukui
Comput. Graph. Forum2
2012 Matting and Compositing for Fresnel Reflection on Wavy Surfaces
abstract
Abstract This paper introduces a framework that can extract an alpha matte from a single image with Fresnel reflection, and that can composite other objects with the image such that plausible reflections are included. Our method handles reflections in a plane with small undulations, for example, a water surface with waves or a glossy tabletop. During the matting stage, our method first estimates the transmission color, which is assumed to be uniform, and then calculates a reflection image and alpha matte based on user markups. However, accurate extraction of the matte becomes challenging when a plane has small undulations because these create perturbations in the matte. We therefore propose a filter that can refine the matte effectively. In the compositing stage, the reflection of a composited object is synthesized by ray tracing in real time. We demonstrate the effectiveness of our method through comparisons with ground‐truth data and results using natural images as inputs.
Yuki Endo 0001, Yoshihiro Kanamori, Yukio Fukui, Jun Mitani
Comput. Graph. Forum1
2011 An interactive design system for pop-up cards with a physical simulation
Satoshi Iizuka, Yuki Endo 0001, Jun Mitani, Yoshihiro Kanamori, Yukio Fukui
Vis. Comput.2