Yuhi Matsuo

dblp:317/5427 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2022
0009-0003-0752-0017ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
image completion
0.612022
Diverse Plausible 360-Degree Image Outpainting for Efficient 3DCG Background Creation · CVPR 2022
Visual content generation and editing
image outpainting
0.612022
Diverse Plausible 360-Degree Image Outpainting for Efficient 3DCG Background Creation · CVPR 2022
Visual content generation and editing › image completion
pluralistic image completion
0.612022
Diverse Plausible 360-Degree Image Outpainting for Efficient 3DCG Background Creation · CVPR 2022
Visual content generation and editing
3d content creation
0.212022
Diverse Plausible 360-Degree Image Outpainting for Efficient 3DCG Background Creation · CVPR 2022

Methods — techniques the papers use, named apart from their topics

transformer · 0.6perceptual loss · 0.6circular inference · 0.6
YearPublicationVenuePosition
2022 Diverse Plausible 360-Degree Image Outpainting for Efficient 3DCG Background Creation
abstract
We address the problem of generating a 360-degree image from a single image with a narrow field of view by estimating its surroundings. Previous methods suffered from overfitting to the training resolution and deterministic generation. This paper proposes a completion method using a transformer for scene modeling and novel methods to improve the properties of a 360-degree image on the output image. Specifically, we use CompletionNets with a transformer to perform diverse completions and Adjust-mentNet to match color, stitching, and resolution with an input image, enabling inference at any resolution. To improve the properties of a 360-degree image on an output image, we also propose WS-perceptual loss and circular inference. Thorough experiments show that our method out-performs state-of-the-art (SOTA) methods both qualitatively and quantitatively. For example, compared to SOTA methods, our method completes images 16 times larger in resolution and achieves 1.7 times lower Fréchet inception distance (FID). Furthermore, we propose a pipeline that uses the completion results for lighting and background of 3DCG scenes. Our plausible background completion enables perceptually natural results in the application of inserting virtual objects with specular surfaces.
Naofumi Akimoto, Yuhi Matsuo, Yoshimitsu Aoki
CVPR2
2022 Document Shadow Removal with Foreground Detection Learning From Fully Synthetic Images
abstract
Shadow removal for document images is a major task for digitized document applications. Recent shadow removal models have been trained on pairs of shadow images and shadow-free images. However, obtaining a large-scale and diverse dataset is laborious and remains a great challenge. Thus, only small real datasets are available. To create relatively large datasets, a graphic renderer has been used to synthesize shadows, nonetheless, it is still necessary to capture real documents. Thus, the number of unique documents is limited, which negatively affects a network’s performance. In this paper, we present a large-scale and diverse dataset called fully synthetic document shadow removal dataset (FSDSRD) that does not require capturing documents. The experiments showed that the networks (pre-)trained on FSDSRD provided better results than networks trained only on real datasets. Additionally, because foreground maps are available in our dataset, we leveraged them during training for multitask learning, which provided noticeable improvements. The code is available at: https://github.com/IsHYuhi/DSRFGD.
Yuhi Matsuo, Naofumi Akimoto, Yoshimitsu Aoki
ICIP1