Ryugo Morita

dblp:334/4113 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
0009-0007-6324-9291ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TKG-DM: Training-free Chroma Key Content Generation Diffusion Model
abstract
Diffusion models have enabled the generation of high-quality images with a strong focus on realism and textual fidelity. Yet, large-scale text-to-image models, such as Stable Diffusion, struggle to generate images where foreground objects are placed over a chroma key background, limiting their ability to separate foreground and background elements without fine-tuning. To address this limitation, we present a novel Training-Free Chroma Key Content Generation Diffusion Model (TKG-DM), which optimizes the initial random noise to produce images with foreground objects on a specifiable color background. Our proposed method is the first to explore the manipulation of the color aspects in initial noise for controlled background generation, enabling precise separation of foreground and background without fine-tuning. Extensive experiments demonstrate that our training-free method outperforms existing methods in both qualitative and quantitative evaluations, matching or surpassing fine-tuned models. Finally, we successfully extend it to other tasks (e.g., consistency models and text-to-video), highlighting its transformative potential across various generative applications where independent control of foreground and background is crucial.
Ryugo Morita, Stanislav Frolov, Brian B. Moser, Takahiro Shirakawa, Ko Watanabe 0001, Andreas Dengel 0001, Jinjia Zhou
CVPR1
2025 Bidirectional Learned Facial Animation Codec for Low Bitrate Talking Head Videos
abstract
In this paper, we propose a novel bidirectional learned animation codec that generates natural facial videos by using past and future keyframes. First, we introduce a compact auxiliary stream for non-keyframes, which is enhanced by adaptively selecting one of two keyframes (past and future) in the BRG-ASE process. This stream improves video quality with a slight increase in bitrate. Then, we animate the adaptively selected keyframe and reconstruct the target frame using both the animated keyframe and the auxiliary frame in the BRG-VRec process. In our bidirectional frame reconstruction method, the future keyframe is used as the past keyframe in the next group of pictures. Therefore, it is temporarily stored in the decoder.
Riku Takahashi, Ryugo Morita, Fuma Kimishima, Kosuke Iwama, Jinjia Zhou
DCC2
2025 Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos
abstract
Talking head video compression has advanced with neural rendering and keypoint-based methods, but challenges remain, especially at low bit rates, including handling large head movements, suboptimal lip synchronization, and distorted facial reconstructions. To address these problems, we propose a novel audio-visual driven video codec that integrates compact 3D motion features and audio signals. This approach robustly models significant head rotations and aligns lip movements with speech, improving both compression efficiency and reconstruction quality. Experiments on the CelebV-HQ dataset show that our method reduces bitrate by 22% compared to VVC and by 8.5% over state-of-the-art learning-based codec. Furthermore, it provides superior lip-sync accuracy and visual fidelity at comparable bitrates, highlighting its effectiveness in bandwidth-constrained scenarios.
Riku Takahashi, Ryugo Morita, Jinjia Zhou
ICMR2
2024 Learned Measurement Interpolation for Scalable Compressive Sensing
abstract
Deep learning-based joint compressive sensing image acquisition and reconstruction has attracted great attention in the field of image compression and computational sensors. Although these methods show excellent results in execution speed and image quality, each compression ratio needs one specific deep-learning model which causes training costs and memory pressure. To solve the above problems, we propose a Measurements Interpolation strategy (MI) that can be placed before a DL-based reconstruction model. MI takes original measurements, which are the compressed data of the input image, and outputs interpolated measurements with an arbitrary compression ratio by interpolating the original measurements. The experimental results show that our proposal can be integrated with the existing DL-based compressive sensing reconstruction works to achieve scalable sampling and reconstruction while keeping the image quality.
Manato Shirai, Fuma Kimishima, Jinjia Zhou, Ryugo Morita
IJCNN4
2024 Visual question answering based evaluation metrics for text-to-image generation
abstract
Text-to-image generation and text-guided image manipulation have received considerable attention in the field of image generation tasks. However, the mainstream evaluation methods for these tasks have difficulty in evaluating whether all the information from the input text is accurately reflected in the generated images, and they mainly focus on evaluating the overall alignment between the input text and the generated images. This paper proposes new evaluation metrics that assess the alignment between input text and generated images for every individual object. Firstly, according to the input text, chatGPT is utilized to produce questions for the generated images. After that, we use Visual Question Answering(VQA) to measure the relevance of the generated images to the input text, which allows for a more detailed evaluation of the alignment compared to existing methods. In addition, we use Non-Reference Image Quality Assessment(NR-IQA) to evaluate not only the text-image alignment but also the quality of the generated images. Experimental results show that our proposed evaluation approach is the superior metric that can simultaneously assess finer text-image alignment and image quality while allowing for the adjustment of these ratios.
Mizuki Miyamoto, Ryugo Morita, Jinjia Zhou
ISCAS2
2023 BATINeT: Background-Aware Text to Image Synthesis and Manipulation Network
abstract
Background-Induced Text2Image (BIT2I) aims to generate foreground content according to the text on the given background image. Most studies focus on generating high-quality foreground content, although they ignore the relationship between the two contents. In this study, we analyzed a novel Background-Aware Text2Image (BAT2I) task in which the generated content matches the input background. We proposed a Background-Aware Text to Image synthesis and manipulation Network (BATINet), which contains two key components: Position Detect Network (PDN) and Harmonize Network (HN). The PDN detects the most plausible position of the text-relevant object in the background image. The HN harmonizes the generated content referring to background style information. Finally, we reconstructed the generation network, which consists of the multi-GAN and attention module to match more user preferences. Moreover, we can apply BATINet to text-guided image manipulation. It solves the most challenging task of manipulating the shape of an object. We demonstrated through qualitative and quantitative evaluations on the CUB dataset that the proposed model outperforms other state-of-the-art methods.
Ryugo Morita, Jinjia Zhou
ICIP1
2023 Dynamic Unilateral Dual Learning for Text to Image Synthesis
abstract
Dual learning trains two inverse processes tasks dually to further improve the selected tasks’ performance. There are currently two training paradigms in dual learning. One is to directly utilize two existing models for training in a dual manner to improve the selected models’ performance. However, it cannot effectively guarantee the improvement of the selected models. Another is that the networks of both parties are manually designed. Nevertheless, the network performance of both parties will be poor in the initial stage of training, which will easily lead to unsatisfactory training. Besides, most of dual learning researches can only be used for the conversion between the same data types, but it is powerless for the conversion of different data types. To address the above issues, a paradigm called unilateral dual learning (UDL) is proposed and verified in the text-to-image (T2I) synthesis field. UDL allows one party to design the network manually, and the other party calls the pre-trained model to promote the training of the manually designed network to achieve satisfactory training. Experimental results on the Oxford-102 flower and Caltech-UCSD Birds datasets demonstrate the feasibility of our proposed UDL paradigm in the T2I field, and it achieves excellent performance qualitatively and quantitatively.
Jiayao Xu, Ryugo Morita, Wenxin Yu 0001, Jinjia Zhou
ICIP3
2023 Block based Adaptive Compressive Sensing with Sampling Rate Control
abstract
Compressive sensing (CS), acquiring and reconstructing signals below the Nyquist rate, has great potential in image and video acquisition to exploit data redundancy and greatly reduce the amount of sampled data. To further reduce the sampled data while keeping the video quality, this paper explores the temporal redundancy in video CS and proposes a block based adaptive compressive sensing framework with a sampling rate (SR) control strategy. To avoid redundant compression of non-moving regions, we first incorporate moving block detection between consecutive frames, and only transmit the measurements of moving blocks. The non-moving regions are reconstructed from the previous frame. In addition, we propose a block storage system and a dynamic threshold to achieve adaptive SR allocation to each frame based on the area of moving regions and target SR for controlling the average SR within the target SR. Finally, to reduce blocking artifacts and improve reconstruction quality, we adopt a cooperative reconstruction of the moving and non-moving blocks by referring to the measurements of the non-moving blocks from the previous frame. Extensive experiments have demonstrated that this work is able to control SR and obtain better performance than existing works.
Kosuke Iwama, Ryugo Morita, Jinjia Zhou
MMAsia2
2023 Interactive Image Manipulation with Complex Text Instructions
abstract
Recently, text-guided image manipulation has received increasing attention in the research field of multimedia processing and computer vision due to its high flexibility and controllability. Its goal is to semantically manipulate parts of an input reference image according to the text descriptions. However, most of the existing works have the following problems: (1) text-irrelevant content cannot always be maintained but randomly changed, (2) the performance of image manipulation still needs to be further improved, (3) only can manipulate descriptive attributes. To solve these problems, we propose a novel image manipulation method that interactively edits an image using complex text instructions. It allows users to not only improve the accuracy of image manipulation but also achieve complex tasks such as enlarging, dwindling, or removing objects and replacing the background with the input image. To make these tasks possible, we apply three strategies. First, the given image is divided into text-relevant content and text-irrelevant content. Only the text-relevant content is manipulated and the text-irrelevant content can be maintained. Second, a super-resolution method is used to enlarge the manipulation region to further improve the operability and to help manipulate the object itself. Third, a user interface is introduced for editing the segmentation map interactively to re-modify the generated image according to the user’s desires. Extensive experiments on the Caltech-UCSD Birds-200-2011 (CUB) dataset and Microsoft Common Objects in Context (MS COCO) datasets demonstrate our proposed method can enable interactive, flexible, and accurate image manipulation in real-time. Through qualitative and quantitative evaluations, we show that the proposed model outperforms other state-of-the-art methods.
Ryugo Morita, Man M. Ho, Jinjia Zhou
WACV1