Yusuke Kato

dblp:14/9268 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-6848-0683ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 53% Vision and language · 21% Segmentation and scene understanding · 12%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 23 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.532025
LaViDa: A Large Diffusion Language Model for Multimodal Understanding · NeurIPS 2025
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025
Aligning Diffusion Models by Optimizing Human Utility · NeurIPS 2024
Machine learning › Generative modeling › multimodal generation
any-to-any generation
0.912025
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.912025
Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection · ICCV 2025
Machine learning › Generative modeling › diffusion model › discrete diffusion model
discrete diffusion language model
0.912025
LaViDa: A Large Diffusion Language Model for Multimodal Understanding · NeurIPS 2025
Machine learning › Generative modeling › multimodal generation
multimodal generative model
0.912025
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
SegLLM: Multi-round Reasoning Segmentation with Large Language Models · ICLR 2025
Computer vision › Vision and language
multimodal understanding
0.912025
LaViDa: A Large Diffusion Language Model for Multimodal Understanding · NeurIPS 2025
Computer vision › Vision and language › visual grounding › language-guided segmentation
reasoning segmentation
0.912025
SegLLM: Multi-round Reasoning Segmentation with Large Language Models · ICLR 2025
Machine learning › Generative modeling › diffusion model
rectified flow
0.912025
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection · ICCV 2025
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.912025
Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection · ICCV 2025
Natural language and speech › Language models and text generation › alignment
preference alignment
0.812024
Aligning Diffusion Models by Optimizing Human Utility · NeurIPS 2024
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model alignment
0.812024
Aligning Diffusion Models by Optimizing Human Utility · NeurIPS 2024
Computer vision › Segmentation and scene understanding › image segmentation
hierarchical segmentation
0.712023
Hierarchical Open-vocabulary Universal Image Segmentation · NeurIPS 2023
Computer vision › Segmentation and scene understanding
image segmentation
0.712023
Hierarchical Open-vocabulary Universal Image Segmentation · NeurIPS 2023
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.712023
Hierarchical Open-vocabulary Universal Image Segmentation · NeurIPS 2023
Computer vision › Vision and language › multimodal fusion
text-image fusion
0.712023
Hierarchical Open-vocabulary Universal Image Segmentation · NeurIPS 2023
Natural language and speech › Question answering and dialogue systems
multi-turn dialogue
0.312025
SegLLM: Multi-round Reasoning Segmentation with Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation
text generation
0.312025
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025
Audio and music processing
sound synthesis
0.312025
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025
Human-robot interaction › robot navigation
robot approach behavior
0.212015
May I help you?: Design of Human-like Polite Approaching Behavior · HRI 2015
Human-robot interaction
intention recognition
0.112015
May I help you?: Design of Human-like Polite Approaching Behavior · HRI 2015
Human-robot interaction
service robot
0.112015
May I help you?: Design of Human-like Polite Approaching Behavior · HRI 2015

Methods — techniques the papers use, named apart from their topics

multimodal transformer · 1.7classifier-free guidance · 1.7timestep shifting · 0.9prefix KV cache · 0.9mask-aware multimodal LLM · 0.9inference-time scaling · 0.9in-context reflection · 0.9conversational memory · 0.9complementary masking · 0.9binary feedback · 0.8field study · 0.2
YearPublicationVenuePosition
2025 OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
abstract
We introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis. OmniFlow advances the rectified flow (RF) framework used in text-to-image models to handle the joint distribution of multiple modalities. It outperforms previous any-to-any models on a wide range of tasks, such as text-to-image and text-to-audio synthesis. Our work offers three key contributions: First, we extend RF to a multi-modal setting and introduce a novel guidance mechanism, enabling users to flexibly control the alignment between different modalities in the generated outputs. Second, we propose a novel architecture that extends the text-to-image MMDiT architecture of Stable Diffusion 3 and enables audio and text generation. The extended modules can be efficiently pretrained individually and merged with the vanilla text-to-image MMDiT for fine-tuning. Lastly, we conduct a comprehensive study of the design choices of rectified flow transformers for large-scale audio and text generation, providing valuable insights into optimizing performance across various modalities. Code is available at https://github.com/jacklishufan/OmniFlows.
Konstantinos Kallidromitis, Akash Gokul, Zichun Liao, Yusuke Kato, Kazuki Kozuka, Aditya Grover
CVPR5
2025 Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru, Yusuke Kato, Kazuki Kozuka, Aditya Grover
ICCV5
2025 SegLLM: Multi-round Reasoning Segmentation with Large Language Models
abstract
We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual outputs. By leveraging a mask-aware multimodal LLM, SegLLM re-integrates previous segmentation results into its input stream, enabling it to reason about complex user intentions and segment objects in relation to previously identified entities, including positional, interactional, and hierarchical relationships, across multiple interactions. This capability allows SegLLM to respond to visual and text queries in a chat-like manner. Evaluated on the newly curated MRSeg benchmark, SegLLM outperforms existing methods in multi- round interactive reasoning segmentation by over 20%. Additionally, we observed that training on multi-round reasoning segmentation data enhances performance on standard single-round referring segmentation and localization tasks, resulting in a 5.5% increase in cIoU for referring expression segmentation and a 4.5% improvement in [email protected] for referring expression localization.
Xudong Wang 0007, Shaolun Zhang, Konstantinos Kallidromitis, Yusuke Kato, Kazuki Kozuka, Trevor Darrell
ICLR6
2025 LaViDa: A Large Diffusion Language Model for Multimodal Understanding
abstract
Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning. In real-world scenarios, desirable properties for VLMs include fast inference and controllable generation (e.g., constraining outputs to adhere to a desired format). However, existing autoregressive (AR) VLMs like LLaVA struggle in these aspects. Discrete diffusion models (DMs) offer a promising alternative, enabling parallel decoding for faster inference and bidirectional context for controllable generation through text-infilling. While effective in language-only settings, DMs' potential for multimodal tasks is underexplored. We introduce LaViDa, a family of VLMs built on DMs. We build LaViDa by equipping DMs with a vision encoder and jointly fine-tune the combined parts for multimodal instruction following. To address challenges encountered, LaViDa incorporates novel techniques such as complementary masking for effective training, prefix KV cache for efficient inference, and timestep shifting for high-quality sampling. Experiments show that LaViDa achieves competitive or superior performance to AR VLMs on multi-modal benchmarks such as MMMU, while offering unique advantages of DMs, including flexible speed-quality tradeoff, controllability, and bidirectional reasoning. On COCO captioning, LaViDa surpasses Open-LLaVa-Next-8B by +4.1 CIDEr with 1.92x speedup. On bidirectional tasks, it achieves +59% improvement on Constrained Poem Completion. These results demonstrate LaViDa as a strong alternative to AR VLMs. Code and models is available at https://github.com/jacklishufan/LaViDa
Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Jason Kuen, Kai-Wei Chang 0001, Aditya Grover
NeurIPS5
2024 Speech Recognition for Indigenous Language Using Self-Supervised Learning and Natural Language Processing
Satoshi Tamura, Tomohiro Hattori, Yusuke Kato, Naoki Noguchi
ICPRAM3
2024 Aligning Diffusion Models by Optimizing Human Utility
abstract
We present Diffusion-KTO, a novel approach for aligning text-to-image diffusion models by formulating the alignment objective as the maximization of expected human utility. Unlike previous methods, Diffusion-KTO does not require collecting pairwise preference data nor training a complex reward model. Instead, our objective uses per-image binary feedback signals, e.g. likes or dislikes, to align the model with human preferences. After fine-tuning using Diffusion-KTO, text-to-image diffusion models exhibit improved performance compared to existing techniques, including supervised fine-tuning and Diffusion-DPO, both in terms of human judgment and automatic evaluation metrics such as PickScore and ImageReward. Overall, Diffusion-KTO unlocks the potential of leveraging readily available per-image binary preference signals and broadens the applicability of aligning text-to-image diffusion models with human preferences.
Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato, Kazuki Kozuka
NeurIPS4
2023 Hierarchical Open-vocabulary Universal Image Segmentation
abstract
Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions. However, complex visual scenes can be naturally decomposed into simpler parts and abstracted at multiple lev4 els of granularity, introducing inherent segmentation ambiguity. Unlike existing methods that typically sidestep this ambiguity and treat it as an external factor, our approach actively incorporates a hierarchical representation encompassing different semantic-levels into the learning process. We propose a decoupled text-image fusion mechanism and representation learning modules for both “things” and “stuff”. Additionally, we systematically examine the differences that exist in the textual and visual features between these types of categories. Our resulting model, named HIPIE, tackles HIerarchical, oPen-vocabulary, and unIvErsal segmentation tasks within a unified framework. Benchmarked on diverse datasets, e.g., ADE20K,COCO, Pascal-VOC Part, and RefCOCO/RefCOCOg, HIPIE achieves the state-of14 the-art results at various levels of image comprehension, including semantic-level (e.g., semantic segmentation), instance-level (e.g., panoptic/referring segmentationand object detection), as well as part-level (e.g., part/subpart segmentation) tasks.
Xudong Wang 0007, Konstantinos Kallidromitis, Yusuke Kato, Kazuki Kozuka, Trevor Darrell
NeurIPS4
2021 Foot-Based 6-DOF Haptic Interface with Force Feedback Capability for Third Arm Manipulation
abstract
In this paper, we propose a 6-DOF haptic interface with force feedback capability for foot-based interaction. To direct a "third arm" to any position that the operator wants to reach while both hands are busy, the controller needs 6-DOF input from a modality that does not rely on the hands. We focus on foot-based operation and propose a device that is composed of a parallel link and omni wheel. We evaluated the operating performance of this device during a robot manipulation task, an obstacle avoidance task, and a Fitts’ Law task. The results show the precision of the 6-DOF manipulations to be 1.83 cm. The performance of the obstacle avoidance task is significantly higher with force feedback. The performance measure of the Fitts’ Law task, the throughput, differed depending on the plane of operation. In comparison with prior study, the throughput of 1.21 bits/s, the average of all planes, was higher than the prior study. The proposed device could be effective as an interface to control the third arm.
Seigo Okada, Yasunao Okazaki, Yusuke Kato, Jun Ozawa, Takeshi Ando
SMC3
2019 Adjusting Weight of Action Decision in Exploration for Logistics Warehouse Picking Learning
abstract
The purpose of this study is for a robot to learn picking motions in a logistics warehouse environment. The picking operation performed by a robot often fails owing to the inclination of items placed on a shelf, as well as the minimum clearance between the products and their vinyl packaging. Therefore, we considered acquiring a specific motion trajectory by reinforcement learning. However, because numerous types of items are handled in logistics warehouses, efficient learning is required. Therefore, in this research, we propose a method to efficiently exploration for learning picking an object by determining a focus exploration area for learning based on previous results of different objects.
Yusuke Kato, Tomoaki Nakamura, Takayuki Nagai, Natsuki Yamanobe, Kazuyuki Nagata, Jun Ozawa
IROS1
2018 A Convolution Neural Network Based Nursing-Care Text Classification Model with a New Filter for Expressing Dependency Relations of Words
abstract
In this paper, a convolution neural network (CNN) based text classification method is proposed. CNNs show strong performance for computer vision and speech recognition applications. Recently, in some researches, CNNs have been applied to sentence classification applications. Currently, we have studied nursing-care text classification to improve nursing-care quality in Japan. In our former works, several types of feature definitions have been proposed and examined by some classification models like SVMs. In this paper, a single layer CNN is used for classifying nursing-care texts. Each nursing-care text is represented as a concatenated word vectors. Each word is represented as a fixed length word vector which is obtained by the word2vec [1]-[4]. Then, nursing-care texts are classified using a two-dimensional CNN-based classification method. The proposed CNN has a new kind of filters which extracts dependency relation between words. From our experimental results, the proposed CNN-based method obtained better performance than our former works.
Manabu Nii, Yuya Tuchida, Yusuke Kato, Atsuko Uchinuno, Reiko Sakashita
SMC3
2017 Sit-to-stand assistance system based on using EMG to predict movement
abstract
We propose herein a method to predict the sit-to-stand movement before a user leaves their seat. The proposed method is evaluated by using it for sit-to-stand and noisy movements, and the sit-to-stand movement is predicted with an average accuracy of 99.5%. Furthermore, based on this proposed method, we develop a prototype system to assist the sit-to-stand movement. To verify the effectiveness of this system, we test it with seven subjects. The results show that, based on the predicted movement, the assist starts about 130.4 ms before the user leaves the seat. In addition, the results confirm that, when using this assistance system, muscle activity is reduced by about 46% compared with the unassisted sit-to-stand movement.
Takahiro Hiyama, Yusuke Kato, Tsuyoshi Inoue
RO-MAN2
2015 May I help you?: Design of Human-like Polite Approaching Behavior
abstract
When should service staff initiate interaction with a visitor? Neither simply-proactive (e.g. talk to everyone in a sight) nor passive (e.g. wait until being talked to) strategies are desired. This paper reports our modeling of polite approaching behavior. In a shopping mall, there are service staff members who politely approach visitors who need help. Our analysis revealed that staff members are sensitive to "intentions" of nearby visitors. That is, when a visitor intends to talk to a staff member and starts to approach, the staff member also walks a few steps toward the visitors in advance to being talked. Further, even when not being approached, staff members exhibit "availability" behavior in the case that a visitor's intention seems uncertain. We modeled these behaviors that are adaptive to pedestrians' intentions, occurred prior to initiation of conversation. The model was implemented into a robot and tested in a real shopping mall. The experiment confirmed that the proposed method is less intrusive to pedestrians, and that our robot successfully initiated interaction with pedestrians.
Yusuke Kato, Takayuki Kanda 0001, Hiroshi Ishiguro
HRI1
2005 Construction method of acoustic models dealing with various background noises based on combination of HMMs
Motoyuki Suzuki, Yusuke Kato, Akinori Ito, Shozo Makino
INTERSPEECH2