Kai Katsumata

dblp:280/6265 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0001-9729-2588ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-PrefDrive: Optimizing Large Language Models for Autonomous Driving Through Multi-Preference Tuning
abstract
This paper introduces Multi-PrefDrive, a framework that significantly enhances LLM-based autonomous driving through multidimensional preference tuning. Aligning LLMs with human driving preferences is crucial yet challenging, as driving scenarios involve complex decisions where multiple incorrect actions can correspond to a single correct choice. Traditional binary preference tuning fails to capture this complexity. Our approach pairs each chosen action with multiple rejected alternatives, better reflecting real-world driving decisions. By implementing the Plackett-Luce preference model, we enable nuanced ranking of actions across the spectrum of possible errors. Experiments in the CARLA simulator demonstrate that our algorithm achieves an 11.0% improvement in overall score and an 83.6% reduction in infrastructure collisions, while showing perfect compliance with traffic signals in certain environments. Comparative analysis against DPO and its variants reveals that Multi-PrefDrive’s superior discrimination between chosen and rejected actions, which achieving a margin value of 25, and such ability has been directly translates to enhanced driving performance. We implement memory-efficient techniques including LoRA and 4-bit quantization to enable deployment on consumer-grade hardware and will open-source our training code and multi-rejected dataset to advance research in LLM-based autonomous driving systems. Project Page (https://liyun0607.github.io/).
Ehsan Javanmardi, Kai Katsumata, Alex Orsholits, Manabu Tsukada
IROS4
2025 PrefDrive: Enhancing Autonomous Driving Through Preference-Guided Large Language Models
abstract
This paper presents PrefDrive, a novel frame-work that integrates driving preferences into autonomous driving models through large language models (LLMs). While recent advances in LLMs have shown promise in autonomous driving, existing approaches often struggle to align with specific driving behaviors (e.g., maintaining safe distances, smooth acceleration patterns) and operational requirements (e.g., traffic rule compliance, route adherence). We address this challenge by developing a preference learning framework that combines multimodal perception with natural language understanding. Our approach leverages Direct Preference Optimization (DPO) to fine-tune LLMs efficiently on consumer-grade hardware, making advanced autonomous driving research more accessible to the broader research community. We introduce a comprehensive dataset of 74,040 sequences, carefully annotated with driving preferences and driving decisions, which, along with our trained model checkpoints, is made publicly available https://github.com/LiYun0607/PrefDrive/ to facilitate future research. Through extensive experiments in the CARLA simulator, we demonstrate that our preference-guided approach significantly improves driving performance across multiple metrics, including distance maintenance and trajectory smoothness. Results show up to 28.1% reduction in traffic light violations and 8.5% improvement in route completion while maintaining appropriate distances from obstacles. The framework demonstrates robust performance across different urban environments, showcasing the effectiveness of preference learning in autonomous driving applications.
Ehsan Javanmardi, Kai Katsumata, Alex Orsholits, Manabu Tsukada
IV4
2024 A Compact Dynamic 3D Gaussian Representation for Real-Time Dynamic View Synthesis
Kai Katsumata, Duc Minh Vo, Hideki Nakayama
ECCV (86)1
2024 Soft Curriculum for Learning Conditional GANs with Noisy-Labeled and Uncurated Unlabeled Data
abstract
Label-noise or curated unlabeled data are used to compensate for the assumption of clean labeled data in training the conditional generative adversarial network; however, satisfying such an extended assumption is occasionally laborious or impractical. As a step towards generative modeling accessible to everyone, we introduce a novel conditional image generation framework that accepts noisy-labeled and uncurated unlabeled data during training: (i) closed-set and open-set label noise in labeled data and (ii) closed-set and open-set unlabeled data. To combat it, we propose soft curriculum learning, which assigns instance-wise weights for adversarial training while assigning new labels for unlabeled data and correcting wrong labels for labeled data. Unlike popular curriculum learning, which uses a threshold to pick the training samples, our soft curriculum controls the effect of each training instance by using the weights predicted by the auxiliary classifier, resulting in the preservation of useful samples while ignoring harmful ones. Our experiments show that our approach outperforms existing semi-supervised and label-noise robust methods in terms of both quantitative and qualitative performance. In particular, the proposed approach matches the performance of (semi-)supervised GANs even with less than half the labeled data.1
Kai Katsumata, Duc Minh Vo, Tatsuya Harada, Hideki Nakayama
WACV1
2024 Revisiting Latent Space of GAN Inversion for Robust Real Image Editing
abstract
We present a generative adversarial network (GAN) inversion with high reconstruction and editing quality. GAN inversion algorithms with expressive latent spaces produce near-perfect inversion but are not robust to editing operations in a latent space, leading to undesirable edited images, a phenomenon known as the trade-off between reconstruction and editing quality. To cope with the trade-off, we revisit the hyperspherical prior of StyleGANs $\mathcal{Z}$ and propose to combine an extended space of $\mathcal{Z}$ with highly capable inversion algorithms. Our approach maintains the reconstruction quality of seminal GAN inversion methods while improving their editing quality owing to the constrained nature of $\mathcal{Z}$. Through comprehensive experiments with several GAN inversion algorithms, we demonstrate that our approach enhances the image editing quality in 2D/3D GANs.1
Kai Katsumata, Duc Minh Vo, Bei Liu 0001, Hideki Nakayama
WACV1
2024 Label Augmentation as Inter-class Data Augmentation for Conditional Image Synthesis with Imbalanced Data
abstract
Conditional image synthesis performs admirably when trained on well-constructed and balanced datasets. However, in practice, training datasets frequently contain minorities (i.e., a class with a few samples), known as imbalanced data, which causes difficulties in learning generative models. To address conditional image synthesis with imbalanced data, we analyze a diversity issue of label-preserving data augmentation and an affinity issue of non-label-preserving data augmentation. From this observation, we present label augmentation, which works as inter-class data augmentation that effectively augments data by predicting a new label for a given image using the prediction of a pretrained image classification model (i.e., probabilities for each class). We incorporate our label augmentation into the discriminator of a seminal conditional generative adversarial network (GAN) model, proposing Softlabel-GAN. Using class probabilities extracts class-invariant and shared features between similar classes, achieving data augmentation with high affinity and diversity. Our experiments on imbalanced datasets show that Softlabel-GAN produces images with high quality and diversity while being hardly affected by the number of samples in each class. Code: https://github.com/raven38/softlabel-gan.
Kai Katsumata, Duc Minh Vo, Hideki Nakayama
WACV1
2022 OSSGAN: Open-Set Semi-Supervised Image Generation
abstract
We introduce a challenging training scheme of conditional GANs, called open-set semi-supervised image generation, where the training dataset consists of two parts: (i) labeled data and (ii) unlabeled data with samples belonging to one of the labeled data classes, namely, a closed-set, and samples not belonging to any of the labeled data classes, namely, an open-set. Unlike the existing semi-supervised image generation task, where unlabeled data only contain closed-set samples, our task is more general and lowers the data collection cost in practice by allowing open-set samples to appear. Thanks to entropy regularization, the classifier that is trained on labeled data is able to quantify sample-wise importance to the training of cGAN as confidence, allowing us to use all samples in un-labeled data. We design OSSGAN, which provides decision clues to the discriminator on the basis of whether an unlabeled image belongs to one or none of the classes of interest, smoothly integrating labeled and unlabeled data during training. The results of experiments on Tiny ImageNet and ImageNet show notable improvements over supervised Big-GAN and semi-supervised methods. Our code is available at https://github.com/raven38/OSSGAN.
Kai Katsumata, Duc Minh Vo, Hideki Nakayama
CVPR1
2021 Semantic Image Synthesis from Inaccurate and Coarse Masks
abstract
Semantic image synthesis is an image-to-image translation problem where the goal is to learn mapping from semantic segmentation masks to corresponding photorealistic images. However, conventional semantic image synthesis methods require numerous pairs of correct semantic masks and real images, and collecting these pairs is not always possible. To address this issue, we propose a smoothing method, which we call local label smoothing (LLS), that incorporates label smoothing per small patch of an input mask to learn mapping from masks to images even when semantic masks are inaccurate. Furthermore, we also propose an extended method for coarse masks. We demonstrate the advantage of the proposed methods over existing methods to deal with noisy masks on several datasets.
Kai Katsumata, Hideki Nakayama
ICASSP1
2021 Open-Set Domain Generalization VIA Metric Learning
abstract
In this study, we address open-set domain generalization, which aims to reject unknown class samples while classifying known class samples in unseen domains. Conventional domain generalization has the problem of unknown class samples being classified as known classes because domain generalization methods align feature distributions without distinction between known and unknown classes. To tackle this problem, we propose a decoupling loss that diffuses the feature representations of unknown samples. The loss allows us to construct a feature space that can better distinguish unknown samples. We demonstrate the effectiveness of decoupling loss using open-set domain generalization benchmarks.
Kai Katsumata, Ikki Kishida, Ayako Amma, Hideki Nakayama
ICIP1