Yu Mitsuzumi

dblp:226/3395 · DBLP profile ↗
← Back
11ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0001-8803-8748ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Accelerating Graph Construction for MIPS without Search Accuracy Loss
Yasuhiro Fujiwara, Ángel López García-Arias, Yu Mitsuzumi, Yasutoshi Ida, Atsutoshi Kumagai, Masahiro Nakano, Makoto Nakatsuji, Akisato Kimura
EDBT3
2026 Model-free Domain Adaptation for Concealed Multimodal Large-Language Models
abstract
Multimodal large-language models (MLLMs) exhibit remarkable capability for various vision tasks but still struggle with the domain-shift problem, in which their performance degrades for data from unfamiliar domains. Since the latest MLLMs often conceal their model resources (i.e., data, parameters, and outputs) from training purposes, current domain adaptation methods cannot satisfactorily address this problem due to their dependence on those resources. To this end, we introduce a novel domain adaptation setup, "model-free domain adaptation (MFDA)" of MLLMs, to investigate whether we can address domain adaptation problems without using any resources of the concealed models. As a proof of concept for MFDA, we built a method named model-transferable domain-adaptable visual prompting (MTDA-VP). In the training, this method executes cross-model visual prompting on surrogate models with a domain adaptation objective so that the visual prompts simultaneously acquire model transferability and domain adaptability. In the testing, we can adapt the concealed MLLMs to the target domain by just inputting the test images with the trained prompt into the models. Besides, we developed two techniques, cross-model pseudo labeling (CMPL) and cross-model gradient alignment (CMGA), to further enhance model transferability and domain adaptability of the visual prompts. We empirically confirmed that MFDA-VP improved the performance of several MLLMs with large margins on two datasets.
Yu Mitsuzumi, Akisato Kimura, Hisashi Kashima
WACV1
2025 Flexible Source-free Domain Generalization via Domain Prompt-Discriminator Collaborative Learning
abstract
Source-free domain generalization (SFDG) is an emerging paradigm of domain generalization problems that aims to train a model generalizable to unseen target domains without any actual source data. Existing methods have pragmatically tackled this challenging problem by using large-scale vision-language pre-trained models (VLMs). However, they cannot fully exploit the potential of VLMs. Although the favorable configurations should inherently differ for each input domain, those methods employ a domain-shared configuration (e.g., classifier, representation) for data inputs from any domain, which undermines the flexibility for various domains. In this paper, we propose a novel SFDG method called Domain Prompt-Discriminator Collaborative Learning. Our method endows the model with domain flexibility by jointly training the following two modules: (1) various domain-specific prompts to enhance generalizability to unseen domains and (2) a domain discriminator to choose favorable domains for the input image. Notably, this training can be carried out in the form of simple domain classification learning. Our method design also opens the door to a simple yet effective test-time adaptation technique that further boosts recognition accuracy. We empirically demonstrate that our method consistently outperforms state-of-the-art SFDG methods on four domain generalization benchmarks.
Yu Mitsuzumi, Akisato Kimura, Hisashi Kashima
IJCNN1
2024 Understanding and Improving Source-Free Domain Adaptation from a Theoretical Perspective
abstract
Source-free Domain Adaptation (SFDA) is an emerging and challenging research area that addresses the problem of unsupervised domain adaptation (UDA) without source data. Though numerous successful methods have been proposed for SFDA, a theoretical understanding of why these methods work well is still absent. In this paper, we shed light on the theoretical perspective of existing SFDA methods. Specifically, we find that SFDA loss functions comprising discriminability and diversity losses work in the same way as the training objective in the theory of self-training based on the expansion assumption, which shows the existence of the target error bound. This finding brings two novel insights that enable us to build an improved SFDA method comprising 1) Model Training with Auto-Adjusting Diversity Constraint and 2) Augmentation Training with Teacher-Student Framework, yielding a better recognition performance. Extensive experiments on three benchmark datasets demonstrate the validity of the theoretical analysis and our method.
Yu Mitsuzumi, Akisato Kimura, Hisashi Kashima
CVPR1
2024 Cross-Action Cross-Subject Skeleton Action Recognition Via Simultaneous Action-Subject Learning With Two-Step Feature Removal
abstract
In this paper, we tackle a novel skeleton-based action recognition problem named Cross-Action Cross-Subject (CACS) Skeleton Action Recognition, where we can access the data of only a part of the target action classes for each training subject. Existing skeleton-based action recognition methods suffer from solving this problem because there are scarce clues to resolve the cross-entanglement of action and subject information, and the trained model will confuse those two features. To solve this challenging problem, we propose a method that consists of simultaneous action-subject learning with feature removal. In our method, 1) we use two data augmentation techniques, Bone Randomization and Phase Randomization, to roughly remove unnecessary features for respective recognitions, and then, 2) we introduce a debiased learning approach to remove the confusing features by minimizing mutual information with an action-subject-shared discriminator network. Extensive experiments on three datasets demonstrate that our method is consistently effective for several CACS problems.
Yu Mitsuzumi, Akisato Kimura, Go Irie, Atsushi Nakazawa
ICIP1
2024 Phase Randomization: A data augmentation for domain adaptation in human action recognition
Yu Mitsuzumi, Go Irie, Akisato Kimura, Atsushi Nakazawa
Pattern Recognit.1
2021 Generalized Domain Adaptation
abstract
Many variants of unsupervised domain adaptation (UDA) problems have been proposed and solved individually. Its side effect is that a method that works for one variant is often ineffective for or not even applicable to another, which has prevented practical applications. In this paper, we give a general representation of UDA problems, named Generalized Domain Adaptation (GDA). GDA covers the major variants as special cases, which allows us to organize them in a comprehensive framework. Moreover, this generalization leads to a new challenging setting where existing methods fail, such as when domain labels are unknown, and class labels are only partially given to each domain. We propose a novel approach to the new setting. The key to our approach is self-supervised class-destructive learning, which enables the learning of class-invariant representations and domain-adversarial classifiers without using any domain labels. Extensive experiments using three benchmark datasets demonstrate that our method outperforms the state-of-the-art UDA methods in the new setting and that it is competitive in existing UDA variations as well.
Yu Mitsuzumi, Go Irie, Daiki Ikami, Takashi Shibata 0001
CVPR1
2021 Learning with Selective Forgetting
abstract
Lifelong learning aims to train a highly expressive model for a new task while retaining all knowledge for previous tasks. However, many practical scenarios do not always require the system to remember all of the past knowledge. Instead, ethical considerations call for selective and proactive forgetting of undesirable knowledge in order to prevent privacy issues and data leakage. In this paper, we propose a new framework for lifelong learning, called Learning with Selective Forgetting, which is to update a model for the new task with forgetting only the selected classes of the previous tasks while maintaining the rest. The key is to introduce a class-specific synthetic signal called mnemonic code. The codes are "watermarked" on all the training samples of the corresponding classes when the model is updated for a new task. This enables us to forget arbitrary classes later by only using the mnemonic codes without using the original data. Experiments on common benchmark datasets demonstrate the remarkable superiority of the proposed method over several existing methods.
Takashi Shibata 0001, Go Irie, Daiki Ikami, Yu Mitsuzumi
IJCAI4
2020 A Generative Self-Ensemble Approach To Simulated+Unsupervised Learning
abstract
In this paper, we consider Simulated and Unsupervised (S+U) learning which is a problem of learning from labeled synthetic and unlabeled real images. After translating the synthetic images to real ones, existing S+U learning methods use only the labeled synthetic images for training a predictor (e.g., a regression function) and ignore the target real images, which may result in unsatisfactory prediction performance. Our approach utilizes both synthetic and real images to train the predictor. The main idea of ours is to involve a self-ensemble learning framework into S+U learning. More specifically, we require the prediction results for an unlabeled real image to be consistent between “teacher” and “student” predictors, even after some perturbations are added to the image. Furthermore, aiming at generating diverse perturbations along the underlying data manifold, we introduce one-to-many image translation between synthetic and real images. Evaluation experiments on an appearance-based gaze estimation task demonstrate that the proposed ideas can improve the prediction accuracy and our full method can outperform existing S+U learning methods.
Yu Mitsuzumi, Go Irie, Akisato Kimura, Atsushi Nakazawa
ICIP1
2018 Eye Contact Detection Algorithms Using Deep Learning and Generative Adversarial Networks
abstract
Eye contact (mutual gaze) is a foundation of human communication and social interactions; therefore, it is studied in many fields such as psychology, social science, and medicine. Our group have been studied wearable vision-based eye contact detection techniques using a first person camera for the purpose of evaluating the gaze skills in the tender dementia care. In this work, we search for deep learning-based eye contact detection techniques from small number of labeled images. We implemented and tested two eye contact detection algorithms: naïve deep-learning-based algorithm and generative adversarial networks (GAN)-based semi supervised learning (SSL) algorithm. These methods are learned and verified by using Columbia Gaze Dataset, Facescrub and our original datasets. The results show the effectiveness and limitations of the deep-learning-based and GAN-based approaches. Interestingly, we found the bilateral difference of the accuracy of eye contact detection with respect to the facial pose with respect to the camera, which is expected to be caused by the learning datasets.
Yu Mitsuzumi, Atsushi Nakazawa
SMC1
2017 DEEP eye contact detector: Robust eye contact bid detection using convolutional neural network
Yu Mitsuzumi, Atsushi Nakazawa, Toyoaki Nishida
BMVC1