EDBT 2026 Demo / reviewers in the wild / expert
Qing Yu 0013
dblp:92/3937-13
· DBLP profile ↗
19ranked-venue papers
9as first author
13since 2021 · last 2025
0000-0001-6965-9581ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal ModelsabstractThis paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed Unsolvable Problem Detection (UPD). Multiple-choice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM’s ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs. Atsuyuki Miyai, Jingyang Zhang, Yifei Ming, Qing Yu 0013, Go Irie, Yixuan Li 0001, Hai Li 0001, Ziwei Liu 0002, Kiyoharu Aizawa |
ACL (1) | 5 |
| 2025 | A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language ModelsabstractOut-of-distribution (OOD) detection is a task that detects OOD samples during inference to ensure the safety of deployed models. However, conventional benchmarks have reached performance saturation, making it difficult to compare recent OOD detection methods. To address this challenge, we introduce three novel OOD detection benchmarks that enable a deeper understanding of method characteristics and reflect real-world conditions. First, we present ImageNet-X, designed to evaluate performance under challenging semantic shifts. Second, we propose ImageNet-FS-X for full-spectrum OOD detection, assessing robustness to covariate shifts (feature distribution shifts). Finally, we propose Wilds-FS-X, which extends these evaluations to real-world datasets, offering a more comprehensive testbed. Our experiments reveal that recent CLIP-based OOD detection methods struggle to varying degrees across the three proposed benchmarks, and none of them consistently outperforms the others. We hope the community goes beyond specific benchmarks and includes more challenging conditions reflecting real-world scenarios. The code is https://github.com/hoshi23/OOD-X-Benchmarks. Shiho Noda, Atsuyuki Miyai, Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
ICIP | 3 |
| 2025 | Open-set domain adaptation with visual-language foundation models
Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
Comput. Vis. Image Underst. | 1 |
| 2025 | GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution DetectionabstractAbstract Zero-shot OOD detection is a task that detects OOD images during inference with only in-distribution (ID) class names. Existing methods assume ID images contain a single, centered object, and do not consider the more realistic multi-object scenarios, where both ID and OOD objects are present. To meet the needs of many users, the detection method must have the flexibility to adapt the type of ID images. To this end, we present Global-Local Maximum Concept Matching (GL-MCM), which incorporates local image scores as an auxiliary score to enhance the separability of global and local visual features. Due to the simple ensemble score function design, GL-MCM can control the type of ID images with a single weight parameter. Experiments on ImageNet and multi-object benchmarks demonstrate that GL-MCM outperforms baseline zero-shot methods and is comparable to fully supervised methods. Furthermore, GL-MCM offers strong flexibility in adjusting the target type of ID images. The code is available via https://github.com/AtsuMiyai/GL-MCM . Atsuyuki Miyai, Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
Int. J. Comput. Vis. | 2 |
| 2024 | Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models
Kent Fujiwara, Mikihiro Tanaka, Qing Yu 0013 |
ECCV (58) | 3 |
| 2024 | Self-Labeling Framework for Open-Set Domain Adaptation With Few Labeled SamplesabstractUnsupervised domain adaptation (UDA) is extremely effective for transferring knowledge from a label-rich source domain to a label-scarce target domain. Because the target domain is unlabeled and may contain additional novel classes, open-set domain adaptation (ODA) has been suggested as a possible solution to detect these novel classes in the training phase. However, existing ODA methods rely heavily on abundant fully labeled source data, which are expensive to collect in specific applications and may also contain novel classes. In this study, we propose a novel self-labeling framework with prototypical contrastive learning and mutual information maximization to achieve ODA even when the amount of labeled data is very small, which is a new problem setting named few-shot ODA (FODA). We use self-supervised prototypical contrastive learning to train the network to learn the representations of source and target samples and maximize the mutual information between labels and input data to simultaneously recognize known and novel classes in the source and target domains. We evaluated our strategy in several domain adaptation environments and found that our method performed far better than existing approaches. Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
IEEE Trans. Multim. | 1 |
| 2023 | Frame-Level Label Refinement for Skeleton-Based Weakly-Supervised Action RecognitionabstractIn recent years, skeleton-based action recognition has achieved remarkable performance in understanding human motion from sequences of skeleton data, which is an important medium for synthesizing realistic human movement in various applications. However, existing methods assume that each action clip is manually trimmed to contain one specific action, which requires a significant amount of effort for annotation. To solve this problem, we consider a novel problem of skeleton-based weakly-supervised temporal action localization (S-WTAL), where we need to recognize and localize human action segments in untrimmed skeleton videos given only the video-level labels. Although this task is challenging due to the sparsity of skeleton data and the lack of contextual clues from interaction with other objects and the environment, we present a frame-level label refinement framework based on a spatio-temporal graph convolutional network (ST-GCN) to overcome these difficulties. We use multiple instance learning (MIL) with video-level labels to generate the frame-level predictions. Inspired by advances in handling the noisy label problem, we introduce a label cleaning strategy of the frame-level pseudo labels to guide the learning process. The network parameters and the frame-level predictions are alternately updated to obtain the final results. We extensively evaluate the effectiveness of our learning approach on skeleton-based action recognition benchmarks. The state-of-the-art experimental results demonstrate that the proposed method can recognize and localize action segments of the skeleton data. Qing Yu 0013, Kent Fujiwara |
AAAI | 1 |
| 2023 | Noise-Avoidance Sampling for Annotation Missing Object DetectionabstractExcellent results can be achieved using object detection with fully supervised training on large well-annotated datasets. However, the problem of missing annotations in real-world datasets can considerably reduce the performance of object detectors. In this study, we thoroughly analyze the effect of missing annotations on both positive and negative samples in object detector training. To mitigate the negative impact caused by annotation missing problem, we propose a simple yet effective method, noise-avoidance sampling, to distinguish noisy training samples and subsequently reduce their negative impact. Experiments are conducted on the PASCAL VOC 07+12 dataset with varying levels of missing annotations. The results reveal that the proposed method achieves comparable or superior performance with state-of-the-art methods. Jiafeng Mao, Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
ICIP | 2 |
| 2023 | LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt LearningabstractWe present a novel vision-language prompt learning approach for few-shot out-of-distribution (OOD) detection. Few-shot OOD detection aims to detect OOD images from classes that are unseen during training using only a few labeled in-distribution (ID) images. While prompt learning methods such as CoOp have shown effectiveness and efficiency in few-shot ID classification, they still face limitations in OOD detection due to the potential presence of ID-irrelevant information in text embeddings. To address this issue, we introduce a new approach called $\textbf{Lo}$cal regularized $\textbf{Co}$ntext $\textbf{Op}$timization (LoCoOp), which performs OOD regularization that utilizes the portions of CLIP local features as OOD features during training. CLIP's local features have a lot of ID-irrelevant nuisances ($\textit{e.g.}$, backgrounds), and by learning to push them away from the ID class text embeddings, we can remove the nuisances in the ID class text embeddings and enhance the separation between ID and OOD. Experiments on the large-scale ImageNet OOD detection benchmarks demonstrate the superiority of our LoCoOp over zero-shot, fully supervised detection methods and prompt learning methods. Notably, even in a one-shot setting -- just one label per class, LoCoOp outperforms existing zero-shot and fully supervised detection methods. The code is available via https://github.com/AtsuMiyai/LoCoOp. Atsuyuki Miyai, Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
NeurIPS | 2 |
| 2023 | Rethinking Rotation in Self-Supervised Contrastive Learning: Adaptive Positive or Negative Data AugmentationabstractRotation is frequently listed as a candidate for data augmentation in contrastive learning but seldom provides satisfactory improvements. We argue that this is because the rotated image is always treated as either positive or negative. The semantics of an image can be rotation-invariant or rotation-variant, so whether the rotated image is treated as positive or negative should be determined based on the content of the image. Therefore, we propose a novel augmentation strategy, adaptive Positive or Negative Data Augmentation (PNDA), in which an original and its rotated image are a positive pair if they are semantically close and a negative pair if they are semantically different. To achieve PNDA, we first determine whether rotation is positive or negative on an image-by-image basis in an unsupervised way. Then, we apply PNDA to contrastive learning frameworks. Our experiments showed that PNDA improves the performance of contrastive learning. The code is available at https://github.com/AtsuMiyai/rethinking_rotation. Atsuyuki Miyai, Qing Yu 0013, Daiki Ikami, Go Irie, Kiyoharu Aizawa |
WACV | 2 |
| 2022 | Self-Labeling Framework for Novel Category Discovery over DomainsabstractUnsupervised domain adaptation (UDA) has been highly successful in transferring knowledge acquired from a label-rich source domain to a label-scarce target domain. Open-set domain adaptation (open-set DA) and universal domain adaptation (UniDA) have been proposed as solutions to the problem concerning the presence of additional novel categories in the target domain. Existing open-set DA and UniDA approaches treat all novel categories as one unified unknown class and attempt to detect this unknown class during the training process. However, the features of the novel categories learned by these methods are not discriminative. This limits the applicability of UDA in the further classification of these novel categories into their original categories, rather than assigning them to a single unified class. In this paper, we propose a self-labeling framework to cluster all target samples, including those in the ''unknown'' categories. We train the network to learn the representations of target samples via self-supervised learning (SSL) and to identify the seen and unseen (novel) target-sample categories simultaneously by maximizing the mutual information between labels and input data. We evaluated our approach under different DA settings and concluded that our method generally outperformed existing ones by a wide margin. Qing Yu 0013, Daiki Ikami, Go Irie, Kiyoharu Aizawa |
AAAI | 1 |
| 2021 | Noisy Annotation Refinement for Object Detection
Jiafeng Mao, Qing Yu 0013, Yoko Yamakata, Kiyoharu Aizawa |
BMVC | 2 |
| 2021 | Divergence Optimization for Noisy Universal Domain AdaptationabstractUniversal domain adaptation (UniDA) has been proposed to transfer knowledge learned from a label-rich source domain to a label-scarce target domain without any constraints on the label sets. In practice, however, it is difficult to obtain a large amount of perfectly clean labeled data in a source domain with limited resources. Existing UniDA methods rely on source samples with correct annotations, which greatly limits their application in the real world. Hence, we consider a new realistic setting called Noisy UniDA, in which classifiers are trained with noisy labeled data from the source domain and unlabeled data with an unknown class distribution from the target domain. This paper introduces a two-head convolutional neural network framework to solve all problems simultaneously. Our network consists of one common feature generator and two classifiers with different decision boundaries. By optimizing the divergence between the two classifiers’ outputs, we can detect noisy source samples, find "unknown" classes in the target domain, and align the distribution of the source and target domains. In an extensive evaluation of different domain adaptation settings, the proposed method outperformed existing methods by a large margin in most settings. Qing Yu 0013, Atsushi Hashimoto 0001, Yoshitaka Ushiku |
CVPR | 1 |
| 2020 | Multi-task Curriculum Framework for Open-Set Semi-supervised Learning
Qing Yu 0013, Daiki Ikami, Go Irie, Kiyoharu Aizawa |
ECCV (12) | 1 |
| 2020 | Noisy Localization Annotation Refinement For Object DetectionabstractThe production of finely annotated datasets for object detection tasks is labor-intensive, therefore, cloud sourcing is often used to create datasets, which leads to these datasets tending to contain incorrect annotations such as inaccurate localization bounding boxes. In this study, we highlight a problem of object detection with noisy bounding box annotations and show that these noisy annotations are harmful to the performance of deep neural networks. To solve this problem, we further propose a framework to allow the network to modify the noisy datasets by alternating refinement. The experimental results demonstrate that our proposed framework can significantly alleviate the influences of noise on model performance. Jiafeng Mao, Qing Yu 0013, Kiyoharu Aizawa |
ICIP | 2 |
| 2020 | Unknown Class Label Cleaning For Learning With Open-Set Noisy LabelsabstractDeep neural networks (DNNs) trained on large-scale annotated datasets have achieved impressive results in the area of image classification. Many large-scale datasets have been collected from websites; however, such data are inevitably corrupted with noise. In this study, we researched the open-set noisy label problem, where some outliers are contained in a dataset and annotated through a noisy label but do not belong to any class of training data. To address this problem, we propose a novel unknown class label cleaning framework for the training of DNNs with open-set noisy labels. In addition to general image classification, we also estimate the probability of an input being from an unknown class by assigning a pseudo unknown label to all of the data and correct these labels through an alternating update of the network parameters and labels. The results of experiments conducted on the noisy CIFAR-10 datasets demonstrate that our approach can robustly train DNNs with a high proportion of noisy labels. Qing Yu 0013, Kiyoharu Aizawa |
ICIP | 1 |
| 2020 | The Aleatoric Uncertainty Estimation Using a Separate Formulation with Virtual ResidualsabstractWe propose a new optimization framework for aleatoric uncertainty estimation in regression problems. Existing methods can quantify the error in the target estimation, but they tend to underestimate it. To obtain the predictive uncertainty inherent in an observation, we propose a new separable formulation for the estimation of a signal and of its uncertainty, avoiding the effect of overfitting. By decoupling target estimation and uncertainty estimation, we also control the balance between signal estimation and uncertainty estimation. We conduct three types of experiments: regression with simulation data, age estimation, and depth estimation. We demonstrate that the proposed method outperforms a state-of-the-art technique for signal and uncertainty estimation. Takumi Kawashima, Qing Yu 0013, Akari Asai, Daiki Ikami, Kiyoharu Aizawa |
ICPR | 2 |
| 2019 | Unsupervised Out-of-Distribution Detection by Maximum Classifier DiscrepancyabstractSince deep learning models have been implemented in many commercial applications, it is important to detect out-of-distribution (OOD) inputs correctly to maintain the performance of the models, ensure the quality of the collected data, and prevent the applications from being used for other-than-intended purposes. In this work, we propose a two-head deep convolutional neural network (CNN) and maximize the discrepancy between the two classifiers to detect OOD inputs. We train a two-head CNN consisting of one common feature extractor and two classifiers which have different decision boundaries but can classify in-distribution (ID) samples correctly. Unlike previous methods, we also utilize unlabeled data for unsupervised training and we use these unlabeled data to maximize the discrepancy between the decision boundaries of two classifiers to push OOD samples outside the manifold of the in-distribution (ID) samples, which enables us to detect OOD samples that are far from the support of the ID samples. Overall, our approach significantly outperforms other state-of-the-art methods on several OOD detection benchmarks and two cases of real-world simulation. Qing Yu 0013, Kiyoharu Aizawa |
ICCV | 1 |
| 2018 | Food Image Recognition by Personalized ClassifierabstractSince the development of food diaries could enable people to develop healthy eating habits, food image recognition is in high demand to reduce the effort in food recording. Previous studies have worked on this challenging domain with datasets having fixed numbers of samples and classes. However, in the real-world setting, it is impossible to include all of the foods in the database because the number of classes of foods is large and increases continually. In addition to that, inter-class similarity and intra-class diversity also bring difficulties to the recognition. In this paper, we attempted to solve these problems by using deep convolutional neural network features to build a personalized classifier which incrementally learns the user's data and adapts to the user's eating habit. As a result, we achieved the state-of-the-art accuracy of food image recognition by the personalization of 300 food records per user. Qing Yu 0013, Masashi Anzawa, Sosuke Amano, Makoto Ogawa, Kiyoharu Aizawa |
ICIP | 1 |