Jiahao Xiao

dblp:238/4029 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Learning paradigms · 77% Vision and language · 7% Knowledge representation and reasoning · 6%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms
semi-supervised learning
3.042026
Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Dual-Decoupling Learning and Metric-Adaptive Thresholding for Semi-supervised Multi-label Learning · ECCV (52) 2024
Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning · NeurIPS 2023
Machine learning › Learning paradigms
multi-label classification
2.742024
Counterfactual Reasoning for Multi-Label Image Classification via Patching-Based Training · ICML 2024
Dual-Decoupling Learning and Metric-Adaptive Thresholding for Semi-supervised Multi-label Learning · ECCV (52) 2024
Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning · NeurIPS 2023
Machine learning › Learning paradigms › semi-supervised learning
pseudo-labeling
2.232026
Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning · NeurIPS 2023
Label-Aware Global Consistency for Multi-Label Learning with Single Positive Labels · NeurIPS 2022
Machine learning › Learning paradigms › multi-label classification
semi-supervised multi-label learning
1.012026
Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Large multimodal models evaluation: a survey · Sci. China Inf. Sci. 2025
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.812024
Counterfactual Reasoning for Multi-Label Image Classification via Patching-Based Training · ICML 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning
0.812024
Counterfactual Reasoning for Multi-Label Image Classification via Patching-Based Training · ICML 2024
Machine learning › Learning theory
generalization bounds
0.312026
Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Learning paradigms › multi-label classification
label correlation modeling
0.212024
Counterfactual Reasoning for Multi-Label Image Classification via Patching-Based Training · ICML 2024

Methods — techniques the papers use, named apart from their topics

regularization · 1.0pseudo-labeling · 1.0class-aware thresholding · 1.0patching-based training · 0.8metric-adaptive thresholding · 0.8dual-decoupling learning · 0.8counterfactual reasoning · 0.8generalization error bound · 0.7class-aware thresholds · 0.7consistency regularization · 0.6
YearPublicationVenuePosition
2026 Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning
abstract
Pseudo-labeling has emerged as a popular and effective approach for utilizing unlabeled data. However, in the context of semi-supervised multi-label learning (SSMLL), conventional pseudo-labeling methods encounter difficulties when dealing with instances associated with multiple labels and an unknown label count. These limitations often result in the introduction of false positive labels or the neglect of true positive ones. To overcome these challenges, this paper proposes a novel solution called Class-distribution-Aware Pseudo-labeling (CAP) that performs pseudo-labeling in a class-aware manner. The proposed approach introduces a regularized learning framework incorporating class-aware thresholds, which effectively control the assignment of positive and negative pseudo-labels for each class. Notably, even with a small proportion of labeled examples, our observations demonstrate that the estimated class distribution serves as a reliable approximation. Motivated by this finding, we develop a class-distribution-aware thresholding (CAT) strategy to ensure the alignment of pseudo-label distribution with the true distribution. Moreover, we extend CAT into a label decision method, aiming to improve the model's classification performance during the testing phase. The correctness of the estimated class distribution is theoretically verified, and a generalization error bound is provided for our proposed method. Extensive experiments on multiple benchmark datasets confirm the efficacy of CAP in addressing the challenges of SSMLL problems.
Ming-Kun Xie, Jiahao Xiao, Hao-Zhe Liu, Gang Niu 0001, Masashi Sugiyama, Sheng-Jun Huang
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Adaptive Federated Distillation for Multi-Domain Non-IID Textual Data
abstract
The widespread success of pre-trained language models has established a new training paradigm, where a global PLM is fine-tuned using task-specific data from local clients. The local data are highly different from each other and can not capture the global distribution of the whole data in real world. To address the challenges of non-IID data in real environments, privacy-preserving federated distillation has been proposed and highly investigated. However, previous experimental non-IID scenarios are primarily identified with the label (output) diversity, without considering the diversity of language domains (input) that is crucial in natural language processing. In this paper, we introduce a comprehensive set of multi-domain non-IID scenarios and propose a unified benchmarking framework that includes diverse data. The benchmark can be used to evaluate the federated learning framework in a real environment. To this end, we propose an Adaptive Federated Distillation (AdaFD) framework designed to address multi-domain non-IID challenges in both homogeneous and heterogeneous settings. Experimental results demonstrate that our models capture the diversity of local clients and achieve better performance compared to the existing works. The code for this paper is available at: https://github.com/jiahaoxiao1228/AdaFD.
Jiahao Xiao, Jiangming Liu
IJCNN1
2025 Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai
Sci. China Inf. Sci.9
2024 Dual-Decoupling Learning and Metric-Adaptive Thresholding for Semi-supervised Multi-label Learning
Jiahao Xiao, Ming-Kun Xie, Heng-Bo Fan, Gang Niu 0001, Masashi Sugiyama, Sheng-Jun Huang
ECCV (52)1
2024 Counterfactual Reasoning for Multi-Label Image Classification via Patching-Based Training
abstract
The key to multi-label image classification (MLC) is to improve model performance by leveraging label correlations. Unfortunately, it has been shown that overemphasizing co-occurrence relationships can cause the overfitting issue of the model, ultimately leading to performance degradation. In this paper, we provide a causal inference framework to show that the correlative features caused by the target object and its co-occurring objects can be regarded as a mediator, which has both positive and negative impacts on model predictions. On the positive side, the mediator enhances the recognition performance of the model by capturing co-occurrence relationships; on the negative side, it has the harmful causal effect that causes the model to make an incorrect prediction for the target object, even when only co-occurring objects are present in an image. To address this problem, we propose a counterfactual reasoning method to measure the total direct effect, achieved by enhancing the direct effect caused only by the target object. Due to the unknown location of the target object, we propose patching-based training and inference to accomplish this goal, which divides an image into multiple patches and identifies the pivot patch that contains the target object. Experimental results on multiple benchmark datasets with diverse configurations validate that the proposed method can achieve state-of-the-art performance.
Ming-Kun Xie, Jiahao Xiao, Pei Peng 0005, Gang Niu 0001, Masashi Sugiyama, Sheng-Jun Huang
ICML2
2023 Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning
abstract
Pseudo-labeling has emerged as a popular and effective approach for utilizing unlabeled data. However, in the context of semi-supervised multi-label learning (SSMLL), conventional pseudo-labeling methods encounter difficulties when dealing with instances associated with multiple labels and an unknown label count. These limitations often result in the introduction of false positive labels or the neglect of true positive ones. To overcome these challenges, this paper proposes a novel solution called Class-Aware Pseudo-Labeling (CAP) that performs pseudo-labeling in a class-aware manner. The proposed approach introduces a regularized learning framework incorporating class-aware thresholds, which effectively control the assignment of positive and negative pseudo-labels for each class. Notably, even with a small proportion of labeled examples, our observations demonstrate that the estimated class distribution serves as a reliable approximation. Motivated by this finding, we develop a class-distribution-aware thresholding strategy to ensure the alignment of pseudo-label distribution with the true distribution. The correctness of the estimated class distribution is theoretically verified, and a generalization error bound is provided for our proposed method. Extensive experiments on multiple benchmark datasets confirm the efficacy of CAP in addressing the challenges of SSMLL problems.
Ming-Kun Xie, Jiahao Xiao, Hao-Zhe Liu, Gang Niu 0001, Masashi Sugiyama, Sheng-Jun Huang
NeurIPS2
2022 Label-Aware Global Consistency for Multi-Label Learning with Single Positive Labels
abstract
In single positive multi-label learning (SPML), only one of multiple positive labels is observed for each instance. The previous work trains the model by simply treating unobserved labels as negative ones, and designs the regularization to constrain the number of expected positive labels. However, in many real-world scenarios, the true number of positive labels is unavailable, making such methods less applicable. In this paper, we propose to solve SPML problems by designing a Label-Aware global Consistency (LAC) regularization, which leverages the manifold structure information to enhance the recovery of potential positive labels. On one hand, we first perform pseudo-labeling for each unobserved label based on its prediction probability. The consistency regularization is then imposed on model outputs to balance the fitting of identified labels and exploring of potential positive labels. On the other hand, by enforcing label-wise embeddings to maintain global consistency, LAC loss encourages the model to learn more distinctive representations, which is beneficial for recovering the information of potential positive labels. Experiments on multiple benchmark datasets validate that the proposed method can achieve state-of-the-art performance for solving SPML tasks.
Ming-Kun Xie, Jiahao Xiao, Sheng-Jun Huang
NeurIPS2