Jun-Yi Hang

dblp:299/4577 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0002-0345-8637ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Label-Specific Feature Learning for Multi-Label Classification: A Survey
abstract
In multi-label classification, each instance can be associated with multiple class labels simultaneously. However, each label is supposed to possess specific characteristics of its own and thus naturally shows distinct discriminative preferences on features. With consideration on this property, label-specific feature learning has emerged as a promising strategy for multi-label classification, which constructs features specific to each label to facilitate multi-label discrimination process. This article aims to provide a timely review on this emerging modeling strategy, focusing on main progress made during the last decade. Firstly, fundamentals on label-specific feature learning including formal definition and key challenges are provided. Then, six representative label-specific feature learning algorithms are scrutinized under a concise taxonomy with necessary discussions on algorithmic properties. Lastly, open research problems in label-specific feature learning are summarized to provide possible directions for future studies.
Jun-Yi Hang, Min-Ling Zhang
ACM Trans. Knowl. Discov. Data1
2025 Mitigating Local Cohesion and Global Sparseness in Graph Contrastive Learning with Fuzzy Boundaries
abstract
Graph contrastive learning (GCL) aims at narrowing positives while dispersing negatives, often causing a minority of samples with great similarities to gather as a small group. It results in two latent shortcomings in GCL: 1) local cohesion that a class cluster contains numerous independent small groups, and 2) global sparseness that these small groups (or isolated samples) dispersedly distribute among all clusters. These shortcomings make the learned distribution only focus on local similarities among partial samples, which hinders the ability to capture the ideal global structural properties among real clusters, especially high intra-cluster compactness and inter-cluster separateness. Considering this, we design a novel fuzzy boundary by extending the original cluster boundary with fuzzy set theory, which involves fuzzy boundary construction and fuzzy boundary contraction to address these shortcomings. The fuzzy boundary construction dilates the original boundaries to bridge the local groups, and the fuzzy boundary contraction forces the dispersed samples or groups within the fuzzy boundary to gather tightly, jointly mitigating local cohesion and global sparseness while forming the ideal global structural distribution. Extensive experiments demonstrate that a graph auto-encoder with the fuzzy boundary significantly outperforms current state-of-the-art GCL models in both downstream tasks and quantitative analysis.
Yuena Lin, Hai-Chun Cai, Jun-Yi Hang, Haobo Wang 0001, Zhen Yang 0004, Gengyu Lyu
ICML3
2025 Partial multi-label learning via label-specific feature corrections
Jun-Yi Hang, Min-Ling Zhang
Sci. China Inf. Sci.1
2025 Dual Perspective of Label-Specific Feature Learning for Multi-Label Classification
abstract
Label-specific features work as an effective supervised feature manipulation strategy to account for distinct discriminative properties of each class label in multi-label classification. Existing approaches implement this strategy in its primal form, i.e., finding the most pertinent features specific to each class label and directly inducing classifiers on these features. Instead of such a straightforward implementation, a dual perspective for label-specific feature learning is investigated in this article. As a dual problem of existing primal one, we consider label-specific discriminative properties by identifying non-informative features for each class label and making the discrimination process immutable to variations of identified features. Accordingly, a perturbation-based approach Dela is presented, which endows classifiers with immutability on simultaneously identified non-informative features by solving a probabilistically relaxed expected risk minimization problem. Furthermore, we touch the realistic issue of label-specific feature learning in a weakly supervised scenario via extending Dela to accommodate to multi-label data with missing labels. Comprehensive experiments show that our approach outperforms the state-of-the-art counterparts.
Jun-Yi Hang, Min-Ling Zhang
ACM Trans. Knowl. Discov. Data1
2024 Binary Decomposition: A Problem Transformation Perspective for Open-Set Semi-Supervised Learning
abstract
Semi-supervised learning (SSL) is a classical machine learning paradigm dealing with labeled and unlabeled data. However, it often suffers performance degradation in real-world open-set scenarios, where unlabeled data contains outliers from novel categories that do not appear in labeled data. Existing studies commonly tackle this challenging open-set SSL problem with detect-and-filter strategy, which attempts to purify unlabeled data by detecting and filtering outliers. In this paper, we propose a novel binary decomposition strategy, which refrains from error-prone procedure of outlier detection by directly transforming the original open-set SSL problem into a number of standard binary SSL problems. Accordingly, a concise yet effective approach named BDMatch is presented. BDMatch confronts two attendant issues brought by binary decomposition, i.e. class-imbalance and representation-compromise, with adaptive logit adjustment and label-specific feature learning respectively. Comprehensive experiments on diversified benchmarks clearly validate the superiority of BDMatch as well as the effectiveness of our binary decomposition strategy.
Jun-Yi Hang, Min-Ling Zhang
ICML1
2024 Learning Label-Specific Multiple Local Metrics for Multi-Label Classification
Junxiang Mao, Jun-Yi Hang, Min-Ling Zhang
IJCAI2
2024 Multi-Label Open Set Recognition
abstract
In multi-label learning, each training instance is associated with multiple labels simultaneously. Traditional multi-label learning studies primarily focus on closed set scenario, i.e. the class label set of test data is identical to those used in training phase. Nevertheless, in numerous real-world scenarios, the environment is open and dynamic where unknown labels may emerge gradually during testing. In this paper, the problem of multi-label open set recognition (MLOSR) is investigated, which poses significant challenges in classifying and recognizing instances with unknown labels in multi-label setting. To enable open set multi-label prediction, a novel approach named SLAN is proposed by leveraging sub-labeling information enriched by structural information in the feature space. Accordingly, unknown labels are recognized by differentiating the sub-labeling information from holistic supervision. Experimental results on various datasets validate the effectiveness of the proposed approach in dealing with the MLOSR problem.
Jun-Yi Hang, Min-Ling Zhang
NeurIPS2
2023 Can Label-Specific Features Help Partial-Label Learning?
abstract
Partial label learning (PLL) aims to learn from inexact data annotations where each training example is associated with a coarse candidate label set. Due to its practicability, many PLL algorithms have been proposed in recent literature. Most prior PLL works attempt to identify the ground-truth labels from candidate sets and the classifier is trained afterward by fitting the features of examples and their exact ground-truth labels. From a different perspective, we propose to enrich the feature space and raise the question ``Can label-specific features help PLL?'' rather than learning from examples with identical features for all classes. Despite its benefits, previous label-specific feature approaches rely on ground-truth labels to split positive and negative examples of each class and then conduct clustering analysis, which is not directly applicable in PLL. To remedy this problem, we propose an uncertainty-aware confidence region to accommodate false positive labels. We first employ graph-based label enhancement to yield smooth pseudo-labels and facilitate the confidence region split. After acquiring label-specific features, a family of binary classifiers is induced. Extensive experiments on both synthesized and real-world datasets are conducted and the results show that our method consistently outperforms eight baselines. Our code is released at https://github.com/meteoseeker/UCL
Ruo-Jing Dong, Jun-Yi Hang, Tong Wei 0001, Min-Ling Zhang
AAAI2
2023 Partial Multi-Label Learning with Probabilistic Graphical Disambiguation
abstract
In partial multi-label learning (PML), each training example is associated with a set of candidate labels, among which only some labels are valid. As a common strategy to tackle PML problem, disambiguation aims to recover the ground-truth labeling information from such inaccurate annotations. However, existing approaches mainly rely on heuristics or ad-hoc rules to disambiguate candidate labels, which may not be universal enough in complicated real-world scenarios. To provide a principled way for disambiguation, we make a first attempt to explore the probabilistic graphical model for PML problem, where a directed graph is tailored to infer latent ground-truth labeling information from the generative process of partial multi-label data. Under the framework of stochastic gradient variational Bayes, a unified variational lower bound is derived for this graphical model, which is further relaxed probabilistically so that the desired prediction model can be induced with simultaneously identified ground-truth labeling information. Comprehensive experiments on multiple synthetic and real-world data sets show that our approach outperforms the state-of-the-art counterparts.
Jun-Yi Hang, Min-Ling Zhang
NeurIPS1
2023 Learning label-specific features for decomposition-based multi-class classification
Bin-Bin Jia 0001, Jun-Ying Liu 0001, Jun-Yi Hang, Min-Ling Zhang
Frontiers Comput. Sci.3
2022 End-to-End Probabilistic Label-Specific Feature Learning for Multi-Label Classification
abstract
Label-specific features serve as an effective strategy to learn from multi-label data with tailored features accounting for the distinct discriminative properties of each class label. Existing prototype-based label-specific feature transformation approaches work in a three-stage framework, where prototype acquisition, label-specific feature generation and classification model induction are performed independently. Intuitively, this separate framework is suboptimal due to its decoupling nature. In this paper, we make a first attempt towards a unified framework for prototype-based label-specific feature transformation, where the prototypes and the label-specific features are directly optimized for classification. To instantiate it, we propose modelling the prototypes probabilistically by the normalizing flows, which possess adaptive prototypical complexity to fully capture the underlying properties of each class label and allow for scalable stochastic optimization. Then, a label correlation regularized probabilistic latent metric space is constructed via jointly learning the prototypes and the metric-based label-specific features for classification. Comprehensive experiments on 14 benchmark data sets show that our approach outperforms the state-of-the-art counterparts.
Jun-Yi Hang, Min-Ling Zhang, Yang-He Feng, Xiaocheng Song
AAAI1
2022 Dual Perspective of Label-Specific Feature Learning for Multi-Label Classification
abstract
Label-specific features serve as an effective strategy to facilitate multi-label classification, which account for the distinct discriminative properties of each class label via tailoring its own features. Existing approaches implement this strategy in a quite straightforward way, i.e. finding the most pertinent and discriminative features for each class label and directly inducing classifiers on constructed label-specific features. In this paper, we propose a dual perspective for label-specific feature learning, where label-specific discriminative properties are considered by identifying each label’s own non-informative features and making the discrimination process immutable to variations of these features. To instantiate it, we present a perturbation-based approach DELA to provide classifiers with label-specific immutability on simultaneously identified non-informative features, which is optimized towards a probabilistically-relaxed expected risk minimization problem. Comprehensive experiments on 10 benchmark data sets show that our approach outperforms the state-of-the-art counterparts.
Jun-Yi Hang, Min-Ling Zhang
ICML1
2022 Submodular Feature Selection for Partial Label Learning
abstract
Partial label learning induces a multi-class classifier from training examples each associated with a candidate label set where the ground-truth label is concealed. Feature selection improves the generalization ability of learning system via selecting essential features for classification from the original feature set, while the task of partial label feature selection is challenging due to ambiguous labeling information. In this paper, the first attempt towards partial label feature selection is investigated via mutual-information-based dependency maximization. Specifically, the proposed approach SAUTE iteratively maximizes the dependency between selected features and labeling information, where the value of mutual information is estimated from confidence-based latent variable inference. In each iteration, the near-optimal features are selected greedily according to properties of submodular mutual information function, while the density of latent label variable is inferred with the help of updated labeling confidences over candidate labels by resorting to kNN aggregation in the induced lower-dimensional feature space. Extensive experiments over synthetic as well as real-world partial label data sets show that the generalization ability of well-established partial label learning algorithms can be significantly improved after coupling with the proposed feature selection approach.
Wei-Xuan Bao, Jun-Yi Hang, Min-Ling Zhang
KDD2
2022 Collaborative Learning of Label Semantics and Deep Label-Specific Features for Multi-Label Classification
abstract
In multi-label classification, the strategy of label-specific features has been shown to be effective to learn from multi-label examples by accounting for the distinct discriminative properties of each class label. However, most existing approaches exploit the semantic relations among labels as immutable prior knowledge, which may not be appropriate to constrain the learning process of label-specific features. In this paper, we propose to learn label semantics and label-specific features in a collaborative way. Accordingly, a deep neural network (DNN) based approach named CLIF, i.e. Collaborative Learning of label semantIcs and deep label-specific Features for multi-label classification, is proposed. By integrating a graph autoencoder for encoding semantic relations in the label space and a tailored feature-disentangling module for extracting label-specific features, CLIF is able to employ the learned label semantics to guide mining label-specific features and propagate label-specific discriminative properties to the learning process of the label semantics. In such a way, the learning of label semantics and label-specific features interact and facilitate with each other so that label semantics can provide more accurate guidance to label-specific feature learning. Comprehensive experiments on 14 benchmark data sets show that our approach outperforms other well-established multi-label classification algorithms.
Jun-Yi Hang, Min-Ling Zhang
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Partial Label Dimensionality Reduction via Confidence-Based Dependence Maximization
abstract
Partial label learning deals with training examples each associated with a set of candidate labels, among which only one is valid. Most existing works focus on manipulating the label space by estimating the labeling confidences of candidate labels, while the task of manipulating the feature space by dimensionality reduction has been rarely investigated. In this paper, a novel partial label dimensionality reduction approach named CENDA is proposed via confidence-based dependence maximization. Specifically, CENDA adapts the Hilbert-Schmidt Independence Criterion (HSIC) to help identify the projection matrix, where the dependence between projected feature information and confidence-based labeling information is maximized iteratively. In each iteration, the projection matrix admits closed-form solution by solving a tailored generalized eigenvalue problem, while the labeling confidences of candidate labels are updated by conducting kNN aggregation in the projected feature space. Extensive experiments over a broad range of benchmark data sets show that the predictive performance of well-established partial label learning algorithms can be significantly improved by coupling with the proposed dimensionality reduction approach.
Wei-Xuan Bao, Jun-Yi Hang, Min-Ling Zhang
KDD2